<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Inequality | Macro Paper Warehouse</title><link>https://macropaperwarehouse.com/topics/inequality/</link><atom:link href="https://macropaperwarehouse.com/topics/inequality/index.xml" rel="self" type="application/rss+xml"/><description>Inequality</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 01 Jan 2026 00:00:00 +0000</lastBuildDate><item><title>A Macro Study of the Unequal Effects of Climate Change</title><link>https://macropaperwarehouse.com/papers/a-macro-study-of-the-unequal-effects-of-climate-change/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/a-macro-study-of-the-unequal-effects-of-climate-change/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper develops a macro heterogeneous-agent model to quantify the distributional welfare impacts of higher temperatures from climate change across income groups in the United States. The motivation is that existing macro climate-economy models either abstract from heterogeneity entirely or focus on spatial heterogeneity across regions rather than income heterogeneity within regions. The paper fills this gap by modeling how the welfare consequences of temperature change depend on both the region a household lives in and its position in the income distribution.&lt;/p&gt;
&lt;p&gt;The model is calibrated to the US using five data sources: NIPA accounts from the BEA (averaged 1997–2020), the 2015 Residential Energy and Consumption Survey (RECS), PRISM climate data (1950–2022), a proprietary product-level data set of over 1,000 heaters, air conditioners, and heat pumps scraped from ecomfort.com in fall 2023, and county-level climate projections for year 2100 under RCP 8.5 from Rasmussen et al. (2016). The US is divided into five regions (cold, cool, mild, warm, and hot) of approximately equal population based on average county temperature. The quantitative exercise compares two stationary equilibria: a contemporary equilibrium using the current temperature distribution and a climate-change equilibrium using the projected 2100 distribution under RCP 8.5 (a no-large-scale-climate-policy scenario). Welfare is measured using the consumption-housing equivalent variation (CHEV), defined as the percent increase in consumption and housing a household would require in every period in the contemporary equilibrium to be indifferent between the two equilibria.&lt;/p&gt;
&lt;p&gt;Households adapt to temperature through two channels: an intensive margin (adjusting energy use for heating and cooling given existing equipment) and an extensive margin (deciding whether to purchase a heater, air conditioner, or heat pump, each carrying a fixed cost). The production functions for heating and cooling are estimated by OLS on the product-level data set, yielding equipment exponents of 0.35 (air conditioners), 0.28 (heaters), and 0.27 (heat pumps), and energy exponents of 0.77, 0.86, and 0.85, respectively, with R-squared values of 0.97, 0.79, and 1.00. A key analytical insight from a stylized model is that the outdoor temperature acts as a &amp;ldquo;transfer from nature&amp;rdquo; to households — warmer days in cold weather and cooler days in hot weather reduce the energy households must purchase, augmenting real income. Because this transfer is a larger share of income for lower-income households, its changes are distributionally regressive when the transfer falls (hotter regions warming further) and progressive when it rises (colder regions warming).&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. Among middle- and high-income households, climate change generates progressive welfare gains in colder regions — ranging from +0.71 percent of consumption-and-housing for households in the third income decile in the cool region to near-zero for the highest income households — and regressive welfare losses in hotter regions, ranging from −1.85 percent for third-decile households in the warm region to near-zero for high-income households. These patterns are driven by the intensive margin (changes in transfers from nature). For low-income households, the pattern reverses: low-income households in colder regions suffer welfare losses (the dominant effect is that climate change forces them to purchase their first air conditioner), while some low-income households in hotter regions experience welfare gains (they can forgo purchasing a heater). Climate change raises the Gini coefficient on lifetime welfare by 1.02, 1.01, and 0.50 percent in the cold, cool, and mild regions, and reduces it by 0.09 and 0.21 percent in the warm and hot regions. Aggregate welfare effects from the heterogeneous-agent model substantially exceed what a representative-agent model would imply: for example, in the mild region, climate change reduces aggregate welfare by 0.65 percent in the baseline but only 0.17 percent in the representative-agent version.&lt;/p&gt;
&lt;p&gt;Policy experiments reveal: (1) Fully offsetting the welfare costs of climate change for the lowest-income households would require government spending on energy assistance to more than double (a factor of 2.2 increase), with the largest increases concentrated in colder regions. (2) A universal heat-pump mandate eliminates the extensive-margin channel, producing monotonically progressive welfare gains in colder regions and monotonically regressive welfare losses in hotter regions across all income deciles. (3) Heat-pump cost parity with heaters largely increases adoption and moderates welfare costs, but low-income households in the hot region see limited improvement because they still prefer air conditioners. (4) Accounting for temperature effects on the labor productivity of outdoor workers (roughly 8 percent of the workforce, concentrated at lower incomes) amplifies welfare costs in hotter regions and moderates them in colder regions, with magnitudes tied to the share of workers affected.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper is a calibrated structural model rather than an empirical identification exercise. Identification in the sense of parameter estimation comes from two sources: (1) OLS estimation of heating and cooling production functions on cross-sectional product-level data, where manufacturers measure capacity and efficiency under standardized conditions, limiting TFP endogeneity concerns that plague aggregate production function estimation; and (2) internal calibration of remaining parameters to match a set of moments from RECS 2015 and NIPA. Threats to the structural analysis include the assumption that households treat housing and equipment as flow (rental) choices rather than durable stocks, abstracting from switching costs and adjustment costs over the transition — the paper explicitly notes this limits the analysis to long-run stationary equilibria. The small-open-economy assumption for capital removes domestic capital-market clearing as a constraint. The calibration uses 2015 RECS (not 2020) to avoid COVID-19 distortions to cooling budget shares. The paper abstracts from amenity values of outdoor temperature, mortality from temperature exposure (approximately 0.04 percent of US deaths from 1999–2020), and spatial migration responses.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-core-mechanisms-and-how-are-they-distinguished"&gt;Q2. What are the two core mechanisms and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;The two mechanisms are the intensive margin (how much energy to use given existing equipment) and the extensive margin (whether to purchase heating or cooling equipment at all). The paper distinguishes them analytically using the simple model, which isolates the intensive margin by assuming all households have equipment. The intuition from the simple model — outdoor temperature as a transfer from nature — explains why welfare effects are progressive in regions where climate change makes temperatures more moderate (transfers rise) and regressive where temperatures become more extreme (transfers fall). The extensive margin is then added in the quantitative model through fixed costs of heater, air conditioner, and heat pump equipment. The paper shows that climate change affects specialization favorability (the degree to which a temperature distribution favors concentrating on only heating or only cooling equipment), and that this extensive-margin channel is most important for lower-income households who are near a corner solution of specializing in only one type of equipment. The heat-pump-mandate counterfactual is used to isolate the intensive-margin channel: when all households use heat pumps in both equilibria, the extensive-margin decision is unchanged by climate change, and all welfare effects are driven purely by transfers from nature.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-across-income-groups-and-regions"&gt;Q3. What heterogeneity is documented across income groups and regions?&lt;/h3&gt;
&lt;p&gt;Welfare effects vary dramatically in both sign and magnitude. Among middle- and high-income households, climate change generates progressive welfare gains in colder regions (e.g., +0.71 percent CHEV for third-decile households in the cool region, falling toward zero at the top) and regressive welfare losses in hotter regions (e.g., −1.85 percent CHEV for third-decile households in the warm region, again near-zero at the top). For low-income households, the pattern reverses: they experience welfare losses in colder regions (forced to buy first air conditioner) and welfare gains or smaller losses in hotter regions (can forgo purchasing a heater). Figure 2 in the paper shows these crossing patterns by income decile for all five regions simultaneously. The Gini coefficient changes by +1.02% (cold), +1.01% (cool), +0.50% (mild), −0.09% (warm), and −0.21% (hot). Migration incentives also differ: high-income households gain incentives to move to cooler regions (driven by transfers from nature), while low-income households gain incentives to move to warmer regions (driven by specialization changes).&lt;/p&gt;
&lt;h3 id="q4-what-is-the-transfers-from-nature-concept-and-why-does-it-produce-differential-welfare-effects"&gt;Q4. What is the &amp;rsquo;transfers from nature&amp;rsquo; concept and why does it produce differential welfare effects?&lt;/h3&gt;
&lt;p&gt;The paper formalizes the idea that outdoor temperature provides free heating or cooling that substitutes for costly purchased energy. On a cold day with outdoor temperature ζ, nature provides ζ degrees of heating for free, effectively augmenting household income by p_eh * ζ (the value of that heating at market prices). This transfer is identical in absolute terms for all households regardless of income, but it is a larger fraction of income for low-income households, so its loss or gain has greater proportional welfare impact on them. This parallels the progressivity of lump-sum transfers in public finance: losing a dollar matters more when income is lower. Consequently, when climate change moves a region to more moderate temperatures (colder regions), the resulting increase in transfers from nature is progressive — lower-income households gain proportionally more. When climate change moves a region to more extreme temperatures (hotter regions), the decrease in transfers is regressive — lower-income households lose proportionally more. The amenity value of outdoor temperature (distinct from the heating/cooling transfer) is abstracted from in the quantitative model on the grounds that, per the simple model, it does not affect the cross-income distribution of welfare changes if preferences over amenities are uncorrelated with income.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-extensive-margin-generate-the-reversal-of-welfare-effects-for-low-income-households"&gt;Q5. How does the extensive margin generate the reversal of welfare effects for low-income households?&lt;/h3&gt;
&lt;p&gt;The extensive margin works through what the paper calls &amp;lsquo;specialization favorability.&amp;rsquo; When a temperature distribution is dominated by cold days, households can optimally purchase only heater equipment, avoiding the additional fixed cost of an air conditioner; the reverse holds in hot climates. Climate change reduces the specialization favorability index in colder regions by adding more hot days, and increases it in hotter regions by reducing cold days. The welfare impact of moving between a corner solution (one type of equipment) and an interior solution (two types of equipment, or a heat pump) tends to be larger than moving between two interior solutions. In the cold region, climate change causes the majority of households in the bottom three income deciles to transition from not having air conditioning to having it (Figure 5, left panel). The fixed cost of buying an air conditioner for the first time exceeds the intensive-margin gains from more moderate temperatures, producing net welfare losses. In the hot region, many second-through-fourth decile households move from having heat in the contemporary equilibrium to not having heat in the climate-change equilibrium (Figure 5, right panel), saving the fixed cost and producing net welfare gains despite more extreme temperatures.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-model-calibrated-and-what-is-the-quality-of-fit"&gt;Q6. How is the model calibrated and what is the quality of fit?&lt;/h3&gt;
&lt;p&gt;Externally calibrated parameters include: capital income share α = 0.26 (Kiyotaki et al., 2011), depreciation rate δ = 0.066, interest rate r* = 0.04, CRRA coefficient σ = 2, bliss point temperature ζ* = 18°C, labor productivity process (ρ = 0.97, σ²_ε = 0.02, σ²_ξ = 0.66 from Kaplan, 2012), and production function exponents estimated from the ecomfort.com data. Internally calibrated parameters are jointly chosen to match: wealth-to-output ratio (3.0), housing-to-non-housing capital ratio (0.88), average heating budget share for non-heat-pump households (0.014), average cooling budget share (0.0055), energy budget share for heat-pump households (0.014), fractions of households with heating (0.95), cooling (0.86), and heat pumps (0.09), the ratio of energy budget shares between the fifth and first income quintile (0.12), the ratio of energy expenditures between high and low income (1.72), and energy assistance as a fraction of energy expenditures (0.83). Table 3 shows the model matches all targeted moments closely. External validation (untargeted moments) shows the model also replicates the associations between heating/cooling degree days and budget shares, equipment ownership, and indoor temperature choices, with similar signs and magnitudes to RECS 2015 data. One limitation is that the model overstates heat pump adoption (17% in model vs. 9% in 2015 RECS, though 14% in 2020 RECS), because it treats modern cold-weather-capable heat pumps as the default.&lt;/p&gt;
&lt;h3 id="q7-what-do-the-policy-counterfactuals-show"&gt;Q7. What do the policy counterfactuals show?&lt;/h3&gt;
&lt;p&gt;Four policy experiments are analyzed. First, scaling energy assistance proportionally to energy needs under climate change reduces assistance by 24% in cold and 20% in cool regions (where transfers from nature increase) and raises it by 9%, 36%, and 79% in mild, warm, and hot regions. Government spending increases by 25%, but the program remains smaller than 0.02% of output. This scaling partially offsets but does not eliminate the distributional distortions. Fully eliminating welfare costs for the lowest-income households would require multiplying energy assistance spending by a factor of 2.2. Second, a universal heat-pump mandate (analogous to natural gas bans like New York, Washington DC, or California&amp;rsquo;s post-2030 ban on natural gas furnaces) eliminates all extensive-margin effects because all households hold heat pumps in both equilibria. Under this mandate, climate change produces monotonically progressive welfare gains across all income groups in colder regions and monotonically regressive welfare costs in hotter regions. Third, heat-pump cost parity with heaters drives near-universal heat pump adoption and broadly moderates welfare costs relative to baseline, but the lowest-income households in the hot region see limited improvement because they still prefer air conditioners over heat pumps even at cost parity (air conditioners are cheaper and heat pumps&amp;rsquo; heating advantage is less valuable in an already-hot, increasingly-hotter climate). Fourth, the labor productivity extension (using the Richardson construction cost database adjustment factor of 1% per degree outside 40°F–85°F) implies that climate change raises low-income productivity by 2% in cold and 0.9% in cool regions and reduces it by 0.1%, 1.1%, and 2.2% in mild, warm, and hot regions. These labor-productivity changes modestly moderate welfare costs in colder regions and amplify them in hotter regions for low-income households.&lt;/p&gt;
&lt;h3 id="q8-why-does-income-heterogeneity-matter-for-aggregate-welfare-calculations"&gt;Q8. Why does income heterogeneity matter for aggregate welfare calculations?&lt;/h3&gt;
&lt;p&gt;The paper demonstrates that a representative-agent model substantially underestimates the aggregate welfare cost of climate change in all regions except the hot region. In the cold region, the aggregate CHEV is −1.03% in the baseline but the average (seventh-decile) household experiences small positive welfare effects (+0.19%), and the representative-agent model yields −0.00%. In the mild region, the aggregate is −0.65% but the representative-agent model gives −0.17%. The discrepancy arises because the welfare distribution is skewed: large losses for low-income households in colder regions are not offset by small or negative gains for high-income households, so the average is dominated by the tails. In the hot region the direction reverses: the baseline aggregate benefit (+0.24%) is driven by large gains at the bottom that the representative-agent model (−0.43%) misses entirely. This finding parallels the broader macroeconomics literature showing that income heterogeneity affects the aggregate welfare cost of business cycles, inflation, and asset pricing.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q9. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of two literatures. The macro climate-economy literature (Acemoglu et al., 2012; Golosov et al., 2014; Barrage, 2020) typically uses representative-agent models that abstract from heterogeneity. The spatial heterogeneity literature (Cruz and Rossi-Hansberg, 2024; Bilal and Rossi-Hansberg, 2023; Rudik et al., 2022) studies how welfare consequences vary across regions based on their income levels and exposures but not within-region income differences. The within-region inequality literature (Dennig et al., 2015; Kornek et al., 2021; Belfori and Macera, 2022; Douenne et al., 2023) adds heterogeneous fixed income types to integrated assessment models, but does not model endogenous income and wealth distributions. Blanz (2023) is the closest precursor: it uses a standard incomplete-markets model to study food-price effects of climate change in developing countries, but does not model the temperature-equipment-energy production technology. The empirical literature (Hsiang et al., 2017; Park et al., 2018; Doremus et al., 2022) estimates reduced-form relationships between temperature and energy spending by income group, but cannot decompose intensive vs. extensive margin mechanisms or conduct structural policy counterfactuals. The key novel contributions are: (1) endogenous income and wealth heterogeneity within the Bewley-Huggett-Aiyagari tradition, (2) explicit modeling of both margins of temperature adaptation with estimated production functions, and (3) the ability to separately identify the roles of transfers from nature and specialization favorability.&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-are-conducted"&gt;Q10. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The paper reports several robustness checks. First, the main calibration uses the housing exponent γ = 0.1, but Appendix Figure D.1 shows results with γ = 0.4 (the upper bound implied by the RECS regression of energy on square footage, before controlling for quality), finding broadly similar qualitative results. Second, the 2015 RECS is used instead of the 2020 RECS due to COVID-19 distortions to cooling budget shares; the paper notes heating budget shares are similar between the two surveys while cooling shares are materially higher in 2020. Third, external validation of the model on untargeted moments (associations between HDD/CDD and heating/cooling budget shares, equipment ownership, and indoor temperatures) confirms the model&amp;rsquo;s predictive validity. Fourth, the welfare results are computed for both the main five-region model and a representative-agent version, documenting the magnitude of the aggregation bias. Fifth, the labor productivity extension bounds the relevant population (bottom 3% vs. bottom 16% of workers) to bracket the Occupational Requirements Survey estimate of 8% of workers constantly or frequently exposed outdoors.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-scope-conditions-and-limitations-of-the-main-results"&gt;Q11. What are the scope conditions and limitations of the main results?&lt;/h3&gt;
&lt;p&gt;Several important scope conditions apply. The analysis focuses exclusively on the direct effects of higher temperatures in the US; it does not cover other forms of climate damage (sea level rise, storm frequency, drought, wildfire) or effects in other countries. The model is solved for stationary equilibria, so it cannot speak to transition dynamics or the welfare costs of adjustment during the period when households are switching equipment. Housing and equipment are modeled as flow (rental) choices, abstracting from switching costs, adjustment frictions, and the interaction between homeownership and equipment decisions. The model abstracts from the amenity value of outdoor temperature (e.g., preference for pleasant weather), temperature-related mortality (about 0.04% of US deaths, 1999–2020, heavily concentrated among the unhoused population outside the model), and behavioral adaptation beyond energy and equipment choices (migration is analyzed only as a partial equilibrium incentive calculation, not as an equilibrium outcome). The capital market operates as a small open economy, so general equilibrium effects on interest rates are absent. Labor productivity effects of temperature are only explored for low-income workers in the outdoor sector, not for higher-income or indoor workers.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-migration-findings-and-their-caveats"&gt;Q12. What are the migration findings and their caveats?&lt;/h3&gt;
&lt;p&gt;The paper shows that climate change increases incentives for high-income households to migrate to cooler regions (driven by the transfers-from-nature channel — cooler regions offer larger increases in transfers) and increases incentives for low-income households to migrate to warmer regions (driven by the specialization channel — warmer regions allow forgoing heater equipment). The magnitude of the change in migratory pressure for high-income households is much smaller (order of magnitude roughly 0.15 on the paper&amp;rsquo;s scale) than for low-income households (order of magnitude roughly 3 on the same scale). The authors explicitly caveat that this is a partial equilibrium exercise: the model abstracts from the amenity value of temperature (which would reduce pressure to move to warmer regions by reducing the attractiveness of hot destinations) and from other dimensions of climate change (storm risk, fire risk) that would affect migration incentives independently.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Transfers from nature&lt;/strong&gt;: In this paper&amp;rsquo;s framework, outdoor temperature acts as a subsidy equivalent to income: on a cold day, nature provides degrees of heating for free, augmenting household real income by the value of that heating energy; on a hot day, it provides degrees of cooling. The transfer is the same in absolute terms for all households but represents a larger fraction of income for lower-income households, making changes in temperature distributionally progressive (when transfers rise) or regressive (when transfers fall).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extensive margin of temperature adaptation&lt;/strong&gt;: The binary decision of whether to purchase temperature-control equipment — a heater, air conditioner, or heat pump — each carrying a fixed cost. Households at the extensive margin may optimally forego one type of equipment entirely (complete specialization), and climate change can force them to acquire equipment they previously lacked or allow them to drop equipment they previously held.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive margin of temperature adaptation&lt;/strong&gt;: The continuous decision of how much energy to purchase to operate existing heating and cooling equipment in order to achieve a desired indoor temperature, conditional on having that equipment. Changes in the outdoor temperature distribution affect energy expenditures along this margin for all households that already own equipment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Specialization favorability index&lt;/strong&gt;: A region-level index S_n ∈ [0,1] defined as the absolute difference between total degrees of heating need and total degrees of cooling need, divided by their sum. Higher values indicate that the temperature distribution is more dominated by either heating or cooling demand, making it more efficient for households to specialize in a single type of temperature-control equipment rather than purchasing both. Climate change reduces specialization favorability in colder regions and increases it in hotter regions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption-housing equivalent variation (CHEV)&lt;/strong&gt;: The paper&amp;rsquo;s welfare metric: the percentage by which a household&amp;rsquo;s consumption and housing would need to increase in every period of the contemporary equilibrium for the household to be indifferent between remaining in the contemporary equilibrium and living in the climate-change equilibrium. Negative CHEV values indicate welfare losses from climate change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Temperature damage function D(T)&lt;/strong&gt;: A function mapping the deviation of indoor temperature from the bliss point to the fraction of full utility the household receives from housing services. D equals 1 when indoor temperature equals the bliss point (18°C in calibration) and falls below 1 as indoor temperature deviates in either direction, with the rate of decline governed by parameter χ. This function creates the motive to use energy for heating and cooling.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;RCP 8.5&lt;/strong&gt;: As used in this paper, a climate scenario from the CMIP archive representing emissions in the absence of large-scale climate policy, used to construct the 2100 temperature distribution in the climate-change equilibrium. County-level projections come from Rasmussen et al. (2016), probability-weighted across climate models.&lt;/p&gt;</description></item><item><title>A Tractable Income Process for Business Cycle Analysis</title><link>https://macropaperwarehouse.com/papers/a-tractable-income-process-for-business-cycle-analysis/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/a-tractable-income-process-for-business-cycle-analysis/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Guvenen, McKay, and Ryan estimate a stochastic income process for US male workers that simultaneously matches five empirical regularities from Social Security Administration administrative panel data covering 1978–2011: (i) flat and acyclical variance of income growth rates, (ii) volatile and procyclical Kelley skewness, (iii) very high kurtosis — targeted at 20 for one-year changes and 12 for five-year changes — (iv) a near-linear rise in cross-sectional log-income variance from age 25 to 55, and (v) a systematic factor structure in business cycle incidence whereby income losses during recessions are predictably related to a worker&amp;rsquo;s pre-recession income rank. All five facts are drawn from Guvenen et al. (2014) and Guvenen et al. (2021), which document them from SSA records on individual income histories.\n\nThe income process adds three key departures to the workhorse persistent-plus-transitory Gaussian specification. First, transitory &amp;ldquo;nonemployment&amp;rdquo; shocks — arriving annually with approximately 45% probability and drawn from an exponential distribution — create fat tails through their arrival (large income losses) and departure (large income gains), and leave a persistent &amp;ldquo;scarring&amp;rdquo; residue through a passthrough parameter ψ estimated at 9.4% in the baseline nonemployment model. Each year, roughly 8.6% of workers experience income declines of 50% or more from the nonemployment shock alone, and 1.8% fall to effectively zero income. The scarring mechanism makes the left tail of the income growth density fatter than the right tail, consistent with the data (left-tail log-density slope 1.4, right-tail slope –2.2). Second, innovations to the persistent AR(1) component are drawn from a time-varying three-component normal mixture — with the dominant central component realized with about 83% probability and near-zero standard deviation (~1%), flanked by left-tail and right-tail components with probabilities of ~10.9% and ~6.2% and standard deviations of ~16.4% and ~19.2% — whose means shift with contemporaneous aggregate wage income growth (xt = β·Δwt). This mean-shifting mechanism generates procyclical skewness under an acyclical variance, because it redistributes probability mass between the tails without altering mixture probabilities or component variances. Third, a piecewise-linear factor structure makes each individual&amp;rsquo;s income sensitivity to aggregate fluctuations depend on the persistent component of income (γi + zi,t), with a kink separating two slope regimes. In the Great Recession, workers at the 10th percentile of pre-recession income lost approximately 18 percentage points more than workers at the 90th percentile; both the bottom and top deciles were more exposed than the middle of the distribution, producing a V-shaped incidence pattern.\n\nEstimation uses simulated method of moments (SMM) with 360,000 simulated individuals per year, a 1947 burn-in start, and optimization via the TikTak global algorithm. Six models of increasing complexity are estimated, each requiring only one individual state variable (the persistent component z) — matching the parsimony of the standard model. The workhorse Gaussian model (Model 1) understates the variance of one-year log income changes by 60–80%; introducing nonemployment shocks (Model 2) largely resolves this, matching one-year variance exactly and narrowing the five-year shortfall to 30%. Adding the time-varying normal mixture (Model 3) generates procyclical skewness and acyclical variance. Adding the factor structure (Model 4) captures differential recession exposure. Models 5 and 6 introduce Heterogeneous Income Profiles (HIP, σκ = 0.015) and estimate AR(1) persistence freely, obtaining ρ ≈ 0.80, which better captures the right tail of the income growth distribution.\n\nThe paper recommends Model 5 as a general-purpose benchmark (without the factor structure), Model 4 when differential business cycle incidence is central, and Model 3 when maximum parsimony is needed. The richer income dynamics documented here have direct implications for quantifying the welfare cost of business cycles, the value of social insurance, the design of automatic stabilizers, the distribution of marginal propensities to consume, and asset pricing under heterogeneous agents.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-estimation-procedure-and-what-data-does-it-use"&gt;Q1. What is the estimation procedure and what data does it use?&lt;/h3&gt;
&lt;p&gt;The paper uses simulated method of moments (SMM), targeting approximately 120+ moments derived from Social Security Administration administrative panel data on individual income histories of US male workers over 1978–2011 (from Guvenen et al. 2014 and 2021). The simulation panel contains 360,000 individuals per year, initialized in 1947 with a burn-in period. Optimization uses the TikTak global algorithm (Arnoud et al., 2019). Moments targeted include the 10th, 50th, and 90th percentiles of one-, three-, and five-year income growth averaged across 1979–2011 (nine moments); kurtosis at one-year and five-year horizons (two moments); cross-sectional variance of log income at ages 25, 35, 45, and 55 (four moments); left- and right-tail mass and log-density slopes from the 1995–1996 income growth distribution (four moments); the full time series of Kelley skewness for one-, three-, and five-year changes (93 moments); and piecewise-linear slopes of the factor structure for seven business cycle episodes — four recessions and three expansions covering 1979–2010 (14 moments). Moments are weighted approximately equally, with skewness moments down-weighted collectively.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-key-departures-from-the-workhorse-gaussian-model-and-what-feature-does-each-address"&gt;Q2. What are the three key departures from the workhorse Gaussian model and what feature does each address?&lt;/h3&gt;
&lt;p&gt;First, transitory &amp;rsquo;nonemployment&amp;rsquo; shocks drawn from an exponential distribution, arriving with ~45% annual probability, along with a scarring parameter ψ that loads a fraction of the transitory shock onto the persistent state — this generates the high kurtosis, thick tails, and asymmetry (steeper right than left tail) of the income growth distribution. Second, a three-component time-varying normal mixture for persistent innovations — the component means shift with the aggregate wage component xt = β·Δwt — producing procyclical skewness and acyclical variance simultaneously. Third, a piecewise-linear factor structure f(γi + zi,t) mediating each individual&amp;rsquo;s exposure to aggregate fluctuations, capturing the V-shaped relationship between pre-recession income rank and recession income loss.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-scarring-mechanism-and-how-large-is-it-empirically"&gt;Q3. What is the scarring mechanism and how large is it empirically?&lt;/h3&gt;
&lt;p&gt;Transitory nonemployment shocks ζi,t are assigned with probability (1 − pζ) each year and drawn from an exponential distribution with parameter λ, where ℓi,t ∈ [0,1] represents the income fraction lost. A fraction ψ of this transitory shock flows permanently into the persistent state zi,t via ˜ηi,t = ηi,t + ψζi,t. In Model 2, the annual probability of receiving a nonemployment shock is 45% (pζ ≈ 0.55), λ = 3.357 (mean income loss fraction ≈ 0.30), and ψ = 9.4%. Each year, 8.6% of workers experience income declines of 50% or more from the nonemployment shock alone, and 1.8% effectively lose all income (full-year nonemployment). The scarring makes the right tail steeper than the left tail in the income growth distribution, as re-employed workers do not return to their pre-shock income level.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-time-varying-normal-mixture-generate-procyclical-skewness-without-changing-variance"&gt;Q4. How does the time-varying normal mixture generate procyclical skewness without changing variance?&lt;/h3&gt;
&lt;p&gt;The three normal mixture components for the persistent innovation η are: a central component (probability ~83%, standard deviation ~1%), a left-tail component (~10.9%, ~16.4% sd), and a right-tail component (~6.2%, ~19.2% sd). Their means shift via the latent variable xt = β·Δwt: the central and left-tail means move with xt while the right-tail mean does not. A normalization ensures xt has zero mean-income effect. In recessions (xt &amp;lt; 0, Δwt &amp;lt; 0), the left-tail component&amp;rsquo;s mean shifts down and the right-tail component&amp;rsquo;s mean shifts up relative to the central, generating more left-skewed draws without changing the probabilities or variances of the components — hence acyclical variance and procyclical skewness. Alternative designs (cyclical mixture probabilities or variances) did not generate both patterns simultaneously.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-factor-structure-and-how-non-monotonic-is-it"&gt;Q5. What is the factor structure and how non-monotonic is it?&lt;/h3&gt;
&lt;p&gt;In deep recessions the factor structure is broadly monotone decreasing over the bulk of the distribution (lower-income workers lose more), with the 10th percentile losing about 18 percentage points more than the 90th percentile in the Great Recession (2007–2010). However, the pattern reverses for the top 10% of the income distribution: high earners also face large losses in financial-market-driven recessions, producing a V-shape. The piecewise-linear model f(q) with a kink at q-bar and slopes α1 (below) and α2 (above) captures this. The model fits the Great Recession V-shape and the mild 1990–1992 and 2000–2002 recessions (where the pattern is flatter, consistent with smaller drops in wt), but struggles to fit the large top-income losses in 2000–2002 without an additional stock-market-correlated factor.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-levels-vs-differences-puzzle-and-how-is-it-resolved"&gt;Q6. What is the levels-vs-differences puzzle and how is it resolved?&lt;/h3&gt;
&lt;p&gt;The canonical persistent-plus-transitory Gaussian model (Model 1) faces a fundamental tension: it can fit the cross-sectional variance of log income levels at each age, but it then understates the variance of one-year and five-year log income changes by 60–80% (squared standard deviations from Figures 8a and 9a). This tension was documented by Heathcote, Perri, and Violante (2010). Introducing the nonemployment shocks in Model 2 largely resolves it: the one-year variance of log income changes is matched exactly, and the five-year understatement narrows to about 30%. The nonemployment shock contributes high-frequency variance in income changes without requiring a comparably large increase in the variance of the persistent state, because it is mostly transitory.&lt;/p&gt;
&lt;h3 id="q7-what-role-does-hip-play-and-what-tensions-does-it-create"&gt;Q7. What role does HIP play and what tensions does it create?&lt;/h3&gt;
&lt;p&gt;Heterogeneous Income Profiles (HIP, σκ = 0.015 from Baker 1997 and Guvenen et al. 2021) allow AR(1) persistence ρ to be estimated freely rather than restricted to 1. The estimated ρ falls to 0.80 in Models 5 and 6. HIP provides a convex component to the lifecycle variance profile (from dispersion in individual growth-rate slopes κi) that offsets the concave contribution of mean-reverting persistent shocks, maintaining a near-linear age-variance profile at ρ &amp;lt; 1. Lower persistence better fits the right tail of annual income growth and the standard deviation of five-year changes. However, in Model 6 HIP worsens the fit to the factor structure, because mean reversion at ρ &amp;lt; 1 already generates faster income growth for low-income workers in expansions, reducing the work the factor structure needs to do in booms while resisting the factor structure&amp;rsquo;s ability to generate large losses for low-income workers in recessions.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-and-alternative-specifications-are-estimated"&gt;Q8. What robustness checks and alternative specifications are estimated?&lt;/h3&gt;
&lt;p&gt;The paper estimates two supplementary models reported in Appendix B. Model 2&amp;rsquo; removes the scarring component (ψ ≡ 0) from Model 2, finding a worse fit particularly in the histogram, kurtosis, and lifecycle inequality moments. Model 3&amp;rsquo; replaces the time-varying mixture with a static normal mixture (β ≡ 0), still improving over Model 2 (objective falls from 2.44 to 2.26) via better tail fit and average skewness, but without capturing the procyclical skewness time series. Model 4&amp;rsquo; removes time variation from the innovation distribution (β ≡ 0) while retaining the factor structure, showing that the factor structure fit survives without time variation in skewness. Additionally, the paper discusses a special parsimony case: under ρ = 1, homothetic preferences, and no factor structure, z can be normalized away entirely, leaving no individual state variable.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-prior-work-on-non-gaussian-income-processes"&gt;Q9. How does this paper relate to and differ from prior work on non-Gaussian income processes?&lt;/h3&gt;
&lt;p&gt;Kaplan, Moll, and Violante (2018) capture leptokurtic income growth but include no business cycle variation and no factor structure. McKay (2017), McKay and Reis (2021), and Catherine (2021) allow for procyclical skewness in income risk but do not target high kurtosis or a factor structure. Bhandari, Evans, Golosov, and Sargent (2021) allow for a factor structure but do not match higher-moment properties of income risk. Other work documenting the relevant facts includes Guvenen, Ozkan, and Song (2014) for countercyclical skewness in US SSA data; Guvenen, Karahan, Ozkan, and Song (2021) for lifecycle earnings dynamics from the same source; Harmenberg (2021) and Kramarz, Nimier-David, and Delemotte (2021) for related European evidence; and Guvenen, Schulhofer-Wohl, Song, and Yogo (2017) for factor structure evidence labeled &amp;lsquo;worker betas.&amp;rsquo; This paper is the first to jointly target and fit all four properties within a single tractable process that adds only one state variable.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-and-structural-implications-highlighted-by-the-paper"&gt;Q10. What are the policy and structural implications highlighted by the paper?&lt;/h3&gt;
&lt;p&gt;Leptokurtic income risk (high kurtosis, fat tails) has quantitatively important effects on the value of social insurance and optimal redistribution (Saez, 2001; Golosov, Troshkin, and Tsyvinski, 2016) and interacts with borrowing constraints to shape the distribution of wealth and marginal propensities to consume (Kaplan, Moll, and Violante, 2018). Cyclical variation in income risk — the procyclical skewness feature — matters for the welfare cost of business cycles (Storesletten, Telmer, and Yaron, 2001; Krebs, 2003, 2007) and for the optimal design and welfare value of automatic stabilizers (McKay and Reis, 2021; Bhandari et al., 2021). The factor structure is relevant for cyclical variation in income inequality and for asset pricing under household heterogeneity (Mankiw, 1986; Constantinides and Duffie, 1996; Constantinides and Ghosh, 2016). The scope condition throughout is male US workers in the SSA administrative data; no direct results are provided for female workers, self-employed individuals, or other countries, though the modeling framework is general.&lt;/p&gt;
&lt;h3 id="q11-what-practical-guidance-does-the-paper-provide-for-incorporating-the-process-into-dynamic-models"&gt;Q11. What practical guidance does the paper provide for incorporating the process into dynamic models?&lt;/h3&gt;
&lt;p&gt;The paper provides explicit Bellman equation structure: cash on hand m and the persistent income state z are the two endogenous individual state variables (z being the single income-process state variable), with individual parameters γ and κ treated as fixed effects. Income at each node requires evaluating a closed-form expression from Equation 1. Expectations over next-period z and ζ are handled via quadrature, with the time-varying mixture of normals requiring quadrature nodes that shift with the aggregate state S and S′ — following McKay and Reis (2021). Under the special case ρ = 1, homothetic preferences, and no factor structure, all variables can be normalized by exp(z + γ), eliminating z as a state variable and reducing the problem to one with no idiosyncratic income state. The authors note that a perpetual-youth demographic structure avoids tracking age as a state variable.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Procyclical skewness&lt;/strong&gt;: In the paper&amp;rsquo;s sense: the Kelley skewness of the cross-sectional distribution of one-year and five-year income growth rates falls significantly during every NBER recession (distribution shifts left — more large negative shocks, fewer large positive ones) and rises during expansions, while the standard deviation of that distribution shows no discernible cyclical pattern. This is a feature of the income shock distribution itself, not of average income levels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nonemployment shock with scarring&lt;/strong&gt;: A transitory income loss event modeled as an exponential random variable ℓi,t ∈ [0,1] (representing the fraction of income lost) arriving with probability ~45% per year. A fraction ψ of this transitory shock is loaded permanently onto the persistent income state — the &amp;lsquo;scarring&amp;rsquo; effect — so that re-employed workers do not fully return to their pre-shock income trajectory. In the paper&amp;rsquo;s model this single mechanism generates high kurtosis, thick double-Pareto tails, and asymmetric tail slopes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time-varying normal mixture for persistent innovations&lt;/strong&gt;: A three-component mixture of normals for the AR(1) innovation η in which the component means (not probabilities or variances) shift proportionally to contemporaneous aggregate wage income growth via a loading parameter β. A mean-preserving normalization ensures no effect on average income. This mean-shifting mechanism moves probability mass between the central and tail components of the innovation distribution, generating procyclical skewness while keeping income growth variance acyclical.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Factor structure in business cycle incidence&lt;/strong&gt;: A systematic, pre-determined relationship between a worker&amp;rsquo;s position in the persistent income distribution and the magnitude of income change experienced during a given recession or expansion. Modeled as a piecewise-linear function f(γi + zi,t) that multiplies the aggregate income component wt, with slopes that differ below and above an estimated kink point. Empirically, the factor structure produces a V-shaped incidence pattern: income losses in deep recessions are largest at both the bottom and top of the pre-recession income distribution, and smallest in the middle.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Income scarring parameter (ψ)&lt;/strong&gt;: The fraction of a transitory nonemployment shock ζi,t that is permanently loaded onto the persistent income state zi,t via the equation ˜ηi,t = ηi,t + ψζi,t. Estimated at 9.4% in Model 2 and 15.1% in Model 3. Controls the degree to which transitory shocks generate long-lasting income effects and determines the relative steepness of the left versus right tails of the annual income growth distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneous Income Profiles (HIP)&lt;/strong&gt;: Individual-specific linear deterministic growth-rate slopes κi distributed with standard deviation σκ = 0.015 (calibrated from Baker 1997 and Guvenen et al. 2021), representing permanent heterogeneity in the steepness of individual income trajectories over the lifecycle. Introducing HIP allows the AR(1) persistence parameter ρ to be estimated below 1 (≈0.80 in Models 5–6) while preserving the near-linear age-variance profile, because the convex variance contribution of heterogeneous slopes offsets the concavity induced by mean-reverting persistent shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Kelley skewness&lt;/strong&gt;: In the paper&amp;rsquo;s use: a robust, percentile-based measure of skewness defined as [(P90 − P50) − (P50 − P10)] / (P90 − P10), which the paper prefers for income growth distributions because it is less sensitive to extreme outliers than moment-based skewness. Used as the primary target for capturing business cycle variation in the shape of the income growth distribution.&lt;/p&gt;</description></item><item><title>Are Targeted Matching Schemes Effective in Stimulating Retirement Savings?</title><link>https://macropaperwarehouse.com/papers/are-targeted-matching-schemes-effective-in-stimulating-retirement-savings/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/are-targeted-matching-schemes-effective-in-stimulating-retirement-savings/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Governments across ten-plus countries — including Australia, the United States, Germany, and New Zealand — have introduced matching schemes to encourage low- and middle-income earners to contribute voluntarily to private pensions, motivated by the concern that progressive tax systems give these groups weaker incentives to save for retirement than high-income earners. Whether such schemes actually raise retirement savings is theoretically ambiguous: by reducing the cost of contributing they produce a substitution effect favoring more contributions, but the government payment also raises anticipated retirement income, reducing the desire to save further (a retirement income effect). The sign of the net effect depends on the distribution of contributions that would have occurred in the scheme&amp;rsquo;s absence, and it is especially unclear for those who would already have contributed above the matching ceiling.&lt;/p&gt;
&lt;p&gt;This paper tests the full set of theoretical predictions from a two-period intertemporal savings model using Australia&amp;rsquo;s Superannuation Co-contribution Scheme as a clean natural experiment. The scheme matches personal after-tax superannuation contributions up to $1,000 per year at a single, flat matching rate that varied over time — 100% in 2003-04 and 2009-10 to 2011-12, 150% in 2004-05 to 2008-09, and 50% from 2012-13 onward — and eligibility is phased out smoothly with income (no sharp income discontinuity, unlike the US Saver&amp;rsquo;s Credit), removing incentives for income manipulation. The maximum co-contribution payment was accordingly $1,000, $1,500, or $500 depending on the period. Estimation uses the ATO Longitudinal Information Files (ALife), a 10% random sample of all registered Australian tax filers linked longitudinally since 1990-91, covering 1,416,622 individual-year observations from 1999-2000 to 2016-17. The authors employ a first-differenced estimator exploiting within-individual variation in eligibility and match rates across years, conditioning on income, income squared, demographic controls, and year fixed effects.&lt;/p&gt;
&lt;p&gt;On the extensive margin, eligibility is associated with statistically significant but small increases in the probability of making any voluntary after-tax contribution: 0.6 percentage points at the 50% match rate, 0.9 percentage points at 100%, and 2.7 percentage points at 150%. Bunching at the salient $1,000 eligible maximum rises monotonically with the match rate: 0.23, 0.84, and 1.4 percentage points, respectively. Below $1,000, the probability of contributing in that range increases by 1.2, 1.6, and 2.7 percentage points — consistent with the substitution effect drawing in non-contributors and low contributors. Above $3,000, however, the probability of contributing falls significantly at all match rates: -0.66 pp (50%), -0.91 pp (100%), and -0.98 pp (150%), consistent with a retirement income windfall effect inducing high contributors to reduce their contributions toward the kink at $1,000.&lt;/p&gt;
&lt;p&gt;These opposing forces mean that average personal after-tax contributions (intensive margin) fall under all match-rate regimes: by $24.0 (50%), $24.6 (100%), and $6.49 (150%) per person-year, all significant. The attenuation of the fall at the 150% rate is consistent with substitution effects beginning to overshoot the eligible maximum and partially offsetting the income effect. When the government co-contribution payment itself is included, the combined personal-plus-government contribution rises ($40 at 100%, $126 at 150%), but these gains are partly offset by crowding out of voluntary concessional (salary sacrifice, pre-tax) contributions: eligibility is associated with 1.1 percentage point and 0.8 percentage point reductions in the proportion making voluntary concessional contributions at the 50% and 100% match rates respectively.&lt;/p&gt;
&lt;p&gt;Symmetry tests show no evidence of persistent habit formation: increases and decreases in treatment intensity produce contributions changes of roughly equal and opposite magnitudes on the extensive margin (gains +1.3 pp, losses -1.4 pp), ruling out the hypothesis that temporary eligibility establishes lasting savings behavior.&lt;/p&gt;
&lt;p&gt;Heterogeneity analysis reveals that the small average response reflects constrained liquidity. The response is largest for partnered females (+2.7 pp on the extensive margin), who have more discretionary income as secondary earners, and for those in the top permanent-income quintile (+3.6 pp), compared with bottom quintile (+0.4 pp) and second quintile (+0.7 pp). Responses increase with age and with lagged superannuation balance, with those holding balances above $100,000 responding at around 2.5 pp versus only 0.6 pp for those with balances below $25,000. There is no evidence that information is the binding constraint: respondents who use a tax consultant respond no more than those who self-file, and survey data document approximately 80% scheme awareness among superannuants.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central policy conclusion is that even a simple, transparent, and generous co-contribution scheme fails to meaningfully raise contributions of those it targets. The negative intensive margin arises because the scheme acts as a windfall for existing high contributors rather than newly inducing saving. These findings raise doubts about analogous reforms under discussion for the US Saver&amp;rsquo;s Credit.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-key-threats-to-it"&gt;Q1. What is the identification strategy and what are the key threats to it?&lt;/h3&gt;
&lt;p&gt;The primary estimator is a first-differenced OLS regression exploiting within-individual, year-on-year changes in co-contribution eligibility and match rates. Because the income thresholds shift over time and individuals&amp;rsquo; income fluctuates, the same person can move in and out of eligibility or across match-rate regimes, providing 16 distinct combinations of year-on-year changes in treatment status that identify the three match-rate coefficients. The key identification assumption is that first-differenced treatment indicators are contemporaneously uncorrelated with first-differenced idiosyncratic shocks. The main threat is income endogeneity — treatment is inversely related to income, and unobserved preferences to save may correlate with income. The authors address this by differencing out individual fixed effects and including income and income-squared as controls. They also test whether income manipulation around thresholds is occurring (it is not, unlike the US Saver&amp;rsquo;s Credit): frequency distributions of income show no bunching at the eligibility thresholds. The only income bunching observed is at the top of the lowest tax bracket (~$37,000), unrelated to scheme thresholds. As a robustness check, the authors also estimate individual fixed-effects models; results are broadly consistent, except for a theoretically inconsistent anomaly on the extensive margin for the 50% rate in the fixed-effects version, which the authors attribute to that model&amp;rsquo;s stricter exogeneity assumption being more likely violated in a life-cycle context.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-decompose-income-and-substitution-effects-and-what-is-the-empirical-test-for-each"&gt;Q2. How does the paper decompose income and substitution effects, and what is the empirical test for each?&lt;/h3&gt;
&lt;p&gt;The paper uses a two-period intertemporal model to show that the scheme creates a kinked budget constraint at the maximum eligible contribution (pmax). Those who would have contributed below pmax in the absence of the scheme face a lower cost of saving (substitution effect) and may increase contributions up to pmax. Those who would have contributed above pmax receive the co-contribution as a pure retirement income windfall, face no substitution incentive (the matching rate applies only below pmax), and respond only via a negative income effect by reducing contributions toward pmax. The empirical decomposition tests these predictions by estimating contribution probabilities in three ranges: contributions up to $1,000 (captures substitution effect), contributions between $1,001 and $3,000 (theoretically ambiguous — outflow from above $3,000 may offset inflow to $1,000), and contributions above $3,000 (captures negative income effect, as this range sits entirely above pmax). In Figure 5, the paper plots cumulative distribution function effects for each match rate across $100 increments from $0 to $10,000, showing negative effects on the CDF below $1,000 (substitution draws people above zero) and positive effects at and above $1,000 (income effect shifts mass below the maximum). The sign pattern is consistent with theory across all three match rates, and is more pronounced at higher match rates.&lt;/p&gt;
&lt;h3 id="q3-what-does-the-paper-find-about-bunching-at-the-1000-maximum-eligible-contribution"&gt;Q3. What does the paper find about bunching at the $1,000 maximum eligible contribution?&lt;/h3&gt;
&lt;p&gt;Eligibility is associated with significantly increased probability of contributing exactly $1,000, rising with the match rate: 0.23 pp at 50%, 0.84 pp at 100%, and 1.4 pp at 150%. The alternative specification distinguishing full eligibility (income below lower threshold, pmax = $1,000) from part eligibility (income in the tapered zone, pmax &amp;lt; $1,000) shows that part-eligible individuals also bunch significantly at $1,000 despite being entitled to match payments only for contributions below $1,000. This highlights the salience of the nominal maximum — people in the tapered zone treat $1,000 as the focal contribution amount rather than computing their individual optimal eligible contribution. The ATO online calculator does not report the maximum eligible contribution for part-eligible individuals, which likely reinforces this behavioral pattern.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-crowding-out-effects-on-unmatched-concessional-contributions"&gt;Q4. What are the crowding-out effects on unmatched (concessional) contributions?&lt;/h3&gt;
&lt;p&gt;The co-contribution scheme is associated with reductions in the use of voluntary concessional contributions (salary sacrifice, which are pre-tax and thus ineligible for matching). Using data from 2009-10 to 2016-17 (when salary sacrifice can be separated from compulsory employer contributions), the authors find that eligibility reduces the proportion of people making voluntary concessional contributions by 1.1 pp at the 50% match rate and 0.8 pp at the 100% match rate (both statistically significant). The data do not allow estimation at the 150% match rate because salary sacrifice records are unavailable before 2010. This crowding out compounds the scheme&amp;rsquo;s limited impact on total retirement savings: the net addition to retirement income from voluntary contributions is even smaller than the after-tax contribution estimates suggest. The mechanism attributed is the income windfall effect — for those who already made after-tax contributions in the absence of the scheme, the matching payment reduces their need for additional voluntary pre-tax saving.&lt;/p&gt;
&lt;h3 id="q5-is-there-evidence-of-asymmetry-in-scheme-effects--do-people-who-gain-eligibility-respond-differently-from-those-who-lose-it"&gt;Q5. Is there evidence of asymmetry in scheme effects — do people who gain eligibility respond differently from those who lose it?&lt;/h3&gt;
&lt;p&gt;The symmetry test in Equation (6) separates increases in treatment intensity (becoming eligible or moving to a higher match rate) from decreases (losing eligibility or moving to a lower rate). On the extensive margin, the effects are approximately symmetric: gaining intensity raises the contribution rate by 1.3 pp on average, while losing intensity reduces it by 1.4 pp. This rules out the &amp;rsquo;early targeting&amp;rsquo; hypothesis that short-term scheme exposure establishes lasting contribution habits that persist after eligibility ends. There is, however, some distributional asymmetry: bunching at $1,000 and the negative income effect above $3,000 are weaker in response to decreases in treatment intensity than to increases, suggesting some stickiness — people whose treatment falls may sustain slightly higher contributions for a period because prior co-contributions made them feel wealthier. But on the intensive margin, the reduction in average contributions is significant when treatment increases and statistically indistinguishable from zero when treatment decreases. The overall conclusion is no meaningful asymmetry that would justify life-cycle &amp;lsquo;seeding&amp;rsquo; arguments for young-age eligibility phased out later.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-in-responses-is-documented-and-what-does-it-imply-about-who-benefits"&gt;Q6. What heterogeneity in responses is documented, and what does it imply about who benefits?&lt;/h3&gt;
&lt;p&gt;Responses are largest among groups with greater discretionary income relative to their current consumption needs. Partnered females respond at 2.7 pp on the extensive margin (versus 1.2 pp for partnered males, 1.1 pp for single females, and 0.6 pp for single males). The interpretation is that partnered females are more likely to be secondary earners whose income is discretionary, reducing the liquidity cost of foregoing current consumption. The extensive margin response increases monotonically with permanent income quintile: 0.4 pp (bottom), 0.7 pp (2nd), 1.3 pp (3rd), 1.8 pp (4th), and 3.6 pp (top). Those in the top quintile are eligible only when their transitory income is temporarily low, and they appear to have both the liquid assets and the foresight to exploit the scheme. Responses increase with age, consistent with older workers facing lower liquidity constraints and having stronger retirement income motives. Lagged superannuation balance matters: those with balances above $100,000 respond at ~2.5 pp versus ~0.6 pp for those with balances below $25,000 — the scheme does not help low-balance individuals catch up. Importantly, there is no evidence that scheme uptake is constrained by information: tax-agent filers and self-filers respond at similar rates (~1.3 pp vs ~1.9 pp), and external surveys show roughly 80% public awareness. This rules out information provision as a policy lever likely to substantially raise the scheme&amp;rsquo;s impact.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-study-relate-to-and-differ-from-prior-evaluations-of-the-us-savers-credit-and-german-riester-schemes"&gt;Q7. How does this study relate to and differ from prior evaluations of the US Saver&amp;rsquo;s Credit and German Riester schemes?&lt;/h3&gt;
&lt;p&gt;Prior work on the Saver&amp;rsquo;s Credit (Duflo et al. 2007, Ramnath 2013, Heim and Lurie 2014) found small or null effects, attributed mainly to the scheme&amp;rsquo;s complexity — non-refundable tax credit with match rates of 11%, 25%, or 100% depending on income thresholds that create sharp discontinuities and strong income manipulation incentives. The Riester scheme (Corneo et al. 2009, 2010) showed zero effects on total savings, attributed to its complex co-contribution formula where the effective match rate depends on income and number of children, making the true incentive opaque. This paper&amp;rsquo;s contribution is to evaluate a scheme explicitly designed to avoid those complexities: a single flat match rate, co-contribution paid directly to the pension account, eligibility smoothly phased out with no discontinuities, and near-universal institutional coverage through mandatory superannuation. This design is analogous to the Duflo et al. (2006) H&amp;amp;R Block field experiment (which found 5–11 pp increases in contribution rates for 20–50% match rates), and the paper can be read as asking whether those larger field-experiment effects generalize to a national, ongoing program at comparable design simplicity. The answer is no: the national scheme produces responses an order of magnitude smaller than the field experiment. The paper attributes this partly to the field experiment&amp;rsquo;s &amp;lsquo;one-time-only&amp;rsquo; nature (creating urgency), potential interaction with Saver&amp;rsquo;s Credit tax refunds, and selection of H&amp;amp;R Block clients. The Australian study also goes beyond prior work by estimating distributional effects (contribution ranges), crowding out of unmatched contributions, and symmetry tests — none of which were examined in the prior national scheme evaluations.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-papers-policy-implications-and-their-scope-conditions"&gt;Q8. What are the paper&amp;rsquo;s policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary implication is that co-contribution matching schemes, even when simple, generous, and widely known, are likely to produce small effects on retirement savings of low- and middle-income earners. The mechanism is that many in the eligible population already contributed more than the scheme maximum and treat the matching payment as a windfall, reducing personal contributions. The scheme is particularly ineffective for the lowest permanent-income earners, who face binding liquidity constraints and respond least even when they are aware of the scheme. This is directly relevant to proposed US reforms of the Saver&amp;rsquo;s Credit (the Retirement Security and Savings Act considered by Congress at time of writing) that would convert it to a direct co-contribution more like Australia&amp;rsquo;s scheme — the paper&amp;rsquo;s results suggest such simplification may not yield large savings increases. A scope condition concerns institutional context: Australia has near-universal mandatory superannuation with employer contributions at 9.5% of earnings, which may reduce the marginal value of voluntary contributions. The authors acknowledge that responses might be higher in countries without mandatory employer coverage, though the finding that lower-balance individuals respond least makes this qualification weak. A second scope condition is that the scheme excludes compulsory employer contributions from the matching base, so the results speak specifically to voluntary behavior. Future research is identified on whether tightening access to public pensions (raising the pension access age) would increase voluntary contributions among low-income earners who currently rely on public pensions as their retirement backstop.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-conducted"&gt;Q9. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The authors report four main robustness exercises. First, they estimate an individual fixed-effects model alongside the first-differenced model; results are broadly consistent, with the noted exception of a theoretically inconsistent anomaly at the 50% match rate for the extensive margin in the fixed-effects version, attributed to violation of the strict exogeneity assumption. This validates the first-differenced approach as the preferred specification. Second, they extend the base model to distinguish full eligibility (income at or below the lower threshold, pmax = $1,000) from part eligibility (income in the tapered zone, pmax &amp;lt; $1,000), confirming that even partial eligibility generates bunching at the salient $1,000 level. Third, they examine distributional predictions by estimating the model for 100 incremental contribution thresholds from $0 to $10,000 (Figure 5), verifying that the CDF-effect pattern is consistent with the theoretical predictions across all three match rates. Fourth, information access is tested by interacting scheme response with whether a tax agent was used to lodge the return; the absence of any significant difference between tax-agent filers and self-filers, combined with documented high public awareness, eliminates information deficiency as an explanation for the small response.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Co-contribution matching scheme&lt;/strong&gt;: A government program that pays a specified fraction (the matching rate) of the individual&amp;rsquo;s voluntary personal pension contributions up to a maximum eligible contribution ceiling, credited directly to the individual&amp;rsquo;s retirement account — as distinct from a tax credit that may not reach the account.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Retirement income effect (windfall effect)&lt;/strong&gt;: The tendency of matching payments to reduce voluntary personal contributions among those who would have contributed above the scheme maximum in the scheme&amp;rsquo;s absence: because the government contribution supplements their retirement income regardless of their own effort, they rationally reduce personal saving to the eligible maximum.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Substitution effect (in this scheme)&lt;/strong&gt;: The scheme&amp;rsquo;s reduction in the effective cost of contributing by raising the return to each dollar contributed, inducing those who previously contributed below the eligible maximum to increase contributions toward that maximum.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bunching at the eligible maximum&lt;/strong&gt;: Mass concentration of contributions at exactly $1,000 (the scheme&amp;rsquo;s nominal maximum eligible contribution), drawing both from below (via the substitution effect) and from above (via the income/windfall effect), and reinforced by the salience of the round-number maximum even for part-eligible individuals whose true eligible maximum is below $1,000.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Permanent income (in this context)&lt;/strong&gt;: The predicted value of long-run log total personal income estimated from a Mincer-style regression including individual fixed effects, used to distinguish individuals who are structurally low-income (and face genuine liquidity constraints) from those whose transitory income is temporarily low and who are high-permanent-income individuals exploiting the scheme.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crowding out of concessional contributions&lt;/strong&gt;: The reduction in voluntary pre-tax (salary sacrifice) superannuation contributions associated with scheme eligibility, reflecting the income windfall from the matching payment reducing the need for supplementary retirement saving through the pre-tax channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Symmetry of scheme effects&lt;/strong&gt;: The property that the contribution response to gaining eligibility (or a higher match rate) is equal in magnitude and opposite in sign to the response to losing eligibility (or a lower match rate); symmetry implies no lasting habit formation from scheme exposure and rules out &amp;rsquo;early targeting&amp;rsquo; strategies aimed at establishing lifetime saving patterns.&lt;/p&gt;</description></item><item><title>Armed conflict exposure and trust: evidence from a natural experiment</title><link>https://macropaperwarehouse.com/papers/armed-conflict-exposure-and-trust-evidence-from-a-natural-experiment/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/armed-conflict-exposure-and-trust-evidence-from-a-natural-experiment/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks how individual-level exposure to internal armed conflict shapes social capital, specifically trust in institutions and trust in people. The question matters because trust is a core component of social capital that underpins cooperation, economic growth, financial development, political participation, and post-conflict recovery; yet the empirical literature is split between studies finding conflict erodes trust and studies finding &amp;ldquo;post-traumatic growth&amp;rdquo; that enhances pro-sociality. The authors argue prior work cannot cleanly identify causal effects because of non-random selection into exposure, attrition from migration/death, and confounding conflict-induced changes in the socio-economic environment.&lt;/p&gt;
&lt;p&gt;The empirical strategy exploits a natural experiment in Turkey: mandatory conscription assigns every male citizen, via a lottery, to a military base, and a significant share are randomly sent to bases in the eastern/south-eastern conflict zone where the state has fought the PKK since 1984. By sampling ex-recruits who live in peaceful western districts, exposure during military service is the respondents&amp;rsquo; only personal contact with the conflict, isolating individual-level effects from environmental confounds. Data come from a field survey of 5,024 randomly selected adult males in 29 western districts in summer/fall 2019 (response rate 83%); eligible men had completed service between 1984 and 2014. Only 5 respondents did not answer the military-service questions.&lt;/p&gt;
&lt;p&gt;Two exposure measures are built. ACE (Exposure to Armed Conflict Environment) is the standardized number of combatant casualties in the county and during the period of a respondent&amp;rsquo;s service, drawn from the Turkish State-PKK Conflict Event Database; its variation comes from four exogenous components (birthdate-driven timing, regulation-driven duration, clash intensity, and lottery-assigned location). TDE (Traumatic Direct Experiences) is a binary indicator equal to 1 if the respondent was wounded in armed clashes or had someone around them killed/hurt; 2% reported being wounded and 15% reported others around them killed or hurt. ACE and TDE correlate only 0.25. Two trust outcomes: Institutional Trust (average of 14 five-point items: army, judiciary, parliament, TV, newspapers, parties, clergy, universities, environmental orgs, charities, police, banks, private companies, EU) and Social Trust (trust in unfamiliar people / strangers). The army was the most trusted institution (~75% high trust vs. 43% for courts, 35% for parliament). Estimation is OLS with age, education, and minority controls, standard errors clustered at the living-block level.&lt;/p&gt;
&lt;p&gt;Main findings: the two exposure types have opposing effects. In the preferred specification including both measures, ACE raises Institutional Trust (about 0.02, significant at 5%) and Social Trust (about 0.03, significant at 5%), while TDE lowers Institutional Trust (about -0.15, 5%) and Social Trust (about -0.11, 1%). ACE is insignificant when TDE is omitted because it then pools traumatized and non-traumatized recruits, biasing it toward zero. There is no significant ACE-by-TDE interaction, so the negative trauma effect is independent of conflict intensity. Effects are similar in sign and magnitude across both trust dimensions, indicating an encompassing change rather than institution-specific distrust. Interactions with time-since-service are insignificant, implying the effects are permanent.&lt;/p&gt;
&lt;p&gt;Mechanism: the authors invoke Janoff-Bulman&amp;rsquo;s (1992) &amp;ldquo;shattered assumptions&amp;rdquo; theory. TDE is positively associated with depression and insecurity indexes, which in turn correlate negatively with both trust measures; ACE is not significantly related to depression/insecurity. There is no significant relationship between exposure and trust in the army, ruling out an accountability mechanism. Heterogeneity by in-group: TDE raises trust in family (coping mechanism) but, like strangers, friends show positive ACE and (insignificant) negative TDE effects, arguing against parochialism as the main driver. Implications: distinguish contextual from direct exposure; design psychological recovery programs for veterans; estimates are likely conservative given the limited 6-18 month exposure window.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Identification relies on Turkey&amp;rsquo;s conscription lottery, which randomly assigns drafted men to military bases, a significant share of which lie in the eastern/south-eastern conflict zone. Because the sample is drawn only from peaceful western districts, service is the respondents&amp;rsquo; sole exposure to the conflict, isolating individual-level effects from conflict-induced changes in the socio-economic environment. ACE&amp;rsquo;s variation comes from four exogenous components: birthdate-driven timing of service, regulation-driven service duration (18 months in the 80s, 15 in 1992, 18 in 1995, 15 in 2003, 12 in 2014), clash intensity around the base, and lottery-assigned location. Threats: (1) non-random base assignment - addressed by balance tests (Table 2) showing no systematic differences in age, ethnicity, or height by conflict-zone assignment; education differs because college graduates are slightly skewed toward western bases (40% of non-college-grads served in the east vs. 30% of college grads), but the difference vanishes when college graduates (9.3% of sample) are excluded, education is controlled in all specs, and a no-college-grad sample (Table A2) is robust; (2) self-selection into dangerous tasks/violence for TDE - addressed by the fact that task assignments are made by command at the start of service before behavior is observed, and Table 3 balance tests show wounded vs. non-wounded respondents do not differ on pre-military characteristics; an alternative TDE (observing a fellow soldier hurt/killed, immune to own risk-taking) yields similar results (Table A1).&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The proposed mechanism is a transformation of fundamental world assumptions (benevolence, meaning, safety of the world) per Janoff-Bulman (1992). Distinguishing tests: (1) TDE affects a broad range of trust dimensions but is NOT significantly related to trust in the army, ruling out an accountability interpretation (which would predict distrust concentrated on state security institutions) and a comradeship interpretation (which would predict effects only on social trust). (2) TDE is positively and significantly associated with depression and insecurity indexes (Tables 7-8), and these indexes are themselves negatively and significantly related to both trust measures, consistent with shattered world assumptions. (3) ACE is not significantly associated with depression/insecurity; the authors note these scales are worded to detect negative states and may miss the positive feelings ACE could elicit, and that indirect environmental exposure plausibly has weaker effects on fundamental beliefs than direct trauma.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;The central heterogeneity is by exposure type: contextual exposure (ACE) raises trust, direct trauma (TDE) lowers it. No significant ACE-by-TDE interaction, so trauma&amp;rsquo;s effect does not depend on conflict intensity. No significant moderation by time since service (Table 6), implying permanent effects. In-group heterogeneity (Table 9, ordered logit): TDE significantly raises trust in family (coefficient 0.26, 5%), interpreted as a coping mechanism of retreating to closest networks; trust in friends shows positive ACE (0.07, 5%) and negative but insignificant TDE, mirroring the stranger result. The similar pattern for strangers and friends argues against parochialism as the primary driver.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;(1) Alternative TDE defined as observing a fellow soldier hurt/killed, more immune to own risk-taking (Table A1) - results unchanged. (2) Excluding college graduates (Table A2) - results unchanged. (3) Tobit specification accounting for the censored nature of trust measures (Table A3) - similar results. (4) Including a conflict-zone dummy and base-district fixed effects (Tables A4-A5) to absorb unobserved location heterogeneity (though the authors note these likely absorb part of the ACE variation, so they are not in the baseline). (5) Separate results for each of the 14 institutional-trust dimensions (Table A6) and excluding one dimension at a time from the composite index - results stable. (6) Alternative standard-error clustering at home-district or region levels - unchanged.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds on the draft-lottery natural-experiment tradition (Angrist 1990 on Vietnam; Angrist-Chen 2011; Galiani et al. 2011; Grossman et al. 2015) and the conflict-and-social-capital literature (Rohner et al. 2013; Cassar et al. 2013; Bauer et al. 2016; Kijewski-Freitag 2018). It differs by: (1) cleanly identifying causal effects free of environmental confounds, since trust is measured in untouched western locations rather than in transformed post-conflict settings; (2) carefully separating contextual from direct exposure, which many studies cannot; (3) proposing a novel individual-level psychological mechanism (shattered world assumptions) rather than the economic/institutional-legacy channels (Besley-Reynal-Querol 2014; Nunn-Wantchekon 2011; Grosjean 2014) or the inter-group-competition/parochialism explanation (Bauer et al. 2016). The authors argue the heterogeneity they document can help reconcile the conflicting positive and negative findings in prior literature - prior &amp;lsquo;pro-social&amp;rsquo; effects may reflect coping-driven re-creation of safe social space (consistent with Grosjean&amp;rsquo;s (2014) &amp;lsquo;dark nature&amp;rsquo; of conflict-induced pro-sociality), not genuine restoration of trust.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Two main implications: (1) researchers and policy advisers should carefully distinguish contextual from direct conflict exposure when studying behavioral outcomes; (2) the findings inform the design of psychological and social recovery programs for combat veterans and victimized post-conflict populations. Scope conditions: the study is specific to the Turkish conflict setting and limited to male ex-combatants; it remains open whether effects generalize to women, civilians, or other countries. Because exposure lasted only a pre-determined 6-18 months after which recruits returned to peaceful lives, the authors argue estimates are conservative relative to populations living in protracted conflict environments.&lt;/p&gt;
&lt;h3 id="q7-what-additional-findings-or-caveats-are-noted"&gt;Q7. What additional findings or caveats are noted?&lt;/h3&gt;
&lt;p&gt;The authors report (results not shown) that individuals with traumatic experiences are more likely to participate in political organizations, and cite Kibris-Nelson (2021) that such individuals are more likely to start their own businesses (while being less successful at it), consistent with coping strategies of creating a controllable environment. They concede the mechanism evidence for the positive ACE effect is &amp;lsquo;somewhat less clear&amp;rsquo; than for TDE, and offer an alternative possibility that whether intense-environment survival raises trust may be moderated by how heroically the veteran&amp;rsquo;s social network views his service. The depression subscale is the 6-item Brief Symptoms Inventory; insecurity is an 8-item scale. Roughly 6.5 million of the 15 million men drafted since 1984 are estimated to have served in the conflict zone.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Exposure to Armed Conflict Environment (ACE)&lt;/strong&gt;: A standardized, individual-specific measure of contextual conflict exposure equal to the number of combatant casualties in the county and during the time period of a respondent&amp;rsquo;s military service. It captures immersion in the conflict environment with high geo-temporal precision and is treated as exogenous because its components (birthdate-driven timing, regulation-driven duration, clash intensity, lottery-assigned location) are outside the individual&amp;rsquo;s control.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Traumatic Direct Experiences (TDE)&lt;/strong&gt;: A binary indicator equal to 1 if a respondent was personally wounded in armed clashes or had someone around them killed or hurt during military service. It captures direct, personal experience of violence as distinct from mere presence in a conflict environment; in the sample 2% were wounded and 15% had others around them hurt/killed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Institutional Trust&lt;/strong&gt;: In the paper&amp;rsquo;s sense, the simple average of a respondent&amp;rsquo;s 5-point Likert trust ratings across 14 public and private organizations (army, judiciary, parliament, media, parties, clergy, universities, environmental orgs, charities, police, banks, private companies, EU) - deliberately broad so as not to over-weight state institutions directly tied to the conflict.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Trust&lt;/strong&gt;: A generalized form of trust measured by how much a respondent trusts people they are not familiar with (strangers), rather than the vaguer &amp;lsquo;most people&amp;rsquo; wording, chosen to minimize in-group/out-group and ethnic associations and isolate generalized trust in others.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Shattered assumptions&lt;/strong&gt;: The paper&amp;rsquo;s operative mechanism, drawn from Janoff-Bulman (1992): people hold core assumptions that the world is benevolent, meaningful, and safe; traumatizing experiences shatter these positive assumptions, eroding deeply rooted trust - whereas surviving a dangerous environment without mishap can instead reinforce them. Trust, depression, and insecurity are treated as observable implications of these otherwise-unobservable world assumptions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parochialism / parochial altruism&lt;/strong&gt;: The rival hypothesis (associated with Bauer et al. 2016) that conflict exposure increases in-group favoritism while eroding out-group trust. The paper tests and largely rejects it as the primary driver because ACE raises trust in both strangers and friends and the in-group (family) pattern does not match parochial predictions.&lt;/p&gt;</description></item><item><title>Balancing Work and Care: How Workplace Factors Can Mitigate the Gendered Impacts of Caregiving</title><link>https://macropaperwarehouse.com/papers/balancing-work-and-care-how-workplace-factors-can-mitigate-the-gendered-impacts-of-caregiving/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/balancing-work-and-care-how-workplace-factors-can-mitigate-the-gendered-impacts-of-caregiving/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper examines how workplace environments shape the economic consequences that fall on mothers — but not fathers — when a child is diagnosed with cancer. The motivation is a gap in the caregiving-and-labor-markets literature: while the earnings penalties from childbirth are well-documented, less is known about caregiving shocks that arrive later in childhood, or about whether and how the firm, occupation, or industry a parent works in moderates those penalties.&lt;/p&gt;
&lt;p&gt;The empirical setting is Australia. The authors use the ABS Person Level Integrated Data Asset (PLIDA), a longitudinal administrative database linking tax records (ATO, 2005–2022), Medicare health records, and 2011 Census occupation and hours data. A distinctive feature is matched employer-employee identifiers, enabling construction of workplace characteristics at the firm, occupation, and industry levels. The sample comprises 3,258 families in which a child (age 4–18, average age 12.98) began chemotherapy between 2012 and 2023 and both parents were employed two years before treatment. Pre-diagnosis average earnings are $37,639 for mothers and $79,702 for fathers (CPI-adjusted to 2012).&lt;/p&gt;
&lt;p&gt;The identification strategy is a dynamic difference-in-differences (DiD) model following Fadlon and Nielsen (2019, 2021). The treatment group consists of parents whose children started chemotherapy between 2012 and 2017; the control group consists of parents whose children will receive the same diagnosis later, between 2018 and 2023, with placebo treatment assigned six years before actual treatment. Individual fixed effects absorb time-invariant heterogeneity; year fixed effects absorb common trends. Childhood cancer — specifically chemotherapy-requiring cancer — is treated as a largely random shock with no pre-trend in earnings or employment between treated and control families before diagnosis.&lt;/p&gt;
&lt;p&gt;Main findings on the average effects: Maternal earnings fall by $5,608 in the year chemotherapy begins (14.9% of baseline earnings). The earnings decline persists for at least three years even as measured caregiving intensity (child healthcare service use) returns to baseline by year 3, leaving earnings approximately 9.7% below baseline in year 3 (−$3,645). The primary mechanism is a reduction in hours worked rather than outright job exit: employment falls by 4.9 percentage points in year 0, peaking at a decline of 5.6 percentage points two years post-treatment, a modest reduction relative to the earnings loss. Job-to-job transitions are not significantly elevated. Mental health service use (therapy, antidepressants, anxiolytics) shows no significant change for either parent, ruling out a mental health channel and reinforcing that caregiver time demands drive the result. Fathers experience no statistically significant change in earnings, employment, or job transitions across all specifications.&lt;/p&gt;
&lt;p&gt;Subgroup heterogeneity: The earnings penalty is substantially larger for mothers of younger children (under 12): −$9,443 in year 0, equivalent to 25.8% of that subgroup&amp;rsquo;s baseline earnings. For children with above-median healthcare utilization, the year-0 penalty is −$7,826 (21.6%).&lt;/p&gt;
&lt;p&gt;Workplace moderation — three dimensions are examined at the firm, occupation, and industry levels:&lt;/p&gt;
&lt;p&gt;(1) Gender pay gap: Mothers in occupations with below-average gender pay gaps face lower earnings losses ($5,782 vs $8,409; 16.5% vs 18.1%). The effect is significant at the occupation level but not at the firm or industry level.&lt;/p&gt;
&lt;p&gt;(2) Work hour intensity: Mothers in firms with below-median weekly hours face a year-0 earnings loss of $3,240 (9.9%) versus $7,159 (15.6%) in high-hours firms — a difference of $3,919, significant at the firm level. A parallel gap holds at the occupation level. When both firm and occupation are low-hours, the combined loss equals $2,519; when both are high-hours, it reaches $9,357 — a fourfold difference.&lt;/p&gt;
&lt;p&gt;(3) Female representation in the top 20% of earners: Mothers at firms where women are the majority of top-20%-earners suffer a penalty of $3,856 (8.3%) versus $7,799 (23.4%) elsewhere — a $3,943 mitigation at the firm level. At the occupation level the corresponding figures are $4,240 (9.2%) versus $8,356 (25.0%). Female representation in middle or bottom earnings tiers carries no significant moderating effect.&lt;/p&gt;
&lt;p&gt;In the combined specification (all firm- and occupation-level variables simultaneously), female representation in the top 20% and work hour intensity remain jointly significant; the gender pay gap loses significance, consistent with these variables being correlated. In the polar comparison between fully supportive jobs (low hours, high female senior representation, low occupation gender pay gap) and fully unsupportive jobs (opposite), the difference is dramatic: mothers in supportive jobs suffer a −$6,280 year-0 earnings hit that recovers fully by year 1, while mothers in unsupportive jobs face −$10,416 in year 0 widening to −$13,882 in year 3 before partially recovering in year 4.&lt;/p&gt;
&lt;p&gt;Policy implications (with scope conditions): The results support policies that reduce greedy-work norms and increase female representation in senior roles as instruments for attenuating the gendered economic cost of caregiving shocks. The study does not isolate specific workplace policies (e.g., formal paid leave) but identifies observable correlates of supportive environments. Effects are identified among working parents of children requiring chemotherapy; they do not generalize to cancer not requiring chemotherapy or other types of caregiving shocks without further evidence. Notably, fathers&amp;rsquo; outcomes are unresponsive to workplace factors, suggesting that social norms or intra-household bargaining — not workplace barriers per se — are the primary constraints on paternal caregiving adjustment.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The authors use a later-treated dynamic DiD, comparing parents whose children began chemotherapy 2012–2017 (treated) to parents whose children will begin the same treatment 2018–2023 (control), with the control group&amp;rsquo;s placebo treatment assigned six years before their actual treatment. Individual fixed effects absorb time-invariant heterogeneity; year fixed effects absorb macro shocks. The parallel trends assumption is validated by showing: (1) no statistically significant differences in pre-cancer demographic, socioeconomic, or workplace characteristics between treated and control groups (Figure 1); and (2) no pre-trend in earnings or employment in years -4 and -3 relative to baseline (Table A3, estimates small and insignificant). The main threats acknowledged are (a) non-random selection into workplace types — mothers who anticipate greater caregiving loads may sort into more family-friendly jobs — and (b) differences in baseline wage levels across job types. On (a), the authors argue the direction of selection bias goes the wrong way: if selection were driving results, mothers in supportive workplaces (who selected there due to caregiving preferences) would have weaker labor market attachment and larger post-shock earnings declines; instead the opposite is found. On (b), the authors show that absolute dollar declines in less-supportive workplaces also correspond to larger percentage declines relative to baseline, so the pattern is not an artifact of higher baseline wages in high-hour jobs (though Appendix Table A2 confirms mothers in high-hour and high-senior-female firms do have higher baseline earnings of around $46,000–$50,000 vs $32,000–$33,000).&lt;/p&gt;
&lt;h3 id="q2-how-is-the-caregiving-shock-defined-and-what-does-this-imply-for-external-validity"&gt;Q2. How is the caregiving shock defined and what does this imply for external validity?&lt;/h3&gt;
&lt;p&gt;The shock is defined as initiation of chemotherapy by the child, identified from Medicare prescription records using ATC codes beginning with L01 (excluding methotrexate L01BA01) and adding immunomodulators with chemotherapy-like effects. Chemotherapy initiation is treated as a reliable, time-consistent marker because it typically follows immediately from diagnosis of cancers such as acute lymphoid leukemia, astrocytoma, and neuroblastoma. The authors note explicitly that estimates do not represent the effects of childhood cancer not requiring chemotherapy (e.g., early-stage cancers treated with surgery, radiation, or immunotherapy alone). This restriction to chemotherapy-requiring cancers likely selects a sample with above-average caregiving intensity.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-main-mechanism-through-which-the-earnings-decline-operates"&gt;Q3. What is the main mechanism through which the earnings decline operates?&lt;/h3&gt;
&lt;p&gt;The primary mechanism is a reduction in hours worked rather than outright job exit. The employment decline (approximately 4.5–5.0 percentage points in years 0–2 per Table A3) is modest relative to the earnings loss of $5,608. A back-of-envelope calculation in footnote 6 shows that if 5% of mothers left the labor market at average earnings, the implied earnings drop would be only $1,882, far below the observed $5,608. Job-to-job transitions (probability of switching employer) are not significantly elevated. Mental health service use (psychological therapy, antidepressant/anxiolytic/antipsychotic prescriptions) shows no significant change for either parent (Appendix Figure A4), ruling out mental health deterioration as a channel. The persistence of earnings losses beyond the period of peak healthcare service use (which returns to baseline by year 3, per Appendix Figure A2) is consistent with stalled career trajectories — foregone promotions or skill development — or with continued but less-measured caregiving demands.&lt;/p&gt;
&lt;h3 id="q4-at-which-organizational-level-firm-occupation-or-industry-do-workplace-moderators-operate-most-strongly"&gt;Q4. At which organizational level (firm, occupation, or industry) do workplace moderators operate most strongly?&lt;/h3&gt;
&lt;p&gt;Firm and occupation levels are the dominant levels; industry-level measures are consistently insignificant for all three moderating variables. The authors interpret this as follows: industry-level measures are too broad to capture the specific work arrangements and norms that affect caregiving balance. At the occupation level, structural characteristics — profession-wide agreements, flexibility of task-based roles, part-time feasibility — directly govern how feasible it is to reduce hours without exiting employment. At the firm level, immediate workplace culture and specific HR policies apply. The relative contribution of firm vs occupation varies by the moderator: work hour intensity effects are significant at both firm and occupation levels, female senior representation is significant at both, while the gender pay gap effect is significant only at the occupation level.&lt;/p&gt;
&lt;h3 id="q5-why-does-female-representation-in-senior-roles-top-20-of-earners-mitigate-the-earnings-penalty-while-middle-and-bottom-tier-representation-does-not"&gt;Q5. Why does female representation in senior roles (top 20% of earners) mitigate the earnings penalty while middle and bottom tier representation does not?&lt;/h3&gt;
&lt;p&gt;The authors argue that women in the top-20% of earners — effectively leadership positions — are better positioned to advocate for and implement caregiving-supportive policies (paid leave, flexible scheduling). Representation in lower tiers may be indicative of a caregiving-friendly workforce composition but lacks the organizational power to shape policies. This is supported empirically: the moderating interaction is significant and economically large for top-20% female representation at both the firm (mitigating the penalty by $3,943) and occupation levels (mitigating by $4,116), while interactions for the middle 50–80% and bottom 50% earnings tiers are not statistically significant in most specifications.&lt;/p&gt;
&lt;h3 id="q6-why-does-the-occupational-gender-pay-gap-matter-for-the-earnings-penalty-but-not-the-firm-level-or-industry-level-gap"&gt;Q6. Why does the occupational gender pay gap matter for the earnings penalty but not the firm-level or industry-level gap?&lt;/h3&gt;
&lt;p&gt;The authors offer two explanations. First, occupations define the day-to-day nature of work — task structure, required hours, flexibility — in ways that make caregiving more or less compatible. Occupations that accommodate part-time and flexible scheduling tend to attract more women and develop norms that support caregiving, which in turn narrows occupational gender pay gaps. At the firm level, the same firm often contains diverse occupations with heterogeneous norms, so firm-level gender pay gap is a noisier signal. At the industry level, the measure is too aggregated. Second, narrow occupational gender pay gaps may reflect the collective bargaining power of women in female-dominated occupations (e.g., nursing), which translates into formal caregiving protections. A firm or industry may exhibit a wide gender pay gap due to male dominance in senior or high-earning roles even when specific female-dominated occupations within that firm/industry have caregiving-friendly norms. However, in the combined specification including all workplace factors simultaneously, the gender pay gap variable loses statistical significance, suggesting its initial effect was partly mediated by correlated factors (hours intensity and female senior representation).&lt;/p&gt;
&lt;h3 id="q7-how-does-the-combined-supportive-vs-unsupportive-comparison-work-and-what-does-it-show"&gt;Q7. How does the combined &amp;lsquo;supportive vs unsupportive&amp;rsquo; comparison work and what does it show?&lt;/h3&gt;
&lt;p&gt;Supportive jobs are defined as those satisfying all three criteria: low work hour intensity at both firm and occupation levels, high female representation in the top 20% of earners at both firm and occupation levels, and low gender pay gap at the occupation level (N = 2,708 mother-years). Unsupportive jobs are the opposite on all criteria (N = 2,339). Event study estimates (Table A9, Figure 3) show stark divergence. In supportive jobs, the year-0 penalty is −$6,280, and earnings recover quickly to statistically insignificant levels by years 1–4. In unsupportive jobs, the year-0 penalty is −$10,416, it widens to −$10,658 in year 2 and −$13,882 in year 3, before partially recovering in year 4. Pre-treatment estimates are not significantly different from zero in both subsamples, supporting parallel trends within each group.&lt;/p&gt;
&lt;h3 id="q8-what-heterogeneity-is-documented-by-child-and-family-characteristics"&gt;Q8. What heterogeneity is documented by child and family characteristics?&lt;/h3&gt;
&lt;p&gt;Appendix Figure A3 presents two subgroup analyses. Mothers of children under age 12 at diagnosis experience a year-0 earnings loss of −$9,443 (25.8% of baseline earnings of $36,567), substantially larger than the average. Mothers of children with above-median healthcare utilization (measured by number of medical appointments in the year following treatment initiation) experience a year-0 loss of −$7,826 (21.6% of baseline earnings of $36,278). These patterns are consistent with the interpretation that caregiving intensity — driven by child age and treatment severity — scales the maternal earnings penalty.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-conducted"&gt;Q9. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s main robustness arguments are: (1) pre-trend validation (Figures 1 and 2, Table A3) confirming no anticipatory effects and balanced pre-characteristics; (2) the selection-direction argument for workplace heterogeneity — the selection story would predict larger penalties in supportive workplaces but the opposite is found; (3) showing that absolute earnings declines in less-supportive workplaces also represent larger proportional declines relative to baseline, ruling out a level-effect interpretation; (4) the mental health non-result (Appendix Figure A4) confirming earnings effects are not confounded by parental mental health deterioration; (5) separate combined specification (Table A8) testing all workplace moderators simultaneously to address multicollinearity. The paper does not report explicit placebo tests using alternative shocks or falsification samples, nor does it report results restricted to narrow geographic areas or specific cancer types.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-prior-literature-on-caregiving-shocks"&gt;Q10. How does this paper relate to prior literature on caregiving shocks?&lt;/h3&gt;
&lt;p&gt;The paper builds most directly on three prior studies using Nordic or European administrative data: Eriksen et al. (2021, Journal of Health Economics) on childhood health shocks and parental labor supply; Breivik and Costa-Ramon (2024, Review of Economics and Statistics) on children&amp;rsquo;s health shocks and parental earnings and mental health; and Vaalavuo et al. (2023, Demography) on gender inequality from child health shocks on parental trajectories. All three find significant maternal earnings or employment losses and no or small paternal effects. The present paper&amp;rsquo;s contribution relative to these is the explicit examination of how firm-, occupation-, and industry-level workplace characteristics moderate the maternal penalty — a dimension the prior literature has not addressed. It also connects to Fadlon and Nielsen (2019, 2021) on the methodology and to the broader child-penalty literature reviewed by Cortes and Pan (2023, Journal of Economic Literature). On workplace mechanisms it connects to Goldin (2014) on &amp;lsquo;greedy jobs&amp;rsquo; and Goldin and Katz (2016) on pharmacy as a family-friendly profession.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The findings suggest that maternal earnings losses from caregiving shocks can be substantially mitigated by workplace environments characterized by lower work hour intensity and higher female representation in senior earnings tiers. This points to policies promoting: (1) reduced greedy-work norms — discouraging long-hours cultures and enabling part-time flexibility without disproportionate wage penalties; (2) greater female representation in leadership and high-earning positions, which appears to create cultural and policy environments more accommodating of caregiving. Scope conditions: the results apply to working mothers (and fathers) of children requiring chemotherapy in Australia, where Medicare provides universal healthcare coverage and existing social insurance exists. The paper explicitly does not identify specific causal mechanisms (e.g., it cannot isolate the effect of formal paid leave from culture). On fathers, the implication is that workplace factors alone are unlikely to induce fathers to increase caregiving, pointing instead to the need to shift social norms around paternal caregiving and intra-household bargaining.&lt;/p&gt;
&lt;h3 id="q12-how-do-the-australian-institutional-context-and-data-compare-to-european-studies"&gt;Q12. How do the Australian institutional context and data compare to European studies?&lt;/h3&gt;
&lt;p&gt;Australia&amp;rsquo;s PLIDA dataset is exceptional in combining population-level coverage, employer-employee identifiers (enabling firm-level workplace measures), and Medicare healthcare records (enabling both shock identification via chemotherapy and caregiving-intensity proxying via healthcare utilization). The employer identifiers are critical for this paper&amp;rsquo;s contribution — most comparable European studies cannot construct firm-level workplace characteristics. The Australian context differs from Nordic studies in terms of family policy generosity (less universal paid parental leave), but Medicare provides universal healthcare access. Pre-diagnosis earnings ($37,639 for mothers vs $79,702 for fathers) indicate a large pre-existing earnings gap, consistent with a majority-male breadwinner household structure in the sample.&lt;/p&gt;
&lt;h3 id="q13-do-fathers-outcomes-respond-to-any-workplace-factor"&gt;Q13. Do fathers&amp;rsquo; outcomes respond to any workplace factor?&lt;/h3&gt;
&lt;p&gt;In almost all specifications, fathers&amp;rsquo; earnings, employment, and job changes show no statistically significant effects of the caregiving shock and no significant interactions with workplace characteristics (Appendix Tables A4 and A6). One exception: in Table A4, the interaction between the cancer shock and working at a firm with above-median work hours is negative and significant at the 5% level for fathers, suggesting that fathers who work in high-hours firms do experience some earnings reduction — consistent with them reducing hours in an environment that penalizes deviations from long hours. However, the authors note the effect is substantially smaller relative to baseline earnings than the corresponding maternal effect. The broader pattern implies that workplace flexibility does not appear to be the binding constraint preventing fathers from taking on more caregiving; social norms and intra-household bargaining are posited as more important.&lt;/p&gt;
&lt;h3 id="q14-what-are-the-data-limitations-and-caveats"&gt;Q14. What are the data limitations and caveats?&lt;/h3&gt;
&lt;p&gt;First, work hours at the firm and occupation levels are constructed from the 2011 Census, which is a single cross-section; work hour norms may have shifted between 2011 and the 2012–2023 sample period. Occupation and industry codes also come from the 2011 Census, so parents who changed occupation between 2011 and their baseline year may be misclassified. Second, employment status is inferred from positive ATO earnings in a financial year, a coarser measure than actual employment spells. Third, the sample is restricted to firms with at least 10 employees, which excludes small-firm workers. Fourth, the analysis uses dollar earnings levels, not log earnings, which means baseline wage differences across workplace types can affect the interpretation of absolute dollar results (though the authors show percentage effects are also larger in less-supportive workplaces). Fifth, the study identifies workplace correlates of smaller penalties but does not isolate the causal effect of any specific policy. Sixth, the paper covers only cancer requiring chemotherapy — typically more intensive cancers — so results may overstate average caregiving-shock effects.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Caregiving shock&lt;/strong&gt;: In this paper, a sudden, largely unanticipated increase in caregiving demands on parents triggered by a child&amp;rsquo;s initiation of chemotherapy. Distinguished from the chronic caregiving burden of childbirth; specifically refers to health events that arrive later in childhood and impose large, time-intensive care requirements.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Later-treated dynamic DiD&lt;/strong&gt;: The paper&amp;rsquo;s identification design, following Fadlon and Nielsen (2019, 2021), in which the control group consists of parents who will receive the same treatment (child&amp;rsquo;s cancer diagnosis) at a later date. The control group&amp;rsquo;s placebo treatment year is set six years before their actual treatment, enabling estimation of time-path effects relative to diagnosis while accounting for pre-existing differences via individual fixed effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Work hour intensity&lt;/strong&gt;: Median weekly hours worked by employees at a given firm or in a given occupation (from the 2011 Census), used as a proxy for &amp;lsquo;greedy job&amp;rsquo; characteristics — workplaces that reward continuous long-hours presence and penalize deviations. High work hour intensity captures both above-full-time norms and the likely presence of evening and weekend work requirements.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Female representation in the top 20% of earners&lt;/strong&gt;: A binary indicator equal to one when women are the majority (above 50%) of workers in the top quintile of earnings at a given firm or occupation. The paper distinguishes this from female representation in middle and lower earnings tiers to isolate the effect of women&amp;rsquo;s presence in positions with organizational power to influence workplace policies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Supportive job&lt;/strong&gt;: As defined operationally in this paper: a job in which the worker&amp;rsquo;s firm and occupation both have below-median work hour intensity, both have majority female representation in the top 20% of earners, and the occupation has a below-average gender pay gap. Mothers in supportive jobs suffer smaller and shorter-lived earnings penalties following a caregiving shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Greedy occupation&lt;/strong&gt;: Borrowed from Goldin (2014), and used in this paper to describe occupations that disproportionately reward workers who supply long, often inflexible, hours. In the paper&amp;rsquo;s empirical framework, these are occupations with above-median work hour intensity, which are shown to amplify maternal earnings losses after a caregiving shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Caregiving intensity&lt;/strong&gt;: The time-varying burden of care associated with a child&amp;rsquo;s illness, proxied in this paper by the volume of child healthcare service utilization (Medicare items: GP visits, specialist consultations, diagnostic imaging, prescriptions). Caregiving intensity peaks at year 0 (treatment initiation), declines significantly by year 2, and returns to baseline by year 3 — yet maternal earnings penalties persist beyond this return to baseline.&lt;/p&gt;
&lt;!-- flags: Employment figures cited in the text (4.9 pp in year 0; peak of 5.6 pp in year 2) differ slightly from Table A3 values (-0.045 = 4.5 pp in year 0; -0.050 = 5.0 pp in year 2). This is a within-paper discrepancy in the IZA working paper version. Layer 1 reports the text-stated figures as authored. --&gt;</description></item><item><title>Carbon Pricing and Inequality: A Normative Perspective</title><link>https://macropaperwarehouse.com/papers/carbon-pricing-and-inequality-a-normative-perspective/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/carbon-pricing-and-inequality-a-normative-perspective/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper quantifies the sources and distributional consequences of unexpected carbon price changes for European households using a money-metric welfare framework. The motivation is stark: while carbon taxes enjoy broad support among economists, they face persistent public opposition — exemplified by Australia&amp;rsquo;s 2014 repeal, France&amp;rsquo;s 2018 Yellow Vest protests, and the 2025 rollback of Canada&amp;rsquo;s consumer carbon tax. The authors ask whether average welfare losses are unusually large, and whether the burden falls disproportionately on vulnerable groups, both questions with direct implications for understanding and reducing political resistance.&lt;/p&gt;
&lt;p&gt;The empirical approach rests on the &amp;ldquo;feasible set approach&amp;rdquo; of Del Canto et al. (2025), which applies the Envelope Theorem to show that the first-order welfare impact of a shock on any household is fully summarized by how the shock changes the discounted present value of their future budget sets — through consumption-basket prices, labor income, financial wealth (asset prices and dividends), and government transfers. This money-metric welfare change is preference-free up to first order: behavioral responses drop out, and the measure is independent of specific utility-function assumptions. The framework is appropriate for policy shocks (supply-side) but not for preference shocks.&lt;/p&gt;
&lt;p&gt;The geographic focus is euro-area countries (excluding the Netherlands and Austria due to data gaps) over 1999–2019. The identification strategy follows Känzig (2023): high-frequency shifts in EU ETS carbon futures prices around regulatory events affecting allowance supply are used as instruments in an external-instruments VAR to isolate plausibly exogenous carbon policy shocks. These shocks are then projected onto a wide array of household-level outcomes using local projections (Jordà 2005). The normalization throughout is a 1% increase in the HICP energy component on impact, which corresponds to roughly a 2.5-euro (or about 20%) increase in EU ETS carbon prices. Cross-sectional household budget data come from three Eurostat/ECB surveys: the Household Budget Survey (HBS, 2015 wave) for consumption baskets, EU-SILC (from 2004) for labor and transfer income by demographic group, and the Household Finance and Consumption Survey (HFCS) for household portfolio positions. Demographics are grouped by four age brackets (25–34, 35–49, 50–64, 65+), two education levels (college vs. non-college), three income brackets (bottom quartile = low, middle 50% = mid, top quartile = high), and four geographic regions (Southern, Western, Northern, Eastern Europe).&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. First, aggregate welfare losses are large: a 1% carbon-policy-induced energy price increase causes an average welfare loss of approximately 250 euros, corresponding to about 0.5% of a household&amp;rsquo;s three-year consumption (68% confidence band: 0.06% to 0.94%). Second, decomposing by channel, the direct consumption-price effect accounts for 0.19% of three-year consumption (68% CI: 0.02% to 0.35%); the labor income channel for 0.43% (68% CI: –0.08% to 0.93%); the portfolio channel for –0.04% (a welfare gain; 68% CI: –0.10% to 0.01%); and the transfer income channel for –0.07% (a welfare gain; 68% CI: –0.15% to 0.02%). Labor income is thus the dominant driver — both in aggregate and in the distributional patterns.&lt;/p&gt;
&lt;p&gt;Third, distributional heterogeneity is pervasive and statistically significant (joint F-tests reject uniformity with p-value = 0.00 across all demographic groupings). Non-college-educated households bear welfare losses of roughly 0.6% of three-year consumption, versus roughly 0.3% for college graduates — a gap concentrated in the labor income channel, not the consumption channel (which is broadly similar across groups at around 0.2%). By income, the pattern is U-shaped: young, low-income households suffer the largest losses, exceeding 1% of three-year consumption, while middle-income and older households are the most insulated; high-income households also experience significant losses (around the 0.5% average), driven by their own labor income exposure. Households aged 65 and over suffer welfare losses of only around 0.15%, largely because they are retired from the labor market.&lt;/p&gt;
&lt;p&gt;Fourth, regional heterogeneity is stark. Southern Europe bears the highest burden, with welfare losses of 0.5% to 0.8% for working-age households; Eastern Europe also faces substantial losses; Western Europe stands at around 0.2% to 0.3%; Northern Europe is the most insulated, with losses below 0.2% and not statistically significant. The labor income channel is the primary driver of these regional differences, consistent with more rigid labor markets in Southern and Eastern Europe (stronger employment protection, less flexible wage-setting). Northern Europe is protected partly by its high share of renewable energy, which mutes the carbon-price pass-through. Eastern Europe benefited from disproportionate free ETS allowance allocations over the sample period, dampening direct price impacts.&lt;/p&gt;
&lt;p&gt;These results collectively suggest that public opposition to carbon taxes may stem from legitimate distributional concerns rather than mere ideological resistance or ignorance. The authors conclude with three policy implications: (1) compensation schemes focused only on consumption prices will be insufficient because the dominant channel is labor income; (2) expansionary (green) monetary policy could ease the income burden, though at some inflationary cost; and (3) redistribution should run from older to younger households, since working-age groups bear the disproportionate burden while retirees are largely insulated.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-carbon-policy-shock-and-what-are-the-main-threats-to-identification"&gt;Q1. What is the identification strategy for the carbon policy shock, and what are the main threats to identification?&lt;/h3&gt;
&lt;p&gt;The instrument is the high-frequency shift in EU ETS carbon futures prices around regulatory events affecting allowance supply (following Känzig 2023). The logic is that economic conditions are already priced in prior to the regulatory news, so futures-price movements in a tight window around those events reflect only policy surprises. This instrument is then used in an external-instruments VAR to identify a monthly structural carbon policy shock series (1999–2019). The local projections use 6 lags for monthly outcomes and 2 lags for quarterly outcomes, plus a linear trend and a dummy for the euro sovereign debt crisis (July 2011–March 2012). The main identification threats are: (a) if economic conditions are not fully priced into carbon futures before the regulatory events, the instrument could be correlated with macroeconomic conditions; (b) the framework assumes no preference shocks, which rules out COVID-style demand shifts; (c) the small-noise approximation underlying the feasible-set approach is less suitable for large aggregate shocks.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-feasible-set-approach-not-require-specific-preference-assumptions-and-what-are-its-limitations"&gt;Q2. Why does the feasible-set approach not require specific preference assumptions, and what are its limitations?&lt;/h3&gt;
&lt;p&gt;By the Envelope Theorem applied to household optimization, first-order welfare effects depend only on how the policy changes the prices and quantities in the household&amp;rsquo;s budget constraint — not on how preferences are shaped. Behavioral responses drop out at first order. The welfare metric is money-metric: the willingness-to-pay to avoid the shock, expressed in euros (income units). Limitations: (1) It is a small-noise approximation around a zero-risk limit; large aggregate shocks are not well-handled. (2) It is valid for shocks from the production or policy side but not for preference shocks (e.g., discount rate changes). (3) Accounting properly for idiosyncratic risk requires covariance weights (Theta terms in Proposition 1 of the appendix); Del Canto et al. (2025) estimate these at –0.1 to –0.4, implying somewhat attenuated welfare levels but no meaningful change to the distributional comparisons. (4) Carbon emissions-reduction benefits are excluded from the welfare calculation by design, since the paper focuses on the pecuniary costs side only.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-mechanism-behind-the-labor-income-channel-and-how-does-it-vary-across-demographic-groups-and-regions"&gt;Q3. What is the mechanism behind the labor income channel, and how does it vary across demographic groups and regions?&lt;/h3&gt;
&lt;p&gt;Carbon price increases raise production costs for energy-intensive sectors, reduce output and employment, and depress aggregate wages — a general equilibrium effect that transmits to household labor income over multiple quarters. The average labor income response peaks at around 1% below trend. For non-college-educated households the peak fall exceeds 1%, while for college graduates the response is more muted. By income group, low-income households face the sharpest falls — around 2–4% over the three-year horizon — whereas middle-income households fall by approximately 0.5–1% and high-income households by about 1%. These effects are larger than those estimated by Del Canto et al. (2025) for oil price shocks on US households (approximately 0.3% welfare loss from labor income after a 10% oil price increase), which the authors attribute to more rigid European labor markets: strong employment protection limits wage cuts but discourages hiring and prolongs unemployment spells, amplifying extensive-margin adjustments. In Southern and Eastern Europe, rigidities are most pronounced, generating the largest regional labor-income responses. Northern and Western Europe show more muted responses.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-role-of-the-portfolio-channel-and-who-gains-or-loses-through-it"&gt;Q4. What is the role of the portfolio channel, and who gains or loses through it?&lt;/h3&gt;
&lt;p&gt;Stock prices fall by a peak of about 5% and dividends decline by about 3% after a carbon policy shock. Bond prices initially decline then partially recover. House prices decline substantially but with a lag. The welfare effect of asset price changes depends on whether a household is a net buyer or net seller of the asset. Younger households in the accumulation phase gain from falling asset prices (they can buy cheaply); older households planning to dis-save lose. The portfolio channel is quantitatively modest: average welfare gain of about 0.04%, most pronounced for younger college-educated households. The channel is not large enough to offset labor income or consumption-price losses for any group.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-role-of-the-transfer-income-channel-and-which-groups-benefit-most"&gt;Q5. What is the role of the transfer income channel, and which groups benefit most?&lt;/h3&gt;
&lt;p&gt;Transfer income — which the paper splits into inflation-indexed pension income and other government transfers (unemployment, sickness, disability, education benefits) — generates a welfare gain of about 0.07% on average. Pensions are indexed to inflation and rise as carbon pricing lifts headline prices; this benefit accrues primarily to older households (aged 65+), who have large pension income. Other transfers show an increase post-shock but the responses are generally not statistically significant at conventional levels. High-income households show a negative transfer response. Northern and Southern Europe benefit more from the transfer channel, consistent with more generous welfare programs; Eastern Europe shows little or negative transfer response, consistent with weaker automatic stabilizers.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-u-shaped-pattern-of-welfare-losses-by-income-and-what-explains-it"&gt;Q6. What is the U-shaped pattern of welfare losses by income, and what explains it?&lt;/h3&gt;
&lt;p&gt;The paper finds that low-income and young households suffer the largest losses (exceeding 1% of three-year consumption), middle-income and older households are most insulated, and high-income households also face significant losses (broadly around the 0.5% average). The U-shape arises from the labor income channel: low-income households are concentrated in sectors and employment types most exposed to carbon pricing contractions; high-income households also have substantial labor income (in absolute terms) that contracts; middle-income households appear more buffered, possibly due to sector composition or greater employment stability. The consumption channel contributes approximately uniformly across income groups (around 0.2%), so does not generate the U-shape.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-differ-methodologically-from-prior-distributional-studies-of-carbon-taxes"&gt;Q7. How does this paper differ methodologically from prior distributional studies of carbon taxes?&lt;/h3&gt;
&lt;p&gt;Prior work such as Andersson and Atkinson (2020) and Beznoska et al. (2012) focused on direct consumption-price incidence, following Poterba (1989) and using static input-output methods or cross-sectional spending data to estimate first-round price effects. The present paper differs in three ways: (1) it instruments for unexpected carbon price shocks, isolating exogenous variation; (2) it incorporates indirect channels — labor income, asset prices, and transfers — in addition to direct consumption prices; (3) it estimates dynamic IRFs directly, capturing the persistence of effects over a three-year horizon. The key novel finding is that indirect labor income effects are the dominant driver of both the level and the distribution of welfare losses, and that neglecting these indirect channels substantially understates both the size and the regressiveness of carbon pricing.&lt;/p&gt;
&lt;h3 id="q8-why-are-regional-differences-in-welfare-loss-so-large-and-what-drives-northern-europes-relative-insulation"&gt;Q8. Why are regional differences in welfare loss so large, and what drives Northern Europe&amp;rsquo;s relative insulation?&lt;/h3&gt;
&lt;p&gt;Regional differences are driven primarily by differential pass-through from carbon prices to consumer prices and by differential labor market rigidity. Northern Europe sources a large share of energy from renewables, so a carbon price increase has a smaller pass-through to domestic energy costs. Eastern Europe was allocated disproportionate free ETS allowances over the 1999–2019 sample period, also dampening direct price impacts — consistent with Känzig and Konradt (2024). Southern and Eastern Europe have more rigid labor markets (stronger employment protection, less flexible wage-setting), amplifying the labor-income contraction. Northern and Western Europe have more flexible labor markets. Additionally, Northern and Southern Europe have more generous welfare programs that partially cushion losses via the transfer channel; Eastern Europe lacks this buffer.&lt;/p&gt;
&lt;h3 id="q9-what-data-sources-does-the-paper-combine-and-what-are-the-key-sample-restrictions"&gt;Q9. What data sources does the paper combine, and what are the key sample restrictions?&lt;/h3&gt;
&lt;p&gt;The paper combines three Eurostat/ECB household surveys: (1) the Household Budget Survey (HBS), 2015 wave, for consumption basket shares by COICOP categories for demographic groups; (2) EU-SILC (2004 onward for some countries, 2005 for most) for annual labor income and transfer income time series by group, converted to quarterly frequency via Chow-Lin interpolation; (3) HFCS (conducted every 4 years by the ECB) for household portfolio positions. Time-series macro data on HICP components, house prices, bond prices, stock prices, and dividends come from Eurostat and ECB/Bloomberg. The sample covers euro-area countries (excluding Netherlands and Austria for data reasons) over 1999–2019. Households are restricted to ages 25–75; top and bottom 1% by net worth are excluded from portfolio statistics. The base year for all life-cycle variables is 2015.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Three main implications are drawn: (1) Public resistance to carbon taxes is not merely ideological — the estimated welfare losses are sizable (about 0.5% of three-year consumption for a 1% energy-price increase), so opposition reflects genuine economic concerns. (2) Standard compensation via energy-bill rebates or consumption-basket adjustments is insufficient because the dominant channel is labor income (0.43% vs. 0.19% for consumption). Compensation schemes should include labor-market policies; the authors also suggest expansionary (green) monetary policy as a tool to ease the income burden, though at some inflationary cost. (3) The intergenerational dimension is important: working-age households (especially young, less-educated, lower-income ones) bear the brunt while retirees are largely shielded. Redistribution should run from old to young, not just from rich to poor. Scope conditions: the estimates are derived from the EU ETS context (European carbon market, euro area, 1999–2019), rely on a small-shock linear approximation, and focus on short-to-medium-run impacts (three-year horizon). The benefits of reduced carbon emissions are excluded from the welfare calculation.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-handle-inference-given-the-short-time-series-and-estimation-uncertainty"&gt;Q11. How does the paper handle inference given the short time series and estimation uncertainty?&lt;/h3&gt;
&lt;p&gt;The sample runs from 1999 to 2019, which is relatively short for the IRF exercises. The paper reports 68% and 90% confidence bands throughout (rather than the conventional 95%), using the lag-augmentation approach of Montiel Olea and Plagborg-Møller (2021) to account for serial correlation. For the money-metric welfare calculations, inference uses a parametric bootstrap that draws from the estimated distribution of IRFs (assuming block-wise uncorrelatedness across variables, justified by low cross-residual correlations averaging 0.16). Cross-sectional group shares are treated as given. The authors explicitly acknowledge considerable uncertainty: the 68% confidence band on the aggregate welfare loss spans 0.06% to 0.94%. They conduct joint F-tests for homogeneity of welfare effects across demographic groups; in all cases the null is rejected with p-value = 0.00. Only 68% bands are reported for welfare calculations given short sample and estimation uncertainty.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-heterogeneous-labor-income-irf-magnitudes-for-different-groups-and-are-they-statistically-significant"&gt;Q12. What are the heterogeneous labor income IRF magnitudes for different groups, and are they statistically significant?&lt;/h3&gt;
&lt;p&gt;Average labor income falls by about 1% at the peak (imprecisely estimated). Non-college-educated peak fall exceeds 1%; college-educated peak fall is more muted. By income group: low-income households see falls of roughly 2–4% over three years; high-income households see a fall of about 1%; middle-income households fall by approximately 0.5–1%. These effects are noted to be larger than analogous results for oil shocks in the US (Del Canto et al. 2025), attributed to European labor market rigidity. The responses are described as featuring &amp;lsquo;a considerable degree of persistence but only imprecisely estimated&amp;rsquo; at the average level. The welfare calculations based on these IRFs have wide confidence bands, reflecting this imprecision.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-consumer-price-dynamics-following-a-carbon-policy-shock"&gt;Q13. What are the consumer price dynamics following a carbon policy shock?&lt;/h3&gt;
&lt;p&gt;Energy prices (HICP energy component) rise by 1% on impact and remain elevated for approximately one year before returning toward baseline. Housing and utilities experience a significant, persistent increase, remaining approximately 0.5% above baseline three years after the shock. Transport prices increase by 0.5% on impact but revert within a year. Food prices rise to a lesser extent. Restaurants and hotels, recreation and culture, and clothing also show significant impact-period increases, though most effects become insignificant after 12 months. Two exceptions at 12 months: housing and utilities remain significantly elevated; education and communication prices actually fall, possibly reflecting adverse general-equilibrium wage and employment effects.&lt;/p&gt;
&lt;h3 id="q14-how-is-the-welfare-analysis-limited-to-short-to-medium-run-effects-and-what-longer-run-effects-are-left-unaddressed"&gt;Q14. How is the welfare analysis limited to short-to-medium-run effects, and what longer-run effects are left unaddressed?&lt;/h3&gt;
&lt;p&gt;The welfare calculations are restricted to a three-year horizon because statistical power in the local projections declines beyond that point given the available sample (1999–2019). The paper explicitly notes that the estimates may miss unemployment hazard effects (i.e., transitions into and out of employment), borrowing cost effects induced by carbon taxes, and any long-run structural adjustments (sectoral reallocation, green investment, capital formation). The benefits of reduced carbon emissions — which may be very large in welfare terms but are realized over much longer horizons — are also excluded by design.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Feasible Set Approach&lt;/strong&gt;: A welfare-measurement methodology (from Del Canto et al. 2025) that applies the Envelope Theorem to show that the first-order welfare impact of any shock on a household equals the change in the discounted present value of that household&amp;rsquo;s budget set — encompassing consumption prices, labor income, asset income, and transfers. The measure is preference-free at first order and is expressed in money-metric (income-equivalent) units.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Money-Metric Welfare Loss&lt;/strong&gt;: In this paper, the number of euros a household would be willing to pay to avoid exposure to the carbon policy shock, computed as a share of total three-year consumption. It is derived from the feasible-set formula and expressed in income units, making it directly interpretable and comparable across demographic groups without requiring preference parameters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Carbon Policy Shock&lt;/strong&gt;: An exogenous, unexpected change in carbon prices driven by regulatory events affecting the supply of EU ETS emission allowances, identified via high-frequency shifts in carbon futures prices around those events used as instruments in an external-instruments VAR. Distinguished from demand-driven carbon price fluctuations correlated with the business cycle.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor Income Channel&lt;/strong&gt;: The indirect welfare effect of a carbon price shock that operates through general-equilibrium changes in aggregate wages and employment. It is the dominant welfare channel in the paper (0.43% of three-year consumption on average, versus 0.19% for direct consumption-price effects), and the primary driver of both the aggregate welfare loss and the distributional heterogeneity across education, income, and regional groups.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption Channel (Direct Effect)&lt;/strong&gt;: The welfare impact arising from higher prices for goods in the household&amp;rsquo;s consumption basket following a carbon price increase. Weighted by the household&amp;rsquo;s nominal expenditure on each good. Broadly similar across demographic groups (clustering around 0.2% of three-year consumption), so it does not generate the observed distributional heterogeneity — in contrast to the labor income channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Portfolio Channel&lt;/strong&gt;: The welfare effect transmitted through changes in asset prices (equities, bonds, housing) after a carbon shock. The sign depends on whether a household is a net buyer or net seller of the asset: younger households in the accumulation phase gain from falling asset prices; older households in the dis-saving phase lose. Quantitatively small on average (net welfare gain of about 0.04%), most pronounced for younger, college-educated households.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transfer Channel&lt;/strong&gt;: The welfare effect operating through government transfer income (unemployment and other social benefits) and inflation-indexed pension payments. Because pensions are indexed to the price level, carbon-induced inflation raises pension income and benefits older households. Other transfer income tends to rise post-shock but the responses are generally imprecisely estimated. On average the channel generates a modest welfare gain (about 0.07% of three-year consumption), primarily for the elderly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Greenflation&lt;/strong&gt;: The phenomenon, documented empirically by Bettarelli et al. (2025) and referenced in this paper, whereby carbon-tax shocks contribute to broader consumer price inflation beyond the direct energy-price impact — through pass-through to housing, transport, food, and other categories, and by raising inflation expectations and triggering tighter monetary policy, which in turn depresses bond and house prices.&lt;/p&gt;</description></item><item><title>Diet, Economic Development and Climate Change</title><link>https://macropaperwarehouse.com/papers/diet-economic-development-and-climate-change/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/diet-economic-development-and-climate-change/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Food production accounts for roughly one-third of global greenhouse gas (GHG) emissions, and richer nations contribute disproportionately through meat-intensive diets and input-intensive farming. This paper asks how much of that disparity will be exported to the developing world as it grows, and which policies can most cost-effectively reduce agricultural emissions during that transition. The answer requires separately identifying two distinct channels—demand-side dietary change and supply-side technological change—and tracing their general equilibrium consequences through global food markets.&lt;/p&gt;
&lt;p&gt;The authors build a quantitative multi-country general equilibrium model calibrated to 90 countries (plus a rest-of-world aggregate) and 47 food products for 2010. The demand side features nested non-homothetic CES preferences, which allow income elasticities to differ across food products—the core mechanism of the nutrition transition. The supply side, built on Farrokhi and Pellegrina (2023), operates at a granular grid-cell level covering the Earth&amp;rsquo;s surface, with producers on each plot choosing both which crop to grow and whether to use a modern, input-intensive (higher-GHG) technology or a traditional, labor-intensive one—the core mechanism of agricultural modernization. GHG emissions are tracked from both production and transportation. Data on calorie intake come from FAO Food Balance Sheets; emissions from Poore and Nemecek (2018) and EDGAR-FOOD; yields from FAO-GAEZ (approximately 1.1 million fields).&lt;/p&gt;
&lt;p&gt;A key methodological contribution is an identification result for income elasticities that requires no price data. In open-economy models, trade shares provide a sufficient statistic for consumer prices, so the model&amp;rsquo;s implicit Marshallian demand equations can be estimated using only expenditure shares and bilateral trade flows—a cleaner identification than prior closed-economy approaches. Structural elasticity estimates are validated against reduced-form regressions that regress product-level log absorption on log GDP per capita interacted with the product&amp;rsquo;s GHG intensity; the cross-method correlation has a slope of 0.64–0.77 and R² of 0.93–0.95.&lt;/p&gt;
&lt;p&gt;Four empirical patterns motivate the model. First, diet composition alone drives large variation in emissions: if the whole world adopted the US diet (holding total calories fixed), the food share of global GHG emissions would rise from 30% to 42%; adopting the Argentinian diet would raise it to 74%; adopting the Ethiopian diet would lower it to 12%. Second, GHG emissions per capita from food rise strongly with GDP per capita (elasticity 0.39 in the cross-section); about one-third of this is a pure scale effect (more calories) and two-thirds is a compositional shift toward higher-emission foods (elasticity of emissions per calorie with respect to GDP per capita is 0.23–0.28). Third, products with higher GHG emissions per calorie have higher income elasticities; a 1% rise in a product&amp;rsquo;s GHG intensity is associated with a 0.17–0.21% higher income elasticity, robust to excluding all meat products. Fourth, emissions from fertilizers and energy use as a share of total agricultural emissions rise with GDP per capita (slope 0.82), indicating that agricultural modernization independently amplifies GHG emissions within each crop.&lt;/p&gt;
&lt;p&gt;Model decompositions reveal that about two-thirds of the cross-sectional correlation between food emissions per capita and GDP per capita is attributable to intrinsic dietary preferences (culture, religion, demographics) rather than to income itself, and about one-half of the correlation for emissions per calorie. This implies that the causal effect of economic growth on emissions is substantially smaller than raw correlations suggest.&lt;/p&gt;
&lt;p&gt;Policy counterfactuals (Table 4) are the paper&amp;rsquo;s centerpiece. A uniform 10% TFP shock across all modern agricultural, non-agricultural, and input producers raises global welfare by 14.9% and increases global agricultural GHG emissions by 5.0% (approximately 0.6 Gt CO₂ from production, 0.004 Gt from transport). Shutting down the nutrition transition channel reduces this emission increase by 28%; shutting down agricultural modernization reduces it by a further 16%; shutting both down reduces it by 42%—so the two mechanisms together account for more than one-third of the growth-induced emission increase. Crucially, ignoring general equilibrium supply responses would overstate the emission impact of economic growth by 100%: higher food demand raises production prices, which dampens both consumption growth and further technology adoption.&lt;/p&gt;
&lt;p&gt;For dietary restrictions: a global no-beef mandate would reduce agricultural GHG emissions by 20%, at a global welfare cost of 0.6%, with large concentrated losses in major beef-producing and consuming countries (Argentina −3–5%; Uruguay −4%). A global vegetarian mandate would reduce emissions by 30% (approximately the same 20% figure is given in the abstract with apparent inconsistency but Table 4 column 3 shows −20% for no-beef and −30% for vegetarian), at a welfare cost of 2.8% globally and with greater inequality impacts for developing countries. Back-of-the-envelope calculations that ignore general equilibrium overstate the emission reductions from dietary restrictions by roughly one-third.&lt;/p&gt;
&lt;p&gt;For food trade policy: raising trade costs enough to cut transportation emissions by 75% reduces total agricultural GHG emissions by 11.9%, but at a global welfare cost of 17.8%—a ratio far worse than dietary policies. The welfare loss is highly unequal: countries in the bottom quartile of the GDP per capita distribution face welfare losses of up to 41% (the abstract states this figure; Table 4 col. 2 shows the Q4/Q1 inequality worsening by 4.9 percentage points in the eat-local scenario). The conclusion is that dietary policies dominate food trade policies on both effectiveness and equity grounds.&lt;/p&gt;
&lt;p&gt;Transportation emissions account for only about 5% of agricultural GHG (0.7 Gt CO₂ vs. 16.5 Gt from production), so policies targeting transport emissions alone have limited aggregate impact.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-for-income-elasticities-and-why-is-it-novel"&gt;Q1. What is the core identification strategy for income elasticities, and why is it novel?&lt;/h3&gt;
&lt;p&gt;Standard non-homothetic CES estimation requires price data because the demand equation depends on price indices. In a closed economy this problem is severe. The authors show that in an open economy, bilateral trade shares provide a sufficient statistic for variety price indices: averaging trade shares across a country&amp;rsquo;s import partners yields a geometric mean of production prices that can be differenced out using fixed effects. The key estimating equation (40) regresses an adjusted expenditure share on log income per capita, with fixed effects absorbing production-price variation through the set of import partners. No price data is needed. This is exact—not an approximation—unlike the approximate methods in Comin et al. (2021) or Caron and Fally (2022), which either impose additional assumptions about price variation across consumer groups or require proxies for crop-specific trade costs such as gravity variables.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-threats-to-identification-and-how-are-they-addressed"&gt;Q2. What are the main threats to identification and how are they addressed?&lt;/h3&gt;
&lt;p&gt;The key concern is that income is correlated with prices and preference shifters that also affect food expenditure shares. In the reduced-form regressions (equation 1), country-year and product-year fixed effects control for country-specific factors (including regional technology change) and global product-specific factors (including product-specific technological progress). In the structural estimation (equation 40), the model&amp;rsquo;s functional form is used to control fully for endogeneity arising through prices, since trade shares substitute out unobservable price indices exactly. The close agreement between reduced-form and structural income elasticity estimates (slope 0.64–0.77, R² 0.93–0.95 in cross-validation) is reassuring that the two quite different identifying assumptions yield similar results. One remaining concern is unobservable preference shifters (ai,k and ã_i,s), which appear as residuals; identification requires income variation orthogonal to these shifters, and the authors follow the precedent of assuming fixed effects are sufficient. Household-level data from Brazil&amp;rsquo;s Consumer Expenditure Survey (POF) bolster the reduced-form patterns using within-country income variation.&lt;/p&gt;
&lt;h3 id="q3-how-are-the-nutrition-transition-and-agricultural-modernization-distinguished-empirically-and-in-the-model"&gt;Q3. How are the nutrition transition and agricultural modernization distinguished empirically and in the model?&lt;/h3&gt;
&lt;p&gt;These are fundamentally different economic mechanisms. The nutrition transition operates through demand: as incomes rise, consumers shift toward food products that, for reasons of taste or nutrition, happen to have higher GHG emissions per calorie. It is a between-product phenomenon captured by non-homothetic income elasticities. Agricultural modernization operates through supply: as wages rise, producers substitute away from labor-intensive traditional technologies toward input-intensive modern technologies (fertilizers, machinery) that emit more GHG per calorie of output, for any given crop. It is a within-product phenomenon captured by the endogenous technology-choice margin in the agricultural production model. In the counterfactual decompositions, the authors shut down each channel independently: the nutrition transition is shut down by setting all within-sector income elasticity parameters (ε_k) equal; agricultural modernization is shut down by fixing the land share in each technology exogenously. Doing so reveals that the nutrition transition accounts for 28% and modernization for 16% of the emission increase from a 10% TFP shock (jointly 42%), with the remainder attributable to scale effects and general equilibrium price responses.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-role-of-general-equilibrium-supply-responses-and-why-do-they-matter-so-much"&gt;Q4. What is the role of general equilibrium supply responses and why do they matter so much?&lt;/h3&gt;
&lt;p&gt;A central finding is that ignoring supply-side equilibrium price responses would overstate the emission impact of economic growth by 100%. The mechanism is straightforward: economic growth raises income and thus food demand, which pushes up production prices (because agricultural supply is upward-sloping due to limited land and heterogeneous productivity across grid cells). Higher prices dampen consumption, which partially offsets the demand-driven emission increase. For dietary restriction policies, back-of-the-envelope calculations that simply remove the GHG attributable to banned food products overstate the emission reduction by roughly one-third, because consumers substitute toward other food products and global agricultural production reorganizes. The model&amp;rsquo;s general equilibrium structure is therefore essential for obtaining credible policy counterfactuals, and a main conclusion of the paper is that the literature&amp;rsquo;s existing back-of-the-envelope calculations in environmental science substantially overstate both the emission risks from growth and the emission benefits from dietary policies.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-across-countries-and-products"&gt;Q5. What heterogeneity is documented across countries and products?&lt;/h3&gt;
&lt;p&gt;Across countries: diet composition varies enormously. Counterfactual calculations show that if all countries adopted the Argentinian diet (holding total calories fixed), the global food share of total emissions would rise to 74%; adopting the Ethiopian diet would lower it to 12%, compared to the factual 30%. The income elasticity of the agricultural sector as a whole is 0.39, close to Comin et al. (2021)&amp;rsquo;s 0.37. Rich countries have a higher share of modern technology in production, higher fertilizer and energy use per unit of land, higher food GHG per capita, and higher food GHG per calorie. About two-thirds of the cross-sectional gradient in food GHG per capita is attributable to intrinsic preferences rather than income per se. Religion is documented as one driver: Islamic-majority countries show lower preference for pork; Hindu-majority countries show higher preference for lamb, mutton, and poultry relative to other meats. Across products: GHG emissions per 1,000 kcal range from above 35 kg CO₂ for beef and coffee to below 5 kg CO₂ for wheat and rye. Income elasticity parameters (ε_k) range from lowest for staples (yams, sweet potatoes, millet, sorghum, rice) to highest for luxury fruits and vegetables (berries, asparagus, cucumbers, watermelon). Notably, the income-GHG gradient persists after excluding all meat products: vegetables and fruits have higher GHG per calorie than staples, so the nutrition transition is broader than a simple meat-consumption story.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-diet-restriction-and-food-trade-policy-counterfactuals-compare-on-welfare-and-effectiveness"&gt;Q6. How do the diet restriction and food trade policy counterfactuals compare on welfare and effectiveness?&lt;/h3&gt;
&lt;p&gt;Diet restriction (no-beef): global GHG emissions fall 20%, global welfare falls 0.6%. The welfare effect is highly concentrated—Argentina experiences −3–5% welfare loss, Uruguay approximately −4% in the no-beef scenario, because they are large meat producers and exporters. Inequality between rich (Q4) and poor (Q1) countries worsens by 1.0 percentage point. Diet restriction (vegetarian): global GHG emissions fall 30%, global welfare falls 2.8%. Inequality worsens by 6.0 percentage points, indicating developing countries bear more of the cost because a larger share of their income goes to food, and their income sources (agriculture) are more directly affected. Food trade policy (&amp;rsquo;eat local&amp;rsquo;, raising trade costs to cut transportation emissions by 75%): global GHG emissions fall 11.9%, but global welfare falls 17.8%—roughly 25–30 times the welfare cost per percentage point of emission reduction compared to dietary policies. Inequality worsens substantially more: Q4/Q1 ratio worsens by 4.9 percentage points. Countries in the bottom GDP quartile face welfare losses up to 41%. The paper concludes that dietary restrictions are both substantially more effective in reducing GHG emissions and far more equitable in their welfare consequences than food trade policies.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-share-of-agricultural-ghg-from-transportation-versus-production-and-what-are-the-implications"&gt;Q7. What is the share of agricultural GHG from transportation versus production, and what are the implications?&lt;/h3&gt;
&lt;p&gt;In the 2010 data, GHG emissions from food transportation account for approximately 5% of total agricultural GHG (0.7 Gt CO₂ out of approximately 17.2 Gt total). Production accounts for 95% (16.5 Gt CO₂). This has two implications. First, in the economic growth counterfactual, transportation emissions increase by 2.2%, but because transportation is only 5% of total, its contribution to total emission growth (0.004 Gt) is negligible. Second, it implies that policies targeting food &amp;lsquo;food miles&amp;rsquo; or local eating are poorly targeted: even a dramatic 75% reduction in transportation emissions only mechanically eliminates 4.6% of total agricultural GHG, and the actual general equilibrium reduction (11.9%) comes mostly from production effects (agricultural trade restrictions reduce global production and consumption), accompanied by very large welfare costs.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-and-validation-exercises-are-conducted"&gt;Q8. What robustness checks and validation exercises are conducted?&lt;/h3&gt;
&lt;p&gt;The paper provides several validation exercises. (1) The reduced-form income elasticity regressions are run both with all crops and excluding all meat products (beef, lamb and mutton, pig meat, poultry), yielding nearly identical coefficients of 0.176 and 0.175 (columns 1 and 2 of Table 1), and with country-year and product-year fixed effects (columns 3–4), showing similar results across specifications. (2) The structural income elasticities are compared to the reduced-form estimates, with a cross-method slope of 0.64–0.77 and R² of 0.93–0.95, reassuring given the two methods make different identifying assumptions. (3) Model fit is checked against six untargeted empirical regularities (Figure 6): declining agricultural employment share, rising input cost share, rising modern technology land share, rising food GHG per capita, rising calories per capita, and rising food GHG per calorie—all with GDP per capita. The model matches the sign and approximate magnitude of each relationship. (4) Household-level estimates using Brazil&amp;rsquo;s POF survey replicate the cross-country finding that higher-GHG products have higher income elasticities, controlling for fixed effects, food price proxies, and excluding meat. (5) The decomposition of the cross-sectional income-emissions gradient shows that equalizing comparative advantage (column 3) or trade costs (column 4) across countries leaves the gradient approximately unchanged, supporting the focus on preferences and technology.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-prior-work-and-where-does-it-depart-from-it"&gt;Q9. How does this paper relate to prior work and where does it depart from it?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of several literatures. It builds on Farrokhi and Pellegrina (2023) for the granular grid-cell production model with technology choice; on Costinot, Donaldson, and Smith (2016) for the agricultural field structure; and on Comin, Lashkari, and Mestieri (2021) for non-homothetic CES preferences and the identification of income elasticities. Key departures: (a) Relative to Comin et al. (2021), the authors extend identification to nested CES preferences and to an open-economy without requiring price data—their method is exact rather than approximate. (b) Relative to the environmental science literature (e.g., Hoolohan et al., 2013; Perignon et al., 2017; Tilman et al., 2011), the paper endogenizes general equilibrium supply responses, which the authors show dramatically attenuate the effect of both income growth and dietary policies on emissions. (c) Relative to prior quantitative spatial models of climate change (e.g., Shapiro 2016 on trade costs and CO₂), this paper focuses on agricultural emissions specifically and introduces nutrition transition and technology choice. (d) The authors claim to be the first to analyze both dietary restrictions and food trade policies on agricultural emissions within quantitative trade models. (e) Relative to Chen et al. (2022), who use a computable general equilibrium model with general equilibrium supply adjustments, this paper includes far more food products (47 vs. their smaller set) and endogenizes technology choice, both of which are quantitatively important for capturing the nutrition transition.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-papers-mechanism-for-why-vegetable-and-fruit-consumption-also-raises-ghg-emissions-as-income-rises-even-without-meat"&gt;Q10. What is the paper&amp;rsquo;s mechanism for why vegetable and fruit consumption also raises GHG emissions as income rises, even without meat?&lt;/h3&gt;
&lt;p&gt;The paper notes in footnote 1 that the positive correlation between income elasticities and GHG emissions per calorie persists even when meat products are excluded from the sample (Table 1, columns 3–4). The reason is that vegetables and fruits—which become more preferred as countries grow richer—emit more GHG per calorie than staple foods such as yams and potatoes. Staples require little processing or refrigeration and are typically produced with traditional, low-input technologies. By contrast, fresh fruits and vegetables (especially high-value items such as berries, asparagus, grapes, and coffee) require more energy-intensive transportation, storage, and sometimes greenhouse production. This means that the nutrition transition generates rising emissions not merely through the beef channel emphasized in much of the public debate, but through a broader shift away from calorie-dense staples toward diverse, lower-calorie-density products that happen to have higher GHG footprints per calorie.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-model-imply-about-the-environmental-kuznets-curve-for-food-emissions"&gt;Q11. What does the model imply about the Environmental Kuznets Curve for food emissions?&lt;/h3&gt;
&lt;p&gt;The paper explicitly tests for and finds no evidence of an Environmental Kuznets Curve (EKC) in food emissions—that is, no inverse-U shape in which emissions per capita eventually decline as countries become very rich, as might be expected if wealthy nations adopt more sustainable diets or stricter environmental regulations. The income-emission relationship is found to be approximately log-linear across all levels of development (footnote 8). This is consistent with the broader empirical literature on the EKC (cited survey by Dinda, 2004). The implication is that there is no automatic &amp;lsquo;greening&amp;rsquo; of diets as countries develop; active policy intervention would be needed.&lt;/p&gt;
&lt;h3 id="q12-how-is-economic-development-modeled-in-the-policy-counterfactuals-and-what-are-the-scope-conditions"&gt;Q12. How is economic development modeled in the policy counterfactuals, and what are the scope conditions?&lt;/h3&gt;
&lt;p&gt;Economic development is modeled as a uniform 10% increase in TFP for three types of agents: (i) modern agricultural producers, (ii) non-agricultural producers, and (iii) agricultural input producers (fertilizers, machinery, pesticides). Traditional agricultural technology is not subject to productivity growth, following Gollin, Parente, and Rogerson (2007). This creates both income effects (via higher wages) and substitution effects (via changes in relative input prices that favor modern, input-intensive technology). The scope conditions are important: the results apply specifically to a uniform global TFP shock, not to individual-country development. For individual-country TFP shocks, the analytical decomposition (equation 34) shows that general equilibrium income spillovers to foreign countries can attenuate the nutrition transition if foreign incomes fall (e.g., due to terms-of-trade effects). The model does not incorporate dynamics (it is a static model calibrated to 2010), so it cannot directly speak to transition paths or time horizons for emission convergence.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-welfare-implications-for-developing-countries-under-different-policies-and-why-do-dietary-policies-dominate"&gt;Q13. What are the welfare implications for developing countries under different policies, and why do dietary policies dominate?&lt;/h3&gt;
&lt;p&gt;Under economic growth (10% TFP shock), global welfare rises 14.9% with a modest increase in Q4/Q1 inequality of 0.4 percentage points, indicating relatively even welfare gains. Under no-beef, global welfare falls 0.6% but inequality worsens by 1.0 pp; under vegetarian, welfare falls 2.8% and inequality worsens by 6.0 pp—developing countries lose more because more of their income is spent on food and the agricultural sector is a larger share of their economy. Under eat-local (food trade restrictions), welfare falls 17.8% and the Q4/Q1 ratio worsens by 4.9 pp, with countries in the bottom GDP quartile facing losses up to 41%. The stark dominance of dietary policies over trade policies reflects two structural features: (a) food trade restrictions reduce the gains from comparative advantage in food production, which are particularly large for food-exporting developing countries; and (b) the welfare cost per unit of GHG reduction is far higher for trade policies because they distort production allocation without addressing the underlying demand-side emissions driver.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Nutrition Transition&lt;/strong&gt;: As defined and used in this paper: the demand-side process by which rising income causes consumers to shift their caloric intake away from staple foods (yams, potatoes, rice, millet) toward food products with higher GHG emissions per calorie (meats, fruits, vegetables, coffee). The transition is captured in the model by non-homothetic income elasticity parameters ε_k that are higher for more emissions-intensive products and is operative even after excluding all meat products.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Agricultural Modernization&lt;/strong&gt;: As defined and used in this paper: the supply-side process by which rising wages induce producers to substitute from traditional, labor-intensive agricultural technology (τ=0, no purchased intermediate inputs) toward modern, input-intensive technology (τ=1, fertilizers, machinery, pesticides), which emits more GHG per calorie of output. This operates within each crop and is captured in the model by endogenous technology choice at the plot level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-Homothetic CES Preferences (Nested)&lt;/strong&gt;: A three-tier preference structure in which the expenditure share of a food product k depends on income through a product-specific parameter ε_k that governs how fast the product&amp;rsquo;s preference weight grows with utility. Products with higher ε_k have higher income elasticities; the overall income elasticity of the agricultural sector (0.39 in this paper&amp;rsquo;s calibration) is an expenditure-weighted average of the ε_k values. The nested structure allows the agricultural sector&amp;rsquo;s income elasticity relative to non-agriculture to be determined separately from the income elasticities of individual food products within agriculture.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implicit Marshallian Demand&lt;/strong&gt;: The demand equation derived from non-homothetic CES preferences by substituting out unobservable price indices using a base good, yielding a demand specification that depends on observable expenditure shares and income rather than on prices directly. In this paper&amp;rsquo;s open-economy extension, trade shares further substitute out unobservable variety price indices, making the estimation equation fully price-data-free.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GHG Emission Intensity (per calorie)&lt;/strong&gt;: In this paper: the parameter φ_k (crop-specific) and φ_τ (technology-specific), where φ_kτ = φ_k × φ_τ is the kg CO₂-equivalent emitted per 1,000 kcal of crop k produced under technology τ. This is the key cross-product heterogeneity that, combined with income elasticity heterogeneity, drives the environmental consequences of the nutrition transition. In the data: ranges from below 5 kg CO₂ per 1,000 kcal for wheat and rye to above 35 kg for beef and coffee.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Grid-Cell Production Model&lt;/strong&gt;: A representation of the agricultural supply side in which the Earth&amp;rsquo;s land surface is divided into approximately 1.1 million fields (FAO-GAEZ), each with agro-climatically determined potential yields by crop and technology that are independent of market conditions. Within each field, a continuum of plots is allocated to crops and technologies via Fréchet productivity draws, yielding smooth aggregate supply functions and allowing for realistic specialization patterns and technology gradients across geography.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Back-of-the-Envelope (Demand Mechanism) Benchmark&lt;/strong&gt;: In this paper: a partial-equilibrium counterfactual calculation that takes observed or baseline food demand quantities and simply attributes changes to them from a policy without allowing supply prices, production, or trade flows to adjust. The paper systematically compares model general equilibrium results against this benchmark (column 9 of Table 4) to quantify how much supply-side adjustments matter, finding that the back-of-the-envelope approach overstates the emission impact of economic growth by approximately three times, and overstates the emission reduction from dietary policies by roughly one-third.&lt;/p&gt;</description></item><item><title>Distortions, Producer Dynamics, and Aggregate Productivity: A General Equilibrium Analysis</title><link>https://macropaperwarehouse.com/papers/distortions-producer-dynamics-and-aggregate-productivity-a-general-equilibrium-analysis/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/distortions-producer-dynamics-and-aggregate-productivity-a-general-equilibrium-analysis/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks how institutional distortions to factor markets affect not only the static allocation of inputs across farms but also the dynamic choices — crop selection and productivity-enhancing investment — that determine the long-run distribution of farm productivities and hence aggregate agricultural TFP. The question matters because prior work on misallocation has largely treated the productivity distribution as exogenous; this paper endogenizes it, showing that the dynamic channels can be quantitatively larger than the static factor-misallocation channel.&lt;/p&gt;
&lt;p&gt;The empirical foundation is the Vietnam Access to Resources Household Survey (VARHS), a balanced panel of 2,118 farm households surveyed biennially from 2006 to 2016 across twelve provinces in north and south Vietnam. Vietnam provides a natural laboratory: post-1986 reforms decollectivized agriculture nationally, but deeply divergent pre-reform institutions (collective agriculture in the north for more than three decades; private household farming in the south throughout) produced durable differences in land-market functioning, crop-choice restrictions, and property-rights security. Measured TFP is more than 2.5 times higher in the south than the north (the observed log TFP ratio implies roughly a 2.5-fold level difference). The elasticity of land use with respect to farm TFP is 0.554 in the south versus 0.152 in the north, and the elasticity of labor use is 0.382 versus 0.122 — three to four times larger in the south — indicating far more efficient resource allocation in the south. The share of perennial-crop farmers (high-value cash crops, especially coffee) is 33% in the south and roughly 5% in the north. Average biennial TFP growth is 6.2% in the south versus 2.6% in the north.&lt;/p&gt;
&lt;p&gt;The authors build a dynamic general equilibrium model of heterogeneous farm managers (following Lucas 1978) in which farm productivity has four components: a permanent farmer-specific component, a random transitory component, an endogenous managerial ability component accumulated through investment, and a crop-specific component tied to endogenous crop choice. Institutional distortions are modeled as idiosyncratic revenue taxes correlated with farm productivity (following Restuccia and Rogerson 2008), with the key parameter being the elasticity of distortions with respect to farm productivity (rho). A higher rho means more-productive farms face proportionately larger distortions, which (i) compresses the gap between large and small farms in equilibrium factor use, and (ii) reduces the private return to investing in ability. The model also incorporates government-imposed crop restrictions that force a fraction of farms to grow rice regardless of profitability. The model is calibrated to south Vietnam moments: average TFP growth, dispersion in TFP and growth, the land-size distribution, the measured elasticity of distortions, and crop shares. Measurement error in output and inputs is explicitly modeled following Bils, Klenow, and Ruane (2021); the estimated BKR statistic is 0.906 for the south and 0.987 for the north, indicating relatively limited measurement error by manufacturing-sector standards.&lt;/p&gt;
&lt;p&gt;The main counterfactual imposes north Vietnam distortion parameters on the south-calibrated benchmark economy. Three distortion parameters differ: (1) the distortion elasticity rho rises from 0.79 (south) to 0.91 (north); (2) crop-specific distortions flip sign — in the south perennials face lower effective taxes than rice (phi_perennial = 1.61 &amp;gt; 1), while in the north perennials face higher effective taxes than rice (phi_perennial = 0.68 &amp;lt; 1); (3) the share of farms subject to government-imposed crop restrictions rises from 23% to 43%.&lt;/p&gt;
&lt;p&gt;The counterfactual experiment produces four main quantitative results. First, aggregate TFP falls by 41% relative to the benchmark, accounting for 61% of the observed productivity gap between north and south Vietnam (the observed ratio is 0.42; the counterfactual ratio is 0.59). Second, the average biennial farm TFP growth rate falls by 1.6 percentage points (from 6.23% to 4.60%), accounting for just under half of the observed 3.6 percentage-point north-south gap. Third, TFP dispersion (standard deviation of log TFP) falls by 8 percentage points, more than half of the 14-percentage-point lower dispersion observed in the north. Fourth, the share of perennial farmers collapses from 33% to 9%, closely matching the observed 5% in the north.&lt;/p&gt;
&lt;p&gt;Channel decomposition reveals that static factor misallocation alone reduces output by 19.4% (one-third of the total 40.8% gap, proportionately allocated), while the crop-choice channel reduces output by 8.0% and the farm-ability channel (endogenous investment) reduces output by 31.6%. Together, the dynamic channels (crop choice plus farm ability) account for approximately two-thirds of the total productivity loss, more than doubling the contribution of static misallocation. Among individual distortions, the distortion elasticity rho alone accounts for a 38.3% output reduction, crop-specific distortions account for 7.5%, and government crop restrictions account for only 1.4%. The key mechanism is that a small increase in rho (from 0.79 to 0.91) has large productivity consequences because the productivity cost is convex in rho and accelerates as rho approaches one — at rho = 1, distortions fully absorb all incremental profits from higher ability, eliminating investment incentives entirely.&lt;/p&gt;
&lt;p&gt;The paper shows that measurement error has limited impact on the north-south comparison (since the main experiment is a within-survey, within-country comparison), but substantially inflates the level gains from removing all distortions: removing measurement error from the model more than doubles the estimated gains from moving to a first-best economy, underscoring that measurement error matters most in cross-economy level comparisons.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-the-main-threats-to-validity"&gt;Q1. What is the identification strategy and the main threats to validity?&lt;/h3&gt;
&lt;p&gt;The identification exploits within-country, within-survey variation between north and south Vietnam, which share a common currency, survey instrument, and price measurement methodology. The main threat is that technology and geography differ across regions beyond institutions. The paper addresses this in two ways. First, it restricts comparisons to the two rice-growing delta regions — the Red River Delta (north) and Mekong Delta (south) — where technology and geographic differences are minimal, and shows the same patterns hold: measured distortion elasticity in the Mekong Delta is 0.79 versus 0.94 in the Red River Delta, and growth is higher and productivity more dispersed in the south. Second, the paper uses FAO Global Agro-Ecological Zones data to show land quality differences are negligible between north and south and, if anything, slightly favor the north; when scaled through the production function (land share times span-of-control = 0.35), land quality cannot account for the observed TFP gap. A second threat is measurement error inflating wedge dispersion and the estimated distortion elasticity. The paper addresses this by embedding explicit measurement error in the calibration and by using the Bils-Klenow-Ruane (2021) methodology, finding BKR statistics of 0.91 (south) and 0.99 (north), suggesting measurement error is modest in agriculture relative to manufacturing. The calibrated true distortion elasticity for the south is rho = 0.79, versus a measured elasticity of 0.86, a bias of around 0.06 — consistent with BKR estimates.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-productivity-channels-and-how-is-each-measured"&gt;Q2. What are the three productivity channels and how is each measured?&lt;/h3&gt;
&lt;p&gt;The three channels are (1) static factor misallocation, (2) crop distribution, and (3) farm ability. Each is isolated by a sequential decomposition: for factor misallocation, counterfactual distortions rho and phi are imposed while holding the crop and ability distributions fixed at benchmark-economy values, yielding an output loss of 19.4%. For crop distribution, the crop shares are adjusted to the counterfactual economy while holding within-crop ability distributions fixed at benchmark values; output falls by 8.0%. For farm ability, the ability distribution conditional on crop type is adjusted to the counterfactual while holding crop shares fixed; output falls by 31.6%. The sum (59.0%) exceeds the total gap (40.8%) because of negative interactions among channels — factor misallocation has a smaller bite when the productivity distribution is more compressed, as in the counterfactual.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-role-of-the-distortion-elasticity-parameter-rho-and-why-does-it-generate-outsized-productivity-losses-from-a-small-change"&gt;Q3. What is the role of the distortion elasticity parameter rho and why does it generate outsized productivity losses from a small change?&lt;/h3&gt;
&lt;p&gt;The parameter rho governs the extent to which more productive farms face proportionately larger distortions. At rho = 0, distortions are orthogonal to productivity; at rho = 1, distortions grow one-for-one with productivity, fully taxing away any incremental profit from increasing ability. The investment return to moving up the ability ladder is proportional to the incremental profit gained, which equals (1 - rho) times the increment in revenue. As rho rises toward 1, this return collapses toward zero. Because the South&amp;rsquo;s calibrated rho is already 0.79 — close to 1 on the relevant scale — a further increase to 0.91 is disproportionately large in terms of investment disincentives. The paper demonstrates this asymmetry explicitly in Appendix C.6: a symmetric increase and decrease of rho by 0.1 (set to the observed North-South difference in measured elasticity) reduces output by 42% when rho rises but only 39% when rho falls, driven primarily by the farm-ability channel (27 log points versus 21 log points difference in log output).&lt;/p&gt;
&lt;h3 id="q4-how-do-crop-specific-distortions-and-government-crop-restrictions-work-and-what-is-their-quantitative-contribution"&gt;Q4. How do crop-specific distortions and government crop restrictions work and what is their quantitative contribution?&lt;/h3&gt;
&lt;p&gt;Crop-specific distortions phi_i create wedges that differ across crop types. In the south, phi_perennial = 1.61 (perennial growers face lower effective taxes than rice farmers), while in the north phi_perennial = 0.68 (perennial growers face higher effective taxes). This reversal in relative distortions discourages perennial farming in the north both directly (lower profits) and dynamically (perennial farmers, who tend to be higher-ability, invest less). Unilaterally imposing north crop-specific distortions on the south benchmark reduces output by 7.5%. Government-imposed crop restrictions force a fraction omega of farms to grow rice regardless of profitability, with omega rising from 23% to 43% north-south. This channel has the smallest impact (1.4% output loss) because: (a) a large fraction of restricted farmers would have chosen rice anyway, and (b) back-of-envelope calculation shows the loss amounts to reducing productivity of only about 7% of farmers (the 20 percentage-point change in omega times the 33% perennial share) by about 20% (measured perennial-rice TFP gap).&lt;/p&gt;
&lt;h3 id="q5-what-empirical-evidence-motivates-the-endogenous-investment-mechanism"&gt;Q5. What empirical evidence motivates the endogenous investment mechanism?&lt;/h3&gt;
&lt;p&gt;Table 3 shows that in both north and south Vietnam, farm investment (cash or labor investment in irrigation or soil/water conservation) and extension-service participation are positively correlated with farm TFP and negatively correlated with farm-level distortion wedges, indicating that more distorted farms invest less. In the south, both investment and extension services are significantly positively associated with subsequent TFP growth. In the north, only extension-service participation is positively associated with future growth, while physical investment is not — suggesting the return to investment is suppressed in the north. The data also document a life-cycle profile (Figure 3) in which farm TFP rises steeply for young farms and then levels off, much more sharply in the south than in the north, consistent with faster ability accumulation in the less-distorted south.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-is-documented-across-crop-types-within-each-region"&gt;Q6. What heterogeneity is documented across crop types within each region?&lt;/h3&gt;
&lt;p&gt;In the south, perennial farmers have higher average output (log output 10.6 vs. 9.9 for rice), more land (3.9 acres vs. 2.4), more labor, higher TFP (above mean relative to rice), and far higher biennial TFP growth (10.9% vs. 4.9%). In the north, the pattern reverses: perennial farmers underperform relative to rice farmers in output (-0.583 log points, significant), land, labor, and TFP (-0.413 log points). This reversal occurs because crop-specific distortions disproportionately penalize perennial farming in the north. Despite the average gaps, there is substantial productivity overlap across crop types within both regions (Figure A.1), with many unproductive perennial farmers and productive rice farmers coexisting. This overlap motivates the paper&amp;rsquo;s modeling of crop selection as a utility-cost decision with idiosyncratic taste heterogeneity (Frechet distribution), rather than a pure productivity-cutoff rule.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-conducted-and-what-do-they-show"&gt;Q7. What robustness checks are conducted and what do they show?&lt;/h3&gt;
&lt;p&gt;Four robustness exercises are conducted. First, re-calibrating with fixed quadratic investment-cost curvature (zeta = 2) instead of the estimated 1.74 yields a counterfactual output ratio of 58.8%, similar to the baseline 59.2%. Second, lowering the targeted average growth rate by 2 percentage points (addressing the concern that aggregate TFP growth partly reflects economy-wide technology rather than ability investment) produces a counterfactual output ratio of 58.6% — essentially unchanged. Third, lowering the targeted growth rate by 4 percentage points produces 62.4%, still economically large. Fourth, two model extensions are explored: (a) incorporating a hump-shaped life-cycle profile with a young-to-old transition produces a 43% productivity loss, similar to the 41% baseline; (b) allowing entrants to draw ability from a distribution dependent on the exiting predecessor&amp;rsquo;s ability produces a 57% productivity loss — larger than baseline because investment creates positive spillovers to future entrants.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-the-prior-misallocation-literature"&gt;Q8. How does this paper relate to and differ from the prior misallocation literature?&lt;/h3&gt;
&lt;p&gt;The paper builds on Restuccia and Rogerson (2008) and Hsieh and Klenow (2009), who model static misallocation via idiosyncratic wedges. It contributes three extensions. First, it endogenizes the farm productivity distribution by adding investment and crop choice, so that the same wedges that generate static misallocation also distort dynamics — this doubles the productivity cost. Second, the experiment is a within-country comparison between two regions rather than a comparison against a hypothetical undistorted economy, avoiding the criticism that the undistorted benchmark is unrealistic. The re-calibrated north model accounts for 100% of the observed north-south TFP ratio (40.7% model vs. 42% data). Third, the dynamic model generates falsifiable predictions about farm TFP growth rates, TFP dispersion, and crop distributions — all of which move in the right directions — providing a richer validation test than static models allow. The paper also relates to Hsieh and Klenow (2014), who document faster life-cycle productivity growth in less distorted economies (India and Mexico vs. US), and to Adamopoulos and Restuccia (2020), who study land reform in Vietnam but with exogenous productivity distributions; the current paper finds that endogenizing productivity distributions significantly amplifies the costs of distortions. The measurement-error treatment follows Bils, Klenow, and Ruane (2021) and Adamopoulos et al. (2022).&lt;/p&gt;
&lt;h3 id="q9-what-does-the-model-imply-about-a-hypothetical-undistorted-economy"&gt;Q9. What does the model imply about a hypothetical undistorted economy?&lt;/h3&gt;
&lt;p&gt;Removing all distortions (rho = 0, phi_i = 1 for all crops, omega = 0, sigma_epsilon = 0) increases TFP by a factor of 3.37 relative to the south benchmark (Appendix C.5, Table C.11), meaning the first-best economy is more than three times as productive. Static reallocation gains alone (holding the productivity distribution fixed) account for roughly 70% of this gap. The remaining gains come from the endogenous shift in the ability distribution — in the undistorted economy, lower ability farmers invest less (because higher general equilibrium wages lower profits) but higher ability farmers invest more (because distortions no longer claw back incremental profits). The net result is a more polarized ability distribution with a heavier right tail, consistent with the concentrated structure of agriculture in advanced economies. Importantly, the paper cautions that abstracting from measurement error inflates the estimated undistorted-economy gains by more than a factor of two: a model without measurement error yields gains more than twice as large as the calibrated model that accounts for it.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central policy implication is that institutions distorting factor markets — particularly those that generate a positive correlation between farm productivity and the effective tax rate (captured by rho) — reduce agricultural TFP through three compounding channels, with two-thirds of the loss arising from dynamic distortions (investment suppression and crop selection) rather than static factor reallocation. This means that standard static calculations of misallocation costs substantially understate the true costs. Land accumulation restrictions that prevent productive farms from expanding (the historical legacy in north Vietnam, where 82.8% of Red River Delta agricultural land was state-allocated) are particularly costly because they are the empirical analog of high rho. The scope conditions are: (1) the analysis applies to the Vietnamese agricultural context in 2006-2016, a period well after initial reform but still characterized by persistent institutional differences; (2) the model abstracts from occupational choice and structural transformation, which other work has shown amplify distortion costs further; (3) the main results are robust to the north-south within-country design but level estimates (gains from the first-best) are sensitive to measurement error treatment. The paper suggests that reducing the productivity-distortion correlation — e.g., through secure land titles and functioning land rental markets — would unlock gains exceeding what static misallocation calculations imply.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Distortion elasticity (rho)&lt;/strong&gt;: The parameter governing how strongly institutional distortions — modeled as idiosyncratic revenue taxes — are correlated with farm-level productivity. A higher rho means more productive farms face proportionately larger distortions, compressing both static factor allocation and the dynamic return to investing in ability. In the paper&amp;rsquo;s calibration, rho = 0.79 for south Vietnam and 0.91 for north Vietnam; the difference accounts for the majority of the measured North-South productivity gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Managerial ability ladder&lt;/strong&gt;: The endogenous component of farm productivity that farmers accumulate through investment. A farmer at ability node h has productivity phi^h; investing e units of output raises ability to the next node with probability x = (e/a)^(1/zeta). The investment return depends on the incremental profit gain from higher ability, which is suppressed when the distortion elasticity rho is large, creating a tight link between static institutional distortions and dynamic farm growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crop-specific distortion (phi_i)&lt;/strong&gt;: A factor in the distortion specification that captures institutional barriers differentially affecting specific crops. In south Vietnam, phi_perennial = 1.61, meaning perennial-crop growers face lower effective taxes than rice farmers; in north Vietnam, phi_perennial = 0.68, reversing the ranking. This parameter embeds market-access barriers, infrastructure gaps, and regulatory disadvantages specific to particular crops.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Government-imposed crop restriction (omega)&lt;/strong&gt;: The share of farms legally required to grow rice regardless of relative profitability or household preferences, reflecting Vietnamese national food-security policies. The restriction is more prevalent in the north (43% of farms) than the south (23%). Unlike idiosyncratic distortions, crop restrictions enter the model as a direct constraint on the discrete crop-choice decision rather than as a tax on revenue, and the paper finds their productivity cost is relatively small (1.4% output loss) because many restricted farmers would have chosen rice anyway.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic misallocation&lt;/strong&gt;: The productivity losses arising from distortions&amp;rsquo; effects on farms&amp;rsquo; forward-looking decisions — specifically the choice of crop (crop selection) and investment in managerial ability — as opposed to the static misallocation of given factor inputs across farms with fixed productivities. In the paper, dynamic misallocation accounts for two-thirds of the total productivity gap, more than doubling the contribution of static factor misallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;BKR measurement-error statistic&lt;/strong&gt;: A diagnostic from Bils, Klenow, and Ruane (2021) that estimates the ratio of true wedge dispersion to observed wedge dispersion using the cross-term in a regression of log output changes on log wedge, log input, and their interaction. Values near one indicate little measurement error; values near zero indicate the observed wedge is mostly noise. The paper finds BKR = 0.906 for south Vietnam and 0.987 for north Vietnam, indicating measurement error is modest and is unlikely to confound the north-south comparison.&lt;/p&gt;</description></item><item><title>Distributional Consequences of Becoming Climate-Neutral</title><link>https://macropaperwarehouse.com/papers/distributional-consequences-of-becoming-climate-neutral/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/distributional-consequences-of-becoming-climate-neutral/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how the EU&amp;rsquo;s Fit-for-55 climate package will affect aggregate output and distribute its costs across the income distribution. The question matters because energy is a necessity good — poorer households devote a larger share of spending to energy — so policies that raise energy prices are regressive in their first-order incidence. Despite a large literature on the aggregate macroeconomics of the green transition, distributional consequences have received limited attention.&lt;/p&gt;
&lt;p&gt;The authors build a parsimonious dynamic general-equilibrium model with two infinitely-lived households (rich and poor), a standard output-producing firm that treats energy as a complementary CES input alongside the capital-labor aggregate, and an energy-producing sector that combines a carbon-intensive brown technology with a carbon-free green technology as imperfect substitutes (CES with elasticity of substitution calibrated to 3 following Papageorgiou et al. 2017). The novel feature is Price Independent Generalized Linearity (PIGL) non-homothetic preferences following Boppart (2014), which generate nonlinear Engel curves: the poor agent&amp;rsquo;s energy expenditure share exceeds the rich agent&amp;rsquo;s, matching Eurostat Household Finance and Consumption Survey data (2015) showing the bottom income quintile has more than twice the energy expenditure share of the top quintile. The model targets an 18% energy expenditure share for the poor agent and 7.5% for the rich agent. The rich agent holds all financial wealth; the poor agent lives on labor income alone. The government taxes the brown technology and recycles revenue as a green-technology subsidy under a balanced budget, representing the ETS. Agents have perfect foresight. The paper simulates perfect-foresight transitions from an initial steady state to a new climate-neutral steady state, with the transition path endogenously determining the new steady state — a nonstandard feature arising from non-homothetic preferences.&lt;/p&gt;
&lt;p&gt;In the baseline scenario (linear tax ramp over 25 years), achieving an 85% reduction in brown energy use requires a 168% tax on the brown technology. This drives the price of energy services up by 49%, GDP down by 9.3% in the new steady state, energy as a production input down by 10.9%, and capital input down by 9.3%, while the real wage falls by roughly 7% and the real interest rate is nearly unchanged (dropping by only 0.02 percentage points transiently). The welfare cost measured in expenditure-equivalent terms is a 10.8% loss for the rich agent and a 16.2% loss for the poor agent — the poor agent suffers approximately 50% more. To finance consumption during the transition the poor agent accumulates debt equal to 38.8% of annual income.&lt;/p&gt;
&lt;p&gt;Results are highly sensitive to the brown-green substitution elasticity: raising it from 3 to 5 roughly halves the required tax (to 78.6%) and halves GDP losses (to 4.7%); lowering it to 2 roughly doubles the tax (to 354%) and GDP losses (to 17.7%). Non-homothetic preferences matter quantitatively: switching to homothetic preferences (while preserving different expenditure shares) shrinks aggregate GDP losses by 26% and eliminates nearly all distributional disparity, confirming that the non-homotheticity — not merely different expenditure levels — is the operative distributional mechanism. If the Fit-for-55 energy efficiency improvement target of 1.49% per year is simultaneously achieved, the required tax falls to 136%, the price of energy actually declines by 5.5%, and GDP rises by 1.1% in the new steady state, with the poor agent benefiting slightly more and accumulating assets (4% of annual income) rather than debt.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-modeling-and-calibration-strategy-and-what-are-the-main-threats"&gt;Q1. What is the core modeling and calibration strategy, and what are the main threats?&lt;/h3&gt;
&lt;p&gt;The paper is a quantitative theory exercise with no econometric identification. Calibration targets HFCS Eurostat data (2015) for energy expenditure shares by income quintile, the Papageorgiou et al. (2017) estimate of the brown-green substitution elasticity (ρE = 3), and stylized facts on wealth and income distribution from Krueger, Mitman, and Perri (2016). The main threat is parameter uncertainty around ρE, which the paper acknowledges is poorly identified empirically and which drives the results almost one-for-one. The sensitivity analysis explores ρE ∈ {2, 3, 5}, a range the paper concedes is narrow relative to the literature&amp;rsquo;s full dispersion.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-generating-the-distributional-gap-between-rich-and-poor"&gt;Q2. What are the main mechanisms generating the distributional gap between rich and poor?&lt;/h3&gt;
&lt;p&gt;Three reinforcing channels: (1) Non-homothetic preferences give the poor agent a higher energy expenditure share (18% vs. 7.5%), so the 49% energy price increase hits the poor&amp;rsquo;s budget much harder as a share of income. (2) The poor agent cannot buffer the shock through wealth drawdowns (holding zero net assets initially), forcing it to accumulate debt of 38.8% of annual income. (3) Non-homothetic preferences alter the labor supply response: as expenditures fall, the poor agent&amp;rsquo;s labor supply declines less than the rich agent&amp;rsquo;s (the rich agent decreases labor supply by 0.2 percentage points more), reflecting that leisure is a luxury good in this preference system. In the new steady state the rich agent&amp;rsquo;s consumption of the consumption good drops sharply while the rich agent front-loads consumption at the announcement, immediately jumping 2% higher.&lt;/p&gt;
&lt;h3 id="q3-how-are-non-homothetic-preferences-distinguished-empirically-and-in-the-model-from-simply-having-different-expenditure-shares"&gt;Q3. How are non-homothetic preferences distinguished empirically and in the model from simply having different expenditure shares?&lt;/h3&gt;
&lt;p&gt;Section 4.4 runs a counterfactual with homothetic preferences (ε = 0) but preserves identical initial expenditure shares for each agent (7.5% and 18%) by making ν agent-specific. Under homotheticity the expenditure shares do not vary with income as the transition unfolds. The comparison shows that GDP losses shrink by 26% (from 9.3% to 6.9%) and the distributional gap nearly vanishes — both agents experience almost identical welfare losses. This decomposition isolates the effect of non-homotheticity itself: it is the income-dependent adjustment of expenditure shares during the transition, not merely the different initial levels, that drives both larger aggregate losses and the distributional disparity.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented-and-along-what-dimensions"&gt;Q4. What heterogeneity is documented and along what dimensions?&lt;/h3&gt;
&lt;p&gt;Heterogeneity is modeled along two dimensions: initial wealth (rich holds all assets; poor holds zero) and energy expenditure shares (18% for poor, 7.5% for rich) arising from non-homothetic preferences. The model produces no within-group heterogeneity by construction (two-agent framework). The paper documents the time paths of consumption, expenditures, expenditure equivalents, energy expenditure shares, and wealth shares for each agent separately along the transition, showing that both agents cut energy consumption by roughly 15% while the poor agent cuts consumption-good spending by substantially more than the rich agent.&lt;/p&gt;
&lt;h3 id="q5-what-alternative-transition-timing-paths-are-explored-and-what-do-they-imply"&gt;Q5. What alternative transition timing paths are explored and what do they imply?&lt;/h3&gt;
&lt;p&gt;Three alternatives supplement the linear baseline: tax introduction after 1 year, after 12.5 years, and after 25 years of the announcement. Key findings: (a) the required final tax rate is nearly insensitive to timing — the 25-year-delayed scenario requires 172% vs. 168% in the baseline; (b) conditional on excluding climate damages, it is always welfare-superior to delay implementation, with the poor agent gaining close to 3.5 percentage points in expenditure equivalent welfare by delaying to 25 years vs. implementing after 1 year; (c) gradual vs. immediate introduction yields similar welfare outcomes in the benchmark without adjustment costs, but with investment adjustment costs (χ = 10) a sudden implementation causes a brief sharp drop in the real interest rate without large quantity effects.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-gdp-measure-differ-from-aggregate-output-in-the-model"&gt;Q6. How does the GDP measure differ from aggregate output in the model?&lt;/h3&gt;
&lt;p&gt;GDP is defined to exclude the share of final output used as input into energy production. Aggregate output Y falls 7.3% in the new steady state, but GDP falls 9.3%. The gap (approximately 2 percentage points) reflects the increased resource cost of energy production under the green transition: because the brown and green technologies are imperfect substitutes, satisfying the emission reduction target requires devoting a larger share of final output to producing energy services, a real resource drain captured in the GDP definition but excluded from raw output Y.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-energy-efficiency-scenario-imply-and-what-is-its-key-caveat"&gt;Q7. What does the energy efficiency scenario imply, and what is its key caveat?&lt;/h3&gt;
&lt;p&gt;If energy efficiency improves at 1.49% per year over 25 years (a 45% cumulative gain in energy-producing-firm total factor productivity), the required tax falls to 136.3%, the price of energy declines by 5.5% (rather than rising 49%), and GDP rises 1.1% rather than falling 9.3%. The poor agent benefits more from the efficiency gains and accumulates assets worth 4% of annual income rather than debt. The critical caveat is that the efficiency improvement is modeled as purely exogenous and costless. The paper explicitly acknowledges that achieving these efficiency gains may require investment that is not modeled, so the results should be interpreted as an upper bound on the offsetting potential.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-relate-to-and-differ-from-the-most-closely-related-prior-work"&gt;Q8. How does the paper relate to and differ from the most closely related prior work?&lt;/h3&gt;
&lt;p&gt;Ascari et al. (2025) is the closest related paper (developed independently). Differences: (i) Ascari et al. use a Bewley-type incomplete-markets model generating heterogeneity through random discount factors, whereas this paper uses a two-agent complete-markets construct with exogenously fixed initial wealth; (ii) this paper allows endogenous labor supply, which increases short-run flexibility; (iii) this paper does not consider transfer schemes to redistribute away from distributional consequences. Results are described as broadly consistent. Fried, Novan, and Peterman (2018) and Boehl and Budianto (2024) use OLG models and find inequality implications but focus on inter-generational rather than intra-generational distributional effects.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The core implications are: (1) the Fit-for-55 emission tax alone is regressive — the poor bear a welfare loss 50% larger than the rich and end up with 38.8% of annual income in additional debt; (2) delaying tax implementation (with early announcement) is welfare-improving in the absence of climate damage modeling — the welfare difference is nearly 3.5 percentage points for the poor between fastest and latest implementation; (3) if energy efficiency targets are met exogenously, the transition is nearly costless and distributional concerns vanish; (4) the regressive result is conditional on the government recycling tax revenues to green-technology subsidies rather than to household transfers. All these implications are conditional on European economies where climate damages are plausibly small and the model abstracts from open-economy dynamics, endogenous technology, and within-income-group heterogeneity.&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-are-reported"&gt;Q10. What robustness checks are reported?&lt;/h3&gt;
&lt;p&gt;Five robustness exercises are reported: (1) investment adjustment costs raised from χ = 0 to χ = 10 — minimal effect on welfare or quantities in the smooth baseline, though sudden tax introduction produces a brief interest-rate plunge; (2) homothetic preferences counterfactual while maintaining initial expenditure shares (Section 4.4); (3) elasticity of substitution between brown and green technology at ρE = 2 and ρE = 5 (Section 4.3, Table 2); (4) alternative transition timing (1 year, 12.5 years, 25 years post-announcement; Section 4.2); (5) simultaneous energy efficiency improvement of 1.49% per year (Section 4.5). A New Keynesian extension with Rotemberg price adjustment costs and a Taylor rule (Appendix B) is also provided for robustness on inflation dynamics.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-main-caveats-or-limitations-acknowledged-by-the-authors"&gt;Q11. What are the main caveats or limitations acknowledged by the authors?&lt;/h3&gt;
&lt;p&gt;Climate damages are excluded, so the paper understates the case for early action and cannot provide a full welfare comparison between acting early and acting late. Energy efficiency improvement is modeled as exogenous and costless, overstating the net gain from that channel. The two-agent framework abstracts from within-group heterogeneity and overlapping generations. Open-economy dynamics are not modeled; the brown-technology structure serves as a reduced-form for energy imports but does not capture international price feedback. The elasticity of substitution between brown and green technology is uncertain, and results are nearly proportional to this parameter. The model has no endogenous innovation or directed technical change, limiting applicability to long-run transition analysis.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Non-homothetic PIGL preferences&lt;/strong&gt;: Preferences of the Price Independent Generalized Linearity class (Boppart 2014) where energy expenditure shares depend on income level, making energy a necessity good (share declining in income) and consumption goods a luxury. Parameter ε ∈ (0,1) controls non-homotheticity; ε = 0 recovers homothetic preferences. The paper calibrates γ = 0.639 from CEX data, implying an elasticity of substitution between consumption and energy goods of approximately 0.4.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Brown vs. green technology&lt;/strong&gt;: Two imperfectly substitutable technologies for producing energy services within the model&amp;rsquo;s energy sector. The brown technology converts units of final output into energy services using a carbon-intensive (emission-producing) process; the green technology is emission-free. They enter a CES aggregator for energy production with elasticity ρE calibrated to 3. Imperfect substitutability means the green transition raises the cost of energy services even with subsidies to green technology.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Expenditure equivalent loss&lt;/strong&gt;: The welfare metric used in the paper: the percentage change in expenditures in the initial steady state (without any tax) that would make an agent indifferent between remaining in the initial steady state and living through the actual transition path. Defined implicitly by equating flow utility at scaled initial expenditures to flow utility along the transition. Baseline results: -10.8% for the rich agent and -16.2% for the poor agent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tax on the brown technology&lt;/strong&gt;: The policy instrument modeled as capturing the essence of EU ETS and national carbon schemes. It raises the unit cost of the emission-intensive energy input; revenue is recycled as a subsidy to the green technology within a balanced government budget rather than distributed to households. A 168% tax achieves the 85% emission reduction target in the baseline, implying fossil fuel prices nearly triple.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous final steady state&lt;/strong&gt;: The model&amp;rsquo;s new steady state after the green transition is not predetermined; it depends on the wealth distribution that emerges endogenously during the transition. Because markets are complete and preferences are non-homothetic, different transition paths generate different terminal wealth distributions and therefore different aggregate outcomes in the new steady state. This prevents backward solution and requires a fully nonlinear transition path solver.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Energy expenditure share by income quintile&lt;/strong&gt;: The empirical regularity, documented from Eurostat HFCS data (2015), that the bottom income quintile devotes more than twice the fraction of disposable income to energy (electricity, gas, fuels for personal transport) as the top quintile. This fact calibrates the non-homotheticity of preferences (targeting 18% for the poor agent and 7.5% for the rich agent) and motivates the paper&amp;rsquo;s focus on distributional consequences.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Elasticity of substitution between brown and green technology (ρE)&lt;/strong&gt;: The key production-side parameter governing how easily the energy sector can switch from fossil-fuel to clean inputs. Calibrated to ρE = 3 from Papageorgiou et al. (2017). Results are nearly proportional to this parameter: ρE = 5 halves and ρE = 2 roughly doubles the required tax, GDP losses, and welfare costs. The paper identifies this as the dominant source of quantitative uncertainty.&lt;/p&gt;</description></item><item><title>Entrepreneurial Investment Dynamics and the Wealth Distribution</title><link>https://macropaperwarehouse.com/papers/entrepreneurial-investment-dynamics-and-the-wealth-distribution/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/entrepreneurial-investment-dynamics-and-the-wealth-distribution/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how the illiquidity of entrepreneurial capital shapes investment dynamics and wealth inequality. The central question is whether entrepreneurship drives wealth heterogeneity or merely attracts the already-wealthy — and, specifically, whether the investment behavior of nascent entrepreneurs can be rationalized by frictions on capital reallocation rather than financial constraints alone.&lt;/p&gt;
&lt;p&gt;The empirical foundation is the restricted Kauffman Firm Survey (KFS), a single-cohort panel of 3,140 U.S. firms founded in 2004 and tracked through 2011. The key measurement is the log average revenue product of capital (log ARPK), residualized on two-digit NAICS industry fixed effects and time dummies. Two striking facts emerge. First, the cross-sectional distribution of log ARPK is left-skewed (skewness approximately -0.33, mean -0.49, standard deviation 1.75, kurtosis 5.7). Second, the distribution shows asymmetric persistence: the autocorrelation of log ARPK in the bottom quintile (ρ₁ = 0.897) is statistically significantly larger than in the top quintile (ρ₅ = 0.443), and the diagonal entry of the estimated transition matrix for the first quintile (0.614) substantially exceeds that for the fifth (0.568). These facts are inconsistent with standard models: a frictionless dynamic investment model with time-to-build predicts i.i.d. ARPK; one with collateral constraints predicts right-skewness and right-tail persistence.&lt;/p&gt;
&lt;p&gt;The model extends Cagetti and De Nardi (2006) by distinguishing between liquid bonds and illiquid entrepreneurial capital. Capital adjustment generates four friction types: a proportional fixed cost (fs) on upward investment, a proportional transaction cost (λ) on downsizing, an additional proportional cost (ζ) on exit, and a minimum capital requirement on entry. The model is calibrated via indirect inference to identifying moments from the KFS (persistence and skewness of log ARPK, investment rate distribution, share of employer firms, entry and exit rates) plus economy-wide targets (entrepreneur fraction, interest rate of 3–4%).&lt;/p&gt;
&lt;p&gt;The FULL-sample calibration yields λ = 0.43 (43% loss on capital sold by continuing entrepreneurs) and ζ = 0.55 (additional 55% write-down upon exit), with a proportional fixed cost fs = 0.035 (3.5%). The effective net collateral constraint is approximately 44% of the real capital value. These frictions are quantitatively large: eliminating them under general equilibrium raises aggregate TFP in the entrepreneurial sector by 23.3% and average welfare by 23.1% in consumption equivalent variation terms. Decomposing the welfare losses relative to a complete-markets benchmark shows that approximately 89% of the total welfare loss (relative to full frictions) is attributable to market incompleteness and financial frictions, with the remaining 11% directly attributable to the illiquidity frictions — that is, frictions alone account for roughly 7.15 percentage points of a total 64.8% lifetime consumption welfare loss.&lt;/p&gt;
&lt;p&gt;A key finding on wealth inequality contradicts prior literature. When calibrated to KFS micro-data, the model generates a Gini coefficient of 0.65 (FULL sample) or 0.53 (NAICS54), well below the empirical U.S. Gini of approximately 0.8. The top 1% hold only 26% of wealth in the FULL calibration versus roughly 30% empirically. This contrasts with Quadrini (2000) and Cagetti and De Nardi (2006), who match the wealth distribution by calibrating to PSID or SCF household survey data. The reason for the gap is the left-skewed, illiquidity-depressed returns to entrepreneurship in the KFS: the calibrated returns to scale (ν = 0.79 FULL, 0.82 NAICS54) and the transaction costs together suppress the variance of capital income returns. Removing illiquidity frictions raises the Gini from 0.65 to 0.77 (fixed-r partial equilibrium) or 0.72 (general equilibrium), demonstrating that capital illiquidity compresses the wealth distribution by depressing average entrepreneurial returns.&lt;/p&gt;
&lt;p&gt;Three policy experiments — credit expansion (reducing borrowing spreads à la SBA 7(a) programs), a government buyer-of-last-resort for used capital (Resale I), and exit-cost reduction (Fire sale) — all raise welfare by 0.07–0.15% in consumption equivalent terms and TFP by 0.5–0.9% relative to benchmark. Resale policies are preferred by entrepreneurs; workers prefer the credit policy. All three policies benefit lower-wealth households more than wealthy ones (the richest decile suffers welfare losses due to the savings tax used to finance the programs). The paper concludes that policies addressing capital illiquidity can yield welfare gains comparable to or exceeding standard credit provision programs, and that the distinction between illiquidity risk and financial constraint risk has first-order importance for policy design.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-two-core-empirical-facts-from-the-kfs-that-motivate-the-paper-and-why-do-standard-models-fail-to-generate-them"&gt;Q1. What are the two core empirical facts from the KFS that motivate the paper, and why do standard models fail to generate them?&lt;/h3&gt;
&lt;p&gt;First, the cross-sectional distribution of log ARPK among KFS firms is left-skewed (skewness ≈ -0.33), not symmetric or right-skewed. Second, log ARPK shows higher persistence in the left tail (autocorrelation ρ₁ = 0.897 for bottom-quintile firms) than in the right tail (ρ₅ = 0.443). A frictionless dynamic model with time-to-build predicts i.i.d. log ARPK that inherits the distribution of TFP innovations, generating no skewness under Gaussian shocks and no persistence. Models with collateral constraints (as in Cagetti and De Nardi 2006) generate right-skewed ARPK with right-tail persistence, because constrained firms operate below optimal scale, pushing ARPK above the unconstrained optimum. Neither class of models can produce the left-skewed, left-tail-persistent pattern in the KFS.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-mechanism-by-which-partial-irreversibility-generates-left-skewness-and-left-tail-persistence"&gt;Q2. What is the mechanism by which partial irreversibility generates left-skewness and left-tail persistence?&lt;/h3&gt;
&lt;p&gt;Partial irreversibility creates an asymmetry between the purchase price and the resale price of capital (the resale price being 1 − λ per unit). When a bad productivity shock hits, the option value of waiting to recover is higher than the cost of holding excess capital, so entrepreneurs adopt a &amp;lsquo;wait-and-see&amp;rsquo; attitude and maintain oversized firms rather than downsizing immediately. This creates a left tail of low-ARPK, large-capital firms. Moreover, since the incentive to wait is itself persistent (the transitory bad shock must resolve before the entrepreneur will downsize), the left tail displays higher autocorrelation. The exit cost ζ amplifies this for the exit margin: entrepreneurs with poor draws stay in business longer than is efficient, further extending the left tail. The right tail is not symmetrically elongated because entrepreneurs seeking to expand face a different option value (the call option value of capital rises), leading them to invest to smaller sizes, slightly thickening the right tail — but not enough to overcome the left-tail extension.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-calibration-strategy-and-which-parameters-are-identified-by-which-moments"&gt;Q3. What is the calibration strategy, and which parameters are identified by which moments?&lt;/h3&gt;
&lt;p&gt;Eleven parameters are jointly calibrated to KFS moments via indirect inference. The key mappings are: the downsizing transaction cost λ is identified by the asymmetric left-tail persistence of log ARPK (the ratio ρ₁/ρ₅ increases monotonically in λ); the exit cost ζ is identified by the skewness of log ARPK (higher ζ monotonically increases left skewness); the collateral constraint ϕ also affects skewness but has no monotone effect on ρ₁/ρ₅, aiding separation; the returns to scale ν is identified by the coefficient from a log-revenue on log-capital regression for employer firms; the fixed investment cost fs is identified by the fraction reporting positive investment; TFP shock autocorrelation ρ_z is identified by investment rate autocorrelation; the shock standard deviation σ_z by the coefficient of variation of investment rates; and the worker signal distortion and entrepreneur signal distortion parameters control entry and exit rates respectively. The discount factor β pins down the interest rate. Two separate calibrations are run: one targeting full KFS sample moments (FULL) and one targeting the modal industry — Professional, Scientific and Technical Services (NAICS54, 24.7% of the sample) — as a robustness check.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-main-calibrated-parameter-values-and-how-do-they-compare-across-the-full-and-naics54-calibrations"&gt;Q4. What are the main calibrated parameter values and how do they compare across the FULL and NAICS54 calibrations?&lt;/h3&gt;
&lt;p&gt;For the FULL calibration: λ = 0.43, ζ = 0.55, ϕ = 0.92, fs = 0.035, ρ_z = 0.66, σ_z = 0.43, ν = 0.79, β = 0.9265, α_e = 0.63. For NAICS54: λ = 0.53, ζ = 0.75, ϕ = 0.035, fs = 0.23, ρ_z = 0.66, σ_z = 0.43, ν = 0.82, β = 0.94, α_e = 0.50. The illiquidity parameters (λ and ζ) are larger in NAICS54 than in FULL. The collateral constraint parameter ϕ differs substantially (0.92 FULL versus 0.035 NAICS54), though the net effective collateral constraint (accounting for λ and depreciation) converges to a similar range in both calibrations.&lt;/p&gt;
&lt;h3 id="q5-how-are-the-illiquidity-and-financial-friction-channels-distinguished-both-theoretically-and-empirically"&gt;Q5. How are the illiquidity and financial friction channels distinguished both theoretically and empirically?&lt;/h3&gt;
&lt;p&gt;Theoretically, collateral constraints (parameterized by ϕ) make the lower support of log ARPK truncated from the left (log ARPK ≥ log(r+δ) - log α), generating right-skewness and right-tail persistence. Illiquidity frictions (λ and ζ), by contrast, induce a wait-and-see option value that extends the left tail of ARPK while leaving the right tail relatively thinner, generating left-skewness and left-tail persistence. Empirically, the paper proposes using the sign and magnitude of the skewness of log ARPK (negative implies illiquidity dominates; positive implies financial frictions dominate) and the ratio of left-tail to right-tail persistence (ρ₁/ρ₅ &amp;gt; 1 indicates illiquidity frictions, &amp;lt; 1 indicates financial frictions) as discriminating statistics. Separately, the portfolio composition of entrepreneurs offers a further discriminating test: increasing illiquidity drives entrepreneurs to hold more liquid assets (flight to liquidity), while tightening collateral constraints pushes entrepreneurs toward more illiquid assets in their portfolios.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-aggregate-tfp-and-welfare-findings-from-the-counterfactual-analysis"&gt;Q6. What are the aggregate TFP and welfare findings from the counterfactual analysis?&lt;/h3&gt;
&lt;p&gt;Under general equilibrium, removing all illiquidity frictions (λ = ζ = fs = 0) raises entrepreneurial sector TFP by 23.3% and average economy-wide welfare by 23.1% in consumption equivalent variation. Under partial equilibrium (fixed interest rate), welfare gains are even larger: 24.8% (entrepreneur subgroup) and 58.3% (worker subgroup), for an economy-wide average of 16.6%. The GE result is somewhat lower because the interest rate adjusts when more capital flows into entrepreneurship. The average productivity of entrepreneurs (conditional on being an entrepreneur) is 8.8% higher in the no-friction world than in the benchmark. The TFP gains arise from both extensive-margin selection (higher-productivity entrepreneurs enter; lower-productivity ones exit) and intensive-margin reallocation (high-productivity firms operate closer to optimal scale; low-productivity firms downsize rather than persist).&lt;/p&gt;
&lt;h3 id="q7-how-does-the-paper-decompose-total-welfare-losses-between-market-incompleteness-and-the-illiquidity-distortions"&gt;Q7. How does the paper decompose total welfare losses between market incompleteness and the illiquidity distortions?&lt;/h3&gt;
&lt;p&gt;Following Buera and Shin (2011), the paper computes welfare as a fraction of lifetime consumption relative to a complete-markets benchmark (a social planner&amp;rsquo;s problem where the planner allocates occupational choice and capital optimally). Relative to complete markets, the economy with no illiquidity frictions but with market incompleteness loses approximately 57.7% of lifetime consumption. The benchmark economy (with all frictions) loses approximately 64.8% of lifetime consumption relative to complete markets. The difference — approximately 7.15 percentage points — is attributed to the illiquidity frictions. As a share of the total frictional loss, about 89% is attributable to market incompleteness and financial frictions, and 11% to the illiquidity frictions. While 11% may seem small as a fraction, in absolute terms it is economically non-trivial.&lt;/p&gt;
&lt;h3 id="q8-why-does-the-paper-find-that-entrepreneurship-cannot-match-the-empirical-wealth-distribution-when-calibrated-to-the-kfs"&gt;Q8. Why does the paper find that entrepreneurship cannot match the empirical wealth distribution when calibrated to the KFS?&lt;/h3&gt;
&lt;p&gt;The model generates a Gini of 0.65 (FULL) or 0.53 (NAICS54) against a U.S. empirical Gini of approximately 0.8. The top 1% holds roughly 26% of wealth in the FULL calibration versus around 30% empirically. Two factors suppress capital income risk in the KFS-calibrated model. First, the calibrated returns to scale (ν = 0.79 FULL, 0.82 NAICS54) are lower than those used by Cagetti and De Nardi (2006) (ν ≈ 0.88), which were calibrated to PSID/SCF data on large-ish successful firms. Lower ν translates exponentially into lower variance of capital income. Second, the illiquidity frictions directly depress average returns to entrepreneurship by raising the user cost of capital and forcing entrepreneurs into suboptimal firm sizes. These two forces together prevent the model from generating the thick right tail of wealth needed to match empirical distributions. The paper argues that the KFS captures &amp;lsquo;broad&amp;rsquo; small-scale entrepreneurship, not the high-growth, high-return entrepreneurs who likely account for the top of the wealth distribution.&lt;/p&gt;
&lt;h3 id="q9-how-does-capital-illiquidity-affect-the-wealth-distribution-conditional-on-holding-returns-to-scale-fixed"&gt;Q9. How does capital illiquidity affect the wealth distribution conditional on holding returns to scale fixed?&lt;/h3&gt;
&lt;p&gt;More illiquid capital (higher λ or ζ) compresses the wealth distribution and lowers the Gini coefficient. The Gini rises from 0.65 (benchmark FULL calibration) to 0.77 under partial equilibrium without illiquidity frictions, and to 0.72 under general equilibrium without illiquidity frictions (while holding the net collateral constraint constant). The NAICS54 benchmark Gini is 0.53, rising to 0.76 (PE) or 0.68 (GE) without illiquidity frictions. The mechanism is that illiquid capital depresses the average return to entrepreneurial wealth, which compresses the income process and reduces the variance of wealth accumulation. Additionally, illiquid capital forces entrepreneurs to hold more bonds as a liquidity buffer, reducing the overall scale of their business investment and thus their lifetime income.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-three-policy-experiments-and-their-comparative-findings"&gt;Q10. What are the three policy experiments and their comparative findings?&lt;/h3&gt;
&lt;p&gt;The three policies are all financed by a proportional tax on bond savings returns. (1) Credit expansion: the government subsidizes borrowing intermediation costs (analogous to SBA 7(a)/CDC 504 programs), reducing the spread between the saving and borrowing rate. Economy-wide welfare rises by about 0.147%; TFP rises by about 0.9% relative to benchmark. Workers benefit more (0.169%) than entrepreneurs (-0.006% average for all entrepreneurs, since most wealthy entrepreneurs do not borrow and pay the tax). (2) Resale policy I (Buyer of last resort for all used capital): government offers a higher resale price q ≥ 1 − λ. Economy-wide welfare rises about 0.076%; TFP rises 0.6%. Entrepreneurs gain (0.084%) while workers also gain (0.074%) indirectly through the option value of future entrepreneurship. (3) Fire-sale (exit cost reduction only, Resale II): government subsidizes exiting entrepreneurs&amp;rsquo; capital resale. Economy-wide welfare rises 0.073%; TFP rises 0.5%. Workers prefer credit; entrepreneurs prefer resale policies. Wealthiest decile suffers welfare losses under all three policies. All welfare numbers are in consumption equivalent variation.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-relate-to-cagetti-and-de-nardi-2006-and-where-does-it-diverge"&gt;Q11. How does the paper relate to Cagetti and De Nardi (2006) and where does it diverge?&lt;/h3&gt;
&lt;p&gt;The paper builds directly on the Cagetti and De Nardi (2006) framework of occupational choice and incomplete markets with collateral constraints, extending it by separating liquid bonds from illiquid physical capital. In Cagetti and De Nardi (2006), bonds and capital are perfect substitutes; the sole friction is a collateral constraint that limits investment. The paper shows that this one-asset framework generates right-skewed ARPK and right-tail persistence — inconsistent with KFS facts. The paper&amp;rsquo;s two-asset framework with partial irreversibility generates left-skewed ARPK and left-tail persistence. Furthermore, Cagetti and De Nardi (2006) calibrate to PSID/SCF income data and successfully match the wealth distribution; the paper shows this success partly reflects the higher returns to scale implied by those data. When calibrated directly to KFS firm-level data, the model substantially undershoots the empirical wealth inequality, because the KFS captures a representative sample of small-scale entrepreneurs with genuinely lower returns to scale and significant illiquidity frictions.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-role-of-the-options-value-effect-and-the-collateral-constraint-channel-in-the-model-and-how-do-they-differ"&gt;Q12. What is the role of the options value effect and the collateral constraint channel in the model, and how do they differ?&lt;/h3&gt;
&lt;p&gt;The options value effect is described as the primary distortion. When capital is illiquid (λ or ζ &amp;gt; 0), the put option value of capital falls (selling capital is costly), raising the threshold signal required for workers to enter entrepreneurship, and raising the threshold signal required for incumbents to exit. As a result, entry rates fall, exit rates fall, potential entrepreneurs delay entry, and poorly performing entrepreneurs overstay. Along the intensive margin, the asymmetric purchase/resale price leads entrepreneurs planning to downsize to wait (operating larger-than-optimal firms) and entrepreneurs planning to invest to be more cautious (operating smaller-than-optimal firms). The collateral constraint channel is a secondary effect: illiquid capital reduces the net resale value that can serve as collateral (effective constraint = (1-λ)(1-δ)(ϕ)k&amp;rsquo;), tightening the borrowing constraint even when the formal collateral parameter ϕ is moderate. Crucially, while tighter ϕ forces entrepreneurs to hold more illiquid capital (no flight to liquidity), higher λ forces entrepreneurs to hold more liquid assets (flight to liquidity) — a key empirical distinction.&lt;/p&gt;
&lt;h3 id="q13-what-robustness-exercises-does-the-paper-conduct"&gt;Q13. What robustness exercises does the paper conduct?&lt;/h3&gt;
&lt;p&gt;The paper runs two separate full calibrations: one to the entire KFS sample (FULL) and one to the modal industry NAICS54 (Professional, Scientific and Technical Services, 24.7% of the sample). Both calibrations are used to assess the wealth distribution findings. The paper also examines moments at the two-digit industry level (only one industry shows statistically significant results due to small sample size, though most show economically significant signs). An additional measurement error parameter is explored in the appendix, where capital is assumed to be observed with multiplicative log-normal error; this helps improve model fit to the data. All policy experiments are computed under both partial equilibrium (fixed interest rate) and general equilibrium. The paper also analytically proves (in the appendix) the ARPK distribution properties for the four benchmark frameworks (frictionless, time-to-build only, static collateral constraints, and dynamic collateral constraints), establishing the theoretical necessity of partial irreversibility for the facts.&lt;/p&gt;
&lt;h3 id="q14-what-heterogeneity-in-welfare-effects-is-documented-across-the-wealth-distribution"&gt;Q14. What heterogeneity in welfare effects is documented across the wealth distribution?&lt;/h3&gt;
&lt;p&gt;Under all three policy experiments, welfare gains decrease with wealth. The poorest households gain the most in consumption equivalent variation terms because they receive a disproportionate share of the program&amp;rsquo;s benefits (better borrowing conditions, higher resale prices, improved option value of entrepreneurship) while paying a smaller absolute share of the savings tax used to finance the programs. The top 10% richest households — who are the primary taxpayers — experience welfare losses under all three policies. This pattern holds across credit, resale, and fire-sale policies, though the magnitude varies. Separately, entrepreneurs (who are wealthier on average, with over 50% concentrated in the top wealth decile) mostly lose from the credit policy (they fund it but don&amp;rsquo;t directly borrow) while gaining from resale policies (they benefit from higher capital resale prices regardless of wealth position). Workers (who are generally poorer) overwhelmingly gain from credit policies since the option value of switching to entrepreneurship rises substantially.&lt;/p&gt;
&lt;h3 id="q15-what-does-the-paper-imply-for-interpreting-the-literature-on-financial-constraints-and-entrepreneurship"&gt;Q15. What does the paper imply for interpreting the literature on financial constraints and entrepreneurship?&lt;/h3&gt;
&lt;p&gt;The paper issues several cautionary findings. First, the implied formal collateral parameter is relatively loose (ϕ = 0.92), consistent with Hurst and Lusardi (2004), Nanda (2011), and Robb and Robinson (2014) — who find no evidence that average entrepreneurs face severe financial constraints. However, once illiquidity is accounted for, the effective (net) collateral constraint is only about 44% of real capital value, consistent with Evans and Jovanovic (1989) and Cagetti and De Nardi (2006). This suggests that what appears empirically as &amp;lsquo;financial constraint&amp;rsquo; is partly a manifestation of capital illiquidity: banks lend less against entrepreneurial capital because its resale value is low, not primarily because of limited commitment. Second, empirical studies using regional variation in financial conditions to identify financial constraint effects may suffer from omitted variable bias, since resale prices of capital are also highly correlated with local financial conditions. Third, aggregate statistics such as startup rates and investment levels cannot distinguish between illiquidity shocks and financial constraint shocks; portfolio composition (the ratio of liquid to illiquid assets) is a more informative diagnostic.&lt;/p&gt;
&lt;h3 id="q16-what-is-the-papers-contribution-to-the-misallocation-literature-relative-to-hsieh-and-klenow-2009-asker-et-al-2014-and-midrigan-and-xu-2014"&gt;Q16. What is the paper&amp;rsquo;s contribution to the misallocation literature relative to Hsieh and Klenow (2009), Asker et al. (2014), and Midrigan and Xu (2014)?&lt;/h3&gt;
&lt;p&gt;Hsieh and Klenow (2009) and Asker et al. (2014) focus on the dispersion of log MRPK as a measure of misallocation, where adjustment costs (similar to fs and λ here) can generate observed dispersion without implying inefficiency. Midrigan and Xu (2014) focus on financial constraints (similar to ϕ) as the source of misallocation. The paper argues that these frameworks produce observationally equivalent outcomes in terms of log MRPK dispersion alone, making it impossible to distinguish between the two. The paper&amp;rsquo;s contribution is to show that the skewness of log ARPK and the asymmetric tail persistence are additional moments that can discriminate between the two types of frictions: negative skewness and left-tail dominance point to illiquidity frictions, while positive skewness and right-tail dominance point to financial frictions. This provides a new empirical diagnostic tool for decomposing sources of capital misallocation.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Average Revenue Product of Capital (ARPK)&lt;/strong&gt;: In the paper&amp;rsquo;s usage, ARPK = Y_it / K_{i,t-1}, the ratio of a firm&amp;rsquo;s real revenue to its beginning-of-period real capital stock, used as the primary measure of capital productivity. Log ARPK is residualized on two-digit NAICS industry fixed effects and time dummies before analysis, removing industry-level heterogeneity in capital shares and aggregate shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Partial irreversibility&lt;/strong&gt;: The friction arising from an asymmetry between the purchase price of new capital (normalized to 1) and the resale price of used capital (1 − λ for downsizing incumbents, and (1 − ζ)(1 − λ) for exiting entrepreneurs). This is modeled as a proportional transaction cost on capital sales and is interpreted as the difficulty of recouping original investment, analogous to a low resale value of used entrepreneurial equipment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wait-and-see attitude&lt;/strong&gt;: The behavioral response of entrepreneurs facing downside productivity shocks when capital is illiquid: rather than immediately downsizing or exiting upon a bad shock, they maintain larger-than-optimal firm sizes while waiting for conditions to improve. This is optimal because the transaction cost of selling capital makes the option of waiting (and possibly recovering) more valuable than the cost of operating an oversized firm.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Net collateral constraint (effective collateral parameter)&lt;/strong&gt;: Denoted ϕ̃ = (1 − λ)(1 − δ)ϕ, this is the fraction of entrepreneurial capital&amp;rsquo;s real value that can actually be pledged as collateral, after accounting for the reduced resale value from illiquidity (1 − λ) and physical depreciation (1 − δ). The paper distinguishes this from the formal limited-commitment parameter ϕ to show that observed financial constraints partly reflect capital illiquidity rather than contracting failures.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Options value effect&lt;/strong&gt;: The mechanism through which capital illiquidity distorts both the entry/exit decision and the intensive margin of investment. For downsizing incumbents, the put option value of capital (the option to sell it) falls when the resale price is low, inducing them to delay disinvestment. For potential entrants, the call option value of capital (the upside of entering) falls because losses upon exit are larger, raising the productivity signal threshold for entry. This is described as the primary distortion channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Span-of-control parameter (returns to scale, ν)&lt;/strong&gt;: The parameter ν ∈ (0,1) in the entrepreneurial production function y = z(k^{α_e} l^{1-α_e})^ν, capturing the extent to which managerial talent becomes diluted as firm size increases. The paper identifies ν = 0.79 (FULL) from the coefficient of a log-revenue on log-capital regression for employer firms, and shows that ν is the dominant determinant of the variance of capital income returns and hence the model&amp;rsquo;s ability to generate wealth inequality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption equivalent variation (CEV)&lt;/strong&gt;: The welfare metric used throughout the paper. For each household i, CEV µ_i is defined as the percentage increase in reference-economy consumption (or lifetime consumption stream) that makes the household indifferent between the reference economy and the economy of interest. Positive CEV means the new economy is preferred. Aggregate welfare is the distribution-weighted average of individual CEVs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Asymmetric persistence&lt;/strong&gt;: The empirical fact, documented in the KFS, that log ARPK shows higher autocorrelation at the bottom quintile (ρ₁ = 0.897) than at the top quintile (ρ₅ = 0.443), confirmed by both a conditional autocorrelation regression and a quintile transition matrix. This asymmetry is a key moment used to identify and distinguish illiquidity frictions (which produce left-tail persistence) from collateral constraints (which produce right-tail persistence).&lt;/p&gt;</description></item><item><title>Firm dynamics, monopsony, and aggregate productivity differences</title><link>https://macropaperwarehouse.com/papers/firm-dynamics-monopsony-and-aggregate-productivity-differences/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/firm-dynamics-monopsony-and-aggregate-productivity-differences/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; Firms are larger and grow faster over the life cycle in high-income countries, while labor markets in poorer countries are less competitive (employers hold more wage-setting power). The paper asks how important employer labor market power (monopsony) is for explaining cross-country differences in firm dynamics and aggregate productivity. The novelty is that beyond the standard static misallocation-of-workers channel, monopsony also distorts &lt;em&gt;selection into entrepreneurship&lt;/em&gt; and &lt;em&gt;productivity-enhancing technology adoption&lt;/em&gt;, potentially making the losses larger than prior static estimates suggest.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and setup.&lt;/strong&gt; Stylized facts come from the World Bank Enterprise Surveys (WBES), an establishment-level survey of non-agricultural, non-financial private firms with at least 5 full-time permanent employees, covering more than 90 countries from 2006 to 2021, merged with World Development Indicators GDP per capita (2017 constant USD). The estimation sample restricts to countries that ever had GDP per capita above 25,000 USD and to manufacturing firms with non-missing sales/workers/material/capital data, yielding 37,096 firm-year observations across 31 middle- and high-income countries (poorest: Kazakhstan, 19,615 USD in 2009; richest: Ireland, 91,791 USD in 2020). Local labor markets are defined as location-industry (2-digit ISIC v3.1) pairs. The model is a dynamic general-equilibrium neoclassical-monopsony model with occupational choice (entrepreneur vs. wage worker), endogenous productivity investment, and Card-et-al.-style taste-for-employer (amenity) differentiation that gives firms wage-setting power. It is calibrated to the Netherlands (GDP per capita 54,275 USD; median wage markdown 1.301, implying firm-level labor supply elasticity 3.318) via method of simulated moments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main quantitative findings.&lt;/strong&gt; Empirically, moving from poorer to richer countries in the sample, average firm age triples from 11 to nearly 30 years; annualized firm growth rises ~1.6 percentage points per year per doubling of GDP per capita; the share of firms doing R&amp;amp;D more than doubles (from ~15% to &amp;gt;40%); product innovation rises from 20% to 80% and process innovation from 20% to 50%; and median wage markdowns fall (from ~2.25 at 25,000 USD GDP per capita — workers paid ~55% below marginal product — to ~1.25 at 60,000 USD — paid 20-25% below). The calibrated model matches a right-skewed firm-size distribution, life-cycle growth, employer turnover, age distribution, and R&amp;amp;D share (sum of squared deviations between empirical and simulated moments = 1.7%). In counterfactuals raising the markdown from 1.2 to 3, average firm growth shrinks by more than half (from ~150% to ~50%), average firm size falls from ~60 to ~45 employees, the innovating share halves (from ~40% to ~25%), and average firm productivity is ~20% higher in competitive markets. Differences in wage markdown alone account for &lt;strong&gt;25%&lt;/strong&gt; of observed cross-country TFP variation (model TFP std dev 0.051 vs. data 0.201), and &lt;strong&gt;no less than 11%&lt;/strong&gt; across robustness checks. In a Netherlands-vs-Greece decomposition, about &lt;strong&gt;85%&lt;/strong&gt; of the model-implied TFP gap is attributable to lower technology adoption, ~9% to distorted selection into entrepreneurship, and ~6% to static employment reallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanisms and implications.&lt;/strong&gt; Labor market competition acts as a “skill-biased” force favoring high-productivity firms through three channels: (i) static labor reallocation toward high-productivity, low-amenity firms; (ii) improved selection into entrepreneurship (low-productivity high-amenity agents stop being able to profitably attract workers as ϵL rises); and (iii) higher returns to innovation. The policy implication is that raising labor market competition in less-developed economies could yield substantial productivity gains, and that prior static studies understate the cost of monopsony because they omit the dynamic investment/selection channels.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identificationcalibration-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification/calibration strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The model is calibrated to the Netherlands using a mix of externally set and internally estimated (method-of-simulated-moments) parameters. Externally: model period = 1 year; σν (Gumbel scale) normalized to 1; β = 0.961 (4% annual rate); δw = 0.025 (40-year working life); revenue elasticity of labor ξ = 0.333 (estimated via control function in Section 2); labor supply elasticity ϵL = 3.318 backed out from median markdown 1.301 via ϵL = 1/(µ−1). Six parameters {c_f, c_x, p_i, p_n, σ_z, σ_a} are estimated by MSM. The markdown itself is a key input and is estimated as the ratio of marginal revenue product of labor to wage, with revenue elasticity ξ from a standard control-function approach. Threats: the markdown estimate drives the whole quantitative exercise; the WBES sample is truncated at firms with ≥5 employees (biasing toward larger firms), addressed by re-estimating with imputed moments; and the cross-country counterfactual attributes all variation in ϵL to labor market power while holding all other parameters at Netherlands values, so other cross-country differences are not separately identified.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-mechanisms-and-how-are-they-distinguished-quantitatively"&gt;Q2. What are the three mechanisms and how are they distinguished quantitatively?&lt;/h3&gt;
&lt;p&gt;(1) Static labor allocation: lower competition raises marginal factor cost only for sufficiently high-productivity firms, reallocating employment toward less-productive, lower-paying employers. (2) Selection into entrepreneurship: when ϵL is low, amenities matter more for profits, letting low-productivity high-amenity agents profitably self-select into entrepreneurship. (3) Technology adoption: returns to innovation increase with ϵL, so weak competition lowers the share of firms investing. They are distinguished via a decomposition that sequentially fixes policy functions at benchmark levels: ~6% of the TFP loss is from employment allocation alone, ~85% from the distortion to innovation policy, and ~9% from distorted selection into entrepreneurship.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-across-firms-is-documented"&gt;Q3. What heterogeneity across firms is documented?&lt;/h3&gt;
&lt;p&gt;Firms differ in entrepreneurial productivity z and amenity a. Average revenue product of labor rises with productivity and falls with amenities, and this dispersion is much steeper under weak competition: the elasticity of APL with respect to productivity is 0.31 in the baseline (Netherlands) vs 0.79 in the counterfactual (Greece), and with respect to amenities -0.28 vs -0.81. High-productivity, low-amenity firms face the biggest barriers in less-competitive markets and stay inefficiently small; low-productivity, high-amenity firms are propped up. Innovation distortion is concentrated among high-productivity firms.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run-and-what-do-they-show"&gt;Q4. What robustness checks are run and what do they show?&lt;/h3&gt;
&lt;p&gt;Four main checks, each reported as the share of cross-country TFP variation explained (data std dev 0.201): (1) Productivity-amenity correlation — allowing entrants to draw correlated (z,a) with σ_za = 0.296 (matching Sockin 2024’s 0.622 wage-satisfaction correlation) lowers explained variation to ~15% (model std dev 0.030), because correlation reduces scope for reallocation. (2) Costs in terms of labor instead of final goods (per Klenow and Li 2025) gives ~22% (std dev 0.044). (3) Imputed firm-level moments covering all firms (not just ≥5 employees) gives ~14% (std dev 0.028). (4) Over-identified alternative identification using size/age/R&amp;amp;D shares and annualized growth gives ~11% (std dev 0.023). The headline range is therefore 25% baseline, no less than 11% across checks.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds on static monopsony cost estimates: Berger et al. (2022, eliminating US labor market power raises average wage 48%, welfare +6% of lifetime consumption); Armangüé-Jubert et al. (2025, labor market power explains 15% of GDP-per-capita gap over development); Deb et al. (2022, less competition lowered US low/high-skill wages 12% and 11%); Amodio et al. (2025b, eliminating monopsony in Peru raises earnings 26%); Bachmann et al. (2022, monopsony caused a 10% aggregate productivity loss in East Germany). Its contribution is to add the entrepreneurial-selection and innovation channels, yielding larger losses than static studies, and to bridge the monopsony-cost literature with the misallocation literature (Restuccia-Rogerson, Guner et al., Hsieh-Klenow).&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Raising labor market competition (higher firm-level labor supply elasticity) improves allocative efficiency, selection into entrepreneurship, and innovation, raising firm growth and aggregate productivity. Scope conditions: the quantitative results apply to middle- and high-income countries (sample restricted to those ever above 25,000 USD GDP per capita); the 25% headline depends on the assumption that initial productivity and amenities are independent (falls to ~15% under positive correlation); and the decomposition attributing 85% to innovation is specific to the Netherlands-vs-Greece comparison. The model treats labor supply elasticity differences as the sole varying parameter, so the counterfactuals isolate the labor-market-power channel rather than reproducing total cross-country income gaps.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-netherlands-vs-greece-comparison-specifically"&gt;Q7. What is the Netherlands-vs-Greece comparison specifically?&lt;/h3&gt;
&lt;p&gt;Greece has roughly half the GDP per capita of the Netherlands (29,000 vs 54,000 USD) and much weaker competition (wage markdown 2.623 vs 1.301, labor supply elasticity 0.616 vs 3.318). In the Greece counterfactual, average firm size is 26 vs 59 employees, life-cycle growth 84.5% vs 153%, average age 22.5 vs 30 years, and R&amp;amp;D investing share 18% vs 41%. Labor market competition differences explain 29% of the firm-size gap, 27% of the firm-age gap, and 74% of the R&amp;amp;D-share gap between the two countries.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-model-get-right-that-was-not-targeted"&gt;Q8. What does the model get right that was not targeted?&lt;/h3&gt;
&lt;p&gt;The firm size and age distributions are not targeted yet are matched: in the data ~57.6% of firms have &amp;lt;20 employees and ~6.2% have &amp;gt;100; ~60% of firms are under 30 years old and ~10% over 60. The estimated parameters imply investing firms are 15% more likely to grow (p_i=0.649 vs p_n=0.499); innovation and operating costs equal ~43% and ~8% of average incumbent profits respectively; standard errors are small, indicating informative moments.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Firm Heterogeneity, Market Power and Macroeconomic Fragility</title><link>https://macropaperwarehouse.com/papers/firm-heterogeneity-market-power-and-macroeconomic-fragility/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/firm-heterogeneity-market-power-and-macroeconomic-fragility/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Ferrari and Queirós ask why US recoveries have become progressively slower and argue that rising firm heterogeneity and market power — well-documented long-run trends — can substantially increase the probability that a moderate aggregate shock triggers a quasi-permanent slump rather than a transitory recession. They call this probability macroeconomic fragility.&lt;/p&gt;
&lt;p&gt;The theoretical framework is an RBC model with oligopolistic (Cournot) competition, endogenous firm entry, and elastic capital and labor supply (GHH preferences). The economy consists of many product markets; within each market, firms with heterogeneous idiosyncratic TFP compete in quantities, with the marginal firm earning zero net profit. A central complementarity drives the results: more competition raises factor shares and factor prices, which expands factor supply, which in turn allows more firms to enter, sustaining high competition. This complementarity can generate multiple stochastic steady-states — a high-competition, high-output regime and a low-competition, low-output regime.&lt;/p&gt;
&lt;p&gt;Two forces increase fragility by shrinking the basin of attraction around the high steady-state. First, a mean-preserving spread (MPS) in idiosyncratic TFP: the dominant firm expands market share, factor shares fall (market-power effect), the factor price index drops, and smaller firms approach their exit threshold — requiring only a smaller shock to trigger cascading exit. Second, rising fixed production costs: the unstable steady-state shifts toward the high steady-state, narrowing the gap and making downward transitions more likely.&lt;/p&gt;
&lt;p&gt;The model is calibrated three times — to match COMPUSTAT moments in 1975, 1990, and 2007 — varying only the log-normal standard deviation of idiosyncratic productivity (λ = 0.182, 0.213, 0.232) and the fixed cost parameter (c × 10⁻³ = 0.351, 0.691, 0.751). The fixed-to-total-cost ratio in COMPUSTAT rises from 21.9% in 1975 to 31.7% in 1990 to 36.9% in 2007; the standard deviation of log revenues rises from 1.59 to 1.91 to 2.04.&lt;/p&gt;
&lt;p&gt;The quantitative results are stark. The 1975 economy has a unimodal ergodic distribution (one stable steady-state); the 1990 and 2007 economies are bimodal (two stable steady-states). When subjected to the same TFP shock sequence (εt = −σε for four quarters), output falls 4.0% after five quarters in the 1975 economy, 5.1% in 1990, and 5.9% in 2007; after 100 quarters, the 2007 economy remains 6.3% below pre-shock output, against 3.0% for 1990 and 1.3% for 1975. For a larger shock (εt = −2σε for six quarters), only the 2007 economy transitions permanently to the low steady-state, with output 12.5% below trend after 100 quarters. The minimum shock required to trigger a downward transition is 6.84σε for the 1990 economy but only 1.62σε for the 2007 economy. In Monte Carlo simulations, the probability of a recession exceeding 10% of output over a 40-quarter window is 1.7% in 1975, 12.4% in 1990, and 19.6% in 2007. In expectation, the 2007 economy experiences such a recession every 70 years, the 1990 economy every 95 years, and the 1975 economy every 380 years.&lt;/p&gt;
&lt;p&gt;Applying the 2008–09 TFP shocks to the 2007-calibrated model generates a persistent deviation from trend: output is 12.1% below trend by 2019, investment 14.4% below, and hours 9.8% below — closely matching the data (14.2%, 14.7%, and 5.5% respectively). The same shocks applied to the 1975 and 1990 economies produce no permanent transition; by 2040 the 1975 (1990) economy is only 1.5% (4.7%) below trend.&lt;/p&gt;
&lt;p&gt;Cross-industry evidence corroborates the mechanism. Using US Census and BLS data on 791 six-digit NAICS industries, the authors find that a 1 percentage point higher pre-crisis four-firm concentration ratio (CR4) in 2007 is associated with 1.8–1.9 percentage points lower employment growth, 2–3 percentage points lower net firm entry, and a larger decline in the labor share between 2007 and 2016. These qualitative and quantitative patterns are matched by simulated cross-industry regressions from the model.&lt;/p&gt;
&lt;p&gt;On policy, an entry subsidy that eliminates fixed-cost barriers for the approximately 11.8% of markets with positive fixed costs can prevent downward transitions and yields a welfare gain of roughly 10% in consumption-equivalent terms in the 2007 economy. A revenue subsidy applied to all firms achieves welfare gains between 30% and 50% for a 20% subsidy rate, acting as a steady-state selection device by shifting probability mass from the low to the high competition regime. These gains are nonlinear: even a 5% revenue subsidy yields roughly a 20% welfare gain in the 2007 economy. The gains are in line with Edmond et al. (2023), who find welfare costs of markups up to 50%.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-models-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the model&amp;rsquo;s identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper is primarily theoretical and quantitative rather than identification-based in the econometric sense. The causal claim — that rising firm heterogeneity and fixed costs increase macroeconomic fragility — comes from two sources: (1) analytic comparative statics (Propositions 4–6) that formally show fragility rises with a mean-preserving spread on TFP or with fixed costs, and (2) calibration counterfactuals where the 1975, 1990, and 2007 economies face the same shock sequence but differ only in λ and c. The cross-industry regressions are reduced-form and subject to standard endogeneity concerns — pre-crisis concentration could be correlated with industry-specific demand shocks coinciding with 2008. The authors partially address this by including pre-crisis growth trends as controls and sector fixed effects, but do not use an instrumental variable for concentration.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-core-mechanism-linking-firm-heterogeneity-to-fragility-and-how-is-it-distinguished-from-steady-state-multiplicity"&gt;Q2. What is the core mechanism linking firm heterogeneity to fragility, and how is it distinguished from steady-state multiplicity?&lt;/h3&gt;
&lt;p&gt;The mechanism runs through factor markets. When idiosyncratic TFP dispersion rises (MPS), the dominant firm expands market share and charges a higher markup, depressing the aggregate factor share (Proposition 4). This reduces the factor price index and real wages, contracting labor supply. Marginal firms, already earning near-zero profits, move closer to their exit threshold. A smaller aggregate shock suffices to push them out, triggering cascading exit, a further collapse in competition, a further fall in factor prices, and a self-reinforcing transition to the low steady-state. Fragility is distinct from multiplicity: the existence of two steady-states is a necessary but not sufficient condition for fragility. Fragility specifically measures the size of the basin of attraction around the high steady-state from below — how large a shock is needed to trigger a downward transition. An economy can have two steady-states but be highly resilient if the basin is wide.&lt;/p&gt;
&lt;h3 id="q3-what-roles-do-the-three-model-channels-endogenous-market-structure-oligopolistic-markups-elastic-factor-supply-play-quantitatively"&gt;Q3. What roles do the three model channels (endogenous market structure, oligopolistic markups, elastic factor supply) play quantitatively?&lt;/h3&gt;
&lt;p&gt;The authors isolate each channel by shutting it down one at a time and comparing output volatility (Table 8). In the baseline, the standard deviation of log output is 0.063 and autocorrelation is 0.975. Fixing the number of firms (removing the endogenous market structure channel, leaving only elastic factor supply) reduces output standard deviation to 0.035, accounting for 55% of baseline volatility. Replacing oligopoly with monopolistic competition (constant markups, love-for-variety active) recovers 0.049 — approximately 78% of baseline — implying the endogenous markup channel accounts for about one-fourth of total amplification. The love-for-variety channel accounts for another approximately one-fourth. Crucially, all three alternative models exhibit unimodal ergodic distributions, confirming that all three channels are jointly required to generate steady-state multiplicity and the model&amp;rsquo;s nonlinear amplification.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented-and-how-does-it-motivate-the-models-calibration"&gt;Q4. What heterogeneity is documented and how does it motivate the model&amp;rsquo;s calibration?&lt;/h3&gt;
&lt;p&gt;Rising US firm heterogeneity is documented along three dimensions: (1) standard deviation of log revenues (sales) for COMPUSTAT firms, rising from 1.59 in 1975 to 1.91 in 1990 to 2.04 in 2007; (2) the average ratio of fixed (SG&amp;amp;A) to total costs (fixed + COGS), rising from 21.9% in 1975 to 31.7% in 1990 to 36.9% in 2007; (3) sales-weighted average markups for public firms rising from 1.28 in 1975 to 1.37 in 1990 to 1.46 in 2007 (from De Loecker et al., 2020). These moments are the calibration targets for the time-varying parameters λ and c. The structural parameters (elasticities of substitution σI = 1.46 and σG = 11.50) are time-invariant and calibrated jointly to the markup levels across the three years.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-papers-account-of-the-great-recession-differ-from-other-slow-recovery-theories"&gt;Q5. How does the paper&amp;rsquo;s account of the Great Recession differ from other slow-recovery theories?&lt;/h3&gt;
&lt;p&gt;Most related theories attribute slow recovery to (1) the zero lower bound on interest rates and constrained monetary policy (Christiano et al., 2015; Eggertsson et al., 2019; Guerrieri and Lorenzoni, 2017), (2) endogenous TFP decay through R&amp;amp;D decisions (Anzoategui et al., 2019; Bianchi et al., 2019; Queralto, 2020), or (3) declining firm entry per se (Clementi and Palazzo, 2016). Ferrari and Queirós instead argue the 2008 shock was not unusually large — the same shock does not cause a permanent transition in the 1975 or 1990 economies — but rather that the US economy had become structurally more fragile over the preceding decades due to rising concentration and fixed costs. The closest related model is Schaal and Taschereau-Dumouchel (2018), who also use coordination failures among oligopolistic firms to generate multiple steady-states. The key contribution of Ferrari and Queirós relative to that work is the explicit role of cross-sectional firm heterogeneity in determining the probability of transitions, and the empirical documentation that rising heterogeneity preceded the crisis.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-cross-industry-empirical-results-in-detail"&gt;Q6. What are the cross-industry empirical results in detail?&lt;/h3&gt;
&lt;p&gt;The dataset covers 791 six-digit NAICS industries from the US Census, SUSB, and BLS, with the concentration variable defined as CR4/CR50 (top-4 share scaled by top-50 share). Key results: (1) Employment: a 1 pp higher CR4/CR50 in 2007 is associated with 1.77–1.89 pp lower annualized employment growth between 2007 and 2016 (significant at 1%); robust to controlling for pre-crisis employment trends and sector fixed effects. (2) Payroll: similarly negative coefficient of approximately −0.041 on log payroll growth. (3) Net firm entry: a 1 pp higher concentration is associated with 2–3 pp lower post-crisis net entry. (4) Labor share: a negative relationship between 2007 concentration and the change in industry labor share between 2008 and 2016 (coefficient approximately −0.031, significant at 10%). All results are mirrored qualitatively and quantitatively in simulated cross-industry regressions from the model: concentrated markets in the model experience 5.4% larger drops in employment, 3.7% higher firm exit, and 1.1% larger decline in labor share.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-and-extensions-are-reported"&gt;Q7. What robustness checks and extensions are reported?&lt;/h3&gt;
&lt;p&gt;Several extensions and checks are noted: (1) An alternative shock — fluctuations in the fraction of industries with positive fixed costs (xc) rather than TFP shocks — also replicates the medium-run behavior of the US economy, with output falling roughly 15% on impact and remaining −18% below trend in the long run; the cross-sectional implications are unchanged. (2) The 1990 recession counterfactual: applying 1990–1991 recession shocks to the 1990 economy produces no permanent transition, but the same shocks applied to the 2007 economy do, confirming that fragility rather than shock size drove the 2008 outcome. (3) Factor-price-dependent fixed costs: Ferrari and Queirós (2022) show steady-state multiplicity is preserved when fixed costs depend on factor prices. (4) Varying M: results are unchanged for M = 50 and M = 100 potential firms per market. (5) The cross-industry regressions are robust across multiple specifications including controls for the number of firms in 2007, pre-crisis growth, and sector fixed effects (Appendix B.7).&lt;/p&gt;
&lt;h3 id="q8-what-are-the-models-aggregate-predictions-for-labor-share-profit-share-and-markups-post-2008-and-how-do-they-compare-to-data"&gt;Q8. What are the model&amp;rsquo;s aggregate predictions for labor share, profit share, and markups post-2008, and how do they compare to data?&lt;/h3&gt;
&lt;p&gt;Between 2007 and 2016, the model predicts (Table 9): a 0.4 pp decline in the aggregate labor share (data: −2.9 pp decline; the model explains approximately 14% of the total decline, or 17% accounting for the pre-crisis trend); a 0.9 pp increase in the profit share (data: +3.2 pp; model explains 30% of the trend deviation); a 3.7 point increase in sales-weighted markups for COMPUSTAT firms (data: +14.2 points; model explains 26% of the total increase and 58% of the deviation from the pre-crisis trend). The model also predicts a persistent fall in the number of firms in markets with positive fixed costs of 13.4 log points, compared to the observed 15.1 log point decline in the number of US firms with at least one employee. The model understates the magnitude of all these changes, but correctly signs and persists them, consistent with its role in providing a partial explanation.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper studies two interventions: (1) An entry subsidy covering a fraction τf of fixed costs for markets with c &amp;gt; 0 (roughly 11.8% of all markets). A 5% entry subsidy is sufficient to eliminate the welfare costs associated with multiplicity in the 2007 economy; higher subsidies improve allocation within the high steady-state. An entry subsidy large enough to prevent downward transitions yields approximately 10% welfare gain in consumption-equivalent terms. The effect is highly targeted and quantitatively modest per-dollar because only 11.8% of markets are affected. (2) A revenue subsidy τR applied to all firms, equivalent to a fraction of revenues subsidized. Even a 5% revenue subsidy generates approximately 20% welfare gain in the 2007 economy by shifting probability mass from the low to the high competition regime. A 20% revenue subsidy yields gains between 30% and 50% in the 1990 and 2007 economies. The gains are nonlinear in the economies with multiple steady-states, and much smaller in the 1975 economy, which has only one steady-state. A revenue tax has asymmetric large welfare costs in the 1990 economy (which has large output gaps between regimes) relative to the 2007 economy (smaller gap but higher transition probability). The welfare gains come from two sources: reducing static markup distortions and reducing the dynamic cost of transitions (quasi-permanent slumps).&lt;/p&gt;
&lt;h3 id="q10-what-caveats-and-limitations-does-the-paper-acknowledge"&gt;Q10. What caveats and limitations does the paper acknowledge?&lt;/h3&gt;
&lt;p&gt;The authors are explicit about several limitations. First, the model lacks sunk entry costs: all entry decisions are static, which may understate hysteresis and overstate the responsiveness of exit to shocks. Introducing sunk costs with oligopolistic competition poses a computational challenge (20^10 partial equilibria for M=20 and 10 values per firm). Second, idiosyncratic productivities are time-invariant, ruling out Schumpeterian creative destruction within the model. Third, the model features only one-sided market power (product markets only); recent work on labor-market oligopsony could interact with the mechanism. Fourth, the model has no monetary policy channel; the interaction between monetary policy and endogenous market structure is left for future research. Fifth, the model explains only a fraction of the observed post-2008 declines in the labor share (14–17%), profit share (30%), and markup levels (26% of total, 58% of trend deviation), suggesting complementary mechanisms are at work.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-characterize-the-relationship-between-the-great-moderation-and-rising-fragility"&gt;Q11. How does the paper characterize the relationship between the Great Moderation and rising fragility?&lt;/h3&gt;
&lt;p&gt;The paper directly addresses the apparent tension between the Great Moderation (declining aggregate output volatility from 1980 to 2007) and the model&amp;rsquo;s prediction of rising fragility over the same period. The resolution is that aggregate output volatility is the product of exogenous TFP shock volatility and endogenous amplification. If exogenous TFP shocks became less volatile over time (a plausible claim, attributed to demographic shifts and the rising share of low-volatility service industries), then aggregate volatility could have declined even as endogenous amplification increased. Fragility, as defined in the paper, is about the probability of large discrete transitions, not about the variance of the ergodic distribution around a single steady-state. An economy can exhibit lower volatility on average while being more prone to catastrophic (quasi-permanent) downturns.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Macroeconomic Fragility&lt;/strong&gt;: The probability of long slumps, formally measured as the proximity of the high stable steady-state to the preceding unstable steady-state (χ = KU/K*). A higher χ means a smaller negative shock is sufficient to trigger a permanent downward transition. Fragility is distinct from steady-state multiplicity (which is necessary but not sufficient) and distinct from stability (which measures the full basin of attraction in both directions).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Competition-Factor Supply Complementarity&lt;/strong&gt;: The positive feedback loop through which more competitive product markets generate higher factor shares and factor prices, inducing higher labor and capital supply, which in turn allows more firms to enter and compete. This complementarity is the structural foundation for multiple steady-states in the model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mean-Preserving Spread (MPS) on Idiosyncratic TFP&lt;/strong&gt;: An increase in cross-firm productivity dispersion that leaves the average unchanged. In the model&amp;rsquo;s context, an MPS raises aggregate TFP (allocative efficiency effect as output shifts to high-productivity firms) but lowers the factor share and factor price index (market power effect as concentration increases), and shrinks the stable steady-state&amp;rsquo;s capital level while raising the unstable steady-state&amp;rsquo;s capital level — thereby increasing fragility.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Low Competition Trap&lt;/strong&gt;: The low stable steady-state in which the economy becomes trapped following a transition from the high steady-state. Characterized by fewer active firms, higher markups, lower factor shares, lower capital stock, and lower output relative to the high steady-state. In the 2007 calibration, the two steady-states are approximately 21% apart in output terms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous Market Structure&lt;/strong&gt;: The model feature whereby the number of active firms in each product market is determined endogenously by a free-entry condition: the marginal firm exactly breaks even (net profits equal fixed costs). This makes the number of firms — and hence the degree of competition, markups, and factor shares — respond endogenously to aggregate shocks and capital accumulation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Factor Price Index (Θ)&lt;/strong&gt;: A composite of the wage and rental rate representing the minimum cost of one unit of output for a firm with unit productivity. In the model, Θ equals the product of the aggregate factor share and aggregate TFP. It serves as a sufficient statistic for both factor prices and the competitive environment, decreasing with higher firm heterogeneity (via lower factor shares) and increasing with more firms (via higher competition).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Great Deviation&lt;/strong&gt;: The paper&amp;rsquo;s term (following Hall, 2011) for the persistent and widening gap between actual US output and its pre-2007 trend following the 2008–09 recession. In the data, real GDP per capita was 14.2% below its pre-crisis trend as of 2019Q1, a deviation far larger and more persistent than in any prior postwar recession. The paper&amp;rsquo;s model rationalizes this as a transition to the low steady-state.&lt;/p&gt;</description></item><item><title>Global Value Chains and Labor Standards: The Race-to-the-Bottom Problem</title><link>https://macropaperwarehouse.com/papers/global-value-chains-and-labor-standards-the-race-to-the-bottom-problem/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/global-value-chains-and-labor-standards-the-race-to-the-bottom-problem/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Im and McLaren (2025) ask whether globalization induces governments to weaken labor standards for workers — the so-called &amp;ldquo;race to the bottom&amp;rdquo; (RTB) hypothesis. The question has high stakes: advocates point to events such as the 1,136-worker Rana Plaza factory collapse in Bangladesh (2013) and to India&amp;rsquo;s deregulation campaign after 2014 (associated with approximately 6,500 workplace deaths in 2015–2020) as evidence that competition for global capital systematically erodes safety and working conditions. The paper builds a stylized many-country equilibrium model of labor-market integration adapted from the Grossman and Rossi-Hansberg (2008) tasks framework. Output requires a continuum of tasks z in [0,1], performable in any of N countries; labor requirements per task follow a Weibull distribution (shape parameter nu &amp;gt; 0), independently across tasks and countries. Working conditions (kappa_i) enter the cost function multiplicatively — better conditions reduce worker productivity at the relevant margin. Utility is separable in wages and conditions with both components strictly concave, and Assumption 1 (x&lt;em&gt;xi&amp;rsquo;(x) and x&lt;/em&gt;mu&amp;rsquo;(x) strictly decreasing) ensures conditions are normal goods and second-order conditions hold. The unregulated equilibrium task allocation is equivalent to CES cost minimization with elasticity of substitution 1/(1-rho) &amp;gt; 1, rho = nu/(1+nu). Governments set minimum standards non-cooperatively in Nash equilibrium.\n\nThe paper&amp;rsquo;s results fall into two conceptually distinct categories. &amp;ldquo;Globalization in the large&amp;rdquo; (autarky vs. open economy): whether standards are market-determined or government-set, integrating two previously autarkic countries raises labor standards in both (Proposition 1). Under autarky, market and government-optimal conditions coincide — all costs of better standards are borne domestically. Under trade, wages rise (income channel: conditions are a normal good), and governments gain a terms-of-trade incentive: tightening kappa_i makes domestic effective labor scarcer and shifts part of the cost onto foreign consumers, inducing government standards to strictly exceed market standards. Formally, for each country i: autarky level = market level under autarky &amp;lt; market level under integration &amp;lt; government level under integration.\n\n&amp;quot;Globalization at the margin&amp;quot; with symmetric countries (Proposition 2): as more identical countries join (N increasing), both market-set and government-set standards rise monotonically. The terms-of-trade motive does not vanish because each country specializes in an increasingly narrow value-chain slice, retaining market power regardless of N. Government standards exceed market standards for every N &amp;gt;= 2 and grow strictly with N — a race to the top — and are shown to be above the social optimum because each country externally imposes part of its improvement costs on others.\n\n&amp;quot;Globalization at the margin&amp;quot; with a North-South structure (Proposition 3): when Southern host countries (i = 2,&amp;hellip;,N) have perfectly correlated productivity draws (close substitutes for one another), the result reverses for N &amp;gt; 2. Integration of two countries initially raises Southern standards via both channels. But as additional similar Southern competitors join, competition depresses Southern wages and erodes both the income-based demand for better conditions and the terms-of-trade motive (unilateral tightening redirects demand to competitors without cost-shifting benefit). Both market and government standards fall monotonically as N rises beyond 2. As N approaches infinity, both converge to autarky levels. Critically, however, for any finite N, Southern standards remain strictly above their autarky levels — the race to the bottom, even when operative, never fully materializes while integration is incomplete. The efficiency implication is counter-intuitive: government-set standards are inefficiently strict under GVCs because each country over-provides standards by externalizing costs onto trading partners.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-models-formal-structure-and-how-does-it-generate-tractable-results"&gt;Q1. What is the model&amp;rsquo;s formal structure and how does it generate tractable results?&lt;/h3&gt;
&lt;p&gt;The model adapts Grossman and Rossi-Hansberg (2008). Output requires a unit measure of tasks; labor requirement for task z in country i is A_i * a^i_z, where A_i = bar_A_i * kappa_i, so working conditions raise unit labor costs. Each a^i_z is drawn Weibull(nu, 1) independently. A result (adapted from Anderson et al. 1987, applied by Artuç and McLaren 2015) is that the cost-minimizing task allocation is equivalent to minimizing cost with a CES aggregate of national effective labor supplies, with elasticity of substitution 1/(1-rho) and rho = nu/(1+nu). This reduces the multi-dimensional problem to a standard CES factor-demand problem, yielding closed-form wage equations and tractable Nash equilibrium characterizations.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-channels-driving-globalization-in-the-large-raising-standards-above-autarky"&gt;Q2. What are the two channels driving &amp;lsquo;globalization in the large&amp;rsquo; raising standards above autarky?&lt;/h3&gt;
&lt;p&gt;Two reinforcing channels. First, the income channel: integration raises real wages (gains from specialization), and since working conditions are a normal good under Assumption 1 (utility sufficiently concave), demand for better conditions rises. Second, the terms-of-trade channel: tightening kappa_i makes domestic effective labor more expensive and scarcer; part of the resulting cost increase is borne by foreign consumers and workers via the unit cost identity rather than solely by domestic workers. This cost-shifting gives governments an incentive to tighten standards beyond what the unregulated market sets. The mechanism is formally analogous to the policy externalities in Bagwell and Staiger (2001) and the terms-of-trade motive in Chau and Kanbur (2006), though the latter has no value chains.&lt;/p&gt;
&lt;h3 id="q3-why-does-the-terms-of-trade-motive-for-over-regulation-persist-even-as-the-number-of-symmetric-countries-approaches-infinity"&gt;Q3. Why does the terms-of-trade motive for over-regulation persist even as the number of symmetric countries approaches infinity?&lt;/h3&gt;
&lt;p&gt;As more countries join, each specializes in an increasingly narrow slice of the value chain in which it has comparative advantage. This deepening specialization preserves market power: the wage derivative dw_1/d_kappa_1 converges to a limit proportional to rho*w/kappa (strictly greater than the pure autarky productivity effect -w/kappa) rather than to zero. So even in the limit with infinitely many symmetric countries, each country retains some terms-of-trade gain from tightening its standard, and government standards keep rising above market standards.&lt;/p&gt;
&lt;h3 id="q4-under-what-precise-conditions-does-the-race-to-the-bottom-result-hold"&gt;Q4. Under what precise conditions does the race-to-the-bottom result hold?&lt;/h3&gt;
&lt;p&gt;The RTB result (Proposition 3) requires that competing host countries be close substitutes for one another. The paper operationalizes this with the extreme case of perfectly correlated productivity draws across Southern countries (a^i_z = a^2_z for all i &amp;gt;= 2 and all tasks z). Under this structure, as N increases from 2 onward, Southern market and government standards fall monotonically toward autarky levels. The mechanism: competition among near-identical countries means unilateral tightening of kappa_2 redirects Northern demand to competitors without generating a terms-of-trade gain for Country 2, so the wage falls and conditions deteriorate. The RTB thus requires high substitutability among competitors, not just trade openness.&lt;/p&gt;
&lt;h3 id="q5-does-the-race-to-the-bottom-ever-drive-standards-below-autarky-levels"&gt;Q5. Does the race to the bottom ever drive standards below autarky levels?&lt;/h3&gt;
&lt;p&gt;No. Proposition 3 parts (i) and (ii) establish that for any finite N &amp;gt;= 2, both market-set and government-set standards in Southern countries remain strictly above their autarky levels. The race is toward (but never below) the autarky benchmark. Only in the limit as N approaches infinity do standards converge to the autarky level (Proposition 3, part iii). For any realistic finite degree of globalization, even the worst-case RTB scenario leaves standards strictly above autarky.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-efficiency-implication-of-nash-equilibrium-government-set-standards"&gt;Q6. What is the efficiency implication of Nash equilibrium government-set standards?&lt;/h3&gt;
&lt;p&gt;Government-set standards under GVCs are inefficiently strict. Each government maximizes domestic welfare ignoring the cost its tightening imposes on foreign consumers and workers. Because tightening kappa_i raises costs partly borne abroad, each government over-provides standards relative to the global social optimum. This is a race to the top that generates a negative international externality — the mirror image of the usual RTB externality. The implication is that international coordination, if it occurred, would likely reduce Nash equilibrium standards toward the optimum, not raise them further.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-papers-setting-differ-from-prior-theoretical-work-on-the-race-to-the-bottom"&gt;Q7. How does the paper&amp;rsquo;s setting differ from prior theoretical work on the race to the bottom?&lt;/h3&gt;
&lt;p&gt;Prior RTB models (Chau and Kanbur 2006; Felbermayr et al. 2012; Chen and Dar-Brodeur 2020) model countries competing for export markets — competing to sell goods to a common importer — rather than competing to host tasks in global value chains. The current paper frames globalization as an increase in the number of countries that can supply tasks to a common production process, a qualitatively different competitive margin. Prior work also largely takes the degree of globalization as fixed, while this paper explicitly traces out effects as N changes. The distinction between similar versus different competitors as a determinant of the direction of the RTB is also new. The companion paper Im and McLaren (NBER WP 31363) extends the framework to collective-bargaining rights with an empirical component.&lt;/p&gt;
&lt;h3 id="q8-what-heterogeneity-is-documented-and-what-does-it-imply"&gt;Q8. What heterogeneity is documented and what does it imply?&lt;/h3&gt;
&lt;p&gt;The paper develops two polar cases of country heterogeneity: (1) symmetric countries with independent productivity draws — produces a race to the top as N rises; (2) North-South structure with correlated (identical) Southern productivity draws — produces a race to the bottom as N rises beyond 2. The contrast is the central result: the direction of the marginal effect of globalization on standards depends on the degree of substitutability among competing host countries. The authors connect this to observed patterns — Korean firms relocating only to East Asian affiliates (similar countries) when domestic minimum wages rose, and Chan and Ross (2003) noting that competition is &amp;lsquo;most vicious not between North and South, but among nations of the South.&amp;rsquo;&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The core implication is that trade restrictions justified by RTB concerns lack general theoretical support — globalization relative to autarky always raises standards. However, the model validates a targeted RTB concern: when a country faces competition from many similar low-wage countries (e.g., Mexico competing with China in labor-intensive sectors), standards can erode relative to the peak reached under limited integration. The appropriate response in that case is to integrate with structurally different partners (as Mexico did via NAFTA with the US) rather than restrict trade. Since Nash equilibrium standards already exceed the global optimum, international agreements that ratchet standards up further could be welfare-reducing. The paper explicitly cautions that causation is hard to establish in the Mexico-China-NAFTA example, treating it as suggestive illustration rather than proof.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-main-limitations-and-threats-to-the-conclusions"&gt;Q10. What are the main limitations and threats to the conclusions?&lt;/h3&gt;
&lt;p&gt;The paper is entirely theoretical; no empirical test is conducted for working conditions (the authors cite data scarcity as the reason, having a companion empirical paper on collective-bargaining rights instead). Key assumptions include: (a) Weibull, independent task-productivity draws (ensure tractability but are untested); (b) working conditions always reduce productivity at the margin (rules out the many cases where safety improvements also raise output — e.g., Alfaro-Ureña et al. 2021 find no productivity effect of responsible sourcing in Costa Rica, suggesting the trade-off assumption is plausible but not universal); (c) citizen activism, which empirically affects labor standards (Harrison and Scorse 2010; Koenig and Poncet 2019, 2022), is abstracted away; (d) the model has a single final good and no intermediate goods trade beyond the task-allocation interpretation, limiting applicability to multi-sector settings.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Labor standards (kappa_i)&lt;/strong&gt;: In the paper&amp;rsquo;s specific sense, the quality of working conditions that (i) raise worker utility holding wages fixed and (ii) increase unit labor costs for employers. Explicitly restricted to improvements that involve a trade-off — e.g., safety provisions, clean bathrooms, break times — excluding complementary improvements that raise both utility and productivity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Globalization in the large&lt;/strong&gt;: The paper&amp;rsquo;s term for the comparison of any open-economy equilibrium (N &amp;gt;= 2 countries integrated) against autarky. Result: labor standards are always strictly higher in the open economy whether market-set or government-set, because income rises and the terms-of-trade motive activates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Globalization at the margin&lt;/strong&gt;: The paper&amp;rsquo;s term for the effect on labor standards of adding one more country to an already-integrated economy (increasing N by 1). This effect is ambiguous: it raises standards when new entrants are dissimilar (symmetric model) and lowers them when new entrants are similar (North-South model).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Terms-of-trade effect (labor-standards channel)&lt;/strong&gt;: The mechanism by which tightening a country&amp;rsquo;s labor standard (raising kappa_i) reduces domestic effective labor supply, raises the relative price of domestic tasks, and shifts part of the cost improvement onto foreign consumers and workers. This creates an incentive for governments to set standards above the market level and above the global social optimum — producing standards that are too strict from an efficiency standpoint.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Normal good (working conditions)&lt;/strong&gt;: The property implied by Assumption 1 (both x&lt;em&gt;xi&amp;rsquo;(x) and x&lt;/em&gt;mu&amp;rsquo;(x) strictly decreasing in x) that workers&amp;rsquo; marginal valuation of working conditions relative to wages is higher at higher income levels. This ensures that any source of income gains — including gains from trade — mechanically raises equilibrium demand for better working conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Race to the top&lt;/strong&gt;: The paper&amp;rsquo;s characterization of the symmetric-countries equilibrium: as N increases, both market-set and government-set labor standards rise monotonically, because market power persists through value-chain specialization and the terms-of-trade motive remains strong. Government standards also exceed the social optimum, making this over-regulation an externality imposed on trading partners.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Race to the bottom (conditional)&lt;/strong&gt;: The result in the North-South model where additional similar Southern host countries erode Southern labor standards as N rises beyond 2. The race is toward autarky levels but never below them for finite N. The RTB requires high substitutability among competing host countries and does not hold as a general consequence of globalization.&lt;/p&gt;</description></item><item><title>Health Sector Structural Change</title><link>https://macropaperwarehouse.com/papers/health-sector-structural-change/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/health-sector-structural-change/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;The paper investigates why the U.S. health-services sector has simultaneously experienced a tripling of relative prices since 1948 and a rise in the personal consumption expenditure (PCE) share from under 5% in the late 1940s to 19.6% by 2022. The authors attribute this structural transformation to three candidate drivers: (1) rising relative health-sector markups, (2) unbalanced technological change (differential TFP growth rates across sectors), and (3) changes to the composition of demand from population aging and improving health-investment efficiency.&lt;/p&gt;
&lt;p&gt;The paper proceeds in two stages. First, a growth-accounting decomposition uses a two-sector Dixit-Stiglitz monopolistically competitive model to identify growth rates of relative markups and health-sector TFP directly from sectoral input/output data (NIPA PCE, BEA Fixed Asset Tables, Penn World Tables). The non-health capital intensity is set at 0.40 (from Horenstein and Santos 2019) and health-sector capital intensity at 0.26 (from Donahoe 2000). Because health is more labor-intensive than non-health (alpha_h &amp;lt; alpha_c), GE effects on input prices actually dampen relative price growth. In the baseline decomposition (1954-2019), average annual relative markup growth is estimated at 1.6%, with cumulative growth of approximately 186%. When allowing for a time-varying non-health labor share, relative markup growth rises to 2.0% annually and 255% cumulatively. Average annual health-sector TFP growth is 0.3% (baseline) and 0.2% (time-varying labor share), compared to 0.7% for the non-health sector per Penn World Tables. If no markup growth is assumed, the implied health-sector TFP growth falls to -1.3% annually, implying a 56.3% cumulative decline from 1954 to 2019, which the authors regard as implausible in light of observed healthcare advances. Across all four decomposition exercises, GE effects consistently dampen rather than amplify relative price growth, indicating that demand-side composition shifts from aging play at most a minor role in driving prices.&lt;/p&gt;
&lt;p&gt;Second, the paper builds and calibrates a full general-equilibrium overlapping-generations model (calibration period 1960-2015, in 5-year intervals) with endogenous survival probabilities following Hall and Jones (2007), monopolistic competition, and a PAYG social security system. The model is calibrated to match five time series: relative health price, life expectancy, health expenditure share, capital share in health production, and labor share in health production. The baseline GE model additionally fits the non-targeted decline in average GDP growth rates well. In the baseline calibration, health-sector TFP is estimated to have grown at 0.3% annually from 1950-1970, accelerating to 0.8% (1975-1980), 1.3% (1985-1995), and 1.5% thereafter — faster than the non-health sector’s 0.7% after the mid-1970s. These GE-corrected estimates exceed those from partial-equilibrium exercises because the growth-accounting approach fails to account for factor-input endogeneity; the true GE path requires health-sector TFP to outpace non-health TFP to reconcile observed relative price growth with the magnitude of markup increases.&lt;/p&gt;
&lt;p&gt;Counterfactual simulations isolate each channel. When only demand effects operate (population growth and health-investment efficiency improvements), relative prices rise by only 6.4% compared to 131% in the predicted baseline, and the health share of expenditure rises by 0.004 percentage points versus 0.171 in the baseline — confirming the minor role of aging and demand-composition change. Rising markups alone reproduce nearly all relative price growth but drive expenditure shares up via price rather than quantity increases. Unbalanced TFP growth (with health-sector TFP growing faster post-1975) contributes to real output expansion in the health sector, partially drives up the expenditure share through quantities, supports GDP growth, and — by raising the real value of health services — sustains life-expectancy gains. By 2050, the baseline calibration projects health-sector markups to be approximately 6 times non-health-sector markups if the estimated 1.7% average annual markup growth continues.&lt;/p&gt;
&lt;p&gt;The policy implication is direct: market concentration — documented by HHI levels exceeding 2,500 in the majority of U.S. metropolitan areas, with 19% of MSAs having a single monopolistic hospital provider in 2017 — is the primary driver of rising relative health prices. Antitrust enforcement and policies encouraging technology adoption would together address price growth without sacrificing the real productivity gains that have driven longevity improvements. However, welfare analysis of such policies requires distinguishing between curbing care-provider market power versus pharmaceutical/equipment-manufacturer market power, the latter involving R&amp;amp;D investment incentives that the current aggregate model cannot disentangle.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-relative-markup-growth-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy for relative markup growth, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Relative markup growth is identified from the growth-accounting expression derived from a two-sector Dixit-Stiglitz model: the growth rate of the health-services share of aggregate consumption can be decomposed into relative markup growth, non-health TFP growth (from Penn World Tables), growth in sectoral capital and labor inputs (from BEA and NIPA), and aggregate consumption growth. Taking the capital intensity of the non-health sector as given (alpha_c = 0.40 from Horenstein and Santos 2019) and the data series as known, relative markup growth is backed out residually without requiring knowledge of health-sector TFP or alpha_h. Key threats: (1) the assumption that wages are equalized across sectors (the paper documents supporting evidence in Supplemental Appendix B.6); (2) the constancy of alpha_c, though a time-varying labor-share extension relaxes this; (3) the Dixit-Stiglitz framework abstracts from market selection and endogenous concentration, so markups are characterized as symmetric representative-firm markups rather than firm-distribution markups; (4) the Horenstein and Santos (2019) alternative markups from Compustat cover only publicly traded firms and may understate aggregate markup growth before the 1980s corporatization wave, biasing downward their markup-growth estimates and biasing upward implied TFP-growth estimates for that period.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The three mechanisms are: (1) rising relative markups (supply-side pricing power), (2) unbalanced TFP growth (sector-differential productivity), and (3) changing demand composition (aging and health-investment efficiency). In the partial-equilibrium growth-accounting stage, the three are separated algebraically in equation (8): relative price growth equals relative markup growth plus a GE-effect term (which captures input-price ratio variation and thus embeds demand composition effects) plus relative TFP variation. In the full GE counterfactual stage, channels are separated by switching them off one at a time (fixing gN = gz = gζj = 0 for demand; fixing µt at its 1955 level for markups; fixing gAc = gAh = 0 for TFP), and by activating only one channel at a time. Table 3 presents the counterfactual outcomes for five targeted moments (relative price growth, life-expectancy change, health expenditure share change, capital and labor input shares) under each scenario.&lt;/p&gt;
&lt;h3 id="q3-what-does-the-paper-find-about-health-sector-tfp-growth-and-how-does-this-revise-the-literature"&gt;Q3. What does the paper find about health-sector TFP growth, and how does this revise the literature?&lt;/h3&gt;
&lt;p&gt;The standard view (Triplett and Bosworth 2004; Bates and Santerre 2013) treats health as a ‘cost-disease’ sector with near-zero or negative TFP growth. The paper challenges this: in the baseline partial-equilibrium decomposition, health-sector TFP grows at 0.3% per year on average (1954-2019), compared to -1.3% per year in a model that ignores markup growth entirely. In the full GE model, health-sector TFP growth is higher still — 0.3% (1950-1970), 0.8% (1975-1980), 1.3% (1985-1995), and 1.5% thereafter — eventually exceeding the non-health sector’s 0.7% annual rate. The authors argue this upward revision is correct: partial-equilibrium exercises omit GE feedback effects through factor-input reallocation, and prior studies that did not account for rising markups mechanically attributed all relative price growth to slow TFP growth, biasing health-sector TFP estimates downward.&lt;/p&gt;
&lt;h3 id="q4-what-role-does-population-aging-and-demand-composition-change-play-and-what-is-the-channel"&gt;Q4. What role does population aging and demand-composition change play, and what is the channel?&lt;/h3&gt;
&lt;p&gt;Demand composition changes (population aging and improvements in health-investment efficiency ztζjt) have only a minor role. In GE, such changes can affect input prices (r/w) and thereby health prices only if the health sector uses a different capital intensity than the non-health sector (alpha_h ≠ alpha_c); the elasticity of relative price with respect to the input-price ratio is (alpha_h - alpha_c), which is negative since health is more labor-intensive. This means demand effects actually dampen rather than amplify relative price growth. In the counterfactual where only demand effects operate, relative prices rise by only 6.4% (versus 131% in the predicted baseline from 1960-2015), and the health expenditure share increases by only 0.004 percentage points (versus 0.171 in the predicted baseline). Demand effects do, however, significantly affect life expectancy: shutting them off while allowing only markups produces declining life expectancy, illustrating that income growth and health-investment efficiency improvements are central to longevity gains.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented"&gt;Q5. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;The model features age heterogeneity in three dimensions: (1) age-specific health elasticity θj (how much health expenditure converts to health status); (2) age-specific health output intensity φj; (3) age-specific health-investment productivity ζjt, borrowed from Hall and Jones (2007). These parameters allow older individuals to have lower elasticities of health status with respect to health expenditure, matching the empirical regularity that older patients benefit less per dollar spent on health care. The paper also documents heterogeneity in the sub-components of the health PCE aggregate: over time, prescription drugs and medical appliances have declined in their relative contribution to aggregate health price increases, while hospital services have increased in their relative contribution, consistent with Cooper et al. (2019) on hospital pricing power. Across calibrations, health-sector TFP growth rates vary across four eras, reflecting the different pace of productivity improvements over time.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The authors conduct four different decomposition exercises in the partial-equilibrium stage: (a) baseline with constant non-health capital intensity; (b) time-varying non-health labor share; (c) using Horenstein and Santos (2019) markups from Compustat for publicly traded firms; (d) zero relative markup growth as an extreme baseline. In Supplemental Appendix C.2 they also invert the identification: they set health-sector TFP growth to values from the literature (-0.6% to 0.4% per year) and back out alpha_h, obtaining values between 0.25 and 0.38, consistent with the externally calibrated 0.26. Five full GE calibrations correspond to the five decomposition assumptions. Model fitness is assessed via RMSE across the five targeted moments; the baseline calibration fits best. An untargeted validity check against observed average GDP growth rates over 5-year intervals further supports the baseline model. Results from alternative calibrations’ counterfactuals are presented in Supplemental Appendix D.6.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The closest antecedents are: (1) Horenstein and Santos (2019), who attribute rising U.S. relative health prices to markups and price wedges using Compustat data; the present paper both uses their results and critiques them for under-coverage of non-publicly-traded firms. (2) Hall and Jones (2007), who model health investment, endogenous survival, and the demand side; the present paper embeds their survival technology into a two-sector GE model and adds the supply-side markup and TFP structure. (3) Fonseca et al. (2021, 2023), who account for the rise in health expenditure and cross-country health price differences; the present paper complements them by jointly modeling prices and quantities in a structural change framework. (4) Zhao (2014), who asks why health expenditure shares have risen from a demand side; this paper explores the supply-side (markup and TFP) counterpart. (5) Cost-disease literature (Baumol 1967; Triplett and Bosworth 2004): the paper directly challenges the ‘cost disease’ narrative by showing health-sector TFP is positive and — once GE and markup effects are controlled for — possibly faster than the rest of the economy. Distinctive contributions include the joint treatment of relative prices and real output quantities in structural change, the full GE calibration with endogenous population aging, and the explicit separation of health-care-quantity TFP from health-investment efficiency (the ztζjt composite).&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is that antitrust enforcement targeting market concentration in health services is the most direct lever for reducing relative price growth, since markup growth is almost entirely responsible for rising relative prices. A secondary policy recommendation is to encourage technology adoption in the health sector to sustain the high TFP growth that has benefited consumers through output expansion and life-expectancy improvements. The authors caution, however, that the model uses a broad definition of the health-services sector (encompassing care providers, pharmaceutical companies, and equipment manufacturers), and welfare implications differ sharply depending on whether policies target care-provider pricing power versus pharmaceutical/equipment pricing power, the latter involving R&amp;amp;D investment incentives. The model cannot disaggregate the sources of health-sector productivity growth, so the precise antitrust strategy requires further research. Additionally, the paper abstracts from 2020 short-term fluctuations and focuses on long-run structural change, so findings are most relevant for secular policy rather than cyclical interventions.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-role-of-unbalanced-tfp-growth-for-gdp-and-life-expectancy"&gt;Q9. What is the role of unbalanced TFP growth for GDP and life expectancy?&lt;/h3&gt;
&lt;p&gt;Counterfactual simulations reveal that unbalanced TFP growth — which in the baseline calibration favors the health sector after the mid-1970s — supports aggregate GDP growth. In the counterfactual where TFP growth is turned off (both sectors), GDP grows more slowly because the main remaining driver of income growth is exogenous population growth. The panel (f) of Figure 7 shows that GDP growth is slower without unbalanced TFP variation. For life expectancy, the absence of TFP growth causes life expectancy to rise until the 1980s then stagnate (purple line, panel (b) of Figure 7), since rising income is needed to purchase longevity gains through health investment. The interaction between income growth from TFP and the endogenous demand for health investment is central: health services function as a luxury good in the model, so income growth drives up the quantity demanded and thus survival rates.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-paper-find-about-the-current-level-and-trajectory-of-relative-markups"&gt;Q10. What does the paper find about the current level and trajectory of relative markups?&lt;/h3&gt;
&lt;p&gt;In the baseline calibration, health-sector markups were approximately 1.18 times non-health-sector markups in 1955. By 2010, this ratio had risen to approximately 3.2. The time-varying labor-share model implies even faster growth, from 1.09 in 1955 to 3.9 by 2010. Horenstein and Santos (2019) markups (slowest) go from 1.10 in 1955 to 3.04 in 2010, still a 176% increase. Under the baseline calibration projecting continued markup growth at 1.7% annually, health-sector markups would reach approximately 6 times non-health-sector markups by 2050. These projections are corroborated by micro evidence: HHI for managed care exceeds 2,500 in all California counties (Tawil and DiGiorgio 2022); national MSA-level hospital-bed HHI rose from 5,426 in 2007 to 5,808 in 2017; and 19% of MSAs had a single monopolistic provider in 2017 (Johnson and Frakt 2020).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Relative markup&lt;/strong&gt;: The ratio of the health-sector markup (price over marginal cost in a Dixit-Stiglitz monopolistically competitive equilibrium) to the non-health-sector markup; variation in this ratio is identified from sectoral input/output data and is almost entirely responsible for rising relative health-services prices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Unbalanced technical change&lt;/strong&gt;: Differential rates of TFP growth across the health and non-health sectors; in models with homothetic preferences and identical factor intensities, relative prices move inversely with relative TFP, but in the paper’s GE setting with different capital intensities the relationship is modified by GE input-price effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Health-investment efficiency (ztζjt^θj)&lt;/strong&gt;: An age-specific and time-varying composite productivity term governing how effectively a dollar of health-services expenditure (hjt) converts into improved health status and survival probabilities; it captures environmental, behavioral, and knowledge-based factors orthogonal to health-sector TFP (Aht), and is borrowed from Hall and Jones (2007).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GE (general equilibrium) effect&lt;/strong&gt;: In the price-decomposition framework, the term (αh − αc)(gLh,t − gKh,t) capturing how changes in the economy-wide capital-labor ratio — driven by demographic change, markup growth, and TFP changes — feed back into relative sector input prices and thereby into relative health prices; because αh &amp;lt; αc, this effect consistently dampens relative health-price growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cost disease&lt;/strong&gt;: The Baumol (1967) hypothesis that labor-intensive sectors like health services experience slow TFP growth, causing their relative prices to rise as economy-wide wages grow; the paper challenges this characterization by showing health-sector TFP growth is positive and, once GE and markup effects are controlled for, exceeds that of the non-health sector after the mid-1970s.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Corporatization of health services&lt;/strong&gt;: The historical transition of health-services providers from not-for-profit and public-sector organizations to for-profit investor-owned corporations (including private equity-backed systems), which the paper argues has driven the increase in aggregate health-sector markups and whose timing explains why Compustat-based markup estimates from the 1970s understate long-run markup growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous population / survival rate&lt;/strong&gt;: In the model, survival probabilities are functions of individual health-services expenditure (hjt) and health-investment efficiency; this makes population aging partly endogenous to health-sector pricing and productivity, linking structural change in health to aggregate life-expectancy dynamics and GDP growth within a unified OLG framework.&lt;/p&gt;</description></item><item><title>How Costly Are Cartels?</title><link>https://macropaperwarehouse.com/papers/how-costly-are-cartels/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/how-costly-are-cartels/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Moreau and Panon ask how much cartels cost the aggregate economy — in terms of both total factor productivity and welfare — and find the losses are considerably larger than the received wisdom from Harberger (1954) would suggest. The paper&amp;rsquo;s motivation is the mounting evidence that markups are large and growing, combined with a near-total absence of macroeconomic quantification of collusion as one micro-origin of those markups.&lt;/p&gt;
&lt;p&gt;The empirical foundation is an original firm-level database for France covering the period 1994–2007, assembled by scraping all written decisions of the French Competition Authority (ADLC). The final dataset contains 174 cartels and more than 1,000 firms before matching. These cartel records are merged to administrative balance-sheet and income-statement data covering the universe of French firms (BRN and RSI regimes). Key facts documented: average cartel duration is 4.5 years (median 3 years); average cartel size is 6.3 members (median 4); cartels are prevalent across construction, manufacturing, wholesale, retail, and transportation. Crucially, cartel members are empirically shown to be dramatically larger than non-members even within narrowly defined 4-digit industries — roughly 1,900% more sales, a market share premium of 4 percentage points, 1,150% more employment, and 37% higher labor productivity. Firms within a cartel are also substantially more homogeneous in productivity than the overall within-industry distribution: the interquartile productivity ratio across cartel members is only 1.4-to-1, versus 2-to-1 across all non-cartel firms in the same industry.&lt;/p&gt;
&lt;p&gt;The theoretical framework extends the static heterogeneous-firm oligopoly model of Atkeson and Burstein (2008) by introducing collusion microfounded via the cross-ownership framework of O&amp;rsquo;Brien and Salop (1999). A single collusion-intensity parameter κ ∈ [0,1] governs how much each cartel member internalizes the profits of other members. When κ = 0 the model reduces to competitive Cournot oligopoly; when κ = 1 all cartel members jointly maximize profits. In equilibrium, markups rise with firm market share, generating endogenous markup dispersion. Adding collusion causes cartel members to face a lower effective demand elasticity — their own market share augmented by the weighted market shares of co-conspirators — and to charge supracompetitive markups (overcharges). Critically, the effect of cartels on aggregate productivity is theoretically ambiguous: the output contraction of colluding firms redirects demand toward non-colluding firms. If the cartel is composed of the largest (most productive) firms, demand shifts toward less productive non-members, reducing productivity. If the cartel is composed of the least efficient firms, demand shifts toward large non-members, potentially improving allocation.&lt;/p&gt;
&lt;p&gt;The model is calibrated to match six moments from French data in 2007 — aggregate markup, cartel overcharge, the slope of the inverse-markup-on-HHI regression, the median number of firms per sector, the median number of cartel members, and the distribution of relative sales. The key calibrated parameters are: within-sector elasticity of substitution ρ = 10.19; across-sector elasticity η = 1.86; collusion intensity κ = 0.79. The cartel overcharge target is set to 10%, consistent with the OECD benchmark used by antitrust authorities and with Laborde (2021).&lt;/p&gt;
&lt;p&gt;Main quantitative findings (baseline calibration, cartels composed of top producers):&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Eliminating all cartels raises aggregate TFP by 1.1%.&lt;/li&gt;
&lt;li&gt;The productivity cost of markups with respect to the efficient allocation is 70% higher in the model with collusion (3.67%) than in the calibrated competitive oligopoly (2.16%), because collusion generates additional markup dispersion on top of the dispersion inherent in firm heterogeneity.&lt;/li&gt;
&lt;li&gt;Eliminating cartels brings the economy 30% closer to the efficient allocation.&lt;/li&gt;
&lt;li&gt;The aggregate markup falls by approximately 1.5 percentage points when cartels are eliminated.&lt;/li&gt;
&lt;li&gt;Consumption-equivalent welfare gains from eliminating cartels equal 2%.&lt;/li&gt;
&lt;li&gt;Larger cartels (market share above median) account for roughly 80% of the productivity gains; dismantling only large cartels yields a 0.88% TFP gain and 1.97% consumption-equivalent welfare gain; smaller cartels yield 0.23% TFP and 0.54% welfare.&lt;/li&gt;
&lt;li&gt;Umbrella pricing — non-cartel members raise their markups because the cartel&amp;rsquo;s higher prices provide cover — dampens aggregate gains quantitatively but only slightly: fixing non-members&amp;rsquo; markups yields 1.14% productivity gain versus 1.11% in the benchmark.&lt;/li&gt;
&lt;li&gt;Reducing collusion intensity from κ = 0.79 to κ ≈ 0.4 (roughly a 50% reduction) still generates TFP gains of 0.54% and welfare gains of 0.85%, demonstrating that tougher antitrust enforcement at the intensive margin (forcing cartels to soften, not dissolve) yields substantial gains.&lt;/li&gt;
&lt;li&gt;These estimates are one order of magnitude above Harberger&amp;rsquo;s (1954) 0.1% dead-weight loss estimate; the paper shows this discrepancy arises because Harberger uses sectoral data and near-unit demand elasticities, both of which suppress markup dispersion within sectors.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The paper&amp;rsquo;s scope conditions are explicit: results reflect the static cost of cartels; dynamic effects (entry deterrence, innovation incentives) are acknowledged but not quantified; only domestic, detected cartels are covered, so estimates likely understate the true cost; the channel through geographic markup dispersion is excluded.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-primary-identification-strategy-and-what-are-its-main-limitations"&gt;Q1. What is the paper&amp;rsquo;s primary identification strategy, and what are its main limitations?&lt;/h3&gt;
&lt;p&gt;The paper does not rely on a natural experiment or difference-in-differences design. Instead, it uses a structural calibration approach: a heterogeneous-firm oligopoly model with collusion is calibrated to match French data moments, and the cost of cartels is computed as the difference between the calibrated cartel equilibrium and a counterfactual competitive Nash-Cournot equilibrium. The main threats to this strategy are: (1) the sample of cartels consists only of detected cartels, which may not be representative of the latent population — discovered cartels could be either more or less severe than undiscovered ones; (2) no firm-level price data are available, so markups cannot be estimated directly; (3) the counterfactual is a calibrated competitive model rather than an empirically observed post-cartel state; (4) the model abstracts from entry and exit, which may dampen or amplify the true gains from cartel dissolution.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-through-which-cartels-affect-aggregate-productivity-and-how-are-they-distinguished"&gt;Q2. What are the main mechanisms through which cartels affect aggregate productivity, and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;Two channels operate simultaneously. First, the direct price effect: cartel members raise markups above the competitive level (overcharges), reducing their output. In the presence of markup dispersion, this disproportionately contracts output from high-markup (high-productivity) firms, increasing misallocation. Second, the demand reallocation effect: as cartel members contract output and raise prices, non-cartel members gain market share and increase their markups via the umbrella pricing mechanism. The net effect on productivity depends on which firms gain market share. When cartels consist of top producers, reallocation goes toward less productive non-members, reducing aggregate TFP. When cartels consist of the least efficient firms, reallocation goes toward larger non-members, potentially improving allocation. The two channels are not empirically separated in the data; rather, the model disentangles them analytically and then disciplines the net effect via calibration to observed cartel overcharges.&lt;/p&gt;
&lt;h3 id="q3-why-do-the-authors-assume-cartels-are-composed-of-the-most-productive-firms-and-what-is-the-evidence-for-this"&gt;Q3. Why do the authors assume cartels are composed of the most productive firms, and what is the evidence for this?&lt;/h3&gt;
&lt;p&gt;The assumption is motivated by three pieces of evidence. First, empirical regressions on the matched administrative data show that cartel members within their 4-digit industries have roughly 1,900% more sales, 1,150% more employment, and 37% higher labor productivity than non-members. Second, firms within a cartel are much more homogeneous than the overall within-industry distribution: the interquartile productivity ratio within a cartel is 1.4-to-1, versus approximately 2-to-1 for all non-cartel firms in the same industry, and the 90-10 ratio is 1.7-to-1 within a cartel versus over 4-to-1 across the industry. Third, only the top-producer composition assumption, combined with a collusion intensity κ = 0.79, can generate a cartel overcharge of 10% consistent with the calibration target. All other composition configurations (least efficient, all-inclusive, random top-10%) yield either implausibly small overcharges or implausibly large ones.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-umbrella-pricing-effect-and-how-large-is-it-quantitatively"&gt;Q4. What is the umbrella pricing effect and how large is it quantitatively?&lt;/h3&gt;
&lt;p&gt;Umbrella pricing refers to the mechanism by which cartel members&amp;rsquo; higher prices raise the sectoral price index, allowing non-cartel members to expand output and raise their own markups without reducing their market share. Proposition 1 of the model shows that collusion increases the markups of all firms — cartel and non-cartel — with non-cartel members experiencing markup increases that are larger for larger non-members. Quantitatively, when non-cartel members are held to fixed markups (so the umbrella effect is turned off), the aggregate TFP gain from eliminating cartels rises from 1.11% to 1.14% — a difference of 0.03 percentage points, or less than 3% of the total effect. The welfare effect is similarly small: 2.01% versus 2.00%. The umbrella pricing channel thus dampens aggregate gains but is quantitatively minor.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-in-cartel-effects-is-documented"&gt;Q5. What heterogeneity in cartel effects is documented?&lt;/h3&gt;
&lt;p&gt;Three dimensions of heterogeneity are explored. First, cartel size matters: large cartels (those with cumulated market share above the median) account for roughly 80% of the aggregate TFP gain from eliminating all cartels (0.88 percentage points out of 1.11%), while small cartels account for only 0.23 percentage points. Second, cartel composition is critical: top-producer cartels amplify misallocation, all-inclusive cartels generate very large overcharges and dramatically higher misallocation, least-efficient-firm cartels barely affect allocation, and random-top-10% cartels can slightly improve allocation. Third, collusion intensity matters monotonically: across the range κ = 0.1 to κ = 0.4, TFP gains from elimination fall from 0.99% to 0.54%, and welfare gains fall from 1.70% to 0.85%.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run-and-how-do-the-results-change"&gt;Q6. What robustness checks are run, and how do the results change?&lt;/h3&gt;
&lt;p&gt;The paper runs six main robustness experiments, all recalibrating the model: (1) Alternative overcharge target of 15% (versus 10% baseline): requires κ = 1.28, yields TFP gains of 1.63% and welfare gains of 2.77%. (2) Low aggregate markup target M = 1.1: TFP gain of 1.37%, welfare gain of 2.07%. (3) High aggregate markup target M = 1.3: TFP gain of 0.90%, welfare gain of 1.96%. (4) Bertrand rather than Cournot competition: TFP gain of 0.55%, welfare gain of 1.35% — smaller because Bertrand generates less markup dispersion, though the reduction in distance to the efficient allocation is larger (39%). (5) Heterogeneous κ across cartels drawn from a truncated normal with four variance levels: TFP gains range from 0.84% to 1.11% and welfare gains from 1.53% to 1.99%, close to the benchmark of 1.11% and 2.00%. (6) The cartel screen regression yields an estimated κ of 0.70 from data on colluding firms, close to the calibrated benchmark of 0.79.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-model-generate-a-cartel-detection-screen-and-what-does-it-find"&gt;Q7. How does the model generate a cartel detection screen, and what does it find?&lt;/h3&gt;
&lt;p&gt;The model&amp;rsquo;s equilibrium first-order conditions imply a regression of a cartel member&amp;rsquo;s labor share (a proxy for the inverse markup under log-linear production) on its own market share and the total cartel market share. The ratio of the estimated coefficient on cartel market share to the sum of both coefficients recovers the collusion intensity κ. Running this regression on the sample of detected cartel firms, the authors find a coefficient on own market share of -0.53 and an intercept of 0.70, both significant at 1%. Adding the cartel joint market share, its coefficient is negative and significant at 1%; the estimated κ from this specification is 0.70, close to the benchmark of 0.79. Results are qualitatively robust to including year fixed effects, though estimates become slightly noisier.&lt;/p&gt;
&lt;h3 id="q8-how-do-the-authors-explain-the-large-discrepancy-with-harberger-1954"&gt;Q8. How do the authors explain the large discrepancy with Harberger (1954)?&lt;/h3&gt;
&lt;p&gt;Harberger&amp;rsquo;s classic estimate of the deadweight loss from monopoly is approximately 0.1% of GDP. The authors show that their model can reproduce estimates close to this when (a) the model is aggregated to the sectoral level, eliminating within-sector markup dispersion — in that case, the TFP gain from eliminating cartels falls to 0.08%; or (b) demand elasticities are set close to unity as in Harberger&amp;rsquo;s sectoral data — the TFP gain falls to 0.24%. The key reason for the discrepancy is that Harberger&amp;rsquo;s framework suppresses both the within-sector dispersion of markups (which in the baseline model amplifies allocative losses) and the endogenous markup response to market share changes (which is large when ρ is substantially greater than 1). Using disaggregated firm-level data and calibrated high-within-sector elasticities restores the large estimated costs.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper implies that antitrust enforcement against horizontal price-fixing cartels can yield aggregate TFP gains of 1.1% and welfare gains of 2% in consumption-equivalent terms — figures the authors describe as conservative, because (i) the estimate is static (no dynamic gains from entry or innovation effects are included), (ii) only domestic detected cartels are captured and international cartels are excluded, (iii) geographic markup dispersion is abstracted from, and (iv) the calibration uses a conservative overcharge target of 10%. Importantly, the gains from targeting the intensive margin (forcing cartels to reduce overcharges rather than dissolving them entirely) are also substantial: a 50% reduction in κ still yields 0.54% TFP and 0.85% welfare gains. The results further imply that industrial policy and trade liberalization reforms that ignore competition enforcement may be partially undermined if new market power enables cartelization. The scope condition most critical to the quantitative magnitude is cartel composition: results depend on cartels being composed of top producers; the sign and magnitude of productivity effects can flip for alternative compositions. The authors also note that if cartels spur long-run innovation (through higher profits), their static welfare cost estimates would overstate the net social cost.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-differ-from-edmond-midrigan-and-xu-2022-and-baqaee-and-farhi-2020"&gt;Q10. How does this paper differ from Edmond, Midrigan, and Xu (2022) and Baqaee and Farhi (2020)?&lt;/h3&gt;
&lt;p&gt;Edmond et al. (2022) and Baqaee and Farhi (2020) quantify the total welfare and productivity cost of markups relative to the efficient allocation — the gap between the current economy (with all its markup dispersion from firm heterogeneity) and the first-best. Moreau and Panon instead isolate the cost of one specific, policy-relevant source of excess markup dispersion — collusion — by computing the gap between the cartel equilibrium and the competitive (but still imperfect) Nash-Cournot equilibrium. They also show that competitive oligopoly models of the Edmond et al. type understate the total misallocation cost of markups by approximately 70% when cartels are present and composed of top producers, because competitive models are calibrated to match the same aggregate markup data but attribute all markup dispersion to firm heterogeneity rather than to collusion. The papers are thus complementary: Edmond et al. bound the full cost of all markup distortions, while Moreau and Panon bound the portion attributable to cartels and amenable to competition enforcement.&lt;/p&gt;
&lt;h3 id="q11-what-caveats-and-limitations-do-the-authors-acknowledge"&gt;Q11. What caveats and limitations do the authors acknowledge?&lt;/h3&gt;
&lt;p&gt;The authors flag several important limitations. (1) The analysis is static: dynamic effects — including entry deterrence by cartels, barriers to exit for inefficient firms, and the innovation-competition relationship — are not modeled. The relationship between competition and innovation is hump-shaped (Aghion et al., 2005), so cartels could in principle spur or dampen innovation; the authors treat their estimates as an upper bound if cartels raise innovation. (2) Only detected French domestic cartels are in the sample; international cartels (investigated by the European Commission) and undetected cartels are excluded, likely causing understatement of total costs. (3) The selection of detected cartels is non-random: the direction of bias from using only discovered cartels is unclear — discovered cartels may be unusually large (biasing costs upward) or undiscovered large cartels may exist (biasing costs downward). (4) The model abstracts from geographic markup dispersion and from vertical arrangements across industries. (5) The model has no entry or exit of firms, which could amplify or dampen transition dynamics. (6) Firm-level prices are unavailable, so markups cannot be directly measured and must be inferred from the model or from labor shares.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Collusion intensity parameter (κ)&lt;/strong&gt;: A scalar in [0,1] that governs the weight each cartel member assigns to co-conspirators&amp;rsquo; profits when choosing output. When κ = 0, behavior is competitive Cournot; when κ = 1, members jointly maximize aggregate cartel profits. In the baseline calibration κ = 0.79, chosen to match a 10% median cartel overcharge in French data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cartel overcharge&lt;/strong&gt;: The percentage difference in cartel members&amp;rsquo; average markups between the cartel equilibrium and the competitive Nash-Cournot equilibrium. Computed as the median overcharge across cartels in the model. In the baseline calibration it is 10%, consistent with the OECD benchmark and Laborde (2021). The overcharge increases with both collusion intensity (κ) and the cartel&amp;rsquo;s total market share.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Umbrella pricing&lt;/strong&gt;: The mechanism by which a cartel&amp;rsquo;s higher prices raise the sectoral price index, enabling non-cartel members to expand demand, gain market share, and charge higher markups than they would in the absence of the cartel. In the model, umbrella pricing implies that the introduction of collusion increases the markups of all firms in cartelized sectors, not just cartel members; quantitatively, the effect dampens but does not reverse the aggregate productivity gains from cartel dissolution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distance to efficient allocation&lt;/strong&gt;: The ratio of the productivity gain from eliminating cartels (Acartel → Acomp) to the total productivity gain from eliminating all markup dispersion (Acomp → Aeff or equivalently from Acartel → Aeff). In the baseline, eliminating cartels reduces this distance by 30%, meaning cartels are responsible for roughly 30% of the gap between the actual economy and the first-best efficient allocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous markups (size-related)&lt;/strong&gt;: In the Atkeson-Burstein framework embedded in this model, a firm&amp;rsquo;s equilibrium markup is a harmonic average of within- and between-sector demand elasticities weighted by the firm&amp;rsquo;s own market share. More productive firms endogenously hold larger market shares and thus face lower demand elasticities, charging higher markups. Collusion further distorts this by augmenting the effective market share with co-members&amp;rsquo; shares, yielding supracompetitive overcharges.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cartel composition&lt;/strong&gt;: The identity of firms within a cartel — specifically, where they sit in the within-industry productivity distribution. The paper shows this is the single most important determinant of whether cartels amplify or dampen aggregate misallocation. Empirically, discovered French cartels are composed of the largest, most productive firms (nearly 1,900% more sales than non-members), and this is the only composition configuration that can match observed 10% overcharges in the calibrated model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive versus extensive margin of cartel policy&lt;/strong&gt;: The extensive margin refers to whether a cartel exists (zero versus positive κ); the intensive margin refers to the degree of collusion among existing cartel members (high versus low κ). The paper shows both margins are quantitatively important: breaking down all cartels (extensive margin) yields 1.11% TFP gain, while halving κ without dissolution (intensive margin) yields 0.54% TFP gain and 0.85% welfare gain.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cartel screen&lt;/strong&gt;: A regression of cartel members&amp;rsquo; labor shares on their own market share and the joint cartel market share, derived directly from the model&amp;rsquo;s equilibrium first-order conditions. The collusion intensity κ can be recovered as the ratio of the joint market share coefficient to the sum of both market share coefficients. Applied to French data on detected cartel firms, this screen yields κ̂ = 0.70, close to the calibrated value of 0.79.&lt;/p&gt;</description></item><item><title>Illuminating the Global South</title><link>https://macropaperwarehouse.com/papers/illuminating-the-global-south/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/illuminating-the-global-south/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Satellite nighttime lights (luminosity) are the dominant remote-sensing proxy for local economic conditions in low-income countries, yet their accuracy at fine spatial scales and over time has remained contested. This paper by Chiovelli, Michalopoulos, Papaioannou, and Regan makes two linked contributions. First, it constructs a standardized, annual, global panel of nighttime lights from 1992 to 2023, integrating the legacy DMSP-OLS satellite series (1992–2013) with the higher-quality VIIRS series (2013–onward) after applying three adjustments to the noisier DMSP data: cross-sensor inter-calibration (following Li et al. 2020), top-coding correction (following Bluhm and Krause 2022, using a truncated Pareto distribution to replace pixels with Digital Number ≥ 55), and blooming correction (following Cao et al. 2019, modeling light spillover as spatial decay and subtracting predicted pseudo-light). VIIRS is then downgraded to DMSP-comparable units using an ensemble machine-learning method — extremely randomized trees trained on the single year of full overlap (2013) — yielding an out-of-sample RMSE of 1.50 versus 3.27 for the Li et al. sigmoid approach and 1.57 for the Nechaev et al. convolutional neural network; the F1 score for the binary lit/unlit classification is 0.72 versus 0.51 and 0.71 for those alternatives, with recall = 0.95 and precision = 0.58 against an actual lit-pixel share of only 8.6 percent globally. At the cross-country level — a sample of 173 countries — the adjusted series retains an elasticity of luminosity to GDP of approximately 0.85 and an R² around 0.9 in cross-section; for Africa specifically the elasticity is 0.7 and R² remains around 0.9. In long-difference panel regressions over 1992–2019, the luminosity-GDP elasticity is approximately 0.25–0.24, broadly consistent with Henderson et al. (2012)&amp;rsquo;s estimate of 0.30–0.33, while at the five-year panel frequency the elasticity is around 0.15–0.17. The second contribution is a systematic validation of the new series against multiple local development proxies across four low-income settings. Using 139 georeferenced DHS surveys from 34 African countries (gridcells of ~28km × 28km), the adjusted series yields cross-sectional coefficients of approximately 0.6 standard deviations for schooling, electricity access, and improved sanitation, and approximately 1 standard deviation for the composite wealth index, between lit and unlit gridcells; in within-gridcell panel regressions, the adjusted log-lights coefficient on schooling is approximately double that of the unadjusted series (~0.02 versus ~0.01), and lit/unlit panel coefficients are statistically significant only with the adjusted series — gridcells turning lit see schooling rise by ~0.05 standard deviations (~0.125 schooling years), wealth index rise by ~0.05 SD, and electricity access rise by ~0.05 SD. In Mozambique, using all post-civil-war censuses (1997, 2007, 2017) across 1,126 admin-4 localities, schooling and non-agricultural employment are at least 0.5 standard deviations higher in lit than unlit localities, equivalent to approximately 0.5 years of schooling and 10 percentage points of non-agricultural employment; within-locality changes in lights co-move significantly with schooling changes, with the difference in schooling gain between localities that turn lit versus stay unlit being about half a year even controlling for admin-3 fixed effects. In Indonesia, panel estimates for public goods across more than 60,000 PODES villages show the adjusted series yields a positive and significant coefficient on the composite wealth index while the unadjusted series yields a counterintuitively negative coefficient. In India, across more than 550,000 SHRUG villages and towns, the adjusted series consistently produces stronger cross-sectional and panel associations with non-farm, manufacturing, and services employment. A key empirical regularity across all settings is that the adjusted series outperforms the unadjusted one most sharply at finer spatial resolutions and in over-time (panel) comparisons, while at coarse aggregation levels (large administrative units or large grid squares) differences between the two series are minor, as spatial averaging attenuates measurement error in the unadjusted data too. Blooming correction delivers most of the improvement in the African context, where top-coding is rare (fewer than 2% of lit DMSP pixels in Africa approach the 63 DN ceiling). The paper also replicates three canonical studies — Michalopoulos and Papaioannou (2013) on precolonial ethnic institutions, Michalopoulos and Papaioannou (2014) on national institutions and split ethnic homelands, and Hodler and Raschky (2014) on regional favoritism — confirming that qualitative conclusions are robust to the data revision while documenting that the adjusted series sharpens several estimates, particularly those exploiting within-region over-time variation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper is a measurement and validation study rather than a causal identification exercise. Its core design is correlational: it regresses local development proxies on nighttime luminosity across gridcells and administrative units, conditioning on country-year fixed effects in cross-section and on unit fixed effects in panel regressions. The main threats are (a) reverse causation (luminosity and development are jointly determined), which the authors acknowledge but do not attempt to address — they are explicit that the goal is proxy validation, not causal estimation; (b) measurement error in both the luminosity variable and the development outcomes (DHS wealth index, census schooling, PODES public goods), which the paper addresses by comparing adjusted versus unadjusted luminosity series and interpreting attenuation bias reduction as evidence of improved measurement; (c) the binary transformation of luminosity (lit/unlit) produces non-classical measurement error — an explicit point drawn from econometric theory (Aigner 1973; Meyer and Mittag 2017) — which partly motivates the adjusted continuous series; and (d) spatial autocorrelation and systematic geographic patterns in prediction error, which the authors check by regressing prediction errors on latitude and longitude and find that the ERT-downgraded series reduces the latitude coefficient to 10% of its magnitude in the unadjusted VIIRS specification for log lights and to 35% for the lit indicator.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-dmsp-deficiencies-corrected-and-what-are-the-specific-methods-used"&gt;Q2. What are the three DMSP deficiencies corrected and what are the specific methods used?&lt;/h3&gt;
&lt;p&gt;Cross-sensor inter-calibration: DMSP data come from six satellites; Li et al. (2020) supply a cross-calibrated series using a second-order polynomial fitted on overlapping satellite years, which the paper adopts as its &amp;lsquo;unadjusted&amp;rsquo; baseline. Top-coding: DMSP records 8-bit Digital Numbers (DN) 0–63, so radiance above a ceiling is truncated. Pixels with DN ≥ 55 are subject to &amp;lsquo;implicit&amp;rsquo; top-coding (averages of potentially top-coded sub-readings). The correction uses the radiance-calibrated (RC) vintage available for seven years, ranks the top-coded pixels by the RC series from the nearest year, then replaces them with &amp;lsquo;structural values&amp;rsquo; drawn from a truncated Pareto distribution with parameters α = 1.5, L = 55, H = 2000. Blooming: the DMSP sensor stretches edge pixels and can be spatially displaced up to 3 km, causing light spillover. Following Cao et al. (2019), pseudo-light pixels (PLPs) — lit pixels neighboring at least one dark pixel — are identified. An OLS regression of PLP light on the inverse-squared-distance weighted sum of neighbors&amp;rsquo; light within a 7 × 7 window is estimated separately for broad global regions. The predicted blooming contribution is subtracted from each lit pixel, negative residuals are set to zero, and a local 3 × 3 mean smoothing is applied. Globally, the blooming correction raises the share of unlit pixels from 92% to 95% in 1992 and from 88% to 91% in 2012.&lt;/p&gt;
&lt;h3 id="q3-how-is-viirs-downgraded-and-harmonized-with-dmsp-and-what-does-extremely-randomized-trees-mean"&gt;Q3. How is VIIRS downgraded and harmonized with DMSP, and what does &amp;rsquo;extremely randomized trees&amp;rsquo; mean?&lt;/h3&gt;
&lt;p&gt;Because VIIRS records 14-bit DN at 15-arc-second resolution with far superior sensor quality, it is not directly comparable to the 8-bit, 30-arc-second DMSP. The authors&amp;rsquo; preferred approach downgrades VIIRS to match the DMSP scale. They use an ensemble machine-learning method called &amp;rsquo;extremely randomized trees&amp;rsquo; (Geurts et al. 2006), a variant of random forests that, instead of choosing the best splits from the training sample, picks split thresholds randomly, which further reduces variance and improves computational efficiency. Features used to predict DMSP-like values from VIIRS include: pixel statistics (mean, median, min, max of the four VIIRS sub-pixels within each DMSP 30-arc-second cell), statistics of neighboring pixels within windows of 3, 4, 7, 9, 11, 13, 17, and 21 pixel widths, and regional dummies for broad world regions. The model is trained on 2013 (the one full year of DMSP-VIIRS overlap) and its out-of-sample performance is assessed by retraining on 2012 and predicting 2013. Four merged series are produced corresponding to the four versions of DMSP (unadjusted; blooming only; top-coding only; both). The authors&amp;rsquo; approach outperforms both the Li et al. (2020) sigmoid-function method (RMSE 3.27 globally vs. 1.50) and the Nechaev et al. (2021) CNN approach (RMSE 1.57), especially in the low-to-middle luminosity range most relevant for low-income countries.&lt;/p&gt;
&lt;h3 id="q4-what-development-proxies-are-used-in-validation-and-across-what-samples"&gt;Q4. What development proxies are used in validation and across what samples?&lt;/h3&gt;
&lt;p&gt;Africa (DHS, 34 countries, 139 surveys, ~28km × 28km gridcells): mean years of schooling (respondents aged 15–39), DHS composite household wealth index, share of households with improved sanitation, share with electricity connection. All outcomes are standardized to mean zero, SD one. Mozambique (Census 1997, 2007, 2017, 1,126 admin-4 localities): mean years of schooling (aged 15–39) and non-agricultural employment (aged 15–24 or 19–24). Indonesia (PODES village census waves 1996–2018, 60,000+ villages): binary measures for garbage disposal, toilet use, drinking water access, gas/electricity for cooking, paved roads, and counts of kindergartens, primary, middle, and secondary schools — aggregated into a first principal component (eigenvalue ~3.5, capturing ~1/3 of variance). India (SHRUG dataset, 550,000+ towns and villages, Population Censuses 1991/2001/2011, Economic Censuses 1990/1998/2005/2013): population count, total non-farm employment, manufacturing employment, services employment.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented"&gt;Q5. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Spatial resolution: adjusted series outperforms unadjusted most at fine resolutions (2×2 gridcell blocks, ~56km × 56km at the equator); at coarse levels (12×12 blocks, ~336km × 336km), both series yield similar coefficients, as spatial aggregation attenuates noise in the unadjusted series. Urban vs. rural: cross-sectional estimates are similarly significant in urban and rural DHS samples. Panel estimates are statistically significant only with the adjusted series; urban panel coefficients are consistently larger than rural ones, echoing Asher et al. (2021)&amp;rsquo;s India finding. The adjustment matters more in rural areas than in urban areas in cross-section. Local variation (spatial RDD / fine fixed effects): with unadjusted series, panel wealth-index coefficients are statistically indistinguishable from zero until spatial fixed effects cover areas at least 7×7 gridcells (~200km × 200km at equator); with the adjusted series, coefficients remain significantly positive at all fixed-effect sizes including the finest 2×2 blocks. Top-coding vs. blooming: most of the improvement in Africa derives from blooming correction; top-coding correction has minor impact because fewer than 2% of lit African DMSP pixels approach the DN ceiling. Country-ethnic homelands (large areas, avg. 25,547 km²): adjustments matter little because spatial averaging already reduces noise. Applications replication: the precolonial institutions result (Michalopoulos and Papaioannou 2013) is robust and essentially unchanged because the units are very large. The national-institutions-at-border result (Michalopoulos and Papaioannou 2014) is strengthened in within-ethnicity specifications (coefficient marginally significant at 90% with adjusted series vs. p ≈ 0.15 with unadjusted); capital-proximity heterogeneity is sharpened. The regional-favoritism result (Hodler and Raschky 2014) strengthens: the log-lights lagged-leader coefficient rises from 0.038 to 0.058, and the lit-probability coefficient rises from ~3 to ~7 percentage points.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-and-specification-variations-are-run"&gt;Q6. What robustness checks and specification variations are run?&lt;/h3&gt;
&lt;p&gt;The paper compares four luminosity series (unadjusted Li et al.; blooming only; top-coding only; both combined + VIIRS fusion) to isolate each correction&amp;rsquo;s contribution. It checks the luminosity-GDP nexus at annual, five-year, and long-difference frequencies. It examines seven African countries&amp;rsquo; co-evolution of the harmonized series with electrification share (Kenya, DRC, Ghana, Tanzania, Nigeria, Mozambique, and one other) and finds no discontinuity at the 2012/2013 DMSP-VIIRS transition year. Spatial aggregation robustness: coefficients are computed across aggregation blocks ranging from 2×2 to 12×12 gridcells, showing stability in cross-section (~0.18) and mild size dependence in panel (~0.075, slightly rising with coarser units). Local variation robustness: fixed effects of increasing spatial coverage (2×2 to 12×12 cells) are added while the outcome remains at the gridcell level. Results replicated for schooling and electricity access (Appendix Section B.2) beyond the primary wealth-index outcome. Confounding by latitude in the ML model is assessed via regressions of prediction errors on latitude and longitude with and without country fixed effects. Median regressions confirm the OLS elasticity estimates at the cross-country level. The India analysis is replicated for both towns (urban) and villages (rural) separately.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Henderson et al. (2012): pioneer the use of luminosity as a cross-country GDP proxy and estimate a long-difference elasticity of 0.30–0.33 across 188 countries; this paper estimates 0.25–0.24 over a comparable specification, consistent but slightly lower. Gibson et al. (2021): show that VIIRS is superior to DMSP but find weak GDP-lights correlations outside cities for the early DMSP period in China, Indonesia, and South Africa; this paper addresses the concern by adjusting DMSP and merging it with VIIRS. Asher et al. (2021): validate luminosity as a strong proxy in India and find stronger urban-luminosity links; this paper replicates and extends those findings to Africa, Mozambique, and Indonesia and shows the adjusted series strengthens the Asher et al. patterns. Chen et al. (2024): find strong cross-sectional but weak panel associations; this paper&amp;rsquo;s adjusted series substantially strengthens panel associations. Bluhm and Krause (2022): provide the top-coding correction method adopted here. Cao et al. (2019): provide the blooming correction method. Nechaev et al. (2021): propose a CNN-based DMSP-VIIRS fusion but apply it to the unadjusted DMSP; this paper outperforms their RMSE slightly (1.50 vs. 1.57) and improves on their F1 score (0.72 vs. 0.71), with greater advantage in low-light regions. Li et al. (2020): propose a sigmoid-based fusion calibrated for high-light pixels; this paper substantially outperforms it (RMSE 1.50 vs. 3.27) particularly in low-luminosity areas. The paper thus synthesizes and extends multiple strands: it unifies the corrections of Bluhm-Krause and Cao et al., pairs them with state-of-the-art ensemble ML fusion, and provides by far the most comprehensive multi-country, multi-context validation of the resulting series.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is methodological: researchers studying development in low-income countries should use the adjusted and harmonized nighttime lights series rather than raw DMSP data, and should be especially careful at fine spatial scales (e.g., spatial regression discontinuity designs, granular village-level analyses) and in panel specifications. The gains from adjustment are largest precisely where applied development research is moving — toward local identification strategies and over-time variation. For practitioners and statistical agencies, the series provides a low-cost annual proxy for local economic conditions in environments with weak administrative data, particularly across sub-Saharan Africa, South Asia, and Southeast Asia. Scope conditions: (a) Correlations are far from perfect — binary lit/unlit classification misses much variation in the many-zeros low-income context. (b) At large aggregate units (admin-1, country-ethnic homelands), the adjustments yield minimal additional improvement since noise averages out. (c) The series does not resolve the fundamental limitation that most of sub-Saharan Africa remains unlit (98.4% of DMSP pixels in Africa in 1992), so it captures variation among already-lit areas better than the development gradient at the zero-light frontier. (d) Future research blending nighttime lights with daytime imagery (traffic, built structures) is flagged as a promising extension, though daytime data are often proprietary.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-main-findings-from-the-three-replication-exercises"&gt;Q9. What are the main findings from the three replication exercises?&lt;/h3&gt;
&lt;p&gt;Michalopoulos and Papaioannou (2013) — precolonial ethnic institutions and contemporary development: Replication across 682 country-ethnic homelands confirms that areas with higher precolonial political centralization (as measured by a 0–4 jurisdictional hierarchy index) have significantly higher contemporary luminosity, conditional on country constants and geographic controls. With the adjusted series, the unlit share among homelands rises from 24% to 29% (because blooming correction removes spurious light), but the coefficients on political centralization are still highly significant, somewhat smaller in magnitude, and similar qualitatively. The main conclusion is robust because the units are large and spatial averaging already reduces noise in the raw series. Michalopoulos and Papaioannou (2014) — national institutions and split-border ethnic development: Replication across 38,427 gridcells of 220 systematically partitioned ethnic homelands. Cross-sectional results show a one-point increase in the rule-of-law index (range −2.5 to 2.5) is associated with a ~10 pp higher probability of a gridcell being lit. The within-ethnicity coefficient drops by more than half (~0.025). With the adjusted series, this within-ethnicity coefficient is marginally significant at 90% versus a p-value of ~0.15 with unadjusted. Spatial RDD coefficients remain small and insignificant regardless of adjustment. Capital-proximity heterogeneity: the positive association between rule of law and luminosity is significant only for ethnically split groups where both portions are close to their respective capitals, and this finding is more precisely estimated with the adjusted series; the effect is nil far from capitals in both series. Hodler and Raschky (2014) — regional favoritism: Panel replication across 38,427 subnational regions in 126 countries, 1992–2009. The lagged-leader dummy coefficient (log lights specification) rises from 0.038 to 0.058 with the adjusted series. The linear-probability-model lit indicator rises from ~3 to ~7 percentage points. All specifications with the adjusted series are at least two standard errors above zero, matching or exceeding the precision of the original.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-limitations-and-caveats-acknowledged-by-the-authors"&gt;Q10. What are the limitations and caveats acknowledged by the authors?&lt;/h3&gt;
&lt;p&gt;First, the correlations between luminosity and development are &amp;lsquo;far from perfect&amp;rsquo; — the binary lit/unlit transformation in particular fails to capture the significant continuous variation in assets, education, and public goods across regions that are all formally &amp;rsquo;lit.&amp;rsquo; Second, bottom-coding (under-recording of low-light areas) is acknowledged but not corrected; no existing method addresses it, though the authors note that their corrections nonetheless improve elasticities even in rural African regions with very low light. Third, downgrading VIIRS to DMSP by construction sacrifices some of the VIIRS data quality; the long-difference VIIRS elasticity for Africa (0.4) shrinks to 0.35 in the downgraded series. Fourth, daytime satellite imagery and combinations with nighttime lights (Jean et al. 2016; Yeh et al. 2020; Rossi-Hansberg and Zhang 2025) can better capture local wealth but are often proprietary and not replicable in standard economic research. Fifth, the top-coding correction in Africa is minor because very few pixels approach the DN=63 ceiling (0.98–1.7% of lit pixels in 1992–2012), so the main African improvement comes from blooming; other regions with denser urban cores may benefit more from top-coding correction. Sixth, the cross-sensor inter-calibration step is taken &amp;lsquo;off-the-shelf&amp;rsquo; from Li et al. (2020) and further investigation of sensor calibration is left to future work.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Top coding (DMSP)&lt;/strong&gt;: The truncation of Digital Number values at the 8-bit ceiling of 63 in DMSP-OLS data, caused by sensor calibration for cloud detection. Pixels with DN ≥ 55 also suffer &amp;lsquo;implicit&amp;rsquo; top coding because they represent averages of multiple potentially top-coded sub-readings. The paper corrects this by replacing top-coded pixels with structural values drawn from a truncated Pareto distribution, using the radiance-calibrated DMSP vintage to rank pixels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Blooming (spatial spillover of light)&lt;/strong&gt;: A measurement artifact in DMSP data whereby light from bright pixels spills into neighboring dark areas due to the sensor&amp;rsquo;s imprecise spatial accuracy and possible displacement of up to 3 km. The paper identifies pseudo-light pixels (lit pixels adjacent to at least one dark pixel), models the spillover as an inverse-squared-distance weighted function of neighboring lights, and subtracts the predicted blooming from each lit pixel. This correction raises the global unlit pixel share from 92% to 95% in 1992.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extremely randomized trees (ERT)&lt;/strong&gt;: An ensemble machine-learning method used to downgrade VIIRS luminosity data to the DMSP scale. Unlike standard random forests that find the best split thresholds within a random feature subset, ERT selects split thresholds randomly, reducing variance and improving computational efficiency. The authors train it on pixel statistics (mean, median, min, max) and neighborhood statistics within windows of varying sizes to predict DMSP-like values for 2014 onward from VIIRS readings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Harmonized (adjusted + fused) luminosity series&lt;/strong&gt;: The authors&amp;rsquo; main output: an annual global panel of nighttime lights from 1992 to 2023 that applies inter-sensor calibration, top-coding correction, and blooming correction to DMSP data (1992–2013), then uses the ERT ensemble model to convert post-2013 VIIRS data into DMSP-comparable units, yielding four variants (unadjusted, blooming only, top-coding only, both corrections) merged into a continuous time series at 30-arc-second (~1 km²) resolution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pseudo-light pixels (PLPs)&lt;/strong&gt;: In the blooming correction procedure, PLPs are defined as lit pixels (DN &amp;gt; 0) that have at least one dark neighbor (DN = 0). They are the pixels most likely to contain spurious light from neighboring bright areas. PLP light values are regressed on the inverse-squared-distance weighted sum of surrounding pixels to estimate the blooming decay function.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DHS composite wealth index&lt;/strong&gt;: Used in the validation analysis as a local development proxy: a principal-component aggregation of household characteristics including roof quality and ownership of consumer assets, constructed by the Demographic and Health Surveys program across African countries. The paper standardizes this and other outcomes to mean zero and standard deviation one for cross-outcome coefficient comparisons.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Spatial RDD (regression discontinuity design) using nighttime lights&lt;/strong&gt;: As applied in Michalopoulos and Papaioannou (2014) and referenced throughout, a design that restricts estimation to gridcells within a narrow band (e.g., 50 km) of a political or administrative border to compare otherwise similar areas on opposite sides, using luminosity as the outcome. The paper notes that such fine-resolution, localized comparisons are exactly the setting where measurement error in the unadjusted DMSP series is most consequential and where the adjusted series yields the largest improvement.&lt;/p&gt;</description></item><item><title>Labor Market Discrimination and the Racial Unemployment Gap: Can Monetary Policy Make a Difference?</title><link>https://macropaperwarehouse.com/papers/labor-market-discrimination-and-the-racial-unemployment-gap-can-monetary-policy-make-a-difference/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/labor-market-discrimination-and-the-racial-unemployment-gap-can-monetary-policy-make-a-difference/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper addresses two connected questions: why do Black workers face persistently higher and more volatile unemployment than white workers, and can the Federal Reserve&amp;rsquo;s August 2020 shift from a symmetric &amp;ldquo;Deviations&amp;rdquo; rule to a &amp;ldquo;Shortfalls&amp;rdquo; rule narrow the resulting racial unemployment gap? The authors build a New Keynesian search and matching model with endogenous separations (Mortensen-Pissarides) and add employer taste-based discrimination, calibrated to U.S. Current Population Survey microdata from January 1976 to December 2019.&lt;/p&gt;
&lt;p&gt;The empirical motivation is stark. In CPS data, the Black unemployment rate averages 12.0 percent against 5.5 percent for whites — a gap of 6.5 percentage points that is largely unexplained by observable characteristics such as age, education, marital status, and state of residence (Cajner et al. 2017). The racial gap is also strongly countercyclical: its cyclical correlation with the aggregate unemployment rate is 0.77. A Shimer (2012)-style flow decomposition shows that the separation rate margin accounts for approximately two-thirds (67 percent) of the mean gap and 60 percent of its cyclical variance, with the job-finding rate contributing 20 percent of the mean and 27 percent of variance.&lt;/p&gt;
&lt;p&gt;The model features two types of representative households that differ only in a non-productive attribute (race). Firms incur a per-period perceived cost κ₁ of employing a type-1 (Black) worker, following Becker (1971). This cost is time-invariant and not directly affected by monetary policy. Search is random (firms cannot direct search by race, consistent with anti-discrimination law). The model also incorporates Calvo price rigidities and an effective lower bound (ELB) on the nominal interest rate, solved via Dynare&amp;rsquo;s extended path method. Two aggregate shocks drive dynamics: a risk-premium (demand) shock and a productivity (supply) shock. The discriminatory parameter is calibrated to κ₁ = 0.0292 — equivalent to 3.6 percent of the steady-state average wage — to match the 6.4 percentage-point mean racial unemployment gap.&lt;/p&gt;
&lt;p&gt;The baseline model (under the symmetric Deviations rule) generates four untargeted results that match the data: (1) higher mean separation rates and lower mean job-finding rates for Black workers, with the ratio of Black-to-white separation rates at 2.3 in the model (1.9 in data); (2) higher cyclical volatility of Black unemployment, driven by higher separation-rate volatility; (3) a strongly countercyclical racial gap (near-unit correlation with aggregate unemployment in the model); and (4) positively skewed unemployment distributions for both groups — skewness that arises endogenously from the ELB constraint, which is absent when the ELB is removed. The mechanism is geometric: because Black workers face a higher reservation productivity threshold (due to κ₁ &amp;gt; 0), more Black workers cluster near that threshold. A given aggregate shock therefore moves a larger mass of Black workers across the threshold, amplifying their unemployment response relative to whites.&lt;/p&gt;
&lt;p&gt;Novel model-based discrimination measures — workers not hired or fired solely due to being Black — average 5.86 percent of the Black labor force under the Deviations rule and are strongly countercyclical (correlation with aggregate unemployment = 0.99 in the model vs. 0.64 in EEOC race-charge data). The welfare gap between white and Black households averages 2.4 percent in consumption-equivalent terms.&lt;/p&gt;
&lt;p&gt;Shifting to the Shortfalls rule — which responds to unemployment shortfalls symmetrically but only tightens policy when unemployment is above its steady-state level — strengthens expansions by keeping interest rates lower. The aggregate unemployment rate falls by 0.7 percentage point, from 6.37 percent to 5.65 percent. Because Black workers are more cyclically sensitive, they benefit disproportionately: Black unemployment falls by 1.1 percentage points and white unemployment falls by 0.7 percentage points, narrowing the racial gap by 0.5 percentage point (from 6.50 to 6.03 percent). Model-based discrimination also declines (aggregate measure from 5.86 to 5.52 percent). The downside is a 0.5 percentage-point rise in average inflation, from 1.9 percent to 2.4 percent. The negative skewness in the racial unemployment rate gap is essentially eliminated under the Shortfalls rule, so the distribution shifts toward a lower mean with fewer episodes of extreme gaps.&lt;/p&gt;
&lt;p&gt;From a welfare perspective, however, the gains are quantitatively trivial. Both households experience slightly positive welfare gains under the Shortfalls rule — consumption rises by 0.62 percent for Black households and 0.64 percent for white households — but the differences are effectively indistinct from zero in consumption-equivalent terms. Crucially, the consumption-equivalent welfare wedge between the two groups actually widens slightly, because white wages rise more than Black wages under the Shortfalls rule (average productivity of Black employed workers falls more as the lower reservation threshold admits marginal workers). The authors note their welfare analysis is a lower bound, given within-group consumption insurance, the absence of liquidity constraints, and non-expiring unemployment benefits in the model.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses a structural calibration approach rather than quasi-experimental identification. The model is calibrated to match 10 aggregate moments (1976-2019 CPS data) with all parameters common across racial groups except κ₁. The racial unemployment gap in steady state is the sole targeted moment for racial differences; all other racial outcomes are untargeted predictions. Threats include: (1) the model attributes all cross-race labor market differences to discrimination, ruling out unobserved productivity heterogeneity; (2) the representative firm with taste-based discrimination abstracts from market-selection forces that, in Becker&amp;rsquo;s classic model, would erode discrimination in the long run (the authors cite Black 1995, Rosen 1997, Sasaki 1998 for equilibrium justifications); (3) the model is solved under perfect foresight (extended path), not fully stochastic, though Dynare&amp;rsquo;s method approximates stochastic dynamics; (4) the Shortfalls rule is a reduced-form approximation of the FOMC&amp;rsquo;s 2020 framework, not a structural representation.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-through-which-discrimination-generates-the-observed-racial-unemployment-patterns"&gt;Q2. What are the main mechanisms through which discrimination generates the observed racial unemployment patterns?&lt;/h3&gt;
&lt;p&gt;The core mechanism is that κ₁ &amp;gt; 0 raises the reservation productivity threshold for Black workers at both hiring (firms require higher expected productivity to justify the cost) and separations (existing matches must clear a higher bar to survive). Because idiosyncratic productivity is log-normally distributed, more Black workers cluster near their higher reservation threshold than white workers do near the lower white threshold. This concentration in the density means that any aggregate shock — moving both thresholds — shifts a proportionally larger mass of Black workers across the destruction margin, amplifying the volatility of Black unemployment and separations. The countercyclical racial gap arises because aggregate downturns raise both reservation thresholds, but since more Black workers are near their threshold, more are destroyed. The authors show that the separation-rate margin dominates: in the model it explains 92 percent of the mean gap and 81 percent of its cyclical variance, somewhat overstating the empirical 67 percent and 60 percent, because variation in the job-finding rate comes mostly from the common job-meeting probability.&lt;/p&gt;
&lt;h3 id="q3-how-do-the-two-types-of-discrimination-in-the-model--hiring-discrimination-and-separation-discrimination--work-quantitatively"&gt;Q3. How do the two types of discrimination in the model — hiring discrimination and separation discrimination — work quantitatively?&lt;/h3&gt;
&lt;p&gt;The hiring discrimination measure Df_t counts the fraction of Black job-seekers who are not hired because their idiosyncratic productivity draw falls above the white reservation threshold but below the (higher) Black threshold. The separation discrimination measure Dλ_t counts the fraction of employed Black workers who are endogenously separated for the same reason. Under the Deviations rule with ELB, the hiring margin averages 0.64 percent and the separation margin averages 5.22 percent of the Black labor force, for a total Dt of 5.86 percent. Both measures are strongly countercyclical (correlations with aggregate unemployment of 0.80 and 0.95 respectively). Under the Shortfalls rule, these fall to 0.56 and 4.95 percent (total 5.52 percent), and their skewness toward high discrimination levels is significantly reduced.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-aggregate-macroeconomic-effects-of-switching-from-the-deviations-rule-to-the-shortfalls-rule"&gt;Q4. What are the aggregate macroeconomic effects of switching from the Deviations rule to the Shortfalls rule?&lt;/h3&gt;
&lt;p&gt;The Shortfalls rule keeps nominal interest rates lower during periods of below-target unemployment (its asymmetry means it does not tighten in expansions unless inflation rises). This raises average output and consumption. The aggregate unemployment rate falls by 0.7 percentage point (from 6.37 to 5.65 percent), driven by both a lower average separation rate (3.36 to 3.10 percent) and a higher average job-finding rate (50.14 to 56.99 percent). Average inflation rises by 0.5 percentage point (from 1.88 to 2.40 percent annually). The Shortfalls rule increases the volatility of all labor market variables (it has lower stabilization properties) but essentially eliminates the positive skewness in the aggregate unemployment rate. The probability of a binding ELB falls from 10.6 percent to 8.5 percent under the Shortfalls rule. The correlation between inflation and unemployment strengthens from -0.32 to -0.51.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-shortfalls-rule-differentially-affect-black-and-white-workers"&gt;Q5. How does the Shortfalls rule differentially affect Black and white workers?&lt;/h3&gt;
&lt;p&gt;Black workers benefit disproportionately because their unemployment is more cyclically sensitive. The unemployment rate falls by 1.1 percentage points for Black workers (from 11.89 to 10.78 percent) versus 0.7 percentage points for white workers (from 5.39 to 4.74 percent). The racial gap narrows by 0.5 percentage point (from 6.50 to 6.03 percent). Separation rates fall more for Black workers (6.53 to 6.29 vs. 2.90 to 2.65 for whites). Average wages for Black workers increase by 0.43 percent and for white workers by 0.48 percent. The slight relative wage disadvantage under the Shortfalls rule arises because the lower reservation threshold for Black workers admits workers with lower average productivity, pulling down average Black wages relative to whites.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-welfare-implications-of-the-policy-change-and-why-are-they-small"&gt;Q6. What are the welfare implications of the policy change, and why are they small?&lt;/h3&gt;
&lt;p&gt;Both households gain welfare under the Shortfalls rule, but the gains are quantitatively very small in consumption-equivalent terms (effectively indistinct from zero). The aggregate benefit — lower average unemployment — is partially offset by the cost of higher average inflation (price dispersion loss in the Calvo framework). Consumption rises by about 0.62 percent for Black households and 0.64 percent for white households. The consumption-equivalent welfare wedge between Black and white households (2.4 percent under the Deviations rule) actually widens slightly under the Shortfalls rule, because white wages increase more than Black wages. The authors emphasize several reasons their welfare analysis understates true racial inequality: (1) within-group consumption insurance prevents individual unemployment spells from being welfare-costly; (2) no liquidity constraints; (3) unemployment benefits do not expire; (4) the model abstracts from labor force participation margins and involuntary part-time employment. These features, if relaxed, would likely reveal larger welfare differences between the two groups.&lt;/p&gt;
&lt;h3 id="q7-what-role-does-the-effective-lower-bound-elb-on-nominal-interest-rates-play"&gt;Q7. What role does the effective lower bound (ELB) on nominal interest rates play?&lt;/h3&gt;
&lt;p&gt;The ELB is essential to generating positively skewed unemployment distributions in the model. Without the ELB, the model produces essentially symmetric (near-zero skewness) distributions for both aggregate and racial unemployment outcomes. With the ELB, the baseline model matches the observed positive skewness of the unemployment rate (1.25 aggregate; 1.23 for Black workers, 1.26 for whites). The ELB also raises the mean unemployment rate by about 0.25 percentage point and slightly amplifies labor market volatilities. It introduces a deflationary bias (inflation averages 1.88 percent vs. the 2.0 percent steady-state target). Critically, the main results — the 0.5 pp narrowing of the racial gap and 0.7 pp fall in aggregate unemployment under the Shortfalls rule — are robust to removing the ELB constraint (Appendix B.2.2), confirming they are not artifacts of the nonlinearity introduced by the ELB.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-are-conducted"&gt;Q8. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;Key robustness exercises include: (1) removing the ELB constraint, which confirms the main results hold (aggregate unemployment falls 0.7 pp, racial gap narrows 0.5 pp, inflation rises 0.5 pp without the ELB; Table A.8-A.9); (2) extending the unemployment flow decomposition to a three-state system (employed, unemployed, out of labor force), which confirms that the employment-to-unemployment (EU) transition is the primary driver of the racial gap even accounting for labor force participation transitions (Appendix A.2); (3) verifying that employer-to-employer transition rates are similar across racial groups (2.20 percent for Blacks vs. 1.96 percent for whites, 2004-2019), supporting the assumption of equal exogenous separation rates; (4) confirming that inflation experiences are similar between Black and white households using the Chicago Fed IBEX data (2.80 percent for Blacks vs. 2.87 percent for whites, 1983-2013), supporting the equal-inflation assumption; (5) presenting impulse response functions under both a productivity shock and a demand shock, in models with and without monetary policy inertia.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q9. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The paper contributes to four literatures. First, versus Cajner et al. (2017) on empirical racial labor market gaps, it provides a structural explanation rather than documenting gaps. Second, versus search-and-matching discrimination models (Bartel 1995, Bowlus-Eckstein 2002, Rosen 2003, Flabbi 2010, Borowczyk-Martins et al. 2017), the key contributions are: (a) endogenous separations (prior models used exogenous exit), which the authors view as essential since separation rates dominate the gap&amp;rsquo;s dynamics; and (b) incorporating nominal rigidities and an ELB, enabling analysis of monetary policy. Third, versus Ravenna-Walsh (2012) and Bergman et al. (2022), who embed worker heterogeneity in New Keynesian search models, this paper differs by modelling heterogeneity as discrimination rather than productivity differences, and by studying the Deviations-to-Shortfalls rule change specifically. Fourth, versus Bundick-Petrosky-Nadeau (2021) who study the same Deviations/Shortfalls comparison for the aggregate economy, this paper adds the racial dimension. Versus Lee et al. (2022), Nakajima (2023), and Ait Lahcen et al. (2023) — all of which also study monetary policy and racial inequality — the contribution is generating racial disparities endogenously from discrimination rather than taking them as given, and including endogenous separations.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-paper-find-about-the-countercyclicality-of-racial-discrimination"&gt;Q10. What does the paper find about the countercyclicality of racial discrimination?&lt;/h3&gt;
&lt;p&gt;Both the model and the data exhibit strongly countercyclical discrimination. In the data, EEOC race-based discrimination charges (normalized per non-white labor force member) have a contemporaneous correlation of 0.65 with the cyclical component of the aggregate unemployment rate from 1997 to 2019. In the model, the aggregate discrimination measure Dt has a correlation of 0.99 with aggregate unemployment. The countercyclical pattern arises mechanically from the higher density of Black workers near the reservation productivity threshold: during recessions, both thresholds rise, destroying proportionally more Black matches and blocking more Black hires. The model-based discrimination measure also shows positive skewness (1.13 aggregate skewness under the Deviations rule with ELB), consistent with the asymmetric incidence of recessions.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-quantitative-scope-conditions-and-limitations-the-authors-themselves-identify"&gt;Q11. What are the quantitative scope conditions and limitations the authors themselves identify?&lt;/h3&gt;
&lt;p&gt;The authors identify several scope conditions and limitations: (1) the model abstracts from labor force participation, so it misses the racial gap in participation rates and involuntary part-time employment; (2) within-group consumption insurance and no liquidity constraints imply welfare estimates are a lower bound on true racial inequality — the consumption-equivalent wedge of 2.4 percent would be larger with incomplete insurance or borrowing constraints; (3) the welfare analysis assumes equal inflation rates across racial groups, which is empirically supported but abstracts from possible differences in consumption baskets; (4) the discriminatory parameter κ₁ is time-invariant and unresponsive to monetary policy, so all channels are indirect (through business cycle dynamics); (5) the model assumes a representative firm with taste-based discrimination, abstracting from firm heterogeneity in discrimination and from customer or statistical discrimination; (6) the Shortfalls rule is a reduced-form approximation of the FOMC&amp;rsquo;s 2020 framework and may not capture all aspects of the actual policy change.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Shortfalls rule&lt;/strong&gt;: A Taylor-type monetary policy rule that responds symmetrically to inflation deviations from target but responds to unemployment deviations from steady state only when unemployment is above its steady-state level — not when it is below. This captures, in reduced form, the FOMC&amp;rsquo;s August 2020 revision from &amp;lsquo;deviations&amp;rsquo; to &amp;lsquo;shortfalls&amp;rsquo; of employment from maximum.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Deviations rule&lt;/strong&gt;: A symmetric Taylor-type interest rate rule that responds to deviations of both inflation and unemployment from their respective steady-state values, regardless of the direction of the unemployment deviation. The baseline monetary policy in the model before the 2020 FOMC framework change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Taste-based discrimination (κ₁)&lt;/strong&gt;: A per-period perceived cost κ₁ borne by employers for each period they employ a Black worker, following Becker (1971). In this model, κ₁ = 0.0292 (≈3.6 percent of the steady-state wage), is time-invariant, and is not directly altered by monetary policy — only indirectly through business cycle conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reservation productivity threshold (zRi)&lt;/strong&gt;: The minimum idiosyncratic productivity level at which it is profitable for a firm to either hire or retain a worker of type i. Because of κ₁, the Black reservation threshold exceeds the white threshold, generating higher endogenous separation rates and lower job-finding rates for Black workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model-based discrimination measures (Df_t, Dλ_t)&lt;/strong&gt;: Novel measures of the fraction of the Black labor force that is not hired (Df_t, hiring margin) or is fired (Dλ_t, separation margin) solely due to discrimination — i.e., workers whose idiosyncratic productivity exceeds the white reservation threshold but falls below the Black threshold. These are expressed as fractions of the Black labor force and compared to EEOC race-based charge data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption-equivalent welfare wedge (Ψ_t)&lt;/strong&gt;: The percentage increase in per-period consumption that must be given to Black households every period to equalize their welfare with that of white households, given the same stochastic future. Under the Deviations rule, this averages 2.4 percent. The change under the Shortfalls rule is effectively zero in quantitative terms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous separation&lt;/strong&gt;: A separation that occurs because a matched worker-firm pair draws an idiosyncratic productivity below the reservation threshold — as distinct from exogenous separations (random layoffs unrelated to productivity). The dominance of the separation margin in explaining the racial unemployment gap motivates the use of endogenous separations as a key model ingredient; prior search-and-discrimination models assumed exogenous exit.&lt;/p&gt;</description></item><item><title>Labor Share, Markups, and Input-Output Linkages – Evidence from the U.S. National Accounts</title><link>https://macropaperwarehouse.com/papers/labor-share-markups-and-input-output-linkages-evidence-from-the-u.s.-national-accounts/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/labor-share-markups-and-input-output-linkages-evidence-from-the-u.s.-national-accounts/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;The paper asks why the U.S. labor share has declined over the postwar period, and whether rising markups or capital deepening (automation, falling capital prices) is the primary driver. The authors argue that the existing literature lacks consensus partly because micro-level studies weight producers by sales shares rather than Domar weights, which are gross-output-to-GDP ratios that correctly capture how sectoral changes propagate through the input-output structure. When intermediate inputs are themselves marked up by their producers and then re-marked-up by downstream firms (&amp;ldquo;double marginalization&amp;rdquo;), a modest sectoral markup increase is amplified into a substantially larger aggregate effect.&lt;/p&gt;
&lt;p&gt;The empirical framework is a two-sector (goods versus services) multisector extension of the Farhi-Gourio (2018) model with Cobb-Douglas production functions and monopolistic competition in the Dixit-Stiglitz tradition. The model is calibrated to three balanced-growth-path subperiods — 1957–1973, 1984–2000, and 2001–2016 — using U.S. NIPA data covering gross output, intermediate inputs, compensation, capital stocks, investment, and sectoral price-dividend ratios from Kenneth French&amp;rsquo;s data library. The unobservable user cost of capital, which is needed to separate normal capital returns from markups (factorless income), is backed out from the model&amp;rsquo;s Euler equation via the Gordon growth formula applied to sectoral price-dividend ratios and includes a risk premium.&lt;/p&gt;
&lt;p&gt;Main quantitative findings: The aggregate labor share fell 4.9 percentage points (pp) from 1957–1973 to 2001–2016. Aggregate markups rose 6.6 pp (from 1.072 to 1.138), more than either sector&amp;rsquo;s standalone increase, because double marginalization through input-output linkages amplifies sectoral markups into a larger aggregate effect. Sectoral gross-output markups rose approximately 3.8 pp in goods (1.039 to 1.077) and 3.3 pp in services (1.034 to 1.067). In the top-down counterfactual holding markups constant at their 1957–73 levels, the labor share falls only 0.5 pp instead of 4.9 pp — markups account for 4.4 pp of the total 4.9 pp decline. Holding labor output elasticities constant instead yields only a 2.2 pp decline; holding materials elasticities constant reduces the decline by 0.8 pp; holding structural change (sector output weights) constant causes the labor share to fall 7.7 pp — meaning structural reallocation to services offset 2.8 pp of the decline. In a bottom-up Taylor decomposition, the first-order direct effects of rising markups account for 5.0 pp and falling labor output elasticities account for 4.4 pp — together nearly twice the actual 4.9 pp decline, confirming the Grossman-Oberfield (2021) observation that individual candidate forces over-explain the total. The offsetting effects that reconcile the over-explanation are: (i) the interaction of falling goods-sector labor elasticities with structural change toward services (which have a higher and slightly rising labor elasticity) offsets 3.6 pp, and (ii) the interaction of rising markups with changing sector weights offsets a further 0.8 pp; the aggregate labor output elasticity αL barely changes (0.794 to 0.788) because capital deepening in goods (goods value-added labor elasticity fell from 0.807 to 0.700) is fully offset by reallocation to services (services value-added labor elasticity rose from 0.790 to 0.814). The final-output share of goods fell by more than half, from 0.460 to 0.194. Materials intensities rose in both sectors (goods non-intermediate factor share fell from 0.374 to 0.353; services from 0.625 to 0.572), amplifying double marginalization over time. The user cost of capital declined from roughly 13.8% to 12.3% in aggregate, driven by falling expected discount rates (from ~6.1% to ~3.4%), partially offset by rising depreciation rates. When IPP capital is excluded from NIPA measurement, the aggregate labor share declines by only 1 pp (from 0.743 to 0.732), consistent with Koh et al. (2021).&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s core implication for the debate is that forces concentrated in the goods sector — capital deepening, automation, globalization, declining union power — cannot account for the aggregate labor share decline because the goods sector shrank dramatically and structural change to services largely offsets goods-specific capital deepening. A credible candidate explanation must affect both goods and services with similar strength, and rising markups do: sectoral gross-output markups increased by similar amounts in both sectors (roughly 3.3–3.8 pp each), and input-output linkages amplify their aggregate impact substantially.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-model-structure-and-why-does-it-differ-from-one-sector-models"&gt;Q1. What is the model structure and why does it differ from one-sector models?&lt;/h3&gt;
&lt;p&gt;The model is a two-sector (goods, services) extension of Farhi-Gourio (2018) with Epstein-Zin preferences, Dixit-Stiglitz aggregation of varieties within each sector, Cobb-Douglas production in capital, labor, and intermediate inputs from both sectors, and sector-specific markups under monopolistic competition. The two-sector structure is essential because (i) labor shares differ substantially across sectors at any point in time, (ii) they evolve differently over time, and (iii) goods production is far more materials-intensive than services. A one-sector model cannot capture the double marginalization amplification, the input-output linkages between sectors, or the offsetting effects of structural change.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identification-strategy-for-separating-output-elasticities-from-markups"&gt;Q2. What is the identification strategy for separating output elasticities from markups?&lt;/h3&gt;
&lt;p&gt;A well-known challenge is that one must split the residual between capital&amp;rsquo;s normal return and pure profit (markup). The authors do not use micro production data. Instead they calibrate the user cost of capital from the model&amp;rsquo;s balanced-growth-path Euler equation: ρj is inferred from the Gordon growth formula applied to sectoral price-dividend ratios from Kenneth French&amp;rsquo;s data library. Given ρj, the depreciation-plus-capital-loss term δj + γQ is inferred from the sectoral investment-capital ratio. The markup then equals sectoral gross output value divided by the sum of all observed factor payments (labor compensation, materials costs) plus the imputed capital cost (user cost times capital stock). Output elasticities of each factor equal their respective cost shares in total factor payments, a standard Cobb-Douglas result.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-threats-to-identification-and-how-are-they-addressed"&gt;Q3. What are the main threats to identification and how are they addressed?&lt;/h3&gt;
&lt;p&gt;Three main threats are addressed. First, the price-dividend ratio (the key input for ρj) covers only listed corporations, not all private firms; the authors note listed firms account for about 60% of business capital, and Atkeson, Heathcote, and Perri (2025) find very similar rates of return using a broader measure. Second, the shift toward share repurchases rather than cash dividends may understate payout yield and overstate the fall in ρ, inflating markups; the authors rerun the model using Boudoukh et al. (2007) repurchase-adjusted yields and find aggregate markups still increase by 5 pp (vs. 7 pp in the baseline). Third, the balanced-growth-path assumption imposes constant ratios within each subperiod, which may be violated; the robustness exercise recalibrating with 2016 end-of-sample values yields nearly identical conclusions.&lt;/p&gt;
&lt;h3 id="q4-how-is-double-marginalization-measured-and-why-does-it-matter-so-much"&gt;Q4. How is double marginalization measured and why does it matter so much?&lt;/h3&gt;
&lt;p&gt;Double marginalization arises because approximately half of U.S. gross output value is materials costs, and those inputs are purchased from monopolistically competitive suppliers who charge a markup. When the downstream firm marks up its own price, it marks up the cost of already-marked-up inputs a second time. Formally, the aggregate markup exceeds any sectoral markup because intermediate goods get embedded in final goods through the Leontief inverse (Domar weights). The paper proves in Proposition 3 that aggregate markups are the same in gross-output and value-added models, but sectoral value-added markups are always larger than gross-output markups; this means taking simple cost- or revenue-weighted averages of sectoral value-added markups overstates the implied market power and misrepresents the channel. Materials intensities rose in both sectors over the sample, so double marginalization has itself increased over time, adding to the aggregate markup rise beyond what sectoral gross-output markups alone would imply.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-across-sectors"&gt;Q5. What heterogeneity is documented across sectors?&lt;/h3&gt;
&lt;p&gt;Goods and services differ in three key respects that are quantified: (i) Materials intensity: goods gross-output materials share is roughly 0.60 versus 0.40 for services (2001-16 averages), making double marginalization far stronger in goods. (ii) Capital deepening: the goods value-added labor elasticity ˜αLg fell from 0.807 to 0.700 (-10.7 pp) while services ˜αLs rose from 0.790 to 0.814 (+2.3 pp). (iii) Domar weights: the goods Domar weight Φg fell from 1.018 to 0.572 while services Φs rose from 0.933 to 1.289, reflecting the shift of economic activity toward services. Despite these differences, sectoral gross-output markups increased by similar amounts in both sectors (3.8 pp goods, 3.3 pp services), which is the main reason markups can explain the aggregate decline while sector-specific capital deepening cannot.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Four sets of robustness exercises are presented. (1) Balanced-growth-path assumption: the model is recalibrated using only 2016 end-of-sample values for the final period; results are nearly unchanged, though markups come out slightly higher. (2) Dividend measurement: repurchase-adjusted payout yields from Boudoukh et al. (2007) are used; aggregate markups still increase by 5 pp rather than 7 pp, so the conclusion is unchanged though the magnitude is modestly smaller. (3) Missing capital — organizational capital: using Crouzet-Eberly (2021) estimates, including organizational capital reduces the markup level (from 1.138 to 1.092 in 2001-16) but barely changes the markup increase (from 4.6 pp to 3.9 pp). (4) Missing capital — land: industrial and commercial land values over 2002-16 averaged roughly $2 trillion versus a private non-real-estate capital stock of $16.6 trillion; eliminating the markup increase would require a 27% rise in the capital-output ratio, but land can provide at most a 12% increase even under extremely counterfactual assumptions. (5) Alternative user cost: using Barkai (2020)&amp;rsquo;s Aaa interest rate yields aggregate markups increasing from 1.101 to 1.151 (1984-2000 to 2001-16), similar to baseline. (6) IPP capital: excluding IPP from NIPA yields only a 1 pp labor share decline, consistent with Koh et al. (2021). (7) Intangible capital generally: including intangible capital in the NIPA does not change the importance of markups.&lt;/p&gt;
&lt;h3 id="q7-how-do-the-authors-reconcile-their-low-gross-output-markups-with-the-much-higher-firm-level-markups-found-by-de-loecker-eeckhout-and-unger-2020"&gt;Q7. How do the authors reconcile their low gross-output markups with the much higher firm-level markups found by De Loecker, Eeckhout, and Unger (2020)?&lt;/h3&gt;
&lt;p&gt;The reconciliation has two parts. First, weighting: De Loecker et al. use sales-weighted markups, whereas the model-correct weighting in this context is harmonic cost-weighting (Hasenzagl and Perez, 2023); cost-weighted markups in this paper grow only 5 pp (from 1.193 to 1.246) versus 21 pp for sales-weighted markups, substantially narrowing the gap. Second, returns to scale and fixed costs: the paper assumes constant returns to scale and no fixed costs, so all markup revenue is pure profit. Firm-level studies assume fixed costs exist, meaning their markups must cover both pure profits and overhead, so markups are mechanically larger. A fixed cost share of about 15% of production costs accounts for the remaining difference between the two estimates. Importantly, both approaches produce similar economic profit rates: this paper finds sales-weighted profit rates of 4.5% (1984-2000) rising to 6.6% (2001-16), similar to De Loecker et al.&amp;rsquo;s finding of profit rates rising from 1% in 1980 to 8% in 2016.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-role-of-structural-change-and-why-does-it-offset-capital-deepening-but-not-markups"&gt;Q8. What is the role of structural change, and why does it offset capital deepening but not markups?&lt;/h3&gt;
&lt;p&gt;Structural change — the reallocation of final-output expenditure shares away from goods toward services — acts as a natural counterweight when a factor depresses labor share only in the shrinking sector. Capital deepening (falling goods labor elasticity) is concentrated in goods; as goods&amp;rsquo; expenditure share fell from 0.460 to 0.194, the weight placed on goods in the aggregate labor share shrank, largely undoing the direct effect of capital deepening on aggregate labor share. The second-order interaction term in the bottom-up decomposition confirms this: the interaction of falling labor elasticities with changing sector weights offsets 3.6 pp. In contrast, markups rose by similar amounts in both goods and services, so there is no equivalent shrinking-sector effect to offset the markup increase; summing the direct markup effect (−5.0 pp) with the markup-weight interaction (+0.8 pp) gives approximately the full observed decline.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-treat-the-possibility-of-labor-monopsony-as-an-explanation"&gt;Q9. How does the paper treat the possibility of labor monopsony as an explanation?&lt;/h3&gt;
&lt;p&gt;The paper acknowledges that labor monopsony (markdowns over wages) could in principle produce a gap between price and marginal cost similar to product markups. However, the authors argue the evidence does not support a role for increasing markdowns in driving the aggregate labor share trend: Yeh, Macaluso, and Hershbein (2022) find large markdowns in manufacturing but no role for them in explaining the time series of manufacturing labor share; Deb et al. (2022), allowing for both markups and markdowns, attribute changes in the price-marginal-cost gap to markups; Kirov and Traina (2023) find similar evidence in manufacturing. The authors therefore interpret their factorless-income estimates as markups rather than markdowns.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy implication is that explanations for the labor share decline should focus on product market power (markups) operating across both the goods and services sectors, not primarily on capital-deepening forces such as automation or falling capital prices. The paper does not directly propose policy remedies, but the results imply that policies targeting capital deepening or trade-induced displacement alone cannot fully explain or reverse the aggregate labor share trend. The analysis is scoped to the U.S. private economy excluding real estate, 1957–2016, and the two-sector decomposition. The authors acknowledge the NIPA-based approach cannot directly speak to the firm-level sources of increasing markups (market concentration, fixed costs, intangibles), leaving the microeconomic explanation for rising sectoral markups to future research. Extension to finer industry disaggregations and other countries is flagged as a direct next step.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-papers-contribution-relative-to-farhi-gourio-2018-and-karabarbounis-neiman-2014"&gt;Q11. What is the paper&amp;rsquo;s contribution relative to Farhi-Gourio (2018) and Karabarbounis-Neiman (2014)?&lt;/h3&gt;
&lt;p&gt;Farhi-Gourio (2018) is a one-sector model calibrated at the aggregate level; this paper extends it to two sectors with explicit input-output linkages, allowing the decomposition to distinguish sector-specific from aggregate forces and to quantify double-marginalization amplification. Karabarbounis-Neiman (2014) attributed the labor share decline primarily to falling relative prices of capital (capital deepening) driven by an elasticity of substitution between capital and labor exceeding one; this paper&amp;rsquo;s calibration finds that the aggregate output elasticity of labor barely changes (0.794 to 0.788), which is inconsistent with capital deepening as the dominant aggregate force, and notes that evidence from Herrendorf et al. (2015) and Oberfield-Raval (2021) suggests the elasticity of substitution is below one in most of the goods sector. Moreira (2022) also uses an input-output model but does not allow markups by intermediate producers, which this paper shows is quantitatively crucial.&lt;/p&gt;
&lt;h3 id="q12-why-does-the-paper-use-nipa-data-rather-than-firm--or-establishment-level-data"&gt;Q12. Why does the paper use NIPA data rather than firm- or establishment-level data?&lt;/h3&gt;
&lt;p&gt;NIPA data have four advantages in this context: (i) they cover all market activity rather than just publicly listed or large firms; (ii) they capture inter-sectoral input-output linkages that micro datasets lack; (iii) they include broad coverage of intangible assets (IPP) following the 1999 and 2013 revisions; and (iv) they respect standard accounting adding-up constraints, ensuring that sectoral forces aggregate consistently to the macro level. The NIPA-based calibration also has limited data requirements, making it feasible to extend the analysis back to the late 1950s and, potentially, to other countries. The main limitation is that NIPA data are available only at the two-sector level of aggregation for the full postwar period, due to the switch from SIC to NAICS classification in 1997 and the aggregated reporting of some items like proprietors&amp;rsquo; income.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Domar weight&lt;/strong&gt;: The ratio of a sector&amp;rsquo;s gross output value to aggregate final output (GDP). Unlike expenditure weights, Domar weights exceed one when summed across sectors because they capture how a sector&amp;rsquo;s output is both a direct contributor to final demand and an indirect contributor through its use as intermediate inputs elsewhere. The paper uses Domar weights as the correct aggregation weights for sectoral labor shares, showing that properly accounting for input-output linkages through these weights is essential for connecting sectoral forces to aggregate outcomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double marginalization&lt;/strong&gt;: The amplification of sectoral markups at the aggregate level that occurs because intermediate inputs are priced above marginal cost by their producers (first markup) and then purchased and re-priced above marginal cost by downstream firms (second markup). In this paper&amp;rsquo;s model, double marginalization causes the aggregate markup to exceed either sector&amp;rsquo;s standalone gross-output markup; with roughly half of U.S. gross output being materials cost, the amplification is quantitatively large (aggregate markups of 6.6 pp increase versus sectoral increases of only 3.3–3.8 pp).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gross-output markup&lt;/strong&gt;: The ratio of a sector&amp;rsquo;s gross output value to the sum of all factor payments at the gross-output level (capital user costs times capital stock, plus labor compensation, plus the cost of intermediate inputs from all sectors). Under perfect competition this ratio equals one; deviations above one represent market power. This differs from value-added markups, which divide value added by only capital and labor payments, and are therefore mechanically inflated in materials-intensive sectors via double marginalization even when gross-output markups are identical across sectors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Risky balanced growth path (RBGP)&lt;/strong&gt;: The equilibrium concept used for calibration, extending the standard balanced growth path to allow for rare disaster shocks (Farhi-Gourio). Along the RBGP, expected variables grow at constant rates, but occasional level shifts occur when the rare disaster shock materializes. This allows the model to have realistic risk premia embedded in the discount rate ρ while maintaining analytically tractable solutions; the calibration avoids modeling transitional dynamics and instead compares the RBGP parameters across sub-periods.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Output elasticity of labor (αL)&lt;/strong&gt;: The Cobb-Douglas coefficient on labor in the production function, equal under the paper&amp;rsquo;s calibration to each factor&amp;rsquo;s cost share in total factor payments. Changes in αL represent capital deepening (automation, falling capital prices) when αL falls because capital displaces labor in production. The key finding is that αL barely changes at the aggregate level (0.794 to 0.788) over the full 1957–2016 period because capital deepening in goods is offset by structural change toward services.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Factorless income&lt;/strong&gt;: Income that remains after subtracting payments to labor (at market wages) and payments to capital (at normal user cost rates) from gross output. In this model, factorless income equals markup revenue (the portion of output value above total factor payments). Rising factorless income / markups are the mirror image of the declining labor share when the output elasticity of labor does not change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural change (in this paper&amp;rsquo;s sense)&lt;/strong&gt;: The reallocation of final-output expenditure shares across goods and services sectors over time, captured by changes in the expenditure weights ϕj. The paper documents that the goods final-output share fell by more than half (from 0.460 to 0.194) over 1957–2016. Structural change acts as a counterweight to any force concentrated in the goods sector: as goods&amp;rsquo; weight shrinks, the aggregate labor share becomes more determined by services.&lt;/p&gt;</description></item><item><title>Labour Market Power and the Effects of Fiscal Policy</title><link>https://macropaperwarehouse.com/papers/labour-market-power-and-the-effects-of-fiscal-policy/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/labour-market-power-and-the-effects-of-fiscal-policy/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper proposes a novel fiscal transmission channel through which government spending expansions reduce employer monopsony power in the labor market, generating larger fiscal multipliers and stronger distributional consequences than standard models predict.&lt;/p&gt;
&lt;p&gt;Standard New Keynesian models rely on two transmission channels with contested empirical support: a negative wealth effect on labor supply (which moves workers to supply more hours when taxes rise) and countercyclical price markups (which fall in booms, raising labor demand). The evidence on both is ambiguous. This paper introduces a third channel — countercyclical monopsony power — that operates independently of, and interacts with, the other two.&lt;/p&gt;
&lt;p&gt;The theoretical framework is a Two-Agent New Keynesian (TANK) model, extending Cantore and Freund (2021). There are two household types: workers (fraction λ = 0.8), who supply labor and have limited financial market access, and capitalists (fraction 1 − λ = 0.2), who earn profit income. Intermediate-good firms compete monopsonistically in local labor markets, paying wages below the marginal revenue product. The wage markdown μ = η/(η+1), where η is the wage elasticity of labor supply to the individual firm. Workers value both pay and non-pay job characteristics (firm location, culture, flexibility), with heterogeneous idiosyncratic preferences drawn from a type-1 extreme value distribution. This differentiation, following Card et al. (2018), gives firms wage-setting power because they cannot observe individual preferences.&lt;/p&gt;
&lt;p&gt;The key mechanism is that η depends endogenously on workers&amp;rsquo; labor earnings (wt·nt) and their marginal utility of income (uW_c,t): η = θ·uW_c,t·wt·nW_t + 1/φ. When government spending rises, it increases both labor income and — because higher current or future taxes reduce lifetime net income — workers&amp;rsquo; marginal valuation of income. Both forces unambiguously raise η, flattening the firm-level labor supply curve, reducing the marginal cost of labor for firms seeking to attract workers, and driving wages up toward the marginal revenue product. Employment and output rise; profits fall and are redistributed toward workers.&lt;/p&gt;
&lt;p&gt;In the calibrated baseline (steady-state markdown μ = 2/3, i.e., wages at two-thirds of marginal revenue products, calibrated to Yeh et al. 2022), the impact fiscal multiplier is approximately 0.6 under monopsonistic competition compared to slightly less than 0.4 under perfect competition — a difference attributable entirely to the countercyclical-monopsony channel. The wage markdown rises by approximately 0.3 percentage points on impact following a 1% of GDP government spending shock, roughly twice the response observed when the steady-state markdown is 0.9 rather than 0.67.&lt;/p&gt;
&lt;p&gt;The amplification from countercyclical monopsony is strongest when the wealth effect on hours worked is near zero — the baseline calibration consistent with Schmitt-Grohé and Uribe (2012) and Galí et al. (2012). As the wealth elasticity of hours increases, the markdown and output response to spending shocks weaken, because a larger hours response implies a smaller consumption response, which reduces the marginal utility channel. The degree of price stickiness has little effect on the markdown response.&lt;/p&gt;
&lt;p&gt;The channel is amplified when workers bear more of the fiscal burden — either through profit redistribution to workers (amplification rises from approximately 0.25 in the no-redistribution baseline to approximately 0.4 when half of profit income is redistributed to workers) or through regressive taxation. Progressively redistributing the tax burden toward capitalists weakens the countercyclical-monopsony channel, which runs counter to the standard cyclical-inequality channel (Bilbiie 2020) that predicts larger multipliers with progressive taxation.&lt;/p&gt;
&lt;p&gt;The empirical validation uses an expectations-augmented VAR estimated on quarterly U.S. data from 1981Q3 to 2019Q4 (macroeconomic variables) and 2000Q4 to 2019Q4 (monopsony measure). Government spending shocks are identified via recursive ordering (government spending ordered first), controlling for professional forecasters&amp;rsquo; spending growth expectations (following Auerbach-Gorodnichenko 2012), the real interest rate using the Wu-Xia shadow policy rate, and the average tax rate. The inverse monopsony measure — the wage elasticity of worker-firm separations — is estimated by extending Langella and Manning (2021) to quarterly frequency using SIPP microdata, controlling for demographics, industry, occupation, human capital, and time effects via complementary log-log regressions month by month. The VAR impulse responses confirm the model&amp;rsquo;s central prediction: government spending expansions raise the wage elasticity of separations (reducing employer market power), raise labor income, reduce profits, and generate substantial output increases.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-in-the-empirical-var-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy in the empirical VAR and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses a recursive (Cholesky) identification scheme with government spending ordered first, following Blanchard and Perotti (2002). The identifying assumption is that government spending does not respond to economic conditions within the same quarter due to decision and implementation lags. Anticipation effects are addressed by including a fiscal news variable — professional forecasters&amp;rsquo; one-period-ahead spending growth forecast from the Survey of Professional Forecasters — following Auerbach and Gorodnichenko (2012). The innovation in government spending orthogonal to this forecast is taken as the exogenous surprise shock. The real interest rate (Wu-Xia shadow federal funds rate, which captures unconventional monetary policy at the zero lower bound) and the average tax rate are included to control for monetary policy stance and financing mix. A key threat the paper acknowledges concerns the separation elasticity estimates: the monopsony literature recognizes biases from insufficient controls for alternative wage offers, unobserved heterogeneity, and lack of firm-level exogenous wage variation. The authors follow Langella and Manning (2021) in arguing that these biases are roughly constant over time, so changes in the estimated separation elasticity still reflect changes in true monopsony power.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-key-mechanism-through-which-government-spending-reduces-monopsony-power"&gt;Q2. What is the key mechanism through which government spending reduces monopsony power?&lt;/h3&gt;
&lt;p&gt;Two reinforcing forces simultaneously raise the wage elasticity of labor supply to individual firms (η). First, higher government spending raises labor income, which increases the dollar magnitude of pay differences between firms, making workers more responsive to relative pay. Second, higher current or future taxes reduce workers&amp;rsquo; lifetime net income, raising their marginal valuation of income (marginal utility of consumption, uW_c,t). Workers facing a tighter budget place greater relative weight on pay versus non-pay job characteristics, further increasing their responsiveness to firm-level wages. Both effects increase η unambiguously for government spending shocks (unlike productivity shocks, where the two forces can offset each other). Higher η flattens the firm-level labor supply curve, compresses the gap between the marginal cost of labor and the wage, and induces firms to raise wages toward the marginal revenue product. Employment and output rise while profits decline, redistributing income from capitalists to workers.&lt;/p&gt;
&lt;h3 id="q3-how-is-monopsony-modeled-and-why-does-the-paper-use-a-discrete-choice-rather-than-ces-approach"&gt;Q3. How is monopsony modeled, and why does the paper use a discrete choice rather than CES approach?&lt;/h3&gt;
&lt;p&gt;The paper adopts a discrete workplace choice model following Card et al. (2018), where workers draw idiosyncratic preferences over non-pay job characteristics from a type-1 extreme value distribution each period. Firms cannot observe individual preferences and set a posted wage. Standard logit calculations yield the wage elasticity of firm-level labor supply as η = θ·uW_c,t·wt·nW_t + 1/φ, where θ is the inverse importance of non-pay characteristics and 1/φ is the intensive-margin (hours) elasticity. Under CES preferences (used by Berger et al. 2022, Alpanda and Zubairy 2021), the wage markdown is constant in equilibrium — analogous to constant price markups under CES monopolistic competition — which eliminates the time variation in monopsony power that is the paper&amp;rsquo;s central object of study. The discrete choice framework generates endogenous variation in η through the endogenous terms wt·nW_t and uW_c,t. Berger et al. (2022) show that the CES approach is a special case of the discrete choice model under restrictive assumptions about individual hours responses; the paper intentionally avoids those assumptions.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-across-calibrations-is-documented-regarding-the-strength-of-the-monopsony-channel"&gt;Q4. What heterogeneity across calibrations is documented regarding the strength of the monopsony channel?&lt;/h3&gt;
&lt;p&gt;The paper documents several dimensions of heterogeneity: (1) Steady-state markdown: the relationship between the steady-state markdown and the markdown&amp;rsquo;s response to government spending is hump-shaped (inverted U-shape). At the baseline value of 0.67, the markdown rises by approximately 0.3 percentage points; at a steady-state markdown of 0.9, the response is roughly half as large. Perfect competition (markdown = 1) and maximum monopsony (markdown → 0) both imply no response. (2) Wealth effect on labor supply (χ): as χ increases from zero (baseline, near-GHH preferences) to one (strong wealth effect), the markdown response and the output amplification decline monotonically. With a near-zero wealth effect (baseline), amplification relative to the perfect-competition counterfactual is approximately 0.25 percentage points of steady-state GDP; it diminishes substantially as χ rises. (3) Profit redistribution (φd): output amplification rises from approximately 0.25 (no redistribution, baseline) to approximately 0.4 when half of profits are redistributed to workers. (4) Tax progressivity (φτ): the channel is stronger under regressive taxation (more of the burden falling on workers) and weaker under progressive taxation, in contrast to the cyclical-inequality channel. (5) Degree of tax financing (φg): higher contemporaneous tax financing strengthens the channel because it raises workers&amp;rsquo; current marginal valuation of income more directly. (6) Price stickiness (ξ): changing price adjustment costs has little effect on the markdown response and the countercyclical-monopsony amplification.&lt;/p&gt;
&lt;h3 id="q5-how-is-the-separation-elasticity-measured-and-linked-to-the-models-concept-of-monopsony-power"&gt;Q5. How is the separation elasticity measured and linked to the model&amp;rsquo;s concept of monopsony power?&lt;/h3&gt;
&lt;p&gt;The separation elasticity γ is the wage elasticity of worker-firm separations: the percentage change in a firm&amp;rsquo;s separation rate in response to a 1% change in the wage. In the model, γ is shown to be proportional to η − 1/φ (the extensive-margin component of labor supply elasticity to the firm), because firm size and separation rate are linked through a constant elasticity derived from the logit choice structure. Empirically, the paper extends Langella and Manning (2021) to quarterly frequency using SIPP data from 2000Q4 to 2019Q4. Month-by-month complementary log-log regressions of separation dummies on residualized log hourly wages (purged of demographic, industry, occupation, human capital, and time effects) yield time-varying quarterly estimates of γ. A higher γ (less negative, since separations fall with higher wages) indicates lower monopsony power. The VAR incorporates this time-varying series as the inverse monopsony measure.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-countercyclical-monopsony-channel-interact-with-the-wealth-effect-and-price-markup-channels"&gt;Q6. How does the countercyclical-monopsony channel interact with the wealth effect and price markup channels?&lt;/h3&gt;
&lt;p&gt;The three channels interact in both complementary and partially offsetting ways. The wealth effect on hours worked (χ &amp;gt; 0) independently shifts the market labor supply curve rightward when taxes rise, increasing employment. However, a larger hours response implies a smaller consumption response, which reduces the increase in workers&amp;rsquo; marginal utility of consumption. Since uW_c,t is a key driver of η, a stronger wealth effect on hours dampens the countercyclical-monopsony channel. Similarly, the countercyclical price markup channel (ξ &amp;gt; 0) raises the marginal revenue product of labor when government spending pushes up demand, boosting employment through an independent channel that also raises labor income — which in turn reinforces η. Yet changing price stickiness has quantitatively little effect on the markdown response in the calibrated model. Income redistribution between agent types mediates the interaction: when capitalists bear most of the tax burden (progressive taxation), workers&amp;rsquo; marginal utility of income rises less, weakening the monopsony channel. When workers bear the burden (regressive taxation or profit redistribution), the monopsony channel is strengthened.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-distributional-consequences-of-the-countercyclical-monopsony-channel"&gt;Q7. What are the distributional consequences of the countercyclical-monopsony channel?&lt;/h3&gt;
&lt;p&gt;When government spending rises, the reduction in employer market power forces firms to pay wages closer to the marginal revenue product, increasing labor income and decreasing profits. This redistribution from capitalists (profit recipients) to workers operates through the wage markdown declining (i.e., markup rising toward one). Under monopsonistic competition with endogenous employer market power, this redistribution is stronger than under perfect competition, where only the price markup channel operates. The VAR evidence confirms these distributional predictions: government spending shocks reduce corporate profits (after taxes) and raise labor income in U.S. data. In the model, this redistribution also feeds back into the mechanism: workers facing declining after-tax income (or receiving a portion of declining profits) place greater weight on pay in their workplace choices, further eroding employer market power.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-cantore-and-freund-2021-and-the-tank-literature-on-fiscal-multipliers"&gt;Q8. How does this paper relate to Cantore and Freund (2021) and the TANK literature on fiscal multipliers?&lt;/h3&gt;
&lt;p&gt;The paper extends the worker-capitalist TANK model of Cantore and Freund (2021), who introduced capitalists that do not participate in the labor market to avoid the criticism (Broer et al. 2019, 2021) that the Bilbiie (2008, 2020) cyclical-inequality channel relies on countercyclical profit income inducing rich households to supply more labor. The Cantore-Freund framework delivers income redistribution between high-MPC workers and low-MPC capitalists without relying on labor supply responses of the rich. This paper adds monopsonistic competition to that framework, introducing a new form of cyclical variation in inequality through time-varying wage markdowns. The interaction with the Bilbiie cyclical-inequality channel is analyzed formally: in particular, tax progressivity has opposing effects under the two channels — progressive taxation amplifies the Bilbiie effect (redistribution to high-MPC workers) but weakens the monopsony channel (capitalists bear more of the tax burden, reducing workers&amp;rsquo; marginal valuation of income).&lt;/p&gt;
&lt;h3 id="q9-what-robustness-is-discussed-or-implied-regarding-the-empirical-var"&gt;Q9. What robustness is discussed or implied regarding the empirical VAR?&lt;/h3&gt;
&lt;p&gt;The paper addresses robustness primarily through the following design choices: (1) Use of the Wu-Xia shadow federal funds rate rather than the actual federal funds rate, to capture monetary policy stance during the zero lower bound period; (2) inclusion of the spending growth forecast variable to control for anticipation effects; (3) inclusion of the average tax rate as a control for fiscal financing; (4) detrending all VAR variables as deviations from linear trends. The separation elasticity itself is shown to be robustly procyclical across three detrending methods (linear, linear-quadratic, and HP-filter with λ=1600), with R² values of 49.9%, 43.6%, and 17.1%, respectively, and regression slopes of 1.52, 1.40, and 1.51 in each case. The paper notes that standard biases in separation elasticity estimation (from unobserved heterogeneity, inadequate controls for alternative offers, absence of firm-level exogenous wage variation) are likely roughly constant over time, which validates using changes in the estimated elasticity as changes in true monopsony power, following Langella and Manning (2021, p. 2942). The sample for the monopsony series (2000Q4–2019Q4) is shorter than the macro VAR sample (1981Q3–2019Q4) due to data availability.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-analytical-results-from-the-simplified-model"&gt;Q10. What are the analytical results from the simplified model?&lt;/h3&gt;
&lt;p&gt;Under flexible prices (no price markup channel), no wealth effect on hours worked (χ = 0), no financial market access for workers (ψW → ∞), full tax financing, and no profit redistribution, the paper derives closed-form expressions for output, labor income, and profits following a government spending shock. Output and labor earnings respond positively to spending only when θ is finite (workers value both pay and non-pay characteristics, so η is endogenous). When θ = ∞ (workers only care about pay → perfect competition with constant η) or θ = 0 (workers only care about non-pay → constant η again), government spending has zero output effect. The parameter Γ = 0 in both limiting cases. For intermediate θ, Γ &amp;gt; 0, government spending raises output and redistributes income from capitalists to workers. This establishes that the countercyclical-monopsony channel is the sole mechanism at work in the simplified model and that it requires intermediate values of workers&amp;rsquo; preference for non-pay characteristics.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper implies that fiscal multipliers may be larger than standard New Keynesian models predict if labor markets exhibit significant employer monopsony power — calibrated to produce a steady-state wage markdown of 2/3 (wages at two-thirds of marginal revenue products), consistent with empirical estimates for the U.S. The countercyclical-monopsony channel provides expansionary effects of government spending even in models where the wealth effect on labor supply is negligible and price markups do not decline. The distributional consequences of fiscal expansions are also stronger under monopsony: income shifts from profit recipients (capitalists) to wage earners more substantially. Scope conditions include: the channel is weaker with stronger wealth effects on hours worked; it is stronger when government spending is financed through current taxes rather than deficit (more tax financing raises workers&amp;rsquo; marginal valuation of income more sharply); it is stronger under regressive rather than progressive taxation; and it is stronger when profit income is redistributed to workers. Progressivity of taxation affects the monopsony and cyclical-inequality channels in opposing directions, implying that the optimal tax structure from a fiscal multiplier perspective depends on which channel is quantitatively dominant.&lt;/p&gt;
&lt;h3 id="q12-what-prior-empirical-literature-on-cyclical-monopsony-power-does-this-paper-build-on-and-extend"&gt;Q12. What prior empirical literature on cyclical monopsony power does this paper build on and extend?&lt;/h3&gt;
&lt;p&gt;The paper builds on three prior empirical findings. First, substantial employer market power in U.S. labor markets (Berger et al. 2022; Langella and Manning 2021; Yeh et al. 2022). Second, unconditional countercyclicality of employer market power — Hirsch et al. (2018) for Germany, Bassier et al. (2022) for Oregon, and Webber (2022) for the U.S. all document that firms hold more monopsony power in slack labor markets. The paper&amp;rsquo;s own descriptive analysis confirms this procyclicality of the separation elasticity across multiple detrending methods. Third, Langella and Manning (2021) provide the estimation methodology for the separation elasticity using SIPP data. The paper&amp;rsquo;s extension is twofold: (a) it extends the Langella-Manning estimates to quarterly frequency and expands the sample to 2019Q4; and (b) it examines the conditional cyclicality of employer market power — specifically, how monopsony power responds to identified government spending shocks — which prior literature had not done.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Countercyclical monopsony channel&lt;/strong&gt;: The novel fiscal transmission mechanism proposed by the paper: government spending expansions endogenously reduce employer monopsony power by raising both labor income and workers&amp;rsquo; marginal valuation of income, which makes workers more responsive to relative pay differences across firms (higher η), compresses wage markdowns, and raises employment and output. The channel is &amp;lsquo;countercyclical&amp;rsquo; in that employer market power falls as spending rises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wage markdown (µ)&lt;/strong&gt;: The ratio of the wage paid to workers to the marginal revenue product of labor, defined as µ = η/(η+1), bounded between zero and one. A smaller µ implies a larger wedge between pay and marginal product, i.e., greater monopsony power. Perfect competition corresponds to µ = 1. In the baseline calibration µ = 2/3, meaning wages equal two-thirds of the marginal revenue product.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wage elasticity of labor supply to the individual firm (η)&lt;/strong&gt;: The key measure of firms&amp;rsquo; monopsony power in the model. Defined as η = θ·uW_c,t·wt·nW_t + 1/φ, where 1/φ is the intensive-margin (hours) elasticity. The extensive-margin component θ·uW_c,t·wt·nW_t determines how strongly a firm can attract workers from competitors by raising pay. Higher η means less monopsony power (wages closer to marginal revenue product); lower η means greater power.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Separation elasticity (γ)&lt;/strong&gt;: The empirical proxy for inverse monopsony power: the wage elasticity of worker-firm separations, measuring how steeply a firm&amp;rsquo;s separation rate falls when it pays higher wages. In the model, γ is proportional to the extensive-margin component of η. Estimated from SIPP microdata via month-by-month complementary log-log regressions of separation dummies on residualized log wages, following Langella and Manning (2021).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;New classical (idiosyncrasy) monopsony&lt;/strong&gt;: The modeling approach used in the paper, following Card et al. (2018), in which monopsony power arises from workers&amp;rsquo; heterogeneous preferences over non-pay job characteristics (location, culture, flexibility) rather than from search frictions or geographic isolation. Firms differ in non-pay attributes, and because firms cannot observe individual preferences, they have wage-setting power even with frictionless worker flows between firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cyclical-inequality channel&lt;/strong&gt;: A fiscal transmission mechanism from the HANK/TANK literature (Bilbiie 2008, 2020): government spending redistributes income from low-MPC capitalists to high-MPC workers, amplifying the fiscal multiplier. The paper shows this channel interacts with the countercyclical-monopsony channel in conflicting ways — progressive taxation strengthens the cyclical-inequality channel but weakens the monopsony channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wealth effect on labor supply (χ)&lt;/strong&gt;: Parameterized via the Jaimovich-Rebelo (2009) utility function, χ governs how strongly a decline in household lifetime income (due to higher taxes) induces workers to supply more hours. The baseline calibration sets χ → 0, consistent with near-GHH preferences and estimates in Schmitt-Grohé and Uribe (2012). A higher χ dampens the countercyclical-monopsony channel by reducing the consumption response and thereby the marginal utility response.&lt;/p&gt;</description></item><item><title>Macroeconomic Effects of 'Free' Secondary Schooling in the Developing World</title><link>https://macropaperwarehouse.com/papers/macroeconomic-effects-of-free-secondary-schooling-in-the-developing-world/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/macroeconomic-effects-of-free-secondary-schooling-in-the-developing-world/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether publicly funded (&amp;ldquo;free&amp;rdquo;) secondary schooling in developing countries raises GDP per capita. The question is policy-relevant because many low-income countries — including Ghana, Kenya, Tanzania, Uganda, and others listed in the paper&amp;rsquo;s appendix — have recently adopted or are considering such policies, motivated by the combination of low secondary enrollment (roughly one-third of secondary-school-age children enrolled in the poorest countries, versus near-universal enrollment in rich countries) and evidence that credit constraints keep talented students out of school.&lt;/p&gt;
&lt;p&gt;The analysis is built around an overlapping-generations (OLG) model with heterogeneous households and credit constraints, estimated to match experimental evidence from a randomized controlled trial (RCT) in Ghana (Duflo, Dupas, and Kremer, 2021). The RCT randomly offered full four-year scholarships covering 100 percent of tuition and fees to approximately two thousand poor but high-ability students who had passed the Basic Education Certificate Examination (BECE) but had not enrolled in Senior High School (SHS). Scholarship winners were 27 percentage points more likely to complete secondary school than the control group, scored 0.16 standard deviations (equivalent to 7.6 percent wage gains in the model) higher on math and literacy tests, and experienced a 10.6 percent decline in fertility after 12 years.&lt;/p&gt;
&lt;p&gt;The model departs from standard human capital OLG models in three ways. First, it incorporates an explicit opportunity cost of schooling: teenagers who attend SHS forgo labor income during ages 15–19, which is economically significant given that secondary-school-age individuals are near their prime working years in developing countries. Second, the model includes a merit-based entrance exam (the BECE), so that removing the exam requirement as part of free schooling causes negative selection — the new marginal students induced to attend have lower average ability than those already attending. Third, the model features education-dependent fertility: more-educated households have fewer children (estimated fertility of 2.07 per less-educated family vs 1.19 per more-educated family, in line with Ghanaian Demographic and Health Survey data). The model also incorporates imperfect substitutability between skilled and unskilled labor (elasticity of substitution set to 4, following long-run cross-country estimates), savings wedges that match low liquid asset holdings, and Ghana&amp;rsquo;s actual progressive income tax schedule.&lt;/p&gt;
&lt;p&gt;The model is estimated using the Simulated Method of Moments (SMM) targeting ten moments — five non-experimental (aggregate population growth rate of 2.2 percent per year, aggregate SHS completion rate, SHS completion in the top and bottom test-score quartiles of the control group, and variance of the permanent component of log wages) and five experimental or quasi-experimental (RCT treatment effects on human capital, fertility, overall SHS completion, the Q4 vs Q1 difference in SHS completion, and the intergenerational schooling correlation from administrative data).&lt;/p&gt;
&lt;p&gt;The central quantitative finding is that nationwide free secondary schooling — eliminating both fees and the entrance-exam requirement — raises secondary school completion by about 12 percentage points (from 30 percent to 42 percent of the population) but reduces GDP per capita by approximately 1 percent in the long run. The 95 percent confidence interval for the GDP effect excludes any positive value (lower bound -4.2 percent, upper bound -0.7 percent), so the model can statistically reject any positive GDP impact. The direct fiscal cost of the policy is 1.4 percent of GDP, implying a total cost (direct cost plus lost GDP) of approximately 2.4 percent of GDP. Taxes per capita increase by 1.4 percent. Adult earnings rise by about 1.2 percent, but this is more than offset by a 7.5 percent decline in child earnings (the opportunity cost of schooling for newly enrolled students). The skilled-to-unskilled wage ratio falls by about 10 percent, reflecting general-equilibrium wage compression from the expanded supply of secondary graduates.&lt;/p&gt;
&lt;p&gt;Three counterfactual experiments decompose the negative GDP result. (i) Eliminating the opportunity cost of schooling reverses the GDP effect from -1.0 percent to +2.9 percent, a swing of nearly 4 percentage points — the dominant channel. (ii) Holding the ability distribution of new secondary attendees to match the experimental sample (removing negative selection) moves GDP from -1.0 percent to essentially 0, accounting for about 1 percentage point of the gap. (iii) Holding fertility constant for new secondary attendees moves GDP from -1.0 percent to +1.2 percent, contributing about 2.2 percentage points. When all three channels are shut down simultaneously, GDP rises by 6.9 percent — close to the naive back-of-the-envelope projection of 6 percent based on the RCT&amp;rsquo;s test-score estimates.&lt;/p&gt;
&lt;p&gt;As a policy comparison, an economy-wide improvement in schooling quality that raises test scores by 0.1 standard deviations (a conservative estimate consistent with randomized teacher-incentive interventions in India and Kenya) raises GDP per capita by 2.7 percent and increases SHS completion by 13.8 percentage points — more than free schooling and at lower fiscal cost (the policy pays for itself in equilibrium). Improving schooling quality avoids the negative selection and opportunity-cost channels because it raises human capital for both new and inframarginal students.&lt;/p&gt;
&lt;p&gt;On welfare and distribution, the policy is predominantly redistributive. The bottom 25 percent of parents gain welfare equivalent to a 7.3 percent increase in lifetime consumption, while the top 25 percent lose 4.2 percent. For children, the bottom 25 percent gain 23 percent in consumption-equivalent welfare, while the top 75 percent lose about 5.3 percent. These distributional predictions are validated against a new nationally representative survey of 3,500 Ghanaian households (conducted by the authors in August–September 2022): households with at most a JHS education were 3.1 percentage points more likely to support the policy than average, while those with SHS education or more were 5.2 percentage points less likely — remarkably close to the model&amp;rsquo;s predicted values of 2.6 and 5.9 percentage points, respectively. The authors conclude that free secondary schooling in developing countries is primarily a redistributive policy and not an efficient path to economic growth at current levels of schooling quality.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses a two-step strategy. First, it estimates the OLG model using SMM, with the experimental moments from Duflo, Dupas, and Kremer&amp;rsquo;s (2021) RCT serving as the key identifying variation. The RCT randomly assigned scholarships to poor but high-ability students in Ghana who had passed the BECE but had not enrolled in SHS, making the treatment effect on schooling completion, test scores, and fertility credibly causal in partial equilibrium. Second, the estimated model is used to compute general-equilibrium counterfactuals for a nationwide policy. The main threats to validity are: (a) external validity of the RCT sample to the general population — the sample is explicitly &amp;lsquo;smart kids from poor families,&amp;rsquo; which the authors account for through the negative-selection counterfactual; (b) the model misses on the intergenerational schooling correlation (model: 0.32 vs data: 0.45) and on the treatment effect on SHS completion (model: 21.3 pp vs data: 27 pp), though the authors show in Appendix C that forcing the model to match these moments does not reverse the negative GDP conclusion (a 40 percent higher schooling cost parameter yields a -0.8 percent GDP result vs -1.0 percent baseline; a 15 percent higher ability-persistence parameter yields -2.0 percent); (c) abstracting from human capital externalities (Lucas 1988 type spillovers) and crime reduction effects of education — the authors note these omissions but argue the low estimated effects of the policy make them unlikely to matter quantitatively; and (d) partial equilibrium of the RCT itself — the authors assume no general-equilibrium effects of the experiment since it covered only 2,064 students.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-three-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the three main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The three channels are (i) opportunity cost — attendees ages 15–19 forgo labor income; (ii) negative selection — removing the BECE requirement means new marginal students have lower average ability than current attendees; (iii) differential fertility — newly educated households reduce fertility, shifting the long-run population distribution toward less-educated (higher-fertility) households, diluting the share of educated workers over time. The paper isolates each channel through sequential counterfactual experiments: (i) is isolated by eliminating the option for ages-15–19 children to work (forcing the choice between schooling and idleness), which raises the GDP effect from -1.0 to +2.9 percent; (ii) is isolated by artificially boosting the ability of new secondary attendees to match the experimental sample&amp;rsquo;s ability distribution, which moves GDP from -1.0 to approximately 0; (iii) is isolated by setting new attendees&amp;rsquo; fertility to the uneducated-household level, which moves GDP from -1.0 to +1.2 percent. The magnitudes reveal that the opportunity cost channel is the largest (approximately 4 pp swing), followed by the fertility channel (approximately 2.2 pp), and then the selection channel (approximately 1 pp).&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Several dimensions of heterogeneity are documented. In the experimental sample, the treatment effect on SHS completion is not particularly skewed toward high-ability students: the difference in treatment effects between the top and bottom test-score quartiles is only 4 percentage points in the data (and 3 in the model), implying broadly similar gains across the ability distribution within the selected sample. In the estimated model&amp;rsquo;s misallocation analysis, the attendance probability plot (Figure 3) shows that the highest-ability children are fairly likely to attend SHS even when born to low-ability parents — suggesting relatively low misallocation in the estimated model compared to the stylized high-misallocation case. On welfare, the paper documents large heterogeneity by income quartile: the bottom 25 percent of parents gain 7.3 percent in consumption-equivalent welfare while the top 25 percent lose 4.2 percent; for children the bottom 25 percent gain 23 percent while the top 75 percent lose about 5.3 percent. Welfare also differs across generations: gains for grandchildren who always exist are smaller (9 percent) than for children (12 percent), reflecting the compounding fertility effect. The survey confirms these patterns across urban/rural, male/female, and across the Volta (42.3 percent average support for free SHS) and Ashanti (78.2 percent average support) regions of Ghana.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The authors report three robustness checks in Appendix C. First, they increase the schooling cost parameter ΨS by 40 percent to force the model to match the (currently undershot) treatment effect on SHS completion; the free schooling policy then produces a -0.8 percent GDP result (vs -1.0 percent baseline) and a 14 percent increase in attendance (vs 12 percent baseline) — the conclusion is unchanged. Second, they increase the ability-persistence parameter ρ by 15 percent to match the intergenerational schooling correlation; the result is a -2.0 percent GDP decline and a 4 percent attendance increase — the GDP decline is larger, so if anything the baseline is too generous to free schooling. Third, they experiment with lower values of the elasticity of substitution between skilled and unskilled labor (down to 1.4 from the baseline value of 4) and report no substantive change in conclusions. The authors also use bootstrapped 95 percent confidence intervals for all aggregate predictions, which is unusual in general-equilibrium counterfactual exercises in macroeconomics.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does the paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The paper is most closely related to Abbott, Gallipoli, Meghir, and Violante (2019) and Daruich (2020), both of which study public education expansions in the United States and find largely positive effects on GDP and welfare. The authors argue the contrast with their pessimistic findings reflects lower school quality in developing countries — in a rich-country setting, opportunity costs are lower relative to the returns to schooling. Hendricks and Schoellman (2014) find similar negative selection of college students in the US as enrollment expands, lending support to the selection channel. Khanna (2023) documents substantial declines in the relative wages of skilled workers after an education expansion in India, consistent with the model&amp;rsquo;s 10 percent skilled-to-unskilled wage compression, though Khanna&amp;rsquo;s short-run effects are larger due to lower short-run elasticity of substitution. In terms of methodology, the paper follows Daruich (2020) in using RCT evidence to discipline an OLG model, and is the first paper to do so for the macroeconomic effects of education policy in the developing world. The paper also builds on the macro-development literature emphasizing school quality (Hanushek and Woessmann, 2007; Schoellman, 2012) over average years of schooling as the proximate cause of low human capital in poor countries.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central policy implication is that free secondary schooling in developing countries, at current low levels of schooling quality, is primarily redistributive rather than growth-enhancing. Countries considering free schooling should expect secondary enrollment to rise substantially (by around 12 percentage points in the baseline) but GDP per capita to fall or stay flat. The alternative of improving schooling quality — modeled as a 0.1 standard deviation increase in test scores, using teacher incentives or additional teachers at a cost of approximately US$5.78 per student per year (based on Mbiti et al. 2019 in Tanzania) — raises GDP by 2.7 percent and schooling enrollment by even more (13.8 percentage points), while paying for itself in equilibrium. A key scope condition: the negative GDP finding is driven by the combination of high opportunity costs of schooling (secondary-school-age workers have economically significant labor income in developing countries), negative selection from removing merit requirements, and low schooling quality that limits the human capital return per year of schooling. In rich countries where these conditions do not hold, the same policy has been found to be beneficial. The paper also shows (Table 6) that maintaining the entrance-exam requirement alongside free schooling substantially mitigates the GDP decline (-0.3 percent vs -1.0 percent), and that keeping both the test and a positive fee results in approximately zero GDP change — suggesting that the test-requirement component of the policy design is important.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-paper-find-about-misallocation-in-the-estimated-model"&gt;Q7. What does the paper find about misallocation in the estimated model?&lt;/h3&gt;
&lt;p&gt;The estimated model exhibits relatively low misallocation. The misallocation concept refers to situations where high-ability children of poor parents are kept out of secondary school by borrowing constraints even though the net-present-value of additional schooling exceeds the cost. The paper shows (Figure 2) that economies can have similar aggregate secondary enrollment rates of around 30 percent but very different degrees of misallocation — one where enrollment is low because returns are low (low-misallocation case), and one where enrollment is low because high-ability children are credit-constrained (high-misallocation case). The estimated model falls closer to the low-misallocation case (Figure 3), with the highest-ability children fairly likely to attend SHS even if born to low-ability parents. This finding is consistent with the modest increase in SHS completion induced by free schooling (12 percentage points) relative to the experimental treatment effect on the selected sample (27 percentage points): most high-ability children are already attending, so there is limited room for a free schooling policy to reduce misallocation.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-welfare-analysis-reveal-about-the-puzzle-of-large-welfare-gains-alongside-a-gdp-decline"&gt;Q8. What does the welfare analysis reveal about the puzzle of large welfare gains alongside a GDP decline?&lt;/h3&gt;
&lt;p&gt;The paper documents an apparent puzzle: the free schooling policy reduces long-run GDP per capita by 1 percent but produces large positive welfare gains for parents (average 3.9 percent in consumption-equivalent welfare) and even larger gains for children (average 12.4 percent). The resolution is that (a) welfare gains for parents come entirely from redistribution — the very poor gain 7.3 percent while the rich lose 4.2 percent, and the progressive tax schedule is the mechanism; (b) the welfare gains for the children&amp;rsquo;s generation partially reflect large gains to the small number of previously misallocated children who now attend secondary school (the bottom 25 percent of children gain 23 percent, primarily through income gains for those who previously could not afford school); and (c) these gains erode across generations — grandchildren who always exist gain less (9 percent vs 12 percent for children), because the grandchildren who would only have existed without the free schooling policy (i.e., the &amp;lsquo;unborn&amp;rsquo; due to reduced fertility among educated households) would have experienced disproportionately large gains (almost 17 percent). The composition of the population thus shifts toward those experiencing smaller gains, compounding over generations and producing the long-run GDP decline.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-role-of-the-entrance-exam-design-in-free-schooling-policy-outcomes"&gt;Q9. What is the role of the entrance exam design in free schooling policy outcomes?&lt;/h3&gt;
&lt;p&gt;The paper shows that how access is structured matters as much as whether schooling is free. In the main analysis, free schooling eliminates both fees and the BECE entrance requirement, consistent with Ghana&amp;rsquo;s 2017 policy. In alternative simulations (Table 6), free schooling that maintains the existing entrance requirement (a &amp;lsquo;relaxed test&amp;rsquo; policy) produces a GDP decline of only -0.3 percent instead of -1.0 percent. Free schooling that keeps the test at full stringency (so fewer new students gain access) produces essentially no change in GDP (-0.0 percent), but also a much smaller increase in secondary attendance (3.0 pp vs 11.8 pp). Eliminating only the test requirement while keeping a positive fee produces a -0.4 percent GDP decline. These results confirm that the negative selection channel is a quantitatively important driver of the adverse GDP effect and is specifically activated by the removal of the merit requirement.&lt;/p&gt;
&lt;h3 id="q10-how-is-the-model-estimated-and-what-moments-does-each-parameter-primarily-identify"&gt;Q10. How is the model estimated and what moments does each parameter primarily identify?&lt;/h3&gt;
&lt;p&gt;The model is estimated by SMM minimizing the sum of squared differences between model moments and their data counterparts, using a vector of 10 parameters (fertility parameters νJ and νS; schooling efficiency ηS; goods cost of schooling ΨS; intergenerational altruism b; exam score noise σε; Gumbel taste-shock scale θ; savings wedge χ; ability persistence ρ; ability shock standard deviation συ). Six parameters are chosen directly from the literature or normalization (A, α, β, r*, λ, σζ). Ten moments are targeted: population growth rate (primarily identifies νJ, νS), aggregate SHS completion rate and quartile completion rates (identify ηS, b, ΨS, χ), variance of the permanent component of wages (identifies συ, ρ), and five experimental moments from the Duflo et al. RCT (treatment effects on human capital, fertility, SHS completion, the Q4–Q1 completion difference, and the intergenerational schooling correlation). Confidence intervals are bootstrapped by re-sampling the five experimental moments 100 times, treating the non-experimental moments as fixed. The Jacobian matrix (Appendix Table C.1) and sensitivity matrix (Appendix Table C.2) are computed following Kaboski and Townsend (2011) and Andrews, Gentzkow, and Shapiro (2017) to document identification.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-survey-design-details-and-how-well-does-it-validate-the-model"&gt;Q11. What are the survey design details and how well does it validate the model?&lt;/h3&gt;
&lt;p&gt;The authors conducted a new nationally representative household survey in Ghana in August–September 2022, covering 3,500 households selected via two-stage cluster sampling from seven regions accounting for about 61 percent of the Ghanaian population. Respondents were asked whether eight categories of government expenditure should be abolished, cut substantially, cut somewhat, maintained, or expanded. For free SHS, respondents with at most a JHS education were 3.1 percentage points more likely to support the policy than average; those with SHS education or more were 5.2 percentage points less likely. These empirical patterns align closely with the model&amp;rsquo;s predicted values of 2.6 and 5.9 percentage points respectively. The pattern is robust across urban/rural subsamples, male/female subsamples, and across the Volta and Ashanti regions (which differ substantially in overall support levels — 42.3 percent vs 78.2 percent — but maintain the same qualitative pattern of lower-educated households being more supportive). The one discrepancy is that the model over-predicts the support of JHS-educated households who have children enrolled in SHS.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Opportunity cost of schooling&lt;/strong&gt;: In this paper&amp;rsquo;s model, the foregone labor income of teenagers aged 15–19 who attend secondary school rather than work. This cost persists even when the school fee is eliminated by government policy and is identified as the single largest channel explaining why free secondary schooling reduces rather than raises GDP per capita in developing countries, contributing approximately 4 percentage points to the adverse GDP effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negative selection of new students&lt;/strong&gt;: The reduction in average ability of the marginal students who enter secondary school once both fees and the merit-based entrance exam are eliminated. The existing pool of secondary attendees was positively selected by the entrance exam, so broadening access induces a lower-ability pool of new entrants, reducing the average human capital gain per new graduate. The paper estimates this channel accounts for approximately 1 percentage point of the adverse GDP gap relative to the back-of-the-envelope projection.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Differential fertility by education&lt;/strong&gt;: The model feature by which secondary-educated households have significantly fewer children (parameter νS = 0.19 implying 2.4 children per family) than non-secondary-educated households (νJ = 1.07 implying 4.1 children per family). When free schooling induces more households to obtain secondary education, aggregate fertility falls, and crucially the share of high-ability households in the long-run population declines because those households now have fewer children, reducing the long-run supply of educated workers and contributing approximately 2.2 percentage points to the adverse GDP gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Misallocation of talent&lt;/strong&gt;: In this paper&amp;rsquo;s sense: the situation in which high-ability children of poor parents are prevented by borrowing constraints from attending secondary school even though the net-present-value of additional schooling exceeds the combined goods and opportunity costs. The paper finds that the estimated model of Ghana corresponds more closely to a low-misallocation economy (Figure 3), meaning the highest-ability children attend SHS at fairly high rates regardless of parental income, so the scope for free schooling to reduce misallocation is limited.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Balanced growth path&lt;/strong&gt;: In this paper: a recursive competitive equilibrium in which aggregate population grows at a constant rate while the relative distribution of households across individual states (ability, education, assets) is stationary, and household policy functions are independent of the aggregate population level. All policy counterfactuals are conducted by introducing a policy into the balanced growth path and computing transition dynamics to the new balanced growth path.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Schooling quality (ηS)&lt;/strong&gt;: The efficiency parameter governing how much human capital a student of given ability acquires from a year of secondary schooling, defined in the production function h(z,S) = z · ηS. In the estimated model, ηS = 5.66, implying an annual return to education of 7.9 percent for the experimental sample. The paper shows that a policy raising ηS (schooling quality) by enough to increase average test scores by 0.1 standard deviations raises GDP by 2.7 percent and expands SHS enrollment by 13.8 percentage points, outperforming free schooling on both counts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Savings wedge (χ)&lt;/strong&gt;: A wedge between the international market rate of return on capital (r*) and the return available to households in the model (r = r* - χ), calibrated to match the low savings rates observed in low-income economies. In the estimated model χ = 0.09, implying households earn approximately 2 percent per year on savings. Together with the borrowing constraint (no borrowing against children&amp;rsquo;s future income), this ensures that poor parents cannot save their way out of the constraint preventing them from sending high-ability children to school.&lt;/p&gt;</description></item><item><title>Manipulation of information in times of crisis: evidence from Covid excess mortality</title><link>https://macropaperwarehouse.com/papers/manipulation-of-information-in-times-of-crisis-evidence-from-covid-excess-mortality/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/manipulation-of-information-in-times-of-crisis-evidence-from-covid-excess-mortality/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Karlinsky and Shayo ask which governments manipulate public information, in which direction, and by how much — questions that are normally intractable because the ground truth is unobservable. The Covid-19 pandemic supplies an unusual opportunity: all countries faced a broadly similar crisis simultaneously, and all-cause mortality — collected by national statistical offices as a routine bureaucratic function independently of Covid — provides a manipulation-resistant benchmark against which officially-reported Covid deaths can be evaluated.&lt;/p&gt;
&lt;p&gt;The authors hand-collect all-cause mortality data for 134 countries and territories from national statistical offices, population registries, health ministries, and, in some cases, right-to-information requests facilitated by local journalists. Data span 2015–2021 at weekly, monthly, or annual frequency. Their sample covers 93 percent of countries with at least 75 percent Death Registration Completeness. They compute, for each country, a Misreporting Rate (MRR) defined as estimated Covid deaths minus officially reported Covid deaths, normalised by expected total deaths derived from pre-pandemic trends. Estimated Covid deaths equal excess mortality — itself estimated from a country-specific model with weekly/monthly fixed effects and an annual trend (R² = 0.997 in pre-pandemic prediction) — minus adjustments for excess deaths attributable to conflicts, natural disasters, and other identifiable non-Covid causes. Those adjustments are small: the mean total adjustment across the sample is 0.04 percent of expected deaths.&lt;/p&gt;
&lt;p&gt;Six main findings emerge. First, between 45 and 55 percent of the 134 countries misreported Covid deaths. Second, the direction of manipulation is overwhelmingly one-sided: of 131 countries with sufficient data to estimate confidence intervals, 59 reported accurately, 62 significantly underreported, and only 10 overreported. The theoretical prediction that governments might exaggerate a crisis — to rally populations, legitimise repressive measures, or attract foreign aid — finds no empirical support. Third, the magnitude of underreporting is large: the sample reported 5.08 million Covid deaths in 2020–2021 while estimated actual Covid deaths were 12.47 million, nearly 2.5 times the official figure; the implied global MRR is 12.8 percent. Among the 62 underreporting countries, the average MRR is 14.5 percent of expected total deaths and the median is 12 percent. Individual-country MRRs range from above 37 percent (Bolivia, Nicaragua) downward, with Russia at 24 percent. Fourth, state capacity in counting and registering deaths explains some but far from most cross-country variation; the R² of the best capacity-only regression is 0.115. Chile and Russia have virtually identical Death Registration Completeness and Percent Well-Certified Death Registrations, yet Chile accurately reported while Russia&amp;rsquo;s MRR is 24 percent. Fifth, the extent of underreporting is strongly associated with constraints on governmental power. In individual regressions conditioning on capacity, each of three institutional constraint measures — Clean Elections, Executive Constraints, and Freedom of the Press — is associated with a 0.4–0.5 standard deviation lower MRR per one standard deviation stronger constraint. In a joint model including all 12 factors from four domains (macroeconomic incentives, culture, audience sophistication, institutions), institutional constraints are the strongest predictor (partial R² ≈ 0.11), followed by audience sophistication (partial R² ≈ 0.04–0.06). Macroeconomic incentives — tourism reliance, unemployment, foreign direct investment — are not jointly significant. Cultural factors (trust, individualism, religiosity) lose significance once other factors are controlled. The full model explains more than 50 percent of MRR variation. Sixth, countries with a communist legacy (defined as having had a communist or socialist regime for at least 10 years, covering 34 countries) show significantly higher misreporting even holding current institutional and cultural conditions constant. Countries that held elections during 2020–2021 also show significantly higher misreporting.&lt;/p&gt;
&lt;p&gt;The results are robust to alternative expected-mortality models, alternative MRR normalisations, the inclusion of Bangladesh, China, and Indonesia (treated separately due to data quality concerns), year-by-year (2020 vs. 2021) splits, controls for age structure and GDP per capita, and alternative manipulation measures (underdispersion, Benford&amp;rsquo;s law deviations). The evidence that manipulation cannot be attributed to varying standards for false-positive attribution of cause of death is direct: four pre-pandemic measures of a country&amp;rsquo;s tendency to use unspecified cause-of-death categories are uncorrelated with MRR and individually account for less than 1 percent of its variation.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s contribution to the economics of information manipulation is methodological as well as empirical: it provides a comparable, country-level measure of governmental misinformation based on actual observable actions regarding a policy issue of central importance, covering a large and diverse cross-section of countries.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The strategy compares officially-reported Covid deaths (the variable that attracted political attention and over which governments had strong incentives and ability to intervene) with estimated Covid deaths derived from excess all-cause mortality (a statistic collected routinely by national bureaucracies under very different incentive structures, harder to manipulate, and less visible publicly during the pandemic). The identifying assumption is that all-cause mortality data are not themselves systematically manipulated in response to Covid. The authors defend this on four grounds: (1) all-cause mortality has long been collected independently of Covid; (2) ascertaining that someone died is far easier than attributing a cause of death; (3) Covid figures attracted vastly more public attention, making their manipulation more urgent; (4) when governments appear to have discovered the evidential value of excess mortality, their response has been to delay publication of all-cause data rather than to alter it (Belarus is cited as an example). The main remaining threat is that the adjustment for non-Covid excess deaths (conflicts, disasters, traffic accidents, suicides, homicides) is imperfect in countries with poor data on those causes. The authors note this caveat but show mean adjustments are tiny (0.04% of expected deaths) and the largest individual adjustments (Armenia 6.1%, Azerbaijan 3.2%) are driven by the Nagorno-Karabakh war and are handled explicitly.&lt;/p&gt;
&lt;h3 id="q2-how-is-excess-mortality-estimated-and-how-sensitive-are-the-results-to-modelling-choices"&gt;Q2. How is excess mortality estimated, and how sensitive are the results to modelling choices?&lt;/h3&gt;
&lt;p&gt;Country-specific models are estimated using 2015–2019 all-cause mortality data, including country-specific weekly or monthly fixed effects and a country-specific annual trend to capture seasonality and long-run factors (population ageing, improvements in health care, etc.). The model achieves R² = 0.997 in predicting pre-pandemic mortality. The authors report in Supplementary Material B that alternative expected-mortality approaches from the literature yield very similar results, as do alternative normalisations of the MRR. Sensitivity to model choice is low because the discrepancies between excess and reported deaths in weak-institution countries are so large that they persist across methodological variants.&lt;/p&gt;
&lt;h3 id="q3-how-do-the-authors-distinguish-intentional-manipulation-from-limited-state-capacity"&gt;Q3. How do the authors distinguish intentional manipulation from limited state capacity?&lt;/h3&gt;
&lt;p&gt;They use two pre-pandemic, capacity-specific measures: (1) Death Registration Completeness (DRC) — the share of deaths captured by the vital registration system — and (2) Percent of Well-Certified Death Registrations (PWC) — the share with proper cause-of-death attribution. Both are computed before the pandemic so they are not contaminated by Covid-era behaviour. Regressions confirm that capacity predicts MRR negatively (R² up to 0.115), but the residual variation remains large. The clearest illustration is Chile vs. Russia: both have complete DRC and near-identical high PWC, yet Chile reports accurately and Russia has an MRR of 24 percent. All subsequent analysis of correlates conditions on these capacity measures.&lt;/p&gt;
&lt;h3 id="q4-how-do-the-authors-rule-out-the-possibility-that-differences-in-false-positive-aversion-rather-than-manipulation-explain-mrr-variation"&gt;Q4. How do the authors rule out the possibility that differences in false-positive aversion (rather than manipulation) explain MRR variation?&lt;/h3&gt;
&lt;p&gt;They construct four pre-pandemic measures from WHO Mortality Database ICD-10 cause-of-death data: (1) number of ICD codes reported; (2) share of specific-viral deaths among all viral deaths; (3) share of specific-infection deaths among all infection deaths; (4) share of specific-respiratory deaths among all respiratory deaths. A country more averse to false positives would report less specific causes. None of the four measures is significantly associated with MRR, and none accounts for more than 1 percent of its variation. This rules out differences in diagnostic/reporting standards as a driver of the observed discrepancies.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-direction-of-manipulation-and-what-does-this-imply-for-theories-of-governmental-information-behaviour"&gt;Q5. What is the direction of manipulation and what does this imply for theories of governmental information behaviour?&lt;/h3&gt;
&lt;p&gt;Of 131 countries with estimable confidence intervals, 62 significantly underreported and only 10 overreported. The four main theoretical channels for overreporting — rally-around-the-flag effects, legitimising repression, attracting foreign aid, and inducing flight-to-safety compliance — find no empirical support. The authors argue that the rally-around-the-flag mechanism requires an outgroup-related threat (Covid, unlike a foreign military attack, was not easily framed this way), that Covid mortality does not signal repressive capacity, and that international economic actors appear sufficiently sophisticated to be sceptical of inflated figures. The pattern is consistent instead with governments downplaying to project competence, reduce accountability, and justify inadequate responses.&lt;/p&gt;
&lt;h3 id="q6-what-factors-are-most-strongly-associated-with-misreporting-and-how-are-they-ranked"&gt;Q6. What factors are most strongly associated with misreporting, and how are they ranked?&lt;/h3&gt;
&lt;p&gt;In joint regressions with all 12 factors from four domains, after conditioning on capacity: (1) Institutional constraints (Clean Elections, Executive Constraints, Freedom of the Press) have the highest partial R² (approximately 0.11 for Executive Constraints alone) and are jointly significant at p &amp;lt; 0.001; each standard deviation of stronger institutional constraint is associated with roughly 0.4–0.5 standard deviations lower MRR. (2) Audience Sophistication (tertiary education, HDI Education Index, internet access) is the second strongest domain (partial R² in the range of 0.04–0.06 per variable; jointly significant at p &amp;lt; 0.05). (3) Cultural factors (trust, individualism, religiosity) are individually significant in bivariate regressions but lose significance when institutional and other factors are controlled. (4) Macroeconomic incentives (tourism, unemployment, net FDI) are not jointly significant in any specification. Specification-curve analysis across all combinations of controls confirms that Executive Constraints is the single most robust predictor, retaining sign, magnitude, and significance across all models. The full model (Table 4, column 1) has R² exceeding 0.50.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-communist-legacy-finding-and-how-is-it-interpreted"&gt;Q7. What is the communist legacy finding and how is it interpreted?&lt;/h3&gt;
&lt;p&gt;Countries defined as having had a communist or socialist regime for at least 10 years (34 countries) show significantly higher MRRs even after conditioning on contemporary institutional constraints, audience sophistication, culture, and capacity. The coefficient is statistically significant at p &amp;lt; 0.05 or better in the main and most robustness specifications. The authors point to Harrison (2017) on the pervasiveness of information manipulation in communist states as a historical precedent, and interpret the finding as a persistent legacy operating through channels not fully captured by current measures. This suggests that historical exposure to a political culture of systematic information manipulation may have durable effects on bureaucratic behaviour or political norms that current V-Dem indices do not fully absorb.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-elections-finding"&gt;Q8. What is the elections finding?&lt;/h3&gt;
&lt;p&gt;Countries holding national parliamentary or presidential elections during 2020–2021 (76 of 134 countries) show significantly higher misreporting, consistent with electoral incentive theories of information manipulation. This finding is robust to including controls for GDP per capita, population age structure, and other domains, and is stable across the 2020-only and 2021-only sub-samples.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-performed"&gt;Q9. What robustness checks are performed?&lt;/h3&gt;
&lt;p&gt;The authors conduct: (1) specification-curve analysis across all combinations of covariates; (2) a joint model with all 12 individual factors; (3) principal component analysis within each domain to recover common variation and reduce dependence on specific measurement choices; (4) alternative expected-mortality models (Supplementary Material B.1); (5) alternative MRR normalisations (Supplementary Material B.2); (6) separate year-by-year analysis for 2020 and 2021; (7) inclusion of Bangladesh, China, and Indonesia as robustness cases despite lower data reliability; (8) addition of GDP per capita to check whether the institution-misreporting link is proxying for development; (9) analysis using underdispersion (Kobak 2022) and Benford&amp;rsquo;s law deviations as alternative manipulation measures; (10) exploration of colonial legacy as an additional historical variable (no significant effect found). The primacy of institutional constraints is robust across all of these.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-authors-treat-china-bangladesh-and-indonesia"&gt;Q10. How do the authors treat China, Bangladesh, and Indonesia?&lt;/h3&gt;
&lt;p&gt;These three large countries are excluded from the main analysis because their all-cause mortality data come from surveys (Bangladesh, China) rather than vital registration systems, or are very incomplete (Indonesia), making excess mortality estimation unreliable. They are included in a robustness regression (Table 4, column 6) and results are described as qualitatively similar. The authors flag that China&amp;rsquo;s data may itself be informative as a potential indicator of data suppression.&lt;/p&gt;
&lt;h3 id="q11-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q11. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;The paper is closest in spirit to Olken (2007), who uses the gap between reported and actual infrastructure spending to measure corruption, and Martinez (2022), who compares GDP growth to night-time-light-implied growth and finds autocracies overstate growth by more than a third. The authors extend this approach to a different domain (health/mortality) with broader country coverage. Prior Covid-specific work documented anomalies — underdispersion (Kobak 2022) and Benford&amp;rsquo;s law deviations (Kapoor et al. 2020; Kilani 2021) — and noted that autocratic regimes reported lower-than-expected deaths (Annaka 2021; Cassan and Van Steenvoort 2021), but these studies relied on regime type as the sole or primary explanatory variable and did not systematically rank competing factors. Neumayer and Plümper (2022) and Wigley (2024) used the authors&amp;rsquo; own World Mortality Dataset to test data manipulation. This paper is distinctive in that it: (a) provides what the authors describe as the most systematic estimates to date of Covid mortality and misreporting; (b) examines a broad range of factors across four domains without a priori privileging any; (c) directly tests and rejects capacity and false-positive aversion as alternative explanations; and (d) identifies communist legacy and elections as additional significant correlates.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Three implications are highlighted. First, unconstrained regimes appear to manipulate not only economic statistics but also health information during the most salient public policy event of the era; travel restrictions and multilateral actions during the pandemic relied on reported Covid figures, so manipulation had direct international externalities. This raises broader questions about the credibility of official data from such governments across domains — foreign aid targeting, climate action, vaccination campaigns. Second, the MRR provides a comparable cross-country measure of institutional quality grounded in actual governmental behaviour, potentially useful as an input to studies of institutions, conflict, electoral outcomes, and economic performance. Third, some countries that score respectably on conventional executive constraint indices — Albania, El Salvador, India, Serbia — show high MRRs, suggesting these rates may be leading indicators of democratic erosion not yet captured by standard measures. The scope condition the authors flag is external validity: if pandemic mortality is an extreme case with unique incentive structures (tourism, investment, aid eligibility), then findings about determinants of manipulation may not generalise beyond crisis settings. The authors argue against this interpretation on the grounds that macroeconomic factors — which would be pandemic-specific — are not significant, while institutional constraints — which reflect general governmental behaviour — are.&lt;/p&gt;
&lt;h3 id="q13-what-limitations-do-the-authors-acknowledge"&gt;Q13. What limitations do the authors acknowledge?&lt;/h3&gt;
&lt;p&gt;First, the analysis is explicitly descriptive rather than causal; factors are correlates, not proven determinants. Second, the MRR may understate true manipulation if all-cause mortality data are themselves selectively withheld or manipulated; the authors argue this is probably modest but acknowledge it cannot be fully ruled out. Third, important large countries — Pakistan, Nigeria, Ethiopia, Venezuela — cannot be scored because sufficient all-cause mortality data are not publicly available; the authors note this absence may itself be informative but cannot be quantified. Fourth, data on other causes of excess deaths (traffic accidents, suicides, homicides) are patchy in many countries, though the scale of these adjustments is very small. Fifth, some capacity controls (PWC) use data from as early as 2003, introducing measurement error. The paper does not claim to fully separate the channels through which institutions reduce manipulation (electoral accountability, press scrutiny, judicial oversight, professional agency independence), treating them as joint constraints rather than separately identified mechanisms.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Misreporting Rate (MRR)&lt;/strong&gt;: The paper&amp;rsquo;s central measure, defined as (estimated Covid deaths minus officially reported Covid deaths) divided by expected total deaths for the country in the same period based on pre-pandemic trends. A positive MRR indicates underreporting; a negative MRR indicates overreporting. Normalising by expected total deaths rather than by reported Covid deaths accounts for differences in population size, age structure, and baseline mortality across countries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Excess mortality&lt;/strong&gt;: The number of deaths above and beyond what would have been expected in the absence of the pandemic, estimated from country-specific models with weekly or monthly fixed effects and an annual trend fitted to 2015–2019 data. Used as the primary building block for estimated Covid deaths after subtracting excess deaths due to identified non-Covid causes (conflict, natural disasters, traffic accidents, homicides, suicides).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Death Registration Completeness (DRC)&lt;/strong&gt;: In this paper&amp;rsquo;s usage, the share of all deaths in a country captured by its vital registration system each year, measured using pre-pandemic data. Treated as the most basic indicator of a country&amp;rsquo;s capacity to count deaths. Used as a control to separate capacity constraints from intentional manipulation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Percent of Well-Certified Death Registrations (PWC)&lt;/strong&gt;: The share of death certificates in a country that carry a properly specified cause of death, measured using pre-pandemic data. Used alongside DRC as a second capacity control capturing not just whether deaths are registered but whether causes are correctly attributed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Informational Autocrat&lt;/strong&gt;: Following Guriev and Treisman (2022), the paper uses this concept to describe executives in countries where formal and informal checks and balances are weak, who systematically manipulate public information to project competence and reduce accountability. The paper&amp;rsquo;s empirical results are interpreted as evidence that such executives behave as informational autocrats not only in economic statistics but also in health data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;False-positive aversion&lt;/strong&gt;: The tendency of some countries to apply a higher evidentiary bar before attributing a death to a specific cause — such as Covid — rather than leaving the cause unspecified, independently of capacity or intention to deceive. The paper operationalises this using pre-pandemic ICD-10 data on specificity of reported causes of death and shows it is uncorrelated with MRR, ruling it out as a driver of observed discrepancies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Communist legacy&lt;/strong&gt;: The paper&amp;rsquo;s binary indicator for countries that had a communist or socialist regime for at least 10 consecutive years (34 countries). The variable captures historical exposure to a political culture of systematic information manipulation and is found to be a significant positive predictor of MRR even after conditioning on current institutional constraints, consistent with persistent norms or bureaucratic practices.&lt;/p&gt;</description></item><item><title>Means-Tested Transfers in the US: Facts and Parametric Estimates</title><link>https://macropaperwarehouse.com/papers/means-tested-transfers-in-the-us-facts-and-parametric-estimates/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/means-tested-transfers-in-the-us-facts-and-parametric-estimates/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Guner, Rauh, and Ventura document the scope, generosity, distributional impact, and time evolution of means-tested transfers to working-age US households, and provide parametric estimates of transfer functions for use in applied macroeconomics and public finance. The paper addresses three questions: How large are these transfers? How do they affect income inequality? How have they changed over time? The contribution is descriptive and empirical rather than structural; the paper does not estimate behavioral effects but rather characterizes the effective transfer schedule that households face.&lt;/p&gt;
&lt;p&gt;The data source is the Survey of Income and Program Participation (SIPP), using five waves spanning 1998 to 2016. The benchmark analysis uses the 2014 wave (years 2013–2016). The sample is restricted to household-years in which the head is aged 25–54, is not self-employed, and does not switch marital status within the year — yielding 18,612 households and 38,375 household-year observations. Six programs are covered: TANF, SNAP, WIC, SSI, housing assistance, and Medicaid. For TANF, SNAP, WIC, and SSI, transfer values are observed directly. Medicaid values are imputed using regional HMO premium costs; housing values are imputed as the difference between Fair Market Rent and actual rent paid.&lt;/p&gt;
&lt;p&gt;In the 2013–2016 benchmark period, approximately 35% of working-age households receive some means-tested transfer in a given year, and, conditional on receipt, the average household receives about $17,000 (in 2016 dollars), exceeding one-fourth of average household income. Unconditional total transfers decline steeply with income but in a non-monotone way: households with zero non-transfer income receive $7,500 in non-medical and $13,700 in Medicaid transfers ($21,000 total, or 26% of mean household income). Transfers dip for households with small positive incomes (creating a hump shape), then rise slightly before declining again. At the bottom income decile (0–10%), households receive on average $4,125 in non-medical transfers and $14,141 total. At the median income decile (50–60%), households receive $425 non-medical and $3,006 total. In the top decile, non-medical transfers are negligible ($169) and total transfers are $1,200. The decline in unconditional transfers with income is driven primarily by reduced coverage: conditional on receipt, transfer amounts are relatively stable across income levels, remaining above 15% of mean household income throughout the distribution. The extensive margin of coverage is 82% for zero-income households, 70% for the bottom decile, 29% at the median, and still 5% (non-medical) to 11% (including Medicaid) in the top decile.&lt;/p&gt;
&lt;p&gt;Medicaid is the dominant program throughout. For zero-income households, Medicaid transfers are more than six times larger than the next-largest program (SNAP). Medicaid&amp;rsquo;s share of total transfers rises with income. As a single program, Medicaid reaches 31% of working-age households with an average conditional benefit of about $15,000 per recipient. SNAP covers 18% of households with conditional benefits of about $3,000.&lt;/p&gt;
&lt;p&gt;Transfers substantially compress inequality. The pre-transfer Gini coefficient is 0.48 and falls to 0.42 when all transfers (including Medicaid) are included, and to 0.46 with non-medical transfers only. The pre-transfer 50-10 income ratio of 10.2 drops to 3.0 with all transfers and to 5.6 with non-medical transfers only. The variance of log income falls by nearly 36% (47 log points) with all transfers and by 21% with non-medical transfers. These equalizing effects are concentrated at the bottom of the distribution; for households at 10% of average pre-transfer income, total transfers more than double disposable income.&lt;/p&gt;
&lt;p&gt;Between 1998–1999 and 2013–2016, total unconditional transfers per household quadrupled from approximately 2% to 7.3% of mean household income (from about $1,535 to $6,000). Household coverage rose from 19% to 35%. The expansion is driven almost entirely by Medicaid; non-medical transfers rose only marginally in magnitude (from about 1.3% to 1.8% of mean income), though their coverage increased from 16% to 24% of households. Notably, over this period the concentration of non-medical transfers shifted upward in the income distribution: households with zero income received a smaller relative share in 2013–2016 than in 1998–1999, while shares for households in the second, third, and fourth deciles increased. Pre-transfer income inequality rose substantially over the period, with the Gini increasing from 0.40 to 0.48; the post-transfer Gini rose more moderately, from 0.38 to 0.42, indicating that transfer growth largely offset rising market-income inequality at the bottom.&lt;/p&gt;
&lt;p&gt;For the parametric section, the paper estimates a flexible four-parameter Ricker-style function T(I) = exp(alpha) * exp(beta_0 * I) * I^beta_1 for positive income I (normalized by mean income), with a separate level parameter gamma at I = 0. This captures the hump-shaped pattern at low incomes and the rapid decline thereafter. Implicit benefit reduction rates derived from these estimates are large: earning one additional dollar when starting from zero income reduces total transfers by more than $11,000, as crossing from zero into positive income sharply reduces program eligibility. A more realistic $10,000 income increase reduces total transfers by more than $5,000 — an implicit marginal tax penalty exceeding 50%. Non-medical transfer penalties are somewhat smaller: the first dollar earned reduces non-medical transfers by more than $4,500, and a $10,000 income increase reduces them by about $3,300.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper is descriptive, not causal — there is no causal identification strategy in the traditional sense. The authors document reduced-form facts about transfer receipt by income level and demographic group using SIPP microdata. The main methodological choices and data limitations are: (1) Medicaid and housing assistance values are imputed rather than directly observed — Medicaid is valued at regional HMO premiums, which may not accurately reflect the value recipients place on coverage; housing benefits are valued at the difference between state Fair Market Rent and actual rent paid, which can produce negative values (2.7% of cases, set to zero). (2) SIPP is known to under-report income at the top of the distribution relative to the CPS; the paper documents that income shares of the top quintile differ by about five percentage points between SIPP and CPS, largely due to SIPP&amp;rsquo;s poor measurement of asset income. This means the effective transfer schedule at the top of the income distribution may be somewhat distorted. (3) The SIPP was overhauled after 2016, precluding analysis of more recent waves and meaning the trends analysis ends in 2013–2016. (4) Self-employed households are excluded (~7% of households) as their income measurement is noisier.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-handle-the-non-linear-hump-shaped-pattern-in-transfers-at-low-income-levels"&gt;Q2. How does the paper handle the non-linear hump-shaped pattern in transfers at low income levels?&lt;/h3&gt;
&lt;p&gt;The paper documents a hump-shaped pattern: transfers are positive at zero income, fall sharply at very low positive income (around the bottom 1% of the distribution), then increase modestly before declining monotonically. This arises because crossing from zero income to any positive income can reduce eligibility for several programs simultaneously. The parametric functional form — the Ricker function from fisheries biology — is specifically chosen to capture this pattern: for I &amp;gt; 0, T(I) = exp(alpha) * exp(beta_0 * I) * I^beta_1, where the beta_0 term governs the initial decline/rise and beta_1 allows further curvature. The zero-income level gamma is estimated separately as a discontinuity. The tight confidence intervals around observed income-percentile averages confirm that the fitted function closely tracks the data.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-by-demographic-group-is-documented"&gt;Q3. What heterogeneity by demographic group is documented?&lt;/h3&gt;
&lt;p&gt;The paper documents heterogeneity along three dimensions — marital status, number of children, and age of children — in each case reporting both unconditional and conditional transfer amounts and coverage by income decile. Key findings: (a) Marital status: Single-woman households with zero income receive 12% of mean household income in non-medical transfers and about 31% in total transfers. Married households with zero income receive 27% total, and single men receive 17.9% total. At higher income levels, married households can receive more in total transfers than single women, because Medicaid coverage is broader for families. Single-woman households show the highest coverage at very low incomes (88% receive some transfer), but married households lead in coverage at middle income levels. Single men show surprisingly high coverage even at relatively high incomes. (b) Number of children: Transfers increase substantially with children. A first-decile married household without children receives about 1.7% of average income in non-medical transfers and 9% total; with two or more children, non-medical transfers rise nearly five-fold for single-woman households in the same decile. (c) Age of children: Transfers decline as children age, but the magnitude of the age gradient is smaller than the number-of-children gradient.&lt;/p&gt;
&lt;h3 id="q4-how-do-conditional-and-unconditional-transfers-compare-across-the-income-distribution"&gt;Q4. How do conditional and unconditional transfers compare across the income distribution?&lt;/h3&gt;
&lt;p&gt;Unconditional transfers (averaged over all households including non-recipients) decline steeply with income, driven primarily by falling coverage rates. Conditional transfers (among recipients only) are much more stable. For zero-income households, total conditional transfers average $26,500 (32% of mean income) versus $21,000 unconditionally. In the bottom decile, conditional total transfers are about $21,000 or 26% of mean income. After the third income decile, conditional transfer levels stabilize and remain above 15% of mean income throughout most of the distribution. This means that once a household is enrolled in the transfer system, the amounts received are relatively constant regardless of where in the distribution they fall; the intensive margin differences are largely accounted for by Medicaid, which has high conditional values even at middle income levels.&lt;/p&gt;
&lt;h3 id="q5-what-role-does-medicaid-play-relative-to-non-medical-programs"&gt;Q5. What role does Medicaid play relative to non-medical programs?&lt;/h3&gt;
&lt;p&gt;Medicaid dominates the transfer system for working-age households by every measure. It reaches 31% of households in the benchmark period (the next largest program, SNAP, covers 18%). For zero-income households, Medicaid transfers are more than six times larger than SNAP (the next largest non-medical program). Medicaid&amp;rsquo;s share of total transfers grows with income: for zero-income households, total transfers are less than three times non-medical transfers; for households in the 50–60th percentile, this ratio exceeds six. In terms of aggregate spending, Medicaid rose from below 1% of GDP in 1980 to more than 3% in 2022, while non-medical transfers declined from 1.6% to about 1% of GDP over the same period. Almost the entire growth in household transfers between 1998 and 2016 is attributable to Medicaid expansion. Medicaid is also the most important single contributor to measured inequality reduction.&lt;/p&gt;
&lt;h3 id="q6-how-do-transfers-affect-income-inequality-and-how-has-this-changed-over-time"&gt;Q6. How do transfers affect income inequality and how has this changed over time?&lt;/h3&gt;
&lt;p&gt;In the 2013–2016 benchmark, total transfers reduce the Gini coefficient by 6 points (from 0.48 to 0.42) and the variance of log income by nearly 36%. The 50-10 income ratio falls from 10.2 to 3.0. Non-medical transfers alone reduce the Gini by 2 points (to 0.46) and the 50-10 ratio to 5.6. The impact is concentrated at the bottom of the distribution: transfers more than double total income of households with pre-transfer income around 10% of the mean. Over time, pre-transfer inequality rose sharply, with the Gini going from 0.40 (1998–1999) to 0.48 (2013–2016) and the 50-10 ratio doubling from 4.19 to 10.2. Post-transfer inequality rose more mildly: the Gini increased from 0.38 to 0.42 (all transfers), and the 50-10 ratio remained stable at around 3 throughout. Excluding Medicaid, the moderating effect is weaker; the Gini rose from 0.39 to 0.46 on a post-non-medical-transfer basis.&lt;/p&gt;
&lt;h3 id="q7-how-has-the-concentration-of-transfers-across-income-groups-evolved-over-time"&gt;Q7. How has the concentration of transfers across income groups evolved over time?&lt;/h3&gt;
&lt;p&gt;A notable distributional shift occurred between 1998–1999 and 2013–2016. For non-medical transfers, the share accruing to households with zero income declined substantially — from receiving about $9 per $100 of total transfers distributed in 1998–1999 to about $4 in 2013–2016. Similarly, the relative share for the bottom decile declined. In contrast, the share going to households in the second, third, and fourth income deciles increased. For total transfers including Medicaid, the pattern is similar but the shift is less pronounced, partly because Medicaid expansion was broad and reached middle-income working families. The authors interpret this as reflecting the design changes in the transfer system: TANF (which targeted the very bottom) declined sharply while Medicaid expansion (which reaches further up the distribution) grew.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-implicit-benefit-reduction-rates-and-why-do-they-matter"&gt;Q8. What are the implicit benefit reduction rates and why do they matter?&lt;/h3&gt;
&lt;p&gt;The paper derives implicit benefit reduction rates from the estimated parametric transfer functions. At zero income, earning the first dollar of income triggers a very large decline in transfers because eligibility for several programs is lost simultaneously. Specifically, earning $1 reduces non-medical transfers by more than $4,500 and total transfers by more than $11,000. This enormous implicit marginal tax reflects the discontinuity at zero income. For more realistic income increments, earning an additional $10,000 when starting from zero income reduces total transfers by more than $5,000 (over 50% implicit tax rate) and non-medical transfers by about $3,300. These findings are directly relevant for quantitative macroeconomic models that study labor supply and welfare, since the effective marginal tax on low-income workers entering employment is substantially higher than the statutory rate.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-differ-from-prior-work-on-parametric-tax-and-transfer-functions"&gt;Q9. How does the paper differ from prior work on parametric tax and transfer functions?&lt;/h3&gt;
&lt;p&gt;The closest antecedents are Gouveia and Strauss (1994), Heathcote, Storesletten, and Violante (2017) (who use the Benabou log-linear tax function), and Guner, Kaygusuz, and Ventura (2014) (who provide effective income tax estimates). Prior work either focused on taxes only or combined taxes and transfers into a single progressivity measure. This paper is the first to estimate effective transfer functions separately from the tax system, decomposed by program, by marital status, and by number of children. Relative to Guner et al. (2023), which assumed transfers decline linearly with income, this paper estimates a more flexible non-linear function that captures the hump at very low incomes. Relative to Ferriere et al. (2023), who propose a transfer function that increases then decreases with income, the current paper provides empirical estimates rather than a theoretical prescription. The functional form (a Ricker-style function with a separate parameter at zero income) is also more flexible than prior approximations.&lt;/p&gt;
&lt;h3 id="q10-what-data-limitations-are-noted-and-how-do-they-affect-comparability-with-other-sources"&gt;Q10. What data limitations are noted and how do they affect comparability with other sources?&lt;/h3&gt;
&lt;p&gt;The paper compares SIPP income distributions with the CPS. Both surveys yield similar Gini coefficients and variance of log income, but SIPP shows higher income shares for the bottom quantiles and lower shares for the top quintile (a discrepancy of about five percentage points). This reflects SIPP&amp;rsquo;s weaker measurement of asset income, which is a larger component of total income as one moves up the distribution. The analysis excludes self-employed households (~7%) because their income is harder to measure. The SIPP was overhauled after 2016, making cross-wave comparisons infeasible for later years; this means the paper cannot characterize the effects of post-2016 Medicaid expansion, the COVID-19 pandemic transfer surge, or recent SNAP reforms. For Medicaid, the imputation using regional HMO costs does not capture the insurance value as households themselves perceive it, a standard limitation in this literature also noted by Ben-Shalom et al. (2012) and Scholz et al. (2009) whose methods the paper follows.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-of-the-findings"&gt;Q11. What are the policy implications of the findings?&lt;/h3&gt;
&lt;p&gt;Several implications follow with scope conditions: (1) The transfer system substantially reduces income inequality, but the lion&amp;rsquo;s share of the reduction comes from Medicaid. Policies that reduce Medicaid coverage would substantially raise measured inequality, particularly at the bottom of the distribution. (2) The implicit benefit reduction rates documented — above 50% for a $10,000 income gain at the bottom — generate large effective marginal taxes on low-income households entering employment, relevant for evaluating welfare-to-work policies and for calibrating labor supply elasticities in quantitative models. (3) Despite the large size of the system, the decline in TANF spending (from above 1% of GDP to 0.1%) means that unrestricted cash assistance to the very poorest has fallen sharply; the system has shifted toward in-kind and medical programs that provide less flexibility to recipients. (4) The shift in transfer concentration away from zero-income households toward the second through fourth deciles suggests that the system increasingly supports the working poor rather than the non-working poor — a structural change in the composition of welfare that quantitative models should incorporate. These implications pertain to households headed by working-age adults (25–54), are based on pre-2016 data, and exclude the institutionalized population and self-employed households.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-key-features-of-the-parametric-function-and-how-well-does-it-fit-the-data"&gt;Q12. What are the key features of the parametric function and how well does it fit the data?&lt;/h3&gt;
&lt;p&gt;The estimated function has the form T(I) = exp(alpha) * exp(beta_0 * I) * I^beta_1 for I &amp;gt; 0 and T(0) = gamma, estimated by non-linear least squares on income-percentile averaged data. The function is flexible enough to capture: (a) a strictly positive level at zero income; (b) an initial increase then decrease at very low positive incomes (the hump); (c) a decay toward zero at high incomes that can be faster or slower depending on beta_1. The fit is shown to be close — Figure 7 documents tight confidence intervals around mean transfers by percentile, confirming that a smooth function well approximates the data. Parameter estimates are provided for each individual program, for non-medical aggregates, for total transfers, and separately for married and single households and by number of children (in appendix tables C10–C12). The zero-income gamma parameter is notably small for TANF (0.00) and large for Medicaid (0.24) and total transfers (0.26), consistent with the descriptive findings on coverage.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Means-tested transfer&lt;/strong&gt;: In this paper, a government transfer program for which eligibility and benefit amounts are conditioned on household income and assets, targeting the non-retired working-age population. The six programs studied are TANF, SNAP, WIC, SSI, housing assistance, and Medicaid.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive margin of coverage&lt;/strong&gt;: The fraction of months in a given calendar year during which a household receives a positive transfer amount, as distinct from the extensive margin (whether the household receives any transfer at all during the year). The paper documents both margins separately.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implicit benefit reduction rate (implicit penalty)&lt;/strong&gt;: The reduction in transfer payments associated with a marginal increase in non-transfer income, expressed as the derivative of the estimated transfer function with respect to income. In this paper the implicit penalty at zero income is very large because moving from zero to any positive income simultaneously triggers loss of eligibility in multiple programs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Unconditional vs. conditional transfer&lt;/strong&gt;: Unconditional transfers are averages computed over all households at a given income level, including non-recipients. Conditional transfers are averages computed only among households that actually receive a positive amount. The paper shows that the steep decline in unconditional transfers with income is almost entirely a coverage effect; conditional amounts remain relatively stable across the distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ricker transfer function&lt;/strong&gt;: The parametric functional form T(I) = exp(alpha) * exp(beta_0 * I) * I^beta_1 adopted by the paper to fit the non-linear relationship between normalized household income and normalized transfer receipt for I &amp;gt; 0, with a separate parameter gamma for I = 0. Borrowed from the Ricker (1954) stock-recruitment model in fisheries biology and chosen for its flexibility in capturing the hump-shaped pattern at very low incomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-medical transfers&lt;/strong&gt;: The aggregate of TANF, SNAP, WIC, SSI, and housing assistance — the programs that provide cash or in-kind support excluding health insurance. The paper distinguishes these from total transfers throughout to separate the role of Medicaid, which dominates all other programs in magnitude.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Medicaid imputation&lt;/strong&gt;: The procedure used to assign a monetary value to Medicaid enrollment, following Scholz et al. (2009) and Ben-Shalom et al. (2012). Each enrolled household member is assigned the cost of a single HMO policy in their Census region (from the Kaiser Foundation Employer Health Benefits survey), with family policies or sums of individual policies used for multi-member households, and a 2.5× multiplier for elderly or disabled individuals to reflect higher medical needs.&lt;/p&gt;</description></item><item><title>Medical innovation and health disparities</title><link>https://macropaperwarehouse.com/papers/medical-innovation-and-health-disparities/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/medical-innovation-and-health-disparities/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks why medical innovation can widen health disparities even when it unambiguously improves health for everyone who takes it. The authors argue that the standard access-versus-preferences dichotomy is a false one: disadvantaged patients can rationally forgo effective medications because treatment side effects interfere with work, and the income cost of not working is particularly severe for low-education workers who hold physically demanding, inflexible jobs. Health-maximizing and welfare-maximizing behavior are therefore not the same thing, and the gap between the two is systematically larger for lower-education individuals.&lt;/p&gt;
&lt;p&gt;The empirical setting is the introduction of Highly Active Antiretroviral Therapy (HAART) for HIV in the mid-1990s. HAART was substantially more effective than prior mono- and combo-therapy at preventing AIDS progression and death, but it produced harsh physical side effects (fatigue, diarrhea, headache, fever). Data come from the Multi-Center AIDS Cohort Study (MACS), a semi-annual panel of men who have sex with men in Baltimore, Chicago, Pittsburgh, and Los Angeles, covering 1991–2003. After sample restrictions, the analysis uses 11,290 person-visit observations for 1,201 HIV-positive individuals aged 30–64, approximately 63% of whom hold a college degree or more. The study dichotomizes education into less-than-college versus college-or-more and tracks treatment choices, labor supply, immune-system health (CD4 count, with AIDS threshold at 250), physical ailments, income, insurance, and out-of-pocket medical expenditures.&lt;/p&gt;
&lt;p&gt;The structural model is a lifecycle discrete-choice dynamic programming framework in which forward-looking individuals simultaneously choose treatment (no treatment, monotherapy, combotherapy, and post-1995 HAART) and full-time work or non-work each half-year period to maximize expected lifetime utility. Health and survival evolve stochastically as functions of prior health, treatment, and age. Utility is a function of consumption (income minus out-of-pocket expenses), ailments, and labor supply, with utility parameters allowed to differ by education. The model is estimated via maximum likelihood using nested backwards induction; the quasi-experimental introduction of HAART as an unanticipated shock helps identify utility parameters.&lt;/p&gt;
&lt;p&gt;Key quantitative results: (1) HAART drastically reduced mortality for both groups—six-month mortality fell from 9% to 2% for less-educated men and from 6% to 1% for college graduates—and raised the probability of maintaining a high CD4 count from 62% to 78% (less-educated) and 68% to 83% (college+). (2) Despite equivalent access (both groups face roughly 91-95% insurance coverage and similarly low out-of-pocket costs), lower-educated men adopted HAART at a lower rate (58% of post-HAART visits versus 66% for college graduates) and approximately five months later. (3) The structural utility parameters confirm that while the direct disutility of ailments is not significantly different across education groups, the disutility of working while experiencing ailments is substantially larger in magnitude for less-educated men (estimated parameter -2.73) than for college graduates (-1.97). (4) Measured as expected lifetime utility, HAART&amp;rsquo;s introduction increased value for low-CD4 men by 236.1% (less-educated) versus 176.6% (college+), but in absolute utility units the gains were larger for college graduates—establishing that HAART increased welfare inequality. (5) Decompositions show the largest single driver of the education gap in HAART value is the differential survival process; income differences also matter but financial access variables (insurance, out-of-pocket costs) explain little. (6) A simulated six-month HAART mandate improves health—by 1.7 percentage points more for less-educated men—but reduces expected lifetime value by 2.8% for the less-educated versus 1.4% for college graduates, and reduces employment by 4.1% versus 1.6%, as mandated HAART forces men into ailment-producing treatment whose side effects they cannot manage alongside work. (7) A counterfactual $10,000-per-six-months non-labor income subsidy (similar to COVID-19 transfer policies) reduces work by 31–49% for less-educated men and by 25–39% for college graduates, while inducing an 81.2% increase in HAART take-up among less-educated men in good health who were not previously on treatment (from 5% to 9% baseline probability), and a 44.5% increase for similar college graduates (8% to 11%). For men with AIDS-level CD4 counts not on treatment, the policy raises the probability of being healthy next period by 12.6% for less-educated men and 5.3% for college graduates.&lt;/p&gt;
&lt;p&gt;The central mechanism is a wedge between health and welfare that is steeper for disadvantaged workers: occupational conditions make it harder to work while experiencing side effects, so the opportunity cost of HAART compliance is higher. This means effective medical innovation—precisely by creating more severe side effects than older regimens—can widen welfare inequality even as it compresses mortality gaps. Clinical trials that randomize assignment to treatment and measure health outcomes will register the innovation as a success while masking the distributional welfare costs. Policy interventions that reduce the cost of not working (income transfers, labor market restructuring) can simultaneously increase HAART take-up and improve health, with effects concentrated among the disadvantaged.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-main-identification-strategy-and-what-are-the-key-threats-to-identification"&gt;Q1. What is the main identification strategy and what are the key threats to identification?&lt;/h3&gt;
&lt;p&gt;The model is estimated by maximum likelihood using nested backwards induction over observable state variables. A key identifying variation is the quasi-experimental, unanticipated introduction of HAART in 1995, which shifts the choice set mid-panel and allows the authors to trace behavioral responses to an exogenous change in treatment efficacy and side-effect profiles. Disutility of ailments and work parameters are identified by conditional choice probabilities given state variables (health, ailment status, prior treatment) and by comparing behavior before and after HAART availability. The authors follow Magnac and Thesmar (2002) to establish that under the distributional assumptions (Type I EV shocks, fixed discount factor β=0.95) and the normalization imposed, the likelihood has a unique maximum. The main threats are: (a) the assumption that individuals were surprised by HAART (no forward-looking anticipation), which simplifies the model but is explicitly noted—Hamilton et al. (2021) show that incorporating individual expectations substantially complicates the framework; (b) the exclusion of unobserved heterogeneity in the utility function, though specifications including it produce very small probabilities of a second type (below 5%); (c) the absence of borrowing and saving, which could allow more educated individuals to smooth consumption across treatment cycles—the authors note this would bias downward the disutility of working with ailments for higher-educated individuals, meaning the estimated cross-education difference in that parameter is a lower bound; (d) the sample is restricted to white men in four cities, limiting external validity; and (e) the education dichotomy collapses heterogeneity within education groups.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-through-which-education-moderates-the-health-welfare-tradeoff-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms through which education moderates the health-welfare tradeoff, and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The paper identifies two nested channels. First, the estimated structural utility parameter for working while experiencing ailments is larger in magnitude for less-educated men (θ = -2.73) than for college graduates (θ = -1.97), indicating greater disutility from combining work and side effects. The paper argues this reflects occupational sorting: lower-education men are significantly more likely to hold manual occupations (occupation score 5.12 versus 4.49 for college graduates, where higher scores indicate more manual tasks per Autor et al. 2003), making physical side effects especially incompatible with job performance. Second, lower-educated men have lower incomes ($15,373 versus $22,290 per half-year for less-educated versus college-educated, pre-HAART), so the income cost of not working is larger in relative terms, creating stronger incentives to maintain employment even at the cost of forgoing treatment. The authors decompose the relative contribution of these mechanisms in the non-labor income subsidy simulation: when they give lower-educated men the income process of higher-educated men (Appendix Figure A1), the gap in behavioral response narrows but does not close; when they give lower-educated men the disutility parameters of higher-educated men (Figure A2), similarly the gap narrows but remains. Both mechanisms are jointly operative.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-in-haart-take-up-and-welfare-value-is-documented"&gt;Q3. What heterogeneity in HAART take-up and welfare value is documented?&lt;/h3&gt;
&lt;p&gt;Education is the primary heterogeneity dimension examined. Post-HAART, lower-educated men used HAART in 58% of observations versus 66% for college graduates, were slower to start (5 months later on average), and less likely to ever use it (67% versus 81%). Health status interacts with education: low-CD4 men gain more in percentage terms from HAART because they are more in need of its health-improving effects (236.1% gain for less-educated low-CD4 versus 176.6% for college-educated low-CD4; 85.7% versus 76.3% for high-CD4 men, with college graduates gaining more in absolute utility units throughout). The welfare cost of a treatment mandate is higher for less-educated men (2.8% lifetime value decline versus 1.4%), and the employment reduction induced by the mandate is also larger for them (4.1% versus 1.6%). In the income subsidy simulation, low-CD4 men not on any medication show the largest health response. The paper does not examine race/ethnicity heterogeneity, having excluded non-white individuals from the analysis due to sampling methodology concerns.&lt;/p&gt;
&lt;h3 id="q4-what-does-the-value-decomposition-reveal-about-why-haart-benefited-more-educated-men-more"&gt;Q4. What does the value decomposition reveal about why HAART benefited more-educated men more?&lt;/h3&gt;
&lt;p&gt;Table A17 sequentially replaces the processes and parameters of lower-educated agents with those of higher-educated agents. Giving lower-educated men the income process of college graduates narrows but does not close the gap—income is not the primary driver. Replacing the insurance and medical expenditure processes slightly reduces value for less-educated men relative to giving them only the income process, because more-educated individuals actually have somewhat higher out-of-pocket costs. Changing the health and ailments processes has modest positive effects. The largest single contributor to closing the education gap is the survival process: less-educated men face much higher baseline mortality, which depresses the expected present value of all future flows including the gains from HAART. This suggests that policies targeting survival differentials (e.g., access to other health services) could partially close the HAART welfare gap. Finally, replacing the utility parameters mechanically closes the remaining gap, but preferences are less amenable to direct policy intervention than the survival process.&lt;/p&gt;
&lt;h3 id="q5-what-do-the-treatment-mandate-simulations-show-and-why-do-they-matter-for-evaluating-clinical-trials"&gt;Q5. What do the treatment mandate simulations show, and why do they matter for evaluating clinical trials?&lt;/h3&gt;
&lt;p&gt;A six-month HAART mandate mimics randomized assignment to treatment in a clinical trial. It improves health—the probability of high CD4 rises by 1.7 percentage points more for less-educated men than baseline (reflecting a larger baseline gap in HAART use)—which would appear a policy success from a health-only perspective. However, expected lifetime utility falls by 2.8% for less-educated men and 1.4% for college graduates, because mandated HAART forces individuals into ailment-inducing treatment they would not have chosen, inhibiting labor supply. Employment falls by 4.1% for less-educated men versus 1.6% for college graduates. Appendix analyses removing the ailment-producing properties of treatment largely eliminate both the welfare cost and the employment effect, confirming that ailments are the mediating channel. This shows that clinical trials—which typically report health endpoints and do not measure welfare or distributional consequences—can mask the costs that effective but side-effect-heavy treatments impose, and that those costs fall disproportionately on less-advantaged patients.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-non-labor-income-subsidy-simulation-show-and-which-groups-respond-most"&gt;Q6. What does the non-labor income subsidy simulation show, and which groups respond most?&lt;/h3&gt;
&lt;p&gt;A permanent $10,000-per-six-months increase in non-employment income (approximately 50% of median income, calibrated to COVID-era transfer policies) induces labor force exit across all groups but concentrates its health-promoting effects among disadvantaged men who were not already on HAART. Among relatively healthy (high-CD4) less-educated men not using any medication, HAART take-up rises by 81.2% (from 5% to 9%); the corresponding figure for college graduates is 44.5% (from 8% to 11%). Among men with AIDS-level (low) CD4 not on treatment, the probability of being healthy next period increases by 12.6% for less-educated men and 5.3% for college graduates. Men already on HAART—who are unlikely to change treatment regardless—show little response. The policy has small but positive health externalities beyond the immediate recipients, since people on antiretrovirals have lower viral loads and lower transmission risk. Decomposition simulations (Appendix Figures A1–A2) show that both the income-level channel and the disutility-of-work-with-ailments channel independently contribute to the larger lower-education response, with neither alone sufficient to fully explain the differential.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q7. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The paper is most closely related to Papageorge (2016, Quantitative Economics), which uses the same MACS data and setting to link non-uptake of HAART to labor supply and side effects. The key difference is scope: Papageorge (2016) focuses on individual-level mechanisms; the present paper&amp;rsquo;s goal is to characterize distributional differences in the health-welfare tradeoff across education groups and to show that innovation can exacerbate existing inequality. Chan, Hamilton, and Papageorge (2016, Review of Economic Studies) also use the MACS setting to study the value of medical innovation, and Hamilton, Hincapié, Miller, and Papageorge (2021, International Economic Review) examine the diffusion of HAART. Relative to the sociological fundamental cause theory literature (Link and Phelan 1995; Phelan et al. 2010), which documents that medical innovations tend to widen health disparities, the present paper provides a structural quantification of the specific mechanisms and their relative magnitude. Relative to papers attributing health disparities primarily to access barriers (insurance, cost), the paper provides evidence that for this sample—where insurance coverage exceeds 91% even for less-educated men and HIV drugs are inexpensive—access explains little of the educational disparity in HAART use or health outcomes.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The core implication is that policies reducing the cost of not working—income transfers, disability benefits, worker protections—can raise HAART adoption and improve health among disadvantaged patients, precisely the group for whom standard health-access policies have limited traction. The non-labor income subsidy simulation suggests that the health improvements are modest in absolute magnitude (a 0.2% rise in probability of being healthy next period for the best-responding group among high-CD4 non-HAART users, and 13% for low-CD4 non-HAART users), but there are unmodeled positive externalities through reduced transmission risk that would multiply the social return. Scope conditions: (1) The sample is white men who have sex with men in four U.S. cities during 1991–2003, enrolled in a prospective cohort study; generalizability to other populations (women, racial minorities, other diseases) is uncertain. (2) The income subsidy that triggers HAART take-up must be large enough to induce labor force exit; a $10,000 per-six-months transfer is needed to generate the simulated behavioral response, larger for higher-income workers. (3) The paper explicitly notes that drug costs and insurance are not binding constraints in this sample, and the policy conclusions may differ in settings with weaker drug coverage. (4) Mental health is excluded from the model; the paper shows depression variables have smaller effects on treatment choice than the physical mechanisms included, but mental health could independently affect some populations&amp;rsquo; response. The paper&amp;rsquo;s conclusions extend to other conditions where effective treatment has disabling side effects and disadvantaged patients hold inflexible physical jobs—the authors invoke COVID-19 as a contemporary analog.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-conducted"&gt;Q9. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The authors report several robustness exercises. Treatment transition results are shown to be robust to defining the HAART introduction period as survey visit 23 or 25 rather than 24. Ailment specifications are noted to be robust to varying the type or frequency of ailments counted (citing Papageorge 2016 for this). Specifications including unobserved heterogeneity in the utility function produce very small second-type probabilities (below 5%), arguing against its inclusion. The treatment mandate simulations are run under three alternative shock-assignment methods (2 draws, 8 draws, and the preferred 2-draw approach), with results consistent across methods on the main welfare-versus-health asymmetry. Appendix Tables A19 and A20 remove ailments from all medications and from HAART only, respectively, confirming that the welfare cost of mandates is driven by treatment-induced ailments. Appendix Figures A1 and A2 mechanically decompose the education-differential response to the income subsidy by replacing income processes and disutility parameters separately, confirming that both channels are active. The model fit (Table A9) shows overall employment (66% model, 66% data) and HAART use (33% model, 36% data) closely matching, though the model slightly over-predicts medication use among low-CD4 individuals.&lt;/p&gt;
&lt;h3 id="q10-why-does-the-paper-focus-on-white-men-only-and-what-does-this-imply-for-interpretation"&gt;Q10. Why does the paper focus on white men only, and what does this imply for interpretation?&lt;/h3&gt;
&lt;p&gt;The authors drop 1,098 observations from 390 non-white individuals because of concerns about the sampling methodology used to recruit the refresher sample for those individuals—specifically, non-white participants entered the panel via a different selection process that could confound estimates. The paper does not investigate racial disparities in HAART take-up, which are also well-documented in the literature. This is a significant limitation because HIV/AIDS has disproportionately affected Black men in the United States, and the mechanisms the paper identifies—occupational sorting, income constraints, disutility of working with ailments—may operate differently or more intensely along racial lines. The authors acknowledge this limitation and note that the structural framework could in principle be applied to other groups if appropriate data were available.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Health-welfare tradeoff&lt;/strong&gt;: In this paper, the wedge between the action that maximizes health (taking effective medication despite side effects) and the action that maximizes lifetime utility (avoiding medication to remain employed and maintain income). The tradeoff is not a bias or error but a rational response to economic constraints, and it is wider for less-educated individuals whose occupational conditions make working with side effects especially costly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HAART (Highly Active Antiretroviral Therapy)&lt;/strong&gt;: A combination antiretroviral HIV treatment introduced in the mid-1990s, far more effective than prior mono- or combo-therapy at improving CD4 count and preventing AIDS-level immune decline and death. In this paper&amp;rsquo;s model, HAART serves as the innovation whose adoption the authors study: it is more efficacious but produces harsher side effects than earlier treatments, and its introduction is treated as an unanticipated aggregate shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Disutility of working with ailments&lt;/strong&gt;: A structural utility parameter (θ_2,f=0) capturing how much worse-off an agent feels from working while experiencing physical ailments (fatigue, diarrhea, headache, fever). Estimated at -2.73 for less-educated men and -1.97 for college graduates, this parameter is the primary driver of the differential health-welfare tradeoff across education groups and explains why side-effect-bearing treatments like HAART are disproportionately avoided by lower-education workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Treatment mandate simulation&lt;/strong&gt;: A counterfactual in which all agents are assigned to HAART for six months (eliminating choice among other treatment options), used to mimic randomized assignment in a clinical trial. The simulation is designed specifically to illustrate that health improvements observable in a clinical trial coexist with welfare reductions and employment disruptions that would not be captured in standard trial endpoints.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fundamental cause theory&lt;/strong&gt;: A sociological framework (Link and Phelan 1995) arguing that socioeconomic status is a &amp;lsquo;fundamental cause&amp;rsquo; of health disparities that persists despite or is even amplified by medical innovation, because more advantaged individuals are better positioned to adopt and benefit from new treatments. The paper provides structural economic microfoundations for this theory by quantifying the mechanisms through which HAART&amp;rsquo;s introduction widened the welfare gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-labor income subsidy&lt;/strong&gt;: A counterfactual policy simulation in which non-employment income is raised by $10,000 per six months (approximately 50% of the median person&amp;rsquo;s income), modeled after COVID-19 transfer policies. In the paper&amp;rsquo;s model this policy reduces employment but increases HAART take-up and health improvements particularly for less-educated HIV-positive men who were previously forgoing treatment to maintain income from work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Source text origin&lt;/strong&gt;: Not a paper-specific concept but denoted here: the full working paper text was obtained from the NBER Working Paper (No. 28864), not from abstract-only, satisfying the GUARD requirement.&lt;/p&gt;</description></item><item><title>Non-Tariff Barriers in the U.S.-China Trade War</title><link>https://macropaperwarehouse.com/papers/non-tariff-barriers-in-the-u.s.-china-trade-war/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/non-tariff-barriers-in-the-u.s.-china-trade-war/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Chen, Hsieh, and Song study the use of unofficial non-tariff barriers (NTBs) by China during the U.S.-China trade war of 2018–2019 and in the first year of the Phase 1 purchase agreement (2020). The central motivation is that much prior analysis of the trade war focused on announced tariff hikes, yet abundant anecdotal evidence — permit requirements for U.S. pet food, pest-inspection orders on U.S. apples and lumber, changes to pig-feed formulas reducing soybean content — points to a parallel, opaque regulatory channel. The critical puzzle the paper highlights is that China&amp;rsquo;s purchases of U.S. goods rose by 156 percent between 2019 and 2020 without any reduction in tariffs, which is only explicable if NTBs were used in reverse to favour U.S. exporters during the Phase 1 period.&lt;/p&gt;
&lt;p&gt;The paper uses Chinese customs administrative data from 2015 to July 2020, covering 946 HS-6 products aggregated by state-owned versus non-state importer and by source country. Tariff data are constructed from official Customs Tariff Commission documents listing each round of retaliatory hikes beginning April 2018. The empirical strategy proceeds in three steps. First, demand (elasticity of substitution across source countries, epsilon) and supply (gamma) elasticities are estimated by regressing changes in import quantities and CIF prices on changes in tariff rates, using product-country fixed effects so identification comes from within-product, cross-country variation in tariff changes. The identifying assumption — that tariff changes across countries are orthogonal to NTB changes and foreign supply shifts — is validated empirically. The estimated demand elasticity is epsilon = 3.36 for agriculture and 2.34 for manufacturing; supply elasticities of 42 (agriculture) and 71 (manufacturing) imply near-horizontal foreign supply curves, so essentially all the incidence of Chinese trade barriers falls on Chinese consumers.&lt;/p&gt;
&lt;p&gt;Second, NTBs are inferred as a residual: the change in U.S. import quantities relative to imports from other countries of the same HS-6 product, after netting out the estimated price and tariff effect. A normalisation sets the import-weighted average NTB change on non-U.S. source countries to zero, so the residual is attributed to U.S.-specific barriers. This procedure is run separately for non-state and state importers. The tariff-equivalent of NTBs on U.S. agricultural products faced by non-state importers rose by 0.73 log points between 2017 and 2019, while NTBs on state importers were essentially unchanged (Table 4). The weighted average NTB increase for agriculture was 0.60 log points, compared to a tariff increase of 17 percentage points (from 7.5% to 24.5%). For manufactured goods, average NTBs rose by only 0.16 log points versus a tariff increase of 9 percentage points (5.6% to 14.6%). NTBs were highly concentrated: the tariff equivalent rose by 1.0 log points for oil seeds, 1.5 log points for cereals, and 1.1 log points for ores, slag and ash. The variance of tariff-adjusted import growth across HS-6 products increased 18-fold from 0.296 (2015–2017) to 5.31 (2017–2019), and controlling for state versus non-state ownership accounts for 38% of that increase.&lt;/p&gt;
&lt;p&gt;Third, welfare effects are computed using a three-nest CES model (HS-6 products, importer firms, source countries). Tariffs harm welfare via dispersion of tariff rates across source countries; NTBs harm welfare via both the mean and dispersion of NTBs across source countries, firm types, and products, and also because — unlike tariffs — NTBs generate no fiscal revenue. The total welfare loss to China in 2019 relative to 2017 is estimated at $40 billion, of which 92% is attributable to NTBs rather than tariffs (Table 7). For agricultural products alone, NTBs account for 86% of the $12.7 billion welfare loss; for manufacturing they account for 94.1% of the $27.2 billion loss. Crucially, for a given dollar reduction in U.S. imports, NTBs impose approximately six times the welfare cost of equivalent tariff hikes (the Figure 2 text says &amp;ldquo;five times&amp;rdquo;), because NTBs (i) generate no revenue and (ii) create misallocation by applying to some importers (non-state) but not others (state-owned). By 2020 China&amp;rsquo;s welfare loss relative to 2017 widened further to $48.11 billion, as NTB reversals in agriculture were partial and manufacturing NTBs were not reversed at all. The paper also documents that the Chinese government&amp;rsquo;s choice of instrument was strategic: tariff hikes were smaller in sectors with a larger pre-war state importer share, while NTB hikes on non-state importers were larger in those same sectors, consistent with a government pursuing dual objectives of punishing U.S. exporters while protecting state-firm profits.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-its-key-assumption"&gt;Q1. What is the core identification strategy and its key assumption?&lt;/h3&gt;
&lt;p&gt;The demand elasticity (epsilon) and supply elasticity (gamma) are estimated from a system of two equations: the change in log import quantity and the change in log CIF price, both regressed on the change in log tariff rates, with product-country fixed effects and year fixed effects. The identifying assumption is that tariff changes across source countries are orthogonal to NTB changes and foreign supply shifts — i.e., China&amp;rsquo;s retaliatory tariff schedule was not systematically targeted at products where NTBs were also rising or where foreign supply conditions were deteriorating. The authors validate this assumption in two ways: (1) Appendix Figure A2 shows near-zero correlation between imputed NTB changes and tariff changes across HS-6 product-country pairs (OLS coefficient 0.014); (2) Appendix Figure A3 shows near-zero correlation between pre-war import growth (2015–2017) and post-war tariff changes (OLS coefficient -0.02), arguing against correlated foreign supply trends.&lt;/p&gt;
&lt;h3 id="q2-how-exactly-are-ntbs-measured-and-what-normalization-is-required"&gt;Q2. How exactly are NTBs measured and what normalization is required?&lt;/h3&gt;
&lt;p&gt;NTBs are inferred as a structural residual. From the CES demand function, the change in non-state imports of a U.S. product relative to the same product from another source country equals minus epsilon times the relative change in tariff-inclusive CIF price, minus epsilon times the relative NTB. Given estimated epsilon and data on prices and tariffs, the relative NTB (U.S. vs. other countries) is identified. To convert this into the absolute NTB on U.S. goods, the paper normalizes the import-expenditure-weighted average NTB change on all non-U.S. source countries to zero. State-importer NTBs are then backed out from the ratio of state to non-state import growth for U.S. products, using equation (7), which relies on the elasticity of substitution between state and non-state firm types (eta = 3, borrowed from Khandelwal, Schott and Wei 2013).&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-threats-to-identification-and-how-are-they-addressed"&gt;Q3. What are the main threats to identification and how are they addressed?&lt;/h3&gt;
&lt;p&gt;Three threats are discussed. (1) Quality or supply changes specific to U.S. products: if imputed NTBs reflect deteriorating U.S. product quality rather than Chinese regulatory barriers, U.S. exports to non-China markets should also fall for the same HS-6 products. Appendix Figure A1 shows no such correlation (OLS slope 0.016, SE 0.007), confirming NTBs are China-specific. (2) Endogenous targeting of tariffs toward products also receiving NTBs (violating the orthogonality assumption): Appendix Figure A2 directly shows near-zero correlation. (3) Correlated pre-trends: Appendix Figure A3 shows no correlation between 2015–2017 import growth and 2017–2019 tariff changes, so pre-existing trends do not appear to have driven the targeting of tariffs.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-across-firm-ownership-is-documented"&gt;Q4. What heterogeneity across firm ownership is documented?&lt;/h3&gt;
&lt;p&gt;NTBs fell almost entirely on non-state importers of U.S. agricultural products. Non-state NTBs rose by 0.73 log points (2017–2019) while state NTBs were essentially unchanged (Table 4, column 3 vs. column 4). The state share of Chinese agricultural imports from the U.S. roughly doubled from 19.3% in 2017 to 39.8% in 2019 (Table 2), before returning to ~20% in 2020. For imports from the rest of the world, the state share remained stable at ~20% throughout. In manufacturing, state-importer NTBs declined slightly (-0.066) while non-state NTBs rose modestly (0.023). The divergence between state and non-state importers accounts for 38% of the 18-fold increase in variance of tariff-adjusted import growth.&lt;/p&gt;
&lt;h3 id="q5-what-product-level-heterogeneity-is-found-in-the-use-of-ntbs-vs-tariffs"&gt;Q5. What product-level heterogeneity is found in the use of NTBs vs. tariffs?&lt;/h3&gt;
&lt;p&gt;NTBs were highly product-concentrated compared to tariffs. Table 5 shows the largest NTB increases in oil seeds (+1.006 log points), cereals (+1.492), and food industry residues (+0.688), all products where the U.S. held large pre-war import shares. For manufactured goods, the largest NTB increases occurred in ores, slag and ash (+1.106) and vehicles (+0.366). By contrast, tariff hikes were distributed more broadly across products. Table 9 shows that, across HS-6 products, (a) tariff increases were significantly smaller for products with a higher pre-war state importer share (OLS coefficient -0.202) and (b) non-state importer NTB increases were significantly larger for those same products (OLS coefficient +4.431). Both patterns hold when controlling for the U.S. import share in total imports of the product.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-welfare-framework-and-what-are-its-scope-conditions"&gt;Q6. What is the welfare framework and what are its scope conditions?&lt;/h3&gt;
&lt;p&gt;Welfare is derived from a three-level CES utility function over HS-6 products (elasticity sigma), importer firms (elasticity eta), and source countries (elasticity epsilon). Tariff revenue is rebated to consumers; NTB costs are not. The welfare cost operates through three channels: (1) tariffs raise dispersion of prices across source countries, reducing welfare with elasticity epsilon; (2) NTBs affect both the mean and the dispersion of import prices, with no offsetting revenue effect; (3) differential NTBs across firm types (state vs. non-state) add a misallocation channel scaled by eta. The framework accounts for expenditure reallocation across source countries within an HS-6 product and across HS-6 products, but not between imported and domestic Chinese goods. This last restriction means welfare losses are likely understated, as the model does not capture the cost of switching from foreign to domestic substitutes.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-quantitative-welfare-results-and-how-do-they-decompose"&gt;Q7. What are the quantitative welfare results and how do they decompose?&lt;/h3&gt;
&lt;p&gt;Total welfare loss in 2019 relative to 2017: $40 billion. Agriculture: $12.7 billion (of which tariffs account for $1.7B and average NTBs for an additional $9.3B; differential state/non-state NTBs add a further $1.7B). Manufacturing: $27.2 billion (of which tariffs account for only $1.6B; average NTBs add $23.5B and differential NTBs a further $2.1B). NTBs&amp;rsquo; share: 92% of total (86% for agriculture, 94% for manufacturing). By 2020, the overall welfare loss widened to $48.11 billion, because partial NTB reversal in agriculture was more than offset by continued welfare losses from manufacturing NTBs.&lt;/p&gt;
&lt;h3 id="q8-why-are-ntbs-so-much-more-costly-per-dollar-of-import-reduction-than-tariffs"&gt;Q8. Why are NTBs so much more costly per dollar of import reduction than tariffs?&lt;/h3&gt;
&lt;p&gt;Two mechanisms. First, tariffs generate revenue that is assumed to be rebated to consumers, partially offsetting their welfare cost; NTBs generate no government revenue. Second, because NTBs are unofficial and opaque, they can be and were applied selectively to non-state importers but not to state importers, creating misallocation: within an HS-6 product, some importers face artificially high effective prices while others (state firms) do not, so the aggregate consumption basket becomes inefficient. The welfare elasticity with respect to import value is approximately five to six times larger for NTBs than for tariffs (Figure 2; the abstract states six times, the Figure 2 text states five times — a minor internal discrepancy).&lt;/p&gt;
&lt;h3 id="q9-what-does-the-paper-show-about-the-phase-1-purchase-agreement-2020"&gt;Q9. What does the paper show about the Phase 1 purchase agreement (2020)?&lt;/h3&gt;
&lt;p&gt;In 2020 China agreed to increase purchases of U.S. goods without reducing tariffs. The paper shows this was accomplished by partially reversing NTBs. The average NTB for agricultural products fell from +0.60 log points (2017–2019) to +0.14 log points over the full 2017–2020 period, implying substantial 2020 reversal. This reversal applied exclusively to non-state importer NTBs on agricultural products; state importer NTBs and manufacturing NTBs were not reversed. The U.S. share of Chinese agricultural imports rose from 13.7% in 2019 to 17.2% in 2020 despite unchanged tariffs (Table 1), directly confirming the NTB reversal interpretation. Welfare in 2020 from agricultural imports partly recovered but remained $7.3 billion below 2017 baseline; manufacturing welfare loss persisted, yielding an overall 2020 welfare loss of $48.11 billion.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-prior-work-on-the-us-china-trade-war"&gt;Q10. How does this paper relate to prior work on the U.S.-China trade war?&lt;/h3&gt;
&lt;p&gt;The paper builds most directly on Fajgelbaum et al. (2019), borrowing their IV procedure to estimate demand and supply elasticities (using tariff variation across source countries as instruments) and replicating their finding of near-horizontal foreign supply curves. It differs in focusing on Chinese consumers rather than American consumers and in measuring NTBs in addition to tariffs. It also extends Khandelwal, Schott and Wei (2013), whose analysis of state-firm export quotas motivated the state/non-state ownership dimension; the current paper inverts the logic to study selective barriers on non-state importers. Benguria and Safdie (2021) similarly find product variation in U.S. exports to China correlated with state ownership, but do not impute NTBs structurally or quantify welfare. Ma, Ning and Xu (2021) and Liu (2020) use Chinese customs data to document tariff effects on imports but do not examine NTBs. Chor and Li (2021) use night-lights data to estimate aggregate tariff exposure effects.&lt;/p&gt;
&lt;h3 id="q11-what-robustness-checks-are-conducted-and-what-do-they-show"&gt;Q11. What robustness checks are conducted and what do they show?&lt;/h3&gt;
&lt;p&gt;Three main robustness exercises. (1) Falsification test: for products where high NTBs are imputed, U.S. exports to non-China markets do not fall (Appendix Figure A1, slope 0.016, SE 0.007), confirming NTBs are China-specific rather than reflecting U.S.-side supply deterioration. (2) Orthogonality check: Appendix Figure A2 shows near-zero correlation between imputed NTBs and tariff changes across product-country pairs. (3) Alternative country normalization: NTBs are estimated for the four largest non-U.S. exporters to China (Brazil, Canada, Thailand, Australia), assuming barriers on the remaining countries average zero. Brazil, Canada, and Thailand show essentially zero imputed NTB changes 2017–2019, consistent with the identifying normalization. Australia shows a modest NTB increase consistent with documented retaliations after Australia&amp;rsquo;s 2018 national security law, but far smaller than the U.S. NTB increase. Additionally, Appendix Tables A1-A3 re-run all estimates with alternative parameter values: sigma = 1 (instead of 1.47/1.25) and eta = 5 (instead of 3). All qualitative results survive: NTBs exceed tariffs in magnitude, fall disproportionately on non-state importers, and impose far larger welfare costs per dollar of import reduction.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy implication is that opaque regulatory tools are an unusually costly instrument of trade retaliation — approximately five to six times more costly per unit of import reduction than equivalent tariffs — because they neither generate revenue nor require the same importer to bear equal costs. If the Chinese government&amp;rsquo;s objective was to punish U.S. exporters, it chose a particularly self-damaging instrument. A secondary implication concerns the Phase 1 deal: the deal&amp;rsquo;s purchase commitments were met not through tariff reductions but through NTB reversals, and those reversals were partial, selective (agriculture but not manufacturing; non-state but not state), and left China&amp;rsquo;s welfare substantially below the 2017 baseline. Scope conditions: the welfare model does not account for import-to-domestic substitution, so welfare costs are likely understated. The elasticity estimates assume CES preferences and a particular nesting structure. The NTB measurement relies on the normalisation that average barriers on non-U.S. sources did not change, which is validated but not directly observable.&lt;/p&gt;
&lt;h3 id="q13-what-does-the-paper-reveal-about-the-strategic-logic-of-chinas-instrument-choice"&gt;Q13. What does the paper reveal about the strategic logic of China&amp;rsquo;s instrument choice?&lt;/h3&gt;
&lt;p&gt;Section 7 shows that Chinese authorities&amp;rsquo; instrument choice is consistent with a dual-objective government: punish U.S. exporters while protecting state-firm profits. Tariffs, which apply uniformly to all importers, harm state firms importing from the U.S. as much as non-state firms. NTBs, being unofficial and selectively enforced, can exempt state importers. Regression evidence (Table 9) confirms: tariff hikes were systematically smaller for products with higher pre-war state importer shares (coefficient -0.202, SE 0.042), while NTB hikes on non-state importers were systematically larger for the same products (coefficient +4.431, SE 0.655). These patterns hold controlling for the U.S. product share in total Chinese imports.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Non-tariff barrier (NTB)&lt;/strong&gt;: In this paper, unofficial and opaque regulatory measures — health inspections, permit requirements, informal directives to importers — that function as trade barriers but are not publicly disclosed as such and are not uniformly applied to all importing firms. Measured in tariff-equivalent units as the residual change in U.S. import share after controlling for tariff and price effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tariff-equivalent of NTBs&lt;/strong&gt;: The ad-valorem tariff rate that would produce the same reduction in import demand as the estimated NTB, derived from the structural demand equation. Expressed in log points (e.g., 0.60 log points for average agricultural NTBs in 2017–2019).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Misallocation from selective NTBs&lt;/strong&gt;: The welfare loss that arises specifically because NTBs are applied to non-state importers but not state importers within the same HS-6 product category. This within-product dispersion of effective prices across firms generates an allocative inefficiency absent when tariffs are used, since tariffs apply uniformly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Phase 1 purchase agreement&lt;/strong&gt;: The January 2020 U.S.-China trade deal in which China committed to purchasing specified amounts of U.S. goods in 2020–2021. The paper shows that China fulfilled these commitments by reversing NTBs rather than reducing tariffs, and that the reversal was partial, concentrated in agricultural imports by non-state firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Elasticity of substitution across source countries (epsilon)&lt;/strong&gt;: The parameter governing how sensitive Chinese import demand for an HS-6 product from a given country is to that country&amp;rsquo;s relative price. Estimated at 3.36 for agriculture and 2.34 for manufacturing using tariff variation as an instrument.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;State vs. non-state importer&lt;/strong&gt;: The ownership classification of Chinese importing firms in the customs data. State-owned importers were largely exempt from NTBs during the trade war, while non-state (private) importers bore nearly all of the NTB increases on U.S. agricultural products. This differential application is the central mechanism generating misallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Welfare channel distinction: tariffs vs. NTBs&lt;/strong&gt;: Tariffs affect welfare only through the dispersion of prices across source countries (revenue is rebated). NTBs affect welfare through both the mean and dispersion of prices across source countries, firm types, and products, with no revenue offset. This structural distinction is why the paper finds NTBs impose approximately five to six times greater welfare cost per dollar of import reduction.&lt;/p&gt;
&lt;!-- flags: Minor internal discrepancy in paper: abstract and conclusion state NTBs impose ~6x the welfare cost of equivalent tariffs per dollar of import reduction; Figure 2 text states ~5x. Both figures are in the source text; the summary uses 'approximately six times' per the abstract/conclusion. --&gt;</description></item><item><title>On the Effects of Monetary Policy Shocks on Income and Consumption Heterogeneity</title><link>https://macropaperwarehouse.com/papers/on-the-effects-of-monetary-policy-shocks-on-income-and-consumption-heterogeneity/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/on-the-effects-of-monetary-policy-shocks-on-income-and-consumption-heterogeneity/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks how conventional and informational monetary policy shocks affect the cross-sectional distributions of labor earnings, consumption, and financial income in the United States. The motivation is the growing concern, particularly in the aftermath of the global financial crisis, about distributional consequences of central bank actions. Existing studies either include scalar inequality statistics in standard VARs — losing information about the full distribution — or rely on indirect approaches that hold household portfolio compositions fixed. Chang and Schorfheide instead apply the functional VAR (fVAR) framework developed in Chang, Chen, and Schorfheide (2024, JPE forthcoming) that stacks macroeconomic aggregates alongside the full time-varying cross-sectional density, represented as a log probability density function approximated via a cubic-spline sieve. This allows simultaneous, internally-consistent IRFs for percentiles, Gini coefficients, 90-10 ratios, standard deviations, and other distributional statistics without the risk of quantile crossings.&lt;/p&gt;
&lt;p&gt;The earnings analysis uses monthly micro data from the Current Population Survey (CPS), sample period 1990:M2 to 2016:M12. The consumption and financial income analyses use quarterly Consumer Expenditure Survey (CEX) data from 1990:Q2 to 2016:Q4. Monetary policy shocks are identified via the Jarocinski-Karadi (2020) high-frequency instruments — surprises in the three-month fed funds futures and in S&amp;amp;P 500 index — used as internal instruments in the structural VAR. The instruments isolate (a) conventional monetary policy shocks (interest rate surprise, stock price opposite direction) and (b) informational shocks (interest rate and stock price surprise in the same direction). Sign restrictions set-identify the two shocks. Bayesian estimation uses a Chan (2022) Normal-Inverse Gamma prior suitable for high-dimensional VARs; model selection (sieve order K, lag length p, hyperparameters) is done by maximizing the marginal data density (MDD). The shock normalization corresponds to an unanticipated 25-basis-point cut in the three-month federal funds rate.&lt;/p&gt;
&lt;p&gt;Main quantitative findings:&lt;/p&gt;
&lt;p&gt;Earnings (conventional shock): An expansionary shock reduces earnings inequality, primarily through the employment (extensive) margin. At the posterior median, the 10th earnings percentile rises by up to 5% relative to steady state, the 20th percentile by up to 1%, while the 80th and 90th percentiles are essentially unaffected. The Gini coefficient for labor earnings falls from approximately 0.431 to 0.428 over a 36-month horizon. The 90-10 earnings ratio falls from approximately 12.27 to 11.76 after 36 months. These effects are driven almost entirely by individuals moving from unemployment into employment (the point mass at zero in the earnings distribution falls as the unemployment rate drops by approximately 0.3 percentage points at the posterior median after three years). When the unemployed point mass is excluded from the inequality computation, the inequality effect is small and short-lived, confirming that the employment channel dominates. The estimated Gini drop of 0.001–0.003 is broadly consistent with the HANK model of Ma (2021) with indivisible labor, which predicts a drop of approximately 0.001 for a comparable shock.&lt;/p&gt;
&lt;p&gt;Consumption (conventional shock): The expansionary shock generates a weakly positive (inequality-increasing) effect on consumption inequality at the posterior median, but with wide credible bands that span both positive and negative values. The cross-sectional standard deviation of consumption, the 90-10 ratio, and the Gini coefficient all peak upon impact and remain above steady state. The slight increase appears concentrated in durable goods expenditure; nondurable and service consumption inequality shows little response at the posterior median. The contrast with the earnings result reflects: (i) only labor income is captured in the earnings analysis, while wealthy households&amp;rsquo; capital income (rising with equity and bond prices) also rises; (ii) potentially higher interest-rate sensitivity of high-consumption households.&lt;/p&gt;
&lt;p&gt;Financial income (conventional shock): No statistically significant effect on financial income inequality. The cross-sectional standard deviation and Gini coefficient of financial income do not respond to the shock. An important caveat is that the CEX misses the top-10 percent of households by financial income (visible from CDF comparison with the Survey of Consumer Finances in 2012). The households most likely to benefit from equity and bond price appreciation — captured in other studies — are absent from the sample.&lt;/p&gt;
&lt;p&gt;Informational shock: A negative informational shock (unexpected simultaneous drop in interest rates and stock prices, signaling worse-than-expected output) increases earnings inequality, mainly via a rise in unemployment. The 10th earnings percentile drops by about 2% at the posterior median. Consumption inequality, by contrast, shows the opposite pattern: the 90-10 ratio and Gini coefficient for consumption decrease, and the posterior median responses are negative, though uncertainty is substantial.&lt;/p&gt;
&lt;p&gt;Policy implication: The authors conclude that earnings inequality effects of conventional monetary policy are well-proxied by the unemployment rate response, so standard macro indicators subsume the distributional information for earnings. The small and highly uncertain responses of consumption and financial income inequality provide, in their view, support for central banks continuing to focus primarily on macroeconomic aggregates.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-monetary-policy-shocks-and-what-are-the-main-threats-to-validity"&gt;Q1. What is the identification strategy for monetary policy shocks and what are the main threats to validity?&lt;/h3&gt;
&lt;p&gt;The paper uses the Jarocinski-Karadi (2020) high-frequency instruments as internal instruments in a structural VAR. The two instruments are surprises in the three-month federal funds futures (ff4_hf) and surprises in the S&amp;amp;P 500 index (sp500_hf), measured in narrow windows around FOMC announcements. Sign restrictions separate two shocks: a conventional shock is identified by an interest rate increase combined with a stock price fall; an informational shock by both increasing. The key assumptions are instrument relevance (the instruments are correlated with the policy shocks) and instrument validity (the instrument innovations are uncorrelated with non-policy structural shocks). As a robustness check the authors also use the Nakamura-Steinsson (2018) instruments and report very similar results. The main threat to validity is the standard one for external-instrument SVARs: the instruments may capture other economic news released simultaneously with FOMC decisions, violating the exclusion restriction. The informational shock identification partially addresses this by explicitly modeling the central bank&amp;rsquo;s information revelation.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-functional-var-approach-and-why-is-it-preferred-over-simpler-alternatives"&gt;Q2. What is the functional VAR approach and why is it preferred over simpler alternatives?&lt;/h3&gt;
&lt;p&gt;The functional VAR stacks macroeconomic aggregates Yt with the time-varying cross-sectional log-density of micro outcomes. The log-density is approximated by a finite-dimensional linear sieve (cubic spline basis of order K). Sieve coefficients are estimated period-by-period by maximum likelihood from the cross-section, then treated as observations in a standard VAR. The MDD selects K, lag order p, and Minnesota-type hyperparameters jointly. Compared to simply including a few inequality statistics in a VAR, the functional approach (a) derives a single coherent model from which arbitrarily many distributional statistics can be computed without quantile crossings; (b) achieves tighter credible intervals by efficiently compressing cross-sectional information through the sieve; (c) avoids the problem of internally inconsistent forward projections of stacked quantile VARs. Compared to indirect approaches (e.g., McKay-Wolf 2023), it does not require the assumption that household income or portfolio composition is fixed in response to the shock. Compared to panel approaches, it does not require high-frequency panel data, which are unavailable for the US at relevant horizons.&lt;/p&gt;
&lt;h3 id="q3-how-is-the-earnings-distribution-modeled-to-handle-unemployment"&gt;Q3. How is the earnings distribution modeled to handle unemployment?&lt;/h3&gt;
&lt;p&gt;The earnings distribution is treated as a mixture of a point mass at zero (representing unemployed individuals, whose weight equals the CPS-based unemployment rate) and a continuous part (the density of positive earnings of employed individuals, normalized to integrate to one minus the unemployment rate). The sieve density is estimated only from the positive-earnings observations, with a top-coding adjustment for right-censored values. The unemployment rate is included separately as an aggregate variable in the Yt vector. This mixture representation allows the analysis to separately identify the extensive-margin (employment) channel — changes in the probability mass at zero — from the intensive-margin channel (changes within the positive-earnings density). The key finding is that inequality effects are driven almost entirely by the extensive margin.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-in-earnings-responses-is-documented"&gt;Q4. What heterogeneity in earnings responses is documented?&lt;/h3&gt;
&lt;p&gt;In percentage terms, the expansionary monetary policy shock has the largest impact at the 10th earnings percentile (posterior median response of 0 to 5%), capturing workers moving out of unemployment. The 20th percentile rises by 0 to 1%. The 80th and 90th percentiles show essentially zero response. Earnings above 2 times GDP per capita (roughly twice the labor share of GDP per capita) are essentially unaffected. When the point mass at zero is excluded and only the continuous part of the earnings distribution is analyzed, the effect on inequality statistics (Gini, 90-10 ratio) is small and short-lived, confirming that the heterogeneous response across the full distribution is driven almost entirely by the employment transition at the bottom.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-in-consumption-responses-is-documented-and-why-might-consumption-inequality-rise-while-earnings-inequality-falls"&gt;Q5. What heterogeneity in consumption responses is documented, and why might consumption inequality rise while earnings inequality falls?&lt;/h3&gt;
&lt;p&gt;At the posterior median, both the 10th and 20th consumption percentiles initially rise above steady state (h=1), then fall 0.9% to 1.3% below baseline from h=5 onwards. The 80th and 90th percentile responses are quantitatively similar in shape but slightly larger in magnitude, leading to a weakly positive net inequality effect. The Gini coefficient and 90-10 ratio for consumption peak upon impact and stay above steady state. The authors offer two explanations for the inequality-increasing result despite earnings inequality falling: (i) wealthy households also earn substantial capital income (equities, bonds) that rises with the expansionary shock, boosting their total resources and hence consumption, a channel not captured by earnings alone; (ii) higher-consumption households may have more interest-rate-sensitive consumption decisions (larger direct Euler-equation effect), or may be wealthy hand-to-mouth consumers with high MPCs. The component analysis shows the increase is concentrated in durable goods, while nondurable and services Gini responses are near zero at the posterior median.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-financial-income-analysis-find-and-what-data-limitation-is-most-important"&gt;Q6. What does the financial income analysis find and what data limitation is most important?&lt;/h3&gt;
&lt;p&gt;The financial income distribution estimated from the CEX shows no statistically significant response to either the level or inequality of financial income following a conventional monetary policy shock. The cross-sectional standard deviation and Gini coefficient of financial income are essentially flat. The most important caveat is that the CEX substantially underrepresents high-financial-income households. A CDF comparison with the Survey of Consumer Finances for 2012 shows that the CEX misses the top-10 percent of households by financial income. These are precisely the households most likely to experience capital gains from equity and bond price appreciation following an interest rate cut. The fraction of households with essentially zero financial income (the point mass κt) fluctuates between 0.65 and 0.82 over the sample, so the analysis is largely capturing the lower 65–82 percent of the financial income distribution.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-informational-shock-and-how-do-its-distributional-effects-differ-from-the-conventional-shock"&gt;Q7. What is the informational shock and how do its distributional effects differ from the conventional shock?&lt;/h3&gt;
&lt;p&gt;An informational shock is defined as an unanticipated change in interest rates that conveys private central-bank information about the state of the economy — for example, a rate cut that signals the central bank expects worse output and prices than the public. It is identified by the simultaneous drop in interest rates and stock prices, the opposite pattern from the conventional shock. Aggregate effects: real GDP drops approximately 20 basis points and unemployment rises up to 0.15 percentage points after one year. Earnings distributional effects are roughly the mirror image of the conventional shock: the 10th earnings percentile drops about 2% at the posterior median, while other percentiles change little. The Gini coefficient and 90-10 ratio for earnings rise in the long run, driven by the increase in unemployment. Consumption distributional effects are different: relative consumption at the 10th and 20th percentiles rises, while the 90th percentile falls slightly, so consumption inequality (90-10 ratio, Gini) decreases. However, since aggregate consumption also falls, the rise in relative consumption at the bottom does not imply an absolute gain.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-coibion-gorodnichenko-kueng-and-silvia-2017"&gt;Q8. How does this paper relate to and differ from Coibion, Gorodnichenko, Kueng, and Silvia (2017)?&lt;/h3&gt;
&lt;p&gt;CGKS (2017) include inequality statistics directly in a VAR and use the Romer-Romer shock measure. For earnings, they find the Gini coefficient rises by about 0.0025 per 100bp contractionary shock (i.e., falls by 0.0025 for an expansionary shock); adjusting for shock size this is slightly smaller than the Chang-Schorfheide estimate of a 0.001–0.003 Gini drop per 25bp expansionary shock (which scales to 0.004–0.012 per 100bp). For consumption, CGKS find that inequality decreases in response to an expansionary shock, the opposite sign from Chang-Schorfheide&amp;rsquo;s posterior-median result (weakly increasing). The discrepancy may reflect: (i) the functional approach&amp;rsquo;s more flexible modeling of the full distribution versus using a single Gini; (ii) differences in shock identification (Romer-Romer vs. JK instruments); (iii) sample period differences. The wide credible bands in the consumption result mean the two findings are not statistically inconsistent.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-conducted"&gt;Q9. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;The authors run the following robustness exercises: (i) Nakamura-Steinsson (2018) instruments instead of Jarocinski-Karadi (2020) for the earnings VAR — results are very similar. (ii) Model selection across sieve order K ∈ {4,6,8,10} and lag length p ∈ {1,2,3,4} via MDD maximization, confirming that results are robust to the choice of approximation order. (iii) For the earnings inequality analysis, the paper explicitly separates the contribution of the employment margin from the wage distribution within employment, by recomputing inequality statistics excluding the point mass at zero — confirming that the employment channel dominates. (iv) Comparison of aggregate IRFs across all four model specifications (aggregate VAR, earnings fVAR, consumption fVAR, financial income fVAR) showing that inclusion of cross-sectional data does not substantially alter inference about aggregate variables. (v) Comparison with time-aggregated monthly-to-quarterly rescaled IRFs to validate that monthly and quarterly specifications produce consistent results.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-scope-conditions-and-limitations-of-the-findings"&gt;Q10. What are the scope conditions and limitations of the findings?&lt;/h3&gt;
&lt;p&gt;Key scope conditions: (a) The sample runs through 2016:Q4/M12, so the post-2016 period and the 2020 pandemic episode are excluded. (b) The paper uses repeated cross-sections rather than a panel, so it directly estimates how the cross-sectional distribution evolves but cannot separately identify cohort effects, individual trajectories, or nonlinearities in unit-level histories. (c) The CEX substantially misses high-financial-income households, making the financial income results inapplicable to the top 10% of the financial income distribution. (d) The functional VAR models the unconditional distribution; it does not identify heterogeneous responses by subgroup in the sense of comparing specific groups (e.g., mortgagors vs. owners) as pseudo-panel approaches do. (e) The approach identifies the average linear response to a 25bp shock; nonlinear or asymmetric effects (large shocks, ZLB periods) are not modeled. (f) The simultaneous drop in earnings inequality and (weakly) rising consumption inequality cannot be fully reconciled without a complete model including capital income; the paper acknowledges this limitation explicitly.&lt;/p&gt;
&lt;h3 id="q11-how-do-the-quantitative-results-compare-to-the-ma-2021-hank-model-benchmark"&gt;Q11. How do the quantitative results compare to the Ma (2021) HANK model benchmark?&lt;/h3&gt;
&lt;p&gt;Ma (2021) incorporates an indivisible labor supply mechanism into a HANK model and shows that an expansionary monetary policy shock raises wages, inducing low-productivity workers to enter the labor market, raising earnings in the left tail. His calibration produces a Gini coefficient drop of approximately 0.001 for a comparable shock (scaled from his Figure 3: −0.4/(4×100) = −0.001 on a 0-to-1 scale for a 100bp shock). The Chang-Schorfheide empirical estimate is a drop of between 0.001 and 0.003 for a 25bp shock, which is broadly consistent with Ma&amp;rsquo;s model. The qualitative mechanism — earnings inequality reduction driven by low-productivity workers transitioning out of unemployment — is also consistent with the Chang-Kim (2006) heterogeneous-agent model with indivisible labor, which generates a negative correlation between idiosyncratic productivity and reservation wage.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-for-central-banks"&gt;Q12. What are the policy implications for central banks?&lt;/h3&gt;
&lt;p&gt;The paper provides semi-structural empirical evidence relevant for central banks concerned about distributional effects. The main conclusion is that for labor earnings inequality, the distributional effect of conventional monetary policy is well-summarized by the unemployment rate response: reducing unemployment compresses earnings inequality, and a central bank that targets unemployment de facto targets earnings inequality. The small, uncertain, and sometimes-positive effects on consumption and financial income inequality suggest that tracking these additional distributional statistics adds little actionable information beyond what standard macro aggregates already convey. The authors therefore conclude that there is an empirical case for central banks to continue focusing on macroeconomic aggregates. An important qualifier is that the financial income results are constrained by CEX top-coding, so the analysis cannot speak to very-high-income households&amp;rsquo; welfare.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Functional VAR (fVAR)&lt;/strong&gt;: A vector autoregression in which macroeconomic aggregates are stacked with the full cross-sectional log-probability density function of micro outcomes. The log-density is approximated by a finite-dimensional sieve (cubic spline basis), with sieve coefficients estimated period-by-period from cross-sectional data and then entered as observations in a linear VAR. This yields coherent IRFs for the entire distribution — percentiles, Gini, 90-10 ratio, etc. — from a single model, avoiding the quantile-crossing inconsistency of stacked-quantile approaches.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Employment channel (extensive margin)&lt;/strong&gt;: In this paper, the mechanism by which an expansionary monetary policy shock lowers earnings inequality: it reduces the unemployment rate, moving workers from a point mass of zero earnings into the positive-earnings distribution. The paper distinguishes this from the intensive margin (changes in wage rates conditional on employment), and finds empirically that the extensive margin dominates the inequality response of labor earnings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Informational shock (central bank information shock)&lt;/strong&gt;: As defined following Jarocinski-Karadi (2020): an unanticipated change in short-term interest rates that conveys the central bank&amp;rsquo;s private assessment of economic conditions. Identified by the simultaneous movement of interest rates and stock prices in the same direction, opposite to a conventional monetary policy shock. A negative informational shock (rates and equity prices both fall) signals that the central bank expects weaker output and prices than the public, and leads in this paper to rising earnings inequality via higher unemployment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Point mass at zero (earnings distribution)&lt;/strong&gt;: The concentration of probability mass at zero earnings, corresponding to the fraction of individuals in the labor force who are unemployed (the CPS-based unemployment rate). The total earnings density is modeled as a mixture of this point mass and a continuous density for positive earnings. The IRF for the point mass is the IRF for the unemployment rate; including it in inequality computations is necessary to capture the full distributional effect of employment transitions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Log probability density function (log-pdf) sieve representation&lt;/strong&gt;: The modeling device that represents each period&amp;rsquo;s cross-sectional distribution as the logarithm of a probability density, approximated by a finite linear combination of cubic spline basis functions (order K chosen by MDD). Working in log-pdf space avoids non-negativity and monotonicity constraints, enabling coherent linear propagation through the VAR law of motion; the density is recovered by exponential normalization in each period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marginal data density (MDD) model selection&lt;/strong&gt;: The Bayesian integrated likelihood used in this paper to jointly select the sieve approximation order K, lag length p, and Minnesota-type hyperparameters. The MDD balances in-sample fit (the log-spline likelihood) against a dimensionality penalty, thereby avoiding overfitting. A key result is that the preferred earnings fVAR uses K = 10 with a single lag, while the smoother consumption distribution is adequately captured with K = 6.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;κt (financial income point mass)&lt;/strong&gt;: The time-varying fraction of households in the CEX with financial income below a threshold x (set at the 10th percentile of pooled standardized financial income ≈ 0.0014 of the capital share of per-capita GDP). κt fluctuates between 0.65 and 0.82 over 1990–2016, meaning 65–82 percent of households have negligible financial income in a given quarter. The CEX data constraint — missing the top-10 percent of high-financial-income households — is the principal limitation on the financial income analysis.&lt;/p&gt;</description></item><item><title>Optimal Fiscal Policy in a Climate-Economy Model with Heterogeneous Households</title><link>https://macropaperwarehouse.com/papers/optimal-fiscal-policy-in-a-climate-economy-model-with-heterogeneous-households/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-fiscal-policy-in-a-climate-economy-model-with-heterogeneous-households/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether inequality and redistributive taxation should make climate policy more or less ambitious, and how optimal carbon taxes interact with optimal income taxes when households differ in productivity, wealth, and energy demand. The motivation is twofold: equity considerations belong at the center of normative climate analysis, and the distributional consequences of environmental policies are increasingly recognized as critical for their political feasibility — as illustrated by the Yellow Vests episode in France. The paper extends Barrage (2020)&amp;rsquo;s representative-agent dynamic climate-Ramsey model to a heterogeneous-agent setting, using the Werning (2007) technique to characterize the Ramsey optimum in terms of aggregate variables. The government maximizes utilitarian social welfare choosing linear taxes on labor income, capital income, energy, and pollution plus a uniform lump-sum transfer. The climate module is calibrated to DICE 2016 (Nordhaus, 2017). Household heterogeneity is calibrated to US data: ten productivity groups from SCF 2013 hourly wages ranging from $6.44 (bottom decile) to $101.35 (top decile), yielding a model consumption Gini of 0.33, very close to the empirical value of 0.32 (Heathcote et al., 2010). Tax rates are set at effective US rates from Trabandt and Uhlig (2012): capital income tax of 41.1% and labor income tax of 25.5%. The model period is five years beginning in 2015, and the discount factor follows DICE at beta = 1/(1.015) per year, with inverse IES sigma = 1.45. The main quantitative exercise compares optimal policy to a climate-skeptic planner who sets carbon taxes to zero. Key findings: (i) Tax distortions have a negligible effect on the optimal carbon tax in the heterogeneous-agent setting. The second-best carbon tax is initially only 0.5% below the social cost of carbon (SCC) and subsequently fluctuates within about 0.2% above or below it — in sharp contrast to Barrage (2020), who finds tax distortions reduce optimal carbon taxes by 8% in the representative-agent setting. The key mechanism is that, with heterogeneous agents, the government optimally levies distortionary taxes for redistributive purposes (not merely to finance public spending), so the marginal cost of public funds (MCF) averages to 1 over time and its temporal deviations are quantitatively trivial. (ii) Income inequality only slightly reduces the optimal carbon tax: residual consumption inequality after optimal income-tax redistribution lowers the SCC by 3.9% in the baseline. The mechanism is that inequality raises the average marginal utility of consumption (because the marginal utility function is convex), increasing the opportunity cost of abatement; this effect dominates when IES &amp;lt; 1 (sigma &amp;gt; 1 in the calibration). (iii) The optimal carbon tax path starts at $21.7/tCO2 in 2020 and reaches $229.2/tCO2 one century later — levels consistent with Barrage (2020) and Nordhaus (2017/2018) but insufficient to achieve the Paris +2°C target under baseline damages. (iv) Comparing optimal policy to the climate-skeptic baseline, the additional carbon tax revenue is split nearly equally: the present value of labor taxes falls by 0.7% of GDP, while transfers rise by 0.8% of GDP. This violates the weak double-dividend hypothesis, which prescribes using carbon tax revenue exclusively to cut distortionary taxes. (v) The optimal policy has progressive welfare effects in the 21st century, because increased tax progressivity benefits lower-income households. The average discounted welfare gain is 5.8% of consumption under baseline damages. In the long run, gains become regressive because richer households (with IES &amp;lt; 1) are willing to pay proportionally more in consumption to avoid temperature increases. By contrast, a representative-agent double-dividend policy — using all carbon revenue to cut labor taxes — is regressive from the outset, with low-income households bearing a net cost even in the short run. The 3.9% inequality effect on the SCC is robust to changes in fiscal pressure and damage calibration but is sensitive to sigma: with sigma = 2, inequality reduces optimal carbon taxes by 16.2% rather than 3.9%. Extensions with wealth heterogeneity, heterogeneous energy demand (calibrated to CEX), and heterogeneous environmental damage sensitivity confirm that the MCF remains negligible and the inequality effect on carbon taxes remains small in quantitative terms.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-theoretical-result-on-the-optimal-carbon-tax-and-why-does-it-differ-from-barrage-2020"&gt;Q1. What is the core theoretical result on the optimal carbon tax and why does it differ from Barrage (2020)?&lt;/h3&gt;
&lt;p&gt;The optimal carbon tax is approximately Pigouvian — set equal to the social cost of carbon — because the MCF averages to 1 over time with balanced-growth preferences when households are heterogeneous and the government can optimize a uniform lump-sum transfer. In Barrage (2020)&amp;rsquo;s representative-agent model, the government cannot choose the level of lump-sum taxes or transfers because there is no redistribution motive, so distortionary taxes are the only way to finance public spending and the MCF exceeds 1, reducing optimal carbon taxes by 8%. With heterogeneous agents, the government optimally provides lump-sum transfers for redistribution, so the constraint on transfers is barely binding and the MCF is close to 1 even when the ability to adjust transfers is removed.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-mechanism-by-which-inequality-affects-the-optimal-carbon-tax-and-what-is-the-sign"&gt;Q2. What is the mechanism by which inequality affects the optimal carbon tax, and what is the sign?&lt;/h3&gt;
&lt;p&gt;Inequality reduces the optimal carbon tax when IES &amp;lt; 1 (sigma &amp;gt; 1). The mechanism operates through the Pigouvian tax formula: pollution abatement reduces aggregate consumption, and the welfare cost of this reduction depends on the social marginal utility of consumption (Vc,t). With inequality, Vc,t is affected by two opposing forces. First, the average marginal utility of consumption is higher because of Jensen&amp;rsquo;s inequality (convex marginal utility function), increasing the opportunity cost of abatement and pushing the pollution tax down. Second, additional consumption goes disproportionately to richer households with lower marginal utilities, reducing Vc,t and pushing the tax up. When IES &amp;lt; 1, the first (higher average marginal utility) effect dominates, so inequality unambiguously reduces the SCC and hence the optimal pollution tax. When IES = 1, the two effects exactly cancel and inequality has no effect.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-mcf-and-why-does-it-average-to-1-in-the-heterogeneous-agent-setting"&gt;Q3. What is the MCF and why does it average to 1 in the heterogeneous-agent setting?&lt;/h3&gt;
&lt;p&gt;The MCF is defined as the ratio of the public (planner&amp;rsquo;s Lagrange multiplier on the resource constraint) to the private (aggregate welfare-weighted) marginal utility of consumption. It measures the social cost of transferring resources from the private to the public sector. The MCF averages to 1 because the first-order condition for the uniform lump-sum transfer implies that the sum of the Lagrange multipliers on agents&amp;rsquo; implementability constraints is zero. With balanced-growth preferences, this implies the welfare-weighted average MCF equals 1 from period 0. The temporal covariance between type-specific shadow costs (theta_i) and the type-specific implementability term (I_{c,i,t}) averages to zero over time, so while the MCF can deviate temporarily from 1, it is 1 on average.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-double-dividend-hypothesis-and-how-does-the-papers-optimal-policy-relate-to-it"&gt;Q4. What is the double-dividend hypothesis and how does the paper&amp;rsquo;s optimal policy relate to it?&lt;/h3&gt;
&lt;p&gt;The weak double-dividend hypothesis holds that it is optimal to use carbon tax revenue exclusively to reduce distortionary taxes, yielding both environmental and efficiency dividends. The paper shows this does not hold with heterogeneous agents: at the optimum, the welfare gain from a marginal reduction in tax distortions equals the welfare loss from increased inequality, so the government splits carbon revenue between cutting distortionary taxes and increasing redistribution. In the baseline quantification, the split is roughly equal: present-value labor taxes fall by 0.7% of GDP and lump-sum transfers rise by 0.8% of GDP. By contrast, following the double-dividend prescription — using all carbon revenue to reduce labor taxes without raising transfers — generates a strongly regressive policy in which low-income households bear net welfare costs even in the short run.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-calibration-strategy-and-how-does-the-model-match-us-inequality-data"&gt;Q5. What is the calibration strategy and how does the model match US inequality data?&lt;/h3&gt;
&lt;p&gt;The economic side is calibrated to the US, while the climate side uses DICE 2016. The discount factor follows DICE (beta = 1/(1.015) per year), and sigma = 1.45 (IES = 1/1.45). Household productivity is calibrated using SCF 2013 hourly wage deciles, yielding ten equal-sized groups with hourly wages from $6.44 (bottom) to $101.35 (top), normalized so that the productivity-weighted average is 1. Although productivity inequality is directly targeted rather than moments of the consumption distribution, the model correctly predicts the consumption Gini of 0.33, close to the empirical 0.32 (Heathcote et al., 2010). Capital and labor income tax rates are from Trabandt and Uhlig (2012): 41.1% and 25.5% respectively. Government debt-to-GDP is approximately 111% (average 2011-2015, IMF). The Frisch elasticity of labor supply is targeted at 0.75 (Chetty et al., 2011). Production in both sectors is Cobb-Douglas with energy share nu = 0.04 from Golosov et al. (2014).&lt;/p&gt;
&lt;h3 id="q6-what-happens-to-optimal-income-taxes-in-the-model"&gt;Q6. What happens to optimal income taxes in the model?&lt;/h3&gt;
&lt;p&gt;The optimal labor income tax roughly doubles from its calibrated level of 25% to about 50% in the first period and stabilizes there. Revenue from these taxes is rebated via the uniform lump-sum transfer, achieving most of the desired redistribution. Because optimal labor income taxes are approximately constant over time, the associated intertemporal distortions are small, and the optimal capital income tax converges to zero quickly after the second period. The mechanism is that, with access to lump-sum transfers, the only reason to tax capital income is to mitigate intertemporal distortions created by labor income taxation; when labor taxes are roughly constant, this motive is weak.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-sensitivity-analysis-reveal-about-the-robustness-of-the-39-inequality-effect"&gt;Q7. What does the sensitivity analysis reveal about the robustness of the 3.9% inequality effect?&lt;/h3&gt;
&lt;p&gt;The effect of inequality on optimal carbon taxes is robust along several dimensions but sensitive to sigma. Under the high-damage scenario (cubic rather than quadratic damage function, yielding an SCC about four times larger), the inequality effect falls to 2.6% rather than 3.9%, because higher carbon taxes reduce warming and thus the share of utility (rather than production) damages. The effect is roughly proportional to the degree of productivity inequality: half the inequality implies about half the effect on the carbon tax. The effect changes more than proportionally with sigma: with sigma = 2 (IES = 0.5), inequality reduces carbon taxes by 16.2%, versus 3.9% with the DICE value of sigma = 1.45. With sigma = 1, the effect is exactly zero. Government expenditure levels and fiscal pressure have negligible effects on the results. The share of damages entering utility directly matters: if only 10% of damages affect utility directly (versus the baseline 26%), the inequality effect falls to 1.8%; if 40% affect utility directly, it rises to 5.2%.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-role-of-initial-wealth-inequality"&gt;Q8. What is the role of initial wealth inequality?&lt;/h3&gt;
&lt;p&gt;Initial wealth inequality (studied in Section 6.1) creates an additional motive for deviating from Pigouvian taxation in period 0 only. Because the planner cannot use the period-0 capital tax to expropriate initial wealth (it is fixed at 41.1%), higher damages would reduce interest rates and thereby partially mitigate wealth inequality (a subtle indirect redistribution mechanism), calling for lower pollution taxes in period 0. Quantitatively, this produces a significant reduction in the initial-period optimal carbon tax. However, from period 1 onward, the optimal tax rules are unaffected by initial wealth heterogeneity, and the effects of MCF and income inequality remain very similar to the baseline. Welfare gains from carbon taxation in the wealth-heterogeneity extension are U-shaped with income but strictly increasing in initial wealth.&lt;/p&gt;
&lt;h3 id="q9-how-does-energy-demand-heterogeneity-stone-geary-extension-affect-the-results"&gt;Q9. How does energy-demand heterogeneity (Stone-Geary extension) affect the results?&lt;/h3&gt;
&lt;p&gt;The extension introduces a second dirty consumption good with Stone-Geary preferences, calibrated using CEX data to match the average energy expenditure share of 10.8% and the observed distribution of energy budget shares across and within income groups. Target emissions share from household energy consumption is 30%. The optimal pollution tax formula remains a modified Pigouvian rule (the MCF structure is unchanged), and the MCF effect remains negligible. The inequality effect on carbon taxes stays near 3.9%, rising marginally to 4.1% with identical energy necessity and 4.1% with heterogeneous energy necessity. Theoretically, the optimal excise tax on the energy good is zero when energy preferences are homogeneous; with heterogeneous necessity levels calibrated to the US, the optimal energy excise tax is quantitatively tiny: about -0.4% of energy prices (a small subsidy). The negative sign arises because within-income-group heterogeneity in energy needs means that energy-intensive households (who are valued more by the planner on average) can be partially targeted via a subsidy. Under the double-dividend scenario with energy inequality, regressive effects are magnified: the poorest, most energy-intensive households actually lose in welfare terms even accounting for long-run climate mitigation benefits.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-paper-establish-theoretically-about-heterogeneous-environmental-damages"&gt;Q10. What does the paper establish theoretically about heterogeneous environmental damages?&lt;/h3&gt;
&lt;p&gt;Proposition 6 (Section 6.3) shows that with additively separable environmental utility and a utilitarian planner, heterogeneous marginal utility damages from pollution have no effect on the optimal pollution tax: they enter the welfare criterion symmetrically and cancel in the aggregate. The pollution tax increases relative to the utilitarian benchmark only if the planner&amp;rsquo;s welfare weights are positively correlated with marginal utility damages — that is, if the planner cares relatively more about the households that are more exposed. A Rawlsian planner would set a higher pollution tax if and only if the least-well-off household is also more sensitive to environmental degradation.&lt;/p&gt;
&lt;h3 id="q11-what-are-third-best-policy-results-when-either-income-tax-is-fixed"&gt;Q11. What are third-best policy results when either income tax is fixed?&lt;/h3&gt;
&lt;p&gt;The paper analyzes policies where either the labor or capital income tax is fixed at its current calibrated level (studied in Appendix E, with results referenced in the main text). These constraints introduce an additional fiscal interaction effect on the optimal carbon tax — the carbon tax is pushed below its second-best Pigouvian level when the fixed tax is set at a sub-optimally low level, and above it when the fixed tax is sub-optimally high. The roles of the MCF and income inequality remain similar to the second-best baseline under these third-best constraints.&lt;/p&gt;
&lt;h3 id="q12-how-does-the-paper-relate-to-and-differ-from-the-double-dividend-and-pollution-taxation-literatures"&gt;Q12. How does the paper relate to and differ from the double-dividend and pollution taxation literatures?&lt;/h3&gt;
&lt;p&gt;The paper builds on three earlier pillars. First, Pigou (1920) established first-best Pigouvian taxation. Second, a large literature (Sandmo, 1975; Bovenberg and de Mooij, 1994; Bovenberg and Goulder, 1996) showed that in representative-agent second-best settings the MCF exceeds 1 and optimal pollution taxes fall below the Pigouvian level. Barrage (2020) is the closest dynamic general-equilibrium predecessor, finding the 8% reduction from tax distortions. Third, Jacobs and de Mooij (2015) and Jacobs and van der Ploeg (2019) showed in static models with heterogeneous agents and a uniform lump-sum transfer that the MCF equals 1. This paper extends this insight to a fully dynamic climate-economy framework with general equilibrium and a rich model of household heterogeneity. The key innovation relative to Barrage (2020) is agent heterogeneity, which both provides microfoundations for distortionary taxation and significantly changes the quantitative implications for optimal carbon taxes. Relative to Jacobs and de Mooij (2015), the contribution is the dynamic setting, the linkage to the DICE climate module, and the full quantitative characterization including distributional welfare analysis and multiple sources of heterogeneity.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q13. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is that a carbon tax should be set approximately equal to the SCC (Pigouvian level) and the associated revenue should be split roughly equally between increasing lump-sum transfers and reducing distortionary labor taxes — rather than following the double-dividend prescription of using all revenue to reduce distortionary taxes. This combination is both more efficient (the MCF argument) and more equitable (progressive in the short run). The scope conditions are: (a) the result applies under a utilitarian welfare criterion with linear income taxes and a uniform lump-sum transfer; (b) it requires that the government can optimize the level of lump-sum transfers for redistribution; (c) the approximately Pigouvian result is quantitatively robust to alternative damage functions, fiscal pressure, and energy demand heterogeneity, but the degree to which inequality lowers the carbon tax depends sensitively on the IES/inequality aversion parameter sigma; (d) the calibration is designed to capture US conditions assuming that the US internalizes the full global impact of its emissions (strategic considerations are abstracted away); (e) heterogeneous environmental damage sensitivity does not affect the utilitarian optimum, but would increase the optimal carbon tax under a more inequality-averse social planner.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Marginal Cost of Public Funds (MCF)&lt;/strong&gt;: The ratio of the public (planner&amp;rsquo;s shadow price on the resource constraint) to the private (aggregate welfare-weighted) marginal utility of consumption. In this paper, it captures the divergence between second-best and first-best pollution taxes due to fiscal distortions. With heterogeneous agents and an optimized uniform lump-sum transfer, the MCF averages to 1 over time under balanced-growth preferences, implying that tax distortions do not systematically push the carbon tax below the Pigouvian level — unlike in the representative-agent setting where the MCF exceeds 1.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pigouvian tax (second-best)&lt;/strong&gt;: In this paper&amp;rsquo;s context, the Pigouvian tax refers to the pollution tax equal to the social cost of pollution (the discounted present value of marginal production and utility damages), evaluated at the second-best allocation rather than the first-best. When the MCF equals 1 (as it approximately does in the heterogeneous-agent setting), the second-best optimal pollution tax is equal to this second-best Pigouvian level, which may itself differ from the first-best Pigouvian level due to residual consumption inequality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Cost of Carbon (SCC)&lt;/strong&gt;: The present discounted value of marginal climate damages (both production and utility losses) from emitting one additional ton of CO2, converted into consumption units using the social marginal utility of consumption. In the paper, the SCC corresponds to the case where the MCF is set to 1 in every period, and it is affected by consumption inequality through its effect on the social marginal utility of consumption. With sigma &amp;gt; 1, residual inequality raises the opportunity cost of abatement, reducing the SCC by 3.9% in the baseline calibration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double-dividend hypothesis (weak)&lt;/strong&gt;: The claim that it is optimal to use the entire proceeds of a carbon tax to reduce existing distortionary taxes, yielding both an environmental dividend (less pollution) and an efficiency dividend (lower tax distortions). The paper shows this does not hold with heterogeneous agents: because distortionary taxes serve a redistributive purpose, reducing them at the margin has a welfare cost (increased inequality), so the planner optimally splits revenue between tax reduction and increased transfers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ramsey problem (climate-economy)&lt;/strong&gt;: The government&amp;rsquo;s optimization problem in this paper: maximizing utilitarian social welfare over an infinite horizon by choosing paths for linear taxes on labor income, capital income, energy, and pollution, plus a uniform lump-sum transfer, subject to households&amp;rsquo; optimality conditions (implementability constraints), resource constraints, climate dynamics from DICE, and abatement technology constraints. The approach extends Werning (2007) to a dynamic climate-economy context.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implementability condition&lt;/strong&gt;: The constraint in the Ramsey problem that captures each household&amp;rsquo;s lifetime budget constraint in terms of aggregate variables and market weights. It requires that the present value of a household&amp;rsquo;s consumption minus labor income equals its initial assets plus its share of the present value of lump-sum transfers, evaluated using the social marginal utilities implied by the planner&amp;rsquo;s choice of taxes. The shadow cost of this constraint for each household type (theta_i) determines the MCF through its covariance with a fiscal externality term.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Residual inequality&lt;/strong&gt;: The level of inequality that remains after the planner has optimally set all income taxes and the lump-sum transfer — i.e., the inequality that cannot be eliminated because individualized lump-sum transfers are not feasible and only linear instruments are available. In the paper, it is this residual inequality (not total inequality) that affects the optimal carbon tax: the carbon tax responds to the inequality that income-tax policy cannot address, not to the underlying productivity or wealth dispersion per se.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Balanced-growth preferences&lt;/strong&gt;: A preference specification of the form u(c, h, Z) = [c(1 - varsigma*h)^gamma]^(1-sigma)/(1-sigma) + u_hat(Z), with 1/sigma the intertemporal elasticity of substitution. This specification ensures that the economy admits a balanced growth path and plays a key role in the paper&amp;rsquo;s theoretical results: under balanced-growth preferences, the welfare-weighted average MCF equals 1 from period 0, and when IES = 1 (sigma = 1) the MCF is exactly 1 in every period.&lt;/p&gt;</description></item><item><title>Population and Welfare: Measuring Growth when Life is Worth Living</title><link>https://macropaperwarehouse.com/papers/population-and-welfare-measuring-growth-when-life-is-worth-living/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/population-and-welfare-measuring-growth-when-life-is-worth-living/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;The paper asks how much economic progress looks different when one applies a total utilitarian welfare criterion — counting every person&amp;rsquo;s flow utility — rather than the standard per-capita consumption measure. The motivation is both philosophical and practical: philosophers have long debated whether the number of people matters for social welfare alongside average living standards, yet growth economists have almost exclusively used the per-capita approach. The authors do not adjudicate the debate; they quantify its stakes across a broad cross-country sample.&lt;/p&gt;
&lt;p&gt;The framework is parsimonious. Under total utilitarianism, flow social welfare is W = N·u(c). Consumption-equivalent (CE) welfare growth is gλ = v(c)·gN + gc, where gN is population growth, gc is per-capita consumption growth, and v(c) = u(c)/[u&amp;rsquo;(c)·c] is the value of a year of life in units of per-capita consumption. Diminishing marginal utility guarantees v(c) &amp;gt; 1: each percentage point of population growth is worth more than a percentage point of per-capita consumption growth. The baseline utility function is u(c) = ū + log(c). The key parameter ū is calibrated to the U.S. Environmental Protection Agency&amp;rsquo;s Value of Statistical Life (VSL) of $7.4 million (2006 prices): dividing by remaining life expectancy (~40 years) and U.S. per-capita consumption of $38,000 gives v(c_US,2006) ≈ 4.87. Normalizing U.S. 2006 consumption to 1 sets ū = 4.87. Under log utility, v(c) = ū + log(c) rises with living standards: it averaged roughly 2 in 1820 for the U.S. and nearly 5 by 2019, and ranges from about 2 (Ethiopia) to 5 (U.S.) across countries in 2019. The world-sample average of v(c) over 1960–2019 is 2.7.&lt;/p&gt;
&lt;p&gt;Applying the formula to Penn World Table 10.0 data for 101 countries over 1960–2019 yields the following main findings. CE welfare growth averages 6.2% per year versus 2.1% per year for per-capita consumption growth; at 2.1% growth per-capita consumption doubles every 33 years, but under the CE measure social welfare doubles every 12 years. Population growth (averaging 1.8% per year) accounts for 66% of CE welfare growth unweighted across countries, and 51% weighting by country population (which gives China a large weight). For the United States specifically, CE welfare growth averages 6.5% per year versus 2.2% for per-capita consumption growth. Country rankings shift dramatically. Mexico rises from the 35th to the 88th percentile (CE welfare growth: 8.6% per year; population contribution: 79%). South Africa and Kenya similarly move up sharply. Germany falls to the 11th percentile, Japan to the 32nd, and China to the 44th — all below the United States. The cross-country correlation between CE welfare growth and per-capita consumption growth is 0.51; with population growth, 0.29. Over the very long run (1500–2018, Maddison data), per-capita consumption rose 20-fold (0.6% per year) while CE welfare rose 3,700-fold (1.6% per year) due to population growing at 0.5% per year scaled by v(c).&lt;/p&gt;
&lt;p&gt;Robustness checks confirm the core result. Halving the baseline VSL (setting ū = 2.4) still leaves population contributing 38% of CE welfare growth on average. Incorporating within-country consumption inequality under a log-normal distribution lowers CE welfare growth by an average of just 10 basis points (from 6.1% to 6.0% for 1980–2007). Attributing migrants to source rather than destination countries produces a correlation of 0.92 between adjusted and baseline CE welfare growth rates. Decomposing population growth, roughly three-quarters of actual population growth in a 24-country subsample reflected increases in the number of lives lived (i.e., births), not longevity extension — so the welfare contribution of births exceeds that of rising longevity.&lt;/p&gt;
&lt;p&gt;An extended model adds leisure, parental altruism toward children&amp;rsquo;s consumption and human capital, and endogenous fertility. Using time-use data from six countries (U.S. 2003–2019; Netherlands 1975–2006; Japan 1991–2016; South Korea 1999–2019; Mexico 2006–2019; South Africa 2000–2010), the extension modestly reduces the population share of CE welfare growth in most countries. The main reason is that parental altruism &amp;ldquo;double-counts&amp;rdquo; children&amp;rsquo;s consumption in the social welfare function, making consumption growth relatively more valuable and thus scaling down the weight on population growth. Rising quality of children (human capital) roughly offsets falling fertility in most countries, leaving net CE welfare growth little changed. Mexico is the sharpest exception: under extended preferences, CE welfare growth falls from 6.5% to 3.3% because of sharply declining leisure and little offsetting gain in children&amp;rsquo;s quality. Japan and South Korea also see smaller population shares under the extended model. The qualitative conclusion — that population growth is a major contributor to CE welfare growth — survives across all specifications.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-identification-strategy-and-what-does-it-rely-on"&gt;Q1. What is the paper&amp;rsquo;s identification strategy and what does it rely on?&lt;/h3&gt;
&lt;p&gt;This is a welfare accounting exercise rather than a causal identification exercise. There is no identification problem in the traditional econometric sense: the authors are computing a welfare index given a social welfare function and observed data on population and consumption. The two key inputs are (1) data on population and consumption per capita from the Penn World Table 10.0 for 101 countries over 1960–2019, and (2) a calibrated value of the parameter ū, which is the value of a year of life measured in units of per-capita consumption. The calibration of ū is anchored to external VSL estimates (EPA&amp;rsquo;s $7.4 million in 2006 prices), divided by life expectancy and per-capita consumption. The paper is explicit that it cannot make causal policy recommendations because it says nothing about the production side of the economy or externalities (pollution, ideas, human capital spillovers).&lt;/p&gt;
&lt;h3 id="q2-what-is-vc-and-why-does-it-matter-so-much-for-the-results"&gt;Q2. What is v(c) and why does it matter so much for the results?&lt;/h3&gt;
&lt;p&gt;v(c) = u(c)/[u&amp;rsquo;(c)·c] is the value of a year of life measured in consumption-equivalent units — specifically, how many years&amp;rsquo; worth of per-capita consumption an individual would require as compensation for losing one year of life. Under log utility u(c) = ū + log(c), v(c) = ū + log(c), so it rises with the log of consumption. The key implication is that each percentage point of population growth is worth v(c) percentage points of per-capita consumption growth. Since v(c) empirically ranges from about 2 (Ethiopia) to 5 (rich countries), and averages 2.7 over 1960–2019 across 101 countries, population growth receives a substantial weight in the CE welfare measure. Without this amplification (i.e., if v = 1 so CE welfare equals aggregate consumption growth), population would still account for 36% of all growth.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-distinguish-its-approach-from-simply-using-aggregate-total-consumption-growth"&gt;Q3. How does the paper distinguish its approach from simply using aggregate (total) consumption growth?&lt;/h3&gt;
&lt;p&gt;Using aggregate consumption growth is equivalent to setting v(c) = 1 in the CE welfare formula — that is, weighting population growth and consumption growth equally. The paper shows that, under a total utilitarian welfare function with diminishing marginal utility, the correct weight on population growth is v(c) &amp;gt; 1, not 1. So aggregate consumption growth systematically understates the contribution of population growth to welfare: in a country with average v(c) = 2.7, a percentage point of population growth should receive 2.7 times the weight of a percentage point of consumption growth, not equal weight.&lt;/p&gt;
&lt;h3 id="q4-what-threats-to-the-baseline-calibration-of-vc-does-the-paper-address"&gt;Q4. What threats to the baseline calibration of v(c) does the paper address?&lt;/h3&gt;
&lt;p&gt;The paper addresses four main threats. First, VSL uncertainty: it considers halving and raising the baseline VSL by 50%, yielding ū = 2.4 and ū = 7.3 respectively. Population&amp;rsquo;s share of CE welfare growth remains 38% even under the low VSL. Second, functional form: it considers CRRA utility with risk-aversion γ = 2 rather than log (γ = 1), which lowers the population share to 40% (from 53% baseline, population-weighted). Third, whether v(c) should be constant rather than income-varying: rows 7–9 of Table 3 test constant v = 4.87, v = 2.7, and v = 1. Even v = 1 (aggregate consumption growth) gives population a 36% share. Fourth, whether the marginal VSL used to calibrate the model overstates the average value of a birth (since a birth produces a new life from the start, not an added year for a middle-aged person). The paper acknowledges this concern but treats the calibration as a natural baseline and explores lower ū as a robustness check.&lt;/p&gt;
&lt;h3 id="q5-how-does-within-country-inequality-affect-the-results"&gt;Q5. How does within-country inequality affect the results?&lt;/h3&gt;
&lt;p&gt;Under log utility and a log-normal distribution of individual consumption, CE welfare growth becomes gλ = [ū + log(c_t) - (1/2)σ²_t]·gN + gc - σ²_t·gσ, where σ² is the cross-sectional variance of log consumption. Inequality enters in two ways: it reduces the weight on population growth (because average utility is lower than utility of average consumption under concavity), and increases in inequality directly reduce CE welfare growth. Implementing this for 90 countries over 1980–2007, the mean adjustment is -10 basis points (6.1% to 6.0%), with a mean absolute deviation of 18 basis points. The adjustment is sizable for South Africa (−0.83 pp, due to very high inequality relative to U.S. 2006 baseline) and small or positive for Brazil (falling inequality over the period) and Ethiopia.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-treat-migration-and-does-it-matter"&gt;Q6. How does the paper treat migration, and does it matter?&lt;/h3&gt;
&lt;p&gt;The baseline credits population growth to the country of residence. The migration-adjusted measure reassigns migrants to their country of birth: it adds the flow utility of out-migrants (at destination-country consumption levels) and subtracts the flow utility of in-migrants (at destination-country consumption levels) from each country&amp;rsquo;s welfare. Using the World Bank Global Bilateral Migration Database for 81 countries over 1960–2000, migration-adjusted and baseline CE welfare growth rates have a correlation of 0.92. The adjustment matters most for specific countries — it raises welfare growth for net out-migrant countries like Mexico and the Philippines (since their emigrants consume more abroad) and lowers it for net in-migrant countries — but it does not alter the broad conclusion that population growth matters greatly.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-decomposition-of-population-growth-into-fertility-and-longevity-effects"&gt;Q7. What is the decomposition of population growth into fertility and longevity effects?&lt;/h3&gt;
&lt;p&gt;For a 24-country subsample (from the Human Mortality Database combined with World Bank migration data), the authors compute counterfactual population growth holding age-specific death rates constant at their initial-period values. Population-weighted, actual annual population growth is 0.72% versus a counterfactual of 0.53% with fixed longevity. So roughly three-quarters of population growth (and therefore three-quarters of the CE welfare contribution of population growth) reflected an increase in the number of lives lived (births minus deaths under fixed mortality), not gains in longevity. Italy and Japan are outliers: falling death rates (i.e., longevity gains) account for about three-quarters of their population growth. For context, Jones and Klenow (2016) attribute ~1% per year of CE welfare growth to rising longevity for 128 countries over 1980–2007; the total population growth contribution here (~3% per year, population-weighted from Table 1) substantially exceeds the longevity-only benchmark.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-extended-model-with-parental-preferences-add-and-what-are-its-main-results"&gt;Q8. What does the extended model with parental preferences add, and what are its main results?&lt;/h3&gt;
&lt;p&gt;The extended model incorporates adult leisure, parental altruism toward children&amp;rsquo;s consumption and human capital, endogenous fertility, and children&amp;rsquo;s utility as separate welfare contributors. Social welfare is W = N_p·u(c_p, l, c_k, h_k, b) + N_k·ũ(c_k), where b is fertility per adult, l is adult leisure, and h_k is children&amp;rsquo;s human capital. CE welfare growth is computed using first-order conditions from parents&amp;rsquo; utility maximization — specifically, the MRS between leisure/fertility/human capital and consumption can be measured from time-use data, which provides the welfare weights on each term. Key parameters: parental altruism weight α = 2/3 (calibrated to USDA household spending data), diminishing-returns-to-fertility parameter θ = 0.8, and children&amp;rsquo;s human capital elasticity η = 0.21 (from Mincer estimates in Lee, Roys, and Seshadri 2024). Main results: (1) Population growth remains an important contributor to CE welfare growth in most countries. (2) The population share falls somewhat because parental altruism double-counts children&amp;rsquo;s consumption, raising the relative weight on consumption growth. (3) Rising children&amp;rsquo;s quality (human capital, measured via real wage growth) roughly offsets falling fertility in most countries. (4) Mexico is the main exception: CE welfare growth drops from 6.5% to 3.3% due to falling leisure and little offset from rising children&amp;rsquo;s quality. (5) Japan&amp;rsquo;s population share falls further, turning slightly negative in some specifications.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-philosophical-foundation-and-what-is-the-repugnant-conclusion-objection"&gt;Q9. What is the philosophical foundation and what is the &amp;lsquo;repugnant conclusion&amp;rsquo; objection?&lt;/h3&gt;
&lt;p&gt;The total utilitarian social welfare function W = N·u(c) follows from three axioms: same-number Pareto (welfare ordering respects Pareto improvements for fixed populations), non-anti-egalitarianism (society does not prefer inequality), and mere addition (adding a person who values living, holding others&amp;rsquo; utilities constant, does not reduce welfare). These axioms, as surveyed by Kuruc, Budolfson, and Spears (2022), together imply total utilitarianism and rule out diminishing-returns-to-population approaches (e.g., W = N^α·u(c) for α &amp;lt; 1). The repugnant conclusion (Parfit 1984) holds that total utilitarianism could justify very large populations of people whose lives are barely worth living. The authors respond that their calculations are local — reflecting only actual births and deaths over 1960–2019 — not arbitrary expansions. They also note that 29 philosophers and economists (Zuber et al. 2021) have argued the repugnant conclusion is not a reason to reject totalism. The per-capita approach has its own problems: it implies one should remove people whose utility is valuable but below average (and implies the &amp;lsquo;sadistic conclusion&amp;rsquo; under certain conditions).&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-jones-and-klenow-2016"&gt;Q10. How does this paper relate to Jones and Klenow (2016)?&lt;/h3&gt;
&lt;p&gt;Jones and Klenow (2016) is the closest predecessor. That paper computes CE welfare measures incorporating consumption, leisure, life expectancy, and inequality, but in a per-capita framework — it measures individual living standards, not aggregate social welfare. The key difference here is moving from per-capita utility to total utilitarian welfare by multiplying individual utility by population, which introduces the v(c)·gN term. The current paper&amp;rsquo;s baseline is also simpler (consumption only) with an extended version that adds leisure and parental preferences. Jones and Klenow attribute ~1% per year of CE welfare growth to rising longevity for 128 countries over 1980–2007; the present paper shows total population growth (birth + longevity channels combined) contributes ~3% per year (population-weighted), substantially more.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper explicitly states it cannot make policy recommendations because it says nothing about the production side of the economy or about externalities (pollution, ideas externalities, human capital spillovers). Whether fertility rates are &amp;rsquo;too low&amp;rsquo; or the demographic transition raised or reduced social welfare requires estimating these externalities, which is beyond the paper&amp;rsquo;s scope. The paper is a measurement exercise, not an optimal policy analysis. Nonetheless, the results have implications for policy questions that depend on which welfare criterion is adopted: optimal fertility policy, the welfare cost of HIV/AIDS or other mortality shocks, the assessment of China&amp;rsquo;s One Child Policy, the welfare calculus of climate change mitigation, and the social returns to nonrival knowledge (which benefit a larger future population under totalism). The scope condition throughout is that the paper evaluates actual births and deaths over a historical period; the results do not directly speak to the desirability of population expansion beyond what occurred.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-main-robustness-checks-run-and-what-do-they-show"&gt;Q12. What are the main robustness checks run and what do they show?&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;VSL calibration: halving (ū = 2.4) or raising by 50% (ū = 7.3) the baseline VSL — population share falls to 38% or rises to higher levels, but population remains important in all cases. 2. CRRA utility with γ = 2 (more concave): population share falls to 40% population-weighted (from 53%). 3. Constant v(c): results with v = 4.87 (U.S. 2006 level), v = 2.7 (world average), and v = 1 (aggregate consumption growth) all confirm that population growth matters, with v = 1 still giving a 36% population share. 4. Inequality: mean absolute adjustment of 18 basis points; largest adjustment for South Africa (−0.83 pp). 5. Migration: correlation 0.92 between adjusted and baseline. 6. Birth vs. longevity decomposition: ~75% of population growth (population-weighted) is from net new lives, not longevity. 7. Extended preferences (time-use data): qualitative results survive; population share falls modestly except for Japan, South Korea, and Mexico.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="q13-what-heterogeneity-across-countries-and-time-periods-is-documented"&gt;Q13. What heterogeneity across countries and time periods is documented?&lt;/h3&gt;
&lt;p&gt;Cross-country: CE welfare growth ranges from just above 2% per year for the slowest-growing countries to more than 10% per year for the fastest. The correlation between CE welfare growth and per-capita consumption growth is 0.51; with population growth, 0.29. The value of v(c) ranges from about 2 (Ethiopia) to 5 (U.S.) in 2019, tracking consumption levels. Countries with high population growth (Mexico, Brazil, South Africa, Kenya, Sub-Saharan Africa more broadly) move up sharply in the growth rankings; countries with slow population growth (Germany, Japan, China) fall sharply. Within time: Japan shows CE welfare growth falling from 9.7% per year in the 1960s to −0.3% in the 2010s as both consumption growth and population growth slowed and then turned negative. China&amp;rsquo;s CE welfare growth fell more modestly from a 7.0% peak in the 1990s to 5.7% in the 2010s because rising v(c) partly offset slower population growth. Sub-Saharan Africa maintained stable population growth (~2.5% per year across all decades) and saw rising consumption in the 2000s and 2010s, producing CE welfare growth above 8% in the 2010s. Extended-model results (six-country sample): Mexico is a major outlier with falling leisure driving CE welfare growth down to 3.3% (from 6.5% baseline); Japan and South Korea have very small or near-zero population shares under extended preferences.&lt;/p&gt;
&lt;h3 id="q14-what-does-the-very-long-run-historical-exercise-show"&gt;Q14. What does the very long-run historical exercise show?&lt;/h3&gt;
&lt;p&gt;Using Maddison Project data (de Pleijt and van Zanden 2020) from 1500 to 2018 for the world as a whole, per-capita consumption rose by a factor of 20 (0.6% per year) and aggregate consumption rose by a factor of 163 (1.1% per year). CE welfare rose by a factor of 3,700 — a 1.6% per year average annual growth rate. The power of compounding over 500 years causes a difference of only 1 percentage point per year between CE welfare growth (1.6%) and per-capita consumption growth (0.6%) to produce a ratio of 185:1 in cumulative outcomes (3,700-fold versus 20-fold). Population growth accounts for 61% of CE welfare growth over this very long run.&lt;/p&gt;
&lt;h3 id="q15-what-data-sources-are-used-and-what-are-the-sample-restrictions"&gt;Q15. What data sources are used and what are the sample restrictions?&lt;/h3&gt;
&lt;p&gt;Baseline: Penn World Table 10.0 (Feenstra, Inklaar, and Timmer 2015) for 101 countries over 1960–2019 (starting from 111 countries and dropping those flagged as outliers in any year). Consumption is private plus government consumption. Inequality data: Jones and Klenow (2016), 90 countries, 1980–2007. Migration data: World Bank Global Bilateral Migration Database (1960, 1970, 1980, 1990, 2000), 81 countries. Birth/death decomposition: Human Mortality Database combined with World Bank migration data, 24 countries. Long run: Maddison Project (de Pleijt and van Zanden 2020), 1500–2018, with consumption proxied as 0.8 times per-capita GDP. Extended model: time-use surveys for U.S. (2003–2019), Netherlands (1975–2006), Japan (1991–2016), South Korea (1999–2019), Mexico (2006–2019), South Africa (2000–2010); Penn World Table for population, consumption, and hours worked; World Bank for number of children (0–19 years); USDA spending-on-children data (Lino 2011) for parental altruism calibration; Lee, Roys, and Seshadri (2024) Mincer estimates for human capital elasticity.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Consumption-Equivalent (CE) Welfare Growth&lt;/strong&gt;: The rate at which per-capita consumption would need to grow, holding population constant, to produce the same increase in total utilitarian social welfare as the observed combination of population growth and per-capita consumption growth. Formally gλ = v(c)·gN + gc. It is analogous to GDP in being a flow measure (not a present-discounted sum across generations).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Value of a Year of Life, v(c)&lt;/strong&gt;: The ratio u(c)/[u&amp;rsquo;(c)·c], equal to individual utility divided by the marginal utility of consumption times consumption. It converts the value of being alive for one year into consumption-equivalent units. Under log utility u(c) = ū + log(c), v(c) = ū + log(c), so it rises with living standards. It is calibrated from Value of Statistical Life estimates: v(c_US,2006) ≈ 4.87, meaning a year of American life in 2006 was worth approximately 4.87 years&amp;rsquo; worth of per-capita consumption.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Total Utilitarian Social Welfare Function&lt;/strong&gt;: W = N·u(c): social welfare is the sum of all individuals&amp;rsquo; flow utilities. It treats every person&amp;rsquo;s utility symmetrically and linearly in population, so adding a person who values living always increases welfare. This contrasts with the per-capita (average utilitarian) approach (which implicitly sets the welfare weight on population to zero) and intermediate approaches that weight population with diminishing returns (W = N^α·u(c), α &amp;lt; 1).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mere Addition (Axiom)&lt;/strong&gt;: One of three axioms (with same-number Pareto and non-anti-egalitarianism) whose conjunction implies the total utilitarian social welfare function for variable populations. It states that, holding the utilities of existing persons constant, adding a new person who values living does not reduce social welfare. The axiom directly rules out the average (per-capita) utilitarian criterion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Repugnant Conclusion&lt;/strong&gt;: Parfit&amp;rsquo;s (1984) critique of total utilitarianism: maximizing the sum of utilities could in principle justify an extremely large population of people whose lives are just barely worth living (positive but tiny utility), since total utility could exceed that of a smaller population with high per-capita welfare. The paper responds that its welfare calculations are local (reflecting actual historical births and deaths), not global maximization exercises, and cites the Zuber et al. (2021) consensus that this is not a decisive objection.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parental Altruism Weight (α, θ)&lt;/strong&gt;: Parameters governing how parents value children&amp;rsquo;s consumption relative to their own in the extended model. Under Assumption 1, u includes the term αb^θ·log(c_k): α governs the overall weight on children&amp;rsquo;s consumption and θ governs diminishing returns as the number of children rises. Calibrated to α = 2/3 (from USDA household spending ratios with two children) and θ = 0.8 (from cross-family variation in per-child spending). Parental altruism causes double-counting of children&amp;rsquo;s consumption in the social welfare function, upweighting consumption growth and reducing the relative contribution of population growth to CE welfare.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double-Counting of Children&amp;rsquo;s Consumption&lt;/strong&gt;: When parents are altruistic, their utility depends on children&amp;rsquo;s consumption as well as their own; and children receive direct utility from their consumption too. In the CE welfare growth formula, this means a rise in children&amp;rsquo;s consumption raises welfare through two channels simultaneously (parental and child utility), so each unit of consumption growth is &amp;lsquo;worth more&amp;rsquo; relative to population growth. This is why the extended model&amp;rsquo;s population term is smaller than the baseline&amp;rsquo;s: consumption growth is valued more heavily under parental altruism, scaling down the consumption-equivalent weight on population growth.&lt;/p&gt;</description></item><item><title>Populism and the Skill-Content of Globalization</title><link>https://macropaperwarehouse.com/papers/populism-and-the-skill-content-of-globalization/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/populism-and-the-skill-content-of-globalization/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how the skill structure of globalization shocks — rather than globalization per se — drives the long-run evolution of populism across countries, making a unified empirical case that what gets imported or who immigrates matters as much as how much.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; The literature has documented that trade exposure and immigration fuel populist voting, but prior work has studied these channels separately, used narrow time windows, and relied on binary party classifications that cannot capture shifts in populism across the full party landscape. Rodrik&amp;rsquo;s (2018) widely-cited hypothesis holds that trade shocks drive left-wing populism (as in Latin America) and immigration drives right-wing populism (as in Europe). The authors examine whether this hypothesis survives when skill content is explicitly disaggregated and both channels are studied jointly in a unified long-panel setting.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data, sample, and empirical strategy.&lt;/strong&gt; The authors construct a new continuous, time-varying populism score for 3,860 party-election pairs covering 1,206 unique parties across 628 national elections in 55 countries from 1960 to 2018. The score is built from the Manifesto Project Database (MPD) using two dimensions identified in the political-science literature: an anti-establishment stance (AES) and a commitment-to-protect stance (CTP). A two-stage polychoric PCA extracts synthetic indices for each dimension and then combines them into a single populism score. The paper defines populist parties as those scoring more than one standard deviation above the mean (threshold validated by comparison with four external databases — Van Kessel, Swank, PopuList, GPop 1 — with ratios of accurate forecasts ranging from 80 to 91 percent). Two dependent variables are studied: (i) the volume margin of populism, the vote share of classified populist parties, estimated with PPML given many zero observations (about 60 percent of the full sample); and (ii) the mean margin of populism, the vote-weighted average populism score of all parties, estimated with OLS. Globalization regressors are skill-specific: imports of low-skill and high-skill labor-intensive goods (as shares of GDP, sourced from Feenstra et al. 2005 and UN Comtrade) and immigration inflows of low-skill and high-skill workers (from Abel 2018, skill-level imputed from dyadic migrant-stock selection ratios). To address reverse causality — populist governments restrict trade and immigration, biasing OLS downward — the authors implement a gravity-based IV strategy: a zero-stage PPML regression predicts bilateral flows using time-invariant dyadic fixed effects interacted with a post-1990 dummy and origin-country-year fixed effects, then aggregates to the destination level; these predicted flows serve as instruments. For the volume margin, a reduced-form IV approach replaces actual with predicted flows (to avoid the incidental-parameter problem in PPML with fixed effects). For the mean margin, standard 2SLS is used; the Kleibergen-Paap F-statistic is around 10–12, reasonable given four instruments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main quantitative findings.&lt;/strong&gt; (All claims below are with country and year fixed effects throughout; IV results reinforce baseline OLS/PPML results.)&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Low-skill labor-intensive imports raise total and right-wing populism along both the volume margin and the mean margin. In the OLS mean-margin specification the coefficient on low-skill imports is approximately 4, implying a 1 percentage-point increase in the import-to-GDP ratio for low-skill goods is associated with a 0.04 increase in the mean margin of populism (scaled in standard deviations of the populism score). The 2SLS coefficient on the total mean margin is approximately 5.0 (significant at 5%), and on the right-wing mean margin approximately 4.1 (significant at 5%). For the volume margin, the reduced-form IV coefficient on low-skill imports is 0.91 (significant at 10%) for total and 1.82 (significant at 5%) for right-wing populism. These effects are larger by a factor of approximately 1.3 when IV is used relative to OLS/PPML, consistent with downward bias from reverse causality. Low-skill imports do not significantly affect left-wing populism in baseline estimates; a left-wing response cannot be ruled out during severe crises, when shocks are persistent, or among EU countries specifically.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;High-skill labor-intensive imports reduce the volume of populism, especially right-wing populism. In the reduced-form IV specification the coefficient on high-skill imports is -1.22 (significant at 10%) for total volume and -2.14 (significant at 5%) for right-wing volume. The mean-margin effect of high-skill imports is insignificant.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Low-skill immigration induces a transfer of votes from left-wing to right-wing populist parties, leaving total volume and the mean margin unchanged. The baseline PPML coefficient on low-skill immigration is 1.52 (significant at 1%) for right-wing volume and -1.78 (significant at 1%) for left-wing volume. In the reduced-form IV the right-wing volume coefficient is 1.97 (significant at 1%) and the left-wing coefficient is -1.70 (significant at 10%). The mean margin of total populism is not significantly affected by low-skill immigration in any specification.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;High-skill immigration reduces the volume of right-wing populism (PPML coefficient -1.32, significant at 1%; IV coefficient -2.02, significant at 5%) and generates a weak substitution toward left-wing populism in the baseline.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Descriptive findings: populism fluctuated since the 1960s, peaking after major economic crises (the oil shocks of the 1970s, deep crises of the 1990s, and after 2008). Right-wing populism reached an all-time high in the EU after 2005. The share of elections with at least one right-wing populist party rose from about 5 percent to more than 50 percent in EU member states over the study period.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Mechanisms.&lt;/strong&gt; Decomposing the volume margin into extensive (number of populist parties) and intensive (average vote share per party) sub-margins reveals that: the trade channel operates primarily through the intensive margin (existing populist parties gaining more votes); the immigration channel operates through the extensive margin (new right-wing populist parties with moderate scores entering parliament). Low-skill trade and immigration never increase the populism score of parties that have never been classified as populist, indicating that globalization shifts the composition of the party system rather than radicalizing mainstream parties.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Amplifiers and heterogeneity.&lt;/strong&gt; The right-wing populism response to low-skill imports is amplified during periods of de-industrialization and when internet coverage is high. Diversity in the origin mix of imported goods dampens the right-wing response. The populism response to low-skill immigration is not amplified by cultural distance between natives and immigrants; if anything, high cultural distance slightly reduces the centrist and left-wing populist responses. The effects on volume margin are primarily driven by EU28 countries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions and caveats.&lt;/strong&gt; Analysis is at the country level; party-level repositioning dynamics are left for further research. The unified trade-plus-immigration framework is new, but the long panel setting, unbalanced sample, and aggregate data impose limits on identifying specific mechanisms. The finding that globalization does not affect never-populist parties&amp;rsquo; scores limits concerns about contamination through party contagion in the short run. These results only partially confirm Rodrik&amp;rsquo;s (2018) hypothesis — left-wing populism is not robustly driven by trade shocks at the aggregate level, and trade&amp;rsquo;s effects are not confined to non-European contexts.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The identification relies on a two-stage approach. In the first stage (zero-stage gravity model), the authors predict bilateral flows of low- and high-skill goods and migrants using (i) time-invariant dyadic fixed effects interacted with a post-1990 structural-break dummy and (ii) origin-country-year fixed effects capturing time-varying push factors at the source. Critically, destination-country-time characteristics are excluded from the zero-stage, so the predicted aggregated flows capture only supply-side variation and bilateral connectivity — not demand-side populism dynamics in the destination. These predicted flows are then used as instruments. For the mean margin, standard 2SLS is implemented; for the volume margin, a reduced-form IV approach replaces actual flows with predicted flows to avoid the incidental-parameter problem in a PPML model with many fixed effects. The main threats are: (1) correlated origin shocks — if a push shock in origin country j simultaneously triggers populism in destination i through channels other than trade/migration (e.g., financial contagion), the exclusion restriction is violated; the authors cannot fully rule this out but note that including year fixed effects absorbs common global shocks; (2) the post-1990 structural break is used as an additional source of variation for bilateral dyadic ties, but the Berlin Wall dummy simultaneously captures many unobserved structural changes; (3) imputation of the skill structure of migration flows from census-round selection ratios (1990, 2000, 2010) introduces measurement error, though the authors show robustness to using only the year-2000 ratio; (4) Kleibergen-Paap F-statistics are around 10–12 when all four endogenous variables are instrumented simultaneously, which is modest; the authors show values are substantially larger when instrumenting one or two variables at a time.&lt;/p&gt;
&lt;h3 id="q2-how-are-trade-and-immigration-distinguished-empirically-and-how-is-the-skill-content-measured"&gt;Q2. How are trade and immigration distinguished empirically, and how is the skill content measured?&lt;/h3&gt;
&lt;p&gt;Trade data come from Feenstra et al. (2005) for 1962–2000 and UN Comtrade for 2001–2015. Product categories at the SITC 3-digit level are classified by skill and technology intensity following the Trade and Development Report (2002), yielding five categories: primary commodities, labor-intensive/resource-based, and manufacturing with low-, medium-, and high-skill labor intensity. The baseline uses only the low-skill and high-skill manufacturing ends; medium-skill goods are tested in robustness (their inclusion causes collinearity that kills volume-margin significance while preserving mean-margin results). Migration data come from Abel (2018) — five-year bilateral migration flow estimates interpolated to annual frequency. The skill level of migration flows is imputed by applying census-round skill-selection ratios (ratio of college graduates in the dyadic migrant stock to the native pre-migration population, from the closest available census round of 1990, 2000, or 2010) to the interpolated flows. Both trade and immigration variables enter as percentages — imports as share of GDP, immigration as share of destination population — averaged over the election year and the preceding year.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-difference-between-the-volume-margin-and-the-mean-margin-of-populism-and-why-does-it-matter"&gt;Q3. What is the difference between the volume margin and the mean margin of populism, and why does it matter?&lt;/h3&gt;
&lt;p&gt;The volume margin is the aggregate vote share of parties classified as populist (using a binary threshold of one standard deviation above mean in the populism score); it equals zero in elections with no populist party (about 60 percent of observations). The mean margin is the vote-weighted average populism score of all parties — populist and non-populist alike — so it is always defined and continuous. The mean margin captures the average ideological &amp;rsquo;exposure&amp;rsquo; of voters to populist ideas in a given election, including the spillover of populist ideas into mainstream parties. The distinction matters because globalization can affect the political landscape through multiple channels: it may shift votes toward existing populist parties (intensive margin of the volume margin), it may encourage new populist parties to enter (extensive margin), or it may shift the policy positions of all parties toward more populist stances (captured by the mean margin). The paper finds that low-skill trade raises both margins, but through different mechanisms — the volume effect operates through the intensive margin while the mean-margin effect partly reflects score increases among centrist populist parties. Low-skill immigration raises only the volume margin (through extensive-margin changes, not the mean margin).&lt;/p&gt;
&lt;h3 id="q4-how-is-the-populism-score-constructed-and-how-is-it-validated"&gt;Q4. How is the populism score constructed, and how is it validated?&lt;/h3&gt;
&lt;p&gt;The score is built from the Manifesto Project Database, which counts quasi-sentences associated with specific political topics as shares of party manifestos. Six MPD variables are selected, grouped into two dimensions: anti-establishment stance (AES — political corruption mentions and anti-pluralism/political authority mentions) and commitment-to-protect stance (CTP — protectionism, internationalism, EU institutions, and nationalization). A polychoric PCA within each dimension extracts the first principal component (by Kaiser criterion — eigenvalues above one). The two synthetic indices are then combined into a single populism score by equal weighting. A party is classified as populist if its score exceeds one standard deviation above the mean. This threshold maximizes the partial correlation with three of four external databases and maximizes accurate-forecast rates across all four databases. Probit regressions of existing binary classifications (Van Kessel 2015, Swank 2018, PopuList 2019, GPop 1 2020) on the continuous score yield ratios of accurate forecasts between 80 and 91 percent. OLS correlations with continuous external measures (GPop 2 leader-speech scores, CHES expert survey) are positive and significant. Unsupervised k-means clustering on the (AES, CTP) space confirms that parties above the one-SD threshold cluster distinctly in a well-separated region of the two-dimensional space. Extended scores using more MPD variables do not improve fit, confirming parsimony.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-across-left-wing-and-right-wing-populism-is-documented"&gt;Q5. What heterogeneity across left-wing and right-wing populism is documented?&lt;/h3&gt;
&lt;p&gt;The paper systematically decomposes results by political orientation (terciles of the RILE left-right index from MPD). Key heterogeneities: (1) Low-skill imports raise total and right-wing populism but not left-wing populism along the volume margin — this holds in baseline PPML and reduced-form IV. The mean-margin result is also concentrated in total and right-wing. (2) Low-skill immigration shifts votes from left-wing to right-wing populism (with opposing-sign PPML coefficients of 1.52 and -1.78, both significant at 1%), leaving total populism unchanged. High-skill immigration reverses this — it reduces right-wing and weakly increases left-wing populism. (3) High-skill imports reduce right-wing populism particularly (PPML -1.30, IV -2.14) and weakly shift votes toward left-wing populism. (4) Descriptively, the average populism score of right-wing populist parties increased since 2005 and reached 1.7 (2.1 standard deviations) in 2018, while left-wing populist parties&amp;rsquo; average score declined to 1.4 (1.75 standard deviations) — for the first time since the 1960s, radical-right populism is more intense than radical-left. (5) The volume-margin effects of globalization are primarily driven by EU28 countries. Among non-EU countries or when Latin America is excluded, results are directionally preserved but sometimes less precisely estimated.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The authors conduct an extensive battery documented in Appendix D: (1) Lag structure — the globalization variables are redefined using flows at t, t-1, t-2, average of t and t-1 (baseline), and the sum between elections; results on immigration are robust across lags; trade significance holds except at very short (election year) or very long (between elections) windows. (2) Populism threshold — results are preserved at the lax (0.9 SD) threshold and mostly preserved at the strict (1.1 SD) threshold, though some become insignificant when well-known parties like Syriza, M5S, and La France Insoumise exit the classification. (3) Skill imputation for immigration — using only year-2000 selection ratios yields similar results; interactions with migrant-stock quartile dummies are mostly insignificant. (4) Skill content of imports — adding labor-intensive and medium-skill imports does not disturb the baseline; collinearity from medium-skill imports kills volume-margin trade significance. (5) Origin-country income level — positive populism responses are concentrated in flows from low-income countries on the volume margin, but the mean-margin positive response is more driven by North-North movements. (6) Sub-samples — results are not driven by post-1990 years alone (interaction with post-1990 dummy attenuates but does not eliminate effects), not by Latin American countries (exclusion leaves results unchanged), and not by the unbalanced panel structure (restricting to countries present since 1970 confirms results). (7) Turnout — globalization variables do not significantly predict turnout, and results are robust to controlling for turnout. (8) Electoral system — results hold when controlling for electoral system; proportional representation systems show a significant effect of low-skill imports on left-wing populism volume. (9) Exports and emigration — including skill-specific export and emigration flows does not substantially alter the main coefficients; export and emigration effects are less significant and robust than import and immigration effects. (10) Vote-share normalization — results are robust to normalizing vote shares to sum to 100 percent.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work-especially-autor-et-al-2020-and-the-immigration-literature"&gt;Q7. How does this paper relate to and differ from closely related prior work, especially Autor et al. (2020) and the immigration literature?&lt;/h3&gt;
&lt;p&gt;Autor, Dorn, Hanson, and Majlesi (2020) study the electoral consequences of the China trade shock in the US, documenting polarization effects concentrated in a specific trade shock and a narrow time frame. The present paper extends this by: (1) spanning 60 years and 55 countries (vs. US-focused short panels); (2) studying trade and immigration jointly in one specification; (3) using continuous populism scores rather than party platforms; (4) distinguishing left- vs. right-wing populism responses; (5) examining skill content rather than origin-country GDP growth. On immigration, Edo et al. (2019) and Moriconi et al. (2022, 2019) document that the skill structure of immigration matters for voting — high-skill immigration reduces far-right votes while low-skill immigration raises them. The present paper confirms these findings in a much larger multi-decade panel and adds the novel result that low-skill immigration does not affect total populism but merely shuffles votes between left-wing and right-wing populism. On Rodrik&amp;rsquo;s (2018) taxonomy, the paper only partially confirms his hypothesis: left-wing populism is not robustly driven by trade shocks in the cross-country aggregate (only under specific amplifying conditions), and trade&amp;rsquo;s effects are not confined to non-European settings. A key novelty vs. the entire prior literature is the simultaneous inclusion of skill-specific trade and immigration flows — no prior cross-country long-panel study had done this.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The skill-content result implies that globalization&amp;rsquo;s effect on populism depends critically on whether economic integration predominantly involves low-skill or high-skill goods and workers. Policies that shift the composition of globalization toward high-skill activities — skill-upgrading policies, investment in education and retraining, managed migration policies that attract high-skill workers — could mechanically reduce populist pressures. The finding that low-skill immigration transfers votes from left to right without increasing total populism has a nuanced implication: reducing low-skill immigration may primarily benefit left-wing parties at the expense of right-wing ones rather than reducing aggregate political instability. The amplification by de-industrialization and internet access suggests that the populist dividend of adverse trade shocks is largest precisely when affected regions are also losing manufacturing jobs and when social media spreads grievance discourse. The attenuation by diversity in imported goods suggests that more geographically diversified trade may reduce the cultural-threat salience of any single origin. Scope conditions: the volume-margin effects are largely driven by EU28 countries, so the quantitative magnitudes may not generalize to other institutional contexts with different electoral systems; the analysis is at the country level and abstracts from regional labor-market dynamics; party-level repositioning of mainstream parties is not modeled.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-handle-the-measurement-challenge-of-comparing-populism-scores-across-countries-and-time"&gt;Q9. How does the paper handle the measurement challenge of comparing populism scores across countries and time?&lt;/h3&gt;
&lt;p&gt;This is a central methodological concern. The authors use party manifestos, which are available consistently across the 55 countries and the full 1960–2018 period in the Manifesto Project Database, allowing a principled content-based scoring without relying on expert surveys (which are available only for limited periods) or dichotomous external classifications (which are time-invariant in some datasets and country-limited in others). The two-stage PCA with polychoric principal components ensures that the dimensions are extracted from the structure of the data without imposing cardinal interpretations on ordinal quasi-sentence counts. The populism score has zero mean by construction with a standard deviation of 0.81, making cross-country and cross-time comparisons meaningful within the sample. The authors validate cross-country comparability by showing that the GPop 1 classification (which spans 1960–2018 for 36 countries) is well predicted by the score even though the score was not calibrated to that dataset specifically. An unsupervised clustering algorithm (k-means on the two dimensions) independently recovers the same set of parties as those above the one-SD threshold, without using any external label. The authors acknowledge that deliberate exclusion of immigration and multiculturalism variables from the score construction prevents mechanical correlation between the populism measure and the globalization regressors, which is an important design choice for the causal analysis.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-trends-in-the-right-left-decomposition-of-populism-over-the-study-period"&gt;Q10. What are the trends in the right-left decomposition of populism over the study period?&lt;/h3&gt;
&lt;p&gt;Descriptively (Section 3): the number of left-wing populist parties (as counted by the extensive margin) increased more than right-wing populist parties in the most recent period, partly because centrist parties are entering the populist bucket. However, the vote share gains (intensive margin) are dominated by right-wing populist parties. The share of elections with at least one left-wing populist party rose from about 15 to 30 percent globally over the study period. The share of elections with at least one right-wing populist party rose from about 5 to more than 50 percent in the EU and from about 10 to 25 percent in the rest of the world. The average populism score of right-wing populist parties increased since 2005, reaching 1.7 (about 2.1 standard deviations) in 2018, while the average score of left-wing populist parties declined to 1.4 (about 1.75 standard deviations). This means that for the first time since the 1960s, right-wing populist parties are on average more populist (by their own score) than left-wing populist parties. The gap between populist and non-populist parties&amp;rsquo; average scores has widened since 2008, consistent with the within-country Theil inequality increase after the financial crisis.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Volume margin of populism&lt;/strong&gt;: The aggregate vote share obtained by parties classified as populist (those with a populism score exceeding one standard deviation above the mean). Estimated with PPML given the large share of zero observations (about 60 percent of the sample). Captures whether populist parties win more votes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mean margin of populism&lt;/strong&gt;: The vote-weighted average populism score of all parties that obtained at least one seat in an election, regardless of whether they are classified as populist. Captures the average ideological &amp;rsquo;exposure&amp;rsquo; of voters to populist ideas, including spillovers into mainstream parties. Estimated with OLS.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Anti-establishment stance (AES)&lt;/strong&gt;: One of two dimensions underlying the paper&amp;rsquo;s populism score. Measured from Manifesto Project Database quasi-sentences on political corruption and anti-pluralism (political authority), capturing the core populist premise that the people are virtuous and the ruling class corrupt, leaving no room for pluralism or minority protection.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Commitment-to-protect stance (CTP)&lt;/strong&gt;: The second dimension underlying the populism score. Measured from Manifesto Project Database quasi-sentences on protectionism, internationalism, EU institutions, and nationalization, capturing populists&amp;rsquo; claim to shield &amp;rsquo;the people&amp;rsquo; from external or alien economic and cultural threats.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Skill-content of globalization&lt;/strong&gt;: The decomposition of import flows into goods intensive in low-skill vs. high-skill labor (using the SITC 3-digit classification from the Trade and Development Report 2002), and of immigration inflows into low-skill and high-skill workers (using dyadic skill-selection ratios from census rounds). The key empirical innovation of the paper: it is the skill content, not the size, of globalization flows that determines the direction and ideological valence of populist responses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gravity-based IV strategy&lt;/strong&gt;: An instrumentation approach that predicts bilateral skill-specific flows of goods and migrants using a zero-stage PPML regression with time-invariant dyadic fixed effects (interacted with a post-1990 structural-break dummy) and origin-country-year fixed effects, then aggregates predicted flows to the destination level. Excludes destination-country-time characteristics to purge reverse causality (populist governments restricting trade and immigration) and omitted variable bias.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extensive vs. intensive margin of the volume margin&lt;/strong&gt;: The decomposition of the total vote share for populist parties into the number of populist parties running (extensive margin) and the average vote share per populist party (intensive margin). Low-skill imports primarily affect the intensive margin (existing populist parties gain more votes); low-skill immigration primarily affects the extensive margin (new right-wing populist parties enter parliament).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Vote-transfer mechanism of low-skill immigration&lt;/strong&gt;: The paper&amp;rsquo;s finding that low-skill immigration reallocates votes between left-wing and right-wing populist parties without changing total populism. The authors interpret this as low-skill immigration enabling new right-wing populist parties with moderate populism scores to gain at least one seat in parliament (an extensive-margin effect), while simultaneously reducing the vote share and/or number of left-wing populist parties.&lt;/p&gt;</description></item><item><title>Property rights, fiscal capacity, and social capacity: The lasting impact of the Taiping Rebellion</title><link>https://macropaperwarehouse.com/papers/property-rights-fiscal-capacity-and-social-capacity-the-lasting-impact-of-the-taiping-rebellion/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/property-rights-fiscal-capacity-and-social-capacity-the-lasting-impact-of-the-taiping-rebellion/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: How do civil wars affect long-term development, and through which institutional mechanisms? The paper studies the Taiping Rebellion (1850-1864) in Qing China, one of history&amp;rsquo;s deadliest civil wars (at least ~20 million deaths, with some estimates of 70-100 million), as a critical juncture in China&amp;rsquo;s path to modernity. It matters because the rebellion generated large, persistent regional institutional variation that can help explain what the authors call the &amp;ldquo;Intra-China Divergence&amp;rdquo; — regional GDP-per-capita gaps as large as 27-to-1 (Dongguan vs. Tianshui, 2010) that rival the world&amp;rsquo;s largest inter-regional gaps.&lt;/p&gt;
&lt;p&gt;Data and design: A prefecture-level (occasionally county-level) panel covering 266 prefectures in China proper (1820 delineation). 55 prefectures fell under Taiping control (treatment) — split into 37 &amp;ldquo;Early Taiping&amp;rdquo; prefectures (occupied up to 1859, in Anhui/Jiangxi/Hubei, ambiguous land rights) and 18 &amp;ldquo;Late Taiping&amp;rdquo; prefectures (occupied from 1860, in Jiangsu/Zhejiang, stronger land rights) — and 211 control prefectures. Population is observed at seven points (1820, 1851, 1880, 1910, 1953, 1982, 2000). The core strategy is difference-in-differences (1820 reference year, prefecture and year fixed effects), supplemented by propensity-score matching (135-prefecture matched sample), a spatial autoregressive (SAR) model, and an instrumental-variable strategy using the longitude of the prefectural seat (motivated by the Taiping Navy&amp;rsquo;s eastward-along-the-Yangtze military strategy; first-stage F-statistics above 20).&lt;/p&gt;
&lt;p&gt;Main quantitative findings (with scope conditions): (1) Population: The rebellion caused large, permanent population losses. The Taiping DID coefficient is -0.45 in 1880 (a 36% lower population growth rate vs. control) and -0.51 in 1953 (40% lower) — no convergence. Crucially, in the matched sample Late Taiping areas recovered (no significant long-run population gap vs. control) while Early Taiping areas did not (an immediate ~30% drop in 1880 plus further decline). (2) Property rights: In 1915 county data, the idle-land share is 3.6 percentage points higher in Early Taiping than control counties, while Late Taiping is not significantly different from control — supporting the property-rights hypothesis. (3) Fiscal capacity (likin): Taiping areas collected ~12 times (e^2.5) as much likin per 1,000 sq km as control areas in 1869-1879, still 3.7 times as much in 1922-1925. Late Taiping areas had even higher intensity (22.2x in 1869-1879; 6.1x in 1922-1925) than Early Taiping (9.0x; 2.7x). (4) Social capacity (charities): On average the rebellion had no significant effect, but Late Taiping areas saw charity growth ~56 percentage points (44 log points) above control by 1880, rising to ~78 percentage points (58 log points) by mid-20th century. (5) Long-term development: Driven entirely by Late Taiping areas — 1982 agricultural+industrial output per capita 90% higher (64 log points), 2010 GDP per capita 87% higher (63 log points), and 2010 fiscal revenue per capita 203% higher (111 log points) than control; Early Taiping is statistically indistinguishable from control. Late Taiping counties also show higher post-1895 industrial firm entry. (6) Civic outcomes and resilience: Using CGSS 2010, Late Taiping residents show higher trust in personal networks and greater civic engagement (political attention, local participation). During the Great Famine (1959-1961), Taiping areas had 6.9% larger survivor cohorts; the effect is 28% stronger in Late Taiping (8.4%) than Early Taiping (6.5%).&lt;/p&gt;
&lt;p&gt;Implications: Violent conflict can leave lasting positive institutional imprints — through property rights, decentralized local fiscal capacity (&amp;ldquo;war made the state&amp;rdquo; at the local level), and elite-led social capacity — conditional on favorable initial conditions (strong gentry, wealthier commercial regions). The authors argue cultivating civil society and social capacity could yield large payoffs given China&amp;rsquo;s strong-state/weak-society configuration.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the core identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The baseline is a difference-in-differences comparing Taiping vs. control prefectures over 1820-2000, with prefecture and year fixed effects and 1820 as the reference year. Identification rests on parallel pre-trends: the Taiping coefficient in 1851 (pre-rebellion) is small and insignificant, indicating no differential selection conditional on controls. The main threats are: (i) the binary Taiping measure aligning with provincial boundaries and picking up broad regional dynamics; (ii) control-group contamination because some control prefectures were temporarily conquered (but not governed) by the Taiping Army; (iii) spatial spillovers between neighbors (Tobler&amp;rsquo;s law / Kelly 2019 critique); (iv) omitted subsequent historical events; and (v) omitted variables differing systematically between treated and control areas. The authors address these with dosage measures (battles, occupation months), matching, a SAR model, an IV (longitude), explicit controls for the Taiping conquest, an adjacent-treatment indicator, leave-one-province-out checks, and controls for many other historical events.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-instrumental-variable-strategy-work-and-why-might-longitude-be-valid"&gt;Q2. How does the instrumental-variable strategy work and why might longitude be valid?&lt;/h3&gt;
&lt;p&gt;Longitude of the prefectural seat instruments for the Taiping dummy. Relevance: the Taiping leaders&amp;rsquo; July 1852 military plan was to march eastward along the Yangtze, capture Jiangning (Nanjing), and expand from there using their dominant navy — so eastern (higher-longitude) prefectures were far more likely to fall under Taiping rule (Table 1 confirms Taiping prefectures have significantly larger longitudes; first-stage F-statistics above 20, Shea&amp;rsquo;s partial R-squared above 0.1). Exclusion: prefecture fixed effects absorb time-invariant geographic advantages, and year-dummy interactions with key geography (distances to coastline, Grand Canal, Yangtze) allow flexible time-varying geographic effects; conditional on these, longitude is argued to be excludable. IV estimates are larger in magnitude than OLS but qualitatively confirm a persistent negative population effect (robust to Anderson-Rubin weak-IV inference). The authors caution that omitted determinants correlated with longitude cannot be fully ruled out.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-four-hypotheses-and-how-are-they-distinguished-empirically"&gt;Q3. What are the four hypotheses and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;(1) Property-rights hypothesis: Late Taiping areas (post-1860 &amp;lsquo;direct tenant payment&amp;rsquo; system creating de facto/de jure tenant ownership) had better-defined land rights than Early Taiping areas (collapsed landlord system, lost deeds, anti-rent movements), so should have less idle land and faster population recovery — tested via the 1915 idle-land cross-section and the Early-vs-Late population DID. (2) Likin-as-fiscal-capacity hypothesis: Qing fiscal decentralization and the likin tax (introduced 1853) strengthened local fiscal capacity, persistently higher in Taiping (especially Late Taiping) areas — tested via the likin-intensity DID. (3) Social-change hypothesis: elite-led militias and reconstruction spurred charities (&amp;lsquo;benevolent halls&amp;rsquo;/shantang) as bridging social capital, especially in Late Taiping areas — tested via charity-stock DID and by adding charities as a mediator in long-term regressions. (4) Social-cohesion-and-civic-engagement hypothesis: forged social capital persists, raising modern trust/civic engagement and reducing Great Famine deaths — tested via CGSS 2010 and famine-survivor cohort ratios.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented"&gt;Q4. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;The central heterogeneity is Early vs. Late Taiping. Early Taiping areas (Anhui/Jiangxi/Hubei) suffered permanent population loss, higher idle land (+3.6pp), only modest likin gains, no charity growth, no long-term development advantage, and weaker famine resilience. Late Taiping areas (Jiangsu/Zhejiang) recovered population, had no excess idle land, far higher likin intensity (22x early period), large charity growth (+56 to +78pp), strong long-term development gains (90%/87%/203% in output/GDP/fiscal revenue), higher modern trust and civic engagement, and the strongest famine resilience (8.4% vs 6.5%). Industrialization heterogeneity is also temporal: no Early/Late firm-entry difference before 1895, but after the 1895 Treaty of Shimonoseki liberalized private industry, Late Taiping counties had more entry and Early Taiping fewer.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;For the population results: dosage interactions (log battles, log occupation months); excluding six most-intense-fighting prefectures (Wuchang, Songjiang, Anqing, Jiangning, Suzhou, Hangzhou); controlling for newly selected jinshi (civil-service quota channel); a SAR spatial model (after Pesaran cross-sectional-dependence tests); PSM matched sample; longitude IV with Anderson-Rubin inference; controls for seven other historical events (Guangxu Drought, Hui Revolt, Nian Rebellion, early-Republic conflicts, Sino-Japanese War, Chinese Civil War, missionary activity); explicit controls for Taiping conquest vs. regime; an adjacent-treatment indicator (Butts 2021) for spillovers; and leave-one-province-out exclusion. Long-term development results add SAR, matching, historical-event controls including the Cultural Revolution, and an &amp;lsquo;intermediate-term&amp;rsquo; 1930s industrialization check. Famine results are robust to alternative famine-severity measures, SAR, matching, and historical-event controls.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-mediation-analysis-handled-and-what-does-it-show"&gt;Q6. How is the mediation analysis handled and what does it show?&lt;/h3&gt;
&lt;p&gt;The authors add likin intensity (1880) and average charities (1880-1941) to cross-sectional long-term regressions, explicitly flagging these as endogenous &amp;lsquo;bad controls&amp;rsquo; (Angrist-Pischke 2009; Imai et al. 2011) to be interpreted cautiously as descriptive mediation. Findings: a one-SD increase in likin intensity is associated with +1.7pp middle-school completion, +4.8pp literacy, +5.3% schooling, and +12.2% (11.5 log points) GDP per capita in 2010. A one-SD increase in charities is associated with +15% 1982 output, +20% 2010 GDP, and +55% 2010 fiscal revenue per capita. Once charities are netted out, Late Taiping advantages in output, GDP, and fiscal revenue are attenuated by about 17%, 14%, and 22% respectively — highlighting the social-capacity channel.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-great-famine-resilience-result-connect-to-the-rebellion"&gt;Q7. How does the Great Famine resilience result connect to the rebellion?&lt;/h3&gt;
&lt;p&gt;Famine severity is measured by &amp;lsquo;Famine Control&amp;rsquo; = ratio of cohort size born during the famine (1959-1961) to cohort size born pre-famine (1954-1957) from the 1990 census 1% sample (higher = less severe). Taiping areas had a 6.9% larger survivor cohort than non-Taiping; the effect is 8.4% in Late Taiping vs. 6.5% in Early Taiping. Back-of-envelope, the Late Taiping experience would have &amp;lsquo;saved&amp;rsquo; ~31,374 people in an average prefecture (17% of the 1959-1961 cohort) vs. ~24,145 (13%) for Early Taiping. Controlling for political radicalism (reverse party-member density, -1*PMD, after Yang 1996) does not change the result. The mechanism: higher social capital made local officials more sympathetic/less radical in grain procurement and citizens better able to act collectively (paralleling Cao-Xu-Zhang 2022 on clan density and Hu-Yao-You 2023 on home-county officials).&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q8. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Prior Taiping studies examined narrower consequences: civil-service exam quotas (Li 2014), demographic and industrialization effects (Li and Ma 2016), migration and public goods (Hao and Xue 2017), and late-Qing power distribution (Bai, Jia, and Yang 2023). None addressed the rebellion&amp;rsquo;s enduring impacts on modern development, social trust, and Great Famine responses, nor the property-rights/fiscal-capacity/social-capacity mechanism triad. It complements Xue (2021) on Qing charities, generalized trust, and political participation, but extends to development outcomes. Against the European state-building literature (war strengthens central state capacity via centralization), this paper&amp;rsquo;s distinctive claim is that the Taiping Rebellion strengthened LOCAL fiscal capacity through DECENTRALIZATION, and expanded local social capacity that constrained the central state.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The benefits of war-induced institutions are conditional, not universal: they appeared chiefly in Late Taiping areas with a strong gentry class and favorable initial conditions for modern sectors (the wealthier, more commercial Lower Yangtze). The likin/fiscal-capacity benefits are explicitly stated to be conditional on strong gentry and good modern-sector initial conditions. The broad implication is that, given China&amp;rsquo;s very strong state but still weak society today, cultivating civil society and strengthening social capacity could yield particularly large long-term payoffs. The authors also caution (Appendix F.1) that likin could be distortionary taxation rather than fiscal capacity, arguing the fiscal-capacity interpretation is more relevant for long-term development.&lt;/p&gt;
&lt;h3 id="q10-what-significant-caveats-does-the-paper-acknowledge"&gt;Q10. What significant caveats does the paper acknowledge?&lt;/h3&gt;
&lt;p&gt;Long-term mechanisms cannot be exhaustively identified — likin and charities are endogenous outcomes, so mediation magnitudes are descriptive, not causal. History contains near-infinite interrelated events, so confounding cannot be fully eliminated (a fundamental limitation of all history-based work). The IV may have omitted correlates of longitude. Some 2SLS estimates for development outcomes were largely insignificant. The charity-stock measure assumes charities persisted once founded (no closure dates in the data). On property-rights persistence: using 2005 World Bank Enterprise Survey data they find no association between modern firms&amp;rsquo; perceived property-rights protection and Taiping regimes, suggesting the channel works through income effects rather than persistence of property rights per se.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Early vs. Late Taiping areas&lt;/strong&gt;: Early Taiping = prefectures occupied by the rebels up to 1859 (Anhui, Jiangxi, Hubei), where the old landlord system collapsed and land rights stayed ambiguous; Late Taiping = prefectures occupied from 1860 (Jiangsu, Zhejiang), where the Taiping introduced a &amp;lsquo;direct tenant payment&amp;rsquo; (作佃交粮) system and issued new deeds, granting tenants de facto/de jure ownership. This distinction is the paper&amp;rsquo;s central source of institutional variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Likin (lijin)&lt;/strong&gt;: A local tax on trade and commerce introduced in 1853 (a transit tax on travelling merchants&amp;rsquo; goods plus a business tax on resident merchants), collected in a decentralized, province-specific way. In the paper it is the operational measure of local fiscal capacity (likin revenue per 1,000 sq km), not central state capacity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social capacity&lt;/strong&gt;: In the paper&amp;rsquo;s sense, the ability of society to act collectively, constrain the state, and empower its members — operationalized empirically by the stock of local charity organizations (&amp;lsquo;benevolent halls&amp;rsquo;/shantang) that functioned as bridging social capital across classes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Likin-as-fiscal-capacity hypothesis&lt;/strong&gt;: The claim that the rebellion-induced likin system durably raised LOCAL fiscal capacity (an instance of Tilly&amp;rsquo;s &amp;lsquo;war made the state&amp;rsquo; operating locally rather than centrally), which improved public-goods provision and long-run development — conditional on strong gentry and favorable modern-sector initial conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stationary bandit (applied to Late Taiping rulers)&lt;/strong&gt;: Borrowing Olson (1993): in Late Taiping areas the consolidated, longer-horizon Taiping regime behaved like a stationary bandit, lowering effective tax rates, encouraging land registration, and securing tenant property rights to expand the tax base and promote production, unlike the looting/confiscation of the early stage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Famine Control&lt;/strong&gt;: The paper&amp;rsquo;s local famine-severity measure: the ratio of the cohort born during the Great Famine (1959-1961) to the cohort born pre-famine (1954-1957) in the 1990 census; a higher value means less severe famine and more survivors, and it is less vulnerable to government understatement of famine deaths.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intra-China Divergence&lt;/strong&gt;: The authors&amp;rsquo; term for China&amp;rsquo;s persistent, very large regional disparities in economic performance (up to 27-to-1 in GDP per capita) despite all regions historically sharing similar Malthusian income levels — the macro puzzle the rebellion&amp;rsquo;s institutional legacy helps explain.&lt;/p&gt;</description></item><item><title>Resource Misallocation in European Firms: The Role of Constraints, Firm Characteristics and Managerial Decisions</title><link>https://macropaperwarehouse.com/papers/resource-misallocation-in-european-firms-the-role-of-constraints-firm-characteristics-and-managerial-decisions/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/resource-misallocation-in-european-firms-the-role-of-constraints-firm-characteristics-and-managerial-decisions/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates why firms in the European Union exhibit wide dispersion in marginal revenue products (MRP) of capital and labor — a direct indicator of resource misallocation — and asks how much aggregate productivity the EU forfeits as a result. The research question is motivated by the persistent productivity gap between the EU and the United States, by evidence that within-country MRP dispersion in Europe has been trending upward since the mid-1990s, and by an institutional context in which the EU single market (launched in 1993) has not eliminated cross-country factor market frictions even three decades later.&lt;/p&gt;
&lt;p&gt;The primary data source is the EIB Investment Survey (EIBIS), a stratified random survey of non-financial enterprises conducted annually since 2016 across all 28 EU member states, covering manufacturing, services, utilities, and construction (NACE categories C–J). The analysis uses three waves (2016–2018), with approximately 12,500 firms per wave and a panel component of roughly 2,000 firms appearing in all three waves. Survey responses are matched to Orbis administrative data; the correlation between log employment in EIBIS and Orbis is 0.91, confirming data quality. MRP of capital (MRPK) is measured as the capital cost share times revenue divided by fixed assets; MRP of labor (MRPL) is the labor cost share times revenue divided by employment. Cost shares are calibrated from OECD STAN and Eurostat national accounts at the country–year–industry level.&lt;/p&gt;
&lt;p&gt;The theoretical framework is a dynamic model of a profit-maximizing firm with Cobb-Douglas production, isoelastic demand, and quadratic adjustment costs. Under the assumption that pure economic profits are small and that the labor output distortion is negligible (following Hsieh-Klenow 2009), the model implies that log MRPK and log MRPL can be approximated by observable average revenue products. The empirical strategy is a Mincerian regression of log MRPK (and log MRPL) on a rich vector of firm-level characteristics — firm demographics, input quality, capacity utilization, investment constraints, dynamic adjustment variables, and financing sources — plus country, industry, and year fixed effects (and their interactions). Because regressors are endogenous, the R² from OLS is interpreted as an upper bound on the share of MRP variance attributable to each factor (formally shown to dominate the IV R²). Marginal R² increments when a variable block is added identify the contribution of that block to the variance in MRP, which is then mapped into productivity gains via the Hsieh-Klenow formula.&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. Raw dispersion is large: the standard deviation of log MRPK is 1.43 and of log MRPL is 1.19 (and 1.63 for log MRPL minus log MRPK), all substantially exceeding comparable US figures (0.98 for capital and 0.58 for labor from Asker et al. 2014 and Bartelsman et al. 2013). The R² in the full regression is 0.14 (without fixed effects) and 0.49 (with country × industry × year fixed effects) for MRPK, and 0.29 and 0.74 respectively for MRPL. Among firm-characteristic blocks, the &amp;ldquo;adjustment&amp;rdquo; (dynamic investment and employment growth) and &amp;ldquo;demographics&amp;rdquo; (firm size, age, subsidiary and exporter status) blocks carry the largest marginal R² contributions; the &amp;ldquo;obstacles to investment&amp;rdquo; block (direct reports of constraints) contributes modestly by comparison. Country fixed effects alone explain R² = 0.052 for MRPK and R² = 0.445 for MRPL, while industry fixed effects alone explain R² = 0.239 for MRPK and R² = 0.268 for MRPL. The combined country–industry–year fixed-effects R² reaches 0.275 for MRPK and 0.611 for MRPL; adding the full interaction yields 0.492 and 0.736 respectively.&lt;/p&gt;
&lt;p&gt;Treating the &amp;ldquo;distortions&amp;rdquo; block of variables as genuine frictions, removing them would raise EU aggregate productivity by more than 40 percent (computed as 1.5 × 1.42 × 0.186 + 0.13 × 2.66 × 0.134 = 0.442). If all variables in X are treated as distortions, the implied gain is approximately 72 percent (0.715 in log points). Removing cross-country inequality in average MRPs (equalizing country fixed effects) would imply a 102 percentage log-point gain in productivity under the Hsieh-Klenow formula; removing barriers between industries and countries could raise productivity by at least 143 percentage log points.&lt;/p&gt;
&lt;p&gt;A Machado-Mata distributional decomposition comparing Germany (σ(log MRPK) = 0.92, σ(log MRPL) = 0.61) and Greece (σ(log MRPK) = 1.64, σ(log MRPL) = 0.91) reveals that the primary driver of Greece&amp;rsquo;s higher dispersion is the &amp;ldquo;prices&amp;rdquo; (regression coefficients reflecting institutional and policy environment), not the &amp;ldquo;endowments&amp;rdquo; (firm characteristics). Giving Greece German institutional &amp;ldquo;prices&amp;rdquo; reduces the counterfactual standard deviation of Greek MRPK from 1.66 to 0.94. This pattern generalizes across EU countries: German b (coefficients) tends to reduce MRPK dispersion for most countries, while German X (firm characteristics) tends to increase it, because Germany has more heterogeneous firms but an environment that prices those characteristics in a way that equalizes returns. This finding constitutes large-scale microeconomic evidence that institutions matter — cross-country differences in MRP dispersion reflect how business, institutional, and policy environments translate firm heterogeneity into outcomes, more than they reflect differences in firm characteristics per se.&lt;/p&gt;
&lt;p&gt;The policy implication is that deep institutional reform — not merely changes in firm composition — is required to narrow EU resource misallocation. The scope condition is that these estimates are upper bounds, and some observed MRP dispersion likely reflects compensating differentials (e.g., higher-quality capital commanding a higher MRPK) rather than pure distortions.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper does not attempt causal identification. Instead, it uses OLS to estimate equilibrium (Mincerian-type) regressions of log MRPK and log MRPL on firm characteristics plus fixed effects. The key insight is that OLS R² provides an upper bound on the share of MRP variance causally attributable to each regressor, because simultaneity or omitted variables can only inflate OLS R² above the true IV R². The main threats are: (1) endogeneity of regressors — a growing firm facing red tape will have high MRPK and a binding constraint simultaneously, inflating the R² attributed to constraints; (2) classical measurement error in survey responses, which attenuates R² toward zero (so OLS actually understates causal effects in this direction); (3) omitted variable bias via unobserved firm quality (managerial talent, etc.); (4) use of same variables (employment, fixed assets) on both left and right sides, addressed by cross-checking with Orbis data as instruments. The authors argue these threats are mostly conservative — they overstate, not understate, the upper bound.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-theoretical-justification-for-using-average-revenue-products-to-measure-marginal-revenue-products"&gt;Q2. What is the theoretical justification for using average revenue products to measure marginal revenue products?&lt;/h3&gt;
&lt;p&gt;Under the assumption that the share of pure economic profits is small (following Basu and Fernald 1997), the optimality conditions of the dynamic model imply that MRPK ≈ (capital cost share) × (revenue / capital) and MRPL ≈ (labor cost share) × (revenue / employment). These are average revenue products scaled by factor cost shares, matching Hsieh and Klenow (2009). The distortion framework further implies that the variance of log MRPK and log MRPL, when distortions are log-normally distributed and uncorrelated, maps directly into the Hsieh-Klenow productivity-loss formula, linking the regression R² to quantitative welfare calculations.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-role-of-compensating-differentials-versus-true-distortions-in-interpreting-the-results"&gt;Q3. What is the role of compensating differentials versus true distortions in interpreting the results?&lt;/h3&gt;
&lt;p&gt;The paper emphasizes that not all dispersion in MRPs reflects inefficient distortions. Some dispersion — particularly from &amp;lsquo;quality of capital,&amp;rsquo; &amp;lsquo;capacity utilization,&amp;rsquo; and &amp;lsquo;dynamic adjustment&amp;rsquo; — may reflect compensating differentials: firms that invest in higher-quality capital rationally face higher costs, demanding a higher MRPK in equilibrium, analogous to how more educated workers earn higher wages in a Mincerian framework. If these variables reflect compensating differentials rather than frictions, using &amp;lsquo;raw&amp;rsquo; MRP dispersion overstates misallocation. Conversely, if all variables proxy for distortions, the productivity gains from reform are even larger (72 percent versus 40 percent). The paper presents both interpretations explicitly, making the framework &amp;lsquo;highly portable&amp;rsquo; for different views of what drives observed dispersion.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-in-mrp-dispersion-is-documented-across-eu-countries-and-industries"&gt;Q4. What heterogeneity in MRP dispersion is documented across EU countries and industries?&lt;/h3&gt;
&lt;p&gt;Dispersion is notably lower in Germany (σ(log MRPK) = 0.92, σ(log MRPL) = 0.61) than in Greece (1.64 and 0.91) or smaller countries such as Malta, Luxembourg, and Cyprus. Country fixed effects explain R² = 0.445 of MRPL variation but only R² = 0.052 of MRPK variation, meaning labor is more segmented across countries than capital. Industry fixed effects explain R² = 0.239 for MRPK versus R² = 0.268 for MRPL, indicating capital is more segmented across industries than across countries. Core EU countries (France, Denmark) are relatively insensitive to counterfactual substitution of German coefficients, while periphery countries (Portugal, Ireland) show large movements. Romania, which resembles Slovenia in raw MRPK dispersion, looks much more like the Netherlands after controlling for firm characteristics — illustrating that observed dispersion rankings can be misleading without adjustment.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-machado-mata-decomposition-reveal-and-how-is-it-implemented"&gt;Q5. What does the Machado-Mata decomposition reveal, and how is it implemented?&lt;/h3&gt;
&lt;p&gt;The Machado-Mata (2005) decomposition separates the distribution of MRP into an &amp;rsquo;endowments&amp;rsquo; component (due to the values of firm characteristics X) and a &amp;lsquo;prices&amp;rsquo; component (due to the regression coefficients b, which capture how the institutional and policy environment translates X into outcomes). The decomposition draws B = 10,000 bootstrap samples from the empirical distribution of X for each country, combines them with quantile regression coefficients estimated separately for each country, and constructs counterfactual distributions. Applying Greek X with German b reduces Greece&amp;rsquo;s counterfactual σ(log MRPK) from 1.66 to 0.94 — close to Germany&amp;rsquo;s actual 0.92 — while applying German X with Greek b increases dispersion. The main finding is that differences in &amp;lsquo;prices&amp;rsquo; (institutional environment) dominate differences in &amp;rsquo;endowments&amp;rsquo; (firm characteristics) in explaining cross-country variation in within-country MRP dispersion. This pattern holds generally across EU countries: gains from &amp;lsquo;importing&amp;rsquo; German institutions are correlated with poor World Bank Governance Indicators and International Country Risk Guide scores.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-papers-estimates-of-eu-misallocation-compare-to-us-benchmarks"&gt;Q6. How do the paper&amp;rsquo;s estimates of EU misallocation compare to US benchmarks?&lt;/h3&gt;
&lt;p&gt;The EU standard deviations of log MRPK (1.43) and log MRPL (1.19) substantially exceed comparable US figures of 0.98 for capital (Asker et al. 2014) and 0.58 for labor (Bartelsman et al. 2013). The paper discusses three caveats for this comparison: (1) EIBIS uses revenue rather than value added, which affects dispersion (approximately +0.16 log points for MRPL, -0.21 for MRPK) — insufficient to explain the full gap; (2) survey measurement error is present but small — averaging over multiple waves reduces the standard deviation of log MRPK by only 8–12 percent; (3) EIBIS measures firms (not plants), and since about two-thirds of within-firm MRPK variance occurs across plants within firms (Kehrig and Vincent 2017), the EU–US comparison likely understates the true difference. Qualitatively, the greater EU dispersion is consistent with lower EU aggregate TFP relative to the US.&lt;/p&gt;
&lt;h3 id="q7-what-specific-regression-results-are-reported-for-individual-variable-blocks"&gt;Q7. What specific regression results are reported for individual variable blocks?&lt;/h3&gt;
&lt;p&gt;The full R² (without / with country × industry × year fixed effects) is 0.14 / 0.49 for MRPK and 0.29 / 0.74 for MRPL. Among variable blocks, the &amp;lsquo;adjustment&amp;rsquo; (investment, employment growth, past and planned investment) and &amp;lsquo;demographics&amp;rsquo; (size, age, subsidiary, exporter) blocks have the largest marginal R². The &amp;lsquo;obstacles to investment&amp;rsquo; (direct constraint reports) block contributes modestly, with some coefficients not statistically significant. Within regression coefficients (from Table A.4): older, exporting, high-utilization firms have higher MRPK and MRPL; investment is strongly negatively associated with MRPK (movement down the MRPK curve as capital rises) and positively with MRPL (labor becomes relatively scarcer); employment growth is positively associated with MRPK and negatively with MRPL (symmetric logic); credit-constrained status is negatively correlated with both MRPK and MRPL.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-are-run"&gt;Q8. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper reports: (1) &amp;lsquo;between&amp;rsquo; regressions on multi-year firm averages to reduce transitory variation and measurement error — results are qualitatively similar with slightly larger productivity gains; (2) restricting the sample to firms appearing in all three survey waves (Appendix Table A.5) — qualitatively similar results; (3) estimating equation (4) for each wave separately — similar results; (4) using Orbis employment and investment as regressors instead of EIBIS responses to address mechanical measurement-error correlation — nearly identical results (Appendix Table A.17); (5) replacing log(1+investment) with an indicator for positive investment (Appendix Table A.7) — similar results; (6) using industry-specific rather than country–year–industry cost shares — similar results; (7) confirming that measurement error can account for only a portion of the EU–US dispersion difference (8–12 percent reduction in standard deviation when averaging over waves). The paper also reports separate coefficient estimates for three blocs of EU countries (North/West, South, Center/East) in Appendix Tables A.10–A.16.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-relate-to-and-differ-from-hsieh-and-klenow-2009-and-related-prior-work"&gt;Q9. How does the paper relate to and differ from Hsieh and Klenow (2009) and related prior work?&lt;/h3&gt;
&lt;p&gt;The paper extends Hsieh and Klenow (2009) in several directions. First, while Hsieh-Klenow use administrative census-type data for India and China restricted to manufacturing, this paper uses a consistent cross-country survey covering all sectors in 28 EU countries, enabling direct cross-country comparison. Second, Hsieh-Klenow implicitly assume all MRP dispersion reflects distortions; this paper explicitly distinguishes distortions from compensating differentials and shows the distinction matters quantitatively. Third, this paper develops the Mincerian regression approach to apportion the variance in MRPs across observable factors — analogous to labor economists decomposing wage dispersion — and shows OLS R² provides a valid upper bound without requiring exogenous variation. Fourth, unlike country-level distortion measures (Gamberoni et al. 2016), tight theoretical restrictions (David and Venkateswaran 2017), or specific reforms (Rotemberg 2019), this paper draws on firm-level survey data with minimal restrictions and maintains high external validity. Fifth, the Machado-Mata distributional decomposition adds a new dimension absent from Hsieh-Klenow: decomposing cross-country differences into endowments vs. institutional &amp;lsquo;prices.&amp;rsquo;&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary policy implication is that EU productivity could rise by more than 40 percent if distortions to resource allocation were removed — and up to 72 percent if all observed MRP variation is attributed to distortions. A more modest goal of equalizing within-industry MRP dispersion across countries (i.e., making Germany and Greece similar within industries) implies gains of approximately 31–53 percent depending on interpretation. The decomposition evidence implies that institutional reform (changing how environments price firm characteristics) is more important than directly changing firm composition. The scope conditions are: (1) these are upper bounds derived from OLS; (2) some dispersion reflects compensating differentials that should not be counted as losses; (3) the EIBIS covers firms with at least 5 employees, so very small firms are excluded; (4) the framework assumes log-normal, uncorrelated distortions and constant returns to scale — relaxing these can increase estimated losses further (Jones 2011); (5) the estimates do not account for firm-level markup heterogeneity, which could overstate or understate other channels.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-paper-contribute-to-the-literature-on-measurement-error-in-mrp-studies"&gt;Q11. What does the paper contribute to the literature on measurement error in MRP studies?&lt;/h3&gt;
&lt;p&gt;The paper shows formally (Appendix D) that classical measurement error in regressors attenuates OLS R² toward zero, so OLS provides a conservative upper bound from this direction. It also shows that averaging across multiple survey waves reduces measurement error while also attenuating transitory adjustment-cost variation, so multi-year averages likely overstate the role of measurement error. Crucially, the paper validates EIBIS against Orbis administrative data, finding a 0.91 correlation for log employment, similar standard deviations of log MRPK (1.44 in Orbis vs. 1.37 in EIBIS) and log MRPL (1.07 in Orbis vs. 1.30 in EIBIS) for matched firms, and a mean absolute log difference in standard deviations of approximately 2 percent across countries. This contributes to the debate initiated by Bils et al. (2017) on whether measured MRP dispersion reflects mismeasurement, and corroborates that surveys can be reliable substitutes for census-type administrative data in cross-country analysis.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-paper-find-about-the-role-of-credit-constraints-specifically"&gt;Q12. What does the paper find about the role of credit constraints specifically?&lt;/h3&gt;
&lt;p&gt;Credit constraint status (defined as loan rejection, discouragement from applying, or receiving a loan that was too small or too expensive) is negatively correlated with both MRPK and MRPL in the full regression. This is consistent with credit-constrained firms being unable to invest to the point where MRPK is equalized with the cost of capital, but the negative sign also raises the interpretive caveat noted by the authors: cross-sectional equilibrium relationships can have signs inconsistent with causal priors because constraints may be more binding for firms that are already performing poorly. The &amp;lsquo;source of funds&amp;rsquo; block (share of investment from internal vs. external sources, and credit constraint) is grouped with &amp;lsquo;distortions&amp;rsquo; in the paper&amp;rsquo;s preferred decomposition.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Marginal Revenue Product (MRPK/MRPL)&lt;/strong&gt;: In this paper, the marginal revenue product of capital (MRPK) and labor (MRPL) are measured as observable average revenue products — the capital or labor cost share times revenue divided by the stock of capital or employment. Under the paper&amp;rsquo;s model assumptions, these approximate the shadow cost of inputs and serve as the primary measure of firm-level resource allocation efficiency. A firm with a high MRPK relative to its cost of capital is under-capitalized; dispersion of MRPK across firms signals misallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Compensating differentials (in the MRP context)&lt;/strong&gt;: The paper adapts the Mincerian concept of compensating differentials from labor markets to the firm side: some observed dispersion in MRPK and MRPL may reflect optimal responses to heterogeneity in input quality, capital utilization, or adjustment dynamics — not inefficient distortions. For example, a firm with state-of-the-art machinery may face a higher MRPK reflecting the quality premium, not a barrier to investment. Because such dispersion is rational, it should be subtracted from productivity-loss calculations rather than counted as welfare-reducing misallocation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Machado-Mata decomposition&lt;/strong&gt;: A distributional decomposition technique (Machado and Mata 2005) applied here to attribute cross-country differences in the dispersion of MRPK and MRPL to two components: &amp;rsquo;endowments&amp;rsquo; (the empirical distribution of firm characteristics X in a given country) and &amp;lsquo;prices&amp;rsquo; (the regression coefficients b, which capture how the country&amp;rsquo;s business, institutional, and policy environment translates those characteristics into marginal revenue products). The decomposition constructs counterfactual MRP distributions by combining one country&amp;rsquo;s X with another country&amp;rsquo;s b.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mincerian productivity regression&lt;/strong&gt;: The paper&amp;rsquo;s core empirical framework, modeled explicitly on Mincer&amp;rsquo;s (1958) wage regression: just as wages are regressed on worker characteristics (education, experience) to decompose earnings dispersion, log MRPK and log MRPL are regressed on firm characteristics (demographics, quality, utilization, adjustment, constraints, financing) to decompose MRP dispersion. OLS R² in this regression is an upper bound on the share of MRP variance attributable to each regressor.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;EIB Investment Survey (EIBIS)&lt;/strong&gt;: An annual firm-level survey administered by Ipsos MORI on behalf of the European Investment Bank since 2016, covering all 28 EU member states with a stratified random sample of approximately 12,500 non-financial enterprises per wave (minimum 5 employees, NACE C–J). Unique features include consistent cross-country design, merger with Orbis administrative data, and questions on investment plans, capital quality, capacity utilization, perceived obstacles, and financing sources — all directly informative about sources of MRP variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Institutional &amp;lsquo;prices&amp;rsquo; on firm characteristics&lt;/strong&gt;: In the Machado-Mata framework as applied here, &amp;lsquo;prices&amp;rsquo; refer to the country-specific regression coefficients b in the MRP regression — how steeply a country&amp;rsquo;s environment (regulations, institutions, policies) translates a given unit of firm heterogeneity in X into a difference in marginal revenue products. Countries with smaller b magnitudes (like Germany) achieve more equalization of MRPs across heterogeneous firms, reflecting an efficient institutional environment; countries with large b (like Greece) amplify firm-level heterogeneity into large MRP dispersion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Upper-bound R² approach to productivity gains&lt;/strong&gt;: The paper&amp;rsquo;s portable method for quantifying productivity gains from removing a friction: the marginal R² increment in an OLS regression of log MRPK (or log MRPL) when a friction variable is added is an upper bound on the share of MRP variance attributable to that friction. This bound, multiplied by the variance of log MRP and the Hsieh-Klenow productivity-loss formula parameters, gives an upper-bound estimate of the aggregate TFP gain from eliminating that friction. The method does not require exogenous variation or tight structural assumptions.&lt;/p&gt;</description></item><item><title>Selection, Structural Transformation, and the Cost Disease of Services</title><link>https://macropaperwarehouse.com/papers/selection-structural-transformation-and-the-cost-disease-of-services/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/selection-structural-transformation-and-the-cost-disease-of-services/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether worker self-selection, rather than slow technological progress, can explain the low measured labor productivity growth in the U.S. service sector — a phenomenon known as Baumol&amp;rsquo;s cost disease. The conventional view, associated with Young (2014), is that as workers reallocate from manufacturing into services, the incoming workers are less skilled than incumbents, mechanically depressing measured productivity; on that view, the cost disease might be a transient mismeasurement artifact rather than a permanent technological fact. Shu challenges this interpretation by showing that the selection pattern differs sharply across service sub-sectors and is far weaker in aggregate than the conventional model predicts.&lt;/p&gt;
&lt;p&gt;The empirical foundation is the Outgoing Rotation Group of the U.S. Current Population Survey (1989–2020), linked longitudinally to track workers who switch sectors between consecutive years. The sample contains 1,406,674 matched worker-year observations. Cross-country patterns from the GGDC 10-Sector Database (nine developed countries, 1989–2009) provide motivating evidence: over that period, labor productivity grew 56 log points in manufacturing, 77 log points in professional services (finance, real estate, professional and business services), and only 5 log points in EHP (education, health, and public administration). The cross-country correlation between employment-share growth and labor productivity growth is +0.48 for professional services — the opposite sign from the conventional selection story — and −0.14 for EHP, which conforms to it.&lt;/p&gt;
&lt;p&gt;At the micro level, a regression of log real weekly earnings on previous-sector dummies (with year and county fixed effects, standard errors clustered by county) yields a key asymmetry: workers who move from manufacturing into professional services earn 4.8 log points (approximately 4.9%) more than incumbent professional services workers (coefficient 0.048, se 0.010), while workers who move from EHP into professional services earn 14.3 log points less (coefficient −0.143, se 0.008). Workers switching from manufacturing into EHP earn 8.7 log points less than EHP incumbents (coefficient −0.087, se 0.023). The first fact — that incoming workers from manufacturing outperform incumbents in professional services — cannot be generated by conventional Roy models based on independent Fréchet skill distributions, which force skill levels in an expanding sector to fall.&lt;/p&gt;
&lt;p&gt;To accommodate these patterns, Shu builds a three-sector general-equilibrium Roy model with a non-homothetic CES demand structure (following Comin, Lashkari and Mestieri 2021). The skill distribution is parameterized by allowing absolute advantage in professional services to depend on comparative advantages in manufacturing (parameter αm) and EHP (αe), conditional on the comparative advantage quantiles following a Gumbel distribution. The model is estimated via simulated method of moments, targeting the three observed earnings premia and the variance of log income. The estimated parameters confirm αm = 0.055 &amp;gt; 0 (workers with higher comparative advantage in manufacturing also have higher absolute productivity in professional services) and αe = −0.123 &amp;lt; 0 (workers with higher comparative advantage in EHP are less productive in professional services).&lt;/p&gt;
&lt;p&gt;The main quantitative results for the full 1990–2020 sample are: selection raises labor productivity in professional services by 1.2 log points and lowers it in EHP by 0.7 log points, for a net effect of zero on aggregate services. By contrast, the conventional independent Fréchet model predicts selection effects of −8.7 log points for professional services and −3.0 log points for EHP, summing to −5.2 log points for aggregate services. The discrepancy for professional services alone is 9.9 log points — a difference of more than seven-fold in magnitude and opposite in sign. Consequently, the conventional model overpredicts true technology growth in professional services by over one-third relative to the baseline. The implied true technology growth rates over 1990–2020 are 88.1 log points for manufacturing, 27.3 for professional services, and −0.6 for EHP, leaving a large and unexplained productivity gap between manufacturing and services that selection cannot close. This directly refutes Young&amp;rsquo;s (2014) claim that selection accounts for virtually all of the measured gap, and confirms that Baumol&amp;rsquo;s cost disease reflects genuinely low technology growth in EHP and moderately lower growth in professional services.&lt;/p&gt;
&lt;p&gt;A forward-looking simulation extending the implied technology growth rates (2.9% p.a. for manufacturing, 0.9% for professional services, 0% for EHP) over fifty years produces similar welfare gains under both specifications (29.4 vs. 29.2 log points), but through very different mechanisms: the conventional model reaches its welfare estimate through counterfactually large selection effects in both directions that cancel, while the baseline model generates more modest and empirically grounded reallocation dynamics.&lt;/p&gt;
&lt;p&gt;The unexplained portion of the manufacturing-to-professional-services earnings premium is explored through an extensive set of micro-regressions controlling for education, experience, hours, occupation, age, race, and gender. Gender composition is the single most important observable channel: workers switching from manufacturing into professional services are 17.7 percentage points more male than the incumbent professional services workforce, and male workers earn roughly 40% more, implying a composition-driven premium of about 7.1 log points. Even after controlling for all observables, approximately one-quarter of the 4.8 log-point premium remains unexplained. Among college-educated female workers, the unexplained manufacturing premium is 4.5 log points — as large as the unconditional estimate — which Shu flags for future investigation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-micro-level-selection-patterns-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy for the micro-level selection patterns, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses the longitudinal structure of the CPS Outgoing Rotation Group to observe the same worker in two consecutive years and identify their origin sector and destination sector. The income gap between incoming workers and incumbents in the same sector-year cell (conditional on year and county fixed effects, with county-clustered standard errors) provides the key moments. The main threats are: (1) workers may self-select into switching for unobserved reasons correlated with productivity (e.g., those with better outside options move), but the direction of such bias is ambiguous; (2) the paper explicitly focuses on direct sector-to-sector transitions to isolate long-run structural reallocation from short-run labor supply fluctuations — a design choice distinguishing it from Young (2014), who used aggregate defense spending as an IV but thereby conflated unemployment and non-participation dynamics with genuine sector reallocation. The paper does not employ a separate instrument for the selection into switching; instead, it uses the income-gap moments as identified empirical objects to discipline the structural model.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-differ-from-young-2014-and-why-does-it-reach-opposite-conclusions"&gt;Q2. How does the paper differ from Young (2014), and why does it reach opposite conclusions?&lt;/h3&gt;
&lt;p&gt;Young (2014) uses industry-level employment and output data and estimates a uniform, negative elasticity of &amp;lsquo;worker efficacy&amp;rsquo; with respect to employment share across all industries, concluding that selection explains away essentially all of the manufacturing–services productivity gap. Three key differences drive Shu&amp;rsquo;s opposite conclusion. First, Shu uses worker-level panel data that allow distinct selection patterns to be estimated separately for professional services versus EHP, rather than imposing a common pattern. Second, Shu documents that the conventional pattern (incoming workers earn less than incumbents) holds for EHP but fails for professional services, where workers from manufacturing earn about 4.9% more than incumbents — a fact Young&amp;rsquo;s approach cannot detect. Third, Young&amp;rsquo;s IV (defense spending-to-GDP ratio) is used for demand shocks on aggregate employment, which mixes short-run unemployment and non-participation adjustments with the long-run structural reallocation that is relevant for selection; Shu&amp;rsquo;s design isolates workers who transition directly between sectors and thus captures only the long-run phenomenon.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-role-of-the-relationship-between-absolute-and-comparative-advantages-in-the-model-and-how-does-the-paper-generalize-prior-work"&gt;Q3. What is the role of the relationship between absolute and comparative advantages in the model, and how does the paper generalize prior work?&lt;/h3&gt;
&lt;p&gt;Standard Roy models (including those using independent Fréchet distributions as in Lagakos and Waugh 2013, Bryan and Morten 2019, and Hsieh et al. 2019) implicitly assume that workers&amp;rsquo; absolute advantage in a sector increases with their comparative advantage in the same sector. This restriction forces labor productivity of any expanding sector to fall. Adão (2016) and Alvarez-Cuadrado, Amodio and Poschke (2019) made the theoretical point that the sign of αm (the correlation between comparative advantage in manufacturing and absolute advantage in professional services) is the key determinant of whether selection helps or hurts professional services productivity. Shu&amp;rsquo;s paper generalizes Adão&amp;rsquo;s two-sector log-linear framework to three sectors, introduces the explicit parameterization via the Gumbel conditional distribution, and crucially provides a parametric method to quantify the contribution of selection to measured labor productivity by estimating αm and αe from worker-level moments. The estimated αm = 0.055 &amp;gt; 0 is what generates the positive selection effect for professional services.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-calibrated-technology-growth-rates-implied-by-the-model-and-what-do-they-imply-for-baumols-cost-disease"&gt;Q4. What are the calibrated technology growth rates implied by the model and what do they imply for Baumol&amp;rsquo;s cost disease?&lt;/h3&gt;
&lt;p&gt;Over 1990–2020, the calibrated model implies cumulative technology growth of 88.1 log points in manufacturing, 27.3 log points in professional services, and −0.6 log points in EHP. These numbers confirm that technology growth in EHP has been essentially zero over three decades, and that professional services, despite having high measured labor productivity growth, has grown at roughly one-third the rate of manufacturing in true technology terms. The 93.5 log-point difference in measured output per worker between manufacturing and aggregate services is broken down as: 15.6 log points attributable to the selection effect on manufacturing (outgoing workers are below-average) and essentially zero attributable to selection in aggregate services, leaving a true technology gap of approximately 77.9 log points. The conclusion is that the cost disease — specifically the stagnation of EHP — is a real technological phenomenon, not a mismeasurement artifact.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-conventional-independent-fréchet-model-compare-quantitatively-to-the-baseline-and-where-do-the-specifications-diverge-most"&gt;Q5. How does the conventional independent Fréchet model compare quantitatively to the baseline, and where do the specifications diverge most?&lt;/h3&gt;
&lt;p&gt;The comparison is presented in Table 7. For professional services, the baseline finds a selection effect of +1.2 log points while the conventional model finds −8.7 log points — a difference of 9.9 log points, more than seven-fold in magnitude and reversed in sign. For EHP the baseline finds −0.7 versus −3.0 under the conventional model. For aggregate services the baseline finds 0.0 versus −5.2 for the conventional model. In the implied technology growth, the conventional model overpredicts professional services technology growth by over one-third relative to the baseline (37.3 versus 27.3 log points), and for aggregate services overpredicts by more than 50% (15.4 versus 10.2 log points). In the 50-year forward projection, both models produce nearly identical welfare changes (29.4 vs. 29.2 log points) but through opposite and partially offsetting selection effects in manufacturing versus services under the Fréchet model — a result Shu flags as an artifact of the conventional model&amp;rsquo;s internally inconsistent mechanism.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-in-selection-patterns-is-documented-at-the-micro-level"&gt;Q6. What heterogeneity in selection patterns is documented at the micro level?&lt;/h3&gt;
&lt;p&gt;Three dimensions of heterogeneity are documented. First, the direction of selection differs by sub-sector: incoming manufacturing workers earn more than incumbents in professional services (+4.9%) but less than incumbents in EHP (−8.7%). Second, the role of observables differs: in professional services, none of the standard controls (education, experience, hours, occupation, age, race) eliminate the manufacturing premium, while gender composition accounts for roughly three-quarters of it. In EHP, the same set of controls explains the income gaps well, consistent with conventional selection. Third, the premium within professional services is concentrated among college graduates: among workers with college degrees, the manufacturing premium is 2.7%; among those without degrees, it is statistically indistinguishable from zero. College-educated female workers from manufacturing show a particularly strong premium of 4.5 log points, larger than most subgroups. Male workers switching from manufacturing constitute over 60% of the inflow for most of the sample, compared to roughly 50% male share among incumbents (the male share of incumbents rises over time as the inflow changes the composition).&lt;/p&gt;
&lt;h3 id="q7-what-role-does-gender-play-in-explaining-the-manufacturing-earnings-premium-in-professional-services"&gt;Q7. What role does gender play in explaining the manufacturing earnings premium in professional services?&lt;/h3&gt;
&lt;p&gt;Gender is the quantitatively dominant observable channel. Workers reallocating from manufacturing into professional services are on average 17.7 percentage points more male than the incumbent professional services workforce. Male workers earn roughly 40% (log 0.407) more than female workers within professional services. A back-of-envelope calculation: a 17.7 percentage-point male-share gap times a 40% earnings premium implies a composition-driven premium of approximately 7.1 log points, which matches the difference between the unconditional coefficient (0.048) and the gender-conditioned coefficient (−0.022). Adding the gender dummy to the regression turns the manufacturing premium negative and marginally significant (−0.022, Table 10 column 4), confirming that the premium is largely a composition effect. However, Table 11&amp;rsquo;s full specification (including all observable controls) still leaves a positive residual of 1.2 log points (statistically significant), suggesting approximately one-quarter of the original 4.8 log-point premium is genuinely unexplained. The paper identifies non-pecuniary sorting preferences (Goldin 2014; Faberman, Mueller and Şahin 2025) and sector-specific human capital as candidate explanations for future research.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-are-run-and-what-is-the-sensitivity-of-results-to-parameter-choices"&gt;Q8. What robustness checks are run, and what is the sensitivity of results to parameter choices?&lt;/h3&gt;
&lt;p&gt;The paper compares the baseline model to an independent Fréchet specification (shape parameter 2.7, consistent with Bryan and Morten 2019 and Lagakos and Waugh 2013) as the main alternative parameterization. It notes in a footnote that lower shape parameters (Hsieh et al.&amp;rsquo;s ~2, or Young&amp;rsquo;s implied ~1.33) would produce even stronger negative selection effects, making the Fréchet comparison conservative. At the micro level, the earnings regressions are extended through five successive specifications in Tables 9, 10, 11, and 12, each adding further controls, to verify the robustness of the manufacturing premium in professional services. The premium survives across all specifications for workers with college degrees. The paper also notes that its selection effect is identified entirely from worker-level income data and does not depend on the measured numbers of labor productivity, so measurement errors in sectoral output data (discussed in Triplett and Bosworth 2004) do not contaminate the core finding. The paper excludes workers under 25 to ensure the sector choices are long-run-oriented rather than early-career experiments.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-future-projection-exercise-show-and-what-are-its-scope-conditions"&gt;Q9. What does the future projection exercise show, and what are its scope conditions?&lt;/h3&gt;
&lt;p&gt;The exercise projects structural transformation over 2020–2070 by feeding the sample-period-implied technology growth rates (2.9% p.a. for manufacturing, 0.9% for professional services, 0% for EHP) into both specifications, starting from 2020 equilibrium conditions. Under the baseline model, manufacturing employment share declines by 9.1 percentage points, professional services by 5.2 points, and EHP rises by 14.3 points — reflecting that stagnant EHP technology must absorb more workers to meet demand. The conventional Fréchet model produces less contraction in professional services (−2.2 points) and more in manufacturing (−11.5 points). Both specifications predict similar welfare gains (~29 log points). The scope condition is that these projections treat technology growth rates as exogenous and constant at their sample-period averages; they abstract from endogenous innovation, feedback between human capital reallocation and technology, and from demand-side shifts (which Duernecker, Herrendorf, and Valentinyi 2024 and Sen 2021 emphasize).&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-and-differ-from-the-broader-structural-roy-model-literature"&gt;Q10. How does this paper relate to and differ from the broader structural Roy model literature?&lt;/h3&gt;
&lt;p&gt;The paper is in direct dialogue with Lagakos and Waugh (2013), who use a two-sector Roy model with independent Fréchet marginals to explain cross-country agricultural/non-agricultural productivity gaps; Bryan and Morten (2019) and Hsieh et al. (2019), who use multivariate Fréchet to evaluate productivity gains from reducing labor market frictions; and Adamopoulos et al. (2022), Pulido and Świecki (2019), and Gai et al. (2025), who use multivariate normal distributions for similar questions. All these papers find that sector expansion is accompanied by falling average worker quality — a consequence of the parametric restriction that comparative advantage aligns positively with absolute advantage in the same sector. Adão (2016) and Alvarez-Cuadrado, Amodio and Poschke (2019) showed theoretically that this alignment is the key sufficient condition for the conventional result, and found non-parametric evidence against it in some sectors. Shu&amp;rsquo;s contribution is to provide a tractable parametric framework (Gumbel conditional on quantile ranks) that relaxes this restriction, estimate it with the relevant micro moments (earnings gap between incumbents and switchers), and show quantitatively that the relaxation matters enormously — reversing the sign of the selection effect for professional services.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q11. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The primary implication is that policies aimed at accelerating technology growth in EHP (education, health, public administration) are warranted because the cost disease there is genuine and not a mismeasurement artifact. The paper explicitly confirms that low labor productivity growth in services reflects slow true technology growth, especially in EHP where the calibrated 30-year technology growth is essentially zero. The positive selection effect for professional services (1.2 log points over 30 years) is quantitatively small and does not materially offset the technology disadvantage. A secondary implication is that conventional models used in trade and development economics (with independent Fréchet skill distributions) systematically overstate the adverse selection effect of sectoral expansion, leading to overprediction of implied technology growth in professional services by over one-third. Studies using such models to evaluate, for example, gains from reducing labor market frictions should interpret their implied technology parameters with caution. Scope conditions: the model takes technology as exogenous and abstracts from endogenous responses of innovation to worker quality, from demand-side dynamics studied elsewhere, and from industry-level heterogeneity within the broad sub-sectors.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Selection effect (on labor productivity)&lt;/strong&gt;: In this paper, the change in average skill level of workers in a sector induced by reallocation — measured as the difference between measured labor productivity growth and true technology growth. A positive selection effect means incoming workers are more skilled than incumbents on average; a negative effect means they are less skilled. The paper distinguishes the selection effect from the conventional presumption that expansion always produces negative selection.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Absolute advantage (in professional services)&lt;/strong&gt;: A worker&amp;rsquo;s log skill level in professional services, a(i) ≡ ln z_p(i), which determines output contribution to that sector independently of what the worker could earn elsewhere. In the model, absolute advantage is distributed Gumbel conditional on the worker&amp;rsquo;s comparative advantages, with mean α(q_m, q_e) = α_m ln q_m + α_e ln q_e.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Comparative advantage (between sectors)&lt;/strong&gt;: The log ratio of a worker&amp;rsquo;s skill in one sector relative to professional services: s_m(i) ≡ ln(z_m(i)/z_p(i)) for manufacturing and s_e(i) ≡ ln(z_e(i)/z_p(i)) for EHP. A worker&amp;rsquo;s comparative advantage determines which sector they choose when wage rates are equalized, while the relationship between comparative and absolute advantage determines the productivity of workers on the margin of switching.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;α_m and α_e parameters&lt;/strong&gt;: The key parameters governing whether incoming workers from manufacturing (α_m) or EHP (α_e) are more or less productive in professional services than incumbents. When α_m &amp;gt; 0, workers with a high comparative advantage in manufacturing also have high absolute advantage in professional services, so that reallocation from manufacturing raises average quality in professional services. When α_e &amp;lt; 0, workers with high comparative advantage in EHP have low absolute advantage in professional services, so inflows from EHP lower quality. Estimated values: α_m = 0.055, α_e = −0.123.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Baumol&amp;rsquo;s cost disease&lt;/strong&gt;: Used in this paper to refer to the phenomenon whereby the service sector&amp;rsquo;s true technology growth is persistently low relative to manufacturing — implying that resources must continuously be reallocated to services to maintain consumption of service output, raising the relative price of services. The paper confirms this is a genuine technology fact, not a mismeasurement artifact from selection, especially for EHP where 30-year cumulative technology growth is calibrated at essentially −0.6 log points.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Income premium of switching workers&lt;/strong&gt;: The difference in log real weekly earnings between workers who transitioned from a given source sector in the prior year and workers who were already in the destination sector (incumbents), estimated by regression with year and county fixed effects. This premium is the paper&amp;rsquo;s primary empirical moment and the main target for identifying the skill-distribution parameters. Positive premium (MFG→PROF: +0.048) indicates incoming workers are more productive; negative premium (EHP→PROF: −0.143; MFG→EHP: −0.087) indicates they are less productive.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Independent Fréchet specification (benchmark)&lt;/strong&gt;: The conventional parametric Roy model in which each worker&amp;rsquo;s sector-specific skills are drawn independently from Fréchet marginal distributions. This specification implies that workers&amp;rsquo; absolute advantage in a sector is negatively correlated with their comparative advantage — an implicit restriction that forces average skill in any expanding sector to decline with employment share. The paper uses this as the comparison case, with shape parameter 2.7 following Bryan and Morten (2019) and Lagakos and Waugh (2013), and shows it mispredicts the selection effect for professional services by 9.9 log points and reverses its sign.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-homothetic CES preference&lt;/strong&gt;: The demand structure from Comin, Lashkari and Mestieri (2021) used in the model, which allows income elasticities to differ across sectors and vary with aggregate consumption. It governs how structural transformation proceeds on the demand side as incomes grow. Calibrated parameters imply professional services demand is most income-elastic (ξ_p = 1.382) and EHP demand is least income-elastic (ξ_e = 0.644), so growth shifts expenditure toward professional services and eventually toward EHP as incomes rise further.&lt;/p&gt;</description></item><item><title>Serial Entrepreneurship in China</title><link>https://macropaperwarehouse.com/papers/serial-entrepreneurship-in-china/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/serial-entrepreneurship-in-china/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper studies entrepreneurship and new firm creation in China through the lens of serial entrepreneurs (SEs) — individuals who establish more than one firm — contrasting them with non-serial entrepreneurs (Non-SEs). The central question is whether serial entrepreneurs are selected on persistent productive skill or on non-skill advantages such as preferential access to finance, because the two mechanisms have opposite implications for resource allocation: skill-driven serial entrepreneurship raises aggregate productivity, while favoritism-driven serial entrepreneurship generates misallocation.\n\nThe empirical foundation is two administrative datasets for Chinese firms: the Business Registry of China (SAIC), covering the universe of all firms since 1949 with a 2015 snapshot, used for the period 1995–2015; and the Inspection Database (SAIC), providing firm-level income-statement and balance-sheet data, used for 2008–2012 due to data quality constraints. The sample focuses on individually-owned firms (with the largest shareholder being a natural person), covering roughly 17 million entrepreneurs and 20 million firms by 2015. SE firms constitute approximately one-third of all individual-owned firms throughout the period and hold nearly half of all registered capital — making serial entrepreneurship quantitatively central to the Chinese private sector. SE firms have on average about twice the registered capital of Non-SE firms (e.g., 3.22 million yuan vs. 1.91 million yuan in 1995).\n\nTo organize empirical findings the authors develop a two-period Hopenhayn (1992)-style model with collateral-constrained borrowing (k ≤ λe, where k is capital and e is equity). The model generates two competing predictions. If TFP draws across firms started by the same entrepreneur are persistent (AR(1) with autocorrelation ρ), SEs outperform Non-SEs on TFP and the second firm outperforms the first. If instead some entrepreneurs are &amp;ldquo;favored&amp;rdquo; with a less binding collateral constraint (higher λ) and persistence is low, favored entrepreneurs enter more readily, pushing SE TFP below Non-SE TFP while installing more capital conditional on TFP.\n\nEmpirically, the average evidence favors persistent skills: 1st-SE firms are 9% more productive than Non-SE firms (within 2-digit industry, province, and year) and 2nd-SE firms are 18% more productive, both significant at the 1% level. In terms of assets, 1st-SE firms are 40% larger and 2nd-SE firms are 66% larger than Non-SE firms.\n\nThis average premium, however, conceals critical heterogeneity driven by industry-switching behavior. Two-thirds of SEs (67%) start the second firm in a different 2-digit input-output industry (switchers); one-third stay in the same industry (stayers). Stayers&amp;rsquo; 1st-SE and 2nd-SE firms are respectively 49% and 70% more productive than Non-SE firms — accounting for the entire average SE premium. Switchers&amp;rsquo; 1st-SE and 2nd-SE firms are respectively 9% and 11% less productive than Non-SE firms. Despite their TFP deficit, switchers hold at least 7% more capital in both firm generations than stayers. TFP persistence (autocorrelation of log TFP across 1st- and 2nd-SE firms) is twice as high for stayers (0.29) as for switchers (0.14), confirming the model&amp;rsquo;s key identifying assumption that within-industry persistence exceeds cross-industry persistence. The model interprets switchers&amp;rsquo; low-TFP/high-capital profile as the empirical signature of favored entrepreneurs.\n\nThe model further predicts that equity-constrained entrepreneurs should close the first firm when the second is substantially more productive (opportunity cost of capital). Consistently, 1st-SE firms that are shut when the 2nd starts have 32% lower TFP and 13% lower equity than those run concurrently; 2nd-SE firms operated non-concurrently have 8% higher TFP and 22% lower equity than those run alongside the first.\n\nBeyond learning, the paper documents two additional industry-choice motives for switchers. First, a diversification motive: a one-standard-deviation increase in the covariance of returns between the 1st- and 2nd-SE firm industries raises 2nd-SE TFP by 20%, consistent with entrepreneurs demanding a risk premium to enter correlated industries. Second, an input-output complementarity motive: serial entrepreneurs are significantly more likely to choose industries that are upstream-integrated (coefficient 0.46), downstream-integrated (0.47), or complementary (0.41) with the first industry (all significant at 1%), consistent with transaction-cost motives for co-owning trading partners.\n\nThe policy implication is that China&amp;rsquo;s private sector harbors both dynamism — embodied in highly productive stayer SEs driven by persistent skills — and distortion — embodied in low-productivity switcher SEs who enter and accumulate capital through preferential credit access. Since SE firms account for roughly one-third of all firms and nearly half of all capital, the aggregate productivity costs of favoritism-driven serial entrepreneurship are likely significant. Results apply to individually-owned private firms in China over 1995–2015 and may not extend to settings with more uniform financial markets or state-owned firm dynamics.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The paper does not use a natural experiment or instrumental variables for the main TFP comparisons. It relies on a structural model to interpret conditional correlations, with TFP measured relative to province-industry-year cell averages (2-digit industry, province, and year fixed effects). The theoretical identification comes from the fact that two distinct mechanisms — persistent skills and favoritism — generate opposite predictions on the joint TFP/capital relationship: skill dominance predicts higher TFP for SEs while favoritism predicts lower TFP combined with higher capital. The paper shows both signatures in data for distinct subgroups (stayers and switchers respectively), lending internal consistency. The concurrent/non-concurrent distinction provides an additional layer: the model predicts concurrency depends on equity and the TFP gap between firms, and the data confirm these predictions precisely (Table 7). The main threat is selection on unobservables: entrepreneurs who choose to start second firms may differ from non-SEs along dimensions not captured by the model, such as risk preferences, managerial talent, or social connections, and these could confound the TFP comparisons even within industry-province-year cells.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;Two mechanisms are posited. (1) Persistent skills (ρ &amp;gt; 0 in an AR(1) for TFP across an entrepreneur&amp;rsquo;s firms): positive selection makes SEs more productive and the 2nd-SE more productive than the 1st-SE. (2) Favoritism/credit access heterogeneity (heterogeneous collateral multiplier λ): favored entrepreneurs enter at lower TFP thresholds, so they are over-represented among SEs but have lower TFP and more capital conditional on TFP. The mechanisms are empirically distinguished by using industry switching as a proxy for favoritism. The learning model predicts low-first-period-TFP entrepreneurs switch industry (they do better by searching elsewhere), so favored individuals, who also have low TFP, should be concentrated among switchers. The data show switchers have both lower TFP than Non-SEs and more capital — a pattern only rationalized by favoritism. Stayers exhibit high TFP consistent with persistent skills. TFP persistence (autocorrelation) is twice as high within-industry (stayers, 0.29) as across-industry (switchers, 0.14), confirming the structural assumption separating the two mechanisms.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-across-se-types"&gt;Q3. What heterogeneity is documented across SE types?&lt;/h3&gt;
&lt;p&gt;First, stayer vs. switcher heterogeneity is the dominant finding: stayers&amp;rsquo; 1st-SE TFP is 49% above Non-SE and 2nd-SE TFP is 70% above Non-SE; switchers&amp;rsquo; 1st-SE TFP is 9% below Non-SE and 2nd-SE TFP is 11% below Non-SE. Switchers have more assets, equity, and registered capital than stayers despite lower TFP (at least 7% more capital). Second, concurrent vs. non-concurrent heterogeneity: 47.5% of SE firms in the 2008–2012 sample are operated concurrently. Non-concurrent 1st-SE firms have 32% lower TFP and 13% lower equity; non-concurrent 2nd-SE firms have 8% higher TFP and 22% lower equity, consistent with equity-constrained optimal capital reallocation. Third, generational heterogeneity: 2nd-SE firms are consistently larger and more productive than 1st-SE firms across all measures (TFP +18% vs. +9%; assets +66% vs. +40%), consistent with high ρ and positive selection into the second firm. Fourth, geographic stability: 72.3% of SEs locate the 2nd firm in the same prefecture as the first, suggesting local knowledge and networks matter for firm creation.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-and-data-restrictions-are-applied"&gt;Q4. What robustness checks and data restrictions are applied?&lt;/h3&gt;
&lt;p&gt;The paper trims the top and bottom 1% of assets and TFP before computing relative TFP. It excludes the 2007–2008 period from return-to-capital calculations (financial crisis concern). It excludes post-2014 registry data because of a registry reform that inflated new registrations and depressed measured exit. It confirms the covariance-TFP diversification result holds when including SE firms not run concurrently. It excludes entrepreneurs who established more than 20 firms (542 individuals, 188,266 firms) to avoid chain-store effects. The paper does not report instrumental-variable estimates, placebo tests, or alternative TFP measures as formal robustness exercises.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Prior work on serial entrepreneurship (Holmes and Schmitz 1990, 1995; Lafontaine and Shaw 2016 for US; Rocha et al. 2015 for Portugal; Shaw and Sørensen 2019, 2022 for Denmark; Felix et al. 2021) uniformly finds SEs are more productive or larger than Non-SEs and attributes this to ability or learning. This paper confirms the average finding but is the first to demonstrate that the premium fully disappears and reverses for industry switchers, and to link this reversal to capital market distortions and favoritism rather than skill. The use of a comprehensive universe of firms (not manufacturing-only or survey-based samples) distinguishes it empirically. The misallocation literature (Hsieh and Klenow 2009; Buera, Kaboski, Shin 2011; Midrigan and Xu 2014; Moll 2014) analyzes distortions across all firms but does not analyze serial entrepreneurship. Song, Storesletten and Zilibotti (2011) and Hsieh and Song (2015) focus on state vs. private sector differences; this paper shows distortions exist within the private sector among individual-owned firms. Contemporaneous work by Shaw and Sørensen (2022) on Denmark documents similar properties of SE firms to the Chinese average findings.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-models-key-structural-propositions"&gt;Q6. What are the model&amp;rsquo;s key structural propositions?&lt;/h3&gt;
&lt;p&gt;Proposition 1: entrepreneurs enter iff TFP z ≥ z*(e), where the entry threshold is decreasing in equity e. Proposition 2: without financial frictions and with ρ &amp;gt; 0, 1st-SE and 2nd-SE firms have higher expected TFP than Non-SE, and 2nd-SE &amp;gt; 1st-SE for sufficiently large ρ. Proposition 3: with frictions, the 2nd-period entry threshold Z(z1, e) is increasing in z1 (opportunity cost of first firm&amp;rsquo;s capital) and decreasing in e. Proposition 4: with frictions and Assumption 1 (equity monotone in TFP) and sufficiently large ρ, SE firms are more productive than Non-SE. Proposition 5: with ρ = 0 and heterogeneous λ, favored entrepreneurs are over-represented among SEs, which then have lower average TFP but more capital conditional on TFP. Proposition 6: concurrent operation is increasing in equity and decreasing in |z2 − z1|. Proposition 7: entrepreneurs stay in the same industry iff 1st-firm TFP exceeds the unconditional mean; stayers have higher TFP than switchers for both SE firms. Proposition 8: with a risk diversification motive, the probability of choosing industry s&amp;rsquo; for the 2nd firm is decreasing in Cov(δs&amp;rsquo;, δs); conditional on choosing s&amp;rsquo;, 2nd-SE TFP is increasing in Cov(δs&amp;rsquo;, δs).&lt;/p&gt;
&lt;h3 id="q7-what-are-the-diversification-and-input-output-linkage-findings"&gt;Q7. What are the diversification and input-output linkage findings?&lt;/h3&gt;
&lt;p&gt;For diversification, the authors construct an industry-level return-on-assets covariance matrix using 2010–2012 Inspection Data (excluding the financial crisis year). A one-standard-deviation increase in the covariance of returns between 1st and 2nd SE firm industries increases 2nd-SE TFP by 20% (significant at 1%), meaning entrepreneurs require a TFP risk premium to enter a correlated industry. In the excess-probability regression for industry choice, the covariance has a coefficient of -0.11 (significant at 1%), confirming switchers prefer industries negatively correlated with their first industry. For linkages, using 2007 Chinese Input-Output tables and Fan-Lang (2000) methodology, the authors find excess probability of industry choice is significantly higher for downstream-integrated industries (0.47), upstream-integrated industries (0.46), and complementary industries (0.41), all at the 1% level in a joint regression. These results hold controlling for 1st-SE industry fixed effects and year of establishment.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper implies that China&amp;rsquo;s private sector suffers from a specific type of misallocation: entrepreneurs with preferential credit access (favored individuals, proxied by industry switchers) establish and expand firms despite lower productivity, crowding out more productive entrepreneurs. Reducing distortions in credit access — leveling the collateral constraint across entrepreneurs — would shift resources toward skill-driven serial entrepreneurs (stayers) and raise aggregate productivity. The scale of the problem is meaningful: SE firms hold roughly half of all capital in the individual-owner sector. Scope conditions: these findings apply to individually-owned private firms in China during 1995–2015, a period characterized by rapid private-sector growth, underdeveloped financial markets, and significant political-economic favoritism. The results abstract from cross-regional and cross-industry variation in financial frictions; if such variation matters (as Brandt, Kambourov and Storesletten 2023 suggest), the aggregate distortion estimates could differ. The paper does not quantify the aggregate TFP losses from misallocation in a counterfactual exercise.&lt;/p&gt;
&lt;h3 id="q9-what-data-limitations-and-caveats-apply"&gt;Q9. What data limitations and caveats apply?&lt;/h3&gt;
&lt;p&gt;The Inspection Data lack employment information, so the authors impute labor input from the labor first-order condition under competitive wages within province-industry-year cells — a valid proxy only if factor market prices are equalized within cells. Revenue is used as a proxy for value added, valid only if intermediate input shares are constant within industry-province-year cells. The registry snapshot is from end-2015, so ownership history must be inferred; the authors note that for over 80% of individual-owned firms the founding owner coincides with the exit-period or current owner. Post-2014 data are excluded due to registry reform contamination. The analysis excludes entrepreneurs who established more than 20 firms (542 individuals, 188,266 firms) to avoid chain-store effects. The analysis excludes SEs who start a 2nd firm through an enterprise they control (expanding the definition would add 300,400 such cases). Concurrent/non-concurrent classification uses the Inspection Data&amp;rsquo;s 2008–2012 window, which may misclassify some firms. The TFP measure is relative within province-industry-year cells, so cross-cell TFP comparisons are not made.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Serial entrepreneur (SE)&lt;/strong&gt;: In this paper, an individual investor who is or has been the largest shareholder in at least two separate firms over the observation period, not necessarily concurrently; 1st-SE refers to the entrepreneur&amp;rsquo;s first firm and 2nd-SE to all subsequent firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-serial entrepreneur (Non-SE)&lt;/strong&gt;: An individual investor who is or was the largest shareholder in exactly one firm over the entire observation window; the benchmark category for TFP and size comparisons.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stayer&lt;/strong&gt;: A serial entrepreneur whose 2nd-SE firm is in the same 2-digit input-output industry as the 1st-SE firm; interpreted in the model as evidence of high industry-specific comparative advantage and high TFP persistence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Switcher&lt;/strong&gt;: A serial entrepreneur whose 2nd-SE firm is in a different 2-digit input-output industry from the 1st-SE firm; interpreted as evidence of either low first-period TFP (learning/Jovanovic motive) or preferential credit access (favoritism motive); empirically identified by lower TFP than Non-SEs combined with more capital.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Favored entrepreneur&lt;/strong&gt;: In the model, an entrepreneur with a less binding collateral constraint (higher λ), representing individuals with preferential access to bank credit or other non-skill advantages; they enter at lower TFP thresholds, are over-represented among SEs, and display the signature pattern of lower TFP combined with more capital conditional on TFP.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Collateral constraint&lt;/strong&gt;: A borrowing limit of the form k ≤ λe, where k is installed capital, e is equity, and λ ≥ 1 is the collateral multiplier; the central financial friction in the model, generating the observed co-movement between TFP, assets, and debt-equity ratios in the data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Concurrent vs. non-concurrent SE operation&lt;/strong&gt;: Whether the entrepreneur&amp;rsquo;s 1st and 2nd firms are both operating simultaneously (concurrent) or the 1st firm is closed before or when the 2nd begins (non-concurrent); the model predicts non-concurrent operation is optimal when equity is scarce and the TFP gap between firms is large, rationalizing the observed pattern that non-concurrent 2nd-SE firms have higher TFP and lower equity.&lt;/p&gt;</description></item><item><title>Sources of rising student debt in the U.S.: College costs, wage inequality, and delinquency</title><link>https://macropaperwarehouse.com/papers/sources-of-rising-student-debt-in-the-u.s.-college-costs-wage-inequality-and-delinquency/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/sources-of-rising-student-debt-in-the-u.s.-college-costs-wage-inequality-and-delinquency/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;U.S. outstanding student debt rose roughly 20-fold, from about $50 billion in 1985 to nearly $1 trillion in 2014 (about 7% of GDP), making it the second-largest form of household debt after mortgages. Kim and Kim ask how much of this growth in &lt;em&gt;undergraduate&lt;/em&gt; loans can be explained by three forces: rising college costs, rising wage inequality, and the option to become delinquent. They build a partial-equilibrium incomplete-markets overlapping-generations (OLG) model with a three-stage life cycle (college, work, retirement, ages 18-85, annual periods). Individuals are endowed with heterogeneous ability (decile distribution of demeaned log AFQT80) and correlated parental transfers, and choose college attendance, government student-loan borrowing, and whether to repay or become delinquent (90+ days past due, carrying a skill-specific utility cost). College lasts 4 years; lower-ability students face a dropout probability at year 2 (aggregate enrollment-to-non-completion is ~54%). Loans follow a fixed 10-year repayment schedule (nT=10), accrue interest at rb=6.1% (risk-free r=3%), with a cumulative borrowing limit of $23,000 (raised to $31,000 from 2008) and a cap of 70% of tuition.&lt;/p&gt;
&lt;p&gt;The model is calibrated to the 1985 steady state, mainly with NLSY79 (plus NLSY97 for transfers/costs and PSID for the experience premium and wage-shock process). Transitional dynamics 1985-2014 feed in three time-varying inputs: rising college costs (net cost rises from $5,859 in 1985 to $12,000 in 2014), rising wage inequality (persistent-shock variance rises from 0.015 to 0.03 and transitory from 0.05 to 0.08; college wage premium from 1.2 to 1.37; skilled ability premium from 0.89 to 1.33; shock persistence ρ=0.9791), and a growing preference for college (a declining psychic cost calibrated to reproduce rising attainment).&lt;/p&gt;
&lt;p&gt;Main results: the benchmark economy raises aggregate undergraduate debt from $37 billion (1985) to $351 billion (2014), a $314 billion increase that explains about 64% of the observed U.S. rise — without being calibrated to the debt increase. Rising college costs are the primary driver of higher borrowing; rising income risk and declining average student ability drive higher delinquency (the aggregate delinquency rate more than triples 1985-2014; 16% of borrowers delinquent in 2014). In a decomposition (Table 3), fixing college costs cuts the debt rise to +$33B; fixing ability premia leaves it roughly unchanged (+$317B); fixing the college wage premium lowers it by $49B (to +$265B); and fixing wage-shock variances &lt;em&gt;raises&lt;/em&gt; it to +$418B (less risk means less delinquency but more borrowing). Removing the delinquency option entirely cuts the debt rise to $178 billion, so delinquency accounts for about 43% of the transitional increase. Delinquency works through a mechanical channel (missed payments plus accrued interest) and an incentive channel (delinquency as insurance encourages borrowing, the Domar-Musgrave effect); roughly one-third of the benchmark/no-delinquency gap is mechanical and two-thirds incentive. Finally, an income-driven repayment (IDR) plan (10% of discretionary income) cuts delinquency from 5.0% to 2.2% and slows debt growth to a $169 billion rise over the transition, because IDR substitutes for delinquency as insurance.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-model-and-the-identificationquantification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the model and the identification/quantification strategy, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;It is a partial-equilibrium incomplete-markets OLG model solved as two steady states (1985 and 2014) with a transition path. Identification of the aggregate-debt contribution is not econometric but quantitative: the model is calibrated to 1985 cross-sectional moments (and a few transition-path moments) WITHOUT targeting the aggregate debt increase, then exogenous time-varying inputs (college costs, wage inequality, college preference) are fed in and the resulting debt path is compared to data, explaining ~64% of the rise. The main threats are: (i) the model is partial equilibrium, taking costs/inequality/preferences as exogenous (general-equilibrium feedback, e.g. tuition responding to inequality per Cai-Heathcote 2022, is abstracted from); (ii) the residual 36% is unexplained and could reflect omitted forces such as private loans, for-profit institutions, or graduate-school spillovers; (iii) the &amp;lsquo;preference for college&amp;rsquo; is a reduced-form declining psychic cost that absorbs many unmodeled drivers (job amenities, over-optimism about graduation) rather than being separately identified.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-channels-through-which-delinquency-raises-debt-and-how-are-they-distinguished"&gt;Q2. What are the two channels through which delinquency raises debt, and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;The mechanical channel: missed scheduled payments plus accrued interest are added directly to the outstanding balance. The incentive channel: the option to delay payment acts as insurance against adverse post-college income shocks, encouraging students to borrow more ex ante (the Domar-Musgrave effect). They are separated with a &amp;lsquo;mechanical effect counterfactual&amp;rsquo; that removes delinquency but holds borrowing fixed at benchmark levels: the gap between benchmark and this counterfactual is the mechanical effect, and the gap between the mechanical counterfactual and the full no-delinquency economy is the incentive effect. The incentive effect dominates — roughly two-thirds of the benchmark/no-delinquency gap — because the mechanical effect operates only through the small share of delinquent borrowers (16% in 2014), while the incentive effect shapes all college students&amp;rsquo; borrowing. The incentive channel grows over time as income risk rises.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Borrowing increases with ability and (weakly) with parental transfers, driven by consumption smoothing: high-ability individuals anticipate higher lifetime earnings and borrow more against future income. Notably, in the 1985 simulation, average earnings during college exceed college costs across all ability groups, so most students could self-finance but still borrow. Dropout probability declines sharply with ability (so ~54% of enrollees do not complete). Delinquency rates differ by skill: 7% for college graduates vs 25% for college dropouts in 2010 (calibration targets). The stronger college preference draws more low-ability students into college over time, lowering average student ability and raising delinquency. Under IDR, the rise in borrowing participation (34%-&amp;gt;40%) is driven primarily by low-ability students.&lt;/p&gt;
&lt;h3 id="q4-what-robustnessvalidation-checks-are-run"&gt;Q4. What robustness/validation checks are run?&lt;/h3&gt;
&lt;p&gt;Validation (not targeted): the model reproduces the rising trend in average annual borrowing 1993-2014 (NPSAS), the cross-sectional borrowing distribution by ability tercile and parental-transfer quartile in 1997 (NLSY97), the more-than-tripling of the aggregate 90+ day delinquency rate (FRBNY), and ~8% of borrowers behind on payments 10 years after graduation (Table D1). It also replicates the untargeted population distribution across ability/transfer cells. Robustness: results are stable with 10 or more ability grid points; the implied ~12% decline in average student ability between the 1960s and 1990s cohorts is consistent with Hendricks-Schoellman (2014). An alternative delinquency definition using 270-day default plus wage garnishment (Appendix C) yields similar aggregate effects, with delinquency explaining about 33% of the debt increase (vs 43% in the 90-day benchmark). A weakness flagged by the authors: the model generates flat college costs across parental-transfer quartiles and so misses the non-monotonic (U-shaped) cost pattern in the data, because ability and transfers are positively correlated.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds directly on Abbott, Gallipoli, Meghir, Violante (2019), whose framework of government grants/loans and college attainment it extends by adding an endogenous delinquency choice on student debt to capture debt amplification. It differs from Ionescu (2008, 2009), which evaluate specific loan-policy reforms (lock-in interest, flexible repayment, eligibility) for enrollment/default, by focusing on the &lt;em&gt;dynamics of the aggregate debt stock&lt;/em&gt; rather than direct policy evaluation. It connects to the credit-constraints/family-income literature (Belley-Lochner 2007, Lochner-Monge-Naranjo 2011, Carneiro-Heckman 2002, Keane-Wolpin 2001) by jointly modeling parental transfers and borrowing, and to the repayment/default-determinants literature (Looney-Yannelis 2015, Lochner-Monge-Naranjo 2015, Deming-Goldin-Katz 2012). It remains agnostic about private loans (only 6-7% of outstanding debt and structurally different, per Ionescu-Simpson 2016).&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;IDR is identified as an effective instrument for managing student-loan burdens: capping payments at 10% of discretionary income reduces delinquency sharply (5.0%-&amp;gt;2.2% in steady state) and slows the transitional debt rise from $314B to $169B, because formal repayment flexibility substitutes for informal insurance via delinquency. Scope conditions: IDR also &lt;em&gt;increases&lt;/em&gt; loan participation (34%-&amp;gt;40%), so the slowdown in debt comes from the delinquency-reduction effect dominating the borrowing-increase effect; in steady state total debt falls only $3 billion, the larger effect being on the transition. The result holds in partial equilibrium with no model re-calibration and assumes borrowers choose labor supply anticipating 10%-of-income repayment; general-equilibrium and fiscal-cost (loan-forgiveness) implications are not modeled. Take-up was low over 1985-2014 (11% of undergraduate borrowers in 2010, 24% by 2017), so IDR is treated as a forward-looking policy extension rather than a driver of the historical debt rise.&lt;/p&gt;
&lt;h3 id="q7-what-other-significant-findings-or-caveats-appear"&gt;Q7. What other significant findings or caveats appear?&lt;/h3&gt;
&lt;p&gt;Fixing wage-shock variances counterintuitively raises debt (+$418B vs +$314B) because lower income risk reduces delinquency but encourages more borrowing — illustrating that inequality&amp;rsquo;s net effect on debt runs partly through the insurance/incentive channel rather than just borrowing need. The annual flow of newly delinquent debt rose from about $200 million (1985) to $5.5 billion (2015) in the benchmark (Figure D9). The number of borrowers and average debt per borrower both rose (borrowers from 8% of population in 2004 to 14% in 2014; average debt per borrower from $15,106 to $21,677). The model abstracts from endogenous dropout during college (no idiosyncratic risk in college) and from graduate loans, focusing on undergraduate debt as the largest component.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Taxation and Entrepreneurship in the United States</title><link>https://macropaperwarehouse.com/papers/taxation-and-entrepreneurship-in-the-united-states/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/taxation-and-entrepreneurship-in-the-united-states/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how the level and progressivity of personal income taxes shape entrepreneurial activity in the United States, contributing empirical evidence, theoretical intuition, and a structural quantitative evaluation. The motivation is both descriptive — entrepreneurs own more than 40% of total capital and hire more than half of private-sector workers, yet their share of the population varies substantially across states and time — and normative, given growing policy interest in more redistributive taxation. The central question is whether a more progressive tax system, which simultaneously reduces the risk and the return to entrepreneurship, produces more or fewer entrepreneurs in practice.&lt;/p&gt;
&lt;p&gt;The empirical analysis draws on CPS microdata from 1962 to 2019 (entrepreneurs defined as households where the head or spouse is self-employed, averaging 11.7% of the national population), County Business Pattern data from 1986 to 2018, Business Dynamics Statistics, and NBER TAXSIM. Tax measures — both average tax rates at the 50th, 90th, and 95th percentiles of the national earnings distribution and a parametric Benabou (2002) tax function with a level parameter theta_0 and a progressivity parameter theta_1 — are constructed by applying TAXSIM to a fixed 2010 CPS cross-section across all 51 state-year cells from 1977 to 2019, thereby limiting endogeneity from tax-code-induced changes in the observed income distribution. The benchmark panel regression includes state and year fixed effects, state-level economic and demographic controls, lagged local business cycle variables, and local non-linear time trends; the benchmark outcome is measured two years after the tax change. Instrumental variables — lagged state tax rates plus contemporaneous federal rates — are used to further address endogeneity.&lt;/p&gt;
&lt;p&gt;The core empirical findings are strongly negative across all measures of entrepreneurship and all tax measures. A one-percentage-point increase in the average tax rate at median income reduces the number of entrepreneurs by 4.5% (coefficient -0.0449, significant at 1%); a one-standard-deviation increase in that tax rate (about 2.35 percentage points) implies roughly 9.7% fewer entrepreneurs. Negative effects also hold for college-educated entrepreneurs and for firm-side proxies (number of small establishments, employment at small establishments). For tax progressivity, holding tax level constant, a one-percentage-point increase in the average tax rate at twice average earnings reduces the number of entrepreneurs by about 15%. Using the parametric progressivity measure, an increase in theta_1 of 0.01 (about 60% of the cross-state standard deviation) reduces the total number of entrepreneurs by approximately 10% and the number of small establishments by about 2.5%. These results hold under additional lagged controls, different horizons (negative and significant through about nine years for the count of entrepreneurs, more persistent for firm-side measures), and IV estimation (IV magnitudes are one to three times larger than OLS, with first-stage F-statistics of 136 and 112 for the progressivity instrument). A subsample analysis around major federal tax reform years (1988, 1991–1993, 2001) finds consistent signs but smaller and noisier estimates given the reduced sample size.&lt;/p&gt;
&lt;p&gt;To explain these patterns, the paper develops a life-cycle overlapping-generations incomplete-markets model in the spirit of Quadrini (2000) and Cagetti and De Nardi (2006). Households are heterogeneous in age, innate ability, idiosyncratic labor and entrepreneurial productivity shocks, risk aversion (distributed uniformly over three values), and asset holdings. Entrepreneurs face a collateral constraint (capital bounded by theta times assets), a fixed operating cost each period, and a switching cost when exiting to wage employment. The same progressive tax function applies to both workers and entrepreneurs. The model is calibrated to U.S. data: exogenous parameters include an inverse Frisch elasticity of 1, labor productivity persistence of 0.929 and standard deviation of 0.227 (from Chang and Kim 2007), a 45-year working life, and returns to scale in entrepreneurship of 0.85. Eight parameters — including the discount factor, entrepreneurial productivity persistence and dispersion, operating cost, switching cost, and risk-aversion dispersion — are estimated via simulated method of moments, matching 21 moments including the entrepreneur population share, income and wealth shares of entrepreneurs, fraction of entrepreneurs with negative profits, and aggregate wealth distribution. The model matches the data well on targeted and untargeted moments.&lt;/p&gt;
&lt;p&gt;The main structural counterfactual holds average tax rates constant and varies progressivity. Converting to a flat tax (theta_1 = 0) increases the number of entrepreneurs by about 15% in general equilibrium. Aggregate output rises by about 11% and the capital stock falls by about 27% when progressivity doubles from 0.13 to 0.26 (relative to the benchmark of theta_1 = 0.13). The return effect — more progressive taxes compress the expected return to entrepreneurship relative to wage work — quantitatively dominates the insurance effect (more progressive taxes reduce the variance of entrepreneurial income). The distributional analysis shows that medium-productivity entrepreneurs are more sensitive to tax changes than high-productivity ones; older, wealthier entrepreneurs are also more responsive. For welfare, the socially optimal progressivity level — measured by ex-ante expected lifetime welfare of unborn agents in steady state — is theta_1 = 0.109, only about 16% less progressive than the current U.S. benchmark of 0.13. The welfare gains from this reform are described as tiny. The welfare-optimal policy reflects the trade-off between efficiency losses (from reduced entrepreneurship and output) and distributional gains (from redistribution to below-average-income households, who benefit from more progressive taxation). Raising the average tax level while holding progressivity constant also reduces output and capital, with capital falling by roughly 40% and output by about 10% when the level parameter doubles; these effects interact with progressivity in non-linear ways captured only through the structural model.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-in-the-empirical-analysis-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy in the empirical analysis and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The benchmark strategy is a state-year panel regression with state and year fixed effects, state-level economic and demographic controls (real GDP per capita, sector employment shares), and lagged local GDP growth rates and unemployment rates over four years before the tax measure. The dependent variable is measured two years after the tax change to allow recognition lags. IV instruments are constructed as the sum of the lagged (by two years) state tax rate at the relevant income percentile and the current federal marginal tax rate at that percentile, following Akcigit et al. (2018); for progressivity, lagged theta_1 and theta_0 are used as instruments, with first-stage F-statistics of 136 and 112 respectively, ruling out weak instruments. A further alternative IV constructs hypothetical tax parameters by applying current federal rates to state-level rates lagged by two years via TAXSIM. Main threats are (1) endogeneity of state tax policy to local economic conditions — addressed through the rich set of lagged business cycle controls, state-specific quadratic trends, and IV; (2) income-composition endogeneity in estimating the tax function — addressed by fixing the CPS 2010 sample and scaling incomes by average wage growth rather than using the contemporaneous distribution; (3) short sample periods around major reform years, which make the reform-event analysis underpowered.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-through-which-taxes-affect-entrepreneurial-choice-and-how-are-they-distinguished"&gt;Q2. What are the main mechanisms through which taxes affect entrepreneurial choice, and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;The paper identifies two opposing forces from greater tax progressivity. The return effect: higher progressivity reduces the average after-tax payoff to entrepreneurship, because entrepreneurs earn above-average incomes and the progressive schedule compresses post-tax profits relative to wages. The insurance effect: higher progressivity also reduces the variance of after-tax entrepreneurial income, making entrepreneurship less risky and potentially more attractive to risk-averse agents. The simple theoretical models (mean-variance utility with lognormal profits and CRRA utility) show that the sign of the net effect is theoretically ambiguous. In the quantitative model — and in the data — the return effect dominates: flatter taxes raise entrepreneurial entry. The two effects are separated analytically in the simple model (Section 4) and quantitatively in the structural model by examining partial-equilibrium versus general-equilibrium effects and by isolating the capital demand response (sensitive to progressivity) from the labor demand response (less sensitive).&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-across-entrepreneurs-and-along-the-life-cycle-is-documented"&gt;Q3. What heterogeneity across entrepreneurs and along the life cycle is documented?&lt;/h3&gt;
&lt;p&gt;Empirically, the negative tax effect is larger for college-educated entrepreneurs than for non-college entrepreneurs when measured by high-income tax rates (90th and 95th percentiles), consistent with higher-educated entrepreneurs having higher incomes. In the structural model, medium-productivity entrepreneurs lose the most when progressivity rises: when theta_1 doubles, the medium-productivity group&amp;rsquo;s share falls by 0.84 percentage points from a base of 9.08%, while the high-productivity group falls by only 0.11 points from 3.47%. Older and wealthier households are more sensitive to progressivity changes because the return effect matters more relative to the insurance effect for those who have accumulated wealth. Risk aversion heterogeneity (modeled as uniform dispersion around 2.5) affects saving and occupational choice; more risk-averse households are more sensitive to the variance reduction from progressive taxes, but the model shows this does not reverse the dominance of the return effect in aggregate.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The robustness battery includes: (1) adding state-specific quadratic time trends and longer lags of local business cycle variables; (2) two IV strategies — lagged state tax rates plus current federal rates, and hypothetical tax measures constructed from TAXSIM with lagged state and current federal components; (3) controlling for lagged entrepreneurial activity levels (log number of entrepreneurs and establishments lagged two years); (4) examining effects at horizons from t+0 to t+10 via local projection methods, finding effects most pronounced in the short run and diminishing over about nine years for entrepreneur counts but more persistent for establishment and employment measures; (5) restricting the sample to years around major federal tax reforms (1988, 1991–1993, 2001) and finding consistent negative signs even though magnitudes are weaker given the smaller sample; (6) using alternative measures of progressivity (differences between tax rates at multiples of average earnings) as a robustness check on the parametric theta_1 measure; (7) structural model sensitivity analysis varying each estimated parameter individually to confirm monotonic identification of moments.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-prior-empirical-and-structural-work"&gt;Q5. How does this paper relate to and differ from prior empirical and structural work?&lt;/h3&gt;
&lt;p&gt;Empirically, it extends Gentry and Hubbard (2000), who used PSID data 1978–1993 to document that progressive marginal rates discourage self-employment, and Cullen and Gordon (2007), who used IRS cross-sectional data to study the role of tax incentives in business formation. The current paper uses a much larger micro-level dataset (CPS, CBP, BDS), covers both cross-sectional and time-series variation across all U.S. states from 1962 to 2019, examines a broader set of entrepreneurial outcomes (count, employment, establishment dynamics), and controls rigorously for local trends and business cycles. Structurally, it is in the tradition of Quadrini (2000), Cagetti and De Nardi (2006), and Kitao (2008), but uniquely combines a life-cycle OLG framework with empirically estimated tax progressivity and a novel SMM estimation of key entrepreneurial parameters including risk-aversion dispersion. Unlike Meh (2005), which studies switching from progressive to proportional tax in a similar model, this paper brings empirical discipline via state-level identification and explicitly estimates the optimal progressivity. Unlike Brüggemann (2017), which focuses on optimal top marginal rates, this paper studies the full distribution and links it to state-level quasi-experimental evidence. Scheuer (2014) studies optimal taxation with endogenous entry theoretically; this paper complements that with quantitative general-equilibrium analysis.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy implication is that tax progressivity has a quantitatively large negative effect on entrepreneurship and output: converting to a flat tax (holding average tax revenue constant) would increase the number of entrepreneurs by about 15% and GDP by about 11%. However, the welfare-optimal progressivity is only marginally less than the current U.S. level (optimal theta_1 of 0.109 versus benchmark of 0.13, about 16% less progressive), implying the welfare gains from flattening taxes are tiny. This is because redistribution from high-income entrepreneurs to below-average-income workers and retirees is welfare-improving even as it reduces aggregate output. The results hold in both general equilibrium (where wages and interest rates adjust) and in partial equilibrium (more relevant for state-level comparisons, where PE effects are somewhat stronger). The scope conditions include: the model abstracts from age-dependent taxation, occupational-specific tax treatment, endogenous human capital accumulation by entrepreneurs, wealth taxes, and the distinction between corporate and pass-through taxation. These omitted features could alter the optimal progressivity result.&lt;/p&gt;
&lt;h3 id="q7-what-do-the-general-equilibrium-versus-partial-equilibrium-comparisons-reveal"&gt;Q7. What do the general equilibrium versus partial equilibrium comparisons reveal?&lt;/h3&gt;
&lt;p&gt;Partial equilibrium effects (constant wages and interest rates, approximating the small open economy view of U.S. states) are somewhat stronger than general equilibrium effects. This is consistent with the empirical panel estimates, which more closely correspond to PE since state economies face roughly fixed factor prices from the national market. When progressivity doubles in PE (adjusting average tax), the entrepreneur share falls more than in GE, and the optimal progressivity in PE is higher than in GE because in GE there is an additional channel: lower capital stock from reduced entrepreneurship depresses wages, imposing an additional cost on workers that is absent in PE. This comparison validates using PE as the interpretive benchmark for the empirical regressions.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-model-say-about-the-interaction-between-tax-level-and-tax-progressivity"&gt;Q8. What does the model say about the interaction between tax level and tax progressivity?&lt;/h3&gt;
&lt;p&gt;The model reveals a non-linear interaction that cannot be separated in empirical analysis. When tax progressivity is held at zero (flat tax), the entrepreneur share declines smoothly as the tax level rises. At benchmark progressivity, the entrepreneur share exhibits a non-monotonic relationship with the level: for very low tax levels the share is high, it falls as taxes rise, but at sufficiently high levels the entrepreneur share may rise again because workers&amp;rsquo; wealth effects lead to higher labor supply, partially offsetting the dampening of entrepreneurial returns. At doubled progressivity, the non-monotonicity is more pronounced. Tax revenue also exhibits a Laffer-curve pattern with respect to the level parameter across all progressivity scenarios, though this is not the paper&amp;rsquo;s primary focus.&lt;/p&gt;
&lt;h3 id="q9-what-quantitative-moments-does-the-calibrated-model-match-and-where-does-it-fall-short"&gt;Q9. What quantitative moments does the calibrated model match, and where does it fall short?&lt;/h3&gt;
&lt;p&gt;The model matches an aggregate capital-to-output ratio of 2.716 (data: 2.650), entrepreneur population share of 12.6% (data: 12.1%), employment hired by entrepreneurs of 55.9% (data: 56.0%), share of entrepreneurs with negative profits of 12.2% (data: 11.0%), average exit rate of 9.4% (data: 17.0%, a notable miss), average age of entrepreneurs of 44.4 (data: 49.2, another miss), entrepreneur income and wealth shares across the distribution, and top household wealth shares. The model overshoots capital and wealth shares for the top decile relative to data but matches the middle of the distribution well. The average age and exit rate mismatches are acknowledged; the operating-cost and switching-cost parameters are the primary levers for these, and the paper notes that exit costs (rather than entry costs) are more effective at generating entrepreneurs with negative profits.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Tax progressivity (theta_1)&lt;/strong&gt;: The progressivity parameter in the Benabou (2002) tax function ya/AE = theta_0*(y/AE)^(1-theta_1): a higher theta_1 means after-tax income rises less than proportionally with pre-tax income, implying marginal rates increase with income. In the paper&amp;rsquo;s measure, theta_1 = 0 is a flat tax and the U.S. benchmark is estimated at 0.13. Progressivity is measured separately from the average tax level (controlled by theta_0), allowing the two to vary independently in both empirics and counterfactuals.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Return effect vs. insurance effect&lt;/strong&gt;: The two opposing forces through which tax progressivity affects entrepreneurial choice. The return effect is the compression of average after-tax entrepreneurial profits relative to wages — since entrepreneurs earn above-average incomes, progressive taxes reduce the relative net payoff to entrepreneurship. The insurance effect is the reduction in after-tax income variance for entrepreneurs — progressive taxes act as partial insurance against bad profit realizations. The paper finds the return effect quantitatively dominates in both the simple theoretical models and the calibrated quantitative model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Collateral constraint&lt;/strong&gt;: The restriction k &amp;lt;= Theta*a in the model, where k is the entrepreneur&amp;rsquo;s capital input and a is her asset holdings. This models credit market frictions: an entrepreneur can borrow and invest no more than Theta - 1 times her own wealth in the business. Set to Theta = 0.35 in calibration (following Midrigan and Xu 2014), this constraint links entrepreneurial capital demand to wealth accumulation, making the tax-wealth-capital nexus a central quantitative mechanism.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Entrepreneur switching cost (Gamma_s)&lt;/strong&gt;: A cost paid by an entrepreneur who exits to wage employment in the current period. In the calibrated model, Gamma_s = 1.005 (in units of average earnings). This switching cost generates inertia in occupational choice: entrepreneurs with temporarily low productivity may remain rather than exit, generating the empirical share of entrepreneurs with zero or negative profits. It also contributes to life-cycle patterns of entrepreneurship by raising the bar for exit among older, wealthier incumbents.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ex-ante welfare measure&lt;/strong&gt;: The paper&amp;rsquo;s social welfare criterion: the expected lifetime utility of an unborn agent at the beginning of life (age 1), averaging over all initial states (innate ability, initial labor and entrepreneurial productivity draws), and taking the maximum of the worker and entrepreneur value functions. This differs from ex-post welfare (which conditions on realized occupational choice) and is the basis for the optimal tax progressivity calculation. The welfare-maximizing theta_1 = 0.109 uses this criterion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Progressivity wedge (PW)&lt;/strong&gt;: A summary statistic for tax progressivity defined as PW(y1, y2) = 1 - (1 - T&amp;rsquo;(y2))/(1 - T&amp;rsquo;(y1)) for pre-tax incomes y1 &amp;lt; y2. Under the Benabou tax function, the wedge is uniquely determined by theta_1 and equals zero for a flat tax, approaching 1 as the marginal tax rate at the higher income approaches 100%. This measure allows comparison of progressivity across tax systems independently of the level of tax rates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simulated method of moments (SMM)&lt;/strong&gt;: The estimation procedure used for eight model parameters (discount factor beta, entrepreneurial productivity persistence rho_z and dispersion sigma_z, operating cost Gamma_f, switching cost Gamma_s, labor disutility chi, aggregate productivity A, and risk-aversion dispersion sigma_U). The procedure minimizes the weighted distance between 21 model-implied moments and their data counterparts, with a diagonal weighting matrix that puts larger weights on the aggregate capital-to-output ratio and the overall entrepreneur population share.&lt;/p&gt;</description></item><item><title>Taxation of Capital: Capital Levies and Commitment</title><link>https://macropaperwarehouse.com/papers/taxation-of-capital-capital-levies-and-commitment/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/taxation-of-capital-capital-levies-and-commitment/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Barro and Chari (2024) revisit the long-standing debate over optimal capital income taxation, unifying the Chamley-Judd zero-tax result, the Straub-Werning positive-tax amendment, and the Chari-Nicolini-Teles (2020) commitment-based framework into a single coherent analysis centered on the treatment of the &amp;ldquo;period-zero problem.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The research question is fundamental: under what commitment assumptions is the optimal long-run tax rate on capital income zero, positive, or negative, and does optimal policy require special treatment of the initial period? The paper operates entirely within a deterministic neoclassical growth model with a representative household whose preferences are time-separable, separable between consumption and labor, and homothetic — the &amp;ldquo;standard preferences&amp;rdquo; of Chari et al. (2020). The government&amp;rsquo;s tax instruments are proportional consumption tax rates (τ_t^c), proportional asset-income tax rates (τ_t^k), and possibly a one-time proportional levy on initial assets (l_0 ≤ 1). No empirical estimation is performed; the contribution is analytical and quantitative through calibrated simulation.&lt;/p&gt;
&lt;p&gt;The central theoretical finding is that the transitional dynamics of Chamley-Judd and the fully positive long-run capital taxes of Straub-Werning both derive from the same source: the period-zero Ramsey planner&amp;rsquo;s incentive to impose capital levies on assets that happen to exist at the start of the optimization. In Chamley et al., direct levies are precluded (l_0 = 0) and the capital-income tax rate is capped at 100%, so the planner engineers indirect levies via positive future τ_t^k (possibly forever, as Straub-Werning show) and time-varying consumption taxes. In the Chari-Nicolini-Teles (2020) formulation, the planner instead faces a constraint that household initial wealth in utility units (W_0) must meet a designated threshold (W̃_0). Under this constraint, the optimal policy features a one-time direct capital levy l_0 in period zero, zero asset-income taxes in all periods (τ_t^k = 0 for t ≥ 0), and a uniform consumption tax for all t ≥ 0. The level of l_0 and the consumption tax rate are jointly determined to satisfy the wealth constraint and the government budget.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s main contribution is extending the Chari et al. period-zero commitment to all periods, thereby achieving time-consistency and eliminating period zero&amp;rsquo;s special status. If each period-t policymaker faces a wealth constraint W_t ≥ W̃_t with W̃_t set high enough that the policymaker voluntarily chooses l_t = 0, the full sequence of policies is time-consistent and accords with Woodford&amp;rsquo;s (1999) &amp;ldquo;timeless perspective&amp;rdquo;: period zero is like any other period, capital-income tax rates are always zero, and consumption taxes are constant.&lt;/p&gt;
&lt;p&gt;The appendix provides quantitative validation using a U.S.-calibrated model: government consumption = 20% of output, capital-income tax rate = 38% (initial steady state, from Barro-Furman 2018), public debt = 70% of output, labor-income tax rate = 26%, discount factor β = 0.97 (implying a 3% real interest rate), capital share α = 0.34, and depreciation δ = 0.08. Welfare gains from switching to the Ramsey policy (with the wealth-in-utility constraint set to the pre-reform steady-state value) are 0.82% of steady-state consumption under standard preferences, 0.76% under balanced-growth preferences, and 0.62% under zero-wealth-effect preferences. Under balanced-growth preferences, the capital stock rises monotonically to a new steady state approximately 12% higher, government debt rises about 6 percentage points, the labor-income tax rate stays essentially constant at approximately 30% (roughly 4 percentage points above the old steady state), and the capital-income tax rate is approximately 1% in the first period and then drops quickly to zero. Under zero-wealth-effect preferences, the initial capital-income tax rate is slightly higher at approximately 7% before dropping sharply. Under an extreme scenario with the initial capital stock at half its steady-state level and public debt at twice its normal ratio, the capital-income tax rate starts at approximately 3% and gradually approaches zero. In all three cases, constraining the capital-income tax rate to zero and holding the labor-income tax rate constant yields welfare indistinguishable from the unconstrained Ramsey optimum. The paper concludes that zero taxation of capital income is approximately optimal across all three preference specifications, and that the apparent necessity of positive long-run capital taxes in existing literature is an artifact of the period-zero commitment asymmetry.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-period-zero-problem-and-why-is-it-central-to-the-papers-argument"&gt;Q1. What is the &amp;lsquo;period-zero problem&amp;rsquo; and why is it central to the paper&amp;rsquo;s argument?&lt;/h3&gt;
&lt;p&gt;The period-zero problem refers to the asymmetry in the standard Ramsey formulation whereby the period-zero policymaker can commit to all future tax rates but is not bound by any commitments made in the past. Because assets already in existence at period zero are inelastically supplied ex post, the planner has a strong incentive to expropriate them via a capital levy — directly (l_0) or indirectly through high early tax rates on asset income or non-constant consumption tax rates. Chamley-Judd and Straub-Werning results, while superficially different, both arise from this same incentive. The Barro-Chari paper argues that period zero is in reality just an arbitrary starting point for analysis, not a date on which commitment ability uniquely materializes, and that correctly accounting for this eliminates the period-zero problem.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-chari-nicolini-teles-2020-formulation-differ-from-chamley-et-al-and-what-does-it-imply"&gt;Q2. How does the Chari-Nicolini-Teles (2020) formulation differ from Chamley et al., and what does it imply?&lt;/h3&gt;
&lt;p&gt;Chamley et al. preclude direct capital levies (l_0 = 0) and cap τ_t^k ≤ 1, so the planner engineers indirect capital levies via positive future asset-income taxes and time-varying consumption taxes. Chari et al. (2020) instead constrain the household&amp;rsquo;s initial wealth in utility units (W_0) to be at least a designated threshold W̃_0, but leave all tax instruments unrestricted. Under this constraint, the optimal policy selects a one-time direct capital levy l_0, zero asset-income taxes forever, and uniform consumption taxes. The critical difference is that when l_0 = 0 is the outcome under the Chari et al. formulation, it is an optimizing response to a high W̃_0 rather than an arbitrary restriction, so there is no incentive for indirect levies.&lt;/p&gt;
&lt;h3 id="q3-how-is-time-consistency-achieved-and-what-is-the-timeless-perspective"&gt;Q3. How is time-consistency achieved, and what is the &amp;rsquo;timeless perspective&amp;rsquo;?&lt;/h3&gt;
&lt;p&gt;Time-consistency fails if future policymakers are unconstrained because they will repeat the period-zero capital levy logic for their own &amp;lsquo;initial&amp;rsquo; period. The paper shows that introducing a series of per-period wealth constraints — W_t ≥ W̃_t for all t ≥ 0, where W_t is period-t household wealth in utility units — achieves time-consistency if each W̃_t is set high enough that each policymaker voluntarily chooses l_t = 0. The required sequence of W̃_t corresponds exactly to the wealth path generated by the period-0 policymaker&amp;rsquo;s committed Ramsey plan. When this holds, the analysis conforms to Woodford&amp;rsquo;s (1999) &amp;rsquo;timeless perspective&amp;rsquo;: each policymaker adopts the program that would have been committed to far in the past, period zero is not special, capital-income taxes are always zero, and consumption taxes are constant.&lt;/p&gt;
&lt;h3 id="q4-what-role-do-restrictions-on-tax-instruments-play-and-why-does-the-paper-prefer-wealth-constraints-over-direct-instrument-restrictions"&gt;Q4. What role do restrictions on tax instruments play, and why does the paper prefer wealth constraints over direct instrument restrictions?&lt;/h3&gt;
&lt;p&gt;Direct instrument restrictions — such as banning capital levies (l_t = 0) or forcing τ_t^k = 0 and constant consumption taxes — are vulnerable to circumvention through other instruments. For example, time-varying labor-income tax rates (τ_t^n) introduce intertemporal wedges equivalent to indirect capital levies, so a prohibition on capital-income taxes can be undone by varying labor taxes. Constraints on household wealth in utility units (Eqs. 7 and 8) are robust to this vulnerability because any tax instrument that reduces household utility-unit wealth below the threshold violates the constraint, regardless of which specific instrument is used.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-partial-commitment-interpretation-of-the-per-period-wealth-constraints"&gt;Q5. What is the &amp;lsquo;partial commitment&amp;rsquo; interpretation of the per-period wealth constraints?&lt;/h3&gt;
&lt;p&gt;The paper offers two interpretations. The first is that the sequence of W̃_t was set at the founding of a country (e.g., 1789 for the United States). The more palatable &amp;lsquo;partial commitment&amp;rsquo; interpretation is that each period-t policymaker specifies the wealth commitment W̃_{t+1} for the next policymaker, in exchange for adhering to the commitment W̃_t set by the preceding policymaker. This bilateral exchange generates the same sequence of wealth constraints that would have been set arbitrarily far into the past.&lt;/p&gt;
&lt;h3 id="q6-what-happens-in-the-stochastic-extension-of-the-model"&gt;Q6. What happens in the stochastic extension of the model?&lt;/h3&gt;
&lt;p&gt;In a stochastic setting with fluctuations in government spending, technology, war and peace, etc. (as in Chari et al. 2020, proposition 3), choices of capital levies and tax rates become state-contingent rules, following the Lucas-Stokey (1983) framework. Non-zero direct capital levies are optimal under emergency conditions such as war, pandemic, or major financial crisis, and correspondingly below average during non-emergencies. Consumption and labor-income tax rates follow random-walk-like processes, analogous to the tax-rate smoothing predictions of Barro (1979, 1990) that apply when state-contingent capital levies are unavailable.&lt;/p&gt;
&lt;h3 id="q7-how-is-the-covid-inflation-episode-interpreted-within-this-framework"&gt;Q7. How is the COVID inflation episode interpreted within this framework?&lt;/h3&gt;
&lt;p&gt;The paper interprets the post-2020 rise in the U.S. price level through the fiscal theory of the price level (Cochrane 2023; Barro-Bianchi 2023; Bianchi-Faccini-Melosi 2023). The surge in &amp;lsquo;unfunded&amp;rsquo; government spending during and after the COVID pandemic was financed by the inflation that eroded the real value of nominally-denominated government bonds. This constitutes a state-contingent capital levy on bondholders. A cautionary note is added: the availability of such a mechanism may encourage excessive spending, analogous to Ricardo&amp;rsquo;s (1820) argument for balanced-budget war finance.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-role-of-heterogeneity-among-households-in-potentially-generating-commitment"&gt;Q8. What is the role of heterogeneity among households in potentially generating commitment?&lt;/h3&gt;
&lt;p&gt;The paper discusses two sources. First, drawing on Broner-Martin-Ventura (2010), if the government cares about domestic holders of its bonds but not foreign holders, and if bonds can be traded on secondary markets so the two groups cannot be separated, then default becomes unattractive ex post because it harms domestic residents. This gives the government an incentive to promote secondary markets as a commitment device against sovereign default — potentially extensible to capital taxation commitments. Second, the distinction between old and new capital (e.g., via investment tax credits) partially limits the attractiveness of high capital-income taxes by tying the tax rate on old capital to the rate on new capital, which creates investment disincentives. However, as Straub-Werning demonstrate, this commitment may be too weak to drive the optimal capital-income tax to zero.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-calibration-targets-and-preference-specifications-used-in-the-quantitative-experiments"&gt;Q9. What are the calibration targets and preference specifications used in the quantitative experiments?&lt;/h3&gt;
&lt;p&gt;The model is calibrated to represent the U.S. economy with: government consumption = 20% of output, capital-income tax rate = 38% (from Barro-Furman 2018), public debt = 70% of output, labor fraction of time endowment = 1/3, discount factor β = 0.97 (3% real interest rate), capital share α = 0.34, depreciation δ = 0.08. Three preference specifications are explored: (1) standard preferences (time-separable, separable, homothetic in c and n); (2) balanced-growth preferences with consumption-leisure Cobb-Douglas aggregator and IES = 0.5; (3) zero-wealth-effect preferences. The wealth constraint W̃_0 is set to match the pre-reform steady-state wealth in utility terms.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-detailed-quantitative-results-across-preference-specifications"&gt;Q10. What are the detailed quantitative results across preference specifications?&lt;/h3&gt;
&lt;p&gt;Under standard preferences: capital-income tax rate is always exactly zero, labor-income tax rate is constant, welfare gain = 0.82% of steady-state consumption. Under balanced-growth preferences (IES = 0.5): initial capital-income tax ≈ 1%, quickly drops to zero; capital stock rises ≈ 12% to new SS; government debt rises ≈ 6 pp; labor-income tax ≈ 30% (constant, ≈ 4 pp above old SS of 26%); welfare gain = 0.76%; steady-state public debt under zero-capital-tax policy = 33% of output; initial capital levy l_0 = 0.126; new SS labor tax = 0.297. Under zero-wealth-effect preferences: initial capital-income tax ≈ 7%, drops sharply; welfare gain = 0.62%; l_0 = 0.160; new SS labor tax = 0.301; maximum capital tax rate = 0.070. Under extreme initial conditions (balanced-growth, capital stock at half SS level, debt at twice normal ratio): capital-income tax ≈ 3% initially, approaches zero; l_0 = 0.033; new SS labor tax = 0.400. Across all cases, constraining capital-income tax to zero with constant labor tax yields welfare nearly identical to the unconstrained Ramsey optimum.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-scope-of-the-zero-capital-tax-result-and-what-preference-conditions-support-it"&gt;Q11. What is the scope of the zero-capital-tax result and what preference conditions support it?&lt;/h3&gt;
&lt;p&gt;The zero-capital-tax result holds exactly under standard preferences (time-separable, separable between consumption and labor, and homothetic in consumption and labor), which satisfy the Diamond-Mirrlees-Sandmo-Sadka conditions for uniform taxation of goods. Under balanced-growth preferences, it holds with σ = 1 but not necessarily when σ ≠ 1. Under zero-wealth-effect preferences it does not hold if V is strictly concave. However, the quantitative experiments show that deviations from zero are small and short-lived under all three specifications, so zero capital taxation is approximately optimal across the board.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-relationship-between-the-papers-results-and-tax-rate-smoothing-models"&gt;Q12. What is the relationship between the paper&amp;rsquo;s results and tax-rate smoothing models?&lt;/h3&gt;
&lt;p&gt;Barro (1979, 1990) showed that optimal income-tax rates follow a random walk when capital levies are unavailable. The present paper shows that, once state-contingent capital levies are available (the Lucas-Stokey stochastic extension), consumption and labor-income tax rates also exhibit random-walk-like behavior, as realizations of spending and technology shocks move the optimal tax rates. This provides a unified framework connecting capital levy theory and tax-rate smoothing.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-survivalinstitutional-arguments-for-why-commitment-constraints-might-exist-in-practice"&gt;Q13. What are the survival/institutional arguments for why commitment constraints might exist in practice?&lt;/h3&gt;
&lt;p&gt;The paper suggests a selection argument: societies that fail to maintain commitments of the form W_t ≥ W̃_t severely under-accumulate capital because anticipating capital levies causes households and firms not to invest, potentially causing the economy to effectively disappear. This selection pressure may explain why functioning market economies tend to develop institutions (constitutions, property rights, secondary markets) that approximate the required commitments. Major regime changes, such as the Bolshevik revolution (100% default on Czarist bonds), can destroy these commitments, but many regime changes (e.g., France after World War II) do not fully repudiate prior obligations.&lt;/p&gt;
&lt;h3 id="q14-how-does-this-paper-relate-to-and-differ-from-the-three-main-antecedents-chamley-judd-straub-werning-and-chari-et-al-2020"&gt;Q14. How does this paper relate to and differ from the three main antecedents (Chamley-Judd, Straub-Werning, and Chari et al. 2020)?&lt;/h3&gt;
&lt;p&gt;Chamley (1986) and Judd (1985, 1999) showed zero long-run capital-income tax is optimal under the Ramsey formulation with l_0 = 0 and τ_t^k ≤ 1. Straub-Werning (2020) showed that positive capital-income taxes can be optimal even in the steady state under the same constraints when the IES is below one. Chari et al. (2020) replaced instrument restrictions with a utility-wealth constraint for period zero, obtaining a direct capital levy in period zero plus zero capital-income taxes thereafter. Barro-Chari extend Chari et al.&amp;rsquo;s period-zero constraint to all periods, achieving time-consistency and removing period zero&amp;rsquo;s special status. The novel contribution is the multi-period, time-consistent version of the Chari et al. framework and the quantitative demonstration that zero capital taxation is approximately optimal across preference specifications.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Period-zero problem&lt;/strong&gt;: The asymmetry in the standard Ramsey formulation in which the period-zero policymaker can commit to all future tax rates but faces no commitments from the past, creating a strong incentive to expropriate existing assets via capital levies (direct or indirect); the paper&amp;rsquo;s central target of critique.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital levy&lt;/strong&gt;: A proportional confiscation of asset holdings (l_t), distinct from ongoing taxes on the flow of asset income; a direct capital levy takes a fraction of the stock outright, while indirect capital levies are engineered through high asset-income tax rates or time-varying consumption taxes that reduce the real value of existing wealth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wealth constraint in utility units (W_t ≥ W̃_t)&lt;/strong&gt;: A commitment device, following Chari-Nicolini-Teles (2020) and Armenter (2008), that requires each period&amp;rsquo;s policymaker to leave households with at least a threshold level of wealth measured in units of utility rather than goods; instrumental in eliminating the period-zero problem without directly restricting tax instruments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Timeless perspective&lt;/strong&gt;: Woodford&amp;rsquo;s (1999) principle that the policymaker should adopt the behavior that would have been committed to far in the past contingent on current events, rather than optimizing from the current period taking past expectations as given; the paper shows its Ramsey results conform to this principle once per-period wealth constraints are imposed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time-consistency (in optimal taxation)&lt;/strong&gt;: The property that a tax plan chosen at date 0 will be voluntarily continued by each subsequent policymaker; fails in the Chari et al. (2020) baseline formulation when future policymakers are unconstrained because each will want to re-impose a &amp;lsquo;period-zero&amp;rsquo; capital levy, achieved here only when per-period wealth constraints W_t ≥ W̃_t are sufficient to deter direct levies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Indirect capital levy&lt;/strong&gt;: The engineering of a de facto reduction in the real value of existing wealth through policy instruments other than a direct asset levy — specifically positive tax rates on future asset income (τ_t^k &amp;gt; 0) or non-constant consumption tax rates that alter the present value of after-tax consumption; the mechanism underlying both Chamley-Judd transitional dynamics and Straub-Werning permanent positive capital taxes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Standard preferences&lt;/strong&gt;: Preferences that are time-separable, separable between consumption and labor, and homothetic in consumption and labor (Eq. 1 in the paper: u(c,n) = [c^{1-σ}/(1-σ)] − η·n^{1+Ψ}); the class under which uniform taxation of consumption at all dates and zero tax rates on asset income are exactly optimal, satisfying Diamond-Mirrlees-Sandmo-Sadka conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;State-contingent capital levy&lt;/strong&gt;: In the stochastic extension (following Lucas-Stokey 1983), a capital levy whose magnitude depends on the realized state of the world (e.g., war, pandemic, financial crisis); optimal under emergencies when emergency government spending must be financed, and below average during normal times — the paper interprets post-2020 U.S. inflation as an implicit state-contingent levy on nominal government bonds via the fiscal theory of the price level.&lt;/p&gt;</description></item><item><title>Taxing Top Wealth: Migration Responses and their Aggregate Economic Implications</title><link>https://macropaperwarehouse.com/papers/taxing-top-wealth-migration-responses-and-their-aggregate-economic-implications/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/taxing-top-wealth-migration-responses-and-their-aggregate-economic-implications/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Proposals to tax top wealth (e.g., Saez and Zucman, 2019) face a recurring objection in public debate: that the wealthy will emigrate en masse and, because many are entrepreneurs, their departure will inflict large negative spillovers (&amp;ldquo;trickle-down&amp;rdquo;) on the broader economy, making wealth taxes self-defeating. Credible evidence on international migration responses to wealth taxes has been scarce due to data limitations and a lack of clean identifying variation. This paper provides such evidence and quantifies the aggregate economic implications.&lt;/p&gt;
&lt;p&gt;Data and setting: The authors use exhaustive administrative data from Sweden (wealth tax register Förmögenhetsregistret 1993-2007, LISA, matched employer-employee RAMS, K10 closely-held-business filings, and the Serrano ownership-network data that maps indirect ownership) and Denmark (used for out-of-sample validation). A key strength is observing all wealth components without top-coding and linking individuals to firms they control directly and indirectly. They exploit three large reforms: the unexpected 2007 repeal of the Swedish wealth tax (statutory top marginal rate fell from 1.5% to 0%; effective average rate on the top 2% was ~0.5%), and Danish reforms of 1989 (rate cut from 2.2% to 1%) and 1996/1997 (abolition). Business assets were exempt in Sweden but fully taxed in Denmark.&lt;/p&gt;
&lt;p&gt;Empirical strategy: A two-step procedure. Step 1 estimates migration elasticities using difference-in-differences around the reforms (treated = top 2% of net wealth; baseline control = top 20% to top 10%), with treatment assigned on predicted wealth to avoid endogeneity post-2007. Step 2 estimates the effect of migration on individual-, firm-, and market-level outcomes via event studies (never-movers with placebo dates as controls), independent of the tax reforms. The two are combined, weighted by the wealthy&amp;rsquo;s share of aggregate activity (decomposition in equation 1).&lt;/p&gt;
&lt;p&gt;Main quantitative findings: A 1pp increase in the top wealth tax rate raises the out-migration rate by 0.17pp and reduces in-migration by 0.05pp; the 2007 repeal cut wealthy out-migration propensity by ~30% (about one-third of top-2% expatriations were tax-induced). Danish elasticities are statistically indistinguishable. Net flow semi-elasticity is -0.22pp per 1pp. Flow effects cumulate to a modest stock elasticity: the elasticity of the wealthy population w.r.t. the net-of-tax rate is 1.77 (s.e. 0.47) — a 1% rise in the net-of-tax rate raises the stock by under 2%. The implied income-net-of-tax migration elasticity is ~0.05, comparable to top-income cross-border elasticities. Firms controlled by the top 2% account for ~9% of Swedish employment, 15% of value added, 12% of investment, 19% of tax payments (and ~10% employment / 15% value added per the intro). When a top-2% owner out-migrates, directly-controlled firms see employment fall ~33%, gross investment ~22%, value added ~34%, and tax payments ~51%, driven almost entirely by the extensive margin of firm disappearance (effects near zero conditional on survival). But 45% of &amp;ldquo;closed&amp;rdquo; firms are absorbed via mergers/acquisitions; displaced workers lose only 4.3% in earnings and face a 0.6pp higher unemployment probability; market-level spillovers are small and insignificant even for granular firms.&lt;/p&gt;
&lt;p&gt;Aggregate and policy implications: Combining steps, a 1pp rise in the top wealth tax rate reduces aggregate employment by 0.022%, investment by 0.065%, and value added by 0.103% in the long run — modest despite the wealthy&amp;rsquo;s large economic footprint, because migration flows are small. Fiscally, each $1 raised loses only $0.22 to migration responses vs. $0.54 to intensive-margin responses (savings/avoidance/evasion, using Jakobsen et al. 2020), so $0.76 total. Migration responses are far from the Laffer bound but, because the MCPF is highly nonlinear, they nearly double it from ~2.2 to ~4.2. Migration threats, while salient in debate, matter less for welfare and policy than intensive-margin responses.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-migration-elasticity-and-what-are-the-main-threats"&gt;Q1. What is the identification strategy for the migration elasticity and what are the main threats?&lt;/h3&gt;
&lt;p&gt;A difference-in-differences design around the 2007 Swedish wealth tax repeal, comparing out-migration of the treated top-2% group to a control group in the top 20% to top 10%. The non-contiguous control avoids contamination bias (households near the threshold anticipating future liability; less than 1% of controls reach the top 2% by 2006). The main threat is the parallel-trends assumption given a control group lower in the distribution; the authors show no differential pre-trends in out-migration and that effective capital-income and labor-income tax rates evolved similarly across groups (only wealth-inclusive tax rates diverged). The 2007 inheritance tax abolition is ruled out as a confounder because inheritance tax had little bite and strict residency rules made it hard to avoid by migrating (10-year non-residence required at death). Treatment is assigned on predicted wealth (from pre-reform variables) to avoid endogenous post-2007 wealth measurement. 2SLS specification (4) instruments the log net-of-tax rate with the treatment-by-post interaction.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-aggregate-effect-identified-separately-from-the-migration-channel-and-why-not-use-the-reform-directly"&gt;Q2. How is the aggregate effect identified separately from the migration channel, and why not use the reform directly?&lt;/h3&gt;
&lt;p&gt;National wealth tax reforms cannot identify general-equilibrium/aggregate effects because treatment and control groups share the same aggregate economy, the exclusion restriction fails (wealth taxes also affect savings, capital accumulation, avoidance/evasion), and they are underpowered (small stock changes are hard to detect). The two-step procedure circumvents this: event studies of migration events (specification 7, with randomly-assigned placebo dates for never-movers, no matching) give the effect of migration on outcomes independent of the tax reform, and these are combined with the reform-based migration elasticity, weighted by the wealthy&amp;rsquo;s share of each aggregate outcome (equation 1).&lt;/p&gt;
&lt;h3 id="q3-what-is-the-role-of-the-late--marginal-mover-correction"&gt;Q3. What is the role of the LATE / marginal-mover correction?&lt;/h3&gt;
&lt;p&gt;The two-step procedure requires the population whose migration impact is measured (event studies) to match the population whose migration responds to the tax (compliers). Using methods from the insurance-selection literature (Hendren et al., 2021) and the fact that 30% of pre-reform wealthy migrants were tax compliers, they recover the characteristics and treatment effects of marginal movers. Tax-induced movers (compliers) are slightly younger, slightly more likely entrepreneurs, slightly wealthier, around the 65th-70th skill percentile, but their firms are not selected. Event-study estimates pre vs post reform are similar (not statistically different), so treatment-effect heterogeneity is limited; column (5) double-difference LATE estimates for compliers are the preferred inputs.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-firm-level-evidence-and-how-is-reallocation-distinguished-from-genuine-destruction"&gt;Q4. What is the firm-level evidence and how is reallocation distinguished from genuine destruction?&lt;/h3&gt;
&lt;p&gt;Owner out-migration causes a ~30pp drop in firm survival (firm-identifier disappearance) and large declines in employment (~33%), value added (~34%), investment (~22%), turnover, and tax payments (~51%), almost entirely extensive-margin. The authors distinguish destruction from reallocation using Bolagsverket merger/closure-reason data: 45% of closures are linked to mergers (the firm is absorbed), 55% are liquidations/bankruptcies. Accounting for buy-outs cuts the firm-existence and employment effects by ~40%. Worker-level event studies show displaced employees lose only 4.3% in earnings and 0.6pp higher unemployment, indicating workers reallocate. Including indirectly-held firms, five-year effects are employment -19%, value added -33%, turnover -28%, investment -19%, tax payments -45%.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented"&gt;Q5. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Migration semi-elasticities do not vary much by age or education; entrepreneurs&amp;rsquo; out-migration semi-elasticity is larger but less precisely estimated (their effective tax rate dropped less because business assets were exempt; their out-migration fell ~0.14pp, roughly 50%, within a year). Firm-level migration effects show limited heterogeneity by owner age or children; effects are smaller for larger firms and especially for the top-10 largest moves (multi-billion-SEK businesses), where effects are considerably below average. In-migration effects mirror out-migration with opposite sign but are smaller for value added, turnover, investment, and tax payments.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Estimates are robust to alternative control groups closer to the treatment group; to assumptions on the regeneration/replacement rate of the wealthy population and to dynastic effects (detectable but small); and to tax evasion — using Alstadsæter et al. (2019) and Boas et al. (2024) bounds, the stock elasticity ranges 1.85 (lower) to 1.92 (upper) vs. 1.77 baseline. Firm outcomes are robust to winsorization choices (Appendix Table IV.3); with no winsorization, value added/investment/tax effects turn positive-insignificant due to one outlier firm. Market-level spillovers are insignificant across alternative market definitions. Alternative aggregate calibrations (including accounting for buy-outs) imply smaller effects, so the baseline is a conservative upper bound.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-paper-relate-to-and-differ-from-prior-work"&gt;Q7. How does the paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;It builds on the wealth-tax behavioral-response literature (Seim 2017; Jakobsen et al. 2020; Brülhart et al. 2022) which is largely silent on international migration, and on the tax-migration literature (Kleven et al. 2013/2014/2020; Akcigit et al. 2016) which focuses on income taxes and within-country mobility. It is the first systematic evidence on international migration responses to wealth taxes and their trickle-down. Versus the CEO/owner death-and-retirement literature (Smith et al. 2019: -26pp firm survival, -82% profits per worker, -45% even conditional on survival; Jäger and Heining 2022), migration effects are much smaller and nearly zero conditional on survival, because owners often retain control or restructure rather than shut down. Findings echo Bach et al. (2023) for France.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Migration-driven fiscal externality is $0.22 per $1 raised, vs. $0.54 for intensive-margin responses, $0.76 combined — below the Laffer bound. Because the MCPF is nonlinear, migration roughly doubles it from ~2.2 to ~4.2; wealth taxation would be welfare-improving if revenue funds projects with MVPF above 4.2 (e.g., programs for low-income children, often above 5 per Hendren and Sprung-Keyser 2020). Scope conditions: estimates come from reforms that only cut rates, so asymmetric responses to increases cannot be ruled out; the elasticity depends on destination-country taxes (Swedish movers went to low-tax UK non-dom, Switzerland, Austria), so responses could be more muted if all neighbors taxed wealth heavily; results are for small open economies with low wealth inequality and weaker agglomeration than the US, suggesting the estimates are upper bounds; computations reflect 1990s-2000s Scandinavia where offshoring/evasion mattered, and depend on tax base, enforcement, and exit-tax design.&lt;/p&gt;
&lt;h3 id="q9-how-is-the-stock-elasticity-derived-from-flow-elasticities"&gt;Q9. How is the stock elasticity derived from flow elasticities?&lt;/h3&gt;
&lt;p&gt;Using a simple OLG framework, the population stock elasticity ≈ net-flow semi-elasticity times (T+1)/2, where T is the average &amp;rsquo;lifespan&amp;rsquo; of wealthy individuals (the inverse of the regeneration/birth rate into the wealthy population). Longer lifespan means slower regeneration, so lost migrants are harder to replace and the stock effect is larger. This yields a stock elasticity of 1.77 (s.e. 0.47); the effect stays modest because top-of-distribution migration flow rates are very small.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-magnitudes-of-migration-flows-and-tax-payment-effects-and-any-caveats-on-persistence"&gt;Q10. What are the magnitudes of migration flows and tax-payment effects, and any caveats on persistence?&lt;/h3&gt;
&lt;p&gt;Top-decile out-migration is ~0.2% per year in Sweden (vs. ~0.65% in the bottom half) and ~0.1% in Denmark, rising in the extreme tail; taxable wealth of wealth-tax-liable out-migrants is only 0.09% of total taxable wealth; net migration is small and slightly positive. One year after out-migration, total tax payments fall ~66% (wealth tax -59%, income tax -68%; income taxes are ~90% of the wealthy&amp;rsquo;s payments, implying large fiscal externalities on income tax). Effects attenuate over time: ~40% reduction at five years because ~40% of out-migrants return within five years (migration is persistent but return migration is common). Taxable wealth in Sweden falls 94% one year out; real estate is typically sold, and financial wealth falls at extensive (-21%) and intensive (-15%) margins, confirming real rather than purely fiscal-residence responses.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>The Aggregate Costs of Uninsurable Business Risk</title><link>https://macropaperwarehouse.com/papers/the-aggregate-costs-of-uninsurable-business-risk/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-aggregate-costs-of-uninsurable-business-risk/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question and motivation.&lt;/strong&gt; A large literature argues that credit constraints are the dominant financial friction holding private businesses below their optimal scale, so that easing credit access would yield large aggregate efficiency gains. This paper challenges that view. Private businesses are also poorly diversified — their owners bear undiversifiable business-income risk — and the authors argue the macroeconomic costs of this lack of diversification are far larger than those of credit constraints. The crux is that entrepreneurs can limit risk exposure by operating at a smaller scale, so productive-but-poor entrepreneurs choose an inefficiently low scale and are unwilling to borrow to expand. Firm size is thus limited by risk, not by credit availability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and setup.&lt;/strong&gt; The empirical analysis uses the historical Orbis dataset (Moody&amp;rsquo;s Bureau van Dijk), 1995–2019, focusing on Spain (best coverage; results extend to Italy, France, Norway, Portugal, Slovakia in the appendix). Output is value added; the sample is partnerships and private limited companies, excluding FIRE, public administration, defense, education. The final sample is 622,883 firms (6,298,358 firm-year observations), observed on average 10 years; the mean (median) firm has 12 (5) workers and 486 (151) thousand EUR value added. The Spanish Survey of Household Finances (EFF, 2008–2020) provides entrepreneur wealth/prevalence and consumption data. The model is a small-open-economy model of entrepreneurial dynamics (à la Quadrini 2000; Cagetti–De Nardi 2006) with two frictions: each firm is owned by a single (undiversified) entrepreneur, and a collateral constraint k&amp;rsquo; ≤ a&amp;rsquo;/(1−ξ). Key modeling choices: capital AND labor are chosen before productivity is observed (time-to-build), and productivity has persistent and transitory shocks drawn from fat-tailed mixtures of normals. Parameters are estimated by simulated method of moments (9 parameters, 16 moments; objective 0.013, ~1.3% average deviation).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main quantitative findings.&lt;/strong&gt; Profit shares fluctuate sharply: 5% of firms have losses exceeding 20% of output, against an average profit share of 0.13; the 5th percentile of profit-share deviations is −0.33 and the 95th is +0.47. Output growth is fat-tailed (s.d. 0.48, IQR/s.d. ratio 0.65 vs 1.35 Gaussian; excess kurtosis 10.7). Inputs do not track output: regressing wage-bill growth on output growth gives 0.40 (capital 0.16); restricting to |Δlog y|&amp;lt;0.5 gives 0.58 and 0.31. A change in profit share on output growth has slope 1.56 (0.46 in the restricted sample). The headline result: eliminating both frictions would raise output by 15.8%; eliminating the risk wedge alone raises output by 15.4%, while eliminating the credit wedge alone raises output by only 0.4%. Misallocation losses are 10.8% (11.0% due to risk, 0.2% due to credit). Aggregate wedges are equivalent to a 12.8% tax on labor and 14.9% on capital. Wage losses are 27.8% (26.4% risk, 0.4% credit).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanisms and implications.&lt;/strong&gt; Two wedges distort choices: a risk wedge (from the covariance of consumption and productivity) that distorts both labor and capital, and a credit wedge (from the binding collateral constraint) that distorts only capital. The credit wedge falls quickly with wealth (vanishing once unconstrained), but the risk wedge declines only gradually and persists even for wealthy entrepreneurs. Aggregate losses are governed by the distribution of wedges weighted by efficient firm size (Hopenhayn 2014): risk wedges are large precisely for high-ability entrepreneurs who would be large under efficiency, whereas credit-constrained firms are mostly unproductive with small efficient size. Policy implication: improving credit access has limited impact unless it also improves risk sharing. The findings also imply firm profits largely reflect compensation for risk (75% of the aggregate profit share), and dispersion in returns to business wealth largely reflects risk compensation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-model-and-how-are-parameters-pinned-down"&gt;Q1. What is the identification strategy for the model, and how are parameters pinned down?&lt;/h3&gt;
&lt;p&gt;Parameters ϑ=(β,α,η,ρ,σu,σε,s,p,ϕ) are estimated by simulated method of moments, minimizing a weighted distance between 16 empirical and model moments scaled by 1+empirical moment (objective = 0.013, ~1.3% average deviation). Intuitively: β is pinned by the entrepreneur wealth-to-income ratio (12.5 in data and model); α and η by the capital-output ratio (1.22 vs 1.21), labor share (0.72 vs 0.71) and profit share (0.13 vs 0.14); ρ, σu, σε by output autocorrelations at horizons 1–3, the cross-sectional s.d. of output, and the s.d. of output growth at horizons 1–3; the tail parameters s and p by the IQR of output growth relative to its s.d.; and ϕ by the entrepreneurship rate. Three assigned parameters: δ=0.10, r=0.02, θ=2, with ξ=0.408 set to match the aggregate debt-to-capital ratio of 0.408. Standard errors (bootstrapped) are small because the firm sample is very large.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-main-mechanism-and-how-are-the-risk-wedge-and-credit-wedge-distinguished"&gt;Q2. What is the main mechanism, and how are the risk wedge and credit wedge distinguished?&lt;/h3&gt;
&lt;p&gt;Because labor and capital are chosen before productivity is realized and risk is undiversified, the entrepreneur weights future states by their own stochastic discount factor. The risk wedge τ (&amp;gt;1) arises from the negative covariance between marginal utility of consumption and productivity and distorts both labor and capital equally. The credit wedge ω (&amp;gt;1 when the collateral constraint binds) distorts only capital. As wealth rises, the credit wedge falls rapidly and vanishes once the firm is unconstrained, but the risk wedge declines only gradually and never disappears. The two are isolated quantitatively by setting ω=1 (to get the role of risk) or τ=1 (to get the role of credit) in the productivity-loss mapping (eq. 13).&lt;/p&gt;
&lt;h3 id="q3-why-does-risk-dominate-credit-in-the-aggregate-even-though-most-firms-are-credit-constrained"&gt;Q3. Why does risk dominate credit in the aggregate even though most firms are credit-constrained?&lt;/h3&gt;
&lt;p&gt;Aggregate outcomes depend on the distribution of wedges weighted by efficient firm size n_it (Hopenhayn 2014). Weighted by efficient size, the risk wedge ranges from 1.27 (10th pct) to 1.61 (90th pct), while the credit wedge is essentially 1 except at the very top (1.02 at the 90th pct). Unweighted, the risk wedge is only 1.12 at the 90th pct and the credit wedge is positive for more than half of firms — but those constrained firms are unproductive with small efficient size. Risk wedges are large precisely for high-ability entrepreneurs who would be large under the efficient allocation, so they drive the aggregate.&lt;/p&gt;
&lt;h3 id="q4-why-is-the-result-robust-to-the-form-of-the-collateral-constraint"&gt;Q4. Why is the result robust to the form of the collateral constraint?&lt;/h3&gt;
&lt;p&gt;The authors consider two extremes: no borrowing at all (ξ=0) and unlimited borrowing (ξ=1, no credit limit). With no borrowing, misallocation losses rise only from 10.8% to 11.7%, still mostly risk-driven (8.3% risk vs 1.4% credit). With no credit limit, risk wedges remain nearly as large as baseline and removing credit frictions has negligible effects. Intuitively, risk leads entrepreneurs to operate small and accumulate precautionary wealth, so they self-finance most desired capital and credit wedges stay small even without credit.&lt;/p&gt;
&lt;h3 id="q5-which-three-ingredients-are-essential-to-the-risk-dominates-result-and-what-happens-without-each"&gt;Q5. Which three ingredients are essential to the risk-dominates result, and what happens without each?&lt;/h3&gt;
&lt;p&gt;(1) Fat-tailed productivity shocks, (2) transitory productivity shocks, and (3) labor chosen before productivity is realized. Removing each in isolation (with re-estimation) reverses the conclusion so that credit becomes the primary driver: without fat tails, misallocation losses fall to 2.1% (credit 1.5%, risk 0.3%); without transitory shocks, losses are 12.1% (credit 10.9%, risk 0.4%); with flexible labor, losses fall to 3.3% (credit 2.4%, risk 0.1%). The flexible-labor case matters because risk then distorts only capital, whose share is smaller than labor&amp;rsquo;s, reducing income volatility and pushing firms to expand and hit the credit constraint. In all three counterfactuals, the 1st percentile of profit-share deviations ranges −0.21 to −0.43, far smaller in magnitude than the data (−1.66) or baseline model (−1.92).&lt;/p&gt;
&lt;h3 id="q6-is-the-result-driven-by-high-risk-aversion"&gt;Q6. Is the result driven by high risk aversion?&lt;/h3&gt;
&lt;p&gt;No. The baseline uses relative risk aversion θ=2. Re-estimating with θ=0.5 (low end of usual values) still yields sizable, risk-dominated losses: productivity losses 6.4%, output losses 9.2%, wage losses 16.7% — roughly three-fifths of the baseline — and again primarily driven by risk rather than credit.&lt;/p&gt;
&lt;h3 id="q7-what-untargeted-moments-does-the-model-match-model-validation"&gt;Q7. What untargeted moments does the model match (model validation)?&lt;/h3&gt;
&lt;p&gt;The model reproduces the distribution of profit-share deviations (10th pct −0.17 data vs −0.16 model; 1st pct −1.66 data vs −1.92 model), the full distribution of output growth rates, the low wage-bill/output comovement (0.58 data vs 0.55 model in the restricted sample), the profit-share/output comovement (0.46 vs 0.42; falling to 0.10 vs 0.06 when holding the labor share constant), and the persistence/volatility of capital and labor (e.g., wage-bill growth s.d. 0.36 vs 0.32). Critically, it matches the low comovement of entrepreneur consumption with profits: regressing Δc on Δπ gives a slope of 0.02 in both data and model (data based on 799 EFF observations, three-year changes).&lt;/p&gt;
&lt;h3 id="q8-what-heterogeneity-and-external-validity-does-the-paper-document"&gt;Q8. What heterogeneity and external validity does the paper document?&lt;/h3&gt;
&lt;p&gt;The motivating facts hold for Italy, France, Norway, Portugal and Slovakia, and for Spanish public firms; for young (age≤5) and old firms; for small and large firms (top decile of value added vs rest); and across the five largest sectors (manufacturing, construction, wholesale/retail, accommodation/food, professional activities). Output-growth kurtosis ranges roughly 11–18 across countries. On diversification: 12% of households are entrepreneurs; 93% of entrepreneurs own exactly one business; multi-business owners hold 71% of their business wealth in their main business; the average ownership share is 83%, and 71% own 100% of their main business.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-extensive-margin-and-unconstrained-firm-results"&gt;Q9. What are the extensive-margin and unconstrained-firm results?&lt;/h3&gt;
&lt;p&gt;Extensive margin: when the planner can also choose who becomes an entrepreneur, it cuts the entrepreneurship rate from 13.2% to 1.2%, but because marginal entrepreneurs are low-ability the gains are small — productivity, output and wage losses relative to the unconstrained planner are 10.8%, 16% and 27.8%, very close to the intensive-margin numbers. Unconstrained firms: adding a frictionless sector calibrated to match the 58.7% output share of public firms in Orbis leaves misallocation losses at 10.5% (vs 10.8% baseline), still mostly risk-driven (risk 10.1%, credit 0.1%); wage losses fall to about three-fifths of baseline because the unconstrained sector reduces the aggregate labor wedge.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-implications-for-profits-and-returns-to-wealth"&gt;Q10. What are the implications for profits and returns to wealth?&lt;/h3&gt;
&lt;p&gt;Decomposing the profit share into span-of-control, risk and credit components: risk accounts for 75% of the aggregate profit share (0.11/0.146), with the rest from span of control; credit contributes little. Risk also drives most of the profit-share dispersion (s.d. 5.5%, essentially all from risk; credit contributes only 1%). For excess returns to wealth, the mean of 2.2% is almost entirely accounted for by risk, and risk drives most of the dispersion (s.d. 5.5%). This implies dispersion in returns to private business wealth — a driver of wealth inequality — largely reflects compensation for risk rather than credit constraints.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-working-capital-robustness-check"&gt;Q11. What is the working-capital robustness check?&lt;/h3&gt;
&lt;p&gt;Adding a working-capital constraint where a fraction ϑ=0.25 of the wage bill is paid in advance (à la Mendoza 2010), evaluated at baseline parameters, gives misallocation losses of 11.1% (vs 10.8% baseline), with risk still accounting for the bulk (9.4%) and credit less important (1.3%); risk accounts for 13.4% of the 16.3% total output losses. So even when credit frictions can also distort labor, risk remains dominant.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The central implication is that policies expanding firms&amp;rsquo; access to credit will have limited aggregate impact unless they also improve risk sharing. This holds within the scope of the model — undiversified private businesses with single owners, where risk exposure is endogenously chosen via scale and can be partly self-insured through wealth, labor income, and occupational switching. The authors note their framework assumes (rather than micro-founds) the lack of diversification, and suggest future work should model the moral-hazard or informational frictions preventing diversification, and broaden redistributive tax analysis to incorporate uninsurable-risk distortions (as in Di Tella et al. 2024).&lt;/p&gt;
&lt;h3 id="q13-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q13. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It contributes to the misallocation literature (Hsieh-Klenow 2009; Buera et al. 2011; Moll 2014; Midrigan-Xu 2014; Gopinath et al. 2017). Prior work on risk and investment (Tan 2018; Robinson 2021; David et al. 2022a) studies how risk distorts investment; this paper instead emphasizes how risk distorts LABOR choices, relating it to Arellano et al. (2019) and David et al. (2022b). It differs from the credit-constraint-centric tradition by showing credit matters little once undiversified risk and the three key ingredients are present. Di Tella et al. (2024), partly motivated by these findings, study optimal policy under uninsurable risk and show it is the opposite of optimal policy when misallocation stems from markups.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Risk wedge (τ)&lt;/strong&gt;: In the paper&amp;rsquo;s sense, the gap between the expected marginal product of an input and its price arising from undiversifiable business risk. It equals [1 + COV(c^{-θ}, zε)/(E c^{-θ} · E zε)]^{-1}, generally &amp;gt;1 because of the negative covariance between the entrepreneur&amp;rsquo;s marginal utility of consumption and productivity. It distorts both labor and capital, declines only gradually with wealth, and persists even for wealthy entrepreneurs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Credit wedge (ω)&lt;/strong&gt;: The distortion from a binding collateral constraint, ω=1+(1−ξ)μ/R, where μ is the multiplier on the constraint k&amp;rsquo;≤a&amp;rsquo;/(1−ξ). It exceeds one only when the constraint binds, distorts only capital, falls rapidly with wealth, and vanishes once the entrepreneur is unconstrained.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Profit share&lt;/strong&gt;: In this paper, the ratio of profits to output (value added), π_it/y_it, where profit is output net of the wage bill and the user cost of capital. Its average is 0.13; the paper studies its large transitory firm-level fluctuations as the empirical signature of uninsurable risk.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time-to-build (inputs chosen before productivity)&lt;/strong&gt;: The assumption that both capital and labor are chosen before the firm observes its productivity shock. This parsimoniously generates the imperfect high-frequency comovement between inputs and output and makes wealth affect employment as well as investment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Efficient-size-weighted wedge distribution&lt;/strong&gt;: The paper&amp;rsquo;s organizing device (following Hopenhayn 2014): aggregate productivity losses depend on the distribution of risk and credit wedges weighted by each firm&amp;rsquo;s efficient size n_it. Because high-ability firms have large efficient size and large risk wedges, risk dominates the aggregate even though most firms are credit-constrained.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Self-financing&lt;/strong&gt;: The mechanism by which entrepreneurs, operating at small scale and saving for precautionary reasons because of risk, accumulate enough wealth to finance most of their desired capital — so credit wedges stay small even in an economy with no credit, rendering the borrowing limit nearly irrelevant for aggregates.&lt;/p&gt;</description></item><item><title>The Efficiency-Equity Tradeoff of the Corporate Income Tax: Evidence from the Tax Cuts and Jobs Act</title><link>https://macropaperwarehouse.com/papers/the-efficiency-equity-tradeoff-of-the-corporate-income-tax-evidence-from-the-tax-cuts-and-jobs-act/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-efficiency-equity-tradeoff-of-the-corporate-income-tax-evidence-from-the-tax-cuts-and-jobs-act/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper estimates the firm- and worker-level effects of the corporate income tax cuts in the 2017 Tax Cuts and Jobs Act (TCJA) — the largest corporate tax cut in U.S. history — to inform the long-running efficiency-versus-equity debate over corporate taxation. The question matters because federal corporate tax reforms are rare, prior credible evidence comes mostly from subnational or small-economy variation (where factors are more mobile and the tax base smaller), and theory predicts alternate instruments behave differently, so existing estimates may not extrapolate to a major reform in a large advanced economy.&lt;/p&gt;
&lt;p&gt;Identification exploits that TCJA cut the top C-corporation rate from 35% to 21% (a 40% reduction) while cutting the implied top rate for S corporations far less — from 39.6% to 37%, and to 29.6% for many via the new 20% Qualified Business Income deduction (a cumulative ~25% reduction). The authors use employer-employee matched federal tax records (corporate SOI files merged with W-2 and individual returns), tax years 2013-2019, on a balanced panel of large firms (&amp;gt;=50 employees and &amp;gt;=$1M sales each pre-period year): 15,490 firms and 108,430 firm-year observations. The main design is an event study / 2SLS comparing similarly sized C and S corps in the same industry-size bin, with firm and industry-size-year fixed effects and standard errors clustered by firm; entity-switchers are dropped. The identifying assumption is parallel trends absent the tax change (as in Yagan 2015), not random C/S assignment.&lt;/p&gt;
&lt;p&gt;First stage: C corps&amp;rsquo; marginal tax rate fell ~5.0 pp (s.e.=0.2) relative to S corps, raising the log net-of-tax rate ~6.6% (s.e.=0.2); C corps paid ~$2,100 (s.e.=341) less tax per worker. Real effects: C-corp sales rose 3.9 pp (s.e.=1.2) relative to S corps; pre-tax profits +3.0 pp (s.e.=0.7); after-tax profits +4.0 pp (s.e.=0.7); total payouts +21.9% intensive (s.e.=2.9) and +3.0 pp extensive (s.e.=0.5); employment +2.3% (s.e.=0.8); payrolls +3.4% (s.e.=0.8); net investment +2.9% (s.e.=0.4). The benchmark corporate elasticity of taxable income (pre-tax profits) is 0.46 (s.e.=0.11); after-tax-profit elasticity 0.61 (s.e.=0.11); investment elasticity 0.45 (s.e.=0.07). Worker earnings are flat for the bottom 90% (median wp50 coefficient -0.001, s.e.=0.004) but rise for the top 10%: +1.3% at the 95th percentile (s.e.=0.4), +4.8% at the 99th, and +4.8% for executives (top-5 paid; s.e.=0.7, earnings elasticity 0.73). Executive-pay gains barely shrink when controlling for firm performance (4.8% to 4.5%) and are concentrated among incumbents, consistent with rent-sharing rather than productivity.&lt;/p&gt;
&lt;p&gt;Responses concentrate in capital-intensive industries and are not larger for cash-constrained firms, pointing to a cost-of-capital channel rather than liquidity. Via a stylized model, a $1 marginal cut in corporate tax revenue generates $0.44 in additional output; revenue falls $0.85 per $1 mechanical loss (total -$86 billion, 0.40% of GDP). Factor incidence: 51% of gains to firm owners, 10% to executives, 38% to high-paid workers, 0% to low-paid workers. Across the income distribution, 80% of gains accrue to the top 10% and 20% to the bottom 90%, with gains concentrated in the Northeast/West and large high-income cities. The corporate tax is ~twice as inefficient as the personal income tax but similarly progressive, suggesting margin-of-efficiency gains from shifting toward personal income taxation. Results are short-run and abstract from public-goods provision and deficit financing.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The strategy is a difference-in-differences/event study (and 2SLS) comparing C corporations to S corporations in the same industry-size bin before and after TCJA, instrumenting the change in the log net-of-tax rate with pre-existing C/S entity status, with firm and industry-size-year fixed effects and firm-clustered standard errors. The identifying assumption is parallel trends in outcomes absent the tax change (not random C/S assignment), supported by (a) flat pre-trends in the event studies, (b) Yagan (2015) showing C and S trends were statistically indistinguishable 1996-2008, (c) the unexpected nature of TCJA before the 2016 elections limiting anticipation, and (d) industry-size-year fixed effects matching firms in similar product markets. Main threats: anticipatory/intertemporal tax shifting (some rate decline already in 2017; executive pay also trends up in 2017); other concurrent TCJA provisions (bonus depreciation, DPAD repeal, NOL/interest limitation, international); endogenous entity switching; differential industry-size composition; and general-equilibrium/SUTVA violations where C-corp gains could be S-corp mirror-image losses or where common wage effects are absorbed by time fixed effects.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The authors argue the dominant mechanism is a reduction in the cost of capital from the permanent rate cut, not liquidity relief and not primarily bonus depreciation. Evidence: (1) responses are larger in capital-intensive industries (profits and investment), consistent with the cost-of-capital first-order condition; (2) high-cash firms are if anything more responsive than low-cash firms, ruling out liquidity constraints (and thus income effects); (3) bonus depreciation is downweighted because many eligible firms do not claim it, much capital (intangibles, structures) is never fully expensed, C and S corps had near-identical expensing exposure (so the design differences them out), and the investment response is driven almost entirely by short-lived assets rather than the long-lived assets where accelerated depreciation is most valuable. A complementary dynamic-adjustment-cost model (Auerbach-Hassett 1992 with Foertsch 2018 cost-of-capital inputs) yields elasticities very similar to the benchmark.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;By capital intensity: C corps in capital-intensive industries show significantly larger profit and investment responses (supporting the cost-of-capital channel). By liquidity: high-cash firms are no less (if anything more) responsive than low-cash firms, contrasting with Zwick and Mahon (2017). By firm size: no clear pattern in profits, median earnings, or investment, with only suggestive evidence that high-income-worker gains are larger in smaller firms. By worker position: earnings gains are concentrated entirely in the top 10% of the within-firm distribution and especially in executives, with zero gains below the 90th percentile. By worker tenure: gains are driven by incumbents, not new hires (consistent with rent-sharing). Geographically: gains concentrate in the Northeast and West and in large high-income commuting zones (e.g., ~3x the median CZ gain in New York City, ~5x in the San Francisco Bay Area).&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Alternate specifications (Table 7): cohort(age)-by-year FE, state-by-year FE, firm-specific pretrend controls, 6-digit NAICS industries, reweighting S to match the C industry-size distribution, inverse-propensity-weighting, log-transformed outcomes, winsorizing at 5th/95th percentiles, and 2016-sales/payroll weighting — elasticities are stable. Alternate samples (Table 8): excluding firms with &amp;gt;$1B sales or &amp;gt;10,000 employees, excluding mismatched industries (C share &amp;gt;80% or &amp;lt;20%), excluding manufacturing (trade-war exposure), unbalanced panel, excluding public firms, excluding industries most exposed to DPAD/NOL/interest-limitation/bonus-depreciation provisions, excluding multinationals, dropping tax years 2017-2018 (anticipation/shifting), and dropping single-owner S corps (wage/profit reclassification). Entity switching rose only from ~0.1% to ~0.3% (profit-weighted) and is negligible. Most estimates stay within the benchmark confidence intervals.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q5. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It builds on the C-vs-S comparison design of Yagan (2015) but studies marginal corporate rate cuts rather than the 2003 dividend tax cut. It obtains an investment elasticity (0.45) very close to Chodorow-Reich et al. (2023)&amp;rsquo;s 0.52 despite a different identification strategy and sample. Its corporate ETI (0.46) is below state/local estimates (Giroud-Rauh ~0.50; Suarez Serrato-Zidar ~0.9; Bachas-Soto 3.0-5.0 in Costa Rica) but above typical personal-income ETIs (Saez et al. central 0.25), consistent with distortions scaling with factor mobility. Its incidence finding — that the corporate tax falls on capital and high-income workers — differs from Fuest et al. (2018), who find German municipal corporate tax hikes fall on low-skilled/marginally-attached workers (the authors note possible asymmetry between hikes and cuts and small-firm effects), and aligns with Risch (2024). It uses directly observed owner returns and the full earnings distribution, requiring weaker assumptions than Suarez Serrato-Zidar (2016, who infer owner returns structurally) and Fuest et al. (who assume negligible rental-rate changes).&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;On efficiency: a $1 cut in corporate tax revenue yields $0.44 of additional output, and current U.S. top corporate rates appear below the revenue-maximizing rate (revenue falls only $0.85 per $1 mechanical loss). The corporate tax is ~twice as inefficient as the personal income tax but similarly progressive, and 3-4x more progressive than the payroll tax while being 2-3x as inefficient — implying that shifting the federal revenue mix toward personal income taxes could raise efficiency without much loss of progressivity. On equity: the cuts are regressive in the short run, with 80% of gains to the top 10% (24% to the top 1%, 56% to the 90-99th percentiles), 0% to low-paid workers, and 17% flowing to foreign equity holders. Scope conditions: estimates are short-run (through 2019, pre-COVID); they hold welfare equal to output (ignoring utility curvature); they assume a representative consumer (no consumer-price channel) and equal redistribution of revenue; they abstract from deficit financing, public-goods provision, and long-run productivity/wage effects; and the very largest C corps have no S-corp analogue, so their responses are not well identified.&lt;/p&gt;
&lt;h3 id="q7-what-other-significant-findings-extensions-or-caveats-appear"&gt;Q7. What other significant findings, extensions, or caveats appear?&lt;/h3&gt;
&lt;p&gt;Employment increases reflect predominantly reallocation of workers across sectors rather than net new hiring, which the authors account for in the aggregate analysis (and is why incidence focuses on wages, not employment). New investment gains are in short-life assets (e.g., computers), with no change in long-life machinery or structures. Firms returned excess profits via dividends and buybacks but did not increase equity or debt issuance, and shareholder-payout results are robust to excluding multinationals (so the repatriation holiday is not the driver). Executive pay shifted forward into 2017 (bonuses) to be deducted at the higher pre-cut rate. Caveats flagged by the authors: rent-sharing tests are suggestive not dispositive (conditioning on post-treatment outcomes; unobserved hours/effort; short two-year horizon); private-income components are precisely estimated but the welfare confidence interval includes zero (up to ~0.4% of GDP); and long-run channels (productivity, lower prices, real wages) and offsetting cuts to public services/transfers are outside the analysis.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;C corporation vs. S corporation&lt;/strong&gt;: The two legal entity types whose divergent TCJA tax treatment provides identification. C corps pay corporate income tax directly (rate cut 35% to 21%) and their dividends are taxed at the shareholder level; S corps pass income through to up to 100 individual U.S. shareholders who pay ordinary income tax (top rate cut 39.6% to 37%, or 29.6% with QBI), with no corporate-level or dividend tax.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Implied marginal tax rate (for S corps)&lt;/strong&gt;: Because S corps pay no entity-level tax, their firm marginal rate is constructed as the ownership-share-weighted average of the individual marginal income tax rates of the firm&amp;rsquo;s owners, computed from linked personal returns (e.g., two equal owners at 25% and 35% imply 30%).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Corporate elasticity of taxable income (ETI)&lt;/strong&gt;: The percent change in the corporate tax base (pre-tax profits) per percent change in the net-of-tax rate; the paper&amp;rsquo;s benchmark is 0.46. Following Feldstein (1999), it summarizes the deadweight loss / efficiency cost of the tax under negligible income shifting and income effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Net-of-tax rate&lt;/strong&gt;: One minus the marginal tax rate, ln(1-tau); the object firms optimize against, used to scale reduced-form effects into elasticities. TCJA raised C corps&amp;rsquo; log net-of-tax rate by ~6.6% relative to S corps.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cost-of-capital channel&lt;/strong&gt;: The mechanism by which a lower tax rate (or higher expensing parameter theta) reduces the user cost of capital phi = r(1-theta*tau)/(1-tau), raising capital demand, labor demand, and firm scale — the paper&amp;rsquo;s preferred interpretation, distinguished from liquidity effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marginal excess burden&lt;/strong&gt;: dW/dT, the change in welfare (output, defined as private income plus tax revenue) per dollar of corporate tax revenue; estimated so that $1 of foregone corporate revenue generates $0.44 of additional output.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Incidence across the income distribution&lt;/strong&gt;: An extension of factor incidence that assigns owners&amp;rsquo; capital gains back to workers using the Distributional Financial Accounts (since many workers hold equity and many owners work), yielding the result that 80% of tax-cut gains accrue to the top 10% of earners.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rent-sharing&lt;/strong&gt;: The channel whereby earnings gains accrue to incumbent high-paid workers and executives rather than to new hires (the marginal unit of labor), with executive pay only weakly tied to firm performance — interpreted as workers/executives capturing a share of excess after-tax profits.&lt;/p&gt;</description></item><item><title>The Lost Marie Curies and Foregone Economic Growth</title><link>https://macropaperwarehouse.com/papers/the-lost-marie-curies-and-foregone-economic-growth/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-lost-marie-curies-and-foregone-economic-growth/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Women accounted for only 3% of U.S. inventors in 1976 and still just 14% in 2023, a pace of convergence far slower than in law (3% to 49%) or medicine (6% to 46%) over the same period. Under the natural assumption of no innate gender differences in inventive potential, this persistent underrepresentation reveals a misallocation of talent. The paper asks how costly this misallocation is for aggregate productivity and welfare.&lt;/p&gt;
&lt;p&gt;Brouillette develops an overlapping-generations (OLG) model of semi-endogenous growth in the spirit of Jones (1995), in which individuals with heterogeneous innate inventive talent choose sequentially among three decisions: (1) whether to pursue a STEM education (the prerequisite for research), (2) whether to work in research or production, and (3) whether to have children. Three gendered barriers can deter women from their comparative advantage. First, a labor market distortion, modeled as a tax on research earnings, captures discrimination in pay and credit attribution. Second, a child penalty distortion reduces mothers&amp;rsquo; hours in research relative to fathers, amplified by the &amp;ldquo;greedy job&amp;rdquo; nature of research (a premium on long hours). Third, an exposure distortion, modeled as a Bernoulli random variable, captures the probability of ever encountering inventive career opportunities — driven empirically by the absence of female role models.&lt;/p&gt;
&lt;p&gt;The model is calibrated to the U.S. economy using two data sources: PatentsView (all USPTO patents since 1976, covering roughly 1.7 million inventors and 3.7 million patents, with gender inferred from first names) and the U.S. Decennial Census/ACS (demographic and occupational data). Across these sources, female inventors exhibit only marginally higher research productivity than men (consistent with modest positive selection from the earnings tax), while mothers in research work approximately 4.5% fewer hours per week than childless female researchers (fathers work 2.7% more). The small productivity gap and modest hours gap together imply that neither the earnings tax nor the child penalty is the dominant driver; the exposure distortion is inferred as the residual, calibrated to a benchmark female share in research of 23% (average of 19% from PatentsView and 27% from Census/ACS). The resulting distortion estimates are: labor market tax 3.3%, child penalty 7%, and exposure barrier 79%.&lt;/p&gt;
&lt;p&gt;Counterfactual elimination of all three distortions raises U.S. income per person by 14.2% in the long run, compared with only 1.5% from a 30% R&amp;amp;D subsidy in a distortion-free economy. The gain materializes slowly, with a half-life of approximately 76 years, reflecting the semi-endogenous structure (where reallocating talent shifts the level but not the long-run growth rate of living standards) and the OLG structure (where career choices are irreversible, slowing labor reallocation). Aggregate research labor increases by 49% within the first 50 years of the transition — women&amp;rsquo;s research labor more than quadruples while men&amp;rsquo;s shrinks by about 10% — but almost all of the productivity gain operates through the intensive rather than the extensive margin: the aggregate share of inventors barely rises, because exposure barriers blocked many talented women entirely rather than only marginal ones, so lifting them introduces very high-quality new researchers who crowd out less talented men. If the underrepresentation were instead attributed entirely to selection-based barriers (labor market or child penalty), long-run consumption would rise by only 3.6%, less than a quarter of the baseline 14.2%.&lt;/p&gt;
&lt;p&gt;Taking transition dynamics into account, eliminating all distortions is equivalent to permanently raising everyone&amp;rsquo;s consumption by 7.2% (lower than 14.2% because the transition is slow and future gains are discounted back at a rate exceeding the low projected U.S. population growth). Of this welfare gain, 95% comes from higher mean consumption; the remainder comes from reduced consumption inequality and utility from children. The distribution of gains is unequal across time and demographic groups: future cohorts experience an 8.6% permanent consumption increase versus only 1% for surviving cohorts. Among the current generation of inventors, women gain the equivalent of a 1.3% permanent consumption increase while men lose 1.7%, a distributional tension that complicates implementation when current costs are concentrated and future benefits diffuse.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-three-distortions-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy for the three distortions, and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The three distortions are identified from three moments, each theoretically linked to a specific distortion through the model&amp;rsquo;s aggregation. The labor market distortion (earnings tax) is identified from the research productivity gender gap: positive selection under this tax implies women should be marginally more productive, and the magnitude of the observed (small) gap pins down a distortion of 3.3%. The child penalty distortion is identified from gender differences in hours worked between parent and non-parent researchers: mothers work 4.5% fewer hours than childless women while fathers work 2.7% more; after normalizing male distortions to zero, the model recovers a child penalty distortion of 7%. The exposure distortion is identified as the residual that explains remaining underrepresentation (23% female share in research) after accounting for the other two mechanisms; it is estimated at 79%. Key threats: (1) The gender productivity gap is measured from PatentsView, which uses name-based gender attribution and citation-weighted patents — both susceptible to gender bias (women are documented to receive 30% fewer citations than men with common names, and are 59% less likely to be credited with authorship on patents they contributed to), so the paper uses stock market valuation and textual similarity of patents as bias-resistant alternatives. (2) The exposure distortion is a residual and could capture other forces not in the model, including occupational preferences, gendered barriers to human capital retention, or mismeasurement of the female researcher share. (3) The model abstracts from the direction of innovation (unlike Einïo, Feng, and Jaravel 2022), so welfare effects through consumption-cost inequality across groups are not captured.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The three mechanisms operate through distinct theoretical channels, which allows moment-based identification. The labor market distortion works through selection on talent: if only highly talented women choose research despite earning below their marginal product, the female researcher pool should be right-shifted in the talent distribution, implying modestly higher measured productivity for women. The empirical counterpart is the gender gap in patent output (quality-weighted patents per career year), controlling for field fixed effects and team size. The child penalty works through hours worked: a higher opportunity cost of childbearing in research (amplified by greedy-job premiums) reduces mothers&amp;rsquo; time in research. The empirical counterpart is the gender gap in hours worked between parents and non-parents in research, from the Census/ACS. The exposure distortion works through the extensive margin of talent — it is a binary probability of ever having access to research as a career path, so it can block even the most talented women, unlike the other two distortions which induce selection. It is identified as the residual after the other two are estimated. The insight that the productivity gap is small and the hours gap is modest together rule out the first two as primary drivers, placing most explanatory weight on the exposure distortion.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-semi-endogenous-growth-framework-differ-from-an-endogenous-growth-approach-and-what-are-the-implications-for-the-results"&gt;Q3. How does the semi-endogenous growth framework differ from an endogenous growth approach, and what are the implications for the results?&lt;/h3&gt;
&lt;p&gt;In semi-endogenous growth (Jones 1995), the long-run per-capita growth rate equals n/[(sigma-1)(1-phi)], determined by population growth and idea difficulty, not by the quantity or quality of researchers. A reallocation of inventive talent therefore cannot raise the long-run growth rate but can raise the level of per-capita consumption by shifting the cumulative stock of ideas and thus the entire trajectory of living standards upward. This stands in contrast to endogenous growth models where reallocating talent can permanently raise the growth rate. The author justifies the semi-endogenous approach on two grounds: (1) despite sustained researcher-population growth in most advanced economies, the per-capita growth rate has not trended up; (2) the framework is qualitatively and quantitatively consistent with the documented fact that &amp;lsquo;ideas are getting harder to find&amp;rsquo; (Bloom et al. 2020, which estimates phi = -2.1 for the aggregate U.S. economy). The implication is that the paper finds more modest effects on productivity growth than prior endogenous-growth models, with the gain materializing entirely as a level shift with a long half-life of ~76 years. Einïo, Feng, and Jaravel (2022), using an endogenous growth model, find that barriers to female innovation reduce the growth rate by 1.4 percentage points; this paper&amp;rsquo;s semi-endogenous model finds a 14.2% level gain with no permanent growth rate effect.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-in-the-gender-gap-is-documented-empirically"&gt;Q4. What heterogeneity in the gender gap is documented empirically?&lt;/h3&gt;
&lt;p&gt;Field heterogeneity: Between the 1990 and 2020 inventor cohorts, the female share in chemistry and metallurgy rose from 13% to approximately 30%, while in fixed constructions and mechanical engineering it rose from under 5% to about 10%. Despite this, male-dominated fields accounted for about 53% of total patents granted in 2023. Importantly, when the inventive productivity gender gap is plotted against the female share across technological fields and cohorts, there is no significant relationship (the slope is -0.09 with a standard error of 0.2), implying selection-based barriers are not the primary driver of field-level disparities. Cohort heterogeneity: By cohort, the female share among new inventors rose from 7.5% (1990 cohort) to 17.6% (2020 cohort). Life-cycle heterogeneity: The inventive productivity gender gap (with women slightly ahead) is primarily a cohort effect rather than a within-career pattern; more recent cohorts show a somewhat larger productivity advantage for women at career onset, but the magnitude remains modest, which argues against gendered human capital depreciation as a leading explanation. Parental status heterogeneity: The fraction of female researchers who are mothers converged to the fraction of male researchers who are fathers over time (both around 40% by 2023, down from an 80% male vs. 40% female gap in 1960), suggesting research has become more accommodating. The child penalty in research (hours worked differential between parents and non-parents) has also narrowed over time and is smaller in research than in non-research occupations.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-conducted"&gt;Q5. What robustness checks are conducted?&lt;/h3&gt;
&lt;p&gt;Five sets of robustness exercises are reported. (1) Degree of increasing returns to scale (gamma): Jones (2002) estimates gamma from 0.05 to 0.33; Peters (2021) estimates 0.6. Across this range, the long-run consumption gain from eliminating all distortions ranges from about 2% to almost 27% for gamma going from 0.05 to 0.6. (2) Talent signal shape parameter (theta_s): With theta_s raised to 2 from the baseline 1.26 (implying greater scarcity of superstar inventors, so fewer marginal researchers are displaced), the long-run gain falls to 8.7% from 14.2%. (3) Demographic parameters (retirement rate d and entry rate b): Setting d to match expected working lives of 20 and 40 years (versus baseline 30) shifts the transition half-life by roughly 6-8 years, leaving long-run income unchanged but moving welfare gains slightly (7.6% or 6.9% vs. baseline 7.2%). (4) Knowledge spillover parameter (phi): Values of 0.5 and -6.2 (lower bound of Bloom et al.) are tested with sigma adjusted to hold gamma constant; long-run income gains remain at 14.2%, while the half-life varies modestly and welfare gains shift by at most 24 basis points. (5) Patent quality metrics: Three alternative measures of patent quality are used — stock market valuation (Kogan et al. 2017), textual &amp;lsquo;importance&amp;rsquo; (Kelly et al. 2021), forward citations, and unweighted counts. Results are consistent across measures, with the bias-resistant metrics (stock market valuation and textual importance) ruling out citation-based bias as a confound.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-relate-to-and-differ-from-einïo-feng-and-jaravel-2022"&gt;Q6. How does this paper relate to and differ from Einïo, Feng, and Jaravel (2022)?&lt;/h3&gt;
&lt;p&gt;Einïo et al. (2022) is the closest antecedent. That paper develops a two-sector endogenous growth model with heterogeneous consumer tastes and unequal access to innovation across sociodemographic groups including gender, finding that barriers to female innovation are responsible for an 18.2% difference in the cost of living between women and men and reduce the economic growth rate by 1.4 percentage points. Brouillette&amp;rsquo;s paper uses a semi-endogenous growth framework and arrives at a 14.2% long-run level gain in income per person and a 7.2% consumption-equivalent welfare gain, with no permanent effect on the growth rate. Beyond the growth framework, the paper extends the analysis to include labor market discrimination and a child penalty for female researchers, which Einïo et al. do not model. However, Brouillette&amp;rsquo;s model abstracts from the direction of innovation — the idea that women and men produce inventions differently tailored to different users&amp;rsquo; needs — which Einïo et al. show is quantitatively important for cost-of-living inequality. The two papers are therefore treated as providing complementary insights.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-role-model-externality-extension-and-how-does-it-change-the-results"&gt;Q7. What is the role-model externality extension, and how does it change the results?&lt;/h3&gt;
&lt;p&gt;In the baseline model, the exposure distortion is a fixed parameter representing the probability of ever encountering inventive career opportunities. In the extension, this probability is multiplied by a technology friction that depends on the fraction of same-gender and opposite-gender inventors in prior generations, with elasticities calibrated from Bell et al. (2018): own-gender elasticity 0.24 for girls, cross-gender elasticity approximately 0 (statistically insignificant in the underlying regression). This creates a positive externality: current inventors increase exposure probabilities for future cohorts of the same gender, but they are not compensated for this spillover, constituting a market failure. In the extended model, some of what was previously captured as the exposure distortion is now attributed to the technological friction from role model scarcity, and the residual exposure distortion is smaller. The counterfactual elimination of all distortions yields a more modest long-run income gain of 10.6% and a consumption-equivalent welfare gain of 3.8% (compared to 14.2% and 7.2% in the baseline). The role model externality also opens a rationale for temporarily gender-differentiated wage subsidies for female researchers as transitional optimal policy: a welfare-maximizing planner might accept a slightly worse talent allocation today in order to accelerate the expansion of the female role model base, reaching the efficient allocation sooner.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q8. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s central policy implication is that interventions targeting exposure to innovation for girls earlier in the pipeline — before entry into the labor market — offer far larger aggregate productivity returns than either conventional R&amp;amp;D subsidies or policies aimed at reducing workplace discrimination or the child penalty in isolation. A 30% R&amp;amp;D subsidy yields only 1.5% long-run income per capita growth versus 14.2% from full elimination of female research barriers. Within those barriers, the exposure distortion alone accounts for the bulk of the gain: if the underrepresentation were entirely due to the labor market or child penalty distortions (selection-based mechanisms), long-run gains would be only 3.6%. Scope conditions and caveats: (1) The framework is calibrated to the U.S. and to patent-based inventors plus Census-classified researchers, so generalization to other settings requires re-estimation of distortions. (2) The semi-endogenous structure implies that gains are level effects, not growth rate effects, and the half-life of ~76 years means that most gains accrue to future rather than current generations. (3) Distributional effects are asymmetric: the current generation of male inventors suffers a 1.7% consumption loss, while future cohorts broadly gain 8.6%; this temporal and demographic incidence complicates implementation. (4) The model abstracts from the direction of innovation, so welfare effects through differential cost-of-living impacts on men and women are not captured. (5) The role model externality extension suggests that affirmative action policies for female researchers may be warranted on efficiency grounds, but the exact form of optimal transitional policy is not fully characterized.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-greedy-job-mechanism-and-how-is-it-quantified"&gt;Q9. What is the &amp;lsquo;greedy job&amp;rsquo; mechanism and how is it quantified?&lt;/h3&gt;
&lt;p&gt;The &amp;lsquo;greedy job&amp;rsquo; concept (Goldin 2021) refers to occupations where extended, inflexible hours are compensated at a premium, making it suboptimal for couples to share labor supply equally and thus imposing a larger effective cost of parenthood on whoever reduces hours (in practice, more often women). In the model, an individual researcher&amp;rsquo;s effective labor supply is proportional to alpha^(1+delta) when they have children (where alpha = 0.93 is the fraction of time parents spend working and delta &amp;gt; 0 governs the additional return to hours in research). This magnifies the talent threshold required for a parent to prefer research over production. The parameter delta is estimated empirically by regressing log hourly wages on log hours worked, an indicator for research occupation, and their interaction (plus controls for age, experience, education, occupation, state, race, marital status, year, gender, and occupation-by-gender fixed effects), using the Census/ACS with over 11.8 million observations. The estimated delta for researchers is 0.004, statistically significant but modest — implying research is a &amp;lsquo;modestly greedy job,&amp;rsquo; less so than law (0.011) or medicine (0.006). This small value of delta constrains the child penalty distortion&amp;rsquo;s aggregate impact and helps explain why the exposure distortion dominates empirically.&lt;/p&gt;
&lt;h3 id="q10-how-is-research-productivity-measured-and-what-biases-are-addressed"&gt;Q10. How is research productivity measured, and what biases are addressed?&lt;/h3&gt;
&lt;p&gt;Research productivity is measured as average quality-weighted patents granted per year over an inventor&amp;rsquo;s career, with experience fixed effects removed before averaging across years. Three patent quality metrics are used: (1) stock market valuation (Kogan et al. 2017), inferred from abnormal stock returns around patent grant announcements — chosen for its resistance to gender bias because it reflects market assessments rather than subjective citation choices; (2) &amp;lsquo;importance&amp;rsquo; (Kelly et al. 2021), measured from textual similarity between patent pairs, rewarding novelty relative to prior patents and influence on subsequent ones, and also robust to citation bias because it would require precise paraphrase rather than mere omission; (3) forward citation counts, acknowledged as potentially biased (Jensen et al. 2018 show women with common names receive 30% fewer citations, while women with rare names receive 20% more); (4) unweighted patent counts. All metrics are adjusted for 3-digit CPC class fixed effects and co-inventorship team size. The results are consistent across all four measures, with women slightly ahead in all cases, suggesting that citation bias does not qualitatively alter the productivity comparison. A further concern is attribution bias: Ross et al. (2022) show women are 59% less likely to be credited with authorship on patents they contributed to, meaning PatentsView may undercount the true female inventor population.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-paper-say-about-the-stem-education-gender-gap-specifically"&gt;Q11. What does the paper say about the STEM education gender gap specifically?&lt;/h3&gt;
&lt;p&gt;Women account for approximately 35% of employed STEM graduates aged 25 to 45 in the Census/ACS data (and less than 20% of engineering graduates). However, this STEM gap alone explains only 7% of the patenting gender gap (Hunt et al. 2013, using the 2003 NSCG which recorded patenting in the prior five years); a substantial 78% of the gap stems from differences in patenting behavior among STEM graduates themselves. Furthermore, since the early 2000s, female researchers have been more likely than male researchers to hold a college degree, ruling out educational attainment differences as the primary driver. The model addresses STEM underrepresentation not through a gendered STEM education cost but through the exposure distortion, on the grounds that: (1) exposure to role models is well-documented as influencing girls&amp;rsquo; decisions to pursue STEM (Carrell et al. 2010; Breda et al. 2023; Bell et al. 2018); and (2) if a higher STEM cost were the primary barrier, the model would predict women to be substantially more productive than men (strong positive selection), which the data does not support.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Semi-endogenous growth&lt;/strong&gt;: A growth framework in which the long-run per-capita growth rate is determined by population growth and the difficulty of finding new ideas (the knowledge spillover parameter phi), not by the quantity or quality of researchers. Reallocating inventive talent shifts the level of living standards permanently but cannot alter the long-run growth rate; &amp;lsquo;ideas are getting harder to find&amp;rsquo; (phi &amp;lt; 0 in the paper&amp;rsquo;s calibration, phi = -2.1) is an integral feature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exposure distortion&lt;/strong&gt;: A Bernoulli random variable with mean (1 - tau_E_gk) governing whether an individual of gender g and cohort k ever encounters inventive career opportunities, regardless of their talent. In the baseline model it captures the aggregate probability of not having relevant role models or other enabling conditions during formative years; it is estimated at 79% for women (meaning only 21% of women are exposed to research as a potential career path). Unlike selection-based distortions, it blocks access to the innovation system even for the most talented women.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor market distortion&lt;/strong&gt;: A proportional tax tau_L on the research earnings of female inventors, representing discrimination in compensation, credit attribution, promotions, and rent-sharing from intellectual property. It induces positive selection: under this tax, only sufficiently talented women prefer research over production, making the average female researcher marginally more productive than the average male researcher. Estimated at 3.3%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Child penalty distortion&lt;/strong&gt;: A proportional reduction tau_C in the effective research hours of mothers, capturing the disproportionate burden of childcare and household responsibilities on women&amp;rsquo;s research careers. Combined with the &amp;lsquo;greedy work&amp;rsquo; parameter delta (the premium on long hours in research), it raises the talent threshold above which a woman who wants children will still choose a research career. Estimated at 7%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Greedy job&lt;/strong&gt;: An occupation, in the sense of Goldin (2021), where working long and inflexible hours is rewarded at a premium over and above what a simple proportional-hours model would predict. In the model, captured by the parameter delta &amp;gt; 0 in the research labor supply function. Estimated at delta = 0.004 for researchers (modest relative to lawyers at 0.011 or doctors at 0.006), implying that research is a modestly greedy job, amplifying the child penalty but not dominating the exposure distortion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive vs. extensive margin of research labor&lt;/strong&gt;: The extensive margin refers to the number (fraction) of people who choose research careers; the intensive margin refers to the average quality (talent-weighted hours) of researchers. The paper&amp;rsquo;s key finding is that the 14.2% long-run income gain from eliminating gender barriers is achieved almost entirely on the intensive margin: the aggregate share of inventors barely rises, but average researcher quality increases substantially because exposure barriers had been blocking the most talented women entirely.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption-equivalent welfare variation&lt;/strong&gt;: The permanent proportional adjustment lambda to every person&amp;rsquo;s consumption in the distorted economy that would make utilitarian social welfare equal to that in the undistorted economy. A lambda of 1.072 (7.2% gain) means permanently raising everyone&amp;rsquo;s consumption by 7.2% would compensate for remaining in the distorted equilibrium rather than transitioning to the undistorted one. It is lower than the 14.2% long-run income gain because the slow transition and the discounting of future population growth reduce the present value of future gains.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inventive productivity gender gap&lt;/strong&gt;: The difference in average quality-weighted patents per year between female and male inventors, after controlling for technological field fixed effects, experience, and co-inventorship team size. Measured across multiple patent quality metrics (stock market valuation, textual importance, forward citations, unweighted counts). In the paper&amp;rsquo;s data, the gap is positive but small — women are slightly more productive — which is the key empirical moment used to identify the (small) labor market distortion and to rule out large selection-based barriers as the primary driver of underrepresentation.&lt;/p&gt;</description></item><item><title>The Unequal Costs of Carbon Pricing: Economic and Political Effects Across European Regions</title><link>https://macropaperwarehouse.com/papers/the-unequal-costs-of-carbon-pricing-economic-and-political-effects-across-european-regions/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-unequal-costs-of-carbon-pricing-economic-and-political-effects-across-european-regions/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether carbon pricing through the EU Emissions Trading System (EU ETS) imposes economic costs that are unequally distributed across European regions, and whether those economic costs translate into political costs in the form of votes for extremist and populist parties. The motivation is both practical — political opposition has blocked or rolled back climate policies in several countries — and analytical: no prior study had systematically estimated the political consequences of carbon pricing at the subnational level.&lt;/p&gt;
&lt;p&gt;The authors build a panel dataset covering 224 NUTS2 regions from 20 European countries (covering 97% of EU GDP, plus Norway) over 2000–2019. Economic data come from the European Commission&amp;rsquo;s ARDECO database; emission data from EDGAR (aggregate GHG) and the EU ETS Transaction Log (verified ETS emissions from regulated installations, mapped to NUTS2 via zip codes); voting data from the EU-NED dataset with party classifications from The PopuList. Household expectations are measured from 34 Eurobarometer survey waves (2004–2019). The dataset spans 114 elections (110 national, four European Parliament).&lt;/p&gt;
&lt;p&gt;Identification rests on the carbon policy shocks of Kanzig (2023), constructed from high-frequency movements in EU carbon allowance futures prices around 126 regulatory events between 2005 and 2019, instrumented in a monthly VAR and aggregated to annual frequency. These shocks are orthogonal to contemporaneous economic conditions by construction, and are normalized so that the on-impact effect equals a 1% rise in Euro Area HICP energy prices. The main estimator is Jorda (2005) local projections in a panel with region fixed effects, lagged controls, and Driscoll-Kraay standard errors, estimated over a four-year horizon.&lt;/p&gt;
&lt;p&gt;Main economic findings (average region): A 1%-energy-price-equivalent carbon shock reduces real GDP by approximately 0.7% — a contraction that persists for four years. Employment, real net disposable household income, real GVA, real compensation, real investment, and hours worked all decline significantly and persistently. GHG emissions fall by roughly 1% one year after the shock, confirming the policy&amp;rsquo;s effectiveness.&lt;/p&gt;
&lt;p&gt;Main political findings: The combined extremist vote share (far-left plus far-right) rises by 0.3 to 0.4 percentage points two years after the shock and remains elevated. Populist and Eurosceptic vote shares also rise significantly in the medium term. Political fragmentation (1 minus the HHI) increases persistently. The shift is primarily toward far-right parties.&lt;/p&gt;
&lt;p&gt;Survey-based expectations: The share of respondents citing environmental issues as a top concern falls by approximately 2 percentage points and remains depressed for four years. Respondents become significantly more pessimistic about national economic and employment prospects and their own financial situation.&lt;/p&gt;
&lt;p&gt;Role of the economic channel: Using the Holm-Paul-Tischbirek (2021) decomposition, up to two thirds of the total rise in the extremist vote share over the four-year horizon is attributed to the decline in GDP, employment, and household income. The first year is more dominated by non-economic attribution effects (roughly 25% of the effect is explained by the economic channel at h=1), consistent with voters initially blaming the government&amp;rsquo;s policy choice rather than responding to realized economic deterioration.&lt;/p&gt;
&lt;p&gt;Regional heterogeneity and inequality: Regions one standard deviation above mean ETS emission intensity experience a meaningfully larger output contraction and a 20–50% larger and more persistent rise in the extremist vote share relative to the average region. Regions receiving fewer free ETS allowances face analogously larger economic and political costs. The within-country 90–10 ratio of real disposable household income rises by approximately 0.05 percentage points, with widening concentrated at the lower tail (the median-to-10th-percentile gap), meaning poorer regions bear disproportionate costs. These heterogeneous effects imply that carbon pricing contributes to regional inequality within countries.&lt;/p&gt;
&lt;p&gt;Policy implication: The EU ETS lacks direct redistribution mechanisms. The authors argue that progressive revenue recycling — household rebates calibrated to income — is necessary to cushion vulnerable regions, limit inequality, and rebuild public support for climate policy. These concerns are especially pressing given the EU ETS&amp;rsquo;s scheduled expansion to buildings and transportation in 2027.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The key identifying assumption is that the carbon policy shocks of Kanzig (2023) are exogenous with respect to regional economic conditions. The shocks are constructed from high-frequency daily movements in EU carbon allowance futures prices on days of regulatory announcements, relative to wholesale electricity prices on the prior day; the narrow event window ensures that confounding macroeconomic factors are already priced in. The shocks are then instrumented in a monthly VAR to extract structural shocks with a higher signal-to-noise ratio before being aggregated to annual frequency. The main threat would be if major regulatory announcements coincidentally coincided with other economic news. The authors defend against this by showing robustness to controlling for unemployment, stock market indices, monetary policy rates, oil prices, and a global financial crisis dummy. For the heterogeneity analysis, ETS intensity and free allowance share are fixed at their pre-sample values (end of ETS pilot phase, 2008) to rule out reverse causality from carbon pricing to the exposure measures.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-economic-voting-channel-distinguished-empirically-from-other-channels"&gt;Q2. How is the economic voting channel distinguished empirically from other channels?&lt;/h3&gt;
&lt;p&gt;The authors use the decomposition approach of Holm, Paul, and Tischbirek (2021). They re-estimate the extremist vote share local projection while controlling for the contemporaneous path of GDP, employment, and household income over the same h-year horizon. The residual coefficient on the carbon shock captures voting effects not attributable to economic deterioration. Comparing the controlled and uncontrolled responses shows that over the full four-year horizon, roughly two thirds of the voting increase is explained by economic variables. In the first year, the economic channel explains only about 25% of the response, consistent with non-economic attribution effects — voters blaming a government policy choice rather than an exogenous shock — being more prominent early on.&lt;/p&gt;
&lt;h3 id="q3-what-additional-evidence-distinguishes-ets-driven-political-effects-from-other-energy-price-effects"&gt;Q3. What additional evidence distinguishes ETS-driven political effects from other energy price effects?&lt;/h3&gt;
&lt;p&gt;Two benchmarks are used. First, national carbon taxes, which prior literature shows have muted economic effects, produce no statistically significant response in either real GDP or the extremist vote share (Appendix A.2), consistent with the economic channel being essential for the political response. Second, oil supply news shocks (Kanzig, 2021), constructed with a comparable high-frequency methodology and producing a similarly sized GDP decline, generate a statistically significantly smaller increase in the extremist vote share over the first two years (Appendix A.3). The excess political response to carbon shocks over oil shocks is interpreted as reflecting voters attributing policy-driven economic pain to the government, analogously to Gabriel, Klein, and Pessoa (2023) finding that austerity-induced recessions elicit stronger political responses than general downturns.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-across-regions-is-documented-and-how-is-it-measured"&gt;Q4. What heterogeneity across regions is documented and how is it measured?&lt;/h3&gt;
&lt;p&gt;Two exposure dimensions are explored. First, ETS emission intensity (verified ETS emissions scaled by GDP) captures direct agglomeration of installations covered by the carbon market. Second, the share of freely allocated ETS allowances relative to verified emissions captures the effective carbon price faced by firms in the region. Regions one standard deviation above mean ETS intensity experience meaningfully larger output and employment contractions, and 20–50% larger and more persistent increases in the extremist vote share. Regions with fewer free allowances bear analogously larger costs. Results hold when GHG intensity (covering non-ETS sectors) replaces ETS intensity, and when sectoral composition is controlled in the free allowance analysis. A country-level inequality analysis using local projections on the 90–10 ratio of regional household income shows that carbon pricing raises within-country dispersion by approximately 0.05 percentage points, driven primarily by widening of the lower tail (50th to 10th percentile gap), indicating that poorer regions suffer most.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Vote share results are robust to: (a) excluding parties coded as borderline by The PopuList; (b) excluding European Parliament elections and using only national elections; (c) averaging national and European election outcomes in years when both occur; (d) a minimal control set of only lagged dependent variable and region fixed effects; (e) an expanded control set adding country-level unemployment rate, stock market index, monetary policy rate, Brent oil price, and a GFC dummy variable. The inequality results are robust to using the 75–25 ratio and the Gini coefficient in addition to the 90–10 ratio. The heterogeneity results are robust to including time fixed effects, which absorb the aggregate carbon shock but preserve cross-sectional variation, confirming that heterogeneous responses are not driven by aggregate confounders. Driscoll-Kraay standard errors are used throughout to allow for cross-sectional and serial dependence; clustering at region-year level delivers nearly identical results.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q6. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Most directly related is Mangiante (2024), which documents that regions in poorer Euro Area countries are more exposed to carbon policy shocks. The present paper complements this by identifying within-country variation driven by ETS intensity and free allowance allocation, and by adding the political dimension. Kanzig and Konradt (2024) establish country-level economic effects of EU ETS shocks; this paper confirms those findings carry to the regional level and confirms comparable magnitudes. Gabriel, Klein, and Pessoa (2023) use the same econometric approach to study the political costs of austerity in European regions; the present paper finds analogous results for carbon pricing and attributes the political response similarly to economic deterioration. The finding that national carbon taxes lack economic or political bite echoes Metcalf and Stock (2023) and Konradt and Weder di Mauro (2023). The paper adds to the globalization-and-populism literature (Funke et al., 2016; Pastor and Veronesi, 2021; Colantone and Stanig, 2018) by identifying carbon pricing as another channel through which economic shocks drive extremist voting.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-direction-of-the-political-shift--toward-far-right-or-far-left"&gt;Q7. What is the direction of the political shift — toward far right or far left?&lt;/h3&gt;
&lt;p&gt;The decomposition in Appendix A.2 shows the increase in the combined extremist vote share is driven primarily by far-right parties. The far-right vote share rises significantly, while the far-left vote share shows a smaller and less precisely estimated increase. This is consistent with prior literature (Funke, Schularick, and Trebesch, 2016) documenting that far-right parties disproportionately benefit from recessions. A small decline in voter turnout is also documented, which may amplify measured increases in extremist vote shares by reducing the denominator (valid votes).&lt;/p&gt;
&lt;h3 id="q8-what-do-the-results-imply-for-environmental-concern-and-the-political-sustainability-of-climate-policy"&gt;Q8. What do the results imply for environmental concern and the political sustainability of climate policy?&lt;/h3&gt;
&lt;p&gt;Eurobarometer data show that the share of respondents ranking environmental issues among the two most important problems facing their country falls by approximately 2 percentage points following a carbon policy shock, a persistent decline lasting four years. The authors interpret this as a self-interest crowding-out effect: when carbon pricing imposes economic costs, concern for the environment is displaced by concern for living standards, consistent with Douenne and Fabre (2022). This creates a potential self-undermining dynamic: carbon pricing erodes the popular support needed to sustain and strengthen climate policy over time, particularly given that carbon-intensive regions — which suffer most economically — also see the largest decline in public support for environmental issues.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-scope-conditions-on-the-policy-implications"&gt;Q9. What are the scope conditions on the policy implications?&lt;/h3&gt;
&lt;p&gt;The findings pertain to ETS-style cap-and-trade pricing based on regulatory-driven supply restriction, not to national carbon taxes, which the paper shows have much smaller economic and political footprints. The sample covers 20 European countries with NUTS2 regional data over 2000–2019. The carbon policy shocks are derived from EU ETS regulatory events and are specific to that institutional context; generalization outside the EU ETS requires caution. Political effects operate primarily over a two-to-four-year horizon coinciding with electoral cycles. The paper&amp;rsquo;s redistribution prescription (progressive revenue recycling) presupposes a policy instrument capable of targeting household income; the EU ETS currently lacks such a mechanism, which is precisely the gap the authors flag as most urgent given the ETS expansion to buildings and transportation scheduled for 2027.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Carbon policy shock&lt;/strong&gt;: A series of exogenous regulatory surprises in EU ETS carbon allowance markets, constructed by Kanzig (2023) from high-frequency futures price movements around 126 regulatory events (2005–2019), instrumented in a monthly VAR, and normalized to produce a 1% on-impact increase in Euro Area HICP energy prices. Distinct from carbon price levels or oil shocks; isolates policy-driven changes in the supply of emission allowances, orthogonal to contemporaneous economic conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;ETS emission intensity&lt;/strong&gt;: Verified ETS emissions from regulated industrial installations in a NUTS2 region, scaled by regional GDP. The primary measure of a region&amp;rsquo;s direct exposure to EU carbon pricing; regions with higher ETS intensity experience larger economic contractions and larger shifts toward extremist parties when carbon prices rise.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Share of free allowances&lt;/strong&gt;: The ratio of freely allocated ETS emission permits to a region&amp;rsquo;s verified ETS emissions, used as a second regional exposure measure. A higher share implies a lower effective carbon price faced by firms; regions with fewer free allowances bear larger economic and political costs from carbon policy shocks. Free allowances were originally granted to protect energy- and trade-intensive sectors from rapid cost increases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extremist vote share&lt;/strong&gt;: The combined vote share of far-left and far-right parties in a region-election observation, using party classifications from The PopuList expert-coding database. The primary political outcome variable in the paper; empirically driven mainly by the far-right component in response to carbon policy shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Political fragmentation&lt;/strong&gt;: Defined in the paper as one minus the Herfindahl-Hirschman Index computed over all parties&amp;rsquo; vote shares in an election (1 − sum of squared vote shares). Captures the dispersion of votes across parties beyond the extremist vote share; used as a summary indicator of political polarization.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Economic voting channel&lt;/strong&gt;: The mechanism by which voters respond to carbon-pricing-induced economic deterioration — falling GDP, employment, and household income — by shifting support away from mainstream parties toward extremist alternatives. Isolated empirically via the Holm-Paul-Tischbirek (2021) decomposition; accounts for approximately two thirds of the total extremist voting response over the four-year impulse response horizon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regional inequality (90–10 ratio)&lt;/strong&gt;: Within-country dispersion of regional real disposable household income (or employee compensation) measured as the difference between the 90th and 10th percentile NUTS2 regions. Carbon pricing raises this measure persistently, with widening concentrated at the lower tail (the median-to-10th-percentile gap), indicating that poorer regions bear disproportionate economic costs.&lt;/p&gt;</description></item><item><title>The Winners and Losers of Climate Policies: A Sufficient Statistics Approach</title><link>https://macropaperwarehouse.com/papers/the-winners-and-losers-of-climate-policies-a-sufficient-statistics-approach/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-winners-and-losers-of-climate-policies-a-sufficient-statistics-approach/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks who wins and loses from climate policies — carbon taxes, renewable subsidies, and carbon tariffs — across 193 heterogeneous countries, and by how much. The motivation is that the standard IAM literature aggregates welfare into a global number, obscuring the distributional structure that determines political feasibility. Without knowing which countries gain and lose, and through which channels, it is impossible to understand why international cooperation is so difficult or which club structures can sustain themselves.&lt;/p&gt;
&lt;p&gt;The authors build a static Integrated Assessment Model (IAM) with heterogeneous countries, international trade in goods (Armington CES), international trade in fluid fossil (oil and gas), locally traded coal, and locally supplied renewables. Production uses a nested CES combining labour with a composite of three energy types. A reduced-form climate system maps world emissions linearly to global temperature, then to country-specific local temperatures, which damage TFP through a quadratic damage function. The key methodological contribution is a first-order (log-linear) decomposition of welfare around the current equilibrium, which expresses welfare changes analytically as a function of five observable sufficient statistics: (i) direct TFP damage, (ii) export terms-of-trade, (iii) import price index, (iv) energy cost effects (change in energy prices faced by producers), and (v) energy rent effects (change in profits of domestic fossil and renewable producers). This decomposition requires no model simulation; it reads off welfare directly from observables and a small set of elasticities.&lt;/p&gt;
&lt;p&gt;Two sets of structural parameters are estimated. First, a structural damage function is estimated using bilateral trade data from the ITPD-E dataset (2000–2016, 169 countries) via a Poisson pseudo-maximum-likelihood gravity regression that instruments temperature shocks against within-trading-partner variation in import penetration, controlling for energy market effects. The preferred specification recovers a global peak temperature of T* = 14.02°C and a damage slope parameter γ = 0.012. This strategy is designed to be robust to the Lucas critique: unlike reduced-form GDP regressions, it nets out general-equilibrium spillovers through trade and energy channels. Second, country-specific energy supply elasticities for oil-gas and coal are estimated from time-series variation in fossil rent shares and international prices (1985–2019 data), using OLS country-by-country and then an empirical Bayes shrinkage procedure with a truncated-normal prior that enforces positive elasticities. Coal is found to be substantially more elastically supplied than oil-gas; OPEC nations (e.g., Saudi Arabia) have near-inelastic oil-gas supply, while the US has relatively elastic supply.&lt;/p&gt;
&lt;p&gt;Key quantitative results from the policy experiments follow. (1) Business-as-usual: a 3°C warming by 2100 generates a 17% loss in consumption-equivalent world welfare under utilitarian weights, implying a Social Cost of Carbon of $203/tCO₂ at the current equilibrium point-of-approximation, rising to $302/tCO₂ if computed at 3°C of warming. Under Negishi (income-proportional) weights, the SCC falls to $3.31, reflecting that damages are concentrated in low-income countries with high marginal utility. Winners include Canada and Russia; losers are concentrated in Africa, Latin America, and South-East Asia. (2) Unilateral carbon tax (China, $50/tonne): global emissions rise by less than 0.07% (not fall) because China&amp;rsquo;s carbon tax shifts its energy mix from coal toward oil-gas (coal is ~1.44× dirtier per unit of energy), raising the international oil-gas price by approximately 5%, which boosts fossil exporters&amp;rsquo; rents and induces other countries to substitute back to coal. Global utilitarian welfare falls by 0.2%. China itself gains on net through falling coal prices and improved terms of trade. EU nations lose from higher energy import costs. (3) Unilateral carbon tax (USA, $50/tonne): global emissions fall by 0.8%; US welfare effects are small but positive (energy cost increases largely offset by terms-of-trade gains with Canada and Europe). (4) Renewable subsidies (42.6%, calibrated to produce the same average relative-price shift as a $50 carbon tax): on average substantially less effective than carbon taxation and more harmful to welfare because subsidies push countries up their upward-sloping domestic renewable supply curves, wasting resources on costly domestic generation (especially in countries with high baseline renewable shares such as France). (5) EU climate club ($50 carbon tax + CBAM tariffs): global emissions fall by 3%; global utilitarian welfare rises by around 5% (1% under Negishi weights), but the EU itself is a net loser — only Southern Europe (Spain, Portugal, Italy) gains; Germany and Scandinavian nations lose both from direct policy costs and from cooling that harms countries that benefit from warming. Oil-gas price falls by 4.6% within the club. (6) ASEAN climate club (same structure): global emissions fall by 0.5%; global utilitarian welfare rises by about 0.8% (0.2% Negishi); ASEAN members broadly benefit because they are already losers from climate change and the carbon-reduction benefit outweighs policy costs. Oil-gas price falls by 0.6%. (7) Global $50 carbon tax (all 193 countries): global emissions fall by 3.82%; global oil-gas price rises by 0.96% (substitution from coal toward oil-gas under a global carbon tax); global utilitarian welfare rises by about 6% (1% Negishi). Most of the utilitarian gain reflects reduced international inequality, since benefits concentrate in low-income tropical countries. Fossil exporters such as Saudi Arabia and Nigeria see energy rents rise as coal is substituted for by oil-gas globally.&lt;/p&gt;
&lt;p&gt;The central mechanism finding is that leakage operates primarily through energy trade, not goods trade: energy market effects are consistently larger than goods-market terms-of-trade effects across all policy experiments. This quantifies why unilateral climate policy is so limited in effectiveness. International coordination through climate clubs overcomes leakage but creates winners and losers within member coalitions depending on each member&amp;rsquo;s energy mix, trade exposure, and baseline climate damage.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-for-the-structural-damage-function-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy for the structural damage function and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The authors estimate the damage function using a Poisson pseudo-maximum-likelihood gravity regression on bilateral import penetration ratios (Xij/Xii) as a function of temperature differences between exporters and importers (and their squares), with country-pair fixed effects and year fixed effects. Controls for GDP/capita (polynomial), oil rent share, and renewable energy share proxy for the time-varying component of factory-gate prices driven by energy prices and wages. The key identifying assumption is that conditional on these controls and fixed effects, temperature shocks are uncorrelated with time-varying bilateral preference or cost shifters. Threats include: (1) confounding time-varying bilateral shocks correlated with temperature, such as ENSO events or specific geopolitical shocks; (2) the possibility that global (rather than local) temperature drives damages, which the paper cannot address given limited time-series variation and potential spurious correlation concerns (following Goulet Coulombe and Klieber, 2025); (3) the treatment of θ = 5 as a known parameter in computing γ from the regression coefficient, which propagates calibration error. The authors argue their strategy is robust to the Lucas critique because it nets out general-equilibrium effects on GDP that would contaminate GDP-based damage regressions.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-papers-welfare-decomposition-work-and-what-are-its-five-channels"&gt;Q2. How does the paper&amp;rsquo;s welfare decomposition work and what are its five channels?&lt;/h3&gt;
&lt;p&gt;The welfare decomposition is a first-order log-linearisation of the indirect utility around the current equilibrium. Changes in consumption-equivalent welfare for country i decompose into: (i) direct climate TFP damage (change in Dy_i); (ii) export terms-of-trade effect (change in domestic good price p_i); (iii) import price-index effect (change in price index P_i); (iv) energy cost effects (changes in oil-gas price q^f, coal price q^c_i, and renewable price q^r_i weighted by their shares in production); and (v) energy rent effects (changes in profits from fossil, coal, and renewable extraction weighted by their shares in household income). The key insight is that none of these five terms requires solving the full model; each can be computed from observable data moments (energy mix, energy rent shares, trade shares) and a small number of estimated or calibrated elasticities.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-in-climate-damages-is-documented-and-what-drives-it"&gt;Q3. What heterogeneity in climate damages is documented and what drives it?&lt;/h3&gt;
&lt;p&gt;Winners from climate change (3°C warming) are primarily cold countries: Canada, Russia, Scandinavian nations. Losers are concentrated in Africa (Djibouti, Niger, Burkina Faso, Sudan), Latin America, and South-East Asia. The heterogeneity arises from: (1) differences in baseline temperature relative to the estimated global peak productivity temperature T* = 14.02°C; countries hotter than T* lose productivity with further warming, while colder countries gain; (2) partial local adaptation (αT = 0.5) so each country&amp;rsquo;s effective peak temperature is halfway between T* and its current local temperature; (3) indirect effects through trade networks — cold, open economies can lose if major trading partners are damaged; (4) energy rent effects — fossil exporters lose energy rents as warming reduces global energy demand, partially offsetting their direct productivity gains.&lt;/p&gt;
&lt;h3 id="q4-why-does-chinas-unilateral-carbon-tax-at-50tonne-raise-global-emissions-rather-than-lower-them"&gt;Q4. Why does China&amp;rsquo;s unilateral carbon tax at $50/tonne raise global emissions rather than lower them?&lt;/h3&gt;
&lt;p&gt;China relies heavily on coal, which has a carbon concentration ratio of approximately ξc/ξf ≈ 1.44 (coal is ~44% dirtier per unit energy than oil-gas). A carbon tax on both fuels raises the effective cost of coal more than oil-gas, inducing China to substitute toward oil-gas imports. This raises the international oil-gas price by approximately 5%, which: (1) increases energy rents for fossil exporters (Gulf states, Russia) and (2) makes oil-gas costlier for other countries, incentivising them to substitute back toward coal. The net effect on global emissions is a slight increase of less than 0.07%, rather than a decline. This is the carbon leakage effect operating through energy trade.&lt;/p&gt;
&lt;h3 id="q5-why-are-renewable-subsidies-substantially-less-effective-than-carbon-taxes"&gt;Q5. Why are renewable subsidies substantially less effective than carbon taxes?&lt;/h3&gt;
&lt;p&gt;Several mechanisms distinguish the two policies. First, a carbon tax directly raises the relative price of all fossil fuels versus renewables and pushes production up the upward-sloping renewable supply curve only modestly. A renewable subsidy instead directly subsidises a reduction in the cost of renewables, which expands renewable supply — but this requires moving up the domestic renewable supply curve, wasting real resources in countries where the marginal renewable site is expensive (e.g., France with over 40% baseline renewable share). Second, a carbon tax creates a reallocation from coal to oil-gas (since the tax raises the coal price more per unit of energy), which can inadvertently raise oil-gas prices and redistribute income to exporters. A renewable subsidy does not have this feature in the same way. Third, the lump-sum financing of subsidies has a direct income cost, while carbon tax revenues are rebated, so only general equilibrium price effects matter for welfare. On average across countries, renewable subsidies cause more harm and generate smaller emission reductions per dollar.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-distinction-between-the-eu-and-asean-climate-clubs-and-why-do-outcomes-differ-so-substantially"&gt;Q6. What is the distinction between the EU and ASEAN climate clubs, and why do outcomes differ so substantially?&lt;/h3&gt;
&lt;p&gt;The EU club ($50 carbon tax + CBAM on imports from non-members) reduces global emissions by 3%, raises global utilitarian welfare by about 5%, but makes EU members net losers on average. The reason is that EU countries include many cold nations (Germany, Scandinavia) that benefit from warming; by cooling the climate, the policy harms them. Additionally, energy cost effects within the EU are heterogeneous — energy costs rise in France but fall in Poland and Germany — and Ireland is harmed through goods trade with Great Britain. The ASEAN club reduces global emissions by only 0.5% (ASEAN is smaller and less fossil-intensive in global terms), raises global utilitarian welfare by 0.8%, and ASEAN members broadly benefit because: (1) all ASEAN members are in the tropical/sub-tropical zone and thus lose from warming; (2) reducing global temperature yields direct productivity gains for members; (3) the energy rent loss for fossil exporters within ASEAN (Brunei, Indonesia) is outweighed by the climate benefit for others. The key structural difference is that the ASEAN club&amp;rsquo;s members are already losers from warming and hence have aligned incentives for carbon reduction.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-social-cost-of-carbon-computed-in-this-framework-and-how-does-it-vary-with-assumptions"&gt;Q7. What is the Social Cost of Carbon computed in this framework and how does it vary with assumptions?&lt;/h3&gt;
&lt;p&gt;Under utilitarian Pareto weights (ωi = 1, equal weight per person) and a 3°C warming by 2100, the global consumption-equivalent welfare loss is 17%, implying SCC = $203/tCO₂ at the current baseline temperature. Changing the point of linearisation to the 3°C warmer world raises the SCC to $302/tCO₂, indicating that damages accelerate as warming progresses and that the baseline approximation understates future costs. Under Negishi weights (proportional to income, ωi ∝ 1/u&amp;rsquo;(ci)), the SCC falls dramatically to $3.31/tCO₂, because damages are concentrated in low-income countries which receive little weight under income-proportional welfare aggregation. The authors note their static, log-linearised model provides a lower bound: fully dynamic IAMs with nonlinearities, uncertainty, or catastrophic-tail risks would further raise the SCC.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-estimate-energy-supply-elasticities-and-what-are-the-key-findings"&gt;Q8. How does the paper estimate energy supply elasticities and what are the key findings?&lt;/h3&gt;
&lt;p&gt;The authors regress changes in the oil-gas rent share of GDP on changes in the international oil-gas price (and changes in GDP as a control) country-by-country using first differences, recovering country-specific supply elasticities. Because some OLS estimates are noisy, negative, or below 1 (implying negative supply elasticity, inconsistent with theory), the authors apply an empirical Bayes shrinkage procedure: they impose a truncated-normal prior (truncated below 1) whose hyperparameters come from a pooled regression, and compute the posterior mean for each country. Key findings: oil-gas supply is nearly inelastic in OPEC nations (Saudi Arabia) and Russia and China, consistent with market power compressing effective supply elasticity; the US has relatively elastic oil-gas supply. Coal supply is substantially more elastic on average than oil-gas; the US and India have relatively inelastic coal supply; Russia and China have more elastic coal supply. Coal rents never exceed 1% of GDP even in the largest producers, consistent with near-competitive flat supply curves. These spatial patterns matter significantly for which countries gain or lose from energy price changes induced by climate policy.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-main-mechanism-through-which-leakage-operates--energy-trade-or-goods-trade--and-how-is-this-established"&gt;Q9. What is the main mechanism through which leakage operates — energy trade or goods trade — and how is this established?&lt;/h3&gt;
&lt;p&gt;The paper establishes that energy market effects are consistently larger in magnitude than goods-market terms-of-trade effects across all policy experiments (see Appendix Table A3). Leakage through energy trade operates because: (1) a domestic carbon tax reduces domestic demand for fossil fuels, lowering the international price of oil-gas (for small countries) or shifting demand between fuels; (2) lower oil-gas prices benefit importing countries and encourage them to use more fossil fuels, partially offsetting the original emission reduction. Goods-market leakage (productivity and competitiveness effects through the trade network) exists but is secondary. This finding has implications for policy: carbon border adjustment mechanisms (CBAMs) target goods trade leakage, but the model suggests the larger channel — energy trade leakage — is not addressed by CBAM alone.&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-or-sensitivity-analyses-does-the-paper-report"&gt;Q10. What robustness checks or sensitivity analyses does the paper report?&lt;/h3&gt;
&lt;p&gt;The paper reports several robustness exercises: (1) The damage function estimation reports results under OLS (Columns 1-2) and Poisson (Columns 3-4), with separate or restricted coefficients on importer and exporter temperatures; the preferred Poisson specification with restricted coefficients yields T* = 14.02 and γ = 0.012, and the separate-coefficient specification yields statistically indistinguishable estimates. (2) The SCC is computed at two points of approximation — the current baseline and a 3°C warmer world — yielding $203 and $302/tCO₂ respectively, giving a sense of nonlinearity bias from log-linearisation. (3) Welfare is reported under both utilitarian (ωi = 1) and Negishi (ωi ∝ 1/u&amp;rsquo;(ci)) weights throughout, and the results differ sharply, highlighting how inequality weighting matters. (4) The partial local adaptation parameter αT = 0.5 nests pure global peak (αT = 1) and pure local baseline (αT = 0) damage specifications. (5) Appendix Table A3 provides a comprehensive decomposition of welfare into climate, energy, and trade effects for all six policy scenarios (BAU, global carbon tax, China tax, US tax, EU club, ASEAN club), enabling consistency checks across experiments.&lt;/p&gt;
&lt;h3 id="q11-how-does-this-paper-relate-to-the-broader-literature-on-iams-and-sufficient-statistics"&gt;Q11. How does this paper relate to the broader literature on IAMs and sufficient statistics?&lt;/h3&gt;
&lt;p&gt;The paper makes three connections. First, it is related to the large IAM literature (Nordhaus and Yang 1996; Barrage and Nordhaus 2024; Cruz and Rossi-Hansberg 2024) but differs by explicitly decomposing welfare into observable sufficient statistics, avoiding the need to solve a large dynamic system. Second, it is related to the sufficient statistics literature in trade (Lashkaripour 2021 on trade wars; Baqaee and Farhi 2024 on trade barriers; Kleinman, Liu, and Redding 2024 on productivity shocks in trade models) — the paper extends this approach to a broad set of climate instruments in a model with detailed energy markets. Third, it differs from Bourany (2025) — a companion paper by one author — which solves for optimal climate agreement design; the present paper instead uses sufficient statistics to evaluate many given policies, trading optimality for analytical tractability and decomposability. The paper also distinguishes from Krusell and Smith (2022), which does not allow cross-border energy trade, and from Cruz and Rossi-Hansberg (2024), which does not model heterogeneous energy rents across space.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-scope-conditions-and-limitations-of-the-approach"&gt;Q12. What are the scope conditions and limitations of the approach?&lt;/h3&gt;
&lt;p&gt;Scope conditions and limitations are significant. (1) The model is static, so it cannot capture dynamic considerations: optimal intertemporal extraction paths, green paradox effects (whether carbon taxes accelerate fossil extraction), directed innovation toward renewables, adaptation capital accumulation, or dynamic leakage in energy markets. (2) The first-order log-linearisation abstracts from nonlinearities in the climate system, making the results most relevant as marginal effects near the current equilibrium rather than for large climate-policy changes or for evaluating policies at future, warmer states of the world. (3) The paper does not model market power in international energy markets (OPEC behaviour), abstracting from strategic behaviour by fossil exporters. (4) Labour is internationally immobile, so migration as a margin of adaptation is excluded. (5) Utility damages from climate change (mortality, amenity loss) are excluded — only productivity (TFP) damages are modelled; including utility damages would amplify gains and losses proportionally. (6) The framework cannot evaluate dynamic policy environments such as climate coordination with commitment problems or intergenerational redistribution from carbon taxation.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-policy-implications-of-the-papers-findings"&gt;Q13. What are the policy implications of the paper&amp;rsquo;s findings?&lt;/h3&gt;
&lt;p&gt;Several policy implications follow from the paper&amp;rsquo;s results, with important scope conditions. (1) Unilateral climate policy is largely ineffective for reducing global emissions and can even increase them (as in China&amp;rsquo;s carbon tax case); the standard free-rider analysis understates the problem because energy-market leakage can reverse the direction of emissions. (2) Renewable energy subsidies are generally a worse policy instrument than carbon taxes, because they push countries up costly domestic supply curves rather than reallocating away from fossil fuels through price signals; policy prescriptions that favour subsidies (such as the US Inflation Reduction Act) should account for this comparative inefficiency. (3) Climate clubs with both a domestic carbon tax and carbon tariffs (CBAMs) can overcome leakage effects and yield positive global welfare gains, but impose net costs on members whose composition makes them net losers from cooling (cold, energy-exporting member nations). This suggests club membership incentives are heterogeneous even within a bloc and require side payments or complementary redistribution to be stable. (4) ASEAN-style clubs where all members are hot-country losers from warming can achieve a Pareto-improvement for members while also improving global welfare, making them potentially more robust to free-riding than clubs like the EU where some members prefer a warmer climate. (5) The SCC estimated under utilitarian weights ($203/tCO₂) is substantially higher than under Negishi weights ($3.31/tCO₂), implying that the appropriate SCC for policy depends critically on how inequality across countries is weighted in the social welfare function.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Sufficient statistics (for climate policy)&lt;/strong&gt;: In this paper&amp;rsquo;s sense, a set of observable data moments and estimable elasticities — specifically nations&amp;rsquo; energy mix (shares of oil-gas, coal, renewables), energy rent shares of GDP, bilateral trade shares, energy supply and demand elasticities, and damage parameters — that fully characterise, to the first order, the welfare impact of a climate policy change without requiring the full model to be solved. The approach follows Chetty (2009) and extends it from tax incidence to climate policy in an IAM with trade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Carbon leakage&lt;/strong&gt;: In this paper&amp;rsquo;s framework, the phenomenon by which a unilateral domestic carbon tax reduces domestic fossil demand and lowers the international price of oil-gas, inducing countries outside the policy to increase their fossil fuel consumption, partly or fully offsetting the original emission reduction. The paper shows leakage operates primarily through energy trade (oil-gas price channel) rather than through goods trade competitiveness effects, with energy effects consistently dominating in magnitude across all policy experiments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Cost of Carbon (LCC)&lt;/strong&gt;: The country-specific welfare cost of an additional unit of global carbon emissions, measured in monetary units as the negative of the partial derivative of country i&amp;rsquo;s welfare with respect to aggregate emissions, divided by the marginal utility of consumption. Distinct from the global Social Cost of Carbon (SCC), which aggregates LCCs across countries with Pareto weights. Countries whose productivity is harmed more by warming have a higher LCC; cold countries may have a negative LCC (they benefit from marginal warming).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural damage function&lt;/strong&gt;: The function Dy_i(E) mapping world cumulative emissions E to country i&amp;rsquo;s TFP via a quadratic temperature-productivity relationship with peak temperature T* and slope parameter γ, estimated in this paper from bilateral trade data (import penetration ratios and temperature differences) rather than from GDP-temperature regressions. The estimation is designed to be robust to the Lucas critique by netting out general-equilibrium propagation through trade and energy markets that would bias GDP-based estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Climate club&lt;/strong&gt;: In this paper&amp;rsquo;s usage (following Nordhaus 2015), a coalition of countries that jointly impose a domestic carbon tax on their own emissions and levy carbon tariffs (carbon border adjustment mechanism, CBAM) on imports from non-member countries scaled by the carbon intensity of those imports. The paper studies EU and ASEAN climate clubs and finds they differ sharply in welfare distribution: the EU club creates net losers among members (because some EU countries benefit from warming), while the ASEAN club delivers welfare gains for all members because all are hot-country losers from climate change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Energy rent effect&lt;/strong&gt;: The component of the welfare decomposition arising from changes in profits of domestic energy producers (fossil extractors, coal producers, renewable firms) due to changes in energy prices. Captured in the sufficient statistics formula as the profit share of GDP weighted by the relevant price change. Fossil-fuel-exporting countries have large positive exposure to oil-gas price increases (gains from price rises) and are harmed when global carbon policy reduces the fossil price — this is a key redistribution channel distinct from both climate damages and goods trade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Bayes shrinkage (energy supply elasticities)&lt;/strong&gt;: In this paper, a procedure that estimates country-specific fossil and coal supply elasticities by first running OLS regressions of rent share changes on price changes country-by-country, then shrinking noisy or negative estimates toward a pooled mean by imposing a truncated-normal prior (truncated below 1 to enforce positive elasticities) and computing posterior means. Used because country-level time series are short and noisy, while the prior encodes the theoretical constraint that supply must be upward-sloping.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negishi weights vs. utilitarian weights&lt;/strong&gt;: Two distinct social welfare aggregation methods used throughout the paper to aggregate country-level welfare changes into global welfare. Utilitarian weights (ωi = 1 per person) put equal importance on each person globally, so welfare gains in low-income tropical countries count fully; this yields high SCCs ($203/tCO₂) and large global welfare gains from carbon taxation. Negishi weights (ωi ∝ 1/u&amp;rsquo;(ci), proportional to income) downweight poor countries and upweight rich ones, yielding dramatically lower SCCs ($3.31/tCO₂) and smaller measured global welfare gains because damages concentrate in low-income countries that receive little weight.&lt;/p&gt;</description></item><item><title>Train to Opportunity: the Effect of Infrastructure on Intergenerational Mobility</title><link>https://macropaperwarehouse.com/papers/train-to-opportunity-the-effect-of-infrastructure-on-intergenerational-mobility/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/train-to-opportunity-the-effect-of-infrastructure-on-intergenerational-mobility/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether proximity to transport infrastructure can sever the occupational tie between parents and children — a question with direct bearing on the debate over place-based versus people-based policies. The authors exploit the nineteenth-century expansion of the railroad network across England and Wales, a setting where the First and Second Industrial Revolutions were remaking the occupational structure at the same time that the railroad was knitting together local labor markets and enabling geographic mobility.&lt;/p&gt;
&lt;p&gt;The empirical strategy centers on a novel dataset of close to 980,848 father-son pairs constructed from the full digitized population censuses of England and Wales in 1851, 1881, and 1911 (I-CeM project). Individuals are tracked across consecutive censuses using the Abramitzky-Mill-Perez (2019) linking procedure, which achieves match rates of 43–50% for men aged 40–52. Crucially, each individual is geolocated to the street level by matching census addresses to the GB1900 gazetteer, allowing railroad access to be measured as the straight-line distance from the childhood residence to the nearest train station — a finer measure than the district-level presence indicators used in prior work. Sons&amp;rsquo; occupations are observed at ages 40–52; fathers&amp;rsquo; occupations are measured 30 years earlier when sons were aged 10–22. Occupational mobility uses two complementary scales: HISCO categories (farming, laborer, services, sales, clerical, managerial, professional) and the continuous HISCAM social-interaction-distance ranking (scores 28–99, mean 50, SD 10).&lt;/p&gt;
&lt;p&gt;The key endogeneity problem is that railroad companies targeted low-density, cheap land, and that wealthy landowners and local politicians influenced station placement. To isolate exogenous variation, the authors construct a dynamic least-cost path (DLCP) network connecting 53 major towns identified by their 1801 populations (top 10% of the population distribution, threshold 9,172 inhabitants). The DLCP assigns slope costs to 50x50 meter grid cells and finds the minimum-cost path between every town pair. Lines are ranked by betweenness centrality to separate &amp;ldquo;early&amp;rdquo; 1851 lines from &amp;ldquo;late&amp;rdquo; 1881 lines, giving a time-varying instrument. Proximity to the nearest DLCP line is used as the instrument for proximity to the nearest actual train station, with standard errors clustered at the parish level. Controls include county and census-year fixed effects, distance to the nearest 1801 major town and its population, distance to Roman roads, ancient ports, and navigable waterways, plus household characteristics (number of servants as a wealth proxy, household size, and father&amp;rsquo;s foreign birth).&lt;/p&gt;
&lt;p&gt;Main results (preferred IV specification with full controls): sons who grew up one standard deviation — approximately 5 km, or about one hour&amp;rsquo;s walk — closer to a train station were 11 percentage points more likely to work in an occupation category different from their father&amp;rsquo;s. They were 5 percentage points more likely to be upwardly mobile, defined as a son&amp;rsquo;s HISCAM score exceeding his father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s distribution. The downward mobility estimate is 3 percentage points — positive but smaller in magnitude — indicating that railroad access raises occupational churn asymmetrically, predominantly upward. First-stage F-statistics exceed the Staiger-Stock threshold comfortably (135–414 across specifications). OLS estimates are uniformly smaller than IV estimates, consistent with historical evidence that the railroad targeted areas with weaker growth trajectories.&lt;/p&gt;
&lt;p&gt;The occupational transitions underlying these results run strongly out of farming and into professional, clerical, sales, and services categories, regardless of the father&amp;rsquo;s own occupation (Table IV). Sons growing up closer to the railroad were 19 percentage points less likely to work in a declining occupation and 16 percentage points more likely to work in a growing occupation. The distributional pattern shows an inverted-U relationship with father&amp;rsquo;s occupational decile for occupation-category switching and rank divergence, with the greatest gains concentrated among sons of middle-ranking fathers. For upward mobility specifically, the benefits diminish monotonically as father&amp;rsquo;s rank rises — sons from blue-collar backgrounds gained more (upward mobility coefficient 0.064) than sons from white-collar backgrounds (0.031).&lt;/p&gt;
&lt;p&gt;The authors decompose the total railroad effect on intergenerational mobility into three channels using a structural decomposition applied to a sample of 342,715 brothers: (1) changes in local labor-market opportunities, estimated as the effect on mobility for stayers; (2) changes in the returns to spatial mobility, estimated via a within-family comparison of brothers who moved versus stayed; and (3) changes in the rate of spatial mobility itself. Better railroad access raised the probability of moving away from the birth county by 15 percentage points. However, the estimated return to spatial mobility — the extra boost from actually moving — was reduced by railroad access (negative interaction between proximity and mover status), meaning the railroad decreased the relative advantage of leaving. The decomposition (Table C.6) shows that changes in local opportunities account for the great majority of the total mobility effect. Parish-level evidence confirms the local opportunity mechanism: better-connected parishes saw population growth, more industrial chimneys, more entrepreneurs, higher shares of skilled and literate workers, higher Gini coefficients, and higher median occupational ranks — consistent with agglomeration, industrialization, and skill-biased structural change.&lt;/p&gt;
&lt;p&gt;The policy implication is that transport infrastructure investment can reduce intergenerational persistence in occupational status, primarily by restructuring the local labor market rather than by enabling workers to exit. The caveat is that these gains were unevenly distributed — middle- and lower-ranking families benefited most, and the railroad simultaneously raised local inequality alongside local mobility.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-and-what-are-the-main-threats-it-addresses"&gt;Q1. What is the core identification strategy and what are the main threats it addresses?&lt;/h3&gt;
&lt;p&gt;The authors use a &amp;lsquo;dynamic least-cost path&amp;rsquo; (DLCP) instrument. They connect 53 major English and Welsh towns (defined as the top 10% of the 1801 population distribution, with at least 9,172 inhabitants) via least-cost routes computed over a 50×50 meter terrain grid that assigns slope-based costs to each cell. The instrument is proximity from the childhood residence to the nearest line in this DLCP network. The logic is that individuals incidentally located near the geographic route between major historical towns are more likely to be near an actual railroad — but the DLCP route is based purely on terrain costs, not on local demand, local resources, or the political lobbying that shaped where stations were actually placed. The strategy addresses: (a) reverse causality from high-growth areas attracting railroad placement; (b) sorting of ambitious or wealthy households toward connected parishes; (c) railroad companies&amp;rsquo; demand-driven routing choices. The exclusion restriction could be violated if location along least-cost paths between 1801 major towns is directly correlated with intergenerational mobility for reasons other than the railroad. The paper addresses this by controlling for distance to the nearest 1801 major town and its population (proximity to nodes), proximity to Roman roads, ancient ports, and navigable waterways (pre-existing trade routes), and household wealth proxies.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-instrument-made-dynamic-and-why-does-this-matter"&gt;Q2. How is the instrument made dynamic, and why does this matter?&lt;/h3&gt;
&lt;p&gt;The authors divide the hypothetical network into &amp;rsquo;early&amp;rsquo; (1851) and &amp;rsquo;late&amp;rsquo; (1881) lines by ranking lines in decreasing order of betweenness centrality — the number of times a line connects major towns via shortest paths — until the total cost of the 1851 observed network is exhausted. This dynamic structure means the instrument varies across both space and census cohorts (sons measured in 1851-1881 versus 1881-1911). Without the dynamic feature, the instrument could conflate the effects of lines that were built early (and thus had decades to affect local economies) with lines built later. The temporal variation bolsters the plausibility of the exclusion restriction and is shown to be robust in alternative specifications using static least-cost paths and slope-free least-cost paths.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-four-dependent-variables-and-how-is-intergenerational-mobility-defined"&gt;Q3. What are the four dependent variables and how is intergenerational mobility defined?&lt;/h3&gt;
&lt;p&gt;The paper uses four measures: (1) an indicator equal to one if the son works in a different HISCO occupation category than his father; (2) the absolute value of the difference in HISCAM scores between son and father; (3) &amp;lsquo;upward mobility,&amp;rsquo; an indicator equal to one if the son&amp;rsquo;s HISCAM score exceeds his father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s score distribution; (4) &amp;lsquo;downward mobility,&amp;rsquo; the symmetric indicator for a decline greater than one standard deviation. Sons&amp;rsquo; occupations are observed when sons are 40–52 years old; fathers&amp;rsquo; occupations are measured 30 years earlier when sons were 10–22. The HISCAM scale is held constant over the period (national GB scale, 1800–1938) so that rankings reflect fixed social stratification positions rather than period-specific prestige. The paper also uses time-varying HISCAM, HISCLASS, Woollard, and Armstrong classifications as robustness checks.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-first-stage-performance-of-the-instrument"&gt;Q4. What is the first-stage performance of the instrument?&lt;/h3&gt;
&lt;p&gt;The first-stage relationship between proximity to the nearest DLCP line and proximity to the nearest actual train station is positive and statistically significant across all specifications. The Sanderson-Windmeijer F-statistic is 414 in the specification without controls, 136 with county and year fixed effects and full controls, and remains well above the conventional threshold of 10. The first-stage coefficient drops from 0.640 to 0.339 when full controls are added, indicating that a portion of the geographic correlation between the DLCP and the actual network reflects the pre-existing economic importance of towns and travel routes — which is precisely what the controls absorb.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-main-mechanisms-and-how-are-they-distinguished-empirically"&gt;Q5. What are the main mechanisms and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The paper decomposes the total IV effect on intergenerational mobility using a three-part decomposition: (1) Changes in local opportunities, measured as the effect of proximity on mobility for sons who stayed in their birth county (stayers); (2) Changes in the returns to spatial mobility, estimated by comparing brothers who moved with brothers who stayed (using family fixed effects), and interacting this comparison with railroad proximity; (3) Changes in the rate of spatial mobility itself, estimated from the effect of proximity on the probability of county-to-county migration. Table C.6 shows that local opportunities account for the dominant share of the total effect. The railroad raised the migration probability by 15 percentage points (Table VI), so spatial mobility channels exist — but the railroad decreased the relative advantage of actually moving (negative interaction term in Table V), meaning the local opportunity channel more than offsets the spatial channel. Supporting evidence from parish-level regressions (Table VII) shows that better-connected parishes experienced significantly higher population growth, more industrial chimneys, more entrepreneurs per 100 square meters, higher shares of skilled and literate workers, higher Gini coefficients, and higher median occupational ranks — consistent with agglomeration and skill-biased industrialization.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-is-documented-by-fathers-occupation-and-position-in-the-distribution"&gt;Q6. What heterogeneity is documented by father&amp;rsquo;s occupation and position in the distribution?&lt;/h3&gt;
&lt;p&gt;The effects are heterogeneous by the father&amp;rsquo;s occupational position. Figure 6 shows an inverted-U pattern for occupation-category switching and absolute rank divergence: sons of middle-ranking fathers benefit most from railroad access. For upward mobility (Figure 6c), the benefits diminish monotonically from the lower end of the father&amp;rsquo;s distribution — sons of low-ranking fathers are most likely to move up. Sons of white-collar fathers see smaller (and sometimes statistically insignificant) upward mobility gains (0.031) compared with sons of blue-collar fathers (0.064), while the occupation-category switching benefit is also larger for blue-collar sons (0.108 vs. 0.057) (Table C.1). Separate transition matrices by HISCO category (Table IV) show that railroad access reduces the probability of farming for sons of all father types, and raises probabilities of clerical, sales, and services occupations. Effects on becoming a laborer are heterogeneous: for sons of farmers, proximity raises the probability of becoming a laborer; for sons in service occupations, it decreases it.&lt;/p&gt;
&lt;h3 id="q7-what-robustness-checks-are-run"&gt;Q7. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper performs an extensive battery. (1) Alternative connectivity measures: distance to the nearest railroad line, indicator variables for train station within 5, 10, and 15 km, and parish-level station presence. (2) Alternative mobility thresholds: 0.5, 1.5, and 2 standard deviations for upward and downward mobility; time-varying HISCAM to account for changing occupational prestige. (3) Removing railroad-specific occupations (train conductors, controllers) to check for mechanical effects. (4) Alternative specifications: second-order polynomials, parish fixed effects (10,419 parishes), and fully nonparametric covariate controls via k-means clustering (500 clusters). (5) Alternative instruments: a slope-free DLCP and a static (non-dynamic) least-cost path. (6) Geolocation robustness: using parish centroids instead of street-level addresses. (7) Linking bias: controlling for the individual probability of being linked using cubic polynomials on linkage probability and surname-frequency dummies; also checking that the railroad network explains little of the share of linked individuals at the parish level. (8) Subsamples: by census year (1851-1881 vs. 1881-1911), by county (leave-one-out), by rural/urban status, by father&amp;rsquo;s age, by son&amp;rsquo;s age, by birth order, by native/first-/second-generation immigrant status, by whether the son was born in the same county he grew up in, and by whether the father was in farming. (9) Causal response weighting: the Loken-Mogstad-Wiswall decomposition shows positive IV weights across the entire proximity distribution, consistent with a LATE interpretation. Results are stable across all checks.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-handle-the-selection-into-migration-problem-in-estimating-returns-to-spatial-mobility"&gt;Q8. How does the paper handle the selection-into-migration problem in estimating returns to spatial mobility?&lt;/h3&gt;
&lt;p&gt;The authors follow Abramitzky, Boustan, and Eriksson (2012) and use a within-family comparison of brothers — a subsample of 342,715 sons from 157,369 households who grew up in the same household but one moved county while the other stayed. Family fixed effects absorb the shared household characteristics (wealth, motivation, family networks, financial constraints) that jointly determine the propensity to migrate and the baseline mobility trajectory. The railroad-proximity interaction with mover status is instrumented using the interaction of the DLCP instrument with the mover indicator, via a control function approach. The estimated baseline return to spatial mobility (the mover premium) is positive and significant — movers have higher occupation-category divergence and shift more in both directions — but the railroad-induced change in return to mobility is negative, meaning that proximity to the railroad reduced the additional mobility benefit of actually migrating. This finding is the core of the conclusion that local opportunities, not spatial mobility, dominate.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-paper-document-about-local-labor-market-changes-induced-by-the-railroad"&gt;Q9. What does the paper document about local labor market changes induced by the railroad?&lt;/h3&gt;
&lt;p&gt;Parish-level IV regressions (Table VII) show that better proximity to the 1851 network (instrumented by the DLCP) is associated with: significantly higher population growth between 1851 and 1881; a significantly larger number of industrial chimneys (proxying factory concentration, sourced from Heblich-Trew-Zylberberg (2021)); more entrepreneurs per 100 square meters (from the British Business Census of Entrepreneurs); higher shares of high-skilled and literate workers; a higher Gini coefficient over occupational ranks; and a higher median occupational rank. Additionally, sons in better-connected parishes were 19 percentage points less likely to work in a declining occupation and 16 percentage points more likely to work in a growing occupation (Table C.3). Sons were also 3 percentage points more likely to be literate and 7 percentage points more likely to work in a non-manual occupation (Table C.5). These findings collectively point to agglomeration, industrialization, skill-biased technological change, and the creation of a new entrepreneur class as the mechanisms by which the railroad transformed local labor market structure.&lt;/p&gt;
&lt;h3 id="q10-what-prior-work-does-this-paper-relate-to-most-closely-and-what-distinguishes-it"&gt;Q10. What prior work does this paper relate to most closely, and what distinguishes it?&lt;/h3&gt;
&lt;p&gt;The paper sits at the intersection of the railroad-infrastructure and intergenerational-mobility literatures. In the infrastructure tradition, it relates closely to Donaldson (2018, AER) on railroads in India, Donaldson and Hornbeck (2016, QJE) on US market access, Bogart et al. (2022, JUE) on population and structural change in England and Wales, and Heblich-Redding-Sturm (2020, QJE) on London commuting and urban growth. The closest prior paper is Perez (2017) on nineteenth-century Argentina, who finds railroad access shifted children from agricultural into white-collar and skilled blue-collar occupations; this paper provides similar evidence for England and Wales at individual level and adds a full mechanism decomposition. In the intergenerational mobility tradition it relates to Long and Ferrie (2013, AER) and Long (2013, ERH) on census-based occupational mobility in Victorian Britain. The key methodological advantages of the current paper are: (a) use of the full (not 2%) census for all three years, yielding close to 1 million father-son pairs with match rates of 43–50% versus 15–33% in prior work; (b) street-level geolocation enabling individual-level rather than district-level measurement of railroad access; (c) the explicit three-way mechanism decomposition separating local opportunities, returns to migration, and migration rates; and (d) documenting rich heterogeneity by father&amp;rsquo;s occupational rank and occupation category.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-policy-implications-and-what-scope-conditions-limit-their-external-validity"&gt;Q11. What are the policy implications and what scope conditions limit their external validity?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s core policy message is that transport infrastructure investment can be an effective mechanism for reducing intergenerational occupational persistence — primarily by creating new local labor market opportunities rather than by enabling low-income workers to reach distant job centers. This provides historical support for place-based policies of the sort embodied in the Biden &amp;lsquo;Build Back Better&amp;rsquo; infrastructure proposals or the UK HS2 high-speed railway project (mentioned in the paper). The main scope conditions limiting generalizability are: (1) The setting is nineteenth-century England and Wales during the Industrial Revolution, when the occupational structure was shifting rapidly from farming to industry and commerce — the railroads arrived at a moment of latent demand for new labor market structures; (2) The benefits were not evenly distributed: middle-ranking families (by father&amp;rsquo;s occupational rank) gained most in absolute occupational switching and rank divergence, while the lowest-ranked families gained most specifically in upward mobility; (3) The railroad simultaneously raised local inequality alongside local mobility, suggesting infrastructure investment can be inequality-increasing in the cross-sectional distribution of wages even as it reduces intergenerational persistence; (4) The effects are highly localized — even 5 km of additional distance matters — implying that the placement of stations relative to where low-income families actually live is crucial for achieving distributional goals.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-paper-document-about-the-baseline-patterns-of-intergenerational-mobility-in-the-sample"&gt;Q12. What does the paper document about the baseline patterns of intergenerational mobility in the sample?&lt;/h3&gt;
&lt;p&gt;In the full sample of 980,848 father-son pairs covering 1851-1881 and 1881-1911, 80% of sons do not remain in the same HISCO occupation category as their father. The correlation between father&amp;rsquo;s and son&amp;rsquo;s HISCAM ranks is 0.28. Among sons, 18% experienced upward mobility (son&amp;rsquo;s HISCAM rank more than one SD higher than father&amp;rsquo;s) and 15% experienced downward mobility (more than one SD lower). About 31% of sons moved to a different county from where they grew up, settling on average 100 km away. Sons grew up on average 3.28 km from the nearest train station (SD 5.45 km). These descriptives reveal strong spatial clustering in intergenerational mobility patterns at the parish level.&lt;/p&gt;
&lt;h3 id="q13-does-the-late-interpretation-hold-and-what-does-the-weighting-function-show"&gt;Q13. Does the LATE interpretation hold and what does the weighting function show?&lt;/h3&gt;
&lt;p&gt;The authors verify the LATE interpretation via two approaches. First, following Loken-Mogstad-Wiswall (2012), they compute the causal response weighting function as the covariance between each discrete proximity indicator and the DLCP instrument, divided by the covariance between the proximity measure and the DLCP instrument. They find positive weights across the entire distribution of proximity to the nearest train station, concentrated most heavily for individuals residing 0.5 to 1.5 proximity units (approximately 2.7 to 8.1 km) from a train station — these are the individuals whose proximity is most affected by incidental location along the DLCP. The absence of negative weights indicates the IV estimate does not mix complier and never/always-taker effects in a sign-reversing way. Second, following Blandhol et al. (2022), a fully nonparametric specification using 500 k-means clusters for covariates yields estimates very close to the parametric baseline, consistent with a LATE interpretation of the linear IV estimator.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Dynamic Least-Cost Path (DLCP) Network&lt;/strong&gt;: The paper&amp;rsquo;s instrument for railroad access. A hypothetical railroad network connecting England and Wales&amp;rsquo;s 53 largest towns in 1801 via routes that minimize geographic cost (distance plus slope-based terrain costs), ignoring all demand-side factors. Lines are classified as &amp;rsquo;early&amp;rsquo; (1851) or &amp;rsquo;late&amp;rsquo; (1881) by betweenness centrality until the cost budget of the actual 1851 network is exhausted. Proximity from childhood residence to the nearest DLCP line instruments proximity to the nearest actual train station.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intergenerational Occupational Mobility&lt;/strong&gt;: In this paper, the degree to which a son&amp;rsquo;s adult occupation differs from his father&amp;rsquo;s, measured both categorically (same versus different HISCO category) and cardinally (difference in HISCAM scores). Upward (downward) mobility is specifically defined as the son&amp;rsquo;s HISCAM score exceeding (falling below) the father&amp;rsquo;s by more than one standard deviation of the son&amp;rsquo;s HISCAM distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HISCAM Score&lt;/strong&gt;: A continuous occupational ranking (range 28–99, mean 50, SD 10) derived from the frequency of social interactions — marriages, friendships, parent-child links — between occupations in historical data. Higher scores indicate a more advantageous position in the social stratification structure. The paper uses the national Great Britain scale, held constant for 1800–1938, to make rankings comparable across census years.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Opportunities Channel&lt;/strong&gt;: The mechanism by which railroad access improved intergenerational mobility through restructuring the local labor market — enabling commuting, attracting factories and entrepreneurs, spurring urbanization and industrialization, and creating new occupations requiring new skills — without requiring sons to migrate away from their birth county. Identified empirically as the effect of railroad proximity on mobility outcomes for sons who stayed in their birth county (stayers).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Returns to Spatial Mobility&lt;/strong&gt;: The additional intergenerational mobility benefit (or penalty) associated with actually migrating to another county, estimated using within-family variation among brothers — one who moved and one who stayed — to net out shared household-level determinants of mobility. The paper finds that railroad access reduced (made more negative) the returns to spatial mobility, meaning that the relative advantage of leaving shrank as local opportunities expanded.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inconsequential Place IV Approach&lt;/strong&gt;: An identification strategy (following Chandra-Thompson 2000 and Michaels 2008) in which the instrument for infrastructure access is constructed from the geographic convenience of locations lying between endpoints of a planned network, rather than from demand-side factors at those locations. The DLCP instrument in this paper is a specific implementation: individuals living between 1801 major towns incidentally receive railroad access because the low-cost route between towns passes near their residence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Occupational Tie (Father-Son)&lt;/strong&gt;: The tendency for sons to remain in the same occupation category or same position in the occupational ranking as their father. In this paper, severing the occupational tie means a son moves to a different HISCO category and/or achieves a HISCAM score meaningfully different from his father&amp;rsquo;s. The railroad&amp;rsquo;s main effect is framed as reducing this tie, with upward mobility being the dominant direction of change.&lt;/p&gt;</description></item><item><title>Unconventional Monetary Policies and Inequality</title><link>https://macropaperwarehouse.com/papers/unconventional-monetary-policies-and-inequality/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/unconventional-monetary-policies-and-inequality/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether the Federal Reserve&amp;rsquo;s unconventional monetary policies (UMP) — specifically quantitative easing (QE) and forward guidance — exacerbated income and welfare inequality in the United States during the effective lower bound (ELB) episode following the Great Recession (2009–2015). The question is empirically and theoretically contested: QE raises profits and equity prices, benefiting wealthy households who hold most equity, while simultaneously reducing unemployment, which benefits poorer households who rely almost entirely on labor income. Resolving the net effect requires a unified framework that captures both channels simultaneously, with empirically realistic responses of profits, wages, and unemployment to monetary policy.&lt;/p&gt;
&lt;p&gt;The paper builds a medium-scale Heterogeneous Agent New Keynesian (HANK) model that incorporates: (i) a two-asset structure (liquid deposits and illiquid equity) with portfolio adjustment costs; (ii) three working statuses — employed, unemployed, and business owner — with endogenous job-finding rates determined by a search-and-matching labor market; (iii) a banking sector modeled after Gertler and Karadi (2011), with a moral-hazard leverage constraint; (iv) a substantial fixed cost in production that, combined with wage rigidity, generates procyclical profit responses to monetary policy shocks — a feature absent from standard New Keynesian models and critical for capturing benefits to wealthy households; and (v) an occasionally binding ELB constraint with QE modeled as central bank asset purchases and forward guidance modeled as exogenous expected ELB durations following Jones (2017). The model is calibrated to match the 2007 Survey of Consumer Finances (SCF), targeting the top decile&amp;rsquo;s share of wealth (~70%), income composition across wealth groups, and standard labor market and financial sector moments. Remaining parameters are estimated using Bayesian methods on U.S. quarterly data from 1992 Q1 to 2018 Q4, using ten observables (output, consumption, investment, inflation, nominal interest rate, real wage, unemployment, lump-sum transfers, profits, and Federal Reserve assets), with the ELB regime handled via an inversion filter and the Kulish-Jones method for exogenous ELB durations.&lt;/p&gt;
&lt;p&gt;At the posterior mode, the model attributes the Great Recession primarily to a series of large negative risk premium shocks around 2008–2009, causing investment to fall by more than 20% relative to the pre-crisis level. The central counterfactual compares the actual ELB episode (with UMP) against a scenario where the central bank held its balance sheet constant and allowed ELB durations to be determined endogenously by fundamentals. Between 2009 and 2015, UMP on average produced: a 3.3% increase in profits, a 0.9% increase in equity prices, a 1.5 percentage-point reduction in the unemployment rate, and only a 0.1% increase in real wages (reflecting high estimated wage rigidity). Output and investment were higher by approximately 1% and 3% respectively on average, with profits rising as much as 8% during the ELB episode.&lt;/p&gt;
&lt;p&gt;These aggregate effects translated into non-linear distributional outcomes. For the Gini index, lower unemployment reduced the income Gini by up to 0.6 percentage points, but this was offset by about 80% by the increase in profits and equity prices — leaving only a marginal net Gini reduction of 0.04 percentage points on average. When computed for the bottom 90% alone, the Gini reduction was more pronounced because that group relies overwhelmingly on labor income. However, the income share of the top 10% rose by an average of 0.17 percentage points, driven mainly by higher profits and equity prices. Thus the answer to whether UMP raised inequality is measure-dependent: UMP reduced within-bottom-90% inequality while widening the top-decile income gap.&lt;/p&gt;
&lt;p&gt;Welfare gains (consumption equivalents over the ELB episode) were U-shaped across the wealth distribution: the average gain was 0.27% of lifetime consumption, but households at both extremes gained more than the middle. The bottom 10% benefited from higher job-finding rates (gaining ~0.3%), the top 10% from profits and equity prices (also ~0.3%), and the top 1% gained ~0.33%. The middle 60% gained only ~0.26%. By working status, business owners gained the most (0.82%), followed by the unemployed (0.35%) and the employed (0.27%).&lt;/p&gt;
&lt;p&gt;Decomposing UMP into QE and forward guidance, the paper finds that forward guidance accounted for approximately 55% of total UMP stimulus. Forward guidance amplified both the aggregate and distributional effects of asset purchases: QE alone raised the top 10% income share by about 0.1 percentage point, and forward guidance added a further 0.09 percentage point increase. Forward guidance lowered the overall Gini by about 0.05 percentage points more than QE alone around 2013, and reduced the bottom-90% Gini by an additional 0.2 percentage points during the same period. The interaction intensified what the paper calls a &amp;ldquo;hollowing out&amp;rdquo; of the middle class: forward guidance further reduced middle-60% income shares while leaving bottom-10% shares nearly unchanged, because the additional stimulus disproportionately raised profits and equity prices (by about 2% and 1%, respectively, between 2011 and 2014).&lt;/p&gt;
&lt;p&gt;Comparing QE with a hypothetical conventional monetary policy (CMP) that would have allowed the nominal rate to drop to approximately -1%, the paper finds that CMP would have produced larger aggregate stimulus than QE but more adverse distributional effects. Under CMP, lower financing costs disproportionately boosted bank net worth, indirectly raising profits and benefiting wealthy households even more than QE did. Under QE, central bank asset purchases crowded out private bank investment by reducing expected equity returns even as they raised equity prices, partially dampening the profitability gains to the financial sector. Consequently, CMP would have delivered above-average welfare gains only to the bottom 1% (debtors benefiting from lower real rates) and the top 10% (through larger bank profit effects), while the broad middle class would have fared no better and in some dimensions worse.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s key methodological contribution is the first Bayesian estimation of a HANK model with an occasionally binding ELB constraint. Its key substantive finding is that standard NK models, which generate countercyclical profits, systematically understate the benefits that expansionary monetary policy delivers to wealthy households, producing a misleading or incomplete picture of the distributional effects of monetary policy.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-models-identification-strategy-and-how-is-the-elb-period-handled-in-estimation"&gt;Q1. What is the model&amp;rsquo;s identification strategy and how is the ELB period handled in estimation?&lt;/h3&gt;
&lt;p&gt;The model is estimated with Bayesian methods using an inversion filter (following Guerrieri and Iacoviello 2017 and Cuba-Borda et al. 2019) on ten quarterly observables from 1992 Q1 to 2018 Q4. The key identification challenge is the occasionally binding ELB constraint. The paper follows Kulish et al. (2014) and Jones (2017), treating the ELB as a temporary alternative regime with exogenous expected durations. These expected durations are themselves estimated as latent variables, with priors informed by the New York Fed&amp;rsquo;s primary dealer survey. The Metropolis-Hastings algorithm is used for structural parameters (treating ELB durations as fixed in each draw), while ELB durations are drawn separately using a discrete uniform proposal density. To make estimation computationally feasible given the large idiosyncratic state space, the paper follows Bayer and Luetticke (2020) and updates only the subset of the model Jacobian corresponding to &amp;lsquo;aggregate&amp;rsquo; and &amp;lsquo;summary&amp;rsquo; equations during each iteration, leaving the &amp;lsquo;idiosyncratic&amp;rsquo; blocks fixed across estimated parameters.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-by-which-ump-affects-inequality-and-how-does-the-model-distinguish-them-empirically"&gt;Q2. What are the main mechanisms by which UMP affects inequality and how does the model distinguish them empirically?&lt;/h3&gt;
&lt;p&gt;The paper identifies four main channels: (1) Profit and equity price channel — QE raises equity prices and reduces financing costs, increasing profits and the dividend rate on illiquid assets. Because the top decile holds ~70% of total wealth overwhelmingly in the form of equity, with capital and business income accounting for ~50% of their income, this channel benefits the wealthy disproportionately. (2) Unemployment channel — lower interest rates stimulate demand and raise the job-finding rate. Because households at the bottom of the wealth distribution are more likely to be unemployed at the onset of the ELB episode (8.75% of the bottom decile vs. 6.54% in the middle quintile in 2009 Q1), this channel is progressive. (3) Wage channel — nominal and real wage rigidity (only one-fifth of the real wage adjusts to labor productivity changes) means that the wage channel is very weak; average real wages rose by only 0.1% due to UMP. (4) Inflation/redistribution channel — forward guidance generates inflationary expectations that compress real rates, redistributing from savers to debtors. The empirical decomposition is performed by first isolating QE alone (endogenizing ELB durations) and then comparing to the full UMP scenario (exogenous ELB durations), attributing the residual effect to forward guidance.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-key-modeling-innovation-regarding-profits-and-why-does-it-matter-for-inequality"&gt;Q3. What is the key modeling innovation regarding profits, and why does it matter for inequality?&lt;/h3&gt;
&lt;p&gt;Standard New Keynesian models generate countercyclical profit responses to monetary policy shocks: when demand rises, price rigidity keeps prices sticky while factor prices (wages) adjust upward, squeezing markups and reducing profits. This contradicts empirical evidence from structural VARs, which show procyclical profits. The paper introduces three interacting features that resolve this: (a) a substantial fixed cost of production calibrated to roughly 20% of steady-state output, so that average production cost falls even as marginal cost rises, boosting net profits; (b) wage rigidity with search-and-matching frictions, so that real wages respond very weakly to monetary shocks; and (c) a banking sector with a financial accelerator, so that rising equity prices boost banks&amp;rsquo; net worth and their investment demand, further amplifying profits. Without procyclical profits, the model would understate the benefits wealthy households (whose income depends heavily on profits and equity returns) gain from expansionary monetary policy, producing an incomplete picture of distributional effects.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-in-households-balance-sheets-and-income-composition-is-documented-and-how-does-it-shape-distributional-results"&gt;Q4. What heterogeneity in households&amp;rsquo; balance sheets and income composition is documented, and how does it shape distributional results?&lt;/h3&gt;
&lt;p&gt;Using the 2007 SCF, the paper documents stark composition differences. The bottom 80% of the wealth distribution derives ~80% of income from labor, with transfer income making up most of the rest. The top 10% derives about 50% from labor and 50% from capital (equity and business income). For the top 0.1%, labor income is only 16% and capital/business income is about 83–85%. In the model, the top 10% hold about 70% of total wealth, overwhelmingly in illiquid equity. These composition differences mean that any policy raising profits and equity prices is strongly progressive at the top and neutral-to-mild at the bottom, while any policy reducing unemployment is strongly progressive at the bottom. The interplay of these two forces explains why UMP simultaneously reduces bottom-90% inequality (through the unemployment channel) and widens the top-vs.-rest gap (through the profit and equity channel), and why welfare gains are U-shaped rather than monotone.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-welfare-accounting-methodology-and-what-are-the-key-welfare-findings"&gt;Q5. What is the welfare accounting methodology and what are the key welfare findings?&lt;/h3&gt;
&lt;p&gt;Welfare gains are measured as consumption equivalents — the fraction of lifetime consumption that a household in the counterfactual (no UMP) scenario would be willing to forgo to enjoy the UMP outcome. Households are sorted into wealth groups based on their 2009 Q1 wealth position (so group composition is not affected by UMP), and the same households are followed throughout the episode. Beyond the sample end (2018 Q4), no further shocks are assumed. The average welfare gain at the posterior mode is 0.27% of lifetime consumption. Bottom 10%: ~0.3% (driven by higher job-finding rates). Top 10%: ~0.3% (driven by profits and equity gains). Top 1%: ~0.33%. Middle 60%: ~0.26%. Business owners: 0.82%. The unemployed: 0.35%. The employed: 0.27%. Critically, the welfare gaps between extremes and middle are smaller than the income gaps, because anticipated tapering after the sample implies lower future profits and equity prices for wealthy households, narrowing their long-term advantage.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-contributions-of-qe-and-forward-guidance-compare-in-aggregate-and-distributional-terms"&gt;Q6. How do the contributions of QE and forward guidance compare in aggregate and distributional terms?&lt;/h3&gt;
&lt;p&gt;Forward guidance accounted for approximately 55% of the total UMP stimulus at the posterior mode. Exogenous expected ELB durations exceeded endogenous (fundamentals-based) durations by 1–2 quarters on average, and sometimes by up to 8 quarters, with the divergence widening from 2011 onward. In distributional terms, QE alone initially reduced the bottom-90% Gini and raised the top 10% income share by about 0.1 percentage point. Forward guidance amplified both effects: it lowered the overall Gini by an additional ~0.05 pp and the bottom-90% Gini by an additional 0.2 pp around 2013, but also added a further ~0.09 pp to the top 10% income share between 2011 and 2014. The amplification occurred because forward guidance raised profits and equity prices by about 2% and 1% respectively during that window, intensifying the income concentration at the top while also stimulating job creation at the bottom. The middle class saw its income share further compressed.&lt;/p&gt;
&lt;h3 id="q7-how-does-qe-compare-with-conventional-monetary-policy-in-terms-of-aggregate-and-distributional-effects"&gt;Q7. How does QE compare with conventional monetary policy in terms of aggregate and distributional effects?&lt;/h3&gt;
&lt;p&gt;In the counterfactual CMP scenario, the nominal policy rate drops to approximately -1% and remains negative for an extended period. CMP produces larger aggregate stimulus than QE: the stimulus effects of QE were partly crowded out by general equilibrium effects, specifically QE reduced banks&amp;rsquo; expected return on equity even as it raised equity prices, discouraging private bank investment. Under CMP, lower nominal rates instead benefit banks through lower financing costs, boosting bank net worth via an accelerator mechanism more strongly than under QE. This difference has distributional consequences: CMP would have delivered higher welfare gains only to the bottom 1% (low-wealth debtors benefiting from lower real rates on their liabilities) and the top 10% (benefiting from larger bank profits). Households in the broad middle — already employed, holding limited equity, neither heavy borrowers nor large business income recipients — would have been no better off and in some dimensions worse off under CMP. The paper thus concludes that QE had less adverse distributional effects than CMP would have had, absent the ELB constraint.&lt;/p&gt;
&lt;h3 id="q8-what-robustness-checks-and-sensitivity-analyses-are-conducted"&gt;Q8. What robustness checks and sensitivity analyses are conducted?&lt;/h3&gt;
&lt;p&gt;The paper checks results against: (a) the full 10th–90th percentile range of the posterior distribution for all key findings on aggregate effects, income inequality, welfare gains, and QE vs. CMP comparisons, showing that qualitative findings are robust to parameter uncertainty; (b) a comparison between rigid-wage and flexible-wage model variants (Table A1), showing that the flexible-wage version generates countercyclical profits, a weak unemployment response, and a strong real wage response — inconsistent with empirical SVAR evidence — validating the modeling choice of high wage rigidity; (c) a structural VAR analysis on U.S. data confirming procyclical profits, weak real wage responses, and significant unemployment responses to monetary policy shocks; (d) a comparison of the OccBin method (endogenous ELB durations, Guerrieri and Iacoviello 2015) vs. the Kulish-Jones method (exogenous durations) for solving the occasionally binding constraint; (e) a check that wages implied by the calibrated wage function always remain in the bargaining set, validating the equilibrium wage assumption.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-key-differences-between-this-paper-and-the-closest-prior-work"&gt;Q9. What are the key differences between this paper and the closest prior work?&lt;/h3&gt;
&lt;p&gt;Kaplan, Moll, and Violante (2018) and Bayer et al. (2020) have two-asset HANK models but omit frictional labor markets, so they cannot capture how monetary policy affects employment and thus the progressive unemployment channel. Gornemann et al. (2016) include search-and-matching labor markets but only one asset, so they cannot capture the capital income benefits to wealthy households. Broer et al. (2019) and Auclert et al. (2023) identify the countercyclical profit problem but their solutions (wage rigidity alone) produce procyclical profits that are too weak quantitatively. This paper combines fixed costs, wage rigidity, and a banking sector to produce procyclical profits quantitatively consistent with SVAR evidence. On unconventional policy specifically, Lenza and Slacalek (2018) and Casiraghi et al. (2018) study ECB QE with partial equilibrium methods and find inequality-reducing effects; Bivens (2015) and Montecino and Epstein (2015) reach opposite conclusions for U.S. QE. This paper is the first to study both QE and forward guidance jointly in a Bayesian-estimated HANK model with an explicitly binding ELB, and is to the author&amp;rsquo;s knowledge the first to estimate a HANK model with an occasionally binding ELB constraint.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-main-policy-implications-and-their-scope-conditions"&gt;Q10. What are the main policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;First, UMP&amp;rsquo;s inequality effects are measure-dependent: policies that simultaneously stimulate employment and profits can reduce within-bottom-90% inequality while widening the top-vs.-rest gap. Policymakers who cite Gini reductions and those who cite rising top-income shares are both correct, pointing to different parts of the distribution. Second, forward guidance amplifies inequality effects as much as it amplifies aggregate effects, so its use carries a distributional cost concentrated at the top of the distribution. Third, QE had less adverse distributional effects than conventional monetary policy would have had, suggesting that concerns about QE&amp;rsquo;s inequality effects should be placed in context of the ELB constraint — the relevant comparison is not QE vs. no policy but QE vs. CMP with the ELB absent. Fourth, models that generate countercyclical profits will systematically understate benefits to the wealthy and potentially reach qualitatively different conclusions about whether monetary policy raises or reduces inequality. These findings are scoped to the U.S. Great Recession ELB episode, estimated with the specific HANK model structure and Bayesian posterior; findings may differ for different financial structures, more generous unemployment insurance, or different asset price dynamics.&lt;/p&gt;
&lt;h3 id="q11-what-drives-the-great-recession-in-the-model-and-how-is-ump-modeled-mechanically"&gt;Q11. What drives the Great Recession in the model and how is UMP modeled mechanically?&lt;/h3&gt;
&lt;p&gt;At the posterior mode, the Great Recession is primarily attributed to a series of large negative risk premium shocks (shocks to banks&amp;rsquo; discount factor) around 2008–2009, which caused banks to sharply contract their investment, leading to the investment collapse (&amp;gt;20% below pre-crisis). QE is modeled following Gertler and Karadi (2011): the central bank issues bonds (sold to the private sector) and uses proceeds to purchase equity directly, converting non-productive asset demand into productive capital demand and raising equity prices and investment. Forward guidance is modeled as setting exogenous expected ELB durations longer than would be implied endogenously by the Taylor rule fundamentals, effectively mimicking future negative interest rate shocks and inducing inflationary pressure via intertemporal substitution. The expected ELB durations at the posterior mode range from 6 to 8 quarters through 2013, falling sharply to 1–2 quarters by late 2014–2015.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneous Agent New Keynesian (HANK) model&lt;/strong&gt;: As used in this paper, a DSGE model where households differ ex-post in idiosyncratic productivity, asset holdings (liquid deposits and illiquid equity), and employment status; combined with search-and-matching labor markets, a banking sector with leverage constraints, and a zero lower bound on the policy rate. The heterogeneity in wealth composition and income sources determines how aggregate policy shocks translate into distributional outcomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Procyclical profits&lt;/strong&gt;: The property, established empirically via SVAR and reproduced in the model, that firm profits rise in response to expansionary monetary policy shocks. Standard New Keynesian models generate the opposite (countercyclical profits) because price rigidity compresses markups when demand rises. In this paper, the combination of large fixed costs in production, wage rigidity, and a banking sector financial accelerator is required to generate quantitatively realistic procyclical profit responses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Effective lower bound (ELB) episode&lt;/strong&gt;: The period from 2009 Q1 to 2015 Q4 during which the Federal Reserve&amp;rsquo;s policy rate was constrained at zero. In the model, this is treated as a temporary alternative regime with exogenous expected durations; when the policy rate hits the ELB, the central bank can only affect the economy through asset purchases (QE) and forward guidance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Forward guidance (as exogenous expected ELB durations)&lt;/strong&gt;: In this paper&amp;rsquo;s framework, forward guidance is operationalized as the central bank committing to maintain the policy rate at zero for a longer period than the endogenous (fundamentals-based) Taylor rule would prescribe. This is parameterized as an exogenous expected ELB duration that exceeds the endogenous one, creating anticipations of future negative interest rate shocks and thus stimulating activity through intertemporal substitution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption equivalent welfare gain&lt;/strong&gt;: The fraction of lifetime consumption that a household in the counterfactual scenario (no UMP) would be willing to forgo in order to instead experience the outcomes under UMP. Used to compare welfare across heterogeneous households in a cardinal, utility-based metric rather than income alone.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Business owner working status&lt;/strong&gt;: A third working status (alongside employed and unemployed), following Bayer et al. (2019), in which households receive a fixed fraction of aggregate profits as income without supplying labor. Business owners transition into and out of this status exogenously and are the highest-income group in the model, calibrated to match the top-decile&amp;rsquo;s share of liquid assets and the income composition data showing that capital and business income dominate the very top of the wealth distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inversion filter&lt;/strong&gt;: The likelihood evaluation method used in this paper for Bayesian estimation, following Guerrieri and Iacoviello (2017). Rather than running a Kalman filter, structural shocks are backed out directly by inverting the linear solution of the model given the observed data and a given set of expected ELB durations. This avoids continuously updating the large state-transition matrix and makes estimation computationally feasible.&lt;/p&gt;</description></item><item><title>Understanding High-Wage Firms: Monopoly, Monopsony, and Bargaining Power</title><link>https://macropaperwarehouse.com/papers/understanding-high-wage-firms-monopoly-monopsony-and-bargaining-power/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/understanding-high-wage-firms-monopoly-monopsony-and-bargaining-power/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Why do some firms pay persistently higher wages for observably similar workers, and what role do firms&amp;rsquo; product-market power (monopoly/markups), labor-market power (monopsony/markdowns), and workers&amp;rsquo; collective bargaining power play in shaping wages and welfare? Prior literature studies labor-market power as a driver of wages/profits but abstracts from product-market power and bargaining, while the markups literature abstracts from imperfect labor competition and bargaining. The paper unifies all three in one structural framework.&lt;/p&gt;
&lt;p&gt;Central theoretical insight: A firm&amp;rsquo;s wage equals its marginal revenue product of labor (MRPL) times a &amp;ldquo;labor wedge&amp;rdquo; (the share of MRPL workers receive). The labor wedge decomposes into three components — price-cost markups, monopsony markdowns, and bargaining power — via equation (3): Lambda = kappa*(product market rents term) + (1-kappa)*lambda. With positive bargaining power (kappa&amp;gt;0) workers capture a share of markup-generated rents, so the labor wedge rises with markups (rent-sharing); this nests pure monopsony as the kappa=0 special case.&lt;/p&gt;
&lt;p&gt;Data and setting: French administrative micro-data. Firm balance sheets (FARE, 2008-2019, DGFiP); firm-product output prices (EAP survey, 2009-2019, INSEE, manufacturing firms &amp;gt;=20 employees or sales &amp;gt;5m euros); matched employer-employee data (DADS, 1995-2018) which crucially includes hours worked. Firm wage premia estimated via a k-means/BLM grouped AKM regression (Bonhomme, Lamadon, Manresa 2019). Markups and labor wedges estimated with the production-function/production approach (De Loecker-Warzynski 2012; Yeh et al. 2022) using translog functions and an Ackerberg-Frazer-Caves control function, separating the two by noting markups distort all input demands while labor wedges distort only labor demand.&lt;/p&gt;
&lt;p&gt;Two key empirical facts a standard monopsony model cannot explain: (i) high-wage firms charge higher output prices and markups; (ii) high-wage firms pay a larger share of MRPL as wages (higher labor wedges). Both persist within narrow industries and conditional on TFP, pointing to product quality and positive bargaining power.&lt;/p&gt;
&lt;p&gt;Main quantitative findings (French manufacturing, 2016 unless noted): Median markup 1.32 (IQR 1.14-1.60). Median labor wedge 0.62 (median monopsony markdown 0.46) — the gap is due to bargaining power and markups. Workers capture about 12% of firm profits (bargaining power kappa ~ 0.12-0.14; falls to ~0.05-0.13 under IV correction). Median markdown 0.46 implies a median firm-specific labor supply elasticity of 0.85. Accounting for hours matters: median labor wedge is 0.62 with effective hours, 0.65/0.68/0.71 across specifications, rising to 0.71 when labor is measured by employment (near Yeh et al.&amp;rsquo;s 0.70-0.73 US figures) — so omitting hours upward-biases labor wedges.&lt;/p&gt;
&lt;p&gt;Quantitative GE model (oligopoly/oligopsony, nested-CES, Atkeson-Burstein/Berger et al.): A 1% productivity shock has wage passthrough 0.97-0.99 versus 0.23 for an equal quality shock (because varieties are close substitutes, sigma=5.17), though quality still generates more wage-premium dispersion. Markups and markdowns reduce welfare by 46% in consumption-equivalent terms, with markups alone accounting for over 80%; misallocation explains about 63% of the markup welfare cost. Equalizing markups raises average wages 39% and wage variance 99% and welfare 24% (output-restriction effect dominates rent-sharing, so equalizing markups raises wage dispersion). Raising bargaining power from 0.12 to 0.50 matches the wage gains of removing markups but yields only 10% welfare gain (vs 38%); full bargaining power (kappa=1) raises welfare 13%, under one-third of the planner&amp;rsquo;s 46% gain. Bargaining power offsets the uniform-tax and misallocation distortions on labor demand but cannot fix markup distortions to capital/material demand.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-identification-strategy-for-separating-markups-from-labor-wedges-and-what-are-its-main-assumptionsthreats"&gt;Q1. What is the core identification strategy for separating markups from labor wedges, and what are its main assumptions/threats?&lt;/h3&gt;
&lt;p&gt;The author applies the production approach: estimate translog production functions per 2-digit manufacturing sector (via two-step GMM with an Ackerberg-Frazer-Caves control function for unobserved productivity) to recover firm-specific output elasticities. Markups distort the demand for ALL inputs while labor wedges distort ONLY labor demand, so choosing materials as a flexible, price-taken input lets markups be identified from the material cost share (mu = alpha_m * PY/(Pm*M)) and labor wedges from the wage-bill-to-materials ratio scaled by elasticity ratios (eq. 4). Key assumptions/threats: materials must be a flexible input firms take prices for (examined in Appendix B.7-B.8); unobserved productivity must satisfy scalar unobservability and monotonicity in material demand; unobserved output and input prices bias elasticities — addressed using observed EAP output prices (measuring output in quantities) plus the De Loecker et al. (2016) input-price control function, and additionally controlling for firm wage premia because monopsony markdowns create unobserved labor-price variation. Markup variation driven by idiosyncratic demand uncorrelated with TFP is controlled via export status, market shares, firm age, and a 3rd-order price polynomial. Gandhi-Navarro-Rivers concerns about identifying material elasticities are addressed in Appendix B.9.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-new-identification-challenge-for-estimating-bargaining-power-and-how-is-it-solved"&gt;Q2. What is the new identification challenge for estimating bargaining power, and how is it solved?&lt;/h3&gt;
&lt;p&gt;The rent-sharing literature estimates bargaining power kappa by regressing wages on quasi-rents using instruments (export demand, patent shocks) assumed orthogonal to the worker&amp;rsquo;s reservation wage. But in this model, when kappa=0 workers earn an endogenous monopsony wage (lambda*MRPL) that moves with the SAME firm-specific shocks (productivity, quality, amenities) that shift quasi-rents — so standard instruments violate the exclusion restriction. The solution: instead of the wage equation, exploit the labor-wedge equation (3), which relates labor wedges to markups and avoids unobserved monopsony wages. Conditional on markdowns, variation in product-market rents identifies kappa (when kappa=0 product-market rents do not affect the labor wedge). This shifts the core challenge from unobserved monopsony wages to unobserved amenities (mirroring IC3 in the rent-sharing literature), handled by a theory-consistent control function in which employment and the wage bill jointly proxy for amenities under a monotonicity assumption (labor supply increasing in amenities). Under multiplicative separability of wages and amenities, markdowns do not depend directly on amenities, so unobserved amenities do not bias kappa at all.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-bargaining-power-estimates-across-specifications"&gt;Q3. What are the bargaining-power estimates across specifications?&lt;/h3&gt;
&lt;p&gt;Pooled OLS gives ~0.135; adding firm fixed effects ~0.124; adding the amenity control function (columns 3-4) ~0.124-0.135, indicating amenities have little direct effect on markdowns; instrumenting product-market rents with their lags to correct correlated measurement error (columns 5-6) gives 0.130 and 0.059. Baseline kappa is taken as ~0.12 (specification 4). All 2-digit sectors have kappa below 0.3. These align with the rent-sharing literature&amp;rsquo;s typical 0.05-0.15, though external innovation-based instruments tend to find ~0.30.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-paper-measure-firm-wage-premia-and-why-not-use-standard-akm"&gt;Q4. How does the paper measure firm wage premia and why not use standard AKM?&lt;/h3&gt;
&lt;p&gt;Standard AKM firm effects assume time-invariant firm effects and rely on worker mobility; short panels yield noisy estimates with upward-biased variance. The author needs time-varying premia (to measure effective labor over time). He uses the BLM (Bonhomme, Lamadon, Manresa 2019) k-means approach: cluster firms by the similarity of their internal wage distributions (by 2-digit sector over overlapping 2-year windows), then run an AKM-style regression with firm-GROUP effects that vary by year, identified by workers switching between firm-groups — greatly increasing the number of switchers. DADS-Postes is used for clustering (broad coverage) and DADS-Panel for the wage-premium regression.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-across-firms"&gt;Q5. What heterogeneity is documented across firms?&lt;/h3&gt;
&lt;p&gt;Firm wage premia dispersion accounts for 5.2% of wage dispersion; the 90-10 premium gap is ~30% (about 4 euros/hour, 25% of the median worker&amp;rsquo;s hourly wage), IQR 15%. Markdowns increase with firm wage premia (flat gradient) but DECREASE with firm size — larger firms have more monopsony power, consistent with oligopsony models. Firm-specific labor supply elasticities are 0.54/0.85/1.33 at the 25th/50th/75th percentiles. About 7% of firms have labor wedges above 1, and these tend to have much higher markups (rationalized by kappa&amp;gt;0). In the GE model, top-decile high-wage firms are ~15% more productive but have over 100% greater product quality than bottom-decile firms; amenities rise slightly more steeply with premia than productivity. Passthrough is substantially smaller for 90th-percentile firms (0.74 productivity, 0.18 quality) than for median/10th-percentile firms (~1.06/~0.26).&lt;/p&gt;
&lt;h3 id="q6-how-is-the-dispersion-of-wage-premia-decomposed-across-sources-of-firm-heterogeneity"&gt;Q6. How is the dispersion of wage premia decomposed across sources of firm heterogeneity?&lt;/h3&gt;
&lt;p&gt;Introducing one source at a time into the GE model and comparing variance to baseline (Table 6): varying only product quality reproduces 161.5% of baseline variance, only TFP 153.3%, and only amenities 40.8%. Product quality is the largest single contributor to wage-premium dispersion, closely followed by productivity, then amenities.&lt;/p&gt;
&lt;h3 id="q7-why-does-the-productivity-passthrough-differ-so-much-from-the-quality-passthrough"&gt;Q7. Why does the productivity passthrough differ so much from the quality passthrough?&lt;/h3&gt;
&lt;p&gt;Total passthrough is 0.97 for a 1% productivity shock vs 0.23 for an equal quality shock (~4x). The decomposition (Table 5) attributes most of the gap to the direct effect (1.07 vs 0.26): with high within-market substitutability (sigma=5.17), consumers are very price-sensitive, so productivity (which lowers price) moves sales and labor demand far more than quality. Higher sigma raises productivity passthrough but lowers quality passthrough. For sufficiently low sigma the ranking can reverse. The variable-market-power channel also matters: higher productivity raises markups, increasing rent-sharing (+0.06 via labor wedge) but also output restriction (-0.09 via markup), with output restriction dominating; firm-size effects (sectoral price -0.10, sectoral wage +0.03) further adjust passthrough. Amenity shocks have direct effect -0.26 (mirror of quality) but total -0.28, amplified because better amenities lower hiring costs and expand the firm.&lt;/p&gt;
&lt;h3 id="q8-how-does-worker-bargaining-power-affect-welfare-and-what-are-the-limits"&gt;Q8. How does worker bargaining power affect welfare, and what are the limits?&lt;/h3&gt;
&lt;p&gt;Bargaining power offsets two distortions firm market power imposes on aggregate labor demand: a uniform tax (Lambda/mu, lowering labor demand proportionally) and a misallocation tax (Theta, from dispersion in wedges). There exists a kappa-bar that exactly cancels the uniform tax, and kappa-bar falls as markups rise (high markups make bargaining more effective). With full bargaining power and common markups, the markdown-driven misallocation tax is fully neutralized. BUT bargaining only acts through labor demand; markups also distort capital and material demand, which bargaining cannot fix. Quantitatively: raising kappa from 0.12 to 0.50 matches the wage gain of removing markups but yields only 10% welfare gain (vs 38%) and far less dispersion increase; full kappa=1 raises welfare 13%, under one-third of the planner&amp;rsquo;s 46% gain. So bargaining power is a partial, not full, remedy for firm market power.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-welfare-accounting-for-markups-vs-markdowns"&gt;Q9. What is the welfare accounting for markups vs markdowns?&lt;/h3&gt;
&lt;p&gt;Comparing the decentralized economy to the social planner&amp;rsquo;s (Table 7, column 3): eliminating both markups and markdowns raises wage-premium dispersion 113%, average wages 303%, and welfare 46% (consumption-equivalent). Over 80% of the welfare gain comes from removing markups. Equalizing markups alone (column 4) gives 24% welfare, +39% wages, +99% wage variance, implying ~63% of the markup welfare cost is misallocation. Equalizing markdowns alone (column 5) has little welfare effect (2%), though a wide markdown level reduces welfare significantly (column 2).&lt;/p&gt;
&lt;h3 id="q10-what-robustness-checks-and-caveats-does-the-author-flag"&gt;Q10. What robustness checks and caveats does the author flag?&lt;/h3&gt;
&lt;p&gt;Caveats: (1) Multiplication bias — mismeasured output elasticities enter both labor wedges and product-market rents multiplicatively, mechanically biasing kappa upward (Appendix B.10); IV with lags only fixes classical, not serially-correlated, measurement error. (2) Labor adjustment costs get absorbed into the labor wedge and bias kappa; firm fixed effects do not fully fix this (Appendix B.11). (3) The markdown estimation imposes that all markdown variation reflects firm size and amenities — more general than kappa=0 approaches but restrictive in this dimension. (4) The model uses collective (not individual) bargaining and abstracts from sequential-auction wage-setting (Cahuc-Postel-Vinay-Robin); robustness to hiring-wages-only following Di Addario et al. (2020) is shown (Appendix B). (5) Worker types assumed perfect substitutes; an Appendix E two-skill extension gives similar results. (6) Empirical patterns hold without TFPQ controls (Figure D.3) and by firm size (Figure D.4).&lt;/p&gt;
&lt;h3 id="q11-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q11. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;Versus the labor-market-power literature (Berger et al. 2022; Lamadon et al. 2022) it adds product-market power and bargaining, showing their pure-monopsony labor wedge is a kappa=0 special case. Versus the markups/welfare literature (De Loecker et al. 2020; Edmond et al. 2023) it adds imperfect labor competition and bargaining. Versus recent integrated product+labor power models that use wage-posting and no bargaining (Kroft et al. 2024; Deb et al. 2024), it adds the rent-sharing channel where markups raise (not just lower) the labor wedge. Versus production-approach markdown estimation (Yeh et al. 2022; Mertens 2020), it shows their estimates are labor wedges (not markdowns) once kappa&amp;gt;0, and that omitting hours upward-biases them. Versus the rent-sharing literature (Card et al. 2018; Kline et al. 2019; Van Reenen 1996), it shows their instruments violate exclusion under endogenous monopsony wages and proposes the labor-wedge-equation alternative. The closest exception incorporating unions is Azkarate-Askasua and Zerecero (2025).&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Strengthening worker collective bargaining power can raise welfare mainly by offsetting markup-induced distortions to labor demand and redistributing rents, but it raises between-firm wage inequality and cannot restore full efficiency because it leaves markup distortions to capital/material untouched (full kappa closes under one-third of the planner gap). The wage effects of innovation depend on whether it improves productivity or quality and on the degree of product differentiation. Scope conditions: estimates are for French manufacturing under firm-level collective bargaining institutions (firms &amp;gt;=50 employees legally bargain annually); results rely on the production-approach assumptions (flexible/price-taken materials, scalar unobservability) and on data including hours and output prices that many countries lack — researchers should interpret labor-wedge/markup moments cautiously without hours data.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Universal Daycare and Mothers' Working Lifetime</title><link>https://macropaperwarehouse.com/papers/universal-daycare-and-mothers-working-lifetime/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/universal-daycare-and-mothers-working-lifetime/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper estimates the causal effects of universal daycare access on mothers&amp;rsquo; labor force participation, full-time employment, hours worked, and earnings across 34 years after the birth of their first child — the longest window examined in this literature. The motivation is twofold: the existing evidence base is overwhelmingly short-run, and the human capital channel (reduced depreciation of skills, accumulation of experience) implies that early labor market attachment during child-rearing years could compound over decades in ways that short-run estimates miss entirely.&lt;/p&gt;
&lt;p&gt;The identification exploits Denmark&amp;rsquo;s 1964 reform that converted a targeted (means-tested) childcare system into a universal one, which triggered a staggered geographic roll-out of daycare centers from 1966 onward across the country&amp;rsquo;s 2,033 neighborhoods nested in 277 municipalities. The paper combines digitized historical daycare yearbooks (1964–1975), the 1970 census, and administrative registers from Statistics Denmark covering 370,602 mothers who had their first child between 1964 and 1975. Employment is measured via annual contributions to the Supplementary Pension Fund (ATP); earnings from tax records are available from 1980 through 2015, adjusted to 2016 USD. The empirical strategy is a difference-in-differences design comparing mothers in neighborhoods with versus without daycare within the same municipality over time. Daycare availability when the first-born child turns four is used as the fixed treatment indicator for the long-run regressions. Municipality fixed effects absorb cross-sectional confounders; year-of-first-birth dummies capture macro trends.&lt;/p&gt;
&lt;p&gt;The contemporaneous effects are already substantial. Once year and municipality fixed effects and covariates are included, daycare availability raises the probability of participation by 1.5 percentage points when the child is two, rising to 5.3–5.7 percentage points for years three through six — translating to roughly 9 percent more likely to participate relative to the mean. Full-time employment rises by 9–12 percent relative to the mean for years three through six; hours worked increase by 0.27 hours per week (1.8 percent) when the child is four.&lt;/p&gt;
&lt;p&gt;The long-run effects persist throughout the entire working life. Relative to the sample mean, mothers with daycare access are 9.7 percent more likely to participate when the first child turns four, declining to 5.7 percent at child age 14, 3.1 percent at child age 22, and still 1.2 percent at child age 34 (when the average mother is approximately 57.7 years old). Full-time employment effects follow a parallel trajectory: 11 percent higher at child age four, 8.2 percent at child age 14, and 4.4 percent at child age 34. Log earnings (conditional on employment) range between 3 and 6 percent higher throughout the observation window; mothers earn 5.3 percent more when the child is 16 and 4.2 percent more when the child is 34.&lt;/p&gt;
&lt;p&gt;Heterogeneity by education is a central finding. For low-educated mothers (no post-secondary education, 50 percent of the sample), participation effects are 10.1 percent at child age 10, 5.1 percent at child age 17, and remain statistically significant through 32 years. For higher-educated mothers, participation effects are 3.9 percent at child age 10, fall below 1 percent by child age 17, and become statistically indistinguishable from zero by child age 23. Employment effects are thus larger and more persistent for low-educated mothers. Earnings effects, however, are more closely aligned across education groups and show a distinctive pattern for higher-educated mothers: earnings effects persist and remain significant long after employment effects have faded, suggesting that sustained attachment during child-rearing years translates into qualitative career advancement (not just more years worked) for the more educated group.&lt;/p&gt;
&lt;p&gt;Potential mediators include reduced secondary fertility and increased parental separation. Daycare for children aged three to six reduces the total number of children by 0.036 (1.6 percent relative to the mean of 2.2), reduces the probability of having more than two children by 1.8 percentage points (6.0 percent), and increases birth spacing by 0.137 years, making mothers 2.2 percentage points less likely to have a second child within two years. Additionally, mothers with daycare access are 2 percentage points more likely to live apart from the first-born child&amp;rsquo;s father when that child turns 16 — consistent with greater female economic independence. These mediator effects do not vary systematically by education level. Daycare access does not affect additional educational attainment after first birth, ruling out re-skilling as a channel.&lt;/p&gt;
&lt;p&gt;The policy implication is that subsidized universal daycare is not merely a short-run labor supply intervention but a persistent investment in female human capital accumulation, with effects that compound over careers and remain economically meaningful into near-retirement ages.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-key-threats-to-it"&gt;Q1. What is the identification strategy and what are the key threats to it?&lt;/h3&gt;
&lt;p&gt;The paper uses a staggered difference-in-differences design. The key variation is the timing of daycare center openings across neighborhoods within municipalities following the 1964/1966 Danish reform. Daycare availability in the year the first-born child turns four is the fixed treatment indicator for long-run regressions; current-year daycare availability is used for contemporaneous regressions. Municipality fixed effects absorb time-invariant local differences; year-of-first-birth dummies absorb aggregate time trends. The main threat is non-random placement of daycare centers — if centers opened in areas where female labor force participation was already rising, the estimates would be upward biased. The paper addresses this with (1) an event study at the neighborhood level using data from 1960 through 2003 showing no pre-reform differential trends between neighborhoods that later received daycare and those that did not (compared against placebo neighborhoods assigned fictitious opening dates mimicking the actual distribution), and (2) a selective migration check showing that mothers who moved longer distances from their birthplace were no more likely to reside in a neighborhood with daycare once the full conditioning set is included. A residual concern is that for mothers having their first child before 1970, neighborhood assignment is measured post-birth (1970 census), which is addressed by a robustness check excluding the pre-1970 first-birth cohort.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-deal-with-heterogeneous-treatment-effects-and-two-way-fixed-effects-bias"&gt;Q2. How does the paper deal with heterogeneous treatment effects and two-way fixed effects bias?&lt;/h3&gt;
&lt;p&gt;The paper acknowledges the recent literature on TWFE bias under treatment effect heterogeneity (De Chaisemartin and d&amp;rsquo;Haultfoeuille 2020; Callaway and Sant&amp;rsquo;Anna 2021; Sun and Abraham 2021; Borusyak et al. 2024). It replicates the pre-reform event study using the Borusyak et al. (2024) imputation estimator, which is robust to heterogeneous treatment effects and allows for covariates, and finds similar results to the standard TWFE event study (Appendix Figure A.2). The main long-run regressions fix the treatment indicator to daycare availability when the child is four, so there is no variation in treatment timing within a regression, limiting but not eliminating TWFE concerns for the long-run estimates.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-main-mechanism-behind-the-persistent-effects"&gt;Q3. What is the main mechanism behind the persistent effects?&lt;/h3&gt;
&lt;p&gt;The paper attributes the persistence to human capital dynamics: labor force participation during the child-rearing years reduces depreciation of previously accumulated human capital (from education and prior work experience) and enables new on-the-job human capital accumulation through the current job. For low-educated mothers, the primary channel appears to be the extensive margin — daycare moves mothers who would otherwise become homemakers into paid employment, and the employment effects persist because once labor market attachment is established, it is durable. For higher-educated mothers, the earnings-employment gap is the key signal: employment effects fade within roughly 23 years (consistent with convergence once children are no longer preschool age and informal care becomes feasible), yet earnings remain elevated for decades, suggesting that the women who maintained employment during child-rearing years accrued qualitatively better positions — more experience, better job-match, more promotions — compared to those who did not.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-main-mediators-and-how-are-they-distinguished"&gt;Q4. What are the main mediators and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;Three mediators are examined. First, secondary fertility: daycare for children aged 3–6 reduces number of children by 0.036, probability of a third child by 1.8 percentage points, and probability of a fourth child by 0.5 percentage points. The effect operates through daycare for children 3–6 (not 0–2), consistent with the main employment effects operating when the child is three or older. The fertility reduction increases the opportunity cost interpretation — daycare raises the effective wage, making additional children more costly in terms of foregone earnings. Second, birth spacing: mothers with daycare access wait 0.137 more years between first and second child, and are 2.2 percentage points less likely to have the second child within two years, allowing longer uninterrupted work spells. Third, parental separation: mothers with daycare access are 2 percentage points more likely to live apart from the child&amp;rsquo;s father at child age 16, consistent with greater economic independence from labor market participation reducing barriers to separation. Additional educational attainment after first birth is tested and found to be an insignificant channel (no significant effect overall, a marginal effect only for low-educated mothers), ruling out re-skilling as a mediator.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented-beyond-the-education-split"&gt;Q5. What heterogeneity is documented beyond the education split?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s primary heterogeneity analysis is by maternal education level (low: no post-secondary education versus higher: any post-secondary education including vocational training, college, or university). The education split produces the most substantive finding: employment effects are larger and more persistent for low-educated mothers, while the earnings-employment divergence is the distinctive feature for higher-educated mothers. No other dimensions of heterogeneity (by birth cohort, by municipality type beyond the urban indicator, by parity) are formally reported in the main results, though geographic robustness checks (exclusion of three largest cities, exclusion of suburbs) implicitly test whether effects are concentrated in particular settings and find they are not.&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run"&gt;Q6. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Four main sets of robustness checks are reported. First, selective migration: regressions of daycare availability on distance moved from birthplace (linear, quadratic, and IHST-transformed) with the full conditioning set show no significant relationship, ruling out systematic sorting into daycare neighborhoods. Second, pre-1970 cohort exclusion: restricting to mothers with first birth after 1970 (for whom the 1970 census address is predetermined relative to birth) yields qualitatively similar results, though participation effect sizes are somewhat smaller. Third, urban geography: excluding the three largest municipalities (Copenhagen, Frederiksberg, Aarhus, Odense) and separately excluding suburbs of Copenhagen and Aarhus both leave the main results intact. Fourth, differential time trends: allowing the most populous neighborhood within each municipality to have its own set of time dummies (to capture potentially faster urban trend evolution) does not change the finding that participation and earnings effects persist beyond 30 years. The paper also shows that results are robust to an alternative participation definition based solely on ATP contributions for all years (versus mixing ATP pre-1980 and earnings post-1980).&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-prior-work-and-what-is-its-main-contribution"&gt;Q7. How does this paper relate to prior work and what is its main contribution?&lt;/h3&gt;
&lt;p&gt;The prior literature falls into two camps. The short-run camp (Havnes and Mogstad 2011 for Norway; Carta and Rizzica 2018 for Italy; Bettendorf et al. 2015 for Netherlands; Cascio 2009 and Fitzpatrick 2012 for the US) documents modest to moderate employment effects during the preschool years. The medium-run camp (Lefebvre et al. 2009 and Haeck et al. 2015 for Quebec; Nollenberger and Rodriguez-Planas 2015 for Spain; Herbst 2017 for the US Lanham Act) tracks effects up to about 11–17 years. This paper&amp;rsquo;s first contribution is extending the window to 34 years — covering the majority of the working life — using Danish administrative data that allow continuous observation rather than decennial census snapshots. The second contribution is documenting the earnings-employment divergence for higher-educated mothers specifically, which was not visible in shorter windows. The third contribution is the simultaneous analysis of fertility, spacing, and parental separation as mediators using the same administrative data and identification strategy, rather than treating these as separate exercises in different papers.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-scope-conditions-and-policy-implications"&gt;Q8. What are the scope conditions and policy implications?&lt;/h3&gt;
&lt;p&gt;Several scope conditions qualify the policy implications. First, the context is a universal reform in a Nordic welfare state with strong labor market institutions and universal access; the results may not directly generalize to settings with low baseline female employment or weak formal sector employment. Second, the relevant margin for the 1960s–70s cohorts was daycare for children aged three to six; the paper notes that by recent decades the relevant margin has shifted to children under two (consistent with Simonsen 2010 finding effects for younger children in 2001 data), possibly reflecting changing cultural norms or the fact that 1960s–70s mothers had multiple children before returning to work. Third, the employment effects are larger for low-educated mothers, so the labor market attachment argument applies most forcefully to this group. Fourth, the negative fertility effects mean that the total welfare calculation must weigh labor market gains against reductions in desired family size. The policy implication the paper emphasizes is that universal daycare is an investment in long-run economic output, not merely a short-run participation subsidy, because the labor market attachment it induces during child-rearing years compounds over careers through human capital accumulation.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-sample-and-data-structure"&gt;Q9. What is the sample and data structure?&lt;/h3&gt;
&lt;p&gt;The sample consists of 370,602 mothers who had their first child between 1964 and 1975 and were resident in Denmark in 1970 (from the census), after excluding women with immigrant backgrounds (2.2 percent) and those who died or emigrated before the first child turned 16 (0.6 percent). Employment is observed from the birth of the first child through 34 years after (1964–2009 approximately); earnings from 1980 through 2015. The daycare panel is constructed from historical yearbooks (1964–1975) and administrative registers (1976–1993) and provides yearly neighborhood-level data on daycare availability. The average mother in the sample was born in 1945, was 23.7 years old at first birth, had 10.8 years of education, and had 2.2 children total. The sample is split roughly 50/50 between low-educated and higher-educated mothers.&lt;/p&gt;
&lt;h3 id="q10-why-do-effects-appear-only-when-the-child-is-three-not-earlier"&gt;Q10. Why do effects appear only when the child is three, not earlier?&lt;/h3&gt;
&lt;p&gt;The paper finds that contemporary participation effects are small and statistically insignificant for years zero through two, then jump sharply at year three. The paper attributes this to two factors: (1) the universal daycare reform primarily expanded slots for children aged three to six, with nurseries for children under three expanding much more slowly through the 1980s and 1990s (Figure A.1 in the paper); and (2) cultural norms and the multi-child fertility pattern of this cohort — mothers in the 1960s–70s were more likely to have multiple children before returning to work, implying that the eldest child often reached age three or four before the mother re-entered employment. This contrasts with more recent periods (Simonsen 2010 uses 2001 data) where the relevant margin has shifted to children under two.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Universal daycare&lt;/strong&gt;: In the paper&amp;rsquo;s sense, daycare centers open to children from all socioeconomic backgrounds (not means-tested), with building costs fully publicly funded and operating costs split among state, municipality, and parents (with parents paying 30 percent), following the 1964 Danish reform. Contrasted with the pre-reform &amp;rsquo;targeted&amp;rsquo; system that only subsidized institutions where two-thirds of children came from low-income families.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Working lifetime effects&lt;/strong&gt;: The paper&amp;rsquo;s central object of analysis: the causal impact of early daycare access on maternal labor outcomes measured annually across 34 years after the birth of the first child, covering the majority of the working life. Distinguished from short-run (0–7 year) and medium-run (up to 11–17 year) effects documented in prior work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor market attachment&lt;/strong&gt;: As used in the paper, the sustained connection to paid employment during the child-rearing years (when children are of preschool age). The paper argues that attachment during this period is the mechanism for long-run effects because it reduces human capital depreciation and enables on-the-job accumulation of experience and job-specific skills.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;ATP (Supplementary Pension Fund) contributions&lt;/strong&gt;: The paper&amp;rsquo;s primary employment measure for years before 1980. Annual ATP contributions are proportional to hours worked: one-third contribution corresponds to 10–19 hours/week, two-thirds to 20–29 hours/week, and full contribution to 30 or more hours/week. Used to construct both a participation dummy and a full-time employment dummy (full ATP contribution = at least 30 hours/week). Crucially, the unemployed, self-employed, and those outside the labor force made no ATP contributions during this period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Human capital depreciation channel&lt;/strong&gt;: The mechanism by which absence from the labor market during child-rearing years erodes previously accumulated skills (from education and prior work). The paper uses this concept, following Adda et al. (2017) and Lefebvre et al. (2009), to explain why participation effects on earnings can persist long after direct employment effects have diminished: mothers who worked during preschool years entered subsequent career phases with a larger, less-depreciated human capital stock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Secondary fertility decisions&lt;/strong&gt;: The paper&amp;rsquo;s term for fertility choices conditional on already having a first child, i.e., the decision to have additional children. Examined on the intensive margin (number of additional children, spacing between births) rather than extensive margin (whether to have any children), because the sample consists entirely of women who already have at least one child.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Daycare for 3–6 year olds vs. 0–2 year olds&lt;/strong&gt;: The paper distinguishes between two types of daycare that expanded at different speeds: daycare for children aged 3–6 expanded rapidly from 1966, while nurseries for children under 3 (crèches) expanded only from the 1980s–1990s. All significant effects in the paper — on employment, fertility, and parental separation — load onto access to daycare for children aged 3–6, not 0–2, consistent with the historical timing of the expansion.&lt;/p&gt;</description></item><item><title>University Research and the Market for Higher Education</title><link>https://macropaperwarehouse.com/papers/university-research-and-the-market-for-higher-education/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/university-research-and-the-market-for-higher-education/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper proposes that university R&amp;amp;D is determined endogenously by competition for tuition and talented students in the market for higher education, and asks why universities fund research internally with tuition despite negligible returns to patenting. Motivation: between 2000 and 2018 U.S. universities accounted for 13% of aggregate R&amp;amp;D spending and 53% of all basic-research spending, yet in 2018 over 25% of university research was internally funded (25.54% in 2018; federal government 52.97%) while between 1991 and 2018 the median university earned patent licensing revenue totaling less than 2% of its R&amp;amp;D expenditure. Internal funds therefore come essentially from tuition.&lt;/p&gt;
&lt;p&gt;Approach: (1) four stylized facts from administrative microdata (IPEDS, NSF HERD survey covering 916 universities / 99.1% of sector R&amp;amp;D, AUTM patent-licensing survey, Web of Science / Leiden bibliometrics); (2) a causal natural experiment; (3) a general-equilibrium model of the higher-education sector with heterogeneous universities choosing teaching and research, calibrated to U.S. data; and (4) policy counterfactuals.&lt;/p&gt;
&lt;p&gt;Causal evidence: the authors exploit the 1998-2003 doubling of the NIH budget (from $13.6bn to $27.1bn) using a Bartik shift-share instrument built from each university&amp;rsquo;s pre-period (1993-1997) share of federal life-science grants, regressing the change in net tuition (1993-1997 to 2004-2008) on the instrumented change in R&amp;amp;D per student, with state-clustered standard errors and state-specific trends. The benchmark estimate is that a $1.00 increase in R&amp;amp;D spending per student raises tuition by $0.15 (s.e. 0.05) — universities recoup up to 15% of R&amp;amp;D through higher tuition. Across specifications the effect ranges $0.10-$0.15; it is driven by research universities (non-liberal-arts), is statistically insignificant for liberal arts colleges, and a placebo using student-amenities spending shows no significant effect. The point estimate is about 60% larger at private non-profits than publics, but that difference is not statistically significant.&lt;/p&gt;
&lt;p&gt;Model and mechanism: education quality q = k^ωk * z̄^ωz * eT^ωe depends on intangible knowledge capital k (accumulated via research, k&amp;rsquo; = k^γk * eR^γe), peer ability z̄, and teaching spending. Universities maximize discounted education quality, funding research from tuition. Equilibrium features an endogenous college hierarchy with two-dimensional sorting by ability and family income. The research share sR rises with the steepness of the college quality-ladder Σq/Σk; when students are highly stratified or tuition rises sharply with rank, universities invest in research even if the direct contribution to teaching (ωk) is small — research persists even as ωk→0 (acting as a pure signal). Incentives fall when intangible capital is highly dispersed across colleges.&lt;/p&gt;
&lt;p&gt;Calibration matches the joint distribution of research, tuition, and student ability, plus untargeted R&amp;amp;D dispersion; simulated NIH expansion yields $0.18 per $1 in steady state and $0.11 along the transition, bracketing the empirical $0.10-$0.15.&lt;/p&gt;
&lt;p&gt;Policy findings (long-run, vs baseline): removing all need-based federal tuition subsidies cuts university research by 8.1% (replacing progressive with revenue-neutral flat tuition subsidy: -2.2%); progressive aid compresses revenue dispersion, steepens the quality-ladder, and raises the research share (+0.8 pp). Removing all federal research grants cuts research by 69.1% — only 6.9 pp below the government&amp;rsquo;s 76% funding share, implying crowding-out: the meritocratic grant structure concentrates funds at top schools, flattening the ladder and cutting the research share by 16.4 pp. A revenue-neutral flat research subsidy would instead raise research by 14.8%, human capital by 9.6%, and output by 11.1%.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;A Bartik/shift-share IV exploiting the 1998-2003 NIH budget doubling. Each university&amp;rsquo;s change in R&amp;amp;D is instrumented by its pre-period (1993-1997) share of all federal life-science research grants. Relevance: NIH was the bulk of federal life-science funding before the shock and did not substantially change award criteria, so high-share schools received mechanically larger funding increases. Exogeneity requires that universities did not systematically invest in life-science research in the pre-period in anticipation of the expansion. The estimation is in long-differences comparing steady states; standard errors are clustered at the state level with state-specific tuition trends. Threats: the NIH expansion occurs at a common point in time, so it may correlate with other contemporaneous market changes; initially larger or higher-quality research universities might have raised tuition for reasons unrelated to R&amp;amp;D. The authors address this with group-specific time trends (public/private, pre-existing life-science status, school size, initial quality via faculty-student ratio) and pre-trend controls (1987-1992 faculty-student ratio, FTE size, life-science status). A limitation the authors acknowledge: they cannot test the effect on subsequent student ability because ability proxies are only available after the intervention.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-and-how-are-they-distinguished"&gt;Q2. What are the main mechanisms and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;The college quality-ladder Σq/Σk (the cross-sectional elasticity of education quality with respect to intangible capital) is the sufficient statistic for research incentives. Equation (14) decomposes it into three channels: (i) the direct teaching contribution of research ωk; (ii) attracting better students, ωz × Σz̄/Σk; and (iii) charging higher tuition, ωe × ΣR/Σk. Channels (ii) and (iii) flow from competition for talented students and tuition and can dominate even when ωk is tiny. Empirically, Σz̄/Σk maps to the cross-sectional elasticity of student ability w.r.t. research (Figure 3) and ΣR/Σk to the elasticity of tuition w.r.t. research (Figure 4), so the calibration disciplines these channels with observable cross-sectional relationships.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;The tuition effect is concentrated in research universities (non-liberal-arts), with a larger, highly significant point estimate; for liberal arts colleges the NIH shock has no statistically significant effect on tuition (the authors caution the LAC sample is smaller — ~32% of institutions, ~24% of FTE — and more heterogeneous, so power may be insufficient). The effect appears ~60% stronger at private non-profits than publics, but the difference is not statistically significant. Across the model, top schools and bottom schools both invest less in research when intangible capital is highly dispersed (top schools face weak incentives to improve already-secure rank; bottom schools find climbing too costly).&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Empirically: adding pre-trend controls (column 3) leaves estimates intact; splitting by NLA vs LAC; and a placebo replacing R&amp;amp;D with student-services (amenities) spending, which yields no significant effect, rejecting spurious cross-category correlation. In the model: (1) the limiting case ωk→0 where research is a pure signal — the research share falls from 8.8% to 2.4% of tuition but stays strictly positive, and policy effects retain 50% (tuition-subsidy removal: -0.4 pp vs -0.8) and 66% (research-subsidy removal: +10.8 vs +16.4 pp) of their magnitude; (2) allowing some teaching expenditure to also enter intangible-capital production (γT&amp;gt;0), where the research share falls from 8.8% to 4.7% and policy effects moderate (-0.4 pp and +7.1 pp). In both, existing tuition policies still boost research and federal research grants still crowd it out.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-relate-to-and-differ-from-prior-work"&gt;Q5. How does this relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;It builds on equilibrium higher-education models — Epple, Romano &amp;amp; Sieg (2006) (quality maximization, exogenous endowment hierarchy, finite universities with market power) and Cai &amp;amp; Heathcote (2022) (competitive, constant-returns technology) — but endogenizes university R&amp;amp;D alongside teaching. A theoretical contribution is proving existence of a unique dynamic equilibrium with quality maximization and an endogenous college-quality hierarchy with a continuum of colleges; Cai &amp;amp; Heathcote argued no quality-maximization equilibrium exists when colleges are ex-ante identical (all want to be at the top), which this paper resolves via the endogenous knowledge hierarchy. It contributes to the economics of science / university-R&amp;amp;D literature by adding market-driven incentives, and to the basic-research-subsidy literature (Akcigit et al.) by showing universities have private incentives to do basic research, implying the need for government subsidy may be smaller than the standard Nelson/Arrow/Rosenberg view holds.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q6. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Two main implications. First, a novel complementarity between equity and innovation: progressive need-based tuition aid compresses revenue dispersion across colleges, makes them more similar, steepens the quality-ladder, and raises research (+8.1% relative to a no-subsidy world; flat subsidy gives only ~one-quarter of that, +2.2%). Second, current meritocratic federal research grants partially crowd out internal research and raise educational inequality by concentrating resources at top schools; removing them cuts research by 69.1% (only 6.9 pp below the 76% federal share, the gap being the crowding-out). A revenue-neutral flat research subsidy would raise research by 14.8%, human capital 9.6%, and output 11.1%, eliminating the equity-innovation trade-off because it lowers research cost without altering market structure. Scope conditions: these are long-run steady-state comparisons in a calibrated model of 4-year public and private non-profit U.S. institutions; magnitudes depend on the hard-to-measure ωk and on the research-technology specification, as the robustness exercises show.&lt;/p&gt;
&lt;h3 id="q7-why-do-universities-fund-research-from-tuition-rather-than-patents-and-does-the-model-rationalize-it"&gt;Q7. Why do universities fund research from tuition rather than patents, and does the model rationalize it?&lt;/h3&gt;
&lt;p&gt;Because patent licensing is too small (median &amp;lt;2% of R&amp;amp;D, 1991-2018) to fund the &amp;gt;25% of R&amp;amp;D that is internal, and unrestricted operating funds are composed almost entirely of tuition (much of it from unrecovered facilities-and-administration costs on sponsored projects — roughly $7bn in 2018). The model rationalizes diverting tuition to research because research raises education quality and thus students&amp;rsquo; willingness to pay, so in a competitive sector students accept it. The model also replicates the joint pattern that higher-R&amp;amp;D universities are higher-ranked, attract wealthier and abler students, and charge higher tuition.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-sources-of-inefficiency-in-the-model"&gt;Q8. What are the sources of inefficiency in the model?&lt;/h3&gt;
&lt;p&gt;Two. First, borrowing constraints prevent efficient sorting of students by ability (a social planner would send the ablest to the best colleges, but students are limited by parental capacity to pay). Second, university knowledge has positive spillovers to the real economy (calibrated ιk = 0.1) that colleges do not internalize, causing under-investment; however, quality-maximizing colleges face extra competitive incentives to do research, so net under- or over-investment is ambiguous and depends on stratification relative to spillover strength.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;College quality-ladder (Σq/Σk)&lt;/strong&gt;: The equilibrium cross-sectional elasticity of education quality with respect to a university&amp;rsquo;s intangible knowledge capital — a sufficient statistic for a university&amp;rsquo;s private incentive to invest in research. Steeper ladder (more stratification, tuition rising more with rank) means stronger research incentives.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intangible (knowledge) capital k&lt;/strong&gt;: Institution-specific intangible capital accumulated by investing in research (k&amp;rsquo; = k^γk eR^γe). It is primarily frontier knowledge and ideas exposed to students, but also networks, recruiting, labs, and methods; it can act purely as a reputation signal in the limiting case ωk→0.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Research share (sR)&lt;/strong&gt;: The share of a university&amp;rsquo;s tuition revenue allocated to research in equilibrium (≈8.8% under existing policies). It increases with college forward-lookingness (βc) and the steepness of the quality-ladder, and decreases with the dispersion of intangible capital across colleges.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crowding-out of internal research&lt;/strong&gt;: In the paper&amp;rsquo;s sense, the phenomenon whereby federal grants, by concentrating funds at top schools, raise the dispersion of research (Σk), flatten the quality-ladder (Σq/Σk), lower the research share, and thereby reduce universities&amp;rsquo; internal research spending — so total research rises less than the government&amp;rsquo;s funding share (69.1% decline vs 76% share on removal).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Equity-innovation complementarity&lt;/strong&gt;: The model&amp;rsquo;s finding that progressive need-based tuition aid, by compressing revenue dispersion and making colleges more similar, steepens competition and raises university research — so equity-promoting policy also boosts basic research, rather than trading off against it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Education-innovation gap (ωk calibration)&lt;/strong&gt;: Biasi &amp;amp; Ma&amp;rsquo;s (2021) measure of how frontier-current a university&amp;rsquo;s curriculum is, interpreted in the model as log(k). A one-unit decrease is associated with a 0.011% rise in graduate income; normalized by its school-level standard deviation of 0.85, it is used to pin down ωk via ωk·α = .011/.85·Σk.&lt;/p&gt;</description></item><item><title>Warming with Borders: Forced Climate Migration and Carbon Pricing</title><link>https://macropaperwarehouse.com/papers/warming-with-borders-forced-climate-migration-and-carbon-pricing/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/warming-with-borders-forced-climate-migration-and-carbon-pricing/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper asks how the threat of forced climate migration — international displacement driven by climate-induced natural disasters — should alter optimal carbon taxation. The motivation is twofold. First, climate change is intensifying natural disasters that disproportionately afflict developing nations, generating large cross-border population flows that existing integrated assessment models (IAMs) ignore. Second, migration and climate policy are simultaneously among the most contested political issues, yet their interaction has received almost no joint economic analysis.&lt;/p&gt;
&lt;p&gt;The paper proceeds in two stages. First, it documents empirically that natural disasters cause international migration. Using a global annual panel (165 countries, 1980–2013) from EM-DAT and UN migration flow tables, the paper estimates a fixed-effects regression of log-migration flows from developing (origin) to developed (host) countries on disaster frequency, controlling for GDP per capita and population. The key coefficient implies a semi-elasticity of approximately 2.3%: a unit increase in natural-disaster occurrence is associated with a 2.3% rise in migration to host regions. To link disaster frequency to carbon concentrations, a time-series cointegration analysis yields an elasticity of 13.49 for climatological and hydrological disasters (6.74 when meteorological disasters are added), implying an overall elasticity of climate refugees to CO2 concentrations of 11.87 (5.93 with meteorological events).&lt;/p&gt;
&lt;p&gt;Second, these empirical estimates calibrate a quantitative multi-region integrated assessment model (IAM) in which energy-related emissions generate two externalities simultaneously: output damage through temperature, and population reallocation from origin to host regions. The model features a North–South structure (Kyoto Annex I countries as host; rest of world as origin), Cobb-Douglas production with capital, labor, and energy (coal-proxy), region-specific climate damage parameters drawn from Hassler et al. (2019), and a climate module following Golosov et al. (2014). Social welfare in host regions can optionally include a direct disutility from immigration (parameterized using data on European Pay-to-Go programs and the 2016 EU–Turkey Agreement). The model is simulated over 300 years starting from 2015, with 10-year periods.&lt;/p&gt;
&lt;p&gt;The paper then analytically characterizes and quantitatively estimates optimal carbon prices under three policy regimes: (1) unilateral host-only action, (2) globally cooperative (first-best), and (3) a Nash equilibrium with all regions active.&lt;/p&gt;
&lt;p&gt;The central quantitative finding is an asymmetry across policy regimes. Under unilateral host-region action, accounting for forced climate migration raises the optimal carbon price by approximately 22% (from $44.72 to $54.73 per ton of carbon when calibrated to climatological and hydrological disasters only; to $49.77, an 11% increase, when meteorological events are included). The dominant mechanism is the &amp;ldquo;Labor Effect&amp;rdquo;: migrants move without capital and dilute per capita income in host regions because environmental resources and capital are finite, making the negative welfare consequences exceed the positive labor-supply benefit under a Cobb-Douglas technology with climate damages. The social cost of immigration (disutility of anti-immigration sentiment) adds only marginally to the carbon price ($54.99 vs. $54.73 per ton under the Pay-to-Go calibration). When border control is modeled explicitly, a planner facing US-calibrated deportation costs ($4.6 × 10^5 per immigrant) prefers tightening the carbon tax over using border control, validating the main finding. Only when border control is costless does the optimal strategy switch to low carbon taxes and restricted immigration.&lt;/p&gt;
&lt;p&gt;In contrast, the globally optimal SCC is nearly unchanged by forced climate migration ($118.62 without FCM vs. $123.03 with FCM), because the Global Labor Effect balances out: costs of population growth in the host are offset by the adaptation benefit of relocating people to less climate-vulnerable areas. Under Nash equilibrium, host SCCs rise modestly ($44.72 to $49.89 under C&amp;amp;H disasters), while origin SCCs fall slightly ($73.81 to $72.51) as migrants, once relocated, face lower climate damages. The welfare cost to host-region natives from applying the no-FCM policy when FCM is in fact present amounts to a 0.193% permanent consumption equivalent.&lt;/p&gt;
&lt;p&gt;Policy implication: in the absence of a global climate agreement (the prevalent situation), developed countries have substantially stronger unilateral incentives to price carbon than existing IAMs suggest, because they indirectly bear the economic costs of climate-induced immigration. The global SCC, however, is not materially affected, so the case for international coordination rests on the same foundation as before.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-empirical-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the empirical identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The empirical strategy exploits the quasi-random timing of natural disasters within an origin country using a two-way fixed-effects (country and year) panel regression. The dependent variable is the log of annual unilateral migration flows from each origin country to the pooled group of host countries (43 OECD-type destinations). The independent variable is the frequency (or log frequency) of climate-related natural disasters in the origin country in the same year. Country fixed effects absorb time-invariant push/pull factors; year fixed effects absorb common global shocks. Main threats discussed: (1) Endogeneity of contemporaneous GDP and population, addressed by using first lags of controls. (2) Reporting bias in EM-DAT (disasters in early years may be under-recorded), addressed by computing the ratio of warming-related to geophysical disasters (reporting bias should be type-orthogonal) and by restricting to large disasters (&amp;gt;=1,000 affected or &amp;gt;=100 deaths). (3) The paper focuses exclusively on the contemporaneous (same-year) migration response, treating lagged effects as lower bounds. (4) The semi-elasticity estimates are used as calibration inputs, not as causal estimates of structural parameters — the author acknowledges the causal chain from concentrations to disasters is not fully established.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-four-theoretical-components-of-the-unilateral-host-scc-and-how-do-they-combine"&gt;Q2. What are the four theoretical components of the unilateral host SCC and how do they combine?&lt;/h3&gt;
&lt;p&gt;The unilateral host SCC (equation 12) is the sum of: (1) Standard Output Damages — the present discounted value of climate damage to final output, the only component in standard IAMs; (2) Emissions Reallocation — the reduction in origin-region emissions as migrants move to the host, which lowers global concentrations and benefits the host, making this component negative (it reduces the carbon price); (3) Immigration Social Cost — the direct disutility of newly arrived immigrants borne by host natives (parameterized by gamma), which adds to the carbon price when gamma &amp;gt; 0; and (4) Labor Effect — the net welfare consequence of a larger host labor force, which comprises a positive externality (higher output) and a negative externality (dilution of per capita consumption due to finite environmental resources and capital). Under Cobb-Douglas production with climate damages and capital (Result 1), the net Labor Effect is always a negative externality that raises the carbon price. In the quantitative exercise, the Labor Effect dominates all other FCM-related components and accounts for essentially the entire 22% increase in the unilateral SCC.&lt;/p&gt;
&lt;h3 id="q3-why-does-the-global-scc-remain-nearly-unchanged-when-forced-climate-migration-is-included"&gt;Q3. Why does the global SCC remain nearly unchanged when forced climate migration is included?&lt;/h3&gt;
&lt;p&gt;The global planner internalizes the welfare of both host and origin regions. The &amp;lsquo;Global Labor Effect&amp;rsquo; contains two offsetting terms: costs to host natives from capital dilution and per capita income reduction, and benefits to origin-region emigrants who move to a less climate-vulnerable, more economically developed area. These effects largely cancel. In addition, migration reallocates economic activity away from high-damage origin regions, lowering expected global climate damages. Migration costs calibrated to equalize consumption per capita across regions (absent climate change) prevent the global planner from strategically using pollution to trigger welfare-improving migration. Quantitatively, the global SCC rises only slightly, from $118.62 to $123.03 per ton of carbon (less than 4%), and may even fall after roughly four decades as the adaptation benefit grows.&lt;/p&gt;
&lt;h3 id="q4-how-is-the-social-cost-of-immigration-anti-immigrant-sentiment-parameterized-and-calibrated"&gt;Q4. How is the social cost of immigration (anti-immigrant sentiment) parameterized and calibrated?&lt;/h3&gt;
&lt;p&gt;The parameter gamma represents the marginal social cost of immigration to native households — their willingness to pay to prevent a marginal unit of immigration. Two calibration approaches are used: (A) Pay-to-Go programs: using data on European Assisted Voluntary Return programs in 2015, the paper derives gamma = 7.1 × 10^3 (in terms of final good per billion migrants). (B) EU-Turkey Agreement: using costs from the 2016 deal managing the Syrian refugee influx, the paper derives gamma = 7.3 × 10^3. The similarity of the two estimates provides cross-validation. The baseline quantitative exercise disables this feature (gamma = 0), treating it as a sensitivity; a UK Brexit-era survey value implies a four-fold increase in the unilateral SCC but is judged unrepresentative of permanent preferences. The paper is explicit that these are positive descriptions of political preferences, not normative endorsements.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-in-the-migration-response-is-documented-empirically"&gt;Q5. What heterogeneity in the migration response is documented empirically?&lt;/h3&gt;
&lt;p&gt;Three dimensions of heterogeneity are explored: (1) Income: Unlike for slow-onset climate migration (where middle-income countries drive the response), poorer countries show a stronger migration response to disasters (positive and significant interaction between disaster frequency and a poor-country dummy, column 4 of Table B.1). This is interpreted as evidence that migration costs are less binding when disaster severity forces departure. (2) Disaster type: Climatological and hydrological disasters have higher and statistically significant migration-response coefficients than meteorological disasters (Table B.5). This differential is why the paper presents results under two calibrations (C&amp;amp;H disasters vs. C&amp;amp;H&amp;amp;M disasters). (3) Disaster severity: Restricting to large disasters (&amp;gt;=1,000 affected or &amp;gt;=100 deaths) yields an even larger migration response (column 5 of Table B.1).&lt;/p&gt;
&lt;h3 id="q6-what-robustness-checks-are-run-on-the-empirical-results"&gt;Q6. What robustness checks are run on the empirical results?&lt;/h3&gt;
&lt;p&gt;The paper runs an extensive set of checks reported in Online Appendix B: (1) Zero-inflated negative binomial (ZINB) model to handle zeros in the dependent variable. (2) Bilateral migration flows with origin-destination fixed effects. (3) Three-year non-overlapping windows (to reduce zero mass in independent variable), which more than doubles the estimated coefficients. (4) Per capita migration as the dependent variable. (5) Disaster frequency weighted by share of affected population. (6) Inverse hyperbolic sine (IHS) transformation. (7) Excluding China and India. (8) Excluding Singapore and South Korea. (9) Controlling for conflict (battle-related deaths). (10) Controlling for a climate vulnerability index. (11) Controlling for the second lag of disasters. (12) Polynomial regression to check for acceleration. (13) Poisson specification. (14) Checking that an upward trend in disaster ratios relative to geophysical events is not attributable to reporting bias. Results are consistent across all specifications.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-nash-equilibrium-result-and-how-does-it-differ-from-both-the-unilateral-and-first-best-settings"&gt;Q7. What is the Nash equilibrium result, and how does it differ from both the unilateral and first-best settings?&lt;/h3&gt;
&lt;p&gt;In the Nash equilibrium, each region implements its own best-response carbon policy. Host regions&amp;rsquo; NE SCC resembles the unilateral SCC (Section 4) except that the &amp;lsquo;Emissions Reallocation&amp;rsquo; component drops out, because when all regions are strategically active, the host cannot treat origin emissions as exogenously reduced by migration. Quantitatively, host NE SCC rises from $44.72 (no FCM) to $49.89 (with FCM, C&amp;amp;H disasters) — a roughly 11.5% increase. Origin region NE SCC falls slightly from $73.81 to $72.51, because origin planners care about the welfare of their emigrants who now live in lower-damage host regions. Without FCM, the origin SCC is 1.6 times higher than the host SCC (reflecting greater vulnerability and larger population in origin). With FCM, this gap narrows. The NE global SCC is lower than the first-best because each region only partially internalizes the global externality.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-border-control-extension-interact-with-the-optimal-carbon-tax"&gt;Q8. How does the border control extension interact with the optimal carbon tax?&lt;/h3&gt;
&lt;p&gt;When the host planner can choose both a carbon tax and a border control stringency (share of migrants admitted), the optimal carbon tax with FCM is lower than in the no-border-control case, because restricting migration inflows reduces both the Labor Effect cost and the Immigration Social Cost. At the same time, restricting inflows reduces the Emissions Reallocation benefit. In equilibrium, the marginal cost of deportation equals the net benefit of keeping an additional immigrant out. Quantitatively, when border control costs are calibrated to US Department of Homeland Security data ($4.6 × 10^5 per detained immigrant), the carbon tax remains essentially equal to the no-border-control case and migration inflows are also nearly unchanged — the planner finds it optimal to abate emissions rather than pay deportation costs. Only when border control is costless does the planner switch to a low carbon tax and high migration restriction. This sensitivity analysis validates the main finding under realistic border enforcement costs.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-cruz-and-rossi-hansberg-2024"&gt;Q9. How does this paper relate to, and differ from, Cruz and Rossi-Hansberg (2024)?&lt;/h3&gt;
&lt;p&gt;Cruz and Rossi-Hansberg (2024) use a highly spatially disaggregated model with endogenous migration to quantify welfare costs of climate change under an exogenous global carbon tax. The key differences are: (1) This paper derives optimal carbon taxes — both globally and regionally — rather than taking them as exogenous. (2) This paper provides closed-form analytical characterizations of the SCC under multiple policy regimes, enabling clear decomposition of mechanisms. (3) Migration in this paper is exclusively &amp;lsquo;forced&amp;rsquo; (disaster-driven), not microfounded by economic incentives (though Appendix F relaxes this); Cruz and Rossi-Hansberg treat migration as fully endogenous to economic conditions. (4) This paper explicitly analyzes strategic interactions (Nash equilibrium) between regions. (5) This paper can account for anti-immigration sentiment (gamma) and border control policies. The approaches are thus complementary: Cruz and Rossi-Hansberg offer richer spatial geography and fully endogenous migration; this paper offers analytical tractability and policy-regime analysis.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q10. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The principal implication is that developed countries (host regions) have approximately 22% stronger unilateral incentives to impose a carbon tax than existing IAMs indicate, once climate-induced international displacement is accounted for. This result holds under climatological and hydrological disasters calibration and US-level border enforcement costs; it is smaller (~11%) when meteorological events are added and even smaller when border control is assumed freely available. The global SCC is barely affected, so the normative case for a global agreement is not strengthened or weakened in magnitude, but the analytical structure of the globally optimal tax is qualitatively different. Scope conditions: the model abstracts from internal migration, micro-founded voluntary migration, endogenous TFP growth, and capital mobility across regions. Results are robust to Stern discounting, more catastrophic damage functions, and Negishi weights. The welfare cost of ignoring FCM in policy design is modest in magnitude (0.193% consumption equivalent) but positive and policy-relevant as a systematic downward bias in host-country incentives.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-microfounded-migration-extension-show"&gt;Q11. What does the microfounded migration extension show?&lt;/h3&gt;
&lt;p&gt;Online Appendix F relaxes the forced-migration-only assumption by introducing economically motivated migration: individuals in the origin choose migration based on consumption differentials across regions, subject to migration costs calibrated to eliminate non-climate migration at steady state. The host unilateral SCC rises to $79.52 per ton of carbon under microfounded migration, compared to $54.73 under forced-only climate migration and $44.72 with no migration (Table F.1). This indicates the 22% increase in the main analysis is a lower bound: broader climate-related migration (including voluntary economic responses to climate shocks) would generate even larger incentives for host regions to tighten carbon pricing. However, this extension sacrifices analytical tractability and closed-form solutions.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-welfare-cost-of-ignoring-fcm"&gt;Q12. What is the welfare cost of ignoring FCM?&lt;/h3&gt;
&lt;p&gt;Table 6 reports the welfare cost of applying the sub-optimal &amp;rsquo;no FCM&amp;rsquo; carbon tax to a world in which FCM is actually occurring. The cost is measured as the percentage increase in consumption in every period that would be needed to make host-region natives as well-off as they would be under the correctly calibrated FCM-inclusive policy. Without immigration disutility, the cost is 0.193%. With the Pay-to-Go disutility calibration, it is 0.195%. These figures are small but positive and increasing in the social cost of immigration. They represent the aggregate efficiency loss to host-region natives from the systematic underestimation of the unilateral SCC in existing IAMs.&lt;/p&gt;
&lt;h3 id="q13-how-is-the-migrationconcentrations-link-empirically-constructed-for-model-calibration"&gt;Q13. How is the migration–concentrations link empirically constructed for model calibration?&lt;/h3&gt;
&lt;p&gt;The paper uses an elasticity decomposition: the elasticity of climate refugees to CO2 concentrations is the product of two elasticities. The first — the elasticity of migration to disaster frequency — is estimated from the panel regression and equals 0.88 after pooling countries into two regions. The second — the elasticity of disaster frequency to carbon concentrations — is estimated from a time-series cointegration analysis following Thomas and Lopez (2015), yielding 13.49 for climatological and hydrological disasters alone and 6.74 when meteorological events are included. The product gives overall elasticities of 11.87 and 5.93 respectively. These are then used to calibrate the linear migration function B (the flow of migrants per unit change in carbon concentrations), using historical average concentration increases, average migration flows relative to host population, and the elasticities. B = 5.03 × 10^-5 (C&amp;amp;H disasters) or 2.52 × 10^-5 (C&amp;amp;H&amp;amp;M disasters).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Forced Climate Migration (FCM)&lt;/strong&gt;: In the paper&amp;rsquo;s usage, the specific subset of climate migrants who are forced to move internationally because of climate change-induced natural disasters (rapid-onset events such as floods, storms, and heatwaves), as distinct from voluntary economic migration or migration driven by slow-onset climate variables such as temperature trends.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Cost of Carbon (SCC)&lt;/strong&gt;: The monetary value of the present and future economic damage caused by a marginal one-unit increase in carbon emissions today, which under the Pigouvian framework equals the optimal carbon tax. The paper distinguishes three variants: the unilateral host-region SCC, the globally optimal (first-best) SCC, and the Nash-equilibrium SCCs for host and origin regions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor Effect&lt;/strong&gt;: A novel component of the unilateral SCC in the model, capturing the net welfare consequence of a larger host-region labor force due to FCM. It contains a positive sub-term (higher labor raises output) and a negative sub-term (capital dilution and reduction in per capita consumption because environmental goods are finite). Under Cobb-Douglas production with climate damages and capital, the net Labor Effect is always negative (raises the carbon price), as shown in Result 1.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Emissions Reallocation&lt;/strong&gt;: The reduction in origin-region emissions that mechanically follows when population — and therefore emission-generating activity — moves from the high-emission-intensity origin region to the host region. This component enters the unilateral SCC with a negative sign (it reduces the carbon price), because the host planner benefits from lower global concentrations induced by fewer emitters in the origin.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Cost of Immigration&lt;/strong&gt;: The direct disutility experienced by host-country natives from the arrival of immigrants in the current period, parameterized by gamma, representing the native household&amp;rsquo;s marginal willingness to pay to prevent an additional unit of immigration. It is calibrated using data on European Pay-to-Go programs and the EU–Turkey Agreement. It adds to both the unilateral and Nash-equilibrium host SCCs, but quantitatively contributes only a small increment above the Labor Effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;North-South Calibration&lt;/strong&gt;: The paper&amp;rsquo;s two-region parameterization in which &amp;lsquo;host&amp;rsquo; corresponds to Kyoto Annex I countries (most European nations, the United States, Canada, Australia, New Zealand) and &amp;lsquo;origin&amp;rsquo; corresponds to the rest of the world. Host regions have higher GDP per capita, lower climate vulnerability parameters (theta), and higher emissions per capita; origin regions are more exposed to climate damages and more densely populated.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nash Equilibrium (non-cooperative) SCC&lt;/strong&gt;: The carbon price chosen by a local planner as the best response to other regions&amp;rsquo; optimal strategies, without the Emissions Reallocation component (since other regions&amp;rsquo; emissions are now also strategically set). In this setting, host SCCs rise relative to the no-FCM benchmark but less than under unilateral action; origin SCCs fall slightly because origin planners account for the welfare of emigrants residing in host regions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Integrated Assessment Model (IAM) with FCM&lt;/strong&gt;: The paper&amp;rsquo;s quantitative framework that combines a neoclassical multi-region growth model, a climate module following GHKT (Golosov et al. 2014), region-specific damage functions, and an endogenous migration flow driven by carbon concentrations. The model is solved by direct optimization over savings rates and energy-labor shares, simulated for 300 years, with each period representing 10 years.&lt;/p&gt;</description></item><item><title>Who Buys High and Sells Low: Trading against Expected Returns and Wealth Inequality</title><link>https://macropaperwarehouse.com/papers/who-buys-high-and-sells-low-trading-against-expected-returns-and-wealth-inequality/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/who-buys-high-and-sells-low-trading-against-expected-returns-and-wealth-inequality/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Wealth in the US is far more concentrated than income, even among the bottom 99%. In 2013, the next-49% (above the bottom 50%) earned 4.7 times the income of the bottom 50% but held 6.5 times the net worth (SCF 2013). Since housing is most Americans&amp;rsquo; primary vehicle of wealth accumulation, differences in housing returns could amplify wealth gaps. Prior work studied heterogeneity in risk-taking in housing; this paper instead studies the timing (mistiming) of housing trades: do some households consistently &amp;ldquo;buy high and sell low&amp;rdquo; relative to EXPECTED asset returns, and what does that do to portfolio returns and wealth inequality? Theory is ambiguous: pro-cyclical credit supply (Mian-Sufi, Rajan) predicts poorer, credit-constrained households buy more in booms (when expected returns are low); extrapolative expectations (Barberis et al., Kaplan-Mitman-Violante) predict richer, less-constrained households buy more in booms. So it is an open empirical question.&lt;/p&gt;
&lt;p&gt;Data and method: The author builds a novel annual balanced panel of real-estate ownership from CoreLogic (formerly DataQuick) assessor file (a 2012-2013 cross section, ~104 million records, ~94% of US population) plus transaction-deed records, working backwards from 2012-2013 to assign owners by year (owner on Dec 31). Owners&amp;rsquo; wealth/permanent-income is imputed from surnames: household wage income averaged at the surname level in the 1940 full-count Census (the latest full Census and first to ask income) is a strong predictor of those surnames&amp;rsquo; 2012-2013 wealth (Henry de Frahan and Sakong 2023). Surname population counts and racial shares come from the 2000 Census tabulations (in 2000, 151,671 surnames with 100+ people, covering 242M of 282M people = 85.8%). Two samples: a &amp;ldquo;long&amp;rdquo; sample 1988-2013 (148 counties, 674 jurisdictions, 11 states, ~21-25% of US population) and a &amp;ldquo;wide&amp;rdquo; sample 1998-2013 (36 states, &amp;gt;60% of US population). Expected asset returns are estimated following Cochrane (2011) by regressing one-year-ahead realized housing returns on the log rent-to-price ratio (rents from BLS owner-equivalent rent or imputed from IRS local income; house prices from CoreLogic HPI, with Case-Shiller and FHFA for robustness), at aggregate, CBSA, county and zip-code levels, using common or area-specific (heterogeneous) coefficients. The key estimand is the covariance between (residualized) log housing quantity held by a wealth group and the log expected asset return — the &amp;ldquo;active&amp;rdquo; timing component, decomposed via a lognormal first-order approximation (Calvet-Campbell-Sodini-style passive/active split). Specifications include group, time, and group-time-trend fixed effects to isolate cyclical-frequency timing from long-run trends and new construction.&lt;/p&gt;
&lt;p&gt;Main findings (with magnitudes): (1) Over 1988-2013, lower-wealth (lower 1940-income-percentile) surnames consistently held more housing pro-cyclically — buying when expected returns were low and selling when high. Portfolio expected returns from active trades are increasing in wealth (decreasing in pro-cyclicality), especially pronounced for the bottom 20% of the 1940 income distribution. (2) Using more disaggregated expected returns raises the estimated gradient almost monotonically: the coefficient on surname 1940 income percentile rises from 0.089 bp (aggregate) to 0.180 bp per percentile (zip code, heterogeneous coefficients, wide sample — the preferred specification). Aggregate returns bias the estimate downward toward zero. (3) The gradient is larger where expected-return volatility is higher: a one-standard-deviation higher expected-return volatility roughly doubles the wealth gradient (Table 3a, zip codes); meanwhile the extent of buy-high-sell-low behavior itself is statistically unrelated to volatility (Table 3b, near zero). (4) The positive overall return-on-wealth slope is driven by BETWEEN-race differences (non-White groups own housing highly pro-cyclically, consistent with Kermani-Wong); WITHIN race, portfolio expected returns are slightly DECREASING in wealth. (5) Quantitatively, projecting 1940 income percentiles onto the 2013 wealth distribution (via average home value and a housing Engel curve from the 2013 SCF), a 10% rise in net-worth percentile is associated with ~13 bp higher annual portfolio expected return; across the interquartile range this is a 65-basis-point per year differential — about two-thirds of the ~1% total realized-return spread Fagereng et al. (2020) find for financial wealth in Norway, here from timing alone. (6) A back-of-the-envelope calculation (APC out of labor income cy≈0.25 from PSID, wealth-to-labor-income ratio W/Y≈10 from SCF) implies the 65 bp differential raises the wealth share ~9% above the income share, accounting for roughly 20% (a fifth) of residual wealth concentration above income concentration across the interquartile range. Implication: time-series volatility of housing markets widens wealth inequality beyond income inequality; dynamic trade timing, not just average returns or asset heterogeneity, matters for wealth levels.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-conceptual-distinction-the-paper-insists-on-and-why-does-it-use-expected-rather-than-realized-returns"&gt;Q1. What is the core conceptual distinction the paper insists on, and why does it use expected rather than realized returns?&lt;/h3&gt;
&lt;p&gt;The paper measures &amp;lsquo;buying high and selling low&amp;rsquo; as the negative co-movement between the QUANTITY of an asset held and the EXPECTED asset return on it — not realized returns on completed trades. Three reasons: (1) Over a finite period some households get lucky/unlucky on unpredictable realized returns, but those wash out over the long run; only co-movement with the PREDICTABLE (expected) component survives to affect long-run wealth accumulation. (2) Expected returns are imputed as a log-linear function of the local rent-to-price ratio, observable at local levels, rather than realized returns on a specific property. (3) It computes returns on the whole stock of housing owned, not only traded units, because non-traders earning 0% realized return must be averaged in for wealth-inequality purposes. Example given: from 2007, aggregate housing had a realized return of -8% (-20% vs the 12% time-series average) but a +8% one-year expected return (-4% vs average); the paper focuses on the -4% expected, not the -20% realized.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identificationmeasurement-strategy-and-what-are-the-main-threats"&gt;Q2. What is the identification/measurement strategy and what are the main threats?&lt;/h3&gt;
&lt;p&gt;Identification rests on (a) imputing owner wealth from surname-level 1940 Census average wage income, validated against 2000 Census zip-code incomes (Table 1: strong, expected correlations, e.g., owner-occupant 1940 log wage loads ~1.6-1.8 on Census median income; investment-home owners&amp;rsquo; residence income loads positively even controlling for property-site income), and (b) estimating the covariance of residualized log quantity held with log expected asset returns at cyclical frequency, with group, time, and group-specific-trend fixed effects (equations 7-8) to strip out level differences, differential new construction, and long-run population/inequality/homeownership trends. Threats: surname-level estimates require additional assumptions to map to family-level behavior (handled via Henry de Frahan and Sakong 2023 framework; the author deliberately avoids 2010s surname income/consumption to prevent reverse causality with 1988-2013 trading); the samples are not nationally representative (more urban, larger boom-busts); expected returns are imprecisely estimated for short local time series; and new construction cyclicality could confound who-owns-when (argued orthogonal because the outcome is the portfolio expected-return differential — even if poorer residents buy new units in booms, they are acquiring risky assets when expected returns are low).&lt;/p&gt;
&lt;h3 id="q3-what-are-the-two-competing-theoretical-mechanisms-and-does-the-paper-claim-to-distinguish-which-one-operates"&gt;Q3. What are the two competing theoretical mechanisms, and does the paper claim to distinguish which one operates?&lt;/h3&gt;
&lt;p&gt;Mechanism A: pro-cyclical credit supply (market- or government-driven, Rajan 2011; Mian-Sufi 2009) relaxes constraints in booms, so credit-constrained POORER households buy/own more housing in booms (when expected returns are low). Mechanism B: extrapolative expectations (Barberis et al. 2015; Kaplan-Mitman-Violante 2017) make booms coincide with optimism, and RICHER, less-constrained households are better positioned to add exposure, so they own more in booms. The two give opposite cross-sectional predictions. The paper emphasizes that its quantification of the wealth-inequality impact does NOT depend on WHICH mechanism drives the pattern or why households buy high — it measures the covariance regardless. Empirically it finds the poorer-buy-in-booms pattern dominates, consistent with the credit-supply channel, but does not structurally separate the mechanisms.&lt;/p&gt;
&lt;h3 id="q4-what-heterogeneity-is-documented"&gt;Q4. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Three dimensions. (1) Geographic volatility: areas with more volatile expected returns (California, Florida prominently) show steeper wealth gradients in portfolio expected returns; one SD higher volatility roughly doubles the gradient (Table 3a). (2) Time period: the positive wealth slope holds both pre-subprime (1988-2002) and during the boom-bust, but is larger during the more-volatile subprime boom-bust. (3) Race: the overall positive slope of portfolio expected return on wealth is driven by BETWEEN-race variation — non-White groups own housing highly pro-cyclically (consistent with Kermani-Wong 2021, who attribute lower Black realized returns largely to foreclosures) — while WITHIN-race the gradient is slightly decreasing in wealth. The bottom 20% of the 1940 income distribution shows the most pronounced pro-cyclicality.&lt;/p&gt;
&lt;h3 id="q5-what-robustness-checks-are-run"&gt;Q5. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Quantity units: results robust to using number of properties (baseline), number of bedrooms, or square footage. Price indices: aggregate results similar using CoreLogic HPI, Case-Shiller, and FHFA (Table 2a columns: 0.080, 0.063, 0.057 bp). Samples: long (1988-2013) vs wide (1998-2013) give similar aggregate estimates. Rent source: BLS owner-equivalent rent vs IRS-income-imputed rents both yield strong predictability and similar gradients. Estimation of expected returns: common vs heterogeneous (area-specific) prediction coefficients both work, with heterogeneous generally larger. Validation of surname-wealth mapping via three sets of Census 2000 regressions (Table 1). Geographic disaggregation robustness (aggregate to CBSA to county to zip) shows monotone increase, and restricting to CBSA counties with BLS rent for apples-to-apples comparison (Online Appendix Table OA.3a) preserves results.&lt;/p&gt;
&lt;h3 id="q6-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q6. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It complements contemporaneous work on heterogeneity in REALIZED portfolio returns along income/race (Goldsmith-Pinkham-Shue 2020; Xavier 2021; Kermani-Wong 2021; Martinez-Toledano 2022; Wolff 2022) and the wealth-returns literature finding returns increasing in wealth (Bach-Calvet-Sodini in Sweden; Fagereng et al. in Norway; Garbinti-Goupille-Lebret-Piketty in France; Kuhn-Rios-Rull, Wolff in US). It differs by focusing on EXPECTED returns and the TIMING (covariance) channel rather than realized returns or asset heterogeneity, and by isolating the active-trade timing component on the whole housing stock. Its 65 bp interquartile differential from timing alone is ~two-thirds of Fagereng et al.&amp;rsquo;s ~1% total realized financial-return differential, highlighting that timing matters even absent asset heterogeneity. It also relates to cyclical homeownership-by-demographic literature (Goodman-Mayer 2018; Mabille 2023).&lt;/p&gt;
&lt;h3 id="q7-what-are-the-policytheoretical-implications-and-their-scope-conditions"&gt;Q7. What are the policy/theoretical implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Implication: because expected housing returns are time-varying and predictable, and lower-wealth households trade against them, trade timing widens wealth inequality beyond income inequality — and areas/periods with more volatile housing markets amplify this. Dynamic, asset-price-driven mechanisms (not just average returns) matter for wealth LEVELS, not merely their cyclicality. Scope conditions: the result requires expected returns to be genuinely time-varying and predictable (if EtR were constant, the covariance term vanishes); the lognormal approximation requires positive asset quantities (holds for housing, would fail for risk-free borrowing); the quantification depends on cy≈0.25 (PSID), W/Y≈10 (SCF), and the housing Engel-curve projection; samples are urban-skewed and not nationally representative; and the cross-sectional volatility-inequality prediction is only suggestively, not rigorously, tested (data limits on local wealth inequality).&lt;/p&gt;
&lt;h3 id="q8-what-does-the-formal-decomposition-propositions-2-3-deliver"&gt;Q8. What does the formal decomposition (Propositions 2-3) deliver?&lt;/h3&gt;
&lt;p&gt;Proposition 2 decomposes long-run average wealth return into (i) a participation term — the product of differences in average asset shares times expected returns (the focus of the risky-participation literature) — and (ii) a covariance term between asset shares and expected returns (this paper&amp;rsquo;s focus). The covariance term is nonzero only if expected returns are time-varying and asset shares vary across households. Proposition 3 splits the share-return covariance into a &amp;lsquo;passive&amp;rsquo; part (price changes mechanically move shares opposite to expected returns) and an &amp;lsquo;active&amp;rsquo; part (deliberate quantity adjustment), via a first-order lognormal approximation; a sufficiently contrarian active change can flip the covariance positive. The paper targets the active component, equation (4): E(mu) times cov(residual log quantity, log expected return).&lt;/p&gt;
&lt;h3 id="q9-what-are-the-key-caveats-the-author-flags"&gt;Q9. What are the key caveats the author flags?&lt;/h3&gt;
&lt;p&gt;(1) Estimates are fundamentally at the surname level; family/household interpretation needs extra assumptions. (2) Expected returns are noisily estimated, especially locally with short series; heterogeneous coefficients add error but allow meaningful heterogeneity. (3) The wealth-inequality quantification is explicitly &amp;lsquo;back-of-the-envelope&amp;rsquo; and depends on approximations (APC, W/Y ratio, Engel curve, household-vs-surname extrapolation assumption). (4) During the subprime boom-bust, realized returns were far more volatile than rent-to-price-predicted expected returns (Online Appendix Fig OA.1), so the expected-return measure deliberately understates realized volatility. (5) Aggregate expected returns bias the gradient toward zero, so even the preferred zip-code estimate is likely a lower bound if returns are heterogeneous at finer-than-zip levels. (6) Samples cover urban areas with larger boom-busts and are not US-representative.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Within-Firm Pay Inequality and Productivity</title><link>https://macropaperwarehouse.com/papers/within-firm-pay-inequality-and-productivity/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/within-firm-pay-inequality-and-productivity/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper investigates how within-firm pay inequality relates to firm-level labor productivity, using a novel linkage of three confidential U.S. Census Bureau datasets covering millions of workers at hundreds of thousands of firms from 2003 to 2015.&lt;/p&gt;
&lt;p&gt;The motivating puzzle is that the dramatic rise in U.S. wage inequality since the 1970s is well documented, but the firm-side determinants of within-firm pay dispersion have been difficult to study due to the absence of comprehensive matched employer-employee data in the United States. The paper asks whether firms&amp;rsquo; own productivity levels can explain the structure of pay inequality within firms, and whether rising aggregate productivity can account for the secular increase in the CEO-to-median-worker pay gap.&lt;/p&gt;
&lt;p&gt;The data come from three linked sources. The Longitudinal Employer-Household Dynamics (LEHD) program provides quarterly earnings for essentially all UI-covered workers from 2003 to 2015, covering all 50 states and Washington, D.C. These earnings encompass salaries, wages, bonuses, and exercised stock options, making them comprehensive for top earners. The Longitudinal Business Database (LBD) supplies annual firm-level revenue and employment, from which the key productivity measure — real revenue per worker, deflated to 2010 dollars using the PCE deflator — is constructed. The Management and Organizational Practices Survey (MOPS), a supplement to the Annual Survey of Manufactures conducted in 2010 and 2015, provides structured management scores (scaled 0 to 1) measuring the intensity of performance monitoring, target-setting, and incentive use across manufacturing firms. The main analysis sample restricts to firms with at least 100 full-year &amp;ldquo;6-quarter sandwich&amp;rdquo; workers to ensure clean measurement of annual earnings; it covers approximately 443,000 firm-year observations and 73,000 unique firms. A supplementary Execucomp sample (4,681 firms, 2006–2016) validates results for large publicly traded firms.&lt;/p&gt;
&lt;p&gt;Three main findings are reported. First, employees at more productive firms earn more across the entire within-firm pay distribution — from the 1st to the 99th percentile. A 10 percent increase in productivity is associated with a 0.7 percent increase in average worker pay (elasticity 0.068). Moving from the 10th to the 90th percentile of the firm productivity distribution projects an 18 percent increase in average pay.&lt;/p&gt;
&lt;p&gt;Second, the pay-productivity relationship is steeper at higher pay ranks — it strengthens monotonically with seniority. For a given doubling of firm productivity, the top-paid employee (likely the CEO) sees approximately 15 percent more pay, while the median-paid employee sees approximately 7 percent more. Equivalently, the pay-productivity elasticity is 0.15 for the top earner and 0.07 for the median earner. At the percentile level, a 10 percent productivity increase predicts a 0.86 percent pay increase at the 90th percentile but only 0.53 percent at the 10th percentile. Consequently, more productive firms have higher within-firm inequality: a 10 percent productivity increase widens the top-earner-to-median-worker log pay gap by 0.9 percent, and moving from the 10th to the 90th percentile of productivity projects a 23.1 percent increase in this gap. These cross-sectional results survive firm fixed effects, demographic controls (sex, education, age), industry fixed effects at the 6-digit NAICS level, and 2SLS instrumentation with industry exposures to seven major currencies, oil prices, and economic policy uncertainty (Alfaro, Bloom, and Lin 2024). Within-worker, within-firm estimates confirm the pattern dynamically: when a firm&amp;rsquo;s productivity doubles, workers earning $45,000–$65,000 expect roughly a 1 percent pay increase while workers earning above $300,000 expect nearly a 2 percent increase. The pay-productivity relationship is roughly twice as strong for top earners at publicly traded firms as at private firms (coefficient of 0.22 vs. 0.13 for rank-1 earners), while workers outside the top 50 ranks show similar coefficients across ownership types.&lt;/p&gt;
&lt;p&gt;Third, the mechanism is traced to performance-based pay. More productive firms exhibit higher within-year pay volatility (measured as the standard deviation of quarterly log earnings within a year), particularly for top earners, consistent with larger bonus payments. Firms with higher structured management scores — capturing more intensive performance monitoring, goal-setting, and incentive pay — also show higher pay levels and higher pay volatility for top earners, with the gradient across ranks matching the productivity results.&lt;/p&gt;
&lt;p&gt;Finally, a back-of-the-envelope calculation applies the estimated pay-productivity elasticities to observed aggregate productivity growth. Aggregate U.S. labor productivity roughly doubled (96 percent compounded growth) from 1980 to 2013. The top-earner-to-median-worker pay ratio at firms with at least 100 employees rose from 7.55 in 1980 to 8.69 in 2013 (an increase of 1.14). Applying the paper&amp;rsquo;s elasticities for rank-1 (0.1534) and rank-50 (0.0657) earners to the observed productivity doubling predicts a ratio of 8.01 in 2013 — accounting for 40 percent of the actual increase. The authors interpret this as evidence that rising productivity, channeled through differential performance pay, is a quantitatively important driver of rising within-firm inequality.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-primary-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q1. What is the primary identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;The core cross-sectional estimates in models (1) and (2) regress percentile- or rank-specific pay on log revenue per worker, controlling for a quadratic expansion of firm-level worker demographic composition (sex, education, age and their interactions), year fixed effects, and 6-digit NAICS industry fixed effects. The main threat is omitted variable bias: unobserved firm characteristics correlated with both productivity and pay (e.g., high-skill worker sorting into high-productivity firms) could inflate estimates. The paper addresses this in three ways. First, specifications with firm fixed effects (Appendix Figure A.1) use only within-firm changes in productivity and pay, producing similar convex-across-ranks patterns. Second, the within-worker, within-firm change specification (model 4, Figure 2) holds individual workers fixed and relates earnings growth to productivity growth. Third, a 2SLS approach instruments log productivity (and its interaction with rank) using industry-level exposures to seven currency pairs, oil prices, and economic policy uncertainty constructed from rolling 10-year daily stock-return regressions by Alfaro, Bloom, and Lin (2024); the logic is that industries have idiosyncratic exposure to these aggregate shocks, so productivity movements attributable to the instruments are exogenous to individual pay-setting. The 2SLS results are broadly similar to OLS in sign and pattern, though first-stage F-statistics are approximately 3, which is weak by conventional standards. Additional tests using lagged productivity (Appendix Table A.3) show if anything stronger relationships, consistent with productivity causally passing through to pay rather than pay determining past productivity.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-main-mechanisms-proposed-and-how-are-they-distinguished-empirically"&gt;Q2. What are the main mechanisms proposed and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;The primary mechanism proposed is performance-based pay (bonuses and incentive compensation) that is disproportionately concentrated among senior managers at more productive firms. The paper cannot directly observe bonus pay in the LEHD, which reports total quarterly earnings. Instead, it uses within-year pay volatility — the standard deviation of log quarterly earnings within a calendar year — as a proxy for bonus income (most visibly fourth-quarter bonus payments). Figure 4 shows that top earners at more productive firms have significantly higher pay volatility, and this relationship is steeper at higher ranks, exactly paralleling the pay-level results. The management channel is examined separately: Figure 5 shows that firms with higher MOPS structured management scores (capturing explicit monitoring, target-setting, and incentive-pay practices) display higher pay levels and higher pay volatility for top earners, again with the gradient increasing at the top. The public-vs.-private ownership comparison is a further diagnostic: if performance-based executive compensation is the mechanism, it should be stronger at publicly traded firms, where stock grants, option awards, and formal incentive contracts are more prevalent. Panel a of Figure 3 confirms the top-earner pay-productivity coefficient is 0.22 at public firms and 0.13 at private firms, while workers outside the top 50 show similar coefficients across ownership type. This asymmetry is robust to reweighting public firms to match the employment distribution of private firms (panel b of Figure 3), ruling out pure size effects as the explanation.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented-across-sectors-firm-age-and-ownership-type"&gt;Q3. What heterogeneity is documented across sectors, firm age, and ownership type?&lt;/h3&gt;
&lt;p&gt;Across sectors (Appendix Figure A.2), the positive and convex pay-productivity gradient across earnings ranks is present in nearly all 18 two-digit NAICS sectors. Shallower (less convex) patterns appear in utilities, finance and insurance, and health, which the authors attribute to heavy regulation limiting scope for differential performance pay across ranks. Across firm age groups (Appendix Figure A.3), the pattern holds across firms younger than 10 years, between 10 and 25 years, and 25 or more years. Across ownership, the pay-productivity relationship for top earners is roughly twice as large in publicly traded firms as in privately held firms, while the relationship for workers outside the top 50 is similar. Within publicly traded firms, the LEHD top-earner coefficients closely match those for named executives in the Compustat Execucomp data (Figure 3, panel a), validating both the LEHD measure of top earnings and the Execucomp-based executive pay literature.&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;The paper runs the following robustness checks: (1) Full demographic controls — a quadratic expansion of firm-level shares by sex, education category, and age group, plus interactions — included in all baseline regressions to account for worker sorting. (2) 6-digit NAICS industry fixed effects to net out cross-industry pay and productivity variation. (3) Firm fixed effects (Appendix Figure A.1): the convex pattern across ranks survives when only within-firm variation in productivity and pay is used. (4) Sector heterogeneity analysis (Appendix Figure A.2): the main pattern holds across nearly all 18 two-digit NAICS sectors. (5) Firm age heterogeneity (Appendix Figure A.3): results hold across all age groups. (6) Reweighting public firms to match private firms&amp;rsquo; employment distribution (Figure 3, panel b): the stronger pay-productivity gradient for top earners at public firms is not explained by their greater average size. (7) Size controls: including log total LEHD employment does not eliminate the pattern. (8) 2SLS with macroeconomic instruments: similar signs and pattern to OLS, supporting causal interpretation despite weak first stages. (9) Lagged productivity (Appendix Table A.3): if anything, the pay-productivity relationship by rank is slightly stronger when using prior-year productivity, reducing reverse-causality concerns. (10) Comparison to Execucomp: the LEHD public-firm top-earner coefficients align with those from Execucomp named executives. (11) Analysis of sandwich-worker selection (Appendix Table A.1): workers at more productive firms are marginally more likely to remain sandwich workers the following year, with this pattern slightly stronger at lower earnings ranks; the paper discusses this selection and argues it does not drive the main results.&lt;/p&gt;
&lt;h3 id="q5-what-exactly-is-the-lehd-earnings-measure-and-how-does-it-capture-bonuses"&gt;Q5. What exactly is the LEHD earnings measure and how does it capture bonuses?&lt;/h3&gt;
&lt;p&gt;The LEHD is based on state unemployment insurance (UI) wage records submitted by employers. It captures total quarterly earnings, including salaries, wages, bonuses, stock option exercises, and restricted stock awards when vested. Qualified (incentive) stock options are not subject to UI tax and are excluded, but these are capped and the paper judges them immaterial for top earners. The quarterly frequency of the data allows the paper to construct within-year pay volatility (the standard deviation of log quarterly earnings in a year) as a proxy for bonus income, since bonus payments typically appear as spikes in Q4. The paper uses only non-imputed demographic characteristics from ancillary LEHD sources; imputed values (e.g., education, which is imputed for 88 percent of individuals) are replaced with a constant and flagged with a missing-value indicator.&lt;/p&gt;
&lt;h3 id="q6-how-exactly-is-firm-productivity-measured-and-what-are-its-limitations"&gt;Q6. How exactly is firm productivity measured and what are its limitations?&lt;/h3&gt;
&lt;p&gt;Productivity is measured as real revenue per worker (log scale), with nominal revenue deflated to 2010 dollars using the PCE deflator. Revenue and employment come from the LBD, which covers all non-farm sectors from 1997 onward. This is a revenue-based labor productivity measure, not total factor productivity, and no industry-level price deflators are used beyond the economy-wide PCE; instead, 6-digit NAICS industry fixed effects control for cross-industry differences in revenue-per-worker levels. The LBD&amp;rsquo;s revenue coverage may be biased toward older, more stable firms, but the paper argues this has minimal impact because its sample is already restricted to large firms (at least 100 full-year workers). The paper explicitly contrasts its broad economy-wide measure with more granular TFP measures available only for manufacturing and in Economic Census years.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-structured-management-score-and-what-does-it-measure"&gt;Q7. What is the structured management score and what does it measure?&lt;/h3&gt;
&lt;p&gt;The structured management score is derived from 16 core questions in the MOPS asking plant managers about practices in three domains: performance monitoring, target setting, and incentivization of workers. Each question is scored 0 to 1, where 0 reflects least structured (less explicit, formal, frequent, or specific) and 1 reflects most structured (more explicit, formal, frequent, or specific). The firm-level score is an employment-weighted average of establishment-level scores (requiring at least 10 non-missing responses per establishment). It ranges from 0 to 1 and follows the methodology of Bloom et al. (2019), who establish that higher scores predict higher establishment-level productivity. Because MOPS targets manufacturing establishments surveyed in the ASM, the management sample is a 2.5 percent subset of the main sample, resulting in wider standard errors for management-related estimates. The paper treats this score as an indirect proxy for the adoption of performance-based incentive systems.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-song-et-al-2019-and-the-broader-between-firm-vs-within-firm-inequality-literature"&gt;Q8. How does this paper relate to and differ from Song et al. (2019) and the broader between-firm vs. within-firm inequality literature?&lt;/h3&gt;
&lt;p&gt;Song et al. (2019), also using linked LEHD-LBD data, document that the rise in U.S. earnings inequality between 1978 and 2013 was driven predominantly by increases in between-firm pay dispersion, with within-firm inequality rising more modestly. This paper takes the within-firm inequality result as a starting point and asks what firm characteristics predict cross-sectional and dynamic variation in within-firm inequality. The key addition is connecting within-firm pay dispersion to revenue labor productivity and to management practices, neither of which Song et al. (2019) directly analyze. The paper uses Song et al.&amp;rsquo;s published aggregate statistics on top-earner and median-earner pay (from their Figure VI) as the benchmark for the back-of-the-envelope calculation linking rising productivity to rising inequality. More broadly, the paper contributes to a cross-country literature (Barth et al. (2016), Card, Heining, and Kline (2013), Faggio, Salvanes, and Van Reenen (2010), Mueller, Ouimet, and Simintzi (2017)) that documents firms as the locus of increasing wage dispersion, by providing a specific firm-level mechanism — productivity and performance-pay practices.&lt;/p&gt;
&lt;h3 id="q9-how-does-this-paper-relate-to-and-differ-from-the-ceo-pay-literature"&gt;Q9. How does this paper relate to and differ from the CEO pay literature?&lt;/h3&gt;
&lt;p&gt;The CEO pay literature (Gabaix and Landier (2008), Frydman and Jenter (2010), Kaplan (2013), Edmans and Gabaix (2016)) debates whether rising CEO pay reflects performance, firm size, or rent extraction, but typically studies only the named top executives at large publicly traded firms covered by Execucomp. This paper&amp;rsquo;s key innovation is extending the analysis to all workers across the full within-firm pay distribution, for millions of U.S. workers at firms of all sizes and ownership types. It finds that the pay-productivity gradient is present across all earnings ranks, not only at the CEO level, though it is steeper at the top. The paper validates its LEHD-based top-earner results against Execucomp, finding close agreement for publicly traded firms, and interprets the public-vs.-private differential as consistent with formal performance-based executive contracts being more prevalent at public firms — a finding consistent with Gao and Li (2015), who show CEO pay-performance sensitivity is greater at public firms.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-aggregate-inequality-implications-and-how-robust-is-the-40-percent-estimate"&gt;Q10. What are the aggregate inequality implications and how robust is the 40 percent estimate?&lt;/h3&gt;
&lt;p&gt;The 40 percent figure comes from a back-of-the-envelope calculation in Table 4. Using Song et al.&amp;rsquo;s (2019) data, the top-earner-to-median-worker pay ratio rose from 7.55 in 1980 to 8.69 in 2013 (a change of 1.14). Aggregate U.S. labor productivity grew 96 percent compounded over this period (sourced from FRED series PRS85006092). The paper applies the pay-productivity elasticities for rank-1 (0.1534) and rank-50 (0.0657) earners from Figure 1 to this productivity growth to predict earnings levels in 2013. The predicted top-earner mean earnings is $224,357 (versus actual $301,614) and predicted median mean is $28,013 (versus actual $34,702), yielding a predicted ratio of 8.01 and an explained change of 0.46, which is 40.13 percent of the actual change of 1.14. The authors label this a &amp;lsquo;simple back-of-the-envelope&amp;rsquo; calculation and do not claim it as a structural decomposition. Key caveats: (i) the cross-sectional elasticities from 2003–2015 are applied to a 1980–2013 trend, assuming stability of these relationships over time; (ii) aggregate productivity growth may also shift the productivity distribution of firms, which the calculation does not fully model; (iii) the calculation attributes none of the remaining 60 percent, which could include technology, globalization, changing labor market institutions, or other forces.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-role-of-firm-size-in-explaining-the-results"&gt;Q11. What is the role of firm size in explaining the results?&lt;/h3&gt;
&lt;p&gt;Publicly traded firms in the sample are substantially larger than private firms on average (mean 7,763 versus 491.7 full-year employees). To ensure the stronger pay-productivity gradient at public firms is not simply a size artifact, the paper reweights public firms to match the employment distribution of private firms (using ventile-based inverse-probability weights) and finds the differential persists (panel b of Figure 3). The paper also includes log total LEHD employment as a control in additional specifications and reports similar results. The large-firm pay premium literature (Brown and Medoff (1989), Oi and Idson (1999)) posits that large firms pay more due to compensating differentials, monitoring difficulties, or rent-sharing. The paper&amp;rsquo;s finding that pay is higher at more productive firms across the entire earnings distribution is interpreted as more supportive of the rent-sharing explanation, since compensation-based and monitoring-based explanations would not apply uniformly to all workers.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The main policy-relevant implication is that rising productivity — itself associated with technology adoption and innovation — contributes substantially (estimated 40 percent) to the CEO-to-median-worker pay gap that the Dodd-Frank Act requires publicly traded firms to disclose annually from 2018. This implies that policies targeting within-firm pay inequality may need to grapple with the fact that a significant share of observed inequality is tied to real productivity differences and performance-pay practices, not purely to governance failures or rent extraction. However, several scope conditions limit this implication: the 40 percent figure is an economy-wide back-of-the-envelope estimate with caveats about stability of elasticities over time; the paper does not assess whether performance pay practices are optimally structured or reflect rent-seeking; the mechanism analysis uses pay volatility and management scores as proxies rather than direct observation of bonus contracts; and the remaining 60 percent of the inequality increase is left unaccounted for, potentially reflecting factors outside the paper&amp;rsquo;s framework.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-key-data-limitations-and-potential-measurement-concerns"&gt;Q13. What are the key data limitations and potential measurement concerns?&lt;/h3&gt;
&lt;p&gt;Several limitations are acknowledged or implicit. (1) Revenue labor productivity is used rather than TFP; the measure conflates product demand and productivity shocks and does not adjust for industry-specific output price variation. (2) LEHD earnings exclude qualified (incentive) stock options not subject to UI tax; the paper argues these are capped and immaterial for top earners, but this may understate total compensation for senior executives, especially at technology firms. (3) Within-year pay volatility is used as a proxy for bonus income rather than direct bonus data. (4) The management sample is confined to firms with at least one manufacturing establishment in the MOPS, covering only 2.5 percent of main-sample firm-year observations, limiting precision. (5) Education is imputed for 88 percent of individuals in the LEHD; the paper uses only non-imputed values and controls for missingness, but this reduces demographic control precision. (6) The IV first-stage F-statistics are approximately 3, suggesting weak instruments, so 2SLS standard errors are wide and the causal interpretation should be taken cautiously. (7) The sample is restricted to firms with at least 100 full-year workers, so results do not speak to smaller firms, which employ a large share of the U.S. workforce.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Revenue labor productivity&lt;/strong&gt;: Real revenue per worker at the firm level, computed from LBD annual revenue deflated to 2010 dollars using the PCE deflator and divided by total firm employment; the paper&amp;rsquo;s primary measure of firm performance, entered in log form in all regressions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pay-productivity elasticity (by rank)&lt;/strong&gt;: The regression coefficient on log firm productivity in a regression of mean log annual earnings for a given within-firm earnings rank or percentile; the paper documents that this elasticity rises monotonically from approximately 0.07 for the median earner to 0.15 for the top earner (rank 1), producing a convex schedule across ranks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Within-firm earnings inequality&lt;/strong&gt;: Dispersion in annual earnings among full-year workers within a single firm in a given year; measured variously as the 90th-10th percentile log earnings gap, the 99th-10th gap, the top-earner-to-50th-percentile gap, and the top-earner-to-10th-percentile gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Within-year pay volatility&lt;/strong&gt;: The standard deviation of log quarterly earnings within a calendar year for a given worker rank; used as a proxy for variable (bonus) compensation since it captures deviations from a constant salary path, particularly fourth-quarter bonus payments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structured management score (MOPS)&lt;/strong&gt;: A continuous index bounded between 0 and 1 derived from 16 MOPS survey questions on performance monitoring, target-setting, and worker incentivization practices; higher values indicate more explicit, formal, frequent, and specific management practices, following the scoring methodology of Bloom et al. (2019).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;6-quarter sandwich worker&lt;/strong&gt;: An individual who is employed at and earns above the minimum wage at the same firm in all four quarters of the current year, the fourth quarter of the prior year, and the first quarter of the following year; the restriction ensures that measured annual earnings reflect genuine full-year employment rather than partial-year spells or job transitions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DHS (Davis-Haltiwanger-Schuh) growth rate&lt;/strong&gt;: A symmetric growth rate measure defined as (x_t - x_{t-1}) / (0.5 * (x_t + x_{t-1})), bounded between -2 and 2; used in the within-worker, within-firm change analysis to measure both earnings growth and productivity growth while accommodating entry and exit.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Top-earner-to-median-worker pay ratio&lt;/strong&gt;: The ratio of mean annual earnings of the highest-paid worker to mean annual earnings of the median-paid worker within firms, aggregated across firms of different sizes using employment weights; the Dodd-Frank Act metric that publicly traded firms have been required to disclose annually since 2018, and the paper&amp;rsquo;s primary metric for the aggregate inequality calculation.&lt;/p&gt;</description></item><item><title>Zero-hours Contracts in a Frictional Labour Market</title><link>https://macropaperwarehouse.com/papers/zero-hours-contracts-in-a-frictional-labour-market/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/zero-hours-contracts-in-a-frictional-labour-market/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Dolado, Lalé, and Turon build a structural equilibrium model of the U.K. low-wage labour market to evaluate zero-hours contracts (ZHCs), employment agreements under which firms are not required to guarantee any minimum working hours and workers may decline any hours offered. The paper&amp;rsquo;s central question is whether ZHCs raise or lower welfare in general equilibrium, and through which channels. The model features two-sided heterogeneity in a random-search-and-matching environment: firms differ in the volatility of their labour demand, workers differ in their relative preferences for flexible versus regular employment, and wages are fixed at or near the statutory minimum wage. Three mechanisms operate simultaneously. First, a job-creation effect: firms facing highly volatile demand that cannot profitably hire under regular terms enter the market only because ZHCs exist. Second, a substitution effect: some firms that could hire under regular contracts instead post ZHC vacancies, crowding out regular employment. Third, a labour-force-participation effect: workers with a strong preference for flexible schedules join the labour force specifically because ZHCs exist and would withdraw if ZHCs were banned.&lt;/p&gt;
&lt;p&gt;The model is calibrated to U.K. Labour Force Survey data for the low-pay segment (roughly 16 percent of total employment), covering September 2018 through March 2020, with a sample of 9,342 individuals aged 16 to 69. A mixture-of-exponentials approach due to Karlis and Xekalaki (1999) applied to job-tenure and unemployment-duration distributions reveals statistically exactly two worker types in both ZHC employment and unemployment, and only one in regular employment, consistent with the presence of R-best workers (who prefer regular employment but accept ZHCs as a stepping stone) and Z-only workers (who would exit the labour force without ZHCs) but not R-only or Z-best workers. Calibrated parameters include a biweekly job-finding rate of λ(θ) = 0.051, a job-destruction probability of δ = 0.005, an on-the-job search efficiency of x = 0.352, and a share of R-best workers of ζ_{R-best} = 0.969. The matching function elasticity ψ is estimated to be 0.65 from U.K. occupation-level hiring and vacancy data (range 0.60–0.70 across specifications). ZHC employment accounts for 6.5 percent of the low-wage employment stock but 19.4 percent of vacancies, because higher turnover in ZHC jobs causes them to be re-advertised more frequently.&lt;/p&gt;
&lt;p&gt;A ban on ZHCs — simulated as an extreme tightening of flexible-work regulation — raises the unemployment rate by 2.0 to 2.7 percentage points depending on the assumed volatility of ZHC firms&amp;rsquo; demand. When ZHC workers have a low enough disutility of labour that they remain in the workforce after a ban (accepting regular jobs instead), the employment rate falls by the same 2.0 to 2.7 p.p., and sectoral GDP falls by only 0.02 to 0.14 percent, because higher average hours per employed worker partially offset the employment decline. When ZHC workers&amp;rsquo; disutility is high enough that they withdraw from the labour force, the employment-rate fall is larger — 4.8 to 5.4 p.p. — and sectoral GDP falls by 2.9 to 3.2 percent. Decomposing via the model&amp;rsquo;s analytical formula (Proposition 4a), lower job creation alone would reduce regular employment by almost 30 percent in isolation (λ(tilde-θ)/λ(θ) = 0.71), but this is partially offset by reduced vacancy competition (+24 percent, ceteris paribus) and improved search efficiency for regular jobs (+15 percent, ceteris paribus) after the ban.&lt;/p&gt;
&lt;p&gt;Welfare effects are measured in consumption-equivalent variation units. In general equilibrium, R-best workers (those who prefer regular jobs but sometimes hold ZHCs as a stepping stone) suffer welfare losses of −0.5 to −0.6 percent of consumption from a ZHC ban, driven primarily by longer expected unemployment spells. Yet in a partial equilibrium experiment that converts their ZHC jobs to regular jobs while holding all other equilibrium objects fixed, these same workers gain approximately +0.2 percent: the substitution effect is genuinely welfare-improving for them in isolation, but the job-creation channel dominates in general equilibrium and more than reverses that gain. Z-only workers — those who would exit the labour force if ZHCs were banned — suffer general-equilibrium welfare losses of −1.7 to −2.0 percent (low-disutility scenario) or approximately −1.8 to −2.1 percent (high-disutility scenario). These losses exceed the losses to R-best workers because Z-only workers are also forced into a type of employment they strictly prefer to avoid. The paper concludes that a ZHC ban is welfare-reducing for all workers in general equilibrium, and proposes that policy instead target ZHC use toward matches where workers voluntarily choose flexibility (Recommendation P1) and toward small firms that cannot diversify demand volatility across many positions (Recommendation P2).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-models-core-structure-and-what-frictions-drive-the-results"&gt;Q1. What is the model&amp;rsquo;s core structure and what frictions drive the results?&lt;/h3&gt;
&lt;p&gt;The model is a discrete-time steady-state random-search-and-matching model with two-sided heterogeneity. Workers are heterogeneous in their flow payoffs from regular employment (ω^i_R), flexible ZHC employment (ω^i_Z), and non-employment (ω^i_N), with these payoffs shaped by CRRA utility over consumption and a type-specific disutility of hours worked (α^i). Firms are heterogeneous in the volatility of their demand shock (σ_j), which determines the expected profit flow under each contract type. Flow profits depend on how actual hours h deviate from a stochastic target h-tilde via a quadratic loss specification. Market tightness θ is determined endogenously by free entry. The key friction is random search: workers cannot direct their search to their preferred contract type, so R-best workers sometimes end up in ZHCs and must search on-the-job to move to regular employment.&lt;/p&gt;
&lt;h3 id="q2-how-are-worker-types-identified-empirically-and-why-only-two-types"&gt;Q2. How are worker types identified empirically, and why only two types?&lt;/h3&gt;
&lt;p&gt;The paper adapts a mixture-of-exponential distributions procedure from Karlis and Xekalaki (1999), applied separately to the duration distribution of ZHC employment, regular employment, and unemployment in LFS data. A bootstrapped sequential hypothesis test determines the number of latent classes M* that best fits the survival function. For ZHC employment, two exponential components are needed (p-value for M=1 vs. M≥2 is 0.01; for M=2 vs. M≥3 it is 0.74). For regular employment, one component suffices (p-value for M=1 vs. M≥2 is 0.99). For unemployment, again two components (p-values 0.01 and 0.93 respectively). Cross-referencing which types are present in which states using the model&amp;rsquo;s theoretical exit-rate table rules out R-only and Z-best workers, leaving only R-best and Z-only workers as consistent with all three distributions simultaneously.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-identification-strategy-and-what-are-the-main-threats-to-it"&gt;Q3. What is the identification strategy and what are the main threats to it?&lt;/h3&gt;
&lt;p&gt;Identification rests on three steps. First, the mixture-of-exponentials procedure identifies the number of worker types from shape of duration distributions; this step relies on recalled job tenure and unemployment duration, which the authors acknowledge may suffer from recall bias and heaping (rounding to salient durations). Second, the turnover parameters are calibrated by minimizing distance between model-implied and empirical transition matrices across U, Z, and R states from the longitudinal LFS; the main limitation noted is that the two moments (transitions and durations) are not jointly consistent because they come from different measurement processes. Third, flow profits and payoffs are calibrated to external moments (minimum wage, replacement rate, business creation costs) and the preference for ZHC hours; the hours volatility parameter σ_Z has no direct empirical counterpart and is varied across scenarios. The model abstracts from wage bargaining, treating wages as fixed at the minimum wage, which reduces scope for confounding but is an approximation even in the low-wage sector.&lt;/p&gt;
&lt;h3 id="q4-how-are-the-three-channels--job-creation-substitution-and-labour-force-participation--distinguished-in-the-quantitative-analysis"&gt;Q4. How are the three channels — job creation, substitution, and labour-force participation — distinguished in the quantitative analysis?&lt;/h3&gt;
&lt;p&gt;The job-creation channel is captured by Z-only firms (firms with σ_Z = 6 such that regular employment is not profitable): removing ZHCs forces these firms out of the market entirely, reducing labour market tightness θ and hence the aggregate job-finding rate λ(θ). The substitution channel is captured by Z-best firms (σ_Z = 3): these firms could profitably hire under regular contracts but choose ZHCs, and after a ban they convert vacancies to regular posts, with incomplete crowd-out due to general equilibrium adjustment. The labour-force-participation channel is captured by Z-only workers: those with disutility α^i above the threshold (WTP &amp;gt; £7.9 per week to avoid regular work) withdraw from the labour force when ZHCs are banned, while those below the threshold remain and take regular jobs. The paper runs scenarios that vary both the firm side (low vs. high volatility) and the worker side (low vs. high disutility) to disentangle the magnitude of each channel.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-decomposition-of-the-effect-on-regular-employment-proposition-4a"&gt;Q5. What is the decomposition of the effect on regular employment (Proposition 4a)?&lt;/h3&gt;
&lt;p&gt;Under the calibrated parameters (no Z-best workers), regular employment in the baseline relative to the no-ZHC counterfactual equals the product of three multiplicative terms. The job-creation term is λ(θ)/λ(tilde-θ) = 1/0.71 ≈ 1.41, meaning that ZHCs raise the job-finding rate by about 41 percent relative to the no-ZHC counterfactual. The vacancy-competition term vR/v ≈ 0.81 (80.6 percent of vacancies are for regular jobs, while the remaining 19.4 percent for ZHC jobs dilute the pool). The search-efficiency term captures the fact that some R-best workers are in ZHC employment and search on-the-job at reduced intensity x &amp;lt; 1. The ceteris paribus decomposition at the ban scenario indicates: job creation alone would cut regular employment by 29 percent; competition reduction adds 24 percent; and search-efficiency gains add 15 percent — so the post-ban equilibrium has higher regular employment despite worse job creation overall.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-handle-the-partial-versus-general-equilibrium-distinction-for-welfare"&gt;Q6. How does the paper handle the partial versus general equilibrium distinction for welfare?&lt;/h3&gt;
&lt;p&gt;For R-best workers, the PE experiment replaces their ZHC jobs with regular jobs while keeping all other equilibrium objects (tightness θ, vacancy composition, etc.) fixed. This isolates the substitution effect and yields a welfare gain of approximately +0.15 to +0.18 percent for R-best workers. In general equilibrium, the full ban requires θ to fall (less job creation), which extends unemployment spells, and the net welfare effect is −0.50 to −0.62 percent. The difference between GE and PE therefore quantifies the job-creation externality that ZHCs provide — approximately 0.65 to 0.80 percentage points of consumption equivalent variation for R-best workers. For Z-only workers, the PE experiment replaces ZHC jobs with non-employment (their next-best option in the baseline), yielding PE welfare changes of −2.94 to −3.28 percent, which overstates the GE loss (−1.65 to −2.0 percent) because GE adjustment allows some Z-only workers to take regular jobs, partially compensating for the loss of ZHC access.&lt;/p&gt;
&lt;h3 id="q7-what-heterogeneity-is-documented-in-the-data-for-uk-zhc-workers"&gt;Q7. What heterogeneity is documented in the data for U.K. ZHC workers?&lt;/h3&gt;
&lt;p&gt;ZHC employment is concentrated at both ends of the age distribution: workers aged 16–29 are over-represented, as are workers aged 55–69, relative to regular employment. Mean age is 40.8 years for ZHC workers vs. 46.3 for regular workers. Gender composition is similar: 56.5 percent female in ZHCs vs. 60.4 percent female in regular employment, a difference that is not statistically significant. Educational attainment distributions are similar: 21.9 percent of ZHC workers hold a degree vs. 18.0 percent of regular workers. By industry, ZHC employment is heavily concentrated in Accommodation and food services (19.9 percent), Health and social work (20.5 percent), and Arts, entertainment and recreation (6.7 percent). Average hours worked are 18.4 per week for continuously employed ZHC workers vs. 28.1 for regular contract workers; the standard deviation of hours is 7.8 vs. 7.2. 16.6 percent of ZHC workers report wanting more hours vs. 10.1 percent in regular contracts, and 18.2 percent of ZHC workers are looking for another/additional job vs. 5.0 percent of regular workers, suggesting a minority are in involuntary underemployment while a majority are not actively seeking to change.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-key-calibrated-parameter-values-and-how-do-they-compare-to-the-broader-literature"&gt;Q8. What are the key calibrated parameter values and how do they compare to the broader literature?&lt;/h3&gt;
&lt;p&gt;The biweekly job-finding rate λ(θ) = 0.051; the biweekly job-destruction probability δ = 0.005; on-the-job search efficiency x = 0.352 (authors note this is on the high end but consistent with estimates accounting for flexible work); share of R-best workers ζ_{R-best} = 0.969; share of type-R vacancy-posting firms γ_R = 0.950. The matching function elasticity ψ = 0.65 (estimated from U.K. data, range 0.60–0.70, higher than the commonly used 0.50 but consistent with bias-corrected estimates from Borowczyk-Martins et al. 2013). The job-filling rate is 0.21 per biweek, consistent with Kuhn et al. (2021) U.K. estimates of 0.35–0.38 per month. The vacancy posting cost κ = £36.3 per week and startup cost K = £4,376, the latter close to the £4,500 implied by U.K. business creation data. Non-employment income b = £148.8 per week (replacement ratio 80 percent). The minimum wage is set to £7.50 per hour (2017 U.K. National Living Wage); labour productivity p = £8.25, implying a 10 percent productivity premium over the minimum wage.&lt;/p&gt;
&lt;h3 id="q9-what-robustness-checks-are-run-and-do-the-main-results-change"&gt;Q9. What robustness checks are run, and do the main results change?&lt;/h3&gt;
&lt;p&gt;The authors run three main robustness analyses. First, they vary the hours parameters: an alternative calibration uses σ_Z = 4.5 for both firm types but differentiates by mean hours (µ_Z = 20 for Z-best, µ_Z = 16 for Z-only); employment and unemployment effects are modestly smaller than the baseline but welfare effects are nearly identical. Second, they hold µ_Z = 18 and vary σ_Z to 1.0 (low) and 8.0 (high); results move in the expected direction and remain broadly consistent. Third, they vary the targeted job-filling rate: at λ(θ)/θ = 0.16 (25 percent lower than baseline), the unemployment response to a ZHC ban is only 0.33–0.51 p.p. and GDP effects are positive in the low-disutility case; at λ(θ)/θ = 0.26 (25 percent higher), unemployment rises by 4.1–5.5 p.p. and sectoral GDP falls by up to 6 percent. The authors conclude that the baseline calibration of 0.21 is the most plausible. The qualitative conclusions — that GE welfare effects are negative for all workers — are robust across specifications.&lt;/p&gt;
&lt;h3 id="q10-how-does-this-paper-relate-to-and-differ-from-closely-related-prior-work"&gt;Q10. How does this paper relate to and differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;The closest model-based study is Scarfe (2019) on casual work in Australia. Scarfe&amp;rsquo;s model features homogeneous agents ex ante, with contract choice driven by luck (stochastic match productivity), while Dolado et al. emphasise ex ante heterogeneity in preferences/profitability as the primary source of variation. The empirical study of Datta et al. (2019) documents U.K. ZHC characteristics using LFS, online survey, and matched employer-employee data from the social care sector; Dolado et al. use the LFS but impose structural discipline to recover preference parameters and conduct GE welfare analysis. The paper differs from the dual labour market literature (Cahuc et al. 2016, 2020; Créchet 2022) in that temporary jobs in that literature have a fixed expiration date, whereas ZHCs are jobs with potentially long tenure but endogenously lower expected duration due to on-the-job search quit-outs, not contractual termination. Mas and Pallais (2017) and Angelici and Profeta (2020) use field experiments to estimate workers&amp;rsquo; valuation of flexibility; Dolado et al. instead recover this from duration distributions, allowing for general equilibrium job-creation and participation effects that field experiments cannot capture.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-sorting-patterns-in-the-equilibrium-and-what-sustains-zhc-jobs"&gt;Q11. What are the sorting patterns in the equilibrium, and what sustains ZHC jobs?&lt;/h3&gt;
&lt;p&gt;In the baseline equilibrium, 66.8 percent of filled ZHC jobs are held by R-best workers (workers who prefer regular employment but accept ZHCs as a stepping stone). Only 4.8 percent of employed R-best workers are in ZHCs at any point in time, because most vacancies are for regular jobs (80.6 percent of vacancies). This sorting has a crucial implication: ZHC vacancies would not be viable without the presence of R-best workers, because Z-only workers alone are too few to sustain the ZHC sector in equilibrium. A firm posting a ZHC vacancy accepts a higher worker-turnover risk (R-best workers quit on-the-job once they find a regular vacancy) in exchange for the profit advantage of hours flexibility; the trade-off is viable only because the random search pool contains enough R-best workers willing to take ZHC jobs temporarily.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q12. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;The paper identifies four recommendations. P1: restrict ZHCs to matches where the worker voluntarily chooses the flexible contract when offered a choice; this would protect R-best workers who currently end up in ZHCs due to search frictions from the substitution effect without eliminating the job-creation channel. P2: prioritise access to ZHCs for small firms (as a proxy for inability to diversify demand shocks), limiting substitution by large firms while preserving genuine job creation by high-volatility operators. P3: recognise that the allocation of hours-flexibility between firms and workers is often an implicit and incomplete contract rather than an explicit one. P4: regulate the sharing of hours flexibility — specifically, who controls the timing and quantity of work — to reduce the income uncertainty that generates the main political objections to ZHCs. The scope conditions for all recommendations are: the low-wage sector of the U.K. labour market; the results do not directly apply to higher-wage workers with more bargaining power, or to markets where exclusivity clauses remain common.&lt;/p&gt;
&lt;h3 id="q13-what-key-empirical-facts-about-zhc-flows-does-the-paper-document"&gt;Q13. What key empirical facts about ZHC flows does the paper document?&lt;/h3&gt;
&lt;p&gt;From the transition matrix estimated from LFS data: 11 percent of exits from unemployment are to ZHC employment. The rate of transition to unemployment is almost 50 percent larger in ZHC employment than in regular employment (6.2 percent vs. 4.4 percent semi-annually). Job-to-job transitions from ZHC to regular employment are 6.5 percent semi-annually; the reverse (regular to ZHC) is only 0.5 percent. Nearly half of ZHC workers report job tenures longer than two years. 9.2 percent of ZHC workers were recruited in the last three months vs. 3.4 percent of regular workers; 30.3 percent of ZHC workers have been with their employer less than one year vs. 14.3 percent in regular contracts. The non-employment rate for this low-pay segment is 11.2 percent; ZHCs account for 4.6 percent of the overall sample (5.2 percent of employees), about 1.5 times the aggregate U.K. incidence rate.&lt;/p&gt;
&lt;h3 id="q14-what-does-the-model-say-about-time-spent-out-of-regular-employment-following-a-zhc-ban"&gt;Q14. What does the model say about time spent out of regular employment following a ZHC ban?&lt;/h3&gt;
&lt;p&gt;Despite higher aggregate unemployment rates after the ban, R-best workers spend less total time out of regular employment: the duration of non-regular-employment spells decreases by 7 weeks. This is because ZHCs, by acting as a stepping stone, expose workers to more frequent labour market transitions — they cycle through unemployment, ZHC employment, and regular employment rather than simply unemployment and regular employment. The ban removes the ZHC stepping stone, so workers face longer individual unemployment spells but avoid the ZHC-employment phase, and on net spend more time in regular employment. However, this does not translate into a welfare gain because (a) ZHC employment, even if imperfect, provides utility above the unemployment level, and (b) the longer unemployment spells that do occur under a ban are more costly than the shorter ZHC spells they replace.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Zero-hours contract (ZHC)&lt;/strong&gt;: In the paper&amp;rsquo;s sense, an employment arrangement under which the employer is not obligated to provide any minimum guaranteed hours of paid work, and the worker is not required to accept any hours offered. Workers on ZHCs in the U.K. hold &amp;lsquo;worker&amp;rsquo; status (between employee and self-employed), entitling them to holiday pay, minimum wage protections, and Universal Credit, but not redundancy pay. The key feature for the model is that actual hours worked equal the firm&amp;rsquo;s demand realisation, eliminating the quadratic deviation costs that arise under fixed-hours regular contracts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;R-best workers&lt;/strong&gt;: In the paper&amp;rsquo;s worker taxonomy, individuals for whom the asset value of regular employment strictly exceeds that of ZHC employment, which in turn exceeds the asset value of non-employment (W^i_R &amp;gt; W^i_Z &amp;gt; N^i). These workers accept ZHCs as a stepping stone when regular jobs are unavailable, and search on-the-job (at reduced efficiency x) for regular vacancies. They constitute 96.9 percent of the low-wage sector in the calibration and account for two-thirds of filled ZHC jobs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Z-only workers&lt;/strong&gt;: Workers for whom the asset value of ZHC employment exceeds both the value of regular employment and non-employment (W^i_Z &amp;gt; N^i &amp;gt; W^i_R, or W^i_Z &amp;gt; W^i_R &amp;gt; N^i), and who prefer non-employment to regular work. Without ZHCs, these workers&amp;rsquo; participation in the labour market depends on whether their disutility parameter α^i implies ω^i_R &amp;gt; ω^i_N. A subset — those with high disutility (WTP &amp;gt; £7.9 per week to avoid regular work) — exit the labour force if ZHCs are banned, generating the participation effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Z-only firms&lt;/strong&gt;: In the paper&amp;rsquo;s firm taxonomy, firms with high demand volatility (σ_Z = 6 in the calibration) for which regular employment is not profitable (V^j_R &amp;lt; 0 &amp;lt; V^j_Z). These firms can only operate and post vacancies because ZHCs allow them to set actual hours equal to realised demand. A ban on ZHCs causes Z-only firms to exit entirely, generating the pure job-creation loss.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Z-best firms&lt;/strong&gt;: Firms with moderate demand volatility (σ_Z = 3 in the calibration) that could profitably post regular vacancies (V^j_R &amp;gt; 0) but prefer ZHC vacancies because the hours-flexibility profit advantage outweighs the higher quit risk from R-best workers. A ban redirects these firms to regular contracts, constituting the substitution effect on the firm side.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stepping-stone effect&lt;/strong&gt;: The mechanism by which R-best workers accept ZHC employment when unemployed, using it as a bridge to search on-the-job for regular employment. ZHCs therefore simultaneously reduce unemployment duration and extend the time workers spend out of regular employment. The paper documents that a ZHC ban reduces total time out of regular employment by 7 weeks for R-best workers despite raising the unemployment rate, precisely because the stepping-stone pathway — which adds a ZHC phase before reaching regular employment — is eliminated.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption equivalent variation (welfare measure)&lt;/strong&gt;: The percentage permanent change in consumption that would make a worker indifferent between the baseline equilibrium (with ZHCs) and the counterfactual (ZHC ban). The paper uses this metric to express welfare effects: R-best workers suffer losses of −0.50 to −0.62 percent, and Z-only workers suffer losses of −1.65 to −2.0 percent, in general equilibrium following a ZHC ban.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mixture-of-exponentials identification of worker types&lt;/strong&gt;: A statistical procedure adapted from Karlis and Xekalaki (1999) that fits the empirical distribution of job tenure or unemployment duration as a mixture of M exponential distributions. Each component corresponds to a latent class of workers exiting the labour market state at a distinct rate. The optimal number of components M* is chosen via a bootstrapped sequential hypothesis test. Applied to U.K. LFS data, the procedure identifies M* = 2 for ZHC employment and unemployment, and M* = 1 for regular employment, which the model interprets as evidence for R-best and Z-only worker types.&lt;/p&gt;</description></item><item><title>CBDC as Imperfect Substitute to Bank Deposits: A Macroeconomic Perspective</title><link>https://macropaperwarehouse.com/papers/cbdc-as-imperfect-substitute-to-bank-deposits-a-macroeconomic-perspective/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/cbdc-as-imperfect-substitute-to-bank-deposits-a-macroeconomic-perspective/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: As central banks worldwide explore retail central bank digital currency (CBDC), the macroeconomic consequences depend heavily on how CBDC interacts with bank deposits. Prior work spans a wide range of conclusions — from &amp;ldquo;no effect&amp;rdquo; (Brunnermeier and Niepelt 2019) to disintermediation that reduces lending and output (Keister and Sanches 2022; Chiu et al. 2022) to large output gains (Barrdear and Kumhof 2021, +3% GDP). Bacchetta and Perazzi argue these differences hinge on (i) how substitutable CBDC is with checking deposits, (ii) how easily banks replace lost deposits with other funding, (iii) the interest rate on CBDC, and (iv) the competitive structure of banking. The paper provides quantitative welfare estimates in a model where CBDC and deposits are imperfect substitutes and banks are in monopolistic competition.&lt;/p&gt;
&lt;p&gt;Model setup: A closed-economy steady-state model (akin to Gali 2015 and Del Negro-Sims 2015) with households, &amp;ldquo;bank owners,&amp;rdquo; firms, banks, government, and central bank. Money reduces a transaction cost on consumption (Schmitt-Grohe-Uribe 2004 style). Deposits and CBDC combine via a CES composite liquid asset characterized by three CBDC design dimensions: its interest rate (rc), its relative liquidity (alpha_c/alpha_b, the CES weight), and its substitutability with deposits (elasticity epsilon_cb). Crucially, with monopolistic competition each bank takes the average deposit rate as given, so the equilibrium deposit rate is unaffected by CBDC (Lemma 1); and because firms can fund at the risk-free rate, bank credit extension and loan rates are also unaffected by CBDC in steady state. Calibration (US-based): risk-free rate 4%, deposit spread 2%, loan spread 1%, reserve ratio 5%, deposit management cost 25 bps, interest semi-elasticity of money demand -0.05, inverse Frisch elasticity gamma=1, wealth/consumption=4. The two extreme ownership cases are zeta=1 (&amp;ldquo;case a,&amp;rdquo; households fully own banks) and zeta=0 (&amp;ldquo;case b,&amp;rdquo; a zero-measure set of bankers receives all profits).&lt;/p&gt;
&lt;p&gt;Main findings (welfare in consumption-equivalent basis points): Welfare can improve via three channels — (1) seigniorage allowing lower distortionary labor taxes, (2) a lower opportunity cost of holding money (raising money holdings, cutting transaction costs, stimulating labor and consumption), and (3) redistribution of bank deposit rents from bankers to the general population. The optimal CBDC rate trades off seigniorage versus opportunity-cost reduction and is decreasing in the labor tax rate and decreasing in the share of banks owned by households (Proposition 3). The first two channels alone yield only modest gains: +9 bps at a 25% labor tax and +20 bps at 45%. Adding the redistribution channel (&amp;ldquo;case b&amp;rdquo;) raises non-bankers&amp;rsquo; welfare to +54 bps (25% tax) and +59 bps (45% tax); the headline maximum is about 60 bps. From Table 2 (epsilon_cb=20, equal liquidity): consumption rises +27 bps (case a) / +54 bps (case b) at 25% tax, and +41 / +62 bps at 45% tax. All benefits require historically normal interest rates (baseline 4%); near the zero lower bound seigniorage, money&amp;rsquo;s opportunity cost, and deposit rents all vanish, so the welfare gain falls roughly linearly to zero with the deposit spread.&lt;/p&gt;
&lt;p&gt;Policy/theoretical implications: CBDC is a tool to mitigate two distortions — distortionary taxation and the gap between the opportunity cost and the (low) production cost of money — plus a redistributive lever against the concentration of bank rents. The pure efficiency gains are modest; the larger gains come from redistribution and are larger where labor taxes (e.g., EU-14 averaging &amp;gt;40% vs. US ~25%), the Frisch elasticity, or the interest semi-elasticity of money demand are higher.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-models-identificationderivation-strategy-since-this-is-a-theoretical-paper-rather-than-an-empirical-one"&gt;Q1. What is the model&amp;rsquo;s identification/derivation strategy, since this is a theoretical paper rather than an empirical one?&lt;/h3&gt;
&lt;p&gt;There is no econometric identification; results come from a calibrated closed-economy steady-state general equilibrium model. The &amp;lsquo;identification&amp;rsquo; of the welfare channels is analytical: three propositions (proved in an online appendix) characterize how seigniorage and the optimal CBDC rate depend on CBDC liquidity (alpha_c), substitutability (epsilon_cb), and the labor tax rate, and numerical experiments on a US-calibrated economy quantify the welfare changes. The key structural assumption enabling the results is monopolistic competition in banking plus a financial-market funding alternative for banks at the risk-free rate.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-introduction-of-cbdc-leave-the-deposit-rate-and-bank-lending-unchanged-in-this-model"&gt;Q2. Why does the introduction of CBDC leave the deposit rate and bank lending unchanged in this model?&lt;/h3&gt;
&lt;p&gt;Lemma 1: under monopolistic competition each individual bank takes the aggregate deposit rate as given and does not internalize how aggregate deposit demand shifts with CBDC, so its optimal deposit rate (eq. 30) is invariant to CBDC&amp;rsquo;s interest rate or liquidity. CBDC lowers aggregate deposit demand, so banks simply rely more on other liabilities (bonds/equity). Lending is unaffected because the marginal cost of bank funding remains the risk-free rate (banks can borrow from the market), so the loan rate (eq. 32) and quantity of loans do not change. This contrasts with monopoly/Cournot banking (Andolfatto 2021; Chiu et al. 2022) where CBDC moves the deposit rate.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-three-welfare-channels-and-how-is-each-maximized"&gt;Q3. What are the three welfare channels and how is each maximized?&lt;/h3&gt;
&lt;p&gt;(1) Seigniorage: higher central-bank seigniorage finances lower distortionary labor taxes; maximized by setting rc to raise seigniorage revenue (peak occurs at rc &amp;lt; rb in the cases analyzed). (2) Opportunity cost of money: paying high interest on CBDC raises money holdings and cuts the transaction cost, stimulating labor and consumption; maximized by setting rc equal to the risk-free rate so households drop deposits entirely and drive the transaction cost toward zero. (3) Redistribution: CBDC lets non-bankers capture deposit rents previously held by bankers (via tax cuts or interest on CBDC), maximal when zeta=0 and rc near the risk-free rate. Channels (1) and (2) conflict, generating the optimal-rate tradeoff.&lt;/p&gt;
&lt;h3 id="q4-what-does-seigniorage-look-like-as-a-function-of-the-cbdc-rate-and-what-do-propositions-1-2-say"&gt;Q4. What does seigniorage look like as a function of the CBDC rate, and what do Propositions 1-2 say?&lt;/h3&gt;
&lt;p&gt;Seigniorage is non-monotonic in rc: a higher rc lowers seigniorage per unit of CBDC but raises CBDC demand. Proposition 1 (under alpha_b^{epsilon_cb}*epsilon_cb &amp;gt; 1 and negligible CBDC management cost): the seigniorage-maximizing rc exceeds the deposit rate rb; if epsilon_cb&amp;gt;1.5 the optimal rc decreases in CBDC liquidity alpha_c; and the peak seigniorage rises with both alpha_c and epsilon_cb. Proposition 2: within that parameter region, maximum seigniorage is achieved as epsilon_cb to infinity (perfect substitutes) with rc set infinitesimally above rb — i.e., outcompete deposits. In the numerical cases shown, the seigniorage peak occurs at rc &amp;lt; rb, moving closer to rb as CBDC liquidity rises.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity--cross-country-variation-does-the-paper-document"&gt;Q5. What heterogeneity / cross-country variation does the paper document?&lt;/h3&gt;
&lt;p&gt;Two dimensions. (i) Labor tax level: US ~25% vs EU-14 averaging &amp;gt;40% (Trabandt-Uhlig 2011). Higher taxes raise the value of the seigniorage/tax-cut channel, lower the optimal CBDC rate, and raise welfare gains (efficiency gains +9 bps at 25% to +20 bps at 45%). (ii) Bank ownership (zeta): &amp;lsquo;case a&amp;rsquo; (households own banks) gives small gains (7-8 bps at 20% tax to 18-20 bps at 45%); &amp;lsquo;case b&amp;rsquo; (bankers own banks) gives large gains (52-53 bps at 20% to 58-60 bps at 45%) via redistribution. The optimal CBDC rate is higher in case b than case a and rises with the tax rate (Proposition 3 / Figure 3).&lt;/p&gt;
&lt;h3 id="q6-what-robustness--alternative-parameter-checks-are-run-table-3"&gt;Q6. What robustness / alternative-parameter checks are run (Table 3)?&lt;/h3&gt;
&lt;p&gt;Frisch elasticity (gamma=0.25 i.e. Frisch=4, and gamma=4 i.e. Frisch=0.25): higher Frisch raises case-a gains (e.g., +28 bps at 25% tax) but case-b gains are roughly independent of Frisch. Interest semi-elasticity of money demand set to -0.12 (Benati et al. 2021 for Switzerland): with 45% taxes, gains reach +35 bps (case a) and +85 bps (case b) — this parameter has the biggest impact. Other variations with small effects: deposit/loan management costs, reserve ratio (0% vs 10%), bank-profit tax tau_b (15% vs 35%; lower tau_b means more inequality and larger CBDC gain), loan elasticity epsilon_l, working-capital share phi, wealth/consumption ratio (2 vs 4). Loan-side parameters and household wealth essentially do not matter because lending is unaffected by CBDC. With lump-sum (non-distortionary) taxes, case-a gains shrink (the seigniorage-tax channel is inactive) while case-b gains are essentially unchanged. At the zero lower bound the welfare gain is approximately linear in the deposit spread and zero when the spread (net of management cost) is zero.&lt;/p&gt;
&lt;h3 id="q7-how-does-this-paper-relate-to-and-differ-from-the-closest-prior-work"&gt;Q7. How does this paper relate to and differ from the closest prior work?&lt;/h3&gt;
&lt;p&gt;Versus Barrdear and Kumhof (2021): shares the transaction-cost money-demand approach but estimates a much smaller welfare benefit; their large +3% GDP gain comes mainly from the central bank buying public debt and lowering the government bond rate — a channel absent here. Versus Brunnermeier-Niepelt (2019): they get equivalence (no effect) under specific funding conditions; here CBDC does affect outcomes through seigniorage, opportunity cost, and redistribution. Versus Andolfatto (2021, monopoly bank) and Chiu et al. (2022, Cournot): in those the CBDC rate moves the deposit rate, whereas monopolistic competition here insulates the deposit rate (Lemma 1). Versus Chiu-Davoodalhosseini (2021): the opportunity-cost channel is shared. The paper abstracts from cyclical issues (cf. Burlon et al. 2022 DSGE; Piazzesi et al. 2022 monetary-policy use of rc) by focusing on steady state.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-main-caveats-and-scope-conditions-on-the-welfare-results"&gt;Q8. What are the main caveats and scope conditions on the welfare results?&lt;/h3&gt;
&lt;p&gt;(1) Steady-state only — no transitional or cyclical analysis. (2) Requires historically normal interest rates; near the ZLB all three channels are inert. (3) Liquidity and substitutability are treated as fixed design constraints in the welfare optimization, with only rc as the policy lever, because they may be technologically hard to set. (4) The headline ~60 bps gain relies on the extreme &amp;lsquo;case b&amp;rsquo; (zero-measure bankers own all banks) and on the welfare function ignoring bankers — i.e., it is largely a redistribution result, not a pure efficiency result. (5) The model deliberately shuts down CBDC effects on bank lending (banks fund at the risk-free rate), so disintermediation-of-credit channels stressed elsewhere are absent by construction. (6) Bank profits in the model equal net interest income (~1.5-2% of consumption), comparable to US bank NII but higher than actual bank profits.&lt;/p&gt;
&lt;h3 id="q9-is-cash-incorporated-and-does-it-change-the-conclusions"&gt;Q9. Is cash incorporated, and does it change the conclusions?&lt;/h3&gt;
&lt;p&gt;The baseline model excludes cash, but an appendix adds cash as a third zero-interest money in a nested CES (cash and CBDC combine, then that composite substitutes for deposits). The paper shows that if the &amp;lsquo;composite interest&amp;rsquo; of cash-plus-CBDC equals the rc of the two-instrument baseline, economic outcomes are unchanged: households rebalance across the three instruments so the equilibrium transaction cost and total cost of holding money are the same.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Go big or buy a home: The impact of student debt on career and housing choices</title><link>https://macropaperwarehouse.com/papers/go-big-or-buy-a-home-the-impact-of-student-debt-on-career-and-housing-choices/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/go-big-or-buy-a-home-the-impact-of-student-debt-on-career-and-housing-choices/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: Folch and Mazzone ask how undergraduate student debt shapes three intertwined post-college decisions — whether to pursue a post-bachelor (graduate) degree, the trajectory of earnings, and whether/when to buy a home. The motivation is the steep rise in student borrowing: between 1993 and 2016 the share of undergraduates who ever borrowed rose from 45% to 68%, and median cumulative borrowing rose from $14,329 to $29,115 (2020 dollars). The puzzle the paper resolves is why debt strongly distorts education and earnings yet has a negligible net effect on home ownership timing.&lt;/p&gt;
&lt;p&gt;Data and empirical strategy: The authors use restricted-use Baccalaureate and Beyond Longitudinal Study (B&amp;amp;B) data, focusing on the B&amp;amp;B:08/18 cohort (followed up to ten years post-graduation), merged with college-level IPEDS/College Scorecard data. The sample is restricted to US citizens/residents who earned a bachelor&amp;rsquo;s at ages 21-25, first enrolled 2001-2004, did not transfer, and excludes private for-profit colleges (~9,000 graduates in B&amp;amp;B:08/18; ~8,000 in B&amp;amp;B:16/17). In 2008, 72% of graduates held debt averaging $23,640; in 2016, 66% averaging $28,843. To address endogeneity of debt, they instrument with the change during enrollment in an institution-level grant-to-aid ratio (institutional grants / (grants + loans)), exploiting supply-side shifts in grants unlikely to be anticipated at application. The first stage is strong: one SD increase in grant-to-aid while enrolled predicts an ~18% decline in debt (about $4,250 lower balances), with F-statistics around 22-29.&lt;/p&gt;
&lt;p&gt;Main quantitative findings: Increasing debt balances by 10% ($2,364 relative to average $23,640) reduces the probability of obtaining a post-bachelor degree by about 1 percentage point (from a baseline of 22% four years after graduation and 45% ten years after). The same 10% increase raises initial post-graduation earnings — about +3.6% four years out ($1,440) and +$1,392 one year out — but reverses to a 5.3% decline ($2,828) ten years out. Graduate-school enrollment falls by about 0.85% (1 year) and 0.83% (4 years) per 10% debt increase. The net effect on first-time home ownership timing is statistically insignificant.&lt;/p&gt;
&lt;p&gt;Mechanisms: A life-cycle Roy model (Borjas 1987) with Ben-Porath (1967) human capital accumulation, housing, and financial frictions rationalizes this. Debt affects home ownership through two offsetting channels: (1) a traditional wealth effect that deters ownership, and (2) discouragement of further education that pushes graduates into early labor-market entry, accelerating ownership for that subgroup; these roughly cancel. Education choices are especially wealth-sensitive because post-bachelor attendance carries large non-monetary (amenity) returns valued at $3,929 on average (vs. $1,155 housing amenity), while the medium-run graduate wage premium is roughly 30% controlling for ability and human capital.&lt;/p&gt;
&lt;p&gt;Policy implications: Traditional mortgage-style fixed repayment imposes high burdens right after graduation, distorting human capital investment. Income-based repayment (modeled on PAYE, 10% of discretionary income, 20-year term with forgiveness) raises post-bachelor enrollment (from 35% to 42.4%) and home ownership, but adversely sorts lower-ability workers into graduate school via the implicit subsidy and dampens human capital investment through a Ben-Porath labor-supply/tax channel. The assessment is partial equilibrium.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-the-main-threats-to-it"&gt;Q1. What is the identification strategy and the main threats to it?&lt;/h3&gt;
&lt;p&gt;OLS of outcomes on log cumulative undergraduate debt is biased because unobservables (ability, true family contribution) drive both debt and outcomes. The authors instrument debt with the change during enrollment in an institution-level grant-to-aid ratio = institutional grants/(grants+loans). They use the CHANGE rather than the level (Eq. 2) because students may sort into colleges on the level of grants; mid-enrollment changes are unlikely anticipated. The exclusion concern is that grant-to-aid correlates with unobserved student characteristics affecting outcomes. They address relevance (first-stage F ~22-29; one SD raises grant-to-aid predicts ~18%/$4,250 lower debt) and conduct a balancing test (Table A.2) regressing the instrument on predetermined attributes — only financial need is significant (at 5%), and an F-test fails to reject joint insignificance. A residual threat is that idiosyncratic grant fluctuations could contract graduate slots at the same institution (supply-side); only 3.9% pursue graduate study at their undergrad institution, and splitting by Carnegie research vs. non-research institutions (Table A.8) leaves results intact. Another threat — relocation driving the housing/grad-school substitution — is addressed by re-estimating on 2009 and 2018 (years with state of residence): non-movers are 79% and 64%, and results closely mirror the full sample (Table A.7).&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-channels-through-which-debt-affects-home-ownership-and-how-are-they-distinguished"&gt;Q2. What are the two channels through which debt affects home ownership, and how are they distinguished?&lt;/h3&gt;
&lt;p&gt;Channel 1 is the traditional wealth effect: debt reduces wealth available for a downpayment, deterring ownership. Channel 2 is an indirect education channel: debt discourages graduate enrollment, pushing graduates into earlier labor-market entry where higher savings and lower balances facilitate earlier purchase, raising ownership for that subgroup. The two nearly cancel, yielding a negligible net effect. Empirically they are distinguished via ability sub-populations (Table 5): the housing response is negative for low-ability students but positive for high-ability students, and high-ability students cut enrollment more in response to debt. The structural model confirms it: for graduates who will not attend graduate school (Table A.10 Panel A), housing responds positively to debt; the substitution is also visible in life-cycle profiles where indebted bachelor holders have higher early ownership that reverses by age 30.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Ability heterogeneity is central. Two proxies are used: high-school grades, and time-to-degree (graduating within four years = high ability, five-plus years = low ability, following Hendricks and Leukhina 2018). High-ability graduates respond more in enrollment to debt; the housing response is positive for high-ability and negative for low-ability graduates (Table 5). In the model, the non-monetary value of graduate school is highly heterogeneous across the income distribution: poorer workers weigh almost only monetary returns, while high-income graduates value graduate school at the equivalent of hundreds of thousands of dollars in lifetime income, and debt shifts this distribution sharply leftward, especially for less wealthy individuals (Fig. 4).&lt;/p&gt;
&lt;h3 id="q4-what-robustness-checks-are-run"&gt;Q4. What robustness checks are run?&lt;/h3&gt;
&lt;p&gt;Restricting the instrument sample to institutions with at least 6 observed graduates (preferred spec, dropping 5-10% of obs; robust to alternative cutoffs); a balancing test (Table A.2); relocation/non-mover re-estimation for 2009/2018 (Table A.7); splitting by Carnegie research vs. non-research institutions (Table A.8); testing completion conditional on enrollment (no detectable effect, Table A.6); home value conditional on ownership (insignificant, Table A.9); a binary &amp;rsquo;ever borrowed&amp;rsquo; instrument specification implying smaller income effects (Table A.1); varying max sample age to 23 or 30 (similar results); age-dependent unemployment risk calibration leaving results unaffected; and a gradual house-price-trend exercise (1.4%/yr for 12 years, Table A.17) confirming the baseline.&lt;/p&gt;
&lt;h3 id="q5-how-does-this-relate-to-and-differ-from-prior-work"&gt;Q5. How does this relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;On earnings, the paper aligns with Rothstein and Rouse (2011), Luo and Mongey (2019), Field (2009), and Alon et al. (2023) showing debt raises initial earnings (their ~$500 per $1,000 is larger than Rothstein-Rouse&amp;rsquo;s ~$200, Luo-Mongey&amp;rsquo;s $70-160, and Alon et al.&amp;rsquo;s ~$210 — attributed to their Great Recession entry cohort and pre-ICL period); the ten-year reversal of ~$1,200 per $1,000 is close to Alon et al.&amp;rsquo;s ~$1,270. On graduate school, it complements Zhang (2013) and Chakrabarti et al. (2023); they find a $10,000 debt increase reduces probability of a post-graduate degree by 3.4%. On home ownership, it contrasts with Mezza et al. (2020), who find ~1pp reduction per $1,000; the null is attributed to sampling — excluding for-profit and two-year programs and dropouts (over one-fourth of US graduates) selects higher-ability, lower-debt individuals for whom the education-substitution channel offsets the wealth channel. The structural contribution extends the initial-conditions/lifetime-inequality literature (Huggett et al. 2011; Griffy 2021) by modeling multiple wealth dimensions and graduate-education choice.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-structural-model-add-and-how-well-does-it-fit"&gt;Q6. What does the structural model add and how well does it fit?&lt;/h3&gt;
&lt;p&gt;The model lets the authors control for ability explicitly and run the &amp;lsquo;ideal&amp;rsquo; regression on simulated data (Table 9): indebted graduates have 0.22% higher earnings per 1% additional borrowing one year out but 0.11% lower ten years out, qualitatively replicating data point estimates within/near the 95% CIs. It fits earnings profiles, enrollment (slightly over a third pursue further education), and home ownership (reaching ~85% by age 50 in model and data). The model attributes excess sensitivity of education to wealth to the amenity value of graduate school operating as a luxury good (parameter xi). Quantitatively, discrete-choice effects are somewhat stronger than data, partly because only one graduate-school type exists and bequests/inter-vivo transfers are omitted, steepening the home-ownership profile.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-ibr-policy-results-and-their-scope-conditions"&gt;Q7. What are the IBR policy results and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Under universal PAYE-style income-based repayment (tau=10% of discretionary income above a threshold, capped at the 10-year Stafford payment, 20-year term with forgiveness), post-bachelor enrollment rises from 35% to 42.4% and home ownership grows (50-plus ownership up &amp;gt;13%), but total retirement wealth rises only ~3% — the ownership gain is mostly a shift from liquid to housing wealth driven by reduced precautionary saving. Enrollment among non-indebted graduates falls from above 60% to ~40% (because the implicit subsidy is decreasing in income), while the most-indebted tercile&amp;rsquo;s enrollment jumps from ~3.5% to ~42%. IBR adversely sorts lower-ability workers into graduate school and dampens human capital investment via a Ben-Porath/proportional-tax channel (consistent with de Silva 2025, Fu et al. 2025). Fiscally, ~4% of individuals (6% of borrowers) get forgiveness averaging &lt;del&gt;$55,000 (&lt;/del&gt;$42,000 net of 24% tax), about $1,700 averaged across the cohort, or ~$20 per half-year period — small enough that behavioral feedback is negligible. SCOPE: the assessment is partial equilibrium, abstracting from general-equilibrium wage, return-to-education, and aggregate-demand adjustments.&lt;/p&gt;
&lt;h3 id="q8-why-does-the-earnings-effect-reverse-sign-over-time"&gt;Q8. Why does the earnings effect reverse sign over time?&lt;/h3&gt;
&lt;p&gt;Higher debt (lower net wealth) shifts the trade-off between current and future income: indebted graduates front-load earnings — choosing higher-paying occupations or careers rather than working more hours (labor-supply evidence is weak, Table A.5) — to ease debt payments on current consumption. The &amp;lsquo;smoking gun&amp;rsquo; for the later decline is that debt reduces graduate-school enrollment both short- and long-run, forgoing the ~30% graduate wage premium and reduced human-capital accumulation; the model adds that early career sorting is hard to reverse because re-enrolling entails partial loss of accumulated human capital.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Merger guidelines for the labor market</title><link>https://macropaperwarehouse.com/papers/merger-guidelines-for-the-labor-market/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/merger-guidelines-for-the-labor-market/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation. Antitrust review of mergers has historically focused almost entirely on harm to consumers (product-market monopoly), ignoring harm to workers (labor-market monopsony). Following the July 2021 White House executive order and the DOJ&amp;rsquo;s monopsony-based challenge to the Penguin Random House (PRH)/Simon &amp;amp; Schuster (SS) publishing merger, the agencies are now putting buyer power at the center of policy. The paper asks: how should Herfindahl-based merger-review thresholds, designed for product markets, perform if applied to local labor markets, and what efficiency gains would a merger need to leave workers unharmed?&lt;/p&gt;
&lt;p&gt;Model and data. The authors extend Berger, Herkenhoff, and Mongey (2022, &amp;ldquo;BHM&amp;rdquo;) to allow multi-plant (post-merger) ownership. The model has a representative household supplying labor through a nested-CES system (within-market substitutability governed by eta, across-market by theta, with eta &amp;gt; theta &amp;gt; 0), firms competing in quantities (Cournot/oligopsony), heterogeneous firm productivity, decreasing returns to scale, and capital. Firms set wages as a variable markdown on the marginal revenue product of labor; the markdown depends on the firm&amp;rsquo;s local payroll share. Markets are defined as 3-digit NAICS by commuting zone. Calibration is taken directly from BHM using confidential US Census data (LBD). Key estimated values: theta = 0.42 and eta = 10.85 (the elasticity-substitution parameters; the paper also reports theta = 0.45 in one passage), productivity dispersion sigma_z, returns to scale alpha, etc. The average market has 113 firms, an HHI of 0.11 (about nine equal firms), the average firm share is ~0.02, and the employment-weighted average markdown is 0.72 (workers paid 72% of marginal revenue product), equivalent to a labor-supply elasticity of 2.57.&lt;/p&gt;
&lt;p&gt;Theory. Proposition 1 shows that, absent efficiency gains, a within-market merger equalizes the two merged plants&amp;rsquo; markdowns at the level implied by their combined share, depresses both merging plants&amp;rsquo; wages, lowers the market wage index and employment, and reduces total worker pay. Non-merging firms&amp;rsquo; shares rise and they expand, so the actual rise in concentration is smaller than a &amp;ldquo;naive&amp;rdquo; calculation (adding pre-merger shares) would predict. Under the monopsony limit (infinitely many firms, or eta = theta), mergers have no effect.&lt;/p&gt;
&lt;p&gt;Main quantitative findings. (1) Model validation: replicating Arnold (2020), the model generates a change in log employment of -9.0 (vs Arnold -14.4, about three-fifths), log earnings -0.7 (vs -0.8), log payroll -10.5 (vs -12.1); earnings fall -4.4% in high-concentration vs -1.1% in medium-concentration markets (Arnold: -3.1% and -0.8%); the naive-concentration regression coefficient is 0.893 (Arnold 0.834), both below one. (2) PRH/SS simulation (PRH 37% share, SS 12%): with no efficiency gains the merger cuts author wages by 5%; the Required Efficiency Gain (REG) for worker-surplus neutrality is 17%. A merger of the two largest publishers gives -10% wages and a 30% REG; the two smallest Big Five give a 13% REG. (3) Applying product-market thresholds to labor markets via a 200,000-market simulation: under the stricter 1982 guidelines (block if post-merger HHI &amp;gt; 1800 and Delta-HHI &amp;gt; 100), the average REG of permitted mergers is 4.68%; under the looser 2010 guidelines (HHI &amp;gt; 2500, Delta-HHI &amp;gt; 200) it is 5.96%. Thus at the standard assumed 5% efficiency gain, 1982-permitted mergers raise the wage index (+0.04%) while 2010-permitted mergers lower it (-0.14%) and harm workers. (4) The Gross Downward Wage Pressure Index (GDWPI) equals (1/theta - 1/eta) times the other plant&amp;rsquo;s payroll share. Among mergers with GDWPI &amp;gt; 5% at both plants, more than 80% require a REG of at least 5.8% (20th-percentile REG = 5.8%, median 6.4%); among GDWPI &amp;gt; 10% at both plants, more than 80% generate a welfare loss under an assumed 5% efficiency gain.&lt;/p&gt;
&lt;p&gt;Implications. Product-market thresholds are too lenient for labor markets because labor is harder to substitute than products (low theta). The framework lets regulators trade off Type I error tolerance and efficiency-gain priors to set concentration thresholds.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identificationestimation-strategy-for-the-key-parameters-and-what-are-the-threats-to-it"&gt;Q1. What is the identification/estimation strategy for the key parameters, and what are the threats to it?&lt;/h3&gt;
&lt;p&gt;The model is not separately estimated; calibration is inherited wholesale from BHM (2022). The crucial labor-supply substitution parameters theta (across-market) and eta (within-market) are estimated in BHM from tradeable firms&amp;rsquo; market-share-dependent employment responses to corporate tax changes, identifying how much firms with different market shares move employment when after-tax returns change. Productivity dispersion sigma_z matches the payroll-weighted HHI, alpha matches labor&amp;rsquo;s share, gamma the capital share, Z mean firm size, and phi mean worker earnings. Main threats: (i) theta and eta are estimated from tradeable (largely manufacturing) firms and held fixed economy-wide, while the authors acknowledge no economy-wide substitutability estimates exist outside manufacturing; (ii) markets are defined by NAICS3-by-CZ rather than occupation (the conceptually preferred unit), because occupation codes are unavailable for the universe of workers; (iii) the whole exercise relies on the calibrated structure being the right laboratory.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-model-validated-out-of-sample"&gt;Q2. How is the model validated out of sample?&lt;/h3&gt;
&lt;p&gt;By replicating Arnold (2020), who estimates causal labor-market effects of US mergers. The authors draw and merge two firms per market, impose a pre-merger employment cutoff (tilde-n = 46, about five times average firm size) so that median pre-merger employment matches Arnold&amp;rsquo;s sample (116), and run Arnold&amp;rsquo;s exact regressions on simulated data. The model reproduces the sign and roughly the magnitude of employment and wage declines, the concentration interaction (effects more than three times larger in high-concentration markets), and the sub-one naive-concentration coefficient. This is out-of-sample because none of these moments were targeted in calibration.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-central-welfare-metric-and-policy-quantity"&gt;Q3. What is the central welfare metric and policy quantity?&lt;/h3&gt;
&lt;p&gt;Worker Surplus Neutrality: a merger is worker-surplus neutral if the market-level wage index W_j is unchanged (using a household problem in which profits are NOT rebated, to mirror the product-market consumer-surplus standard). The key policy object is the Required Efficiency Gain (REG, Delta-star): the common post-merger productivity gain at both plants needed to keep W_j constant. By Proposition 1.5 the REG is always positive.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-main-mechanisms-and-what-is-downward-wage-pressure-specifically"&gt;Q4. What are the main mechanisms, and what is downward wage pressure specifically?&lt;/h3&gt;
&lt;p&gt;Market power comes from costly worker mobility within (eta) and across (theta) markets. When two plants merge, hiring at Plant 1 raises the market wage and thus the wage the merged firm must pay its inframarginal workers at Plant 2 (and vice versa). The merged firm internalizes this cross-plant cost, which acts like a per-worker &amp;rsquo;labor cannibalization tax,&amp;rsquo; lowering the marginal benefit of hiring at both plants, so it hires less and pays less. Downward wage pressure at Plant 1 equals n_2j times the derivative of w_2j with respect to n_1j; in share form DWP_1j = w_1j (1/theta - 1/eta) s_2j. The GDWPI normalizes this by the wage: GDWPI_1j = (1/theta - 1/eta) s_2j, bounded in [0, theta^-1 - eta^-1], interpretable as a wage tax rate. Larger partner share and higher within-market substitutability (eta) raise downward pressure.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneity-is-documented"&gt;Q5. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Effects vary strongly with concentration: earnings fall -4.4% in high-concentration markets vs -1.1% in medium-concentration markets (model). Effects depend on the merging firms&amp;rsquo; shares: assuming a 5% efficiency gain, fewer than 12.1% of mergers in which the smaller firm&amp;rsquo;s payroll share exceeds 5% yield a worker-surplus gain. REGs differ across publisher pairings in the PRH case (17% for PRH+SS, 30% for the two largest, 13% for the two smallest). The model also generates wide firm-level variation in markdowns (small firms near competitive, large firms marked down well below 0.72).&lt;/p&gt;
&lt;h3 id="q6-what-do-the-confidencethreshold-figures-show"&gt;Q6. What do the confidence/threshold figures show?&lt;/h3&gt;
&lt;p&gt;Fixing a 5% efficiency gain, the simulation reports the fraction of mergers yielding a worker-surplus gain by concentration cell. 89.5% of mergers with post-merger HHI &amp;lt; 500 and Delta-HHI &amp;lt; 50 yield gains. Under the 2010 highly-concentrated definition (HHI &amp;gt; 2500, Delta-HHI &amp;gt; 100 in the cited cell), fewer than 34.8% yield gains. A merger with small-firm share 4% and large-firm share 18% has a 69.7% chance of a worker-surplus gain at 5% efficiency, rising to 97.7% at a 10% efficiency gain. This lets a regulator pick thresholds for a desired Type I error tolerance.&lt;/p&gt;
&lt;h3 id="q7-how-sensitive-are-results-to-the-assumed-efficiency-gain"&gt;Q7. How sensitive are results to the assumed efficiency gain?&lt;/h3&gt;
&lt;p&gt;Highly. Under 1982 guidelines, permitted mergers change average W_j by -0.40% at 1% efficiency, &amp;hellip; up to +0.04% at 5% efficiency; blocked mergers fall -7.39% (1%) to -5.99% (5%). Under 2010 guidelines, permitted mergers fall -0.63% (1%) to -0.14% (5%); blocked mergers fall -10.37% (1%) to -8.61% (5%). The 5% benchmark (Farrell-Shapiro) is itself questioned: Blonigen and Pierce (2016) find roughly zero or negative merger productivity gains, implying even the 1982 thresholds may be too lenient.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-differ-from-closely-related-prior-work"&gt;Q8. How does this paper differ from closely related prior work?&lt;/h3&gt;
&lt;p&gt;It extends BHM by adding multi-plant ownership and merger analysis. Relative to Nocke and Schutz (2018a,b) and Nocke and Whinston (2022), who derive product-market merger comparative statics under Bertrand competition (and, for Nocke-Whinston, CRS), this paper derives results for the LABOR market under nested-CES supply, Cournot competition, decreasing returns to scale, and endogenous household income. Relative to Naidu, Posner, Weyl (2018) and Marinescu-Hovenkamp (2019), who translate downward-wage-pressure concepts but assume symmetric firms, this paper provides a downward-wage-pressure test with firm heterogeneity across and within markets and shows it can be computed from readily available payroll shares and existing eta/theta estimates. It empirically benchmarks to Arnold (2020) and Prager-Schmitt (2021).&lt;/p&gt;
&lt;h3 id="q9-what-are-the-policy-implications-and-their-scope-conditions"&gt;Q9. What are the policy implications and their scope conditions?&lt;/h3&gt;
&lt;p&gt;Product-market HHI thresholds are too lenient when applied to labor markets: at an assumed 5% efficiency gain, 1982 thresholds (1800/100) keep permitted mergers worker-surplus neutral while 2010 thresholds (2500/200) do not. Scope conditions: (i) results hinge on the assumed efficiency gain (which empirical evidence suggests may be well below 5%); (ii) the framework treats product-market effects as &amp;lsquo;out of market&amp;rsquo; and should be combined with consumer-harm analysis; (iii) parameters are economy-wide benchmarks that may not fit a specific industry; (iv) market definition (NAICS3-by-CZ) matters, though the low estimated theta makes it consistent with a hypothetical-monopsonist test. The framework can be modified to add monopolistic pricing or variable markups (e.g., Deb et al. 2022).&lt;/p&gt;
&lt;h3 id="q10-are-there-internal-inconsistencies-a-reader-should-note"&gt;Q10. Are there internal inconsistencies a reader should note?&lt;/h3&gt;
&lt;p&gt;Yes. Table 1 reports theta = 0.42 (and 1.49 as the data moment), but the text at one point states &amp;rsquo;theta = 0.45, and eta = 10.85, giving theta^-1 - eta^-1 = 2.29.&amp;rsquo; The 2010 threshold is described in the abstract/Section 3 as Delta-HHI &amp;gt; 200 but the headline simulation result (4.68% vs 5.96%) compares &amp;lsquo;1800/100&amp;rsquo; against &amp;lsquo;2500/200&amp;rsquo;, and one passage lists the 2010 thresholds as (2500, 200) while the highly-concentrated text uses Delta-HHI of 200 for presumption and 100 in a figure cell. These are presentational; the substantive ranking (1982 stricter, 2010 more lenient) is robust.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;!-- flags: Internal parameter inconsistency: Table 1 reports theta=0.42 but text states theta=0.45 in the GDWPI bound passage (theta^-1 - eta^-1 = 2.29)., Threshold reporting: 1982 simulation uses Delta-HHI&gt;100 while Section 3 text also references Delta-HHI thresholds of 100/200; the headline comparison is 1800/100 vs 2500/200., Efficiency-gain assumption of 5% (Farrell-Shapiro) is load-bearing for the 'workers harmed under 2010 guidelines' conclusion; paper itself notes empirical evidence (Blonigen-Pierce 2016) of near-zero gains. --&gt;</description></item><item><title>Studying Generational Risk in a Large-Scale Life-Cycle Model</title><link>https://macropaperwarehouse.com/papers/studying-generational-risk-in-a-large-scale-life-cycle-model/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/studying-generational-risk-in-a-large-scale-life-cycle-model/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Hasanhodzic and Kotlikoff ask a question prior work assumed away: how large is generational risk, and can pay-go Social Security actually mitigate it? Earlier studies (Diamond, Bohn, Krueger-Kubler, etc.) presumed generational risk is large enough to merit policy and showed Social Security can in principle share it, but did not directly measure its size. This paper measures it directly, with and without Social Security, in a realistically large overlapping-generations (OLG) model.&lt;/p&gt;
&lt;p&gt;Model setup: an 80-period annual OLG model with aggregate shocks. Agents work 45 periods (retire at R=45) and live 80, have isoelastic (CRRA) preferences with risk aversion gamma=2 (gamma=5 under the extra-large shocks calibration), annual discount factor beta=0.96 (quarterly 0.99). Production is Cobb-Douglas; log TFP is trend-stationary AR(1) (quarterly rho=0.95, sigma=0.01; annualized rho=0.814, sigma=0.019). Two calibrations add a normal capital-depreciation shock. Households invest in risky capital or one-period safe bonds (zero net supply); &amp;ldquo;soft&amp;rdquo; increasing borrowing costs (Chen-Mangasarian function, slope b) shut down private risk-sharing to expose generational risk in its purest form while still delivering a realistic risk and growth premium. Policy is pay-go Social Security with a fixed payroll tax tau=15% (also tested at 1%). The model is solved to high precision via a projection method (building on Marcet 1988; Judd, Maliar, Maliar 2011) over an 81-variable state space (79 cohort cash-on-hand values plus the TFP and depreciation shocks). Generational risk measures are evaluated 300 years into the transition; cohort utility uses generations born after year 300 of a 750-year run. The U.S. data targets cover the return to national wealth and one-month Treasuries, 1947-2015, and detrended NNP/consumption, 1929-2020.&lt;/p&gt;
&lt;p&gt;Four calibrations: (1) baseline (TFP shock only, matched to output/consumption variability); (2) larger shocks (adds depreciation shock to match variability of the return to national wealth); (3) extra-large shocks (bigger depreciation shock to match U.S. equity-market return variability, a la Krueger-Kubler); (4) negative risk-free-rate baseline (steeper borrowing costs giving a roughly negative 2% safe rate, to test Blanchard 2019).&lt;/p&gt;
&lt;p&gt;Main findings (compensating-consumption differentials needed to reach long-run average lifetime utility): generational risk is 1.396% under baseline, 2.128% under larger shocks, and 15.303% under extra-large shocks (without Social Security). The authors view baseline 1.396% as small (on the order of a good-sized distortion) and prefer the baseline calibration. Social Security slightly WORSENS baseline generational risk (rising to 1.462%), but reduces it by 8% in the larger-shocks and 19% in the extra-large-shocks calibrations. So Social Security&amp;rsquo;s risk-pooling value depends on calibration. Contemporaneous risk (absolute consumption adjustment for full risk sharing among living cohorts) is tiny: 0.206% baseline, 0.933% larger shocks, 0.437% extra-large; Social Security raises it to 0.310% in baseline but lowers it under the other two.&lt;/p&gt;
&lt;p&gt;On welfare and Blanchard&amp;rsquo;s conjecture: pay-go Social Security at a 15% tax cuts long-run expected utility by 18% in baseline and larger-shocks, and by 56% in extra-large shocks, via crowding out (long-run capital falls 28% baseline, 56% extra-large). Under the negative-safe-rate calibration there is still an 18% long-run welfare loss; the average growth rate is zero in all simulations. The authors find no support for Blanchard&amp;rsquo;s (2019) claim that deficits can be Pareto-improving when safe rates run below growth: even under Blanchard-favorable conditions, crowding out swamps risk sharing (e.g., 17.83% utility loss at 15% tax, 1.17% at 1% tax). Macro shocks are second-order for policy: the capital transition under Social Security with shocks closely tracks the no-shock (deterministic) path, echoing Lucas (1987).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-exactly-is-the-papers-primary-measure-of-generational-risk"&gt;Q1. What exactly is the paper&amp;rsquo;s primary measure of generational risk?&lt;/h3&gt;
&lt;p&gt;It is the average absolute percentage adjustment to a cohort&amp;rsquo;s annual consumption needed to equate that cohort&amp;rsquo;s realized lifetime utility to the long-run cross-cohort average realized lifetime utility. Formally, for each generation born in period t they compute lambda_t = U-bar / U_t (U_t is realized lifetime utility, U-bar the average over generations born in years 301-750), then take the mean absolute deviation of lambda from 1. It captures both being born in a bad state and being hit by a bad sequence of lifetime shocks. A value near zero means birth date barely matters.&lt;/p&gt;
&lt;h3 id="q2-why-does-annualizing-to-80-periods-matter-relative-to-two-period-models"&gt;Q2. Why does annualizing to 80 periods matter relative to two-period models?&lt;/h3&gt;
&lt;p&gt;With one year per period, an agent experiences 45 annual wage shocks and 79 annual investment-return shocks that largely average out, and can self-insure by adjusting saving annually. In a two-period model a single negative TFP shock hits a worker&amp;rsquo;s entire lifetime earnings or a retiree&amp;rsquo;s whole old-age return. The authors note, however, that because TFP shocks are positively autocorrelated, amplifying multi-period shocks could in principle generate more risk, not less, so the result is not mechanical.&lt;/p&gt;
&lt;h3 id="q3-how-is-private-risk-sharing-handled-and-why-shut-it-down"&gt;Q3. How is private risk-sharing handled, and why shut it down?&lt;/h3&gt;
&lt;p&gt;In three of four calibrations the authors impose &amp;lsquo;soft&amp;rsquo; increasing borrowing costs (Chen-Mangasarian function, parameter b) calibrated so the marginal borrowing cost is 15-20 times the safe rate (b=28 baseline, 25 larger shocks, 45 for negative-safe-rate cases). This nearly closes the bond market, isolating generational risk with no private or public mitigation. The extra-large calibration omits borrowing costs because its large depreciation shock alone delivers a realistic risk premium (and to match Krueger-Kubler). Notably, adding borrowing constraints has little impact on key macro aggregates.&lt;/p&gt;
&lt;h3 id="q4-why-does-social-security-increase-generational-risk-in-the-baseline-single-tfp-shock-case"&gt;Q4. Why does Social Security INCREASE generational risk in the baseline (single-TFP-shock) case?&lt;/h3&gt;
&lt;p&gt;Five reasons given: (1) benefits depend on the prevailing wage, so autocorrelated TFP wage shocks now interact with capital-return shocks through retirement, extending nonlinear discounting past retirement; (2) crowding out lowers wages and raises risky returns, so the same percentage TFP shock is larger in absolute terms, making realized resources more variable; (3) Social Security is a random floor on old-age living standards, encouraging less risk-averse consumption and a higher propensity to consume; (4) positive TFP autocorrelation (high benefits today predict high benefits tomorrow) further raises the propensity to consume; (5) Social Security alters the stochastic distribution of the 79 cohort cash-on-hand state variables, producing complex consumption changes. This echoes Rios-Rull&amp;rsquo;s (1994) paradox that better micro insurance can amplify macro fluctuations.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-paper-test-blanchards-2019-deficits-may-be-free-conjecture-and-what-does-it-find"&gt;Q5. How does the paper test Blanchard&amp;rsquo;s (2019) &amp;lsquo;deficits may be free&amp;rsquo; conjecture and what does it find?&lt;/h3&gt;
&lt;p&gt;It uses Blanchard&amp;rsquo;s own ex-ante Pareto criterion but with 80 periods (vs his 2), realistic risk aversion, and dropping his assumption that half of wages are perfectly safe. Calibrations engineered with negative safe rates and large growth premiums (e.g. risky ~2%, safe ~negative 2%) still show Social Security reducing long-run expected utility: 17.83% loss at a 15% tax (1.17% at 1%) in the standard-premium case, falling to 12.51%/12.582% (15% tax) under even-larger growth premiums, but always negative. Crowding out dominates any risk-sharing gains. The authors find no support for the conjecture. They note Blanchard&amp;rsquo;s Pareto gains, when they arise, depend critically on his assumption that half of wages are certain, leaving workers ideally placed to insure the elderly.&lt;/p&gt;
&lt;h3 id="q6-what-heterogeneity-across-cohorts-is-documented"&gt;Q6. What heterogeneity across cohorts is documented?&lt;/h3&gt;
&lt;p&gt;Baseline generational risk has mean 1.396%, s.d. 1.293%, max 4.949% (no Social Security). Decomposed: generations with worst luck need roughly +5.0% positive adjustment; those with best luck need roughly negative 5.1%. Extra-large shocks produce extreme spread: max positive adjustment 66.14%, max negative 44.10%. A separate exercise (Table 8) shows the cost of uncertainty depends on birth state due to mean reversion: those born with low capital actually prefer uncertainty (negative 1.482%) because capital and wages will rise, while those born with high capital would pay 2.374% to lock in their state.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-welfare-cost-of-uncertainty-and-precautionary-saving-findings"&gt;Q7. What are the welfare-cost-of-uncertainty and precautionary-saving findings?&lt;/h3&gt;
&lt;p&gt;Under larger shocks, the compensating variation between the stochastic steady state and a no-shocks steady state is only 1.12% (newborns would need 1.12% more consumption each year to match a never-shocked long run), despite that calibration overstating macro variability. This is small because precautionary saving raises the stochastic economy&amp;rsquo;s average capital stock 18.4% above the no-shocks steady state: the uncertain long run is &amp;lsquo;riskier, but richer.&amp;rsquo; A decomposition removing the 0.77% average age-specific consumption difference leaves a 0.34% residual (about one quarter of 1.12%) reflecting age-pattern and cohort-sequence heterogeneity.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-build-on-and-differ-from-krueger-kubler-2006"&gt;Q8. How does this paper build on and differ from Krueger-Kubler (2006)?&lt;/h3&gt;
&lt;p&gt;Five differences: (1) many more periods (80 vs 9) permit better shock-averaging and more precise autocorrelation treatment plus more self-insurance opportunities; (2) two calibrations the authors view as more realistic than KK (who chose theirs partly to favor a Pareto improvement), using borrowing costs rather than excessively large depreciation shocks to get a realistic risk premium; (3) ex-ante rather than ex-interim expected utility; (4) explicit measurement of generational risk with and without Social Security; (5) testing whether a large growth premium can sustain an intergenerational Ponzi scheme at scale. Like KK, they find a negative net long-run welfare impact of pay-go Social Security.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-model-deliberately-omit-and-why"&gt;Q9. What does the model deliberately omit, and why?&lt;/h3&gt;
&lt;p&gt;It is &amp;lsquo;intentionally bare bones to maximize the potential for generational risk&amp;rsquo;: no variable labor supply (which would help cohorts self-insure), no progressive income taxation (which redistributes from winning to losing generations), and no social insurance other than Social Security. It also omits capital-adjustment costs (which would raise asset-return volatility) because incomplete markets make firm investment policy ill-defined when differently-aged shareholders disagree; the depreciation shock is a crude proxy for adjustment-cost-driven asset-return shocks. The authors flag correlated idiosyncratic shocks (Harenberg-Ludwig) as important future work.&lt;/p&gt;
&lt;h3 id="q10-how-well-does-each-calibration-match-the-data"&gt;Q10. How well does each calibration match the data?&lt;/h3&gt;
&lt;p&gt;Baseline matches output (model 3.72% vs data 3.33%) and consumption (2.10% vs 1.75%) variability but understates the s.d. of the return to national wealth by an order of magnitude (0.14% vs 4.89%). Larger shocks reproduces the return-to-wealth s.d. (4.61-4.62% vs 4.89%) and a realistic wage/return correlation (negative 0.054) but overstates macro-aggregate variability. Extra-large shocks matches equity Sharpe ratio (model 0.333 vs target 0.286; risk premium 4.63%, return s.d. 13.92%) but overstates return-to-capital variability nearly three-fold and consumption variability sixteen-fold. The model&amp;rsquo;s overall risk premium ranges 3.55-6.03% vs 5.43% in data.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-role-of-the-bond-market-across-calibrations"&gt;Q11. What is the role of the bond market across calibrations?&lt;/h3&gt;
&lt;p&gt;The one-period bond market only operates in the extra-large shocks calibration (borrowing costs close it in the others). There, the young short bonds and the old lend: because the young&amp;rsquo;s resources are mostly human capital (less risky than, and negatively correlated with, stock returns), the young use bonds to insure the old. Workers effectively borrow to hold equity, which the authors rationalize via student loans, credit cards, mortgages alongside 401(k) equity, or implicit long-term firm contracts.&lt;/p&gt;
&lt;h3 id="q12-what-policy-implications-follow-and-what-are-their-scope-conditions"&gt;Q12. What policy implications follow, and what are their scope conditions?&lt;/h3&gt;
&lt;p&gt;If macro shocks are calibrated to realistic macro-aggregate volatility (the authors&amp;rsquo; preferred baseline), generational risk is small (about 1.4%) and pay-go Social Security slightly worsens it while imposing an 18% long-run welfare loss via crowding out; deterministic models (e.g. Auerbach-Kotlikoff 1987) then suffice to capture the long-run impact of intergenerational redistribution. Social Security&amp;rsquo;s risk-mitigation value emerges only under calibrations that overstate macro volatility (larger/extra-large shocks). The scope condition is decisive: the case for Social Security as generational insurance hinges on which calibration one finds realistic, and the authors&amp;rsquo; preferred reading implies a weak case. They also caution the conclusions may not extend to models with correlated idiosyncratic risk.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Time Averaging Meets Heckman, Lochner, and Taber and Ben-Porath</title><link>https://macropaperwarehouse.com/papers/time-averaging-meets-heckman-lochner-and-taber-and-ben-porath/</link><pubDate>Wed, 01 Jan 2025 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/time-averaging-meets-heckman-lochner-and-taber-and-ben-porath/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;Research question and motivation: How does endogenizing retirement (career-length) choice change the labor-supply and human-capital implications of the canonical Heckman, Lochner, and Taber (1998a, HLT) life-cycle general-equilibrium model, and what does this imply for social-security reform, labor-income taxation, aggregate labor-supply elasticities, and inequality? HLT already contains two ingredients of Ljungqvist-Sargent (2006) &amp;ldquo;time-averaging&amp;rdquo; models — credit markets and within-period labor-supply indivisibilities — but shuts time-averaging down by assuming inelastic labor supply until a mandatory retirement age of 65. The authors &amp;ldquo;activate&amp;rdquo; time-averaging by letting workers choose when to retire and by adding a pay-as-you-go social security system. This matters because the micro-foundation of the high aggregate labor-supply elasticity that Prescott invoked (switching from Rogerson&amp;rsquo;s employment lotteries to time-averaging) hinges on whether workers sit at corner solutions for career length.&lt;/p&gt;
&lt;p&gt;Model setup: A perfect-foresight OLG model in discrete annual time; agents live from age 18 to 80. Eight agent types index four innate ability levels (theta in {1,2,3,4}) crossed with two education levels (high school S=1, college S=2). Each type has a Ben-Porath (1967) human-capital technology. An aggregate CES/Cobb-Douglas production function combines physical capital and two human-capital aggregates. Within-period labor is indivisible (work full time omega=1 or not omega=0). Utility is time-separable with intertemporal elasticity 1/gamma and a fixed disutility B of working. The baseline social security program has payroll tax rate tau_p=0.10, eligibility age eta_p=65, and benefit P=8 (about 40% of average earnings), paid only to retirees; collecting nothing while working after 65 creates an implicit tax that pins all workers to a corner at age 65.&lt;/p&gt;
&lt;p&gt;Calibration: Most parameters are borrowed or backed out from HLT (delta=0.96, gamma rounded from 0.9 to 1, tau_l=tau_k=0.15, tuition zeta=1.02 thousand 1992 dollars). New parameters: disutility B=0.8, fraction of capital held by in-model agents kappa=0.388, efficiency-decline logistic parameters phi1=0.2, phi2=75. The model targets a capital-output ratio of 4 and an after-tax interest rate of 0.05; the calibrated model reproduces HLT&amp;rsquo;s baseline and post-skill-biased-technological-change (SBTC) steady states closely (e.g., baseline interest rate 0.0588 matched; aggregate human capital H1≈274/249, H2≈280/287 in HLT/our model).&lt;/p&gt;
&lt;p&gt;Main quantitative findings (with scope conditions): (1) Social security reform that pays benefits from 65 regardless of work removes the implicit tax wedge. At fixed prices all workers extend careers (high school +2.4 years on average; college +7.6 years to age 72.6); in general equilibrium effects are attenuated — high school workers actually retire ~1 year early (average 63.9) while college workers retire later (average 70.8). (2) Tax experiment along Prescott (2002) lines: raising tau_l with revenue rebated lump-sum produces a Laffer curve peaking at tau_l=0.54; without rebates the Laffer curve peaks at tau_l=0.73 (general equilibrium) and the small-open-economy version is nearly linear. (3) The aggregate labor-supply elasticity is zero at low tax rates (corner at 65), then rises above 1 and levels around 1.2 over a wide middle range before rising again past tau_l=0.7. (4) Ben-Porath nonconvexities create &amp;ldquo;tipping points&amp;rdquo;: e.g., high school ability-3 workers are indifferent between two starkly different career strategies over tax range 0.42-0.52, and at high tax rates workers jump discretely from long careers with high human capital to much shorter careers with little/no on-the-job investment.&lt;/p&gt;
&lt;p&gt;Implications: College-educated (steeper-earnings-profile) workers&amp;rsquo; labor supplies are more resilient to tax and social-security reforms than high school workers&amp;rsquo;. High tax rates with lump-sum rebates can produce a &amp;ldquo;dual labor market&amp;rdquo; / bifurcation, raising lifetime earnings inequality (Gini) while welfare conditioned on schooling converges, all at a growing efficiency cost.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-methodological-contribution-relative-to-hlt"&gt;Q1. What is the core methodological contribution relative to HLT?&lt;/h3&gt;
&lt;p&gt;The authors retain HLT&amp;rsquo;s primitives (credit markets, indivisible within-period labor, Ben-Porath human capital, aggregate production) but replace HLT&amp;rsquo;s exogenous mandatory retirement at 65 with endogenous career-length choice, and add a pay-as-you-go social security system. The social security system with an implicit tax on working past 65 puts all workers at a corner solution at age 65, so the model reproduces HLT&amp;rsquo;s outcomes. This provides a choice-theoretic rationalization for retirement behavior that HLT hard-wired. They state HLT could have used this time-averaging model with endogenous retirement to obtain the same quantitative findings.&lt;/p&gt;
&lt;h3 id="q2-why-is-there-no-separate-identificationempirical-strategy-in-the-usual-sense"&gt;Q2. Why is there no separate identification/empirical strategy in the usual sense?&lt;/h3&gt;
&lt;p&gt;This is a calibrated/quantitative general-equilibrium model, not a reduced-form causal study. Parameters are borrowed or &amp;lsquo;backed out&amp;rsquo; from HLT (who estimated human-capital technologies via nonlinear least squares on NLSY 1979-1993 earnings profiles for white male civilians, plus CPS 1963-1993 and NIPA aggregates). New parameters are calibrated to be compatible with HLT: B and the efficiency-decline parameters (phi1, phi2) are jointly set so all agents retire at 65 in baseline; kappa=0.388 is set to match HLT&amp;rsquo;s interest rate given a capital-output ratio of 4; sigma (dispersion of nonpecuniary college cost) is calibrated to match the 8% rise in the relative college skill price between HLT&amp;rsquo;s two steady states; ability-specific means mu_theta target college enrollment rates from Taber (2002, Table 1).&lt;/p&gt;
&lt;h3 id="q3-what-are-the-three-forces-that-make-high-school-workers-retire-earlier-than-college-workers-under-the-social-security-reform"&gt;Q3. What are the three forces that make high school workers retire earlier than college workers under the social security reform?&lt;/h3&gt;
&lt;p&gt;First, the social security system redistributes from high-ability to low-ability agents (equal benefit, proportional payroll tax), and the income effect on low-ability (mostly high school) workers reduces their labor supply; removing social security entirely (recalibrating kappa from 0.388 to 0.767) shows lowest-ability high school workers extend careers most. Second, per Ljungqvist-Sargent (2014), the more elastic an earnings profile to accumulated work, the longer the career; giving high school workers college workers&amp;rsquo; more productive human-capital technology lengthens their careers. Third, a time-averaging &amp;lsquo;apprenticeship&amp;rsquo; effect: college is treated as a fixed pre-work requirement Z tacked onto an optimal working span, so at an interior solution optimal career length = baseline length + Z; this accounts for roughly a 4-year career-length difference between high school and college workers in the relevant perturbed economy.&lt;/p&gt;
&lt;h3 id="q4-how-do-the-effects-of-a-labor-tax-increase-depend-on-how-revenue-is-spent-and-what-is-the-mechanism"&gt;Q4. How do the effects of a labor tax increase depend on how revenue is spent, and what is the mechanism?&lt;/h3&gt;
&lt;p&gt;Following Prescott (2002): if revenue is rebated lump-sum (a good substitute for private consumption), the income effect of the tax is suppressed and the substitution effect dominates, sharply reducing labor supply (Laffer peak at tau_l=0.54). If revenue is squandered or spent on poor substitutes, income and substitution effects roughly cancel under balanced-growth preferences, so labor supply is little affected (Laffer peak at tau_l=0.73 in GE; nearly linear/flat in the small-open-economy version where capital inflows hold the interest rate constant at 0.059). With lump-sum rebates the equilibrium interest rate is U-shaped in the tax rate and the Laffer curve eventually approaches zero (output collapses); without rebates the interest rate rises monotonically to offset what would otherwise be capital inflows.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-ben-porath-nonconvexities-and-the-tipping-points"&gt;Q5. What are the Ben-Porath nonconvexities and the &amp;rsquo;tipping points&amp;rsquo;?&lt;/h3&gt;
&lt;p&gt;Returns to on-the-job human-capital investment can only be harvested over a long enough career, so the value function over retirement ages can become non-concave with two local maxima: a long career with high end-of-life human capital versus a short career with little/no investment. As a determinant (tax rate, disutility, technology productivity) changes incrementally, the optimal response can be discontinuous — a discrete jump to a much shorter career and much less human-capital accumulation. Example: at tau_l=0.45 high school ability-3 workers have two optima, retirement at 65 (high human capital) and early retirement at age 50 (low human capital); they are indifferent over tax range 0.42-0.52. The nonconvexity is intrinsic to the Ben-Porath technology and arises even in a laissez-faire economy with interior career-length solutions, not only because of the social-security corner.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-indifference-between-career-strategies-handled-in-equilibrium-heterogeneity-and-computation"&gt;Q6. How is the indifference between career strategies handled in equilibrium (heterogeneity and computation)?&lt;/h3&gt;
&lt;p&gt;When otherwise-identical agents become indifferent between two career strategies, the regularity condition of a unique solution fails. The authors extend the equilibrium definition to allow equilibrium fractions of identical agents choosing different strategies; market clearing pins down these fractions (a &amp;lsquo;convexification&amp;rsquo;). Computationally they identify the &amp;lsquo;most indifferent&amp;rsquo; worker type (smallest gap between the two local maxima; threshold 0.05%) and vary the fraction retiring at each age until GE conditions are satisfied. They also introduce continuous retirement ages via cubic-spline interpolation of the value function, validated against a closed-form analytical formula for agents who do not accumulate human capital (largest deviation only about half a month at tau_l=0.61).&lt;/p&gt;
&lt;h3 id="q7-what-heterogeneity-is-documented-across-the-eight-worker-types"&gt;Q7. What heterogeneity is documented across the eight worker types?&lt;/h3&gt;
&lt;p&gt;College enrollment rises with ability in baseline (about 0.11, 0.34, 0.56, 0.86 for ability groups 1-4 in the authors&amp;rsquo; model). Group 4 has the second-highest average disutility of attending college, so 14% of group 4 become high school workers despite large advantages, and group 4&amp;rsquo;s enrollment falls most sharply with higher taxes. Group 1 has the highest disutility and lowest college human capital, so only ~11% attend college, falling below 1% above tau_l=0.45. End-of-life human capital of lower ability groups (1,2) falls monotonically with taxes, while higher ability groups (3,4) initially raise human capital as the interest rate falls. High school ability-1 workers eventually stop working entirely at the highest tax rates, with lifetime labor earnings falling to zero, relying on lump-sum transfers and social security.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-paper-find-for-aggregate-labor-supply-elasticity-and-why-is-12-notable"&gt;Q8. What does the paper find for aggregate labor-supply elasticity, and why is ~1.2 notable?&lt;/h3&gt;
&lt;p&gt;With lump-sum rebates, after an initial range of zero elasticity (all at the corner of retiring at 65), the elasticity quickly rises above 1 and levels around 1.2 over a substantial middle range, then rises again after tau_l=0.7 (as physical capital gets scarce and the interest rate rises steeply). The ~1.2 is notable because in the Ljungqvist-Sargent (2014) framework with the same utility, the analytical aggregate elasticity is exactly one regardless of the learning-by-doing wage exponent; the model obtains ~1.2 despite college workers being stuck at the corner until tau_l≈0.6, because falling college enrollment shifts would-be college workers into earlier-retiring high school careers. Without rebates the elasticity is suppressed.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-inequality-findings"&gt;Q9. What are the inequality findings?&lt;/h3&gt;
&lt;p&gt;Two measures: present value of lifetime labor earnings and lifetime utility. The pre-tax earnings Gini is roughly flat for the first five percentage points above baseline (all still retiring at 65), then rises nearly one-to-one with the tax rate until tau_l=0.65, flattens as college ability groups 2 and 3 switch to short careers, drops when group 4 (highest earners) switches, then rises again as college workers&amp;rsquo; relative earnings surge (driven by the rising college skill premium compensating for tuition and nonpecuniary costs). Using the Holter-Ljungqvist-Sargent-Stepanchuk (2025) ex post-ex ante welfare measure, higher taxes with lump-sum transfers shrink welfare inequality conditional on schooling even as income inequality grows, at an efficiency cost that accelerates above tau_l=0.4.&lt;/p&gt;
&lt;h3 id="q10-how-do-taxation-results-differ-under-the-social-security-reform-versus-the-baseline-social-security-system"&gt;Q10. How do taxation results differ under the social security reform versus the baseline social security system?&lt;/h3&gt;
&lt;p&gt;Laffer curves under the reform (Figure 12a) closely resemble the baseline (Figure 2a). The key difference is that under the reform workers are at interior career-length solutions, so high school workers&amp;rsquo; average retirement age falls with the very first tax increments (rather than staying stuck at 65), and college workers raise average retirement ages over a mid-range of taxes. At sufficiently high taxes the two economies become identical (above tau_l=0.74 with, 0.72 without rebates), because the implicit post-65 tax wedge becomes irrelevant once everyone retires early. Under the reform, college workers&amp;rsquo; careers are &amp;lsquo;anchored&amp;rsquo; near the age where human-capital efficiency depreciates rapidly rather than by the official retirement age.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-relate-to-and-differ-from-fan-seshadri-and-taber-2024"&gt;Q11. How does the paper relate to and differ from Fan, Seshadri, and Taber (2024)?&lt;/h3&gt;
&lt;p&gt;FST (2024) independently endogenize career lengths in a Ben-Porath model estimated on SIPP data for male high school graduates, with nine worker types differing in disutility B(theta), learning ability A(theta), and initial human capital H(theta). A key difference: FST impose identical Ben-Porath exponents across all workers, so the Ljungqvist-Sargent force (more elastic earnings profiles imply longer careers) is largely absent; and FST do not impose balanced-growth preferences, so income effects of higher wages do not cancel. The authors suspect the sharp declines in career length with higher productivity in FST&amp;rsquo;s first two rows reflect income effects, and that time-averaging strengthens income effects. In the authors&amp;rsquo; own balanced-growth model, the level of wages does not affect labor supply — only the terms on which human capital can be accumulated.&lt;/p&gt;
&lt;h3 id="q12-what-robustnesssensitivity-checks-and-appendices-are-reported"&gt;Q12. What robustness/sensitivity checks and appendices are reported?&lt;/h3&gt;
&lt;p&gt;Appendix C: sensitivity analysis of disutility B and the efficiency-decline function e(n); searching over (B, phi1) that keep all agents retiring at 65 yields end-point coordinates approximately (0.59, 0.09) and (0.9, 0.31), with the baseline (B=0.8, phi1=0.2) chosen as an intermediate pair subject to no noticeable efficiency decline before the 60s. Appendix D: alternative social security reforms raising benefits — college workers keep retiring at 65 while high school workers retire ever earlier. Appendix F.1: elasticity of the aggregate human-capital composite Q. Appendix G: replacing the Ben-Porath technology with exogenous earnings-experience profiles yields less polarization (lower Gini) and a lower aggregate labor-supply elasticity. The authors also note an unresolved discrepancy: their present-value earnings are 6.9-7.0% (high school) and 7.1-7.2% (college) lower than HLT&amp;rsquo;s Table II, but college enrollment is little affected since differences are similar across schooling.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-main-caveats-and-policy-scope-conditions"&gt;Q13. What are the main caveats and policy scope conditions?&lt;/h3&gt;
&lt;p&gt;Results depend on balanced-growth preferences (income/substitution effects of wage levels cancel), on HLT&amp;rsquo;s estimated human-capital technologies and nonpecuniary college-cost distributions, and on the auxiliary kappa device for targeting the capital-output ratio. The disutility B and efficiency-decline parameters are not pinned down by data when workers sit at the 65 corner, hence only a sensitivity analysis. Limited heterogeneity (only 8 types) means aggregate smoothness comes from convexification rather than from a continuum of switching agents. The central policy warning — that high enough tax wedges or distortions can dislodge even high-productivity workers into a &amp;lsquo;dual labor market&amp;rsquo; with earlier retirement and less human-capital accumulation, risking an implosion of activity — applies within this calibrated structure.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>Who bears the costs of inflation? Euro area households and the 2021-2023 shock</title><link>https://macropaperwarehouse.com/papers/who-bears-the-costs-of-inflation-euro-area-households-and-the-2021-2023-shock/</link><pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/who-bears-the-costs-of-inflation-euro-area-households-and-the-2021-2023-shock/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper measures the heterogeneous first-order welfare effects of the 2021-2023 inflation surge across households in the four largest euro area countries (Germany, France, Italy, Spain). Motivation: euro area headline HICP inflation peaked at 10.6% (year-on-year) in October 2022, driven mainly by energy and food prices following Russia&amp;rsquo;s invasion of Ukraine; cumulatively over 2021-23 the price index rose roughly 14% in France and Spain, 16% in Italy and 20% in Germany. The classic question—who wins and who loses from surprise inflation, and through which channels—is the focus.&lt;/p&gt;
&lt;p&gt;Method: The authors build a tractable two-period overlapping-generations framework and use the envelope theorem to decompose the &amp;ldquo;money-metric&amp;rdquo; welfare change (in euros) into four additive, observable components requiring no functional-form or structural-parameter assumptions: (1) a direct component (raw inflation before fiscal support, holding wages and asset prices fixed; captures heterogeneous consumption baskets and the Fisher revaluation of net nominal positions, labor income, dividends and capital gains); (2) an unconventional fiscal policy component (ad-hoc energy price interventions and transfers); (3) an indirect component (short-run responses of nominal wages, pensions, taxes/fiscal drag, and asset prices); (4) a long-run adjustment component (relative prices returning to pre-shock ratios). They combine micro data—Household Budget Survey (2015 wave) for expenditure shares, HICP micro data for good-specific price changes (20 COICOP-based categories), the 2017 Household Finance and Consumption Survey (HFCS) for budget-constraint components, the Bruegel dataset for fiscal responses, and IMF (Dao et al. 2023) counterfactual prices—with event-study/high-frequency identification (on German HICP release days) for wage, pension, house, stock and bond price responses. Households are sorted into 15 groups: three age classes (25-44 young, 45-64 middle-aged, 65+ retirees) and five consumption (permanent-income proxy) quintiles per country. Welfare is expressed as a share of triennial (3-year) disposable income.&lt;/p&gt;
&lt;p&gt;Main findings: (i) Average country-level welfare losses were sizable and heterogeneous: around 3% of triennial income in France and Spain, 7% in Germany, and 9% in Italy. (ii) The episode resembles an age-dependent tax: retirees lost up to 14% (German and Italian high-income retirees), while roughly half of 25-44 year-olds were net winners; young French households gained up to 7% (about EUR 4,000 on average), young Spanish broke even; middle-aged households lost roughly 2-11%. Overall about one quarter of euro area households were net winners. (iii) Losses were quite uniform across consumption quintiles because rigid (sticky) rents hedged the poor; excluding rents, the poor suffer more due to higher energy/food exposure. (iv) Nominal net positions (NNP) were the key driver of cross-household heterogeneity—retirees hold large positive nominal assets, the young hold nominal mortgage debt. (v) Energy prices generated vast individual-inflation-rate variation, but unconventional fiscal policy (especially energy price caps, more so in France where it cut inflation ~2 p.p.) shielded households, reducing first-stage welfare costs by about one-fifth on average. Estimated asset-price elasticities to a 10% inflation surprise: house prices -1.38% (beta x delta = -3.995 x 0.035 = -0.138), stocks -0.410, bonds -0.726. Pensions, being indexed, rose faster than wages; fiscal drag taxed away gains in Italy and Spain (unindexed brackets), much less in France/Germany. The counterpart of household losses is a large government gain from eroded real public debt: governments in France, Italy and Spain were net winners (Italy +4.5 to 5.1% of triennial GDP), while Germany roughly broke even. Policy implication: in a monetary union where monetary policy cannot address country-specific dynamics, fiscal policy was crucial; and redistributing government inflation gains to households could substantially offset their losses.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identificationmeasurement-strategy-and-what-are-its-main-threats"&gt;Q1. What is the identification/measurement strategy and what are its main threats?&lt;/h3&gt;
&lt;p&gt;The core strategy is an envelope-theorem decomposition that yields analytical &amp;lsquo;sufficient-statistic&amp;rsquo; formulas for money-metric welfare change, requiring only observable budget-constraint quantities and price changes—no structural parameters or functional forms. The key assumption is that, to first order, substitution in consumption baskets and portfolio rebalancing after the shock have only second-order welfare effects, so observed pre-shock quantities (2015 HBS shares, 2017 HFCS positions) can be used. Four structural assumptions define the shock: (1) it is unanticipated; (2) the price-level jump is permanent but inflation is temporary (returns to zero from t=1); (3) the shock is long-run neutral in aggregate and across the distribution—all nominal variables and relative prices realign one-to-one with the new price level by t=1; (4) the government budget constraint accommodates either via the price level (active/FTPL) or via future real surpluses (passive). For asset-price responses they use high-frequency identification: regressing daily REIT, stock and bond returns on the inflation surprise (daily change in 1-year inflation-linked swaps) on German HICP release days, controlling for stock returns. Main threats: the first-order/second-order approximation could fail if substitution effects are large (the authors note that pre/post high-frequency micro data—unavailable to them—could test this); the use of 2015 expenditure shares and 2017 balance sheets to represent the pre-shock state; reliance on counterfactual price series (IMF, OMIE) for what prices would have been absent intervention; and the assumption that relative prices fully return to pre-shock ratios in the long run.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-four-channels-and-how-are-they-distinguished-empirically"&gt;Q2. What are the four channels and how are they distinguished empirically?&lt;/h3&gt;
&lt;p&gt;(1) Direct component: raw inflation effect on cost of living before fiscal support and before wage/asset-price adjustment; split into average inflation, the &amp;lsquo;pi difference&amp;rsquo; from heterogeneous baskets (C), net income/labor-income purchasing power (Y), net nominal positions (NNP), and dividends+capital gains (K). (2) Unconventional fiscal policy (UFP): energy price interventions (changes in good-specific tax/subsidy wedges, requiring counterfactual no-intervention price indices) plus ad-hoc transfers to households. (3) Indirect: short-run changes in nominal wages, minimum wages, pensions, fiscal drag, and asset prices (house, stock, bond) plus the direct effect of monetary-policy-driven interest-rate changes on deposits and debt. (4) Long-run: welfare from relative prices realigning to the new price level, discounted to t=0. They are computed sequentially in stages so each component&amp;rsquo;s contribution is isolated. NNP is the dominant driver of age heterogeneity; Y is the largest single contributor to losses but is fairly uniform across groups; C matters mainly for poor elderly in Italy and Spain.&lt;/p&gt;
&lt;h3 id="q3-what-heterogeneity-is-documented"&gt;Q3. What heterogeneity is documented?&lt;/h3&gt;
&lt;p&gt;Age is the most pronounced dimension: retirees lose most (driven by large positive nominal asset holdings), the young least (often net winners via mortgage debt revaluation). German and Italian retirees lost up to 14% of triennial income; high-income retirees lost more than EUR 10,000 on average. By contrast, the consumption-quintile (permanent-income) gradient is weak because sticky rents hedge low-income renters; excluding rents reveals a negative inflation-income gradient (poor face higher inflation via energy/food). Cross-country: Italy highest cost (~9%), France lowest (~3%), due to (i) bigger raw price shock in Italy (energy import dependence/market structure), (ii) more effective fiscal offset in France, (iii) nominal wages lagging inflation much more in Italy, (iv) Italian middle-aged/elderly holding larger nominal positions while the young borrow less than in France. Within-bin heterogeneity (homeowners with mortgages vs renters) means about a quarter of households are winners overall; more than half of the young in France and Spain, ~50% in Germany, ~30% in Italy, and ~50% of Spanish retirees (extensive pension indexation) are winners.&lt;/p&gt;
&lt;h3 id="q4-what-role-did-unconventional-fiscal-policy-play"&gt;Q4. What role did unconventional fiscal policy play?&lt;/h3&gt;
&lt;p&gt;Fiscal interventions reduced first-stage welfare losses by about one-fifth on average across countries and household types. Energy price caps were more important than transfers, especially in 2022 when caps were active in all countries. In France, interventions reduced the measured inflation rate by about 2 p.p.; in Italy interventions came ex-post via bonuses/transfers and so did not lower recorded inflation. Retirees benefited most, consistent with their higher energy/food shares and targeted measures. Government fiscal support outlays were approximately 1% of triennial GDP in all four countries, though in Italy and Spain a larger share (above 35% of costs) went to firms versus 14% (Germany) and 5% (France).&lt;/p&gt;
&lt;h3 id="q5-how-are-asset-prices-treated-and-what-are-the-estimated-elasticities"&gt;Q5. How are asset prices treated and what are the estimated elasticities?&lt;/h3&gt;
&lt;p&gt;House prices: a two-step approach—daily REIT (FTSE EPRA NAREIT Eurozone Residential) returns regressed on inflation surprises (beta = -3.995 on the swap surprise) on German HICP release days, then quarterly house-price returns (2006Q1-2023Q4) regressed on lagged REIT returns (delta = 0.035); the product beta x delta = -0.138 means a 10% inflation surprise lowers house prices ~1.38%. Stock and bond elasticities are larger and negative: -0.410 and -0.726 respectively. The asset-price channel is quantitatively negligible in welfare terms because house elasticity is small and stock/bond holdings are concentrated only at the very top of the consumption distribution. Housing and stocks are therefore not good inflation hedges when inflation has a large cost-push component.&lt;/p&gt;
&lt;h3 id="q6-what-about-wages-pensions-and-fiscal-drag-in-the-indirect-channel"&gt;Q6. What about wages, pensions, and fiscal drag in the indirect channel?&lt;/h3&gt;
&lt;p&gt;Nominal wage increases were modest, generating a welfare gain of only about 3% of disposable income against a direct loss on nominal wages of about 9.5%. Wages rose faster in France (sectoral agreements, over 4% vs 2-3% elsewhere) and for low-quintile German workers (large minimum-wage rise in October 2022). Pensions, being indexed to past inflation, grew more than wages in all four countries, so retirees gained substantially from the indirect channel, especially in Spain (pensions up 9.5% for most pensioners in 2023). However, fiscal drag (unindexed tax brackets in Italy and Spain) taxed away nominal gains—up to 2.5% for higher-quintile pensioners—whereas France and Germany had near-real-time bracket indexation, so drag was small. Higher ECB interest rates (tightening from July 2022) raised mortgage payments for young Spanish households with adjustable-rate mortgages, partly wiping out their NNP gains; the effect was small elsewhere (fixed-rate mortgages, limited deposit-rate pass-through).&lt;/p&gt;
&lt;h3 id="q7-what-does-the-sectoral-government-and-foreign-analysis-show"&gt;Q7. What does the sectoral (government and foreign) analysis show?&lt;/h3&gt;
&lt;p&gt;Using Euro Area Sector Financial Accounts (2017), the household sector holds positive net nominal positions (total NNP/triennial GDP: 0.28 Germany, 0.31 France, 0.35 Italy, 0.13 Spain), governments hold negative positions, and the foreign sector is a creditor against all except Germany. From the NNP channel alone the household sector lost (as % of triennial GDP): -3.8 Germany, -2.9 France, -3.9 Italy, -0.5 Spain; governments gained +3.5, +4.8, +7.5, +4.5; the foreign sector gained +0.3 in Germany but lost -1.9, -3.6, -3.9 in France, Italy, Spain. Adding fiscal drag (revenue), fiscal support cost (~1% GDP), higher pension cost (~1% GDP, peak 1.7% Italy), and higher government energy purchase cost, total government gains were: Germany -0.6 to +0.5 (roughly breaks even), France +1.3 to 2.1, Italy +4.5 to 5.1, Spain +1.6 to 2.2% of triennial GDP. Cross-country differences in government gains are driven mainly by the outstanding stock of public debt. Redistributing these government gains to households could substantially offset household losses.&lt;/p&gt;
&lt;h3 id="q8-how-does-this-paper-relate-to-and-differ-from-prior-work"&gt;Q8. How does this paper relate to and differ from prior work?&lt;/h3&gt;
&lt;p&gt;It applies the envelope-theorem money-metric approach used by Auclert (2019), Slacalek et al. (2020), Fagereng et al. (2022) and Del Canto et al. (2023), but studies a specific historical episode as an event study rather than identified shocks. It builds directly on Cardoso et al. (2022), who quantify the direct channel for Spain using bank-account data, by adding the other three channels (fiscal, indirect, long-run) and covering four countries. It contributes to the inflation-heterogeneity literature (Kaplan-Schulhofer-Wohl, Jaravel, Hobijn-Lagakos, Argente-Lee) by documenting inflation-rate differentials an order of magnitude larger than pre-pandemic US estimates, and confirms Doepke-Schneider (2006) that age is the key dimension via life-cycle net nominal positions. Unlike fully specified HANK models (Pugsley-Rubinton, Olivi et al., Yang), the sufficient-statistic approach cannot evaluate policy counterfactuals. Most contemporaneous euro-area papers stop at measuring differential inflation; this one quantifies full welfare.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-main-caveats-and-robustness-considerations"&gt;Q9. What are the main caveats and robustness considerations?&lt;/h3&gt;
&lt;p&gt;The framework is first-order: it assumes consumption and portfolio adjustments have only second-order welfare effects, which the authors flag as testable with high-frequency micro data they lacked. Survey-based (HFCS) nominal asset measures are 2-3 times smaller than financial-account measures because surveys undersample the very rich, so the Section 4 micro results best represent the population excluding the wealth top. Expenditure weights come from the 2015 HBS (judged stable using 2005/2015 HBS and credit-card evidence); inflation expectations (0.4-1.7%/year) come from Consensus Economics early 2021. A robustness note: assuming 0.75%/year trend productivity growth (so part of nominal wage rises reflects trend, not catch-up) increases welfare losses by roughly 1.5% of disposable income. The retiree/young housing trade is modeled as selling/buying one tenth of housing (3/30 over the 3-year long run). The conclusion notes the episode coincided with high pandemic excess savings that cushioned purchasing-power erosion, and that the inflation tax effectively redistributes from retirees to the young, partially offsetting future fiscal adjustment.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;</description></item><item><title>"Compensate the Losers?" Economic Policy and the Origins of U.S. Partisan Realignment</title><link>https://macropaperwarehouse.com/papers/compensate-the-losers-economic-policy-and-the-origins-of-u.s.-partisan-realignment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/compensate-the-losers-economic-policy-and-the-origins-of-u.s.-partisan-realignment/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Why have less-educated voters in the United States abandoned the Democratic Party over recent decades? The paper argues that the Democratic Party&amp;rsquo;s evolution on &lt;em&gt;economic policy&lt;/em&gt; — specifically its retreat from &amp;ldquo;predistribution&amp;rdquo; — is a central, previously understudied driver of partisan realignment by education.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conceptual Framework.&lt;/strong&gt; The authors distinguish between two categories of egalitarian economic policy: (1) &lt;em&gt;predistribution&lt;/em&gt; — policies that alter the pre-tax-and-transfer earnings distribution, including job guarantees, minimum wage increases, union support, and protectionist trade policies (following Hacker 2011); and (2) &lt;em&gt;redistribution&lt;/em&gt; — taxes and transfers. The paper&amp;rsquo;s central claim is that these two types of policy have sharply different educational gradients among voters, and that the Democratic Party moved away from predistribution beginning in the 1970s, triggering educational realignment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Methodology.&lt;/strong&gt; The authors harmonize over 1,000 surveys (N ≈ 2.2 million observations) spanning 1942–2020, drawn from Gallup, ANES, GSS, CCES, and historical survey archives housed at iPoll/Cornell. Education is translated into a common metric (adjusted years of schooling) using Census data, controlling for sex, race, year, and birth cohort to address the changing selectivity of educational categories over time. Congressional roll-call data come from the Comparative Agendas Project (CAP). Campaign finance data come from FEC filings, Congressional hearing records, and watchdog sources. DLC membership data are compiled from official Democratic Leadership Council records (available for 1985, 1986, 1991, 1993, and 1997 onward) and DLC-aligned Congressional caucus lists. House election returns are taken from King and Palmquist (1997) at the minor-civil-division-group (MCDG) level (~60 units per Congressional district), matched to 1980 Census demographic data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Voter preferences (demand side):&lt;/em&gt; The educational gradient for predistribution is large and negative: averaged across the four predistribution questions (job guarantee, minimum wage, union support, trade protection), each additional year of education reduces support by 0.044 standard deviations (p &amp;lt; 0.001). A college graduate relative to a high school graduate supports predistribution 0.176 standard deviations less — equivalent to roughly half the average Democrat-Republican gap in predistribution support (which is 0.34 standard deviations). This gradient has been stable since at least the 1940s. By contrast, the educational gradient for redistribution (higher taxes on the rich, views on own taxes, welfare spending) is close to zero (summary β = 0.004, not distinguishable from zero in the full sample). The difference between the two gradients is statistically significant (p &amp;lt; 0.001). These results replicate in white-only samples. Notably, the educational gradient on social issues — measured across nine questions on racial attitudes, gender roles, sexual norms — is positive (more education predicts more liberal positions) but has been largely &lt;em&gt;stable&lt;/em&gt; since the 1940s, not increasing, conditional on the long-run sample.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Party supply (supply side):&lt;/em&gt; Before 1976, predistribution topics accounted for roughly one-quarter of Democratic House roll-call votes when Democrats controlled the chamber. After 1976 (taking Jimmy Carter&amp;rsquo;s presidency as the start of the &amp;ldquo;New Democrat&amp;rdquo; era), this share falls by approximately nine to ten percentage points, while the redistribution share of votes holds steady. Between 1968 and 1980, the union share of total PAC donations to Democratic Congressional candidates falls from approximately 90 percent to 40 percent, coincident with 1970s campaign finance reforms that placed union and corporate PACs on equal legal footing and allowed corporations to exploit their naturally deeper pockets. Corporate PAC share of Democratic donations correspondingly rises from approximately 10 percent to 45 percent over the same period. In individual contributions to primary elections (data beginning in 1980), Democratic primaries rely on increasingly more-educated census tracts relative to Republican primaries; by 2018 Democratic primaries are financed from census tracts averaging 0.41 more years of education than Republican primaries (against a within-year standard deviation of 1.56 years).&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The New Democrat/DLC faction:&lt;/em&gt; The authors identify the anti-predistribution faction through official DLC membership records and aligned caucus lists. DLC membership as a share of Democratic House seats grows from near zero in the mid-1970s to approximately half by the early 2000s. Roll-call voting analysis (N = 3,428,405 vote-observations) shows DLC members are more conservative than other Democrats overall, and &lt;em&gt;especially&lt;/em&gt; so on predistribution: for a 10-percentage-point increase in the share of Republicans voting for a bill, the probability a DLC member votes in favor increases 36 percent more on predistribution bills than on other bills. DLC members show no differential conservatism on redistribution. They are also significantly more socially conservative — more likely than other Democrats to support the Defense of Marriage Act (by 16 pp), the Partial-Birth Abortion Ban (by 7 pp), and restrictive immigration bills (by 10 pp). DLC candidates receive significantly less from labor PACs and significantly more from corporate PACs, and draw their out-of-district individual donations from census tracts averaging more than 0.1 years more educated than non-DLC Democrats.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Voter reaction and the inflection point:&lt;/em&gt; Using the N ≈ 2.2 million partisan identification dataset, the authors estimate a structural break in the education-party identification gradient. From the 1940s through the mid-1970s, each additional year of education reduces the probability of identifying as a Democrat by approximately 3 percentage points. A Chow breakpoint test identifies 1976 as the inflection point. Since 1976, the gradient steadily rises; by 2000 it reaches zero; and today (as of the sample period end ~2020) each additional year of education &lt;em&gt;increases&lt;/em&gt; Democratic identification by approximately 3 percentage points — an almost exact reversal. The breakpoint for Republican identification occurs later, in 1992, consistent with the Democratic agenda changing first. A Gallup prosperity question (&amp;ldquo;which party will better keep the country prosperous?&amp;rdquo;) shows a parallel pattern: controlling for views on parties&amp;rsquo; economic performance explains approximately 44 percent of partisan realignment, interpreted as an upper bound on economic policy&amp;rsquo;s contribution.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Factional tests — hypothetical elections and actual results:&lt;/em&gt; In hypothetical general-election matchups from 1972–1992 Democratic primaries (in which most contests pitted a &amp;ldquo;New Democrat&amp;rdquo; against an &amp;ldquo;Old Democrat&amp;rdquo;), a voter with a college degree is roughly 3 percentage points &lt;em&gt;more&lt;/em&gt; likely to vote Democratic when the candidate is a New Democrat rather than an Old Democrat. In 1980s actual House elections using MCDG-level data, DLC candidates out-perform other Democrats in more educated neighborhoods by a magnitude large enough to erase approximately 90 percent of the general Democratic underperformance in highly educated areas. Combining these estimates, the party&amp;rsquo;s shift toward the DLC accounts for a lower bound of approximately 20 percent, and an upper bound (from the prosperity question) of approximately 50 percent, of educational realignment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; The analysis focuses on the United States, 1942–2015 (with some post-2015 discussion in the conclusion). The faction analysis focuses on the Democratic side; Republican faction changes are discussed but not the primary focus. The paper is explicit that between 20–50 percent of realignment is explained, leaving room for other factors, including social issues. The analysis ends mostly before 2016 to avoid complications from the closure of the DLC in 2011 and shifting post-2010 party dynamics.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-central-conceptual-innovation-and-how-does-it-differ-from-prior-realignment-research"&gt;Q1. What is the paper&amp;rsquo;s central conceptual innovation, and how does it differ from prior realignment research?&lt;/h3&gt;
&lt;p&gt;The paper separates egalitarian economic policies into &amp;ldquo;predistribution&amp;rdquo; (pre-tax-and-transfer market interventions such as minimum wages, job guarantees, union support, and protectionism) and &amp;ldquo;redistribution&amp;rdquo; (taxes and transfers) and shows these two types have sharply different educational gradients. Prior work typically aggregated all economic policies into a single index, which the authors argue masks essential heterogeneity. By documenting that the educational gradient is large and negative for predistribution but close to zero for redistribution — a pattern stable since the 1940s — the paper reframes the &amp;ldquo;voting against economic interest&amp;rdquo; puzzle: less-educated voters leaving the Democratic Party may be responding rationally to changes in the supply of the type of economic policy they actually prefer.&lt;/p&gt;
&lt;h3 id="q2-how-large-and-stable-is-the-educational-gradient-on-predistribution-and-how-does-it-compare-to-social-issues"&gt;Q2. How large and stable is the educational gradient on predistribution, and how does it compare to social issues?&lt;/h3&gt;
&lt;p&gt;The average coefficient on adjusted years of schooling across the four predistribution questions is -0.044 (p &amp;lt; 0.001), stable over eight decades. A four-year difference in education (high school vs. college) shifts an individual&amp;rsquo;s support for predistribution by 0.176 standard deviations in the conservative direction — about half the average Democrat-Republican gap in predistribution support (0.34 standard deviations). For social issues, the summary gradient is positive (+0.028, p &amp;lt; 0.001 for the full sample), but this gradient has been largely &lt;em&gt;stable&lt;/em&gt; since the 1940s across nine social issue questions, not increasing over time. This stability undermines the interpretation that rising social liberalism among the educated is a new phenomenon driving realignment, at least through the supply of parties&amp;rsquo; social positions.&lt;/p&gt;
&lt;h3 id="q3-what-happened-to-predistribution-as-a-share-of-the-democratic-house-agenda-after-the-1970s"&gt;Q3. What happened to predistribution as a share of the Democratic House agenda after the 1970s?&lt;/h3&gt;
&lt;p&gt;Using the Comparative Agendas Project classification, predistribution topics (labor regulation, industrial policy, public works, trade) accounted for roughly one-quarter of all House roll-call votes during years Democrats controlled the Speakership before 1977. After 1977, this share falls by approximately 9–10 percentage points (a decline of nearly half from its pre-1977 share), and the decline is statistically significant (p &amp;lt; 0.001). The redistribution share of votes holds essentially constant. Party platform data from Hopkins et al. (2022) show a sharp decline in Democratic use of terms like &amp;ldquo;minimum wage,&amp;rdquo; &amp;ldquo;full employment,&amp;rdquo; and labor-relations language beginning in the 1970s and 1980s, while Republican platforms use these terms sparingly throughout.&lt;/p&gt;
&lt;h3 id="q4-how-did-1970s-campaign-finance-reforms-change-the-financial-composition-of-the-democratic-party"&gt;Q4. How did 1970s campaign finance reforms change the financial composition of the Democratic Party?&lt;/h3&gt;
&lt;p&gt;Before the early 1970s, unions enjoyed substantially more freedom than corporations under separate legal regimes governing PAC donations; mid-1970s reforms placed them on equal legal footing, enabling corporations to exploit their deeper pockets. The union share of total PAC donations to Democrats fell from approximately 90 percent in 1968 to approximately 40 percent by 1980, while the corporate share rose from approximately 10 percent to 45 percent. For Republicans, both series barely changed: unions had never donated substantially to the GOP, and the corporate share rose only modestly (from approximately 70 to 80 percent). The authors note the rapid decline cannot be attributed to falling union density in the economy, since both union and corporate PAC donations grew in absolute terms during this period; the relative shift was the result of the regulatory change.&lt;/p&gt;
&lt;h3 id="q5-who-are-the-new-democrats--dlc-and-when-did-they-emerge"&gt;Q5. Who are the &amp;ldquo;New Democrats&amp;rdquo; / DLC, and when did they emerge?&lt;/h3&gt;
&lt;p&gt;The DLC officially operated from 1985 to 2011, but members who would join it began entering Congress in large numbers in the 1970s (&amp;ldquo;Watergate Babies&amp;rdquo; of 1974, &amp;ldquo;Atari Democrats&amp;rdquo;). The DLC grew to approximately half of all Democratic House seats by the early 2000s. Members were drawn from suburban, affluent districts; their founder Al From explicitly criticized all four predistribution policies the paper studies (minimum wage, job guarantees, unions, and protectionism). The breakpoint test on DLC share in Congress identifies 1975 as the pivotal year — one year before the 1976 inflection point in partisan identification.&lt;/p&gt;
&lt;h3 id="q6-how-do-dlc-members-vote-differently-from-other-democrats-and-how-is-this-differential-conservatism-distributed-across-policy-types"&gt;Q6. How do DLC members vote differently from other Democrats, and how is this differential conservatism distributed across policy types?&lt;/h3&gt;
&lt;p&gt;In roll-call regressions (N = 3,428,405 observations, with roll-call fixed effects), a 10 pp increase in the Republican vote share for a bill increases the probability a DLC member votes in favor by 1.48 pp more than for other Democrats (baseline result for all bills). For predistribution-classified bills, this excess alignment with Republicans is 36 percent larger than for non-predistribution bills. Crucially, DLC members are no more conservative than other Democrats on redistribution-classified votes (the interaction with redistribution is near zero and insignificant). DLC members are also differentially more conservative on social issues, a result that proves useful in separating economic from social-issue explanations of realignment.&lt;/p&gt;
&lt;h3 id="q7-do-dlc-members-finance-differently-from-other-democrats"&gt;Q7. Do DLC members finance differently from other Democrats?&lt;/h3&gt;
&lt;p&gt;Yes. In primary elections, DLC candidates receive approximately 9.7 pp less of their PAC financing from labor unions and approximately 6.7 pp more from corporate PACs (with state fixed effects) relative to non-DLC Democrats. Out-of-district individual contributions to DLC primary candidates come from census tracts averaging more than 0.1 years more educated than those for non-DLC Democrats, while within-district contributions show no significant difference (0.060 years, insignificant). This pattern suggests educated out-of-district donors, rather than local constituency demands, drive DLC candidates&amp;rsquo; anti-predistribution orientation.&lt;/p&gt;
&lt;h3 id="q8-when-precisely-did-educational-realignment-in-democratic-party-identification-begin-and-what-does-the-inflection-point-analysis-show"&gt;Q8. When precisely did educational realignment in Democratic party identification begin, and what does the inflection-point analysis show?&lt;/h3&gt;
&lt;p&gt;Using N ≈ 2.2 million observations from 1,006 surveys, a Bai-Perron breakpoint test on the year-by-year education gradient in Democratic party identification identifies 1976 as the inflection point (with robustness to alternative specifications yielding breakpoints of 1978–1980 for white-only samples and unadjusted years of schooling). Before 1976, each additional year of education reduces the probability of Democratic identification by approximately 3 percentage points (a stable, significantly negative relationship since the 1940s). After 1976, the gradient steadily rises; it reaches zero around 2000 and today is approximately +3 percentage points per year of education — nearly an exact reversal of the baseline. The corresponding Republican inflection point occurs in 1992, about 16 years later, consistent with the Democratic Party&amp;rsquo;s agenda changing first.&lt;/p&gt;
&lt;h3 id="q9-how-do-hypothetical-presidential-matchup-surveys-test-the-dlc-mechanism"&gt;Q9. How do hypothetical presidential matchup surveys test the DLC mechanism?&lt;/h3&gt;
&lt;p&gt;The authors identify six Democratic primaries from 1972–1992 where a &amp;ldquo;New Democrat&amp;rdquo; and an &amp;ldquo;Old Democrat&amp;rdquo; were the top two contenders (e.g., Hart vs. Mondale in 1984, Clinton vs. Brown in 1992). Gallup and other surveys asked all respondents — regardless of party — whom they would vote for if either the New or the Old Democrat faced the eventual Republican nominee. A voter with a college BA is approximately 3 percentage points more likely to vote for the Democrat when the candidate is a New Democrat versus an Old Democrat (the &amp;ldquo;difference in differences&amp;rdquo; of hypothetical vote shares). This holds after controlling for state × election fixed effects and in five of the six election cycles studied (the 1976 exception is attributed to Mo Udall&amp;rsquo;s low name recognition, with 28 percent of respondents unfamiliar with him in a May 1976 poll). The result is attenuated but remains marginally significant when excluding non-white respondents, consistent with New Democrats&amp;rsquo; success with white voters due in part to their more conservative civil rights positioning.&lt;/p&gt;
&lt;h3 id="q10-what-do-actual-house-election-results-mcdg-level-data-show-about-dlc-electoral-performance-by-neighborhood-education"&gt;Q10. What do actual House election results (MCDG-level data) show about DLC electoral performance by neighborhood education?&lt;/h3&gt;
&lt;p&gt;Using 1980s House returns at the MCDG level (~60 neighborhoods per Congressional district), the authors regress Democratic vote share on neighborhood years of education interacted with a DLC candidate indicator, with Congressional district fixed effects. More-educated neighborhoods generally depress Democratic vote share (reflecting the still-negative overall educational gradient in the 1980s), but DLC candidates dramatically out-perform other Democrats in educated areas: the interaction coefficient is positive and significant, and its magnitude is large enough to erase approximately 90 percent of the general Democratic underperformance in highly educated neighborhoods. This result is robust to including District × Year fixed effects (so the identification comes from within-election, cross-neighborhood variation) and to adding controls for share white and share under age 35.&lt;/p&gt;
&lt;h3 id="q11-how-much-of-educational-realignment-can-the-papers-mechanism-account-for-and-how-is-this-calculated"&gt;Q11. How much of educational realignment can the paper&amp;rsquo;s mechanism account for, and how is this calculated?&lt;/h3&gt;
&lt;p&gt;Two bounding estimates are provided. Upper bound (~44–50%): controlling for a respondent&amp;rsquo;s view on which party is better for economic prosperity (from Gallup since 1950) explains approximately 44 percent of the change in the education-party identification gradient (specifically, the total difference in the unconditional gradient between the 1948–1967 baseline and 2001–2020 is 2.411 pp per year of schooling; after controlling for the prosperity question, the unexplained residual is 1.342 pp, leaving a share explained of 44.3 percent). Lower bound (~20%): the difference in the education gradient between matchups involving New versus Old Democrats in Table 4 (~0.75 pp) divided by the total realignment shift (~4 pp from pre-1976 to post-2008 for presidential voting) implies the faction shift accounts for at least approximately one-fifth of realignment. The authors interpret these as bounds because the prosperity question may partly capture party identification itself (upper bound concern), while the hypothetical matchup estimate misses the broader ideological shift not captured in a single election (lower bound).&lt;/p&gt;
&lt;h3 id="q12-can-social-issues-civil-rights-realignment-or-republican-changes-better-explain-the-1970s-inflection-point"&gt;Q12. Can social issues, Civil Rights realignment, or Republican changes better explain the 1970s inflection point?&lt;/h3&gt;
&lt;p&gt;Three alternative explanations are addressed. (1) &lt;em&gt;Civil Rights:&lt;/em&gt; Regional analysis shows that educated white Southerners &lt;em&gt;left&lt;/em&gt; the Democrats in the 1940s–1960s (not the 1970s), consistent with their realignment being driven by Democrats&amp;rsquo; liberal turn on civil rights rather than economic policy. After the 1960s, the South follows all other regions in the pace of educational realignment. (2) &lt;em&gt;Republican changes:&lt;/em&gt; The Republican party identification inflection point occurs in 1992, about 16 years after the Democratic inflection in 1976. Reagan elections in 1980 and 1984 do not appear to have differentially attracted less-educated voters (the &amp;ldquo;Reagan Democrats&amp;rdquo; were not differentially less educated). (3) &lt;em&gt;Social issues:&lt;/em&gt; The New Democrats were actually &lt;em&gt;more&lt;/em&gt; socially conservative than other Democrats (more likely to vote for DOMA, anti-abortion bills, restrictive immigration legislation), yet they disproportionately attracted educated voters. This internal inconsistency rules out a pure social-issues explanation for why educated voters preferred the DLC faction. (4) &lt;em&gt;Religion:&lt;/em&gt; Flexibly controlling for religious affiliation explains essentially none of partisan realignment (Appendix Figure A.24).&lt;/p&gt;
&lt;h3 id="q13-what-is-the-role-of-out-of-district-individual-donors-in-shifting-democratic-party-positions"&gt;Q13. What is the role of out-of-district individual donors in shifting Democratic Party positions?&lt;/h3&gt;
&lt;p&gt;Out-of-district primary donors are analytically important because they influence candidate supply without being able to vote in the election, isolating the &amp;ldquo;within-party&amp;rdquo; financial influence of educated supporters. By 1980, out-of-district primary donors to Democratic candidates already come from census tracts more educated than those for Republican candidates, even as local Democratic voters and within-district donors remain less educated than Republican counterparts. Democratic candidates also receive a substantially higher share of out-of-district contributions than Republican candidates — by almost 10 percentage points (Appendix Table A.7). Out-of-district donors thus represent a channel through which educated, anti-predistribution preferences are transmitted into the Democratic Party&amp;rsquo;s candidate supply before the electoral realignment is visible in vote totals.&lt;/p&gt;
&lt;h3 id="q14-are-predistribution-policies-becoming-less-popular-overall-which-might-independently-push-democrats-away-from-them"&gt;Q14. Are predistribution policies becoming less popular overall, which might independently push Democrats away from them?&lt;/h3&gt;
&lt;p&gt;The paper tests this alternative in Appendix Table A.9 and finds no evidence that predistribution has become less popular relative to redistribution over time. Predistribution appears on average more popular than redistribution across the sample period. If anything, support for predistribution has held steady or slightly risen relative to redistribution over time, conditional on the paper&amp;rsquo;s survey harmonization. The stability of the educational gradient (shown in Appendix Table A.10 to be unchanged even using educational rank within cohort rather than raw years of schooling) further suggests the negative education-predistribution relationship is a relative, not absolute, phenomenon — consistent with rising average education and stable preferences by education rank.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Predistribution:&lt;/strong&gt; Policies that aim to change the distribution of earnings or income &lt;em&gt;before&lt;/em&gt; taxes and transfers are applied. In this paper, this comprises government job guarantees, minimum wage increases, support for unions and collective bargaining, and protectionist trade policies. Distinguished from redistribution in that it operates on pre-tax market income rather than post-tax outcomes. The paper uses this term following Hacker (2011): &amp;ldquo;a focus on market reforms that encourage a more equal distribution of economic power and rewards even before government collects taxes or pays out benefits.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Redistribution:&lt;/strong&gt; Policies that change post-market income through the tax and transfer system, including higher taxes on the rich, views on own tax burden, prioritization of tax cuts, and transfers to the poor (welfare spending). In the paper&amp;rsquo;s usage, redistribution is analytically distinct from predistribution and has a near-zero educational gradient, in contrast to predistribution&amp;rsquo;s strongly negative gradient.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Educational Gradient:&lt;/strong&gt; The coefficient on adjusted years of schooling in a regression of an outcome variable (policy preference or partisan identification) on education, estimated separately by time period. The paper&amp;rsquo;s core finding is that the educational gradient for predistribution is stably negative (approximately -0.044 per year of schooling over the full sample), while the gradient for redistribution is close to zero, and the gradient for Democratic party identification shifts from approximately -0.03 to +0.03 per year of schooling between the 1940s and 2020.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;New Democrats / DLC (Democratic Leadership Council):&lt;/strong&gt; An explicitly anti-predistribution faction within the Democratic Party, identified through official DLC membership records and affiliated Congressional caucus lists. Founded formally in 1985 (operating through 2011), the DLC arose in part from the &amp;ldquo;Watergate Babies&amp;rdquo; cohort of 1974. DLC members were more conservative than other Democrats &lt;em&gt;especially&lt;/em&gt; on predistribution and social issues, relying differentially on corporate PACs and educated out-of-district donors. The paper treats DLC membership as a proxy for an anti-predistribution faction that gained bargaining power within the Democratic Party from the 1970s onward.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Adjusted Years of Schooling (AdjYearsEduc):&lt;/strong&gt; The paper&amp;rsquo;s harmonized education variable across more than 1,000 surveys spanning eight decades. Because raw educational categories change over time and represent different selectivity (e.g., in 1940 only one-quarter of adults had completed twelfth grade, versus nearly 90 percent today), the authors use Census microdata to predict years of schooling as a function of self-reported educational category, sex, race, year, and birth cohort in ten-year bins. This provides a common unit of measurement across surveys with incompatible category systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inflection Point (1976):&lt;/strong&gt; The structural break in the trend of the education-Democratic identification gradient, estimated using Bai-Perron (1998) methods on N ≈ 2.2 million observations. The data select 1976 as the year at which the previously stable negative gradient begins its upward trajectory. The corresponding Republican inflection point occurs in 1992. The paper argues that identification of this inflection point — not previously documented in the realignment literature — is made possible only by the large historical dataset assembled.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Minor Civil Division Group (MCDG):&lt;/strong&gt; The granular geographic unit used in the House election analysis for the 1980s, with approximately sixty MCDGs per Congressional district. Matched to 1980 Census demographic data to assign average years of education. Used to test whether DLC candidates out-perform other Democrats in more-educated neighborhoods, within the same Congressional district and election year, to address the concern that DLC candidates sort into more-educated districts.&lt;/p&gt;</description></item><item><title>(Not) Thinking About the Future: Financial Information and Maternal Labor Supply</title><link>https://macropaperwarehouse.com/papers/not-thinking-about-the-future-financial-information-and-maternal-labor-supply/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/not-thinking-about-the-future-financial-information-and-maternal-labor-supply/</guid><description>&lt;p&gt;This paper investigates whether information constraints — rather than fully forward-looking choices — contribute to mothers&amp;rsquo; reduced labor supply after childbirth, a key driver of gender inequality. The authors deploy two complementary methods in Switzerland: a representative descriptive survey of Swiss mothers aged 25–50, and a large-scale randomized controlled trial (RCT) among approximately 2,400 female public school teachers with children who work part-time.&lt;/p&gt;
&lt;p&gt;The descriptive survey first establishes that long-term financial factors are not top of mind for mothers making labor supply decisions: only about 11% of mothers spontaneously mention pensions or long-term career considerations when asked about their post-childbirth employment choices, compared to roughly half who mention child or own well-being. Beyond salience, the survey documents substantial misperceptions: 62% of women over-estimate pension receipt under part-time work by more than 10%, and a similar share believes wage growth under low part-time hours (40% FTE) is at least as high as under 80% employment. The authors label mothers with overly optimistic beliefs on both dimensions &amp;ldquo;cost-unaware&amp;rdquo;; 42% of the sample qualifies. Cost-unawareness is more prevalent among less-educated mothers and correlates with less financial interest and more gender-conservative attitudes.&lt;/p&gt;
&lt;p&gt;The RCT tests whether providing objective, individualized information shifts financial planning and labor supply. Teachers in treatment schools (two-thirds of all schools) were individually randomized into a treatment group viewing an informational video about the long-run earnings, pension, and life-event consequences of sustained part-time employment, plus access to a Future Calculator tool, or a placebo video on unrelated financial topics. The two-stage randomization (school-level first, then individual within treated schools) allows identification of both direct treatment effects and spillovers. Outcomes are measured in a Wave 1 post-video survey, a follow-up survey two months later, and linked administrative personnel records from the Department of Education one year post-intervention.&lt;/p&gt;
&lt;p&gt;Main findings: treated teachers are 31.26 percentage points (58% over the pure control mean) more likely to correctly rank the relative magnitude of long- versus short-term financial factors. Demand for financial planning tools rises by 0.39 standard deviations (SD) overall and by 0.31 SD among cost-unaware women specifically. In terms of stated labor supply plans, the treatment raises planned employment for the next academic year by 1.69 percentage points (ppt) in the full sample and by 4.95 ppt (9% over the pure control mean) among cost-unaware women. These plan effects persist two months later for cost-unaware women but fade for the full sample.&lt;/p&gt;
&lt;p&gt;Critically, stated plans translate into verified behavior: linked administrative data one year post-intervention show that cost-unaware teachers increase their contracted employment level by 3.87 ppt, or 7% over the pure control mean of 53.30% FTE. Cost-aware and overly pessimistic women do not reduce their labor supply upon learning they are better off than feared, an asymmetry consistent with agents responding more to perceived losses than gains. If the 3.87 ppt increase were sustained from age 40 onward, cost-unaware teachers would accumulate an additional 130,000 CHF in lifetime income and 40,000 CHF in pension wealth, shrinking the gender gap in lifetime income and pension receipt among teachers by approximately 18% each.&lt;/p&gt;
&lt;p&gt;The paper is scoped to Swiss female public school teachers — a population with linear pay scales, no part-time promotion penalty, and relatively low adjustment barriers — meaning the measured lifetime earnings and pension losses likely represent a lower bound relative to other occupations. Short-term RCT findings replicate among a sample of pregnant women in the general Swiss population, and the paper argues that similar labor supply adjustment magnitudes are feasible for a broader segment of part-time working mothers.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question and why does it matter?
A: The paper asks whether mothers&amp;rsquo; post-childbirth reduction in labor supply is partly driven by information constraints — specifically, whether mothers fail to account for the full long-term financial consequences of working reduced hours. This matters because if the child penalty partly reflects uninformed choices rather than deliberate tradeoffs, standard policy tools (parental leave, childcare subsidies) may underperform precisely because their long-term financial benefits are not internalized.&lt;/p&gt;
&lt;p&gt;Q: How prevalent is cost-unawareness among Swiss mothers?
A: 62% of mothers in the descriptive survey over-estimate pension receipt under part-time work by more than 10%, a similar share believes wage growth under low part-time (40% FTE) is at least as high as under 80% employment, and 42% are overly optimistic on both dimensions simultaneously. Cost-unawareness follows an education gradient: 77% of low-education women over-estimate pension receipt versus 51% of high-education women.&lt;/p&gt;
&lt;p&gt;Q: What share of mothers spontaneously considers long-term financial factors when deciding on their labor supply?
A: Only about 11% of mothers mention any long-term financial factor (pensions, financial independence, long-term career considerations) in open-ended responses; the share is similarly low across education groups (6% low, 12% mid, 13% high). About 50% mention child or own well-being; roughly 30% raise short-term financial factors such as current childcare costs.&lt;/p&gt;
&lt;p&gt;Q: What are the actual long-term financial stakes of the average female teacher&amp;rsquo;s part-time employment pattern in Switzerland?
A: Compared to full-time employment, the average female teacher&amp;rsquo;s employment trajectory produces a 35% reduction in potential lifetime earnings (approximately 3.34 million CHF versus 5.12 million CHF). Monthly pension receipt under the part-time scenario is 31% lower overall and 43% lower from the occupational second-pillar scheme specifically — a gap comparable to the average 47.5% gender pension gap observed in the second pillar in Switzerland in 2024.&lt;/p&gt;
&lt;p&gt;Q: How was the RCT designed and what populations were included?
A: The study recruited 2,359 part-time working mothers employed as public school teachers in a German-speaking Swiss canton. A two-stage randomization assigned two-thirds of schools to treatment schools (within which teachers were individually randomized 50/50 to treatment or spillover control) and one-third to pure control schools. This design allows estimation of direct treatment effects and spillover effects. The intervention was timed to precede December–January, the period when teachers communicate their preferred employment levels for the next school year.&lt;/p&gt;
&lt;p&gt;Q: What was the treatment intervention?
A: Treated teachers watched an informational video following a representative female teacher considering an employment-level increase, covering the impact of part-time work on lifetime earnings, monthly pension receipt, and financial exposure after adverse events such as divorce; it also benchmarked these magnitudes against childcare costs. Treated teachers additionally received individualized access to the Future Calculator, an online projection tool developed with a Swiss bank, calibrated to teachers&amp;rsquo; deterministic salary and pension schedules.&lt;/p&gt;
&lt;p&gt;Q: Did treated teachers understand and retain the treatment information?
A: Yes. Treated teachers were 31.26 ppt (58% over the pure control mean) more likely immediately after the intervention to correctly rank long- versus short-term financial factors in a vignette. Two months later, the treatment group remained significantly more likely to apply the information correctly (22.63 ppt higher), indicating the knowledge was not short-lived.&lt;/p&gt;
&lt;p&gt;Q: How did demand for financial planning tools respond to the treatment?
A: The treatment raised a financial information/tools index by 0.39 SD overall. For cost-unaware women specifically, demand for financial tools rose by 0.31 SD; cost-aware and pessimistic women showed no significant change. There was no significant average treatment effect on sign-up for an incentivized financial consultation.&lt;/p&gt;
&lt;p&gt;Q: How large were the labor supply plan effects in the survey, and did they persist?
A: For the full sample, treated teachers planned a 1.69 ppt higher employment level for the next school year immediately after the treatment, and 3.13 ppt higher in 10 years. For cost-unaware women, the short-run planned increase was 4.95 ppt (9% over the pure control mean of about 55%), and plans for 5 and 10 years into the future rose by approximately 4 ppt (6–7% over the mean). The short-run effects for cost-unaware women persisted to the two-month follow-up, while full-sample short-run effects faded.&lt;/p&gt;
&lt;p&gt;Q: What do the linked administrative data show about actual labor supply one year post-intervention?
A: Cost-unaware women in the treatment group increased their contracted employment level by 3.87 ppt relative to the pure control group (7% over the pure control mean of 53.30% FTE), closely matching the planned increase stated immediately after the treatment. Cost-aware women and the full sample showed no statistically significant shift in actual hours.&lt;/p&gt;
&lt;p&gt;Q: What asymmetry did the authors observe between cost-unaware and cost-aware women?
A: Cost-unaware (overly optimistic) women increased their labor supply upon learning the true financial costs; cost-aware and overly pessimistic women did not reduce their labor supply upon learning they were better off than expected. The authors interpret this as consistent with agents responding more to perceived losses (bad news for cost-unaware women) than to gains (good news for pessimistic women), and with cost-aware women already having incorporated the financial logic into their decisions even without precise estimates.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated lifetime impact of the observed labor supply adjustment?
A: If cost-unaware teachers maintain the 3.87 ppt employment increase from age 40 to retirement, they accumulate an additional 130,000 CHF in lifetime income and 40,000 CHF in pension wealth on average. This would reduce the gender gap in both lifetime income and pension receipt among teachers by approximately 18% each.&lt;/p&gt;
&lt;p&gt;Q: What emotional and social mechanisms did the paper document?
A: The treatment initially produced significantly negative emotional responses (−0.41 SD on an emotions index overall; −0.68 SD for cost-unaware women), consistent with cognitive dissonance from information conflicting with prior beliefs. Two months later, the treatment group reported feeling more in control and less stressed, and cost-unaware women returned to a neutral emotional baseline. Treated women were also 19.61 ppt more likely to have discussed the topic with anyone, with the largest effect on conversations with partners or family.&lt;/p&gt;
&lt;p&gt;Q: Did the treatment affect household-level labor supply — specifically, did partners reduce their hours?
A: No. The authors found no evidence that partners of cost-unaware women planned to work less in response to the treatment, and women did not plan to adjust future fertility. This suggests the observed hours increase by treated cost-unaware women was not offset by partner adjustments within the household.&lt;/p&gt;
&lt;p&gt;Q: Were there social spillover effects within schools?
A: Treated teachers were 11.59 ppt more likely to report having discussed the video with colleagues. Two months later, cost-unaware control teachers in treated schools (the spillover group) showed some evidence of absorbing the general treatment message and adjusting short-term labor supply plans upward, and a noisy increase in actual employment of roughly one-third the magnitude of the direct treatment effect, though these estimates were imprecise.&lt;/p&gt;
&lt;p&gt;Q: Why might cost-unaware women be uninformed in the first place?
A: In both the descriptive survey and the RCT sample, cost-unaware women lean more gender-conservative in their attitudes and report less interest in financial topics. The authors interpret this as suggesting a lack of information (rather than mere salience or forgetting) drives cost-unawareness, implying that passive information delivery through employers or pension funds could be effective.&lt;/p&gt;
&lt;p&gt;Q: What constraints to labor supply adjustment did the authors explore?
A: In a hypothetical scenario exercise, the scenario producing the largest desired employment increase for both treatment and control groups was if the partner were more engaged (roughly double the adjustment relative to a scenario of higher pay for additional hours). The treatment group adjusted their desired employment level by an additional 0.62–2.03 ppt relative to pure control across all scenarios except relaxing conservative gender norms.&lt;/p&gt;
&lt;p&gt;Q: How generalizable are the findings beyond the teacher sample?
A: The short-term RCT findings replicated among a sample of pregnant women in the general Swiss population. The authors also document that potential net gains from increasing labor supply — net of additional childcare costs — are large for the broader population of part-time working Swiss mothers, supporting feasibility of similar-magnitude adjustments outside teaching. The teaching context likely represents a lower bound for lifetime earnings and pension losses in other professions due to the absence of a part-time promotion penalty in teaching.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications?
A: The findings suggest that default exposure to individualized financial information about the long-term costs of part-time work — delivered by employers, pension funds, or the state — could improve decision quality and labor supply. More broadly, the results imply that policies designed to increase female labor supply (parental leave reforms, childcare subsidies) may underperform if mothers do not fully internalize the financial benefits of additional hours; ensuring that families solve the correct optimization problem is a precondition for unlocking the full potential of such policies.&lt;/p&gt;
&lt;p&gt;Child Penalty: The large and persistent reduction in women&amp;rsquo;s labor force participation and income following the birth of a first child, identified in the paper as the key driver of remaining gender inequality in the labor market in industrialized countries and a source of profound life-cycle financial consequences including reduced lifetime earnings and pension savings.&lt;/p&gt;
&lt;p&gt;Cost-Unaware: The authors&amp;rsquo; term for women who hold overly optimistic expectations about the financial consequences of part-time work — specifically, who over-estimate pension receipt under low part-time employment by more than 10% and who believe wage growth under low part-time is at least as high as under higher employment levels. In the descriptive survey 42% of mothers qualify on both dimensions.&lt;/p&gt;
&lt;p&gt;Future Calculator: An online individualized projection tool developed by the authors in cooperation with a Swiss bank, calibrated to teachers&amp;rsquo; deterministic salary and pension schedules, allowing users to estimate the long-term financial implications of different employment levels. Used both in the descriptive survey vignette and as part of the RCT treatment.&lt;/p&gt;
&lt;p&gt;Second Pillar (Occupational Pension Scheme, PP): Switzerland&amp;rsquo;s occupational pension scheme, the pillar most heavily affected by part-time work because contributions are directly proportional to earnings above a minimum annual earnings threshold. The paper documents an average gender pension gap of 47.5% in this pillar in 2024 and a 43% lower monthly pension receipt for the average female teacher&amp;rsquo;s part-time trajectory relative to full-time employment.&lt;/p&gt;
&lt;p&gt;Two-Stage Randomization: The experimental design used to separate direct treatment effects from spillover effects within schools. One-third of schools are assigned to a pure control group; in the remaining two-thirds, teachers are individually randomized into treatment or spillover control (untreated teachers in treated schools), enabling identification of both causal treatment impacts and social learning channels.&lt;/p&gt;
&lt;p&gt;Information Constraint: The paper&amp;rsquo;s central mechanism — mothers&amp;rsquo; failure to spontaneously account for the full long-term financial implications of reduced labor supply when making employment decisions, distinct from deliberate forward-looking tradeoffs. The authors document this both through the absence of long-term financial factors in open-ended decision narratives (only 11% of mothers mention them) and through systematic misperceptions of pension and wage outcomes.&lt;/p&gt;
&lt;p&gt;Cognitive Dissonance (as used in the paper): The authors use this term to describe the initial negative emotional response (−0.41 SD overall, −0.68 SD for cost-unaware women) when treated women learn that the true financial costs of part-time work are higher than they expected — information that conflicts with prior beliefs and prior choices, producing unpleasant emotions that subsequently reverse into lower stress levels two months later.&lt;/p&gt;</description></item><item><title>A Cognitive Theory of Reasoning and Choice</title><link>https://macropaperwarehouse.com/papers/a-cognitive-theory-of-reasoning-and-choice/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/a-cognitive-theory-of-reasoning-and-choice/</guid><description>&lt;p&gt;Bordalo, Gennaioli, Lanzani, and Shleifer develop a cognitive theory of choice in which a decision maker&amp;rsquo;s attention to the features of options is determined by her categorization of the current problem against a memory database of problems she solved in the past. The core claim is that before solving a problem, the decision maker asks &amp;ldquo;what kind of problem is this?&amp;rdquo; and resolves it by selecting the category — indexed by a prototype attention-plus-context vector and a time-discounted frequency — whose similarity to the current problem is maximized. This problem recognition step then pins down which features (price, quality, probabilities) receive attention, which in turn shapes valuation and choice.&lt;/p&gt;
&lt;p&gt;The model formalizes two-step choice. In step one (recognition), the decision maker jointly chooses an attention vector alpha_P and a category c* to maximize a separable similarity function S[(alpha_P, kappa_P), (alpha_c, kappa_c)] weighted by category frequency F_c, plus a Type I extreme-value shock that yields a logit probability over categories. In step two, she maximizes perceived value over the menu using the endogenously determined weights. Perceived hedonic value of feature i shrinks toward the menu average when alpha_{P,i} &amp;lt; 1; perceived probabilities compress toward uniform when the event-attention weight falls below 1, producing probability overweighting of unlikely events. Full attention recovers expected utility.&lt;/p&gt;
&lt;p&gt;The model yields three structural predictions that hold without changing tastes or information. First, within-person multi-modal attention: because categorization is stochastic, the same person can cluster on entirely different features (e.g., the base rate vs. the likelihood in an inference problem) across otherwise identical choice occasions. Second, systematic context-driven instability: when an irrelevant context feature kappa_{P,i} drifts away from a category&amp;rsquo;s diagnostic kappa_{c,i}, the probability of that category falls discontinuously, causing a discrete switch in the attention profile and hence in valuation. Third, experience-driven heterogeneity: people more frequently exposed to a category (higher F_c) are more likely to use it, producing persistent differences in price elasticities or probability weighting at constant income and tastes.&lt;/p&gt;
&lt;p&gt;Applied to riskless consumer choice, the paper introduces two categories — &amp;ldquo;buying&amp;rdquo; (full attention to price, partial to quality: alpha_{M_g}=1 &amp;gt; alpha_{Q_g}=alpha) and &amp;ldquo;consuming&amp;rdquo; (full attention to quality, partial to price: alpha_{Q_g}=1 &amp;gt; alpha_{M_g}=alpha). A jam problem categorized as buying yields valuation v = alpha&lt;em&gt;q - eta&lt;/em&gt;p; categorized as consuming, v = q - alpha&lt;em&gt;eta&lt;/em&gt;p. The valuation jumps discontinuously as context crosses a threshold kappa*, which shifts when relative category frequency F_{buy}/F_{con} changes. This framework accounts for context-dependent price elasticities (Wakefield and Inman 2003), poverty-driven excess price focus (Shah et al. 2018), de-commoditization through advertising, and mental accounting anomalies including opportunity cost neglect and the sunk cost fallacy — both arising because con neglects capital gains (alpha_{con,Delta_M}=0) and buy neglects quality shocks (alpha_{buy,Delta_Q}=0).&lt;/p&gt;
&lt;p&gt;Applied to statistical judgment, the paper introduces two categories — &amp;ldquo;frequency estimation&amp;rdquo; (attention alpha_1=1 to a single i.i.d. draw from a known DGP) and &amp;ldquo;agnostic inference&amp;rdquo; (attention alpha_S=1 to the share of heads as a sufficient statistic). The threshold N* separates recognition: for sequence length N_P &amp;lt; N*(F_{freq}/F_{inf}), the decision maker categorizes as frequency and correctly assesses odds; for N_P &amp;gt;= N*, she switches to inference and overweights balanced sequences, producing the Gambler&amp;rsquo;s Fallacy. The same competition between categories also accounts for base rate neglect, conjunction fallacy, and correlation neglect, with the bias strengthening as sequences grow longer.&lt;/p&gt;
&lt;p&gt;Applied to risky choice, bottom-up salience — sensory prominence and contrast — interacts with categorization. A publicity shock drawing attention to a low-probability contamination risk raises similarity to &amp;ldquo;consuming,&amp;rdquo; triggering a category switch that amplifies attention to quality broadly and reduces attention to price, producing large valuation drops disproportionate to the actual probability shift. This mechanism generates the framing effects of prospect theory without a stable S-shaped utility function: gains and losses frames correspond to different contexts activating different categories.&lt;/p&gt;
&lt;p&gt;Scope conditions: the theory applies when features and their values are fully known to the decision maker (no uncertainty about attributes), so the distortions take the form of altered sensitivity to known features rather than missing information. The set of categories C is taken as given in the formal analysis, though the authors discuss endogenization as future work.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s central departure from standard rational inattention and noisy-perception models?&lt;/p&gt;
&lt;p&gt;A: Standard models (Sims 2003, Woodford 2012, Enke and Graeber 2023) produce unimodal, stably weighted valuations — the decision maker&amp;rsquo;s weighting of features is a smooth function of payoff-relevant costs or priors. In this paper, the weighting is determined by problem recognition, which is discrete and stochastic, producing within-person multi-modal attention: the same person can cluster on entirely different features across identical problems. The authors cite direct evidence from Bordalo, Conlon, Gennaioli, Kwon, and Shleifer [20] showing bimodal clustering on base rates vs. likelihoods in statistical problems, a pattern inconsistent with stable-weighting models.&lt;/p&gt;
&lt;p&gt;Q: How is perceived value distorted when the attention weight on a hedonic feature is below 1?&lt;/p&gt;
&lt;p&gt;A: The perceived value of hedonic feature i is u_i(alpha_P) = alpha_{P,i} * u_i + (1 - alpha_{P,i}) * u_bar_i, where u_bar_i is the average value of that feature across options in the menu. An attention weight of zero collapses perceived variation in that feature to zero; full attention recovers the true value. The implication is that under-attention shrinks the decision maker&amp;rsquo;s effective sensitivity to a known attribute, causing systematic under- or over-valuation relative to a rational benchmark while tastes (marginal utilities) are held fixed.&lt;/p&gt;
&lt;p&gt;Q: How is perceived probability distorted?&lt;/p&gt;
&lt;p&gt;A: With attention weight alpha_{P,W} on event W, the perceived probability of event e is P(e)^{alpha_{P,W}} / sum_{e&amp;rsquo;} P(e&amp;rsquo;)^{alpha_{P,W}}, which compresses the distribution toward uniform as alpha_{P,W} falls toward 0 and recovers the true distribution at alpha_{P,W}=1. In the jam example, under-attention to the small probability of spoilage causes the decision maker to overestimate the risk of contamination. For multi-dimensional event vectors the formula generalizes multiplicatively, allowing &amp;ldquo;editing out&amp;rdquo; of entire event dimensions (e.g., urn selection in a balls-and-urns problem) when their attention weight hits zero.&lt;/p&gt;
&lt;p&gt;Q: What is the mechanism for context-dependent price elasticity?&lt;/p&gt;
&lt;p&gt;A: When context kappa_P is below threshold kappa*(F_{buy}/F_{con}), the decision maker categorizes the problem as &amp;ldquo;buying&amp;rdquo; and her valuation is v = alpha&lt;em&gt;q - eta&lt;/em&gt;p, giving a high price sensitivity (coefficient eta) and attenuated quality sensitivity (coefficient alpha &amp;lt; 1). Above kappa*, she categorizes as &amp;ldquo;consuming&amp;rdquo; and valuation is v = q - alpha&lt;em&gt;eta&lt;/em&gt;p, reversing the emphasis. Because the threshold kappa* is increasing in relative frequency F_{buy}/F_{con}, a decision maker with more buying experience has a higher threshold and thus acts as more price-elastic at any given context level. These elasticity differences arise without any change in the true marginal utility of money eta or quality q.&lt;/p&gt;
&lt;p&gt;Q: How does the model generate the sunk cost fallacy and opportunity cost neglect as a unified phenomenon?&lt;/p&gt;
&lt;p&gt;A: Both anomalies arise because buying and consuming categories selectively neglect shocks. In the football example, recognizing the problem as &amp;ldquo;buying&amp;rdquo; activates alpha_{buy,Delta_Q}=0, so the blizzard quality shock Delta_q&amp;lt;0 is ignored and the decision maker drives to the game as if the shock did not occur — the sunk cost fallacy. In the wine example, recognizing the problem as &amp;ldquo;consuming&amp;rdquo; activates alpha_{con,Delta_M}=0, so the capital gain Delta_p is ignored and the decision maker reports a zero or purchase-price cost — opportunity cost neglect. The unifying mechanism is that each category attends only to the features diagnostic of its prototypical experiences: buying attends to price paid and normal quality; consuming attends to realized quality and partly to price, but not to capital gains.&lt;/p&gt;
&lt;p&gt;Q: What comparative static does the model predict for sunk cost susceptibility based on experience?&lt;/p&gt;
&lt;p&gt;A: People with higher F_{buy} (more buying experiences, e.g. poverty experiences or having recently purchased but not yet consumed the good) exhibit more sunk cost fallacy and less opportunity cost neglect. Conversely, season ticket holders face many consuming experiences relative to one buying event, raising F_{con} and thus reducing susceptibility to the sunk cost fallacy for sports events. Making the blizzard more salient in the description shifts similarity toward &amp;ldquo;consuming,&amp;rdquo; also reducing the sunk cost fallacy through a different channel (bottom-up salience rather than experience).&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s explanation for the Gambler&amp;rsquo;s Fallacy, and what distinguishes it from prior accounts?&lt;/p&gt;
&lt;p&gt;A: The Gambler&amp;rsquo;s Fallacy arises when sequence length N_P exceeds threshold N*(F_{freq}/F_{inf}), causing the decision maker to switch from the frequency category (which attends to the 50:50 fairness of the coin) to the inference category (which attends to the share of heads). Under inference, the decision maker treats balanced and unbalanced sequences as representatives of their &amp;ldquo;share of heads equivalence class,&amp;rdquo; and the class of balanced sequences is larger, so balanced sequences receive higher estimated probability — the Gambler&amp;rsquo;s Fallacy. This differs from Rabin and Vayanos (2010), where the bias stems from a belief that the coin is drawn from a pool; here the decision maker knows the coin is fair (kappa_{P,U}=0.5) but the inference representation causes question substitution rather than a wrong model of the DGP.&lt;/p&gt;
&lt;p&gt;Q: How does the model make the Gambler&amp;rsquo;s Fallacy testable beyond length effects?&lt;/p&gt;
&lt;p&gt;A: The model predicts the bias is stronger for decision makers who recently solved many inference problems (lower F_{freq}/F_{inf}), and weaker when the 50:50 nature of flips is made bottom-up salient in the choice context (because salience raises similarity to the frequency category, hindering recognition of inference). These cognitive proxies — experience frequencies and bottom-up salience — are orthogonal to the statistical content of the problem and thus allow identification of the mechanism separately from changes in information or incentives.&lt;/p&gt;
&lt;p&gt;Q: How does the model produce framing effects in risky choice without a stable S-shaped utility function?&lt;/p&gt;
&lt;p&gt;A: Gains and losses frames are modeled as different context vectors kappa_P that differentially increase similarity to a &amp;ldquo;safe outcome&amp;rdquo; category or a &amp;ldquo;risk&amp;rdquo; category. Recognizing the problem as the safe-outcome category shifts attention toward the certain option; recognizing it as the risk category shifts attention toward variance. The reversal of preferences between gain and loss frames (the Asian Disease problem, Tversky and Kahneman 1981) thus emerges from context-driven re-categorization rather than from a fixed probability weighting function. The novel prediction is that framing effects should be stronger for decision makers with more experience with the category activated by each frame, and weaker when bottom-up salience of the alternative frame&amp;rsquo;s features is raised.&lt;/p&gt;
&lt;p&gt;Q: How does bottom-up salience interact with top-down categorization in the contamination example?&lt;/p&gt;
&lt;p&gt;A: A publicity shock alpha_{delta,Q_b}&amp;gt;0 raises baseline attention to the spoiled-jam quality feature, increasing the similarity of the current problem to the &amp;ldquo;consuming&amp;rdquo; category (where quality is focal). This triggers a category switch for marginal agents, activating the full consuming attention profile — which attends to quality broadly, not just to contamination specifically, and reduces attention to price. The resulting valuation drop is therefore disproportionate to the actual probability of contamination and exhibits price insensitivity, because re-categorization shifts the entire attention profile rather than just updating a single probability.&lt;/p&gt;
&lt;p&gt;Q: How does the model relate to and distinguish itself from case-based decision theory (Gilboa and Schmeidler 1995) and analogical reasoning (Mullainathan 2002, Fryer and Jackson 2008)?&lt;/p&gt;
&lt;p&gt;A: In Gilboa-Schmeidler and related models, the decision maker uses past cases to resolve uncertainty about unknown attributes of current options; attention is full and the mechanism is extrapolation of payoffs from similar cases. In Mullainathan (2002) memory-based model, categories again serve to fill in missing information. In this paper, there is no uncertainty about attributes — features and their values are fully known — and the distortion instead takes the form of altered sensitivity to known features through selective attention. This allows the model to produce biases even in simple problems with full data disclosure, and to explain phenomena like base rate neglect and price insensitivity that are not primarily about missing information.&lt;/p&gt;
&lt;p&gt;Q: What does the model predict about within-person versus across-person distributions of valuations?&lt;/p&gt;
&lt;p&gt;A: Within a person, attention is multi-modal (bimodal in the two-category case) because categorization is stochastic. However, if many categories are possible across the population, the aggregate distribution of valuations can appear approximately unimodal even though each individual&amp;rsquo;s distribution is not. This distinction is empirically important: a researcher observing average choices may incorrectly infer smooth preference heterogeneity when the underlying mechanism is discrete category switching.&lt;/p&gt;
&lt;p&gt;Q: What cognitive proxies does the model propose for empirical identification?&lt;/p&gt;
&lt;p&gt;A: The theory links endogenous attention and choice to three observable (or measurable) proxies: (1) past experience frequencies F_c, measurable from administrative histories, surveys about past exposure, or experimental manipulation of training; (2) contextual similarity, measurable from field or experimental variation in irrelevant context features; and (3) bottom-up salience, experimentally controllable via prominence or contrast manipulations. The key identification logic is that these proxies are payoff-irrelevant — they do not change tastes, information, or the objective choice problem — yet predict systematic shifts in choice through their effect on recognition.&lt;/p&gt;
&lt;p&gt;Problem Recognition: The first step in the decision maker&amp;rsquo;s choice process, in which she jointly selects an attention vector alpha_P and a category c* by maximizing weighted similarity between the current problem (characterized by its context vector kappa_P) and the prototype of a past category (alpha_c, kappa_c), multiplied by the category&amp;rsquo;s time-discounted frequency F_c. Recognition is not about resolving uncertainty over attributes but about selecting which known attributes to attend to.&lt;/p&gt;
&lt;p&gt;Category: A partition element of the decision maker&amp;rsquo;s memory database, indexed by a prototype attention-plus-context vector (alpha_c, kappa_c) and a frequency scalar F_c. The prototype encodes both the context features diagnostic of experiences in that category (binary alpha_{c,i} for i in Phi_K) and the attention to hedonic and event features (alpha_{c,i} for i in Phi_H union Phi_E) used when solving problems in that category. Examples in the paper: &amp;ldquo;buying&amp;rdquo; and &amp;ldquo;consuming&amp;rdquo; for riskless choice; &amp;ldquo;frequency estimation&amp;rdquo; and &amp;ldquo;agnostic inference&amp;rdquo; for statistical judgment.&lt;/p&gt;
&lt;p&gt;Attention Weight (alpha_{P,i}): A scalar in [0,1] assigned to feature i of the current problem P. For hedonic features, alpha_{P,i}&amp;lt;1 collapses perceived variation toward the menu average; for event features, alpha_{P,i}&amp;lt;1 compresses perceived probabilities toward uniform. Full attention alpha_{P,i}=1 recovers expected utility. Attention weights are the endogenous output of the recognition step, not fixed preference parameters.&lt;/p&gt;
&lt;p&gt;Contextual Similarity S: A separable function measuring how close the current problem (alpha_P, kappa_P) is to a category prototype (alpha_c, kappa_c). It decreases in discrepancies in the attention vector (measured by a strictly increasing, convex function d) and in discrepancies in the values of context features diagnostic of the category (d_i(kappa_{P,i}, kappa_{c,i}) * alpha_{c,i}). Endogenous attention to context is set to reduce sensitivity to discrepancies, not to eliminate them.&lt;/p&gt;
&lt;p&gt;Mental Accounting (as categorization): In the paper&amp;rsquo;s account, non-fungibility, sunk cost fallacy, and opportunity cost neglect all arise because buying and consuming categories selectively attend to different monetary and quality features. The sunk cost effect is alpha_{buy,Delta_Q}=0; opportunity cost neglect is alpha_{con,Delta_M}=0. Mental accounts are not separate budget constraints but the by-product of category-specific attention profiles that were calibrated to normal-state experiences and do not generalize to shocks.&lt;/p&gt;
&lt;p&gt;Bottom-up Salience: Exogenous attention to a feature driven by sensory prominence (described by alpha_{delta,i} in the problem&amp;rsquo;s presentation vector) or payoff contrast (the DM attends more to features where her option&amp;rsquo;s value deviates more from the menu average relative to total menu variance). Bottom-up salience raises baseline attention to a feature before top-down categorization acts, and can trigger a category switch by raising similarity to the category for which that feature is focal.&lt;/p&gt;
&lt;p&gt;Gambler&amp;rsquo;s Fallacy via Question Substitution: In the model, the Gambler&amp;rsquo;s Fallacy arises when a long sequence length kappa_{P,N} causes recognition of the &amp;ldquo;agnostic inference&amp;rdquo; category, which focuses attention on the share of heads alpha_S=1. The decision maker then treats sequences as representatives of a &amp;ldquo;share of heads equivalence class,&amp;rdquo; and since the balanced class is larger than the unbalanced class, balanced sequences are assigned higher estimated probability. This is not a belief that the coin is unfair; it is question substitution induced by the inference representation.&lt;/p&gt;</description></item><item><title>All Along the Watchtower: Military Landholders and Serfdom Consolidation in Early Modern Russia</title><link>https://macropaperwarehouse.com/papers/all-along-the-watchtower-military-landholders-and-serfdom-consolidation-in-early-modern-russia/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/all-along-the-watchtower-military-landholders-and-serfdom-consolidation-in-early-modern-russia/</guid><description>&lt;p&gt;This paper investigates the origins of serfdom in early modern Russia, arguing that the institution consolidated primarily through political economy dynamics between the crown and a landholding military class, rather than from economic fundamentals such as labor scarcity, land-labor ratios, or grain trade opportunities. The central argument is that the prolonged defense of Russia&amp;rsquo;s southern frontier against Crimean Tatar nomadic raids generated a class of military landholders who possessed both the coercive capacity and the political leverage to press the state into restricting peasant labor mobility.&lt;/p&gt;
&lt;p&gt;The mechanism runs as follows. The Russian state, lacking the fiscal capacity to pay soldiers directly, granted frontier lands along the Tula defense line to high-ranked soldiers in exchange for military service under the pomest&amp;rsquo;e system. These lands were selected for their defensive rather than agricultural value and sat on the forest-steppe boundary roughly 180 km south of Moscow. Since soldiers could not farm while on duty and could not compete in free labor markets given the area&amp;rsquo;s low agricultural attractiveness, the arrangement was only sustainable if peasants were bound to the land. Military landholders collectively petitioned the Tsar repeatedly — with petition volumes peaking during urban uprisings (9 petitions in 1648, 13 in 1682) when the government&amp;rsquo;s political vulnerability increased the military&amp;rsquo;s bargaining power — until serfdom was codified in the Law Code of 1649.&lt;/p&gt;
&lt;p&gt;The authors test this theory using newly digitized data from the 1678 household census, which records male population by six legally distinct peasant categories across 172 districts of Muscovy, combined with data on landholder estate counts and sizes. The primary empirical finding is that districts on the Tula defense line had approximately 40% of their population composed of serfs, compared to roughly 14% nationally — a difference of about 25 percentage points that survives the inclusion of geographic and climatic controls (grain suitability, temperature seasonality, precipitation, terrain ruggedness, river location, distance to Moscow, and regional fixed effects). Placebo tests confirm this pattern is specific to the most legally dependent peasant groups: the defense line is negatively associated with royal peasants and statistically insignificant for church peasants, free peasants, and non-Russian peasants.&lt;/p&gt;
&lt;p&gt;To address potential endogeneity of the defense line&amp;rsquo;s location, the authors construct an instrumental variable using a novel geospatial algorithm. The algorithm computes optimal nomadic invasion routes from Crimea to Moscow via topographic cost rasters (using flow accumulation values as proxies for river-crossing barriers), then intersects these routes with the historically stable forest-steppe boundary (identified through FAO/UNESCO soil types — Podzoluvisols versus Chernozems). Districts at this intersection were 70 percentage points more likely to host the actual defense line. Two-stage least squares estimates confirm and slightly exceed the OLS magnitudes, supporting the causal interpretation.&lt;/p&gt;
&lt;p&gt;The paper further tests two canonical alternative explanations and finds them insufficient. Domar&amp;rsquo;s (1970) labor-scarcity hypothesis predicts serfdom should be higher where population density is lower; the data show the opposite sign, contradicting this prediction. The Baltic grain trade hypothesis yields only a small, unstable positive interaction between river access to the Baltic and grain suitability, which disappears when the defense line variable is included. A horse race including all variables simultaneously shows the defense line coefficient at approximately 24 percentage points remains stable while alternative predictors become insignificant.&lt;/p&gt;
&lt;p&gt;Mechanism tests show that defense line districts had 3.2 more estates per 100 square kilometers than the national average of 2.3, with the excess concentrated in very small (up to 5 serf households) and small (6–25 households) estates — consistent with the state&amp;rsquo;s strategy of maximizing soldier count by allocating the minimum serf labor sufficient to sustain a cavalryman. A bigram similarity analysis of collective petitions versus the 1649 Law Code yields a correlation coefficient of 0.7 for the top twenty bigrams between a 1637 petition and Chapter 11 (restricting peasant mobility), with no comparable similarity to other chapters. Persistence is documented through 1719, 1795, and 1858 censuses: defense line districts maintained the highest serf concentration through to three years before emancipation in 1861.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-central-argument-about-the-origins-of-russian-serfdom"&gt;Q1. What is the paper&amp;rsquo;s central argument about the origins of Russian serfdom?&lt;/h3&gt;
&lt;p&gt;A: The paper argues that serfdom consolidated primarily due to political economy dynamics: the crown&amp;rsquo;s dependence on a landholding military class for frontier defense against steppe nomads gave that class sufficient political leverage to secure the legal restriction of peasant labor mobility. The military landholders&amp;rsquo; coercive capacity and proximity to their small estates made labor coercion a viable complement to their military function. This explanation dominates alternative accounts based on labor scarcity, grain trade, or soil quality in all specifications tested.&lt;/p&gt;
&lt;h3 id="q2-what-was-the-tula-defense-line-and-why-was-it-located-where-it-was"&gt;Q2. What was the Tula defense line and why was it located where it was?&lt;/h3&gt;
&lt;p&gt;A: The Tula defense line (Great Abatis Line) was a chain of about 40 fort towns stretching over 500 km east-west, centered on Tula approximately 180 km south of Moscow, erected in the 1560s using felled trees, earth mounds, ditches, and watchtowers. Its location on the forest-steppe boundary was determined by two military-logistical constraints: it had to block the main nomadic invasion routes from Crimea, and it had to lie within the forest zone where timber was the cheapest construction material and which provided natural shelter. The paper documents that the defense line area did not differ from the rest of Muscovy in agricultural suitability, annual precipitation, seasonality, or terrain ruggedness — its distinctive feature was purely defensive.&lt;/p&gt;
&lt;h3 id="q3-how-large-is-the-estimated-effect-of-defense-line-proximity-on-serf-concentration"&gt;Q3. How large is the estimated effect of defense line proximity on serf concentration?&lt;/h3&gt;
&lt;p&gt;A: In the unconditional specification, defense line districts had a 30 percentage point higher share of serfs than the rest of the country. After adding geographic controls (grain suitability, seasonality, precipitation, terrain ruggedness, river dummy, distance to Moscow, and regional fixed effects), the coefficient stabilizes at approximately 25 percentage points. Given that serfs averaged about 14% of total population nationally but about 40% in defense line districts, the estimated effect is substantial relative to the baseline.&lt;/p&gt;
&lt;h3 id="q4-how-do-the-authors-address-endogeneity-of-the-defense-line-location"&gt;Q4. How do the authors address endogeneity of the defense line location?&lt;/h3&gt;
&lt;p&gt;A: They construct an instrumental variable defined as the intersection of two variables: districts lying on the computed optimal nomadic invasion routes (covering 98 of 172 districts, or 57% of the sample), and districts on the forest-steppe soil boundary (38 districts, or 22% of the sample). Their interaction covers 23 districts and is the excluded instrument. In the first stage, this interaction term raises a district&amp;rsquo;s probability of hosting the actual defense line by 70 percentage points, while the linear terms become essentially zero once the interaction is included. The 2SLS second-stage estimates of the serf-share effect are slightly higher than OLS and statistically significant, confirming the direction and approximate magnitude of the OLS results.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-paper-find-about-domars-labor-scarcity-hypothesis"&gt;Q5. What does the paper find about Domar&amp;rsquo;s labor-scarcity hypothesis?&lt;/h3&gt;
&lt;p&gt;A: The paper finds no support for Domar&amp;rsquo;s (1970) prediction that serfdom should be more prevalent where labor is scarcer (lower population density). Controlling for grain suitability and geographic factors, population density enters with a positive and statistically significant coefficient at the 5% level — the opposite sign from what Domar&amp;rsquo;s theory predicts. When the defense line dummy is added, population density becomes insignificant while the defense line coefficient remains at approximately 25 percentage points, consistent with the baseline.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-paper-find-about-the-baltic-grain-trade-hypothesis"&gt;Q6. What does the paper find about the Baltic grain trade hypothesis?&lt;/h3&gt;
&lt;p&gt;A: An exogenous measure of Baltic trade potential — a dummy for districts with river access to the Baltic, interacted with grain suitability — yields a small and marginally positive effect on serf share in Baltic districts with higher grain suitability. However, this effect disappears when the defense line dummy is included, and is also sensitive to alternative spatial clustering (becoming insignificant at the 300 km clustering radius even without the defense line dummy). The authors interpret this instability as inconsistent with grain trade being a primary driver of serfdom.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-evidence-for-the-estate-size-mechanism"&gt;Q7. What is the evidence for the estate-size mechanism?&lt;/h3&gt;
&lt;p&gt;A: Defense line districts had on average 3.2 more estates per 100 square kilometers than the national average of 2.3 per 100 square kilometers. Among estate-size brackets, very small (up to 5 serf households) and small (6–25 serf households) estates were disproportionately concentrated in defense line districts, while the location of medium-sized and large estates was statistically independent of the defense line. This pattern is consistent with the state&amp;rsquo;s strategy of allocating minimum viable serf endowments to maximize the number of soldiers supportable along the line.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-textual-evidence-linking-military-petitions-to-the-1649-law-code"&gt;Q8. What is the textual evidence linking military petitions to the 1649 Law Code?&lt;/h3&gt;
&lt;p&gt;A: A bigram similarity analysis between a 1637 collective petition and Chapter 11 of the 1649 Law Code reveals a correlation coefficient of 0.7 for the top twenty bigrams. The five most common bigrams appear in both texts: &amp;ldquo;runaway peasants,&amp;rdquo; &amp;ldquo;commoner peasants,&amp;rdquo; &amp;ldquo;census books,&amp;rdquo; &amp;ldquo;search years,&amp;rdquo; and &amp;ldquo;tsar&amp;rsquo;s decree.&amp;rdquo; This correlation does not extend to other chapters of the Law Code that regulate non-peasant matters, establishing specificity of the legislative influence.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-timing-of-collective-petitions-relate-to-political-crises"&gt;Q9. How does the timing of collective petitions relate to political crises?&lt;/h3&gt;
&lt;p&gt;A: Over a corpus of 96 petitions between 1608 and 1698, landholders petitioned on average once per year, but activity spiked sharply during domestic uprisings: 9 petitions in 1648 (the &amp;ldquo;Salt Riot&amp;rdquo; urban uprising) and 13 petitions in 1682 (the musketeers&amp;rsquo; revolt). These peaks coincide with moments when the government&amp;rsquo;s political vulnerability increased the military&amp;rsquo;s bargaining power, and in both cases were followed by legislative concessions — the 1649 Law Code and new decrees in 1683–85 on harsher punishment for harboring runaways, respectively.&lt;/p&gt;
&lt;h3 id="q10-what-do-the-placebo-tests-show"&gt;Q10. What do the placebo tests show?&lt;/h3&gt;
&lt;p&gt;A: Regressions of non-serf peasant shares on the defense line dummy show that the defense line is negatively associated with royal peasants and statistically insignificant for church peasants, free peasants, and non-Russian peasants. A placebo test replacing military landholders with merchants and artisans shows no significant defense line effect on the latter group, while Moscow has an 11 percentage point higher merchant/artisan share. The specificity of the defense line effect to legally dependent peasants and military landholders supports the military-political mechanism rather than a generic frontier-area effect.&lt;/p&gt;
&lt;h3 id="q11-how-persistent-was-the-spatial-distribution-of-serfdom-after-1649"&gt;Q11. How persistent was the spatial distribution of serfdom after 1649?&lt;/h3&gt;
&lt;p&gt;A: The authors estimate their baseline equation with serf share from the 1719, 1795, and 1858 censuses as dependent variables. Defense line districts maintained disproportionately higher serf densities in all three periods, including when the sample is restricted to the original Muscovite districts to exclude post-18th century territorial acquisitions. By 1858, three years before emancipation, the spatial distribution of serfs remained similar to that observed 200 years earlier at the time of serfdom&amp;rsquo;s consolidation — despite the defense line having been militarily obsolete for over a century.&lt;/p&gt;
&lt;h3 id="q12-what-explains-the-persistence-of-serfdom-beyond-its-original-military-rationale"&gt;Q12. What explains the persistence of serfdom beyond its original military rationale?&lt;/h3&gt;
&lt;p&gt;A: The persistence reflects a mutually beneficial exchange between the crown and former military landholders. Landholders provided local state capacity — overseeing tax collection, administering military conscription, and adjudicating peasant disputes through estate courts — in lieu of a centralized bureaucracy. In return, the crown granted successive expansions of landholder rights: Peter I equalized military landholdings with hereditary estates in 1714, and Peter III in 1762 freed landholders from military service obligations while retaining their property rights over land and serfs. This fiscal-administrative dependency is also cited as a reason for the late timing and unfavorable-to-peasants terms of the 1861 emancipation reform.&lt;/p&gt;
&lt;h3 id="q13-how-does-this-papers-explanation-relate-to-easternwestern-european-institutional-divergence"&gt;Q13. How does this paper&amp;rsquo;s explanation relate to Eastern/Western European institutional divergence?&lt;/h3&gt;
&lt;p&gt;A: The paper argues that while the military revolution in Western Europe generated fiscally capable centralized states with regular infantry armies, Russia&amp;rsquo;s peripheral nomadic threat prolonged the feudal cavalry model supported by land grants and serf labor. This delayed the formation of Weberian bureaucracy and entrenched what the authors term a &amp;ldquo;garrison state&amp;rdquo; — one whose institutions and social structure were shaped primarily by military-security considerations. The paper positions military factors alongside existing divergence explanations emphasizing land property rights, political institutions, demographic regimes, and Enlightenment ideas.&lt;/p&gt;
&lt;h3 id="q14-what-is-the-methodological-contribution-of-the-optimal-invasion-route-algorithm"&gt;Q14. What is the methodological contribution of the optimal invasion route algorithm?&lt;/h3&gt;
&lt;p&gt;A: The algorithm uses flow accumulation rasters (proportional to river width and basin size) as a cost function to compute the lowest-cost travel paths from Crimea to Moscow, iteratively penalizing cells within 15 km of each computed route and re-running the path search to generate four distinct routes per origin point (eight total, including routes from the Don River steppe). This produces a high-resolution, geographically continuous measure of military threat exposure that the authors argue provides statistical power in contexts where terrain ruggedness or simple distance measures lack variation — particularly relevant for flat plains with a single threat origin correlated with other variables.&lt;/p&gt;
&lt;p&gt;Pomest&amp;rsquo;e system: The institutional arrangement by which the Russian state granted frontier lands to high-ranked soldiers in exchange for military service, under the rule that &amp;ldquo;the land must not leave the service.&amp;rdquo; Unlike hereditary estates, pomest&amp;rsquo;e holdings were conditional on active service and could not be passed to heirs unless sons continued military service. This system enabled the formation of a permanent cavalry force despite the state&amp;rsquo;s low fiscal capacity, but required binding peasants to the land to make the arrangement viable for the soldier-landholders.&lt;/p&gt;
&lt;p&gt;Serfs (bobyli and dvorovye): In the paper&amp;rsquo;s 1678 census framework, serfs are defined as the two most legally dependent subgroups of private peasants — cotters (bobyli), who owned no property and worked full-time for their landlord in exchange for payment in kind, and servants (dvorovye), who performed household and support functions on the estate. These groups constituting about 14% of total population nationally were totally dependent on their landlord and could not retain the marginal product of any part of their labor. After the 1649 Law Code, villeins (krest&amp;rsquo;yane) gradually converged to this status as well.&lt;/p&gt;
&lt;p&gt;Collective petitions (chelobitnye): The primary institutional channel through which the military landholder class communicated collective interests and applied political pressure on the crown in 17th-century Muscovy. The paper documents 96 such petitions between 1608 and 1698, showing that their volume, timing (peaking during urban uprisings), and textual content (closely matching Chapter 11 of the 1649 Law Code) were the proximate mechanism by which landholders converted military leverage into legal codification of serfdom.&lt;/p&gt;
&lt;p&gt;Optimal defense line (instrumental variable): The paper&amp;rsquo;s constructed instrument, defined as the intersection of computed optimal nomadic invasion routes (based on topographic cost rasters approximating river-crossing barriers) and the forest-steppe soil boundary (Podzoluvisols/Chernozems boundary from the FAO/UNESCO Soil Map). This instrument captures the geographically and militarily determined placement of defensive fortifications, purging variation in actual defense line location that might reflect agricultural or economic value.&lt;/p&gt;
&lt;p&gt;Garrison state: Used by the authors (adapting Lasswell&amp;rsquo;s term) to describe a state whose institutions and social structure are shaped primarily by military security considerations. In the Russian context, this refers to the persistence of a feudal cavalry system, land-grant-based military compensation, and labor coercion that together delayed centralized state formation and Weberian bureaucracy relative to Western European states undergoing the military revolution toward regular infantry armies.&lt;/p&gt;
&lt;p&gt;Labor coercion complementarity: The paper&amp;rsquo;s mechanism whereby employers with high coercive capacity (proximity to weapons, military training) can deploy that same capacity to restrict workers&amp;rsquo; outside options and extract labor surplus. In the defense line context, soldiers&amp;rsquo; military skills and armament made them effective at preventing serf flight and enforcing labor obligations — creating a complementarity between military capacity and serfdom that was absent among merchants or church institutions with comparable landholdings elsewhere.&lt;/p&gt;</description></item><item><title>An endogenous gridpoint method for distributional dynamics</title><link>https://macropaperwarehouse.com/papers/an-endogenous-gridpoint-method-for-distributional-dynamics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/an-endogenous-gridpoint-method-for-distributional-dynamics/</guid><description>&lt;p&gt;This paper introduces the Distributional Endogenous Gridpoint Method (DEGM), a novel numerical technique for solving the distributional dynamics that arise in heterogeneous agent macroeconomic models. The core problem is how to efficiently update the distribution of agents over the state space as the economy evolves. The dominant existing approach — the &amp;ldquo;lottery method&amp;rdquo; of Young (2010) — discretizes the state space and represents policy functions as lotteries over nearby gridpoints, producing a transition matrix that is linear in optimal policies. This linearity renders the lottery method incapable of capturing nonlinear effects in distributional dynamics, a limitation that becomes quantitatively significant for higher-order perturbation solutions.&lt;/p&gt;
&lt;p&gt;DEGM extends Carroll&amp;rsquo;s (2006) endogenous gridpoint method from individual optimization to the distributional level. Rather than discretizing the density and integrating forward, DEGM works directly on the cumulative distribution function (CDF). The key insight is that when the policy function is monotone — as savings functions typically are — the endogenous gridpoints generated by the policy function trace out exact points on the post-policy CDF without requiring integration. Specifically, if A*_{i,j} = a*(A_i, Y_j) are optimal asset choices from grid point A_i at income Y_j, then the CDF values at those endogenous points are known analytically as F_t(A_i | Y_j). An interpolant using shape-preserving splines constructed through these points allows evaluation of the updated CDF at any point without integration. The income transition step is handled separately via standard quadrature over the discretized income process.&lt;/p&gt;
&lt;p&gt;The paper demonstrates DEGM&amp;rsquo;s performance with two applications. First, in the Aiyagari (1994) economy, DEGM converges to the stationary equilibrium an order of magnitude faster than the lottery method in terms of gridpoints. At nk=40 gridpoints, the lottery method deviates from the benchmark capital stock by 1.72% and the wealth Gini by 2.24% (for nh=5), while DEGM deviates by only 0.09% and 0.12% respectively. Both methods converge to the same solution as the number of gridpoints increases, but DEGM reaches this limit far faster.&lt;/p&gt;
&lt;p&gt;Second, the authors introduce a Krusell-Smith style model with aggregate investment risk (capital depreciation shocks calibrated following Barro, 2006, as a 0.4% quarterly probability of 7.5% capital destruction causing a 10% annual GDP drop) as a new baseline for studying aggregate nonlinearities with household heterogeneity. This model overcomes the near-linearity of aggregate capital dynamics in the original Krusell-Smith specification. Using a third-order perturbation solution with DEGM, aggregate investment risk lowers the capital stock by 5 to 11 basis points and increases wealth inequality by up to 11 basis points relative to the non-stochastic steady state, depending on idiosyncratic income risk calibration. The lottery method systematically mispredicts these effects: it always predicts a decrease in wealth inequality in the presence of investment risk, while DEGM predicts an increase. At third order, the lottery method predicts wealth Gini changes of +2.0 bp (persistent calibration) and -149.7 bp (transitory calibration), while DEGM predicts +10.7 bp and +2.1 bp respectively.&lt;/p&gt;
&lt;p&gt;The mechanism for increased inequality under investment risk is heterogeneous: for less wealthy households the substitution effect dominates (they reduce saving more in response to risky returns), while for wealthy households the income effect is stronger and precautionary saving motives dominate. The lottery method, by making the distributional transition matrix linear in policies, zeros out the second derivative of the transition matrix with respect to the policy function, missing the term capturing how the density at the pre-image of each asset level is affected nonlinearly. DEGM&amp;rsquo;s cubic spline interpolant captures all nonlinearities up to third order, enabling economically meaningful results that qualitatively differ from lottery-method predictions on wealth inequality.&lt;/p&gt;
&lt;p&gt;Q: What is the fundamental numerical problem that DEGM solves?
A: Evolving the distribution of agents forward over time in heterogeneous agent models requires evaluating a Kolmogorov forward equation, which naively demands numerical integration. The lottery method avoids integration by discretizing the state space and expressing transitions as a linear matrix operation, but this forces the distributional dynamics to be linear in optimal policies. DEGM avoids integration by exploiting policy function monotonicity: the endogenous policy gridpoints are the interpolation nodes, so the CDF update requires only interpolation, not integration. This preserves nonlinear effects up to the order of the splines used.&lt;/p&gt;
&lt;p&gt;Q: How does DEGM handle the borrowing constraint and the resulting mass point?
A: Savings policy functions are typically weakly monotone: constant at the borrowing constraint for sufficiently poor households, then strictly monotone above a threshold. DEGM accommodates this by starting the endogenous grid at the EGM solution corresponding to the borrowing constraint (the threshold a_j above which the policy is strictly monotone), restoring strict monotonicity on the relevant domain. The mass point at the borrowing constraint is captured by evaluating F_t(a_j, Y_j). Echoes of the borrowing constraint diminish as the number of income states increases, and in practice 10 income gridpoints are sufficient to smooth them.&lt;/p&gt;
&lt;p&gt;Q: How much faster does DEGM converge relative to the lottery method for the stationary equilibrium?
A: In the Aiyagari economy with nk=40 asset gridpoints, the lottery method&amp;rsquo;s capital stock deviates from the benchmark by 1.72% and the wealth Gini by 2.24% (nh=5), while DEGM deviates by only 0.09% and 0.12% respectively — roughly a 20-fold improvement in accuracy for the same gridpoints. At nk=80, the lottery method still shows 0.56%/0.78% deviations while DEGM shows 0.03%/0.00%. Although for a fixed number of gridpoints the lottery method is faster in wall-clock time (0.35s vs 0.82s at nk=40, nh=20), DEGM is faster for a given level of accuracy because it requires far fewer gridpoints.&lt;/p&gt;
&lt;p&gt;Q: Why does the lottery method fail at higher-order perturbations?
A: The lottery method constructs its transition matrix as a piecewise linear function of the optimal policy a*, so its second derivative with respect to a* is zero. As a result, it misses the second term in the second-order derivative of the end-of-period CDF: the term involving the derivative of the density at the pre-image of each asset level times the squared linear policy effect. This missing nonlinearity becomes quantitatively important at second and third order. DEGM&amp;rsquo;s cubic hermitian spline interpolant captures all nonlinearities up to third order, allowing it to correctly represent how the distribution responds nonlinearly to aggregate shocks.&lt;/p&gt;
&lt;p&gt;Q: What does the paper find about the effect of aggregate investment risk on the capital stock and wealth inequality?
A: Using a third-order perturbation solution with DEGM, aggregate investment risk lowers the capital stock by 5 to 11 basis points from the non-stochastic steady state, depending on whether income risk is persistent or transitory (DEGM third-order: -4.7 bp persistent, -11.4 bp transitory). Wealth inequality increases by up to 11 basis points (DEGM third-order: +10.7 bp persistent, +2.1 bp transitory). The lottery method diverges dramatically at third order, predicting Gini changes of +2.0 bp and -149.7 bp for the persistent and transitory calibrations respectively, compared to DEGM&amp;rsquo;s +10.7 bp and +2.1 bp.&lt;/p&gt;
&lt;p&gt;Q: What is the mechanism through which aggregate investment risk increases wealth inequality?
A: The mechanism operates through heterogeneous saving responses across the wealth distribution. For less wealthy households, capital income is a small share of total income, so the substitution effect of risky returns dominates: higher investment risk reduces their incentive to save. For wealthy households, capital income is central, so the income effect is stronger and precautionary saving motives intensify. A capital depreciation shock upon realization compresses the wealth distribution, but the risk of such a shock increases inequality on average because it disproportionately reduces saving among poorer households.&lt;/p&gt;
&lt;p&gt;Q: How do the authors extend DEGM to handle aggregate risk and higher-order perturbations?
A: The authors follow Reiter (2009) in including the distribution and value functions in the state space, defining a nonlinear difference equation over these objects. Higher-order perturbation of this system proceeds using the algorithms of Andreasen et al. (2018) and Levintal (2017), with second-order terms solved via a generalized Sylvester equation using Kim et al.&amp;rsquo;s (2008) doubling algorithm. The implementation handles up to 3,200 variables at second order and 220 variables at third order. For the second-order solution, the Bayer-Luetticke (2020) state-space reduction and its refinement in Bayer et al. (2024) yield results identical to the full unreduced system.&lt;/p&gt;
&lt;p&gt;Q: What is the state-space reduction procedure and how much does it compress the system?
A: The full system uses 402 states and 412 controls (persistent calibration). A copula representation of the distribution reduces this to 213 states and 412 controls; adding DCT compression of the value function gives 213 states and 98 controls; further adding a factor representation from the first-order solution yields 111 states and 98 controls — a 75% reduction. The R-squared-like IRF statistic remains 1.00 across all reductions, and ergodic moments are identical (capital: 25.54, Gini: 0.61 for the persistent calibration).&lt;/p&gt;
&lt;p&gt;Q: Does DEGM produce different first-order impulse responses than the lottery method?
A: For first-order perturbations, DEGM and the lottery method converge to the same solution as the number of gridpoints increases, but DEGM converges faster. For the first-order dynamics of the wealth distribution (wealth Gini IRFs), DEGM reaches convergence with nk=40 gridpoints while the lottery method requires nk=160. For aggregate capital stock IRFs, both methods converge quickly at first order. Quantitative differences become significant only at second and higher orders.&lt;/p&gt;
&lt;p&gt;Q: What calibration is used for the investment risk model?
A: Capital depreciation deviates from its steady-state value by a shock with second moment sigma_delta = 0.005 and third moment tau_delta = 0.012. This corresponds to a 0.4% quarterly probability that a disaster destroys 7.5% of the capital stock and causes a 10% drop in annual GDP, consistent with the evidence in Barro (2006). The model is solved under both a persistent income calibration (beta=0.98, rho=0.98, sigma_epsilon=0.14, implied Gini=0.66) and a transitory income calibration (beta=0.99, rho=0.88, sigma_epsilon=0.18, implied Gini=0.42).&lt;/p&gt;
&lt;p&gt;Distributional Endogenous Gridpoint Method (DEGM): A numerical method for evolving the joint CDF of agents over the state space by constructing an interpolant at endogenous gridpoints A*_{i,j} = a*(A_i, Y_j) — the optimal policy values — at which CDF values are known analytically as F_t(A_i | Y_j), thus updating the distribution through interpolation rather than integration and preserving nonlinearities up to the order of the spline.&lt;/p&gt;
&lt;p&gt;Lottery Method (LM): Young&amp;rsquo;s (2010) standard technique that replaces the continuous distribution with a discrete counterpart and represents optimal policy functions as probability weights over nearby gridpoints, yielding a single transition matrix A* such that f_{t+1} = f_t * A*. The transition matrix is linear in optimal policies, which zeroes out the second derivative of the distributional dynamics with respect to policies and causes systematic misprediction of distributional dynamics under higher-order perturbation.&lt;/p&gt;
&lt;p&gt;Kolmogorov Forward Equation (Distributional Dynamics): The law of motion for the joint CDF F_t(a, y) describing how the distribution of households over assets and income evolves given optimal policies and the income transition process. In DEGM, this equation is split into a sub-period for asset choices (where endogenous gridpoints allow integration-free updating) and a sub-period for income transitions (handled by quadrature over the discretized income process).&lt;/p&gt;
&lt;p&gt;Higher-Order Perturbation Solution: A Taylor expansion of the model&amp;rsquo;s nonlinear equilibrium conditions around the non-stochastic steady state beyond first order. Second-order solutions capture precautionary motives and mean deviations from the steady state; third-order solutions additionally capture asymmetric effects of shocks, requiring DEGM&amp;rsquo;s nonlinear distributional representation to produce accurate results.&lt;/p&gt;
&lt;p&gt;Aggregate Investment Risk (Capital Depreciation Shocks): Shocks to the aggregate capital depreciation rate calibrated following Barro (2006) as a 0.4% quarterly probability of a disaster that destroys 7.5% of the capital stock and causes a 10% annual GDP drop. Proposed as a replacement for near-linear Krusell-Smith aggregate productivity shocks to generate genuine nonlinearities in aggregate capital dynamics while remaining equally parsimonious.&lt;/p&gt;
&lt;p&gt;State-Space Reduction: A sequence of compression techniques — copula representation of the wealth distribution, discrete cosine transform (DCT) compression of the value function, and factor representation from the first-order solution — that reduce the Reiter (2009) system from 402 states and 412 controls to 111 states and 98 controls (a 75% reduction) with no measurable loss of accuracy in impulse responses or ergodic moments.&lt;/p&gt;
&lt;p&gt;Shape-Preserving Interpolation: Interpolation methods (linear spline or piecewise cubic hermitian splines) that maintain the monotonicity of the CDF when constructing the interpolant from endogenous gridpoints. Cubic hermitian splines additionally preserve differentiability, making the distributional dynamics smooth enough for third-order perturbation and capturing all nonlinear effects that the lottery method misses.&lt;/p&gt;</description></item><item><title>An Equilibrium Analysis of the Effects of Neighborhood-Based Interventions on Children</title><link>https://macropaperwarehouse.com/papers/an-equilibrium-analysis-of-the-effects-of-neighborhood-based-interventions-on-children/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/an-equilibrium-analysis-of-the-effects-of-neighborhood-based-interventions-on-children/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; How should governments design neighborhood-based policies to improve long-run outcomes for children, once one accounts for general equilibrium (GE) forces—endogenous rents, neighborhood quality, wages, and distortionary taxation—that small-scale experimental studies cannot identify?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The paper embeds neighborhood effects into a quantitative, heterogeneous-agent overlapping-generations (OLG) model with endogenous location choice and child skill development. The economy has three building blocks: (1) a dynastic life-cycle structure in which parents choose a neighborhood (from two options: a disadvantaged n=1 and an advantaged n=2) and allocate time to child development, with child skills produced by a nested CES aggregator combining parental time and neighborhood quality (proxied by per-capita income in the tract); (2) a GE Aiyagari incomplete-markets framework with endogenous labor supply, wage uncertainty, and progressive labor taxation; and (3) a government that finances housing vouchers or place-based wage subsidies by adjusting the labor income tax parameter, with all additional net expenses fully offset by tax revenue. Housing supply is upward-sloping (elasticity 1.75, from Saiz 2010), so rents are endogenous.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and calibration.&lt;/strong&gt; The model is estimated by simulated method of moments to match U.S. data from the 2000s, drawing on the PSID, NLSY, ATUS, the 2012–2016 ACS, and the Opportunity Atlas (Chetty et al. 2018). Neighborhoods are mapped to Census tracts divided into bottom-10-percent and top-90-percent median household income groups within each commuting zone. Key targeted moments include the income gap between neighborhoods (108 percent higher mean individual income in n=2), the 30 percent higher incomes for children from low-income families raised in the better neighborhood, and a 32 percent gap in weekly parental time with children across neighborhoods.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Validation.&lt;/strong&gt; Before policy counterfactuals, the calibrated model is validated against two bodies of reduced-form evidence. First, a simulated small-scale, single-generation, partial-equilibrium voucher experiment generates 23 percent higher income for children—close to the 31 percent MTO experimental estimate from Chetty et al. (2016), with the difference largely explained by a smaller poverty-rate contrast (18 vs. 22 percentage points) in the simulation. Second, a simulated 20 percent place-based wage subsidy generates 17–21 percent earnings gains for adult residents of n=1, consistent with Busso et al.&amp;rsquo;s (2013) quasi-experimental EZ estimates of 17–24 percent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings — housing vouchers.&lt;/strong&gt; The welfare-maximizing voucher program features a 100 percent subsidy rate, targets households with children and wages below the 80th percentile (fourth quintile), and is financed by progressive labor taxes. In the long-run steady state this policy raises 12.5 percent more children in the advantaged neighborhood, increases labor productivity by 1.1 percent, reduces income inequality (variance of log after-tax lifetime earnings) by 6.3 percent—comparable in magnitude to the Sweden–U.S. after-tax inequality gap—and raises upward mobility by 27.7 percent (roughly half its standard deviation across U.S. Census tracts). The average marginal tax rate must increase by 15.7 percent to fund the program. Despite this, long-run welfare rises by 3.4 percent in consumption equivalence units. A decomposition shows that intergenerational dynamics add 11.5 percentage points to welfare (relative to a short-run, single-generation scenario), while taxation subtracts 10.2 percentage points, and rent plus neighborhood-quality effects together subtract only 1.4 percentage points—leaving the net long-run GE gain similar to the short-run partial-equilibrium gain of 3.5 percent. Crucially, non-targeting children generates welfare losses of 5.0 percent, confirming that restriction to households with children is essential.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings — place-based wage subsidies.&lt;/strong&gt; A 12 percent wage subsidy to workers in the disadvantaged neighborhood yields the highest steady-state welfare gain of 0.7 percent. This is approximately one-fifth of the gain achievable with the optimal voucher. The subsidy induces substantial resorting toward n=1, reducing the share of children in n=2 by 6.7 percent while raising neighborhood quality in n=1 by 19.7 percent. Income inequality falls by 8.7 percent and upward mobility rises by 20.4 percent. However, in a short-run partial-equilibrium setup, the wage subsidy has a negative welfare effect of −1.0 percent because it draws parents (and their children) into the disadvantaged area; the positive net effect only emerges through long-run intergenerational channels (+2.5 percentage points) and equilibrium neighborhood-quality adjustments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Political economy.&lt;/strong&gt; Because voucher gains are concentrated among young cohorts (those aged 16–43 at introduction), only 33 percent of incumbent adults would rationally vote for the housing voucher program. In contrast, the place-based wage subsidy provides positive average welfare gains for all age cohorts alive at introduction, yielding estimated majority support from over 63 percent of adults. This creates a fundamental political economy tradeoff: the policy with the larger long-run social gains lacks majority democratic support, while the policy with broader support delivers smaller long-run gains.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-two-market-frictions-that-justify-government-intervention-in-the-model"&gt;Q1. What are the two market frictions that justify government intervention in the model?&lt;/h3&gt;
&lt;p&gt;A1: The first friction is the absence of intergenerational borrowing markets: parents cannot borrow against their child&amp;rsquo;s future income, which limits the parent&amp;rsquo;s willingness to pay the higher rent in n=2 to give their child a developmental advantage. Housing vouchers act as a tax-financed substitute for this missing contract by paying the rent premium and recovering the cost through taxes on the high-earning adults the children become. The second friction is a neighborhood externality: individuals do not internalize the effect of their own income on the neighborhood quality experienced by neighbors&amp;rsquo; children. Place-based wage subsidies partially correct this externality by subsidizing work in the disadvantaged area, raising local income per capita and thereby improving the neighborhood quality index for all children resident there.&lt;/p&gt;
&lt;h3 id="q2-how-is-neighborhood-quality-defined-and-modeled-and-why-is-this-specification-chosen"&gt;Q2. How is neighborhood quality defined and modeled, and why is this specification chosen?&lt;/h3&gt;
&lt;p&gt;A2: Neighborhood quality sn is defined as total income per capita (the sum of labor and capital income) for all residents of neighborhood n, including non-workers. This specification is intended to capture multiple mechanisms: school quality (which depends on local tax bases), role-model effects from productive adults, and social organization effects through adult supervision of children. The formulation includes retired and non-working residents, which means the arrival of children mechanically reduces neighborhood quality per capita in the model, partially capturing a crowding channel. Formally, the neighborhood spillover function takes the power form f(sn) = A * sn^ζ, where ζ governs the elasticity of child development to neighborhood quality.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-validate-the-models-key-mechanism--the-neighborhood-effect-on-children"&gt;Q3. How does the paper validate the model&amp;rsquo;s key mechanism — the neighborhood effect on children?&lt;/h3&gt;
&lt;p&gt;A3: The validation mimics the MTO RCT within the calibrated model: the government provides a 100 percent rent voucher usable only in n=2 to households in n=1 with incomes below the 10th percentile, holding prices and neighborhood qualities fixed (as in a small-scale experiment). The model generates 25 percent voucher take-up and a 23 percent increase in children&amp;rsquo;s income in their late 20s. This compares to the experimental MTO estimate of approximately 31 percent. The paper attributes most of the gap to the smaller poverty-rate contrast in the simulation (18 percentage points) relative to MTO (22 percentage points), and shows that plotting the simulated result against the site-specific MTO estimates in a scatterplot of child income gains against neighborhood poverty reductions places the model prediction on the fitted line through the experimental data.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-quantitative-role-of-long-run-intergenerational-dynamics-in-the-voucher-program-relative-to-other-ge-channels"&gt;Q4. What is the quantitative role of long-run intergenerational dynamics in the voucher program, relative to other GE channels?&lt;/h3&gt;
&lt;p&gt;A4: The decomposition in Table 5 isolates four GE channels. Starting from a short-run partial-equilibrium welfare gain of 3.5 percent (for the children of a single treated generation), allowing the economy to operate for the long run while holding prices and taxes fixed raises welfare to 15.0 percent — an increase of 11.5 percentage points — because improved skills in one generation create higher-skilled, higher-income parents who invest more in the next generation. Introducing housing market price adjustments (rents rise by 3.9 percent in n=2) reduces welfare by only 0.6 percentage points. Allowing neighborhood quality to adjust (quality in n=2 falls by 4 percent as lower-income families move in) reduces welfare by an additional 0.8 percentage points. Adding full taxation to balance the government budget reduces welfare by 10.2 percentage points, from 13.6 to 3.4 percent. The four channels nearly cancel, leaving the long-run GE steady-state gain close to the short-run single-generation gain.&lt;/p&gt;
&lt;h3 id="q5-why-does-the-optimal-voucher-program-require-targeting-to-families-with-children-and-what-happens-without-this-restriction"&gt;Q5. Why does the optimal voucher program require targeting to families with children, and what happens without this restriction?&lt;/h3&gt;
&lt;p&gt;A5: When the voucher is extended to all households regardless of children (Column 6 of Table 4), nearly 82.6 percent of the population receives a subsidy, pushing almost everyone to n=2. Rents in n=2 rise by 5.3 percent. To finance this much broader program, the average marginal tax rate must increase by 44 percent, far exceeding the 15.7 percent required for the children-targeted program. The large tax increase suppresses labor supply and income, which reduces neighborhood quality in n=2 by 11.6 percent. The net effect is a welfare loss of 5.0 percent. The intuition is that the benefit of the voucher program flows primarily through child skill development, so subsidizing adults without children is fiscally expensive without producing the intergenerational gains that justify the cost.&lt;/p&gt;
&lt;h3 id="q6-what-drives-the-difference-in-long-run-welfare-gains-between-vouchers-34-percent-and-place-based-wage-subsidies-07-percent"&gt;Q6. What drives the difference in long-run welfare gains between vouchers (3.4 percent) and place-based wage subsidies (0.7 percent)?&lt;/h3&gt;
&lt;p&gt;A6: The primary channel is labor productivity. The optimal voucher program raises labor productivity by 1.1 percent by increasing the average neighborhood quality to which children are exposed by 1.2 percent. The wage subsidy raises productivity by only 0.2 percent because it induces resorting toward the disadvantaged neighborhood, meaning children&amp;rsquo;s average neighborhood quality actually decreases by 0.2 percent despite large improvements in n=1&amp;rsquo;s quality (up 19.7 percent), since fewer children reside in n=1 after the subsidy draws their parents there. Inequality reduction is not the source of the gap: the wage subsidy actually reduces inequality more (8.7–8.9 percent) than the voucher (6.3 percent), but this inequality effect does not translate into larger aggregate welfare because productivity effects dominate.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-wage-subsidy-produce-positive-long-run-welfare-when-it-generates-negative-welfare-in-the-short-run"&gt;Q7. How does the wage subsidy produce positive long-run welfare when it generates negative welfare in the short run?&lt;/h3&gt;
&lt;p&gt;A7: In the short run, the wage subsidy draws parents into the disadvantaged neighborhood to exploit higher wages, which reduces the share of children in the advantaged neighborhood n=2 and lowers children&amp;rsquo;s late-life productivity (welfare of −1.0 percent for treated children in the single-generation scenario). Two long-run channels flip the sign. First, the subsidy is permanent, so children themselves receive it as adults, providing a direct wage income benefit. Second, the sustained presence of higher-income workers in n=1 raises neighborhood quality there durably (by 19.7 percent at the steady state), which benefits the children who reside in n=1. Together these intergenerational effects add 2.5 percentage points to welfare, while taxation costs reduce it by only 1.4 percentage points, yielding a net gain of 0.7 percent.&lt;/p&gt;
&lt;h3 id="q8-what-determines-the-political-economy-divide-between-the-two-policies"&gt;Q8. What determines the political economy divide between the two policies?&lt;/h3&gt;
&lt;p&gt;A8: For the housing voucher, welfare gains are concentrated among younger incumbent adults (ages 16–43), particularly those who are about to have or already have children, while older adults tend to lose because they face higher taxes without benefiting from improved neighborhood quality for their (now independent) children. This concentration implies only 33 percent of incumbent adults would support the voucher under the model&amp;rsquo;s welfare metric. For the place-based wage subsidy, average welfare gains are positive for every age cohort alive at introduction (though larger for younger cohorts), because the wage subsidy raises incomes for workers in n=1 immediately and benefits from equilibrium rent declines in n=1 that allow all residents to benefit. Over 63 percent of adults would support the wage subsidy. The paper notes that if the government could borrow to initially finance the voucher program and pay for it later (as in Daruich 2020 for early childhood programs), majority support for the voucher could potentially be achieved.&lt;/p&gt;
&lt;h3 id="q9-how-sensitive-are-the-welfare-results-to-the-key-calibrated-parameters"&gt;Q9. How sensitive are the welfare results to the key calibrated parameters?&lt;/h3&gt;
&lt;p&gt;A9: The sensitivity analysis (Table 9, following Andrews et al. 2017) shows that individual parameters would need to change substantially to overturn the conclusion that vouchers generate larger steady-state welfare gains than wage subsidies. For example, the altruism parameter β̃ would need to increase by 22 percent to eliminate the voucher welfare gain, which would require average parental transfers to rise to 198 percent of income — far from the empirical target of 125.4 percent. Using the more conservative tract-level housing supply elasticity from Baum-Snow and Han (2021) of 0.3–0.4 (about 80 percent below the baseline Saiz 2010 estimate of 1.75) would reduce the voucher welfare gain from 3.37 to approximately 2.57 percent, not reversing the qualitative conclusion. The parameters with the largest influence on welfare gains are the labor disutility parameter µ and the altruism parameter β̃; the housing supply elasticity matters more for the voucher than the wage subsidy because easier housing supply accommodates growth in n=2 without displacement under the voucher.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-transition-path-of-the-voucher-program-look-like-and-why-do-welfare-gains-initially-dip-before-recovering"&gt;Q10. What does the transition path of the voucher program look like, and why do welfare gains initially dip before recovering?&lt;/h3&gt;
&lt;p&gt;A10: When the voucher is unexpectedly introduced, the first newborn cohort gains approximately 4 percent welfare, but gains for subsequent cohorts initially dip to around 3 percent before stabilizing at 3.4 percent by the 20th post-introduction cohort. The dip occurs because moving costs slow resorting: immediately after introduction, rents in n=2 begin rising and neighborhood quality there begins falling as low-income families move in, but the capital stock adjustment (which would counteract these effects by raising GDP) lags the resorting. The rebound comes as capital accumulates in n=2 over time and as intergenerational productivity gains build through successive cohorts of better-skilled parents. Labor productivity jumps noticeably for the first cohort born to parents who received the voucher (approximately 28 years after introduction) and again for the first cohort born to grandparents who received it, visibly demonstrating the intergenerational mechanism. In contrast, the wage subsidy&amp;rsquo;s welfare gains are approximately constant at 0.7 percent across all cohorts because the key channels (neighborhood quality improvement in n=1 and wage gains) materialize rapidly and remain stable throughout the transition.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Neighborhood quality (sn):&lt;/strong&gt; In this paper, neighborhood quality is not school quality or amenities in a generic sense but is explicitly defined as total income per capita — the sum of labor income and capital income — for all residents of neighborhood n, including non-workers. This endogenous measure rises when higher-income or more productive residents move in and falls when lower-income residents or additional children arrive.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intergenerational borrowing constraint:&lt;/strong&gt; The inability of parents to borrow against their child&amp;rsquo;s future income, modeled as a non-negativity constraint on the monetary transfer from parent to child (transfer ≥ 0). This is the paper&amp;rsquo;s first key market friction: without it, a poor parent who moved to a better neighborhood would smooth consumption across generations by having the high-earning child compensate the parent. The constraint prevents this, reducing parental investment below the socially efficient level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption equivalence (veil of ignorance):&lt;/strong&gt; The welfare metric used throughout the policy analysis. It is defined as the percentage change in consumption that would make a newborn individual indifferent between the pre-policy and post-policy steady states, computed before knowing their position in the skill or income distribution. This is the paper&amp;rsquo;s measure of long-run steady-state welfare.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parental investment aggregator (CES):&lt;/strong&gt; A nested constant-elasticity-of-substitution function that determines how parental time τ and neighborhood quality sn combine to form the effective investment input I into child skill development: I = Ā[αI f(sn)^γ + (1 − αI)τ^γ]^(1/γ). The elasticity parameter 1/(1 − γ), estimated at 0.41, governs the degree of complementarity between time and neighborhood quality; a lower elasticity (γ = −1.43) implies the two inputs are complements, so parents with children in better neighborhoods also spend more time with them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Place-based wage subsidy:&lt;/strong&gt; A neighborhood-specific wage premium (denoted w̃s) paid to all workers who both live and work in the disadvantaged neighborhood n=1, raising their effective wage to w1 = (1 + w̃s)w2. This policy targets the neighborhood externality by increasing the income of residents in n=1, which raises neighborhood quality and provides an incentive for higher-skilled workers to relocate to (or remain in) the disadvantaged area.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Upward mobility:&lt;/strong&gt; Measured in this paper as the probability that a child born to parents in the bottom 20 percent of the income distribution reaches the top 20 percent of the income distribution during the working stage of their own life. This is distinct from mean income rank measures; it specifically tracks cross-quintile transitions in the model&amp;rsquo;s stationary distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Equilibrium decomposition:&lt;/strong&gt; A simulation-based method in which GE channels are progressively activated. Starting from a short-run, partial-equilibrium, single-generation baseline (analogous to an RCT), the authors sequentially allow: (i) long-run intergenerational dynamics while holding prices fixed; (ii) housing market price adjustments; (iii) neighborhood quality adjustments; (iv) tax and production-price adjustments. Each step&amp;rsquo;s change in outcomes identifies the quantitative contribution of that specific channel.&lt;/p&gt;</description></item><item><title>Are Inflationary Shocks Regressive? A Feasible Set Approach</title><link>https://macropaperwarehouse.com/papers/are-inflationary-shocks-regressive-a-feasible-set-approach/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/are-inflationary-shocks-regressive-a-feasible-set-approach/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; The paper asks whether inflationary shocks are regressive, and demonstrates that the answer depends critically on the &lt;em&gt;source&lt;/em&gt; of the shock. A single aggregate inflation statistic conceals radically different distributional consequences depending on whether inflation is driven by an oil supply contraction or by expansionary monetary policy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Framework.&lt;/strong&gt; The authors develop a &amp;ldquo;feasible set approach&amp;rdquo; grounded in the envelope theorem. They show that the first-order money-metric welfare effect of any macroeconomic shock on a household is summarized by the present discounted value of changes to five components of the household&amp;rsquo;s budget constraint: (1) consumption prices, (2) wage income, (3) asset dividends, (4) asset prices, and (5) government transfers. Because the envelope theorem implies that endogenous substitution responses are not welfare-relevant to a first order, no assumption about the utility function&amp;rsquo;s form or the economy&amp;rsquo;s general equilibrium structure is required. The framework is valid for generic stationary shocks that do not directly shift household preferences.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Strategy.&lt;/strong&gt; The welfare formula requires two inputs: (i) impulse response functions (IRFs) for all prices, dividends, wages, and unemployment, estimated using internal-instrument SVAR methods applied to two identified shocks — the Kanzig (2021) oil supply news shock (instrumented by oil futures surprises around OPEC announcements) and the Gertler-Karadi (2015) monetary policy shock (instrumented by fed funds futures surprises in 30-minute windows around FOMC announcements) — and (ii) cross-sectional data on consumption bundles, labor income, and asset portfolios from the CEX, CPS, SCF, and SIPP for three education groups (high school or less, some college, college-educated) across the full lifecycle. The baseline cross-section uses 2019 data. Shocks are normalized to produce comparable aggregate inflation responses: a 10% WTI oil price increase and a 25 basis point decline in the one-year Treasury yield each generate roughly 15–16 basis points of CPI-U inflation on impact, rising to approximately 34–35 basis points after two quarters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings.&lt;/strong&gt; Oil supply contractions are regressive and monetary expansions are progressive, and this divergence is primarily driven by the asset price channel, not the consumption price or labor income channels.&lt;/p&gt;
&lt;p&gt;For the 10% oil supply shock: middle-aged households with high school education or less must be paid approximately $870 (around 2% of annual consumption) to be made whole relative to their pre-shock utility; college-educated middle-aged households, by contrast, gain the equivalent of approximately $833 (1.1% of annual consumption). Younger college-educated households (still net equity accumulators) gain around $572.&lt;/p&gt;
&lt;p&gt;For the 25 basis point monetary rate cut: low-education households approximately break even (net welfare effect near $23), while middle-aged college-educated households must be paid approximately $4,051 (around 5.5% of annual consumption) to restore their pre-shock utility. Older college-educated households must be paid approximately $851.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why asset prices dominate.&lt;/strong&gt; Oil supply contractions reduce equity prices (S&amp;amp;P500 falls approximately 2% one year post-shock) and depress dividends (approximately 82 basis points), while leaving house prices and bond prices largely unaffected. Because middle-aged college-educated households are the primary accumulators of equities, they benefit from the price decline (cheaper future accumulation), making oil shocks progressive through this channel — but regressive overall once the consumption and labor income channels (both mildly regressive) are included. Monetary expansions do the opposite: equity prices rise approximately 3 percentage points on impact, house prices rise approximately 1.5% after three years, and dividends increase. These asset price increases hurt those in the accumulation phase — disproportionately middle-aged college-educated households — creating a progressive distributional pattern.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption and labor income channels.&lt;/strong&gt; Both shocks generate disproportionate inflation in motor fuel and fuel and utilities, and low-education households spend a larger share of their budget on these goods, making the consumption channel mildly regressive for both shocks. The labor income channel differs sharply: oil shocks raise unemployment (approximately 0.15 log points for low-education households two years post-shock) and reduce weekly earnings by 0.2–0.6 log points, mildly harming low-education workers; monetary expansions reduce unemployment (approximately 0.83 log points for low-education workers one year post-shock) and similarly benefit low-education households through the labor market, pushing toward progressivity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; Results apply to short-run first-order welfare effects of identified stationary macroeconomic shocks (four-year horizon). The framework does not incorporate uncertainty shocks, preference shocks, or the role of hedging motives in portfolio choice. Results concern policy shocks rather than policy rules.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Robustness.&lt;/strong&gt; Qualitative conclusions hold across six alternative specifications: incorporating borrowing constraints (with or without empirical death rates), adjusting for unemployment insurance replacement rates (approximately 6% true average replacement rate), allowing for log-linear trends in no-shock choices, and dropping aggregate CPI controls from IRF estimation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-feasible-set-approach-and-how-does-it-differ-from-prior-work-on-inflation-incidence"&gt;Q1. What is the &amp;ldquo;feasible set approach&amp;rdquo; and how does it differ from prior work on inflation incidence?&lt;/h3&gt;
&lt;p&gt;A: The feasible set approach measures welfare effects through changes in the household&amp;rsquo;s entire budget constraint — consumption prices, wage income, asset dividends, asset prices, and government transfers — rather than focusing on any single channel. Prior work either examined the Fisher channel (net nominal positions), or consumption price heterogeneity, or labor income responses in isolation. The key insight is that the envelope theorem implies substitution responses are not welfare-relevant to a first order, so the money-metric welfare change is simply the discounted sum of changes in the five budget constraint components evaluated at pre-shock choices, without requiring knowledge of the utility function&amp;rsquo;s form or the economy&amp;rsquo;s general equilibrium structure.&lt;/p&gt;
&lt;h3 id="q2-why-is-the-asset-price-channel--rather-than-consumption-prices--the-dominant-channel-in-both-shocks"&gt;Q2. Why is the asset price channel — rather than consumption prices — the dominant channel in both shocks?&lt;/h3&gt;
&lt;p&gt;A: Asset holdings are large relative to annual consumption (net worth averages $1.5 million for college-educated and $260,000 for high-school-educated households in 2019), so even modest percentage movements in asset prices generate large dollar welfare effects. By contrast, the budget shares on the goods most responsive to both shocks (motor fuel, fuel and utilities) are relatively modest, so the consumption channel, while mildly regressive, is quantitatively small relative to the portfolio channel. The portfolio channel accounts for roughly 0.5% of consumption gains for middle-aged college-educated households under the oil shock, while the consumption channel produces losses of only about 0.1% for college-educated and 0.25% for low-education households.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-direction-of-the-equity-price-response-differ-between-oil-and-monetary-shocks-and-why-does-this-create-opposite-distributional-effects"&gt;Q3. How does the direction of the equity price response differ between oil and monetary shocks, and why does this create opposite distributional effects?&lt;/h3&gt;
&lt;p&gt;A: An oil supply contraction reduces equity prices (approximately 2% decline one year post-shock) and dividends (approximately 82 basis points decline), while a monetary expansion raises equity prices (approximately 3 percentage points on impact, approximately 4% higher after four quarters) and increases dividends. The welfare effect of asset price changes falls on those who &lt;em&gt;trade&lt;/em&gt; the asset, not those who merely hold it at a constant level: middle-aged college-educated households are the primary net &lt;em&gt;accumulators&lt;/em&gt; of equity, so falling prices benefit them (they can buy more cheaply) while rising prices hurt them. This is the principal reason oil shocks appear progressive through the portfolio channel — but regressive overall — while monetary expansions are regressive through the portfolio channel and progressive overall.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-precise-welfare-numbers-for-oil-supply-shocks-by-education-group-baseline-ages-2265"&gt;Q4. What are the precise welfare numbers for oil supply shocks by education group (baseline, ages 22–65)?&lt;/h3&gt;
&lt;p&gt;A: From Table 3 (baseline row, lifecycle-weighted averages for ages 25–65): households with high school or less experience a welfare loss of approximately $798; those with some college experience a loss of approximately $816; and college-educated households experience a welfare &lt;em&gt;gain&lt;/em&gt; of approximately $494. These numbers reflect the sum of the consumption, labor income, portfolio, and transfer channels over a 16-quarter horizon, discounted at the one-year Treasury yield.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-precise-welfare-numbers-for-monetary-policy-shocks-by-education-group-baseline-ages-2565"&gt;Q5. What are the precise welfare numbers for monetary policy shocks by education group (baseline, ages 25–65)?&lt;/h3&gt;
&lt;p&gt;A: From Table 3 (baseline row): households with high school or less experience a small welfare &lt;em&gt;gain&lt;/em&gt; of approximately $23; those with some college experience a welfare loss of approximately $1,278; and college-educated households experience a welfare loss of approximately $3,055. These losses for college-educated households are driven overwhelmingly by rising equity and house prices that raise the cost of planned asset accumulation.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-life-cycle-interact-with-the-distributional-incidence-of-both-shocks"&gt;Q6. How does the life cycle interact with the distributional incidence of both shocks?&lt;/h3&gt;
&lt;p&gt;A: There is substantial heterogeneity within education groups across the life cycle because asset accumulation and decumulation patterns are age-dependent. Under oil shocks, younger college-educated households (who are net equity accumulators) gain approximately $572, middle-aged college-educated households gain approximately $833, while older college-educated households lose approximately $69 (because they hold large equity positions and lose dividend income). Under monetary shocks, middle-aged college-educated households lose the most (approximately $4,051) because they are simultaneously accumulating equities and housing, both of which become more expensive. Older college-educated households lose less (approximately $851) because rising dividends on existing holdings partially offset the asset price cost. Low-education households are approximately flat across the life cycle under monetary shocks.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-consumption-channel-compare-across-education-groups-and-across-the-two-shocks"&gt;Q7. How does the consumption channel compare across education groups and across the two shocks?&lt;/h3&gt;
&lt;p&gt;A: The consumption channel is mildly regressive for both shocks, but of similar absolute magnitude across the two shocks because both generate similar inflation in motor fuel and fuel and utilities — the goods with the largest price response. Low-education households spend a larger share on motor fuel and fuel and utilities; as a result, they lose approximately 0.25% of consumption from the consumption channel under the oil shock, compared with less than 0.1% for college-educated households. For monetary shocks, the consumption channel affects all household types roughly equally in proportional terms.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-labor-income-channel-differ-between-oil-and-monetary-shocks-across-education-groups"&gt;Q8. How does the labor income channel differ between oil and monetary shocks across education groups?&lt;/h3&gt;
&lt;p&gt;A: Oil shocks raise unemployment disproportionately for low-education workers (approximately 0.15 log point increase after two years, roughly 0.68 standard deviations, compared with near-zero response for college-educated workers) and reduce weekly earnings by 0.2–0.6 log points across groups. Monetary expansions reverse this: a 25 basis point rate cut reduces log unemployment by approximately 0.83 log points for low-education workers and approximately 1.96 log points for college-educated workers after one year, with limited response in conditional wages. Thus the labor income channel pushes toward regressive incidence for oil shocks and toward progressive incidence for monetary expansions, though in both cases it is quantitatively smaller than the portfolio channel.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-role-of-housing-in-the-portfolio-channel"&gt;Q9. What is the role of housing in the portfolio channel?&lt;/h3&gt;
&lt;p&gt;A: Housing behaves simultaneously as a durable consumption good and a financial asset. A house price increase raises welfare for households planning to &lt;em&gt;decumulate&lt;/em&gt; (sell) housing (primarily older households) through the portfolio channel, but also raises the implicit rental cost for those who &lt;em&gt;use&lt;/em&gt; housing — a negative consumption-side effect. Monetary expansions raise house prices by approximately 1.5% after three years. College-educated households accumulate housing at a faster rate and earlier in the life cycle than low-education households, making them more exposed to the cost of rising house prices during the accumulation phase. This amplifies the progressive pattern of monetary shocks through the portfolio channel.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-paper-handle-the-dual-role-of-durable-goods-vehicles-and-housing"&gt;Q10. How does the paper handle the dual role of durable goods (vehicles and housing)?&lt;/h3&gt;
&lt;p&gt;A: Durable goods are treated as both a consumption good and a financial asset. The utility-relevant consumption price of a durable is proportional to the price times the depreciation rate per unit of use, capturing the &amp;ldquo;implicit rent&amp;rdquo; of ownership. On the asset side, the durable enters the portfolio channel like a zero-dividend financial asset. This allows the framework to correctly attribute, for example, that a rise in house prices hurts net accumulators (through the portfolio channel) while also raising the implicit cost of housing services (through the consumption channel), rather than treating house price appreciation as an unambiguous welfare gain for homeowners.&lt;/p&gt;
&lt;h3 id="q11-what-happens-to-the-main-conclusions-when-borrowing-constraints-are-introduced"&gt;Q11. What happens to the main conclusions when borrowing constraints are introduced?&lt;/h3&gt;
&lt;p&gt;A: Incorporating net worth constraints (with either constant or empirical death rates) dampens the portfolio channel for young and middle-aged college-educated households, because rising asset prices relax borrowing constraints for these households, partially offsetting the welfare cost of more expensive accumulation. Under constant death rates with borrowing constraints, college-educated households&amp;rsquo; oil shock welfare gain falls from +$494 to +$76; under empirical death rates, it becomes a loss of -$394. For monetary shocks, the college-educated loss falls from -$3,055 to -$1,718 (constant death rate) or -$1,036 (empirical death rates). Despite these quantitative changes, the qualitative conclusion — oil shocks are regressive, monetary expansions are progressive — holds across all specifications.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-implication-of-these-findings-for-the-policy-interaction-between-oil-shocks-and-monetary-tightening"&gt;Q12. What is the implication of these findings for the policy interaction between oil shocks and monetary tightening?&lt;/h3&gt;
&lt;p&gt;A: If the monetary authority responds to oil-price-induced inflation with unexpected interest rate increases, it may exacerbate the distributional consequences of the initial oil shock. An oil supply contraction is already regressive (harming low-education households through consumption prices and labor market effects); a disinflationary monetary tightening would additionally harm low-education households through the labor income channel (higher unemployment, lower wages) while partially benefiting college-educated households through lower asset prices. The paper notes this policy interaction as noteworthy, while cautioning that the results concern identified policy &lt;em&gt;shocks&lt;/em&gt; rather than policy &lt;em&gt;rules&lt;/em&gt;.&lt;/p&gt;
&lt;h3 id="q13-how-are-the-two-shocks-calibrated-to-be-comparable"&gt;Q13. How are the two shocks calibrated to be comparable?&lt;/h3&gt;
&lt;p&gt;A: The oil shock is normalized to a 10% increase in WTI crude oil prices (approximately one standard deviation of monthly oil price growth). The monetary shock is normalized to a 25 basis point decline in the one-year Treasury yield — chosen because it generates approximately the same aggregate CPI-U inflation response as the oil shock (approximately 15–16 basis points on impact, rising to approximately 34–35 basis points after two quarters). This normalization allows the paper to attribute the different distributional outcomes to the &lt;em&gt;source&lt;/em&gt; of inflation rather than to differences in the aggregate inflation magnitude.&lt;/p&gt;
&lt;h3 id="q14-what-role-does-the-transfer-channel-play-and-for-whom"&gt;Q14. What role does the transfer channel play, and for whom?&lt;/h3&gt;
&lt;p&gt;A: The transfer channel is small relative to the other three channels for the vast majority of working-age households, because transfer income is less than $100 per month for most households under age 65. Social Security payments — the bulk of transfer income — are explicitly indexed to the CPI; the paper models them as moving with CPI with a one-year lag. The transfer channel exclusively benefits older households (those receiving Social Security), and its quantitative effect is modest even there. Transfer income is more than 20 times smaller than labor and asset income for prime-age households of all education groups.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Feasible set approach.&lt;/strong&gt; The paper&amp;rsquo;s organizing framework, in which the first-order welfare impact of a macroeconomic shock is measured by how the shock changes the household&amp;rsquo;s budget constraint (consumption prices, wage income, asset dividends, asset prices, and government transfers) evaluated at the household&amp;rsquo;s pre-shock choices. Substitution responses are not welfare-relevant to a first order by the envelope theorem.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Money-metric welfare gain.&lt;/strong&gt; The willingness-to-pay measure used throughout: the welfare change from a shock divided by the household&amp;rsquo;s marginal utility of consumption at time zero, expressed in time-zero dollars. Interpreted as an equivalent variation — the amount the household must be paid or would give up to be indifferent to receiving the shock. Used because it places households with very different utility functions on a common dollar scale.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Portfolio channel.&lt;/strong&gt; The component of the welfare formula capturing the effect of asset price and dividend changes on household welfare. Asset price changes are welfare-relevant only for households that &lt;em&gt;trade&lt;/em&gt; (accumulate or decumulate) the asset: rising prices benefit sellers and harm buyers; falling prices benefit buyers and harm sellers. This is distinct from the &amp;ldquo;Fisher channel&amp;rdquo; in prior literature, which focuses on net nominal positions rather than on which households are in the accumulation versus decumulation phase.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Internal instrument SVAR.&lt;/strong&gt; The time-series estimation procedure used throughout: the pre-estimated identified shock series (oil supply news or monetary policy surprise) is included as a variable ordered first in a recursive structural VAR for each outcome variable. This separates shock identification (using the published instruments and controls from Kanzig 2021 and Gertler-Karadi 2015) from IRF estimation for each outcome variable, allowing the use of the full available sample for each outcome series.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Oil supply news shock (Kanzig 2021).&lt;/strong&gt; An identified supply shock to oil markets, constructed from changes in oil price futures in tight windows around OPEC production announcements. Used to capture exogenous cost-push inflation driven by supply constraints rather than demand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Monetary policy shock (Gertler-Karadi 2015).&lt;/strong&gt; An identified demand-side shock, constructed from federal funds rate futures surprises in 30-minute windows around FOMC announcements, instrumented into a monetary SVAR. Captures exogenous interest rate cuts that generate aggregate demand expansion and inflation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Borrowing constraint wedge.&lt;/strong&gt; An additional term that appears in the welfare formula when households face net worth constraints. Proportional to the Lagrange multiplier on the net worth constraint, it discounts future periods more heavily when constraints bind, and adds a term for the welfare value of relaxed constraints when asset prices rise. Identified from deviations from perfect consumption smoothing using CEX lifecycle consumption data.&lt;/p&gt;</description></item><item><title>Automated credit limit increases and consumer welfare</title><link>https://macropaperwarehouse.com/papers/automated-credit-limit-increases-and-consumer-welfare/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/automated-credit-limit-increases-and-consumer-welfare/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Should regulators restrict banks from proactively raising credit card limits using machine-learning algorithms, and if so, how? The paper asks: to what extent are bank-initiated credit limit increases directed toward revolving borrowers (those who carry interest-accruing balances month-to-month), and what are the welfare consequences of policies that constrain such increases?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The empirical analysis uses the Federal Reserve&amp;rsquo;s Capital Assessments and Stress Testing (Y-14M) regulatory data, January 2014 to December 2024, covering monthly account-level records for all credit cards issued by large stress-tested banks (assets &amp;gt; $100B). The 26 banks in the sample collectively represent more than 70% of U.S. credit card balances. A 0.5% sample yields more than 150 million observations across more than 3.6 million unique active credit cards. A key advantage of Y-14 over credit bureau data is that it identifies whether each limit change was bank-initiated or consumer-initiated — a distinction not available in other datasets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stylized Facts.&lt;/strong&gt; Credit limit increases are an important and understudied source of consumer credit. During the post-pandemic period, limit increases generate more than $40 billion of additional available credit per quarter, roughly 60% of the approximately $70 billion coming from new card originations; prior to the pandemic the figure was about $30 billion, or roughly half of new issuance. The number of accounts undergoing a limit increase each quarter is on average 30% higher than the number of new cards issued. Consistent with &amp;ldquo;low-and-grow&amp;rdquo; lending strategies, limit increases are disproportionately important for lower credit-score borrowers: average subprime credit limits rise from $700 at origination to $2,700 by five years after origination (a 285% increase) and to nearly $5,000 by eight years, while average superprime limits rise only from approximately $12,000 to $15,000 (a 25% increase). About 30% of total revolving balances are made possible by limit increases, with the share reaching 60% for subprime borrowers but only 12% for superprime borrowers. Approximately 75–80% of all limit increases — both by dollar amount and by number of cards — are bank-initiated rather than consumer-initiated. Banks that more frequently reference &amp;ldquo;artificial intelligence&amp;rdquo; or &amp;ldquo;machine learning&amp;rdquo; in their 10-K filings support a larger share of revolving balances through limit increases. Bank-initiated increases are roughly 1.5–2 times more prevalent among accounts that have revolved in the prior three months, whereas consumer-initiated increases show essentially no differential by revolving status.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Analysis.&lt;/strong&gt; Using a linear probability model with card-portfolio-group fixed effects, month fixed effects, and controls for credit score, income, prior limit changes, and other account characteristics, the authors show that the probability of a bank-initiated limit increase follows an inverse-U shape in revolving utilization: accounts with revolving utilization in the moderate range (roughly 0.2–0.7) are most likely to receive an increase, while those near zero or near 1.0 are not. An account with revolving utilization in the (0.2, 0.3] bin is approximately as likely to receive a limit increase as an account whose credit score just rose by 66 points. Transacting utilization, by contrast, follows a logistic growth pattern: the probability rises monotonically until about a utilization of 0.3 and is flat above that. An event study shows that after a bank-initiated limit increase, revolving utilization rebounds to its pre-increase level within approximately 8 months; on average, revolving balances increase by about 40% of the limit increase, with approximately 30% of the limit increase going toward revolving balances. This rebound occurs even for accounts with revolving utilization below the pre-increase mean of 0.28, indicating that the effect is not confined to liquidity-constrained borrowers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The authors develop a life-cycle consumption–saving model with credit card borrowing, uninsurable income and employment risk, potential default (Chapter 7 style), and heterogeneous preferences following Nakajima (2017) and Gul–Pesendorfer (2001, 2004). Two household types coexist: 60% with standard exponential-discounting preferences (calibrated β = 0.92) and 40% with temptation preferences (calibrated β = 0.96, temptation parameter λ = 0.28 from Kovacs et al., 2021). The credit limit increase function is calibrated using Y-14M data via a latent-variable formulation, replicating the empirical inverted-U relationship between revolving utilization and limit increase probability. The four internally calibrated targets are: share of households with revolving credit card debt (data: 45%, model: 41.8%); utilization rate conditional on debt (data: 35%, model: 28.9%); default probability (data: 0.94%, model: 0.94%); debt-to-income ratio (data: 8.6%, model: 6.8%).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings — Baseline.&lt;/strong&gt; Through the model, tempted agents are disproportionately likely to receive credit limit increases because they are more likely to revolve. For customers with utilization above 50%, the majority of credit limit increases are detrimental from the borrower&amp;rsquo;s own perspective. Standard agents almost always benefit from higher credit limits.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Counterfactual 1 — UK-style (prohibit limit increases for revolving borrowers).&lt;/strong&gt; This policy reduces the annual probability of limit increases from roughly 5.5% to approximately 1.0%. The default probability falls from about 0.9% to near zero. The debt-to-income ratio declines by roughly 2 percentage points. Aggregate welfare improves by 1.12% in consumption equivalent variation (CEV) when the social planner internalizes the psychological cost of temptation (0.98% without). Standard households incur a modest welfare loss of 0.21% from reduced consumption-smoothing flexibility, while tempted households gain approximately 3.12% in CEV.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Counterfactual 2 — Canada/EU-style (require consumer consent).&lt;/strong&gt; This policy reduces the annual limit-increase probability from 5.5% to approximately 1.9%. Aggregate welfare improves by 1.16% in CEV (1.04% without psychological costs). Standard households lose 0.19%, while tempted households gain approximately 3.19%. Under the baseline assumption of sophisticated tempted households, results are nearly identical to the UK-style policy. However, when the fraction of naïve tempted households is large, the consent-based policy becomes ineffective (naïve consumers accept limit increases they will regret), whereas the UK-style revolving-borrower ban remains welfare-improving regardless of the naïve share.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Robustness.&lt;/strong&gt; When the firm is allowed to re-optimize its credit limit increase policy, it endogenously reallocates more limit increases toward standard consumers. Welfare gains remain positive but are attenuated: the UK-style policy yields 0.21% CEV (vs. 1.12% in the baseline calibration) and the consent-based policy yields 0.27% CEV.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy Implications.&lt;/strong&gt; The U.S. lacks regulation of bank-initiated proactive credit limit increases (existing rules under ECOA and ability-to-pay provisions are largely non-binding for this purpose). The authors conclude that banks&amp;rsquo; revealed preference for targeting revolvers constitutes an implicit targeting of consumers with self-control issues, and that if a meaningful share of households have self-control issues, there are strong consumer protection grounds for regulating algorithmic credit limit increases.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-do-the-authors-use-y-14m-data-rather-than-credit-bureau-data-and-what-does-this-data-uniquely-enable"&gt;Q1. Why do the authors use Y-14M data rather than credit bureau data, and what does this data uniquely enable?&lt;/h3&gt;
&lt;p&gt;A: The Y-14M dataset allows the authors to distinguish between bank-initiated and consumer-initiated credit limit changes — a distinction not observable in credit bureau data. It also contains actual payment information enabling identification of revolvers (those carrying interest-accruing balances) rather than just total balances. The sample covers more than 70% of U.S. credit card balances and more than 150 million monthly observations over the January 2014 to December 2024 period.&lt;/p&gt;
&lt;h3 id="q2-how-large-are-credit-limit-increases-relative-to-new-card-originations-in-the-us-credit-card-market"&gt;Q2. How large are credit limit increases relative to new card originations in the U.S. credit card market?&lt;/h3&gt;
&lt;p&gt;A: During the post-pandemic period, limit increases produce more than $40 billion of additional available credit per quarter, roughly 60% of the approximately $70 billion created by new card originations. Prior to the pandemic the figure was approximately $30 billion, or about half of new issuance. On a count basis, the number of cards undergoing a limit increase each quarter is on average 30% higher than the number of new cards issued.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-low-and-grow-strategy-and-how-large-is-the-subsequent-credit-expansion"&gt;Q3. What is the &amp;ldquo;low-and-grow&amp;rdquo; strategy, and how large is the subsequent credit expansion?&lt;/h3&gt;
&lt;p&gt;A: The low-and-grow strategy involves originating higher-risk borrowers at low initial credit limits and then expanding limits based on observed borrowing behavior. For the average subprime credit card, the initial limit of $700 grows to $2,700 by five years after origination (a 285% increase) and to nearly $5,000 by eight years. For superprime borrowers, the initial limit of approximately $12,000 grows only to $15,000 (a 25% increase) by five years and then is approximately unchanged.&lt;/p&gt;
&lt;h3 id="q4-how-does-a-borrowers-revolving-status-affect-the-probability-of-receiving-a-bank-initiated-limit-increase"&gt;Q4. How does a borrower&amp;rsquo;s revolving status affect the probability of receiving a bank-initiated limit increase?&lt;/h3&gt;
&lt;p&gt;A: Bank-initiated increases are approximately 1.5–2 times more prevalent among accounts that have revolved at least once in the prior three months, compared to non-revolving accounts. By contrast, consumer-initiated increases show essentially no differential between revolvers and non-revolvers. This reveals a bank-side revealed preference for targeting revolvers.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-shape-of-the-relationship-between-revolving-utilization-and-the-probability-of-a-bank-initiated-limit-increase-and-how-large-is-its-economic-magnitude"&gt;Q5. What is the shape of the relationship between revolving utilization and the probability of a bank-initiated limit increase, and how large is its economic magnitude?&lt;/h3&gt;
&lt;p&gt;A: The relationship follows an inverted-U shape. Accounts with revolving utilization in bins between approximately 0.2 and 0.7 have the highest probability of receiving an increase; accounts near zero or near full utilization are as unlikely to receive an increase as zero-utilization accounts. The effect of being in the (0.2, 0.3] revolving utilization bin has approximately the same positive effect on the probability of receiving a limit increase as a 66-point increase in credit score, making it economically large relative to standard risk signals.&lt;/p&gt;
&lt;h3 id="q6-how-does-transacting-utilization-relate-to-bank-initiated-limit-increases-and-how-does-this-differ-from-revolving-utilization"&gt;Q6. How does transacting utilization relate to bank-initiated limit increases, and how does this differ from revolving utilization?&lt;/h3&gt;
&lt;p&gt;A: Transacting utilization follows a logistic growth pattern rather than an inverted-U. The probability of receiving a limit increase rises monotonically with transacting utilization until about a utilization of 0.3, above which the probability does not vary with utilization. This contrasts with revolving utilization, where very high utilization (above 0.9) is actually no more predictive than zero utilization.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-event-study-show-about-borrowing-behavior-following-credit-limit-increases"&gt;Q7. What does the event study show about borrowing behavior following credit limit increases?&lt;/h3&gt;
&lt;p&gt;A: After a bank-initiated limit increase, revolving utilization (as a share of the credit limit) drops mechanically but then rebounds to pre-increase levels within approximately 8 months. On average, revolving balances increase by about 40% of the amount of the limit increase, with approximately 30% of each dollar of new credit limit going toward revolving balances. These magnitudes are somewhat larger than the 13% (Gross and Souleles, 2002) and 18% (Aydin, 2022) found in prior work, which the authors attribute to the non-causal nature of their event study, higher average utilization in their sample, and their focus on revolving rather than total utilization.&lt;/p&gt;
&lt;h3 id="q8-is-the-post-increase-borrowing-rebound-driven-by-liquidity-constrained-borrowers"&gt;Q8. Is the post-increase borrowing rebound driven by liquidity-constrained borrowers?&lt;/h3&gt;
&lt;p&gt;A: No. The authors show that limiting the sample to accounts with revolving utilization below the pre-increase mean of 0.28 — accounts that are unlikely to be liquidity constrained — yields very similar results. This finding is consistent with the presence of self-control issues rather than binding credit constraints.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-key-modeling-assumptions-about-household-types-and-how-were-the-share-parameters-calibrated"&gt;Q9. What are the key modeling assumptions about household types, and how were the share parameters calibrated?&lt;/h3&gt;
&lt;p&gt;A: The model features two types: 60% with standard exponential-discounting preferences (estimated discount factor β = 0.92) and 40% with temptation preferences (β = 0.96, temptation parameter λ = 0.28 set from Kovacs et al., 2021). The 40% tempted share is internally estimated via the Method of Simulated Moments targeting four aggregate moments: share with revolving credit card debt (45% in data, 41.8% in model), utilization rate conditional on debt (35% vs. 28.9%), default probability (0.94% vs. 0.94%), and debt-to-income ratio (8.6% vs. 6.8%).&lt;/p&gt;
&lt;h3 id="q10-how-do-tempted-and-standard-households-differ-in-their-credit-card-usage-within-the-model"&gt;Q10. How do tempted and standard households differ in their credit card usage within the model?&lt;/h3&gt;
&lt;p&gt;A: In the model, 76% of tempted agents carry revolving credit card debt, with an average utilization rate of 73.6%, a debt-to-income ratio of 15.4%, and a default probability of 2.22%. Standard agents carry debt only 18.9% of the time, with average utilization of 4.1%, a debt-to-income ratio of 1.1%, and a default probability of 0.08%. Tempted agents also pay a substantially higher share of income on credit card interest.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-model-capture-the-mechanism-by-which-credit-limit-increases-harm-tempted-households"&gt;Q11. How does the model capture the mechanism by which credit limit increases harm tempted households?&lt;/h3&gt;
&lt;p&gt;A: The Gul–Pesendorfer temptation utility function makes household welfare depend on both actual consumption and the most tempting consumption alternative available (the budget-set maximum). When credit limits rise, the most tempting alternative ˜c_t increases, which raises the utility cost of self-restraint even for households that do not succumb to temptation. This mechanism is distinct from hyperbolic discounting: temptation imposes a psychic cost even on those who ultimately choose not to over-borrow.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-quantitative-welfare-effects-of-the-uk-style-policy-prohibiting-limit-increases-for-revolving-borrowers"&gt;Q12. What are the quantitative welfare effects of the UK-style policy prohibiting limit increases for revolving borrowers?&lt;/h3&gt;
&lt;p&gt;A: The policy yields an overall welfare gain of 1.12% in consumption equivalent variation (CEV) when the social planner internalizes the psychological cost of temptation (0.98% without). Standard households suffer a modest welfare loss of 0.21% from reduced consumption-smoothing flexibility. Tempted households gain approximately 3.12% in CEV, because the benefit from reduced temptation and lower interest expenditure outweighs the cost of reduced credit access.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-quantitative-welfare-effects-of-the-canadaeu-style-consent-required-policy"&gt;Q13. What are the quantitative welfare effects of the Canada/EU-style consent-required policy?&lt;/h3&gt;
&lt;p&gt;A: The consent-based policy yields an overall welfare gain of 1.16% in CEV (1.04% without psychological costs). Standard households lose 0.19%, and tempted households gain approximately 3.19%. Under the baseline assumption of fully sophisticated tempted households, results are nearly identical to the UK-style ban.&lt;/p&gt;
&lt;h3 id="q14-how-sensitive-are-the-two-policy-counterfactuals-to-the-share-of-naïve-unaware-of-their-self-control-issues-tempted-households"&gt;Q14. How sensitive are the two policy counterfactuals to the share of naïve (unaware of their self-control issues) tempted households?&lt;/h3&gt;
&lt;p&gt;A: The UK-style ban on limit increases for revolving borrowers remains welfare-improving regardless of whether tempted households are sophisticated or naïve — the welfare impact is approximately flat as the naïve fraction rises from zero to one. The consent-based policy, by contrast, exhibits a negative linear relationship between the naïve fraction and welfare impact, with welfare gains disappearing as the naïve fraction approaches one. Naïve consumers accept limit increases they would regret, so the policy&amp;rsquo;s effectiveness depends on households accurately recognizing their own self-control issues.&lt;/p&gt;
&lt;h3 id="q15-what-happens-when-the-firm-is-allowed-to-re-optimize-its-credit-limit-increase-policy-in-response-to-regulation"&gt;Q15. What happens when the firm is allowed to re-optimize its credit limit increase policy in response to regulation?&lt;/h3&gt;
&lt;p&gt;A: With firm re-optimization, both counterfactual policies continue to improve welfare but the magnitudes are attenuated. The UK-style policy yields 0.21% CEV overall (tempted: 0.89%) and the consent-based policy yields 0.27% overall (tempted: 0.98%), compared to 1.12% and 1.16% without re-optimization. The re-optimizing firm reallocates more limit increases toward standard consumers, which reduces the number directed at tempted households but also limits the welfare gains from regulation.&lt;/p&gt;
&lt;h3 id="q16-what-do-lenders-10-k-filings-reveal-about-the-role-of-aiml-in-targeting-revolvers-for-limit-increases"&gt;Q16. What do lenders&amp;rsquo; 10-K filings reveal about the role of AI/ML in targeting revolvers for limit increases?&lt;/h3&gt;
&lt;p&gt;A: Banks that mention &amp;ldquo;artificial intelligence&amp;rdquo; or &amp;ldquo;machine learning&amp;rdquo; above the median number of times in their 2024 10-K filings support a higher share of revolving balances through credit limit increases, for all credit score groups. This difference is not driven by differences in credit limits at origination between higher-AI and lower-AI lenders, suggesting that AI/ML adoption affects the targeting of limit increases toward revolvers rather than the initial credit allocation.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Revolving utilization.&lt;/strong&gt; In this paper, revolving utilization is defined as the portion of overall credit card utilization attributable to balances that the borrower carries from one month to the next without full repayment, thereby accruing interest. It is measured as revolving balances divided by credit limit, averaged over the prior three months. This is distinct from transacting utilization (new purchases as a share of limit) and is the primary signal banks use — implicitly, via their algorithms — to select accounts for proactive limit increases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bank-initiated vs. consumer-initiated credit limit increase.&lt;/strong&gt; A bank-initiated limit increase is one in which the lender proactively raises a borrower&amp;rsquo;s credit limit without a request from the borrower. A consumer-initiated increase is one explicitly requested by the borrower. The Y-14M data uniquely identify the source of each change. The paper documents that approximately 75–80% of all limit increases are bank-initiated, and that bank-initiated increases are strongly correlated with revolving utilization whereas consumer-initiated increases are not.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Low-and-grow strategy.&lt;/strong&gt; The practice of originating higher-risk borrowers at low initial credit limits and then expanding those limits over time based on observed borrowing behavior. In the paper this is a documented empirical pattern, not an assumption: subprime accounts start at an average $700 limit at origination and reach nearly $5,000 by eight years, a 285% increase versus only 25% for superprime accounts over the same horizon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Temptation preferences (Gul–Pesendorfer).&lt;/strong&gt; A utility framework in which household welfare depends not only on actual consumption but also on the most tempting consumption alternative within the budget set. The disutility from temptation arises even when the household does not succumb — it reflects the psychological cost of self-restraint. In the paper, λ (set to 0.28) parameterizes the weight of this temptation cost relative to standard utility. Temptation preferences are time-consistent, which facilitates welfare analysis, and are preferred to hyperbolic discounting in this setting because they predict that individuals may pay to have tempting options removed even without acting on them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Revealed preference for targeting revolvers.&lt;/strong&gt; The paper&amp;rsquo;s characterization of banks&amp;rsquo; credit limit increase behavior as reflecting a systematic preference for giving increases to revolving borrowers, inferred from the empirical pattern in the Y-14M data (the inverted-U shape between revolving utilization and limit increase probability). Because banks&amp;rsquo; algorithms are proprietary and unobserved, the paper interprets the observed allocation of limit increases as a revealed preference, consistent with banks&amp;rsquo; profit motive since revolvers generate the majority of credit card interest income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption equivalent variation (CEV).&lt;/strong&gt; The welfare metric used throughout the paper&amp;rsquo;s counterfactual analysis. CEV is defined as the percentage change in consumption in every period and state that would make households indifferent between the baseline policy regime and the counterfactual policy. A positive CEV indicates that the counterfactual policy improves welfare; a negative CEV indicates harm. The paper considers two versions: one in which the social planner internalizes the psychological cost of temptation (consistent with tempted households&amp;rsquo; actual preferences), and one in which the planner ignores that cost (λ = 0 for the planner) but households still face temptation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Persistent revolving debt (UK regulatory definition).&lt;/strong&gt; In the UK Financial Conduct Authority&amp;rsquo;s framework, a borrower is considered in &amp;ldquo;persistent revolving debt&amp;rdquo; when the cumulative amount paid toward interest and fees exceeds the cumulative amount of principal repaid over a 12-month period. The UK rule prohibits lenders from increasing credit limits for borrowers meeting this definition. The paper models a stylized version: any account currently carrying a revolving balance is ineligible for a bank-initiated limit increase in the UK-style counterfactual.&lt;/p&gt;</description></item><item><title>Automation and Rent Dissipation</title><link>https://macropaperwarehouse.com/papers/automation-and-rent-dissipation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/automation-and-rent-dissipation/</guid><description>&lt;p&gt;Acemoglu and Restrepo examine the effects of automation in economies where labor market distortions cause some workers to earn rents—wages above their opportunity cost or outside option. The central question is how the interplay between automation and these distortions shapes wages, inequality, and productivity. The paper makes three contributions: a theoretical framework identifying a rent dissipation mechanism, reduced-form empirical evidence using US data from 1980 to 2016, and a general equilibrium quantification of automation&amp;rsquo;s aggregate effects.&lt;/p&gt;
&lt;p&gt;The theoretical framework extends the task model of Acemoglu and Restrepo (2022) to incorporate task-specific wage wedges. In this setup, a firm employing labor of type g in task x pays a wage equal to the base wage multiplied by an exogenous wedge capturing rents from efficiency wages, bargaining, licensing, regulations, or norms. Because these wedges artificially inflate labor costs in high-rent tasks, firms have a stronger incentive to automate precisely those tasks—automation saves more in labor costs where rents are highest. Proposition 3 establishes that endogenous adoption decisions are tilted toward high-rent tasks: the rent distribution in automated tasks first-order stochastically dominates the rent distribution across all tasks. This targeting generates the rent dissipation mechanism. The equilibrium is inefficient on both the intensive margin (too little employment in high-rent tasks) and the extensive margin (excessive automation of high-rent tasks that a social planner would prefer to keep labor-intensive).&lt;/p&gt;
&lt;p&gt;The rent dissipation mechanism has three consequences identified theoretically. First, it amplifies average wage losses for exposed groups beyond what displacement alone would produce, pushing displaced workers toward lower-paying jobs. Second, it compresses within-group wage dispersion by concentrating losses at higher percentiles of the within-group distribution, generating a U-shaped pattern of wage changes: workers at low percentiles earn no rents and experience only base-wage adjustments, while workers between the 70th and 95th percentiles face the steepest declines due to loss of high-rent jobs. Third, it is inefficient: because the tasks targeted by automation are not those where wages reflect scarcity or skill but rather distortionary rents, a planner would have preferred more labor allocated to these tasks, and rent dissipation offsets part or all of the cost-saving productivity gains from automation.&lt;/p&gt;
&lt;p&gt;The empirical analysis covers 500 detailed demographic groups defined by education (five levels), gender, five age groups, five race/ethnicity groups, and nativity. Task displacement is measured as a weighted sum of industry-level automation exposure using three proxies: adjusted industrial robot penetration, specialized software services, and dedicated machinery in value added. Workers in the middle and lower-middle of the wage distribution lost 15–20% of their tasks to automation between 1980 and 2016, while post-college workers saw few tasks automated.&lt;/p&gt;
&lt;p&gt;A 10 percentage point increase in task displacement is associated with a 24% decline in group-level relative wages (β = −2.36, s.e. = 0.13), falling to 19% after controlling for gender, education, sectoral demand, and rent shifters (β = −1.90, s.e. = 0.29). The U-shaped pattern in within-group wage changes is clearly visible: wages decline by 25–30% per 10 percentage point task displacement at the 70th–90th percentiles, compared to only 16% at the 5th–40th percentiles. Decomposing the average wage effect, the base-wage component is β = −1.53 (s.e. = 0.33) and the rent-dissipation component is β = −0.37 (s.e. = 0.11), implying a rent dissipation rate of approximately 37%. Across multiple proxies for rents—inter-industry/occupation wage differentials, wage losses after job displacement, and quit rates—the average estimated rent dissipation rate is approximately 35%. Rent dissipation accounts for one-fifth of the overall relative wage decline experienced by groups exposed to automation.&lt;/p&gt;
&lt;p&gt;In the general equilibrium quantification (with elasticity of substitution λ = 0.5, average cost savings π = 30%, and average rent in automated tasks of 35%), automation accounts for 52% of the rise in between-group wage inequality since 1980: 42 percentage points via baseline displacement effects on labor demand, and 10 percentage points via rent dissipation. Cost savings from automation increased TFP by approximately 3% between 1980 and 2016, but inefficient rent dissipation offsets 60–90% of these gains, leaving net TFP gains of only 0.3–1.3% and net aggregate consumption gains of only 0.45–1.95% over the 36-year period.&lt;/p&gt;
&lt;p&gt;Q: What is the rent dissipation mechanism, and why does it arise?
A: Rent dissipation arises because labor market wedges make high-rent tasks artificially costly to staff with workers, giving firms a stronger incentive to automate precisely those tasks. When automation displaces workers from high-rent jobs, workers lose the premium above their opportunity cost that those jobs paid, amplifying wage losses beyond what displacement alone would cause. The mechanism is endogenous: firms do not randomly automate tasks but disproportionately target tasks where rents are highest, since doing so saves the most in labor costs. Proposition 3 formalizes this as first-order stochastic dominance of the rent distribution in automated tasks over the rent distribution in all tasks.&lt;/p&gt;
&lt;p&gt;Q: Why is rent dissipation inefficient?
A: In a distorted economy, high-rent tasks already feature too little employment at the equilibrium—firms under-hire in these tasks because the wage wedge makes labor artificially expensive. A social planner would want to allocate more labor to these tasks, not less. When automation further removes labor from high-rent tasks, it moves the economy further from the efficient allocation, dissipating rents that reflect distortions rather than true scarcity. The TFP formula shows that this inefficient targeting offsets part or all of the cost-saving gains from automation, and can even reduce aggregate productivity if the cost savings are small relative to the rent losses.&lt;/p&gt;
&lt;p&gt;Q: What is the U-shaped pattern of within-group wage changes, and what does it indicate?
A: The U-shaped pattern means that wage declines due to automation are smallest at the bottom percentiles of a group&amp;rsquo;s within-group wage distribution, largest in the 70th–95th percentile range, and then smaller again at the very top. Workers at low percentiles earn no rents, so they experience only the base-wage adjustment from reduced labor demand. Workers in the middle-upper range of the distribution hold the high-rent jobs that are disproportionately automated, so they lose both the base-wage component and the rent component of their wages. This pattern is directly visible in US data 1980–2016, with declines of 25–30% per 10 percentage point task displacement at the 70th–90th percentiles versus 16% at the 5th–40th percentiles.&lt;/p&gt;
&lt;p&gt;Q: How is task displacement measured, and which groups are most exposed?
A: Task displacement is measured as a weighted sum of industry-level automation exposure, accounting for each demographic group&amp;rsquo;s specialization in routine tasks within industries. Three proxies are used: the adjusted penetration of industrial robots, the increase in specialized software services, and the increase in dedicated machinery in value added. Workers in the middle and lower-middle of the wage distribution—broadly corresponding to non-college workers—lost 15–20% of their tasks to automation between 1980 and 2016. Post-college degree workers saw few tasks automated.&lt;/p&gt;
&lt;p&gt;Q: How large is the rent dissipation rate, and how robust is this estimate?
A: The baseline estimate from the U-shaped within-group wage change decomposition implies a rent dissipation rate (μ_Ag/μ_g − 1) of approximately 37% (β = −0.37, s.e. = 0.11). Using inter-industry and occupation wage differentials as a proxy for rents, the estimate is 39% (β = −0.39, s.e. = 0.11). Using wage losses after job displacement, the estimate is 20% (β = −0.20, s.e. = 0.04). After purging compensating differentials from the wage differential proxy the estimate remains 37%; after purging from the displacement-loss proxy it falls to 19%. Quit-rate evidence is consistent with rent dissipation: automation shifts workers toward higher-quit-rate jobs, which are lower-rent jobs. The average across proxies is approximately 35%.&lt;/p&gt;
&lt;p&gt;Q: How much of between-group wage inequality since 1980 does automation explain, and what share is due to rent dissipation specifically?
A: Automation accounts for 52% of the rise in between-group wage inequality in the US since 1980. Of this 52 percentage points, 42 percentage points are attributable to the baseline displacement effect working through reduced labor demand for exposed groups. The remaining 10 percentage points are attributable to rent dissipation—automation pushing exposed groups away from high-rent tasks into lower-paying employment. Rent dissipation thus accounts for roughly one-fifth (10/52) of automation&amp;rsquo;s total contribution to between-group inequality.&lt;/p&gt;
&lt;p&gt;Q: How large are the productivity gains from automation, and how much does rent dissipation offset them?
A: Cost savings from automation increased TFP by approximately 3% between 1980 and 2016. However, inefficient rent dissipation offsets 60–90% of these gains, because automation disproportionately targets high-rent tasks rather than tasks where the efficiency case is strongest. The net TFP increase attributable to automation is only 0.3–1.3% over the 36-year period, and the corresponding net increase in aggregate consumption is only 0.45–1.95%.&lt;/p&gt;
&lt;p&gt;Q: How does automation affect within-group versus between-group inequality, and why is this notable?
A: Automation increases between-group inequality by reducing relative wages of exposed groups (largely non-college workers) relative to unexposed groups, accounting for 52% of the rise in between-group inequality since 1980. At the same time, automation reduces within-group wage dispersion for exposed groups by compressing wages at higher percentiles. This contrasts with the standard view that inequality is fractal—rising at all levels of aggregation due to skill-biased demand—and helps explain why within-group inequality has risen steadily for college workers since the 1980s while remaining flat and then declining for non-college workers since the 1990s.&lt;/p&gt;
&lt;p&gt;Q: What do the propagation matrix and rent-impact matrix represent in the general equilibrium analysis?
A: The propagation matrix encodes how task reallocation due to automation in one demographic group creates competition for marginal tasks across other groups, transmitting the wage effects of automation to groups not directly displaced. The rent-impact matrix encodes how this task reallocation changes the rent composition of employment across groups. Both matrices are estimated from US data on task shares and group-level wage elasticities and are used to translate partial-equilibrium estimates of task displacement and rent dissipation into general equilibrium effects on wages and productivity for all demographic groups simultaneously.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of inefficient rent dissipation?
A: Because rent dissipation is inefficient, the social value of automation is lower than what firms and consumers are willing to pay—firms capture all the labor cost savings but do not internalize the welfare cost of destroying high-rent jobs that the distorted equilibrium already under-supplies. Second-best interventions should address the underlying distortions generating rents rather than trying to slow automation directly. The paper suggests that strengthening labor market institutions supporting worker rents in non-automatable tasks could partially counteract the adverse distributional consequences of automation.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to Bound and Johnson (1992) and Borjas and Ramey (1995)?
A: Bound and Johnson (1992) decompose changes in the US wage structure between 1979 and 1988 into technology, supply, and rent components (modeled as exogenous industry wedges), finding that 10–20% of between-group wage changes reflect rent losses. Borjas and Ramey (1995) estimate that trade increased the college premium by 1.3–2.6 log points between 1976 and 1990, with 15–33% due to loss of rents from trade-exposed jobs. Both are comparable to this paper&amp;rsquo;s finding that rent dissipation accounts for one-fifth of the wage effect of automation, though Bound and Johnson&amp;rsquo;s estimates include all factors affecting rents while this paper isolates automation specifically.&lt;/p&gt;
&lt;p&gt;Worker rents: Wages above a worker&amp;rsquo;s opportunity cost or outside option, arising from efficiency wages, bargaining, licensing, regulations, or norms. Modeled as task-specific multiplicative wedges (μ_gx ≥ 1) that force firms to pay more than the base wage for labor in particular tasks. Explicitly excludes compensating differentials and skill premia.&lt;/p&gt;
&lt;p&gt;Rent dissipation: The loss of above-opportunity-cost wages experienced by workers displaced from high-rent tasks into lower-paying employment. Occurs because automation endogenously targets high-rent tasks where labor is most expensive, and pushes workers into tasks where rents are lower. Quantified as the ratio of average rents in automated tasks to average rents across all tasks, minus one (approximately 35% in US data 1980–2016).&lt;/p&gt;
&lt;p&gt;Task displacement: The share of tasks performed by a demographic group that are automated away, measured as a weighted sum of industry-level automation exposure accounting for the group&amp;rsquo;s specialization in routine tasks. Distinct from employment loss because it captures reallocation of tasks from labor to capital within the production function.&lt;/p&gt;
&lt;p&gt;U-shaped within-group wage change profile: The pattern whereby automation generates the largest wage declines at intermediate-to-upper percentiles (70th–95th) of an exposed group&amp;rsquo;s within-group wage distribution, with smaller declines at the bottom, because high-percentile workers disproportionately hold high-rent jobs targeted by automation. Predicted theoretically and confirmed empirically in US data 1980–2016.&lt;/p&gt;
&lt;p&gt;Propagation matrix: A matrix estimated from US data on task shares and group-level wage elasticities that encodes how automation of tasks performed by one demographic group creates competition for marginal tasks with other groups, transmitting wage effects across the demographic distribution in general equilibrium.&lt;/p&gt;
&lt;p&gt;Inefficient automation targeting: The mechanism by which labor market distortions cause firms to automate high-rent tasks that a social planner would prefer to keep labor-intensive, since the distorted equilibrium already features too little employment in those tasks. Results in rent dissipation offsetting 60–90% of automation&amp;rsquo;s direct TFP gains from cost savings.&lt;/p&gt;
&lt;p&gt;Rent-impact matrix: A matrix that encodes how task reallocation due to automation changes the rent composition of employment across demographic groups, used alongside the propagation matrix to compute general equilibrium effects of automation on wages and productivity accounting for distortions.&lt;/p&gt;</description></item><item><title>Bargaining and Inequality in the Labor Market</title><link>https://macropaperwarehouse.com/papers/bargaining-and-inequality-in-the-labor-market/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/bargaining-and-inequality-in-the-labor-market/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; How prevalent is individual wage bargaining in the labor market, what determines firms&amp;rsquo; bargaining strategies, how do bargaining encounters unfold for workers, and does heterogeneity in bargaining behavior translate into wage inequality—including the gender wage gap?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Setting.&lt;/strong&gt; The paper develops and validates novel linked survey data for Germany. A firm survey was fielded by the ifo Institute to senior HR professionals and managers in two waves (September 2021 and January 2022), yielding 772 complete responses across all major sectors and regions. These responses were linked—with consent obtained from 72% of firms—to German Social Security records (the Integrated Employment Biographies, IEB) covering 416,821 full-time employees at matched firms in 2020, and to Orbis balance sheet data for firm productivity proxies. A separate worker survey was fielded by the IAB to 135,000 full-time German workers, with 9,756 completing it; nearly 10,000 responses were used for analysis, with 7,079 workers employed at surveyed firms. The worker survey elicited detailed bargaining histories for workers who had received an outside offer in the prior six months, bargaining at the start of current employment (for workers with tenure of three years or less), and responses to a hypothetical salary expectation scenario.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Definition of Individual Bargaining.&lt;/strong&gt; The authors define a firm as having a &amp;ldquo;bargaining strategy&amp;rdquo; if it differentiates pay between workers in the same position it perceives to have similar productivity—encompassing both variation in initial offers (which may reflect firms using information on workers&amp;rsquo; salary expectations) and back-and-forth negotiation. Elicitation distinguishes four employee groups (recent labor market entrants, experienced non-managers, managers, and bottleneck-occupation workers) and two contexts (new external hires and incumbent workers who receive an outside offer).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prevalence of Bargaining.&lt;/strong&gt; Approximately 50% of surveyed firms are willing to differentiate base wages for recent labor market entrants, more than 80% for experienced non-managers and managers, and nearly all for workers in bottleneck occupations they are struggling to fill. For incumbent workers facing outside offers, 57% of firms would increase pay for recent entrants, and more than 80% for experienced incumbents, managers, and bottleneck workers. In total, 80% of workers in the sample are in positions where individual bargaining is possible.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Magnitude of Wage Differentiation.&lt;/strong&gt; For new external hires, the typical firm expects a gap between the highest and lowest offers of 3% for recent entrants, 5% for experienced non-managers, and 10% for managers (conditional on a gap: 6%, 10%, and 12% respectively). For incumbent workers responding to outside offers, the typical firm will adjust pay by 3% for recent entrants, 6% for experienced non-managers, and 10% for managers (conditional on responding: 6%, 7%, and 14% respectively). Forty-four percent of firms report that variation in initial offers is at least as important as back-and-forth negotiation in determining workers&amp;rsquo; final pay.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Predictors of Firm Bargaining Strategies.&lt;/strong&gt; Contrary to models predicting more productive firms are more likely to bargain (Doniger 2015; Postel-Vinay and Robin 2004; Flinn and Mullins 2021), firms that bargain are not more productive—as proxied by firm age, size, or assets per employee—nor do they pay higher mean wages. A variance decomposition shows that employee-group dummies alone explain 33% of variation in bargaining strategies for new hires, comparable to more than 500 firm dummies. Labor market factors—particularly whether a position is hard to fill—are systematically associated with bargaining willingness. Collective bargaining agreement (CBA) coverage and East German location are negatively correlated with bargaining flexibility.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How Bargaining Unfolds.&lt;/strong&gt; In 57% of worker-firm interactions, the worker provides salary expectations before the firm makes its initial offer; 29% of firms require this information. About one-third of applicants ask for more after the initial offer, requesting on average a 3% increase; conditional on asking, about half of firms raise the offer, but fewer than one-third match what was requested, with the typical worker improving the offer by 1.5%. The majority of outside offers are rejected: only 9% of workers who received an outside offer in the prior six months chose to move to a new firm. Of the 91% who remained at their incumbent firm, 13% successfully renegotiated their pay. Back-and-forth dynamics—where offers are accepted or rejected only after multiple rounds—are consistent with models of two-sided incomplete information.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Worker Heterogeneity and Wage Inequality.&lt;/strong&gt; Workers with better self-assessed outside options are 9 percentage points more likely to ask for an increase after the initial offer and 7 percentage points more likely to successfully negotiate a raise, relative to same-occupation coworkers with worse outside options. Women are 6 percentage points less likely to successfully negotiate their pay upward and show lower salary expectation provision rates, including in a hypothetical scenario in which pay range information is equalized. These gender differences in bargaining are not explained by women negotiating more over non-wage amenities; controlling for outside options and risk tolerance shrinks the female coefficient by at most 15%. Among surveyed workers, after controlling for occupation-establishment fixed effects, there is no gender wage gap at firms that do not bargain, but a 4–5 percentage point gender wage gap at firms that do bargain. Across specifications, firms that engage in individual bargaining have a 3 percentage point higher gender wage gap. A simple decomposition suggests that at surveyed firms, 44% of the residual gender pay gap can be attributed to bargaining. For workers at bargaining firms, a 10 percentage point higher pay premium at the prior firm is associated with 0.5 percent higher pay at the current firm, conditional on occupation-establishment fixed effects; this relationship is statistically insignificant for workers at non-bargaining firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; Results apply to full-time private-sector workers in Germany between ages 25 and 50, with the firm sample over-representing medium and large firms (median size 50–249 employees). CBA coverage in the sample (41%) reflects Germany&amp;rsquo;s institutional context where firms retain the right to pay above CBA floors. Results are robust to re-weighting to match the overall distribution of German firm size and sector.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-how-do-the-authors-define-individual-bargaining-and-why-is-this-definition-broader-than-standard-labor-economics-usage"&gt;Q1. How do the authors define &amp;ldquo;individual bargaining&amp;rdquo; and why is this definition broader than standard labor economics usage?&lt;/h3&gt;
&lt;p&gt;The authors define a firm as having a bargaining strategy if it differentiates pay between workers in the same position it perceives to have similar productivity, covering both tailoring of initial offers and back-and-forth negotiation. Standard labor economics definitions typically condition on wages being set ex post once outside options are revealed, and focus on back-and-forth negotiation alone. The authors&amp;rsquo; definition is most analogous to standard definitions of price discrimination. Empirically, the vast majority of firms that differentiate initial offers (93%) are also willing to engage in back-and-forth negotiation.&lt;/p&gt;
&lt;h3 id="q2-how-was-the-firm-survey-designed-to-elicit-bargaining-strategies-reliably-and-what-is-the-protocol-question"&gt;Q2. How was the firm survey designed to elicit bargaining strategies reliably, and what is the &amp;ldquo;protocol question&amp;rdquo;?&lt;/h3&gt;
&lt;p&gt;The protocol question asked: &amp;ldquo;How much more could a person maximally receive compared to the fixed compensation you would have offered based on the person&amp;rsquo;s qualification/fit for the position alone?&amp;rdquo; with options ranging from &amp;ldquo;0%/no adjustments possible&amp;rdquo; to &amp;ldquo;more than 40%.&amp;rdquo; Wording was developed through over 100 conversations with HR professionals; &amp;ldquo;qualifications and fit&amp;rdquo; was the phrase most closely aligned with HR professionals&amp;rsquo; concept of productivity. The survey was fielded by the ifo Institute—an organization with decades of experience surveying this population—with a 51% response rate, 83% completion rate, and median response time of 11 minutes.&lt;/p&gt;
&lt;h3 id="q3-what-validation-exercises-support-the-reliability-of-the-elicited-firm-bargaining-measures"&gt;Q3. What validation exercises support the reliability of the elicited firm bargaining measures?&lt;/h3&gt;
&lt;p&gt;Four exercises are reported. First, intra-respondent reliability: the cross-tabulations between the protocol and incidence questions show most mass on or below the diagonal (incidence-implied spread no greater than the protocol-implied flexibility). Second, inter-respondent reliability: among 37 firms with multiple respondents, there is significant overlap in independently provided answers. Third, external validity using publicly available data: for 90% of firms reporting no CBA, no CBA evidence is found; for 99% reporting no pay information in job ads, none is found in online postings; for 82% reporting no salary expectation elicitation, no evidence of it appears in online application forms. Fourth, the elicited firm strategies are highly correlated with the matching workers&amp;rsquo; survey responses—e.g., workers at firms stating they elicit salary expectations are significantly more likely to report having provided these expectations.&lt;/p&gt;
&lt;h3 id="q4-is-firm-productivity-associated-with-whether-a-firm-engages-in-individual-bargaining"&gt;Q4. Is firm productivity associated with whether a firm engages in individual bargaining?&lt;/h3&gt;
&lt;p&gt;No. Firms that bargain and those that do not are similar with respect to firm size, firm age, and total assets per employee, and they also do not differ significantly in their AKM wage premium. These findings are inconsistent with theoretical models predicting that more productive firms are more likely to set pay via bargaining (Doniger 2015; Postel-Vinay and Robin 2004; Flinn and Mullins 2021). The result holds for both binary and continuous measures of bargaining, and is not overturned by machine learning prediction attempts.&lt;/p&gt;
&lt;h3 id="q5-what-firm-characteristics-other-than-productivity-predict-bargaining-strategies"&gt;Q5. What firm characteristics other than productivity predict bargaining strategies?&lt;/h3&gt;
&lt;p&gt;CBA coverage is negatively correlated with wage flexibility—CBA-covered firms report less flexibility even for managers who are typically exempt from CBAs and for groups not covered by CBAs, suggesting institutional norms or culture matter. Firms headquartered in East Germany are less likely to bargain with workers in all groups. Publicly traded firms (stock-based corporations) are more likely to set wages flexibly. These correlations are consistent with the view that managerial style and firm culture (rather than productivity) shape wage-setting strategies.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-variance-decomposition-say-about-the-relative-importance-of-firm-versus-market-factors-in-predicting-bargaining-strategies"&gt;Q6. What does the variance decomposition say about the relative importance of firm versus market factors in predicting bargaining strategies?&lt;/h3&gt;
&lt;p&gt;Employee-group dummies alone explain 33% of the variation in bargaining strategies for new hires. After adjusting for the number of fixed effects used, four employee-group dummies explain as much variation as more than 500 firm dummies. Adding firm characteristics or coarse industry dummies does not significantly improve the adjusted R-squared relative to a model containing only group dummies. This supports models emphasizing market-level factors (worker replaceability, labor market tightness) over firm-level factors.&lt;/p&gt;
&lt;h3 id="q7-how-common-is-it-for-workers-to-provide-salary-expectations-before-receiving-an-initial-offer-and-what-do-firms-do-with-this-information"&gt;Q7. How common is it for workers to provide salary expectations before receiving an initial offer, and what do firms do with this information?&lt;/h3&gt;
&lt;p&gt;In 57% of worker-firm interactions, the worker provides salary expectations before the firm makes its initial offer. Twenty-nine percent of firms require this information; most ask for it. Forty-four percent of firms report that variation in initial offers is at least as important as subsequent back-and-forth negotiations in determining workers&amp;rsquo; final pay. HR professionals and prior research indicate firms interpret variation in stated expectations as reflecting outside options rather than productivity.&lt;/p&gt;
&lt;h3 id="q8-what-fraction-of-outside-offers-are-rejected-and-what-happens-when-workers-stay-at-the-incumbent-firm"&gt;Q8. What fraction of outside offers are rejected, and what happens when workers stay at the incumbent firm?&lt;/h3&gt;
&lt;p&gt;Only 9% of workers who received one or more outside offers in the prior six months chose to move to a new firm. Of the 91% who remained at the incumbent firm, 13% successfully renegotiated their pay at the incumbent. A follow-up survey fielded in spring 2024 corroborates this finding, showing approximately 80% of workers who received an outside offer remained at the incumbent firm; even recoding all job-to-job transitions as accepted offers implies no more than 26% of offers lead to a transition.&lt;/p&gt;
&lt;h3 id="q9-what-do-the-back-and-forth-dynamics-imply-for-appropriate-theoretical-models-of-wage-bargaining"&gt;Q9. What do the back-and-forth dynamics imply for appropriate theoretical models of wage bargaining?&lt;/h3&gt;
&lt;p&gt;That many offers are accepted or rejected only after multiple rounds of negotiation is difficult to rationalize with models assuming either firms or workers have perfect information, which typically predict immediate acceptance or rejection. The patterns are consistent with models of two-sided incomplete information (Perry 1986; Chatterjee and Samuelson 1983). Sixty-nine percent of HR professionals in the survey report that decision-makers at their firm only have market-level information on wages, not specific information on what competitors pay.&lt;/p&gt;
&lt;h3 id="q10-how-do-outside-options-predict-worker-bargaining-behavior-and-outcomes-controlling-for-occupation-establishment-fixed-effects"&gt;Q10. How do outside options predict worker bargaining behavior and outcomes, controlling for occupation-establishment fixed effects?&lt;/h3&gt;
&lt;p&gt;Workers who rated it &amp;ldquo;easy&amp;rdquo; or &amp;ldquo;very easy&amp;rdquo; to obtain a better outside offer are 9 percentage points more likely to ask for an increase after the initial offer and 7 percentage points more likely to successfully negotiate a raise relative to same-occupation-establishment coworkers who rated it &amp;ldquo;difficult&amp;rdquo; or &amp;ldquo;very difficult.&amp;rdquo; The same pattern persists during the employment spell: workers with better outside options are 9 percentage points more likely to initiate and 8 percentage points more likely to succeed in renegotiation. These workers are not more likely to receive raises without asking.&lt;/p&gt;
&lt;h3 id="q11-how-does-risk-tolerance-predict-bargaining-and-how-does-it-compare-to-outside-options"&gt;Q11. How does risk tolerance predict bargaining, and how does it compare to outside options?&lt;/h3&gt;
&lt;p&gt;Workers with greater risk tolerance (those rating themselves 7 or above on a 10-point scale) are more likely to engage in wage negotiations and more likely to succeed both at the start of and during employment spells. Gaps in successful negotiations are somewhat larger than gaps in attempted negotiations, suggesting risk-tolerant workers also negotiate more effectively. However, outside options explain more of the between-worker variation in bargaining behavior than risk tolerance does.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-gender-differences-in-bargaining-behavior-and-can-they-be-explained-by-differences-in-outside-options-or-risk-tolerance"&gt;Q12. What are the gender differences in bargaining behavior, and can they be explained by differences in outside options or risk tolerance?&lt;/h3&gt;
&lt;p&gt;Women are less likely to engage in back-and-forth negotiations and are 6 percentage points less likely to successfully negotiate pay upward during an employment spell. Women are also less likely to provide salary expectations and provide lower expectations as a fraction of their current salary in the hypothetical scenario, including when the salary range is provided—women are 6 percentage points less likely to provide expectations above the top of the stated range. Controlling for outside options and risk tolerance shrinks the female coefficient by at most 15%. There is no evidence that women substitute toward negotiating for non-wage amenities. The pattern is most consistent with women finding negotiation uncomfortable, not with a belief that it will not pay off or fear of backlash.&lt;/p&gt;
&lt;h3 id="q13-what-is-the-estimated-gender-wage-gap-attributable-to-individual-bargaining"&gt;Q13. What is the estimated gender wage gap attributable to individual bargaining?&lt;/h3&gt;
&lt;p&gt;Among surveyed workers, after controlling for occupation-establishment fixed effects, there is no gender wage gap at firms without individual bargaining (coefficient closes to zero), while a 4–5 percentage point gender wage gap persists at firms with individual bargaining. This difference is robust across measures of pay (total daily pay, base pay, pay conditioning on hours worked), alternative fixed effect specifications, and to including non-surveyed workers at surveyed firms. A simple decomposition suggests 44% of the residual gender pay gap at surveyed firms can be attributed to bargaining. Across the interaction specifications, bargaining firms have a 3 percentage point higher gender wage gap and—in one key specification—a 6 percentage point difference between the gender gaps at bargaining and non-bargaining firms.&lt;/p&gt;
&lt;h3 id="q14-how-does-a-workers-prior-firm-wage-premium-affect-current-wages-and-does-bargaining-status-matter"&gt;Q14. How does a worker&amp;rsquo;s prior firm wage premium affect current wages, and does bargaining status matter?&lt;/h3&gt;
&lt;p&gt;In a regression of log current wages on the AKM wage premium of the prior firm (conditional on occupation-establishment fixed effects), a 10 percentage point higher pay premium at the prior firm is associated with 0.5 percent higher pay at the new firm for workers at bargaining firms. For workers whose pay is not set via individual bargaining, the relationship between the prior firm&amp;rsquo;s pay premium and current pay is statistically insignificant. The result is consistent with the idea that during negotiations with a new firm, workers use their prior firm&amp;rsquo;s pay policy as an outside option.&lt;/p&gt;
&lt;h3 id="q15-how-do-akm-person-effects-relate-to-bargaining-behavior"&gt;Q15. How do AKM person effects relate to bargaining behavior?&lt;/h3&gt;
&lt;p&gt;Higher-person-effect individuals are more likely to have provided salary expectations when applying to their current firm and ask for a larger fraction of their current salary in the hypothetical scenario (conditional on their wage). These differences persist when controlling for occupation-establishment fixed effects and age and experience. Higher-person-effect workers are not more likely to receive raises without asking. These results are inconsistent with AKM person effects reflecting only productivity differences and instead suggest that fixed differences in individual bargaining behavior contribute to the variance in person effects—which Card, Heining, and Kline (2013) estimated explains a large share (40%) of the growth in German wage inequality.&lt;/p&gt;
&lt;h3 id="q16-are-the-bargaining-patterns-found-at-surveyed-firms-representative-of-bargaining-more-broadly"&gt;Q16. Are the bargaining patterns found at surveyed firms representative of bargaining more broadly?&lt;/h3&gt;
&lt;p&gt;Two robustness exercises support broader representativeness. First, similar bargaining dynamics are found when including a random sample of German workers employed at non-surveyed firms. Second, re-weighting the sample to match the overall distribution of firm size and sector in Germany yields similar results. Because medium and large firms are over-represented in the firm sample, and because small firms hire infrequently and are less likely to have formal bargaining strategies, the true prevalence of individual bargaining among all German firms may be somewhat lower.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Individual Bargaining Strategy (firm-level).&lt;/strong&gt; A firm has an individual bargaining strategy if it differentiates pay between workers in the same position that it perceives to have similar productivity. This definition encompasses both tailoring of initial offers (based on, e.g., workers&amp;rsquo; stated salary expectations) and back-and-forth negotiation. It is analogous to price discrimination rather than to the standard labor economics distinction between wage posting and Nash bargaining.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Protocol Question.&lt;/strong&gt; The main survey measure of firm bargaining strategies: firms are asked the maximum percentage by which pay could be increased for a new hire above the fixed compensation the firm would have offered based on qualifications and fit alone, with response bins from &amp;ldquo;0%/no adjustments&amp;rdquo; to &amp;ldquo;more than 40%.&amp;rdquo; A zero response is used to classify a firm as not bargaining.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Incidence Question.&lt;/strong&gt; A supplementary survey measure eliciting the expected spread (between highest and lowest offers) that the firm would make to ten candidates with identical qualifications and fit but differing stated salary expectations and competing offers. Used to validate the protocol question and to quantify the importance of initial-offer differentiation relative to back-and-forth negotiation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bottleneck Occupation.&lt;/strong&gt; A firm-defined category of workers in positions that are particularly difficult to fill, drawing on an official German Federal Employment Agency designation. In the paper, bargaining willingness is systematically higher for workers in these positions than for other workers at the same firm, providing evidence that labor market tightness drives bargaining strategies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Outside Offer Renegotiation.&lt;/strong&gt; Wage renegotiation at the incumbent firm triggered by a worker receiving an outside offer, without a change in job tasks. The paper documents this is empirically more common than actual job-to-job transitions: of workers receiving outside offers, 91% remain at the incumbent firm, and 13% of those who remain successfully renegotiate their pay.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AKM Person Effect.&lt;/strong&gt; A worker fixed effect estimated from a two-way fixed effects regression of log wages on worker and firm fixed effects (following Abowd, Kramarz, and Margolis 1999). In this paper, AKM person effects are taken from Bellmann et al. (2020), estimated over 2010–2017 German population data. The paper provides evidence that these effects capture, in part, fixed differences in individual bargaining behavior rather than solely differences in productivity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AKM Firm Effect (Wage Premium).&lt;/strong&gt; The firm fixed effect from the same two-way fixed effects regression, representing the pay premium a firm pays relative to what would be expected given its workforce composition. The paper uses the prior firm&amp;rsquo;s AKM effect as a measure of a worker&amp;rsquo;s outside option quality when testing whether prior-firm pay policy influences current pay under individual bargaining.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Salary Expectations (Gehaltsvorstellungen).&lt;/strong&gt; The wage figure a worker provides to a prospective employer, typically before the firm&amp;rsquo;s initial offer. Legally, German firms (like most US states) cannot ask for salary history but can ask for salary expectations. In the paper, 57% of worker-firm interactions begin with the worker providing expectations; firms report using these to tailor initial offers, interpreting variation in stated expectations as reflecting outside options rather than productivity.&lt;/p&gt;</description></item><item><title>Biased expectations and labor market outcomes: Evidence from German survey data and implications for the East–West wage gap</title><link>https://macropaperwarehouse.com/papers/biased-expectations-and-labor-market-outcomes-evidence-from-german-survey-data-and-implications-for-the-eastwest-wage-gap/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/biased-expectations-and-labor-market-outcomes-evidence-from-german-survey-data-and-implications-for-the-eastwest-wage-gap/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; The paper asks two questions: (1) How do workers&amp;rsquo; biased expectations about job finding and job separation shape the labor market equilibrium and wages? (2) Are differences in expectation biases across workers a quantitatively important driver of wage differentials, specifically the East–West German wage gap?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The empirical analysis uses the German Socio-Economic Panel (SOEP), a nationally representative longitudinal survey of approximately 30,000 participants per wave. The working-age sample (ages 25–65) covers nine biennial survey waves from 1999 to 2015, yielding 67,772 observations for job separation expectations and 6,423 for job finding expectations. Perceived transition probabilities are reported on a 0–100 scale in steps of 10 percentage points. Actual (statistical) transition probabilities are constructed by estimating probit models that predict realized transitions within 24 months using a rich set of individual, job, and employer characteristics, and are rounded to the nearest decile for consistency with the survey scale.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main empirical findings.&lt;/strong&gt; Employed workers in Germany overestimate their job separation probability by 6.4 percentage points on average (perceived: 19.8%; actual: 13.3%), a pessimistic bias significant at the 1% level. Unemployed workers overestimate their job finding probability by 8.2 percentage points on average (perceived: 57.0%; actual: 48.8%), an optimistic bias also significant at the 1% level. The East–West divergence is striking. East German workers exhibit a pessimistic job separation bias of 12.1 percentage points, compared to only 4.7 percentage points in the West, despite broadly similar actual separation rates (15.1% vs. 12.8%). For job finding, West Germans overestimate their probability by 12.9 percentage points, while East Germans overestimate by only 2.0 percentage points — meaning East Germans are also substantially less optimistic about re-employment. These East–West differences survive controls for compositional differences and alternative definitions of job separation (dismissals only; selected reasons; spell-based) and job finding (including those out of the labor force). The biases are stable over the 1999–2015 sample period with no discernible trend. A cohort analysis shows that the excess pessimism in East Germany is concentrated among cohorts who were already in the labor market at the time of German reunification (born in the 1950s and 1960s), consistent with persistent effects of the communist GDR experience. Individuals do not systematically learn over time: mean changes in individual-level absolute deviations between consecutive waves are close to zero. Individual deviations between perceived and actual rates have statistically significant but quantitatively negligible predictive power for subsequent transitions (a 1 pp higher perceived job separation is associated with only a 0.001 pp higher realized separation rate), ruling out private information as a first-order explanation for the biases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The authors extend the Diamond–Mortensen–Pissarides (DMP) frictional labor market framework by (i) allowing workers to hold biased perceived transition rates (λw for job finding, σw for job separation) while firms have rational expectations, and (ii) introducing wage contracts of explicit length T periods after which parties re-bargain. Common knowledge of each party&amp;rsquo;s perceived values is assumed, and generalized Nash bargaining is applied. The contract length T is a key parameter: there exists a critical threshold T* such that a pessimistic job separation bias raises the equilibrium wage for T &amp;lt; T* (the continuation-value effect dominates) and lowers it for T ≥ T* (the within-contract discounting effect dominates). An optimistic job finding bias unambiguously raises the equilibrium wage by inflating the perceived value of unemployment and hence the reservation wage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative results.&lt;/strong&gt; The model is calibrated to East Germany. The job separation bias (∆σ = 0.0194) and job finding bias (∆λ = 0.0044) are set to SOEP-based estimates. The critical threshold implied by calibrated parameter values is T* = 10 quarters. The baseline contract length, constructed from the share of permanent (88%) and temporary (12%) contracts in SOEP and average remaining tenure until retirement, is T = 67 quarters (a lower bound). This exceeds T*, so the pessimistic separation bias depresses wages in the baseline. A counterfactual experiment assigns West German bias levels to East German workers, while holding all other parameters fixed. For the preferred calibration range (γ ∈ {0.35, 0.50}, T ∈ {67, 106, 159}), East German wages rise by 1.07 to 2.36 percent. This corresponds to a reduction in the conditional East–West German wage gap (23 percent) of 4.6 to 10.6 percent, and a reduction in the unconditional gap (30 percent) of 3.6 to 7.9 percent. Although wages rise, equilibrium unemployment increases by 0.70 to 1.01 percentage points, widening the already large East–West unemployment gap (approximately 7 percentage points). Net of the unemployment effect, expected lifetime income (computed at actual, unbiased transition rates) rises by 0.7 to 1.88 percent for East German workers under West German biases, implying an unambiguous welfare gain. Under a biennial calibration (robustness), wages increase by up to 3.3 percent and expected lifetime income rises by up to 2.23 percent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; Results apply to a stationary environment (no aggregate fluctuations). Firms are assumed to have rational expectations; an extension shows results hold provided firm bias is smaller than worker bias. Workers are assumed homogeneous in their bias levels; learning is abstracted from. The quantitative magnitudes are sensitive to the workers&amp;rsquo; bargaining power γ and the contract length T, both of which are subject to uncertainty in calibration.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-how-are-actual-statistical-transition-probabilities-constructed-and-why-are-probit-predicted-probabilities-preferred-over-realized-sample-means"&gt;Q1. How are actual (statistical) transition probabilities constructed, and why are probit-predicted probabilities preferred over realized sample means?&lt;/h3&gt;
&lt;p&gt;A: Realized transition rates in the sample mix transitions for various idiosyncratic reasons that vary substantially across population groups, so raw sample means do not reflect the probability a given individual faces at interview time. The authors estimate probit models separately for job separation (employed sample) and job finding (unemployed sample), including a rich set of covariates — age, gender, education, tenure, firm size, unemployment experience, industry, survey year, and East Germany indicator, among others — and predict individual-level probabilities at the time of the interview. For consistency with the survey&amp;rsquo;s discrete response format, probit-predicted probabilities are rounded to the nearest decile (0%, 10%, &amp;hellip;, 100%). The bias is computed as the individual-level difference between perceived and probit-predicted actual probabilities, averaged over the sample.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-magnitude-and-direction-of-the-aggregate-expectation-biases-in-germany"&gt;Q2. What is the magnitude and direction of the aggregate expectation biases in Germany?&lt;/h3&gt;
&lt;p&gt;A: Employed workers overestimate job separation by 6.4 percentage points on average (perceived 19.8% vs. actual 13.3%), a pessimistic bias significant at the 1% level. Unemployed workers overestimate job finding by 8.2 percentage points (perceived 57.0% vs. actual 48.8%), an optimistic bias also significant at the 1% level. Both directions are statistically robust across alternative definitions of separation and finding, as well as to trimming extreme responses (0% and 100% answers) and adjusting for directional rounding.&lt;/p&gt;
&lt;h3 id="q3-how-large-are-the-eastwest-differences-in-expectation-biases-and-do-they-survive-controls-for-compositional-differences"&gt;Q3. How large are the East–West differences in expectation biases, and do they survive controls for compositional differences?&lt;/h3&gt;
&lt;p&gt;A: East German workers exhibit a pessimistic job separation bias of 12.1 percentage points, more than 2.5 times the West German level of 4.7 percentage points, despite actual separation rates being broadly comparable (15.1% vs. 12.8%). For job finding, West Germans are optimistic by 12.9 percentage points while East Germans are optimistic by only 2.0 percentage points, a difference of 10.9 percentage points. The paper states these differences persist after accounting for compositional differences between regions, and are robust across all alternative definitions of job separation (Dismissals, Selected, Spell) and job finding (out of U or O). The table of robustness results (Table 2) confirms that in all specifications, the pessimistic separation bias is substantially larger in the East and the optimistic finding bias is substantially smaller.&lt;/p&gt;
&lt;h3 id="q4-what-cohort-analysis-is-conducted-to-explore-the-origins-of-greater-east-german-pessimism"&gt;Q4. What cohort analysis is conducted to explore the origins of greater East German pessimism?&lt;/h3&gt;
&lt;p&gt;A: The authors conduct a regression of the individual-level bias on birth-cohort indicators, controlling for age, demographic, and economic characteristics. They find that the pessimistic job separation bias is most pronounced among cohorts born in the 1950s and 1960s — those who experienced adult working life in the communist GDR and lived through reunification — and is smaller for cohorts born before 1950 and substantially smaller for cohorts born after 1970. For job finding, the optimistic bias is comparably low among cohorts born in the 1960s and earlier, but rises significantly for later-born East German cohorts. This cohort pattern is consistent with a long-lasting &amp;ldquo;experience effect&amp;rdquo; of communist institutions and the reunification shock on beliefs, analogous to findings in the broader literature on the persistent effects of communism.&lt;/p&gt;
&lt;h3 id="q5-is-there-evidence-that-individuals-update-their-biased-expectations-over-time"&gt;Q5. Is there evidence that individuals update their biased expectations over time?&lt;/h3&gt;
&lt;p&gt;A: To assess learning, the authors use the panel dimension and compute for each individual in two consecutive survey waves the absolute value of the deviation between perceived and actual transition probabilities, then examine the change in this absolute deviation between waves. The histograms of individual-level changes show substantial dispersion but means close to zero in all four sub-groups (East/West, job separation/finding), indicating no systematic convergence of beliefs toward actual rates. Biases are also stable in the time-series dimension, with perceived and actual rates moving largely in parallel across survey waves from 1999 to 2015, leaving the aggregate bias level roughly constant.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-model-rule-out-private-information-as-an-alternative-explanation-for-the-biases"&gt;Q6. How does the model rule out private information as an alternative explanation for the biases?&lt;/h3&gt;
&lt;p&gt;A: If biases reflected private information about idiosyncratic risk not captured by observable characteristics, individual-level deviations between perceived and actual rates should predict subsequent realized transitions. The authors add the individual-level deviation as an additional regressor in the probit transition models. The estimated coefficients are statistically significant and positive, but quantitatively negligible: a 1 percentage point higher expected job separation probability is associated with only a 0.001 percentage point higher realized separation probability, and a 1 percentage point higher expected job finding probability with a 0.002 percentage point higher realized finding probability. These magnitudes are too small to materially alter the interpretation of the biases as reflecting systematic expectation errors rather than private information.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-role-of-contract-length-t-in-the-model-and-what-is-the-critical-threshold-t"&gt;Q7. What is the role of contract length T in the model, and what is the critical threshold T*?&lt;/h3&gt;
&lt;p&gt;A: The wage contract length T determines which of two opposing effects of pessimistic job separation expectations dominates in bargaining. The first (negative wage) effect: a pessimistic worker discounts future wages within the current contract more heavily than the firm does, so the worker values the contract less and accepts a lower wage. The second (positive wage) effect: a pessimistic worker also discounts the continuation value of future contracts more heavily, making it less attractive to remain in the match, so the firm must offer a higher wage to retain the worker. For short contract lengths (T &amp;lt; T*), the second (positive) effect dominates, so the pessimistic bias raises wages. For long contracts (T ≥ T*), the first (negative) effect dominates, so the pessimistic bias depresses wages. The critical threshold T* is the smallest positive integer such that T*/λw(θ) &amp;lt; β times a weighted sum involving σw and T*. Using calibrated parameter values for East Germany, T* = 10 quarters (2.5 years). The baseline contract length is T = 67 quarters (approximately 16.8 years), well above T*, placing the economy in the regime where pessimism depresses wages.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-optimistic-job-finding-bias-affect-equilibrium-wages-and-unemployment"&gt;Q8. How does the optimistic job finding bias affect equilibrium wages and unemployment?&lt;/h3&gt;
&lt;p&gt;A: An optimistic job finding bias (λw &amp;gt; p(θ)) raises the perceived value of unemployment U because workers expect to escape unemployment sooner. A higher value of unemployment raises the worker&amp;rsquo;s outside option in bargaining, increases the reservation wage, and thereby pushes up the bargained wage. In general equilibrium, the job creation condition (which is unaffected by worker expectations) is unchanged, so the upward rotation of the wage curve reduces labor market tightness θ, raises equilibrium unemployment, and extends average unemployment duration. This comparative static holds unambiguously for any contract length T.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-quantitative-results-of-the-counterfactual-experiment-assigning-west-german-biases-to-east-german-workers"&gt;Q9. What are the quantitative results of the counterfactual experiment assigning West German biases to East German workers?&lt;/h3&gt;
&lt;p&gt;A: The counterfactual assigns West German bias levels (smaller pessimistic separation bias, larger optimistic finding bias) to East German workers while holding all other parameters at East German calibrated values. For the preferred calibration with γ ∈ {0.35, 0.50} and T ∈ {67, 106, 159}, wages in East Germany rise by 1.07 to 2.36 percent. This implies a reduction in the conditional East–West wage gap (23 percent) of 4.6 to 10.6 percent and a reduction in the unconditional gap (30 percent) of 3.6 to 7.9 percent. Equilibrium unemployment in East Germany rises by 0.70 to 1.01 percentage points as a side effect. Net of the unemployment effect, ex-ante unbiased expected lifetime income rises by 0.7 to 1.88 percent, confirming a positive welfare effect of reducing East German pessimism to West German levels. Under the biennial calibration robustness check, wage increases reach up to 3.3 percent, the conditional wage gap narrows by up to 11 percent, and lifetime income rises by up to 2.23 percent.&lt;/p&gt;
&lt;h3 id="q10-how-is-the-bargaining-power-parameter-γ-calibrated-and-why-does-it-matter-for-the-results"&gt;Q10. How is the bargaining power parameter γ calibrated and why does it matter for the results?&lt;/h3&gt;
&lt;p&gt;A: The paper considers a range γ ∈ {0.35, 0.50, 0.65}, rather than a single calibrated value, because γ plays a crucial role in the sensitivity of wages to expectation biases. Lower bargaining power reduces the equilibrium wage directly; however, because lower wages spur job creation, the model requires a higher vacancy cost κ to match the empirical job finding rate, which in turn increases the elasticity of wages with respect to the bias (see the wage equation, which shows that the bias effect scales with κθ/p(θ)). The paper argues that γ = 0.65 is inconsistent with the empirical wage–bias relationship estimated in SOEP data (which is negative and about twice as negative in East Germany as in the West), while γ ∈ {0.35, 0.50} is consistent. Lower bargaining power is also argued to be realistic for East Germany given weaker union representation there relative to the West.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-empirical-relationship-between-the-job-separation-bias-and-wages-serve-as-a-model-validation-target"&gt;Q11. How does the empirical relationship between the job separation bias and wages serve as a model validation target?&lt;/h3&gt;
&lt;p&gt;A: Using SOEP data, the authors regress log hourly wages on the individual-level difference between perceived and actual job separation rates, controlling for individual fixed effects and other covariates, and allow the slope to differ between East and West Germany. They find a statistically significant and negative relationship in both regions, with the effect approximately twice as large in East Germany as in the West. The estimate implies that if East German workers&amp;rsquo; job separation pessimism were reduced to West German levels, hourly wages in the East would be about 1 percent higher. This empirical gradient is used as an external validation check — not a calibration target — to assess which combinations of (γ, T) in the model are quantitatively plausible.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-model-predict-about-the-general-equilibrium-effects-on-unemployment-from-reducing-east-german-pessimism"&gt;Q12. What does the model predict about the general equilibrium effects on unemployment from reducing East German pessimism?&lt;/h3&gt;
&lt;p&gt;A: Reducing East German pessimism — both the pessimistic separation bias and the low optimistic finding bias — shifts the wage curve upward in equilibrium. Because the job creation condition is unaffected by worker beliefs (firms have rational expectations), higher wages reduce the firm&amp;rsquo;s incentive to post vacancies, lowering labor market tightness θ. This leads to higher equilibrium unemployment and longer average unemployment duration. The counterfactual with West German biases implies that East German unemployment would rise by 0.70 to 1.01 percentage points, further widening the approximately 7 percentage point East–West unemployment gap. The authors note this is a welfare-relevant trade-off, but show that the wage gain dominates the unemployment cost in terms of expected lifetime income.&lt;/p&gt;
&lt;h3 id="q13-what-robustness-checks-are-performed-on-the-quantitative-results"&gt;Q13. What robustness checks are performed on the quantitative results?&lt;/h3&gt;
&lt;p&gt;A: The paper considers (i) a narrower definition of job separation (dismissals only) to match the most likely interpretation of the survey question; (ii) targeting the officially reported East German unemployment rate (14.5% average from the Federal Employment Agency) rather than the SOEP-implied rate of 8.6% as a calibration target; (iii) a biennial calibration frequency instead of quarterly. The main results — wage increases and narrowing of the wage gap — are quantitatively similar across these alternatives, with one exception: the biennial calibration yields substantially larger wage increases (up to 3.3%), a larger reduction in the conditional wage gap (up to 11%), and larger lifetime income gains (up to 2.23%).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Expectation bias (job separation / job finding).&lt;/strong&gt; In this paper, a bias in expectations is defined as a systematic average difference between an individual&amp;rsquo;s perceived transition probability and the actual (statistically predicted) transition probability for their demographic and job group. A pessimistic job separation bias means workers overestimate the probability of losing their job (σw &amp;gt; σ); an optimistic job finding bias means unemployed workers overestimate the probability of re-employment (λw &amp;gt; p(θ)). Biases are not attributed to private information but to systematic expectation errors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Actual (statistical) transition probability.&lt;/strong&gt; The paper defines actual transition probabilities not as raw sample transition rates but as individual-level predicted probabilities from probit models estimated on realized transitions within 24 months, conditional on a comprehensive set of individual, job, and employer characteristics observed at interview time. These are rounded to the nearest decile for comparability with the survey&amp;rsquo;s discrete response format.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wage contract length (T).&lt;/strong&gt; The contract length T is the number of periods for which a bargained wage is fixed before the match parties re-bargain. A job match consists of a sequence of consecutive wage contracts of length T. The paper departs from the standard DMP assumption of period-by-period bargaining (T = 1) and shows that T is central to how job separation expectations feed into the bargained wage. A permanent job approximates T → ∞.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;Critical contract length (T&lt;/em&gt;).&lt;/em&gt;* A theoretically derived threshold: the pessimistic job separation bias raises equilibrium wages for contract lengths T &amp;lt; T* and depresses wages for T ≥ T*. Specifically, T* is the smallest positive integer such that T*/λw(θ) &amp;lt; β times a weighted sum involving β, σw, and T*. In the East German calibration, T* = 10 quarters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Generalized Nash bargaining with common knowledge / agree to disagree.&lt;/strong&gt; The model assumes that both the worker and the firm know each other&amp;rsquo;s perceived values of the job match and outside options and accept them as the basis for bargaining, even though they differ. Workers use their biased perceived transition rates to value employment and unemployment; firms use actual rates. There is no private information. The paper refers to this as workers and firms &amp;ldquo;agreeing to disagree.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ex-ante unbiased expected lifetime income (EI_{W,U}).&lt;/strong&gt; A welfare measure defined as the present discounted value of income for an individual entering the economy, computed at actual (unbiased) job separation and job finding probabilities rather than at workers&amp;rsquo; perceived (biased) rates. This measure captures the net welfare effect of changing expectation biases because it correctly accounts for actual employment transitions, even though the behavioral responses in equilibrium are driven by biased perceptions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Effective discount factor (β(1 − σw)).&lt;/strong&gt; When a worker holds pessimistic job separation expectations, future payoffs within the current contract are discounted not at the pure time discount factor β but at β(1 − σw), which is smaller when σw is larger. A more pessimistic worker therefore effectively discounts future wage payments more steeply, and this differential discounting relative to the firm (which uses β(1 − σ)) is the key mechanism generating the contract-length dependence of the wage effect.&lt;/p&gt;</description></item><item><title>Bridges</title><link>https://macropaperwarehouse.com/papers/bridges/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/bridges/</guid><description>&lt;p&gt;This paper measures the causal effects of land transport infrastructure on economic activity, exploiting quasi-experimental variation in bridge construction over the Mississippi and Ohio Rivers in the United States. The central empirical puzzle motivating the study is a hump-shaped relationship between per capita income and distance to major land transport routes in contemporary U.S. data: income peaks around 5 km from a transport route, with an elasticity of 0.072 closer than 4.1 km and -0.096 at greater distances, so that 85% of Americans live where local income increases with distance to transport routes rather than decreasing. The question is whether this pattern reflects causal effects of infrastructure, selection, or sorting.&lt;/p&gt;
&lt;p&gt;The paper develops two complementary identification strategies. The first exploits tributary confluences — where smaller rivers join larger rivers, sharply raising downstream flow rates and bridge construction costs — to generate quasi-random variation in bridge location. Because bridge construction costs increase convexly with river flow (maximum bending moment scales with span length squared), bridges are disproportionately built just upstream of confluences. The median upstream census tract lies 0.7 km from a bridge versus 2.3 km for the median downstream tract, making upstream tracts on average 60% closer to bridges and 27% closer to the nearest major land transport route. This asymmetry dates to at least 1880 and persists to 2010. Despite this persistent connectivity advantage, by 2010 upstream tracts have 13% lower per capita incomes and 63% higher population densities than downstream neighbours. The implied elasticity of per capita income with respect to distance to land transport, scaling the income effect by the distance-to-transport effect, is approximately 0.44. Income density (income per unit area) is higher upstream, though the difference is not statistically significant. Historical placebo tests using pre-bridge-construction data show no asymmetry in land values or population upstream versus downstream, supporting the identification assumption.&lt;/p&gt;
&lt;p&gt;The second strategy exploits variation in the timing of bridge construction. Because major bridge projects involve decades of planning, financing, design, and construction — the Wheeling Suspension Bridge was chartered in 1816 but opened in 1849 — the precise opening date is argued to be exogenous to short-run deviations from local growth trends. Using a county-level panel from 1860 to 2010 (432 counties, 14–19 states), the paper estimates event-study regressions around the first time a county experiences a 50% reduction in distance to a bridge. After such a reduction, farm land values (the best available consistent proxy for total economic activity in historical data) rise immediately and cumulatively by approximately 9% over 30 years. Population rises by approximately 5% over the same period. The proportionally larger rise in land values than population implies higher per capita economic activity in better-connected counties after 30 years.&lt;/p&gt;
&lt;p&gt;These two sets of results are reconciled through a narrative account of development. Better bridge access drives industrialization — manufacturing employment shares rise in counties experiencing improved connectivity — and urbanization. Cities form around historical transport routes and expand. Richer households then sort away from historical city centres into lower-density suburban areas, while lower-income households remain near or selectively migrate to the historical transport corridors. This within-city sorting produces the observed cross-sectional gradient: areas nearest transport routes end up with higher population density but lower per capita incomes. The negative local income effect of proximity to transport routes is larger in more urbanized areas and areas with higher income inequality, and is concentrated among non-white and low-education populations.&lt;/p&gt;
&lt;p&gt;The paper also contributes a new dataset covering every road and rail bridge (237 total) ever constructed over the Mississippi and Ohio Rivers from 1849 to 2010, assembled from the National Bridge Inventory and extensively cross-checked with satellite imagery and historical sources.&lt;/p&gt;
&lt;p&gt;Q: What is the motivating empirical puzzle about transport infrastructure and income?&lt;/p&gt;
&lt;p&gt;A: In contemporary U.S. census data, per capita income does not monotonically increase with proximity to land transport routes. Instead, the relationship is hump-shaped: income peaks around 5 km from a major transport route, with a positive elasticity of 0.072 within 4.1 km and a negative elasticity of -0.096 beyond that distance. Population density, by contrast, falls monotonically with distance to transport routes. As a result, 85% of Americans live in places where local mean income increases with distance to transport infrastructure rather than decreasing.&lt;/p&gt;
&lt;p&gt;Q: How does the tributary confluence identification strategy work?&lt;/p&gt;
&lt;p&gt;A: Tributary confluences — where smaller rivers join the main river — cause sharp, localized increases in river flow rates and thus in bridge construction costs, because cost scales convexly with required span length. This makes bridges systematically more likely to be built just upstream of confluences than just downstream. The strategy compares census tracts located upstream versus downstream of the 27 major tributary confluences identified on the Mississippi and Ohio Rivers, controlling for nearest-tributary fixed effects and distance to the confluence.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the connectivity difference between upstream and downstream census tracts?&lt;/p&gt;
&lt;p&gt;A: Upstream census tracts are approximately 60% closer to a bridge than downstream tracts (coefficient of 0.91 in log distance to bridge, p &amp;lt; 0.01), and consequently 27% closer to the nearest major land transport route (coefficient of 0.32, p &amp;lt; 0.10). This asymmetry is established by 1880 and persists through 2010. The advantage arises approximately equally from proximity to railroads and primary roads.&lt;/p&gt;
&lt;p&gt;Q: What are the causal effects of this connectivity advantage on per capita income and population density?&lt;/p&gt;
&lt;p&gt;A: Despite being better connected, upstream census tracts have 13% lower per capita incomes (coefficient 0.14 on the downstream indicator in log per capita income, p &amp;lt; 0.05) and 63% higher population densities (coefficient -0.49 on the downstream indicator in log population density, p &amp;lt; 0.05) in 2010. Income density is higher upstream, but the difference is not statistically distinguishable from zero. Scaling the income effect by the effect on distance to land transport implies an elasticity of approximately 0.44.&lt;/p&gt;
&lt;p&gt;Q: What pre-bridge-era placebo tests support the identifying assumption for the tributary confluence strategy?&lt;/p&gt;
&lt;p&gt;A: Matching modern census tracts to county-level historical data from 1840 and 1850 (before substantive bridge construction began), the paper finds no statistically significant asymmetry in land values or population density upstream versus downstream of tributary confluences. Asymmetric patterns emerge only after bridge construction begins. Ferry crossing locations, traced through place names in the USGS Geographic Names database, also appear equally frequently upstream and downstream, suggesting ferries did not differentially locate upstream.&lt;/p&gt;
&lt;p&gt;Q: How does the timing-based identification strategy work, and what is its key assumption?&lt;/p&gt;
&lt;p&gt;A: The strategy uses a county-level panel from 1860 to 2010 and estimates event-study regressions around the first time a county experiences a 50% reduction in distance to a bridge. County fixed effects and county-specific quadratic time trends absorb all fixed differences across counties and average changes in trends. The key assumption is that the exact opening date of a bridge is exogenous to short-run deviations from local long-run growth trends — supported by the argument that major bridges involve decades-long planning processes that evolve independently of local economic fluctuations. Pre-trend tests show no significant differences in outcomes before the event.&lt;/p&gt;
&lt;p&gt;Q: What are the quantitative effects of a major improvement in bridge access on land values and population?&lt;/p&gt;
&lt;p&gt;A: After a county first experiences a 50% reduction in distance to a bridge, farm land values rise immediately and cumulatively by approximately 9% (cumulative effect on log land values of about 0.09) over 30 years, relative to counties with no such change. Population rises by approximately 5% (cumulative log effect of about 0.05) over the same period. The proportionally larger effect on land values than on population implies that per capita economic activity is higher in better-connected counties 30 years after the event. The divergence between land value and population effects grows over time, suggesting productivity advantages accumulate.&lt;/p&gt;
&lt;p&gt;Q: Why does the paper use farm land values rather than other income measures in the historical panel?&lt;/p&gt;
&lt;p&gt;A: Farm land values — the total value of farm land and buildings — are the best consistently measured proxy for total economic activity available throughout the 1860–2010 census panel. The paper notes explicitly that as the economy industrializes and urbanizes, farm land values increasingly miss urban land values, implying that the estimated effects on farm land values are likely lower bounds on the true effects on total economic activity.&lt;/p&gt;
&lt;p&gt;Q: How does the paper address the concern that bridge timing might reflect anticipated local growth?&lt;/p&gt;
&lt;p&gt;A: The paper shows that results hold when restricting to counties whose distance to a bridge is only affected by bridges constructed in other counties, addressing the concern that local planners might time construction in anticipation of local growth. The results are also insensitive to controlling for pre-period trends, and outcomes of interest are uncorrelated with future changes in distance to a bridge in preferred specifications.&lt;/p&gt;
&lt;p&gt;Q: How does the paper reconcile the negative local income effect (tributary confluence strategy) with the positive aggregate effect (timing strategy)?&lt;/p&gt;
&lt;p&gt;A: The reconciliation proceeds through a narrative account combining industrialization, urbanization, and within-city sorting. Better bridge access drives a shift toward manufacturing employment and attracts population, consistent with a productivity advantage enabling exploitation of economies of scale. Cities form around historical transport routes. As cities mature and expand, richer households sort into lower-density suburban areas further from the historical transport corridor, while lower-income households remain near or migrate to the city centre. This within-city sorting produces lower per capita incomes near transport routes even as aggregate economic activity is higher in better-connected areas.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports the within-city sorting mechanism specifically?&lt;/p&gt;
&lt;p&gt;A: The negative income effect of proximity to transport routes is larger in more urbanized areas and in areas with higher income inequality. The effect is concentrated in areas that were more rapidly urbanizing in the 19th century, and it is stronger for non-white and low-education populations. Upstream census tracts simultaneously show higher manufacturing employment shares and higher population densities, consistent with cities having formed around transport routes, followed by residential sorting away from the core.&lt;/p&gt;
&lt;p&gt;Q: What are the two novel identification strategies and their broader applicability?&lt;/p&gt;
&lt;p&gt;A: The tributary confluence strategy exploits discontinuities in bridge construction costs generated by sharp increases in river flow rates at confluences; it requires only that bridges are more likely built upstream of confluences than downstream, an asymmetry the paper shows is detectable elsewhere in the world from satellite imagery. The timing strategy exploits the multi-decade planning and construction process for major bridges as a source of near-exogenous variation in opening dates. Both strategies can be applied in other settings where major rivers form substantial barriers to land transport networks.&lt;/p&gt;
&lt;p&gt;Q: What does the paper contribute to the debate about whether early U.S. transport infrastructure followed or led economic development?&lt;/p&gt;
&lt;p&gt;A: The results support the view that early investments in land transport infrastructure led to meaningful changes in economic geography rather than merely following pre-existing growth patterns. However, the paper finds a moderate level of responsiveness — population density responds to bridge access over several decades, not immediately — consistent with a broader literature documenting sluggish population responses to changes in economic conditions.&lt;/p&gt;
&lt;p&gt;Tributary confluence: A location where a smaller river (tributary) joins a larger river, causing a sharp, localized increase in downstream flow rates and therefore a discontinuous increase in bridge construction costs, generating the quasi-experimental variation in bridge location exploited in the paper.&lt;/p&gt;
&lt;p&gt;Within-city sorting: The process by which, as cities expand around historical transport routes, richer households differentially relocate to lower-density suburban areas further from the transport corridor while lower-income households remain near or migrate to the historical city centre, reversing the income gradient at small spatial scales.&lt;/p&gt;
&lt;p&gt;Income density: The product of population density and per capita income, corresponding to total economic activity per unit area; the paper finds income density is higher in better-connected upstream census tracts even when per capita income is lower, reflecting the dominant effect of higher population density.&lt;/p&gt;
&lt;p&gt;Farm land values: The total value of farm land and buildings, used as the best consistently available proxy for total economic activity in the 1860–2010 historical county panel; the paper treats estimated effects on farm land values as lower bounds on effects on total economic activity because farm values increasingly miss urban land as the economy industrializes.&lt;/p&gt;
&lt;p&gt;Structural transformation: The shift in the composition of employment away from agriculture and toward manufacturing, which the paper documents occurring in counties that experience improved bridge access, interpreted as evidence that transport infrastructure provides a productivity advantage attracting industrial activity.&lt;/p&gt;
&lt;p&gt;Distance to a bridge (as proxy for land transport access): In the study area along the Mississippi and Ohio Rivers, where all land has comparable water access, distance to the nearest bridge strongly predicts distance to the nearest major land transport route (rail or primary road), allowing bridge distance to serve as a consistent measure of transport connectivity throughout the entire study period.&lt;/p&gt;
&lt;p&gt;Market access: A measure of economic connectivity that captures both the state of the transport network and the size of accessible markets; the paper notes that log distance to a bridge explains 46% of the variation in market access in 1890 (from Donaldson and Hornbeck&amp;rsquo;s data) with an elasticity of approximately 0.1, and that halving distance to a bridge increases market access by approximately 7%.&lt;/p&gt;</description></item><item><title>Changing Opportunity: Sociological Mechanisms Underlying Growing Class Gaps</title><link>https://macropaperwarehouse.com/papers/changing-opportunity-sociological-mechanisms-underlying-growing-class-gaps/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/changing-opportunity-sociological-mechanisms-underlying-growing-class-gaps/</guid><description>&lt;p&gt;This paper documents sharp divergent trends in intergenerational economic mobility by race and class in the United States across the 1978 to 1992 birth cohorts, and investigates the causal mechanisms driving those changes. The core empirical facts are two: between 1978 and 1992 birth cohorts, the earnings gap between white children from high-income versus low-income families grew by approximately 28–30% (the &amp;ldquo;white class gap&amp;rdquo;), while the earnings gap between white and Black children from low-income families shrank by approximately 27–30% (the &amp;ldquo;white-Black race gap&amp;rdquo;). These twin trends — growing class gaps and shrinking race gaps — appear consistently across earnings, employment rates, educational attainment, SAT/ACT scores, incarceration, marriage, and mortality, and they hold in nearly every region of the country.&lt;/p&gt;
&lt;p&gt;The data are drawn from de-identified federal income tax returns linked to decennial census records and the Numident database, covering 57 million children born between 1978 and 1992, with information on parental and child incomes, employment, marital status, mortality, and residential location, supplemented by ACS educational attainment and linked SAT/ACT records covering 24.8 million students. Children&amp;rsquo;s outcomes are measured primarily as household income percentile ranks at age 27.&lt;/p&gt;
&lt;p&gt;In dollar terms, the white class gap (mean income difference between children raised at the 25th vs. 75th parental income percentile) grew from $17,720 to $20,950 in real 2023 dollars, while the white-Black race gap for low-income families fell from $20,810 to $14,910. The intergenerational rank-rank slope for white children increased from 0.23 to 0.29. The racial gap in intergenerational persistence of poverty — the probability of a child born to the bottom income quintile remaining there — shrank from 14.7 percentage points to 4.1 percentage points (a 72% reduction), driven roughly equally by improvement in Black children&amp;rsquo;s chances of escaping poverty and deterioration in low-income white children&amp;rsquo;s chances. The white class gap in early-adulthood mortality more than doubled, while the white-Black race gap in mortality fell by 77%.&lt;/p&gt;
&lt;p&gt;The paper systematically rules out three alternative explanations. Observable family characteristics (parental education, wealth, occupation, and marital status) explain only 7% of the growing white class gap and none of the shrinking white-Black race gap. Neighborhood-level common shocks, tested by including childhood county or Census tract-by-cohort fixed effects, similarly explain only 7% of the class gap and none of the race gap. The divergent trends persist even among children raised in the same Census tract, pointing to forces that operate differentially across race and class groups within the same neighborhood.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central finding is that changes in children&amp;rsquo;s outcomes across cohorts are strongly and positively correlated (r = 0.91 across subgroups) with changes in parental employment rates within the child&amp;rsquo;s social community, defined as families sharing the same race, class, and childhood county. Low-income white communities experienced sharp relative declines in parental employment rates; low-income Black communities experienced relative improvements. These community-level parental employment changes account for nearly all of the divergent trends.&lt;/p&gt;
&lt;p&gt;To establish causation, the paper exploits variation in the age at which children move to counties with changing parental employment rates. Children who moved at younger ages (before age 8) to counties where parental employment was increasing experienced larger improvements in earnings than those who moved at older ages (after age 13), consistent with a causal exposure effect with greater impact for longer durations of exposure. Sibling comparisons — comparing outcomes of younger versus older siblings who moved together — confirm that the age gradient reflects causal exposure rather than family-level selection.&lt;/p&gt;
&lt;p&gt;The social interaction mechanism is supported by two sources of variation: children&amp;rsquo;s outcomes are more strongly related to parental employment rates of their own birth cohort than adjacent cohorts (cohort specificity unlikely to be explained by resources), and outcomes are primarily driven by the employment rates of same-race, same-class community members, with cross-racial influence appearing only in counties where cross-racial interaction is greater (counties with small Black population shares or higher interracial marriage rates). The unified explanation the paper proposes is that children&amp;rsquo;s outcomes mimic those of the adults in their social communities, following Borjas (1992).&lt;/p&gt;
&lt;p&gt;Q: What are the precise magnitudes of the growing white class gap and shrinking white-Black race gap in income percentile ranks?
A: The white class gap — the difference in mean household income ranks between white children raised at the 25th versus 75th parental income percentiles — increased from 11.1 to 14.1 percentile ranks between the 1978 and 1992 birth cohorts, a 28% increase. The white-Black race gap for children from low-income families fell from 14.9 to 10.9 percentile ranks, a 27% decrease. The intergenerational rank-rank slope for white children increased from 0.23 to 0.29 (a 28% rise in persistence).&lt;/p&gt;
&lt;p&gt;Q: How did the trends in poverty persistence versus upward mobility differ?
A: The convergence in white-Black outcomes was driven almost entirely by changes in poverty persistence rather than upward mobility. The racial gap in the probability of remaining in the bottom income quintile shrank from 14.7 percentage points to 4.1 percentage points (a 72% reduction), with roughly half from Black children being less likely to remain at the bottom and half from white children being more likely to remain. By contrast, the white-Black gap in the probability of rising from the bottom quintile to the top quintile fell by only 1.9 percentage points (17%).&lt;/p&gt;
&lt;p&gt;Q: How widespread geographically were the divergent trends?
A: Outcomes declined for low-income white families in nearly every county, but the largest declines occurred in historically high-mobility areas such as the Great Plains and the coasts. For low-income Black families, outcomes improved in most areas, with the largest gains in historically low-mobility regions including the Southeast and the industrial Midwest. The correlation between county-level changes for low-income white versus low-income Black children is a positive 0.58, meaning the areas where Black families improved most tended to be areas where white families declined least, not most.&lt;/p&gt;
&lt;p&gt;Q: Do the trends persist when using non-rank, inflation-adjusted dollar outcomes?
A: Yes. The white class gap in mean household income grew from $17,720 to $20,950 in real 2023 dollars, and the white-Black race gap for low-income families narrowed from $20,810 to $14,910. The paper also reports similar patterns for individual earnings (as opposed to household income), ruling out changes in household composition as a driver.&lt;/p&gt;
&lt;p&gt;Q: What do the pre-labor-market outcomes show?
A: The divergent trends emerge before children enter the labor market. The white class gap in educational attainment grew by 20%, driven by growing gaps in four-year college completion. The white-Black race gap in educational attainment disappeared by the 1992 cohort, driven by narrowing gaps in high school graduation. The white class gap in the share of students taking the SAT/ACT increased by 12.1 percentage points between the 1980 and 1991 birth cohorts, while the white-Black race gap in SAT/ACT-taking decreased by 20.3 percentage points. The white class gap in mean SAT/ACT scores grew by 62% between the 1980 and 1997 birth cohorts among test-takers.&lt;/p&gt;
&lt;p&gt;Q: How large is the mortality dimension of these trends?
A: The white class gap in early-adulthood mortality (ages 24–27) more than doubled between the 1978 and 1992 birth cohorts, while the white-Black race gap in early-adulthood mortality decreased by 77%. These non-monetary outcomes are invariant to inflation and income measurement choices, confirming the robustness of the broader trends.&lt;/p&gt;
&lt;p&gt;Q: How much do family-level characteristics explain?
A: Controlling jointly for parental education, wealth, occupation, and marital status reduces the estimated growth in the white class gap by only 7% (from 3.37 to 3.13 percentile ranks). The same controls do not explain the shrinking white-Black race gap — the estimated reduction in the race gap actually becomes slightly larger (4.56 rather than 4.16 percentiles) after controlling for family characteristics, indicating that observable family factors work against the observed convergence.&lt;/p&gt;
&lt;p&gt;Q: How much do neighborhood-level common shocks explain?
A: Including childhood county fixed effects interacted with birth cohort explains only 7% of the growing white class gap and none of the shrinking white-Black race gap. Including Census tract fixed effects yields essentially identical results. The divergent trends persist among children growing up in the same Census tract, ruling out explanations based on differential exposure to neighborhood-level economic shocks.&lt;/p&gt;
&lt;p&gt;Q: What is the community-level parental employment correlation, and what does it explain?
A: Changes in children&amp;rsquo;s earnings, SAT/ACT scores, and educational attainment across cohorts are strongly positively correlated with changes in parental employment rates within the child&amp;rsquo;s community (same race, same class, same county), controlling for the employment status of the child&amp;rsquo;s own parents. The correlation between changes in children&amp;rsquo;s outcomes and changes in community parental employment rates across all race and class subgroups is 0.91. This single community-level factor — as proxied by parental employment rates — accounts for nearly all of the divergent trends by race and class.&lt;/p&gt;
&lt;p&gt;Q: What is the quasi-experimental design for estimating causal effects, and what does it assume?
A: The paper compares outcomes of children who moved to counties with increasing parental employment rates at younger versus older ages, across earlier versus later birth cohorts. The identification assumption is &amp;ldquo;constant selection by age&amp;rdquo;: any selection of families into moving to a given county in years when parental employment is higher may differ across cohorts, but those selection differences must not themselves vary systematically with the age at which children move. The paper treats this as a &amp;ldquo;constant selection by age&amp;rdquo; assumption standard in the neighborhood effects literature.&lt;/p&gt;
&lt;p&gt;Q: What do the causal exposure results show?
A: Children who moved before age 8 to communities where parental employment was increasing show systematically higher earnings in later birth cohorts, while children who made the same move after age 13 show little difference in earnings across cohorts. This pattern — larger effects at younger ages — is consistent with a causal exposure effect of growing up in an improving community, with effects proportional to the duration of exposure.&lt;/p&gt;
&lt;p&gt;Q: How do sibling comparisons validate the identification assumption?
A: When siblings move together to a community with increasing parental employment rates, the younger sibling — who receives more years of exposure to the higher-employment environment — earns significantly more than the older sibling. The earnings difference is proportional to the age gap between siblings. This rules out explanations based on fixed unobserved family characteristics and supports the constant-selection-by-age assumption.&lt;/p&gt;
&lt;p&gt;Q: What evidence distinguishes social interaction mechanisms from economic resource mechanisms?
A: Two sources of variation are used. First, children&amp;rsquo;s outcomes are much more strongly related to the parental employment rates of peers in their own birth cohort than peers in adjacent cohorts — a cohort-specificity that is implausible for economic resource channels (school budgets, local tax bases) which would not vary sharply across adjacent cohorts. Second, outcomes of low-income white children are driven primarily by the employment rates of low-income white parents, not by low-income Black or high-income white parents&amp;rsquo; employment rates, and vice versa for low-income Black children — consistent with interaction patterns being stratified by race and class.&lt;/p&gt;
&lt;p&gt;Q: What role does cross-racial interaction play?
A: In counties where Black children constitute a small share of the population (making cross-racial interaction more likely), Black children&amp;rsquo;s outcomes are also related to low-income white parental employment rates. Similarly, in counties with higher interracial marriage rates (a proxy for cross-racial interaction), Black children&amp;rsquo;s outcomes are related to white parental employment rates even after controlling for racial composition. This cross-sectional variation supports the interpretation that the influence channel is social interaction rather than parallel economic shocks.&lt;/p&gt;
&lt;p&gt;Q: How do the findings for Hispanic, Asian, and AIAN children compare?
A: Changes in economic mobility for Hispanic, Asian, and AIAN children between 1978 and 1992 birth cohorts were much more modest than for white and Black children. For children from low-income families, mean household income ranks were essentially unchanged for Asian children and rose by only about 0.5 percentiles for Hispanic and AIAN children. However, the same community-level parental employment rate mechanism explains the (smaller) changes for these groups as well; the correlation between changes in children&amp;rsquo;s outcomes and changes in community parental employment rates is 0.91 across all subgroups.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s unified theoretical account of all the divergent trends?
A: The paper concludes that a parsimonious theory — that children&amp;rsquo;s outcomes mimic those of the parents in their social communities, following Borjas (1992) — explains the divergent trends by race and class. Because social interaction is stratified by race and class even within neighborhoods, changes in parental outcomes in the parent generation propagate differentially to white versus Black and high-income versus low-income children, producing growing class gaps and shrinking race gaps through the same underlying mechanism.&lt;/p&gt;
&lt;p&gt;Q: What does the paper imply about the malleability of economic mobility disparities?
A: Because the causal exposure effects of community environments on children&amp;rsquo;s outcomes can be detected within a 14-year span (1978 to 1992 birth cohorts), the paper implies that differences in economic mobility by race and class may be malleable in policy-relevant timeframes. This is despite the fact that long-standing disparities partly trace back to historical factors such as slavery, Jim Crow laws, redlining, and the Great Migration.&lt;/p&gt;
&lt;p&gt;White class gap: The difference in mean household income ranks in adulthood for white children born to families at the 25th versus 75th percentiles of the national parental income distribution; increased from 11.1 to 14.1 percentile ranks (28%) between the 1978 and 1992 birth cohorts.&lt;/p&gt;
&lt;p&gt;White-Black race gap: The difference in mean household income ranks in adulthood for white versus Black children born to families at the 25th percentile of the national parental income distribution; decreased from 14.9 to 10.9 percentile ranks (27%) between the 1978 and 1992 birth cohorts.&lt;/p&gt;
&lt;p&gt;Social community: In this paper&amp;rsquo;s usage, other families who share the same race, class category, and childhood county as a given child; the unit within which community-level parental employment rates are measured and found to be predictive of children&amp;rsquo;s outcomes.&lt;/p&gt;
&lt;p&gt;Causal exposure effect: The effect on a child&amp;rsquo;s adult outcomes of an additional year spent growing up in a community with higher parental employment rates, estimated quasi-experimentally by comparing children who moved to counties with changing parental employment rates at younger versus older ages; larger effects at younger ages imply a causal, duration-sensitive exposure channel.&lt;/p&gt;
&lt;p&gt;Constant selection by age: The identification assumption underlying the quasi-experimental design; requires that any systematic differences in the types of families who move to a county when parental employment is high versus low do not themselves vary with the age at which children move to that county.&lt;/p&gt;
&lt;p&gt;Intergenerational rank-rank slope: The OLS slope coefficient from regressing child income percentile rank on parental income percentile rank; for white children, increased from 0.23 in the 1978 birth cohort to 0.29 in the 1992 birth cohort, indicating greater persistence of economic status.&lt;/p&gt;
&lt;p&gt;Cohort-specificity of community effects: The empirical pattern that children&amp;rsquo;s outcomes are more strongly related to the parental employment rates of peers in their own birth cohort than those of adjacent cohorts, used in the paper as evidence favoring social interaction over economic resource channels as the mediating mechanism.&lt;/p&gt;</description></item><item><title>Civil War–Induced Displacement and Human Capital</title><link>https://macropaperwarehouse.com/papers/civil-warinduced-displacement-and-human-capital/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/civil-warinduced-displacement-and-human-capital/</guid><description>&lt;p&gt;This paper examines the impact of conflict-driven forced displacement on human capital accumulation using the Mozambican civil war (1977–1992) as the empirical setting. During this war, over four million civilians — roughly a third of the population — fled to rural areas, cities, neighboring countries, or UN-managed refugee camps. The study advances on prior work in three dimensions: it uses the full post-war population census (12 million individuals) rather than a small survey; it studies multiple displacement trajectories in a single framework; and it separately identifies place-based exposure effects from a general uprootedness effect.&lt;/p&gt;
&lt;p&gt;The primary data source is the 1997 Mozambican census, which records each individual&amp;rsquo;s place of birth, residence in 1992 (the war&amp;rsquo;s end), and residence in 1997. Key outcomes are educational attainment and sectoral employment (agricultural versus services). The authors supplement the census with digitized colonial road and school maps, georeferenced conflict events, and landmine contamination data.&lt;/p&gt;
&lt;p&gt;The main identification strategy compares approximately 135,000 siblings (from 45,000 families) separated during the war, using the sibling who stayed behind as a within-family counterfactual. This design controls for household-level characteristics including religious and ethnic background, aspirations, and exposure to violence.&lt;/p&gt;
&lt;p&gt;The key findings are as follows. First, rural-born IDPs displaced to cities have a 7.3 percentage point higher likelihood of attending primary school and 0.53 more years of schooling compared to their siblings who stayed behind — roughly one-third of the non-displaced mean. Rural-born IDPs displaced to other rural areas also show gains, with a 3 percentage point higher likelihood of attending school and 0.24 additional years, supporting the uprootedness hypothesis even for displacements that did not reach urban centers. Urban-born IDPs forcibly relocated to the countryside — primarily through FRELIMO&amp;rsquo;s villagization scheme — experienced 9 percentage point lower primary school attendance and approximately 0.5 fewer years of schooling relative to siblings who remained in cities.&lt;/p&gt;
&lt;p&gt;External displacement (to camps in Malawi or Zimbabwe) generated no significant schooling gains relative to staying siblings, despite UN-built schools in camps, likely because scarce employment opportunities reduced perceived returns to education.&lt;/p&gt;
&lt;p&gt;Second, the paper jointly estimates place-based and uprootedness effects in a single within-family framework. Place effects are statistically significant: displacement to a district one standard deviation more developed than one&amp;rsquo;s birthplace raises schooling likelihood by approximately 3 percentage points (OLS) to 5 percentage points (2SLS reduced form). Crucially, a residual uprootedness effect of approximately 2–4 percentage points persists even after controlling fully for destination-origin differences in development and conflict intensity. This uprootedness effect is quantitatively comparable to being displaced to a district one standard deviation more developed than one&amp;rsquo;s birthplace.&lt;/p&gt;
&lt;p&gt;Third, a primary survey of 208 Nampula residents conducted in early 2020 — three decades after the war — confirms lasting educational gains. IDPs displaced to Nampula have a 10 percentage point higher likelihood of completing primary school relative to their siblings who stayed in the countryside, and their educational attainment converged to levels of urban-born, never-displaced residents despite large urban-rural education gaps. However, IDPs report significantly lower social capital, civic participation, and community trust than urban-born respondents, and score significantly worse on mental health indicators, including depression, loneliness, and pessimism. These psychosocial costs persist three decades after the war&amp;rsquo;s end.&lt;/p&gt;
&lt;p&gt;The findings apply to a low-income, post-colonial African setting characterized by widespread illiteracy (over 60%) and subsistence agriculture (over 85% of employment) at the war&amp;rsquo;s close. The results are robust to alternative age restrictions, extended family comparisons, dropping the oldest sibling, same-sex sibling pairs, and narrowing the age gap between sibling pairs to as few as two years.&lt;/p&gt;
&lt;p&gt;Q: What is the core identification strategy and why is it preferred over cross-sectional estimates?
A: The authors compare siblings within the same household who experienced different displacement trajectories during the war. Because siblings share household-level characteristics — parental preferences for education, ethnic and religious background, wealth, and local conflict exposure — the within-family design controls for confounders that would bias cross-sectional estimates. The within-family estimates are systematically smaller than cross-sectional ones (e.g., 7.3 pps vs. 24–30 pps for rural-to-urban displacement in primary school attendance), confirming that sorting was present even in the unpredictable civil war setting.&lt;/p&gt;
&lt;p&gt;Q: What do the results show for rural-born IDPs displaced to urban centers?
A: Within the sibling-pair framework, rural-born IDPs displaced to cities and towns have a 7.3 percentage point higher likelihood of attending primary school and 0.53 more years of schooling compared to their siblings who stayed in rural birthplaces, against a non-displaced sibling mean of approximately 20% primary school access and one year of formal schooling. These IDPs also show a 4 percentage point higher likelihood of non-agricultural employment five years after the war&amp;rsquo;s end.&lt;/p&gt;
&lt;p&gt;Q: What do the results show for rural-born IDPs displaced to other rural areas?
A: Even displacement to a different rural district — not a city — generates modest but statistically significant gains: a 3 percentage point higher likelihood of attending school and 0.24 additional years of schooling relative to siblings staying in their birthplace rural district. The authors interpret this as evidence for the uprootedness hypothesis, since rural Mozambique at the time was among the most impoverished and insecure environments in the world, meaning destination quality alone cannot explain the gain.&lt;/p&gt;
&lt;p&gt;Q: What do the results show for externally displaced refugees?
A: Refugees displaced to camps and settlements in Malawi, Zimbabwe, Tanzania, Zambia, and Swaziland show schooling levels statistically similar to their siblings who remained in their rural birthplaces, despite UN-built primary schools in camps. The authors attribute the absence of gains to low perceived returns to education stemming from scarce employment opportunities at displacement destinations. Externally displaced individuals do show a 5 percentage point lower likelihood of agricultural employment relative to staying siblings.&lt;/p&gt;
&lt;p&gt;Q: What are the consequences of urban-to-rural forced displacement?
A: Urban-born individuals forcibly relocated to the countryside — primarily through FRELIMO&amp;rsquo;s villagization and food production programs — have approximately 9 percentage point lower likelihood of attending primary school and 0.5 fewer years of schooling compared to siblings who remained in urban areas. These results indicate that FRELIMO&amp;rsquo;s coercive relocation policies imposed material human capital costs on the displaced.&lt;/p&gt;
&lt;p&gt;Q: How are place-based and uprootedness effects separated empirically?
A: The authors construct principal component indices for destination-origin differences in regional development (aggregating population density, Portuguese-speaking share, offspring mortality, road density, colonial market density, and school density) and conflict intensity (conflict events per capita and landmine contamination per capita). They then include these continuous exposure measures alongside a binary displacement indicator in within-family regressions. The coefficient on the binary displacement indicator — conditional on destination-origin development and conflict differences — isolates the uprootedness effect for individuals displaced to districts with identical characteristics to their birthplace.&lt;/p&gt;
&lt;p&gt;Q: What are the magnitudes of the place-based and uprootedness effects?
A: Under OLS, displacement to a district one standard deviation more developed than one&amp;rsquo;s birthplace raises schooling likelihood by approximately 3 percentage points. The residual uprootedness effect — displacement per se, controlling for destination quality — raises schooling likelihood by approximately 2 percentage points. Under 2SLS (instrumenting destination-origin development differences with the development of districts within 100 km of birthplace), the place-based effect rises to approximately 5 percentage points in the reduced form, and the uprootedness effect remains significant at approximately 4 percentage points. Both the uprootedness and place-based effects are of comparable magnitude.&lt;/p&gt;
&lt;p&gt;Q: What instrument is used in the 2SLS specifications and what is its first-stage strength?
A: The instrument exploits the fact that Mozambique&amp;rsquo;s heavily mined and rudimentary transportation network constrained civilian movement — the median displaced sibling ended up roughly 97 kilometers from birthplace. The authors instrument actual destination-origin development and conflict differences with the predicted differences based on the characteristics of districts within 100 km of the birthplace. The first-stage elasticity between actual and proximity-predicted differences in development is 0.86, and for conflict is 0.88, both precisely estimated.&lt;/p&gt;
&lt;p&gt;Q: What do the long-run survey results from Nampula show about educational persistence?
A: In a 2020 survey of 208 Nampula residents aged over 35, IDPs who fled to Nampula during the war have a 10 percentage point higher likelihood of completing primary school relative to their siblings who stayed in the countryside. Their educational attainment converges to the level of urban-born, never-displaced Nampula residents, despite large historical and contemporary urban-rural education gaps in northern Mozambique. The majority of IDPs (73%) report that extended relatives or friends advised them to attend school upon arriving in the city, and most believed education was necessary for urban employment.&lt;/p&gt;
&lt;p&gt;Q: What are the long-run psychosocial costs documented in the Nampula survey?
A: Even three decades after the war&amp;rsquo;s end, IDPs in Nampula report significantly lower social capital, civic participation, and community trust compared to urban-born never-displaced residents. IDPs also score significantly worse on mental health indicators including depression, loneliness, and pessimism. These findings suggest that forced displacement imposes persistent psychosocial costs that are not remediated by economic or educational convergence.&lt;/p&gt;
&lt;p&gt;Q: What drives displacement in the data, and does selection threaten identification?
A: Linear probability and multinomial logit models show that conflict intensity and geographic proximity (distance to the border for external displacement; distance to cities for urban displacement) are the primary correlates of displacement type, while differences in destination development are uncorrelated with displacement. Nevertheless, the overall explanatory power of these models is low, confirming many idiosyncratic and unpredictable features of the war. The within-family design addresses residual selection on household characteristics, and the 2SLS design addresses selection on destination-specific characteristics.&lt;/p&gt;
&lt;p&gt;Q: How do educational gains translate into sectoral employment outcomes?
A: Across specifications, gains in schooling move in tandem with a shift out of agriculture into services. Rural-to-urban IDPs have a 4 percentage point higher likelihood of non-agricultural employment five years after the war, while externally displaced show a 5 percentage point lower likelihood of agricultural employment. Urban-born IDPs displaced to the countryside are more likely to work in agriculture after the war. The authors interpret this co-movement as suggesting that conflict-driven human capital accumulation may contribute to structural transformation away from subsistence agriculture.&lt;/p&gt;
&lt;p&gt;Q: How robust are the within-family estimates?
A: The authors conduct six sensitivity checks: adding family fixed effects to cross-sectional regressions, restricting to individuals aged 12–18 in 1997 to address co-habitation concerns, extending comparisons to cousins and other relatives, dropping the oldest male sibling to minimize favoritism concerns, restricting to same-sex sibling pairs, and narrowing the age gap to two years. Across all permutations, the qualitative ordering is preserved: refugees show no significant schooling gains, rural-to-urban IDPs show gains of 5–6 percentage points in primary attendance and 0.35–0.5 extra years, rural-to-rural IDPs show small positive gains, and urban-to-rural IDPs show losses.&lt;/p&gt;
&lt;p&gt;Uprootedness hypothesis: The idea, traced in the paper to Stigler and Becker (1977) and earlier scholars, that forced displacement incentivizes human capital investment precisely because education is a mobile asset that cannot be expropriated — distinct from place-based effects of destination quality.&lt;/p&gt;
&lt;p&gt;Place-based (exposure) effects: The impact on human capital outcomes attributable to differences between the development level and conflict intensity of the displacement destination and the individual&amp;rsquo;s birthplace, measured as destination-origin differences in a principal component index of regional development.&lt;/p&gt;
&lt;p&gt;Separated siblings design: An identification strategy that compares siblings from the same household who experienced different displacement trajectories during the war, holding constant all household-level characteristics including parental preferences, ethnicity, religion, wealth, and local conflict exposure.&lt;/p&gt;
&lt;p&gt;Internal displacement (IDP): Conflict-driven movement within national borders to either rural areas or urban centers, constituting approximately 60% of global forced displacement and the majority of displacement in the Mozambican civil war context.&lt;/p&gt;
&lt;p&gt;Source text origin: A categorization of the working paper text used for summarization — distinguishing full PDF or HTML text from abstract-only text. Abstract-only text is a hard block for summary generation in the pipeline.&lt;/p&gt;
&lt;p&gt;Structural transformation: In this paper&amp;rsquo;s usage, the shift of workers out of subsistence agriculture into services associated with human capital accumulation triggered by conflict-driven displacement, treated as a potential mechanism of post-conflict recovery.&lt;/p&gt;
&lt;p&gt;Psychosocial costs of displacement: Long-run deficits in social capital, civic engagement, community trust, and mental health (depression, loneliness, pessimism) reported by IDPs three decades after displacement, persisting despite convergence in educational attainment and employment.&lt;/p&gt;</description></item><item><title>Closing Gender Gaps Through Workplace Diversity: The Intergenerational Effects of World War I</title><link>https://macropaperwarehouse.com/papers/closing-gender-gaps-through-workplace-diversity-the-intergenerational-effects-of-world-war-i/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/closing-gender-gaps-through-workplace-diversity-the-intergenerational-effects-of-world-war-i/</guid><description>&lt;p&gt;This paper asks whether exposure to greater female representation in the workplace can persistently reduce intergenerational gender gaps in labor market outcomes. The authors exploit the sudden, city-by-department variation in female employment within the U.S. federal government triggered by World War I mobilization. Using the Official Registers of the United States — biennial personnel rosters covering the near-universe of federal employees from 1913 to 1921 — linked to full-count decennial censuses (1900–1940), they construct a granular measure of each office&amp;rsquo;s (city × department) change in female share between 1915 and 1919, then trace labor force outcomes for the children of incumbent civil servants in the 1940 Census.&lt;/p&gt;
&lt;p&gt;WWI caused the female share of the federal civilian workforce to jump by 13 percentage points — a doubling within two years (1917–1919). These wartime female entrants were younger, more likely to be single, more educated, more geographically mobile, and less likely to have been previously employed than their male counterparts, suggesting the war mobilized a previously untapped labor pool. The increase was driven almost entirely by clerical positions: the female share of the federal clerical workforce rose from roughly 30% to nearly 70% within two years.&lt;/p&gt;
&lt;p&gt;The main finding is that a one standard deviation (SD) increase in parental exposure to female co-workers reduces the gender gap in labor force participation (LFP) among children of incumbent civil servants by 4.1–4.6 percentage points in the within-city, within-department specification — a decline in the mean gender LFP gap of approximately 8.6–9.6% by 1940. This effect is entirely driven by a higher propensity of daughters to work; sons&amp;rsquo; LFP is unaffected. The intergenerational effect operates primarily through exposed fathers, including fathers without working wives, identifying a channel beyond the mother-to-daughter vertical transmission emphasized in prior literature. Children who were teenagers at the time of parental exposure show the largest effects, consistent with formative-years malleability. A placebo test using civil servants who left the same offices before the wartime shock shows no comparable effect, ruling out time-invariant office-level selection.&lt;/p&gt;
&lt;p&gt;Parental exposure extends beyond the public sector: the private sector LFP effect is comparable in magnitude to the public sector effect. The gender earnings gap among children of exposed civil servants narrows by 12%, driven by daughters moving into higher-paying, previously male-dominated positions rather than by differences in hours or weeks worked. Marriage, fertility, and schooling differences only partially mediate the LFP effect, with a residual exposure effect remaining after controlling for these proximate determinants.&lt;/p&gt;
&lt;p&gt;At the aggregate level, a 1 SD increase in city-level exposure to female federal workers raises overall female LFP by 0.9–1.0 percentage points, with no effect on male LFP, and the effect persists through 1940. A back-of-envelope calculation implies each additional female wartime civil service entrant generated approximately 2.4 additional women entering the workforce — a multiplier effect. Neighborhood-level analysis shows LFP gains are concentrated in enumeration districts where wartime female civil servants resided, and cities with greater female federal employment exposure also saw faster women&amp;rsquo;s club membership growth after WWI.&lt;/p&gt;
&lt;p&gt;The scope conditions are important: the sample covers 70 cities and 8 federal departments with meaningful pre-war staffing; children must have been born by 1917; and the 1940 outcomes reflect adulthood labor decisions in a labor market shaped by subsequent decades of change. The design relies on within-city and within-department residual variation in female share change being conditionally exogenous, supported by lack of correlation with pre-war office characteristics.&lt;/p&gt;
&lt;p&gt;Q: What was the scale of the WWI shock to female federal employment?
A: The U.S. entry into WWI in April 1917 triggered a near-doubling of total federal civilian employment from roughly 150,000 to over 300,000 workers by 1919. Within this expansion, the share of female civil servants increased by 13 percentage points — a doubling of the female share within two years. The increase was driven almost entirely by clerical positions, where the female share rose from around 30% to nearly 70%.&lt;/p&gt;
&lt;p&gt;Q: How do the authors measure parental exposure to female co-workers?
A: Exposure is measured as the change in the share of female civil servants at the city-by-department (&amp;ldquo;office&amp;rdquo;) level between 1915 and 1919. The sample is restricted to offices with at least 20 civil servants in 1915 and cities with at least two federal departments, yielding 70 cities and 8 departments. The interquartile range of exposure across offices is approximately 10 percentage points, and cross-city and cross-department variation explains 58% of the overall variation, leaving substantial residual office-level variation for identification.&lt;/p&gt;
&lt;p&gt;Q: What is the main intergenerational finding and its magnitude?
A: A 1 SD increase in parental exposure to female co-workers increases the relative likelihood that a daughter works (compared to a son) by 2 percentage points in the baseline specification, and by 4.1–4.6 percentage points in the preferred within-city and within-department specification. Since daughters of civil servants are on average 48 percentage points less likely than sons to be in the labor force in 1940, this corresponds to closing the mean gender LFP gap by approximately 8.6–9.6%.&lt;/p&gt;
&lt;p&gt;Q: Does the effect operate through daughters or sons?
A: The effect is entirely driven by daughters. Parental exposure to female co-workers has no statistically discernible impact on the labor force participation of sons. The decline in the gender LFP gap is thus attributable to a higher propensity of daughters of exposed civil servants to work.&lt;/p&gt;
&lt;p&gt;Q: What is the key placebo test, and what does it show?
A: The authors exploit high-frequency personnel records to identify civil servants who selected into the same offices that would later be exposed but who left before the wartime shock occurred. These pre-departure leavers show no intergenerational exposure effects on their children&amp;rsquo;s LFP, ruling out the interpretation that time-invariant selection into particular offices drives the results.&lt;/p&gt;
&lt;p&gt;Q: Which parent serves as the primary channel of transmission?
A: Exposed fathers are the primary conduit. The effect for daughters is precise and sizable even when restricting the sample to fathers without working wives, suggesting the channel does not depend on children observing maternal employment. While the estimated effect through mothers is positive, it is imprecise — likely due to the small sample of female incumbent civil servants. This identifies fathers as a new channel of vertical intergenerational norm transmission, beyond the mother-to-daughter pathway emphasized in prior literature.&lt;/p&gt;
&lt;p&gt;Q: How does children&amp;rsquo;s age at the time of parental exposure moderate the effect?
A: The exposure effects are concentrated among children who were teenagers at the time of parental exposure during WWI. Children who were older and more likely to have already left the household or formed fixed beliefs show little to no detectable effect. This pattern is consistent with the formative-years hypothesis that experiences during adolescence shape lifetime economic behavior.&lt;/p&gt;
&lt;p&gt;Q: Does the intergenerational effect extend beyond the public sector?
A: Yes. The private sector LFP effect for daughters is comparable in magnitude to the public sector effect, with a 1 SD increase in parental exposure having approximately equal effects on LFP within public and private employment. There is also no measurable shift toward clerical occupations specifically, suggesting the channel is a broader change in attitudes toward women working, not transmission of information about specific government or clerical jobs.&lt;/p&gt;
&lt;p&gt;Q: What is the effect on the gender earnings gap?
A: A 1 SD increase in parental exposure to female co-workers closes the gender earnings gap among children of civil servants by 12%. This is not driven by differences in weeks or hours worked, but rather by daughters of exposed parents selecting into higher-paying and previously male-dominated occupations.&lt;/p&gt;
&lt;p&gt;Q: How do the authors address the possibility that the results reflect local labor market conditions rather than parental exposure per se?
A: By 1940, 67% of civil servant children lived in a city different from their parent&amp;rsquo;s WWI-era city. Even among children who moved to the same destination city — and thus face identical labor market conditions — variation in parental exposure at the origin city-by-department remains highly predictive of daughters&amp;rsquo; LFP. Comparing children moving from the same origin city to the same destination city, those with parents in higher-exposure departments still show higher LFP, pointing to cultural transmission rather than local labor market demand.&lt;/p&gt;
&lt;p&gt;Q: What do the marriage and fertility results indicate about mechanisms?
A: Daughters of more exposed civil servants are less likely to be married (a 1 SD increase in parental exposure reduces the relative likelihood of daughters being married by 3.7 percentage points) and tend to have fewer children by 1940. A mediation exercise shows these observable differences in marriage, fertility, and education only partially explain the LFP increase; a statistically significant and economically large residual exposure effect remains, consistent with parental exposure shifting broader gender norms rather than only proximate determinants of labor supply.&lt;/p&gt;
&lt;p&gt;Q: What does the spousal work decision evidence contribute?
A: A 1 SD increase in male civil servants&amp;rsquo; exposure to female co-workers increases the propensity of their subsequent wife to work by 0.5 percentage points after WWI. The effect is driven by marriages formed after the exposure and is not mechanically explained by men marrying their female co-workers. This revealed preference measure supports the interpretation that exposure changed men&amp;rsquo;s attitudes toward women&amp;rsquo;s work.&lt;/p&gt;
&lt;p&gt;Q: What do naming patterns suggest about changing attitudes?
A: Exposed parents are more likely to give daughters names that are less feminine — specifically, names with a lower share of vowels or less likely to end with a vowel — for daughters born after WWI. No comparable effect is observed for sons&amp;rsquo; names. This provides supplementary evidence of a shift in paternal attitudes following workplace exposure to female co-workers.&lt;/p&gt;
&lt;p&gt;Q: What are the aggregate city-level effects on female LFP?
A: In a difference-in-differences design using cross-city variation in female federal worker exposure before and after WWI, a 1 SD increase in city-level exposure raises aggregate female LFP by 0.9–1.0 percentage points, with no effect on male LFP. The effect is persistent through 1940 and city-level exposure is uncorrelated with female LFP prior to WWI. A back-of-envelope calculation implies each additional female wartime entrant generated approximately 2.4 additional women entering the broader workforce — a social multiplier.&lt;/p&gt;
&lt;p&gt;Q: Is there evidence of horizontal (non-family) transmission?
A: Yes. The aggregate LFP gains are concentrated almost entirely in census enumeration districts where female wartime civil servants resided; neighboring districts without female entrants do not see comparable gains. Cities with greater increases in female federal employees also experienced faster growth in women&amp;rsquo;s club memberships, with this pattern appearing only after WWI and coinciding with the rise in female LFP. Both findings are consistent with social learning operating through residential proximity and community networks.&lt;/p&gt;
&lt;p&gt;Q: How robust are the results to potential selection bias from imperfect census linking?
A: The propensity of a civil servant&amp;rsquo;s child to be linked to the 1940 Census is — conditional on city and department fixed effects — uncorrelated with the parental exposure measure. The authors apply inverse probability weighting (IPW) to ensure the matched sample is balanced on baseline characteristics, and results remain virtually identical. Estimates are also stable across different linking strategies individually.&lt;/p&gt;
&lt;p&gt;Q: What instrumental variable strategy is used and what does it find?
A: The authors instrument for office-level female share change using the interaction of the 1915 clerical workforce share and an indicator for war-related departments — a pre-determined source of variation in the capacity and demand for female clerical workers. The IV estimates are consistent with the OLS main specification: parental exposure to female co-workers closes the children&amp;rsquo;s gender LFP gap.&lt;/p&gt;
&lt;p&gt;Q: What is the policy implication regarding public sector hiring?
A: The paper suggests that increasing gender representation within public sector employment can have labor market implications that extend well beyond the organization itself — across generations through vertical intergenerational transmission and across the broader community through horizontal social spillovers. The findings imply that public sector diversity policies can serve as a lever for broader, persistent reductions in gender gaps in the private labor market.&lt;/p&gt;
&lt;p&gt;Office-level exposure: The city-by-department measure of the change in female share of civil servants between 1915 and 1919, capturing the granular intensity of each workplace unit&amp;rsquo;s contact with wartime female entrants; the interquartile range across offices is approximately 10 percentage points.&lt;/p&gt;
&lt;p&gt;Intergenerational gender gap in LFP: The difference in labor force participation rates between daughters and sons of incumbent civil servants measured in 1940 adulthood, used as the primary outcome to capture whether parental workplace exposure transmits to children&amp;rsquo;s labor supply decisions.&lt;/p&gt;
&lt;p&gt;Vertical transmission: The intergenerational channel through which exposed parents — identified here primarily as fathers, including those without working wives — convey changed attitudes or information about female work to their children, closing the gender LFP gap.&lt;/p&gt;
&lt;p&gt;Horizontal transmission: The community-level channel through which the increased presence of female civil servants in a city spreads changed norms or information about women&amp;rsquo;s work to women who are not daughters of exposed co-workers, operating through residential proximity and social networks such as women&amp;rsquo;s clubs.&lt;/p&gt;
&lt;p&gt;Social multiplier: The amplification of the direct effect of hiring female workers through behavioral spillovers; the authors&amp;rsquo; back-of-envelope calculation estimates that each additional female wartime civil service entrant generated approximately 2.4 additional women entering the workforce.&lt;/p&gt;
&lt;p&gt;Formative years: The period of adolescence during which children are argued to be most malleable in forming preferences and beliefs; exposure effects in this paper are concentrated among children who were teenagers at the time of parental exposure, with older children showing little effect.&lt;/p&gt;
&lt;p&gt;Source text origin: The authors&amp;rsquo; classification of whether a summary is based on full working paper text (pdf or oa-html) vs. abstract only; in this workflow, abstract-only is a hard block for summary generation.&lt;/p&gt;</description></item><item><title>Community Engagement and Public Safety: Evidence from Crime Enforcement Targeting Immigrants</title><link>https://macropaperwarehouse.com/papers/community-engagement-and-public-safety-evidence-from-crime-enforcement-targeting-immigrants/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/community-engagement-and-public-safety-evidence-from-crime-enforcement-targeting-immigrants/</guid><description>&lt;p&gt;This paper studies how immigration enforcement affects public safety, asking two questions: (1) what is the effect of increased enforcement on criminal victimization, and (2) how does increased enforcement affect victims&amp;rsquo; willingness to report crimes to police? The authors exploit the staggered rollout of the U.S. Secure Communities (SC) program — the largest expansion of interior immigration enforcement in U.S. history — across counties between 2008 and 2013. SC expanded information sharing between local police and federal immigration authorities, causing ICE honored detainer requests to increase by over 50% following program activation.&lt;/p&gt;
&lt;p&gt;The primary data source is the restricted-access National Crime Victimization Survey (NCVS), which measures victimizations independently of whether they were reported to police and includes respondent ethnicity. This allows the authors to separately estimate effects on underlying crime incidence and on reporting behavior for Hispanic and non-Hispanic individuals. The empirical strategy uses a staggered difference-in-differences design following Sun and Abraham (2021), comparing earlier-treated counties to the last 25% of counties to activate SC, with estimates run separately by ethnicity.&lt;/p&gt;
&lt;p&gt;The main findings run contrary to the stated policy goal of improving public safety. Among Hispanic individuals, SC caused a statistically significant 0.15 percentage point increase in monthly victimization — a 16% increase relative to the pre-period baseline of 0.9 percentage points — implying approximately 1.3 million additional crimes against Hispanics in the two years following program activation. The increase is concentrated primarily in property crimes (a statistically significant 15% increase), with a similarly sized but imprecisely estimated 15% increase in violent crime victimizations. The victimization increase is larger for Hispanic females (0.23 percentage points, or 25%) and in counties with higher shares of non-citizen Hispanic residents.&lt;/p&gt;
&lt;p&gt;Simultaneously, SC caused a 9.5 percentage point decline in the likelihood that Hispanic victims report incidents to police — a 30% decline relative to the pre-period mean reporting rate of 33 percentage points. This reporting decline is primarily driven by a 34% decline in the reporting of property offenses. No changes in victimization or reporting are found for non-Hispanic individuals in the aggregate, though non-Hispanic individuals in neighborhoods with high Hispanic population shares do experience higher victimization rates after SC.&lt;/p&gt;
&lt;p&gt;Critically, reported crime rates (the product of victimization and reporting) are unchanged for both Hispanic and non-Hispanic individuals, explaining why prior studies using administrative reported-crime data found null effects of SC. The null effect on reported crime masks two large, opposing causal forces.&lt;/p&gt;
&lt;p&gt;The authors provide evidence that the decline in crime reporting is the primary driver of the increase in victimization. Cohorts with larger reporting declines experienced larger victimization increases, and a decomposition exercise shows the reporting decline is substantially more important than concurrent SC-induced changes in unemployment, wages, female-headed household shares, and the male immigrant share. Supporting data from 75 police departments confirm no change in 911 call volumes or total arrest volumes, while showing a decline in the Hispanic share of arrestees in both Hispanic and non-Hispanic neighborhoods — consistent with reduced reporting leading to reduced apprehension of offenders, with offending shifting toward non-Hispanic individuals.&lt;/p&gt;
&lt;p&gt;Scope conditions: results are estimated for the population residing in counties exceeding 100,000 residents (representing 61% of total U.S. population and 69% of the Hispanic population), excluding southern border counties and states that actively resisted SC implementation (Illinois, Massachusetts, New York). Effects apply to all Hispanic respondents — citizens and non-citizens — consistent with prior evidence that citizen Hispanics respond to immigration enforcement out of concern for non-citizen contacts.&lt;/p&gt;
&lt;p&gt;Q: What was the Secure Communities program and how was it implemented?
A: SC was a federal program launched in 2008 that required fingerprints of individuals booked into local jails to be forwarded not only to the FBI but also to the Department of Homeland Security, enabling automatic screening for immigration violations. Local authorities could not prevent federal officials from learning of an arrestee&amp;rsquo;s immigration status. The program rolled out county-by-county between October 2008 and January 2013 due to technological constraints and resource bottlenecks, generating the staggered variation used for identification.&lt;/p&gt;
&lt;p&gt;Q: How large was the first-stage effect on actual immigration enforcement?
A: County-level honored ICE detainer requests increased by over 50% following SC activation, with a similar 40% increase in all detainer requests. The number of honored detainers nationwide doubled between 2008 and 2012. Over 90% of detainers and removals in any given month were for individuals of Hispanic ethnicity.&lt;/p&gt;
&lt;p&gt;Q: What is the main finding on Hispanic victimization?
A: SC caused a 0.15 percentage point increase in monthly Hispanic victimization rates, a 16% increase relative to the pre-period baseline of 0.9 percentage points. This translates to approximately 1.3 million additional crimes against Hispanics over two years following program activation, calculated by multiplying the monthly effect by 24 months and the 35.3 million Hispanics in the sample counties.&lt;/p&gt;
&lt;p&gt;Q: What is the main finding on Hispanic crime reporting?
A: SC caused a 9.5 percentage point decline in the likelihood that Hispanic victims report incidents to police, a 30% decline relative to the pre-period mean reporting rate of 33 percentage points. This decline occurred relatively quickly after activation and was concentrated in property offenses, where reporting fell by 34%.&lt;/p&gt;
&lt;p&gt;Q: Why do reported crime rates show no change despite large shifts in victimization and reporting?
A: Reported crime rates — the probability of being victimized and reporting the crime — are unchanged because the 16% increase in victimization and the 30% decline in reporting are approximately offsetting in magnitude. This explains why prior work using administrative police data (Miles and Cox 2014; Treyger et al. 2014; Hines and Peri 2019) found null effects of SC on reported crime: those data sources cannot separately identify the two underlying changes.&lt;/p&gt;
&lt;p&gt;Q: Does SC affect non-Hispanic individuals?
A: In the aggregate, SC has no statistically significant effect on non-Hispanic victimization or reporting. However, non-Hispanic individuals living in neighborhoods with high Hispanic population shares do experience victimization increases, and in those neighborhoods their reporting rates also decline slightly. Re-weighting non-Hispanic respondents to match the county composition of Hispanic respondents yields an 8% increase in non-Hispanic victimization, suggesting spillover effects in Hispanic-dense areas.&lt;/p&gt;
&lt;p&gt;Q: What mechanism links the reporting decline to the victimization increase?
A: The authors argue that reduced victim reporting lowers the probability that offenders are apprehended, thereby reducing the cost of committing crimes. They demonstrate this through two analyses: first, cohorts of counties with larger reporting declines experienced larger victimization increases; second, a decomposition shows the reporting channel is substantially more important than concurrent SC-induced changes in unemployment, wages, female-headed household shares, and the male immigrant share of the population.&lt;/p&gt;
&lt;p&gt;Q: What do the police administrative data show about offender composition?
A: Data from 75 police departments show no change in 911 call volumes or total arrest volumes following SC — consistent with the NCVS finding of unchanged reported crime rates. However, the Hispanic share of arrestees declined after SC, with a 1.5 percentage point drop in Hispanic neighborhoods (off a base of 54%), suggesting the rise in offending was more concentrated among non-Hispanic offenders as reduced reporting lowered expected punishment probabilities.&lt;/p&gt;
&lt;p&gt;Q: How does the victimization effect vary by gender?
A: The victimization point estimate for Hispanic males is 0.085 percentage points and imprecisely estimated (SE = 0.088). For Hispanic females, the effect is over 2.5 times larger at 0.23 percentage points, a 25% increase. The decline in reporting is comparable in magnitude across male and female Hispanic victims, suggesting fear of enforcement is similar by gender but that females disproportionately bear the crime burden.&lt;/p&gt;
&lt;p&gt;Q: How does the victimization effect vary by neighborhood non-citizen Hispanic share?
A: Victimization effects for Hispanics are relatively constant across neighborhood types but are higher — around 25% — in neighborhoods with the highest shares of non-citizen Hispanics. Counties with higher non-citizen Hispanic shares also exhibit higher ICE removal rates, indicating greater total enforcement, and these counties have higher victimization effects. Reporting declines among Hispanics appear relatively uniform across neighborhood types.&lt;/p&gt;
&lt;p&gt;Q: Could survey attrition or compositional changes explain the results?
A: The authors rule this out through several tests. First, SC has no statistically significant effect on household survey response rates, even in Census tracts above the 90th percentile of Hispanic share. A worst-case bias calculation implies attrition could account for at most 26% of the victimization effect. Second, re-estimating using predicted victimization (based on pre-SC demographics) yields precise null effects, indicating the increase is not driven by compositional change. Third, results are stable when restricting to respondents present at all survey waves or using individual fixed effects.&lt;/p&gt;
&lt;p&gt;Q: Could the reporting decline be mechanical — reflecting a change in the types of crimes committed rather than behavioral change?
A: The authors test this by constructing predicted reporting rates using pre-SC incident characteristics. The largest alternative estimate is -1.45 percentage points, over six times smaller than the estimated main reporting effect of 9.5 percentage points, ruling out crime composition change as the primary explanation. Results also hold when focusing on always-respondents and using individual fixed effects, ruling out entry of low-reporting individuals into the survey.&lt;/p&gt;
&lt;p&gt;Q: How robust are the results to alternative empirical strategies?
A: Results are robust to including states that resisted SC (with somewhat smaller magnitudes as expected), alternative population cutoffs, TWFE specifications, the Borusyak et al. (2021) and Callaway and Sant&amp;rsquo;Anna (2021) estimators (which yield larger point estimates), a triple-differences specification using non-Hispanics as an additional control group, and the inclusion of time-varying unemployment rates. The dynamic event-study plots show parallel pre-trends across all specifications.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the null effect on aggregate victimization?
A: The authors estimate that the policy ruled out declines in aggregate victimization larger than 3.3%, indicating SC did not generate meaningful improvements in aggregate public safety. This contradicts the stated mission of immigration enforcement agencies. The findings imply that policies targeting immigrant communities can generate public safety costs through trust erosion that outweigh any deterrence or incapacitation benefits.&lt;/p&gt;
&lt;p&gt;Secure Communities (SC): A federal program launched in 2008 requiring automatic sharing of fingerprints from local jail bookings with the Department of Homeland Security, enabling identification of unauthorized immigrants among local arrestees and triggering ICE detainer requests; the largest expansion of interior immigration enforcement in U.S. history.&lt;/p&gt;
&lt;p&gt;Chilling effect: The mechanism by which immigration enforcement raises the perceived cost of contacting law enforcement for immigrant victims and witnesses — through fear that they, a family member, or neighbor will be detained or deported — thereby reducing willingness to report crimes independently of any change in underlying criminality.&lt;/p&gt;
&lt;p&gt;Victimization rate: The likelihood that an individual is the victim of a crime in a given period, measured via the NCVS independently of whether the crime was reported to police; the paper&amp;rsquo;s primary measure of public safety.&lt;/p&gt;
&lt;p&gt;Reporting rate: The likelihood that a criminal victimization is reported by the victim to the police, measured as a share of all crime incidents; distinct from victimization rate and central to the paper&amp;rsquo;s decomposition of reported crime into its two components.&lt;/p&gt;
&lt;p&gt;Reported crime rate: The joint probability of being victimized and reporting the crime, analogous to measures available in administrative police data such as the FBI UCR; this outcome masks the opposing effects of SC on victimization and reporting.&lt;/p&gt;
&lt;p&gt;Honored detainer: An ICE detainer request that results in a transfer of the arrested individual to ICE custody; the paper&amp;rsquo;s preferred measure of immigration enforcement intensity because it is available both before and after SC activation and is more directly linked to deportation actions than all detainer requests.&lt;/p&gt;
&lt;p&gt;Decomposition of victimization increase: The paper&amp;rsquo;s procedure for quantifying the relative importance of the reporting-channel (reduced probability of apprehension) versus other SC-induced social and economic changes (unemployment, wages, female-headed households, male immigrant share) in explaining the rise in Hispanic victimization.&lt;/p&gt;</description></item><item><title>Consumer Credit and the Incidence of Tariffs: Evidence from the Auto Industry</title><link>https://macropaperwarehouse.com/papers/consumer-credit-and-the-incidence-of-tariffs-evidence-from-the-auto-industry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/consumer-credit-and-the-incidence-of-tariffs-evidence-from-the-auto-industry/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Do import tariffs affect consumer credit terms, and does focusing solely on goods prices understate tariff pass-through to consumers? The paper also asks whether vertical integration &amp;ndash; specifically, the ownership of a captive finance subsidiary &amp;ndash; expands the channels through which manufacturers can pass on cost shocks, and whether tariff incidence falls disproportionately on consumers with less elastic credit demand or in areas with lower credit market competition.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting.&lt;/strong&gt; The Trump administration&amp;rsquo;s 2018 metal tariffs &amp;ndash; a 25 percent tariff on steel and a 10 percent tariff on aluminum &amp;ndash; created a large and largely unanticipated cost shock for US auto manufacturers who are heavy consumers of both metals across their supply chains. Crucially, auto manufacturers own captive finance subsidiaries (e.g., Ford Credit, GM Financial, Honda Finance) that originate consumer auto loans alongside independent noncaptive lenders (banks, credit unions, independent finance companies). Because noncaptive lenders had no direct exposure to the metal tariffs, they serve as a natural control group in a difference-in-differences design.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The primary data source is Regulation AB II, which requires issuers of public auto loan asset-backed securities to report loan-level information monthly to the SEC. The final sample covers 1,973,639 auto loans originated between January 2017 and December 2018 across 14 lenders (8 captive, 6 noncaptive). Vehicle invoice price data come from Regulation AB II; consumer sales price data come from the Texas Department of Motor Vehicles (covering approximately 3.9 million vehicle transactions in 2017-2018). Population credit bureau data from Equifax are used for representativeness checks and HHI construction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Strategy.&lt;/strong&gt; The baseline difference-in-differences compares captive auto loans to otherwise-identical noncaptive auto loans originated in the same state, the same quarter, for the same vehicle make-model-condition, and to borrowers in similar income and credit score bins. Parallel pre-trends tests confirm no economically meaningful differential pre-trends across captive and noncaptive lenders for any outcome variable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings.&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Interest Rate Pass-Through.&lt;/strong&gt; Relative to noncaptive lenders, captive lenders increased average interest rates by 26 basis points following the tariff announcement, representing a 10 percent increase relative to the pretreatment captive mean of 252 basis points. This corresponds to an average present value increase in total loan payments of $179 per loan (discounted at 5 percent for an average $26,914 principal with 66-month maturity). By the fourth quarter of 2018, the dynamic estimate reaches 48 basis points &amp;ndash; nearly double the pooled average &amp;ndash; as metal prices continued to rise. The increase is concentrated among more-exposed captive lenders (those whose manufacturers operate two or more domestic production plants), not less-exposed captive lenders (primarily BMW, Mercedes-Benz, Volkswagen), ruling out captive-specific omitted variables.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Non-Price Loan Terms.&lt;/strong&gt; There is no economically significant change in captive loan amounts, maturities, or loan-to-value ratios following the tariffs. Captive lenders responded to the tariff shock exclusively by raising interest rates, consistent with prior evidence that auto loan demand is less sensitive to interest rates than to non-price terms.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Vehicle Prices.&lt;/strong&gt; Invoice prices for makes with greater domestic production rose by approximately 1.0 percent (relative to makes with less domestic production), and consumer sales prices rose by approximately 0.7 percent ($225 average increase relative to a pretreatment mean of $32,206) for these same makes.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Relative Magnitude of Pass-Through Channels.&lt;/strong&gt; After accounting for estimated spillover effects on noncaptive lenders of 7 basis points, the spillover-adjusted estimate implies captive interest rates rose by 33 basis points on average, corresponding to $227 per loan in present value terms. Interest rate pass-through is estimated to be almost two-thirds as large as vehicle price pass-through, meaning that focusing solely on vehicle prices would underestimate tariff incidence on consumers by approximately 37 percent. The population-weighted average cost increase per vehicle is $146 &amp;ndash; roughly equally split between higher vehicle prices ($74) and higher financing costs ($72).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Intensive vs. Extensive Margin.&lt;/strong&gt; The composition of captive borrowers did not deteriorate following the tariffs: average household incomes of captive borrowers increased slightly (economically small), credit scores were unchanged, and future default rates showed no significant change. This confirms that the interest rate increase reflects tariff pass-through to inframarginal borrowers along the intensive margin, not a shift in borrower composition.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity by Credit Demand Elasticity.&lt;/strong&gt; Pass-through via interest rates was higher for borrowers with lower incomes (33 basis points vs. 20 basis points for higher-income consumers), lower credit scores (36 basis points vs. 15 basis points), and smaller loan amounts (36 basis points vs. 12 basis points). These groups are proxies for less elastic credit demand, consistent with theoretical predictions that cost pass-through is larger where demand is less price sensitive.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity by Market Competition.&lt;/strong&gt; Tariff pass-through via interest rates was higher in states with lower credit market competition (as measured by state-level Herfindahl-Hirschman Index). Consumers in the lowest competition decile experienced an average captive interest rate increase of 41 basis points, compared to 24 basis points for consumers in the highest competition decile. This 17 basis point differential implies that interest rate pass-through was approximately 88 percent as large as vehicle price pass-through in less competitive markets, versus 57 percent in more competitive markets.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-a-captive-finance-subsidiary-and-why-does-it-create-a-novel-channel-for-tariff-pass-through"&gt;Q1. What is a captive finance subsidiary, and why does it create a novel channel for tariff pass-through?&lt;/h3&gt;
&lt;p&gt;A captive finance subsidiary is a wholly owned lending unit of an auto manufacturer (e.g., Ford Credit, GM Financial, American Honda Finance) whose primary purpose is to finance the sale of the manufacturer&amp;rsquo;s vehicles. Because the captive lender and the manufacturing unit share a parent company, a cost shock to the manufacturing side &amp;ndash; such as higher steel and aluminum prices from the tariffs &amp;ndash; can be passed on to consumers not only through higher vehicle prices but also through worse financing terms offered by the captive. Prior studies documented tariff pass-through to goods prices but found limited evidence of pass-through to consumer prices; this paper shows that the bundling of a product with captive financing creates a second, previously unmeasured channel. The institutional structure also facilitates &amp;ldquo;price shrouding&amp;rdquo;: because consumers are less attentive to financing costs than vehicle sticker prices, captive lenders can exploit this inattention to pass on cost shocks along the financing margin.&lt;/p&gt;
&lt;h3 id="q2-why-is-the-auto-loan-market-a-particularly-suitable-setting-for-studying-this-question"&gt;Q2. Why is the auto loan market a particularly suitable setting for studying this question?&lt;/h3&gt;
&lt;p&gt;The auto loan market provides three key advantages. First, both captive lenders (directly exposed to metal tariffs via manufacturing) and noncaptive lenders (with no direct tariff exposure) compete for the same borrowers on the same vehicle purchases, creating a clean within-vehicle, within-period control group. Second, the Regulation AB II data contain vehicle make-model-condition information, allowing the authors to hold vehicle choice fixed and isolate tariff pass-through to loan terms separately from any vehicle switching by consumers. Third, the indirect dealer-intermediated financing process means that consumers typically do not observe the full set of lender bids, weakening their ability to actively arbitrage between captive and noncaptive loan offers.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-regulation-ab-ii-data-and-how-representative-is-it"&gt;Q3. What is the Regulation AB II data, and how representative is it?&lt;/h3&gt;
&lt;p&gt;Under Regulation AB II (effective November 2016), issuers of publicly offered auto loan asset-backed securities must report monthly loan-level data to the SEC, including interest rates, loan amounts, maturities, vehicle characteristics, borrower credit scores and incomes, and loan performance. The final sample covers approximately 8 percent of all open auto loans in the United States and around 30 percent of the total auto loan portfolios of the 14 sampled lenders. Average loan characteristics in the Regulation AB II data closely match population credit bureau data from Equifax, indicating that securitization selection is not a major concern. Average credit scores and incomes are slightly higher in Regulation AB II than in the population, primarily because small banks and credit unions that serve riskier borrowers do not access public securitization markets.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-baseline-empirical-specification-and-what-identifying-variation-does-it-use"&gt;Q4. What is the baseline empirical specification and what identifying variation does it use?&lt;/h3&gt;
&lt;p&gt;The baseline is a difference-in-differences regression comparing captive loans (treated) to noncaptive loans (control) before and after January 2018 (the date of the Department of Commerce&amp;rsquo;s initial tariff recommendation, chosen conservatively). The regression includes lender fixed effects, vehicle make-model-condition x origination quarter fixed effects, state x origination quarter fixed effects, $25,000 income bin x origination quarter fixed effects, and 10-point credit score bin x origination quarter fixed effects. The coefficient of interest is estimated using within-lender variation after netting out common vehicle-level shocks, state-level shocks, and shocks common across income and credit score cells. This granular fixed effect structure ensures that the estimate compares captive and noncaptive loans for exactly the same vehicle, in the same state, in the same quarter, to borrowers with similar incomes and credit scores.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-main-coefficient-estimates-on-interest-rates-and-how-do-they-evolve-dynamically"&gt;Q5. What are the main coefficient estimates on interest rates, and how do they evolve dynamically?&lt;/h3&gt;
&lt;p&gt;In the full sample, the pooled difference-in-differences estimate is 26 basis points (t = 2.75), representing a 10 percent increase relative to the pretreatment captive mean of 252 basis points. Excluding subvented (subsidized) loans, the estimate is 29 basis points (t = 2.85). Dynamically, captive interest rates started rising within one quarter of the treatment date and continued increasing alongside metal prices, reaching a terminal coefficient of 48 basis points in the fourth quarter of 2018 &amp;ndash; nearly double the pooled average. Consistent with the parallel trends assumption, there is no economically significant evidence of differential pre-trends across captive and noncaptive loans in the pretreatment period.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-authors-validate-that-noncaptive-lenders-constitute-a-valid-counterfactual"&gt;Q6. How do the authors validate that noncaptive lenders constitute a valid counterfactual?&lt;/h3&gt;
&lt;p&gt;Four alternative specifications are presented. First, when splitting captive lenders by tariff exposure (more exposed: Ford, GM-AmeriCredit, Honda, Toyota; less exposed: BMW, Mercedes-Benz, Volkswagen), only more-exposed captive lenders show a significant increase in interest rates (30 basis points; t = 3.37), while less-exposed captive lenders show no significant increase (-18 basis points; t = -1.33). This rules out captive-specific correlated omitted variables. Second, the authors add interactions of the treatment indicator with changes in the Fed Funds rate and 1-, 5-, and 10-year Treasury yields; results are unchanged in magnitude, ruling out differential sensitivity to the rising interest rate environment of 2018. Third, using CarMax (a noncaptive that also sells and finances vehicles but does not participate in DealerTrack) as the sole control group yields similar results. Fourth, lender-specific borrowing cost controls do not attenuate the estimates.&lt;/p&gt;
&lt;h3 id="q7-did-captive-lenders-adjust-any-non-price-loan-terms-in-response-to-the-tariffs"&gt;Q7. Did captive lenders adjust any non-price loan terms in response to the tariffs?&lt;/h3&gt;
&lt;p&gt;No. Columns 2-4 of Table 3 document that loan amounts, maturities, and loan-to-value ratios showed no economically significant changes for captive lenders relative to noncaptive lenders following the tariffs. Some coefficient estimates in the full sample are statistically significant but economically small, and they lose significance or flip signs once subvented loans are excluded. The event study plots confirm no meaningful pre-trends and no meaningful post-treatment changes in non-price terms. The authors note that this is consistent with prior evidence that auto loan demand is less sensitive to interest rates than to maturity, making interest rates the optimal margin along which to pass through costs.&lt;/p&gt;
&lt;h3 id="q8-how-do-the-authors-rule-out-that-the-increase-in-captive-interest-rates-reflects-a-change-in-borrower-composition-rather-than-intensive-margin-pass-through"&gt;Q8. How do the authors rule out that the increase in captive interest rates reflects a change in borrower composition rather than intensive-margin pass-through?&lt;/h3&gt;
&lt;p&gt;The authors estimate a separate regression (equation 4) with log household income, log credit score, and future default rate as outcomes. Relative to noncaptive borrowers, captive borrowers experienced a small but positive increase in average household income (Gamma = 0.012, t = 3.25), no significant change in credit scores (Gamma = 0.001, t = 1.13), and no significant change in 12-month or 24-month default rates. The income increase is of the wrong sign and too small in magnitude to explain the observed interest rate increase from a risk-based pricing perspective. Additionally, captive loan origination volumes declined 6.7 percent after the tariffs, inconsistent with a demand surge driving the interest rate increase.&lt;/p&gt;
&lt;h3 id="q9-how-do-the-authors-rule-out-alternative-explanations-including-demand-surges-borrowing-cost-increases-securitization-changes-and-dealer-markup-changes"&gt;Q9. How do the authors rule out alternative explanations including demand surges, borrowing cost increases, securitization changes, and dealer markup changes?&lt;/h3&gt;
&lt;p&gt;For demand surges: vehicle sales volumes showed no noticeable increase following the tariff announcement, and captive loan originations actually declined. For differential borrowing costs: controlling for lender-specific CDS spreads and other borrowing cost measures does not attenuate the main estimate. For securitization changes: combining Regulation AB II and credit bureau data, the authors find no significant change in captive lenders&amp;rsquo; securitization rates, the ratio of securitized to total loan amounts, maturities, or monthly payments. For dealer markup changes: noncaptive loans are also subject to dealer markups, so common changes are absorbed in the DiD; additionally, subvented loans (which dealers cannot mark up) also show higher captive interest rates post-tariff, ruling out differential markup changes. For interest rate sensitivity differentials: controlling for changes in risk-free rates does not alter results. For prepayment responses: 12-month and 24-month prepayment rates show no significant change for captive loans.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-authors-measure-vehicle-price-pass-through-and-what-data-do-they-use"&gt;Q10. How do the authors measure vehicle price pass-through, and what data do they use?&lt;/h3&gt;
&lt;p&gt;To measure invoice price pass-through, the authors use Regulation AB II data (which contains the invoice price for new vehicles) and estimate a regression comparing the change in log invoice prices for makes with a higher proportion of US-assembled vehicles versus those with lower domestic production, controlling for vehicle make-model fixed effects and price bin x quarter fixed effects. Invoice prices rose approximately 1.0 percent for more-exposed makes. For consumer sales price pass-through, the authors use Texas DMV data (1,819,498 new and 2,105,938 used vehicle transactions in 2017-2018) with the same identification strategy. Sales prices rose approximately 0.7 percent ($225 average increase) for more-exposed makes. Both effects are robust to defining exposure at either the make level or the make-model level.&lt;/p&gt;
&lt;h3 id="q11-how-is-the-overall-pass-through-rate-decomposed-between-the-interest-rate-and-vehicle-price-channels"&gt;Q11. How is the overall pass-through rate decomposed between the interest rate and vehicle price channels?&lt;/h3&gt;
&lt;p&gt;The authors define total tariff pass-through as the sum of interest rate pass-through (change in aggregate captive financing costs divided by aggregate production cost increase) and vehicle price pass-through (change in aggregate new vehicle sales revenue divided by aggregate production cost increase). Taking the ratio of these two components allows them to estimate the relative importance of each channel without needing to directly measure production costs. With a captive loan penetration rate (M) of 0.59, a per-loan present value financing cost increase of $179 (unadjusted) or $227 (adjusted for 7 basis point spillover effect on noncaptives), and a $225 average vehicle price increase, the spillover-adjusted estimate implies interest rate pass-through is almost two-thirds as large as vehicle price pass-through. Focusing solely on vehicle prices would underestimate tariff incidence on consumers by approximately 37 percent. The population-weighted average total cost increase is $146 per vehicle, roughly equally split between vehicle prices ($74) and financing costs ($72).&lt;/p&gt;
&lt;h3 id="q12-how-large-is-the-estimated-aggregate-impact-of-the-tariffs-on-consumer-financing-costs"&gt;Q12. How large is the estimated aggregate impact of the tariffs on consumer financing costs?&lt;/h3&gt;
&lt;p&gt;Using population data of approximately 50 million vehicles sold annually in the United States and a population-weighted average financing cost increase of $72 per vehicle, the authors estimate that the tariffs resulted in approximately $3.6 billion (= 50,000,000 x $72) in additional present value financing costs each year. For reference, Flaaen, Hortacsu, and Tintelnot (2020) estimated that the 2018 tariffs on washing machines led to $1.5 billion in additional annual consumer costs.&lt;/p&gt;
&lt;h3 id="q13-which-borrowers-bore-a-disproportionate-share-of-the-interest-rate-pass-through-and-by-how-much"&gt;Q13. Which borrowers bore a disproportionate share of the interest rate pass-through, and by how much?&lt;/h3&gt;
&lt;p&gt;The triple-differences results show monotonically higher pass-through for borrowers with less elastic credit demand. Lower-income borrowers (below median) experienced an average captive interest rate increase of 33 basis points versus 20 basis points for higher-income borrowers. Lower-credit-score borrowers experienced an increase of 36 basis points versus 15 basis points for higher-credit-score borrowers. Borrowers with smaller loan amounts (below median) experienced an increase of 36 basis points versus 12 basis points for larger loan amounts. Within income quartiles, consumers in the lowest income quartile experienced a 37 basis point increase compared to 17 basis points in the highest quartile. These patterns are not driven by changes in borrower composition, as default rates show no significant change across any of these subgroups.&lt;/p&gt;
&lt;h3 id="q14-how-does-credit-market-competition-affect-tariff-pass-through-via-interest-rates"&gt;Q14. How does credit market competition affect tariff pass-through via interest rates?&lt;/h3&gt;
&lt;p&gt;States with lower credit market competition (higher Herfindahl-Hirschman Index, constructed from pretreatment lender market shares) experienced higher interest rate pass-through. Comparing above- versus below-median HHI states, the difference is 5 basis points (28 vs. 23 basis points), statistically significant at the 10 percent level. When restricting to the tails of the competition distribution, the difference is substantially larger: consumers in the lowest competition decile experienced an average increase of 41 basis points versus 24 basis points for consumers in the highest competition decile &amp;ndash; a 17 basis point differential. This implies interest rate pass-through was 88 percent as large as vehicle price pass-through in less competitive markets versus 57 percent in more competitive markets, consistent with theoretical predictions that firm-specific cost shocks generate higher pass-through when competition is weaker.&lt;/p&gt;
&lt;h3 id="q15-why-do-captive-lenders-spread-interest-rate-increases-broadly-across-vehicle-types-rather-than-targeting-directly-tariff-exposed-new-vehicle-models"&gt;Q15. Why do captive lenders spread interest rate increases broadly across vehicle types rather than targeting directly tariff-exposed new vehicle models?&lt;/h3&gt;
&lt;p&gt;The authors find that captive interest rates increased for both new and used vehicles, and that within more-exposed captive lenders, interest rate increases were not concentrated in domestically produced vehicle models. This is consistent with the hypothesis that firms spread cost shocks across multiple goods and business segments (as documented in the industrial organization literature for multiproduct firms). The authors argue this occurs because vehicles of different makes and models are substitutes for each other (making vehicle-specific price increases costlier in terms of demand loss), whereas auto loans are complementary to vehicle purchases and are offered as an add-on to the sales transaction. This bundled structure, combined with consumer inattention to financing terms, makes it optimal to spread the cost shock across the loan book rather than concentrating it in specific vehicle models.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Captive Finance Subsidiary&lt;/strong&gt;: A wholly owned lending unit of a manufacturer (e.g., Ford Credit, GM Financial) whose primary purpose is to originate loans and leases to finance the sale of the manufacturer&amp;rsquo;s own products. Unlike independent noncaptive lenders, captive lenders are vertically integrated with the manufacturing unit and can, in principle, use financing terms as an additional margin to pass through manufacturing-side cost shocks to consumers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tariff Pass-Through (Interest Rate Channel)&lt;/strong&gt;: The extent to which an input cost increase caused by an import tariff is transmitted to consumers via higher interest rates charged by captive lenders, rather than (or in addition to) higher goods prices. The paper defines interest rate pass-through as the ratio of the aggregate present value increase in captive financing costs to the aggregate increase in manufacturing production costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive vs. Extensive Lending Margin&lt;/strong&gt;: The distinction between raising loan prices charged to existing (inframarginal) borrowers (intensive margin) versus changing the pool of borrowers served or lending standards (extensive margin). The paper argues that the observed increase in captive interest rates reflects intensive-margin pass-through because borrower incomes, credit scores, and future default rates did not change significantly after the tariffs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Price Shrouding&lt;/strong&gt;: The practice of making price increases less salient to consumers by embedding them in a less-scrutinized component of a bundled transaction. In the auto market, because consumers are documented to be less sensitive to increases in financing costs than to vehicle sticker prices, captive lenders can pass on cost shocks through interest rates with less demand response than if they raised vehicle prices by an equivalent amount. The paper treats this as a key mechanism enabling the financing pass-through channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Subvented (Subsidized) Loan&lt;/strong&gt;: A promotional auto loan offered at a below-market interest rate, often tied to specific vehicle models or sales events (e.g., &amp;ldquo;1.99 percent APR for well-qualified borrowers&amp;rdquo;). Subvented loans are typically fixed by the manufacturer and cannot be marked up by dealers. The paper uses the subsample of non-subvented loans as a robustness check and to isolate tariff pass-through from seasonal variation in promotional financing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Captive Loan Penetration Rate (M)&lt;/strong&gt;: The ratio of captive auto loans originated to new vehicles produced and sold, used in the paper&amp;rsquo;s decomposition of total tariff pass-through into the interest rate and vehicle price channels. Estimated at approximately 0.59 from population data, this parameter determines how the aggregate present value financing cost increase scales relative to the aggregate vehicle sales price increase when computing the relative importance of the two pass-through channels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Herfindahl-Hirschman Index (HHI) as Market Competition Measure&lt;/strong&gt;: The paper constructs state-level HHIs based on pretreatment lender market shares in each state using population credit bureau data, as an inverse measure of credit market competition. Local (direct) auto lending markets exhibit meaningful geographic variation in HHI, in contrast to the largely national scope of indirect (dealer-arranged) lending. The paper uses this variation to test whether pass-through is higher in less competitive credit markets, consistent with theoretical predictions for firm-specific cost shocks.&lt;/p&gt;</description></item><item><title>Consumer durables and monetary policy according to HANK</title><link>https://macropaperwarehouse.com/papers/consumer-durables-and-monetary-policy-according-to-hank/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/consumer-durables-and-monetary-policy-according-to-hank/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;h3 id="research-question"&gt;Research Question&lt;/h3&gt;
&lt;p&gt;Consumer durables account for a disproportionately large share of household expenditure fluctuations despite their small share of total private consumption. Two stylized facts motivate the paper: (1) durable expenditure is far more interest-rate sensitive than nondurable expenditure following monetary policy shocks, and (2) durable and nondurable expenditures comove positively and persistently—both reaching trough in the same quarter. Standard two-sector New Keynesian models struggle to generate this positive conditional comovement because asymmetric sectoral price rigidity induces large relative-price movements that push the two sectors in opposite directions. This paper asks what model features are necessary and sufficient to reproduce both the sectoral comovement pattern and the hump-shaped aggregate dynamics observed in the data, and how the answer changes across households sorted by liquid asset holdings.&lt;/p&gt;
&lt;h3 id="data-and-methodology"&gt;Data and Methodology&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Empirical identification.&lt;/strong&gt; The authors employ a local projection instrumental variables (LP-IV) strategy using Romer-Romer monetary policy shocks updated by Wieland and Yang (2020), over the sample 1969:Q1–2007:Q3. Impulse response functions (IRFs) are normalized to a cumulative 100 basis-point increase in the Federal Funds Rate over five years. Household-level evidence is drawn from the Consumer Expenditure Survey (CEX) and the Survey of Consumer Finances (SCF); households are classified as liquidity-constrained if liquid assets are below $1,000.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The authors develop a two-sector Heterogeneous Agent New Keynesian (HANK) model in which households maximize utility over nondurable consumption and a durable stock (Cobb-Douglas aggregation), face convex adjustment costs on durable purchases, and update expectations infrequently in the Mankiw-Reis sense (probability of not updating: Xi = 0.918 per period). The general equilibrium version features asymmetric Rotemberg price stickiness (Calvo probability 0.671 for nondurables, 0.797 for durables), nominal wage stickiness (Calvo 0.802), and a Taylor rule with inflation coefficient 1.105, output coefficient 1.440, and smoothing 0.988.&lt;/p&gt;
&lt;h3 id="main-findings-and-quantitative-magnitudes"&gt;Main Findings and Quantitative Magnitudes&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Sectoral magnitude gap.&lt;/strong&gt; At trough (approximately 8 quarters after the shock), the durable expenditure response to monetary tightening is an order of magnitude larger than the nondurable response—a fact the calibrated HANK model is designed to match.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Positive comovement.&lt;/strong&gt; Both durable and nondurable expenditures contract and reach trough in the same quarter, contradicting TANK models (Monacelli 2009) in which savers shift portfolios toward durables and generate negative comovement for that group.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Relative-price dynamics.&lt;/strong&gt; The relative price of durables rises following monetary tightening (nondurables deflate more), but the rise is modest and cannot overturn the positive comovement result.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Role of the direct interest-rate effect.&lt;/strong&gt; Across liquid-asset groups, the direct effect accounts for 73–87% of the cumulated durable expenditure response and 37–91% of the cumulated nondurable expenditure response. This direct channel—operating through intertemporal substitution—is quantitatively first-order for durables in a way it is not in standard single-sector HANK models where income effects dominate.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Role of sticky information.&lt;/strong&gt; A full-information HANK variant produces a counterfactually high durable elasticity (35.24 times the baseline) and no hump-shaped dynamics. Infrequent information updating (Xi = 0.918) is essential to match the hump-shaped propagation of both sectoral and aggregate expenditures.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Income effects and fiscal policy.&lt;/strong&gt; For a fiscal subsidy specifically targeting durable purchases, intertemporal substitution incentives generate a large shift toward durables and, without income effects, a counterfactual crowding-out of nondurable spending. Income effects are essential to protect nondurable spending, and the aggregate consumption effect of such a policy is at best modest—consistent with Mian and Sufi&amp;rsquo;s (2012) evidence on cash-for-clunkers.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="scope-conditions"&gt;Scope Conditions&lt;/h3&gt;
&lt;p&gt;All empirical results are conditional on the LP-IV sample 1969:Q1–2007:Q3 and Romer-Romer shocks as instrumented by Wieland-Yang. The household-level comovement result is established for both liquidity-constrained (liquid assets below $1,000) and unconstrained savers using CEX/SCF data. Model quantitative results are specific to the calibration targeting moments from Fagereng et al. (2021) marginal propensities and BEA depreciation data (delta = 0.054).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-empirical-puzzle-the-paper-addresses-and-why-do-standard-models-fail"&gt;Q1. What is the core empirical puzzle the paper addresses, and why do standard models fail?&lt;/h3&gt;
&lt;p&gt;Standard two-sector New Keynesian models predict that asymmetric sectoral price stickiness generates large relative-price movements between durables and nondurables following a monetary shock. These relative-price shifts tend to produce negative conditional comovement—when durables contract, nondurables expand—contradicting the data. The authors document that both categories exhibit positive and persistent comovement, both reaching their trough at approximately 8 quarters, which standard models cannot replicate.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-key-empirical-facts-established-via-lp-iv"&gt;Q2. What are the key empirical facts established via LP-IV?&lt;/h3&gt;
&lt;p&gt;Using Romer-Romer shocks over 1969:Q1–2007:Q3, normalized to a cumulative 100bp Federal Funds Rate increase, the authors find: (1) aggregate expenditure follows a hump-shaped contraction with trough at roughly 8 quarters; (2) the durable expenditure response is an order of magnitude larger than the nondurable response at trough; (3) both categories reach their trough in the same quarter; and (4) the relative price of durables rises modestly after monetary tightening (nondurables deflate more), but not enough to reverse comovement.&lt;/p&gt;
&lt;h3 id="q3-how-is-the-partial-equilibrium-model-calibrated-and-which-moments-does-it-target"&gt;Q3. How is the partial equilibrium model calibrated, and which moments does it target?&lt;/h3&gt;
&lt;p&gt;Key calibrated parameters include CRRA sigma = 2.640, Cobb-Douglas weight on nondurables theta = 0.607 (implying durable expenditure share 0.193), adjustment cost alpha = 8.299, information stickiness Xi = 0.918, depreciation rate delta = 0.054, steady-state real rate r = 0.03/4, discount factor beta = 0.915 (matching a 30% share of liquidity-constrained households with liquid assets-to-income ratio of 0.26), and borrowing wedge kappa = 0.05. Moments matched include quarterly MPC on nondurables (22.94%), quarterly MPX on durables (24.15%), interest-rate elasticity of durable expenditure (3.35, within the empirical range of 1.1–5.0), price elasticity of durable demand (29.59), and durable stock skewness relative to nondurable consumption (0.695, consistent with Bertola et al. 2005).&lt;/p&gt;
&lt;h3 id="q4-how-does-the-paper-decompose-monetary-policy-transmission"&gt;Q4. How does the paper decompose monetary policy transmission?&lt;/h3&gt;
&lt;p&gt;The paper decomposes transmission into three channels: (1) the direct effect of real interest rate changes, which operates through intertemporal substitution and accounts for the quantitatively largest share of the durable response; (2) the relative-price effect, which is modest and redistributive but cannot overturn positive comovement; and (3) pure income effects, which are key for persistence of the nondurable response but not for the sign of comovement.&lt;/p&gt;
&lt;h3 id="q5-what-do-counterfactual-models-reveal-about-the-role-of-each-model-ingredient"&gt;Q5. What do counterfactual models reveal about the role of each model ingredient?&lt;/h3&gt;
&lt;p&gt;A sticky-information RANK produces positive comovement but the dynamics are front-loaded and less inertial than in the data. A sticky-information TANK delivers results similar to RANK—income effects do not qualitatively change the story. A full-information HANK produces a counterfactually high durable interest-rate elasticity (35.24 times the baseline) and no hump-shaped dynamics, demonstrating that sticky information is the ingredient generating realistic propagation, not heterogeneity per se.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-household-level-evidence-from-cex-and-scf-show-about-comovement-across-the-wealth-distribution"&gt;Q6. What does the household-level evidence from CEX and SCF show about comovement across the wealth distribution?&lt;/h3&gt;
&lt;p&gt;Classifying households as liquidity-constrained if liquid assets are below $1,000, the LP-IV estimates show positive comovement between durables and nondurables for both constrained and unconstrained savers. This contradicts TANK models (Monacelli 2009), in which savers shift portfolios toward durables following a monetary shock, generating negative comovement for the saver group. After controlling for income and relative prices, the direct interest-rate effect operates uniformly across financial status groups.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-direct-effect-vary-across-liquid-asset-groups-quantitatively"&gt;Q7. How does the direct effect vary across liquid asset groups quantitatively?&lt;/h3&gt;
&lt;p&gt;Decomposing across four liquid asset groups (below $1k, $1k–$10k, $10k–$20k, above $20k), the direct effect accounts for 73–87% of the cumulated durable expenditure response and 37–91% of the cumulated nondurable expenditure response. Income effects are more important for nondurable spending prolongation among liquidity-constrained households, but the direct channel dominates durable expenditure for all groups.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-general-equilibrium-two-sector-hank-model-differ-from-the-partial-equilibrium-setup"&gt;Q8. How does the general equilibrium two-sector HANK model differ from the partial equilibrium setup?&lt;/h3&gt;
&lt;p&gt;The GE model adds asymmetric sectoral price stickiness (Calvo probabilities 0.671 for nondurables and 0.797 for durables), nominal wage stickiness (Calvo 0.802), a Taylor rule (inflation coefficient 1.105, output coefficient 1.440, smoothing 0.988), and fiscal lump-sum taxes responding to debt (coefficient 0.191). These features generate the relative-price dynamics observed in the data while preserving the positive comovement result.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-fiscal-policy-application-reveal-about-the-role-of-income-effects"&gt;Q9. What does the fiscal policy application reveal about the role of income effects?&lt;/h3&gt;
&lt;p&gt;A fiscal subsidy targeting durable purchases generates a much larger shift in the relative price of durables than monetary policy does. Without income effects, intertemporal substitution dominates and nondurable spending falls—a counterfactual result inconsistent with the data. With income effects present, nondurable spending is protected. The aggregate consumption effect of such a durable-targeted fiscal policy is at best modest, consistent with Mian and Sufi&amp;rsquo;s (2012) evidence from the cash-for-clunkers program.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-broader-implication-for-the-literature-on-hank-versus-rank-transmission"&gt;Q10. What is the broader implication for the literature on HANK versus RANK transmission?&lt;/h3&gt;
&lt;p&gt;In standard single-sector HANK models, income effects (the indirect channel) typically dominate monetary transmission. The presence of consumer durables restores a quantitatively important role for the direct interest-rate channel, which operates through intertemporal substitution in durable purchases. This rebalances the direct-versus-indirect decomposition relative to the conventional HANK wisdom and shows that the durable goods sector is essential to understanding the full transmission mechanism.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Sectoral comovement (conditional on monetary policy shocks)&lt;/strong&gt;
The empirical regularity that durable and nondurable expenditures both contract following monetary tightening and reach their respective troughs in the same quarter. In this paper, comovement is defined conditional on identified monetary policy shocks (LP-IV with Romer-Romer instruments), not unconditionally. Standard two-sector NK models predict negative conditional comovement due to relative-price effects; replicating positive comovement is the central discipline imposed on the model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Direct effect (of real interest rate changes)&lt;/strong&gt;
The component of monetary transmission that operates through the intertemporal substitution incentive induced by changes in the real interest rate, holding income and relative prices fixed. Distinct from the income effect (indirect channel) and the relative-price effect. In this paper&amp;rsquo;s decomposition, the direct effect accounts for 73–87% of the cumulated durable expenditure response across liquid-asset groups.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sticky information (Mankiw-Reis)&lt;/strong&gt;
Households update their information sets infrequently, with probability (1 - Xi) per period; Xi = 0.918 means only about 8.2% of households update each quarter. This mechanism is essential in the model for generating the hump-shaped, inertial impulse response dynamics observed in the data. Without it (full-information HANK), the durable elasticity is counterfactually large (35.24 times baseline) and dynamics are front-loaded.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MPX (Marginal Propensity to Expend on durables)&lt;/strong&gt;
Analogous to the MPC for nondurables, the MPX measures the additional durable expenditure flow induced by an income windfall. Calibrated to 24.15% quarterly, matching estimates from Fagereng et al. (2021). Distinct from the MPC because durable purchases represent investment in a stock, not immediate consumption flow.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Liquidity-constrained households&lt;/strong&gt;
Households with liquid assets below $1,000, identified in the CEX and SCF. In the model, the 30% share of such households is targeted by the discount factor (beta = 0.915) and the borrowing wedge (kappa = 0.05). The paper&amp;rsquo;s key finding is that positive comovement holds for both constrained and unconstrained households, contradicting TANK predictions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;HANK (Heterogeneous Agent New Keynesian model)&lt;/strong&gt;
A New Keynesian general equilibrium model in which households are heterogeneous in their liquid asset holdings (and thus face binding borrowing constraints), so that the distribution of assets matters for aggregate dynamics. Distinguished from RANK (Representative Agent NK) and TANK (Two-Agent NK, which approximates heterogeneity with one unconstrained and one hand-to-mouth agent). In this paper, HANK is extended to a two-sector setting with durables and nondurables.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Convex adjustment costs on durable purchases&lt;/strong&gt;
A cost of adjusting the durable stock that is convex in the size of the adjustment (calibrated parameter alpha = 8.299). This smooths the durable expenditure response and prevents counterfactually sharp jumps in durable purchases following interest rate changes, contributing to realistic propagation dynamics alongside sticky information.&lt;/p&gt;</description></item><item><title>Contract Terms, Employment Shocks, and Default in Credit Cards</title><link>https://macropaperwarehouse.com/papers/contract-terms-employment-shocks-and-default-in-credit-cards/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/contract-terms-employment-shocks-and-default-in-credit-cards/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks two related questions bearing on financial inclusion policy in developing countries: (1) How effective are credit card contract term changes — specifically interest rate reductions and minimum payment increases — in limiting default among new borrowers? (2) How large is the effect of formal-sector job loss on default relative to these contract term interventions, and can the difference in magnitudes be explained by differential cash flow impacts?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and Data&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The study is set in Mexico during 2007–2009 and exploits a large nationwide stratified randomized controlled trial implemented by a major commercial bank (&amp;ldquo;Bank A&amp;rdquo;) on its financial-inclusion credit card — a product that accounted for approximately 15% of all first-time formal-sector loans in Mexico as of 2010. The study card was targeted at borrowers with limited or no formal credit history (the bank&amp;rsquo;s &amp;ldquo;C, C- and D&amp;rdquo; customer segments); 47% of the experimental sample held it as their first formal loan product. A sample of 144,000 pre-existing cardholders was stratified into nine cells based on bank tenure (6–11 months, 12–23 months, 24+ months) and past repayment behavior, then randomly allocated to eight treatment arms combining two minimum payment levels (5% or 10% of the outstanding balance) and four annual interest rates (15%, 25%, 35%, 45%), for 26 months (March 2007 to May 2009). The study sample is representative of the bank&amp;rsquo;s national portfolio of approximately 1.3 million study card customers. Card-level data run through December 2014 — five years after the experiment ended — allowing examination of both short- and long-run effects. The experimental sample is matched to Mexico&amp;rsquo;s Social Security database (IMSS), providing monthly formal employment histories from January 2004 to December 2012 for 59% of the sample; and to credit bureau data, allowing observation of defaults across all formal financial institutions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings with Quantitative Magnitudes&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Result 1 — Interest rate effects are modest in aggregate.&lt;/em&gt; A 30 percentage point (pp) decrease in the annual interest rate (from 45% to 15%, a 67% reduction relative to the baseline rate) decreased cumulative default by 2.5 pp over the 26-month experiment, for a default elasticity of +0.20. Over the same 18-month horizon used for unemployment comparisons, the implied effect is 1.03 pp. These magnitudes are substantially smaller than predictions elicited from Mexican central bank regulators (mean predicted decrease: 8.6 pp) and from participants on the Social Science Prediction Platform (mean predicted decrease: 5 pp). Default continued to decline in the lower-rate arm for approximately three years after the experiment ended, reaching −1 pp by March 2012, after which effects became statistically indistinguishable from zero.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Result 2 — No effect on the newest borrowers.&lt;/em&gt; For the newest borrowers (those with 6–11 months of tenure when the experiment began — the group with a 36% cumulative default rate over 26 months versus 18% for those with 24+ months of tenure), the interest rate reduction has no effect on default over the 26-month period, with point estimates consistently small and statistically indistinguishable from zero. This is in contrast to older borrowers, who are meaningfully responsive.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Result 3 — Minimum payment increases increase short-run default but reduce long-run default.&lt;/em&gt; Doubling the minimum payment from 5% to 10% of outstanding balance increased cumulative default by 0.8 pp by the end of the experiment (26-month elasticity: +0.04; p = 0.016), driven primarily by defaults occurring within the first year. The short-run increase is concentrated among the most liquidity-constrained borrowers — those with the highest baseline debt utilization and those in the minimum-payer stratum (baseline debt utilization rate of 85%). After the experiment ended and all arms were returned to the same 4% minimum payment, the previously higher-minimum-payment arm exhibited persistently lower default, reaching a 1 pp decline by the end of the sample (p = 0.054 at end of study period), relative to a base default rate of 41% at that point.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Result 4 — Job displacement effects are seven times larger than contract term effects.&lt;/em&gt; Formal-sector job displacement (identified using mass layoff events at firms with 50+ employees, defined as year-on-year employment contractions exceeding 30% of prior-year average employment) increased cumulative default by 4.8 pp after 12 months and 7.6 pp after 18 months. This is seven times larger than the effect of a 30 pp interest rate decrease (1.03 pp over 18 months) and nine times larger than the effect of doubling minimum payments (0.8 pp). Formal job loss alone can explain approximately 14% of total study card default during the experiment (calculation: 19.8% of formally employed study card borrowers lose their job at least once in the first 18 months; multiplied by the 7.6 pp default increase per spell, this yields 1.5 pp of the 10.8% base default rate at 18 months). Results are corroborated using a nationally representative matched credit bureau–IMSS sample of 600,339 borrowers, which yields 8,723 mass layoff events and similar estimates.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Per-peso normalization.&lt;/em&gt; A back-of-the-envelope calculation normalizes all three shocks by their respective cash flow impacts. The interest rate decrease reduces cumulative required minimum payments due by 2,917 MXN pesos over 18 months; the minimum payment doubling increases them by 1,325 MXN pesos; formal job loss reduces total labor earnings by an estimated 21,328 MXN pesos (adjusting formal-sector earnings losses of 77,555 MXN pesos downward by 72.5% to reflect that 82% of workers who lose formal employment transition to informal employment in the following quarter, with total earnings falling only 27.5%). The per-peso default effects are: 0.36 pp per 1,000 MXN pesos for the interest rate intervention; 0.51 pp for the minimum payment intervention; and 0.36 pp for job displacement. The null hypothesis that all three per-peso effects are equal cannot be rejected (p = 0.78).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Interpretation&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors present a simple two-period optimizing model emphasizing the role of previously accumulated debt and liquidity constraints. The model generates four testable predictions consistent with the data: (1) lower interest rates decrease default via reduced debt burden; (2) higher minimum payments increase short-run default by tightening liquidity constraints; (3) &amp;ldquo;surprise&amp;rdquo; minimum payment increases (where borrowers anticipated they would continue) reduce post-experiment default via debt reduction; (4) negative income shocks (modeled as first-order stochastic dominance deterioration in period-2 income) increase default. The per-peso normalization supports the interpretation that cash flow impacts — not differential per-peso susceptibility to shocks — drive the relative magnitudes of the three effects.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-is-the-interest-rate-elasticity-of-default-020-so-much-lower-than-prior-estimates-in-the-literature"&gt;Q1. Why is the interest rate elasticity of default (0.20) so much lower than prior estimates in the literature?&lt;/h3&gt;
&lt;p&gt;A: The paper contrasts its 26-month elasticity of +0.20 with estimates from Karlan and Zinman (2019) (1.8) and Adams et al. (2009) (2.2), and notes it falls in the same range as Karlan and Zinman (2009) (0.27) and DeFusco et al. (2021) (0.01). The paper proposes that variation in borrower tenure may partly explain cross-study differences, as default elasticities appear to be increasing in bank tenure. The newest borrowers — the most policy-relevant subgroup — show zero elasticity, pulling the overall estimate down. The paper also argues that in this context, interest-rate-driven moral hazard (all channels: debt burden, concurrent, and dynamic) is collectively small.&lt;/p&gt;
&lt;h3 id="q2-what-mechanism-explains-why-newer-borrowers-are-entirely-unresponsive-to-interest-rate-changes"&gt;Q2. What mechanism explains why newer borrowers are entirely unresponsive to interest rate changes?&lt;/h3&gt;
&lt;p&gt;A: The paper hypothesizes that newer borrowers place a higher continuation value on the card (captured by parameter v in the model) because they have fewer formal credit alternatives; at baseline, only 64% of the 6–11 month stratum held a card with another bank versus 78% of the 24+ month stratum. A higher continuation value implies more muted responses to interest rate changes (formally derived in Appendix E.3). Newer borrowers also respond more strongly to credit limit increases, consistent with tighter liquidity constraints. A regression controlling for age, gender, baseline card ownership, debt utilization, labor force attachment, and earnings cannot explain away the differential treatment effect between new and old borrowers (differential remains significant at p = 0.05), suggesting the tenure gradient in responsiveness is not simply a composition effect.&lt;/p&gt;
&lt;h3 id="q3-why-does-increasing-minimum-payments-raise-short-run-default-but-reduce-long-run-default"&gt;Q3. Why does increasing minimum payments raise short-run default but reduce long-run default?&lt;/h3&gt;
&lt;p&gt;A: In the short run, the doubling of minimum payments tightens liquidity constraints for already-constrained borrowers. The increase in default is concentrated among borrowers in the highest baseline debt-utilization tercile and among minimum-payers (baseline debt utilization of 85%), and is preceded by a sharp rise in delinquencies in months 3–5 (which trigger 350 MXN peso fees per occurrence, further worsening the repayment burden). In the long run, borrowers who anticipated continuing higher minimum payments (the experiment ended without advance notice, so borrowers expected the new terms to persist) chose lower debt levels during the experiment. Since all arms were returned to the same low minimum payment when the experiment ended, the lower-debt borrowers in the higher-minimum-payment arm were better positioned to weather subsequent shocks, producing the 1 pp post-experiment decline in default. The hypothesis that this is driven by habit formation in payment behavior is ruled out by the absence of any effect of past higher minimum payments on post-experimental payment levels.&lt;/p&gt;
&lt;h3 id="q4-how-is-the-mass-layoff-identification-strategy-designed-and-validated"&gt;Q4. How is the mass-layoff identification strategy designed and validated?&lt;/h3&gt;
&lt;p&gt;A: The paper uses the universe of IMSS formal employment records to define a mass layoff at a firm (50+ employees) as the first month in which year-on-year employment declines by more than 30% of average employment in the prior 12 months. An individual is &amp;ldquo;displaced&amp;rdquo; if they lost their job in the same quarter as their employer&amp;rsquo;s mass layoff event. The identification assumption is that, conditional on individual and time fixed effects, the exact timing of the mass layoff is uncorrelated with workers&amp;rsquo; potential default outcomes. This is supported by: (1) mass layoffs occurring in every period, making coincidence with credit market shocks unlikely; (2) time fixed effects absorbing common trends; and (3) the absence of statistically distinguishable pre-trends in default between displaced and non-displaced workers. The paper implements both standard two-way fixed effects and the staggered DiD estimator of de Chaisemartin and D&amp;rsquo;Haultfoeuille (2024), which remains valid under heterogeneous and dynamic effects, and the results are similar across methods.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-paper-account-for-informal-employment-when-estimating-the-cash-flow-impact-of-job-loss"&gt;Q5. How does the paper account for informal employment when estimating the cash flow impact of job loss?&lt;/h3&gt;
&lt;p&gt;A: Formal-sector earnings losses over 18 months post-displacement are estimated at 77,555 MXN pesos using IMSS wage data in an event-study design paralleling the default equation. However, since more than 4/5 of workers who lose formal employment are informally employed in the following quarter (based on Mexico&amp;rsquo;s ENOE labor force survey panel), and total labor earnings fall by only an estimated 27.5% over the three post-displacement quarters, the paper scales the formal earnings loss down to 21,328 MXN pesos (≈ 0.275 × 77,555). This brings the estimated earnings loss closer to prior developed-country estimates of displacement costs and is treated as a lower bound relative to the raw formal-earnings loss figure.&lt;/p&gt;
&lt;h3 id="q6-does-the-cost-of-default-deter-borrowers-from-defaulting-and-what-is-the-cost"&gt;Q6. Does the cost of default deter borrowers from defaulting, and what is the cost?&lt;/h3&gt;
&lt;p&gt;A: The paper argues that defaulters face substantial consequences. Using an instrumental variables strategy (treatment assignment as instrument for default on the study card), the probability of having a new loan one year after default is estimated to be 65 pp lower relative to the non-default counterfactual (p = 0.03). A selection-on-observables approach also shows that study card default is associated with the complete absence of any subsequent credit card for at least four years. These costs should provide strong incentives to remain current, making the high observed default rates primarily attributable to cash flow shocks rather than strategic default. The value of formal credit is further confirmed by the finding that a 100 MXN peso increase in the study card&amp;rsquo;s credit limit translates into 32 MXN pesos of additional debt (instrumental variable estimates are more than twice as large as OLS), and by the comparison of informal loan terms (annual rates averaging 291%, loan amounts of 3,658 MXN pesos, durations of 0.52 years) with formal loan terms (94 pp lower rates, 9,842 MXN peso average amounts, 1.07 year durations).&lt;/p&gt;
&lt;h3 id="q7-are-the-default-treatment-effects-different-across-the-interest-rate-and-minimum-payment-interventions-or-do-they-interact"&gt;Q7. Are the default treatment effects different across the interest rate and minimum payment interventions, or do they interact?&lt;/h3&gt;
&lt;p&gt;A: The paper tests for and cannot reject separability between the two interventions at standard significance levels. At the end of the experiment (May 2009), the p-value for the null that the minimum payment effect is constant across interest rate arms is 0.44; five years later it is 0.65. The null that the interest rate effect is constant across both minimum payment arms yields p = 0.08 at end of experiment and p = 0.411 five years later. The fully saturated specification yields results indistinguishable from the parsimonious linear-separable specification.&lt;/p&gt;
&lt;h3 id="q8-are-there-spillover-effects-from-the-contract-term-changes-onto-other-loans-held-by-study-participants"&gt;Q8. Are there spillover effects from the contract term changes onto other loans held by study participants?&lt;/h3&gt;
&lt;p&gt;A: No spillover effects on default on other loans are found, either during the experiment or after it ended, based on credit bureau data covering all formal-sector loans held by the experimental sample. There is also no evidence of crowd-out or crowd-in from other lenders in terms of new loans or loan closures. The only minor exception is a small decrease in default (3%, or approximately 2 pp out of a 61 pp base) on other Bank A loans in the high minimum payment arm.&lt;/p&gt;
&lt;h3 id="q9-why-does-the-effect-of-unemployment-on-default-exceed-the-models-predictions-from-cash-flow-alone"&gt;Q9. Why does the effect of unemployment on default exceed the model&amp;rsquo;s predictions from cash flow alone?&lt;/h3&gt;
&lt;p&gt;A: The paper&amp;rsquo;s back-of-the-envelope normalization finds that the per-peso effects of all three shocks on default are statistically indistinguishable (p = 0.78 for the null that all three λ estimates are equal), with point estimates of λ_IR = 0.36, λ_MP = 0.51, and λ_U = 0.36 pp per 1,000 MXN pesos. This implies that job loss does not have a larger per-peso effect on default than contract term changes; the larger absolute effect of displacement arises entirely from its larger cash flow impact. Additional consequences of job loss beyond cash flow (health, mental health) do not appear to generate additional default beyond what can be attributed to income loss.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-experimental-results-compare-to-what-experts-predicted"&gt;Q10. How do the experimental results compare to what experts predicted?&lt;/h3&gt;
&lt;p&gt;A: Expert predictions were systematically too large. Mexican central bank regulators predicted a mean decrease of 8.6 pp from a 30 pp interest rate reduction at the 18-month horizon, versus the actual estimated effect of 1.03 pp. Social Science Prediction Platform respondents predicted a mean decrease of 5 pp. For minimum payments, regulators on average predicted a 0.4 pp decrease in default from doubling the minimum payment, whereas the actual effect was a 0.8 pp increase. Three-quarters of SSPP respondents correctly predicted the sign of the minimum payment effect (an increase in default), but the predicted mean increase was 6.4 pp, far larger than the estimated 0.8 pp.&lt;/p&gt;
&lt;h3 id="q11-do-the-job-displacement-results-generalize-beyond-the-experimental-sample"&gt;Q11. Do the job displacement results generalize beyond the experimental sample?&lt;/h3&gt;
&lt;p&gt;A: Yes. The paper repeats the displacement event study on the intersection of the nationally representative credit bureau sample (approximately 600,339 individuals with both credit information and employment histories) with the universe of IMSS data for October 2011–March 2014, yielding 8,723 mass layoff events. This sample is representative of the population of Mexican borrowers with formal employment histories, and the estimated effects on default for any loan in the credit bureau are similar in magnitude to the experimental-sample results, providing a measure of external validity.&lt;/p&gt;
&lt;h3 id="q12-what-do-the-debt-dynamics-during-the-experiment-reveal-about-the-mechanisms-for-interest-rate-effects-on-default"&gt;Q12. What do the debt dynamics during the experiment reveal about the mechanisms for interest rate effects on default?&lt;/h3&gt;
&lt;p&gt;A: The data show that purchases (net of payments) increase in response to interest rate decreases, consistent with downward-sloping demand for credit; yet total debt declines in lower-rate arms. This is consistent with the model&amp;rsquo;s prediction that the mechanical compounding effect (lower rate applied to previously accumulated debt) exceeds the behavioral new-purchase response. Confirmed empirically: the debt elasticity to the interest rate is estimated to be positive, with preferred estimates in the range [+0.18, +0.54]. The decline in default is further concentrated among borrowers with the highest baseline debt utilization rates, those for whom the debt compounding effect is strongest — consistent with the debt channel as the primary mechanism.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Cumulative Default Measure:&lt;/strong&gt; Default is defined as three consecutive monthly payments each below the required minimum payment due, at which point Bank A automatically revokes the card. The outcome variable is coded as Yit = 1 if borrower i has defaulted in any month s ≤ t and 0 otherwise, making it a cumulative (absorbing) measure. This allows estimation on an unchanging sample, avoiding attrition biases that would arise from conditioning on not having defaulted in the prior period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Minimum Payment Due (mpd):&lt;/strong&gt; The paper uses the required minimum payment due to avoid delinquency as its central cash-flow normalization variable. This is a comprehensive measure that incorporates not only the contractually specified fraction of outstanding balance but also interest charges, fees, and endogenous borrower responses (changes in debt and purchases). It serves as the common denominator for benchmarking the cash flow impacts of the two contract term interventions and formal job loss against one another.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Free Cash Flow / Per-Peso Normalization (λ):&lt;/strong&gt; The paper defines per-peso default effects (λ^IR, λ^MP, λ^U) by dividing each intervention&amp;rsquo;s average treatment effect on cumulative default (in percentage points) by the cumulative change in the minimum payment due (or equivalent cash flow impact) induced by that intervention over 18 months. The resulting ratio is expressed as percentage points of default per 1,000 MXN pesos of cash flow change. This normalization is explicitly not treated as an instrumental variable estimate; it is a descriptive back-of-the-envelope calculation intended to equate the scale of the three shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mass Layoff / Displacement:&lt;/strong&gt; A mass layoff at the firm level is defined as the first month in which year-on-year firm employment declines by more than 30% of average employment in the prior 12 months, restricted to firms with 50+ employees. An individual worker is classified as displaced if they lost formal-sector employment in the same calendar quarter as their employer&amp;rsquo;s mass layoff event. This definition follows Jacobson et al. (1993) and subsequent literature and is used to isolate plausibly involuntary (exogenous) separations from voluntary quits or individually driven terminations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Continuation Value (v):&lt;/strong&gt; In the paper&amp;rsquo;s two-period optimizing model, v is the reduced-form utility parameter capturing future flow of card benefits, warm glow from card ownership, or the option value of retaining access to formal credit, experienced only if the card is not in default. The paper uses v to rationalize the zero interest-rate response of newer borrowers: ceteris paribus, higher v implies that borrowers will remain current on the card even when interest rates are high, because they value continued access. Higher v thus implies more muted responses to interest rate changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bank Tenure Strata:&lt;/strong&gt; Borrowers are stratified into three groups based on length of relationship with the study card: &amp;ldquo;new customers&amp;rdquo; (6–11 months), medium-term (12–23 months), and long-term (24+ months). Tenure is used both as a stratification variable for the experiment and as a primary dimension of heterogeneity in treatment effects, reflecting differing default rates (36% vs. 18% at 26 months), labor market vulnerability (1.34× higher job loss probability for new vs. long-term), and interest rate responsiveness (zero for new, significantly positive for long-term borrowers).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Debt Burden Channel vs. Concurrent Moral Hazard:&lt;/strong&gt; The paper distinguishes three channels through which interest rate changes can affect default: (a) the debt burden channel — higher rates mechanically increase the stock of interest-accruing debt, making repayment harder; (b) concurrent moral hazard — higher current interest rates alter the incentive to default on existing obligations, holding debt constant; and (c) dynamic moral hazard — higher future interest rates reduce the benefit of remaining current. The paper&amp;rsquo;s finding of a modest total effect (elasticity 0.20) implies that the sum of all three channels is small in this context, with the debt burden channel being the primary driver of what effect does exist.&lt;/p&gt;</description></item><item><title>Default Options and Retirement Saving Dynamics</title><link>https://macropaperwarehouse.com/papers/default-options-and-retirement-saving-dynamics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/default-options-and-retirement-saving-dynamics/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; Does automatic enrollment (auto-enrollment) in retirement savings plans increase lifetime wealth accumulation and welfare? The prior literature established large short-run participation effects but had not traced the policy&amp;rsquo;s consequences over a full working life.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The paper draws on two primary sources. First, a proprietary panel of 401(k) administrative records from nearly 600 U.S. firms, covering roughly 159,216 first-year employees across 86 firms (for the &amp;ldquo;increasing default&amp;rdquo; fact) and 6,415 employees across 34 firms (for structural estimation), observed between December 2006 and December 2017. Second, 12 successive waves (2006–2017) of the U.K. Annual Survey of Hours and Earnings (ASHE), a 1% nationally representative panel of approximately 200,000 private-sector employees per year, including 37,120 job-switchers, used to exploit the phased rollout of the U.K. Pension Act of 2008.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Methodology.&lt;/strong&gt; The paper proceeds in three steps. (1) Three empirical stylized facts are documented using quasi-experimental variation (comparing employees hired before versus after changes in the default contribution rate within the same firm, and exploiting the staggered employer-size-based rollout of U.K. auto-enrollment). (2) A structural lifecycle model is estimated via the Method of Simulated Moments, using three preference parameters—intertemporal discount factor (δ), elasticity of intertemporal substitution (σ), and opt-out cost (k)—identified from the within-firm default variation in 34 U.S. firms. (3) The estimated model is used for out-of-sample validation and counterfactual welfare analysis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three stylized facts.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Fact I — Increasing the default reduces participation.&lt;/em&gt; Among 159,216 first-year employees in 86 auto-enrollment firms, each percentage-point increase in the default contribution rate reduces 401(k) participation by approximately 1 percentage point and increases contributions strictly below the new default by 1 percentage point. When the default rose from 3% to 6%, workers were 3.2 percentage points more likely to contribute at 1% or 2% of salary. This &amp;ldquo;drop-out&amp;rdquo; pattern is consistent with an opt-out cost model but is inconsistent with loss-aversion and psychological-anchoring theories, both of which predict that raising the default should weakly increase low-end contributions.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Fact II — Non-autoenrolled workers catch up within three years.&lt;/em&gt; In the estimation sample of 34 U.S. firms offering a 50% match up to 6% and an auto-enrollment default of 3%, median cumulative employee 401(k) contributions of non-autoenrolled workers equal those of autoenrolled workers after three years of tenure. Because non-autoenrolled workers compensate for initial non-participation by contributing more later—earning similar cumulative employer match and tax benefits over the full three-year horizon—a modest opt-out cost suffices to explain the observed inertia. Previous studies (which examined only the first year of tenure and did not allow future contribution adjustment) inferred opt-out costs of $1,000–$2,200 or more; the dynamic model implies a cost of only approximately $250.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Fact III — Prior auto-enrollment reduces saving in the next job.&lt;/em&gt; Using the phased U.K. policy rollout, workers who were auto-enrolled in their previous job and then move to a new employer that has not yet implemented auto-enrollment participate 12.8 percentage points less and contribute 0.55% of salary less in the new plan relative to otherwise similar job-switchers from non-auto-enrollment employers. When the new employer also has auto-enrollment, no statistically significant difference is observed. Placebo rollout tests confirm the effect is not a pre-existing selection pattern. This negative spillover contradicts a &amp;ldquo;savings habit&amp;rdquo; hypothesis and suggests that auto-enrollment&amp;rsquo;s short-run boost overstates lifetime savings effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural estimation results.&lt;/strong&gt; The estimated quarterly discount factor is δ = 0.987 (approximately 0.949 annually), and the elasticity of intertemporal substitution is σ = 0.435, both standard in lifecycle models. The opt-out cost is estimated at &lt;strong&gt;$254&lt;/strong&gt; per contribution-rate change (standard error $11). Sensitivity exercises show that combining a short observation window (first year only), sticky contributions (no intra-job adjustment), no income uncertainty, immediate vesting, and penalty-free DC withdrawals yields an opt-out cost of $3,004—broadly matching the range in previous studies. The low baseline estimate is thus driven by the dynamic nature of decisions (ability to compensate later), the illiquidity of retirement accounts (which reduces their perceived value), and income uncertainty (which expands the inaction range).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Long-run wealth effects.&lt;/strong&gt; Simulating a universal 3% auto-enrollment policy, the model predicts that &lt;strong&gt;wealth at retirement changes by less than 2% for the top 7 income deciles&lt;/strong&gt;. For individuals in the top two deciles, total wealth at age 65 is actually reduced by less than 1% because many high earners who would voluntarily contribute above 3% are pulled down to the default. At the &lt;strong&gt;bottom decile&lt;/strong&gt;, however, auto-enrollment raises total retirement wealth by more than &lt;strong&gt;12%&lt;/strong&gt;; savings increases are concentrated in the first 20 years of working life and peak around age 45, where bottom-quintile workers hold an additional 20% of average annual lifetime earnings. Even at the bottom, approximately one-third of the early savings gains are offset by lower contributions after age 45, as the wealth effect dominates. Crowd-out of liquid savings is limited: for bottom-quintile individuals, &lt;strong&gt;89%&lt;/strong&gt; of the increase in retirement savings at age 65 passes through to total wealth; for middle-quintile individuals, &lt;strong&gt;62%&lt;/strong&gt; passes through.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Out-of-sample validation.&lt;/strong&gt; The U.S.-estimated model is not rejected (at the 10% level) in 8 of 11 response moments in the 86-firm sample where defaults were raised between two positive rates, covering over 85% of workers. Recalibrated to U.K. institutions (using δ and σ from the U.S. and k = £160 via the average USD/GBP exchange rate), the model replicates the roughly 30-percentage-point increase in both participation and contributions at the 1% U.K. default. The model also predicts a 9.6-percentage-point drop in participation when workers move from an auto-enrollment to an opt-in employer, close to the empirical 12.8 percentage points.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Welfare and optimal policy.&lt;/strong&gt; Under utilitarian preferences (policymaker shares individuals&amp;rsquo; discount rate, no redistributive motive), the opt-in regime is always preferred to auto-enrollment regardless of policy incidence, because matching and tax incentives already induce over-saving relative to individuals&amp;rsquo; revealed time preferences. Under &lt;strong&gt;paternalistic&lt;/strong&gt; preferences (social discount factor = 1) or &lt;strong&gt;inequality-averse&lt;/strong&gt; preferences (Pareto weights inversely proportional to income, with degree of inequality aversion ν = 1 following Saez 2002), an auto-enrollment default at or near the employer matching threshold (6% of income) maximizes social welfare. A 6% auto-enrollment default improves welfare by 0.3% in lifetime consumption-equivalent for the bottom decile even under a utilitarian policymaker when incidence is on employers. These optimal policy rankings are robust to whether the opt-out cost is treated as fully welfare-relevant (π = 1) or welfare-irrelevant (π = 0), and hold under three incidence scenarios (employer profit reduction, match-rate adjustment, wage adjustment).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-mechanism-by-which-non-autoenrolled-workers-catch-up-at-the-median-and-why-does-this-reduce-the-implied-opt-out-cost-relative-to-prior-estimates"&gt;Q1. What is the core mechanism by which non-autoenrolled workers &amp;ldquo;catch up&amp;rdquo; at the median, and why does this reduce the implied opt-out cost relative to prior estimates?&lt;/h3&gt;
&lt;p&gt;A: Non-autoenrolled workers who do not contribute in their first year are not permanently forgoing employer matching and tax benefits; they can contribute more later in the same job and earn similar cumulative benefits. The paper shows that at the median and 75th percentile, cumulative employee 401(k) contributions among opt-in workers equal those of autoenrolled workers after three years of tenure in 34 U.S. firms offering a 50%-up-to-6% match at a 3% default. This dynamic substitutability means the opportunity cost of initial non-participation is far smaller than one-period back-of-the-envelope calculations suggest. Previous studies, which implicitly or explicitly assumed static contribution decisions or examined only the first year, inferred opt-out costs of $1,000–$2,200; in a fully dynamic model the same inertia requires only ~$254.&lt;/p&gt;
&lt;h3 id="q2-why-does-fact-i-higher-default-reduces-participation-specifically-rule-out-loss-aversion-and-anchoring-as-the-primary-mechanism-and-what-does-it-support-instead"&gt;Q2. Why does Fact I (higher default reduces participation) specifically rule out loss aversion and anchoring as the primary mechanism, and what does it support instead?&lt;/h3&gt;
&lt;p&gt;A: Under loss aversion, contributions above the default feel like losses while contributions below the default feel like gains. Raising the default shifts some contributions from the loss domain into the gain domain, making low contributions relatively less attractive. Proposition 2 demonstrates formally that loss-averse preferences predict a weakly lower fraction contributing below the new (higher) default — the opposite of what is observed. Similarly, Proposition 3 shows that psychological anchoring shifts preferences toward the new default, also predicting more participation at low rates when the default rises. Only the opt-out cost model (Proposition 1) predicts that a higher default causes some workers to incur the cost to switch &lt;em&gt;away&lt;/em&gt; from the default and end up at lower contribution rates, matching the empirical finding that each 1-percentage-point rise in the default increases contributions strictly below the old default by approximately 1 percentage point.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-quantitative-magnitude-of-the-opt-out-cost-and-what-modeling-assumptions-are-responsible-for-it-being-much-smaller-than-prior-estimates"&gt;Q3. What is the quantitative magnitude of the opt-out cost, and what modeling assumptions are responsible for it being much smaller than prior estimates?&lt;/h3&gt;
&lt;p&gt;A: The baseline estimate is $254 per contribution-rate change (s.e. $11), roughly an order of magnitude smaller than prior estimates of $1,000–$3,000+. Table 4 decomposes the sources of the difference: using only first-year data changes the estimate only slightly (to $226). Assuming contributions cannot be changed within a job (&amp;ldquo;sticky contributions&amp;rdquo;) raises the cost to $308 with four years of data or $712 with one year of data. Eliminating income uncertainty raises the estimate to $465. Assuming immediate vesting raises it to $344. Assuming penalty-free DC withdrawals raises it to $609. Combining all these restrictions simultaneously yields $3,004 — closely matching the prior literature. The three key drivers are thus: (1) the ability to adjust contributions over time within a job; (2) the illiquidity of the DC account (early-withdrawal penalties); and (3) income uncertainty widening the inaction range.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-paper-validate-the-structural-model-out-of-sample-and-what-confidence-does-this-provide-in-the-long-run-predictions"&gt;Q4. How does the paper validate the structural model out of sample, and what confidence does this provide in the long-run predictions?&lt;/h3&gt;
&lt;p&gt;A: Two out-of-sample exercises are reported. First, the model estimated on 34 U.S. firms (introduction of auto-enrollment from 0% to 3% default) is used to predict workers&amp;rsquo; response when 86 other firms raised the default from one positive rate to a higher rate. The model prediction cannot be rejected at the 10% level in 8 of 11 response-moment cases, covering 71 of 86 firms and more than 85% of workers. Second, the model is re-calibrated to U.K. institutions (keeping U.S. preference estimates, setting k = £160 via exchange rate) and applied to the phased rollout of the U.K. Pension Act of 2008. The model replicates the roughly 30-percentage-point increase in both participation and contributions at the 1% default following the policy, and predicts a 9.6-percentage-point drop in participation when previously autoenrolled workers move to a new opt-in employer — compared with an empirical estimate of 12.8 percentage points (s.e. 5.5 pp).&lt;/p&gt;
&lt;h3 id="q5-what-are-the-distributional-implications-of-a-universal-3-auto-enrollment-policy-for-wealth-at-retirement"&gt;Q5. What are the distributional implications of a universal 3% auto-enrollment policy for wealth at retirement?&lt;/h3&gt;
&lt;p&gt;A: The effect is concentrated at the bottom. For the top 7 income deciles, retirement wealth at age 65 changes by less than 2% relative to the opt-in counterfactual. For the top two deciles, total wealth at age 65 is actually reduced by less than 1% because high-earning workers who would voluntarily contribute above 3% are pulled down to the default. For the bottom decile, the policy raises total retirement wealth by more than 12%. Even at the bottom, roughly one-third of the early savings gains are later offset by lower contributions after age 45 as the wealth effect dominates, so even 20-year empirical follow-ups may overstate the policy&amp;rsquo;s lifetime effect at the bottom.&lt;/p&gt;
&lt;h3 id="q6-how-large-is-crowd-out-of-liquid-savings-by-auto-enrollment-and-what-explains-the-limited-degree-of-substitution"&gt;Q6. How large is crowd-out of liquid savings by auto-enrollment, and what explains the limited degree of substitution?&lt;/h3&gt;
&lt;p&gt;A: Crowd-out is modest. For bottom-quintile workers, 89% of the increase in retirement savings at age 65 translates into higher total wealth; for middle-quintile workers, 62% passes through. The limited crowd-out arises because liquid assets serve a precautionary motive and DC accounts serve a lifecycle motive — the two assets are not close substitutes. Additionally, as in Kaplan and Violante (2014), the marginal propensity to consume out of liquid assets is high in the model, so autoenrolled workers reduce consumption rather than run down liquid balances. These predictions align with Beshears et al. (2021), who find no significant increase in unsecured debt after four years, and Chetty et al. (2014), who estimate an 80% pass-through to total savings in a different Danish policy.&lt;/p&gt;
&lt;h3 id="q7-why-do-previously-autoenrolled-workers-contribute-less-when-they-switch-to-an-opt-in-employer-and-how-is-this-consistent-with-the-model"&gt;Q7. Why do previously autoenrolled workers contribute less when they switch to an opt-in employer, and how is this consistent with the model?&lt;/h3&gt;
&lt;p&gt;A: The most plausible explanation, and the one consistent with the model&amp;rsquo;s out-of-sample predictions, is a standard wealth effect: workers auto-enrolled early accumulate more retirement wealth and therefore have less incentive to contribute in a new job. The model predicts a 9.6-percentage-point participation drop for AE-to-non-AE movers, close to the empirical 12.8 pp. An alternative explanation — that previously autoenrolled workers rationally expect their new employer to soon adopt auto-enrollment and thus delay active enrollment — is partially ruled out by the finding that the empirical estimate is closer to the model prediction for job-switchers whose new employer is not expected to adopt auto-enrollment in the next 12 months.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-welfare-implications-of-auto-enrollment-under-utilitarian-paternalistic-and-inequality-averse-policymakers-and-how-robust-are-these-to-the-incidence-assumption"&gt;Q8. What are the welfare implications of auto-enrollment under utilitarian, paternalistic, and inequality-averse policymakers, and how robust are these to the incidence assumption?&lt;/h3&gt;
&lt;p&gt;A: Under utilitarian preferences (policymaker shares individuals&amp;rsquo; discount factor, no extra redistributive weight), the opt-in regime is always preferred regardless of whether the policy&amp;rsquo;s cost falls on employer profits, the match rate, or wages. The negative welfare effect is largest when incidence falls on wages (approximately 50% larger than under match-rate reduction). Under paternalistic preferences (social discount factor = 1), a 6% default (equal to the employer matching threshold) is optimal under all three incidence scenarios. Under inequality-averse preferences (ν = 1 Pareto weights), a 6% default is optimal when incidence falls on employers, and a 5% default when incidence falls on workers. These results are identical whether the opt-out cost is treated as fully welfare-relevant (π = 1) or welfare-irrelevant (π = 0). A 6% auto-enrollment default increases welfare by 0.3% in lifetime consumption-equivalent for the bottom income decile even under a utilitarian planner when incidence is on employers.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-address-heterogeneity-in-default-effects-across-age-and-income-groups-within-a-parsimonious-homogeneous-preference-model"&gt;Q9. How does the paper address heterogeneity in default effects across age and income groups within a parsimonious homogeneous preference model?&lt;/h3&gt;
&lt;p&gt;A: The model has only three estimated preference parameters (δ, σ, k), yet it endogenously replicates empirical heterogeneity. Conditional on participating, workers in their 20s are approximately 20 percentage points more likely to stay at the 3% default than workers in their late 50s and early 60s; the model attributes this to the option value of waiting: young workers can compensate for current non-saving by contributing more later, so the cost of opting out is effectively smaller for them. The lowest-income workers are approximately 40 percentage points more likely to remain at the default than the highest-paid; the model explains this primarily because the fixed opt-out cost of $254 represents a larger share of earnings for low-income individuals (and secondarily because high-income workers have more to gain from active contribution decisions due to higher marginal tax rates and a lower Social Security replacement rate). All model-predicted coefficients fall within the 95% confidence intervals of the empirical estimates.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-paper-conclude-about-the-broader-relevance-of-the-dynamic-opt-out-cost-framework-beyond-retirement-saving"&gt;Q10. What does the paper conclude about the broader relevance of the &amp;ldquo;dynamic opt-out cost&amp;rdquo; framework beyond retirement saving?&lt;/h3&gt;
&lt;p&gt;A: The paper argues that wherever individuals can compensate for present inaction with future actions — as in retirement saving — the observed inertia at a default understates the freedom of choice preserved by the nudge, and short-run effects overstate long-term consequences. In contrast, in domains such as healthcare plan choice or school selection, future actions cannot easily offset present inertia; opt-out costs are likely to remain large; and the distinction between a nudge and a hard mandate collapses. The paper therefore argues that the appeal of &amp;ldquo;libertarian paternalism&amp;rdquo; (Thaler and Sunstein 2003) is domain-specific and is strongest precisely where intertemporal adjustment is possible.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Opt-out cost (k).&lt;/strong&gt; In this paper, a utility cost — estimated at $254 per contribution-rate change — that individuals must pay every time they choose a retirement contribution rate different from the current default. The cost is modeled as a consumption reduction and captures both real transaction costs (form-filling, adviser fees) and behavioral costs (cognitive cost of attention and optimal-choice search). It is fixed and homogeneous across individuals, and applies symmetrically in any direction of deviation from the default.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Auto-enrollment default contribution rate.&lt;/strong&gt; The positive contribution rate at which new hires are automatically enrolled in a defined-contribution plan, with the option to opt out by incurring the opt-out cost. In the paper&amp;rsquo;s estimation sample, this is 3% of salary. The default is exogenous at the start of each new job but endogenous thereafter: once established, the default for subsequent periods equals the worker&amp;rsquo;s contribution rate in the previous period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Default eﬀect.&lt;/strong&gt; The empirically observed tendency of workers to remain at the default contribution rate rather than actively choosing a different rate. In this paper, the default effect is explained by opt-out costs rather than loss aversion or psychological anchoring — a distinction identified through the novel prediction that raising the default from a positive rate to a higher positive rate reduces overall participation (the &amp;ldquo;drop-out&amp;rdquo; effect), a pattern consistent only with opt-out costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Drop-out eﬀect.&lt;/strong&gt; The paper&amp;rsquo;s term (following Caplin and Martin 2017) for the empirical finding that increasing the auto-enrollment default contribution rate causes some workers to stop contributing altogether or to contribute at rates strictly below the initial default. This effect is used as a discriminating test between competing theories of the default effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic opt-out cost framework.&lt;/strong&gt; The paper&amp;rsquo;s core modeling insight: that opt-out costs must be estimated in a fully dynamic lifecycle model that allows workers to adjust contributions over time, to hold liquid assets and unsecured debt, and to face labor market risk. In a static or short-horizon model, the opportunity cost of initial non-participation appears large (because the worker permanently forgoes match and tax benefits), requiring large opt-out costs. In the dynamic model, the ability to compensate later shrinks the implied opportunity cost and hence the opt-out cost required to rationalize observed inertia.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crowd-out of liquid savings.&lt;/strong&gt; The extent to which higher DC retirement contributions induced by auto-enrollment reduce liquid asset holdings (or increase unsecured borrowing), rather than increasing total wealth. The paper estimates limited crowd-out (89% pass-through to total wealth for bottom-quintile workers, 62% for middle-quintile workers), attributable to the different roles of liquid assets (precautionary motive) and DC accounts (lifecycle motive) in the model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy incidence.&lt;/strong&gt; The channel through which employers balance their budget in response to higher matching costs created by auto-enrollment. The paper considers three scenarios: employers absorb costs through reduced profits; employers reduce the match rate; employers reduce wages. Optimal policy rankings and welfare magnitudes differ across these scenarios, but the qualitative conclusions — utilitarian policymaker prefers opt-in; paternalistic or inequality-averse policymaker prefers AE at 6% — are robust across incidence assumptions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption-equivalent variation (γ).&lt;/strong&gt; The welfare metric used in the paper: the proportional increase in consumption in every period and every state of the world that would make the policymaker indifferent between an auto-enrollment policy at default d and the opt-in regime. A 6% default increases welfare by 0.3% in consumption-equivalent for the bottom income decile under a utilitarian policymaker when incidence is on employers.&lt;/p&gt;</description></item><item><title>Defying Distance? The Provision of Medical Services in the Digital Age</title><link>https://macropaperwarehouse.com/papers/defying-distance-the-provision-of-medical-services-in-the-digital-age/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/defying-distance-the-provision-of-medical-services-in-the-digital-age/</guid><description>&lt;p&gt;This paper asks whether digital platforms can improve healthcare outcomes by enabling needs-based matching between patients and physicians unconstrained by geography. Amanda Dahlstrand studies digital primary care in Sweden during 2016-2018, exploiting nationwide conditional random assignment between approximately 200,000 patients and 143 doctors employed by Europe&amp;rsquo;s largest digital primary care provider. Patients who selected the &amp;ldquo;first available doctor&amp;rdquo; option (82% of first visits) were effectively randomized to a doctor within each 3-hour shift-by-date stratum, generating quasi-experimental variation free of the patient-doctor sorting that confounds identification in physical primary care.&lt;/p&gt;
&lt;p&gt;The paper defines three observable dimensions of primary care physician skill: (1) identifying risky patients and triaging them to higher levels of care, measured by whether patients subsequently have an avoidable hospitalization within 90 days; (2) providing guideline-consistent treatment, measured by counter-guideline antibiotic prescriptions; and (3) leaving patients sufficiently informed so they do not unnecessarily seek additional in-person care within the following week. Doctor skill in each dimension is estimated via a value-added framework in a hold-out sample (Sample 1, the first 600 randomized consultations per doctor), using empirical Bayes shrinkage to reduce noise. Complementarities between doctor skill and patient risk are then estimated in a disjoint main sample (Sample 2).&lt;/p&gt;
&lt;p&gt;A central finding is that doctor skill is task-specific rather than governed by a single latent ability: skills across the three tasks are not positively correlated, meaning doctors within general practice have individual &amp;ldquo;specializations.&amp;rdquo; A patient ranked in the top 1% of avoidable hospitalization risk who is matched to a doctor ranked in the top 10% at reducing avoidable hospitalizations experiences a 90% reduction in that adverse outcome, relative to a patient with the same risk profile matched to the worst-performing doctor. Patients not estimated as risky show effects indistinguishable from zero when matched to the same high-skilled doctors, establishing a strong complementarity between doctor type and patient risk.&lt;/p&gt;
&lt;p&gt;Using the Average Match Function framework of Graham, Imbens, and Ridder (2014, 2020), the paper evaluates counterfactual reallocation policies. Reallocating only 2% of patients — those in the top 1% of predicted avoidable hospitalization risk — to doctors in the top 10% of triage skill reduces aggregate avoidable hospitalizations by 20% relative to random assignment, without adversely affecting counter-guideline prescriptions or other measured outcomes. Doctor skills across outcomes are not positively correlated, so this reallocation does not generate meaningful trade-offs. The paper benchmarks this matching policy against a selective hiring/expansion policy in which doctors with above-median skill in three tasks expand their hours by up to 70% at the expense of below-median peers; that policy yields no significant reduction in avoidable hospitalizations and only a 4% reduction in counter-guideline prescriptions — smaller gains than matching and harder to implement.&lt;/p&gt;
&lt;p&gt;The paper also documents that physical primary care quality is worse in lower-income and more deprived areas of Sweden (a negative relationship between deprivation index and patient-reported experience is statistically significant at the 1% level in a cross-section of roughly 120-150 primary care centers in Region Skane). Because the estimated risk of avoidable hospitalization and prior avoidable hospitalizations are concentrated in the lower end of the income distribution, needs-based digital matching reallocates triage skill toward lower-income patients, severing the correlation between local area income and service quality. Simulating positive assortative matching on patient income and doctor skill — approximating existing healthcare inequalities — leads to more avoidable hospitalizations than random assignment, because the most vulnerable patients tend to be the poorest. Scope conditions: findings derive from a single digital primary care provider in Sweden, 2016-2018, pre-pandemic, covering conditions amenable to video consultation and a patient pool younger and somewhat more urban than the average Swedish citizen.&lt;/p&gt;
&lt;p&gt;Q: What is the key identification strategy, and why is it valid in this setting but not in physical primary care?
A: Patients who selected the &amp;ldquo;drop in&amp;rdquo; (first available doctor) option — 82% of first visits — were assigned to whichever certified doctor was next in the roster within a 3-hour shift-by-date stratum, a by-product of the first-come-first-served queue. Neither patients nor doctors could intervene in this digital process. The author validates the assumption by regressing doctor characteristics on patient characteristics controlling for shift-by-date fixed effects and finds characteristics are balanced. In physical primary care, endemic patient-doctor sorting means doctors do not meet a common support of patient types, preventing causal identification of doctor effects.&lt;/p&gt;
&lt;p&gt;Q: How are doctor skill estimates constructed and why does the split-sample matter?
A: Doctor skill in each task is estimated as an empirical Bayes-shrunk random effect from a value-added regression on Sample 1, each doctor&amp;rsquo;s first 600 randomized consultations (40% of the sample). Sample 2 (60%) is entirely disjoint and used to estimate complementarities between doctor skill and patient risk. The split-sample design prevents overfitting: doctor skill was estimated on different patients than those in Sample 2. The Durbin-Wu-Hausman test does not reject random effects (p = 0.16).&lt;/p&gt;
&lt;p&gt;Q: What is the main quantitative result on avoidable hospitalization matching?
A: A patient ranked in the top 1% of predicted avoidable hospitalization risk matched to a doctor ranked in the top 10% at reducing avoidable hospitalizations could reduce that patient&amp;rsquo;s avoidable hospitalizations by 90%, relative to the worst-performing doctor in that skill. At the aggregate level, reallocating only 2% of patients (those in the top 1% risk group) to high-triage-skill doctors reduces avoidable hospitalizations across the full patient population by 20% compared to random assignment.&lt;/p&gt;
&lt;p&gt;Q: Does the avoidable hospitalization reallocation harm other outcomes?
A: No. The paper explicitly evaluates the Average Reallocation Effect on counter-guideline prescriptions and additional in-person care seeking when optimizing for avoidable hospitalizations, and finds no significant adverse effects on these other outcomes. The author attributes this to the fact that doctor skills across tasks are not positively correlated, so reallocating triage-skilled doctors does not systematically remove skill from other dimensions.&lt;/p&gt;
&lt;p&gt;Q: How does matching compare to selective hiring and hour expansion as a policy?
A: Even expanding the working hours of doctors with above-median skill across three tasks by as much as 70% yields no significant reduction in avoidable hospitalizations and only a 4% reduction in counter-guideline prescriptions — both smaller gains than the matching policy. Matching outperforms hiring expansion because patients have heterogeneous needs that can be identified from prior healthcare records, and doctors have differentiated skill sets relevant to some patients but not others.&lt;/p&gt;
&lt;p&gt;Q: What is the evidence that doctor skills are task-specific rather than reflecting a single latent ability?
A: The estimated doctor effects across the three tasks — triaging to avoid hospitalizations, guideline-consistent antibiotic prescribing, and minimizing unnecessary follow-up care — are not positively correlated with one another. This means a doctor who is effective at one task is not systematically effective at others, indicating individual specializations within general practice that are not accounted for in standard primary care organization.&lt;/p&gt;
&lt;p&gt;Q: How is patient risk for avoidable hospitalizations measured?
A: A propensity score is estimated from pre-digital physical healthcare data (2013-2015), regressing past number of avoidable hospitalizations on demographic and healthcare utilization variables — including age, a disease index of chronic diagnoses, and previous hospitalizations — all variables already available in patient medical records. The top 1% of predicted risk scores are classified as &amp;ldquo;risky.&amp;rdquo; Patients in the risky group had on average 0.35 avoidable hospitalizations in the prior 3 years, versus 0.01 for non-risky patients.&lt;/p&gt;
&lt;p&gt;Q: What is the distributional (equity) implication of needs-based matching versus income-assortative matching?
A: Estimated risk of avoidable hospitalization and the count of prior avoidable hospitalizations are concentrated in the lower end of the income distribution. Needs-based matching therefore reallocates triage skill toward lower-income patients. Simulating positive assortative matching on patient income and doctor skill — approximating observed inequalities in physical care — produces more avoidable hospitalizations than random assignment, because the most vulnerable patients are often the poorest. Needs-based digital matching can sever the link between local area income and service quality.&lt;/p&gt;
&lt;p&gt;Q: How does digital care usage sort by income and demographics in the data?
A: At the extensive margin, the deprivation index (Care Need Index) is similar among digital users and non-users in Region Skane. However, at the intensive margin, individuals with a higher deprivation index who use the digital service have more appointments in it; similarly, lower-income users use the service more intensively. Digital care users are younger than non-users and are more likely to live in cities than the average Swedish citizen.&lt;/p&gt;
&lt;p&gt;Q: What are avoidable hospitalizations and why are they the primary outcome?
A: Avoidable hospitalizations (also called hospitalizations for ambulatory care sensitive conditions) are hospital admissions defined in the medical literature as preventable by adequate and timely primary care. They are coded using ICD-10 diagnosis codes listed in Page et al. (2007). The most common diagnoses in the 90-day post-consultation window are respiratory and genitourinary, conditions commonly treated in digital care. The outcome is rare (0.2% of patients in the sample), but high-stakes: an estimated 1.1 potential life years are lost per avoidable hospitalization, and in Sweden they cost an estimated SEK 7.1 billion (~$820 million) annually (7% of inpatient curative and rehabilitative care costs).&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the counter-guideline antibiotic prescription outcome?
A: Non-adherence is coded against 16 guidelines from Sweden&amp;rsquo;s strategic programme against antibiotic resistance (Strama 2017, 2019), all designed to limit or narrow antibiotic use. The measured rate of non-adherence is described as quite low by international standards; the CDC estimates 28% of US antibiotic prescriptions are unnecessary, while the author&amp;rsquo;s sample rate is 2%. The guidelines require doctors to sometimes refuse patients who request antibiotics, introducing a behavioral compliance dimension to this skill.&lt;/p&gt;
&lt;p&gt;Q: What are the costs and feasibility considerations for implementing needs-based digital matching?
A: The paper characterizes matching as a &amp;ldquo;resource-neutral&amp;rdquo; policy because it reallocates existing doctors without hiring or training. The primary costs are a small increase in waiting time for some patients and the costs of importing data and developing the matching algorithm. Because the algorithm handles patient-doctor allocation while doctors retain all clinical decision-making, the policy functions as a complement to human skill rather than a substitute, which the author argues makes it less subject to &amp;ldquo;algorithm aversion.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Q: Why does the paper restrict to each patient&amp;rsquo;s first digital consultation only?
A: The first visit is the one subject to conditional random assignment; subsequent visits could reflect endogenous selection by patients who preferred a particular doctor or outcome. Using only first visits eliminates this concern. The restriction reduces the sample from approximately 378,000 to 210,171 patients (56% of the original), paired with 143 doctors who each had at least 600 randomized consultations.&lt;/p&gt;
&lt;p&gt;Conditional random assignment: The allocation mechanism by which patients selecting the &amp;ldquo;first available doctor&amp;rdquo; option in digital primary care were assigned to whichever certified doctor was next in the shift roster, conditional on 3-hour shift-by-date strata — a by-product of the first-come-first-served queue rather than an intended experimental design.&lt;/p&gt;
&lt;p&gt;Average Match Function (AMF): The conditional mean of a patient outcome given observable doctor type and patient type under random assignment, β(x,w) = E[Y|X=x, W=w], which serves as the building block for evaluating counterfactual reallocation policies.&lt;/p&gt;
&lt;p&gt;Average Reallocation Effect (ARE): The difference in expected patient outcomes between a counterfactual doctor-patient assignment and the status quo random assignment, taking into account the externality on the patient from whom a high-skilled doctor is moved.&lt;/p&gt;
&lt;p&gt;Task-specific doctor skill: The paper&amp;rsquo;s finding that primary care physician effectiveness is not governed by a single latent ability but varies across distinct tasks — triage/risk prediction, guideline-consistent prescribing, and minimizing unnecessary follow-up care — with skills across tasks not positively correlated.&lt;/p&gt;
&lt;p&gt;Avoidable hospitalization: A hospital admission coded to a diagnosis (per Page et al. 2007 ICD-10 classification) defined in the medical literature as preventable by adequate and timely primary care, used as the primary high-stakes outcome measure (0.2% incidence in the sample within 90 days of a digital consultation).&lt;/p&gt;
&lt;p&gt;Counter-guideline prescription: A prescription of an antibiotic in violation of one of 16 guidelines from Sweden&amp;rsquo;s Strama antibiotic resistance programme, all of which are designed to limit use or require narrower-spectrum first-line antibiotics; used as the primary guideline-adherence outcome (2% incidence in the sample).&lt;/p&gt;
&lt;p&gt;Empirical Bayes shrinkage: A procedure applied to raw doctor value-added estimates in which the noisy estimate of doctor quality is multiplied by the ratio of signal variance to total (signal plus noise) variance, yielding a best linear predictor of the underlying doctor random effect and reducing noise from small-sample estimation.&lt;/p&gt;</description></item><item><title>Demand Stimulus as Social Policy</title><link>https://macropaperwarehouse.com/papers/demand-stimulus-as-social-policy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/demand-stimulus-as-social-policy/</guid><description>&lt;p&gt;This paper estimates the distributional and social consequences of Department of Defense (DOD) contract spending using a city-level (CBSA) panel dataset spanning 2005–2016. The research question is whether demand stimulus — specifically DOD spending, the largest category of U.S. discretionary government spending — has differential effects across demographic groups and whether it improves social outcomes typically targeted by dedicated government programs. A secondary question is whether these effects are specific to DOD spending or common to any demand shock.&lt;/p&gt;
&lt;p&gt;The empirical strategy exploits variation in DOD contract spending from USAspending.gov, constructing a proxy for outlays over time using contract duration, and instrumenting with a Bartik-type shock (location&amp;rsquo;s average DOD share interacted with aggregate contract spending). The main specification is a two-year differenced panel regression with CBSA and time fixed effects. Social outcomes come primarily from the American Community Survey (ACS), covering 290 CBSAs; mortality data come from the CDC; crime data from the FBI/NACJD. For comparison, the authors construct a general demand shock series using the standard Bartik shift-share approach across two-digit industries, which is nearly uncorrelated with the DOD shock (correlation -0.07).&lt;/p&gt;
&lt;p&gt;Main findings on distributional effects: A 1 percent increase in DOD spending as a share of local earnings raises overall average ACS earnings by 0.43 percent but raises average earnings for households without a bachelor&amp;rsquo;s degree by 0.71 percent, and raises average earnings for Black households by a slightly larger amount, while Whites receive the majority of total income. The employment rate rises by 0.22 percentage points per percent increase in DOD spending. Labor force participation is largely unchanged in aggregate, but rises 0.08 percentage points for the middle-aged (41–61) and 0.14 percentage points for those with a bachelor&amp;rsquo;s degree.&lt;/p&gt;
&lt;p&gt;On social outcomes: The poverty rate falls 0.08 percentage points, driven entirely by those without a bachelor&amp;rsquo;s degree. Food stamp (SNAP) receipt falls 0.08 percentage points. Self-reported disability rates fall, particularly among households without a bachelor&amp;rsquo;s degree. Occupational prestige rises by 0.024 points overall (0.037 for those without a bachelor&amp;rsquo;s degree). Travel time to work falls by 6.7 minutes per day, implying an annual benefit exceeding $558 per worker at a value of time of $10/hour. Marriage rates rise and divorce rates fall for some demographic groups. Homeownership increases significantly for some groups. Mortality falls, with 2.61 fewer deaths per 100,000 among those age 45–65 and 8.49 fewer deaths per 100,000 among those over 65 per percent increase in DOD spending; health-related deaths account for the majority of the decline. Crime is largely unaffected, except for a statistically significant reduction in vehicle theft.&lt;/p&gt;
&lt;p&gt;Comparing DOD to general demand shocks: Although both raise total earnings by similar amounts ($0.56 and $0.63 per dollar of shock, respectively), the general demand shock produces only about half the employment rate response (14.3 vs. 24.5 percentage point increase for households without a bachelor&amp;rsquo;s degree), concentrates earnings gains among already-employed, higher-educated, and White households, produces weaker effects on disability and occupational prestige, increases mortality by approximately 100 deaths per 100,000, and increases crime (vehicle theft and aggravated assault). The differential mortality response is partly attributed to differential pollution effects: general demand shocks raise the median AQI substantially, while DOD shocks do not. The differential employment effects of DOD shocks are explained primarily by city and occupational composition rather than industry composition: DOD shocks are directed toward smaller, lower-earnings cities with lower employment rates and fewer college-educated residents, and toward construction, manufacturing, and production/maintenance occupations with high no-bachelor&amp;rsquo;s shares.&lt;/p&gt;
&lt;p&gt;Scope conditions: Results are identified using CBSA-level variation over 2005–2016. DOD spending is treated as predominantly supply-side-driven and not directly entering household utility or local infrastructure. The social outcome results are local partial-equilibrium estimates and do not account for general equilibrium spillovers across CBSAs.&lt;/p&gt;
&lt;p&gt;Q: What is the core identification strategy, and why is DOD spending considered a valid instrument for demand stimulus?
A: DOD contract data from USAspending.gov are used to construct a proxy for outlays (distributing contract obligations over contract duration), and this measure is instrumented with a Bartik-type shock (location&amp;rsquo;s average DOD share times aggregate contract growth). The Bartik IV isolates the component of DOD contracts associated with new production, addressing endogeneity and the &amp;ldquo;anticipated contracts&amp;rdquo; problem. DOD spending is treated as predetermined relative to local business cycles and does not directly enter household utility or local infrastructure, isolating the aggregate demand channel.&lt;/p&gt;
&lt;p&gt;Q: Which demographic groups receive the most total income from DOD spending, and which see the largest relative gains?
A: In absolute terms, the majority of wage and salary income from DOD spending accrues to Whites and to those without a bachelor&amp;rsquo;s degree. However, adjusting for existing income shares, Black households and households without a bachelor&amp;rsquo;s degree experience the largest proportional increases in average earnings: a 1 percent increase in DOD spending as a share of local earnings raises average earnings for no-bachelor&amp;rsquo;s households by 0.71 percent, compared to a 0.43 percent increase in overall average earnings.&lt;/p&gt;
&lt;p&gt;Q: How does DOD spending affect employment at the extensive margin, and what does this imply about who benefits?
A: A 1 percent increase in DOD spending as a share of local earnings raises the overall employment rate by 0.22 percentage points. The large employment response among those without a bachelor&amp;rsquo;s degree (24.5 percentage points in the comparative analysis) implies that DOD spending disproportionately benefits previously unemployed workers rather than simply raising wages for those already employed.&lt;/p&gt;
&lt;p&gt;Q: Does DOD spending increase labor force participation?
A: There is no detectable aggregate effect on labor force participation rates, suggesting limited effects of demand stimulus on the participation margin over short horizons. However, participation rises 0.08 percentage points for the middle-aged (41–61) and 0.14 percentage points for those with a bachelor&amp;rsquo;s degree. The population response is strongest for those without a bachelor&amp;rsquo;s degree, though the estimate is imprecise.&lt;/p&gt;
&lt;p&gt;Q: What are the poverty and welfare effects of DOD spending?
A: A 1 percent increase in DOD spending as a share of local earnings reduces the poverty rate by 0.08 percentage points, with the entire effect concentrated among households without a bachelor&amp;rsquo;s degree. SNAP (food stamp) receipt falls by 0.08 percentage points. Medicaid receipt falls significantly for young children, while children substitute into private health insurance, leaving overall child health insurance coverage unchanged.&lt;/p&gt;
&lt;p&gt;Q: How does DOD spending affect disability rates?
A: A 1 percent increase in DOD spending leads to a 0.001 percentage point reduction in self-reported disability rates among households without a bachelor&amp;rsquo;s degree. The effect is most apparent for this group, the middle-aged, and Whites. In the comparative analysis, the employment margin accounts for a disability decline of -0.051 for no-bachelor&amp;rsquo;s households, nearly half of the total disability decline of -0.114 for that group.&lt;/p&gt;
&lt;p&gt;Q: What are the occupational prestige and commute time effects?
A: A 1 percent increase in DOD spending raises a city&amp;rsquo;s average occupational prestige score (Siegel score) by 0.024 points, with the effect concentrated among no-bachelor&amp;rsquo;s households (0.037). Commute time falls by 6.7 minutes per day; at a value of time of $10/hour, this implies an annual benefit of approximately $558 per worker.&lt;/p&gt;
&lt;p&gt;Q: How does DOD spending affect household formation outcomes?
A: Marriage rates increase and the likelihood of single parenthood decreases for White households. Divorce rates decrease for middle-aged and Black households. White households become more likely to own homes and less likely to live in multi-family homes. Estimates for Black and Hispanic households are imprecise.&lt;/p&gt;
&lt;p&gt;Q: What are the mortality effects of DOD spending, and how do they compare to general demand shocks?
A: A 1 percent increase in DOD spending as a share of local income leads to 2.61 fewer deaths per 100,000 among those aged 45–65 and 8.49 fewer deaths per 100,000 among those over 65, with health-related deaths accounting for the majority of the decline. This implies the DOD must spend approximately $25 million to save a life aged 45–65, exceeding the typical value of a statistical life. By contrast, a general demand shock increases mortality by approximately 100 deaths per 100,000, consistent with Ruhm&amp;rsquo;s (2000) finding that mortality is procyclical; mortality increases from general shocks are also concentrated among those over 45.&lt;/p&gt;
&lt;p&gt;Q: What explains the divergent mortality effects of DOD and general demand shocks?
A: One mechanism explored is pollution: general demand shocks raise median AQI substantially while DOD shocks leave AQI largely unaffected, consistent with Ruhm&amp;rsquo;s (2000) emphasis on deteriorating health behaviors during expansions. The paper also points to differential occupational and geographic composition: DOD shocks flow to construction, manufacturing, and production/maintenance occupations rather than to higher-pollution or higher-accident-risk activities common in broad economic expansions.&lt;/p&gt;
&lt;p&gt;Q: How do the crime effects differ between DOD and general demand shocks?
A: DOD spending shocks are associated with a statistically significant reduction in vehicle theft but no significant change in other crime categories. General demand shocks, by contrast, appear to increase vehicle theft and aggravated assault. Voter turnout falls substantially in response to a general demand shock; both shock types reduce Democratic vote shares.&lt;/p&gt;
&lt;p&gt;Q: What is the key mechanism explaining why DOD shocks have stronger social effects than general demand shocks?
A: Despite similar average earnings effects for no-bachelor&amp;rsquo;s households (0.71 for DOD vs. 0.69 for general shocks), DOD shocks produce a much larger employment rate increase for that group (24.5 vs. 14.3 percentage points). The authors show that this employment margin accounts for large shares of the differential declines in poverty, food stamp receipt, disability, and improvements in marriage rates and occupational prestige.&lt;/p&gt;
&lt;p&gt;Q: What accounts for the differential employment effects on no-bachelor&amp;rsquo;s households between DOD and general demand shocks?
A: Of the 0.21 percentage point differential employment effect, roughly one quarter is associated with differences in the no-bachelor&amp;rsquo;s share across industries. Differences across cities and across occupations each account for much larger shares. DOD shocks are directed toward smaller, lower-income, lower-employment cities with fewer college-educated residents, while general demand shocks go to larger, richer cities with more elastic housing supply and higher education levels.&lt;/p&gt;
&lt;p&gt;Q: Which industries and occupations drive DOD&amp;rsquo;s stronger employment effects for no-bachelor&amp;rsquo;s workers?
A: Within industries, DOD-induced employment gains for no-bachelor&amp;rsquo;s workers are strongest in construction and manufacturing, with much milder effects from general demand shocks in these industries. The occupations benefiting most are military occupations (broadly defined) and Production and Maintenance occupations, which rank among the lowest in occupational prestige for no-bachelor&amp;rsquo;s workers.&lt;/p&gt;
&lt;p&gt;Q: How does DOD spending compare to targeted social programs in achieving distributional goals?
A: The paper argues that although DOD spending is not designed as social policy, its effects on earnings for households without a bachelor&amp;rsquo;s degree, poverty reduction, disability reduction, homeownership, and occupational upgrading mirror the stated objectives of many targeted programs (job training, housing subsidies, SNAP, Medicaid). At the same time, DOD-induced life savings cost approximately $25–45 million per life, exceeding the typical value of a statistical life, so the mortality benefits cannot alone justify the spending.&lt;/p&gt;
&lt;p&gt;Local DOD earnings multiplier: The dollar amount of earnings for a demographic group produced by a dollar of local DOD spending over a two-year period, estimated using a two-year differenced panel regression with CBSA and time fixed effects, instrumented by a Bartik-type shock.&lt;/p&gt;
&lt;p&gt;Bartik-type IV shock: An instrumental variable constructed as the product of a location&amp;rsquo;s average share of DOD contract spending and aggregate contract spending in a given period; used to isolate the component of DOD contracts associated with new production rather than anticipated or smoothed payments.&lt;/p&gt;
&lt;p&gt;General demand shock: A Bartik shift-share shock constructed from local industry employment shares and national industry-level growth rates across all private-sector industries, used as a comparison series to evaluate whether DOD spending effects are generic or specific to defense contracts (correlation with DOD shock: -0.07).&lt;/p&gt;
&lt;p&gt;Extensive margin of employment: The change in the employment rate (entry from unemployment or non-participation into employment) as distinct from hours or wage adjustments among the already-employed; identified in the paper as the primary mechanism linking DOD shocks to differential social outcomes for no-bachelor&amp;rsquo;s households.&lt;/p&gt;
&lt;p&gt;Deaths of despair: Drug-and-alcohol-related deaths and deaths by suicide, following Case and Deaton (2020); examined here at higher frequency as an outcome of labor market earnings changes induced by aggregate demand stimulus.&lt;/p&gt;
&lt;p&gt;Occupational prestige (Siegel prestige score): A summary measure of job quality based on survey-derived perceptions of occupational standing (Siegel 1971), aggregated to the CBSA level by demographic group; used as a measure of upward job-ladder mobility in response to demand stimulus.&lt;/p&gt;
&lt;p&gt;Source text origin: A classification of the text basis for a paper summary — full PDF or OA-HTML versus abstract-only; the pipeline hard-blocks summaries derived solely from abstract text.&lt;/p&gt;</description></item><item><title>Designing Dynamic Reassignment Mechanisms: Evidence from GP Allocation</title><link>https://macropaperwarehouse.com/papers/designing-dynamic-reassignment-mechanisms-evidence-from-gp-allocation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/designing-dynamic-reassignment-mechanisms-evidence-from-gp-allocation/</guid><description>&lt;p&gt;This paper studies the design of dynamic reassignment mechanisms—centralized systems that must not only provide good initial matches but also accommodate changes in agents&amp;rsquo; preferences over time. The empirical setting is Norway&amp;rsquo;s system for allocating patients to general practitioners (GPs), where every individual is assigned a specific GP whose panel has a binding capacity cap. Since 2016, Norway has allowed patients to join waitlists for oversubscribed GPs while retaining their spot on their current GP&amp;rsquo;s panel, with reassignment proceeding strictly first-come, first-served (FCFS) as vacancies arise.&lt;/p&gt;
&lt;p&gt;The paper makes three contributions. First, it provides direct evidence of unrealized gains from trade: in December 2019, 15 percent of the 133,332 patients then standing on waitlists could have been immediately reassigned via a single run of the Top-Trading Cycles (TTC) algorithm, which identifies not only bilateral swaps but arbitrary cycles. A mechanical simulation holding patient choices fixed shows that running TTC monthly from November 2016 through December 2019 would have left 23 percent fewer patients on waitlists by end-2019, with average waiting times among reassigned patients 29 percent shorter.&lt;/p&gt;
&lt;p&gt;Second, the paper introduces a dynamic TTC mechanism and clarifies why static properties do not carry over. In the static case, TTC is both strategy-proof and Pareto-improving (Shapley and Scarf, 1974; Roth, 1982). In a dynamic setting, neither property holds. Repeated TTC is not strategy-proof because patients&amp;rsquo; GP choices affect how long they wait. More importantly, TTC may leave some patients worse off: a panel slot that would have gone to the first person on a waitlist under FCFS may instead go to a later-arriving patient who can form a trading cycle, effectively de-prioritizing patients whose GPs are undersubscribed. In the mechanical simulation, 4.5 percent of patients face longer waiting times under TTC.&lt;/p&gt;
&lt;p&gt;Third, the paper estimates a structural model of patient attention and GP choice using monthly Norwegian administrative data covering 4.78 million patients and 6,470 GP panels (2014–2019), restricting estimation to the Trondelag region (approximately 8 percent of the country). The model specifies: a Poisson attention process (patients consider switching only when an attention shock arrives); preferences over GPs as a function of travel time, GP fixed effects, and match characteristics; and a belief model mapping observed waitlist lengths into expected waiting times. Parameters are recovered via a Gibbs sampler with Metropolis-Hastings for the discount rate. Key estimates: the annual discount factor is approximately 0.91; a female patient under 45 would travel 7.3 minutes farther to see a female GP (6.3 minutes for a female patient over 45); GP fixed effects have a standard deviation of 31 minutes&amp;rsquo; travel-time equivalent; idiosyncratic taste shocks have a standard deviation of 12.6 minutes.&lt;/p&gt;
&lt;p&gt;The paper then simulates a stationary equilibrium for each counterfactual mechanism. Under the status quo in stationary equilibrium, 9.4 percent of patients are on a waitlist, 82.2 percent of GPs have a waitlist, and average expected waiting time is 16.7 months. Introducing TTC reduces average waiting time to 14.1 months and raises mean patient welfare by the equivalent of 0.75 minutes&amp;rsquo; travel time (more than 13 percent of the gain achievable under a no-capacity-constraints benchmark). Over half of this gain (0.4 minutes) comes directly from patients obtaining geographically closer GPs. Benefits are concentrated among younger patients, female patients, and recent movers; rural patients gain 2.1 minutes. However, patients with undersubscribed GPs face waiting times that rise from 16.7 to 22.8 months and are worse off by the perpetuity equivalent of 0.8 minutes.&lt;/p&gt;
&lt;p&gt;Two modified mechanisms are evaluated. Deferred Acceptance (DA), which strictly respects FCFS priority, achieves essentially no improvement over the status quo, illustrating a fundamental trade-off between eliminating envy and exploiting gains from trade. A &amp;ldquo;TTC with Priority&amp;rdquo; (TTCP) mechanism, which gives priority for panel vacancies to patients with undersubscribed GPs before running TTC, achieves 61 percent of TTC&amp;rsquo;s welfare gains (0.46 minutes flow payoff; 1.08 minutes NPV) while leaving patients with undersubscribed GPs no worse off than under the status quo. A benchmark simulation eliminating waitlists altogether raises mean welfare slightly (0.19 minutes) but lowers median welfare (−0.60 minutes), with gains concentrated among highly mismatched patients.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the core market failure the paper documents?&lt;/strong&gt;
A: Norway&amp;rsquo;s waitlist mechanism assigns panel vacancies strictly first-come, first-served without allowing patients to trade. This creates a &amp;ldquo;double coincidence of wants&amp;rdquo; problem: patients can simultaneously be on each other&amp;rsquo;s waitlists but cannot swap. In December 2019, 15 percent of 133,332 waiting patients could have been immediately reassigned via a single TTC run. A mechanical simulation shows that monthly TTC would have left 23 percent fewer patients on waitlists by end-2019 and reduced average realized waiting times among reassigned patients by 29 percent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does TTC fail to be strategy-proof in a dynamic setting?&lt;/strong&gt;
A: In the static case, TTC gives every agent an assignment at least as good as their endowment, making truthful reporting a dominant strategy. In a dynamic setting, a patient&amp;rsquo;s choice of GP determines not only which GP they receive but also how long they wait — patients who choose less-demanded GPs reach the front of the waitlist faster. This creates incentives to misreport preferences strategically, breaking strategy-proofness. The paper shows this formally and builds it into the equilibrium model by requiring patients to optimize over both GP choice and expected waiting time.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does dynamic TTC harm some patients relative to the status quo?&lt;/strong&gt;
A: Under FCFS, the first person on a waitlist is guaranteed the next available slot on the target GP&amp;rsquo;s panel. Under TTC, a patient who arrived later but whose current GP is oversubscribed can form a trading cycle that redirects that slot, effectively jumping the queue. Patients with undersubscribed GPs — whose panel endowment is not a scarce resource that others want — cannot form cycles and are systematically de-prioritized. In the stationary equilibrium, their expected waiting time rises from 16.7 to 22.8 months, and they are worse off by the perpetuity equivalent of 0.8 minutes&amp;rsquo; travel time.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the main parameter estimates and what do they imply?&lt;/strong&gt;
A: The annual discount factor is estimated at approximately 0.91 once GP fixed effects are included (rising to near 0.95 without them, because more desirable GPs have longer waitlists). Gender homophily is worth 6.3–7.3 minutes of travel time for female patients under 45. Age homophily is worth approximately 1 minute. The standard deviation of GP fixed effects is 31 minutes and idiosyncratic shocks are 12.6 minutes, both in travel-time equivalents, indicating substantial horizontal differentiation across GPs and across patients&amp;rsquo; idiosyncratic tastes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How important are moves as a driver of GP switching?&lt;/strong&gt;
A: Moves are the dominant driver. Among non-movers, older men consider switching just once every 25 years; temporary residents consider switching approximately once every 7.5 years (1.084 percent per month). Among patients who moved more than 30 minutes, a temporary resident has an 18.59 percent monthly probability of considering switching in the month of or month after the move. For a permanent resident making a long-distance move, the cumulative attention probability over the 8 months surrounding the move rises to 34 percent (versus 22 percent for a short-distance move). In the data, 26 percent of waitlist users moved municipality during 2017–2019, versus 6 percent of non-switchers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the stationary equilibrium under the status quo look like?&lt;/strong&gt;
A: In the long-run stationary equilibrium, 9.4 percent of patients are on a waitlist, 82.2 percent of GPs have a waitlist, and the average expected waiting time to switch GPs is 16.7 months. Each month, 2,299 patients on average draw attention shocks; 85.2 percent of these choose to join a waitlist, while the remainder either switch to an open GP or stay with their current GP. The average attentive patient expects to successfully obtain their chosen GP after 16.8 months.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the distributional consequences of TTC across patient subgroups?&lt;/strong&gt;
A: Female patients benefit especially because they are more likely to be attentive (and thus use waitlists) than males. Recent movers gain 2.3 minutes&amp;rsquo; travel-time equivalent. Patients who have never moved still gain 1.0 minutes. Rural patients gain 2.1 minutes (larger than average), reflecting their longer baseline travel times and greater geographic mismatch potential. Urban patients also benefit but less so. The one group that is harmed is patients with undersubscribed GPs, who face longer waits and a welfare loss of 0.8 minutes perpetuity equivalent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does the Deferred Acceptance mechanism fail to improve on the status quo?&lt;/strong&gt;
A: DA strictly respects FCFS waiting-time priority: no patient may be reassigned to a GP for whom another patient has been waiting longer. This means DA can only execute swaps in which all patients ahead of each participant on their respective waitlists are also reassigned in the same month. In practice, this virtually never occurs, so DA reassigns almost no patients earlier than the status quo Waitlists mechanism. The result illustrates a fundamental trade-off: fully respecting FCFS priority eliminates nearly all gains from trade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does TTCP restore fairness while preserving most of the efficiency gains?&lt;/strong&gt;
A: TTCP modifies TTC by prioritizing patients with undersubscribed GPs over those with oversubscribed GPs when assigning panel vacancies, while still respecting the constraint that patients cannot be assigned a GP they prefer less than their current one. This gives patients with undersubscribed GPs a compensating advantage in the queue that offsets their inability to trade via cycles. TTCP achieves 0.46 minutes&amp;rsquo; mean flow payoff improvement versus 0.75 for TTC (61 percent of TTC&amp;rsquo;s gains), and an NPV measure of 1.08 minutes versus 1.25 for TTC. Patients with undersubscribed GPs are left no worse off than under the status quo.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What happens when waitlists are eliminated entirely?&lt;/strong&gt;
A: Under No Waitlists, attentive patients may only choose among GPs with open panels at the moment of attention. Mean welfare rises slightly (0.19 minutes) because patients spend less time mismatched while waiting, but median welfare falls by 0.60 minutes. The gains are concentrated among a minority of highly mismatched patients who prefer limited choice with no waiting over broader choice with long waits, while most patients prefer the option to wait for a more preferred GP. The authors note this may partly explain why formal waitlists are rare in other primary care systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the welfare benchmark and how large are the gains?&lt;/strong&gt;
A: The benchmark is a &amp;ldquo;No Caps&amp;rdquo; scenario in which all panel caps are removed, representing the maximum achievable improvement. The mean welfare gain from TTC (0.75 minutes) represents more than 13 percent of this upper bound. The &amp;ldquo;Truthful TTC&amp;rdquo; benchmark, where patients submit full preference lists, yields 1.04 minutes, but its gains are also concentrated: the median patient is no better off than under the status quo Waitlists mechanism.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the scope conditions for these findings?&lt;/strong&gt;
A: The demand model is estimated on the Trondelag region of Norway (approximately 8 percent of the national population) over 2017–2019, a period when waitlists were growing rapidly rather than in steady state. Counterfactual comparisons are made in a stationary equilibrium calibrated to Trondelag. The model excludes patients under 16 (whose enrollment is managed by parents). The partially capitated payment structure and fixed panel caps are institutional features specific to Norway, though similar systems exist in Canada, the UK, Italy, and Sweden. GP characteristics are held fixed in the model. The analysis abstracts from health outcomes, focusing on preference-based welfare from GP assignment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Top-Trading Cycles (TTC) algorithm&lt;/strong&gt;: A centralized reassignment algorithm that takes agents&amp;rsquo; preference lists and objects&amp;rsquo; priority lists as inputs, has each agent &amp;ldquo;point to&amp;rdquo; their preferred object and each object &amp;ldquo;point to&amp;rdquo; their highest-priority current or waiting agent, identifies cycles of mutual pointing, and executes the trades in those cycles simultaneously. In the paper&amp;rsquo;s static application, TTC is both Pareto-improving (every participant receives an assignment at least as good as their endowment) and strategy-proof. In the dynamic setting studied here, neither property holds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic TTC mechanism&lt;/strong&gt;: A mechanism that runs the TTC algorithm repeatedly at the end of each period after naturally arising vacancies have been filled from waitlists. Because patients&amp;rsquo; GP choices affect how long they wait — not only which GP they receive — this mechanism is not strategy-proof and may leave patients with undersubscribed GPs worse off than under strictly FCFS waitlists.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;TTC with Priority (TTCP)&lt;/strong&gt;: A modified version of dynamic TTC that changes the priority ordering so that patients with undersubscribed current GPs are prioritized above patients with oversubscribed GPs when panel vacancies are allocated. This modification preserves patients&amp;rsquo; endowment rights but compensates the group harmed by standard TTC. In the paper&amp;rsquo;s simulations, TTCP achieves 61 percent of TTC&amp;rsquo;s mean welfare gains while leaving patients with undersubscribed GPs no worse off than under the status quo.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Patient attention model&lt;/strong&gt;: A model in which patients consider switching GPs only when they receive a Poisson-distributed attention shock. Attention rates vary by observable characteristics (age, gender, temporary vs. permanent residency, whether and how far the patient recently moved). The model interprets any switch request as evidence of both an attention shock and a preference for the requested GP over the current one. Patients who do not request switches may be either inattentive or attentive but satisfied.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Horizontal differentiation (GP preference heterogeneity)&lt;/strong&gt;: The extent to which different patients prefer different GPs for reasons unrelated to overall GP quality — primarily driven by geographic proximity, gender homophily (worth 6.3–7.3 travel-time-equivalent minutes for young female patients), and age similarity (approximately 1 minute). Horizontal differentiation is the fundamental source of gains from trade: if all patients preferred the same GP, there would be no mutual-benefit swaps to find.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Deferred Acceptance (DA) algorithm&lt;/strong&gt;: The patient-proposing DA algorithm, which strictly respects FCFS waiting-time priority: no patient may be reassigned ahead of another patient who has been waiting longer for the same GP. In the dynamic context, DA achieves essentially no welfare improvement over the status quo because its strict respect for priority eliminates nearly all trading opportunities, illustrating the trade-off between envy-freeness and efficiency.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double coincidence of wants&lt;/strong&gt;: The situation in which two (or more) patients are simultaneously on each other&amp;rsquo;s waitlists and would mutually benefit from trading GP assignments, but cannot do so under the current mechanism because there is no vacancy on either panel. The paper&amp;rsquo;s direct evidence of this phenomenon — 15 percent of waiters could be immediately reassigned via one TTC run — motivates the counterfactual analysis.&lt;/p&gt;</description></item><item><title>Digital Distractions with Peer Influence</title><link>https://macropaperwarehouse.com/papers/digital-distractions-with-peer-influence/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/digital-distractions-with-peer-influence/</guid><description>&lt;p&gt;This paper estimates the causal effects of mobile app usage on college students&amp;rsquo; academic performance, physical health, and labor market outcomes, while separately identifying behavioral (endogenous) and contextual (exogenous) peer effects in app usage — the first study to do so within a unified empirical framework. The analysis draws on administrative data for three freshman cohorts (2018–2020) at a mid-tier Chinese university, linked to individual-level mobile phone usage records from a major telecommunications carrier covering 6,430 students over four years (excluding COVID semester). High-frequency GPS data, hourly app usage records for the 2020 cohort, and two waves of university surveys supplement the main dataset.&lt;/p&gt;
&lt;p&gt;The identification strategy addresses three challenges: endogeneity of own app usage, endogeneity of peer group formation, and the reflection problem in peer effects. For own usage, two instrumental variables are used: (1) a shift-share instrument interacting the September 2020 launch of the blockbuster game Yuanshen with students&amp;rsquo; pre-college app usage intensity; and (2) China&amp;rsquo;s October 2019 minors&amp;rsquo; game restriction policy (prohibiting under-18s from playing online games 10 p.m.–8 a.m. and capping weekday gaming at 90 minutes/day) interacted with the evolving number of underage pre-college friends. For peer effects, the university&amp;rsquo;s random dormitory assignment within gender-class units provides exogenous peer variation; behavioral peer effects are further isolated using the minors&amp;rsquo; restriction policy interacted with roommates&amp;rsquo; pre-college underage friend networks, an instrument that affects roommates but not the focal student. Contextual peer effects are recovered by subtracting the estimated behavioral component from reduced-form estimates.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. First, app usage is contagious: a one standard deviation (s.d.) increase in roommates&amp;rsquo; in-college total app usage raises a student&amp;rsquo;s own usage by 5.8% (IV). Behavioral peer effects dominate: contextual peer effects are small and statistically insignificant. Second, own app usage severely harms academic performance: a one s.d. increase in total app usage reduces GPA for required courses by 36.2% of a within-cohort-major s.d. (IV), and a one s.d. increase in game app usage alone reduces GPA by 56.6% of a within-cohort-major s.d. The direct disruption effect of roommates&amp;rsquo; app usage reduces GPA by a further 20.6% of a within-cohort-major s.d.; combining the indirect channel (behavioral contagion), the total roommate effect reaches 22.7% of a within-cohort-major s.d., more than 60% of the own-usage effect. Third, the effect on physical education scores is roughly four times larger than on required-course GPA: a one s.d. increase in own app usage reduces PE scores by 2.74 points, while roommates&amp;rsquo; app usage has no direct effect on PE. Fourth, a one s.d. increase in own in-college app usage reduces initial wages upon graduation by 2.3% (12.1% of within-cohort-major wage s.d.); a one s.d. increase in roommates&amp;rsquo; usage reduces wages by 0.9% directly, with a total effect (including the contagion channel) of approximately 1.0% (5.3% of within-cohort-major s.d.). Controlling for cumulative GPA reduces the gaming-to-wage coefficient by roughly one-third, indicating that academic performance is an important but partial mediator.&lt;/p&gt;
&lt;p&gt;A back-of-the-envelope policy simulation extending the minors&amp;rsquo; gaming cap (3 hours/week) to college students — binding for 34.3% of student-month observations — projects an average wage increase of 0.9% at graduation, approximately half the wage premium from one additional year of work experience in developing countries.&lt;/p&gt;
&lt;p&gt;Mechanism evidence from GPS data shows that Yuanshen&amp;rsquo;s launch caused students to arrive at study halls 18.2 minutes later and leave 23.4 minutes earlier per day. High-frequency sleep data show that a one s.d. increase in nighttime app usage reduces sleep duration by approximately 30 minutes and raises the probability of sleeping late by 34 percentage points. Survey evidence indicates that heavy app users recognize the addictive nature of gaming, pointing to self-control problems rather than lack of awareness.&lt;/p&gt;
&lt;p&gt;The scope conditions are: single mid-tier Chinese university; 2018–2020 cohorts; outcomes through initial job placement only; peer group restricted to dormitory roommates; findings rely on IV exclusion restrictions conditional on student and time fixed effects.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question?
A: The paper asks how individual and peer mobile app usage affect college students&amp;rsquo; academic performance, physical health, and early labor market outcomes, and it separately identifies the behavioral (endogenous) versus contextual (exogenous) components of peer influence in app usage. This is claimed as the first study to disentangle these two types of peer effects within a unified empirical framework.&lt;/p&gt;
&lt;p&gt;Q: What data does the paper use?
A: Administrative records for 7,479 undergraduates across three freshman cohorts (2018–2020) at a medium-sized mid-tier Chinese university are linked to monthly mobile app usage records from a telecommunications provider covering 75% of the provincial population; 6,430 students are matched. The dataset also includes GPS location data at 5-minute intervals, hourly app usage for the 2020 cohort (used to infer sleep), and two waves of voluntary annual surveys with 1,798 respondents (24% response rate). Labor market outcomes — employment status, wages, post-graduate admissions — are available for the 2018 and 2019 cohorts.&lt;/p&gt;
&lt;p&gt;Q: How does the paper address the endogeneity of own app usage?
A: Two sets of instruments are used. The first interacts the September 2020 launch of Yuanshen (the most popular game in China, with over 13 million Chinese users by 2021, the majority under age 25) with students&amp;rsquo; pre-college app usage, forming a shift-share instrument under the assumption that the game launch is orthogonal to unobserved GPA determinants conditional on student fixed effects. The second interacts China&amp;rsquo;s October 2019 minors&amp;rsquo; game restriction policy with the evolving count of a student&amp;rsquo;s underage pre-college friends; event studies confirm no pre-trends and a sharp, transitory drop in app usage post-policy that dissipates as friends age out of the restricted group.&lt;/p&gt;
&lt;p&gt;Q: How does the paper solve the reflection problem and separate behavioral from contextual peer effects?
A: Three-step procedure: (1) random dormitory assignment within gender-class units yields reduced-form peer effect estimates using roommates&amp;rsquo; pre-college app usage as the exogenous peer shifter; (2) behavioral peer effects are isolated via an IV using the minors&amp;rsquo; restriction policy interacted with roommates&amp;rsquo; (not the focal student&amp;rsquo;s) underage pre-college friend networks — an instrument that shifts roommates&amp;rsquo; app usage but is orthogonal to the focal student&amp;rsquo;s outcomes; (3) contextual peer effects are recovered as the residual from subtracting the estimated behavioral effect from the reduced-form estimate.&lt;/p&gt;
&lt;p&gt;Q: How large and significant are the behavioral versus contextual peer effects in app usage?
A: A one s.d. increase in roommates&amp;rsquo; in-college total app usage raises own usage by 5.8% (IV estimate, significant). For game apps alone the behavioral spillover is 10.7%, and for games plus video it is 6.5%. Contextual peer effects (identified from roommates&amp;rsquo; pre-college characteristics) are much smaller and statistically insignificant, indicating that peer influence operates primarily through the direct imitation of peers&amp;rsquo; actions rather than their background traits.&lt;/p&gt;
&lt;p&gt;Q: What is the effect of own app usage on GPA?
A: The IV estimate shows a one s.d. increase in total in-college app usage reduces GPA for required courses by 0.716 points, equivalent to 36.2% of a within-cohort-major GPA s.d. (significant at 1%). For game apps alone, a one s.d. increase reduces GPA by 1.119 points, or 56.6% of a within-cohort-major s.d. OLS estimates are biased toward zero, likely because negative health shocks reduce both GPA and app usage simultaneously.&lt;/p&gt;
&lt;p&gt;Q: How large is the total peer effect of roommates&amp;rsquo; app usage on a student&amp;rsquo;s GPA?
A: Roommates&amp;rsquo; app usage directly lowers GPA by 0.408 points (20.6% of within-cohort-major s.d.) through disruption of the dormitory study environment or crowding out of group study. The behavioral contagion channel (5.8% increase in own usage per s.d. of roommates&amp;rsquo; usage) adds an additional 0.042 points, bringing the total effect to approximately 0.450 points, or 22.7% of a within-cohort-major s.d. — over 60% of the own-usage effect.&lt;/p&gt;
&lt;p&gt;Q: What is the effect on physical education (PE) scores, and why do roommates&amp;rsquo; app usage not matter there?
A: A one s.d. increase in own total app usage reduces PE scores by 2.74 points (IV), approximately four times the magnitude of the effect on required-course GPA, consistent with health literature on excessive screen time. Roommates&amp;rsquo; app usage has no statistically significant direct effect on PE, which the authors attribute to the irrelevance of dormitory noise and study disruptions for outdoor physical activity.&lt;/p&gt;
&lt;p&gt;Q: What are the effects of app usage on wages at graduation?
A: Doubling total app usage during college reduces initial wages by approximately 2% (IV). A one s.d. increase in own usage reduces wages by 2.3%, or 12.1% of a within-cohort-major wage s.d. A one s.d. increase in roommates&amp;rsquo; usage directly reduces wages by 0.9% (4.8% of within-cohort-major s.d.); including the behavioral contagion channel, the total roommate effect is approximately 1.0% (5.3% of within-cohort-major s.d.). Controlling for cumulative GPA reduces the game-usage-to-wage coefficient by about one-third, implying GPA is a partial but not complete mediator.&lt;/p&gt;
&lt;p&gt;Q: What does the policy simulation of the gaming cap say?
A: Extending the minors&amp;rsquo; game restriction (3 hours/week cap) to college students would bind for 34.3% of student-month observations, reducing average monthly gaming from 12.1 hours to 8 hours (a one-third decrease). Incorporating the behavioral peer multiplier for gaming (0.078), average gaming further converges to approximately 7.65 hours in steady state. The implied wage gain at graduation is 0.9%, approximately half the wage premium from one additional year of work experience in developing countries (Lagakos et al., 2019 estimate).&lt;/p&gt;
&lt;p&gt;Q: What does the GPS evidence show about time allocation?
A: Following Yuanshen&amp;rsquo;s launch, the average student arrives at the study hall 18.2 minutes later and returns to the dormitory 23.4 minutes earlier per day. The minors&amp;rsquo; restriction reverses this: students with the average number of minor friends arrive at study halls 17.4 minutes earlier and return to the dorm 19.8 minutes later. Both game shocks also shift tardiness and absence rates for major-required courses in the expected directions, and the effects intensify over time with Yuanshen&amp;rsquo;s growing popularity.&lt;/p&gt;
&lt;p&gt;Q: What do the sleep data show?
A: A one s.d. increase in nighttime app usage (9 p.m.–3 a.m.) is associated with roughly 30 minutes less sleep (7% of the mean), a 34 percentage point higher probability of sleeping late, and a 4.5 percentage point higher probability of waking up late. Daytime app usage (8 a.m.–9 p.m.) is also associated with 7.2 fewer minutes of sleep (1.8% of mean) and a 3.7 percentage point higher probability of late wake-up. These results are descriptive (from the 2020 cohort hourly data) rather than IV-based.&lt;/p&gt;
&lt;p&gt;Q: What does the survey evidence show about mechanisms and self-awareness?
A: Heavier app users report worse physical health and higher stress, are less likely to have obtained professional certifications by graduation, submit fewer job applications, and express lower satisfaction with job offers. Notably, heavier users are more likely to acknowledge the addictive nature of apps and games, suggesting a self-control problem rather than informational deficiency. They also report better relationships with roommates and greater likelihood of following roommates&amp;rsquo; advice on post-graduation choices, a potential direct channel for peer labor market effects.&lt;/p&gt;
&lt;p&gt;Q: How representative is the sample, and what are the key scope conditions?
A: The university is a mid-tier institution in southern China with students predominantly from the 30th–80th CEE score percentile among provincial college-admitted applicants; it is less female (42% vs. 53% nationally) and more rural (40% vs. 27% nationally). Survey respondents oversample less advantaged backgrounds and are re-weighted. Findings pertain to dormitory roommates as the peer group; all labor market outcomes are initial wages upon graduation; the sample covers 2018–2021 with COVID semester excluded. The peer effects estimates rest on random dormitory assignment, which the authors verify by showing no within-dorm correlation in pre-college characteristics.&lt;/p&gt;
&lt;p&gt;Behavioral (endogenous) peer effects: The mechanism by which a peer&amp;rsquo;s actual behavior — here, contemporaneous app usage — directly influences a focal individual&amp;rsquo;s own behavior. In this paper, identified via IV using the minors&amp;rsquo; game restriction policy interacted with roommates&amp;rsquo; underage pre-college friend networks, which shifts roommates&amp;rsquo; usage but not the focal student&amp;rsquo;s characteristics.&lt;/p&gt;
&lt;p&gt;Contextual (exogenous) peer effects: The influence of peers&amp;rsquo; pre-determined background characteristics (e.g., pre-college app usage, reflecting motivation, study habits, attitudes toward academics) on a focal individual&amp;rsquo;s outcomes, independent of peers&amp;rsquo; actual in-college behavior. Recovered as the residual after subtracting estimated behavioral peer effects from reduced-form estimates; found to be small and insignificant in this setting.&lt;/p&gt;
&lt;p&gt;Shift-share instrument (Yuanshen): A quasi-experimental instrument constructed by interacting the mid-sample launch date of the blockbuster game Yuanshen (September 2020) with students&amp;rsquo; pre-college app usage intensity, under the assumption that pre-college usage predicts differential susceptibility to the shock while the launch itself is orthogonal to the university&amp;rsquo;s academic environment.&lt;/p&gt;
&lt;p&gt;Minors&amp;rsquo; game restriction policy: China&amp;rsquo;s October 2019 policy prohibiting individuals under 18 from playing online games between 10 p.m. and 8 a.m. and capping weekday gaming at 90 minutes per day (tightened to 3 hours/week in September 2021). Used both as an instrument for own app usage (via underage pre-college friends) and as an instrument for roommates&amp;rsquo; usage (via roommates&amp;rsquo; underage friends) to isolate behavioral peer effects.&lt;/p&gt;
&lt;p&gt;Reflection problem: The identification challenge first articulated by Manski (1993) arising because an individual&amp;rsquo;s behavior both affects and is affected by peers simultaneously, making it impossible to separately identify the direction of influence from observational data without exogenous variation in peer behavior.&lt;/p&gt;
&lt;p&gt;Source text origin: The paper&amp;rsquo;s own data provenance category distinguishing whether summaries are based on full working paper text (pdf or oa-html) versus abstract only — a distinction the paper itself does not use but that is relevant to the review pipeline running this analysis.&lt;/p&gt;
&lt;p&gt;Within-cohort-major GPA standard deviation: The unit used to scale all GPA effect sizes, defined as the standard deviation of GPA within students of the same graduation cohort and declared major. This normalization accounts for systematic differences in grading across fields and years, making effect magnitudes comparable across specifications.&lt;/p&gt;</description></item><item><title>Distributional Growth Accounting: Education and the Reduction of Global Poverty, 1980–2019</title><link>https://macropaperwarehouse.com/papers/distributional-growth-accounting-education-and-the-reduction-of-global-poverty-19802019/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/distributional-growth-accounting-education-and-the-reduction-of-global-poverty-19802019/</guid><description>&lt;h2 id="layer-1--core-argument"&gt;Layer 1 — Core Argument&lt;/h2&gt;
&lt;p&gt;This paper constructs the first estimates of the aggregate and distributional effects of worldwide educational expansion since 1980 by developing a &amp;ldquo;distributional growth accounting&amp;rdquo; framework that isolates the contribution of schooling to economic growth by income group. The framework integrates the canonical labor supply-and-demand model of education and the wage structure (à la Goldin and Katz 2007) with standard growth accounting tools, applied to a new microdatabase covering household surveys in 150 countries and representative of approximately 95% of the world&amp;rsquo;s population, alongside new country-specific estimates of private returns to primary, secondary, and tertiary schooling. Under conservative assumptions — relying on standard Mincerian returns, assuming capital income is unaffected by schooling, and abstracting from human capital externalities — education can account for approximately 50% of global economic growth, 70% of income gains among the world&amp;rsquo;s poorest 20% of individuals, and 40% of extreme poverty reduction since 1980; it also explains over 50% of improvements in the share of labor income accruing to women. A key mechanism is imperfect substitutability between skill groups: as educational expansion raises the supply of skilled workers, their relative wage falls, redistributing income toward low-skilled workers and amplifying education&amp;rsquo;s equalizing effect at the bottom of the distribution — a channel that canonical cross-country growth accounting misses, causing it to underestimate education&amp;rsquo;s contribution to poverty reduction by a factor of approximately three. Combining these indirect investment benefits from education with direct government redistribution (from a companion paper) brings the total contribution of public policies to extreme poverty reduction to at least 50%.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-q-what-is-distributional-growth-accounting-and-how-does-it-differ-from-standard-growth-accounting"&gt;Q1. Q: What is distributional growth accounting and how does it differ from standard growth accounting?&lt;/h3&gt;
&lt;p&gt;A: Standard growth accounting (as in Barro and Lee 2015) combines cross-country data on average years of schooling with a uniform return to derive a counterfactual average income absent educational progress. Distributional growth accounting instead starts from microdata on the joint distribution of income and education within 150 countries, constructs income-group-specific counterfactuals, and accounts for both direct wage effects on individuals whose education changed and general equilibrium supply effects that alter relative wages across all workers. The standard approach is found to underestimate education&amp;rsquo;s contribution to the poorest 20%&amp;rsquo;s income growth by a factor of roughly three (23% vs. 71% in the benchmark specification), because cross-country averages cannot accurately locate the world&amp;rsquo;s poorest individuals and because two key channels — labor income shares being greater at the bottom, and supply-side wage redistribution — are omitted.&lt;/p&gt;
&lt;h3 id="q2-q-how-is-the-counterfactual-world-income-distribution-constructed"&gt;Q2. Q: How is the counterfactual world income distribution constructed?&lt;/h3&gt;
&lt;p&gt;A: In five steps applied to the 150-country microdata. First, education levels are downgraded within each survey until matching the 1980 distribution of educational attainment (using the Barro–Lee database), prioritizing individuals closest to the target level. Second, the earnings of downgraded workers are reduced using the &amp;ldquo;true&amp;rdquo; return to schooling, which lies between the initial return (prevailing before expansion, computed from the CES production function using the 2019 elasticity) and the final return observed in 2019 — for plausible parameterizations, the true return weights initial returns at 50–70%. Third, relative wages are adjusted to reflect supply effects: the increase in skilled-worker supply lowers their relative wage by 1/σ log points per log-point increase in relative supply. Fourth, counterfactual labor income is combined with unchanged capital income to yield counterfactual total income. Fifth, the share of actual income growth attributable to education is computed as the gap between the actual and counterfactual growth rates, expressed as a fraction of actual growth.&lt;/p&gt;
&lt;h3 id="q3-q-what-role-does-imperfect-skill-substitution-play-and-how-is-σ-calibrated"&gt;Q3. Q: What role does imperfect skill substitution play, and how is σ calibrated?&lt;/h3&gt;
&lt;p&gt;A: Imperfect substitution between skill groups (elasticity σ in a CES production function) is the mechanism through which educational expansion redistributes income. When skilled-worker supply rises, their relative wage falls and low-skilled workers&amp;rsquo; relative wage rises, so the income gains from education are shared more broadly than individual returns alone would suggest. With perfect substitutes (σ → ∞), supply effects vanish and education&amp;rsquo;s distributional impact is determined entirely by who directly received schooling. The elasticity is calibrated from the recent macroeconomics literature; in sensitivity analysis, the paper bounds the contribution of education to the poorest 20%&amp;rsquo;s income growth between 60% and 90% across plausible values of σ and private returns.&lt;/p&gt;
&lt;h3 id="q4-q-why-are-the-estimates-described-as-conservative"&gt;Q4. Q: Why are the estimates described as conservative?&lt;/h3&gt;
&lt;p&gt;A: Three reasons, each biasing the estimates downward. First, standard Mincerian returns are used, which are systematically lower than causal estimates from natural experiments — a meta-analysis of 15 papers and the paper&amp;rsquo;s own quasi-experimental validation (India, Indonesia, United States) confirm this; if anything, the framework underestimates schooling&amp;rsquo;s benefits in those settings. Second, capital income is assumed unaffected by schooling, abstracting from potential effects on capital accumulation and returns. Third, human capital externalities — for which there is now substantial empirical evidence — are ignored entirely. These conservative choices are deliberate; relaxing them would increase all headline estimates.&lt;/p&gt;
&lt;h3 id="q5-q-how-does-skill-biased-technical-change-interact-with-the-education-contribution"&gt;Q5. Q: How does skill-biased technical change interact with the education contribution?&lt;/h3&gt;
&lt;p&gt;A: In the CES model, the return to schooling is increasing in the skill bias of technology (AH/AL): a higher skill bias raises the marginal product of skilled workers relative to unskilled, making schooling more profitable. The benchmark counterfactual holds technology fixed at its 2019 value and reduces education to its 1980 level. An alternative counterfactual would hold technology at its 1980 value and increase education to its 2019 level; the difference between these two exercises identifies the contribution of skill-biased technical change in amplifying the benefits of schooling. Because 1980 microdata on the world income distribution are unavailable, this decomposition can only be performed for the subsample of 33 countries with surveys around 2000; for that sample, skill-biased technical change accounts for 20–30% of the income benefits of schooling, meaning education would still have yielded large gains even absent technological progress.&lt;/p&gt;
&lt;h3 id="q6-q-what-do-the-quasi-experimental-validations-in-india-indonesia-and-the-united-states-show"&gt;Q6. Q: What do the quasi-experimental validations in India, Indonesia, and the United States show?&lt;/h3&gt;
&lt;p&gt;A: Three large-scale schooling policy interventions — a school construction program in India (studied in Khanna 2023), Indonesia&amp;rsquo;s INPRES program (Duflo 2001 and 2004), and US compulsory schooling laws (Acemoglu and Angrist 2000) — are used to externally validate the framework. Using regional variation in exposure to each program and rich microdata on the income distribution, the paper documents two findings: (1) educational expansion had large causal effects on aggregate regional incomes comparable in magnitude to individual returns estimated in the same contexts; and (2) all three policies disproportionately benefited low-income earners, substantially reducing inequality. The distributional growth accounting framework reproduces both findings with &amp;ldquo;a remarkable degree of accuracy,&amp;rdquo; and if anything underestimates the benefits of schooling, providing validation of the methodological foundation.&lt;/p&gt;
&lt;h3 id="q7-q-how-does-the-paper-quantify-educations-role-in-gender-inequality-reduction"&gt;Q7. Q: How does the paper quantify education&amp;rsquo;s role in gender inequality reduction?&lt;/h3&gt;
&lt;p&gt;A: The framework is extended to gender by constructing a counterfactual for how large gender labor income gaps would be absent educational improvement since the early 1990s (the period for which female labor income share data are available). The counterfactual accounts for three gender-specific channels: differential educational expansion between men and women, heterogeneous returns to schooling by gender, and differential effects of schooling on female labor force participation. Comparing the counterfactual to actual trends in female labor income shares, education can explain 50–80% of the observed reductions in gender inequality, depending on specification and world region.&lt;/p&gt;
&lt;h3 id="q8-q-how-do-public-policies-as-a-whole-contribute-to-extreme-poverty-reduction"&gt;Q8. Q: How do public policies as a whole contribute to extreme poverty reduction?&lt;/h3&gt;
&lt;p&gt;A: The paper&amp;rsquo;s estimate of education&amp;rsquo;s indirect investment benefits (40% of extreme poverty reduction) is combined with a companion paper&amp;rsquo;s (Gethin 2023) estimates of direct government redistribution — cash and in-kind transfers together accounting for approximately 30% of global poverty reduction since 1980, with in-kind transfers alone accounting for approximately 20%. Because the two contributions overlap (e.g., public education spending is both an indirect investment benefit and an in-kind transfer), the combined lower bound is reported as &amp;ldquo;at least 50%&amp;rdquo; of extreme poverty reduction attributable to public policies.&lt;/p&gt;
&lt;h3 id="q9-q-why-does-the-distributional-approach-yield-such-different-results-from-the-standard-approach-for-the-poorest-20"&gt;Q9. Q: Why does the distributional approach yield such different results from the standard approach for the poorest 20%?&lt;/h3&gt;
&lt;p&gt;A: Two main reasons. First, cross-country data cannot accurately measure the incomes of the world&amp;rsquo;s poorest, because the poorest individuals are not all concentrated in the poorest countries — distributional accounting within countries is necessary to locate them precisely. Second, the standard approach misses two progressive channels: (a) labor income shares are higher at the bottom of the income distribution than average, so gains from schooling translate into larger income increases for the poor; and (b) supply effects redistribute schooling gains from high-skilled to low-skilled workers, a mechanism that is entirely absent from cross-country averages but directly captured in the microdata-based counterfactual.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Distributional growth accounting:&lt;/strong&gt; A framework, introduced in this paper, that combines a model of education and the wage structure with household microdata to construct income-group-specific counterfactuals, isolating the contribution of human capital accumulation to growth at each point of the income distribution rather than at the national-average level.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;True return to schooling (r&lt;/em&gt;):&lt;/em&gt;* In the CES framework with imperfect skill substitution, the &amp;ldquo;true&amp;rdquo; aggregate return to schooling used in the counterfactual lies strictly between the initial return (prevailing before educational expansion, counterfactually higher because skilled-worker supply was lower) and the final return (observed after expansion, lower due to skill-supply pressure). The true return is the return that equates the model&amp;rsquo;s predicted output loss to the actual output loss from reducing education; for plausible parameters it weights initial returns at 50–70%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Supply effects (general equilibrium effects of schooling):&lt;/strong&gt; When the supply of skilled workers rises, their relative wage falls and the relative wage of unskilled workers rises. These wage adjustments are not captured by individual-level Mincerian returns but are modeled via the CES elasticity of substitution σ. Supply effects are central to education&amp;rsquo;s progressive distributional impact: they compress the skill premium and raise earnings at the bottom of the distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Imperfect substitution between skill groups:&lt;/strong&gt; The CES production specification in which skilled (H) and unskilled (L) labor are combined with elasticity σ &amp;lt; ∞. This governs the magnitude of general equilibrium wage effects: a lower σ means a larger wage compression per unit of skilled-supply increase, amplifying the redistributive role of education. The paper calibrates σ from the macroeconomics literature and bounds results over plausible ranges.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Skill-biased technical change (SBTC):&lt;/strong&gt; Technology that raises the marginal product of skilled workers relative to unskilled (captured by the ratio AH/AL in the CES production function). SBTC amplifies returns to schooling; in the subsample of 33 countries with around-2000 surveys, SBTC accounts for 20–30% of schooling&amp;rsquo;s income benefits, but education would still have generated substantial income gains absent SBTC.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conservative assumptions (scope condition):&lt;/strong&gt; All headline quantitative results (50% of aggregate growth, 70% of poorest-20% income gains, 40% of extreme poverty reduction, &amp;gt;50% of gender inequality reduction) are explicitly conditioned on conservative assumptions: Mincerian rather than causal returns, no effect on capital income, and no human capital externalities. The paper argues these assumptions bias all estimates downward.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Summary based on HAL working paper (halshs-04423765v1, Working Paper 2023/25, November 2023). Period covered in working paper text: 1980–2022. AI-assisted, human review pending.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>Diversifying Society's Leaders? Determinants and Causal Effects of Admission</title><link>https://macropaperwarehouse.com/papers/diversifying-societys-leaders-determinants-and-causal-effects-of-admission/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/diversifying-societys-leaders-determinants-and-causal-effects-of-admission/</guid><description>&lt;p&gt;This paper studies why children from high-income families are more likely to attend Ivy-Plus colleges (Ivy League, Stanford, MIT, Duke, Chicago — 12 colleges total) and whether attending these colleges causally improves post-college outcomes. The authors construct a de-identified panel dataset linking federal income tax records, Department of Education college attendance data, College Board and ACT test scores, and application and admissions records from several Ivy-Plus and flagship public colleges covering approximately 2.4 million students across entering classes from 1998–2015.&lt;/p&gt;
&lt;p&gt;The central finding on the input side is that students from families in the top 1% of the income distribution (income above $611,000) are 2.3 times more likely to attend an Ivy-Plus college than middle-class students (defined as the 70th–80th percentiles of the national parental income distribution, approximately $91,000–$114,000) with comparable SAT/ACT scores. Two-thirds of this gap is attributable to higher admissions rates at Ivy-Plus colleges for high-income applicants; conditional on SAT/ACT scores, top-1% applicants are 58% more likely to be admitted than middle-class applicants. The remaining third splits between differences in application rates (roughly 20% of the total attendance gap) and matriculation rates (roughly 12%). In contrast, admissions rates at flagship public colleges are essentially uncorrelated with parental income conditional on test scores.&lt;/p&gt;
&lt;p&gt;Three admissions practices drive the high-income admissions advantage at Ivy-Plus colleges. First, legacy preferences: legacy applicants from the top 1% are admitted at more than five times the rate of non-legacy applicants with comparable test scores, demographics, and admissions ratings; children of alumni of a given Ivy-Plus college are not more likely to be admitted to other Ivy-Plus colleges, confirming that legacy status is not merely a proxy for unobservable credentials. Legacy preferences account for 52 of the estimated 168 &amp;ldquo;extra&amp;rdquo; top-1% students per average Ivy-Plus class (enrollment ~1,650). Second, non-academic ratings: students from the top 1% have markedly stronger non-academic credentials (extracurricular activities, leadership ratings) partly because they disproportionately attend private high schools whose students receive higher non-academic ratings despite no higher academic ratings; this accounts for 35 additional extra top-1% students. Third, athletic recruitment: the share of recruited athletes rises from 5% among admitted students from the bottom 60% to 13% among those from the top 1%, accounting for 27 additional extra top-1% students.&lt;/p&gt;
&lt;p&gt;On the output side, the authors estimate causal effects of attending an Ivy-Plus college using a new research design based on waitlisted applicants. The key identification assumption is that idiosyncratic variation in admissions decisions across waitlisted applicants at one Ivy-Plus college is uncorrelated with admissions decisions at other Ivy-Plus colleges — which the authors verify empirically. Under this assumption, comparisons of admitted vs. rejected waitlisted applicants identify causal effects for marginal students. The marginal student who attends an Ivy-Plus college instead of the average flagship public is approximately 50% more likely to reach the top 1% of the earnings distribution at age 33, nearly twice as likely to attend a highly-ranked graduate school, and 2.5 times as likely to work at a prestigious firm. Attending an Ivy-Plus college increases mean earnings by $101,000 at age 33 relative to a counterfactual mean of $143,000 at state flagships. Effects are concentrated in the upper tail of earnings — the impact on reaching the top quartile is small and statistically insignificant, while impacts on reaching the top 1% far exceed what a constant percentage treatment effect would predict. Effects are larger for students with weaker fallback options (i.e., whose home-state colleges channel fewer students to the top 1%).&lt;/p&gt;
&lt;p&gt;Critically, the three credentials driving the high-income admissions advantage — legacy status, athletic recruitment, and high non-academic ratings — are uncorrelated with or negatively correlated with post-college success once the college attended is held constant. Academic credentials (SAT/ACT scores, academic ratings) remain highly predictive of outcomes.&lt;/p&gt;
&lt;p&gt;Counterfactual simulations show that eliminating all three high-income admissions preferences and replacing those slots with students having the same test score distribution would increase enrollment from the bottom 95% of the parental income distribution by 8.8 percentage points — comparable in magnitude to the effect of race-based affirmative action on Black and Hispanic enrollment shares. Such a policy would have small effects on monetary leadership outcomes (e.g., Fortune 500 CEO share from bottom-95% families rises by only 0.4 pp, because Ivy-Plus graduates are a small fraction of all top earners) but larger effects on non-monetary leadership positions: the share of senators from the bottom 95% would rise by 1.7 pp and the share of Supreme Court justices by 5.4 pp. With need-affirmative policies (giving low-income students preferences comparable to those currently given to legacy applicants), the share of Supreme Court justices from families in the bottom 60% would rise by 17.5 pp. These predictions assume that the causal share of Ivy-Plus attendance in explaining observational differences in leadership outcomes is the same as that estimated for early-career outcomes, and they ignore general equilibrium effects.&lt;/p&gt;
&lt;p&gt;Q: How much more likely are top-1% students to attend an Ivy-Plus college than middle-class students with the same test scores?
A: Students from families in the top 1% (income above $611,000) are 2.3 times more likely to attend an Ivy-Plus college than students from the 70th–80th percentile of the parental income distribution (approximately $91,000–$114,000) with comparable SAT/ACT scores. This &amp;ldquo;missing middle&amp;rdquo; pattern is stable across entering classes from 1998 to 2018 and persists after controlling for race and ethnicity.&lt;/p&gt;
&lt;p&gt;Q: How is the overall attendance gap decomposed into application, admissions, and matriculation?
A: Differences in admissions rates explain two-thirds of the gap in Ivy-Plus attendance between top-1% and middle-class students conditional on test scores. Of the estimated 168 &amp;ldquo;extra&amp;rdquo; top-1% students per average Ivy-Plus class, 87 come from higher admissions rates for non-recruited athletes, 27 from athletic recruitment, and the remaining slack from application rate differences (accounting for roughly 20% of the overall attendance gap) and matriculation differences (roughly 12%).&lt;/p&gt;
&lt;p&gt;Q: How large is the admissions advantage for top-1% applicants at Ivy-Plus colleges?
A: Conditional on SAT/ACT scores, applicants from the top 1% are 58% more likely to be admitted to Ivy-Plus colleges than middle-class applicants. Students from the top 0.1% are 2.5 times more likely to be admitted than middle-class applicants with comparable test scores. At flagship public colleges, admissions rates are essentially constant across the income distribution conditional on test scores.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of legacy preferences and how is it established that legacy is not just a proxy for other credentials?
A: Legacy applicants from the top 1% are admitted at more than five times the rate of otherwise comparable non-legacy applicants at the college their parents attended. The paper isolates the legacy effect by showing that children of alumni at a given Ivy-Plus college are only slightly more likely to be admitted at other Ivy-Plus colleges — and the predicted counterfactual admissions rate for legacy students at other colleges closely matches their actual admissions rate — confirming that legacy status is not merely a proxy for other unobservable credentials. Legacy applicants constitute 2.5% of the overall applicant pool but over 9% of top-1% applicants.&lt;/p&gt;
&lt;p&gt;Q: How do non-academic credentials differ by parental income, and what drives the difference?
A: Top-1% applicants have markedly stronger non-academic ratings (measuring extracurricular participation and leadership traits) compared with other applicants, while the share achieving high academic ratings is essentially constant across the income distribution. Students from the top 1% are much more likely to have attended private high schools, whose applicants receive substantially higher non-academic ratings than students from public high schools with the same SAT/ACT scores. Non-academic ratings account for 35 of the estimated 168 extra top-1% students per Ivy-Plus class.&lt;/p&gt;
&lt;p&gt;Q: What is the research design for estimating causal effects, and what is the key identification assumption?
A: The authors focus on applicants who are waitlisted at a given Ivy-Plus college and compare those ultimately admitted versus rejected from the waitlist. The key identification assumption is that if different colleges&amp;rsquo; admissions committees make correlated assessments of underlying student merit but uncorrelated idiosyncratic admissions errors, then residual variation in admissions outcomes for waitlisted applicants at one college is orthogonal to students&amp;rsquo; long-run potential. The authors validate this empirically by showing that waitlist admission at one Ivy-Plus college is uncorrelated with admissions decisions and internal ratings at other Ivy-Plus colleges.&lt;/p&gt;
&lt;p&gt;Q: What are the causal effects of attending an Ivy-Plus college on post-college outcomes?
A: For the marginal student (one who attends an Ivy-Plus college instead of the average flagship public), attending an Ivy-Plus college increases the probability of reaching the top 1% of the earnings distribution at age 33 by approximately 50%, nearly doubles the probability of attending an elite graduate school, and increases the probability of working at a prestigious firm by approximately 2.5 times. Mean earnings at age 33 increase by $101,000 (relative to a counterfactual mean of $143,000 at state flagships). Effects on reaching the top quartile of earnings are small and statistically insignificant, while effects at the very top tail are disproportionately large.&lt;/p&gt;
&lt;p&gt;Q: Why do the findings differ from Dale and Krueger (2002) and related studies finding little effect of selective college attendance on earnings?
A: The authors replicate the matriculation design of Dale and Krueger (comparing outcomes conditional on the set of colleges to which students were admitted) and obtain estimates statistically indistinguishable from their waitlist design — the research designs are not the source of disagreement. Instead, the differences arise because (1) the authors have direct college fixed effects rather than relying on average test scores as a proxy for college quality, and (2) the authors focus on upper-tail outcomes (top 1% earnings, elite graduate schools, prestigious firms) rather than log mean earnings, where Ivy-Plus colleges have their largest effects.&lt;/p&gt;
&lt;p&gt;Q: Are the credentials that drive the high-income admissions advantage — legacy, athlete status, high non-academic ratings — predictive of better post-college outcomes?
A: No. Recruited athletes, students with higher non-academic ratings, and legacy students have equivalent or lower chances of reaching the upper tail of the income distribution, attending an elite graduate school, or working at a prestigious firm than comparable Ivy-Plus applicants once the college attended is held constant. By contrast, SAT/ACT scores and academic ratings are highly positively predictive of all three post-college outcome measures.&lt;/p&gt;
&lt;p&gt;Q: How much could changing admissions practices diversify Ivy-Plus enrollment and subsequently society&amp;rsquo;s leadership?
A: Eliminating legacy preferences, non-academic rating weights, and the differential recruitment of high-income athletes — and filling those slots with students having the same test score distribution as the current class — would increase enrollment from families in the bottom 95% of the parental income distribution by 8.8 percentage points, a magnitude comparable to race-based affirmative action&amp;rsquo;s effect on Black and Hispanic enrollment shares. For leadership positions, predicted effects are small for monetary outcomes (Fortune 500 CEOs from the bottom 95% would increase by only 0.4 pp) but larger for positions where Ivy-Plus graduates are a larger share: senators from the bottom 95% would increase by 1.7 pp and Supreme Court justices by 5.4 pp. A stronger need-affirmative policy (giving low-income students preferences equivalent to current legacy preferences) would increase the share of Supreme Court justices from the bottom 60% by 17.5 pp.&lt;/p&gt;
&lt;p&gt;Q: How are &amp;ldquo;elite&amp;rdquo; and &amp;ldquo;prestigious&amp;rdquo; employers defined in this study?
A: Elite firms are defined as those that disproportionately employ Ivy-Plus graduates relative to flagship public graduates, pulling firms from the top of that ratio ranking until 25% of Ivy-Plus attendee employment is accounted for. Prestigious employers are defined by the residual of that ratio after controlling for the firm&amp;rsquo;s predicted top-1% income probability — they are firms that disproportionately employ Ivy-Plus graduates conditional on their salaries, capturing high-status jobs that do not necessarily lead to the highest earnings. The paper validates this algorithmic approach against external rankings (Vault.com for law and consulting firms; Scimagoir for hospitals), finding substantial overlap.&lt;/p&gt;
&lt;p&gt;Q: How are treatment effect estimates adjusted for heterogeneity in students&amp;rsquo; fallback options?
A: Causal effects of Ivy-Plus attendance are much larger for students with weaker fallback options — specifically, students whose home-state flagship colleges channel fewer students to the top 1% of earnings. The authors exploit this heterogeneity to estimate the treatment effect for the marginal student who actually switches from a flagship public to an Ivy-Plus college. This heterogeneity also implies that the average causal effect across all admitted students may differ from the effect for the marginal admitted student.&lt;/p&gt;
&lt;p&gt;Q: What share of the overrepresentation of top-1% families at Ivy-Plus colleges is attributable to pre-application factors versus admissions practices?
A: Of the 245 &amp;ldquo;extra&amp;rdquo; top-1% students in an average Ivy-Plus class relative to an unconditionally income-neutral benchmark, 77 (31%) are attributable to the higher test scores of top-1% students (a pre-application factor). The remaining 168 (69%) reflect higher attendance rates conditional on test scores, of which the large majority is attributable to admissions practices (legacy, non-academic ratings, athletic recruitment) rather than application or matriculation rate differences.&lt;/p&gt;
&lt;p&gt;Ivy-Plus colleges: The twelve highly selective private colleges comprising the eight Ivy League institutions plus Stanford, MIT, Duke, and the University of Chicago — the focus group of the study, which together account for more than 10% of Fortune 500 CEOs, a quarter of U.S. senators, and three-fourths of Supreme Court justices appointed in the last half century despite enrolling less than 0.5% of Americans.&lt;/p&gt;
&lt;p&gt;Missing middle: The pattern by which attendance rates at Ivy-Plus colleges conditional on SAT/ACT scores are lowest for students from the middle class (70th–80th percentile of the parental income distribution, approximately $91,000–$114,000) — lower than both the top 1% and, slightly, the bottom 40% — producing a non-monotone income gradient in attendance.&lt;/p&gt;
&lt;p&gt;Legacy preference: An admissions advantage given to applicants whose parent(s) obtained an undergraduate degree from the college to which the student is applying. In the paper&amp;rsquo;s data, legacy applicants from the top 1% are admitted at more than five times the rate of non-legacy applicants with comparable test scores, demographics, and admissions ratings; the preference is college-specific (children of alumni are only slightly more likely to be admitted at other Ivy-Plus colleges).&lt;/p&gt;
&lt;p&gt;Waitlist research design: The paper&amp;rsquo;s primary identification strategy for causal effects, which exploits idiosyncratic variation in admissions decisions among waitlisted applicants. The design&amp;rsquo;s validity rests on the empirical finding that waitlist admissions at one Ivy-Plus college are uncorrelated with admissions decisions and internal ratings at other Ivy-Plus colleges, implying that residual variation conditional on being on the waitlist is orthogonal to students&amp;rsquo; long-run potential outcomes.&lt;/p&gt;
&lt;p&gt;Prestigious employers: Firms defined by the paper&amp;rsquo;s algorithm as disproportionately employing Ivy-Plus graduates conditional on those firms&amp;rsquo; predicted top-1% income probability — capturing high-status employment that does not necessarily lead to the highest earnings (e.g., prominent law firms, consulting firms, elite hospitals). Validated against external rankings (Vault.com, Scimagoir).&lt;/p&gt;
&lt;p&gt;Non-academic ratings: Numerical scores assigned by admissions officers measuring aspects of an application outside academic achievement, such as extracurricular activities and leadership traits. In the paper&amp;rsquo;s data, non-academic ratings differ substantially by parental income — particularly because top-1% applicants disproportionately attend private high schools whose students receive higher non-academic ratings — while academic ratings do not differ across the income distribution.&lt;/p&gt;
&lt;p&gt;Surrogate index: A prediction of later earnings outcomes (specifically, probability of reaching the top 1% at age 33 and mean income rank) constructed from individuals&amp;rsquo; graduate school attendance and employer fixed effects at ages 22–25, used to extend the outcome window for cohorts observed only early in their careers. The approach follows the terminology and methodology of Athey et al. (2019).&lt;/p&gt;</description></item><item><title>Do Financial Concerns Make Workers Less Productive?</title><link>https://macropaperwarehouse.com/papers/do-financial-concerns-make-workers-less-productive/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/do-financial-concerns-make-workers-less-productive/</guid><description>&lt;h2 id="do-financial-concerns-make-workers-less-productive"&gt;Do Financial Concerns Make Workers Less Productive?&lt;/h2&gt;
&lt;h3 id="research-question"&gt;Research Question&lt;/h3&gt;
&lt;p&gt;The paper tests whether financial concerns distract workers sufficiently to meaningfully reduce their productivity, and whether receiving cash — by alleviating those concerns — can raise output even when total compensation is held fixed.&lt;/p&gt;
&lt;h3 id="setting-and-sample"&gt;Setting and Sample&lt;/h3&gt;
&lt;p&gt;The experiment involves 408 low-income male agricultural casual laborers in rural Odisha, India, recruited from 47 villages across five worksites in four districts. The study takes place during the lean agricultural season (March–June 2017 and 2018), when formal employment is scarce (workers found paid wage work on only 1.9 days per week on average). During this period, 86% of workers reported being &amp;ldquo;worried&amp;rdquo; or &amp;ldquo;very worried&amp;rdquo; about their finances, 68–71% carried outstanding loans, and 64–66% said they would have difficulty coming up with Rs. 1,000 (roughly four days of wages) in an emergency. Workers bring these burdens to the job: on a given day, approximately one in two workers reported thinking about financial worries while working.&lt;/p&gt;
&lt;h3 id="experimental-design"&gt;Experimental Design&lt;/h3&gt;
&lt;p&gt;Workers were employed for twelve days in a piece-rate manufacturing task — stitching sal tree leaves into disposable plates for restaurants. The payment-timing manipulation is the core of the identification strategy. Control workers received all accrued earnings as a lump sum on the final day (day 12). Treatment workers received their earnings in two installments: an interim payment of earnings to date on day 8 or 9 (randomly staggered across waves), with the balance paid on day 12. Total compensation was held constant across groups; only the timing of receipt differed. On day 5 (the &amp;ldquo;announcement day&amp;rdquo;), each worker learned his payment schedule individually. The design thus separates the announcement period (days 5 through the interim payment day, when workers know their schedule but have not yet received cash) from the post-pay period (days after the interim payment until the contract end). This enables the authors to test whether productivity effects arise from information about impending cash, or only once cash is physically in hand.&lt;/p&gt;
&lt;h3 id="first-stage-effects-on-financial-strain"&gt;First Stage: Effects on Financial Strain&lt;/h3&gt;
&lt;p&gt;Within three days of receiving the interim payment, treated workers increased loan repayments by Rs. 271, a 287% increase relative to the control group mean (p &amp;lt; 0.001), and were 40 percentage points (222%) more likely to repay any loan (p &amp;lt; 0.001). The majority of repayments occurred on the same evening as the cash disbursement — a 746% single-day increase in loan payments. Household expenditures on food, clothing, and essentials rose by 40% (Rs. 150) over three days (p &amp;lt; 0.001). Treatment workers also reported feeling more focused on the work task (11.5 percentage points more likely, p = 0.032) and were less likely to report thinking about financial worries while making plates (13.7 percentage points, p = 0.044).&lt;/p&gt;
&lt;h3 id="main-productivity-results"&gt;Main Productivity Results&lt;/h3&gt;
&lt;p&gt;In the post-pay period, treated workers increased output by 0.109 SD (6.9%) relative to the control group (p = 0.020). No treatment effect emerged during the announcement period (0.014 SD, p = 0.685); the post-pay and announcement-period effects are statistically distinguishable (p = 0.008). Because work hours are fixed and daily attendance is 98.3% with no treatment effect on attendance, these gains reflect improvements in how quickly workers produce plates per hour of work.&lt;/p&gt;
&lt;p&gt;Effects are concentrated among workers with below-median baseline wealth (fewer assets, less liquidity): for this subgroup, the interim payment increases output by 0.204 SD (13.0%, p = 0.003). For workers with above-median wealth, the effect is close to zero and statistically insignificant (p = 0.819).&lt;/p&gt;
&lt;h3 id="attentiveness-results"&gt;Attentiveness Results&lt;/h3&gt;
&lt;p&gt;Beyond total output, the authors measure attentiveness through three markers embedded in the finished plates: the number of &amp;ldquo;double holes&amp;rdquo; (paired stitching holes indicating a removed mistaken stitch), the number of leaves used, and the number of stitches used. These measures are collected unbeknownst to workers and combined into an &amp;ldquo;attentiveness index.&amp;rdquo; After receiving the interim payment, treated workers&amp;rsquo; attentiveness index increased by 0.077 SD across all workers (p = 0.092); among poorer workers, attentiveness increased by 0.17 SD (p = 0.041). This improvement occurred simultaneously with higher output speed — workers were producing plates faster while also making fewer mistakes, suggesting improved cognitive engagement rather than mere effort intensification.&lt;/p&gt;
&lt;h3 id="piece-rate-comparison"&gt;Piece-Rate Comparison&lt;/h3&gt;
&lt;p&gt;In separate supplementary rounds with 150 experienced workers, the authors varied piece rates (Rs. 2, 3, or 4) while holding overall earnings constant. Each one-rupee increase in the piece rate raised output by 0.020 SD (p = 0.042). Critically, piece-rate increases produced no detectable change in the attentiveness index (point estimate negative, statistically insignificant), and the piece-rate effect on output differs significantly from the attentiveness effect (p = 0.001). This indicates that consciou effort and automatic attentiveness can move independently: higher incentives increase pace but do not reduce attentional lapses, whereas financial relief increases both pace and attentiveness.&lt;/p&gt;
&lt;h3 id="alternative-explanations-ruled-out"&gt;Alternative Explanations Ruled Out&lt;/h3&gt;
&lt;p&gt;The authors systematically address gift exchange/fairness, trust, nutrition, and sleep. Fairness and gift-exchange stories are inconsistent with: (i) no detectable announcement-period effect; (ii) no decline in control-worker effort when treatment workers are paid before them; (iii) the pattern of effects being concentrated among poorer workers; and (iv) attentiveness being affected when it is not a sanctioned quality dimension for payment. Nutritional channels are inconsistent with overnight effect onset (nutritional stock changes are too slow biologically), no treatment effect on breakfast consumption patterns, and productivity effects persisting through the end of each workday. Sleep channels are inconsistent with no treatment effect on hours or quality of sleep.&lt;/p&gt;
&lt;h3 id="scope-conditions-and-implications"&gt;Scope Conditions and Implications&lt;/h3&gt;
&lt;p&gt;The effect operates through the actual arrival of cash, not its anticipation, consistent with a model in which automatic cognitive inputs — unlike consciously chosen effort — respond to current financial strain rather than expected future income. Effects are concentrated among more financially constrained workers within an already-poor sample. The authors do not identify the specific psychological mechanism (worry, anxiety, affect, or rumination) but interpret results as evidence that financial strain, at least partly through psychological channels, reduces earnings exactly when money is most needed.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-does-the-experiment-focus-on-payment-timing-rather-than-an-outright-transfer-of-additional-money"&gt;Q1. Why does the experiment focus on payment timing rather than an outright transfer of additional money?&lt;/h3&gt;
&lt;p&gt;Varying only payment timing — not total pay — holds constant both the piece-rate incentive and total wealth across treatment and control. An outright cash transfer would raise total lifetime income, potentially reducing effort through a neoclassical income effect (more lifetime wealth lowers the marginal utility of current consumption). By holding total compensation fixed and only shifting when it arrives, the design isolates the effect of financial strain per se, separable from any wealth or incentive effect.&lt;/p&gt;
&lt;h3 id="q2-why-is-there-no-treatment-effect-during-the-announcement-period-and-why-does-this-matter"&gt;Q2. Why is there no treatment effect during the announcement period, and why does this matter?&lt;/h3&gt;
&lt;p&gt;Between day 5 (when workers learn their payment schedule) and the interim payment date, treated workers know cash is coming but have not yet received it. Output in this window shows no treatment effect (0.014 SD, p = 0.685), and the announcement effect is significantly smaller than the post-pay effect (p = 0.008). This matters because it rules out mechanisms that should operate on information alone — including gift exchange, trust updating, or effort responses to higher discounted expected income — and is consistent with a model in which financial strain falls only when cash is physically received (e.g., moneylenders do not relent until the loan is actually repaid).&lt;/p&gt;
&lt;h3 id="q3-what-is-the-attentiveness-index-and-how-was-it-constructed"&gt;Q3. What is the attentiveness index and how was it constructed?&lt;/h3&gt;
&lt;p&gt;The attentiveness index averages three plate-level markers: (i) number of &amp;ldquo;double holes&amp;rdquo; — pairs of stitching holes indicating a mistaken stitch was removed; (ii) number of leaves used; and (iii) number of stitches used. Each component was normalized using the control group&amp;rsquo;s post-pay mean and standard deviation, then averaged and reverse-coded so that higher values denote better attentiveness (fewer mistakes, fewer leaves, fewer stitches). Workers were unaware these dimensions were being measured. The index thus captures the number of unforced steps a worker took to complete a plate — a behavioral trace of cognitive lapses.&lt;/p&gt;
&lt;h3 id="q4-how-do-the-piece-rate-rounds-demonstrate-that-effort-and-attentiveness-are-separable"&gt;Q4. How do the piece-rate rounds demonstrate that effort and attentiveness are separable?&lt;/h3&gt;
&lt;p&gt;In supplementary rounds (150 workers, 2019), piece rates were experimentally varied among Rs. 2, 3, and 4 per plate with the base wage adjusted to hold total earnings constant, so financial strain was unchanged. A one-rupee increase in the piece rate raises output by 0.020 SD (p = 0.042), consistent with a standard effort response. The same increase produces no discernible change in the attentiveness index (point estimate: negative but not significant), and the output and attentiveness effects are significantly different from each other (p = 0.001). This shows that workers can speed up via conscious effort without reducing attentional lapses, whereas the cash infusion raises both pace and attentiveness simultaneously — a pattern inconsistent with pure motivation as the mechanism.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-staggered-timing-within-the-treatment-group-wave-a-vs-wave-b-contribute-to-identification"&gt;Q5. What does the staggered timing within the treatment group (Wave A vs. Wave B) contribute to identification?&lt;/h3&gt;
&lt;p&gt;Treatment workers were randomized to receive their interim payment on day 8 (Wave A) or day 9 (Wave B). On day 9, Wave B workers have not yet been paid while Wave A workers have. If fairness concerns drove control workers to reduce effort upon seeing colleagues paid first, control workers on day 9 — having observed Wave A payments the evening before — should work less hard relative to Wave B treatment workers (who have also not yet been paid). The authors find no such pattern: the triple interaction (Cash × Payment Day × Wave B) is close to zero and insignificant, ruling out effort reductions from seeing peers paid earlier.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-magnitudes-and-timing-of-the-spending-response-to-the-cash-infusion"&gt;Q6. What are the magnitudes and timing of the spending response to the cash infusion?&lt;/h3&gt;
&lt;p&gt;Within three days of the interim payment, treatment workers spent Rs. 900 in total — roughly two-thirds of the average interim payment of over Rs. 1,400. On the day of the payment itself, loan repayments rose by Rs. 169 (746% increase), and household expenditures rose by Rs. 70 (68% increase). Over three days, loan repayments increased by Rs. 271 (287%), the probability of repaying any loan rose by 40 percentage points (222%), and total household spending rose by 65% (Rs. 371). These patterns indicate that the two main sources of financial stress cited by workers — outstanding debt and inability to meet household essentials — were directly addressed, suggesting a meaningful reduction in financial strain.&lt;/p&gt;
&lt;h3 id="q7-why-are-the-productivity-effects-concentrated-among-poorer-workers-and-what-are-the-two-interpretations"&gt;Q7. Why are the productivity effects concentrated among poorer workers, and what are the two interpretations?&lt;/h3&gt;
&lt;p&gt;Workers with below-median baseline wealth (fewer assets, lower liquidity) show a 0.204 SD (13.0%) productivity gain, while workers above the median wealth threshold show essentially no effect. The authors offer two interpretations. First, poorer workers may start from a higher level of financial strain, giving the intervention more scope to reduce it. Second, since all workers in the sample are objectively poor and report similar baseline financial worries and loan levels, the more likely explanation is that the interim payment is larger relative to the wealth and income buffer of poorer workers, making the same nominal cash infusion more meaningful for them. Both richer and poorer workers in the sample use the interim payment to repay loans and cover household needs.&lt;/p&gt;
&lt;h3 id="q8-how-do-the-authors-rule-out-nutritional-channels"&gt;Q8. How do the authors rule out nutritional channels?&lt;/h3&gt;
&lt;p&gt;Two tests address nutrition. First, workers were not at subsistence — 94% reported missing no meals the prior week — and increased food spending cannot change the nutritional stock overnight (the medical literature indicates nutritional-stock effects on cognition operate over longer time horizons). Second, and more precisely, all food consumed at the worksite during the workday was provided by the researchers, so differential pre-worksite breakfast consumption is the only plausible same-day biological channel. The authors find no treatment effect on breakfast consumption (whether workers had breakfast, how much, or what they ate). Further, if blood sugar or satiety drove effects, they should attenuate over the workday as all workers are given the same afternoon meal; instead, treatment effects persist and if anything increase through the final hours of the workday.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-self-report-evidence-on-focus-and-worry-show-and-why-is-it-treated-as-suggestive-rather-than-primary"&gt;Q9. What does the self-report evidence on focus and worry show, and why is it treated as suggestive rather than primary?&lt;/h3&gt;
&lt;p&gt;Two days after the interim payment, workers were asked an open-ended question about what they were thinking about while working. Treatment workers were 11.5 percentage points (15.5%) more likely to report feeling focused on the task (p = 0.032) and 13.7 percentage points (32.7%) less likely to report thinking about financial worries (p = 0.044). A supplementary test showed treated workers were 10 percentage points (31%) more likely to generate explanations for a low-income person&amp;rsquo;s negative affect that were unrelated to financial concerns (p &amp;lt; 0.05), suggesting a broadening of cognitive scope. These measures are treated as suggestive because they were collected only at a single point and are self-reported; the primary evidence rests on objective production data because it is more objective and collected at fine hourly resolution throughout the post-pay period.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-paper-say-about-optimal-payment-frequency-as-a-policy-implication"&gt;Q10. What does the paper say about optimal payment frequency as a policy implication?&lt;/h3&gt;
&lt;p&gt;The authors are cautious in drawing a direct policy inference about paying workers more frequently. While the positive productivity effect of early payment points toward more frequent paydays reducing financial strain, this must be weighed against workers&amp;rsquo; self-control problems in consumption. In settings where workers face lumpy expenditure needs (e.g., monthly rent), more frequent payments could cause under-saving and worsen strain at the time of lumpy bills. The authors suggest payment frequency or size that matches expenditure needs, or more generally financial products that allow workers to time income receipts to coincide with expenses, as potentially more robust solutions — noting that such products appear largely absent in these markets.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Financial strain (as used in the paper):&lt;/strong&gt; A psychological burden arising from pressing present needs for resources — defined in the authors&amp;rsquo; model as increasing in both the current marginal utility of consumption (i.e., how valuable an additional rupee would be today) and the level of outstanding debt (including lender harassment pressure). Strain is present-oriented: it responds to current cash-on-hand and debt levels, not to expected future income, which is why anticipating a payment does not fully relieve it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Automatic input (a):&lt;/strong&gt; In the authors&amp;rsquo; behavioral model, one of two inputs into production. Unlike &amp;ldquo;effortful&amp;rdquo; input (e), which the worker consciously controls (speed of hands, consciously directed attention), the automatic input captures cognitive functions that are beyond the worker&amp;rsquo;s full control — background attentional processes that can be degraded by financial strain even when a worker is motivated and exerting high effort. The key behavioral assumption is that a falls when financial strain is high, independently of chosen effort.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Attentiveness index:&lt;/strong&gt; A composite measure constructed from three unincentivized physical markers embedded in completed leaf plates: (i) number of double holes (pairs indicating a stitch was removed to correct a mistake); (ii) number of leaves used; (iii) number of stitches used. The index is normalized to the control group&amp;rsquo;s post-pay distribution and reverse-coded so higher values denote better attentiveness. Workers were unaware these dimensions were measured. The index captures attentional lapses — unforced errors that increase the number of steps and time needed to complete each plate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Announcement period:&lt;/strong&gt; The days between when workers are individually informed of their payment schedule (day 5) and when the interim payment is actually disbursed (day 8 or 9). This window serves as a within-experiment control: if effects arose from information about impending cash (e.g., through discounting, gift exchange, or trust), they should appear here. The consistent absence of treatment effects during this period is a key identification result.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Post-pay period:&lt;/strong&gt; The days from the interim payment until the contract end (day 12). The main productivity and attentiveness treatment effects are estimated in this window, comparing treatment workers (who have received cash) to control workers (who have not yet been paid).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Lean season:&lt;/strong&gt; The months outside the peak agricultural planting and harvesting periods (roughly six to eight months per year in the study area) during which agricultural workers seek intermittent casual employment in manufacturing, construction, and other sectors. Employment rates are low (1.9 paid days per week on average), income is low and variable, and financial strain is correspondingly high. The experiment is intentionally conducted during this period to maximize baseline levels of financial concern.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Piece-rate elasticity of effort:&lt;/strong&gt; The responsiveness of output to changes in the marginal return per unit produced (the piece rate), holding financial strain constant. In the supplementary rounds, a one-rupee increase in the piece rate raises output by 0.020 SD. The authors interpret this as the upper bound on how much pure motivational effort can move output in this task, and use it to benchmark the cash infusion effects, which are roughly five times larger per unit of treatment variation and additionally move attentiveness (which piece-rate changes do not).&lt;/p&gt;</description></item><item><title>Efficiency Criteria, Income Taxation, and Heterogeneous Elasticities</title><link>https://macropaperwarehouse.com/papers/efficiency-criteria-income-taxation-and-heterogeneous-elasticities/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/efficiency-criteria-income-taxation-and-heterogeneous-elasticities/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Can income tax schedules be justified as utilitarian-optimal without adopting extreme normative assumptions about how household welfare should be measured? The paper proposes a welfare criterion strictly stronger than Pareto efficiency—called &lt;em&gt;rationalizability with bounded curvature&lt;/em&gt;—and asks whether observed US income taxes satisfy it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Starting Point.&lt;/strong&gt; Any Pareto-efficient nonlinear income tax schedule can, in principle, be rationalized as utilitarian-optimal under &lt;em&gt;some&lt;/em&gt; cardinalization of household utilities (i.e., some choice of how to measure the cardinal scale of each household&amp;rsquo;s well-being). However, the paper shows that rationalizing Pareto-efficient taxes in this way often requires cardinalizations under which there is &lt;em&gt;no&lt;/em&gt; population upper bound on the curvature of utility with respect to consumption. Equivalently, a utilitarian planner&amp;rsquo;s marginal willingness to transfer resources to households must fall arbitrarily quickly with the size of those transfers—an extreme form of status quo bias violated by virtually all quantitative optimal-tax exercises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Proposed Criterion.&lt;/strong&gt; The authors restrict attention to cardinalizations with &lt;em&gt;locally bounded curvature&lt;/em&gt;: there exists a finite (though potentially arbitrarily large) upper bound on the coefficient of relative risk aversion across the population. This admits two interpretations: (i) ex post, it requires that the social value of transfers not change arbitrarily quickly with transfer size; (ii) ex ante, it corresponds to a decision-maker behind a veil of ignorance with bounded risk aversion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Theoretical Result.&lt;/strong&gt; Within a standard Mirrlees model of nonlinear income taxation with arbitrary preference heterogeneity and intensive-margin labor supply, the paper proves that a tax schedule can be rationalized with bounded curvature if and only if government revenues are both &lt;em&gt;decreasing and concave&lt;/em&gt; (not merely decreasing) with respect to a class of narrowly targeted &amp;ldquo;two-bracket&amp;rdquo; reforms—reforms that raise retention by $1 local to some income level $z$ and zero elsewhere. This contrasts with Pareto efficiency, which requires only that revenues be decreasing in these reforms (Bierbrauer, Boyer, and Hansen 2023). The additional requirement of revenue concavity is what distinguishes the bounded-curvature criterion from pure Pareto efficiency.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sufficient Statistics.&lt;/strong&gt; The paper derives explicit sufficient-statistics expressions for the first- and second-order derivatives of tax revenue with respect to these targeted reforms. The second derivative depends on higher moments of the elasticity distribution, specifically the &lt;em&gt;income-conditional variance&lt;/em&gt; of compensated elasticities of taxable income (ETIs). Revenue convexity—which causes the second-order condition to fail—arises when income-conditional ETI variance is sufficiently high, even holding the mean ETI fixed. The economic mechanism is a &amp;ldquo;sort-and-extort&amp;rdquo; dynamic: a small tax reform sorts higher-elasticity households into income brackets where marginal taxes fall and lower-elasticity households into brackets where marginal taxes rise; repeating the reform then exploits this sorting by differentially taxing households by elasticity, as if applying group-specific tax schedules within a uniform income tax.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Findings.&lt;/strong&gt; Using the NBER panel of US tax returns from 1979 to 1990, the paper estimates income-conditional mean ETIs of approximately 0.2–0.3 at most income levels. Crucially, it estimates a &lt;em&gt;lower bound&lt;/em&gt; on income-conditional ETI variance by comparing elasticities of light versus heavy itemizers (defined by whether a household claims above or below the mean value of deductions in its income bracket). The low-elasticity group has an ETI of approximately zero and the high-elasticity group has an ETI of approximately one, implying a lower bound on ETI variance of roughly 0.2 at most incomes and approximately 0.25 at the top of the distribution. This lower bound is close to—and under plausible assumptions above—the threshold required for the second-order condition to fail. The authors conclude that the US income tax schedule in 1990 was likely Pareto efficient but likely &lt;em&gt;not&lt;/em&gt; rationalizable with bounded curvature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Welfare Gains.&lt;/strong&gt; In a calibrated model with a 50% top marginal tax rate, Pareto-tail shape of 2.5, mean ETI of 0.3, and ETI standard deviation of 0.75 (50% above the estimated lower bound), the planner gains significant welfare from either raising or lowering top marginal taxes. The welfare-maximizing top rate below the baseline is 13.3%, generating social value equivalent to a transfer of $1,966 per top earner. The welfare-maximizing top rate above the baseline is 71.2%, generating social value equivalent to a transfer of $972 per top earner. The revenue-maximizing rate is 80.9% under the baseline calibration, ranging from 74.6% to 86.8% as ETI standard deviation varies by ±25% of the lower bound.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; The theoretical analysis is restricted to intensive-margin labor supply (abstracting from extensive-margin decisions); the empirical application focuses on top incomes where extensive-margin effects are likely small. The empirical period is 1979–1990, covering major federal and state tax reforms. Results concern local efficiency of the tax schedule, not global optimization.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-exactly-is-rationalizability-with-bounded-curvature-and-how-does-it-differ-from-pareto-efficiency"&gt;Q1. What exactly is &amp;ldquo;rationalizability with bounded curvature&amp;rdquo; and how does it differ from Pareto efficiency?&lt;/h3&gt;
&lt;p&gt;A: Pareto efficiency requires that no small reform makes someone better off without making anyone worse off. Rationalizability (with &lt;em&gt;any&lt;/em&gt; cardinalization) is equivalent to Pareto efficiency in this setting. Rationalizability with bounded curvature additionally restricts the cardinalization: there must exist a finite upper bound on the coefficient of relative risk aversion (or equivalently, on the curvature of utility with respect to consumption) across the population. This is a strictly stronger criterion than Pareto efficiency. A schedule can be Pareto efficient but not rationalizable with bounded curvature if the only cardinalizations that rationalize it require unbounded consumption utility curvature.&lt;/p&gt;
&lt;h3 id="q2-why-do-extreme-cardinalizations-with-unbounded-curvature-arise-when-rationalizing-pareto-efficient-taxes"&gt;Q2. Why do &amp;ldquo;extreme&amp;rdquo; cardinalizations with unbounded curvature arise when rationalizing Pareto-efficient taxes?&lt;/h3&gt;
&lt;p&gt;A: When a Pareto-efficient schedule is rationalized as utilitarian, the cardinalization must make the set of feasible, recardinalized utilities convex so it can be separated from the set of Pareto-improving allocations. The paper constructs such a cardinalization explicitly: it takes the form of a function whose second derivative approaches negative infinity as utility approaches its baseline value. This implies the planner&amp;rsquo;s marginal value of transfers to a household falls precipitously as the household is made even slightly better off—an extreme status quo bias. Theorem 2.b establishes that &lt;em&gt;all&lt;/em&gt; cardinalizations rationalizing a schedule with convex revenues must share this pathology.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-sort-and-extort-mechanism-and-how-does-it-generate-revenue-convexity"&gt;Q3. What is the &amp;ldquo;sort-and-extort&amp;rdquo; mechanism and how does it generate revenue convexity?&lt;/h3&gt;
&lt;p&gt;A: When elasticities of taxable income (ETIs) are heterogeneous within an income level and the income density is declining steeply, a reform that lowers marginal taxes around income $z$ brings more households into the local bracket (because there are more households just below $z$ than above). Crucially, it disproportionately attracts households with &lt;em&gt;higher&lt;/em&gt; ETIs, since they respond more strongly to the marginal tax cut and relocate from further away, where the density differs more. Repeating the reform therefore faces a higher-elasticity composition at $z$, generating larger positive behavioral effects—making revenues convex in the size of the reform. The second step (&amp;ldquo;extort&amp;rdquo;) involves raising taxes on the now-concentrated low-elasticity households at adjacent brackets, achieving as-if group-specific taxation within a single income tax schedule.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-precise-relationship-between-revenue-convexity-and-eti-variance"&gt;Q4. What is the precise relationship between revenue convexity and ETI variance?&lt;/h3&gt;
&lt;p&gt;A: The paper shows (Theorem 4) that the second-order revenue derivative with respect to a narrow two-bracket reform around income $z$ equals a positive function of the income density times the expression $-[1-R&amp;rsquo;_0(z)]\varepsilon(z) + [1-R&amp;rsquo;_0(z)]\alpha(z)[\varepsilon^2(z) + \text{var}_h[\varepsilon^h | z^h_0=z]]$. The first term is always negative (pushing toward revenue concavity). The second term, which includes the income-conditional variance of ETIs, can dominate and create revenue convexity when ETI variance is sufficiently large. In the benchmark case with a single household type at each income (no within-income heterogeneity), the variance term vanishes and revenues are always concave whenever decreasing.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-sufficient-statistics-test-for-rationalizability-at-the-top-of-the-income-distribution"&gt;Q5. What is the sufficient statistics test for rationalizability at the top of the income distribution?&lt;/h3&gt;
&lt;p&gt;A: At top incomes (assuming no income effects, no super-elasticities, and CES preferences), taxes are Pareto efficient if and only if $\tau_\text{top} &amp;lt; \frac{1}{1+\alpha_\text{top}\varepsilon_\text{top}}$, and they are rationalizable with bounded curvature if and only if additionally $\tau_\text{top} &amp;lt; \frac{2}{1+\alpha_\text{top}(\varepsilon_\text{top} + \sigma^2_\text{top}/\varepsilon_\text{top})}$, where $\tau_\text{top}$ is the top marginal tax rate, $\alpha_\text{top}$ is the Pareto tail shape, $\varepsilon_\text{top}$ is the mean ETI at the top, and $\sigma^2_\text{top}$ is the income-conditional ETI variance at the top.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-estimate-a-lower-bound-on-income-conditional-eti-variance"&gt;Q6. How does the paper estimate a lower bound on income-conditional ETI variance?&lt;/h3&gt;
&lt;p&gt;A: The authors divide households at each income level into &amp;ldquo;heavy&amp;rdquo; and &amp;ldquo;light&amp;rdquo; itemizers based on whether their total deductions exceed the local income-bracket mean. They then estimate group-specific ETIs using local polynomial regressions of log income changes on log marginal retention changes, interacting tax changes with heavy-itemizer indicators. The within-year difference in elasticities between groups provides a lower bound on within-income ETI variance, since the two-group decomposition captures only a fraction of true variance. The interaction coefficient is allowed to vary by year to isolate within-year, within-income variation in elasticities rather than between-year compositional changes.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-estimated-magnitudes-of-mean-and-variance-of-etis"&gt;Q7. What are the estimated magnitudes of mean and variance of ETIs?&lt;/h3&gt;
&lt;p&gt;A: Income-conditional average ETIs are estimated at between 0.2 and 0.3 at most income levels, consistent with but somewhat below prior literature estimates. The low-elasticity group (light itemizers) has an ETI of approximately zero, while the high-elasticity group (heavy itemizers) has an ETI of approximately one. Given roughly equal group sizes, this implies a lower bound on ETI variance of approximately 0.2 at most incomes and approximately 0.25 at the ninety-fifth percentile. Subdividing the high-elasticity group into two, three, and four subgroups yields a lower bound of approximately 0.25 for variance at the top.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-back-of-the-envelope-calculation-work-to-assess-whether-the-second-order-test-fails"&gt;Q8. How does the back-of-the-envelope calculation work to assess whether the second-order test fails?&lt;/h3&gt;
&lt;p&gt;A: With $\tau_\text{top} \approx 0.5$, $\alpha_\text{top} \approx 2.5$, and $\varepsilon_\text{top} \approx 0.3$ (from prior literature), the second-order condition fails if and only if ETI variance exceeds approximately 0.27. The authors&amp;rsquo; lower bound estimate of ETI variance is already approximately 0.25 (standard deviation approximately 0.5), just below this threshold. The authors note that if the true standard deviation exceeds the lower bound by more than 4%, the second-order condition fails, making it empirically likely that the 1990 US tax schedule was not rationalizable with bounded curvature.&lt;/p&gt;
&lt;h3 id="q9-why-does-the-paper-focus-on-the-top-of-the-income-distribution-for-the-empirical-test"&gt;Q9. Why does the paper focus on the top of the income distribution for the empirical test?&lt;/h3&gt;
&lt;p&gt;A: The second-order condition is most likely to fail at high incomes for three reasons simultaneously: (i) the marginal tax rate is highest, (ii) ETI means are somewhat higher there, and (iii) the Pareto parameter $\alpha(z)$ is largest (income density falls steeply), which amplifies the sort-and-extort mechanism. The authors also note that extensive-margin labor supply responses—which are abstracted away in the theory—are likely small at high incomes.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-calibrated-quantitative-application-reveal-about-optimal-top-tax-policy"&gt;Q10. What does the calibrated quantitative application reveal about optimal top tax policy?&lt;/h3&gt;
&lt;p&gt;A: Calibrated with a 50% initial top marginal tax rate, Pareto tail shape of 2.5, mean ETI of 0.3, and ETI standard deviation of 0.75 (50% above the estimated lower bound), the model finds welfare gains in both directions of reform. The welfare-maximizing rate &lt;em&gt;below&lt;/em&gt; the baseline is 13.3%, yielding equivalent welfare gains of $1,966 per top earner. The welfare-maximizing rate &lt;em&gt;above&lt;/em&gt; the baseline is 71.2%, yielding equivalent gains of $972 per top earner. The revenue-maximizing rate is 80.9%, ranging from 74.6% to 86.8% when ETI standard deviation varies by ±25% of the lower bound. This sensitivity highlights that the optimal direction and magnitude of reform depend substantially on the uncertain degree of ETI heterogeneity.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-relate-to-the-inverse-optimum-literature"&gt;Q11. How does the paper relate to the &amp;ldquo;inverse optimum&amp;rdquo; literature?&lt;/h3&gt;
&lt;p&gt;A: The inverse optimum approach (Bourguignon and Spadaro 2012; Hendren 2020) infers the first-order welfare trade-offs implicit in an observed tax schedule. This paper goes further by inferring from second-order empirical moments—specifically the income-conditional ETI variance—whether taxes are consistent with &lt;em&gt;minimal&lt;/em&gt; requirements on how sensitive the planner&amp;rsquo;s trade-offs are to household welfare levels. Rather than assuming a welfare function, it tests whether &lt;em&gt;any&lt;/em&gt; welfare function with bounded curvature can rationalize the observed schedule.&lt;/p&gt;
&lt;h3 id="q12-is-revenue-convexity-possible-without-within-income-heterogeneity-in-preferences"&gt;Q12. Is revenue convexity possible without within-income heterogeneity in preferences?&lt;/h3&gt;
&lt;p&gt;A: Yes, but only under more specific conditions. The paper provides two supplemental examples. In the first, all households have constant-elasticity labor disutility but differ in both productivity and elasticity across income levels; when lower-income households have higher elasticities, a reform reducing marginal taxes at $z$ attracts higher-elasticity households and raises the average elasticity, leading to convex revenues. In the second, all households have the same initial elasticity but individual elasticities change in response to reforms. However, with the standard additively separable CES preferences and no within-income heterogeneity, revenues are always concave when decreasing—consistent with Werning&amp;rsquo;s (2007) observation that the Pareto planner&amp;rsquo;s problem is convex in this case.&lt;/p&gt;
&lt;h3 id="q13-what-is-the-role-of-random-tax-reforms-in-the-papers-logic"&gt;Q13. What is the role of random tax reforms in the paper&amp;rsquo;s logic?&lt;/h3&gt;
&lt;p&gt;A: Random tax reforms serve as an expository bridge. The paper shows that if the second-order revenue effect of a two-bracket reform is positive at some income $z$, then a &amp;ldquo;randomized&amp;rdquo; reform that applies the reform with equal probability in positive and negative directions generates an expected Pareto improvement—because the convexity of revenues implies expected revenues rise, while for any household with bounded risk aversion the reform&amp;rsquo;s second-order utility effect is also positive when the reform is sufficiently narrow. This establishes that revenue convexity implies random Pareto inefficiency under bounded risk aversion, and then the paper shows the analogous deterministic result for rationalizability.&lt;/p&gt;
&lt;h3 id="q14-what-scope-conditions-attach-to-the-sufficient-conditions-for-rationalizability-theorem-3"&gt;Q14. What scope conditions attach to the sufficient conditions for rationalizability (Theorem 3)?&lt;/h3&gt;
&lt;p&gt;A: Theorem 3 requires Assumptions 1 and 3 plus two boundary conditions: the ratio $\delta\text{Rev}(z)/(zg(z))$ must remain bounded away from zero as income approaches 0 or infinity, and at all incomes there must exist households with low enough compensated elasticities. Assumption 1 requires that average and marginal taxes have upper bounds below one, that marginal taxes have a lower bound, and that $zg(z)$ converges to zero at the boundaries. Assumption 3 is a regularity condition on how conditional moments of the elasticity distribution vary with income. These conditions ensure that the narrow, self-financing reforms considered in the necessity proof cannot generate welfare improvements once revenues are both decreasing and concave.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Rationalizability with Bounded Curvature.&lt;/strong&gt; The property that a tax schedule is utilitarian-optimal under some cardinalization of household utilities in which there exists a finite (though potentially arbitrarily large) upper bound on the curvature of utility with respect to consumption across the population. Formally, there exists a continuous function $\bar{\rho}$ such that, for all households, the absolute value of $[w_h \circ u_h]_{cc} / [w_h \circ u_h]_c$ is bounded by $\bar{\rho}$ evaluated at the household&amp;rsquo;s income. This criterion is strictly stronger than Pareto efficiency and strictly weaker than utilitarian optimality under a fixed cardinalization.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Two-Bracket Reform.&lt;/strong&gt; A targeted tax reform that increases retention (post-tax income) by $1 at incomes local to some level $z$ over a small bracket of width $\ell$, and zero elsewhere (smoothed at the edges). As $\ell \to 0$, this becomes an infinitesimally narrow reform. The first- and second-order revenue effects of these reforms—denoted $\delta\text{Rev}(z)$ and $\delta^2\text{Rev}(z)$—are the paper&amp;rsquo;s key objects: Pareto efficiency requires $\delta\text{Rev}(z) &amp;lt; 0$ for all $z$, and rationalizability with bounded curvature additionally requires $\delta^2\text{Rev}(z) \leq 0$ for all $z$.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Income-Conditional ETI Variance.&lt;/strong&gt; The variance of compensated elasticities of taxable income (ETIs) among households with the same income level, $\text{var}_h[\varepsilon^h | z^h_0 = z]$. This is the paper&amp;rsquo;s primary empirical object of interest and the key determinant of whether revenues are convex or concave in the size of targeted reforms. Unlike the literature&amp;rsquo;s focus on mean ETIs by income bracket, this within-income variance captures heterogeneity among households sharing the same pre-reform income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sort-and-Extort Mechanism.&lt;/strong&gt; The two-step economic mechanism underlying revenue convexity from ETI heterogeneity. In the first step (&amp;ldquo;sort&amp;rdquo;), a marginal tax cut around income $z$ disproportionately attracts higher-ETI households from lower incomes (because they respond more strongly and relocate from further away), shifting the elasticity composition at $z$ upward. In the second step (&amp;ldquo;extort&amp;rdquo;), repeating the reform finds higher-elasticity households concentrated where marginal taxes fall and lower-elasticity households where taxes rise, effectively applying differential tax treatment by elasticity within a single income tax schedule.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Pareto Parameter $\alpha(z)$.&lt;/strong&gt; Defined as $-d\log(zg(z))/d\log z$, where $g(z)$ is the income density. This captures the rate at which the income density is falling in income locally at $z$, and governs the strength of the sort-and-extort mechanism. High $\alpha(z)$ at top incomes (reflecting a steeply declining Pareto-type density) amplifies revenue convexity from ETI heterogeneity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Super-Elasticity.&lt;/strong&gt; A concept that captures how a household&amp;rsquo;s compensated ETI would change if its income were different, holding preferences fixed. Formally, it is the derivative of the household&amp;rsquo;s elasticity with respect to its log income, decomposing into effects from changes in preference curvature and changes in the local curvature of the tax schedule. Super-elasticities are zero in the benchmark case of additively CES preferences and locally CES retention schedules but contribute additional terms to the second-order revenue expression in the general case.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cardinalizing Function.&lt;/strong&gt; A strictly increasing function $w_h$ that maps household $h$&amp;rsquo;s indirect utility $V_h$ to a cardinalized utility level $w_h(V_h)$. The social planner maximizes the expectation of cardinalized utilities. Different choices of ${w_h}_h$ correspond to different stances on interpersonal comparisons, including unbounded curvature (rationalizing any Pareto-efficient schedule) or bounded curvature (the paper&amp;rsquo;s proposed restriction). Rawlsian social welfare is a limit of utilitarian welfare with increasingly concave cardinalizing functions.&lt;/p&gt;</description></item><item><title>Equal Pay for Similar Work</title><link>https://macropaperwarehouse.com/papers/equal-pay-for-similar-work/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/equal-pay-for-similar-work/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper studies the labor market effects of &amp;ldquo;Equal Pay for Similar Work&amp;rdquo; (EPSW) policies — laws that require firms to pay equal wages to workers of different protected-class identities (e.g., different genders) who perform &amp;ldquo;similar&amp;rdquo; work within a firm. EPSW has become increasingly prevalent: as of January 2023, more of the U.S. workforce falls under state EPSW laws than state &amp;ldquo;Equal Pay for Equal Work&amp;rdquo; (EPEW) laws. Despite this spread, the equilibrium consequences of EPSW were previously unknown.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Theoretical Framework&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors develop two theoretical models. The first is a static cooperative game (whose outcomes coincide with the Nash equilibria of a non-cooperative simultaneous-wage-offer game). Homogeneous firms with constant-returns-to-scale production compete for a continuum of heterogeneous workers. Workers belong to one of two groups A or B (e.g., men and women), with group A constituting a β ≥ 1 majority. Each worker&amp;rsquo;s productivity v is drawn from a group-specific distribution (FA or FB); firms&amp;rsquo; willingness to pay equals each worker&amp;rsquo;s productivity, but can embed taste-based discrimination. The analysis is framed as applying &amp;ldquo;within job&amp;rdquo; in a local labor market — only workers performing &amp;ldquo;similar&amp;rdquo; work in the eyes of the law.&lt;/p&gt;
&lt;p&gt;The second model is a dynamic search-and-bargaining framework with an arbitrary number of firms, search frictions, reallocation frictions, and Nash-in-Nash bargaining. EPSW is introduced as a surprise, and constrained firms choose whether to segregate for one group or remain desegregated (paying a common wage to all workers).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Theoretical Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Without EPSW, Bertrand competition among firms drives every worker&amp;rsquo;s wage to equal her productivity; any wage gap between groups A and B exactly reflects the difference in average productivities (EA(v) − EB(v)), whether or not those productivity differences stem from discrimination.&lt;/p&gt;
&lt;p&gt;With EPSW, the equilibrium is qualitatively transformed. In the static model (Proposition 2), firms generically fully segregate their workforces: one firm hires all A-group workers and the other hires all B-group workers. EPSW functions as an enforcement mechanism for this segregation analogous to location choices in Hotelling&amp;rsquo;s model — poaching a worker from the competing firm is costly because EPSW then requires the poaching firm to pay equal wages to all workers it employs. In the core with EPSW (Proposition 3), the wage gap moves in favor of the majority group (A-group, β &amp;gt; 1) in the sense that all core outcomes except one strictly increase the A-group wage advantage. Moreover, firm profits and the magnitude of the wage gap co-move: firms benefit from selecting equilibria with larger wage gaps. The directional conclusion — EPSW benefits the majority group — holds regardless of the distributions of the two groups&amp;rsquo; productivities, conditional only on β &amp;gt; 1 for the wage gap; for the log wage gap the additional regularity condition βEA[v] &amp;gt; EB[v] is required.&lt;/p&gt;
&lt;p&gt;In the dynamic search model (Proposition 4), all firms eventually segregate under any equilibrium, with the long-run wage ratio moving in favor of the group toward which more firms segregate. Under equitable search and sufficiently low reallocation frictions (Proposition 5), more firms segregate toward the majority group when βEA[v] &amp;gt; EB[v]. Firms that are nearly segregated at the time of EPSW enactment segregate sooner than others (Proposition 6).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Setting and Design&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors test these predictions using Chile&amp;rsquo;s 2009 EPSW (Law 20.348), the country&amp;rsquo;s first equal pay law, which prohibited paying women less than men (or vice versa) for similar work. Firms with 10 or more long-term workers at the time of announcement (June 2009) face formal grievance procedures and financial penalties (69–1,384 USD per worker-month of violation); firms below this threshold face no financial penalty, providing a clean threshold-based treatment assignment.&lt;/p&gt;
&lt;p&gt;The data are matched employer-employee administrative records from the Chilean unemployment insurance system covering January 2005 – December 2013, a random sample of approximately 4% of all firms stratified by size. The main estimation sample restricts to firms with 6–13 total workers at announcement (41% of active firms), and the design is a difference-in-differences (event study) comparing treated (≥ 10 long-term workers) to control (&amp;lt; 10 long-term workers) firms. The identifying assumption is parallel trends between similarly sized firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Empirical Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;First, EPSW increases full gender segregation across firms. The share of fully gender-segregated firms increases by 4.4 percentage points (baseline: 34.3% of firms were fully segregated at announcement). Simultaneously, the share of nearly-but-not-fully segregated firms (majority gender share ∈ [0.8, 1)) declines by 4.0 percentage points — a &amp;ldquo;missing mass&amp;rdquo; of near-segregated firms consistent with the search model&amp;rsquo;s prediction that firms on the margin of full segregation segregate most readily (e.g., by separating the sole worker of the &amp;ldquo;wrong&amp;rdquo; gender). Moreover, firms that are nearly segregated at announcement experience an 8.7 percentage point increase in full segregation post-EPSW, compared to 2.8 percentage points for firms not nearly segregated at announcement.&lt;/p&gt;
&lt;p&gt;Second, EPSW shifts the gender wage gap in favor of the local labor market majority group. In male-majority local labor markets (defined by industry × county), EPSW increases the gender wage gap in favor of men by 4.3 percentage points. In female-majority local labor markets, EPSW decreases the gender wage gap (i.e., in favor of women) by 6.2 percentage points. The wage gap change is primarily driven by reductions in minority-group wages: women&amp;rsquo;s average wages in male-majority markets fall by 3.3 percentage points, and men&amp;rsquo;s average wages in female-majority markets fall by 4.5 percentage points; there are no statistically significant changes in majority-group wages. Because men dominate Chile&amp;rsquo;s overall labor market (approximately 5/6 of all workers are employed in majority-male local labor markets), the overall effect of EPSW is to increase the gender wage gap (in favor of men) by 2.7 percentage points. Pre-treatment coefficients are statistically indistinguishable from zero across all specifications, supporting the parallel trends assumption. These findings are robust across six alternative specifications covering different samples, fixed-effect structures, and controls.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Theoretical results apply within a set of &amp;ldquo;similar&amp;rdquo; workers in a given local labor market — the paper does not predict differential effects across job types within a firm (e.g., custodians vs. lawyers) that do not perform similar work. Empirical results are identified for firms with 6–13 workers and pertain to Chile&amp;rsquo;s formal sector (informal labor share ~25% in 2009). Predictions on the wage ratio (log wage gap) require the additional regularity condition βEA[v] &amp;gt; EB[v], which is consistent with the Chilean data.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-mechanism-by-which-epsw-leads-firms-to-fully-segregate-in-the-static-model"&gt;Q1. What is the core mechanism by which EPSW leads firms to fully segregate in the static model?&lt;/h3&gt;
&lt;p&gt;A: EPSW makes cross-group poaching prohibitively costly. If a firm that hires only A-group workers were to hire even a positive measure of B-group workers, EPSW would — by transitivity — require it to pay the same wage to all workers. This eliminates the firm&amp;rsquo;s ability to exploit productivity heterogeneity across workers; it would have to raise all wages to match the highest worker, destroying profit. As a result, firms segregate in equilibrium to avoid the bite of EPSW entirely: each firm caters to one group, and the within-group wage schedule remains unconstrained. The mechanism is analogous to Hotelling&amp;rsquo;s location model: segregation serves as the enforcement device for avoiding the equal-pay constraint.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-equal-profit-condition-generate-a-wage-gap-in-favor-of-the-majority-group"&gt;Q2. How does the equal profit condition generate a wage gap in favor of the majority group?&lt;/h3&gt;
&lt;p&gt;A: In any core outcome under EPSW (Proposition 3), the Equal Profit Condition requires both firms to earn the same total profit. When there are β &amp;gt; 1 A-group workers (more than B-group workers), the firm serving A-group workers must pay higher average wages per worker to extract the same total profit from a larger pool, relative to the firm serving a smaller B-group. This mechanically raises A-group average wages relative to B-group average wages. Crucially, this directional conclusion — EPSW widens the majority-group wage advantage — holds regardless of the shapes of FA and FB, meaning it is robust to any underlying discriminatory or non-discriminatory productivity differences.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-baseline-without-epsw-wage-gap-and-how-does-epsw-change-it"&gt;Q3. What is the baseline (without-EPSW) wage gap, and how does EPSW change it?&lt;/h3&gt;
&lt;p&gt;A: Without EPSW, Proposition 1 establishes that every worker is paid exactly her productivity in any core outcome (full employment, wages = productivity). Therefore, the wage gap equals EA(v) − EB(v) and the wage ratio equals EA(v)/EB(v): any gap reflects only productivity differences (including discrimination embedded in willingness to pay). Under EPSW, Proposition 3 shows that all core outcomes except a single (measure-zero) one strictly widen the wage gap beyond this level. The wage ratio result (Proposition 3, Part 4) requires the additional condition βEA[v] &amp;gt; EB[v] — that the majority group is not sufficiently less productive or more discriminated against to reverse the direction.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-dynamic-search-model-modify-the-static-predictions"&gt;Q4. How does the dynamic search model modify the static predictions?&lt;/h3&gt;
&lt;p&gt;A: In the dynamic model (Proposition 4), full segregation is achieved in finite time T in any equilibrium, not instantaneously. Prior to T, firms make sequential segregation decisions; workers displaced by firm desegregation choices are replaced at rate ρ ∈ [0,1]. The long-run wage ratio is determined by the ratio nA/nB — the number of firms segregating toward group A versus B. If nA &amp;gt; nB, the long-run wage ratio moves in favor of A; if nA = nB, the policy has no long-run effect on the wage ratio. The key departure from the static model is that this outcome depends not only on the majority group size but also on search intensities and reallocation frictions (high firm tenure/low d can make segregating toward the majority costly if the firm already employs many minority-group workers).&lt;/p&gt;
&lt;h3 id="q5-under-what-conditions-does-the-dynamic-model-predict-that-more-firms-segregate-toward-the-majority-group"&gt;Q5. Under what conditions does the dynamic model predict that more firms segregate toward the majority group?&lt;/h3&gt;
&lt;p&gt;A: Proposition 5 states that for sufficiently large d (fast worker turnover / low reallocation frictions) and equitable search (equal search intensity across firms within a group), the number of firms segregating toward A satisfies nA ∈ [xA−1, xA+1], where xA is defined by an equal-profit condition. Moreover, if βEA[v] &amp;gt; EB[v] (the majority group is collectively more valuable), then nA ≥ nB. Without equitable search, the conclusion holds under more stringent conditions: for any search intensity vector r, there exist d* and β* such that for d &amp;gt; d* and β &amp;gt; β*, any equilibrium yields nA &amp;gt; nB. Empirically, 94% of local-labor-market-by-month units in Chile exhibit more firms segregating toward the majority gender post-EPSW, consistent with these conditions being met.&lt;/p&gt;
&lt;h3 id="q6-why-do-firms-that-are-nearly-segregated-at-announcement-respond-most-strongly-to-epsw"&gt;Q6. Why do firms that are nearly segregated at announcement respond most strongly to EPSW?&lt;/h3&gt;
&lt;p&gt;A: Proposition 6 establishes that firms with a low ratio of minority-group to majority-group search intensity (i.e., nearly segregated in employment) segregate earliest, provided the discount rate is sufficiently low. The intuition is that for a nearly segregated firm, the cost of segregating — separating the few minority-group workers — is small relative to the costs of remaining desegregated (paying a common wage that compresses profit, and being unable to poach new workers). Empirically, firms nearly segregated at announcement (majority gender share ∈ [0.8,1) at announcement) show an 8.7 percentage point increase in full segregation post-EPSW, roughly three times larger than the 2.8 percentage point effect for firms not nearly segregated at announcement. This &amp;ldquo;missing mass&amp;rdquo; pattern (decline in near-segregation matched by increase in full segregation) is also consistent with Proposition 6.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-heterogeneous-effect-of-epsw-on-the-wage-gap-by-local-labor-market-type"&gt;Q7. What is the heterogeneous effect of EPSW on the wage gap by local labor market type?&lt;/h3&gt;
&lt;p&gt;A: The empirical design allows the wage gap effect to differ by local labor market (LLM) majority type (male vs. female). In male-majority LLMs (firm industry × county pairs where males comprise more than 50% of workers in June 2009), EPSW increases the gender wage gap in favor of men by 4.3 percentage points (SE = 0.0116). In female-majority LLMs, EPSW decreases the gender wage gap (in favor of women) by 6.2 percentage points (SE = 0.0234). These findings precisely match the theoretical prediction that EPSW benefits whichever group is in the majority of the local labor market. The dynamic event studies show no pre-trends in either subsample; effects begin at announcement (τ = 0) and grow over time.&lt;/p&gt;
&lt;h3 id="q8-what-drives-the-wage-gap-change--majority-wages-rising-or-minority-wages-falling"&gt;Q8. What drives the wage gap change — majority wages rising or minority wages falling?&lt;/h3&gt;
&lt;p&gt;A: The change is primarily driven by a reduction in the minority group&amp;rsquo;s average wages, not an increase in majority wages. Women&amp;rsquo;s average wages in male-majority labor markets fall by 3.29 percentage points (SE = 0.0111) in treated versus control firms post-EPSW. Men&amp;rsquo;s average wages in female-majority labor markets fall by 4.45 percentage points (SE = 0.0178) in treated versus control firms post-EPSW. There are no statistically significant changes in the average wages of the majority group of workers within any LLM type. This is consistent with the model&amp;rsquo;s mechanism: segregation reduces competition for minority-group workers (fewer firms competing for them), depressing their wages.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-aggregate-economy-wide-effect-of-epsw-on-the-gender-wage-gap-in-chile"&gt;Q9. What is the aggregate (economy-wide) effect of EPSW on the gender wage gap in Chile?&lt;/h3&gt;
&lt;p&gt;A: Because approximately 5/6 of all Chilean workers are employed in male-majority local labor markets (men have higher labor force participation, with female labor force participation at roughly 30% in 2009), the overall effect of EPSW is to increase the gender wage gap in favor of men by 2.74 percentage points (SE = 0.0102). This is a net effect that averages the positive (pro-male) gap increase in male-majority markets and the negative (pro-female) gap decrease in female-majority markets, weighted by market sizes.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-identification-strategy-deal-with-anticipation-and-compositional-changes"&gt;Q10. How does the identification strategy deal with anticipation and compositional changes?&lt;/h3&gt;
&lt;p&gt;A: Treatment status is assigned based on firm size at the time of policy announcement (June 2009) rather than enactment (November 2009), creating an intent-to-treat framework: some &amp;ldquo;treated&amp;rdquo; firms may fall below the threshold by enactment, and some &amp;ldquo;control&amp;rdquo; firms may rise above it, both attenuating the estimates (implying estimated effects are plausible lower bounds). The no-anticipation assumption is supported by the absence of statistically significant pre-trends in either the segregation or wage-gap specifications. To address compositional changes in worker characteristics across LLMs induced by EPSW itself, the wage regressions include time fixed effects interacted with human capital dimensions (education, contract type, age decade) and firm comparison groups, controlling for observable composition shifts. Placebo tests at alternative firm-size thresholds find no statistically or economically meaningful effects, supporting the causal interpretation.&lt;/p&gt;
&lt;h3 id="q11-how-does-epsw-in-chile-compare-to-epew-theoretically-and-in-the-literature"&gt;Q11. How does EPSW in Chile compare to EPEW theoretically and in the literature?&lt;/h3&gt;
&lt;p&gt;A: EPEW requires equal pay only for workers doing exactly equal work, which creates an easily exploitable loophole: firms can proliferate job titles or marginally differentiate duties to avoid compliance. EPSW closes this by requiring equal pay across a coarser &amp;ldquo;similar work&amp;rdquo; category, making evasion harder. Theoretically, the prior EPEW literature (Bhaskar et al. 2002, Kaas 2009, Lagerlöf 2020, Lanning 2014) generated ambiguous directional predictions — equal pay laws could either increase or decrease wage disparities within the same paper. The authors attribute this ambiguity to EPEW models&amp;rsquo; requirement that workers be exactly equally productive. By contrast, EPSW applies across workers with heterogeneous productivities, and the authors derive unambiguous predictions: full segregation and a wage gap shift toward the majority group, both of which are confirmed empirically.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-analogy-to-best-price-guarantees-in-product-markets"&gt;Q12. What is the analogy to &amp;ldquo;best-price guarantees&amp;rdquo; in product markets?&lt;/h3&gt;
&lt;p&gt;A: The paper draws a methodological parallel to most-favored-customer (MFC) clauses in product markets. MFC clauses commit firms to rebating past consumers if prices fall, which directly equalizes payments across buyers but unintentionally raises firm market power. In the EPSW setting, the policy plays the role of a best-wage guarantee — but because firms compete for workers, the constraint binds off the equilibrium path. Firms segregate so that no firm is ever exposed to the equal-pay constraint in equilibrium, yet the threat of the constraint (if a firm deviates and hires from both groups) effectively differentiates labor costs across groups, driving the unintended wage effects. This is related to &amp;ldquo;artificial&amp;rdquo; switching costs that create local market power in consumer markets (Klemperer, 1987).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Equal Pay for Similar Work (EPSW):&lt;/strong&gt; A legal constraint requiring that within a firm, workers belonging to different protected-class identities (e.g., different genders) who perform &amp;ldquo;similar&amp;rdquo; work receive equal wages. Distinguished from &amp;ldquo;Equal Pay for Equal Work&amp;rdquo; (EPEW) by its coarser similarity standard, which cannot be evaded by minor job-title differentiation. In the model, this constraint is formalized as: a firm cannot hire positive measures of workers from two different groups such that all workers in one group receive strictly higher wages than all workers in the other group; by transitivity, a firm hiring from both groups must pay almost all workers the same wage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Core Outcome:&lt;/strong&gt; The solution concept used in the static model, drawing on cooperative game theory (Shapley–Shubik assignment game). An outcome (specifying which firm hires each worker and at what wage) is in the core if no firm and subset of workers can form a blocking coalition that makes both the firm and each worker in the coalition strictly better off. The paper uses this concept because its pure-strategy Nash equilibrium outcomes (in the associated non-cooperative simultaneous wage-offer game) exactly coincide with the core outcomes under the restriction that firms pay the same wage to all workers of the same type.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Full Segregation:&lt;/strong&gt; A labor market outcome in which each firm employs workers from only one group (all A-group workers at one firm, all B-group workers at the other). The paper proves (Proposition 2) that EPSW generically forces full segregation in equilibrium, because any deviation to hire from both groups exposes the firm to the equal-pay constraint. Empirically measured as a binary indicator for whether all workers at a given firm in a given month are of the same gender.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Near Segregation:&lt;/strong&gt; A firm-level state in which the majority gender constitutes 80–99% of the firm&amp;rsquo;s workforce (the majority gender share is in [0.8, 1)). The paper uses this as a complementary outcome to full segregation; theory (Proposition 6) predicts a decline in near segregation post-EPSW because firms in this state face the lowest cost of transitioning to full segregation. Empirically, the near-segregation share falls by 4.0 percentage points post-EPSW, mirroring the 4.4 percentage point rise in full segregation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Labor Market (LLM):&lt;/strong&gt; Defined in the empirical analysis as a firm&amp;rsquo;s geographic county interacted with its industry code, creating 321 × 21 potential cells. The LLM is classified as male-majority or female-majority based on the share of female workers across all firms in the industry-county pair in June 2009. This is the unit at which the &amp;ldquo;majority group&amp;rdquo; for Proposition 3&amp;rsquo;s wage gap prediction is defined, and the level at which the heterogeneous wage effects of EPSW are estimated.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Equal Profit Condition:&lt;/strong&gt; A necessary condition of any core outcome (with or without EPSW): both firms must earn the same total profit in equilibrium. Under EPSW with full segregation, this condition determines the relative average wages of the two groups — because firm sizes differ (β A-group workers vs. 1 B-group worker), equal profit requires the firm serving the larger group to pay higher average wages, mechanically moving the wage gap in favor of the majority group.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nash-in-Nash Bargaining:&lt;/strong&gt; The bargaining protocol used in the dynamic search model, following Horn and Wolinsky (1988). Each bilateral worker-firm bargain splits the available surplus in proportion to exogenous bargaining power parameter Δ ∈ (0,1), taking as given the outcome of all other bilateral bargains. A worker&amp;rsquo;s disagreement point is the wage she would receive from bargaining with the next firm in her search order. This generates the result that a worker&amp;rsquo;s realized payoff is increasing in the number of segregated (non-EPSW-constrained) firms competing for her, connecting firm segregation decisions to wage determination.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reallocation Friction:&lt;/strong&gt; In the dynamic search model, represented by a low departure probability d ∈ (0,1) for existing employees. When d is low, firms retain a large fraction of their workforce across periods, making segregation costly because the firm must separate from any existing workers of the &amp;ldquo;wrong&amp;rdquo; group. The paper shows (Proposition 5) that for sufficiently large d (low frictions), the equal-profit condition approximately pins down the number of firms segregating toward each group, and for d above a threshold, the majority group attracts weakly more segregating firms.&lt;/p&gt;</description></item><item><title>Firm Accommodation After Workplace Disability: Labor Market Impacts and Implications for Subsidy Design</title><link>https://macropaperwarehouse.com/papers/firm-accommodation-after-workplace-disability-labor-market-impacts-and-implications-for-subsidy-design/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/firm-accommodation-after-workplace-disability-labor-market-impacts-and-implications-for-subsidy-design/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper studies (1) how firm accommodation decisions respond to financial incentives in the context of workplace disability under workers&amp;rsquo; compensation, (2) what the causal effect of accommodation is on workers&amp;rsquo; subsequent labor market outcomes, and (3) whether the equilibrium level of accommodation is socially efficient, and what the welfare implications of wage subsidies for accommodation are.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Context and Data&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The analysis uses the universe of Oregon workers&amp;rsquo; compensation claims from 2005 through 2017 — over 131,000 disabling claims — linked to longitudinal quarterly earnings records from the Oregon Employment Department. The setting exploits Oregon&amp;rsquo;s Employer at Injury Program (EAIP), which subsidizes employers who provide &amp;ldquo;transitional work&amp;rdquo; accommodations (primarily through wage subsidies) to workers with temporary workplace disabilities. EAIP accounts for roughly 25 percent of claims on average, with the wage subsidy component representing over 96 percent of EAIP expenses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identification Strategy&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors exploit a policy change in July 2013 that reduced the EAIP wage subsidy rate from 50 percent to 45 percent. They construct a firm-level &amp;ldquo;exposure&amp;rdquo; measure — the fraction of a firm&amp;rsquo;s claims that used EAIP in a baseline period (2005–2009) — and estimate a continuous difference-in-differences specification in which the interaction of exposure and a post-2013 indicator instruments for accommodation. The identifying assumption is strong parallel trends: firms with low baseline exposure are unlikely to respond to the subsidy reduction, while high-exposure firms respond more, generating cross-firm variation in accommodation rates after 2013. An MTE framework (Heckman and Vytlacil 2005) is then used to explore heterogeneous treatment effects along an unobserved resistance-to-treatment dimension.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Empirical Findings&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The subsidy reduction from 50% to 45% decreased accommodation rates by &lt;strong&gt;2.9 percentage points&lt;/strong&gt; (9.3 percent) for claims in firms with average exposure, implying a subsidy elasticity of accommodation of 0.9.&lt;/li&gt;
&lt;li&gt;The policy change led to a &lt;strong&gt;0.95 percentage point decrease in employment&lt;/strong&gt; and a &lt;strong&gt;$120 decrease in quarterly earnings&lt;/strong&gt; four quarters after disability for claims in average-exposure firms (roughly 1.3–1.5 percent declines relative to means), with no significant effect on worker turnover to other firms.&lt;/li&gt;
&lt;li&gt;IV estimates of the effect of accommodation itself (using predicted EAIP as instrument) show &lt;strong&gt;accommodation increases the probability of employment four quarters after disability by 33 percentage points&lt;/strong&gt; and &lt;strong&gt;increases quarterly earnings by approximately $4,100&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The MTE analysis reveals &lt;strong&gt;negative selection on gains&lt;/strong&gt;: workers with workplace disabilities who are least likely to receive accommodation have the highest potential gains from it, driven largely by severe disabilities with high accommodation costs.&lt;/li&gt;
&lt;li&gt;Descriptive and IV evidence is consistent with accommodation operating primarily as &lt;strong&gt;general human capital investment&lt;/strong&gt;: accommodation has no statistically significant effect on the probability of moving to a new firm, and earnings gains are not systematically lower for workers who change employers after accommodation.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Structural Model and Counterfactual Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A two-period frictional labor market model with risk-averse workers, risk-neutral firms, Nash bargaining, imperfect experience rating in workers&amp;rsquo; compensation, and firm accommodation as human capital investment is developed and estimated. Two inefficiency sources are identified: (1) a human capital externality — because accommodation builds general human capital, firms cannot capture the full surplus when workers separate, reducing accommodation incentives; and (2) a fiscal externality — imperfectly experience-rated firms do not fully internalize the workers&amp;rsquo; compensation cost savings from accommodation, further depressing it below the efficient level. Counterfactual simulations show:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Eliminating wage subsidies (from 50% to 0%) reduces accommodation rates from &lt;strong&gt;33% to 11%&lt;/strong&gt;, leading to a &lt;strong&gt;7% decline in post-disability employment&lt;/strong&gt; and a &lt;strong&gt;15% decline in post-disability quarterly wages&lt;/strong&gt; (roughly $1,358).&lt;/li&gt;
&lt;li&gt;A revenue-neutral reform eliminating wage subsidies reduces average welfare and the welfare of &lt;strong&gt;more than 90% of workers&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Welfare gains from the subsidy are &lt;strong&gt;larger for low-skilled workers&lt;/strong&gt; than high-skilled workers.&lt;/li&gt;
&lt;li&gt;Conditional on experiencing disability, eliminating wage subsidies decreases welfare by about &lt;strong&gt;10%&lt;/strong&gt;, while increasing the subsidy to 100% raises welfare for disabled workers by around &lt;strong&gt;30%&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Firm profit is maximized at a subsidy rate around 80%, after which higher taxes offset accommodation gains.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-employer-at-injury-program-eaip-and-how-does-it-differ-from-standard-workers-compensation"&gt;Q1. What is the Employer at Injury Program (EAIP), and how does it differ from standard workers&amp;rsquo; compensation?&lt;/h3&gt;
&lt;p&gt;A1: EAIP is an optional component of Oregon&amp;rsquo;s workers&amp;rsquo; compensation system that subsidizes employers for the costs of accommodating workers with temporary disabilities during a transitional return-to-work period. Unlike standard workers&amp;rsquo; compensation premiums (which are experience-rated at the firm level), EAIP is funded through a flat payroll tax on all firms that is not experience-rated — meaning firms that use EAIP do not pay higher premiums. The wage subsidy component accounts for over 96 percent of EAIP expenses; other reimbursable costs (worksite modifications up to $5,000, retraining up to $1,000, clothing up to $400) are rarely used. Eligible employers must be the employer at which the disability occurred, and accommodation is limited to a transitional period during which workers cannot simultaneously receive time-loss benefits.&lt;/p&gt;
&lt;h3 id="q2-how-is-firm-level-exposure-constructed-and-what-is-the-rationale-for-using-it-as-an-instrument"&gt;Q2. How is firm-level &amp;ldquo;exposure&amp;rdquo; constructed, and what is the rationale for using it as an instrument?&lt;/h3&gt;
&lt;p&gt;A2: Exposure is the fraction of a firm&amp;rsquo;s workers&amp;rsquo; compensation claims that used EAIP during a five-year baseline period from 2005 to 2009 — a separate historical period chosen to reduce volatility and avoid mean-reversion. The rationale draws on prior work (Aizawa et al., 2022) showing that firm fixed effects account for nearly 25 percent of variation in accommodation, far more than worker or disability characteristics (1 and 3 percent, respectively), suggesting permanent firm-level heterogeneity in the relative benefits and costs of accommodation. Firms with zero historical exposure are unlikely to change accommodation behavior in response to a subsidy reduction, while high-exposure firms respond more, creating differential quasi-experimental variation in accommodation rates after July 2013.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-first-stage-and-reduced-form-results-from-the-did-specification"&gt;Q3. What are the first-stage and reduced-form results from the DID specification?&lt;/h3&gt;
&lt;p&gt;A3: The first-stage DID coefficient shows that a ten-percentage-point increase in exposure is associated with a one-percentage-point decrease in EAIP take-up after 2013, implying a 2.9 percentage point decrease for claims in firms with average exposure (mean 0.27). The corresponding reduced-form results show a 0.35 percentage point decrease in employment four quarters post-disability and a $45 decrease in quarterly earnings for every ten-percentage-point increase in exposure, scaling to 0.95 percentage points and $120 at average exposure. There is no statistically significant effect on the probability of moving to a new firm. Pre-trend tests show parallel accommodation trends across exposure terciles prior to 2013, supporting the identifying assumption.&lt;/p&gt;
&lt;h3 id="q4-what-do-the-iv-estimates-imply-about-the-causal-effect-of-accommodation-on-labor-market-outcomes"&gt;Q4. What do the IV estimates imply about the causal effect of accommodation on labor market outcomes?&lt;/h3&gt;
&lt;p&gt;A4: Under the exclusion restriction that the subsidy change affects labor market outcomes only through accommodation, the IV estimates imply that receipt of accommodation increases the probability of employment four quarters after disability by &lt;strong&gt;33 percentage points&lt;/strong&gt; (against a mean of 72 percent) and increases quarterly earnings by approximately &lt;strong&gt;$4,100&lt;/strong&gt; (against a mean of $7,807). There is no significant effect on the probability of working at a new firm four quarters later. The authors note these large estimates reflect local average treatment effects for compliers — workers whose accommodation status was changed by the instrument — who disproportionately have high unobserved resistance to treatment and high accommodation returns, explaining the magnitude.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-mte-framework-reveal-about-the-distribution-of-accommodation-effects-and-selection"&gt;Q5. What does the MTE framework reveal about the distribution of accommodation effects and selection?&lt;/h3&gt;
&lt;p&gt;A5: The MTE curves show that workers with the highest unobserved resistance to treatment (least likely to receive accommodation) have the highest potential employment and earnings gains from accommodation. This negative selection on gains arises because these workers tend to have worse employment outcomes in the untreated state, consistent with more severe disabilities commanding higher accommodation costs. IV weights are concentrated at high-resistance values, explaining the large IV estimates. Negative selection on gains is also found along observable dimensions: workers in self-insured firms, healthcare support occupations, women, and those with wounds/cuts/burns show larger gains but lower likelihood of receiving accommodation.&lt;/p&gt;
&lt;h3 id="q6-what-evidence-supports-characterizing-firm-accommodation-as-general-rather-than-firm-specific-human-capital-investment"&gt;Q6. What evidence supports characterizing firm accommodation as general rather than firm-specific human capital investment?&lt;/h3&gt;
&lt;p&gt;A6: Three pieces of evidence point toward general human capital. First, the IV estimate shows accommodation has no statistically significant effect on the probability of working at a new firm four quarters after disability. Second, a triple-interaction specification (DID interacted with new-firm indicator) yields suggestive evidence of even larger earnings gains for workers who move to a new firm post-accommodation, though this is not statistically significant — a pattern inconsistent with firm-specific human capital. Third, the subset of claims that receive non-wage EAIP benefits (worksite modifications, retraining) do show lower mobility, but this comprises fewer than 5 percent of the sample, meaning the predominant form of investment in the context is general in nature.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-two-sources-of-market-inefficiency-in-accommodation-identified-in-the-model"&gt;Q7. What are the two sources of market inefficiency in accommodation identified in the model?&lt;/h3&gt;
&lt;p&gt;A7: The first is a human capital externality operating through worker turnover. Because accommodation builds general human capital that workers carry to new employers, a firm accommodating a worker does not capture the portion of future surplus that accrues to future employers upon separation. In a Nash bargaining framework with lack of commitment, this dynamic inefficiency is larger when industry-wide turnover rates are higher — consistent with the descriptive finding that accommodation rates are strongly negatively associated with industry separation rates. The second is a fiscal externality from imperfect experience rating: firms whose workers&amp;rsquo; compensation premiums are not fully linked to their own claim costs do not fully internalize the cost-savings from accommodation (i.e., reduced time-loss benefit payments), leading them to accommodate at inefficiently low rates.&lt;/p&gt;
&lt;h3 id="q8-how-is-heterogeneity-incorporated-in-the-structural-estimation-and-what-do-the-estimated-parameters-show"&gt;Q8. How is heterogeneity incorporated in the structural estimation, and what do the estimated parameters show?&lt;/h3&gt;
&lt;p&gt;A8: The model incorporates observed heterogeneity (firm insurance status, worker skill type — measured by pre-disability wages — firm baseline exposure, and pre/post policy change) and unobserved heterogeneity mapped to the MTE framework&amp;rsquo;s unobserved resistance to treatment. Indirect inference matches cross-sectional accommodation rates, earnings by subgroup, and the DID coefficients. Key findings: net output during the disability period is negative (accommodation is a costly short-run investment), while post-disability output is higher for accommodated workers. Low-skilled workers experience larger productivity gains from accommodation than high-skilled workers. Accommodation cost shock variance is lower for higher unobserved types, meaning high-gain workers are also more sensitive to subsidy changes, consistent with the large IV estimates. The model fits the DID coefficients for accommodation, employment, and wages well.&lt;/p&gt;
&lt;h3 id="q9-what-do-the-counterfactual-simulations-show-about-the-welfare-effects-of-varying-the-subsidy-rate"&gt;Q9. What do the counterfactual simulations show about the welfare effects of varying the subsidy rate?&lt;/h3&gt;
&lt;p&gt;A9: Eliminating wage subsidies from the current 50% rate reduces the accommodation rate from 33% to 11% and lowers post-disability employment by 7 percentage points and post-disability quarterly wages by 15% ($1,358). From a welfare perspective, eliminating subsidies in a revenue-neutral reform reduces average ex-ante worker welfare and lowers welfare for more than 90% of workers. Conditional on experiencing disability, eliminating subsidies reduces welfare by about 10% while raising the subsidy to 100% increases welfare of disabled workers by around 30%. Firm profit is increasing in the subsidy rate up to about 80%, then decreases. Ex-ante worker welfare gains from the current 50% subsidy relative to no subsidy are modest in consumption-equivalent terms (at most 0.6% increase in consumption), partly because the disability probability is low (2.2%) and because unaccommodated workers still receive two-thirds wage replacement through time-loss benefits.&lt;/p&gt;
&lt;h3 id="q10-what-distributional-implications-do-wage-subsidies-have-across-worker-and-firm-types"&gt;Q10. What distributional implications do wage subsidies have across worker and firm types?&lt;/h3&gt;
&lt;p&gt;A10: Welfare gains from higher wage subsidies are larger for low-skilled workers than high-skilled workers, so the subsidy has a redistributive dimension beyond efficiency correction. Welfare gains are also larger for workers in imperfectly experience-rated firms, where the fiscal externality creates the greater wedge from the efficient level. Self-insured firms, which already internalize workers&amp;rsquo; compensation cost savings and thus accommodate closer to the optimal rate, benefit less from the subsidy and can even be made worse off if subsidies are set very high (since they bear higher flat payroll taxes with smaller marginal accommodation gains). The fraction of worker-firm matches experiencing welfare gains exceeds 90% under the benchmark subsidy level, indicating broad rather than narrowly concentrated gains.&lt;/p&gt;
&lt;h3 id="q11-how-do-the-experience-rating-channel-and-the-worker-turnover-channel-interact-in-comparative-statics"&gt;Q11. How do the experience-rating channel and the worker-turnover channel interact in comparative statics?&lt;/h3&gt;
&lt;p&gt;A11: Model comparative statics show that reducing the job-to-job transition rate of workers with disabilities to one-quarter of its estimated value substantially raises accommodation rates, and this effect is more pronounced for imperfectly experience-rated firms than for self-insured firms. This occurs because self-insured firms already have a strong incentive to accommodate (to reduce workers&amp;rsquo; compensation premiums), so turnover is less marginal for them. Forcing all firms to be self-insured (perfect experience rating) would substantially increase accommodation rates in currently imperfectly rated firms. Lowering the accommodation cost during the disability period (increasing net output during the disability period) also raises accommodation rates for both firm types.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Firm Accommodation (EAIP):&lt;/strong&gt; In this paper&amp;rsquo;s specific sense, accommodation refers to a firm&amp;rsquo;s decision to offer a worker with a temporary workplace disability &amp;ldquo;transitional work&amp;rdquo; — alternative tasks, modified duties, or flexible arrangements — during their recovery period, funded in part through Oregon&amp;rsquo;s Employer at Injury Program wage subsidy. Accommodation is distinct from simple early return to work; it functions as a form of human capital investment by potentially providing skill development opportunities and preventing human capital depreciation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exposure (Instrument):&lt;/strong&gt; A firm-level continuous measure defined as the fraction of a firm&amp;rsquo;s workers&amp;rsquo; compensation claims that used EAIP during a five-year baseline period (2005–2009). Exposure captures permanent, time-invariant firm-level propensity to accommodate, and is used to construct a difference-in-differences instrument for the causal effect of accommodation by interacting exposure with a post-2013 indicator (when the subsidy rate was cut from 50% to 45%).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Imperfect Experience Rating:&lt;/strong&gt; The degree to which a firm&amp;rsquo;s workers&amp;rsquo; compensation insurance premium adjusts to reflect that firm&amp;rsquo;s own claims costs, rather than being set at an industry average. Fully experience-rated (self-insured) firms internalize 100% of claim costs and thus have strong incentives to accommodate. Partially experience-rated firms face a fiscal externality: because their premiums do not fully reflect their own time-loss benefit expenditures, they do not capture all the cost savings from accommodating workers, leading to under-accommodation relative to the social optimum.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Human Capital Externality (Dynamic Inefficiency in Accommodation):&lt;/strong&gt; The mechanism — analogous to Acemoglu and Pischke (1999) and Fang and Gavazza (2011) — by which worker turnover reduces firms&amp;rsquo; incentives to invest in general human capital (here, accommodation). When accommodation raises workers&amp;rsquo; general productivity, part of the future surplus from this investment accrues to future employers upon job-to-job separation. With Nash bargaining and lack of commitment (re-bargaining in the second period), the accommodating firm cannot capture this surplus, creating a dynamic inefficiency that is more severe in high-turnover industries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negative Selection on Gains:&lt;/strong&gt; The empirical finding, established via the MTE framework, that workers with workplace disabilities who are least likely to receive accommodation (highest unobserved resistance to treatment) have the largest potential employment and earnings gains from accommodation. This pattern arises because workers with more severe disabilities have high accommodation costs (making firms unwilling to accommodate them) but also face far worse counterfactual labor market outcomes without accommodation, creating large potential gains.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marginal Treatment Effect (MTE):&lt;/strong&gt; Following Heckman and Vytlacil (2005), the treatment effect of accommodation evaluated at a specific quantile of unobserved resistance to treatment — defined here as the propensity score value at which a worker is indifferent between treatment and non-treatment. The MTE curve maps out the full distribution of treatment effects and reveals who benefits (and by how much), how IV estimates are weighted averages over this distribution, and which compliers drive the large IV estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;General vs. Firm-Specific Human Capital (in Accommodation Context):&lt;/strong&gt; Accommodation is characterized as general human capital investment if the productivity and earnings gains it produces are transferable across employers — i.e., if accommodated workers who move to new firms retain their wage gains. It is firm-specific if gains are tied to the current match. In this paper, general human capital is supported by the null effect of accommodation on new-firm employment probability, suggestive evidence of non-lower (possibly larger) earnings gains for new-firm movers, and the observation that fewer than 5% of claims use non-wage EAIP benefits associated with firm-specific investment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Revenue-Neutral Counterfactual:&lt;/strong&gt; A counterfactual policy experiment in which the wage subsidy rate for accommodation is varied while imposing that both the time-loss benefit program and the EAIP wage subsidy program remain budget-balanced. Higher subsidy rates raise firm accommodation, reduce time-loss benefit payouts (lowering base premiums for imperfectly experience-rated firms), but require a higher flat EAIP payroll tax on all firms, some of which is passed through to workers via lower first-period wages.&lt;/p&gt;</description></item><item><title>Gendered Spheres of Learning and Household Decision-Making over Fertility</title><link>https://macropaperwarehouse.com/papers/gendered-spheres-of-learning-and-household-decision-making-over-fertility/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/gendered-spheres-of-learning-and-household-decision-making-over-fertility/</guid><description>&lt;p&gt;This paper investigates whether information asymmetries within households about maternal health risk can explain persistent spousal disagreement over fertility in a high-fertility, high-maternal-mortality setting. The authors develop a theoretical model and conduct a randomized field experiment among approximately 500 couples in peri-urban Lusaka, Zambia, where the lifetime risk of maternal death is 1 in 59 women and the maternal mortality ratio is 398 deaths per 100,000 live births.&lt;/p&gt;
&lt;p&gt;The central mechanism is a communication barrier that arises from conflicting fertility preferences between spouses. When husbands have higher desired fertility than wives (4.43 vs. 4.19 children on average in the study sample), wives who are better informed about maternal health risk lack the incentive to credibly transmit that information to their husbands. Strategic communication concerns — not a generically lower propensity of men to learn from women — drive this asymmetry. The model predicts a pooling equilibrium in which no informative communication flows from wives to husbands when preference divergence is sufficiently large.&lt;/p&gt;
&lt;p&gt;The experiment randomized whether the maternal mortality information curriculum was delivered to the husband or the wife in each couple, with both spouses in all arms also receiving a family planning curriculum. This design isolates the incremental effect of the maternal mortality information and permits identification of direct versus spillover effects within the household.&lt;/p&gt;
&lt;p&gt;Consistent with the model, treated husbands significantly update their beliefs about maternal health risk factors, and their wives also update — information flows from husbands to wives. By contrast, treated wives update their own beliefs, but their husbands do not update at all. The test that spillover effects are symmetric is rejected (p-value = 0.097 for risk factors index; p-value &amp;lt; 0.001 for direct vs. indirect effects on men). The communication asymmetry is most pronounced among husbands who, at baseline, want a child as soon as possible — precisely the households with the greatest preference conflict.&lt;/p&gt;
&lt;p&gt;Both treatment arms reduce fertility. Households in which the husband is treated experience a 43% reduction in the probability of having a child or being pregnant in the year following the intervention. The fertility reduction is strongest when the wife faces higher ex ante risk based on her birth history, consistent with the model&amp;rsquo;s prediction that treatment effects are concentrated among households with high maternal health costs.&lt;/p&gt;
&lt;p&gt;The transfers evidence is the key differentiator between the two arms. When the wife is treated, fertility declines but is accompanied by a significant reduction in transfers from husband to wife, consistent with the wife updating her own beliefs without being able to convey them to her husband, who then reduces compensation. When the husband is treated, fertility declines without the same reduction in transfers — and treated husbands report higher communication with their spouse about family planning and higher relationship satisfaction. This combination is consistent with the husband treatment resolving the information gap directly, enabling efficient contracting, whereas the wife treatment leaves the information asymmetry in place.&lt;/p&gt;
&lt;p&gt;The study is conducted in informal settlements of Lusaka, a prime-age urban sample in which the average woman is 28 years old with 2.6 children at baseline. Scope conditions: results apply to a setting with very high maternal mortality, large baseline spousal fertility gaps, and strong traditional beliefs (55.5% of men cite marital infidelity as a leading cause of maternal complications). Generalizability to lower-risk or lower-preference-gap settings is explicitly circumscribed by the model&amp;rsquo;s comparative statics.&lt;/p&gt;
&lt;p&gt;Q: What is the baseline gender gap in knowledge of maternal health risk?
A: Men are less likely than women to identify high parity (72.0% vs. 77.7%) and advanced maternal age (74.3% vs. 84.6%) as risk factors. In seven hypothetical scenarios rating complication likelihood on a 0–10 scale, men report lower scores than women in six out of seven cases. Despite Zambia&amp;rsquo;s 1-in-59 lifetime maternal mortality risk, only 27.6% of men (vs. 53.4% of women) report having attempted to discuss maternal health risk with their spouse.&lt;/p&gt;
&lt;p&gt;Q: What drives the gender gap in knowledge?
A: The authors argue the gap stems from &amp;ldquo;gendered spheres of direct and indirect knowledge accumulation of maternal labor and delivery outcomes.&amp;rdquo; Women are embedded in social networks where maternal mortality episodes are more salient: 11.0% of women report knowing a close friend who died giving birth, vs. 6.8% of men knowing a close friend whose wife died. The gap widens with social distance to the victim, suggesting women&amp;rsquo;s networks give them systematically more exposure to maternal mortality events.&lt;/p&gt;
&lt;p&gt;Q: How does the model explain the failure of within-household communication?
A: The model places husband and wife preferences as minimizing the distance between realized fertility and their respective net fertility optima (ideal fertility minus weighted maternal health cost). When the husband&amp;rsquo;s ideal fertility is high enough, he makes transfers to induce the wife to bear more children than her private optimum. Given these incentives, a wife who is informed about high health costs has an interest in exaggerating the cost to extract larger transfers. Because the husband anticipates this, no informative communication occurs in equilibrium — the only equilibrium is a pooling equilibrium where the wife&amp;rsquo;s message is uninformative regardless of her true cost realization.&lt;/p&gt;
&lt;p&gt;Q: What is the specific asymmetry in belief updating observed in the experiment?
A: Among treated husbands, both husbands and their wives update beliefs about maternal risk factors — information flows from husband to wife. Among treated wives, only the wife updates; her husband does not. The Wald test rejects equal direct and indirect effects on men at p &amp;lt; 0.001 and rejects symmetric spillovers at p = 0.097 for the risk factors index. There is no symmetric restriction binding for women&amp;rsquo;s updating across arms.&lt;/p&gt;
&lt;p&gt;Q: How large is the fertility effect and which arm drives it?
A: Households in which the husband is treated experience a 43% reduction in the probability of having a child or being pregnant in the year following the intervention. This effect is described as of the same order of magnitude as other household-level interventions shown to reduce pregnancy (citing Ashraf, Field, and Lee 2014). The fertility reduction is strongest among households where the woman faces higher ex ante risk based on birth history, consistent with the model&amp;rsquo;s Prediction 5 that effects are concentrated where theta_j is high.&lt;/p&gt;
&lt;p&gt;Q: How do transfers differ between the wife-treated and husband-treated arms?
A: When the wife is treated, the fertility decline is accompanied by a significant reduction in transfers from husband to wife. When the husband is treated, the fertility decline is not accompanied by a similar reduction in transfers. The authors interpret this pattern as: wife treatment leaves the husband uninformed, so he reduces transfers when he observes her reducing fertility without understanding why; husband treatment resolves the information gap, allowing efficient renegotiation without penalizing the wife.&lt;/p&gt;
&lt;p&gt;Q: Which husbands fail to update beliefs even when their wife is treated?
A: Husbands who at baseline want a child &amp;ldquo;as soon as possible&amp;rdquo; do not update their beliefs in response to their wife&amp;rsquo;s treatment status. These men also reduce transfers to their wife more than other groups when she is treated. In the model, these are precisely the households with the highest conflict of interest (high alpha_H), where the pooling equilibrium prediction is sharpest.&lt;/p&gt;
&lt;p&gt;Q: What is the role of traditional beliefs about maternal mortality?
A: 55.5% of men and 42.0% of women report (without prompting) marital infidelity as a leading cause of maternal labor and delivery complications — greater weight than assigned to lack of healthcare and poor health status combined. This stigma directly reduces women&amp;rsquo;s willingness to raise concerns about birth complications with their spouse, reinforcing the communication barrier the model formalizes.&lt;/p&gt;
&lt;p&gt;Q: What are the welfare implications of targeting men vs. women with information?
A: The fertility reduction from husband treatment is not inferior to that from wife treatment, but husband treatment also produces improvements in marital surplus — treated husbands report higher communication with spouse about family planning, higher relationship satisfaction, and greater closeness — whereas wife treatment reduces transfers to the wife, indicating she bears a financial cost. The authors argue male-targeted information can reduce unmet need for family planning while enhancing rather than exacerbating household conflict.&lt;/p&gt;
&lt;p&gt;Q: Does this paper provide field experimental evidence on strategic communication models?
A: The authors claim this is the first field experimental evidence directly testing models of strategic communication (Crawford and Sobel 1982; Mailath 1987; Crawford 1998, 2019), wherein persistent preference differences and conflict of interest impede communication and beliefs updating. Prior tests of these models were conducted in the lab; this paper provides the first real-world behavioral test with consequential decisions (fertility) in a high-stakes setting.&lt;/p&gt;
&lt;p&gt;Q: What is the unmet need for family planning in the study sample?
A: Overall, 32% of women in the sample report not using modern contraceptives at baseline. Of the 33% of women who want no more children, 27% are not using any modern contraceptive (8% of the overall sample). Of the 52% of women who wish to delay giving birth by at least one year, 23% are not using any modern contraceptive (12% of the overall sample).&lt;/p&gt;
&lt;p&gt;Q: How does the model characterize the husband&amp;rsquo;s partial internalization of maternal health costs?
A: The husband&amp;rsquo;s utility function includes the maternal health cost theta_j scaled by delta (0 ≤ delta ≤ 1), capturing how much weight he places on his wife&amp;rsquo;s risk. When delta is sufficiently high and the husband&amp;rsquo;s ideal fertility (alpha_H) is sufficiently low, or when his disutility of transfers (gamma) is sufficiently low, informative communication can occur after the husband is treated. When delta is low, the husband discounts his wife&amp;rsquo;s risk and communication barriers are more severe regardless of treatment.&lt;/p&gt;
&lt;p&gt;Maternal health cost (theta): A random variable representing the welfare cost borne by the wife from childbearing, including mortality risk and morbidity. In Zambia, distributed with a higher mean than the worldwide distribution. Enters the wife&amp;rsquo;s utility directly and the husband&amp;rsquo;s utility only scaled by delta, his degree of internalization of her cost.&lt;/p&gt;
&lt;p&gt;Gendered spheres of learning: The paper&amp;rsquo;s term for the systematic differential in experiential exposure to maternal mortality outcomes between men and women, arising from gender-segregated social networks. Women witness maternal mortality events more directly through closer social ties, while men&amp;rsquo;s networks provide systematically less exposure.&lt;/p&gt;
&lt;p&gt;Communication barrier (pooling equilibrium): The equilibrium outcome in the model where no informative signal is transmitted from an informed wife to her uninformed husband about the true realization of maternal health cost. Arises because the wife&amp;rsquo;s incentives to misreport are independent of the true cost realization, making any message uninformative when preference conflict is sufficiently large.&lt;/p&gt;
&lt;p&gt;Intra-household information spillover: The transmission of information learned by one spouse to the other as a consequence of the treated spouse&amp;rsquo;s belief update. The paper documents asymmetric spillovers: information flows from treated husbands to their wives, but not from treated wives to their husbands.&lt;/p&gt;
&lt;p&gt;Husband&amp;rsquo;s demand for children (alpha_H): The husband&amp;rsquo;s ideal fertility level, which governs the degree of preference conflict within the household. Baseline husband desire for a child as soon as possible serves as the empirical proxy for high alpha_H and is the key moderator of spillover and transfer effects.&lt;/p&gt;
&lt;p&gt;Degree of internalization (delta): The parameter in the husband&amp;rsquo;s utility function (0 ≤ delta ≤ 1) capturing how much weight he places on his wife&amp;rsquo;s maternal health cost. When delta is high and gamma (disutility of transfers) is low, communication can occur in equilibrium after the husband is treated.&lt;/p&gt;
&lt;p&gt;Unmet need for family planning: Women who wish to space or limit births but are not using modern contraception. In the study sample, 32% of women report not using modern contraceptives at baseline, with substantial shares among both those wanting no more children and those wishing to delay.&lt;/p&gt;</description></item><item><title>Genetic Prediction and Adverse Selection</title><link>https://macropaperwarehouse.com/papers/genetic-prediction-and-adverse-selection/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/genetic-prediction-and-adverse-selection/</guid><description>&lt;p&gt;This paper asks how much adverse selection would arise in critical illness insurance (CII) markets if consumers can observe polygenic indexes (PGIs) — genetic risk scores derived from millions of genetic variants — while insurers are legally barred from using genetic information. The authors develop an econometric method that measures selection under current PGI technology, then extends identification to expected future PGI accuracy using heritability bounds, even though future PGIs are not yet observable in data.&lt;/p&gt;
&lt;p&gt;The primary dataset is the UK Biobank (UKB), comprising approximately 446,570 genotyped individuals of European-like ancestry linked to NHS electronic health records. The authors study seven single-disease CII contracts (Alzheimer&amp;rsquo;s disease, breast cancer, coronary artery disease, colorectal cancer, prostate cancer, schizophrenia, and type 2 diabetes) and multiple-disease bundled contracts paying a lump sum upon onset. The econometric model assumes a probit disease probability, Gaussian PGI structure, and identification relies on published heritability estimates to pin down future PGI predictive power. The key selection metric is the implicit tax proposed by Hendren (2013): the percentage markup a marginal consumer must pay above her actuarially fair price due to adverse selection. The authors use the minimum implicit tax up to the 80th percentile of risk (t80) as their summary statistic, with market unraveling benchmarked at t80 between 43% and 83% from prior literature.&lt;/p&gt;
&lt;p&gt;The paper reports three main findings, all scoped to a population of 35-year-olds in the standard insurer risk class (those whose predicted risk falls within 0.75–1.25 times the population mean).&lt;/p&gt;
&lt;p&gt;First, under current PGI technology with full consumer adoption, selection is noticeable but heterogeneous across diseases. t80 ranges from 17.9% for coronary artery disease to 117.9% for Alzheimer&amp;rsquo;s disease. Coronary artery disease and colorectal cancer fall in the middle of the no-unraveling range; breast cancer, schizophrenia, and type 2 diabetes fall between the no-unraveling and unraveling ranges; Alzheimer&amp;rsquo;s disease and prostate cancer (t80 = 59.8%) reach or exceed the unraveling range. The current prostate cancer PGI explains 9.9% of liability variance, adding 8.3 percentage points over the 22.9% explained by non-genetic covariates.&lt;/p&gt;
&lt;p&gt;Second, under expected future PGI accuracy — bounded below by SNP heritability and above by twin heritability — selection becomes potentially crippling. Under the lower bound (Scenario 3L), t80 ranges from 57.5% for breast cancer to above 1,000% for Alzheimer&amp;rsquo;s. Under the upper bound (Scenario 3U), t80 exceeds 100% for all seven single-disease contracts and exceeds 1,000% for three of them. For prostate cancer, the reference case, t80 reaches 86.8% under Scenario 3L and 426.9% under Scenario 3U — far above Hendren&amp;rsquo;s unraveling benchmarks. For multiple-disease male contracts, t80 = 30.8% under current technology, rising to 54.4% (Scenario 3L) and 243.9% (Scenario 3U).&lt;/p&gt;
&lt;p&gt;Third, variation in selection across contracts is driven primarily by: the predictive power of the future PGI, the incremental predictive power over non-genetic covariates, and disease prevalence. Alzheimer&amp;rsquo;s and schizophrenia — high heritability, low prevalence — display the highest implicit taxes; breast and colorectal cancer — lower SNP heritability, lower incremental R2 — display the lowest.&lt;/p&gt;
&lt;p&gt;These findings are corroborated by a calibrated Akerlof-Einav-Finkelstein equilibrium model using HRS data: current PGI availability reduces equilibrium market quantity from 30% to 21.4%; future PGI availability drives equilibrium quantity to zero in a full adverse selection death spiral. Partial take-up robustness checks show that even at 50% consumer adoption, selection remains problematically high under future PGI accuracy for most contracts. The analysis is restricted to individuals of European-like ancestry due to data availability constraints.&lt;/p&gt;
&lt;p&gt;Q: What is the core market failure the paper analyzes?
A: The paper analyzes adverse selection arising from an asymmetric information gap: consumers can observe PGI-based disease risk predictions from consumer genetic tests (e.g., 23andMe), while insurers in many jurisdictions are legally prohibited from requesting or using genetic information. This creates a situation where high-risk consumers have private information allowing them to sort into insurance, driving up average claims costs and potentially unraveling the market.&lt;/p&gt;
&lt;p&gt;Q: What is a polygenic index (PGI) and why does it differ from classical genetic testing?
A: A PGI is a weighted sum of millions of genetic variants (typically over one million) each with individually tiny effects, constructed using effect-size estimates from genome-wide association studies (GWASs). This contrasts with traditional genetic testing focused on rare single-gene mutations (e.g., BRCA for breast cancer or PKD for kidney disease), which are rare, explain small shares of population-level disease variance, and can largely be inferred from family history. PGIs target common polygenic diseases and are the primary driver of the adverse selection concern because they aggregate diffuse genetic signals into a meaningful risk prediction.&lt;/p&gt;
&lt;p&gt;Q: What are the current PGI R2 values for the seven diseases studied?
A: Estimated on the liability scale in the UKB, current PGI R2 values are: Alzheimer&amp;rsquo;s disease 7.1%, breast cancer 6.7%, coronary artery disease 2.5%, colorectal cancer 2.2%, prostate cancer 9.9%, schizophrenia 4.9%, and type 2 diabetes 7.4%. These represent the share of liability variance explained by each disease&amp;rsquo;s current PGI in the study sample.&lt;/p&gt;
&lt;p&gt;Q: How does the paper identify the degree of selection under future PGI technology that does not yet exist in the data?
A: The identification strategy combines three elements: the normality of PGI distributions, the relationship between current and future PGIs (the current PGI is modeled as a noisy version of the future PGI with an independent Gaussian error), and published heritability estimates that bound the future PGI&amp;rsquo;s predictive power. Theorem 1 establishes that under five stated assumptions — including a probit disease model and known future R2 from heritability studies — the full joint distribution of loss, current PGI, future PGI, and non-genetic covariates is identified from observed data.&lt;/p&gt;
&lt;p&gt;Q: What heritability bounds are used for the future PGI scenarios, and why two bounds?
A: Scenario 3L sets future PGI R2 equal to each disease&amp;rsquo;s SNP heritability (estimated from common genetic variants), which the authors treat as a conservative lower bound because future PGIs will also incorporate rarer variants with better effect-size precision. Scenario 3U sets future PGI R2 equal to twin heritability, treating it as an upper bound since the theoretical maximum predictive power of a PGI is the trait&amp;rsquo;s narrow-sense heritability. For prostate cancer, these bounds are 18.0% (SNP) and 57.0% (twin); for Alzheimer&amp;rsquo;s, SNP heritability is 33.1% and twin heritability is 58%.&lt;/p&gt;
&lt;p&gt;Q: What is the implicit tax and how is it used as a benchmark?
A: The implicit tax t(r) for a consumer with private risk r equals the percentage by which her insurance cost exceeds her own actuarially fair price when she must pool with all consumers of equal or higher risk. It measures how much the marginal buyer overpays due to adverse selection. The authors follow Hendren (2013) in reporting t80, the minimum implicit tax up to the 80th percentile. Hendren&amp;rsquo;s benchmarks: t80 between 7–35% for markets that did not unravel; t80 between 43–83% for markets that had unraveled.&lt;/p&gt;
&lt;p&gt;Q: What are the single-disease contract results under current PGI technology (Scenario 2)?
A: With full consumer adoption of current PGI technology, t80 ranges from 17.9% for coronary artery disease to 117.9% for Alzheimer&amp;rsquo;s disease. Coronary artery disease (17.9%) and colorectal cancer (26.5%) fall in the middle of Hendren&amp;rsquo;s no-unraveling range. Breast cancer (36.9%), schizophrenia (42.1%), and type 2 diabetes (37.0%) fall between the no-unraveling and unraveling ranges. Alzheimer&amp;rsquo;s disease (117.9%) and prostate cancer (59.8%) reach or exceed the unraveling range.&lt;/p&gt;
&lt;p&gt;Q: What are the single-disease contract results under future PGI technology?
A: Under the lower bound (Scenario 3L, R2 = SNP heritability), t80 ranges from 57.5% for breast cancer to above 1,000% for Alzheimer&amp;rsquo;s disease. Under the upper bound (Scenario 3U, R2 = twin heritability), t80 exceeds 100% for all seven contracts and exceeds 1,000% for three (Alzheimer&amp;rsquo;s, schizophrenia, and at least one other). These figures substantially exceed Hendren&amp;rsquo;s unraveled-market benchmarks for virtually all contracts.&lt;/p&gt;
&lt;p&gt;Q: What drives cross-disease variation in the implicit tax?
A: The authors identify three main drivers: the expected accuracy of future PGI (higher heritability → higher implicit tax), the incremental predictive power of the future PGI over non-genetic covariates observable by insurers (more incremental information → more adverse selection), and disease prevalence (lower prevalence concentrates risk heterogeneity, amplifying selection). Alzheimer&amp;rsquo;s disease and schizophrenia — high heritability and low prevalence — have the highest implicit taxes. Breast and colorectal cancers — lower SNP heritability and lower incremental R2 — have the lowest.&lt;/p&gt;
&lt;p&gt;Q: What do the multiple-disease bundled contract results show?
A: For the male multiple-disease contract under Scenario 2 (current PGI), t80 = 30.8%, comparable to Hendren&amp;rsquo;s no-unraveling range. Under Scenario 3L, t80 = 54.4%; under Scenario 3U, t80 = 243.9%, both in or above the unraveling range. The female contract yields qualitatively similar results. Implicit taxes in bundled contracts are generally lower than in single-disease contracts, suggesting some diversification of genetic risk across diseases.&lt;/p&gt;
&lt;p&gt;Q: What does the calibrated equilibrium model find?
A: Using an Akerlof (1970) / Einav-Finkelstein-Cullen (2010) supply-and-demand model calibrated to match a 30% market participation rate and a 50% loss ratio in the UK CII market, and using HRS data on individual risk aversion, the model finds that current PGI availability reduces equilibrium quantity from 30% to 21.4%. Future PGI availability (both Scenario 3L and 3U) drives equilibrium quantity to zero — a complete adverse selection death spiral with no trade.&lt;/p&gt;
&lt;p&gt;Q: How robust are results to partial consumer adoption of genetic testing?
A: At 10% consumer take-up, selection is low regardless of PGI accuracy. At 50% take-up, selection remains problematically high for all single-disease contracts under future PGI accuracy (Scenarios 3L and 3U). For multiple-disease contracts at 50% take-up, t80 falls just below Hendren&amp;rsquo;s unraveling threshold under Scenario 3L but enters the unraveling range under Scenario 3U. This suggests market problems would materialize once predictive power exceeds the SNP heritability bound and take-up exceeds roughly 50%.&lt;/p&gt;
&lt;p&gt;Q: What role do risk preferences play, and do they confound the results?
A: The authors test whether risk tolerance correlates with disease risk in the UKB using a self-reported general risk tolerance measure. They find extremely low correlations between risk tolerance and each disease. This is consistent with low correlation between relative risk aversion and disease risk in the HRS calibration, and supports the finding that correlation between risk and risk preferences is unlikely to meaningfully affect the main results.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s assessment of preventive treatment as a mitigating factor?
A: The authors acknowledge that genetic testing could enable personalized preventive medicine, which would reduce actual disease incidence among high-risk individuals. However, they argue this is unlikely to substantially affect their main findings because the most commonly covered diseases under CII are cancers, for which preventive behaviors have bounded effectiveness.&lt;/p&gt;
&lt;p&gt;Q: What are the paper&amp;rsquo;s policy implications?
A: The paper situates the genetic information problem within the standard regulatory framework for selection markets, distinguishing laissez-faire (allow genetic underwriting — efficient but potentially unfair to high-risk consumers), government provision (unattractive for non-essential CII), and managed competition (community rating combined with subsidies and risk adjustment). The authors argue that a full ban on genetic underwriting — the current policy in many countries — may become untenable as PGI accuracy improves, because it generates potentially crippling adverse selection. Some level of community rating may remain desirable for redistribution, but needs to be paired with subsidies or risk adjustment to prevent market collapse.&lt;/p&gt;
&lt;p&gt;Q: What are the main data and scope limitations?
A: The analysis is restricted to individuals of European-like ancestry because most large GWASs were conducted in European ancestry samples and PGIs perform poorly across ancestries. The UKB sample was aged 40–69 at recruitment and the analysis adjusts for age-dependent covariates; the HRS replication uses approximately 20,000 individuals. The equilibrium model ignores moral hazard and uses a parsimonious binary loss framework. The paper does not specify a timeline for when PGI accuracy will reach heritability bounds.&lt;/p&gt;
&lt;p&gt;Polygenic Index (PGI): A weighted sum of an individual&amp;rsquo;s genetic variants across the genome (typically over one million variants), constructed using effect-size estimates from a genome-wide association study (GWAS) conducted in an independent sample. It is a noisy proxy for the individual&amp;rsquo;s true additive genetic factor for a disease, and its predictive power is bounded above by the trait&amp;rsquo;s narrow-sense heritability.&lt;/p&gt;
&lt;p&gt;Implicit Tax: A measure of adverse selection defined by Hendren (2013) as the percentage by which a consumer with private risk r must overpay relative to her own actuarially fair price if she is pooled with all consumers of equal or higher risk. The minimum implicit tax up to the 80th percentile of risk (t80) serves as the paper&amp;rsquo;s primary summary statistic; t80 above roughly 43% is associated with market unraveling in prior literature.&lt;/p&gt;
&lt;p&gt;SNP Heritability: The share of variance in a disease&amp;rsquo;s liability attributable to the set of common genetic variants (SNPs) used in heritability estimation. Used in this paper as a conservative lower bound on the predictive power of future PGIs, because future PGIs will additionally capture rarer variants.&lt;/p&gt;
&lt;p&gt;Twin Heritability: An estimate of a trait&amp;rsquo;s narrow-sense (additive) heritability computed by comparing resemblance of monozygotic twins (sharing 100% of their genomes) to dizygotic twins (sharing ~50% on average). Used as an upper bound on future PGI predictive power, since heritability is the theoretical maximum R2 for a PGI.&lt;/p&gt;
&lt;p&gt;Standard Risk Class: The set of consumers whose predicted disease risk (based on non-genetic covariates observable to insurers) falls between 0.75 and 1.25 times the population-wide average risk, following standard insurance underwriting practice. Insurers charge the same premium to all consumers in this class; any variation in risk within the class due to private genetic information constitutes the source of adverse selection analyzed in this paper.&lt;/p&gt;
&lt;p&gt;Private Risk Function: The probability rho(g, w) of contracting the disease conditional on both the consumer&amp;rsquo;s observed PGI g and non-genetic factors w. Contrasted with the non-genetic private risk function pi(w), which conditions only on non-genetic covariates. The dispersion of the private risk distribution across consumers in the same risk class determines the degree of adverse selection.&lt;/p&gt;
&lt;p&gt;Adverse Selection Death Spiral: The Akerlof (1970) mechanism in which high-risk consumers disproportionately purchase insurance, causing insurers to raise premiums, which deters low-risk consumers, which further raises the average risk of purchasers, ultimately driving equilibrium quantity to zero. The paper&amp;rsquo;s calibrated equilibrium model finds this outcome under future PGI accuracy for the HRS CAD contract.&lt;/p&gt;</description></item><item><title>Germs in the Family: The Short- and Long-Term Consequences of Intra-Household Disease Spread</title><link>https://macropaperwarehouse.com/papers/germs-in-the-family-the-short-and-long-term-consequences-of-intra-household-disease-spread/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/germs-in-the-family-the-short-and-long-term-consequences-of-intra-household-disease-spread/</guid><description>&lt;p&gt;This paper studies the short- and long-term consequences of intra-household respiratory disease transmission from older to younger siblings in Danish families. The central research questions are: (1) how do respiratory illnesses spread from preschool-aged older siblings to younger infant siblings during the first year of life, and (2) how does respiratory disease exposure during infancy causally affect younger siblings&amp;rsquo; long-term economic, human capital, and health outcomes?&lt;/p&gt;
&lt;p&gt;The study uses population-level Danish administrative data covering 1,230,180 children from 37 birth cohorts (1981–2017), linking records from the National Patient Register, income and labor market registers, education registers, and psychiatric care registers. The identification strategy combines birth order variation in respiratory disease vulnerability with within-municipality variation in local respiratory disease prevalence among children aged 13–71 months. The authors construct a municipality-level disease exposure index—cumulative respiratory hospitalizations per 100 children aged 13–71 months in a child&amp;rsquo;s municipality over their first 12 months of life—and estimate the differential effect of this index on younger versus older siblings, controlling for municipality fixed effects, birth year-month fixed effects, and an extensive set of individual and family background characteristics.&lt;/p&gt;
&lt;p&gt;The descriptive findings are stark: younger siblings have 2–3 times higher rates of hospitalization for acute respiratory conditions during their first year of life compared to older siblings at the same age, with the gap largest at ages two and three months. The gap is larger for winter births, shorter birth spacing, and when older siblings attend childcare centers—all patterns consistent with the older sibling serving as a disease vector.&lt;/p&gt;
&lt;p&gt;On the causal estimates, moving from the 25th to the 75th percentile of the disease exposure index distribution increases the younger sibling&amp;rsquo;s acute respiratory hospitalizations in the first year of life by 0.023 (32.9 percent above the sample mean), with effects more than twice as large for exposure in the first six months compared to the second six months.&lt;/p&gt;
&lt;p&gt;In the long run, an interquartile increase in first-year respiratory disease exposure reduces younger siblings&amp;rsquo; wage earnings (conditional on employment) at ages 25–32 by 0.8 percent and total income by 0.8 percent, and reduces their income percentile rank by 0.3 percentage points. There is no significant effect on labor force participation at the extensive margin. Effects on earnings are approximately twice as large when exposure is measured in the first six months of life. These earnings effects are comparable in magnitude to those from a 10 percent reduction in birth weight or a 9 percent increase in ambient air pollution at birth, and correspond to roughly two-thirds of the adult earnings impact of in utero exposure to the 1918 Spanish Influenza. When the disease index interaction is included, the main birth order coefficient declines by approximately 70 percent, suggesting intra-household disease transmission is an important channel underlying the documented birth order earnings disadvantage.&lt;/p&gt;
&lt;p&gt;Additional findings include: a 0.5 percentage point reduction in high school graduation and a 0.6 percentage point reduction in college graduation (interquartile effects); a 0.01 standard deviation penalty in ninth grade Danish test scores; a 20 percent increase (0.016 per hundred per year) in chronic respiratory hospitalizations at ages 16–26; and a 6.1 percent increase (0.5 additional visits per hundred per year) in psychiatric clinic visits at ages 16–26. Breastfeeding mitigates short-term effects, with 15 months of breastfeeding sufficient to entirely offset the elevated hospitalization risk.&lt;/p&gt;
&lt;p&gt;Scope conditions: findings apply to second-born relative to first-born children in Danish sibling pairs with at least 11 months birth spacing; long-term estimates are net of parental compensatory responses and any immunity benefits, and thus represent lower bounds of the uncompensated biological impact of respiratory illness in infancy.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the birth order gap in acute respiratory hospitalizations during infancy, and what patterns support an intra-household transmission mechanism?
A: Younger siblings have 2–3 times higher hospitalization rates for acute respiratory conditions in the first year of life compared to older siblings at the same age, with the gap especially large at ages two and three months. The gap is larger for winter births (when respiratory viruses circulate more), for siblings with shorter birth spacing, and when the older sibling attends a childcare center. Hospitalizations for non-infectious digestive diseases and injuries show no analogous birth order differences, ruling out differential parental healthcare-seeking as an explanation.&lt;/p&gt;
&lt;p&gt;Q: How is the disease exposure index constructed and what variation does it exploit?
A: The index is the cumulative count of acute respiratory hospitalizations per 100 children aged 13–71 months in a child&amp;rsquo;s municipality over their first 12 months of life, with the older sibling excluded from the count when applicable. It exploits irregular spatial and temporal waves of respiratory viruses (such as RSV and influenza) across Danish municipalities. The interquartile range of this index captures meaningful variation in community disease burden faced by infants across different places and years.&lt;/p&gt;
&lt;p&gt;Q: What is the first-stage relationship between the disease index and infant hospitalizations?
A: Moving from the 25th to the 75th percentile of the disease index increases younger siblings&amp;rsquo; acute respiratory hospitalizations in the first year of life by 0.023 (a 32.9 percent increase relative to the sample mean), while the effect on older siblings is substantially smaller. The interaction coefficient in the preferred specification implies that one additional hospitalization per 100 community children aged 13–71 months raises the younger sibling&amp;rsquo;s hospitalization count by 0.012 more than the older sibling&amp;rsquo;s. Effects are more than twice as large for exposure in the first compared to the second six months of life.&lt;/p&gt;
&lt;p&gt;Q: What are the estimated long-term effects on adult earnings, and how do they compare to benchmarks in the literature?
A: An interquartile increase in first-year respiratory disease exposure reduces younger siblings&amp;rsquo; wage earnings at ages 25–32 by 0.8 percent and total income by 0.8 percent, with a 0.3 percentage point reduction in income percentile rank. These magnitudes are comparable to a 1 percent earnings reduction from a 10 percent birth weight reduction (Black et al., 2007), a 1 percent earnings reduction from a 9 percent increase in ambient air pollution (Isen et al., 2017b), and roughly two-thirds of the in utero Spanish Influenza effect (Almond, 2006).&lt;/p&gt;
&lt;p&gt;Q: Does the birth order earnings disadvantage reflect intra-household disease transmission?
A: When the interaction between birth order and the disease index is excluded, the regression finds a 1.9 percent birth order earnings disadvantage for second-born children (consistent with Black et al., 2005 range of 1.2–4.2 percent). When the interaction is included, the main birth order coefficient declines by approximately 70 percent, suggesting that disease transmission from older to younger siblings is an important channel driving the birth order earnings penalty.&lt;/p&gt;
&lt;p&gt;Q: Are effects larger for exposure in the first versus second six months of life?
A: Yes, consistently across all outcomes. The interaction coefficient for acute respiratory hospitalizations is more than twice as large when exposure is measured in the first versus second six months. Effects on wage earnings are approximately 60 percent larger for first-half exposure, and effects on income rank are two to three times larger. This is consistent with biomedical evidence that infants&amp;rsquo; immune systems mature around six months when solid food introduction begins.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on educational outcomes?
A: An interquartile increase in first-year respiratory disease exposure reduces the likelihood of high school graduation by 0.5 percentage points (0.6 percent at the sample mean) and college graduation by 0.6 percentage points (1.7 percent at the sample mean), with effects approximately 60 percent larger when measuring first-half exposure. A 0.01 standard deviation reduction in ninth grade Danish test scores is also found. A back-of-the-envelope calculation using Danish returns to schooling suggests the reduction in educational attainment can explain approximately half of the estimated earnings effect.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on chronic respiratory and mental health outcomes?
A: An interquartile increase in first-year exposure increases chronic respiratory hospitalizations (asthma, COPD) at ages 16–26 by 0.016 per hundred per year (20 percent above the sample mean), with significant increases also apparent at ages one to two. For mental health, the same exposure is associated with 0.5 additional psychiatric clinic visits per hundred per year at ages 16–26 (6.1 percent above the sample mean), with effects becoming more significant in the early twenties. Effects on mental health from this paper are smaller than those estimated for more extreme fetal and early childhood shocks such as Ramadan exposure or maternal bereavement.&lt;/p&gt;
&lt;p&gt;Q: What does the acute respiratory trajectory look like beyond infancy?
A: Elevated acute respiratory hospitalizations persist at age one, then there is a reduction at ages two to three consistent with an immunity formation hypothesis, but this protective effect disappears by age four. There is no significant increase or decrease in acute respiratory hospitalizations at older ages, in contrast to the persistent increase found for chronic respiratory conditions.&lt;/p&gt;
&lt;p&gt;Q: What heterogeneity is found in short-term effects?
A: Effects on infant respiratory hospitalizations are larger for low birth weight children, for male infants (consistent with the fragile male hypothesis), for siblings with shorter birth spacing, and for sibling pairs where the older child attends childcare. The monotonic decline in effect size with increasing birth spacing is the opposite of what would be predicted if differential parental time investment were the main mechanism, supporting intra-household disease spread as the operative channel.&lt;/p&gt;
&lt;p&gt;Q: What is the role of breastfeeding as a moderator?
A: Using supplementary data on breastfeeding duration (covering 2009–2016, matched to 7.6 percent of the sample), the authors find that the impact of disease exposure on younger siblings&amp;rsquo; infancy hospitalizations declines significantly with longer breastfeeding duration. A linear specification implies that 15 months of breastfeeding entirely offsets the elevated hospitalization risk from higher disease exposure. Second-born children breastfed for less than half a month are particularly vulnerable to acute respiratory infections.&lt;/p&gt;
&lt;p&gt;Q: How do the authors validate the identifying assumption?
A: Three validation exercises are used. First, results are robust to adding municipality-specific linear and quadratic trends and maternal fixed effects. Second, using family background characteristics as outcomes in the interaction regression, at most two of fourteen coefficients are significant in any specification, and all effect sizes are less than one percent of sample means. Third, using alternative disease indices based on non-infectious digestive diseases and injuries shows no differential effects for younger siblings, ruling out a parental healthcare-seeking confound.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications?
A: The authors highlight breastfeeding support policies (paid family leave, workplace lactation accommodations), RSV vaccination campaigns for pregnant women and monoclonal antibody prophylaxis for infants, sick pay regulations, and childcare attendance policies as levers to reduce infant respiratory disease burden. They argue that current cost-benefit evaluations of such policies likely undercount the long-term human capital and earnings benefits. The COVID-19 pandemic illustrates the mechanism: restrictions reduced RSV spread during 2020 potentially benefiting infants with older siblings, while the subsequent RSV surge in 2021–2022 may have exposed later cohorts to above-average disease burden.&lt;/p&gt;
&lt;p&gt;Respiratory Disease Exposure Index: A municipality-level cumulative measure of acute respiratory hospitalizations per 100 children aged 13–71 months assigned to each child over their first 12 months of life (or first and second six months separately), designed to proxy for community respiratory disease burden faced by infants from slightly older children, with the child&amp;rsquo;s own older sibling excluded from the count.&lt;/p&gt;
&lt;p&gt;Intra-Household Disease Transmission: The mechanism by which preschool-aged older siblings, exposed to respiratory viruses in group childcare settings, bring home those viruses and infect younger infant siblings who are in a vulnerable stage of immune and brain development, creating a within-family externality in health outcomes.&lt;/p&gt;
&lt;p&gt;Differential Birth Order Effect (Identification): The quasi-experimental design exploits the interaction between birth order (younger siblings are more exposed to older siblings&amp;rsquo; illnesses) and local disease prevalence variation to identify causal impacts, netting out the main effects of both birth order and local disease environment through municipality and birth year-month fixed effects.&lt;/p&gt;
&lt;p&gt;Immunity Formation Hypothesis: The conjecture that early respiratory disease exposure may have a protective effect on later acute respiratory illness through immune system training; supported in the data by reduced acute hospitalizations at ages two to three, though this protection disappears by age four and does not prevent chronic respiratory disease development.&lt;/p&gt;
&lt;p&gt;Dynamic Complementarities with Sibling Health Spillovers: An extension of the Cunha-Heckman framework: while standard models incorporate investment complementarities across time periods for a given child, this paper&amp;rsquo;s findings imply that sibling health spillovers create differential returns to early-life health investments by birth order, since disease asymmetries between older and younger siblings are not incorporated in existing theoretical models.&lt;/p&gt;
&lt;p&gt;Net Long-Term Effects: The estimated long-run impacts incorporate not only the direct biological effects of respiratory illness on the younger sibling but also any parental compensatory responses and immunity benefits; thus they represent lower bounds of the uncompensated biological impact, as parental compensation would attenuate the measured sibling difference.&lt;/p&gt;</description></item><item><title>Health Shocks, Health Insurance, Human Capital, and the Dynamics of Earnings and Health</title><link>https://macropaperwarehouse.com/papers/health-shocks-health-insurance-human-capital-and-the-dynamics-of-earnings-and-health/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/health-shocks-health-insurance-human-capital-and-the-dynamics-of-earnings-and-health/</guid><description>&lt;p&gt;Capatina and Keane build and calibrate a life-cycle model of labor supply and savings for U.S. men that incorporates health shocks, endogenous human capital accumulation via learning-by-doing, employer-sponsored health insurance (ESHI), means-tested social insurance, and endogenous medical treatment decisions. The model is calibrated to White males using the Medical Expenditure Panel Survey (MEPS) for 2000–2013, supplemented by CPS, HRS, and PSID data; separate calibrations are presented for Black and Hispanic men with high school or less education.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central research question is how health shocks affect labor supply, earnings, and earnings inequality over the life cycle, and through which mechanisms. Four channels are identified and quantified: (1) the direct labor supply effect — sick days and reduced tastes for work caused by health shocks; (2) the human capital effect — reduced work experience from health-shock-induced employment exits, which deteriorates future job and wage offers in a snowball dynamic; (3) the health-productivity effect — reduced functional health directly lowering wage offers; and (4) the behavioral effect — anticipation of health risk induces low-skill workers lacking ESHI to curtail labor supply to maintain means-tested transfer eligibility.&lt;/p&gt;
&lt;p&gt;The key quantitative findings from eliminating serious health shocks for working-age men (ages 25–64) are: the expected present value of lifetime earnings (PVE) for White men rises by 11% on average, and inequality in PVE falls by 12% (coefficient of variation). For White men with high school or less education the increase in PVE is 17.9%. For the typical White male the four channels contribute 5.7%, 2.7%, 1.4%, and 0.8% respectively. For low-skill White high school men the same channels contribute 10.7%, 14.8%, 1.3%, and 9.8% — with the human capital and behavioral effects dramatically larger for the low-skill group. For comparison, a severe health shock at age 40 reduces the present value of remaining lifetime earnings by 5.6% (approximately $53.9k) for a typical college man and by 11.5% (approximately $55.0k) for a typical high school man.&lt;/p&gt;
&lt;p&gt;Human capital amplification operates through employment persistence: a major health shock causes full-time employment to drop by 12 percentage points one year after the shock for the average man, and by 20 percentage points for high school men, with recovery still incomplete eight years later (employment remains 7.8 pp and 10 pp below baseline, respectively). Holding human capital fixed as in the pre-shock baseline causes employment to recover quickly, confirming that persistent wage-offer deterioration is the mechanism.&lt;/p&gt;
&lt;p&gt;On health insurance policy, the model evaluates providing public insurance to all workers lacking ESHI. This substantially increases medical utilization, improves health and life expectancy (survival to age 65 rises from 82% to 87% when health shocks are eliminated, as a related benchmark), reduces Medicaid and free-care costs, and raises labor supply among low-skill workers by weakening means-tested transfer incentives. The net program cost in a balanced budget simulation is modest, and all agent types are ex ante better off. By contrast, expanding Medicaid access creates perverse labor supply disincentives — workers reduce labor supply to maintain eligibility — does little to improve health, and makes almost all agents worse off in a balanced budget scenario.&lt;/p&gt;
&lt;p&gt;Scope conditions: the primary calibration covers non-institutionalized civilian White males; results for Blacks and Hispanics are presented only for the high school or less education group due to small samples. The model period ends at 2013, before ACA implementation.&lt;/p&gt;
&lt;p&gt;Q: What is the model&amp;rsquo;s overall estimate of how much health shocks reduce lifetime earnings for White men?
A: Eliminating serious health shocks at working ages (25–64) would increase the expected present value of lifetime earnings (PVE) for the average White male by 11% and reduce inequality in PVE by 12% as measured by the coefficient of variation. For White men with high school or less education the PVE gain is larger at 17.9%.&lt;/p&gt;
&lt;p&gt;Q: What are the four channels through which health shocks affect earnings, and how large is each for the average White male versus a low-skill high school male?
A: The four channels are (1) direct labor supply via sick days and reduced tastes for work, (2) human capital deterioration from lost work experience worsening future job/wage offers, (3) reduced health productivity lowering wage offers, and (4) behavioral responses to health risk reducing labor supply to preserve transfer eligibility. For the average White male the contributions to PVE are 5.7%, 2.7%, 1.4%, and 0.8%, respectively. For low-skill White high school men the same channels contribute 10.7%, 14.8%, 1.3%, and 9.8% — the human capital and behavioral effects are roughly five to twelve times larger for the low-skill group.&lt;/p&gt;
&lt;p&gt;Q: Why is the human capital effect so much larger for low-skill high school men than for college men?
A: Low-skill high school men are much more likely to exit full-time employment following a major health shock and are slow to return. Lifetime work years decline by 1.89 for the typical high school man versus only 0.84 for the typical college man following a major shock at age 40. Because job offer probabilities depend on lagged employment, absence from the labor market creates a snowball effect that persistently depresses offer quality; human capital accounts for 42% of the earnings decline for high school men versus 34% for college men.&lt;/p&gt;
&lt;p&gt;Q: How does the paper characterize the persistent employment effects of a major health shock?
A: For the average man, full-time employment drops by 12 percentage points one year after a severe shock and remains 7.8 pp below baseline after eight years. For high school men the initial drop is 20 pp, still 10 pp below baseline after eight years; for college men the figures are 7 pp and 3 pp. When human capital is held fixed at the pre-shock baseline — so wage and job offers do not deteriorate due to lost experience — employment recovers quickly for workers of all skill levels, confirming the human capital mechanism drives the persistence.&lt;/p&gt;
&lt;p&gt;Q: How does the behavioral effect operate for low-skill workers?
A: Workers without ESHI who face health risk have an incentive to maintain sufficiently low income and assets to qualify for means-tested social insurance, which provides a consumption floor approximating Medicaid, Food Stamps, SSDI, and SSI. This perverse incentive leads low-skill workers to curtail labor supply preemptively. When health risk is eliminated, this incentive disappears and labor supply rises, generating the behavioral effect of 9.8% of PVE for low-skill high school men versus only 0.8% for the average White male.&lt;/p&gt;
&lt;p&gt;Q: How does the paper correct for under-reporting of health shocks among the uninsured?
A: The measurement model assumes health shocks are correctly measured for the treated, but uninsured workers who do not seek treatment only record a shock with a shock-specific probability less than one. A key identifying assumption is that, conditional on health status, risk factors, age, and education, the true frequency of health shocks does not differ by insurance status per se — ruling out ex ante moral hazard. The measurement model parameters are calibrated to match observed frequencies of health shocks and high risk in MEPS for the uninsured.&lt;/p&gt;
&lt;p&gt;Q: What does the model estimate regarding the effect of a severe health shock on cumulative earnings relative to existing reduced-form evidence?
A: The model predicts an average cumulative (non-discounted) earnings loss of $42.8k over ten years following a severe shock for men aged 50, compared with Smith&amp;rsquo;s (2004) estimate of $37k from the HRS. The paper argues Smith&amp;rsquo;s estimate identifies effects on workers who actually experience shocks, who are a selected sample with low baseline earnings (as untreated shocks are more likely to be severe, and non-treaters tend to have low earnings). The model&amp;rsquo;s &amp;ldquo;average effect&amp;rdquo; — comparing a world where everyone experiences the shock to one where no one does — yields a substantially higher loss of $59.8k.&lt;/p&gt;
&lt;p&gt;Q: What are the key findings from the public insurance experiment (providing insurance to the uninsured)?
A: Providing public insurance to all workers lacking ESHI substantially increases medical utilization among the previously uninsured, who are intrinsically less healthy. This improves health and life expectancy, raising Social Security costs. However, it also generates positive labor supply incentives for low-skill workers (reducing their reliance on means-tested transfers), substantially reduces Medicaid and free-care costs, and increases tax revenue. On balance, the net program cost in a balanced budget simulation is modest, and all types of workers are ex ante better off.&lt;/p&gt;
&lt;p&gt;Q: Why does expanding Medicaid access produce perverse results in contrast to providing public insurance?
A: Medicaid is means-tested, so expanded access requires workers to maintain sufficiently low income and assets to remain eligible. This creates disincentives to work and save — workers reduce labor supply to preserve eligibility. The result is reduced earnings, lower tax revenue, little improvement in health (as access to care depends on maintaining low income), and almost all agents being worse off in a balanced budget scenario.&lt;/p&gt;
&lt;p&gt;Q: What role does insurance play beyond consumption smoothing in this model?
A: Beyond lowering out-of-pocket (OOP) costs and smoothing consumption, insurance grants access to care: in the US system, proof of insurance is often required before treatment, so uninsured workers may not have the option to treat at all. The model captures three distinct option sets for the uninsured — all options available, treatment not available, or default not available — each motivated by different real-world contexts. Non-treatment worsens health transition probabilities, so the access-granting role of insurance independently affects health trajectories beyond its cost-reducing role.&lt;/p&gt;
&lt;p&gt;Q: What explains the observed positive association between education, income, insurance, and health transitions in the data, and how does the model generate this without education entering the health production function directly?
A: The association between education and health is largely driven by the positive correlation between education and latent health types; controlling for latent health type in a descriptive logit largely eliminates the education coefficient. The association between insurance and health transitions is driven by the fact that the insured are more likely to receive treatment; controlling for treatment and true shocks eliminates the insurance coefficient. Education affects health indirectly through its effects on treatment decisions — via wages, job offers with ESHI, and consumption capacity — without appearing as a direct argument in the health production function.&lt;/p&gt;
&lt;p&gt;Q: How large are the effects of health shocks on key population health statistics according to the model?
A: Eliminating serious health shocks at working ages would increase the fraction of working-age men in good health from 60% to 75% and raise the probability of survival to age 65 from 82% to 87%. Average annual sick days of 16.42 would be eliminated, implying a 6% increase in work days for employed workers and an employment rate increase from 88% to 91%. Average annual medical costs would fall from $4,618 to $1,132.&lt;/p&gt;
&lt;p&gt;Q: How do the results for Black and Hispanic men compare to White men?
A: The results are qualitatively similar, but the magnitudes for Black men are somewhat larger. Eliminating health shocks would raise PVE for Whites, Blacks, and Hispanics with high school or less education by 17.9%, 23.7%, and 17.7%, respectively. Separate access-to-care probabilities are calibrated for each group, reflecting racial disparities in access that explain part of the observed differences in health outcomes and treatment rates.&lt;/p&gt;
&lt;p&gt;Q: What is the role of the consumption floor (means-tested social insurance) in shaping equilibrium outcomes for low-skill workers?
A: The consumption floor guarantees a minimum household consumption level approximating Medicaid, Food Stamps, SSDI, and SSI. It shields low-skill workers from the full cost of health shocks, reducing both the consumption-smoothing value of ESHI and precautionary saving incentives. However, it also creates a powerful disincentive for low-skill workers without ESHI to work, as earning above the eligibility threshold would eliminate benefits. This mechanism amplifies earnings inequality by generating perverse labor supply behavior concentrated among low-skill, uninsured workers.&lt;/p&gt;
&lt;p&gt;Functional Health (H): A discrete stock variable (Poor, Fair, or Good) measuring aspects of health that directly affect worker productivity and tastes for work; distinguished from asymptomatic health risk. Transitions depend on lagged health, latent health type, age, persistent health shocks, and whether shocks are treated.&lt;/p&gt;
&lt;p&gt;Asymptomatic Health Risk (R): A binary state (low or high) capturing risk factors such as obesity, high cholesterol, and hypertension that increase the probability of future health shocks but do not affect current productivity.&lt;/p&gt;
&lt;p&gt;Human Capital Effect: The channel by which health shocks reduce lifetime earnings not directly but indirectly — by causing employment exits that slow work experience accumulation, which in turn deteriorates future job offer probabilities and wage offers in a persistent, self-reinforcing (snowball) dynamic.&lt;/p&gt;
&lt;p&gt;Behavioral Effect: The reduction in labor supply — and associated earnings loss — that occurs because workers facing health risk and lacking ESHI have an incentive to keep income and assets low enough to maintain eligibility for means-tested social insurance, even absent any contemporaneous health shock.&lt;/p&gt;
&lt;p&gt;Tied Wage-Hours-Insurance Offer: The model&amp;rsquo;s labor market structure in which employment offers jointly specify a wage rate, hours (no offer, part-time, or full-time), and whether the offer includes ESHI; workers accept or reject the bundle rather than choosing hours and insurance independently.&lt;/p&gt;
&lt;p&gt;Source Text Origin: The paper&amp;rsquo;s own term distinguishing how the full text of a paper was obtained (PDF, OA-HTML, or abstract-only); used in the summarization pipeline. [Note: this concept is from the summarization pipeline metadata, not from the paper itself — omitting.]&lt;/p&gt;
&lt;p&gt;Treatment/Payment Options: The set of decisions available to a worker after a health shock occurs — whether to seek treatment and, if treated, whether to pay the out-of-pocket cost or default on bills. The available choice set differs by insurance status and context: the uninsured may face denial of access (option to treat unavailable) or required prepayment (default unavailable), or may have all options including free care.&lt;/p&gt;
&lt;p&gt;Latent Health Type: An unobserved permanent individual characteristic capturing innate biological resilience and pre-age-25 health investments; determines baseline transition probabilities for functional health conditional on shocks. Positively correlated with latent skill type within education groups.&lt;/p&gt;</description></item><item><title>Homeownership, Polarization, and Inequality</title><link>https://macropaperwarehouse.com/papers/homeownership-polarization-and-inequality/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/homeownership-polarization-and-inequality/</guid><description>&lt;p&gt;This paper asks why job polarization and income inequality are higher in large U.S. cities, and proposes a novel housing-market mechanism that operates independently of — but interacts with — the skill-biased technical change (SBTC) explanations dominant in the existing literature.&lt;/p&gt;
&lt;p&gt;The core argument is that large cities have experienced faster growth in house prices relative to both wages (price-wage ratio) and rents (price-rent ratio) since 1980. This excess price growth has priced middle-income households out of homeownership in expensive cities. Because low-income households cannot afford to own anywhere and high-income households can afford to own everywhere, it is specifically middle-income (middle-skilled) households whose location choice becomes entangled with their tenure choice. These households increasingly sort toward smaller, more affordable cities where they can purchase a home. This selective out-migration hollows out the middle of the income distribution in large cities, producing greater employment polarization and income inequality there.&lt;/p&gt;
&lt;p&gt;Empirically, the paper uses Census and ACS data from 1980 to 2019 covering 465 commuting zones (CZs). Polarization is measured following Autor and Dorn (2013) by assigning 3-digit occupations to income percentiles fixed at 1980 levels; inequality is measured by the Gini coefficient and variance of log annual wages. Housing costs are captured by hedonic price and rent indices and three derived ratios. OLS and IV results (instrumented using the interaction of land unavailability and long-run changes in real interest rates) show that doubling of prices is associated with a 1 percentage point decline in the middle-skilled employment share; doubling of the price-rent ratio is associated with an 11.3 percentage point decline; doubling of the price-wage ratio with a 5.3 percentage point decline. Inequality follows the same pattern: doubling prices raises 100x the variance of log wages by 2.3 points; doubling the price-rent ratio raises it by 11.7 points; doubling the price-wage ratio by 7.7 points.&lt;/p&gt;
&lt;p&gt;The migration mechanism is documented using 2001–2019 CPS ASEC data, which — uniquely among available sources — reports reasons for moving. A doubling of the price index, price-wage ratio, or price-rent ratio in the origin state relative to the destination raises the probability that a middle-income (2nd–4th quintile) household moves for housing-related reasons by approximately 5–10 percentage points in absolute terms, implying a 50–80% relative increase compared with low- or high-income households making a housing-related move.&lt;/p&gt;
&lt;p&gt;The theoretical framework extends the standard spatial equilibrium (Rosen-Roback) model with two additions: skill heterogeneity and housing tenure choice. Households face a minimum house size constraint and a payment-to-income (PTI) constraint (calibrated at lambda = 0.308). These constraints create distinct skill thresholds for homeownership that vary by city; the interaction between location and tenure choices applies only to middle-skilled households who can afford ownership in cheap but not expensive cities.&lt;/p&gt;
&lt;p&gt;In the quantitative model, calibrated separately for 1980 and 2019 with two locations (top 30 CZs vs. the rest), counterfactual experiments show that holding price-wage ratios at their 1980 levels reduces the excess polarization gap between large and small CZs by 93% and the excess inequality gap by 40%. Holding price-rent ratios constant reduces the polarization gap by 96% and the inequality gap by 27%. By contrast, shutting down SBTC entirely reduces the polarization gap by only 54% and the inequality gap by 73%. These results establish that while SBTC is an important driver, its effect on polarization and inequality is substantially amplified by faster house price growth in large cities; without the housing affordability channel, the effect of SBTC on disproportionate polarization would be 63–81% smaller and on the inequality gap 18–36% smaller.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s central research question?
A: The paper asks why job polarization and income inequality are systematically higher in large U.S. cities than in small ones. Prior literature attributed this to skill-biased technical change, external labor demand shocks, or IT-driven displacement of routine jobs; this paper proposes a complementary, housing-market-based explanation that does not rely on features of the production technology.&lt;/p&gt;
&lt;p&gt;Q: What is the core mechanism linking house prices to polarization?
A: When price-wage and price-rent ratios are higher in large cities, middle-income households face binding minimum-size and payment-to-income constraints that prevent them from owning a home there but not in cheaper cities. Because homeownership carries financial advantages, these households sort toward smaller, more affordable cities. Low-income households cannot afford ownership anywhere and high-income households can afford it anywhere, so only the middle group&amp;rsquo;s location choice is distorted by tenure considerations. This selective out-migration hollows out the middle of the income distribution in expensive large cities.&lt;/p&gt;
&lt;p&gt;Q: What empirical patterns in CZ-level data motivate the paper?
A: Doubling CZ size is associated with a 1.9 percentage point greater fall in the middle-skilled employment share and a 2.7 point higher growth in 100x the variance of log wages from 1980 to 2019. Larger CZs also experienced 3.4% higher price growth, 3.1% higher price-wage ratio growth, and a 10% greater increase in price-rent ratios. These associations persist after controlling for initial CZ size and other characteristics.&lt;/p&gt;
&lt;p&gt;Q: What do the OLS and IV results show about house prices and polarization?
A: A doubling of house prices is associated with a 1 percentage point decline in the middle-skilled share; a doubling of the price-rent ratio with an 11.3 percentage point decline; and a doubling of the price-wage ratio with a 5.3 percentage point decline. IV results using the interaction of land unavailability and the change in real interest rates as an instrument confirm the negative relationship remains statistically significant, suggesting a causal interpretation is plausible.&lt;/p&gt;
&lt;p&gt;Q: What do the OLS and IV results show about house prices and income inequality?
A: A doubling of prices is associated with a 2.3 point increase in 100x the variance of log wages; a doubling of the price-rent ratio with an 11.7 point increase; and a doubling of the price-wage ratio with a 7.7 point increase. IV results suggest a causal relationship between price growth and income inequality at the CZ level.&lt;/p&gt;
&lt;p&gt;Q: What evidence does the paper provide for the migration mechanism?
A: Using 2001–2019 CPS ASEC data (which reports stated reasons for moving, unlike the ACS), the paper estimates logit regressions of interstate migration for housing-related reasons. A doubling of the price index in the origin state relative to the destination raises the probability of a housing-related move for middle-income (2nd–4th quintile) households by 5–6 percentage points; a doubling of the price-wage ratio raises it by 6–7 percentage points; and a doubling of the price-rent ratio raises it by 7–10 percentage points. These effects imply a 50–80% relative increase in housing-related migration probability for the middle quintiles compared with the bottom or top quintile. Housing-related movers constitute over 12% of all interstate migrants in the sample.&lt;/p&gt;
&lt;p&gt;Q: What is the key finding about homeownership rates?
A: There is no statistically significant relationship between the change in homeownership rates and the growth in prices, price-rent, or price-wage ratios from 1980 to 2019. This is consistent with the model&amp;rsquo;s mechanism, in which middle-income households who cannot afford ownership in large cities move away rather than simply switching to renting there — so aggregate local ownership rates need not fall.&lt;/p&gt;
&lt;p&gt;Q: How does the theoretical model generate the polarization result?
A: The model extends the Rosen-Roback spatial equilibrium framework with skill heterogeneity and housing tenure choice. Two skill thresholds — one for minimum-size-constrained ownership and one for unconstrained ownership — interact with the price-wage and price-rent ratios of each city. Proposition 1 proves that a city with higher price-wage and price-rent ratios will have a lower middle-skilled share, because middle-skilled workers (those who can afford to own in cheap but not expensive cities) are drawn to cheaper locations. Proposition 2 shows that in a world with only renters or only owners, skill shares would be identical across cities regardless of price differences — the polarization result requires heterogeneity in tenure choice.&lt;/p&gt;
&lt;p&gt;Q: What does the no-SBTC counterfactual show?
A: Holding the parameters governing local returns to skills at their 1980 levels (shutting down skill-biased technical change) reduces the difference in the decline in the middle-skilled share between large and small CZs by 54% and the gap in the increase in the variance of log wages by 73%. This is broadly consistent with prior literature attributing the bulk of disproportionate polarization and inequality in big cities to SBTC.&lt;/p&gt;
&lt;p&gt;Q: What do the constant price-ratio counterfactuals show?
A: When price-wage ratios are held at 1980 levels (but SBTC is allowed to operate), the excess polarization gap between large and small CZs falls by 93% and the excess inequality gap by 40%. When price-rent ratios are held at 1980 levels, the polarization gap falls by 96% and the inequality gap by 27%. When both are held constant simultaneously, the polarization gap falls by 89% and the inequality gap by 27%. These results show that the effect of SBTC on polarization would be 63–81% smaller in the absence of the housing affordability amplification channel.&lt;/p&gt;
&lt;p&gt;Q: Who are the largest losers from rising price-wage ratios in large cities?
A: The counterfactual welfare analysis identifies middle-skilled workers with skill levels between approximately 0.29 and 0.80 as the primary losers. In the counterfactual with fixed price-wage ratios, workers with skills from 0.29 to 0.57 who previously could not afford ownership in large cities are now able to own there, and those with skills from 0.57 to 0.80 spend a smaller share of income on housing. This group either lost homeownership opportunities or was induced to move to less productive CZs by the actual price growth that occurred.&lt;/p&gt;
&lt;p&gt;Q: How is the quantitative model calibrated and structured?
A: The model is calibrated separately for 1980 and 2019 as two stationary spatial equilibria. It features two locations (the top 30 CZs, which account for 49.3% of employment, and the remaining CZs). Key parameters include a Frechet elasticity of 6.1, an agglomeration externality of 0.04, a PTI constraint of 0.308, and an annual discount factor of 0.96. Land shares differ between large and small CZs (0.3965 vs. 0.2239). The model finds that the price-rent ratio was relatively stable in large cities but fell in small ones, while the price-wage ratio increased much more in large CZs — both indicators point to purchasing a home becoming relatively more expensive in large CZs.&lt;/p&gt;
&lt;p&gt;Q: What are the paper&amp;rsquo;s policy implications?
A: Zoning reforms and other policies that increase housing supply in large, unaffordable cities could produce a more efficient spatial allocation of labor, greater aggregate productivity, and more economically diverse — less polarized and less unequal — cities, while also reducing the wealth gap between owners and renters. Policies that promote homeownership by reducing the cost of owning without raising housing supply may reduce local polarization and inequality but could lower aggregate output and do not necessarily increase homeownership rates.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to existing explanations for city-level polarization?
A: The paper&amp;rsquo;s housing-market mechanism is explicitly complementary to SBTC-based explanations (Baum-Snow, Freedman, and Pavan, 2018; Cerina et al., 2023), external demand shock explanations (Davis, Mengus, and Michalski, 2020), and IT-displacement explanations (Eeckhout, Hedtrich, and Pinheiro, 2024). The paper&amp;rsquo;s key added contribution is that even if SBTC were the primary driver of disproportionate polarization, its measured effect would be substantially smaller in the absence of faster house price growth in large cities — the housing market amplifies rather than replaces the technology channel.&lt;/p&gt;
&lt;p&gt;Job polarization (city-level): The hollowing out of middle-income employment shares in a commuting zone, measured as the change in the share of workers in occupations assigned to the 21st–80th income percentile (using the 1980 occupation-to-percentile mapping fixed over time). In this paper, polarization is greater in cities where price-wage and price-rent ratios grew faster, attributed to selective out-migration of middle-skilled households.&lt;/p&gt;
&lt;p&gt;Price-wage ratio: The ratio of hedonic house prices to median annual wages in a commuting zone, constructed from Census and ACS data. A higher price-wage ratio tightens the payment-to-income constraint on potential homebuyers and is the primary driver of the skill threshold for homeownership in the model.&lt;/p&gt;
&lt;p&gt;Price-rent ratio: The ratio of hedonic house prices to rents in a commuting zone. In the model, a higher price-rent ratio reduces the financial advantage of owning over renting, raising the skill threshold at which ownership becomes optimal. The paper treats price-rent and price-wage ratios as distinct channels that both independently amplify polarization.&lt;/p&gt;
&lt;p&gt;Housing tenure choice: The household decision to own or rent, modeled as a discrete choice made at the start of life that interacts with location choice. Ownership requires satisfying both a minimum house size constraint and a payment-to-income (PTI) constraint (lambda = 0.308). The interaction between tenure and location choices is the paper&amp;rsquo;s key model innovation; it exists only for middle-skilled workers whose income is sufficient for ownership in cheap but not expensive cities.&lt;/p&gt;
&lt;p&gt;Skill threshold for homeownership (s*_i): The minimum skill level at which a worker in city i chooses to own rather than rent, defined by Lemma 2. This threshold is decreasing in local labor productivity and increasing in price-wage and price-rent ratios. Workers with skill below s*_i in all cities always rent; those with skill above s*_i in all cities always own; those in between face city-dependent tenure choice that distorts their location decision.&lt;/p&gt;
&lt;p&gt;Skill-biased technical change (SBTC): In the paper&amp;rsquo;s quantitative model, SBTC is represented by faster growth in the skill dispersion parameter (alpha_it) in large CZs, reflecting differential productivity growth concentrated at the top of the skill distribution. The paper finds SBTC accounts for 54% of the polarization gap and 73% of the inequality gap in its counterfactual, but argues its effect is amplified 4–5x by the housing affordability channel.&lt;/p&gt;
&lt;p&gt;Payment-to-income (PTI) constraint: The constraint that a homebuyer cannot spend more than a fraction lambda (calibrated at 0.308) of annual labor earnings on the annual housing payment (user cost times price times quantity). This constraint, together with the minimum house size, determines the income threshold for ownership and makes location and tenure choices interdependent for middle-skilled workers.&lt;/p&gt;</description></item><item><title>How Do You Identify a Good Manager?</title><link>https://macropaperwarehouse.com/papers/how-do-you-identify-a-good-manager/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/how-do-you-identify-a-good-manager/</guid><description>&lt;p&gt;This paper develops a novel experimental method to identify the causal contribution of managers to team performance, and uses it to evaluate which characteristics predict managerial effectiveness and how manager selection mechanisms affect organizational outcomes.&lt;/p&gt;
&lt;p&gt;The core identification challenge is that managers are not randomly assigned to teams in the field, and field managers are a highly non-random sample, making it difficult to infer which traits genuinely predict managerial performance. The authors address this by repeatedly randomly assigning managers to multiple teams in a controlled laboratory experiment, then estimating each manager&amp;rsquo;s average causal contribution to group output after conditioning on group members&amp;rsquo; individual productive skills. The intuition is that a good manager is someone who consistently causes their team to produce more than the sum of their parts.&lt;/p&gt;
&lt;p&gt;The experiment was conducted at the University of Essex lab with 555 participants (46% female, mean age 25, ethnically diverse) forming 728 groups of three across four rounds. Each group consisted of one manager and two workers who performed a Collaborative Production Task requiring coordination across three problem-solving modules (numerical, spatial, and analytical reasoning). The team score was the minimum module score — a weakest-link structure making coordination essential. Prior to group testing, all participants completed individual assessments of task-specific skill, fluid intelligence (CFIT), emotional perceptiveness (Reading the Mind in the Eyes Test, RMET), economic decision-making skill (the Assignment Game, which measures resource allocation under comparative advantage), Big 5 personality, and demographic characteristics. Manager selection was randomly varied at the session level: in 20 sessions, the participant with the strongest preference for leadership became manager (self-promotion); in 19 sessions, managers were assigned by lottery.&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. First, there are large, stable, and statistically significant manager effects: a manager one standard deviation above average improves team performance by approximately 0.23 standard deviations (p = 0.04). This estimate is roughly 90% the size of the combined productive skill coefficient for the two workers (approximately 0.26 sd), indicating that a good manager is roughly twice as valuable as a good individual worker. Manager contributions predict out-of-sample group performance in a leave-one-out procedure (p &amp;lt; 0.01).&lt;/p&gt;
&lt;p&gt;Second, among randomly assigned managers, only two predictors significantly explain managerial performance: fluid intelligence (CFIT) and economic decision-making skill (Assignment Game scores), both significant at below the 1% level. Gender, age, and ethnicity do not predict managerial performance.&lt;/p&gt;
&lt;p&gt;Third, self-promoted managers perform substantially worse than lottery-assigned managers, by approximately 0.10 standard deviations — roughly equivalent to being assigned a manager with fluid intelligence one full standard deviation below average. The mechanism is overconfidence: people who strongly prefer management roles are significantly more overconfident (d = 0.41 sd, p &amp;lt; 0.01) and exhibit a strong negative correlation between self-reported social skills and actual emotional perceptiveness on the RMET (r = -0.37, p &amp;lt; 0.001). Among self-promoted managers, self-reported extraversion and political skill are negatively correlated with managerial performance (rho = -0.24 and -0.26, p &amp;lt; 0.05); no such negative relationship appears among lottery managers.&lt;/p&gt;
&lt;p&gt;Fourth, selecting managers on economic decision-making skill rather than self-promotion improves average manager quality by 0.6 standard deviations — equivalent to replacing an average worker in every group with a worker at the 99th percentile of individual productivity.&lt;/p&gt;
&lt;p&gt;The three mechanisms through which good managers improve performance are: (1) monitoring — good managers (1 sd above average) cut monitoring errors from 16% to 8%; (2) optimal task allocation according to comparative advantage — groups with optimally assigned workers score 0.52 sd higher (p &amp;lt; 0.01); (3) worker motivation in late-stage effort — teams led by a 1-sd-above-average manager solve 0.6 more problems in the final two minutes versus only 0.3 more in the first two minutes.&lt;/p&gt;
&lt;p&gt;The experiment was conducted in a university lab in the UK, and the sample skews toward graduate students with limited work experience. Generalizability to field settings is supported by prior evidence that peer productivity spillover experiments yield similar magnitudes in lab versus field settings, and that the estimated manager effects are similar to Lazear et al. (2015) estimates from a large employer dataset.&lt;/p&gt;
&lt;p&gt;Q: What is the core methodological innovation of this paper?
A: The paper requires repeated random assignment of managers to multiple teams, combined with controls for individual productive skill measured prior to group work. This allows identification of each manager&amp;rsquo;s average causal contribution to group output, rather than confounding management quality with team composition or individual worker ability. The key estimand is the standard deviation of individual manager effects (sigma_alpha), interpreted as the impact of having a manager one standard deviation above average.&lt;/p&gt;
&lt;p&gt;Q: How large is the estimated manager effect, and how does it compare to worker effects?
A: A manager one standard deviation above average improves team performance by approximately 0.23 standard deviations (p = 0.04 by randomization inference). This is roughly 90% the size of the combined productive skill effect of both workers together (approximately 0.26 sd), implying a good manager is nearly twice as valuable as a good individual worker. Without conditioning on production skills, the manager effect rises to 0.29 sd.&lt;/p&gt;
&lt;p&gt;Q: What characteristics predict managerial performance among randomly assigned managers?
A: Only two measures predict managerial performance in the lottery arm: fluid intelligence (CFIT) and economic decision-making skill (scores on the Assignment Game), both significant at below the 1% level. These predictors are robust to controls for demographics, education, work experience, emotional perceptiveness, and personality traits. Gender, age, and ethnicity do not predict managerial performance.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;Assignment Game&amp;rdquo; and why is it a strong predictor?
A: The Assignment Game (Caplin et al., 2024) places participants in a simulated managerial role where they must assign fictional workers to tasks. Performing well requires understanding comparative advantage intuitively, managing an attentionally demanding numerical environment, and avoiding biases such as anchoring. The paper argues its strong predictive power reflects that good managers excel at allocating workers according to comparative advantage — which the experiment directly identifies as a key mechanism.&lt;/p&gt;
&lt;p&gt;Q: How do self-promoted managers perform relative to lottery-assigned managers?
A: Self-promoted managers perform approximately 0.10 standard deviations below lottery managers, and this gap is robust across model specifications. The performance deficit is roughly equivalent to being assigned a manager whose fluid intelligence is one full standard deviation below average. This finding implies that common organizational practice of selecting managers partly via self-nomination actively reduces team productivity.&lt;/p&gt;
&lt;p&gt;Q: Why do self-promoted managers underperform?
A: The paper attributes underperformance primarily to overconfidence. People strongly preferring management roles are significantly more overconfident than those without strong preferences (d = 0.41 sd, p &amp;lt; 0.01). Self-promoted managers specifically overestimate their social skills: among them, self-reported people skills are strongly negatively correlated with actual emotional perceptiveness on the RMET (r = -0.37, p &amp;lt; 0.001), and self-reported extraversion and political skill are negatively correlated with managerial performance (rho = -0.24 and -0.26, p &amp;lt; 0.05). None of these negative relationships appear among lottery managers.&lt;/p&gt;
&lt;p&gt;Q: Who wants to be a manager, and does it differ by gender?
A: The three variables most strongly correlated with wanting to be in charge are extraversion, risk appetite, and being male. The relationship between high extraversion and preference for management is driven largely by men. Women are much less likely to nominate themselves for leadership roles despite being equally or more effective on average — a finding consistent with broader experimental evidence on gender and leadership self-selection.&lt;/p&gt;
&lt;p&gt;Q: How large are the potential gains from skill-based manager selection?
A: Compared to self-promotion, selecting managers based on economic decision-making skill yields managers who are 0.6 standard deviations better in terms of estimated manager effects. In terms of group performance, this is equivalent to replacing an average worker in every group with a worker at the 99th percentile of individual productivity. Selecting on both economic decision-making and fluid intelligence outperforms random assignment, selection on social skills, or selection on worker task performance (the Peter Principle).&lt;/p&gt;
&lt;p&gt;Q: What are the three mechanisms through which good managers improve team performance?
A: First, monitoring: good managers (1 sd above average) reduce monitoring errors — defined as having a worker on a module substantially above the minimum score at task end — from 16% to 8% (bivariate correlation with manager performance = -0.40, p &amp;lt; 0.001). Second, optimal task allocation: the probability of finding the optimal comparative-advantage-based assignment is positively associated with manager performance (rho = 0.19, p &amp;lt; 0.01), and groups with always-optimal starting assignments score 0.52 sd higher than those with never-optimal assignments (p &amp;lt; 0.01). Third, worker motivation: team performance in the final two-minute period is about 50% more influential for overall outcomes than the first two minutes (p = 0.038), and 1-sd-above-average managers generate 0.6 more problems solved in the final period versus 0.3 in the first, consistent with differential motivational effects emerging over time.&lt;/p&gt;
&lt;p&gt;Q: What is the Peter Principle, and how does this paper relate to it?
A: The Peter Principle refers to the practice of promoting employees based on their performance as line workers rather than their suitability for management — promoting individuals to their level of incompetence. Benson et al. (2019) document this selection pattern empirically. This paper shows that selecting managers on worker task skill is inferior to selecting on economic decision-making skill or fluid intelligence, confirming that task skill is not the right criterion for manager selection even if it predicts individual worker output.&lt;/p&gt;
&lt;p&gt;Q: How does the paper validate that manager effects are real and not noise?
A: The paper uses randomization inference with 5,000 simulated allocations to compute p-values, obtaining p = 0.04 for the main manager effect. Robustness checks include controlling for pre-existing social relationships, manager risk appetite, variance of individual scores, and granular skill measures — all yielding estimates near 0.22 sd. A leave-one-out out-of-sample prediction test confirms manager contributions significantly predict held-out group performance (p &amp;lt; 0.01), while the analogous worker out-of-sample estimate is less than half the magnitude and not statistically significant.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions on the experimental results?
A: The experiment is conducted in a university lab in the UK with graduate students averaging 25 years of age and two years of work experience, limiting direct generalizability to experienced workers or senior management. The task lasts approximately 15 minutes, which may not capture longer-run managerial dynamics. Compensation equalized average earnings between managers and workers, which differs from most real-world settings. The authors note their effect-size estimates closely match Lazear et al. (2015) from a large employer, and that Herbst and Mas (2015) find lab peer-productivity experiments generalize to the field.&lt;/p&gt;
&lt;p&gt;Manager Effect (sigma_alpha): The standard deviation of individual managers&amp;rsquo; average causal contributions to group performance, estimated via repeated random assignment and conditioning on individual productive skill. Represents the impact of having a manager one standard deviation above average, estimated at approximately 0.23 standard deviations of group output.&lt;/p&gt;
&lt;p&gt;Collaborative Production Task: A novel lab group task in which a manager and two workers solve problems across three modules (numerical, spatial, analytical reasoning), with team score defined as the minimum module score (weakest-link structure). Managers are responsible for worker assignment, monitoring, and motivation; workers face no financial performance incentives.&lt;/p&gt;
&lt;p&gt;Economic Decision-Making Skill: Defined by Caplin et al. (2024) as the ability to make good resource allocation decisions, assessed via the Assignment Game in which participants must optimally assign workers to tasks under comparative advantage. The single strongest predictor of managerial performance in the lottery arm.&lt;/p&gt;
&lt;p&gt;Monitoring Failure: Defined in the paper as having any group member working on a module at task end whose score is substantially greater (e.g., 10 points higher) than the minimum module score — meaning the worker&amp;rsquo;s effort is not contributing to the group score. Occurs in 16% of groups overall; managers one sd above average reduce this to 8%.&lt;/p&gt;
&lt;p&gt;Self-Promotion (as selection mechanism): A treatment condition in which the participant with the strongest stated preference for being manager (on a 1-10 scale) is assigned the managerial role. Contrasted with lottery assignment; self-promoted managers perform approximately 0.10 sd worse than lottery managers.&lt;/p&gt;
&lt;p&gt;Overconfidence (in managerial context): The gap between self-assessed skill (particularly social/interpersonal skill) and objectively measured skill (e.g., RMET score). Self-promoters are significantly more overconfident (d = 0.41 sd), and overconfidence is strongly negatively correlated with actual emotional perceptiveness (r = -0.33, p &amp;lt; 0.001).&lt;/p&gt;
&lt;p&gt;Comparative Advantage Allocation: The practice of assigning each worker to the module in which they have the highest relative (not absolute) performance advantage. Captured via whether a manager selects the optimal one-to-one assignment given pre-measured individual module scores; groups with always-optimal allocation score 0.52 sd higher.&lt;/p&gt;</description></item><item><title>Ideas Have Consequences: The Impact of Law and Economics on American Justice</title><link>https://macropaperwarehouse.com/papers/ideas-have-consequences-the-impact-of-law-and-economics-on-american-justice/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/ideas-have-consequences-the-impact-of-law-and-economics-on-american-justice/</guid><description>&lt;p&gt;This paper quantifies the effect of the Manne Economics Institute for Federal Judges — an intensive two-week economics training program run by the Law and Economics Center from 1976 to 1998 — on the decision-making of U.S. federal judges. The research question is whether exposure to a coherent set of economic ideas can directly shift the policy decisions of sitting policymakers, as distinct from effects operating through partisan affiliation or formal legal rules.&lt;/p&gt;
&lt;p&gt;The program trained nearly half of all federal judges over its two decades of operation. By 1990, forty percent of federal judges had attended; by the late 1990s, roughly half of circuit court cases had a Manne-trained judge on the panel. Instructors included Milton Friedman, Armen Alchian, Harold Demsetz, Martin Feldstein, Paul Samuelson, and Orley Ashenfelter, covering supply-and-demand theory, the Coase Theorem, externalities, property rights, and criminal deterrence following Becker (1968). The program was funded by pro-business foundations and had a recognized conservative-leaning orientation, though it invited both Republican- and Democrat-appointed judges and was popular across party lines.&lt;/p&gt;
&lt;p&gt;The identification strategy is a differences-in-differences design exploiting staggered attendance timing. Because the program was oversubscribed and admitted judges on a first-come-first-served basis — with applicants bumped to later cohorts when capacity was reached — the timing of attendance within the ever-attending population has a quasi-random component. The preferred control group consists exclusively of other ever-attending judges who had not yet attended, rather than never-attenders, because never-attenders differ systematically on observables and show a pre-existing positive trend in economics language use, likely from ambient diffusion through clerks, law schools, and organizations such as the Federalist Society. Judge fixed effects and circuit-by-year (or courthouse-by-year) fixed effects absorb time-invariant judge characteristics and court-level time trends. Elastic-net-selected covariates predicting attendance timing, fully interacted with year fixed effects, are added as robustness controls. Standard errors are clustered by judge.&lt;/p&gt;
&lt;p&gt;The data cover approximately 200,000 published circuit court opinions (1970–2005) from Bloomberg Law, a 5% random sample of circuit cases hand-coded for ideological direction from the Songer-Auburn database, machine-coded regulatory agency outcomes, a newly collected antitrust case dataset, and approximately 1.03 million district court criminal sentencing records (1992–2003) from TRAC.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. First, after attending the Manne program, judges increase their use of economics language in written opinions by approximately one-third of a standard deviation, measured via word-embedding similarity to an economics lexicon; this effect is statistically significant in the short-run event-study window but does not persist over the full career. Second, Manne attendance raises conservative voting in economics-related cases (labor and regulation) by approximately one-quarter of a standard deviation — corresponding to judges deciding in the conservative direction about 20 percent more often relative to the mean — with no significant effect on non-economics cases; the interaction effect is robust across specifications including never-attenders. Third, post-Manne judges vote more frequently against federal labor and environmental regulatory agencies, a result that is statistically significant and economically meaningful with no detectable pre-trends. Fourth, post-Manne judges impose longer and more frequent prison sentences, with no increase in sentencing harshness for drug crimes — consistent with Manne instructors having explicitly advocated drug legalization — and with the harshness gap between Manne and non-Manne judges widening after the 2005 Booker decision expanded judicial sentencing discretion. Fifth, there is some evidence of increased voting against antitrust enforcement, though this result is more sensitive to specification. Persuasion rates computed following DellaVigna and Gentzkow (2010) are slightly larger than those estimated for partisan media interventions such as Fox News and are closest to the effect of a 10-week Washington Post subscription on Democratic governor vote share. Neither the legalist model (judges follow statutes mechanically) nor the attitudinal model (judges follow party affiliation) can explain these within-judge, within-party shifts.&lt;/p&gt;
&lt;p&gt;Q: What is the central identification challenge and how do the authors address it?
A: The key threat is that judges who chose to attend the Manne program — or who attended at a particular time — may differ systematically from non-attenders in ways correlated with their decision trajectories. The authors address this in two steps. First, they restrict the control group to other ever-attending judges who had not yet attended, exploiting the first-come-first-served oversubscription rule that created quasi-random variation in timing among applicants. Second, they use judge fixed effects plus circuit-by-year fixed effects, and add elastic-net-selected biographical covariates (e.g., birth cohort indicators) interacted with year fixed effects as a robustness check. Republican affiliation — the most salient ideological predictor of attendance — is not a statistically significant predictor of attendance timing, supporting the exclusion restriction.&lt;/p&gt;
&lt;p&gt;Q: Why are never-attenders excluded from the preferred control group?
A: Never-attenders differ from attenders on observables including political party and show a positively trending use of economics language in their opinions even before any treatment, suggesting ambient diffusion of economics ideas through law clerks, law school curricula, and organizations such as the Federalist Society. Including never-attenders in the control group produces a near-zero coefficient on the language outcome, which the authors interpret as reflecting spillovers rather than a true null effect; the coefficient on conservative voting in the interaction specification, however, remains positive and significant even when never-attenders are included.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the effect on economics language use?
A: The within-judge effect of Manne attendance on the word-embedding similarity between judicial opinions and an economics lexicon is approximately one-third of a standard deviation, statistically significant in the short-run event-study window (covering six years before and after attendance). The effect shrinks and becomes non-significant when the full career of Manne judges is examined (rather than just the event-study window), consistent with broad diffusion of economics language across the judiciary over time rather than a persistent individual-level treatment effect.&lt;/p&gt;
&lt;p&gt;Q: How large is the effect on conservative voting, and is it concentrated in particular case types?
A: Post-Manne attendance raises conservative voting in economics-related cases (labor and regulation) by approximately one-quarter of a standard deviation, corresponding to judges deciding in the conservative direction about 20 percent more often relative to the mean liberal-conservative decision rate. There is no statistically significant effect on non-economics cases. The interaction coefficient — the differential effect on economics versus non-economics cases — is positive and significant across all specifications including the full sample with never-attenders, making this the most robust directional result in the paper.&lt;/p&gt;
&lt;p&gt;Q: What is the effect on regulatory agency voting?
A: Post-Manne judges vote more frequently against federal labor agencies (National Labor Relations Board, OSHA, Department of Labor, Federal Labor Relations Authority, Office of Worker&amp;rsquo;s Compensation Programs) and the Environmental Protection Agency. The event study shows a positive and significant increase that persists across the event-study window with no detectable pre-trends. This result is robust to both the baseline specification and the elastic-net-controls specification.&lt;/p&gt;
&lt;p&gt;Q: What is the effect on criminal sentencing, and what heterogeneity is found?
A: Post-Manne judges impose both more frequent prison sentences and longer sentences, consistent with Becker&amp;rsquo;s deterrence framework taught in the program&amp;rsquo;s criminal law curriculum. The sentencing effects are absent for drug crimes, consistent with Manne instructors — including Milton Friedman — having explicitly advocated against the drug war and for drug legalization. The gap in sentencing harshness between Manne and non-Manne judges widens after the 2005 United States v. Booker decision, which made the Federal Sentencing Guidelines advisory rather than mandatory; this is consistent with the program having shaped latent judicial preferences that are expressed more fully when formal constraints are relaxed.&lt;/p&gt;
&lt;p&gt;Q: How do the persuasion rates compare to benchmark media studies?
A: The persuasion rates computed following DellaVigna and Gentzkow (2010) are slightly larger than those estimated for partisan media interventions such as Fox News (DellaVigna and Kaplan, 2007) and are closest to the persuasion rates implied by a 10-week subscription to the Washington Post on Democratic governor vote share (Gerber et al. 2009). The comparison contextualizes the Manne program as a moderately high-intensity ideational intervention relative to documented cases of political persuasion.&lt;/p&gt;
&lt;p&gt;Q: What do the results imply for theories of judicial behavior?
A: The findings are inconsistent with both the legalist/formalist model — under which judges apply statutes and precedent without regard to extra-legal factors, predicting zero effect — and the attitudinal model — under which judges simply follow partisan preferences, also predicting zero effect since the program attended judges of both parties. The within-judge, within-party shifts point to a third channel: judicial worldviews and economic ideas, independent of formal law and partisan affiliation, shape high-stakes precedent-setting decisions.&lt;/p&gt;
&lt;p&gt;Q: Can the authors distinguish between a pedagogical (informational) and an ideological persuasion mechanism?
A: They cannot definitively distinguish between the two. Both mechanisms predict increased economics language, more conservative rulings in economics cases, deregulatory voting, and harsher non-drug sentences. The drug-crime heterogeneity is somewhat more consistent with a nuanced pedagogical channel, since Manne instructors explicitly discussed drug legalization, but this pattern is also consistent with complex ideological effects. Evidence on decision quality (citation rates, judicial promotion) is mixed and not robust, providing no clean test of the informational mechanism.&lt;/p&gt;
&lt;p&gt;Q: What does the antitrust evidence show?
A: Post-Manne judges tend to vote against antitrust claimants (i.e., in favor of less antitrust enforcement), but this result is more sensitive to specification than the regulatory agency and sentencing results and is not always statistically significant across specifications. The authors treat it as suggestive rather than conclusive.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to the literature on economics education and normative beliefs?
A: Prior work finds that economics students are less redistributive (Selten and Ockenfels 1998), view surge prices more favorably (Frey and Meier 2005), favor profit maximization (Rubinstein 2006), and that economics professors are less ideologically liberal than other social scientists (Jelveh et al. 2018). The present paper extends this literature by studying established professionals (judges) making high-stakes real-world decisions, and by documenting a direct policy impact rather than a change in survey responses or experimental choices.&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the dataset and the program coverage?
A: The circuit court dataset covers approximately 200,000 published opinions from 1970 through 2005. The district court sentencing dataset covers approximately 1.03 million cases from 1992 through 2003 (event study sample). The Manne program ran from 1976 to 1998, with roughly twenty judges per cohort; by 1990 forty percent of federal judges had attended, and by the late 1990s roughly half of circuit court cases had a Manne-trained panelist. Biographical information comes from the Federal Judicial Center; program attendance lists come from Butler (1999) supplemented by FOIA-obtained annual reports.&lt;/p&gt;
&lt;p&gt;Manne Economics Institute for Federal Judges: An intensive two-week economics training program for sitting U.S. federal judges, run by the Law and Economics Center from 1976 to 1998, covering supply-and-demand theory, the Coase Theorem, externalities, property rights, deterrence theory, and related topics; funded by pro-business foundations; admitted judges on a first-come-first-served basis and trained nearly half of all federal judges over its operation.&lt;/p&gt;
&lt;p&gt;Word-embedding economics language measure: A continuous measure of how closely a judicial opinion&amp;rsquo;s vocabulary aligns with a lexicon of law-and-economics phrases, constructed using word2vec embeddings (Mikolov et al. 2013) trained on the corpus of judicial opinions; measures the semantic proximity of opinion text to the Ellickson (2000) economics lexicon in embedding space, capturing implicit and contextual use of economics reasoning rather than raw phrase counts.&lt;/p&gt;
&lt;p&gt;Deterrence theory (Becker model): The framework, drawn from Becker (1968), taught in the Manne program&amp;rsquo;s criminal law curriculum, which holds that optimal crime deterrence requires setting the expected penalty — the economic cost of punishment times the probability of detection — high enough to outweigh the expected benefits of crime; treated in the paper as the theoretical basis for predicting harsher sentencing among post-Manne judges, and contrasted with retribution- or rehabilitation-based sentencing rationales that dominated before its diffusion.&lt;/p&gt;
&lt;p&gt;Conservative judicial decision (economics cases): In the paper&amp;rsquo;s usage, a ruling against the liberal/pro-plaintiff position in a case involving labor or regulation, as hand-coded by the Songer-Auburn database; includes ruling against a labor agency, rejecting a regulatory claimant, or voting against antitrust enforcement; the paper finds Manne attendance shifts judges in this direction in economics cases but not in non-economics cases.&lt;/p&gt;
&lt;p&gt;First-come-first-served oversubscription: The admission rule of the Manne program during its oversubscribed heyday (from the second cohort in 1977 through the late 1980s), under which applicants who did not secure a spot were bumped to the next year&amp;rsquo;s cohort; the authors argue this rule generates quasi-random variation in the timing of attendance among ever-attending judges, conditional on applying, providing the identifying variation for the differences-in-differences design.&lt;/p&gt;
&lt;p&gt;Persuasion rate: A summary statistic, following DellaVigna and Gentzkow (2010), measuring the fraction of the &amp;ldquo;persuadable&amp;rdquo; population that is convinced by a treatment; used in the paper to benchmark the Manne program&amp;rsquo;s effect size against documented media persuasion interventions such as Fox News and Washington Post subscriptions.&lt;/p&gt;</description></item><item><title>Identification and Estimation of Dynamic Random Coefficient Models</title><link>https://macropaperwarehouse.com/papers/identification-and-estimation-of-dynamic-random-coefficient-models/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/identification-and-estimation-of-dynamic-random-coefficient-models/</guid><description>&lt;p&gt;This paper studies linear panel data models where regression coefficients are individual-specific (random coefficients) and regressors may be predetermined — that is, sequentially exogenous rather than strictly exogenous, as occurs when a lagged dependent variable appears on the right-hand side. The canonical example is the AR(1) model Yit = gamma_i + beta_i * Yi,t-1 + epsilon_it, where both the intercept and the autoregressive coefficient vary across individuals. The setting is short panels (small T), which rules out learning about individual-level coefficient values.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central finding, building on Chamberlain (1993, 2022), is that the mean of the coefficient distribution is not point-identified in this dynamic setting. Chamberlain established this for discrete regressors; the paper&amp;rsquo;s Proposition 1 extends the non-identification result to continuous regressors under stronger assumptions. The paper then characterizes finite lower and upper bounds for the mean, variance, and CDF of the random coefficient distribution. The identification strategy recasts the problem as an infinite-dimensional linear program and exploits the dual representation of that program (following Galichon and Henry (2009) and Schennach (2014)) to derive tractable closed-form bounds for the mean and optimization-based bounds for the variance and CDF.&lt;/p&gt;
&lt;p&gt;For the mean parameter, the bounds take a closed-form expression involving the individual OLS estimator, the pooled OLS estimator, and cross-sectional moments of the data. The bounds remain finite even when the data are unbounded, provided certain moments of the data are finite. Tighter (refined) bounds are available when instrumental variables are brought in as additional unconditional moment restrictions. A numerical illustration shows how the outer identified set for E(beta_i) with a true value of 0.5 shrinks as T increases: at T=3 the outer set is approximately [0.216, 0.617]; at T=5 it narrows to approximately [0.306, 0.613]; the corresponding sharp identified sets (available for T=3 through T=5) range from [0.401, 0.593] at T=3 to [0.473, 0.532] at T=5.&lt;/p&gt;
&lt;p&gt;The paper proposes computationally tractable inference procedures matched to each parameter. For mean parameters, the closed-form bounds permit a delta-method asymptotic approach augmented with Stoye&amp;rsquo;s (2020) smooth approximation to handle cases where the sample analog of the bound width can be negative (due to overidentification or mild misspecification). The resulting confidence intervals are valid and robust to overidentification. For the variance and CDF of the coefficient distribution, the paper uses the Andrews and Shi (2017) procedure for inference on a continuum of moment inequalities, which remains computationally feasible.&lt;/p&gt;
&lt;p&gt;The empirical application estimates a generalization of Guvenen&amp;rsquo;s (2007, 2009) lifecycle earnings models using the Panel Study of Income Dynamics (PSID). Where Guvenen compared a restricted income profile (RIP, homogeneous persistence rho) against a heterogeneous income profile (HIP, heterogeneous time trend beta_i), this paper allows persistence rho itself to vary across households (rho_i). The key empirical findings are: (1) under both the RIP and HIP specifications, the estimated average earnings persistence E(rho_i) is significantly below 1; (2) the two specifications produce similar mean-persistence estimates once heterogeneity in rho_i is permitted, suggesting that misspecifying HIP as RIP or vice versa may not cause serious model misspecification when earnings persistence is allowed to vary; (3) the identified sets for the variance of rho_i provide evidence of genuine heterogeneity in earnings persistence across households, implying that households face different levels of earnings risk, which in turn contributes to heterogeneity in their consumption and savings behavior.&lt;/p&gt;
&lt;p&gt;Q: Why is the mean of the random coefficient not point-identified in a short dynamic panel?
A: Chamberlain (1993, 2022) first established this non-identification for discrete regressors. The paper&amp;rsquo;s Proposition 1 extends the result to continuous regressors under stronger assumptions. The fundamental obstacle is Lemma 1: E(beta_i) is point-identified if and only if there exists an unbiased estimator of beta_i in the individual time series, and no such estimator exists in short panels where T is small relative to the number of individual parameters.&lt;/p&gt;
&lt;p&gt;Q: How does the paper characterize the identified set for the mean parameter?
A: The identification problem is recast as an infinite-dimensional linear program. Using the dual representation (Galichon and Henry, 2009; Schennach, 2014), Theorem 1 yields a closed-form interval [L, U] = [BR - (1/2)&lt;em&gt;sqrt(ER&lt;/em&gt;DR), BR + (1/2)&lt;em&gt;sqrt(ER&lt;/em&gt;DR)], where BR is a weighted average of the individual OLS estimator and the pooled OLS estimator, ER is a non-negative term capturing cross-sectional variation in design matrices, and DR is a non-negative term related to residual variation. The bounds are finite whenever the relevant moments of the data are finite, even with unbounded data.&lt;/p&gt;
&lt;p&gt;Q: How are the bounds tightened using instruments?
A: Proposition 2 introduces refined bounds [LS, US] by incorporating additional unconditional moment restrictions from instruments Sit. The refined bounds use a larger set of restrictions and are weakly tighter than the baseline bounds. The empirical application employs up to 59 regressors with homogeneous coefficients (handled by Proposition 3), and instruments from lagged earnings levels and differences, substantially increasing the number of moment conditions.&lt;/p&gt;
&lt;p&gt;Q: How are the variance and CDF of the coefficient distribution identified?
A: Theorem 2 provides a general duality result for any parameter theta of the coefficient distribution. The lower bound is the maximum of E[min_{b} {m(Wi,b) + sum_k lambda_k phi_k(Wi,b)}] over Lagrange multipliers lambda, and the upper bound is the minimum of the corresponding maximum. Proposition 5 and Proposition 6 specialize this to the second moment (variance) of beta_i, with the upper bound requiring an eigenvalue assumption (Assumption 9) that the smallest eigenvalue of the individual design matrix R&amp;rsquo;R is bounded away from zero. Proposition 7 derives lower and upper bounds for the CDF P(e&amp;rsquo;Bi &amp;lt;= c) using a two-step optimization that separates the support into two regions.&lt;/p&gt;
&lt;p&gt;Q: What guarantees computational tractability of the optimization problems?
A: Proposition 4 establishes that GL(lambda, w) is globally concave in lambda for every w, and GU(lambda, w) is globally convex in lambda for every w. This means the optimization problems for the lower and upper bounds are concave maximization and convex minimization problems respectively, which can be solved with standard convex optimization methods.&lt;/p&gt;
&lt;p&gt;Q: How does the inference procedure for mean parameters handle overidentification and misspecification?
A: In finite samples, the sample analog of the bound-width term D_hat_S can be negative, which would make the estimated bounds degenerate. The paper adopts Stoye&amp;rsquo;s (2020) approach using the smooth approximation s(x,y) = sqrt((xy + sqrt((xy)^2 + r^2))/2). The (1-alpha)-level confidence interval combines a standard bound-based interval with an interval for a pseudo-true parameter mu*_e, ensuring validity under both correct specification and mild overidentification or misspecification.&lt;/p&gt;
&lt;p&gt;Q: How does this paper&amp;rsquo;s approach to inference on the variance and CDF differ from that for the mean?
A: For the mean, closed-form bounds permit a straightforward delta-method asymptotic argument and explicit confidence intervals. For the variance and CDF, the paper uses the Andrews and Shi (2017) procedure for inference on a continuum of moment inequalities, constructing a test statistic TAS(theta) = sup_{lambda} max{sqrt(N)&lt;em&gt;(mu_hat_GL - theta)/sigma_hat_GL, sqrt(N)&lt;/em&gt;(theta - mu_hat_GU)/sigma_hat_GU}^2, 0, with the confidence set being the set of theta values not rejected. This procedure is computationally more demanding but remains feasible.&lt;/p&gt;
&lt;p&gt;Q: What are the main empirical findings from the PSID application?
A: In both the RIP and HIP specifications extended to allow heterogeneous persistence rho_i, the estimated average earnings persistence E(rho_i) is significantly below 1. Both specifications produce similar mean-persistence estimates once rho_i heterogeneity is permitted, suggesting that the HIP vs. RIP misspecification debate may be less consequential when persistence itself varies across households. The identified sets for the variance of rho_i provide evidence of genuine unobserved heterogeneity in earnings persistence.&lt;/p&gt;
&lt;p&gt;Q: What is the economic significance of heterogeneous earnings persistence?
A: Heterogeneity in earnings persistence rho_i means households face different levels of earnings risk: a household with high rho_i experiences earnings shocks that are more persistent, reducing its ability to smooth consumption over time and strengthening its motive for precautionary savings. The paper argues this heterogeneity contributes directly to heterogeneity in consumption and savings behavior, making rho_i a first-order parameter in lifecycle consumption models such as those of Hall and Mishkin (1982), Blundell, Pistaferri, and Preston (2008), and Arellano, Blundell, and Bonhomme (2017).&lt;/p&gt;
&lt;p&gt;Q: How does the paper situate itself relative to Guvenen (2007, 2009)?
A: Guvenen showed that allowing for heterogeneity in the time trend of earnings (HIP: heterogeneous income profile) yields estimated persistence significantly below 1, whereas imposing no such heterogeneity (RIP: restricted income profile) yields persistence near 1. This paper generalizes both models by additionally allowing persistence itself to vary across households (rho_i). The finding that both HIP and RIP deliver similar E(rho_i) estimates significantly below 1 suggests that Guvenen&amp;rsquo;s contrast may be partly an artifact of restricting persistence to be homogeneous.&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the identification results?
A: The results apply to short panels (small T, large N), accommodate discrete, continuous, and unbounded data, and require the idiosyncratic error epsilon_it to be mean-independent of the full history of strictly exogenous regressors and of the current history of predetermined regressors. The bounds for the mean are finite under finite moment conditions on the data. The bounds for the variance additionally require the eigenvalue assumption (Assumption 9). The paper notes that the results extend to probit and logit models with individual-specific coefficients, panel VAR models, and systems of panel data regressions, though these extensions are not developed in detail.&lt;/p&gt;
&lt;p&gt;Dynamic random coefficient model: A linear panel data model in which both the intercept and slope coefficients are individual-specific (gamma_i, beta_i), the regressor is predetermined (sequentially exogenous rather than strictly exogenous), and T is small — so individual coefficient values cannot be estimated from the time series alone.&lt;/p&gt;
&lt;p&gt;Partial identification: The property that a parameter of interest (such as E(beta_i)) cannot be consistently estimated from the data (it is not point-identified), but finite lower and upper bounds on its value can be characterized. The paper shows this is the generic situation for dynamic random coefficient models in short panels.&lt;/p&gt;
&lt;p&gt;Dual representation of infinite-dimensional linear programs: The technique, following Galichon and Henry (2009) and Schennach (2014), of converting an infinite-dimensional linear programming problem (which arises when data or coefficients are continuous) into an equivalent dual problem that yields tractable closed-form or convex-optimization-based bounds.&lt;/p&gt;
&lt;p&gt;Refined bounds (instrument-augmented bounds): Tighter identified sets for the mean parameter obtained by incorporating additional unconditional moment restrictions from instruments Sit, beyond the baseline moment conditions. These correspond to Proposition 2 and make the identification interval weakly narrower.&lt;/p&gt;
&lt;p&gt;Sequential exogeneity (predetermined regressor): The assumption E(epsilon_it | gamma_i, beta_i, Zi1,&amp;hellip;,ZiT, Xi1,&amp;hellip;,Xit) = 0, which allows the regressor Xit (e.g., Yi,t-1) to be correlated with future errors but not current or past errors. This is weaker than strict exogeneity and is what makes the model dynamic and identification challenging.&lt;/p&gt;
&lt;p&gt;Heterogeneous income profile (HIP) vs. restricted income profile (RIP): In Guvenen&amp;rsquo;s framework, HIP allows the time trend of earnings to vary across individuals (heterogeneous beta_i), while RIP does not. The paper extends both by also allowing the AR(1) persistence parameter rho to vary across individuals (rho_i), yielding an empirically more general earnings process.&lt;/p&gt;
&lt;p&gt;Earnings persistence (rho_i): The individual-specific autoregressive coefficient in the lifecycle earnings process. High rho_i means earnings shocks last longer, increasing earnings risk, reducing the household&amp;rsquo;s ability to smooth consumption, and strengthening precautionary savings motives. The paper finds evidence that rho_i varies meaningfully across U.S. households in the PSID.&lt;/p&gt;</description></item><item><title>Identification of Time-Inconsistent Models: The Case of Insecticide-Treated Nets</title><link>https://macropaperwarehouse.com/papers/identification-of-time-inconsistent-models-the-case-of-insecticide-treated-nets/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/identification-of-time-inconsistent-models-the-case-of-insecticide-treated-nets/</guid><description>&lt;p&gt;This paper addresses two related problems: the formal identification of time-inconsistent preferences in dynamic discrete choice models with unobserved heterogeneous types, and the structural estimation of those preferences using data from a health intervention in rural Orissa, India. The identification challenge is fundamental — even the standard exponential discount factor delta is generically not identified in dynamic choice models (Rust 1994; Magnac and Thesmar 2002), and this non-identification extends a fortiori to the hyperbolic (beta, delta) parameterization. The paper&amp;rsquo;s first contribution is constructing identification conditions that overcome these results through two exclusion restrictions: a variable z that affects utility only through the perceived value of future states (played in the application by elicited beliefs about state evolution), and a variable r that acts as an imperfect signal of agent type but is uninformative about choices conditional on type.&lt;/p&gt;
&lt;p&gt;The general model accommodates a finite but unknown number of agent types — time-consistent (beta=1), time-inconsistent naive (beta&amp;lt;1, unaware of future present-bias), and time-inconsistent sophisticated (beta&amp;lt;1, aware of future present-bias) — as well as sub-types within each class. The paper proceeds in four identification steps when types are unobserved: identifying the total number of types (via the rank of an observable matrix), recovering type-specific choice probabilities, assigning type identities, and recovering preference parameters. For time-consistent and sophisticated agents, both beta and delta are point-identified. For naive agents, the parameters are set-identified in general, with point identification available under a monotonicity condition (Assumption 14) or by imposing a common exponential discount factor across types (Assumption 15).&lt;/p&gt;
&lt;p&gt;The empirical application studies demand for insecticide-treated nets (ITNs) and their periodic retreatment — a health-protective technology with low up-front cost but substantial future benefits — among households in malarious areas of rural Orissa. A key design feature is that households were offered either a standard ITN contract (with the option to purchase retreatment later) or a commitment contract bundling two consecutive retreatments, allowing the commitment product choice to serve as a noisy type signal r. Elicited beliefs about future state variables serve as the excluded z variable.&lt;/p&gt;
&lt;p&gt;The main empirical findings are: approximately 21% of the population is time-consistent, 49% are naive time-inconsistent, and 30% are sophisticated time-inconsistent — so time-inconsistent agents account for approximately 79% of the sample. The preferred estimates of the hyperbolic parameter beta are 0.16 for naive agents and 0.08 for sophisticated agents, indicating substantial present-bias in both groups. These estimates of the population type distribution and type-specific beta parameters are described as new to the literature.&lt;/p&gt;
&lt;p&gt;A counterfactual exercise quantifies the welfare cost of present-bias: the median undiscounted additional expected total cost of malaria during the study period attributable to under-investment in ITNs exceeds the price of a treated net by a factor of approximately six. However, because time-inconsistent households heavily discount future malaria costs, the discounted total costs of malaria are low for many inconsistent agents relative to the ITN price, explaining low demand from the agents&amp;rsquo; own subjective perspective. The paper also finds that commitment products are not disproportionately chosen by sophisticated agents — take-up of the commitment contract is actually higher among naive households — contradicting the deterministic mapping from commitment product purchase to sophistication that is commonly assumed in the literature. Finally, differences in per-period utilities across agent types exist but are not substantively important in explaining differential outcomes in the sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the core identification problem the paper addresses, and why is it hard?&lt;/strong&gt;
A: Even the standard exponential discount factor delta is generically not identified in dynamic discrete choice models (Rust 1994; Magnac and Thesmar 2002). This non-identification extends a fortiori to both beta and delta in the hyperbolic (beta, delta) model. When agents are also heterogeneous in unobserved type, the additional problem of identifying the population distribution of types — itself a key policy parameter — must be solved jointly with preference identification.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What two exclusion restrictions provide the key identifying variation?&lt;/strong&gt;
A: The first restriction is a variable z that affects utility only via the perceived value of future states but not per-period utility (Assumption 3); in the application this is played by elicited subjective beliefs about future state evolution. The second is a variable r that predicts agent type but, conditional on type and observables, provides no additional information about choices (Assumption 16); in the application r includes elicited time-preference indicators and the choice of the commitment versus standard ITN contract.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does the paper require at least three periods?&lt;/strong&gt;
A: Three periods are the minimum required to capture the notions of time-inconsistency studied here: with only two periods, no time-inconsistency problem would arise. Three periods allow the researcher to separately observe how an agent plans in period 1, how the agent actually behaves in period 2 (potentially deviating from the period-1 plan), and how the agent behaves in the terminal period 3 where the problem reduces to a static discrete choice.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is point-identified versus set-identified across agent types?&lt;/strong&gt;
A: For time-consistent agents, all per-period utilities and the (single) discount factor delta are point-identified. For sophisticated agents, both beta and delta are separately point-identified under the rank conditions in Assumptions 10-11. For naive agents, the parameters are in general only set-identified (Lemma 4 provides sharp bounds); point identification holds under either a monotonicity condition (Assumption 14) or the assumption that naive and sophisticated agents share the same exponential discount factor (Assumption 15).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the paper identify the total number of types in the population?&lt;/strong&gt;
A: The number of types equals the rank of a directly identified matrix P formed from the joint distribution of actions and states in adjacent time periods (Proposition 1). The rank provides a lower bound in general and equals the true number of types when the state space is sufficiently rich and type-specific choice probabilities vary sufficiently across the state space (Assumptions 17 and 19).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the paper distinguish naive from sophisticated agents among the identified type-specific choice probabilities?&lt;/strong&gt;
A: A key diagnostic is the function delta_hat_tau(x2,z2), which compares an agent&amp;rsquo;s period-1 view of the future against what would be expected given period 2-3 choices. For time-consistent and sophisticated agents, this function is constant across the state space (x2,z2); for naive agents it varies across the state space (Lemma 7, Proposition 2). This variation arises because naive agents incorrectly anticipate their future behavior in period 1, generating a wedge between planned and actual continuation values that shifts with the state.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What fraction of the sample is time-inconsistent, and what are the estimated beta parameters?&lt;/strong&gt;
A: Approximately 79% of the sample is time-inconsistent: 49% are naive and 30% are sophisticated. The preferred estimates of the hyperbolic (present-bias) parameter beta are 0.16 for naive agents and 0.08 for sophisticated agents. Both estimates indicate substantial present-bias. The paper states that these estimates of the population type distribution and the type-specific beta values are new to the literature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the welfare cost of present-bias in terms of malaria risk?&lt;/strong&gt;
A: Present-bias leads to lower ITN purchases and fewer retreatments, which increases the likelihood of contracting malaria. The median undiscounted additional expected total cost of malaria during the study period attributable to under-investment in ITNs exceeds the price of a treated net by a factor of approximately six. However, because inconsistent agents heavily discount future health costs, the discounted total costs of malaria are low relative to the ITN price for many such agents, which explains low demand from the agents&amp;rsquo; own subjective perspective despite large social costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the paper find about commitment products and agent sophistication?&lt;/strong&gt;
A: The commitment contract — bundling two consecutive retreatments — was designed to appeal to sophisticated present-biased agents who anticipate their future self-control problems. Contrary to the deterministic mapping from commitment product purchase to agent sophistication commonly assumed in the literature, take-up of the commitment contract is actually higher among naive households than sophisticated ones. The paper argues this is possible because the model allows commitment product choice to only imperfectly predict type, enabling a richer analysis than prior work that rules out type heterogeneity by assumption.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Are differences in per-period utilities across types an important alternative explanation for observed behavior?&lt;/strong&gt;
A: Per-period utilities do vary across agent types, but the paper finds they are not substantively important in explaining differential outcomes in the sample. This finding supports the interpretation that time-inconsistent preferences — rather than heterogeneity in static preferences over states — are the primary driver of the behavioral differences observed across agent types in this context.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the role of elicited beliefs in the identification strategy?&lt;/strong&gt;
A: Elicited beliefs about the future evolution of state variables serve as the excluded variable z that shifts the forward-looking component of the value function while leaving per-period utility unchanged. The use of expectational data, as advocated by Manski (2004), provides a natural and interpretable source of identifying variation for the discount parameters. The paper argues that this plausible exclusion restriction contributes to the encouraging Monte Carlo simulation results relative to other work in the identification literature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What happens to identification under partial sophistication?&lt;/strong&gt;
A: When agents are partially sophisticated — aware of some but not all of their future present-bias, so that beta_tilde in [beta, 1] rather than exactly equal to beta or 1 — the three time-preference parameters (delta, beta, beta_tilde) are not point-identified in general (Proposition 4 provides a set identification result). Point identification requires that the exponential discount factor delta be identified separately. The paper shows that partial and complete sophistication can be distinguished from time-consistency by whether the function delta_hat varies across the state space, and partially sophisticated types can be distinguished from fully sophisticated types under an additional variability condition (Assumption 23, Proposition 3).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hyperbolic (beta-delta) discounting:&lt;/strong&gt; A model of time-inconsistent preferences in which future utility at time s discounted from time t carries the factor beta*delta^(s-t), where beta&amp;lt;1 introduces an additional present-bias relative to pure exponential discounting. The parameter beta governs the wedge between the discount rate applied to immediate versus purely future tradeoffs; delta governs the intertemporal rate of substitution between any two future periods.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sophisticated vs. naive agents:&lt;/strong&gt; Both types are time-inconsistent (beta&amp;lt;1) and both are aware of their current present-bias. Sophisticated agents (tau_S) also correctly anticipate the extent of their future present-bias (beta_tilde = beta), while naive agents (tau_N) incorrectly believe their future self will behave as if beta_tilde = 1. This difference in beliefs about future behavior drives distinct choice dynamics across the three periods, providing the key observable variation used to distinguish the two types.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exclusion restriction (z variable):&lt;/strong&gt; A state variable that enters the transition probabilities and thus the value of future states but does not enter the current per-period utility function (Assumption 3). Variation in z shifts the forward-looking component of the Bellman equation while holding current utility fixed, providing the identifying variation needed to separately recover discount parameters from per-period utility parameters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Type indicator / type proxy (r):&lt;/strong&gt; An observed variable that is informative about an agent&amp;rsquo;s time-preference type but, conditional on type and other observables, provides no additional information about choices (Assumption 16). In the application, r includes elicited time-preference indicators and whether the agent chose the commitment versus standard ITN contract. Critically, the mapping from r to type is imperfect, so r does not directly reveal type for each individual.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conditional choice probability (CCP) inversion:&lt;/strong&gt; Following Hotz and Miller (1993), the type-specific conditional choice probabilities P_tau(a_t|x_t, z_t) — directly identified from data given type — can be inverted to recover per-period utility differences and combinations of discount parameters without solving the full dynamic programming problem. This approach underpins the constructive identification arguments throughout the paper.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Commitment contract:&lt;/strong&gt; A product design in which two consecutive ITN retreatments are bundled at purchase, intended to mitigate the time-inconsistency problem by removing the future self-control decision about retreatment. The commitment contract is theoretically predicted to be preferred by sophisticated present-biased agents; the paper finds this prediction fails empirically, with naive households showing higher take-up.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Present-bias welfare cost:&lt;/strong&gt; The undiscounted additional expected total cost of malaria attributable to under-investment in ITNs driven by present-bias. The paper estimates this cost exceeds the price of a treated net by a factor of approximately six at the median, capturing the gap between the social planner&amp;rsquo;s valuation of ITN adoption and the discounted valuation of time-inconsistent agents.&lt;/p&gt;</description></item><item><title>Inequality and asset prices during Sudden Stops</title><link>https://macropaperwarehouse.com/papers/inequality-and-asset-prices-during-sudden-stops/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/inequality-and-asset-prices-during-sudden-stops/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper studies the cross-sectional dimension of Fisher&amp;rsquo;s (1933) debt-deflation mechanism as it operates during Sudden Stop crises — episodes characterized by large, abrupt reversals in the current account. The central question is how the distribution of wealth and leverage across households shapes the macroeconomic dynamics of financial crises, and whether greater inequality makes Sudden Stops more or less severe.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The empirical analysis uses panel microdata from the Mexican Family Life Survey (MxFLS) across three waves (2002, 2005, 2009), covering a representative sample of approximately 8,400 households in 150 localities. The 2009 wave captures a Sudden Stop in which Mexico&amp;rsquo;s current account reversed by 1.5 percentage points of GDP, per capita consumption fell 7 percent, and housing prices fell 4 percent below pre-crisis trend by 2010. Households are sorted by net wealth and leverage ratio — defined as total debt divided by total assets — to identify how balance sheet heterogeneity drove differentiated asset-holding dynamics during the crisis.&lt;/p&gt;
&lt;p&gt;The theoretical framework is a Bewley small open economy model with heterogeneous agents, incomplete markets, aggregate risk (simultaneous shocks to the international interest rate and total factor productivity), and an occasionally-binding loan-to-value (LtV) collateral constraint. Households hold two assets: a one-period risk-free international bond and a risky domestic collateralizable asset (land). Households face persistent non-insurable idiosyncratic risk in both labor income and dividend returns; the latter creates an endogenous risk-wealth tradeoff, since larger asset holdings raise future income volatility while simultaneously expanding debt capacity. The model is calibrated to Mexican data — matching the leverage ratio distribution in 2005 (10 percent of households financially constrained) and a net foreign asset position of −35 percent of GDP — and solved using the FiPIt algorithm combined with the Krusell-Smith stochastic-simulation approach.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings with Quantitative Magnitudes&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The empirical evidence from Mexico&amp;rsquo;s 2009 crisis reveals sharply divergent asset dynamics across the household balance sheet distribution. Wealthy households (top net-wealth decile) with low leverage increased their real estate holdings by 61.4 percent (annualized, relative to the average) between 2005 and 2009, consistent with a crisis-dampening effect whereby unconstrained agents absorb fire-sales. Wealthy households in the top decile of both net wealth and leverage ratio — financially constrained — reduced their real estate holdings by 36.6 percent, consistent with a crisis-amplifying effect. Cross-country descriptive evidence shows that Sudden Stop episodes are associated with significantly larger contractions in consumption and GDP in more unequal economies (Gini index, World Bank data, 58 Sudden Stop episodes identified by Bianchi and Mendoza 2020).&lt;/p&gt;
&lt;p&gt;In the calibrated model, the crisis-dampening effect dominates relative to the representative agent baseline: the heterogeneous-agents economy produces a smaller decline in asset prices (−0.99 percent vs. −2.57 percent in the representative agent model during crisis episodes), but a larger and more persistent consumption decline (−2.97 percent vs. −1.17 percent) and current account reversals (1.56 percentage points vs. 0.09 percentage points). The wealth Gini index generated by the calibrated model is 0.61, close to the untargeted 2005 Mexican estimate of 0.73. The aggregate equity premium generated is 5.1 percent, close to the data estimate of 6.5 percent; of this, 55.3 percent is attributable to the risk component, 35.9 percent to the persistence effect, and 8.6 percent to the constraint effect.&lt;/p&gt;
&lt;p&gt;When comparing the baseline emerging economy (wealth Gini 0.61) to an advanced economy calibration in which idiosyncratic dividend risk is set to zero (wealth Gini 0.29), crises are milder and less frequent in the more equal economy: consumption drops 1.0 percentage point less, asset prices drop 0.2 percentage points less, and the net foreign debt position is 6.2 percentage points larger relative to GDP. The implied slope coefficient from the model relating consumption declines during Sudden Stops to the income Gini (−11.1) closely matches the cross-country empirical estimate (−11.5). An economy with an income Gini index 0.10 points lower experiences a decline in consumption 1.1 percentage points smaller during a crisis.&lt;/p&gt;
&lt;p&gt;An impulse response to a two-standard-deviation aggregate shock confirms that, conditional on starting from a perfectly equal (symmetric) initial distribution via complete redistribution, declines in consumption and asset prices are approximately 0.5 percentage points smaller than in the baseline economy with the stationary ergodic distribution as initial condition.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Redistributive Dividend Tax&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A flat 30 percent dividend income tax, redistributed as lump-sum transfers, reduces Sudden Stop severity by lowering average asset prices by 9.6 percent relative to the benchmark, which shrinks effective debt capacity and limits bond adjustment during crises. The average current account reversal during a crisis falls by 0.54 percentage points, and aggregate consumption falls by 0.63 percentage points less than in the benchmark. Crisis probability under the benchmark threshold falls from 4.3 to 1.83 percent (less than half). Average welfare improves by a gain equivalent to 2.8 percent of consumption. However, 26.7 percent of households — those more leveraged and three times wealthier than the beneficiaries — experience welfare losses averaging 6.8 percent of consumption, due to asset price declines and tighter financial conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Overall Conclusion&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Both the empirical evidence and the model suggest that economies with lower inequality, whether due to reduced idiosyncratic risk (as in advanced versus emerging economy calibrations) or wealth redistribution across agents with identical idiosyncratic risk processes, experience less severe Sudden Stop crises.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-two-cross-sectional-channels-through-which-household-heterogeneity-affects-the-debt-deflation-mechanism-and-in-which-direction-do-they-move-asset-prices"&gt;Q1. What are the two cross-sectional channels through which household heterogeneity affects the debt-deflation mechanism, and in which direction do they move asset prices?&lt;/h3&gt;
&lt;p&gt;A1: The dampening effect operates when unconstrained wealthy households — who hold diversified portfolios and have precautionary savings in bonds — purchase fire-sold assets from constrained households, relieving downward pressure on asset prices. The amplifying effect operates when highly leveraged households, once pushed into binding credit constraints by declining asset prices, must further liquidate asset positions, deepening the price decline and tightening the collateral constraint for additional households via the pecuniary externality. These two effects move in opposite directions, so the net effect of inequality on crisis severity is theoretically ambiguous and depends on calibration.&lt;/p&gt;
&lt;h3 id="q2-what-specific-empirical-evidence-from-mexicos-2009-sudden-stop-supports-both-cross-sectional-effects"&gt;Q2. What specific empirical evidence from Mexico&amp;rsquo;s 2009 Sudden Stop supports both cross-sectional effects?&lt;/h3&gt;
&lt;p&gt;A2: Using MxFLS microdata, Table 1 in the paper shows that wealthy households (top net-wealth decile) with low leverage (deciles I–VII of leverage) increased their real estate holdings by 61.4 percent between 2005 and 2009 — evidence for the dampening effect. Wealthy households in the top decile of both net wealth and leverage reduced their real estate holdings by 36.6 percent — evidence for the amplifying effect. Between 2005 and 2009, the share of financially constrained households (leverage ratio above 0.168, the 90th percentile) increased by 1.7 percentage points, while the share of financial savers dropped by 5.0 percentage points. The pre-crisis period (2002–2005) shows no comparable divergence, ruling out a mechanical mean-reversion explanation.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-risk-wealth-tradeoff-and-why-is-it-central-to-generating-a-realistic-wealth-and-leverage-distribution-in-the-model"&gt;Q3. What is the risk-wealth tradeoff, and why is it central to generating a realistic wealth and leverage distribution in the model?&lt;/h3&gt;
&lt;p&gt;A3: The risk-wealth tradeoff arises because idiosyncratic dividend risk is endogenous to asset holdings: holding more risky domestic assets increases debt capacity (relaxing borrowing constraints) but also raises future income volatility, since the variance of household flow income is convex in asset holdings. For households earning high dividend realizations, there exists a threshold beyond which precautionary savings motives — driven by rising income risk — dominate the benefit from expanded debt capacity, causing these households to begin accumulating bonds and eventually become net savers. This mechanism generates an empirically plausible distribution in which some households are financially constrained at the LtV limit, others are unconstrained borrowers, and a fraction are net savers holding both domestic assets and positive international bonds.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-model-calibration-match-the-stationary-distribution-of-mexican-households"&gt;Q4. How does the model calibration match the stationary distribution of Mexican households?&lt;/h3&gt;
&lt;p&gt;A4: Three parameters governing the dividend income risk process (average dividend yield, autocorrelation, and standard deviation) are jointly calibrated to match three statistics from the MxFLS 2005 distribution of households: 14.1 percent financial savers (data: 14.2 percent), 75.9 percent unconstrained indebted (data: 75.8 percent), and 10.0 percent financially constrained (data: 10.0 percent). The collateral fraction κ = 0.168 is set equal to the 90th percentile of the leverage ratio distribution in 2005, reflecting that the average delinquency rate for commercial bank household credit was 10.3 percent between 2004 and 2008. The discount factor β = 0.90 matches the average net foreign asset position relative to GDP of −35 percent for Mexico.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-heterogeneous-agents-model-compare-to-the-representative-agent-model-in-terms-of-crisis-dynamics"&gt;Q5. How does the heterogeneous-agents model compare to the representative agent model in terms of crisis dynamics?&lt;/h3&gt;
&lt;p&gt;A5: In the heterogeneous-agents benchmark, the average current account reversal during a Sudden Stop is 1.56 percentage points, consumption falls 2.97 percent, and asset prices fall 0.99 percent below the steady state. In the representative agent model with the same average leverage ratio (κ = 0.12), the current account reversal is only 0.09 percentage points, consumption falls 1.17 percent, and asset prices fall 2.57 percent. The crisis-dampening effect in the heterogeneous economy produces a smaller asset price drop but a larger consumption decline, because leveraged households must make larger consumption adjustments when hit by negative idiosyncratic shocks in addition to the aggregate shock. Impulse response analysis shows the heterogeneous-agents economy generates current account reversals 1.9 percentage points larger than the representative agent, and consumption responses approximately four times larger.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-mechanism-by-which-comparing-emerging-and-advanced-economy-calibrations-shows-that-lower-inequality-leads-to-less-severe-crises"&gt;Q6. What is the mechanism by which comparing emerging and advanced economy calibrations shows that lower inequality leads to less severe crises?&lt;/h3&gt;
&lt;p&gt;A6: The advanced economy calibration sets idiosyncratic dividend risk to zero, eliminating the risk-wealth tradeoff and resulting in a wealth Gini of 0.29 (compared to 0.61 in the baseline). Without dividend risk, households have weaker incentives to accumulate assets as a precautionary buffer against income volatility, so they hold less debt on average and the long-run net foreign debt relative to GDP is 6.2 percentage points larger (i.e., less debt). During a Sudden Stop under this calibration, consumption drops 1.0 percentage point less, asset prices drop 0.2 percentage points less, and the economy is less frequently in crisis. The model-implied slope of consumption decline on income Gini is −11.1, matching the cross-country empirical estimate of −11.5.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-impulse-response-analysis-reveal-about-the-effect-of-wealth-redistribution-on-crisis-severity-holding-idiosyncratic-risk-constant"&gt;Q7. What does the impulse response analysis reveal about the effect of wealth redistribution on crisis severity, holding idiosyncratic risk constant?&lt;/h3&gt;
&lt;p&gt;A7: The impulse response analysis compares the baseline heterogeneous-agents economy (with the stationary ergodic distribution as the initial condition) against a version in which all households are given a perfectly symmetric initial distribution — identical bond and asset holdings equal to long-run averages — while retaining the same idiosyncratic risk processes. The symmetric initial condition corresponds to a complete redistribution of wealth without changing fundamentals. In the first three periods after a two-standard-deviation aggregate shock, the symmetric economy shows declines in consumption and asset prices approximately 0.5 percentage points smaller than the baseline. This demonstrates that even holding the risk environment constant, reducing wealth dispersion mitigates crisis severity.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-equity-premium-decomposition-work-in-the-heterogeneous-agents-model-and-which-components-are-quantitatively-most-important"&gt;Q8. How does the equity premium decomposition work in the heterogeneous-agents model, and which components are quantitatively most important?&lt;/h3&gt;
&lt;p&gt;A8: The aggregate equity premium is decomposed into five components (Equation 7 in the paper): a constraint effect (positive, increasing in the measure and intensity of constrained households), a risk effect (positive, from the negative covariance between the individual stochastic discount factor and individual equity return, weighted more heavily on constrained households), a persistence effect (positive, from the covariance between idiosyncratic dividend return and asset holdings, since high-dividend households accumulate more assets), a trading cost effect (approximately zero in aggregate), and a no-short-sales effect (negative, since households at the short-sales constraint add to asset demand without increasing the marginal benefit of saving). In the calibrated model, the equity premium is 5.1 percent; the risk effect accounts for 55.3 percent, the persistence effect for 35.9 percent, and the constraint effect for 8.6 percent.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-mechanism-by-which-the-dividend-income-tax-reduces-crisis-severity"&gt;Q9. What is the mechanism by which the dividend income tax reduces crisis severity?&lt;/h3&gt;
&lt;p&gt;A9: A flat 30 percent dividend income tax lowers average after-tax dividend returns, reducing households&amp;rsquo; incentive for precautionary accumulation of domestic assets and weakening the risk-wealth tradeoff. As a result, households demand fewer domestic assets and fewer international bonds in normal times. The reduced demand for the domestic asset lowers the equilibrium asset price by 9.6 percent on average relative to the benchmark, which — through the pecuniary externality embedded in the LtV constraint — tightens borrowing constraints, raising the share of financially constrained households from 5.6 to 7.8 percent. Nevertheless, the reduction in equilibrium debt positions means that during a crisis, bond adjustments and consumption drops are more limited: the average current account reversal during crises falls by 0.54 percentage points, and aggregate consumption falls by 0.63 percentage points less than in the benchmark. Crisis probability under the benchmark threshold falls from 4.3 to 1.83 percent.&lt;/p&gt;
&lt;h3 id="q10-who-benefits-and-who-loses-from-the-dividend-income-tax-and-by-how-much"&gt;Q10. Who benefits and who loses from the dividend income tax, and by how much?&lt;/h3&gt;
&lt;p&gt;A10: Among the simulated population, 73.3 percent of households experience welfare gains averaging 6.2 percent of consumption in consumption-equivalent terms, while 26.7 percent experience welfare losses averaging 6.8 percent of consumption. The average welfare gain across all households is equivalent to 2.8 percent of consumption. The households experiencing losses are more leveraged and three times wealthier on average than those that benefit; the policy reduces their net worth through lower asset prices and tightens their financial constraints. The welfare analysis accounts for the transition to the new tax policy.&lt;/p&gt;
&lt;h3 id="q11-why-does-the-representative-agent-model-miss-the-cross-sectional-effects-that-are-central-to-the-papers-mechanism"&gt;Q11. Why does the representative agent model miss the cross-sectional effects that are central to the paper&amp;rsquo;s mechanism?&lt;/h3&gt;
&lt;p&gt;A11: In the representative agent model, all households behave identically and either collectively want to buy or sell assets, but since there is no one to trade with domestically, actual asset holdings remain unchanged by cross-sectional forces. Additionally, the average debt constraint multiplier in the representative agent equals the single household&amp;rsquo;s multiplier, whereas in the heterogeneous model a small fraction of highly constrained households can have much larger individual multipliers, amplifying the aggregate debt-deflation effect. In the calibrated stationary model, 10 percent of constrained households own 7.7 percent of assets and have a consumption share of 9.0 percent, while 75.9 percent of unconstrained indebted households hold 88.1 percent of assets with a consumption share of 78.1 percent — distributional features invisible to a representative agent.&lt;/p&gt;
&lt;h3 id="q12-what-robustness-does-the-model-validation-provide-for-the-quantitative-results"&gt;Q12. What robustness does the model validation provide for the quantitative results?&lt;/h3&gt;
&lt;p&gt;A12: The model reproduces the untargeted net wealth and asset distributions across deciles from MxFLS 2005 closely, with slight underestimation at the top deciles; the exception is the bottom decile of debt (where the model cannot generate households with negative net wealth since default is not modeled). The aggregate law of motion for the Krusell-Smith algorithm fits with R² = 0.99 for bond position and R² = 0.93 for asset price, and Den Haan (2010) accuracy checks show maximum forecast errors of 2.8 (current account) and 1.1 (asset price). The model replicates the untargeted magnitude of current account reversals observed in Mexican Sudden Stops. The wealth Gini of 0.61 is close to the untargeted 2005 Mexican estimate of 0.73, and the equity premium of 5.1 percent is close to the data estimate of 6.5 percent.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Sudden Stop&lt;/strong&gt;: An episode characterized by a large, abrupt reversal in the current account, typically triggered by a sudden halt in foreign capital inflows. In this paper, Sudden Stops are modeled as endogenous crises that arise from the interaction of a negative aggregate shock (simultaneous rise in the international interest rate and decline in total factor productivity) with an occasionally-binding LtV collateral constraint. The paper follows Bianchi and Mendoza (2020) in identifying 58 such episodes over the past four decades.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Debt-deflation mechanism (cross-sectional dimension)&lt;/strong&gt;: The paper studies Fisher&amp;rsquo;s (1933) debt-deflation spiral — in which declining asset prices tighten credit constraints, forcing further asset sales, further depressing prices — through the lens of household heterogeneity. The cross-sectional dimension refers to the fact that different households (wealthy unconstrained vs. highly leveraged constrained) respond differently to price declines, generating two opposing effects: dampening (wealthy buyers absorb fire-sales) and amplifying (constrained households fire-sell additional assets).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Risk-wealth tradeoff&lt;/strong&gt;: A novel feature of the model in which holding more risky domestic assets simultaneously (a) expands debt capacity by relaxing the LtV constraint and (b) increases future income volatility through higher exposure to idiosyncratic dividend risk, since the variance of household flow income is convex in asset holdings. This tradeoff generates the endogenous transition of households from indebted to net-saver status and gives rise to the empirically plausible distribution of savers, unconstrained borrowers, and constrained households.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Loan-to-value (LtV) collateral constraint&lt;/strong&gt;: A borrowing limit requiring that households&amp;rsquo; international debt (negative bond holdings) cannot exceed a fixed fraction κ of the market value of their domestic asset holdings. In the paper, κ = 0.168 (the 90th percentile of the Mexican leverage ratio distribution in 2005). The constraint is occasionally binding and generates a pecuniary externality: households fail to internalize that their individual portfolio choices affect the aggregate asset price, which in turn determines the borrowing limits of all other households.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pecuniary externality&lt;/strong&gt;: The externality arising from the LtV constraint in which each household&amp;rsquo;s choice of asset holdings affects the equilibrium asset price, thereby changing the borrowing limits of all households simultaneously. This externality drives the debt-deflation spiral and is the source of Sudden Stop crises in the model: no single household internalizes the aggregate impact of its fire-sales on credit conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fire-sale&lt;/strong&gt;: In the context of this paper, the forced liquidation of domestic asset holdings by financially constrained households during a crisis. Fire-sales are triggered when the LtV constraint becomes binding, forcing households to sell assets to reduce debt; the resulting price decline tightens the constraint further, producing additional fire-sales. The paper documents that, during Mexico&amp;rsquo;s 2009 Sudden Stop, wealthy constrained households (top decile of both net wealth and leverage) reduced real estate holdings by 36.6 percent, while wealthy unconstrained households increased holdings by 61.4 percent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dampening and amplifying effects&lt;/strong&gt;: Two opposing cross-sectional effects on asset prices during a crisis. The dampening effect: unconstrained wealthy households purchase depressed assets fire-sold by constrained households, relieving downward pressure on prices and weakening the debt-deflation spiral. The amplifying effect: highly leveraged households that are pushed into binding constraints by falling prices must also fire-sell assets, further depressing prices and tightening financial conditions. The net impact on crisis severity depends on which effect dominates, which the paper establishes empirically and quantitatively is inequality-dependent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Equity premium decomposition&lt;/strong&gt;: A decomposition derived in the paper (Equation 7) that expresses the aggregate excess return on the risky domestic asset as the sum of five components: a constraint effect (positive, from the measure and intensity of binding LtV constraints), a risk effect (positive, from the covariance of individual stochastic discount factors with individual equity returns), a persistence effect (positive, from the covariance of idiosyncratic dividend returns with asset holdings due to return persistence), a trading cost effect (approximately zero in aggregate), and a no-short-sales effect (negative). In the calibrated model, the risk and persistence effects account for 91 percent of the 5.1 percent equity premium.&lt;/p&gt;</description></item><item><title>Insuring Peace: Index-Based Livestock Insurance, Droughts, and Conflict</title><link>https://macropaperwarehouse.com/papers/insuring-peace-index-based-livestock-insurance-droughts-and-conflict/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/insuring-peace-index-based-livestock-insurance-droughts-and-conflict/</guid><description>&lt;p&gt;This paper provides quasi-experimental evidence that Index-Based Livestock Insurance (IBLI) — a remote-sensing-triggered, automated payout scheme for pastoralists — substantially reduces drought-induced conflict in Kenya over the 2001–2020 period.&lt;/p&gt;
&lt;p&gt;The research question is whether a market-based financial instrument can mitigate the causal chain running from drought shocks to violent conflict between nomadic pastoralists and sedentary farmers and other land users. The authors motivate the study by documenting that droughts force pastoralists out of their traditional grazing grounds and into mixed-land-use areas (farms, ranches, urban settlements, nature reserves), where miscoordination with other land users escalates into violence. A case study of the Samburu-Laikipia-Isiolo-Meru region in central Kenya — drawing on georeferenced survey data from Lengoiboni et al. (2010) and ACLED conflict events — validates this spatial mechanism: during droughts, roughly 60–90% of non-pastoral land users report encounters with pastoralists, and conflicts accumulate precisely where drought migration routes cross into non-pastoral land.&lt;/p&gt;
&lt;p&gt;The empirical design combines two sources of variation: (1) plausibly exogenous changes in rainfall deficits at the 0.1 × 0.1-degree grid-cell level (roughly 10 × 10 km), derived from NASA GPM satellite data; and (2) the staggered, five-wave rollout of IBLI across 146 insurance districts in Kenya from 2010 onward, which the authors argue was driven primarily by technical challenges rather than pre-existing conflict or drought patterns. The unit of observation is 94,300 cell-periods. Because conflicts due to pastoralist drought migration occur in the neighborhood of affected areas rather than within them, both drought and IBLI coverage are measured as inverse-distance-weighted averages over surrounding cells. The estimating equation is a linear probability model with cell and period fixed effects, interacting neighborhood rainfall deficit with neighborhood IBLI coverage; the coefficient on this interaction term (delta3) is the parameter of interest.&lt;/p&gt;
&lt;p&gt;The main finding is that a one-standard-deviation increase in neighborhood IBLI coverage reduces the semi-elasticity of neighborhood rainfall deficit on conflict probability by approximately 23%. In absolute terms, a one-percentage-point increase in the rainfall deficit raises the probability of conflict by 6.92 percentage points at average IBLI coverage; with one additional standard deviation of neighborhood IBLI, that same deficit raises conflict probability by only 5.34 percentage points — a reduction of 1.58 percentage points against a baseline conflict probability of roughly 2.5%.&lt;/p&gt;
&lt;p&gt;Scope conditions: the effect is estimated for Kenya specifically, over a pastoralist-heavy population of approximately 8.8 million out of 53 million Kenyans, during 2001–2020. The conflict-mitigating effect is approximately four times larger in mixed-land-use areas (nine times when rollout-cluster-times-period fixed effects are included), consistent with the theoretical expectation that IBLI matters most where pastoralists are most likely to encounter other land users during drought migration.&lt;/p&gt;
&lt;p&gt;Two mechanisms are identified. First, IBLI reduces migratory pressure: when pastoral homelands have IBLI coverage, the distance between the ethnic homeland centroid and conflict events involving that group decreases, indicating reduced drought migration. Second, IBLI smooths incomes — corroborated with Afrobarometer geo-coded data — raising the opportunity cost of fighting. An instrumental-variable specification finds that actual IBLI payouts in the neighborhood reduce conflict probability by approximately 150% relative to the baseline risk.&lt;/p&gt;
&lt;p&gt;A cost-effectiveness analysis finds that even using conservative World Health Organization or World Bank estimates of the value of statistical life, IBLI delivers fatality savings of between 10 and 22 cents per dollar spent on government subsidies for the program, making it a cost-effective complement to political and institutional conflict-mitigation approaches.&lt;/p&gt;
&lt;p&gt;Q: What is the core causal mechanism linking droughts to conflict that IBLI interrupts?&lt;/p&gt;
&lt;p&gt;A: Droughts deplete forage in pastoralists&amp;rsquo; traditional grazing grounds, forcing them to migrate into mixed-land-use areas — farms, ranches, urban settlements, and nature reserves — where encounters with other land users are more likely to escalate into violence. Without insurance, pastoralists hold excess livestock as precautionary savings, amplifying the extent of necessary migration during dry periods. IBLI payouts allow pastoralists to purchase forage locally, reducing migration distance and intensity, and also smooth income, raising the opportunity cost of engaging in violence.&lt;/p&gt;
&lt;p&gt;Q: How does IBLI work technically, and why does it overcome problems of traditional livestock insurance?&lt;/p&gt;
&lt;p&gt;A: IBLI uses satellite remote sensing to calculate whether a district-specific drought threshold has been crossed; if so, automated payments are triggered immediately without requiring direct loss assessment or field inspections. This design eliminates moral hazard and adverse selection problems inherent in traditional indemnity insurance, reduces monitoring costs, and enables fast delivery via mobile payment platforms such as MPESA even to remote households. The Kenyan government rebranded the program as the Kenyan Livestock Insurance Program (KLIP) in 2015 and fully subsidizes coverage for up to five tropical livestock units per household.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the main conflict-mitigation result?&lt;/p&gt;
&lt;p&gt;A: A one-standard-deviation increase in neighborhood IBLI coverage reduces the semi-elasticity of the neighborhood rainfall deficit on conflict probability by approximately 23% (delta3/delta1 = -0.0158/0.0692). In absolute terms, this translates to a reduction from a 6.92 percentage-point increase in conflict probability per one-percentage-point rainfall deficit to a 5.34 percentage-point increase — a decline of 1.58 percentage points against a mean conflict probability of roughly 2.5%.&lt;/p&gt;
&lt;p&gt;Q: Why do the authors use a neighborhood rather than cell-level treatment measure?&lt;/p&gt;
&lt;p&gt;A: Drought-induced pastoralist conflicts occur primarily not in the pastoral home areas themselves but in neighboring regions where drought migration routes cross into non-pastoral land. The case study documents this pattern directly: ACLED conflict events accumulate where migration routes from Namelok, Lodungokwe, and Ngaremara communities intersect urban or agricultural areas, not within the pastoral zones. The neighborhood approach, using inverse-distance-weighted averages, captures both the probability of migration from surrounding cells and the declining probability of migration with distance.&lt;/p&gt;
&lt;p&gt;Q: What is the main identification concern and how do the authors address it?&lt;/p&gt;
&lt;p&gt;A: The main concern is that the timing of the IBLI rollout is endogenously determined — areas with a higher latent drought-conflict elasticity might receive coverage earlier or later, biasing the interaction coefficient. The authors show that the pre-treatment drought-conflict elasticity has no systematic correlation with either IBLI eligibility or the timing of coverage receipt. Placebo tests interacting the neighborhood rainfall deficit with pre-treatment eligibility or eventual coverage indicators yield positive, statistically insignificant coefficients, suggesting any bias would run in the direction of underestimating the mitigation effect. A permutation test randomly reassigning IBLI coverage across the six rollout clusters finds the actual point estimate is in the bottom 2.2% of the simulated distribution, indicating it is unlikely to arise from cluster-level confounders.&lt;/p&gt;
&lt;p&gt;Q: How do the authors rule out that other programs — cash transfers or development aid — explain the result?&lt;/p&gt;
&lt;p&gt;A: The authors control for cell-level and neighborhood-level coverage of Kenya&amp;rsquo;s Hunger Safety Net Programme (HSNP), which provides unconditional cash transfers to vulnerable households and covers most IBLI-eligible areas, as well as for World Bank agricultural aid projects. Across these specifications, the estimated conflict mitigation ranges from -19.16% to -42.24%, with the baseline estimate of -22.79% remaining robust, indicating neither HSNP nor development aid is a plausible alternative explanation.&lt;/p&gt;
&lt;p&gt;Q: What is the alternative identification strategy using within-rollout-cluster variation?&lt;/p&gt;
&lt;p&gt;A: The authors exploit pre-determined (1984 government land-use map) variation in mixed-land-use status across cells within the same IBLI rollout cluster-period, including rollout-cluster-times-period fixed effects that absorb any omitted variable related to the potentially endogenous rollout steps. The conflict-mitigating effect of IBLI is approximately four times larger in mixed-land-use cells, and approximately nine times larger in the most restrictive specification with rollout-cluster-times-period fixed effects, consistent with the prediction that IBLI matters most where pastoralists encounter other land users.&lt;/p&gt;
&lt;p&gt;Q: How do the authors establish the migratory pressure mechanism?&lt;/p&gt;
&lt;p&gt;A: Following Eberle et al. (2023), the authors match conflict actors to ethnic homelands using Murdock (1967) boundaries and test whether IBLI coverage in a homeland reduces the distance between the homeland centroid and conflict events involving that group. They find that it does, indicating that IBLI coverage reduces the spatial range of pastoralist drought migration and thus the probability of conflict-generating encounters with other land users.&lt;/p&gt;
&lt;p&gt;Q: How do the authors establish the income-smoothing mechanism?&lt;/p&gt;
&lt;p&gt;A: Using geo-coded Afrobarometer survey data, the authors show that IBLI coverage is associated with higher reported incomes among pastoralist households, consistent with Jensen et al. (2017). Higher incomes raise the opportunity cost of fighting (following Grossman, 1991), contributing to the overall conflict-mitigating effect alongside reduced migratory pressure.&lt;/p&gt;
&lt;p&gt;Q: What does the instrumental variable specification find?&lt;/p&gt;
&lt;p&gt;A: The authors instrument inverse-distance-weighted IBLI payouts in the neighborhood with the interaction of neighborhood rainfall deficit and neighborhood IBLI coverage. The first stage confirms that rainfall deficits trigger payouts conditional on coverage. The second stage finds that the occurrence of payouts in the neighborhood reduces the probability of conflict by approximately 150% relative to the baseline risk, corroborating the reduced-form results.&lt;/p&gt;
&lt;p&gt;Q: How do the authors assess cost-effectiveness?&lt;/p&gt;
&lt;p&gt;A: The authors predict plausible drought-induced conflict fatalities in Kenya over the pre-treatment period and calculate yearly lives saved from the main estimates, then compare the monetary value of saved lives to government subsidy expenditures on IBLI. Using conservative VSL estimates from the WHO and World Bank, IBLI delivers between 10 and 22 cents of pure fatality savings per dollar of public subsidy expenditure.&lt;/p&gt;
&lt;p&gt;Q: How robust are the results to alternative drought and conflict measures?&lt;/p&gt;
&lt;p&gt;A: Results are qualitatively similar using an Aridity Index or Dry Matter Productivity (DMP) as drought proxies instead of rainfall deficit. The estimated interaction effect maintains a t-statistic above two for spatial decay functions ranging from distance^-0.5 to distance^-1.5 and for Conley standard error cutoffs from 200 km up to 400 km. Results also hold when restricting to conflict events not involving the government, or to battles, riots, and violence against civilians only, and when excluding the pre-IBLI period (2000–2009) entirely.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications regarding scalability?&lt;/p&gt;
&lt;p&gt;A: Pastoralism covers 43% of the African landmass across 36 countries, supporting approximately 268 million people (FAO, 2018). The World Bank and private equity were planning to invest close to 900 million dollars in East African pastoralist programs over 2023–2027. The authors argue that IBLI&amp;rsquo;s cost structure — high fixed costs of technology and setup but low marginal costs of expansion — gives it a scalability advantage over cash transfer programs or public works schemes that require sustained state capacity. Market-based IBLI complements rather than substitutes for political and institutional reforms.&lt;/p&gt;
&lt;p&gt;Index-Based Livestock Insurance (IBLI): A financial instrument that uses satellite remote sensing to automatically trigger preemptive cash payouts to pastoralists when a pre-determined district-specific drought threshold is crossed, bypassing direct loss assessment and thereby eliminating moral hazard and adverse selection problems inherent in traditional indemnity insurance.&lt;/p&gt;
&lt;p&gt;Drought-conflict semi-elasticity: The percentage-point change in the probability of conflict associated with a one-percentage-point increase in the rainfall deficit; the paper&amp;rsquo;s main outcome quantity, estimated at 6.92 percentage points at mean IBLI coverage, reduced by 23% for a one-standard-deviation increase in neighborhood IBLI coverage.&lt;/p&gt;
&lt;p&gt;Neighborhood approach: An empirical strategy that measures both drought severity and IBLI coverage as inverse-distance-weighted averages over all surrounding grid cells, reflecting the authors&amp;rsquo; finding that pastoralist drought-migration generates conflicts not in the pastoral home area but in neighboring mixed-land-use zones where migration routes intersect other land users.&lt;/p&gt;
&lt;p&gt;Migratory pressure: The mechanism by which drought forces pastoralists — who hold excess livestock as precautionary savings in the absence of insurance — to migrate farther from traditional grazing grounds into mixed-land-use areas, increasing the probability of encounters and violent miscoordination with farmers, urban dwellers, and protected-area managers.&lt;/p&gt;
&lt;p&gt;Mixed land use: Areas, designated using a 1984 Kenyan government land-use map, where pastoral grazing zones are proximate to farms, ranches, urban settlements, or nature reserves; the paper identifies these as the locations with the highest expected treatment intensity, where IBLI coverage reduces drought-induced conflict approximately four to nine times more than elsewhere.&lt;/p&gt;
&lt;p&gt;Tropical Livestock Unit (TLU): The standard unit of account for IBLI contracts in Kenya; one TLU corresponds to one head of cattle or ten goats or sheep; the Kenyan government fully subsidizes IBLI for up to five TLUs per household.&lt;/p&gt;
&lt;p&gt;Rollout-cluster-times-period fixed effects: A restrictive set of fixed effects included in the alternative identification strategy that absorbs all omitted variables varying at the level of the six IBLI spatial rollout clusters over time, allowing the authors to identify the conflict-mitigating effect purely from within-cluster variation in mixed-land-use exposure.&lt;/p&gt;</description></item><item><title>Intergenerational Impacts of Secondary Education: Experimental Evidence from Ghana</title><link>https://macropaperwarehouse.com/papers/intergenerational-impacts-of-secondary-education-experimental-evidence-from-ghana/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/intergenerational-impacts-of-secondary-education-experimental-evidence-from-ghana/</guid><description>&lt;p&gt;This paper provides experimental evidence on the intergenerational impacts of secondary education subsidies in a low-income context, leveraging a randomized controlled trial (RCT) conducted in rural Ghana with a 15-year longitudinal follow-up. The study exploits a 2008 scholarship lottery in which 682 students — drawn from 2,064 rural youth who had been admitted to public senior high school but had not enrolled due to financial constraints — were randomly selected to receive four-year secondary school scholarships covering full tuition and fees. Scholarship receipt increased senior high school completion by 27–28 percentage points for both men and women (from 39.8% to 67.2% for women; from 49.7% to 77.9% for men), and raised average years of education by 1.33 years.&lt;/p&gt;
&lt;p&gt;The central research question is whether secondary education subsidies generate intergenerational benefits — specifically, whether children of scholarship recipients have better survival and cognitive development outcomes — and what mechanisms drive any such effects.&lt;/p&gt;
&lt;p&gt;For female scholarship recipients, the scholarship significantly altered fertility timing and partnership. By 2013, female recipients were 6.9 percentage points less likely to have ever been pregnant (on a control-group base of 48.3%), with the decline driven almost entirely by a 7 percentage point (17%) reduction in unwanted pregnancies. Though total fertility eventually caught up by 2022, recipients were still less likely to be married or cohabiting as of 2019 and were significantly more likely to have a partner with tertiary education.&lt;/p&gt;
&lt;p&gt;Children of female scholarship recipients experienced substantially lower mortality. Among control-group female respondents, 3.5% of children died before age one and 4.0% before age three. These rates fell to 1.7% (p=0.028) and 2.2% (p=0.065) respectively among children of female recipients — a roughly 45–51% reduction in under-one and under-three mortality.&lt;/p&gt;
&lt;p&gt;Child cognitive development gains emerge only once children reach school age. Children of female recipients show no significant cognitive score differences at 18 months, 2.5 years, or 3.5 years, but score 0.238 standard deviations higher at age five (p=0.005) and 0.252 standard deviations higher at age seven (p=0.035). Effects span language, math and numeracy, spatial reasoning, and executive function, but not socio-cognitive development. These effect sizes fall between the 75th and 80th percentile of RCT-based educational intervention effect sizes in low- and middle-income countries.&lt;/p&gt;
&lt;p&gt;The primary mechanism is not higher income or greater monetary investment in children. The study finds no significant treatment effect on household SES index (0.107 SDs, p=0.103), no impact on formal schooling inputs, and no difference in parental aspirations or knowledge of child stimulation&amp;rsquo;s importance. Instead, more-educated mothers seek more prenatal care, engage in more preventive health behaviors, and — critically — spend more time interacting with their children in stimulating ways. Day-long LENA (Language Environment Analysis) recordings at 18 months confirm 20% more adult-child conversational turns per minute (effect size 0.068, p=0.005) and 17% more child vocalizations per minute (effect size 0.32, p=0.014) for children of female recipients.&lt;/p&gt;
&lt;p&gt;For male scholarship recipients, no analogous intergenerational benefits appear. Their partners are not more educated (in fact slightly less educated on tertiary rates), their children show no mortality improvement, and cognitive scores are if anything negative at age five (point estimate -0.22, p=0.069). The absence of effects is attributed to male scholarship recipients having caregivers — overwhelmingly mothers — with no more education than in the control group, and to children of male recipients being 8.7 percentage points less likely to live with their father.&lt;/p&gt;
&lt;p&gt;A cost-benefit analysis finds internal rates of return (IRR) of 27%–76% for a female-only means-tested scholarship program and 20%–51% for a mixed-gender program. The cost per under-three death averted ($15,184 for female-only) places the scholarship program within the range of the 10th-percentile most cost-effective WHO-recommended child health interventions.&lt;/p&gt;
&lt;p&gt;Scope conditions: the study estimates effects for students who qualified for senior high school but faced binding financial constraints in rural Ghana in 2008 — a population that is well-prepared academically but economically disadvantaged. Results may not generalize to students who would not have qualified for secondary school or to contexts where financial barriers are not binding.&lt;/p&gt;
&lt;p&gt;Q: What was the experimental design and who was in the study sample?
A: In 2008, 2,064 rural Ghanaian students who had been admitted to senior high school (SHS) but had not enrolled — typically due to inability to pay fees — were sampled. After a baseline survey, 682 were randomly selected (approximately one-third) by lottery to receive a four-year scholarship covering full tuition and fees for a day (non-boarding) student, stratified by district, school, gender, and exam-year cohort. The two-thirds comparison group received no scholarship. Students were on average 17 years old at baseline and just over 31 at the last follow-up in Spring 2023.&lt;/p&gt;
&lt;p&gt;Q: How large was the scholarship&amp;rsquo;s effect on educational attainment?
A: Scholarship receipt raised SHS completion from 39.8% to 67.2% among women (a 69% increase) and from 49.7% to 77.9% among men (a 57% increase). Overall, the scholarship led to an average of 1.33 more years of education. For women only, it also significantly raised tertiary education: by 2023, scholarship receipt increased tertiary completion by 10.8 percentage points for women, but had no significant tertiary effect for men.&lt;/p&gt;
&lt;p&gt;Q: What were the effects on fertility and family formation for female scholarship recipients?
A: By 2013, female recipients were 6.9 percentage points less likely to have ever been pregnant (base: 48.3% in control), driven almost entirely by a 7 percentage point (17%) reduction in unwanted pregnancies. By 2019, recipients were still 6 percentage points less likely to have started childbearing and had 0.152 fewer children on average (p=0.065). Total fertility eventually caught up by 2022. By 2016, female recipients were 12.1 percentage points (24% of control mean) less likely to have ever lived with a partner, and by 2019 were 6.2 percentage points less likely to be married or cohabiting. Conditional on having a partner, they were significantly more likely to have a partner who completed tertiary education (p=0.071).&lt;/p&gt;
&lt;p&gt;Q: What were the effects on fertility and family formation for male scholarship recipients?
A: Male recipients showed few changes in fertility or marriage behavior. They were 7.8 percentage points (30% of control mean) more likely to still be living with their parents as of 2019. Their partners were not more educated; in the cognitive games subsample, treatment actually reduced the share of partners with tertiary education by 3.6 percentage points from a control base of 4.3%.&lt;/p&gt;
&lt;p&gt;Q: What were the child mortality results for children of female scholarship recipients?
A: Among children of female control respondents, 3.5% died before age one and 4.0% before age three. These fell to 1.7% (p=0.028) and 2.2% (p=0.065), respectively, among children of female recipients — approximately a halving of under-one and under-three mortality. These point estimates are robust to varying the covariates (linear vs. fixed effects for birth year, dropping or adding controls). After multiple-hypothesis testing adjustment using the Romano-Wolf step-down procedure, the p-value for survived-to-one rises from 0.028 to 0.119.&lt;/p&gt;
&lt;p&gt;Q: What were the child mortality results for children of male scholarship recipients?
A: The estimated effects for children of male recipients were smaller and statistically insignificant: a 1.4 percentage point increase in survived-to-one (p=0.161) and 0.9 percentage points in survived-to-three (p=0.549). These estimates are not significantly different from those for female recipients. Results were sensitive to sample perturbations given the smaller sample: only 26 of 1,016 children of male respondents died before age one.&lt;/p&gt;
&lt;p&gt;Q: What child cognitive development gains did children of female scholarship recipients show, and at what ages?
A: No significant differences emerged at 18 months (-0.066 SDs, p=0.489), 2.5 years (-0.024 SDs, p=0.850), or 3.5 years (0.026 SDs, p=0.736). Significant gains appeared at age five (0.238 SDs, p=0.005) and age seven (0.252 SDs, p=0.035). Effects span language (0.15 SDs at five; 0.27 SDs at seven), math and numeracy (0.15 SDs; 0.26 SDs), spatial reasoning (0.20 SDs; 0.12 SDs), and executive function (0.25 SDs; 0.20 SDs), but not socio-cognitive development. These effect sizes fall between the 75th and 80th percentile of educational RCT effect sizes in low- and middle-income countries.&lt;/p&gt;
&lt;p&gt;Q: What cognitive development effects did children of male scholarship recipients show?
A: No significant positive effects emerged at any age. Point estimates were negative at all ages except 18 months, and marginally significantly negative at age five (-0.22 SDs, p=0.069). The difference in treatment effects between children of male and female recipients is statistically significant at age five (p=0.005).&lt;/p&gt;
&lt;p&gt;Q: Why do cognitive gains appear only at age five and not earlier?
A: The authors offer three interpretations: first, that the cognitive tests for younger children are noisier instruments (cross-sectional and longitudinal correlations within domains are much lower for 1.5-year tests than 5-year tests); second, that impacts on cognitive development may take time to materialize; third, that marginal survivors in the treatment group may start with a cognitive deficit (e.g., surviving a cerebral malaria episode), and maternal education effects require time to overcome this initial handicap. Gains concentrate on skills underlying literacy and numeracy, consistent with more educated mothers bridging home and school environments.&lt;/p&gt;
&lt;p&gt;Q: What is the primary mechanism driving intergenerational effects?
A: The primary mechanism is changes in parenting behaviors, not income. Female recipients do not invest more money in children (no significant difference in SES index or child investment index). Instead, they seek more prenatal care, engage in significantly more preventive health behaviors, and interact more with their children in cognitively stimulating ways. Day-long LENA recordings at 18 months show 20% more conversational turns per minute (effect size 0.068, p=0.005) and 17% more child vocalizations per minute (effect size 0.32, p=0.014). Caregiver reports confirm more playing, singing, and doing simple mathematics with children.&lt;/p&gt;
&lt;p&gt;Q: Does the income effect of scholarship receipt explain the child outcomes?
A: No. Duflo et al. (2024) find no significant earnings impacts until 2019 or later, meaning children tested at ages five and seven by 2023 largely grew up before their mothers&amp;rsquo; earnings improved. The household SES index shows only a 0.107 SD gain (p=0.103), indistinguishable from the effect for children of male recipients. There is also no evidence of a quality-quantity trade-off: caregivers of scholarship recipients do not have fewer children to care for.&lt;/p&gt;
&lt;p&gt;Q: Does the increase in maternal age at birth explain the child mortality reduction?
A: It is not the primary driver. Maternal age at birth increases by only 0.349 years on average (p=0.142) for children of female recipients, and 0.64 years for first-born children (p=0.040). Point estimates on mortality for first-born children are somewhat smaller than for the full sample, suggesting maternal age is not the main channel. Moreover, maternal age at birth falls for children of male recipients yet their survival point estimates are positive, which further argues against maternal age as the primary mechanism.&lt;/p&gt;
&lt;p&gt;Q: How does the education of the primary caregiver mediate the results?
A: For 84% of children in the sample, the primary caregiver is the child&amp;rsquo;s mother. Children of female scholarship recipients have caregivers who are 25 percentage points more likely to have completed secondary school and 5 percentage points more likely to have completed tertiary education. Children of male scholarship recipients have caregivers with no more education than the control group, because the recipients&amp;rsquo; partners — the typical caregivers — are not more educated. Treatment effects for female recipients are not altered when father&amp;rsquo;s education is added as a control, confirming maternal education as the main driver.&lt;/p&gt;
&lt;p&gt;Q: What threat to validity arises from co-residence of the father?
A: Children of male scholarship recipients are 8.7 percentage points less likely to live with their father (p=0.024), compared to no such effect for children of female recipients (92% of whom live with their scholarship-recipient mother). LENA recordings show negative treatment effects for children of male recipients — fewer adult words and conversational turns — consistent with father absence mechanically reducing auditory engagement and possibly leaving single mothers less time to verbally interact with each child.&lt;/p&gt;
&lt;p&gt;Q: How are multiple-hypothesis testing concerns addressed?
A: The pre-analysis plan pre-specified child survival and child cognitive development as primary outcomes. The authors apply the Romano-Wolf step-down procedure for multiple hypothesis testing adjustment. After adjustment, the p-value for survived-to-one for children of female recipients rises from 0.028 to 0.119; the cognitive development effects at age five and seven remain significant.&lt;/p&gt;
&lt;p&gt;Q: How does the study address potential sample selection bias in the child outcomes sample?
A: The authors use entropy balancing (Hainmueller, 2012) to reweight observations so that baseline (2008) characteristics are balanced between treatment and control within the subsample of recipients who had children. Results are qualitatively unchanged for both female and male recipients. The authors also note that children of female recipients are younger on average (4.71 months, p=0.067), which is why the study collects data at fixed age windows (14-22 months, 2.5 years, 3.5 years, 5 years, 7 years) rather than in a single cross-sectional wave.&lt;/p&gt;
&lt;p&gt;Q: What is the cost-effectiveness and cost-benefit result for secondary school scholarships?
A: Social costs are estimated at $585 per recipient for a mixed-gender program and $505 for a female-only program (combining school fees, materials, and foregone wages). The cost per under-three death averted is $23,582 for mixed-gender and $15,184 for female-only — placing the female-only program within the range of the 10th-percentile most cost-effective WHO-recommended child health interventions. The IRR is 27%–76% for a female-only means-tested scholarship program and 20%–51% for a mixed-gender program. These are likely conservative, as they exclude welfare gains from avoiding unwanted pregnancies, greater female agency, and recipient health benefits.&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the experiment and to what population do findings generalize?
A: The study estimates ITT effects for students in rural Ghana who qualified for SHS on exam performance but faced binding financial constraints in 2008 — a population that is academically prepared but economically disadvantaged. Results do not directly apply to students who would not have qualified, to contexts without binding financial barriers, or to settings where secondary school quality or the marriage market differs substantially. The study also cannot yet observe complete fertility, since scholarship-lottery participants were only 31 years old on average at last follow-up.&lt;/p&gt;
&lt;p&gt;LENA (Language Environment Analysis): A day-long recording device worn by a child that uses speech recognition software to generate count-based metrics — adult word count, adult-child conversational turns, and child vocalizations per minute — providing an objective measure of the child&amp;rsquo;s auditory environment and caregiver engagement quality without reliance on self-report.&lt;/p&gt;
&lt;p&gt;IRT Score (Item Response Theory Score): A latent-trait measure of child cognitive ability estimated from a one-parameter logistic model applied to binary correct/incorrect responses across cognitive game questions, assigned a difficulty level to each question and a latent ability to each child, then standardized. Used as the primary cognitive development outcome across age windows.&lt;/p&gt;
&lt;p&gt;Incarceration Effect: The hypothesis that education delays fertility mechanically only while students are in school (analogous to incarceration preventing activity), with no persistent effect once they exit. The authors rule this out by showing that the fertility gap between female treatment and control groups persists well after the majority of scholarship recipients have graduated.&lt;/p&gt;
&lt;p&gt;Quality-Quantity Trade-off (Becker 1991): The economic framework predicting that more educated parents, facing higher opportunity costs of children and lower costs of investing in child quality, will have fewer but better-invested-in children. The authors find delayed and reduced fertility but do not find that recipients have fewer children to care for in the cognitive assessment sample, suggesting the child quality gains operate primarily through parenting practices rather than resource concentration.&lt;/p&gt;
&lt;p&gt;Intent-to-Treat (ITT) Effect: The treatment effect estimated by comparing all lottery winners to all losers regardless of whether winners actually enrolled, which captures the effect of the scholarship offer (including compliance costs). The cost-benefit analysis uses ITT estimates, so the cost of subsidizing inframarginal students who would have attended anyway is incorporated.&lt;/p&gt;
&lt;p&gt;Entropy Balancing: A reweighting procedure (Hainmueller, 2012) that assigns weights to observations in the control group so that the weighted distribution of baseline covariates matches that of the treatment group, used to assess whether imbalances in the subsample of participants who had children drive the results. The authors apply this as a robustness check for both mortality and cognitive development outcomes.&lt;/p&gt;
&lt;p&gt;Unwanted Pregnancy: A pregnancy reported by the respondent as unplanned at the time of conception, which the authors use to distinguish fertility reduction from a change in desired fertility versus a reduction in unintended out-of-wedlock pregnancies. The scholarship&amp;rsquo;s early fertility impact is almost entirely a reduction in unwanted pregnancies (7 percentage point decline, 17% reduction).&lt;/p&gt;</description></item><item><title>Investing in Influence: Investors, Portfolio Firms, and Political Giving</title><link>https://macropaperwarehouse.com/papers/investing-in-influence-investors-portfolio-firms-and-political-giving/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/investing-in-influence-investors-portfolio-firms-and-political-giving/</guid><description>&lt;p&gt;This paper investigates whether institutional investors influence the political activities of their portfolio firms, using political action committee (PAC) giving as a window into the broader question of whether institutional investors can leverage their concentrated ownership to extract benefits from portfolio firms for their own interests rather than those of their clients.&lt;/p&gt;
&lt;p&gt;The sample covers 574 institutional investors (those with at least $100 million in assets under management, i.e., 13-F filers) matched to 2,456 portfolio firms that had PACs, over the period 1980–2018. The primary source of variation is the first acquisition by an institutional investor of at least one percent of a portfolio firm&amp;rsquo;s outstanding shares, yielding 68,387 large acquisition events. PAC giving data come from FEC records matched by name to investor and firm entities. The main regression specification examines how the relationship between investor and firm PAC contributions to the same congressional district changes after such an acquisition, using a saturated set of fixed effects including firm × investor, firm × congressional district, firm × election cycle, investor × congressional district, investor × election cycle, and district × election cycle.&lt;/p&gt;
&lt;p&gt;The central finding is that, following a large block purchase, a firm&amp;rsquo;s PAC giving mirrors more closely that of the acquiring investment management company. In the preferred specification (column 8 of Table 2), the probability that a portfolio firm gives to a politician supported by its investor&amp;rsquo;s PAC increases by 31 percent after an acquisition. Using a cosine similarity measure of investor-firm PAC giving, the mean similarity of 0.10 at the acquisition cycle rises by 0.02–0.03 (a 20–30 percent increase) by the fourth post-acquisition election cycle.&lt;/p&gt;
&lt;p&gt;A key identification concern is that acquisitions may be driven by shared political preferences rather than representing a causal effect. To address this, the authors exploit stock index inclusions as exogenous shifters of institutional investor block purchases: when a firm is added to an index for the first time, passive indexers are compelled to rebalance toward that firm regardless of political alignment. Restricting to 5,601 index-inclusion acquisitions by passive investors, the authors find near-identical effect sizes (beta1 = 0.0132 in column 8 versus 0.0135 in the full sample), and an event study shows no pre-trend in giving convergence for the index subsample, in contrast to a slight pre-trend in the full sample. Divestment events exhibit the symmetric negative pattern: the interaction of post-divestment and investor PAC giving falls by between -0.074 and -0.058 across specifications.&lt;/p&gt;
&lt;p&gt;The authors argue that investors drive the convergence rather than portfolio firms adjusting investor preferences. Around acquisition dates, firms exhibit a larger drop in between-election-cycle cosine similarity than investors do. In a difference-in-differences comparison of the acquisition period relative to the preceding period, the difference in stability between investors and firms is 0.075 (significant at the 1 percent level), indicating that firms shift their giving more than investors. Investors obtaining a board seat at the portfolio firm amplifies the effect: in the preferred specification, the board-seat interaction is more than twice as large as the acquisition-alone interaction.&lt;/p&gt;
&lt;p&gt;Heterogeneity analysis provides evidence that the convergence reflects investors&amp;rsquo; partisan tastes rather than coordinated profit-maximizing political strategy. Acquisitions by more partisan investors (those whose giving is more skewed toward one party) produce a convergence coefficient roughly twice as large (0.020) as less partisan investors (0.010). Private fund families show more than twice the convergence effect of publicly owned fund families. The partisan composition of firm giving also shifts: a firm acquired by an investor giving exclusively to Republicans sees its Republican share increase by 2.8 percentage points relative to a baseline of 47.4 percent (a 5.9 percent increase).&lt;/p&gt;
&lt;p&gt;Finally, higher overall institutional ownership is associated with an increase in total PAC giving at the firm level, and this expanded giving does not go disproportionately to politicians on committees overseeing issues the firm actively lobbies — suggesting the ownership-driven increment in political spending is non-strategic from the firm&amp;rsquo;s profit standpoint and likely serves investors&amp;rsquo; own interests.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the central research question and why does it matter?&lt;/strong&gt;
The paper asks whether institutional investors influence the political giving of portfolio firms, motivated by the broader concern that the rise of institutional ownership — from 6 percent of U.S. public equities in 1950 to 65 percent in 2017 — concentrates not only economic but also political power in the hands of a small number of asset managers. This matters because if investors shape firms&amp;rsquo; PAC giving to serve investors&amp;rsquo; own preferences rather than firms&amp;rsquo; profit interests, it represents a misuse of corporate resources and a potential amplification of a small group&amp;rsquo;s political voice.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What data are used and how is the sample constructed?&lt;/strong&gt;
The analysis draws on 13-F filings (investors with at least $100M AUM) from Thomson-Reuters, matched to FEC PAC records via fuzzy and manual name matching. The resulting sample contains 574 investors with PACs and 2,456 portfolio firms with PACs, spanning 1980–2018. The Cartesian product of investor-firm pairs is restricted to those connected by at least one large acquisition event (defined as first acquisition of at least 1 percent of outstanding shares), yielding 68,387 such events. PAC contributions are measured at the investor- and firm-congressional-district-election-cycle level, linked to House of Representatives winners using MIT Election Data files.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the baseline regression and what does it find?&lt;/strong&gt;
The baseline regression (equation 1) interacts Log Investor PAC with a Post indicator (equal to 1 after the first large acquisition and while the stake is maintained) at the investor-firm-congressional-district-election-cycle level, with a saturated set of fixed effects. The coefficient on the interaction (beta1) is positive and highly significant (p &amp;lt; 0.001) across all eight specifications, ranging from 0.013 to 0.032. In the preferred specification, the increase in giving similarity is 31 percent relative to the pre-acquisition baseline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do the authors establish causality and rule out endogenous acquisitions?&lt;/strong&gt;
The primary identification strategy uses first-time inclusions of firms in stock indices (approximately 1,000 indices tracked in the sample) as exogenous shifters: passive indexers must rebalance toward the included firm regardless of political alignment. This subsample of 5,601 index-inclusion acquisitions produces near-identical coefficient estimates (0.0132 versus 0.0135 in the full sample), and the event study for this subsample shows no pre-trend in giving convergence, unlike the slight pre-trend in the full sample. Equality of the two coefficients cannot be rejected at standard significance levels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What evidence shows it is firms adjusting to investors rather than the reverse?&lt;/strong&gt;
The authors compute between-election-cycle cosine similarity separately for investors and firms around acquisitions. On average, investors exhibit more stable giving than firms at acquisition dates (Cos(xi,t, xi,t+1) &amp;gt; Cos(xf,t, xf,t+1)). The difference-in-differences estimate — comparing the acquisition period to the preceding period — is 0.075 (significant at 1 percent), indicating a relatively larger break in firm giving. Over a two-cycle window, the difference-in-differences estimate is 0.083, again indicating convergence is driven by firms shifting toward investors rather than the reverse.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What role does board representation play?&lt;/strong&gt;
In approximately 5 percent of acquisitions in the sample, the investor obtains a board seat. In specifications that include both the acquisition effect (Post × Log Investor PAC) and a board-membership interaction (Board × Log Investor PAC), both terms are positive and significant at the 1 percent level. In the preferred specification, the board-seat interaction is more than twice as large as the acquisition-alone interaction, indicating that a direct governance channel — board representation — substantially amplifies the convergence in political giving.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the divestment analysis show?&lt;/strong&gt;
Symmetric to the acquisition results, divestment events (where an investor exits a stake of at least 1 percent held for at least one election cycle) are associated with a decline in investor-firm PAC giving correlation. Post-divestment interaction coefficients range from -0.074 to -0.058 across specifications, and an event study confirms the correlation falls sharply after the divestment cycle.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does investor partisanship affect the magnitude of influence?&lt;/strong&gt;
Yes. Classifying investors as &amp;ldquo;More Partisan&amp;rdquo; (above-mean absolute deviation from 50/50 party split) versus &amp;ldquo;Less Partisan,&amp;rdquo; the interaction coefficient for More Partisan investors (0.020) is roughly twice that of Less Partisan investors (0.010). After a large acquisition by a fully Republican-giving investor, the acquired firm&amp;rsquo;s giving to that politician increases by 23.5 percent; the comparable figure for a Less Partisan investor is 7.6 percent. This pattern holds in both the full sample and the index-inclusion subsample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do private versus public fund families differ in their influence?&lt;/strong&gt;
Private fund families (e.g., Vanguard, Fidelity) show more than twice the convergence coefficient of publicly owned fund families (e.g., BlackRock, State Street, Invesco). The authors attribute this to private fund managers facing less outside scrutiny, allowing their giving to more readily reflect the preferences of owners and managers. Private investors also show greater partisan polarization: the 10th–90th percentile Republican-giving range for private investors is 6.3–100 percent, versus 21.7–88.3 percent for public investors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does increased institutional ownership expand overall firm PAC spending?&lt;/strong&gt;
Yes. In firm-year level regressions, institutional ownership is a positive and significant predictor of total firm PAC giving (significant at at least the 5 percent level in both cross-sectional and firm-fixed-effects specifications). Total corporate political expenditure by sample firms increased by nearly a factor of six over 1980–2018. The authors note that while many factors contribute, increased institutional ownership may be at least partly responsible for this expansion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does the additional giving driven by institutional ownership go to strategically important politicians for the firm?&lt;/strong&gt;
No. Regressions relating institutional ownership to giving to politicians on congressional committees overseeing issues the firm actively lobbies (a standard measure of politicians&amp;rsquo; strategic importance to firms) yield near-zero and statistically weak point estimates. In the preferred firm-fixed-effects specification, the share of total PAC giving devoted to such strategically relevant politicians is negatively associated with institutional ownership at marginal significance (p &amp;lt; 0.10), consistent with the interpretation that ownership-driven incremental political spending is non-strategic from the firm&amp;rsquo;s own profit perspective and expands total giving rather than displacing strategic giving.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the policy and legal implications?&lt;/strong&gt;
The authors flag three concerns: (i) the ownership-driven increment in political spending may represent a misuse of corporate resources that does not serve portfolio firm shareholders; (ii) it may constitute an illegal activity, since using a firm&amp;rsquo;s PAC to reimburse or proxy for an investor&amp;rsquo;s own political preferences can run afoul of campaign finance law; and (iii) it is a channel through which unequal resources amplify the political voice of a small number of fund managers at the expense of dispersed ultimate investors who are likely unaware of and do not sanction these contributions. The findings challenge the Supreme Court&amp;rsquo;s premise in Citizens United that corporate political speech reflects shareholder profit maximization.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PAC comovement (investor-firm giving similarity):&lt;/strong&gt; The increase in the probability that a portfolio firm&amp;rsquo;s PAC donates to a politician also supported by an acquiring investor&amp;rsquo;s PAC, measured as the interaction coefficient between Log Investor PAC and a Post-acquisition indicator in the baseline regression. In the preferred specification this represents a 31 percent increase relative to the pre-acquisition baseline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cosine similarity (cross-time and cross-entity):&lt;/strong&gt; A measure defined as the Euclidean dot product between two vectors of PAC giving (either the same entity across adjacent election cycles, or investor versus firm in the same cycle), taking values between 0 and 1, where 1 indicates identical giving patterns. Used both to confirm convergence post-acquisition and to attribute that convergence to firm rather than investor adjustment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Index-inclusion acquisition:&lt;/strong&gt; A large block purchase that results from a firm being added for the first time to a stock index tracked by a passive institutional investor, used as an exogenous shifter of investor stakes that is orthogonal to investor-firm political alignment. There are 5,601 such events in the sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Partisanship (investor):&lt;/strong&gt; Classified as &amp;ldquo;More Partisan&amp;rdquo; if an investor&amp;rsquo;s absolute deviation from a 50/50 party split in PAC donations is above the sample mean. More partisan investors produce roughly twice the convergence effect on portfolio firm giving compared to less partisan investors, used as evidence that personal political preferences rather than profit-maximizing business strategy drive the convergence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Post indicator (Postift):&lt;/strong&gt; A binary variable equal to 1 for all election cycles following an investor&amp;rsquo;s first acquisition of at least 1 percent of a portfolio firm&amp;rsquo;s outstanding shares, and remaining 1 as long as the investor holds any stake in the firm. The key source of temporal variation in the baseline regression.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Strategically important politicians:&lt;/strong&gt; Members of Congress sitting on committees that oversee issues on which a firm actively lobbies, identified by crosswalking lobbying reports from the Senate Office of Public Records to relevant committee jurisdictions. Used to test whether ownership-driven political giving displaces or supplements firm-profit-motivated giving.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Board seat channel:&lt;/strong&gt; The mechanism through which investor influence on firm political giving is amplified when the investor obtains representation on the portfolio firm&amp;rsquo;s board of directors (present in approximately 5 percent of acquisitions). The board interaction coefficient is more than twice the acquisition-alone coefficient in the preferred specification.&lt;/p&gt;</description></item><item><title>Labor Market Competition and the Assimilation of Immigrants</title><link>https://macropaperwarehouse.com/papers/labor-market-competition-and-the-assimilation-of-immigrants/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/labor-market-competition-and-the-assimilation-of-immigrants/</guid><description>&lt;h2 id="labor-market-competition-and-the-assimilation-of-immigrants"&gt;Labor Market Competition and the Assimilation of Immigrants&lt;/h2&gt;
&lt;h3 id="research-question"&gt;Research Question&lt;/h3&gt;
&lt;p&gt;Why have immigrant-native wage gaps widened substantially across arrival cohorts in the United States since the 1960s, and why has the speed of wage convergence slowed? The paper argues that the existing literature, which attributes these trends entirely to declining immigrant cohort quality, omits a critical general-equilibrium channel: labor market competition arising from imperfect substitutability between immigrants and natives. The paper quantifies how much of the observed deterioration in wage assimilation profiles can be attributed to (i) increasing immigrant cohort sizes raising labor market competition, (ii) secular shifts in relative skill demand, and (iii) genuine changes in immigrant cohort quality.&lt;/p&gt;
&lt;h3 id="data-and-methodology"&gt;Data and Methodology&lt;/h3&gt;
&lt;p&gt;The analysis uses U.S. Census microdata for 1970, 1980, 1990, and 2000, combined with American Community Survey (ACS) data pooled for 2009–2011 (labeled 2010) and 2018–2019 (labeled 2020), all drawn from IPUMS-USA. The sample covers individuals aged 25–64 who are employed in the civilian sector, not self-employed, not in group quarters, and report positive earnings. Immigrant cohort sizes grew from approximately 800,000 individuals in the 1960s cohort to 2.3 million in the 1980s cohort and 4.6 million in the 2000s cohort.&lt;/p&gt;
&lt;p&gt;The theoretical framework is a constant elasticity of substitution (CES) production function in which workers supply two types of skills: &amp;ldquo;general&amp;rdquo; skills portable across countries and &amp;ldquo;specific&amp;rdquo; skills particular to the host country (including language proficiency and knowledge of cultural and institutional environment). Immigrants arrive with the same general skills as observationally equivalent natives but only a fraction of their specific skills; they accumulate specific skills over time. Because immigrants disproportionately supply general skills upon arrival, increasing immigrant inflows raise the relative supply of general skills, depress the relative price of general skills, and thereby widen the immigrant-native wage gap. This mechanism operates only when immigrants and natives are imperfect substitutes (elasticity of substitution σ &amp;lt; ∞).&lt;/p&gt;
&lt;p&gt;The model is estimated in two steps using nonlinear least squares (NLS). First, productivity factor parameters are estimated from native wages year by year, with state dummies identifying state-level skill prices. Second, specific skill accumulation parameters and the elasticity of substitution σ are jointly identified from immigrant wage differences across labor markets (defined as U.S. states) and over time. The demand shift parameter δ_t, which captures changes in the relative demand for specific skills (e.g., technology that favors communication over manual tasks), enters as a linear time trend in the baseline specification.&lt;/p&gt;
&lt;h3 id="main-findings-with-quantitative-magnitudes"&gt;Main Findings with Quantitative Magnitudes&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Competition effect:&lt;/strong&gt; Immigration-induced increases in labor market competition explain 14.2, 43.9, and 40.8 percent of the increase in the initial wage gap of the 1970s, 1980s, and 1990s cohorts relative to the 1960s cohort, respectively. Averaged across all years spent in the United States, the competition effect alone accounts for 14.1, 22.4, and 20.4 percent — approximately one fifth overall.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Competition plus demand effect:&lt;/strong&gt; Adding secular shifts in relative skill demand raises these figures to 24.8, 68.3, and 109.5 percent at arrival and 21.2, 33.6, and 36.4 percent averaged across years — approximately one third overall.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Elasticity of substitution:&lt;/strong&gt; The baseline estimate of σ (elasticity of substitution between general and specific skills) is 0.020 (s.e. 0.002), implying an inverse elasticity of approximately 50.5. The relative supply of general skills increased by 1.67 log points between 1970 and 2020, producing a predicted increase in the relative price of specific skills of approximately 59.6 log points. The demand shift trend is estimated at 1.3 log points per year.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cohort quality:&lt;/strong&gt; Once competition and demand effects are netted out, the remaining deterioration in assimilation profiles is entirely attributable to observable changes in immigrants&amp;rsquo; educational attainment and country-of-origin composition. Conditional on these two observable characteristics, unobservable skill quality improved across cohorts (consistent with English language proficiency trends), reversing the conventional narrative of declining cohort quality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Specific skills gap at arrival:&lt;/strong&gt; The 1960s cohort faced a specific skills gap of approximately 52.4 percent relative to native equivalents; this narrowed to 41.8 percent for the 1970s cohort, 35.6 percent for the 1980s cohort, and 17.6 percent for the 1990s cohort, conditional on origin and education. After 20–30 years, all cohorts reach 83.7–92.0 percent of their native counterparts&amp;rsquo; specific skill levels.&lt;/p&gt;
&lt;h3 id="scope-conditions"&gt;Scope Conditions&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;The analysis focuses on employed men in the main text (women are analyzed in an Online Appendix, showing qualitatively similar but quantitatively smaller patterns).&lt;/li&gt;
&lt;li&gt;Labor markets are defined at the U.S. state level in the baseline; robustness checks use state-education and state-gender cells.&lt;/li&gt;
&lt;li&gt;The decomposition covers the period from the 1960s to the 1990s arrival cohorts.&lt;/li&gt;
&lt;li&gt;Results are robust to corrections for selective outmigration, undercounting of undocumented immigrants, immigrant network effects, alternative demand shift specifications, alternative labor market definitions, and endogenous immigrant location choice (using shift-share instruments in the spirit of Card, 2001).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-theoretical-mechanism-by-which-increasing-immigrant-inflows-widen-the-immigrant-native-wage-gap"&gt;Q1. What is the core theoretical mechanism by which increasing immigrant inflows widen the immigrant-native wage gap?&lt;/h3&gt;
&lt;p&gt;A: Because immigrants disproportionately supply general (country-portable) skills upon arrival, while natives disproportionately supply specific (host-country) skills, an increase in immigrant inflows raises the ratio of general to specific skills in the economy. Under imperfect substitutability (σ &amp;lt; ∞), this lowers the relative price of general skills and raises the relative price of specific skills, thereby widening the wage gap between immigrants (who earn predominantly from general skills) and natives (who earn more from specific skills). The effect is larger in the early years after arrival when immigrants&amp;rsquo; specific skill endowment s is small, and diminishes as immigrants accumulate specific skills over time.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-model-immigrants-skill-accumulation-and-how-do-accumulation-profiles-differ-across-groups"&gt;Q2. How does the paper model immigrants&amp;rsquo; skill accumulation, and how do accumulation profiles differ across groups?&lt;/h3&gt;
&lt;p&gt;A: Immigrants&amp;rsquo; specific skill endowment s(·) upon arrival and over time is modeled as a flexible polynomial in years since migration, interacted with dummies for region of origin, education, cohort of entry, and potential experience abroad. Mexican high school dropouts (the reference group) are estimated to arrive with approximately 80 percent of the specific skills of equivalent natives. Immigrants from Latin America, Asia, and other regions arrive with lower specific skills than Western immigrants, who arrive near native parity. Higher-educated immigrants arrive relatively less similar to equivalently educated natives than low-educated immigrants, reflecting the greater importance of language-intensive skills in high-skill occupations. Conditional on origin and education, more recent cohorts arrive with narrower specific skill deficits: the 1990s cohort faces a gap of 17.6 percent at arrival compared to 52.4 percent for the 1960s cohort.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-estimated-technology-parameters-and-how-are-they-interpreted"&gt;Q3. What are the estimated technology parameters, and how are they interpreted?&lt;/h3&gt;
&lt;p&gt;A: The elasticity of substitution between general and specific skills is estimated at σ = 0.020 (s.e. 0.002), with a confidence interval of [0.017, 0.024]. This implies an inverse elasticity of approximately 50.5, meaning a one percent increase in the relative supply of general skills raises the relative price of specific skills by about 50.5 percent. The implied elasticity of substitution between natives and immigrants (evaluated at market-level averages) is approximately 0.013 in 1990, 0.020 in 2000, and 0.025 in 2010 — in the same range as the Ottaviano and Peri (2012) benchmark of 0.034 (s.e. 0.008). The demand shift trend is estimated at δ̃ = 0.013 (s.e. 0.001) log points per year, reflecting secular increases in the relative demand for specific (host-country) skills.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-paper-identify-the-elasticity-of-substitution-σ-and-the-skill-accumulation-parameters-separately"&gt;Q4. How does the paper identify the elasticity of substitution σ and the skill accumulation parameters separately?&lt;/h3&gt;
&lt;p&gt;A: The estimation proceeds in two steps. First, productivity factor parameters (returns to education and experience) are estimated from native wage regressions, with state-year dummies absorbing state-specific skill prices. Second, skill accumulation parameters θ are identified from wage differences between immigrants with different characteristics working in the same labor market, while σ and the demand shift δ̃ are identified from variation in immigrant wage gaps across states (which have different immigrant population shares) and over time. Specifically, states with higher immigrant shares display lower relative prices of general skills, providing the identifying variation for σ.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-quantitative-magnitudes-of-the-competition-effect-for-specific-cohorts-at-different-time-horizons"&gt;Q5. What are the quantitative magnitudes of the competition effect for specific cohorts at different time horizons?&lt;/h3&gt;
&lt;p&gt;A: At the time of arrival, the competition effect explains 14.2 percent (1970s cohort), 43.9 percent (1980s cohort), and 40.8 percent (1990s cohort) of the increase in initial wage gaps relative to the 1960s cohort. After 10 years, these figures are 17.1, 22.7, and 22.2 percent respectively. After 20 years, they are 12.2, 16.9, and 16.2 percent. After 30 years, 10.9, 15.3, and 13.7 percent. The declining share across years reflects the fact that as immigrants accumulate specific skills, their wages become less sensitive to equilibrium skill prices. Averaged across all years since migration, the competition effect accounts for 14.1, 22.4, and 20.4 percent for the three cohorts.&lt;/p&gt;
&lt;h3 id="q6-how-does-labor-market-competition-affect-the-speed-of-wage-assimilation-and-does-it-prevent-full-convergence"&gt;Q6. How does labor market competition affect the speed of wage assimilation, and does it prevent full convergence?&lt;/h3&gt;
&lt;p&gt;A: The effect on assimilation speed is theoretically ambiguous and depends on whether future cohorts are larger or smaller than the reference cohort, and whether immigrants fully converge to native skill levels. In the stylized examples, a one-time permanent increase in competition raises both the initial wage gap and the speed of subsequent convergence (since the gap between immigrant and native skill levels is larger and therefore more responsive to changes in skill prices). However, continuous inflows of increasingly large cohorts counteract this speedup by continuously shifting the wage profile downward — the &amp;ldquo;dynamic competition effect.&amp;rdquo; For immigrants who fully converge (s → 1), competition delays but does not prevent convergence; for those who only partially converge (s → &amp;lt; 1), competition permanently widens the long-run wage gap. Quantitatively, the paper finds the effect on assimilation speed to be small in the full-sample decomposition.&lt;/p&gt;
&lt;h3 id="q7-what-do-the-illustrative-examples-for-specific-immigrant-groups-reveal-about-heterogeneous-competition-effects"&gt;Q7. What do the illustrative examples for specific immigrant groups reveal about heterogeneous competition effects?&lt;/h3&gt;
&lt;p&gt;A: For a Mexican male high school dropout (1960s cohort skills), facing the same competition level as the 1990s cohort would widen the initial wage gap by 10.2 log points; facing 2010 competition levels would widen it by 21.1 log points. However, because this group fully converges (s → 1), the effect dissipates entirely after approximately 25 years, and long-run wage assimilation is not prevented. For a Latin American male high school graduate who only partially converges (s → &amp;lt; 1), facing 1990s competition would widen the initial gap by 17.4 log points and leave a 3.8 log-point larger long-run wage gap. For a Western college graduate who arrives near native skill parity, competition effects are negligible throughout.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-changes-in-absolute-wage-gaps-documented-in-the-baseline-data"&gt;Q8. What are the changes in absolute wage gaps documented in the baseline data?&lt;/h3&gt;
&lt;p&gt;A: The 1960s cohort arrived with an initial wage gap of approximately 17.2 log points relative to natives. The 1970s cohort arrived with a gap of 30.1 log points, the 1980s cohort 29.2 log points, and the 1990s cohort 20.8 log points. Under the no-competition counterfactual, these initial gaps narrow to 13.6, 24.7, 20.3, and 15.7 log points respectively. Removing both competition and demand effects further narrows them to 13.7, 23.4, 17.5, and 13.3 log points.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-paper-find-about-the-role-of-observable-versus-unobservable-immigrant-quality"&gt;Q9. What does the paper find about the role of observable versus unobservable immigrant quality?&lt;/h3&gt;
&lt;p&gt;A: Once competition and demand effects are accounted for, all remaining cohort differences in assimilation profiles are attributable to observable changes in immigrants&amp;rsquo; educational attainment and country-of-origin composition. Conditional on these two observable characteristics, immigrants in more recent cohorts display higher levels of unobservable skills (smaller specific skill deficits conditional on origin and education), consistent with rising English language proficiency across cohorts. This reverses the standard interpretation that unobservable immigrant quality has declined.&lt;/p&gt;
&lt;h3 id="q10-how-do-aggregate-skill-supplies-and-relative-skill-prices-evolve-over-the-sample-period"&gt;Q10. How do aggregate skill supplies and relative skill prices evolve over the sample period?&lt;/h3&gt;
&lt;p&gt;A: Between 1970 and 2020, the total supply of general skills from immigrants grew by a factor of 16.3, while the supply of specific skills grew by a factor of 15.0. The resulting increase in the relative supply of general skills caused the relative price of general skills to fall from 0.89 to 0.38. Accounting for growing relative demand for specific skills (the δ_t trend), the ratio of relative skill prices fell further to 0.20 by 2020. At the state level, relative prices of general skills are well below 0.3 in high-immigration states like California, Florida, and New York, and approach 1.0 in states with low immigrant shares.&lt;/p&gt;
&lt;h3 id="q11-are-the-results-robust-to-selective-outmigration-undocumented-immigrants-and-alternative-specifications"&gt;Q11. Are the results robust to selective outmigration, undocumented immigrants, and alternative specifications?&lt;/h3&gt;
&lt;p&gt;A: Yes. Across twelve robustness checks covering selective outmigration corrections (using Borjas and Bratsberg 1996 or Rho and Sanders 2021 outmigration rates, and synthetic cohort reweighting), undocumented immigrant undercounting corrections, immigrant network controls (share and stock of compatriots in the same state), alternative demand shift specifications (quadratic and time dummies), alternative labor market definitions (state-education and state-gender cells), and endogenous immigrant location choice (GMM with shift-share instruments), the estimated elasticity of substitution σ ranges from 0.017 to 0.033 and the average competition effects remain stable. Averaged across all robustness checks, competition effects are 1.3 log points (1960s cohort), 3.0 log points (1970s), 5.2 log points (1980s), and 4.3 log points (1990s), compared to baseline values of 1.4, 3.1, 5.5, and 4.6 log points.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-highlighted-by-the-authors"&gt;Q12. What are the policy implications highlighted by the authors?&lt;/h3&gt;
&lt;p&gt;A: First, since assimilation and competition effects are intertwined, the wage impact of immigration on natives is intrinsically dynamic: newly arrived immigrants initially compete relatively little with natives but increasingly substitute for them as their specific skills grow. Second, labor market competition may reduce immigrants&amp;rsquo; incentives to invest in host-country-specific skills, a channel not modeled in most existing structural models. Third, dispersal policies (such as those used during refugee crises) that reallocate immigrants across regions will affect local skill price ratios and therefore alter wage assimilation trajectories — a potentially unintended consequence of geographic allocation policies.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;General skills:&lt;/strong&gt; Skills that are portable across countries and can be used productively in any labor market. In the paper&amp;rsquo;s framework, general skills are those required for tasks (such as manual or physical labor) that are similar across national contexts. Upon arrival, immigrants are assumed to supply the same amount of general skills as observationally equivalent natives, making immigrants&amp;rsquo; relative supply of general skills high at arrival.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Specific skills (host-country-specific skills):&lt;/strong&gt; Skills particular to the host country, including language proficiency (English in the U.S. context) as well as familiarity with the institutional and cultural environment. Immigrants arrive with only a fraction s of the specific skills of comparable natives; this fraction evolves over time as immigrants spend time in the host country. The level of specific skills governs how substitutable a given immigrant worker is with native workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor market competition effect:&lt;/strong&gt; The mechanism by which increasing immigrant inflows affect relative wages through equilibrium changes in skill prices rather than through individual skill accumulation. When immigrants and natives are imperfect substitutes, rising immigrant inflows raise the relative supply of general skills, depress the relative price of general skills, and widen the immigrant-native wage gap. This effect is larger for recently arrived immigrants (small s) and diminishes as immigrants assimilate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic competition effect:&lt;/strong&gt; The combined effect on a given cohort&amp;rsquo;s observed assimilation profile of continuous, growing immigrant inflows over its time in the country. Unlike a one-time permanent increase in competition (which would raise both the initial gap and assimilation speed), continuously growing inflows both widen the initial gap and exert a continuous downward shift on the cohort&amp;rsquo;s wage profile, with an ambiguous net effect on the speed of convergence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Demand shift (δ_t):&lt;/strong&gt; A time-varying parameter in the CES production function capturing secular changes in the relative demand for specific versus general skills beyond what is explained by standard skill-biased technological change. A positive trend in δ_t (estimated at 1.3 log points per year in the baseline) reflects technological change that favors communication-intensive (specific-skill-intensive) tasks over manual (general-skill-intensive) tasks, and amplifies the competition effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Elasticity of substitution between general and specific skills (σ):&lt;/strong&gt; The key technology parameter governing the degree of imperfect substitutability between natives and immigrants in equilibrium. Estimated at σ = 0.020 in the baseline. When σ = ∞, immigrants and natives are perfect substitutes and labor market competition has no effect on relative wages. As σ decreases, the competition effect on relative wages becomes stronger for a given change in relative skill supplies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Specific skill accumulation function s(·):&lt;/strong&gt; A flexible parametric function of years since migration, interacted with region of origin, education level, cohort of entry, and potential experience at arrival, that governs the rate at which immigrants acquire host-country-specific skills over time. The intercept of s(·) at arrival (relative to a native s = 1) measures the initial specific skill deficit; the polynomial in years since migration captures how quickly this deficit closes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wage assimilation profile:&lt;/strong&gt; The trajectory of the immigrant-native log wage gap as a function of years spent in the host country, conditional on a cohort of arrival. The paper distinguishes between changes in the level of the profile (the initial wage gap) and changes in its slope (the speed of convergence), and decomposes both dimensions into competition effects, demand effects, and cohort quality effects.&lt;/p&gt;</description></item><item><title>Latent Heterogeneity in the Marginal Propensity to Consume</title><link>https://macropaperwarehouse.com/papers/latent-heterogeneity-in-the-marginal-propensity-to-consume/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/latent-heterogeneity-in-the-marginal-propensity-to-consume/</guid><description>&lt;p&gt;Lewis, Melcangi, and Pilossoph estimate the unconditional distribution of the marginal propensity to consume (MPC) using the 2008 Economic Stimulus Act (ESA) rebate payments, deploying Gaussian mixture linear regression (GMLR) — a clustering regression approach — rather than the standard practice of interacting the rebate with observable household characteristics. The key methodological departure is that households are assigned to groups not by any presupposed observable, but by how well estimated group-specific MPCs describe each household&amp;rsquo;s actual consumption response; this allows recovery of the full unconditional MPC distribution, including heterogeneity driven by latent (unobservable) factors.&lt;/p&gt;
&lt;p&gt;Data come from the 2008 Consumer Expenditure Survey (CEX), which contains household-level expenditure data and supplemental questions on ESA payments. Identification exploits the quasi-random timing of rebate receipt, determined by the last two digits of recipients&amp;rsquo; Social Security Numbers, following the design of Parker, Souleles, Johnson, and McClelland (2013). The specification is updated following Borusyak et al. (2024) to avoid &amp;ldquo;forbidden comparisons&amp;rdquo; in staggered treatment settings. The number of groups G is selected by BIC, which selects G = 3 for total expenditures, confirmed by K-fold cross-validation.&lt;/p&gt;
&lt;p&gt;The main finding is substantial MPC heterogeneity. For total expenditures, the three estimated group-level MPCs are 0.04, 0.23, and 1.33, with population shares of 30%, 48%, and 23% respectively. The implied aggregate (share-weighted average) MPC is 0.42, compared to 0.24 in the homogeneous Parker et al. (2013) specification estimated on the same data. Splitting by consumption category: for nondurables, two groups have MPCs of 0.09 and 0.18, with roughly equal population shares, and the lower bound of 0.09 is statistically distinguishable from zero — evidence against strict adherence to the Permanent Income Hypothesis even among the lowest-MPC group. For durables, the MPC distribution is dichotomous: about 29% of households have a durable MPC statistically indistinguishable from zero, while 21% have an MPC of 0.67. The cross-good correlation between household-level nondurable and durable predicted MPCs is only 0.13, ruling out strong substitution but indicating weak complementarity.&lt;/p&gt;
&lt;p&gt;Turning to observable determinants, the paper finds that many household characteristics are individually correlated with estimated MPCs — including homeownership, mortgage status, income, and the average propensity to consume (APC) — despite the fact that the same dataset and similar identification strategies previously yielded insignificant relationships. Homeowners have significantly higher MPCs than renters; households with a mortgage have even higher MPCs than outright homeowners. In salary income, households in the top tercile spend 0.17 more per rebate dollar than the baseline group; households in the top tercile of non-salary income spend 0.19 more. However, in joint regressions, only two characteristics remain robustly and positively correlated with MPCs: total income (both salary and non-salary components) and the APC. The APC relationship is particularly notable: a one-percentage-point higher prior spending rate is associated with 0.19 additional cents spent per rebate dollar in the full multivariate specification.&lt;/p&gt;
&lt;p&gt;The paper identifies three groups in the joint income-APC space: &amp;ldquo;poor savers&amp;rdquo; (low income, low APC, lowest MPCs), an intermediate group (high income or high APC but not both), and &amp;ldquo;rich spenders&amp;rdquo; (high income and high APC, highest MPCs). The &amp;ldquo;rich spender&amp;rdquo; group has received little prior attention in consumption-savings models.&lt;/p&gt;
&lt;p&gt;Critically, observable characteristics jointly explain at most 8% of MPC variation (adjusted R-squared from a measurement-error correction). With 92% of MPC heterogeneity unexplained by standard observables, the authors conclude that a substantial share of variation reflects latent household traits — plausibly heterogeneity in discount rates or intertemporal elasticities of substitution. This finding also limits the practical scope for government targeting of fiscal transfers: because observable characteristics predict little MPC variation, any targeting strategy can exploit only a small fraction of the overall distribution.&lt;/p&gt;
&lt;p&gt;Scope conditions: results apply to household expenditure responses (marginal propensities to spend, not to consume in the strict sense) within one quarter of rebate receipt. The income-MPC positive correlation is confined to households within the income range eligible for the 2008 ESA (phased out above $150,000 for joint filers). The sample excludes the top and bottom 1.5% of consumption changes as outliers.&lt;/p&gt;
&lt;p&gt;Q: What is the core methodological innovation of this paper?
A: The paper applies Gaussian mixture linear regression (GMLR) to the 2008 tax rebate setting, jointly estimating group-level MPCs and household group membership probabilities without imposing any prior restriction on which observable characteristics drive heterogeneity. Because groups are determined by how well group-specific MPCs explain consumption patterns rather than by presupposed observables, the method recovers the full unconditional distribution of MPCs, including latent heterogeneity. This contrasts with sample-splitting approaches that can only recover co-variation with chosen characteristics.&lt;/p&gt;
&lt;p&gt;Q: What are the three group-level MPCs for total expenditures, and what shares of the population do they represent?
A: The three estimated MPCs are 0.04 (30% of households), 0.23 (48%), and 1.33 (23%), all with precisely estimated group shares (standard errors of 0.01). The largest MPC of 1.33 is statistically significant at the 1% level. The lowest MPC of 0.04 is not statistically different from zero even under the more favorable conditional standard errors that treat group assignment as known.&lt;/p&gt;
&lt;p&gt;Q: How does the average MPC implied by the GMLR distribution compare to the homogeneous specification?
A: The share-weighted average MPC from the three-group GMLR is 0.42, compared to 0.24 from the homogeneous (G=1) specification on the same data and identification strategy. This gap arises partly because the homogeneous estimate averages across households with very heterogeneous responses, and partly because the distribution has a right-skewed tail with a meaningful mass at MPC above 1.&lt;/p&gt;
&lt;p&gt;Q: What are the MPC distributions for nondurable and durable goods separately?
A: For nondurables, BIC selects two groups with MPCs of 0.09 and 0.18 and roughly equal population shares (48% and 52%); crucially, the lower bound of 0.09 is statistically distinguishable from zero at the 5% level, providing evidence that no household strictly follows the Permanent Income Hypothesis for nondurables. For durables, BIC selects three groups: MPCs of 0.03 (not distinguishable from zero, 29% of households), 0.15 (50%), and 0.67 (21%), reflecting the discrete, lumpy nature of durable goods purchases.&lt;/p&gt;
&lt;p&gt;Q: How correlated are nondurable and durable MPCs at the household level?
A: The correlation between household-level posterior predicted MPCs for nondurables and durables is 0.13, statistically significant at the 1% level. This rules out substitution between goods categories, but the positive complementarity is quantitatively small. The authors interpret this as possibly reflecting a small share of &amp;ldquo;spender&amp;rdquo; types who adjust multiple consumption categories in response to transitory income shocks.&lt;/p&gt;
&lt;p&gt;Q: Which observable characteristics are individually correlated with MPCs?
A: Homeowners have significantly higher MPCs than renters; households with a mortgage display even greater MPCs than outright homeowners. Both salary and non-salary income are positively correlated: households in the top tercile of salary income have MPCs about 0.13 higher than the omitted group, and top-tercile non-salary income households have MPCs about 0.015 higher (though the latter is individually less precisely estimated). The average propensity to consume (APC) is significantly positively correlated with the MPC, with a coefficient of 0.075 in univariate regression and 0.166 in the full joint specification.&lt;/p&gt;
&lt;p&gt;Q: Which observable characteristics remain significant in the joint (multivariate) regression?
A: When all household characteristics are included jointly, only income (both salary and non-salary components) and the APC remain robustly and positively correlated with MPCs. Top-tercile salary income is associated with 0.112 higher MPCs and top-tercile non-salary income with 0.049 higher MPCs, while the APC coefficient rises to 0.166 (from 0.075 univariate). Homeownership, age, education, and most demographic controls become statistically insignificant in the joint specification.&lt;/p&gt;
&lt;p&gt;Q: What fraction of MPC variation is explained by observable characteristics?
A: The adjusted R-squared from the full multivariate regression of predicted MPCs on all observable characteristics is approximately 6%. After a measurement-error correction proposed in Supplement A.6 to account for noise in estimated posterior MPCs, the corrected R-squared rises to 8%. Either way, the vast majority — over 90% — of MPC heterogeneity is unexplained by standard observables, implicating latent household traits such as heterogeneous discount rates or intertemporal elasticities of substitution.&lt;/p&gt;
&lt;p&gt;Q: How does the extent of MPC heterogeneity recovered by GMLR compare to sample-splitting on observables?
A: Table 4 shows that splitting by age terciles yields MPC estimates ranging from 0.13 to 0.34; splitting by total income yields a range of 0.18 to 0.45; splitting by the APC yields 0.06 to 0.21. All of these ranges are far narrower than the GMLR-recovered range of 0.04 to 1.33. The authors argue that sample-splitting on individual observables, which are noisy and correlated with only a portion of MPC heterogeneity, systematically understates the true extent of heterogeneity.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;rich spender&amp;rdquo; finding and why is it theoretically notable?
A: Households with both high total income and a high prior average propensity to consume have the largest MPCs. This &amp;ldquo;rich spender&amp;rdquo; group is poorly accommodated by standard consumption-savings models: the canonical one-asset incomplete markets model typically predicts a negative MPC-APC correlation conditional on income, and the two-asset Kaplan-Violante (2014) model can generate wealthy hand-to-mouth households with high income and high MPCs, but not necessarily high APCs. Preference heterogeneity — e.g., heterogeneous intertemporal elasticities of substitution as in Aguiar, Boar, and Bils (2019) — can rationalize the positive income-APC-MPC nexus.&lt;/p&gt;
&lt;p&gt;Q: What explains the positive income-MPC correlation, and how does the paper relate it to the prior literature?
A: The paper notes that this positive correlation is consistent with Kueng (2018), who finds higher spending propensities among high-income recipients of Alaska Permanent Fund payments, and rationalizes it via near-rationality or mental accounting: when a rebate is small relative to income, the perceived cost of deviating from consumption smoothing is low. The authors also note that low-income households still exhibit large absolute MPCs, suggesting sizable deviations from consumption smoothing at the bottom of the income distribution, even if relatively lower than for high-income households.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications for targeting fiscal transfers?
A: The paper finds that the 2008 ESA increased spending for all households in partial equilibrium (minimum group MPC of 0.04, nondurable lower bound 0.09, all statistically positive or near-positive). Among observable characteristics, targeting relatively higher-income households (including retirees and entrepreneurs via non-salary income) would maximize aggregate consumption effects. However, since observables explain only 8% of MPC variation, any targeting strategy can exploit only a small fraction of the overall heterogeneity; the government faces fundamental limits on feasible targeting. This also implies a tension between stimulus and distributional/insurance motives for transfer programs.&lt;/p&gt;
&lt;p&gt;Q: How does the paper confirm that recovered heterogeneity is not spurious?
A: The authors generate 250 Monte Carlo samples from the estimated homogeneous model, impose G=3, and re-run the GMLR and observable regressions; they find significant relationships with observable characteristics in virtually none of these samples. Additionally, applying the BIC to homogeneous Monte Carlo samples, the BIC selects G=1 in all 250 samples, confirming that the selected G=3 in actual data reflects genuine heterogeneity rather than overfitting.&lt;/p&gt;
&lt;p&gt;Q: How does GMLR compare to quantile regression for recovering the MPC distribution?
A: Quantile regression (as used by Misra and Surico (2014) on the same data) recovers relationships at percentiles of the overall conditional distribution of consumption changes, so the ranking of households is driven by all sources of variation in consumption, not just the rebate response. If factors unrelated to the rebate dominate the conditional distribution, MPC heterogeneity will be underestimated in the presence of noise. The authors illustrate this formally in Supplement B and note that Misra and Surico (2014) find a substantial share of MPCs at or below zero for nondurables, in contrast to the GMLR lower bound of 0.09 that is statistically positive.&lt;/p&gt;
&lt;p&gt;Q: What do the longer-run (lagged) MPC estimates show?
A: The specification includes up to two lags of rebate indicators, allowing measurement of spending responses in subsequent quarters after rebate receipt. The paper reports these results (Section 4.4) but the text provided does not fully detail them; the heterogeneous structure is maintained across horizons.&lt;/p&gt;
&lt;p&gt;Gaussian Mixture Linear Regression (GMLR): A probabilistic clustering regression approach that jointly estimates group-specific regression coefficients (here, MPCs) and population group shares by maximizing an expected log-likelihood via the EM algorithm. Households receive continuous posterior weights (gamma_{jg}) reflecting uncertainty about their group membership rather than binary hard assignment, with identification from a Gaussianity assumption on within-group errors.&lt;/p&gt;
&lt;p&gt;Unconditional MPC Distribution: The full marginal distribution of MPCs across all households in the population, capturing heterogeneity from both observable and latent (unobservable) sources. Contrasted in the paper with the conditional distributions recovered by sample-splitting on observables, which by construction can only reflect co-variation with the chosen splitting variable.&lt;/p&gt;
&lt;p&gt;Posterior Predicted MPC: For each household, the expectation of the group-specific MPC weighted by the household&amp;rsquo;s posterior group membership probabilities (lambda-tilde_{0,j} = sum_g gamma_{jg} lambda_{0g}). This object is the optimal (MSE-minimizing) individual-level MPC prediction and is the relevant input for targeted fiscal policy design.&lt;/p&gt;
&lt;p&gt;Latent Heterogeneity: MPC variation that cannot be attributed to any observable household characteristic and is instead driven by unobserved traits — plausibly heterogeneous discount rates, intertemporal elasticities of substitution, or other preference parameters. Operationalized as the share of MPC variance unexplained by observable regressors (approximately 92% in this paper).&lt;/p&gt;
&lt;p&gt;Rich Spenders: A group identified jointly in the APC-income space: households with both high total income and a high average propensity to consume, displaying the largest marginal propensities to consume out of the rebate. This group is not well-accommodated by standard one-asset or two-asset incomplete markets models under homogeneous preferences.&lt;/p&gt;
&lt;p&gt;Average Propensity to Consume (APC): Defined empirically as average lagged consumption expenditures divided by total income, intended to capture persistent preference heterogeneity — a &amp;ldquo;spender type&amp;rdquo; — by measuring how much of income a household habitually spends before receiving the rebate. A one-percentage-point higher APC is associated with 0.19 additional cents spent per rebate dollar in the full multivariate specification.&lt;/p&gt;
&lt;p&gt;Forbidden Comparisons: A bias identified by Borusyak et al. (2024) in event-study designs with staggered treatment, arising when newly treated units are compared to previously treated units rather than true controls. The paper addresses this by regressing consumption changes on rebate receipt indicators (iota_{jl}) directly rather than on rebate amounts, and including lagged rebate indicators to account for persistent effects.&lt;/p&gt;</description></item><item><title>Leveraging Virtual Contact and Social Networks to Foster Interethnic Harmony</title><link>https://macropaperwarehouse.com/papers/leveraging-virtual-contact-and-social-networks-to-foster-interethnic-harmony/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/leveraging-virtual-contact-and-social-networks-to-foster-interethnic-harmony/</guid><description>&lt;p&gt;This paper investigates whether virtual contact — exposure to an outgroup through a documentary film — can promote interethnic harmony, and whether targeting network-central individuals amplifies effects on untreated community members. The study addresses a context of deep, historically rooted discrimination: the Santal ethnic minority in northwestern Bangladesh have faced colonial-era land dispossession, ongoing violence, labor market discrimination, and structural exclusion by the Bengali ethnic majority. The Santals are the second-largest ethnic-minority group in Bangladesh; in the study villages, their share ranges from 13% to 83% of the population.&lt;/p&gt;
&lt;p&gt;The authors conducted a cluster-randomized field experiment across 121 multiethnic villages in the Rajshahi and Naogaon districts of Bangladesh, involving over 3,300 households. Villages were randomly assigned to three arms: a random treatment arm (RR, 40 villages, N=562 Bengalis) in which approximately 14 randomly selected ethnic-majority households per village watched a 45-minute documentary film (&amp;ldquo;Ami Santal&amp;rdquo; / &amp;ldquo;I Am Santal&amp;rdquo;) portraying Santal culture, economic hardships, and aspirations; a central treatment arm (41 villages) in which approximately 7 randomly selected Bengalis (RC) and 7 network-central Bengalis identified via a diffusion-centrality nomination exercise (CC) watched the same film; and a control arm (40 villages) in which households watched a placebo documentary on flower farming. The documentary, costing approximately $13 per participant, was screened individually at participants&amp;rsquo; homes on tablets. Data were collected at baseline (September–October 2022), first end line approximately 3 months post-screening (February–March 2023), and a casual-work field experiment second end line approximately 4.5–5 months post-screening (April–May 2023). Outcomes were measured via lab-in-the-field experiments (dictator game, solidarity game), an experimentally validated interethnic trust survey item (Falk et al. 2018), self-reported behaviors, administrative police complaint data, and facial emotion detection during screening.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. First, treated Bengalis in the central arm (RC) gave 14.7% more in the dictator game (p &amp;lt; .01) and exhibited 21.7% greater trust toward Santals (p &amp;lt; .01) compared to controls; RR participants showed a 7.1% increase in solidarity game giving (p &amp;lt; .10) and 11.8% greater trust (p &amp;lt; .01). Effects on reducing negative stereotypes and discriminatory opinions were not statistically significant, suggesting that affective components of prejudice are more responsive to the intervention than cognitive components. About 82% of treated Bengalis reported acquiring new information about Santals, primarily regarding occupational struggles, educational aspirations, and economic potential. Facial expression analysis using emotion-detection software found sadness to be significantly more prevalent among viewers (p &amp;lt; .05), particularly among network-central participants, consistent with an empathetic response.&lt;/p&gt;
&lt;p&gt;Second, untreated Bengalis in the central arm — who never watched the documentary — showed 20.9% higher altruism (p &amp;lt; .10), 27.3% higher solidarity (p &amp;lt; .05), and 8.1% higher trust (p &amp;lt; .05) toward Santals relative to controls. No significant effects on untreated Bengalis were found in the random arm. Untreated Santals in both arms exhibited greater trust toward Bengalis (11% increase in random arm, p &amp;lt; .05; 21.7% increase in central arm, p &amp;lt; .01) and higher subjective well-being (p &amp;lt; .01 in both arms). Village-level administrative data show a significant reduction in Bengali police complaints against Santals post-intervention (p &amp;lt; .05), but only in the central arm.&lt;/p&gt;
&lt;p&gt;Third, in the casual-work field experiment, multiethnic pairs jointly produced paper bags under piece-rate compensation. Overall productivity increased approximately 5% (p &amp;lt; .05) in the central arm only. Both Bengali and Santal workers increased productivity specifically in the finisher role — the most critical role for determining earnings — in the central arm. The authors interpret Bengali productivity gains as reflecting increased prosociality toward Santal co-workers, and Santal productivity gains as reflecting conformism or peer pressure in response to Bengali effort. The scope of all effects is limited to multiethnic villages in northwestern Bangladesh, a context of historically severe and ongoing majority-minority inequality; the intervention deliberately did not challenge the socioeconomic hierarchy of the villages.&lt;/p&gt;
&lt;p&gt;Q: What was the documentary film&amp;rsquo;s content and design rationale?
A: The 45-minute film &amp;ldquo;Ami Santal&amp;rdquo; featured three narrative layers: Santal culture (rituals, cuisine, the Baha festival), economic hardships (housing, water access, low incomes, labor market struggles, educational barriers), and aspirational stories of Santals who achieved success. All stories were narrated by non-actor local Santals, filmed outside the study region, and deliberately avoided attributing blame to Bengalis. The film was designed under the supervision of anthropologists at the University of Rajshahi to maintain ethnographic authenticity and a non-moralistic, observational tone (moral judgment language was much lower than in comparison Bangladeshi documentaries and general films, per LIWC-22 analysis).&lt;/p&gt;
&lt;p&gt;Q: How were network-central individuals identified and why might targeting them matter?
A: In central-arm villages, enumerators surveyed approximately 18–20 randomly selected passers-by at village markets and asked them to nominate the 15 people most effective at disseminating information. The seven most consistently and highly ranked individuals per village were selected as network-central (CC). These individuals were expected to have high diffusion centrality — meaning information they receive spreads widely — so targeting them with the documentary could shift attitudes and behavior among untreated community members through persuasion, visibility, credibility, or diffusion (the paper cannot separately identify which mechanism operates).&lt;/p&gt;
&lt;p&gt;Q: What were the primary behavioral effects on treated Bengalis (the ethnic majority who watched the film)?
A: Randomly selected participants in the central arm (RC) gave 14.7% more in the dictator game (p &amp;lt; .01) and 8% more in the solidarity game (not statistically significant), and exhibited 21.7% greater trust toward Santals (p &amp;lt; .01), all relative to controls. In the random arm (RR), participants showed a 6.4% increase in dictator game giving (not statistically significant), a 7.1% increase in solidarity game giving (p &amp;lt; .10), and 11.8% greater trust toward Santals (p &amp;lt; .01). Effects on self-reported behaviors — interethnic friendships, social interactions, amount charged to minorities for water — were not statistically significant.&lt;/p&gt;
&lt;p&gt;Q: Did the intervention change Bengali stereotypes or discriminatory opinions toward Santals?
A: No. Despite treated Bengalis acquiring substantial new information (approximately 82% reported learning new things, primarily about Santal occupational struggles and educational aspirations), the authors find no significant effects on the stereotypes index or the discriminatory-opinions index among treated Bengalis. They propose two explanations: cognitive components of prejudice (stereotypes) are harder to change through indirect contact than affective components (emotions, prosocial behavior), consistent with Tropp and Pettigrew (2005) and Turner, Crisp, and Lambert (2007); and a single documentary may be insufficient to counter deeply ingrained generational biases due to resistance to change.&lt;/p&gt;
&lt;p&gt;Q: What emotional responses did the documentary elicit, and how was this measured?
A: Field assistants took candid photographs of participants&amp;rsquo; faces at a random point during the screening; these were analyzed using Emotimeter software (machine learning-based emotion detection) that assigns scores across seven emotion categories summing to 100%. Sadness was significantly more prevalent among documentary viewers compared to placebo viewers (p &amp;lt; .05), particularly among network-central participants (CC). The authors interpret this as consistent with an empathetic response to the film&amp;rsquo;s content about Santal hardships, and connect it to increased prosocial behavior via emotion-regulation mechanisms (alleviating sadness through prosocial action).&lt;/p&gt;
&lt;p&gt;Q: What were the spillover effects on untreated Bengalis in the central arm?
A: Untreated Bengalis in central-arm villages — who never watched the documentary — showed 20.9% higher altruism (p &amp;lt; .10), 27.3% higher solidarity (p &amp;lt; .05), and 8.1% higher trust toward Santals (p &amp;lt; .05) relative to controls. By contrast, untreated Bengalis in random-arm villages showed no statistically significant effects on any of these outcomes. The authors attribute the central-arm spillovers to the presence of network-central individuals being treated in those villages, though whether these patterns reflect persuasion, visibility, credibility, or information diffusion cannot be separately identified.&lt;/p&gt;
&lt;p&gt;Q: How did the intervention affect the Santal ethnic minority (who never watched the documentary)?
A: Untreated Santals in both arms exhibited greater trust toward Bengalis: an 11% increase in the random arm (p &amp;lt; .05) and a 21.7% increase in the central arm (p &amp;lt; .01) compared to controls. Santals in both arms also reported higher subjective well-being (p &amp;lt; .01). A weakly significant increase in food security was observed among Santals in the central arm (p &amp;lt; .10), possibly reflecting increased material support from Bengalis. No statistically significant effects were found on Santal altruism or solidarity.&lt;/p&gt;
&lt;p&gt;Q: What did the village-level administrative complaint data show?
A: Using data collected from two police stations covering all 121 villages, the authors find a significant reduction in Bengali complaints against Santals post-intervention in the central arm (p &amp;lt; .05). No significant reduction was found in Santals&amp;rsquo; complaints against Bengalis (p &amp;gt; .10) in any arm. Data from village counselors&amp;rsquo; offices (shalish arbitration complaints) showed no significant change in any arm. The distinction matters because police complaints involve more serious, violent matters, while village-counselor complaints involve routine arbitration.&lt;/p&gt;
&lt;p&gt;Q: How was the casual-work field experiment designed, and what did it find?
A: Approximately 4.5 months after the documentary screenings, 720 participants (360 Bengalis, 360 Santals) drawn equally from the three study arms were paired into multiethnic dyads to jointly produce paper bags for a local supplier under piece-rate compensation, with earnings split equally. One worker was randomly assigned the preparer role and the other the finisher role; roles were switched halfway through the three-hour session. The paper finds an approximately 5% overall productivity increase (p &amp;lt; .05) in the central arm only, concentrated in the finisher role (the role most critical for final output). Bengalis and Santals both increased productivity specifically as finishers in the central arm.&lt;/p&gt;
&lt;p&gt;Q: What mechanisms explain the productivity effects in the casual-work experiment?
A: For Bengali finishers, the productivity gain is interpreted as prosocial behavior: treated Bengalis who showed greater altruism toward Santals worked harder to increase the earnings of their Santal co-workers. For Santal finishers, the productivity gain is interpreted as conformism or peer pressure: Santals increased effort more when they worked as finisher after swapping roles (i.e., after observing Bengalis&amp;rsquo; higher effort as finisher first), suggesting responsiveness to the higher productivity of Bengalis rather than an independent prosocial motivation. The authors present a simple theoretical model to formalize these interpretations, citing Rotemberg (1994) on prosocial effort and Kandel and Lazear (1992) and Mas and Moretti (2009) on peer pressure mechanisms.&lt;/p&gt;
&lt;p&gt;Q: Why was virtual rather than direct contact used in this intervention?
A: The authors argue that encouraging direct contact between Bengalis and Santals in this setting carries specific risks: the unequal status of the groups may generate anxiety during interactions, potentially limiting engagement or provoking backlash. By contrast, the documentary provides an indirect, low-cost ($13 per participant) form of contact that presents Santal lives without disrupting the socioeconomic hierarchy of the villages and without attributing blame to Bengalis. The film&amp;rsquo;s entertaining veneer and emotional storytelling make it more scalable and logistically feasible in contexts where direct contact is socially difficult or impractical.&lt;/p&gt;
&lt;p&gt;Q: What are the primary limitations acknowledged by the authors?
A: The authors acknowledge that the study&amp;rsquo;s sampling protocol relied on a door-to-door skip procedure without systematic records of approached households, raising the possibility of convenience or snowball-type recruitment and potential deviations from random sampling — this is reflected in some imbalances in baseline characteristics across arms. CC-control comparisons are explicitly descriptive (not causal) because network-central individuals were selected on centrality. Differential attrition was found among untreated Santals (both treatment arms had significantly lower attrition than control, p &amp;lt; .05), which could bias estimates for that subgroup. The authors cannot separately identify the mechanisms (persuasion, visibility, credibility, diffusion) underlying spillover effects in central villages.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of this study?
A: The findings suggest that media-based virtual contact interventions are a low-cost, scalable tool for improving interethnic prosociality even in contexts of deep-rooted discrimination where direct contact may be socially impractical. Targeting network-central individuals — identified via a simple nomination exercise requiring no pre-existing network data — amplifies village-wide effects, including among untreated community members and the minority group itself. The productivity gains in multiethnic work teams imply that improved interethnic relations can have tangible economic consequences beyond attitudinal change. However, the null effects on stereotypes and discriminatory opinions suggest that single documentary interventions may not be sufficient to alter deep-seated cognitive biases, and more intensive or repeated interventions may be needed to achieve durable attitude change.&lt;/p&gt;
&lt;p&gt;Virtual contact: Indirect exposure to an ethnic outgroup through a documentary film, as distinct from direct intergroup contact; posited to influence majority-group attitudes and behavior by increasing empathy and identification with the outgroup without requiring face-to-face interaction.&lt;/p&gt;
&lt;p&gt;Diffusion centrality: A network measure of how effectively an individual can spread information through a community, operationalized via a nomination exercise in which community members identify those best positioned to disseminate information; used to select the seven highest-ranked individuals per village for targeted treatment.&lt;/p&gt;
&lt;p&gt;Prosociality (altruism and solidarity): Measured using incentivized lab-in-the-field games — the dictator game (unilateral allocation of an endowment to a passive outgroup recipient) and the solidarity game (precommitted transfers to an outgroup member who may incur a random loss) — capturing willingness to benefit non-coethnic others at personal cost.&lt;/p&gt;
&lt;p&gt;Affective versus cognitive components of prejudice: A distinction between emotional aspects of prejudice (feelings, empathy) — which the authors find to be more responsive to the documentary intervention — and cognitive aspects (negative stereotypes, discriminatory opinions) — which show no significant change despite new information acquisition.&lt;/p&gt;
&lt;p&gt;Spillover effects (untreated individuals): Changes in behavior or attitudes among community members who did not directly receive the intervention (did not watch the documentary), attributed to the influence of treated individuals in their village, particularly network-central individuals in the central arm.&lt;/p&gt;
&lt;p&gt;Piece-rate casual-work field experiment: A second end line in which multiethnic pairs of Bengali and Santal workers jointly produced paper bags for a local supplier, with individual earnings determined by joint piece-rate output; designed to measure whether improved interethnic attitudes translated into higher workplace productivity in ethnically mixed teams.&lt;/p&gt;
&lt;p&gt;Source text origin: The provenance classification of the text used to generate a paper summary (full PDF, open-access HTML, or abstract only); the paper&amp;rsquo;s pipeline rules impose a hard block on abstract-only summarization.&lt;/p&gt;</description></item><item><title>Life-Cycle Wages and Human Capital Investments: Selection and Missing Data</title><link>https://macropaperwarehouse.com/papers/life-cycle-wages-and-human-capital-investments-selection-and-missing-data/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/life-cycle-wages-and-human-capital-investments-selection-and-missing-data/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 &amp;ndash; Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks how wage inequalities build up over the life cycle when individual wage trajectories are plagued by interruptions in private-sector participation, and when the standard Missing At Random (MAR) assumption used to handle those gaps may be violated. Specifically, it asks: what is the causal effect of career interruptions on both the level and the dispersion of wages after twenty years of potential experience, and does endogeneity of those interruptions matter for the dispersion result?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Sample&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The empirical analysis uses the 2011 DADS Grand Format-EDP panel, a French administrative dataset merging social security records (DADS) and census extracts (EDP). The working sample covers males who entered the private sector between 1985 and 1992, aged 16-30 at entry, and observed through 2011. The authors require at least 15 years of observed private-sector wages, yielding a working sample of 7,004 males and 137,315 person-year observations. Education is grouped into four levels (high-school dropouts, high-school graduates, some college, college graduates). Participation outside the private sector &amp;ndash; including public-sector employment, self-employment, unemployment, and non-employment &amp;ndash; constitutes the &amp;ldquo;alternative sector&amp;rdquo; and generates missing wage observations. On average, cumulative duration outside the private sector is 3.7 years, and the average number of interruptions is 1.44.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper builds on a structural Ben Porath (1967) human capital model extended to two sectors (private sector and an alternative sector), yielding a reduced-form log-wage equation with five individual-specific coefficients: an intercept (initial human capital), a linear trend in potential experience (growth rate), a curvature term in potential experience (Mincer concavity), the cumulative years of interruptions, and a curvature term in interruptions. Because parameters are individual-specific, the wage equation is a random-coefficient model estimated with a fixed-effects approach.&lt;/p&gt;
&lt;p&gt;Selection into the private sector is addressed not by a standard MAR assumption but by a weaker &amp;ldquo;Missing At Random Conditionally On Factors&amp;rdquo; (MARCOF) assumption. Sector-preference shocks, human capital prices, and depreciation rates are each decomposed into a common factor (time-varying) and an individual factor loading, plus a residual that is mean-independent of factors and loadings. Conditional on factors and factor loadings, wage residuals and sector choices are independent, making covariates &amp;ndash; including the interruption variables &amp;ndash; exogenous. The preferred specification includes two unobserved factors, selected by four of six Bai-Ng (2002) information criteria.&lt;/p&gt;
&lt;p&gt;Estimation proceeds via an Expectation-Maximization (EM) algorithm adapted from Bai (2009) and Song (2013), with initial values from Moon and Weidner (2018)&amp;rsquo;s nuclear-norm convex estimator. Because individual parameters converge at rate sqrt(T) and summary statistics of their distributions suffer from incidental-parameter bias, the authors use bias-correction methods from Jochmans and Weidner (2019) for quantiles and inter-decile ranges, and from Arellano and Bonhomme (2012) for variances. Monte Carlo experiments confirm that variances remain poorly corrected even when T &amp;gt; 20, so the paper focuses on inter-decile ranges as the dispersion measure.&lt;/p&gt;
&lt;p&gt;Counterfactual &amp;ldquo;average structural functions&amp;rdquo; (Blundell and Powell, 2003) are constructed by holding individual parameters fixed and manipulating the history of interruptions. These compare four scenarios: the observed benchmark, the counterfactual with no interruptions (potential wage), the counterfactual with no current-period selection, and both combined.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Downward bias from omitting interruptions and factors.&lt;/em&gt; Omitting interruption variables and unobserved factors strongly downward biases estimated returns to experience after 20 years. Most of this bias is attributable to interruptions rather than to the interactive factor effects: selectivity is mainly captured through the interruption channel, not through residual factor structure.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Effect on mean wages.&lt;/em&gt; Potential experience increases log wages by approximately 65% over 20 years, consistent with cross-country evidence from homogeneous Mincer equations. The average cost of interruptions after 20 years is approximately 10% of log wages. Reassigning interruptions to the beginning of the working life has a persistent negative effect on mean log wages that never fully recovers over 20 years, while reassigning them to the end increases mean wages above the no-interruption benchmark at every experience level.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Effect on wage dispersion &amp;ndash; a new stylized fact.&lt;/em&gt; Interruptions decrease, not increase, the inter-decile range of log wages after 20 years. After 20 years, with an average interruption duration of 2.47 years, interruptions decrease the inter-decile range by 0.52 log points (approximately 38%). This compression operates differentially: the 90th percentile falls by 0.34 and the 10th percentile rises by 0.18.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Endogeneity explains the dispersion compression.&lt;/em&gt; When years of interruption are randomly reassigned across time (holding total interruption years fixed), the inter-decile range diverges upward from the observed benchmark after about 5 years. This shows that the dispersion-reducing effect of actual interruptions is due to the endogenous timing of those interruptions &amp;ndash; specifically to the negative correlation between the timing of interruptions and potential log wages &amp;ndash; rather than to the correlation between the structural coefficients on interruptions and potential wages (which is also negative, with a Spearman rank correlation of -0.32 between eta_i1 and eta_i3). Endogenously chosen interruptions smooth inequality over time.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Current-period selection is negligible.&lt;/em&gt; Current-period selection into private-sector employment has no statistically significant effect on median, mean, variance, or inter-decile range of wages at any experience level, as confirmed by the small inter-decile range of the interactive factor component.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Results pertain to cohorts of French males entering the private sector between 1985 and 1992, restricted to those with at least 15 observed private-sector years. The French context is distinctive: wage inequality in the working population was stable over 1985-2011, driven in part by minimum wage policy and payroll tax exemptions for lower-skilled workers, in contrast to rising inequality in the United States and Germany. Results on timing of interruptions (eta_i3 and eta_i4) are identified only for individuals with at least two interruptions followed by re-entry (roughly those with K_T &amp;gt;= 2). The paper does not analyze female wages.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-structural-model-and-how-does-it-generate-a-reduced-form-wage-equation"&gt;Q1. What is the structural model and how does it generate a reduced-form wage equation?&lt;/h3&gt;
&lt;p&gt;The model is a Ben Porath (1967) two-sector human capital model in which individuals divide time between investing in human capital and earning wages in either the private sector (e) or an alternative sector (n). Human capital accumulation in each sector has a sector-specific return rate (rho^s) and depreciation (lambda^s_t). Period utility is log income minus a quadratic investment cost, plus a sector preference shock. Solving the dynamic program backwards (because of log-linearity) yields closed-form optimal investments that are linear in the individual-specific terminal value of human capital (kappa). The resulting log-wage equation (Proposition 5) is a function of five terms: an intercept (eta_i0), a linear trend in potential experience t (eta_i1), a geometric curvature term beta^{-t} (eta_i2), cumulative years of interruptions x^(3)_it (eta_i3), and a curvature in interruptions x^(4)_it (eta_i4), all with individual-specific coefficients. This provides a tractable random-coefficient structure.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-marcof-assumption-and-why-is-it-weaker-than-mar"&gt;Q2. What is the MARCOF assumption and why is it weaker than MAR?&lt;/h3&gt;
&lt;p&gt;MARCOF &amp;ndash; Missing At Random Conditionally On Factors &amp;ndash; posits that sector-preference shocks, human capital prices, and depreciation rates each follow factor structures: a common time-varying factor (phi_t) multiplied by an individual loading (theta_i) plus an i.i.d. residual. The residuals are assumed mean-independent of factors and loadings, and independent over time. Under standard MAR, missingness is assumed independent of outcomes conditional on observables alone. Under MARCOF, residuals in the wage equation and the sector choice equation are independent conditional on (unobserved) factors and factor loadings. This is weaker than MAR because it allows the unobservable determinants of wages and participation to share common factors, accommodating the high persistence observed in human capital stocks (20-year lag correlation of 0.28, far above the geometric decay benchmark of 0.024).&lt;/p&gt;
&lt;h3 id="q3-how-are-the-individual-specific-parameters-identified"&gt;Q3. How are the individual-specific parameters identified?&lt;/h3&gt;
&lt;p&gt;Under exogenous selection (or, under MARCOF, conditional on factors), identification of eta_i0, eta_i1, and eta_i2 requires variation in potential experience within the individual&amp;rsquo;s time series. Identification of eta_i3 and eta_i4 separately requires individuals to experience at least two spells out of the private sector each followed by re-entry (at least four transitions, so K_T &amp;gt;= 2). An individual with only one interruption spell generates proportional variation in x^(3) and x^(4), so only a linear combination of eta_i3 and eta_i4 is identified. The &amp;ldquo;flat spot&amp;rdquo; approach &amp;ndash; using the observed fact that individuals aged 50-55 have stopped investing in human capital &amp;ndash; separately identifies time, cohort, and age effects and provides the restriction that factors are orthogonal to the level, trend, and curvature in potential experience.&lt;/p&gt;
&lt;h3 id="q4-what-do-the-distributions-of-estimated-individual-specific-coefficients-look-like"&gt;Q4. What do the distributions of estimated individual-specific coefficients look like?&lt;/h3&gt;
&lt;p&gt;Focusing on the main (two-factor) specification with bias correction: the median of the growth parameter eta_i1 is positive (consistent with rising wages with experience) and the median of the curvature parameter eta_i2 is negative (consistent with concavity). However, heterogeneity is substantial: the 90th percentile of eta_i1 is 6.2 times the median, and the first quartile of eta_i1 is negative (implying declining potential wages for a non-negligible share). For the interruption coefficients eta_i3 (year of interruptions) and eta_i4 (curvature), bias-corrected medians are close to zero in the sub-sample with &amp;gt;=2 interruptions, but dispersion is large and symmetric around zero. Bias correction reduces the 90th percentile of eta_i1 by approximately 20% and reduces the absolute 10th percentile of eta_i3 by approximately 27%.&lt;/p&gt;
&lt;h3 id="q5-how-important-are-interruptions-relative-to-potential-experience-and-factors-in-explaining-wage-variation"&gt;Q5. How important are interruptions relative to potential experience and factors in explaining wage variation?&lt;/h3&gt;
&lt;p&gt;A wage decomposition using inter-decile ranges (preferred over variance due to bias) shows that the potential experience component is the largest contributor to wage dispersion, followed by the interruption component (described as &amp;ldquo;sizable&amp;rdquo;), while factors play a minor role. Crucially, the potential experience and interruption components are highly negatively rank-correlated: the Spearman rank correlation between the growth coefficient eta_i1 and the interruption coefficient eta_i3 is -0.32. This negative correlation is central to understanding why interruptions compress dispersion rather than expanding it.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-finding-on-the-effect-of-interruptions-on-mean-wages-and-what-does-the-timing-experiment-show"&gt;Q6. What is the finding on the effect of interruptions on mean wages, and what does the timing experiment show?&lt;/h3&gt;
&lt;p&gt;After 20 years, the average cost of interruptions (relative to a counterfactual of no interruptions) is approximately 10% of log wages. The timing of interruptions matters: reassigning interruptions to the beginning of the working life causes a persistent loss in mean log wages that does not fully recover over the 20-year horizon, while reassigning them to the end raises mean log wages above the no-interruption level at every experience level. For median wages, the early-interruption loss is eventually recovered (median log wages do catch up), but the mean does not catch up. These asymmetries are consistent with early interruptions having a larger negative effect on human capital accumulation due to the geometric structure of investment returns.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-key-finding-on-wage-dispersion-and-what-explains-it"&gt;Q7. What is the key finding on wage dispersion and what explains it?&lt;/h3&gt;
&lt;p&gt;Interruptions compress the inter-decile range of log wages by 0.52 log points (approximately 38%) after 20 years, with average interruption duration of 2.47 years. This compression is asymmetric: the 90th percentile of wages falls by 0.34 and the 10th percentile rises by 0.18. The dispersion-reducing effect is established by comparing the benchmark (observed interruptions) to the counterfactual of no interruptions. When interruptions are instead randomly reassigned across time (holding total interruption duration fixed), the inter-decile range diverges upward from the benchmark starting around 5 years of experience. This demonstrates that the compression is due to the endogenous timing of interruptions &amp;ndash; individuals who have high potential wages tend to time their interruptions in ways that reduce the measured spread of actual wages &amp;ndash; rather than to the negative structural coefficient (eta_i3 &amp;lt; 0 for high-wage workers on average).&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-handle-the-incidental-parameter-problem-for-distributional-statistics"&gt;Q8. How does the paper handle the incidental parameter problem for distributional statistics?&lt;/h3&gt;
&lt;p&gt;Because individual parameters are estimated at rate sqrt(T) and the panel is unbalanced (some individuals observed for as few as 15 years while the model has up to 7 individual parameters), standard distributional statistics like the variance suffer from substantial incidental parameter bias. Monte Carlo experiments show that bias-corrected variance estimates remain strongly biased even at T &amp;gt; 20. Inter-decile ranges are better behaved and the Jochmans and Weidner (2019) bias-correction procedure reduces their bias satisfactorily. This is why the paper reports inter-decile ranges as its primary dispersion measure rather than variances. The bias in corrected inter-decile ranges is at most approximately 10% of the uncorrected estimate.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-paper-show-about-the-mar-assumption-in-the-context-of-this-data"&gt;Q9. What does the paper show about the MAR assumption in the context of this data?&lt;/h3&gt;
&lt;p&gt;The results directly challenge the MAR assumption that is standard in the life-cycle earnings literature. Under MAR, interruptions would be treated as random conditional on observables, and their endogeneity would be ignored. The paper shows that treating interruptions as endogenous (through the MARCOF + structural model approach) substantially changes estimated returns to experience (there is a strong downward bias when interruptions and factors are omitted) and reverses the sign of the effect of interruptions on dispersion (under exogenous interruptions, randomly reassigned, dispersion would be higher than observed; the actual compression is an artifact of endogenous timing). The conclusion is that MAR assumptions produce systematically misleading pictures of life-cycle wage inequality dynamics.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-robustness-and-external-validity-considerations"&gt;Q10. What are the robustness and external validity considerations?&lt;/h3&gt;
&lt;p&gt;The working sample excludes individuals observed fewer than 15 years. A robustness exercise compares the subsample observed 10-14 years to a censored version of the 20+ subsample with matched marginal distributions of observation counts. Median profiles for the uncensored and censored 20+ samples are similar, and inter-decile ranges are slightly more dispersed in the censored sample only for potential experience greater than 7. However, the 10-14 year sample shows substantially different patterns &amp;ndash; larger median gaps between benchmark and no-interruption cases, and a larger inter-decile range &amp;ndash; consistent with lower private-sector returns to human capital for that group. The authors conclude that selection into the 15+ working sample matters, and results are explicitly restricted to that working sample. The French context (stable aggregate wage inequality, minimum wage policy) limits direct comparability to countries with rising inequality.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;MARCOF (Missing At Random Conditionally On Factors):&lt;/strong&gt; The paper&amp;rsquo;s central identifying assumption, weaker than standard MAR. It posits that sector-preference shocks, human capital prices, and depreciation rates follow factor structures (common time-varying factor x individual loading + i.i.d. residual), and that residuals are mean-independent of factors, loadings, and their own histories. Conditional on factors and loadings, wage residuals and sector-choice residuals are independent, making selection exogenous.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Interactive effects / factor structure for selection:&lt;/strong&gt; An approach in which unobserved confounders are modeled as a bilinear product of time-varying common factors (phi_t) and individual factor loadings (theta_i). This allows flexible correlation between wage processes and participation choices without requiring exclusion restrictions or instrumental variables. The paper&amp;rsquo;s preferred specification uses two unobserved factors identified by Bai-Ng information criteria.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Average structural functions:&lt;/strong&gt; Objects defined by Blundell and Powell (2003) that integrate counterfactual outcomes (wages evaluated at a manipulated interruption history) over the distribution of individual-specific parameters. They allow estimation of the causal impact of a change in interruption timing or presence while holding individual structural parameters fixed, under identification conditions analogous to those of Chernozhukov et al. (2013).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Individual-specific coefficients (random coefficients):&lt;/strong&gt; The five parameters (eta_i0, eta_i1, eta_i2, eta_i3, eta_i4) governing each individual&amp;rsquo;s wage equation, with structural interpretations: initial log human capital, return to potential experience, curvature (Mincer concavity), effect of cumulative interruption years, and curvature in interruptions. Their individual-specificity is the source of the incidental parameter problem for distributional statistics.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Flat spot approach:&lt;/strong&gt; An identification device (from Heckman, Lochner, and Taber, 1998; Bowlus and Robinson, 2012) that uses median wages of workers aged 50-55 &amp;ndash; who are assumed to have stopped investing in human capital &amp;ndash; as consistent estimates of human capital prices by education group and year. This separates the volume of human capital from its price, and provides the restriction identifying the level, trend, and curvature factors from the time-varying unobserved factors phi_t.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Interruption variables x^(3) and x^(4):&lt;/strong&gt; Reduced-form variables derived from the structural model summarizing the history of private-sector participation gaps. x^(3)_it is the cumulative number of periods spent in the alternative sector prior to date t; x^(4)_it is a geometric-weighted version of those interruptions that reflects the timing (early vs. late) through the discount factor beta. They enter the wage equation with individual-specific coefficients that are identified only for workers with at least two complete interruption spells.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mincer dip:&lt;/strong&gt; A U-shaped profile in wage variance (or inter-decile range) over potential experience, predicted by the Ben Porath model because high-return workers invest more at the start of their careers (reducing current wages), causing their wage profile to cross below then above low-return workers. Estimated in this paper at approximately 5 years of potential experience under the main specification.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Incidental parameter bias in distributional statistics:&lt;/strong&gt; The bias that arises when estimating moments or quantiles of the distribution of individual-specific parameters that converge at rate sqrt(T) rather than sqrt(N). The paper shows through Monte Carlo experiments that variance estimates remain substantially biased even after Arellano-Bonhomme (2012) correction when T &amp;gt;= 20, while inter-decile ranges corrected by Jochmans-Weidner (2019) are more reliable.&lt;/p&gt;</description></item><item><title>Linking Social and Personal Preferences: Theory and Experiment</title><link>https://macropaperwarehouse.com/papers/linking-social-and-personal-preferences-theory-and-experiment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/linking-social-and-personal-preferences-theory-and-experiment/</guid><description>&lt;p&gt;This paper asks whether an individual&amp;rsquo;s attitude toward risk in the personal domain (choices affecting only oneself) can be linked to that same individual&amp;rsquo;s attitude toward risk in the social domain (choices affecting both oneself and others). The authors provide a theoretical answer in the form of necessary and sufficient conditions, and then test those conditions experimentally.&lt;/p&gt;
&lt;p&gt;The formal model posits a decision maker (DM) with a preference relation over lotteries on a set of social states, where a distinguished subset of states are personal (consequences for the DM alone). The authors assume preferences satisfy Completeness, Transitivity, Continuity, and State Monotonicity — the last being equivalent to respect for First-Order Stochastic Dominance (FOSD), a condition weaker than the Expected Utility Independence Axiom and satisfied by virtually all extant decision theories including Weighted Expected Utility, Rank-Dependent Utility, and Prospect Theory. The key theoretical result (Theorem 1) establishes that the full preference relation over all social lotteries can be uniquely deduced from the partial observations of (i) riskless social choices and (ii) risky personal choices if and only if the DM finds every social state indifferent to some personal state. When this condition fails, there exist social lotteries whose ranking cannot be recovered from the partial data.&lt;/p&gt;
&lt;p&gt;For two empirically relevant preference types, this condition generates directly testable predictions: for selfish subjects (who allocate nothing to others in deterministic social choices), risky personal preferences must coincide with risky social preferences; for impartial subjects (who treat self and other symmetrically in deterministic social choices), riskless social preferences must coincide with risky social preferences.&lt;/p&gt;
&lt;p&gt;The experiment was conducted at the University of Bergen and NHH Norwegian School of Economics with 276 undergraduate subjects. Each subject faced 50 budget-line choice problems in each of three domains: Personal Risk (equiprobable binary lotteries over own payoffs only), Social Choice (deterministic splits between self and an anonymous other), and Social Risk (equiprobable binary lotteries over symmetric payout pairs for self and other). The graphical interface of Choi et al. (2007b) was used throughout. One randomly selected decision per domain was paid out; each token was worth 1.2 NOK (approximately 0.2 USD), with average earnings of approximately 270 NOK.&lt;/p&gt;
&lt;p&gt;Within-domain consistency, measured by the Critical Cost Efficiency Index (CCEI), is high: mean CCEIs are 0.959, 0.952, and 0.902 in the Personal Risk, Social Choice, and Social Risk domains respectively. At the CCEI &amp;gt; 0.90 threshold, 89.9%, 85.9%, and 69.9% of subjects pass in the three domains. Using a 0.95 share-to-self threshold, 103 subjects (37.3%) are classified as selfish; using revealed-preference criteria at the 5% significance level, 33 subjects (12.0%) are classified as impartial.&lt;/p&gt;
&lt;p&gt;Testing is done via an individual-level nonparametric permutation test that draws 10,000 random data sets per subject and compares simulated CCEI distributions to actual cross-domain CCEIs, with Bonferroni correction. At the 1% significance level, the null that Personal Risk and Social Risk preferences coincide is rejected for only 5.9%–9.3% of selfish subjects (varying by classification threshold), compared with 14.7%–16.3% rejection rates for non-selfish subjects. For impartial subjects at the 1% level, the null that Social Choice and Social Risk preferences coincide is rejected for 0.0%–11.1%, compared with 19.8%–26.8% for non-impartial subjects. The theory&amp;rsquo;s predictions are thus supported for a large majority of both selfish and impartial subjects.&lt;/p&gt;
&lt;p&gt;A theoretical extension (Theorem 2) shows that if one additionally observes comparisons between social states and personal lotteries, unique deduction of the full preference relation requires that preferences in both personal and social domains satisfy Expected Utility (Independence Axiom) and that every social state is indifferent to some personal lottery — a strictly stronger set of conditions.&lt;/p&gt;
&lt;p&gt;Q: What is the central theoretical question and why does it matter?
A: The paper asks whether preferences over risky social choices (lotteries over outcomes for self and others) can be deduced from observing only riskless social choices and risky personal choices. This matters because people frequently observe or predict the risky social choices of leaders and representatives, but may have access only to those leaders&amp;rsquo; personal risk-taking behavior and their expressed social preferences under certainty.&lt;/p&gt;
&lt;p&gt;Q: What is the main theoretical result (Theorem 1)?
A: Under Completeness, Transitivity, Continuity, and State Monotonicity, the unique extension of the partial preference relation (over social states and personal lotteries) to the full domain of social lotteries exists if and only if every social state is indifferent to some personal state. When this condition is not met, multiple distinct preference relations can extend the partial observations, making deduction impossible.&lt;/p&gt;
&lt;p&gt;Q: What is State Monotonicity and how does it relate to standard axioms?
A: State Monotonicity requires that if each social state in one lottery dominates the corresponding state in another lottery, then the first lottery is weakly preferred. The paper shows this is equivalent to respect for First-Order Stochastic Dominance (FOSD) given the other axioms, and is strictly weaker than the von Neumann–Morgenstern Independence Axiom. It is satisfied by Weighted Expected Utility, Rank-Dependent Utility, and Prospect Theory, making it a broadly applicable assumption.&lt;/p&gt;
&lt;p&gt;Q: What are the testable predictions for selfish subjects?
A: Proposition 2 establishes that if a subject&amp;rsquo;s Social Choice preferences are selfish — meaning any bundle (x, y) is indifferent to (0, y), so the subject is indifferent between keeping x for self and giving it to other — then preferences in the Personal Risk domain must coincide with preferences in the Social Risk domain. In the experiment, selfish subjects are those allocating more than 95% of tokens to themselves in the Social Choice domain (103 of 276 subjects, or 37.3%).&lt;/p&gt;
&lt;p&gt;Q: What are the testable predictions for impartial subjects?
A: Proposition 3 establishes that if a subject&amp;rsquo;s Social Choice preferences are symmetric — meaning (x, y) is indifferent to (y, x) for all pairs — then preferences in the Social Choice domain must coincide with preferences in the Social Risk domain, implying risk neutrality toward social lotteries. The intuition is that such a subject treats self and other identically, so risky splits are evaluated by expected value alone. In the experiment, 33 subjects (12.0%) are classified as impartial by the revealed-preference criterion at the 5% significance level.&lt;/p&gt;
&lt;p&gt;Q: How does the experiment measure within-domain rationality?
A: Choices within each domain are evaluated using the Critical Cost Efficiency Index (CCEI, following Afriat 1967), which measures how much a budget constraint must be relaxed to remove all GARP violations. Mean CCEIs are 0.959 (Personal Risk), 0.952 (Social Choice), and 0.902 (Social Risk). At the CCEI &amp;gt; 0.90 threshold, 248 subjects (89.9%), 237 (85.9%), and 193 (69.9%) pass in the three domains respectively, compared to a simulated mean CCEI of only 0.585 for subjects randomizing uniformly.&lt;/p&gt;
&lt;p&gt;Q: How does the cross-domain test work and why is it nonparametric?
A: The test uses individual-level permutation inference: under the null that preferences in domains I and J are identical, any 50-element subset drawn from the pooled 100 choices should satisfy GARP as well as the actual domain-specific choices. For each subject, 10,000 such random draws are generated, their CCEI scores are computed, and the distribution is compared to the actual cross-domain CCEI with Bonferroni correction. The test makes no functional form assumptions about utility and accommodates the observed within-domain errors without parametric error modeling.&lt;/p&gt;
&lt;p&gt;Q: What are the rejection rates for the selfish-subject prediction?
A: At the 1% significance level, the null that Personal Risk and Social Risk preferences coincide is rejected for only 5.9%–9.3% of selfish subjects (range across four classification thresholds from 0.99 to 0.90 share-to-self), compared to 14.7%–16.3% for non-selfish subjects. At the 5% level, rejection rates rise to 20.4%–25.6% for selfish and 22.4%–31.8% for non-selfish subjects.&lt;/p&gt;
&lt;p&gt;Q: What are the rejection rates for the impartial-subject prediction?
A: At the 1% significance level, the null that Social Choice and Social Risk preferences coincide is rejected for 0.0%–11.1% of impartial subjects (range depending on threshold and classification method), compared to 19.8%–26.8% for non-impartial subjects. At the 5% and 10% levels, rejection rates for impartial subjects range from 0.0% to 22.2%.&lt;/p&gt;
&lt;p&gt;Q: Does the theory predict how risk aversion should map across domains for non-selfish, non-impartial subjects?
A: The theory does not directly produce testable cross-domain predictions for subjects who are neither selfish nor impartial without additional parametric assumptions, because the specific personal-state equivalent of each social state depends on the form of preferences. The paper restricts its nonparametric tests to the two polar cases where the equivalence mapping is determinate from social choice behavior alone.&lt;/p&gt;
&lt;p&gt;Q: What is the extended result (Theorem 2) and what stronger conditions does it require?
A: When one additionally observes comparisons between social states and personal lotteries (not just within each domain separately), unique deduction of the full preference relation is possible if and only if preferences in both the personal and social domains are consistent with an Expected Utility representation and every social state is indifferent to some personal lottery. This requires the Independence Axiom — a strictly stronger condition than State Monotonicity — highlighting that the main Theorem 1 result exploits the weaker observational structure.&lt;/p&gt;
&lt;p&gt;Q: What is the distribution of social preferences in the sample?
A: Of 276 subjects, 103 (37.3%) are classified as selfish at the 0.95 share-to-self threshold. Only 6 subjects (2.2%) kept fewer than 0.45 of tokens on average, making purely altruistic subjects rare. In the Personal Risk domain, 41 subjects (14.9%) allocated more than 95% to the cheaper account (consistent with risk neutrality), while 9 (3.3%) allocated fewer than 55% (consistent with infinite risk aversion). In the Social Risk domain, 30 subjects (10.9%) are consistent with utilitarianism in money and 9 (3.3%) with Rawlsianism in money.&lt;/p&gt;
&lt;p&gt;Q: How does the Social Risk domain compare to the Personal Risk and Social Choice domains in terms of rationality scores?
A: The Social Risk domain shows lower consistency than the other two: mean CCEI is 0.902 versus 0.959 and 0.952, and only 69.9% of subjects exceed the 0.90 threshold versus 89.9% and 85.9%. The CCEI distribution is shifted left for Social Risk, suggesting the novel combined dimension of social and risky choice introduces more decision complexity or error.&lt;/p&gt;
&lt;p&gt;Q: What is the relationship to the prior experimental literature on social and risk preferences?
A: The Personal Risk domain replicates the symmetric risk experiment of Choi et al. (2007a), and the Social Choice domain replicates the linear two-person dictator experiment of Fisman et al. (2007). The Social Risk domain is new to this paper. The theoretical framework connects to Saito (2013) on social preferences under risk, and to the preference extension literature of Grant et al. (1992) and Nishimura et al. (2017).&lt;/p&gt;
&lt;p&gt;State Monotonicity: The axiom requiring that if each social state in one lottery weakly dominates the corresponding social state in another lottery, the first lottery is weakly preferred. The paper proves this is equivalent to respect for First-Order Stochastic Dominance given Completeness, Transitivity, and Continuity, and distinguishes it from the stronger Independence Axiom by noting that Independence compares lotteries over lotteries while State Monotonicity only compares lotteries over states.&lt;/p&gt;
&lt;p&gt;Selfish preferences (in the paper&amp;rsquo;s sense): Preferences in the Social Choice domain such that (x, y) is indifferent to (0, y) for all bundles — the subject is indifferent between receiving x themselves versus giving x to the other person. Operationally measured as allocating more than a threshold share (e.g., 95%) of tokens to self across Social Choice decisions.&lt;/p&gt;
&lt;p&gt;Impartial preferences (in the paper&amp;rsquo;s sense): Preferences in the Social Choice domain such that (x, y) is indifferent to (y, x) for all bundles — the subject treats self and other symmetrically. Operationally identified by the revealed preference criterion that choices in the Social Choice domain satisfy GARP and are consistent with symmetric treatment.&lt;/p&gt;
&lt;p&gt;Unique extension (deducibility): The property that there exists exactly one complete preference relation over all social lotteries that is consistent with the axioms and agrees with the observed partial relation over social states and personal lotteries. Theorem 1 identifies the necessary and sufficient condition for unique extension under State Monotonicity.&lt;/p&gt;
&lt;p&gt;Personal state indifference condition: The condition that for every social state omega in Omega minus P, there exists some personal state in P to which the DM is indifferent. This is the necessary and sufficient condition in Theorem 1 for deducibility of the full preference relation. Interpreted as: for every proposed social allocation, there exists a &amp;ldquo;bribe&amp;rdquo; — a personal allocation with nothing for others — that the DM finds equally desirable.&lt;/p&gt;
&lt;p&gt;Critical Cost Efficiency Index (CCEI): A measure of how much budget constraints must be scaled down to eliminate all GARP violations in a dataset of choices from budget lines (following Afriat 1967). A CCEI of 1 indicates perfect rationality; the paper uses 0.90 as a practical threshold. Mean values are 0.959, 0.952, and 0.902 in the Personal Risk, Social Choice, and Social Risk domains respectively.&lt;/p&gt;
&lt;p&gt;Nonparametric permutation test: The individual-level test used to assess consistency across choice domains. Under the null that preferences are identical in domains I and J, any random 50-element draw from the pooled 100 choices should achieve CCEI scores no worse than the actual domain scores. The test draws 10,000 permuted datasets per subject and uses the Bonferroni correction for multiple comparisons, making no assumptions about the functional form of utility.&lt;/p&gt;</description></item><item><title>Making the Invisible Hand Visible: Managers and Worker Allocation</title><link>https://macropaperwarehouse.com/papers/making-the-invisible-hand-visible-managers-and-worker-allocation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/making-the-invisible-hand-visible-managers-and-worker-allocation/</guid><description>&lt;p&gt;This paper asks why managers matter for firm performance, and specifically whether managers improve productivity by matching workers to better-suited jobs inside firms rather than through supervision, motivation, or selection out of the firm. The setting is the internal labor market of a large private consumer goods multinational enterprise (MNE) operating in more than 100 countries, with annual turnover exceeding EUR 50 billion. The data cover the universe of white-collar workers and managers at the firm — 200,000 workers and 30,000 managers observed monthly over 11 years (January 2011 to December 2021) — linked to payroll, performance ratings, organizational chart, digital platform activity, employee surveys, and an independent sales productivity series for field sales workers in 15 countries.&lt;/p&gt;
&lt;p&gt;The paper confronts two identification challenges. First, the author constructs a measure of manager quality — &amp;ldquo;high flyers&amp;rdquo; — defined as managers who were promoted to the first managerial work level (WL2) by age 30. This threshold yields 26.2% of managers classified as high flyers. The measure is defined entirely ex ante, before the manager ever supervises the worker under study, which addresses reverse causality. It is validated against ex post performance metrics including future salary growth, probability of promotion to WL3, performance ratings, and anonymous subordinate feedback. Second, to identify causal effects of manager quality on workers, the author exploits the firm&amp;rsquo;s long-standing policy of rotating WL2 managers laterally across teams as part of their career development, a practice implemented for several decades. Using an event-study design centered on the worker&amp;rsquo;s first manager transition, the author compares workers who transition from a low-flyer to a high-flyer manager (LtoH) against workers who transition from one low-flyer to a different low-flyer (LtoL), netting out the effect of the transition itself. Pre-event parallel trends are confirmed empirically.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. Gaining a high-flyer manager causes substantial reallocation of workers within the firm through lateral job transfers: seven years after the manager transition event, cumulative lateral moves are 40% higher for workers who gained a high-flyer manager relative to those who gained another low-flyer. These lateral moves are not confined to a single organizational margin — transfers rise within-team, across teams in the same function, and across functions — and they involve meaningfully larger shifts in task content, as measured by angular separation across O*NET cognitive, routine, and social task intensity dimensions, with cumulative task distance becoming statistically distinguishable from zero approximately seven quarters post-transition. These gains in lateral mobility translate into persistent wage growth: seven years after the manager transition, workers supervised by a high-flyer earn salaries 13% higher than the comparison group, with divergence beginning only after the transition date. Using independent sales bonus data, three years after gaining a high-flyer manager workers&amp;rsquo; sales productivity increases by 0.347 standard deviations, ruling out the interpretation that wage gains merely reflect manager favoritism rather than genuine productivity improvement. Establishment-level data further show that sites with a higher share of workers under high-flyer managers display higher output per worker and lower operational costs per unit.&lt;/p&gt;
&lt;p&gt;Effects are asymmetric: gaining a good manager has large positive effects, but losing one (comparing HtoL with HtoH transitions) produces no corresponding negative effects, implying that a single exposure to a high-flyer manager generates durable benefits that survive a subsequent downgrade in manager quality. A mediation analysis finds that 64% of the salary gain is explained by lateral job changes, though the author notes this understates the full allocation channel because it excludes vertical transfers and the gains from remaining well-matched in the current role. These findings hold under multiple robustness checks including restricting to new hires, using the Sun and Abraham (2021) interaction-weighted estimator, varying the age threshold for high-flyer classification, using a tenure-based alternative, and placebo tests with randomly assigned manager types.&lt;/p&gt;
&lt;p&gt;The scope conditions are specific to white-collar workers at a large, organizationally homogeneous consumer goods multinational. All workers hold college degrees, mean firm tenure is 8.5 years, team sizes average five workers, and the firm has the same organizational structure across all countries, functions, and years.&lt;/p&gt;
&lt;p&gt;Q: How does the paper define &amp;ldquo;high flyer&amp;rdquo; managers and what share of managers receive this classification?
A: High flyers are managers who achieved the first managerial work level (WL2) by age 30, a threshold derived from continuous age estimates constructed from 10-year age bands in the personnel records. This definition yields 26.2% of managers classified as high flyers. The measure is time-invariant and defined ex ante relative to any interaction with the workers whose outcomes are studied.&lt;/p&gt;
&lt;p&gt;Q: What validates the high-flyer measure as capturing genuine managerial ability rather than noise?
A: The high-flyer classification is significantly positively correlated with multiple ex post performance metrics recorded after the manager&amp;rsquo;s own promotion: future salary growth, probability of subsequent promotion to WL3 (director level), annual performance ratings, and anonymous upward feedback scores from subordinates on leadership. High flyers are also 14.5 percentage points less likely to be mid-career recruits, suggesting they are internally developed talent rather than external hires.&lt;/p&gt;
&lt;p&gt;Q: What is the source of identifying variation and how does the event-study design address endogeneity?
A: The firm has operated a decades-long policy of rotating WL2 managers laterally across teams to broaden their experience and to screen candidates for promotion to WL3. These rotations are asserted by firm executives and HR representatives to be orthogonal to worker and team characteristics. The author verifies this empirically by showing that a wide range of team characteristics measured over the two years before a transition — including team performance, inequality, transfer rates, and team diversity — cannot predict the type of incoming manager. The event-study design compares workers who receive a high-flyer replacement (LtoH) against workers who receive another low-flyer replacement (LtoL), netting out any generic effect of a managerial change, and confirms parallel pre-trends.&lt;/p&gt;
&lt;p&gt;Q: What is the effect of gaining a high-flyer manager on lateral job mobility?
A: Seven years after the manager transition, workers assigned to a high-flyer manager exhibit lateral moves that are 40% higher relative to workers assigned to another low-flyer. These lateral moves occur across all organizational margins: within the same team, across teams within the same function (the largest contributor), and across functions. Beyond frequency, lateral moves under high-flyer managers also involve larger task-content shifts, with cumulative task distance (measured using O*NET cognitive, routine, and social task dimensions via angular separation) becoming statistically distinguishable from zero approximately seven quarters after the transition.&lt;/p&gt;
&lt;p&gt;Q: What is the wage effect of gaining a high-flyer manager and when does it materialize?
A: Workers who transition from a low-flyer to a high-flyer manager earn a salary 13% higher than workers who transition to another low-flyer, measured seven years after the transition event. The divergence begins only after the transition date, consistent with the pre-event parallel trends assumption, and accumulates gradually rather than appearing as an immediate jump.&lt;/p&gt;
&lt;p&gt;Q: Does the wage gain reflect genuine productivity improvement or simply managerial favoritism in pay decisions?
A: The author uses an independent sales bonus series — based on monthly targets set by supply chain demand planning teams, not by managers — for 5,604 field sales workers in 15 countries from 2018 to 2021. Three years after gaining a high-flyer manager, workers&amp;rsquo; sales productivity increases by 0.347 standard deviations. This confirms that pay gains correspond to actual productivity improvement rather than inflated ratings for unchanged performance.&lt;/p&gt;
&lt;p&gt;Q: How much of the wage gain is attributable to the lateral reallocation channel specifically?
A: A mediation analysis attributes 64% of the 13% salary gain to lateral job changes. The author cautions that this is a lower bound because the mediation excludes vertical transfers (which mechanically raise salary) and does not capture gains for workers who remain in their current job because it represents a good match rather than requiring reallocation.&lt;/p&gt;
&lt;p&gt;Q: Are the effects symmetric — does losing a high-flyer manager reverse the gains?
A: No. Comparing workers who transition from a high-flyer to a low-flyer manager (HtoL) against workers who transition from a high-flyer to another high-flyer (HtoH) reveals no corresponding negative effects. The gains from a single prior exposure to a high-flyer manager are persistent and are not undone by a subsequent low-quality manager. The author interprets this as evidence that a good match, once created, endures independently of the manager who created it.&lt;/p&gt;
&lt;p&gt;Q: Does gaining a high-flyer manager raise the rate of worker exit from the firm?
A: No. There is no statistically detectable effect on either voluntary exits (quits) or involuntary exits (layoffs), with null results that are not masked by heterogeneity across high- and low-performing workers. This rules out the interpretation that high-flyer managers improve measured outcomes of retained workers by selecting out underperformers.&lt;/p&gt;
&lt;p&gt;Q: Do workers move into roles connected to their high-flyer manager&amp;rsquo;s prior network or follow their manager when the manager moves?
A: No. There is no evidence that workers move into roles connected to the high-flyer manager&amp;rsquo;s prior colleagues; if anything, subordinates of high-flyer managers are less likely to make such moves. Workers also do not follow their high-flyer managers when those managers subsequently rotate to a different team. These findings rule out favoritism, social network access, and information-advantage explanations as primary drivers.&lt;/p&gt;
&lt;p&gt;Q: How does the paper rule out on-the-job teaching (human capital transmission) as the primary mechanism?
A: If high-flyer managers improved worker outcomes primarily by teaching workers to be more productive in their current job, the prediction would be reduced lateral mobility (workers become too productive to leave their current role). The observed pattern — substantially higher rates of lateral reallocation under high-flyer managers — is the opposite of this prediction, making teaching as the dominant channel unlikely.&lt;/p&gt;
&lt;p&gt;Q: What does the manager behavior evidence show about how high flyers spend their time?
A: Time-use data from a random sample of approximately 600 WL2 managers in 2019 show that high-flyer managers spend 19% more time in one-on-one meetings with subordinates and engage more in communication and multitasking activities relative to low-flyer managers. Their skill profiles also differ: high flyers are more likely to have strengths in strategy and talent management rather than project management, consistent with a more coordination-intensive and people-development-oriented style.&lt;/p&gt;
&lt;p&gt;Q: What heterogeneity is there in who benefits from high-flyer managers?
A: Effects are larger when managers and workers are in the same physical office (proximity facilitates talent assessment), when the organizational unit has a more diverse set of job roles (more matching opportunities), and for younger workers who are still discovering their comparative advantages. Critically, benefits are not concentrated among high-baseline performers: workers with low initial pay growth experience gains comparable to those of high performers, suggesting high-flyer managers uncover and deploy hidden talent broadly rather than accelerating only already-visible stars.&lt;/p&gt;
&lt;p&gt;Q: Does high-flyer management aggregate to establishment-level productivity?
A: Yes. Establishments where a higher share of workers are supervised by high-flyer managers show higher output per worker (tons per FTE) and lower operational costs per unit of output (operational costs per ton), measured using establishment-year data across approximately 150 sites globally over 2019-2021. This is consistent with the individual-level allocation mechanism producing aggregate productivity gains.&lt;/p&gt;
&lt;p&gt;Q: What are the organizational design implications of the asymmetric effects?
A: Because the gains from a single exposure to a high-flyer manager persist even after a subsequent manager downgrade, firms do not need each worker to be continuously supervised by a high-flyer. It is sufficient to rotate high-flyer managers across teams so that each worker receives at least one exposure. This makes the allocation mechanism resource-neutral relative to hiring, firing, or formal training programs.&lt;/p&gt;
&lt;p&gt;High flyer (paper&amp;rsquo;s definition): A manager who achieved the first managerial work level (WL2) at the firm by age 30 — a time-invariant, ex ante classification representing the firm&amp;rsquo;s revealed-preference assessment of leadership potential, validated against subsequent salary growth, promotion probability, performance ratings, and subordinate feedback. Constitutes 26.2% of managers in the sample.&lt;/p&gt;
&lt;p&gt;Internal labor market (paper&amp;rsquo;s usage): The system within the firm through which workers are allocated to jobs via lateral transfers and vertical promotions, mediated by managers rather than by external price mechanisms; the institutional context within which manager-worker matching produces wage growth and productivity gains.&lt;/p&gt;
&lt;p&gt;Lateral transfer (paper&amp;rsquo;s usage): A horizontal reallocation of a worker to a different job title, team, subfunction, or function at the same work level, as distinct from a vertical promotion. Captured monthly in personnel records; operationalized as moves involving changes in task content measured by O*NET task distances.&lt;/p&gt;
&lt;p&gt;Task distance (paper&amp;rsquo;s usage): The angular separation between origin and destination occupations across three O*NET task dimensions (cognitive, routine, and social intensity), ranging from zero (identical task profiles) to one (completely distinct profiles), used to characterize the substantive scope of lateral moves induced by high-flyer managers.&lt;/p&gt;
&lt;p&gt;Manager rotation (paper&amp;rsquo;s usage): The firm&amp;rsquo;s longstanding policy of reassigning WL2 managers laterally across teams within a subfunction, designed to broaden managerial experience and screen for promotion to WL3; treated in the empirical strategy as generating plausibly exogenous variation in the manager type each worker encounters.&lt;/p&gt;
&lt;p&gt;Allocation mechanism (paper&amp;rsquo;s usage): The process by which managers discover workers&amp;rsquo; specific skills and match them to specialized jobs inside the firm, operating through lateral reallocation rather than through hiring, firing, or on-the-job training; identified in the paper as the primary channel through which high-flyer managers generate persistent wage and productivity gains.&lt;/p&gt;
&lt;p&gt;Asymmetric persistence (paper&amp;rsquo;s usage): The empirical pattern in which the gains from gaining a high-flyer manager are large and durable, while losing a high-flyer manager (transitioning to a low-flyer) produces no corresponding negative effects on the outcomes of previously well-matched workers, implying that good matches, once formed, survive a change in manager quality.&lt;/p&gt;</description></item><item><title>Manager Pay Inequality and Market Power</title><link>https://macropaperwarehouse.com/papers/manager-pay-inequality-and-market-power/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/manager-pay-inequality-and-market-power/</guid><description>&lt;p&gt;This paper asks whether managers are paid for market power. Bao, De Loecker, and Eeckhout build a general equilibrium model in which firms compete oligopolistically in goods markets (following Atkeson and Burstein 2008) while managers are allocated to firms through a competitive matching market (following Gabaix and Landier 2008 and Tervio 2008). The model identifies two distinct channels through which market power and firm size jointly determine executive compensation: a market power channel, whereby a more productive firm charges a higher markup given its output level, and a firm size channel, whereby higher total factor productivity expands output given markups. Because manager ability and firm type are complementary inputs into TFP, assortative matching arises: high-ability managers sort into high-type firms, amplifying both productivity dispersion and markup dispersion across firms.&lt;/p&gt;
&lt;p&gt;The authors estimate the model year-by-year using Simulated Method of Moments on Compustat data covering 1994 to 2019, targeting ten moments including the average salary share, markup distribution, employment, and manager compensation levels. Firm-level markups are estimated using the production approach of De Loecker, Eeckhout, and Unger (2020). The ExecuComp variable TDC1 — encompassing salary, bonus, restricted stock grants, and option grant values — measures manager pay. Finance, insurance, and real estate sectors (SIC 6000–6799) are excluded.&lt;/p&gt;
&lt;p&gt;Main findings: market power accounts for on average 45.8% of total manager pay over the sample period, rising from 38.0% in 1994 to 48.8% in 2019. Over the full period, average CEO compensation (net of reservation utility) roughly doubled, from approximately $2.94 million to $6.43 million. Of the $3.49 million cumulative increase, $2.02 million (57.8%) is attributed to rising market power, with the remainder ($1.47 million) due to the firm size channel. The market power channel&amp;rsquo;s dominance is concentrated among top managers: for the highest-ranked managers in 2019, 80.3% of pay is attributable to market power, and nearly all of their pay growth since 1994 stems from the market power channel. For lower-ranked managers, pay is determined primarily by the firm size channel and has been roughly flat over the period.&lt;/p&gt;
&lt;p&gt;Within the market power channel, changes in technology — specifically increasing dispersion in firm-level TFP — are the dominant factor, contributing $1.33 million (65.9% of total market power channel growth). The increasing importance of manager ability (rising parameter alpha) contributes an additional $1.14 million through the market power channel. Within the firm size channel, TFP change accounts for 70.1% ($1.03 million) of growth, but the large effects from rising alpha and rising complementarity (gamma) are substantially offset by increasing dispersion in firm type. Structural estimates confirm that the average number of firms per market declines from 4.40 to 3.15, and firm-type dispersion (sigma_z) rises from 0.51 to 0.77, both consistent with rising market power over the period.&lt;/p&gt;
&lt;p&gt;A counterfactual economy with no market power — firms priced at marginal cost — would yield a social welfare gain of 58.4% on average. The welfare cost of market power in 1994 could be offset by a 33.8% TFP increase; by 2019 the required TFP offset had risen to 51.7%. Without any market power, even the most talented managers would earn only their reservation utility, because firms earn zero profits regardless of productivity, eliminating the complementarity-driven matching surplus that makes top managers valuable. This confirms that superstar manager pay is intrinsically tied to the existence of market power in goods markets, not solely to firm size.&lt;/p&gt;
&lt;p&gt;Scope conditions: the model applies to publicly listed US firms covered by Compustat and ExecuComp. The mechanism relies on Cournot competition within oligopolistic markets, assortative matching between managers and firms, and complementarity between manager ability and firm type (elasticity of substitution gamma estimated to be negative throughout the sample). The findings on market power share apply to CEOs specifically; the authors argue the same logic extends to all managerial positions with span-of-control over other workers, which encompasses roughly one-fifth of the workforce.&lt;/p&gt;
&lt;p&gt;Q: What are the two channels through which manager pay is determined in the model, and how do they differ mechanically?
A: The market power channel captures how a given level of TFP translates into higher markups — more productive firms charge more above marginal cost — thereby increasing profits per unit of output. The firm size channel captures how higher TFP expands the quantity of output a firm produces, increasing total profits through scale rather than through price-cost margin. Both channels raise profits and thus the marginal product of managers, but they operate through distinct economic mechanisms: one through pricing power and the other through productive scale.&lt;/p&gt;
&lt;p&gt;Q: What is the empirical magnitude of the market power channel&amp;rsquo;s contribution to manager pay levels and growth?
A: Market power accounts for an average of 45.8% of total manager pay over 1994–2019, rising monotonically from 38.0% in 1994 to 48.8% in 2019. For the total pay increase of $3.49 million over the period, $2.02 million (57.8%) is due to the increase in market power, with the remaining $1.47 million attributable to the firm size channel.&lt;/p&gt;
&lt;p&gt;Q: How does the market power channel&amp;rsquo;s importance vary across the manager ability distribution?
A: For the highest-ranked managers, 80.3% of total pay in 2019 is attributable to market power, and nearly all of their pay growth since 1994 runs through the market power channel. For the lowest-ranked managers, pay is almost entirely explained by the firm size channel and has been approximately flat over the period. This heterogeneity arises because top managers sort into high-markup firms through assortative matching, making their compensation disproportionately dependent on those firms&amp;rsquo; market power.&lt;/p&gt;
&lt;p&gt;Q: How does the model generate assortative matching between manager ability and firm type?
A: Manager ability and firm type are complementary inputs into TFP (the CES aggregator with elasticity of substitution gamma less than one), which makes the matching output supermodular. In a frictionless matching market with transferable utility, supermodularity guarantees that high-ability managers match with high-type firms in equilibrium (Proposition 1). This positive assortative matching then amplifies productivity and markup dispersion, since the most productive firms become even more productive and gain larger market shares.&lt;/p&gt;
&lt;p&gt;Q: What structural changes drive the rising importance of market power in manager pay over time?
A: The dominant factor within the market power channel is changes in technology, specifically increasing firm-type dispersion (sigma_z rising from 0.51 to 0.77), which contributes $1.33 million or 65.9% of market power channel growth. The rising importance of manager ability (alpha, the weight on manager ability relative to firm type in the TFP aggregator) contributes another $1.14 million. The number of firms per market declines from an average of 4.40 to 3.15, further reducing competitive pressure and amplifying the markup premium for high-productivity firms.&lt;/p&gt;
&lt;p&gt;Q: What does the counterfactual with no market power (first-best pricing) imply for manager pay and social welfare?
A: Without market power, firms price at marginal cost and earn zero profits regardless of productivity, which eliminates the surplus from manager-firm matching. All managers would earn only their reservation utility, which is negligible relative to actual compensation. Social welfare would increase by 58.4% on average. The efficiency cost of market power — measured as the TFP increase needed to offset welfare losses — rose from 33.8% in 1994 to 51.7% in 2019, indicating a worsening welfare distortion over the period.&lt;/p&gt;
&lt;p&gt;Q: How are markups measured, and what is their trend in the data?
A: Markups are not directly observable and are estimated using the production approach of De Loecker, Eeckhout, and Unger (2020), which recovers firm-level price-cost margins from production data without requiring price data. Average markups in the Compustat sample rose from 1.53 in 1994 to 1.78 in 2019. The reduced-form elasticity of manager pay with respect to markups (controlling for firm characteristics, year, and firm fixed effects) increased substantially: in 2019 a one-percent increase in firm-level markup raises manager pay by 0.41 percent, which is 70.1% larger than the effect estimated in 1994.&lt;/p&gt;
&lt;p&gt;Q: How does the paper handle the identification challenges inherent in regressing manager pay on markups?
A: The reduced-form regression (with firm fixed effects, year effects, and interactions of year dummies with markups) documents a robust positive correlation but cannot establish causality due to reverse causality and omitted-variable bias. The paper addresses this by embedding the markup-manager pay relationship in a structural model where both are jointly determined by primitives — technology, market structure, and manager ability — and estimating those primitives via Simulated Method of Moments. The quantitative decomposition into market power and firm size channels derives from the model structure rather than from identifying variation in an instrumental variables sense.&lt;/p&gt;
&lt;p&gt;Q: What do the matching model estimates reveal about manager-firm complementarity over time?
A: The estimated elasticity of substitution between manager ability and firm type (gamma) is negative throughout the sample, confirming complementarity. Gamma was relatively stable before declining sharply from -2.22 in 2014 to -3.55 in 2019, indicating that manager ability and firm type became substantially more complementary in the latter part of the sample. The importance-of-manager parameter alpha is small (consistent with Gabaix and Landier 2008) but generally increasing, suggesting managers play an expanding role in determining firm-level TFP over time.&lt;/p&gt;
&lt;p&gt;Q: What are the broader macroeconomic and distributional implications of the findings?
A: Because approximately one-fifth of workers supervise other workers, the market-power-driven premium in managerial pay has implications beyond CEO compensation for the shape of the earnings distribution. The rise in top-1-percent income is identified as an efficiency concern, not just an equity concern: the best managers are hired by high-markup firms where they generate profits for shareholders but disproportionately little additional social value. Assortative matching between top managers and top firms widens the productivity gap between competitors, increasing market power and deadweight loss — the social return to managerial talent is therefore below the private return in equilibrium.&lt;/p&gt;
&lt;p&gt;Market Power Channel: The component of manager pay attributable to how a firm&amp;rsquo;s TFP raises its markup — the ratio of output price to marginal cost — given the level of output. Distinct from the firm size channel; operates through pricing power rather than scale.&lt;/p&gt;
&lt;p&gt;Firm Size Channel: The component of manager pay attributable to how a firm&amp;rsquo;s TFP expands output quantity given markups. Increasing output scale raises total profits and thus the marginal product of the manager even absent any change in price-cost margins.&lt;/p&gt;
&lt;p&gt;Assortative Matching: The equilibrium allocation of high-ability managers to high-type firms, arising because manager ability and firm type are complementary inputs into TFP (supermodular matching output). Matching is determined in a frictionless market with transferable utility.&lt;/p&gt;
&lt;p&gt;Markup: The ratio of output price to marginal cost, equal to the inverse of the price elasticity of demand under the nested CES preference structure. Endogenously determined by the firm&amp;rsquo;s sales share within its oligopolistic market and the elasticities of substitution within markets (eta) and across markets (theta).&lt;/p&gt;
&lt;p&gt;Manager-Firm Complementarity: The property that manager ability and firm type are imperfect substitutes with elasticity of substitution gamma less than one in the TFP aggregator. Complementarity is the necessary condition for positive assortative matching and for the supermodularity of matching surplus.&lt;/p&gt;
&lt;p&gt;Span of Control (Lucas 1978): The mechanism by which a manager raises the productivity of all workers under supervision, so that a more able manager generates a proportionally larger productivity gain the larger the firm. Provides the microfoundation for why firm size amplifies the value of manager ability.&lt;/p&gt;
&lt;p&gt;Market Structure: The number of firms in each oligopolistic sub-market (Ij), which varies across markets and over time. Together with the distribution of firm-level TFP within a market, market structure determines how much competitive pressure limits markup extraction. Average firms per market declines from 4.40 to 3.15 over 1994–2019.&lt;/p&gt;</description></item><item><title>Marginal Propensity to Consume and Personal Characteristics: Evidence from Bank Transaction Data and Survey</title><link>https://macropaperwarehouse.com/papers/marginal-propensity-to-consume-and-personal-characteristics-evidence-from-bank-transaction-data-and-survey/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/marginal-propensity-to-consume-and-personal-characteristics-evidence-from-bank-transaction-data-and-survey/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; This paper asks whether heterogeneity in the marginal propensity to consume (MPC) stems from &lt;em&gt;temporary circumstances&lt;/em&gt; (e.g., transient wealth shocks that tighten liquidity) or &lt;em&gt;persistent personal characteristics&lt;/em&gt; (e.g., high time discount rates or strong risk aversion that permanently shape saving behavior). Because liquidity constraints are endogenous — they can reflect either bad luck or impatient preferences — disentangling these two sources requires independently measured individual characteristics, which are not available in standard transaction datasets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Setting.&lt;/strong&gt; The study combines two data sources drawn from Mizuho Bank, one of Japan&amp;rsquo;s three largest banks (approximately 24 million individual accounts). First, weekly bank account transaction data for January 2019 to November 2022 covering all outflows (ATM withdrawals, credit card debits, utility payments, interbank transfers) for the approximately 5,282 survey respondents. Second, a bespoke survey conducted in November–December 2022 among 400,000 randomly selected salary-receiving account holders (response rate 1.32%, yielding 5,282 usable observations). The survey elicits the Arrow–Pratt measure of absolute risk aversion, quantitative time discount rates for one-week, one-year, and ten-year horizons, self-reported liquidity constraints, homeownership, education, age, and gender, among other variables.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three Income Shocks.&lt;/strong&gt; MPC is estimated against three distinct income events: (1) the Japanese government&amp;rsquo;s Special Cash Payments (SCP) — a 100,000 JPY (approximately 800 USD) per-person lump-sum transfer during COVID-19, likely transitory, unexpected, and nearly randomly timed across municipalities due to administrative bottlenecks; (2) regular salary receipts (recurring, expected in both timing and amount); and (3) semi-annual bonus payments (received twice yearly, with timing known in advance but amount largely unknown — intermediate between SCP and salary in terms of expectedness).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Estimation Strategy.&lt;/strong&gt; A two-way fixed effects regression with event-study leads and lags (windows of five weeks before and after each income event) is used to estimate consumption responses. Individual and week fixed effects absorb time-invariant heterogeneity and aggregate shocks (including COVID-19 emergency declarations). Standard errors are clustered at the individual level. For heterogeneity analysis, the income shock variable is interacted with individual characteristics from the survey (treated as proxies for persistent characteristics) and with time-varying log wealth and a liquidity constraint dummy (wealth below one-twelfth of annual income, proxying temporary circumstances).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings — Average MPC.&lt;/strong&gt; Across all three income types, the on-impact MPC (week of receipt) is approximately 0.2: specifically γ₀ = 0.23 for the SCP (significant at 5%), 0.20 for salary, and 0.22 for bonus. When estimated jointly in a single regression, coefficients are γ_SCP = 0.21, γ_salary = 0.19, and γ_bonus = 0.21. This uniformity holds despite the sharply different properties of these shocks (transitory-unexpected vs. regular-expected vs. semi-known).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings — Heterogeneity.&lt;/strong&gt; Significant heterogeneity in MPC is found primarily in the bonus subsample, where statistical power is greatest. The following cross-term coefficients are significant at the 5% level in the multivariate specification: (a) &lt;em&gt;liquidity constraint dummy&lt;/em&gt; — positive and significant, indicating that individuals temporarily below one month&amp;rsquo;s income in deposits spend a larger fraction of their bonus, with a one standard deviation increase raising MPC by 0.094 (9.4 percentage points); (b) &lt;em&gt;time discount rate&lt;/em&gt; (quantitative measure) — positive and significant, with a one standard deviation increase in impatience raising MPC by 0.084; (c) &lt;em&gt;risk aversion&lt;/em&gt; (quantitative Arrow–Pratt measure) — positive and significant, conditional on controlling for wealth and liquidity, with a one standard deviation increase raising MPC by 0.031; (d) &lt;em&gt;education&lt;/em&gt; — negative and significant irrespective of wealth/liquidity controls, with a one standard deviation increase in education reducing MPC by 0.041.&lt;/p&gt;
&lt;p&gt;These magnitude estimates are sizable relative to the baseline MPC of approximately 0.2. For SCP and salary shocks, cross-term coefficients are uniformly insignificant at the 5% level, which the author attributes partly to smaller sample sizes and shorter observation windows for the SCP subsample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; The sample consists of Mizuho Bank account holders who receive salary payments directly into their Mizuho account, overrepresenting metropolitan areas and salaried workers relative to the national census. Wealth at Mizuho captures only deposits at that institution and excludes securities accounts, postal savings, and intra-household transfers. Age and gender do not yield significant cross-term coefficients in any specification; the self-reported survey measure of liquidity constraints (ability to cover one month&amp;rsquo;s income by drawing on savings, assets, or borrowing) is also insignificant, in contrast to the transaction-based liquidity constraint dummy.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-is-separating-temporary-circumstances-from-persistent-characteristics-important-for-mpc-estimation"&gt;Q1. Why is separating temporary circumstances from persistent characteristics important for MPC estimation?&lt;/h3&gt;
&lt;p&gt;Liquidity constraints — the standard proximate predictor of high MPC — are endogenous. An individual may be liquidity-constrained because of a temporary adverse income shock (bad luck) or because of persistently high impatience (high time discount rate) that leads to chronically low saving. If policy evaluation treats all constrained households symmetrically, it conflates these two very different channels. The paper follows Jappelli and Pistaferri (2020), Gelman (2021), and Aguiar, Bils, and Boar (2021) in arguing that both channels matter and that their relative contributions need empirical separation.&lt;/p&gt;
&lt;h3 id="q2-why-are-japanese-bonuses-particularly-well-suited-to-identifying-mpc-heterogeneity"&gt;Q2. Why are Japanese bonuses particularly well-suited to identifying MPC heterogeneity?&lt;/h3&gt;
&lt;p&gt;Bonuses are paid semi-annually to most regular employees in Japan (accounting for roughly 15–30% of annual income), with timing known in advance but amount largely unknown until receipt. This intermediate nature — partially anticipated in timing but uncertain in magnitude — provides meaningful variation in consumption responses across individuals while maintaining a clean event-study design. The bonus subsample (3,722 individuals who received a bonus at least once) is also large enough to detect cross-term effects that are statistically insignificant in the SCP subsample (2,446 individuals) and in the salary analysis, likely due to greater statistical power.&lt;/p&gt;
&lt;h3 id="q3-how-is-the-arrowpratt-measure-of-risk-aversion-constructed-from-the-survey"&gt;Q3. How is the Arrow–Pratt measure of risk aversion constructed from the survey?&lt;/h3&gt;
&lt;p&gt;Respondents are asked whether they would purchase a lottery ticket at prize value Z = 100,000 JPY and price p = 10,000 JPY for varying winning probabilities α. The threshold α at which a respondent switches from accepting to rejecting identifies their risk attitude. The absolute risk aversion σ = −U&amp;rsquo;&amp;rsquo;/U&amp;rsquo; is then calculated as (αZ² − 2αZp + p²) / (2(αZ − p)). This yields σ ranging from −4.5 (when α = 0.01, i.e., risk-loving) to 0.891 (when α = 1, i.e., refusing to buy even at a 90% win probability). Risk neutrality corresponds to σ = 0 (at α = 0.1).&lt;/p&gt;
&lt;h3 id="q4-how-are-time-discount-rates-measured-and-what-is-the-range"&gt;Q4. How are time discount rates measured, and what is the range?&lt;/h3&gt;
&lt;p&gt;Respondents are asked the minimum amount X they would require to wait one week, one year, or ten years to receive a payment instead of receiving 100,000 JPY one week from now (using a one-week anchor to address hyperbolic discounting). The discount rate is calculated as r = X/100,000. The range is 0.01 (X = 100 JPY) to 100 (X = 10,000,000 JPY, i.e., would not wait even for 1,100,000 JPY in ten years). The unweighted average across one-week, one-year, and ten-year horizons is used as the composite discount rate in the multivariate specifications.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-transaction-based-liquidity-constraint-dummy-and-how-does-it-differ-from-the-survey-based-measure"&gt;Q5. What is the transaction-based liquidity constraint dummy, and how does it differ from the survey-based measure?&lt;/h3&gt;
&lt;p&gt;The transaction-based dummy equals one if end-of-month deposits at Mizuho Bank (the previous month) are below one-twelfth of the individual&amp;rsquo;s annual income — i.e., if the individual holds less than one month&amp;rsquo;s equivalent income in liquid deposits. This is a time-varying measure. The survey-based measure asks respondents to self-report whether they could cover one month&amp;rsquo;s income by drawing on savings, selling assets, or borrowing. The transaction-based measure is significant at the 5% level in the bonus and salary heterogeneity regressions, while the survey-based measure is insignificant, indicating that the precise definition and data source of the liquidity constraint measure matters materially for detecting its effect on MPC.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-estimated-on-impact-mpc-values-for-each-income-shock-and-how-stable-are-they-across-robustness-checks"&gt;Q6. What are the estimated on-impact MPC values for each income shock, and how stable are they across robustness checks?&lt;/h3&gt;
&lt;p&gt;The point estimates from the event-study regression (γ₀) are: 0.23 for SCP in the baseline sample (SCP recipients in 2020, N = 2,446 individuals), 0.20 for salary (all 5,282 survey respondents), and 0.22 for bonus (3,722 bonus recipients). In a robustness specification restricting to only year-2020 data for the SCP, γ₀ = 0.235; using cash withdrawals from ATMs as a proxy for consumption instead of total outflows, γ₀ = 0.162 for SCP. In a joint regression including all three income types simultaneously, γ_SCP = 0.21, γ_salary = 0.19, and γ_bonus = 0.21. The SCP MPC for the smaller second-wave subsample (200 individuals, 2021–22) is 0.104 and insignificant, consistent with insufficient statistical power rather than a structural difference.&lt;/p&gt;
&lt;h3 id="q7-why-is-the-similarity-in-mpc-across-the-three-shock-types-potentially-surprising-and-what-does-the-paper-say-about-it"&gt;Q7. Why is the similarity in MPC across the three shock types potentially surprising, and what does the paper say about it?&lt;/h3&gt;
&lt;p&gt;Standard theory predicts divergent MPCs: transitory unexpected windfalls (SCP) should have a higher MPC than permanent salary changes under the permanent income hypothesis, while Ricardian equivalence might reduce the MPC to fiscal transfers like the SCP if households anticipate future tax increases. The paper finds the MPCs are approximately equal (around 0.2 across all three types), and if anything the SCP MPC is slightly higher than the salary MPC. The paper acknowledges this uniformity without offering a structural explanation, using it primarily as a robustness check on the baseline estimate rather than a substantive puzzle to resolve.&lt;/p&gt;
&lt;h3 id="q8-which-personal-characteristics-are-significantly-associated-with-higher-mpc-and-in-which-income-shock-samples"&gt;Q8. Which personal characteristics are significantly associated with higher MPC, and in which income shock samples?&lt;/h3&gt;
&lt;p&gt;In the multivariate heterogeneity regression, significant cross-term coefficients at the 5% level are found exclusively in the bonus subsample (columns 5–6 of Table 6): the quantitative risk aversion measure (positive, coefficient 0.042–0.049), the quantitative discount rate (positive, coefficient 0.004), and education (negative, coefficient −0.034 to −0.037). The liquidity constraint dummy (transaction-based) is also positive and significant for bonuses. In the univariate robustness regressions (Table 7), the own-house dummy is negative and significant at 5% for bonuses (controlled and uncontrolled); discount rates for one-week and ten-year horizons are positive and significant at 5% for bonuses; risk aversion A (direct self-report) is negative and significant at 5% for SCPs in the uncontrolled specification.&lt;/p&gt;
&lt;h3 id="q9-do-age-and-gender-matter-for-mpc-heterogeneity"&gt;Q9. Do age and gender matter for MPC heterogeneity?&lt;/h3&gt;
&lt;p&gt;No. In all specifications across all three income shock types, the cross-term coefficients on age and the male dummy are uniformly insignificant at the 5% level. The lack of significance for age and gender is noted as a notable result, since both are commonly used demographic proxies in heterogeneous agent models that assume they reflect economically meaningful differences in consumption behavior.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-paper-quantify-the-economic-magnitude-of-each-significant-heterogeneity-factor"&gt;Q10. How does the paper quantify the economic magnitude of each significant heterogeneity factor?&lt;/h3&gt;
&lt;p&gt;Table 8 reports the product of each cross-term coefficient and the standard deviation of the corresponding variable. For the bonus subsample: a one standard deviation increase in the liquidity constraint dummy raises MPC by 0.094 (9.4 percentage points); a one standard deviation increase in the discount rate raises MPC by 0.084; a one standard deviation increase in risk aversion raises MPC by 0.031; and a one standard deviation increase in education reduces MPC by 0.041. All four magnitudes are described as sizable relative to the baseline MPC of approximately 0.2 (20%).&lt;/p&gt;
&lt;h3 id="q11-why-does-the-paper-focus-on-bonuses-for-the-heterogeneity-analysis-rather-than-the-scp"&gt;Q11. Why does the paper focus on bonuses for the heterogeneity analysis rather than the SCP?&lt;/h3&gt;
&lt;p&gt;The SCP events provide cleaner identification of transitory, exogenous income shocks (near-random timing due to municipal administrative bottlenecks, as documented by Kubota, Onishi, and Toyama 2021), but the subsample of SCP recipients is smaller (2,446 in 2020, 200 in the second wave), reducing statistical power for detecting heterogeneity in cross-term coefficients. The salary sample is large (5,282 individuals) but salaries are expected, recurring, and may partially update permanent income, complicating interpretation of cross-term estimates. Bonuses offer a balance: a relatively large subsample (3,722) and a partially unexpected income component, making them the most informative sample for heterogeneity analysis.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-main-caveats-and-limitations-the-paper-identifies"&gt;Q12. What are the main caveats and limitations the paper identifies?&lt;/h3&gt;
&lt;p&gt;Four caveats are noted. First, the personal characteristics from the survey — including time discount rates and risk aversion — are treated as exogenous, but they may themselves be endogenous to economic circumstances or short-term conditions at the time of the survey. Second, only Mizuho Bank deposits are observed; financial assets at other institutions (securities, postal savings) are missing, meaning the liquidity constraint measure understates true wealth for some respondents. Third, the sample is tilted toward metropolitan salaried workers and toward wealthier individuals compared to the full Mizuho customer base (median log wealth of 7.4 vs. 5.9 in Kubota et al. 2021). Fourth, the multiple-testing problem is acknowledged: with many cross-term tests conducted, some rejections of the null at the 5% level may be spurious.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Marginal Propensity to Consume (MPC, on-impact).&lt;/strong&gt; In this paper, MPC is operationalized as the coefficient γ₀ from the two-way fixed effects event-study regression — specifically, the fraction of an income shock spent during the &lt;em&gt;same week&lt;/em&gt; the shock is received, estimated from total bank account outflows. This is a weekly, within-account measure, not a lifetime or annual consumption response.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Arrow–Pratt Absolute Risk Aversion (σ).&lt;/strong&gt; A quantitative measure of risk preferences computed from the paper&amp;rsquo;s survey by eliciting the probability threshold α at which a respondent is indifferent between buying and not buying a lottery with prize Z = 100,000 JPY and price p = 10,000 JPY. Calculated as σ = (αZ² − 2αZp + p²) / (2(αZ − p)). Ranges from −4.5 to 0.891 in the sample, with σ = 0 indicating risk neutrality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time Discount Rate (r).&lt;/strong&gt; Measured by asking respondents the minimum additional amount X (beyond 100,000 JPY) they would require to delay receipt by one week, one year, or ten years, with r = X/100,000. The paper uses the unweighted average of three horizon-specific rates as a composite measure. Ranges from 0.01 to 100 in the sample. Used as a proxy for impatience or myopia — a persistent personal characteristic.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Liquidity Constraint Dummy (transaction-based).&lt;/strong&gt; A time-varying binary indicator that equals one if individual i&amp;rsquo;s end-of-month Mizuho Bank deposit balance in month t−1 is below one-twelfth of annual income at t−1 — i.e., less than one month&amp;rsquo;s equivalent income in liquid deposits. Distinguished in the paper from a survey-based self-report of liquidity constraints, which is found to be insignificant.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Special Cash Payment (SCP).&lt;/strong&gt; The Japanese government&amp;rsquo;s COVID-19 pandemic transfer program, providing 100,000 JPY (approximately 800 USD) per person in 2020 (universal) and 100,000 JPY per child in 2021–22 (restricted to households with children under 18 and income below 9.6 million JPY annually). Used in this paper as a transitory, salient, and largely unexpected income shock because municipal administrative bottlenecks made the exact timing unpredictable and nearly random across households.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Two-Way Fixed Effects Event-Study Regression.&lt;/strong&gt; The paper&amp;rsquo;s primary estimator, which includes individual fixed effects (controlling for time-invariant person-level heterogeneity) and week fixed effects (absorbing aggregate shocks such as COVID-19 emergency declarations and seasonal patterns). Event-study leads and lags (k = −5 to +5 weeks around each income receipt) allow pre-trend testing and tracing of the dynamic consumption response. Normalized to γ_{−1} = 0.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MPC Heterogeneity Cross-Term.&lt;/strong&gt; A regression augmentation (equation 3 in the paper) in which the contemporaneous income shock X⁰_{it} is interacted with individual characteristic Z_{it}. The coefficient δ on this cross-term identifies how the MPC varies with Z — the marginal effect of characteristic Z on the MPC. Persistent characteristics (e.g., risk aversion, discount rate, education from the survey) and temporary circumstances (e.g., log wealth, liquidity constraint dummy from transaction data) are included as separate Z variables.&lt;/p&gt;</description></item><item><title>Marginal Returns to Public Universities</title><link>https://macropaperwarehouse.com/papers/marginal-returns-to-public-universities/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/marginal-returns-to-public-universities/</guid><description>&lt;p&gt;This paper asks whether enrolling in an American public university generates positive net returns for marginal students — those who barely qualify for admission — and whether those returns justify public expenditures. The question is policy-relevant because marginal students have weak academic preparation, face high dropout risk, and the net returns to expanding admission margins are theoretically ambiguous.&lt;/p&gt;
&lt;p&gt;The author assembles administrative records spanning all 35 public universities in Texas, covering the universe of Texas public high school graduates from 2004–2014 (approximately 2.7 million students). Texas public universities collectively enroll over 10 percent of all American public university students. The data link high school records (test scores, demographics, coursework, attendance, disciplinary infractions) to college application and admission records, postsecondary enrollment and degree completion records, financial aid packages, institutional expenditure data from IPEDS, and quarterly earnings records from the Texas Workforce Commission unemployment insurance system.&lt;/p&gt;
&lt;p&gt;The identification strategy exploits hundreds of decentralized SAT/ACT score cutoffs in university admissions — varying across schools and application years — that generate sharp discontinuities in admission probability. A fuzzy regression discontinuity design compares applicants just above versus just below each cutoff. On average, crossing a cutoff raises the probability of admission by 27 percentage points and the probability of enrolling at the target university by 15 percentage points. Density tests and pre-college covariate balance validate the smoothness assumptions. The typical cutoff complier is more disadvantaged than the average college applicant but comparable to the average Texas high school graduate.&lt;/p&gt;
&lt;p&gt;Roughly half of cutoff compliers would fall back to another, typically less selective, four-year institution if rejected; 43 percent would fall back to a two-year community college; and only about 6 percent would forgo higher education entirely. The pooled estimates therefore blend intensive-margin effects (more selective versus less selective four-year college) with extensive-margin effects (four-year college versus community college or no college).&lt;/p&gt;
&lt;p&gt;Main causal findings for enrollment compliers: the typical marginally admitted student completes approximately one additional year of credits in the four-year sector and becomes 12 percentage points more likely to ever earn a bachelor&amp;rsquo;s degree from any institution. About half of the additional four-year credits are offset by 15 fewer credits in the two-year sector, and associate degree or certificate completion falls by 7 percentage points. All bachelor&amp;rsquo;s degree gains are in non-STEM fields; STEM degree completion shows no detectable increase. Compliers become about 3 percentage points more likely to hold a graduate degree by 10 years out.&lt;/p&gt;
&lt;p&gt;On earnings, admitted compliers earn less than rejected counterparts in the first five years due to continued enrollment. Year six is the crossover point; by years 8–12, compliers earn a stable 8.6 percent earnings premium in log terms (8.2 percent in dollar ratio terms, representing a LATE of $3,339 against an untreated complier mean of $40,829), with earnings ranks rising approximately 4 percentiles from a base near the 50th percentile.&lt;/p&gt;
&lt;p&gt;Marginally admitted students pay no additional net tuition on average: $4,600 in additional gross tuition is nearly fully offset by grant aid, though they take on $5,300 more in student loans. Society incurs approximately $10,000 in additional educational expenditures per complier. Internal rates of return are 26 percent for students, 16 percent for society, and 7 percent for the government budget. At a 3 percent discount rate, the lifetime net present value of enrolling the typical marginal applicant is approximately $80,000 — $70,000 accruing to the student and $10,000 to taxpayers.&lt;/p&gt;
&lt;p&gt;Earnings gains are similar across institutions of varying selectivity, but significantly smaller for low-income compliers, who spend more time enrolled, complete fewer degrees, and major in less lucrative fields. A bounding method shows that extensive-margin compliers (those who would otherwise not attend any four-year college) experience larger effects than intensive-margin compliers.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question and why is credible evidence scarce?
A: The paper asks whether enrolling marginal students in American public universities generates positive net returns — private, social, and fiscal — and what drives heterogeneity in those returns. Credible evidence is scarce because most existing work is correlational and fails to account for selection bias: individuals with more college education may have had pre-existing advantages, confounding college&amp;rsquo;s causal effect with systematic sorting into it. Even if average returns are positive, the policy-relevant question is whether the marginal student — who has weak preparation and high dropout risk — represents a good investment.&lt;/p&gt;
&lt;p&gt;Q: What is the regression discontinuity design, and what does the first stage look like?
A: The author infers hundreds of decentralized SAT/ACT score cutoffs across approximately 700 application cells (combinations of university, year, GPA quartile, and test type) by searching for the score value with the largest discontinuity in admission and enrollment within each cell. This procedure delivers a superconsistent estimator of each cell&amp;rsquo;s true cutoff. Pooled across all cells, crossing a cutoff raises the probability of admission by 27 percentage points and the probability of enrollment at the target university by a precisely estimated 15 percentage points. The density of applicants and a rich set of pre-college characteristics run smoothly through the cutoffs, supporting the exclusion restriction.&lt;/p&gt;
&lt;p&gt;Q: Who are the cutoff compliers, and are they representative of any broader population?
A: Compliers — applicants who enroll in the target university if and only if they barely cross its cutoff — comprise approximately 15 percent of marginal applicants. In observable characteristics, compliers are roughly representative of the broader population of marginal applicants at the cutoff. They are significantly more disadvantaged than the average public university applicant, but broadly comparable to the average Texas public high school graduate in terms of academic preparation and family income.&lt;/p&gt;
&lt;p&gt;Q: What are the next-best alternatives for marginal applicants who are rejected?
A: Approximately 47 percent of compliers would fall back to another Texas four-year college (mostly public), 43 percent to a two-year community college, and approximately 9 percent would not enroll in any Texas institution. National Student Clearinghouse data for the 2008–2014 cohorts confirm that only 4 percent of untreated compliers attend a college outside the THECB universe, meaning approximately 6 percent of all compliers truly forgo higher education altogether if rejected. The empirically relevant extensive margin is therefore between the four-year sector and the two-year sector, not between college and no college.&lt;/p&gt;
&lt;p&gt;Q: How does cutoff crossing change the institutional characteristics a complier experiences?
A: Compliers are propelled into substantially better-resourced environments: the average math test score of college peers rises by half a standard deviation; peers are 12 percentage points less likely to have been low-income; gross tuition rises by $2,400 (a 42 percent increase over the untreated complier mean of $5,700); educational spending per student rises by $3,200 (43 percent over the untreated mean); peers&amp;rsquo; 10-year BA completion rate rises by 28 percentage points; and peer mean earnings 8–12 years after college entry are $6,700 higher.&lt;/p&gt;
&lt;p&gt;Q: What are the educational attainment effects?
A: Cutoff crossing causes compliers to complete approximately 28 additional credits at any four-year institution (roughly one full year of a four-year program) and increases the probability of ever earning a bachelor&amp;rsquo;s degree by 12 percentage points, raising the completion rate from approximately 40 percent to just above 50 percent. About 15 fewer two-year sector credits are offset against the four-year gains, and associate degree or certificate completion falls by 7 percentage points. All bachelor&amp;rsquo;s degree gains are in non-STEM fields; there is no detectable increase in STEM degrees. Graduate degree completion rises by approximately 3 percentage points by 10 years out.&lt;/p&gt;
&lt;p&gt;Q: What is the earnings trajectory, and when does the premium materialize?
A: Admitted compliers earn less than rejected counterparts in the first five years after application because they remain enrolled longer. Year six is the crossover point. By years 8–12, the earnings premium stabilizes at approximately 8.6 percent in log terms and 8.2 percent in dollar ratio terms (a LATE of $3,339 against an untreated complier mean of $40,829). Earnings rank rises by approximately 4 percentiles from a base near the 50th percentile. These results are robust across sandwich earnings, all-quarters-with-earnings, and zero-imputed specifications.&lt;/p&gt;
&lt;p&gt;Q: What does the cost-benefit analysis show?
A: Marginally admitted students pay no additional net tuition on average: $4,600 in additional gross tuition is nearly fully offset by additional grant aid. They do borrow $5,300 more in student loans, likely financing higher room, board, and consumption costs at four-year colleges. From society&amp;rsquo;s perspective, compliers generate approximately $10,000 in additional educational expenditures. Cumulative undiscounted earnings benefits surpass costs after 8 years for students, 11 years for society, and 19 years for taxpayers. At a 3 percent discount rate, the lifetime net present value is approximately $80,000 total — $70,000 accruing to the student and $10,000 to taxpayers — with internal rates of return of 26 percent for students, 16 percent for society, and 7 percent for the government budget.&lt;/p&gt;
&lt;p&gt;Q: Does selectivity of the admitting institution predict larger earnings returns?
A: No. Compliers at more selective institutions experience substantially larger increases in peer quality than those at less selective institutions, but they are also less likely to be on the extensive margin of four-year enrollment and experience smaller BA attainment gains. These factors roughly offset, producing no systematic difference in earnings gains across institutions of varying selectivity. More selective institutions also impose no additional cumulative cost on society, while compliers actually pay slightly less in additional net tuition at more selective schools.&lt;/p&gt;
&lt;p&gt;Q: How does the commonly used measure of college value-added (mean peer earnings) compare to actual complier returns?
A: Mean peer earnings overpredicts actual value-added for marginal students by a factor of two: compliers attend an institution with $6,700 higher average peer earnings as a result of admission but gain only $3,300 themselves. The measure also overpredicts the earnings return to selectivity by a factor of three: a 100-SAT-point increase in target school selectivity predicts $3,000 higher peer earnings but only a statistically insignificant $900 higher gain in the complier&amp;rsquo;s own earnings.&lt;/p&gt;
&lt;p&gt;Q: How do earnings returns differ by family income?
A: Compliers from low-income families experience significantly smaller earnings gains compared to higher-income compliers. The gap is not explained by differential changes in college quality induced by admission. Instead, low-income compliers gain fewer degrees despite spending more time in college and major in less lucrative fields, consistent with related findings in the literature on family income gaps in degree completion and major choice.&lt;/p&gt;
&lt;p&gt;Q: How do earnings returns differ by gender and by race?
A: Female and male compliers eventually earn similar log earnings and earnings rank gains, but women reach their gains more quickly — likely because men take longer to finish college. White and Asian compliers experience similar earnings gains and BA completion improvements as Black and Hispanic compliers, despite white and Asian students experiencing larger increases in college selectivity and spending per student as a result of admission.&lt;/p&gt;
&lt;p&gt;Q: What is the method for separating intensive- and extensive-margin effects?
A: The two complier types are not directly distinguishable in the data. The author first uses an endogenous but strong stratification variable — having at least one other Texas public university admission offer — to identify some mean potential outcomes for each type. He then imposes an empirically-informed rank assumption to bound the remaining unknown mean potential outcomes, delivering tightly informative upper and lower bounds on each margin&amp;rsquo;s effects without requiring full nonparametric identification. The results show that pooled effects are driven by larger returns for extensive-margin compliers who would not have attended any four-year college, with smaller contributions from intensive-margin compliers shifting between four-year institutions.&lt;/p&gt;
&lt;p&gt;Q: How do this paper&amp;rsquo;s earnings estimates compare to prior studies, and what explains the differences?
A: This paper&amp;rsquo;s 8 percent earnings gain is smaller than the 17–26 percent reported in prior studies (Zimmerman 2014: 22%; Kozakowski 2023: 26%; Smith, Goodman, and Hurwitz 2025: 17%; Bleemer 2024: 21%; Hoekstra 2009: 20%). The differences are likely explained by the much larger educational attainment and institutional quality gains induced by those studies&amp;rsquo; natural experiments: in Zimmerman (2014), enrollment compliers gain roughly three additional years of four-year education versus one year in this paper; in Bleemer (2024), compliers experience roughly $30,000 more in institutional spending per student versus approximately $3,000 in this paper.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions for these results?
A: The results pertain to marginal applicants to Texas public universities (excluding UT-Austin, which uses holistic admission with no detectable SAT/ACT cutoffs) from the 2004–2014 high school graduation cohorts. The identified effects are local average treatment effects for compliers — applicants who would enroll in the target university if and only if they barely crossed its admission cutoff — and do not represent effects for always-takers or infra-marginal students. Earnings are measured only for Texas-based workers covered by the state unemployment insurance system, which captures an estimated 90 percent of the civilian labor force.&lt;/p&gt;
&lt;p&gt;Cutoff complier: An applicant who enrolls in their target university if and only if their SAT/ACT score barely exceeds that university&amp;rsquo;s admission cutoff. Compliers are the population whose behavior — and thus whose treatment effects — are identified by the fuzzy RD design. They comprise approximately 15 percent of marginal applicants and are more disadvantaged than the average public university applicant but broadly comparable to the average high school graduate.&lt;/p&gt;
&lt;p&gt;Extensive versus intensive margin: The extensive margin refers to the contrast between attending any four-year college versus falling back to a two-year community college or no college. The intensive margin refers to the contrast between attending a more selective versus a less selective four-year institution. Approximately half of cutoff compliers are on each margin; the paper treats them as economically distinct parameters requiring separate identification.&lt;/p&gt;
&lt;p&gt;Fuzzy regression discontinuity (RD) design: An identification strategy that uses the discontinuous jump in admission probability at a test score cutoff as an instrument for enrollment, recovering the LATE for compliers via the ratio of the reduced-form discontinuity in outcomes to the first-stage discontinuity in enrollment. &amp;ldquo;Fuzzy&amp;rdquo; refers to the fact that crossing the cutoff changes admission and enrollment probabilities with a discrete jump rather than with certainty.&lt;/p&gt;
&lt;p&gt;Internal rate of return (IRR): The discount rate at which the net present value of an investment equals zero — here, the discount rate equating the discounted stream of earnings benefits to the discounted stream of costs. The paper estimates IRRs separately for students (26 percent), society (16 percent), and the government budget (7 percent), reflecting different cost and benefit definitions from each perspective.&lt;/p&gt;
&lt;p&gt;Rank assumption (bounding method): An empirically-informed assumption about the ordering of mean potential outcomes across latent complier types (extensive vs. intensive margin) that, combined with partial identification from a strong endogenous stratification variable, yields tight upper and lower bounds on each margin&amp;rsquo;s causal effects without requiring full nonparametric identification.&lt;/p&gt;
&lt;p&gt;Net tuition: Gross tuition charges minus grant aid. For the typical marginal complier, gross tuition rises by $4,600 but is nearly fully offset by additional grant aid, yielding approximately zero additional net tuition cost — meaning the private financial cost of attending a public university for marginal students is effectively zero on net, though they take on $5,300 more in student loans to finance room, board, and consumption.&lt;/p&gt;
&lt;p&gt;Sandwich earnings measure: A procedure applied to quarterly state earnings data that retains only quarters with positive earnings sandwiched between other quarters with positive earnings, discarding high-variance transition quarters between employment spells. Annualized by multiplying the quarterly average by four; used to reduce noise from entry and exit transitions in administrative earnings records.&lt;/p&gt;</description></item><item><title>Marriage, Fertility, and Cultural Integration in Italy</title><link>https://macropaperwarehouse.com/papers/marriage-fertility-and-cultural-integration-in-italy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/marriage-fertility-and-cultural-integration-in-italy/</guid><description>&lt;p&gt;Bisin and Tura study the cultural integration of immigrants in Italy by estimating a structural model of marital matching embedded with intra-household decisions — fertility, socialization of children, and divorce — along cultural-ethnic lines. The central research question is how to decompose the demand for integration (from immigrants) and the supply of cultural acceptance (from natives) in explaining the pace and heterogeneity of cultural convergence.&lt;/p&gt;
&lt;p&gt;The empirical analysis exploits administrative individual-level data from ISTAT&amp;rsquo;s ADELE Laboratory covering the universe of marriages formed in Italy from 1995 to 2012 and the universe of births and separations over the same period. After matching marriage, birth, and separation records, the final sample comprises more than 4 million marriages, representing 92.6% of all marriages celebrated in Italy over the period. Seven cultural-ethnic groups are studied: Italian (majority), Europe-EU15, Other Europe, North Africa–Middle East, Sub-Saharan Africa, East Asia, and Latin America. The model is a transferable-utility (TU) frictionless marriage market in which the joint marital surplus depends on a systematic component — itself the outcome of a collective household decision problem — and an idiosyncratic component capturing unobserved individual heterogeneity (following Choo and Siow, 2006). Parameters are estimated via method of moments, with identification drawing on cross-sectional variation across ethnic-group pairings and across Italy&amp;rsquo;s 20 administrative regions. Cultural socialization is proxied by language transmission (whether Italian is spoken at home with children).&lt;/p&gt;
&lt;p&gt;The data confirm strong positive assortative mating along cultural-ethnic lines, with particularly high homogamy rates for Sub-Saharan African and East Asian minorities. Homogamous minority households show notably lower rates of Italian-language use at home — for East Asian parents, 20% in a homogamous marriage versus 92% in a heterogamous marriage. Heterogamous marriages have higher separation rates (7.5% for mixed families with at least one Italian spouse versus 6.4% for homogamous Italian couples) and lower fertility.&lt;/p&gt;
&lt;p&gt;The estimated cultural intolerance parameters — measuring the psychological value a parent places on socializing a child to his/her own ethnic identity relative to a child acquiring a different identity — are strictly positive, asymmetric across directions, and highly heterogeneous across groups. North Africa–Middle East immigrants exhibit the highest minority intolerance (estimated at 97.85), more than six times that of Europe-EU15 immigrants (6.69). Latin America (93.13), Sub-Saharan Africa (87.08), and East Asia (81.22) also show high intolerance. On the native side, Italian intolerance is highest toward Sub-Saharan African immigrants (78.23) and lowest toward Europe-EU15 immigrants.&lt;/p&gt;
&lt;p&gt;Long-run simulations over successive generations show that all minorities eventually converge to the Italian majority along the language dimension, but at heterogeneous rates. Seventy-five percent of second-generation immigrants speak Italian at home with their children (one-generation integration rate). Europe-EU15 and Other Europe minorities converge almost completely within a single generation. Latin America shows the slowest path, with only 70% integration after four generations. East Asia and Sub-Saharan Africa also integrate more slowly, driven respectively by high fertility rates and strong selection into homogamous marriages.&lt;/p&gt;
&lt;p&gt;A counterintuitive counterfactual result is central to the paper: if Italian cultural intolerance were reduced to zero (full acceptance), cultural integration of minorities would slow by 15 percentage points over a generation (from 93% to 78% by the third generation). The mechanism is that greater native acceptance enables immigrants to sustain their own language even within heterogamous (mixed) marriages, increasing demand for such marriages and raising minority fertility, thereby preserving cultural distinctiveness.&lt;/p&gt;
&lt;p&gt;Finally, doubling immigration inflows while holding population shares constant reduces third-generation integration from 93% to 86% (a 7-percentage-point reduction). Effects are concentrated among Sub-Saharan African (20-percentage-point reduction) and East Asian (6-percentage-point reduction) minorities, with little impact on European and North African minorities. When inflows are reweighted toward Sub-Saharan African and East Asian groups, integration losses for those minorities range from 20 to 60 percentage points by the third generation.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s core methodological contribution?
A: The paper embeds a collective household decision problem — covering fertility, socialization, and divorce — within a transferable-utility frictionless marriage matching framework. This allows marital utility to emerge endogenously from intra-household decisions rather than being specified exogenously. The key innovation is that socialization incentives and technologies differ systematically between homogamous and heterogamous marriages, and these differences feed back into marital matching and long-run cultural dynamics.&lt;/p&gt;
&lt;p&gt;Q: What does &amp;ldquo;cultural intolerance&amp;rdquo; mean in this model, and how is it identified?
A: Cultural intolerance is the psychological value a parent obtains from socializing a child to his/her own ethnic identity, relative to having a child adopt a different cultural-ethnic identity. It is the main parameter driving socialization effort and resistance to cultural integration. Identification relies on two sources of cross-sectional variation: differences in matching patterns, fertility, separation, and socialization rates across cultural-ethnic group pairings, and exogenous variation in the ethnic composition of the regional population across Italy&amp;rsquo;s 20 administrative regions.&lt;/p&gt;
&lt;p&gt;Q: How heterogeneous are the estimated cultural intolerance parameters across minority groups?
A: The parameters are highly heterogeneous. North Africa–Middle East immigrants have the highest estimated minority intolerance (97.85), more than six times the EU15 estimate (6.69). Latin America (93.13), Sub-Saharan Africa (87.08), and East Asia (81.22) are also substantially higher than EU15. The matrix is asymmetric: Italian intolerance toward Sub-Saharan Africans (78.23) is higher than toward North Africans (67.88), even though those two groups show comparable minority intolerance levels.&lt;/p&gt;
&lt;p&gt;Q: What are the three mechanisms beyond intolerance parameters that explain heterogeneous integration dynamics?
A: First, selection into homogamous marriages: Sub-Saharan Africa&amp;rsquo;s particularly strong selection into homogamy gives those households access to superior coordinated socialization technology, sustaining cultural heterogeneity despite similar intolerance levels to other groups. Second, fertility rates: East Asian minorities have particularly high estimated fertility, which amplifies the transmission of their cultural identity across generations. Third, socialization effectiveness in heterogamous marriages: Latin American immigrants are uniquely able to socialize children to their own language even when married to native Italians, making their integration the slowest despite being in many mixed marriages.&lt;/p&gt;
&lt;p&gt;Q: What is the counterintuitive result about Italian cultural intolerance and integration speed?
A: Lowering Italian cultural intolerance to zero would reduce minority integration by 15 percentage points over one generation, with third-generation integration falling from 93% to 78%. The intuition is that higher native acceptance enables immigrants to maintain their own language more effectively within heterogamous marriages, which in turn increases immigrant demand for intermarriage with natives and raises minority fertility — both of which slow cultural convergence rather than accelerating it.&lt;/p&gt;
&lt;p&gt;Q: How do divorce dynamics differ between homogamous and heterogamous households?
A: Heterogamous households exhibit higher separation rates than culturally homogeneous unions: 7.5% for mixed families with at least one Italian spouse versus 6.4% for homogamous Italian couples. In the model, divorce by heterogamous households can be a strategic choice by mothers with high cultural intolerance, since custody grants single mothers greater unilateral control over socialization. Divorce probabilities are decreasing in the number of children for both family types. Interestingly, heterogamous households invest more in socialization when divorced than when married, because the high-intolerance parent can act without spousal opposition.&lt;/p&gt;
&lt;p&gt;Q: How well does the model fit the data?
A: The raw correlation between predicted and observed gains to marriage is 0.84. The correlation between predicted and observed foreign-language socialization rates is 0.83, for both homogamous and heterogamous families. The dataset covers 92.5% of all marriages in Italy from 1995 to 2012, representing over 4 million marriages matched with birth and separation records at a 98.5% one-to-one match rate.&lt;/p&gt;
&lt;p&gt;Q: What happens to cultural integration when immigration inflows are doubled with an overweighting of North Africa–Middle East, Sub-Saharan Africa, and East Asian immigrants?
A: North Africa–Middle East immigrants reduce third-generation convergence by only 4 percentage points. By contrast, East Asian and Sub-Saharan African minorities produce integration losses ranging from 20 to 60 percentage points by the third generation. This wide range reflects how the interaction between high fertility, strong homogamy selection, and effective socialization in heterogamous marriages amplifies cultural persistence when these groups constitute a larger share of inflows.&lt;/p&gt;
&lt;p&gt;Q: What is the one-generation cultural integration rate, and which groups diverge most from it?
A: Seventy-five percent of second-generation immigrants speak Italian at home with their children, constituting the one-generation baseline integration rate. Europe-EU15 and Other Europe minorities converge almost completely within one generation, as does North Africa–Middle East. Latin America diverges most sharply downward, with only 70% integration even after four generations, and shows a partial retreat from integration in the first generation. Sub-Saharan Africa and East Asia also fall below the 75% one-generation benchmark.&lt;/p&gt;
&lt;p&gt;Q: How does the paper relate to the debate on native labor market effects of immigration?
A: The paper notes that sizeable negative labor market effects of immigration on natives are far from well-documented in the empirical literature, with results ranging from negative wage effects (Borjas) to positive or heterogeneous effects (Card, Ottaviano-Peri, Dustmann et al.). The authors therefore focus on the cultural externalities channel, which they argue better explains voter opposition to immigration, and study cultural integration structurally rather than examining wage outcomes.&lt;/p&gt;
&lt;p&gt;Cultural intolerance: The psychological value a parent obtains from socializing a child to his/her own ethnic identity, relative to having a child adopt a different cultural-ethnic identity. It is specific to the household type (homogamous vs. heterogamous) and is the primary parameter measuring the strength of a group&amp;rsquo;s resistance to cultural integration.&lt;/p&gt;
&lt;p&gt;Cultural socialization / language transmission: The costly investments parents make to transmit their own cultural-ethnic traits to children. In the empirical model, socialization is proxied by whether a parent speaks his/her own non-Italian language at home with children. Socialization technologies are more efficient in homogamous (same-ethnicity) marriages than heterogamous ones.&lt;/p&gt;
&lt;p&gt;Homogamous vs. heterogamous marriage: A homogamous marriage is one in which both spouses share the same cultural-ethnic identity; a heterogamous marriage is one in which spouses differ. The distinction is load-bearing throughout the model: homogamous households have coordinated socialization incentives and superior technology, higher fertility, and lower separation rates.&lt;/p&gt;
&lt;p&gt;Transferable utility (TU) matching: A marriage market framework in which utility is transferable between spouses, so that the equilibrium allocation maximizes aggregate marital surplus and equilibrium transfers are determined by outside options. The model is frictionless, meaning matching is driven purely by preferences over the characteristics of potential spouses.&lt;/p&gt;
&lt;p&gt;Cultural integration (language dimension): In the paper&amp;rsquo;s long-run simulations, cultural integration is defined as the share of second- (or later-) generation immigrants who speak Italian at home with their own children. It is the empirical outcome used to track convergence to the majoritarian culture across generations.&lt;/p&gt;
&lt;p&gt;Assortative mating along cultural-ethnic lines: The tendency for individuals to match with spouses of the same cultural-ethnic group. The paper finds positive assortative mating for all groups, with particularly strong homogamy for Sub-Saharan African and East Asian minorities, and explains it as the equilibrium outcome of the TU matching model given cultural intolerance preferences.&lt;/p&gt;
&lt;p&gt;Socialization technology asymmetry: The model&amp;rsquo;s assumption that homogamous married parents hold a more efficient socialization technology than heterogamous parents, but that divorced heterogamous households invest more in socialization than married heterogamous ones, because the high-intolerance parent can act unilaterally without spousal opposition.&lt;/p&gt;</description></item><item><title>Measuring and Mitigating Racial Disparities in Tax Audits</title><link>https://macropaperwarehouse.com/papers/measuring-and-mitigating-racial-disparities-in-tax-audits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/measuring-and-mitigating-racial-disparities-in-tax-audits/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Do Black taxpayers face higher IRS audit rates than non-Black taxpayers, despite race-blind audit selection? And if so, why — and what would mitigation look like?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Methodology.&lt;/strong&gt; The authors use comprehensive administrative microdata covering approximately 148 million individual income tax returns and 780,627 operational audits for tax year 2014, supplemented with 71,878 research audits from the IRS National Research Program (NRP) pooled over 2010-2014. Because neither the researchers nor the IRS observe taxpayer race, the authors employ Bayesian Improved First Name Surname Geocoding (BIFSG), which imputes the probability that a taxpayer is Black from first name, surname, and Census Block Group. They develop a novel partial identification strategy: two estimators (a probabilistic estimator and a linear estimator) that, under conditions verified using a matched North Carolina voter-registration dataset containing self-reported race, asymptotically bound the true racial audit disparity from below and above respectively. To address the selective labels problem — underreporting is observable only for audited returns — the authors combine operational audit data with NRP random-sample audits to simulate counterfactual audit selection algorithms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Magnitude of the disparity.&lt;/em&gt; The probabilistic estimator implies a racial audit disparity of 0.81 percentage points; the linear estimator implies 1.34 percentage points. Against a base audit rate of 0.54% for the overall U.S. population in 2014, these bounds imply that Black taxpayers are audited at between 2.9 and 4.7 times the rate of non-Black taxpayers.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Role of the EITC.&lt;/em&gt; The disparity is concentrated among EITC claimants. The estimated disparity within the EITC population is 1.96 to 2.90 percentage points, compared to only 0.10 to 0.18 percentage points among non-EITC claimants. In relative terms, Black EITC claimants are audited at 2.9 to 4.4 times the rate of non-Black EITC claimants. A formal decomposition attributes 70-73% of the overall disparity to higher audit rates among Black EITC claimants, 20-21% to racial differences in EITC claiming rates, and 7-8% to differential audit rates among non-EITC filers. Within EITC claimants, 78.5% of the observed audit disparity is attributable to the Dependent Database (DDb) program.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Source of the disparity — algorithmic objective.&lt;/em&gt; Using counterfactual audit selection algorithms estimated on NRP data, the authors find that allocating EITC audits to maximize detected total underreporting (from any source) would produce audit rates of 0.74% for Black EITC claimants versus 1.63% for non-Black EITC claimants — reversing the disparity. In contrast, the status quo, which prioritizes detecting overclaimed refundable credits, yields 3.00% for Black claimants versus 1.04% for non-Black claimants. The primary driver is a difference in the types of noncompliance that are more prevalent by race: dependent-claiming errors are more common among Black EITC claimants (dependent error rate of 26.6% vs. 16.3% for non-Black), while the highest underreporting via business income underreporting is disproportionately concentrated among non-Black EITC claimants. An algorithm focused on refundable credit overclaims implicitly targets dependent errors and therefore selects Black taxpayers at higher rates.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Prediction model bias.&lt;/em&gt; Even conditional on the refundable-credit objective, the status quo disparity (1.96 p.p.) exceeds the disparity that would arise under an oracle that uses actual rather than predicted refundable credit overclaims (1.08 p.p.), suggesting that prediction errors are unevenly distributed by race. The refundable credit prediction algorithm generates a disparity of 1.75 p.p., approximately 60% larger than the oracle. The authors find suggestive evidence of missingness in birth certificate data (paternal information is disproportionately missing for children claimed on Black taxpayers&amp;rsquo; returns) and differential predictive accuracy in the DDb risk score across race.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Operational consequences.&lt;/em&gt; Switching the objective from refundable credit overclaims to total underreporting would shift the composition of audited returns from predominantly dependent-eligibility issues (80% of refundable credit oracle-selected returns contain a dependent error) toward business income (86% of total-underreporting oracle-selected returns have business income underreporting). EITC returns with substantial business income (gross receipts above $25,000) cost on average $369.70 to audit versus $23.09 for other EITC returns. Holding the audit rate fixed, the switch would raise average examination costs by nearly an order of magnitude, while also increasing detected underreporting (mean adjustment of $22,578 per return under the total underreporting oracle versus $9,595 under the refundable credit oracle).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; Results pertain primarily to tax year 2014. The paper finds similar patterns for tax years 2010, 2012, 2016, and 2018. The analysis covers Black versus non-Black taxpayers; disparities for other racial and ethnic groups are not the focus. The selective labels identification strategy relies on the NRP random-audit sample and the bounding conditions verified in the North Carolina matched data.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-cant-the-disparity-be-attributed-simply-to-black-taxpayers-being-more-likely-to-claim-the-eitc-combined-with-eitc-claimants-facing-higher-audit-rates-generally"&gt;Q1. Why can&amp;rsquo;t the disparity be attributed simply to Black taxpayers being more likely to claim the EITC, combined with EITC claimants facing higher audit rates generally?&lt;/h3&gt;
&lt;p&gt;The authors test this directly by estimating racial audit disparities separately within EITC claimants and non-claimants. If differential EITC claiming rates were the full explanation, the within-EITC disparity would be close to zero. Instead, the disparity among EITC claimants (1.96-2.90 p.p.) is larger in absolute terms than the overall disparity (0.81-1.34 p.p.), indicating that Black EITC claimants face substantially higher audit rates than non-Black EITC claimants even holding EITC claimant status fixed. The formal decomposition attributes 70-73% of the overall disparity to differential audit rates within the EITC claimant population, not to differential claiming rates across the population.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-partial-identification-strategy-work-and-what-are-its-key-identifying-assumptions"&gt;Q2. How does the partial identification strategy work, and what are its key identifying assumptions?&lt;/h3&gt;
&lt;p&gt;The authors derive two estimators of the racial audit disparity that use BIFSG-imputed race probabilities rather than observed race. The probabilistic estimator weights each taxpayer&amp;rsquo;s contribution by their estimated probability of being Black; it is downward-biased when there is a positive residual covariance between audits and true race after conditioning on imputed race (E[Cov(Y,B|b)] &amp;gt; 0). The linear estimator regresses audit status on imputed race probability; it is upward-biased when there is a positive residual covariance between audits and imputed race after conditioning on true race (E[Cov(Y,b|B)] &amp;gt; 0). When both covariance terms are positive, the probabilistic and linear estimates bound the true disparity from below and above. The authors verify both conditions are positive and statistically significant (p &amp;lt; 0.01) in the matched North Carolina dataset, for the full population and the EITC population specifically.&lt;/p&gt;
&lt;h3 id="q3-does-the-racial-audit-disparity-within-eitc-claimants-disappear-when-comparing-taxpayers-with-similar-levels-of-underreporting"&gt;Q3. Does the racial audit disparity within EITC claimants disappear when comparing taxpayers with similar levels of underreporting?&lt;/h3&gt;
&lt;p&gt;No. The authors use NRP data to estimate audit rates by race within each underreporting decile among EITC claimants. Within every decile of the underreporting distribution, the estimated audit rate for Black taxpayers exceeds that for non-Black taxpayers. An oracle algorithm that selects returns in descending order of actual underreporting produces an audit rate of 0.74% for Black EITC claimants and 1.63% for non-Black EITC claimants — the opposite of the status quo pattern (3.00% for Black, 1.04% for non-Black). This rules out total-dollar underreporting as the primary driver of the observed disparity.&lt;/p&gt;
&lt;h3 id="q4-why-does-focusing-audit-selection-on-refundable-credit-overclaims-specifically-lead-to-higher-audit-rates-for-black-taxpayers"&gt;Q4. Why does focusing audit selection on refundable credit overclaims specifically lead to higher audit rates for Black taxpayers?&lt;/h3&gt;
&lt;p&gt;Two mechanisms operate simultaneously. First, EITC eligibility is linked to children, so detecting erroneously claimed dependents generates large refundable credit adjustments. The dependent error rate is higher among Black EITC claimants than non-Black EITC claimants (26.6% vs. 16.3% in the probabilistic estimate, or 30.8% vs. 15.4% in the linear estimate). Second, the highest-dollar noncompliance via underreported business income is disproportionately concentrated among non-Black EITC claimants: among EITC claimants in the top 1% of business income underreporting, the probabilistic estimate shows 0.05% are Black compared to 0.21% non-Black. An algorithm aimed at refundable credit overclaims implicitly targets dependent errors and therefore selects Black taxpayers at higher rates; one aimed at total underreporting would prioritize business income underreporting instead and therefore select non-Black taxpayers at higher rates.&lt;/p&gt;
&lt;h3 id="q5-how-do-the-simulated-algorithms-compare-to-the-actual-irs-algorithms"&gt;Q5. How do the simulated algorithms compare to the actual IRS algorithms?&lt;/h3&gt;
&lt;p&gt;The authors cannot directly replicate the IRS&amp;rsquo;s confidential DDb algorithm, but they provide three pieces of evidence that their refundable credit prediction algorithm is a reasonable proxy. First, public governmental documents describe DDb&amp;rsquo;s stated goal as identifying taxpayers who do not meet refundable credit eligibility requirements. Second, when selecting audits based on predicted refundable credit overclaims using largely the same features available to IRS, the authors generate a disparity (1.75 p.p.) close to the status quo disparity (1.96 p.p.). Third, operational audits of EITC returns are strongly associated with their predicted refundable credit overclaims measure but show a much weaker association with predicted total underreporting.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-status-quo-disparity-exceeding-the-refundable-credit-oracle-disparity-reveal-about-prediction-model-design"&gt;Q6. What does the status quo disparity exceeding the refundable credit oracle disparity reveal about prediction model design?&lt;/h3&gt;
&lt;p&gt;The status quo disparity (1.96 p.p.) is approximately 80% larger than the disparity that would arise if the IRS were perfectly informed about actual refundable credit overclaims and selected accordingly (oracle disparity: 1.08 p.p.). The refundable credit prediction algorithm generates a disparity of 1.75 p.p., approximately 60% larger than the oracle. This gap between the oracle and prediction disparity is consistent with prediction errors being distributed unevenly by race. The authors find that birth certificates of children claimed on Black taxpayers&amp;rsquo; returns are substantially more likely to be missing paternal identity information, which may reduce the predictive accuracy of the DDb model for this population. They provide suggestive evidence that modifying the predictive features used could reduce the disparity without substantially degrading credit overclaim detection.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-downstream-operational-consequences-of-switching-the-algorithmic-objective"&gt;Q7. What are the downstream operational consequences of switching the algorithmic objective?&lt;/h3&gt;
&lt;p&gt;Switching from refundable credit overclaims to total underreporting would shift audited issues from dependent eligibility (80% of refundable credit oracle-selected returns have a dependent error) toward business income (86% of total underreporting oracle-selected returns have business income underreporting). Auditing business income returns is substantially more resource-intensive: $369.70 per return on average for returns with gross receipts above $25,000, versus $23.09 for other EITC returns. Holding the current EITC audit rate fixed, the share of audited returns with substantial business income would rise from 3% to 93%, raising total examination costs by nearly an order of magnitude. However, because total detected underreporting per audited return would also rise substantially (mean of $22,578 vs. $9,595), the increase in detected noncompliance would exceed the increase in audit costs, and the qualitative pattern persists even when accounting for higher per-return costs.&lt;/p&gt;
&lt;h3 id="q8-is-the-disparity-consistent-across-years-and-is-it-driven-by-a-particular-audit-type"&gt;Q8. Is the disparity consistent across years, and is it driven by a particular audit type?&lt;/h3&gt;
&lt;p&gt;The authors find comparable audit disparities for tax years 2010, 2012, 2016, and 2018, confirming the 2014 results are not year-specific. The disparity is concentrated in correspondence audits: the estimated disparity in correspondence audit rates is 0.804-1.328 p.p. for the full population, while the disparity in field/office audit rates is only 0.010-0.016 p.p. The disparity is present in both pre-refund and post-refund audits, though pre-refund audits show a larger disparity even among correspondence audits alone. Among EITC claimants, the correspondence audit channel is nearly entirely responsible for the group-level disparity.&lt;/p&gt;
&lt;h3 id="q9-what-heterogeneity-exists-within-eitc-claimants"&gt;Q9. What heterogeneity exists within EITC claimants?&lt;/h3&gt;
&lt;p&gt;The disparity is especially pronounced among unmarried male EITC claimants with dependents: among this subgroup, the audit rate for Black men exceeds the audit rate for non-Black men by more than 4 percentage points, and both are an order of magnitude above the overall U.S. population audit rate. Disparities are smaller among joint filers, unmarried women, and unmarried men without dependents, though the ratio of Black to non-Black audit rates remains substantial across all subgroups. The concentration of the disparity among unmarried men with dependents is consistent with the role of dependent-claiming errors, which are more likely to arise in family structures characterized by nonmarital cohabitation — a pattern more prevalent among Black Americans due to lower marriage rates.&lt;/p&gt;
&lt;h3 id="q10-can-the-disparity-be-attributed-to-disparate-treatment--ie-race-conscious-selection"&gt;Q10. Can the disparity be attributed to disparate treatment — i.e., race-conscious selection?&lt;/h3&gt;
&lt;p&gt;The authors rule out disparate treatment for the EITC population. The DDb audit selection process for EITC returns is automated (no manual review), and IRS does not use race or geography as an input into audit selection. The disparity is therefore the product of disparate impact: race-neutral selection criteria interact with racially correlated patterns of tax return characteristics to produce differential audit rates. For higher-income non-EITC taxpayers, where audit selection may involve human classifiers, the authors cannot rule out disparate treatment.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Audit Disparity (D).&lt;/strong&gt; Defined in the paper as D = E[Y|B=1] - E[Y|B=0], the difference in audit rates between Black taxpayers (B=1) and non-Black taxpayers (B=0). This is a group-level difference in selection rates, not conditional on any other characteristic, and is the primary estimand throughout.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Probabilistic Disparity Estimator.&lt;/strong&gt; An estimator that calculates group-specific audit rates by weighting each taxpayer&amp;rsquo;s contribution by their BIFSG-imputed probability of being Black (or non-Black). It is shown to be downward-biased when E[Cov(Y,B|b)] &amp;gt; 0, i.e., when there is residual positive association between true race and audits after conditioning on imputed race.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Linear Disparity Estimator.&lt;/strong&gt; An estimator based on regressing audit status (Y) on BIFSG-imputed race probability (b). It is shown to be upward-biased when E[Cov(Y,b|B)] &amp;gt; 0, i.e., when imputed race probability predicts audits even after conditioning on true race. Together, the probabilistic and linear estimators form bounds on the true disparity under conditions verified empirically.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;BIFSG (Bayesian Improved First Name Surname Geocoding).&lt;/strong&gt; A probabilistic race imputation method that uses Bayes rule under a conditional independence assumption (first name, surname, and geography are independent given race) to compute Pr[Black | first name, surname, Census Block Group]. Applied here to all 148 million tax returns; calibrated and validated against matched North Carolina voter registration data with self-reported race.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Selective Labels Problem.&lt;/strong&gt; The problem that noncompliance (underreporting) is observed only for returns selected for audit, not for the full filing population. In this paper it means the IRS cannot directly observe the underreporting distribution for unaudited returns. The authors address this using NRP random-audit data, which allows estimation of the unaudited underreporting distribution and construction of counterfactual selection algorithms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Algorithmic Objective.&lt;/strong&gt; The paper distinguishes between (1) the prediction component of audit selection — which model to use to forecast noncompliance — and (2) the objective component — what type of noncompliance to predict and pursue (overclaimed refundable credits versus total underreporting from any source). The paper finds that the objective, not just prediction error, is an independent driver of the racial audit disparity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dependent Database (DDb) Program.&lt;/strong&gt; The IRS&amp;rsquo;s primary EITC audit selection program, responsible for approximately 75% of audited EITC returns in 2014. DDb flags returns based on rules, heuristics, and proprietary risk scores, with the stated goal of identifying taxpayers who do not meet refundable credit eligibility requirements. Selection through DDb is fully automated, without human classifier review.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;National Research Program (NRP).&lt;/strong&gt; A stratified random sample audit program through which the IRS conducts near-line-by-line examinations of a small fraction of the filing population each year (approximately 2% of audited returns in 2014). The paper pools 71,878 NRP audits from 2010-2014 to identify the distribution of underreporting in the full EITC filing population and to estimate counterfactual selection algorithms.&lt;/p&gt;</description></item><item><title>Minimum Wages, Efficiency, and Welfare</title><link>https://macropaperwarehouse.com/papers/minimum-wages-efficiency-and-welfare/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/minimum-wages-efficiency-and-welfare/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; Can minimum wages improve welfare through efficiency — by correcting monopsony-driven under-employment — and, if so, by how much? What is the optimal minimum wage, and how much of the welfare gain from a higher minimum wage comes from efficiency versus redistribution?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model and methodology.&lt;/strong&gt; The paper develops a tractable general equilibrium oligopsony model with heterogeneous workers (four types: non-high-school, high-school, college workers, and capital owners) and heterogeneous firms (varying in total factor productivity), embedded in a continuum of local labor markets where firms compete strategically in Cournot fashion. Firms face downward-sloping labor supply curves; their market power generates wages below the marginal revenue product of labor (markdowns). The model is calibrated to US data using the Census Longitudinal Business Database (LBD, 2014), the Bureau of Labor Statistics Current Population Survey (CPS, 2019), and the Survey of Consumer Finances (SCF). Key calibration targets include: average firm size of 22.83 workers (LBD), 29 percent of workers earning below $15/hr (CPS), labor and capital income shares, and household-level earnings and capital income ratios. The model is validated by quantitatively replicating four strands of empirical evidence: (i) reallocation effects of the German minimum wage introduction (Dustmann et al., 2021); (ii) employer spillover responses to Amazon&amp;rsquo;s voluntary $15 minimum wage (Derenoncourt et al., 2021); (iii) wage distribution compression evidence from Brazil (Engbom and Moser, 2021); and (iv) heterogeneous employment effects by market concentration (Azar et al., 2019).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three channels for efficiency gains.&lt;/strong&gt; The model identifies three mechanisms through which a minimum wage can improve efficiency under oligopsony: (1) a &lt;em&gt;direct effect&lt;/em&gt; in which constrained firms with monopsony markdowns increase wages and expand employment toward the competitive level (Region II firms); (2) a &lt;em&gt;spillover effect&lt;/em&gt; in which unconstrained competitor firms narrow their own markdowns in response to constrained firms&amp;rsquo; increased wages and market shares; (3) a &lt;em&gt;reallocation effect&lt;/em&gt; in which employment is shifted away from low-productivity firms (which enter Region III — constrained on labor demand) toward more productive firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings on efficiency versus redistribution.&lt;/strong&gt; Under the $15.12/hr minimum wage that maximizes social welfare under utilitarian weights (population-share weights), less than 5 percent of the welfare gains come from improved efficiency, while more than 95 percent come from redistribution. When the government is additionally given access to budget-neutral lump-sum transfers that fully address redistribution goals, the efficiency-maximizing minimum wage narrows to a range of approximately $7.50–$10.00 per hour, which is robust across social welfare weight specifications. The welfare gains attributable to efficiency alone are approximately 0.16–0.20 percent in consumption-equivalent terms, representing only about 1–2 percent of the welfare gains achievable in an economy with no labor market power at all (which would be 15.26 percent in consumption-equivalent terms under the same conditions with optimal transfers).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why efficiency gains are small.&lt;/strong&gt; Three structural reasons limit efficiency gains: (i) low-productivity firms — which are the firms most affected by a binding minimum wage in Region II — have endogenously narrow markdowns even absent a minimum wage, because they face more elastic labor supply and command small market shares; (ii) the calibrated production function has relatively flat marginal revenue product of labor schedules (decreasing returns parameter α = 0.940), so once firms enter Region III, employment rationing occurs rapidly; (iii) the large, high-productivity firms with the widest markdowns are not materially affected by the minimum wages of their small, low-wage competitors because those competitors have small market shares — making spillovers quantitatively negligible even though the model matches empirical cross-employer wage elasticities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Optimal minimum wages under alternative frameworks.&lt;/strong&gt; Without transfers and under utilitarian weights, the optimal minimum wage is $15.12. Without transfers but under Negishi weights (which rationalize the observed competitive equilibrium and load approximately 62 percent of weight on college workers and owners versus their 35 percent population share), the optimal is $6.97. Under a 97 percent weight on high-school graduates, the optimal rises to $18.32. With optimal lump-sum transfers, the optimal collapses to $7.76–$10.11 regardless of social welfare weights — a range robust across Frisch elasticity variants (ϕ ∈ {0.30, 0.62, 0.86}), regional decompositions (low, medium, and high income US states), short-run capital-fixed scenarios (where the optimum declines by approximately $1 under utilitarian weights), and the removal of household heterogeneity entirely (which yields an optimum of $7.74).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distributional proxies versus welfare.&lt;/strong&gt; Wage inequality (college–non-college log wage premium, cross-sectional variance of log wages) and the labor income share are monotonically improving as the minimum wage rises, even as welfare is hump-shaped and eventually declining. A rise in the minimum wage from $7.50 to $15 reduces the college–non-college log wage premium from 0.53 to 0.43 (roughly one-fifth), reduces the cross-sectional variance of log wages by nearly half, and raises the aggregate labor income share by approximately 3 percentage points — all while welfare (under utilitarian weights with no transfers) reaches its maximum at $15.12 and then declines. These standard proxies therefore do not reliably indicate welfare.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; All results are long-run steady-state comparisons unless otherwise noted. Results assume no price passthrough and a unit elasticity of substitution between capital and labor. The paper abstracts from capital–labor substitution responses and occupational choice. The redistribution channel quantified here is specific to the utilitarian welfare criterion and to the existing distribution of capital and profit income, in which owners (6 percent of households) earn 92 percent of dividends.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-three-regions-of-firm-behavior-in-response-to-a-binding-minimum-wage-and-what-are-their-efficiency-implications"&gt;Q1. What are the three regions of firm behavior in response to a binding minimum wage, and what are their efficiency implications?&lt;/h3&gt;
&lt;p&gt;A: A firm can be in one of three regions. In Region I the minimum wage is not binding: the firm pays its optimal monopsony wage and employment is inelastically below the competitive level. In Region II the minimum wage binds and exceeds the firm&amp;rsquo;s optimal monopsony wage, but labor supply at the minimum wage still falls short of labor demand: employment and efficiency improve as the shadow markdown narrows. In Region III the minimum wage exceeds the competitive wage, so unconstrained labor supply would exceed demand: the firm rations employment and the rationing constraint binds, reducing efficiency. At the boundary of Region II and Region III, the shadow markdown equals one and the firm is at its efficient employment level. Only a firm-specific minimum wage targeting each firm&amp;rsquo;s competitive wage could deliver economy-wide efficiency.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-define-and-use-shadow-wages-to-characterize-equilibrium"&gt;Q2. How does the paper define and use &amp;ldquo;shadow wages&amp;rdquo; to characterize equilibrium?&lt;/h3&gt;
&lt;p&gt;A: The shadow wage for a firm is the effective wage that rationalizes equilibrium employment given rationing constraints. Formally, when a firm rations employment (Region III), households act as if facing a shadow wage equal to the actual minimum wage multiplied by a rationing factor p &amp;lt; 1 (the Lagrange multiplier on the rationing constraint, normalized as a fraction). Shadow wages aggregate across firms into market- and type-level shadow wages via CES aggregation. The key insight is that shadow wages, not observed wages, are allocative: aggregate labor supply for each worker type is determined by the type-level shadow wage, not by the minimum wage that firms actually pay. This allows the paper to express aggregate efficiency via two wedges — the aggregate shadow markdown (capturing average market power) and a misallocation term — without tracking all firm-specific constraints individually.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-two-aggregate-efficiency-wedges-and-how-do-they-behave-as-the-minimum-wage-rises"&gt;Q3. What are the two aggregate efficiency wedges and how do they behave as the minimum wage rises?&lt;/h3&gt;
&lt;p&gt;A: The two wedges are: (i) the aggregate shadow markdown µ̃, which is a productivity-weighted average of firm-level shadow markdowns and measures the extent to which aggregate wages fall short of marginal revenue products; and (ii) the misallocation term ω, which measures whether employment is allocated toward more productive firms and equals one when all shadow markdowns are identical. As the minimum wage rises from zero, µ̃ initially narrows (improving efficiency) because firms in Region II expand toward their competitive employment level and constrained firms&amp;rsquo; market shares rise, tightening the residual labor supply of unconstrained competitors and narrowing their markdowns. But as the minimum wage rises further, Region III rationing causes shadow markdowns to widen rapidly — first for low-productivity firms and then progressively for more productive ones — so µ̃ turns back downward. The misallocation term ω first improves as low-productivity firms are pushed out, but then worsens because rationing at intermediate-productivity firms redirects employment from high- to medium-productivity firms.&lt;/p&gt;
&lt;h3 id="q4-what-does-the-model-validation-exercise-on-the-german-minimum-wage-dlsub-2021-show"&gt;Q4. What does the model validation exercise on the German minimum wage (DLSUB 2021) show?&lt;/h3&gt;
&lt;p&gt;A: The paper calibrates the model to the German context by setting a minimum wage of $8.95/hr equivalent to 48 percent of the pre-reform median wage — matching Germany&amp;rsquo;s 8.50 euro introduction in 2015, where 15 percent of workers earned below the threshold. The model produces employment effects that are slightly positive (consistent with empirical findings of no disemployment), average wage increases consistent with both constrained and unconstrained firms raising wages, a negative elasticity of the number of operating firms with respect to minimum wage exposure (correctly signed, moderately smaller than data), and a positive elasticity of average firm size with respect to exposure (slightly larger than the data). The reallocation direction — small unproductive firms shrinking and workers moving to larger, more productive firms — matches the data qualitatively and within the range of data estimates across specifications.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-amazon-spillover-replication-dnwt-2021-show-and-what-does-it-imply-about-the-minimum-wage-spillover-channel"&gt;Q5. What does the Amazon spillover replication (DNWT 2021) show, and what does it imply about the minimum wage spillover channel?&lt;/h3&gt;
&lt;p&gt;A: Derenoncourt et al. (2021) estimate a cross-employer wage elasticity of 0.26: when Amazon raised wages by approximately 18.1 percent, competitors raised wages by 4.7 percent on average. The model replicates this by treating Amazon as the largest (or second-largest) firm in each market, exogenously narrowing its markdown by a fraction ζ calibrated to deliver an 18.1 percent wage increase. Competitors in the model raise wages through the strategic interaction mechanism: Amazon&amp;rsquo;s higher wage and market share tightens competitors&amp;rsquo; residual supply curves, inducing them to narrow their own markdowns. The model matches the 0.26 cross-employer elasticity when Amazon is the largest firm in markets with at least 36 competitors, or the second-largest in markets with at least 12. Critically, the authors note that this empirical evidence concerns responses to a &lt;em&gt;large&lt;/em&gt; firm raising wages; for minimum wages the question is whether &lt;em&gt;large&lt;/em&gt; firms respond to their small wage competitors, which the model shows they do not substantially, because small firms have negligible market shares.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-separate-efficiency-from-redistribution-and-what-is-the-key-methodological-innovation"&gt;Q6. How does the paper separate efficiency from redistribution, and what is the key methodological innovation?&lt;/h3&gt;
&lt;p&gt;A: The paper gives the government access to budget-neutral, unrestricted lump-sum transfers across households in addition to the minimum wage. With transfers available, the government can use them to meet any redistributive objective encoded in arbitrary social welfare weights. Whatever is left for the minimum wage to do must be purely efficiency-improving. The paper shows (via aggregation theorems) that optimal lump-sum transfers can be computed in closed form for any social welfare weights, and that the social welfare maximizing allocation subject to transfers can be decentralized by transfers that sum to zero across households. Under this framework, the efficiency-maximizing minimum wage lies between $7.50 and $10.00 per hour regardless of whether utilitarian, Negishi, or 97 percent high-school-weighted social welfare functions are used — collapsing the original $0–$31 range to a tight interval.&lt;/p&gt;
&lt;h3 id="q7-how-are-negishi-weights-computed-and-why-are-they-important-for-interpreting-the-results"&gt;Q7. How are Negishi weights computed, and why are they important for interpreting the results?&lt;/h3&gt;
&lt;p&gt;A: The Negishi weights are the social welfare weights under which a planner would choose the observed competitive equilibrium with zero lump-sum transfers. They are computed by inverting the planner&amp;rsquo;s first-order conditions: for the competitive equilibrium to be optimal under some set of weights, the implied consumption ratios must match observed data. The calibrated Negishi weights assign a combined weight of approximately 62 percent to college workers and owners, who constitute only 35 percent of the population. This means the competitive equilibrium is disproportionately aligned with higher-income households. A utilitarian planner, which weights households by population shares, therefore sees large scope for redistribution toward non-college workers — which is exactly why the utilitarian-optimal minimum wage is $15.12 and why 94 percent of its welfare gains come from redistribution rather than efficiency.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-quantitative-welfare-gains-from-the-efficiency-maximizing-minimum-wage-and-how-small-are-they-relative-to-the-potential-gains-from-eliminating-monopsony"&gt;Q8. What are the quantitative welfare gains from the efficiency-maximizing minimum wage, and how small are they relative to the potential gains from eliminating monopsony?&lt;/h3&gt;
&lt;p&gt;A: With optimal lump-sum transfers, the welfare gains from the efficiency-maximizing minimum wage are approximately 0.16–0.20 percent in consumption-equivalent terms, robust across social welfare weight specifications, Frisch elasticity variations, and regional decompositions. The welfare gains associated with an economy in which all firms&amp;rsquo; markdowns are set to one (no labor market power at all), also evaluated with optimal transfers, are 15.26 percent in consumption-equivalent terms. The efficiency-maximizing minimum wage therefore recovers approximately 1–2 percent of the potential welfare gains from eliminating monopsony. Equivalently, the efficiency gains correspond to roughly a 0.1 percent increase in TFP. These gains are small despite the model matching all empirical evidence on the channels through which efficiency gains could occur.&lt;/p&gt;
&lt;h3 id="q9-how-do-employment-effects-of-minimum-wages-vary-by-market-concentration-and-why"&gt;Q9. How do employment effects of minimum wages vary by market concentration, and why?&lt;/h3&gt;
&lt;p&gt;A: In concentrated markets (upper tercile of HHI), firms have larger monopsony markdowns, so a binding minimum wage pushes them into Region II — where employment expands — over a wider range of minimum wage values before entering Region III. This produces large, positive employment effects in concentrated markets. In less concentrated markets, firms already have narrow markdowns (they are closer to competitive), so even small minimum wage increases push them into Region III, where employment contracts. The model replicates the statistically significant positive effects in high-concentration markets and negative effects in low-concentration markets documented by Azar et al. (2019), for initial minimum wages below approximately $8/hr. At higher initial minimum wages, however, even high-concentration markets exhibit negative employment effects as more firms enter Region III.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-robustness-exercise-for-mississippi-reveal"&gt;Q10. What does the robustness exercise for Mississippi reveal?&lt;/h3&gt;
&lt;p&gt;A: Mississippi has the lowest per capita income in the US, and a $15 minimum wage would bind for 41.3 percent of its workers (versus 29.4 percent nationally). Despite this, the model finds that Mississippi would benefit from a $15 federal minimum wage under utilitarian weights, and the Mississippi-specific optimal minimum wage is $14.89 — nearly identical to the national optimum. The reason is an offsetting compositional effect: while Mississippi has lower average wages (pushing toward a lower optimal), it has a larger share of high-school graduates (63 percent versus 52.8 percent nationally) who prefer higher minimum wages (around $17 in the model). These two forces wash out, producing a stable optimal close to the national figure.&lt;/p&gt;
&lt;h3 id="q11-what-happens-to-common-empirical-proxies-for-inequality-and-worker-power-as-the-minimum-wage-rises"&gt;Q11. What happens to common empirical proxies for inequality and worker power as the minimum wage rises?&lt;/h3&gt;
&lt;p&gt;A: The college–non-college log wage premium declines from 0.53 to 0.43 (a fall of roughly one-fifth) as the minimum wage rises from $7.50 to $15. The cross-sectional variance of log wages falls by nearly half over this range, driven equally by declining within- and between-type inequality. The aggregate labor income share rises by approximately 3 percentage points, and the share of output created in non-high-school jobs paid to non-high-school workers rises by 7 percentage points. All of these proxies are monotonically improving in the minimum wage throughout, even as aggregate welfare under the model&amp;rsquo;s social welfare function is hump-shaped and declining past the optimum. The paper concludes that observations of declining inequality or a rising labor share are consistent with falling welfare, so these proxies cannot serve as reliable welfare indicators.&lt;/p&gt;
&lt;h3 id="q12-how-does-the-short-run-fixed-capital-analysis-differ-from-the-long-run-baseline"&gt;Q12. How does the short-run (fixed-capital) analysis differ from the long-run baseline?&lt;/h3&gt;
&lt;p&gt;A: In the short run, capital at each firm is fixed at the type-specific level chosen under a zero minimum wage. This creates sharper decreasing returns in labor (parameter γα rather than α̃), overhead costs that can make operation unprofitable, and a narrower range of minimum wages over which firms remain in Region II. The result is that firms in the short run enter Region III at lower minimum wages than in the long run, limiting the range of efficiency gains. Quantitatively, the efficiency-maximizing optimal minimum wage declines by approximately $1 under utilitarian weights (from about $10 to about $9 in the short-run exercise) and by only about $0.20 under Negishi weights. The robustness conclusion is that the difference between short- and long-run optimal minimum wages is modest, and the main finding that efficiency gains are small is preserved.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Shadow wage (w̃ᵢⱼ):&lt;/strong&gt; The effective wage that rationalizes a firm&amp;rsquo;s equilibrium employment in the presence of a minimum wage. When labor is rationed at firm ij (Region III), the shadow wage equals the actual minimum wage multiplied by a rationing factor pᵢⱼ &amp;lt; 1, where pᵢⱼ is derived from the Lagrange multiplier on the household&amp;rsquo;s rationing constraint. The shadow wage is allocative — it determines labor supply decisions — while the observed minimum wage wage is not. When the rationing constraint is slack (Regions I and II), the shadow wage coincides with the observed wage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Shadow markdown (µ̃ᵢⱼ):&lt;/strong&gt; The ratio of a firm&amp;rsquo;s shadow wage to its marginal revenue product of labor. In Region I (unconstrained), this equals the standard monopsony markdown. In Region II (constrained, on the labor supply curve), the shadow markdown narrows as the minimum wage increases, moving the firm toward its efficient employment level. In Region III (constrained, on the labor demand curve), the shadow markdown equals the rationing multiplier pᵢⱼ and widens, reflecting efficiency losses from rationing. An aggregate shadow markdown µ̃ is computed as a productivity-weighted average of firm-level shadow markdowns across all firms in the economy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Misallocation wedge (ω):&lt;/strong&gt; A productivity-weighted measure of how well employment is allocated across firms. In an efficient allocation with identical shadow markdowns, ω = 1. When high-productivity firms have wider markdowns than low-productivity firms (the baseline oligopsony outcome), ω &amp;lt; 1 because employment is directed away from productive firms. A minimum wage can improve ω by shrinking low-productivity firms but worsens it when high-productivity firms enter Region III and are over-rationed relative to medium-productivity firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Oligopsony with Cournot competition:&lt;/strong&gt; The specific form of labor market power in this model. In each local labor market (defined as a NAICS 3-digit industry × commuting zone cell), a finite number of firms compete strategically in employment quantities, taking their competitors&amp;rsquo; employment levels as given (Cournot assumption). Each firm has an upward-sloping labor supply curve derived from nested CES household preferences, and exercises a markdown on the marginal revenue product of labor. This differs from monopsony (one firm) or perfect competition (infinitely many firms), and generates both direct effects and spillover effects of minimum wages.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negishi weights:&lt;/strong&gt; The vector of social welfare weights under which the observed competitive equilibrium allocation would be the solution to a social planner&amp;rsquo;s problem with zero lump-sum transfers. In this model, the calibrated Negishi weights assign roughly 62 percent combined weight to college workers and owners (who constitute only 35 percent of the population), reflecting the fact that the market equilibrium allocates a disproportionate share of consumption to high-income households. The Negishi weights are used both to identify the gap between market outcomes and utilitarian objectives (motivating redistribution) and as one alternative normative benchmark.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Efficiency-maximizing minimum wage:&lt;/strong&gt; The minimum wage that maximizes social welfare when the government additionally has access to budget-neutral lump-sum transfers across households. Because transfers can be optimized to handle any redistributive objective encoded in any arbitrary social welfare weights, the minimum wage under this framework serves solely to improve productive efficiency. In the calibrated model, the efficiency-maximizing minimum wage is approximately $7.50–$10.00 per hour, robust to social welfare weight specifications, Frisch elasticity variations (ϕ ∈ {0.30, 0.86}), and regional income differences.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rationing constraint (n̄ᵢⱼₖ):&lt;/strong&gt; A firm-specific, type-specific upper bound on the labor a household may supply to a firm in equilibrium. These constraints are taken as given by households and determined in equilibrium by firms&amp;rsquo; labor demand decisions. When the minimum wage is above the firm&amp;rsquo;s competitive wage (Region III), the firm&amp;rsquo;s labor demand is less than what households would want to supply at that wage, so the rationing constraint binds. The binding rationing constraint generates the shadow wage discount (pᵢⱼ &amp;lt; 1) and is the mechanism by which high minimum wages reduce efficiency in the model.&lt;/p&gt;</description></item><item><title>Mis(sed) Diagnosis: Physician Decision Making and ADHD</title><link>https://macropaperwarehouse.com/papers/missed-diagnosis-physician-decision-making-and-adhd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/missed-diagnosis-physician-decision-making-and-adhd/</guid><description>&lt;p&gt;This paper develops and estimates a structural model of ADHD diagnosis to decompose the mechanisms driving the observed 2.3:1 male-to-female diagnostic difference in the United States. The research question is: to what extent does the large gender gap in ADHD diagnosis reflect true differences in symptom prevalence, versus patient-side utilization costs, versus physician decision-making under uncertainty? The setting is particularly well-suited to this question because DSM-V diagnostic guidelines for ADHD are explicitly gender-neutral, making any gender difference in physician thresholds a detectable deviation from uniform clinical rules.&lt;/p&gt;
&lt;p&gt;The data come from de-identified electronic health records from a large Arizona healthcare system covering January 2014 through September 2017. The sample encompasses 36,193 unique encounters for approximately 11,070 pediatric patients. The raw male-to-female diagnostic ratio in the data is 2.32:1 (7.2% of males vs. 3.1% of females receive a clinical ADHD diagnosis). This gap persists after controlling for demographics, general healthcare utilization, and mental health utilization in reduced-form regressions, motivating the structural approach.&lt;/p&gt;
&lt;p&gt;Because two key variables — whether a patient received a behavioral assessment (Qi) and the ADHD match signal observed by the physician (xi) — are not directly recorded in the EHR, the author constructs them from clinical doctor note text. A random forest machine learning classifier trained on labeled appointments predicts behavioral assessment take-up for unlabeled encounters; approximately 20.8% of children are predicted to have received a behavioral assessment (23.2% of males vs. 18.3% of females). The ADHD match signal is constructed via an adjusted Bag-of-Words cosine similarity measure comparing each patient&amp;rsquo;s aggregated note text to the DSM-V symptom list, rescaled to [0,1]. The average signal is 0.319 overall, with males averaging 0.326 and females 0.311.&lt;/p&gt;
&lt;p&gt;The structural model has three stages. First, patients/caregivers decide whether to schedule a behavioral assessment, a function of underlying latent ADHD risk (vi) and mental healthcare utilization costs (ci). Second, conditional on assessment, the physician receives a noisy signal of vi and updates beliefs via Bayesian learning; signal quality ρ governs diagnostic uncertainty. Third, the physician diagnoses ADHD if posterior risk exceeds a gender-specific diagnostic threshold τ. Population mean ADHD risk (μ) is identified using regression-adjusted initial primary care provider referral rates as a quasi-exogenous cost-shifter — patients of high-referral-rate providers select into assessment less selectively, so their observed signals approach population mean risk. This extrapolation approach follows Arnold et al. (2022).&lt;/p&gt;
&lt;p&gt;The structural parameter estimates reveal that male and female children have similar but slightly different mean ADHD risk (μm = 0.290 vs. μf = 0.262) and similar mean utilization costs (cm = 0.116 vs. cf = 0.109). The most striking differences are in physician parameters: signal quality is lower for male patients (ρm = 0.479 vs. ρf = 0.552), indicating higher diagnostic uncertainty for boys; and diagnostic thresholds are substantially lower for male patients (τm = 0.257 vs. τf = 0.312), meaning physicians are willing to diagnose ADHD in boys with lower posterior risk.&lt;/p&gt;
&lt;p&gt;Counterfactual decomposition simulations attribute approximately 20–25% of the 2.32:1 diagnostic gap to underlying differences in ADHD risk, approximately 20% to differences in selection into behavioral assessments, and the remaining majority — approximately 55–60% — to physician decision-making. Within physician decision-making, differences in diagnostic thresholds alone account for roughly two-thirds of the overall diagnostic gap.&lt;/p&gt;
&lt;p&gt;The paper offers economic rationales for why gender-specific thresholds may be consistent with physician rationality despite uniform guidelines: higher diagnostic uncertainty for boys justifies lower thresholds under Bayesian updating; hyperactive/impulsive symptoms predominant in boys impose larger classroom externalities (Aizer, 2008); and female patients show higher rates of internalizing co-morbidities (anxiety, depression) that may reduce the marginal benefit of an additional ADHD diagnosis. A type-specific threshold extension finds that for male patients the threshold for hyperactive/impulsive symptoms is significantly lower than for inattentive symptoms, consistent with salience of externally disruptive behaviors. These rationalizations do not vindicate the gap as fully guideline-consistent, but suggest physicians may be responding to real heterogeneity in external costs and co-morbidity patterns.&lt;/p&gt;
&lt;p&gt;Q: What is the main research question and why is ADHD a useful setting?
A: The paper asks what mechanisms produce the 2.3:1 male-to-female ADHD diagnostic difference: true symptom prevalence, patient utilization costs, or physician decision-making. ADHD is well-suited because (1) clinical guidelines (DSM-V) are explicitly gender-neutral and require the same symptom count threshold regardless of sex; (2) diagnosis is based on subjective behavioral assessment rather than objective testing, creating substantial physician discretion; and (3) both missed and excess diagnosis carry meaningful costs — missed diagnosis limits educational accommodations; excess diagnosis exposes children to Schedule II controlled substances.&lt;/p&gt;
&lt;p&gt;Q: What data does the paper use and what are the key descriptive facts?
A: The data are de-identified electronic health records from a large Arizona healthcare system, 2014–2017, covering 36,193 encounters for 11,070 pediatric patients aged 5 and above. Overall ADHD diagnosis rate is 5.2%, with males at 7.2% and females at 3.1%, a 2.32:1 ratio that matches national levels. Approximately 49.5% of the sample is Hispanic, which the author notes contributes to a below-national-average overall diagnosis rate. The gender diagnostic gap persists even after controlling for demographics, general healthcare utilization, and mental health utilization in reduced-form regressions.&lt;/p&gt;
&lt;p&gt;Q: How does the paper construct the behavioral assessment indicator (Qi) and the ADHD match signal (xi)?
A: Qi is constructed using a random forest classifier trained on doctor notes from appointments where assessment status is known with near-certainty (ADHD diagnosis or DSM-V comorbid diagnosis = positive; non-mental-health diagnosis code for patients with no mental health history = negative). The classifier uses 41 features including note length and top-20 word frequencies for each label class. xi is constructed via an adjusted Bag-of-Words cosine similarity between each patient&amp;rsquo;s combined behavioral assessment notes and the DSM-V symptom list, separately for inattentive and hyperactive/impulsive sub-types, taking xi = max{xi1, xi2}. The average xi is 0.319 (males 0.326, females 0.311) in the behavioral assessment subsample.&lt;/p&gt;
&lt;p&gt;Q: What is the identification strategy for recovering population mean ADHD risk (μ)?
A: Because xi is observed only for endogenously selected patients, the observed sample mean overestimates population mean risk. The author uses regression-adjusted referral rates of each patient&amp;rsquo;s initial primary care provider (IPCP) as a quasi-exogenous cost-shifter satisfying (a) relevance — IPCP referral intensity lowers patient scheduling costs — and (b) independence from patient ADHD risk vi, since IPCPs are typically chosen before behavioral symptoms develop and only 28% of IPCPs in the sample ever diagnose ADHD themselves. Population mean risk is then recovered by extrapolating the relationship between IPCP referral propensity and average observed xi to propensity = 1, following Arnold et al. (2022). The maximum observed IPCP referral propensity is only about 0.75, so the estimate requires extrapolation beyond the observed support.&lt;/p&gt;
&lt;p&gt;Q: What are the estimated structural parameters and what do they imply?
A: Mean ADHD risk is μm = 0.290 vs. μf = 0.262 — males have modestly higher underlying risk. Mean utilization costs are cm = 0.116 vs. cf = 0.109 — nearly identical across genders. Signal quality (diagnostic certainty) is lower for males: ρm = 0.479 vs. ρf = 0.552, indicating physicians face more diagnostic uncertainty when assessing boys. Most importantly, diagnostic thresholds are lower for males: τm = 0.257 vs. τf = 0.312, meaning physicians diagnose ADHD in boys at a lower required posterior risk level, consistent with viewing missed diagnosis as relatively more costly for male patients.&lt;/p&gt;
&lt;p&gt;Q: How much of the 2.32:1 diagnostic gap can be attributed to each mechanism?
A: Counterfactual simulations decompose the gap as follows: differences in underlying ADHD risk distribution account for approximately 20–25% of the diagnostic difference; differences in selection into behavioral assessments (utilization costs operating through assessment rates) account for approximately 20%; and physician decision-making differences account for the remaining majority, approximately 55–60%. Within physician factors, differences in diagnostic thresholds (τm &amp;lt; τf) are the single largest contributor, explaining roughly two-thirds of the overall male/female diagnostic gap.&lt;/p&gt;
&lt;p&gt;Q: What do the type-specific threshold estimates reveal?
A: When the baseline model is extended to allow separate diagnostic thresholds for inattentive vs. hyperactive/impulsive symptom sub-types, male patients show significantly lower thresholds for hyperactive/impulsive symptoms relative to inattentive symptoms (τ^HI_m &amp;lt; τ^Inatt_m). This is consistent with the hypothesis that more externally salient and disruptive symptoms carry larger classroom externalities, which physicians may implicitly factor into diagnosis decisions (following Aizer, 2008). For female patients, the threshold differences across symptom types are smaller and less statistically significant.&lt;/p&gt;
&lt;p&gt;Q: What economic rationales does the paper offer for gender-specific diagnostic thresholds despite uniform guidelines?
A: Three mechanisms are identified. First, higher diagnostic uncertainty for males (lower ρm) implies that under symmetric costs, Bayesian-rational physicians should set lower thresholds when the signal is noisier — this alone partially rationalizes the threshold gap. Second, hyperactive/impulsive symptoms predominant in boys impose greater externalities on classroom peers (Aizer, 2008), increasing the social benefit of diagnosis for boys on the margin. Third, females show substantially higher rates of co-morbid internalizing conditions (anxiety, depression) whose treatment may mitigate ADHD-related behaviors or whose interaction with stimulant medication makes the marginal ADHD diagnosis less beneficial for girls (Currie et al., 2014). These factors together suggest physicians may be responding to genuine heterogeneity in net diagnosis benefits, even if their behavior deviates from gender-neutral clinical guidelines.&lt;/p&gt;
&lt;p&gt;Q: What share of the 2.3:1 national diagnostic gap is consistent with genuine symptom prevalence differences?
A: Simulations indicate that only about 20–25% of the 2.32:1 male/female diagnostic difference can be explained by the underlying difference in ADHD risk distributions. The majority — roughly 75–80% — reflects factors beyond true prevalence: selection into care and, most substantially, physician decision-making differences including both signal quality and diagnostic thresholds.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications?
A: The findings suggest that targeted interventions in physician awareness and clinical training are likely more effective than generic awareness campaigns, since the dominant driver of the diagnostic gap is physician threshold-setting rather than symptom prevalence. Structured decision support tools or updated training that make physicians aware of gender-specific diagnostic patterns could reduce medically unwarranted diagnostic differences. Policies targeting patient-side access barriers (the ~20% explained by selection) remain relevant but secondary. The roughly 20–25% of the gap attributable to genuine symptom prevalence differences is, by construction, guideline-consistent and should not be targeted for elimination.&lt;/p&gt;
&lt;p&gt;Q: What are the methodological contributions?
A: The paper makes three methodological contributions. First, it develops a structural model of mental health diagnosis that explicitly incorporates endogenous patient selection — a feature absent from standard physician decision-making models — which is shown empirically important. Second, it applies machine learning and NLP to clinical doctor note text to construct key unobserved clinical variables (behavioral assessment indicator and ADHD match signal) that are unavailable as structured data in EHRs. Third, the identification of population mean health risk uses a quasi-exogenous variation approach (IPCP referral rates) analogous to Arnold et al. (2022)&amp;rsquo;s method for measuring racial discrimination in bail decisions, adapted here to a continuous health risk setting with endogenous selection.&lt;/p&gt;
&lt;p&gt;Diagnostic threshold (τ_θ): The gender-specific posterior ADHD risk level above which a physician chooses to diagnose ADHD. Set ex-ante, it reflects the physician&amp;rsquo;s perceived tradeoff between the costs of over-diagnosis (misdiagnosis) and under-diagnosis (missed diagnosis). A lower threshold implies the physician views missed diagnosis as relatively more costly for that patient group. By construction, uniform clinical guidelines imply a single threshold independent of patient gender.&lt;/p&gt;
&lt;p&gt;ADHD match signal (x_i): A physician-observed, noisy signal of a patient&amp;rsquo;s true latent ADHD risk (v_i), observed only conditional on the patient receiving a behavioral assessment. In estimation, it is proxied via a cosine similarity measure between the patient&amp;rsquo;s aggregated clinical doctor note text and the DSM-V symptom list, constructed separately for inattentive and hyperactive/impulsive sub-types.&lt;/p&gt;
&lt;p&gt;Signal quality / diagnostic uncertainty (ρ_θ): The correlation between the physician&amp;rsquo;s observed ADHD match signal and the patient&amp;rsquo;s true ADHD risk. Higher ρ means the physician&amp;rsquo;s signal is more informative and diagnostic uncertainty is lower. In the Bayesian updating framework, higher ρ implies the physician places more weight on the observed signal relative to the prior.&lt;/p&gt;
&lt;p&gt;Mental healthcare utilization cost (c_i): The composite of all patient/caregiver factors that affect the decision to schedule a behavioral assessment net of child symptom level. Includes non-monetary barriers such as time constraints, distance, stigma, and information from primary care providers during wellness visits; does not include monetary out-of-pocket costs since insurance typically covers behavioral assessments.&lt;/p&gt;
&lt;p&gt;Initial Primary Care Provider (IPCP) referral rate: The regression-adjusted share of a given PCP&amp;rsquo;s patients who ultimately receive a behavioral assessment at some point in the sample. Used as a quasi-exogenous cost-shifter that influences patient scheduling costs without being correlated with patient ADHD risk, enabling identification of population mean ADHD risk via extrapolation.&lt;/p&gt;
&lt;p&gt;Latent ADHD risk (v_i): An unobserved continuous measure of a child&amp;rsquo;s underlying ADHD-related behavioral symptoms, drawn from a gender-specific normal distribution N(μ_θ, σ²_θ). A child&amp;rsquo;s true ADHD status is Si = 1(v_i &amp;gt; v̄), where v̄ is the DSM-V minimum symptom threshold, defined identically for boys and girls.&lt;/p&gt;
&lt;p&gt;Adjusted Bag-of-Words (BOW) cosine similarity: The NLP method used to construct the ADHD match signal proxy. Patient notes are tokenized into uni-grams and bi-grams after preprocessing (spell check, abbreviation replacement, part-of-speech tagging, synonym replacement), and tf-idf weighted. The cosine similarity between the resulting document vector and the DSM-V symptom text vector is computed separately for each ADHD sub-type and rescaled to [0,1].&lt;/p&gt;</description></item><item><title>On the Nature of Entrepreneurship</title><link>https://macropaperwarehouse.com/papers/on-the-nature-of-entrepreneurship/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/on-the-nature-of-entrepreneurship/</guid><description>&lt;p&gt;This paper uses a novel longitudinal administrative dataset drawn from U.S. Internal Revenue Service (IRS) and Social Security Administration (SSA) records to characterize income dynamics and the determinants of entrepreneurial entry for pass-through business owners — sole proprietors, partners, and S corporation owners — who collectively account for over 50 percent of all U.S. business net income. The sample covers 2000–2015 and includes up to 1.3 billion person-year observations for individuals aged 25–65. The authors construct balanced panels using birth cohorts 1950–1975, impute education (college attainment) and skill (cognitive, interpersonal, manual) via machine-learning classifiers trained on CPS and O*NET data, and estimate life-cycle income profiles using a three-component model that separates individual fixed effects, group-specific time effects, and group-cohort-specific age effects.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central departure from prior work is coverage of the full income distribution, including the high-earning right tail that household surveys such as the CPS misrepresent due to top-coding and small samples. When the IRS and CPS samples are compared on a consistent classification basis, median self-employment income is lower in the IRS data at all ages, consistent with the survey literature&amp;rsquo;s emphasis on the &amp;ldquo;typical&amp;rdquo; self-employed individual. However, mean incomes diverge sharply: the IRS shows mean self-employment income rising from $23 thousand at age 25 to $93 thousand at age 55, whereas the CPS (with incorporated owners reclassified) shows a rise from only $41 thousand to $73 thousand. Roughly 80 percent of self-employment income in the IRS data accrues to individuals above the $100 thousand threshold, compared to 42–53 percent in the CPS. The IRS-CPS gap is dominated by the right tail and concentrated in professional services and health care. For paid-employed individuals, the IRS and CPS medians and means are close at all ages, confirming the discrepancy is specific to self-employment.&lt;/p&gt;
&lt;p&gt;The life-cycle estimation finds that individuals who have &amp;ldquo;tried self-employment&amp;rdquo; — a group earning virtually all self-employment income — start at similar average incomes to primarily paid-employed peers at age 25 but reach $134 thousand by age 55, compared with $79 thousand for paid-employed peers with the same observable characteristics. Age effects for the self-employed are 63 percent higher than for the paid-employed at age 26 and remain elevated until age 55. Time effects show dramatically greater cyclical volatility for the self-employed: income growth declined by $9,655 (2008) and $8,785 (2009) for the self-employed versus $373 and $1,583 for paid-employed in the same years, concentrated in real estate and construction.&lt;/p&gt;
&lt;p&gt;On the determinants of entry, the paper finds: (i) no evidence that house-price appreciation raises entry rates, contra collateral-constraint hypotheses; (ii) most entrants have lower asset incomes than future entrants with the same characteristics, arguing against a liquid-wealth precondition; (iii) most entrants have higher prior labor income than future entrants, consistent with entry being driven by on-the-job experience rather than fallback from low-paid work; (iv) almost all founders report positive individual tax income in their first year of operation despite negative business net income and no external debt financing. Self-employed income growth exhibits greater dispersion — a 10th-to-90th percentile range roughly 2.5 times wider than for the paid-employed — and a Kelly skewness about 0.1 higher. A standard consumption-risk model calibrated with household-finance estimates of risk aversion rationalizes the patterns if individuals are insured against the most adverse downside shocks. Entry and exit rates are stable across the sample period, including the Great Recession, and the entrepreneurship share does not decline.&lt;/p&gt;
&lt;p&gt;The subgroup congruent with non-pecuniary motivation — primarily self-employed individuals earning less than paid-employed peers with matching characteristics — comprises roughly 57 percent of primarily self-employed by count but earns only 16 percent of total self-employment income.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-do-irs-and-cps-data-give-such-different-pictures-of-self-employment-income"&gt;Q1. Why do IRS and CPS data give such different pictures of self-employment income?&lt;/h3&gt;
&lt;p&gt;The CPS suffers from top-coding of high incomes and small samples that underrepresent high earners in key industries. The IRS-CPS mean income gap for the self-employed is dominated by the right tail: in the main IRS sample, individuals above the $100 thousand threshold earn roughly 80 percent of all self-employment income, versus 42 percent in the comparable CPS sample. The average income of top earners above $100 thousand is $355 thousand in the IRS versus $218 thousand in the CPS. The gap is concentrated in professional services and health care and persists across all income thresholds and sample definitions tested. No analogous discrepancy exists for paid-employed individuals, where IRS and CPS medians and means are close at all ages.&lt;/p&gt;
&lt;h3 id="q2-what-does-the-comparison-look-like-at-the-median-versus-the-mean"&gt;Q2. What does the comparison look like at the median versus the mean?&lt;/h3&gt;
&lt;p&gt;At the median, IRS self-employment income is lower than both CPS samples at all ages, with the gap largest for younger owners and those with incorporated businesses — a pattern consistent with the survey-based &amp;ldquo;self-employment discount&amp;rdquo; narrative. At the mean, the IRS shows much higher income at older ages: by age 55, IRS mean self-employment income is $93 thousand versus $73 thousand in the CPS sample that includes reclassified incorporated-owner wages. The divergence arises because the mean is sensitive to the right tail, which the CPS systematically underrepresents.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-estimate-life-cycle-income-profiles-while-separating-age-time-and-cohort-effects"&gt;Q3. How does the paper estimate life-cycle income profiles while separating age, time, and cohort effects?&lt;/h3&gt;
&lt;p&gt;Individual income is decomposed into an individual fixed effect (permanent latent ability and preferences), a group-specific time effect (business-cycle fluctuations common to a group), and a group-cohort-specific age effect (life-cycle income growth). Identification exploits the overlapping cohort structure of the 16-year panel: age effects are assumed equal across cohort bins of size at least two, allowing time and age effects to be separately identified. The model is estimated in levels rather than logs to accommodate business losses. Groups are defined as a Cartesian product of 32,256 subgroups based on education, three skill dimensions, industry (21 two-digit NAICS codes), demographics (gender, cohort, marital status, children), and employment-status history.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-headline-life-cycle-income-profile-findings-for-self--versus-paid-employed"&gt;Q4. What are the headline life-cycle income profile findings for self- versus paid-employed?&lt;/h3&gt;
&lt;p&gt;Among the &amp;ldquo;primarily employed&amp;rdquo; group, those who have tried self-employment and those who are primarily paid-employed have similar average incomes at age 25. By age 55 the self-employed reach an estimated $134 thousand (2012 dollars) versus $79 thousand for paid-employed peers with identical observable characteristics. The estimated age effect for the self-employed is 63 percent higher than for the paid-employed at age 26 and remains higher through age 55. These gaps would widen further if incomes were adjusted upward for the BEA-estimated net misreporting rates of 46 percent for unincorporated owners and 14 percent for S corporation owners.&lt;/p&gt;
&lt;h3 id="q5-how-large-is-the-group-consistent-with-non-pecuniary-motivation-and-how-much-income-does-it-earn"&gt;Q5. How large is the group consistent with non-pecuniary motivation, and how much income does it earn?&lt;/h3&gt;
&lt;p&gt;The non-pecuniary subgroup — primarily self-employed individuals (at least 12 years in self-employment) who earn less on average than primarily paid-employed peers matched on gender, education, skills, and other characteristics — is numerically larger, comprising approximately 57 percent of primarily self-employed by count. However, this group earns only 16 percent of total self-employment income. Adjusting for paid-employed fringe benefits and self-employed income misreporting can change the group&amp;rsquo;s size but does not alter the finding that it accounts for a small income share. The paper concludes that non-pecuniary motives may guide occupational choice for many individuals but are not the driver of the typical dollar earned in self-employment.&lt;/p&gt;
&lt;h3 id="q6-how-does-idiosyncratic-income-risk-compare-between-self--and-paid-employed"&gt;Q6. How does idiosyncratic income risk compare between self- and paid-employed?&lt;/h3&gt;
&lt;p&gt;Self-employed income changes are substantially more dispersed: the 10th-to-90th percentile range of income growth is roughly 2.5 times wider for the self-employed than for the paid-employed. Income changes for the self-employed are also more right-skewed, with a Kelly skewness difference of approximately 0.1. When a standard consumption-risk model — augmented with a lower bound on consumption growth to allow for external insurance — is parameterized with risk-aversion estimates from the household finance literature, the observed patterns are rationalized if individuals are insured against the most adverse downside shocks, i.e., the attractive aspect of self-employment is large potential upside with insured downside.&lt;/p&gt;
&lt;h3 id="q7-what-happened-to-self-employed-income-and-exit-rates-during-the-great-recession"&gt;Q7. What happened to self-employed income and exit rates during the Great Recession?&lt;/h3&gt;
&lt;p&gt;Time effects show steep income growth declines for the self-employed of -$9,655 in 2008 and -$8,785 in 2009, compared with much more modest declines of -$373 and -$1,583 for paid-employed peers. The aggregate income declines are concentrated in cyclically sensitive self-employed subgroups in real estate and construction, with their paid-employed counterparts experiencing only modest declines. Despite these large income shocks, exit rates from self-employment showed little change during the Great Recession, either in aggregate or in the cyclically sensitive sectors. Entry rates were likewise stable, and the share of entrepreneurs in the population did not decline over the full sample period.&lt;/p&gt;
&lt;h3 id="q8-does-the-evidence-support-collateral-constraints-as-a-binding-barrier-to-entrepreneurial-entry"&gt;Q8. Does the evidence support collateral constraints as a binding barrier to entrepreneurial entry?&lt;/h3&gt;
&lt;p&gt;No. The paper tests the hypothesis, standard in the liquidity-constraints literature, that entry rates should be higher for homeowners experiencing house-price appreciation (which raises collateral value). The IRS data do not support this prediction. Separately, comparing asset incomes (interest, dividends, capital gains) of current entrants and future entrants with the same characteristics, the paper finds that most current entrants have lower asset incomes and less liquid wealth than those who switch later, which also argues against a liquid-wealth precondition for entry.&lt;/p&gt;
&lt;h3 id="q9-what-does-prior-labor-income-reveal-about-why-people-enter-self-employment"&gt;Q9. What does prior labor income reveal about why people enter self-employment?&lt;/h3&gt;
&lt;p&gt;Current entrants have higher prior labor income than matched future entrants with the same characteristics, indicating they enter with accumulated on-the-job experience rather than being pushed into self-employment as a fallback after failure in paid work. This is consistent with self-employment being a deliberate, experience-driven career transition for most entrants rather than a last resort for low earners. The paper interprets this as positive evidence for the role of experience-based human capital in driving entrepreneurial choice.&lt;/p&gt;
&lt;h3 id="q10-how-do-founders-finance-startup-costs-if-most-have-negative-business-net-income-in-early-years"&gt;Q10. How do founders finance startup costs if most have negative business net income in early years?&lt;/h3&gt;
&lt;p&gt;Almost all founders in the sample report positive income on their personal (individual) tax form in the first year of operation, even though most report negative business net income and carry no external debt financing. This pattern suggests founders rely on personal income sources — prior savings, part-time paid employment, or spousal income — to cover startup costs rather than external debt, implying that formal credit-market financing constraints are not the primary barrier to entry for most entrants in the sample.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-scope-conditions-and-key-limitations"&gt;Q11. What are the scope conditions and key limitations?&lt;/h3&gt;
&lt;p&gt;The sample covers pass-through owners (sole proprietors, partners, S corporation owners) and excludes C corporation shareholders, whose entrepreneurial income does not flow to individual returns until distributed. Income measures exclude most employer fringe benefits; capital gains are excluded from self-employment income, and the authors note their inclusion would strengthen the main findings. The analysis covers 2000–2015 for cohorts born 1950–1975, and income is reported before taxes and transfers. Baseline estimates are not adjusted for misreporting, though BEA-implied adjustments of 46 percent for unincorporated owners and 14 percent for S corporation owners would widen the income gaps further.&lt;/p&gt;
&lt;p&gt;Pass-through business owner: An individual who owns a sole proprietorship, partnership, or S corporation, such that business net income flows directly onto the owner&amp;rsquo;s personal tax return; excludes C corporation shareholders whose income appears only upon dividend or capital-gains distributions.&lt;/p&gt;
&lt;p&gt;Tried self-employment: The paper&amp;rsquo;s primary self-employed comparison group within the &amp;ldquo;primarily employed&amp;rdquo; category — individuals with any years in self-employment (including frequent switchers and those with most years in self-employment) — who collectively earn virtually all self-employment income.&lt;/p&gt;
&lt;p&gt;Group-specific age effect: The paper&amp;rsquo;s estimate of how individual income changes with age within a defined subgroup (determined by education, skill, industry, demographics, and employment history), identified by exploiting overlapping birth cohorts in the 16-year panel and separated from individual fixed effects and business-cycle time effects.&lt;/p&gt;
&lt;p&gt;Primarily employed: Individuals with at least 12 of 16 sample years in either self- or paid-employment, with at most one intermediate year of non-employment; the paper&amp;rsquo;s main analytical focus for life-cycle income comparisons.&lt;/p&gt;
&lt;p&gt;SOI Databank: The Statistics of Income Databank, a de-identified balanced panel combining SSA demographic records with IRS tax filing data for all living U.S. individuals with a Social Security number over 1996–2015; the paper&amp;rsquo;s primary data source providing Schedule C, K-1, W-2, and related filing information.&lt;/p&gt;
&lt;p&gt;Kelly skewness: A robust measure of distributional asymmetry used by the paper to characterize income growth; the paper reports that Kelly skewness of self-employed income changes exceeds that of paid-employed by approximately 0.1, indicating greater right-skewness in self-employment income dynamics.&lt;/p&gt;
&lt;p&gt;Non-pecuniary motivation subgroup: Primarily self-employed individuals who earn less on average than primarily paid-employed peers matched on observable characteristics, taken by the paper as consistent with non-wage job amenities (autonomy, flexibility) driving occupational choice; found to be 57 percent of primarily self-employed by count but earning only 16 percent of total self-employment income.&lt;/p&gt;</description></item><item><title>Online Business Models, Digital Ads, and User Welfare</title><link>https://macropaperwarehouse.com/papers/online-business-models-digital-ads-and-user-welfare/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/online-business-models-digital-ads-and-user-welfare/</guid><description>&lt;p&gt;Acemoglu, Huttenlocher, Ozdaglar, and Siderius develop a two-sided platform model to study the welfare consequences of digital advertising as an online business model. The platform intermediates between a firm selling a horizontally differentiated product and a continuum of users who derive utility from both entertaining content and informative signals about product quality embedded in ads. Users have a two-dimensional type: a sophistication dimension (sophisticated with probability lambda, naïve with probability 1-lambda) and a product-quality dimension (high quality with prior probability q). The central departure from the standard informational-advertising literature is that sophisticated users hold the correct model of the ad signal process, while naïve users underestimate the false-positive rate — the probability that a low-quality product generates a positive ad signal (phi_0). Naïve users perceive this false-positive rate to be phi_{0,N} = omega_N * omega_P * phi_0, where omega_N &amp;lt;= 1 captures inherent naïveté and omega_P &amp;lt;= 1 captures failure to understand personalized targeting, so phi_{0,N} &amp;lt; phi_0. The equilibrium concept is Berk-Nash equilibrium (Esponda and Pouzo 2016), meaning all agents are Bayesian given their subjective model.&lt;/p&gt;
&lt;p&gt;The platform chooses ad load alpha (Poisson rate of ad displays), subscription fees, and the monetary transfer from the firm; the firm sets product price p after observing the platform&amp;rsquo;s contract. The central finding (Proposition 2) is that when the objective false-positive rate phi_0 exceeds a threshold phi-hat_0(lambda, phi_1, phi_{0,N}) — which is increasing in lambda and phi_{0,N} and decreasing in the true-positive rate phi_1 — the unique equilibrium is an advertising-based plan that fully segments the market: naïve users receive an ad load that extracts all their surplus, while sophisticated users are excluded entirely. In this regime the firm charges a strictly higher price p-hat* &amp;gt; p-bar*, where p-bar* = (beta*q + c)/2 is the monopoly price without advertising. The ad-based equilibrium emerges precisely when ads are more misleading (larger gap between phi_0 and phi_{0,N}), not when they are more informative — a comparative static the authors describe as paradoxical.&lt;/p&gt;
&lt;p&gt;Welfare consequences (Proposition 4) are unambiguous in the advertising regime: both naïve and sophisticated users are strictly worse off than the baseline without any platform. Naïve users over-purchase due to inflated posteriors from misread signals; sophisticated users are harmed through the price channel — the firm&amp;rsquo;s higher profit-maximizing price p-hat* applies to all buyers. In the fully rational benchmark (phi_{0,N} = phi_0), the unique equilibrium is subscription-based and user welfare equals the no-platform baseline (Proposition 3).&lt;/p&gt;
&lt;p&gt;These results extend to richer menus (Proposition 5), mixed subscription-plus-advertising plans (Proposition 7), and to multi-firm and multi-platform competition (Propositions 9-12). Digital ads soften Bertrand competition by generating endogenous horizontal differentiation among otherwise identical firms, so equilibrium prices can exceed marginal cost even with two competing firms. Platform competition similarly fails to restore welfare: platforms compete away subscription fees but both adopt ad-based plans targeting naïfs when phi_1 exceeds a threshold, maintaining the welfare loss.&lt;/p&gt;
&lt;p&gt;On policy, the first best (planner observes types) cannot be decentralized because naïve users prefer more ads than is socially optimal, inverting the usual self-selection constraint. The second best (planner subject to incentive-compatibility constraints) is a single pooling plan with an intermediate ad load alpha^{SB} in [alpha^{FB}_N, alpha^{FB}_S] and yields average welfare above the no-platform baseline, though below first best (Proposition 13). This second best can be decentralized with a nonlinear digital ad tax, a per-unit product subsidy, and a platform subscription subsidy (Proposition 14). A simpler flat tax on digital ad revenues — above a threshold gamma-bar &amp;lt; 1 — also improves welfare relative to the ad-based equilibrium, though it does not restore the second best (Proposition 15).&lt;/p&gt;
&lt;p&gt;Four robustness extensions are developed: endogenous manipulation (platform always chooses the most manipulative environment, lowest phi_{0,N}); naïve learning dynamics (learning raises the sophisticate share in steady state, making ad-based models less profitable but not overturning the main results); imperfect price discrimination by the firm (naïfs are unambiguously worse off, threshold for advertising equilibrium shifts down); and an added price-sensitivity dimension (the platform runs a 2x2 menu separating by both sophistication and price sensitivity, preserving the result that naïve users tolerate and receive more ads than sophisticates in every stratum).&lt;/p&gt;
&lt;p&gt;Q: What is the key asymmetry between naïve and sophisticated users that drives the main results?
A: Sophisticated users hold the correct Bayesian model of the ad signal process and thus correctly account for the false-positive rate phi_0 when updating beliefs from positive ad signals. Naïve users perceive the false-positive rate as phi_{0,N} = omega_N * omega_P * phi_0 &amp;lt; phi_0, so they treat positive signals as stronger evidence of high product quality than they actually are. Because naïve users overestimate the informativeness of ads, their (interim) subjective valuation of an ad-based plan is higher, making them more tolerant of ad loads and more willing to join platforms with heavy advertising. This asymmetry is what makes it profitable to target naïfs with high ad loads while excluding or charging subscription fees to sophisticates.&lt;/p&gt;
&lt;p&gt;Q: Why does advertising to sophisticated users generate no additional firm profit, while advertising to naïve users does?
A: Lemma 1 establishes that with linear-quadratic utility the firm extracts no surplus from advertising to sophisticates: because sophisticated agents are fully Bayesian, their expected posterior equals the prior (E_S[pi_i] = q), so expected demand after advertising is identical to demand before advertising. By contrast, Lemma 2 shows that the firm&amp;rsquo;s profit from naïve agents is positive and strictly increasing in ad load alpha, because naïve users&amp;rsquo; average demand curve drifts upward as alpha rises — their inflated perceived informativeness of ads causes them to over-update on positive signals, systematically raising their willingness to pay. The platform captures this surplus from the firm via the advertising transfer m*.&lt;/p&gt;
&lt;p&gt;Q: What is the threshold condition determining whether the equilibrium is subscription-based or advertising-based?
A: Proposition 2 identifies a threshold phi-hat_0(lambda, phi_1, phi_{0,N}) that is increasing in the sophisticate share lambda and in the naïve false-positive perception phi_{0,N}, and decreasing in the true-positive rate phi_1. When the objective false-positive rate phi_0 is below this threshold, the profit-maximizing business model is subscription-based with price P* = T - v and product price p* = p-bar* = (beta&lt;em&gt;q + c)/2. When phi_0 exceeds the threshold, the advertising model dominates: the platform sets a high ad load alpha-hat&lt;/em&gt; that makes naïve users exactly indifferent between participating and their outside option v, excludes sophisticates, and the firm charges p-hat* &amp;gt; p-bar*. The threshold falls with phi_1, meaning more informative ads expand the range of phi_0 over which the advertising equilibrium obtains.&lt;/p&gt;
&lt;p&gt;Q: How does allowing the platform to offer menus change the results relative to the baseline two-plan case?
A: Proposition 5 shows that with menus the platform can simultaneously serve both user types: sophisticates receive a subscription plan at P* = T - v and naïve users receive an ad-based plan with the same high load alpha-hat* as in the baseline. The threshold for the advertising equilibrium shifts down to phi*&lt;em&gt;0(lambda, phi_1, phi&lt;/em&gt;{0,N}) &amp;lt; phi-hat_0, so advertising business models arise for a strictly larger set of parameters. Welfare consequences are unchanged (Corollary 1): when phi_0 &amp;gt; phi*_0, both types have welfare strictly below the no-platform baseline. Proposition 6 further shows consumer welfare is monotonically decreasing in both phi_0 and phi_1: higher phi_1 (more informative true-positive signals) also reduces welfare because any surplus from greater informativeness is fully captured by the platform.&lt;/p&gt;
&lt;p&gt;Q: What is the welfare ranking across the three regimes: no platform, advertising equilibrium, and subscription equilibrium?
A: In the subscription equilibrium (regime (a) of Proposition 2 or 4), user welfare for both types equals the no-platform base case W_base(tau) — the platform captures all surplus it creates and users are no better or worse off. In the advertising equilibrium (regime (b)), both naïve and sophisticated users are strictly worse off than with no platform: W-hat*(tau) &amp;lt; W_base(tau) for both tau in {S, N}. The first-best, where a planner controls ad loads separately by type, yields W^{FB}(tau) &amp;gt; W_base(tau) for both types because informative ads can genuinely improve sophisticated users&amp;rsquo; decisions and a constrained amount improves naïve users&amp;rsquo; decisions too.&lt;/p&gt;
&lt;p&gt;Q: How does firm-level competition interact with digital advertising to affect prices and welfare?
A: Without advertising, two ex ante identical firms compete à la Bertrand and price at marginal cost (p*_1 = p*_2 = c). Proposition 9 establishes that when phi_1 &amp;gt; phi^F_1 and phi_0 &amp;gt;= phi^F_0(phi_1), the platform offers an ad-based plan and equilibrium prices p-hat*_1 and p-hat*_2 are both strictly above p-bar* — the monopoly price without advertising. The mechanism is endogenous horizontal differentiation: users who see positive ad signals for one firm&amp;rsquo;s product form higher valuations for that product, so the two products become differentiated in the eyes of consumers even though they are ex ante identical, breaking Bertrand logic. Example 1 further illustrates that advertising can be more prevalent with competition than without: a second firm&amp;rsquo;s entry can push the equilibrium from no-advertising to separating.&lt;/p&gt;
&lt;p&gt;Q: Does platform competition protect users from the welfare losses associated with digital advertising?
A: Not fully. Proposition 11 shows that with two competing platforms (M=2, N=1) and no advertising, platforms compete away both subscription fees and ad loads, and welfare reaches the fully rational benchmark. However, when phi_1 exceeds threshold phi^P_1, both platforms adopt ad-based plans targeting naïve users, charge no subscription fees, and the product price rises to p-hat*_P &amp;gt; p-bar* (Proposition 12). Competition reduces subscription fees to zero but does not eliminate the incentive to target naïfs with heavy ads, because naïve users&amp;rsquo; over-valuation of ads means they remain willing to join ad-heavy plans. The fundamental inefficiency from naïve users&amp;rsquo; misspecified model persists under platform competition.&lt;/p&gt;
&lt;p&gt;Q: Why is the first-best allocation not implementable as a decentralized equilibrium?
A: Proposition 13 explains the obstacle: the social planner would ideally offer naïve users fewer ads (alpha^{FB}_N) than sophisticated users (alpha^{FB}_S), with alpha^{FB}_N &amp;lt;= alpha^{FB}_S. However, naïve users have a higher subjective valuation for ads than sophisticates because they believe ads are more informative. If offered a menu with both options, naïve users would self-select into the plan with the higher ad load alpha^{FB}_S — the exact opposite of what the planner wants. The incentive-compatibility constraints therefore force the planner toward a single pooling plan with an intermediate ad load alpha^{SB} in [alpha^{FB}_N, alpha^{FB}_S]. Average welfare under the second best exceeds the no-platform baseline, confirming that some advertising is socially valuable, but falls short of the first best whenever alpha^{FB}_N &amp;gt; 0.&lt;/p&gt;
&lt;p&gt;Q: How does a flat digital ad tax improve welfare, and what are its limitations?
A: Proposition 15 establishes that whenever the equilibrium features an ad-based plan, a flat tax on digital ad revenues at rate gamma &amp;gt; gamma-bar &amp;lt; 1 improves welfare by discouraging advertising-based business models and inducing the platform to shift toward subscription-based plans. The mechanism is that taxing ad revenue reduces the platform&amp;rsquo;s marginal gain from increasing ad load, making the subscription plan relatively more profitable. However, the flat tax does not achieve the second best because it operates linearly rather than targeting the nonlinear distortion: the optimal nonlinear tax-subsidy scheme (Proposition 14) requires a threshold-style ad tax at rate mu &amp;gt; mu-bar combined with a per-unit product subsidy delta* and a platform subscription subsidy eta &amp;gt; eta-bar.&lt;/p&gt;
&lt;p&gt;Q: What happens when the platform can endogenously choose how manipulative its ads are?
A: Proposition 16 shows that a profit-maximizing platform always chooses the lowest feasible phi_{0,N} = phi-bar — the most manipulative environment. Two reinforcing channels drive this: the pricing channel (lower phi_{0,N} amplifies naïve demand shifts per positive signal, so the downstream firm raises price and sales, increasing ad revenues extracted by the platform) and the participation channel (lower phi_{0,N} raises naïve users&amp;rsquo; perceived informational value of ads, relaxing their participation constraint and permitting a higher ad load alpha). Platform competition constrains the equilibrium ad load through tighter participation constraints but does not alter the choice of phi_{0,N} = phi-bar, so competition limits ad quantity but not ad manipulativeness.&lt;/p&gt;
&lt;p&gt;Q: How do naïve learning dynamics affect the main results?
A: Proposition 17 introduces a birth-death environment where exposure to disconfirming evidence gradually converts naïve agents to sophisticates. A unique steady-state sophisticate share lambda*(alpha_N, phi_0) exists; both higher ad load alpha_N and higher phi_0 accelerate the conversion of naïfs, raising future sophisticate share and reducing future ad revenues. This creates a new intertemporal trade-off that constrains the platform&amp;rsquo;s choice of ad loads relative to the static case. The key result (part ii) is that the main characterization of Proposition 7 carries through under a modified cutoff phi-tilde^{dynamic}&lt;em&gt;0 &amp;gt;= phi-tilde_0(lambda-tilde, phi_1, phi&lt;/em&gt;{0,N}), so learning dynamics make the ad-based business model less likely but do not overturn the fundamental welfare results.&lt;/p&gt;
&lt;p&gt;Q: How does imperfect price discrimination by the firm affect naïve users?
A: Proposition 18 considers a firm that observes a user&amp;rsquo;s sophistication type with probability kappa in [0,1]. With price discrimination, the firm sets type-specific prices satisfying p*_N &amp;gt;= p* &amp;gt;= p*_S, moving toward the type-specific monopoly levels. Naïfs are unambiguously worse off: when identified (with probability kappa), they face the higher price p*_N and a higher equilibrium ad load. The threshold for the advertising equilibrium also shifts down relative to the baseline, meaning advertising business models emerge for a larger parameter range when price discrimination is possible.&lt;/p&gt;
&lt;p&gt;Q: How does the paper define and measure user welfare, and why is ex post rather than interim welfare the relevant concept?
A: User welfare W(tau_i) is defined as ex post utility, which depends on the actual product quality theta_i realized after consumption, not on interim beliefs formed after viewing ads. Naïve users&amp;rsquo; interim assessment inflates expected product quality, but their ex post utility depends on whether the product is genuinely high quality for them (theta_i = 1 with probability q, theta_i = 0 with probability 1-q). Because naïve users over-purchase due to misread signals — consuming more than optimal when theta_i = 0 — their ex post utility is strictly lower than their interim expected utility, and strictly lower than the no-platform baseline in the advertising equilibrium. The ex post welfare concept is the relevant one precisely because it captures the actual material consequences of manipulation, not the subjectively perceived gains from ads.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Naïve vs. Sophisticated Users&lt;/strong&gt;: The paper&amp;rsquo;s primary user heterogeneity dimension. Sophisticated users hold the correct model of the ad signal process, setting phi_{0,S} = phi_0 (the true false-positive rate). Naïve users hold a misspecified model with phi_{0,N} = omega_N * omega_P * phi_0 &amp;lt; phi_0, underestimating the probability that a low-quality product generates a positive ad signal, due to inherent naïveté (omega_N) and failure to understand personalized targeting (omega_P).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ad Load (alpha)&lt;/strong&gt;: The Poisson rate at which ads are displayed to a user per unit time. Total ad displays follow a Poisson(alpha*T) distribution. Higher ad load means less time on entertaining content — expected entertainment time is (1-alpha)&lt;em&gt;T — and a higher probability (1 - exp(-alpha&lt;/em&gt;T)) that the user sees the ad at least once. The platform chooses alpha as its primary instrument for extracting surplus from naïve users.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;False-Positive Rate (phi_0)&lt;/strong&gt;: The objective probability that a low-quality product (theta_i = 0) generates a positive (&amp;ldquo;good&amp;rdquo;) ad signal. The gap between phi_0 (objective) and phi_{0,N} (naïve users&amp;rsquo; perceived rate) is the key parameter driving all welfare results: a larger gap implies greater de facto manipulation and a stronger incentive for the platform to adopt an advertising-based model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Berk-Nash Equilibrium&lt;/strong&gt;: The solution concept from Esponda and Pouzo (2016), used to model agents with misspecified subjective models. All agents are Bayesian conditional on their own subjective model. Sophisticates&amp;rsquo; subjective model equals the objective model (standard Bayesian), while naïfs update using the misspecified phi_{0,N}. Perfection requires sequential rationality at each information set given beliefs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;De Facto Manipulation&lt;/strong&gt;: The paper&amp;rsquo;s term for a situation in which the platform and firm exploit naïve users&amp;rsquo; misspecified model to boost demand and extract surplus, without requiring any outright deception in the formal sense. It arises because naïve users voluntarily choose high-ad-load plans (believing ads to be highly informative) and voluntarily over-purchase (having updated on what they mistakenly think are strong positive signals). The manipulation is &amp;ldquo;de facto&amp;rdquo; because it operates through the users&amp;rsquo; own rational (but misspecified) decision-making.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Separating Equilibrium&lt;/strong&gt;: An equilibrium in which naïve and sophisticated users self-select into distinct platform plans. In the advertising equilibrium, naïve users join an ad-heavy plan (extracting all their surplus via inflated willingness to pay for ads) while sophisticated users are either excluded or placed on a subscription plan. This separation is the vehicle through which the platform maximizes revenue from naïf manipulation while limiting the disciplining force of sophisticates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-Best Allocation&lt;/strong&gt;: The welfare-maximizing allocation subject to the incentive-compatibility constraints that users self-select into plans. Because naïve users prefer more ads than sophisticated users (the inverse of what the planner desires), the second best is a single pooling plan with an intermediate ad load alpha^{SB} in [alpha^{FB}_N, alpha^{FB}_S]. This is strictly worse than the first best but achieves average welfare above the no-platform baseline, and can be decentralized with a nonlinear ad tax, product subsidy, and platform subscription subsidy.&lt;/p&gt;</description></item><item><title>Optimal Taxation and Market Power</title><link>https://macropaperwarehouse.com/papers/optimal-taxation-and-market-power/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-taxation-and-market-power/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether and how optimal income taxation should change when firms have market power. The question is motivated by the documented rise in economy-wide markups since 1980, which has compressed the labor share, widened the gap between worker and entrepreneurial income, and generated allocative inefficiency through excessive pricing.&lt;/p&gt;
&lt;p&gt;The authors develop a Mirrleesian optimal taxation framework augmented with three features absent from the canonical literature: (i) oligopolistic intermediate goods markets with endogenous, variable markups, (ii) heterogeneous firm productivities, and (iii) two occupational groups—wage-earning workers and profit-earning entrepreneurs—whose abilities are private information. Entrepreneurs strategically set prices under Cournot competition, which means that the tax system affects profits both through a firm&amp;rsquo;s own behavior and through the responses of its competitors. This strategic interaction is the critical novelty relative to prior work that assumes monopolistic competition.&lt;/p&gt;
&lt;p&gt;The main theoretical contribution is the derivation of optimal tax formulas for both labor income and profit income that decompose into four named components: (i) the Mirrleesian incentive component, which reflects the standard trade-off between redistribution and labor supply distortions; (ii) the Pigouvian component, which corrects for the externality from market power by subsidizing labor and entrepreneurial effort to offset the output shortfall from high markups; (iii) the Reallocation Effect (RE), which shifts the profit tax to redirect labor inputs from low-markup firms to high-markup firms where labor is inefficiently scarce, and which emerges only under heterogeneous markups; and (iv) the Indirect Redistribution Effect (IRE), which uses changes in competitors&amp;rsquo; product prices—a channel present only under oligopolistic (not monopolistic) competition—to redistribute income between entrepreneurs.&lt;/p&gt;
&lt;p&gt;For the labor income tax, the dominant force is the Pigouvian component. As average markups rise, the Pigouvian subsidy to labor supply grows, mechanically reducing optimal labor income tax rates. The profit tax is shaped by all four components in opposing directions; the net quantitative effect is resolved empirically.&lt;/p&gt;
&lt;p&gt;The model is calibrated to match distributions of labor income (from the Current Population Survey), profits (from Compustat-based data in De Loecker, Eeckhout, and Unger 2020), and firm-level markups (also from De Loecker, Eeckhout, and Unger 2020, using the cost-minimization approach) for the US in 1980 and 2019. The cost-weighted average markup rose from 1.25 in 1980 to 1.33 in 2019, with the increase concentrated at the top of the markup distribution.&lt;/p&gt;
&lt;p&gt;The central quantitative prescription is that the optimal labor income tax rate should decline by 7.7 percentage points between 1980 and 2019 (average optimal rate falls from 22.0 percent to 14.3 percent), while the optimal profit tax rate should rise by 2.2 percentage points on average (from 58.4 percent to 60.5 percent) and by 29.1 percentage points at the top. The decline in the labor income tax is driven primarily by the rise in average markups reducing the Pigouvian component. The increase in the profit tax, especially at the top, is driven primarily by the Mirrleesian component operating through the skill gap, which rises because higher markups reduce profit elasticity. The Pigouvian and reallocation components push in the opposite direction on the profit tax, but the Mirrleesian effect dominates.&lt;/p&gt;
&lt;p&gt;The optimal profit tax structure is regressive for large, high-markup firms—reflecting the RE, which requires lower tax rates for high-markup firms to incentivize labor reallocation toward them—but less regressive in 2019 than in 1980, reflecting the distributional tightening from rising markup inequality.&lt;/p&gt;
&lt;p&gt;Robustness checks across parameter values for the social welfare curvature k, the span of control ξ, and the elasticity of substitution σ confirm that the directional results hold: labor income tax rates decrease and profit tax rates increase from 1980 to 2019 across all parameter configurations. Extensions to nonlinear sales taxes and conditioning on markups confirm that even when the planner can observe markups directly, the first-best is not achievable because markups are endogenous to entrepreneurs&amp;rsquo; unobservable decisions.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-fundamental-difference-between-this-papers-model-and-prior-work-on-optimal-taxation-with-market-power"&gt;Q1. What is the fundamental difference between this paper&amp;rsquo;s model and prior work on optimal taxation with market power?&lt;/h3&gt;
&lt;p&gt;Prior work using monopolistic competition (e.g., Gürer 2021; Boar and Midrigan 2019) assumes each entrepreneur holds monopoly power in its own market, so no strategic interaction exists between firms. Under monopolistic competition, entrepreneurs price to maximize utility given competitors&amp;rsquo; choices, and the envelope theorem implies that tax changes have no first-order effect on prices or utility through the pricing channel—the Indirect Redistribution Effect (IRE) disappears. In this paper, entrepreneurs compete in Cournot oligopolistic markets with a finite number of firms I, so each firm&amp;rsquo;s pricing depends on competitors&amp;rsquo; output. A change in one firm&amp;rsquo;s output (induced by taxation) shifts competitors&amp;rsquo; prices, opening a redistribution channel through product markets that is entirely absent in monopolistic competition. Additionally, the Reallocation Effect (RE) emerges only when firm-level markups are heterogeneous, which requires oligopolistic (not perfectly competitive) markets.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-four-components-of-the-optimal-tax-formula-and-how-does-each-relate-to-market-power"&gt;Q2. What are the four components of the optimal tax formula and how does each relate to market power?&lt;/h3&gt;
&lt;p&gt;The optimal tax wedge for both labor and profit income decomposes into four components. First, the Mirrleesian component reflects the standard trade-off between redistribution and the efficiency cost of taxation; in the presence of market power, it is modified because the skill gap for entrepreneurs depends on markups through the profit elasticity. Second, the Pigouvian component corrects the externality from market power, which causes prices to exceed marginal cost and output to be inefficiently low; it implies a subsidy to both worker and entrepreneurial effort, scaled by the reciprocal of the average markup (for the labor tax) or firm-level markup (for the profit tax). Third, the Reallocation Effect (RE) applies only to the profit tax and reflects that labor should be shifted toward high-markup firms where it is inefficiently underemployed; it reduces the tax rate for firms whose markup exceeds the average. Fourth, the Indirect Redistribution Effect (IRE) captures redistribution through competitor price changes under oligopolistic interaction; it can either raise or lower the profit tax rate depending on the distribution of social welfare weights and the cross-inverse demand elasticity.&lt;/p&gt;
&lt;h3 id="q3-what-happens-to-the-labor-income-tax-formula-as-average-markups-rise"&gt;Q3. What happens to the labor income tax formula as average markups rise?&lt;/h3&gt;
&lt;p&gt;The labor income tax formula contains a Pigouvian component equal to the reciprocal of the employment-weighted average markup. As average markups rise, this reciprocal falls, reducing the optimal labor income tax rate. Quantitatively, the optimal average labor income tax rate declines from 22.0 percent in 1980 to 14.3 percent in 2019, a decrease of 7.7 percentage points. In a purely competitive benchmark economy, the top labor income tax rate would be around 60 percent (consistent with Saez 2001); in the calibrated model with market power, it is 34.2 percent in 1980 and 28.7 percent in 2019. The Pigouvian component accounts for essentially the entire difference because the Mirrleesian component, when calibrated to the same labor income distribution, is unchanged.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-mirrleesian-component-cause-the-top-profit-tax-rate-to-rise-with-market-power"&gt;Q4. How does the Mirrleesian component cause the top profit tax rate to rise with market power?&lt;/h3&gt;
&lt;p&gt;The Mirrleesian component of the profit tax is driven by the skill gap, defined as the proportional rate of change in the composite entrepreneur ability measure. The skill gap depends on markups through the profit elasticity: as markups rise, profit elasticity falls (since profit elasticity is approximately the reciprocal of markup minus the span-of-control parameter minus the inverse of the labor supply elasticity term), which increases the skill gap. A higher skill gap amplifies the income divergence across entrepreneur types, increasing the Mirrleesian incentive to redistribute at the top. Quantitatively, Figure 5 shows that the rise in the skill gap from 1980 to 2019 tracks almost exactly the change in the inverse of profit elasticity, confirming that markup changes—not changes in the ability distribution—are the primary driver of increased Mirrleesian pressure on top profit taxes.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-reallocation-effect-influence-the-structure-progressivity-of-the-profit-tax"&gt;Q5. How does the Reallocation Effect influence the structure (progressivity) of the profit tax?&lt;/h3&gt;
&lt;p&gt;The RE term equals the ratio of the average markup to the firm-level markup minus one: RE(θe) = μ/μ(θe) − 1. For firms with markups above the average, RE is negative, reducing their optimal tax rate; for firms below the average, RE is positive, increasing it. This implies that the optimal profit tax should be regressive relative to markup (i.e., high-markup firms face lower marginal tax rates), even though the overall profit tax rises on average. This provides a novel rationale for why the profit tax schedule in practice is less progressive—or even regressive—for large firms. As markups rise across the distribution, the reallocation effect pushes down the top profit tax but does not offset the larger increase from the Mirrleesian component in the quantitative exercise.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-indirect-redistribution-effect-and-why-does-it-disappear-under-monopolistic-competition"&gt;Q6. What is the Indirect Redistribution Effect and why does it disappear under monopolistic competition?&lt;/h3&gt;
&lt;p&gt;The IRE captures the change in entrepreneurial utility that arises because a tax reduction for one entrepreneur increases their output, which reduces the prices of substitute goods produced by competitors, thereby lowering competitors&amp;rsquo; incomes. Under oligopolistic competition with I &amp;gt; 1 firms per market, the cross-inverse demand elasticity is nonzero, so competitor prices are sensitive to any one firm&amp;rsquo;s output decision, and this redistribution channel is open. Under monopolistic competition (I = 1), each entrepreneur is the sole producer in its market; competitors&amp;rsquo; prices do not depend on the firm&amp;rsquo;s output, the cross-inverse demand elasticity is zero, and the IRE vanishes by the envelope theorem. The IRE is also absent in perfectly competitive economies. Empirical evidence for the US suggests the hazard ratio of profits is sufficiently high that the IRE generally pushes toward a lower top profit tax rate, but the Mirrleesian effect dominates in the quantitative results.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-quantitative-effect-of-rising-markups-on-the-optimal-tax-rates-and-what-drives-the-net-change-in-the-profit-tax"&gt;Q7. What is the quantitative effect of rising markups on the optimal tax rates, and what drives the net change in the profit tax?&lt;/h3&gt;
&lt;p&gt;The model calibrated to 1980 and 2019 US data prescribes a decline in the optimal average labor income tax rate of 7.7 percentage points (from 22.0 to 14.3 percent) and an increase in the optimal average profit tax rate of 2.2 percentage points (from 58.4 to 60.5 percent). At the top of the profit distribution, the increase is 29.1 percentage points. The net profit tax increase results from four opposing forces: the Pigouvian component falls (pushing toward lower taxes) and the RE decreases for high-markup firms (also pushing down the top rate), while the IRE and especially the Mirrleesian component rise (pushing up top rates). The Mirrleesian effect is the dominant force, driven by rising markup inequality reducing profit elasticity and widening the skill gap for top entrepreneurs.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-counterfactual-analysis-isolate-the-role-of-markups-from-productivity-changes"&gt;Q8. How does the counterfactual analysis isolate the role of markups from productivity changes?&lt;/h3&gt;
&lt;p&gt;The counterfactual fixes the markup distribution at its 1980 level while holding the 2019 productivity distribution constant, then solves for optimal taxes. The result is that high-profit entrepreneurs would face lower optimal tax rates under 1980 markups than under 2019 markups, while low-profit entrepreneurs would face higher rates. Decomposing the difference, the Pigouvian component and the RE are larger for high incomes under 1980 (lower) markups, making the profit tax more regressive, while the IRE and the Mirrleesian component are smaller under 1980 markups, producing a lower top rate. The increase in the Mirrleesian component due to the markup increase from 1980 to 2019 is identified as the primary reason top profit taxes rise. This isolates the markup channel from the productivity channel in accounting for changes in optimal taxes.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-robustness-analysis-reveal-about-parameter-sensitivity"&gt;Q9. What does the robustness analysis reveal about parameter sensitivity?&lt;/h3&gt;
&lt;p&gt;The main qualitative result—labor income taxes decline and profit taxes rise from 1980 to 2019—holds across a broad parameter space. The optimal profit tax rate is largely insensitive to the social welfare curvature parameter k: across k ∈ {0.77, 1, 3}, the average optimal profit tax rate is approximately 58 percent in 1980 and 61 percent in 2019. The optimal average labor income tax rate is more sensitive to k: for k = 0.7, 1, and 3, the 1980 rates are 20.3, 26.7, and 44.6 percent, and the 2019 rates are 12.5, 19.4, and 39.1 percent, respectively. Changes in the span-of-control parameter ξ and the substitution elasticity σ do not affect the labor income tax wedge schedule directly but do influence it indirectly through the markup distribution. The directional results are confirmed for all tested parameter configurations.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-role-of-the-additivity-property-from-prior-externality-literature-and-why-does-it-fail-here"&gt;Q10. What is the role of the &amp;ldquo;additivity property&amp;rdquo; from prior externality literature, and why does it fail here?&lt;/h3&gt;
&lt;p&gt;The additivity property from the Pigouvian externality literature (see Kopczuk 2003; Sandmo 1975) states that the Pigouvian correction is separable from other components of the optimal tax formula, implying that rising markups would simply decrease the optimal tax rate (since 1/μ falls). This property holds under simplifying assumptions that abstract from the general equilibrium and incentive effects of market power. In the present model, the additivity property does not hold because markups enter all four components of the optimal tax formula—not just the Pigouvian term—through the skill gap (Mirrleesian component), the RE, and the IRE. As a result, rising markups can increase the optimal profit tax rate even though the Pigouvian component falls, because the skill gap and Mirrleesian force dominate.&lt;/p&gt;
&lt;h3 id="q11-can-the-government-attain-the-first-best-by-conditioning-taxes-on-markups"&gt;Q11. Can the government attain the first-best by conditioning taxes on markups?&lt;/h3&gt;
&lt;p&gt;No. The paper demonstrates that even if the planner can observe and condition taxes on firm-level markups, the first-best is not achievable. The reason is that markups are endogenous to the entrepreneurs&amp;rsquo; unobservable decisions: an entrepreneur&amp;rsquo;s markup depends on their privately known type and chosen output. When the planner designs a mechanism that conditions on markup, the incentive constraint facing entrepreneurs remains the same as in the benchmark model, because the promise-keeping constraints are independent of the entrepreneur&amp;rsquo;s true type when markups are observable. The optimal allocation with markup-conditioned taxes is shown to be equivalent to the second-best with nonlinear sales taxes, which still falls short of the first-best.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-for-the-design-of-the-profit-tax-schedule"&gt;Q12. What are the policy implications for the design of the profit tax schedule?&lt;/h3&gt;
&lt;p&gt;The model yields three concrete prescriptions for the joint design of labor and profit income taxes in the context of rising market power. First, labor income taxes should be reduced and top profit taxes should be increased as market power rises. Second, for large, high-productivity firms the profit tax should be designed to be appropriately regressive to enhance allocative efficiency through the Reallocation Effect—this provides a new normative justification for why profit tax schedules observed in practice are often less progressive than labor income taxes. Third, while profit taxes should be regressive for large firms, the degree of regressivity should decrease as market power rises, reflecting the trade-off between efficiency and equality: higher markups increase the Mirrleesian pressure for redistribution at the top, reducing the optimal regressivity.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Mirrleesian component (of the optimal tax formula):&lt;/strong&gt; The standard incentive component of the optimal tax, capturing the trade-off between direct redistribution and the efficiency cost of taxation. In the presence of market power, this component is modified because the skill gap for entrepreneurs depends on markups through the profit elasticity: higher markups reduce profit elasticity, widen the skill gap, and amplify the Mirrleesian force toward higher top profit taxes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pigouvian component:&lt;/strong&gt; The correction in the optimal tax formula for the externality from market power. Because oligopolistic pricing causes output to be inefficiently low, the optimal tax subsidizes both worker and entrepreneurial labor supply. In the labor income tax formula, the Pigouvian component is the reciprocal of the employment-weighted average markup; in the profit tax formula, it is the reciprocal of the firm-level markup. As average markups rise, the Pigouvian component reduces the optimal labor income tax rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reallocation Effect (RE):&lt;/strong&gt; A component of the optimal profit tax formula that captures the efficiency gain from reallocating labor inputs from low-markup firms (where labor&amp;rsquo;s marginal product is high relative to value) to high-markup firms (where labor demand is inefficiently low). It equals the ratio of the average markup to the firm-level markup minus one. It implies a lower optimal marginal tax rate for firms with markups above the average, producing a regressive structure in the profit tax for large firms. This effect is absent under monopolistic competition (uniform markups) and in competitive markets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Indirect Redistribution Effect (IRE):&lt;/strong&gt; A component of the optimal profit tax formula specific to oligopolistic competition, capturing redistribution through competitor prices. Lowering the marginal tax rate of a high-productivity entrepreneur raises their output, which reduces the prices of substitutable goods produced by their competitors, thereby lowering competitors&amp;rsquo; incomes and redistributing toward workers who benefit from lower prices. This effect is present only when the cross-inverse demand elasticity is nonzero—i.e., only under oligopolistic (Cournot) competition with multiple firms per market—and vanishes under monopolistic competition and in the limit as the number of firms grows to infinity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Skill gap (for entrepreneurs):&lt;/strong&gt; The proportional rate of change in the composite entrepreneur ability measure with respect to entrepreneur type, analogous to the Mirrleesian skill gap for workers. Under market power, the entrepreneur skill gap depends on the markup through the profit elasticity: as firm-level markups rise, profit elasticity falls, the skill gap increases, and the income dispersion across entrepreneurs widens, which amplifies the Mirrleesian incentive to redistribute at the top and raises the optimal top profit tax rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Symmetric Cournot Competitive Tax Equilibrium (SCCTE):&lt;/strong&gt; The equilibrium concept used in the paper. It is a combination of a tax system, symmetric allocation, and symmetric price system such that all agents (final goods producer, entrepreneurs of each type, workers) are optimizing, strategic interaction in the intermediate goods market is a Cournot Nash equilibrium within each granular market, and all commodity and labor markets clear. Strategic interaction is restricted to within each granular market (firms in the same market compete), so decisions across markets are taken as given.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Composite ability:&lt;/strong&gt; A combined measure of entrepreneur productivity that determines equilibrium allocations and optimal taxation in the nested-CES economy. It aggregates the entrepreneur&amp;rsquo;s raw ability (affecting output capacity) and the demand parameter (affecting the market-level markup). The markup-relevant component and the quantity-relevant component are not perfect substitutes in the composite, since equilibrium prices depend on their specific composition while equilibrium quantities depend only on their combined value.&lt;/p&gt;</description></item><item><title>Peer Effects and Rank Concerns in the Classroom</title><link>https://macropaperwarehouse.com/papers/peer-effects-and-rank-concerns-in-the-classroom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/peer-effects-and-rank-concerns-in-the-classroom/</guid><description>&lt;p&gt;This paper investigates the mechanisms behind peer effects in the classroom using exogenous variation in study disruptions generated by the 2010 Maule mega-earthquake in Chile (magnitude 8.8, the seventh-largest ever instrumentally recorded). The central research question is why classroom peers can shape academic achievement — specifically, whether beyond production complementarities and a desire to conform, a desire to compete for classroom rank can drive peer influence on learning.&lt;/p&gt;
&lt;p&gt;The author constructs a novel dataset linking administrative and survey data from Chile&amp;rsquo;s Ministry of Education (SIMCE test scores, GPA, curriculum coverage, and school expenditure records) for two cohorts of roughly 150,000 eighth-grade students — one measured in 2009 before the earthquake, one measured in 2011 roughly 20–22 months after — to newly constructed measures of housing damage. Damage to each student&amp;rsquo;s home is built in three steps: (1) ground-shaking intensity using an established attenuation formula for the 2010 earthquake; (2) seismic vulnerability of each student&amp;rsquo;s home inferred from a latent-class-analysis model trained on census data linking housing construction materials to vulnerability classes; and (3) a combined expected &amp;ldquo;damage ratio&amp;rdquo; (fraction of home that needs to be rebuilt). Identification uses a difference-in-differences strategy that exploits the differential correlation between pre-existing seismic vulnerability and outcomes across the pre- and post-earthquake cohorts, controlling for socioeconomic composition.&lt;/p&gt;
&lt;p&gt;The main findings, holding fixed a student&amp;rsquo;s own earthquake exposure, are as follows. (1) Own home damage reduced test scores by 0.03 standard deviations (SD) per SD increase in damages (a 4.4 percentage-point increase in collapsed home fraction, approximately USD 3,600) and raised self-reported cost of study effort. GPA effects (–0.02 SD) are statistically insignificant. (2) A 1 SD increase in the mean damage among classroom peers raised test scores by 0.05 SD and GPA by 0.04 SD. School expenditure data (available for the 42% of schools in the preferential subsidy program) show schools responded by reallocating funds away from administrative activities toward educational and psychological support, accounting for this positive effect. (3) A 1 SD increase in the within-classroom standard deviation of peer damages lowered test scores and GPA by approximately 0.085 SD on average, but with sharply heterogeneous effects across the prior-achievement distribution: it lowered test scores and GPA of high-prior-achievement students by 0.08–0.11 SD and raised achievement of low-prior-achievement students, without corresponding changes in those students&amp;rsquo; GPA rank. Neither curriculum-coverage data nor school spending data show significant responses to damage dispersion, pointing to peer-to-peer interactions rather than school mediation.&lt;/p&gt;
&lt;p&gt;The null effect on GPA rank despite heterogeneous GPA effects is the pivotal empirical finding motivating the paper&amp;rsquo;s theory. The author argues that high-achieving students reduced effort in response to a less threatening competitive environment while maintaining their classroom standing — consistent with rank concerns driving effort decisions. Direct survey evidence shows a majority of students agreed they like to do better than classmates.&lt;/p&gt;
&lt;p&gt;Motivated by this evidence, the paper introduces a game-of-status model where each student chooses effort to maximize a utility function combining academic achievement and classroom GPA rank, with rank weighted by a preference parameter lambda &amp;gt; 0. The model admits a unique symmetric Bayesian Nash equilibrium. The model rationalizes all four main empirical patterns: positive mean-damage effects (school compensation); heterogeneous dispersion effects (rank competition changes the density of nearby competitors); null dispersion effects on GPA rank (simultaneous equilibrium adjustment preserves rank ordering); and the survey evidence on competitive preferences.&lt;/p&gt;
&lt;p&gt;The study is confined to Chilean public and subsidized private schools in earthquake-affected, non-coastal regions, with outcomes measured at the 8th grade. The pre/post cohort design removes schools that closed or received earthquake evacuees. Findings apply to a context where classroom rank is observable to peers (GPA) and where competitive preferences are prevalent among students.&lt;/p&gt;
&lt;p&gt;Q: What is the core identification strategy and why does it avoid the usual confounds in peer-effects research?
A: The paper uses a difference-in-differences estimator that exploits the differential relationship between pre-existing seismic vulnerability and outcomes across a pre-earthquake cohort (outcomes measured in 2009) and a post-earthquake cohort (outcomes measured in 2011). Because identification relies on variation in peer disruptions rather than in peer characteristics — and because students did not reallocate across classrooms or schools in response to the earthquake in the estimation sample — the strategy avoids the reflection problem and selection confounds that typically plague peer-effects identification. The identifying assumption is that the relationship between seismic vulnerability and outcomes would have been the same across cohorts absent the earthquake.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports the identifying assumption?
A: The paper provides three pieces of supporting evidence. First, the fraction of students switching schools or classrooms between grades 7 and 8 is identical across the pre- and post-earthquake cohorts in the estimation sample, indicating no earthquake-induced reallocation. Second, pre-trend tests show precise zero effects of own damage, mean peer damage, and SD of peer damage on lagged (4th-grade) test scores and GPA. Third, placebo tests using students in regions unaffected by the earthquake show no significant differential relationships between seismic vulnerability measures and outcomes across cohorts.&lt;/p&gt;
&lt;p&gt;Q: How was housing damage measured, and why does this matter for identification?
A: Damage is estimated in three steps: ground-shaking intensity at the student&amp;rsquo;s town is calculated from a validated attenuation formula; seismic vulnerability of the home is predicted using a latent-class-analysis model trained on pre-earthquake census housing data and then applied to student records; and the two are combined into a damage ratio (fraction of home to be rebuilt) using structural engineering damage-grade distributions. This constructed measure is not self-reported and is determined by physical and housing-quality factors largely predetermined before the earthquake, which supports exogeneity. Coastal towns are excluded because the accompanying tsunami caused damages not captured by the damage-ratio formula, and results are robust to different definitions of coastal proximity.&lt;/p&gt;
&lt;p&gt;Q: What were the effects of damage to a student&amp;rsquo;s own home on achievement?
A: A 1 SD increase in own home damages (corresponding to a 4.4 percentage-point increase in the collapsed fraction of the home, or roughly USD 3,600) reduced test scores by 0.03 SD. GPA fell by 0.02 SD but this was not statistically significant. Survey data show that own-home damages raised students&amp;rsquo; self-reported cost of study effort, suggesting this effort channel may mediate the achievement effects. These negative effects did not vary significantly across the baseline achievement distribution.&lt;/p&gt;
&lt;p&gt;Q: What were the effects of mean peer damage on own achievement, and what mechanism explains them?
A: A 1 SD increase in mean peer home damage raised own test scores by 0.05 SD and GPA by 0.04 SD. School spending data from SEP-program schools (42% of the sample) show that schools responded to higher average student damage by reallocating expenditures away from administrative activities (recruitment of non-teaching staff, equipment purchases) toward educational support and psychological support activities. This reallocation more than offset potential negative peer-environment effects, generating positive net achievement effects that were approximately uniform across the prior-achievement distribution.&lt;/p&gt;
&lt;p&gt;Q: What were the effects of within-classroom damage dispersion on achievement, and how do they vary across students?
A: A 1 SD increase in the within-classroom standard deviation of peer damages lowered average test scores and GPA by approximately 0.085 SD. These average effects mask sharp heterogeneity: high-prior-achievement students experienced losses of 0.08–0.11 SD in test scores and GPA, while low-prior-achievement students saw gains. For some students the dispersion effect was comparable to or larger than the effect of damage to their own home.&lt;/p&gt;
&lt;p&gt;Q: Why is the null effect of damage dispersion on GPA rank theoretically important?
A: Students with high prior achievement experienced drops in GPA in classrooms with more dispersed damages, but without an accompanying drop in their GPA rank. The paper argues this is inconsistent with students passively absorbing a changed study environment: instead, students appear to have adjusted effort precisely enough to maintain their classroom standing. This equilibrium pattern — GPA changes that leave rank ordering intact — is the paper&amp;rsquo;s key empirical signature of rank-motivated competition as a mechanism for peer influence.&lt;/p&gt;
&lt;p&gt;Q: What direct survey evidence is presented on rank concerns?
A: Survey data from the post-earthquake cohort show that a majority of students agreed with the statement &amp;ldquo;I like to do better than my classmates in school,&amp;rdquo; providing direct evidence that students value classroom rank. Additionally, students with higher initial achievement reported reductions in self-reported ability to engage with course content in classrooms with more dispersed damages, consistent with these students reducing effort when the competitive environment became less threatening to their rank.&lt;/p&gt;
&lt;p&gt;Q: Do schools mediate the damage-dispersion spillovers?
A: The available data on curriculum coverage and school spending do not show statistically significant responses to within-classroom damage dispersion (as distinct from mean damage). Emergency reconstruction funds were also allocated by schools based on overall damage severity, not its within-classroom dispersion. This absence of a detectable school-mediation channel for dispersion effects strengthens the interpretation that the heterogeneous achievement effects of dispersion reflect peer-to-peer interactions rather than differential school responses.&lt;/p&gt;
&lt;p&gt;Q: How does the game-of-status model rationalize the empirical findings?
A: In the model, each student maximizes a utility function over academic achievement and GPA rank, with rank weighted by lambda &amp;gt; 0. Students choose effort simultaneously, and their cost-of-effort type is shaped by prior test scores, socioeconomic characteristics, and earthquake damage. The model admits a unique symmetric Bayesian Nash equilibrium. In this equilibrium: schools&amp;rsquo; compensating inputs in response to mean damage raise achievement uniformly (rationalizing positive mean-damage effects); changes in damage dispersion alter the density of nearby types differently for high- and low-cost-effort students, changing the marginal benefit of exerting effort to overtake competitors (rationalizing heterogeneous GPA effects); and because all students adjust effort simultaneously, the rank ordering is approximately preserved (rationalizing null rank effects).&lt;/p&gt;
&lt;p&gt;Q: What is the mechanism by which damage dispersion produces heterogeneous effort incentives?
A: The key mechanism is that when students derive utility from rank, the marginal benefit of a unit of additional effort depends on how many competitors are &amp;ldquo;nearby&amp;rdquo; in the effort-cost distribution. When dispersion increases, the density of types just below a high-achiever (low-cost-effort student) decreases, reducing the gain from exerting more effort to maintain rank over nearby rivals; high-achievers therefore reduce effort and GPA falls. Conversely, when dispersion increases, low-achievers face a distribution where they can more effectively compete for higher ranks, raising their effort incentive and GPA.&lt;/p&gt;
&lt;p&gt;Q: How does this paper&amp;rsquo;s theory differ from prior theories of peer influence?
A: Prior theories have emphasized two mechanisms: production complementarities (peer ability directly improves own learning) and a desire to conform (students prefer to match their peers&amp;rsquo; effort or achievement). Both rationalize a linear-in-means model that captures only mean peer characteristics. This paper&amp;rsquo;s theory is the first in the peer-effects literature to rationalize why higher-order moments of the peer distribution (specifically dispersion) affect learning, through a competitive rank-concern mechanism that is parsimonious and does not require extensions to production technology or preferences beyond adding rank to the utility function.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the competitive-motive theory?
A: The theory implies that classroom composition policies affecting the dispersion of student ability — such as ability tracking, gifted programs, or reshuffling policies — can have heterogeneous and potentially perverse effects: policies that reduce ability dispersion may concentrate competitive incentives in ways that harm some students while benefiting others. Standard linear-in-means models of peer effects, which capture only mean peer characteristics, would not predict these distributional consequences. The author argues this means the competitive mechanism has been largely unexplored despite its intuitive appeal, and calls for structural estimation and policy analysis in future work.&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the empirical findings?
A: The findings apply to 8th-grade students in Chilean public and private subsidized schools located in earthquake-affected, non-coastal regions, with outcomes observed approximately 20–22 months post-earthquake. The sample excludes schools that closed due to the earthquake and schools that received evacuees. The paper notes that while the theory is formulated around an earthquake shock, the competitive-motive mechanism applies whenever the dispersion of students&amp;rsquo; cost-of-effort types changes — including through classroom assignment policies or other shocks — and is not specific to the natural-disaster context.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Damage ratio&lt;/strong&gt;: The fraction of a student&amp;rsquo;s home that needs to be rebuilt, constructed by combining geocoded ground-shaking intensity (via the Astroza et al. attenuation formula for the 2010 Chilean earthquake) with the predicted seismic vulnerability class of the home (derived from a latent-class-analysis model trained on census housing data). Used as the paper&amp;rsquo;s measure of disruption to each student&amp;rsquo;s environment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exogenous peer effect&lt;/strong&gt; (in the sense of Manski 1993): The reduced-form impact on a student&amp;rsquo;s outcome of a change in the distribution of an exogenous characteristic — here, earthquake damage — among classroom peers, holding fixed the student&amp;rsquo;s own characteristics. Distinguished in the paper from endogenous peer effects (best-response functions).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rank concern&lt;/strong&gt;: Students&amp;rsquo; utility derived from their position (rank) in the classroom GPA distribution, irrespective of whether that rank is formally rewarded. The paper treats rank concern as a preference parameter (lambda &amp;gt; 0 in the utility function) and identifies it as a mechanism for peer influence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Game-of-status model&lt;/strong&gt;: The paper&amp;rsquo;s theoretical framework, in which students simultaneously choose study effort to maximize utility over own academic achievement and GPA rank. The model admits a unique symmetric Bayesian Nash equilibrium. The central insight is that the density of nearby competitors in the effort-cost distribution determines the marginal benefit of effort, generating heterogeneous incentives when peer cost-of-effort types become more dispersed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Effort-cost type&lt;/strong&gt;: Each student&amp;rsquo;s marginal cost of exerting study effort, shaped by prior test scores, socioeconomic characteristics, and earthquake damages to the student&amp;rsquo;s own home. The key primitive of the model that links individual disruptions to equilibrium effort choices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;SEP (Subvencion Escolar Preferencial)&lt;/strong&gt;: Chile&amp;rsquo;s preferential school subsidy program for disadvantaged students, which requires participating schools (42% of the sample) to submit detailed annual spending reports to the Ministry of Education. The paper uses these reports to identify school spending responses to mean and dispersed peer damages.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Seismic vulnerability class&lt;/strong&gt;: A classification of a home&amp;rsquo;s resistance to earthquake damage based on its construction materials (exterior walls, roof, floor), assigned using a logistic latent-class-analysis model estimated on census data. Found to align strongly with household socioeconomic status, enabling prediction of housing vulnerability from administrative student records.&lt;/p&gt;</description></item><item><title>Peer Effects and the Gender Gap in Corporate Leadership</title><link>https://macropaperwarehouse.com/papers/peer-effects-and-the-gender-gap-in-corporate-leadership/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/peer-effects-and-the-gender-gap-in-corporate-leadership/</guid><description>&lt;p&gt;This paper investigates whether exposure to a larger share of female peers during an MBA program causally affects the gender gap in senior corporate leadership positions. The research question is motivated by the persistent underrepresentation of women in top management: in S&amp;amp;P 1500 companies, women hold only 6% of CEO positions despite comprising 40% of the workforce.&lt;/p&gt;
&lt;p&gt;The authors merge administrative data from a top-10 U.S. business school (graduating classes 2000–2018, excluding 2009) with public LinkedIn profile data covering full employment histories, firm-level data from multiple sources including InHerSight crowdsourced female-employee ratings, and a 2023–2024 alumni survey of female graduates. Senior management is defined as Vice President, Director, Senior Vice President, or C-level executive, identified from exact job titles in LinkedIn CVs.&lt;/p&gt;
&lt;p&gt;Identification exploits the quasi-random assignment of incoming MBA students to one of eight sections of approximately 60 students each, based on alphabetical order with balance checks on gender, undergraduate institution, and ethnicity. This assignment generates exogenous variation in the share of female section peers (mean 34%, standard deviation 4 percentage points). Randomization tests following Guryan et al. (2009) and Caeyers and Fafchamps (2021) confirm the assignment is as good as random. The estimating equation is a linear-in-means model with class, year, and class-by-year fixed effects interacted with gender, plus individual and section-level controls.&lt;/p&gt;
&lt;p&gt;The paper first documents a baseline gender gap: despite 96% of both male and female MBA graduates entering management within 15 years, women are 24% less likely than men to hold senior management positions. This gap emerges immediately after graduation, persists for at least 15 years, and is partly attributable to lower promotion rates from first-level management (43% of women in first-level management transition to senior management within five years, versus 57% of men).&lt;/p&gt;
&lt;p&gt;The main causal finding is that a 4 percentage point (1 SD) increase in the share of female MBA section peers increases the probability of a woman holding a senior management position by 8.4% (a 3.3 percentage point increase off a 39.1% baseline), equivalent to a 26% reduction in the management gender gap. There is no corresponding effect for men. The effect emerges as early as two years post-graduation, peaks around year seven, and persists through the 15-year horizon.&lt;/p&gt;
&lt;p&gt;The increase is concentrated in female-friendly firms, defined as those with above-median ratings on InHerSight metrics including maternity leave generosity, flexible work schedules, and professional support. Women with more female peers are significantly more likely to transition into female-friendly firms 6 to 10 years after graduation — a period coinciding with prime childbearing years — where they subsequently attain senior management roles. The effect on senior management in female-friendly firms is statistically distinguishable from the null effect in non-female-friendly firms (p-value = 0.03). The results are largest in male-dominated industries (consulting, tech, finance) where women face greater barriers to informal networks.&lt;/p&gt;
&lt;p&gt;A survey of 283 female MBA alumnae (10% response rate) reveals three mechanisms: (i) information sharing, especially gender-specific advice about employer policies and culture; (ii) higher ambitions and self-confidence through role modeling and emotional support; and (iii) increased perceived support from male MBA peers as female section representation rises. Corroborating the information-sharing channel, women with more female peers are more likely to work at the same firms as their female section peers, particularly when those firms are female-friendly.&lt;/p&gt;
&lt;p&gt;A counterfactual exercise shows that reallocating the existing stock of female students so that all sections have at least 34% women would yield 2 to 5 additional female senior managers per graduating class (a 2.4% to 8.4% increase), holding the total number of female students fixed.&lt;/p&gt;
&lt;p&gt;Q: What is the baseline gender gap in senior management among MBA graduates, and how does it evolve over time?
A: Female MBA graduates are 24% less likely than male graduates to hold senior management positions in the 15 years after graduation. The gap emerges immediately after the MBA and persists for at least 15 years without closing. At year 15, 74% of men hold a senior management position compared to 59% of women.&lt;/p&gt;
&lt;p&gt;Q: How is female peer share defined and what is its distribution across sections?
A: Female peer share is the proportion of female students in an individual&amp;rsquo;s assigned MBA section of approximately 60 students, excluding the individual themselves. The average section female share is 34% with a standard deviation of 4 percentage points. The distribution ranges from 19% at the 1st percentile to 45% at the 99th percentile, with the interquartile range spanning approximately 32% to 36%.&lt;/p&gt;
&lt;p&gt;Q: What is the main causal estimate of female peers on women&amp;rsquo;s senior management probability?
A: A 4 percentage point (1 SD) increase in female section peer share increases the probability of a woman holding a senior management position by 8.4% (3.3 percentage points off a 39.1% mean), averaged across the 15 post-MBA years. This translates to a 26% reduction in the management gender gap. There is no statistically significant effect on men.&lt;/p&gt;
&lt;p&gt;Q: When does the effect of female peers emerge and how does it evolve dynamically?
A: The effect on women emerges as early as two years after MBA graduation and grows over time, peaking around seven years post-graduation. The effect is persistent across the 15-year horizon studied. Estimates become less precise toward the end of the sample period as recent cohorts contribute fewer observations.&lt;/p&gt;
&lt;p&gt;Q: How do female-friendly firms mediate the main result?
A: The main effect is entirely concentrated in female-friendly firms (those with above-median InHerSight ratings). The coefficient on female peer share is positive and significant for senior management in female-friendly firms, and statistically indistinguishable from zero in non-female-friendly firms. The difference between the two coefficients is significant at p = 0.03.&lt;/p&gt;
&lt;p&gt;Q: What is the mechanism linking female peers to female-friendly firm transitions?
A: Women with more female peers are significantly more likely to be employed at female-friendly firms 6 to 10 years after graduation, a window corresponding to prime childbearing years. This suggests female peers facilitate sorting into supportive firm environments when family-work tradeoffs become most acute. Once at female-friendly firms, women attain senior management positions at higher rates.&lt;/p&gt;
&lt;p&gt;Q: Does the increase in female senior managers reflect easier paths (smaller firms, lower pay, non-P&amp;amp;L roles)?
A: No. The effect is significant for both small (under 500 employees) and large (over 5,000 employees) firms, with no significant effect on the firm size of employment itself. There is no consistent pattern of women being promoted in firms with higher or lower average compensation. The increase in female senior managers includes those with Profit and Loss responsibilities, indicating these are substantive management positions.&lt;/p&gt;
&lt;p&gt;Q: In which industries is the effect largest, and what does this imply?
A: The effect is concentrated in male-dominated industries (consulting, tech, finance), with no significant effect in female-dominated industries (consumer goods, healthcare). The difference between coefficients is significant at the 3% level. Entry rates into male-dominated industries are not significantly affected, suggesting the mechanism is higher promotion rates within these industries rather than differential sorting into them. The authors interpret this as evidence that female MBA networks are most valuable where women face greater barriers to informal workplace networks.&lt;/p&gt;
&lt;p&gt;Q: What does the survey evidence reveal about mechanisms?
A: Among 283 survey respondents (10% response rate), three mechanisms emerge: information sharing about gender-specific employer attributes and policies; raising ambitions and self-confidence through role modeling; and increased perceived support from male MBA peers as section female share rises. Women with more female peers are also more likely to work at the same firms as their female section peers, especially female-friendly ones, consistent with referral and information-sharing channels.&lt;/p&gt;
&lt;p&gt;Q: Does the effect operate through greater attachment to the corporate pipeline (fewer career breaks, higher entry into management)?
A: No. Female peers do not significantly affect employment rates, career break incidence, entry into first-level management positions, or self-employment rates. The results thus reflect higher promotion rates from first-level management into senior management, not changes in pipeline attachment.&lt;/p&gt;
&lt;p&gt;Q: What do the randomization tests show about identification validity?
A: Two randomization tests confirm as-good-as-random assignment. Following Guryan et al. (2009), the section-level leave-out mean female share is not significantly different from zero after controlling for the class-level leave-out mean. Following Caeyers and Fafchamps (2021), after netting out the asymptotic exclusion bias, the female share coefficient is insignificant across all specifications. A simulation test (Bietenbeck 2020) finds no statistically significant difference between the actual and simulated within-class female share distributions.&lt;/p&gt;
&lt;p&gt;Q: What placebo tests are conducted and what do they show?
A: Two placebo tests are run. First, 1,000 random reassignments of students to sections within the same class show the true estimated effect for women lies outside the distribution of placebo effects, while the null effect for men lies within it. Second, estimating the main equation for up to three years before MBA enrollment finds no consistent pre-treatment effect of female share on future female graduates, supporting the identification strategy.&lt;/p&gt;
&lt;p&gt;Q: What is the counterfactual policy exercise and what does it imply?
A: Holding the total number of female students fixed, reallocating them so that all sections contain at least 34% women would yield 2 to 5 additional female senior managers per graduating class (a 2.4% to 8.4% increase). This assumes nonlinearity in the relationship and suggests meaningful gains from rebalancing section composition without increasing overall female enrollment.&lt;/p&gt;
&lt;p&gt;Q: How do the results compare to the Thomas (2021) finding that more male peers raise female MBA earnings?
A: The authors note several differences: Thomas (2021) focuses on starting earnings while this paper studies senior management positions over 15 years; the two studies use different universities and time periods; and this paper employs gender-by-cohort fixed effects to account for time trends in female labor market outcomes. The authors suggest these design and outcome differences explain the divergent findings.&lt;/p&gt;
&lt;p&gt;Section peers: Students assigned to the same MBA section of approximately 60 students who take core classes together and form the primary peer network; sections are assigned quasi-randomly based on alphabetical order with balance adjustments, generating exogenous variation in gender composition.&lt;/p&gt;
&lt;p&gt;Female-friendly firms: Firms with above-median ratings on InHerSight, a crowdsourced platform where female employees rate employers on metrics including maternity leave generosity, flexible work schedules, mentorship programs, and female representation in management; defined in this paper&amp;rsquo;s own terms as firms whose cultures and policies help women balance work-family responsibilities and support career advancement.&lt;/p&gt;
&lt;p&gt;Senior management: Positions defined as Vice President (VP), Director, Senior Vice President (SVP), or C-level executive, identified using keyword matching on exact job titles from LinkedIn CVs; distinguished from first-level management (managers and supervisors) and representing the upper rungs of the corporate management ladder.&lt;/p&gt;
&lt;p&gt;Female share (treatment variable): The proportion of female students among an individual&amp;rsquo;s section peers, excluding the individual themselves (leave-out mean); averaged 34% with a 4 percentage point standard deviation across sections, after residualizing by graduating class.&lt;/p&gt;
&lt;p&gt;Management gender gap: The 24 percentage point (24%) difference in the likelihood of female versus male MBA graduates holding senior management positions within 15 years of graduation; emerges immediately post-MBA and does not close over the observed horizon.&lt;/p&gt;
&lt;p&gt;Information sharing mechanism: The channel through which female MBA peers provide gender-specific advice and information about employer policies, culture, and female-friendliness that is otherwise difficult to observe; evidenced by the co-location of women with more female peers at the same female-friendly firms as their section peers.&lt;/p&gt;
&lt;p&gt;Exclusion bias: The systematic negative correlation between an individual&amp;rsquo;s own characteristic and her leave-out peer mean that arises mechanically when individuals cannot be their own peer under assignment without replacement; addressed via the Caeyers and Fafchamps (2021) correction in randomization tests.&lt;/p&gt;</description></item><item><title>Place-Based Redistribution</title><link>https://macropaperwarehouse.com/papers/place-based-redistribution/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/place-based-redistribution/</guid><description>&lt;h2 id="place-based-redistribution-overview"&gt;Place-Based Redistribution: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Should national governments redistribute income to residents of poor areas through place-based transfers, or should redistribution rely solely on place-blind (income-only) taxes? The longstanding view in urban economics—&amp;ldquo;help poor people, not poor places&amp;rdquo;—holds that place-based aid is inefficient because it channels activity to less productive locations. This paper challenges that view by formalizing the conditions under which place-based redistribution improves on purely income-based transfers, using tools from optimal tax theory embedded in a spatial equilibrium model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper develops a two-location model (&amp;ldquo;Distressed&amp;rdquo; and &amp;ldquo;Elsewhere&amp;rdquo;) with a unit mass of heterogeneous households who differ in skill level (θ) and idiosyncratic preference for living in Distressed (φ). Households choose where to live and how much to earn, facing competitive labor and housing markets in each location. Locations may differ in amenity levels, wage schedules (which may embody skill-specific comparative advantage), and housing costs. A utilitarian planner sets location-specific income tax schedules—observed earnings and location are the only signals of unobserved skill—maximizing a weighted average of household utilities and landlord profits subject to a budget constraint.&lt;/p&gt;
&lt;p&gt;The paper proceeds in three steps. First, it derives closed-form conditions for the optimality of a lump-sum place-based transfer under a fixed income tax. Second, it characterizes fully general optimal nonlinear, location-specific marginal tax rate (MTR) schedules (Proposition 2). Third, it calibrates the model numerically, anchoring to the U.S. Empowerment Zone (EZ) program.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three Sorting Mechanisms and Their Policy Implications&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper identifies three polar mechanisms that generate sorting of lower-skill households into Distressed:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;em&gt;Skill-taste correlation&lt;/em&gt;: higher-skill households have stronger tastes for Elsewhere, independent of wages or rents.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Comparative advantage&lt;/em&gt;: higher-skill workers are relatively more productive in Elsewhere.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Income-based sorting&lt;/em&gt;: because Elsewhere is more expensive, lower-income households are priced into Distressed.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Under skill-taste correlation, place-based transfers to Distressed are unambiguously welfare-improving even when income taxes are already optimal, because high-skill households prefer Elsewhere for reasons that are orthogonal to income. Under comparative advantage, the direction of the optimal transfer depends on migration elasticities: low migration elasticities favor transfers to Distressed, while high migration elasticities can reverse the sign. Under pure income-based sorting (with homogeneous locational preferences), the conditions for superfluous commodity taxation (Atkinson-Stiglitz 1976) are satisfied, and optimal place-based transfers are zero—though idiosyncratic preference heterogeneity restores non-zero optimal transfers even in this case.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Numerical simulations use Census data and ACS moments calibrated to EZ areas. With high migration responsiveness (κ = 0.5, approximating urban EZs) and skill-taste correlation as the sole sorting driver, the optimal average place-based transfer to Distressed is &lt;strong&gt;$4,805&lt;/strong&gt;, with about 40% ($1,943) arising from lower MTRs rather than a higher demogrant. With low migration responsiveness (κ = 4, approximating rural EZs), the optimal transfer more than doubles to &lt;strong&gt;$10,918&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;When comparative advantage alone drives sorting and migration is low (κ = 4), the optimal transfer to Distressed is &lt;strong&gt;$7,091&lt;/strong&gt;, with a $3,740 larger demogrant. With high migration and comparative advantage, the transfer reverses to &lt;strong&gt;−$2,763&lt;/strong&gt; (i.e., Elsewhere receives the subsidy). For intermediate migration under comparative advantage (e.g., κ ≈ 1), the optimal policy is nonlinear: the poorest Distressed residents receive a place-based transfer of &lt;strong&gt;$1,254&lt;/strong&gt;, while high-skill Distressed residents face a place-based tax of &lt;strong&gt;$12,398&lt;/strong&gt; at the 99th percentile.&lt;/p&gt;
&lt;p&gt;In the empirically calibrated &lt;strong&gt;urban EZ baseline&lt;/strong&gt; (migration elasticity 0.82, rent ratio 0.86, sorting driven by skill-taste correlation and income effects), the optimal average place-based transfer is &lt;strong&gt;$3,143&lt;/strong&gt;, roughly matching the magnitude of actual EZ wage tax credits (~$3,000 for full-time eligible workers). The demogrant advantage for Distressed is &lt;strong&gt;$1,462&lt;/strong&gt;, with just over half of the transfer arising from lower MTRs.&lt;/p&gt;
&lt;p&gt;In the &lt;strong&gt;rural EZ baseline&lt;/strong&gt; (migration elasticity 0.20, rent ratio 0.54, comparative advantage and income effects), the optimal average transfer rises to &lt;strong&gt;$4,329&lt;/strong&gt;, concentrated in lower MTRs rather than a larger demogrant. Halving the migration elasticity from the rural baseline raises the optimal transfer to &lt;strong&gt;$6,906&lt;/strong&gt;, while doubling it reduces the transfer to near zero (&lt;strong&gt;$573&lt;/strong&gt;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;All results are derived under the assumption of &lt;em&gt;no market failures&lt;/em&gt;; the model deliberately excludes agglomeration spillovers or other Pigouvian motives, attributing the case for place-based redistribution purely to redistributive goals.&lt;/li&gt;
&lt;li&gt;The planner observes only earnings and location, not skill type directly.&lt;/li&gt;
&lt;li&gt;Household Pareto weights are set equal to one across types in the simulations, so redistribution is driven solely by diminishing marginal utility of consumption.&lt;/li&gt;
&lt;li&gt;The model abstracts from interactions with subnational governments, local public services, and endogenous amenities.&lt;/li&gt;
&lt;li&gt;Results on the desirability of transfers to Distressed hinge critically on the motive for sorting, not simply on the existence of spatial income inequality.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-equity-efficiency-tradeoff-formula-for-a-lump-sum-place-based-transfer-and-what-does-it-reveal"&gt;Q1. What is the equity-efficiency tradeoff formula for a lump-sum place-based transfer, and what does it reveal?&lt;/h3&gt;
&lt;p&gt;Lemma 1 shows that the first-order welfare effect of a small per-capita transfer from Elsewhere to Distressed starting from a place-blind tax system is dSWF/dt = (λ̄₁ − λ̄₀) + Eθ{m(0)·[T(z₁*) − T(z₀*)]}. The equity gain (λ̄₁ − λ̄₀) is positive when Distressed households have higher average social marginal welfare weights, which holds when their skill distribution is first-order stochastically dominated by Elsewhere&amp;rsquo;s. The fiscal cost equals the earnings-tax-revenue loss from movers: households induced to migrate to Distressed who earn less there generate lower tax payments. This formula identifies the earnings response to migration as a sufficient statistic for the efficiency cost of place-based policy.&lt;/p&gt;
&lt;h3 id="q2-what-characterizes-the-optimal-lump-sum-transfer-t-in-proposition-1"&gt;Q2. What characterizes the optimal lump-sum transfer t* in Proposition 1?&lt;/h3&gt;
&lt;p&gt;Proposition 1 shows t* = [λ̄₁(t*) − λ̄₀(t*) + Eθ{m(t*)·[T(z₁*) − T(z₀*)]}] / (Eθ[m(t*)] / [L₀(t*)L₁(t*)]). The optimal transfer is larger when (i) the average social marginal welfare weight gap between Distressed and Elsewhere is greater, (ii) migration responses m(t*) are small, and (iii) the earnings difference between locations for marginal movers is small. This formula holds regardless of whether the income tax schedule T(·) is itself set optimally.&lt;/p&gt;
&lt;h3 id="q3-under-skill-taste-correlation-why-are-place-based-transfers-always-welfare-improving-even-under-an-optimal-income-tax"&gt;Q3. Under skill-taste correlation, why are place-based transfers always welfare-improving even under an optimal income tax?&lt;/h3&gt;
&lt;p&gt;When sorting is driven by skill-taste correlation (high-skill households have stronger preferences for Elsewhere despite identical wages and rents), the equity gain λ̄₁ − λ̄₀ is positive because low-skill households concentrate in Distressed. A small positive transfer starting from t = 0 also incurs zero fiscal cost because movers between locations face identical wages and do not change their earnings. Thus, welfare unambiguously increases. The key insight is that skill-taste correlation violates the Atkinson-Stiglitz condition: high earners would still prefer Elsewhere even if forced to earn less, so location serves as a proxy for skill not captured by income taxes alone.&lt;/p&gt;
&lt;h3 id="q4-under-comparative-advantage-why-can-the-sign-of-the-optimal-transfer-reverse-with-migration-elasticity"&gt;Q4. Under comparative advantage, why can the sign of the optimal transfer reverse with migration elasticity?&lt;/h3&gt;
&lt;p&gt;When higher-skill workers are more productive in Elsewhere, movers to Distressed experience wage and earnings reductions, generating a fiscal externality. When migration elasticities are high (low κ), this fiscal cost is large and can dominate the equity gain, making transfers to Elsewhere optimal (simulated optimal transfer of −$2,763 at κ = 0.5). When migration elasticities are low (high κ), the fiscal cost is small and equity considerations dominate, yielding transfers to Distressed ($7,091 at κ = 4). At intermediate elasticities, the optimal policy is nonlinear, redistributing to poor Distressed residents while taxing rich Distressed residents more.&lt;/p&gt;
&lt;h3 id="q5-why-are-place-based-transfers-superfluous-under-pure-income-based-sorting-with-homogeneous-locational-preferences"&gt;Q5. Why are place-based transfers superfluous under pure income-based sorting with homogeneous locational preferences?&lt;/h3&gt;
&lt;p&gt;Example 6 (and its formal proof in Appendix B.3.5) demonstrates that when sorting arises solely from higher rents in Elsewhere and preferences over location are homogeneous (no idiosyncratic φ heterogeneity), the Atkinson-Stiglitz sufficient condition for commodity tax superfluousness is met: hypothetically forcing high earners to earn less would not change their preferred consumption bundle relative to low earners. Hence a place-blind income tax implements optimal redistribution without spatial supplements. As the variance of idiosyncratic location preferences κ shrinks toward zero, Figure 3 confirms that optimal place-based transfers tend toward zero across all three sorting motives.&lt;/p&gt;
&lt;h3 id="q6-what-new-terms-appear-in-the-optimal-location-specific-mtr-formulas-proposition-2-relative-to-a-standalone-economy-optimum"&gt;Q6. What new terms appear in the optimal location-specific MTR formulas (Proposition 2) relative to a standalone-economy optimum?&lt;/h3&gt;
&lt;p&gt;The optimal MTR schedules in Proposition 2 contain two new terms beyond the standard Mirrlees (1971)/Saez (2001) formula. The term Δτ+(θ) captures the fiscal externality from migration: raising Elsewhere&amp;rsquo;s MTR at skill level θ and above induces movers to Distressed who change their tax revenue by T₁(z₁*(s)) − T₀(z₀*(s)). The term (λ_L − 1)Δr+(θ) captures the equilibrium rent effect: MTR changes shift households between locations, altering rents in both communities and redistributing between renters and landlords. When λ_L &amp;lt; 1 (landlords are weighted less than average households), the rent term creates additional motives for spatial redistribution depending on the ratio of rents to housing supply elasticities across locations.&lt;/p&gt;
&lt;h3 id="q7-how-do-housing-supply-elasticities-affect-the-optimal-spatial-transfer-and-why-does-the-sign-differ-between-urban-and-rural-settings"&gt;Q7. How do housing supply elasticities affect the optimal spatial transfer, and why does the sign differ between urban and rural settings?&lt;/h3&gt;
&lt;p&gt;The rent redistribution term Δr+(θ) has sign determined by r₁/ϱ₁ − r₀/ϱ₀. For urban EZs, where Distressed has lower rents but also lower housing supply elasticity than Elsewhere (ϱ₁ = 0.24, ϱ₀ = 0.34 in the baseline), this ratio is positive, meaning transfers to Distressed shift households into relatively inelastic markets, raising rents there and generating landlord income. When λ_L &amp;lt; 1, this reduces the desirability of transfers to Distressed. For rural EZs, Distressed has higher housing supply elasticity (ϱ₁ = 0.60), so the ratio is negative: transfers shift households to more elastic markets where rents rise minimally. When λ_L &amp;lt; 1, this actually motivates more transfers to rural Distressed areas. In the 75%-landlord-weight sensitivity, optimal urban transfers fall by ~$1,000 while rural transfers rise by ~$1,000, illustrating this asymmetry.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-urban-ez-baseline-calibration-find-about-optimal-transfers-and-how-does-it-compare-to-actual-ez-policy"&gt;Q8. What does the urban EZ baseline calibration find about optimal transfers and how does it compare to actual EZ policy?&lt;/h3&gt;
&lt;p&gt;The urban baseline targets a migration elasticity of 0.82 (from Busso et al. 2013), a Distressed-to-Elsewhere rent ratio of 0.86, and 56% of Distressed residents earning under $50,000. The calibrated κ is 0.44. At the optimum, Distressed residents receive an average place-based transfer of $3,143, with $1,462 as a higher demogrant and the remainder from lower MTRs. By comparison, actual EZs provide a wage tax credit of approximately $3,000 per eligible full-time worker. The paper concludes that the magnitude—but not the capped, flat structure—of EZ transfers approximates the optimal level.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-rural-ez-calibration-find-and-how-sensitive-are-results-to-migration-assumptions"&gt;Q9. What does the rural EZ calibration find, and how sensitive are results to migration assumptions?&lt;/h3&gt;
&lt;p&gt;The rural baseline targets a migration elasticity of 0.20 (from Sprung-Keyser et al. 2022), a rent ratio of 0.54, and 60% of Distressed residents earning under $50,000, with sorting attributed to comparative advantage and income effects. The calibrated κ is 4.06. The optimal average transfer is $4,329, primarily arising from lower MTRs rather than a higher demogrant ($532). Doubling the migration elasticity reduces the optimal transfer to near zero ($573); halving it raises it to $6,906. The direction and magnitude of optimal transfers are therefore highly sensitive to the assumed level of migration responsiveness, highlighting the empirical importance of estimating migration elasticities—particularly heterogeneity in migration by income level and earnings changes for marginal movers.&lt;/p&gt;
&lt;h3 id="q10-do-within-income-transfers-arising-from-differences-in-marital-and-parental-status-across-communities-effectively-constitute-place-based-redistribution"&gt;Q10. Do within-income transfers arising from differences in marital and parental status across communities effectively constitute place-based redistribution?&lt;/h3&gt;
&lt;p&gt;Online Appendix A investigates this by estimating the implicit place-based transfer induced by marital and parental status differences between EZ communities and the rest of the country. Using ACS tract-level data merged with Piketty-Saez-Zucman distributional national accounts (DINA), the authors find that marital status and parental status have offsetting effects: marital status raises taxes on single households (common in Distressed), while parental status increases transfers to households with children (also common in Distressed). Across all preferred CPS-adjusted estimates, net within-earnings transfers are below $1,000 in magnitude, and the two factors essentially cancel. The authors conclude that marital and parental status differences do not yield substantial de facto place-based redistribution within income levels.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-mtr-decomposition-table-3-reveal-about-why-sorting-motives-generate-different-mtr-patterns"&gt;Q11. What does the MTR decomposition (Table 3) reveal about why sorting motives generate different MTR patterns?&lt;/h3&gt;
&lt;p&gt;The decomposition separates the optimal MTR into a within-community component (standard equity-efficiency tradeoff) and a between-community component (fiscal externality from migration). Under skill-taste correlation with high migration (κ = 0.5), both components contribute positively to the Distressed MTR (0.246 within + 0.234 between = 0.479), yielding lower MTRs in Distressed (0.479) than in Elsewhere (0.510). Under comparative advantage with high migration, the within-community component is negative (−0.111) because high MTRs at the optimum reduce the concentration of high-skill types in Distressed, depressing the standard revenue-raising benefit of MTRs. The large positive between-community component (0.655) reflects the large fiscal externality from movers and overcomes this, yielding higher Distressed MTRs (0.544 vs. 0.509 in Elsewhere). With low migration (κ = 4), between-community components shrink substantially, and MTRs in Distressed fall below Elsewhere in all sorting scenarios.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-crosswalk-from-urban-to-rural-baseline-reveal-about-which-assumptions-drive-the-change-in-optimal-transfers"&gt;Q12. What does the crosswalk from urban to rural baseline reveal about which assumptions drive the change in optimal transfers?&lt;/h3&gt;
&lt;p&gt;Table 5 traces the urban-to-rural transition step by step. Starting from the urban baseline ($3,143 average transfer), replacing the migration elasticity target with the rural value of 0.20 triples the optimal transfer to $9,870. Subsequently replacing skill-taste correlation with comparative advantage as the sorting mechanism reduces the transfer by roughly half ($6,402). Adjusting rent to match the rural ratio (0.54) reduces it further to $2,780, as lower Distressed rent reduces the marginal utility of consumption at the bottom and increases income-based sorting. Targeting the rural income share (60% below $50K) raises it back to $4,140, and incorporating rural housing supply elasticities yields the rural baseline result of $4,329. This decomposition reveals that lower migration responsiveness is the single largest driver of higher optimal transfers in rural settings.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Place-based redistribution&lt;/strong&gt;: Transfer schemes in which economic benefits or tax burdens are conditioned on the geographic location of residence, as distinct from place-blind income taxes that condition only on earned income. In this paper, modeled as location-specific tax schedules T_j(z) that may differ across communities j.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Skill-taste correlation&lt;/strong&gt;: A source of spatial sorting in which households with higher skill levels (θ) have systematically stronger preferences for the &amp;ldquo;Elsewhere&amp;rdquo; location, independently of wage or rent differences. Formally, the conditional distribution G_θ(φ) of locational tastes given skill is weakly increasing in θ. This correlation breaks the Atkinson-Stiglitz sufficient condition for commodity tax superfluousness and generates unambiguously positive optimal transfers to Distressed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Comparative advantage (spatial)&lt;/strong&gt;: A sorting mechanism in which higher-skill workers are disproportionately more productive in Elsewhere than in Distressed, captured by the wage elasticity with respect to skill being higher in Elsewhere (γ₀(θ) &amp;gt; γ₁(θ)). Households with skill above a threshold sort into Elsewhere even with homogeneous locational preferences. The existence of spatial comparative advantage means that migrants to Distressed earn less, creating a fiscal externality for place-based transfers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Income-based sorting&lt;/strong&gt;: Sorting of lower-income, lower-skill households into Distressed arising purely from the higher cost of living in Elsewhere, without any systematic skill-taste correlation or comparative advantage. Because high-skill households are less sensitive to rent differences, they sort into Elsewhere when rents there are higher. When this is the sole sorting mechanism and locational preferences are homogeneous, the Atkinson-Stiglitz commodity tax superfluousness conditions are satisfied and optimal place-based transfers are zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fiscal externality (migration)&lt;/strong&gt;: The change in income tax revenue caused by migration responses to place-based policy changes, not by changes in incentives for stayers. When movers from Elsewhere to Distressed earn less in their new location, they generate lower tax payments, imposing a first-order cost on the government budget. This externality is measured by Δτ+(θ) in the optimal MTR formulas and equals the earnings-tax-revenue loss from movers across all skill levels above θ. This term is a &amp;ldquo;sufficient statistic&amp;rdquo; for the efficiency cost of place-based transfers in the sense of Chetty (2009).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Demogrant (∆₀)&lt;/strong&gt;: The difference in lump-sum transfers provided to zero-earners across the two locations (−T₀(0) − (−T₁(0)) = T₀(0) − T₁(0)). A positive ∆₀ means Distressed provides a larger transfer to non-earners. It represents the place-based redistribution that occurs at the bottom of the earnings distribution, independently of MTR differences. In the paper&amp;rsquo;s decomposition, total optimal place-based redistribution (∆_z) exceeds ∆₀ when Distressed also has lower MTRs, meaning redistribution grows with income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Income-constant average tax difference (∆_z)&lt;/strong&gt;: The paper&amp;rsquo;s preferred summary measure of the average place-based transfer, defined as an equally weighted average of two tax-difference indices: the tax difference evaluated at Elsewhere earnings levels and the tax difference evaluated at Distressed earnings levels. This measure isolates tax schedule differences from productivity differences across locations, avoiding conflation of tax policy and wage effects on measured income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Landlord welfare weight (λ_L)&lt;/strong&gt;: The social marginal welfare weight assigned to landlords relative to the multiplier on the government budget constraint. When λ_L &amp;lt; 1, the planner values a marginal dollar of public funds more than a marginal dollar to landlords, creating a motive to use place-based taxes to shift rent incidence. The rent redistribution effect on optimal MTRs operates through the term (λ_L − 1)Δr+(θ), which has opposite signs in urban (positive) and rural (negative) distressed areas because of their different housing supply elasticities.&lt;/p&gt;</description></item><item><title>Policy Diffusion and Polarization across U.S. States</title><link>https://macropaperwarehouse.com/papers/policy-diffusion-and-polarization-across-u.s.-states/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/policy-diffusion-and-polarization-across-u.s.-states/</guid><description>&lt;p&gt;DellaVigna and Kim study the innovation and diffusion of policies across U.S. states using a dataset of over 700 state laws spanning seven decades. The central question is what predicts whether a state adopts a policy — and how those predictors have changed over time. The paper draws on two primary data sources: the State Policy Innovation and Diffusion (SPID) Database (Boehmke et al., 2020), covering 676 policies, and a hand-collected sample of 57 policies from 91 NBER working papers (April 2012–September 2021) that feature state-level policy variation. The combined dataset covers 733 policies adopted from the 1950s onward across the contiguous 48 states.&lt;/p&gt;
&lt;p&gt;On policy innovation, the paper finds that state capacity plays only a small role: larger and richer states are only slightly more likely to introduce new policies, innovation originates from both Republican and Democratic states, and the patterns are largely idiosyncratic with respect to observable state characteristics. California is the most frequent innovator, but large states like Florida and Texas rank in the middle.&lt;/p&gt;
&lt;p&gt;For policy diffusion, the paper employs both a static Geary&amp;rsquo;s C clustering statistic (measuring whether the first 10 adopting states cluster geographically or politically relative to a random-diffusion benchmark) and a dynamic logit hazard model estimated separately by decade. The hazard model identifies three similarity channels — geographic, demographic, and political — and allows their coefficients to vary over time.&lt;/p&gt;
&lt;p&gt;The central finding is a structural break in diffusion patterns around 2000. From the 1950s to the 1990s, geographic proximity is the dominant predictor of policy adoption: the coefficient on geographic similarity is 0.34 in the 1970s and remains roughly constant at 0.33 in the most recent decade. Demographic similarity is consistently positive and stable (approximately 0.20 in the 1980s, 0.22 in the 2010s). Political similarity — measured by closeness in Republican vote-share from the most recent presidential election — is a modest predictor before 2000, with coefficients between 0.14 (1970s) and 0.17 (1990s). Since 2000, the political similarity coefficient triples: 0.46 in the 2000s and 0.52 in the 2010s, making it by far the strongest predictor. The overall pseudo R-squared rises from 0.13 in the 1970s to 0.19 in the 2010s.&lt;/p&gt;
&lt;p&gt;These patterns are more pronounced for policies studied by economists: in the NBER subsample, the political similarity coefficient reaches 0.66 (s.e.=0.09) in the most recent two decades, versus 0.42 (s.e.=0.04) in the SPID sample.&lt;/p&gt;
&lt;p&gt;The paper tests whether the increased role of political similarity reflects correlated voter preferences, learning, or competition versus party discipline. Against correlated-preferences explanations: adding cross-state migration flows as a similarity measure reduces geographic predictive power but leaves the political similarity coefficient entirely unchanged; and typical policy-outcome variables (poverty rate, opioid mortality, income) have not become more correlated among politically similar states over time. In favor of party discipline: similarity in unified state government has zero predictive power through the 1990s but a coefficient of 0.42 (s.e.=0.06) in the 2000s–2010s. An event study of switches to unified party control confirms this causally for 1991–2020: switching to unified government raises the probability of passing ideologically aligned laws by approximately 2 percentage points in the four years following the switch, with no pre-trends and no effect on neutral-leaning laws; the same event study for 1950–1990 yields no detectable effect.&lt;/p&gt;
&lt;p&gt;COVID policies (77 state laws since October 2019) show strong political similarity in adoption; historical vaccination mandate policies (28 laws since 1975) show no political similarity effect. The paper concludes that rising party polarization at the state level — detectable from the 2000s onward, lagging the Congressional trend by roughly four to five decades — is the primary driver of the shift in diffusion patterns. The authors additionally classify each of the 57 NBER-sample policies by type of diffusion as an input for difference-in-differences research design assessment.&lt;/p&gt;
&lt;p&gt;Q: What data do the authors use and what is its scope?
A: The main source is the SPID Database (Boehmke et al., 2020), covering 676 policies over seven decades. The authors supplement this with 57 policies hand-collected from 91 NBER working papers (2012–2021) that use state-level policy variation. The combined sample covers 733 policies adopted from the 1950s onward in the contiguous 48 states, with the SPID sample averaging 23 adopting states per policy and the NBER sample averaging 29.&lt;/p&gt;
&lt;p&gt;Q: Do states with more resources or larger populations systematically innovate more policies?
A: The evidence for a state-capacity hypothesis is weak. There is only suggestive evidence that higher per-capita income predicts being in the top-20% of innovators, and no clear difference in population between the top and bottom innovators. Innovations arise from both Republican and Democratic states. One consistent correlate is urban population share, but overall innovation is largely idiosyncratic with respect to observable characteristics.&lt;/p&gt;
&lt;p&gt;Q: What was the dominant predictor of policy diffusion before 2000?
A: Geographic proximity was the dominant predictor. The coefficient on geographic similarity in the hazard model is 0.34 in the 1970s and remains stable at approximately 0.33 in the 2010s. Demographic similarity contributes consistently at approximately 0.20. Political similarity before 2000 is modest, ranging from 0.14 in the 1970s to 0.17 in the 1990s — roughly one-third to one-half the magnitude of the geographic coefficient.&lt;/p&gt;
&lt;p&gt;Q: How dramatically does political similarity change after 2000, and is this finding robust?
A: The political similarity coefficient triples, rising from 0.17 in the 1990s to 0.46 in the 2000s and 0.52 in the 2010s, making it the largest single predictor in recent decades. This pattern is robust across linear probability models, alternative measures of political similarity, alternative thresholds for &amp;ldquo;closest&amp;rdquo; states (closest fifth, fourth, third, or half all yield comparable coefficients), and alternative ways of computing adoption counts.&lt;/p&gt;
&lt;p&gt;Q: Is the shift toward political diffusion stronger for policies economists study?
A: Yes. In the NBER subsample, the political similarity coefficient reaches 0.66 (s.e.=0.09) in the 2000s–2010s, compared to 0.42 (s.e.=0.04) in the SPID sample. Geographic similarity also has somewhat higher coefficients in the NBER sample throughout the period. This implies that the policies most studied for difference-in-differences evaluation are also those most subject to politically-driven diffusion.&lt;/p&gt;
&lt;p&gt;Q: What does the Medicaid case study illustrate about political polarization?
A: ACA Medicaid expansion spread almost exclusively along partisan lines, with Republican vote-share accurately predicting the year of adoption. Crucially, the states that delayed or declined adoption — higher Republican vote-share states — had a higher share of population that would benefit from the expansion and therefore face a worse policy-need match. By contrast, the original 1966 Medicaid rollout showed no relationship between state political leaning and timing of adoption, and neither did the 1960s–1970s food stamp program expansion.&lt;/p&gt;
&lt;p&gt;Q: How do the authors distinguish party discipline from correlated voter preferences as the mechanism?
A: Two tests point away from correlated preferences: (1) cross-state migration flows, when added as a similarity measure, absorb geographic predictive power but leave the political similarity coefficient entirely unaffected; (2) typical policy-outcome variables (opioid mortality, poverty rate, income, etc.) have not become more correlated among politically similar states over time, contradicting the hypothesis that local needs or environments have become politically correlated.&lt;/p&gt;
&lt;p&gt;Q: What is the direct evidence for party discipline as the operative mechanism?
A: The authors construct a measure of similarity based on unified party control (governor and both chambers of the same party). This variable has zero predictive power through the 1990s (point estimate near zero). In the 2000–2020s, the coefficient for unified-government similarity is 0.42 (s.e.=0.06), making it the strongest single predictor of adoption in those decades. States with divided governments show no predictive power of adoption by other divided-government states, further isolating the role of party control.&lt;/p&gt;
&lt;p&gt;Q: What does the event-study of switches to unified party control show?
A: Switches to unified party control in 1991–2020 produce a statistically significant increase of approximately 2 percentage points in the probability of adopting ideologically aligned laws within four years of the switch, relative to the year before. The effect emerges in year t+1 and is persistent, with no pre-trends, and the effect on neutral-leaning laws is zero, ruling out a simple reduced-gridlock story. The same event study for 1950–1990 detects no effect.&lt;/p&gt;
&lt;p&gt;Q: How do COVID state policies compare to historical vaccination policies in terms of political diffusion?
A: COVID policies (77 state laws, October 2019–August 2021) show significant political similarity in adoption, consistent with the recent-decade patterns. Vaccination mandate laws (28 policies since 1975) show no political similarity effect whatsoever, with demographic and modest geographic similarity being the relevant predictors. This contrast underscores that political polarization in policy adoption is a recent phenomenon that has spread even to policy areas without prior partisan patterning.&lt;/p&gt;
&lt;p&gt;Q: How does partisan polarization at the state level compare temporally to polarization in Congress?
A: Congressional polarization (measured by DW-NOMINATE) has been rising since the 1950s. State-level policy polarization, as documented here, does not emerge until the 2000s — a lag of roughly four to five decades. The paper notes it has risen rapidly and has already reached policy domains (such as COVID mandates) that showed no political patterning historically.&lt;/p&gt;
&lt;p&gt;Q: Does the diffusion pattern vary across policy types?
A: Yes. For economic policies, geography and demographics decline in importance over time with a smaller increase in political predictors. For non-economic (social) policies, geographic importance remains stable while political polarization is especially strong. Political polarization is strongest in Republican-leaning and Democratic-leaning states, and weaker among battleground states, consistent with a party-driven model where ideologically extreme states adopt from each other.&lt;/p&gt;
&lt;p&gt;Q: How do the authors classify individual NBER-sample policies by diffusion type?
A: Using Geary&amp;rsquo;s C statistics computed separately for geographic and political clustering for each of the 57 NBER policies, the authors identify three approximate clusters: (1) primarily politically-clustered (e.g., Medicaid expansion); (2) jointly geographically and politically clustered (e.g., ban on asking about past salary history); and (3) largely idiosyncratic, neither geographically nor politically clustered (e.g., anti-bullying laws). This classification has direct implications for assessing identification threats in difference-in-differences designs.&lt;/p&gt;
&lt;p&gt;Q: What does the overall predictability of policy adoption look like over time?
A: The pseudo R-squared from the logit hazard model rises from 0.13 in the 1970s to 0.19 in the 2010s. The increase in political similarity is large enough not only to surpass geographic similarity as a predictor but to make the overall process of state policy adoption more predictable over time.&lt;/p&gt;
&lt;p&gt;Policy diffusion: The process by which a policy adopted in one state subsequently spreads to other states; measured here along geographic, demographic, and political dimensions using a logit hazard model estimated by decade.&lt;/p&gt;
&lt;p&gt;Geary&amp;rsquo;s C statistic: A ratio of weighted to unweighted average pairwise squared differences in adoption status, adapted from spatial statistics (Geary, 1954). Values below 1 indicate clustering; values above 1 indicate anti-clustering. The paper reports 1−C so higher values mean more clustering among similar states.&lt;/p&gt;
&lt;p&gt;Policy innovation: First-year adoption of a law in any state; a state is an &amp;ldquo;innovator&amp;rdquo; if it adopts in the first year the policy appears anywhere. The paper distinguishes innovation (origination) from diffusion (spread).&lt;/p&gt;
&lt;p&gt;Logit hazard model: A discrete-time logit model estimated at the state-year-policy level for all states that have not yet adopted a given policy, with policy-decade fixed effects as a baseline hazard and three time-varying similarity measures (geographic, demographic, political) as key predictors.&lt;/p&gt;
&lt;p&gt;Political similarity: Closeness of two states&amp;rsquo; Republican vote-shares from the most recent presidential election; the closest third of states in this dimension are used to construct the diffusion measure. Shown to be independent of — and to have grown far more predictive than — geographic similarity since 2000.&lt;/p&gt;
&lt;p&gt;Unified party control: A state government in which the governor and both state legislative chambers belong to the same party. The paper shows this is the variable most predictive of politically-driven policy diffusion in the 2000s–2020s, with a coefficient of 0.42 where it was effectively zero before 2000.&lt;/p&gt;
&lt;p&gt;Party discipline / party polarization: The paper&amp;rsquo;s preferred explanation for post-2000 patterns: state politicians increasingly vote and adopt policies along party lines beyond what voter preferences alone would predict, with the effect detectable since the 2000s at the state level, lagging the Congressional polarization trend by roughly four decades.&lt;/p&gt;</description></item><item><title>Politics at Work</title><link>https://macropaperwarehouse.com/papers/politics-at-work/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/politics-at-work/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Do individual political views shape firm behavior and labor market outcomes in the private sector? Specifically, do business owners sort copartisan workers into their firms, and does employers&amp;rsquo; political discrimination drive this sorting?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Setting&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper studies the complete Brazilian formal labor market over 2002–2019, assembling a novel longitudinal worker-firm-owner-party matched dataset from three administrative sources: (1) RAIS (Relação Anual de Informações Sociais), the universe of formal-sector workers (87 million unique workers, 7.6 million unique firms); (2) the Receita Federal do Brazil (RFB) and Cadastro Nacional de Empresas (CNE), containing business ownership structures for all registered firms; and (3) the Tribunal Superior Eleitoral (TSE) registry of all party members (19.3 million individuals) over 2002–2019. Matching these sources yields political affiliation for 11.4% of all private-sector owners and 7.8% of all private-sector workers in the sample. Party affiliation in Brazil requires an active registration step and is interpreted as a signal of strong and visible political views, distinguishing affiliated from unaffiliated individuals who likely hold milder views. The 35 parties in the sample are highly fragmented; the top 7 account for nearly 70% of all party members.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Political assortative matching.&lt;/em&gt; Using a likelihood ratio index (Eika et al., 2019; Chiappori et al., 2020), the paper finds that workers and owners belonging to the same party are on average about twice as likely to match in the labor market relative to random matching. Once within-municipality geographical sorting is accounted for, this figure falls to approximately 55% excess probability of copartisan matching, and increases over time: from 1.41 in 2002–2006 to 1.67 in 2016–2019. A dyadic regression approach — constructing all worker-firm dyads within industry-municipality labor markets and controlling for shared gender, race, age, and education — confirms the result: across all years, a politically affiliated worker is between 41% and 75% more likely to be employed by a copartisan owner than by an owner affiliated with a different party. Political assortative matching is driven both by higher hiring probabilities (range: 32%–59% more likely for copartisans, hiring margin only) and by longer tenure: copartisan workers stay in the firm roughly 5.5% longer than otherwise comparable workers of a different party, even within the same firm and hire-year (column 3 of Table 2). In every year and by every method, the degree of political assortative matching exceeds that of gender (15%–31% excess probability under dyadic approach) and race (approximately 3.4%), which are themselves both positive and significant.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Mechanisms: political discrimination.&lt;/em&gt; Three sets of evidence point to employer political discrimination as a relevant driver. First, in the administrative micro-data: assortative matching decreases strongly with firm size — it is more than twice as large in firms with up to 10 employees than in medium firms and more than six times as large as in firms with more than 50 employees — and is stronger for higher occupational layers and for jobs requiring above-median social skills or interpersonal relationships. Political assortative matching is, if anything, larger for parties not in power locally, inconsistent with a patronage mechanism. An event study of 5,262 owners who switched party finds a sharp increase of about 0.2 standard deviations in hires from the new party and a corresponding drop in hires from the old party at the time of the switch, with the share of workers from the new party rising by roughly 5 percentage points persistently. Second, an incentivized resume rating (IRR) field experiment (150 business owners; nondeceptive design) shows that owners rate copartisan resumes 0.213 points higher on a 1–7 Likert scale (a 7.4% increase relative to the mean rating for different-party resumes, statistically significant at p &amp;lt; 0.05), with no significant effect on perceived candidate acceptance probability. Third, a representative survey of 891 owners and 1,003 workers finds that belief-based and taste-based discrimination are ranked as the leading explanations by both groups; 47% of owners and 58% of workers agree with the belief-based discrimination statement. Additionally, 29% of surveyed owners (22% say &amp;ldquo;Yes&amp;rdquo; and 7% &amp;ldquo;In some cases&amp;rdquo;) explicitly reveal that political views affect their hiring decisions.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Real consequences.&lt;/em&gt; Conditional on employment, copartisan workers are promoted faster: they are 0.448 percentage points more likely to be promoted from white-collar to managerial positions (against a base rate of 2.58%) and 0.44 percentage points more likely to be promoted from blue-collar to white-collar positions (base rate 2.98%). Workers from a different party than the owner face a promotion penalty of 0.104–0.180 percentage points for white-collar-to-manager promotions. On wages, copartisan workers earn 3.9% more than unaffiliated coworkers within the same firm and year (firm-year FE specification); the effect is 2.8% when restricting to the same occupation within the firm. Workers from a different party earn 1.6% less. Decomposing by tier: managers (copartisan premium 1.6%), white-collar workers (3.4%), blue-collar workers (1.5%). Despite better outcomes, copartisan workers are 2.1 percentage points (2.3% relative to the mean) less likely to be educationally qualified for their occupation, conditional on firm-year and controlling for a full set of demographics. Finally, a higher share of copartisan workers in the prior year is associated with lower firm employment growth (estimated β = −0.071), corresponding to approximately a 1 percentage point gap in annual growth rate for a one-standard-deviation difference in copartisan share — substantial relative to an average annual growth rate of 10%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;All findings pertain to the formal private sector in Brazil over 2002–2019. Political affiliation in the Brazilian system requires an active step and signals strong views; results apply to the approximately 7.8%–11.4% of workers and owners who are party-registered. The field experiment sample is limited to 150 business owners affiliated with major Brazilian parties who were actively seeking to hire. The firm growth result is explicitly characterized as suggestive, without a source of exogenous variation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-likelihood-ratio-index-and-what-does-it-show-for-political-matching-in-brazil"&gt;Q1. What is the likelihood ratio index and what does it show for political matching in Brazil?&lt;/h3&gt;
&lt;p&gt;The likelihood ratio index measures how many times more likely a match between a worker and owner of the same party is, relative to the expected frequency under random matching (conditional on the population shares of each party). Across 2002–2019, the unconditional index ranges from 1.56 to 1.85, implying workers and employers of the same party are on average about twice as likely to match as under random matching. After accounting for geographic sorting within municipalities, the index ranges from approximately 1.41 (2002–2006 average) to 1.67 (2016–2019 average), showing a clear increasing trend. The corresponding gender and race indexes average about 1.2 and 1.35, respectively, in the basic specification, both significantly lower than the party index in every year of the sample.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-dyadic-regression-estimates-control-for-omitted-characteristics-and-what-do-they-find"&gt;Q2. How do the dyadic regression estimates control for omitted characteristics, and what do they find?&lt;/h3&gt;
&lt;p&gt;The dyadic regression constructs all possible worker-firm pairs within each municipality-industry labor market in a given year. The dependent variable is an indicator for whether worker i is employed by firm f. The key coefficient of interest is the differential probability of employment for a copartisan pair relative to a different-party pair, controlling for indicators for shared gender, race, age bracket, and education level, as well as worker occupation fixed effects and experience. This controls for the concern that politically affiliated individuals share non-political traits that correlate with employment choices. After these controls, a politically affiliated worker is 41%–75% more likely (depending on year) to be employed by a copartisan owner than by a different-party owner. The effect stems primarily from copartisan workers being preferentially hired (not just from unaffiliated owners preferring any affiliated worker indiscriminately). The analogous dyadic estimate for shared gender is 15%–31% and for shared race is approximately 3.4%, both lower than the party estimate in all years.&lt;/p&gt;
&lt;h3 id="q3-how-is-political-assortative-matching-decomposed-into-hiring-versus-retention-margins"&gt;Q3. How is political assortative matching decomposed into hiring versus retention margins?&lt;/h3&gt;
&lt;p&gt;To isolate the hiring margin, the authors estimate the dyadic regression restricting to newly hired workers (not present in the firm in year t-1). They find that the probability of being hired by a copartisan owner is 32%–59% higher than by a different-party owner across years. The retention (tenure) margin is estimated by regressing the share of subsequent years a worker remains at the firm on partisan alignment at the time of hire. In the most stringent specification (year-of-hire × firm fixed effects), copartisan hires stay 5.5 percentage points longer (as a share of post-hire years) than different-party hires from the same firm and hire-year cohort. Both margins are significant, and both exhibit stronger political sorting than equivalent estimates for gender or race.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-evidence-against-political-patronage-as-the-primary-driver-of-political-assortative-matching"&gt;Q4. What is the evidence against political patronage as the primary driver of political assortative matching?&lt;/h3&gt;
&lt;p&gt;If political patronage (parties pressuring owners to hire copartisans) were the main driver, we would expect political assortative matching to be stronger when the owner&amp;rsquo;s party is in power locally, as those parties have greater leverage over business owners. The authors estimate a modified dyadic regression distinguishing between cases where the owner&amp;rsquo;s party is in the ruling coalition of the municipal mayor or state governor versus not in power. The results show that political assortative matching is, if anything, larger for parties not in power. This is inconsistent with patronage being the dominant mechanism and consistent with the discrimination channel being driven by owner preferences rather than external political pressure.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-event-study-of-owner-party-changes-show"&gt;Q5. What does the event study of owner party changes show?&lt;/h3&gt;
&lt;p&gt;The event study tracks 5,262 owners who switch party affiliation during 2002–2019, comparing their firms to control firms in the same market whose owners remain affiliated to the original party. At the time of the switch, there is a sharp increase of approximately 0.2 standard deviations in hires from the owner&amp;rsquo;s new party and a corresponding sharp decrease in hires from the old party. Hires from other parties and unaffiliated hires also decline modestly. The share of the workforce affiliated with the new party increases by roughly 5 percentage points and remains elevated in subsequent years. Because nonpolitical network ties (shared school, neighborhood, sports team) are unlikely to dissolve abruptly when an owner changes party, this design provides additional evidence that the change in hiring is driven by a direct change in the owner&amp;rsquo;s political preferences rather than by network overlap.&lt;/p&gt;
&lt;h3 id="q6-what-was-the-design-of-the-incentivized-resume-rating-experiment-and-why-does-it-identify-political-discrimination"&gt;Q6. What was the design of the incentivized resume rating experiment and why does it identify political discrimination?&lt;/h3&gt;
&lt;p&gt;The experiment was conducted with 150 Brazilian business owners recruited from the administrative data (who are already known to be affiliated with one of six major parties), targeting owners with active hiring interest through a leading job platform. Owners rated 20 synthetic resumes with fully randomized features (education, experience, training, skills, formatting). Sixteen resumes had no partisan cues; two contained cues signaling copartisanship with the rating owner; two signaled a party from the opposite side of the political spectrum. Incentives were provided by committing to send respondents real job-seeker profiles from the platform chosen by machine learning based on revealed preferences. Because all resume features other than the partisan cue were randomized, the experiment shuts down shared nonpolitical networks and patronage as explanations; the only channel is the employer&amp;rsquo;s direct preference for the candidate&amp;rsquo;s partisan affiliation. The response rate was 11% and the survey was conducted March–May 2022.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-quantitative-magnitude-of-the-field-experiment-result"&gt;Q7. What is the quantitative magnitude of the field experiment result?&lt;/h3&gt;
&lt;p&gt;Owners rate copartisan resumes 0.213 points higher on the 1–7 Likert scale relative to resumes from the opposite side of the political spectrum (statistically significant at p &amp;lt; 0.05), representing a 7.4% increase relative to the mean rating of different-party resumes (2.950). When resume-level controls (gender, high-skill experience flag, years of experience, programming skills, training) are added, the estimate is 0.254. There is no statistically significant effect on owners&amp;rsquo; perceived likelihood that a candidate would accept a job offer (coefficient 0.150–0.158, not significant), suggesting that the observed difference in interest ratings reflects a genuine direct preference for copartisans, not an expectation that copartisans are more likely to accept.&lt;/p&gt;
&lt;h3 id="q8-what-do-the-survey-findings-add-about-mechanisms-and-the-prevalence-of-political-discrimination"&gt;Q8. What do the survey findings add about mechanisms and the prevalence of political discrimination?&lt;/h3&gt;
&lt;p&gt;The survey of 891 owners and 1,003 workers (response rate 26.84%) presents five candidate mechanisms and asks respondents to evaluate each. Both groups rank belief-based discrimination (owners believe copartisans would be more productive) as the most likely explanation: 47% of owners and 58% of workers partially or strongly agree. Taste-based discrimination is second (36% owners, 52% workers agree), followed by networks (39% owners, 49% workers). Patronage and workers&amp;rsquo; preferences attract little agreement from either group. Among owners ranked by single strongest agreement, 29.7% most strongly agree with belief-based discrimination and 22.0% with taste-based, while 29% of all surveyed owners explicitly stated that political views do affect their hiring decisions. These patterns are broadly similar regardless of the respondent&amp;rsquo;s own political affiliation status.&lt;/p&gt;
&lt;h3 id="q9-how-large-are-the-political-promotion-and-wage-premia-and-how-do-they-compare-to-gender-and-race-effects"&gt;Q9. How large are the political promotion and wage premia, and how do they compare to gender and race effects?&lt;/h3&gt;
&lt;p&gt;For promotions, copartisan white-collar workers are 0.448 percentage points more likely to be promoted to manager (relative to unaffiliated co-workers hired in the same firm-year), against a base promotion rate of 2.58% — an effect of approximately 17% of the mean. For blue-collar-to-white-collar promotion, the copartisan premium is 0.44 percentage points against a base rate of 2.98%. For wages, copartisans earn 3.9% more than unaffiliated co-workers within the same firm and year; restricting to the same occupation within the firm, the premium is 2.8%. The political wage premium (3.9%) exceeds the gender wage premium (1.5%) and the race wage premium (1.0%) in the same specification. Workers from a different party than the owner earn 1.6% less than unaffiliated co-workers within the same firm-year.&lt;/p&gt;
&lt;h3 id="q10-are-copartisan-workers-better-qualified-than-those-they-displace-and-what-does-this-imply-for-firm-performance"&gt;Q10. Are copartisan workers better qualified than those they displace, and what does this imply for firm performance?&lt;/h3&gt;
&lt;p&gt;Copartisan workers are significantly less qualified in terms of education relative to their occupation: they are 2.1 percentage points less likely to be educationally qualified for their position than their unaffiliated co-workers within the same firm-year (2.3% relative to the mean qualification rate of 93.2%), with the largest effects for managers. Workers of a different party show only a small and economically negligible qualification gap. The fact that copartisans are paid more, promoted faster, and yet are less qualified is consistent with political discrimination substituting for competence in personnel decisions. The qualification shortfall is specifically attributed to copartisanship and not to shared gender, race, age, or education between owner and worker, as those coefficients are economically small.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-evidence-on-firm-growth-and-what-are-the-limitations-of-that-evidence"&gt;Q11. What is the evidence on firm growth and what are the limitations of that evidence?&lt;/h3&gt;
&lt;p&gt;Firms with a higher share of copartisan workers in the prior year grow less. The estimated coefficient β = −0.071, and a one-standard-deviation difference in the copartisan share is associated with approximately a 1 percentage point gap in annual employment growth, relative to a mean growth rate of 10%. The specification compares firms of the same size and with the same number of affiliated workers in the same year. The result is robust to adding municipality and municipality-industry fixed effects. The authors explicitly characterize this evidence as suggestive, noting the absence of an exogenous source of variation in political discrimination. The negative association is more consistent with taste-based discrimination (Becker, 1957) — in which politically homogeneous firms sacrifice productivity for the owners&amp;rsquo; amenity of employing copartisans — than with accurate belief-based discrimination.&lt;/p&gt;
&lt;h3 id="q12-how-is-political-assortative-matching-distributed-across-parties-and-does-it-depend-on-party-ideology"&gt;Q12. How is political assortative matching distributed across parties and does it depend on party ideology?&lt;/h3&gt;
&lt;p&gt;The likelihood ratio index shows large assortative matching across the entire political spectrum. For most years, relatively more ideologically extreme parties — on the left (PT, PDT) and on the right (PP, DEM) — display higher assortative matching than more centrist parties (PMDB, PSDB). This pattern is consistent with stronger partisan identity at the extremes leading to stronger preferences for copartisan workers, but the paper does not formally model the mechanism behind this heterogeneity.&lt;/p&gt;
&lt;h3 id="q13-what-is-the-role-of-workers-preferences-as-opposed-to-employers-discrimination-and-how-can-wages-distinguish-them"&gt;Q13. What is the role of workers&amp;rsquo; preferences as opposed to employers&amp;rsquo; discrimination, and how can wages distinguish them?&lt;/h3&gt;
&lt;p&gt;If workers have a preference for working with copartisan owners (treating this as a job amenity), compensating differentials theory would predict a negative wage premium for copartisan workers — they would accept lower wages in exchange for working with like-minded owners. The data show the opposite: copartisan workers earn significantly more, not less, than their unaffiliated co-workers. This evidence is inconsistent with workers&amp;rsquo; preferences being the primary driver of political assortative matching, and is instead consistent with employers&amp;rsquo; discrimination. The survey evidence corroborates this: both owners and workers assign low priority to the &amp;ldquo;workers&amp;rsquo; preferences&amp;rdquo; mechanism.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Political assortative matching&lt;/strong&gt;: The phenomenon by which workers and business owners belonging to the same political party are matched in the labor market at rates significantly exceeding what would occur under random matching within the local labor market. Measured via the likelihood ratio index and dyadic regressions that control for shared demographic characteristics. In this paper, political assortative matching is larger in magnitude than assortative matching along gender or racial lines.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Likelihood ratio index (S)&lt;/strong&gt;: A measure of assortative matching defined as the weighted sum of the ratios of observed same-party co-occurrence probabilities to their expected probabilities under random matching. S &amp;gt; 1 indicates positive assortative matching. The paper uses both a basic version and a geography-adjusted version that computes the index within municipalities to control for geographic concentration of party membership.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dyadic regression&lt;/strong&gt;: A regression approach that constructs all possible worker-firm pairs within a defined labor market (municipality × 2-digit industry) to estimate the differential probability that a worker is employed by a copartisan firm relative to a different-party firm. The key advantage is the ability to control simultaneously for multiple shared demographic characteristics between worker and owner, accounting for the correlation of assortative criteria.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Incentivized resume rating (IRR) experiment&lt;/strong&gt;: A nondeceptive field experiment design (following Kessler et al., 2019) in which business owners rate synthetic resumes with fully randomized characteristics. Truthful rating is incentivized because respondents are told that their revealed preferences will be used to select real job-seeker profiles sent to them by a partner platform via machine learning. This design allows direct identification of employer preference for copartisan candidates while ruling out alternative channels such as shared nonpolitical networks or patronage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Political wage premium&lt;/strong&gt;: The percentage wage difference earned by copartisan workers relative to unaffiliated co-workers within the same firm-year (and occupation), after controlling for a full set of socio-demographic characteristics. A positive political wage premium is the paper&amp;rsquo;s primary piece of evidence that workers&amp;rsquo; compensating differentials cannot explain political assortative matching, since amenity-based sorting would predict a negative premium.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Political promotion premium&lt;/strong&gt;: The differential probability that a copartisan worker is promoted to a higher organizational layer (blue-collar to white-collar, or white-collar to manager) relative to an unaffiliated co-worker hired in the same firm and year, net of demographic controls.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Educational mismatch (Qualified)&lt;/strong&gt;: An indicator variable equal to one if a worker&amp;rsquo;s educational level meets or exceeds the educational level required by their specific occupation in the CBO (Classificação Brasileira de Ocupações) classification. Used to assess whether politically favored (copartisan) workers are less competent along this observable dimension.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Belief-based discrimination vs. taste-based discrimination&lt;/strong&gt;: Two distinct theoretical channels for employer political discrimination. Belief-based discrimination (Phelps, 1972; Arrow, 1973) occurs when employers perceive copartisans to be more productive — e.g., because shared political views reduce intra-firm conflict. Taste-based discrimination (Becker, 1971) occurs when employers have a direct utility-affecting preference for copartisan workers, independent of productivity beliefs. The paper treats these as observationally distinct from patronage and network overlap, and uses the negative correlation between political homogeneity and firm growth as suggestive evidence favoring the taste-based channel.&lt;/p&gt;</description></item><item><title>Praying for Rain</title><link>https://macropaperwarehouse.com/papers/praying-for-rain/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/praying-for-rain/</guid><description>&lt;p&gt;This paper studies rainmaking as an instrumental religious belief. The central research question is: why do people believe that prayer can bring rain, even though it does not work? The authors develop a model of cultural evolution in which a religious leader prays for rain at an arbitrary time, and people update their beliefs about whether the leader can cause rainfall based on whether rain follows. The key mechanism is the local rainfall hazard function — the probability of rain conditional on how many days have passed since the last rainfall. In environments where the hazard is increasing (rain becomes more likely the longer a drought continues), a leader who prays during a drought will tend to be followed by rain, creating the illusion of efficacy. In environments with a flat or declining hazard, prayer cannot be systematically followed by rain in a persuasive way. The model yields five predictions: rain ritual traditions will select for prayers correlated with rainfall; the level of average rainfall does not determine persuasiveness; constant-hazard environments cannot support persuasive prayer; increasing-hazard environments are more likely to adopt rainmaking; and higher net benefits of rainfall (e.g., settled agriculture) further increase the likelihood of ritual.&lt;/p&gt;
&lt;p&gt;The authors test these predictions with two empirical strategies. First, they use daily data from the Catholic church in Murcia, Spain, covering 1600 to 1836. Church records provide the daily timing of pro pluvia rogations (prayers for rain), while municipal council records — kept independently of the church — record notable rainfall events. Murcia&amp;rsquo;s rainfall hazard is estimated to be increasing after long dry spells: the hazard rate after a long drought is roughly double the hazard rate two months after the last rainfall. The main finding is that a prayer for rain in the last 30 days predicts a 0.144 percentage-point higher daily probability of notable rainfall (standard error 0.057 pp), relative to a baseline mean daily rainfall probability of 0.203 pp — a 71% increase in the predicted probability. Prayer also Granger-causes rainfall conditional on lags of recent rainfall, and the predictive power holds within a given calendar month, ruling out a purely seasonal coincidence.&lt;/p&gt;
&lt;p&gt;Second, the authors construct an original dataset covering rainmaking practice for 1,208 ethnic groups drawn from the Ethnographic Atlas (Murdock, 1967), coded from 370 anthropological sources. They match each ethnic group to its nearest weather station and estimate the rainfall hazard function each group faces in its ancestral location. Of the 1,208 groups, 33% face an increasing rainfall hazard, and 39% of all groups practice rain ritual. The main global finding is that ethnic groups facing an increasing rainfall hazard are 14 percentage points more likely to practice rainmaking (standard error 3.7 pp), relative to a base rate of 30% among groups facing a non-increasing hazard — a 47% increase. This result is robust to continent fixed effects, geographic and climatic controls (longitude, latitude, elevation, distance to coast, ruggedness, mean temperature, mean rainfall, coefficient of variation of rainfall, maximum dry spell length, and the Giuliano-Nunn 2021 climatic variability measure), alternative hazard estimation methods, and linguistic family fixed effects. Crucially, lower average rainfall, longer droughts, and greater climatic variability are not associated with more rain ritual conditional on hazard shape — it is specifically the shape of the hazard function, not aridity or variability per se, that drives adoption.&lt;/p&gt;
&lt;p&gt;A second global finding concerns demand: groups dependent on agriculture are 11 pp more likely to practice rainmaking; those dependent on intensive agriculture, 21 pp more likely; and those dependent on intensive irrigated agriculture, 32 pp more likely (on a base of 32%). The scope of the findings is the pre-modern or traditional period captured by the Atlas; the Murcia case covers 1600–1836. The authors conclude that some environments create an illusion of efficacy that sustains instrumental religious belief through cultural selection, without requiring that believers be irrational.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s central theoretical claim about why rainmaking beliefs persist?
A: The paper argues that in environments where the rainfall hazard is increasing during a drought, a leader who begins praying during a dry spell will tend to be followed by rain, because the probability of rain rises as the drought lengthens. People who cannot observe the counterfactual hazard (what rainfall would have been without prayer) interpret this coincidence as evidence that prayer works. Cultural selection then favors leaders whose prayer timing is more persuasive, causing the belief to persist across generations even though prayer does not actually cause rain.&lt;/p&gt;
&lt;p&gt;Q: What is the rainfall hazard function, and why does its shape determine whether prayer can be persuasive?
A: The hazard function h(t) gives the instantaneous probability of rain at time t days after the last rainfall. If the hazard is flat, the probability of rain is the same regardless of whether prayer was offered or not, so there is no systematic correlation between prayer and rainfall to exploit. If the hazard is declining, prayer during a drought will be followed by lower-than-average rainfall probability, undermining the leader. Only if the hazard is increasing does prayer during a long dry spell systematically coincide with a higher probability of rain, creating a persuasive correlation.&lt;/p&gt;
&lt;p&gt;Q: What do Propositions 2 and 3 of the model establish?
A: Proposition 2 establishes that if the hazard rate is constant and a person&amp;rsquo;s prior belief that prayer works is below 0.5, then no prayer start time can persuade them to support the leader. Proposition 3 establishes the converse: if the hazard rate is increasing and the prior is below 0.5, there exists a meaningful belief for which a person will support the leader for any prayer start time. Together these propositions identify the increasing hazard as the necessary and sufficient structural condition for persuasive prayer.&lt;/p&gt;
&lt;p&gt;Q: What is the main quantitative finding from Murcia, and what identification strategy supports it?
A: A prayer for rain in the last 30 days predicts a 0.144 percentage-point higher daily probability of notable rainfall (standard error 0.057 pp) relative to a baseline mean of 0.203 pp, a 71% increase. The authors additionally demonstrate that prayer Granger-causes rainfall conditional on lags of recent rainfall, and that the effect holds within a given calendar month, ruling out the explanation that prayer simply tracks the rainy season. The prayer and rainfall records are kept by independent institutions (church and municipal council), reducing the risk of strategic recording.&lt;/p&gt;
&lt;p&gt;Q: How does the hazard rate in Murcia behave, and does it satisfy the model&amp;rsquo;s key condition?
A: The hazard of rainfall in Murcia is initially high just after rain, declines to a minimum roughly two months after the last rainfall, and then increases significantly thereafter, reaching or exceeding its initial level after a long drought. The fluctuations are large: the hazard after a long dry spell is roughly double the hazard two months after rainfall. This U-shaped pattern means the hazard is increasing during a prolonged drought, satisfying the model&amp;rsquo;s key condition for persuasive prayer.&lt;/p&gt;
&lt;p&gt;Q: How was the global rainmaking dataset constructed, and what is its coverage?
A: The authors used the Ethnographic Atlas (Murdock, 1967) as a template, covering 1,290 ethnic groups, and combed 370 anthropological sources — primarily group-specific ethnographic monographs — to code rainmaking practice for 1,208 groups. A group is coded as practicing rain ritual only if there is clear evidence of a practice specifically intended to bring rain through supernatural means. The authors treat their measure as a lower bound. They find that 39% of the 1,208 groups practice rainmaking, across every settled continent.&lt;/p&gt;
&lt;p&gt;Q: What is the main global regression result and how robust is it?
A: Ethnic groups facing an increasing rainfall hazard are 14 percentage points more likely to practice rain ritual (standard error 3.7 pp) relative to a base rate of 30%, a 47% proportional increase. This coefficient is positive and statistically significant across all specifications, including those adding continent fixed effects, a full battery of geographic and climatic controls (longitude, latitude, elevation, distance to coast, ruggedness, mean temperature, mean rainfall, coefficient of variation of rainfall, maximum dry spell length, and the Giuliano-Nunn 2021 climatic variability measure), alternative hazard estimation methods, linguistic family fixed effects, and restrictions to groups with high-quality rainfall data.&lt;/p&gt;
&lt;p&gt;Q: Does aridity or climatic variability explain rainmaking adoption?
A: No. Lower average rainfall, longer droughts, and greater climatic variability (measured using the Giuliano-Nunn 2021 index) are not associated with more rain ritual practice, conditional on the shape of the hazard function. This rules out the naive hypothesis that people pray for rain simply because they do not get enough, or because their rainfall is unreliable. It is specifically the shape of the hazard — whether it is increasing during a drought — that drives adoption, not the level or volatility of rainfall.&lt;/p&gt;
&lt;p&gt;Q: How does demand for rainfall, proxied by agricultural subsistence, affect rainmaking adoption?
A: Groups dependent on agriculture are 11 percentage points more likely to practice rainmaking relative to other subsistence modes. Groups dependent on intensive agriculture are 21 percentage points more likely, and groups dependent on intensive irrigated agriculture are 32 percentage points more likely, all on a base of 32%. This gradient is consistent with Proposition 5 and 6 of the model: settled, location-specific agricultural investment raises the net benefit of rainfall control, increasing support for rain ritual independently of the persuasion channel.&lt;/p&gt;
&lt;p&gt;Q: What does the model&amp;rsquo;s cultural evolution mechanism (Proposition 4) predict about how prayer timing changes over generations?
A: Proposition 4 states that rituals with high support are more likely to persist. In increasing-hazard environments, random variation in prayer timing means some leaders gain more support than others; those with more persuasive timing are more likely to persist. Each generation then adopts a policy at least as persuasive as the prior generation, so support rises over time and prayers gradually converge toward the timing that maximizes persuasiveness. This mechanism does not require deliberate optimization by any individual leader.&lt;/p&gt;
&lt;p&gt;Q: How does the paper&amp;rsquo;s finding relate to the long-standing anthropological debate between the traditional and revisionist schools on rainmaking?
A: The traditional school (following Frazer 1890) holds that belief is instrumental — people engage in rainmaking to make rain, and belief responds to empirical evidence. The revisionist school (Wittgenstein, Durkheim) argues that religious belief and rationality are fundamentally separate, and religious practice is performative rather than evidence-responsive. The paper&amp;rsquo;s finding that rainmaking is more prevalent precisely where it is more persuasive — i.e., where the environment makes prayer appear to work — supports the traditional, instrumental interpretation that belief responds to evidence of efficacy.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions for the paper&amp;rsquo;s conclusions?
A: The Murcia case study covers the period 1600–1836, ending when the abolition of tithes reduced the church&amp;rsquo;s funding and influence; it applies to a sophisticated Catholic institutional context. The global analysis covers traditional practices of pre-modern ethnic groups as recorded in the Ethnographic Atlas and anthropological literature; it does not speak to modern religious practice or to religions after substantial modernization. The persuasion mechanism requires that people cannot directly observe what rainfall would have been without prayer, a condition satisfied in pre-scientific contexts.&lt;/p&gt;
&lt;p&gt;Rainfall hazard function: In this paper&amp;rsquo;s usage, the function h(t) = f(t)/(1-F(t)) giving the instantaneous probability of rainfall at time t days since the last rainfall. Its shape — whether flat, declining, or increasing during a drought — determines whether prayer can be persuasive, not the overall level of rainfall.&lt;/p&gt;
&lt;p&gt;Increasing hazard: A hazard rate that rises as the length of a dry spell increases, so that rain becomes more likely the longer the drought has continued. The paper defines this specifically as the derivative of the hazard function evaluated at the 99th percentile of spell length. This is the necessary structural condition for prayer to seem efficacious.&lt;/p&gt;
&lt;p&gt;Instrumental religious belief: Belief directed at achieving a worldly outcome (here, rainfall), as opposed to purely expressive or social belief. The paper treats belief as instrumental if it responds to perceived evidence of efficacy and is adopted where it appears to work.&lt;/p&gt;
&lt;p&gt;Persuasion (in the model): The process by which a leader&amp;rsquo;s prayer timing causes people to update their belief that prayer works, by generating a correlation between prayer and subsequent rainfall that exceeds what people expect from the background hazard rate. Persuasion is possible only when the hazard is increasing.&lt;/p&gt;
&lt;p&gt;Pro pluvia rogations: The Catholic church&amp;rsquo;s formal prayers for rain, practiced in Murcia since at least the 14th century. In the paper&amp;rsquo;s data, these prayers follow a pattern of escalation — increasing in number and intensity — during prolonged droughts, consistent with the model&amp;rsquo;s prediction about prayer timing.&lt;/p&gt;
&lt;p&gt;Cultural evolution: The paper&amp;rsquo;s framework (drawing on Henrich 2015) in which religious leaders act as cultural entrepreneurs; leaders whose prayer timing happens to be more persuasive gain greater support and are more likely to survive across generations, so prayer traditions drift toward more persuasive timing without deliberate design.&lt;/p&gt;
&lt;p&gt;Rain ritual (global measure): A binary indicator coded as one for an ethnic group if the anthropological literature contains clear evidence of a practice specifically intended to bring rain through supernatural means, including dances, sacrifices, prayers, and petitioning of rain deities. Treated by the authors as a lower bound on actual prevalence.&lt;/p&gt;</description></item><item><title>Racial disparities in crime and wealth</title><link>https://macropaperwarehouse.com/papers/racial-disparities-in-crime-and-wealth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/racial-disparities-in-crime-and-wealth/</guid><description>&lt;p&gt;This paper asks whether racial differences in labor income can simultaneously explain both the crime gap and the wealth gap between Black and White individuals in the United States. The authors build a large-scale overlapping generations (OLG) model in which property crime is endogenously determined — agents choose whether to steal alongside their consumption and savings decisions — while drug-related incarcerations are treated as exogenous, reflecting evidence that racial profiling distorts enforcement independently of offending behavior. The model is calibrated to match several well-documented racial disparities: Black individuals comprise 12.36% of the adult population but 33.8% of the incarcerated population; 42.7% of Black individuals fall in the bottom wealth quintile (below $3,400 in assets) versus 15.1% of White individuals; the median Black-White wealth gap is 89.5% (SCF 2019). Data sources include the Survey of Consumer Finances (SCF 2019), Uniform Crime Reports (UCR 1996–2011), NLSY79, PSID (1968–2021), and MORG (2000–2019).&lt;/p&gt;
&lt;p&gt;The model incorporates four dimensions of labor market disparity between Black and White agents: educational attainment, unemployment risk and duration, age-earnings profiles, and idiosyncratic income shock processes. It also incorporates race-skill-specific survival probabilities (life expectancy at birth: 73 years for Black, 78 years for White), scarring effects from incarceration on future labor income, a progressive income tax, means-tested transfers, and accidental bequests distributed within race groups.&lt;/p&gt;
&lt;p&gt;The benchmark model successfully replicates key data moments. Black individuals constitute 34.3% of the incarcerated population (data: 33.8%). The model-generated median wealth gap is 83.6% (data: 89.5%). The share of Black individuals in the bottom wealth quintile is 37.7% in the model versus 42.7% in the data. The model does not match the average wealth gap: the model-generated gap is 58.9% versus 84.4% in the SCF.&lt;/p&gt;
&lt;p&gt;The main counterfactual experiments yield three findings. First, equalizing labor market conditions — particularly age-earnings profiles — is the dominant driver of both racial wealth and crime disparities. When all labor market conditions are equalized, the Black crime rate falls by 66.25% (from 11.97% to 4.04%), the median wealth gap declines by 69.6% (from 83.58% to 25.4%), and the share of Black individuals in the bottom wealth quintile falls from 37.73% to 20.75%. Equalizing age-earnings profiles alone accounts for the largest single-factor effect: the median wealth gap declines from 83.58% to 44.16% and the Black crime rate from 11.97% to 7.59%. The resource cost of equalizing age-earnings profiles is estimated at 3.29% of GDP for the No-HS group and 22.2% of GDP for the HS group.&lt;/p&gt;
&lt;p&gt;Second, higher crime and incarceration rates among Black individuals do not significantly contribute to their lower wealth. When crime is entirely eradicated, the share of Black individuals in the bottom quintile barely moves (37.73% to 37.61%), and the median wealth gap falls only from 83.58% to 82.6%. The mechanism is that most crimes are committed by young, already-poor individuals who are not saving in any case; income loss during incarceration is not large enough to affect wealth accumulation meaningfully.&lt;/p&gt;
&lt;p&gt;Third, equalizing life expectancy generates a 25.39% reduction in the median wealth gap and a 12.4% decline in the share of Black individuals in the bottom wealth quintile, with negligible effect on crime rates.&lt;/p&gt;
&lt;p&gt;The paper also validates the model against Cesarini et al. (2023), who find a small, statistically insignificant effect of lottery wealth on criminal behavior in Sweden. The model replicates this finding: a $150,000 windfall reduces incarceration risk over seven years by 0.81 percentage points. The mechanism is that lottery winnings displace means-tested transfers, winnings gradually dissipate as low income persists, and individuals eventually return to poverty and resume criminal activity.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question and why does the paper treat property crime and drug crime differently?
A: The paper asks whether racial labor income differences can simultaneously account for both crime and wealth disparities. Property crimes are modeled endogenously because offending behavior responds rationally to economic incentives. Drug crime incarcerations are exogenous to capture evidence that racial profiling in enforcement — rather than differential offending alone — drives racial disparities in drug arrests: Beck and Blumstein (2018) show differential offending explains only about 52% of the drug imprisonment gap, versus over 70% for overall imprisonment.&lt;/p&gt;
&lt;p&gt;Q: What are the benchmark model&amp;rsquo;s key calibration targets and how well does it fit the data?
A: The benchmark targets Black individuals as 34.3% of the incarcerated population (data: 33.8%), a median Black-White wealth gap of 83.6% (data: 89.5%), and 37.7% of Black individuals in the bottom wealth quintile (data: 42.7%). The model does not match the average wealth gap: the model-generated gap is 58.9% versus 84.4% in the SCF, which the authors acknowledge explicitly.&lt;/p&gt;
&lt;p&gt;Q: What is the quantitative effect of equalizing all labor market conditions?
A: Experiment 5 (equalize educational attainment, unemployment risk, and age-earnings profiles jointly) reduces the Black crime rate by 66.25% (from 11.97% to 4.04%), the median wealth gap by 69.6% (from 83.58% to 25.4%), and the share of Black individuals in the bottom quintile from 37.73% to 20.75%. Equalizing all factors including life expectancy drives the median wealth gap to 0%, with the bottom-quintile share for Black individuals at 19.31%.&lt;/p&gt;
&lt;p&gt;Q: Which single labor market factor matters most for the wealth gap and crime rate?
A: Equalizing age-earnings profiles (Experiment 3) is the single most important factor, reducing the median wealth gap from 83.58% to 44.16% and the Black crime rate from 11.97% to 7.59%. By contrast, equalizing educational attainment or unemployment risk each reduces the median wealth gap only to 76.71%, with smaller crime effects.&lt;/p&gt;
&lt;p&gt;Q: Does education-group heterogeneity matter for interpreting the age-earnings equalization effect?
A: Yes, substantially. Equalizing age-earnings profiles for the No-HS group reduces the Black crime rate by 21% with little effect on the median wealth gap. Equalizing profiles for the HS group reduces the median wealth gap by approximately 40% with a much smaller effect on crime rates. The earnings channel to crime operates primarily at the bottom of the education distribution, while the earnings channel to wealth accumulation operates more strongly in the high school group.&lt;/p&gt;
&lt;p&gt;Q: Why does crime have so little effect on the wealth distribution?
A: Criminals are predominantly young and already-poor individuals who are not accumulating savings. Because these individuals have minimal assets and rely heavily on means-tested transfers for consumption, the income loss during incarceration does not reduce their wealth meaningfully. When crime is completely eradicated, the share of Black individuals in the bottom quintile falls only from 37.73% to 37.61% and the median wealth gap declines from 83.58% to only 82.6%.&lt;/p&gt;
&lt;p&gt;Q: What is the effect of eliminating drug-related incarcerations on the Black wealth distribution?
A: Experiment 4 (eliminating drug crime incarcerations) reduces the share of Black individuals in the bottom quintile only slightly, from 37.73% to 37.30%. Eliminating the scarring effect of all incarcerations likewise has negligible effects on the bottom-quintile share (37.53% versus 37.73% in the benchmark) and the zero-assets share (32.56% versus 33.24%). Neither the direct incarceration penalty nor its labor market scarring meaningfully affects wealth accumulation.&lt;/p&gt;
&lt;p&gt;Q: What happens to crime and wealth when the property crime clearance rate changes?
A: Doubling the clearance rate from 17.2% to 34.4% reduces the Black crime rate from 11.97% to 1.72% and the White rate from 3.05% to 0.52%, with minimal change in the wealth distribution (Blacks in bottom quintile: 37.83%). Halving the clearance rate to 8.6% more than doubles Black crime to 27.53% and White crime to 9.55%, and increases the share of Black individuals in the bottom quintile by about 11% to 42.01%. This asymmetry — crime reduction barely helps wealth but crime increase does hurt — is consistent with the poverty-trap mechanism.&lt;/p&gt;
&lt;p&gt;Q: How does the model validate against the Cesarini et al. (2023) Swedish lottery study?
A: Cesarini et al. find a small, statistically insignificant negative effect of a $150,000 lottery windfall on conviction rates. The model replicates this: simulating 34,709 individuals per skill-race group, a $150,000 windfall reduces incarceration risk over the following seven years by 0.81 percentage points. When the authors use model-generated property crime records rather than incarceration records as the dependent variable, they find a statistically significant effect more than twice as large, suggesting incarceration data systematically understates the crime-reducing effect of wealth shocks.&lt;/p&gt;
&lt;p&gt;Q: What is the mechanism by which lottery winnings have minimal persistent effects on crime?
A: Lottery winners in the model are disproportionately drawn from low-income, low-wealth individuals who also receive means-tested transfers. After winning, these individuals lose transfer eligibility, so winnings substitute for lost transfers rather than being invested. With income levels remaining low, winnings dissipate over time, individuals return to poverty, and resume criminal activity. Larger lottery prizes extend the crime-free interval but do not permanently alter behavior.&lt;/p&gt;
&lt;p&gt;Q: What is the role of life expectancy differences in racial wealth and crime gaps?
A: Equalizing survival probabilities generates a 25.39% reduction in the median wealth gap and a 12.4% reduction in the share of Black individuals in the bottom quintile, with virtually no change in crime rates. The channel operates through savings incentives: a shorter expected lifetime (73 years for Black versus 78 for White) reduces the return to wealth accumulation independently of income.&lt;/p&gt;
&lt;p&gt;Q: What are the fiscal resource requirements implied by the income equalization experiments?
A: Implementing equalized age-earnings profiles for the No-HS group would require resources equal to 3.29% of total GDP, while equalization for the HS group would require 22.2% of GDP. These figures reflect the scale of redistribution needed to close earnings profiles and serve as a benchmark for assessing policy feasibility.&lt;/p&gt;
&lt;p&gt;Q: How does incarceration scarring affect lifetime income in the benchmark, and how does this validate against external data?
A: A Black high school graduate who experiences at least one incarceration earns 16.8% less over his lifetime than one who is never incarcerated; for White high school graduates the gap is 28.7%. Gordon et al. (2023) report corresponding empirical estimates of 18.6% for Black and 32.7% for White high school graduates, closely validating the model&amp;rsquo;s scarring calibration.&lt;/p&gt;
&lt;p&gt;Endogenous property crime: A rational choice by working-age agents who weigh the expected gain from stealing (fraction γ = 6.4% of average labor income y) against the probability of apprehension (clearance rate πa = 17.2%), the loss of means-tested transfers, scarring of future labor income, and the minimum consumption floor in jail. Retired agents face no such choice.&lt;/p&gt;
&lt;p&gt;Exogenous drug incarceration: Incarceration for drug possession modeled as an exogenous shock with race-age-specific probabilities, not responsive to individual optimization, capturing the possibility that racial profiling in enforcement generates disparities in drug arrests independently of offending behavior.&lt;/p&gt;
&lt;p&gt;Scarring effect: Post-incarceration labor income penalty modeled as a higher probability of drawing a lower idiosyncratic income shock state upon labor market re-entry, calibrated so the model reproduces lifetime income gaps between ever-incarcerated and never-incarcerated individuals by race-skill group (18.6% for Black HS, 32.7% for White HS per Gordon et al. 2023).&lt;/p&gt;
&lt;p&gt;Age-earnings profile (ε^{i,ζ}_j): The deterministic, skill-race-age-specific component of labor income estimated from PSID data for each of six race-education groups. The gap between Black and White age-earnings profiles is identified as the dominant driver of both the racial wealth gap and racial crime disparities, accounting for the largest single-factor reduction in both outcomes across all counterfactual experiments.&lt;/p&gt;
&lt;p&gt;Means-tested transfer floor: A consumption support program that fills the gap between an agent&amp;rsquo;s post-tax income plus assets and a minimum threshold κ (5.8% of average net tax income and assets). This transfer is a critical mechanism linking wealth shocks to crime: lottery winnings and other wealth gains displace transfer eligibility, causing winnings to be consumed rather than saved, and eventually exhausted.&lt;/p&gt;
&lt;p&gt;Median wealth gap: The percentage difference between median White and median Black wealth — 89.5% in the 2019 SCF, 83.6% in the benchmark model — used as the primary scalar summary of racial wealth disparity, chosen because the model does not match the average wealth gap (model: 58.9%, data: 84.4%).&lt;/p&gt;
&lt;p&gt;Victimization probability (πv(Y)): A step-wise decreasing function of taxable income capturing spatial concentration of property crime in low-income neighborhoods; in equilibrium this equals the aggregate property crime rate χp, ensuring market clearing in the crime sector and implying that poorer agents face higher victimization risk.&lt;/p&gt;</description></item><item><title>Racial Disparities in Federal Sentencing: Evidence from Drug Mandatory Minimums</title><link>https://macropaperwarehouse.com/papers/racial-disparities-in-federal-sentencing-evidence-from-drug-mandatory-minimums/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/racial-disparities-in-federal-sentencing-evidence-from-drug-mandatory-minimums/</guid><description>&lt;p&gt;This paper studies racial disparities in federal criminal sentencing by analyzing abnormal bunching in the distribution of crack-cocaine amounts recorded at sentencing. The identifying variation comes from the Fair Sentencing Act (FSA) of 2010, which raised the 10-year mandatory minimum threshold for crack-cocaine from 50 grams to 280 grams. Because the new 280g threshold was set at a point with essentially zero pre-existing bunching, the author implements a difference-in-bunching design (following Kleven 2016) that compares the pre-2010 distribution of charged drug amounts — treated as the counterfactual — to the post-2010 distribution. The primary data are case-level records from the United States Sentencing Commission (USSC) covering all federal drug cases sentenced 1999–2015, restricted to crack-cocaine offenses (approximately 50,273 cases, of which 83.3% involve black defendants, 9.2% Hispanic, and 7.6% white).&lt;/p&gt;
&lt;p&gt;The main finding is that after 2010, the fraction of cases charged with amounts in the 280–290g range increases by 3.3 percentage points overall. This increase is disproportionately concentrated among minority defendants: black and Hispanic offenders are more than 2.5 times as likely as white offenders to be charged with 280–290g after the threshold shifts to that level. Approximately 80% of the excess mass at 280g is drawn from cases that had previously been charged in the 50–280g range, indicating that prosecutors are moving cases upward to cross the new threshold rather than negotiating downward from above it. For black and Hispanic offenders specifically, cases from the 50–280g range account for 88% of the increase at the new threshold.&lt;/p&gt;
&lt;p&gt;The author rules out differential drug involvement as an explanation. The pre-2010 distributions of charged amounts from 60–280g are nearly identical across racial groups; a Kolmogorov-Smirnov test fails to reject equality (p-value = 0.792). This implies the post-2010 racial disparity in bunching is a conditional disparity — arising not from differences in underlying drug involvement but from differential treatment of similarly situated defendants.&lt;/p&gt;
&lt;p&gt;The paper then traces the bunching to prosecutorial discretion specifically. Drug seizure records (NIBRS, DEA STRIDE), survey data on drug use and selling (NSDUH), and state-level conviction records from Florida all show no change in drug quantities or behaviors at the offender or law enforcement level coinciding with the FSA. Critically, there is no bunching at 280g in drug seizure data, pointing to decisions made after arrest. By contrast, case management files from the Executive Office of the US Attorney (EOUSA) show the fraction of cases recorded in the 280–290g range increases by 7.8 percentage points after 2010. Approximately 22–30% of prosecutors (depending on the detection method) are responsible for the rise in 280g cases. Bunching patterns persist across districts and mandatory minimum thresholds for the same prosecutors, indicating it reflects a prosecutor-level characteristic.&lt;/p&gt;
&lt;p&gt;The Supreme Court&amp;rsquo;s 5-4 decision in Alleyne v. United States (June 2013) raised the evidentiary standard for facts that trigger mandatory minimums and shifted that factual determination to juries. The share of EOUSA cases recorded in the 280–290g range fell from 9.1% (2011–2013) to 6.8% (2014–2016) after Alleyne, and a difference-in-discontinuities design confirms that bunching was partially reined in by this decision.&lt;/p&gt;
&lt;p&gt;On the question of discrimination, the racial disparity in bunching cannot be explained by observable defendant characteristics — education, sex, age, criminal history, seized drug amount, or other offense elements. Approximately 70% of the disparity persists after controlling for state-by-post fixed effects and 60% after district-by-post fixed effects. The disparity can be largely explained by a state-level measure of racial animus based on Google search data (Stephens-Davidowitz 2014): prosecutors operating in higher-animus states apply more disparate treatment, a pattern consistent with taste-based rather than statistical discrimination.&lt;/p&gt;
&lt;p&gt;Cases charged just above the 280g threshold receive longer sentences than those just below it in the post-2010 period, confirming that prosecutorial bunching has real consequences for sentence length.&lt;/p&gt;
&lt;p&gt;Q: What is the central empirical strategy of the paper?
A: The paper uses a difference-in-bunching design exploiting the Fair Sentencing Act of 2010, which shifted the 10-year mandatory minimum threshold for crack-cocaine from 50g to 280g. Because the 280g point had essentially zero bunching before 2010, the pre-2010 distribution of charged drug amounts serves as an empirical counterfactual for the post-2010 distribution absent the threshold change. The design allows the author to isolate bunching caused by the new threshold and to test whether that bunching is racially disparate.&lt;/p&gt;
&lt;p&gt;Q: What is the main quantitative finding on bunching?
A: After 2010, offenders sentenced for crack-cocaine are 3.3 percentage points more likely to be charged with amounts in the 280–290g range (Column 1, Table 2). Black and Hispanic offenders are more than 2.5 times as likely as white offenders to be charged with 280–290g after the threshold change (Column 2, Table 2). This racial gap is the central disparity the paper investigates.&lt;/p&gt;
&lt;p&gt;Q: Does the racial disparity in bunching reflect genuine differences in drug involvement?
A: No. The pre-2010 distributions of charged amounts from 60–280g are nearly identical across racial groups; a Kolmogorov-Smirnov test fails to reject equality with a p-value of 0.792. Because these pre-period distributions are taken as reflecting true drug involvement, their similarity by race implies the post-2010 disparity is a conditional racial disparity — arising from differential treatment of similarly situated defendants, not from differential drug involvement.&lt;/p&gt;
&lt;p&gt;Q: Where in the criminal justice process does the bunching originate?
A: The bunching originates in prosecutorial decisions, not at the arrest or law enforcement stage. Drug seizure records (NIBRS and DEA STRIDE) show no bunching at 280g, and survey data (NSDUH) show no post-FSA change in drug use or selling by minority defendants. Florida state-level records show no shift in the share of high drug-weight cases. By contrast, EOUSA case management files — which capture quantities recorded by prosecutors — show an increase of 7.8 percentage points in the fraction of cases in the 280–290g range after 2010.&lt;/p&gt;
&lt;p&gt;Q: What fraction of prosecutors engage in this bunching behavior?
A: Approximately 29.7% of prosecutors have a higher-than-normal percentage of cases at 280–290g after 2010 under a straightforward outlier criterion. Using the outlier detection procedure from Ridgeway and MacDonald (2009), approximately 22% are flagged as outliers. A Bayesian shrinkage method estimates approximately 30% (SE = 0.042) of prosecutors engage in this bunching. The behavior persists across districts and across multiple mandatory minimum thresholds for the same prosecutors, indicating it is a durable prosecutor-level characteristic.&lt;/p&gt;
&lt;p&gt;Q: What evidence links the bunching to upward manipulation rather than downward negotiation?
A: Approximately 80% of the excess mass at 280g is drawn from cases previously charged in the 50–280g range rather than from cases above 290g. For black and Hispanic offenders the share is 88%. This pattern indicates prosecutors are pushing amounts upward past the new threshold to secure longer sentences, not negotiating amounts downward from above the threshold — reversing the direction assumed in prior qualitative discussions.&lt;/p&gt;
&lt;p&gt;Q: What was the effect of Alleyne v. United States on bunching?
A: The Supreme Court&amp;rsquo;s 5-4 decision in Alleyne (June 2013) raised the evidentiary standard for facts triggering mandatory minimums and assigned those factual determinations to juries rather than judges. The share of EOUSA cases in the 280–290g range fell from 9.1% in 2011–2013 to 6.8% in 2014–2016. A difference-in-discontinuities design confirms that bunching expanded in the run-up to Alleyne and was partially curtailed afterward, providing additional evidence that the bunching reflects prosecutorial manipulation rather than genuine drug amounts.&lt;/p&gt;
&lt;p&gt;Q: Can observable defendant characteristics explain the racial disparity in bunching?
A: No. The racial disparity in bunching persists after controlling for education, sex, age, criminal history, seized drug amount, and other offense elements. Approximately 70% of the disparity remains after controlling for state-by-post fixed effects and 60% after controlling for district-by-post fixed effects. The disparity exists among observably similar defendants, ruling out the hypothesis that it is driven by correlated case characteristics.&lt;/p&gt;
&lt;p&gt;Q: What evidence distinguishes taste-based from statistical discrimination?
A: The racial disparity in bunching is largely explained by a state-level measure of racial animus constructed from Google search data (Stephens-Davidowitz 2014): prosecutors in higher-animus states apply more racially disparate treatment. Because statistical discrimination would predict disparate outcomes based on informative case characteristics rather than on the ambient racial attitudes of the jurisdiction, the correlation with racial animus is more consistent with taste-based discrimination than with statistical discrimination.&lt;/p&gt;
&lt;p&gt;Q: Does bunching at 280g have real consequences for sentence length?
A: Yes. Cases charged just above the 280g threshold receive longer sentences than those charged just below it in the post-2010 period, confirming that the mandatory minimum threshold is binding and that prosecutorial bunching translates into materially longer sentences for the affected defendants.&lt;/p&gt;
&lt;p&gt;Q: How does this paper contribute relative to Rehavi and Starr (2014)?
A: Rehavi and Starr (2014) linked arrest to sentencing records to show black offenders receive harsher sentences, driven by prosecutorial charging of mandatory minimums, but acknowledged that unobserved differences in criminal conduct within offense codes remained a concern. This paper addresses that concern by using the pre-2010 distribution of charged amounts as a counterfactual for drug involvement, documenting near-identical pre-period distributions by race, and tracing the post-FSA disparity through multiple data sources to isolate prosecutorial decisions specifically. The paper also quantifies the fraction of prosecutors involved and tests discrimination mechanisms.&lt;/p&gt;
&lt;p&gt;Q: What is the relationship between this paper&amp;rsquo;s findings and the policy goals of the Fair Sentencing Act?
A: The FSA achieved its stated goal of narrowing racial gaps attributable to the crack-powder disparity in mandatory minimum thresholds, and in line with prior work the author confirms a net decline in sentences after 2010. However, the increase in bunching at 280g by prosecutors — disproportionately applied to black and Hispanic defendants — dampened the FSA&amp;rsquo;s effectiveness. The paper thus documents a strategic response by a subset of prosecutors that partially offset the reform&amp;rsquo;s intended benefits for minority defendants.&lt;/p&gt;
&lt;p&gt;Q: How robust are the main bunching estimates?
A: The 3.3 percentage point overall increase and the 2.5x racial disparity are robust to various sample restrictions, inclusion of state fixed effects, time trends, state-specific time trends, offender-level controls, Logit/Probit/Poisson models, wider bunching range definitions (e.g., 280–380g), inclusion of cases with weights coded as a range, and alternative standard error calculations. Including range-coded cases actually exacerbates the estimated degree of bunching and the racial disparity.&lt;/p&gt;
&lt;p&gt;Bunching (in this paper&amp;rsquo;s sense): An excess mass of cases charged with a drug amount at or just above the mandatory minimum threshold, defined operationally as a disproportionate concentration of cases in the 280–290g range relative to the counterfactual distribution. Bunching reflects discretionary upward adjustment of charged amounts by prosecutors to trigger longer mandatory minimum sentences rather than true drug seizure quantities.&lt;/p&gt;
&lt;p&gt;Difference-in-bunching design: An empirical strategy adapted from Kleven (2016) that compares the actual post-2010 distribution of charged drug amounts to the pre-2010 distribution as a counterfactual for what the post-2010 distribution would have looked like absent the FSA threshold change. The method exploits the fact that the 280g threshold was a point of essentially zero bunching before 2010.&lt;/p&gt;
&lt;p&gt;Conditional racial disparity in bunching: A racial gap in the probability of being charged at 280–290g that remains after conditioning on similar underlying drug involvement, operationalized by the near-identical pre-2010 distributions of charged amounts from 60–280g across racial groups. The conditional disparity isolates differential treatment from differential conduct.&lt;/p&gt;
&lt;p&gt;Prosecutorial discretion (in this context): The legal authority of federal prosecutors to determine the drug quantity attributed to a defendant for sentencing purposes, which is not strictly bound to the amount physically seized at arrest. Prosecutors can rely on informant testimony, conspiracy attribution, or approximations to establish amounts above what was seized, giving them effective control over whether the mandatory minimum threshold is crossed.&lt;/p&gt;
&lt;p&gt;Taste-based discrimination: Racially disparate prosecutorial behavior that cannot be explained by observable case characteristics or informative statistical inference about defendant conduct, and that correlates instead with ambient state-level racial animus. In this paper&amp;rsquo;s framing, taste-based discrimination is distinguished from statistical discrimination by its correlation with the Stephens-Davidowitz racial animus measure rather than with defendant or offense characteristics.&lt;/p&gt;
&lt;p&gt;Mandatory minimum threshold (in federal crack-cocaine sentencing): A drug quantity cutoff — set at 50g before 2010 and 280g after the FSA — above which federal law mandates a sentence of at least 10 years unless specific departure conditions are met. The threshold creates a sharp discontinuity in expected sentence length that gives prosecutors an incentive to place cases just above it.&lt;/p&gt;
&lt;p&gt;State-level racial animus measure: A proxy for the prevalence of racially prejudiced attitudes in a state, constructed by Stephens-Davidowitz (2014) from Google Trends search volume data (2004–2007) for a specific racial slur and its plural, normalized by total search volume. Used here as a predictor of the size of the racial disparity in prosecutorial bunching across states.&lt;/p&gt;</description></item><item><title>Racial Disparities in Housing Returns</title><link>https://macropaperwarehouse.com/papers/racial-disparities-in-housing-returns/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/racial-disparities-in-housing-returns/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper estimates the racial/ethnic gap in realized housing returns using administrative data on individual housing transactions, and investigates the mechanisms that generate those gaps. The central question is: why do Black and Hispanic homeowners accumulate less housing wealth than White homeowners, even as minority homeownership rates have risen substantially over the last century?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors merge three primary data sources. First, a nationwide panel of residential property records from ATTOM covering 146.8 million arm&amp;rsquo;s-length home purchases from 1990 to 2020, which records transaction prices, mortgage characteristics, and property-level identifiers. Second, Home Mortgage Disclosure Act (HMDA) records, which contain self-reported race and ethnicity for mortgage applicants. Third, supplementary administrative sources including McDash mortgage servicing records, Equifax credit bureau data, Fannie Mae/Freddie Mac/ABSNet modification records, and the Survey of Income and Program Participation (SIPP). After applying sample restrictions — including requiring an observed purchase price, a linked HMDA record, an arm&amp;rsquo;s-length repeat sale, a combined loan-to-value ratio of at most 102.5%, and an ownership spell of at least 12 months — the baseline analysis sample comprises 13.6 million ownership spells for Black, Hispanic, and White homeowners who purchased homes with a mortgage between 1990 and 2016 in 40 states. Ownership spells unsold by March 2020 have their value imputed using the FHFA county-level house price index, a procedure that is conservative in that it understates racial gaps.&lt;/p&gt;
&lt;p&gt;The authors construct two complementary return measures. The &lt;strong&gt;unlevered return&lt;/strong&gt; compares the annualized ratio of sale price to purchase price. The &lt;strong&gt;levered return&lt;/strong&gt; (internal rate of return) sets the net present value of all homeowner cash flows — down payment, monthly mortgage payments, implicit rent, maintenance, taxes, insurance, transaction costs, and limited liability in foreclosure — equal to zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Among mortgaged home purchases, mean annual unlevered returns are 0.5% for Black homeowners, 0.6% for Hispanic homeowners, and 2.8% for White homeowners, implying Black-White and Hispanic-White gaps of approximately &lt;strong&gt;2.3 percentage points per year&lt;/strong&gt;. Mean annual levered returns are 1.6%, −3.0%, and 6.6% for Black, Hispanic, and White homeowners respectively, yielding gaps of &lt;strong&gt;5.0 and 9.6 percentage points&lt;/strong&gt;. After adjusting for the approximately one-fourth of purchases made in cash (for which no racial gap is found), preferred estimates of the unlevered gap are 1.9 (Black-White) and 1.4 (Hispanic-White) percentage points.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distressed sales — foreclosures and short sales — statistically account for the entire gap in returns.&lt;/strong&gt; Within non-distressed sales, the Black-White gap in annual unlevered returns falls to less than 40 basis points, and the Hispanic-White gap reverses sign. Two distinct factors drive the role of distressed sales: (1) Black and Hispanic homeowners are approximately &lt;strong&gt;twice as likely&lt;/strong&gt; as White homeowners to experience a distressed sale, and (2) minority homeowners live in neighborhoods where distressed sale price discounts are larger — estimated at 39%–40% for Black and Hispanic homeowners versus 28% for White homeowners. A Blinder-Oaxaca decomposition indicates that equalizing distressed sale rates (holding the distressed sale penalty fixed) would eliminate &lt;strong&gt;84.6%&lt;/strong&gt; of the Black-White unlevered returns gap and &lt;strong&gt;133.6%&lt;/strong&gt; of the Hispanic-White gap, confirming that the frequency margin dominates the severity margin.&lt;/p&gt;
&lt;p&gt;A counterfactual wealth-accumulation exercise using PSID data shows that &lt;strong&gt;equalizing housing returns reduces the Black-White gap in housing wealth at retirement by 37%&lt;/strong&gt;. Equalizing first-time purchase rates reduces the gap by only 1%, illustrating that promoting homeownership without addressing the returns gap is largely ineffective. Equalizing both returns and purchase rates reduces the gap by 49%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanisms&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Approximately one-third of the gap in unlevered returns can be explained by purchase year and county fixed effects, with much of this timing effect attributable to the Great Recession. Controlling additionally for income, family structure, gender, and leverage reduces the gap by a further ~0.3 percentage points, leaving a substantial residual. About half of the racial gap in mortgage default can be attributed to observable credit risk (family structure, income, leverage, credit score). The remainder is associated with &lt;strong&gt;unobservable liquidity shortfalls and income instability&lt;/strong&gt;: median liquid wealth among Black and Hispanic homeowners is $2,400 and $5,400 respectively, and minority homeowners are 2–4 percentage points more likely to transition to unemployment conditional on pre-unemployment income. Using quasi-experimental variation from adjustable-rate mortgage resets, the paper shows that in response to a 10% increase in monthly payments, White homeowners increase 90-day mortgage default by 3.0 percentage points after 12 months, while Black and Hispanic homeowners show increases of 4.5 and 7.1 percentage points respectively — excess sensitivity that is not captured by credit scores. The early-2000s credit supply expansion through private securitization and portfolio lending channels (as distinct from GSE/FHA) contributed to &lt;strong&gt;61.5%&lt;/strong&gt; of the 6.2-percentage-point increase in the Black-White distressed-sale gap between the 2002 and 2006 purchase cohorts, and &lt;strong&gt;52.0%&lt;/strong&gt; of the 12.2-percentage-point increase in the Hispanic-White gap. Evidence from the National Survey of Mortgage Originations suggests that Black homeowners hold overoptimistic expectations about future house price growth and income growth relative to their realized outcomes, which may explain why high-risk minority households do not self-select out of homeownership.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Results pertain to mortgaged home purchases (approximately three-fourths of all purchases) by Black, Hispanic, and White homeowners in 40 states (non-disclosure states excluded), with primary coverage from 2000 to 2016. No racial gap in returns is found for cash purchases. The racial gap in non-distressed returns is small and not economically meaningful, so the findings specifically pertain to the realized-return distribution that includes the distressed-sale tail.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-how-large-is-the-racial-gap-in-housing-returns-and-how-does-it-compare-to-previously-documented-racial-disparities-in-housing-costs"&gt;Q1. How large is the racial gap in housing returns, and how does it compare to previously documented racial disparities in housing costs?&lt;/h3&gt;
&lt;p&gt;A: Among mortgaged purchases, Black and Hispanic homeowners each realize annual unlevered returns approximately 2.3 percentage points lower than White homeowners; levered return gaps are 5.0 percentage points (Black-White) and 9.6 percentage points (Hispanic-White). In dollar terms, this translates to a difference of roughly $5,920 per year for the average Black homeowner and $6,762 per year for the average Hispanic homeowner on a ten-year holding horizon. These gaps are an order of magnitude larger than previously documented racial disparities in housing costs, such as post-origination interest rate disparities of about 40 basis points (~$500 annually for a $200,000 home) or inflated property tax assessments amounting to $300–$390 per year.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-role-of-distressed-sales-in-explaining-racial-gaps-in-returns-and-how-do-frequency-versus-severity-contribute"&gt;Q2. What is the role of distressed sales in explaining racial gaps in returns, and how do frequency versus severity contribute?&lt;/h3&gt;
&lt;p&gt;A: Distressed sales statistically account for nearly the entire racial gap in realized housing returns. Within non-distressed sales, the Black-White unlevered gap falls to less than 40 basis points and the Hispanic-White gap inverts. Two channels operate: (1) Black and Hispanic homeowners are approximately twice as likely as White homeowners to experience a distressed sale; and (2) within distressed sales, minority homeowners realize lower returns because they tend to live in neighborhoods with larger distressed-sale price discounts (estimated at 39–40% below imputed market value for Black and Hispanic homeowners, vs. 28% for White homeowners). A Blinder-Oaxaca decomposition indicates that equalizing distressed sale frequency (holding severity fixed) would close 84.6% of the Black-White gap and 133.6% of the Hispanic-White gap, so the frequency margin is quantitatively dominant.&lt;/p&gt;
&lt;h3 id="q3-are-racial-differences-in-house-price-appreciation-responsible-for-the-gap-in-non-distressed-returns"&gt;Q3. Are racial differences in house price appreciation responsible for the gap in non-distressed returns?&lt;/h3&gt;
&lt;p&gt;A: No. Among non-distressed sales, realized returns closely track county-level FHFA house price index growth for Black, Hispanic, and White homeowners alike, essentially one-for-one regardless of race. There is no economically meaningful racial gap in house price appreciation conditional on avoiding a distressed sale. This finding implies that the gap in average realized returns is not generated by differential neighborhood-level appreciation but rather by the incidence of distressed sales and the price penalties they entail.&lt;/p&gt;
&lt;h3 id="q4-how-much-of-the-racial-gap-in-housing-returns-can-be-explained-by-observable-homeowner-characteristics-such-as-income-family-structure-and-leverage"&gt;Q4. How much of the racial gap in housing returns can be explained by observable homeowner characteristics such as income, family structure, and leverage?&lt;/h3&gt;
&lt;p&gt;A: Controlling for county and purchase year fixed effects reduces the raw Black-White and Hispanic-White unlevered returns gaps from 2.3 to 1.5 and 1.6 percentage points, respectively. Additionally controlling for income, family structure (gender and co-applicant status), and leverage reduces the gap by a further ~0.3 percentage points. Even among the ostensibly safest group — high-income couples with low leverage — the Black-White (Hispanic-White) gap in unlevered returns is 0.7 (0.5) percentage points. Among high-leverage, low-income, single-male homeowners the gap is 1.8 (1.7) percentage points. Gaps exist within every demographic subgroup, and neighborhoods (Census tract fixed effects) explain roughly half of the remaining gap for Black homeowners and one-third for Hispanic homeowners, but substantial residual gaps persist even within neighborhood.&lt;/p&gt;
&lt;h3 id="q5-what-observable-credit-risk-characteristics-explain-racial-differences-in-mortgage-default"&gt;Q5. What observable credit risk characteristics explain racial differences in mortgage default?&lt;/h3&gt;
&lt;p&gt;A: Raw racial gaps in 90-day mortgage delinquency are 2.6 percentage points (Black-White) and 1.8 percentage points (Hispanic-White). Controlling for purchase year and county reduces these to 2.2 and 1.6 percentage points respectively. Controlling for family structure, income, leverage, and credit score reduces the gaps to 0.98 and 0.94 percentage points — implying that observable characteristics explain approximately 55% and 41% of the Black-White and Hispanic-White default gaps respectively. Credit scores contribute the most explanatory power among these controls, while mortgage contract characteristics (a test of differential lender treatment) contribute negligibly.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-evidence-that-liquidity-and-income-instability--factors-not-observable-to-lenders--explain-the-residual-racial-gap-in-default"&gt;Q6. What is the evidence that liquidity and income instability — factors not observable to lenders — explain the residual racial gap in default?&lt;/h3&gt;
&lt;p&gt;A: Survey data from SIPP reveal that median liquid wealth (bank accounts, stocks, bonds) for Black and Hispanic homeowners is only $2,400 and $5,400 respectively, while minority homeowners are 2–4 percentage points more likely to transition to unemployment conditional on pre-unemployment income. In SIPP mortgage delinquency regressions, controlling for liquidity, job loss in the prior year, and income reduces the Black-White coefficient by about 30% and the Hispanic-White coefficient by about 41% (and 29% and 70% respectively when also controlling for income level, current loan-to-value, and family composition). In administrative data using ARM payment resets as liquidity shocks, a 10% increase in monthly payments raises 90-day default by 3.0 percentage points for White homeowners, 4.5 percentage points for Black homeowners, and 7.1 percentage points for Hispanic homeowners after 12 months. This excess sensitivity is not substantially reduced by controlling for credit scores, income, or leverage — indicating that the liquidity risk of minority homeowners is largely unobservable to lenders at origination.&lt;/p&gt;
&lt;h3 id="q7-is-there-evidence-that-strategic-default-explains-higher-minority-distress-rates"&gt;Q7. Is there evidence that strategic default explains higher minority distress rates?&lt;/h3&gt;
&lt;p&gt;A: No meaningful evidence supports strategic default as a driver of excess minority distress. Using quasi-experimental variation in ex-post leverage from diverging option ARM indices (following Gupta and Hansman 2022), the paper finds large causal impacts of leverage on default but no evidence that these impacts are larger for minority homeowners. Separate survey evidence from the NSMO shows a statistically insignificant Black-White difference of 0.05 percentage points (s.e. 0.65) in agreement that &amp;ldquo;it is okay to default if it is in the borrower&amp;rsquo;s financial interest&amp;rdquo; (relative to a White mean of 6.1%). The absence of larger leverage-driven default responses combined with the presence of larger payment-shock-driven responses points specifically to liquidity — not strategic behavior — as the relevant mechanism.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-evidence-for-information-frictions-contributing-to-excess-minority-homeownership-risk"&gt;Q8. What is the evidence for information frictions contributing to excess minority homeownership risk?&lt;/h3&gt;
&lt;p&gt;A: Black homeowners in the NSMO report future house price expectations that are 0.07 standard deviations more optimistic than White homeowners, conditional on past price experiences, yet realized house price growth in the subsequent two years is actually 1.1 percentage points lower for Black homeowners. Although Black homeowners are 2.8 percentage points more likely to report past personal financial crises, their stated expectations about future financial crises are similar to those of White homeowners — despite 90-day default rates that are 2.5 percentage points higher in the first two years post-origination. Black homeowners also report income growth expectations 0.3 standard deviations higher than White homeowners, while SIPP and CPS data show minorities are more likely to experience income losses. These patterns of overoptimistic expectations relative to realized outcomes are consistent with information frictions causing high-risk minority households to suboptimally select into homeownership.&lt;/p&gt;
&lt;h3 id="q9-how-much-of-the-racial-gap-in-distress-can-be-attributed-to-the-early-2000s-credit-supply-expansion"&gt;Q9. How much of the racial gap in distress can be attributed to the early-2000s credit supply expansion?&lt;/h3&gt;
&lt;p&gt;A: The paper identifies the expansion as concentrated in portfolio loans and privately securitized mortgages, which are distinct from GSE/FHA mortgages that did not exhibit a comparable supply increase. Between the 2002 and 2006 purchase cohorts, the Black-White gap in distressed sales rose by 6.2 percentage points overall but only 2.4 percentage points among GSE/FHA loans. A decomposition using this contrast attributes 61.5% of the overall 6.2-percentage-point increase to the credit supply expansion. Analogously, 52.0% of the 12.2-percentage-point increase in the Hispanic-White gap between 2002 and 2006 is attributed to credit supply. Within-race decompositions find that credit supply accounts for 42%, 30%, and 35% of the increase in distress relative to 2002 for Black, Hispanic, and White homeowners respectively, for mortgages originated 2004–2006.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-implied-contribution-of-the-returns-gap-to-the-racial-wealth-gap"&gt;Q10. What is the implied contribution of the returns gap to the racial wealth gap?&lt;/h3&gt;
&lt;p&gt;A: Using a simple wealth accumulation model calibrated to PSID data on first-time homebuyer rates and home values (average first home for Black households: $142,587; for White households: $208,621), the paper finds an estimated Black-White gap in housing wealth at retirement of $169,389 versus an observed PSID gap of $182,771. Equalizing housing returns would reduce this gap by 37%. In contrast, equalizing first-time purchase rates alone reduces the gap by only about 1%, because low returns nullify the benefit of purchasing earlier. Equalizing both returns and purchase rates reduces the gap by 49%. Housing wealth in the primary home constitutes 43% of total net wealth for the average retirement-age Black household in PSID, implying the returns gap explains a quantitatively large share of the overall racial wealth gap.&lt;/p&gt;
&lt;h3 id="q11-what-do-the-covid-19-pandemic-forbearance-experience-and-mortgage-modification-evidence-imply-for-policy"&gt;Q11. What do the COVID-19 pandemic forbearance experience and mortgage modification evidence imply for policy?&lt;/h3&gt;
&lt;p&gt;A: Quasi-experimental estimates using servicer-level variation in modification propensity show that mortgage modifications cause economically large increases in housing returns for Black, Hispanic, and White homeowners alike, suggesting that since minority homeowners are more likely to become distressed, expanded modifications would disproportionately benefit them. The pandemic experience provides macroeconomic confirmation: after the onset of COVID-19 forbearance and foreclosure moratoria in March 2020, the Black-White gap in unlevered returns and distressed sales fell by approximately half, while the Hispanic-White gap (whose pre-pandemic distress convergence was already underway) remained comparatively stable. Administratively, Black homeowners who default are already 3–7 percentage points more likely than observationally similar White homeowners to receive a modification, even controlling for neighborhood and servicer, suggesting servicers partially internalize the larger distressed-sale discounts in minority neighborhoods.&lt;/p&gt;
&lt;h3 id="q12-are-neighborhood-level-factors--specifically-distressed-sale-price-discounts-from-illiquid-real-estate-markets--important-for-explaining-racial-heterogeneity-in-returns-conditional-on-distress"&gt;Q12. Are neighborhood-level factors — specifically distressed-sale price discounts from illiquid real estate markets — important for explaining racial heterogeneity in returns conditional on distress?&lt;/h3&gt;
&lt;p&gt;A: Yes. Using MLS data on median days-on-market as a measure of real estate market thickness, the paper shows that distressed sale discounts are substantially larger in less-liquid markets, with discounts experienced by Black homeowners approximately 13 percentage points lower in the least-thick markets relative to the thickest. Black and Hispanic homeowners are disproportionately likely to realize distressed sales in thin markets. Regular sale returns are not affected by market thickness. This establishes that neighborhood market illiquidity is a second-order channel through which neighborhood-level factors contribute to the racial gap — primarily by amplifying the severity of distressed sale penalties rather than by affecting ordinary house price appreciation.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Distressed sale&lt;/strong&gt;: In this paper&amp;rsquo;s usage, an ownership spell that ends in either a foreclosure (where a lender seizes and sells the property after payment default) or a short sale (where the lender allows the homeowner to sell for less than the outstanding mortgage balance without holding the homeowner liable for the deficiency). Distressed sales are the central mediating factor between race and housing returns.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Unlevered return&lt;/strong&gt;: The annualized ratio of sale price to purchase price, capturing property-level capital gains without reference to the financing structure. Computed as (P_sale / P_purchase)^(1/T) − 1. Does not capture leverage amplification or limited homeowner liability in foreclosure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Levered return (internal rate of return)&lt;/strong&gt;: The discount rate that sets the net present value of all homeowner cash flows to zero, including down payment at purchase; monthly payments (principal, interest, taxes, insurance, maintenance); implicit rent; and the net proceeds at sale (property sale price minus outstanding principal balance, subject to a floor of $0.01 capturing limited liability). This measure accounts for both the amplifying effect of leverage on gains and the homeowner&amp;rsquo;s limited liability in underwater foreclosures.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distressed sale frequency versus severity&lt;/strong&gt;: The two distinct components through which distressed sales generate racial gaps. Frequency refers to the higher probability that a minority homeowner&amp;rsquo;s ownership spell terminates in a distressed sale. Severity refers to the larger price discount at distressed sale that minority homeowners experience, concentrated in neighborhoods with illiquid real estate markets. The paper&amp;rsquo;s decomposition finds frequency is the dominant margin.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Unobservable liquidity risk&lt;/strong&gt;: Default risk arising from insufficient liquid wealth (cash, bank deposits, liquid securities) and income instability that is not captured by credit scores or other characteristics observable to lenders at mortgage origination. The paper&amp;rsquo;s ARM-reset event study shows this risk generates excess minority default responses even conditional on credit score and income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Information friction (overoptimism)&lt;/strong&gt;: The tendency of minority homeowners, particularly Black homeowners, to hold expectations about future house prices, personal financial crises, and income growth that are more optimistic than their realized outcomes and than observationally similar White homeowners&amp;rsquo; expectations. The paper uses this to explain why high-risk minority households do not self-select out of homeownership despite the high cost of distressed sales.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Credit supply channel&lt;/strong&gt;: The mechanism by which the early-2000s expansion of private securitization and portfolio lending — channels that exhibited substantially greater growth among Black and Hispanic borrowers than among White borrowers — contributed to increased rates of minority distress during the Great Recession. Distinguished from GSE/FHA channels that did not exhibit comparable credit expansion and serve as the counterfactual.&lt;/p&gt;</description></item><item><title>Rationing by Race</title><link>https://macropaperwarehouse.com/papers/rationing-by-race/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/rationing-by-race/</guid><description>&lt;p&gt;Singh and Venkataramani ask whether resource scarcity causes discriminatory rationing of health care by patient race, with patient death as the starkest possible outcome of biased allocation decisions. They examine 107,221 inpatient admissions from 2015 to 2018 at two large urban academic teaching hospitals (each with over 500 beds) in a Southeastern U.S. city with a sizable Black population. Black patients accounted for 60% of admissions, were on average younger (52 vs. 59 years), more likely to be female (65% vs. 50%), and had similar comorbidity burdens and baseline in-hospital death rates (approximately 2% for both groups), but waited over two hours longer on average for an inpatient bed and were 27% less likely to be admitted to the ICU.&lt;/p&gt;
&lt;p&gt;The authors exploit quasi-exogenous hour-to-hour variation in hospital capacity strain — measured as the share of inpatient beds occupied at the hour of a patient&amp;rsquo;s arrival — which clinical and qualitative literature establishes is difficult to predict even day-to-day. Capacity strain is coded in hospital-specific deciles (beds filled ranged from 69–78% in decile 1 to 91–95% in decile 10). The core regression interacts patient race with strain decile, controlling for hospital-specific hour-of-day, day-of-week, month-of-year, and year fixed effects; physician-of-record fixed effects; and a rich vector of patient characteristics including Elixhauser comorbidity indices, insurance status, and vital signs. Identification rests on the assumption that strain at the hour of arrival is conditionally independent of unobserved patient characteristics correlated with race — an assumption validated through balance tests on demographics, comorbidities, vital signs, machine-learning-derived admission themes, and selective discharge patterns.&lt;/p&gt;
&lt;p&gt;The main finding is that in-hospital mortality rises for Black patients but not for White patients as hospitals approach capacity. At the tenth decile of strain, Black patients face a mortality rate 0.7 percentage points higher than White patients — a 47.6% relative increase over the 1.47% White mortality rate at the same decile. A pooled difference-in-differences estimate implies that approximately 15% of Black patient deaths at high strain (decile 10) would not have occurred had Black patients faced the same strain-mortality relationship as White patients (coefficient 0.0052, p = 0.025). This pattern is concentrated among patients with the greatest ex ante medical need as measured by above-median Elixhauser mortality index scores (a score with AUC of 0.92 for predicting in-hospital mortality) and, in qualitatively similar but less precisely estimated form, by abnormal vital signs at arrival.&lt;/p&gt;
&lt;p&gt;The authors identify wait time for an inpatient bed as the primary mechanism. At all levels of capacity strain, high-need Black patients wait longer than low-need White patients — a pattern the authors characterize as a striking inversion of any need-based allocation principle. Racial disparities in wait times widen further at the highest decile of strain, exactly mirroring the mortality pattern. As an additional, more suggestive mechanism, the authors analyze free-text clinical documentation (the Reason for Admission field) using descriptive text features (time to completion, character count, average word length), sentiment analysis (subjectivity and polarity scores via TextBlob), and adjective counts. Documentation for Black patients exhibits features consistent with lower provider effort at all strain levels — shorter notes, less time deferred to completion — and subjectivity of notes and adjective counts diverge further by race at the highest strain decile, with White patients receiving increasingly detailed and descriptive notes as strain rises.&lt;/p&gt;
&lt;p&gt;The findings are robust across sparse models (age, gender, hospital fixed effects only) through fully saturated specifications (DRG fixed effects, interactions of all controls with race and strain), and to replacing Elixhauser index composites with their 31 individual comorbidity components. The authors explicitly scope their findings to a pre-COVID-19 period (2015–2018), while noting that pandemic-era record capacity strain and racial disparities in health outcomes suggest de facto race-based rationing may have been far more severe during COVID-19.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question and why is the health care setting chosen?
A: The paper asks whether increasing resource scarcity causes discriminatory rationing on the basis of race in consequential, high-stakes real-world decisions. Health care is chosen because it is high-stakes (patient death is the outcome), has a long documented history of racial discrimination at both provider and system levels, and offers uniquely detailed time-stamped electronic health record data that enables identification from hour-to-hour variation in capacity strain — a finer temporal resolution than most prior work.&lt;/p&gt;
&lt;p&gt;Q: How is hospital capacity strain measured and what is the identifying variation?
A: Strain is measured as the total number of patients occupying inpatient beds at the specific hour of a patient&amp;rsquo;s arrival, converted into hospital-specific deciles. The first decile corresponds to 69–78% of beds filled and the tenth decile to 91–95%. The identifying variation is residual hour-to-hour fluctuation in this measure after removing hospital-specific hour-of-day, day-of-week, month-of-year, and year fixed effects, which absorbs all predictable capacity patterns. Clinical and qualitative evidence establishes that even day-to-day strain is difficult to anticipate, making hour-to-hour residual variation plausibly as-if random.&lt;/p&gt;
&lt;p&gt;Q: What are the main mortality findings, and how large are the racial disparities at peak strain?
A: At the tenth decile of capacity strain, Black patients face a mortality rate 0.7 percentage points higher than White patients, representing a 47.6% relative increase over the 1.47% White mortality rate at that decile. The pooled difference-in-differences estimate (comparing decile 10 to deciles 1–9) implies that approximately 15% of Black patient deaths at high strain would not have occurred if Black patients had the same strain-mortality relationship as White patients (coefficient 0.0052, p = 0.025). White patient mortality does not increase at high strain; if anything, small (imprecisely estimated) decreases appear at deciles 7–9.&lt;/p&gt;
&lt;p&gt;Q: Which patients drive the racial mortality disparity?
A: The disparity is concentrated among patients with above-median Elixhauser mortality index scores — the ex ante sickest patients. The Elixhauser Mortality Index has a predictive AUC of 0.92 for in-hospital mortality. At decile 10, high-need Black patients experience a sharp increase in mortality not seen for high-need White patients or for low-need Black patients. Qualitatively similar but less precisely estimated results appear when acute need is measured by abnormal vital signs at arrival, with the difference that the triple interaction (race × strain × high-need vitals) is not statistically significant, consistent with vital signs being noisier proxies for severity than the Elixhauser indices.&lt;/p&gt;
&lt;p&gt;Q: How do the authors validate the identifying assumption that strain is conditionally independent of patient composition by race?
A: They document five types of supporting evidence: (i) the distribution of Black and White patients across hours of arrival and across strain deciles is nearly identical; (ii) regressions of patient demographics, all five Elixhauser comorbidity measures, and five vital signs abnormalities on race × strain interactions show no significant differential selection by race at different strain levels; (iii) machine-learning (Latent Dirichlet Allocation) topic themes from free-text admission notes change similarly by strain for Black and White patients; (iv) there is no evidence of selective discharge to hospice care by race and strain, with point estimates running counter to the hypothesis; and (v) strain is computed at time of arrival to the hospital rather than time of admission to an inpatient bed, preserving exogeneity.&lt;/p&gt;
&lt;p&gt;Q: What is the primary identified mechanism for the mortality finding?
A: Wait time for an inpatient bed is the primary mechanism. Black patients experience greater increases in wait times as strain rises compared to White patients, with the clearest divergence at decile 10 — exactly mirroring the mortality pattern. More strikingly, at every decile of strain (including decile 1, when beds are most abundant), high-need Black patients wait longer for a bed than low-need White patients, implying that the disparity is not solely a product of logistical constraints but reflects ingrained factors in clinical protocols, likely including implicit or explicit provider bias.&lt;/p&gt;
&lt;p&gt;Q: What does the wait time evidence reveal about the role of medical need vs. race in allocation decisions?
A: At lower strain levels, low-need patients appropriately wait longer than high-need patients. However, at higher strain levels (deciles 8–10) this need-based gap almost entirely disappears, while the racial gap in wait times persists. The gap between high-need Black and low-need White patients is larger than the gap between high-need and low-need patients of the same race, meaning race is a stronger predictor of wait times than medical need. This pattern is consistent with the paper&amp;rsquo;s conceptual framework in which increasing strain reduces providers&amp;rsquo; ability to accurately assess medical need while increasing the weight assigned to racial identity.&lt;/p&gt;
&lt;p&gt;Q: How is provider effort measured and what are the findings?
A: Provider effort is inferred from features of free-text Reason for Admission documentation: time to completion, character count, average word length, TextBlob subjectivity and polarity scores, and adjective counts. Across all strain levels, Black patients&amp;rsquo; documentation exhibits features consistent with lower effort — shorter completion times (providers less likely to defer documentation for clinical tasks), shorter notes with fewer characters and shorter words. At the highest strain decile, subjectivity scores for Black patients&amp;rsquo; notes increase relative to White patients&amp;rsquo; (driven by both rising Black and falling White subjectivity), and White patients receive more adjectives as strain rises while Black patients&amp;rsquo; adjective counts do not increase. Polarity scores remain stable by race and strain.&lt;/p&gt;
&lt;p&gt;Q: What do the documentation patterns suggest about compensatory behavior by providers?
A: The authors speculate that providers may anticipate reduced care quality at high strain and compensate by becoming more conscientious with White patients — writing longer, more detailed, more descriptive notes as strain increases, and potentially exerting greater care effort correlated with these documentation improvements. This protective compensatory behavior appears substantially less pronounced or absent for Black patients, which the authors suggest may translate into the small imprecisely estimated decrease in White patient mortality at higher strain deciles. They explicitly characterize this interpretation as speculative and requiring further investigation.&lt;/p&gt;
&lt;p&gt;Q: How robust are the main mortality findings to specification choices?
A: The mortality findings hold across: (i) sparse models with only age, gender, and hospital/year fixed effects; (ii) linear probability and logistic models; (iii) models with DRG fixed effects to compare within-diagnosis; (iv) models interacting all control variables with patient race and strain; (v) models replacing the Elixhauser composite index with its 31 individual comorbidity components; and (vi) models additionally controlling for five individual abnormal vital sign indicators. Results are substantively unchanged across all these specifications.&lt;/p&gt;
&lt;p&gt;Q: What additional care intensity measures are examined and what do they show?
A: The authors also examine ICU admission, ICU length of stay, total inpatient length of stay, and inpatient charges. They find no strain-related racial disparities on these margins. However, they note that unconditionally (across all strain levels), Black patients receive fewer resources on average — they are 27% less likely to be admitted to the ICU. The authors treat these care intensity measures as harder to interpret because both over- and under-provision can harm patients, and thus view them as less informative for their research question.&lt;/p&gt;
&lt;p&gt;Q: What conceptual framework guides the empirical predictions?
A: The framework models providers as assessing perceived medical need N&lt;em&gt;ij(t) = Ni × exp(−γ × S(t)), where the parameter γ captures the diminishing ability to accurately assess true need as strain S(t) rises. Simultaneously, the racial weight R&lt;/em&gt;ij(t) = Ri × φ(S(t)) increases with strain through the parameter φ(S(t)). When γ = 0 and φ = 0, allocation is race-neutral and need-based. When both parameters are positive, increasing strain simultaneously degrades need assessment and amplifies reliance on racial identity in allocation decisions — the paper&amp;rsquo;s core prediction, which is confirmed empirically.&lt;/p&gt;
&lt;p&gt;Q: How do the findings relate to the COVID-19 pandemic?
A: The data predate COVID-19 (2015–2018). The authors argue that pandemic conditions — record hospital capacity strain (especially in hospitals serving Black patients), extreme provider burnout, and documented racial disparities in health access — suggest race-based rationing may have been considerably more severe during COVID-19. The paper also contextualizes its findings within the pandemic-era debate over whether explicit race-based triage protocols were ethical or legal, arguing that de facto rationing by race appears to occur in ordinary care settings under typical stressors irrespective of that normative debate.&lt;/p&gt;
&lt;p&gt;Q: What policy interventions do the authors suggest?
A: The authors propose: increasing provider awareness of implicit biases; developing new algorithms to improve triage decisions for high-mortality-risk patients who might otherwise be overlooked; correcting existing care algorithms with documented racial bias; building provider peer networks to reduce biased treatment decisions; supporting patient self-advocacy; improving capacity prediction systems (as spurred by COVID-19); and creating load-shifting protocols and inter-hospital transfer networks to prevent resources from being stretched beyond capacity during high-strain periods.&lt;/p&gt;
&lt;p&gt;Capacity strain: The state of a hospital when a high share of inpatient beds are occupied, measured here at the hour of patient arrival as hospital-specific deciles of bed occupancy (ranging from 69–78% full at decile 1 to 91–95% full at decile 10); the paper&amp;rsquo;s primary measure of resource scarcity.&lt;/p&gt;
&lt;p&gt;Rationing by race: The paper&amp;rsquo;s term for the phenomenon whereby, as resource scarcity deepens, allocation decisions increasingly reflect patient racial identity rather than medical need — a form of discriminatory rationing that the authors distinguish from explicit (de jure) race-based triage and document as de facto practice.&lt;/p&gt;
&lt;p&gt;Perceived need (N*): In the paper&amp;rsquo;s conceptual framework, the provider&amp;rsquo;s assessment of a patient&amp;rsquo;s medical need, which deviates from true need Ni by the factor exp(−γ × S(t)) as strain S(t) increases; captures the provider team&amp;rsquo;s diminishing ability or willingness to accurately assess true medical need under cognitive and resource constraints.&lt;/p&gt;
&lt;p&gt;Racial weight (R*): The weight assigned to a patient&amp;rsquo;s racial identity in allocation decisions, modeled as Ri × φ(S(t)), where the function φ is increasing in capacity strain; represents the potential for discrimination — from implicit bias, algorithmic bias, reduced patient advocacy, or provider-patient social distance — to intensify as strain rises.&lt;/p&gt;
&lt;p&gt;Wait time inversion: The condition, documented throughout the paper, where high-need Black patients wait longer for an inpatient bed than low-need White patients at every decile of capacity strain, including decile 1 when resources are most abundant — inverting the normative principle that greater medical need should yield faster access to care.&lt;/p&gt;
&lt;p&gt;Elixhauser Mortality Index: A widely validated composite score of patient comorbid conditions used to predict in-hospital mortality (AUC = 0.92); used in this paper as the primary measure of chronic medical need, with patients split at the median into relatively sick (above median) and relatively healthy (below median) groups.&lt;/p&gt;
&lt;p&gt;Provider effort (inferred): An unobserved construct inferred in this paper from features of free-text clinical documentation in the Reason for Admission field, including time to note completion, character count, average word length, TextBlob subjectivity and polarity scores, and adjective counts; features argued to reflect how much attention, detail, and care a provider invested in documenting — and by extension, in assessing — a patient&amp;rsquo;s condition.&lt;/p&gt;</description></item><item><title>Religion, Education, and the State</title><link>https://macropaperwarehouse.com/papers/religion-education-and-the-state/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/religion-education-and-the-state/</guid><description>&lt;p&gt;This paper studies how Indonesia&amp;rsquo;s Islamic education sector responded to one of the largest state-driven mass schooling expansions in history — the SD INPRES program (Sekolah Dasar Presidential Instruction) launched in 1973 — and whether that program achieved its secular nation-building objectives. The research question is three-part: Did Islamic schools enter or exit markets where the state built more primary schools? How did religious school choice shift across cohorts? And did the program advance secular identity formation among exposed individuals?&lt;/p&gt;
&lt;p&gt;The empirical setting is Indonesia in the 1970s onward. Under SD INPRES, the government used windfall oil revenues to build more than 61,000 primary schools between 1973 and 1980, allocating construction across districts proportional to the non-enrolled primary-school-age population. Because Islamic schools were historically more prevalent in underserved areas, this rule produced a strong positive correlation between SD INPRES intensity and pre-existing Islamic school density — the same markets where the state expanded were precisely those with the greatest Islamic education presence.&lt;/p&gt;
&lt;p&gt;The authors use several novel data sources: administrative registries covering nearly 220,000 secular and 160,000 Islamic schools with establishment dates; six rounds of the National Socioeconomic Survey (Susenas) from 2012–18; the Indonesia Family Life Survey (IFLS, 1993–2014); a 2018–19 curriculum timetable registry (SIAP) covering nearly 20% of madrasa; and a 2016 political/religious attitudes survey. Identification relies on difference-in-differences (DID) exploiting cross-district variation in SD INPRES intensity, the synthetic DID approach of Arkhangelsky et al. (2021) for robustness to violations of parallel trends, and a staggered village-level event study using the Borusyak et al. (2024) estimator.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. First, Islamic schools did not exit markets where the state expanded — they entered in greater numbers. A one standard deviation increase in SD INPRES construction led to 1.4 additional Islamic school entries per district above a mean of 1.9 per district in 1972. Entry was competitive at the primary level, where new madrasa (MI) entered at twice the baseline annual rate in the years immediately following INPRES construction, and strategic at the secondary level, where Islamic junior secondary schools (MTs) peaked 6–9 years after INPRES entry as graduates sought continued education. The Islamic sector financed this expansion through waqf (inalienable religious endowments), informal taxation (infaq, zakat), and revenues from a concurrent rice price spike; entry responses were stronger in villages with above-median waqf endowments and above-median potential rice yields.&lt;/p&gt;
&lt;p&gt;Second, rather than converging toward secular curricula, newly established Islamic schools in high-INPRES districts devoted more time to religious content. Each additional SD INPRES is associated with a 1.2 percentage point increase in the religious curriculum share among newly created madrasa, with increases of 1.3 and 2.4 percentage points at the primary and junior secondary levels respectively — the latter equaling 82% of the cross-school standard deviation. Some of this increase came at the expense of Pancasila/civic education and national language instruction.&lt;/p&gt;
&lt;p&gt;Third, while SD INPRES reduced Islamic primary school enrollment by roughly 7%, it increased overall Islamic school attendance: each additional SD INPRES increased the likelihood of attending any Islamic school by approximately 5%, as demand for secondary education outweighed substitution at the primary level. Female students exhibited stronger secondary-level demand effects, amplified in districts with a concurrent state ban on the Islamic veil in public schools.&lt;/p&gt;
&lt;p&gt;Fourth, SD INPRES did not advance its ideological objectives. In the 1977 and 1982 elections, Golkar&amp;rsquo;s vote share fell and the Islamic PPP&amp;rsquo;s rose by 0.5–1.0 percentage points per SD INPRES school in high-INPRES districts. Among exposed cohorts, SD INPRES did not increase Pancasila proficiency, national language use at home, or support for secular governance, but did increase Arabic literacy by approximately 3% per additional SD INPRES. Exposed cohorts also prayed more frequently, fasted more during Ramadan, gave more to charity, and expressed greater pilgrimage intentions. These religious patterns were transmitted to children of exposed cohorts, who were more likely to attend Islamic schools themselves.&lt;/p&gt;
&lt;p&gt;Q: What was the allocation rule for SD INPRES and why did it create confrontation with Islamic schools?
A: Presidential Instruction No. 10/1973 allocated school construction across districts proportional to the non-enrolled primary-school-age population in 1971. Because Islamic schools historically served underserved populations, this rule meant the state built more schools precisely where Islamic education was most prevalent. The paper shows graphically and in Table 1 that the number of SD INPRES schools built is strongly correlated with the pre-existing stock of Islamic schools, conditional on district population and enrollment.&lt;/p&gt;
&lt;p&gt;Q: How large was the Islamic sector&amp;rsquo;s entry response to SD INPRES at the district level?
A: In the standard DID specification (Table 2, panel a), a one standard deviation increase in SD INPRES schools led to 0.013 more Islamic schools per district-year per 1,000 children, equivalent to 1.4 additional Islamic school entries in the average district relative to a mean of 1.9 Islamic schools per district in 1972. The synthetic DID (panel b) delivers positive and slightly larger estimates, indicating the result is not an artifact of diverging pre-trends.&lt;/p&gt;
&lt;p&gt;Q: What was the timing of the Islamic sector entry response at the village level?
A: Using the Borusyak et al. (2024) estimator on a balanced panel from 1960 to 1999, the paper finds (Figure 4) that INPRES construction is followed by a jump in Islamic school entry. Primary madrasa (MI) entered at twice the baseline annual rate in the years immediately following INPRES construction and this elevated rate persisted for six years before reverting to baseline. Islamic junior secondary entry (MTs) peaked around years 6–9 after SD INPRES construction, consistent with newly graduated primary students seeking continued schooling.&lt;/p&gt;
&lt;p&gt;Q: How did the Islamic sector finance its expansion?
A: The sector relied on waqf endowments (inalienable religious land assets), informal faith-based contributions (infaq), and obligatory alms (zakat). Fortuitously, the initial year of SD INPRES coincided with a large spike in the global price of rice, Indonesia&amp;rsquo;s main agricultural commodity, boosting harvest revenues channeled through informal Islamic taxation. Table 3 shows that entry responses were significantly stronger in villages with above-median waqf endowments and above-median potential rice yields, and these heterogeneous effects did not arise in non-INPRES periods or for non-Islamic private schools. Survey data from 2007–13 further show higher rates of informal taxation in villages with Islamic schools built during this period.&lt;/p&gt;
&lt;p&gt;Q: Did Islamic schools converge toward secular curricula under competitive pressure from SD INPRES?
A: No. Table 4 shows that madrasa established in high-INPRES districts after 1972 devote more time to religious content, not less. Each additional SD INPRES is associated with a 1.2 percentage point increase in the share of classroom time devoted to religious subjects among newly created Islamic schools, with increases of 1.3 percentage points at the primary level and 2.4 percentage points at the junior secondary level — the latter equal to 82% of the cross-school standard deviation. Similar patterns hold for Arabic instruction, and the junior secondary increase comes partially at the expense of Pancasila/civic education and national language instruction.&lt;/p&gt;
&lt;p&gt;Q: Did curriculum differentiation responses vary with local religious ideology?
A: Yes. Appendix Table A.14 shows a stronger curriculum differentiation response in markets with greater historical support for conservative Islam, proxied by Islamic political party vote shares in the 1950s elections. The paper also constructs a school-name-based predicted ideology index using a ridge shrinkage estimator and finds (Appendix Table A.15) that madrasa entering high-INPRES districts after the program onset have a more religious ideology on this measure.&lt;/p&gt;
&lt;p&gt;Q: What happened to the formalization of the Islamic sector?
A: Figure 5 and Appendix Table A.6 show that formal madrasa entry increased as a share of all new school entry, while informal Islamic schools (pesantren, diniyah) declined as a share of all new schools and all new Islamic schools. This formalization mirrors the organizational structure of state schools (primary-to-secondary progression), facilitating switching between public and religious schools and providing option value to moderate but still religious families. Crucially, the newly entering formal madrasa introduced more religious curriculum than incumbent madrasa, so formalization did not reduce religious instruction.&lt;/p&gt;
&lt;p&gt;Q: What was the net effect of SD INPRES on Islamic school attendance?
A: Table 5 shows that SD INPRES reduced the likelihood of attending Islamic primary school by roughly 7% per additional SD INPRES school but increased Islamic secondary attendance, with the net effect being a roughly 5% increase in the likelihood of attending any Islamic school (column 4). This finding holds in both DID and synthetic DID. The IFLS validation (Appendix Table A.18) confirms decreased Islamic elementary attendance and increased Islamic junior secondary attendance, consistent with the Susenas results.&lt;/p&gt;
&lt;p&gt;Q: How does selection into secondary education affect the religious schooling results?
A: The authors address selection using parametric (Heckman 1976) and semiparametric (Newey 2009) selection-correction procedures, using exposure to a 1960s pilot compulsory schooling program as an exclusion restriction. Table 6, panels (c) and (d), show that selection-adjusted estimates are broadly consistent with unadjusted estimates, with similar signs and magnitudes. The selection-corrected estimates approximately identify a local average treatment effect among compliers: those induced to attend elementary school were less likely to attend Islamic elementary; those induced to continue to secondary were more likely to attend Islamic secondary.&lt;/p&gt;
&lt;p&gt;Q: How did gender shape the effects of SD INPRES on religious school choice?
A: Table 7 shows that SD INPRES had more limited impacts on total schooling for women than men (consistent with Duflo 2001) but that the secondary-level demand effect toward Islamic schools was stronger for women. Table 8 shows that within high-INPRES areas, the SD INPRES-induced increase in Islamic secondary education is three times larger for women in districts with greater exposure to the 1982 state ban on the Islamic veil in public schools, and this differential is specific to Islamic schooling rather than total schooling.&lt;/p&gt;
&lt;p&gt;Q: Did SD INPRES strengthen or weaken the secular ruling regime&amp;rsquo;s political standing?
A: It weakened it. Table 10 shows that in the 1977 and 1982 elections, Golkar&amp;rsquo;s vote share decreased and the Islamic PPP&amp;rsquo;s vote share increased in high-INPRES districts, in the range of 0.5–1.0 percentage points per SD INPRES school. This represents a 1.5–3.0% change in PPP vote share and a 0.5–1.0% change in Golkar vote share per standard deviation in SD INPRES intensity. The PPP gained most in areas where SD INPRES had the greatest potential to draw students away from Islamic schools.&lt;/p&gt;
&lt;p&gt;Q: Did SD INPRES produce a secular ideological shift among exposed cohorts?
A: No. Table 11 shows that SD INPRES did not increase self-reported Pancasila proficiency, national language use at home, national language literacy, or attitudes in favor of secular governance. By contrast, Arabic literacy increased by approximately 3% per additional SD INPRES among exposed cohorts, indicating that Islamic schooling exposure rather than secular schooling drove literacy gains in that language.&lt;/p&gt;
&lt;p&gt;Q: Did SD INPRES increase religiosity among exposed cohorts?
A: Yes. Table 12 shows that SD INPRES increased prayer frequency, fasting during Ramadan, charitable giving, and pilgrimage intentions among exposed cohorts. These effects on prayer and fasting are stronger among women, consistent with the stronger shift toward Islamic secondary schooling found in Table 7. These outcomes are consistent with greater exposure to Islamic education increasing religiosity rather than the secular curriculum reducing it.&lt;/p&gt;
&lt;p&gt;Q: Were the effects on religious identity and Arabic literacy transmitted to the next generation?
A: Yes. Table 13 shows that SD INPRES increased Arabic literacy among the children of exposed cohorts, and that children of exposed cohorts were more likely to attend Islamic schools themselves. These intergenerational results confirm that the preference for Islamic education instilled during the SD INPRES era persisted into the next generation rather than converging toward secular norms over time.&lt;/p&gt;
&lt;p&gt;Q: What are the key robustness checks for the school entry results?
A: Several checks support causal interpretation. The Roth and Rambachan (2022) procedure finds no systematic pre-trends in the standard DID. Historical Podes data from 1980, 1983, 1990, and 1993 confirm the post-1973 increase in Islamic school entry, addressing survival bias in the 2019 registry. Results are robust to allowing differential trends in waqf endowments, Muslim population share, Islamic party vote shares, historical Arab immigration, Islamist insurgency, and Transmigration resettlement. The heterogeneous entry responses by waqf and rice yield do not appear in non-INPRES periods or for non-Islamic schools.&lt;/p&gt;
&lt;p&gt;Q: What do the results imply for the political economy of education reform more broadly?
A: The paper argues that state capacity to homogenize culture through education is limited when strong non-state actors can mobilize their own resources and provide differentiated alternatives. Rather than crowding out religious schools, state expansion triggered competitive entry, curriculum differentiation, and formalization in the religious sector, producing an equilibrium where both sectors expanded simultaneously with distinct clienteles. The findings imply that the long-run cultural effects of education programs cannot be evaluated without accounting for equilibrium responses by competing non-state providers.&lt;/p&gt;
&lt;p&gt;SD INPRES (Sekolah Dasar Presidential Instruction): Indonesia&amp;rsquo;s 1973 mass public primary school construction program, financed by oil windfalls, which built more than 61,000 elementary schools between 1973 and 1980 by allocating schools to districts proportional to the non-enrolled child population; the program&amp;rsquo;s explicitly secular nation-building objectives brought it into direct confrontation with the Islamic education sector.&lt;/p&gt;
&lt;p&gt;Waqf: Inalienable Islamic religious endowments — of land, agricultural assets, or other property — that under Islamic law can only be used for religious or charitable purposes and cannot be seized or repurposed by the state; in this paper, the pre-existing waqf base in a village serves both as a long-run financing mechanism for Islamic school construction and as an index of Islamic sector organizational capacity.&lt;/p&gt;
&lt;p&gt;Madrasa: Formal day Islamic schools operating at the same primary-to-secondary grade levels as secular state schools, teaching standard academic subjects alongside a religious curriculum (including Islamic law, doctrine, ethics, Qur&amp;rsquo;an, Arabic, and history of the Prophets) that averages 26% of total instruction hours; distinct from the more informal pesantren (boarding schools) and madrasa diniyah (afternoon Qur&amp;rsquo;anic study schools).&lt;/p&gt;
&lt;p&gt;Curriculum differentiation: The strategy by which newly entering madrasa in high-INPRES districts increased the share of classroom time devoted to religious and Arabic instruction rather than converging toward the secular state curriculum; measured as classroom hours devoted to Islamic subjects, Arabic, Pancasila/civic education, and national language instruction from 2018–19 SIAP timetable data.&lt;/p&gt;
&lt;p&gt;Pancasila: The official secular nationalist ideology of the Indonesian state, consisting of five principles (monotheism, humanitarianism, national unity, democracy, and social justice) intended to transcend ethnic and religious divisions; SD INPRES sought to transmit Pancasila through civic education and national language instruction as part of its homogenizing nation-building agenda.&lt;/p&gt;
&lt;p&gt;Synthetic difference-in-differences (SDID): The Arkhangelsky et al. (2021) estimator used throughout the paper, which reweights and matches pre-INPRES trends in Islamic school construction across high- and low-INPRES exposure districts to deliver estimates more robust than standard DID to violations of parallel trends; applied with a binary treatment indicator (districts above the 51st percentile in INPRES intensity).&lt;/p&gt;
&lt;p&gt;Formalization: The documented shift in the composition of the Islamic sector after SD INPRES, whereby formal madrasa (organized along the same grade-level progression as state schools) increased as a share of all new Islamic school entry while informal pesantren and diniyah declined as a share; interpreted as a competitive response that expanded parental option value without sacrificing religious instruction intensity.&lt;/p&gt;</description></item><item><title>Republican Support and Economic Hardship: The Enduring Effects of the Opioid Epidemic</title><link>https://macropaperwarehouse.com/papers/republican-support-and-economic-hardship-the-enduring-effects-of-the-opioid-epidemic/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/republican-support-and-economic-hardship-the-enduring-effects-of-the-opioid-epidemic/</guid><description>&lt;p&gt;This paper establishes a causal connection between the opioid epidemic and the political realignment toward the Republican Party in the United States from the mid-2000s through 2022. The authors—Carolina Arteaga and Victoria Barone—exploit rich geographic variation in Purdue Pharma&amp;rsquo;s initial marketing strategy for OxyContin, drawn from unsealed litigation records, to construct a quasi-exogenous measure of community-level exposure to the epidemic.&lt;/p&gt;
&lt;p&gt;The identification strategy rests on a documented feature of OxyContin&amp;rsquo;s 1996 launch: Purdue initially targeted the established cancer pain market—physicians and patients already using MS Contin—as an entry point into the much larger noncancer pain market. Areas with higher cancer mortality in 1996 received disproportionate pharmaceutical marketing, leading to outsized opioid prescription growth that spilled over from cancer patients to the broader population through shared physicians. The authors use 1996 commuting-zone (CZ) cancer mortality rates as a proxy for this initial targeting, interacted with year fixed effects in an event-study specification with CZ and state-year fixed effects. The sample covers 625 CZs across the continental United States from 1982 to 2022.&lt;/p&gt;
&lt;p&gt;The empirical chain runs through three stages. First, the instrument strongly predicts opioid supply: by 2012, a one-standard-deviation higher 1996 cancer mortality rate led to an additional 0.97 opioid doses prescribed per capita, 65% above the baseline mean. Second, the resulting epidemic caused measurable mortality and economic hardship. A one-standard-deviation increase in 1996 cancer mortality caused drug-induced deaths in 2017 to be 46% above the pre-epidemic average; by 2012 the same increase caused prescription opioid deaths to be 61% higher. Excess mortality was concentrated among individuals under age 55, with no significant effects for those aged 55 and older. The epidemic also raised disability applications: SSDI applications rose by 12% and SSI applications by 7.6% by 2012, effects that persisted through 2020. SNAP enrollment in exposed CZs was 8% higher by 2022, equivalent to a 0.14 standard deviation increase.&lt;/p&gt;
&lt;p&gt;Third, and centrally, the communities that endured these health and economic shocks shifted persistently toward the Republican Party. By the 2022 House elections, a one-standard-deviation increase in 1996 cancer mortality increased the Republican two-party vote share by 4.5 percentage points. Effects of similar magnitude appear in presidential elections (4.6 percentage points) and gubernatorial elections (4.3 percentage points). The vote-share shift is consistent across gender, age, race, and education, with no detectable change in voter turnout. The shift translates into actual seat gains: beginning in 2012, exposed areas consistently elected more Republican House members, moving the chamber&amp;rsquo;s roll-call voting in a more conservative direction. The effect is not driven by anti-incumbent sentiment—results hold regardless of which party held the seat at the time.&lt;/p&gt;
&lt;p&gt;The paper identifies three reinforcing mechanisms. First, the Republican Party repositioned itself during this period as the advocate of &amp;ldquo;forgotten America&amp;rdquo; and working-class economic hardship, a message that resonated acutely in opioid-devastated communities. Second, conservative-leaning newspapers covered the epidemic at higher rates, and their coverage tracked local mortality; liberal-leaning outlets showed no such correlation. Fox News covered opioid stories at 1.5 times the rate of CNN and 1.7 times the rate of MSNBC, emphasizing crime, trafficking, and cartels at twice the frequency of liberal outlets. Third, exposed communities expressed stronger preferences for Republican-favored policy responses: higher police presence, greater sense of safety around law enforcement, and lower support for marijuana legalization on state ballot initiatives.&lt;/p&gt;
&lt;p&gt;Pre-trend tests show no relationship between 1996 cancer mortality and outcomes before OxyContin&amp;rsquo;s launch. Out-of-sample exercises using 1976 cancer mortality find no analogous pattern in the pre-epidemic period (1982–1994). Placebo instruments based on unrelated causes of death yield null results. The baseline findings are robust to controlling for the China import shock, NAFTA, the 1994 Republican Revolution, the 2001 and 2008–2009 recessions, declining unionization, robot adoption, Fox News introduction, deaths of despair, and Southern and rural political realignment.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s central research question?
A: The paper asks whether the opioid epidemic causally increased Republican vote share in communities most severely affected by the crisis. It documents a causal chain from pharmaceutical marketing through drug mortality and economic hardship to political realignment, contributing the first causal estimate of a major public health crisis&amp;rsquo;s effect on partisan voting.&lt;/p&gt;
&lt;p&gt;Q: What is the identification strategy, and why is 1996 cancer mortality a valid instrument?
A: Purdue Pharma explicitly targeted physicians in the cancer pain market at OxyContin&amp;rsquo;s 1996 launch, then used those established relationships to expand into the noncancer pain market. CZs with higher cancer mortality in 1996 received disproportionate marketing, generating differential opioid prescription growth unrelated to pre-existing political or economic trends. Pre-trend tests confirm no differential patterns before 1996, out-of-sample tests using 1976 cancer mortality find no relationship with pre-epidemic outcomes, and placebos using unrelated causes of death yield null results.&lt;/p&gt;
&lt;p&gt;Q: How strong is the first stage linking 1996 cancer mortality to opioid prescriptions?
A: The relationship between 1996 cancer mortality and opioid prescriptions is positive and statistically significant from 1998 through 2020. By 2012—the year prescription rates peaked nationally at 81.3 per 100 persons—a one-standard-deviation higher cancer mortality rate led to an additional 0.97 morphine-equivalent doses prescribed per capita, 65% above the baseline mean. CZs in the highest cancer mortality quartile experienced a 1,800% increase in grams of oxycodone per capita between 1997 and 2010, compared to less than half that in the lowest quartile.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on drug-induced mortality?
A: Drug-induced mortality (a broad measure covering deaths from prescription opioids, heroin, and fentanyl) rose continuously in exposed CZs after 1996. By 2017, a one-standard-deviation increase in 1996 cancer mortality caused drug-induced deaths to be 46% above the pre-epidemic average. By 2012, the same increase caused prescription opioid deaths specifically to be 61% higher relative to the pre-epidemic average. Excess mortality was concentrated among individuals under age 55, with no statistically significant effects for those aged 55 and older.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on disability program take-up?
A: Applications for Social Security Disability Insurance (SSDI) rose by 12% and Supplemental Security Income (SSI) applications rose by 7.6% by 2012 for a one-standard-deviation increase in 1996 cancer mortality. These effects persisted: SSDI recipients grew by 15% and SSI recipients by 3.2% by 2020 in similarly exposed CZs. The increases in disability were concentrated among individuals under age 55, paralleling the mortality effects.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on SNAP enrollment?
A: Exposed CZs showed a continuous increase in SNAP enrollment over two decades following the epidemic&amp;rsquo;s onset. By 2020, a one-standard-deviation increase in 1996 cancer mortality corresponded to an 8% increase in the share of the population receiving SNAP benefits, equivalent to 0.14 standard deviations. By 2022, the corresponding figure remains 8%, indicating persistent economic strain in exposed communities.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the political effects in House elections?
A: A one-unit increase in the 1996 cancer mortality rate yielded a 7.9-percentage-point increase in the 2022 Republican House vote share relative to 1996. Scaled to one standard deviation (0.58 units), this corresponds to a 4.5-percentage-point increase in the Republican two-party vote share by the 2022 midterms. The vote-share shift became statistically significant beginning in 2006, but only translated into consistent seat-level Republican gains starting in 2012.&lt;/p&gt;
&lt;p&gt;Q: When did opioid exposure start winning Republicans additional House seats?
A: Although the Republican vote share in exposed areas began increasing around 2006, actual seat flips did not become consistent until 2012. The paper explains this lag by noting that initial vote-share gains were concentrated in communities with low baseline Republican support, where additional votes did not immediately cross the winning threshold. Starting in 2010, median-baseline-Republican CZs also began shifting, enabling additional seat changes.&lt;/p&gt;
&lt;p&gt;Q: How large are the presidential and gubernatorial election effects?
A: In presidential elections, a one-standard-deviation increase in 1996 cancer mortality raised the Republican vote share by 4.6 percentage points. In gubernatorial elections, the same increase raised the Republican vote share by 4.3 percentage points after approximately six election cycles (corresponding to 2017–2020). These effects are described as comparable in magnitude to the difference in Republican vote share between the top and bottom quartiles of NAFTA vulnerability.&lt;/p&gt;
&lt;p&gt;Q: Does the political shift reflect increased polarization toward extremist candidates?
A: No. The paper finds no increase in the probability of electing candidates at the extremes of the Nokken-Poole ideological scale in any given election year. The ideological shift in the House composition arises from changes in which party wins seats rather than from the election of more extreme Republicans. Campaign donations to Republican candidates did not increase; rather, donations to Democratic candidates declined (in 2016, a one-standard-deviation increase in cancer mortality widened the Republican-Democrat donation gap by 0.44 standard deviations). The shift is interpreted as a change in voting preferences in previously Democratic-leaning areas, not heightened polarization.&lt;/p&gt;
&lt;p&gt;Q: Is the shift driven by anti-incumbent sentiment?
A: The authors test this by splitting the sample by the incumbent&amp;rsquo;s party at the time of each election and by redefining the outcome as the incumbent&amp;rsquo;s vote share. Neither exercise produces evidence of a systematic anti-incumbent response. The changes in Republican vote share are not statistically distinguishable based on whether the incumbent was a Republican or Democrat. If anything, after 2016 there is a slight increase in the likelihood of incumbents retaining their seats.&lt;/p&gt;
&lt;p&gt;Q: Where geographically are the Republican gains largest?
A: Using state-level treatment effects estimated from an in-differences model interacting cancer mortality with state-year indicators, the paper finds a strong positive correlation between the magnitude of the epidemic&amp;rsquo;s effect on economic hardship (measured by SNAP participation) and the magnitude of the Republican vote-share increase. This correlation is strongest with a lag: SNAP effects measured in 2006 are most predictive of vote-share shifts in 2022, indicating that deterioration in community economic fabric preceded and predicted the political realignment.&lt;/p&gt;
&lt;p&gt;Q: How did conservative and liberal media differ in covering the opioid epidemic?
A: Republican-leaning local newspapers covered the opioid epidemic more extensively than Democratic-leaning papers throughout the epidemic period, and their coverage tracked local opioid mortality rates; Democratic-leaning coverage showed no such correlation with local incidence. Fox News covered opioid stories at 1.5 times the rate of CNN and 1.7 times the rate of MSNBC. In terms of content, Republican-leaning newspapers showed 23% higher frequency of economic hardship keywords, 19% higher frequency of illegal activity and crime keywords, and 22% higher frequency of rehabilitation and treatment keywords relative to Democratic-leaning papers. Fox News emphasized crime, drug trafficking, and cartels at double the frequency of more liberal outlets.&lt;/p&gt;
&lt;p&gt;Q: How did voter policy preferences align with Republican versus Democratic platforms?
A: Using 2020 CCES data, the authors find that higher 1996 cancer mortality predicts a greater expressed preference for increasing the number of police officers on the street and a greater reported sense of safety around law enforcement—both consistent with the Republican Party&amp;rsquo;s law enforcement approach. Conversely, exposure to the epidemic predicts lower support for marijuana legalization on state ballot initiatives across 18 states from 2012 to 2023, indicating opposition to a key Democratic harm-reduction policy.&lt;/p&gt;
&lt;p&gt;Q: What role did political actors themselves play in driving the realignment?
A: Relatively little. The opioid epidemic was largely absent from House floor speeches until 2015 and from campaign advertising until 2020. Neither party took a clear legislative lead on the issue during the first two decades of the crisis. The authors interpret the political realignment as driven primarily by the Republican Party&amp;rsquo;s broader repositioning as the champion of working-class economic hardship and by differential media framing, rather than by active legislative competition over opioid policy.&lt;/p&gt;
&lt;p&gt;Q: What major confounds are ruled out?
A: The authors control for exposure to the China import shock, NAFTA, the 1994 Republican Revolution, the 2001 and 2008–2009 recessions, declining unionization, robot adoption, Fox News entry, deaths of despair (which include but are not limited to opioid deaths), and the political realignment of the South, rural areas, evangelicals, and the population over 65. Results remain robust across all these specifications. Placebo instruments using unrelated causes of death yield null results.&lt;/p&gt;
&lt;p&gt;Q: Could the vote-share effects be mechanically driven by opioid-related deaths removing Democratic voters from the electorate?
A: The authors perform a back-of-the-envelope calculation and estimate that even if all opioid-related deaths would have been Democratic votes, the mechanical effect on the Republican vote share is at most 0.22 percentage points relative to the observed 2020 vote share—far smaller than the estimated 4.5-percentage-point shift by 2022. The result is also inconsistent with a turnout mechanism, as voter turnout shows no meaningful change with epidemic exposure.&lt;/p&gt;
&lt;p&gt;Opioid epidemic exposure instrument: The paper measures community-level exposure to the opioid epidemic using cancer mortality rates in 1996, the year OxyContin launched. This instrument is grounded in Purdue Pharma&amp;rsquo;s documented marketing strategy of targeting the cancer pain market first; areas with more cancer patients received disproportionate pharmaceutical marketing, generating differential opioid prescription growth that extended well beyond cancer patients to the broader noncancer population through shared physicians.&lt;/p&gt;
&lt;p&gt;Commuting zone (CZ): The paper&amp;rsquo;s unit of geographic analysis, defined to capture local economic markets. There are 720 CZs in the US, encompassing all metropolitan and nonmetropolitan areas. The authors use 625 CZs with more than 20,000 residents, which account for more than 99% of all opioid deaths and total population.&lt;/p&gt;
&lt;p&gt;Two-party Republican vote share: The ratio of votes for Republican candidates to the total votes for both Republican and Democratic candidates in a given election. The paper tracks this measure for House, presidential, and gubernatorial elections from 1976 or 1982 through 2020 or 2022, depending on data availability.&lt;/p&gt;
&lt;p&gt;Drug-induced mortality: The paper&amp;rsquo;s broadest mortality measure, covering deaths from poisoning and medical conditions caused by legal or illegal drugs, including prescription opioids, heroin, and synthetic opioids such as fentanyl. It is distinguished from the narrower measures of prescription opioid deaths and all opioid deaths.&lt;/p&gt;
&lt;p&gt;Issue ownership: The political science concept, used in the paper to describe how the Republican Party repositioned itself during the epidemic period as the voice of working-class economic hardship, &amp;ldquo;forgotten America,&amp;rdquo; and &amp;ldquo;America left behind.&amp;rdquo; The paper contrasts this with Democratic ownership of income inequality and argues that Republican ownership of the hardship narrative made the party&amp;rsquo;s message especially salient in heavily opioid-affected communities.&lt;/p&gt;
&lt;p&gt;Path dependency in pharmaceutical marketing: Purdue&amp;rsquo;s strategy of concentrating initial OxyContin promotion in cancer-market areas, then later focusing on top-prescribing physicians (the highest three deciles of the distribution), meant that areas receiving high initial cancer-market promotion continued to receive disproportionate promotion as the company expanded to the noncancer market. This created a persistent targeting advantage for high-cancer CZs throughout the epidemic&amp;rsquo;s first wave.&lt;/p&gt;
&lt;p&gt;Nokken-Poole ideological measure: A roll-call-based measure of elected House members&amp;rsquo; ideology along the liberal-conservative dimension. The paper uses this measure to show that the epidemic shifted the composition of the House toward more conservative members, not by electing more extreme candidates in any given election, but by changing which party won seats over time.&lt;/p&gt;</description></item><item><title>Revolutionary Transition: Inheritance Change and Fertility Decline</title><link>https://macropaperwarehouse.com/papers/revolutionary-transition-inheritance-change-and-fertility-decline/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/revolutionary-transition-inheritance-change-and-fertility-decline/</guid><description>&lt;p&gt;Gay, Gobbi, and Goñi test Le Play&amp;rsquo;s (1875) hypothesis that the French Revolution contributed to France&amp;rsquo;s early fertility decline by abolishing impartible inheritance. In 1793, a series of decrees culminating in the Loi de Nivôse (January 6, 1794) abolished testamentary rights and imposed equal partition of assets among all children — partible inheritance — across France, overriding the mosaic of local customs and written laws that had governed inheritance in the Ancien Régime.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central argument is that this reform reduced the economic incentive to have children through indivisibility constraints in agricultural land. Under impartible inheritance, land passed to a single heir undivided, keeping plots above the subsistence productivity threshold even at high fertility. Under partible inheritance, each additional child fragments the land further, potentially pushing plots below the minimum productive size, so households face a strong incentive to limit fertility. A Stone-Geary production function with a minimum land threshold L̄ formalizes this mechanism: when landholdings fall in the binding range (L̄ &amp;lt; L &amp;lt; L̃), fertility is strictly higher under impartible than under partible inheritance.&lt;/p&gt;
&lt;p&gt;The authors construct the first complete map of inheritance rules across France&amp;rsquo;s 435 judicial districts as of 1789, classifying each along two dimensions: partible versus impartible, and whether women were included or excluded. This atlas draws on Brette&amp;rsquo;s (1904) Atlas des Bailliages and the Nouveau Coutumier Général (Bourdot de Richebourg 1724), covering 141 distinct customs. Treatment is defined as municipalities under impartible inheritance before 1793 whose system was altered by the reforms; control municipalities were already under partible inheritance.&lt;/p&gt;
&lt;p&gt;The main identification strategy is a difference-in-differences (DD) design comparing women with varying lengths of remaining fertile years after 1793 — from 0 for women aged 40+ at the reform to 25 for women aged 15 or younger — across treated and untreated municipalities. This is augmented by a regression-discontinuity difference-in-differences (RD-DD) design exploiting sharp discontinuities at judicial district borders. Two independent datasets are used: the Enquête Louis Henry (34,812 women in 39 rural municipalities, family-reconstitution method) and Geni.com crowdsourced genealogies (11,649 women across 2,966 locations after the Blanc 2023 horizontal restriction).&lt;/p&gt;
&lt;p&gt;Each additional fertile year of exposure to the 1793 reforms reduced completed fertility by approximately 1 percent. Over the full 25-year fertile cycle, this corresponds to a reduction of roughly 0.7 children, or 24 percent relative to the pre-reform mean of 2.92 surviving children in treated areas. This magnitude equals the entire pre-reform fertility gap between impartible- and partible-inheritance areas (2.9 versus 2.2 children), meaning the reforms closed this gap entirely. DD and RD-DD estimates are similar and not statistically distinguishable from each other, and results replicate across both datasets. Results hold on both the extensive margin (childlessness) and intensive margin (fertility of mothers).&lt;/p&gt;
&lt;p&gt;The mechanism is most relevant where smallholder landownership is widespread. France — where 40–80 percent of households owned land at the eve of the Revolution — meets this condition. England and Prussia, with more concentrated landownership, would not be expected to show the same response because the indivisibility constraint would not bind even after partition.&lt;/p&gt;
&lt;p&gt;Q: What was France&amp;rsquo;s inheritance system before the Revolution, and how heterogeneous was it?
A: Before 1793, inheritance was governed by 141 distinct customary and written laws applied within 435 judicial districts. The country was broadly divided between the customary-law north (Pays de droit coutumier) and the Roman written-law south (Pays de droit écrit), with substantial local variation within regions. Systems ranged from strictly partible (equal division among all offspring) to impartible (primogeniture, ultimogeniture, or unigeniture). Systems also varied in whether women could inherit or received only a dowry. This geographic variation — rooted in the laws of Germanic peoples after the fall of Rome in 476 CE — is exogenous to late eighteenth-century economic conditions and provides the identifying variation for the paper.&lt;/p&gt;
&lt;p&gt;Q: What exactly did the 1793 reforms change, and were they enforced?
A: The Loi de Nivôse an II (January 6, 1794) abolished testamentary rights entirely and mandated equal partition of assets among all children, including women, throughout France. The reforms came unexpectedly — only 8 of 571 cahiers de doléances analyzed by Goy (1988) mentioned inheritance — and were motivated by the equality principle, legal unification, and the fear that revolutionary sympathizers would be disinherited (Lataste et al. 1901). Offspring quickly asserted their new rights, and by the late 1790s inheritance disputes were the most common cases before family tribunals (Desan 1997; Poumarède 2011).&lt;/p&gt;
&lt;p&gt;Q: What is the model&amp;rsquo;s core mechanism linking inheritance reform to fertility decline?
A: The model uses a Stone-Geary production function with a minimum land threshold L̄ below which output falls to zero. Under impartible inheritance, land passes undivided to a single heir, keeping the farm above L̄ regardless of family size. Under partible inheritance, each child receives an equal share, so adding children risks fragmenting plots below L̄ — a powerful incentive to limit family size. The fertility gap between impartible and partible households is at its maximum when landholdings fall in the intermediate range (L̄ &amp;lt; L &amp;lt; L̃) where the constraint is binding. As land size increases, the constraint becomes less binding but the positive fertility differential persists.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s main quantitative estimate of the reform&amp;rsquo;s effect on completed fertility?
A: Each additional fertile year of exposure to the 1793 reforms reduced completed fertility by approximately 1 percent. Over the full 25-year fertile cycle (ages 15–40), this implies a reduction of roughly 0.7 children, or 24 percent relative to the pre-reform mean of 2.92 surviving children in treated areas. This is nearly identical to the pre-existing fertility gap between impartible- and partible-inheritance areas (0.7 children: 2.9 versus 2.2 surviving children), implying the reforms effectively eliminated the fertility differential.&lt;/p&gt;
&lt;p&gt;Q: Are the DD and RD-DD estimates consistent with each other, and do both datasets agree?
A: Yes. The DD and RD-DD estimates are similar and not statistically different from each other. The RD-DD design compares women born close to judicial district borders where inheritance rules differed, before and after 1793, exploiting the sharp spatial discontinuity at those borders. Consistency across these two designs — which rely on different identifying assumptions — strengthens causal interpretation. Results are also consistent across the Enquête Louis Henry (family-reconstitution) and Geni.com (crowdsourced genealogies) datasets, which are produced by fundamentally different methodologies.&lt;/p&gt;
&lt;p&gt;Q: How do the authors verify the parallel trends assumption?
A: Figure 6 shows that for cohorts who completed their fertile cycle before 1793, fertility trended downward in parallel across partible- and impartible-inheritance areas: a constant gap of approximately 0.7 children was maintained from women born in the early 1700s (3 versus 2.3 children) through women born in the early 1750s (2.7 versus 2.0 children), the last cohorts to complete fertility before the reforms. The convergence — from 0.7 to 0 children — only begins among cohorts fertile after 1793. The authors also include flexible trend controls interacted with municipality-level religiosity, political support for the Revolution, proximity to administrative centers, and wheat prices, and confirm the main estimate is robust.&lt;/p&gt;
&lt;p&gt;Q: What role did the extension of inheritance rights to women play?
A: The extension of rights to women was a companion mechanism distinct from abolishing impartible inheritance. Beyond increasing the number of heirs (which directly reduces land per heir), the right to inherit improves a woman&amp;rsquo;s outside option and postpones entry into marriage, following de Moor and van Zanden (2010). The DD and RD-DD estimates suggest that including women in inheritance and abolishing impartible inheritance had similar effects on fertility. The paper treats these as separate but reinforcing channels.&lt;/p&gt;
&lt;p&gt;Q: How do the authors address potential confounders — mortality, migration, and economic conditions?
A: On mortality: child mortality did not evolve differently after 1793 across areas with different inheritance rules (Appendix Table A3), and baseline adult mortality (age at death, probability of dying before completing the fertile cycle) was balanced across treated and control areas (Table 1). On migration: the authors explicitly rule out that results are driven by migration. On economic conditions: municipality-specific decade-average wheat prices (Ridolfi 2019) are included as controls for local Malthusian dynamics, and results are robust to their inclusion.&lt;/p&gt;
&lt;p&gt;Q: What do the balance tests show?
A: Panel A of Table 1 shows that before the reforms, areas with impartible versus partible inheritance were balanced on 9 of 11 individual-level characteristics — including husband and wife age at death, probability of dying before completing the fertile cycle, probability that parents-in-law were alive at marriage, literacy, data accuracy, and age at marriage. The only systematic pre-reform difference was fertility itself (0.7 children). Municipality-level climatic variables, soil suitability, and proxies for mortality uncertainty were also balanced. This is consistent with the origins of these systems in post-Roman Germanic law, which are unrelated to late eighteenth-century economic conditions.&lt;/p&gt;
&lt;p&gt;Q: What robustness checks are reported?
A: The authors report: (1) permutation tests reshuffling treatment exposure across women and municipalities; (2) non-linear treatment effects across cohorts, showing the heterogeneity required to explain away the baseline estimate is implausibly large per de Chaisemartin and d&amp;rsquo;Haultfoeuille (2020); (3) exclusion of outlier municipalities; (4) a placebo test for cohorts who completed their fertile cycle before 1793; (5) robustness to alternative sample definitions, treatment definitions, outcome variables, and control groups; (6) Cummins (2020) first-name repetition technique to correct for under-reported child deaths in Henry; (7) terrain characteristics including climatic and soil suitability (Galor and Özak 2016) and ruggedness (Nunn and Puga 2012); (8) for RD-DD: alternative bandwidths, running variable specifications, kernel functions, samples, and border-segment fixed effects. All checks support the main finding.&lt;/p&gt;
&lt;p&gt;Q: Why did France experience a fertility decline from inheritance reform while other countries with similar reforms did not?
A: The model rationalizes this through landownership structure. The fertility-reducing mechanism operates through indivisibility constraints that bind only when landholdings are small and fragmented — as in France, where 40–80 percent of households owned their land and plots were small. Where landownership is concentrated (England, Prussia), land per heir remains above L̄ even after partible division, so the indivisibility constraint is non-binding and fertility is unaffected by the reform. This provides a structural reason why France&amp;rsquo;s particular agrarian structure made it uniquely susceptible to this mechanism.&lt;/p&gt;
&lt;p&gt;Q: What is the broader historical significance for understanding France&amp;rsquo;s early demographic transition?
A: France&amp;rsquo;s fertility decline began roughly 50 years before industrialization, making it anomalous relative to standard quantity-quality tradeoff theories linking fertility decline to technological progress and rising returns to human capital. The 1793 reforms provide a legal-institutional explanation for the sharp post-Revolution acceleration visible in Figure 1, which is difficult to attribute to slowly-evolving cultural factors or human capital considerations not yet operative. The estimates imply the reforms brought large impartible-inheritance areas to the low-fertility regime that already characterized partible-inheritance areas, thus sharply accelerating the national transition.&lt;/p&gt;
&lt;p&gt;Impartible inheritance: A system under which parents could designate a single heir (through primogeniture, ultimogeniture, or unigeniture) to receive the bulk of the family estate, preventing fragmentation of wealth; in pre-revolutionary France this was associated with extended family households and higher fertility (2.9 surviving children on average) relative to partible areas (2.2).&lt;/p&gt;
&lt;p&gt;Partible inheritance: A system under which family wealth was divided equally among all offspring upon death; in the paper&amp;rsquo;s model this creates an incentive to limit fertility to prevent land fragmentation below the subsistence productivity threshold L̄.&lt;/p&gt;
&lt;p&gt;Indivisibility constraint (land threshold L̄): In the Stone-Geary production function, a minimum land input below which agricultural output falls to zero; this is the mechanism through which partible inheritance generates fertility-limiting incentives, since dividing a small plot among many heirs risks crossing L̄ into zero production.&lt;/p&gt;
&lt;p&gt;Difference-in-differences (DD) exposure design: The paper&amp;rsquo;s main identification strategy, using remaining fertile years after 1793 as a continuous treatment-intensity variable (0 for cohorts past fertility at the reform date, up to 25 for cohorts entirely within their fertile years), compared between treated municipalities (impartible → partible) and control municipalities (already partible).&lt;/p&gt;
&lt;p&gt;Regression-discontinuity difference-in-differences (RD-DD): An augmented design exploiting the sharp geographic discontinuity at borders between judicial districts with different pre-reform inheritance rules, comparing outcomes on both sides before and after 1793, to address smooth unobserved confounders.&lt;/p&gt;
&lt;p&gt;Completed fertility (net): The number of children surviving to age six, preferred over total births because child mortality before 1800 was high (1–1.5 children per mother did not survive to age six per Houdaille 1984), making net fertility the more economically meaningful measure for inheritance and bequest decisions.&lt;/p&gt;
&lt;p&gt;Horizontal restriction: A sampling correction applied to crowdsourced genealogical data (Blanc 2023a) that retains an observation only if at least one of the four preceding generations has more than one recorded offspring, correcting for the over-representation of single-child families that arises because Geni users tend to record direct ancestors rather than collateral relatives.&lt;/p&gt;</description></item><item><title>Risk Sharing Tests and Covariate Shocks: Drought, Floods, and Pests in Uganda</title><link>https://macropaperwarehouse.com/papers/risk-sharing-tests-and-covariate-shocks-drought-floods-and-pests-in-uganda/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/risk-sharing-tests-and-covariate-shocks-drought-floods-and-pests-in-uganda/</guid><description>&lt;p&gt;This paper identifies and corrects a fundamental flaw in the standard methodology for testing efficient risk-sharing when shocks are covariate (affecting common prices rather than only individual incomes). The standard Townsend (1994) approach infers marginal utilities of expenditure (MUEs) from total expenditures, which implicitly assumes homothetic preferences — specifically Constant Relative Risk Aversion (CRRA) — under which all goods have unitary income elasticities and a single scalar price index captures all price effects. Ligon demonstrates that this assumption causes the standard test to fail when applied to covariate shocks such as droughts, floods, and agricultural pests, because these shocks change relative prices in ways that cannot be captured by a single price index. The perverse consequence is that in Ugandan data, every covariate shock — drought, floods, pests, and adverse prices — appears to improve household welfare under the CRRA specification (significant positive coefficients of 0.046, 0.097, 0.095, and 0.103 respectively, all significant at p&amp;lt;0.01), a result the paper argues is mechanically induced by the mis-specification rather than reflecting reality.&lt;/p&gt;
&lt;p&gt;The paper makes two core theoretical contributions. First, it characterizes the complete class of preferences that permit MUE inference from expenditure data alone — specifically, requiring that item-level expenditures be &amp;ldquo;lambda-separable&amp;rdquo; (additively separable in the MUE and prices). Solving the resulting functional equations yields exactly two families of semiparametric demand systems: Constant Frisch Elasticity (CFE) demands (a generalization of CRRA) and Generalized Stone-Geary demands. Only CFE demands are tractable for panel estimation. Second, the paper shows that under CFE preferences, log expenditures on each good j follow the system: log x^j_it = a_j(p_t) + g_j(z_it) + beta_j * w_it + epsilon^j_it, where beta_j is the good-specific Frisch elasticity and w_it = -log lambda_it is the negative log MUE. This allows price effects to enter flexibly through good-time fixed effects rather than a single index, and MUEs to be recovered via factor analysis on the residual covariance matrix.&lt;/p&gt;
&lt;p&gt;The empirical work uses eight waves of the Ugandan National Panel Surveys (2005–2020), an unbalanced panel of 5,601 distinct households yielding 22,791 usable household-year observations across 41 consumption goods (primarily food items). Uganda is divided into four regional markets, producing 32 market-year cells and 1,312 market-year-good dummies. Estimated Frisch elasticities vary substantially across goods — passion fruit is roughly three times as income elastic as cassava — emphatically rejecting the hypothesis of equal elasticities required by CRRA.&lt;/p&gt;
&lt;p&gt;Using CFE-estimated MUEs, the risk-sharing test shows that none of the four covariate shocks has a significant effect on welfare (CFE coefficients: drought 0.010, floods 0.035, pests 0.041, adverse prices -0.043, all insignificant). The pattern holds across all time windows from 0–12 months: 42 of 52 covariate shock coefficients are significant and positive in the CRRA specification, versus only 4 of 52 in the CFE specification — barely above the 2.6 false positives expected under the null. These findings indicate that the welfare impacts of covariate shocks in Uganda operate primarily through the common price channel rather than through idiosyncratic income variation, meaning they are broadly shared within market-regions. Idiosyncratic income shocks, by contrast, show the expected pattern: they reduce welfare significantly in both specifications (CFE: 0.050***, CRRA: 0.071***), and health shocks are significant only in CFE (−0.059**).&lt;/p&gt;
&lt;p&gt;Q: Why does the standard CRRA risk-sharing test fail for covariate shocks?
A: Under CRRA preferences, MUEs depend on total expenditures only through a single scalar price index pi(p). When a covariate shock raises prices of inelastic goods (primarily food), total food expenditures increase even as actual consumption quantities fall. Because risk-sharing tests based on CRRA total expenditures cannot separate this price effect from a welfare improvement, the shock appears to raise welfare. The disturbance term in the CRRA TWFE regression depends on the very prices affected by covariate shocks, violating the exclusion restriction.&lt;/p&gt;
&lt;p&gt;Q: What is the lambda-separability condition, and why does it matter?
A: Lambda-separability requires that for each good j, some transformation phi_j of expenditures on that good can be written as the sum of a function of prices and a function of the MUE: phi_j(x_j(p,lambda)) = a_j(p) + b_j(lambda). This property is necessary for time fixed effects to absorb price variation and household fixed effects to absorb Pareto weights, which is the identification strategy behind all TWFE risk-sharing tests. Without it, no panel estimator using only expenditure data can consistently recover MUEs.&lt;/p&gt;
&lt;p&gt;Q: What are the two demand families that satisfy lambda-separability, and what distinguishes them?
A: Theorem 1 establishes that rationalizable lambda-separable demands must belong to either the Constant Frisch Elasticity (CFE) family or the Generalized Stone-Geary family. In CFE demands, log expenditures on each good equal the log of a price function minus beta_j times log lambda, where beta_j is a good-specific constant Frisch elasticity. The Stone-Geary family has a more complex nonlinear form that does not lend itself to linear estimation of log MUEs, making CFE the tractable choice. Both families nest CRRA as the special case where all beta_j are equal.&lt;/p&gt;
&lt;p&gt;Q: How are MUEs estimated from the CFE system in practice?
A: Estimation proceeds in two steps. First, log expenditures on each good are regressed on good-time-market effects and household demographic controls to obtain residuals. Second, the covariance matrix of these residuals has the factor structure Sigma = Var(w)&lt;em&gt;beta&lt;/em&gt;beta&amp;rsquo; + Psi, where beta is the vector of Frisch elasticities; the rank-one matrix beta*beta&amp;rsquo; is recovered from the sample covariance matrix via factor analysis, and household-level MUEs are then obtained by regression using the estimated beta as generated regressors.&lt;/p&gt;
&lt;p&gt;Q: What do the estimated Frisch elasticities reveal about preferences in Uganda?
A: The Frisch elasticities beta_j vary substantially across the 41 goods in the Ugandan sample. Starchy staples and salt are least elastic (lowest beta_j), while fresh milk, sweet bananas, coffee, oranges, and passion fruit exhibit high elasticities — passion fruit is roughly three times as income elastic as cassava. The hypothesis that all elasticities are equal (the CRRA restriction) is easily rejected, providing direct evidence against homothetic preferences in this population.&lt;/p&gt;
&lt;p&gt;Q: What direct evidence does the paper provide that droughts, floods, and pests are genuinely covariate and harmful?
A: About 39% of Ugandan households reported drought in the 2005–06 round. Among drought reporters, 92% said it affected their production, 80% said it affected their income, and 50% said it affected their consumption. Drought, pests, and adverse prices (but not floods) led to statistically significant increases in local farmgate prices. Among markets experiencing covariate shocks, 82%, 74%, 44%, and 53% of t-tests rejected equality of relative food prices for drought, floods, pests, and adverse prices respectively. Dietary diversity and intake of vitamin B-12 (from animal-source foods) declined significantly following covariate shocks.&lt;/p&gt;
&lt;p&gt;Q: How do households cope differently with covariate versus idiosyncratic shocks?
A: Households experiencing covariate shocks primarily relied on self-insurance: 51% of drought-affected households reduced consumption and 45% drew on savings, with increased labor supply also reported. In contrast, households experiencing idiosyncratic shocks most often relied on help from friends and family (52%). This behavioral difference is consistent with the finding that covariate shocks affect welfare mainly through common price channels that are not individually insurable through social networks, while idiosyncratic shocks are partially absorbed via informal transfers.&lt;/p&gt;
&lt;p&gt;Q: What do the CFE results imply about the nature of insurance against covariate shocks in Uganda?
A: The CFE regression finds that none of the four covariate shocks (drought, floods, pests, adverse prices) has a statistically significant effect on household MUEs when time-market fixed effects are included. This implies that the welfare impact of covariate shocks is transmitted primarily through common price changes that affect all households in a market-region symmetrically, rather than through idiosyncratic income variation. Effectively, covariate shocks are &amp;ldquo;shared&amp;rdquo; within market-regions — but through price deterioration affecting everyone, not through informal transfers.&lt;/p&gt;
&lt;p&gt;Q: How robust are the results across different shock time windows?
A: Figure 3 shows that for the CRRA specification, any prior covariate shock 3–12 months earlier has a significant positive effect on log consumption in every month, while for the CFE specification no shock window produces a significant effect on w. In the full tabulation across all shock types and windows (Tables 4 and 5), 42 of 52 covariate shock coefficients are significant and positive in CRRA versus only 4 of 52 in CFE — the latter barely exceeding the 2.6 false positives expected under the null hypothesis of full insurance.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of these findings for relief program design?
A: Because covariate shocks affect welfare mainly through common prices within market-regions, relief programs should target communities rather than individual households, since the burden is broadly shared and not concentrated. Policies that integrate markets across regions of Uganda or connect Ugandan markets to broader African or world markets would reduce the price impact of local covariate shocks. Targeted household transfers would be less effective than interventions that stabilize regional prices or supply.&lt;/p&gt;
&lt;p&gt;Q: What broader applicability do CFE MUEs have beyond risk-sharing tests?
A: Since MUE construction is independent of the risk-sharing hypothesis, CFE-estimated MUEs can be used to estimate and test any dynamic life-cycle model that puts structure on the evolution of MUEs over time, including consumption Euler equations, intertemporal marginal rates of substitution calculations, and household bargaining models. The CFE approach requires only the same expenditure data used in the standard CRRA approach and therefore serves as a more general drop-in replacement across all settings where CRRA MUEs are currently employed.&lt;/p&gt;
&lt;p&gt;Marginal Utility of Expenditure (MUE): The Lagrange multiplier lambda on the household budget constraint in the consumer&amp;rsquo;s optimization problem; the object whose proportionality across households (log lambda_it = log mu_t - log theta_i) characterizes efficient risk-sharing. It is a function of budget, prices, and household characteristics — not reducible to a scalar function of total expenditure except under special preference restrictions.&lt;/p&gt;
&lt;p&gt;Lambda-separability: A property of Frischian expenditures on good j such that some transformation phi_j(x_j) can be written as the sum of a function of prices and a function of the MUE alone — phi_j(x_j(p,lambda)) = a_j(p) + b_j(lambda). This is the necessary and sufficient condition for using time fixed effects to control for prices and household fixed effects to control for Pareto weights in a TWFE risk-sharing regression based solely on expenditure data.&lt;/p&gt;
&lt;p&gt;Constant Frisch Elasticity (CFE) expenditure system: The tractable member of the two demand families satisfying lambda-separability, characterized by log x^j_it = a_j(p_t) + g_j(z_it) + beta_j * w_it + epsilon^j_it, where beta_j is a good-specific constant elasticity of expenditures with respect to MUE. Nests CRRA as the special case of equal beta_j across all goods, but admits nonhomothetic preferences and fully flexible relative-price responses.&lt;/p&gt;
&lt;p&gt;Frischian demands: Demands expressed as functions of prices and the MUE lambda rather than prices and budget — f(p, lambda). Homogeneous of degree zero in (p, 1/lambda), equivalently written f(p*lambda). This representation is central to the lambda-separability characterization because it separates the role of the budget (via lambda) from the role of prices directly.&lt;/p&gt;
&lt;p&gt;Covariate shock: In this paper&amp;rsquo;s usage, a shock that affects prices common to all households in a market-region — not merely a shock affecting many households simultaneously. The key analytical distinction is that idiosyncratic shocks change individual budgets without changing prices, while covariate shocks change prices, which is what causes the standard CRRA test to fail.&lt;/p&gt;
&lt;p&gt;Nonhomothetic preferences: Preferences for which expenditure shares vary with income (budget), so no single scalar price index can fully represent the welfare impact of price changes. The paper confirms nonhomotheticity in the Ugandan data through widely varying Frisch elasticities, and argues this is the root cause of the CRRA test&amp;rsquo;s failure for covariate shocks — a problem that does not arise when shocks are idiosyncratic and leave prices unchanged.&lt;/p&gt;</description></item><item><title>Rural Migrants and Urban Informality: Evidence From Brazil</title><link>https://macropaperwarehouse.com/papers/rural-migrants-and-urban-informality-evidence-from-brazil/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/rural-migrants-and-urban-informality-evidence-from-brazil/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Does rural-urban migration increase or decrease urban informality, and through what mechanisms — and does the answer depend on the time horizon?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and Data.&lt;/strong&gt; The paper studies internal migration in Brazil over 2000–2010. The empirical analysis combines: (i) two waves of the Decennial Population Census (2000 and 2010) covering working-age adults (ages 15–64) across 3,548 Minimum Comparable Areas (MCAs); (ii) the universe of formal firms and workers from the matched employer-employee administrative dataset RAIS (1997–2018); (iii) the ECINF informal firm survey (2003); and (iv) the annual National Household Survey (PNAD, 2001–2009) for year-on-year short-run analysis in 700 identifiable municipalities. Internal immigration to the average urban destination was large: 17.6 percent overall over the decade, 7 percent for state-to-state migration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Design.&lt;/strong&gt; The authors use a shift-share instrumental variable (IV) design. The shares are pre-existing migration networks (migrant flows by origin-destination pair, 1995–2000). The shifts are drought shocks constructed from the Standardized Precipitation-Evapotranspiration Index (SPEI) interacted with agricultural crop calendars and the value share of each crop in each origin municipality — accumulated over the 2000–2010 decade. A second independent instrument uses international commodity price shocks as push factors (following a China-analogous construction); the two instruments are nearly uncorrelated across origins (0.007) and only weakly correlated across destinations (-0.3), providing an independent validation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Long-Run Findings (decadal changes, 2000–2010).&lt;/strong&gt; A one-percentage-point increase in the immigration rate (equal to 18.5 percent of a standard deviation):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Increases the share of workers in formal wage employment by &lt;strong&gt;0.27 percentage points&lt;/strong&gt; (a 1.2 percent increase from the mean of 23 percent).&lt;/li&gt;
&lt;li&gt;Decreases the share in informal wage employment by &lt;strong&gt;0.29 percentage points&lt;/strong&gt; (a 2.9 percent decrease from the mean of 10 percent).&lt;/li&gt;
&lt;li&gt;Has no effect on overall wage employment, unemployment, or self-employment — the formalization effect is a reallocation from informal to formal jobs, not net job creation.&lt;/li&gt;
&lt;li&gt;Reduces formal sector wages by &lt;strong&gt;0.6 percent&lt;/strong&gt;, with no effect on informal wages.&lt;/li&gt;
&lt;li&gt;Increases the number of formal establishments by &lt;strong&gt;1.6 percent&lt;/strong&gt; and the number of formal jobs by &lt;strong&gt;2 percent&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Raises gross firm entry by &lt;strong&gt;2.8 percent&lt;/strong&gt; and gross firm exit by &lt;strong&gt;3 percent&lt;/strong&gt; (higher churn), with effects stable or slightly increasing through 2017–18.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These firm-creation effects are not driven by migrants starting businesses: migrants are not more likely to be business owners in high-immigration municipalities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Short-Run Findings.&lt;/strong&gt; Using year-on-year specifications with the PNAD (2001–2009), the authors replicate the results in the prior literature: municipalities receiving more migrants experience a reduction in formal wage employment, with no change in informal employment or non-employment — so the share of informal jobs rises. These short-run informality-increasing effects coexist with the long-run formalization results, and are not a sample artifact (the long-run results are unchanged when restricted to the same 700 PNAD municipalities).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanism — Downward Nominal Wage Rigidity (DNWR).&lt;/strong&gt; DNWR in the formal sector is the key mechanism reconciling short- and long-run effects. In Brazil, nominal wage cuts were illegal, and the national minimum wage rose regularly during the 2000s. Two municipality-level DNWR proxies are used: (i) the Kaitz index (national minimum wage / municipality median wage in 2000); (ii) the share of workers with negative year-on-year nominal wage changes (from RAIS, 1997–2000). In municipalities with higher DNWR: the positive formalization effects of immigration are smaller or fully muted; non-employment increases; and formal wages decline less. These cross-sectional patterns echo the Harris-Todaro-Fields prediction, and are consistent with DNWR being more binding in the short run (when nominal rigidities bind) than in the long run (when inflation and worker turnover allow real wage adjustment).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The paper develops and estimates a dynamic model of firm dynamics and informality, extending the canonical Hopenhayn framework with (i) two margins of informality — the extensive margin (whether a firm registers) and the intensive margin (whether a registered formal firm hires workers formally) — and (ii) heterogeneous long-run productivity parameters (nu) that generate firm-specific life-cycle growth profiles. Formal firms cannot revert to informality; informal firms can formalize by paying the cost differential between formal and informal entry costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Counterfactuals.&lt;/strong&gt; A simulated once-and-for-all 10 percent labor supply shock (approximately the 80th percentile of observed immigration shocks) produces: a 4.1 percent decline in the share of informal workers (IV: 7.5 percent); a 16.1 percent increase in formal firms (IV: 21.1 percent); and a 3.4 percent wage decline (IV: 5 percent). Of the increase in formal firms, &lt;strong&gt;40 percent&lt;/strong&gt; is accounted for by formalization of previously informal firms, highlighting the stepping-stone role of informality that a static or dual-economy model would miss. Average firm productivity declines by 1.4 percent due to worsening firm composition (the share of formal firms in the lowest productivity quartile rises by more than 4 percentage points). A counterfactual that nearly eliminates the extensive margin of informality (via steep enforcement costs) raises total output by 8.6 percent vs. 7 percent in the baseline shock, and increases average firm productivity by 2.1 percent vs. a decline of 1.4 percent — at the cost of displacing the least productive informal firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; Results pertain to internal (not international) migration; drought-induced migrants do not change the skill composition of the labor force at destination, justifying a homogeneous worker assumption. The formalization effects hold for migrants and non-migrants separately, and for high- and low-skilled workers separately. The model is calibrated to the average urban destination in Brazil, not a spatial general equilibrium.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-key-threats-to-validity-the-authors-address"&gt;Q1. What is the identification strategy, and what are the key threats to validity the authors address?&lt;/h3&gt;
&lt;p&gt;The authors use a shift-share IV where shifts are drought shocks at origin municipalities (constructed from SPEI x crop calendar x crop revenue share, accumulated over 2000–2010) and shares are pre-2000 migration networks. Threats addressed: (i) pre-trends — no evidence of differential pre-trends in firm outcomes between 1997–98 and 1999–2000; (ii) demand channel — controlling for local drought shocks and distance-weighted neighboring shocks leaves results unchanged; (iii) capital reallocation — adding a bank-network-based shift-share control (following prior literature) does not change results; (iv) agricultural processing linkages — results hold after excluding agricultural firms and food/beverage/tobacco manufacturers; (v) migration persistence — controlling for baseline log population and 1995–2000 migration rates leaves results unchanged. The commodity-price-shock instrument provides an independent validation, yielding similar results despite near-zero cross-origin correlation with drought shocks and only -0.3 correlation across destinations.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-authors-reconcile-the-long-run-formalization-result-with-the-short-run-informality-increasing-result-and-what-role-does-dnwr-play"&gt;Q2. How do the authors reconcile the long-run formalization result with the short-run informality-increasing result, and what role does DNWR play?&lt;/h3&gt;
&lt;p&gt;DNWR is the key mechanism. Nominal wage cuts are illegal in Brazil&amp;rsquo;s formal sector, and the minimum wage rose through the 2000s, making DNWR binding especially in the short run. In the year-on-year specification (PNAD, 2001–2009), immigration reduces formal wage employment with no change in informal employment, raising the informal share — consistent with prior literature. Over the decade, inflation and worker turnover permit real formal wage adjustment, enabling formal sector expansion. Cross-sectional heterogeneity confirms this: in municipalities with above-median Kaitz index or below-median share of negative wage changes, the formalization effect of immigration is smaller or zero, and non-employment rises — precisely the Harris-Todaro-Fields prediction for rigid-wage environments.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-exact-magnitude-of-the-firm-level-effects-and-how-persistent-are-they"&gt;Q3. What is the exact magnitude of the firm-level effects and how persistent are they?&lt;/h3&gt;
&lt;p&gt;A one-percentage-point increase in the immigration rate increases formal establishments by 1.6 percent, formal jobs by 2 percent, firm entry by 2.8 percent, and firm exit by 3 percent — all decadal effects (1999–2000 to 2011–12). Effects on firms, entry, exit, and jobs remain stable or slightly increasing through 2017–18 as estimated using RAIS panel data, with no evidence of pre-trends (effects near zero in 1997–98 to 1999–2000 period). The effect on firm-level average wages is negative (consistent with the worker-level wage effect) but not statistically significant.&lt;/p&gt;
&lt;h3 id="q4-are-migrants-themselves-the-source-of-new-formal-firm-creation"&gt;Q4. Are migrants themselves the source of new formal firm creation?&lt;/h3&gt;
&lt;p&gt;No. The authors directly test and reject this channel. Migrants are not more likely to be business owners — either of small firms (fewer than 5 employees) or larger firms (6 or more employees) — in municipalities that receive more immigration. The increase in formal firm entry is driven by non-migrants responding to cheaper labor.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-two-margins-of-informality-in-the-model-and-why-does-the-intensive-margin-matter-for-the-migration-formality-nexus"&gt;Q5. What are the two margins of informality in the model, and why does the intensive margin matter for the migration-formality nexus?&lt;/h3&gt;
&lt;p&gt;The extensive margin is whether a firm registers formally (firm-level binary). The intensive margin is whether a formally registered firm hires workers without formal labor contracts (worker-level, within formal firms). The intensive margin is crucial because it links formal firms to migrants: newly arrived migrants may take informal jobs within formal firms, allowing formal firm creation to respond to the immigration shock even before the labor market fully formalizes. In the transition dynamics after an immigration shock with DNWR, new formal firms tend to be small and lower-productivity, and hire a substantial fraction of their workforce informally — so labor informality hovers near its initial level for several years even as firm informality declines quickly.&lt;/p&gt;
&lt;h3 id="q6-what-fraction-of-the-increase-in-formal-firms-in-the-counterfactual-comes-from-stepping-stone-formalization-versus-new-formal-entry"&gt;Q6. What fraction of the increase in formal firms in the counterfactual comes from stepping-stone formalization versus new formal entry?&lt;/h3&gt;
&lt;p&gt;In the baseline 10 percent labor supply counterfactual, approximately &lt;strong&gt;40 percent&lt;/strong&gt; of the increase in the number of formal firms comes from formalization of previously informal firms across their life cycles. The remaining 60 percent comes from new formal firm creation. A static framework would miss the stepping-stone channel entirely and substantially underestimate total formalization.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-models-calibration-pin-down-the-cost-structure-of-informal-vs-formal-firms"&gt;Q7. How does the model&amp;rsquo;s calibration pin down the cost structure of informal vs. formal firms?&lt;/h3&gt;
&lt;p&gt;The model is calibrated using a two-step minimum distance procedure. First-step parameters include the persistence of formal firms&amp;rsquo; productivity process (estimated from RAIS: rho_f = 0.92), and statutory tax rates (payroll tax tau_w = 0.375; revenue VAT tau_y = 0.293). Second-step parameters (12 total, including entry costs, exogenous death rates, productivity dispersion, and cost-function curvatures for both margins of informality) are estimated by minimizing the distance between simulated and observed moments from RAIS (2003 cross-section for static moments; 2000–2011 panel for growth moments) and ECINF (informal firms with up to 5 employees, 2003). Key calibrated values: formal entry costs are more than twice informal entry costs and correspond to over 30 times the 2003 monthly national minimum wage; the informal sector exogenous death rate (delta_i = 0.148) is more than twice the formal rate; productivity variance and persistence are similar across sectors.&lt;/p&gt;
&lt;h3 id="q8-what-happens-to-firm-productivity-and-output-per-worker-in-the-long-run-counterfactual"&gt;Q8. What happens to firm productivity and output per worker in the long-run counterfactual?&lt;/h3&gt;
&lt;p&gt;Average firm productivity declines by 1.4 percent despite lower informality. The composition of formal firms worsens: the share of firms in the lowest productivity quartile rises by more than 4 percentage points, while the share in the top quartile falls by about 3 percentage points. Total output and tax revenues increase (7 and 8.6 percent, respectively), but both decline in per capita terms. The authors note these are likely lower bounds because the model assumes no technological differences between formal and informal sectors and no differential capital access.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-enforcement-counterfactual-reveal-about-the-dual-role-of-informality"&gt;Q9. What does the enforcement counterfactual reveal about the dual role of informality?&lt;/h3&gt;
&lt;p&gt;When the extensive margin of informality is nearly shut down (by making the informal cost function very steep), a 10 percent labor supply shock produces: output increase of 8.6 percent (vs. 7 percent with informality present); average firm productivity increase of 2.1 percent (vs. decline of 1.4 percent); much higher tax revenues due to greater formality. However, this comes at the cost of a sizable reduction in total firm count as the least productive informal firms are displaced. This illustrates the dual role: in the short run, the informal sector acts as an employment buffer and stepping-stone, which is more important when formal wage rigidity is stronger; but in the long run, it dampens aggregate economic benefits from immigration by sheltering low-productivity firms.&lt;/p&gt;
&lt;h3 id="q10-do-the-results-hold-for-both-migrants-and-non-migrants-and-across-skill-levels"&gt;Q10. Do the results hold for both migrants and non-migrants, and across skill levels?&lt;/h3&gt;
&lt;p&gt;Yes. Appendix results show similar employment and wage effects for migrants and non-migrants separately, though formal wage declines are more pronounced for non-migrants. Results are also similar for high- and low-skilled workers — which the authors attribute to the fact that drought-induced migration does not change the skill composition of the workforce at destination (confirmed empirically). Price-shock-induced migrants differ: they are more likely to be young and male, and do change workforce composition, providing a different set of compliers that strengthens external validity.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-relate-to-the-startup-deficit-literature-on-demographic-decline"&gt;Q11. How does the paper relate to the &amp;ldquo;startup deficit&amp;rdquo; literature on demographic decline?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s findings are the mirror image of the US startup deficit literature, which argues that demographic slowdown reduced firm entry, labor reallocation, and employment growth. The magnitudes are comparable in scale: the US startup deficit corresponds to a 5-percentage-point decline in firm entry between 1980 and 2012, while the rural-urban migration shocks studied here produce first-order effects on firm entry of similar or larger magnitude (2.8 percent per percentage point of immigration rate), suggesting labor supply growth is a primary driver of formal firm dynamics in both directions.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Downward Nominal Wage Rigidity (DNWR).&lt;/strong&gt; In the paper&amp;rsquo;s usage, the binding constraint that formal sector wages cannot be cut in nominal terms — in Brazil, both legal prohibition of nominal wage cuts and a rising national minimum wage. DNWR is the paper&amp;rsquo;s central mechanism explaining why immigration increases informality in the short run (wages cannot adjust) but reduces it over the decade (inflation and turnover permit real adjustment). Measured empirically via the municipality-level Kaitz index (national minimum wage / local median wage) and via the share of workers with negative year-on-year nominal wage changes in RAIS.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extensive Margin of Informality.&lt;/strong&gt; Whether a firm is registered with the government (formal) or not (informal). In the model, informal firms can avoid taxes but face a size-increasing cost of informality and the option to formalize by paying the difference in entry costs. This margin captures the firm&amp;rsquo;s legal registration status.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive Margin of Informality.&lt;/strong&gt; Whether a formally registered firm hires individual workers with or without formal labor contracts (signed work booklet, carteira de trabalho). Formal firms face increasing costs for informal hiring but exploit this margin for lower-cost labor, especially when small or young. This margin is critical because it links formal firms to migration-induced informal labor supply and allows formal firms to absorb migrants before full wage adjustment occurs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stepping-Stone Role of Informality.&lt;/strong&gt; The paper&amp;rsquo;s term for the dynamic channel through which the informal sector facilitates transitions to formality for both firms and workers. Informal firms accumulate productivity experience and formalize when productivity crosses the formalization threshold; informal workers within formal firms transition to formal contracts as firms grow. In the counterfactuals, 40 percent of the increase in formal firms following a labor supply shock is attributable to this channel. The stepping-stone role is most valuable during the short-run period of formal wage rigidity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Shift-Share Instrumental Variable.&lt;/strong&gt; The identification design combining pre-existing migration network shares (fraction of prior migrants to destination d from each origin o, computed 1995–2000) with exogenous push shocks at origin (drought shocks or commodity price shocks). The instrument predicts which destination municipalities receive more migrants based purely on exogenous origin-level shocks, purging the endogeneity from migrants self-selecting into prosperous cities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Minimum Comparable Area (MCA).&lt;/strong&gt; The paper&amp;rsquo;s geographic unit of analysis: a harmonized aggregation of Brazilian municipalities whose administrative borders changed during the study period, yielding 3,548 stable units covering all urban destinations studied. The authors call these &amp;ldquo;municipalities&amp;rdquo; for convenience.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Harris-Todaro-Fields Framework.&lt;/strong&gt; The theoretical benchmark against which the paper&amp;rsquo;s results are compared — the view (from Harris and Todaro 1970 and Fields) that rural-urban migration increases urban unemployment or informality because DNWR prevents the formal sector from absorbing migrants, who instead queue for formal jobs or enter the informal sector. The paper shows this prediction holds in the short run and in high-DNWR municipalities, but not in the long run where real wage adjustment occurs.&lt;/p&gt;</description></item><item><title>School Choice and the Housing Market</title><link>https://macropaperwarehouse.com/papers/school-choice-and-the-housing-market/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/school-choice-and-the-housing-market/</guid><description>&lt;p&gt;Grigoryan (2021) develops a unified general-equilibrium framework that jointly models school assignment mechanisms and the housing market to evaluate the welfare and distributional consequences of replacing traditional neighborhood assignment (NA) with the Deferred Acceptance (DA) mechanism. The paper fills a gap in the matching theory literature, where preferences and priorities are typically treated as exogenous, by making residential choices endogenous: families first observe which school assignment mechanism the district announces, then optimally select a neighborhood given market-clearing prices and other families&amp;rsquo; choices, and finally children are assigned to schools through the announced mechanism.&lt;/p&gt;
&lt;p&gt;The model features a continuum of families, each with a type defined by valuations over all neighborhood–school pairs, a finite set of neighborhoods and schools (one school per neighborhood), and competitive equilibrium prices. Three mechanisms are compared: NA (each child attends the neighborhood school), DA without neighborhood priority (DA), and DA with neighborhood priority (DN), where neighborhood residents receive priority at their local school.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s first major result (Theorem 3) is that DN unambiguously generates weakly higher aggregate welfare than NA. The proof exploits the fact that DN preserves NA&amp;rsquo;s option — families can still guarantee admission to the neighborhood school by living there — while additionally allowing families to access seats at other schools that go unclaimed by neighborhood residents. Although price effects under DN can make some individual families worse off relative to NA, aggregate welfare (inclusive of house sellers) is always weakly higher under DN. In simulations with 1,000 students, 10 neighborhoods, and 10 schools, DN yields average aggregate welfare gains of 2.40% relative to NA across the 18 parameter configurations studied.&lt;/p&gt;
&lt;p&gt;The welfare comparison between DA (without neighborhood priority) and NA is ambiguous in the general model: simulations show DA producing gains as large as +5.65% and losses as large as −18.26% relative to NA, depending on the degree of preference alignment across families (parameter α) and the variance in school capacities (parameter γ). DN also dominates DA in aggregate welfare under two sufficient conditions — identical ordinal preference rankings over neighborhoods and schools (Assumption 1 or 2) — though counterexamples exist when these assumptions fail.&lt;/p&gt;
&lt;p&gt;The second major result (Theorem 5, Corollaries 1–2) concerns the welfare of lowest-income families, defined as those with budget (maximum willingness to pay for housing) equal to zero or sufficiently close to zero. Under two jointly sufficient conditions — (1) neighborhoods that are underdemanded (zero-priced) under NA remain underdemanded under DA/DN, and (2) the schools in those underdemanded neighborhoods are themselves underdemanded — both DA and DN generate weakly higher welfare for the lowest-income families than NA. These conditions hold whenever families share common ordinal preference rankings (Corollary 1) and in the uniform economy where each valuation profile is equally likely (Corollary 2). The conditions are shown to be approximately necessary in a robustness sense (Theorem 6): for any economy violating them, an arbitrarily close economy exists in which a positive measure of zero-income families prefer NA. In simulations, DN raises lowest-income welfare by an average of 26.51% and DA by an average of 38.25% relative to NA.&lt;/p&gt;
&lt;p&gt;The paper also proves existence of a competitive equilibrium for the continuum economy under DA and DN via the Schauder-Tychonoff fixed-point theorem (Theorem 2), exploiting the continuity of school assignment probabilities in families&amp;rsquo; neighborhood choices. In discrete economies, assignment externalities can preclude equilibrium existence, but approximate equilibria exist in sufficiently large discrete markets and all welfare comparisons carry over approximately. The existence proof technique applies to general assignment games with externalities including peer preferences and complementarities.&lt;/p&gt;
&lt;p&gt;Scope conditions: results are derived for a model without direct peer externalities or endogenous school quality; a supplementary extension to local public financing finds that the aggregate welfare superiority of DA over NA may not survive when school spending is capitalized into housing prices, though the lowest-income welfare sufficiency conditions of Theorem 5 do extend to that environment.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question and why does the housing market matter for evaluating school choice?&lt;/p&gt;
&lt;p&gt;A: The paper asks how replacing neighborhood assignment with the Deferred Acceptance mechanism affects aggregate welfare and the welfare of the lowest-income families, accounting for the fact that families choose where to live in response to the school assignment mechanism. The housing market matters because under neighborhood assignment families can guarantee enrollment at a preferred school by purchasing a house in that school&amp;rsquo;s neighborhood; switching to DA changes these strategic incentives, alters equilibrium prices, and therefore changes who ends up in which neighborhood before any school assignment takes place. Ignoring residential choices would miss this feedback loop between assignment rules and housing demand.&lt;/p&gt;
&lt;p&gt;Q: What are the three mechanisms compared, and how do they differ?&lt;/p&gt;
&lt;p&gt;A: Neighborhood assignment (NA) assigns each child to the school in their neighborhood with certainty. DA without neighborhood priority allocates seats by student preference rankings and lottery numbers, with market-clearing cutoffs determined iteratively; no residential location confers a priority advantage. DN (DA with neighborhood priority) works like DA but grants neighborhood residents a priority of 1 at their local school and 0 at all other schools, effectively guaranteeing neighborhood families a seat at their local school while filling remaining seats by lottery among non-neighborhood applicants.&lt;/p&gt;
&lt;p&gt;Q: What does Theorem 3 establish, and what is the intuition for why DN dominates NA in aggregate welfare?&lt;/p&gt;
&lt;p&gt;A: Theorem 3 establishes that for any competitive equilibrium under DN and any competitive equilibrium under NA, aggregate welfare is weakly higher under DN. The intuition is that DN preserves all options available under NA — a family can always choose the neighborhood corresponding to its most-valued school and be guaranteed admission there — while additionally providing access to seats at other schools not claimed by their own neighborhood residents. The proof maps DN&amp;rsquo;s CE onto a Walrasian equilibrium of a continuum assignment game and invokes the welfare-maximization property of such equilibria from Gretsky, Ostroy, and Zame (1992).&lt;/p&gt;
&lt;p&gt;Q: Why is the welfare comparison between DA and NA ambiguous?&lt;/p&gt;
&lt;p&gt;A: Under NA, families with the highest cardinal valuations for a particular school can guarantee admission by purchasing a house in that neighborhood, and this targeted sorting can raise aggregate welfare when preferences over schools are strongly aligned. Under DA (without neighborhood priority), no location guarantees school admission, so families lose this signaling device; but DA allows families to live in preferred neighborhoods without sacrificing school quality, which raises welfare when preferences are heterogeneous. Neither effect dominates in general: in simulations, DA ranges from −18.26% to +5.65% relative to NA across the parameter space.&lt;/p&gt;
&lt;p&gt;Q: What role do neighborhood priorities play as a &amp;ldquo;signaling device,&amp;rdquo; and when does DN dominate DA?&lt;/p&gt;
&lt;p&gt;A: Neighborhood priorities allow families to credibly signal high valuations for a school by choosing to live in that school&amp;rsquo;s neighborhood, analogously to signaling devices in matching markets without money. When families have identical ordinal preference rankings over neighborhoods and schools (Assumptions 1 or 2), DN generates weakly higher aggregate welfare than DA because any DA assignment probability can be replicated under DN by mixing over neighborhoods, but the converse is not true. Counterexamples exist when preference rankings differ across families, so the DN-over-DA dominance is not universal.&lt;/p&gt;
&lt;p&gt;Q: What are the sufficient conditions for lowest-income families to prefer DA/DN to NA, and how tight are they?&lt;/p&gt;
&lt;p&gt;A: The two joint conditions are: (1) neighborhoods that have zero price (are underdemanded) under NA also have zero price under DA or DN after the mechanism switch; and (2) the schools located in those underdemanded neighborhoods are themselves underdemanded (have zero admission cutoffs) under DA/DN. Condition (1) reflects that the poorest neighborhoods are unlikely to become highly sought-after merely because the assignment mechanism changed. Condition (2) is consistent with the empirical finding of Owens and Candipan (2019) that in large US metropolitan areas the poorest neighborhoods typically have underperforming schools. Theorem 6 shows these conditions are approximately necessary: any economy violating them is arbitrarily close to one where a positive measure of zero-budget families prefer NA, so robustness requires them.&lt;/p&gt;
&lt;p&gt;Q: What do the simulations show about the magnitude of welfare effects for lowest-income families?&lt;/p&gt;
&lt;p&gt;A: In simulations with 10 lowest-income families (budgets of 0.05) among 1,000 total, DN raises lowest-income welfare by an average of 26.51% relative to NA and DA raises it by an average of 38.25% relative to NA, across the 18 parameter configurations. The gains are larger when preferences for neighborhoods and schools are less correlated (lower α) and when school capacities are more uniform (higher γ). DA consistently outperforms DN for lowest-income families in the simulations, even though DN dominates NA in aggregate welfare more reliably.&lt;/p&gt;
&lt;p&gt;Q: How does the paper handle equilibrium existence given the externalities created by residential choices?&lt;/p&gt;
&lt;p&gt;A: Because a family&amp;rsquo;s expected utility from a neighborhood depends on other families&amp;rsquo; neighborhood choices (through their effect on school assignment probabilities), standard existence results for assignment games do not directly apply. For the continuum economy, the author proves that school assignment probabilities under DA/DN are equicontinuous in families&amp;rsquo; neighborhood choices, which enables application of the Schauder-Tychonoff fixed-point theorem to guarantee the existence of a competitive equilibrium (Theorem 2). In finite discrete economies, assignment externalities can prevent equilibrium existence (illustrated by an example in Appendix B), but approximate equilibria exist for sufficiently large discrete markets, and all welfare comparisons hold approximately.&lt;/p&gt;
&lt;p&gt;Q: How does the paper&amp;rsquo;s model relate to and extend prior theoretical work on school choice and welfare?&lt;/p&gt;
&lt;p&gt;A: Prior theoretical work (e.g., Calsamiglia et al. 2015; Xu 2019; Avery and Pathak 2020) uses stylized models with single-parameter family types, identical ordinal school rankings, supermodular valuations, and no preferences over neighborhoods. This paper allows an unrestricted preference domain — families have arbitrary valuations over all neighborhood–school pairs — which generates novel findings: in the general model, lowest-income families do not necessarily benefit from DA (contrary to Calsamiglia et al. and Xu), aggregate welfare comparisons between DA and NA are ambiguous (whereas they are trivially resolved in the special cases of prior work), and neighborhood priorities can be welfare-improving even relative to DA without priorities.&lt;/p&gt;
&lt;p&gt;Q: Does the paper address the extension to endogenous school quality or local public financing?&lt;/p&gt;
&lt;p&gt;A: In Supplementary Appendix B, the model is extended to allow school spending to be financed by local property taxes, making school quality endogenous to neighborhood housing values. In that environment, the aggregate welfare superiority of DA/DN over NA may not hold: DA attracts non-neighborhood applicants to high-priced neighborhoods, and if those schools are a poor match for those applicants absent the spending, social welfare may fall — a result analogous to Barseghyan et al. (2013). However, the paper reports that the sufficiency conditions for lowest-income family welfare comparisons (Theorem 5) do extend to the local public financing environment, preserving the distributional results.&lt;/p&gt;
&lt;p&gt;Q: What does the paper say about alternative mechanisms such as Immediate Acceptance (Boston mechanism) and Top Trading Cycles?&lt;/p&gt;
&lt;p&gt;A: The Supplementary Appendix studies these alternatives. For Immediate Acceptance (IA), the paper shows that when there are neighborhood priorities, lowest-income families may prefer DA to IA, echoing the finding that IA is not strategyproof and may disproportionately hurt low-income families who are worse at gaming the system or have worse outside options (Pathak and Sonmez 2008; Calsamiglia et al. 2015). Top Trading Cycles and further extensions are also analyzed in the Supplementary Appendix, though detailed results are not developed in the main text.&lt;/p&gt;
&lt;p&gt;Neighborhood Assignment (NA): The baseline mechanism under which each family&amp;rsquo;s child is automatically enrolled in the school located in their chosen residential neighborhood, with no option to attend schools outside that neighborhood.&lt;/p&gt;
&lt;p&gt;Deferred Acceptance without Neighborhood Priority (DA): A strategyproof centralized assignment mechanism in which seats are allocated by families&amp;rsquo; stated preference rankings and lottery numbers via market-clearing admission cutoffs; residential location confers no priority advantage at any school.&lt;/p&gt;
&lt;p&gt;Deferred Acceptance with Neighborhood Priority (DN): A version of DA in which families residing in a neighborhood receive priority 1 at their neighborhood school and priority 0 at all other schools, guaranteeing neighborhood residents a seat at their local school before remaining seats are allocated by lottery to non-neighborhood applicants.&lt;/p&gt;
&lt;p&gt;Competitive Equilibrium (CE): A pair of neighborhood choices and a price vector such that (1) each family optimally selects the neighborhood maximizing expected utility net of price (subject to budget), (2) neighborhood capacities are not exceeded, and (3) neighborhoods with excess capacity are priced at zero.&lt;/p&gt;
&lt;p&gt;Underdemanded Neighborhood/School: A neighborhood whose equilibrium price is zero (excess housing supply) or a school whose admission cutoff is zero (excess capacity), meaning any applicant who lists it can gain admission.&lt;/p&gt;
&lt;p&gt;Assignment Externality: The indirect dependence of a family&amp;rsquo;s expected utility on other families&amp;rsquo; neighborhood choices, which operates through the effect of the population distribution across neighborhoods on the family&amp;rsquo;s school assignment probabilities under DA or DN. This externality can preclude competitive equilibrium existence in discrete economies.&lt;/p&gt;
&lt;p&gt;Aggregate Welfare: The utilitarian sum of all families&amp;rsquo; expected utilities from their neighborhood–school assignments, not netting out neighborhood prices (so it includes the welfare of house sellers as passive agents); the comparison criterion for Theorems 3 and 4.&lt;/p&gt;
&lt;p&gt;Signaling Device (neighborhood priority as): The interpretation that neighborhood priorities allow families to credibly reveal high valuations for a school by choosing to live in that school&amp;rsquo;s neighborhood, analogously to signaling instruments in matching markets without monetary transfers; the mechanism through which DN can improve welfare relative to DA.&lt;/p&gt;</description></item><item><title>Screening and Segmenting: A Consumer Surplus Perspective</title><link>https://macropaperwarehouse.com/papers/screening-and-segmenting-a-consumer-surplus-perspective/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/screening-and-segmenting-a-consumer-surplus-perspective/</guid><description>&lt;p&gt;Bergemann, Heumann, and Wang study consumer surplus when a monopolist simultaneously engages in second-degree price discrimination (screening consumers within each market segment through quality-differentiated menus) and third-degree price discrimination (offering different menus across segments). The central question is which market segmentation maximizes aggregate consumer surplus, and under what conditions any segmentation benefits consumers at all.&lt;/p&gt;
&lt;p&gt;The model features a monopolist selling vertically differentiated goods of quality q at strictly convex cost c(q) to a continuum of buyers with privately known values v drawn from an aggregate market m*. A segmentation is any decomposition of m* into submarkets, each receiving a profit-maximizing screening menu. The seller observes segment identity but not individual values. The problem of finding the consumer-optimal segmentation is, on its face, an optimization over distributions of distributions — an infinite-dimensional object.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central methodological contribution is a dramatic dimensional reduction. Theorem 1 establishes that the maximum consumer surplus achievable by any segmentation equals the maximum of the expected local information rent, u(v,h) = h·Q(v−h), over all inverse hazard rate functions h satisfying a majorization constraint h ≺ h* (where h* is the aggregate market&amp;rsquo;s inverse hazard rate). The local information rent captures both the extensive margin (h measures the mass of higher-value buyers per unit of value-v buyers who earn rent from v&amp;rsquo;s allocation) and the intensive margin (Q(v−h) is the quality allocated to value v, decreasing in h as distortion increases). The two margins trade off: raising h widens the base of rent-earning buyers but worsens allocative distortion, making u(v,h) hump-shaped in h with an interior maximizer h̄(v).&lt;/p&gt;
&lt;p&gt;The consumer-optimal segmentation has a striking structural property: every buyer of a given value v receives the same quality in every segment in which they appear, even though the monopolist could in principle offer different qualities across segments. Prices, however, differ across segments for identical buyers. This holds because the optimal segmentation is always a uniform segmentation — one in which the inverse hazard rate hm(v) is equalized across all segments containing value v.&lt;/p&gt;
&lt;p&gt;Under log-concavity of both aggregate demand (equivalently, a non-increasing aggregate inverse hazard rate h*(v), satisfied by uniform, normal, logistic, and exponential distributions) and the supply function Q(v) (equivalent to c&amp;rsquo;&amp;rsquo;&amp;rsquo;(q)q/c&amp;rsquo;&amp;rsquo;(q) ≥ −1, satisfied by all power cost functions), the optimal segmentation takes a transparent two-regime form (Proposition 3): for values below a threshold v̂ where h*(v̂) = h̄(v̂), the inverse hazard rate is reduced to h̄(v) by concentrating low-value buyers; for values above v̂, the aggregate market is left unchanged. The resulting segments are nested convex intervals [vm, v̄], all sharing the same upper bound v̄, with pricing differing across segments only by a quality-independent base price Tm that increases with vm (Theorem 2).&lt;/p&gt;
&lt;p&gt;Corollary 3 delivers the sharpest policy-relevant finding: under log-concave demand and supply, zero segmentation is optimal — any segmentation harms consumers — if and only if h*(v̲) ≤ h̄(v̲) at the lowest value v̲. For iso-elastic costs c(q) = q^γ/γ (γ &amp;gt; 1), this becomes η*(v̲) ≤ γ/(1−γ), where η*(v̲) is the aggregate demand elasticity at the bottom of the distribution. When demand is sufficiently elastic relative to supply, the monopolist&amp;rsquo;s screening already provides near-optimal consumer rents and no redistribution of buyers across segments can improve them. More elastic supply (lower γ) shrinks the set of markets where zero segmentation is optimal (Proposition 4, Zγ&amp;rsquo; ⊂ Zγ for γ&amp;rsquo; &amp;lt; γ); more inelastic supply (higher γ) expands it, and in the limit γ → ∞ zero segmentation is suboptimal only when the aggregate allocation itself is efficient.&lt;/p&gt;
&lt;p&gt;For iso-elastic costs, the optimal segmentation assigns each segment a Pareto distribution below v̂ with shape parameter α = γ/(γ−1), and the aggregate market above v̂ (Corollary 1). Each segment&amp;rsquo;s demand elasticity equals the constant γ/(1−γ) below v̂ and the aggregate elasticity above (Corollary 2): the supply elasticity 1/(γ−1) determines how elastic demand must be made within segments to counteract monopoly distortions. The paper also extends the framework to adverse selection (where seller cost rises with buyer type), with the full reduction to inverse hazard rate optimization preserved when the rate of increase in adverse selection satisfies τ&amp;rsquo;&amp;rsquo;(v)v/τ&amp;rsquo;(v) ∈ [0,1] (Proposition 5).&lt;/p&gt;
&lt;p&gt;Q: What is the local information rent and why is it central?
A: The local information rent is u(v,h) = h·Q(v−h), where h is the inverse hazard rate at value v and Q is the inverse marginal cost (supply) function (equation 9). The factor h captures the extensive margin — the mass of higher-value buyers per unit of value-v buyers who earn rent from v&amp;rsquo;s quality allocation — while Q(v−h) captures the intensive margin — the quality allocated to v via the virtual value v−h, which falls as h rises. Because u is hump-shaped in h, there is an interior rent-maximizing inverse hazard rate h̄(v) for each value. Lemma 2 establishes that in every regular market, total consumer surplus equals the integral of u(v,hm(v))dFm(v), so the entire segmentation problem reduces to choosing h.&lt;/p&gt;
&lt;p&gt;Q: What is the majorization constraint and what does it exactly characterize?
A: The majorization constraint h ≺ h* requires that for all v ∈ V, the integral from v̲ to v of [h*(t) − h(t)]dF*(t) ≥ 0 (equation 18). Proposition 1 shows that for any segmentation σ, the average inverse hazard rate hσ must satisfy hσ ≺ h*. A partial converse holds: given h ≺ h* under regularity conditions, a uniform segmentation implementing h exists. The constraint is strictly weaker than the pointwise bound h ≤ h* available in the binary case because it permits h to exceed h* at some values (dilution) provided it falls sufficiently below h* at higher values (concentration) to maintain the cumulative inequality.&lt;/p&gt;
&lt;p&gt;Q: What are concentration and dilution, and how do they interact?
A: Concentration gathers buyers of a given value into fewer segments, lowering their inverse hazard rate below h*(v). Dilution raises the inverse hazard rate of value v by placing v in segments where immediately higher values are missing — creating gaps in the support — thereby increasing the support increment Δm(v) and hence hm(v) (equation 12). Dilution at v requires that values just above v have already been concentrated elsewhere to create the gaps; concentration thus enables dilution, linking the two tools. With only binary values, only concentration is available; with a continuum, dilution can strictly expand achievable consumer surplus by permitting h to exceed h* at low values.&lt;/p&gt;
&lt;p&gt;Q: What does Theorem 1 establish and why is it a major simplification?
A: Theorem 1 states that the maximum consumer surplus over all segmentations of m* equals the maximum of ∫u(v,h(v))dF*(v) over all h satisfying the majorization constraint h ≺ h* (equation 25). The original problem maximizes over distributions on the infinite-dimensional space of probability measures on V; the reduced problem is a standard optimal control problem over a single real-valued function h: V → R+, amenable to Karush-Kuhn-Tucker methods and often yielding closed-form solutions. Furthermore, every optimal segmentation is a uniform segmentation implementing some h solving the reduced problem, so the reduction is exact. The optimal h always satisfies regularity (h&amp;rsquo;(v) ≤ 1), meaning v − h(v) is non-decreasing, which ensures segments in the optimal uniform segmentation are themselves regular.&lt;/p&gt;
&lt;p&gt;Q: What is the structural property of consumer-optimal segmentations regarding quality across segments?
A: In any consumer-optimal segmentation, every buyer of value v receives the same quality in every segment in which they appear (the uniform quality property following from Theorem 1). This holds because the optimal inverse hazard rate h(v) is equalized across segments (uniform segmentation), and quality in a regular market is qm(v) = Q(v − hm(v)), which depends on the market only through hm(v). Prices, however, differ across segments for identical buyers: the monopolist does not redesign its product line across segments but adjusts only quality-independent base prices. This is counterintuitive because nothing in the monopolist&amp;rsquo;s problem requires quality uniformity — it emerges purely from the consumer surplus maximization.&lt;/p&gt;
&lt;p&gt;Q: What conditions guarantee the simple two-regime convex segmentation structure?
A: Log-concavity of aggregate demand — equivalently, h*(v) non-increasing in v, satisfied by uniform, normal, logistic, and exponential families — and log-concavity of the supply function Q(v), equivalent to c&amp;rsquo;&amp;rsquo;&amp;rsquo;(q)q/c&amp;rsquo;&amp;rsquo;(q) ≥ −1, together guarantee the structure of Proposition 3 and Theorem 2. Under these conditions, h̄(v) is strictly increasing in v (log-concave supply) while h*(v) is decreasing (log-concave demand), so they cross exactly once at v̂. The optimal h equals h̄(v) below v̂ and h*(v) above. Only concentration (not dilution) is ever used because log-concave supply makes u concave in h and log-concave demand ensures monotone ordering of marginal local information rents across values, so the binding majorization constraint becomes the pointwise constraint at the bottom.&lt;/p&gt;
&lt;p&gt;Q: What is the structure of convex segmentations and their menus (Theorem 2)?
A: Under log-concave demand and supply, the consumer-optimal segmentation consists of segments m with absolutely continuous supports [vm, v̄] for varying lower bounds vm ≤ v̂, all sharing the same upper bound v̄ (Part 1 of Theorem 2). Pricing across these segments differs only by a quality-independent base price Tm that is increasing in vm — more concentrated segments (lower vm) face a lower base price and carry higher information rents — while the quality menu p(q) is uniform across segments (Part 2). Equivalently, the monopolist offers nested menus all sharing the same efficient upper bound quality Q(v̄), differing in how far down the menu is extended and in the price of the lowest offered quality.&lt;/p&gt;
&lt;p&gt;Q: What do Corollaries 1 and 2 say for iso-elastic cost functions?
A: With iso-elastic cost c(q) = q^γ/γ (γ &amp;gt; 1) and log-concave demand, the consumer-optimal segmentation assigns each segment a Pareto distribution with shape parameter α = γ/(γ−1) below the threshold v̂, and the aggregate distribution above v̂ (Corollary 1). This delivers a constant demand elasticity of γ/(1−γ) within each segment below v̂, matching the aggregate market&amp;rsquo;s elasticity above v̂ (Corollary 2). The Pareto shape — and thus the degree of demand manipulation — is determined entirely by the supply elasticity 1/(γ−1): more elastic supply (lower γ) mandates a higher shape parameter α and more elastic within-segment demand to counteract larger monopoly distortions.&lt;/p&gt;
&lt;p&gt;Q: When is zero segmentation optimal, and what is the precise elasticity condition?
A: Under log-concave demand and supply, zero segmentation is optimal if and only if h*(v̲) ≤ h̄(v̲) — the aggregate inverse hazard rate at the lowest value already lies at or below its rent-maximizing level (Corollary 3). Since h* is decreasing under log-concavity, this condition at v̲ implies it holds everywhere, so the designer cannot improve rents at any value. For iso-elastic cost, the condition becomes η*(v̲) ≤ γ/(1−γ): aggregate demand elasticity at the bottom must be at least as large in magnitude as one plus the supply elasticity. For a Pareto aggregate distribution with shape parameter α, zero segmentation is optimal when α ≥ γ/(γ−1).&lt;/p&gt;
&lt;p&gt;Q: How does supply elasticity govern the scope for beneficial segmentation (Proposition 4)?
A: Proposition 4 establishes that for iso-elastic cost, the set of markets Zγ where zero segmentation is optimal is strictly nested increasing in γ: for any γ&amp;rsquo; &amp;lt; γ, Zγ&amp;rsquo; ⊂ Zγ. More elastic supply (lower γ) amplifies monopoly distortions and enlarges the set of markets where segmentation benefits consumers; more inelastic supply (higher γ) makes quality provision rigid, reducing segmentation&amp;rsquo;s scope. In the limit γ → ∞ (approaching unit demand), zero segmentation is suboptimal only if the aggregate allocation is already efficient — but this limit also means very inelastic supply, so the potential benefits from segmentation have shrunk toward zero simultaneously.&lt;/p&gt;
&lt;p&gt;Q: How does this paper compare to and depart from Haghpanah and Siegel (2023)?
A: Haghpanah and Siegel (2023) showed that in generic markets with a finite number of goods, some segmentation always improves consumer surplus relative to the aggregate market. This paper shows that with a continuum of qualities, this universal improvement result fails: Corollary 3 identifies a large, non-degenerate class of markets satisfying Haghpanah and Siegel&amp;rsquo;s genericity conditions where zero segmentation is optimal for consumers. The discrepancy arises because the log-concave supply condition (equation 27) is violated in finite-good environments — Haghpanah and Siegel explicitly provide a counterexample showing their result fails with a continuum of goods. This paper characterizes exactly when the finite-good gains vanish as the quality space becomes continuous, providing the precise elasticity conditions.&lt;/p&gt;
&lt;p&gt;Q: What changes and what is preserved when extending to adverse selection?
A: In the adverse selection specification, buyer net value v is private and the seller&amp;rsquo;s cost per unit is τ(v) − v, increasing in v when τ&amp;rsquo;(v) &amp;gt; 1. The local information rent becomes w(v,h) = u(v, τ&amp;rsquo;(v)·h), where adverse selection enters by amplifying the effective inverse hazard rate by τ&amp;rsquo;(v) (equation 40). Proposition 5 confirms that the full reduction to majorization-constrained optimization over h goes through, and the optimal segmentation features more elastic within-segment demand when adverse selection is more severe. The reduction requires τ&amp;rsquo;&amp;rsquo;(v)v/τ&amp;rsquo;(v) ∈ [0,1] (equation 39), bounding the rate of increase of adverse selection severity; if this fails, the key inequality (35) driving the optimality of uniform segmentations may break down.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications for regulation of price discrimination?
A: The results imply that blanket restrictions on market segmentation may harm consumers by preventing welfare-enhancing price discrimination in markets where demand is sufficiently inelastic relative to supply (the region outside the zero-segmentation condition). In markets satisfying η*(v̲) ≤ γ/(1−γ), allowing segmentation yields no consumer benefit, so restrictions are harmless to consumers. The key policy-relevant primitives are demand and supply elasticities, which are in principle measurable. The findings also imply that the welfare effects of data-driven personalized pricing depend critically on the interaction between consumer heterogeneity (demand shape) and cost structure (supply elasticity), rather than on the degree of segmentation per se.&lt;/p&gt;
&lt;p&gt;Local information rent: u(v,h) = h·Q(v−h), the total consumer surplus generated per unit mass of buyers at value v as a function of the inverse hazard rate h. The factor h is the extensive margin (mass of higher-value buyers per unit of value-v buyers who earn rent) and Q(v−h) is the intensive margin (quality allocated to v via the virtual value v−h). It is hump-shaped in h with interior maximizer h̄(v), and the segmentation problem reduces entirely to maximizing its expectation.&lt;/p&gt;
&lt;p&gt;Inverse hazard rate hm(v): in a continuous market, (1−Fm(v))/fm(v); generalized to accommodate atoms and support gaps (equation 12). It simultaneously determines the virtual value ϕm(v) = v − hm(v) (governing allocative distortion) and the scaled mass of higher-value buyers per unit of value-v buyers (governing the extensive margin of rents). The dual role requires both a continuum of qualities and endogenous segmentation.&lt;/p&gt;
&lt;p&gt;Majorization constraint h ≺ h*: for all v, the cumulative integral of [h*(t)−h(t)]dF*(t) from v̲ to v is non-negative (equation 18). It is the exact characterization of inverse hazard rate functions achievable by some segmentation of m*, strictly weaker than the pointwise bound h ≤ h* of the binary case because it permits h to exceed h* at some values (dilution) provided it falls sufficiently below h* at higher values (concentration).&lt;/p&gt;
&lt;p&gt;Uniform segmentation: a segmentation in which every buyer of value v faces the same inverse hazard rate hm(v) = hσ(v) in every segment containing v (equation 22). Theorem 1 establishes that every consumer-optimal segmentation is uniform; this class converts the double integral over segments and values into a single integral against F*, enabling the dimensional reduction of Theorem 1.&lt;/p&gt;
&lt;p&gt;Concentration and dilution: the two tools by which segmentation modifies inverse hazard rates. Concentration gathers buyers of a given value into fewer segments, lowering hm(v) below h*(v). Dilution raises hm(v) above h*(v) by placing value v in segments where immediately higher values are absent, creating support gaps. Dilution requires prior concentration of adjacent higher values, so the two tools are linked; under log-concave demand and supply, only concentration is used in the optimal segmentation.&lt;/p&gt;
&lt;p&gt;Convex segmentation: a segmentation whose constituent segments have nested convex interval supports [vm, v̄] all sharing the same upper bound v̄, with varying lower bounds vm. This is the consumer-optimal structure under log-concave demand and supply (Theorem 2). For iso-elastic cost, each segment below the threshold v̂ corresponds to a Pareto distribution with shape parameter α = γ/(γ−1) determined by cost convexity γ.&lt;/p&gt;
&lt;p&gt;Zero-segmentation condition: the condition under which no segmentation can improve consumer surplus over the aggregate market. Under log-concave demand and supply with iso-elastic cost c(q) = q^γ/γ, it is η*(v̲) ≤ γ/(1−γ): aggregate demand elasticity at the lowest value must be at least as large in magnitude as one plus the supply elasticity (Corollary 3). When this holds, any redistribution of buyers across segments strictly reduces consumer surplus.&lt;/p&gt;</description></item><item><title>Should Monetary Policy Care about Redistribution? Optimal Monetary and Fiscal Policy with Heterogeneous Agents</title><link>https://macropaperwarehouse.com/papers/should-monetary-policy-care-about-redistribution-optimal-monetary-and-fiscal-policy-with-heterogeneous-agents/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/should-monetary-policy-care-about-redistribution-optimal-monetary-and-fiscal-policy-with-heterogeneous-agents/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Should monetary policy deviate from price stability to address redistributive concerns in an economy with heterogeneous agents? The paper jointly solves for optimal monetary and fiscal policy under commitment in a Heterogeneous Agent New Keynesian (HANK) environment with incomplete insurance markets for idiosyncratic risk, nominal frictions (Rotemberg price adjustment costs), and aggregate technology shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Framework.&lt;/strong&gt; The model is a Bewley-style incomplete-markets economy populated by a continuum of agents who differ in their idiosyncratic labor productivity histories. Agents save in two assets — nominal public debt and real capital shares — and face nominal borrowing constraints. Intermediate firms operate under monopolistic competition and face quadratic price adjustment costs. The government has up to five fiscal instruments: linear taxes on real capital income, on nominal asset income, and on labor income; lump-sum transfers; and one-period public nominal debt. Monetary policy controls the path of the nominal interest rate, and thereby inflation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three fiscal regimes are analyzed:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regime 1 — Full optimal fiscal policy.&lt;/strong&gt; When both capital taxes (on real and nominal asset returns) and a labor tax are freely optimizable and time-varying, the paper proves analytically (Proposition 1) that optimal monetary policy implements exact price stability at all periods. The intuition is that linear capital taxes replicate all direct redistributive channels of inflation (return effects and Fisher effects), while the labor tax replicates all indirect general-equilibrium channels (real wage effects). Hence fiscal tools are sufficient substitutes for any redistributive role of inflation, and the Rotemberg price-adjustment loss makes any deviation from zero inflation strictly costly. This equivalence result extends Correia et al. (2008) to environments with heterogeneous asset holdings, capital, and both real and nominal assets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regime 2 — Exogenous fiscal rules (constant or modestly time-varying taxes).&lt;/strong&gt; Using a standard quarterly calibration for the US (capital tax 36%, labor tax 28%, transfers 8% of GDP; Frisch elasticity 0.5; price adjustment cost κ=100; TFP shock persistence 0.95, standard deviation 0.31% per quarter; wealth Gini 0.73), the paper solves for optimal inflation dynamics numerically via a &amp;ldquo;timeless perspective&amp;rdquo; — i.e., around the long-run equilibrium. Under Fiscal Rule 1 (constant marginal tax rates, debt-stabilizing transfer rule), the maximum change in the inflation rate following a one-standard-deviation negative TFP shock is &lt;strong&gt;0.01%&lt;/strong&gt;, and the annualized standard deviation of inflation is &lt;strong&gt;0.020%&lt;/strong&gt;. Under Fiscal Rule 2 (labor tax falls by 0.2 percentage points on impact from 28% to 27.8%, capital tax rises by 0.2 percentage points from 36% to 36.2%), inflation volatility is &lt;strong&gt;slightly lower&lt;/strong&gt; and aggregate consumption volatility is also reduced, confirming that even simple time-varying fiscal rules dominate optimal inflation as an insurance device. The aggregate welfare gain from implementing optimal inflation relative to constant inflation (Π=1) is &lt;strong&gt;0.002%&lt;/strong&gt; in consumption-equivalent terms, with the gain concentrated among low-productivity agents (up to 0.01%), while high-productivity agents who can self-insure experience a near-zero gain.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regime 3 — Constrained-optimal fiscal policy.&lt;/strong&gt; Holding the capital tax constant while optimizing over the labor tax (or vice versa), and calibrating Pareto weights via an inverse-optimal-taxation approach to match the observed US steady-state fiscal system, the paper finds that optimal inflation volatility remains small at a standard deviation of &lt;strong&gt;0.01%&lt;/strong&gt;, again confirming the dominance of fiscal over monetary instruments for redistribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Robustness.&lt;/strong&gt; A simple two-agent economy calibrated closer to Bhandari et al. (2021b) — with a steeper Phillips curve (κ=20, slope ~6%), higher IES (1/σ=1/2), and highly unequal profit distribution (parameter ν=10 so high-productivity agents receive nearly all profits) — generates an inflation response on impact of &lt;strong&gt;0.17%&lt;/strong&gt;. Introducing a countercyclical fiscal rule (even a simple one) in this more volatile calibration reduces optimal inflation volatility by one order of magnitude, from &lt;strong&gt;0.68% to 0.07%&lt;/strong&gt;, and the on-impact response from &lt;strong&gt;0.15% to less than 0.01%&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Methodological contribution.&lt;/strong&gt; The analysis relies on two innovations: (i) a Lagrangian approach adapted from Marcet and Marimon (2019) that introduces the concept of &amp;ldquo;net social value of liquidity&amp;rdquo; for each agent, greatly simplifying first-order conditions; and (ii) a truncation method (LeGrand and Ragot 2022a,c) that represents incomplete-market heterogeneity by grouping agents by their last N periods of idiosyncratic history (truncation length N=5, giving 727 active histories), yielding a finite state space tractable for optimal policy computation. Results are validated against the Reiter (2009) histogram method.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; The equivalence result holds with commitment, a timeless perspective, and requires one distinct tax instrument per asset class (a separate tax on nominal and real returns). It holds under general period utility (not only separable forms). The result does not hold if the nominal asset tax is constrained to equal the real capital tax, in which case inflation would partially substitute for the missing instrument. The quantitative findings on small optimal inflation volatility are specific to the timeless perspective; a time-0 problem can generate larger deviations due to the ability to surprise agents with an initial inflation jump.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-central-equivalence-result-and-under-what-exact-conditions-does-it-hold"&gt;Q1. What is the central equivalence result and under what exact conditions does it hold?&lt;/h3&gt;
&lt;p&gt;When the government has access to time-varying linear taxes on real capital income, on nominal asset income, and on labor income — in addition to lump-sum transfers and public debt — optimal monetary policy implements exact price stability (gross inflation Πt = 1 at all dates). The conditions are: Ramsey commitment, both real and nominal asset taxes available as distinct instruments, and the Rotemberg price adjustment friction. The equivalence holds in the timeless perspective and the time-0 perspective, and does not require separability of the utility function.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-availability-of-capital-and-labor-taxes-render-inflation-redundant-as-a-redistributive-tool"&gt;Q2. Why does the availability of capital and labor taxes render inflation redundant as a redistributive tool?&lt;/h3&gt;
&lt;p&gt;Monetary policy operates through five channels identified in the HANK literature: three direct channels (substitution effect on returns, Fisher effect on nominal assets, wealth effect from unhedged interest-rate exposure) and two indirect channels (general-equilibrium labor income effects, heterogeneous exposure to income variation). The real capital tax — by affecting returns on all savings proportionally — can replicate any allocation achievable through the direct channels. The labor tax — by creating a wedge between the firm&amp;rsquo;s marginal cost of labor and household labor income — can replicate any allocation achievable through the indirect channels. With both instruments available, inflation&amp;rsquo;s only remaining effect is to destroy resources via Rotemberg adjustment costs, so the planner optimally sets Πt = 1.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-net-social-value-of-liquidity-and-how-does-it-simplify-the-analysis"&gt;Q3. What is the &amp;ldquo;net social value of liquidity&amp;rdquo; and how does it simplify the analysis?&lt;/h3&gt;
&lt;p&gt;The net social value of liquidity for agent i at date t, ψ̂i,t = ψi,t − μt, equals the planner&amp;rsquo;s benefit from transferring one unit of consumption to agent i net of its fiscal cost. It combines the agent&amp;rsquo;s marginal utility of consumption with the planner&amp;rsquo;s internalization of effects on saving incentives (through real and nominal Euler equations) and on labor supply (through the labor Euler equation). Expressing the Ramsey first-order conditions in terms of ψ̂i,t reduces them to Euler-like smoothing conditions that closely parallel the individual agents&amp;rsquo; Euler equations, making both algebra and economic interpretation substantially more transparent.&lt;/p&gt;
&lt;h3 id="q4-how-large-is-the-optimal-inflation-response-in-the-baseline-quantitative-calibration-and-how-does-it-decompose"&gt;Q4. How large is the optimal inflation response in the baseline quantitative calibration, and how does it decompose?&lt;/h3&gt;
&lt;p&gt;Under the baseline US calibration (κ=100, quarterly period, standard fiscal rules with constant marginal tax rates), the optimal inflation response to a one-standard-deviation negative TFP shock reaches a maximum of 0.01% (ten basis points on an annualized basis or less). The annualized standard deviation of inflation is 0.020%. Inflation rises on impact and then declines back to steady state. The correlation of optimal inflation with output is 0.20, indicating mild countercyclicality. The difference in aggregate consumption volatility between the optimal-inflation economy (Economy 1) and the constant-inflation economy (Economy 2) is small; the std of consumption is 1.33% vs. 1.34% of the mean.&lt;/p&gt;
&lt;h3 id="q5-what-welfare-gains-does-optimal-inflation-deliver-and-how-do-they-vary-across-the-productivity-distribution"&gt;Q5. What welfare gains does optimal inflation deliver, and how do they vary across the productivity distribution?&lt;/h3&gt;
&lt;p&gt;The average welfare gain from implementing optimal inflation relative to constant inflation (Π=1) is 0.002% in consumption-equivalent terms. This aggregate figure conceals heterogeneity: low-productivity agents experience a welfare gain of up to 0.01% because they benefit disproportionately from the reduction in consumption volatility (inflation acts as a partial Fisher-effect transfer to debtors who are credit-constrained). High-productivity agents experience a near-zero gain because they can self-insure through portfolio choice. All productivity groups experience a positive but modest welfare gain.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-effect-of-introducing-a-simple-time-varying-fiscal-rule-fiscal-rule-2-on-optimal-inflation-dynamics"&gt;Q6. What is the effect of introducing a simple time-varying fiscal rule (Fiscal Rule 2) on optimal inflation dynamics?&lt;/h3&gt;
&lt;p&gt;Fiscal Rule 2 sets the labor tax to fall from 28% to 27.8% on impact after a negative TFP shock (a decline of 0.2 percentage points), while the capital tax rises from 36% to 36.2%. The public debt path is roughly unchanged relative to Fiscal Rule 1. Compared to the constant-tax baseline, Fiscal Rule 2 yields slightly lower inflation volatility (standard deviation 0.018% vs. 0.020%) and lower aggregate consumption volatility (std 1.31% vs. 1.33% of mean). These results confirm that even a small, simple exogenous fiscal rule dominates inflation as an insurance device against aggregate TFP shocks.&lt;/p&gt;
&lt;h3 id="q7-under-what-calibration-does-the-optimal-inflation-response-become-quantitatively-sizable-and-how-does-a-fiscal-rule-affect-it-in-that-case"&gt;Q7. Under what calibration does the optimal inflation response become quantitatively sizable, and how does a fiscal rule affect it in that case?&lt;/h3&gt;
&lt;p&gt;A combination of a steep Phillips curve (κ=20 rather than 100, implying a slope of about 6% rather than 2%), a higher intertemporal elasticity of substitution (IES = 1/σ = 1/2 rather than 1), and highly unequal profit distribution (parameter ν=10, so high-productivity agents receive nearly all profits) generates an on-impact inflation response of approximately 0.15%–0.17% after a 1% negative TFP shock, and an inflation volatility of 0.68%. Introducing a countercyclical fiscal rule in this environment reduces inflation volatility by one order of magnitude to 0.07%, and the on-impact response from 0.15% to less than 0.01%, while also reducing aggregate consumption volatility.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-role-of-profit-distribution-in-determining-the-sign-and-magnitude-of-the-optimal-inflation-response"&gt;Q8. What is the role of profit distribution in determining the sign and magnitude of the optimal inflation response?&lt;/h3&gt;
&lt;p&gt;The distribution of firms&amp;rsquo; profits to households is a key driver of optimal inflation. When profits are distributed predominantly to high-productivity agents (ν=10), optimal inflation rises on impact after a negative TFP shock, because higher inflation benefits low-productivity credit-constrained agents through the Fisher effect and the real-wage channel. When profits are distributed equally across agents (ν=0), the optimal inflation response reverses sign and becomes negative on impact (−0.13% instead of +0.17%), because decreasing inflation raises firms&amp;rsquo; profits and, since those profits are equally shared, acts as a progressive transfer to credit-constrained low-income agents who consume a larger fraction at the margin.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-constrained-optimal-fiscal-policy-scenario-regime-3-affect-inflation-dynamics"&gt;Q9. How does the constrained-optimal fiscal policy scenario (Regime 3) affect inflation dynamics?&lt;/h3&gt;
&lt;p&gt;In Regime 3, a Pareto-weight social welfare function is calibrated via an inverse-optimal-taxation approach so that the observed US fiscal steady state (36% capital tax, 28% labor tax, 8% transfers/GDP) is an interior optimal. The planner then jointly optimizes either the labor tax path (holding capital tax constant) or the capital tax path (holding labor tax constant) together with the inflation path. The resulting optimal inflation standard deviation is 0.01%, confirming that even partial fiscal flexibility is sufficient to drive inflation volatility close to zero.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-timeless-perspective-differ-from-a-time-0-problem-in-generating-inflation-deviations"&gt;Q10. How does the timeless perspective differ from a time-0 problem in generating inflation deviations?&lt;/h3&gt;
&lt;p&gt;In a time-0 problem the planner can exploit initial surprise: at date 0, unexpected inflation can redistribute real wealth through the Fisher effect on pre-existing nominal debt holdings, a mechanism immune to the time-consistency constraint. This creates a larger initial inflation front-loading. In the timeless perspective — the paper&amp;rsquo;s main framework — the economy is assumed to have been running under the optimal commitment rule for a long time, so no such surprise mechanism is available, and the planner&amp;rsquo;s only inflationary tool is the recurrent business-cycle insurance motive. As a result, inflation volatility in the timeless perspective is substantially smaller than in a time-0 problem.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-truncation-method-and-how-does-the-paper-validate-its-accuracy"&gt;Q11. What is the truncation method and how does the paper validate its accuracy?&lt;/h3&gt;
&lt;p&gt;The truncation method (LeGrand and Ragot 2022a,c) groups agents by their last N periods of idiosyncratic productivity history, creating a finite state space. With N=5 and 5 idiosyncratic states, there are 5^5=3,125 possible histories, of which 727 have positive probability. A &amp;ldquo;refined&amp;rdquo; variant (LeGrand and Ragot 2022c) applies longer truncation lengths to more common histories while keeping total history count linear rather than exponential in Nmax. The paper sets Nmax=20 for the refined truncation as a robustness check and finds impulse responses and second-order moments nearly identical to the N=5 baseline. Results are also compared against the Reiter (2009) histogram method, showing close agreement in both impulse response functions and second-order moments.&lt;/p&gt;
&lt;h3 id="q12-how-does-the-paper-relate-to-the-equivalence-results-of-correia-et-al-2008"&gt;Q12. How does the paper relate to the equivalence results of Correia et al. (2008)?&lt;/h3&gt;
&lt;p&gt;Correia et al. (2008) show that in a representative-agent economy without capital, a time-varying consumption tax can implement price stability regardless of nominal frictions. The current paper extends this to an environment with heterogeneous asset holdings (both real and nominal), capital accumulation, and an incomplete insurance market. The extension requires one distinct tax instrument per asset class (separate taxes on nominal and real returns), rather than a single consumption tax. The equivalence result would break down if the nominal asset tax were forced to equal the real capital tax, because inflation would then be needed to partially substitute for the missing degree of freedom.&lt;/p&gt;
&lt;h3 id="q13-what-three-mechanisms-shape-the-optimal-inflation-first-order-condition-when-fiscal-policy-is-exogenous"&gt;Q13. What three mechanisms shape the optimal inflation first-order condition when fiscal policy is exogenous?&lt;/h3&gt;
&lt;p&gt;When tax rates follow exogenous fiscal rules, the planner&amp;rsquo;s first-order condition for inflation balances three forces: (1) the Rotemberg resource-destruction cost of price adjustment (μt·κ·(Πt−1)), which penalizes any deviation from Πt=1; (2) the ability to manipulate the real wage through the New-Keynesian Phillips curve (a term involving the lead and lag of the Phillips-curve multiplier γt), which can transfer resources across households; and (3) the gain from reducing the real interest payment on existing nominal public debt through unexpected inflation (a term involving fund multipliers Γt and Υt, scaled by the outstanding debt Bt−1). The balance among these three forces determines the sign and magnitude of the optimal inflation response.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Net Social Value of Liquidity (ψ̂i,t).&lt;/strong&gt; The planner&amp;rsquo;s benefit from transferring one unit of consumption to agent i net of its fiscal cost (μt). Formally ψ̂i,t = ψi,t − μt, where ψi,t captures the agent&amp;rsquo;s marginal utility of consumption adjusted for the planner&amp;rsquo;s internalization of savings distortions through real and nominal Euler equations and the labor supply equation. This concept is introduced in the paper to simplify Ramsey first-order conditions in incomplete-market environments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Equivalence Result (Proposition 1).&lt;/strong&gt; The theoretical finding that, when the government has access to time-varying linear taxes on both nominal and real asset returns and on labor income, the planner can exactly reproduce the flexible-price allocation and optimal monetary policy is to implement zero net inflation at all dates. The equivalence holds because the fiscal instruments can replicate every redistributive channel of monetary policy at no resource cost, while any inflation deviation destroys output through price adjustment costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Timeless Perspective.&lt;/strong&gt; A solution concept for Ramsey optimal policy in which the economy is assumed to have been operating under the optimal commitment rule for a long time, so initial conditions no longer matter. As described in the paper (following Woodford, 1999, and McCallum and Nelson, 2000), this is &amp;ldquo;the closest notion to optimal policy making according to a rule&amp;rdquo; and eliminates the time-0 front-loading bias that arises when the planner can surprise agents with an initial inflation jump.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Truncation Method.&lt;/strong&gt; A method (LeGrand and Ragot 2022a,c) that approximates the infinite-dimensional heterogeneous-agent state space by grouping agents by their last N periods of idiosyncratic productivity history. Within each truncated history, agents are pooled with history-specific heterogeneity parameters (ξh) capturing wealth dispersion from histories prior to the aggregation window. The refined variant assigns different truncation lengths to different histories to keep the total number of histories linear in Nmax rather than exponential.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Direct vs. Indirect Channels of Monetary Policy.&lt;/strong&gt; Following Kaplan et al. (2018) and Auclert (2019), the paper distinguishes: (i) direct channels — the substitution effect on real returns, the Fisher effect on nominal asset values, and the wealth effect from unhedged interest-rate exposure — which operate through changes in asset returns; and (ii) indirect channels — heterogeneous labor income effects and heterogeneous income exposure — which operate through general-equilibrium effects on wages and employment. The paper&amp;rsquo;s equivalence result shows that capital taxes replicate the direct channels and the labor tax replicates the indirect channels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fiscal Rule (Bohn-type, affine structure).&lt;/strong&gt; An exogenous rule specifying that marginal tax rates on capital and labor respond linearly to current and lagged TFP deviations from steady state, while transfers respond to TFP deviations and public debt deviations from target. The paper uses two such rules: Fiscal Rule 1 (constant marginal tax rates, debt-stabilizing transfer) and Fiscal Rule 2 (countercyclical labor tax and procyclical capital tax with the same debt path), to assess whether simple time-varying fiscal policies substitute for optimal inflation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rotemberg Price Adjustment Cost.&lt;/strong&gt; A quadratic cost κ/2·(pj,t/pj,t−1 − 1)^2·Yt incurred by each intermediate firm when it changes its price, used as the nominal friction generating the New-Keynesian Phillips curve. In the paper&amp;rsquo;s model, any deviation of gross inflation Πt from 1 destroys real output, making this the welfare cost of using inflation as a policy instrument.&lt;/p&gt;</description></item><item><title>Silence to Solidarity: How Communication About a Minority Affects Discrimination</title><link>https://macropaperwarehouse.com/papers/silence-to-solidarity-how-communication-about-a-minority-affects-discrimination/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/silence-to-solidarity-how-communication-about-a-minority-affects-discrimination/</guid><description>&lt;p&gt;This paper examines how two types of communication about a minority group affect discriminatory behavior: (i) horizontal communication between majority-group members, and (ii) top-down communication from agents of authority such as the legal system. The setting is urban Chennai, India, where the paper measures discrimination against thirunangai — a community of transgender women who are India&amp;rsquo;s most visible LGBTQ+ group — in a field experiment with 3,397 participants.&lt;/p&gt;
&lt;p&gt;Discrimination is measured using incentivized hiring choices. Participants are offered a free grocery delivery and make 10 binary choices over which worker will carry out the delivery, with worker gender (cisgender male, cisgender female, or transgender) varying across options. The stakes are real: one choice is randomly selected and implemented 2–9 weeks later. Participants in the control condition are highly discriminatory: they are 19 percentage points (32%) less likely to hire a transgender worker than a non-transgender worker (p&amp;lt;0.001), and are willing to sacrifice grocery items worth 1.9 times their median daily per capita food expenditure to avoid a 15-minute interaction with a transgender worker.&lt;/p&gt;
&lt;p&gt;The first main treatment involves randomly assigning participants to a 3-person group discussion with two neighbors, in which they discuss and make collective hiring choices over the same options. The key outcome is participants&amp;rsquo; subsequent private, individual hiring choices. The discussion eliminates anti-transgender discrimination on average: participants in the discussion arm are 17 percentage points (42%) more likely to select a transgender worker in their private post-discussion choices relative to the control group (p&amp;lt;0.001), so that discrimination is no longer statistically distinguishable from zero (p=0.30). The discussion&amp;rsquo;s effect is partially persistent: approximately one month later, discussion participants are still 4 percentage points more likely to select transgender workers in hypothetical hiring choices (p=0.03), representing roughly 25% of the short-run effect.&lt;/p&gt;
&lt;p&gt;The second main treatment cross-randomizes a video shown before hiring choices. The legal rights video informs participants of a Supreme Court ruling affirming that transgender people hold the same fundamental constitutional rights as other citizens. This reduces discrimination by 10.3 percentage points (p&amp;lt;0.001). A rights messaging video — which argues that transgender people should have equal rights without invoking legal authority — reduces discrimination by a smaller 5.8 percentage points (p=0.001), and there is some evidence the legal-authority version is more effective (p of difference in [0.01, 0.12]). However, the legal rights video&amp;rsquo;s effect is only 59% as large as the discussion&amp;rsquo;s effect (p of difference in [0.002, 0.04]), and it does not persist at the one-month follow-up (p in [0.12, 0.51]).&lt;/p&gt;
&lt;p&gt;The paper rules out two candidate mechanisms for the discussion&amp;rsquo;s effects and supports a third. First, the discussion does not work primarily through correcting misperceived norms: while control-group participants do overestimate peer discrimination by 5 percentage points, the discussion reduces predicted discrimination by 24 percentage points — far more than a corrected misperception could explain (at most 21% of the effect under generous assumptions). Second, the discussion does not work through virtue signaling alone: a &amp;ldquo;No discussion (public)&amp;rdquo; arm in which participants make individually-visible choices shows no reduction in discrimination on average (p=0.83). Third, the paper provides affirmative evidence for a persuasion channel: participants in a &amp;ldquo;listener&amp;rdquo; arm, who silently observe a 2-person discussion without participating, discriminate 13 percentage points less than the control group (p&amp;lt;0.001), an effect that is highly persistent at the 2–9 week follow-up (11 percentage points, p&amp;lt;0.001). The persuasion mechanism is further supported by the finding that pro-trans participants are more vocal: each additional transgender worker chosen in post-discussion private choices is associated with a 32% higher probability of speaking first (p=0.03) and a 27% higher probability of dominating the discussion (p=0.02). Statements about transgender workers during discussions were 5.7 times more likely to be positive than negative. Listeners who heard moral argumentation about equality, rights, and giving opportunities subsequently discriminated less (p&amp;lt;0.001).&lt;/p&gt;
&lt;p&gt;Scope conditions: the study is conducted among urban Chennai residents (85% female), where transgender identity is visually recognizable and socially salient, awareness of the 2014 Supreme Court ruling is low (36% could not identify a single legal right transgender people hold), and a wedge exists between descriptive norms (high actual discrimination) and prescriptive norms (93% of the control group rate explicit discrimination as wrong). The model&amp;rsquo;s &amp;ldquo;sweet spot&amp;rdquo; logic implies these effects may not generalize to settings where discrimination is either near-universal (no privately pro-trans individuals to be vocal) or already minimal (no incentive to persuade).&lt;/p&gt;
&lt;p&gt;Q: How is anti-transgender discrimination measured in the experiment?
A: Participants make 10 incentive-compatible binary hiring choices over grocery delivery workers, with one choice randomly selected and implemented 2–9 weeks later. Discrimination is defined as the reduction in the probability of selecting the alternative worker when that worker is transgender versus non-transgender, conditional on other option characteristics such as items offered and reliability score. Participants are told they will have a 15-minute conversation with the selected worker, ensuring anticipated social contact. The design is framed as market research to obfuscate the study&amp;rsquo;s purpose; only 8% correctly guessed the true focus.&lt;/p&gt;
&lt;p&gt;Q: How large is baseline discrimination in the control group?
A: In the No discussion (private) control condition, participants are 19 percentage points (32%) less likely to hire a transgender worker than a non-transgender worker (p&amp;lt;0.001). In willingness-to-pay terms, participants sacrifice grocery items worth 1.9 times their median daily per capita food expenditure (Rs. 127 on a base of Rs. 67) to avoid selecting a transgender worker. Even when a transgender worker dominates on both items and reliability score, participants in the control group still select the non-transgender worker 47% of the time.&lt;/p&gt;
&lt;p&gt;Q: What is the main effect of the 3-person group discussion on subsequent discrimination?
A: Participants who engage in a group discussion with two neighbors are 17 percentage points more likely to select a transgender worker in their subsequent private individual choices (p&amp;lt;0.001). This eliminates average discrimination entirely: in the discussion arm, the probability of selecting a transgender worker is not statistically distinguishable from the probability of selecting a non-transgender worker (p=0.30). The willingness-to-pay to avoid a transgender worker falls from Rs. 127 to Rs. 13 (p of difference &amp;lt; 0.001), and is no longer significantly different from zero (p=0.265).&lt;/p&gt;
&lt;p&gt;Q: How persistent are the effects of the group discussion?
A: At the 2–9 week follow-up survey (mean 35 days), discussion participants are approximately 4 percentage points more likely to select transgender workers in hypothetical hiring choices (p=0.03). This represents approximately 25% of the short-run 17 percentage point effect, a decay rate comparable to the persistence of US political advertising effects in the political science literature (Hill et al., 2013, estimate 10–15% remaining after 30 days).&lt;/p&gt;
&lt;p&gt;Q: What is the effect of the legal rights video, and how does it compare to the discussion?
A: The legal rights video — informing participants of the Supreme Court ruling affirming transgender people&amp;rsquo;s fundamental constitutional rights — increases the probability of selecting a transgender worker by 10.3 percentage points (p&amp;lt;0.001). The rights messaging video, which argues that transgender people should have equal rights without invoking legal authority, increases it by 5.8 percentage points (p=0.001). The legal rights video&amp;rsquo;s effect is only 59% as large as the discussion&amp;rsquo;s 17 percentage point effect (p of difference in [0.002, 0.04]), and unlike the discussion, neither video&amp;rsquo;s effect is detectable at the one-month follow-up (p in [0.12, 0.51]).&lt;/p&gt;
&lt;p&gt;Q: Does the legal rights video work through a different channel than the rights messaging video?
A: There is evidence that the legal authority of the Supreme Court matters beyond the content of the rights message. The legal rights video is more effective than the rights messaging video at reducing discrimination (p of difference in [0.01, 0.12]), and the legal rights video (but not the rights messaging) affects participants&amp;rsquo; beliefs about the legal status of transgender people (as measured by a summary index). Both videos shift perceived descriptive norms — participants predict others will select transgender workers more, by 2–6 percentage points — but neither significantly affects attitudes as measured by a list experiment or disapproval questions.&lt;/p&gt;
&lt;p&gt;Q: Does the discussion work through correcting misperceived norms?
A: This channel can account for at most a small fraction of the effect. Control-group participants do overestimate peer discrimination by 5 percentage points in incentivized predictions (p&amp;lt;0.001, as measured by predicted probability of selecting a transgender worker). However, the discussion reduces predicted discrimination by 24 percentage points (p&amp;lt;0.001), far exceeding the initial misperception. Even under generous assumptions in which the misperception is precisely corrected, this mechanism could account for no more than 21% of the discussion&amp;rsquo;s treatment effect (95% CI: [8.9%, 32.5%]).&lt;/p&gt;
&lt;p&gt;Q: Does the discussion work through virtue signaling?
A: The evidence rules out virtue signaling as the primary channel. The &amp;ldquo;No discussion (public)&amp;rdquo; treatment arm makes participants&amp;rsquo; individual hiring choices visible to their group members, exogenously increasing social image concerns in the absence of a discussion. This has no detectable average effect on discrimination (p=0.83), indicating that social image concerns alone — without the persuasive content of an actual discussion — do not explain the reduction in discrimination generated by the group discussion.&lt;/p&gt;
&lt;p&gt;Q: What is the evidence for the persuasion mechanism?
A: The &amp;ldquo;listener&amp;rdquo; treatment arm provides direct evidence. In this arm, one participant silently observes a 2-person discussion without speaking, then makes private individual choices. Listeners discriminate 13 percentage points less than the control group (p&amp;lt;0.001), an effect statistically indistinguishable from full discussion participants. Since listeners changed their behavior based solely on what they heard and saw, this constitutes evidence of persuasion. The listener effect is highly persistent at the 2–9 week follow-up (11 percentage points, p&amp;lt;0.001) and holds on a robustness outcome designed to be completely private. The implied persuasion rate is 29%, described as high relative to values in the literature (DellaVigna &amp;amp; Gentzkow, 2010).&lt;/p&gt;
&lt;p&gt;Q: Why do pro-trans participants persuade others — what drives the discussion&amp;rsquo;s content?
A: Pro-trans participants are disproportionately vocal. Each additional transgender worker chosen in post-discussion private choices (a proxy for pro-trans private attitudes) is associated with a 32% higher probability of speaking first (p=0.03) and a 27% higher probability of dominating the discussion (p=0.02), but only when discussing a choice involving a transgender worker. The overall tone of discussions is strongly pro-trans: statements about transgender workers are 5.7 times more likely to be positive than negative. Participants who hear moral argumentation about equality, rights, and giving opportunities subsequently discriminate significantly less (p&amp;lt;0.001).&lt;/p&gt;
&lt;p&gt;Q: Does the discussion work by changing statistical (belief-based) discrimination?
A: Partially, baseline discrimination in the control group is partly statistical: despite transgender workers having the same average reliability scores as others, participants rate them as less likely to complete a delivery, and revealing the true reliability score makes participants 2.9 percentage points more likely to select a transgender worker (an effect unique to transgender workers). However, the discussion does not significantly affect beliefs about transgender workers&amp;rsquo; reliability, and there is no detected reduction in the belief-based component of discrimination in the discussion arm (though the test is underpowered).&lt;/p&gt;
&lt;p&gt;Q: Are the effects of the discussion and the legal rights video additive?
A: The two interventions appear to combine approximately linearly for the legal rights video: there are no detected interaction effects (p in [0.83, 0.96]). By contrast, there is weak evidence of a negative interaction between the rights messaging video and the discussion, suggesting these two may be substitutes — consistent with the rights messaging video&amp;rsquo;s content being similar to the pro-trans moral argumentation already present in discussions.&lt;/p&gt;
&lt;p&gt;Q: What alternative explanations are ruled out?
A: The paper tests and finds no support for: (i) photo characteristics such as perceived caste driving results; (ii) social image concerns affecting even post-discussion private choices (the &amp;ldquo;extra private&amp;rdquo; robustness outcome designed to be unobservable by neighbors yields similar results); (iii) increased contemplation or deliberation about choices; (iv) experimenter demand effects or social desirability bias (treatment effects do not differ for the 8% who guessed the study&amp;rsquo;s purpose); (v) increased salience of the transgender category; and (vi) cheap talk from low stakes (choices were incentive-compatible and implemented).&lt;/p&gt;
&lt;p&gt;Q: What is the study&amp;rsquo;s theoretical model for why pro-trans participants speak out?
A: The paper develops a model combining social signaling (people want to fit in with their group; Bénabou &amp;amp; Tirole, 2006) with direct persuasion (participants can change each other&amp;rsquo;s preferences through messages). Under the right conditions, only pro-trans participants send persuasive pro-trans messages. This occurs in a &amp;ldquo;sweet spot&amp;rdquo; range: when average discrimination is not so strong that no one is privately pro-trans, and not so weak that pro-trans participants lack an incentive to persuade (since they are already in the majority). The context in Chennai — high actual discrimination but strong social norms against it — satisfies this sweet spot condition.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications regarding horizontal versus top-down communication?
A: In this context, facilitating horizontal communication between neighbors is a more effective tool for reducing discrimination than top-down communication about legal rights: the discussion&amp;rsquo;s effect is 1.7 times larger than the legal rights video (17 p.p. vs. 10.3 p.p.) and partially persists at one month, whereas the legal rights video&amp;rsquo;s effect does not persist. However, the legal rights video does reduce discrimination relative to the rights messaging video, suggesting that communicating the legal authority of the Supreme Court carries independent weight beyond rights advocacy messaging. Both interventions are complementary when combined.&lt;/p&gt;
&lt;p&gt;Horizontal communication: Communication between members of the majority group about a minority, as distinct from contact between majority and minority groups or top-down communication from authority. In this paper, operationalized as a group discussion among three neighbors who make collective hiring choices.&lt;/p&gt;
&lt;p&gt;Top-down communication: Communication from agents of authority — here, the legal system — about a minority group&amp;rsquo;s rights. Measured via a video informing participants of a Supreme Court ruling affirming transgender people&amp;rsquo;s constitutional rights.&lt;/p&gt;
&lt;p&gt;Anti-transgender discrimination: In the paper&amp;rsquo;s own measurement, the reduction in the probability that a worker is chosen because they are transgender (relative to being non-transgender), conditional on other delivery option characteristics. Measured in incentivized, privately-elicited binary hiring choices.&lt;/p&gt;
&lt;p&gt;Expressive law hypothesis: The theory that changes in the law affect behavior by changing people&amp;rsquo;s perception of the prevailing social norm, not (only) through deterrence. The paper tests this by comparing a legal rights video (invoking Supreme Court authority) to a rights messaging video with identical content but no legal backing, finding the legal-authority version more effective.&lt;/p&gt;
&lt;p&gt;Persuasion channel: The mechanism by which discussion participants change each other&amp;rsquo;s preferences through persuasive messages, particularly moral arguments about equality and rights. Distinguished in the paper from virtue signaling (publicly visible pro-trans behavior) and norm correction (updating misperceived beliefs about peer behavior).&lt;/p&gt;
&lt;p&gt;Pluralistic ignorance: A setting in which people misperceive how common discriminatory attitudes are among their peers, potentially hiding genuine minority support for the discriminated group. The paper tests this as a candidate mechanism and finds it can account for at most 21% of the discussion effect.&lt;/p&gt;
&lt;p&gt;Sweet spot condition: The range of average group discrimination levels in which pro-trans participants have both the motivation and opportunity to speak out persuasively — discrimination is not so universal that no one is privately pro-trans, and not so minimal that the pro-trans participants feel no need to persuade others. The paper argues the Chennai context satisfies this condition.&lt;/p&gt;</description></item><item><title>Skill-Replacing Technology and Bottom-Half Inequality</title><link>https://macropaperwarehouse.com/papers/skill-replacing-technology-and-bottom-half-inequality/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/skill-replacing-technology-and-bottom-half-inequality/</guid><description>&lt;p&gt;This paper proposes a model of skill-replacing routine-biased technological change (SR-RBTC) to explain patterns in U.S. bottom-half wage inequality that standard RBTC models cannot account for. The central departure from prior models (e.g., Acemoglu and Autor 2011; Cortes 2016) is that technology substitutes the usage of skill within routine occupations rather than replacing routine workers wholesale. Formally, SR-RBTC is characterized by epsilon &amp;lt; 0, where epsilon = d² log phi_R / (d theta_i d tau), meaning productivity gains in the routine occupation are disproportionately concentrated among lower-skilled workers, compressing skill-wage gradients within that occupation.&lt;/p&gt;
&lt;p&gt;The paper addresses three stylized facts that skill-neutral RBTC models leave unexplained. First, wage polarization concentrated around the median rather than the entire bottom half, even though routine workers are dispersed across the full bottom half of the wage distribution. SR-RBTC explains this because the largest wage drops accrue to the highest-skilled routine workers, who were empirically concentrated near the middle of the overall distribution. Second, the decline in middle wages stopped around 2000 even as routine employment continued falling. The model accounts for this through a two-phase mechanism: once the return to skill in routine occupations falls below that in manual occupations, the routine occupation attracts the lowest-skilled workers, shifting negative wage pressure to the bottom rather than the middle of the distribution. Third, average wages in routine occupations did not fall substantially despite large employment declines; in SR-RBTC, wage losses for higher-skilled routine workers are partially offset by gains for lower-skilled ones, leaving average routine wages relatively stable.&lt;/p&gt;
&lt;p&gt;The paper tests two new predictions using an Interactive Fixed-Effects Model (IFEM) estimated on Panel Study of Income Dynamics (PSID) data for 1980–2017. The IFEM regresses log wages on occupation-year fixed effects, experience controls, and worker fixed effects (capturing unobserved skill theta_i) interacted with occupational category and year, instrumenting the fixed effects with years of schooling to correct attenuation bias. Results confirm both predictions. The return to skill in routine occupations declined sharply from the late 1980s onward: log alpha_{R,t} fell by more than 0.7, corresponding to a greater-than-50 percent reduction between its 1987 peak and 2017, while manual and abstract occupations showed no comparable decline. Average skill in routine occupations also fell steadily, dropping from near the population mean in the early 1980s to approximately -0.2 by the end of the sample, such that by 2015 routine workers had lower average skill than manual workers.&lt;/p&gt;
&lt;p&gt;To quantify SR-RBTC&amp;rsquo;s contribution to overall wage polarization, the paper introduces a skewness decomposition. Because SR-RBTC violates the ignorability assumption underlying standard decomposition methods (e.g., DiNardo et al. 1996; Firpo et al. 2009), prior approaches could not capture the within-occupation inequality changes central to the mechanism. The skewness decomposition partitions the third central moment of log wages into a within-occupation component, a between-occupation component, and a covariance component (correlation between occupation mean wages and occupation wage inequality). Using Current Population Survey Outgoing Rotation Group (CPS-ORG) data focused on 1992–2002, the paper finds that 93 percent of the rise in skewness is related to occupational trends (the within component explains only 7 percent). Of that, 78 percent of the total increase in skewness is driven by the covariance component — rising inequality in higher-paying abstract occupations combined with falling inequality in lower-paying routine occupations — consistent exclusively with SR-RBTC rather than skill-neutral RBTC. The paper concludes that SR-RBTC can account for the large majority of U.S. bottom-half wage polarization trends from the late 1980s through the early 2000s.&lt;/p&gt;
&lt;p&gt;Q: What is the core distinction between SR-RBTC and standard (skill-neutral) RBTC?&lt;/p&gt;
&lt;p&gt;A: In standard RBTC, technology raises productivity uniformly for all routine workers regardless of skill (epsilon = 0), so wage effects are identical across the skill distribution within routine occupations. In SR-RBTC (epsilon &amp;lt; 0), technology and skill are substitutes, so higher-skilled routine workers experience proportionally smaller productivity gains — or relative wage declines — while lower-skilled routine workers may benefit. This means SR-RBTC compresses the within-routine wage distribution rather than shifting it uniformly downward.&lt;/p&gt;
&lt;p&gt;Q: How does SR-RBTC generate wage polarization concentrated at the median rather than across the full bottom half?&lt;/p&gt;
&lt;p&gt;A: Because the largest wage drops fall on the highest-skilled workers within the routine occupation, and those workers were empirically concentrated near the middle of the overall wage distribution, SR-RBTC disproportionately reduces wages around the median. Skill-neutral RBTC, by contrast, would reduce wages equally for all routine workers who are spread across the full bottom half, predicting wage declines throughout the bottom 50 percent rather than just near the 50th percentile.&lt;/p&gt;
&lt;p&gt;Q: Why does the model predict a non-monotonic relationship between technological progress and bottom-half inequality?&lt;/p&gt;
&lt;p&gt;A: In Phase 1, routine occupations employ middle-skilled workers; SR-RBTC reduces wages most for the highest-earning (highest-skilled) routine workers, compressing the bottom half of the distribution. In Phase 2, once the return to skill in routine occupations falls below that in manual occupations, the comparative advantage of middle-skilled workers shifts away from routine jobs, and routine occupations come to employ the lowest-skilled workers. Further SR-RBTC then concentrates negative wage pressure at the bottom of the distribution, potentially increasing bottom-half inequality. The transition between these phases corresponds empirically to the reversal around 2000.&lt;/p&gt;
&lt;p&gt;Q: What does the IFEM find about the return to skill in routine versus other occupations?&lt;/p&gt;
&lt;p&gt;A: Log alpha_{R,t} (the return to unobserved skill in routine occupations) fell by more than 0.7 log points between its 1987 peak and 2017, representing a greater-than-50 percent reduction. Manual occupations remained stable at approximately log alpha_{M,t} = -0.3. Abstract occupations saw a smaller and later decline, largely after 1994, consistent with evidence on a reversal in demand for cognitive skills (Beaudry et al. 2016) but far less pronounced than the routine occupation decline. The ranking of return to skill between routine and manual occupations reversed during the 1990s, matching the model&amp;rsquo;s Phase 2 threshold condition (Theorem 5).&lt;/p&gt;
&lt;p&gt;Q: What does the IFEM find about the skill composition of routine workers over time?&lt;/p&gt;
&lt;p&gt;A: Average estimated skill (theta_hat_i) in routine occupations declined from near zero (the population average) in the early 1980s to approximately -0.2 by the end of the sample. By 2015, average skill in routine occupations fell below that of manual workers, a reversal not seen for abstract or manual occupations over the same period. The decline in routine skill composition was primarily driven by fewer middle-skilled workers entering the labor force into routine jobs: the share of middle-skilled new entrants going into routine occupations fell from nearly 50 percent in the early 1980s to around 33 percent after 2010, at a rate of 0.53 percentage points per year.&lt;/p&gt;
&lt;p&gt;Q: What is the skewness decomposition and why is it needed?&lt;/p&gt;
&lt;p&gt;A: Skewness — the third standardized moment of the log wage distribution — measures asymmetry and captures wage polarization (rising top-half inequality alongside falling bottom-half inequality). It decomposes into three components: within-occupation (residual skewness not explained by occupational structure), between-occupation (skewness from differences in group means), and a covariance component (correlation between occupation-level mean wages and occupation-level wage inequality). Standard decomposition methods (Juhn et al. 1993; DiNardo et al. 1996; Firpo et al. 2009) rely on ignorability, which fails when the within-occupation wage distribution itself changes — as SR-RBTC predicts. The covariance component of skewness captures exactly these within-occupation structural changes without requiring ignorability.&lt;/p&gt;
&lt;p&gt;Q: What do the skewness decomposition results show about the driver of wage polarization?&lt;/p&gt;
&lt;p&gt;A: Decomposing the rise in skewness between 1992 and 2002 using 3-digit occupational coding, 93 percent of the total increase is attributable to occupational trends (only 7 percent is explained by the within-occupation component unrelated to occupational structure). Of the total skewness increase, 78 percent is accounted for by the covariance component — rising inequality in high-paying abstract occupations combined with declining inequality in low-paying routine occupations. This pattern is precisely what SR-RBTC predicts and cannot be generated by skill-neutral RBTC, which would predict the rise to come primarily from the between-occupation component (declining average routine wages).&lt;/p&gt;
&lt;p&gt;Q: Why did prior decomposition methods fail to detect the SR-RBTC mechanism?&lt;/p&gt;
&lt;p&gt;A: Prior methods (e.g., Autor et al. 2005; Firpo et al. 2013) operated under the ignorability assumption: the conditional distribution of wages given observables (e.g., occupation) is unchanged when the distribution of observables changes. This holds under skill-neutral RBTC (uniform wage effects within routine occupations) but fails under SR-RBTC, where the within-occupation wage structure itself changes. Consequently, prior methods only captured the (modest) decline in average routine wages — too small to explain observed polarization — and missed the inequality compression within routine occupations, which is the primary driver.&lt;/p&gt;
&lt;p&gt;Q: What are the two micro-foundations offered for SR-RBTC?&lt;/p&gt;
&lt;p&gt;A: The first (Appendix B.1) models technology as automating a subset of tasks within routine occupations, freeing workers to spend more time on remaining tasks. SR-RBTC arises when the automated task is more skill-intensive than the average task (e.g., arithmetic calculations for cashiers); automating a relatively skill-intensive task disproportionately helps lower-skill workers. The second (Appendix B.2) models technology as improving the quality or quantity of capital (computers, robots) that substitutes for skill; SR-RBTC arises when the elasticity of substitution between skill and technology exceeds a threshold, making skill and technology gross substitutes.&lt;/p&gt;
&lt;p&gt;Q: How does SR-RBTC explain the absence of large average wage declines in routine occupations despite large employment declines?&lt;/p&gt;
&lt;p&gt;A: Under SR-RBTC, wages fall for the highest-skilled workers in the routine occupation but may rise (or fall less) for lower-skilled routine workers, since the technology reduces the skill premium rather than depressing all wages uniformly. The compositional shift — higher-skilled workers exiting routine occupations — further mitigates measured average wage declines by replacing the departing high earners with lower-skilled entrants who earn closer to the (now-compressed) routine wage floor. As a result, quantity (employment) adjusts more than price (average wage), consistent with the observed data.&lt;/p&gt;
&lt;p&gt;Q: What is the quantitative magnitude of the skill-level change in routine occupations?&lt;/p&gt;
&lt;p&gt;A: Given that the return to skill in routine occupations in 2017 (alpha_{R,2017}) was approximately 0.3 (corresponding to -1.2 in log units), and average skill in routine occupations fell by approximately 0.2 units, the paper calculates that if routine workers in 2017 had maintained the same average skill level as in 1980, their wages would have been approximately 6 percent higher.&lt;/p&gt;
&lt;p&gt;Q: What alternative explanations does the paper evaluate, and how does it rule them out?&lt;/p&gt;
&lt;p&gt;A: The paper considers minimum wage increases (Piketty 2014) and declining unionization (Firpo et al. 2013) as potential contributors. The skewness decomposition implies these explanations are limited: since 93 percent of the skewness increase is driven by occupational trends and 78 percent by the covariance component (within-occupation inequality changes), mechanisms that operate through uniform group-level wage shifts — as minimum wage or union explanations would — can account for only a small fraction of the overall trend. The IFEM further rules out that the decline in within-routine inequality reflects worker composition becoming more homogeneous rather than a genuine decline in return to skill, as the sensitivity analysis shows alpha_jt changes are driven almost entirely by workers staying within each occupational category.&lt;/p&gt;
&lt;p&gt;Skill-Replacing RBTC (SR-RBTC): A variant of routine-biased technological change in which technology substitutes the usage of skill within routine occupations (epsilon &amp;lt; 0), reducing the return to skill and compressing within-occupation wage inequality, as distinct from skill-neutral RBTC (epsilon = 0) which shifts wages uniformly and skill-enhancing RBTC (epsilon &amp;gt; 0) which widens skill gaps.&lt;/p&gt;
&lt;p&gt;Interactive Fixed-Effects Model (IFEM): An extension of the standard fixed-effects panel wage regression in which worker fixed effects (capturing unobserved permanent skill theta_i) are interacted with both occupational category and year, allowing the estimated return to skill alpha_jt to vary across occupations and over time; worker fixed effects are instrumented with years of schooling to correct attenuation bias.&lt;/p&gt;
&lt;p&gt;Skewness Decomposition: A decomposition of the third central moment of the log wage distribution (skewness) into three components — within-occupation, between-occupation, and a covariance term (the covariance between occupation-level mean wages and occupation-level wage inequality) — that, unlike standard decomposition methods, does not require the ignorability assumption and can therefore capture changes in the within-occupation wage structure.&lt;/p&gt;
&lt;p&gt;Ignorability Assumption: The assumption, required by standard decomposition methods (e.g., DiNardo et al. 1996; Firpo et al. 2009), that the conditional distribution of wages given observables (here, occupations) does not change when the distribution of observables changes; violated under SR-RBTC because the within-occupation wage structure itself shifts as skill-replacing technology advances.&lt;/p&gt;
&lt;p&gt;Comparative Advantage (Occupational Sorting): The mechanism by which workers sort into occupations based on their skill level theta_i relative to occupation-specific return-to-skill schedules; SR-RBTC shifts occupational thresholds by compressing the routine occupation&amp;rsquo;s skill premium, causing higher-skilled workers to exit routine jobs and lower-skilled workers to enter.&lt;/p&gt;
&lt;p&gt;Two-Phase Dynamics: The non-monotonic relationship between technological progress and bottom-half inequality in the SR-RBTC model; Phase 1 (late 1980s–2000) sees middle wages decline as the highest-skilled (middle-of-distribution) routine workers experience the largest wage drops; Phase 2 (2000 onward) sees bottom wages fall as the routine occupation shifts to employing the lowest-skilled workers once the routine skill premium falls below the manual skill premium.&lt;/p&gt;</description></item><item><title>Spatial Implications of Telecommuting</title><link>https://macropaperwarehouse.com/papers/spatial-implications-of-telecommuting/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/spatial-implications-of-telecommuting/</guid><description>&lt;p&gt;Delventhal and Parkhomenko build a quantitative spatial model of the United States to study how the rise of telecommuting reshapes the distribution of residents, jobs, and housing costs across and within cities. The model divides the continental U.S. into 4,502 locations (defined as intersections of Census PUMAs and counties) and allows each worker to choose any residence-job pair. Workers differ by education (college vs. non-college) and occupation type (telecommutable vs. non-telecommutable). Telecommutable workers can split labor time between on-site and remote work; their remote-work intensity responds endogenously to relative remote productivity, a work-from-home aversion parameter, home floorspace costs, and commute time.&lt;/p&gt;
&lt;p&gt;The model is calibrated to pre-2020 U.S. data (2012–2016 ACS, 2018 SIPP, 2017 NHTS). Key calibrated facts include: 33.6% of workers have telecommutable jobs (40.6% of non-college, 72.7% of college workers); remote work is nearly as productive as on-site work (relative productivity 0.99–1.00); elasticities of substitution between work modes range from 3.48 to 5.05; and work-from-home aversion parameters range from 2.48 to 3.35, indicating large non-pecuniary barriers especially for non-college workers in non-tradable sectors.&lt;/p&gt;
&lt;p&gt;The counterfactual simulates a permanent increase in remote work driven by an 8–10% rise in remote productivity and a fall in work-from-home aversion, guided by Barrero, Bloom, and Davis (2021) survey evidence. Results show net reallocation of jobs and residences equivalent to nearly 5% of the population.&lt;/p&gt;
&lt;p&gt;Main spatial findings exhibit a non-monotonic pattern. Telecommutable residents move away from dense, high-cost locations toward sparser areas with lower housing costs and better amenities. Non-telecommutable residents partially counteract this by centralizing — moving toward denser areas as housing costs fall near job centers. Non-tradable jobs follow telecommuters outward. Tradable jobs move in both directions: some firms relocate to low-density areas with newly accessible remote worker pools; others expand in the largest, most productive city centers as office space costs fall and the catchment area of workers widens.&lt;/p&gt;
&lt;p&gt;In aggregate: the average worker lives 47% farther (in commuting time) from their workplace but spends 25% less time commuting, because average remote-work frequency rises by 1.1 days per week. The share of workers living in one commuting zone and working in another increases from 24.6% to 34%. Average income falls marginally by 1%, masking large gains for telecommutable workers and losses for non-telecommutable workers. Average floorspace prices fall by 2%; non-tradable prices rise by 2.6%. Overall welfare increases by an average of 12.7%, driven by gains for telecommutable workers, while non-telecommutable workers experience net losses.&lt;/p&gt;
&lt;p&gt;The model predicts a partial reversal of the &amp;ldquo;Great Divergence&amp;rdquo;: skill sorting falls both within and across commuting zones, residential income inequality across CZs falls, and house price dispersion falls both within and across cities. These predictions are directionally consistent with 2019–2023 data.&lt;/p&gt;
&lt;p&gt;Scope conditions: results are for a permanent shock to the full-time U.S. workforce as modeled in 2012–2016; the model does not predict the end of big cities but rather a reallocation at the margin. The model shows that the introduction of telecommuting narrows the parameter range guaranteeing a unique spatial equilibrium, because remote-capable firms can draw from a broader worker catchment area, amplifying agglomeration forces.&lt;/p&gt;
&lt;p&gt;Q: What are the four stylized facts about pre-2020 telecommuting that discipline the model?
A: Fact 1: telecommutability is higher for college workers and those in tradable industries — 68.8% of college-tradable workers can work from home versus 18.9% of non-college non-tradable workers. Fact 2: among telecommutable workers, uptake is also higher for college-tradable workers (38% actually work from home at least one day per week) than for non-college non-tradable workers (21%). Fact 3: the distribution of remote-work frequency is bimodal — most workers are either fully on-site or fully remote, with the bimodality less pronounced for college-tradable workers where hybrid (1–4 days/week) accounts for over 11% of paid workdays. Fact 4: there is a positive relationship between work-from-home frequency and distance from the job site, consistent with telework reducing effective commuting costs.&lt;/p&gt;
&lt;p&gt;Q: How is the counterfactual shock calibrated and what drives it?
A: The counterfactual raises remote-work productivity by 8–10% across all worker types and simultaneously reduces work-from-home aversion, guided by Barrero, Bloom, and Davis (2021) survey evidence that 25–30% of paid workdays will be remote post-pandemic, compared to about 8% in 2018. The authors consider both a technology shock (productivity increase) and a preference shock (aversion decrease) as mechanisms, consistent with their view that multiple hypotheses about the COVID-19 telework shock are plausible and non-exclusive.&lt;/p&gt;
&lt;p&gt;Q: How do residents reallocate in response to the rise in telecommuting?
A: Net reallocation of residents equivalent to nearly 5% of the population occurs. Telecommutable residents decentralize — moving to less dense areas with lower housing costs and better amenities — because the cost of choosing a residence far from work falls. Non-telecommutable residents partially centralize, moving toward denser locations in larger metro areas, because housing costs fall in locations with short commutes, making them more affordable.&lt;/p&gt;
&lt;p&gt;Q: How do jobs reallocate?
A: Non-tradable jobs follow the decentralization of residents (their source of demand) monotonically to less dense locations. Tradable jobs move in both directions: some firms relocate to low-density areas that can now access a larger pool of remote workers at lower real estate costs; others expand operations in the highest-productivity city centers, benefiting from both an expanded catchment of remote workers and a decline in the high cost of office space.&lt;/p&gt;
&lt;p&gt;Q: What are the aggregate commuting implications?
A: The average worker lives 47% farther in commuting time from their workplace in the counterfactual, yet spends 25% less time commuting, because average remote-work frequency increases by 1.1 days per week. The share of workers living in one commuting zone and working in another rises from 24.6% to 34%, which the authors note may call into question current administrative definitions of commuting zones and have major impacts on travel patterns.&lt;/p&gt;
&lt;p&gt;Q: What are the welfare and income effects?
A: Overall welfare increases by an average of 12.7%, but this masks very unequal distribution: telecommutable workers experience large gains while non-telecommutable workers suffer losses. Average worker income falls marginally by 1%, reflecting sizable gains for remote-capable workers offset by losses for those who cannot telecommute. Average floorspace prices fall by 2%, while non-tradable goods prices rise by 2.6%.&lt;/p&gt;
&lt;p&gt;Q: What does the model predict for the &amp;ldquo;Great Divergence&amp;rdquo;?
A: The model predicts a significant re-convergence across multiple dimensions: skill sorting falls both within and across commuting zones, residential wage inequality across CZs falls, and house price dispersion falls both within and across cities. The authors find that commuting zones with higher college shares in 2019 experienced slower growth in college shares 2019–2023, and that there is a negative correlation between average wages by CZ in 2019 and wage growth 2019–2023 — both consistent with model predictions.&lt;/p&gt;
&lt;p&gt;Q: How does the model validate against post-2019 data?
A: The authors show that their counterfactual results are positively correlated with observed changes in population, jobs, and housing rents since 2019. Within-city price variance has already converged in 2019–2023 data, consistent with model predictions. CZ-level patterns of skill concentration and wage growth also move in the direction the model predicts.&lt;/p&gt;
&lt;p&gt;Q: Is the COVID-19 shock better described as a technology shock or a preference shock?
A: The authors test both. To replicate observed changes in remote-work frequency using only a productivity shock requires a 55–99% jump in remote productivity, which yields implausibly large wage gains for remote-capable workers of 47–82%. The preference-based scenario yields results more consistent with observed data, supporting the view that a preference shock — changes in norms, attitudes, and institutional policies — is the primary driver.&lt;/p&gt;
&lt;p&gt;Q: What happens to real estate prices when supply and amenities are held fixed?
A: When real estate supply, productivity, and amenities are all held fixed, residential prices jump by 16% and commercial prices fall by 16%. The authors note this mimics the bifurcated shift in real estate values observed during the pandemic years, suggesting that supply responses and amenity adjustments are important for dampening the price effects in the full model.&lt;/p&gt;
&lt;p&gt;Q: How does the model handle the uniqueness of spatial equilibrium, and how does telecommuting affect it?
A: In a standard quantitative spatial model, agglomeration forces are dampened by the finite pool of workers willing to commute daily to a productive location. When telecommuting is introduced, productive locations can draw workers from a much broader catchment area, amplifying agglomeration forces and narrowing the range of parameter values for which a unique equilibrium is guaranteed. The authors establish conditions under which uniqueness is preserved.&lt;/p&gt;
&lt;p&gt;Q: What are the model&amp;rsquo;s three main advantages over more stylized spatial models of remote work?
A: First, by including 4,502 locations, the model can predict how far telecommuters will move from their jobs — a key variable for real estate markets and commuting patterns. Second, it can represent changes in the distribution of workers across different work-from-home frequencies, which is crucial as hybrid work has emerged as the dominant post-pandemic arrangement. Third, it predicts how the location of jobs (not just residents) changes, which has important implications for city centers.&lt;/p&gt;
&lt;p&gt;Q: What is the overall welfare conclusion regarding non-telecommutable workers and income inequality?
A: Non-telecommutable workers suffer welfare losses from the rise of remote work, even as overall average welfare rises by 12.7%. The overall income inequality — as opposed to spatial wage dispersion — does not fall. The authors note this means the spatial re-convergence does not translate into a broader reduction in income inequality, which they flag as an important limitation for policy.&lt;/p&gt;
&lt;p&gt;Telecommutability: the ability of a worker&amp;rsquo;s occupation to be performed from home, measured using Dingel and Neiman (2020) occupational classifications; varies by education and industry, with 68.8% of college-tradable workers telecommutable versus 18.9% of non-college non-tradable workers.&lt;/p&gt;
&lt;p&gt;Work-from-home aversion (ς): a preference parameter representing tastes, norms, and institutional policies that create non-pecuniary barriers to remote work; calibrated to range from 2.48 to 3.35 across worker types, higher for non-college workers in non-tradable sectors.&lt;/p&gt;
&lt;p&gt;Hybrid work: an arrangement in which a telecommutable worker splits paid workdays between on-site and remote work (1–4 days per week from home); the model&amp;rsquo;s bimodal distribution of work-from-home frequency replicates the empirical observation that most workers are either fully on-site or fully remote, with hybrid most prevalent among college-tradable workers.&lt;/p&gt;
&lt;p&gt;Catchment area: the pool of workers from which a firm can practically hire, which widens under telecommuting because workers no longer need to commute daily; this widening amplifies agglomeration forces and narrows the parameter range guaranteeing a unique spatial equilibrium.&lt;/p&gt;
&lt;p&gt;Great Divergence: the multi-decade trend (documented in Moretti 2012 and related work) of spatially concentrating talent, income, and housing costs in a small number of large, high-skill cities; the paper predicts a partial reversal — &amp;ldquo;Great Re-Convergence&amp;rdquo; — driven by the rise of telecommuting.&lt;/p&gt;
&lt;p&gt;Productive externalities (agglomeration): local productivity in the model depends on employment density; remote workers participate in these externalities only partially (parameter ψ ∈ [0,1]), so the shift to remote work can reduce agglomeration benefits in city centers.&lt;/p&gt;
&lt;p&gt;Source text origin: the paper&amp;rsquo;s own classification of the text on which a summary is based (full PDF, open-access HTML, or abstract-only); the paper&amp;rsquo;s CLAUDE.md rules mandate that abstract-only summaries are blocked.&lt;/p&gt;</description></item><item><title>Supply, Demand, Institutions, and Firms: A Theory of Labor Market Sorting and the Wage Distribution</title><link>https://macropaperwarehouse.com/papers/supply-demand-institutions-and-firms-a-theory-of-labor-market-sorting-and-the-wage-distribution/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/supply-demand-institutions-and-firms-a-theory-of-labor-market-sorting-and-the-wage-distribution/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; How do workforce composition (labor supply), labor demand, and minimum wage policy jointly determine the wage distribution in imperfectly competitive labor markets, and what were the quantitative contributions of each force to the dramatic decline in Brazilian wage inequality between 1998 and 2012?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Motivation.&lt;/strong&gt; Brazil&amp;rsquo;s formal-sector wage inequality fell sharply over this period. Three candidate shocks are well-documented: (1) a large increase in educational attainment — the share of adults completing at least secondary school rose by 20 percentage points (a 68 percent increase) between 1998 and 2012; (2) labor demand shocks, primarily the commodities boom of the 2000s; and (3) a 93.7 percent (66.1 log point) real increase in the federal minimum wage. Existing frameworks analyze these shocks separately — competitive supply/demand models on one side and imperfectly competitive minimum wage models on the other — and therefore cannot detect interactions or jointly explain all observed patterns, including the novel finding that assortative matching between high-wage workers and high-wage establishments rose in 104 out of 151 microregions, a fact inconsistent with the predictions of leading minimum wage models.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The paper uses the RAIS (Relação Anual de Informações Sociais), a confidential linked employer-employee dataset covering the Brazilian formal sector, together with Brazilian Census data for 1991, 2000, and 2010. Statistics are computed for 151 microregions (analogous to US commuting zones) with at least 15,000 workers in RAIS in both base years and at least 1,000 formal workers per educational group. The final sample covers 73 percent of the adult population. Firm wage premiums and assortative matching are measured via AKM two-way fixed effects regressions using the bias-corrected KSS (Kline, Saggio, Sølvsten 2018) estimator, run separately for each microregion and period on three-year panels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Theoretical framework.&lt;/strong&gt; The paper develops a unified general-equilibrium model featuring: (i) a task-based production function with distance-dependent complementarity between worker types; (ii) monopsony power arising from idiosyncratic worker preferences for firms, generating constant firm-level labor supply elasticity β (calibrated at 4, implying markdowns of 20 percent); (iii) heterogeneous firms differentiated by their production &amp;ldquo;blueprints&amp;rdquo; (the complexity of tasks they require), with blueprint shape parameterized as a Gamma distribution; and (iv) free firm entry, endogenous participation, and goods market general equilibrium with CES consumer preferences (elasticity σ). A key result is that firms with different blueprints exhibit different within-firm substitution patterns: worker types that are substitutes at low-skill, low-wage firms may be complements at high-skill, high-wage firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Estimation.&lt;/strong&gt; A parsimonious parameterization is estimated by simultaneous-equation nonlinear least squares, targeting 26 endogenous outcomes per region (13 per period) including between- and within-group wage inequality, variance of establishment effects, covariance of worker and establishment effects, formal employment rates by education, and minimum wage bindingness. The model requires solving for equilibrium more than 15,000 times per optimization step (151 regions × 2 periods × 53 Jacobian columns). The elasticity of substitution between goods is estimated at σ = 8.36 (significantly above 1), and the aggregate labor supply parameter λ implies formal-sector elasticities of approximately 0.6–0.7 for college workers and around 1.1 for less-than-secondary workers. The model fits the data well, with R² above 0.5 for most targeted moments and perfect fit for the six moments used in the inversion procedure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings.&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Demand shocks and the minimum wage are the primary drivers of falling inequality.&lt;/strong&gt; In counterfactual simulations, the minimum wage alone (a 66.1 log point increase) reduces the variance of log wages by 0.13. Demand shocks reduce it by a further 0.18. Supply shocks (rising education) increase the variance by 0.04, leaving their net inequality-reducing contribution negligible.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Supply shocks increase assortative matching despite compressing within-firm skill premiums.&lt;/strong&gt; Within-firm task reassignment would reduce the variance of log wages by 0.221 and the correlation between worker and establishment effects by 0.165, holding production levels and firm entry fixed. However, scale, entry, and price adjustments — driven by the large estimated σ = 8.36 &amp;gt; β + 1 = 5 — reallocate skilled labor toward high-wage, skill-intensive firms, counteracting within-firm compression and raising assortative matching by 0.189. These two channels largely offset each other.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Concurrent supply and demand changes attenuate minimum wage impacts by roughly half.&lt;/strong&gt; When the minimum wage is the only shock, it would have reduced the variance of log wages by 0.13; in the presence of supply and demand changes, its incremental contribution is approximately 0.07. Minimum wage effects on sorting (which would reduce assortative matching when acting alone) disappear when accompanied by supply and demand transformations.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Minimum wage effects are concentrated in the bottom two productivity deciles.&lt;/strong&gt; Wage effects for workers in productivity deciles three through ten from the minimum wage are approximately 1 percent or less once all channels are considered. Strong wage gains are concentrated at the bottom, primarily through the monopsony channel. The wage-posting channel (within-firm returns to skill) reduces wages for low- and middle-skill workers and raises them at the top two deciles due to the reallocation of low-skilled workers toward high-wage firms, which reduces those workers&amp;rsquo; marginal products there.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cross-firm differences in substitution patterns generate non-standard minimum wage spillovers.&lt;/strong&gt; Conditional on the task demands of the firm employing them, a pair of worker types may be substitutes in low-skill firms and complements in high-skill firms. This firm-heterogeneity channel causes minimum wage impacts to be non-monotone across the productivity distribution, contrasting with the smooth inequality-reducing effects predicted by both competitive task-based models and frictional minimum wage models.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-novel-empirical-fact-that-motivates-the-unified-framework"&gt;Q1. What is the novel empirical fact that motivates the unified framework?&lt;/h3&gt;
&lt;p&gt;A: Using KSS bias-corrected AKM decompositions performed separately for each of 151 microregions, the paper documents that assortative matching — measured as the correlation between worker and establishment fixed effects — rises in 104 out of 151 regions between 1998 and 2012. The covariance term accounts for less than 7 percent of the average decline in the variance of log wages. This finding is inconsistent with the leading imperfectly competitive minimum wage model (Engbom and Moser 2022), in which minimum wages reduce assortative matching. It is also inconsistent with purely competitive supply/demand models, which have no role for firm wage premiums or sorting. The divergence from prior national-level studies (which do not find rising sorting) is explained by the fact that national-level sorting conflates geographical sorting with supply-demand dynamics.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-key-mechanism-through-which-the-task-based-production-function-generates-cross-firm-differences-in-substitution-patterns"&gt;Q2. What is the key mechanism through which the task-based production function generates cross-firm differences in substitution patterns?&lt;/h3&gt;
&lt;p&gt;A: In the task-based production function, each firm assigns workers to tasks assortatively — lower types handle lower-complexity tasks, higher types handle higher-complexity tasks, with cutoff thresholds determined by the firm&amp;rsquo;s blueprint. When a firm has a blueprint concentrated in complex tasks (a high-skill, high-wage firm), adjacent worker types are more differentiated in the tasks they perform, making them complements. When a firm has a blueprint concentrated in simple tasks (a low-skill, low-wage firm), adjacent worker types are assigned to a narrow, similar range of tasks and are therefore closer substitutes. The elasticity of complementarity between any pair of worker types is thus endogenous, depending on which tasks the firm uses and, in the monopsony case, on the firm&amp;rsquo;s skill intensity — a prediction validated empirically using nonroutine cognitive task content data for Brazilian occupations.&lt;/p&gt;
&lt;h3 id="q3-under-what-conditions-can-a-positive-supply-shock-rising-educational-attainment-widen-the-aggregate-skill-wage-premium-rather-than-compress-it"&gt;Q3. Under what conditions can a positive supply shock (rising educational attainment) widen the aggregate skill wage premium rather than compress it?&lt;/h3&gt;
&lt;p&gt;A: The paper&amp;rsquo;s Proposition 4 and Corollary 2 show that a supply shock that increases the relative supply of skilled workers can widen the aggregate skill wage premium when the elasticity of substitution between goods (σ) exceeds the firm-level elasticity of labor supply plus one (β + 1). Intuitively, when σ is large, the reduction in prices for skill-intensive goods generated by the supply shock shifts consumption toward those goods, causing net entry of skill-intensive firms. If the gains in firm wage premiums earned by skilled workers reallocated to those firms outweigh the compression in within-firm productivity differentials, the aggregate skill premium can rise. This mechanism does not require non-convexities from endogenous innovation; it operates through imperfect competition and firm entry alone. In the estimated Brazilian model, σ = 8.36 substantially exceeds β + 1 = 5, so this condition holds, explaining why rising education increases rather than compresses assortative matching in the data.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-model-generate-positive-employment-effects-from-minimum-wages-and-how-do-these-interact-with-reallocation"&gt;Q4. How does the model generate positive employment effects from minimum wages, and how do these interact with reallocation?&lt;/h3&gt;
&lt;p&gt;A: In the monopsonistic baseline without a minimum wage, firms post wages below workers&amp;rsquo; marginal revenue products, causing some workers to choose non-employment. A minimum wage increase raises posted wages at constrained firms, shifting some workers from non-employment (or home production) to formal employment, generating positive employment effects at the margin where the minimum wage binds. Simultaneously, minimum wages price out the least productive workers at low-wage firms (disemployment), while workers in the intermediate productivity range reallocate from low- to high-wage firms, because high-wage firms have higher revenue productivity and can profitably hire workers that low-wage firms can no longer afford. The net employment elasticity for the lowest productivity decile with respect to the log minimum wage is −0.61 (Table 7), while the mean wage for that decile rises substantially through the monopsony channel.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-three-channels-through-which-the-minimum-wage-affects-wages-and-employment-in-the-model-and-what-does-each-channel-contribute"&gt;Q5. What are the three channels through which the minimum wage affects wages and employment in the model, and what does each channel contribute?&lt;/h3&gt;
&lt;p&gt;A: The paper decomposes minimum wage effects into three channels. Channel 1 (monopsony): mechanical wage increases, positive employment effects at firms where the minimum wage binds, disemployment of very low-productivity workers, and reallocation from low- to high-wage firms, holding posted wage schedules, prices, and entry fixed. This channel accounts for nearly all of the strong wage effects at the bottom two productivity deciles. Channel 2 (wage posting): firms reoptimize earnings schedules following changes in worker composition and marginal products induced by Channel 1, holding prices and entry fixed. This channel reduces wages for low- and middle-skill workers (productivity deciles 1–7) by approximately 0.01–0.02 log points and increases wages for top deciles (decile 9: +0.04, decile 10: +0.11), because reallocation of low-skill labor to high-wage firms lowers those workers&amp;rsquo; marginal products there. Channel 3 (general equilibrium): firm entry and price responses. The fall in low-wage-firm profits causes entry of high-wage, skill-intensive firms, while the price of low-skill goods falls. General equilibrium effects generate modest positive wage effects for most workers but negative effects for very low-productivity workers due to reduced aggregate demand for low-skill labor.&lt;/p&gt;
&lt;h3 id="q6-why-do-the-minimum-wages-inequality-reducing-effects-diminish-when-accompanied-by-concurrent-supply-and-demand-changes"&gt;Q6. Why do the minimum wage&amp;rsquo;s inequality-reducing effects diminish when accompanied by concurrent supply and demand changes?&lt;/h3&gt;
&lt;p&gt;A: The paper documents that, under concurrent supply and demand transformations, the minimum wage&amp;rsquo;s reduction of the variance of log wages is approximately 0.07, roughly half the 0.13 reduction it would achieve acting alone. The attenuation occurs through interactions: supply and demand shocks raise the average productivity level of the labor market and shift workers toward high-wage, skill-intensive firms. In this altered equilibrium, the minimum wage binds less tightly (or hits a different part of the distribution), and the reallocation effects of the minimum wage that would normally reduce assortative matching are offset by the sorting-increasing effects of supply and demand changes. The estimated model shows that interactions between the minimum wage and supply/demand changes (columns 6, 7, 8 of Table 5) are economically meaningful, something undetectable without a unified framework.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-models-prediction-regarding-minimum-wage-spillovers-differ-from-engbom-and-moser-2022-and-what-explains-the-difference"&gt;Q7. How does the model&amp;rsquo;s prediction regarding minimum wage spillovers differ from Engbom and Moser (2022), and what explains the difference?&lt;/h3&gt;
&lt;p&gt;A: Engbom and Moser (2022) find that the Brazilian minimum wage hike had significant wage effects extending far up the worker productivity distribution, while this paper&amp;rsquo;s model finds negligible effects (approximately 1 percent) beyond the bottom two productivity deciles. Two structural differences explain this divergence. First, Engbom and Moser (2022) assume perfect substitutability between worker types within firms, so a minimum wage increase at low-wage firms mechanically raises posted wages at all other firms to maintain relative competitiveness. In this paper&amp;rsquo;s framework, wage-posting responses at high-wage firms can be negative for low-skill workers because the inflow of reallocated low-skill workers reduces their marginal products — a channel absent under perfect substitution. Second, Engbom and Moser (2022) use a national model, allowing displaced low-skill workers to reallocate to top-productivity firms anywhere in the country, dampening disemployment; this paper&amp;rsquo;s local labor markets approach restricts reallocation to within-region boundaries, consistent with low rates of interregional migration documented for Brazil by Dix-Carneiro and Kovak (2017).&lt;/p&gt;
&lt;h3 id="q8-how-are-firm-wage-premiums-generated-in-the-model-and-why-do-differences-in-physical-productivity-between-firms-not-generate-wage-differentials"&gt;Q8. How are firm wage premiums generated in the model, and why do differences in physical productivity between firms not generate wage differentials?&lt;/h3&gt;
&lt;p&gt;A: Proposition 3 establishes that wage dispersion for similar workers across firms requires either (i) differences in blueprint shapes (firm heterogeneity in skill intensity) or (ii) differences in entry costs. Differences in physical productivity (z_g) or consumer taste parameters alone are insufficient, because with equal entry costs, differences in productivity lead to additional firm entry until the marginal revenue product of labor is equalized across firm types. Wage premiums proportional to entry costs arise because optimal firm creation requires larger-scale operation for higher-entry-cost firms, and hiring more workers forces those firms to post higher wages. Additionally, skill-intensive firms (firms with blueprints tilted toward complex tasks) pay relative wage premiums for the worker types they use most intensively, and if skill intensity and entry costs co-vary, all workers at high-skill firms may receive a wage premium.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-estimation-procedure-handle-unobserved-regional-heterogeneity-in-labor-demand"&gt;Q9. How does the estimation procedure handle unobserved regional heterogeneity in labor demand?&lt;/h3&gt;
&lt;p&gt;A: Demand shocks are not directly observed; they are inferred as a residual from changes in targeted outcomes after accounting for observed supply (education shares from Census) and minimum wage changes. Five region-time-specific demand parameters — TFP (z), blueprint complexities (θ₁, θ₂), relative entry costs (F₂/F₁), and relative consumer preferences (γ₂/γ₁) — are modeled as linear functions of 1998 regional covariates (educational shares, agricultural share, manufacturing share, and initial minimum wage bindingness) with time-specific coefficients. This formulation allows unobserved demand shifters to correlate with initial educational levels, preventing incorrect attribution of demand-supply correlations to causal supply effects. Region-specific parameters (TFP in each period, education-group-specific formal employment shifters) are inverted exactly from six targeted moments within each region, eliminating incidental parameter bias.&lt;/p&gt;
&lt;h3 id="q10-what-micro-level-empirical-validations-does-the-paper-conduct-for-the-task-based-models-mechanisms"&gt;Q10. What micro-level empirical validations does the paper conduct for the task-based model&amp;rsquo;s mechanisms?&lt;/h3&gt;
&lt;p&gt;A: The paper tests four micro-level predictions using nonroutine cognitive task content data for Brazilian occupations. First, skill-intensive firms have greater demand for complex tasks (consistent with Figure 1 of the model). Second, within firms, more skilled workers are assigned to more complex tasks (Lemma 1). Third, workers who move to more skill-intensive firms are assigned more complex tasks (Lemma 2, consistent with the monopsony model&amp;rsquo;s mismatch prediction). Fourth, wage gaps between high- and low-skill firms are larger for skilled workers (Proposition 3). The paper reports finding strong support for all four predictions in the data, lending credibility to the theoretical structure and quantitative results.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Task-based production function (paper&amp;rsquo;s definition):&lt;/strong&gt; A production function in which a firm produces output by assigning workers of different types to tasks indexed by complexity. The assignment is assortatively optimal: lower-type workers handle lower-complexity tasks, with unique threshold complexities separating adjacent worker types. The critical property is distance-dependent complementarity — any pair of worker types that are &amp;ldquo;close&amp;rdquo; in skill rank are substitutes, while pairs distant in skill rank are complements. This differs from CES production functions where the elasticity of complementarity is the same for all pairs; in the task-based version, substitutability depends on endogenous assignment and thus on the firm&amp;rsquo;s blueprint.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Blueprint (paper&amp;rsquo;s definition):&lt;/strong&gt; A function b_g(x) that specifies the density of tasks of each complexity level x required to produce one unit of good g. It is the fundamental source of firm heterogeneity in the model: firms producing goods with blueprints tilted toward complex tasks are more skill-intensive, hire workers of higher average type, and pay higher wages. The paper parameterizes blueprints as Gamma distributions with shape parameter θ_g indexing average task complexity; firms with higher θ_g are more skill-intensive.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Firm wage premium (paper&amp;rsquo;s definition):&lt;/strong&gt; The component of wages at a given establishment that accrues equally to all workers at that firm regardless of their type, measured as the establishment fixed effect ψ_j in AKM two-way fixed effects regressions. In this model, firm wage premiums arise from heterogeneity in blueprints (skill intensity) and entry costs, not from differences in TFP or consumer tastes. Under monopsony, firms with higher entry costs must operate at larger scale and post higher wages; blueprint heterogeneity generates differential wage premiums by skill type.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sorting / assortative matching (paper&amp;rsquo;s definition):&lt;/strong&gt; The correlation between the worker fixed effect (ν_i,r capturing worker skill) and the establishment fixed effect (ψ_j capturing firm wage premium) in the AKM decomposition, measured as Cov(ν_i,r, ψ_{J(i,r,τ)} | r). In this paper&amp;rsquo;s framework, sorting arises because firms with blueprints demanding complex tasks (high-wage firms) have a comparative advantage in employing high-skill workers; labor market sorting can therefore change over time due to supply, demand, or minimum wage shocks, even without changes in search frictions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Monopsony power / markdown (paper&amp;rsquo;s definition):&lt;/strong&gt; Arising from idiosyncratic worker preferences for firms (modeled as a nested logit), firms face upward-sloping labor supply curves with constant firm-level elasticity β. Optimal posted wages equal a constant markdown β/(β+1) of the marginal revenue product of labor, set to β = 4 (implying a 20 percent markdown). The macro elasticity of formal sector labor supply is governed by a separate parameter λ, estimated from the data, yielding aggregate formal-sector supply elasticities of approximately 0.6–0.7 for college workers and around 1.1 for less-educated workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wage posting responses (paper&amp;rsquo;s definition):&lt;/strong&gt; The second channel of minimum wage effects, in which firms reoptimize their entire earnings schedule following the wage-composition changes induced by the minimum wage&amp;rsquo;s mechanical and reallocation effects (Channel 1), while keeping goods prices and firm entry fixed. Because task-based production functions are concave, changes in factor proportions (due to reallocation of low-skill workers to high-wage firms) alter marginal products of all worker types within those firms, causing firms to adjust all posted wages — not just those directly constrained by the minimum wage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distance-dependent complementarity (paper&amp;rsquo;s definition):&lt;/strong&gt; The property, proven as a Corollary to Proposition 1, that for a fixed worker type h, the partial elasticity of complementarity between h and any other type h&amp;rsquo; is strictly increasing in h&amp;rsquo; for h&amp;rsquo; ≥ h (more distant high types are stronger complements) and strictly decreasing in h&amp;rsquo; for h&amp;rsquo; ≤ h (more distant low types are weaker substitutes / stronger complements). This pattern results from the division of labor: adding a very different worker type allows specialization gains that do not arise when adding similar-type workers competing for the same tasks.&lt;/p&gt;</description></item><item><title>Talent Hoarding in Organizations</title><link>https://macropaperwarehouse.com/papers/talent-hoarding-in-organizations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/talent-hoarding-in-organizations/</guid><description>&lt;p&gt;This paper provides the first empirical evidence of talent hoarding in organizations — the practice whereby managers deliberately suppress workers&amp;rsquo; internal mobility to retain productive team members, thereby serving their own performance-based compensation interests at the expense of firm-wide talent allocation. The research question is whether managers with misaligned incentives hoard talent, how this can be measured, and what consequences it have for worker career outcomes and organizational efficiency.&lt;/p&gt;
&lt;p&gt;The study uses personnel records from a large German manufacturing firm with over 200,000 employees worldwide, focused on more than 30,000 white-collar and management employees in Germany, covering over 300,000 employee-by-quarter observations from 2015 to 2018. This is supplemented by a manager survey (62% response rate, over 3,000 responses) and an employee survey (50% response rate, over 15,000 responses), plus the universe of internal job application and hiring data covering over 16,000 job openings and over 200,000 applicants.&lt;/p&gt;
&lt;p&gt;The conceptual framework formalizes talent hoarding as a moral hazard problem: managers observe worker productivity and are compensated based on team performance, but are tasked with identifying and developing talent for promotion. When a high-productivity worker leaves, team productivity falls. The framework predicts that hoarding intensity increases with worker productivity, team vulnerability to departures (smaller teams), and manager-level hoarding incentives (performance-related pay, low talent visibility).&lt;/p&gt;
&lt;p&gt;The key administrative measure of hoarding is the systematic gap between managers&amp;rsquo; private performance ratings (not shared outside the team) and public potential ratings (widely circulated within the firm). Managers who suppress potential ratings relative to what would be predicted given worker performance are interpreted as strategically reducing worker visibility. Managers with a 1 percentage point higher share of performance-related pay are 0.19 percentage points more likely to hoard talent; a one-person increase in team size reduces hoarding probability by 1.3 percentage points; and managers in low-visibility functional areas are 4.0 percentage points more likely to hoard. Survey-based hoarding measures yield directionally identical patterns.&lt;/p&gt;
&lt;p&gt;To identify causal effects on workers, the paper exploits quasi-random manager rotations. When a manager learns they will move to a different team — typically two to three quarters before the actual transition — their hoarding incentive ceases. This creates a temporary window of reduced hoarding. During this window, worker application rates increase by 2.3 percentage points, representing a 78% increase over the baseline application rate of 2.9%. An event study confirms flat pre-trends prior to the announcement period, supporting the identifying assumption.&lt;/p&gt;
&lt;p&gt;Using manager rotations as an instrument for worker applications, marginal applicants — those induced to apply only by the manager rotation — face a 49.1% likelihood of receiving a new position, compared to an average hiring likelihood of 27.6%. This positive selection implies that many deterred applicants would have been successful and that talent hoarding meaningfully degrades the quality of the internal applicant pool. Gender analysis reveals that women are 22% more likely to rely on manager career guidance and 26% more likely to prioritize preserving a good manager relationship. Marginal female applicants are more positively selected on education, past performance, and hiring probability for higher-level positions. The counterfactual reduction in the gender pay gap from eliminating talent hoarding is estimated at 86%.&lt;/p&gt;
&lt;p&gt;Scope conditions: the firm is a large European manufacturer with long average tenures (13 years), an application-based internal labor market, and centralized online job portal. Results apply most directly to white-collar and management employees in Germany. External validity is supported by comparisons to German workforce surveys and by the fact that 83% of top publicly listed German companies and half of 665 global organizations in industry surveys report talent hoarding as a significant organizational friction.&lt;/p&gt;
&lt;p&gt;Q: How is talent hoarding formally defined in this paper?
A: Talent hoarding is defined as actions taken by managers that lower the likelihood that a worker applies for and receives a promotion or any internal transfer outside the team. In the formal framework, a manager chooses hoarding intensity β ≥ 0, where β &amp;gt; 0 reduces the equilibrium probability that a worker gets promoted. The definition encompasses all forms of managerial action that reduce worker departure probability, including suppressing visibility, restricting access to trainings, explicit discouragement, and threats.&lt;/p&gt;
&lt;p&gt;Q: Why do managers have an incentive to hoard talent?
A: Managers are compensated based on team performance, so losing a high-productivity worker (whose replacement is a random draw from an outside distribution with expected productivity ᾱ) reduces team performance and thus manager compensation. The framework shows that when a worker&amp;rsquo;s productivity αi exceeds the expected productivity of an outside hire ᾱ, the manager optimally sets β* &amp;gt; 0. The cost of hoarding (parameterized as φm) is convex and varies across managers, capturing altruism, reputation risk, or detection probability.&lt;/p&gt;
&lt;p&gt;Q: What share of managers in the survey self-report talent hoarding?
A: 75% of managers reported that they sometimes find themselves in situations where they need to dissuade a team member from exploring opportunities in another department due to immediate team needs or performance goals. Additionally, 45% cite the risk of losing talent as a reason not to invest in employee career development, and 66% cite the need to prioritize short-term performance targets over long-term employee development.&lt;/p&gt;
&lt;p&gt;Q: How are misaligned incentives documented in the manager survey?
A: 55% of managers agree or strongly agree that talent development entails a conflict of interest because more developed workers are more likely to leave the team. While 96% believe their direct intervention has a large impact on workers&amp;rsquo; career development, only 36% perceive that impact to be valued by the firm as much as team performance impact. Similarly, 87% say talent development is a high-impact area for the firm, but only 40% believe a track record in talent development matters for their own compensation and promotion.&lt;/p&gt;
&lt;p&gt;Q: How is the administrative measure of talent hoarding constructed?
A: The measure is the residual from an OLS regression of a worker&amp;rsquo;s potential rating (a public signal of promotion readiness, widely circulated within the firm) on their performance rating (a private signal of current task performance, not shared outside the team) and worker characteristics including age, education, gender, and tenure. The manager-level measure is the average of these residuals across all workers and quarters under that manager. Managers in the top tercile (mean deviation above 0.1036) are classified as hoarding-prone.&lt;/p&gt;
&lt;p&gt;Q: Does the hoarding measure respond to the incentive proxies as predicted by the framework?
A: Yes. A 1 percentage point higher share of performance-related compensation is associated with a 0.19 percentage point increase in the probability of being classified as hoarding-prone (p = 0.000), corresponding to a 13 percentage point difference between the 90th and 10th percentiles of the financial incentive distribution. A one-person increase in team size reduces hoarding probability by 1.3 percentage points (p = 0.000), again a 13 percentage point difference across percentiles. Managers in low-visibility functional areas are 4.0 percentage points more likely to hoard (p = 0.002) relative to high-visibility areas.&lt;/p&gt;
&lt;p&gt;Q: Is the training-based hoarding measure consistent with the potential-rating measure?
A: Yes. A complementary measure based on managers restricting worker access to high-visibility in-person trainings yields nearly identical patterns: a 1 percentage point increase in performance-related pay increases hoarding probability by 0.20 percentage points (p = 0.000); a one-person increase in team size reduces it by 1.4 percentage points (p = 0.000); low-visibility areas increase hoarding by 2.98 percentage points (p = 0.021). The direction and economic magnitudes are highly similar across both administrative measures and the survey-based measures.&lt;/p&gt;
&lt;p&gt;Q: How are manager rotations used to identify causal effects on workers?
A: When a manager learns they will move to a different position — typically two to three quarters before the rotation — their incentive to hoard workers on their current team ceases. This creates a quasi-random window of reduced talent hoarding for workers on that team. An event study with worker and quarter fixed effects shows flat pre-trends in application rates beyond three quarters before the rotation, consistent with the identifying assumption that managers do not yet know about their rotation in that earlier window. Balance tests confirm workers exposed to rotations are observationally similar on demographics and past performance to non-exposed workers.&lt;/p&gt;
&lt;p&gt;Q: How large is the effect of manager rotations on worker applications?
A: Manager rotations increase worker application rates by 2.3 percentage points in the quarter of rotation, representing a 78% increase over the baseline application rate of 2.9%. The effect is transitory: application rates return to baseline within one quarter after the new manager settles in. The effect is not driven by managers taking subordinates with them (97% of applications are to positions outside both the current team and the manager&amp;rsquo;s new team).&lt;/p&gt;
&lt;p&gt;Q: Does the rotation effect vary with predicted hoarding intensity as the framework requires?
A: Yes. The rotation effect is larger for workers with higher productivity, those whose replacement would be costlier (consistent with the prediction that workers harder to replace face more hoarding), and those working under managers with lower utility costs of hoarding. The paper tests these cross-sectional predictions using continuous interactions between the rotation indicator and standardized proxies for hoarding intensity, and all patterns are consistent with the talent hoarding mechanism rather than alternative explanations.&lt;/p&gt;
&lt;p&gt;Q: How successful would the deterred applicants have been?
A: Marginal applicants — those induced to apply by the manager rotation who would not otherwise have applied, identified via IV assumptions — face a hiring probability of 49.1%, compared to the average hiring likelihood of 27.6% across all applicants. This large positive selection implies that a substantial share of deterred applicants would have been successful, and that talent hoarding meaningfully degrades the quality and quantity of the firm&amp;rsquo;s internal applicant pool and the firm&amp;rsquo;s ability to promote high-productivity workers.&lt;/p&gt;
&lt;p&gt;Q: Does talent hoarding have differential effects by gender?
A: Yes. Women are 22% more likely to place high value on preserving a good relationship with their manager and 26% more likely to rely on manager career guidance when making career decisions. Consistent with this, marginal female applicants are more positively selected on educational qualifications, past performance, and hiring probability for higher-level positions than marginal male applicants. When comparing potential earnings outcomes, both men and women would earn more in the absence of talent hoarding, but the larger earnings gains for women imply a counterfactual reduction in the gender pay gap of 86%.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports external validity of the findings?
A: The firm&amp;rsquo;s employee demographics closely match those of large manufacturing firms in the German BiBB workforce survey across gender, age, citizenship, and marital status. The firm&amp;rsquo;s internal labor market design is standard for large German firms, where 83% of top publicly listed companies cite talent hoarding as a key organizational friction. Industry surveys also report that half of 665 global organizations report managers hoarding talent by discouraging worker mobility, and talent hoarding occurs through many of the same behaviors documented in this study.&lt;/p&gt;
&lt;p&gt;Q: How does the paper rule out confounding mechanisms for the rotation effect?
A: The paper tests and rules out several alternatives: worker-manager specific match effects (the effect does not depend on characteristics of the incoming or outgoing manager); finite project timelines driving a rush to apply; and workers being recruited by managers to their new teams (97% of applications are outside the current team and not to the manager&amp;rsquo;s new team). Balance tests show workers exposed to rotations are observationally similar to non-exposed workers, and event studies confirm absence of pre-trends in team-level outcomes including absenteeism.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: The findings suggest firms forgo productivity gains when hoarded workers are not allocated to positions where they would be most productive. Potential organizational responses include monitoring or rewarding managers for promoting talent, reducing performance-related pay tied to team composition, or structuring career development activities in ways that cannot easily be suppressed by individual managers. The paper notes that firms generally do not compensate managers for promoting workers, partly due to practical difficulties of such contracts, and that the misalignment between what managers believe benefits the firm and what is recognized in their own compensation is particularly pronounced for talent development relative to all other managerial responsibilities.&lt;/p&gt;
&lt;p&gt;Talent hoarding: Actions taken by managers that lower the likelihood that a worker applies for and receives a promotion or internal transfer outside the team, driven by managers&amp;rsquo; incentive to retain productive workers to protect team performance and manager compensation. Distinct from mere neglect — it is strategic and deliberate.&lt;/p&gt;
&lt;p&gt;Potential rating: A public signal of a worker&amp;rsquo;s future potential for higher-level positions, assigned by the direct supervisor and widely circulated within the firm (e.g., via HR lists of high-potential workers); distinguished from performance ratings by its visibility outside the worker&amp;rsquo;s current team, making it a lever for strategic manipulation by hoarding managers.&lt;/p&gt;
&lt;p&gt;Performance rating: A private, task-specific signal of a worker&amp;rsquo;s past performance in their current position, not shared with other units in the firm; used as the baseline against which potential ratings are compared in the paper&amp;rsquo;s administrative hoarding measure.&lt;/p&gt;
&lt;p&gt;Visibility suppression (hoarding measure): The manager-level average residual from a regression of workers&amp;rsquo; potential ratings on their performance ratings and worker characteristics; a positive average residual indicates the manager systematically assigns lower potential ratings than predicted, suppressing worker visibility outside the team in a manner consistent with strategic talent hoarding.&lt;/p&gt;
&lt;p&gt;Manager rotation: An event in which a manager leaves their current team for a different internal position within the firm, temporarily eliminating their hoarding incentive for current team workers and creating the paper&amp;rsquo;s quasi-experimental source of variation in hoarding exposure.&lt;/p&gt;
&lt;p&gt;Marginal applicant: In the IV framework, a worker who applies for an internal position only because their manager is rotating and would not have applied otherwise; estimated via complier analysis (Abadie 2003) and used to characterize the counterfactual quality and hiring probability of workers deterred by talent hoarding.&lt;/p&gt;
&lt;p&gt;Utility cost of hoarding (φm): A manager-level parameter capturing the convex private cost to a manager of engaging in talent hoarding; may reflect altruism, detection risk, or reputational consequences; managers with lower φm hoard more intensively, and variation in φm is proxied empirically by performance-related pay, team size, and functional-area talent visibility.&lt;/p&gt;</description></item><item><title>The Architecture of Social Networks and the Diffusion of Innovations</title><link>https://macropaperwarehouse.com/papers/the-architecture-of-social-networks-and-the-diffusion-of-innovations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-architecture-of-social-networks-and-the-diffusion-of-innovations/</guid><description>&lt;p&gt;This paper examines how the architecture of social networks shapes the success or failure of technology diffusion when adoption decisions exhibit strategic complementarities. The research question is: which structural feature of a network determines whether a new technology spreads or fails, and in which direction does that feature work?&lt;/p&gt;
&lt;p&gt;The paper builds on the canonical threshold diffusion model of Morris (2000) and Granovetter (1978), in which an agent adopts a new technology if the share of his neighbors who have adopted exceeds a threshold Q in [0,1]. The key innovation is the addition of a second structural object — a set of decision-making units C — that captures the empirically common phenomenon that subsets of agents (friends, family, neighbors, colleagues) can coordinate and make joint adoption decisions. The model is purely theoretical; the paper derives characterizations and comparison theorems rather than estimating parameters from data.&lt;/p&gt;
&lt;p&gt;The central structural concept introduced is insularity: the extent to which agents concentrate their connections to a narrow set of other agents, rather than distributing connections broadly. A formal partial order over networks is defined: network {w̃} is less insular than network {w} if there is no local increase in insularity in {w̃} relative to {w}, where a local increase in insularity occurs when one agent&amp;rsquo;s proportionate connections to a narrow set S are strictly higher and another agent&amp;rsquo;s proportionate connections to a superset R are strictly lower (with the first agent&amp;rsquo;s share of S weakly exceeding the second agent&amp;rsquo;s share of R). Moving from a network toward a convex combination with the complete network strictly reduces insularity under this definition (Lemma 3).&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s main characterization result (Proposition 1) establishes that the set of non-adopters of technology Q is precisely SQ — the maximal (1−Q)-subgroup-cohesive set — defined as the largest set in which every decision-making unit C contained in SQ has at least one agent with at least fraction (1−Q) of his connections inside SQ. This extends Morris&amp;rsquo;s (2000) cohesion characterization to the joint-decision setting.&lt;/p&gt;
&lt;p&gt;The main theorem (Theorem 1) establishes that for any two societies sharing the same decision-making structure C but differing in network insularity, there exists a cutoff threshold mu in [0,1] such that: (i) for technologies with Q &amp;lt; mu, adoption is weakly higher in the less insular network; and (ii) for technologies with Q &amp;gt;= mu, adoption is weakly lower in the less insular network. The direction reversal at mu reflects two competing mechanisms. Insular connections hinder singleton diffusion: an agent over-connected to a narrow set will not adopt individually until others in that set adopt, blocking entry of the technology from outside. But insular connections facilitate joint adoption: the same over-connectedness makes it profitable for the group to adopt together if they can coordinate, because each member already has a high share of neighbors within the group. High-threshold technologies depend crucially on joint adoption cascades and so benefit from insularity; low-threshold technologies spread person-to-person and are impeded by insularity when agents cannot coordinate.&lt;/p&gt;
&lt;p&gt;Proposition 2 establishes a complementary monotonicity result: expanding the set of decision-making units (C subset of C&amp;rsquo;) weakly increases adoption for any technology and any network, because joint decision-making resolves local coordination failures.&lt;/p&gt;
&lt;p&gt;The main result is extended to heterogeneous thresholds (Section 7). Proposition 3 shows that Theorem 1 continues to hold when agent-specific idiosyncratic components theta_i are bounded within an interval [−gamma/2, gamma/2] for some gamma &amp;gt; 0. Proposition 4 characterizes the necessary conditions for the main result to break: the specification fails only if there exist two agents i and j with theta_i &amp;gt; theta_j + Q2 − Q1, meaning the idiosyncratic gap between them exceeds the difference between the two technology thresholds being compared.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s central research question?
A: The paper asks how the architecture of a social network — specifically the structure of agents&amp;rsquo; connections — determines whether a new technology spreads widely or fails to diffuse. It focuses on technologies with strategic complementarities, where an agent&amp;rsquo;s benefit from adopting depends on neighbors adopting and those neighbors&amp;rsquo; benefit depends on their neighbors, creating potential for both snowballing and coordination failure.&lt;/p&gt;
&lt;p&gt;Q: What is the key modeling innovation relative to the standard threshold model?
A: The paper adds a set of decision-making units C, a collection of subsets of agents each of which can make a joint adoption decision. In the standard Morris (2000) model, only individual agents decide; here, groups such as friends, family, or neighbors can collectively agree to adopt, resolving their local coordination problem. The set C is subject only to closure under subsets and inclusion of all singletons, making the framework highly flexible.&lt;/p&gt;
&lt;p&gt;Q: How does the diffusion process work formally?
A: At each period t &amp;gt;= 1, agent i adopts if either: (1) more than fraction Q of his neighbors adopted in period t−1 (singleton adoption), or (2) i belongs to a decision-making unit C not yet adopted, and for every j in C the fraction of j&amp;rsquo;s neighbors in A_{t−1} union C exceeds Q (joint adoption). Actions are irreversible, and Appendix C proves this irreversibility assumption is without loss of generality for the final adoption set under myopic best-response dynamics.&lt;/p&gt;
&lt;p&gt;Q: What is the characterization of non-adopters (Proposition 1)?
A: The set of agents who do not adopt technology Q equals SQ, the unique maximal (1−Q)-subgroup-cohesive set — the largest set S such that every decision-making unit C contained in S has at least one member i with Pi(S minus C) &amp;gt;= (1−Q), meaning at least fraction (1−Q) of i&amp;rsquo;s connections remain inside S outside of C. This extends Morris (2000)&amp;rsquo;s p-cohesion concept: when C contains only singletons, (1−Q)-subgroup cohesion collapses to (1−Q)-cohesion in Morris&amp;rsquo;s sense.&lt;/p&gt;
&lt;p&gt;Q: What does the simple eight-agent example illustrate?
A: With two four-clique subgraphs (agents 1-4 and 5-8), Network A has agents 1, 3, 5, 7 each holding 3/4 of their connections within their four-agent group; Network B reduces those within-group shares to 5/8 by weakening two within-group links from weight 1 to weight 1/2 and adding cross-group links of weight 1/2. For Q = 3/10: in Network B all eight agents adopt (group {1,2,3,4} adopts jointly at t=1, then agents 5 and 7 adopt as singletons at t=2, agents 6 and 8 at t=3), while in Network A only {1,2,3,4} adopt (agents 5-8 each have only 1/4 of neighbors adopted, below Q = 3/10). For Q = 7/10: in Network A group {1,2,3,4} adopts jointly (each has 3/4 &amp;gt; 7/10 of neighbors adopting), while in Network B there is zero adoption (agent 3 has only 5/8 &amp;lt; 7/10 of neighbors in the joint group). This is the concrete illustration of the threshold-dependent reversal in Theorem 1.&lt;/p&gt;
&lt;p&gt;Q: What is insularity and how is it formally defined?
A: Insularity is the extent to which agents concentrate their connections to a narrow set of others. A local increase in insularity in {w} relative to {w̃} occurs when, for some agents i and j and sets S subset of R: (1) Pi(S) is strictly higher in {w} and Pj(R) is strictly lower in {w}, and (2) Pi(S) &amp;gt;= Pj(R) in {w}. Network {w̃} is less insular than {w} if no local increase in insularity exists in {w̃} relative to {w}. Lemma 3 establishes that the lambda-convex combination of any non-complete network with the complete network is strictly less insular.&lt;/p&gt;
&lt;p&gt;Q: What is the main theorem (Theorem 1) and its precise statement?
A: For two societies sharing the same decision-making structure C but differing in network insularity — with {w̃} strictly less insular than {w} — there exists a cutoff mu in [0,1] such that: for Q &amp;lt; mu, adoption is weakly higher in the less insular network; and for Q &amp;gt;= mu, adoption is weakly lower in the less insular network. The cutoff mu depends on the specific networks and decision-making structure. The result is a clean reversal: less insular is better for low-threshold technologies and worse for high-threshold technologies.&lt;/p&gt;
&lt;p&gt;Q: What are the two competing mechanisms driving Theorem 1?
A: First, insular connections hinder individual diffusion: an agent with a high share of connections concentrated inside a set will not adopt as a singleton until others in that set adopt, blocking entry of the technology from outside via individual contagion. Second, insular connections facilitate joint adoption: precisely because an agent has a high share of connections to a narrow group, jointly adopting with that group is profitable — each member has enough neighbors already within the group to exceed the threshold when the group adopts together. For high-threshold technologies, joint adoption is the only viable mechanism, so the second effect dominates; for low-threshold technologies, singleton diffusion suffices and the first effect dominates.&lt;/p&gt;
&lt;p&gt;Q: How does joint decision-making affect adoption (Proposition 2)?
A: Expanding the set of decision-making units from C to any C&amp;rsquo; containing C weakly increases adoption of technology Q for any network and any Q. The proof shows that the non-adopter set SQ under C&amp;rsquo; is also (1−Q)-subgroup cohesive under C, making it a subset of non-adopters under C. The economic logic is that any group able to make a joint decision can solve its local coordination problem: agents who individually would not adopt because too few neighbors have adopted may collectively adopt if each would benefit from group adoption.&lt;/p&gt;
&lt;p&gt;Q: How robust is Theorem 1 to heterogeneous thresholds?
A: Proposition 3 shows that Theorem 1 extends with the same cutoff structure when each agent i has an idiosyncratic threshold component theta_i in [−gamma/2, gamma/2] for sufficiently small gamma &amp;gt; 0. Proposition 4 establishes the necessary condition for the result to break with unbounded heterogeneity: there must exist agents i and j with theta_i &amp;gt; theta_j + Q2 − Q1, meaning the idiosyncratic gap must strictly exceed the technology threshold gap being compared. The underlying intuition of Theorem 1 persists even when the precise specification fails.&lt;/p&gt;
&lt;p&gt;Q: What are the policy and managerial implications?
A: A firm with a low-threshold technology should target less insular societies to maximize uptake, while a firm with a high-threshold technology should target more insular societies; the paper cites Facebook&amp;rsquo;s initial launch within closed university networks as consistent with the high-threshold logic. Policymakers and firms can increase adoption by encouraging joint decision-making — sanitation campaigns that organize neighborhood workshops, family mobile-plan discounts, or online coordination platforms all work through this channel. Conversely, governments trying to suppress collective action such as protest can prohibit in-person gatherings or online communication to prevent joint decision-making. The paper notes results abstract from seeding, leaving optimal seeding under joint decision-making as a future research direction.&lt;/p&gt;
&lt;p&gt;Insularity: The extent to which agents concentrate their connections to a narrow set of other agents rather than distributing connections broadly; formally defined via a partial order based on local increases in agents&amp;rsquo; proportionate connections to nested sets S subset of R.&lt;/p&gt;
&lt;p&gt;Decision-making unit: A set C of agents who can make a joint decision to adopt together; the collection C of all decision-making units is closed under subsets and contains all singletons, capturing informal group coordination among friends, family, or neighbors.&lt;/p&gt;
&lt;p&gt;p-Subgroup cohesion: A set S is p-subgroup cohesive if every decision-making unit C contained in S (of any size, including singletons) is p-connected in S — meaning at least one agent in C has at least fraction p of his connections to S minus C; the paper&amp;rsquo;s generalization of Morris (2000)&amp;rsquo;s p-cohesion to settings with joint decision-making.&lt;/p&gt;
&lt;p&gt;Threshold of adoption (Q): A parameter Q in [0,1] summarizing a technology&amp;rsquo;s strategic complementarities, such that an agent is better off adopting if and only if more than fraction Q of his neighbors adopt; low Q means the technology is valuable even with few adopters, high Q means it requires near-universal neighborhood adoption.&lt;/p&gt;
&lt;p&gt;Local increase in insularity: A pairwise comparison between two networks: {w} exhibits a local increase in insularity relative to {w̃} when one agent&amp;rsquo;s proportionate connections to narrow set S are strictly higher and another agent&amp;rsquo;s proportionate connections to superset R are strictly lower in {w}, with the first agent&amp;rsquo;s share of S weakly exceeding the second agent&amp;rsquo;s share of R in {w}.&lt;/p&gt;
&lt;p&gt;SQ (maximal non-adopter set): The unique maximal (1−Q)-subgroup-cohesive set in a society, constituting exactly the agents who do not adopt technology Q in the final outcome; it is the union of all (1−Q)-subgroup-cohesive sets and is itself (1−Q)-subgroup-cohesive (Lemma 2, Proposition 1).&lt;/p&gt;</description></item><item><title>The Confederate Diaspora</title><link>https://macropaperwarehouse.com/papers/the-confederate-diaspora/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-confederate-diaspora/</guid><description>&lt;p&gt;This paper investigates how white migration out of the postbellum South diffused Confederate culture and entrenched racial norms across the United States during a critical juncture of westward expansion and post-Civil War reconciliation. The central question is whether the &amp;ldquo;Confederate diaspora&amp;rdquo; — Southern white migrants who left the former Confederacy from 1870 to 1900 — causally shaped the geography of Confederate memorialization, white supremacist organizations, racial violence, and long-run racial inequity outside the South.&lt;/p&gt;
&lt;p&gt;Using complete-count U.S. Census records from 1870–1900 and linked Census records from the Census Linking Project, the authors track nearly one million white migrants from former Confederate states, including more than 61,000 former enslavers and 127,000 of their household kin, who settled outside the South by 1900. By 1900, migrants from the former Confederacy comprised on average 2.2% of the population in destination counties. Four outcomes measuring Confederate culture at the county level are constructed: Confederate memorialization (monuments, place names, schools), United Daughters of the Confederacy (UDC) chapters, Ku Klux Klan (KKK) chapters, and lynchings of Black people.&lt;/p&gt;
&lt;p&gt;The primary identification strategy is a shift-share instrumental variable (SSIV) that combines the cross-sectional distribution of Southern white migrants across non-Southern counties in 1870 (shares) with predicted migration flows out of each Southern state between 1870 and 1900 (shifts). The predicted shifts are constructed from origin-county economic and ideological push factors estimated via LASSO, insulating the IV from endogenous location sorting. Conditional on the 1870 Southern white population share, the SSIV identifies the distinct causal influence of the postbellum Confederate diaspora.&lt;/p&gt;
&lt;p&gt;Main findings are large relative to the diaspora&amp;rsquo;s modest population share. Moving from zero to the mean Confederate diaspora share implies an 8 percentage point (p.p.) increase in the likelihood of KKK activity relative to a mean prevalence of 35% in non-Southern counties. Effects on post-1900 lynching events are even larger proportionally: a 4 p.p. increase in likelihood relative to a mean of only 5%. IV estimates for Confederate memorialization show that a 1 p.p. increase in the Southern white share in 1900 raised the likelihood of memorialization by 3.4 p.p. (after controlling for the 1870 share), relative to a baseline prevalence of 25% outside the South. Effects on UDC chapters are similarly large given the organization&amp;rsquo;s limited non-Southern footprint (present in only 10% of counties). IV estimates consistently exceed OLS estimates, consistent with economic sorting biasing OLS downward.&lt;/p&gt;
&lt;p&gt;Beyond Confederate symbolism, the diaspora also contributed to a novel form of racial exclusion: the &amp;ldquo;sundown town.&amp;rdquo; A 1 p.p. increase in the Confederate diaspora share in 1900 led to a 2.4 p.p. increase in the likelihood of Black depopulation (defined as towns with at least 25 Black residents in 1870 having zero Black residents after 1900).&lt;/p&gt;
&lt;p&gt;Former slaveholders, though only about 6% of Confederate migrants, played an outsized role. They disproportionately sorted into frontier counties and into positions of public authority — more than twice as likely to work as lawyers or judges and nearly three times as likely to work in public administration as the average non-slaveholding Southern white migrant. Their cultural influence was especially pronounced in frontier communities where institutions were weak and norms malleable. In Denver, first-generation Southern white migrants were 11% more likely to join the KKK than men with no Southern heritage, with a similar differential observed for second-generation migrants.&lt;/p&gt;
&lt;p&gt;The diaspora&amp;rsquo;s effects persist into the 21st century: counties with larger Confederate diasporas in 1900 exhibit larger racial wage gaps, greater residential segregation, higher rates of Black incarceration, higher rates of police-induced Black mortality, and more conservative racial attitudes among whites, as measured in modern survey data. These long-run findings are identified using the same county-level SSIV strategy. Scope conditions: effects are larger in frontier counties (weaker institutions, more malleable norms), in counties with fewer Union Army enlistees, and in newly incorporated areas with fewer than 2 residents per square mile in 1860.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question and why does it matter?
A: The paper asks whether postbellum Southern white migration causally diffused Confederate culture — memorialization, organized white supremacy, and racial violence — beyond the South, and whether this early cultural transplantation has persistent effects on racial inequity today. It matters because Confederate monuments and persistent Black disadvantage in labor, housing, and policing are often attributed to the legacies of slavery within the South; this paper shows the mechanism by which those norms spread nationally through internal migration at a critical juncture of westward expansion and post-war reconciliation.&lt;/p&gt;
&lt;p&gt;Q: How large was the Confederate diaspora, and who comprised it?
A: Estimates from linked Census records suggest that nearly one million whites left the former Confederacy for the rest of the U.S. in the three decades after the war, including more than 61,000 former enslavers and 127,000 of their household kin. By 1900, migrants from the former Confederacy averaged 2.2% of the population in non-Southern destination counties. The diaspora hailed primarily from the upper South — Virginia, Tennessee, and North Carolina — and later from Texas, Arkansas, and Oklahoma.&lt;/p&gt;
&lt;p&gt;Q: How do the authors construct the shift-share instrumental variable, and what identifying assumption does it require?
A: The SSIV multiplies each Southern origin state&amp;rsquo;s 1870 settlement shares across non-Southern counties (the shares) by predicted total Southern white outflows from 1870 to 1900 (the shifts), where the predicted shifts are constructed by summing LASSO-selected origin-county push factors — economic conditions, cotton and tobacco potential, Civil War battle locations, Black population share — rather than actual flows. The exclusion restriction requires that these predicted push-factor-driven outflows affect destination county outcomes only through the Confederate diaspora they deliver, not through direct economic linkages with origin counties. Conditioning on the 1870 Southern white share absorbs time-invariant destination heterogeneity correlated with antebellum settlement.&lt;/p&gt;
&lt;p&gt;Q: What are the IV estimates for Confederate memorialization and UDC chapters?
A: A 1 p.p. increase in the Southern white share in 1900 raised the likelihood of Confederate memorialization by 3.4 p.p. after controlling for the 1870 share (relative to a baseline prevalence of 25% outside the South). For UDC chapters, which were present in only 10% of non-Southern counties, IV estimates show similar or larger proportional effect sizes. IV estimates are consistently more than twice the size of OLS estimates, consistent with downward bias from economic sorting of Southern whites toward productive, culturally-diverse destinations.&lt;/p&gt;
&lt;p&gt;Q: What are the IV estimates for KKK activity and Black lynchings, and how are they interpreted?
A: A 1 p.p. increase in the Southern white share in 1900 raised the likelihood of KKK chapter presence by 3.5 p.p. (controlling for 1870 shares), relative to a mean KKK prevalence of 37% in non-Southern counties, implying that moving from zero to the mean diaspora share is associated with an 8 p.p. increase in the probability of KKK activity. For Black lynchings, the corresponding IV estimate is 1.5 p.p. (column 5), with the effect rising when earlier migration is controlled, against a mean prevalence of only 5% — implying moving from zero to the mean raises lynching likelihood by 4 p.p. Critically, the authors find no diaspora effect on white lynchings, which distinguishes racially-targeted violence from a generalized Southern culture of violence.&lt;/p&gt;
&lt;p&gt;Q: What is a &amp;ldquo;sundown town&amp;rdquo; and what does the paper find about the diaspora&amp;rsquo;s role in producing them?
A: Sundown towns, described in historical research by Loewen (2005), are all-white towns where Black residents and other minorities were excluded from residing after sunset, spreading throughout the non-South from 1890 to 1960 and representing a novel form of racial exclusion distinct from de jure Jim Crow institutions. The authors find that a 1 p.p. increase in the size of the Confederate diaspora in 1900 led to a 2.4 p.p. increase in the likelihood of Black depopulation — defined as towns with at least 25 Black residents in 1870 having zero Black residents after 1900 — changing the geography of Black settlement throughout the 20th century.&lt;/p&gt;
&lt;p&gt;Q: What role did former slaveholders specifically play, and how are their effects separately identified?
A: Former slaveholders comprised just over 6% of the Confederate migrant sample but played an outsized role: they were about 50% more likely than the average Southern white migrant to work in any public-facing authority occupation, more than twice as likely to work as lawyers or judges, and nearly three times as likely to work in public administration. Their effects are identified using an analogous SSIV that, conditional on the instrumented overall diaspora, draws on distinct identifying variation in slaveholder-specific push factors. Former slaveholders gravitated toward Western, lower-density, cotton-suitable counties with higher Breckinridge vote shares and fewer Union Army soldiers, consistent with seeking to reconstruct antebellum hierarchies in malleable frontier spaces.&lt;/p&gt;
&lt;p&gt;Q: Why were effects stronger in frontier counties?
A: The paper finds that diaspora impacts on Confederate culture diffusion were significantly larger in counties along the frontier, where state institutions were weak and cultural norms not yet deeply ingrained. Restricting the sample to counties with fewer than 2 residents per square mile in the 1860 Census yields somewhat larger estimates than baseline, and the differential sorting of Southern whites (especially former slaveholders) into these nascent communities suggests that institutional malleability amplified the cultural entrepreneurs&amp;rsquo; influence. Fewer Union Army enlistees in destination counties also amplified effects, as those families might otherwise have opposed resurgent Confederate ideology.&lt;/p&gt;
&lt;p&gt;Q: How did the diaspora transmit its norms to subsequent generations and non-Southern neighbors?
A: In the Denver metropolitan area, using newly digitized KKK membership records, first-generation Southern migrants were 11% more likely to join the KKK than men with no Southern heritage, and a similar differential holds for second-generation migrants (born in the diaspora), with patterns holding within Census enumeration blocks. White men without Southern heritage living next door to first- or second-generation Southern whites were significantly more likely to join the KKK, consistent with horizontal cultural spillovers. For naming patterns, non-Southern white parents who moved to counties with a larger Confederate diaspora gave their later-born children names more evocative of Confederate heroes than those given to earlier-born children — providing direct evidence of cultural spillovers beyond the diaspora.&lt;/p&gt;
&lt;p&gt;Q: What long-run effects of the diaspora are documented through the 21st century?
A: Using the county-level SSIV strategy, the paper finds that a larger Confederate diaspora in 1900 is associated with larger racial wage gaps, greater residential segregation, higher rates of Black incarceration, and higher rates of police-induced Black mortality through the 21st century. These disparities are mirrored in more conservative racial attitudes among whites in these counties as measured in modern survey data. These persistent effects suggest that, despite racially progressive national policy reform since the 1960s, locally institutionalized mechanisms reinforced by a culture of racial animus continue to generate inequity.&lt;/p&gt;
&lt;p&gt;Q: How robust are the main estimates to alternative specifications?
A: The authors show robustness across: (i) alternative spatial standard errors using Conley (1999) distance-based clustering and Adao et al. (2019) shift-share inference corrections; (ii) Belloni et al. (2014) double LASSO control selection; (iii) replacing predicted shifts with actual shifts; (iv) a random-shifts placebo where fewer than 5% of coefficients are significant; (v) dropping individual origin or destination states one-by-one (all estimates remain significant with 97% positive Rotemberg weights); (vi) excluding border states with antebellum slavery (Delaware, Kentucky, Maryland, Missouri, West Virginia), which actually increases estimates; and (vii) restricting to newly incorporated counties with near-zero 1860 populations, which yields somewhat larger effects.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s contribution to the culture-institutions literature?
A: The paper uses granular data on migration, occupational choices, and local governance to shed light on the historical process by which Confederate &amp;ldquo;cultural entrepreneurs&amp;rdquo; captured early institutions across America, illustrating how culture and institutions reinforce each other during critical junctures of nation-building. The findings suggest that laws to reduce racial discrimination may have limited impact where a culture of racial animus is ingrained in local institutions — an institutionalized persistence mechanism that helps explain the gap between formal legal reforms and observed racial outcomes. The paper also identifies a prestige-biased cultural transmission channel, consistent with Henrich and Gil-White (2001), wherein non-elite masses emulate former slaveowners in positions of power.&lt;/p&gt;
&lt;p&gt;Confederate diaspora: The approximately one million white migrants, including more than 61,000 former enslavers and 127,000 of their household kin, who left former Confederate states for the rest of the U.S. in the three decades after the Civil War, comprising on average 2.2% of destination county populations by 1900 and retaining strong cultural attachments to the Confederacy.&lt;/p&gt;
&lt;p&gt;Confederate culture: A cluster of symbolic and material expressions that coalesced in the postbellum South, encompassing Lost Cause narratives (glorifying Confederate figures and reframing secession as a defense of states&amp;rsquo; rights rather than slavery), public memorialization (monuments, place names, school names), United Daughters of the Confederacy chapters, Ku Klux Klan activity, and lynchings of Black people — together functioning as technologies to transmit white supremacist norms and maintain racial hierarchies.&lt;/p&gt;
&lt;p&gt;Lost Cause: A revisionist narrative emerging after the Civil War that sought to redeem the image of the South by offering noble rationalizations for secession — emphasizing Northern aggression and states&amp;rsquo; rights while downplaying slavery — and portraying enslaved people as content and slaveowners as generously paternalistic; central to the ideology propagated by the UDC and to Confederate memorialization.&lt;/p&gt;
&lt;p&gt;Shift-share instrumental variable (SSIV): An identification strategy that combines the 1870 distribution of Southern white migrants across non-Southern counties (shares, reflecting historical migration networks) with predicted total Southern white outflows from 1870 to 1900 constructed from origin-county push factors via LASSO (shifts), to isolate exogenous county-level variation in Confederate diaspora exposure that is insulated from endogenous location sorting.&lt;/p&gt;
&lt;p&gt;Sundown town: An all-white municipality where Black residents and other minorities were excluded from residing after sunset, spreading throughout the non-South from 1890 to 1960, operationalized in this paper as towns with at least 25 Black residents in 1870 having zero Black residents after 1900 (Black depopulation), representing a novel form of racial exclusion distinct from de jure Jim Crow institutions associated with the Confederacy.&lt;/p&gt;
&lt;p&gt;Prestige-biased cultural transmission: An evolutionary transmission mechanism, formalized in Henrich and Gil-White (2001), in which non-elite populations emulate culturally salient leaders; invoked in this paper to explain how former slaveholders in positions of authority could diffuse Confederate norms to non-Southern whites who had no direct connection to the Confederacy.&lt;/p&gt;
&lt;p&gt;Cultural entrepreneur: A migrant (especially a former slaveholder) who, by sorting into positions of public-facing authority — judges, lawyers, law enforcement, clergy, public administrators — at early stages of community formation when institutions are most malleable, actively embeds cultural norms into nascent local institutions, amplifying influence beyond their small population share.&lt;/p&gt;</description></item><item><title>The Earnings and Labor Supply of U.S. Physicians</title><link>https://macropaperwarehouse.com/papers/the-earnings-and-labor-supply-of-u.s.-physicians/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-earnings-and-labor-supply-of-u.s.-physicians/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; What do U.S. physicians earn, how is that earnings variation structured across geography and specialty, and how much does government healthcare payment policy shape those earnings and — through them — physicians&amp;rsquo; labor supply and long-run talent allocation?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The paper builds a novel administrative panel by merging the universe of U.S. federal individual income tax returns (2005–2017) with: the National Plan and Provider Enumeration System (NPPES) physician registry; Medicare billing records with procedure-level Relative Value Unit (RVU) rates (2012–2017); restricted-use American Community Survey responses; Social Security Administration demographic records; and medical school ranking and graduation data. The main sample covers 11.6 million physician-year observations for 965,000 unique physicians aged 20–70.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Earnings Facts.&lt;/strong&gt; In 2017, average physician total individual income was $350,000 (median $265,000); the distribution is right-skewed — the top 1% of age-40–55 physicians averages $4.0 million. Physicians in aggregate earned $297 billion in pre-tax dollars, equaling 8.6% of total U.S. healthcare spending. The age-earnings profile is steep: earnings are approximately $60,000 during residency, rise to roughly $185,000 by the early thirties, and peak near $425,000 at age 50. Business income — systematically underreported in survey data (ACS estimates are approximately $140,000 lower than tax data during peak career years, almost entirely due to non-reporting of business income) — accounts for nearly one-quarter of earnings at age 50. Earnings differ sharply across specialties: primary care physicians average $201,200 (ages 40–55), about half the sample mean, while surgeons earn roughly twice as much.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Geographic Pattern.&lt;/strong&gt; Contrary to the pattern for lawyers and workers broadly, physician earnings are not highest on the coasts. A movers-based event study (physicians who changed commuting zones once during 2005–2017) finds that roughly 70% of the cross-location income difference is driven by place rather than worker composition. A two-way fixed-effects variance decomposition reveals pronounced negative physician-location sorting: high-earning physicians tend to locate in lower-income commuting zones, while lower-earning physicians locate in higher-income areas — the opposite of the pattern for lawyers. Medicare&amp;rsquo;s relatively weak adjustment of reimbursement rates for local costs (the empirical elasticity of the Geographic Adjustment Factor to median household income is 0.09, versus 0.33 for a broader local price index) can, by the authors&amp;rsquo; estimates, account for approximately one-third of this unusual geographic earnings pattern.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Government Influence — Medicare Price Changes.&lt;/strong&gt; Using procedure-specific RVU changes as a simulated instrument for each physician&amp;rsquo;s Medicare price exposure, the authors find that a 10% increase in the Medicare price instrument leads to a 2.4% increase in professional earnings of physicians aged 40–55. The behavioral supply response is substantial: physicians bill 4.4% more RVUs (supply elasticity of 0.4 after netting out the mechanical component), of which 3.9% reflects more unique procedures and the rest a shift toward higher-paid procedures. Nearly all of the procedure-level supply increase (3.4 out of 3.8 percentage points) comes from treating additional patients rather than more frequent treatment of existing patients. Converting to pass-through: physicians retain $62 of each $100 in additional Medicare spending directly, or approximately $25 of each $100 of any insurance spending once Medicare&amp;rsquo;s documented spillover into private insurance rates is accounted for. For physicians aged 56–70, a 10% increase in earnings driven by reimbursement changes reduces retirement probability by 0.5 percentage points in that year.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Government Influence — ACA Insurance Expansion.&lt;/strong&gt; Using county-level variation in pre-ACA uninsurance rates (as of 2013) as a source of differential exposure to the ACA&amp;rsquo;s Medicaid expansions and Marketplace subsidies (in 24 states expanding Medicaid in 2014 or early 2015), the authors estimate that a 10 percentage point higher baseline uninsurance rate led to 3.9% higher physician earnings four years post-expansion. Scaling by the first stage (a 10 p.p. higher uninsurance rate translating to 4.96 p.p. higher insurance coverage post-expansion), the implied elasticity of physician earnings to the insurance rate is 0.41. The ACA expansion also reduced retirement probability — a 10 p.p. higher insurance coverage rate leads to a 1 p.p. decline in retirement probability — consistent with a medium-run retirement-to-income elasticity of approximately −1.1. In aggregate, 6% of the $110 billion in annual ACA insurance expansion spending accrued to physicians personally, slightly below their 8.6% baseline share of healthcare spending.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Talent Allocation.&lt;/strong&gt; Specialty choice is sticky and entry-restricted. The authors estimate a discrete-choice model of specialty choice using graduates of top-5 medical schools — physicians with effectively unconstrained specialty access — and an aggregate model using USMLE Step 1 score buckets as ability proxies. At the top of the ability distribution, higher specialty earnings strongly attract physicians: increasing primary care physicians&amp;rsquo; hourly income from $98 to $168 per hour (the level of medicine subspecialists) would raise the share of top-5 medical school graduates choosing primary care by approximately 20 percentage points (nearly doubling their representation in primary care). Moving down the USMLE score distribution, the earnings coefficient falls monotonically and turns negative for the lowest score groups — consistent with the model&amp;rsquo;s prediction that entry restrictions cause higher-paying specialties to displace lower-ability applicants as earnings rise, rather than simply attracting more entrants. A more modest counterfactual — raising internal medicine earnings to dermatology levels — raises the average USMLE score in internal medicine by 10 points (from 230.2 to 239.6).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; The earnings estimates are for the period 2005–2017. Pass-through estimates use a short-run price instrument; long-run pass-through may differ depending on private market spillovers and entry. The ACA analysis is restricted to 24 early-expanding states. The specialty-choice model is estimated on medical graduates entering the residency match; the extensive margin of entering medicine itself is not modeled. Health outcome effects of changing physician ability distributions are not estimated.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-level-and-composition-of-physician-earnings-in-the-tax-data-and-how-do-they-compare-to-survey-based-estimates"&gt;Q1. What is the level and composition of physician earnings in the tax data, and how do they compare to survey-based estimates?&lt;/h3&gt;
&lt;p&gt;In 2017, average physician total individual income was $350,000 and median was $265,000; the top 1% of age-40–55 physicians earned $4.0 million on average, more than twice the average of the top 5%. Business income constitutes nearly one-quarter of earnings at age 50 and is concentrated among top earners: 80% of physicians in the top 1% have business income exceeding $25,000, versus 35% overall. ACS survey data for the same physicians underestimate earnings by approximately $140,000 (roughly one-third of the administrative mean) during peak career years, driven entirely by non-reporting of business income on the extensive margin.&lt;/p&gt;
&lt;h3 id="q2-what-share-of-total-us-healthcare-spending-do-physician-earnings-represent-and-what-does-this-imply-for-policy"&gt;Q2. What share of total U.S. healthcare spending do physician earnings represent, and what does this imply for policy?&lt;/h3&gt;
&lt;p&gt;Physicians in aggregate earned $297 billion pre-tax in 2017, equaling 8.6% of total U.S. healthcare spending (approximately $913 of the average American&amp;rsquo;s $10,611 annual healthcare expenditure). After applying a 30% income tax rate, after-tax physician earnings equal approximately 6% of total healthcare spending, or roughly 1% of GDP. The authors note this provides an upper bound on the magnitude of savings available from policies aimed at reducing physician incomes as a strategy for lowering overall healthcare spending.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-age-earnings-profile-of-physicians-evolve-and-what-drives-growth-during-peak-years"&gt;Q3. How does the age-earnings profile of physicians evolve, and what drives growth during peak years?&lt;/h3&gt;
&lt;p&gt;Physician earnings average approximately $60,000 during residency, rise to roughly $185,000 by the early thirties, and peak near $425,000 at age 50, before declining gradually to approximately $270,000 in the late 60s. Growth during peak earning years (ages 40–55) is driven almost entirely by business income: average wages are approximately flat at $285,000 across this age range, while business income and the probability of filing Schedule C rise steadily.&lt;/p&gt;
&lt;h3 id="q4-how-large-and-unusual-is-the-geographic-pattern-of-physician-earnings-and-what-is-the-causal-role-of-location"&gt;Q4. How large and unusual is the geographic pattern of physician earnings, and what is the causal role of location?&lt;/h3&gt;
&lt;p&gt;Physician earnings are highest in lower-income states (not on the coasts), unlike lawyers and the broader workforce. A movers event study finds that approximately 70% of the cross-commuting-zone income difference is attributable to location rather than worker characteristics; within specialty the estimate rises to approximately 85%. A two-way fixed-effects variance decomposition (with limited-mobility-bias corrections following Andrews et al. 2008 and Kline et al. 2020) reveals pronounced negative physician-location sorting, with the corrected covariance between individual and location effects being 0.6–0.8 times the variance of location effects in magnitude but opposite in sign — a pattern that reverses to positive sorting when the same methods are applied to lawyers.&lt;/p&gt;
&lt;h3 id="q5-what-instrument-is-used-to-identify-the-causal-effect-of-medicare-price-changes-on-physician-earnings-and-why-is-it-valid"&gt;Q5. What instrument is used to identify the causal effect of Medicare price changes on physician earnings, and why is it valid?&lt;/h3&gt;
&lt;p&gt;The authors construct a physician-year &amp;ldquo;Medicare price instrument&amp;rdquo; by fixing each physician&amp;rsquo;s service mix at its 2012–2017 average and then multiplying those fixed quantities by annually-updated RVU rates, summing over services. Because the fixed quantity weights exclude behavioral responses, and because national RVU changes from CMS periodic reviews affect physicians differentially according to their pre-determined service mix, variation across physicians and over time is plausibly exogenous to individual physicians&amp;rsquo; income shocks. Year-by-specialty fixed effects absorb common specialty-level price trends.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-magnitudes-of-the-earnings-and-labor-supply-responses-to-medicare-price-changes"&gt;Q6. What are the magnitudes of the earnings and labor supply responses to Medicare price changes?&lt;/h3&gt;
&lt;p&gt;A 10% increase in the Medicare price instrument raises earnings of 40–55 year-old physicians by 2.4% (reduced-form), with a 2SLS elasticity of income to billed RVUs of 0.17. The total-RVU billing coefficient of 1.437 implies a supply elasticity of 0.437 (subtracting 1 for the mechanical component). At the procedure level, a 10% price increase for a specific code leads to 3.8% more billings for that code, of which 3.4 percentage points reflects treating additional patients. For physicians aged 56–70, a 10% earnings increase reduces that year&amp;rsquo;s retirement probability by 0.5 percentage points.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-aca-insurance-expansion-affect-physician-earnings-and-retirement-and-what-is-the-implied-pass-through"&gt;Q7. How does the ACA insurance expansion affect physician earnings and retirement, and what is the implied pass-through?&lt;/h3&gt;
&lt;p&gt;Counties with a 10 percentage point higher pre-ACA uninsurance rate saw 3.9% higher physician earnings by 2017 (four years post-expansion). Scaled by the first stage (4.96 p.p. higher coverage), the elasticity of physician earnings to insurance coverage is 0.41. A 10 p.p. higher insurance coverage rate leads to a 1 p.p. lower retirement probability post-expansion (medium-run elasticity of retirement to income of approximately −1.1). In aggregate, 6% of $110 billion in annual ACA expansion spending — roughly $7.1 billion, or about $8,400 per physician — accrued to physicians.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-earnings-specialty-choice-relationship-vary-across-the-physician-ability-distribution"&gt;Q8. How does the earnings-specialty choice relationship vary across the physician ability distribution?&lt;/h3&gt;
&lt;p&gt;In the individual-level discrete-choice model estimated on top-5 medical school graduates (likely unconstrained in specialty choice), the coefficient on hourly earnings is 0.014. In the aggregate score-group model, the implied earnings coefficient is 0.016 for USMLE scores above 260 and declines monotonically to −0.008 for scores at or below 190. This negative coefficient for low scorers is consistent with the theoretical prediction that higher earnings attract high-ability physicians, leaving fewer slots for lower-ability applicants due to binding entry restrictions — not a reversal of preferences.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-quantitative-implications-for-specialty-choice-if-primary-care-incomes-were-raised-to-subspecialty-levels"&gt;Q9. What are the quantitative implications for specialty choice if primary care incomes were raised to subspecialty levels?&lt;/h3&gt;
&lt;p&gt;Raising primary care hourly income from $98 to $168 (the level of medicine subspecialists) would increase the share of top-5 medical school graduates choosing primary care by approximately 20 percentage points (about 48% would enter primary care, versus the current share), nearly doubling their representation. Nearly half of these reallocations would come from procedural specialties. An analogous exercise raising internal medicine earnings to dermatology levels shifts the average USMLE score in internal medicine from 230.2 to 239.6 — a 10-point increase — as higher-scoring applicants displace lower-scoring ones within a fixed slot constraint.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-pass-through-from-medicare-reimbursements-to-physician-earnings-and-how-does-it-compare-to-rent-sharing-elsewhere"&gt;Q10. What is the pass-through from Medicare reimbursements to physician earnings, and how does it compare to rent-sharing elsewhere?&lt;/h3&gt;
&lt;p&gt;Direct estimates imply physicians retain $62 of each $100 in additional Medicare spending. Accounting for Medicare&amp;rsquo;s documented spillover into private insurance rates (following Clemens and Gottlieb 2017), the pass-through drops to $25 per $100 of total insurance spending. The authors note this is substantially higher than the modest rent-sharing found for average workers in response to firm-level shocks (Card et al. 2018), but comparable to rent-sharing with high-skilled workers benefiting from patent rents (Kline et al. 2019).&lt;/p&gt;
&lt;h3 id="q11-can-medicares-geographic-pricing-policy-explain-the-unusual-geographic-earnings-pattern-for-physicians"&gt;Q11. Can Medicare&amp;rsquo;s geographic pricing policy explain the unusual geographic earnings pattern for physicians?&lt;/h3&gt;
&lt;p&gt;The elasticity of Medicare&amp;rsquo;s Geographic Adjustment Factor (GAF) to commuting zone median household income is 0.09, compared to 0.33 for a broader local price index. Using the authors&amp;rsquo; short-run estimate that a 10% increase in Medicare prices raises earnings by 2.4%, a counterfactual simulation shows that if the GAF-to-income elasticity rose to 0.33 (aligning Medicare rates with the general cost-of-living gradient), the geographic physician earnings pattern would more closely resemble that of lawyers. The authors estimate that the gap in Medicare&amp;rsquo;s local cost adjustment explains approximately one-third of the unusual physician earnings geography, conditional on the short-run pass-through estimate.&lt;/p&gt;
&lt;h3 id="q12-how-does-the-theoretical-model-of-specialty-choice-and-entry-restrictions-guide-the-empirical-predictions"&gt;Q12. How does the theoretical model of specialty choice and entry restrictions guide the empirical predictions?&lt;/h3&gt;
&lt;p&gt;The model features a unit mass of physicians with heterogeneous ability (Pareto-distributed) and idiosyncratic specialty preferences (exponentially distributed). Physicians choose whether to specialize in period 1; government sets reimbursement rates in period 2; physicians choose labor supply in period 3. With a fixed number of residency slots, higher specialty earnings raise the ability cutoff for entry (rationing by ability). This generates a key nonmonotonic empirical prediction: higher-ability physicians respond positively to earnings increases (choosing a specialty more frequently), while lower-ability physicians respond negatively (displaced by the shift upward in the ability cutoff). The model also implies that demand shocks are not moderated by contemporaneous entry, so incumbents capture the full rent — motivating the estimated pass-through.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Medicare Price Instrument (Simulated RVU Instrument).&lt;/strong&gt; A physician-year measure of Medicare payment exposure constructed by holding each physician&amp;rsquo;s service mix fixed at its 2012–2017 average and multiplying those fixed quantities by time-varying national RVU rates, then summing across services. This purges the instrument of behavioral responses, creating exogenous cross-physician variation in price exposure arising from the interaction of fixed service mix with national RVU policy changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Relative Value Unit (RVU).&lt;/strong&gt; The unit by which Medicare defines and reimburses each physician service in the Physician Fee Schedule. RVUs are intended to reflect the time, effort, and resources required to provide each service, but are subject to periodic review by CMS&amp;rsquo;s RVU Update Committee (RUC) and influenced by political factors. Changes in RVUs translate directly into changes in Medicare reimbursement rates for affected services.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pass-Through (Reimbursement to Earnings).&lt;/strong&gt; The share of an additional dollar of Medicare (or insurance) spending that accrues to physicians personally as earnings, after accounting for practice costs, intermediaries, and behavioral responses. The paper estimates $62 per $100 of direct Medicare spending or $25 per $100 of total insurance spending (the latter accounting for Medicare&amp;rsquo;s spillover into private rates).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negative Physician-Location Sorting.&lt;/strong&gt; The empirical finding — robust to limited-mobility-bias corrections — that higher-ability (higher-earning) physicians disproportionately locate in lower-income commuting zones, while lower-earning physicians concentrate in higher-income areas. This is the opposite of the pattern for lawyers and for worker-firm matching in the broader labor literature. The paper attributes part of this pattern to Medicare&amp;rsquo;s incomplete geographic adjustment of reimbursement rates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ability Cutoff (am) in Residency Matching.&lt;/strong&gt; In the paper&amp;rsquo;s theoretical model, the minimum ability level required to gain entry into a restricted-entry specialty. Because the number of residency slots is fixed, the cutoff rises when a specialty&amp;rsquo;s relative earnings increase (attracting more high-ability applicants), displacing lower-ability physicians who would otherwise have entered. This makes the earnings-specialty relationship nonmonotonic across the ability distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Business Income (Pass-Through Entity Income).&lt;/strong&gt; Income from physician-owned practices organized as sole proprietorships, S-corporations, or partnerships, reported on Schedule C or through pass-through entities rather than on Form W-2. In the tax data, business income accounts for nearly one-quarter of physician earnings at career peak and is the main source of earnings for top physicians, but is systematically underreported in survey data (ACS), leading to a roughly one-third underestimate of total earnings during peak years.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Geographic Adjustment Factor (GAF).&lt;/strong&gt; A Medicare policy parameter that multiplies the national RVU rate to adjust physician reimbursements for local input costs (specifically physicians&amp;rsquo; work, practice expenses, and malpractice). The paper documents that the GAF&amp;rsquo;s elasticity to local median household income is 0.09 — far below the 0.33 elasticity of the general local price index — constituting an effective subsidy to rural and lower-income markets relative to higher-income areas.&lt;/p&gt;</description></item><item><title>The Effect of Education Policy on Crime: An Intergenerational Perspective</title><link>https://macropaperwarehouse.com/papers/the-effect-of-education-policy-on-crime-an-intergenerational-perspective/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effect-of-education-policy-on-crime-an-intergenerational-perspective/</guid><description>&lt;p&gt;This paper studies the intergenerational effects of education policy on crime, asking whether a compulsory schooling reform that reduced crime among those directly exposed also reduced crime among their children. The authors exploit the staggered municipal rollout of Sweden&amp;rsquo;s comprehensive school reform, implemented gradually between 1949 and 1962 across more than 1,000 municipalities, which increased compulsory schooling by one to two years, abolished tracking into academic and vocational streams after 6th grade, and introduced a uniform national curriculum. The parent generation consists of all individuals born in Sweden between 1945 and 1955 (approximately 447,000 men and 450,000 women), and their children form the child generation (426,721 sons observed from age 15 to 29). Crime is measured by administrative conviction records from the Swedish National Council for Crime Prevention covering 1973–2010.&lt;/p&gt;
&lt;p&gt;The empirical strategy is difference-in-differences, comparing changes in conviction rates across cohorts in municipalities that implemented the reform at different times, with treatment assigned based on the parent&amp;rsquo;s birth municipality to avoid endogenous sorting bias. Standard errors are clustered at the municipality level. Parallel trends validity is supported by three tests: results are unchanged when municipality-specific linear trends are included, placebo tests using incorrect reform dates yield effects indistinguishable from zero, and residuals from crime regressions show no correlation with municipality-specific trends.&lt;/p&gt;
&lt;p&gt;The main finding is a significant 0.79 percentage point (pp) decline in conviction rates among sons of fathers exposed to the reform (p-value &amp;lt; 0.002), representing a 3.4 percent reduction relative to baseline. The decline spans multiple crime types: violent crime fell by 0.27 pp, traffic-related crime by 0.45 pp, fraud by 0.22 pp, and other offenses by 0.41 pp — percentage reductions of three to six percent across categories. Multiple convictions fell by 0.43 pp (5.8 percent). These second-generation effects are driven entirely by paternal exposure: the impact of maternal reform exposure is an order of magnitude smaller and statistically insignificant, and the difference between paternal and maternal effects is itself significant (p-value 0.048 for any conviction, 0.009 for multiple convictions). Effects on daughters in the child generation are much smaller, with only the residual &amp;ldquo;other crime&amp;rdquo; category showing a significant 0.129 pp (15.5 percent) decline.&lt;/p&gt;
&lt;p&gt;The asymmetry between paternal and maternal transmission is explained by the first-generation effects of the reform. For men, the reform increased schooling by 0.32 years, earnings by approximately 1 percent, the probability of white-collar employment by 1.2 percent, cognitive skills by 0.14 standard deviations, noncognitive skills by 0.17 standard deviations, spousal earnings by 1,022 SEK per year, and overall household income by approximately 1 percent. For women, the reform increased education by 0.21 years but did not raise earnings, household income, or white-collar employment, and did not reduce their already low crime rates. Only 13 percent of women in the 1945–55 cohorts were at or below the compulsory schooling threshold, versus 20 percent of men, substantially limiting the reform&amp;rsquo;s bite for women.&lt;/p&gt;
&lt;p&gt;A mediation analysis decomposes the intergenerational transmission through three channels: fathers&amp;rsquo; education accounts for 64.8 percent of the indirect effect, the decline in paternal crime accounts for 18.5 percent, and the increase in household disposable income accounts for 16.7 percent. The direct effect (unexplained by these mediators) accounts for 48 percent of the total effect. The paper also documents that children of treated fathers attended schools with lower peer crime rates and lived in neighborhoods with lower youth crime rates, supporting a neighborhood and peer effects channel alongside human capital and role-model channels.&lt;/p&gt;
&lt;p&gt;Scope conditions: the study covers male children observed to age 29 in Sweden; results apply to a context of near-universal administrative records, a specific postwar schooling reform, and cohorts born 1945–1955 in a Nordic welfare state.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the intergenerational crime reduction caused by the reform?&lt;/p&gt;
&lt;p&gt;A: Sons of fathers exposed to the reform experienced a 0.79 pp decline in conviction rates (p-value &amp;lt; 0.002), corresponding to a 3.4 percent reduction relative to the baseline conviction rate of approximately 24 percent for the child generation by age 29. Multiple convictions fell by 0.43 pp, a 5.8 percent reduction. These magnitudes are similar in percentage terms to the direct crime reduction the reform caused among fathers themselves.&lt;/p&gt;
&lt;p&gt;Q: Does the reform&amp;rsquo;s intergenerational effect on crime differ by the sex of the treated parent?&lt;/p&gt;
&lt;p&gt;A: Yes. The intergenerational effect is driven entirely by paternal exposure to the reform: the effect of maternal exposure is an order of magnitude smaller and insignificant at any conventional significance level. The difference between paternal and maternal effects is statistically significant, with p-values of 0.048 for any conviction and 0.009 for multiple convictions. The paper attributes this asymmetry to the much weaker first-generation effects of the reform on women&amp;rsquo;s earnings, household income, crime rates, and neighborhood sorting.&lt;/p&gt;
&lt;p&gt;Q: Which crime types declined significantly among sons of treated fathers?&lt;/p&gt;
&lt;p&gt;A: Significant declines were found in violent crime (−0.27 pp, Romano-Wolf p-value 0.09), traffic-related crime (−0.45 pp, RW p-value 0.057), fraud (−0.22 pp, RW p-value 0.09), and other offenses (−0.41 pp, RW p-value 0.047), each representing a three-to-six percent reduction relative to the mean incidence of that crime type. Property crime and drug-related crime did not show significant declines.&lt;/p&gt;
&lt;p&gt;Q: What were the direct effects of the reform on the parent generation&amp;rsquo;s human capital?&lt;/p&gt;
&lt;p&gt;A: For men, the reform increased schooling by 0.32 years, earnings by approximately 1 percent, the probability of white-collar employment by 1.2 percent, cognitive skills by 0.14 standard deviations, and noncognitive skills by 0.17 standard deviations, all measured at military enlistment. Spousal earnings increased by 1,022 SEK per year and overall household income rose by approximately 1 percent. For women, education increased by 0.21 years and marriage market matches improved, but earnings, household income, and white-collar employment probability did not increase significantly.&lt;/p&gt;
&lt;p&gt;Q: Why did the reform have stronger first-generation effects on men than on women?&lt;/p&gt;
&lt;p&gt;A: The average share of individuals at or below the compulsory schooling threshold — the margin at which the reform was binding — was 20 percent for men but only 13 percent for women in the 1945–55 cohorts. Because fewer women were constrained by the old compulsory schooling limit, the reform increased their education by less and produced smaller downstream effects on earnings and labor market outcomes.&lt;/p&gt;
&lt;p&gt;Q: What are the three channels through which the reform reduces child crime, and what is the relative contribution of each?&lt;/p&gt;
&lt;p&gt;A: The paper identifies three channels: (1) the human capital channel, whereby increased parental education raises household income and child human capital; (2) the role model channel, whereby reduced paternal crime participation directly reduces son&amp;rsquo;s crime; and (3) the neighborhood and peer effects channel, whereby higher income enables sorting into lower-crime neighborhoods and better schools. The mediation analysis attributes 64.8 percent of the indirect effect to fathers&amp;rsquo; increased education, 18.5 percent to the decline in paternal crime, and 16.7 percent to the increase in household disposable income. The direct effect unexplained by these three mediators accounts for 48 percent of the total effect.&lt;/p&gt;
&lt;p&gt;Q: What is the role model effect, and how strong is it in the parent generation?&lt;/p&gt;
&lt;p&gt;A: The role model channel operates through the strong intergenerational persistence in crime participation: sons are 2.06 times more likely to participate in crime if their fathers have been convicted (Hjalmarsson and Lindquist, 2012). The reform reduced the incidence of any conviction among treated men by 1.5 pp and repeat convictions by 1.5 pp — the latter representing an approximately 8 percent decline from a lower base. For women, the reform produced no reduction in crime, providing no analogous role model improvement through the maternal channel.&lt;/p&gt;
&lt;p&gt;Q: How does neighborhood and school peer quality change for children of treated fathers versus treated mothers?&lt;/p&gt;
&lt;p&gt;A: Sons of fathers exposed to the reform moved to neighborhoods with lower youth crime rates (−0.087 pp) and attended schools with lower peer crime rates (−0.077 pp). In contrast, sons of mothers exposed to the reform experienced higher neighborhood crime rates (p-value 0.06) and higher school peer crime rates (p-value 0.01), the opposite direction. This asymmetry helps explain why only paternal treatment generates significant second-generation crime reductions.&lt;/p&gt;
&lt;p&gt;Q: What happens to other outcomes for children of treated fathers beyond crime?&lt;/p&gt;
&lt;p&gt;A: Sons experienced a 1.2 percentile increase in school GPA (RW p-value 0.05), a 2.3 pp increase in employment (RW p-value 0.04), a matching 2.3 pp decline in unemployment benefit receipt, a reduction in hospitalization of 2.4 days (17 percent, RW p-value 0.02), and a decline in prescribed drugs of 31 doses (2.8 percent, RW p-value 0.09). The decline in prescribed drugs for sons is driven by nervous system drugs and painkillers, pointing to improved mental health. Daughters of treated fathers show a significant reduction in welfare dependency but no other significant improvements.&lt;/p&gt;
&lt;p&gt;Q: How does the paper validate the parallel trends assumption?&lt;/p&gt;
&lt;p&gt;A: Three tests are reported. First, including municipality-specific linear trends leaves the main coefficient unchanged (p-value 0.85 for the trend terms themselves). Second, placebo contrasts using incorrect reform implementation dates produce effects indistinguishable from zero for all tested dates. Third, graphical inspection of regression residuals shows no correlation with municipality-specific trends. Together these provide strong support for the identifying assumption.&lt;/p&gt;
&lt;p&gt;Q: Are the results sensitive to using a linear probability model instead of a nonlinear model?&lt;/p&gt;
&lt;p&gt;A: A Monte Carlo experiment was conducted replicating observed crime rates across municipalities and imposing the estimated average treatment effect. Assuming the true data-generating process is a probit model, the linear probability model biases the estimated average effect upward by only 5 percent — a difference that is statistically indistinguishable from zero in the actual data — validating the OLS approach.&lt;/p&gt;
&lt;p&gt;Q: What is the broader policy implication of the findings?&lt;/p&gt;
&lt;p&gt;A: The results show that well-designed education policies can reduce crime not only among the directly treated generation but also among their children, amplifying the social benefits of reform across generations. The authors interpret this as consistent with the theoretical framework of Becker and Tomes (1979) on intergenerational transmission of human capital, and suggest that education policy evaluations that focus only on the treated generation substantially understate total social returns.&lt;/p&gt;
&lt;p&gt;Intergenerational transmission of education reform effects: the phenomenon whereby an education policy that raises parental human capital produces improvements in children&amp;rsquo;s outcomes — including crime — through multiple channels including resource increases, parental role modeling, and neighborhood sorting, beyond any direct policy exposure of the child generation.&lt;/p&gt;
&lt;p&gt;Comprehensive school reform (Sweden, 1949–1962): a nationally mandated restructuring of compulsory schooling that extended required attendance by one to two years, abolished selection into academic and vocational tracks after 6th grade, and introduced a uniform national curriculum, rolled out staggered across 1,055 Swedish municipalities.&lt;/p&gt;
&lt;p&gt;Human capital channel: the mechanism by which increased parental education raises earnings and household income, enabling greater investments in children&amp;rsquo;s development and exploiting complementarity between parental and child human capital in the skill production function, thereby raising children&amp;rsquo;s opportunity cost of crime.&lt;/p&gt;
&lt;p&gt;Role model channel: the mechanism by which reduced parental crime participation directly reduces children&amp;rsquo;s crime, operating through the transmission of norms and information across generations; identified empirically by the strong intergenerational correlation in convictions (sons with convicted fathers are 2.06 times more likely to be convicted themselves).&lt;/p&gt;
&lt;p&gt;Neighborhood and peer effects channel: the mechanism by which increased parental income from the reform enables sorting into residential neighborhoods and schools with lower youth crime rates, exposing children to peers less involved in illegal activities and thereby reducing their own crime participation.&lt;/p&gt;
&lt;p&gt;Mediation analysis: a decomposition method following Heckman, Pinto, and Savelyev (2013) that quantifies the share of a total treatment effect accounted for by specific intermediate variables (here: fathers&amp;rsquo; education, fathers&amp;rsquo; crime participation, and household disposable income) versus the direct unexplained effect.&lt;/p&gt;
&lt;p&gt;Conviction rate: the proportion of individuals in a given generation and observation window who received at least one criminal conviction in Swedish administrative records; used as the primary outcome measure because it captures offenses that led to a court appearance, excluding minor infractions resolved by direct fine.&lt;/p&gt;</description></item><item><title>The Effect of Provider Diversity on Racial Health Disparities: Evidence from the Military</title><link>https://macropaperwarehouse.com/papers/the-effect-of-provider-diversity-on-racial-health-disparities-evidence-from-the-military/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effect-of-provider-diversity-on-racial-health-disparities-evidence-from-the-military/</guid><description>&lt;p&gt;This paper asks whether racial concordance between patients and medical providers — specifically, whether Black patients are treated by Black physicians — improves use of preventive care and reduces mortality among patients with chronic, manageable diseases. The authors argue that trust and communication deficits along racial lines cause Black patients to underuse low-cost, life-saving preventive care, and that increasing the share of Black providers addresses this deficit.&lt;/p&gt;
&lt;p&gt;The authors use data from the Military Health System (MHS) Data Repository covering fiscal years 2003–2013, encompassing roughly 9.6 million beneficiaries. A distinctive feature of the MHS is that active-duty providers are themselves MHS beneficiaries, so their race is observed in the same eligibility files used for patients — overcoming the typical absence of provider-race data in claims databases. The study focuses on four chronic, deadly but manageable conditions: diabetes, hypertension, hypercholesterolemia, and clinical atherosclerotic cardiovascular disease. Preventive care is measured by medication fill-days for condition-appropriate generic drugs, HEDIS-recommended Comprehensive Diabetes Care compliance, and (for a subset) blood pressure control. Mortality is tracked across the full sample period.&lt;/p&gt;
&lt;p&gt;The identification strategy exploits quasi-random variation in provider racial composition induced by across-base moves. The MHS setting generates abundant moves driven by DoD personnel management needs — not by patient health or preferences. Using a movers-only differences specification (analogous to Finkelstein et al. 2016), the authors compare differential changes in outcomes for Black versus non-Black patients who move to bases with larger versus smaller increases in the share of Black providers. This design includes fixed effects for both sending and receiving bases, controlling flexibly for regional quality differences. The estimand is an intent-to-treat effect among patients living within 10 miles of a base (who use on-base care 66% of the time).&lt;/p&gt;
&lt;p&gt;The findings are consistent across all four disease samples. For diabetes, a move-induced one-standard-deviation increase in the share of Black diabetes providers is associated with a roughly 6 additional metformin fill-days per year (approximately 16% relative to the mean) and a 3 percentage-point increase (roughly 8% relative to the mean) in Comprehensive Diabetes Care compliance for Black relative to non-Black patients. Mortality falls by 0.4 percentage points — a 33% relative decline — for Black relative to non-Black diabetes patients following such a move.&lt;/p&gt;
&lt;p&gt;Pooling across all four chronic-disease samples, a one-standard-deviation move-induced increase in the Black provider share is associated with approximately 3 additional fill-days of relevant preventive medication and a roughly 0.2 percentage-point reduction in mortality — approximately 15% relative to the mean mortality rate — for Black relative to non-Black patients.&lt;/p&gt;
&lt;p&gt;A decomposition analysis combining the paper&amp;rsquo;s estimates with medical-literature parameters on the mortality effects of preventive medications finds that between 55% and 69% of the concordance mortality effect across the four disease samples can be attributed to improved medication adherence alone, with the remainder attributed to other aspects of the provider-patient relationship (e.g., lifestyle effects, other preventive care).&lt;/p&gt;
&lt;p&gt;Scope conditions: results are local to MHS movers, who are on average slightly younger and healthier than non-movers, potentially understating concordance benefits for the full population. The MHS covers over 3% of all Black U.S. residents, but beneficiaries may differ from the general population. The paper measures Black patient / Black provider concordance specifically; it does not establish a symmetric concordance effect for non-Black patients. The concordance effect estimated is relative — it captures how much Black patients benefit more than non-Black patients from moving to a higher Black-provider-share base. A system-wide spillover mechanism (non-Black providers improving care for Black patients when working alongside more Black providers) cannot be ruled out and would also be consistent with the core concordance motivation.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question and why is the MHS an advantageous setting?
A: The paper asks whether racial concordance between providers and patients causes Black patients to use more preventive care and achieve better health outcomes, focusing on the trust and communication channel. The MHS is advantageous because active-duty providers are themselves MHS beneficiaries, making their race observable — a feature absent in most claims databases. Across-base moves are driven by DoD staffing needs rather than patient health or preferences, providing quasi-random variation in provider racial composition. The system offers complete claims data covering both on- and off-base care, allowing full mortality tracking.&lt;/p&gt;
&lt;p&gt;Q: How does the empirical strategy address selection concerns that plague prior concordance studies?
A: Prior studies face selection problems from Black patients choosing different doctors than white patients and from residential segregation concentrating Black patients and Black physicians in regions with distinct care quality. The movers-based differences specification directly addresses both problems: it uses only patients who move across bases, comparing how the same individual&amp;rsquo;s outcomes change relative to non-Black patients experiencing the same move, as a function of the move-induced change in the Black provider share. Inclusion of fixed effects for both sending and receiving bases accounts flexibly for regional quality differences. Balance tests on observable patient characteristics show no differential sorting of Black versus non-Black patients toward high-Black-provider-share bases.&lt;/p&gt;
&lt;p&gt;Q: What specific preventive care and outcome measures are used for each disease?
A: For diabetes, the primary measures are annual metformin fill-days and Comprehensive Diabetes Care (CDC) compliance — defined as receiving HbA1c testing, a retinal eye exam, and medical attention for nephropathy in the focal year — plus blood pressure control (available only from 2009 onward for on-base patients). For hypertension, the measures are annual fill-days of WHO-recommended antihypertensives (thiazides, ACEs/ARBs, or long-acting dihydropyridine CCBs) and blood pressure control. For hypercholesterolemia, the measure is fill-days of antilipemic agents, bile acid sequestrants, and statins. For atherosclerotic cardiovascular disease, the HEDIS statin therapy receipt indicator is used. Mortality is tracked across all four samples.&lt;/p&gt;
&lt;p&gt;Q: What are the main quantitative results for the diabetes sample?
A: A move-induced one-standard-deviation increase in the share of Black diabetes providers is associated with approximately 6 additional metformin fill-days annually for Black relative to non-Black patients (roughly 16% relative to the mean). Compliance with Comprehensive Diabetes Care increases by 3 percentage points for Black relative to non-Black patients (roughly 8% relative to the mean). Mortality falls by 0.4 percentage points for Black relative to non-Black patients — a 33% relative decline — in connection with the same one-standard-deviation increase in Black provider share.&lt;/p&gt;
&lt;p&gt;Q: What are the pooled results across all four chronic-disease samples?
A: Pooling across diabetes, hypertension, hypercholesterolemia, and atherosclerotic cardiovascular disease, a one-standard-deviation move-induced increase in the Black provider share is associated with approximately 3 additional preventive medication fill-days per year for Black relative to non-Black patients. The pooled mortality effect is a 0.2 percentage-point reduction — roughly 15% relative to the mean mortality rate — for Black relative to non-Black patients.&lt;/p&gt;
&lt;p&gt;Q: How much of the concordance mortality effect operates through medication adherence?
A: The decomposition combines the paper&amp;rsquo;s estimated concordance effects on medication fill-days with medical-literature estimates of the mortality impact of each additional fill-day. For the diabetes sample, increased metformin adherence (4.2 additional fill-days) explains approximately 58.8% of the 0.4 percentage-point concordance mortality effect, with the residual 41.2% attributed to other channels such as lifestyle changes or other preventive care. Across all four disease samples, the medication fill-day channel explains between 55% and 69% of the respective concordance mortality effects.&lt;/p&gt;
&lt;p&gt;Q: What specification checks do the authors conduct to validate causal identification?
A: The authors conduct five main checks. First, balance regressions show that move-induced changes in Black provider share are not differentially related to baseline patient characteristics for Black versus non-Black patients. Second, regressions of the probability of moving on initial Black provider share and its interaction with patient race yield a near-zero concordance coefficient (0.008, SE 0.023), indicating no differential sorting. Third, regressions of post-move on-base care share on the concordance interaction term yield a near-zero coefficient (0.002, SE 0.003), indicating no differential race-specific selection into on-base care. Fourth, a distance falsification test shows that concordance coefficients are near zero and statistically insignificant for patients living more than 10 miles from the base. Fifth, event-study dynamics show no pre-move divergence in preventive care adherence between Black and non-Black patients, with a positive divergence emerging only after the move to a higher Black-provider-share base.&lt;/p&gt;
&lt;p&gt;Q: How does the paper separate a concordance effect from a pure Black-physician-quality effect?
A: The paper estimates a &amp;ldquo;first stage&amp;rdquo; specification on the subsample receiving on-base care (where provider race is observed), regressing the change in the probability of visiting a Black provider on the move-induced change in Black provider density. The results show an approximately one-to-one relationship between higher Black provider availability and increased visits to Black providers for all patients, with only a modest differential by patient race. This confirms that non-Black patients also see more Black providers when Black provider density rises, allowing the interaction specification to isolate concordance from a pure physician-quality effect.&lt;/p&gt;
&lt;p&gt;Q: How do the authors assess the potential role of spillover effects?
A: The authors acknowledge they cannot rule out that some of the estimated concordance effect arises through system-wide spillovers — for instance, non-Black providers on bases with more Black colleagues may improve their care for Black patients through peer learning or information transmission. They note that even if such a spillover mechanism operates, it is still consistent with the paper&amp;rsquo;s core concordance motivation, because provider-knowledge deficiencies about treating Black patients are among the theorized channels of racial discordance.&lt;/p&gt;
&lt;p&gt;Q: What do the results imply for the overall racial mortality gap?
A: Among MHS beneficiaries aged 20–65, Black beneficiaries are roughly 38% more likely to have diabetes and die over the sample period than non-Black beneficiaries; this gap appears driven primarily by higher diabetes prevalence rather than a within-diabetes mortality gap. Applying the diabetes concordance mortality estimate (a 0.4 percentage-point reduction), the authors calculate that a one-standard-deviation increase in the Black provider share would reduce the overall diabetes mortality gap from 38% to approximately 21% — a substantial narrowing driven by the concordance effect operating through conditional-on-prevalence outcomes.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: The results imply that investments in increasing physician workforce diversity could meaningfully reduce racial mortality disparities in the United States, particularly for chronic diseases manageable through preventive medication. The paper notes the results are relevant to affirmative action policies in medical school admissions, specifically the pending Supreme Court cases Students for Fair Admissions v. University of North Carolina and Students for Fair Admissions v. Harvard at the time of writing. The MHS population covered in the study includes over 3% of all Black U.S. residents, so the policy stakes extend substantially beyond the military context.&lt;/p&gt;
&lt;p&gt;Q: What are the limitations of the study regarding generalizability?
A: Movers in the chronic-disease samples are on average about four years younger and 0.2 percentage points less likely to die than non-movers, suggesting the local average treatment effect for movers may understate concordance benefits for the full population. The MHS population may be healthier overall than the general population, though conditioning on chronic-disease patients mitigates this concern. The paper covers only Black-patient/Black-provider concordance; concordance effects for other racial and ethnic groups are not estimated. The estimate of the concordance coefficient technically captures how much the Black patient / Black provider concordance effect exceeds the non-Black patient / non-Black provider concordance effect, meaning the absolute magnitude of Black concordance benefits is understated if non-Black concordance effects are also positive.&lt;/p&gt;
&lt;p&gt;Racial concordance: In this paper&amp;rsquo;s usage, the match between the race of a patient and their treating physician — specifically Black patient / Black provider pairing — theorized to improve care through trust, communication, and reduced provider knowledge deficiencies about Black patients.&lt;/p&gt;
&lt;p&gt;Provider Black share: The fraction of outpatient office visits for a given chronic condition at a given military base that are attended by Black active-duty providers, used as the base-level treatment variable; varies across bases from zero to approximately 20 percentage points in the pooled sample.&lt;/p&gt;
&lt;p&gt;Movers-based differences specification: An identification strategy that restricts to patients who relocate across military bases exactly once during the sample period and estimates the differential change in outcomes for Black versus non-Black patients as a function of the move-induced change in the base&amp;rsquo;s Black provider share, including fixed effects for both the sending and receiving base.&lt;/p&gt;
&lt;p&gt;Intent-to-treat (ITT) effect: The concordance estimate as applied to all patients living within 10 miles of a base — regardless of whether they actually received on-base care — to avoid selection bias from differential race-specific decisions to seek care on versus off base.&lt;/p&gt;
&lt;p&gt;Comprehensive Diabetes Care (CDC): A HEDIS composite measure requiring receipt of all three of the following in the focal year: HbA1c testing, a retinal eye exam, and medical attention for nephropathy (via microalbumin exam, ACE/ARB therapy, or nephropathy treatment).&lt;/p&gt;
&lt;p&gt;Medication fill-days: Annual days of supply dispensed for condition-appropriate generic medications (metformin for diabetes; thiazides/ACEs/ARBs/CCBs for hypertension; antilipemic agents, bile acid sequestrants, and statins for hypercholesterolemia; statins for atherosclerotic cardiovascular disease), used as the primary preventive care adherence measure.&lt;/p&gt;
&lt;p&gt;Decomposition of concordance mortality effect: A calculation that uses the paper&amp;rsquo;s estimated concordance effect on medication fill-days, combined with medical-literature estimates of the mortality impact per fill-day, to determine what share of the total concordance mortality effect passes through medication adherence versus other channels (lifestyle, other preventive care).&lt;/p&gt;</description></item><item><title>The Effects of Gender Integration on Men</title><link>https://macropaperwarehouse.com/papers/the-effects-of-gender-integration-on-men/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effects-of-gender-integration-on-men/</guid><description>&lt;p&gt;Greenberg, Wasserman, and Weber (2024/2026) ask whether men negatively respond—in terms of job performance, behavior, and workplace perceptions—when women first enter an exclusively male occupation. They exploit the staggered 2017-onward integration of women into U.S. Army infantry and armor combat companies following the 2016 rescission of the Ground Combat Exclusion Policy. The setting offers unusually clean causal identification: integration timing within Brigade Combat Teams was neither systematic nor data-driven, the Army&amp;rsquo;s rigid pay scales meant integration posed no displacement or wage threat to incumbent men, and roughly 391 companies are observed over 2012–2020. The empirical strategy is a staggered difference-in-differences design with company fixed effects, BCT-by-year-of-arrival fixed effects, and month-of-year fixed effects, applied to an individual-level sample of newly arrived male soldiers. Outcomes come from monthly administrative personnel records (retention, misconduct separations, demotions, criminal investigations, drug tests, medical profiles, physical fitness scores) and the Defense Organizational Climate Survey (DEOCS), a congressionally mandated annual survey with response rates above 50% covering organizational effectiveness, equal opportunity, and sexual assault prevention and response. The main finding is that integrating women into previously all-male combat companies does not negatively affect men&amp;rsquo;s performance or behavioral outcomes. Estimates are precise enough to rule out small detrimental effects: two years post-integration, the authors can rule out a 3% increase in attrition, a 5% increase in demotions, and a 4% increase in criminal investigations relative to their respective means. One behavioral outcome shows a statistically significant improvement: integration reduces separations for misconduct by 1.3 percentage points (16% of the mean). Drug test positivity also declines. The sole potential negative administrative finding is a 1.8-point decline in physical fitness scores (0.7% of the mean, roughly 5% of a standard deviation), but this does not affect pass rates and becomes statistically insignificant when scores are imputed using observable covariates. An aggregate Performance and Behavior Index rules out reductions of 0.8% of a standard deviation; the No Adverse Outcomes measure rules out a 1.2 percentage point increase (3% of the mean). Despite these null-to-positive performance effects, survey data reveal that integration causes a 5% of a standard deviation decline in men&amp;rsquo;s overall perceptions of workplace quality. This perception decline is concentrated in companies that received a female officer shortly after integration. Among companies integrated only with female enlisted soldiers (no female officer), men&amp;rsquo;s workplace attitudes actually improve by 14.7% of a standard deviation. Two mechanisms are examined: increased male awareness of pre-existing workplace problems (supported by higher reported observations of bullying, hazing, and unwanted comments, especially among male officers in female-officer-integrated companies), and negative reactions to women in positions of authority (supported by broader declines in organizational effectiveness perceptions not confined to equal-opportunity items). Crucially, the perception decline does not translate into retaliatory behavior or performance deterioration; companies integrated with a female officer show some performance gains, and female enlisted soldiers in those companies report fewer workplace problems. Scope conditions: findings apply to a high-stakes, traditionally male-dominated, hierarchical occupational setting during 2017–2020, a period when U.S. deployment missions were primarily advise-and-assist rather than direct combat. Integration increased female representation by approximately 4.7 percentage points on average.&lt;/p&gt;
&lt;p&gt;Q: What was the policy change studied and why does it offer causal leverage?
A: In December 2015, Secretary of Defense Ashton Carter announced that all U.S. military occupations, including infantry and armor combat roles, would open to women starting in 2016. Women did not begin arriving at operational companies until 2017 due to training timelines. Within BCTs, the selection of which companies to integrate was neither systematic nor data-driven, and baseline characteristics of integrated and non-integrated companies are similar after conditioning on BCT and company-type fixed effects, supporting a parallel trends assumption.&lt;/p&gt;
&lt;p&gt;Q: What are the main administrative performance findings?
A: Integration has a positive but statistically insignificant effect on retention, and reduces misconduct separations by 1.3 percentage points (significant at the 5% level), representing a 16% reduction relative to the mean. Demotions, criminal investigations (including sex-related and domestic violence), and medical profiles show no significant negative effects, with precision sufficient to rule out 5% increases in demotions and 4% increases in criminal investigations. Physical fitness scores decline by 1.8 points (0.7% of mean, approximately 5% of a standard deviation), but pass rates are unaffected and the estimate becomes insignificant when scores are imputed with observable covariates.&lt;/p&gt;
&lt;p&gt;Q: What does the aggregate performance index show?
A: The Performance and Behavior Index—an equally weighted z-score average of retention, misconduct separations, demotions, criminal investigations, medical profiles, promotions to Sergeant, and physical fitness outcomes—shows a positive but insignificant effect of integration, ruling out reductions of 0.8% of a standard deviation. The No Adverse Outcomes measure rules out a 1.2 percentage point increase (3% of the mean incidence of adverse outcomes).&lt;/p&gt;
&lt;p&gt;Q: How do men&amp;rsquo;s workplace perceptions change after integration?
A: The overall workplace quality index constructed from all DEOCS Likert-scale items declines by 5% of a standard deviation following integration, spanning perceptions of organizational effectiveness, workplace inclusivity, and sexual assault prevention and response. This average effect masks critical heterogeneity by the rank composition of integrating women.&lt;/p&gt;
&lt;p&gt;Q: What is the key heterogeneity in survey responses?
A: The decline in men&amp;rsquo;s perceptions is entirely driven by companies that received a female officer shortly after integration. In companies integrated only with female enlisted soldiers (17% of integrating companies did not receive a female officer within a month), men&amp;rsquo;s perceptions improve by 14.7% of a standard deviation. Male officers show a larger negative shift than male enlisted soldiers in officer-integrated companies, and this difference is statistically significant.&lt;/p&gt;
&lt;p&gt;Q: What mechanisms explain the negative perception response to female officers?
A: Two mechanisms are investigated. First, increased awareness: male soldiers—especially male officers—report observing more bullying, hazing, and unwanted comments after a female officer is integrated but not after integration with only female enlisted, and the decline in perceptions of sexual assault prevention and response is significantly larger among male officers than enlisted men, consistent with shared leadership roles amplifying awareness of workplace problems. Second, negative reactions to female authority: declines in perceptions are more pronounced on organizational effectiveness questions than on equal-opportunity items and extend to issues unrelated to women, suggesting broader dissatisfaction with female leadership alongside heightened awareness.&lt;/p&gt;
&lt;p&gt;Q: Is the decline in perceptions related to actual differences in female officer qualifications or preferential treatment?
A: No. Female and male officers have similar baseline characteristics including educational background and experience. Companies integrated with female officers perform at least as well as non-integrated companies or those integrated only with enlisted women on administrative metrics. There is no evidence that male officers waited longer for leadership assignments relative to female colleagues, ruling out perceived preferential treatment as a driver.&lt;/p&gt;
&lt;p&gt;Q: Do men&amp;rsquo;s negative perceptions of female officers translate into retaliatory behavior toward women?
A: No. Administrative misconduct metrics show some improvements in male behavior when a female officer is present. Female enlisted soldiers in female-officer-integrated companies report fewer workplace problems on the climate survey than female enlisted soldiers in companies integrated without a female officer, indicating that the presence of a female officer generates benefits for female enlisted soldiers rather than backlash against them.&lt;/p&gt;
&lt;p&gt;Q: Does heterogeneity by integration intensity or women&amp;rsquo;s rank affect administrative outcomes for men?
A: Integration intensity (number of women initially integrated) and rank composition (female officers vs. only female enlisted) do not produce negative administrative outcomes in any subgroup. The aggregate Performance and Behavior Index shows a positive effect when a female officer is included. Effects also do not vary with male soldiers&amp;rsquo; rank (enlisted vs. officer) or their tenure in the company.&lt;/p&gt;
&lt;p&gt;Q: What happens in units that deploy to combat zones?
A: Approximately one in five integrated companies deployed to a combat zone within two years of integration. Integration does not negatively affect retention, behavior, or performance of men in deploying units. Declines in workplace perceptions are larger for deploying units and are most pronounced when integration occurs shortly after return from deployment, consistent with deployment strengthening in-group identity among male soldiers rather than women performing poorly during combat-zone service.&lt;/p&gt;
&lt;p&gt;Q: What do the findings imply for theories of identity economics and the pollution theory of discrimination?
A: The null-to-positive behavioral and performance responses to women&amp;rsquo;s entry contradict the predictions of Akerlof and Kranton&amp;rsquo;s (2000) identity economics model and Goldin&amp;rsquo;s (2014) pollution theory of discrimination, which predict retaliatory or otherwise unproductive behaviors when women enter a male-dominated occupation. The paper shows that, to the extent identity concerns shape male responses, these are confined to subjective perceptions and do not manifest in diminished performance, retention, or conduct.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications for employers considering gender integration?
A: The paper provides evidence against the argument that men will become less productive when women enter previously male-only occupations, a justification sometimes offered for excluding women from such jobs. The finding that performance and behavior are unaffected—and misconduct actually declines—allows policymakers and employers to weigh these results against concerns about operational or productivity costs of integration. The perception gap between men&amp;rsquo;s attitudes and actual outcomes points to a need for targeted leadership and organizational interventions, particularly around the introduction of female leaders.&lt;/p&gt;
&lt;p&gt;Ground Combat Exclusion Policy (GCEP): The U.S. military policy, rescinded in 2013 and fully eliminated by Secretary of Defense Carter in 2016, that precluded women from serving in infantry and armor positions; the policy whose removal is the source of the integration shock studied. | Staggered difference-in-differences: The empirical strategy exploiting the sequential, non-systematic integration of women into combat companies across years 2017–2023, using never-yet-treated companies as a comparison group with company fixed effects and BCT-by-year-of-arrival fixed effects. | Performance and Behavior Index: An equally weighted average of z-scored administrative outcomes (retention, no misconduct separations, no demotions, no criminal investigations, no medical profiles, promotion to Sergeant, physical fitness pass/fail and score), constructed for enlisted soldiers, oriented so higher values indicate better outcomes. | Leaders First policy: An Army requirement that a female officer be assigned to a combat company before or alongside female junior enlisted soldiers to ensure female leadership presence at integration; adherence was not universal, with 17% of integrating companies not following it within one month. | Defense Organizational Climate Survey (DEOCS): A congressionally mandated, annually administered, anonymous survey of military unit members covering organizational effectiveness, equal opportunity, and sexual assault prevention and response; the source of workplace perception outcomes. | Pollution theory of discrimination: Goldin&amp;rsquo;s (2014) theory that men may seek to exclude women from occupations because women&amp;rsquo;s presence is perceived to diminish the occupation&amp;rsquo;s prestige or status, potentially leading to retaliatory or unproductive behaviors among incumbent male workers. | Perception-performance wedge: The paper&amp;rsquo;s central finding that men&amp;rsquo;s subjective workplace quality perceptions decline with integration—especially when a female officer is present—even as objective administrative performance and behavior metrics show null to positive effects, a divergence between attitudes and measurable outcomes.&lt;/p&gt;</description></item><item><title>The Effects of Mandatory Profit-Sharing on Workers and Firms</title><link>https://macropaperwarehouse.com/papers/the-effects-of-mandatory-profit-sharing-on-workers-and-firms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effects-of-mandatory-profit-sharing-on-workers-and-firms/</guid><description>&lt;p&gt;This paper studies the causal effects of mandatory profit-sharing on workers and firms using a quasi-experimental design arising from a 1990 French reform that lowered the eligibility threshold for mandatory profit-sharing from 100 to 50 employees. The institutional setting is the French RSP (Réserve Spéciale de Participation), a profit-sharing scheme in place since 1967 that requires firms above the threshold to distribute a fraction of their excess profits — defined as net income above 5% of book equity — to employees according to a formula scaled by the firm&amp;rsquo;s labor share. For the median firm, this amounts to roughly 10.5% of pre-tax income transferred to workers.&lt;/p&gt;
&lt;p&gt;The authors employ two primary empirical strategies. First, a bunching analysis exploits the pre-reform distribution of firm employment around the 100-employee threshold as a revealed-preference test of whether firms perceive profit-sharing as a net cost. Second, a difference-in-differences design compares treated firms (55–85 employees in 1989–1990, who become newly subject to the regulation after 1991) against two control groups: small firms (35–45 employees, likely never subject) and large firms (120–300 employees, already subject). Data come from the universe of French corporate tax files (FICAS) and a linked employer-employee panel (DADS) covering approximately 4% of private-sector workers, spanning 1985–1997.&lt;/p&gt;
&lt;p&gt;The bunching analysis documents a 22.3% excess density in the 95–99 employee bin before the reform, which disappears after 1991. Three tests — comparing wage bills per employee across the threshold, cross-checking with DADS employment records, and examining profitability patterns — collectively support the conclusion that bunching reflects genuine employment reductions rather than under-reporting. The implied employment loss is approximately 1.67% of total employment among affected firms.&lt;/p&gt;
&lt;p&gt;The difference-in-differences results yield the following firm-level findings: (a) the total compensation share (wages plus profit-sharing divided by value added) rises by 1.8 percentage points for firms with positive excess profits; (b) 77% of this increase comes at the expense of firm owners — the profit share falls by 1.37 percentage points; (c) the remainder is borne by the government through a reduction in the corporate income tax share; (d) the wage share (base wages only) is unaffected, indicating that owners do not reduce wages to offset the cost of profit-sharing; (e) investment and total factor productivity show no statistically significant change — effects on productivity are bounded below ±1% for several TFP measures; and (f) the capital-labor ratio shows a small, mostly insignificant negative effect, consistent with a model-implied increase in the cost of capital of only 0.43 percentage points.&lt;/p&gt;
&lt;p&gt;Worker-level analysis using the linked employer-employee data confirms that average total compensation rises by approximately 3.5% for workers in treated firms, with no decline in base wages. Critically, this average conceals distributional heterogeneity across the skill spectrum. For low- and medium-skill workers (blue-collar workers, clerks, supervisors, skilled technicians), total compensation rises while base wages are unchanged — consistent with wage rigidity binding for these groups. For high-skill workers (managers, engineers, executives), base wages fall by enough to leave total compensation unchanged, consistent with more flexible wages at the upper end of the skill distribution. This pattern implies that mandatory profit-sharing is a progressive policy within firms, redistributing excess profits predominantly to lower-skill workers.&lt;/p&gt;
&lt;p&gt;The paper concludes that France&amp;rsquo;s mandatory profit-sharing scheme, as implemented, functions as a non-distortive redistributive tool: it transfers excess profits from shareholders to lower-skill workers without generating measurable productivity losses or large investment distortions. The fiscal cost is non-trivial: each dollar transferred to workers costs approximately 20 cents in foregone corporate income tax. The scheme also has an inherent inequality in its redistribution since it exclusively benefits workers in profitable firms, and firms&amp;rsquo; excess profits are highly persistent.&lt;/p&gt;
&lt;p&gt;Q: What is the French RSP and how does the formula work?
A: The RSP (Réserve Spéciale de Participation) is a mandatory profit-sharing fund established by executive order in 1967. The formula is RSP = 0.5 × (wage bill / value added) × max(net income − 5% × book equity, 0). The 5% deduction represents lawmakers&amp;rsquo; view of fair compensation to shareholders; any excess is split between shareholders and workers, with the split scaled by the firm&amp;rsquo;s labor share. For the median firm in the sample — ROE of 12%, labor share of 0.52, corporate tax rate of 37% — the formula yields roughly 9.5% of pre-tax income, and in post-1991 data the realized average is 10.5% of pre-tax income for firms with positive excess profits.&lt;/p&gt;
&lt;p&gt;Q: Why can&amp;rsquo;t a standard regression discontinuity be used at the 100-employee threshold?
A: Because firms strategically control their position relative to the threshold — the bunching analysis itself demonstrates this. When firms sort non-randomly around the cutoff, the local randomization assumption underlying RD is violated. The authors instead use a difference-in-differences design exploiting the time variation introduced by the 1990 reform.&lt;/p&gt;
&lt;p&gt;Q: How large is the pre-reform bunching and what does it imply?
A: The distribution of employment shows 22.3% excess density in the 95–99 employee bin relative to the post-reform counterfactual distribution. Interpreting this as real employment reduction (supported by three empirical tests), the implied employment loss is approximately 1.67% of total employment among firms in the 85–120 employee range. Dynamic bunching analysis shows this is persistent rather than temporary — the 100-employee threshold significantly constrained three-year employment growth for firms in the 85–99 range in the pre-reform period.&lt;/p&gt;
&lt;p&gt;Q: How do the authors establish that bunching is real rather than under-reporting of employment?
A: Three tests are conducted. First, wage bills per employee show no discontinuity around the 100-employee threshold in either period, ruling out systematic under-reporting of headcount while truthfully reporting wages. Second, employment from DADS payroll records — harder to manipulate — shows only a statistically insignificant gap of roughly 0.5 employees relative to tax-file employment just below the threshold, far too small to shift firms across the 100-employee bin. Third, profitability and value added per employee are significantly higher just below the threshold, consistent with more profitable firms having stronger incentives to bunch through genuine employment reductions.&lt;/p&gt;
&lt;p&gt;Q: What is the main identification strategy for the firm-level analysis?
A: A difference-in-differences design where treated firms have 55–85 employees in both 1989 and 1990 (newly subject to the mandate after 1991), compared to small control firms with 35–45 employees (likely never subject) and large control firms with 120–300 employees (likely always subject). Specifications include firm fixed effects and county-by-year and industry-by-year fixed effects. Parallel pre-trends are confirmed graphically and in event-study regressions. The design is intent-to-treat: by 1997, 26.7% of treated firms had shrunk below 50 employees and did not actually pay profit-sharing. LATE estimates are obtained via 2SLS.&lt;/p&gt;
&lt;p&gt;Q: What are the main firm-level findings on compensation and profit shares?
A: For treated firms with positive excess profits, the total compensation share rises by 1.8 percentage points. The wage share (base wages only, excluding profit-sharing) is precisely estimated at zero — owners do not reduce wages. The profit share falls by 1.37 percentage points, accounting for 77% of the increase in total compensation. The remaining approximately 23% is borne by the tax authority through a reduction in the corporate income tax share, since profit-sharing reduces the corporate income tax base. These findings are robust to balanced vs. unbalanced samples and to alternative control group definitions.&lt;/p&gt;
&lt;p&gt;Q: Does mandatory profit-sharing raise or lower firm productivity?
A: Across five different TFP estimators (Olley-Pakes, Olley-Pakes with Ackerberg-Caves-Frazer correction, Wooldridge, Levinsohn-Petrin, and Ackerberg-Caves-Frazer), the effect of mandatory profit-sharing on productivity is a precisely estimated zero. For several measures, effects larger than ±1% in magnitude can be rejected. Softer measures of effort — sick leave rates and the probability of working extra hours — also show no significant change. This null finding contrasts with the literature on voluntary profit-sharing adoption, which typically finds 3–5% productivity gains, likely reflecting selection bias in that literature.&lt;/p&gt;
&lt;p&gt;Q: Does mandatory profit-sharing distort investment?
A: The effect on investment is small and mostly statistically insignificant. The theoretical model shows why: the profit-sharing formula is based on excess profits (net income minus 5% of book equity), not total profits. When the firm&amp;rsquo;s actual cost of equity approximately equals the regulatory 5% benchmark, the distortion to the cost of capital is zero. The calibrated distortion to the user cost of capital is only 0.43 percentage points — approximately 1.9% of the standard user cost — implying an investment ratio reduction of about 0.84 percentage points using estimated elasticities from Chodorow-Reich et al. (2024). Empirically, capital-labor ratios show a small, largely insignificant negative effect.&lt;/p&gt;
&lt;p&gt;Q: How does profit-sharing incidence differ across the skill distribution?
A: The worker-level DADS analysis reveals that the average 3.5% increase in total compensation masks sharp heterogeneity. For low- and medium-skill workers (blue-collar workers, clerks, supervisors, skilled technicians), total compensation rises while base wages are unchanged. For high-skill workers (managers, engineers, executives), base wages decline sufficiently to leave their total compensation unchanged. The authors interpret this pattern as consistent with wage rigidity being more binding for lower-skill workers — due to the federal minimum wage and collective agreements — than for managers whose pay is more flexibly set.&lt;/p&gt;
&lt;p&gt;Q: Why does profit-sharing not affect base wages for low-skill workers?
A: Two candidate explanations are considered. The risk channel — that profit-sharing is risky and thus less valuable to risk-averse workers, who demand wage compensation — is rejected empirically because profit-sharing only marginally increases the variability of workers&amp;rsquo; total earnings. The wage rigidity channel is supported: France&amp;rsquo;s binding federal minimum wage and widespread collective agreements constrain downward adjustment in base wages for lower-skill workers, so firms cannot pass through profit-sharing costs as lower wages for this group.&lt;/p&gt;
&lt;p&gt;Q: What is the fiscal cost of the profit-sharing scheme?
A: Each dollar transferred to workers through mandatory profit-sharing costs approximately 20 cents in reduced corporate income tax receipts, since profit-sharing payments are deductible from taxable income. The paper notes this is a partial fiscal evaluation; a full assessment would also require analyzing personal income tax implications, which are left for future work.&lt;/p&gt;
&lt;p&gt;Q: How does this scheme compare to a corporate income tax as a redistributive tool?
A: Both instruments reduce firm profits and can benefit workers, but differ in three key respects. First, the tax base differs: profit-sharing targets excess profits above 5% of book equity whereas the corporate income tax applies to all corporate earnings, generating different distortions to investment. Second, profit-sharing goes directly to workers in the same firm, whereas corporate tax revenues are redistributed through general government spending — making the incidence more direct and more closely monitored by workers. Third, workers have stronger incentives to monitor firm compliance with profit-sharing (each euro of diverted excess profit reduces workers&amp;rsquo; collective income by roughly 10–15 cents) than with corporate taxes.&lt;/p&gt;
&lt;p&gt;Q: How does this paper compare to findings on mandatory profit-sharing in Peru?
A: Tolentino (2022) studies a mandatory profit-sharing scheme in Peru exploiting a 20-employee eligibility threshold and finds larger distortions — reductions in both investment and productivity. The authors attribute this difference to two features: the Peruvian scheme applies to the entirety of post-tax profits rather than excess profits above an equity deduction, creating a broader and more distortionary base; and there is pre-existing bunching at the Peruvian threshold even before the scheme was introduced, suggesting confounding pre-existing regulations.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions on the external validity of the findings?
A: The findings apply specifically to mandatory profit-sharing under the French RSP formula — which exempts a 5% equity return from the profit-sharing base, limiting distortions — during 1985–1997, for firms in the 55–300 employee range. The null productivity effect may not generalize to voluntary schemes, where selection on anticipated gains likely produces positive correlations. The redistributive finding (benefiting lower-skill workers) is specific to a context with binding minimum wages and collective agreements that constrain wage adjustment for that group. The fiscal cost calculation also excludes personal income tax effects.&lt;/p&gt;
&lt;p&gt;Excess profits: Defined in the paper as net income minus 5% of book equity — the amount above what lawmakers considered fair compensation to shareholders. Only excess profits (not total profits) are subject to the mandatory profit-sharing formula.&lt;/p&gt;
&lt;p&gt;RSP formula (Réserve Spéciale de Participation): The statutory formula RSP = 0.5 × (wage bill / value added) × max(net income − 5% × book equity, 0), scaled by the firm&amp;rsquo;s labor share to reflect labor&amp;rsquo;s contribution to production. Unchanged since 1967.&lt;/p&gt;
&lt;p&gt;Total compensation share: The ratio of (wage bill plus profit-sharing) to value added — the paper&amp;rsquo;s primary measure of workers&amp;rsquo; overall claim on firm output, as distinct from the wage share (wage bill alone divided by value added).&lt;/p&gt;
&lt;p&gt;Wage incidence parameter (λ): The fraction of profit-sharing that firms pass through to workers as lower base wages. λ = 1 means full incidence (workers&amp;rsquo; total compensation unchanged); λ = 0 means no incidence (workers fully benefit). The paper&amp;rsquo;s empirical findings are consistent with λ ≈ 0 for low-skill workers and λ ≈ 1 for high-skill workers.&lt;/p&gt;
&lt;p&gt;Bunching: The empirical phenomenon whereby firms cluster employment just below the 100-employee regulatory threshold to avoid mandatory profit-sharing. The paper uses the pre- vs. post-reform shift in the employment distribution as a revealed-preference test of whether firms perceive the scheme as a net cost.&lt;/p&gt;
&lt;p&gt;Intent-to-treat (ITT) design: The empirical design comparing firms that were in the newly eligible size range (55–85 employees) just before the 1990 reform against firms that were either always or never eligible, regardless of whether treated firms actually ended up paying profit-sharing post-reform. LATE estimates are obtained via 2SLS to recover effects on actual compliers.&lt;/p&gt;
&lt;p&gt;Distortion to user cost of capital: The additional cost of capital induced by profit-sharing, equal to ϕ × γ(1−λ) / [1 − γ(1−τ)] × (re − ρ), where ρ = 5% is the regulatory equity benchmark. When the firm&amp;rsquo;s actual cost of equity equals the 5% benchmark, this distortion is zero — a feature that distinguishes the French scheme from a standard corporate income tax.&lt;/p&gt;</description></item><item><title>The Effects of Medical Debt Relief: Evidence from Two Randomized Experiments</title><link>https://macropaperwarehouse.com/papers/the-effects-of-medical-debt-relief-evidence-from-two-randomized-experiments/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effects-of-medical-debt-relief-evidence-from-two-randomized-experiments/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks whether relieving downstream medical debt — debt that has been sold to third-party debt collectors — causes improvements in financial outcomes, mental and physical health, and healthcare utilization for recipients. The question is motivated by a large correlational literature documenting strong associations between medical debt and adverse outcomes, and by the rapid expansion of government and private debt relief programs that, as of mid-2024, had committed or planned over $14.6 billion in relief.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Design&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors partnered with RIP Medical Debt (a non-profit that purchases and forgives medical debt for government and private donors) to conduct two randomized controlled trials between March 2018 and October 2020. In total the experiments relieved medical debt with a face value of $169 million for 83,401 people.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Hospital debt experiment&lt;/strong&gt;: RIP purchased a random subset of debt from a large for-profit hospital system at the juncture when the hospital would normally sell accounts to a debt collector (approximately one year after the medical service). The purchase price was 5.5 cents per dollar of face value. The treatment group consisted of 14,377 people who received $19 million in face-value relief (average of $1,321 per person). The 61,496-person control group had their debt pursued by the collector under normal protocol.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Collector debt experiment&lt;/strong&gt;: RIP purchased a random subset of older debt already under collection on the secondary market for several years, at a price of less than one cent per dollar. The treatment group consisted of 69,024 people who received $150 million in face-value relief (average of $2,167 per person). The 68,014-person control group retained their debt.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Credit reporting sub-experiment&lt;/strong&gt;: Partway into the collector debt experiment, the debt collector ceased reporting medical debt to the credit bureaus, reflecting an industry-wide trend. The authors isolate 2,761 accounts (6.8% of wave 1) that were reported prior to treatment assignment to estimate the effects of debt relief when accounts would have been counterfactually reported, compared to the subsequent no-reporting environment.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Outcomes are tracked using quarterly depersonalized credit bureau data from TransUnion (spanning at least four quarters before to four quarters after treatment), collections account data on future bill accrual, and a multimodal survey of 2,888 hospital debt experiment respondents measuring mental and physical health, healthcare utilization, and financial wellness. The primary credit-bureau outcome is the number of accounts past due; the primary survey outcome is the share with at least moderate depression (PHQ-8).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Credit market outcomes (main experiments)&lt;/strong&gt;: In both the hospital and collector debt experiments — where there is no counterfactual credit bureau reporting — debt relief has no average effect on financial distress, credit access, or credit utilization. The effect on the number of accounts past due is -0.01 (statistically insignificant; 95% CI excludes effects smaller than -0.04, relative to a control mean of 1.20). Effects on credit card balances (95% CI: -$42 to $47 relative to a mean of $1,481) and auto loan balances (95% CI: -$235 to $148 relative to a mean of $8,020) are similarly precise nulls. These null effects hold for the hospital debt sample (younger debt, 1.3 years old on average) and the collector debt sample (older debt, 7.0 years old on average), and across all preregistered subgroups.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Credit reporting sub-experiment&lt;/strong&gt;: When control group accounts are counterfactually reported, debt relief immediately raises credit scores by an economically small average of 3.4 points (p-value 0.021), with a larger 13.8-point increase (p-value 0.008) for persons with no other debt in collections. Credit limits grow gradually, reaching $340 (15.3% of the post-reporting control mean of $2,231; p-value 0.010) after the no-reporting period begins, with larger effects for those with no other debt in collections. Once control group reporting ceases, both the credit score and credit limit effects converge to zero for those with other debts in collections. No effects on borrowing or financial distress measures are detected in this sub-experiment.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Collections account outcomes (bill repayment)&lt;/strong&gt;: Debt relief causes a statistically significant 1.1 percentage-point increase in the probability of having another unpaid bill sent to collections (6.6% of the control mean of 16.2%; p-value &amp;lt; 0.05) and a $15 increase in the dollar amount of future medical debt sent to collections (7.2% of the control mean of $208). The increase is almost entirely attributable to pre-relief medical services, indicating reduced repayment of existing bills rather than greater healthcare utilization.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Survey outcomes&lt;/strong&gt;: There are no detectable average effects on depression (primary outcome), anxiety, stress, subjective well-being, or general health. Debt relief raises the share with at least moderate depression by a statistically insignificant 3.2 percentage points (p-value 0.097; control mean 45.0%); a 95% CI rules out a reduction of more than 0.6 percentage points, well below the 7.0 percentage-point improvement predicted by the median expert respondent. There are similarly null effects on healthcare utilization and financial wellness as measured in the survey.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The study focuses specifically on downstream medical debt in collections — debt that has already been through the hospital billing cycle and sold to third-party collectors. Results do not necessarily apply to upstream debt relief (e.g., financial assistance programs applied closer to the time of the medical event), nor to populations with different baseline financial profiles. The credit reporting results are most relevant to the prior regime of widespread reporting; under the current environment in which most medical debt has been removed from credit reports, the credit-access channel is largely foreclosed.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-did-the-authors-focus-specifically-on-downstream-medical-debt-in-collections-and-how-does-this-define-the-scope-of-their-study"&gt;Q1. Why did the authors focus specifically on downstream medical debt in collections, and how does this define the scope of their study?&lt;/h3&gt;
&lt;p&gt;The authors focus on downstream medical debt because this is the target of essentially all large-scale government and private relief programs working with RIP Medical Debt, and because it is the category of debt that is most comprehensively observable. Downstream medical debt is defined as bills that have been or are about to be sold by the healthcare provider to a third-party debt collector. This focus excludes upstream unpaid bills still held by the hospital, bills being paid over time, and medical expenses charged to credit cards. The distinction matters because prior literature on hospital financial assistance programs finds substantial benefits from upstream interventions that relieve debt closer to the precipitating medical event; the authors&amp;rsquo; null results are explicitly scoped to the downstream, post-collection stage.&lt;/p&gt;
&lt;h3 id="q2-why-did-the-purchase-price-of-medical-debt-55-cents-per-dollar-for-hospital-debt-less-than-1-cent-per-dollar-for-collector-debt-suggest-caution-about-expected-financial-impacts-ex-ante"&gt;Q2. Why did the purchase price of medical debt (5.5 cents per dollar for hospital debt, less than 1 cent per dollar for collector debt) suggest caution about expected financial impacts ex ante?&lt;/h3&gt;
&lt;p&gt;The authors argue that in a competitive market, the purchase price of medical debt reflects the sum of expected recovery rates and collection costs. A price of 5.5 cents per dollar implies that actual recovery (what collectors expect to collect from patients) is very low. Even if all of the expected recovery is passed through to the patient as a financial benefit, the direct liquidity gain from debt forgiveness is a small fraction of the debt&amp;rsquo;s face value. For the collector debt experiment, where the purchase price is less than 1 cent per dollar, the expected direct financial benefit to recipients is even smaller. The authors note that survey respondents expected to pay 54% of their outstanding medical debt and thought it fair to pay 37%, suggesting that perceived (rather than actual) payment obligations may be what connects medical debt to financial behavior.&lt;/p&gt;
&lt;h3 id="q3-how-was-random-assignment-implemented-in-the-hospital-debt-experiment-and-what-design-features-ensure-the-validity-of-the-experiment"&gt;Q3. How was random assignment implemented in the hospital debt experiment, and what design features ensure the validity of the experiment?&lt;/h3&gt;
&lt;p&gt;Within each of 18 waves between August 2018 and October 2020, RIP received a portfolio of unpaid bills from the hospital system. Persons were grouped at the individual level and stratified by the amount of debt, state of residence, insurance status, and a collections score predicting repayment likelihood. Within strata, persons were randomly assigned to treatment or control, with approximately 20% treated per wave (varying with donor funding). The hospital was unaware of the intervention, eliminating scope for selection of particularly uncollectible accounts. Treatment notification occurred via two letters sent approximately three and six weeks post-purchase. Balance tests confirm successful randomization: all p-values on baseline characteristics are above 0.05, and F-tests fail to reject joint balance.&lt;/p&gt;
&lt;h3 id="q4-what-was-the-credit-reporting-sub-experiment-and-how-was-it-identified"&gt;Q4. What was the credit reporting sub-experiment and how was it identified?&lt;/h3&gt;
&lt;p&gt;The debt collector in the collector debt experiment historically reported medical debt to the credit bureaus but largely ceased doing so before the first intervention wave (March 2018), reflecting broader industry concerns about CFPB enforcement and data integrity risk. However, a subset of accounts — 2,761 accounts (6.8% of wave 1, with virtually identical match rates across treatment and control) — were still being reported until 2019 Q1 (three quarters after wave 1 and one quarter after wave 2). This created a natural sub-experiment: for this subset, treatment group accounts were removed from credit reports immediately upon debt relief, while control group accounts continued to be reported for three more quarters before also being removed. The authors identify reported accounts by matching dollar amounts in collections account data to credit bureau tradeline data in the four quarters prior to intervention, and use this variation to estimate effects separately for the &amp;ldquo;reporting&amp;rdquo; and &amp;ldquo;no-reporting&amp;rdquo; periods.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-exact-estimated-effects-on-credit-scores-and-credit-limits-in-the-credit-reporting-sub-experiment"&gt;Q5. What are the exact estimated effects on credit scores and credit limits in the credit reporting sub-experiment?&lt;/h3&gt;
&lt;p&gt;During the three quarters when control group accounts are still reported to credit bureaus, debt relief raises credit scores by an average of 3.4 points (p-value 0.021) for the full reporting subsample. The effect is concentrated among those with no other debt in collections: 13.8 points (p-value 0.008) versus 1.2 points (p-value 0.440) for those with other debt in collections. Credit limits increase gradually, reaching $340 (15.3% of the post-reporting control mean of $2,231; p-value 0.010) by the four quarters after control group reporting ceases. Among persons with no other debt in collections, this credit limit effect grows to $922 (23% of the control mean; p-value 0.070). Once control group reporting stops, both the credit score effect and the credit limit growth converge to zero for persons with other debts in collections. The event study coefficients show the credit limit effect growing approximately linearly over five quarters post-intervention before leveling out.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-rule-out-the-possibility-that-medical-debt-relief-increases-healthcare-utilization-thereby-causing-more-future-medical-bills"&gt;Q6. How does the paper rule out the possibility that medical debt relief increases healthcare utilization, thereby causing more future medical bills?&lt;/h3&gt;
&lt;p&gt;The collections account analysis separates future debt accrual into debt associated with pre-relief medical services (which can only result from reduced repayment of existing bills) and post-relief medical services (which could reflect either increased utilization or changed repayment of new bills). Panel B of Table VI shows that virtually all of the increased debt sent to collections — a $15 increase and 1.1 percentage-point increase in the probability of any future collection — is attributable to pre-relief services. Panel C shows statistically insignificant increases in future debt from post-relief services. The authors therefore attribute the effect to reduced payment of existing bills and conclude they &amp;ldquo;cannot rule in or rule out effects on healthcare utilization&amp;rdquo; for the post-relief services channel, but the dominant mechanism is behavioral change in repayment of already-incurred debt.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-three-mechanisms-proposed-to-explain-the-reduction-in-repayment-of-existing-medical-bills-and-which-mechanism-is-rejected"&gt;Q7. What are the three mechanisms proposed to explain the reduction in repayment of existing medical bills, and which mechanism is rejected?&lt;/h3&gt;
&lt;p&gt;The authors offer three candidate mechanisms for the 6.6% relative increase in the probability of future bill collections: (i) an expectations mechanism, in which beneficiaries reduce payments because they anticipate future debt relief from similar charitable programs; (ii) a targeting mechanism, drawing on Dobkin et al. (2018), in which patients tolerate a certain level of indebtedness — relieving some debt creates &amp;ldquo;room&amp;rdquo; in their debt budget, so they reduce payment of remaining bills to return to that target level; and (iii) a confusion mechanism, in which recipients mistakenly believe the relief applied to non-forgiven bills (the notification letter explicitly stated &amp;ldquo;the forgiveness is for this outstanding bill only&amp;rdquo; but patients may not have internalized this). The income effect or &amp;ldquo;flypaper&amp;rdquo; mechanism — the idea that financial relief of existing debt frees up mental-account resources for paying medical bills, thereby increasing repayment — is explicitly rejected by the data, as the effect goes in the direction of less repayment, not more.&lt;/p&gt;
&lt;h3 id="q8-what-did-the-expert-survey-predict-and-how-did-those-predictions-compare-to-the-experimental-estimates"&gt;Q8. What did the expert survey predict, and how did those predictions compare to the experimental estimates?&lt;/h3&gt;
&lt;p&gt;An expert survey conducted between April and May 2022 — after the interventions were completed but before results were released — asked academics, non-profit staff, hospital revenue-cycle practitioners, and policymakers to predict the impact of the hospital debt experiment. The median expert predicted a 7.0 percentage-point reduction in depression (8.0 points when weighted by confidence), a 10.2 percentage-point reduction in borrowing (13.7 points when confidence-weighted), and meaningful improvements in healthcare access. In total, 75.6% of respondents predicted medical debt relief is at least a moderately valuable use of charity resources, and 51.1% thought it very or extremely valuable. The authors estimate a statistically insignificant 3.2 percentage-point increase in depression (not a decrease), and a 95% confidence interval that rules out a reduction in depression of more than 0.6 percentage points — far below the 7.0 percentage-point expert prediction.&lt;/p&gt;
&lt;h3 id="q9-what-survey-methodology-was-used-and-what-response-rate-was-achieved"&gt;Q9. What survey methodology was used, and what response rate was achieved?&lt;/h3&gt;
&lt;p&gt;The survey, administered by NORC at the University of Chicago, targeted a random subset of 14,922 hospital debt experiment participants who entered the study after September 2019 (waves 6-18) and owed at least $500. The protocol spanned 13 weeks and included five postal mailings (including a $2 upfront incentive and a $5 incentive with the paper survey), twice-weekly email reminders, certified mail delivery of the full survey instrument, and telephone interviews by a US-based call center. Respondents received a $50 completion incentive. The protocol achieved a 19.4% response rate, with 68% responding via web, 10% via telephone, and 23% via mail. The survey was titled &amp;ldquo;Health and Financial Wellness Study&amp;rdquo; and made no reference to RIP Medical Debt to avoid priming respondents. Respondents were surveyed on average 13 months after treatment assignment (interquartile range 10 to 17 months).&lt;/p&gt;
&lt;h3 id="q10-what-heterogeneity-in-survey-outcomes-was-detected-and-how-do-the-authors-interpret-the-anomalous-depression-finding-for-high-debt-recipients"&gt;Q10. What heterogeneity in survey outcomes was detected, and how do the authors interpret the anomalous depression finding for high-debt recipients?&lt;/h3&gt;
&lt;p&gt;Across all four preregistered heterogeneity dimensions (medical debt amount, age of debt, age of person, amount of other debt in collections), null effects on survey outcomes were found in 15 of 16 subgroups. The exception is persons in the fourth quartile of medical debt eligible for relief, for whom debt relief caused a statistically significant 12.4 percentage-point increase in depression (p-value 0.002) relative to a control mean of 45.9%, with similar patterns for anxiety, stress, subjective well-being, and general health. The authors consider this may be a statistical fluke given the null results across all other 15 groups. They also note potential parallels with findings from unconditional cash transfer experiments, where the receipt of transfers raised the salience of financial deprivation without addressing its underlying causes. A charity-stigma mechanism (recipients did not request the assistance) is also considered. The authors caution against giving this result undue weight in the overall assessment.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-position-downstream-debt-relief-relative-to-upstream-interventions-and-what-does-prior-evidence-suggest-about-upstream-alternatives"&gt;Q11. How does the paper position downstream debt relief relative to upstream interventions, and what does prior evidence suggest about upstream alternatives?&lt;/h3&gt;
&lt;p&gt;The authors highlight that their null results do not extend to upstream medical debt relief. Adams et al. (2022), studying a hospital financial assistance program at Kaiser Permanente that bundled debt relief with reductions in cost-sharing close to the time of the medical event, found substantial increases in high-value healthcare utilization. The Oregon Health Insurance Experiment (Baicker et al. 2013) found that Medicaid reduced depression by 9 percentage points among low-income uninsured adults. The authors suggest several reasons why downstream relief may fail: the intervention occurs too late after the precipitating event (approximately 15 months after the medical service in the hospital debt experiment, and about 7 years in the collector debt experiment), patients may have habituated to the stress of debt collections, the relief amount may be too small relative to overall financial distress, and the direct financial benefit is inherently limited by the low market price of collections-stage debt.&lt;/p&gt;
&lt;h3 id="q12-how-do-the-authors-address-concerns-about-differential-survey-response-and-external-validity"&gt;Q12. How do the authors address concerns about differential survey response and external validity?&lt;/h3&gt;
&lt;p&gt;Treated persons were a statistically insignificant 1.3 percentage points more likely to respond to the survey (p-value 0.056). The authors address this in two ways. First, they estimate specifications that (i) add rich observable controls and (ii) use speed of survey response as a proxy for unobserved response propensity; neither exercise changes the estimates meaningfully. Second, to probe external validity, they test for heterogeneous effects by predicted response propensity (from a logistic regression of a response indicator on baseline characteristics) and by speed of response; neither yields evidence of differential effects for non-respondents. They also compare credit bureau treatment effects for the full hospital debt sample, the survey outreach sample, and the survey respondent sample and find similar estimates across all three groups.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Downstream medical debt&lt;/strong&gt;: Medical bills that have already been sent to third-party debt collectors by the healthcare provider after the initial billing cycle, as distinguished from upstream unpaid bills still held by the hospital at or near the time of the medical event. The paper studies debt at this late stage specifically because it is the target of most large-scale relief programs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Credit reporting sub-experiment&lt;/strong&gt;: An embedded quasi-experiment within the collector debt RCT, exploiting the fact that a subset of accounts (6.8% of wave 1) were still being reported to credit bureaus at the time of intervention while the debt collector had already ceased reporting for the remaining accounts. This allows separate estimation of debt relief effects with and without counterfactual credit bureau reporting, using the period until 2019 Q1 (when the collector stopped reporting entirely) as the &amp;ldquo;reporting&amp;rdquo; window.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Downstream bill repayment effect&lt;/strong&gt;: The paper&amp;rsquo;s finding that debt relief increases the probability of a subsequent unpaid medical bill being sent to collections. The paper attributes this primarily to reduced repayment of existing pre-relief medical bills rather than to increased healthcare utilization, consistent with an expectations, targeting, or confusion mechanism — and inconsistent with an income or flypaper effect that would increase repayment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Targeting a level of indebtedness&lt;/strong&gt;: A behavioral model (drawn from Dobkin et al. [2018]) in which patients implicitly target a certain level of indebtedness. Under this model, relieving some debt creates headroom in the patient&amp;rsquo;s implicit debt budget, leading to reduced repayment of remaining bills to restore the targeted level of total indebtedness.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Expert survey (pre-results)&lt;/strong&gt;: A structured elicitation of predicted treatment effects conducted between April and May 2022 — after the interventions were completed but before results were released — from academics, non-profit practitioners, hospital revenue-cycle managers, and policymakers. Used as a benchmark to quantify how far the causal estimates fall below prevailing beliefs, and to document that the null results were ex ante surprising to informed observers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PHQ-8 (Patient Health Questionnaire-8)&lt;/strong&gt;: An eight-item validated clinical screen for depression, used as the paper&amp;rsquo;s primary preregistered survey outcome. An indicator for &amp;ldquo;at least moderate depression&amp;rdquo; on the PHQ-8 is the main mental health measure against which the debt relief treatment effect is estimated.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Multimodal survey&lt;/strong&gt;: A survey protocol combining five postal mailings, twice-weekly email reminders, certified mail delivery of a paper survey instrument, and US-based call center telephone interviews, designed to maximize response rates in a hard-to-reach low-income population with medical debt in collections.&lt;/p&gt;</description></item><item><title>The Future in Mind: Aspirations and Long-Term Outcomes in Rural Ethiopia</title><link>https://macropaperwarehouse.com/papers/the-future-in-mind-aspirations-and-long-term-outcomes-in-rural-ethiopia/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-future-in-mind-aspirations-and-long-term-outcomes-in-rural-ethiopia/</guid><description>&lt;p&gt;This paper tests whether a light-touch behavioral intervention targeting aspirations can produce persistent economic effects on a poor rural population. The research question is whether changing how poor people perceive their future opportunities — by raising aspirations — alters their investment decisions in ways that persist over a multi-year horizon. The authors conduct a randomized controlled trial in Doba, a remote mountainous district in rural Ethiopia roughly 380 kilometers from Addis Ababa, selected partly because its extreme isolation meant residents had almost no exposure to television or media, making even a single video screening a memorable event.&lt;/p&gt;
&lt;p&gt;The sample consists of 1,152 households (2,112 individuals) across 64 villages. Households were randomly assigned to one of three conditions: a treatment group shown four 15-minute documentaries featuring real rural individuals from similar communities who escaped poverty through goal-setting and hard work; a placebo group shown an Ethiopian entertainment comedy with no aspirational content; and a within-village control group who were only surveyed. Both the household head and spouse in treatment and placebo groups were invited to attend. Compliance was very high, with only 2 percent of individuals not complying with their assigned condition. Data were collected at baseline (2010), six months after screening (2011), and five years after baseline (2015–2016). Attrition was notably low: 96 percent of households were re-interviewed at the five-year endline, and 94 percent of individual respondents.&lt;/p&gt;
&lt;p&gt;Five years after the screening, treated households show meaningfully larger investment across three domains relative to the control group, with all headline results significant at 5 percent or less and robust to multiple hypothesis testing. First, on agricultural effort and investment: treated household heads and spouses work approximately one extra hour per day on their own farms (roughly 8.6 percent of the control mean per spouse). Treated households are 10 percentage points more likely to have adopted modern crop inputs (improved seeds, inorganic fertilizer) and 10 percentage points more likely to have invested in modern livestock inputs (feed, veterinary supplies). Holdings of productive tools are 20 percent higher than in the control group. Second, on educational investment: treated households spend approximately 36 percent more on children&amp;rsquo;s schooling than the control group. Among children who were of school-going age at the time of the intervention (aged 11–15 then, 16–20 at endline), the number completing full primary school is nearly double the control rate (0.16 per household versus 0.07 in the control). Third, on living standards: treated households experienced 0.33 to 0.38 fewer months of food insecurity in the previous year. Their holdings of consumer durables (furniture, kitchenware, phones) are 29 percent higher than the control group in value. Estimated house values are 27 percent higher. However, there is no statistically significant effect on measured food or frequent non-food consumption expenditure, a finding the authors interpret as consistent with households continuing to divert resources toward future-oriented investments rather than current consumption.&lt;/p&gt;
&lt;p&gt;The intervention&amp;rsquo;s effects appear to operate primarily through aspirations — defined in this paper as desired goals for the future that motivate investment and effort. Treated households report significantly higher aspirations and expectations for income, assets, and children&amp;rsquo;s education five years later. By contrast, the paper finds no persistent changes in time preferences, risk preferences, grit, or beliefs about returns to technology. Locus of control shifted six months after the intervention but did not persist to the five-year endline, and the authors argue that if locus of control were the operative mechanism, investment effects would also have dissipated. The placebo group shows no significant effects relative to the control, ruling out screening exposure or social attention as mechanisms.&lt;/p&gt;
&lt;p&gt;The paper is explicit about scope conditions. The study area was deliberately chosen for its extreme remoteness and media isolation, and the authors caution that this may have amplified the intervention&amp;rsquo;s salience and persistence relative to less isolated populations. External validity beyond comparable settings is uncertain. A back-of-the-envelope cost-effectiveness calculation finds that increases in durable asset holdings alone outweigh intervention costs by a factor of approximately two at reasonable scale.&lt;/p&gt;
&lt;p&gt;Q: What was the intervention and what made it distinct from other role model studies?
A: Treated households were invited to watch four 15-minute documentary films featuring real rural individuals from similar socioeconomic backgrounds who had escaped poverty through goal-setting, perseverance, and hard work. The films were produced in Oromiffa, the local language, and featured two male and two female role models depicting achievable actions such as installing irrigation or starting a small business. Unlike studies that vary exposure to in-person mentors or peers, participants received no ongoing mentorship, financial resources, or support of any kind beyond the single video screening, isolating the aspirations channel from material or informational transfers.&lt;/p&gt;
&lt;p&gt;Q: How were aspirations measured and validated?
A: Aspirations were measured using locally validated survey instruments (Bernard and Taffesse, 2014) that asked respondents what level of annual income, asset wealth, and oldest child&amp;rsquo;s education they would like to achieve in their lifetime. Test-retest reliability over two weeks produced within-respondent correlations of 0.77 to 0.98 across domains, which the authors benchmark against Angrist and Krueger (1999) standards for reliable income and education measures. The measures correlated in expected directions with wealth: mean income aspirations in the upper wealth tercile were 1.5 times those in the lower tercile, and asset aspirations in the upper tercile were 1.9 times those in the lower tercile.&lt;/p&gt;
&lt;p&gt;Q: What were the five-year effects on agricultural effort and investment?
A: Treated household heads and spouses worked approximately half an hour more per day each on their own farms relative to control, implying roughly one extra hour per day across the typical household&amp;rsquo;s adult members — an 8.6 percent increase over the control mean. Treated households were 10 percentage points more likely to have adopted modern crop inputs and 10 percentage points more likely to have invested in modern livestock inputs. Holdings of productive tools were 20 percent higher in value than in the control group. The overall agricultural investment index increased by 0.21 standard deviations relative to the control and 0.18 standard deviations relative to the placebo.&lt;/p&gt;
&lt;p&gt;Q: What were the five-year effects on children&amp;rsquo;s education?
A: Among children aged 16 to 20 at endline (who were 11 to 15, upper primary school age, at the time of the intervention), the number per household completing full primary school nearly doubled: 0.16 in the treatment group versus 0.07 in the control. These children in treated households also spent on average 33 minutes more per day attending school than the control group. Across all children, schooling expenditures in the treatment group were 36 percent higher than in the control and 30 percent higher than in the placebo. The education index increased by 0.25 standard deviations relative to the placebo and 0.21 standard deviations relative to the control.&lt;/p&gt;
&lt;p&gt;Q: Why did consumption expenditure not increase despite improvements in assets and food security?
A: The authors argue that the consumption result is theoretically ambiguous: if treated households continue to divert resources toward future-oriented investments (savings, productive assets, durable goods, housing), intertemporal substitution effects could offset income effects within the five-year observation window. The measured consumption variables — food and frequent non-food spending — do not capture the service flow value of accumulated durables or housing improvements, both of which increased substantially. The authors interpret this as evidence that households were still in an investment phase rather than having converted accumulated wealth into current consumption by endline.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports aspirations as the operative mechanism rather than alternative channels?
A: The treatment group had significantly higher aspirations and expectations for income, assets, and children&amp;rsquo;s education at the five-year endline, while the placebo group did not. Measured time preferences, risk preferences, grit, and beliefs about returns to technology were all statistically unchanged for treated households. Locus of control shifted six months post-intervention but did not persist to five years, and the authors note that if locus of control were the driver, investment effects would also have dissipated alongside it. The null placebo effect rules out screening exposure, social attention, or information salience from outside facilitators as mechanisms.&lt;/p&gt;
&lt;p&gt;Q: How were locus of control and fatalistic beliefs assessed in this population?
A: The sample scored twice as high as Western samples on the classic Levenson (1981) fatalism scale. On the Feagin (1975) scale of perceived causes of poverty, the sample was more likely to attribute poverty to structural or fatalistic explanations than Western samples, and both measures of fatalistic beliefs were higher among poorer households within the sample. The study region&amp;rsquo;s worldview — rooted in traditional Waaqeffannaa religion, local variants of Orthodox Christianity (Fekade Egziabher), and Islam (Qadar) — emphasizes deference to authority, predestination, and resistance to change, providing qualitative grounding for the aspirations deficit being targeted.&lt;/p&gt;
&lt;p&gt;Q: What were the effects on food insecurity and subjective wellbeing?
A: Treated households reported 0.33 fewer months of food insecurity in the previous year relative to the control group (from a base of 2.71 months in the control), and 0.38 fewer months relative to the placebo. Treated participants scored approximately a quarter of a step higher on the Cantril ladder of self-reported wellbeing than the control group. There was no significant difference on the USDA food insecurity questionnaire, which the authors attribute to that scale&amp;rsquo;s unsuitability for households that consume largely from own production.&lt;/p&gt;
&lt;p&gt;Q: What were the effects on durable goods and housing?
A: Treated households reported 29 percent higher value of consumer durables (furniture, kitchenware, phones) than the control group and 32 percent higher than the placebo. Estimated house replacement values were 27 percent higher than the control and 21 percent higher than the placebo. Enumerators directly observed that treated households were more likely to have their own toilet facility, though this result was not significant relative to the placebo. There were no effects on the probability of having a non-organic roof, which the authors note is an especially expensive upgrade.&lt;/p&gt;
&lt;p&gt;Q: How does the paper rule out spillover effects from treated to control households?
A: The authors collected data on a supplementary sample of non-treated villages to serve as a &amp;ldquo;pure control&amp;rdquo; and used this to run a suggestive test for spillovers from treated households to untreated households within the same village. They found little evidence of large spillover effects, although they acknowledge limitations in the power of these tests. The physical design of the screenings — held in rooms with shuttered windows, requiring tickets for entry, conducted separately from placebo screenings — also minimized contamination during the intervention itself.&lt;/p&gt;
&lt;p&gt;Q: What were the early (six-month) results and what do they suggest about the timing of effects?
A: At six months, the shorter follow-up found increases in savings and investment in education, consistent with behavioral change beginning soon after treatment. Aspirations showed positive but noisier effects at immediate post-screening and six-month follow-ups, which the authors interpret as consistent with aspirations increasing gradually as people experiment with alternative futures (Appadurai, 2004) or as demotivating beliefs shift incrementally (Carvalho et al., 2023), rather than changing abruptly. This gradual pattern is consistent with a learn-by-doing dynamic where small initial investments generate returns that further raise aspirations.&lt;/p&gt;
&lt;p&gt;Q: How does this study&amp;rsquo;s attrition and follow-up compare to the literature?
A: The five-year attrition rate was very low: 96 percent of baseline households were re-interviewed and 94 percent of individual respondents. The authors cite Bouguen et al. (2019) as a benchmark, noting this is a high tracking rate relative to recent long-run RCT follow-ups in low- and middle-income countries. The low attrition strengthens confidence that endline estimates are not contaminated by selective dropout.&lt;/p&gt;
&lt;p&gt;Q: What is the cost-effectiveness of the intervention?
A: A back-of-the-envelope calculation indicates that increases in durable asset holdings alone outweigh the costs of the intervention by a factor of approximately two at reasonable implementation scale. The authors present this as a proof-of-concept estimate, not a full social cost-benefit analysis, and caution that cost-effectiveness may differ in settings with higher baseline media exposure or less extreme isolation.&lt;/p&gt;
&lt;p&gt;Q: What are the key scope conditions limiting external validity?
A: The study district (Doba) was chosen specifically for its extreme remoteness: at baseline, only 11 percent of respondents watched TV at least weekly and no household owned a television. The authors argue this isolation likely made the screening event especially salient and memorable, potentially amplifying effects relative to what would be expected in less isolated contexts. They are explicit that the findings represent a proof of concept for the aspirations mechanism and that effect magnitudes should not be assumed to replicate in settings with higher baseline media exposure or different cultural belief systems.&lt;/p&gt;
&lt;p&gt;Aspirations: Defined in this paper as desired goals for the future that motivate investment and effort in order to attain them (following Bandura, 1977; Locke and Latham, 1990). Measured via validated survey instruments asking respondents the level of income, assets, or children&amp;rsquo;s education they would like to achieve in their lifetime — distinct from expectations (what one expects to achieve) and from the village maximum (what one believes the most successful person in the village could achieve).&lt;/p&gt;
&lt;p&gt;Aspirations gap: The difference between an individual&amp;rsquo;s aspired level of income, assets, or education and their current reported level. Median aspirations gaps in the sample are 55 percent of median wealth aspirations and 58 percent of median income aspirations, indicating that aspirations exceed current levels by meaningful but not unrealistic margins.&lt;/p&gt;
&lt;p&gt;Capacity to aspire: Drawn from Appadurai (2004), defined as a navigational capacity — the ability to read and navigate a map of a journey into the future. In contexts of poverty, this capacity is described as more brittle because poorer individuals have narrower social networks, fewer role models, and less material slack for experimentation with alternative futures.&lt;/p&gt;
&lt;p&gt;Role model: A real individual from a similar socioeconomic background whose documented experience of escaping poverty through goal-setting and effort provides vicarious experience that allows audience members to imagine what is possible for people like them. Role models are most effective when their success appears attainable and when the steps to achieve it are visible.&lt;/p&gt;
&lt;p&gt;Zero-sum beliefs: The belief that gains for one individual come at the expense of others in the community, documented in the study area as part of a broader fatalistic, deterministic belief system. These beliefs can suppress effort and future-oriented investment by making individual advancement appear normatively transgressive or materially impossible.&lt;/p&gt;
&lt;p&gt;Source text origin: A classification in the paper&amp;rsquo;s pipeline framework distinguishing whether a summary is based on a full working paper PDF or HTML text versus abstract-only text. Abstract-only summaries are blocked as they miss scope conditions, quantitative results, and the full argument structure.&lt;/p&gt;
&lt;p&gt;Placebo group: Households randomly invited to watch an Ethiopian comedy entertainment program (with no aspirational content) rather than the role model documentaries. Used to separate the effect of the aspirations content from the effects of the screening event itself, exposure to outside facilitators, or social attention accompanying selection for the intervention.&lt;/p&gt;</description></item><item><title>The Geography of job creation and job destruction</title><link>https://macropaperwarehouse.com/papers/the-geography-of-job-creation-and-job-destruction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-geography-of-job-creation-and-job-destruction/</guid><description>&lt;p&gt;This paper asks why unemployment rates differ so persistently across local labor markets, and what role job creation and job destruction play in generating those differences. The authors document a comprehensive set of spatial labor market facts using administrative and survey microdata from Germany, the United States, and the United Kingdom, then build and calibrate a quantitative theoretical framework that accounts for all documented regularities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and scope.&lt;/strong&gt; For Germany, the authors use administrative data from the German employment office (universe of vacancies and unemployed, 1999–2020) and the IAB social security sample (SIAB, 2% of all workers, 2000–2017) aggregated to 194 commuting zones. For the U.S., they use BLS Local Area Unemployment Statistics (2000–2019) at commuting zones, CPS worker flows at metropolitan areas, and JOLTS vacancy data for the 18 largest MSAs (covering roughly 40% of the U.S. labor force). For the UK, they use Nomis data and Jobcentre Plus vacancy records (2004–2006) for 378 Local Authority Districts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical findings.&lt;/strong&gt; Spatial unemployment rate differences are large and highly persistent. In Germany, the correlation of local unemployment rates across commuting zones over a 19-year span is 0.84 (West) and 0.77 (East). In the U.S., the correlation between 2000 and 2019 unemployment rates is 0.81; in the UK it is 0.76. In all three countries, local labor markets with lower unemployment are tighter (more vacancies per unemployed worker) and less productive. Firms in low-unemployment markets fill vacancies more slowly — in Germany, vacancy duration ranges from approximately 35 days in high-unemployment locations to approximately 65 days in low-unemployment locations, roughly an 85% difference.&lt;/p&gt;
&lt;p&gt;A formal steady-state decomposition reveals that across all three countries, differences in job-separation rates account for approximately two-thirds of the cross-sectional variation in unemployment rates, while differences in job-finding rates account for roughly one-third. Specifically: Germany 62.4% separations / 33.2% job-finding; U.S. 72.0% / 32.8%; UK 64.3% / 35.8%. This primacy of separation rates in the cross-section stands in stark contrast to business-cycle dynamics, where job-finding rates account for 50–60% of unemployment fluctuations (Fujita and Ramey, 2009).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Theory.&lt;/strong&gt; The authors embed a Diamond-Mortensen-Pissarides (DMP) model with endogenous separations — following Den Haan, Ramey, and Watson (2000) — into a Rosen-Roback spatial equilibrium framework. Locations differ in exogenous productivity; workers and firms are freely mobile; cost-of-living differences sustain the spatial equilibrium. The model is calibrated to the U.S. median-unemployment labor market (separation rate 0.0128, job-finding rate 0.2368, vacancy-filling rate 0.7365) plus the productivity differential between the 5th and 95th percentile unemployment locations (4.8% higher and 3.0% lower productivity than median, respectively). The baseline model, imposing the Hosios condition, matches the spatial patterns of separation rates, job-finding rates, tightness, vacancy duration, wages, and cost of living without targeting most of these. The decomposition in the calibrated baseline model attributes 33.5% of spatial unemployment variation to job-finding rates, compared to 32.8% in the data.&lt;/p&gt;
&lt;p&gt;The baseline model generates a counterfactual upward-sloping Beveridge curve and cannot explain why job-finding rates dominate business-cycle fluctuations. Introducing on-the-job search (with 12% of employed workers searching each period, calibrated from Faberman et al., 2017) resolves both problems. In the extended model, job-to-job transition rates are virtually constant across local labor markets (matching the data) but strongly procyclical over the business cycle. This asymmetry amplifies the response of vacancies and job-finding rates to aggregate productivity shocks while muting the cyclical variation in separation rates. The extended model&amp;rsquo;s business-cycle decomposition attributes 54.4% of unemployment volatility to job-finding rates, within the empirical 50–60% range.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy implications.&lt;/strong&gt; Under the Hosios condition, the decentralized equilibrium is efficient — large spatial differences in unemployment, tightness, and wages are efficient outcomes, not signs of mismatch. The relevant policy benchmark is not deviation of tightness from the national average but deviation from the model&amp;rsquo;s location-specific prediction conditional on local productivity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the central empirical puzzle the paper addresses?&lt;/strong&gt;
A: Spatial unemployment differences are large and persistent — in Germany, unemployment rates ranged from 1.9% to 11.9% across commuting zones even after 15 years of decline. These differences are not well understood theoretically, and the crucial missing empirical piece was data on job creation and vacancy filling across locations, which this paper provides for three countries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How large and persistent are cross-sectional unemployment differences in each country?&lt;/strong&gt;
A: In Germany, commuting-zone unemployment ranged from 3.6% to 24.0% in 2000 and persisted with a 19-year correlation of 0.84 (West) and 0.77 (East). In the U.S., the 2000–2019 correlation is 0.81, with unemployment as low as 1.5% and as high as 16.9% in 2000. In the UK, the 2004–2018 correlation is 0.76, with 2004 unemployment ranging from 1.8% to 13.1%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What do the data show about the relationship between unemployment and labor market tightness across locations?&lt;/strong&gt;
A: In all three countries, lower-unemployment labor markets are tighter — they have more vacancies per unemployed worker. This is documented for Germany using the universe of registered vacancies, for the U.S. using JOLTS data for 18 large MSAs, and for the UK using Jobcentre Plus administrative data. The relationship holds after controlling for local labor market composition (age, gender, education, occupation, industry shares).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What do vacancy-filling rates look like across locations, and how large are the differences?&lt;/strong&gt;
A: Vacancy-filling rates are lower in low-unemployment (tight) labor markets. In Germany, the monthly probability of filling a vacancy is approximately 50% higher in high-unemployment markets than in low-unemployment markets. Completed vacancy duration ranges from about 35 days in high-unemployment locations to about 65 days in low-unemployment locations — a difference of approximately 85%. The UK data show a strikingly similar elasticity of vacancy-filling rates with respect to unemployment rates to Germany.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the formal decomposition reveal about the sources of spatial unemployment differences?&lt;/strong&gt;
A: In a steady-state two-state decomposition, separation rates account for 62.4% (Germany), 72.0% (U.S.), and 64.3% (UK) of cross-sectional unemployment variation, while job-finding rates account for 33.2%, 32.8%, and 35.8%, respectively, with small residuals. This consistently assigns primary importance to separation rates across all three countries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why is the primacy of separation rates in the cross section surprising, and what literature does it contrast with?&lt;/strong&gt;
A: The business-cycle literature (Fujita and Ramey, 2009; Shimer, 2012) finds that job-finding rate variation accounts for 50–60% of unemployment fluctuations over the cycle, roughly twice the contribution of separation rates. The spatial pattern is the mirror image: separations dominate. Any credible theory of spatial unemployment must rationalize both patterns simultaneously — a challenge the paper explicitly takes up.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the baseline DMP model with endogenous separations generate the spatial patterns?&lt;/strong&gt;
A: Higher-productivity locations feature higher match surpluses. Higher surplus induces more vacancy creation and tighter markets, raising job-finding rates and lowering vacancy-filling rates. Crucially, a higher surplus means idiosyncratic shocks must be more negative to make the joint surplus negative, so fewer matches dissolve — separation rates are lower. The calibrated model reproduces the 32.8% job-finding / ~67% separation decomposition without targeting it (model yields 33.5% job-finding).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the calibration targets and key parameter values in the baseline model?&lt;/strong&gt;
A: The model is calibrated monthly to the U.S. economy. Median-unemployment-location targets: separation rate 0.0128, job-finding rate 0.2368, vacancy-filling rate 0.7365. Productivity targets: the 5th-percentile-unemployment location is 4.8% more productive than median, and the 95th-percentile-unemployment location is 3.0% less productive. Key calibrated values include matching elasticity alpha = 0.4711 (equal to worker bargaining power under Hosios), matching efficiency m = 0.4371, vacancy posting cost kappa = 0.3070, and flow nonmarket value z = 0.9072.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the two shortcomings of the baseline model, and how does on-the-job search resolve them?&lt;/strong&gt;
A: The baseline model generates a counterfactual upward-sloping Beveridge curve and cannot generate the asymmetry between cross-sectional and business-cycle drivers of unemployment. Adding on-the-job search (fraction phi = 0.12 of employed workers searching, calibrated from Faberman et al., 2017) resolves both. It corrects the Beveridge curve by allowing the model to match the spatial vacancy-unemployment relationship, and it introduces procyclical job-to-job mobility that amplifies the cyclical response of job-finding rates while dampening cyclical separation rate variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do job-to-job transition rates differ across space versus over the business cycle, and why does this matter?&lt;/strong&gt;
A: Job-to-job rates are virtually constant across the cross-section of local labor markets (the extended model is calibrated to match this). But they are strongly procyclical — high in booms, low in recessions, about as volatile as job-finding rates over the cycle. In a boom, more employed workers search, spurring vacancy creation, which raises both vacancy-filling probability (making vacancies easier to fill) and job-finding probability for the unemployed, amplifying the cyclical job-finding rate response while muting the cyclical separation rate response.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the extended model predict for business-cycle dynamics?&lt;/strong&gt;
A: The model with on-the-job search and aggregate productivity shocks (parameterized following Hagedorn and Manovskii, 2008) generates unemployment and vacancy rates that are an order of magnitude more volatile than productivity — matching the data. Labor market tightness is about twice as volatile as unemployment, as in the data. The Fujita-Ramey decomposition in the model attributes 54.4% of unemployment volatility to job-finding rates, which falls within the empirical range of 50–60%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the paper&amp;rsquo;s efficiency result and its policy implication?&lt;/strong&gt;
A: Under the Hosios condition (imposed in calibration), the decentralized equilibrium is efficient: job creation and destruction are privately efficient in each market, and free mobility of workers and firms ensures efficient spatial allocation. Therefore, large observed differences in unemployment, tightness, and wages across locations are not evidence of inefficiency. The relevant signal for policy is not deviation from the national average but deviation from the model&amp;rsquo;s location-specific prediction conditional on productivity. Locations where data deviate from model predictions are candidates for policy intervention.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Do the spatial patterns survive controls for worker and firm composition?&lt;/strong&gt;
A: Yes. The authors regress labor market tightness and vacancy-filling rates on local unemployment rates and a full set of composition controls (age, gender, education, occupation, and industry shares) derived from the IAB microdata for Germany, along with year fixed effects. The relationship between local unemployment and both tightness and job-filling rates remains highly statistically and economically significant after these controls, for both Germany and the U.S.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the model handle wages and cost of living, and does it match the data?&lt;/strong&gt;
A: Wages are determined by state-contingent generalized Nash bargaining with worker bargaining power eta. Cost-of-living differences are backed out as the values needed to sustain the spatial equilibrium (Rosen-Roback). Neither wages nor costs of living are calibration targets in the cross section, yet the model closely matches the empirically observed wage gradient across local labor markets and the negative correlation between cost of living and local unemployment (using Economic Policy Institute Family Budget Calculator data).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor market tightness:&lt;/strong&gt; The ratio of vacancies posted in a local labor market to the number of unemployed workers in that market; the paper documents that tightness is systematically higher (more vacancies per unemployed worker) in lower-unemployment locations across all three countries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Job-separation rate (EU rate):&lt;/strong&gt; The share of employed workers who transition from employment to unemployment in a period; in the paper&amp;rsquo;s framework, this is endogenously determined by the idiosyncratic match productivity threshold below which the joint match surplus turns negative, and it is the primary driver of spatial unemployment differences (accounting for roughly two-thirds of cross-sectional variation).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Job-finding rate (UE rate):&lt;/strong&gt; The share of unemployed workers who transition from unemployment to employment in a period; in the paper&amp;rsquo;s framework, this is higher in tighter (lower-unemployment) markets, but accounts for only roughly one-third of spatial unemployment variation — the opposite of its dominant role in business-cycle fluctuations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Spatial Beveridge curve:&lt;/strong&gt; The cross-sectional relationship between vacancy rates and unemployment rates across local labor markets; in the data it is downward sloping (low-unemployment locations have both high vacancies and low unemployment), which the baseline model fails to capture but the extended model with on-the-job search reproduces.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous separation threshold:&lt;/strong&gt; The location-specific minimum idiosyncratic match productivity below which the joint match surplus becomes negative and the worker-firm pair dissolves; this threshold is lower (tolerates a wider range of idiosyncratic shocks) in higher-productivity locations because the average surplus is larger, generating lower separation rates in more productive locations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Spatial equilibrium (Rosen-Roback):&lt;/strong&gt; The equilibrium condition in which differences in local costs of living adjust to make workers and firms indifferent across locations, sustaining persistent productivity-driven differences in wages and unemployment as equilibrium outcomes rather than disequilibrium phenomena.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Procyclical on-the-job search:&lt;/strong&gt; The mechanism by which the fraction of employed workers actively searching — and thus the rate of job-to-job transitions — is approximately constant across the cross-section of local labor markets but strongly procyclical over the business cycle. This asymmetry is the key to reconciling why job-finding rates drive business-cycle unemployment variation while separation rates drive spatial unemployment variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hosios condition:&lt;/strong&gt; The parametric restriction equating the unemployment elasticity of the matching function (alpha) and the workers&amp;rsquo; Nash bargaining weight (eta); when satisfied, job creation is efficient in every local labor market. The paper imposes this condition deliberately to demonstrate that the decentralized equilibrium is efficient despite large spatial differences in outcomes.&lt;/p&gt;</description></item><item><title>The Illiquidity of Water Markets</title><link>https://macropaperwarehouse.com/papers/the-illiquidity-of-water-markets/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-illiquidity-of-water-markets/</guid><description>&lt;p&gt;Donna and Espín-Sánchez investigate whether a market (sequential English auction) or a non-market institution (fixed quota) more efficiently allocates an intermediate good — irrigation water — when some buyers are liquidity constrained. The setting is Mula, a city in southeastern Spain, where farmers used an unregulated water auction continuously from 1244 until August 1, 1966, when the institution was replaced by a fixed quota system. This 700-year natural experiment, combined with the fact that water demand for a given crop is pinned down by the crop&amp;rsquo;s production function rather than by farmer wealth, allows the authors to separately identify liquidity constraints from unobserved heterogeneity in productivity.&lt;/p&gt;
&lt;p&gt;The empirical context has four features the authors exploit. First, the pre-1966 auction was entirely unregulated, so price differences directly reflect valuations without the confounds of regulatory changes. Second, water is an intermediate good for apricot production; conditional on plot area, tree count, and crop type, demand is determined by the apricot tree&amp;rsquo;s biological water requirements — not by the farmer&amp;rsquo;s wealth — so wealthy and poor farmers growing the same bulida apricot variety share the same underlying demand up to an idiosyncratic productivity shock. Third, farmers are classified as wealthy if they held positive urban real estate (non-agricultural wealth) in 1955 tax records; wealthy farmers&amp;rsquo; average annual urban rental income (5,702 pesetas) far exceeded their average annual water expenditure (500 pesetas, rising to 1,619 in the highest-expenditure year, 1963), supporting the assumption that wealthy farmers were never liquidity constrained. Fourth, the 1966 institutional shift to quotas — under which each farmer received a fixed water allotment (tanda) every three weeks proportional to plot size, paying only a small annual maintenance fee after the critical season — provides the counterfactual.&lt;/p&gt;
&lt;p&gt;The authors build a structural dynamic demand model with three key features: storability (irrigation raises soil moisture, creating intertemporal substitution between periods because water evaporates partially), liquidity constraints (poor farmers cannot always afford water during the critical season when prices peak), and weather seasonality (the critical season, corresponding to apricot fruit growth stages II–III and the Early Post-Harvest period, spans roughly weeks 18–32 and is when trees most need water). Farmers are forward-looking and form expectations about future prices and rainfall. The model&amp;rsquo;s production function, drawn from the agricultural engineering literature (Torrecillas et al., 2000; Allen et al., 2006), transforms soil moisture into apricot output via a transformation rate parameter gamma, a hydric stress coefficient, and a seasonal dummy.&lt;/p&gt;
&lt;p&gt;Demand parameters are estimated using a two-step conditional choice probability (CCP) estimator (Hotz et al., 1994) on wealthy farmers only, then projected onto poor farmers&amp;rsquo; welfare calculations. The sample consists of 24 single-crop apricot farmers observed in weekly auction records from January 1955 to July 1966, embedded in a market with over 500 total participants.&lt;/p&gt;
&lt;p&gt;The main finding is that the institutional change from auction to quota increased total efficiency. Welfare increased by 23.4 real pesetas per farmer per tree, a 6 percent increase in total apricot production relative to the market. This gain arises because: (1) farmers were relatively homogeneous in productivity (small idiosyncratic shocks), so the primary source of misallocation was not productivity heterogeneity but wealth heterogeneity; (2) liquidity constraints prevented poor farmers from purchasing water during the critical season when their valuation was high, causing them instead to buy earlier (at lower prices but with partial evaporation loss) or later (when their trees had already experienced hydric stress); and (3) the apricot production function is concave in water, so uniform quota allocation is more efficient than market allocation when farmers are approximately homogeneous. The paper provides the first empirical demonstration that liquidity constraints can reverse the standard efficiency ranking of markets over quotas.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question?
A: The paper asks whether a free market (water auction) or a non-market institution (fixed quota) more efficiently allocates an intermediate good when some buyers are liquidity constrained. The theoretical ranking is ambiguous when agents are heterogeneous in both productivity and wealth, making this an empirical question. The authors find that quotas dominated the auction in the specific Mula setting.&lt;/p&gt;
&lt;p&gt;Q: What was the historical water market in Mula and when did it end?
A: From 1244 to 1966 — over 700 years — Mula farmers used a sequential ascending-price (English) auction to allocate river water. The auctioneer sold water in discrete units called cuartas (each representing 3 hours of canal flow, or approximately 432,000 liters), holding 40 units per weekly Friday session. Farmers paid in cash on auction day. On August 1, 1966, the farmers&amp;rsquo; union (Sindicato de Regantes) replaced the auction with a fixed quota system, having secured a credit line to purchase water property rights share by share.&lt;/p&gt;
&lt;p&gt;Q: How did the quota system work, and how did it eliminate liquidity constraints?
A: Under the quota, each plot of land received a fixed water allotment (tanda) every three weeks, proportional to plot size. Farmers paid only a small annual maintenance fee to the Sindicato at year-end, after the critical season harvest. Because payment occurred after farmers collected harvest revenue, no farmer was liquidity constrained under the quota. The fee was substantially lower than the per-unit average price under the market.&lt;/p&gt;
&lt;p&gt;Q: How do the authors identify liquidity constraints separately from unobserved heterogeneity in productivity?
A: The key insight is that water is an intermediate good whose demand is determined by the apricot tree&amp;rsquo;s biological production function, not by farmer wealth. Two farmers growing the same bulida apricot variety with the same number of trees should have the same water demand up to an idiosyncratic shock. The authors use wealthy farmers (those with positive urban real estate in 1955 tax records) to estimate preferences, under the assumption that wealthy farmers are never liquidity constrained. They then verify that outside the critical season, wealthy and poor farmers purchase similar amounts of water; the purchasing divergence appears only during the high-price critical season, consistent with a cash constraint rather than a preference difference.&lt;/p&gt;
&lt;p&gt;Q: What empirical evidence shows poor farmers were liquidity constrained rather than simply less interested in water?
A: Poor farmers display a bimodal purchasing pattern inconsistent with the apricot tree&amp;rsquo;s biological water needs: they buy water before the critical season (when prices are low) in anticipation of not being able to afford it during the critical season, and again after the critical season (when prices fall) to prevent their trees from withering from dehydration. Wealthy farmers, by contrast, delay purchases strategically to the critical season when trees most need water (weeks 18–32). Regression analysis confirms that wealthy farmers purchase significantly more water per tree during the critical season than poor farmers growing identical bulida apricots, while the difference outside the critical season is not statistically significant.&lt;/p&gt;
&lt;p&gt;Q: How were wealthy farmers defined and why does their wealth validate the non-constrained assumption?
A: A farmer is defined as wealthy if the value of their urban real estate (from 1955 urban tax records) is positive, and as poor if it is zero. Urban real estate constitutes non-agricultural wealth uncorrelated with the apricot production function. Wealthy farmers&amp;rsquo; average annual urban rental income was 5,702 pesetas, while their average annual water expenditure was only 500 pesetas (rising to 1,619 pesetas in 1963, the highest-expenditure sample year). This large gap supports the assumption that wealthy farmers could always afford water purchases.&lt;/p&gt;
&lt;p&gt;Q: What is the model&amp;rsquo;s treatment of soil moisture dynamics and why does it matter?
A: Soil moisture (M_it) evolves according to an agricultural engineering formula: it increases with rainfall and irrigation purchases (each unit adding 432,000 liters divided by plot area) and decreases via evapotranspiration (ET), subject to a full-capacity ceiling (FC) and a permanent wilting point (PW) lower bound. This storage structure creates intertemporal substitution — water purchased early partially substitutes for future purchases, but at a cost (evaporative loss). The dynamics mean poor farmers who pre-buy water before the critical season lose some of that investment to evaporation, generating a real efficiency loss relative to the quota that delivers water closer to when it is biologically needed.&lt;/p&gt;
&lt;p&gt;Q: What are the two sources of potential inefficiency the authors identify?
A: The first is inefficiency due to heterogeneity: if farmers differ in ex-post productivity (captured by idiosyncratic shocks epsilon_it), allocating water to a less productive farmer at a given moment is wasteful. Markets correct this inefficiency (they direct water to highest-valuation buyers) while quotas do not. The second is inefficiency due to decreasing marginal returns (DMR): because the production function is concave in water, giving water to a farmer with already-high soil moisture is less productive than giving it to a farmer with low moisture. Quotas naturally avoid DMR inefficiency by allocating uniformly; markets with liquidity constraints exacerbate DMR inefficiency by directing scarce critical-season water to wealthy farmers who may have already accumulated moisture from prior purchases.&lt;/p&gt;
&lt;p&gt;Q: What is the main quantitative result of the welfare analysis?
A: Switching from the market auction to the fixed quota system increased welfare by 23.4 real pesetas per farmer per tree, representing a 6 percent increase in total apricot production relative to the market counterfactual. This is computed as the difference in yearly mean welfare per tree per farmer (net of irrigation costs, excluding water expenditures which are transfers) between the quota and market allocations using the estimated structural model.&lt;/p&gt;
&lt;p&gt;Q: Under what conditions is a quota more efficient than a market with liquidity constraints?
A: Quotas dominate markets when three conditions hold simultaneously: (1) farmers are relatively homogeneous in productivity (so the market&amp;rsquo;s advantage of directing water to high-valuation buyers is small), (2) liquidity constraints are significant (so the market misallocates water away from constrained high-valuation farmers), and (3) the production function is concave in water (so uniform allocation is efficient when farmers are homogeneous). The authors find all three conditions hold in Mula. Conversely, markets dominate quotas when heterogeneity in productivity is large relative to heterogeneity in wealth.&lt;/p&gt;
&lt;p&gt;Q: How is the transformation rate parameter gamma estimated and interpreted?
A: The transformation rate gamma measures how soil moisture above the permanent wilting point converts into apricot output (in pesetas) during the critical season, via the production function h() = gamma * (M_it - PW) * KS(M_it) * Z(w_t). It is identified from variation in purchasing patterns across seasons and variation in moisture across farmers within the same season. The preferred specification (column 3 of Table 3) yields gamma_L = 0.05. With average moisture per tree (accounting for the hydric stress coefficient) of 873.93 during the critical season, a farmer earns on average 29.09 pesetas per tree per week during the critical season, or 407.25 pesetas per tree per year.&lt;/p&gt;
&lt;p&gt;Q: How does ignoring liquidity constraints bias demand estimates?
A: If one estimates demand using the full sample (poor and wealthy farmers pooled), a decrease in demand during the critical season when prices rise conflates two effects: (1) the standard price effect (fewer farmers have valuations above the price) and (2) the liquidity constraint effect (some farmers with valuations above the price still cannot buy because they lack cash). Attributing the second effect to price sensitivity overstates the demand elasticity, biasing its absolute value upward.&lt;/p&gt;
&lt;p&gt;Q: What robustness checks do the authors provide against unobserved heterogeneity?
A: The authors provide four pieces of evidence that wealthy and poor farmers do not have systematically different underlying preferences: (1) wealthy and poor farmers are not geographically sorted into different locations (both groups appear in subareas 1, 2, 4, and 7); (2) wealthy and poor farmers grow the same bulida apricot variety; (3) outside the critical season, wealthy and poor farmers purchase statistically similar amounts of water; and (4) the purchasing divergence is significant only during the critical season when prices are high, precisely the pattern predicted by the liquidity constraint mechanism.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications for water allocation in developing countries?
A: The paper implies that before introducing water markets in regions where farmers may be liquidity constrained, policymakers should assess the magnitude of those constraints. If liquidity constraints are significant and farmers are relatively homogeneous in productivity, a quota system or a market supplemented with credit provision may deliver higher efficiency than a pure market. The standard presumption that markets outperform quotas can reverse when poor farmers cannot access credit to purchase water at the times they most need it.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to Che et al. (2013)?
A: Che, Gale, and Kim (2013) assume agents consume at most one unit with linear utility and find that markets always dominate quotas, though some non-market mechanisms with resale outperform markets. Donna and Espín-Sánchez extend this framework by allowing multiple discrete units, a concave utility function, and intertemporal dynamics. Under these extensions, the efficiency ranking between markets and quotas is theoretically indeterminate, and the authors show empirically that quotas can dominate markets. Both papers agree that non-market mechanisms with resale outperform both markets and simple quotas.&lt;/p&gt;
&lt;p&gt;Liquidity constraint (paper&amp;rsquo;s sense): A farmer is liquidity constrained when they lack sufficient cash to purchase water at the prevailing auction price, even if their valuation (marginal productivity of water) exceeds that price. In Mula, poor farmers without urban real estate income faced this constraint during the critical season when prices peaked, because they had already spent their harvest proceeds from the prior year and lacked access to credit markets.&lt;/p&gt;
&lt;p&gt;Soil moisture (M_it): The state variable measuring water accumulated in a farmer&amp;rsquo;s plot, computed using the agricultural engineering evapotranspiration formula. Moisture increases with rainfall and irrigation purchases (each auction unit contributing 432,000 liters divided by plot area) and decreases via evapotranspiration. It is bounded below by the permanent wilting point (PW) — below which trees die — and above by field capacity (FC). Moisture creates intertemporal substitution in demand.&lt;/p&gt;
&lt;p&gt;Critical season: The period corresponding to apricot fruit growth stages II and III and the Early Post-Harvest (EPH) period, spanning approximately weeks 18–32 (early May to early August). This is when the bulida apricot tree transforms water into fruit at the most rapid rate, when water demand peaks biologically, and when auction prices rise to their highest levels. It is the season during which liquidity constraints are binding.&lt;/p&gt;
&lt;p&gt;Transformation rate (gamma): The parameter in the apricot production function that measures the rate at which excess soil moisture (above the permanent wilting point) converts into apricot output (measured in real pesetas) during the critical season. Estimated at gamma_L = 0.05 in the preferred specification (column 3). It is identified from cross-seasonal variation in purchasing patterns and cross-farmer variation in moisture levels.&lt;/p&gt;
&lt;p&gt;Inefficiency due to decreasing marginal returns (DMR): One of two sources of allocation inefficiency identified in the paper. It arises when a farmer with already-high soil moisture receives water, yielding less additional output than if that water had gone to a farmer with lower moisture, given the concavity of the production function. Quotas avoid this inefficiency by allocating uniformly; markets with liquidity constraints exacerbate it by directing critical-season water to wealthy farmers who may have accumulated moisture from earlier purchases.&lt;/p&gt;
&lt;p&gt;Cuarta (quarter): The unit of water sold at Mula auctions, representing the right to use water flowing through the main channel for three hours. At approximately 40 liters per second of flow, each cuarta carried approximately 432,000 liters of water. Water rights and land rights were held independently; farmers who participated in auctions owned only land, while waterlords separately owned canal usage rights.&lt;/p&gt;
&lt;p&gt;Conditional choice probability (CCP) estimator: The two-step estimation procedure used to recover demand parameters from wealthy farmers&amp;rsquo; purchasing choices. In Step 1, transition probability matrices for observable state variables (moisture, week, price, rainfall) are computed and CCP is estimated via multinomial logit. In Step 2, the value function is forward-simulated using these transition matrices and parameters are estimated by GMM, following Hotz et al. (1994).&lt;/p&gt;</description></item><item><title>The Impact of Incarceration on Employment, Earnings, and Tax Filing</title><link>https://macropaperwarehouse.com/papers/the-impact-of-incarceration-on-employment-earnings-and-tax-filing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-impact-of-incarceration-on-employment-earnings-and-tax-filing/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;This paper estimates the causal effect of incarceration on employment, wage earnings, self-employment, and tax filing behavior using administrative criminal justice data linked to Internal Revenue Service (IRS) records for approximately half a million felony defendants in two U.S. states: North Carolina and Ohio. The study period covers cases filed from the early 2000s through 2014, with outcomes tracked through 2020 using IRS W-2 and 1040 records.&lt;/p&gt;
&lt;h3 id="research-question"&gt;Research Question&lt;/h3&gt;
&lt;p&gt;The central question is whether incarceration itself — as distinct from arrest, conviction, and other criminal justice interactions that precede or accompany it — causes lasting reductions in defendants&amp;rsquo; labor market outcomes. The paper explicitly holds fixed upstream interactions (conviction, arrest) to isolate the effect of the incarceration sentence.&lt;/p&gt;
&lt;h3 id="data-and-sample"&gt;Data and Sample&lt;/h3&gt;
&lt;p&gt;Criminal justice records from Ohio (Common Pleas courts in Franklin, Cuyahoga, and Hamilton counties, covering Columbus, Cleveland, and Cincinnati) and North Carolina (Administrative Office of the Courts and Department of Public Safety) are linked to de-identified IRS records via name, date of birth, sex, address, and partial Social Security Numbers. Match rates are 92% in Ohio and 95% in North Carolina. The sample is restricted to defendants aged 18–50 at time of offense with cases filed 2002–2014. IRS records include employer-reported W-2 wages (regardless of individual tax filing), self-employment income from Schedule C/SE, non-employee compensation (1099-MISC), and gig-economy earnings from 1099 returns. All dollar figures are adjusted to 2016 dollars using the PCE deflator.&lt;/p&gt;
&lt;h3 id="empirical-strategy"&gt;Empirical Strategy&lt;/h3&gt;
&lt;p&gt;Two independent quasi-experimental research designs are used:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;North Carolina — Sentencing guideline discontinuities&lt;/strong&gt;: North Carolina&amp;rsquo;s structured sentencing guidelines map offense class (E through I, the five least severe felony classes) and prior record points (a numerical criminal history score) into permissible punishment types (incarceration vs. probation) and sentence lengths. Allowable punishment types change discretely at five cell boundaries, generating discontinuities in incarceration sentences for otherwise similar defendants. The paper uses these five boundary discontinuities as excluded instruments in a parameterized regression discontinuity design stacked across offense classes. First-stage F-statistic = 115.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Ohio — Random assignment to judges&lt;/strong&gt;: Cases are randomly assigned by computer to judges at arraignment in the three counties studied. Judge leave-out mean sentence length is used as an instrument for individual sentence length. The design follows Norris et al. (2021) and yields F-statistic = 321. The instrument shifts sentences along both the extensive margin (any vs. no incarceration) and intensive margin (longer vs. shorter sentences).&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Both designs produce complier populations for whom at least 37–45% are shifted on the extensive margin (from no incarceration to some incarceration), based on partial identification bounds using linear programming.&lt;/p&gt;
&lt;h3 id="main-findings"&gt;Main Findings&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s central finding is that incarceration generates &lt;strong&gt;large short-run reductions&lt;/strong&gt; in labor market activity during the incapacitation period, but &lt;strong&gt;no detectable long-run reductions&lt;/strong&gt; in annual employment or earnings once defendants have been released and the incapacitation effects have dissipated.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;In the first year after case filing, when incarceration rates peak (roughly 75–100 additional days incarcerated for a 12-month sentence), employment falls by approximately &lt;strong&gt;10 percentage points&lt;/strong&gt; and total W-2 earnings contract commensurately.&lt;/li&gt;
&lt;li&gt;Within 3–4 years of filing, employment effects return to near zero and are statistically insignificant in both states.&lt;/li&gt;
&lt;li&gt;Five to nine years after filing, when effects on contemporaneous incarceration have dissipated, the estimated effect of a 12-month sentence on annual earnings is &lt;strong&gt;positive or near zero&lt;/strong&gt; in both states. The combined 95% confidence interval rules out reductions in annual wages greater than &lt;strong&gt;$231&lt;/strong&gt; (approximately 5% of the untreated complier mean) and rules out any adverse employment effects.&lt;/li&gt;
&lt;li&gt;Despite no long-run level effects, losses during incapacitation are never recouped. A one-year sentence reduces &lt;strong&gt;cumulative earnings over five years by approximately $2,914&lt;/strong&gt; — a 13% reduction relative to the complier mean.&lt;/li&gt;
&lt;li&gt;Effects on self-employment, independent contracting, 1040 filing, adjusted gross income, EITC take-up, and interstate migration are similarly null in the long run.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="incapacitation-vs-post-release-scarring"&gt;Incapacitation vs. Post-Release Scarring&lt;/h3&gt;
&lt;p&gt;The paper provides two tests for whether short-run earnings losses reflect incapacitation alone or also post-release scarring (e.g., human capital depreciation, employer discrimination, or discouragement effects):&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;A &amp;ldquo;visual IV&amp;rdquo; regression of year-t earnings effects on year-t days-incarcerated effects yields an R² of 0.83–0.85 across states, with the intercept near zero (positive and small), indicating that virtually all dynamic earnings impacts flow through contemporaneous incapacitation and not through a post-release channel.&lt;/li&gt;
&lt;li&gt;Constructed outcomes that impose the null of pure incapacitation (scaling pre-case average earnings or covariate-predicted earnings by the share of the year free from prison) closely track actual earnings effects in both states, further confirming that incapacitation is the dominant mechanism.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="pre-existing-labor-market-detachment"&gt;Pre-existing Labor Market Detachment&lt;/h3&gt;
&lt;p&gt;A key scope condition is defendants&amp;rsquo; severe labor market disadvantage prior to their case. Fewer than 50–60% of defendants are employed in the year before filing; average pre-case W-2 earnings (including zeros) are below $6,000. Among employed defendants, only 10% earn more than $22,000 per year. Untreated complier means for earnings in the year after case filing are below $4,000, with virtually no earnings or employment growth over the following nine years. The paper concludes that returning to pre-filing earnings levels is sufficient for incarcerated defendants to match their non-incarcerated peers — a low bar that is readily met.&lt;/p&gt;
&lt;h3 id="policy-implications"&gt;Policy Implications&lt;/h3&gt;
&lt;p&gt;Back-of-envelope aggregation implies incapacitation losses of approximately &lt;strong&gt;$6.16 billion per year&lt;/strong&gt; in foregone earnings for the U.S. prison population, concentrated in communities heavily affected by incarceration. However, a marginal reduction in incarceration rates would increase average earnings by only &lt;strong&gt;$51 for white men&lt;/strong&gt; and &lt;strong&gt;$213 for black men&lt;/strong&gt;, suggesting incarceration&amp;rsquo;s direct contribution to labor market inequality is modest relative to the $21,100 black-white earnings gap estimated by Bayer and Charles (2018). The paper concludes that upstream factors — other criminal justice interactions, human capital deficits, and broader socioeconomic disadvantage — are more plausibly responsible for low earnings among the formerly incarcerated.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-exact-treatment-variable-and-what-is-the-counterfactual"&gt;Q1. What is the exact treatment variable, and what is the counterfactual?&lt;/h3&gt;
&lt;p&gt;The treatment variable is months of incarceration sentenced in the focal case (a continuous, weakly positive ordered treatment). The counterfactual for non-incarcerated defendants in North Carolina is probation (all defendants are convicted by construction under structured sentencing guidelines). In Ohio, the authors cannot reject that all compliers who do not receive a prison sentence are still convicted, implying the counterfactual is also conviction and probation. All compliers therefore acquire a criminal record regardless of sentence. The treatment effect is thus the effect of incarceration conditional on conviction, holding fixed the criminal record.&lt;/p&gt;
&lt;h3 id="q2-how-are-effects-interpreted-given-multiple-instruments-and-continuous-treatment"&gt;Q2. How are effects interpreted given multiple instruments and continuous treatment?&lt;/h3&gt;
&lt;p&gt;Under a &amp;ldquo;weakly positive ordered treatment&amp;rdquo; assumption and standard LATE conditions, the 2SLS estimates can be interpreted as Average Causal Responses (ACRs) — weighted averages of the marginal dose effects (12 vs. 11 months, 6 vs. 5 months, 1 vs. 0 months, etc.) for complier subgroups shifted by each instrument. In North Carolina with five parameterized RD instruments, the estimate averages ACRs weighted by first-stage strength. In Ohio with a leave-out mean instrument, the estimate is a convex average of ACRs under the assumption that the linear first-stage model is a good approximation. Dosage weights for both states put mass on a wide range of sentence lengths including both extensive and intensive margins, though Ohio&amp;rsquo;s weights are more skewed toward shorter sentences.&lt;/p&gt;
&lt;h3 id="q3-how-large-are-the-first-stage-effects-and-how-strong-is-the-instrument"&gt;Q3. How large are the first-stage effects, and how strong is the instrument?&lt;/h3&gt;
&lt;p&gt;In North Carolina, sentences jump by 50% or more at sentencing guideline cell boundaries where allowable punishment types change to include incarceration. The first-stage F-statistic is 115. In Ohio, defendants assigned to the most severe judge receive incarceration sentences approximately six months longer than those assigned to the least severe judge (roughly 30% of the average non-zero sentence), with a slope of approximately 0.8 in the first-stage regression; F-statistic = 321. At least 37% of compliers in North Carolina and 45% in Ohio are shifted on the extensive margin (from no incarceration to some positive incarceration), with upper bounds as high as 95%.&lt;/p&gt;
&lt;h3 id="q4-what-evidence-supports-instrument-validity-exclusion-restriction-and-independence"&gt;Q4. What evidence supports instrument validity (exclusion restriction and independence)?&lt;/h3&gt;
&lt;p&gt;Instrument validity is tested by estimating 2SLS &amp;ldquo;effects&amp;rdquo; on pre-case outcomes measured 2–4 years before the focal case. In both states, the instruments show no relationship with pre-case employment, W-2 wages, total days previously incarcerated, or binary severe prior incarceration. The probability of being matched to IRS records and the quality of the match are also uncorrelated with the instruments. In Ohio, potential exclusion restriction violations from judges affecting conviction (not just sentence) are addressed empirically: nearly 90% of defendants are convicted, the most severe judge is only 0.7 p.p. more likely to convict than the least severe judge (t-stat = 1.53), and the estimated conviction rate among untreated compliers is 0.972 (s.e. 0.018), so one cannot reject that all non-incarcerated compliers are convicted.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-paper-test-for-the-incapacitation-mechanism-against-post-release-scarring"&gt;Q5. How does the paper test for the incapacitation mechanism against post-release scarring?&lt;/h3&gt;
&lt;p&gt;Two complementary exercises are conducted. First, a &amp;ldquo;visual IV&amp;rdquo; plot regresses year-t earnings effects on year-t days-incarcerated effects across all post-filing years. If incapacitation is the sole channel, all points should lie on a line through the origin. The R² is 0.83 in North Carolina and 0.85 in Ohio, the estimated intercept is near zero (positive and small) in both states, and the slope (earnings lost per day incarcerated) is approximately $12. This implies cumulative earnings losses of $12 × 268 days = $3,216, very close to the directly estimated $2,914. Second, constructed outcomes that scale pre-case earnings or covariate-predicted earnings by the share of the year not incarcerated closely track actual earnings effects throughout the post-filing period, and both converge to zero as incapacitation effects fade — consistent with pure incapacitation and no net scarring.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-long-run-59-years-earnings-and-employment-estimates-and-how-precisely-are-null-effects-ruled-out"&gt;Q6. What are the long-run (5–9 years) earnings and employment estimates, and how precisely are null effects ruled out?&lt;/h3&gt;
&lt;p&gt;Averaged across both states using inverse-variance weights, the estimated effect of a 12-month sentence on annual W-2 earnings five to nine years after filing is positive but statistically indistinguishable from zero. The 95% confidence interval rules out reductions in annual wages greater than $231 (approximately 5% of the untreated complier mean of roughly $4,500–$5,000). The 95% CI also rules out any adverse employment effects. The untreated complier mean for employment 5–9 years post-filing is approximately 40% in North Carolina and slightly above 40% in Ohio.&lt;/p&gt;
&lt;h3 id="q7-what-happens-to-cumulative-earnings-over-five-years-despite-null-long-run-level-effects"&gt;Q7. What happens to cumulative earnings over five years despite null long-run level effects?&lt;/h3&gt;
&lt;p&gt;Even though long-run annual earnings are unaffected, earnings losses during incapacitation are never made up. A one-year sentence reduces cumulative employment (measured as years with any W-2) and cumulative earnings over five years by approximately $2,914 — a 13% reduction relative to the complier mean. This reflects the mechanical loss of earnings during the period of physical incapacitation, without a subsequent compensating period of higher earnings after release.&lt;/p&gt;
&lt;h3 id="q8-do-defendants-with-stronger-pre-case-labor-market-attachment-show-different-long-run-patterns"&gt;Q8. Do defendants with stronger pre-case labor market attachment show different long-run patterns?&lt;/h3&gt;
&lt;p&gt;The sample is split between defendants employed in at least 2 of the 4 years prior to the case (53–57% of the sample across states) and those less attached. Both groups show zero long-run earnings and employment effects. Previously employed defendants experience much larger short-run earnings drops — more than three times larger in the first year post-filing — and their earnings recover more slowly, reaching zero effect approximately six years after filing (vs. three years for the previously unemployed). For a stricter cut (pre-case average earnings above $15,000, representing only 12–15% of the sample), the long-run earnings effect is −$1,426 (8% of the untreated complier mean), significant only at the 10% level, and partly attributable to residual incapacitation (19.6 additional days incarcerated 5–9 years post-filing). For defendants with pre-case earnings below $15,000, incarceration slightly increases long-run employment (2.4 pp, p = 0.01) and earnings ($400, p = 0.03), possibly reflecting rehabilitative benefits (GED or educational programs) for labor-market-detached individuals.&lt;/p&gt;
&lt;h3 id="q9-does-first-time-incarceration-extensive-margin-exposure-have-larger-long-run-effects-than-repeat-exposure"&gt;Q9. Does first-time incarceration (extensive-margin exposure) have larger long-run effects than repeat exposure?&lt;/h3&gt;
&lt;p&gt;The paper tests this by splitting the sample into defendants with and without prior incarceration history. Among defendants with no prior incarceration, the instruments generate large differences in lifetime exposure: a 12-month sentence increases the probability of ever being incarcerated over the next 5–9 years by 26 p.p. (North Carolina) and 41 p.p. (Ohio). Among those not receiving a sentence, 48% (North Carolina) and 19% (Ohio) are eventually incarcerated anyway, implying treatment causes a 52 and 81 p.p. increase in lifetime incarceration probability for extensive-margin compliers. Despite these large differences in lifetime exposure, long-run earnings and employment effects remain small and statistically insignificant in both subsamples. The difference in long-run effects between previously and never incarcerated defendants is not statistically significant (p = 0.29 for employment, p = 0.82 for earnings).&lt;/p&gt;
&lt;h3 id="q10-are-there-heterogeneous-effects-by-race-sex-or-criminal-history"&gt;Q10. Are there heterogeneous effects by race, sex, or criminal history?&lt;/h3&gt;
&lt;p&gt;There is no evidence of long-run scarring for any demographic or criminal history subgroup. Effects for black and non-black defendants are both positive for long-run earnings and employment. Non-black defendants show somewhat larger cumulative losses (consistent with marginally higher counterfactual earnings), but differences are not statistically significant. Estimates for women are imprecise due to small sample size. Among defendants with and without prior felony charges in the four years preceding the case, there are neither economically nor statistically significant long-run earnings or employment effects. Cumulative losses are somewhat larger for defendants without prior felony charges (p = 0.07), reflecting their higher pre-case earnings.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-handle-potential-migration-bias-in-outcomes"&gt;Q11. How does the paper handle potential migration bias in outcomes?&lt;/h3&gt;
&lt;p&gt;Tax filing and W-2 receipt in the state of sentencing are used to proxy for whether defendants remain in the same state. Among untreated compliers, 88% of those with a tax footprint maintain it in the state of sentencing. No statistically significant effects of incarceration on migration (measured as filing or receiving a W-2 in North Carolina or Ohio) are detected, suggesting prior studies of recidivism measured within-state are unlikely to be severely biased by migration responses.&lt;/p&gt;
&lt;h3 id="q12-what-effect-does-incarceration-have-on-mortality"&gt;Q12. What effect does incarceration have on mortality?&lt;/h3&gt;
&lt;p&gt;Incarceration reduces five-year mortality by approximately 0.8 percentage points (about 20% of the untreated mean). The authors note this is too small to explain the null long-run labor market effects: even if all defendants whose death was averted were employed, removing them from the employment count would reduce the employment effect of a 12-month sentence only to approximately zero.&lt;/p&gt;
&lt;h3 id="q13-how-do-the-papers-findings-compare-to-prior-studies-particularly-mueller-smith-2015"&gt;Q13. How do the paper&amp;rsquo;s findings compare to prior studies, particularly Mueller-Smith (2015)?&lt;/h3&gt;
&lt;p&gt;Mueller-Smith (2015) finds large and persistent negative incarceration effects on labor market outcomes in Texas using a structural decomposition and Lasso-based judge-covariate interactions as instruments. The paper argues methodological differences are the likely explanation: the Lasso-selected interacted instruments can be susceptible to many-weak instruments bias toward OLS. It notes that Mueller-Smith&amp;rsquo;s simpler 2SLS specifications (analogous to those used here) show no statistically significant earnings effects. North Carolina and Ohio are documented to be broadly similar to Texas (and the U.S. average) in rehabilitation program participation, recidivism rates, and incarceration rates, reducing the likelihood that genuine geographic heterogeneity explains the divergence.&lt;/p&gt;
&lt;h3 id="q14-what-is-the-papers-aggregate-extrapolation-of-incapacitation-earnings-losses"&gt;Q14. What is the paper&amp;rsquo;s aggregate extrapolation of incapacitation earnings losses?&lt;/h3&gt;
&lt;p&gt;Scaling the estimated $2,914 cumulative loss per 12-month sentence by the ratio of days exposed to total days in a year gives a per-day loss of approximately $12. Applied to the 1,435,500 people incarcerated in U.S. prisons on any given day in 2019 (excluding the more than 700,000 in jail), the implied aggregate yearly earnings loss from incapacitation is approximately $6.16 billion.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Incapacitation effect&lt;/strong&gt;: The mechanical reduction in earnings and employment that occurs while a defendant is physically confined in prison and unable to work, as distinct from any post-release scarring effect. The paper shows this is the dominant — and essentially sole — causal channel through which incarceration affects labor market outcomes in their sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Post-release scarring&lt;/strong&gt;: Persistent reductions in earnings or employment that persist after a defendant is released from prison, caused by mechanisms such as employer discrimination based on incarceration history, human capital depreciation, loss of job contacts, or psychological discouragement effects. The paper finds no evidence of scarring in either state.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Average Causal Response (ACR)&lt;/strong&gt;: The weighted average of the marginal dose effects of incarceration (e.g., effect of 12 vs. 11 months, 1 vs. 0 months) for groups of defendants whose sentence lengths are shifted by a given instrument. Contrasted with a binary LATE, the ACR averages across the full dosage distribution for compliers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Complier&lt;/strong&gt;: An individual whose incarceration sentence is shifted by the instrument — either from zero to some positive sentence (extensive margin) or from a shorter to a longer sentence (intensive margin). Counterfactual outcome means for compliers sentenced to zero months provide the baseline for evaluating effect magnitudes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sentencing guideline discontinuity&lt;/strong&gt;: The discrete jump in permissible punishment types and minimum sentence lengths at specific criminal history score thresholds within North Carolina&amp;rsquo;s structured sentencing grid. Defendants just above a threshold are more likely to be incarcerated than otherwise similar defendants just below, generating quasi-experimental variation exploited as an instrument.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Leave-out mean judge instrument&lt;/strong&gt;: In Ohio, each defendant&amp;rsquo;s assigned judge&amp;rsquo;s average incarceration sentence length computed over all other cases that judge handles (excluding the defendant&amp;rsquo;s own case), residualized on court-by-month fixed effects. Because judges are randomly assigned to cases, this measure is conditionally independent of defendant potential outcomes and serves as an instrument for sentence length.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Control complier mean&lt;/strong&gt;: The estimated mean potential outcome for compliers under the counterfactual of receiving zero months of incarceration. Used as a benchmark to evaluate the magnitude of treatment effects and to characterize how low the earnings baseline is for the population driving the causal estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extensive vs. intensive margin of incarceration&lt;/strong&gt;: The extensive margin refers to the binary shift from receiving no prison sentence to receiving any prison sentence; the intensive margin refers to increasing sentence length conditional on some incarceration. The paper argues that neither margin appears to produce long-run labor market scarring, and uses linear programming bounds to estimate that at least 37–45% of compliers in each state are shifted on the extensive margin.&lt;/p&gt;</description></item><item><title>The Impact of Unions on Nonunion Wage Setting: Threats and Bargaining</title><link>https://macropaperwarehouse.com/papers/the-impact-of-unions-on-nonunion-wage-setting-threats-and-bargaining/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-impact-of-unions-on-nonunion-wage-setting-threats-and-bargaining/</guid><description>&lt;p&gt;This paper estimates the impact of unions on nonunion wage setting in the United States over the period 1980–2010, distinguishing two channels through which unions affect nonunion wages: (1) a traditional threat channel, in which nonunion firms raise wages to preempt unionization by making workers indifferent between forming a union and remaining nonunion (an &amp;ldquo;emulation wage&amp;rdquo;); and (2) a bargaining channel, in which nonunion workers use the availability of high-paying union jobs as part of their outside option when bargaining individually with their employer, so that a decline in union job prevalence or the union wage premium erodes nonunion bargained wages even at firms that face no direct unionization threat.&lt;/p&gt;
&lt;p&gt;The authors build a search-and-bargaining model grounded in Nash bargaining, with endogenous union formation and, in the most complete version, the possibility of nonunion firm responses to the threat of unionization. Workers in this model can be employed at simple nonunion firms, union firms, or union-emulating firms. The model is embedded in a multi-industry, multi-city framework following Beaudry, Green, and Sand (2012), which formalizes the mechanism by which higher-rent jobs in a city raise outside options and therefore wages for workers in all other jobs throughout that local labor market. This cross-city, within-industry variation is the primary source of identification.&lt;/p&gt;
&lt;p&gt;The empirical implementation uses Current Population Survey Merged Outgoing Rotation Groups (1983–2020) and CPS May extracts (1978–1982), pooling observations around 1980, 1990, 2000, 2010, and 2020 across 43 cities and 51 industries. To address endogeneity of outside option variables — which may be correlated with unobserved local productivity shocks — the authors construct Bartik-style instruments based on start-of-period local industry and union employment composition interacted with national changes in industry growth, industry wage premia, and union job transition probabilities. The threat channel is identified by the interaction of the probability a firm in a given industry-city cell faces a union election (proxied using NLRB data) with the outside option value of union workers. The authors derive a model-based overidentifying restriction, test it, and cannot reject it, providing support for their identification strategy.&lt;/p&gt;
&lt;p&gt;The central quantitative finding is that de-unionization accounts for approximately 38% of the 16% decline in the mean real (composition-constant) wage in a typical US city between 1980 and 2010. One-third of that de-unionization effect arises from a standard shift-share component — workers moving from higher-paying union jobs to lower-paying nonunion jobs — while two-thirds arises from spillover channels affecting nonunion wage setting. The spillover effects are almost entirely attributable to the bargaining channel rather than the traditional threat channel; the threat probability was too low, even in 1980, to generate large emulation effects in the aggregate. The total impact of a one-dollar increase in the outside option value for the mean wage in industry i is estimated at 1.78 dollars once within-industry feedback loops are included.&lt;/p&gt;
&lt;p&gt;The paper finds no evidence of bargaining spillovers in the 1980s specifically, the decade of the sharpest unionization declines. The offsetting forces were declining probabilities of finding union jobs and simultaneously rising union wage premia — with the model explaining the premium increase as a consequence of nonunion firms no longer needing to emulate union wages once the threat of their shop being organized receded substantially. After 1990 the threat stabilized at a low level, the premium declined, and the outside-option effect of declining unionization became the dominant force.&lt;/p&gt;
&lt;p&gt;Heterogeneity results show that spillover effects are larger for women than men, and that de-unionization accounts for 43% of the real wage decline for women versus 27% for men. For workers without post-secondary education, de-unionization accounts for 43% of their real wage decline. The traditional threat effect is statistically insignificant in states with Right-to-Work laws, consistent with the interpretation that identification captures emulation responses to unionization threat.&lt;/p&gt;
&lt;p&gt;Q: What are the two channels through which unions affect nonunion wages in this model?
A: The traditional threat channel operates when nonunion firms raise wages to make workers indifferent between unionizing and remaining nonunion, thereby forestalling a costly union election. The bargaining channel operates because nonunion workers can credibly point to available union jobs when bargaining individually; a decline in union job prevalence or the union wage premium therefore weakens nonunion workers&amp;rsquo; outside options and lowers their bargained wages even at firms that face no direct unionization threat.&lt;/p&gt;
&lt;p&gt;Q: How large is the overall contribution of de-unionization to the US wage decline between 1980 and 2010?
A: The paper estimates that de-unionization accounts for 38% of the approximately 16% decline in the mean composition-constant real wage in a typical US city between 1980 and 2010. One-third of that 38% arises from the direct shift-share effect of workers moving from higher-paying union to lower-paying nonunion employment; the remaining two-thirds arises from spillover effects on nonunion wages.&lt;/p&gt;
&lt;p&gt;Q: Which spillover channel dominates in the decomposition, and why?
A: The bargaining channel dominates almost entirely. The traditional threat channel is statistically significant but quantitatively small because the probability that any given nonunion firm faced a union election was low even in 1980, so the scope for emulation to affect aggregate wages was limited. The bargaining channel, by contrast, operates through the outside options of all nonunion workers searching across many industries and cities, giving it broader aggregate reach.&lt;/p&gt;
&lt;p&gt;Q: Why was there no measurable bargaining spillover in the 1980s despite the decade&amp;rsquo;s large drop in union density?
A: During the 1980s, two forces offset each other: the probability of a nonunion worker finding a union job fell sharply, but the union wage premium rose substantially over the same period, so the expected value of the union outside option changed little. The paper explains the rising premium as a consequence of nonunion firms reducing their emulation wages as the threat of unionization receded, causing nonunion wages to fall faster than union wages and thus mechanically widening the premium. After 1990, when the threat stabilized at a low level, the premium declined and the net outside-option effect of continued de-unionization became the dominant spillover force.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated multiplier effect of an improvement in outside options on nonunion wages?
A: The total impact of a one-dollar increase in the outside option value on the mean wage in a given industry is estimated at 1.78 dollars once within-industry feedback loops — in which an improved outside option raises wages, which in turn improves outside options elsewhere — are accounted for.&lt;/p&gt;
&lt;p&gt;Q: How do the authors address endogeneity of the outside option variables?
A: They construct Bartik-style instruments based on start-of-period local industry and union employment composition interacted with national-level changes in industry growth, industry wage premia, and the probability of transitioning to a union job. This strategy isolates variation in local outside options that is driven by predetermined compositional exposure rather than contemporaneous local shocks. They derive a model-based overidentifying restriction, test it in the data, and cannot reject it, supporting the validity of the instrument.&lt;/p&gt;
&lt;p&gt;Q: How do the authors address selection bias arising from the changing composition of union and nonunion workers as unionization declines?
A: They implement a generalized Heckman two-step approach, including a quartic in the change in the proportion unionized to control for selectivity. After this correction, they cannot reject the null of no selectivity effects, and the main estimated coefficients change very little, indicating that compositional selection is not the primary driver of their results.&lt;/p&gt;
&lt;p&gt;Q: What heterogeneity is found across gender groups?
A: Both the bargaining and traditional threat effects are larger for women than for men. Men experienced a decline in mean real wages between 1980 and 2010 more than double that experienced by women, but spillover effects are of identical size, so de-unionization accounts for a larger share of women&amp;rsquo;s wage decline (43%) than men&amp;rsquo;s (27%).&lt;/p&gt;
&lt;p&gt;Q: What heterogeneity is found by education level?
A: For workers with a high school education or less, the traditional threat effect estimate is twice as large as the bargaining effect, while the reverse holds for workers with post-secondary education. Workers without post-secondary education experienced real wage declines nearly triple those of the more educated group, and de-unionization accounts for 43% of the lower-educated group&amp;rsquo;s wage decline.&lt;/p&gt;
&lt;p&gt;Q: How do the authors validate that they are identifying the threat channel rather than some other effect?
A: The traditional threat effect is estimated to be statistically insignificant in states with Right-to-Work (RTW) laws, where the legal environment substantially reduces the ability of workers to organize and therefore reduces the credible threat of unionization that would induce nonunion firms to emulate union wages. This pattern is consistent with the interpretation that the identified effect captures firm emulation responses to a genuine unionization threat.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the distinction between the two channels?
A: The traditional threat effect can only be activated by increasing union power directly, since it depends on a credible risk of a firm&amp;rsquo;s workforce voting to unionize. The bargaining channel, however, is not union-specific: any policy that raises workers&amp;rsquo; outside option values — such as eliminating non-compete agreements or expanding access to higher-paying jobs in a local labor market — can generate similar wage spillovers. Unions are one powerful mechanism for doing this, but not the only one.&lt;/p&gt;
&lt;p&gt;Q: What is the theoretical model structure, and what distinguishes it from Taschereau-Dumouchel (2020)?
A: The model is built on TD&amp;rsquo;s search-and-bargaining framework with endogenous union formation, in which unions can threaten to withdraw the entire workforce from production whereas individual nonunion workers can only threaten to withdraw their own labor. The key modifications are: (1) the hiring-channel mechanism of TD (firms skew toward skilled workers who dislike unions) is replaced with a direct wage-emulation mechanism; (2) the BGS multi-industry, multi-city framework is incorporated to allow outside options to vary with the composition of jobs across industries in a locality; and (3) a single skill level with multiple industries is used, keeping the model tractable for empirical implementation.&lt;/p&gt;
&lt;p&gt;Q: What data sources are used and over what period?
A: The primary dataset is the Current Population Survey Merged Outgoing Rotation Groups for 1983–2020 combined with CPS May extracts for 1978–1982, covering workers aged 25–65 not enrolled in school. The sample is organized into 93 geographic areas (43 cities), 51 industries based on 1980 Census classification, and analyzed at 10-year intervals (1980, 1990, 2000, 2010, 2020) with three-year pooling windows to reduce noise. NLRB case data on union elections proxies for unionization threat probabilities, and County Business Patterns data are used in constructing emulation probabilities.&lt;/p&gt;
&lt;p&gt;Traditional threat effect: The mechanism by which nonunion firms raise wages to an &amp;ldquo;emulation wage&amp;rdquo; — the level that makes workers indifferent between unionizing and remaining nonunion — in order to preempt the costs of a union election, thereby reducing the net benefit of unionization below the threshold required for workers to vote for a union.&lt;/p&gt;
&lt;p&gt;Bargaining channel (bargaining spillover effect): The mechanism by which the availability of union jobs in a local labor market raises the outside option of nonunion workers during individual Nash bargaining, so that declines in union job prevalence or the union wage premium lower nonunion bargained wages even at firms not directly facing a unionization threat.&lt;/p&gt;
&lt;p&gt;Outside option: In the model&amp;rsquo;s Nash bargaining framework, the value a worker (or firm) obtains if negotiations break down — for nonunion workers, this is the expected value of searching across both nonunion and union jobs weighted by transition probabilities and wage rents in each sector.&lt;/p&gt;
&lt;p&gt;Emulation wage: The wage a nonunion firm sets that is just high enough to make workers indifferent between unionizing and remaining nonunion, determined by the firm&amp;rsquo;s calculation of the threshold below which workers would prefer to bear the costs of unionization.&lt;/p&gt;
&lt;p&gt;Union formation (endogenous): In the model, unionization occurs when the surplus workers gain from collective bargaining exceeds the costs of organizing; firms can influence this calculus through wage emulation or direct anti-union actions, making union formation an equilibrium outcome rather than an exogenous event.&lt;/p&gt;
&lt;p&gt;Bartik-style instrument (outside option instrument): An instrument for local outside option values constructed by interacting start-of-period local employment composition across industries with national-level changes in industry growth, industry wage premia, and union job transition probabilities, isolating variation in outside options driven by predetermined exposure to national trends rather than local demand shocks.&lt;/p&gt;
&lt;p&gt;Shift-share (between) component: The portion of the aggregate wage effect of de-unionization attributable to the direct reallocation of workers from higher-paying union jobs to lower-paying nonunion jobs, distinct from spillover effects on nonunion wage setting itself.&lt;/p&gt;</description></item><item><title>The Optimal Taxation of Couples</title><link>https://macropaperwarehouse.com/papers/the-optimal-taxation-of-couples/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-optimal-taxation-of-couples/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; What is the optimal joint nonlinear earnings tax schedule for married couples? How should one spouse&amp;rsquo;s marginal tax rate depend on the other&amp;rsquo;s earnings? When is individual earnings-based (separable) taxation optimal versus family-income-based taxation, and what determines the sign and magnitude of &amp;ldquo;jointness&amp;rdquo; — the dependence of one spouse&amp;rsquo;s marginal tax on the other&amp;rsquo;s earnings?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The paper studies a canonical unitary household model in which each couple consists of two spouses who jointly maximize utility subject to a joint budget constraint. Spousal productivities are drawn from a joint distribution F with arbitrary dependence structure. The planner maximizes a weighted sum of couples&amp;rsquo; utilities, with Pareto weights that are decreasing functions of productivities. Utility takes a quasi-linear form in consumption and labor disutility with constant labor supply elasticity parameter γ (implying earnings elasticity γ/(γ-1)). The tax problem is equivalent to a two-dimensional mechanism design problem in which the planner chooses allocations as functions of reported productivity types, subject to incentive compatibility and budget feasibility. Because spousal productivities are two-dimensional, the problem is a multi-dimensional screening problem whose properties are poorly understood in general.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Methodology.&lt;/strong&gt; The authors proceed in two directions. First, they establish conditions under which the first-order approach (FOA) — restricting attention to local incentive constraints — is valid in this bi-dimensional setting. They show, for the special case of the benchmark economy (symmetric, independent types, separable Pareto weights), that FOA validity is equivalent to convexity of a certain transformation of the value function, and derive necessary and sufficient conditions that are strictly weaker than their unidimensional analogs — so the FOA is more likely to hold in two dimensions than in one. For the general economy, they invoke an Implicit Function Theorem argument in Hölder space to show that the FOA holds for Pareto weights sufficiently close to utilitarian (i.e., when the planner is not &amp;ldquo;too redistributive&amp;rdquo;). Second, assuming FOA validity, they characterize optimal taxes via a second-order nonlinear PDE. Since this PDE cannot be solved analytically in general, they apply the Coarea Formula to derive closed-form expressions for conditional averages of optimal tax distortions over various subsets of the type space, expressed entirely in terms of structural primitives (labor supply elasticities, Pareto weights, and elasticities of the joint distribution of productivities).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings.&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Average distortions and assortativeness.&lt;/strong&gt; Average optimal distortions on married individuals are ranked by the degree of positive quadrant dependence (PQD) in spousal productivities: more assortative matching implies higher optimal tax rates. Optimal distortions on married individuals are always weakly lower than on single individuals with the same productivity, same elasticities, and same marginal productivity distribution — strictly so unless matching is perfectly positively assortative. The intuition is that when couples pool resources, intra-family redistribution already occurs, and distortionary taxation crowds this out; more random matching produces more within-family redistribution, reducing the marginal social value of public redistribution through taxation.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Optimality of separable (individual earnings-based) taxation.&lt;/strong&gt; In the benchmark economy with independent types, optimal taxes are exactly separable (individual earnings-based), and optimal distortions on married individuals equal precisely one-half of those on comparable single individuals. With separable Pareto weights and independent types more generally, taxes remain separable. Once types are positively dependent, however, the planner optimally introduces jointness even under separable social weights.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Jointness and tail (in)dependence.&lt;/strong&gt; Optimal jointness — whether one spouse&amp;rsquo;s marginal tax rate increases or decreases in the other&amp;rsquo;s earnings — depends critically on tail dependence of the joint productivity distribution, captured by the copula and survival copula elasticities. For right-tail dependent distributions (so that extremely productive individuals are likely to be matched with extremely productive partners), positive jointness is optimal at the top (raising taxes on high earners whose partners are also high earners) and negative at the bottom. For right-tail independent distributions (such as the Gaussian copula, which is tail-independent for any finite ρ), the distortion-reducing motive dominates: optimal jointness is negative at the top and positive at the bottom, conditional on standard convergence conditions.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Primary vs. secondary earners.&lt;/strong&gt; The secondary earner (lower-productivity spouse) faces on average higher optimal distortions than the primary earner when the planner values redistribution to couples with a very unproductive spouse (α(w,0) ≥ 1), because the phasing out of transfers targeted to such couples generates high marginal tax rates on secondary earners. Family earnings-based taxation is optimal only when total family productivity and relative spousal productivity are independent, and when social weights are measurable only with respect to total family output.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Restricted taxation.&lt;/strong&gt; Optimal distortions under any of the three restricted tax regimes (anonymous, separable, family earnings-based) exactly equal the relevant conditional average of unrestricted optimal distortions. This establishes that the welfare difference between the restricted and unrestricted optimum stems solely from the planner&amp;rsquo;s inability to tag taxes to individual productivity types within the restricted class.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Findings (calibrated to 2020 CPS data on U.S. married couples, ages 25-65, worked ≥ 20 weeks).&lt;/strong&gt; Spousal productivities are positively but not perfectly dependent, with Kendall&amp;rsquo;s tau = 0.21 and Pearson correlation = 0.25 for productivities (0.21 for earnings). The joint distribution is well approximated by a Gaussian copula (ρ = 0.33) with Pareto-lognormal marginals (a = 2.95, Gini = 0.31). The Gaussian copula is tail-independent, so consistent with analytical results, optimal jointness is positive for low earners and negative for high earners (the latter arising at earnings above approximately $8.5 million in the benchmark specification). The quantitative magnitude of optimal jointness is small — marginal taxes for one spouse change by at most several percentage points as a function of the other spouse&amp;rsquo;s earnings. Individual earnings-based taxation provides a good approximation to the unrestricted optimum. By contrast, family earnings-based (joint) taxation is a poor approximation in all specifications, with marginal taxes on family income varying substantially with the earnings share of the secondary earner, and this conclusion holds even when Pareto weights explicitly favor family earnings-based taxation (k = 0 case). The implied top marginal tax rate converges toward approximately 55 percent (corresponding to limiting distortion of ≈1.35 = 1/γa with γ = 0.25, a = 2.95) but the convergence is slow, so optimal marginal rates remain substantially below this limit even at earnings of $300,000.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-mechanism-design-formulation-and-why-is-foa-validity-a-key-concern-in-the-bi-dimensional-setting"&gt;Q1. What is the mechanism design formulation, and why is FOA validity a key concern in the bi-dimensional setting?&lt;/h3&gt;
&lt;p&gt;A: The planner&amp;rsquo;s problem is cast as a direct mechanism in which couples report their two-dimensional productivity type (w1, w2) and receive allocations (consumption, earnings). Incentive compatibility requires that no couple prefers to misreport. In one-dimensional models (Mirrlees 1971), restricting attention to local incentive constraints (the FOA) yields the standard ODE characterization of optimal taxes and is valid for a broad class of primitives. In two dimensions, solutions to multi-dimensional screening problems generically display &amp;ldquo;bunching&amp;rdquo; (Rochet-Choné 1998, Armstrong 1996), and the FOA may fail. The key difference exploited in this paper is the absence of participation constraints in the public finance setting, which eliminates the main force driving FOA failure in industrial organization models.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-necessary-and-sufficient-conditions-for-foa-validity-in-the-benchmark-economy-with-independent-types"&gt;Q2. What are the necessary and sufficient conditions for FOA validity in the benchmark economy with independent types?&lt;/h3&gt;
&lt;p&gt;A: (Proposition 1) In the benchmark economy (symmetric, independent types, separable Pareto weights), FOA validity is equivalent to the condition that x·(1 + λ̃(x^{-γ})/2) is increasing in x, where λ̃(t) = [∫_t^∞ (1-α̃(w))g(w)dw] / (γtg(t)). The unidimensional analog requires x·(1 + λ̃(x^{-γ})) to be increasing. Since the bi-dimensional condition multiplies λ̃ by 1/2 rather than 1, the set of primitives satisfying it is strictly larger: every (G, α̃, γ) for which the unidimensional FOA holds also satisfies the bi-dimensional condition, but not vice versa. Economically, the FOA holds as long as the planner is not &amp;ldquo;too redistributive&amp;rdquo; — i.e., Pareto weights on low types are not so high as to violate these monotonicity conditions.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-coarea-formula-result-equation-27-and-why-is-it-the-central-technical-tool"&gt;Q3. What is the Coarea Formula result (equation 27) and why is it the central technical tool?&lt;/h3&gt;
&lt;p&gt;A: Given that the optimality conditions form a PDE system that cannot generally be solved pointwise, the authors integrate the optimality condition (equation 20) over subsets of the type space defined by level sets of an arbitrary function Q(w1, w2). The Coarea Formula allows them to express the result as: E[Σ_i λ*_i γ_i (∂lnQ/∂lnw_i) | Q=t] = [1 − E[α|Q≥t]] / [−∂ln P(Q≥t)/∂ln t]. By choosing different Q functions (e.g., Q = w_i, Q = max{k_1 w_1, k_2 w_2}, Q = R(w) for total family productivity, Q = I(w) for relative productivity), the formula delivers closed-form expressions for distinct conditional averages of optimal distortions, all expressed in terms of exogenous primitives. This contrasts with variational approaches (Golosov et al. 2014, Spiritus et al. 2022) that express optimal taxes in terms of endogenous moments.&lt;/p&gt;
&lt;h3 id="q4-how-do-optimal-distortions-on-married-individuals-compare-to-those-on-single-individuals-and-what-is-the-exact-quantitative-relationship-in-the-independent-types-benchmark"&gt;Q4. How do optimal distortions on married individuals compare to those on single individuals, and what is the exact quantitative relationship in the independent-types benchmark?&lt;/h3&gt;
&lt;p&gt;A: (Proposition 4) In the benchmark economy with independent types, the optimal distortion on spouse i with productivity t equals exactly one-half of the optimal distortion λ^{sng,&lt;em&gt;}(t) in the corresponding unidimensional economy: λ&lt;/em&gt;&lt;em&gt;i(t, w&lt;/em&gt;{-i}) = (1/2)λ^{sng,*}(t), and this is independent of the partner&amp;rsquo;s productivity w_{-i}. The intuition: the deadweight cost of taxing any individual depends only on her own characteristics (elasticity, productivity, density), not on whom she is married to. However, the redistributive benefit of taxation depends on matching — when matching is random, every high-productivity individual is married on average to an average person, so the incremental social benefit of extracting tax revenue from her is exactly half of what it would be if she were single (since half the benefit goes to a partner who is already average). More generally (Proposition 5 and Corollary 2), average distortions are weakly lower for married individuals than for singles as long as matching is not perfectly positively assortative.&lt;/p&gt;
&lt;h3 id="q5-what-is-average-jointness-and-how-is-it-measured"&gt;Q5. What is average jointness and how is it measured?&lt;/h3&gt;
&lt;p&gt;A: Average jointness J_i(t) is defined as the ratio of average distortions on spouse i conditional on the partner having above-t productivity to average distortions conditional on the partner having below-t productivity, minus one. Jointness is positive if the marginal tax rate on spouse i is on average increasing in the partner&amp;rsquo;s productivity, negative if decreasing, and zero for separable (individual earnings-based) taxes. The paper characterizes jointness through auxiliary functions H_i(t) (conditional distortion relative to unconditional average), whose behavior is determined by the copula elasticities η_i and survival copula elasticities η̄_i — the percentage change in the conditional quantile of the partner&amp;rsquo;s productivity when one spouse&amp;rsquo;s productivity quantile increases by 1%.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-role-of-tail-dependence-in-determining-the-sign-of-optimal-jointness"&gt;Q6. What is the role of tail dependence in determining the sign of optimal jointness?&lt;/h3&gt;
&lt;p&gt;A: (Proposition 7, Lemma 4) For right-tail dependent distributions — where the probability that an extremely productive person is married to an extremely productive partner remains bounded away from zero as productivity → ∞ — the redistributive benefit of positive jointness (targeting taxes to the richest couples) dominates its distortionary cost, so optimal average jointness is positive at the top. For right-tail independent distributions (where this probability converges to zero), the distortionary cost of positive jointness dominates, and optimal jointness is negative at the top. Exactly symmetric logic applies at the bottom using the survival copula and left-tail dependence. The bivariate lognormal/Gaussian copula is right-tail independent for any finite correlation ρ, while a distribution with perfect assortative matching in the tails would be right-tail dependent. The speed of convergence to tail independence, measured by κ = lim_{u→0} ln(u)/ln(C(u,u)) ∈ [1/2, 1), also matters: slower convergence (κ closer to 1) implies smaller optimal jointness under tail independence.&lt;/p&gt;
&lt;h3 id="q7-when-is-individual-earnings-based-separable-taxation-optimal-and-when-is-family-earnings-based-taxation-optimal"&gt;Q7. When is individual earnings-based (separable) taxation optimal, and when is family earnings-based taxation optimal?&lt;/h3&gt;
&lt;p&gt;A: (Propositions 4, 8, Corollary 1) Individual earnings-based taxation is optimal when Pareto weights are separable and spousal productivities are independent. When types are positively dependent, the planner introduces jointness even with separable social weights, because conditioning taxes on both spouses&amp;rsquo; earnings facilitates redistribution across couple types. Family earnings-based taxation is optimal when: (i) social weights are measurable only with respect to total family productivity r (i.e., the planner cares only about total family output, not the identity or relative productivity of individual spouses), and (ii) total family productivity r and relative spousal productivity ι are statistically independent. When r and ι are not independent, even a planner with an intrinsic preference for family earnings-based taxation will find it optimal to depart from it.&lt;/p&gt;
&lt;h3 id="q8-what-does-proposition-9-corollary-7-establish-about-the-relationship-between-restricted-and-unrestricted-optimal-taxes"&gt;Q8. What does Proposition 9 (Corollary 7) establish about the relationship between restricted and unrestricted optimal taxes?&lt;/h3&gt;
&lt;p&gt;A: (Corollary 7) For each restricted tax regime (anonymous, individual earnings-based, family earnings-based), the optimal distortions under the restricted tax equal the corresponding conditional average of unrestricted optimal distortions. Specifically: optimal individual earnings-based distortions equal E[λ*_i | w_i = t] (the average unrestricted distortion at productivity t); optimal family earnings-based distortions equal E[weighted average of λ*_i | R(w) = r]. This reveals that the unrestricted and restricted planners solve the same tradeoff between redistribution benefits and distortionary costs, but the restricted planner must apply a single tax rate to groups of couples that cannot be distinguished under the restriction. The welfare loss from restriction comes entirely from this forced bunching, not from a different objective or a different first-order condition.&lt;/p&gt;
&lt;h3 id="q9-what-do-the-quantitative-results-say-about-the-goodness-of-approximation-of-separable-vs-family-earnings-based-taxation"&gt;Q9. What do the quantitative results say about the goodness of approximation of separable vs. family earnings-based taxation?&lt;/h3&gt;
&lt;p&gt;A: In the calibrated benchmark economy (Gaussian copula, ρ = 0.33, Pareto-lognormal marginals, γ = 0.25, m = 0.35), optimal jointness is quantitatively small — the marginal tax rate on one spouse changes by at most several percentage points as a function of the other spouse&amp;rsquo;s earnings over the plotted range. Individual earnings-based (separable) taxation therefore provides a good approximation to the unrestricted optimum across all specifications considered. By contrast, family earnings-based taxation is a poor approximation: the marginal tax rate on family income varies substantially with the earnings share of the secondary earner (the ratio min{y1,y2}/(y1+y2)), and the deviation from the optimal unrestricted tax is large. This finding is robust across different Pareto weight specifications (m ∈ {0.35, 1.5}, k ∈ {0, 1, 2}) and holds even when k = 0, i.e., when the planner&amp;rsquo;s social weights inherently prefer family earnings-based taxation.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-calibration-results-relate-to-the-analytical-comparative-statics-predictions"&gt;Q10. How do the calibration results relate to the analytical comparative statics predictions?&lt;/h3&gt;
&lt;p&gt;A: The calibration validates the analytical predictions quantitatively. The analytical result (Proposition 5) that optimal distortions in the U.S. lie between those under random matching (1/2 of single-individual rates) and perfect assortative matching (same as single-individual rates) is confirmed: optimal tax rates for married individuals in the calibrated economy lie between the independence and perfect-dependence gray-line benchmarks in Figure 6. The analytical prediction (Proposition 7) that the Gaussian copula implies positive jointness at the bottom and negative at the top is confirmed, with the switch to negative jointness occurring above approximately $8.5 million in earnings. The slow convergence of the Gaussian copula to tail independence (κ = (1+ρ)/2 ≈ 0.665) explains the small magnitude of optimal jointness relative to the FGM copula (which has κ = 1/2, faster convergence, and exhibits more pronounced jointness as shown in the appendix). The analytical limiting distortion of E[λ*_i | w_i = t] → 1/(γa) ≈ 1.35 as t → ∞ (corresponding to a top marginal tax rate of approximately 55 percent) is confirmed, though convergence is slow and rates remain substantially below this limit at $300,000 in earnings.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-relate-to-and-advance-beyond-kleven-kreiner-and-saez-20072009"&gt;Q11. How does the paper relate to and advance beyond Kleven, Kreiner, and Saez (2007/2009)?&lt;/h3&gt;
&lt;p&gt;A: Kleven et al. (2009) studied couples taxation but avoided the multi-dimensional screening complexity by restricting the secondary earner to binary labor supply. The working paper by Kleven et al. (2007) considered the continuous setting but noted the difficulty of the FOA and derived several special-case insights. The current paper extends KKS in several systematic ways: it provides the first formal proof that the FOA conditions are strictly weaker in bi-dimensional than unidimensional settings; generalizes the formula for average distortions to arbitrary joint distributions (not just independent types); characterizes optimal jointness under positive dependence (not just independence); establishes the role of tail (in)dependence in determining the sign of jointness; compares optimal taxes for married vs. single individuals; and derives conditions under which family earnings-based or individual earnings-based taxation is optimal. It also shows that the KKS result on jointness sign (determined by the third derivative of the SWF) applies only under independence and can be reversed even with arbitrarily small positive dependence, as demonstrated with the Gaussian copula example.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;First-Order Approach (FOA) in multi-dimensional taxation.&lt;/strong&gt; The restriction of the mechanism design problem to local incentive constraints only — dropping global (non-local) incentive compatibility conditions and solving a relaxed problem. In the paper&amp;rsquo;s context, FOA validity is equivalent to convexity of a specific transformation vx* of the optimal utility function in the &amp;ldquo;linearized&amp;rdquo; type space X. The paper shows that the condition for FOA validity is strictly weaker (i.e., a strictly larger set of primitives satisfies it) in the bi-dimensional couples setting than in the corresponding unidimensional model, because the absence of participation constraints eliminates the main force driving FOA failure in industrial organization multi-dimensional screening.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;Optimal tax distortion λ&lt;/em&gt;_i(w).&lt;/em&gt;* The monotone transformation of the marginal tax rate defined by λ_i(w) = [∇_i T(y(w))] / [1 − ∇_i T(y(w))], where ∇_i T is the partial derivative of the tax function with respect to spouse i&amp;rsquo;s earnings. This transformation maps [−∞, ∞] marginal tax rates to (−1, ∞) distortions. The optimal tax schedule is characterized by the function λ* satisfying a system of PDEs; the paper studies conditional averages of λ* rather than λ* pointwise.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Coarea Formula.&lt;/strong&gt; A mathematical result from geometric measure theory that, in this context, converts an integral of the PDE optimality condition over a two-dimensional domain into an integral over the level sets of an arbitrary function Q(w). Applied to equation (20), it yields: E[Σ_i λ*_i γ_i (∂lnQ/∂lnw_i) | Q=t] = [1 − E[α|Q≥t]] / [−∂ln P(Q≥t)/∂ln t]. By choosing different Q functions, the formula delivers conditional averages of optimal distortions over different subsets of the type space, all in terms of exogenous primitives. This is the paper&amp;rsquo;s principal analytical tool for characterizing optimal taxes without solving the PDE explicitly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Jointness (positive/negative).&lt;/strong&gt; The dependence of the optimal marginal tax rate on one spouse&amp;rsquo;s earnings on the other spouse&amp;rsquo;s earnings. Taxes are positively jointed at w if ∂²T/∂y_1∂y_2 &amp;gt; 0 (so raising one spouse&amp;rsquo;s earnings increases the marginal tax rate on the other); negatively jointed if this cross-partial is negative; disjointed (separable) if it is zero. Average jointness J_i(t) at productivity t is measured as the ratio of conditional average distortions above and below the partner&amp;rsquo;s productivity threshold, minus one. Optimal jointness is the paper&amp;rsquo;s primary policy object for understanding how taxes on one spouse should respond to the other&amp;rsquo;s earnings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Copula and survival copula elasticities (η_i, η̄_i).&lt;/strong&gt; Defined as η_i(t) = ∂ln C(u)/∂ln u_i and η̄_i(t) = ∂ln C̄(u)/∂ln ū_i, where C is the copula of the joint productivity distribution, C̄ is the survival copula, and u_i = G_i(t_i), ū_i = 1−G_i(t_i) are the corresponding quantiles. These elasticities measure the percentage change in the conditional quantile of the partner&amp;rsquo;s productivity when one spouse&amp;rsquo;s productivity quantile increases by 1%. They quantify the additional distortionary cost introduced by jointness relative to a separable tax schedule: smaller elasticities (stronger dependence) correspond to larger distortionary costs of jointness at the boundaries of probability mass.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tail (in)dependence.&lt;/strong&gt; A joint distribution F is right-tail dependent if lim_{t→∞} P(w_{-i}≥t | w_i≥t) &amp;gt; 0, i.e., extremely productive individuals have a positive probability of being matched with equally extreme partners. It is right-tail independent if this limit is zero. The speed of convergence to tail independence is measured by κ = lim_{u→0} ln(u)/ln(C(u,u)) ∈ [1/2, 1). Tail dependence determines the sign of optimal average jointness in the tails: right-tail dependence favors positive jointness at the top; right-tail independence favors negative jointness at the top. The Gaussian copula is right-tail independent for any finite ρ; a perfectly assortative matching distribution is right-tail dependent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Positive quadrant dependence (PQD) order.&lt;/strong&gt; A partial ordering on joint distributions with the same marginals: F^b ≥_{PQD} F^a if F^b(w) ≥ F^a(w) for all w, equivalently if Cov(φ_1(w_1), φ_2(w_2)) ≥ 0 for any two increasing functions. The paper uses this order to rank economies by the &amp;ldquo;assortativeness&amp;rdquo; of matching, and shows that optimal average distortions are monotone in this order (Proposition 5): more assortative matching implies weakly higher optimal tax distortions on each married individual.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pareto-lognormal (PLN) distribution.&lt;/strong&gt; Used in the calibration to model the marginal distribution of spousal productivities. Defined as G(t) = Φ((ln t − μ)/σ) − a·exp(aμ + a²σ²/2)·Φ((ln t − μ)/σ − aσ), parameterized by location μ, scale σ, and tail parameter a. The PLN family has a lognormal body and a Pareto tail with tail parameter a, making it suitable for capturing the empirical finding of a thin left tail (implying optimal marginal taxes approaching zero as earnings → 0) and a thick right tail (implying a positive limiting marginal tax rate of approximately 1/(1 + 1/(γa)) as earnings → ∞).&lt;/p&gt;</description></item><item><title>The Power of Proximity to Coworkers</title><link>https://macropaperwarehouse.com/papers/the-power-of-proximity-to-coworkers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-power-of-proximity-to-coworkers/</guid><description>&lt;p&gt;This paper studies how physical proximity to coworkers affects on-the-job training and productivity, using software engineers at a Fortune 500 online retailer observed from 2019 to 2024. The authors exploit two quasi-experimental shocks to proximity: the office closures of 2020, which eliminated proximity differentials that previously existed across team types, and the firm&amp;rsquo;s subsequent return-to-office (RTO) mandates in 2022 and 2023, which restored proximity for co-located teams while leaving geographically-distributed teams apart. The core identification strategy is a difference-in-differences design comparing engineers whose teams were co-located in a single headquarters building to those whose teams were split across two buildings a ten-minute walk apart — a distinction that became immaterial once offices closed.&lt;/p&gt;
&lt;p&gt;The central finding is that sitting near teammates substantially increases the digital feedback engineers receive on their code. Before the office closures, engineers on co-located teams received 23.9% (1.92 comments per program) more code review feedback than engineers on multi-building teams. Once offices closed, this advantage narrowed by 18.3% (1.47 comments per program, p-value = 0.0026). The lost comments were disproportionately those predicted by a machine-learning classifier to be helpful, actionable, well-reasoned, and impactful, with high-quality comments declining by 21–23% — exceeding the overall volume decline. Face-to-face and digital communication are complements, not substitutes: proximate engineers drew on a wider pool of reviewers and asked 48.4% more follow-up questions, a differential that vanished once offices closed.&lt;/p&gt;
&lt;p&gt;Proximity&amp;rsquo;s effects are highly heterogeneous. Gains in feedback are concentrated among less-tenured, younger, and female engineers — those with the most to learn. Junior engineers on co-located teams lost 2.03 more comments per program upon office closure than junior engineers already on distributed teams (p-value = 0.001); young engineers lost 2.47 more comments (p-value = 0.0001). Female engineers lost 38.9% more comments than their distributed female counterparts (p-value &amp;lt; 0.0001), partly because women stop asking as many people for feedback when they cannot do so in person.&lt;/p&gt;
&lt;p&gt;Proximity improves code quality for inexperienced engineers. Around the second RTO (three days per week), engineers on co-located teams became 2.2 percentage points less likely to add files subsequently deleted — a measure of churn — and 1.4 pp less likely to introduce bugs, relative to distributed teams (p-values of 0.041 and 0.022 respectively). These gains were roughly twice as large for less-tenured and younger engineers. The benefits persist: engineers who spent more pre-closure time on co-located teams continued to write higher-quality code during the fully remote period.&lt;/p&gt;
&lt;p&gt;However, mentorship is costly for those who provide it. Senior engineers on co-located teams wrote 0.76 fewer programs per month in the main codebase before closures (p-value = 0.0005), a gap that closed when offices did and widened again during the second RTO. The firm faces a fundamental tradeoff: proximity accelerates junior engineers&amp;rsquo; human capital development while reducing experienced engineers&amp;rsquo; immediate coding output.&lt;/p&gt;
&lt;p&gt;These dynamics shape hiring. The firm shifted toward hiring older, more experienced engineers during closures — buying talent it could no longer build in-house — and back toward younger hires once offices reopened. Nationally, young college graduates in remotable occupations (classified per Dingel and Neiman, 2020) experienced a 0.88 pp increase in unemployment between 2017–2019 and 2022–2024, while older graduates saw a marginal decline of 0.11 pp. A triple-difference estimate finds a 0.65 pp greater increase in young workers&amp;rsquo; unemployment in remotable versus non-remotable occupations (p-value = 0.029), a pattern that predates generative AI diffusion and is robust to controlling for AI exposure. Back-of-the-envelope, remote work accounts for an estimated 64% of the total unemployment increase among young college graduates over this period.&lt;/p&gt;
&lt;p&gt;The paper also documents that proximity is fragile: a ten-minute walk between two buildings reduces feedback as much as being multiple states away, and even a single distant teammate imposes negative externalities on those who remain co-located, reducing their feedback by 1.71 comments per program (p-value = 0.095) via a &amp;ldquo;one Zoom, all Zoom&amp;rdquo; norm.&lt;/p&gt;
&lt;p&gt;Q: What is the main identification strategy for the office-closure analysis, and what is the key parallel-trends evidence?&lt;/p&gt;
&lt;p&gt;A: The authors compare engineers on co-located teams (all members in one headquarters building) to those on multi-building teams (split across two buildings a ten-minute walk apart), before and after the March 2020 office closures. Co-located teams lost more proximity when offices closed, while multi-building teams experienced a smaller shock, enabling a difference-in-differences design. Pre-closure trends in feedback are parallel across the two team types (Figure I), supporting the identifying assumption. Standard errors are clustered by team, the unit of treatment assignment.&lt;/p&gt;
&lt;p&gt;Q: How large is the effect of proximity on total code review feedback, and how is it broken down by feedback source?&lt;/p&gt;
&lt;p&gt;A: Before closure, co-located engineers received 23.9% (1.92 comments per program) more feedback than multi-building engineers. The DiD estimate indicates that losing proximity reduced feedback by 18.3% (1.47 comments per program, p-value = 0.0026, Column 3 of Table II). This decline stems entirely from reduced feedback from teammates; there is no detectable effect on feedback from engineers on other teams — a placebo check that supports the identification strategy and rules out explanations based on differential project complexity.&lt;/p&gt;
&lt;p&gt;Q: How does proximity affect the quality — not just the quantity — of code review comments?&lt;/p&gt;
&lt;p&gt;A: Using a gradient-boosted decision tree trained on 5,377 human-labeled comments, the authors predict comment quality across all 174,014 comments. Losing proximity reduced comments predicted to be helpful, well-reasoned, actionable, and likely to change the code by 21–23% — exceeding the 18.3% overall volume decline. The residual comments were lower quality: 2.9 pp fewer were helpful (p-value = 0.039), 1.7 pp fewer explained their reasoning (p-value = 0.094), and 1.9 pp fewer were likely to change the code (p-value = 0.072).&lt;/p&gt;
&lt;p&gt;Q: What mechanisms drive the complementarity between face-to-face interaction and digital feedback?&lt;/p&gt;
&lt;p&gt;A: Proximity increases feedback on both the extensive and intensive margins. On the extensive margin, co-located engineers draw on a wider pool of reviewers, returning less frequently to the same commenter. On the intensive margin, losing proximity reduces follow-up questions by 48.4% (0.12 questions per program, p-value = 0.0083), accounting for roughly half of the total feedback decline. The other half comes from reduced initial reviewer feedback. References to other communication channels (e.g., Slack) within code reviews also decline when proximity is lost, confirming that face-to-face and digital communication are complements.&lt;/p&gt;
&lt;p&gt;Q: How small a physical barrier is sufficient to reduce feedback substantially?&lt;/p&gt;
&lt;p&gt;A: A ten-minute walk between two buildings on the same headquarters campus reduces feedback by as much as being multiple states away — both groups receive significantly less feedback than engineers whose entire team sits in the same building (Figure Ib). This finding aligns with research on academics showing that different floors or buildings reduce coauthorship, and extends it to daily teammates sharing projects.&lt;/p&gt;
&lt;p&gt;Q: What are the externality effects of a single distant teammate?&lt;/p&gt;
&lt;p&gt;A: Through the firm&amp;rsquo;s implicit &amp;ldquo;one Zoom, all Zoom&amp;rdquo; norm, even one teammate in a different location shifts all team meetings to video calls. Engineers in the same building exchange 14.5% less feedback when even one teammate is in another building versus when all teammates are co-located (p-value = 0.037). When a new hire transforms a co-located team into a multi-building one, feedback between the original co-located teammates drops by 1.71 comments per program (p-value = 0.095); adding a new co-located hire produces no such decline.&lt;/p&gt;
&lt;p&gt;Q: How does the effect of proximity on feedback differ by engineer tenure, age, and gender?&lt;/p&gt;
&lt;p&gt;A: Less-tenured engineers on co-located teams lost 2.03 more comments per program upon closure than less-tenured engineers on distributed teams (p-value = 0.001). Young engineers (under 29) on co-located teams lost 2.47 more comments per program than young distributed engineers (p-value = 0.0001). Female engineers on co-located teams lost 38.9% (3.71) more comments than female engineers on distributed teams (p-value &amp;lt; 0.0001), partly because women draw feedback from 14.7% fewer people when proximity is lost (p-value = 0.0078), compared to a negligible 2.6% decline for men. The extra feedback women receive in person is of higher quality, not rude or condescending.&lt;/p&gt;
&lt;p&gt;Q: How is the effect of proximity on code quality identified using the RTO design, and what are the magnitudes?&lt;/p&gt;
&lt;p&gt;A: The RTO design compares engineers on co-located (same-city) teams to geographically-distributed teams across three periods: full closure, first RTO (two days per week), and second RTO (three days per week). The authors predict γ_closed ≈ 0 (office assignment irrelevant when closed) and γ_2nd_RTO &amp;gt; γ_1st_RTO (more in-office days means more proximity). Both predictions are confirmed. During the second RTO, co-located engineers were 2.2 pp less likely to add files later deleted (p-value = 0.041) and 1.4 pp less likely to introduce bugs (p-value = 0.022), with effects roughly twice as large for less-tenured and younger engineers.&lt;/p&gt;
&lt;p&gt;Q: Does the benefit of co-location on code quality persist after remote work resumes?&lt;/p&gt;
&lt;p&gt;A: Yes. After all engineers returned to remote work, those who had been on co-located teams pre-closure were 2.37 pp less likely to write disposable code (p-value = 0.013) and 3.09 pp less likely to introduce bugs (p-value = 0.0012). Code quality improves monotonically with the number of pre-closure months spent on co-located teams (Figure A.5). These gaps persist when including current team fixed effects, meaning within the same post-closure team, the previously co-located engineer writes higher-quality code.&lt;/p&gt;
&lt;p&gt;Q: What is the cost of mentorship for senior engineers, and how does it manifest in coding output?&lt;/p&gt;
&lt;p&gt;A: Senior engineers on co-located teams wrote 0.76 fewer programs per month in the main codebase when offices were open (p-value = 0.0005). Once offices closed, this gap disappeared, and senior engineers who lost proximity to their teammates saw a relative increase in output of 0.58 programs per month (p-value = 0.0014). During the second RTO, engineers with more than sixteen months of tenure on co-located teams wrote fewer programs, while no significant difference emerged for less-tenured engineers. Overall, the DiD estimate indicates losing proximity to teammates increases immediate output by 0.48 programs per month (p-value = 0.0002).&lt;/p&gt;
&lt;p&gt;Q: How does the firm&amp;rsquo;s hiring age distribution respond to changes in proximity?&lt;/p&gt;
&lt;p&gt;A: When offices were closed, the firm shifted toward hiring older engineers: the share of hires under age 29 fell from over half pre-closure to less than a third during the closure. After the RTOs, the firm shifted back toward younger hires. Geographic variation reinforces this: headquarters-campus hires were 7–10 years younger than those hired into distributed roles when offices were open; this gap narrowed substantially during closures when everyone was far from teammates.&lt;/p&gt;
&lt;p&gt;Q: Does proximity affect which engineers are poached by other firms?&lt;/p&gt;
&lt;p&gt;A: Yes. During the office closures, 1.2% of co-located engineers were poached per month, compared to 0.9% of multi-building engineers of similar tenure, age, and engineering group (p-value = 0.044). By the end of the closure period, nearly a quarter of co-located engineers had been poached versus a sixth of multi-building engineers. There is a dose response: more pre-closure time on co-located teams predicts higher poaching rates. The effect is concentrated among younger and female engineers, consistent with their feedback building more transferable general human capital. Tenure does not moderate the poaching effect, consistent with less-tenured engineers&amp;rsquo; feedback being more firm-specific.&lt;/p&gt;
&lt;p&gt;Q: What does national unemployment data show about the scarring effects of remote work on young workers?&lt;/p&gt;
&lt;p&gt;A: Between 2017–2019 and 2022–2024, young college graduates (under 29) in remotable occupations experienced a 0.88 pp increase in unemployment (p-value &amp;lt; 0.00001), while older graduates in the same occupations saw a marginal decline of 0.11 pp (p-value = 0.053). A triple-difference regression finds a 0.65 pp greater increase in young workers&amp;rsquo; unemployment in remotable versus non-remotable occupations (p-value = 0.029). Back-of-the-envelope, scaling this estimate by the 61% share of young graduates in remotable jobs predicts a 0.4 pp increase in young college graduates&amp;rsquo; overall unemployment — equal to 64% of the realized 0.63 pp increase.&lt;/p&gt;
&lt;p&gt;Q: Is the unemployment increase among young workers in remotable jobs driven by generative AI rather than remote work?&lt;/p&gt;
&lt;p&gt;A: The authors argue against AI as the primary driver on two grounds. First, the uptick in young workers&amp;rsquo; unemployment in remotable occupations predates the rapid diffusion of generative AI. Second, the differential increase is not concentrated among occupations with the highest AI task exposure. The triple-difference estimate is robust to controlling for occupational AI exposure using the Eisfeldt, Schubert and Zhang (2023) index. The authors acknowledge that AI may become more important as it diffuses further.&lt;/p&gt;
&lt;p&gt;Q: How do young workers&amp;rsquo; own office attendance decisions reflect the value of proximity?&lt;/p&gt;
&lt;p&gt;A: At the partner firm, engineers under 29 were 8.8 pp (37.6%) more likely to come into the office during the RTOs than older engineers when on co-located teams (solid line in Figure VIIa). This difference was roughly halved on geographically-distributed teams (p-value of difference = 0.0085), indicating that the draw is specifically proximity to teammates. Co-located managers raised attendance by 2.6 pp, while co-located teammates raised it by 5.1 pp. Nationally, Stack Overflow survey data show nearly half of engineers under 25 are in the office each day, versus a quarter of older engineers (p-value &amp;lt; 0.00001).&lt;/p&gt;
&lt;p&gt;Q: What does the paper imply about why remote work was rare before the pandemic despite workers&amp;rsquo; stated preferences for it?&lt;/p&gt;
&lt;p&gt;A: The paper offers a resolution: firms may have recognized that the value of the office lies in training for tomorrow and improving the quality — not the quantity — of work today. Remote work boosts immediate output, especially for experienced workers, but it reduces mentorship and long-run skill development. The tradeoff between current and future productivity, and between individual and collective returns to human capital, explains why firms historically resisted remote work even when workers preferred it and short-run output was unaffected.&lt;/p&gt;
&lt;p&gt;Q: What are the implications for gender equity in remote work?&lt;/p&gt;
&lt;p&gt;A: The findings suggest remote work has ambiguous gender effects. While remote work may help working mothers remain in the workforce, it appears costly for young women&amp;rsquo;s professional development, which is especially sensitive to physical proximity. Women receive substantially more high-quality feedback when co-located, draw feedback from a wider network in person, and lose disproportionately more feedback when proximity is lost. Young female engineers on co-located teams were also disproportionately poached — suggesting their human capital gains from co-location are more general and transferable.&lt;/p&gt;
&lt;p&gt;Code review feedback: The digital comments engineers exchange when reviewing each other&amp;rsquo;s code before it is merged into the live codebase; the paper&amp;rsquo;s primary measure of on-the-job training and mentorship investment, distinct from mere volume because the authors also classify comments by helpfulness, reasoning, actionability, and expected impact using supervised machine learning.&lt;/p&gt;
&lt;p&gt;Co-located team: A team in which all members are assigned to the same office building; the treatment group in the difference-in-differences designs, distinguished from multi-building teams (split across two headquarters buildings, a ten-minute walk apart) and geographically-distributed teams (members in different cities or permanently remote).&lt;/p&gt;
&lt;p&gt;One Zoom, all Zoom norm: The implicit team practice of holding all meetings virtually if any single teammate cannot be physically present; the mechanism by which one distant colleague generates negative externalities for the remaining co-located teammates, reducing their in-person interaction and feedback.&lt;/p&gt;
&lt;p&gt;Proximity fragility: The finding that even small physical barriers — a ten-minute walk between buildings — reduce feedback as much as being multiple states away, implying that the relationship between physical distance and mentorship is highly nonlinear near zero.&lt;/p&gt;
&lt;p&gt;Churn (disposable code): Files that are added by an engineer but deleted within the subsequent six months, either because the code was poorly structured or because it introduced a feature later abandoned; used as one of two code quality proxies in the RTO analysis (occurring in 15% of programs).&lt;/p&gt;
&lt;p&gt;Bugs (immediate reversions): Programs that are immediately and fully reverted after being merged, typically indicating the engineer&amp;rsquo;s changes precipitated an emergency requiring rollback to an earlier version; used as the more serious of the two code quality proxies (occurring in 3.5% of programs).&lt;/p&gt;
&lt;p&gt;Scarring effects: The persistent adverse impact on young workers&amp;rsquo; human capital and labor market outcomes from reduced mentorship during the remote work period; manifested both as lower code quality at the individual level and higher unemployment rates nationally among young college graduates in remotable occupations.&lt;/p&gt;
&lt;p&gt;Remotable occupation: An occupation classified by Dingel and Neiman (2020) as feasibly performed from home; used to construct the national triple-difference analysis comparing age gaps in unemployment across remotable and non-remotable jobs before and after the pandemic.&lt;/p&gt;</description></item><item><title>The Productivity of Professions: Evidence from the Emergency Department</title><link>https://macropaperwarehouse.com/papers/the-productivity-of-professions-evidence-from-the-emergency-department/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-productivity-of-professions-evidence-from-the-emergency-department/</guid><description>&lt;p&gt;This paper studies the productivity of nurse practitioners (NPs) versus physicians performing overlapping tasks in Veterans Health Administration (VHA) emergency departments (EDs), exploiting a quasi-experiment created by the VHA&amp;rsquo;s December 2016 grant of full practice authority to NPs. The identification strategy instruments patient assignment to NPs versus physicians using quasi-random variation in the number of NPs on duty on a given ED-day, conditional on ED-by-time-category fixed effects. The sample covers 1.1 million ED visits across 44 VHA EDs from January 2017 to January 2020, seen by 1,348 physicians and 156 NPs. The instrument is validated by demonstrating balance in patient observable characteristics across values of the instrument, stability of IV estimates across 256 combinations of patient covariate controls, and absence of spillover effects from NP presence onto physician performance.&lt;/p&gt;
&lt;p&gt;On average in the ED setting, NPs increase patient length of stay by 11 percent (approximately 18 additional minutes) and raise the cost of the ED visit by 7 percent (approximately $66 per visit). NPs raise the 30-day preventable hospitalization rate by 0.25 percentage points, a 20 percent increase relative to the mean. No statistically significant effect on 30-day mortality is detected (95 percent confidence interval: -0.34 to 0.11 percentage points). OLS estimates carry the opposite sign because NPs are assigned healthier patients in observational data; the IV design corrects for this selection.&lt;/p&gt;
&lt;p&gt;The average NP-physician performance gap varies systematically by case complexity and severity. For the highest-complexity quartile of cases (by Elixhauser comorbidities), NPs increase ED costs by 12 percent and length of stay by 28 percent. For cases at or above the 95th percentile of severity (based on 30-day mortality by diagnosis), NPs increase ED costs by 25 percent, length of stay by 99 percent, and admissions by 26 percentage points (42 percent relative to the mean), while reducing 30-day preventable hospitalization by 3 percentage points — suggesting that NPs&amp;rsquo; higher care intensity partially offsets worse intrinsic skill for the most severe cases. For lower-complexity cases, the cost and length-of-stay gaps are smaller, but NPs still significantly raise preventable hospitalizations.&lt;/p&gt;
&lt;p&gt;NPs exhibit clinical decision-making patterns consistent with lower diagnostic skill: they are more likely to order consults (2.6 percentage points, or 11 percent of the mean), CT scans (1.2 percentage points, or 8.3 percent), and X-rays (2.0 percentage points, or 6.9 percent). NPs lower opioid prescriptions by 1.8 percentage points (20 percent of the mean) and raise antibiotic prescriptions by 4.0 percentage points (6.3 percent of the mean), consistent with threshold adjustment under lower diagnostic skill with asymmetric error costs. Downstream, patients treated by NPs incur similar opioid use disorder rates despite lower opioid prescribing, and higher infection-related return visit rates despite higher antibiotic prescribing.&lt;/p&gt;
&lt;p&gt;Counterfactual analysis finds that allocating one quarter of ED patients to NPs increases net spending by $129 million per year to the VHA after accounting for NPs&amp;rsquo; lower wages (approximately half of physicians&amp;rsquo;). However, deploying NPs exclusively to the least-complex quarter of cases reduces net spending to approximately one-fifth of this amount.&lt;/p&gt;
&lt;p&gt;A distributional analysis deconvolving provider-specific IV estimates reveals that within-profession productivity variation substantially exceeds the average between-profession gap. The interquartile range in annual spending attributable to provider productivity within each profession is approximately $900,000, roughly three times the mean annual spending difference between the average NP and the average physician. A randomly chosen NP outperforms a randomly chosen physician in up to 38 percent of pairs. Within professions, individual provider productivity shows essentially no relationship with wages or case complexity assigned, whereas between professions, case assignment and wages are strongly sorted by professional class.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question?
A: The paper asks whether NPs and physicians, who perform overlapping tasks in the ED but differ sharply in training, selectivity, and pay, differ in productivity, and how that average between-profession difference compares to productivity variation within each profession. It also asks what mechanisms drive any observed gap and how case assignment responds to provider skill differences.&lt;/p&gt;
&lt;p&gt;Q: What is the identification strategy and why is it credible?
A: The authors instrument patient assignment to NPs with the number of NPs on duty on the ED-day, conditional on ED-by-year, ED-by-month, ED-by-day-of-week, and ED-by-hour fixed effects. Credibility rests on: provider schedules being set months in advance, decoupling NP availability from arriving patient characteristics; patient characteristics being well balanced across values of the instrument conditional on fixed effects; IV estimates being stable across all 256 covariate-control combinations; and on-duty physician and NP characteristics also being balanced across the instrument.&lt;/p&gt;
&lt;p&gt;Q: What are the main average effects of NPs on resource use?
A: IV estimates show NPs increase patient length of stay by 11 percent (approximately 18 minutes) and ED cost by 7 percent (approximately $66 per visit). There is no significant average effect on inpatient admissions in the overall sample, though NPs significantly raise admissions for high-severity cases.&lt;/p&gt;
&lt;p&gt;Q: What is the effect of NPs on patient health outcomes?
A: NPs raise 30-day preventable hospitalizations by 0.25 percentage points, a 20 percent increase relative to the mean. The 95 percent confidence interval for 30-day mortality is -0.34 to 0.11 percentage points, implying no statistically significant mortality effect in the overall sample.&lt;/p&gt;
&lt;p&gt;Q: Why do OLS and IV estimates have opposite signs?
A: In observational data, NPs treat healthier patients than physicians: NP patients are younger (60.7 versus 62.5 years), have fewer Elixhauser comorbidities (3.2 versus 3.7), and have fewer prior inpatient stays (0.4 versus 0.7). This selection causes OLS estimates of NP effects to be negative. The IV corrects for this by exploiting quasi-random variation in NP availability; IV estimates are stable across all combinations of patient controls, consistent with the instrument being orthogonal to unobservable patient health.&lt;/p&gt;
&lt;p&gt;Q: How does the NP-physician performance gap vary with case complexity and severity?
A: For the highest-complexity quartile, NPs increase length of stay by 28 percent and ED costs by 12 percent without a significant preventable hospitalization effect. For cases at or above the 95th severity percentile, NPs increase length of stay by 99 percent, ED costs by 25 percent, and admissions by 26 percentage points (42 percent relative to the mean), while reducing 30-day preventable hospitalization by 3 percentage points. For lower-complexity quartiles, NPs show smaller cost and length-of-stay effects but significantly raise preventable hospitalizations, suggesting the higher care intensity at high severity compensates for lower skill.&lt;/p&gt;
&lt;p&gt;Q: What does the heterogeneity by severity imply for optimal case assignment?
A: The pattern is consistent with skill-task matching: NPs have a comparative and absolute disadvantage in complex cases, so optimal assignment directs less complex cases to NPs and fewer patients to NPs when physicians are more available. Empirically, NPs are indeed assigned healthier patients from the available pool, and are assigned a modestly smaller share when the ED is less busy.&lt;/p&gt;
&lt;p&gt;Q: What mechanisms explain the average NP-physician gap?
A: Three mechanisms are examined. First, experience: a one-standard-deviation increase in specific experience is associated with a 5.8 percent decline in the NP-physician length-of-stay gap, and general experience with a 10 percent decline; however, experience does not significantly narrow the preventable hospitalization gap. Second, information acquisition: NPs order more consults, CT scans, and X-rays, consistent with compensating for lower diagnostic skill. Third, prescription thresholds: NPs reduce opioid prescribing by 20 percent and raise antibiotic prescribing by 6.3 percent, consistent with threshold adjustment under asymmetric error costs, but downstream outcomes are not improved correspondingly.&lt;/p&gt;
&lt;p&gt;Q: What do prescription patterns and downstream outcomes reveal about NP diagnostic skill?
A: NPs prescribe fewer opioids yet patients treated by NPs obtain similar downstream opioid use disorder rates; NPs prescribe more antibiotics yet patients treated by NPs have higher rates of return visits with infections. This pattern is consistent with NPs exhibiting higher rates of both false positives and false negatives, not merely adjusted thresholds, suggesting genuinely lower diagnostic skill rather than threshold differences alone.&lt;/p&gt;
&lt;p&gt;Q: What do counterfactual cost calculations show?
A: Allocating one quarter of ED patients to NPs raises non-wage spending by $197 million per year to the VHA; after accounting for NP wages being half of physician wages (approximately $120,000 versus $240,000 per year), net cost is still $129 million per year. Restricting NP deployment to the least-complex quarter of cases reduces net spending to approximately one-fifth of this amount, illustrating that targeted case assignment substantially improves NP cost-effectiveness.&lt;/p&gt;
&lt;p&gt;Q: How large is within-profession productivity variation relative to between-profession differences?
A: The interquartile range in annual spending attributable to provider productivity within each profession is approximately $900,000, roughly three times the mean annual spending difference between the average NP and the average physician. A randomly chosen NP outperforms a randomly chosen physician in up to 38 percent of random pairs. The authors conclude that, despite stark differences in training and selection between professions, within-profession variation dominates.&lt;/p&gt;
&lt;p&gt;Q: Is individual provider productivity reflected in wages or case assignment within professions?
A: Within each profession, provider productivity shows essentially no relationship with wages or with the complexity of assigned cases. This contrasts sharply with between-profession patterns, where professional class strongly predicts both wages (NPs earn approximately $120,000 per year versus $240,000 for physicians) and assigned case complexity. The authors interpret this as evidence of informational and organizational frictions in recognizing individual productivity within professional classes, and note that professional class is a far stronger predictor of pay and case assignment than is individual productivity.&lt;/p&gt;
&lt;p&gt;Q: How do complier characteristics relate to the broader patient population?
A: Compliers — cases whose provider type is determined by the instrument — are healthier than the average case: younger, with fewer comorbidities, fewer prior inpatient stays, and lower predicted mortality. Never-takers are riskier than the average case. There are no always-takers since patients cannot be assigned to NPs on days when no NPs are on duty.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to the literature on NP scope-of-practice laws?
A: The scope-of-practice literature estimates general-equilibrium effects of allowing NPs greater autonomy, including labor reallocation between professions. This paper instead estimates the partial-equilibrium causal effect of assigning a patient to an NP versus a physician, holding the broader labor market fixed. The two literatures are complementary: the heterogeneity findings here suggest that scope-of-practice expansions may be more beneficial in lower-complexity primary care settings where the NP-physician performance gap is smaller.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: Three implications are highlighted. First, the efficiency of using NPs depends critically on case assignment: deploying NPs on the least-complex cases reduces net costs to approximately one-fifth of indiscriminate deployment. Second, the substantial overlap between NP and physician productivity distributions provides support for NP use in less complex settings even within the ED context. Third, within-profession productivity variation far exceeding between-profession differences suggests that individual-level productivity assessment, rather than professional class, may be a more accurate guide to case assignment and compensation.&lt;/p&gt;
&lt;p&gt;Quasi-experimental variation in NP availability: The identification strategy exploits day-to-day variation in the number of NPs scheduled to work in a given VHA ED, conditional on ED-by-time-category fixed effects, as an instrument for whether a patient is assigned to an NP versus a physician. Schedules are set months in advance, rendering the NP count orthogonal to arriving patient characteristics conditional on those fixed effects.&lt;/p&gt;
&lt;p&gt;30-day preventable hospitalization: A standardized quality-of-care outcome defined by the Agency for Healthcare Research and Quality, measuring hospitalizations occurring within 30 days of ED discharge that are classified as preventable given adequate prior outpatient management. Used by the paper as the primary downstream health outcome beyond the ED visit itself.&lt;/p&gt;
&lt;p&gt;Elixhauser comorbidities: A set of 31 binary indicators for chronic conditions (e.g., cancer, diabetes) based on medical histories in the prior 365 days, used in this paper to measure and stratify case complexity into quartiles for heterogeneity analysis.&lt;/p&gt;
&lt;p&gt;Productivity distributions within professions: Provider-specific productivity estimates derived from a just-identified IV model that instruments assignment to individual providers by indicators for on-duty providers, then deconvolved into underlying distributions using the Efron (2016) and Kline-Rose-Walters (2022) method. These distributions characterize the spread of productivity within each professional class, separate from measurement error.&lt;/p&gt;
&lt;p&gt;Prescription threshold adjustment: The mechanism, formalized in Chan, Gentzkow, and Yu (2022), by which providers with lower diagnostic skill optimally adjust treatment thresholds in response to asymmetric costs of false-positive versus false-negative errors. In this paper&amp;rsquo;s application, NPs lower the opioid prescription rate (where false positives carry higher costs: addiction and overdose) and raise the antibiotic prescription rate (where false negatives carry higher costs: untreated infection), but downstream outcomes do not improve correspondingly.&lt;/p&gt;
&lt;p&gt;Skill-task matching: The organizational economics principle (Acemoglu and Autor 2011) that efficiency requires assigning more complex tasks to higher-skilled workers. The paper documents that between professions, case assignment broadly follows this principle (NPs receive less complex patients on average), but within professions, essentially no matching between individual provider productivity and case complexity is observed.&lt;/p&gt;
&lt;p&gt;Full practice authority (VHA, December 2016): The VHA policy that allowed NPs to treat patients independently without physician supervision at VHA facilities, superseding state-level restrictions. This policy change defines the start of the paper&amp;rsquo;s sample period and establishes the institutional context in which the quasi-experiment occurs, as it removed the requirement for physician oversight that previously constrained NP independence.&lt;/p&gt;</description></item><item><title>The Social Tax: Redistributive Pressure and Labor Supply</title><link>https://macropaperwarehouse.com/papers/the-social-tax-redistributive-pressure-and-labor-supply/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-social-tax-redistributive-pressure-and-labor-supply/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks whether informal redistributive pressure — the social obligation to share earned income with kin and social networks — distorts labor supply in low-income communities. The authors conceptualize such pressure as a &amp;ldquo;social tax&amp;rdquo; on earnings and develop the first direct causal test of whether it reduces labor supply, output, and earnings among full-time workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and Sample&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The study works with 474 full-time piece-rate factory workers (464 of whom are women) employed in cashew processing plants run by Olam in Côte d&amp;rsquo;Ivoire. Workers are paid biweekly in cash entirely through piece rates for individual nut-peeling output, creating a direct mapping between labor supply and income. At baseline, workers report transferring 25–35% of their income to individuals outside their household, with 77% having made at least one transfer in the previous 3 months. Workers also strongly believe that earning more triggers more transfer requests: 77% agree that if someone starts earning more by working harder, people will ask that person more often for financial help.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intervention&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors introduce a blocked savings account into which workers can deposit any earnings above a self-chosen threshold (set at least as high as their own baseline average earnings). Earnings above the threshold are automatically deposited by the factory directly into the account with the Banque Populaire de Côte d&amp;rsquo;Ivoire; the cash component of pay is unchanged. Funds cannot be withdrawn until the end of the blocked period (9 months in Phase 1; 3 months in Phase 2). The key design feature is that the account reduces the effective social tax rate only on earnings &lt;em&gt;increases&lt;/em&gt; above baseline, thereby eliminating income effects and generating only a pure substitution effect — an unambiguous positive prediction on labor supply if a social tax exists.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Experimental Design&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Workers are randomized into three conditions: (1) Control (no account); (2) Private account (existence unknown to anyone outside the worker); (3) Non-private account (existence and forthcoming unblock date revealed to network members via promotional text messages). The contrast between Private and Non-private isolates the role of redistributive pressure specifically — holding constant all other features of the blocked account product. The experiment runs in two cross-randomized phases conducted between 2018 and 2019.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Take-up of blocked accounts is dramatically higher when accounts are private: 60% in Phase 2 (Private) versus 14% (Non-private), a 77% decline (p&amp;lt;0.001). Among workers who declined Non-private accounts, 96% cite anticipated increases in transfer requests as an important factor.&lt;/p&gt;
&lt;p&gt;Being offered a Private account sharply raises labor supply. Pooling both phases, the Private arm increases average daily earnings by 175.9 FCFA, or &lt;strong&gt;11.4%&lt;/strong&gt; (p=0.012), relative to Control or Non-private arms. This is accompanied by a &lt;strong&gt;6.2 percentage point (9.7%)&lt;/strong&gt; increase in daily work attendance (p=0.023), with the entire attendance effect driven by reduced absenteeism rather than turnover. Effects in Phase 1 (Private vs. Control: +11.3%, p=0.032) and Phase 2 (Private vs. Non-private: +11.5%, p=0.043) are nearly identical in magnitude, indicating the results are not sensitive to cross-phase design. The treatment effect magnitude is equivalent to each worker working an additional 1.19 days in every two-week paycycle. Because 89% of workers have no income outside the factory, these constitute increases in total earned income.&lt;/p&gt;
&lt;p&gt;Heterogeneity is consistent with the hypothesized mechanism: among workers who report difficulty saving due to redistributive pressure, the Private treatment increases earnings by &lt;strong&gt;15.0%&lt;/strong&gt; (p=0.018); among those not reporting such difficulty, the estimated effect is near zero and insignificant (p=0.95). Among workers who report transfers to acquaintances (the most likely social-tax-motivated transfers), the effect is &lt;strong&gt;17.5%&lt;/strong&gt; (p=0.014). Workers without a partner — for whom intra-household redistribution is irrelevant — experience a &lt;strong&gt;15.8%&lt;/strong&gt; earnings increase (p=0.017), indicating that extra-household pressure drives the results.&lt;/p&gt;
&lt;p&gt;Outgoing transfers do not decline. The design leaves cash-on-hand unchanged by construction, and consistent with this, there is no significant change in the likelihood or amount of transfers from treated workers to their networks. Total outgoing transfers are if anything higher among Private account workers (p=0.049), suggesting no loss in redistribution to the network.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Tax Rate Estimation&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Combining the 11.4% treatment effect on output with a labor supply elasticity estimated from an end-of-experiment piece-rate randomization (intensive-margin elasticity of 0.17; total elasticity of approximately 1.11), the authors estimate the social tax rate for the average worker in the sample at &lt;strong&gt;9–14%&lt;/strong&gt;. For the subset who actually take up Private accounts, the implied social tax rate is &lt;strong&gt;19–23%&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Results pertain to full-time female piece-rate workers in formal cashew processing plants in Côte d&amp;rsquo;Ivoire, with average tenure of 1.7 years. Because the intervention lowers the tax only on earnings &lt;em&gt;above&lt;/em&gt; baseline (not on all earnings), the estimates do not directly capture the total distortion from eliminating all redistributive pressure. Alternative confounds — fairness/morale effects, self-control, privacy concerns, goal-setting — are each tested and ruled out as primary drivers.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-theoretical-basis-for-predicting-that-private-accounts-unambiguously-increase-labor-supply"&gt;Q1. What is the theoretical basis for predicting that Private accounts unambiguously increase labor supply?&lt;/h3&gt;
&lt;p&gt;The authors model redistributive pressure as a social tax rate τ₁ on gross earnings. The blocked account reduces this tax to τ₂ &amp;lt; τ₁ only on earnings &lt;em&gt;above&lt;/em&gt; baseline labor supply e₁, creating a kink in the budget constraint. Starting from e₁, the worker faces only a pure substitution effect (no income effect) when τ₂ falls, because her net earnings at e₁ are unchanged. Equation (2) in the paper shows formally that the income effect term drops out, and the derivative of labor supply with respect to τ₂ is unambiguously negative (i.e., reducing τ₂ increases effort). This &amp;ldquo;clean&amp;rdquo; prediction — no income effect, no ambiguity — is the central design advantage relative to simply shielding existing earnings.&lt;/p&gt;
&lt;h3 id="q2-how-do-take-up-rates-differ-between-private-and-non-private-accounts-and-what-do-workers-say-explains-the-difference"&gt;Q2. How do take-up rates differ between Private and Non-private accounts, and what do workers say explains the difference?&lt;/h3&gt;
&lt;p&gt;In Phase 2, take-up of Private accounts is 60% versus only 14% for Non-private accounts — a 77% reduction (p&amp;lt;0.001). Among workers who declined a Non-private account, 96% cite the anticipation of increased transfer requests from network members knowing about the account as an important factor in their decision. Only 5% cite any other reason. This pattern is strong direct evidence that the fear of redistribution — not other features of the accounts — drives take-up differences.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-treatment-effects-on-earnings-and-attendance-and-how-consistent-are-they-across-phases-and-subsamples"&gt;Q3. What are the treatment effects on earnings and attendance, and how consistent are they across phases and subsamples?&lt;/h3&gt;
&lt;p&gt;Pooled across both phases, the Private arm raises daily earnings by 175.9 FCFA (11.4%, p=0.012) and attendance by 6.2 percentage points (9.7%, p=0.023). In Phase 1 alone (Private vs. Control), earnings rise 11.3% (p=0.032). In Phase 2 alone (Private vs. Non-private), earnings rise 11.5% (p=0.043). Restricting to workers not previously treated in Phase 1, the effect is 12.8% (p=0.034); restricting further to workers new to the study in Phase 2 only, the effect is 17.3% (p=0.020). The authors cannot reject that effects across these three Phase 2 subsamples are statistically the same (p=0.427), ruling out sensitivity to the cross-randomized design.&lt;/p&gt;
&lt;h3 id="q4-how-does-treatment-effect-heterogeneity-support-the-redistributive-pressure-mechanism"&gt;Q4. How does treatment effect heterogeneity support the redistributive pressure mechanism?&lt;/h3&gt;
&lt;p&gt;Workers who report difficulty saving because &amp;ldquo;someone else will need it for something urgent&amp;rdquo; see earnings increase by 15.0% (p=0.018) from the Private treatment; those not reporting this difficulty see near-zero, insignificant effects (p=0.95). Workers who make transfers to acquaintances — transfers especially unlikely to reflect altruism — see earnings rise 17.5% (p=0.014). Workers with below-median baseline earnings, potentially those facing the strongest relative disincentive to work, see larger effects. Each of these heterogeneous patterns is in the direction predicted if the social tax is the operative mechanism.&lt;/p&gt;
&lt;h3 id="q5-do-the-treatment-effects-reflect-substitution-away-from-outside-earnings-or-genuine-total-income-gains"&gt;Q5. Do the treatment effects reflect substitution away from outside earnings or genuine total income gains?&lt;/h3&gt;
&lt;p&gt;No. The paper finds no treatment effects on earnings outside the factory. At baseline, 89% of workers report zero outside earnings, and on average 93% of total income comes from factory wages. Consequently, the 11.4% earnings increase represents a near-one-for-one increase in total earned income.&lt;/p&gt;
&lt;h3 id="q6-do-private-accounts-reduce-transfers-to-the-network"&gt;Q6. Do Private accounts reduce transfers to the network?&lt;/h3&gt;
&lt;p&gt;No. The design ensures that cash-on-hand is unchanged by construction — workers receive the same or slightly higher take-home cash pay (the difference is positive but insignificant). Consistent with this, neither the probability of making transfers (p=0.37) nor transfers to family (p=0.35) or non-family (p=0.93) change significantly. Total outgoing transfers in the endline survey are if anything higher in the Private arm (p=0.049, though this may partly reflect redistribution of unblocked savings). The net transfer amount is positive but insignificant (p=0.32). The authors conclude the intervention did not make others in workers&amp;rsquo; networks worse off.&lt;/p&gt;
&lt;h3 id="q7-how-do-the-authors-rule-out-morale-or-fairness-effects-as-an-explanation"&gt;Q7. How do the authors rule out morale or fairness effects as an explanation?&lt;/h3&gt;
&lt;p&gt;Treatment assignment was conducted by lottery with ID numbers drawn in front of workers, clearly dissociating it from employer favoritism. More directly, the authors test for morale effects using the 3–4 week &amp;ldquo;announcement period&amp;rdquo; between treatment disclosure and account activation. If disgruntlement among non-Private workers drove results, output should fall during this period — but estimated announcement effects are near zero (0.8% of control mean, p=0.859 in Phase 2). In contrast, effects arise immediately in the first active paycycle: earnings jump 11.4% (p=0.082) even before workers have seen any deposits occur. The fairness story also cannot explain why effects are concentrated precisely among workers who report more redistributive pressure.&lt;/p&gt;
&lt;h3 id="q8-how-do-the-authors-test-and-rule-out-self-control-as-the-primary-mechanism"&gt;Q8. How do the authors test and rule out self-control as the primary mechanism?&lt;/h3&gt;
&lt;p&gt;Self-control cannot explain why Non-private accounts — which offer the same commitment benefit — have dramatically lower take-up than Private accounts. Separately, the authors test a core prediction of time inconsistency models by surprising workers with an option to opt out of the next deposit, randomly varying whether the offer comes 4 days before payday or on payday itself. Under quasi-hyperbolic preferences, workers should be more likely to opt out on the payday itself. Counter to this prediction, 94% of workers keep their earnings in the account on payday, compared to 86% four days before — and these means are not statistically distinguishable, with the relative magnitudes actually running opposite to time inconsistency predictions.&lt;/p&gt;
&lt;h3 id="q9-how-do-the-authors-address-the-concern-that-non-private-accounts-may-raise-the-tax-rate-above-the-baseline-inflating-treatment-effect-estimates"&gt;Q9. How do the authors address the concern that Non-private accounts may raise the tax rate above the baseline, inflating treatment effect estimates?&lt;/h3&gt;
&lt;p&gt;The concern is that Non-private SMS alerts could make network members more aware of available cash than under the status quo, pushing the effective comparison above the Control level. The authors note that (a) paydays are already publicly known in this setting and workers regularly face transfer requests around them; (b) workers must physically withdraw savings from a bank after the unblock date, and can even re-block funds; and (c) the magnitude of effects when comparing Private to Control is nearly identical to the effect when comparing Private to Non-private (11.3% vs. 11.5%), suggesting the Non-private condition does not materially raise the tax above the status quo.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-authors-rule-out-privacy-concerns-rather-than-redistributive-pressure-as-the-driver-of-low-non-private-take-up-and-treatment-effects"&gt;Q10. How do the authors rule out privacy concerns (rather than redistributive pressure) as the driver of low Non-private take-up and treatment effects?&lt;/h3&gt;
&lt;p&gt;Four arguments are provided. First, Phase 1 effects (Private vs. Control, no Non-private arm) are the same magnitude as Phase 2 effects, yet Phase 1 cannot be confounded by privacy concerns. Second, among workers who refused Non-private accounts, 96% cite transfer request anticipation; none volunteer generic privacy concerns. Third, heterogeneity effects — concentrated among high-redistributive-pressure workers — have no obvious connection to privacy preferences. Fourth, two placebo SMS exercises: 95% of Non-private workers grant permission to send generic bank promotional texts, and 88% of workers who had Phase 1 Private accounts grant permission for messages about their past (already-spent) savings — indicating no inherent aversion to having some financial information shared with networks. Since these workers forgo 11.5% of full-time earnings by refusing Non-private accounts, privacy concerns alone are implausible as a full explanation.&lt;/p&gt;
&lt;h3 id="q11-how-is-the-social-tax-rate-estimated-and-what-does-the-range-look-like"&gt;Q11. How is the social tax rate estimated and what does the range look like?&lt;/h3&gt;
&lt;p&gt;The authors combine the 11.4% ITT treatment effect (used as the ratio e₁/e₂) with a compensated labor supply elasticity ζ estimated from an end-of-experiment piece-rate randomization. The piece-rate experiment (varying piece rates over four values from −15% to +30% of baseline over 6 days) yields an intensive-margin elasticity of 0.17. Using the ratio of attendance to intensive-margin effects from Table 3, the implied extensive-margin elasticity is 0.94, giving ζ ≈ 1.11. With this elasticity and assuming τ₂ = 0 (most conservative), the ITT-implied social tax rate is 9%; assuming τ₂ = 5%, it is 14%. For compliers (workers who actually take up Private accounts), the estimated rate is 19–23%. If instead the lower elasticity estimate of 0.32 (comparable to Goldberg 2016) is used, the ITT tax rate would be at least 29%.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-broader-implications-discussed-by-the-authors"&gt;Q12. What are the broader implications discussed by the authors?&lt;/h3&gt;
&lt;p&gt;The authors propose that if redistributive pressure distorts work incentives, it may also distort other costly income-generating actions: technology adoption, human capital investment, and formal sector participation. They note that 74% of workers believe taking a formal job would increase transfer requests, even though network members could also access such jobs. A speculative but highlighted policy implication is that formal safety nets (health or unemployment insurance) could reduce social tax burdens on non-recipients by absorbing demand for redistribution, potentially generating positive productivity externalities.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Social Tax&lt;/strong&gt;: The paper&amp;rsquo;s central concept. Redistributive pressure from kin and social networks is modeled as a tax rate τ₁ on gross earnings — not altruistic transfers, but transfers made under social pressure that workers would prefer to avoid. The &amp;ldquo;tax&amp;rdquo; analogy captures that the obligation is proportional to visible income and reduces the private return to earning more. The paper explicitly does not take a stance on the underlying microfoundation (risk-sharing, cultural norms, or a mix).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Blocked Savings Account&lt;/strong&gt;: A date-based savings account (implemented with Banque Populaire de Côte d&amp;rsquo;Ivoire) into which any earnings above a worker-chosen threshold are automatically deposited by the factory. Funds are inaccessible until the blocked period ends (3–9 months). Workers cannot withdraw during the period, making deposited earnings unavailable to fulfill transfer requests and therefore effectively reducing the social tax rate on earnings increases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Private vs. Non-private Treatment&lt;/strong&gt;: The paper&amp;rsquo;s key experimental contrast. A Private account&amp;rsquo;s existence is unknown to anyone in the worker&amp;rsquo;s network. A Non-private account triggers SMS messages to network members disclosing that the worker is saving and announcing when the unblock date approaches. The contrast isolates whether the shielding of income from social visibility — not the commitment device per se — drives take-up and labor supply.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Substitution Effect without Income Effect&lt;/strong&gt;: The paper&amp;rsquo;s design deliberately places the tax reduction only on earnings &lt;em&gt;above&lt;/em&gt; baseline, creating a kink in the budget constraint. Starting from the existing labor supply level, there is no change in net earnings at the margin — eliminating the income effect of a tax reduction — so any labor supply response is a pure compensated (substitution) effect. This makes any observed increase in labor supply an unambiguous signal that a distortionary social tax exists.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intent to Treat (ITT) vs. Treatment on the Treated (ToT)&lt;/strong&gt;: The ITT estimate (11.4% earnings increase) reflects the effect of being &lt;em&gt;offered&lt;/em&gt; a Private account on all offered workers, including those who did not take up. The ToT estimate — relevant for workers who actually used the accounts — implies a higher social tax rate (19–23%) because only roughly half of offered workers take up the accounts and only those workers face a materially reduced effective tax rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Compensated (Hicksian) Labor Supply Elasticity (ζ)&lt;/strong&gt;: The ratio used to infer the social tax rate from the observed treatment effect. The paper estimates ζ ≈ 1.11 (extensive margin ζₐ ≈ 0.94, intensive margin ζₑ ≈ 0.17) from an end-of-experiment piece-rate randomization. The social tax rate is recovered as τ₁ = 1 − (1−τ₂)(e₁/e₂)^(1/ζ) from Equation (5).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Piece Rate Setting&lt;/strong&gt;: Workers earn a linear piece rate for every kilogram of cashews peeled, with no fixed pay component. This setting ensures that every unit of additional effort by a worker translates directly into higher earnings, and that any observed earnings changes cleanly reflect labor supply responses rather than hour or schedule effects.&lt;/p&gt;</description></item><item><title>Traditional Institutions in Modern Times: Dowries as Pensions When Sons Migrate</title><link>https://macropaperwarehouse.com/papers/traditional-institutions-in-modern-times-dowries-as-pensions-when-sons-migrate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/traditional-institutions-in-modern-times-dowries-as-pensions-when-sons-migrate/</guid><description>&lt;p&gt;This paper asks whether dowry — a transfer from the bride&amp;rsquo;s family to the groom&amp;rsquo;s household upon marriage, prevalent throughout India — enables male migration by providing liquidity that compensates parents for the old-age support they would otherwise lose when sons leave the village. The core friction is that in patrilocal societies, sons traditionally co-reside with parents and share income in old age; migration disrupts this arrangement and introduces income-sharing frictions (limited commitment, information asymmetries, remittance costs). Dowry attenuates this friction by providing a liquid pool of resources at the time of marriage that the son can transfer to parents, lowering the net return to migration needed for a household to find migration optimal.&lt;/p&gt;
&lt;p&gt;The authors develop a collective household model in which parents and sons jointly maximize a Pareto-weighted utility function. The model yields six testable predictions: (1) net marriage transfers can flow in either direction; (2) parents are more likely to take from the dowry when sons migrate; (3) conditional on migration, the probability of parental taking increases in the son&amp;rsquo;s income and in parental bargaining power; (4) aggregate male migration rates are higher in districts with stronger historical dowry traditions; (5) migration responses to a reduction in migration costs are larger in dowry areas, provided migration rates are relatively low; and (6) parents who receive remittances from migrant sons are more likely to have also taken from the dowry.&lt;/p&gt;
&lt;p&gt;To test predictions 1–3 and the remittance auxiliary prediction, the authors collected two original datasets: a Destination Survey of 557 prime-age men in Gurugram (near Delhi) conducted in 2018, of whom 62% were migrants; and an Origin Survey of 2,541 households across 34 districts in six North Indian states conducted in 2020, covering 3,069 sons, 20% of whom were migrants. These are the first quantitative data on property rights over dowry in India. Across the Destination and Origin surveys, 45% and 27% of grooms&amp;rsquo; parents, respectively, took from the dowry on net. Parents of migrants are 27 percentage points (Destination) and 8 percentage points (Origin) more likely to take than parents of non-migrants. For migrant sons, a doubling of the son&amp;rsquo;s occupational score raises the likelihood of parental taking by 19 percentage points; no such relationship exists for non-migrants. When sons report that parents held veto power over the marriage — a proxy for parental Pareto weight — parents of migrant sons are 28 percentage points more likely to be net takers. Parents whose migrant son sends financial remittances are 17 percentage points more likely to have taken from the dowry (coefficient 0.168, SE 0.074).&lt;/p&gt;
&lt;p&gt;To test predictions 4 and 5, the authors use the Ancestral Characteristics data (Giuliano and Nunn 2018) to construct district-level measures of dowry tradition strength, validated against 1999 REDS and IHDS survey data, where a one-unit increase in the historical dowry measure is associated with 81–109% higher gross or net dowry payments. Using the NSS Round 64 migration module (2007–08), they find that the continuous dowry tradition measure is associated with a 2.7–3.7 percentage point increase in migration probability against a mean of 23.8%. For the highway construction identification strategy, the authors exploit the staggered rollout of the Golden Quadrilateral and North-South/East-West corridor (5,846+ km, $71 billion), using modern staggered-entry difference-in-differences estimators (Borusyak et al. 2021; Callaway and Sant&amp;rsquo;Anna 2020). Young men (ages 15–30) in dowry districts exhibit a large, significant increase in out-migration following highway construction with no pre-trends, while the effect for non-dowry males is indistinguishable from zero. Older males (ages 31–45) show no such effect in either group, consistent with the mechanism operating at marriage. The highway effects are concentrated in inter-district, employment-driven migration.&lt;/p&gt;
&lt;p&gt;Scope conditions: the migration-enabling mechanism operates through marriage-age liquidity and patrilocal support norms; results are specific to male migration in India. The model assumes parents and sons act collectively, matching is based on grooms&amp;rsquo; earning potential, and migration frictions cause income-sharing transfers to be infeasible when the son migrates.&lt;/p&gt;
&lt;p&gt;Q: What is the central hypothesis of the paper?
A: The hypothesis is that dowry, by providing a liquid transfer at the time of marriage, allows sons to compensate parents for the old-age support that would otherwise be lost when sons migrate. Because migration introduces frictions that prevent optimal post-migration income sharing between parents and sons, dowry lowers the minimum net return to migration required for the household to find migration optimal, thereby enabling more migration.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;Seeking&amp;rdquo; versus &amp;ldquo;Satisfied&amp;rdquo; distinction in the model, and why does it matter?
A: &amp;ldquo;Satisfied&amp;rdquo; parents are those whose own income plus the maximum feasible marriage transfer (bounded by the bride&amp;rsquo;s endowment dE when dowry is present) is at least as large as their consumption allocation under no migration; migration then Pareto-improves the household for any non-negative return R. &amp;ldquo;Seeking&amp;rdquo; parents have insufficient income plus endowment, so migration reduces their consumption unless the son&amp;rsquo;s return R exceeds a threshold B(d). Because dowry strictly increases the feasible transfer ceiling, B(d=1) ≤ B(d=0), meaning dowry converts some Seeking households into effectively Satisfied ones and lowers the migration threshold for the rest.&lt;/p&gt;
&lt;p&gt;Q: What share of grooms&amp;rsquo; parents actually take from the dowry, and how does migration status affect this?
A: In the Destination Survey (62% migrants), 45% of parents take from the dowry on net; in the Origin Survey (20% migrants), 27% do. Parents of migrants are 27 percentage points more likely to take in the Destination Survey and 8 percentage points more likely in the Origin Survey, consistent with the model prediction that migration increases net taking.&lt;/p&gt;
&lt;p&gt;Q: How does the son&amp;rsquo;s earnings level affect parental taking, and does this pattern hold for non-migrants?
A: For migrant sons, a 100% increase in the son&amp;rsquo;s occupational score increases the likelihood of parents taking by 19 percentage points. For non-migrant sons, the son&amp;rsquo;s occupational score has no meaningful association with taking. This asymmetry is consistent with prediction 3: when migration occurs and the alpha income-sharing channel is shut down, parents with higher-income migrant sons have a higher relative marginal return to consumption and thus take more of the dowry.&lt;/p&gt;
&lt;p&gt;Q: What is the remittance auxiliary prediction, and is it borne out in the data?
A: The model predicts that parents who receive remittances from migrant sons should also be more likely to have taken from the dowry, because households first exhaust the costless dowry transfer before making costly or risky remittances — so remittance-receiving parents are precisely those Seeking households where dowry was already taken. The data confirm this: parents whose migrant son sends financial remittances are 17 percentage points more likely to have taken from the dowry (coefficient 0.168, SE 0.074, significant at 5%) compared to parents of migrants who do not remit.&lt;/p&gt;
&lt;p&gt;Q: How is the district-level dowry tradition measure constructed and validated?
A: The measure merges the Giuliano and Nunn (2018) Ancestral Characteristics data — which uses ethnographic sources to estimate the share of each district&amp;rsquo;s current population belonging to historically dowry-practicing groups — with district-level demographic data. Validation against the 1999 REDS shows that a one-unit increase in the historical dowry measure is associated with 81% higher gross dowry payments and 109% higher net dowry payments without region fixed effects, with a still-significant 79% for net dowry including region fixed effects. Additional validation in the IHDS confirms the historical measure predicts gold payments at marriage (coefficient 0.152 without state fixed effects, 0.185 with state fixed effects).&lt;/p&gt;
&lt;p&gt;Q: What is the association between historical dowry traditions and migration in nationally representative data?
A: Using the NSS Round 64 migration module (2007–08) for males aged 15–45, against a mean migration rate of 23.8%, the continuous dowry measure is associated with a 2.66 percentage point increase in migration probability with no controls (significant at 1%), and 3.67 percentage points with full controls including state fixed effects, year-of-birth fixed effects, caste fixed effects, distance controls, and education controls (significant at 5%).&lt;/p&gt;
&lt;p&gt;Q: What is the highway construction identification strategy, and what does it show?
A: The authors exploit the staggered construction timing of the Golden Quadrilateral and NS-EW highway corridors (beginning 1999, 5,846+ km, $71 billion investment) across Indian districts, assembling new data on district-level construction timing from a complete capital projects database. Using staggered-entry event study estimators robust to heterogeneous treatment effects, they separately estimate highway effects in districts with and without strong dowry traditions. For young men aged 15–30, dowry districts show a large, significant increase in out-migration after highway construction with no pre-trends; non-dowry districts show an effect indistinguishable from zero. Older men (31–45) show no significant effect in either group.&lt;/p&gt;
&lt;p&gt;Q: Why is the age heterogeneity (15–30 vs. 31–45) in the highway results important for the mechanism?
A: The model predicts that dowry&amp;rsquo;s migration-enabling role operates at the time of marriage, when the liquid transfer is made. Men aged 31–45 at the time of highway construction would largely have already been married before the roads were built, so they cannot retroactively benefit from the new liquidity channel. Young men (15–30) are near or below marriage age and can time their marriages and migration decisions in response to reduced migration costs. The null result for older men and the strong result for younger men together confirm the marriage-time liquidity channel.&lt;/p&gt;
&lt;p&gt;Q: Why is the highway effect concentrated in inter-district rather than intra-district migration?
A: The Golden Quadrilateral connects districts to other districts, and the model&amp;rsquo;s mechanism relies on migration creating income-sharing frictions that are more severe at longer distances. Intra-district moves are shorter, less likely to disrupt co-residence and informal support arrangements, and less likely to require the dowry&amp;rsquo;s compensatory role. The concentration of effects in inter-district migration is directly consistent with the proposed channel.&lt;/p&gt;
&lt;p&gt;Q: How does the paper address concerns about pre-trends and robustness in the highway analysis?
A: The event study plots show no pre-trends in migration for either dowry or non-dowry districts prior to highway construction. Robustness checks include additional geographic controls, caste-by-year fixed effects, time-varying cultural controls, the alternative Callaway-Sant&amp;rsquo;Anna estimator, adjusted age distributions, and varying dowry tradition cutoffs at 1%, 10%, and 25% thresholds. Results are stable across these specifications.&lt;/p&gt;
&lt;p&gt;Q: What do the theory and evidence imply about the modern transformation of dowry&amp;rsquo;s function?
A: While dowry historically served as a pre-mortem bequest to the bride adapted to patrilocal society, the modern practice has evolved so that grooms&amp;rsquo; parents frequently capture the transfer. The evidence is consistent with this reallocation of property rights serving a new function: providing parents with a pension substitute when sons migrate and traditional co-residential support breaks down. The authors speculate this functional evolution may partly explain why dowry prevalence has grown despite legal bans, as declining patrilocality creates rising demand for this type of intergenerational transfer mechanism.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: The paper suggests that policies discouraging dowry — which has many well-documented negative consequences including intimate partner violence, female infant mortality, and adverse resource allocation — may be more effective if paired with expansions of formal pension programs or other mechanisms for old-age support. Without such alternatives, eliminating dowry could inadvertently reduce male migration and associated economic development benefits because the migration-enabling liquidity function of dowry would go unfilled.&lt;/p&gt;
&lt;p&gt;Q: Does the mechanism apply equally to households with both sons and daughters?
A: The theoretical appendix shows that in a household with a son and a daughter, the daughter&amp;rsquo;s dowry outflow partially offsets the son&amp;rsquo;s inflow, reducing but not eliminating the migration-enabling effect. However, the net aggregate effect on male migration remains positive because more sons live in households where sons outnumber daughters, so the dowry inflow for the son exceeds the outflow on average across the population.&lt;/p&gt;
&lt;p&gt;Dowry (in the paper&amp;rsquo;s sense): A transfer from the bride&amp;rsquo;s family accompanying marriage that in the modern Indian context is liquid at the time of the wedding and over which grooms&amp;rsquo; parents frequently exercise property rights — distinct from the traditional anthropological conception of dowry as a pre-mortem bequest to the bride.&lt;/p&gt;
&lt;p&gt;Net Taker: A groom&amp;rsquo;s parent who receives a positive net transfer from the son&amp;rsquo;s dowry (tau &amp;gt; 0 in the model), meaning the flow of dowry resources is from the son/bride&amp;rsquo;s side to the groom&amp;rsquo;s parents.&lt;/p&gt;
&lt;p&gt;Seeking vs. Satisfied parents: Model categories distinguishing parents whose consumption needs can be met from own income plus the maximum feasible marriage transfer (Satisfied, no migration distortion) from those whose needs cannot (Seeking, requiring a minimum migration return threshold B(d) &amp;gt; 0 for migration to be household-optimal).&lt;/p&gt;
&lt;p&gt;Migration friction (alpha = 0 under migration): The modeling assumption that income-sharing transfers between migrant sons and parents are infeasible or prohibitively costly due to limited commitment, information asymmetries, and remittance costs — the friction that dowry&amp;rsquo;s lump-sum transfer at marriage is designed to circumvent.&lt;/p&gt;
&lt;p&gt;Ancestral Characteristics dowry measure: The district-level variable from Giuliano and Nunn (2018) measuring the share of the current population belonging to historically dowry-practicing ethnic groups, used as a proxy for the strength of local dowry traditions.&lt;/p&gt;
&lt;p&gt;Patrilocality: The residential norm in which sons remain with or near their parents after marriage and provide old-age support — the norm whose breakdown via migration creates the income-sharing friction that dowry helps resolve.&lt;/p&gt;
&lt;p&gt;Pareto weight (theta): The weight assigned to parents&amp;rsquo; utility in the collective household problem, capturing parental bargaining power; empirically proxied by whether sons report that parents held veto power over the marriage choice.&lt;/p&gt;</description></item><item><title>Vanguard: Black Veterans and Civil Rights After World War I</title><link>https://macropaperwarehouse.com/papers/vanguard-black-veterans-and-civil-rights-after-world-war-i/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/vanguard-black-veterans-and-civil-rights-after-world-war-i/</guid><description>&lt;p&gt;This paper provides the first causal evidence on how military service shaped Black civil rights activism in the aftermath of World War I. The research question is whether random induction into the segregated National Army caused Black men to join the nascent NAACP and become prominent community leaders during the New Negro era. The authors leverage the WWI draft lottery — in which each registrant&amp;rsquo;s unique serial number was drawn from a bowl to determine induction order — as an instrument for military service, a source of exogenous variation not previously exploited in the literature.&lt;/p&gt;
&lt;p&gt;To support this analysis, Ang and Chinoy construct an unusually rich dataset by digitizing nearly one million Black draft registration cards from the first registration (June 17, 1917), linking them through the 1930 full-count census to 233,517 NAACP member observations across 227 branches from 1912 to 1940, and supplementing with Veterans Administration records, Army Transport Service passenger lists, and biographical dictionaries of prominent African Americans. The instrument — serial number percentile within draft board and race (SNP%) — is validated against all observed pre-draft registrant characteristics and yields a first-stage F-statistic of 1,051 in the preferred specification.&lt;/p&gt;
&lt;p&gt;The main finding is that Black men randomly induced to serve in the military were nearly three times more likely to join the NAACP than observably similar registrants from the same draft board (TSLS coefficient 0.0219, se = 0.0049, against a sample mean NAACP participation rate of 0.8%). The authors estimate that the draft induced more than 10,000 Black men to join the NAACP in total. Military service also raised the probability of appearing in biographical dictionaries of historically prominent African Americans by a factor of roughly 1.6 (TSLS coefficient 0.0027, se = 0.0012, sample mean 0.17%). These results are robust to alternative instruments, flexible polynomial specifications of SNP%, state-year fixed effects, and alternative veteran-status measures from VAMI and ATS records. They are also not explained by differential residential mobility: adding controls for interstate and North-South migration leaves the main coefficient essentially unchanged (0.0217-0.0218).&lt;/p&gt;
&lt;p&gt;In contrast, TSLS estimates for all socioeconomic outcomes — literacy, home ownership, employment, census-predicted income, actual 1940 income, and educational attainment — are small and insignificant, ruling out human capital acquisition as a mechanism. Club involvement measured in the census is likewise unaffected, indicating that NAACP membership reflects specifically civil rights activism rather than generically greater social participation.&lt;/p&gt;
&lt;p&gt;The mechanism the paper identifies is experienced discrimination. Effects on NAACP participation increase monotonically with the racial gap in induction rates across draft boards (significant at p = 0.01). Effects are large and significant for men assigned to camps that restricted Black soldiers&amp;rsquo; access to military training (coefficient 0.0351, se = 0.0104) and to officer promotion (coefficient 0.0360, se = 0.0111), and are large for men in both restriction types simultaneously (coefficient 0.0367, se = 0.0114). In contrast, men attending less discriminatory camps show small and insignificant effects. Among the two all-Black combat divisions, NAACP participation is highest for veterans of the 92nd Division — subjected to constant racial abuse under U.S. command — and lower for the 93rd Division, which served under more hospitable French command. Previously unstudied veteran surveys from Virginia and Connecticut corroborate this narrative: respondents from camps with training and promotion restrictions were more than twice as likely to mention racial injustice, and mentions of injustice were more predictive of postwar civic engagement than any other survey theme.&lt;/p&gt;
&lt;p&gt;The scope of the paper is Black male registrants in the first WWI draft registration (men aged 21-30 as of June 17, 1917), linked to a sample of approximately 300,000 in the 1930 census. Effects are attenuated for men from counties with greater racial hostility — proxied by Confederate state status, Confederate monument density, and county lynching rates — consistent with the interpretation that activism was more feasible in less repressive environments.&lt;/p&gt;
&lt;p&gt;Q: What is the core identification strategy and why was it not feasible to use it before this paper?
A: The paper uses each Black registrant&amp;rsquo;s serial number percentile within his draft board and racial group (SNP%) as an instrument for WWI military service. Unlike the WWII and Vietnam drafts, which used birthday-based lotteries, the WWI lottery assigned induction order by drawing unique serial numbers from a bowl, making serial number rank the source of quasi-random variation. This source had never been exploited in the literature, partly because the serial numbers had to be hand-captured from digitized draft card images.&lt;/p&gt;
&lt;p&gt;Q: How strong is the first stage, and was the lottery truly random?
A: The first-stage F-statistic is 1,051, and a ten-percentile decrease in SNP% is associated with a 34.5 percentage point increase in the probability of serving. Bivariate serial numbers show some non-random patterns — nine of 13 pre-draft characteristics correlate with raw SN% — likely because some Southern boards inflated numbers for white registrants. Conditioning on board fixed effects and using SNP% within board-race cells eliminates these correlations; Panel B of Appendix Table A1 shows the largest standardized coefficient falls to 0.006.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the effect on NAACP membership and how does the causal estimate compare to a naive OLS?
A: The TSLS coefficient is 0.0219 (se = 0.0049) against a sample mean of 0.8%, implying roughly a threefold increase in NAACP membership. The OLS estimate of 0.0116 understates the causal effect, consistent with the marginal man induced by the lottery being observationally weaker than infra-marginal volunteers.&lt;/p&gt;
&lt;p&gt;Q: Does the effect reflect simply that veterans moved to Northern cities where NAACP branches were more accessible?
A: No. Adding indicators for interstate migration and North-South migration leaves the TSLS coefficient essentially unchanged at 0.0218 and 0.0217, respectively. The Great Migration channel is thus not the operative mechanism.&lt;/p&gt;
&lt;p&gt;Q: Did military service improve Black veterans&amp;rsquo; economic outcomes?
A: TSLS estimates for literacy, home ownership, employment, census-predicted income, actual 1940 income, and educational attainment are all small and statistically insignificant. This contrasts sharply with evidence on Black veterans of WWII and Korea (Greenberg et al., 2022) and is consistent with the documented absence of meaningful postwar benefits or training for Black WWI soldiers.&lt;/p&gt;
&lt;p&gt;Q: If it was not human capital or migration, what mechanism does the paper establish?
A: The primary mechanism is exposure to institutional discrimination during military service. Three distinct empirical patterns converge: (1) effects increase monotonically with draft board racial disparities in induction rates; (2) effects are large and significant for men at camps that denied training and promotion, and near zero for men at less discriminatory camps; (3) veteran survey mentions of racial injustice are more common among men from discriminatory camps and are more predictive of postwar NAACP membership than any other survey theme.&lt;/p&gt;
&lt;p&gt;Q: How do the two all-Black combat divisions differ in their postwar NAACP participation, and what does this reveal?
A: Veterans of the 92nd Division, who fought under U.S. command amid constant racial abuse, show the highest NAACP participation rates. Veterans of the 93rd Division, who fought under French command and were received with relative hospitality, show lower (though not statistically significantly lower) participation. Since both divisions received similar formal training and neither group shows socioeconomic gains, the differential reflects discrimination exposure rather than skill acquisition.&lt;/p&gt;
&lt;p&gt;Q: What is the quantitative scale of the effect for the most discriminatory camps?
A: For men assigned to camps with restrictions on both training and promotion, the TSLS coefficient on NAACP membership is 0.0367 (se = 0.0114) — more than 1.5 times the average estimate of 0.0219. Men at camps without restrictions show coefficients that are small and statistically insignificant.&lt;/p&gt;
&lt;p&gt;Q: How does county-level racial hostility moderate the effect?
A: The effects of military service on NAACP membership are larger — more positive — for men from counties with fewer Confederate monuments, lower lynching rates, and non-Confederate state status. This is interpreted as evidence that activism in response to discriminatory military experiences was more feasible in less racially hostile local environments, rather than as evidence that discrimination exposure was lower.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s aggregate policy implication regarding the scale of the draft&amp;rsquo;s effect on the civil rights movement?
A: The authors estimate that the WWI draft induced more than 10,000 Black men to join the NAACP. Veterans accounted for nearly 15% of all male NAACP members, against roughly 8% of Black male adults in the population, and were significantly more likely to appear in biographical dictionaries of prominent African Americans. The draft thus constituted a sizable and measurable contribution to the organizational vanguard of the early civil rights movement.&lt;/p&gt;
&lt;p&gt;Q: How does the paper contribute to the economics of discrimination beyond documenting discriminatory behavior by majority actors?
A: Most economics research on discrimination studies the conduct of white decision-makers (e.g., racial bias in hiring, lending, or bail). This paper examines how experiences of discrimination reshape the political behavior and aspirations of the minority group itself. The results show that institutional betrayal — systematic exclusion, degradation, and denial of training — generated deep discontent that translated into aggressive political mobilization, a dynamic the authors trace through subsequent episodes including the WWII Double V campaign and responses to police killings.&lt;/p&gt;
&lt;p&gt;Serial number percentile within draft board and race (SNP%): The instrument constructed by the authors. Each WWI registrant received a serial number from 1 to the size of his draft board; those numbers were drawn in random order to determine induction priority. SNP% measures where a registrant fell in that draw relative to others in his board and racial group, and serves as the source of quasi-random variation in veteran status.&lt;/p&gt;
&lt;p&gt;New Negro era: The period of invigorated Black political and cultural assertiveness following WWI, characterized by renewed racial pride, economic independence, and progressive politics. The movement spanned the Harlem Renaissance, the Universal Negro Improvement Association, the American Negro Press, and the Brotherhood of Sleeping Car Porters, and represented a rejection of the &amp;ldquo;conservatism, parochialism, and political accommodationism&amp;rdquo; of older Black leaders.&lt;/p&gt;
&lt;p&gt;Draft board racial gap: The authors&amp;rsquo; measure of draft board discrimination, defined as the difference in induction rates between Black and white registrants within a given draft board. The interquartile range spans roughly 0 to 20 percentage points, with a notable fraction of boards exhibiting gaps exceeding 30 percentage points.&lt;/p&gt;
&lt;p&gt;Camp discrimination: The denial of military training and officer promotion opportunities to Black soldiers, documented in War Department reports by military intelligence officers tasked with monitoring the treatment of Black soldiers. The paper classifies each camp as restricted or unrestricted on each dimension and uses this classification to estimate heterogeneous treatment effects.&lt;/p&gt;
&lt;p&gt;Institutional betrayal: The paper&amp;rsquo;s characterization of the U.S. government&amp;rsquo;s treatment of Black WWI soldiers — drafting them at higher rates than whites, denying them training and promotion, and assigning them to menial labor — as generating a profound sense of injustice that motivated postwar political activism rather than loyalty or accommodation.&lt;/p&gt;
&lt;p&gt;NAACP membership as civil rights activism proxy: The paper uses dues-paying membership in local NAACP branches as its primary quantitative measure of civil rights participation. Membership involved active financial cost (annual fees of $1 to $10 at a time when median Black family income was below $500), exposure to harassment and violence in the South, and participation in local protest and legal advocacy, distinguishing it from passive civic engagement.&lt;/p&gt;</description></item><item><title>Voluntary Minimum Wages: The Local Labor Market Effects of National Retailer Policies</title><link>https://macropaperwarehouse.com/papers/voluntary-minimum-wages-the-local-labor-market-effects-of-national-retailer-policies/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/voluntary-minimum-wages-the-local-labor-market-effects-of-national-retailer-policies/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper studies the labor market effects of voluntary minimum wages (VMWs) — company-wide, publicly announced wage floors set by large private employers — in the U.S. low-wage retail and service sector from 2014 to 2023. The central questions are: (1) How do VMWs affect wages and employment at the adopting large retailers? (2) Do VMWs generate wage spillovers to other employers in shared local labor markets?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Setting&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors use anonymized payroll data obtained from a large U.S. credit bureau, covering the wage distributions and employment of over 4,000 firms and approximately 18 million hourly workers (roughly 22–24% of the U.S. hourly workforce) from January 2013 to August 2023. The database is skewed toward retail and service sectors: over a third of covered workers are in retail, and over half in retail and services combined. Critically, the data also include worker flow information — records of individual workers moving between firms — enabling the authors to define shared labor markets via actual employment transitions rather than broad geographic or industry proxies.&lt;/p&gt;
&lt;p&gt;The sample of VMW events consists of &lt;strong&gt;20 voluntary minimum wage policies across 5 large retailers&lt;/strong&gt; (each with over 150,000 employees nationally), restricted to events with no other major wage policy within six months before or after the focal event. Voluntary minimum wage announcements were identified from an inventory maintained by the National Employment Law Project and independently verified through media sources, then matched to anonymized companies using employer size, industry, and observed shifts in the wage distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identification Strategy&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors adapt the &lt;strong&gt;gap design&lt;/strong&gt; from the national minimum wage literature. For each company-by-commuting-zone (CZ) cell, the &amp;ldquo;gap&amp;rdquo; measures the percent increase in average hourly wages that would be required to bring all workers in the area up to the company&amp;rsquo;s new voluntary minimum. The gap is averaged over months −6 to −3 before the event (months −3 to −1 serve as a built-in placebo-in-time check). This variation in bite across CZs — arising because the same nominal VMW level implies different wage increases depending on local wage distributions — is combined with a stacked event study across 20 VMW events. Spillover effects are estimated by regressing log average wages at non-policy establishments on the large retailer&amp;rsquo;s CZ-level gap measure, progressively narrowing the definition of &amp;ldquo;labor market&amp;rdquo; from: (i) all non-policy establishments in the same CZ, to (ii) establishments in industries connected to the large retailer by worker flows (15 three-digit NAICS industries), to (iii) specific establishments with documented pre-event worker flows to or from the large retailer (&amp;ldquo;connected establishments&amp;rdquo;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Own effects:&lt;/em&gt; For $15 VMW events, moving from a CZ gap of 0 to a gap of 1 is associated with an approximately 88 log point increase in average hourly wages in the six months after adoption. Given that the average establishment-level gap for $15 VMWs is 0.11, the implied average wage increase is approximately 10.45% (the authors&amp;rsquo; estimate is 9–10%, consistent with small wage increases even in zero-gap comparison areas). Employment of workers earning under $30 per hour rose by 4.62% after $15 VMW events, 2.01% after major events (affecting ≥30% of workforce), and 1.25% across all 20 events. These employment increases are &lt;strong&gt;entirely attributable to reduced separations&lt;/strong&gt; rather than new hiring: separation rates fell by 0.42, 0.57, and 1.09 percentage points after all, major, and $15 VMW events respectively — equivalent to reductions of 6.57%, 8.73%, and 15.33% relative to pre-period means. Separations specifically to other database companies fell by 0.07–0.19 percentage points (5.63–13.48% relative to base rates). If anything, new hiring fell modestly after VMW adoption. Total monthly base pay and gross compensation both rose after VMWs, indicating increased total take-home pay without compensatory reductions in hours or bonuses. The total employment elasticity with respect to wages ranges from approximately 0.35 to 0.45, while the quit elasticity is 2.20–2.38 (consistent with dynamic monopsony models in which the labor supply elasticity is twice the quit elasticity).&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Spillover effects:&lt;/em&gt; Across all three definitions of the labor market, the paper estimates &lt;strong&gt;precise, economically negligible cross-employer wage spillovers&lt;/strong&gt; in the six months following VMW events. Cross-employer wage elasticities are statistically indistinguishable from zero across all specifications. Among the most narrowly defined sample — establishments with documented pre-event worker flows to or from the large retailer — the upper bound of the confidence interval rules out spillovers greater than 0.2% of wages. No wage spillovers are detected for new hires at non-policy establishments either. These null results are confirmed over a 12-month post-event horizon for the subsample of events with no other major policy nearby.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Mechanism:&lt;/em&gt; The reason for negligible spillovers is that VMWs reduced labor market churn rather than expanding the large retailer&amp;rsquo;s total employment. Hiring away from large retailers by connected non-policy firms falls after VMW adoption — consistent with fewer separations to recruit from — but &lt;strong&gt;overall hiring by non-policy firms does not decline&lt;/strong&gt;, as these firms substitute toward other hiring sources. This substitutability across new hire sources in a thick market is the proximate explanation for the absence of wage pressure on competitor firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Results pertain to large national retailers (&amp;gt;150,000 employees) operating in U.S. commuting zones during 2014–2023. The database covers only employers large enough to participate in credit bureau income verification; smaller employers (representing over 75% of U.S. hourly workers by the BLS comparison) are not observed, and the authors caution that spillover effects on smaller firms cannot be assessed. The authors also explicitly note that their null local spillover results do not rule out national-level strategic wage-setting dynamics — the rapid sequential adoption of VMWs across major retailers may reflect national-level competition rather than local market competition.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-exactly-are-voluntary-minimum-wages-and-how-do-they-differ-from-statutory-minimum-wages"&gt;Q1. What exactly are &amp;ldquo;voluntary minimum wages&amp;rdquo; and how do they differ from statutory minimum wages?&lt;/h3&gt;
&lt;p&gt;Voluntary minimum wages (VMWs) are company-wide, publicly announced wage floors set unilaterally by private employers, typically well above the applicable statutory (federal, state, or local) minimum. Unlike statutory minimums, which bind all employers in a jurisdiction, VMWs apply only to the announcing company across all of its geographic operations in the U.S. The paper studies VMWs adopted by retailers with over 150,000 workers, which include wage floors at levels such as $9, $10, $12, and $15 per hour. $15 VMWs were adopted at a time when few states or localities had yet reached that threshold, meaning the policy bit into the company wage distribution far more deeply than prevailing statutory floors.&lt;/p&gt;
&lt;h3 id="q2-how-were-vmw-events-identified-and-matched-to-anonymized-firms-in-the-payroll-database"&gt;Q2. How were VMW events identified and matched to anonymized firms in the payroll database?&lt;/h3&gt;
&lt;p&gt;VMW events were identified from a database maintained by the National Employment Law Project and verified through an independent review of business news articles. These publicly reported announcements were then matched to the anonymized companies in the credit bureau payroll database using employer size, industry, and the timing of observed shifts in the firms&amp;rsquo; wage distributions. An additional three events were identified directly from data: months where the share of workers earning below a given wage level dropped by at least 15 percentage points (for non-$15 events) or 10 percentage points (for $15 events) while the share at exactly that wage bin jumped by at least 10–20 percentage points. The final sample of 20 events was restricted to those with no other major wage policy in the six months before or after.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-gap-design-work-and-why-does-it-improve-on-the-fraction-affected-approach"&gt;Q3. How does the gap design work and why does it improve on the fraction-affected approach?&lt;/h3&gt;
&lt;p&gt;The gap for a given company, commuting zone, and time period is defined as the total wage increase needed to bring all sub-$30 workers up to the company minimum, divided by total wage costs — formally a labor-share-weighted average shortfall from the new minimum across wage bins. The gap leverages more cross-sectional variation in treatment intensity than the simple fraction of workers below the minimum: for a $15 VMW, an area where all workers earn $10 has a gap of 0.50 while an area where all earn $12 has a gap of 0.25. The gap is averaged over months −6 to −3 before the event. The period months −3 to −1 then serve as a placebo window: genuine VMW effects should appear only after the policy&amp;rsquo;s adoption month, not during the period immediately after the gap is measured. If instead the regression picks up mean reversion in noisy wage data, spurious effects would appear in months −3 to −1 rather than at event time 0.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-magnitude-of-the-wage-effect-on-the-large-retailers-themselves"&gt;Q4. What is the magnitude of the wage effect on the large retailers themselves?&lt;/h3&gt;
&lt;p&gt;For $15 VMW events, the stacked event study estimates that moving from a gap of 0 to a gap of 1 is associated with an approximately 88 log point increase in average hourly wages beginning exactly in the month of policy adoption. Given the average establishment-level gap of 0.11 for $15 VMWs, this implies the average establishment raised wages by approximately 9–10% (the authors compute 10.45% from the average gap, consistent with a slight dampening because zero-gap CZs experienced marginally higher wages too). Wage increases are confirmed persistent at 12 months in robustness checks. For all 20 VMW events pooled, effects are somewhat smaller commensurate with the lower average bite.&lt;/p&gt;
&lt;h3 id="q5-how-did-vmws-affect-total-employment-and-its-components-at-the-large-retailers"&gt;Q5. How did VMWs affect total employment and its components at the large retailers?&lt;/h3&gt;
&lt;p&gt;After $15 VMW events, log total employment of sub-$30 workers rose by 4.62%; after major VMW events (≥30% bite), 2.01%; after all 20 events, 1.25%. The increases are entirely driven by retention gains. Separation rates fell by 1.09 percentage points after $15 VMWs, 0.57 p.p. after major events, and 0.42 p.p. after all events — translating to reductions of 15.33%, 8.73%, and 6.57% relative to pre-period means. Separations to other database companies specifically fell by 0.07–0.19 percentage points (5.63–13.48% relative to the base mean). New hiring — measured as year-on-year log change in hires to control for seasonality — fell after VMW adoption, consistent with a reduced need to replace departing workers.&lt;/p&gt;
&lt;h3 id="q6-what-do-the-labor-supply-elasticities-implied-by-the-vmw-results-look-like"&gt;Q6. What do the labor supply elasticities implied by the VMW results look like?&lt;/h3&gt;
&lt;p&gt;The total employment elasticity with respect to wages ranges from approximately 0.35 to 0.45 across the three event groupings. Under standard dynamic monopsony models, the labor supply elasticity facing the firm equals twice the quit elasticity in steady state (Manning, 2003). The quit elasticity — derived by dividing the proportional reduction in separations by the log wage increase — ranges from 2.20 to 2.38, consistent with the earlier monopsony-based case study of Ford&amp;rsquo;s $5 workday (Raff and Summers, 1987) and implying substantial firm-level wage-setting power.&lt;/p&gt;
&lt;h3 id="q7-did-vmws-increase-total-take-home-pay-or-were-wage-gains-offset-by-reductions-in-hours-or-bonuses"&gt;Q7. Did VMWs increase total take-home pay or were wage gains offset by reductions in hours or bonuses?&lt;/h3&gt;
&lt;p&gt;The paper examines log average monthly base pay and log average gross compensation (which includes bonuses and overtime) as additional outcomes. Both measures rose after $15 VMW events, indicating that the wage floor increase translated into genuine improvements in total take-home pay without compensatory reductions in hours or other non-wage compensation. The monthly gross pay series is an average over calendar year-to-date months, so increases appear gradually rather than as a sharp jump at the adoption month; nevertheless the upward trend is evident and consistent.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-estimated-spillover-effects-on-wages-at-non-policy-employers"&gt;Q8. What are the estimated spillover effects on wages at non-policy employers?&lt;/h3&gt;
&lt;p&gt;Across all three definitions of the labor market — all non-policy establishments in the same CZ, establishments in the 15 connected industries in the same CZ, and establishments with documented pre-event worker flows — the estimated cross-employer wage effects are precise zeros. The stacked event study in the post-period shows coefficients centered on zero with small confidence intervals. The difference-in-differences cross-employer wage elasticity (instrumenting the large retailer&amp;rsquo;s wage change with the gap) is also indistinguishable from zero. Among the most exposed connected establishments, the point estimate is slightly positive but economically negligible; the upper confidence interval bound rules out spillovers greater than 0.2%. Results are confirmed over a 12-month horizon for the clean-event subsample.&lt;/p&gt;
&lt;h3 id="q9-could-the-null-spillover-result-reflect-mean-reversion-bias-rather-than-a-true-zero"&gt;Q9. Could the null spillover result reflect mean reversion bias rather than a true zero?&lt;/h3&gt;
&lt;p&gt;The authors address this concern explicitly. For the policy-company gap design, they build in a placebo-in-time check by measuring the gap over months −6 to −3 and checking that no wage effects appear in months −3 to −1. For the non-policy spillover analysis, they also examine an alternative treatment variable — the gap between non-policy establishments&amp;rsquo; wages and the large retailer&amp;rsquo;s new VMW — and find evidence of mean reversion: wages begin rising in the pre-period in the direction of this gap measure. They correct for this by detrending post-period estimates using a linear extrapolation of the pre-period trend. After detrending, spillover effects remain indistinguishable from zero.&lt;/p&gt;
&lt;h3 id="q10-why-are-spillover-effects-so-limited-if-the-large-retailer-is-drawing-fewer-workers-away-from-competitors"&gt;Q10. Why are spillover effects so limited if the large retailer is drawing fewer workers away from competitors?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s mechanism analysis shows that while the probability of a non-policy firm hiring a worker from the large retailer falls after a VMW event (consistent with fewer separations to recruit from the large retailer), the &lt;strong&gt;overall rate of hiring by non-policy firms does not decline&lt;/strong&gt;. Non-policy firms substitute toward other hiring sources — primarily other non-policy companies — rather than hiring fewer workers overall. This substitutability across recruiting sources in a thick labor market mutes the competitive pressure on competitor wages: since non-policy firms can replace the reduced flow from VMW companies with workers from other sources without changing total employment, they face no pressure to raise wages.&lt;/p&gt;
&lt;h3 id="q11-how-do-the-results-differ-when-focusing-on-czs-where-the-large-retailer-accounts-for-a-larger-employment-share"&gt;Q11. How do the results differ when focusing on CZs where the large retailer accounts for a larger employment share?&lt;/h3&gt;
&lt;p&gt;The authors test whether larger local market presence amplifies spillovers by splitting the sample at the median employment share of the large retailer in the CZ. They find no evidence of positive wage spillovers even in CZs where the large retailer&amp;rsquo;s employment share is above the median, confirming that neither local market size nor market concentration is a mechanism for spillover transmission in this setting.&lt;/p&gt;
&lt;h3 id="q12-how-do-these-vmw-spillover-results-compare-to-prior-evidence-on-employer-wage-setting-spillovers"&gt;Q12. How do these VMW spillover results compare to prior evidence on employer wage-setting spillovers?&lt;/h3&gt;
&lt;p&gt;The main prior U.S. evidence (Staiger et al., 2010) studied a federally mandated wage increase at Veterans Affairs hospitals and found a cross-establishment wage elasticity of approximately 0.19 for registered nurses at neighboring hospitals. The authors note two key differences: first, the VA policy increased both wages and employment at treated facilities, whereas VMWs primarily reduced separations without increasing hiring, so the supply of workers to competitor firms was not squeezed. Second, the market for low-wage retail and service workers is likely thicker (more potential hires available) than the market for registered nurses, allowing competitors to substitute hiring sources without bidding up wages.&lt;/p&gt;
&lt;h3 id="q13-what-do-the-null-local-spillover-results-imply-about-national-level-wage-dynamics"&gt;Q13. What do the null local spillover results imply about national-level wage dynamics?&lt;/h3&gt;
&lt;p&gt;The authors explicitly caution against reading the null local spillover result as implying VMWs have no broader effect on the low-wage labor market. The rapid and successive adoption of VMWs across major retailers during 2021–2022 could reflect national-level strategic wage-setting competition — firms mimicking each other&amp;rsquo;s announcements in an arms-race dynamic during tight labor markets — rather than local competitive transmission. The paper does not test for national-level strategic interactions and calls for further research on this dimension.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Voluntary Minimum Wage (VMW):&lt;/strong&gt; A company-wide, publicly announced wage floor set unilaterally by a private employer, applying across all of the firm&amp;rsquo;s geographic operations in the U.S., typically well above applicable statutory minimums. Distinct from legally mandated minimum wages in that they bind only the announcing firm and arise from the firm&amp;rsquo;s own strategic or reputational motivations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gap Measure:&lt;/strong&gt; Borrowed from the national minimum wage literature (Card, 1992; Draca et al., 2011), this is the percent increase in a firm&amp;rsquo;s average hourly wage that would be required to bring all workers in a given commuting zone up to the company&amp;rsquo;s new voluntary minimum. Formally the labor-share-weighted average shortfall from the VMW across sub-$30 wage bins. A gap of 0 means no workers fall below the new minimum; a gap of 1 means all workers would need to be raised to the minimum, doubling the average wage. Used as a continuous treatment variable capturing the local bite of the policy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stacked Event Study:&lt;/strong&gt; An empirical design in which a separate 12-month panel (6 months pre- and post-event) is constructed for each of the 20 VMW events, these datasets are stacked, and the effect of the continuous gap treatment is estimated jointly across all events, with event-specific indicators interacting all regressors to allow each event to have its own intercept.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Placebo-in-Time Check:&lt;/strong&gt; A robustness test built into the gap design by computing the gap over months −6 to −3 and verifying that wage effects do not appear in months −3 to −1 (the period between gap measurement and VMW adoption). Genuine policy effects should materialize at the adoption month; spurious effects driven by mean reversion in noisy wage data would appear in months −3 to −1 because the gap would mechanically predict wage reversion toward the mean in the period immediately following its measurement.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Connected Establishments / Poaching and Feeder Establishments:&lt;/strong&gt; Specific firm-by-CZ cells identified as sharing a labor market with the large retailer via actual worker flows. &amp;ldquo;Poaching establishments&amp;rdquo; hired at least one worker from the large retailer in the 12 months before the VMW event. &amp;ldquo;Feeder establishments&amp;rdquo; had at least one worker subsequently hired by the large retailer in the same pre-period. These are the most narrowly defined and most economically relevant labor market competitors for testing spillover effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quit Elasticity / Labor Supply Elasticity (Firm-Level):&lt;/strong&gt; The quit elasticity is the percent change in the separation rate divided by the percent change in wages induced by the VMW. Under standard dynamic monopsony models (Manning, 2003), in steady state the recruit elasticity equals the quit elasticity, and the firm-level labor supply elasticity equals twice the quit elasticity. The authors estimate quit elasticities of 2.20–2.38, implying labor supply elasticities of 4.40–4.76 to the firm — consistent with meaningful but not extreme monopsony power.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cross-Employer Wage Elasticity:&lt;/strong&gt; The percent change in wages at a non-policy employer&amp;rsquo;s establishment associated with a 1% change in wages at the large retailer in the same commuting zone, instrumented using the large retailer&amp;rsquo;s gap interacted with the post-event indicator. Estimated to be a precise zero across all market definitions and event groupings in this paper.&lt;/p&gt;</description></item><item><title>What Do Policies Value?</title><link>https://macropaperwarehouse.com/papers/what-do-policies-value/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/what-do-policies-value/</guid><description>&lt;p&gt;This paper asks a fundamental question about policy design: when a program prioritizes one group over another, is that because the group benefits more from the intervention, or because the policy assigns them higher intrinsic welfare weight? Björkegren, Blumenstock, and Knight develop a two-stage method to decompose observed allocation decisions into their underlying components: (i) welfare weights assigned to different types of people, (ii) heterogeneous treatment effects of the intervention, and (iii) relative weights on different outcomes. The key insight is that the same allocation rule can be consistent with very different value systems depending on how much each group actually benefits.&lt;/p&gt;
&lt;p&gt;The method works as follows. In a first stage, the analyst estimates heterogeneous treatment effects — how much each individual benefits on each outcome dimension — using OLS or machine learning methods (e.g., causal forests). In a second stage, the analyst reconciles the observed ranking of beneficiaries with an implicit welfare function using an exploded logit likelihood, recovering welfare weights (who is valued), impact weights (how different outcomes are valued), and a base value for treatment independent of measured outcomes. Identification requires an exclusion restriction: the covariates used to estimate treatment effect heterogeneity must include variables excluded from the welfare weight specification, allowing the analyst to compare households with similar welfare weights but differential treatment effects. Variants of the method that impose known welfare weights or known impact weights can be used without the exclusion restriction.&lt;/p&gt;
&lt;p&gt;The paper demonstrates the method using PROGRESA, Mexico&amp;rsquo;s large conditional cash transfer program launched in 1997. PROGRESA ranked households by a proxy means test poverty score and transferred approximately 197 pesos per month (roughly $20 USD) to eligible poor households, conditional on school attendance and doctor visits. The analysis uses endline survey data on 7,767 households and focuses on three outcomes emphasized in program documents: log per-capita consumption, child sick days (ages 0-5), and school days missed (ages 6-16).&lt;/p&gt;
&lt;p&gt;The program&amp;rsquo;s average treatment effects were: a 0.149 log point increase in monthly consumption (SE=0.015), a 0.165 reduction in sick days per child (SE=0.051), and a near-zero effect on school days missed (-0.0053, SE=0.028). These effects were heterogeneous: indigenous households, for instance, benefited substantially more from the program.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central empirical finding inverts the naive interpretation of PROGRESA&amp;rsquo;s targeting. Indigenous households were ranked 60.6 log points higher in the program&amp;rsquo;s priority order. A simple regression suggests the program favored them. But after accounting for the fact that indigenous households benefit substantially more from treatment, the method finds that the program&amp;rsquo;s implied welfare weight on indigenous households is, if anything, lower by 17.4% relative to non-indigenous households — not higher. The program&amp;rsquo;s prioritization of indigenous households is thus explained by efficiency, not by preferential welfare weighting.&lt;/p&gt;
&lt;p&gt;Because PROGRESA cash transfers relax household budget constraints and outcomes like consumption reflect household choices, the impact weights capture the difference between how the policy values outcomes and how households value them. The estimates strongly reject non-paternalism: the policy implicitly values consumption and potentially health differently from household decision-makers. Of the total welfare impact, approximately 55% is attributed to the base value of the transfer itself (independent of measured outcomes), approximately 45% to consumption impacts, and less than 1% to health and schooling impacts combined. The implied value of providing the transfer independent of outcomes corresponds to 0.16 log points of consumption, or about 23.1 pesos per person per month — slightly below the average transfer of 33.9 pesos per person per month.&lt;/p&gt;
&lt;p&gt;The paper also runs counterfactual exercises showing how alternative preference structures would have changed the allocation. A policy maximizing only educational impacts would have prioritized richer, smaller households; one maximizing only consumption impacts would have further prioritized indigenous households. These counterfactuals are mapped onto a Pareto frontier across the three outcomes. The estimated welfare weights from the implemented policy align closely with preferences elicited in a 2023 survey of 429 Mexican residents, though residents placed higher value on child health relative to what the policy implied.&lt;/p&gt;
&lt;p&gt;Q: What is the core identification challenge the paper addresses?
A: When a policy prioritizes a group, it could be because the group benefits more (efficiency) or because the policy assigns them intrinsically higher value (preference). These two explanations are observationally equivalent from the allocation alone. The paper separates them by first estimating heterogeneous treatment effects and then inverting the allocation to recover residual welfare weights.&lt;/p&gt;
&lt;p&gt;Q: What is the exclusion restriction required for full identification?
A: The covariates used to estimate treatment effect heterogeneity (x-tilde) must include at least some variables excluded from the welfare weight specification (x). This allows the analyst to compare households with similar welfare weights but different predicted treatment effects, pinning down how much of the ranking reflects efficiency versus preference. Without this restriction, one can still recover conditional preferences by imposing known values for either welfare weights or impact weights.&lt;/p&gt;
&lt;p&gt;Q: How does the exploded logit likelihood work in this setting?
A: The analyst observes a single full ranking of all households, rather than partial orderings from multiple decision-makers. The welfare impact of treating household i is modeled as a linear function of predicted treatment effects scaled by welfare and impact weights, plus an extreme-value-distributed shock. The likelihood of observing household i ranked above household i-prime is the ratio of their exponentiated welfare scores, summed over all households ranked below i. Maximum likelihood recovers the welfare weights, impact weights, and base value simultaneously.&lt;/p&gt;
&lt;p&gt;Q: What were PROGRESA&amp;rsquo;s average treatment effects on the three focal outcomes?
A: Average treatment increased log monthly consumption by 0.149 (SE=0.015), reduced child sick days by 0.165 (SE=0.051), and had a near-zero effect on school days missed (-0.0053, SE=0.028). The consumption and health effects are statistically significant; the schooling effect is not distinguishable from zero.&lt;/p&gt;
&lt;p&gt;Q: What does the analysis find about the welfare weight assigned to indigenous households?
A: In the raw ranking regression, indigenous households are ranked 60.6 log points higher, suggesting the program favored them. After accounting for the fact that indigenous households benefit substantially more from treatment, the method finds the implied welfare weight on indigenous households is lower, not higher — specifically, about 17.4% lower than non-indigenous households. The program&amp;rsquo;s higher ranking of indigenous households is explained entirely by their larger treatment effects, not by preferential weighting.&lt;/p&gt;
&lt;p&gt;Q: How are the impact weights on consumption, health, and schooling interpreted given that outcomes reflect household choices?
A: Because PROGRESA relaxes household budget constraints and outcomes like consumption result from household optimization, the estimated impact weights capture the difference between how the policy values outcomes relative to how households value them (internalities), rather than the absolute policy valuation. A nonzero weight implies the policy disagrees with household preferences — paternalism. The positive coefficient on log consumption implies the policy values this outcome more than households do.&lt;/p&gt;
&lt;p&gt;Q: How much of PROGRESA&amp;rsquo;s welfare impact comes from the base transfer value versus measured outcomes?
A: The base value of the transfer (independent of measured impacts on consumption, health, and schooling) accounts for approximately 55% of total implied welfare impact. The impact on consumption accounts for approximately 45%. Impacts on health and schooling together account for less than 1%. The implied value of the base transfer corresponds to 0.16 log points of consumption per capita, or about 23.1 pesos per person per month — somewhat below the average transfer amount of 33.9 pesos per person per month.&lt;/p&gt;
&lt;p&gt;Q: Does the analysis reject egalitarian welfare weights and non-paternalism?
A: Yes, using Wald tests with bootstrapped covariance matrices. The hypothesis of egalitarian weights (all gamma equal to one) is rejected. Non-paternalism (all beta equal to zero) is strongly rejected. The joint hypothesis of egalitarianism and non-paternalism is also rejected across all specifications tested.&lt;/p&gt;
&lt;p&gt;Q: How do the estimated welfare weights compare to stated preferences of Mexican residents?
A: The 2023 survey of 429 Mexican residents elicited preferences using multiple price lists over how to prioritize different household types. The welfare weights implied by the implemented policy are broadly similar to resident preferences, but the policy places relatively higher welfare weight on indigenous households than the median survey respondent does. Survey respondents value child health impacts more than household decision-makers and more than the implemented policy does, consistent with support for paternalism.&lt;/p&gt;
&lt;p&gt;Q: What do counterfactual allocations reveal about the relationship between policy goals and targeting priorities?
A: A policy maximizing only consumption impacts would further prioritize indigenous households with lower income. A policy maximizing only educational impacts would instead prioritize richer, smaller households. A policy maximizing only health impacts would largely preserve indigenous household prioritization while placing less emphasis on lower-education households. These three extreme policies map to the corners of a Pareto frontier, and the implemented PROGRESA policy lies close to the allocation consistent with surveyed resident preferences.&lt;/p&gt;
&lt;p&gt;Q: What changed when Mexico reformed PROGRESA&amp;rsquo;s poverty score in 2003?
A: The 2003 reform increased the priority of older and smaller households. Applying the method to the new poverty score reveals that it implicitly switched to assigning a positive welfare weight to indigenous households (compared to the negative implied weight under the original score), and placed less welfare weight on lower-income and younger households relative to the original design.&lt;/p&gt;
&lt;p&gt;Q: What are the main limitations and scope conditions of the method?
A: Full identification requires an exclusion restriction (some treatment effect heterogeneity predictors excluded from welfare weights) and sufficient variation in treatment effects across household types. If treatment effects are homogeneous, welfare weights and impact weights cannot be separately identified. If correlated unobservables drive the ranking but are not modeled, the method recovers preferences consistent with included variables only, analogous to omitted variable bias in OLS. The method also requires a way to estimate treatment effect heterogeneity, which is most credible with a randomized pilot, though non-experimental methods are in principle applicable.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to the inverse optimum public finance literature?
A: The inverse optimum literature (Bourguignon and Spadaro 2012; Saez and Stantcheva 2016; Hendren 2020) recovers the redistribution preferences consistent with income tax schedules, conditioning on a single covariate (pre-tax income) affecting a single outcome (net-of-tax consumption). This paper generalizes that framework to arbitrary allocation policies conditioning on a vector of covariates and affecting a vector of outcomes, and extends it to settings beyond income taxation where heterogeneous treatment effects can be estimated.&lt;/p&gt;
&lt;p&gt;Q: Can the method be applied when only a binary allocation is observed rather than a full ranking?
A: Yes. A binary allocation corresponds to a ranking with only two levels, and the same exploded logit procedure applies, though with reduced statistical power. The paper provides an empirical illustration of this setting in Section 5.2.1.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Welfare weights (w(x_i)):&lt;/strong&gt; The policy&amp;rsquo;s differential valuation of one household&amp;rsquo;s utility relative to another, expressed as a multiplicative function of household characteristics. Distinct from how much a household benefits — two households may be ranked identically despite different benefits if their welfare weights differ proportionally.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Impact weights (beta_j):&lt;/strong&gt; The policy&amp;rsquo;s relative valuation of different outcome components (consumption, health, schooling). For outcomes that are household choices, impact weights capture the difference between how the policy values the outcome and how the household values it — an internality or paternalistic preference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Base value (alpha):&lt;/strong&gt; The value a policy assigns to providing a treatment independent of its measured impact on any specific outcome. Captures either a direct utility benefit of treatment or the value of relaxing household budget constraints when outcomes are choices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exclusion restriction:&lt;/strong&gt; The requirement that the set of covariates used to estimate treatment effect heterogeneity includes at least some variables excluded from the welfare weight specification. Enables separate identification of efficiency-based and preference-based components of a ranking by comparing households similar in welfare weight but different in predicted treatment effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exploded logit likelihood:&lt;/strong&gt; The econometric procedure used in the second stage, adapted for a single complete ranking of all alternatives rather than partial orderings. Treats the observed ranking of household i as a choice from the set of all households ranked below it, with likelihood given by the softmax of welfare scores.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Value audit:&lt;/strong&gt; A retrospective application of the method that reads the implicit values encoded in an implemented policy&amp;rsquo;s allocation decisions, enabling comparison against stated policy objectives, constituent preferences, or normative benchmarks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Paternalism (in this paper&amp;rsquo;s sense):&lt;/strong&gt; A policy is paternalistic if it assigns nonzero impact weight (beta_j ≠ 0) to outcomes that are household choices — meaning the policy values those outcomes differently from the households making the choices. The envelope theorem implies a non-paternalistic policy would place zero weight on choice outcomes beyond the general constraint relaxation.&lt;/p&gt;</description></item><item><title>What Jobs Come to Mind? Stereotypes About Fields of Study</title><link>https://macropaperwarehouse.com/papers/what-jobs-come-to-mind-stereotypes-about-fields-of-study/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/what-jobs-come-to-mind-stereotypes-about-fields-of-study/</guid><description>&lt;p&gt;Conlon and Patel test whether students stereotype the link between college majors and occupations — that is, whether they exaggerate the likelihood that majors lead to their &amp;ldquo;representative&amp;rdquo; careers (those most overrepresented among a major&amp;rsquo;s graduates relative to other majors, as measured by a likelihood ratio in US census data). The representative career for each major is intuitive: doctors for biology/chemistry, lawyers for political science, counselors for psychology, journalists for communications, artists for art, and so forth.&lt;/p&gt;
&lt;p&gt;The authors draw on three bodies of evidence. First, surveys of first-year undecided undergraduates in Ohio State University&amp;rsquo;s Exploration program (primarily Fall 2020 and Fall 2021 cohorts, ~80% response rate), asking students their beliefs about the share of US graduates in various careers conditional on major, as well as their beliefs about their own likely career. Beliefs are benchmarked against true career shares computed from the 2017–2019 American Community Survey restricted to college graduates aged 30–50. Second, 40+ years (1975–2018) of the CIRP Freshman Survey from UCLA, covering more than nine million nationally representative US college freshmen, which records intended major and intended career. Third, a field experiment embedded in the 2021 OSU survey with an RD design, in which treated students were shown the true share of their top major&amp;rsquo;s representative career before reporting beliefs, intentions, and — via administrative records — actual course enrollments and major declarations up to three years later.&lt;/p&gt;
&lt;p&gt;The main finding is large, systematic overestimation of representative careers. In the OSU survey, students believe 53% of art majors work as artists (true: 17%), 47% of journalism majors work as journalists (true: 4%), 38% of political science majors work as lawyers (true: 16%), and 43% of psychology majors work as counselors (true: 21%). OLS regressions of beliefs on true career frequency and a representative-career indicator yield a stereotyping coefficient θ of 0.32 p.p. (p &amp;lt; 0.01) without career fixed effects and 0.28 p.p. (p &amp;lt; 0.01) with them, meaning students believe representative careers are roughly 28–32 percentage points more common than equally prevalent non-representative careers. These patterns are similar across gender, ethnicity, and first-generation status, replicate in an MTurk sample (θ = 0.30, p &amp;lt; 0.01) and a nationally representative US adult sample (θ = 0.33, p &amp;lt; 0.01).&lt;/p&gt;
&lt;p&gt;In the CIRP data, 63% of biology freshmen expect to become doctors (true: 23%), 62% of psychology freshmen expect to be counselors (true: 21%), 65% of art freshmen expect to be artists (true: 17%), and 42% of communications/journalism freshmen expect to be writers or journalists (true: 4%). The average gap between expected and actual representative-career attainment is 36 p.p., and this gap has been roughly stable since at least the 1970s.&lt;/p&gt;
&lt;p&gt;An implicit association test (IAT) administered to 434 OSU students shows that implicit associations between representative major–career pairs are 0.30–0.36 standard deviations stronger than for non-representative pairs (p &amp;lt; 0.01), and remain 0.24–0.28 SDs stronger (p &amp;lt; 0.01) after controlling for true career frequency. A one-SD increase in individual IAT scores predicts 2.8–4.1 p.p. greater stereotyped beliefs (p &amp;lt; 0.01). Knowing someone with a non-representative major–career combination predicts beliefs 16 p.p. lower for the representative career (p &amp;lt; 0.01) — more than half the stereotyping effect — and also predicts lower IAT scores, suggesting associations arise from personal experience.&lt;/p&gt;
&lt;p&gt;An equilibrium model shows that stereotyping causes students to infer that representative careers have unusually favorable unobservable attributes, and that this inflates enrollment in the representative major among marginal students who are poorly suited to it. Correlational evidence from the NSCG, SIPP, and SHED confirms that majors subject to greater stereotyping are associated with more job dissatisfaction (+6.0% per SD, p &amp;lt; 0.01), greater job-skill mismatch (+3.1%, p &amp;lt; 0.05), more major-career mismatch (+5.4%, p &amp;lt; 0.05), and more regret about field of study (+4.8%, p &amp;lt; 0.05).&lt;/p&gt;
&lt;p&gt;The field experiment shows that correcting beliefs reduces stereotyping and shifts major choices. A 10 p.p. reduction in beliefs about the top major&amp;rsquo;s representative career lowers intentions toward that major by 3.5 p.p. (p &amp;lt; 0.01), reduces enrollment in that major&amp;rsquo;s courses by 0.22 credits in the next semester (p &amp;lt; 0.05), and reduces the probability of declaring that major within one year by 6.1 p.p. (p = 0.23). The same information boosts intentions toward students&amp;rsquo; second-ranked major by 2.1 p.p. (p = 0.17), increases second-major course enrollment by 0.20 credits (p &amp;lt; 0.10), and raises the probability of declaring the second major within a year by 9.9 p.p. (p &amp;lt; 0.01). Treated students also spend on average 0.21 more semesters undecided before declaring a major (p &amp;lt; 0.05). Effects are concentrated in the first year and partially fade over the two-to-three-year follow-up window.&lt;/p&gt;
&lt;p&gt;Q: How do the authors define a major&amp;rsquo;s &amp;ldquo;representative career&amp;rdquo;?
A: The representative career of major M is the career c that maximizes the likelihood ratio R(c, M) = p_{c|M} / p_{c|not-M}, where p_{c|M} is the true share of major-M graduates working in career c and p_{c|not-M} is the share of graduates from all other majors working in c. This ratio captures how much more common a career is among one major&amp;rsquo;s graduates relative to all other graduates. For example, the representative career of communications/journalism is &amp;ldquo;writers and journalists,&amp;rdquo; whose graduates are between 155% and 1,751% more likely to hold their major&amp;rsquo;s representative career than graduates of other majors, even though the absolute frequency of such careers is often modest (ranging from 2% to 60% across fields).&lt;/p&gt;
&lt;p&gt;Q: What is the core model of stereotyped belief formation?
A: The model draws from Bordalo et al. (2016). Let p_{c|M} be the true career share and π_{c|M} the student&amp;rsquo;s belief. The model specifies π_{c|M} = (1 − θ) p_{c|M} + θ · 1[c = c*(M)], where c*(M) is the representative career and θ ∈ [0,1] measures the extent of stereotyping. When θ = 0 the student holds rational beliefs; when θ = 1 beliefs assign all probability mass to the representative career. This formulation implies that students overweight representative careers because those careers come to mind more easily, grounded in a representativeness heuristic based on likelihood ratios.&lt;/p&gt;
&lt;p&gt;Q: What does the regression test for stereotyping find in the OSU survey?
A: The authors regress individual beliefs π_{c|M} on the true frequency p_{c|M} and an indicator for c being the representative career of M, clustering standard errors at the individual and career-by-major level. The estimated θ is 0.32 (p &amp;lt; 0.01) without career fixed effects (Column 1 of Table 1) and 0.28 (p &amp;lt; 0.01) with career fixed effects (Column 2). For self-beliefs about students&amp;rsquo; top-ranked major, the estimates are 0.36–0.43 p.p. (p &amp;lt; 0.01 both with and without career fixed effects). These estimates imply that students regard a major&amp;rsquo;s representative career as 28–43 percentage points more common than an equally prevalent non-representative career for the same major.&lt;/p&gt;
&lt;p&gt;Q: Do the OSU results replicate in other samples?
A: Yes. An MTurk convenience sample of 430 current college students yields a stereotyping coefficient of 0.30 (p &amp;lt; 0.01). A nationally representative sample of US adults yields a coefficient of 0.33 (p &amp;lt; 0.01); this pattern holds separately for college-educated and non-college-educated respondents and for both younger respondents (aged 18–29) and older respondents (aged 30+). The authors also ran a pre-registered 2021 replication survey in a new OSU Exploration cohort and found similar results.&lt;/p&gt;
&lt;p&gt;Q: What does the CIRP Freshman Survey data show about the persistence and scale of stereotyping?
A: Pooling more than nine million US college freshmen surveyed from 1975 to 2018, the CIRP data show that students systematically intend to enter their major&amp;rsquo;s representative career far more often than graduates actually do. Among students who have decided on a major, 63% intend to have their major&amp;rsquo;s representative career while only 27% of college graduates actually attain it — a gap of 36 p.p. (p &amp;lt; 0.01). The specific examples include: 63% of biology freshmen intend to become doctors (true: 23%), 62% of psychology freshmen expect to be counselors (true: 21%), 65% of art freshmen expect to be artists (true: 17%), and 42% of communications/journalism freshmen expect to be writers or journalists (true: 4%). The gap has been stable over the full 40+ year window, with no sign of convergence, and amounts to 40,000–200,000 students per year expecting careers in representative fields that they will not attain.&lt;/p&gt;
&lt;p&gt;Q: Can alternative mechanisms such as overconfidence or motivated reasoning explain the results?
A: The authors argue no, for two reasons. First, students overestimate the prevalence of representative careers not only for majors they plan to pursue (where overconfidence or motivated reasoning might apply) but also for majors they do not plan to pursue — the pattern holds for the gray (population belief) bars across all ten majors in Figure 1. Second, a Shapley-Sharrocks decomposition reported in Table A.V shows that the stereotyping mechanism accounts for a larger share of variance in beliefs than any other mechanism tested. A pre-registered survey also rules out unawareness of non-representative occupations as a driver: students are aware of the overwhelming majority of the 100 most common non-representative occupations, and such unawareness as exists is uncorrelated with stereotyped beliefs.&lt;/p&gt;
&lt;p&gt;Q: What does the IAT reveal about the mechanism behind stereotyping?
A: The IAT was run on 434 OSU Exploration students in Fall 2021, measuring implicit associations between five major–career pairs (Humanities-Writers and Journalists, Sciences-Healthcare, STEM-Business, Social Science-Law, Social Science-Counseling/Education). Participants sorted stimuli faster in &amp;ldquo;matched&amp;rdquo; blocks (where the representative career shares a response key with its major) than in &amp;ldquo;unmatched&amp;rdquo; blocks, yielding DID-IAT effects of 0.30–0.36 SDs (p &amp;lt; 0.01) for all five pairs. After controlling for true career frequency with career and major fixed effects, the effect shrinks only slightly to 0.24–0.28 SDs (p &amp;lt; 0.01), confirming that associations are driven by representativeness beyond base rates. At the individual level, a one-SD increase in DID-IAT scores predicts 4.1 p.p. greater stereotyped beliefs (p &amp;lt; 0.01) without career-by-major fixed effects and 2.8 p.p. (p &amp;lt; 0.01) with them.&lt;/p&gt;
&lt;p&gt;Q: What does the role-model heterogeneity analysis show?
A: Students were asked which major–career combinations they knew personally. Controlling for career-by-major fixed effects, knowing someone with a non-representative major–career combination (i.e., a non-default path) predicts beliefs about the representative career that are 16 p.p. lower (p &amp;lt; 0.01). This is more than half the size of the baseline stereotyping effect (28–32 p.p.). Knowing such a person also predicts lower IAT scores (p &amp;lt; 0.01), implying that personal exposure can reduce both implicit associations and explicit stereotyped beliefs.&lt;/p&gt;
&lt;p&gt;Q: What does the equilibrium model predict about misallocation?
A: The model embeds stereotyped beliefs in a two-stage choice framework: students choose a major first, then choose a career after graduation. It shows two main results (Propositions 1 and 2 in Online Appendix A.1). First, students who perceive the representative career as more common than it is will infer — through a rational expectations mechanism — that the unobservable amenities of that career are particularly favorable, so they will be surprised upon graduation. Second, stereotyping raises misallocation because it draws in marginal students whose career preferences make them poorly matched to the major&amp;rsquo;s representative career, while the inframarginal students who would have chosen the major anyway are better matched. The misallocation effect increases in the extent of stereotyping.&lt;/p&gt;
&lt;p&gt;Q: What correlational evidence links stereotyping to post-graduation mismatch outcomes?
A: Using major-level stereotyping estimates from the OSU data merged with three nationally representative surveys (NSCG, SIPP, SHED), the authors find: a one-SD increase in major-level stereotyping is associated with 6.0% more job dissatisfaction (p &amp;lt; 0.01, NSCG), 3.1% more reports that the job does not fit the worker&amp;rsquo;s skills and experience (p &amp;lt; 0.05, NSCG), 5.4% more reports that the job is unrelated to the field of study (p &amp;lt; 0.05, SIPP), and 4.8% more regret about field of study choice (p &amp;lt; 0.05, SHED). The authors note these are correlational and cannot rule out confounders such as underlying complexity of the career mapping.&lt;/p&gt;
&lt;p&gt;Q: How does the field experiment work and what is its identifying strategy?
A: The experiment was embedded in the second 2021 OSU survey, with students in the treatment group shown the true share of their top major&amp;rsquo;s representative career before reporting beliefs and intentions; control students answered the same questions without receiving this information. The main regression relates outcomes to (True Share − Prior Belief), set to zero for controls. Because students with less accurate prior beliefs may be more likely to choose the relevant major, OLS is potentially inconsistent; the authors use an RD design where the running variable is the information shock (True Share − Prior Belief), with the threshold at zero. Students just above (who overestimated) receive negative news; students just below (who underestimated) receive positive news. The RD estimates are combined with a first-stage estimate of belief updating to produce IV estimates of the effect of a 10 p.p. change in beliefs. Balance tests on predetermined demographics confirm no discontinuities at the threshold.&lt;/p&gt;
&lt;p&gt;Q: What are the first-stage belief-updating results?
A: Students update their posterior beliefs in response to the treatment: in response to information that the representative career is 1 p.p. less likely, students update their posterior beliefs down by 0.37 p.p. (p &amp;lt; 0.01). This under-reaction is consistent with Bayesian updating when priors are informative (Mobius et al. 2022). Students also update beliefs about non-representative careers: a 1 p.p. reduction in the representative career&amp;rsquo;s stated likelihood increases the expected probability of other careers by 0.27 p.p. (p &amp;lt; 0.01).&lt;/p&gt;
&lt;p&gt;Q: What are the effects of the information intervention on major intentions?
A: A 10 p.p. reduction in beliefs about the top major&amp;rsquo;s representative career reduces intentions (stated probability of graduating with that major) by 3.5 p.p. (p &amp;lt; 0.01). This effect is similar across subgroups (Columns 2–4 of Table 2). For students&amp;rsquo; second-ranked major, a 10 p.p. reduction in stereotyping boosts intentions by 2.1 p.p. (p = 0.17), which is imprecisely estimated but consistent in sign with all other outcomes.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on actual course enrollments?
A: In the semester immediately following the experiment, learning that the representative career of the first major is 10 p.p. less likely causes students to enroll in 0.22 fewer credits in that major&amp;rsquo;s field (95% CI: [−0.41, −0.02], p &amp;lt; 0.05), relative to a mean of 0.85 credits. Learning that the representative career of the second major is 10 p.p. less likely causes students to enroll in 0.20 more credits in the second major&amp;rsquo;s field (95% CI: [0.004, 0.40], p &amp;lt; 0.10), relative to a mean of 0.36 credits.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on official major declarations?
A: Within one year of the experiment, students who learned the representative career of their top major is 10 p.p. less likely are 6.1 p.p. less likely to have declared that major (95% CI: [−16.0, 3.8], p = 0.23) and 9.9 p.p. more likely to have declared their second major (95% CI: [2.5, 17.4], p &amp;lt; 0.01); the difference between these two effects is 16.0 p.p. (p &amp;lt; 0.01). By two years out, the effects are more attenuated. Treated students also spend on average 0.21 more semesters undecided before declaring a major (95% CI: [0.02, 0.40], p &amp;lt; 0.05). Effects do not appear to be driven by dropout: treated students are if anything slightly more likely to still be taking classes two to three years later.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Representativeness (likelihood ratio):&lt;/strong&gt; The representativeness R(c, M) of career c for major M is defined as the ratio p_{c|M} / p_{c|not-M} — how much more common career c is among major-M graduates than among graduates of all other majors. This is a relative, not absolute, frequency measure. The representative career (or exemplar) of a major is the career that maximizes this ratio.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stereotyping (as exaggeration of a kernel of truth):&lt;/strong&gt; In this paper&amp;rsquo;s framework, stereotyping means overweighting the representative career when forming beliefs about a major&amp;rsquo;s career distribution. The belief model is π_{c|M} = (1 − θ) p_{c|M} + θ · 1[c = c*(M)], where θ &amp;gt; 0 implies beliefs exaggerate how common the representative career is relative to equally prevalent non-representative careers. This is distinct from overconfidence, motivated reasoning, or simple noise.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DID-IAT score (difference-in-differences implicit association test):&lt;/strong&gt; The paper&amp;rsquo;s adaptation of the standard IAT to measure relative implicit associations between major and career groups. For a focal major–career pair, the DID-IAT score is the difference in the matched-vs-unmatched IAT D-score for the focal major (relative to a comparison major). A positive score indicates the focal major is more strongly associated with the focal career than the comparison major is. This measures implicit memory-based associations rather than deliberate beliefs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Misallocation (as used in the model):&lt;/strong&gt; The welfare loss arising because stereotyped beliefs draw marginal students — those on the margin between choosing the representative major and not — who have career preferences close to the average rather than being the students best suited to that major. These marginal students end up choosing careers other than the representative career after graduation at higher rates, producing major-career mismatch. Misallocation is shown (Proposition 2) to increase in the extent of stereotyping θ.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Information shock:&lt;/strong&gt; In the field experiment, the information shock for a given student and major is the difference between the true share of the major&amp;rsquo;s representative career and the student&amp;rsquo;s prior belief about that share. Positive shocks correspond to students who overestimated (and thus receive bad news); negative shocks correspond to students who underestimated (and receive good news). The RD design uses the threshold at shock = 0 to generate quasi-experimental variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Source text origin (implicit in the paper&amp;rsquo;s design):&lt;/strong&gt; The paper measures beliefs about career distributions benchmarked against American Community Survey data on actual career outcomes of college graduates aged 30–50, restricting to respondents born 1958–1997. This defines the objective ground truth against which stereotyping is measured throughout the paper.&lt;/p&gt;</description></item><item><title>What Works and for Whom? Effectiveness and Efficiency of School Capital Investments Across the U.S.</title><link>https://macropaperwarehouse.com/papers/what-works-and-for-whom-effectiveness-and-efficiency-of-school-capital-investments-across-the-u.s./</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/what-works-and-for-whom-effectiveness-and-efficiency-of-school-capital-investments-across-the-u.s./</guid><description>&lt;h2 id="what-works-and-for-whom-effectiveness-and-efficiency-of-school-capital-investments-across-the-us"&gt;What Works and for Whom? Effectiveness and Efficiency of School Capital Investments Across the U.S.&lt;/h2&gt;
&lt;h3 id="research-question"&gt;Research Question&lt;/h3&gt;
&lt;p&gt;This paper investigates which types of school facility investments benefit students (as measured by test scores) and are valued by homeowners (as measured by house prices), and for which student populations these investments are most effective. Prior state-level studies had reached conflicting conclusions about the returns to school capital spending, and no nationwide evidence had distinguished impacts across spending categories or student backgrounds.&lt;/p&gt;
&lt;h3 id="data-and-methodology"&gt;Data and Methodology&lt;/h3&gt;
&lt;p&gt;The authors assemble a novel panel dataset covering approximately 14,000 school bond referenda in 29 U.S. states and 10,146 districts enrolling 71% of all U.S. students, for the period 1990–2017. The dataset combines: (1) ballot-level bond election records including vote shares, proposed amounts, and ballot text; (2) district-level test scores from the Stanford Education Data Archive (SEDA) extended backward to 2003 for all states and as early as 1995 for some, normalized to a national scale via NAEP; (3) a Census-tract-level house price index (Contat and Larson, 2022) aggregated to school districts; and (4) NCES district finance and demographic data.&lt;/p&gt;
&lt;p&gt;Bond ballot texts are classified into eight spending categories using text-analysis: classroom construction/renovation; HVAC; other infrastructure (plumbing, roofs, furnaces); safety and health (pollutant removal, building safety); STEM equipment and labs; athletic facilities; land purchases; and transportation vehicles.&lt;/p&gt;
&lt;p&gt;The identification strategy exploits quasi-random variation from close bond elections, building on the dynamic regression discontinuity (DRD) framework of Cellini et al. (2010). A key methodological contribution is a stacked DRD design that addresses heterogeneous treatment effects correlated with timing: each treatment cohort (districts that narrowly authorize a bond in year c) is matched against &amp;ldquo;clean controls&amp;rdquo; — districts that also proposed a bond in the same cohort but narrowly failed to authorize it and did not authorize any bond in the following ten years. Cohorts are stacked, and a dynamic RD model is estimated controlling for cohort fixed effects and a district&amp;rsquo;s bond proposal history.&lt;/p&gt;
&lt;h3 id="main-findings-with-quantitative-magnitudes"&gt;Main Findings with Quantitative Magnitudes&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Average effects.&lt;/strong&gt; Bond authorization raises capital spending by approximately $1,650 per pupil cumulatively over five years. Test scores increase gradually, reaching 0.079 standard deviations (sd) higher five to eight years after authorization, and 0.073 sd higher nine to twelve years after. 2SLS estimates, amortizing spending over a 30-year project life at a 9% depreciation rate, imply that a $1,000 increase in the flow value of capital spending raises test scores by 0.048 sd. House prices rise by approximately 9% eight to nine years after authorization. When house price effects are estimated against only locally-financed capital spending (not state aid), the 2SLS estimate is 0.8% per $1,000 — roughly consistent with efficiency — suggesting that the larger reduced-form house price response is driven primarily by state aid that supplements local funds rather than by an inefficiently low ex ante spending level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity by spending category.&lt;/strong&gt; Category-specific estimates reveal that only certain project types raise test scores: HVAC (+0.20 sd, largest effect), safety and health (+0.15 sd), other infrastructure/plumbing/roofs (+0.15 sd), STEM equipment (+0.15 sd implied), and classroom space (+0.10 sd), all measured three to six years post-election. By contrast, bonds for athletic facilities, land purchases, and transportation produce no detectable effects on test scores. The pattern for house prices is the inverse: athletic facilities generate a 17% house price increase; classroom space generates 14%; STEM generates 11% — while HVAC and safety/health bonds produce no significant effect on house prices. The correlation between category-level test score and house price estimates is −0.07, indicating these are largely orthogonal outcomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity by student socioeconomic status.&lt;/strong&gt; Effects are concentrated in districts serving socioeconomically disadvantaged students (top tercile of the share of students eligible for free or reduced-price meals, denoted low-SES). In low-SES districts, bond authorization raises test scores by 0.13 sd after seven years and house prices by 15%; in high-SES districts, neither outcome shows a significant effect. 2SLS estimates confirm that a $1,000 increase in cumulative spending raises test scores by 0.08 sd in low-SES districts but produces no detectable change in high-SES districts. The SES gradient persists after conditioning on spending amounts, spending categories, and baseline capital stock, indicating that students in disadvantaged districts have higher marginal returns to capital improvements independent of these channels. High-minority districts (top tercile of Black and Hispanic share) similarly see a 0.12 sd test score gain and 15% house price gain after seven years, versus 0.04 sd and 3% in low-minority districts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Role of baseline capital stock.&lt;/strong&gt; Among districts with below-median capital stock, test score effects are 0.20 sd in low-SES districts seven years post-election. Even among above-median-stock districts, low-SES districts see house price effects exceeding 10% while high-SES districts see no effect. Differences by SES persist after conditioning on capital stock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy simulation.&lt;/strong&gt; Closing the spending gap between high- and low-SES districts (approximately $1,000 over 10 years) without changing the composition of spending would raise low-SES test scores by roughly 0.08 sd, closing about 8% of the roughly 1 sd achievement gap. Targeting that same additional spending toward HVAC and safety/health (the highest-impact categories) would generate test score increases approximately three times as large, potentially closing up to 25% of the observed achievement gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reconciling prior literature.&lt;/strong&gt; Replicating state-level estimates, the authors show that Ohio&amp;rsquo;s positive effects are explained by a high share of bonds in low-SES districts funding infrastructure, while Texas&amp;rsquo;s near-zero effects reflect a high share of bonds in higher-SES districts funding classrooms and athletic facilities.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-first-stage-effect-of-bond-authorization-on-capital-spending-and-does-it-contaminate-other-spending-categories"&gt;Q1. What is the first-stage effect of bond authorization on capital spending, and does it contaminate other spending categories?&lt;/h3&gt;
&lt;p&gt;A1: Bond authorization raises per-pupil capital spending by approximately $700 per year at two years post-election and $590 at three years, with cumulative spending $1,650 higher over five years in treated districts relative to districts that narrowly failed to authorize a bond. Bond revenues are legally restricted to capital uses, and the paper confirms that non-capital (current) spending and instructional spending are not affected following authorization. This establishes a clean first stage: bond authorization raises only capital outlays.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-standard-drd-estimator-of-cellini-et-al-2010-require-refinement-and-what-problem-does-the-stacked-drd-design-solve"&gt;Q2. Why does the standard DRD estimator of Cellini et al. (2010) require refinement, and what problem does the stacked DRD design solve?&lt;/h3&gt;
&lt;p&gt;A2: The original CFR estimator assumes treatment effects are uncorrelated with the timing of treatment — an assumption potentially violated when, for example, bonds financing HVAC (high-impact) versus athletic facilities (amenity-focused) have different propensities to be proposed at different points in time. The stacked DRD design avoids &amp;ldquo;forbidden comparisons&amp;rdquo; by comparing each treatment cohort only against clean controls that propose but fail to authorize a bond in the same year and do not authorize any bond in the subsequent ten years. This ensures consistency even when treatment effects are heterogeneous across cohorts and correlated with timing.&lt;/p&gt;
&lt;h3 id="q3-how-do-the-authors-validate-the-quasi-random-assignment-assumption-of-the-regression-discontinuity-design"&gt;Q3. How do the authors validate the quasi-random assignment assumption of the regression discontinuity design?&lt;/h3&gt;
&lt;p&gt;A3: Three tests are performed. First, a McCrary (2008) density test on the vote margin distribution shows no discontinuity at the cutoff in the pooled or stacked data (p-values of 0.59 and 0.24, respectively), though discontinuities are found in Arkansas, Missouri, and Oklahoma — those three states are excluded. Second, pre-election district covariates (income, education, SES shares, enrollment, revenues, expenditures) are smooth around the cutoff in both datasets. Third, pre-election trends in test scores and house prices are flat and parallel between marginally approved and marginally rejected districts.&lt;/p&gt;
&lt;h3 id="q4-how-are-the-eight-spending-categories-constructed-and-how-many-bonds-are-successfully-classified"&gt;Q4. How are the eight spending categories constructed, and how many bonds are successfully classified?&lt;/h3&gt;
&lt;p&gt;A4: Categories are drawn from the SchoolBondFinder.com classification produced by The Amos Group, then refined by splitting capital improvements into HVAC versus other infrastructure, splitting construction/renovation into classroom versus athletic facility projects, and adding land purchases as a separate category. Keyword-based text analysis of ballot language successfully assigns 75% of the approximately 14,000 bonds to at least one of the eight categories. More than two-thirds of classified bonds receive multiple category designations, with a mean of 2.9 categories per proposed bond and 3.2 per authorized bond.&lt;/p&gt;
&lt;h3 id="q5-why-do-hvac-bonds-raise-test-scores-but-not-house-prices-while-athletic-facility-bonds-raise-house-prices-but-not-test-scores"&gt;Q5. Why do HVAC bonds raise test scores but not house prices, while athletic facility bonds raise house prices but not test scores?&lt;/h3&gt;
&lt;p&gt;A5: The authors interpret this divergence as reflecting what different types of improvements offer to different stakeholders. HVAC improvements reduce excessive heat and air pollution exposure in classrooms, directly improving students&amp;rsquo; learning experiences — consistent with Park et al. (2020) on heat and Gilraine and Zheng (2022) on air pollution. These improvements are not visibly salient to homeowners without school-age children and carry no amenity value for the broader community. Athletic facilities, by contrast, are highly visible and provide a community amenity valued in the housing market regardless of their impact on academic instruction. The near-zero correlation (−0.07) between category-level test score and house price estimates confirms that the two outcomes respond to largely distinct features of capital investments.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-three-candidate-explanations-for-the-larger-effects-of-bond-authorization-in-low-ses-districts-and-which-explanations-survive-empirical-scrutiny"&gt;Q6. What are the three candidate explanations for the larger effects of bond authorization in low-SES districts, and which explanations survive empirical scrutiny?&lt;/h3&gt;
&lt;p&gt;A6: The three candidates are: (1) larger spending increases after authorization in low-SES districts; (2) a different composition of spending categories (more toward high-impact HVAC and safety); and (3) higher marginal returns per dollar for disadvantaged students, holding spending size and composition fixed. The data confirm all three operate, but the third is the residual: 2SLS estimates show a $1,000 increase raises test scores by 0.08 sd in low-SES districts versus a statistically zero effect in high-SES districts, and within-category estimates show HVAC bonds raise scores by 0.27 sd in low-SES districts but have no detectable effect in high-SES districts. Differences by SES also persist after conditioning on the estimated baseline capital stock, though low capital stock accounts for part of the gap.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-role-of-state-aid-alter-the-interpretation-of-the-house-price-effect-for-spending-efficiency"&gt;Q7. How does the role of state aid alter the interpretation of the house price effect for spending efficiency?&lt;/h3&gt;
&lt;p&gt;A7: A 9% house price increase after bond authorization, if taken at face value under Brueckner&amp;rsquo;s (1979) efficiency test, would suggest the ex ante level of school capital spending was inefficiently low. However, state grants that partly match local bond revenues raise actual spending without raising local property taxes proportionally. When the 2SLS house price effect is estimated against only locally financed capital spending (using proposed bond size as the relevant measure), the implied house price increase is just 0.8% per $1,000 — consistent with rough efficiency on average across the full sample. The authors conclude that the large reduced-form house price response is driven primarily by the capitalization of state aid, not by an undersupply of capital investments at the aggregate level.&lt;/p&gt;
&lt;h3 id="q8-does-household-sorting-account-for-the-observed-test-score-and-house-price-gains-following-bond-authorization"&gt;Q8. Does household sorting account for the observed test score and house price gains following bond authorization?&lt;/h3&gt;
&lt;p&gt;A8: Bond authorization produces small but detectable compositional changes: the share of high-SES students is approximately 3 percentage points higher seven years after an election (a roughly 4% increase relative to an average share of 0.73), while enrollment and the share of white students are largely unaffected. However, controlling for district-by-year shares of each sociodemographic group only slightly attenuates the test score and house price estimates, indicating that sorting accounts for a small share of the observed gains.&lt;/p&gt;
&lt;h3 id="q9-are-the-findings-robust-to-alternative-research-designs"&gt;Q9. Are the findings robust to alternative research designs?&lt;/h3&gt;
&lt;p&gt;A9: The results are robust to five alternative estimation approaches: (1) the original one-step TOT estimator of Cellini et al. (2010); (2) a version of the stacked DRD where clean controls are districts that do not approve any bonds in the full [c−5, c+10] window; (3) a version that matches treated and control districts in each cohort based on bond history; (4) a version not controlling for future bond history; and (5) the extended two-way fixed effects (ETWFE) estimator of Wooldridge (2021). Results are also robust to linear polynomials with different slopes and quadratic polynomials of the vote margin.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-capital-stock-measure-illuminate-mechanism-and-what-are-its-limitations"&gt;Q10. How does the capital stock measure illuminate mechanism, and what are its limitations?&lt;/h3&gt;
&lt;p&gt;A10: The authors construct a district-level capital stock as the 30-year depreciated sum of capital spending from Census of Governments data (1967–2017) at a 5% depreciation rate. This stock is negatively correlated with the share of low-SES students, confirming that more disadvantaged students attend schools in worse structural condition. Conditioning on this proxy, the SES gradient in bond impacts is reduced but remains. Among districts with below-median capital stock, low-SES districts see test score gains of 0.20 sd after seven years, while among above-median-stock districts the gap narrows to approximately 0.10 vs. 0.05 sd. A key limitation is that detailed school-condition data are unavailable nationally, so the capital stock is a proxy only.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-quantitative-policy-implication-of-the-targeting-exercise"&gt;Q11. What is the quantitative policy implication of the targeting exercise?&lt;/h3&gt;
&lt;p&gt;A11: On average, low-SES districts receive about $97 per pupil per year less in capital spending than high-SES districts, so closing this gap over ten years implies approximately $970 in additional cumulative spending. Without changing spending composition, this would raise test scores by roughly 0.08 sd in low-SES districts, closing about 8% of the approximately 1 sd achievement gap between high- and low-SES districts. Redirecting that same additional spending toward the highest-impact categories (HVAC and safety/health) would generate test score gains roughly three times larger, potentially closing up to 25% of the observed achievement gap.&lt;/p&gt;
&lt;h3 id="q12-how-do-the-cross-state-differences-documented-in-prior-literature-map-onto-the-papers-heterogeneity-findings"&gt;Q12. How do the cross-state differences documented in prior literature map onto the paper&amp;rsquo;s heterogeneity findings?&lt;/h3&gt;
&lt;p&gt;A12: The authors replicate earlier state-level estimates and show that Ohio&amp;rsquo;s relatively large positive effects — found by Conlin and Thompson (2017) — are explained by a high concentration of bonds in low-SES districts funding infrastructure, while Texas&amp;rsquo;s near-zero effects — found by Martorell et al. (2016) — reflect a high share of bonds in higher-SES districts funding classrooms and athletic facilities. Wisconsin and Michigan, which showed null effects in earlier studies, similarly have bond compositions and student demographics that predict small impacts under the paper&amp;rsquo;s heterogeneity framework.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Stacked Dynamic Regression Discontinuity (Stacked DRD).&lt;/strong&gt; The paper&amp;rsquo;s primary estimation strategy, which combines the dynamic RD framework of Cellini et al. (2010) with a stacked-cohort design adapted from the staggered difference-in-differences literature. For each treatment cohort (year in which a bond barely passes), &amp;ldquo;clean controls&amp;rdquo; are defined as districts that also proposed a bond in the same year but narrowly failed to authorize it and did not authorize any subsequent bond within ten years. Cohort-specific datasets are stacked and estimated jointly with cohort fixed effects, ensuring that estimates are robust to treatment effect heterogeneity correlated with timing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Clean Controls.&lt;/strong&gt; Districts used as the counterfactual for treated districts in a given cohort: those that propose a bond in the same year as the treated cohort, barely fail to authorize it, and remain untreated for ten subsequent years. Their &amp;ldquo;clean&amp;rdquo; status is quasi-random because their future non-authorization results from narrow electoral loss rather than any endogenous district choice.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bond Spending Categories.&lt;/strong&gt; Eight mutually-non-exclusive classifications of bond spending derived from ballot text using keyword analysis: classroom space; HVAC; other infrastructure (plumbing, roofs, furnaces); safety and health (pollutant removal, compliance upgrades); STEM equipment and labs; athletic facilities; land purchases; and transportation. These categories are defined in the paper not by administrative accounting codes but by the stated intended use of funds in ballot language.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Treatment-on-the-Treated (TOT) Estimator.&lt;/strong&gt; The CFR estimator that captures the effect of bond authorization against the counterfactual of never authorizing a bond in the foreseeable future, achieved by including leads and lags of a district&amp;rsquo;s bond proposal history as controls. This addresses the problem that multiple elections over time make simple treated-vs-control comparisons confounded by past and future bond activity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital Stock (District-Level Proxy).&lt;/strong&gt; A measure of each district&amp;rsquo;s accumulated school facility capital at a given point in time, constructed as the depreciated 30-year running sum of capital expenditures from the Census of Governments, using a 5% annual depreciation rate. Used as a proxy for facility conditions in the absence of nationally available building-quality data, and confirmed to be negatively correlated with district share of low-SES students.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Brueckner Efficiency Test.&lt;/strong&gt; An application of the theoretical framework linking public good provision levels to house price responses. If a spending increase raises house prices, the initial spending level was below the efficient level; if it lowers house prices, spending was too high. In this paper, the test is refined to use only locally-financed capital spending as the explanatory variable, to strip out the capitalization of state aid and isolate the efficiency assessment for locally-determined spending.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Socio-Economic Status (SES) Terciles.&lt;/strong&gt; Districts are ranked by the share of students eligible for free or reduced-price school meals as of 1995. &amp;ldquo;Low-SES districts&amp;rdquo; refers to those in the top tercile of this share (most disadvantaged); &amp;ldquo;high-SES districts&amp;rdquo; refers to those in the bottom tercile (least disadvantaged). Effects are estimated separately for these subsamples throughout.&lt;/p&gt;</description></item><item><title>What's My Employee Worth? The Effects of Salary Benchmarking</title><link>https://macropaperwarehouse.com/papers/whats-my-employee-worth-the-effects-of-salary-benchmarking/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/whats-my-employee-worth-the-effects-of-salary-benchmarking/</guid><description>&lt;p&gt;This paper studies how salary benchmarking tools — products that reveal aggregate market pay statistics for specific job titles — affect employee compensation. The research question is whether firms&amp;rsquo; access to such tools causally changes how they set salaries, and what this implies about information frictions in labor markets and the policy debate over benchmarking regulation.&lt;/p&gt;
&lt;p&gt;The authors collaborated with the largest U.S. payroll processing company (serving 650,000 firms and 20 million workers), exploiting the staggered roll-out of a proprietary Compensation Benchmark Tool. The tool aggregates payroll data into salary benchmarks by standardized job title, with the median base salary as its most prominent statistic. The study draws on three linked administrative datasets: payroll records (January 2017 to July 2021), tool usage logs (September 2019 to August 2021), and historical benchmark snapshots. The main analytical sample covers new hires at 586 treatment firms that gained tool access and 1,419 matched control firms that did not, within a 10-quarter window around each firm&amp;rsquo;s onboarding date.&lt;/p&gt;
&lt;p&gt;The identification strategy is difference-in-differences, exploiting three sources of variation: which firms gain access; the staggered timing of access (driven by the arbitrary order in which sales representatives introduced the tool); and within treatment firms, whether a specific position was actually searched in the tool. New hires are classified into Searched positions (5,266 hires at treatment firms for positions eventually looked up), Non-Searched positions (39,686 hires at treatment firms for positions not looked up), and Non-Searchable positions (156,865 hires at control firms). Event-study analyses confirm flat pre-trends across all groups, supporting causal interpretation.&lt;/p&gt;
&lt;p&gt;The primary finding is that benchmark access reduces salary dispersion around the median market benchmark by 25%. Before onboarding, the average absolute deviation of offered salaries from the median benchmark in Searched positions was 19.8 percentage points (pp). After onboarding, this fell to 14.9 pp — a drop of 5.0 pp using Non-Searched positions as control (p-value &amp;lt; 0.001) and 6.2 pp using Non-Searchable positions as control (p-value &amp;lt; 0.001). Compression runs in both directions: firms previously paying above the benchmark reduce salaries toward the median, and firms previously paying below raise salaries toward the median. The probability of setting a salary within 2.5% of the median benchmark nearly doubled, from 11.6% to 22.1% after onboarding.&lt;/p&gt;
&lt;p&gt;Effects are heterogeneous by skill level. For low-skill positions (approximately 42% of the sample, e.g., bank teller, receptionist), dispersion falls from 14.5 pp to 8.7 pp — a 40% reduction. For high-skill positions (e.g., software developer), dispersion falls from 24.0 pp to 20.5 pp — a 14.6% reduction. For low-skill positions, compression from below dominates, producing a net average salary increase of +5.0% to +6.7% (p-values 0.014 and 0.001 depending on control group). For high-skill positions, the average salary effect is small and statistically insignificant overall. Twelve-month retention rates for low-skill workers increase by 6.6 to 6.8 pp after benchmarking, and the implied retention elasticity is consistent with prior literature estimates.&lt;/p&gt;
&lt;p&gt;The authors propose a theoretical model to rationalize these findings. Firms are assumed uncertain about the wage distribution (aggregate uncertainty), with private information about their own value of filling a position and affiliated valuations across firms. In equilibrium, firms with higher values make higher offers — generating wage dispersion among identical workers without monopsony power, efficiency wages, or amenity differences. When a firm gains benchmark access, it adjusts its offer toward the threshold wage needed to hire, compressing offers from both sides. In the full-information equilibrium where benchmarks are common knowledge, the mean salary is weakly higher than without benchmarks, because the marginal firm had previously underestimated labor market tightness and offered too little, capturing extraordinary profits. Benchmarking eliminates these informational rents, intensifying competition and raising average pay.&lt;/p&gt;
&lt;p&gt;The scope of the empirical findings is restricted to new hires at firms in the top quartile of U.S. firm size by employment, across all industries and U.S. states, over 2017–2020. The estimated effect is the incremental causal impact of one additional high-quality benchmarking source, since most firms already had access to some pay information through other channels.&lt;/p&gt;
&lt;p&gt;Q: What is the main causal finding of the paper?
A: Access to the salary benchmarking tool reduces the absolute deviation of new-hire salaries from the median market benchmark by approximately 25%. Specifically, average dispersion in Searched positions falls from 19.8 pp before onboarding to 14.9 pp after, a drop of 5.0 pp (using Non-Searched controls, p-value &amp;lt; 0.001) or 6.2 pp (using Non-Searchable controls, p-value &amp;lt; 0.001). The two estimates are statistically indistinguishable from each other, and both are robust to a wide range of specification checks.&lt;/p&gt;
&lt;p&gt;Q: How does compression operate — does it raise or lower salaries?
A: Compression operates in both directions. Firms that would otherwise have paid above the median benchmark reduce salaries toward the median (&amp;ldquo;compression from above&amp;rdquo;), and firms that would otherwise have paid below the median benchmark raise salaries toward the median (&amp;ldquo;compression from below&amp;rdquo;). The probability of offering a salary within 2.5% of the median benchmark nearly doubled, from 11.6% before onboarding to 22.1% after.&lt;/p&gt;
&lt;p&gt;Q: What is the identification strategy, and why is the treatment considered as good as random?
A: The authors use a difference-in-differences design with three sources of variation: which firms gain tool access, the staggered timing of access, and whether specific positions were actually searched within a treatment firm. The payroll company introduced the tool through sales representatives contacting clients in an arbitrary order, not in response to firm characteristics or outcomes. This is corroborated by empirical tests: event-study pre-trends for Searched versus Non-Searched (and Non-Searchable) positions are flat and statistically indistinguishable from zero (pre-treatment coefficients of -0.346 and -0.310, p-values 0.749 and 0.604, respectively).&lt;/p&gt;
&lt;p&gt;Q: How large are the effects for low-skill versus high-skill positions?
A: For low-skill positions (approximately 42% of the sample, e.g., bank teller, receptionist), dispersion drops from 14.5 pp to 8.7 pp — a 40% decline (p-value &amp;lt; 0.001). For high-skill positions (e.g., software developer), dispersion drops from 24.0 pp to 20.5 pp — a 14.6% decline (p-value = 0.021). The larger effect for low-skill positions is consistent with anecdotal accounts from compensation managers, who report treating low-skill candidates as interchangeable and therefore wanting to offer exactly the market rate.&lt;/p&gt;
&lt;p&gt;Q: Does benchmarking raise or lower average salaries?
A: On average across all skill levels, the effect on mean salary is small and statistically insignificant: -0.2% (p-value = 0.756) using Non-Searched controls and +1.7% (p-value = 0.308) using Non-Searchable controls. For low-skill positions specifically, average salaries increase by +5.0% (p-value = 0.014) using Non-Searched controls and +6.7% (p-value = 0.001) using Non-Searchable controls. This net increase for low-skill workers reflects compression from below dominating compression from above in that subset.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on employee retention?
A: For low-skill workers, benchmarking increases the probability of remaining employed at the hiring firm 12 months after the hire date by +6.6 pp (p-value = 0.101) using Non-Searched controls and +6.8 pp (p-value = 0.029) using Non-Searchable controls. The implied retention elasticity from the ratio of salary and retention effects is consistent with average estimates in the prior literature (Sokolova and Sorensen, 2021). No retention effects are reported for high-skill positions.&lt;/p&gt;
&lt;p&gt;Q: What is the theoretical mechanism through which aggregate uncertainty generates wage dispersion?
A: The model assumes a unit mass of firms simultaneously making wage offers to a mass Q &amp;lt; 1 of workers, with only the top Q offers accepted. Firms have private information about their value of filling the position, and values are affiliated (correlated in the sense of Milgrom and Weber, 1982). Because each firm is uncertain about what other firms will offer, higher-value firms rationally form higher beliefs about the prevailing wage distribution and make higher offers. This generates equilibrium wage dispersion among identical workers without monopsony power, efficiency wages, or amenity differences.&lt;/p&gt;
&lt;p&gt;Q: What does the model predict about the equilibrium effects of benchmarking when all firms have access?
A: When the benchmark is common knowledge, all firms make offers with full information about the wage distribution. The firms with the highest values win workers at a uniform wage that makes the marginal firm indifferent between hiring and not hiring. The model proves that the mean salary is higher in expectation under the benchmark equilibrium than in the no-benchmark equilibrium. The intuition is that without benchmarks, the marginal firm underestimates labor market tightness, offers less than the full-information competitive wage, and thereby captures extraordinary profits; benchmarking eliminates those rents and intensifies competition.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings regarding antitrust concerns?
A: In 2023, the DOJ and FTC rescinded a long-standing antitrust &amp;ldquo;safety zone&amp;rdquo; for salary benchmarks due to concerns that they could facilitate wage collusion. A 2021 executive order had mandated that agencies consider procompetitive effects as well. The authors&amp;rsquo; model addresses the collusion concern directly: in equilibrium, benchmarking raises (not lowers) average salaries. The empirical evidence is consistent with this — low-skill workers see average salary increases of 5-7% after benchmarking — suggesting a procompetitive justification for the tools.&lt;/p&gt;
&lt;p&gt;Q: How robust are the main results?
A: The main estimates are robust across a wide range of specification checks, including alternative winsorization levels, log-difference and binary (&amp;gt;10% deviation) dependent variables, heteroskedasticity-robust standard errors, exclusion of controls, inclusion of firm fixed effects, exclusion of tipping positions, restriction to Searched positions only, dropping SOC reweighting, and age restrictions. Two additional pieces of evidence corroborate the quasi-experimental findings: a survey experiment with SHRM HR managers shows that hypothetical benchmarks compress stated salary offers from both above and below; and quasi-random benchmark shocks (when large firms abruptly raise a position&amp;rsquo;s base salary by 10% or more) cause firms with tool access to converge to the new benchmark faster than firms without access.&lt;/p&gt;
&lt;p&gt;Q: What does the survey of HR managers reveal about how firms use benchmarks?
A: In a survey of 2,696 HR professionals conducted through SHRM&amp;rsquo;s research panel, 87.6% of those involved in salary-setting report using salary benchmarks. The vast majority (97.4%) use benchmarks to set pay for new hires. The most popular sources are industry surveys (68.0%) and free online data (58.1%), with payroll data services used by 23.2%. The median salary is ranked the most important benchmark statistic by 56.73% of respondents. Most respondents apply filters by state (84.15%) and industry (87.33%) when using the tool.&lt;/p&gt;
&lt;p&gt;Q: What are the main sources of potential attenuation or amplification bias in the estimated effects?
A: Attenuation bias may arise because (1) the benchmark tool studied is among the most advanced available, so firms already had some wage information from other sources, meaning the estimates capture only the incremental effect of one additional high-quality source; and (2) not all positions at treatment firms were searched, so the sample is restricted to positions where firms actually engaged with the benchmark. Potential upward bias could arise if firms adopting the tool were also undergoing broader HR system changes, but the flat event-study pre-trends argue against this explanation.&lt;/p&gt;
&lt;p&gt;Salary Benchmarking: The practice of using aggregated market pay data — provided by third parties such as payroll processors, consulting firms, or online platforms — to identify typical salaries for specific job titles and set internal pay accordingly. In the paper&amp;rsquo;s context, this refers specifically to an online tool that allows employers to look up the median and distributional statistics of base salaries for standardized position titles, filtered by industry and state.&lt;/p&gt;
&lt;p&gt;Aggregate Uncertainty: The paper&amp;rsquo;s label for a distinct source of information friction in which firms are uncertain about the distribution of wages offered by other firms in the market — as opposed to uncertainty about individual worker characteristics. This uncertainty is assumed to be the primitive that generates equilibrium wage dispersion in the model, and its resolution through benchmarking is the mechanism driving the empirical results.&lt;/p&gt;
&lt;p&gt;Salary Dispersion (around the benchmark): Measured empirically as the average absolute percentage difference between a new hire&amp;rsquo;s starting base salary and the median market benchmark for that position, expressed in percentage points. This is the paper&amp;rsquo;s primary outcome variable. Dispersion reflects firms&amp;rsquo; deviation from the market rate in either direction.&lt;/p&gt;
&lt;p&gt;Compression from Above / Compression from Below: Compression from above refers to the reduction in salaries at firms that would otherwise have paid more than the median benchmark after gaining benchmark access. Compression from below refers to the increase in salaries at firms that would otherwise have paid less than the median benchmark. Both directions of adjustment are documented empirically and are predicted by the model.&lt;/p&gt;
&lt;p&gt;Searched / Non-Searched / Non-Searchable Positions: The paper&amp;rsquo;s classification of new hires into three groups for identification purposes. Searched positions are those at treatment firms for which the firm actually looked up the benchmark. Non-Searched positions are at treatment firms but were not looked up, serving as a within-firm control. Non-Searchable positions are at control firms with no tool access, serving as a cross-firm control.&lt;/p&gt;
&lt;p&gt;Affiliation (across firm values): A technical condition borrowed from auction theory (Milgrom and Weber, 1982) used in the paper&amp;rsquo;s model to characterize the correlation structure of firms&amp;rsquo; private valuations of filling a position. Affiliation implies that when one firm has a high value, others are also more likely to have high values, and hence to offer high wages — generating the model&amp;rsquo;s equilibrium wage dispersion.&lt;/p&gt;
&lt;p&gt;Procompetitive Effect of Benchmarking: The paper&amp;rsquo;s term for the welfare-improving property of salary benchmarks identified in the model: by resolving aggregate uncertainty, benchmarks cause the marginal firm to offer closer to the full-information competitive wage, reducing extraordinary profits that arise from informational rents and raising the mean salary in equilibrium. This is the key concept in the paper&amp;rsquo;s contribution to the antitrust policy debate.&lt;/p&gt;</description></item><item><title>Why Doesn't the United States Have National Health Insurance?</title><link>https://macropaperwarehouse.com/papers/why-doesnt-the-united-states-have-national-health-insurance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/why-doesnt-the-united-states-have-national-health-insurance/</guid><description>&lt;p&gt;This paper investigates a critical juncture in the development of national health insurance (NHI) in the United States: the post-World War II period when most peer nations moved to establish comprehensive public coverage while the U.S. did not. The authors examine the causal role of the American Medical Association (AMA), which in 1949 hired Whitaker &amp;amp; Baxter&amp;rsquo;s Campaigns, Inc. — the country&amp;rsquo;s first political public relations firm — to direct a nationwide campaign opposing NHI and promoting private (voluntary) health insurance (PHI).&lt;/p&gt;
&lt;p&gt;The Campaign had two main components. First, a physician outreach component in which AMA members distributed pamphlets to patients warning against &amp;ldquo;socialized medicine&amp;rdquo; and encouraging enrollment in private plans, and acted as liaisons to local civic organizations to solicit resolutions against NHI sent to elected officials (nearly 50 million pieces of material were sent to physicians). Second, a mass newspaper advertising component, in which a standard ad was placed across newspapers nationwide, with an additional $19 million (approximately $240 million in current dollars) in coordinated tie-in advertising from roughly 23,000 corporations and industry associations. The messaging framed NHI as &amp;ldquo;un-American&amp;rdquo; and associated private insurance with &amp;ldquo;freedom&amp;rdquo; and &amp;ldquo;the American way,&amp;rdquo; providing little substantive information about insurance products.&lt;/p&gt;
&lt;p&gt;The authors construct novel measures of Campaign exposure by combining (a) per capita pamphlets distributed by AMA physicians and (b) per capita advertising circulation scaled by local newspaper readership, using archival data from the Whitaker &amp;amp; Baxter Archives (Sacramento), the National Archives (Washington D.C.), digitized AMA Medical Directories, the N.W. Ayer &amp;amp; Son&amp;rsquo;s Newspaper Directory, and newly discovered Blue Shield enrollment data from AMA Council on Medical Service annual reports covering 1946–1954.&lt;/p&gt;
&lt;p&gt;The primary estimation strategy exploits spatial variation in Campaign intensity combined with its timing, using event studies with state and year fixed effects and design controls for income per capita and unionization. The identifying assumption — that Campaign intensity was conditionally as-good-as-randomly assigned — is supported by balance tests showing no pre-Campaign correlation between exposure and enrollment or sociodemographic characteristics (with the exception of Black population share), and by the historical record that the Campaign was organized hastily following Truman&amp;rsquo;s unexpected 1948 electoral victory.&lt;/p&gt;
&lt;p&gt;Main findings: A one standard deviation increase in Campaign exposure explains approximately 20% of the post-Campaign increase in PHI enrollment, corresponding to roughly 14 million additional enrollees — an effect comparable in magnitude to increasing average per capita income by approximately $100 (about 7 percent). On public opinion, a one standard deviation increase in Campaign exposure led to a six percentage point decline in popular support for NHI per Gallup survey wave, a reversal occurring against a backdrop of 69% pre-Campaign approval that was trending upward. For context, this six-point magnitude approximates the entire gap in NHI support between union and non-union households, or one-third the racial gap in support. Campaign intensity also predicts civic organizations passing resolutions favoring PHI, Republican legislators adopting speech semantically similar to Campaign propaganda, and — by 1952 — AMA members being five times more likely to donate to the Eisenhower-Nixon ticket than non-AMA physicians, with donation rates increasing in Campaign intensity.&lt;/p&gt;
&lt;p&gt;Scope conditions: The analysis covers 48 U.S. states from 1946 to 1954, ending at the 1954 IRS tax code change that expanded commercial insurers&amp;rsquo; market share. The enrollment data capture Blue Shield (physician-run) plans specifically; the paper explicitly notes that commercial insurer granular data are unavailable for the main Campaign period. The authors argue that multiple subsequent factors — middle-class acquisition of private coverage reducing demand for a public option, incumbent interests defending the status quo, and the persistent ideological linkage of private insurance with freedom — help explain why NHI was not adopted in subsequent decades, though these persistence mechanisms are outside the paper&amp;rsquo;s direct empirical scope.&lt;/p&gt;
&lt;p&gt;Q: What was the AMA&amp;rsquo;s Campaign, and what prompted it?
A: In response to Harry Truman&amp;rsquo;s unexpected 1948 presidential victory alongside a Democratic Congress — and with a majority of informed voters favoring NHI — the AMA hired Whitaker &amp;amp; Baxter&amp;rsquo;s Campaigns, Inc. to run the National Education Campaign (NEC). The Campaign had two components: physician outreach (pamphlet distribution to patients, liaison to civic organizations) and mass newspaper advertising. The AMA paid Whitaker &amp;amp; Baxter approximately $1.2 million per year in current terms, and coordinated an additional $19 million in 1950 dollars (roughly $240 million today) in tie-in advertising from allied corporations and trade groups.&lt;/p&gt;
&lt;p&gt;Q: How is Campaign exposure measured, and how is it validated as conditionally exogenous?
A: Campaign exposure combines two standardized components: per capita pamphlets distributed by AMA physicians (pamphlet quantity from W&amp;amp;B archives scaled by state AMA membership share) and per capita advertising circulation scaled by local newspaper readership (share of adults with more than five years of schooling). The two components are summed and standardized. Exogeneity is supported by balance tables showing no pre-Campaign correlation between exposure and enrollment or Gallup opinion, by the absence of discontinuous changes in income or unionization at Campaign onset, and by the historical fact that Campaign logistics relied on pre-existing networks assembled hastily in response to Truman&amp;rsquo;s unanticipated victory.&lt;/p&gt;
&lt;p&gt;Q: What is the main effect of the Campaign on private health insurance enrollment?
A: A one standard deviation increase in Campaign exposure is associated with a two percentage point increase in the share enrolled in PHI in the preferred specification (Column 4 of Table 1, which includes income, unionization, state fixed effects, and year fixed effects; coefficient 0.020, se 0.007, significant at 1%). This accounts for approximately 20% of the overall post-Campaign increase in PHI enrollment, corresponding to roughly 14 million new enrollees. The pre-Campaign coefficient is not statistically significant (coefficient 0.004, se 0.005), and the F-test p-value for pre-trends is 0.958.&lt;/p&gt;
&lt;p&gt;Q: What is the effect of the Campaign on public opinion toward NHI?
A: Using Gallup survey data, a one standard deviation increase in Campaign exposure led to an approximately six percentage point decline in individual support for NHI legislation per survey wave, against a pre-Campaign approval level of 69% that was trending upward. The F-test p-value for pre-trends in the Gallup event study is 0.179. This six-point effect is approximately equal to the gap in NHI support between union and non-union households, and approximately one-third the racial gap in support.&lt;/p&gt;
&lt;p&gt;Q: What evidence links the Campaign to civic organizations and the legislative process?
A: The Campaign&amp;rsquo;s archives document all civic organizations &amp;ldquo;on record against compulsory health insurance,&amp;rdquo; meaning they had passed resolutions in favor of PHI. The authors find a positive relationship between Campaign intensity and civic organizations passing such resolutions at the county level. Resolutions sent to elected officials were traced to the Congressional Record and to physical folders in the National Archives; their semantic similarity to AMA-WB propaganda is confirmed. Republican legislators&amp;rsquo; speech in the 81st Congress shows increased similarity to Campaign language in proportion to Campaign intensity in their district or state, while Democrat legislators do not show this pattern. NHI and the AMA experienced spikes in mention frequency in the Congressional Record during this period.&lt;/p&gt;
&lt;p&gt;Q: Did the Campaign affect physician political behavior beyond the clinic?
A: By 1952, when the Republican platform had fully adopted the AMA&amp;rsquo;s position, AMA members were approximately five times more likely to donate to the Eisenhower-Nixon ticket than non-AMA physicians, with donation probability increasing in Campaign intensity. The authors digitized the donor list from the National Professional Committee for Eisenhower (NPCE) — a separate lobbying entity created because the AMA legally could not endorse candidates — and linked approximately 80% of physician donors to the AMA Medical Directory.&lt;/p&gt;
&lt;p&gt;Q: What alternative explanations for PHI growth does the paper address, and how?
A: The standard literature attributes PHI growth to the 1942 Stabilization Act wage freeze (which left benefits unconstrained), collective bargaining rights clarified in the late 1940s, and the 1954 IRS tax exemption for employer-paid premiums. The authors include income per capita and unionization as core design controls and show that their Campaign exposure coefficient is stable across specifications with and without these controls (coefficients of 0.025 and 0.020 in Table 1 Columns 1–2 vs. 3–4, respectively). The analysis stops in 1954 before the tax change, and the authors note that by 1952 roughly 63% of households already had some form of medical expense insurance.&lt;/p&gt;
&lt;p&gt;Q: What is the conceptual mechanism through which the Campaign operated?
A: The authors adapt Sobbrio (2011)&amp;rsquo;s indirect lobbying model. Voters hold uniform priors over whether NHI enactment yields net positive or negative social surplus. The private-sector advocate (AMA-WB) sends messages that shift voters&amp;rsquo; posterior beliefs toward the negative-surplus state and, simultaneously, encourage PHI enrollment, which reduces voters&amp;rsquo; private valuation of a public option. Because citizens were likely unaware of the coordinated tie-in advertising across industries and the financial motivation behind physician messaging, the framing operated through naive belief updating. The public-sector advocate (Truman administration, Committee for the Nation&amp;rsquo;s Health) was vastly outresourced — the CNH raised only $104,000 in 1949 — and faced legal constraints on executive lobbying.&lt;/p&gt;
&lt;p&gt;Q: What advertising tactics specifically characterized the Campaign, and what do they imply about mechanisms?
A: Campaign pamphlets and ads provided little or no substantive information about insurance products (coverage, eligibility, cost) and instead tied health insurance to ideological symbols: &amp;ldquo;freedom,&amp;rdquo; &amp;ldquo;the American way,&amp;rdquo; &amp;ldquo;the voluntary way,&amp;rdquo; and warnings about &amp;ldquo;socialized medicine.&amp;rdquo; Word clouds from Campaign materials confirm &amp;ldquo;America&amp;rdquo; and &amp;ldquo;freedom&amp;rdquo; as dominant terms. The authors connect this to behavioral models of advertising (Mullainathan, Schwartzstein and Shleifer 2008) whereby advertisers create or exploit associations to influence product beliefs. The absence of informational content is consistent with effects operating through ideology and identity rather than rational product evaluation.&lt;/p&gt;
&lt;p&gt;Q: What explains why the U.S. did not adopt NHI in subsequent decades after the immediate Campaign period?
A: The authors offer three mechanisms (discussed outside their main empirical scope): First, as middle-class Americans obtained PHI through employers, demand for a public option diminished — the model formalizes this as reduced private valuation of NHI. Second, incumbents who benefit from the private status quo — Blue Cross Blue Shield, AMA, American Hospital Association, and pharmaceutical companies, which today comprise four of the top ten direct federal lobbyists — actively work to maintain it (Acemoglu, Egorov and Sonin 2021). Third, the Campaign&amp;rsquo;s ideological framing proved durable: ideologically similar rhetoric opposing &amp;ldquo;socialized medicine&amp;rdquo; appeared in campaigns against both Clinton-era and Obama-era reform efforts, and has been linked to increased adverse selection and preventable deaths (Bursztyn et al. 2022; Galvani et al. 2022).&lt;/p&gt;
&lt;p&gt;Q: What are the paper&amp;rsquo;s main contributions to the literature?
A: The paper provides the first causal evidence on the AMA&amp;rsquo;s political role in blocking NHI at the post-WWII juncture, contributing to the economic history of U.S. social insurance development. It contributes to the advertising literature by providing credible estimates of a sustained national campaign combining trusted field agents (physicians) with mass media, and to the lobbying literature by documenting indirect lobbying — persuasion of ordinary citizens — as a distinct and effective tool alongside direct lobbying. It also documents physician behavior outside the clinical setting, showing how rents from supply-side constraints were deployed to shape the market for medical services.&lt;/p&gt;
&lt;p&gt;Indirect lobbying: In the paper&amp;rsquo;s usage, persuasion of ordinary citizens via campaigns — as distinct from direct lobbying of policymakers — used to shift median voter beliefs and behavior to achieve legislative goals. Whitaker &amp;amp; Baxter are credited with creating this field through their work at Campaigns, Inc.&lt;/p&gt;
&lt;p&gt;Campaign exposure: The paper&amp;rsquo;s composite treatment variable, constructed as the sum of two standardized components: per capita pamphlets distributed by AMA physicians (physician outreach) and per capita advertising circulation scaled by local newspaper readership (mass communications), then re-standardized to mean 0, standard deviation 1.&lt;/p&gt;
&lt;p&gt;Tie-in advertising: Coordinated newspaper advertisements by third-party corporations and trade associations placed simultaneously with the main AMA-WB Campaign ad, sharing the &amp;ldquo;Voluntary Way is the American Way&amp;rdquo; slogan. Approximately 60% of newspapers with a main Campaign ad also had tie-in ads, averaging three per issue; third-party spending totaled approximately $19 million in 1950 dollars (~$240 million current).&lt;/p&gt;
&lt;p&gt;Voluntary (private) health insurance: In the paper&amp;rsquo;s framing, the AMA-promoted alternative to NHI — prepaid medical service plans run by state medical societies (Blue Shield) or nonprofit hospitals (Blue Cross) — deliberately labeled &amp;ldquo;voluntary&amp;rdquo; to contrast with &amp;ldquo;compulsory&amp;rdquo; NHI, embedding the product within an ideological frame of free choice.&lt;/p&gt;
&lt;p&gt;National Education Campaign (NEC): The AMA&amp;rsquo;s official name for the anti-NHI campaign directed by Whitaker &amp;amp; Baxter starting in 1949, characterized as &amp;ldquo;educational&amp;rdquo; to provide legal cover; the name itself illustrates the indirect lobbying strategy of framing political advocacy as public information.&lt;/p&gt;
&lt;p&gt;Source text origin / abstract-only block: Not a paper-defined concept; excluded.&lt;/p&gt;
&lt;p&gt;Naive voter updating: The paper&amp;rsquo;s modeling assumption (drawn from Sobbrio 2011) that voters held uniform priors on health insurance policy outcomes and updated beliefs via Bayesian message receipt, without awareness of coordination across industries or the financial motivation of physician messengers — making the ideological framing effective.&lt;/p&gt;
&lt;p&gt;Physician field agents: In the Campaign&amp;rsquo;s design, AMA member physicians served as credible, trusted intermediaries who distributed pamphlets to patients and solicited civic organization resolutions, leveraging their social status to amplify the Campaign&amp;rsquo;s reach into communities where mass advertising alone would be insufficient.&lt;/p&gt;</description></item></channel></rss>