<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Economic-Growth | Macro Paper Warehouse</title><link>https://macropaperwarehouse.com/topics/economic-growth/</link><atom:link href="https://macropaperwarehouse.com/topics/economic-growth/index.xml" rel="self" type="application/rss+xml"/><description>Economic-Growth</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><item><title>A traffic-jam theory of growth</title><link>https://macropaperwarehouse.com/papers/a-traffic-jam-theory-of-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/a-traffic-jam-theory-of-growth/</guid><description>&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Finocchiaro and Weil ask whether financial development necessarily promotes long-run economic growth, or whether congestion externalities in R&amp;amp;D markets can offset — and even reverse — the growth benefits of easier credit access. The paper proposes that the empirical coexistence of expanding financial sectors and roughly constant per-capita GDP growth rates (approximately 2% annually in the United States over the last century) can be explained by the interplay of search frictions in two sequential markets: credit and innovation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Methodology.&lt;/strong&gt; The authors build a continuous-time endogenous growth model in which all growth is innovation-led. Firms must pass through four sequential stages — creation, fund-raising (Stage 0–1), R&amp;amp;D search (Stage 1–2), and high-productivity production (Stage 2–3) — before being exogenously destroyed. Both the credit market (firms searching for banks/venture capitalists) and the innovation market (firms searching for innovators after securing finance) are characterized by constant-returns-to-scale matching functions with endogenous market tightness. Nash bargaining determines the loan repayment, and free entry drives profits to zero in both markets. The model is then calibrated to annual U.S. data, with the risk-free rate r = 3.5%, separation rate s = 4%, symmetric bargaining power ω = 0.5, a productivity jump γ = 0.023 targeting a baseline growth rate of 2%, credit market duration for creditors just below one month and for firms slightly above one year (consistent with Wasmer and Weil, 2004), a two-year average patent approval time (USPTO 2020), 6% employment in finance (BLS 2020), and 0.5% employment in scientific R&amp;amp;D (BLS 2020).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Core Mechanism.&lt;/strong&gt; The paper derives a &amp;ldquo;spillover function&amp;rdquo; Q(p,g) that links the equilibrium probability of finding an innovator (q) to the probability of finding a bank (p) and the growth rate (g). Because free entry holds profits at zero, easier credit — a higher p — forces q downward: if a firm spends less time raising funds, the innovation market becomes more congested (Qp &amp;lt; 0). This negative spillover between the two markets is the paper&amp;rsquo;s central traffic-jam analogy: relieving one bottleneck shifts congestion downstream.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings.&lt;/strong&gt; The GG curve — the locus of (p, g) pairs consistent with equilibrium — is hump-shaped under the symmetric cost condition c = ωn (flow search cost for firms in credit markets equals the firm&amp;rsquo;s share of search costs in innovation markets). Growth is maximized when expected credit search time equals expected innovation search time (1/p = 1/q). Beyond that interior optimum, further financial deepening lowers the growth rate. The calibrated economy sits to the right of the hump in a flat region (p &amp;gt; q), so that reducing credit frictions alone has a marginally negative effect on growth: eliminating credit frictions lowers g from 2.000% to 1.997%, a reduction of 0.003 percentage points. Reducing innovation frictions alone raises g modestly to 2.071% (+0.071 pp). Only a simultaneous reduction of frictions in both markets raises g meaningfully, to 2.122% (+0.122 pp). The quantitative effects are deliberately small, consistent with the near-constancy of long-run growth despite financial deepening.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; The non-monotonicity requires both markets to carry search frictions; when only one friction is present, financial development is unambiguously good for growth (Section 4.3). The hump-shape is established analytically in the symmetric case c = ωn; more generally, the paper shows (via back-of-envelope approximation) that the sign of the finance–growth link depends on whether c/ω is less than or greater than n. The quantitative insensitivity of growth to finance is amplified when the real interest rate is close to the growth rate and when potential growth γ is close to actual growth g: the elasticity of growth with respect to finance is proportional to (γ − g)/γ. Extensions to fixed bank entry costs (introducing a growth-to-finance feedback), endogenous innovator wages (Section 4.2), and frictionless innovation (Section 4.3) all confirm the benchmark conclusions under stated parameter conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q1: What is the paper&amp;rsquo;s central theoretical claim about the finance–growth nexus?&lt;/strong&gt;
The paper claims that the finance–growth relationship is non-monotonic: financial development raises growth when credit is scarce (left of the hump on the GG curve) but lowers it when credit is readily available (right of the hump), because easier financing draws more firms into the innovation market, tightening it and reducing the probability of finding an innovator. This congestion spillover from the credit market to the innovation market is the &amp;ldquo;traffic-jam&amp;rdquo; mechanism. The non-monotonicity vanishes if either market lacks search frictions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q2: What is the &amp;ldquo;spillover function&amp;rdquo; and why is it central to the model?&lt;/strong&gt;
The spillover function Q(p, g) is derived from the free-entry zero-profit condition for firms and expresses the innovation-matching probability q consistent with equilibrium for given credit-matching probability p and growth rate g. It has Qp &amp;lt; 0 (easier credit reduces q) and Qg &amp;lt; 0 (faster growth reduces q), capturing the two-way negative interaction between the markets. It is central because all equilibrium and comparative-statics results flow through it: the GG curve is defined by substituting Q into the growth equation g = γ/(1 + s/p + s/Q(p,g)).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q3: Under what condition is the GG curve hump-shaped, and what is the intuition?&lt;/strong&gt;
The GG curve is hump-shaped when the flow search cost for firms in the credit market c equals the firm&amp;rsquo;s share of innovation search costs ωn (Proposition 4). The intuition mirrors equalizing travel times across two congested roads: growth is maximized when expected credit search time (1/p) equals expected innovation search time (1/q). When credit is very tight (p small), a marginal increase in p raises the share of innovating firms faster than it tightens the innovation market, so growth rises. Once credit is abundant (p large), the congestion effect on innovation dominates and growth falls.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q4: What does the benchmark calibration predict about the quantitative effect of financial development on growth?&lt;/strong&gt;
The benchmark calibration, targeting 2% annual U.S. growth, places the economy to the right of the hump in a flat region of the GG curve (p &amp;gt; q). Eliminating credit market frictions alone reduces the annual growth rate by 0.003 percentage points (from 2.000% to 1.997%) while lengthening expected innovation search time from 2 years to 3.4 years. This marginally negative effect arises because the economy is already well to the right of the optimum. The results are deliberately small and consistent with the empirical near-constancy of growth alongside financial deepening.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q5: What combination of policies does the model recommend for raising growth?&lt;/strong&gt;
Only a simultaneous reduction of frictions in both the credit and the innovation market raises the growth rate meaningfully, to 2.122% in the calibration (+0.122 pp relative to the 2.000% benchmark). Isolated improvements in credit markets have a marginally negative effect; isolated improvements in innovation markets have a marginally positive effect (+0.071 pp). The authors interpret this as supporting the OECD view that growth-stimulating policies should be designed as a system rather than as isolated pro-growth measures.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q6: How does the elasticity of growth to finance depend on the gap between potential and actual growth?&lt;/strong&gt;
The authors show (referenced as available on request) that the elasticity of the growth rate with respect to financial factors is proportional to (γ − g)/γ, where γ is the potential growth rate (the productivity jump per innovation) and g is the actual equilibrium growth rate. When actual growth is close to potential — as in the benchmark calibration with γ = 0.023 and g = 2.000% — this factor is near zero, making growth nearly insensitive to changes in financial conditions. This provides a structural rationale for why empirically measured finance–growth effects are often small or nil in advanced economies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q7: How does introducing fixed bank entry costs (Section 4.1) change the results?&lt;/strong&gt;
When banks bear a fixed licensing cost K (paid each time they enter the credit market), credit market tightness φ becomes an increasing function of (r − g)K: the annuity value of the fixed cost falls as growth rises, inducing more bank entry and reducing credit tightness. This introduces an upward-sloping PP curve (rather than a vertical one) and creates a direct positive feedback from growth to financial deepening. The qualitative conclusions on non-monotonicity are preserved: lower licensing costs shift the PP curve right and steepen it, with the equilibrium effect on growth remaining ambiguous due to the congestion spillover into the innovation market.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q8: What happens to the spillover function when innovators are paid (Section 4.2)?&lt;/strong&gt;
When innovators receive a Nash-bargained wage, the equilibrium wage (Equation 30) is increasing in innovator productivity (πγ), innovation market tightness (θn), and the growth rate, and decreasing in total credit market search costs K(φ). Easier credit raises both expected revenues and innovator wages for the firm. For innovator bargaining power α sufficiently small (and always for α &amp;lt; 1, as shown in the Appendix), the revenue effect dominates so that Qp &amp;lt; 0 is preserved: finance still creates bottlenecks in the innovation market, and the core non-monotonicity result carries through.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q9: What does the model predict when only one market has search frictions?&lt;/strong&gt;
When only the credit market is frictional and innovators are found instantly after financing is secured, improving credit market efficiency unambiguously raises growth (Section 4.3, Figure 4). The GG curve becomes g = γ/(s/p + 1), which is strictly increasing in p, and the PP curve shifts in a way that unambiguously raises equilibrium growth. The paper uses this case to isolate the source of non-monotonicity: the negative spillover from credit ease to innovation congestion requires frictions in both markets to operate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q10: How does the paper relate to the empirical &amp;ldquo;too much finance&amp;rdquo; literature?&lt;/strong&gt;
The paper offers a distinct theoretical mechanism for the inverted-U relationship between credit and productivity growth documented by Arcand et al. (2015), Aghion et al. (2019), and Popov (2018), among others. While Aghion et al. (2019) explain the inverted-U through less-efficient incumbents surviving longer with better credit access, and Malamud and Zucchi (2019) emphasize how financing frictions differentially affect entrant and incumbent composition, Finocchiaro and Weil&amp;rsquo;s mechanism operates through congestion externalities in sequential search markets — a channel not previously formalized in the innovation-led growth literature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Search frictions in credit markets:&lt;/strong&gt; Firms searching for financiers (banks or venture capitalists) and banks searching for firms face a matching technology with constant returns to scale; credit market tightness φ is the ratio of firms searching for banks to banks searching for firms, and the matching probability p(φ) is strictly decreasing in φ. Free entry drives bank profits to zero, pinning equilibrium tightness.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Search frictions in innovation markets:&lt;/strong&gt; After securing financing, firms search for innovators who can upgrade their productivity by factor γ; innovation market tightness θ is the ratio of firms searching for innovators to innovators, and the matching probability q(θ) is strictly decreasing in θ. The number of innovators is held fixed (analogously to fixed labor supply in Mortensen-Pissarides).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Spillover function Q(p, g):&lt;/strong&gt; Derived from the free-entry zero-profit condition for firms, Q expresses the equilibrium innovation-matching probability q as a function of the credit-matching probability p and the growth rate g. It has Qp &amp;lt; 0 and Qg &amp;lt; 0, meaning easier credit and faster growth both reduce q by tightening the innovation market. It is the formal embodiment of the traffic-jam mechanism.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GG curve:&lt;/strong&gt; The locus of (p, g) pairs consistent with the equilibrium growth equation g = γ/(1 + s/p + s/Q(p,g)). Under the symmetric cost condition c = ωn, the GG curve is hump-shaped: it rises from the origin, reaches a maximum interior growth rate, then declines toward an asymptote g∞ &amp;lt; γ. Its shape encodes the non-monotonic relationship between finance and growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PP curve:&lt;/strong&gt; The locus of equilibrium credit-matching probabilities consistent with free entry in the credit market. In the benchmark model it is a vertical line at p* = p(ω/(1−ω) · k/c), independent of q and g. When banks bear a fixed entry cost K, the PP curve becomes upward-sloping, introducing a direct positive feedback from growth to financial deepening.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Potential growth rate γ:&lt;/strong&gt; The productivity jump per successful innovation; in a frictionless world (p = q = ∞) the economy grows at γ. Actual growth g falls below γ to the extent that search frictions delay the delivery of credit and innovation. The elasticity of g to financial factors is proportional to (γ − g)/γ, so when actual and potential growth are close, financial factors matter little for growth.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Congestion externality in R&amp;amp;D:&lt;/strong&gt; The mechanism by which financial deepening — raising p — drives more firms to seek innovators, tightening the innovation market and reducing q. This negative spillover (Qp &amp;lt; 0) is the paper&amp;rsquo;s central departure from models with only a single friction, where finance is always growth-enhancing.&lt;/p&gt;</description></item><item><title>Abundance from Abroad: Migrant Income and Long-Run Economic Development</title><link>https://macropaperwarehouse.com/papers/abundance-from-abroad-migrant-income-and-long-run-economic-development/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/abundance-from-abroad-migrant-income-and-long-run-economic-development/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks how persistent increases in international migrant income prospects affect long-run economic development in migrant-origin areas. The central question is whether Philippine provinces with persistent access to higher-income migration opportunities develop faster than provinces with less attractive migration opportunities, and through which channels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Natural Experiment and Identification Strategy&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors exploit the 1997 Asian Financial Crisis as a large-scale natural experiment. The crisis triggered sharp, heterogeneous, and persistent exchange rate changes across Philippine migrants&amp;rsquo; destination countries — ranging from a 4% depreciation against the Philippine peso (Korea) to a 57% appreciation (Libya), with Japan and Saudi Arabia in between (appreciations of 32% and 52%, respectively). Because Philippine provinces differed in the pre-crisis distribution of migrant income across destinations (measured using unusual POEA/OWWA administrative contract data covering all overseas worker contracts, including migrant incomes, origins, and destinations), these exchange rate shocks generated exogenous, province-level variation in a shift-share instrument: the predicted change in province migrant income per capita due to the 1997 shocks. Identification follows the &amp;ldquo;exogenous shares&amp;rdquo; framework of Goldsmith-Pinkham et al. (2020). Pre-trend tests across up to 12 years of pre-shock panel data find no evidence of differential trends across provinces. The five destinations with the highest Rotemberg weights — Saudi Arabia, Japan, United States, Taiwan, and Hong Kong — collectively account for 75% of the identifying variation. The exchange rate shocks and the exposure weights both exhibit strong persistence over two decades post-1997.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Philippine government administrative data (POEA/OWWA) on all overseas worker contracts, 1992–2015, matched at 95% rate, providing province-of-origin and destination-specific migrant income.&lt;/li&gt;
&lt;li&gt;Philippine Family Income and Expenditure Survey (FIES), up to twelve triennial rounds from 1985–2018 (74 provinces, ~40,000 households per round), for domestic income and expenditure.&lt;/li&gt;
&lt;li&gt;Six rounds of the Philippine Census of Population (1990–2015) for education, migration rates, and sectoral employment shares.&lt;/li&gt;
&lt;li&gt;Province-level consumer price index data (1994–2017) and firm-level export survey data for robustness checks.&lt;/li&gt;
&lt;li&gt;Unit of analysis: 74 Philippine provinces (consistent 1990 borders).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Main Findings with Quantitative Magnitudes&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Six-fold magnification of migrant income&lt;/strong&gt;: Each unit of initial short-run shock (1997–1998) to migrant income per capita is magnified more than six-fold by 2009–2015. A one-standard-deviation shock (0.093) raises long-run migrant income per capita by 14.7% of the baseline mean (PhP 601 per capita, 0.2 standard deviations).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Domestic income gains predominate&lt;/strong&gt;: A one-standard-deviation shock raises domestic income per capita (excluding migrant income and remittances) by 6.4% of the baseline mean (PhP 1,676, 0.18 standard deviations). Remarkably, 73.6% of the long-run global income increase comes from domestic income and only 26.4% from migrant income.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Global income and expenditure&lt;/strong&gt;: A one-standard-deviation shock raises global income per capita by PhP 2,277 (0.2 standard deviations, or 7.5% of the baseline mean) in 2009–2015. Expenditure per capita rises by PhP 1,159 (0.13 standard deviations). Effects emerge gradually over two decades.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Education&lt;/strong&gt;: A one-standard-deviation shock increases the college-educated share of the population by 0.46–0.51 percentage points (0.11–0.12 standard deviations) and secondary completion by 0.63 percentage points. There is no significant effect on primary completion.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Migration rates and skill composition&lt;/strong&gt;: A one-standard-deviation shock increases the migration rate by 0.19 percentage points (0.22 standard deviations), raises the share of skilled migrants by 1.84 percentage points (0.19 standard deviations), and increases average migrant annual salary by PhP 23,703 (0.16 standard deviations). New migration concentrates in higher-education-quartile occupations.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Structural change&lt;/strong&gt;: The shock reduces primary sector employment shares by 1.2 percentage points per standard deviation (0.06 standard deviations), with over 70% of that shift absorbed by non-tradable goods and services sectors. Domestic income gains are driven almost entirely by non-agricultural income, and roughly 55% of the increase in entrepreneurial income is from service sectors.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Education&amp;rsquo;s contribution to income&lt;/strong&gt;: Model-based calculations assign 19.6% of the global income gain, 17.8% of the migrant income gain, and 20.2% of the domestic income gain to educational investments. Exchange rate persistence plus altered migration flows explain an additional 64.6% of the migrant income increase, so together these mechanisms account for 82.3% of the six-fold magnification. A demand multiplier (assuming 64% of migrant income returns to origin economies and a multiplier of 2.9, consistent with estimates from the literature) accounts for approximately 83.3% of the non-education-related portion of the domestic income increase.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Threats to Identification Ruled Out&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Import and export shift-share controls (constructed analogously using bilateral trade data and province-level industry employment shares) are uncorrelated with the migrant income shock and leave coefficient estimates unchanged. Province-level manufactured exports, agricultural income, the CPI, and national-level FDI inflows show no statistically significant response to the shock. Internal migration rates are unaffected. Geographic spillover controls and tourism controls do not alter results. Placebo regressions in the pre-period yield small, statistically insignificant coefficients.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper studies formal, government-regulated temporary labor migration from the Philippines, where migrants sign contracts through POEA-licensed agencies and typically expect to return after one or more contracts. The findings apply specifically to settings where persistent (not transitory) migrant income shocks occur. Approximately 60% of contract migrants are female. The study period spans 1985–2018, with main long-run outcome analyses comparing 1994 (pre-shock) with 2009–2015 (post-shock).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-makes-the-1997-asian-financial-crisis-useful-as-a-natural-experiment-for-this-papers-purposes"&gt;Q1. What makes the 1997 Asian Financial Crisis useful as a natural experiment for this paper&amp;rsquo;s purposes?&lt;/h3&gt;
&lt;p&gt;A1: The crisis was largely unanticipated by policymakers, international organizations, and financial markets, making it implausible that pre-1997 migration destination choices reflected anticipation of the shocks. Exchange rate changes were heterogeneous across destinations (ranging from a 4% depreciation to a 57% appreciation), and crucially, these changes proved highly persistent over two decades — regression coefficients of long-run exchange rate changes on the initial 1997–1998 shock are close to and statistically indistinguishable from 1 in nearly all post-shock periods. Combined with the province-specific variation in migrant destination exposure, this generates persistent, exogenous, and heterogeneous shocks to migrant income prospects across provinces.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-shift-share-variable-and-how-does-it-combine-shifts-and-shares"&gt;Q2. What is the shift-share variable, and how does it combine &amp;ldquo;shifts&amp;rdquo; and &amp;ldquo;shares&amp;rdquo;?&lt;/h3&gt;
&lt;p&gt;A2: The shift-share variable Shiftshareo equals the sum over destinations d of (ωdo0 × ΔRd), where ωdo0 is province o&amp;rsquo;s pre-shock migrant income per capita from destination d (the &amp;ldquo;exposure weight&amp;rdquo; or &amp;ldquo;share&amp;rdquo;), and ΔRd is the fractional change in destination d&amp;rsquo;s exchange rate from before to after the crisis (the &amp;ldquo;shift&amp;rdquo;). It captures the predicted change in province-level migrant income per capita due to the 1997 exchange rate shocks, and is derived directly from a theoretical model of migration. Identification relies on the &amp;ldquo;exogenous shares&amp;rdquo; approach of Goldsmith-Pinkham et al. (2020): the pre-1997 exposure weights are treated as as-good-as-randomly assigned conditional on controls, because they reflect historical migration networks formed well before the crisis.&lt;/p&gt;
&lt;h3 id="q3-why-is-the-six-fold-magnification-of-the-initial-migrant-income-shock-so-striking-and-what-does-the-structural-model-say-about-its-sources"&gt;Q3. Why is the six-fold magnification of the initial migrant income shock so striking, and what does the structural model say about its sources?&lt;/h3&gt;
&lt;p&gt;A3: The coefficient on migrant income per capita (6.463 in Panel D of Table 1) implies that for each unit of initial short-run migrant income shock, migrant income per capita is more than six units higher in 2009–2015 — a far larger response than a one-for-one pass-through would predict. The structural model, which augments a Fréchet-based gravity model of migration with endogenous education investments, accounts for 82.3% of this magnification. Education investments explain 17.8% of the migrant income increase; persistent favorable exchange rates and resulting shifts in migration flows across destinations explain an additional 64.6%. The Fréchet elasticity of migration flows with respect to destination wages is estimated at θ = 3.42 via PPML, implying that even partial reorientation of migrants toward now-higher-wage destinations substantially raises aggregate migrant income.&lt;/p&gt;
&lt;h3 id="q4-what-evidence-supports-the-parallel-trends-assumption-in-the-pre-shock-period"&gt;Q4. What evidence supports the parallel trends assumption in the pre-shock period?&lt;/h3&gt;
&lt;p&gt;A4: The authors present event study diagrams (Figure 2) showing no differential positive pre-trends in either expenditure per capita or domestic income per capita prior to 1997 — for domestic income, there is a statistically insignificant negative trend from 1985–1991 and no trend in 1991–1994. Placebo regressions estimated on the pre-period only (1985, 1988, 1991 as &amp;ldquo;pre,&amp;rdquo; 1994 and 1997 as &amp;ldquo;post&amp;rdquo;) yield small, statistically insignificant coefficients on both domestic income and expenditure. Balance tests focusing on the five high-Rotemberg-weight destination shares (Saudi Arabia, Japan, US, Taiwan, Hong Kong) — which collectively account for 75% of the identifying variation — also show no significant pre-trends in key outcomes across provinces with varying levels of exposure.&lt;/p&gt;
&lt;h3 id="q5-how-do-the-authors-rule-out-trade-flows-as-an-alternative-mechanism-for-the-estimated-income-effects"&gt;Q5. How do the authors rule out trade flows as an alternative mechanism for the estimated income effects?&lt;/h3&gt;
&lt;p&gt;A5: They construct separate import and export shift-share variables, analogous to the &amp;ldquo;China shock&amp;rdquo; of Autor et al. (2013), using baseline bilateral trade values (from COMTRADE, disaggregated to 36 ISIC industries), province-level employment shares in import and export industries (from the 1990 Census), and the same destination exchange rate shocks. These trade shift-share variables are uncorrelated with the migrant income shock after conditioning on baseline controls (Appendix Table A5). Including them as additional controls in Panel D of all main regression tables leaves the migrant income coefficient stable. Further, province-level manufactured exports per capita show no large or statistically significant response to the migrant income shock, agricultural income similarly shows no significant response, and consumer price indices are unresponsive — ruling out import price changes as a confound. FDI inflows at the national level also show no significant relationship with destination-country exchange rate shocks.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-composition-of-the-domestic-income-gains--where-do-they-come-from"&gt;Q6. What is the composition of the domestic income gains — where do they come from?&lt;/h3&gt;
&lt;p&gt;A6: Both wage income and entrepreneurial/rental income rise significantly and in similar magnitude, while &amp;ldquo;other income&amp;rdquo; (pensions, interest, dividends) shows no robust increase (Table 4). Non-agricultural income drives virtually the entire domestic income gain; agricultural income per capita is statistically insignificant (Table 5, columns 1–2). Within entrepreneurial income, approximately 55% of the increase is from service sectors, with manufacturing and primary sector entrepreneurial income showing insignificant effects at the 10% level (Table 5, columns 3–5). These patterns are consistent with the structural change finding: the shock shifts labor from primary sectors toward non-tradable goods and services rather than toward tradable manufacturing.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-global-income-concept-and-what-share-does-each-component-contribute"&gt;Q7. What is the &amp;ldquo;global income&amp;rdquo; concept and what share does each component contribute?&lt;/h3&gt;
&lt;p&gt;A7: Global income per capita is defined as the sum of domestic income per capita (earned within the Philippine economy, excluding all international transfers) and migrant income per capita (the full income earned abroad by a province&amp;rsquo;s international migrants, calculated from contract data). Of the long-run global income increase, 73.6% comes from domestic income and 26.4% from migrant income. A one-standard-deviation shock raises global income by PhP 2,277 per capita in 2009–2015 (0.2 standard deviations, or 7.5% of the baseline mean).&lt;/p&gt;
&lt;h3 id="q8-how-do-education-effects-translate-into-more-and-higher-skilled-migration"&gt;Q8. How do education effects translate into more and higher-skilled migration?&lt;/h3&gt;
&lt;p&gt;A8: A one-standard-deviation migrant income shock increases college completion by 0.46 percentage points and secondary completion by 0.63 percentage points (with no significant effect on primary completion), consistent with the shock raising the return to higher education in the broader population. These better-educated workers then migrate at higher rates: the share of migrants who are skilled (college-educated) rises by 1.84 percentage points per standard deviation. Migration increases are concentrated in the two highest-education quartiles of occupations (engineers, medical professionals, teachers in the 4th quartile; caregivers, restaurant workers, performing artists in the 3rd quartile), with no significant effect in the two lowest quartiles. Average annual migrant salary rises by PhP 23,703 per standard deviation (0.16 standard deviations).&lt;/p&gt;
&lt;h3 id="q9-what-mechanisms-does-the-structural-model-invoke-to-explain-the-domestic-income-gains"&gt;Q9. What mechanisms does the structural model invoke to explain the domestic income gains?&lt;/h3&gt;
&lt;p&gt;A9: The model treats domestic income changes as arising through at least two channels: (1) the education channel, which the model assigns 20.2% of the domestic income increase (using the estimated college completion response of 0.046 per unit shock, baseline skill-migration probabilities, and baseline skill premia for domestic income); and (2) a demand multiplier operating on the portion of migrant income remitted to origin provinces, combined with capital accumulation from sustained migrant income flows. Assuming 64% of migrant income returns to origin economies (estimated indirectly from KNOMAD/ILO and Survey on Overseas Filipinos data) and a multiplier of 2.9 (consistent with estimates from Kenya and India), this demand-plus-investment channel can explain approximately 83.3% of the remaining (non-education-related) domestic income increase of PhP 14.4 per unit shock. Under baseline assumptions (α = 0.64), the stylized dynamic model generates PhP 18.88 of domestic income by 2015 from a PhP 1 initial shock — close to the empirical estimate of PhP 18.02.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-authors-assess-sutva-and-internal-migration"&gt;Q10. How do the authors assess SUTVA and internal migration?&lt;/h3&gt;
&lt;p&gt;A10: They test whether the migrant income shock affects net internal migration rates at the provincial level (Appendix Table A6) and find no large or statistically significant impact. There is a small negative effect on outmigration of young adults (aged 16–24) that the authors judge cannot account for the documented income impacts. The Philippines&amp;rsquo; archipelago geography (over 7,000 islands) is noted as likely limiting inter-provincial economic spillovers; to the extent spillovers occur, they would be positive (demand spillovers from provinces experiencing income gains to neighboring provinces), making estimates conservative lower bounds. Direct tests controlling for the inverse-distance-weighted migrant income shock in neighboring provinces leave main estimates unchanged.&lt;/p&gt;
&lt;h3 id="q11-are-the-exposure-weights-migration-shares-persistent-and-does-this-support-interpreting-the-shock-as-persistent"&gt;Q11. Are the exposure weights (migration shares) persistent, and does this support interpreting the shock as persistent?&lt;/h3&gt;
&lt;p&gt;A11: Yes. Regressions of dyadic migrant income per capita in post-shock years (2009, 2012, 2015) on dyadic migrant income per capita in 1995 yield coefficients ranging from 0.4 to 0.6, each statistically significantly different from zero (and from 1, indicating partial but substantial persistence). The exchange rate shocks ΔRd are even more persistent: regression coefficients on the initial 1997–1998 shock are close to 1 and statistically indistinguishable from 1 in nearly all post-shock periods (with the only exceptions in 2009–2012 during the Great Recession). Both components of the shift-share variable thus show persistence over two decades, supporting interpretation of the long-run effects as responses to a persistent (not transitory) income shock.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-and-how-do-the-authors-connect-findings-to-migration-policy"&gt;Q12. What are the policy implications and how do the authors connect findings to migration policy?&lt;/h3&gt;
&lt;p&gt;A12: The findings suggest migration policy should be an important part of the development policy toolkit. The results are directly relevant to origin-country policies facilitating formal, contract-based labor migration (e.g., regulation of recruitment agencies, educational investments to raise worker skills and competitiveness for overseas employment) and destination-country policies governing legal immigration opportunities. The authors also note implications for overseas development assistance: development agencies could consider supplementing traditional foreign aid with programs that facilitate international labor migration. The paper&amp;rsquo;s context — formal, government-regulated migration through POEA and OWWA — is described as highly policy-relevant, with 94% of developing countries with populations exceeding 1 million having a dedicated government migration agency and 78% having policies promoting migrant remittances.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Shift-share variable (Shiftshareo):&lt;/strong&gt; The paper&amp;rsquo;s primary independent variable, equal to the sum over all overseas destinations d of (ωdo0 × ΔRd) — the province&amp;rsquo;s pre-shock migrant income per capita from each destination (the exposure weight or &amp;ldquo;share&amp;rdquo;) multiplied by that destination&amp;rsquo;s exchange rate shock (the &amp;ldquo;shift&amp;rdquo;). It is the predicted change in province migrant income per capita due to the 1997 Asian Financial Crisis exchange rate shocks, and is derived directly from the theoretical model of migration (Equation A9). Identification treats the exposure weights as exogenous following the &amp;ldquo;exogenous shares&amp;rdquo; approach of Goldsmith-Pinkham et al. (2020).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exposure weights (ωdo0):&lt;/strong&gt; Province o&amp;rsquo;s pre-shock aggregate migrant income per capita earned in destination d, calculated from administrative POEA/OWWA contract data for 1995. These serve as the &amp;ldquo;shares&amp;rdquo; in the shift-share and capture the extent to which a province&amp;rsquo;s residents are exposed to a given destination&amp;rsquo;s exchange rate shock. They reflect historically-formed migration networks rather than anticipation of future shocks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Global income per capita:&lt;/strong&gt; The sum of domestic income per capita and migrant income per capita. Domestic income is household income earned within the Philippine economy (wages, entrepreneurial, and other sources), explicitly excluding all income from international sources including remittances. Migrant income is the full income earned abroad by all international migrants from the province, calculated from contract data (not remittances sent home). Global income thus captures the full resource gain available to a province from the combination of domestic production and international migration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Magnification (of migrant income shock):&lt;/strong&gt; The empirical finding that the long-run coefficient on migrant income per capita (6.463 in Panel D, Table 1) far exceeds 1 — meaning each unit of initial short-run shock becomes more than six units of migrant income per capita in 2009–2015. The paper decomposes this magnification into contributions from persistent exchange rates, educational investments raising skill levels and migration, and shifts in migration flows toward now-higher-wage destinations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Brain gain:&lt;/strong&gt; The paper&amp;rsquo;s term for the process by which improved migrant income prospects raise educational investments among the broader population (not just among migrants), leading to higher skill levels among non-migrants as well. The paper distinguishes this from &amp;ldquo;brain drain&amp;rdquo; (where migration of skilled workers reduces origin-area human capital) and provides evidence of a &amp;ldquo;virtuous cycle&amp;rdquo;: education raises migration rates and migrant skill levels, which in turn raises migrant and domestic incomes, potentially funding further education.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rotemberg weights:&lt;/strong&gt; Province-destination-level weights (following Goldsmith-Pinkham et al. 2020) characterizing which destination-specific exchange rate shocks drive the estimates most. Saudi Arabia (0.20), Japan (0.19), United States (0.18), Taiwan (0.10), and Hong Kong (0.08) together account for 75% of the total Rotemberg weight. These weights guide which destination-specific exposure shares receive the most scrutiny in pre-trend and balance tests.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fréchet elasticity (θ):&lt;/strong&gt; The elasticity of migration flows from an origin province to a destination with respect to destination wages (in Philippine pesos), estimated at 3.42 via PPML using the exchange rate shocks. This parameter governs how much migration flows — and thereby migrant income — respond to the persistent exchange rate changes, and is central to the model&amp;rsquo;s decomposition of the six-fold magnification of migrant income effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Domestic income multiplier:&lt;/strong&gt; The ratio of long-run domestic income increase to the portion of the migrant income shock that returns to origin provinces. Assuming 64% of migrant income returns to origin economies (estimated from multiple administrative data sources), the implicit demand multiplier in the paper&amp;rsquo;s context ranges from about 2.9 to 3.4, consistent with multipliers found in related literature on cash transfers and credit supply shocks in low-income settings.&lt;/p&gt;</description></item><item><title>All Along the Watchtower: Military Landholders and Serfdom Consolidation in Early Modern Russia</title><link>https://macropaperwarehouse.com/papers/all-along-the-watchtower-military-landholders-and-serfdom-consolidation-in-early-modern-russia/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/all-along-the-watchtower-military-landholders-and-serfdom-consolidation-in-early-modern-russia/</guid><description>&lt;p&gt;This paper investigates the origins of serfdom in early modern Russia, arguing that the institution consolidated primarily through political economy dynamics between the crown and a landholding military class, rather than from economic fundamentals such as labor scarcity, land-labor ratios, or grain trade opportunities. The central argument is that the prolonged defense of Russia&amp;rsquo;s southern frontier against Crimean Tatar nomadic raids generated a class of military landholders who possessed both the coercive capacity and the political leverage to press the state into restricting peasant labor mobility.&lt;/p&gt;
&lt;p&gt;The mechanism runs as follows. The Russian state, lacking the fiscal capacity to pay soldiers directly, granted frontier lands along the Tula defense line to high-ranked soldiers in exchange for military service under the pomest&amp;rsquo;e system. These lands were selected for their defensive rather than agricultural value and sat on the forest-steppe boundary roughly 180 km south of Moscow. Since soldiers could not farm while on duty and could not compete in free labor markets given the area&amp;rsquo;s low agricultural attractiveness, the arrangement was only sustainable if peasants were bound to the land. Military landholders collectively petitioned the Tsar repeatedly — with petition volumes peaking during urban uprisings (9 petitions in 1648, 13 in 1682) when the government&amp;rsquo;s political vulnerability increased the military&amp;rsquo;s bargaining power — until serfdom was codified in the Law Code of 1649.&lt;/p&gt;
&lt;p&gt;The authors test this theory using newly digitized data from the 1678 household census, which records male population by six legally distinct peasant categories across 172 districts of Muscovy, combined with data on landholder estate counts and sizes. The primary empirical finding is that districts on the Tula defense line had approximately 40% of their population composed of serfs, compared to roughly 14% nationally — a difference of about 25 percentage points that survives the inclusion of geographic and climatic controls (grain suitability, temperature seasonality, precipitation, terrain ruggedness, river location, distance to Moscow, and regional fixed effects). Placebo tests confirm this pattern is specific to the most legally dependent peasant groups: the defense line is negatively associated with royal peasants and statistically insignificant for church peasants, free peasants, and non-Russian peasants.&lt;/p&gt;
&lt;p&gt;To address potential endogeneity of the defense line&amp;rsquo;s location, the authors construct an instrumental variable using a novel geospatial algorithm. The algorithm computes optimal nomadic invasion routes from Crimea to Moscow via topographic cost rasters (using flow accumulation values as proxies for river-crossing barriers), then intersects these routes with the historically stable forest-steppe boundary (identified through FAO/UNESCO soil types — Podzoluvisols versus Chernozems). Districts at this intersection were 70 percentage points more likely to host the actual defense line. Two-stage least squares estimates confirm and slightly exceed the OLS magnitudes, supporting the causal interpretation.&lt;/p&gt;
&lt;p&gt;The paper further tests two canonical alternative explanations and finds them insufficient. Domar&amp;rsquo;s (1970) labor-scarcity hypothesis predicts serfdom should be higher where population density is lower; the data show the opposite sign, contradicting this prediction. The Baltic grain trade hypothesis yields only a small, unstable positive interaction between river access to the Baltic and grain suitability, which disappears when the defense line variable is included. A horse race including all variables simultaneously shows the defense line coefficient at approximately 24 percentage points remains stable while alternative predictors become insignificant.&lt;/p&gt;
&lt;p&gt;Mechanism tests show that defense line districts had 3.2 more estates per 100 square kilometers than the national average of 2.3, with the excess concentrated in very small (up to 5 serf households) and small (6–25 households) estates — consistent with the state&amp;rsquo;s strategy of maximizing soldier count by allocating the minimum serf labor sufficient to sustain a cavalryman. A bigram similarity analysis of collective petitions versus the 1649 Law Code yields a correlation coefficient of 0.7 for the top twenty bigrams between a 1637 petition and Chapter 11 (restricting peasant mobility), with no comparable similarity to other chapters. Persistence is documented through 1719, 1795, and 1858 censuses: defense line districts maintained the highest serf concentration through to three years before emancipation in 1861.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-central-argument-about-the-origins-of-russian-serfdom"&gt;Q1. What is the paper&amp;rsquo;s central argument about the origins of Russian serfdom?&lt;/h3&gt;
&lt;p&gt;A: The paper argues that serfdom consolidated primarily due to political economy dynamics: the crown&amp;rsquo;s dependence on a landholding military class for frontier defense against steppe nomads gave that class sufficient political leverage to secure the legal restriction of peasant labor mobility. The military landholders&amp;rsquo; coercive capacity and proximity to their small estates made labor coercion a viable complement to their military function. This explanation dominates alternative accounts based on labor scarcity, grain trade, or soil quality in all specifications tested.&lt;/p&gt;
&lt;h3 id="q2-what-was-the-tula-defense-line-and-why-was-it-located-where-it-was"&gt;Q2. What was the Tula defense line and why was it located where it was?&lt;/h3&gt;
&lt;p&gt;A: The Tula defense line (Great Abatis Line) was a chain of about 40 fort towns stretching over 500 km east-west, centered on Tula approximately 180 km south of Moscow, erected in the 1560s using felled trees, earth mounds, ditches, and watchtowers. Its location on the forest-steppe boundary was determined by two military-logistical constraints: it had to block the main nomadic invasion routes from Crimea, and it had to lie within the forest zone where timber was the cheapest construction material and which provided natural shelter. The paper documents that the defense line area did not differ from the rest of Muscovy in agricultural suitability, annual precipitation, seasonality, or terrain ruggedness — its distinctive feature was purely defensive.&lt;/p&gt;
&lt;h3 id="q3-how-large-is-the-estimated-effect-of-defense-line-proximity-on-serf-concentration"&gt;Q3. How large is the estimated effect of defense line proximity on serf concentration?&lt;/h3&gt;
&lt;p&gt;A: In the unconditional specification, defense line districts had a 30 percentage point higher share of serfs than the rest of the country. After adding geographic controls (grain suitability, seasonality, precipitation, terrain ruggedness, river dummy, distance to Moscow, and regional fixed effects), the coefficient stabilizes at approximately 25 percentage points. Given that serfs averaged about 14% of total population nationally but about 40% in defense line districts, the estimated effect is substantial relative to the baseline.&lt;/p&gt;
&lt;h3 id="q4-how-do-the-authors-address-endogeneity-of-the-defense-line-location"&gt;Q4. How do the authors address endogeneity of the defense line location?&lt;/h3&gt;
&lt;p&gt;A: They construct an instrumental variable defined as the intersection of two variables: districts lying on the computed optimal nomadic invasion routes (covering 98 of 172 districts, or 57% of the sample), and districts on the forest-steppe soil boundary (38 districts, or 22% of the sample). Their interaction covers 23 districts and is the excluded instrument. In the first stage, this interaction term raises a district&amp;rsquo;s probability of hosting the actual defense line by 70 percentage points, while the linear terms become essentially zero once the interaction is included. The 2SLS second-stage estimates of the serf-share effect are slightly higher than OLS and statistically significant, confirming the direction and approximate magnitude of the OLS results.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-paper-find-about-domars-labor-scarcity-hypothesis"&gt;Q5. What does the paper find about Domar&amp;rsquo;s labor-scarcity hypothesis?&lt;/h3&gt;
&lt;p&gt;A: The paper finds no support for Domar&amp;rsquo;s (1970) prediction that serfdom should be more prevalent where labor is scarcer (lower population density). Controlling for grain suitability and geographic factors, population density enters with a positive and statistically significant coefficient at the 5% level — the opposite sign from what Domar&amp;rsquo;s theory predicts. When the defense line dummy is added, population density becomes insignificant while the defense line coefficient remains at approximately 25 percentage points, consistent with the baseline.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-paper-find-about-the-baltic-grain-trade-hypothesis"&gt;Q6. What does the paper find about the Baltic grain trade hypothesis?&lt;/h3&gt;
&lt;p&gt;A: An exogenous measure of Baltic trade potential — a dummy for districts with river access to the Baltic, interacted with grain suitability — yields a small and marginally positive effect on serf share in Baltic districts with higher grain suitability. However, this effect disappears when the defense line dummy is included, and is also sensitive to alternative spatial clustering (becoming insignificant at the 300 km clustering radius even without the defense line dummy). The authors interpret this instability as inconsistent with grain trade being a primary driver of serfdom.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-evidence-for-the-estate-size-mechanism"&gt;Q7. What is the evidence for the estate-size mechanism?&lt;/h3&gt;
&lt;p&gt;A: Defense line districts had on average 3.2 more estates per 100 square kilometers than the national average of 2.3 per 100 square kilometers. Among estate-size brackets, very small (up to 5 serf households) and small (6–25 serf households) estates were disproportionately concentrated in defense line districts, while the location of medium-sized and large estates was statistically independent of the defense line. This pattern is consistent with the state&amp;rsquo;s strategy of allocating minimum viable serf endowments to maximize the number of soldiers supportable along the line.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-textual-evidence-linking-military-petitions-to-the-1649-law-code"&gt;Q8. What is the textual evidence linking military petitions to the 1649 Law Code?&lt;/h3&gt;
&lt;p&gt;A: A bigram similarity analysis between a 1637 collective petition and Chapter 11 of the 1649 Law Code reveals a correlation coefficient of 0.7 for the top twenty bigrams. The five most common bigrams appear in both texts: &amp;ldquo;runaway peasants,&amp;rdquo; &amp;ldquo;commoner peasants,&amp;rdquo; &amp;ldquo;census books,&amp;rdquo; &amp;ldquo;search years,&amp;rdquo; and &amp;ldquo;tsar&amp;rsquo;s decree.&amp;rdquo; This correlation does not extend to other chapters of the Law Code that regulate non-peasant matters, establishing specificity of the legislative influence.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-timing-of-collective-petitions-relate-to-political-crises"&gt;Q9. How does the timing of collective petitions relate to political crises?&lt;/h3&gt;
&lt;p&gt;A: Over a corpus of 96 petitions between 1608 and 1698, landholders petitioned on average once per year, but activity spiked sharply during domestic uprisings: 9 petitions in 1648 (the &amp;ldquo;Salt Riot&amp;rdquo; urban uprising) and 13 petitions in 1682 (the musketeers&amp;rsquo; revolt). These peaks coincide with moments when the government&amp;rsquo;s political vulnerability increased the military&amp;rsquo;s bargaining power, and in both cases were followed by legislative concessions — the 1649 Law Code and new decrees in 1683–85 on harsher punishment for harboring runaways, respectively.&lt;/p&gt;
&lt;h3 id="q10-what-do-the-placebo-tests-show"&gt;Q10. What do the placebo tests show?&lt;/h3&gt;
&lt;p&gt;A: Regressions of non-serf peasant shares on the defense line dummy show that the defense line is negatively associated with royal peasants and statistically insignificant for church peasants, free peasants, and non-Russian peasants. A placebo test replacing military landholders with merchants and artisans shows no significant defense line effect on the latter group, while Moscow has an 11 percentage point higher merchant/artisan share. The specificity of the defense line effect to legally dependent peasants and military landholders supports the military-political mechanism rather than a generic frontier-area effect.&lt;/p&gt;
&lt;h3 id="q11-how-persistent-was-the-spatial-distribution-of-serfdom-after-1649"&gt;Q11. How persistent was the spatial distribution of serfdom after 1649?&lt;/h3&gt;
&lt;p&gt;A: The authors estimate their baseline equation with serf share from the 1719, 1795, and 1858 censuses as dependent variables. Defense line districts maintained disproportionately higher serf densities in all three periods, including when the sample is restricted to the original Muscovite districts to exclude post-18th century territorial acquisitions. By 1858, three years before emancipation, the spatial distribution of serfs remained similar to that observed 200 years earlier at the time of serfdom&amp;rsquo;s consolidation — despite the defense line having been militarily obsolete for over a century.&lt;/p&gt;
&lt;h3 id="q12-what-explains-the-persistence-of-serfdom-beyond-its-original-military-rationale"&gt;Q12. What explains the persistence of serfdom beyond its original military rationale?&lt;/h3&gt;
&lt;p&gt;A: The persistence reflects a mutually beneficial exchange between the crown and former military landholders. Landholders provided local state capacity — overseeing tax collection, administering military conscription, and adjudicating peasant disputes through estate courts — in lieu of a centralized bureaucracy. In return, the crown granted successive expansions of landholder rights: Peter I equalized military landholdings with hereditary estates in 1714, and Peter III in 1762 freed landholders from military service obligations while retaining their property rights over land and serfs. This fiscal-administrative dependency is also cited as a reason for the late timing and unfavorable-to-peasants terms of the 1861 emancipation reform.&lt;/p&gt;
&lt;h3 id="q13-how-does-this-papers-explanation-relate-to-easternwestern-european-institutional-divergence"&gt;Q13. How does this paper&amp;rsquo;s explanation relate to Eastern/Western European institutional divergence?&lt;/h3&gt;
&lt;p&gt;A: The paper argues that while the military revolution in Western Europe generated fiscally capable centralized states with regular infantry armies, Russia&amp;rsquo;s peripheral nomadic threat prolonged the feudal cavalry model supported by land grants and serf labor. This delayed the formation of Weberian bureaucracy and entrenched what the authors term a &amp;ldquo;garrison state&amp;rdquo; — one whose institutions and social structure were shaped primarily by military-security considerations. The paper positions military factors alongside existing divergence explanations emphasizing land property rights, political institutions, demographic regimes, and Enlightenment ideas.&lt;/p&gt;
&lt;h3 id="q14-what-is-the-methodological-contribution-of-the-optimal-invasion-route-algorithm"&gt;Q14. What is the methodological contribution of the optimal invasion route algorithm?&lt;/h3&gt;
&lt;p&gt;A: The algorithm uses flow accumulation rasters (proportional to river width and basin size) as a cost function to compute the lowest-cost travel paths from Crimea to Moscow, iteratively penalizing cells within 15 km of each computed route and re-running the path search to generate four distinct routes per origin point (eight total, including routes from the Don River steppe). This produces a high-resolution, geographically continuous measure of military threat exposure that the authors argue provides statistical power in contexts where terrain ruggedness or simple distance measures lack variation — particularly relevant for flat plains with a single threat origin correlated with other variables.&lt;/p&gt;
&lt;p&gt;Pomest&amp;rsquo;e system: The institutional arrangement by which the Russian state granted frontier lands to high-ranked soldiers in exchange for military service, under the rule that &amp;ldquo;the land must not leave the service.&amp;rdquo; Unlike hereditary estates, pomest&amp;rsquo;e holdings were conditional on active service and could not be passed to heirs unless sons continued military service. This system enabled the formation of a permanent cavalry force despite the state&amp;rsquo;s low fiscal capacity, but required binding peasants to the land to make the arrangement viable for the soldier-landholders.&lt;/p&gt;
&lt;p&gt;Serfs (bobyli and dvorovye): In the paper&amp;rsquo;s 1678 census framework, serfs are defined as the two most legally dependent subgroups of private peasants — cotters (bobyli), who owned no property and worked full-time for their landlord in exchange for payment in kind, and servants (dvorovye), who performed household and support functions on the estate. These groups constituting about 14% of total population nationally were totally dependent on their landlord and could not retain the marginal product of any part of their labor. After the 1649 Law Code, villeins (krest&amp;rsquo;yane) gradually converged to this status as well.&lt;/p&gt;
&lt;p&gt;Collective petitions (chelobitnye): The primary institutional channel through which the military landholder class communicated collective interests and applied political pressure on the crown in 17th-century Muscovy. The paper documents 96 such petitions between 1608 and 1698, showing that their volume, timing (peaking during urban uprisings), and textual content (closely matching Chapter 11 of the 1649 Law Code) were the proximate mechanism by which landholders converted military leverage into legal codification of serfdom.&lt;/p&gt;
&lt;p&gt;Optimal defense line (instrumental variable): The paper&amp;rsquo;s constructed instrument, defined as the intersection of computed optimal nomadic invasion routes (based on topographic cost rasters approximating river-crossing barriers) and the forest-steppe soil boundary (Podzoluvisols/Chernozems boundary from the FAO/UNESCO Soil Map). This instrument captures the geographically and militarily determined placement of defensive fortifications, purging variation in actual defense line location that might reflect agricultural or economic value.&lt;/p&gt;
&lt;p&gt;Garrison state: Used by the authors (adapting Lasswell&amp;rsquo;s term) to describe a state whose institutions and social structure are shaped primarily by military security considerations. In the Russian context, this refers to the persistence of a feudal cavalry system, land-grant-based military compensation, and labor coercion that together delayed centralized state formation and Weberian bureaucracy relative to Western European states undergoing the military revolution toward regular infantry armies.&lt;/p&gt;
&lt;p&gt;Labor coercion complementarity: The paper&amp;rsquo;s mechanism whereby employers with high coercive capacity (proximity to weapons, military training) can deploy that same capacity to restrict workers&amp;rsquo; outside options and extract labor surplus. In the defense line context, soldiers&amp;rsquo; military skills and armament made them effective at preventing serf flight and enforcing labor obligations — creating a complementarity between military capacity and serfdom that was absent among merchants or church institutions with comparable landholdings elsewhere.&lt;/p&gt;</description></item><item><title>Artificial intelligence and technological unemployment</title><link>https://macropaperwarehouse.com/papers/artificial-intelligence-and-technological-unemployment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/artificial-intelligence-and-technological-unemployment/</guid><description>&lt;p&gt;Wang and Wong develop a continuous-time labor-search model to assess the dynamic effects of generative AI (GenAI) on labor productivity and unemployment. The paper is motivated by conflicting empirical evidence: micro studies find productivity gains of 14% (Brynjolfsson, Li, and Raymond 2025) and 55.8% faster coding (Peng et al. 2023), while macro estimates suggest modest TFP gains of at most 0.064% annually (Acemoglu 2024), and occupation-level evidence shows a 13% relative employment decline in AI-exposed jobs (Brynjolfsson, Chandar, and Chen 2025).&lt;/p&gt;
&lt;p&gt;The model distinguishes GenAI from earlier automation technologies by its learning-by-using mechanism: AI capability grows at rate µ per employed worker (law of motion dAt/At = µHt − δ), raises employed workers&amp;rsquo; productivity, and creates a displacement threat through renegotiation. When renegotiation fails, AI replaces the worker, generating technological unemployment. Firms renegotiate wages at a rate ρµAt proportional to AI&amp;rsquo;s learning rate and the job&amp;rsquo;s exposure ρ. The joint surplus condition governs whether replacement occurs: AI replaces a worker if and only if πA (AI&amp;rsquo;s net present value per output) exceeds the post-renegotiation joint surplus St.&lt;/p&gt;
&lt;p&gt;The model admits three steady states: (i) a some-AI steady state with finite AI capability, persistent AI adoption (It = 1), expanded job creation but declining employment at H∞ = δ/µ; (ii) an unbounded-AI equilibrium with sustained endogenous growth, no displacement (It = 0), and employment at H∞ = α/(α+σ); and (iii) a no-AI equilibrium reverting to the Mortensen-Pissarides benchmark. In the benchmark model (exogenous job-finding rate, AI-augmented productivity), multiple steady states can coexist—global indeterminacy—when condition (28) holds. In the full model (endogenous job creation via free entry), both global and local indeterminacy are possible, and a continuum of oscillatory transition paths converge to the some-AI steady state.&lt;/p&gt;
&lt;p&gt;Calibrated to U.S. data, targeting a pre-AI unemployment rate of 5%, AI elasticity of productivity εy = 1.069 (from Czarnitzki et al. 2023), initial AI productivity boost of 14% (Brynjolfsson et al. 2025), worker exposure ρ = 0.618 (Brynjolfsson et al. 2018&amp;rsquo;s machine learning suitability index), AI replacement cost ϕ = 0.0043 (from U.S. business GenAI spending), AI learning rate µ = 0.632, and AI error rate δ = 0.462 (Moore&amp;rsquo;s law half-life of 1.5 years), the model converges to a some-AI steady state. The long-run results are: a 23% employment loss (H∞ = 0.732 vs. H0 = 0.95), AI capability improvement of 321%, and labor productivity gain of 366%. Approximately half of the employment loss—11.5 percentage points—occurs within the first five years, alongside a 49.3% output gain and 45.5% AI capability improvement over that period.&lt;/p&gt;
&lt;p&gt;Untargeted moments are validated: the model implies 7.08% labor productivity growth over the first 10 years (consistent with Briggs and Kodnani 2023) and an AI elasticity of vacancies averaging 0.16 over the first five years (consistent with Acemoglu et al. 2022).&lt;/p&gt;
&lt;p&gt;On welfare, equilibria are inefficient even when the Hosios condition holds. AI introduces four externalities beyond standard matching frictions: job destruction via displacement, productivity enhancement for employed workers, feedback from AI learning depending on employment, and direct effects on matching surpluses. A constrained-optimal subsidy to jobs at risk of AI displacement is 26.6% in the short run and exceeds 50% in the long run. In the full model, the Hosios condition requires fixing firm bargaining power θ to the vacancy elasticity of matching ξ, but an additional per-output transfer T = µApωA to firm-worker matches is necessary to correct AI adoption inefficiency.&lt;/p&gt;
&lt;p&gt;Q: What is the core mechanism by which AI generates unemployment in this model?
A: AI capability grows through a learning-by-using process (dAt/At = µHt − δ), improving as it observes employed workers. As capability rises, firms gain a displacement option that arrives at rate ρµAt per matched pair. When renegotiation over wages fails—i.e., when the AI&amp;rsquo;s NPV πA exceeds the joint surplus—firms replace workers with AI, causing unemployment. This creates a feedback loop: higher employment accelerates AI learning, which increases displacement pressure and reduces employment.&lt;/p&gt;
&lt;p&gt;Q: What are the three steady states and what distinguishes them?
A: The some-AI steady state features finite AI capability, persistent displacement (It = 1), and long-run employment H∞ = δ/µ; it involves technological unemployment. The unbounded-AI steady state features infinite AI capability, no displacement (It = 0), endogenous productivity growth, and employment H∞ = α/(α+σ) as in the standard Mortensen-Pissarides model. The no-AI steady state has A∞ = 0 with the same H∞ = α/(α+σ) but no AI contribution. Employment is higher in the unbounded-AI equilibrium than in the some-AI equilibrium.&lt;/p&gt;
&lt;p&gt;Q: What does the calibration imply for long-run employment and productivity?
A: The calibrated full model converges to a some-AI steady state with a 23% employment loss (H∞ = 0.732), a 321% improvement in AI capability, and a 366% gain in labor productivity. The parameters yield a unique equilibrium under the baseline calibration (πA = 1.949 &amp;gt; sAI = 0.8735 confirms some-AI existence). These results reflect a large worker replacement effect under the calibrated AI learning and error rates, while the job creation effect is relatively modest.&lt;/p&gt;
&lt;p&gt;Q: How fast does technological unemployment materialize?
A: Approximately half of the total 23% employment loss occurs within the first five years; specifically, employment falls by 11.5 percentage points over that period. Over the same five years, AI capability improves by 45.5% and output rises by 49.3%. Over the first 10 years, AI capability improvement accumulates to 94.0% and output gain to 103% (approximately double the five-year output gain).&lt;/p&gt;
&lt;p&gt;Q: How does the full model differ from the benchmark model in transition dynamics?
A: In the full model, job-finding rates are endogenous: firms post vacancies until a free-entry condition (κyt = ftΠt) is satisfied, tying job-finding rate αt to the surplus ratio st via αt = α(st). This endogeneity implies that as AI raises labor productivity, firms create more vacancies, slowing the employment decline relative to the benchmark model with a fixed job-finding rate. At the same time, AI capability grows faster in the full model because higher employment accelerates AI learning.&lt;/p&gt;
&lt;p&gt;Q: What is global indeterminacy and when does it arise?
A: Global indeterminacy occurs when both the some-AI and unbounded-AI steady states coexist, so the long-run outcome depends on initial conditions or expectations. In the benchmark model this requires condition (28): 0 &amp;lt; r + σ + α(1−θ) − (1−b)/πA ≤ εy(µα/(α+σ) − δ). In the full model, global indeterminacy is plausible when firm bargaining power rises to θ = 0.95 given the baseline AI replacement cost ϕ = 0.0043. The region of global indeterminacy is larger when firm bargaining power is higher.&lt;/p&gt;
&lt;p&gt;Q: What is local indeterminacy and what does it imply for transition paths?
A: Local indeterminacy means there is a continuum of equilibrium paths converging to the some-AI steady state in the neighborhood of that steady state, rather than a unique saddle path. In the full model, under alternative parameters (θ = 1, ξ = 0.765, εy = 6), the eigenvalues feature a negative real root and two complex roots with negative real parts, yielding oscillatory local dynamics in employment and AI capability. This implies short-run cycles in productivity and unemployment, consistent with the wide range of empirical findings on AI&amp;rsquo;s labor-market effects.&lt;/p&gt;
&lt;p&gt;Q: Why does the Hosios condition fail to deliver efficiency in this model?
A: The Hosios condition eliminates the standard matching externality by setting firm bargaining power to the vacancy elasticity of matching. But AI introduces four additional externalities: (i) job destruction through displacement, (ii) productivity enhancement for employed workers, (iii) feedback from AI learning that depends on aggregate employment, and (iv) direct effects on matching surpluses and job-finding rates. These externalities mean the standard Hosios rule alone is insufficient; additional instruments are required.&lt;/p&gt;
&lt;p&gt;Q: What is the constrained-optimal policy response?
A: In the simple model, the constrained optimal AI adoption threshold differs from the equilibrium threshold because firm bargaining power θ distorts adoption decisions: AI is over-adopted when πA &amp;gt; (1−b)/(r+σ+α(1−θ)) and under-adopted when (1−b)/(r+σ+α) &amp;lt; πA ≤ (1−b)/(r+σ+α(1−θ)). In the full model, constrained optimality requires setting θ = ξ (Hosios) plus a per-output subsidy T = µApωA to firm-worker matches exposed to AI displacement. This targeted subsidy is 26.6% in the short run and exceeds 50% in the long run.&lt;/p&gt;
&lt;p&gt;Q: How does AI compare to computers in this model&amp;rsquo;s counterfactual?
A: The paper reports that exogenous productivity growth from computers reduced unemployment only modestly—by 0.16 percentage points. By contrast, AI&amp;rsquo;s learning-by-using and displacement features imply a nearly 20% long-run employment loss in a comparable counterfactual. The key distinction is that computers lack the self-learning improvement and associated renegotiation-triggered displacement that characterize GenAI in this model.&lt;/p&gt;
&lt;p&gt;Q: How is AI exposure parameterized and what does it capture?
A: The exposure parameter ρ captures the degree to which a job is subject to AI-driven replacement risk. It is calibrated using Brynjolfsson et al. (2018)&amp;rsquo;s suitability for machine learning (SML) index: on a 1–5 scale, SML averages 3.47 across 964 O*NET occupations, translating to (3.47−1)/(5−1) = 61.8%, so ρ = 0.618. The effective exposure measure is ρµ, which is higher when facing a faster-learning AI.&lt;/p&gt;
&lt;p&gt;Q: What is the predator-prey analogy in the model&amp;rsquo;s dynamics?
A: The dynamical system for AI capability (At) and employment (Ht) in the simple model resembles the Lotka-Volterra predator-prey system. Employment (prey) feeds AI learning; as AI capability (predator) grows, it displaces workers faster, reducing employment; lower employment then slows AI learning, causing capability to decay; and the cycle repeats with diminishing magnitude until the steady state is reached. This mechanism operates only when the AI learning rate µ is neither too high nor too low, with the convergence path being a spiral when µα &amp;lt; 4δ²(1 − δ(α+σ)/(µα)).&lt;/p&gt;
&lt;p&gt;Q: What is the labor-share implication of the unbounded-AI equilibrium?
A: In the unbounded-AI steady state, employment is higher than in the some-AI steady state (H^AJJ &amp;gt; H^AI) and labor productivity grows without bound. However, the labor share is lower in the unbounded-AI equilibrium if the firm&amp;rsquo;s bargaining power θ is sufficiently low. This implies that while workers are not fully displaced and rising AI-augmented productivity sustains employment, workers&amp;rsquo; income share may still decline even in the more favorable unbounded scenario.&lt;/p&gt;
&lt;p&gt;Technological unemployment: A phenomenon in which AI adoption raises labor productivity and expands job creation, yet still causes sizable employment losses because the worker displacement effect (driven by renegotiation failure when AI&amp;rsquo;s NPV πA exceeds the joint surplus) dominates the job-creation effect. In the calibrated model this amounts to a 23% employment loss despite a 366% productivity gain.&lt;/p&gt;
&lt;p&gt;Learning-by-using AI: The model&amp;rsquo;s representation of GenAI as a technology whose capability At grows through reinforced learning from employed workers at rate µ per worker, so aggregate AI growth is µHt, offset by deterioration at rate δ. This distinguishes GenAI from earlier automation technologies (computers, robotics) that do not self-improve through usage.&lt;/p&gt;
&lt;p&gt;Some-AI steady state: A long-run equilibrium with finite AI capability (gA∞ = 0), persistent AI adoption (It = 1), and employment pinned at H∞ = δ/µ—the ratio of AI&amp;rsquo;s error rate to its learning rate. Characterized by expanded job creation but lower employment than the no-AI benchmark, constituting the model&amp;rsquo;s primary calibrated outcome.&lt;/p&gt;
&lt;p&gt;Unbounded-AI steady state: A long-run equilibrium with infinite AI capability (A∞ = ∞), no displacement (It = 0), and endogenous growth at rate gA = µH^AJJ − δ. Employment equals the Mortensen-Pissarides level H∞ = α/(α+σ), and labor productivity grows without bound, complementing Aghion, Jones, and Jones (2019)&amp;rsquo;s idea production framework.&lt;/p&gt;
&lt;p&gt;Global indeterminacy: Coexistence of multiple steady states (some-AI and unbounded-AI) such that the long-run equilibrium depends on initial conditions or expectations rather than being uniquely determined. Arises in the benchmark model when condition (28) holds and becomes more likely with higher firm bargaining power θ.&lt;/p&gt;
&lt;p&gt;Local indeterminacy: A continuum of equilibrium transition paths converging to a single steady state from nearby initial conditions, rather than a unique saddle path. Arises in the full model under certain parameter configurations (e.g., θ = 1, ξ = 0.765, εy = 6), implying oscillatory short-run dynamics in employment and AI capability.&lt;/p&gt;
&lt;p&gt;AI exposure (ρ): A firm-level parameter capturing the degree to which a job-match is subject to AI-driven displacement risk. The displacement option arrives at rate ρµAt per matched pair; ρ is calibrated at 0.618 using the average suitability-for-machine-learning score across O*NET occupations. The effective exposure measure is the product ρµ.&lt;/p&gt;
&lt;p&gt;Renegotiation-proof displacement: Proposition 1&amp;rsquo;s result that the joint surplus Snt is independent of the renegotiation round n, so the AI adoption decision It is also round-invariant. This simplifies the model to a single indicator function: AI replaces the worker if and only if πA exceeds the joint surplus St, regardless of how many renegotiation rounds have occurred.&lt;/p&gt;</description></item><item><title>Barriers to Global Capital Allocation</title><link>https://macropaperwarehouse.com/papers/barriers-to-global-capital-allocation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/barriers-to-global-capital-allocation/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Why do observed international investment positions and cross-country differences in rates of return to capital fail to conform to a frictionless capital-market benchmark? The paper asks how large the efficiency and distributional costs of barriers to global capital allocation are, and which frictions — capital income taxes, political risk, and geographic/cultural/linguistic distances — matter most.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The authors develop a multi-country dynamic spatial general equilibrium model in which the entire network of bilateral cross-border investment positions is endogenously determined. Production in each country i follows a three-factor Cobb-Douglas function in reproducible capital, labor, and natural resources, with country-varying income shares. Capital is the only mobile factor. A logit asset demand system governs portfolio shares: the share of country j&amp;rsquo;s savings invested in country i is proportional to the risk-adjusted expected return on capital in i, scaled by the capital stock of i, and inversely proportional to a bilateral portfolio wedge ∆ij. These wedges can be microfounded via either rational inattention (where wedges reflect the precision of prior beliefs about returns) or extreme-value-distributed transaction costs. The model admits multiple microfoundations but yields the same functional form and the same counterfactual welfare calculations regardless of interpretation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Frictions measured.&lt;/strong&gt; Three categories of frictions enter the empirical implementation: (a) bilateral capital income tax rates — a new dataset covering 225 countries (50,625 country pairs), constructed from corporate income tax rates and treaty-adjusted withholding tax rates on dividends and interest, further adjusted for effective tax rates accounting for tax-haven routing; (b) political risk, proxied by an ICRG composite index (excluding socioeconomic conditions) following Alfaro, Kalemli-Ozcan, and Volosovych (2008); (c) geo-political distance, comprising geographic distance, cultural distance (based on 496 World Values Survey questions across 116 countries), and linguistic distance (based on a language-family tree covering 6,737 languages and 242 countries). These distance measures are publicly available at geopoliticaldistance.org. The model covers 96 countries (9,216 dyads), representing 92% of world GDP in 2017.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gravity Estimation.&lt;/strong&gt; Bilateral investment data (restated for tax havens using the nationality-basis methodology of Coppola et al. 2020 and Damgaard et al. 2019) are regressed on cultural, geographic, and linguistic distance with origin and destination fixed effects. In OLS, a one-standard-deviation increase in cultural distance (0.023 units) is associated with a 24.0% decrease in foreign assets; geographic distance (0.977 units in logs) with a 78.6% decrease; linguistic distance (0.174 units) with a 51.5% decrease. These magnitudes are robust across OLS, PPML, and IV (using religious distance as an instrument for cultural distance). Under IV, the standardized effect of cultural distance on log foreign assets rises to −76.5%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tax haven analysis.&lt;/strong&gt; A Tobit regression of the share of bilateral investment routed through tax havens on the estimated tax saving from routing through havens yields coefficients of 0.413–0.999 for equity and 1.001–1.777 for debt (across specifications with varying fixed effects), confirming that tax incentives are a primary driver of the discrepancy between residency-based and nationality-based bilateral positions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model fit (untargeted moments).&lt;/strong&gt; The calibrated baseline model produces: (i) a correlation of 0.658 between model-implied and empirical rates of return to capital (vs. 0.325 for the frictionless benchmark), with a standard deviation of 0.417 (vs. 0.091 frictionless; data: 0.496); (ii) a correlation of 0.947 between model-implied and empirical capital per employee (vs. 0.918 frictionless); (iii) a correlation of 0.94 between model-implied and empirical home bias; the model reproduces the mean home bias of 3.973 vs. 4.006 in data and standard deviation of 1.065 vs. 1.224, while the frictionless benchmark produces exactly zero home bias for all countries. Portfolio-share MSE: 1.16 (baseline) vs. 1.86 (frictionless).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Counterfactual findings.&lt;/strong&gt; Removing all measured barriers raises world GDP by 6.8% relative to the observed equilibrium (equivalent to stating that the distorted equilibrium is 6.8% below the frictionless benchmark). Geo-political distance alone accounts for most of this: when only distance frictions are retained, world GDP is 5.2% below the frictionless level. Capital taxes alone reduce world GDP by 2.6% below frictionless; political risk alone by 0.4%. The standard deviation of log capital per employee is 51.5% higher than it would be without barriers; the standard deviation of log output per employee is 22.5% higher. In the frictionless equilibrium, capital flows from rich to poor countries (the correlation between net foreign assets and development doubles in absolute value), accounting for the Lucas (1990) puzzle. In short-term (one-period) counterfactuals holding wealth fixed, the GDP gain from full barrier removal is 3.6%; the inequality effect remains similar (standard deviation of log capital per employee 48.4% higher with barriers).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; The model focuses on steady-state outcomes; dynamic transition effects are analyzed in extensions but are smaller. Quantitative conclusions are conditioned on: (i) the model sample of 96 countries covering 92% of world GDP in 2017; (ii) the conservative OLS coefficient estimates used for baseline calibration (IV estimates are larger and would amplify results); (iii) the assumption that the logit demand system captures frictions regardless of their microfoundation; (iv) omission of goods-trade frictions from the baseline (when included, the world GDP effect falls to 3.7% and the capital inequality effect to 23.3%).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-theoretical-prediction-about-cross-country-rates-of-return-when-investment-barriers-exist"&gt;Q1. What is the core theoretical prediction about cross-country rates of return when investment barriers exist?&lt;/h3&gt;
&lt;p&gt;A: In the model&amp;rsquo;s frictionless benchmark (Propositions 1 and 2), all origin countries hold identical portfolios and risk-adjusted expected returns are equalized across destinations. When bilateral frictions are introduced, countries that are more &amp;ldquo;peripheral&amp;rdquo; (harder to access for foreign investors due to high geo-political distance or political risk) receive less inward capital and therefore command higher physical rates of return to capital. Countries that are easily accessible (&amp;ldquo;central&amp;rdquo;) attract more capital and exhibit lower rates of return. The Dual Efficiency Theorem establishes that capital is efficiently allocated if and only if marginal products of capital are equalized across countries, which requires that taxes are uniform and that portfolio wedges satisfy a specific cancellation condition.&lt;/p&gt;
&lt;h3 id="q2-how-are-portfolio-wedges-measured-and-what-is-the-identifying-strategy"&gt;Q2. How are portfolio wedges measured, and what is the identifying strategy?&lt;/h3&gt;
&lt;p&gt;A: Portfolio wedges ∆ij are decomposed into a geo-political distance component and a political risk component. The geo-political distance component is specified as a log-linear function of geographic distance, cultural distance, and linguistic distance, with coefficients (β_g, β_c, β_l) estimated from a gravity regression of log bilateral investment on these distances, controlling for origin and destination fixed effects. Because political risk varies only by destination country, it cannot be separately identified from destination fixed effects in the bilateral regression; its elasticity is therefore taken from Alfaro, Kalemli-Ozcan, and Volosovych (2008). The key identification advantage of bilateral data is that origin and destination fixed effects absorb all country-level confounders, so the distance coefficients are identified purely from within-origin, within-destination variation across country pairs.&lt;/p&gt;
&lt;h3 id="q3-what-do-the-ols-gravity-regressions-find-and-are-the-coefficients-stable-across-specifications"&gt;Q3. What do the OLS gravity regressions find, and are the coefficients stable across specifications?&lt;/h3&gt;
&lt;p&gt;A: In the baseline OLS specification (Table 2, column 1), the estimated coefficients on cultural distance, geographic distance, and linguistic distance are −11.944, −1.579, and −4.162 respectively (all significant at the 1% level). In standardized terms, a one-standard-deviation increase in cultural distance reduces foreign assets by 24.0%, geographic distance by 78.6%, and linguistic distance by 51.5%. Adding a rich set of control variables (colonial ties, legal origin, currency pegs, trade agreements, effective tax rates) leaves these magnitudes broadly similar: standardized effects on foreign assets are −26.4%, −80.1%, and −47.6%, respectively. Results are also robust across OLS and PPML specifications and across years 2013–2017. Effects are quantitatively similar for foreign equity and foreign debt, though linguistic distance has a somewhat smaller effect on debt.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-instrumental-variable-strategy-address-reverse-causality-in-cultural-distance-and-what-does-it-find"&gt;Q4. How does the instrumental variable strategy address reverse causality in cultural distance, and what does it find?&lt;/h3&gt;
&lt;p&gt;A: The authors instrument cultural distance with religious distance (based on historical trees of religious affiliation), assuming religious history affects international investment only through its contemporary effect on differences in values and beliefs as captured by the World Values Survey. The instrument is a strong predictor of cultural distance (passes weak-instrument tests comfortably). Under IV, the standardized effect of a one-standard-deviation increase in cultural distance on log foreign assets rises from −24.0% (OLS) to −76.5% (IV). The authors use conservative OLS estimates for their baseline calibration, so the IV results imply the headline counterfactual effects are likely understated.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-model-predict-home-bias-and-how-well-does-it-match-the-data"&gt;Q5. How does the model predict home bias, and how well does it match the data?&lt;/h3&gt;
&lt;p&gt;A: Home bias is defined as the log difference between the domestic portfolio share and the country&amp;rsquo;s share in the world capital stock. In the frictionless model, Proposition 1 implies that all countries hold identical foreign portfolios, so the model produces exactly zero home bias for every country. The baseline model, by incorporating bilateral frictions, generates home bias endogenously without targeting it. The model-implied home bias correlates with the empirically measured home bias at 0.94 across countries and matches both the mean (3.973 model vs. 4.006 data) and standard deviation (1.065 vs. 1.224) closely. The model also predicts, consistent with Lau, Ng, and Zhang (2010), that home bias and rates of return on capital are positively correlated (model-implied ρ = 0.55), and that rates of return on capital correlate negatively with the log of GDP per employee (model-implied ρ = −0.70).&lt;/p&gt;
&lt;h3 id="q6-what-is-the-quantitative-decomposition-of-the-world-gdp-loss-by-type-of-barrier"&gt;Q6. What is the quantitative decomposition of the world GDP loss by type of barrier?&lt;/h3&gt;
&lt;p&gt;A: World GDP in the observed (distorted) equilibrium is measured at $112.9 trillion (PPP), which is 6.8% below the frictionless counterfactual. When all barriers are present except geo-political distance, world GDP is 5.2% below frictionless — meaning distance frictions account for the largest share. When all barriers are present except political risk, world GDP is only 0.4% below frictionless. When all barriers are present except taxes, world GDP is 2.6% below frictionless. These are not exactly additive because the distortions interact; the results confirm that geo-political distance (cultural, linguistic, and geographic) constitutes the dominant source of global capital misallocation among the three measured frictions.&lt;/p&gt;
&lt;h3 id="q7-how-do-barriers-affect-the-cross-country-distribution-of-capital-and-income"&gt;Q7. How do barriers affect the cross-country distribution of capital and income?&lt;/h3&gt;
&lt;p&gt;A: The standard deviation of log capital per employee is 51.5% higher in the distorted equilibrium than in the frictionless counterfactual; the standard deviation of log output per employee is 22.5% higher. When only geo-political distance distortions are maintained, dispersion in log capital per employee is 38.2% higher and in log output per employee 15.9% higher. Maintaining only taxes raises the dispersion in log capital per employee by 12.9% and log output per employee by 6.0%; maintaining only political risk raises them by 7.3% and 3.8%, respectively. In the frictionless equilibrium, the poorest countries gain the most: some of the poorest countries see capital per employee increase by an order of magnitude and income per employee double.&lt;/p&gt;
&lt;h3 id="q8-does-the-model-account-for-the-lucas-puzzle-capital-not-flowing-from-rich-to-poor-countries"&gt;Q8. Does the model account for the Lucas puzzle (capital not flowing from rich to poor countries)?&lt;/h3&gt;
&lt;p&gt;A: Yes. In the observed distorted equilibrium, net foreign asset positions correlate only weakly with the level of development, consistent with Lucas&amp;rsquo;s (1990) observation that capital fails to flow from rich to poor countries. In the frictionless counterfactual, the absolute value of the correlation between net foreign asset positions and log GDP per employee doubles, and capital indeed flows from rich to poor countries as neoclassical theory predicts. The distortions from taxes, political risk, and geo-political distance thus account for the absence of a strong correlation between net positions and development in the data.&lt;/p&gt;
&lt;h3 id="q9-how-do-extensions-incorporating-goods-trade-frictions-capital-controls-and-currency-hedging-costs-affect-the-headline-findings"&gt;Q9. How do extensions incorporating goods-trade frictions, capital controls, and currency hedging costs affect the headline findings?&lt;/h3&gt;
&lt;p&gt;A: Adding goods-trade frictions (country-specific prices for output and capital installation following Monge-Naranjo et al. 2019) reduces the world GDP effect to 3.7% (from 6.8% baseline) and the dispersion of log capital per employee to 23.3% higher (from 51.5%), but the overall pattern of results is preserved. Replacing political risk with capital controls (using Jahan and Wang 2016 de-jure capital account openness) yields a comparable world GDP loss of 6.6% and a geo-political distance effect of 6.2%, very close to the 6.8% and 5.2% in the baseline. Adding currency hedging costs leaves world GDP loss and inequality effects essentially unchanged relative to baseline. None of these extensions materially alters the headline conclusions.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-authors-validate-the-model-against-nationality-based-versus-residency-based-bilateral-investment-data"&gt;Q10. How do the authors validate the model against nationality-based versus residency-based bilateral investment data?&lt;/h3&gt;
&lt;p&gt;A: The model is calibrated to nationality-based positions (restated for tax havens). The MSE for fitting nationality-based external portfolio shares is 1.16, while the MSE for residency-based positions is 1.22. The model was not explicitly designed to distinguish between the two, yet it naturally produces better predictions for nationality-based positions because its frictions incorporate the incentives for indirect investment routing through tax havens. This cross-validation supports the methodological approach of using nationality-restated data and confirms the internal consistency of the model&amp;rsquo;s treatment of tax-haven routing.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-implications-for-global-tax-policy-coordination"&gt;Q11. What are the implications for global tax policy coordination?&lt;/h3&gt;
&lt;p&gt;A: In the presence of information frictions, simple harmonization of capital tax rates across countries does not improve capital allocation efficiency and could worsen it. The Dual Efficiency Theorem implies that efficient capital allocation in a world with information frictions requires that taxes, risk premia, and information frictions satisfy a joint cancellation condition. From a normative perspective, a global social planner maximizing world GDP should impose lower capital tax rates in countries that are &amp;ldquo;peripheral&amp;rdquo; in the network of informational distances, in order to offset the disadvantage created by information frictions for those countries.&lt;/p&gt;
&lt;h3 id="q12-how-is-the-elasticity-parameter-η-calibrated-and-how-sensitive-are-the-results"&gt;Q12. How is the elasticity parameter η calibrated, and how sensitive are the results?&lt;/h3&gt;
&lt;p&gt;A: The elasticity of substitution among countries&amp;rsquo; assets, η, is calibrated at 18.5 based on Koijen and Yogo (2020)&amp;rsquo;s demand-price elasticities for long-term debt (3.1, converted to a gross-return elasticity of approximately 30), short-term debt (25.2, converted to approximately 24.3), and equity (1.3, converted to approximately 14.8), with weights reflecting the composition of global portfolios. The baseline gravity coefficients are calibrated from OLS with controls (cultural: −13.129, geographic: −1.645, linguistic: −3.850), chosen as conservative estimates relative to IV or PPML. Sensitivity analysis using PPML or IV estimates of β yields broadly similar steady-state GDP losses (around 6%), confirming robustness.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Portfolio wedge (∆ij):&lt;/strong&gt; A bilateral distortionary term in the logit asset demand system that captures all frictions reducing the ability of investors from country j to invest in country i. Decomposed empirically into a geo-political distance component and a political risk component. A wedge of 1 means no friction; larger values reduce the share of investment flowing from j to i. Can be interpreted either as prior-belief imprecision under rational inattention or as systematic transaction costs under the extreme-value microfoundation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Geo-political distance:&lt;/strong&gt; A composite of geographic distance (population-weighted geodesic distance), cultural distance (expected disagreement in World Values Survey responses between randomly drawn individuals from two countries, constructed with the &amp;ldquo;flex&amp;rdquo; method using up to 496 questions), and linguistic distance (normalized tree distance in the Ethnologue language family graph, covering 6,737 languages). Distinct from simple physical distance: it captures the informational and transactional barriers that arise from societal dissimilarity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dual Efficiency Theorem:&lt;/strong&gt; A theoretical result (Theorem in Section 2.8) establishing that capital efficient allocation, equalization of marginal products of capital across countries, and uniform taxes combined with a specific cancellation condition on portfolio wedges are mutually equivalent statements in steady-state equilibrium. This is not a restatement of the First Welfare Theorem; it is a statement about GDP (not welfare) and does not require risk premia to be equalized.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Effective bilateral tax rate (τij):&lt;/strong&gt; The composite bilateral tax rate on capital after accounting for tax-haven routing. Firms in the destination country optimally choose the share of capital issued through tax havens (solving a quadratic cost optimization), trading off the lower tax rate available through havens against an increasing quadratic routing cost. The effective rate is therefore lower than the statutory (de jure) rate when the tax-haven rate is lower than the statutory rate, with the gap depending on the estimated βth coefficient from the Tobit regressions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Logit asset demand system:&lt;/strong&gt; A portfolio allocation rule in which the share of country j&amp;rsquo;s savings invested in destination country i is proportional to the risk-adjusted expected return raised to the power η (the elasticity of substitution) times the destination capital stock, divided by the portfolio wedge and summed over all destinations. Microfounded either by rational inattention (Matejka and McKay 2015; Pellegrino 2023) or by extreme-value-distributed transaction costs. Produces portfolio gravity analogous to trade gravity when combined with the market clearing conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Home bias:&lt;/strong&gt; Defined as the log difference between a country&amp;rsquo;s domestic portfolio share (πii, the share of domestic savings invested at home) and that country&amp;rsquo;s share of world capital stock (ki/K). In the frictionless benchmark, home bias is exactly zero for all countries by Proposition 1. The baseline model generates home bias endogenously as a consequence of portfolio wedges and reproduces both the level and cross-sectional distribution of empirically observed home bias without targeting these moments directly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Core-periphery structure:&lt;/strong&gt; An emergent property of international capital markets under investment barriers: countries that are easily accessible to international investors (low geo-political distance, low political risk, favorable tax treatment) are &amp;ldquo;central&amp;rdquo; and attract capital inflows, driving their rates of return to capital lower; &amp;ldquo;peripheral&amp;rdquo; countries that are less accessible have smaller capital stocks and higher rates of return, compensating investors for overcoming barriers. This structure generates persistent capital misallocation and cross-country income inequality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nationality-based vs. residency-based bilateral investment positions:&lt;/strong&gt; Residency-based data (e.g., raw IMF CPIS) attributes investment to the immediate counterparty country, including tax-haven shell companies. Nationality-based data (Coppola et al. 2020; Damgaard et al. 2019; Beck et al. 2024) reattributes investment to the country of the ultimate investor and ultimate issuer, bypassing offshore centers. The model fits nationality-based positions better (MSE 1.16 vs. 1.22 for residency-based) because it incorporates frictions that generate incentives for indirect routing, which is what nationality restatement is designed to undo.&lt;/p&gt;</description></item><item><title>Bridges</title><link>https://macropaperwarehouse.com/papers/bridges/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/bridges/</guid><description>&lt;p&gt;This paper measures the causal effects of land transport infrastructure on economic activity, exploiting quasi-experimental variation in bridge construction over the Mississippi and Ohio Rivers in the United States. The central empirical puzzle motivating the study is a hump-shaped relationship between per capita income and distance to major land transport routes in contemporary U.S. data: income peaks around 5 km from a transport route, with an elasticity of 0.072 closer than 4.1 km and -0.096 at greater distances, so that 85% of Americans live where local income increases with distance to transport routes rather than decreasing. The question is whether this pattern reflects causal effects of infrastructure, selection, or sorting.&lt;/p&gt;
&lt;p&gt;The paper develops two complementary identification strategies. The first exploits tributary confluences — where smaller rivers join larger rivers, sharply raising downstream flow rates and bridge construction costs — to generate quasi-random variation in bridge location. Because bridge construction costs increase convexly with river flow (maximum bending moment scales with span length squared), bridges are disproportionately built just upstream of confluences. The median upstream census tract lies 0.7 km from a bridge versus 2.3 km for the median downstream tract, making upstream tracts on average 60% closer to bridges and 27% closer to the nearest major land transport route. This asymmetry dates to at least 1880 and persists to 2010. Despite this persistent connectivity advantage, by 2010 upstream tracts have 13% lower per capita incomes and 63% higher population densities than downstream neighbours. The implied elasticity of per capita income with respect to distance to land transport, scaling the income effect by the distance-to-transport effect, is approximately 0.44. Income density (income per unit area) is higher upstream, though the difference is not statistically significant. Historical placebo tests using pre-bridge-construction data show no asymmetry in land values or population upstream versus downstream, supporting the identification assumption.&lt;/p&gt;
&lt;p&gt;The second strategy exploits variation in the timing of bridge construction. Because major bridge projects involve decades of planning, financing, design, and construction — the Wheeling Suspension Bridge was chartered in 1816 but opened in 1849 — the precise opening date is argued to be exogenous to short-run deviations from local growth trends. Using a county-level panel from 1860 to 2010 (432 counties, 14–19 states), the paper estimates event-study regressions around the first time a county experiences a 50% reduction in distance to a bridge. After such a reduction, farm land values (the best available consistent proxy for total economic activity in historical data) rise immediately and cumulatively by approximately 9% over 30 years. Population rises by approximately 5% over the same period. The proportionally larger rise in land values than population implies higher per capita economic activity in better-connected counties after 30 years.&lt;/p&gt;
&lt;p&gt;These two sets of results are reconciled through a narrative account of development. Better bridge access drives industrialization — manufacturing employment shares rise in counties experiencing improved connectivity — and urbanization. Cities form around historical transport routes and expand. Richer households then sort away from historical city centres into lower-density suburban areas, while lower-income households remain near or selectively migrate to the historical transport corridors. This within-city sorting produces the observed cross-sectional gradient: areas nearest transport routes end up with higher population density but lower per capita incomes. The negative local income effect of proximity to transport routes is larger in more urbanized areas and areas with higher income inequality, and is concentrated among non-white and low-education populations.&lt;/p&gt;
&lt;p&gt;The paper also contributes a new dataset covering every road and rail bridge (237 total) ever constructed over the Mississippi and Ohio Rivers from 1849 to 2010, assembled from the National Bridge Inventory and extensively cross-checked with satellite imagery and historical sources.&lt;/p&gt;
&lt;p&gt;Q: What is the motivating empirical puzzle about transport infrastructure and income?&lt;/p&gt;
&lt;p&gt;A: In contemporary U.S. census data, per capita income does not monotonically increase with proximity to land transport routes. Instead, the relationship is hump-shaped: income peaks around 5 km from a major transport route, with a positive elasticity of 0.072 within 4.1 km and a negative elasticity of -0.096 beyond that distance. Population density, by contrast, falls monotonically with distance to transport routes. As a result, 85% of Americans live in places where local mean income increases with distance to transport infrastructure rather than decreasing.&lt;/p&gt;
&lt;p&gt;Q: How does the tributary confluence identification strategy work?&lt;/p&gt;
&lt;p&gt;A: Tributary confluences — where smaller rivers join the main river — cause sharp, localized increases in river flow rates and thus in bridge construction costs, because cost scales convexly with required span length. This makes bridges systematically more likely to be built just upstream of confluences than just downstream. The strategy compares census tracts located upstream versus downstream of the 27 major tributary confluences identified on the Mississippi and Ohio Rivers, controlling for nearest-tributary fixed effects and distance to the confluence.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the connectivity difference between upstream and downstream census tracts?&lt;/p&gt;
&lt;p&gt;A: Upstream census tracts are approximately 60% closer to a bridge than downstream tracts (coefficient of 0.91 in log distance to bridge, p &amp;lt; 0.01), and consequently 27% closer to the nearest major land transport route (coefficient of 0.32, p &amp;lt; 0.10). This asymmetry is established by 1880 and persists through 2010. The advantage arises approximately equally from proximity to railroads and primary roads.&lt;/p&gt;
&lt;p&gt;Q: What are the causal effects of this connectivity advantage on per capita income and population density?&lt;/p&gt;
&lt;p&gt;A: Despite being better connected, upstream census tracts have 13% lower per capita incomes (coefficient 0.14 on the downstream indicator in log per capita income, p &amp;lt; 0.05) and 63% higher population densities (coefficient -0.49 on the downstream indicator in log population density, p &amp;lt; 0.05) in 2010. Income density is higher upstream, but the difference is not statistically distinguishable from zero. Scaling the income effect by the effect on distance to land transport implies an elasticity of approximately 0.44.&lt;/p&gt;
&lt;p&gt;Q: What pre-bridge-era placebo tests support the identifying assumption for the tributary confluence strategy?&lt;/p&gt;
&lt;p&gt;A: Matching modern census tracts to county-level historical data from 1840 and 1850 (before substantive bridge construction began), the paper finds no statistically significant asymmetry in land values or population density upstream versus downstream of tributary confluences. Asymmetric patterns emerge only after bridge construction begins. Ferry crossing locations, traced through place names in the USGS Geographic Names database, also appear equally frequently upstream and downstream, suggesting ferries did not differentially locate upstream.&lt;/p&gt;
&lt;p&gt;Q: How does the timing-based identification strategy work, and what is its key assumption?&lt;/p&gt;
&lt;p&gt;A: The strategy uses a county-level panel from 1860 to 2010 and estimates event-study regressions around the first time a county experiences a 50% reduction in distance to a bridge. County fixed effects and county-specific quadratic time trends absorb all fixed differences across counties and average changes in trends. The key assumption is that the exact opening date of a bridge is exogenous to short-run deviations from local long-run growth trends — supported by the argument that major bridges involve decades-long planning processes that evolve independently of local economic fluctuations. Pre-trend tests show no significant differences in outcomes before the event.&lt;/p&gt;
&lt;p&gt;Q: What are the quantitative effects of a major improvement in bridge access on land values and population?&lt;/p&gt;
&lt;p&gt;A: After a county first experiences a 50% reduction in distance to a bridge, farm land values rise immediately and cumulatively by approximately 9% (cumulative effect on log land values of about 0.09) over 30 years, relative to counties with no such change. Population rises by approximately 5% (cumulative log effect of about 0.05) over the same period. The proportionally larger effect on land values than on population implies that per capita economic activity is higher in better-connected counties 30 years after the event. The divergence between land value and population effects grows over time, suggesting productivity advantages accumulate.&lt;/p&gt;
&lt;p&gt;Q: Why does the paper use farm land values rather than other income measures in the historical panel?&lt;/p&gt;
&lt;p&gt;A: Farm land values — the total value of farm land and buildings — are the best consistently measured proxy for total economic activity available throughout the 1860–2010 census panel. The paper notes explicitly that as the economy industrializes and urbanizes, farm land values increasingly miss urban land values, implying that the estimated effects on farm land values are likely lower bounds on the true effects on total economic activity.&lt;/p&gt;
&lt;p&gt;Q: How does the paper address the concern that bridge timing might reflect anticipated local growth?&lt;/p&gt;
&lt;p&gt;A: The paper shows that results hold when restricting to counties whose distance to a bridge is only affected by bridges constructed in other counties, addressing the concern that local planners might time construction in anticipation of local growth. The results are also insensitive to controlling for pre-period trends, and outcomes of interest are uncorrelated with future changes in distance to a bridge in preferred specifications.&lt;/p&gt;
&lt;p&gt;Q: How does the paper reconcile the negative local income effect (tributary confluence strategy) with the positive aggregate effect (timing strategy)?&lt;/p&gt;
&lt;p&gt;A: The reconciliation proceeds through a narrative account combining industrialization, urbanization, and within-city sorting. Better bridge access drives a shift toward manufacturing employment and attracts population, consistent with a productivity advantage enabling exploitation of economies of scale. Cities form around historical transport routes. As cities mature and expand, richer households sort into lower-density suburban areas further from the historical transport corridor, while lower-income households remain near or migrate to the city centre. This within-city sorting produces lower per capita incomes near transport routes even as aggregate economic activity is higher in better-connected areas.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports the within-city sorting mechanism specifically?&lt;/p&gt;
&lt;p&gt;A: The negative income effect of proximity to transport routes is larger in more urbanized areas and in areas with higher income inequality. The effect is concentrated in areas that were more rapidly urbanizing in the 19th century, and it is stronger for non-white and low-education populations. Upstream census tracts simultaneously show higher manufacturing employment shares and higher population densities, consistent with cities having formed around transport routes, followed by residential sorting away from the core.&lt;/p&gt;
&lt;p&gt;Q: What are the two novel identification strategies and their broader applicability?&lt;/p&gt;
&lt;p&gt;A: The tributary confluence strategy exploits discontinuities in bridge construction costs generated by sharp increases in river flow rates at confluences; it requires only that bridges are more likely built upstream of confluences than downstream, an asymmetry the paper shows is detectable elsewhere in the world from satellite imagery. The timing strategy exploits the multi-decade planning and construction process for major bridges as a source of near-exogenous variation in opening dates. Both strategies can be applied in other settings where major rivers form substantial barriers to land transport networks.&lt;/p&gt;
&lt;p&gt;Q: What does the paper contribute to the debate about whether early U.S. transport infrastructure followed or led economic development?&lt;/p&gt;
&lt;p&gt;A: The results support the view that early investments in land transport infrastructure led to meaningful changes in economic geography rather than merely following pre-existing growth patterns. However, the paper finds a moderate level of responsiveness — population density responds to bridge access over several decades, not immediately — consistent with a broader literature documenting sluggish population responses to changes in economic conditions.&lt;/p&gt;
&lt;p&gt;Tributary confluence: A location where a smaller river (tributary) joins a larger river, causing a sharp, localized increase in downstream flow rates and therefore a discontinuous increase in bridge construction costs, generating the quasi-experimental variation in bridge location exploited in the paper.&lt;/p&gt;
&lt;p&gt;Within-city sorting: The process by which, as cities expand around historical transport routes, richer households differentially relocate to lower-density suburban areas further from the transport corridor while lower-income households remain near or migrate to the historical city centre, reversing the income gradient at small spatial scales.&lt;/p&gt;
&lt;p&gt;Income density: The product of population density and per capita income, corresponding to total economic activity per unit area; the paper finds income density is higher in better-connected upstream census tracts even when per capita income is lower, reflecting the dominant effect of higher population density.&lt;/p&gt;
&lt;p&gt;Farm land values: The total value of farm land and buildings, used as the best consistently available proxy for total economic activity in the 1860–2010 historical county panel; the paper treats estimated effects on farm land values as lower bounds on effects on total economic activity because farm values increasingly miss urban land as the economy industrializes.&lt;/p&gt;
&lt;p&gt;Structural transformation: The shift in the composition of employment away from agriculture and toward manufacturing, which the paper documents occurring in counties that experience improved bridge access, interpreted as evidence that transport infrastructure provides a productivity advantage attracting industrial activity.&lt;/p&gt;
&lt;p&gt;Distance to a bridge (as proxy for land transport access): In the study area along the Mississippi and Ohio Rivers, where all land has comparable water access, distance to the nearest bridge strongly predicts distance to the nearest major land transport route (rail or primary road), allowing bridge distance to serve as a consistent measure of transport connectivity throughout the entire study period.&lt;/p&gt;
&lt;p&gt;Market access: A measure of economic connectivity that captures both the state of the transport network and the size of accessible markets; the paper notes that log distance to a bridge explains 46% of the variation in market access in 1890 (from Donaldson and Hornbeck&amp;rsquo;s data) with an elasticity of approximately 0.1, and that halving distance to a bridge increases market access by approximately 7%.&lt;/p&gt;</description></item><item><title>Catastrophes, Delays, and Learning</title><link>https://macropaperwarehouse.com/papers/catastrophes-delays-and-learning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/catastrophes-delays-and-learning/</guid><description>&lt;p&gt;This paper develops a general model of experimentation under catastrophe risk in which the catastrophe is triggered when a stock variable exceeds an unknown threshold, but occurs only after a stochastic delay. The central contribution is the concept of the &amp;ldquo;legacy of the past&amp;rdquo;: at any planning date, past experiments may have already triggered a catastrophe that has not yet materialized, and the planner cannot observe whether triggering has occurred. The legacy is formally defined as the probability, conditional on survival, that a catastrophe was triggered in the past.&lt;/p&gt;
&lt;p&gt;The model unifies two canonical but previously incompatible approaches in the literature. In the hazard-rate approach, the catastrophe is bound to happen and the planner manages its timing and severity. In the unknown-threshold approach, learning is instantaneous and the catastrophe is certainly avoided if the stock has not yet exceeded the threshold. Neither approach captures the intermediate case where the planner remains uncertain about whether the catastrophe is already underway. By introducing a delay governed by an exponential distribution with parameter α, the authors show that both approaches are limiting special cases: as α → ∞ (no delay), the legacy vanishes and the unknown-threshold approach is recovered; when the legacy is set permanently to one (catastrophe triggered with certainty), the hazard-rate approach is recovered.&lt;/p&gt;
&lt;p&gt;Three benchmark stock levels anchor the analysis. QN is the long-run target absent any catastrophe risk. QD (&amp;ldquo;Damages&amp;rdquo;) is the optimal stabilization target when the planner knows a catastrophe was triggered in the past — it lies weakly below QN because the planner trades off current gains against the discounted marginal damage from raising the stock at the moment of eventual catastrophe occurrence. QE (&amp;ldquo;Experimentation&amp;rdquo;) is the stock level below which stabilization is suboptimal when the planner is certain no triggering has occurred — it also lies weakly below QN.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s two main theorems are distinguished by the ranking of QD and QE, which reflects whether mitigation strategies are effective.&lt;/p&gt;
&lt;p&gt;Theorem 1 (QE &amp;lt; QD): When damage is not highly sensitive to the stock level at catastrophe time — so mitigation is relatively ineffective — optimal paths are monotonically increasing and converge to a long-run stock level Q∞ ∈ [QE, QD]. The stopping condition equates the marginal benefit of experimentation to a weighted average of the expected cost under the unknown-threshold approach (weight 1 − π) and the marginal damage under the hazard-rate approach (weight π), where π is the legacy at stopping time. A higher legacy at the stopping time is associated with a higher long-run stock level. A higher initial legacy induces fatalism: since the catastrophe is more likely already triggered, the planner shifts priority toward current consumption rather than caution, leading to more total experimentation.&lt;/p&gt;
&lt;p&gt;Theorem 2 (QD &amp;lt; QE): When damage is highly sensitive to the stock level — so mitigation is valuable — the long-run target is uniquely QE regardless of the initial legacy. However, the short-run path is non-monotonic: for a sufficiently high initial legacy, the planner first reduces the stock sharply (lockdown, emissions cut) to mitigate pending catastrophe damages, then, as the legacy declines because no catastrophe occurs, gradually allows the stock to rise back toward QE. The direction of caution reverses relative to Theorem 1: a higher legacy now induces more caution, not less.&lt;/p&gt;
&lt;p&gt;Applications include pandemic management (stock = infected population, catastrophe = health system collapse) and climate change (stock = cumulative CO2 emissions or atmospheric pollution stock). In the disease control application, whether a planner prioritizes economic production or mortality reduction determines which theorem governs, with the key ratio being production losses relative to mortality increases. For pandemic policy, Theorem 2 produces a formal learning-based rationale for non-monotonic &amp;ldquo;hammer-and-dance&amp;rdquo; policies (strict early lockdown followed by relaxation) that differs from prior explanations in the literature. In the carbon budget application, Proposition 5 formally proves that higher initial legacy raises the optimal carbon budget under Theorem 1 conditions, and can imply unbounded consumption (certainty of catastrophe) above a critical legacy threshold π*. Under Theorem 2 conditions (Proposition 6), the optimal policy can involve first reducing then expanding the stock before stabilizing, with both transition dates increasing in the initial legacy.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;legacy of the past&amp;rdquo; and how is it computed?
A: The legacy πt is defined as the probability, conditional on survival to date t, that a catastrophe was already triggered by past experiments. Formally, πt = 1 − [1 − F(Qt)] / pt, where Qt is the highest stock level ever reached, F is the prior distribution over the threshold, and pt is the survival probability. A past experiment at time t&amp;rsquo; contributes to the current legacy with weight exp[−α(t − t&amp;rsquo;)], so recent experiments matter more than distant ones. As time passes without catastrophe, the legacy of any fixed past experiment declines geometrically at rate α.&lt;/p&gt;
&lt;p&gt;Q: How do the three benchmark stock levels QN, QD, and QE relate to each other?
A: QN is the optimal long-run stock without any catastrophe. QD is defined by the condition where the marginal net benefit of increasing the stock — ν(Q) − [α/(α+δ)]D&amp;rsquo;(Q) — equals zero, and satisfies QD ≤ QN. QE is defined by ν(Q) − [α/(α+δ)]ρ(Q)D(Q) = zero, and also satisfies QE ≤ QN. The ranking between QD and QE depends on whether damage is more sensitive to the marginal increase in stock at catastrophe time (which pushes QD below QE) or to the level of the stock at triggering (which pulls QD above QE).&lt;/p&gt;
&lt;p&gt;Q: What is the key optimality condition in Theorem 1 and how does it unify prior approaches?
A: The stopping condition (equation 15) states: ν(QT) = [α/(α+δ)] × [(1 − πT)ρ(QT)D(QT) + πT D&amp;rsquo;(QT)]. When πT = 0 (no legacy, unknown-threshold limit), this reduces to the experimentation stopping condition of Tsur and Zemel, governed by the hazard rate ρ(QT) times expected loss D(QT). When πT = 1 (full legacy, hazard-rate limit), it reduces to the damage-mitigation condition governed by marginal damage D&amp;rsquo;(QT). The legacy at stopping time thus serves as the mixing weight between the two canonical approaches, embedding both as special cases.&lt;/p&gt;
&lt;p&gt;Q: How does the initial legacy affect total experimentation under Theorem 1 versus Theorem 2?
A: Under Theorem 1 (QE &amp;lt; QD), a higher initial legacy π0 leads to more total experimentation (higher Q∞), because the planner becomes fatalistic — since the catastrophe is more likely already triggered and mitigation is relatively ineffective, current consumption is prioritized. Proposition 5 formally proves this for the carbon budget application: the optimal stopping date T and optimal budget QT are nondecreasing in π0. Under Theorem 2 (QD &amp;lt; QE), a higher legacy triggers more caution in the short run (larger reduction in the stock during the mitigation phase), but the long-run target QE remains the same regardless of π0.&lt;/p&gt;
&lt;p&gt;Q: What generates non-monotonic policies in Theorem 2, and what does this look like in the pandemic application?
A: Non-monotonicity arises because the optimal response to a high legacy is first to reduce the stock sharply to limit catastrophe damages (since damage is sensitive to the stock level), and then, as time passes without catastrophe and the legacy declines, to allow the stock to recover. In the disease control application with high mortality weight, a complete lockdown is optimal in the first phase whenever the legacy is strictly positive. As the legacy declines, the lockdown is gradually relaxed, and eventually the infection level returns to its pre-lockdown level. Figures 3 and 4 show that a higher initial legacy (π0 = 0.1, 0.5, or 0.9) leads to a longer lockdown and slower recovery, though all paths converge to the same long-run infection level.&lt;/p&gt;
&lt;p&gt;Q: How does the model&amp;rsquo;s disease control application determine which theorem governs?
A: Lemma 2 states that if 1 / [1 + (Y(r+d) − Y*) / (wµ&lt;em&gt;dI^D)] &amp;lt; ρ(I^D), then I^E &amp;lt; I^D and Theorem 1 applies; otherwise I^E &amp;gt; I^D and Theorem 2 applies. The key ratio is (Y(r+d) − Y&lt;/em&gt;) / (wµ*d), the production loss relative to mortality increase. A planner who weights economic activity heavily (large production loss ratio) falls under Theorem 1 and tolerates rising infections; a planner who weights mortality heavily falls under Theorem 2 and imposes an initial lockdown.&lt;/p&gt;
&lt;p&gt;Q: What is the carbon budget result under Theorem 1 (Proposition 5)?
A: Under the condition u1 &amp;gt; [α/(α+δ)]v0 (marginal consumption value exceeds discounted marginal damage), Theorem 1 applies and there exists a critical legacy threshold π* such that: below π*, the planner consumes maximally (qt = q-bar) until a finite date T and then stops, with QE &amp;lt; QT &amp;lt; QD; above π*, the planner consumes maximally forever, triggering the catastrophe with certainty. The stopping date T and the optimal budget QT are nondecreasing functions of initial legacy π0, formally proving that higher past emissions (captured through legacy) justify higher future carbon budgets in this model.&lt;/p&gt;
&lt;p&gt;Q: What is the carbon budget result under Theorem 2 (Proposition 6)?
A: Under condition u1 &amp;lt; [α/(α+δ)]v0, QD &amp;lt; QE and Theorem 2 applies. Starting from Q0 above QE, if π0 is small enough (specifically u1 &amp;gt; π0[α/(α+δ)]v0), the optimal policy is to stabilize the stock forever at Q0. Otherwise, there exist two finite dates t1 &amp;lt; t2, both increasing in π0, such that the planner first reduces the stock at maximum rate (qt = q-bar-negative) for t &amp;lt; t1, then expands at maximum rate for t1 &amp;lt; t &amp;lt; t2, then stabilizes at Q0 forever. The optimal carbon budget is Q0 in all cases, showing that the long-run target is independent of legacy under Theorem 2.&lt;/p&gt;
&lt;p&gt;Q: How does the model relate to the hazard-rate literature formally?
A: Papers such as Nordhaus and others that use an exogenous hazard rate h(Qt) for catastrophe — yielding survival probability pt = p0 exp(−∫h(Qτ)dτ) — are shown to be equivalent to the special case where the catastrophe was triggered in the past (legacy = 1 permanently). Their formulation corresponds to assuming α is constant and the legacy is identically one, which reduces the law of motion for pt to pt = p0 exp(−αt). The key difference is that in the hazard-rate approach the planner can reduce the arrival rate by lowering the stock (h is increasing in Q), whereas in the authors&amp;rsquo; model the delay parameter α is constant and policy affects only damages.&lt;/p&gt;
&lt;p&gt;Q: What is the role of the exponential delay distribution assumption?
A: The assumption that the delay τ follows an exponential distribution with parameter α is made for tractability. Under this assumption, the entire past trajectory of the stock (Qt)t≤0 can be summarized by just two state variables — the highest stock on record Q0-bar and the initial legacy π0 — because the exponential &amp;ldquo;memoryless&amp;rdquo; property means that the additional expected waiting time until catastrophe occurrence does not depend on how long the triggering has already been in effect. Without this assumption, the full chronicle of past experiments would be required as a state variable, making the problem intractable.&lt;/p&gt;
&lt;p&gt;Q: What happens when the delay parameter α approaches zero or infinity?
A: When α → ∞ (instantaneous catastrophe upon triggering), pt = 1 − F(Qt) and the legacy is identically zero, recovering the Tsur-Zemel unknown-threshold approach (Proposition 3). The optimal path converges to QE0 from below or stabilizes if already above QE0. When α → 0 (infinite delay, effectively no catastrophe), QE = QD = QN and the problem reduces to the simple stock-flow problem (Proposition 1), with the optimal path converging monotonically to QN.&lt;/p&gt;
&lt;p&gt;Q: Does the model allow for damage mitigation after triggering but before occurrence?
A: Yes, this is a key feature. The continuation payoff after catastrophe occurrence is V(QT) where QT is the stock level at the time of occurrence T, not at triggering time T(S). This means the planner can reduce the stock after triggering to lower damages — analogous to a skater turning back toward shore after the ice first cracks. The assumption that V depends on the stock at occurrence rather than at triggering or at the maximum historical level is what allows this mitigation channel and is explicitly noted as a modeling choice.&lt;/p&gt;
&lt;p&gt;Legacy of the past (πt): The probability, conditional on survival to date t, that past experiments have already triggered a catastrophe. Formally πt = 1 − [1 − F(Qt)] / pt. Recent experiments contribute more to the legacy than distant ones, with contribution decaying at rate α. The legacy is zero when α → ∞ and is the central state variable bridging the paper&amp;rsquo;s two canonical extremes.&lt;/p&gt;
&lt;p&gt;QE (&amp;ldquo;Experimentation&amp;rdquo; threshold): The stock level at which the net marginal gain from further experimentation, defined as ν(Q) − [α/(α+δ)]ρ(Q)D(Q), equals zero, under the assumption that no catastrophe has been triggered. Below QE, stabilization is suboptimal; above QE, the planner does not experiment further when the legacy is zero.&lt;/p&gt;
&lt;p&gt;QD (&amp;ldquo;Damages&amp;rdquo; threshold): The stock level at which the net marginal benefit from holding the stock, defined as ν(Q) − [α/(α+δ)]D&amp;rsquo;(Q), equals zero, under the assumption that the catastrophe is known to have been triggered. QD ≤ QN and represents the optimal long-run target when the hazard-rate approach applies.&lt;/p&gt;
&lt;p&gt;Marginal payoff ν(Q): Defined as uq(0, Q) + (1/δ)uQ(0, Q), it measures the net gain from marginally increasing the flow when the stock is stabilized at Q. It is strictly decreasing in Q under Assumption 1 and equals zero at QN.&lt;/p&gt;
&lt;p&gt;Damage function D(Q): Defined as (1/δ)u(0, Q) − V(Q), it measures the welfare loss from catastrophe occurrence when the stock is Q at occurrence time, relative to permanent stabilization at Q. Assumed weakly positive and weakly increasing in Q.&lt;/p&gt;
&lt;p&gt;Survival probability (pt): The probability, computed from prior beliefs F at the beginning of times, that the catastrophe has not yet occurred by date t. Its law of motion is ṗt = α[1 − F(Qt) − pt], driven solely by the catastrophe parameter α and the current maximum stock Qt.&lt;/p&gt;
&lt;p&gt;Fatalism (under Theorem 1): The policy implication that a higher legacy — meaning a higher probability the catastrophe is already triggered — leads the planner to increase the stock further and accept more experimentation, because mitigation is relatively ineffective (QE &amp;lt; QD) and current consumption must be enjoyed before the catastrophe arrives.&lt;/p&gt;</description></item><item><title>Civil War–Induced Displacement and Human Capital</title><link>https://macropaperwarehouse.com/papers/civil-warinduced-displacement-and-human-capital/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/civil-warinduced-displacement-and-human-capital/</guid><description>&lt;p&gt;This paper examines the impact of conflict-driven forced displacement on human capital accumulation using the Mozambican civil war (1977–1992) as the empirical setting. During this war, over four million civilians — roughly a third of the population — fled to rural areas, cities, neighboring countries, or UN-managed refugee camps. The study advances on prior work in three dimensions: it uses the full post-war population census (12 million individuals) rather than a small survey; it studies multiple displacement trajectories in a single framework; and it separately identifies place-based exposure effects from a general uprootedness effect.&lt;/p&gt;
&lt;p&gt;The primary data source is the 1997 Mozambican census, which records each individual&amp;rsquo;s place of birth, residence in 1992 (the war&amp;rsquo;s end), and residence in 1997. Key outcomes are educational attainment and sectoral employment (agricultural versus services). The authors supplement the census with digitized colonial road and school maps, georeferenced conflict events, and landmine contamination data.&lt;/p&gt;
&lt;p&gt;The main identification strategy compares approximately 135,000 siblings (from 45,000 families) separated during the war, using the sibling who stayed behind as a within-family counterfactual. This design controls for household-level characteristics including religious and ethnic background, aspirations, and exposure to violence.&lt;/p&gt;
&lt;p&gt;The key findings are as follows. First, rural-born IDPs displaced to cities have a 7.3 percentage point higher likelihood of attending primary school and 0.53 more years of schooling compared to their siblings who stayed behind — roughly one-third of the non-displaced mean. Rural-born IDPs displaced to other rural areas also show gains, with a 3 percentage point higher likelihood of attending school and 0.24 additional years, supporting the uprootedness hypothesis even for displacements that did not reach urban centers. Urban-born IDPs forcibly relocated to the countryside — primarily through FRELIMO&amp;rsquo;s villagization scheme — experienced 9 percentage point lower primary school attendance and approximately 0.5 fewer years of schooling relative to siblings who remained in cities.&lt;/p&gt;
&lt;p&gt;External displacement (to camps in Malawi or Zimbabwe) generated no significant schooling gains relative to staying siblings, despite UN-built schools in camps, likely because scarce employment opportunities reduced perceived returns to education.&lt;/p&gt;
&lt;p&gt;Second, the paper jointly estimates place-based and uprootedness effects in a single within-family framework. Place effects are statistically significant: displacement to a district one standard deviation more developed than one&amp;rsquo;s birthplace raises schooling likelihood by approximately 3 percentage points (OLS) to 5 percentage points (2SLS reduced form). Crucially, a residual uprootedness effect of approximately 2–4 percentage points persists even after controlling fully for destination-origin differences in development and conflict intensity. This uprootedness effect is quantitatively comparable to being displaced to a district one standard deviation more developed than one&amp;rsquo;s birthplace.&lt;/p&gt;
&lt;p&gt;Third, a primary survey of 208 Nampula residents conducted in early 2020 — three decades after the war — confirms lasting educational gains. IDPs displaced to Nampula have a 10 percentage point higher likelihood of completing primary school relative to their siblings who stayed in the countryside, and their educational attainment converged to levels of urban-born, never-displaced residents despite large urban-rural education gaps. However, IDPs report significantly lower social capital, civic participation, and community trust than urban-born respondents, and score significantly worse on mental health indicators, including depression, loneliness, and pessimism. These psychosocial costs persist three decades after the war&amp;rsquo;s end.&lt;/p&gt;
&lt;p&gt;The findings apply to a low-income, post-colonial African setting characterized by widespread illiteracy (over 60%) and subsistence agriculture (over 85% of employment) at the war&amp;rsquo;s close. The results are robust to alternative age restrictions, extended family comparisons, dropping the oldest sibling, same-sex sibling pairs, and narrowing the age gap between sibling pairs to as few as two years.&lt;/p&gt;
&lt;p&gt;Q: What is the core identification strategy and why is it preferred over cross-sectional estimates?
A: The authors compare siblings within the same household who experienced different displacement trajectories during the war. Because siblings share household-level characteristics — parental preferences for education, ethnic and religious background, wealth, and local conflict exposure — the within-family design controls for confounders that would bias cross-sectional estimates. The within-family estimates are systematically smaller than cross-sectional ones (e.g., 7.3 pps vs. 24–30 pps for rural-to-urban displacement in primary school attendance), confirming that sorting was present even in the unpredictable civil war setting.&lt;/p&gt;
&lt;p&gt;Q: What do the results show for rural-born IDPs displaced to urban centers?
A: Within the sibling-pair framework, rural-born IDPs displaced to cities and towns have a 7.3 percentage point higher likelihood of attending primary school and 0.53 more years of schooling compared to their siblings who stayed in rural birthplaces, against a non-displaced sibling mean of approximately 20% primary school access and one year of formal schooling. These IDPs also show a 4 percentage point higher likelihood of non-agricultural employment five years after the war&amp;rsquo;s end.&lt;/p&gt;
&lt;p&gt;Q: What do the results show for rural-born IDPs displaced to other rural areas?
A: Even displacement to a different rural district — not a city — generates modest but statistically significant gains: a 3 percentage point higher likelihood of attending school and 0.24 additional years of schooling relative to siblings staying in their birthplace rural district. The authors interpret this as evidence for the uprootedness hypothesis, since rural Mozambique at the time was among the most impoverished and insecure environments in the world, meaning destination quality alone cannot explain the gain.&lt;/p&gt;
&lt;p&gt;Q: What do the results show for externally displaced refugees?
A: Refugees displaced to camps and settlements in Malawi, Zimbabwe, Tanzania, Zambia, and Swaziland show schooling levels statistically similar to their siblings who remained in their rural birthplaces, despite UN-built primary schools in camps. The authors attribute the absence of gains to low perceived returns to education stemming from scarce employment opportunities at displacement destinations. Externally displaced individuals do show a 5 percentage point lower likelihood of agricultural employment relative to staying siblings.&lt;/p&gt;
&lt;p&gt;Q: What are the consequences of urban-to-rural forced displacement?
A: Urban-born individuals forcibly relocated to the countryside — primarily through FRELIMO&amp;rsquo;s villagization and food production programs — have approximately 9 percentage point lower likelihood of attending primary school and 0.5 fewer years of schooling compared to siblings who remained in urban areas. These results indicate that FRELIMO&amp;rsquo;s coercive relocation policies imposed material human capital costs on the displaced.&lt;/p&gt;
&lt;p&gt;Q: How are place-based and uprootedness effects separated empirically?
A: The authors construct principal component indices for destination-origin differences in regional development (aggregating population density, Portuguese-speaking share, offspring mortality, road density, colonial market density, and school density) and conflict intensity (conflict events per capita and landmine contamination per capita). They then include these continuous exposure measures alongside a binary displacement indicator in within-family regressions. The coefficient on the binary displacement indicator — conditional on destination-origin development and conflict differences — isolates the uprootedness effect for individuals displaced to districts with identical characteristics to their birthplace.&lt;/p&gt;
&lt;p&gt;Q: What are the magnitudes of the place-based and uprootedness effects?
A: Under OLS, displacement to a district one standard deviation more developed than one&amp;rsquo;s birthplace raises schooling likelihood by approximately 3 percentage points. The residual uprootedness effect — displacement per se, controlling for destination quality — raises schooling likelihood by approximately 2 percentage points. Under 2SLS (instrumenting destination-origin development differences with the development of districts within 100 km of birthplace), the place-based effect rises to approximately 5 percentage points in the reduced form, and the uprootedness effect remains significant at approximately 4 percentage points. Both the uprootedness and place-based effects are of comparable magnitude.&lt;/p&gt;
&lt;p&gt;Q: What instrument is used in the 2SLS specifications and what is its first-stage strength?
A: The instrument exploits the fact that Mozambique&amp;rsquo;s heavily mined and rudimentary transportation network constrained civilian movement — the median displaced sibling ended up roughly 97 kilometers from birthplace. The authors instrument actual destination-origin development and conflict differences with the predicted differences based on the characteristics of districts within 100 km of the birthplace. The first-stage elasticity between actual and proximity-predicted differences in development is 0.86, and for conflict is 0.88, both precisely estimated.&lt;/p&gt;
&lt;p&gt;Q: What do the long-run survey results from Nampula show about educational persistence?
A: In a 2020 survey of 208 Nampula residents aged over 35, IDPs who fled to Nampula during the war have a 10 percentage point higher likelihood of completing primary school relative to their siblings who stayed in the countryside. Their educational attainment converges to the level of urban-born, never-displaced Nampula residents, despite large historical and contemporary urban-rural education gaps in northern Mozambique. The majority of IDPs (73%) report that extended relatives or friends advised them to attend school upon arriving in the city, and most believed education was necessary for urban employment.&lt;/p&gt;
&lt;p&gt;Q: What are the long-run psychosocial costs documented in the Nampula survey?
A: Even three decades after the war&amp;rsquo;s end, IDPs in Nampula report significantly lower social capital, civic participation, and community trust compared to urban-born never-displaced residents. IDPs also score significantly worse on mental health indicators including depression, loneliness, and pessimism. These findings suggest that forced displacement imposes persistent psychosocial costs that are not remediated by economic or educational convergence.&lt;/p&gt;
&lt;p&gt;Q: What drives displacement in the data, and does selection threaten identification?
A: Linear probability and multinomial logit models show that conflict intensity and geographic proximity (distance to the border for external displacement; distance to cities for urban displacement) are the primary correlates of displacement type, while differences in destination development are uncorrelated with displacement. Nevertheless, the overall explanatory power of these models is low, confirming many idiosyncratic and unpredictable features of the war. The within-family design addresses residual selection on household characteristics, and the 2SLS design addresses selection on destination-specific characteristics.&lt;/p&gt;
&lt;p&gt;Q: How do educational gains translate into sectoral employment outcomes?
A: Across specifications, gains in schooling move in tandem with a shift out of agriculture into services. Rural-to-urban IDPs have a 4 percentage point higher likelihood of non-agricultural employment five years after the war, while externally displaced show a 5 percentage point lower likelihood of agricultural employment. Urban-born IDPs displaced to the countryside are more likely to work in agriculture after the war. The authors interpret this co-movement as suggesting that conflict-driven human capital accumulation may contribute to structural transformation away from subsistence agriculture.&lt;/p&gt;
&lt;p&gt;Q: How robust are the within-family estimates?
A: The authors conduct six sensitivity checks: adding family fixed effects to cross-sectional regressions, restricting to individuals aged 12–18 in 1997 to address co-habitation concerns, extending comparisons to cousins and other relatives, dropping the oldest male sibling to minimize favoritism concerns, restricting to same-sex sibling pairs, and narrowing the age gap to two years. Across all permutations, the qualitative ordering is preserved: refugees show no significant schooling gains, rural-to-urban IDPs show gains of 5–6 percentage points in primary attendance and 0.35–0.5 extra years, rural-to-rural IDPs show small positive gains, and urban-to-rural IDPs show losses.&lt;/p&gt;
&lt;p&gt;Uprootedness hypothesis: The idea, traced in the paper to Stigler and Becker (1977) and earlier scholars, that forced displacement incentivizes human capital investment precisely because education is a mobile asset that cannot be expropriated — distinct from place-based effects of destination quality.&lt;/p&gt;
&lt;p&gt;Place-based (exposure) effects: The impact on human capital outcomes attributable to differences between the development level and conflict intensity of the displacement destination and the individual&amp;rsquo;s birthplace, measured as destination-origin differences in a principal component index of regional development.&lt;/p&gt;
&lt;p&gt;Separated siblings design: An identification strategy that compares siblings from the same household who experienced different displacement trajectories during the war, holding constant all household-level characteristics including parental preferences, ethnicity, religion, wealth, and local conflict exposure.&lt;/p&gt;
&lt;p&gt;Internal displacement (IDP): Conflict-driven movement within national borders to either rural areas or urban centers, constituting approximately 60% of global forced displacement and the majority of displacement in the Mozambican civil war context.&lt;/p&gt;
&lt;p&gt;Source text origin: A categorization of the working paper text used for summarization — distinguishing full PDF or HTML text from abstract-only text. Abstract-only text is a hard block for summary generation in the pipeline.&lt;/p&gt;
&lt;p&gt;Structural transformation: In this paper&amp;rsquo;s usage, the shift of workers out of subsistence agriculture into services associated with human capital accumulation triggered by conflict-driven displacement, treated as a potential mechanism of post-conflict recovery.&lt;/p&gt;
&lt;p&gt;Psychosocial costs of displacement: Long-run deficits in social capital, civic engagement, community trust, and mental health (depression, loneliness, pessimism) reported by IDPs three decades after displacement, persisting despite convergence in educational attainment and employment.&lt;/p&gt;</description></item><item><title>Comment on "Artificial Intelligence and Technological Unemployment" by Wang and Wong</title><link>https://macropaperwarehouse.com/papers/comment-on-artificial-intelligence-and-technological-unemployment-by-wang-and-wong/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/comment-on-artificial-intelligence-and-technological-unemployment-by-wang-and-wong/</guid><description>&lt;p&gt;This comment, written by J. Carter Braxton (University of Wisconsin), discusses the paper &amp;ldquo;Artificial Intelligence and Technological Unemployment&amp;rdquo; by Wang and Wong (2025), which develops and quantifies an equilibrium labor search model to evaluate the employment effects of spreading AI. Wang and Wong&amp;rsquo;s central finding is that improvements in AI quality will increase productivity by a factor of three while reducing employment by 23%, with approximately half of the employment decline occurring within the next five years. Braxton&amp;rsquo;s comment serves two purposes: first, to clarify the model&amp;rsquo;s structural channels through which AI affects employment; and second, to bring empirical evidence from the spread of computers in the 1980s–2000s to bear on the relative magnitude of those channels.&lt;/p&gt;
&lt;p&gt;Braxton identifies two competing forces within Wang and Wong&amp;rsquo;s framework. The &lt;strong&gt;job destruction channel&lt;/strong&gt; arises from endogenous separations: as AI quality improves, firms increasingly replace matched workers with AI, raising outflows from employment. The &lt;strong&gt;job creation channel&lt;/strong&gt; arises from the free-entry condition: rising AI quality increases firm profits on all matches, inducing firms to post more vacancies, which raises workers&amp;rsquo; job-finding rates and employment inflows. Whether aggregate employment rises or falls depends on which channel dominates — a quantitative question the authors resolve through calibration, finding the job destruction channel dominant. Braxton notes that three modeling choices (learning-by-using, the requirement that firms must be matched with a worker to adopt AI, and disembodied technological change) each push &lt;em&gt;against&lt;/em&gt; the job-destruction result, making the authors&amp;rsquo; findings more striking.&lt;/p&gt;
&lt;p&gt;Braxton then evaluates the relative strength of these channels using the historical spread of personal computers. Drawing on Bick, Blandin, and Deming (2024), he notes that workplace AI adoption in 2024 follows nearly the same time trend and income-distribution profile as computer adoption in 1984, making computers a plausible historical analog. Using the CPS Computer Supplement (1984–2003), Braxton measures the change in computer usage by occupation and regresses it against the change in employment-to-unemployment (EU) transition rates by occupation. The estimated coefficient is 0.0146 (robust SE 0.0064), indicating that occupations with higher computer adoption rates saw higher flows into unemployment — confirming that a job destruction channel was active during the computer era. However, regressing the change in log occupation-level employment (1980–2000 Census) on the change in computer usage yields a coefficient of 0.7761 (robust SE 0.2658), with a positive slope indicating that occupations more exposed to computers saw &lt;em&gt;higher&lt;/em&gt; employment growth. For the computer episode, therefore, the job creation channel dominated the job destruction channel — the opposite of Wang and Wong&amp;rsquo;s AI projection.&lt;/p&gt;
&lt;p&gt;Braxton also cites his own prior work showing that even when job creation and destruction balance in aggregate, workers displaced by technological change face lasting earnings losses and elevated permanent income risk, raising the question of how to optimally insure these workers.&lt;/p&gt;
&lt;p&gt;The comment concludes by identifying avenues for future research: introducing occupational heterogeneity (with some occupations more exposed to AI than others) and worker heterogeneity (skills that are complements versus substitutes to AI). The central open question is whether AI is qualitatively different from prior episodes of technological change, and if so, why.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q1: What are the two central channels through which AI quality affects employment in Wang and Wong&amp;rsquo;s model, and how do they operate?&lt;/strong&gt;
The job destruction channel operates through endogenous separations: as AI quality (At) improves, firms that are matched with workers are more likely to replace them with AI at rate ρ, adding the term ρµAt Ht It to outflows from employment in the law of motion for employment. The job creation channel operates through the free-entry condition: higher AI quality raises firm profits on all existing matches (because technological change is disembodied, benefiting matches formed today with future AI gains), inducing firms to post more vacancies, which via free entry reduces the firm&amp;rsquo;s matching probability but raises the worker&amp;rsquo;s job-finding rate αt and thereby increases employment inflows. The net employment effect depends on which channel quantitatively dominates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q2: What is Wang and Wong&amp;rsquo;s quantitative finding about the aggregate employment and productivity effects of AI?&lt;/strong&gt;
Using a calibrated equilibrium labor search model, Wang and Wong find that the spread of AI will increase productivity by a factor of three while reducing employment by 23%. Approximately half of the employment decline is projected to occur within the next five years. A version of the model holding job-finding rates fixed yields a similar result, indicating that through the lens of their model the job creation channel is quantitatively small and the job destruction channel dominates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q3: What three modeling choices push against Wang and Wong&amp;rsquo;s job-destruction result, and why does Braxton view this as making the finding more striking?&lt;/strong&gt;
First, AI improves through &amp;ldquo;learning by using&amp;rdquo; — it learns from all output being produced — which creates an incentive for employment to remain elevated to accelerate AI learning, dampening job destruction. Second, firms can only adopt AI if currently matched with a worker, which creates an incentive for vacancy posting and pushes in favor of job creation. Third, AI improvements are disembodied (raising productivity in all matches, including those formed before the improvement), which increases the value of forming new matches today and strengthens job creation. Because each of these assumptions pushes against the job destruction result, Braxton argues that finding job destruction dominant despite these model features makes the result more striking.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q4: How does Braxton use the historical spread of computers to assess the job destruction and job creation channels?&lt;/strong&gt;
Braxton measures occupation-level computer adoption as the change in the share of CPS Computer Supplement respondents who reported using a computer at work between 1984 and 2003 (denoted ΔCPUo,84–03), using occupation codes from Autor and Dorn (2013). He then regresses the occupation-level change in EU transition rates (ΔEUo,84–03, from monthly CPS micro data) on ΔCPUo,84–03 to measure the job destruction channel, and separately regresses the change in log occupation-level employment (Δlog Eo,80–00, from the 1980 and 2000 Census IPUMS) on ΔCPUo,84–03 to assess the net employment effect. A positive coefficient on the employment regression indicates job creation dominates; a negative coefficient indicates job destruction dominates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q5: What do the regression results show about the job destruction and job creation channels during the computer era?&lt;/strong&gt;
The job destruction regression yields a coefficient of β = 0.0146 (robust SE = 0.0064, R² = 0.0178), indicating that occupations with higher computer adoption rates did see higher employment-to-unemployment transition rates — the job destruction channel was present. However, the employment-level regression yields a coefficient of β = 0.7761 (robust SE = 0.2658, R² = 0.0348), with a positive slope indicating that occupations more exposed to computers experienced &lt;em&gt;higher&lt;/em&gt; employment growth between 1980 and 2000. Thus, for the computer episode, the job creation channel dominated the job destruction channel — the opposite of what Wang and Wong project for AI.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q6: What is the basis for treating the computer episode as a relevant analog to the spread of AI?&lt;/strong&gt;
Braxton cites Bick, Blandin, and Deming (2024), who show that AI adoption in the workplace in 2024 is following nearly the same aggregate time trend as the spread of personal computers in the early 1980s. Moreover, the distribution of AI usage across the income distribution in 2024 is nearly identical to computer usage across the income distribution in 1984: for both technologies, workplace usage peaks between the 80th and 90th percentiles of the income distribution before declining modestly at the top. Bick et al. (2024) also show the similarities hold by education level and age.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q7: Even if job creation and destruction balance in aggregate, what does prior work suggest about the distributional consequences for workers?&lt;/strong&gt;
Braxton and Taska (2023) show that workers in occupations more exposed to technological change (measured by changes in computer and software task requirements) suffered larger earnings losses following displacement. Braxton, Herkenhoff, Rothbaum, and Schmidt (2024, forthcoming AER) show that workers in occupations more exposed to technological change experienced larger increases in permanent income risk between the 1980s and 2010s. These findings imply that even if AI does not reduce aggregate employment, workers who are displaced will face deteriorating labor market prospects, raising the question of how to optimally provide insurance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q8: What policy implication does Braxton draw from the distributional consequences of technological change?&lt;/strong&gt;
Braxton and Taska (2025, forthcoming Review of Economic Dynamics) show that technological change expands the motive for governments to provide retraining subsidies. Braxton argues that if AI represents an acceleration of technological change, even larger retraining subsidies — and potentially other forms of insurance — may be needed for displaced workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q9: What are the main avenues for future research identified in the comment?&lt;/strong&gt;
Braxton identifies two principal directions. First, introducing occupational heterogeneity into the Wang-Wong framework, so that some occupations are more exposed to AI displacement than others, would allow the model to generate richer distributional implications. Second, allowing worker heterogeneity in skills — distinguishing skill dimensions that are complements to AI from those that are substitutes — would permit the model to capture differential effects across the workforce. The overarching research question is whether AI is qualitatively different from prior technological change episodes, and if so, to identify the precise mechanisms that make it different.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Job destruction channel&lt;/strong&gt;: In the Wang-Wong model, the increase in endogenous separations driven by firms replacing matched workers with AI as AI quality improves. Formally, this is the term ρµAt Ht It in the law of motion for employment, representing separations that occur when a firm adopts AI and the worker and firm cannot renegotiate a mutually acceptable wage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Job creation channel&lt;/strong&gt;: The increase in vacancy posting and worker job-finding rates induced by rising AI quality. Because higher AI quality raises firm profits on all matches (via disembodied technological change), the free-entry condition implies firms post more vacancies, lowering the firm&amp;rsquo;s matching probability but raising the worker&amp;rsquo;s job-finding rate αt, increasing employment inflows.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Free-entry condition&lt;/strong&gt;: The equilibrium condition equating the cost of posting a vacancy (κt) to the expected benefit (the probability of matching ft times the firm&amp;rsquo;s match surplus Πt). This condition pins down the job-finding rate for workers: when firms find it more profitable to post vacancies, αt rises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Disembodied technological change&lt;/strong&gt;: The modeling assumption that AI quality improvements raise productivity in all existing matches, not just those formed after the improvement. This means future AI gains benefit matches formed today, increasing the incentive to create new matches and pushing in favor of the job creation channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Learning by using&lt;/strong&gt;: The mechanism in Wang-Wong whereby AI quality (At) improves as a function of current aggregate employment (Ht) and the learning rate µ. Because AI learns from all output being produced, maintaining higher employment accelerates AI improvement, creating a motive that partially offsets the job destruction channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Employment-to-unemployment (EU) transition rate&lt;/strong&gt;: The rate at which employed workers flow into unemployment in a given occupation, used by Braxton as the empirical measure of the job destruction channel during the computer episode. Measured from monthly CPS micro data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capitalization effect&lt;/strong&gt;: The tendency for firms to post more vacancies today in anticipation of future productivity improvements, because the cost of posting is paid upfront while the benefits of a future-better-AI accrue to the match going forward. Referenced by Braxton as relevant to understanding the job creation channel in Wang-Wong&amp;rsquo;s framework (citing Pissarides (2000), Chapter 3).&lt;/p&gt;</description></item><item><title>Customer accumulation, returns to scale, and secular trends</title><link>https://macropaperwarehouse.com/papers/customer-accumulation-returns-to-scale-and-secular-trends/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/customer-accumulation-returns-to-scale-and-secular-trends/</guid><description>&lt;p&gt;This paper asks how rising returns to scale in production contributed to three concurrent U.S. secular trends since 1980: declining business dynamism, rising markups, and growing firm expenditures on customer acquisition. The author constructs a firm dynamics model in the Hopenhayn (1992) tradition with endogenous entry and exit, heterogeneous markups, and customer accumulation grounded in directed search in the product market. Firms compete for customers through both prices and selling activities; larger firms gain a competitive edge when returns to scale rise because their marginal costs fall more than those of smaller firms—even though the technological shift is uniform across firms. This demand-based channel triggers winners-and-losers dynamics and the rise of superstar firms.&lt;/p&gt;
&lt;p&gt;The empirical foundation rests on Compustat data for U.S. publicly traded firms (1977–2014) and Business Dynamics Statistics (BDS) for aggregate and sector-level dynamism measures. Production-function estimation using Ackerberg, Caves, and Frazer (2015) augmented with sales-share controls documents that aggregate returns to scale rose from approximately 1.0 in 1980 to approximately 1.05 by 2014—a within-sector increase, not a reallocation effect. Over the same period, the cost-weighted markup rose by 42%, the firm entry rate fell by 33%, the excess reallocation rate fell by 29%, and selling costs relative to production costs rose by 60%–90% depending on the measure used.&lt;/p&gt;
&lt;p&gt;The model is calibrated to 1980 steady-state moments (firm life-cycle patterns, markups, entry and reallocation rates). A 5% increase in returns to scale—matching the empirical estimate—accounts for: a +15 percentage point rise in the average cost-weighted markup (vs. +42% in the data); a 33% decline in the entry rate (exactly matching the data); a 21% decline in the reallocation rate (vs. 29% in the data); and a 23% increase in selling costs relative to production costs (vs. 60%–90% in the data). The model also generates a 53% rise in the share of firms aged 11 years or older (vs. 50% in the data) and a 58% decline in the employment share of firms aged 5 years or younger (vs. 56% in the data), closely tracking the aging of the U.S. firm population. Firm-level responsiveness to productivity shocks declines by 0.08 in the model, versus about 0.01 in Compustat and 0.09 in Decker et al. (2020).&lt;/p&gt;
&lt;p&gt;Sector-level panel regressions with sector fixed effects confirm the model&amp;rsquo;s directional predictions: within-sector increases in returns to scale are associated with lower entry rates (coefficient −2.89, significant at 1%), lower reallocation rates (−1.16, significant at 1%), higher markups (+3.15, significant at 1%), and higher selling costs relative to production costs (+1.85 for the advertising-based measure; +8.52 for adjusted SG&amp;amp;A).&lt;/p&gt;
&lt;p&gt;A key scope condition is that the model yields a constrained-efficient allocation: directed search and full internalization of returns to scale imply decentralized equilibrium efficiency, making the paper a laboratory for assessing how far efficient firm responses to technological change can explain the secular trends without invoking market failures. The model fits the post-2000 transition dynamics better than the 1980s–1990s period, and explains a substantial but incomplete share of the trends, suggesting complementary—possibly inefficient—forces also contributed.&lt;/p&gt;
&lt;p&gt;Q: What is the core mechanism through which rising returns to scale generate winners-and-losers dynamics?&lt;/p&gt;
&lt;p&gt;A: The marginal cost of production under increasing returns to scale (alpha &amp;gt; 1) is MC(z,n) = l(n,z)^(1−alpha) × (1/alpha) × (W/e^z), which depends on firm size l(n,z). A uniform rise in alpha rotates the marginal cost schedule clockwise by firm size: larger firms see a proportionally larger cost reduction than smaller firms, even though the technological change is identical across all firms. Because firms compete for the same pool of customers, this asymmetric cost advantage allows large firms to offer lower prices while sustaining higher margins, attracting customers away from small firms. The result is a demand-based channel that generates winners-and-losers dynamics and increases market concentration.&lt;/p&gt;
&lt;p&gt;Q: How does the model capture customer accumulation, and why is it central to the paper&amp;rsquo;s argument?&lt;/p&gt;
&lt;p&gt;A: The model introduces directed search in the product market, where firms post advertisements and customers—including those already matched with a firm—choose which submarket to enter by trading off offered utility against matching probability. A constant-returns-to-scale matching function governs match creation; in submarket with tightness theta, customers match with probability m(theta) = theta(1+theta)^(−1) and firms attract customers with probability q(theta) = (1+theta)^(−1). The customer accumulation motive creates an investment-harvest trade-off: firms can either post high promised utility (low prices) to grow their customer base or extract surplus through high prices. Rising returns to scale amplify large firms&amp;rsquo; ability to resolve this trade-off favorably, linking the technological change directly to markup dynamics, entry incentives, and selling expenditures.&lt;/p&gt;
&lt;p&gt;Q: What is the directed search framework&amp;rsquo;s role in ensuring equilibrium uniqueness and efficiency?&lt;/p&gt;
&lt;p&gt;A: The author introduces firm-side commitment contracts—specifying price, separation probability, and continuation utility contingent on productivity realizations—combined with directed search. Because search is directed on both sides and firms fully internalize returns to scale, the decentralized equilibrium is constrained-efficient. This delivers uniquely determined heterogeneous prices in equilibrium (solving the indeterminacy problem common in customer-market models) and establishes the paper&amp;rsquo;s efficient-mechanism benchmark: it tests how far profit-maximizing firm responses to technological change—without any market failure—can account for the secular trends.&lt;/p&gt;
&lt;p&gt;Q: How are prices structured in the model, and what life-cycle pattern do they generate?&lt;/p&gt;
&lt;p&gt;A: Each firm charges two distinct prices in each period: one to incumbent customers (the same for all incumbents, since they are identical conditional on being attached to the same firm) and one to newly acquired customers (which varies based on the promised utility in the submarket searched). Firms that are expanding their customer base offer greater promised utility and therefore charge lower prices to attract customers; firms harvesting their existing base charge higher prices. Because firms enter small and grow, this dynamic generates a price life cycle: young firms invest via low prices and mature firms harvest through higher prices, which the model reproduces as a rising markup pattern over the firm life cycle—an untargeted moment the model fits well.&lt;/p&gt;
&lt;p&gt;Q: What does the calibration target and what untargeted moments does the model reproduce?&lt;/p&gt;
&lt;p&gt;A: The model is calibrated to 1980 using: the number of employees of entrant firms (pinning entry customer base n_e), employees of age-5 firms (pinning convex cost chi_1), share of firms aged 11+ years (pinning chi_2), average firm size (operating cost f), entry rate (entry cost kappa), excess reallocation rate (exit shock delta), and average cost-weighted markup (linear cost c). Untargeted moments reproduced include: a sales-weighted markup of 0.28 (vs. 0.25 in De Loecker et al. 2020), endogenous customer turnover of approximately 9% (vs. 15% in Gourio and Rudanko 2014), and an elasticity of customer base shrinkage to price of 0.08 (within the 0.01–0.16 range from Paciello et al. 2019). The model also matches markup and selling-cost life-cycle patterns that are typically overlooked.&lt;/p&gt;
&lt;p&gt;Q: How large is the quantitative contribution of the 5% rise in returns to scale to each secular trend?&lt;/p&gt;
&lt;p&gt;A: Comparing the 1980 steady state (alpha = 1) to the 2014 steady state (alpha = 1.05): the average cost-weighted markup rises by 15% in the model versus 42% in the data; the entry rate declines by 33% in the model, exactly matching the data; the reallocation rate declines by 21% in the model versus 29% in the data; and selling costs relative to production costs rise by 23% in the model versus 60%–90% in the data. The model thus explains a substantial share of each trend while leaving a residual requiring additional mechanisms.&lt;/p&gt;
&lt;p&gt;Q: How does the model explain the aging of U.S. firms, and how well does it match the data?&lt;/p&gt;
&lt;p&gt;A: The winners-and-losers mechanism shifts activity toward larger, older firms, which mechanically ages the firm population. The model generates a 53% increase in the share of firms aged 11 years or older (vs. 50% in the data) and a 58% decline in the employment share of firms aged 5 years or younger (vs. 56% in the data). This aging arises because rising returns to scale increase the cost of customer acquisition, acting as a barrier to entry that disproportionately hurts new, small firms while allowing large incumbents to remain viable at lower productivity thresholds.&lt;/p&gt;
&lt;p&gt;Q: What is the channel through which rising returns to scale reduce business dynamism specifically?&lt;/p&gt;
&lt;p&gt;A: The unequal reduction in marginal costs intensifies competition for customers and raises customer acquisition costs. This operates through two simultaneous effects on the exit threshold: (i) lower marginal costs allow large firms to remain viable at lower productivity levels despite higher customer acquisition costs; and (ii) heightened competition forces smaller firms to require higher productivity to survive in a market that has become increasingly costly to operate in. Higher customer acquisition costs therefore function as an endogenous barrier to entry, reducing the entry rate and the reallocation of resources across firms.&lt;/p&gt;
&lt;p&gt;Q: Does the model attribute the secular trends entirely to efficient firm behavior, and what does it conclude about residual explanations?&lt;/p&gt;
&lt;p&gt;A: No. The model is explicitly designed as a constrained-efficient benchmark, and the paper finds that while rising returns to scale account for a substantial share of the trends—particularly in magnitude—the transition dynamics show a less accurate fit before the 2000s. The author concludes that complementary mechanisms, likely involving inefficiencies (such as market power from horizontal product differentiation or barriers to entry beyond those captured by the model), played a significant role in the earlier evolution of these trends and in the portion of the trends not explained by the efficient channel.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports the rising returns to scale finding, and what are its limitations?&lt;/p&gt;
&lt;p&gt;A: Production-function estimation using the Ackerberg-Caves-Frazer method with sales-share controls on Compustat data shows returns to scale rising from approximately 1.0 in 1980 to approximately 1.05 by 2014, driven primarily by within-sector increases rather than reallocation toward high-returns sectors. A translog production function finds limited evidence of heterogeneous increases across firm sizes within Compustat. However, Compustat predominantly covers large publicly traded firms; smaller firms outside the sample may have experienced minimal or no increase in returns to scale. If technology adoption involves fixed costs, the aggregate impact could be larger than estimated, meaning the quantitative exercises likely represent a conservative lower bound.&lt;/p&gt;
&lt;p&gt;Q: How does the paper relate to and extend the directed search literature in product markets?&lt;/p&gt;
&lt;p&gt;A: The paper builds on Gourio and Rudanko (2014) and Roldan-Blanco and Gilbukh (2020), where customers are locked in once matched, by introducing labor-search tools from Schaal (2017) to allow: (i) incumbent customer switching between firms at rates of 10%–25% annually (Gourio and Rudanko 2014), and (ii) a non-zero price sensitivity of incumbent customers (Paciello et al. 2019). It also allows firms to invest in demand through selling expenditures, which prior directed search models in product markets typically abstracted from, making it possible to study how technological changes affect customer reallocation and firms&amp;rsquo; cost structures jointly.&lt;/p&gt;
&lt;p&gt;Customer capital: The stock of customers a firm has accumulated through prior selling and pricing decisions; treated as a state variable that firms invest in (by offering low prices and spending on advertisements) or harvest from (by charging high markups), with a customer turnover rate estimated at 10%–25% annually in the literature.&lt;/p&gt;
&lt;p&gt;Directed search in the product market: A market structure in which both firms and customers choose which submarket (indexed by the promised utility level) to enter, trading off match probability against terms; delivers constrained-efficient equilibrium and uniquely determined heterogeneous prices.&lt;/p&gt;
&lt;p&gt;Investment-harvest trade-off: The firm&amp;rsquo;s dynamic choice between offering high promised utility (low prices, low current markups) to grow the customer base versus extracting surplus through high prices from an existing customer base; shaped by the firm&amp;rsquo;s current size, productivity, and the cost structure implied by returns to scale.&lt;/p&gt;
&lt;p&gt;Returns to scale (alpha): The curvature of the production function y = e^z × l^alpha; equals 1.0 under constant returns and approximately 1.05 by 2014 in the empirical estimates; the paper&amp;rsquo;s central technological change parameter, whose rise disproportionately reduces marginal costs for larger firms.&lt;/p&gt;
&lt;p&gt;Winners-and-losers dynamics: The reallocation of customers and market share from small to large firms triggered by the asymmetric cost advantage large firms obtain when returns to scale rise; the demand-based channel through which superstar firms emerge.&lt;/p&gt;
&lt;p&gt;Cost-weighted markup: The average markup aggregated using each firm&amp;rsquo;s costs as weights, as opposed to sales-weighted markup; the primary measure of market power used in the paper, rising by 42% in the data between 1980 and 2014.&lt;/p&gt;
&lt;p&gt;Constrained-efficient allocation: An equilibrium outcome in which, given the frictions present (search-and-matching in the product market), no social planner operating under the same constraints could improve welfare; the paper uses this as a benchmark to assess how far efficient firm responses explain secular trends without invoking market failures.&lt;/p&gt;
&lt;p&gt;Selling costs relative to production costs: The ratio of customer acquisition expenditures (advertising or adjusted SG&amp;amp;A) to cost of goods sold; rose by 60%–90% in the data between 1980 and 2014 and by 23% in the model&amp;rsquo;s steady-state comparison.&lt;/p&gt;</description></item><item><title>Customer Acquisition, Business Dynamism and Aggregate Growth</title><link>https://macropaperwarehouse.com/papers/customer-acquisition-business-dynamism-and-aggregate-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/customer-acquisition-business-dynamism-and-aggregate-growth/</guid><description>&lt;p&gt;This paper asks whether firm-level customer acquisition — distinct from productivity differences — is a quantitatively important driver of aggregate economic growth, and whether ignoring it distorts predictions about growth policy efficacy. The authors build a novel endogenous growth model in which innovating firms must first accumulate customers to sell their products, with two channels of customer acquisition operating simultaneously: costly sales-and-marketing expenditure and below-static-markup pricing (sales-driven accumulation). The model is estimated using indirect inference against a combination of aggregate data (U.S. real GDP per worker growth of 1.43% annually, 1979–2019), Business Dynamics Statistics (BDS) life-cycle profiles, and firm-level data from Compustat matched to Capital IQ&amp;rsquo;s sales-and-marketing expense records covering 1997–2019.&lt;/p&gt;
&lt;p&gt;The benchmark model yields four closed-form propositions. First, a &amp;ldquo;firm-level market size effect&amp;rdquo;: higher customer retention raises a firm&amp;rsquo;s future profit base, strengthening incentives to conduct R&amp;amp;D. Second, an endogenous feedback loop: more productive firms invest more in customer acquisition, which expands their customer base and further strengthens R&amp;amp;D incentives. Third, customer base accumulation raises aggregate growth, but only indirectly — by boosting firm-level innovation rates — since aggregate productivity is a customer-weighted average of firm productivity levels. Fourth, the sensitivity of innovation to R&amp;amp;D subsidies increases with customer base growth, because firms with faster-growing customer bases discount future profits less steeply.&lt;/p&gt;
&lt;p&gt;In the quantitatively estimated full model — which relaxes the benchmark&amp;rsquo;s perfect-scaling restrictions and endogenizes firm entry and exit — the authors conduct two decomposition exercises. In a counterfactual scenario where expected customer retention is reduced to make average customer base growth zero among continuing businesses, firm-level innovation rates fall by approximately 40% relative to the full model. Of this 40% decline, only about 6 percentage points are attributable to the direct firm-level market size effect alone; the vast majority is driven by the endogenous feedback loop between innovation and customer acquisition. In a second decomposition focused on aggregate growth, the firm-level market size effect and a reallocation effect — whereby the feedback loop concentrates customers among high-productivity firms — together account for 44% of aggregate growth in the full model.&lt;/p&gt;
&lt;p&gt;On policy, the authors compare R&amp;amp;D subsidies and operational subsidies in the full model against an otherwise identical model that ignores customer accumulation. R&amp;amp;D subsidies are approximately twice as effective at boosting aggregate growth in the full model as in the model without customer accumulation. Conversely, operational subsidies produce a stronger decline in aggregate growth in the full model than in the benchmark-without-customer-accumulation, because aggregate growth in the full model is a customer-weighted average of firms&amp;rsquo; productivity growth rates, making the joint distribution of productivity and customer bases the relevant object of study.&lt;/p&gt;
&lt;p&gt;Firm-level data support three empirical predictions. Marketing expenditure, R&amp;amp;D intensity, and markups co-move in model-consistent directions both contemporaneously and over the life cycle. The estimated relative weight of marketing versus pricing as channels of customer accumulation is γ = 0.745, indicating marketing is the dominant channel. A model-consistent proxy for the severity of customer-base frictions, estimated in the cross-section of industries, shows that stronger frictions correlate with lower R&amp;amp;D investment, as predicted. The customer-base depreciation rate is estimated at ζ = 0.375, R&amp;amp;D cost scaling at σx = 1.264, and marketing cost scaling at σa = 1.405.&lt;/p&gt;
&lt;p&gt;Q: What is the firm-level market size effect and why does it arise?
A: When a firm retains more customers, successful innovations apply to a larger market, raising the profitability of each unit reduction in production costs. This increases the marginal benefit of R&amp;amp;D investment. In the benchmark model, Proposition 2(a) shows formally that firm-level innovation increases with customer base growth: ∂x/∂(1−ζ) &amp;gt; 0, where ζ is the customer separation rate.&lt;/p&gt;
&lt;p&gt;Q: What is the endogenous feedback loop between innovation and customer accumulation?
A: More productive firms have lower production costs and can therefore afford greater investment in marketing and can set lower markups, both of which attract more customers. A larger customer base raises firm value and strengthens R&amp;amp;D incentives further (Proposition 2(b)). This bidirectional feedback means that productivity growth and customer accumulation are jointly determined in equilibrium, not independent processes.&lt;/p&gt;
&lt;p&gt;Q: How large is the quantitative effect of customer accumulation on firm-level innovation?
A: In the counterfactual where expected customer retention is reduced so that average customer base growth among continuing firms is zero, firm-level innovation rates are approximately 40% lower than in the full model. Of this, only about 6% (of the total drop) is attributable to the direct market size effect in isolation; the feedback loop accounts for the remaining roughly 34 percentage points.&lt;/p&gt;
&lt;p&gt;Q: How much of aggregate growth do customer-acquisition channels explain?
A: The firm-level market size effect and a customer reallocation effect together account for 44% of aggregate growth in the full model. The firm-level market size effect alone reduces aggregate growth by about one-fifth (20%) in the relevant counterfactual. The reallocation effect — by which productive firms accumulate disproportionate market share — contributes the remainder of the 44%.&lt;/p&gt;
&lt;p&gt;Q: What is the reallocation channel for aggregate growth?
A: Because highly productive firms can invest more in customer acquisition, the feedback loop endogenously concentrates customers (market shares) among high-productivity firms. Since aggregate productivity in the model is a customer-weighted average of firm productivity levels (equation 16), this reallocation raises aggregate productivity growth beyond what the firm-level R&amp;amp;D incentive effect alone would produce.&lt;/p&gt;
&lt;p&gt;Q: How does customer accumulation change the efficacy of R&amp;amp;D subsidies?
A: R&amp;amp;D subsidies are approximately twice as effective at raising aggregate growth in the full model (with customer accumulation) as in an otherwise identical model that ignores customer accumulation. The mechanism is Proposition 4(b): faster customer base growth makes firms weight future profits more heavily, increasing their sensitivity to any change in R&amp;amp;D costs, including that brought about by a government subsidy.&lt;/p&gt;
&lt;p&gt;Q: What happens to aggregate growth under operational subsidies in the two models?
A: Operational subsidies lead to a stronger decline in aggregate growth in the full model than in the model without customer accumulation. The reason is that aggregate growth in the full model depends on the joint distribution of firm productivity and customer bases; operational subsidies alter this distribution in ways that reduce the customer-weighted average of productivity growth rates, an effect absent when customer accumulation is ignored.&lt;/p&gt;
&lt;p&gt;Q: How are the two customer-acquisition channels (marketing and pricing) measured empirically?
A: Marketing is measured using sales-and-marketing expenses from Capital IQ, available for 48% of the Compustat sample (34% report directly; an additional 14% report advertising or marketing sub-components). Markups are measured following De Loecker et al. (2020) as the inverse share of variable costs in sales multiplied by the cost-output elasticity, with variation across firms identified from balance sheet data under the assumption that cost-output elasticities are constant within industry-year cells.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated relative strength of marketing versus pricing in customer accumulation?
A: The relative weight on marketing is γ = 0.745, estimated by targeting the coefficient βµ = 0.04 (standard error 0.01) from a reduced-form regression of firm-level sales growth on changes in markups (equation 29). This implies that marketing is the dominant channel, consistent with evidence in Afrouzi et al. (2021) and Fitzgerald et al. (forthcoming).&lt;/p&gt;
&lt;p&gt;Q: What is the estimated customer-base depreciation rate and how is it disciplined?
A: The depreciation rate ζ is estimated at 0.375, targeted to match average firm-level employment growth from the BDS. This falls toward the lower end of existing estimates, which range from about 0.3 to 0.7 across studies.&lt;/p&gt;
&lt;p&gt;Q: How do R&amp;amp;D costs scale with firm size in the estimated model?
A: The R&amp;amp;D cost scaling parameter is σx = 1.264, estimated by targeting the reduced-form coefficient of −0.01 from a regression of log R&amp;amp;D intensity on log sales with industry-time fixed effects (equation 28). This is close to the estimate in Akcigit and Kerr (2018).&lt;/p&gt;
&lt;p&gt;Q: How do marketing costs scale with firm size?
A: The marketing cost scaling parameter is σa = 1.405, estimated by targeting a reduced-form coefficient of −0.01 from a regression of log sales-and-marketing intensity on log sales with industry-time fixed effects (equation 30).&lt;/p&gt;
&lt;p&gt;Q: What empirical co-movement evidence supports the model&amp;rsquo;s predictions?
A: In the cross-section of firms, marketing expenditure, R&amp;amp;D intensity, and markups all co-move in model-predicted directions, for both static (contemporaneous) relationships and dynamic (life-cycle) patterns. Additionally, a model-consistent industry-level proxy for the severity of customer-base frictions shows that stronger frictions are associated with lower R&amp;amp;D investment, as the model predicts.&lt;/p&gt;
&lt;p&gt;Q: How does endogenous firm exit work in the full model and why does it differ from standard models?
A: Firms pay a stochastic per-period operational cost and exit when that cost exceeds a threshold κ*_j = v(q_j, b_j)/W. Unlike standard growth models where exit depends only on productivity, here the exit threshold depends on both productivity and accumulated customers, so customer loss can trigger exit even for relatively productive firms.&lt;/p&gt;
&lt;p&gt;Q: What data sources are used and what are their key limitations?
A: The three primary firm-level sources are the Census Bureau&amp;rsquo;s BDS (broad coverage, employment-focused), Compustat (rich financial data but limited to publicly traded firms and lacking direct customer-acquisition measures), and Capital IQ (sales-and-marketing expenses available from 1997, matched to 91% of the Compustat sample). To address Compustat&amp;rsquo;s non-representativeness, employment-based weights aligning Compustat and BDS firm-size distributions are applied when computing model moments against Compustat targets.&lt;/p&gt;
&lt;p&gt;Firm-level market size effect: The mechanism by which higher customer retention raises a firm&amp;rsquo;s future profit base — because lower production costs from successful innovation apply to a larger market — thereby strengthening incentives to conduct R&amp;amp;D. This is the primary channel linking customer accumulation to innovation.&lt;/p&gt;
&lt;p&gt;Customer base (b_j): The mass of household members consuming a firm&amp;rsquo;s product variety, which varies endogenously across firms. It enters demand directly (equation 4) and serves as a state variable in the firm&amp;rsquo;s value function alongside productivity.&lt;/p&gt;
&lt;p&gt;Endogenous feedback loop: The bidirectional reinforcement between productivity growth and customer accumulation. More productive firms invest more in customers; a larger customer base raises the value of innovation; higher innovation raises productivity further.&lt;/p&gt;
&lt;p&gt;Reallocation effect: The concentration of customers (market shares) toward high-productivity firms that arises endogenously from the feedback loop, contributing to aggregate growth because aggregate productivity is a customer-weighted average of firm-level productivity.&lt;/p&gt;
&lt;p&gt;Customer-base depreciation rate (ζ): The exogenous rate at which a firm loses its existing customers each period, estimated at 0.375 in the paper&amp;rsquo;s calibration. It governs the baseline speed of customer attrition and is the key parameter for the firm-level market size effect.&lt;/p&gt;
&lt;p&gt;Sales-and-marketing expenses: Expenditures on sales force, brand development, customer service, advertising, and customer data acquisition — measured from Capital IQ — that directly drive marketing-based customer accumulation (the dominant channel with estimated weight γ = 0.745).&lt;/p&gt;
&lt;p&gt;Perfect scaling (Assumption 1): The benchmark restriction that R&amp;amp;D and marketing costs, and the sales-driven customer accumulation benefit, all scale one-for-one with a composite of firm productivity and customer base. This assumption enables closed-form solutions and is relaxed in the full model using estimated scaling parameters.&lt;/p&gt;</description></item><item><title>Distributional Growth Accounting: Education and the Reduction of Global Poverty, 1980–2019</title><link>https://macropaperwarehouse.com/papers/distributional-growth-accounting-education-and-the-reduction-of-global-poverty-19802019/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/distributional-growth-accounting-education-and-the-reduction-of-global-poverty-19802019/</guid><description>&lt;h2 id="layer-1--core-argument"&gt;Layer 1 — Core Argument&lt;/h2&gt;
&lt;p&gt;This paper constructs the first estimates of the aggregate and distributional effects of worldwide educational expansion since 1980 by developing a &amp;ldquo;distributional growth accounting&amp;rdquo; framework that isolates the contribution of schooling to economic growth by income group. The framework integrates the canonical labor supply-and-demand model of education and the wage structure (à la Goldin and Katz 2007) with standard growth accounting tools, applied to a new microdatabase covering household surveys in 150 countries and representative of approximately 95% of the world&amp;rsquo;s population, alongside new country-specific estimates of private returns to primary, secondary, and tertiary schooling. Under conservative assumptions — relying on standard Mincerian returns, assuming capital income is unaffected by schooling, and abstracting from human capital externalities — education can account for approximately 50% of global economic growth, 70% of income gains among the world&amp;rsquo;s poorest 20% of individuals, and 40% of extreme poverty reduction since 1980; it also explains over 50% of improvements in the share of labor income accruing to women. A key mechanism is imperfect substitutability between skill groups: as educational expansion raises the supply of skilled workers, their relative wage falls, redistributing income toward low-skilled workers and amplifying education&amp;rsquo;s equalizing effect at the bottom of the distribution — a channel that canonical cross-country growth accounting misses, causing it to underestimate education&amp;rsquo;s contribution to poverty reduction by a factor of approximately three. Combining these indirect investment benefits from education with direct government redistribution (from a companion paper) brings the total contribution of public policies to extreme poverty reduction to at least 50%.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-q-what-is-distributional-growth-accounting-and-how-does-it-differ-from-standard-growth-accounting"&gt;Q1. Q: What is distributional growth accounting and how does it differ from standard growth accounting?&lt;/h3&gt;
&lt;p&gt;A: Standard growth accounting (as in Barro and Lee 2015) combines cross-country data on average years of schooling with a uniform return to derive a counterfactual average income absent educational progress. Distributional growth accounting instead starts from microdata on the joint distribution of income and education within 150 countries, constructs income-group-specific counterfactuals, and accounts for both direct wage effects on individuals whose education changed and general equilibrium supply effects that alter relative wages across all workers. The standard approach is found to underestimate education&amp;rsquo;s contribution to the poorest 20%&amp;rsquo;s income growth by a factor of roughly three (23% vs. 71% in the benchmark specification), because cross-country averages cannot accurately locate the world&amp;rsquo;s poorest individuals and because two key channels — labor income shares being greater at the bottom, and supply-side wage redistribution — are omitted.&lt;/p&gt;
&lt;h3 id="q2-q-how-is-the-counterfactual-world-income-distribution-constructed"&gt;Q2. Q: How is the counterfactual world income distribution constructed?&lt;/h3&gt;
&lt;p&gt;A: In five steps applied to the 150-country microdata. First, education levels are downgraded within each survey until matching the 1980 distribution of educational attainment (using the Barro–Lee database), prioritizing individuals closest to the target level. Second, the earnings of downgraded workers are reduced using the &amp;ldquo;true&amp;rdquo; return to schooling, which lies between the initial return (prevailing before expansion, computed from the CES production function using the 2019 elasticity) and the final return observed in 2019 — for plausible parameterizations, the true return weights initial returns at 50–70%. Third, relative wages are adjusted to reflect supply effects: the increase in skilled-worker supply lowers their relative wage by 1/σ log points per log-point increase in relative supply. Fourth, counterfactual labor income is combined with unchanged capital income to yield counterfactual total income. Fifth, the share of actual income growth attributable to education is computed as the gap between the actual and counterfactual growth rates, expressed as a fraction of actual growth.&lt;/p&gt;
&lt;h3 id="q3-q-what-role-does-imperfect-skill-substitution-play-and-how-is-σ-calibrated"&gt;Q3. Q: What role does imperfect skill substitution play, and how is σ calibrated?&lt;/h3&gt;
&lt;p&gt;A: Imperfect substitution between skill groups (elasticity σ in a CES production function) is the mechanism through which educational expansion redistributes income. When skilled-worker supply rises, their relative wage falls and low-skilled workers&amp;rsquo; relative wage rises, so the income gains from education are shared more broadly than individual returns alone would suggest. With perfect substitutes (σ → ∞), supply effects vanish and education&amp;rsquo;s distributional impact is determined entirely by who directly received schooling. The elasticity is calibrated from the recent macroeconomics literature; in sensitivity analysis, the paper bounds the contribution of education to the poorest 20%&amp;rsquo;s income growth between 60% and 90% across plausible values of σ and private returns.&lt;/p&gt;
&lt;h3 id="q4-q-why-are-the-estimates-described-as-conservative"&gt;Q4. Q: Why are the estimates described as conservative?&lt;/h3&gt;
&lt;p&gt;A: Three reasons, each biasing the estimates downward. First, standard Mincerian returns are used, which are systematically lower than causal estimates from natural experiments — a meta-analysis of 15 papers and the paper&amp;rsquo;s own quasi-experimental validation (India, Indonesia, United States) confirm this; if anything, the framework underestimates schooling&amp;rsquo;s benefits in those settings. Second, capital income is assumed unaffected by schooling, abstracting from potential effects on capital accumulation and returns. Third, human capital externalities — for which there is now substantial empirical evidence — are ignored entirely. These conservative choices are deliberate; relaxing them would increase all headline estimates.&lt;/p&gt;
&lt;h3 id="q5-q-how-does-skill-biased-technical-change-interact-with-the-education-contribution"&gt;Q5. Q: How does skill-biased technical change interact with the education contribution?&lt;/h3&gt;
&lt;p&gt;A: In the CES model, the return to schooling is increasing in the skill bias of technology (AH/AL): a higher skill bias raises the marginal product of skilled workers relative to unskilled, making schooling more profitable. The benchmark counterfactual holds technology fixed at its 2019 value and reduces education to its 1980 level. An alternative counterfactual would hold technology at its 1980 value and increase education to its 2019 level; the difference between these two exercises identifies the contribution of skill-biased technical change in amplifying the benefits of schooling. Because 1980 microdata on the world income distribution are unavailable, this decomposition can only be performed for the subsample of 33 countries with surveys around 2000; for that sample, skill-biased technical change accounts for 20–30% of the income benefits of schooling, meaning education would still have yielded large gains even absent technological progress.&lt;/p&gt;
&lt;h3 id="q6-q-what-do-the-quasi-experimental-validations-in-india-indonesia-and-the-united-states-show"&gt;Q6. Q: What do the quasi-experimental validations in India, Indonesia, and the United States show?&lt;/h3&gt;
&lt;p&gt;A: Three large-scale schooling policy interventions — a school construction program in India (studied in Khanna 2023), Indonesia&amp;rsquo;s INPRES program (Duflo 2001 and 2004), and US compulsory schooling laws (Acemoglu and Angrist 2000) — are used to externally validate the framework. Using regional variation in exposure to each program and rich microdata on the income distribution, the paper documents two findings: (1) educational expansion had large causal effects on aggregate regional incomes comparable in magnitude to individual returns estimated in the same contexts; and (2) all three policies disproportionately benefited low-income earners, substantially reducing inequality. The distributional growth accounting framework reproduces both findings with &amp;ldquo;a remarkable degree of accuracy,&amp;rdquo; and if anything underestimates the benefits of schooling, providing validation of the methodological foundation.&lt;/p&gt;
&lt;h3 id="q7-q-how-does-the-paper-quantify-educations-role-in-gender-inequality-reduction"&gt;Q7. Q: How does the paper quantify education&amp;rsquo;s role in gender inequality reduction?&lt;/h3&gt;
&lt;p&gt;A: The framework is extended to gender by constructing a counterfactual for how large gender labor income gaps would be absent educational improvement since the early 1990s (the period for which female labor income share data are available). The counterfactual accounts for three gender-specific channels: differential educational expansion between men and women, heterogeneous returns to schooling by gender, and differential effects of schooling on female labor force participation. Comparing the counterfactual to actual trends in female labor income shares, education can explain 50–80% of the observed reductions in gender inequality, depending on specification and world region.&lt;/p&gt;
&lt;h3 id="q8-q-how-do-public-policies-as-a-whole-contribute-to-extreme-poverty-reduction"&gt;Q8. Q: How do public policies as a whole contribute to extreme poverty reduction?&lt;/h3&gt;
&lt;p&gt;A: The paper&amp;rsquo;s estimate of education&amp;rsquo;s indirect investment benefits (40% of extreme poverty reduction) is combined with a companion paper&amp;rsquo;s (Gethin 2023) estimates of direct government redistribution — cash and in-kind transfers together accounting for approximately 30% of global poverty reduction since 1980, with in-kind transfers alone accounting for approximately 20%. Because the two contributions overlap (e.g., public education spending is both an indirect investment benefit and an in-kind transfer), the combined lower bound is reported as &amp;ldquo;at least 50%&amp;rdquo; of extreme poverty reduction attributable to public policies.&lt;/p&gt;
&lt;h3 id="q9-q-why-does-the-distributional-approach-yield-such-different-results-from-the-standard-approach-for-the-poorest-20"&gt;Q9. Q: Why does the distributional approach yield such different results from the standard approach for the poorest 20%?&lt;/h3&gt;
&lt;p&gt;A: Two main reasons. First, cross-country data cannot accurately measure the incomes of the world&amp;rsquo;s poorest, because the poorest individuals are not all concentrated in the poorest countries — distributional accounting within countries is necessary to locate them precisely. Second, the standard approach misses two progressive channels: (a) labor income shares are higher at the bottom of the income distribution than average, so gains from schooling translate into larger income increases for the poor; and (b) supply effects redistribute schooling gains from high-skilled to low-skilled workers, a mechanism that is entirely absent from cross-country averages but directly captured in the microdata-based counterfactual.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Distributional growth accounting:&lt;/strong&gt; A framework, introduced in this paper, that combines a model of education and the wage structure with household microdata to construct income-group-specific counterfactuals, isolating the contribution of human capital accumulation to growth at each point of the income distribution rather than at the national-average level.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;True return to schooling (r&lt;/em&gt;):&lt;/em&gt;* In the CES framework with imperfect skill substitution, the &amp;ldquo;true&amp;rdquo; aggregate return to schooling used in the counterfactual lies strictly between the initial return (prevailing before educational expansion, counterfactually higher because skilled-worker supply was lower) and the final return (observed after expansion, lower due to skill-supply pressure). The true return is the return that equates the model&amp;rsquo;s predicted output loss to the actual output loss from reducing education; for plausible parameters it weights initial returns at 50–70%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Supply effects (general equilibrium effects of schooling):&lt;/strong&gt; When the supply of skilled workers rises, their relative wage falls and the relative wage of unskilled workers rises. These wage adjustments are not captured by individual-level Mincerian returns but are modeled via the CES elasticity of substitution σ. Supply effects are central to education&amp;rsquo;s progressive distributional impact: they compress the skill premium and raise earnings at the bottom of the distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Imperfect substitution between skill groups:&lt;/strong&gt; The CES production specification in which skilled (H) and unskilled (L) labor are combined with elasticity σ &amp;lt; ∞. This governs the magnitude of general equilibrium wage effects: a lower σ means a larger wage compression per unit of skilled-supply increase, amplifying the redistributive role of education. The paper calibrates σ from the macroeconomics literature and bounds results over plausible ranges.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Skill-biased technical change (SBTC):&lt;/strong&gt; Technology that raises the marginal product of skilled workers relative to unskilled (captured by the ratio AH/AL in the CES production function). SBTC amplifies returns to schooling; in the subsample of 33 countries with around-2000 surveys, SBTC accounts for 20–30% of schooling&amp;rsquo;s income benefits, but education would still have generated substantial income gains absent SBTC.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conservative assumptions (scope condition):&lt;/strong&gt; All headline quantitative results (50% of aggregate growth, 70% of poorest-20% income gains, 40% of extreme poverty reduction, &amp;gt;50% of gender inequality reduction) are explicitly conditioned on conservative assumptions: Mincerian rather than causal returns, no effect on capital income, and no human capital externalities. The paper argues these assumptions bias all estimates downward.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Summary based on HAL working paper (halshs-04423765v1, Working Paper 2023/25, November 2023). Period covered in working paper text: 1980–2022. AI-assisted, human review pending.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>Enlightenment Ideals and Belief in Progress in the Run-up to the Industrial Revolution</title><link>https://macropaperwarehouse.com/papers/enlightenment-ideals-and-belief-in-progress-in-the-run-up-to-the-industrial-revolution/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/enlightenment-ideals-and-belief-in-progress-in-the-run-up-to-the-industrial-revolution/</guid><description>&lt;p&gt;This paper tests Joel Mokyr&amp;rsquo;s claim that Britain&amp;rsquo;s industrialization was preceded and enabled by a cultural shift — specifically, that Enlightenment ideals produced a &amp;ldquo;progress-oriented&amp;rdquo; view of science that diffused to artisans and craftsmen. The central research question is whether and when the language of science became more progress-oriented in the build-up to the Industrial Revolution, and whether this shift was concentrated in volumes directly linked to industrial production.&lt;/p&gt;
&lt;p&gt;The authors assemble 173,031 unique volumes printed in England and written in English between 1500 and 1900, drawn from the Hathitrust Digital Library. Because copyright law prohibits downloading full text, they use HDL&amp;rsquo;s Extracted-Features &amp;ldquo;bag of words&amp;rdquo; dataset. After removing duplicates and Latin-language volumes from an initial set of 420,081, they apply Latent Dirichlet Allocation (LDA) with cross-validated perplexity minimization to identify an optimal T=60 topics. Topic-pair co-occurrence analysis identifies three categories — science, religion, and political economy — each anchored by three defining topics. Volume-level category weights are derived by multiplying each topic&amp;rsquo;s weight by its category coefficient. The resulting classification yields 50,090 science volumes, 102,565 political economy volumes, and 14,124 religion volumes.&lt;/p&gt;
&lt;p&gt;Progressive sentiment is measured using a seven-word dictionary (progress, improvement, stride, betterment, advance, rise, amelioration) assembled from thesaurus synonyms for &amp;ldquo;progress,&amp;rdquo; manually vetted by all four authors, and restricted to words attested in the Oxford English Dictionary before 1643 (Newton&amp;rsquo;s birth year). Sentiment for each volume equals the count of progress-dictionary words divided by total word count. An analogous optimism-sentiment placebo dictionary is constructed separately.&lt;/p&gt;
&lt;p&gt;Industrial relevance is scored using the digitized indexes of all five volumes of Appleby&amp;rsquo;s Illustrated Handbook of Machinery (1877–1903); the top industrial root words are crane (weight 51), electr (42), weight (37), rope (27), and cost (27). Each volume receives an industry score equal to the weighted occurrence of industrial root words normalized by volume length.&lt;/p&gt;
&lt;p&gt;Three main findings emerge. First, the language of science and religion showed little overlap beginning in the 17th century — that is, the secularization of science predates the onset of industrialization. Science volumes shifted from approximately 40 percent religious content around 1700 to only about 10 percent by 1850, with scientific content rising correspondingly from roughly 40 percent to over 60 percent. This trend was stable from 1650 through 1900.&lt;/p&gt;
&lt;p&gt;Second, while scientific volumes became more progress-oriented during the Enlightenment, this progressive shift was concentrated in volumes at the nexus of science and political economy. Volumes of &amp;ldquo;pure&amp;rdquo; science were largely neutral with respect to progress sentiment, and those at the science-religion nexus had on average negative progress sentiment. The marginal effect of scientific content on progress sentiment was greatest for volumes mixing science and political economy, and most of the increase in predicted sentiment at that nexus occurred during the 18th century, remaining stable thereafter. A placebo test using optimism sentiment finds the opposite pattern: volumes at the science-political economy nexus were among the least optimistic, while the most optimistic language appeared at the religion-political economy nexus. This rules out the interpretation that the measured shift reflects a general increase in positive affect rather than specifically progress-oriented language.&lt;/p&gt;
&lt;p&gt;Third, volumes employing industrial terminology that also sat at the science-political economy nexus were distinctively progressive beginning in the mid-18th century. At the 90th percentile of industry score, predicted progress sentiment at the science-political economy nexus was positive throughout the sample; at zero industry score, it was negative until the mid-18th century. Volumes at the religion-political economy nexus showed modestly positive and time-stable progress sentiment regardless of industry score.&lt;/p&gt;
&lt;p&gt;The paper concludes that it was the pragmatic, applied volumes — those bridging science and political economy, written for artisans and a broader literate public rather than for the human-capital elite alone — that embodied the cultural values Mokyr identifies as central to Britain&amp;rsquo;s industrialization.&lt;/p&gt;
&lt;p&gt;Q: What gap in the existing literature does this paper address?&lt;/p&gt;
&lt;p&gt;A: Prior work on the cultural deep roots of economic growth rarely tracks how culture changes over time, relying instead on cross-sectional variation or qualitative case studies. Quantitative evidence that the language of science itself became more progress-oriented — and that this change reached beyond elite thinkers to artisans and craftsmen — had not been marshaled before. The paper provides inaugural quantitative support by analyzing 173,031 volumes spanning four centuries.&lt;/p&gt;
&lt;p&gt;Q: Why does the paper restrict the progress-sentiment dictionary to words attested before 1643?&lt;/p&gt;
&lt;p&gt;A: Words that entered English only after 1643 (Newton&amp;rsquo;s birth year) could not have appeared in volumes from the early Enlightenment, so including them would bias sentiment scores toward the later part of the sample. The restriction ensures the dictionary is applicable and unbiased across the full 1500–1900 period. The final retained words are: progress, improvement, stride, betterment, advance, rise, amelioration.&lt;/p&gt;
&lt;p&gt;Q: How does LDA classify volumes, and how is T=60 selected?&lt;/p&gt;
&lt;p&gt;A: LDA treats each volume as a bag of words and derives a Dirichlet distribution such that observed documents are generated by repeated topic sampling. The number of topics T is selected by minimizing perplexity on held-out data via 4-fold cross-validation, rotating training and test sets across folds; this procedure yields T=60 as optimal. Each volume is then represented as a mixture over those 60 topics.&lt;/p&gt;
&lt;p&gt;Q: What are the three categories and their anchor topics?&lt;/p&gt;
&lt;p&gt;A: Political Economy is anchored by topics on law/public opinion, governance/parliament, and trade/price/labour. Religion is anchored by topics on church/Christian doctrine, God/faith/sin, and virtue/fame/religion. Science is anchored by topics on engineering/steam/electricity, chemistry/acid/heat, and geometry/equations/trigonometry. These three sets of topics were selected for high corpus-wide importance and mutual independence.&lt;/p&gt;
&lt;p&gt;Q: What does the finding on science-religion separation imply for timing?&lt;/p&gt;
&lt;p&gt;A: The separation of scientific and religious language was already visible by 1600 and firmly established by the mid-17th century, well before the Industrial Revolution conventionally dated to the mid-18th century. This supports Mokyr&amp;rsquo;s argument that the secularization of science was an Enlightenment-era precursor to industrialization rather than a product of it. The trend remained stable from 1650 through 1900.&lt;/p&gt;
&lt;p&gt;Q: How does the progressive sentiment differ between pure science and the science-political economy nexus?&lt;/p&gt;
&lt;p&gt;A: Volumes of pure science were largely neutral with respect to progress-oriented language and in some periods showed slightly negative predicted progress sentiment. The science-religion nexus showed consistently negative progress sentiment. By contrast, volumes at the science-political economy nexus showed the highest level of progressive sentiment beginning in the mid-18th century, and most of this growth in predicted sentiment occurred during the 18th century, after which it remained stable.&lt;/p&gt;
&lt;p&gt;Q: What does the placebo optimism test show?&lt;/p&gt;
&lt;p&gt;A: The optimism sentiment scores are nearly the mirror opposite of the progress scores: the most optimistic language appears at the religion-political economy nexus, while volumes at the science-political economy nexus are among the least optimistic. This dissociation rules out the interpretation that the measured progress-sentiment rise reflects a general shift toward positive language rather than a specific cultural embrace of science as a tool for improving human welfare.&lt;/p&gt;
&lt;p&gt;Q: How is the industrial score constructed and what are the most heavily weighted terms?&lt;/p&gt;
&lt;p&gt;A: The authors digitized the detailed indexes of all five volumes of Appleby&amp;rsquo;s Illustrated Handbook of Machinery (1877–1903), restricted to words attested before 1643, and weighted each industrial root word by its index frequency. Each corpus volume&amp;rsquo;s industry score equals the sum of (word count × index weight) across all industrial words, normalized by volume length, yielding a score between 0 and 1. The top-weighted terms are crane (51), electr (42), weight (37), rope (27), and cost (27).&lt;/p&gt;
&lt;p&gt;Q: What is the key result linking industrial scores to progressive sentiment?&lt;/p&gt;
&lt;p&gt;A: At the science-political economy nexus, volumes with industry scores at the 90th percentile had persistently positive predicted progress sentiment throughout the sample, while volumes at that nexus with zero industry score had negative predicted sentiment until the mid-18th century. The shift to positive sentiment for high-industry volumes at this nexus occurred in the mid-18th century — roughly coinciding with the onset of Britain&amp;rsquo;s industrialization — and those volumes remained the most progress-oriented in the corpus thereafter.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s interpretation of the science-political economy nexus finding in relation to Mokyr?&lt;/p&gt;
&lt;p&gt;A: The authors interpret volumes at the science-political economy nexus as pragmatic, applied works aimed at a broader literate audience including artisans and craftsmen, not exclusively the human-capital elite. These are precisely the volumes Mokyr&amp;rsquo;s &amp;ldquo;Industrial Enlightenment&amp;rdquo; thesis predicts would carry progress-oriented cultural values into the mechanical and artisanal pursuits that drove industrialization. The finding that pure-science volumes were not especially progressive, while applied volumes bridging science and political economy were, is consistent with Mokyr&amp;rsquo;s argument that it was the diffusion of Enlightenment ideals to skilled practitioners — not just to elite scientists — that mattered.&lt;/p&gt;
&lt;p&gt;Q: What qualitative examples support the quantitative findings?&lt;/p&gt;
&lt;p&gt;A: Martin Clare&amp;rsquo;s The Motion of Fluids (1735) explicitly addresses &amp;ldquo;the Unlearned&amp;rdquo; and states in its preface that the work is meant to be &amp;ldquo;of singular Use and Benefit to Mankind&amp;rdquo; — a direct expression of the progress-oriented language the algorithm detects. George Stephenson&amp;rsquo;s 1831 railway report argues that rail infrastructure would allow Ireland to &amp;ldquo;reciprocate with England and with other nations, the products of industry,&amp;rdquo; exemplifying how progress-oriented language pervaded industrial writing by the early 19th century. These examples confirm that the high progress-sentiment scores for industrial volumes at the science-political economy nexus reflect genuine rhetorical content, not measurement artifacts.&lt;/p&gt;
&lt;p&gt;Q: What are the paper&amp;rsquo;s limitations regarding early sample periods?&lt;/p&gt;
&lt;p&gt;A: The corpus is thin in earlier eras, particularly around 1550, so results from the earliest decades must be interpreted with caution. The HDL data derive from digitized scans with OCR output of very old books, introducing errors such as the &amp;ldquo;long-S&amp;rdquo; misread (e.g., &amp;ldquo;juftice&amp;rdquo; for &amp;ldquo;justice&amp;rdquo;) that require manual correction. Additionally, the bag-of-words model discards word order, which may obscure some semantic distinctions.&lt;/p&gt;
&lt;p&gt;Q: What future research directions do the authors identify?&lt;/p&gt;
&lt;p&gt;A: The authors propose applying the same textual analysis techniques to test whether English-language volumes began reflecting greater freedom of expression in the run-up to Britain&amp;rsquo;s economic takeoff, connecting to the literature on European political fragmentation and the marketplace of ideas. They also suggest applying the approach to corpora in other languages — Dutch (following McCloskey&amp;rsquo;s argument about bourgeois values) and Spanish (to examine whether the Counter-Reformation and Spain&amp;rsquo;s economic lag are reflected in cultural attitudes toward progress and science).&lt;/p&gt;
&lt;p&gt;LDA (Latent Dirichlet Allocation): An unsupervised generative statistical model that treats each document as a bag of words and extracts latent topics as multinomial distributions over vocabulary; used here to reduce 173,031 volumes to mixtures of 60 topics without imposing prior scholarly interpretations.&lt;/p&gt;
&lt;p&gt;Progressive Sentiment Score: The fraction of words in a volume belonging to a seven-word dictionary of progress synonyms (progress, improvement, stride, betterment, advance, rise, amelioration), normalized by total word count; measures the cultural orientation toward the betterment of humankind as embedded in text.&lt;/p&gt;
&lt;p&gt;Industrial Score: A volume-level measure equal to the weighted count of industrial root words — derived from the indexes of Appleby&amp;rsquo;s Illustrated Handbook of Machinery (1877–1903) — normalized by volume length; captures the degree to which a volume&amp;rsquo;s vocabulary overlaps with industrial production terminology.&lt;/p&gt;
&lt;p&gt;Science-Political Economy Nexus: The region of the topic simplex where volumes carry substantial weight in both the science and political economy categories but low weight in religion; the paper finds this is where progress-oriented language was most concentrated from the mid-18th century onward, interpreted as applied science aimed at artisans and a broader literate public.&lt;/p&gt;
&lt;p&gt;Industrial Enlightenment: Joel Mokyr&amp;rsquo;s (2009) concept describing the diffusion of Enlightenment ideals about the practical utility of science into the mechanical and artisanal pursuits that drove Britain&amp;rsquo;s industrialization; the paper provides quantitative support for this thesis by showing that industrial volumes at the science-political economy nexus were distinctively progress-oriented.&lt;/p&gt;
&lt;p&gt;Culture of Growth: Mokyr&amp;rsquo;s (2016) broader argument that a pan-European network of elite intellectuals fostered a progress-oriented view of science — the idea that scientific understanding could improve the human condition — and that this cultural norm, in combination with Britain&amp;rsquo;s stock of skilled craftsmen, made industrialization possible.&lt;/p&gt;
&lt;p&gt;Bag of Words: A representation of text that records only word frequencies within a document, discarding word order; used here both because HDL copyright restrictions prevent full-text download and because it is the input format required by LDA.&lt;/p&gt;</description></item><item><title>Growth Experiences and Trust in Government</title><link>https://macropaperwarehouse.com/papers/growth-experiences-and-trust-in-government/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/growth-experiences-and-trust-in-government/</guid><description>&lt;p&gt;This paper investigates whether individuals who have experienced stronger GDP growth over their lifetimes are more likely to trust their national government. The authors — Besley, Dann, and Dray — assemble a newly harmonized global dataset comprising approximately 3.3 million respondents across 166 countries since 1990, drawn from 11 major opinion surveys (Afrobarometer, Americasbarometer, Arabarometer, Asiabarometer, European Social Survey, Gallup World Poll, Integrated Values Survey, Latinobarometer, Life in Transition Survey, South Asia Barometer, and World Justice Project). They supplement this with longer-run U.S. evidence from the American National Election Studies (ANES) going back to 1958, covering respondents born as early as the 1880s, and longitudinal Swiss evidence from the Swiss Household Panel (SHP) which allows individual fixed-effects estimation.&lt;/p&gt;
&lt;p&gt;The core methodological contribution is the exploitation of country-cohort variation in lifetime GDP growth experiences. Following Malmendier and Nagel (2011), the authors construct a weighted average of past growth realizations across an individual&amp;rsquo;s lifetime, with weights decaying linearly over time (lambda = 1), so that more recent growth receives greater weight. The baseline specification includes country fixed effects, cohort-by-subcontinent fixed effects, survey-by-survey-year fixed effects, controls for log GDP per capita at year of birth, and individual characteristics (sex, marital status, education, religious denomination). More demanding specifications add country-by-survey-year and country-by-age fixed effects. For Switzerland, individual fixed effects are included, fully absorbing time-invariant personal characteristics.&lt;/p&gt;
&lt;p&gt;The main finding is that a one standard deviation increase in lifetime GDP growth experience — corresponding to approximately 2 percentage points of additional growth — is associated with a 2.1 percentage point increase in the probability of trusting the national government, significant at the 1 percent level. This corresponds to roughly 0.042 standard deviations of the trust outcome and approximately 5 percent of the global mean trust in government. The effect is quantitatively meaningful: it approximates between one-quarter and one-half of the difference in average trust between older and younger cohorts in India and Italy, respectively. For the U.S. ANES sample, a one standard deviation increase in growth experience (about 0.2 percentage points) increases trust in the federal government by 2.4 percentage points, explaining more than two-thirds of the average trust gap between Baby Boomers (born 1946–1964) and Millennials (born 1981–1996).&lt;/p&gt;
&lt;p&gt;Several scope conditions and heterogeneity findings sharpen the interpretation. First, the growth-trust link is specific to government institutions: there is no statistically significant effect of growth experience on interpersonal trust or trust in religious organizations, indicating the channel runs through perceptions of state performance rather than generalized social capital. Second, a recency heuristic operates: the linearly decaying weighting function (lambda = 1) outperforms both an unweighted lifetime average (lambda = 0) and a formative-years weighting. Growth experienced during formative years (ages 18–25) or before birth has no detectable effect on trust in government; the pre-birth result serves as a placebo test. Third, the positive growth-trust relationship is stronger in democracies than in autocracies, which the authors interpret as democracies producing citizens more responsive to government performance signals. Fourth, a &amp;ldquo;trust paradox&amp;rdquo; emerges: unconditionally, average trust in government is lower in democracies than in autocracies, and longer democratic experience is associated with lower trust, which the authors attribute to democratic institutions generating greater citizen skepticism about government performance. Fifth, core results are robust to controlling for other lifetime politico-economic experiences including inflation, banking and currency crises, epidemics, political unrest, executive turnover, stock market returns, and income inequality. The Swiss evidence further shows that private income growth experience does not drive the result — only aggregate macroeconomic growth does.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s core quantitative finding on the growth-trust relationship?
A: Using the global harmonized dataset of 3.3 million respondents across 166 countries, a one standard deviation increase in lifetime GDP growth experience (corresponding to approximately 2 percentage points of additional growth) is associated with a 2.1 percentage point increase in the probability of trusting the national government, significant at the 1 percent level. Using only the Gallup World Poll subsample (roughly half the observations), the estimated effect is somewhat larger at 3.6 percentage points per standard deviation increase. These estimates remain statistically significant under more demanding specifications with country-by-survey-year and country-by-age fixed effects, though the magnitudes decrease as these interacted fixed effects absorb variation in recent growth experiences.&lt;/p&gt;
&lt;p&gt;Q: How do the authors measure individual lifetime growth experience?
A: The growth experience variable is a weighted average of all past annual GDP per capita growth rates since an individual&amp;rsquo;s birth, with weights that decay linearly over time (lambda = 1 in the Malmendier-Nagel framework). Under this parameterization, the measure simplifies to how much recent economic performance (in the year prior to the survey) exceeds the long-run mean over the respondent&amp;rsquo;s lifetime, scaled by the respondent&amp;rsquo;s midpoint of life. This implies younger individuals are more sensitive to recent growth outcomes because their shorter life histories give recent events relatively greater weight. The authors validate this lambda = 1 choice via a grid search over alternative weighting structures using minimum residual sum of squares as the criterion.&lt;/p&gt;
&lt;p&gt;Q: How is reverse causality addressed?
A: The empirical strategy identifies the relationship using past, cumulative growth experiences measured prior to the survey, so current trust in government cannot cause past growth. Survey-year fixed effects absorb all aggregate time trends simultaneously affecting trust and growth. The authors also conduct a placebo test showing that GDP growth occurring before an individual&amp;rsquo;s birth has a precisely estimated null effect on their trust in government, which would not be the case if unobserved societal trends were jointly driving both growth histories and political perceptions.&lt;/p&gt;
&lt;p&gt;Q: Does growth experience affect interpersonal trust or trust in non-state institutions?
A: No. The estimated coefficient on lifetime growth experience is statistically insignificant at conventional levels when interpersonal trust replaces trust in government as the dependent variable, with narrow confidence intervals indicating a precisely estimated null. Similarly, growth experience has no systematic effect on trust in religious organizations such as churches or mosques. The authors interpret these null results as evidence against the alternative explanation that broad modernizing social changes are jointly driving both growth experiences and political trust.&lt;/p&gt;
&lt;p&gt;Q: What do the U.S. ANES results add?
A: The ANES data, which extends back to 1958 and captures cohorts born as early as the 1880s, provide a within-country test controlling for state fixed effects, generation dummies, and rich individual characteristics including partisan affiliation and partisan strength. A one standard deviation increase in U.S. growth experience (approximately 0.2 percentage points) raises trust in the federal government by 2.4 percentage points, significant at the 1 percent level. This estimate is quantitatively large enough to explain more than two-thirds of the average trust gap between Baby Boomers and Millennials. Results are robust to adding state-by-survey-year fixed effects and birth-state-by-generation fixed effects, and hold for a broader &amp;ldquo;trust in government index&amp;rdquo; covering beliefs about waste, corruption, and responsiveness of the federal government.&lt;/p&gt;
&lt;p&gt;Q: What do the Swiss Household Panel results contribute?
A: The SHP allows individual fixed-effects estimation, exploiting within-person changes in growth experience and trust over time from 1999 onward, which absorbs all time-invariant individual characteristics that could confound the global and U.S. cross-cohort results. The growth experience coefficient remains positive and significant, with a one standard deviation increase yielding a 1.9 percentage point increase in trust in the Swiss federal government (significant at the 1 percent level). The Swiss data also uniquely allow the authors to test whether personal income growth experience drives the result; they find no significant effect of private income growth experience on trust in government, only aggregate macroeconomic growth matters.&lt;/p&gt;
&lt;p&gt;Q: Does the recency heuristic hold — does growth in formative years matter?
A: No. The authors find no detectable effect of growth experienced specifically during formative years (ages 18–25) on trust in government. Additionally, in a grid-search exercise assessing model fit across different lambda values, the linearly decaying weighting scheme (lambda = 1, giving more weight to recent growth) outperforms both equal-weighted lifetime averages (lambda = 0) and weighting schemes that emphasize earlier life experiences (lambda less than 0). The pre-birth placebo result (null effect) and the absence of a formative-years effect together indicate that the operative mechanism is about evaluating current government performance based on recent macroeconomic experience, not the imprinting of long-lasting political dispositions during youth.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;trust paradox&amp;rdquo; and how is it documented?
A: The trust paradox refers to the empirical finding that average trust in government is lower in democracies than in autocracies at the cross-country level, and that longer experience with democratic institutions within countries is associated with lower levels of trust in government in the micro data. This is counterintuitive given the standard view that good institutions should foster confidence in government. The authors suggest the paradox likely reflects democracies cultivating greater citizen skepticism and more critical judgment of government performance, rather than indicating that democratic governance actually performs worse. Importantly, the positive effect of growth experience on trust remains present in democracies, and the growth-trust relationship is actually stronger in democratic regimes, consistent with citizens in democracies being more responsive to government performance signals.&lt;/p&gt;
&lt;p&gt;Q: How is the growth-trust finding related to corruption perceptions and living standards?
A: Using the Gallup World Poll, the authors find that stronger lifetime growth experience is associated with lower perceived corruption in government, greater satisfaction with personal living standards, and higher likelihood of feeling one lives comfortably on one&amp;rsquo;s present income. These results are consistent with citizens attributing economic success to government competence and integrity, and with growth translating into perceptions of improved personal circumstances through both direct income effects and indirect public goods provision.&lt;/p&gt;
&lt;p&gt;Q: Are the results robust to controlling for other lifetime politico-economic experiences?
A: Yes. When the authors include lifetime experience measures for political unrest, executive turnover, epidemic exposure, banking crises, currency crises, and inflation (both levels and volatility) simultaneously in equation (3), the growth experience coefficient remains consistently positive, stable, and significant across all specifications. Among the other experience variables, only lifetime unrest and epidemic exposure are independently negative and statistically significant at conventional levels. F-tests reject the null hypothesis that the crisis and growth experience coefficients are equal in magnitude. The U.S. results are also robust to adding lifetime experiences with S&amp;amp;P 500 returns, unemployment, and top-income-share inequality measures.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: The authors note that sustained economic growth may itself be a mechanism for building political trust, with positive downstream effects for policy compliance — a connection they document has been relevant during the COVID-19 pandemic (where higher-trust societies showed lower mobility during lockdowns and higher vaccine acceptance). The growth-trust channel could have implications for increasing compliance across a range of policy domains including climate action and tax morale. Governments that deliver sustained economic growth can expect citizens to update their trust upward, particularly in democracies where citizens are more performance-responsive, while governments that preside over stagnation or contraction face predictable erosion of political legitimacy across cohorts.&lt;/p&gt;
&lt;p&gt;Growth experience: A weighted average of all past annual GDP per capita growth realizations since an individual&amp;rsquo;s birth, with weights that decay linearly over time following Malmendier and Nagel (2011), so that more recent growth receives greater weight. Under the paper&amp;rsquo;s preferred parameterization (lambda = 1), the measure equals how much last year&amp;rsquo;s GDP per capita exceeds the respondent&amp;rsquo;s lifetime mean, scaled by the respondent&amp;rsquo;s midpoint of life.&lt;/p&gt;
&lt;p&gt;Trust in government: A binary dummy variable equal to one if a survey respondent expresses &amp;ldquo;a great deal&amp;rdquo; or &amp;ldquo;quite a lot&amp;rdquo; of trust or confidence in the national government, constructed from harmonized responses across 11 major opinion surveys. The paper treats this as reflecting respondents&amp;rsquo; perceptions of government performance rather than a deep interpersonal trust relationship.&lt;/p&gt;
&lt;p&gt;Trust paradox: The empirical regularity documented in the paper whereby average trust in government is unconditionally lower in democracies than in autocracies at the cross-country level, and whereby longer democratic experience within countries is associated with lower individual trust in government. The authors attribute this to democratic institutions generating more critical citizen judgment of government performance.&lt;/p&gt;
&lt;p&gt;Recency heuristic: The finding that more recent growth experiences carry greater weight in forming trust in government, as captured by the linear decay weighting scheme (lambda = 1) outperforming equal-weighted or early-life-weighted alternatives. Growth before birth and growth during formative years (ages 18–25) have no detectable effect, while recent macroeconomic performance is the operative signal.&lt;/p&gt;
&lt;p&gt;Cohort-level variation: The within-country differences in lifetime growth experiences across birth cohorts that form the paper&amp;rsquo;s primary identification strategy. Because different cohorts in the same country have lived through different sequences of growth episodes, differences in trust across cohorts within a country can be attributed to differential growth exposure rather than time-invariant country characteristics.&lt;/p&gt;
&lt;p&gt;Formative years effect: The hypothesis, tested and rejected in the paper, that economic experiences during ages 18–25 have a lasting imprint on political attitudes analogous to formative-years effects found in other political behavior literatures. The paper finds no statistically significant association between growth experienced during these years and trust in government.&lt;/p&gt;
&lt;p&gt;Source text origin: In the pipeline context relevant to this paper&amp;rsquo;s acquisition, this refers to whether a summary was generated from full working paper text (&amp;ldquo;pdf&amp;rdquo; or &amp;ldquo;oa-html&amp;rdquo;) versus abstract only (which is hard-blocked). The working paper was obtained from LSE Research Online (eprint 129614), classified as published version under CC BY 4.0.&lt;/p&gt;</description></item><item><title>Heterogeneous innovations and growth under imperfect technology spillovers</title><link>https://macropaperwarehouse.com/papers/heterogeneous-innovations-and-growth-under-imperfect-technology-spillovers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/heterogeneous-innovations-and-growth-under-imperfect-technology-spillovers/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Jo and Kim ask two related questions: (1) How do firms use different types of innovation when learning others&amp;rsquo; technology takes time? (2) How does this process alter the aggregate implications of firm innovation, particularly in the context of increasing competition?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The paper develops a discrete-time infinite-horizon endogenous growth model with multi-product firms pursuing two types of innovation — &amp;ldquo;own-innovation&amp;rdquo; (improving existing product quality) and &amp;ldquo;creative destruction&amp;rdquo; (entering new product markets by displacing incumbents) — subject to a novel friction called &amp;ldquo;imperfect technology spillovers.&amp;rdquo; The friction takes the specific form of lagged learning: creative destruction builds on the one-period-lagged technology of the target market&amp;rsquo;s incumbent, while only the incumbent can observe the current frontier technology level. This one-period lag creates a technology gap (Δ = q_t / q_{t−1}) between the incumbent&amp;rsquo;s frontier and the level available to rivals. Four possible technology gap values arise in equilibrium: Δ₁ = 1 (no gap), Δ₂ = λ (one successful own-innovation), Δ₃ = η (one successful creative destruction), and Δ₄ = η/λ. The step sizes satisfy λ² &amp;gt; η &amp;gt; λ, meaning a single creative destruction improves quality more than a single own-innovation, but two consecutive own-innovations dominate a single creative destruction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Key Mechanisms.&lt;/strong&gt; The learning friction generates two novel mechanisms. First, the &amp;ldquo;market-protection effect&amp;rdquo;: incumbents with a technology advantage (Δ &amp;gt; 1) intensify own-innovation to widen the gap and protect their product lines when competitive pressure rises. Formally, own-innovation probability is highest for Δ₂ products and declines monotonically (z₂ &amp;gt; z₃ &amp;gt; z₄ &amp;gt; z₁), and ∂z₂/∂x &amp;gt; ∂z₃/∂x &amp;gt; 0 while ∂z₁/∂x &amp;lt; 0, conditional on value coefficients. Second, the &amp;ldquo;technological barrier effect&amp;rdquo;: higher overall own-innovation and creative destruction intensity widens the average technology gap across products, reducing rivals&amp;rsquo; conditional probability of successfully taking over a product market. This is distinct from the standard Schumpeterian effect (lower expected future profits) and from the escape-competition effect in step-by-step models (which apply only to neck-and-neck, single-product firms).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Empirical Strategy.&lt;/strong&gt; The empirical analysis combines the USPTO PatentsView database, the Longitudinal Business Database (LBD), the Longitudinal Firm Trade Transactions Database (LFTTD), the Census of Manufactures (CMF), Compustat, and NBER-CES data, covering the universe of U.S. patenting firms from 1976 to 2016, with main analyses from 1982 to 2007. Own-innovation is proxied by the self-citation ratio of patents (the ratio of self-citations to total backward citations); creative destruction by new products added and low-self-citation patents. Exogenous competitive pressure comes from China&amp;rsquo;s WTO accession in 2001, instrumented by the industry-level NTR tariff gap (the gap between non-NTR and NTR rates in 1999) following Pierce and Schott (2016).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Findings.&lt;/strong&gt; Pre-shock (1982–1999): patents with lower self-citation ratios (closer to creative destruction) have significantly longer backward citation gaps (coefficient −2.29 to −2.59, p &amp;lt; 0.01 across specifications), confirming that learning others&amp;rsquo; technology takes more time. Creative-destruction-type patents also have higher market value (Kogan et al. stock return measure) and scientific value (forward citations), with self-citation ratio negatively associated with both (e.g., coefficient on self-citation for market value: −0.289 without firm FE; −0.110 with firm FE, p &amp;lt; 0.01). Conditional on patenting, higher self-citation ratios are negatively associated with employment growth (coefficient −0.256, p &amp;lt; 0.05), number of industries added (−0.158, p &amp;lt; 0.05), and products added (−0.274, p &amp;lt; 0.01).&lt;/p&gt;
&lt;p&gt;Post-shock (DID): foreign competition had no statistically significant effect on overall patent counts, but firms with above-average innovation intensity in industries with high NTR gaps significantly increased their self-citation ratio — indicating a shift toward own-innovation. The triple-interaction coefficient is 0.795 (p &amp;lt; 0.01) with baseline controls. For a firm with average lagged innovation intensity (0.18) in an industry with an average NTR gap (0.291), this corresponds to a 4.2 percentage point increase in the seven-year growth rate of the self-citation ratio, representing a 15.0% increase relative to the average growth rate of 28.2 percentage points. Consistent with the technological barrier effect, firm entry rates are lower in industries with higher TFPR-skewness-based technological barriers (coefficient −0.012 to −0.016, p &amp;lt; 0.05).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Analysis.&lt;/strong&gt; Calibrated to the U.S. manufacturing sector in 1992, the model matches six target moments including average number of products (2.3), products added (0.3), firm entry rate (7.6%), average productivity growth (1.9%), high-growth-firm employment growth (22.5%), and import penetration (15.3%). Creative destruction contributes approximately 1.88 times more to growth per unit than own-innovation (step size ratio 0.075/0.04). The aggregate R&amp;amp;D-to-sales ratio (untargeted) is 4.6% in the model vs. 4.1% in data.&lt;/p&gt;
&lt;p&gt;A counterfactual increasing outside entrants by 83% (matching the rise in import penetration from 15.3% to 25.1% between 1992 and 2007) generates a 1.51% increase in aggregate creative destruction arrival rate x, but firm-level creative destruction probability falls 1.33% and startup creative destruction also falls 1.33%. The aggregate R&amp;amp;D-to-sales ratio falls 1.6% and creative destruction R&amp;amp;D intensity falls 1.2%. Average domestic productivity growth declines 11.0%, with growth from creative destruction falling 13.0% and growth from domestic startups falling 1.7%. The total mass of domestic firms falls 6.4%.&lt;/p&gt;
&lt;p&gt;In economies with creative destruction costs 80 times higher than the U.S. baseline, the same competitive pressure shock raises rather than lowers total R&amp;amp;D (by 1.0%), but domestic growth still falls 9.7%, because the marginal decline in creative destruction impedes the growth contribution and firm entry even when aggregate innovation spending rises.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-key-friction-that-distinguishes-this-model-from-the-existing-multi-product-firm-literature-eg-klette-and-kortum-2004-akcigit-and-kerr-2018"&gt;Q1. What is the key friction that distinguishes this model from the existing multi-product firm literature (e.g., Klette and Kortum 2004; Akcigit and Kerr 2018)?&lt;/h3&gt;
&lt;p&gt;A: The key friction is &amp;ldquo;imperfect technology spillovers,&amp;rdquo; modeled as lagged learning: creative destruction can only build on the one-period-lagged technology of the target product (q_{j,t−1}), while the product&amp;rsquo;s current owner observes the frontier technology (q_{j,t}). In models without this friction — such as Akcigit and Kerr (2018) — rivals can instantly learn and copy frontier technology, so firms have no technological advantage and cannot protect their markets. In the current model, own-innovation by the incumbent widens the gap between q_{j,t} and q_{j,t−1}, creating a barrier that a rival must overcome even after successful creative destruction. This makes own-innovation an endogenous function of the technology gap, a feature absent from existing multi-product firm frameworks.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-model-predict-that-own-innovation-increases-with-the-technology-gap-up-to-a-point-then-decreases"&gt;Q2. Why does the model predict that own-innovation increases with the technology gap up to a point, then decreases?&lt;/h3&gt;
&lt;p&gt;A: From Corollary 1, the ordering z₂ &amp;gt; z₃ &amp;gt; z₄ &amp;gt; z₁ reflects competing forces. Products with gap Δ₂ = λ gain the most from additional own-innovation in terms of reducing the probability of losing the product line (equation 2), so own-innovation is highest there. Products with Δ₃ = η or Δ₄ = η/λ already have substantial technological advantages from prior creative destruction, so the marginal value of own-innovation in reducing market loss probability is lower. Products with Δ₁ = 1 have no advantage at all: if a rival succeeds in creative destruction, the incumbent loses the product regardless of own-innovation (equation 1), so z₁ is lowest. Beyond a certain gap level, the incumbent is sufficiently protected that additional own-innovation has diminishing returns in deterrence.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-market-protection-effect-formally-and-for-which-products-is-it-strongest"&gt;Q3. What is the market-protection effect formally, and for which products is it strongest?&lt;/h3&gt;
&lt;p&gt;A: The market-protection effect (Corollary 2) is the positive response of a firm&amp;rsquo;s own-innovation to an increase in the aggregate creative destruction arrival rate x, conditional on the value coefficients A₁ and A₂ being fixed. It is strongest for products with Δ₂ = λ (∂z₂/∂x is the largest and positive), positive but weaker for Δ₃ = η (∂z₃/∂x &amp;gt; 0), of ambiguous sign for Δ₄ = η/λ, and negative for Δ₁ = 1 (∂z₁/∂x &amp;lt; 0). The asymmetry reflects the asymmetric payoff to own-innovation across gap levels: for Δ₂ products, successful own-innovation can turn a losing situation into a winning one because it shifts the technology gap from Δ₁ to Δ₂ from the rival&amp;rsquo;s perspective, effectively defeating the rival&amp;rsquo;s creative destruction attempt. This mechanism provides a micro-foundation for why frontier firms (like Google or NVIDIA) keep innovating intensely despite their technological leads, a pattern the standard step-by-step model cannot explain.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-technological-barrier-effect-and-how-does-it-differ-from-the-schumpeterian-effect"&gt;Q4. What is the technological barrier effect and how does it differ from the Schumpeterian effect?&lt;/h3&gt;
&lt;p&gt;A: The technological barrier effect refers to the reduction in rivals&amp;rsquo; incentive for creative destruction caused by an increase in the average technology gap across product lines. When incumbents do more own-innovation or when outside firms do more creative destruction, the distribution of technology gaps shifts rightward (density at Δ₁ falls; density at Δ₂, Δ₃, Δ₄ rises). This raises the average technology barrier rivals must overcome to successfully take over a product market, reducing the conditional takeover probability x^{takeover} and the expected value of creative destruction B. In the U.S. counterfactual, the technological barrier effect accounts for 17.0% of the total change in the aggregate creative destruction rate x and 15.0% of the change in startup creative destruction x_e. In contrast, the Schumpeterian effect refers to the reduction in expected future profits from owning a product due to increased displacement risk (through the value coefficient A₂), a mechanism present in standard quality-ladder models. Both operate simultaneously but the technological barrier effect is a novel feature of this framework.&lt;/p&gt;
&lt;h3 id="q5-how-is-own-innovation-vs-creative-destruction-measured-empirically-and-what-validates-this-measure"&gt;Q5. How is own-innovation vs. creative destruction measured empirically, and what validates this measure?&lt;/h3&gt;
&lt;p&gt;A: The self-citation ratio (the share of a patent&amp;rsquo;s backward citations that cite the same assignee&amp;rsquo;s earlier patents) is used as the primary measure: a higher ratio indicates greater reliance on the firm&amp;rsquo;s own prior knowledge, hence a higher probability that the innovation improves an existing product line (own-innovation). This is validated empirically in three ways. First, patents with lower self-citation ratios have significantly larger backward citation gaps (coefficient −2.29 to −2.59 across fixed-effect specifications on 728,721 observations), consistent with creative destruction requiring more time to learn others&amp;rsquo; technology. Second, lower self-citation patents have higher market value and scientific value (forward citations), consistent with η &amp;gt; λ (creative destruction contributes more per event to quality). Third, firm-level regressions show that lower self-citation ratios are associated with higher employment growth, more products added, and more industries entered, consistent with creative destruction contributing more to firm expansion.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-did-identification-strategy-work-and-what-are-the-main-results"&gt;Q6. How does the DID identification strategy work, and what are the main results?&lt;/h3&gt;
&lt;p&gt;A: The identification exploits the removal of trade policy uncertainty (TPU) after China&amp;rsquo;s WTO accession in 2001. The treatment variable is the industry-level NTR gap (the gap between non-NTR and NTR tariff rates in 1999): industries with larger gaps experienced a larger reduction in uncertainty and thus a greater increase in Chinese import competition. The DID compares patenting firms across periods (1992–1999 vs. 2000–2007) and across high- vs. low-NTR-gap industries, with a triple interaction for firm-level innovation intensity (lagged five-year average patents per employee, normalized within two-digit NAICS). The main finding (Table 4): the NTR gap × Post interaction has no significant effect on overall patent counts (coefficient 0.238 without controls, standard error 0.237), but the triple interaction (NTR gap × Post × innovation intensity) has a positive and significant effect on the growth rate of the self-citation ratio (0.732 without controls, p &amp;lt; 0.05; 0.795 with baseline controls, p &amp;lt; 0.01). This implies that innovation-intensive firms in high-competition industries shifted their composition toward own-innovation, while overall patenting was unchanged — consistent with an offsetting rise in own-innovation and fall in creative destruction.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-aggregate-growth-effects-of-increasing-competitive-pressure-in-the-calibrated-model"&gt;Q7. What are the aggregate growth effects of increasing competitive pressure in the calibrated model?&lt;/h3&gt;
&lt;p&gt;A: Using an 83% increase in outside entrants (matching the 1992–2007 rise in import penetration from 15.3% to 25.1%), average domestic productivity growth falls 11.0%. Decomposing: growth from domestic own-innovation falls 11.4%, growth from domestic creative destruction falls 13.0%, and growth from domestic startups falls 1.7% (Table 9). The aggregate R&amp;amp;D-to-sales ratio falls 1.6% and the creative destruction R&amp;amp;D intensity falls 1.2%, indicating that the decline in creative destruction R&amp;amp;D outweighs the rise in own-innovation R&amp;amp;D. The total mass of domestic firms falls 6.4% and the average number of products per firm falls 5.5%.&lt;/p&gt;
&lt;h3 id="q8-how-do-results-differ-in-economies-with-high-creative-destruction-costs-vs-the-us"&gt;Q8. How do results differ in economies with high creative destruction costs vs. the U.S.?&lt;/h3&gt;
&lt;p&gt;A: When creative destruction costs (χ̃) are set 80 times higher than the U.S. baseline, the initial equilibrium has much lower creative destruction: R&amp;amp;D-to-sales ratio is 1.39% (vs. 4.58% in U.S.), creative destruction R&amp;amp;D intensity is 8.6% (vs. 63.9%), average number of products is 1.0 (vs. 2.3), and average domestic productivity growth is 1.4% (vs. 1.9%). Under the same competition shock, total R&amp;amp;D actually rises by 1.0% in this high-CD-cost economy (because own-innovation increases more than creative destruction falls, given the already low baseline of creative destruction), in contrast to the −1.6% in the U.S. However, domestic growth still falls 9.7% even in this economy, driven by reductions in creative destruction by incumbents and startups combined with a decline in the mass of domestic incumbents. This result holds even with a fixed firm mass (Table E5), confirming the mechanism is not solely due to entry/exit dynamics.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-technological-barrier-effects-quantitative-contribution-to-the-decline-in-creative-destruction"&gt;Q9. What is the technological barrier effect&amp;rsquo;s quantitative contribution to the decline in creative destruction?&lt;/h3&gt;
&lt;p&gt;A: In the U.S. counterfactual (Table 8 and associated decomposition), 17.0% of the total change in the aggregate creative destruction arrival rate x and 15.0% of the total change in startup creative destruction x_e are attributable specifically to the technological barrier effect — that is, to the shift in the technology gap distribution µ(Δℓ) holding all else equal. The conditional takeover probability x^{takeover} declines from 73.2% to 73.0%. The density at Δ₁ (the easiest gap to overcome) falls 0.4%, while densities at Δ₃ and Δ₄ rise 1.1% and 1.4% respectively, driven by increased creative destruction by outside firms and intensified own-innovation by incumbents.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-the-paper-draws-from-its-framework"&gt;Q10. What are the policy implications the paper draws from its framework?&lt;/h3&gt;
&lt;p&gt;A: The paper argues that policies evaluating innovation should account for composition, not just aggregate R&amp;amp;D levels or patent counts. Increased overall innovation driven by defensive own-innovation contributes less to economic growth than creative destruction and restricts firm entry — so it is less beneficial than it appears. In low-creativity economies (e.g., European economies with high regulatory barriers to creative destruction), increased foreign competition may raise aggregate R&amp;amp;D while still lowering domestic growth, misleading policymakers who track only total innovation spending. The model also suggests that the mixed empirical findings in the competition-innovation literature (Aghion et al. 2005; Bloom et al. 2016; Autor et al. 2020) can be reconciled by accounting for compositional shifts: the net effect of competition on total innovation is ambiguous because it raises own-innovation for technologically advantaged firms while reducing creative destruction for all firms.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Imperfect Technology Spillovers:&lt;/strong&gt; The novel friction introduced in this paper, modeled as lagged learning: firms attempting creative destruction can only access the one-period-lagged technology of the target product market (q_{j,t−1}), while the incumbent product owner observes and can improve from the current frontier (q_{j,t}). This asymmetry creates a persistent technological advantage for incumbents and enables strategic defensive innovation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Own-Innovation:&lt;/strong&gt; R&amp;amp;D investment by a firm to improve the quality of its existing product lines. Successful own-innovation raises product quality by a step size λ &amp;gt; 1. Own-innovation does not require learning others&amp;rsquo; technology and, in the model, constitutes the incumbents&amp;rsquo; defensive margin against creative destruction. At the aggregate level, it contributes more to total growth than creative destruction because it succeeds more frequently, but per successful event it contributes less (λ &amp;lt; η).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Creative Destruction:&lt;/strong&gt; R&amp;amp;D investment enabling a firm to enter a new product market by displacing the incumbent. Successful creative destruction improves the lagged quality of the target product by a step size η &amp;gt; λ, where λ² &amp;gt; η &amp;gt; λ. It requires learning the incumbent&amp;rsquo;s one-period-lagged technology, takes longer to develop (evidenced empirically by longer backward citation gaps), and contributes more to firm growth and product expansion per event than own-innovation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Technology Gap (Δ):&lt;/strong&gt; The ratio of a product&amp;rsquo;s current-period technology to its previous-period technology (Δ_{j,t} = q_{j,t}/q_{j,t−1}). This gap summarizes the technological advantage the incumbent holds in a product market under imperfect spillovers. Four values are possible in equilibrium: Δ₁ = 1, Δ₂ = λ, Δ₃ = η, Δ₄ = η/λ. The gap determines both the incumbent&amp;rsquo;s own-innovation incentive and the rival&amp;rsquo;s probability of successfully completing a product takeover conditional on creative destruction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Market-Protection Effect:&lt;/strong&gt; The mechanism by which incumbents with a technological advantage (Δ &amp;gt; 1) increase own-innovation in response to heightened competitive pressure (an increase in the aggregate creative destruction arrival rate x). This effect is maximized for products with Δ₂ = λ and positive but diminishing for Δ₃. It is absent for Δ₁ = 1 products (where own-innovation cannot prevent displacement) and is formally distinct from the escape-competition effect in step-by-step innovation models, which applies only to neck-and-neck single-product firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Technological Barrier Effect:&lt;/strong&gt; The reduction in rivals&amp;rsquo; incentive for creative destruction caused by an increase in the average technology gap across the economy&amp;rsquo;s product lines. When incumbents intensify own-innovation and/or when outside creative destruction increases, the distribution of technology gaps shifts toward higher Δ values, reducing the conditional probability that a rival successfully takes over any given product market. This feedback mechanism endogenously suppresses creative destruction and firm entry beyond what the Schumpeterian effect alone would predict.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Self-Citation Ratio:&lt;/strong&gt; The share of a patent&amp;rsquo;s backward citations that cite patents previously owned by the same firm. Used in the paper as a continuous proxy for the likelihood that a patent represents own-innovation vs. creative destruction: a ratio of 1 (100% self-citations) implies 100% probability of own-innovation; a ratio of 0 implies 100% probability of creative destruction. This measure follows Akcigit and Kerr (2018) and is validated in the paper against learning time, quality, and firm growth outcomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;NTR Gap (Trade Policy Uncertainty Shock):&lt;/strong&gt; The industry-level difference between non-NTR (column 2) and NTR (column 1) U.S. tariff rates in 1999, used as an instrument for the exogenous increase in Chinese competitive pressure following China&amp;rsquo;s WTO accession and the U.S. granting of Permanent Normal Trade Relations (PNTR) in 2002. Industries with larger NTR gaps experienced a greater reduction in trade policy uncertainty and thus a larger increase in competitive pressure from foreign firms.&lt;/p&gt;</description></item><item><title>Intergenerational Impacts of Secondary Education: Experimental Evidence from Ghana</title><link>https://macropaperwarehouse.com/papers/intergenerational-impacts-of-secondary-education-experimental-evidence-from-ghana/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/intergenerational-impacts-of-secondary-education-experimental-evidence-from-ghana/</guid><description>&lt;p&gt;This paper provides experimental evidence on the intergenerational impacts of secondary education subsidies in a low-income context, leveraging a randomized controlled trial (RCT) conducted in rural Ghana with a 15-year longitudinal follow-up. The study exploits a 2008 scholarship lottery in which 682 students — drawn from 2,064 rural youth who had been admitted to public senior high school but had not enrolled due to financial constraints — were randomly selected to receive four-year secondary school scholarships covering full tuition and fees. Scholarship receipt increased senior high school completion by 27–28 percentage points for both men and women (from 39.8% to 67.2% for women; from 49.7% to 77.9% for men), and raised average years of education by 1.33 years.&lt;/p&gt;
&lt;p&gt;The central research question is whether secondary education subsidies generate intergenerational benefits — specifically, whether children of scholarship recipients have better survival and cognitive development outcomes — and what mechanisms drive any such effects.&lt;/p&gt;
&lt;p&gt;For female scholarship recipients, the scholarship significantly altered fertility timing and partnership. By 2013, female recipients were 6.9 percentage points less likely to have ever been pregnant (on a control-group base of 48.3%), with the decline driven almost entirely by a 7 percentage point (17%) reduction in unwanted pregnancies. Though total fertility eventually caught up by 2022, recipients were still less likely to be married or cohabiting as of 2019 and were significantly more likely to have a partner with tertiary education.&lt;/p&gt;
&lt;p&gt;Children of female scholarship recipients experienced substantially lower mortality. Among control-group female respondents, 3.5% of children died before age one and 4.0% before age three. These rates fell to 1.7% (p=0.028) and 2.2% (p=0.065) respectively among children of female recipients — a roughly 45–51% reduction in under-one and under-three mortality.&lt;/p&gt;
&lt;p&gt;Child cognitive development gains emerge only once children reach school age. Children of female recipients show no significant cognitive score differences at 18 months, 2.5 years, or 3.5 years, but score 0.238 standard deviations higher at age five (p=0.005) and 0.252 standard deviations higher at age seven (p=0.035). Effects span language, math and numeracy, spatial reasoning, and executive function, but not socio-cognitive development. These effect sizes fall between the 75th and 80th percentile of RCT-based educational intervention effect sizes in low- and middle-income countries.&lt;/p&gt;
&lt;p&gt;The primary mechanism is not higher income or greater monetary investment in children. The study finds no significant treatment effect on household SES index (0.107 SDs, p=0.103), no impact on formal schooling inputs, and no difference in parental aspirations or knowledge of child stimulation&amp;rsquo;s importance. Instead, more-educated mothers seek more prenatal care, engage in more preventive health behaviors, and — critically — spend more time interacting with their children in stimulating ways. Day-long LENA (Language Environment Analysis) recordings at 18 months confirm 20% more adult-child conversational turns per minute (effect size 0.068, p=0.005) and 17% more child vocalizations per minute (effect size 0.32, p=0.014) for children of female recipients.&lt;/p&gt;
&lt;p&gt;For male scholarship recipients, no analogous intergenerational benefits appear. Their partners are not more educated (in fact slightly less educated on tertiary rates), their children show no mortality improvement, and cognitive scores are if anything negative at age five (point estimate -0.22, p=0.069). The absence of effects is attributed to male scholarship recipients having caregivers — overwhelmingly mothers — with no more education than in the control group, and to children of male recipients being 8.7 percentage points less likely to live with their father.&lt;/p&gt;
&lt;p&gt;A cost-benefit analysis finds internal rates of return (IRR) of 27%–76% for a female-only means-tested scholarship program and 20%–51% for a mixed-gender program. The cost per under-three death averted ($15,184 for female-only) places the scholarship program within the range of the 10th-percentile most cost-effective WHO-recommended child health interventions.&lt;/p&gt;
&lt;p&gt;Scope conditions: the study estimates effects for students who qualified for senior high school but faced binding financial constraints in rural Ghana in 2008 — a population that is well-prepared academically but economically disadvantaged. Results may not generalize to students who would not have qualified for secondary school or to contexts where financial barriers are not binding.&lt;/p&gt;
&lt;p&gt;Q: What was the experimental design and who was in the study sample?
A: In 2008, 2,064 rural Ghanaian students who had been admitted to senior high school (SHS) but had not enrolled — typically due to inability to pay fees — were sampled. After a baseline survey, 682 were randomly selected (approximately one-third) by lottery to receive a four-year scholarship covering full tuition and fees for a day (non-boarding) student, stratified by district, school, gender, and exam-year cohort. The two-thirds comparison group received no scholarship. Students were on average 17 years old at baseline and just over 31 at the last follow-up in Spring 2023.&lt;/p&gt;
&lt;p&gt;Q: How large was the scholarship&amp;rsquo;s effect on educational attainment?
A: Scholarship receipt raised SHS completion from 39.8% to 67.2% among women (a 69% increase) and from 49.7% to 77.9% among men (a 57% increase). Overall, the scholarship led to an average of 1.33 more years of education. For women only, it also significantly raised tertiary education: by 2023, scholarship receipt increased tertiary completion by 10.8 percentage points for women, but had no significant tertiary effect for men.&lt;/p&gt;
&lt;p&gt;Q: What were the effects on fertility and family formation for female scholarship recipients?
A: By 2013, female recipients were 6.9 percentage points less likely to have ever been pregnant (base: 48.3% in control), driven almost entirely by a 7 percentage point (17%) reduction in unwanted pregnancies. By 2019, recipients were still 6 percentage points less likely to have started childbearing and had 0.152 fewer children on average (p=0.065). Total fertility eventually caught up by 2022. By 2016, female recipients were 12.1 percentage points (24% of control mean) less likely to have ever lived with a partner, and by 2019 were 6.2 percentage points less likely to be married or cohabiting. Conditional on having a partner, they were significantly more likely to have a partner who completed tertiary education (p=0.071).&lt;/p&gt;
&lt;p&gt;Q: What were the effects on fertility and family formation for male scholarship recipients?
A: Male recipients showed few changes in fertility or marriage behavior. They were 7.8 percentage points (30% of control mean) more likely to still be living with their parents as of 2019. Their partners were not more educated; in the cognitive games subsample, treatment actually reduced the share of partners with tertiary education by 3.6 percentage points from a control base of 4.3%.&lt;/p&gt;
&lt;p&gt;Q: What were the child mortality results for children of female scholarship recipients?
A: Among children of female control respondents, 3.5% died before age one and 4.0% before age three. These fell to 1.7% (p=0.028) and 2.2% (p=0.065), respectively, among children of female recipients — approximately a halving of under-one and under-three mortality. These point estimates are robust to varying the covariates (linear vs. fixed effects for birth year, dropping or adding controls). After multiple-hypothesis testing adjustment using the Romano-Wolf step-down procedure, the p-value for survived-to-one rises from 0.028 to 0.119.&lt;/p&gt;
&lt;p&gt;Q: What were the child mortality results for children of male scholarship recipients?
A: The estimated effects for children of male recipients were smaller and statistically insignificant: a 1.4 percentage point increase in survived-to-one (p=0.161) and 0.9 percentage points in survived-to-three (p=0.549). These estimates are not significantly different from those for female recipients. Results were sensitive to sample perturbations given the smaller sample: only 26 of 1,016 children of male respondents died before age one.&lt;/p&gt;
&lt;p&gt;Q: What child cognitive development gains did children of female scholarship recipients show, and at what ages?
A: No significant differences emerged at 18 months (-0.066 SDs, p=0.489), 2.5 years (-0.024 SDs, p=0.850), or 3.5 years (0.026 SDs, p=0.736). Significant gains appeared at age five (0.238 SDs, p=0.005) and age seven (0.252 SDs, p=0.035). Effects span language (0.15 SDs at five; 0.27 SDs at seven), math and numeracy (0.15 SDs; 0.26 SDs), spatial reasoning (0.20 SDs; 0.12 SDs), and executive function (0.25 SDs; 0.20 SDs), but not socio-cognitive development. These effect sizes fall between the 75th and 80th percentile of educational RCT effect sizes in low- and middle-income countries.&lt;/p&gt;
&lt;p&gt;Q: What cognitive development effects did children of male scholarship recipients show?
A: No significant positive effects emerged at any age. Point estimates were negative at all ages except 18 months, and marginally significantly negative at age five (-0.22 SDs, p=0.069). The difference in treatment effects between children of male and female recipients is statistically significant at age five (p=0.005).&lt;/p&gt;
&lt;p&gt;Q: Why do cognitive gains appear only at age five and not earlier?
A: The authors offer three interpretations: first, that the cognitive tests for younger children are noisier instruments (cross-sectional and longitudinal correlations within domains are much lower for 1.5-year tests than 5-year tests); second, that impacts on cognitive development may take time to materialize; third, that marginal survivors in the treatment group may start with a cognitive deficit (e.g., surviving a cerebral malaria episode), and maternal education effects require time to overcome this initial handicap. Gains concentrate on skills underlying literacy and numeracy, consistent with more educated mothers bridging home and school environments.&lt;/p&gt;
&lt;p&gt;Q: What is the primary mechanism driving intergenerational effects?
A: The primary mechanism is changes in parenting behaviors, not income. Female recipients do not invest more money in children (no significant difference in SES index or child investment index). Instead, they seek more prenatal care, engage in significantly more preventive health behaviors, and interact more with their children in cognitively stimulating ways. Day-long LENA recordings at 18 months show 20% more conversational turns per minute (effect size 0.068, p=0.005) and 17% more child vocalizations per minute (effect size 0.32, p=0.014). Caregiver reports confirm more playing, singing, and doing simple mathematics with children.&lt;/p&gt;
&lt;p&gt;Q: Does the income effect of scholarship receipt explain the child outcomes?
A: No. Duflo et al. (2024) find no significant earnings impacts until 2019 or later, meaning children tested at ages five and seven by 2023 largely grew up before their mothers&amp;rsquo; earnings improved. The household SES index shows only a 0.107 SD gain (p=0.103), indistinguishable from the effect for children of male recipients. There is also no evidence of a quality-quantity trade-off: caregivers of scholarship recipients do not have fewer children to care for.&lt;/p&gt;
&lt;p&gt;Q: Does the increase in maternal age at birth explain the child mortality reduction?
A: It is not the primary driver. Maternal age at birth increases by only 0.349 years on average (p=0.142) for children of female recipients, and 0.64 years for first-born children (p=0.040). Point estimates on mortality for first-born children are somewhat smaller than for the full sample, suggesting maternal age is not the main channel. Moreover, maternal age at birth falls for children of male recipients yet their survival point estimates are positive, which further argues against maternal age as the primary mechanism.&lt;/p&gt;
&lt;p&gt;Q: How does the education of the primary caregiver mediate the results?
A: For 84% of children in the sample, the primary caregiver is the child&amp;rsquo;s mother. Children of female scholarship recipients have caregivers who are 25 percentage points more likely to have completed secondary school and 5 percentage points more likely to have completed tertiary education. Children of male scholarship recipients have caregivers with no more education than the control group, because the recipients&amp;rsquo; partners — the typical caregivers — are not more educated. Treatment effects for female recipients are not altered when father&amp;rsquo;s education is added as a control, confirming maternal education as the main driver.&lt;/p&gt;
&lt;p&gt;Q: What threat to validity arises from co-residence of the father?
A: Children of male scholarship recipients are 8.7 percentage points less likely to live with their father (p=0.024), compared to no such effect for children of female recipients (92% of whom live with their scholarship-recipient mother). LENA recordings show negative treatment effects for children of male recipients — fewer adult words and conversational turns — consistent with father absence mechanically reducing auditory engagement and possibly leaving single mothers less time to verbally interact with each child.&lt;/p&gt;
&lt;p&gt;Q: How are multiple-hypothesis testing concerns addressed?
A: The pre-analysis plan pre-specified child survival and child cognitive development as primary outcomes. The authors apply the Romano-Wolf step-down procedure for multiple hypothesis testing adjustment. After adjustment, the p-value for survived-to-one for children of female recipients rises from 0.028 to 0.119; the cognitive development effects at age five and seven remain significant.&lt;/p&gt;
&lt;p&gt;Q: How does the study address potential sample selection bias in the child outcomes sample?
A: The authors use entropy balancing (Hainmueller, 2012) to reweight observations so that baseline (2008) characteristics are balanced between treatment and control within the subsample of recipients who had children. Results are qualitatively unchanged for both female and male recipients. The authors also note that children of female recipients are younger on average (4.71 months, p=0.067), which is why the study collects data at fixed age windows (14-22 months, 2.5 years, 3.5 years, 5 years, 7 years) rather than in a single cross-sectional wave.&lt;/p&gt;
&lt;p&gt;Q: What is the cost-effectiveness and cost-benefit result for secondary school scholarships?
A: Social costs are estimated at $585 per recipient for a mixed-gender program and $505 for a female-only program (combining school fees, materials, and foregone wages). The cost per under-three death averted is $23,582 for mixed-gender and $15,184 for female-only — placing the female-only program within the range of the 10th-percentile most cost-effective WHO-recommended child health interventions. The IRR is 27%–76% for a female-only means-tested scholarship program and 20%–51% for a mixed-gender program. These are likely conservative, as they exclude welfare gains from avoiding unwanted pregnancies, greater female agency, and recipient health benefits.&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the experiment and to what population do findings generalize?
A: The study estimates ITT effects for students in rural Ghana who qualified for SHS on exam performance but faced binding financial constraints in 2008 — a population that is academically prepared but economically disadvantaged. Results do not directly apply to students who would not have qualified, to contexts without binding financial barriers, or to settings where secondary school quality or the marriage market differs substantially. The study also cannot yet observe complete fertility, since scholarship-lottery participants were only 31 years old on average at last follow-up.&lt;/p&gt;
&lt;p&gt;LENA (Language Environment Analysis): A day-long recording device worn by a child that uses speech recognition software to generate count-based metrics — adult word count, adult-child conversational turns, and child vocalizations per minute — providing an objective measure of the child&amp;rsquo;s auditory environment and caregiver engagement quality without reliance on self-report.&lt;/p&gt;
&lt;p&gt;IRT Score (Item Response Theory Score): A latent-trait measure of child cognitive ability estimated from a one-parameter logistic model applied to binary correct/incorrect responses across cognitive game questions, assigned a difficulty level to each question and a latent ability to each child, then standardized. Used as the primary cognitive development outcome across age windows.&lt;/p&gt;
&lt;p&gt;Incarceration Effect: The hypothesis that education delays fertility mechanically only while students are in school (analogous to incarceration preventing activity), with no persistent effect once they exit. The authors rule this out by showing that the fertility gap between female treatment and control groups persists well after the majority of scholarship recipients have graduated.&lt;/p&gt;
&lt;p&gt;Quality-Quantity Trade-off (Becker 1991): The economic framework predicting that more educated parents, facing higher opportunity costs of children and lower costs of investing in child quality, will have fewer but better-invested-in children. The authors find delayed and reduced fertility but do not find that recipients have fewer children to care for in the cognitive assessment sample, suggesting the child quality gains operate primarily through parenting practices rather than resource concentration.&lt;/p&gt;
&lt;p&gt;Intent-to-Treat (ITT) Effect: The treatment effect estimated by comparing all lottery winners to all losers regardless of whether winners actually enrolled, which captures the effect of the scholarship offer (including compliance costs). The cost-benefit analysis uses ITT estimates, so the cost of subsidizing inframarginal students who would have attended anyway is incorporated.&lt;/p&gt;
&lt;p&gt;Entropy Balancing: A reweighting procedure (Hainmueller, 2012) that assigns weights to observations in the control group so that the weighted distribution of baseline covariates matches that of the treatment group, used to assess whether imbalances in the subsample of participants who had children drive the results. The authors apply this as a robustness check for both mortality and cognitive development outcomes.&lt;/p&gt;
&lt;p&gt;Unwanted Pregnancy: A pregnancy reported by the respondent as unplanned at the time of conception, which the authors use to distinguish fertility reduction from a change in desired fertility versus a reduction in unintended out-of-wedlock pregnancies. The scholarship&amp;rsquo;s early fertility impact is almost entirely a reduction in unwanted pregnancies (7 percentage point decline, 17% reduction).&lt;/p&gt;</description></item><item><title>Leveraging Virtual Contact and Social Networks to Foster Interethnic Harmony</title><link>https://macropaperwarehouse.com/papers/leveraging-virtual-contact-and-social-networks-to-foster-interethnic-harmony/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/leveraging-virtual-contact-and-social-networks-to-foster-interethnic-harmony/</guid><description>&lt;p&gt;This paper investigates whether virtual contact — exposure to an outgroup through a documentary film — can promote interethnic harmony, and whether targeting network-central individuals amplifies effects on untreated community members. The study addresses a context of deep, historically rooted discrimination: the Santal ethnic minority in northwestern Bangladesh have faced colonial-era land dispossession, ongoing violence, labor market discrimination, and structural exclusion by the Bengali ethnic majority. The Santals are the second-largest ethnic-minority group in Bangladesh; in the study villages, their share ranges from 13% to 83% of the population.&lt;/p&gt;
&lt;p&gt;The authors conducted a cluster-randomized field experiment across 121 multiethnic villages in the Rajshahi and Naogaon districts of Bangladesh, involving over 3,300 households. Villages were randomly assigned to three arms: a random treatment arm (RR, 40 villages, N=562 Bengalis) in which approximately 14 randomly selected ethnic-majority households per village watched a 45-minute documentary film (&amp;ldquo;Ami Santal&amp;rdquo; / &amp;ldquo;I Am Santal&amp;rdquo;) portraying Santal culture, economic hardships, and aspirations; a central treatment arm (41 villages) in which approximately 7 randomly selected Bengalis (RC) and 7 network-central Bengalis identified via a diffusion-centrality nomination exercise (CC) watched the same film; and a control arm (40 villages) in which households watched a placebo documentary on flower farming. The documentary, costing approximately $13 per participant, was screened individually at participants&amp;rsquo; homes on tablets. Data were collected at baseline (September–October 2022), first end line approximately 3 months post-screening (February–March 2023), and a casual-work field experiment second end line approximately 4.5–5 months post-screening (April–May 2023). Outcomes were measured via lab-in-the-field experiments (dictator game, solidarity game), an experimentally validated interethnic trust survey item (Falk et al. 2018), self-reported behaviors, administrative police complaint data, and facial emotion detection during screening.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. First, treated Bengalis in the central arm (RC) gave 14.7% more in the dictator game (p &amp;lt; .01) and exhibited 21.7% greater trust toward Santals (p &amp;lt; .01) compared to controls; RR participants showed a 7.1% increase in solidarity game giving (p &amp;lt; .10) and 11.8% greater trust (p &amp;lt; .01). Effects on reducing negative stereotypes and discriminatory opinions were not statistically significant, suggesting that affective components of prejudice are more responsive to the intervention than cognitive components. About 82% of treated Bengalis reported acquiring new information about Santals, primarily regarding occupational struggles, educational aspirations, and economic potential. Facial expression analysis using emotion-detection software found sadness to be significantly more prevalent among viewers (p &amp;lt; .05), particularly among network-central participants, consistent with an empathetic response.&lt;/p&gt;
&lt;p&gt;Second, untreated Bengalis in the central arm — who never watched the documentary — showed 20.9% higher altruism (p &amp;lt; .10), 27.3% higher solidarity (p &amp;lt; .05), and 8.1% higher trust (p &amp;lt; .05) toward Santals relative to controls. No significant effects on untreated Bengalis were found in the random arm. Untreated Santals in both arms exhibited greater trust toward Bengalis (11% increase in random arm, p &amp;lt; .05; 21.7% increase in central arm, p &amp;lt; .01) and higher subjective well-being (p &amp;lt; .01 in both arms). Village-level administrative data show a significant reduction in Bengali police complaints against Santals post-intervention (p &amp;lt; .05), but only in the central arm.&lt;/p&gt;
&lt;p&gt;Third, in the casual-work field experiment, multiethnic pairs jointly produced paper bags under piece-rate compensation. Overall productivity increased approximately 5% (p &amp;lt; .05) in the central arm only. Both Bengali and Santal workers increased productivity specifically in the finisher role — the most critical role for determining earnings — in the central arm. The authors interpret Bengali productivity gains as reflecting increased prosociality toward Santal co-workers, and Santal productivity gains as reflecting conformism or peer pressure in response to Bengali effort. The scope of all effects is limited to multiethnic villages in northwestern Bangladesh, a context of historically severe and ongoing majority-minority inequality; the intervention deliberately did not challenge the socioeconomic hierarchy of the villages.&lt;/p&gt;
&lt;p&gt;Q: What was the documentary film&amp;rsquo;s content and design rationale?
A: The 45-minute film &amp;ldquo;Ami Santal&amp;rdquo; featured three narrative layers: Santal culture (rituals, cuisine, the Baha festival), economic hardships (housing, water access, low incomes, labor market struggles, educational barriers), and aspirational stories of Santals who achieved success. All stories were narrated by non-actor local Santals, filmed outside the study region, and deliberately avoided attributing blame to Bengalis. The film was designed under the supervision of anthropologists at the University of Rajshahi to maintain ethnographic authenticity and a non-moralistic, observational tone (moral judgment language was much lower than in comparison Bangladeshi documentaries and general films, per LIWC-22 analysis).&lt;/p&gt;
&lt;p&gt;Q: How were network-central individuals identified and why might targeting them matter?
A: In central-arm villages, enumerators surveyed approximately 18–20 randomly selected passers-by at village markets and asked them to nominate the 15 people most effective at disseminating information. The seven most consistently and highly ranked individuals per village were selected as network-central (CC). These individuals were expected to have high diffusion centrality — meaning information they receive spreads widely — so targeting them with the documentary could shift attitudes and behavior among untreated community members through persuasion, visibility, credibility, or diffusion (the paper cannot separately identify which mechanism operates).&lt;/p&gt;
&lt;p&gt;Q: What were the primary behavioral effects on treated Bengalis (the ethnic majority who watched the film)?
A: Randomly selected participants in the central arm (RC) gave 14.7% more in the dictator game (p &amp;lt; .01) and 8% more in the solidarity game (not statistically significant), and exhibited 21.7% greater trust toward Santals (p &amp;lt; .01), all relative to controls. In the random arm (RR), participants showed a 6.4% increase in dictator game giving (not statistically significant), a 7.1% increase in solidarity game giving (p &amp;lt; .10), and 11.8% greater trust toward Santals (p &amp;lt; .01). Effects on self-reported behaviors — interethnic friendships, social interactions, amount charged to minorities for water — were not statistically significant.&lt;/p&gt;
&lt;p&gt;Q: Did the intervention change Bengali stereotypes or discriminatory opinions toward Santals?
A: No. Despite treated Bengalis acquiring substantial new information (approximately 82% reported learning new things, primarily about Santal occupational struggles and educational aspirations), the authors find no significant effects on the stereotypes index or the discriminatory-opinions index among treated Bengalis. They propose two explanations: cognitive components of prejudice (stereotypes) are harder to change through indirect contact than affective components (emotions, prosocial behavior), consistent with Tropp and Pettigrew (2005) and Turner, Crisp, and Lambert (2007); and a single documentary may be insufficient to counter deeply ingrained generational biases due to resistance to change.&lt;/p&gt;
&lt;p&gt;Q: What emotional responses did the documentary elicit, and how was this measured?
A: Field assistants took candid photographs of participants&amp;rsquo; faces at a random point during the screening; these were analyzed using Emotimeter software (machine learning-based emotion detection) that assigns scores across seven emotion categories summing to 100%. Sadness was significantly more prevalent among documentary viewers compared to placebo viewers (p &amp;lt; .05), particularly among network-central participants (CC). The authors interpret this as consistent with an empathetic response to the film&amp;rsquo;s content about Santal hardships, and connect it to increased prosocial behavior via emotion-regulation mechanisms (alleviating sadness through prosocial action).&lt;/p&gt;
&lt;p&gt;Q: What were the spillover effects on untreated Bengalis in the central arm?
A: Untreated Bengalis in central-arm villages — who never watched the documentary — showed 20.9% higher altruism (p &amp;lt; .10), 27.3% higher solidarity (p &amp;lt; .05), and 8.1% higher trust toward Santals (p &amp;lt; .05) relative to controls. By contrast, untreated Bengalis in random-arm villages showed no statistically significant effects on any of these outcomes. The authors attribute the central-arm spillovers to the presence of network-central individuals being treated in those villages, though whether these patterns reflect persuasion, visibility, credibility, or information diffusion cannot be separately identified.&lt;/p&gt;
&lt;p&gt;Q: How did the intervention affect the Santal ethnic minority (who never watched the documentary)?
A: Untreated Santals in both arms exhibited greater trust toward Bengalis: an 11% increase in the random arm (p &amp;lt; .05) and a 21.7% increase in the central arm (p &amp;lt; .01) compared to controls. Santals in both arms also reported higher subjective well-being (p &amp;lt; .01). A weakly significant increase in food security was observed among Santals in the central arm (p &amp;lt; .10), possibly reflecting increased material support from Bengalis. No statistically significant effects were found on Santal altruism or solidarity.&lt;/p&gt;
&lt;p&gt;Q: What did the village-level administrative complaint data show?
A: Using data collected from two police stations covering all 121 villages, the authors find a significant reduction in Bengali complaints against Santals post-intervention in the central arm (p &amp;lt; .05). No significant reduction was found in Santals&amp;rsquo; complaints against Bengalis (p &amp;gt; .10) in any arm. Data from village counselors&amp;rsquo; offices (shalish arbitration complaints) showed no significant change in any arm. The distinction matters because police complaints involve more serious, violent matters, while village-counselor complaints involve routine arbitration.&lt;/p&gt;
&lt;p&gt;Q: How was the casual-work field experiment designed, and what did it find?
A: Approximately 4.5 months after the documentary screenings, 720 participants (360 Bengalis, 360 Santals) drawn equally from the three study arms were paired into multiethnic dyads to jointly produce paper bags for a local supplier under piece-rate compensation, with earnings split equally. One worker was randomly assigned the preparer role and the other the finisher role; roles were switched halfway through the three-hour session. The paper finds an approximately 5% overall productivity increase (p &amp;lt; .05) in the central arm only, concentrated in the finisher role (the role most critical for final output). Bengalis and Santals both increased productivity specifically as finishers in the central arm.&lt;/p&gt;
&lt;p&gt;Q: What mechanisms explain the productivity effects in the casual-work experiment?
A: For Bengali finishers, the productivity gain is interpreted as prosocial behavior: treated Bengalis who showed greater altruism toward Santals worked harder to increase the earnings of their Santal co-workers. For Santal finishers, the productivity gain is interpreted as conformism or peer pressure: Santals increased effort more when they worked as finisher after swapping roles (i.e., after observing Bengalis&amp;rsquo; higher effort as finisher first), suggesting responsiveness to the higher productivity of Bengalis rather than an independent prosocial motivation. The authors present a simple theoretical model to formalize these interpretations, citing Rotemberg (1994) on prosocial effort and Kandel and Lazear (1992) and Mas and Moretti (2009) on peer pressure mechanisms.&lt;/p&gt;
&lt;p&gt;Q: Why was virtual rather than direct contact used in this intervention?
A: The authors argue that encouraging direct contact between Bengalis and Santals in this setting carries specific risks: the unequal status of the groups may generate anxiety during interactions, potentially limiting engagement or provoking backlash. By contrast, the documentary provides an indirect, low-cost ($13 per participant) form of contact that presents Santal lives without disrupting the socioeconomic hierarchy of the villages and without attributing blame to Bengalis. The film&amp;rsquo;s entertaining veneer and emotional storytelling make it more scalable and logistically feasible in contexts where direct contact is socially difficult or impractical.&lt;/p&gt;
&lt;p&gt;Q: What are the primary limitations acknowledged by the authors?
A: The authors acknowledge that the study&amp;rsquo;s sampling protocol relied on a door-to-door skip procedure without systematic records of approached households, raising the possibility of convenience or snowball-type recruitment and potential deviations from random sampling — this is reflected in some imbalances in baseline characteristics across arms. CC-control comparisons are explicitly descriptive (not causal) because network-central individuals were selected on centrality. Differential attrition was found among untreated Santals (both treatment arms had significantly lower attrition than control, p &amp;lt; .05), which could bias estimates for that subgroup. The authors cannot separately identify the mechanisms (persuasion, visibility, credibility, diffusion) underlying spillover effects in central villages.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of this study?
A: The findings suggest that media-based virtual contact interventions are a low-cost, scalable tool for improving interethnic prosociality even in contexts of deep-rooted discrimination where direct contact may be socially impractical. Targeting network-central individuals — identified via a simple nomination exercise requiring no pre-existing network data — amplifies village-wide effects, including among untreated community members and the minority group itself. The productivity gains in multiethnic work teams imply that improved interethnic relations can have tangible economic consequences beyond attitudinal change. However, the null effects on stereotypes and discriminatory opinions suggest that single documentary interventions may not be sufficient to alter deep-seated cognitive biases, and more intensive or repeated interventions may be needed to achieve durable attitude change.&lt;/p&gt;
&lt;p&gt;Virtual contact: Indirect exposure to an ethnic outgroup through a documentary film, as distinct from direct intergroup contact; posited to influence majority-group attitudes and behavior by increasing empathy and identification with the outgroup without requiring face-to-face interaction.&lt;/p&gt;
&lt;p&gt;Diffusion centrality: A network measure of how effectively an individual can spread information through a community, operationalized via a nomination exercise in which community members identify those best positioned to disseminate information; used to select the seven highest-ranked individuals per village for targeted treatment.&lt;/p&gt;
&lt;p&gt;Prosociality (altruism and solidarity): Measured using incentivized lab-in-the-field games — the dictator game (unilateral allocation of an endowment to a passive outgroup recipient) and the solidarity game (precommitted transfers to an outgroup member who may incur a random loss) — capturing willingness to benefit non-coethnic others at personal cost.&lt;/p&gt;
&lt;p&gt;Affective versus cognitive components of prejudice: A distinction between emotional aspects of prejudice (feelings, empathy) — which the authors find to be more responsive to the documentary intervention — and cognitive aspects (negative stereotypes, discriminatory opinions) — which show no significant change despite new information acquisition.&lt;/p&gt;
&lt;p&gt;Spillover effects (untreated individuals): Changes in behavior or attitudes among community members who did not directly receive the intervention (did not watch the documentary), attributed to the influence of treated individuals in their village, particularly network-central individuals in the central arm.&lt;/p&gt;
&lt;p&gt;Piece-rate casual-work field experiment: A second end line in which multiethnic pairs of Bengali and Santal workers jointly produced paper bags for a local supplier, with individual earnings determined by joint piece-rate output; designed to measure whether improved interethnic attitudes translated into higher workplace productivity in ethnically mixed teams.&lt;/p&gt;
&lt;p&gt;Source text origin: The provenance classification of the text used to generate a paper summary (full PDF, open-access HTML, or abstract only); the paper&amp;rsquo;s pipeline rules impose a hard block on abstract-only summarization.&lt;/p&gt;</description></item><item><title>Optimal Public Transportation Networks: Evidence from the World's Largest Bus Rapid Transit System in Jakarta</title><link>https://macropaperwarehouse.com/papers/optimal-public-transportation-networks-evidence-from-the-worlds-largest-bus-rapid-transit-system-in-jakarta/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-public-transportation-networks-evidence-from-the-worlds-largest-bus-rapid-transit-system-in-jakarta/</guid><description>&lt;p&gt;This paper studies how commuter preferences over wait times, travel times, and transfers should shape the design of urban bus networks, using the world&amp;rsquo;s largest Bus Rapid Transit (BRT) system — TransJakarta in Jakarta, Indonesia — as the empirical laboratory. The setting provides unusually rich identification: between January 2016 and February 2020, TransJakarta launched 93 new BRT and non-BRT feeder routes in a staggered, city-wide expansion, during which the operating bus fleet more than doubled from roughly 700 to over 1,600 vehicles. The authors combine over 500 million smart-card tap records, GPS tracking of every bus at 5–10 second intervals, and anonymized smartphone location data covering 35 million weekday trips from 2.3 million devices.&lt;/p&gt;
&lt;p&gt;The paper proceeds in three steps. First, the authors classify new route launches into three event types and estimate their causal impact on ridership via difference-in-differences. Event 1: a new direct connection between an origin-destination pair already served by transfer only, with no travel-time improvement — raises BRT ridership by 0.16 log points. Event 2: a new direct connection that also reduces travel time (by 0.29 log points on average) — raises ridership by 0.27 log points. Event 3: additional buses on an already-directly-connected pair, which increases the bus arrival rate by 0.32 log points and reduces wait times — raises ridership by 0.09 log points, implying a ridership elasticity with respect to wait times of approximately −0.29 for BRT. For non-BRT routes the implied wait-time elasticity is −1.05, raising the possibility of multiple equilibria in service levels. Crucially, none of the three event types produce detectable increases in aggregate trip volumes measured by smartphone data, implying the ridership gains reflect modal substitution toward the bus rather than trip generation.&lt;/p&gt;
&lt;p&gt;Second, the authors estimate a structural demand model. At its core is a route-choice model in which bus arrivals follow independent Poisson processes, so wait times are exponentially distributed and idiosyncratic. This formulation avoids the red-bus/blue-bus aggregation problem endemic to logit models. Commuters are also allowed to be partially inattentive to routes whose travel time exceeds the fastest available option by more than an estimated threshold. Structural parameters are recovered by classical minimum distance, matching seven reduced-form moments. Key findings: wait time is valued 2.4 times more than time on the bus for BRT routes, and 4.2 times more for non-BRT routes. There is no additional transfer penalty beyond the wait time and travel time costs of the second leg. Commuters pay significantly less attention to options with travel time more than roughly 34–44 percent above the fastest option in their choice set.&lt;/p&gt;
&lt;p&gt;Third, the authors use the estimated preference parameters to characterize optimal bus networks. Because the optimization problem is high-dimensional (418 grid cells, 1,536 possible edges, yielding on the order of 10^500 configurations) and exhibits neither global convexity nor simple complementarity, they reformulate the social planner&amp;rsquo;s problem as a discrete choice over networks with additive logit shocks — effectively sampling from a multinomial logit distribution via simulated annealing. The result: optimal networks cover approximately 66 percent of grid cells versus 42 percent under the actual TransJakarta network, and would give 91 percent of Jakarta residents bus access versus 73 percent currently. Bus frequency in the city center is somewhat lower in the optimal network. Despite commuters&amp;rsquo; high sensitivity to wait times, the current network concentrates too many buses in the city center where wait times are already short, rather than extending reach to underserved areas. Comparative statics show that doubling the wait-time cost parameter produces much more concentrated optimal networks (23 percent of origin-destination pairs connected, 41 percent fewer than baseline), while increasing the transfer penalty by the equivalent of 15 minutes of wait time raises the direct-connection share of served pairs from 12 to 16 percent.&lt;/p&gt;
&lt;p&gt;Q: What are the three event types and why are they analytically distinct?&lt;/p&gt;
&lt;p&gt;A: Event 1 is the launch of the first direct route between an origin-destination pair already connected by transfer, where the direct route is not faster than the existing transfer option; it isolates the effect of directness absent a travel-time change. Event 2 is the same but with a faster direct route (average reduction of 0.29 log points in travel time), combining directness and speed improvements. Event 3 is the launch of a new route that overlaps an existing direct route, increasing bus frequency and cutting wait times (arrival rate up 0.32 log points) without substantially changing travel time or directness. The three events together provide variation across the key dimensions — directness, speed, and frequency — needed to separately identify commuter preference parameters.&lt;/p&gt;
&lt;p&gt;Q: What are the main ridership effects and how large are they in levels?&lt;/p&gt;
&lt;p&gt;A: For BRT routes, Event 1 raises ridership by 0.16 log points (approximately 19 additional riders per week for a treated origin-destination pair with a baseline of 111 weekly riders), Event 2 by 0.27 log points (approximately 24 additional riders per week), and Event 3 by 0.09 log points (approximately 20 additional riders per week). For non-BRT routes, proportional effects are larger but level effects are similar: Event 1 yields roughly 34 additional weekly riders, Event 2 roughly 21, and Event 3 roughly 15. Event-study graphs show clear, discrete jumps in ridership at route launch with no pre-trends, and some gradual adjustment in the months following.&lt;/p&gt;
&lt;p&gt;Q: What does the paper find about aggregate trip generation versus modal substitution?&lt;/p&gt;
&lt;p&gt;A: Using smartphone location data to measure all trips regardless of mode, the authors find no statistically significant increase in aggregate trip volumes for any of the three event types. For BRT Event 1, the estimated aggregate-trip coefficient is −0.008 with a standard error of 0.051, allowing rejection at the 95 percent level of any positive impact above roughly 0.091 log points — small relative to the precise 0.11 log-point bus ridership effect in the same sample. The authors interpret this as evidence that the ridership gains over the 10-month post-event window reflect substitution from private modes (motorcycles, cars, taxis) toward TransJakarta rather than trip generation, and they use this null result to justify holding destination choices fixed in the structural model.&lt;/p&gt;
&lt;p&gt;Q: How does the model avoid the red-bus/blue-bus aggregation problem?&lt;/p&gt;
&lt;p&gt;A: The paper&amp;rsquo;s route-choice model assumes bus arrivals follow independent Poisson processes, so wait times are exponentially distributed. A key proposition (Proposition 1) proves that splitting one route into two identical routes with half the buses each produces exactly the same choice probabilities and expected utility as the original single route — because the sum of two independent Poisson processes is itself Poisson with the summed rate. Standard logit models fail this invariance because splitting a route creates two options with independent error draws, artificially inflating expected utility. The invariance property is essential for the optimal network design exercise, where the planner freely reallocates buses across routes.&lt;/p&gt;
&lt;p&gt;Q: What are the estimated preference parameters and what do they imply about commuter behavior?&lt;/p&gt;
&lt;p&gt;A: The paper estimates that wait time is valued 2.4 times more than time on the bus for BRT routes and 4.2 times more for non-BRT routes. There is no additional transfer disutility beyond the wait time and travel time costs implied by the extra leg. Commuters become substantially inattentive to routes with travel time more than approximately 34 percent above the fastest available option (BRT threshold) or 44 percent (non-BRT). The high relative cost of waiting versus riding reflects both the discomfort of waiting at exposed non-BRT stops and the fact that TransJakarta runs without a published schedule, so commuters cannot minimize wait time by timing arrivals.&lt;/p&gt;
&lt;p&gt;Q: What explains the non-BRT wait-time elasticity exceeding −1?&lt;/p&gt;
&lt;p&gt;A: For non-BRT routes, Event 3 raises ridership by 0.450 log points while raising the bus arrival rate by 0.425 log points, yielding an implied elasticity of ridership with respect to wait times of −1.05. Because the baseline arrival rate for non-BRT treated pairs is 2–4 times lower than for BRT pairs, the absolute reduction in wait time per additional bus is much larger. An elasticity exceeding −1 in absolute value implies that adding buses on some non-BRT routes could increase ridership enough to maintain or even raise average ridership per bus — the extreme form of the Mohring effect — suggesting the possibility of a high-ridership/low-wait-time equilibrium distinct from the current low-ridership/high-wait-time one.&lt;/p&gt;
&lt;p&gt;Q: How is the optimal network characterized and what algorithm is used?&lt;/p&gt;
&lt;p&gt;A: The social planner chooses a network to maximize utilitarian welfare (average expected utility across all commuters) from the estimated demand model, plus a network-level logit shock capturing cost and other factors outside the model. This transforms the combinatorially explosive optimization into sampling from a multinomial logit distribution over networks, which the authors approximate using simulated annealing. They run the algorithm multiple times to obtain a sample of networks drawn asymptotically from the planner&amp;rsquo;s distribution, then estimate optimal network characteristics and comparative statics from sample analogs. The theoretical framework is general and, the authors note, applicable to other high-dimensional spatial planning problems where welfare differences can be computed for pairs of counterfactuals.&lt;/p&gt;
&lt;p&gt;Q: How does the optimal network differ from the current TransJakarta network?&lt;/p&gt;
&lt;p&gt;A: The typical optimal network covers approximately 66 percent of 2km grid cells versus 42 percent for the actual network, and 91 percent of Jakarta residents would have bus access versus 73 percent currently. The optimal network reduces bus frequency in the city center relative to the current network, accepting longer wait times there in order to extend reach to peripheral areas. The paper finds no tension between distributional and efficiency concerns in this setting — expanding coverage improves both aggregate welfare and access for underserved areas.&lt;/p&gt;
&lt;p&gt;Q: What do the comparative statics reveal about the sensitivity of optimal network design to preference parameters?&lt;/p&gt;
&lt;p&gt;A: Doubling the wait-time cost parameter leads to substantially more concentrated optimal networks: only 23 percent of origin-destination pairs are connected, 41 percent fewer than in the baseline optimal network. This is because higher wait-time costs make it more valuable to concentrate buses on fewer routes to achieve short headways. Increasing the transfer penalty by the equivalent of 15 minutes of wait time raises the share of connected location pairs with a direct (non-transfer) connection from 12 to 16 percent. These comparative statics link micro-level preference parameters to macro-level network topology, clarifying which parameters most influence design choices.&lt;/p&gt;
&lt;p&gt;Q: How does the paper validate the destination imputation from tap-in-only smart card data?&lt;/p&gt;
&lt;p&gt;A: For the subset of BRT stations where tap-out is enforced (36 percent of stations), the authors estimate bivariate regressions of imputed daily ridership shares against actual observed ridership shares, obtaining R-squared of 0.85. They also show robustness by varying the grid cell size from 500 meters to 2 kilometers, finding no systematic decline in treatment effect magnitudes, which rules out large displacement effects within the network as an explanation for the results.&lt;/p&gt;
&lt;p&gt;Q: Does the response to network improvements vary by local poverty rates?&lt;/p&gt;
&lt;p&gt;A: The authors interact all six event types with an indicator for above-median poverty rate at the origin grid cell (from SMERU 2014 data), controlling for population. They find no clear pattern of heterogeneity by income level — richer and poorer areas respond similarly to service improvements. The paper notes this absence of heterogeneity as relevant context for interpreting optimal network design: the case for extending reach is not offset by a differential preference for frequency among poorer commuters.&lt;/p&gt;
&lt;p&gt;Mohring Effect: The externality arising from ridership responsiveness to wait times — more riders justify more buses, which reduce wait times for all riders, further increasing ridership. The paper estimates a BRT wait-time elasticity of −0.29, confirming the effect operates in Jakarta; for non-BRT the elasticity of −1.05 suggests the possibility of multiple equilibria in service levels.&lt;/p&gt;
&lt;p&gt;Negative Exponential Distribution Model (Daganzo 1979): The route-choice model used in the paper, in which bus arrivals on each route follow independent Poisson processes and wait times are exponentially distributed. The model is invariant to aggregation of identical routes (avoids the red-bus/blue-bus problem) and yields tractable closed-form expressions for choice probabilities and expected utility.&lt;/p&gt;
&lt;p&gt;Partial Inattention: The model feature whereby commuters assign near-zero effective arrival rates to bus options whose travel time exceeds the fastest available option by more than an estimated threshold (34–44 percent depending on route type). Captures the empirical finding that commuters in a large, complex network do not appear to consider all available options.&lt;/p&gt;
&lt;p&gt;Event Types (1, 2, 3): The paper&amp;rsquo;s taxonomy of service improvements induced by new route launches. Event 1 isolates the value of directness (new direct route, no speed gain). Event 2 combines directness and speed (new direct route that is also faster). Event 3 isolates the value of frequency (additional buses on an already-direct route, reducing wait time without changing travel time).&lt;/p&gt;
&lt;p&gt;Optimal Network Characterization via Social Planner&amp;rsquo;s Logit: The paper&amp;rsquo;s approach to the combinatorially intractable network optimization problem. The planner is modeled as making a logit discrete choice over all possible networks, with welfare from the demand model plus a network-level idiosyncratic shock. Sampling via simulated annealing yields estimates of optimal network characteristics and comparative statics without requiring identification of a single globally optimal network.&lt;/p&gt;
&lt;p&gt;Network Concentration vs. Extensiveness Tradeoff: The core design tension the paper formalizes — for a fixed bus fleet, concentrating buses on fewer routes reduces wait times on served routes but leaves more areas without coverage, while spreading buses across more routes extends reach at the cost of longer headways. The estimated preference parameters (high wait-time sensitivity) make this tradeoff non-trivial; nonetheless, the paper finds the current network is too concentrated relative to the optimum.&lt;/p&gt;</description></item><item><title>Patent Term, Innovation, and the Role of Technology Disclosure Externalities</title><link>https://macropaperwarehouse.com/papers/patent-term-innovation-and-the-role-of-technology-disclosure-externalities/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/patent-term-innovation-and-the-role-of-technology-disclosure-externalities/</guid><description>&lt;p&gt;This paper examines how anticipated changes in patent term affect R&amp;amp;D and innovation, using the U.S. ratification of the Trade-Related Aspects of Intellectual Property Rights (TRIPs) agreement in 1995 as a quasi-natural experiment. The central research question is whether and how policy anticipation shapes the short- and long-run dynamics of innovative activity, given ambiguous theoretical predictions: news of a patent term reduction could either deter innovation (by signaling lower future returns) or accelerate it (by inducing innovators to file under the more favorable existing regime before it expires).&lt;/p&gt;
&lt;p&gt;The identification strategy exploits a difference-in-differences (DiD) design using two sources of variation across 621 4-digit International Patent Classification (IPC) technological fields. The first is cross-sectional variation in field-specific pending periods — the time between patent application and grant during which monopoly rights are not fully enforceable — which determines whether TRIPs increased or reduced each field&amp;rsquo;s effective patent term (from 17 years post-grant to 20 years post-application minus the pending period). Fields with average pending periods exceeding three years faced expected reductions; those below faced extensions. On average across fields, TRIPs extended patent term by approximately 473 days (about 15 months), but approximately 45% of fields faced greater than 5% probability that individual patents would receive a term reduction. The second source is time variation from two events: a news event at the end of 1992 (when the Blair House Accord substantially reduced uncertainty about TRIPs adoption) and implementation in June 1995. The empirical sample spans 1985Q1–2000Q4 using PATSTAT patent data, augmented by firm-level R&amp;amp;D data from NBER-Compustat for 2,410 listed U.S. firms.&lt;/p&gt;
&lt;p&gt;Three main empirical facts emerge. First (Fact 1), innovation and R&amp;amp;D accelerate more during the anticipation phase (1992Q4–1995Q2) in fields with a higher probability of patent term reduction. A one-percentage-point higher reduction probability corresponds to a 1.4% larger increase in granted patent applications before implementation; a one-month shorter average patent term extension corresponds to a 2.9% larger increase. At the firm level, a one-percentage-point higher reduction probability is associated with a 1.9% increase in annual R&amp;amp;D expenditure (approximately $1.7 million), ruling out the interpretation that rising patent counts merely reflect strategic filing adjustments.&lt;/p&gt;
&lt;p&gt;Second (Fact 2), this heightened innovative activity persists for at least five years after implementation. Two years post-implementation, a one-percentage-point higher reduction probability corresponds to 1.44 additional quarterly patents (+2.7% in Poisson estimates), and a one-month shorter term extension corresponds to 3.3 more patents (+5.9%). This persistence is driven by indirect effects: the anticipation-induced burst in patenting generates additional follow-on innovation through technology disclosure externalities linked to cumulative knowledge creation. The elasticity of post-implementation innovation to news-phase innovation is estimated at approximately 2.1.&lt;/p&gt;
&lt;p&gt;Third (Fact 3), the direct effect of patent term on innovation — estimated by augmenting the DiD specification to control for field-specific innovation histories — is negative for shorter extensions and consistent with prior literature. A one-month shorter patent term extension reduces quarterly patents by 1.7%, and a one-year reduction reduces them by 20.9%. These estimates align with Budish, Roin, and Williams (2015, 2016), who find that a one-year extension of patent monopoly increases R&amp;amp;D by 7%–22% in pharmaceuticals. The identification is supported by the absence of pre-trends, by the finding that pre-news pending period distributions predict realized post-news variation with coefficients near one (0.957–1.104), and by extensive robustness checks.&lt;/p&gt;
&lt;p&gt;Q: What was the effective change in U.S. patent term under TRIPs, and why did it differ across fields?
A: TRIPs shifted patent expiry from 17 years after grant to 20 years after application date. Because monopoly rights are only fully enforceable after grant, the effective term became 20 years minus the pending period. Fields with average pending periods shorter than three years received net extensions; fields with longer average pending periods faced net reductions. Cross-field variation in pending periods arises because applications in different technical fields are reviewed by distinct USPTO technical units with different complexity and backlog levels.&lt;/p&gt;
&lt;p&gt;Q: What was the news event, and how was anticipation established?
A: The paper identifies November 1992 — when the Blair House Accord substantially reduced uncertainty about TRIPs adoption — as the news event, with formal ratification in December 1994 and implementation in June 1995. Documentary evidence confirms anticipation: U.S. business executives were involved in TRIPs negotiations from 1986; the patent term change appeared in a 1991 GATT draft; an Advisory Committee report co-signed by IBM, 3M, Motorola, and others referenced it in August 1992; and a New York Times article noted proposed changes in September 1992.&lt;/p&gt;
&lt;p&gt;Q: How is the probability of patent term reduction (PL_j) constructed, and what is its distribution?
A: PL_j is the fraction of patents in field j granted before the TRIPs news with a pending period exceeding three years, computed using PATSTAT data on U.S. patents granted between January 1990 and May 1992. Approximately 45% of fields faced a reduction probability exceeding 5%, and 15% faced a probability exceeding 10%. Even fields with an average term extension greater than one year had individual-patent reduction probabilities as high as 40%. A 10-percentage-point increase in PL_j corresponds to approximately a four-month shorter average term extension.&lt;/p&gt;
&lt;p&gt;Q: What is Fact 1 and what are its quantitative magnitudes?
A: Fact 1 states that during the news phase, innovation and R&amp;amp;D increase relatively more in fields with higher patent term reduction probability and shorter average term extension. One year after the news (two years before implementation), a one-percentage-point higher reduction probability generates 0.19 additional quarterly patents (+0.5% in Poisson estimates); a one-month shorter average extension generates 0.35 additional units (+0.8%). These effects approximately triple one year before implementation. At the firm level, a one-percentage-point higher probability is associated with a 1.9% increase in annual R&amp;amp;D (~$1.7 million) in 1993.&lt;/p&gt;
&lt;p&gt;Q: Why does news of a potential patent term reduction accelerate rather than deter innovation?
A: Innovators who anticipate a reduction in future patent protection under the new regime have strong incentives to file applications before implementation to secure the longer 17-years-from-grant term while it remains available. The acceleration is therefore consistent with innovators preferring longer protection: they rush to file under the more favorable old regime rather than curtailing innovation. Complementary analyses exploiting within-field dispersion in pending periods find that firms were particularly responsive to scenarios involving adverse policy changes, consistent with loss aversion. The dynamics of the news-phase acceleration are also consistent with an R&amp;amp;D gestation lag of approximately two years, as estimated by Pakes and Schankerman (1984).&lt;/p&gt;
&lt;p&gt;Q: What is Fact 2 and what drives the post-implementation persistence?
A: Fact 2 states that the heightened innovation in fields with higher reduction probability persists for at least five years after June 1995, even though the direct effect of a shorter patent term is innovation-reducing. Two years post-implementation, a one-percentage-point higher reduction probability corresponds to 1.44 additional quarterly patents (+2.7% Poisson) and a one-month shorter extension to 3.3 additional patents (+5.9% Poisson). The persistence is driven by technology disclosure externalities: the news-phase acceleration generates new patented knowledge that subsequent innovations build upon. Fields where new inventions rely more heavily on past innovations from the same field — proxied by backward citation intensity — display stronger post-implementation persistence.&lt;/p&gt;
&lt;p&gt;Q: How does the paper separate direct from indirect (externality-driven) post-implementation effects?
A: Following Angrist and Pischke (2009), the paper augments the baseline DiD specification to control for field-specific innovation histories via a lagged moving average of past outcomes and pre-determined field attributes interacted with quarterly fixed effects. The resulting coefficients capture the effect of patent term variation orthogonal to the news-induced innovation dynamics. The direct effect estimates are negative post-implementation (Fact 3), while the overall estimates are positive (Fact 2), confirming that the indirect externality channel outweighs the direct channel in the post-implementation period.&lt;/p&gt;
&lt;p&gt;Q: What is Fact 3 and how does its magnitude compare to prior literature?
A: Fact 3 states that, controlling for the news shock, a shorter patent term extension leads to a relative decline in innovation post-implementation. The estimated semi-elasticity is 1.7% per one-month increase in patent term and 20.9% per one-year increase. These estimates align with Budish, Roin, and Williams (2015, 2016), who find a 7%–22% increase in pharmaceutical R&amp;amp;D per one-year extension, and with Hemous et al. (2023), whose model implies a 1.2% innovation increase per one-month extension.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated elasticity of post-implementation innovation to news-phase innovation, and what does it imply?
A: Point estimates imply that one additional patent during the news phase generates approximately 5.1 additional patents post-implementation. Given average patent counts of 408.5 during the news phase and 1,000.3 post-implementation, this corresponds to a percent-to-percent elasticity of approximately 2.1. This elasticity captures the technology disclosure externality channel by which transitory accelerations in patenting generate persistent follow-on innovation.&lt;/p&gt;
&lt;p&gt;Q: Why is ignoring anticipation (as in Abrams 2009) a problem for DiD identification?
A: Anticipation inflates patenting in fields with higher reduction probability during the pre-implementation period, violating the DiD assumption that pre-implementation outcomes provide an unaffected baseline. For example, between April 1994 and March 1995, average monthly patents in field C12P (high reduction probability) were 15.1 units above pre-news levels, versus only 2.4 in field E05D (low reduction probability). Using this inflated pre-implementation level as the DiD reference baseline reverses the sign of the estimated implementation effect relative to the specification that uses the unaffected pre-news baseline.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports the technology disclosure externality mechanism over alternative explanations?
A: The paper proxies technological dependence by backward citation intensity at the field level and finds that the news-phase acceleration propagates more strongly into post-implementation innovation in fields where new inventions more heavily cite prior same-field patents. Time-varying measures of technological dependence identify this channel as the primary driver of indirect post-implementation effects. Two alternative mechanisms — changes in technological competition and adjustments in patenting strategies — lack comparable empirical support. The finding is consistent with Hegde, Herkenhoff, and Zhu (2023), who document that permanent increases in knowledge diffusion speed permanently raise follow-on innovation rates.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of jointly considering anticipation and knowledge spillovers?
A: Standard patent term analyses that abstract from anticipation effects and knowledge spillovers may substantially mischaracterize full welfare implications. The paper shows that innovation-policy interventions shape both short- and long-run outcomes, and that near-term variation in innovative activity can itself drive medium- to long-term effects through technological externalities. The estimated semi-elasticities of news, direct, and indirect effects provide empirical calibration targets for normative endogenous growth models used to derive optimal patent term, complementing prior normative recommendations ranging from zero protection (Boldrin and Levine, 2013) to infinite protection (Gilbert and Shapiro, 1990).&lt;/p&gt;
&lt;p&gt;Effective patent term: The duration of legally enforceable monopoly granted by a patent, equal to 17 years after grant under the pre-TRIPs U.S. regime and 20 years after application minus the pending period under the post-TRIPs regime. Because enforcement begins only at grant, the pending period directly erodes effective protection.&lt;/p&gt;
&lt;p&gt;Patent term reduction probability (PL_j): The field-specific fraction of pre-TRIPs patents with a pending period exceeding three years, representing the probability that individual patent applications in that field obtain a net reduction in patent term under the new 20-years-from-filing rule.&lt;/p&gt;
&lt;p&gt;News effect: The incremental change in innovation or R&amp;amp;D at the time of policy announcement, induced by future anticipated changes in patent term, before the new policy enters into force. In this paper&amp;rsquo;s setting, the news effect is positive: higher reduction probability accelerates patenting as innovators rush to file under the favorable existing regime.&lt;/p&gt;
&lt;p&gt;Direct implementation effect: The component of the post-implementation change in innovation attributable to the patent term change itself, isolated by controlling for field-specific innovation histories (i.e., abstracting from the indirect effects of anticipation-induced knowledge accumulation). It is negative for shorter patent term extensions, with a semi-elasticity of 1.7% per one-month increase.&lt;/p&gt;
&lt;p&gt;Technology disclosure externality: The mechanism by which newly patented knowledge, disclosed through the patent system, enables subsequent inventors to build on prior innovations, generating follow-on inventive activity. In this paper, the transitory news-phase burst in patenting generates a persistent externality, particularly in fields with high backward citation intensity.&lt;/p&gt;
&lt;p&gt;Policy anticipation: The phenomenon whereby forward-looking agents adjust behavior in response to credible news about future policy changes before those changes take effect. In this paper, anticipation induces a pre-implementation acceleration in patenting that temporarily pushes innovation in the opposite direction from the direct long-run effect and generates persistent indirect post-implementation effects through knowledge spillovers.&lt;/p&gt;
&lt;p&gt;Pending period: The time between patent application and grant during which USPTO examines the application and during which full monopoly rights are not enforceable. Field-level heterogeneity in pending periods — arising from differences in examination complexity and USPTO unit congestion — is the source of cross-sectional identification in the DiD design.&lt;/p&gt;</description></item><item><title>Praying for Rain</title><link>https://macropaperwarehouse.com/papers/praying-for-rain/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/praying-for-rain/</guid><description>&lt;p&gt;This paper studies rainmaking as an instrumental religious belief. The central research question is: why do people believe that prayer can bring rain, even though it does not work? The authors develop a model of cultural evolution in which a religious leader prays for rain at an arbitrary time, and people update their beliefs about whether the leader can cause rainfall based on whether rain follows. The key mechanism is the local rainfall hazard function — the probability of rain conditional on how many days have passed since the last rainfall. In environments where the hazard is increasing (rain becomes more likely the longer a drought continues), a leader who prays during a drought will tend to be followed by rain, creating the illusion of efficacy. In environments with a flat or declining hazard, prayer cannot be systematically followed by rain in a persuasive way. The model yields five predictions: rain ritual traditions will select for prayers correlated with rainfall; the level of average rainfall does not determine persuasiveness; constant-hazard environments cannot support persuasive prayer; increasing-hazard environments are more likely to adopt rainmaking; and higher net benefits of rainfall (e.g., settled agriculture) further increase the likelihood of ritual.&lt;/p&gt;
&lt;p&gt;The authors test these predictions with two empirical strategies. First, they use daily data from the Catholic church in Murcia, Spain, covering 1600 to 1836. Church records provide the daily timing of pro pluvia rogations (prayers for rain), while municipal council records — kept independently of the church — record notable rainfall events. Murcia&amp;rsquo;s rainfall hazard is estimated to be increasing after long dry spells: the hazard rate after a long drought is roughly double the hazard rate two months after the last rainfall. The main finding is that a prayer for rain in the last 30 days predicts a 0.144 percentage-point higher daily probability of notable rainfall (standard error 0.057 pp), relative to a baseline mean daily rainfall probability of 0.203 pp — a 71% increase in the predicted probability. Prayer also Granger-causes rainfall conditional on lags of recent rainfall, and the predictive power holds within a given calendar month, ruling out a purely seasonal coincidence.&lt;/p&gt;
&lt;p&gt;Second, the authors construct an original dataset covering rainmaking practice for 1,208 ethnic groups drawn from the Ethnographic Atlas (Murdock, 1967), coded from 370 anthropological sources. They match each ethnic group to its nearest weather station and estimate the rainfall hazard function each group faces in its ancestral location. Of the 1,208 groups, 33% face an increasing rainfall hazard, and 39% of all groups practice rain ritual. The main global finding is that ethnic groups facing an increasing rainfall hazard are 14 percentage points more likely to practice rainmaking (standard error 3.7 pp), relative to a base rate of 30% among groups facing a non-increasing hazard — a 47% increase. This result is robust to continent fixed effects, geographic and climatic controls (longitude, latitude, elevation, distance to coast, ruggedness, mean temperature, mean rainfall, coefficient of variation of rainfall, maximum dry spell length, and the Giuliano-Nunn 2021 climatic variability measure), alternative hazard estimation methods, and linguistic family fixed effects. Crucially, lower average rainfall, longer droughts, and greater climatic variability are not associated with more rain ritual conditional on hazard shape — it is specifically the shape of the hazard function, not aridity or variability per se, that drives adoption.&lt;/p&gt;
&lt;p&gt;A second global finding concerns demand: groups dependent on agriculture are 11 pp more likely to practice rainmaking; those dependent on intensive agriculture, 21 pp more likely; and those dependent on intensive irrigated agriculture, 32 pp more likely (on a base of 32%). The scope of the findings is the pre-modern or traditional period captured by the Atlas; the Murcia case covers 1600–1836. The authors conclude that some environments create an illusion of efficacy that sustains instrumental religious belief through cultural selection, without requiring that believers be irrational.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s central theoretical claim about why rainmaking beliefs persist?
A: The paper argues that in environments where the rainfall hazard is increasing during a drought, a leader who begins praying during a dry spell will tend to be followed by rain, because the probability of rain rises as the drought lengthens. People who cannot observe the counterfactual hazard (what rainfall would have been without prayer) interpret this coincidence as evidence that prayer works. Cultural selection then favors leaders whose prayer timing is more persuasive, causing the belief to persist across generations even though prayer does not actually cause rain.&lt;/p&gt;
&lt;p&gt;Q: What is the rainfall hazard function, and why does its shape determine whether prayer can be persuasive?
A: The hazard function h(t) gives the instantaneous probability of rain at time t days after the last rainfall. If the hazard is flat, the probability of rain is the same regardless of whether prayer was offered or not, so there is no systematic correlation between prayer and rainfall to exploit. If the hazard is declining, prayer during a drought will be followed by lower-than-average rainfall probability, undermining the leader. Only if the hazard is increasing does prayer during a long dry spell systematically coincide with a higher probability of rain, creating a persuasive correlation.&lt;/p&gt;
&lt;p&gt;Q: What do Propositions 2 and 3 of the model establish?
A: Proposition 2 establishes that if the hazard rate is constant and a person&amp;rsquo;s prior belief that prayer works is below 0.5, then no prayer start time can persuade them to support the leader. Proposition 3 establishes the converse: if the hazard rate is increasing and the prior is below 0.5, there exists a meaningful belief for which a person will support the leader for any prayer start time. Together these propositions identify the increasing hazard as the necessary and sufficient structural condition for persuasive prayer.&lt;/p&gt;
&lt;p&gt;Q: What is the main quantitative finding from Murcia, and what identification strategy supports it?
A: A prayer for rain in the last 30 days predicts a 0.144 percentage-point higher daily probability of notable rainfall (standard error 0.057 pp) relative to a baseline mean of 0.203 pp, a 71% increase. The authors additionally demonstrate that prayer Granger-causes rainfall conditional on lags of recent rainfall, and that the effect holds within a given calendar month, ruling out the explanation that prayer simply tracks the rainy season. The prayer and rainfall records are kept by independent institutions (church and municipal council), reducing the risk of strategic recording.&lt;/p&gt;
&lt;p&gt;Q: How does the hazard rate in Murcia behave, and does it satisfy the model&amp;rsquo;s key condition?
A: The hazard of rainfall in Murcia is initially high just after rain, declines to a minimum roughly two months after the last rainfall, and then increases significantly thereafter, reaching or exceeding its initial level after a long drought. The fluctuations are large: the hazard after a long dry spell is roughly double the hazard two months after rainfall. This U-shaped pattern means the hazard is increasing during a prolonged drought, satisfying the model&amp;rsquo;s key condition for persuasive prayer.&lt;/p&gt;
&lt;p&gt;Q: How was the global rainmaking dataset constructed, and what is its coverage?
A: The authors used the Ethnographic Atlas (Murdock, 1967) as a template, covering 1,290 ethnic groups, and combed 370 anthropological sources — primarily group-specific ethnographic monographs — to code rainmaking practice for 1,208 groups. A group is coded as practicing rain ritual only if there is clear evidence of a practice specifically intended to bring rain through supernatural means. The authors treat their measure as a lower bound. They find that 39% of the 1,208 groups practice rainmaking, across every settled continent.&lt;/p&gt;
&lt;p&gt;Q: What is the main global regression result and how robust is it?
A: Ethnic groups facing an increasing rainfall hazard are 14 percentage points more likely to practice rain ritual (standard error 3.7 pp) relative to a base rate of 30%, a 47% proportional increase. This coefficient is positive and statistically significant across all specifications, including those adding continent fixed effects, a full battery of geographic and climatic controls (longitude, latitude, elevation, distance to coast, ruggedness, mean temperature, mean rainfall, coefficient of variation of rainfall, maximum dry spell length, and the Giuliano-Nunn 2021 climatic variability measure), alternative hazard estimation methods, linguistic family fixed effects, and restrictions to groups with high-quality rainfall data.&lt;/p&gt;
&lt;p&gt;Q: Does aridity or climatic variability explain rainmaking adoption?
A: No. Lower average rainfall, longer droughts, and greater climatic variability (measured using the Giuliano-Nunn 2021 index) are not associated with more rain ritual practice, conditional on the shape of the hazard function. This rules out the naive hypothesis that people pray for rain simply because they do not get enough, or because their rainfall is unreliable. It is specifically the shape of the hazard — whether it is increasing during a drought — that drives adoption, not the level or volatility of rainfall.&lt;/p&gt;
&lt;p&gt;Q: How does demand for rainfall, proxied by agricultural subsistence, affect rainmaking adoption?
A: Groups dependent on agriculture are 11 percentage points more likely to practice rainmaking relative to other subsistence modes. Groups dependent on intensive agriculture are 21 percentage points more likely, and groups dependent on intensive irrigated agriculture are 32 percentage points more likely, all on a base of 32%. This gradient is consistent with Proposition 5 and 6 of the model: settled, location-specific agricultural investment raises the net benefit of rainfall control, increasing support for rain ritual independently of the persuasion channel.&lt;/p&gt;
&lt;p&gt;Q: What does the model&amp;rsquo;s cultural evolution mechanism (Proposition 4) predict about how prayer timing changes over generations?
A: Proposition 4 states that rituals with high support are more likely to persist. In increasing-hazard environments, random variation in prayer timing means some leaders gain more support than others; those with more persuasive timing are more likely to persist. Each generation then adopts a policy at least as persuasive as the prior generation, so support rises over time and prayers gradually converge toward the timing that maximizes persuasiveness. This mechanism does not require deliberate optimization by any individual leader.&lt;/p&gt;
&lt;p&gt;Q: How does the paper&amp;rsquo;s finding relate to the long-standing anthropological debate between the traditional and revisionist schools on rainmaking?
A: The traditional school (following Frazer 1890) holds that belief is instrumental — people engage in rainmaking to make rain, and belief responds to empirical evidence. The revisionist school (Wittgenstein, Durkheim) argues that religious belief and rationality are fundamentally separate, and religious practice is performative rather than evidence-responsive. The paper&amp;rsquo;s finding that rainmaking is more prevalent precisely where it is more persuasive — i.e., where the environment makes prayer appear to work — supports the traditional, instrumental interpretation that belief responds to evidence of efficacy.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions for the paper&amp;rsquo;s conclusions?
A: The Murcia case study covers the period 1600–1836, ending when the abolition of tithes reduced the church&amp;rsquo;s funding and influence; it applies to a sophisticated Catholic institutional context. The global analysis covers traditional practices of pre-modern ethnic groups as recorded in the Ethnographic Atlas and anthropological literature; it does not speak to modern religious practice or to religions after substantial modernization. The persuasion mechanism requires that people cannot directly observe what rainfall would have been without prayer, a condition satisfied in pre-scientific contexts.&lt;/p&gt;
&lt;p&gt;Rainfall hazard function: In this paper&amp;rsquo;s usage, the function h(t) = f(t)/(1-F(t)) giving the instantaneous probability of rainfall at time t days since the last rainfall. Its shape — whether flat, declining, or increasing during a drought — determines whether prayer can be persuasive, not the overall level of rainfall.&lt;/p&gt;
&lt;p&gt;Increasing hazard: A hazard rate that rises as the length of a dry spell increases, so that rain becomes more likely the longer the drought has continued. The paper defines this specifically as the derivative of the hazard function evaluated at the 99th percentile of spell length. This is the necessary structural condition for prayer to seem efficacious.&lt;/p&gt;
&lt;p&gt;Instrumental religious belief: Belief directed at achieving a worldly outcome (here, rainfall), as opposed to purely expressive or social belief. The paper treats belief as instrumental if it responds to perceived evidence of efficacy and is adopted where it appears to work.&lt;/p&gt;
&lt;p&gt;Persuasion (in the model): The process by which a leader&amp;rsquo;s prayer timing causes people to update their belief that prayer works, by generating a correlation between prayer and subsequent rainfall that exceeds what people expect from the background hazard rate. Persuasion is possible only when the hazard is increasing.&lt;/p&gt;
&lt;p&gt;Pro pluvia rogations: The Catholic church&amp;rsquo;s formal prayers for rain, practiced in Murcia since at least the 14th century. In the paper&amp;rsquo;s data, these prayers follow a pattern of escalation — increasing in number and intensity — during prolonged droughts, consistent with the model&amp;rsquo;s prediction about prayer timing.&lt;/p&gt;
&lt;p&gt;Cultural evolution: The paper&amp;rsquo;s framework (drawing on Henrich 2015) in which religious leaders act as cultural entrepreneurs; leaders whose prayer timing happens to be more persuasive gain greater support and are more likely to survive across generations, so prayer traditions drift toward more persuasive timing without deliberate design.&lt;/p&gt;
&lt;p&gt;Rain ritual (global measure): A binary indicator coded as one for an ethnic group if the anthropological literature contains clear evidence of a practice specifically intended to bring rain through supernatural means, including dances, sacrifices, prayers, and petitioning of rain deities. Treated by the authors as a lower bound on actual prevalence.&lt;/p&gt;</description></item><item><title>Religion, Education, and the State</title><link>https://macropaperwarehouse.com/papers/religion-education-and-the-state/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/religion-education-and-the-state/</guid><description>&lt;p&gt;This paper studies how Indonesia&amp;rsquo;s Islamic education sector responded to one of the largest state-driven mass schooling expansions in history — the SD INPRES program (Sekolah Dasar Presidential Instruction) launched in 1973 — and whether that program achieved its secular nation-building objectives. The research question is three-part: Did Islamic schools enter or exit markets where the state built more primary schools? How did religious school choice shift across cohorts? And did the program advance secular identity formation among exposed individuals?&lt;/p&gt;
&lt;p&gt;The empirical setting is Indonesia in the 1970s onward. Under SD INPRES, the government used windfall oil revenues to build more than 61,000 primary schools between 1973 and 1980, allocating construction across districts proportional to the non-enrolled primary-school-age population. Because Islamic schools were historically more prevalent in underserved areas, this rule produced a strong positive correlation between SD INPRES intensity and pre-existing Islamic school density — the same markets where the state expanded were precisely those with the greatest Islamic education presence.&lt;/p&gt;
&lt;p&gt;The authors use several novel data sources: administrative registries covering nearly 220,000 secular and 160,000 Islamic schools with establishment dates; six rounds of the National Socioeconomic Survey (Susenas) from 2012–18; the Indonesia Family Life Survey (IFLS, 1993–2014); a 2018–19 curriculum timetable registry (SIAP) covering nearly 20% of madrasa; and a 2016 political/religious attitudes survey. Identification relies on difference-in-differences (DID) exploiting cross-district variation in SD INPRES intensity, the synthetic DID approach of Arkhangelsky et al. (2021) for robustness to violations of parallel trends, and a staggered village-level event study using the Borusyak et al. (2024) estimator.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. First, Islamic schools did not exit markets where the state expanded — they entered in greater numbers. A one standard deviation increase in SD INPRES construction led to 1.4 additional Islamic school entries per district above a mean of 1.9 per district in 1972. Entry was competitive at the primary level, where new madrasa (MI) entered at twice the baseline annual rate in the years immediately following INPRES construction, and strategic at the secondary level, where Islamic junior secondary schools (MTs) peaked 6–9 years after INPRES entry as graduates sought continued education. The Islamic sector financed this expansion through waqf (inalienable religious endowments), informal taxation (infaq, zakat), and revenues from a concurrent rice price spike; entry responses were stronger in villages with above-median waqf endowments and above-median potential rice yields.&lt;/p&gt;
&lt;p&gt;Second, rather than converging toward secular curricula, newly established Islamic schools in high-INPRES districts devoted more time to religious content. Each additional SD INPRES is associated with a 1.2 percentage point increase in the religious curriculum share among newly created madrasa, with increases of 1.3 and 2.4 percentage points at the primary and junior secondary levels respectively — the latter equaling 82% of the cross-school standard deviation. Some of this increase came at the expense of Pancasila/civic education and national language instruction.&lt;/p&gt;
&lt;p&gt;Third, while SD INPRES reduced Islamic primary school enrollment by roughly 7%, it increased overall Islamic school attendance: each additional SD INPRES increased the likelihood of attending any Islamic school by approximately 5%, as demand for secondary education outweighed substitution at the primary level. Female students exhibited stronger secondary-level demand effects, amplified in districts with a concurrent state ban on the Islamic veil in public schools.&lt;/p&gt;
&lt;p&gt;Fourth, SD INPRES did not advance its ideological objectives. In the 1977 and 1982 elections, Golkar&amp;rsquo;s vote share fell and the Islamic PPP&amp;rsquo;s rose by 0.5–1.0 percentage points per SD INPRES school in high-INPRES districts. Among exposed cohorts, SD INPRES did not increase Pancasila proficiency, national language use at home, or support for secular governance, but did increase Arabic literacy by approximately 3% per additional SD INPRES. Exposed cohorts also prayed more frequently, fasted more during Ramadan, gave more to charity, and expressed greater pilgrimage intentions. These religious patterns were transmitted to children of exposed cohorts, who were more likely to attend Islamic schools themselves.&lt;/p&gt;
&lt;p&gt;Q: What was the allocation rule for SD INPRES and why did it create confrontation with Islamic schools?
A: Presidential Instruction No. 10/1973 allocated school construction across districts proportional to the non-enrolled primary-school-age population in 1971. Because Islamic schools historically served underserved populations, this rule meant the state built more schools precisely where Islamic education was most prevalent. The paper shows graphically and in Table 1 that the number of SD INPRES schools built is strongly correlated with the pre-existing stock of Islamic schools, conditional on district population and enrollment.&lt;/p&gt;
&lt;p&gt;Q: How large was the Islamic sector&amp;rsquo;s entry response to SD INPRES at the district level?
A: In the standard DID specification (Table 2, panel a), a one standard deviation increase in SD INPRES schools led to 0.013 more Islamic schools per district-year per 1,000 children, equivalent to 1.4 additional Islamic school entries in the average district relative to a mean of 1.9 Islamic schools per district in 1972. The synthetic DID (panel b) delivers positive and slightly larger estimates, indicating the result is not an artifact of diverging pre-trends.&lt;/p&gt;
&lt;p&gt;Q: What was the timing of the Islamic sector entry response at the village level?
A: Using the Borusyak et al. (2024) estimator on a balanced panel from 1960 to 1999, the paper finds (Figure 4) that INPRES construction is followed by a jump in Islamic school entry. Primary madrasa (MI) entered at twice the baseline annual rate in the years immediately following INPRES construction and this elevated rate persisted for six years before reverting to baseline. Islamic junior secondary entry (MTs) peaked around years 6–9 after SD INPRES construction, consistent with newly graduated primary students seeking continued schooling.&lt;/p&gt;
&lt;p&gt;Q: How did the Islamic sector finance its expansion?
A: The sector relied on waqf endowments (inalienable religious land assets), informal faith-based contributions (infaq), and obligatory alms (zakat). Fortuitously, the initial year of SD INPRES coincided with a large spike in the global price of rice, Indonesia&amp;rsquo;s main agricultural commodity, boosting harvest revenues channeled through informal Islamic taxation. Table 3 shows that entry responses were significantly stronger in villages with above-median waqf endowments and above-median potential rice yields, and these heterogeneous effects did not arise in non-INPRES periods or for non-Islamic private schools. Survey data from 2007–13 further show higher rates of informal taxation in villages with Islamic schools built during this period.&lt;/p&gt;
&lt;p&gt;Q: Did Islamic schools converge toward secular curricula under competitive pressure from SD INPRES?
A: No. Table 4 shows that madrasa established in high-INPRES districts after 1972 devote more time to religious content, not less. Each additional SD INPRES is associated with a 1.2 percentage point increase in the share of classroom time devoted to religious subjects among newly created Islamic schools, with increases of 1.3 percentage points at the primary level and 2.4 percentage points at the junior secondary level — the latter equal to 82% of the cross-school standard deviation. Similar patterns hold for Arabic instruction, and the junior secondary increase comes partially at the expense of Pancasila/civic education and national language instruction.&lt;/p&gt;
&lt;p&gt;Q: Did curriculum differentiation responses vary with local religious ideology?
A: Yes. Appendix Table A.14 shows a stronger curriculum differentiation response in markets with greater historical support for conservative Islam, proxied by Islamic political party vote shares in the 1950s elections. The paper also constructs a school-name-based predicted ideology index using a ridge shrinkage estimator and finds (Appendix Table A.15) that madrasa entering high-INPRES districts after the program onset have a more religious ideology on this measure.&lt;/p&gt;
&lt;p&gt;Q: What happened to the formalization of the Islamic sector?
A: Figure 5 and Appendix Table A.6 show that formal madrasa entry increased as a share of all new school entry, while informal Islamic schools (pesantren, diniyah) declined as a share of all new schools and all new Islamic schools. This formalization mirrors the organizational structure of state schools (primary-to-secondary progression), facilitating switching between public and religious schools and providing option value to moderate but still religious families. Crucially, the newly entering formal madrasa introduced more religious curriculum than incumbent madrasa, so formalization did not reduce religious instruction.&lt;/p&gt;
&lt;p&gt;Q: What was the net effect of SD INPRES on Islamic school attendance?
A: Table 5 shows that SD INPRES reduced the likelihood of attending Islamic primary school by roughly 7% per additional SD INPRES school but increased Islamic secondary attendance, with the net effect being a roughly 5% increase in the likelihood of attending any Islamic school (column 4). This finding holds in both DID and synthetic DID. The IFLS validation (Appendix Table A.18) confirms decreased Islamic elementary attendance and increased Islamic junior secondary attendance, consistent with the Susenas results.&lt;/p&gt;
&lt;p&gt;Q: How does selection into secondary education affect the religious schooling results?
A: The authors address selection using parametric (Heckman 1976) and semiparametric (Newey 2009) selection-correction procedures, using exposure to a 1960s pilot compulsory schooling program as an exclusion restriction. Table 6, panels (c) and (d), show that selection-adjusted estimates are broadly consistent with unadjusted estimates, with similar signs and magnitudes. The selection-corrected estimates approximately identify a local average treatment effect among compliers: those induced to attend elementary school were less likely to attend Islamic elementary; those induced to continue to secondary were more likely to attend Islamic secondary.&lt;/p&gt;
&lt;p&gt;Q: How did gender shape the effects of SD INPRES on religious school choice?
A: Table 7 shows that SD INPRES had more limited impacts on total schooling for women than men (consistent with Duflo 2001) but that the secondary-level demand effect toward Islamic schools was stronger for women. Table 8 shows that within high-INPRES areas, the SD INPRES-induced increase in Islamic secondary education is three times larger for women in districts with greater exposure to the 1982 state ban on the Islamic veil in public schools, and this differential is specific to Islamic schooling rather than total schooling.&lt;/p&gt;
&lt;p&gt;Q: Did SD INPRES strengthen or weaken the secular ruling regime&amp;rsquo;s political standing?
A: It weakened it. Table 10 shows that in the 1977 and 1982 elections, Golkar&amp;rsquo;s vote share decreased and the Islamic PPP&amp;rsquo;s vote share increased in high-INPRES districts, in the range of 0.5–1.0 percentage points per SD INPRES school. This represents a 1.5–3.0% change in PPP vote share and a 0.5–1.0% change in Golkar vote share per standard deviation in SD INPRES intensity. The PPP gained most in areas where SD INPRES had the greatest potential to draw students away from Islamic schools.&lt;/p&gt;
&lt;p&gt;Q: Did SD INPRES produce a secular ideological shift among exposed cohorts?
A: No. Table 11 shows that SD INPRES did not increase self-reported Pancasila proficiency, national language use at home, national language literacy, or attitudes in favor of secular governance. By contrast, Arabic literacy increased by approximately 3% per additional SD INPRES among exposed cohorts, indicating that Islamic schooling exposure rather than secular schooling drove literacy gains in that language.&lt;/p&gt;
&lt;p&gt;Q: Did SD INPRES increase religiosity among exposed cohorts?
A: Yes. Table 12 shows that SD INPRES increased prayer frequency, fasting during Ramadan, charitable giving, and pilgrimage intentions among exposed cohorts. These effects on prayer and fasting are stronger among women, consistent with the stronger shift toward Islamic secondary schooling found in Table 7. These outcomes are consistent with greater exposure to Islamic education increasing religiosity rather than the secular curriculum reducing it.&lt;/p&gt;
&lt;p&gt;Q: Were the effects on religious identity and Arabic literacy transmitted to the next generation?
A: Yes. Table 13 shows that SD INPRES increased Arabic literacy among the children of exposed cohorts, and that children of exposed cohorts were more likely to attend Islamic schools themselves. These intergenerational results confirm that the preference for Islamic education instilled during the SD INPRES era persisted into the next generation rather than converging toward secular norms over time.&lt;/p&gt;
&lt;p&gt;Q: What are the key robustness checks for the school entry results?
A: Several checks support causal interpretation. The Roth and Rambachan (2022) procedure finds no systematic pre-trends in the standard DID. Historical Podes data from 1980, 1983, 1990, and 1993 confirm the post-1973 increase in Islamic school entry, addressing survival bias in the 2019 registry. Results are robust to allowing differential trends in waqf endowments, Muslim population share, Islamic party vote shares, historical Arab immigration, Islamist insurgency, and Transmigration resettlement. The heterogeneous entry responses by waqf and rice yield do not appear in non-INPRES periods or for non-Islamic schools.&lt;/p&gt;
&lt;p&gt;Q: What do the results imply for the political economy of education reform more broadly?
A: The paper argues that state capacity to homogenize culture through education is limited when strong non-state actors can mobilize their own resources and provide differentiated alternatives. Rather than crowding out religious schools, state expansion triggered competitive entry, curriculum differentiation, and formalization in the religious sector, producing an equilibrium where both sectors expanded simultaneously with distinct clienteles. The findings imply that the long-run cultural effects of education programs cannot be evaluated without accounting for equilibrium responses by competing non-state providers.&lt;/p&gt;
&lt;p&gt;SD INPRES (Sekolah Dasar Presidential Instruction): Indonesia&amp;rsquo;s 1973 mass public primary school construction program, financed by oil windfalls, which built more than 61,000 elementary schools between 1973 and 1980 by allocating schools to districts proportional to the non-enrolled child population; the program&amp;rsquo;s explicitly secular nation-building objectives brought it into direct confrontation with the Islamic education sector.&lt;/p&gt;
&lt;p&gt;Waqf: Inalienable Islamic religious endowments — of land, agricultural assets, or other property — that under Islamic law can only be used for religious or charitable purposes and cannot be seized or repurposed by the state; in this paper, the pre-existing waqf base in a village serves both as a long-run financing mechanism for Islamic school construction and as an index of Islamic sector organizational capacity.&lt;/p&gt;
&lt;p&gt;Madrasa: Formal day Islamic schools operating at the same primary-to-secondary grade levels as secular state schools, teaching standard academic subjects alongside a religious curriculum (including Islamic law, doctrine, ethics, Qur&amp;rsquo;an, Arabic, and history of the Prophets) that averages 26% of total instruction hours; distinct from the more informal pesantren (boarding schools) and madrasa diniyah (afternoon Qur&amp;rsquo;anic study schools).&lt;/p&gt;
&lt;p&gt;Curriculum differentiation: The strategy by which newly entering madrasa in high-INPRES districts increased the share of classroom time devoted to religious and Arabic instruction rather than converging toward the secular state curriculum; measured as classroom hours devoted to Islamic subjects, Arabic, Pancasila/civic education, and national language instruction from 2018–19 SIAP timetable data.&lt;/p&gt;
&lt;p&gt;Pancasila: The official secular nationalist ideology of the Indonesian state, consisting of five principles (monotheism, humanitarianism, national unity, democracy, and social justice) intended to transcend ethnic and religious divisions; SD INPRES sought to transmit Pancasila through civic education and national language instruction as part of its homogenizing nation-building agenda.&lt;/p&gt;
&lt;p&gt;Synthetic difference-in-differences (SDID): The Arkhangelsky et al. (2021) estimator used throughout the paper, which reweights and matches pre-INPRES trends in Islamic school construction across high- and low-INPRES exposure districts to deliver estimates more robust than standard DID to violations of parallel trends; applied with a binary treatment indicator (districts above the 51st percentile in INPRES intensity).&lt;/p&gt;
&lt;p&gt;Formalization: The documented shift in the composition of the Islamic sector after SD INPRES, whereby formal madrasa (organized along the same grade-level progression as state schools) increased as a share of all new Islamic school entry while informal pesantren and diniyah declined as a share; interpreted as a competitive response that expanded parental option value without sacrificing religious instruction intensity.&lt;/p&gt;</description></item><item><title>Revolutionary Transition: Inheritance Change and Fertility Decline</title><link>https://macropaperwarehouse.com/papers/revolutionary-transition-inheritance-change-and-fertility-decline/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/revolutionary-transition-inheritance-change-and-fertility-decline/</guid><description>&lt;p&gt;Gay, Gobbi, and Goñi test Le Play&amp;rsquo;s (1875) hypothesis that the French Revolution contributed to France&amp;rsquo;s early fertility decline by abolishing impartible inheritance. In 1793, a series of decrees culminating in the Loi de Nivôse (January 6, 1794) abolished testamentary rights and imposed equal partition of assets among all children — partible inheritance — across France, overriding the mosaic of local customs and written laws that had governed inheritance in the Ancien Régime.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central argument is that this reform reduced the economic incentive to have children through indivisibility constraints in agricultural land. Under impartible inheritance, land passed to a single heir undivided, keeping plots above the subsistence productivity threshold even at high fertility. Under partible inheritance, each additional child fragments the land further, potentially pushing plots below the minimum productive size, so households face a strong incentive to limit fertility. A Stone-Geary production function with a minimum land threshold L̄ formalizes this mechanism: when landholdings fall in the binding range (L̄ &amp;lt; L &amp;lt; L̃), fertility is strictly higher under impartible than under partible inheritance.&lt;/p&gt;
&lt;p&gt;The authors construct the first complete map of inheritance rules across France&amp;rsquo;s 435 judicial districts as of 1789, classifying each along two dimensions: partible versus impartible, and whether women were included or excluded. This atlas draws on Brette&amp;rsquo;s (1904) Atlas des Bailliages and the Nouveau Coutumier Général (Bourdot de Richebourg 1724), covering 141 distinct customs. Treatment is defined as municipalities under impartible inheritance before 1793 whose system was altered by the reforms; control municipalities were already under partible inheritance.&lt;/p&gt;
&lt;p&gt;The main identification strategy is a difference-in-differences (DD) design comparing women with varying lengths of remaining fertile years after 1793 — from 0 for women aged 40+ at the reform to 25 for women aged 15 or younger — across treated and untreated municipalities. This is augmented by a regression-discontinuity difference-in-differences (RD-DD) design exploiting sharp discontinuities at judicial district borders. Two independent datasets are used: the Enquête Louis Henry (34,812 women in 39 rural municipalities, family-reconstitution method) and Geni.com crowdsourced genealogies (11,649 women across 2,966 locations after the Blanc 2023 horizontal restriction).&lt;/p&gt;
&lt;p&gt;Each additional fertile year of exposure to the 1793 reforms reduced completed fertility by approximately 1 percent. Over the full 25-year fertile cycle, this corresponds to a reduction of roughly 0.7 children, or 24 percent relative to the pre-reform mean of 2.92 surviving children in treated areas. This magnitude equals the entire pre-reform fertility gap between impartible- and partible-inheritance areas (2.9 versus 2.2 children), meaning the reforms closed this gap entirely. DD and RD-DD estimates are similar and not statistically distinguishable from each other, and results replicate across both datasets. Results hold on both the extensive margin (childlessness) and intensive margin (fertility of mothers).&lt;/p&gt;
&lt;p&gt;The mechanism is most relevant where smallholder landownership is widespread. France — where 40–80 percent of households owned land at the eve of the Revolution — meets this condition. England and Prussia, with more concentrated landownership, would not be expected to show the same response because the indivisibility constraint would not bind even after partition.&lt;/p&gt;
&lt;p&gt;Q: What was France&amp;rsquo;s inheritance system before the Revolution, and how heterogeneous was it?
A: Before 1793, inheritance was governed by 141 distinct customary and written laws applied within 435 judicial districts. The country was broadly divided between the customary-law north (Pays de droit coutumier) and the Roman written-law south (Pays de droit écrit), with substantial local variation within regions. Systems ranged from strictly partible (equal division among all offspring) to impartible (primogeniture, ultimogeniture, or unigeniture). Systems also varied in whether women could inherit or received only a dowry. This geographic variation — rooted in the laws of Germanic peoples after the fall of Rome in 476 CE — is exogenous to late eighteenth-century economic conditions and provides the identifying variation for the paper.&lt;/p&gt;
&lt;p&gt;Q: What exactly did the 1793 reforms change, and were they enforced?
A: The Loi de Nivôse an II (January 6, 1794) abolished testamentary rights entirely and mandated equal partition of assets among all children, including women, throughout France. The reforms came unexpectedly — only 8 of 571 cahiers de doléances analyzed by Goy (1988) mentioned inheritance — and were motivated by the equality principle, legal unification, and the fear that revolutionary sympathizers would be disinherited (Lataste et al. 1901). Offspring quickly asserted their new rights, and by the late 1790s inheritance disputes were the most common cases before family tribunals (Desan 1997; Poumarède 2011).&lt;/p&gt;
&lt;p&gt;Q: What is the model&amp;rsquo;s core mechanism linking inheritance reform to fertility decline?
A: The model uses a Stone-Geary production function with a minimum land threshold L̄ below which output falls to zero. Under impartible inheritance, land passes undivided to a single heir, keeping the farm above L̄ regardless of family size. Under partible inheritance, each child receives an equal share, so adding children risks fragmenting plots below L̄ — a powerful incentive to limit family size. The fertility gap between impartible and partible households is at its maximum when landholdings fall in the intermediate range (L̄ &amp;lt; L &amp;lt; L̃) where the constraint is binding. As land size increases, the constraint becomes less binding but the positive fertility differential persists.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s main quantitative estimate of the reform&amp;rsquo;s effect on completed fertility?
A: Each additional fertile year of exposure to the 1793 reforms reduced completed fertility by approximately 1 percent. Over the full 25-year fertile cycle (ages 15–40), this implies a reduction of roughly 0.7 children, or 24 percent relative to the pre-reform mean of 2.92 surviving children in treated areas. This is nearly identical to the pre-existing fertility gap between impartible- and partible-inheritance areas (0.7 children: 2.9 versus 2.2 surviving children), implying the reforms effectively eliminated the fertility differential.&lt;/p&gt;
&lt;p&gt;Q: Are the DD and RD-DD estimates consistent with each other, and do both datasets agree?
A: Yes. The DD and RD-DD estimates are similar and not statistically different from each other. The RD-DD design compares women born close to judicial district borders where inheritance rules differed, before and after 1793, exploiting the sharp spatial discontinuity at those borders. Consistency across these two designs — which rely on different identifying assumptions — strengthens causal interpretation. Results are also consistent across the Enquête Louis Henry (family-reconstitution) and Geni.com (crowdsourced genealogies) datasets, which are produced by fundamentally different methodologies.&lt;/p&gt;
&lt;p&gt;Q: How do the authors verify the parallel trends assumption?
A: Figure 6 shows that for cohorts who completed their fertile cycle before 1793, fertility trended downward in parallel across partible- and impartible-inheritance areas: a constant gap of approximately 0.7 children was maintained from women born in the early 1700s (3 versus 2.3 children) through women born in the early 1750s (2.7 versus 2.0 children), the last cohorts to complete fertility before the reforms. The convergence — from 0.7 to 0 children — only begins among cohorts fertile after 1793. The authors also include flexible trend controls interacted with municipality-level religiosity, political support for the Revolution, proximity to administrative centers, and wheat prices, and confirm the main estimate is robust.&lt;/p&gt;
&lt;p&gt;Q: What role did the extension of inheritance rights to women play?
A: The extension of rights to women was a companion mechanism distinct from abolishing impartible inheritance. Beyond increasing the number of heirs (which directly reduces land per heir), the right to inherit improves a woman&amp;rsquo;s outside option and postpones entry into marriage, following de Moor and van Zanden (2010). The DD and RD-DD estimates suggest that including women in inheritance and abolishing impartible inheritance had similar effects on fertility. The paper treats these as separate but reinforcing channels.&lt;/p&gt;
&lt;p&gt;Q: How do the authors address potential confounders — mortality, migration, and economic conditions?
A: On mortality: child mortality did not evolve differently after 1793 across areas with different inheritance rules (Appendix Table A3), and baseline adult mortality (age at death, probability of dying before completing the fertile cycle) was balanced across treated and control areas (Table 1). On migration: the authors explicitly rule out that results are driven by migration. On economic conditions: municipality-specific decade-average wheat prices (Ridolfi 2019) are included as controls for local Malthusian dynamics, and results are robust to their inclusion.&lt;/p&gt;
&lt;p&gt;Q: What do the balance tests show?
A: Panel A of Table 1 shows that before the reforms, areas with impartible versus partible inheritance were balanced on 9 of 11 individual-level characteristics — including husband and wife age at death, probability of dying before completing the fertile cycle, probability that parents-in-law were alive at marriage, literacy, data accuracy, and age at marriage. The only systematic pre-reform difference was fertility itself (0.7 children). Municipality-level climatic variables, soil suitability, and proxies for mortality uncertainty were also balanced. This is consistent with the origins of these systems in post-Roman Germanic law, which are unrelated to late eighteenth-century economic conditions.&lt;/p&gt;
&lt;p&gt;Q: What robustness checks are reported?
A: The authors report: (1) permutation tests reshuffling treatment exposure across women and municipalities; (2) non-linear treatment effects across cohorts, showing the heterogeneity required to explain away the baseline estimate is implausibly large per de Chaisemartin and d&amp;rsquo;Haultfoeuille (2020); (3) exclusion of outlier municipalities; (4) a placebo test for cohorts who completed their fertile cycle before 1793; (5) robustness to alternative sample definitions, treatment definitions, outcome variables, and control groups; (6) Cummins (2020) first-name repetition technique to correct for under-reported child deaths in Henry; (7) terrain characteristics including climatic and soil suitability (Galor and Özak 2016) and ruggedness (Nunn and Puga 2012); (8) for RD-DD: alternative bandwidths, running variable specifications, kernel functions, samples, and border-segment fixed effects. All checks support the main finding.&lt;/p&gt;
&lt;p&gt;Q: Why did France experience a fertility decline from inheritance reform while other countries with similar reforms did not?
A: The model rationalizes this through landownership structure. The fertility-reducing mechanism operates through indivisibility constraints that bind only when landholdings are small and fragmented — as in France, where 40–80 percent of households owned their land and plots were small. Where landownership is concentrated (England, Prussia), land per heir remains above L̄ even after partible division, so the indivisibility constraint is non-binding and fertility is unaffected by the reform. This provides a structural reason why France&amp;rsquo;s particular agrarian structure made it uniquely susceptible to this mechanism.&lt;/p&gt;
&lt;p&gt;Q: What is the broader historical significance for understanding France&amp;rsquo;s early demographic transition?
A: France&amp;rsquo;s fertility decline began roughly 50 years before industrialization, making it anomalous relative to standard quantity-quality tradeoff theories linking fertility decline to technological progress and rising returns to human capital. The 1793 reforms provide a legal-institutional explanation for the sharp post-Revolution acceleration visible in Figure 1, which is difficult to attribute to slowly-evolving cultural factors or human capital considerations not yet operative. The estimates imply the reforms brought large impartible-inheritance areas to the low-fertility regime that already characterized partible-inheritance areas, thus sharply accelerating the national transition.&lt;/p&gt;
&lt;p&gt;Impartible inheritance: A system under which parents could designate a single heir (through primogeniture, ultimogeniture, or unigeniture) to receive the bulk of the family estate, preventing fragmentation of wealth; in pre-revolutionary France this was associated with extended family households and higher fertility (2.9 surviving children on average) relative to partible areas (2.2).&lt;/p&gt;
&lt;p&gt;Partible inheritance: A system under which family wealth was divided equally among all offspring upon death; in the paper&amp;rsquo;s model this creates an incentive to limit fertility to prevent land fragmentation below the subsistence productivity threshold L̄.&lt;/p&gt;
&lt;p&gt;Indivisibility constraint (land threshold L̄): In the Stone-Geary production function, a minimum land input below which agricultural output falls to zero; this is the mechanism through which partible inheritance generates fertility-limiting incentives, since dividing a small plot among many heirs risks crossing L̄ into zero production.&lt;/p&gt;
&lt;p&gt;Difference-in-differences (DD) exposure design: The paper&amp;rsquo;s main identification strategy, using remaining fertile years after 1793 as a continuous treatment-intensity variable (0 for cohorts past fertility at the reform date, up to 25 for cohorts entirely within their fertile years), compared between treated municipalities (impartible → partible) and control municipalities (already partible).&lt;/p&gt;
&lt;p&gt;Regression-discontinuity difference-in-differences (RD-DD): An augmented design exploiting the sharp geographic discontinuity at borders between judicial districts with different pre-reform inheritance rules, comparing outcomes on both sides before and after 1793, to address smooth unobserved confounders.&lt;/p&gt;
&lt;p&gt;Completed fertility (net): The number of children surviving to age six, preferred over total births because child mortality before 1800 was high (1–1.5 children per mother did not survive to age six per Houdaille 1984), making net fertility the more economically meaningful measure for inheritance and bequest decisions.&lt;/p&gt;
&lt;p&gt;Horizontal restriction: A sampling correction applied to crowdsourced genealogical data (Blanc 2023a) that retains an observation only if at least one of the four preceding generations has more than one recorded offspring, correcting for the over-representation of single-child families that arises because Geni users tend to record direct ancestors rather than collateral relatives.&lt;/p&gt;</description></item><item><title>Slum Upgrading and Long-Run Urban Development: Evidence from Indonesia</title><link>https://macropaperwarehouse.com/papers/slum-upgrading-and-long-run-urban-development-evidence-from-indonesia/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/slum-upgrading-and-long-run-urban-development-evidence-from-indonesia/</guid><description>&lt;p&gt;This paper estimates the long-term causal effects of the Kampung Improvement Program (KIP), one of the world&amp;rsquo;s largest slum upgrading programs, on urban development in Jakarta, Indonesia. KIP ran from 1969 to 1984 across three staggered waves (Pelita I-III), covered 110 square kilometers (25% of Jakarta&amp;rsquo;s area), and served approximately 5 million residents at a total cost of roughly $500 million (2015 USD). The program provided basic physical upgrades — paved roads and footpaths, sanitation and drainage, and community buildings such as schools and health clinics — along with a verbal non-eviction guarantee for 15 years. Residents were not relocated.&lt;/p&gt;
&lt;p&gt;The central research question is whether preserving slums through upgrading entails long-run dynamic inefficiency: as Jakarta formalizes, do KIP areas lag behind non-KIP areas in ways that generate opportunity costs from land misallocation?&lt;/p&gt;
&lt;p&gt;The authors assemble high-resolution data on KIP policy boundaries, current assessed land values (nearly 20,000 sub-blocks), building heights from a novel photographic survey of 19,518 pixels stratified across Jakarta, and multiple novel measures of informality — a rank-based photographic index (0 to 4), an attributes-based index across fifteen binary characteristics, and administrative data on unregistered land-parcel titles. They also use digitized historical maps from 1937 and 1959 to identify pre-KIP kampung boundaries.&lt;/p&gt;
&lt;p&gt;Two empirical strategies address program selection bias (KIP planners prioritized the worst-condition kampungs first). The first restricts the sample to historical kampungs that existed before KIP and includes locality fixed effects, comparing treated kampungs against nearby untreated ones within the same neighborhood. The second is a boundary discontinuity design (BDD) comparing observations within 200 meters of KIP boundaries. Both strategies include eighteen predetermined controls for historical landmarks, infrastructure, and topography including flood proneness.&lt;/p&gt;
&lt;p&gt;Average effects (robust across both strategies): KIP areas today have land values approximately 14-17 log points (roughly 15%) lower than observably equivalent non-KIP areas, and are about 8-12 percentage points less likely to contain buildings taller than three floors — half the control-group mean of 0.24. KIP areas are more informal across all three informality metrics: the rank-based index is higher by 0.29 standard deviations, the attributes-based index by 0.05 SD units, and the share of unregistered parcels is 3 percentage points higher. Building heights corroborate the land-value finding: imputing the hedonic value of missing tall buildings in KIP accounts for approximately 90% of the aggregate land-value impact ($2.2 billion of $2.4 billion).&lt;/p&gt;
&lt;p&gt;Heterogeneity by real estate potential is a central finding. The authors construct a predicted land index for 2,058 hamlets in Jakarta using non-KIP land values. In the lowest quintile (Q5), KIP areas show a positive and statistically significant effect of +10 log points on land values, consistent with direct capitalization of the upgrades. This effect reverses in higher-potential areas: the estimate reaches -28 log points in Q2 and -30 log points in Q1, as non-KIP neighborhoods formalize while KIP areas lag.&lt;/p&gt;
&lt;p&gt;Surplus calculations integrating land values, building heights, horizontal built-up coverage (35% for KIP vs. 18% for non-KIP), and demand and supply elasticities reveal that 90% of total surplus losses are concentrated in the top two quintiles (Q1 and Q2), which comprise 47% of KIP&amp;rsquo;s coverage area. In Q1, KIP surplus is lower by $2,369 per square meter; in Q2, the gap is $1,044 per square meter. In the bottom two quintiles, KIP delivers greater surplus (up to +$347 per square meter in Q5), covering an estimated 3 million residents across 57 square kilometers.&lt;/p&gt;
&lt;p&gt;Mechanisms consistent with delayed formalization include significantly higher population density in KIP areas (+33 log points, or 39%) and greater land fragmentation (+9 parcels per pixel relative to a non-KIP mean of 19), both of which raise relocation and land assembly costs. The original KIP investments show no differential effect by type or intensity after four decades, consistent with their 15-year projected useful life. Endogenous sorting is ruled out as a confounder: if anything, educational attainment is slightly higher in KIP areas.&lt;/p&gt;
&lt;p&gt;Q: What is the Kampung Improvement Program (KIP) and what did it provide?
A: KIP was a slum upgrading program implemented in Jakarta, Indonesia from 1969 to 1984 across three five-year plan waves (Pelita I, II, III). It covered 110 square kilometers and 5 million residents at a total cost of approximately $500 million (2015 USD). The program provided three categories of basic physical improvements — vehicular and pedestrian road access, sanitation and drainage infrastructure, and community buildings (schools, health clinics) — along with a verbal non-eviction guarantee for 15 years. Crucially, upgrades were designed to be basic, with a planned useful life of only 15 years, to avoid attracting higher-income groups.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question and theoretical concern motivating the paper?
A: The paper asks whether slum upgrading programs, while immediately beneficial to residents, entail dynamic inefficiency by delaying formalization as cities develop. The concern is that preserving slums through upgrades and non-eviction guarantees can create opportunity costs from land misallocation when surrounding areas formalize and redevelop into higher-value formal structures. This is framed as a trade-off between the direct welfare benefits of upgrading (affordable in-situ housing for millions) and the long-run costs to urban land productivity.&lt;/p&gt;
&lt;p&gt;Q: How does the paper address the selection bias problem — KIP targeted the worst-condition kampungs first?
A: Two complementary strategies are used. First, the historical kampung specification restricts the sample to areas that were kampungs before KIP (from 1937 and 1959 maps) and includes locality fixed effects, so treated and control units are compared within the same neighborhood and share the same real estate market by assumption. Second, a boundary discontinuity design (BDD) compares observations within 200 meters of KIP boundaries with boundary fixed effects and quadratic distance controls. A falsification test using sequential KIP waves confirms the approach: the raw data shows a monotonic pattern (Wave I worst: -0.40 log points, Wave II: -0.29, Wave III: -0.17) consistent with selection bias, but this pattern disappears in the historical kampung specification (Wave I: -0.13, Wave II: -0.11, Wave III: -0.14), supporting the identification assumption.&lt;/p&gt;
&lt;p&gt;Q: What are the average effects of KIP on land values and building heights?
A: In the historical kampung specification, KIP areas have land values 14 log points (approximately 15%) lower than non-KIP historical kampungs within the same locality. The BDD estimate is similar at -17 log points. For building heights, KIP areas are 12 percentage points less likely to contain a building taller than three floors in the historical kampung sample (8 percentage points in the BDD), relative to a non-KIP control mean of 0.24 — meaning KIP areas are roughly half as likely to have tall buildings. The average effect on floors is -1.6 floors, relative to a control mean of 5 floors.&lt;/p&gt;
&lt;p&gt;Q: How do the authors validate that land value estimates are not distorted by measurement error in informal areas?
A: The authors impute the hedonic value of missing tall buildings in KIP using a hedonic regression estimated solely on non-KIP historical kampungs. KIP areas have 145 fewer buildings with more than ten floors; combined with a 57% price premium for tall buildings (relative to a base price of 13.4 million Rupiahs per square meter), the implied land value loss from missing buildings above ten floors is approximately $1.3 billion, and from buildings between four and ten floors is $0.9 billion, for a total imputed effect of $2.2 billion. This accounts for approximately 90% of the aggregate land value impact from the historical kampung specification ($2.4 billion), assuaging concerns that lower measured land values in KIP reflect data quality differences rather than true price gaps.&lt;/p&gt;
&lt;p&gt;Q: How does the KIP effect vary across the distribution of real estate potential?
A: The authors construct a predicted land index for 2,058 Jakarta hamlets by regressing non-KIP log land values on hamlet fixed effects, then rank hamlets into quintiles. In Q5 (lowest predicted land values, least likely to formalize), KIP areas show a statistically significant positive effect of +10 log points on land values, consistent with direct capitalization of the upgrades. Moving to higher-potential areas, the effect attenuates and reverses: it is -28 log points in Q2 and -30 log points in Q1, where non-KIP areas have formalized. This cross-sectional pattern traces out the dynamic inefficiency predicted by theory.&lt;/p&gt;
&lt;p&gt;Q: What informality measures does the paper construct and what do they show?
A: The paper constructs three complementary informality metrics. First, a rank-based photographic index (0 = very formal, 4 = very informal) coded by two trained Jakarta-based research assistants from approximately 28,000 hand-coded photographs, with inter-rater correlation of 0.78. Second, an attributes-based index averaging fifteen binary characteristics across vehicular access, neighborhood appearance, and structural permanence, standardized to a z-score. Third, the area share of unregistered land parcels from the Indonesian National Land Agency&amp;rsquo;s 2020 digital land maps. KIP areas score higher on all three: the rank-based index is higher by 0.29 SD units, the attributes-based index by 0.05 SD units, and the unregistered parcel share is higher by 3 percentage points.&lt;/p&gt;
&lt;p&gt;Q: What mechanisms explain why KIP areas remain informal and have lower land values?
A: The paper identifies three mutually reinforcing mechanisms. First, KIP areas have significantly higher population density (+33 log points or 39% in the historical kampung sample, equivalent to 51 more people per pixel), which raises relocation costs. Second, KIP areas have greater land fragmentation, with 9 more parcels per pixel relative to a non-KIP mean of 19, exacerbating holdout problems during land assembly; a back-of-the-envelope calculation attributes a 9% land value effect (60% of the total 15% effect) to this channel. Third, the verbal non-eviction guarantees and improved conditions likely strengthened residents&amp;rsquo; tenure perceptions and encouraged them to stay, leading to sub-division of parcels over time. The original KIP investments show no differential effect by type after four decades, consistent with their designed 15-year useful life, and KIP areas have similar access to public amenities today.&lt;/p&gt;
&lt;p&gt;Q: How does the paper calculate surplus and what are the results?
A: The surplus framework compares KIP (informal, tends to stay informal) against non-KIP counterfactuals (more likely formal) on three dimensions: non-KIP areas have (i) higher land values, (ii) taller structures, but (iii) lower horizontal built-up coverage than slums (18% vs. 35% for KIP). Consumer surplus uses a linear demand approximation with elasticity of 0.2 for non-KIP and 0.16 for KIP (backed out from differences in housing budget shares). Producer surplus integrates a Cobb-Douglas supply curve with elasticities of 1.4 (formal) and 1.3 (informal). In Q1, KIP property value is $1,873 per square meter vs. $3,098 for non-KIP, a difference of $1,225 in value terms and $2,369 in surplus terms. The surplus gap falls to $1,044 in Q2, and halves again in Q3, becoming positive (+$347 per square meter) in Q5. Ninety percent of total surplus losses are concentrated in Q1 and Q2, which cover 47% of KIP&amp;rsquo;s area.&lt;/p&gt;
&lt;p&gt;Q: What do the case studies of kampung clearances illustrate?
A: Three Jakarta kampungs cleared in 2015-2016 are examined. Kampung Bukit Duri (Q5, lowest real estate potential) shows a surplus difference of +$572 per square meter in favor of KIP — meaning clearance there is socially inefficient. Kali Pessangrahan (Q3) shows a surplus difference of -$307. Kalijodo (Q2) shows -$910 per square meter, suggesting sizable societal gains from formalization. However, even in Kalijodo, residents were relocated 24 km away to Marunda (a Q5 area), where consumer surplus is only 46% of Kalijodo&amp;rsquo;s — illustrating that societal gains from formalization do not automatically translate into Pareto improvements for evicted residents.&lt;/p&gt;
&lt;p&gt;Q: What robustness checks address alternative explanations?
A: The paper runs several tests. A placebo BDD using 45 non-KIP historical kampung boundaries finds no significant discontinuity, ruling out the hypothesis that slums generically have persistently lower land values. Bandwidth robustness shows consistent BDD estimates from 150 to 500 meters. Tests for spatial spillovers find no spatial decay pattern in land values near KIP boundaries, consistent with the prevalence of gated communities in formal Jakarta minimizing neighborhood contamination. Endogenous sorting is examined using 2010 Census data on 10 million individuals: educational attainment is slightly higher in KIP, and in-migration is slightly lower (1-2 percentage points below mean) with migrants having slightly more years of schooling — both inconsistent with an explanation based on low-skill sorting into KIP. Direct congestion effects from population density are also ruled out by estimating spatial decay around 45 dense non-KIP informal hamlets, finding no decay large enough to explain the land-value effects.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications for slum upgrading in other developing countries?
A: The paper&amp;rsquo;s framework suggests that slum upgrading&amp;rsquo;s cost-benefit balance depends critically on where the upgraded area sits in the real estate potential distribution. In low-potential areas (bottom quintiles of the land index), upgrading delivers net surplus even decades later and implicitly provides affordable housing at scale to millions of residents. In high-potential areas (top quintiles), the opportunity costs from delayed formalization can be large — up to $2,369 per square meter in surplus terms — and the paper suggests that stronger land market institutions to share surplus with informal residents could partially mitigate these costs. The paper also notes that formalization involves complex institutional and political challenges: relocating millions of kampung residents is logistically difficult, compensation is frequently inadequate or absent, and land assembly faces severe holdout problems.&lt;/p&gt;
&lt;p&gt;Dynamic inefficiency in cities: The phenomenon, in the context of this paper, whereby preserving informal slum settlements through upgrading delays their formalization, generating opportunity costs from land misallocation as surrounding formal areas develop. Distinguished from static inefficiency: KIP may raise resident welfare while simultaneously reducing aggregate land productivity.&lt;/p&gt;
&lt;p&gt;Slum upgrading: A policy providing basic public goods improvements (roads, sanitation, community buildings) and tenure security (typically verbal non-eviction guarantees) to existing slum residents in situ, without relocating them. Contrasted with formalization (redevelopment) and sites-and-services programs.&lt;/p&gt;
&lt;p&gt;Boundary discontinuity design (BDD): The paper&amp;rsquo;s second identification strategy, comparing outcomes for observations within 200 meters on either side of KIP program boundaries, with boundary fixed effects and quadratic distance controls, under the assumption that absent KIP, unobserved real estate potential varies smoothly at program boundaries.&lt;/p&gt;
&lt;p&gt;Predicted land index: A hamlet-level index constructed by regressing non-KIP log land values on hamlet fixed effects across 2,058 Jakarta hamlets, used to proxy real estate market potential and rank neighborhoods into quintiles from highest (Q1) to lowest (Q5) development stage.&lt;/p&gt;
&lt;p&gt;Informal surplus: The surplus generated within the informal housing sector, including built-up volume from high horizontal coverage (35% for KIP kampungs) and low-cost informal structures, which is destroyed upon formalization and must be weighed against the gains from taller, higher-value formal developments.&lt;/p&gt;
&lt;p&gt;Land fragmentation: The number of distinct land parcels per unit area (pixel), measured from Jakarta&amp;rsquo;s 2011 cadastral maps. Higher fragmentation exacerbates holdout problems in land assembly, raising the cost of redevelopment and contributing to delayed formalization.&lt;/p&gt;
&lt;p&gt;Source text origin: A classification in the paper&amp;rsquo;s summarization pipeline indicating whether the paper text derives from a full PDF or open-access HTML (permitting summarization) versus abstract-only text (which blocks summarization). All claims in this summary derive from the full paper text.&lt;/p&gt;</description></item><item><title>State Capacity as an Organizational Problem</title><link>https://macropaperwarehouse.com/papers/state-capacity-as-an-organizational-problem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/state-capacity-as-an-organizational-problem/</guid><description>&lt;p&gt;Mastrorocco and Teso study how the internal organization of a state evolves during national development, framing state capacity as an organizational — specifically a principal-agent — problem. Using a new micro-database covering the U.S. federal bureaucracy from 1817 to 1905, they ask: once rulers have incentives to build a state apparatus, how do they organize it to perform its functions across a vast territory, and what drives transitions between organizational forms?&lt;/p&gt;
&lt;p&gt;The dataset is constructed from every issue of the Official Register of the United States published between 1817 and 1905 (44 biennial volumes, 15,801 pages digitized). It records full name, state of birth, state of appointment, occupation, salary, department, office, and location for 304,410 unique federal employees across 810,942 employee-year observations. The authors reconstruct the bureaucracy&amp;rsquo;s four-layer hierarchy (department → office/bureau → division → local office), link employees over time to track careers, categorize all 11,930 occupation codes into five tiers, and geo-code 9,651 places of employment to 1890 county boundaries.&lt;/p&gt;
&lt;p&gt;The paper first documents three sets of descriptive facts. On growth: the federal workforce expanded very slowly before the 1860s and then rapidly, with geographic expansion accounting for none of state growth before 1859 but roughly 29% after. On location: state presence responded positively to local manufacturing activity (a one standard deviation increase in manufacturing employment share raises presence probability by 1.3 percentage points), but distance from Washington DC significantly attenuated this relationship in 1817–1859 and not in 1861–1905. On organization: before the 1860s, employee turnover was high and spiked sharply at presidential transitions (reaching 72% of employees departing in 1861), supervisors&amp;rsquo; departures strongly predicted subordinates&amp;rsquo; departures (a one-for-one supervisor exit raised subordinate turnover probability by 37% pre-1841), and managerial delegation outside DC was stagnant or declining. After the 1860s, turnover trended down (35% at the 1897 transition), the supervisor-subordinate career link weakened materially, and field managers tripled relative to the 1850s.&lt;/p&gt;
&lt;p&gt;The authors argue that high monitoring costs in the early century made trust-based, personalistic organization the second-best solution to principal-agent problems. The limited supply of sufficiently trusted individuals constrained geographic expansion, delegation, and total size. As railroad and telegraph networks lowered communication and transportation costs, monitoring capacity increased, enabling a transition to a Weberian bureaucracy no longer constrained by trust supply.&lt;/p&gt;
&lt;p&gt;The causal identification strategy uses the staggered expansion of the railroad network. For each county and decade (1820–1900), the authors compute the minimum-travel-time route from the county centroid to DC using Donaldson and Hornbeck (2016) data on railroads, steamboat waterways, coastal routes, and land routes. The specification includes county fixed effects, state-by-decade fixed effects, and controls for local railroad presence in the county and for the county&amp;rsquo;s market access, so the identifying variation comes from distant changes in the network that altered travel time to DC without directly affecting the county&amp;rsquo;s local economy or trade access.&lt;/p&gt;
&lt;p&gt;Results: a one standard deviation decrease in travel time to DC raises the probability of federal state presence by approximately 3 percentage points (about 8% of the mean), raises log employment similarly, raises the probability of observing a local managerial layer by approximately 3 percentage points (about 8% of the mean), and reduces employee turnover by approximately 2 percentage points (about 4% of the mean turnover rate). Placebo tests confirm that travel time to other major economic centers does not predict state presence. Telegraph network data (1845–1852, Wang 2020) yield consistent results. An additional test using the post-Civil War decline in Southern-born employee shares shows that better railroad connection to DC narrowed the North-South employment gap, consistent with monitoring substituting for trust-based selection.&lt;/p&gt;
&lt;p&gt;Scope conditions: the paper covers the civilian executive branch of the federal government, excluding the Postal Office, navy yards, and the engineer department; results are robust to restricting to states already in the union at the start of the sample, ruling out frontier-specific dynamics.&lt;/p&gt;
&lt;p&gt;Q: What is the central theoretical claim of the paper?
A: The paper argues that state capacity is fundamentally an organizational problem shaped by principal-agent constraints. When communication and transportation costs are high, the government cannot effectively monitor distant agents, so the second-best solution is to staff the bureaucracy with trusted individuals connected through personal networks. This personalistic form limits size and delegation because the supply of sufficiently trusted individuals is inherently scarce. Technological reductions in monitoring costs allow a transition to a Weberian bureaucracy based on procedural oversight rather than trust, removing the supply constraint on organizational growth.&lt;/p&gt;
&lt;p&gt;Q: What data source does the study rely on, and what time period does it cover?
A: The study draws on the Official Register of the United States, a biennial government publication listing all federal employees, digitized for every issue from 1817 to 1905. The resulting dataset includes 304,410 unique employees and 810,942 employee-year observations, with each record carrying name, state of birth, state of appointment, occupation, salary, department, office, location, and — through hierarchical reconstruction — position in a four-layer chain of command.&lt;/p&gt;
&lt;p&gt;Q: How did the size of the U.S. federal bureaucracy evolve over the nineteenth century?
A: Growth was slow before the 1860s. The first Register for 1817 listed 1,056 employees across 33 pages; the 1905 volume listed over 120,000 employees across 1,254 pages. Geographic expansion contributed zero to state growth before 1859 — the share of counties with any federal employee hovered around 15% from 1817 to 1859 — but contributed approximately 29% of growth after 1859, when county presence rose to 24% by 1871, 38% by 1881, and 61% by 1905.&lt;/p&gt;
&lt;p&gt;Q: What were the three sources of state growth, and how did their relative importance change?
A: The authors decompose growth into: (1) functions (new offices/bureaus), (2) geographic expansion (new counties), and (3) intensity (more employees per county-office pair). Before 1859, growth was entirely driven by functions (~40%) and intensity (~60%), with zero contribution from geographic expansion. After 1859, geographic expansion accounted for ~29%, intensity for ~32%, and functions for ~39% of growth.&lt;/p&gt;
&lt;p&gt;Q: How did employee turnover behave across the century, and what pattern emerges at presidential transitions?
A: Turnover trended upward through the late 1850s and then declined. During presidential transitions, the rate rose from 52–53% in 1841 and 1845 to 60–63% in 1849 and 1853 and peaked at 72% in 1861; it then fell to 55% in 1869, 44–48% in 1885/1889/1893, and 35% in 1897. Turnover was consistently lower in DC than in the field: controlling for year-bureau-position fixed effects, being employed in DC was associated with a 40% reduction in turnover probability.&lt;/p&gt;
&lt;p&gt;Q: How tight was the link between supervisors&amp;rsquo; and subordinates&amp;rsquo; careers, and how did it change?
A: Before 1841, moving from none to all supervisors leaving an organizational unit increased subordinate turnover probability by 37 percentage points. The effect was similar between 1841 and 1859, then dropped substantially to 22 percentage points in the following twenty-year period, and remained roughly constant after 1881. This pattern is consistent with the early bureaucracy relying on chains of personal trust that broke when a supervisor departed.&lt;/p&gt;
&lt;p&gt;Q: What evidence describes the evolution of delegation outside DC?
A: The number of field managers did not grow between 1817 and 1859 — it actually declined in the 1820s and was flat through the mid-1850s — and then tripled by 1905 relative to the 1850s level. The probability that workers in a local office had an additional managerial layer between them and DC was unchanged between pre-1841 and 1841–1859, increased by 5 percentage points between 1861 and 1881, and by 6 percentage points post-1881.&lt;/p&gt;
&lt;p&gt;Q: How does the paper measure monitoring capacity for the causal analysis?
A: The primary measure is travel time in hours from each county centroid to Washington DC, computed decade by decade (1820–1900) as the minimum-cost route across the available railroad network, steamboat waterways, coastal routes, and land routes, using data from Donaldson and Hornbeck (2016). A second, complementary measure is the number of telegraph connections between a county and DC using data from Wang (2020) for 1845–1852.&lt;/p&gt;
&lt;p&gt;Q: What is the identification strategy for the railroad analysis, and why are controls for local railroads and market access important?
A: The specification includes county fixed effects, state-by-decade fixed effects, an indicator for whether the county itself has railroad (LocalRailroad), and the county&amp;rsquo;s market access. County fixed effects mean beta is identified within-county from changes over time. Controlling for local railroad removes the direct correlation between local construction and local economic growth. Controlling for market access removes the effect of distant rail expansion on trade flows that raised agricultural land values and manufacturing activity. The remaining variation in travel time to DC — coming from distant network changes that altered the DC-county connection without affecting local conditions or broader trade access — is the identifying source.&lt;/p&gt;
&lt;p&gt;Q: What are the main quantitative effects of reduced travel time to DC?
A: A one standard deviation decrease in travel time to DC is associated with: (1) approximately 3 percentage point increase in the probability of federal state presence (~8% of the mean); (2) a similar magnitude increase in log employment conditional on presence; (3) approximately 3 percentage point higher probability of an additional managerial layer (~8% of the mean); and (4) approximately 2 percentage point reduction in employee turnover (~4% of the mean turnover rate).&lt;/p&gt;
&lt;p&gt;Q: How do placebo tests support the monitoring interpretation?
A: The authors show that, conditional on the same controls, travel times from a county to a set of other major economic centers are not associated with larger federal state presence. Since these other cities had no role as monitoring headquarters, the absence of an effect for them and the presence of an effect specifically for DC is consistent with the channel operating through the government&amp;rsquo;s ability to supervise agents from the capital, rather than through generic economic connectivity.&lt;/p&gt;
&lt;p&gt;Q: What does the telegraph evidence add, and what is its limitation?
A: Telegraph data (1845–1852, Wang 2020) show that counties with more telegraph connections to DC have larger state presence, more managerial delegation, and lower turnover, consistent with the monitoring mechanism. The limitation is that the authors have limited ability to address the endogeneity of telegraph network timing — the telegraph analysis is treated as corroborating evidence rather than the primary causal identification.&lt;/p&gt;
&lt;p&gt;Q: How do the Southern-born employee results illuminate the trust mechanism?
A: After the Civil War, the share of Southern-born federal bureaucrats fell sharply, consistent with reduced trust toward individuals from former Confederate states. However, counties that became better connected to DC via railroad expansion experienced a relative increase in the share of Southern-born employees. This shows that when monitoring costs fell, the government was willing to hire individuals from groups with lower baseline trust — monitoring substituted for trust as the mechanism ensuring agent performance.&lt;/p&gt;
&lt;p&gt;Q: Does federal state presence crowd out state and local government?
A: No. The presence of federal bureaucrats is positively correlated with the presence of state and local government employees at the county level, suggesting complementarity rather than substitution across levels of government.&lt;/p&gt;
&lt;p&gt;Q: What alternative mechanisms do the authors consider and how do they address them?
A: Three alternatives are discussed. First, demand shocks (Civil War debt repayment, industrialization) could explain the post-1860s expansion; the empirical specifications control for year fixed effects to absorb aggregate time-varying incentives, and the identification relies on differential cross-county variation in DC connectivity. Second, patronage as an electoral tool is consistent with spoils-driven turnover spikes but cannot explain why better-connected counties show lower turnover before civil service reform. Third, cognitive models of the firm (lower communication costs complement managerial problem-solving even without agency problems) could also predict the positive delegation result; the authors note they cannot empirically distinguish the monitoring and cognitive channels, and both may contribute.&lt;/p&gt;
&lt;p&gt;Q: What are the implications for developing countries today?
A: The authors suggest that their findings from nineteenth-century U.S. history may apply to understanding why modern Weberian bureaucracies remain elusive in many developing countries. Where communication infrastructure is limited and monitoring costs remain high, personalistic organizational forms based on trust networks may persist as constrained optima — not failures of will or design, but rational responses to structural conditions. Infrastructure investment that lowers monitoring costs could be a precondition for bureaucratic modernization.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Personalistic state organization&lt;/strong&gt;: The paper&amp;rsquo;s term for the organizational form that prevails when monitoring costs are high. It is characterized by staffing decisions based on personal character, moral reputation, and relationships of trust between principals and agents — and between supervisors and subordinates — rather than on formal procedural monitoring of performance. Frequent turnover at leadership transitions and constrained delegation are defining features, because the supply of trusted individuals is limited.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Weberian bureaucracy&lt;/strong&gt;: In the paper&amp;rsquo;s usage (following Weber 1978), a modern state organization defined by a fixed hierarchy of officials monitored through procedural rules rather than personal trust, lower turnover, and effective delegation of managerial power to geographically dispersed units. The paper treats this as the organizational form enabled by low monitoring costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Monitoring capacity&lt;/strong&gt;: The principal&amp;rsquo;s (politicians in DC and their cabinets) ability to observe and evaluate the behavior of agents (federal employees) throughout the territory. In the paper&amp;rsquo;s operationalization, monitoring capacity is proxied inversely by travel time and communication cost between DC and the county: lower travel time and more telegraph connections mean higher monitoring capacity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Geographic expansion component&lt;/strong&gt;: One of three decomposed sources of state growth. Defined as the increase in state size attributable to the state becoming present in more county locations. This component contributed zero to federal growth before 1859 and approximately 29% of growth after 1859.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Employee turnover&lt;/strong&gt;: In the paper&amp;rsquo;s measurement, the share of employees who leave the federal bureaucracy in a given year. The paper distinguishes politically-driven spikes at presidential transitions — reaching 72% of employees in 1861 — from the secular trend, which rose through the late 1850s and then declined, reaching 35% by the 1897 transition.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Delegation of managerial power&lt;/strong&gt;: The probability that a local county office has an additional managerial layer between its workers and DC, rather than reporting directly to the bureau-level supervisor in Washington. The paper uses this as its measure of whether decision authority has been decentralized to the field.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Trust substitution&lt;/strong&gt;: The paper&amp;rsquo;s mechanism linking monitoring capacity to organizational form. In the absence of effective monitoring, principals substitute trust for oversight — selecting agents whose personal loyalty, moral character, or political alignment gives the principal confidence they will not shirk or defect. As monitoring costs fall, trust becomes less necessary as a screening device, and the trust-constrained supply limit on organizational growth is relaxed.&lt;/p&gt;</description></item><item><title>Structural Change, Land Use and Urban Expansion</title><link>https://macropaperwarehouse.com/papers/structural-change-land-use-and-urban-expansion/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/structural-change-land-use-and-urban-expansion/</guid><description>&lt;p&gt;This paper asks how cities grow in the process of structural transformation — specifically, whether urban expansion occurs at the intensive margin (higher density within a fixed area) or the extensive margin (larger area). The authors document and explain a persistent decline in urban density in France since 1870, and develop a spatial general equilibrium model in which endogenous land use — land allocated either to agriculture or housing — is the key mechanism linking structural change to urban sprawl.&lt;/p&gt;
&lt;p&gt;The central empirical fact is striking: between 1870 and 2015, the area of the 100 largest French cities increased by a factor of roughly 30, while their population grew by only a factor of about 4, implying that average urban density fell by a factor of roughly 8. This density decline was fastest over 1950–1975, coinciding with the acceleration of structural change (France&amp;rsquo;s rural exodus). Since the mid-nineteenth century, approximately 15% of French land has been reallocated away from agricultural use — more than the total artificially-used land in France today (about 9%).&lt;/p&gt;
&lt;p&gt;The theoretical mechanism operates through the opportunity cost of urban expansion. Agricultural land at the urban fringe must earn its marginal product in the rural sector; this agricultural rent pins down the cost of converting land to urban use. When agricultural productivity is low, farmland is expensive relative to income (the &amp;ldquo;food problem&amp;rdquo;), households devote large shares of resources to food, and cities remain small in area and very dense. As agricultural productivity rises — the engine of structural change — workers leave rural areas, farmland values fall relative to income, and cities can expand cheaply at their fringes. Simultaneously, richer households spend more on housing. Both forces cause urban area to grow faster than urban population, generating a sustained decline in average density.&lt;/p&gt;
&lt;p&gt;The model also predicts a &amp;ldquo;hockey-stick&amp;rdquo; path for housing prices: during structural change, the extensive margin expansion of cities limits the rise in urban land rents despite growing housing demand. Once the reallocation of workers and land out of agriculture slows, urban land values must adjust upward rapidly, producing the pattern documented by Knoll et al. (2017) — relatively flat housing prices until roughly the 1950s, then steep increases.&lt;/p&gt;
&lt;p&gt;The model is a multi-city, multi-sector spatial equilibrium framework with non-homothetic CES preferences (including a subsistence requirement for the agricultural good), endogenous city fringes determined by land market clearing between agricultural and residential uses, and a monocentric commuting structure with endogenous commuting speed (workers adopt faster modes as wages rise). The model is calibrated to French historical data spanning 1840–2015, with 20 regions whose sectoral productivities are estimated to match regional urban populations and local farmland prices.&lt;/p&gt;
&lt;p&gt;Quantitatively, the calibrated model accounts for approximately 70% of the increase in urban area since 1870, most of the decline in average urban density (the factor-of-8 fall), about half of the rise in real housing prices, and most of the reallocation of land values from agricultural to urban. Cross-sectional evidence confirms a core prediction: cities surrounded by more expensive farmland are denser, with an IV-estimated elasticity of urban density with respect to farmland prices of approximately 0.3 (a 10% increase in farmland prices raises urban density by about 3%), consistent with the model&amp;rsquo;s counterpart. Scope conditions include the focus on France as a single country case, reliance on a monocentric urban structure, and the abstraction from within-urban-sector reallocation (manufacturing to services).&lt;/p&gt;
&lt;p&gt;Q: What is the central stylized fact motivating the paper?
A: Between 1870 and 2015, the area of the 100 largest French cities increased by a factor of roughly 30, while their total population grew by a factor of about 4, so average urban density fell by a factor of roughly 8. This density decline was most rapid over 1950–1975, coinciding with France&amp;rsquo;s peak rural exodus, and has barely fallen since — tracking the slowdown of structural change. This pattern is not unique to France; Angel et al. (2010) document persistent urban density decline on a global scale.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s key theoretical mechanism linking structural change to urban sprawl?
A: The rental price of agricultural land at the urban fringe is the opportunity cost of expanding the city into surrounding farmland. When agricultural productivity is low, farmland is expensive relative to income, keeping cities small and dense. As productivity rises and workers migrate to cities, the value of agricultural land falls relative to income, reducing the cost of urban expansion at the fringe. Richer households also devote a larger share of spending to housing, reinforcing the demand for space. These two channels together cause city area to grow faster than city population, generating a sustained decline in average density — even without any improvement in commuting technology.&lt;/p&gt;
&lt;p&gt;Q: How does the paper distinguish between the structural change channel and the commuting cost channel?
A: The model contains both channels: structural change (falling agricultural land values at the fringe) and falling effective commuting costs (rising wages lead workers to adopt faster commuting modes, a wage elasticity of commuting speed calibrated from survey data). Counterfactuals show that without structural change (rural productivity growth set to 4% of baseline), the model cannot replicate the observed density decline. Without faster commutes (setting the income elasticity of commuting speed to unity), the model predicts only about 30% of the baseline density decline. Both channels are necessary; their combined effect exceeds the sum of parts because structural change raises wages, which in turn amplifies the commuting speed mechanism.&lt;/p&gt;
&lt;p&gt;Q: How do the two channels differ in their spatial imprint within cities?
A: Structural change adds new low-density settlements at the urban fringe, so suburban density falls more than average density — the center is relatively less affected. Faster commuting modes, by contrast, induce suburbanization: workers relocate from the center outward, so central density falls more than average density. For Paris, historical data show that central density fell less than average urban density, which is consistent with both mechanisms operating simultaneously — the commuting channel pushing central density down more, but the structural change channel adding fringe expansion that affects suburban density more.&lt;/p&gt;
&lt;p&gt;Q: What is the empirical evidence on the cross-sectional farmland price prediction?
A: Using data on local farmland transaction prices from the French Ministry of Agriculture at the &amp;ldquo;Petite Region Agricole&amp;rdquo; level (over 700 areas), the authors show that cities surrounded by more expensive farmland are denser. A binned scatter plot across 200 French cities shows that moving from the first to last decile of farmland prices raises density by about one third — an effect comparable in magnitude to an increase in population from roughly 25,000 (3rd decile) to 150,000 (9th decile). To address endogeneity (productive cities may inflate nearby farmland prices), the authors instrument farmland prices with soil quality characteristics; the IV elasticity of urban density with respect to farmland prices is approximately 0.3, consistent with the model&amp;rsquo;s predicted counterpart.&lt;/p&gt;
&lt;p&gt;Q: What does the model predict about the time path of housing prices?
A: The model predicts a &amp;ldquo;hockey-stick&amp;rdquo; pattern: housing prices remain relatively flat for decades while structural change is ongoing, because cities expand cheaply at the extensive margin, absorbing growing housing demand without large rent increases. Once the reallocation of workers and land out of agriculture slows, the extensive margin ceases to buffer demand, and urban land values must rise sharply. The calibrated model accounts for about half of the observed rise in real housing prices since the mid-nineteenth century; it matches the qualitative hockey-stick pattern documented by Knoll et al. (2017) and Piketty and Zucman (2014) for France and advanced economies more broadly.&lt;/p&gt;
&lt;p&gt;Q: What happens to the relative values of agricultural versus urban land over the period?
A: Agricultural land values relative to income fall dramatically: the average value of a French agricultural field per unit of land, as a share of per capita income, was divided by a factor of 15 between 1850 and 2015. Meanwhile, urban land values rise. In 1820, agricultural land accounted for more than 70% of total housing and land wealth in France; by 2010 this share had fallen to about 3%. This reallocation of land values from rural to urban is a central prediction the model accounts for, driven by structural change reducing the scarcity premium on farmland.&lt;/p&gt;
&lt;p&gt;Q: How is the model parameterized and calibrated?
A: Preferences are non-homothetic CES with housing preference parameter gamma = 0.22, subsistence consumption for the rural good calibrated to match the 1840 agricultural employment share (about 60%), and substitution elasticity between urban and rural goods sigma = 0.8. The labor share in agriculture is alpha = 0.6. Commuting cost parameters (elasticities to wages and distance) are estimated from the French Labor Force Survey (Enquete Emploi). Region-specific sectoral productivity parameters for 20 regions (40 parameters total) are estimated to match the cross-section of urban populations and local farmland values in the base year 1870. The model is then simulated forward to 2015.&lt;/p&gt;
&lt;p&gt;Q: What share of French land has been reallocated away from agriculture, and how does this relate to urban expansion?
A: About two-thirds of French land was used for agriculture in 1840; by 2015 this fell to 52%, implying roughly 15 percentage points of French territory reallocated away from agricultural use. This 15% exceeds the total land currently under artificial use in France (about 9%). Over the more precisely measured period 1982–2015, artificialized soil increased by about 2 million hectares (3.7% of French territory), representing roughly 70% of the land converted away from agriculture over the same period. Two-thirds of land surrounding French cities is agricultural, confirming that urban expansion occurs at the expense of farmland.&lt;/p&gt;
&lt;p&gt;Q: What are the limitations and directions for future research acknowledged by the authors?
A: The model relies on a monocentric urban structure where all workers commute to a single city center, which is an approximation — commuting distance increases with residential distance to the center but less than one-for-one, suggesting workers sort into nearby jobs. The model also abstracts from within-urban-sector reallocation (the manufacturing-to-services transition), which the authors conjecture matters for the cross-section of cities in recent times. Finally, the model cannot fully replicate the steep recent rise in housing prices, which the authors attribute partly to land-use regulations constraining extensive margin growth — a policy counterfactual the general equilibrium structure is well-suited to analyze.&lt;/p&gt;
&lt;p&gt;Q: How does the paper relate to the Ricardo/Nichols view that land values should rise with economic development?
A: The traditional Ricardian view predicts that a fixed factor like land must rise in value with economic development — counterfactual given the historical data showing farmland values falling sharply relative to income. The authors reconcile this with the data by emphasizing that structural change and agricultural productivity growth reduce the scarcity of farmland even as total income grows, so farmland values fall. Urban land values do rise, but the structural change channel initially dampens this increase by facilitating extensive-margin city growth. The paper thus reconciles the Ricardian fixed-factor view with the commuting technology view (Miles and Sefton, 2020) within a unified spatial structural change framework.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous land use&lt;/strong&gt;: In this paper&amp;rsquo;s framework, land in each region is allocated either to agricultural production or to residential use, with the margin between the two determined in equilibrium by the equality of the rental price of land at the urban fringe and the marginal product of land in the rural sector. This makes the urban-rural land boundary an endogenous object that responds to structural change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Urban fringe (phi_k)&lt;/strong&gt;: The furthest residential location of an urban worker in city k, determined endogenously as the commuting distance at which the opportunity cost of further expansion (the agricultural land rent) equals the willingness of urban workers to pay for land. All workers beyond this fringe produce rural goods without commuting.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural change (in the paper&amp;rsquo;s sense)&lt;/strong&gt;: The reallocation of workers and land away from agriculture driven jointly by non-homothetic preferences with a subsistence consumption requirement for the agricultural good (demand side) and rising sectoral productivity (supply side). Structural change is the primary driver of falling farmland values and urban sprawl in the model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-homothetic CES preferences&lt;/strong&gt;: Household preferences over rural and urban goods that are not homogeneous of degree one in income, specified as a CES aggregate with a subsistence floor for the rural (agricultural) good. At low income levels, households devote large budget shares to food; as income rises, spending shifts toward urban goods and housing. This demand-side non-homotheticity is the channel through which rising income generates structural change.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Food problem (Schultz, 1953)&lt;/strong&gt;: The condition in which low agricultural productivity forces households to devote a large fraction of resources to meeting subsistence food needs, leaving little for housing expenditure. In the paper&amp;rsquo;s model, the food problem makes cities initially small and very dense; as agricultural productivity rises and the food problem relaxes, cities can expand in area.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Commuting cost function tau(l_k)&lt;/strong&gt;: Spatial frictions proportional to the worker&amp;rsquo;s distance from the city center and the urban wage, of the functional form tau(l_k) = a * w_{u,k}^{xi_w} * l_k^{xi_l}, where xi_w in (0,1) captures the endogenous adoption of faster commuting modes as wages rise. Concavity in both arguments is micro-founded by an optimizing commuting mode choice model, ensuring that the share of resources devoted to commuting falls as incomes rise.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hockey-stick housing price path&lt;/strong&gt;: The model&amp;rsquo;s prediction that real housing prices remain relatively flat over the period of active structural change — because city expansion at the extensive margin absorbs rising housing demand without large rent increases — before rising steeply once structural change slows and the extensive margin is exhausted. This prediction matches the empirical pattern documented by Knoll et al. (2017) for France and other advanced economies.&lt;/p&gt;</description></item><item><title>Technology Transfer and Early Industrial Development: Evidence from the Sino-Soviet Alliance</title><link>https://macropaperwarehouse.com/papers/technology-transfer-and-early-industrial-development-evidence-from-the-sino-soviet-alliance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/technology-transfer-and-early-industrial-development-evidence-from-the-sino-soviet-alliance/</guid><description>&lt;p&gt;This paper estimates the causal effect of technology and knowledge transfers on early industrial development using the Sino-Soviet Alliance of the 1950s as a natural experiment. Between 1950 and 1957, the Soviet Union supported the &amp;ldquo;156 Projects&amp;rdquo; — 139 approved civil projects for constructing technologically advanced, large-scale, capital-intensive industrial facilities in China. The intended program comprised two components: a &amp;ldquo;basic&amp;rdquo; transfer of Soviet state-of-the-art machinery and equipment (including blueprints, site surveys, and plant construction assistance), and an &amp;ldquo;advanced&amp;rdquo; know-how transfer involving Soviet experts residing in Chinese plants for roughly three years to train engineers and production supervisors in organizational, technological, and planning methods. Total investment amounted to approximately $80 billion in 2020 figures (45.7% of Chinese GDP in 1949).&lt;/p&gt;
&lt;p&gt;Identification exploits idiosyncratic delays in project completion caused by Soviet production capacity constraints, insufficient experts, translator shortages, and miscommunication — factors documented in historical records as unrelated to project-specific characteristics. When the Sino-Soviet Split in 1960 abruptly ended the program, all 139 plants had been built but differed in what transfers they had received: 46 received both machinery and know-how (advanced), 46 received only machinery (basic), and 47 received neither (comparison). The paper verifies, via ANOVA tests, multinomial logit models, balancing regressions on 26 plant characteristics, pre-trend tests, and Oster (2019) selection-on-unobservables bounds, that the three groups were statistically equivalent prior to receiving the Soviet transfers.&lt;/p&gt;
&lt;p&gt;The primary data source is plant-level annual reports from the Steel Association covering 94 steel firms (1,410 plants) from 1949 to 2000, matched to 304 steel plants across the 156 Projects. Supplementary sources include the declassified 1985 Second Industrial Survey (7,592 largest Chinese firms) and the China Industrial Enterprises database (1998–2013, over 1 million firms).&lt;/p&gt;
&lt;p&gt;Three main results emerge. First, receiving only the basic (machinery) transfer had positive but short-lived effects: output of basic plants peaked at 14.7 percent above comparison plants six years after receiving Soviet machinery, then declined monotonically and became statistically insignificant after 20 years — consistent with the estimated 15–20 year life cycle of Soviet capital. Second, the advanced transfer had large and persistent effects: advanced plants&amp;rsquo; output rose 8.4 percent relative to basic plants within two years, 19.7 percent within 20 years, and 49.5 percent cumulatively after 40 years. TFPQ of advanced plants reached 47.9 percent above basic plants after 40 years. These magnitudes held across industries in 1985 and 1998–2013 data, where value added of advanced firms was 41.4–52.0 percent higher and TFPR 39.5–49.3 percent higher than basic firms. Third, the program generated horizontal spillovers (12.9 percent higher output, 12.4 percent higher productivity for steel plants in counties hosting advanced plants) and vertical spillovers (16.4 percent productivity gain for supply-chain firms in counties of advanced nonsteel plants), with spillover effects conditional on post-1990s market liberalization to materialize in private firms.&lt;/p&gt;
&lt;p&gt;The mechanism driving persistence is the accumulation of organizational and human capital during the advanced transfer, which enabled advanced plants — uniquely — to develop new production processes endogenously, home-fabricate continuous casting furnaces to replace obsolete Soviet open-hearth equipment, and produce export-quality steel. Advanced plants employed more engineers and high-skilled technicians, established professional schools, and their counties had 10.4 percent higher STEM university degree rates and 16.8 percent more technical schools.&lt;/p&gt;
&lt;p&gt;Scope conditions: results apply to large-scale, capital-intensive state-planned industrial facilities in a country at an early stage of industrialization, under conditions of near-complete trade isolation (1960–1978) that prevented basic plants from compensating via imported foreign capital. The estimated aggregate contribution of the program is that, without both transfer types, Chinese real GDP per capita growth between 1953 and 1978 would have been 7 to 19 percent lower.&lt;/p&gt;
&lt;p&gt;Q: What distinguishes the &amp;ldquo;basic&amp;rdquo; from the &amp;ldquo;advanced&amp;rdquo; Soviet transfer?
A: The basic transfer involved duplication of whole Soviet plants through provision of state-of-the-art Soviet machinery, equipment, blueprints, geological surveys, and construction assistance. The advanced transfer added visits of Soviet experts — expected to stay approximately three years — to teach Chinese technicians how to operate the machinery and to provide within-firm training in engineering (math, physics, chemistry, organizational and planning methods) and supervisory management based on &amp;ldquo;scientific management&amp;rdquo; principles including quality-control methods.&lt;/p&gt;
&lt;p&gt;Q: What caused plants to receive different levels of transfer, and why is this variation credible for identification?
A: Delays arose from Soviet production capacity constraints (by 1955, one-third of annual Soviet steel-rolling output was destined for China), insufficient experts, translator shortages, and bilateral miscommunication — all documented in historical records as unrelated to project characteristics. When the 1960 Split ended the program, plants&amp;rsquo; treatment status was determined by where they happened to be in the delivery queue. ANOVA tests find no significant differences in approval year, investment, workforce, equipment value, project length, or capacity across the three groups, and a multinomial logit on province and industry fixed effects shows no group had higher ex-ante probability of receiving either transfer type.&lt;/p&gt;
&lt;p&gt;Q: What were the output effects of the basic transfer, and why did they fade?
A: Output of basic plants was not significantly above comparison plants for the first two years, peaked at 14.7 percent higher six years after receiving Soviet machinery, then declined monotonically and became statistically insignificant after 20 years. This timing corresponds to the estimated 15-year life cycle of Soviet capital goods. TFPQ of basic plants followed the same pattern, peaking at 14.5 percent above comparison plants. Without the know-how component, basic plants could not develop new processes or home-fabricate replacement capital, so productivity advantages disappeared as Soviet equipment became obsolete.&lt;/p&gt;
&lt;p&gt;Q: What were the output and productivity effects of the advanced transfer?
A: Advanced plants&amp;rsquo; output rose 8.4 percent relative to basic plants within two years of the Soviet transfer and 19.7 percent within 20 years, reaching a cumulative effect of 49.5 percent after 40 years. TFPQ of advanced plants increased from 8.3 percent above basic plants two years after the transfer to 47.9 percent after 40 years. These effects were driven by output growth rather than differential input use — the number of workers, coke, and iron were statistically indistinguishable across the three plant types — ruling out government input reallocation as an explanation.&lt;/p&gt;
&lt;p&gt;Q: Did the advanced transfer affect steel quality?
A: Advanced plants produced substantially more crude steel (higher quality, lower carbon content) and less pig iron than basic and comparison plants, and this quality advantage persisted well beyond the 20-year life cycle of Soviet capital. Basic plants also shifted toward crude steel initially but the quality advantage dissipated once Soviet machinery became obsolete, whereas advanced plants maintained the shift through adoption of the basic oxygen process and later continuous casting furnaces.&lt;/p&gt;
&lt;p&gt;Q: What is the main mechanism through which the advanced transfer generated persistent effects?
A: The advanced transfer equipped engineers and supervisors with organizational, technological, and planning knowledge, enabling advanced plants to develop and adopt the basic oxygen steelmaking process independently during China&amp;rsquo;s 1960–1978 period of trade isolation. Advanced plants had a 15.2 percent higher probability of using the basic oxygen process five years after the transfer and a 65.1 percent higher probability twenty years after, relative to basic plants. They also home-fabricated continuous casting furnaces, making them 26.7 to 78.4 percent more likely to use such furnaces 10 to 20 years after the transfer; basic plants showed no differential advantage over comparison plants on this measure.&lt;/p&gt;
&lt;p&gt;Q: What role did trade openness play in the divergence between basic and advanced plants?
A: Once China opened to international trade from 1978, advanced plants relied dramatically less on imported foreign capital than basic plants — likely because they had developed domestic production capabilities. At the same time, advanced plants exported 45.5 percent more steel and produced 51.1 percent more steel above international quality standards than basic plants. Basic plants showed no differential imports of foreign capital or differential exports relative to comparison plants, suggesting that once both types could access foreign machinery, basic plants lost any remaining productivity edge.&lt;/p&gt;
&lt;p&gt;Q: What were the human capital effects of the advanced transfer?
A: Over time, advanced plants opened training schools for high-skilled technicians and offered within-firm training programs for engineers. As a result, advanced plants employed more engineers and high-skilled technicians and fewer low-skilled workers than basic plants, while the human capital composition did not differentially change between basic and comparison plants. At the county level, universities hosting advanced plants were 10.4 percent more likely to offer STEM degrees, had 16.8 percent more technical schools, 14.3 percent more STEM college graduates, and 17.6 percent more high-skilled workers than counties hosting basic plants.&lt;/p&gt;
&lt;p&gt;Q: Did the government differentially favor basic or advanced plants after the Split?
A: The paper finds no evidence of special government favor. Government transfers and loans were not differentially allocated to basic or advanced plants in either the short or long run. Distance from railroads and roads did not change differentially across plant types. Measures of political connection and politician quality at the prefecture level showed no significant differences across the three groups in the 40 years after the Soviet transfer. County-level total investment and investments in related and unrelated industries were also statistically indistinguishable.&lt;/p&gt;
&lt;p&gt;Q: What were the intra-firm spillover effects?
A: Steel plants in the same firm as advanced plants increased their steel production by 24.9 percent and were 22.1 percent more productive relative to plants in the same firm as basic plants, after the Soviet transfer. Plants in the same firm as basic plants showed no differential performance relative to plants in the same firm as comparison plants. The within-firm spillovers appear driven by the transmission of new technologies and production methods through formal within-firm training programs, as supported by historical records.&lt;/p&gt;
&lt;p&gt;Q: What were the horizontal spillover effects across firms?
A: Steel plants in the same counties as advanced plants produced 12.9 percent higher output and were 12.4 percent more productive than those in counties hosting basic plants, after the transfer. They were more likely to adopt basic oxygen converters and continuous casting furnaces, and from 1978 they exported significantly more and produced more steel above international quality standards, mirroring the patterns of the advanced plants themselves.&lt;/p&gt;
&lt;p&gt;Q: What were the vertical spillover effects?
A: Steel plants in counties hosting nonsteel basic plants produced 14.2 percent more steel than those in counties hosting nonsteel comparison plants, suggesting some output spillover from basic machinery. However, only plants in counties of advanced nonsteel plants experienced a productivity increase — estimated at 16.4 percent — relative to plants in counties of basic nonsteel plants. These supply-chain firms were also the only ones to show increased adoption of basic oxygen and continuous casting furnace technology and differential engagement in trade.&lt;/p&gt;
&lt;p&gt;Q: How did market liberalization reforms interact with the spillover effects?
A: Starting in the late 1990s, privatized firms economically related to advanced plants outperformed their counterparts in terms of value added, TFPR, and exports, while state-owned firms in the same counties no longer showed a competitive advantage. New private firms locating in counties that had hosted advanced plants received an additional performance gain. At the county level, counties hosting advanced plants had on average 16.6 percent more private firms and 25.2 percent more privately-produced industrial output than counties hosting basic plants. The mechanism appears to be the stock of industry-specific human capital concentrated in those counties, which private firms could draw on once allowed to compete for workers.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated aggregate contribution of the Soviet transfer to Chinese growth?
A: Province-level regressions show that each additional basic project increased province-level output by 1.1 percent per year on average, and each additional advanced project by 6.2 percent per year. A back-of-the-envelope calculation implies that without both transfer types, Chinese real GDP per capita growth between 1953 and 1978 would have been 7 to 19 percent lower.&lt;/p&gt;
&lt;p&gt;Q: How does the paper rule out selection on unobservable characteristics?
A: Using the Oster (2019) methodology, the paper finds that for the treatment effects to become statistically insignificant, selection on unobserved variables would need to be 8 to 19 times larger than selection on observed variables — a range the authors characterize as implausible given the strong balancing on observables and the historical documentation of delay causes.&lt;/p&gt;
&lt;p&gt;Q: How does this paper differ from Heblich et al. (2020), which also studies Sino-Soviet technology transfer?
A: Heblich et al. (2020) study long-run negative spillovers of the 156 Projects on counties that hosted them relative to counties that were geographically suitable but ultimately not selected, focusing on an outside-the-program comparison. This paper instead exploits within-program variation — differences across the three plant types — using plant-level data to assess short-, medium-, and long-run direct effects and spillover effects of different transfer intensities.&lt;/p&gt;
&lt;p&gt;Basic Transfer: The provision of Soviet state-of-the-art machinery, equipment, blueprints, geological surveys, and plant construction assistance — duplicating a whole Soviet plant — without accompanying human capital or organizational training.&lt;/p&gt;
&lt;p&gt;Advanced Transfer: The full Soviet technology and know-how package: basic machinery provision plus multi-year visits of Soviet experts who taught Chinese engineers and production supervisors organizational, technological, and planning methods based on &amp;ldquo;scientific management&amp;rdquo; principles.&lt;/p&gt;
&lt;p&gt;Comparison Plants: Plants approved under the 156 Projects that received neither Soviet machinery nor technical assistance due to delays compounded by the Split, and which continued operating with traditional domestic technology.&lt;/p&gt;
&lt;p&gt;156 Projects: An array of 139 approved, technologically advanced, large-scale, capital-intensive industrial facilities whose construction the Soviet Union agreed to support between 1950 and 1957 as part of the Sino-Soviet Alliance, representing 45.7% of Chinese GDP in 1949.&lt;/p&gt;
&lt;p&gt;Tacit Knowledge: Industry- and firm-specific knowledge embodied in workers and organizations — including operational methods, quality-control procedures, and process innovation capabilities — that cannot be transferred through capital goods alone and requires extensive on-the-job training from foreign experts.&lt;/p&gt;
&lt;p&gt;Basic Oxygen Process: A steelmaking process innovation that became predominant in the 1960s by blowing oxygen through molten pig iron to reduce carbon content; adopted by advanced plants through endogenous process development, while basic plants showed no differential adoption relative to comparison plants.&lt;/p&gt;
&lt;p&gt;Source Text Origin: The paper&amp;rsquo;s classification scheme for the grounding of evidence — in this case, full working paper text obtained from NBER WP 29455, enabling comprehensive summary of quantitative results, mechanisms, and robustness tests.&lt;/p&gt;</description></item><item><title>The Effect of High-Tech Clusters on the Productivity of Top Inventors: Comment</title><link>https://macropaperwarehouse.com/papers/the-effect-of-high-tech-clusters-on-the-productivity-of-top-inventors-comment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effect-of-high-tech-clusters-on-the-productivity-of-top-inventors-comment/</guid><description>&lt;p&gt;This paper is a comment on Moretti (2021b), which studied agglomeration effects for innovation by testing whether the size of technology clusters causes patenting. The original paper (M21) used US patent data from 1971 to 2007 (Zucker and Darby, 2014) and reported a baseline elasticity of patenting with respect to cluster size of 0.0676, along with event study and instrumental variables (IV) evidence supporting a causal interpretation.&lt;/p&gt;
&lt;p&gt;Wiebe identifies two major methodological problems that undermine M21&amp;rsquo;s causal claims.&lt;/p&gt;
&lt;p&gt;Problem 1 — Misspecified event study. M21&amp;rsquo;s event study (Figure 6) was designed to test for selection bias from &amp;ldquo;rising star&amp;rdquo; inventors sorting into large clusters. The event is inventors moving across cities exactly once. However, M21&amp;rsquo;s specification interacts pre-move average cluster size with pre-move event-time indicators and post-move average cluster size with post-move event-time indicators separately — it does not exploit the change in cluster size generated by the move itself. Following the standard &amp;ldquo;mover&amp;rdquo; design literature (Finkelstein et al., 2016; Molitor, 2018; Cantoni and Pons, 2022), the correct specification uses the change in average cluster size as the treatment variable, interacted with event-time indicators. Wiebe implements this corrected event study and finds no statistically significant pre-trend and no statistically significant treatment effect post-move. Notably, the baseline elasticity estimated on the mover sample using all observed variation is large and significant at 0.3145 (SE 0.0953), but no effect is detected when variation is restricted to that generated by moving. The null result could also partly reflect attenuation bias from misclassified moves, since the dataset does not distinguish inventors who share the same name.&lt;/p&gt;
&lt;p&gt;Problem 2 — Coding error in IV. M21&amp;rsquo;s Table 5 instruments cluster size using variation in the number of inventors in other cities employed by firms also active in the focal inventor&amp;rsquo;s city, with the instrument calculated via first-differencing. Due to a coding error, M21 sorts data by firm, field, and year but not by city before first-differencing, so the differencing is taken across cities rather than within cities. Because firm-field-year is not a unique sorting key, Stata&amp;rsquo;s sort command pseudo-randomly orders observations with tied values, making the results unreproducible across runs. When Wiebe corrects the code to sort by city and compute first-differences within city, the 2SLS estimates become unstable and nonsignificant, with the first-stage F-statistic falling to approximately 7. This means M21 provides no valid IV evidence against confounding from city-field-year shocks such as local subsidies.&lt;/p&gt;
&lt;p&gt;Beyond these two major problems, the Appendix documents seven additional issues. The positive effect of cluster size on patent quality (M21 Table 6) disappears and reverses when the log transformation is corrected from log(y + 0.00001) to log(y + 1) or Poisson regression — the corrected estimate is negative and significant, implying that cluster size reduces citations per patent along the intensive margin and the overall quality effect is negative. Heterogeneous elasticity estimates (M21 Table 8) contain a coding error; corrected estimates show substantial heterogeneity. The distributed lag model (M21 Figure 5) uses an incorrectly defined lag structure in an unbalanced panel; corrected estimates yield nonsignificant contemporaneous effects. Cluster quality estimates (M21 Table A.8) use a cluster size definition differing from the text, and corrected elasticities are approximately half as large. M21&amp;rsquo;s claimed extensive margin effect in Table A.7 is logically unsupported since no zeros are observed. The team size robustness check is conceptually flawed because it controls twice for per-coauthor adjustment. A gap-interpolation coding error in Table A.6 biases estimates downward. Broader computational reproducibility failures arise from many-to-many merges with non-unique sort orders. Wiebe explicitly notes that the null IV and event study results are not evidence against agglomeration effects per se.&lt;/p&gt;
&lt;p&gt;Q: What is the baseline finding in M21 that Wiebe contests?
A: M21 reports a baseline elasticity of patenting with respect to cluster size of 0.0676, estimated from linear regressions with extensive fixed effects including inventor fixed effects. M21 presents an event study and IV strategy as additional evidence supporting a causal interpretation of this elasticity.&lt;/p&gt;
&lt;p&gt;Q: What is wrong with M21&amp;rsquo;s event study specification?
A: M21&amp;rsquo;s event study interacts pre-move average cluster size with pre-move event-time indicators and post-move average cluster size with post-move event-time indicators, but never uses the change in cluster size associated with moving. The standard mover design (Finkelstein et al., 2016; Molitor, 2018) uses the change in average environment as a constant treatment variable interacted with all event-time indicators. Because M21&amp;rsquo;s specification does not exploit moving-induced variation, it would be identified even if moving induced no change in cluster size.&lt;/p&gt;
&lt;p&gt;Q: What does Wiebe&amp;rsquo;s corrected event study find?
A: Wiebe&amp;rsquo;s corrected mover event study shows no statistically significant pre-trend (consistent with no systematic sorting of rising-star inventors into large clusters) and no statistically significant post-move treatment effect. In contrast, the baseline fixed-effects elasticity on the mover sample using all observed variation is 0.3145 (SE 0.0953) — large and significant — indicating the null result is specific to the moving-generated variation.&lt;/p&gt;
&lt;p&gt;Q: What alternative explanation does Wiebe offer for the null event study result?
A: The null result could be partly explained by attenuation bias from misclassified moves. M21&amp;rsquo;s code creates inventor identifiers based on names, but the COMETS dataset does not distinguish inventors who share the same name, so an apparent cross-city move may simply be two different inventors with the same name living in different cities.&lt;/p&gt;
&lt;p&gt;Q: What is the coding error in M21&amp;rsquo;s IV strategy?
A: M21 constructs the instrument by first-differencing a variable measuring inventors in other cities working for firms also active in the focal city. The code sorts by firm, field, and year before differencing, but omits city from the sort key, so first-differencing is computed across cities rather than within cities, generating an instrument that does not match the definition in the text.&lt;/p&gt;
&lt;p&gt;Q: Why does the coding error also cause non-reproducibility?
A: Firm-field-year is not a unique sorting key because multiple cities can share the same firm-field-year values. Stata&amp;rsquo;s sort command pseudo-randomly orders observations with tied values, so each run produces a different city ordering within tied groups and therefore a different instrument and different estimates.&lt;/p&gt;
&lt;p&gt;Q: What do the corrected IV results show?
A: After correcting the sort order to include city and computing first-differences within city, the 2SLS estimates are unstable and nonsignificant. The first-stage F-statistic falls to approximately 7, indicating a weak instrument. This does not constitute evidence against agglomeration effects, but means M21&amp;rsquo;s IV strategy provides no valid evidence against confounding from city-field-level shocks such as local subsidies.&lt;/p&gt;
&lt;p&gt;Q: What happens to the patent quality results when the log transformation is corrected?
A: M21 uses log(citations + 0.00001), which assigns very large weight to the extensive margin. When Wiebe uses log(citations + 1) or Poisson regression instead, the estimated effect of cluster size on patent quality is negative and statistically significant, reversing M21&amp;rsquo;s finding. The corrected result implies that while cluster size may raise the probability of producing any cited patent, it reduces citations per patent for inventors who do produce cited patents, and the overall effect is negative.&lt;/p&gt;
&lt;p&gt;Q: What are the corrected aggregate agglomeration loss estimates?
A: Using the corrected constant elasticity, the estimated output reduction from equalizing cluster sizes is -9.15% (slightly smaller than M21). Using corrected heterogeneous elasticities based on within-field-year size quartiles, the output loss is -23.75% (about twice as large). Using elasticities based on global size quartiles, the loss is -35.11%.&lt;/p&gt;
&lt;p&gt;Q: What is wrong with M21&amp;rsquo;s distributed lag model (Figure 5)?
A: M21&amp;rsquo;s code defines lags and leads using sequential observations in the panel rather than calendar years. Because the inventor-year panel is unbalanced, a coded &amp;ldquo;one-year lag&amp;rdquo; can refer to any number of years prior. When Wiebe restricts to inventors with 11 consecutive years and correctly defines year-based lags, confidence intervals widen substantially and the contemporaneous effect estimate becomes nonsignificant.&lt;/p&gt;
&lt;p&gt;Q: What is the conceptual flaw in M21&amp;rsquo;s team-size robustness check?
A: M21&amp;rsquo;s Table A.8 controls for the number of coauthors on a patent, but the dependent variable is already measured as patents per coauthor. Controlling for team size after already dividing by team size effectively controls for the same variable twice.&lt;/p&gt;
&lt;p&gt;Q: What are the broader computational reproducibility problems in M21?
A: The cleaning code uses many-to-many merges with non-unique sort orders, generating slightly different datasets on each run. For example, when merging inventors with patent assignees, patent identifiers are not unique because multiple firms can be assigned to a single patent. Removing name suffixes also causes distinct inventors (e.g., Paul H. Hamisch Jr. and Sr.) to be assigned the same identifier. Additionally, using reghdfe with the keepsingletons option retains singleton groups explicitly warned against by the package due to biased standard errors.&lt;/p&gt;
&lt;p&gt;Agglomeration elasticity: The elasticity of an inventor&amp;rsquo;s patent output with respect to the size of the technology cluster (city-field-year cell) in which they work; reported as 0.0676 in M21&amp;rsquo;s baseline and 0.3145 on the mover sample with all observed variation.&lt;/p&gt;
&lt;p&gt;Mover event study design: An event study specification in which the treatment variable is the change in an individual&amp;rsquo;s average environment (here, cluster size) before and after a geographic move, interacted with event-time indicators — the standard design used in Finkelstein et al. (2016) and Molitor (2018), which M21&amp;rsquo;s specification does not follow.&lt;/p&gt;
&lt;p&gt;Cluster size: The number of inventors (or cluster density) active in the same city-field-year cell as the focal inventor, used as the key independent variable in M21&amp;rsquo;s regressions.&lt;/p&gt;
&lt;p&gt;First-stage F-statistic: A measure of instrument strength in 2SLS IV estimation; the corrected instrument yields F ≈ 7 (indicating weakness), whereas M21&amp;rsquo;s incorrectly constructed instrument produced a stronger first stage by exploiting spurious cross-city variation.&lt;/p&gt;
&lt;p&gt;Extensive vs. intensive margin (patent quality): The extensive margin captures whether an inventor produces any cited patent; the intensive margin captures citations per patent conditional on having any. M21&amp;rsquo;s log(y + 0.00001) transformation overweights the extensive margin, and the corrected intensive-margin effect of cluster size on quality is negative and significant.&lt;/p&gt;
&lt;p&gt;Computational reproducibility: The property that running code on the same data produces identical results across runs. M21&amp;rsquo;s code fails this standard due to non-unique sort orders in merges and first-differencing steps, causing the IV instrument to differ across runs.&lt;/p&gt;
&lt;p&gt;Rising star sorting: The hypothesized selection mechanism whereby inventors with increasing patent trajectories are preferentially hired into large clusters, which would bias OLS agglomeration elasticity estimates upward; M21&amp;rsquo;s event study was designed to test for this but is incorrectly specified and does not use moving-induced variation.&lt;/p&gt;</description></item><item><title>The Future in Mind: Aspirations and Long-Term Outcomes in Rural Ethiopia</title><link>https://macropaperwarehouse.com/papers/the-future-in-mind-aspirations-and-long-term-outcomes-in-rural-ethiopia/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-future-in-mind-aspirations-and-long-term-outcomes-in-rural-ethiopia/</guid><description>&lt;p&gt;This paper tests whether a light-touch behavioral intervention targeting aspirations can produce persistent economic effects on a poor rural population. The research question is whether changing how poor people perceive their future opportunities — by raising aspirations — alters their investment decisions in ways that persist over a multi-year horizon. The authors conduct a randomized controlled trial in Doba, a remote mountainous district in rural Ethiopia roughly 380 kilometers from Addis Ababa, selected partly because its extreme isolation meant residents had almost no exposure to television or media, making even a single video screening a memorable event.&lt;/p&gt;
&lt;p&gt;The sample consists of 1,152 households (2,112 individuals) across 64 villages. Households were randomly assigned to one of three conditions: a treatment group shown four 15-minute documentaries featuring real rural individuals from similar communities who escaped poverty through goal-setting and hard work; a placebo group shown an Ethiopian entertainment comedy with no aspirational content; and a within-village control group who were only surveyed. Both the household head and spouse in treatment and placebo groups were invited to attend. Compliance was very high, with only 2 percent of individuals not complying with their assigned condition. Data were collected at baseline (2010), six months after screening (2011), and five years after baseline (2015–2016). Attrition was notably low: 96 percent of households were re-interviewed at the five-year endline, and 94 percent of individual respondents.&lt;/p&gt;
&lt;p&gt;Five years after the screening, treated households show meaningfully larger investment across three domains relative to the control group, with all headline results significant at 5 percent or less and robust to multiple hypothesis testing. First, on agricultural effort and investment: treated household heads and spouses work approximately one extra hour per day on their own farms (roughly 8.6 percent of the control mean per spouse). Treated households are 10 percentage points more likely to have adopted modern crop inputs (improved seeds, inorganic fertilizer) and 10 percentage points more likely to have invested in modern livestock inputs (feed, veterinary supplies). Holdings of productive tools are 20 percent higher than in the control group. Second, on educational investment: treated households spend approximately 36 percent more on children&amp;rsquo;s schooling than the control group. Among children who were of school-going age at the time of the intervention (aged 11–15 then, 16–20 at endline), the number completing full primary school is nearly double the control rate (0.16 per household versus 0.07 in the control). Third, on living standards: treated households experienced 0.33 to 0.38 fewer months of food insecurity in the previous year. Their holdings of consumer durables (furniture, kitchenware, phones) are 29 percent higher than the control group in value. Estimated house values are 27 percent higher. However, there is no statistically significant effect on measured food or frequent non-food consumption expenditure, a finding the authors interpret as consistent with households continuing to divert resources toward future-oriented investments rather than current consumption.&lt;/p&gt;
&lt;p&gt;The intervention&amp;rsquo;s effects appear to operate primarily through aspirations — defined in this paper as desired goals for the future that motivate investment and effort. Treated households report significantly higher aspirations and expectations for income, assets, and children&amp;rsquo;s education five years later. By contrast, the paper finds no persistent changes in time preferences, risk preferences, grit, or beliefs about returns to technology. Locus of control shifted six months after the intervention but did not persist to the five-year endline, and the authors argue that if locus of control were the operative mechanism, investment effects would also have dissipated. The placebo group shows no significant effects relative to the control, ruling out screening exposure or social attention as mechanisms.&lt;/p&gt;
&lt;p&gt;The paper is explicit about scope conditions. The study area was deliberately chosen for its extreme remoteness and media isolation, and the authors caution that this may have amplified the intervention&amp;rsquo;s salience and persistence relative to less isolated populations. External validity beyond comparable settings is uncertain. A back-of-the-envelope cost-effectiveness calculation finds that increases in durable asset holdings alone outweigh intervention costs by a factor of approximately two at reasonable scale.&lt;/p&gt;
&lt;p&gt;Q: What was the intervention and what made it distinct from other role model studies?
A: Treated households were invited to watch four 15-minute documentary films featuring real rural individuals from similar socioeconomic backgrounds who had escaped poverty through goal-setting, perseverance, and hard work. The films were produced in Oromiffa, the local language, and featured two male and two female role models depicting achievable actions such as installing irrigation or starting a small business. Unlike studies that vary exposure to in-person mentors or peers, participants received no ongoing mentorship, financial resources, or support of any kind beyond the single video screening, isolating the aspirations channel from material or informational transfers.&lt;/p&gt;
&lt;p&gt;Q: How were aspirations measured and validated?
A: Aspirations were measured using locally validated survey instruments (Bernard and Taffesse, 2014) that asked respondents what level of annual income, asset wealth, and oldest child&amp;rsquo;s education they would like to achieve in their lifetime. Test-retest reliability over two weeks produced within-respondent correlations of 0.77 to 0.98 across domains, which the authors benchmark against Angrist and Krueger (1999) standards for reliable income and education measures. The measures correlated in expected directions with wealth: mean income aspirations in the upper wealth tercile were 1.5 times those in the lower tercile, and asset aspirations in the upper tercile were 1.9 times those in the lower tercile.&lt;/p&gt;
&lt;p&gt;Q: What were the five-year effects on agricultural effort and investment?
A: Treated household heads and spouses worked approximately half an hour more per day each on their own farms relative to control, implying roughly one extra hour per day across the typical household&amp;rsquo;s adult members — an 8.6 percent increase over the control mean. Treated households were 10 percentage points more likely to have adopted modern crop inputs and 10 percentage points more likely to have invested in modern livestock inputs. Holdings of productive tools were 20 percent higher in value than in the control group. The overall agricultural investment index increased by 0.21 standard deviations relative to the control and 0.18 standard deviations relative to the placebo.&lt;/p&gt;
&lt;p&gt;Q: What were the five-year effects on children&amp;rsquo;s education?
A: Among children aged 16 to 20 at endline (who were 11 to 15, upper primary school age, at the time of the intervention), the number per household completing full primary school nearly doubled: 0.16 in the treatment group versus 0.07 in the control. These children in treated households also spent on average 33 minutes more per day attending school than the control group. Across all children, schooling expenditures in the treatment group were 36 percent higher than in the control and 30 percent higher than in the placebo. The education index increased by 0.25 standard deviations relative to the placebo and 0.21 standard deviations relative to the control.&lt;/p&gt;
&lt;p&gt;Q: Why did consumption expenditure not increase despite improvements in assets and food security?
A: The authors argue that the consumption result is theoretically ambiguous: if treated households continue to divert resources toward future-oriented investments (savings, productive assets, durable goods, housing), intertemporal substitution effects could offset income effects within the five-year observation window. The measured consumption variables — food and frequent non-food spending — do not capture the service flow value of accumulated durables or housing improvements, both of which increased substantially. The authors interpret this as evidence that households were still in an investment phase rather than having converted accumulated wealth into current consumption by endline.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports aspirations as the operative mechanism rather than alternative channels?
A: The treatment group had significantly higher aspirations and expectations for income, assets, and children&amp;rsquo;s education at the five-year endline, while the placebo group did not. Measured time preferences, risk preferences, grit, and beliefs about returns to technology were all statistically unchanged for treated households. Locus of control shifted six months post-intervention but did not persist to five years, and the authors note that if locus of control were the driver, investment effects would also have dissipated alongside it. The null placebo effect rules out screening exposure, social attention, or information salience from outside facilitators as mechanisms.&lt;/p&gt;
&lt;p&gt;Q: How were locus of control and fatalistic beliefs assessed in this population?
A: The sample scored twice as high as Western samples on the classic Levenson (1981) fatalism scale. On the Feagin (1975) scale of perceived causes of poverty, the sample was more likely to attribute poverty to structural or fatalistic explanations than Western samples, and both measures of fatalistic beliefs were higher among poorer households within the sample. The study region&amp;rsquo;s worldview — rooted in traditional Waaqeffannaa religion, local variants of Orthodox Christianity (Fekade Egziabher), and Islam (Qadar) — emphasizes deference to authority, predestination, and resistance to change, providing qualitative grounding for the aspirations deficit being targeted.&lt;/p&gt;
&lt;p&gt;Q: What were the effects on food insecurity and subjective wellbeing?
A: Treated households reported 0.33 fewer months of food insecurity in the previous year relative to the control group (from a base of 2.71 months in the control), and 0.38 fewer months relative to the placebo. Treated participants scored approximately a quarter of a step higher on the Cantril ladder of self-reported wellbeing than the control group. There was no significant difference on the USDA food insecurity questionnaire, which the authors attribute to that scale&amp;rsquo;s unsuitability for households that consume largely from own production.&lt;/p&gt;
&lt;p&gt;Q: What were the effects on durable goods and housing?
A: Treated households reported 29 percent higher value of consumer durables (furniture, kitchenware, phones) than the control group and 32 percent higher than the placebo. Estimated house replacement values were 27 percent higher than the control and 21 percent higher than the placebo. Enumerators directly observed that treated households were more likely to have their own toilet facility, though this result was not significant relative to the placebo. There were no effects on the probability of having a non-organic roof, which the authors note is an especially expensive upgrade.&lt;/p&gt;
&lt;p&gt;Q: How does the paper rule out spillover effects from treated to control households?
A: The authors collected data on a supplementary sample of non-treated villages to serve as a &amp;ldquo;pure control&amp;rdquo; and used this to run a suggestive test for spillovers from treated households to untreated households within the same village. They found little evidence of large spillover effects, although they acknowledge limitations in the power of these tests. The physical design of the screenings — held in rooms with shuttered windows, requiring tickets for entry, conducted separately from placebo screenings — also minimized contamination during the intervention itself.&lt;/p&gt;
&lt;p&gt;Q: What were the early (six-month) results and what do they suggest about the timing of effects?
A: At six months, the shorter follow-up found increases in savings and investment in education, consistent with behavioral change beginning soon after treatment. Aspirations showed positive but noisier effects at immediate post-screening and six-month follow-ups, which the authors interpret as consistent with aspirations increasing gradually as people experiment with alternative futures (Appadurai, 2004) or as demotivating beliefs shift incrementally (Carvalho et al., 2023), rather than changing abruptly. This gradual pattern is consistent with a learn-by-doing dynamic where small initial investments generate returns that further raise aspirations.&lt;/p&gt;
&lt;p&gt;Q: How does this study&amp;rsquo;s attrition and follow-up compare to the literature?
A: The five-year attrition rate was very low: 96 percent of baseline households were re-interviewed and 94 percent of individual respondents. The authors cite Bouguen et al. (2019) as a benchmark, noting this is a high tracking rate relative to recent long-run RCT follow-ups in low- and middle-income countries. The low attrition strengthens confidence that endline estimates are not contaminated by selective dropout.&lt;/p&gt;
&lt;p&gt;Q: What is the cost-effectiveness of the intervention?
A: A back-of-the-envelope calculation indicates that increases in durable asset holdings alone outweigh the costs of the intervention by a factor of approximately two at reasonable implementation scale. The authors present this as a proof-of-concept estimate, not a full social cost-benefit analysis, and caution that cost-effectiveness may differ in settings with higher baseline media exposure or less extreme isolation.&lt;/p&gt;
&lt;p&gt;Q: What are the key scope conditions limiting external validity?
A: The study district (Doba) was chosen specifically for its extreme remoteness: at baseline, only 11 percent of respondents watched TV at least weekly and no household owned a television. The authors argue this isolation likely made the screening event especially salient and memorable, potentially amplifying effects relative to what would be expected in less isolated contexts. They are explicit that the findings represent a proof of concept for the aspirations mechanism and that effect magnitudes should not be assumed to replicate in settings with higher baseline media exposure or different cultural belief systems.&lt;/p&gt;
&lt;p&gt;Aspirations: Defined in this paper as desired goals for the future that motivate investment and effort in order to attain them (following Bandura, 1977; Locke and Latham, 1990). Measured via validated survey instruments asking respondents the level of income, assets, or children&amp;rsquo;s education they would like to achieve in their lifetime — distinct from expectations (what one expects to achieve) and from the village maximum (what one believes the most successful person in the village could achieve).&lt;/p&gt;
&lt;p&gt;Aspirations gap: The difference between an individual&amp;rsquo;s aspired level of income, assets, or education and their current reported level. Median aspirations gaps in the sample are 55 percent of median wealth aspirations and 58 percent of median income aspirations, indicating that aspirations exceed current levels by meaningful but not unrealistic margins.&lt;/p&gt;
&lt;p&gt;Capacity to aspire: Drawn from Appadurai (2004), defined as a navigational capacity — the ability to read and navigate a map of a journey into the future. In contexts of poverty, this capacity is described as more brittle because poorer individuals have narrower social networks, fewer role models, and less material slack for experimentation with alternative futures.&lt;/p&gt;
&lt;p&gt;Role model: A real individual from a similar socioeconomic background whose documented experience of escaping poverty through goal-setting and effort provides vicarious experience that allows audience members to imagine what is possible for people like them. Role models are most effective when their success appears attainable and when the steps to achieve it are visible.&lt;/p&gt;
&lt;p&gt;Zero-sum beliefs: The belief that gains for one individual come at the expense of others in the community, documented in the study area as part of a broader fatalistic, deterministic belief system. These beliefs can suppress effort and future-oriented investment by making individual advancement appear normatively transgressive or materially impossible.&lt;/p&gt;
&lt;p&gt;Source text origin: A classification in the paper&amp;rsquo;s pipeline framework distinguishing whether a summary is based on a full working paper PDF or HTML text versus abstract-only text. Abstract-only summaries are blocked as they miss scope conditions, quantitative results, and the full argument structure.&lt;/p&gt;
&lt;p&gt;Placebo group: Households randomly invited to watch an Ethiopian comedy entertainment program (with no aspirational content) rather than the role model documentaries. Used to separate the effect of the aspirations content from the effects of the screening event itself, exposure to outside facilitators, or social attention accompanying selection for the intervention.&lt;/p&gt;</description></item><item><title>The Long-Run Impacts of Public Industrial Investment on Local Development and Economic Mobility: Evidence from World War II</title><link>https://macropaperwarehouse.com/papers/the-long-run-impacts-of-public-industrial-investment-on-local-development-and-economic-mobility-evidence-from-world-war-ii/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-long-run-impacts-of-public-industrial-investment-on-local-development-and-economic-mobility-evidence-from-world-war-ii/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Does government-led construction of large manufacturing plants in previously under-industrialized regions generate long-run improvements in regional economic development and in the lifetime earnings of the incumbent residents who were already living there at the outset? And, if so, through what mechanism — developmental improvements during childhood or expanded adult labor market opportunities?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and Identification.&lt;/strong&gt; The paper exploits the United States industrial mobilization for World War II, specifically the construction of 90 large, government-financed, newly-built manufacturing plants (each costing $10 million or more in contemporary dollars, approximately $150 million in 2020 dollars) in dispersed locations outside the major prewar manufacturing hubs. Strategic and security considerations — not economic optimization — drove the military to insist these plants be sited away from congested industrial centers. Because private firms were unwilling to finance construction in isolated locations with uncertain postwar value, the government built them directly as government-owned, contractor-operated (GOCO) facilities through the Defense Plant Corporation. Site selection within the set of sufficiently populated regions was governed by idiosyncratic, short-run factors — the immediate availability of suitable parcels, informal connections to procurement officers, and expedience — rather than systematic economic characteristics of the receiving counties. The paper documents no systematic association between publicly-funded wartime plant construction and prewar county-level economic or demographic characteristics conditional on population size, and finds parallel prewar trends and balanced outcome levels across treatment and comparison counties in all decades leading up to WWII. A placebo test using 1910-to-1940 intergenerational mobility in matched Census records confirms no differential prewar upward mobility in treatment counties.&lt;/p&gt;
&lt;p&gt;The comparison group consists of 1,400 counties outside the 100 largest prewar manufacturing counties that did not receive large public plants. Treatment assignment for individuals is based on birth county, not adult county of residence, enabling the paper to track outcomes regardless of where individuals ultimately live.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The analysis draws on the 1945 War Production Board data book for plant-level investment; county-level panels from Decennial and Economic Censuses spanning 1900–2000; the SSA NUMIDENT file (birth county and date); IRS Form 1040 individual income tax returns in 1969, 1974, 1979, and 1984 (covering wage earnings and adjusted gross income); the full-count 1940 Census (parent earnings, demographics); the 2000 Census long form (educational attainment); and W-2 earnings histories from the SSA Detailed Earnings Record matched to a CPS-linked subsample, with employer information linked to the Business Register.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regional Effects.&lt;/strong&gt; By 1970, counties receiving large public wartime plants had approximately 30 percent higher manufacturing employment, 20 percent larger populations, and 7–8 percent higher median family income than comparison counties. Manufacturing employment as a share of total employment rose and remained elevated through the 1970s before converging toward parity with the comparison group by 1990. Treated counties were permanently larger — with population stabilizing at a new, persistently higher equilibrium roughly 20 percent above comparison counties by end of century — even after the manufacturing employment share converged, consistent with path dependence and multiple equilibria. Average production worker pay in manufacturing rose by approximately 10 percent, closely tracking value-added per worker, while average retail wages rose by only one-third as much and were not statistically significant in most years. In the 40 years after the war, treated counties saw median family earnings increase by 5–10 percent, concentrated in higher average wages and employment shares in manufacturing and semi-skilled blue-collar occupations, with limited effects on non-manufacturing, white-collar occupations, or female individual income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Individual Earnings Effects.&lt;/strong&gt; Men born in treatment counties in the 18 years before the war (birth cohorts 1922–1940) earned approximately $1,200–$1,300 more per year (2020 dollars) in average wage earnings reported on 1040 returns in 1969, 1974, 1979, and 1984 — an increase of 2.5–3 percent and roughly a one-percentile rise in the national earnings distribution. Effects were largest for children of parents at the bottom of the 1939 earnings distribution: children of the lowest-income parents saw adult wage earnings rise by approximately $1,800–$2,000 per year (3–4 percent), with effects declining linearly by parent rank and effectively vanishing for children of the highest-earning parents. Black men experienced larger average earnings effects (4–6 percent, or $1,500–$2,500 in 2020 dollars) than White men (2–3 percent, or $1,000–$1,500), with the racial earnings gap estimated to have narrowed by about 2 percent in the treatment group. When examining Form 1040 returns (tax-unit level), effects are comparable for men and women, but W-2 individual earnings data from the SSA-CPS subsample show no positive effect on women&amp;rsquo;s own earnings — the 1040 effects for women are entirely driven by their husbands&amp;rsquo; higher earnings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanism.&lt;/strong&gt; The balance of evidence points to access to higher-wage jobs in adulthood as the primary channel, rather than developmental human capital improvements accumulated during childhood. War plants modestly increased male educational attainment — children from the lowest-earning families completed approximately one-quarter of a year more schooling and were 3 percentage points more likely to graduate high school — but education effects are too small to account for the full earnings increase. Critically, there is no gradient in earnings effects by birth cohort: children who were younger at the start of the war and therefore had longer childhood exposure to improved regions did not benefit more, contradicting a childhood exposure-effect mechanism as in Chetty and Hendren (2018b). Adult earnings effects are entirely accounted for by adult location: conditioning on 1979 county of residence eliminates the treatment effect. Stayers in treatment counties show large earnings differences relative to stayers in comparison counties, while movers show none. Men born in treatment counties are also directly documented to have worked in industries with higher wage premiums as adults, with coarse industry classification alone accounting for approximately one-third of the estimated log wage increase.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy Scope Conditions.&lt;/strong&gt; The paper argues these effects are specific to the WWII postwar institutional context — high global demand for U.S. manufactured goods, limited international competition, labor-intensive production techniques, and strong union bargaining power — conditions that no longer hold. Reexamination of &amp;ldquo;million-dollar plant&amp;rdquo; openings in the 1980s and 1990s shows manufacturing employment expanded but average manufacturing wages did not increase, suggesting contemporary plant openings do not generate the same high-wage opportunities. The association between manufacturing employment density and upward mobility visible in 1950 has entirely vanished by the end of the twentieth century.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-exactly-defines-the-treatment-group-and-why-were-these-plants-built-by-the-government-rather-than-private-firms"&gt;Q1. What exactly defines the treatment group, and why were these plants built by the government rather than private firms?&lt;/h3&gt;
&lt;p&gt;A: The treatment group consists of 90 counties outside the 100 largest prewar manufacturing regions that received at least one new, fully publicly-financed manufacturing plant costing $10 million or more (approximately $150 million in 2020 dollars) under the WWII industrial mobilization. Private firms refused to finance construction in dispersed, isolated locations with highly uncertain postwar value; the Air Force historians recorded that &amp;ldquo;industrialists&amp;rsquo; reluctance to invest in dispersed plant facilities was at odds with the government&amp;rsquo;s hope that private capital could finance new inland construction.&amp;rdquo; The government built and owned these facilities as GOCO plants, operated by private firms under contract. The 353 plants meeting the cost threshold (including both large and smaller public plants) account for 70 percent of all spending on new plants during the war.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-authors-establish-that-plant-siting-was-quasi-random-conditional-on-population-size"&gt;Q2. How do the authors establish that plant siting was quasi-random conditional on population size?&lt;/h3&gt;
&lt;p&gt;A: Identification rests on three forms of evidence. First, historical documents show procurement decisions were driven by idiosyncratic factors — availability of a suitable parcel, informal connections to procurement officers, short-run expedience — rather than systematic economic characteristics. Members of Congress had little ability to influence siting, and Rhode et al. (2018) find little evidence that federal politics drove the geographic distribution of wartime spending. Second, balance tests (estimating prewar county characteristics as outcomes in Equation 1) show no significant differences between treatment and comparison counties in earnings levels, demographics, manufacturing development, or industrial composition after conditioning on 1940 population, with a joint p-value of 0.30 (0.36 when also conditioning on geography and infrastructure). Third, a placebo test using children in the 1910 Census matched to the 1940 Census finds no differential economic outcomes or upward mobility rates in counties that would eventually receive treatment plants, conditional on basic region size.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-county-level-effects-on-the-structure-of-the-labor-market-in-the-medium-run"&gt;Q3. What are the county-level effects on the structure of the labor market in the medium run?&lt;/h3&gt;
&lt;p&gt;A: By the 1960s–1970s, treated counties had higher predicted union coverage rates and a greater share of men in semi-skilled production occupations, driven primarily by movement away from farm work and supplemented by higher male labor force participation. Average wages in craftsperson and operator occupations rose by 8 percent in treated counties — more than double the increase in wages for high-skill professional and managerial occupations. Treated counties had 8 percent higher median male individual incomes by 1979. Effects on female median individual income were minimal, and there were no effects on female labor force participation rates.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-estimated-magnitude-of-the-individual-earnings-effects-and-how-do-they-vary-by-parent-income"&gt;Q4. What is the estimated magnitude of the individual earnings effects, and how do they vary by parent income?&lt;/h3&gt;
&lt;p&gt;A: Men born in treatment counties averaged $1,200–$1,300 more per year in real wage earnings (2020 dollars) on 1040 tax returns across the four observation years 1969, 1974, 1979, and 1984, a 2.5–3 percent increase equivalent to roughly one percentile in the national earnings distribution. Heterogeneity by parent rank is pronounced and monotone: children of parents at the very bottom of the 1939 earnings distribution gained approximately $2,000 per year (about 4 percent), while children of the highest-earning parents experienced no significant effect. When county weighting is equalized to eliminate the differential representation of rural (lower-income) counties, effects are roughly constant across the bottom six deciles of the parent earnings distribution and then drop steeply at the top, showing that the earnings gradient was not simply an artifact of plant openings in poorer, smaller counties.&lt;/p&gt;
&lt;h3 id="q5-how-did-effects-differ-by-race"&gt;Q5. How did effects differ by race?&lt;/h3&gt;
&lt;p&gt;A: Wartime plant construction increased annual adult earnings of Black men by 4–6 percent ($1,500–$2,500 in 2020 dollars) and of White men by 2–3 percent ($1,000–$1,500 in 2020 dollars). The racial earnings gap in the treatment group is estimated to have narrowed by about 2 percent. However, the pattern of heterogeneity by parent income differs by race: for White men, effects are largest for children of below-median parents and effectively zero for children of above-median parents. For Black men, the largest effects — 7–10 percent ($4,000–$5,000 in 2020 dollars) — accrue to children of parents with earnings above the pooled-race national median, while effects for lower-income Black families range from 3–6.5 percent, suggesting that Black workers from higher-income backgrounds particularly benefited from wartime anti-discrimination policies and the opening of previously restricted manufacturing occupations.&lt;/p&gt;
&lt;h3 id="q6-why-do-the-1040-returns-show-comparable-effects-for-men-and-women-while-w-2-data-show-no-effect-on-womens-individual-earnings"&gt;Q6. Why do the 1040 returns show comparable effects for men and women, while W-2 data show no effect on women&amp;rsquo;s individual earnings?&lt;/h3&gt;
&lt;p&gt;A: Form 1040 returns are filed at the tax-unit level — for married couples, they report the combined wages of both spouses. Because more than 80 percent of women in the sample are married, an increase in a husband&amp;rsquo;s earnings raises the joint 1040 figure for both spouses. The SSA-CPS subsample with individual W-2 records shows that the entire effect on men&amp;rsquo;s Form 1040 wages directly reflects increases in their own W-2 earnings, while women&amp;rsquo;s own W-2 earnings show no positive treatment effect. This finding is consistent with county-level evidence of no impact on female individual income or female labor force participation, and with Rose (2018) finding that women were almost universally excluded from manufacturing jobs after the war&amp;rsquo;s conclusion despite high wartime female manufacturing employment.&lt;/p&gt;
&lt;h3 id="q7-what-evidence-tests-the-developmental-effects-mechanism"&gt;Q7. What evidence tests the developmental-effects mechanism?&lt;/h3&gt;
&lt;p&gt;A: Three tests argue against childhood developmental effects as the primary driver. First, educational attainment effects — while statistically significant for children of the lowest-income parents (approximately one-quarter of a year more schooling, 3 percentage points more likely to graduate high school) — are too small to account for the earnings increase: a Mincer-equation calculation shows that the education effects can explain less than one-half of the estimated effect on 1979 wages. Second, there is no gradient in earnings effects by birth cohort — children younger at the war&amp;rsquo;s start, who had longer post-treatment childhood exposure, did not benefit more, in direct contrast to the Chetty-Hendren childhood-exposure framework. Third, postwar in-migrants into treatment counties were not drawn from better-educated or higher-income families and did not themselves have more education than in-migrants into comparison regions, ruling out peer effects from selective in-migration.&lt;/p&gt;
&lt;h3 id="q8-what-evidence-directly-implicates-adult-labor-market-access-as-the-operative-mechanism"&gt;Q8. What evidence directly implicates adult labor market access as the operative mechanism?&lt;/h3&gt;
&lt;p&gt;A: Four pieces of evidence point to contemporaneous adult labor market access. First, individuals born in treatment counties lived as adults in counties with 3–4 percent higher median male earnings and higher wages in semi-skilled blue-collar occupations but not in highly-skilled professional occupations — a pattern quantitatively consistent with the individual earnings effects. Second, the entire earnings effect is concentrated among those who remain in their birth counties: stayers in treatment counties show earnings differences of similar magnitude to county-level manufacturing wage effects, while movers show no difference compared to movers from comparison counties. Third, conditioning on 1979 county of residence eliminates the earnings effect entirely (1979 location fixed effects specification). Fourth, using W-2 data matched to the Business Register in the SSA-CPS sample, men born in treatment counties are directly shown to work in industries with higher wage premiums, with coarse industry classification alone accounting for approximately one-third of the log wage increase.&lt;/p&gt;
&lt;h3 id="q9-is-the-persistence-of-regional-effects-driven-by-continued-cold-war-military-spending-at-the-plants"&gt;Q9. Is the persistence of regional effects driven by continued Cold War military spending at the plants?&lt;/h3&gt;
&lt;p&gt;A: No. The paper separates ordnance and ammunition plants — which predominantly became GOCO facilities or Air Force Bases after WWII and received disproportionately more Vietnam War-era defense spending — from general manufacturing plants, which overwhelmingly transitioned to privatized civilian production. Both types of plants show similarly persistent effects on manufacturing employment and comparable impacts on the long-run earnings of local children. Moreover, general manufacturing plants — which did not generate increased postwar military spending — had large permanent effects on overall population growth, while ordnance plants had smaller population effects. The persistence therefore does not appear to reflect continued federal expenditure.&lt;/p&gt;
&lt;h3 id="q10-what-mechanism-explains-the-permanent-population-effect-even-after-manufacturing-employment-shares-converge"&gt;Q10. What mechanism explains the permanent population effect even after manufacturing employment shares converge?&lt;/h3&gt;
&lt;p&gt;A: The authors interpret the permanent population differential — treated counties remain roughly 20 percent larger than comparison counties even at the end of the 20th century, after manufacturing employment shares converge — as evidence of path dependence and multiple equilibria. Once a region reaches a new, larger equilibrium, self-sustaining forces (expanded non-tradable employment, public infrastructure investment) maintain it. Treatment counties are more likely to have been connected to the interstate highway system in subsequent decades and show positive effects on local government capital outlays for utilities. The medium-term persistence is attributed partly to the sunk costs of site establishment (surveying, local approvals, infrastructure connections), which make reinvestment at existing sites more attractive than greenfield construction elsewhere.&lt;/p&gt;
&lt;h3 id="q11-do-smaller-plant-openings-generate-comparable-effects"&gt;Q11. Do smaller plant openings generate comparable effects?&lt;/h3&gt;
&lt;p&gt;A: No. Counties receiving smaller publicly-financed plants costing between $1 and $10 million show no detectable effects on manufacturing employment, population, median family income, or individual adult earnings comparable to those from the large plants. The authors cannot rule out the presence of small effects, but the null results for smaller plants — combined with evidence that the largest effects are in counties with the highest investment intensity per 1940 resident — are consistent with threshold effects (&amp;ldquo;big push&amp;rdquo;) in regional development, though the wide confidence intervals do not allow the authors to conclusively distinguish threshold effects from a linear-in-investment model.&lt;/p&gt;
&lt;h3 id="q12-what-do-modern-million-dollar-plant-openings-reveal-about-the-contemporary-relevance-of-these-findings"&gt;Q12. What do modern &amp;ldquo;million-dollar plant&amp;rdquo; openings reveal about the contemporary relevance of these findings?&lt;/h3&gt;
&lt;p&gt;A: Reexamining plant openings from Greenstone et al. (2010) using an event-study design, the authors find that 1980s–1990s million-dollar plant openings expanded manufacturing employment (consistent with Greenstone et al.) but had no impact on average manufacturing wages — in sharp contrast to the WWII findings. Slattery and Zidar (2020) similarly find no impacts on county-level incomes for plant openings since 2000. The correlation between manufacturing employment density and upward mobility rates visible in 1950 had entirely vanished by the end of the 20th century. The authors attribute the divergent results to the changed institutional environment: contemporary production is highly automated, relies on interchangeable labor from staffing agencies, faces intense international competition, and is conducted under much weaker collective bargaining institutions.&lt;/p&gt;
&lt;h3 id="q13-what-is-the-papers-assessment-of-aggregate-welfare-implications"&gt;Q13. What is the paper&amp;rsquo;s assessment of aggregate welfare implications?&lt;/h3&gt;
&lt;p&gt;A: The paper is explicit that its local estimates do not allow clean conclusions about aggregate effects. Publicly-financed plant construction in peripheral locations may have crowded out private investment that would otherwise have occurred in major manufacturing hubs. If so, the documented regional gains represent geographic reallocation of manufacturing activity rather than a net increase in the aggregate plant stock. Aggregate gains from reallocation would require that the benefits in the selected dispersed locations exceeded what would have occurred in the counterfactual locations — a plausible conjecture given the paper&amp;rsquo;s evidence that effects are larger in counties with lower prewar manufacturing employment shares and lower initial market access, but one the authors cannot demonstrate decisively.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Government-Owned, Contractor-Operated (GOCO) Plants:&lt;/strong&gt; Manufacturing facilities built and owned by a U.S. government agency (typically the Defense Plant Corporation) during WWII but built and operated by private firms under cost-plus contracts. GOCO status meant the government bore full construction risk and that post-war disposition (sale to private buyers at a fraction of construction cost, or continued GOCO operation for ordnance production) was determined by public agencies, not by the constructing firm&amp;rsquo;s investment calculus.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Place-Based Predistribution:&lt;/strong&gt; The paper&amp;rsquo;s term for the mechanism by which wartime plant construction raised the incomes of existing residents — not through ex-post redistribution of income via taxes and transfers, but by expanding the set of high-wage employment opportunities available to incumbent workers in the region, thereby changing the pre-tax, pre-transfer wage structure facing those workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Adult Labor Market Access (vs. Childhood Developmental Exposure):&lt;/strong&gt; A distinction the paper draws in explaining why children born in treated counties had higher adult earnings. The &amp;ldquo;developmental exposure&amp;rdquo; mechanism (as in Chetty and Hendren 2018b) implies benefits scale with the amount of time spent in an improved childhood environment. The &amp;ldquo;adult labor market access&amp;rdquo; mechanism means children benefit irrespective of years of childhood exposure because they can access improved local labor market conditions when they reach working age as adults — what the paper operationalizes through the finding that earnings effects are entirely accounted for by 1979 county of residence and are concentrated among individuals who remain in their birth counties.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Upward Mobility (Absolute and Relative):&lt;/strong&gt; Following Chetty et al. (2014), the paper uses both concepts: absolute upward mobility means children from low-income backgrounds have higher lifetime earnings than comparable children in counterfactual regions; relative upward mobility means their outcomes converge toward those of children from affluent backgrounds. The paper documents both: large earnings effects for the lowest parent-income deciles, declining linearly to zero for the top deciles.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conditional Independence (Plant Siting as Quasi-Random):&lt;/strong&gt; The paper&amp;rsquo;s identification assumption — that among counties with observably similar population sizes and basic geographic/infrastructure characteristics, the specific choice of plant siting locations was driven by idiosyncratic, short-run factors uncorrelated with potential postwar outcomes. This is a level-balance assumption (not merely a parallel-trends assumption), required because individual outcomes are only observed in the post-period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Industry Wage Premium:&lt;/strong&gt; The paper uses Krueger and Summers (1988) estimates of inter-industry wage differentials (the portion of a sector&amp;rsquo;s average wage unexplained by worker characteristics) to classify adult employers of treated individuals. Finding that men born in treatment counties work at employers in higher-premium industries — with industry category alone explaining approximately one-third of the log wage increase — provides direct evidence of the adult labor market access mechanism operating through industry sorting.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Path Dependence / Multiple Equilibria in Regional Development:&lt;/strong&gt; The paper documents that treated counties remain permanently larger in population than comparison counties even after manufacturing employment shares converge and the original plants begin to close. This self-sustaining population differential, inconsistent with a unique spatial equilibrium, is interpreted as evidence that the temporary wartime shock shifted treated regions into a permanently higher equilibrium, sustained by subsequent infrastructure investment and non-tradable sector expansion proportional to the larger population base.&lt;/p&gt;</description></item><item><title>The Macroeconomic Impact of Climate Change: Global Versus Local Temperature</title><link>https://macropaperwarehouse.com/papers/the-macroeconomic-impact-of-climate-change-global-versus-local-temperature/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-macroeconomic-impact-of-climate-change-global-versus-local-temperature/</guid><description>&lt;p&gt;The paper shows that the macroeconomic impact of climate change is &lt;strong&gt;an order of magnitude larger&lt;/strong&gt; than what standard country-level panel estimates suggest. The key identification innovation is to measure the effect of global mean temperature shocks using time-series local projections, rather than using cross-country variation in local temperatures as in the conventional panel literature. A shock to global mean temperature tracks extreme weather events (droughts, heat waves, wind, precipitation anomalies) that affect all countries simultaneously; a local temperature anomaly in one country does not.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical approach&lt;/strong&gt;: The authors estimate local projections of world GDP growth on exogenous global mean temperature shocks. The shock is the innovation to global mean temperature after removing a 2-year autoregressive component and a low-frequency trend, following Hamilton (2018). Two estimation samples: &lt;strong&gt;BU&lt;/strong&gt; (Barro-Ursúa macro history, 43 countries, 1860–2019) and &lt;strong&gt;PWT&lt;/strong&gt; (Penn World Tables, 173 countries, 1960–2019).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Key empirical results&lt;/strong&gt; (Section 3): A 1°C shock to global mean temperature causes world GDP to fall by &lt;strong&gt;14% after 6 years&lt;/strong&gt; in the PWT sample (95% CI: 6%–22%); significant at the 5% level in years 2–8; does not mean-revert within the 10-year sample horizon. In the BU sample, the peak GDP decline is &lt;strong&gt;18% after 5 years&lt;/strong&gt; (95% CI: 6%–30%). Converting the cumulative IRF ratio to a permanent temperature change yields a &lt;strong&gt;22–34% long-run GDP decline per 1°C&lt;/strong&gt; of permanent global warming (PWT and BU respectively). By contrast, local temperature shocks — estimated from a standard cross-country panel with country and year fixed effects — generate effects of &lt;strong&gt;1–3% per °C&lt;/strong&gt;, not statistically significant at the 5% level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why global &amp;gt; local&lt;/strong&gt; (Section 4): Four categories of extreme climatic events (heat waves, droughts, wind, precipitation anomalies) jointly account for roughly &lt;strong&gt;half&lt;/strong&gt; of the estimated global temperature effect on GDP. None of these are strongly correlated with local temperature anomalies because extreme weather reflects ocean-atmosphere dynamics (El Niño/ENSO) that elevate global mean temperature rather than any single country&amp;rsquo;s local temperature. In addition, capital and investment both decline persistently after global temperature shocks (capital response significant at 5% level), and warm/low-income countries are disproportionately affected.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural model&lt;/strong&gt; (Section 5): A parsimonious neoclassical growth model embeds climate change as aggregate TFP changes. Households maximize ∫e^{−ρt}U(C_t)dt; firms use Cobb-Douglas technology Z_t K_t^α L_t^{1−α}. The damage function governing TFP is:&lt;/p&gt;
&lt;p&gt;Z_t = Z_0 exp( ∫&lt;em&gt;0^t ζ_s T̂&lt;/em&gt;{t−s} ds )&lt;/p&gt;
&lt;p&gt;where T̂_t is excess global mean temperature above baseline and ζ_s = A(e^{−Bs} − e^{−Cs}) is the structural damage function. When ζ_s → 0, shocks have level but not growth effects; no statistically significant evidence of growth effects is found in Figure 3 of the paper. The model is calibrated with: risk aversion γ = 1 (log utility), capital share α = 0.33, annual capital depreciation δ = 0.08, and pure time preference ρ = 0.02. &lt;strong&gt;Proposition 1&lt;/strong&gt; (model inversion) shows that, to first order, ŷ_t = ẑ_t + α ∫K_{t,s} ẑ_s ds, where K_{t,s} is the sequence-space Jacobian of the neoclassical growth model. This delivers identification: observed output impulse responses recover the structural TFP damage function ζ_s without imposing functional form on the capital channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Estimation results&lt;/strong&gt; (Section 5.3, Figure 12): The estimated damage function implies a &lt;strong&gt;4% peak short-run productivity decline 2 years after&lt;/strong&gt; a 1°C transitory global temperature shock; the effect decays slowly and remains significant for up to 10 years. The capital response (non-targeted moment) closely matches its empirical counterpart, providing an overidentification check. The local temperature damage function, estimated by targeting the local-panel output IRF, peaks at only &lt;strong&gt;0.5%&lt;/strong&gt; and is &lt;strong&gt;more than 8× smaller&lt;/strong&gt; in cumulative productivity effect; it is not statistically different from zero at the 5% level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Business-as-usual counterfactual&lt;/strong&gt; (Section 6.1–6.2): Temperature rises from 2024, reaching &lt;strong&gt;3°C above preindustrial by 2100&lt;/strong&gt; (asymptoting to 3.3°C), equivalent to 2°C of additional warming since 2024. Under the global temperature damage function:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;World output by 2050: &lt;strong&gt;−28%&lt;/strong&gt; vs. no-warming baseline&lt;/li&gt;
&lt;li&gt;World output by 2100: &lt;strong&gt;−53%&lt;/strong&gt; (accumulated TFP losses reach −40%)&lt;/li&gt;
&lt;li&gt;Capital by 2100: &lt;strong&gt;−51%&lt;/strong&gt; (investment initially rises as households anticipate lower permanent income, then decumulates rapidly)&lt;/li&gt;
&lt;li&gt;Consumption by 2100: &lt;strong&gt;−53%&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;2024 welfare loss (consumption equivalent): &lt;strong&gt;35%&lt;/strong&gt;; welfare continues declining as temperatures rise, eventually reaching &lt;strong&gt;56%&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;95% CI for 2100 output loss: &lt;strong&gt;29%–77%&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;All effects statistically significant at the 5% level&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Under the local temperature damage function with the same warming scenario: long-run output declines only &lt;strong&gt;9%&lt;/strong&gt;, welfare loss is &lt;strong&gt;5%&lt;/strong&gt;, and neither is statistically significant at the 5% or 10% level — consistent with conventional estimates (Nordhaus 1992, Dell et al. 2012, Burke et al. 2015).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Cost of Carbon&lt;/strong&gt; (Section 6.2, Panel F): The SCC is defined as the consumption-equivalent amount households would pay at time 0 to avoid one additional ton of CO2, using the temperature-response function from Dietz et al. (2021a). Baseline result: &lt;strong&gt;$1,207 per ton&lt;/strong&gt; (2024 international dollars), more than &lt;strong&gt;6× larger&lt;/strong&gt; than the $185/ton estimate in Rennert et al. (2022). 95% CI: &lt;strong&gt;$399–$2,015 per ton&lt;/strong&gt;. Climate sensitivity range (half/double median): &lt;strong&gt;$600–$2,400 per ton&lt;/strong&gt;. BU sample (larger damage functions): &lt;strong&gt;&amp;gt;$1,500 per ton&lt;/strong&gt;. Using the local temperature damage function yields an SCC of only &lt;strong&gt;$149/ton&lt;/strong&gt;, consistent with conventional estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sensitivity&lt;/strong&gt; (Section 6.4): Higher time preference ρ &amp;gt; 0.04 lowers welfare losses below 20% and the SCC below 3× conventional high-end estimates — the only scenario where results converge toward prior estimates. Near-Stern discount rates (ρ → 0): welfare loss &amp;gt;40% and SCC &amp;gt;$2,500/ton. A 6°C-by-2100 scenario yields welfare losses &amp;gt;60%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Historical growth accounting&lt;/strong&gt; (Section 6.3): Starting the model in 1960 and imposing the realized 1960–2019 warming path, then holding temperature constant at its 2019 level, reveals:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;World GDP per capita would be &lt;strong&gt;25% higher today&lt;/strong&gt; without warming since 1960&lt;/li&gt;
&lt;li&gt;By 2040, output is &lt;strong&gt;32% below potential&lt;/strong&gt; from past warming — one-quarter of losses from historical warming are yet to materialize (due to delayed damage function and transitional capital dynamics)&lt;/li&gt;
&lt;li&gt;Climate change reduced the annual world growth rate by as much as &lt;strong&gt;a third of baseline&lt;/strong&gt; by the 21st century&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Policy implication&lt;/strong&gt;: Most decarbonization interventions cost ~$80/ton on average (Bistline et al. 2023). Under conventional SCC estimates based on local temperature ($149/ton), the US Domestic Climate Cost (DCC) falls below policy cost, making unilateral emissions reduction prohibitively expensive. Under the paper&amp;rsquo;s global temperature SCC of $1,207/ton, the DCC of the United States exceeds $80/ton even accounting for the fraction of global climate benefits that accrue domestically — &lt;strong&gt;unilateral decarbonization becomes cost-effective for large economies such as the US&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions&lt;/strong&gt;: The neoclassical model abstracts from adaptation, mitigation, trade, urbanization, and endogenous emissions. The identification assumption requires that global mean temperature innovations are uncorrelated with other global economic confounders at business-cycle and trend frequencies; the paper checks robustness against alternative detrending, exclusion of WWII and COVID-19 years, El Niño/ENSO controls, and instrumental variables for temperature based on solar/volcanic forcing. The conversion from medium-run to long-run effects relies on the constrained ζ_s = A(e^{−Bs} − e^{−Cs}) functional form ruling out growth effects — consistent with the data but not formally testable beyond the 10-year horizon. Counterfactuals involve 2–3°C temperature changes substantially beyond the sample&amp;rsquo;s moderate perturbations; the model&amp;rsquo;s extrapolation may understate damages if nonlinearities exist at extreme temperatures (the authors note their conservative constrained-form approach).&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Summary of a forthcoming paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-do-global-temperature-shocks-produce-gdp-effects-an-order-of-magnitude-larger-than-local-temperature-panel-estimates"&gt;Q1. Why do global temperature shocks produce GDP effects an order of magnitude larger than local temperature panel estimates?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Global mean temperature shocks are strongly correlated with extreme weather events — heat waves, droughts, wind storms, and precipitation anomalies — that simultaneously affect all countries; these four event categories jointly account for roughly half of the global temperature effect on GDP.&lt;/strong&gt; Local temperature anomalies in a given country (as measured in standard cross-country panels with year fixed effects absorbed) are not correlated with these same events, because El Niño/ENSO and related ocean-atmosphere dynamics elevate global mean temperature without proportionally elevating any one country&amp;rsquo;s local temperature. Local panel studies also implicitly allow economic activity to shift toward cooler regions within a given year — an option unavailable when global warming affects all locations simultaneously. The resulting bias in local-panel estimates is not &amp;ldquo;aggregation bias&amp;rdquo; in the sense of Jensen&amp;rsquo;s inequality, but rather an identification problem: local panels identify a different object (the effect of temperature relative to other countries in the same year) rather than the aggregate climate impact the paper measures.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identification-strategy-and-what-are-the-main-threats"&gt;Q2. What is the identification strategy and what are the main threats?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The global temperature shock is identified as the innovation to global mean temperature after removing a 2-year AR component and a Hamilton (2018) low-frequency trend, yielding a shock orthogonal to its own recent history and to long-run trends.&lt;/strong&gt; The main threats are: (i) global business-cycle confounders (worldwide recessions that simultaneously lower activity and emissions), addressed by controlling for quadratic time trends and global aggregate demand proxies; (ii) reverse causality (economic expansion warming the atmosphere), addressed by IV estimates using solar/volcanic forcing as instruments; (iii) low-frequency correlation between climate trends and productivity growth, addressed by flexible detrending and robustness to sample period. All major specification checks generate quantitatively similar results, and the paper passes placebo tests for large global confounders (WWII, COVID-19).&lt;/p&gt;
&lt;h3 id="q3-how-does-the-structural-model-translate-medium-run-shock-responses-into-long-run-warming-effects"&gt;Q3. How does the structural model translate medium-run shock responses into long-run warming effects?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Proposition 1 (model inversion) shows that the output impulse response decomposes into a direct TFP effect ẑ_t and a capital channel ŷ_t = ẑ_t + α ∫K_{t,s} ẑ_s ds, where K_{t,s} is the sequence-space Jacobian of the neoclassical growth model (Auclert et al. 2021); this allows recovery of the structural TFP damage function {ζ_s} from the observed 10-year output IRF by non-linear least squares, without having to observe TFP directly.&lt;/strong&gt; The counterfactual for a gradually rising temperature path (BAU scenario with 2°C additional warming since 2024) is then solved via the full nonlinear model — not via the log-linearization used in estimation — because the 2–3°C excursion far exceeds the sample&amp;rsquo;s modest temperature perturbations. The capital response (non-targeted moment) closely tracks its empirical counterpart, providing a strong overidentification check that the model&amp;rsquo;s capital dynamics are correctly specified.&lt;/p&gt;
&lt;h3 id="q4-why-does-capital-initially-rise-in-the-bau-counterfactual-before-declining"&gt;Q4. Why does capital initially rise in the BAU counterfactual before declining?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Following standard permanent-income logic, when households learn at date 0 that global temperatures will rise and future TFP will fall, they temporarily increase saving and investment to accumulate buffer capital before the productivity decline materializes; this front-loads some capital accumulation in the early transition years (2024–2030s), briefly pushing capital above baseline, before the accumulated TFP losses overwhelm the saving motive and capital begins an extended decline.&lt;/strong&gt; The net effect is still a 51% capital shortfall by 2100 because persistently lower TFP reduces the marginal product of capital over decades, depressing investment and allowing the capital stock to drift far below its no-warming balanced growth path.&lt;/p&gt;
&lt;h3 id="q5-how-is-the-social-cost-of-carbon-defined-and-computed"&gt;Q5. How is the Social Cost of Carbon defined and computed?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The SCC is defined as the dollar amount C such that households are indifferent between (a) a world where one additional ton of CO2 is emitted at time 0 and (b) a world in steady-state where the household has paid C at time 0 (equation 7: V^{ss}(K^{ss} − C) = V^{SCC}_0(K^{ss})).&lt;/strong&gt; The temperature response to a 1-ton CO2 pulse is taken from Dietz et al. (2021a) — temperature peaks at 0.002°C after a 1-gigaton pulse and stabilizes. The model generates the productivity path {Z^{SCC}_t} via the structural damage function, solves for equilibrium capital and consumption paths, and computes the value function V^{SCC}_0. The resulting $1,207/ton exceeds prior estimates by 6× because the global-temperature damage function implies 4% peak TFP losses per 1°C transitory shock, compared to the ~0.5% peak implied by local temperature — and the SCC is essentially the capitalized sum of these future productivity losses, so the ratio scales proportionally.&lt;/p&gt;
&lt;h3 id="q6-why-are-historical-climate-losses-so-large-if-year-to-year-warming-is-small"&gt;Q6. Why are historical climate losses so large if year-to-year warming is small?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The key is cumulation: annual warming increments are individually small (tenths of a degree), but the damage function {ζ_s} is persistent (effects last 10+ years), so each year&amp;rsquo;s increment adds a flow of persistent TFP losses that stack on top of prior increments.&lt;/strong&gt; The paper&amp;rsquo;s growth accounting shows that climate change reduced the world growth rate by up to one-third of baseline in the 21st century — a number that appears modest in any single year but, compounded over decades, translates into a 25% GDP per capita shortfall by 2019. Additionally, because the estimated damage function has a 2-year lag before peak TFP impact, a substantial share of past warming&amp;rsquo;s losses are yet to be realized — the paper estimates GDP will be 32% below its potential by 2040 even with no further warming.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-sensitivity-analysis-reveal-about-the-robustness-of-the-results"&gt;Q7. What does the sensitivity analysis reveal about the robustness of the results?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The key sensitivity is the rate of time preference ρ: at ρ = 0.02 (baseline, consistent with secular interest rate decline), welfare loss is 35%; at ρ = 0.04 (above recent market rates), welfare loss is still above 20%; only at implausibly high discount rates does the welfare loss fall below 15%.&lt;/strong&gt; The SCC is more sensitive to ρ than welfare because the SCC is a capitalized stock valuation while welfare is an annualized flow. BU sample damage functions (larger IRF) raise welfare loss to 42% and 2100 GDP loss to 61%; these represent the high end of the estimates. The climate sensitivity range ($600–$2,400/ton for the SCC) reflects uncertainty in the physics of CO2-to-temperature conversion, not in the estimated economic damage function. Across all these dimensions, the global-temperature estimates remain order-of-magnitude larger than local-temperature estimates.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-policy-implication-for-large-economies-considering-unilateral-decarbonization"&gt;Q8. What is the policy implication for large economies considering unilateral decarbonization?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The domestic decarbonization test compares the Domestic Climate Cost (DCC) — the fraction of the global SCC that accrues to the decarbonizing country — against the marginal cost of abatement (~$80/ton average, Bistline et al. 2023).&lt;/strong&gt; Under conventional local-temperature estimates ($149/ton global SCC), the US DCC falls below $80/ton, implying unilateral action destroys domestic value. Under the paper&amp;rsquo;s $1,207/ton global SCC, the US DCC comfortably exceeds $80/ton even if the US only captures a fraction of world welfare gains — because global temperature extremes (hurricanes, heat waves, droughts) strike the US directly, the DCC/SCC ratio is much higher than under local estimates where the US appears less exposed. This fundamentally changes the cost-benefit calculus for large-economy unilateral climate policy.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;global mean temperature shock&lt;/strong&gt;: a time-series innovation to world average surface temperature, identified by Hamilton (2018) detrending; captures ocean-atmosphere climate variability (El Niño/ENSO) correlated with extreme weather events affecting all countries simultaneously; the paper&amp;rsquo;s key identification variable, distinct from local temperature variation used in standard cross-country panels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;global vs. local temperature effect&lt;/strong&gt;: the paper&amp;rsquo;s central finding that the GDP effect per 1°C global mean temperature shock (14–18%) is an order of magnitude larger than the effect per 1°C local temperature shock (1–3%); the gap is explained by extreme climatic events (heat waves, droughts, wind, precipitation) that co-move with global mean temperature but not with individual countries&amp;rsquo; local temperatures.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;structural damage function&lt;/strong&gt; (ζ_s): the kernel relating excess global mean temperature T̂_{t−s} to log TFP at time t, specified as ζ_s = A(e^{−Bs} − e^{−Cs}); estimated from the PWT output impulse response via model inversion (Proposition 1); implies a 4% peak TFP loss 2 years after a 1°C transitory shock, decaying slowly over 10 years; rules out permanent growth effects consistent with the data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Cost of Carbon&lt;/strong&gt; (SCC): the one-time dollar amount households would pay at time 0 to avoid one additional ton of CO2; equals (in the linear limit) the present discounted value of all flow consumption-equivalent welfare losses from the induced warming; paper estimates $1,207/ton (2024 international dollars), more than 6× prior estimates, because the global-temperature damage function implies much larger per-degree productivity losses than local-temperature estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;committed climate losses&lt;/strong&gt;: future GDP shortfalls already locked in by past warming, arising because the estimated damage function has a delayed peak (year 2) and slow decay (10+ years) — temperature rises in recent years continue reducing productivity for the following decade; the paper estimates these committed losses alone will lower GDP 32% below potential by 2040 even with temperature held constant at 2019 levels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;BAU scenario&lt;/strong&gt;: the business-as-usual warming path used for the main counterfactual — global mean temperature reaches 3°C above preindustrial by 2100 (asymptoting to 3.3°C), implying 2°C of additional warming since the 2024 baseline; under this scenario the model implies 53% GDP loss, 51% capital loss, 53% consumption loss, and a 35% consumption-equivalent welfare loss by 2100.&lt;/p&gt;</description></item><item><title>The Macroeconomics of Irreversibility</title><link>https://macropaperwarehouse.com/papers/the-macroeconomics-of-irreversibility/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-macroeconomics-of-irreversibility/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; How does partial capital irreversibility — arising from a wedge between the purchase price and the resale (discounted) price of capital — shape the persistence and amplitude of aggregate capital fluctuations? And what is the quantitative magnitude of the capital price wedge that is needed to simultaneously reconcile micro-level investment behavior with macroeconomic propagation?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Methodology.&lt;/strong&gt; Baley and Blanco build a continuous-time investment model for a continuum of firms facing (i) idiosyncratic productivity shocks (geometric Brownian motion), (ii) fixed capital adjustment costs proportional to productivity, and (iii) a capital price wedge ω, under which firms buy capital at price p and sell at p(1−ω). The key state variable is the log capital-productivity ratio k̂. The optimal policy takes the form of an inaction region with two distinct reset points — one for upsizing (k̂*₋) and one for downsizing (k̂*₊) — instead of the single reset point that arises without the wedge.&lt;/p&gt;
&lt;p&gt;Their central innovation is the Cumulative Impulse Response (CIR): the cumulative deviation of average capital-productivity ratios following a small, permanent, unanticipated aggregate productivity shock. They show the CIR can be expressed analytically through three sufficient statistics derived entirely from the steady-state cross-sectional distribution of k̂ and capital age a: (i) Var[k̂], (ii) Cov[k̂, a], and (iii) an &amp;ldquo;irreversibility term&amp;rdquo; reflecting how idiosyncratic shocks change the anticipated direction of the next adjustment. Because idiosyncratic and aggregate shocks enter the law of motion symmetrically, steady-state moments encode the aggregate propagation.&lt;/p&gt;
&lt;p&gt;To handle the path dependence introduced by the dual reset points, they condition all behavior on the previous reset (upsizing or downsizing) and characterize transitions across reset points via a Markov chain. They then derive explicit mappings from observable microdata — size and direction of investment adjustments, duration of inaction spells, and cross-spell transition probabilities — back to the unobservable capital-productivity distributions and sufficient statistics. These mappings require no revenue or productivity data; investment actions alone suffice.&lt;/p&gt;
&lt;p&gt;They extend the baseline model to a generalized hazard framework (stochastic, asymmetric fixed costs), enabling the model to match the full empirical investment-rate distribution, and apply everything to annual establishment-level manufacturing data from Chile (Encuesta Nacional Industrial Anual, 1980–2011), restricting to plants observed for at least ten years with more than ten workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings.&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Price wedge estimate.&lt;/strong&gt; A capital price wedge of ω = 0.12 (12%) is selected as the preferred value because it maximizes joint consistency between the model&amp;rsquo;s predicted CIR decomposition and the data, while also matching the distribution of investment rates. At ω = 0 the model generates a CIR of 0.92 and a negative covariance term, inconsistent with the data. At ω = 0.18 the aggregate CIR level (2.39) is close to data (2.33) but the decomposition diverges. At ω = 0.12, the CIR is 1.93 and the decomposition into sufficient statistics closely mirrors the data structure.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Irreversibility doubles persistence.&lt;/strong&gt; In the analytically tractable case of zero drift and only a price wedge (no fixed costs), the CIR equals exactly twice the ratio Var[k̂]/σ², compared to the single fixed-cost case. This means irreversibility doubles the persistence of aggregate capital fluctuations for a given cross-sectional dispersion. More generally, under the calibrated model, a 1% decrease in aggregate productivity generates a nearly 2% cumulative deviation of average capital-productivity ratios from steady state. Without irreversibility, the CIR collapses to approximately 1.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Decomposition of the CIR.&lt;/strong&gt; At ω = 0.12, the variance term Var[k̂]/σ² accounts for 72% of the CIR; the covariance term ν·Cov[k̂,a]/σ² accounts for 10%; and the irreversibility term accounts for 18%. The positive covariance (Cov[k̂,a] = 0.152 &amp;gt; 0) reflects that firms subject to downward rigidity accumulate older capital stocks above the economy&amp;rsquo;s average, amplifying persistence. This positive covariance arises because the price wedge&amp;rsquo;s downward-rigidity force dominates the drift&amp;rsquo;s negative effect.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Micro-level evidence.&lt;/strong&gt; In the Chilean data, the inaction rate is 40%. More than 96% of adjustments are positive (upsizing), fewer than 4% are negative. The probability of upsizing after a previous upsize is P⁻⁻ = 0.958; the probability of downsizing after a downsize is P⁺⁺ = 0.124. A logistic regression yields an odds ratio of 3.3, meaning a firm is more than three times as likely to purchase capital following a prior purchase than following a prior sale. The average duration of inaction conditional on a prior purchase is E⁻[τ] = 1.72 years; conditional on a prior sale it is E⁺[τ] = 1.98 years. These patterns are qualitatively consistent with the serial correlation in adjustment sign predicted by the model.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Comparison with existing wedge estimates.&lt;/strong&gt; The calibrated ω = 0.12 lies between micro-level studies based on liquidating firms (Ramey and Shapiro, 2001: ω ≈ 0.72; Kermani and Ma, 2023: ω ≈ 0.65) and structural models calibrated to static moments of investment distributions (Cooper and Haltiwanger, 2006; Khan and Thomas, 2013: ω = 0.025–0.07). The lower value relative to liquidation studies is attributed to selection effects (liquidating firms face fire-sale dynamics) and firm-internal capital reallocation that mitigates irreversibility for continuing firms.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; The analysis is a partial equilibrium characterization of transitional dynamics, maintaining constant interest rates and steady-state investment policies throughout the transition (a general equilibrium extension delivering constant prices as an equilibrium outcome is provided in Appendix D). Results apply to small, permanent, unanticipated aggregate productivity shocks; nonlinearities for shocks below 5% are found to be tiny. The empirical application is specific to Chilean manufacturing establishments, 1980–2011.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-economic-mechanism-by-which-capital-irreversibility-generates-persistence-in-aggregate-capital-fluctuations"&gt;Q1. What is the economic mechanism by which capital irreversibility generates persistence in aggregate capital fluctuations?&lt;/h3&gt;
&lt;p&gt;Irreversibility creates two distinct reset points rather than one. When a negative aggregate productivity shock hits, it shifts more firms into the downsizing region. Downsizing firms, because they have been selling capital sequentially, maintain capital-productivity ratios persistently above the economy&amp;rsquo;s average and continue to do so for multiple periods. This increases the share of firms in a persistent &amp;ldquo;downsizing phase,&amp;rdquo; which prolongs the aggregate deviation from steady state. Two channels compound: first, the population tilts toward more downsizing firms; second, their mean deviations become larger and converge more slowly. Both channels increase the CIR. Crucially, without irreversibility, firms become identical after their first adjustment and there is no additional persistence beyond what fixed costs alone generate.&lt;/p&gt;
&lt;h3 id="q2-how-are-the-three-sufficient-statistics-derived-and-what-does-each-capture"&gt;Q2. How are the three sufficient statistics derived, and what does each capture?&lt;/h3&gt;
&lt;p&gt;The CIR is characterized as a steady-state cross-sectional average of a recursive function m(k̂). Integrating over firms first and then time, and splitting each firm&amp;rsquo;s horizon at its first adjustment, yields three steady-state terms (Proposition 4). The first statistic, Var[k̂]/σ², measures how far firms allow their capital-productivity ratio to drift from the frictionless optimum — the &amp;ldquo;insensitivity of incomplete spells&amp;rdquo; to idiosyncratic productivity shocks. The second statistic, ν·Cov[k̂,a]/σ², is a bias-correction term that removes drift effects from the variance, ensuring only Brownian-shock sensitivity is captured. The third statistic, unique to the irreversibility case, measures how much idiosyncratic shocks alter the anticipated direction of the next adjustment — the &amp;ldquo;insensitivity of complete spells&amp;rdquo; — and equals the difference in expected cumulative deviations between departing and ending points of an inaction spell, scaled by duration.&lt;/p&gt;
&lt;h3 id="q3-why-is-the-cir-exactly-twice-as-large-under-pure-irreversibility-no-fixed-costs-as-under-pure-fixed-costs-for-a-given-level-of-dispersion"&gt;Q3. Why is the CIR exactly twice as large under pure irreversibility (no fixed costs) as under pure fixed costs, for a given level of dispersion?&lt;/h3&gt;
&lt;p&gt;Proposition 5, case (ii) shows that with zero drift and only a price wedge, the CIR = 2 × Var[k̂]/σ², because the first and third sufficient statistics are identical and the covariance term is zero. In contrast, with only fixed costs (case (i)), the CIR = Var[k̂]/σ². The doubling arises because the price wedge generates history-dependence through the dual reset: after a firm adjusts, whether it upsized or downsized predicts its future adjustment direction. This &amp;ldquo;anticipated terminal condition&amp;rdquo; effect (captured by the third statistic) adds an equal contribution to the CIR as the pure inaction effect (the first statistic), doubling total persistence for the same cross-sectional dispersion.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-empirical-strategy-recover-the-capital-price-wedge"&gt;Q4. How does the empirical strategy recover the capital price wedge?&lt;/h3&gt;
&lt;p&gt;The price wedge cannot be identified from the investment rate distribution alone: for any price wedge ω, the generalized hazard framework can find an adjustment hazard function Λ(k̂) such that the product Λ(k̂)·g(k̂) matches the observed investment density h(Δk̂). Instead, the authors use the CIR&amp;rsquo;s sufficient statistics — specifically the covariance term and the irreversibility term — as additional discriminating moments. At ω = 0, the model produces a negative covariance (inconsistent with the positive Cov[k̂,a] = 0.152 in the data) and no irreversibility term. At ω = 0.12, all three sufficient statistics simultaneously align with their data counterparts in relative importance (72%, 10%, 18%), selecting this wedge as preferred. The CIR level at ω = 0.12 is 1.93, somewhat below the data value of approximately 2.54–2.60, but the preferred criterion is mechanistic consistency, not just level matching.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-role-of-the-markov-chain-across-reset-points-in-handling-path-dependence"&gt;Q5. What is the role of the Markov chain across reset points in handling path dependence?&lt;/h3&gt;
&lt;p&gt;Because optimal investment features serial correlation in the sign of adjustment (P⁻⁻ = 0.958 and P⁺⁺ = 0.124 in the data), firms&amp;rsquo; future behavior depends on their most recent reset point. To maintain tractability, the authors condition all densities, durations, and expectations on the previous reset (upsizing g⁻(k̂) or downsizing g⁺(k̂)). The transition matrix P encoding probabilities P⁻⁻, P⁻⁺, P⁺⁻, P⁺⁺ determines the steady-state shares of upsizing and downsizing firms (as the eigenvector of P) and the renewal weights r⁻ and r⁺ that rescale conditional densities to account for observational bias (firms with longer inaction spells contribute more to the cross-section). This Markov structure is sufficient because one adjustment erases all heterogeneity except the direction of adjustment.&lt;/p&gt;
&lt;h3 id="q6-what-do-the-microdata-mappings-recover-and-how-are-the-reset-points-identified"&gt;Q6. What do the microdata mappings recover, and how are the reset points identified?&lt;/h3&gt;
&lt;p&gt;Stage I mappings (Propositions 6–9) recover: drift ν = E[Δk̂]/E[τ]; volatility σ² from cross-spell moment E[(k̂τ&amp;rsquo; + ντ&amp;rsquo;)² − (k̂*)²]/E[τ]; conditional means E±[k̂] as midpoints of inaction spells weighted by relative adjustment size; Var[k̂] from differences in cubed stopped values; Cov[k̂,a] from variance, average age, and the dynamic covariance E[(k̂τ&amp;rsquo; − E[k̂])²τ&amp;rsquo;]/E[τ]; and the irreversibility term from differences in expected deviations at departing vs. ending reset points. Stage II (Proposition 10) recovers the two reset points k̂*₋ and k̂*₊ from optimality conditions that equalize the investment price to the expected discounted marginal product of capital during inaction plus the expected value of undepreciated capital, conditioning on the prior reset. The inner inaction region width k̂*₊ − k̂*₋ = 0.813 in the Chilean data, of which 45% is attributed to the exogenous price wedge and 55% to the endogenous response to the wedge.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-sign-of-covka-depend-on-the-price-wedge-vs-the-drift"&gt;Q7. How does the sign of Cov[k̂,a] depend on the price wedge vs. the drift?&lt;/h3&gt;
&lt;p&gt;With zero price wedge and negative drift ν &amp;lt; 0 (depreciation exceeding productivity growth), firms with older capital have capital-productivity ratios below average, yielding Cov[k̂,a] &amp;lt; 0. The drift makes old capital-productivity ratios negative. Introducing a price wedge creates downward rigidity: unproductive firms delay selling, so old firms accumulate capital-productivity ratios above average, pushing Cov[k̂,a] toward positive values. The covariance turns positive once ω &amp;gt; 0.08 (in the illustrative parametrization in Figure V). In the Chilean calibration at ω = 0.12, Cov[k̂,a] = 0.152 &amp;gt; 0, confirming that the price wedge&amp;rsquo;s effect dominates the drift&amp;rsquo;s negative effect. A positive covariance amplifies the CIR (through the second sufficient statistic with ν &amp;gt; 0).&lt;/p&gt;
&lt;h3 id="q8-what-is-the-generalized-hazard-extension-and-why-is-it-needed"&gt;Q8. What is the generalized hazard extension and why is it needed?&lt;/h3&gt;
&lt;p&gt;The baseline model with a single fixed cost θ generates an investment distribution concentrated at two mass points (purchases and sales of fixed size), which does not match the empirical distribution&amp;rsquo;s coexistence of large and small investment rates and its convex shape. The generalized hazard model replaces the deterministic fixed cost with a stochastic, state-dependent adjustment cost, parameterized by a hazard function Λ(k̂) giving the probability of adjusting per unit time at any capital-productivity ratio in the outer inaction region. This function is recovered non-parametrically from the data by fitting a Gamma distribution to the investment density and inverting the Kolmogorov Forward Equation. The generalized hazard model nests the baseline model, random fixed cost models (Thomas 2002, Khan and Thomas 2008), and asymmetric adjustment models, while preserving the sufficient statistics characterization.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-model-handle-the-problem-with-reinjection-that-arises-from-path-dependence-after-the-first-adjustment"&gt;Q9. How does the model handle the &amp;ldquo;problem with reinjection&amp;rdquo; that arises from path dependence after the first adjustment?&lt;/h3&gt;
&lt;p&gt;Without irreversibility, a firm&amp;rsquo;s initial state k̂₀ does not affect behavior after the first adjustment, because there is a unique reset point; subsequent behavior is independent of the aggregate shock magnitude. With irreversibility, firms only partially absorb the aggregate shock at the first adjustment, since the initial state affects the probability of subsequently upsizing or downsizing. In principle, one must track firms through infinitely many adjustments. The paper&amp;rsquo;s resolution (Proposition 2) is to note that the first adjustment erases all heterogeneity except the direction (upsizing vs. downsizing), allowing subsequent behavior to be summarized by just two numbers m(k̂*₋) and m(k̂*₊), combined with the transition probabilities P⁻(k̂₀) and P⁺(k̂₀). This yields a recursive formulation for m(k̂) governed by an HJB equation with two boundary conditions at the reset points, making the problem tractable.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-role-of-the-stationarity-condition-in-pinning-down-the-cir"&gt;Q10. What is the role of the stationarity condition in pinning down the CIR?&lt;/h3&gt;
&lt;p&gt;The HJB for m(k̂) has infinitely many solutions (m(k̂) + a for any constant a). The stationarity condition, requiring that the cross-sectional average of m(k̂) in steady state is zero (no fluctuations without shocks), pins down the unique solution. Economically, it says that average cumulative deviations from complete upsizing spells and complete downsizing spells must exactly balance the deviations from incomplete inaction spells. For upsizing firms, deviations are negative (they hold too little capital relative to average); for downsizing firms, deviations are positive (they hold too much capital). The stationarity condition imposes a linear relationship between m(k̂*₋) and m(k̂*₊) that together with the HJB uniquely determines the solution.&lt;/p&gt;
&lt;h3 id="q11-how-are-the-results-extended-to-assess-nonlinearities-and-robustness"&gt;Q11. How are the results extended to assess nonlinearities and robustness?&lt;/h3&gt;
&lt;p&gt;Appendix G studies nonlinearities numerically in the generalized hazard model for different signs and magnitudes of the aggregate productivity shock. The authors find tiny nonlinearities and asymmetries for productivity shocks below ε = 5%, validating the first-order approximation used throughout. Appendix E.7 provides comparative statics on the output-capital elasticity α. The model is estimated with an inaction threshold of ι = 0.01 (investment rates below 1% in absolute value are treated as inaction), consistent with Cooper and Haltiwanger (2006). The investment distribution is truncated at the 2nd and 98th percentiles to remove outliers.&lt;/p&gt;
&lt;h3 id="q12-what-broader-applicability-do-the-authors-claim-for-the-cir-sufficient-statistics-framework"&gt;Q12. What broader applicability do the authors claim for the CIR sufficient statistics framework?&lt;/h3&gt;
&lt;p&gt;The authors argue the framework applies wherever path-dependent lumpy adjustments occur, including: inventory management (with two types of ordering decisions), durable goods consumption, and labor markets with sticky wages. The key requirement is the existence of a finite number of reset points and sufficient microdata to discipline the transition probabilities across them. Future extensions noted in the paper include: analysis of other aggregate shocks (profitability, capital prices, interest rates); corporate tax reform; monetary policy interacting with investment frictions; time-varying and endogenous price wedges in secondary markets; and higher-order cross-sectional moment responses (variance, skewness of capital-productivity ratios) by choosing different functions f(k̂) for the generalized CIR.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Capital price wedge (ω).&lt;/strong&gt; The fractional discount between the purchase price of capital p and its resale price p(1−ω). In the model this creates two distinct reset points for investment (one for buying at price p, one for selling at the discounted price) and represents the core source of irreversibility. It reflects asset specificity, adverse selection, intermediary fees, and obsolescence. The preferred calibrated value for Chilean manufacturing is ω = 0.12.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cumulative Impulse Response (CIR).&lt;/strong&gt; The integral over all future dates of the impulse response function of the average capital-productivity ratio following a small, permanent, unanticipated aggregate productivity shock. It summarizes both the impact and persistence of aggregate capital fluctuations in a single scalar. Without investment frictions, the CIR is zero (firms adjust instantaneously); the calibrated CIR at ω = 0.12 is 1.93, meaning a 1% aggregate shock generates a 1.93% cumulative deviation.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;Dual reset points (k̂&lt;/em&gt;₋ and k̂&lt;/em&gt;₊).** The two levels to which firms reset their capital-productivity ratio upon adjustment: k̂*₋ after a capital purchase (upsizing) and k̂*₊ after a capital sale (downsizing). With a price wedge, k̂*₊ &amp;gt; k̂*₋, creating an &amp;ldquo;inner inaction region&amp;rdquo; [k̂*₋, k̂*₊] with path-dependent behavior. The inner inaction region width is 0.813 in the Chilean data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sufficient statistics for the CIR.&lt;/strong&gt; Three steady-state cross-sectional moments that together fully characterize the CIR up to first order: (i) Var[k̂]/σ², the scaled cross-sectional variance of capital-productivity ratios (captures insensitivity of incomplete spells to idiosyncratic shocks); (ii) ν·Cov[k̂,a]/σ², the scaled covariance of capital-productivity ratios with capital age (a drift-bias correction); (iii) the &amp;ldquo;irreversibility term&amp;rdquo; measuring how idiosyncratic shocks change the anticipated direction of the next adjustment (unique to the irreversibility case, zero without a price wedge).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Serial correlation in adjustment sign.&lt;/strong&gt; The property, implied by the dual-reset structure, that a firm is more likely to purchase capital following a prior purchase and more likely to sell following a prior sale. In the Chilean data, P⁻⁻ = 0.958 (probability of upsizing after a prior upsize) vs. P⁺⁺ = 0.124 (probability of downsizing after a prior downside), and a logistic regression yields an odds ratio of 3.3.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Generalized hazard function Λ(k̂).&lt;/strong&gt; A state-dependent adjustment probability per unit time, allowing for stochastic and asymmetric fixed costs, that generates the full empirical investment rate distribution. It replaces the single deterministic fixed cost of the baseline model. The hazard function is recovered non-parametrically from microdata by fitting a Gamma distribution to the investment density and inverting the Kolmogorov Forward Equation, conditional on the price wedge.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Renewal weights (r⁻, r⁺).&lt;/strong&gt; Weights used to construct the unconditional density of capital-productivity ratios from the two conditional densities (conditional on prior purchase g⁻(k̂) and prior sale g⁺(k̂)). They rescale adjustment shares by relative average duration, correcting for the observational bias that firms with longer inaction spells are over-represented in the cross-section: r± = (N±/N) × (E±[τ]/E[τ]).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous irreversibility.&lt;/strong&gt; The component of the inner inaction region width (k̂*₊ − k̂*₋) that arises not from the exogenous price wedge directly but from firms&amp;rsquo; endogenous responses to the wedge — specifically, the differences in expected marginal products and user costs across the two types of inaction spells. At ω = 0.12, 45% of the inner inaction region is attributed to the exogenous wedge and 55% to endogenous amplification.&lt;/p&gt;</description></item><item><title>The Origins and Control of Forest Fires in the Tropics</title><link>https://macropaperwarehouse.com/papers/the-origins-and-control-of-forest-fires-in-the-tropics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-origins-and-control-of-forest-fires-in-the-tropics/</guid><description>&lt;p&gt;This paper studies the economics of illegal tropical forest fires in Indonesia, framed as a modern counterpart to Pigou&amp;rsquo;s canonical externality example of sparks from railway engines. The central research question is whether private firms adjust their fire-setting behavior depending on the degree to which the costs of fire spread fall on themselves versus others, and what enforcement architecture shapes that adjustment.&lt;/p&gt;
&lt;p&gt;The empirical setting is Indonesia&amp;rsquo;s national forest estate, where palm oil and wood fiber concession holders use fire as a cheap land-clearance method — burning primary forest costs 44–70% less than mechanical clearance — despite the practice being illegal. The paper assembles a novel dataset of 107,334 fires across Indonesia&amp;rsquo;s major forested islands from October 2000 to January 2016, constructed from NASA MODIS daily satellite hotspot data (1 km resolution, four flyovers per day). Fire ignitions and spread paths are traced by linking contiguous pixels burning on adjacent days. This fire data is merged with geocoded concession boundaries (logging, palm oil, wood fiber), land-use classifications (protected forest, unleased productive forest, areas outside the forest estate), annual deforestation data from Hansen et al. (2013) at 30 m resolution, daily wind speed data from NOAA NCEP-DOE Reanalysis 2 interpolated to each 1 km pixel, and data on firms investigated by the Indonesian government following the 2015 fires. The main analytical sample focuses on the 39,077 fires started inside wood fiber and palm oil concessions.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s identification strategy exploits two intersecting sources of variation: (1) temporal and spatial variation in monthly wind speed, which predicts the probability and extent of fire spread — a one-standard-deviation increase in wind speed (approximately 5 km/hr) increases fire spread area by 287%; and (2) cross-sectional variation in the land-type composition of the area surrounding each ignition pixel, which determines whether spread costs would fall on the fire-setter or on others. The interaction of these two factors identifies whether firms are more cautious about igniting fires on windy days when surrounding land is their own versus when it belongs to others.&lt;/p&gt;
&lt;p&gt;Three main findings emerge. First, fires are systematically human-caused and linked to industrial land clearance. Fires are eight times more likely per hectare in oil palm and wood fiber concessions than in logging concessions. Completely deforesting a 1 km pixel increases the probability of fire ignition in that pixel in the subsequent year by 279%, and this effect reverses in the year after (two years post-deforestation), ruling out natural flammability as the explanation and confirming a deliberate slash-and-burn cycle. Fire use following deforestation falls by approximately 38% in oil palm concessions during district election years, consistent with tighter enforcement when political incentives favor suppression.&lt;/p&gt;
&lt;p&gt;Second, firms partially internalize the externalities from fire-setting. They are significantly less likely to set fires on windy days when surrounding pixels belong to their own concession rather than to others. A buffer zone entirely owned by the same concession holder reduces ignitions by 8–25% at mean wind speed, and by 22–61% at the 95th-percentile wind speed. However, firms treat neighboring concession land and unleased productive forest similarly — suggesting Coasian bargaining between concession holders is not occurring.&lt;/p&gt;
&lt;p&gt;Third, the government&amp;rsquo;s enforcement pattern shapes firm behavior. Using data on firms investigated after the 2015 fires, the paper shows the government disproportionately investigates firms whose fires burned protected areas or high-population-density land, but not those whose fires damaged other private concessions. The relative weights firms place on different land types when deciding whether to ignite fires align closely with this government punishment function, consistent with firms responding to implicit Pigouvian incentives.&lt;/p&gt;
&lt;p&gt;Counterfactual simulations show that broadening enforcement to treat all land types as the government currently treats populated areas would reduce fires by 80%; treating all land like protected forest would reduce fires by 67%. By contrast, fully Coasian property-rights solutions yield only 14% reductions, and tort reform allowing concession holders to recover damages from neighbors yields only 6%.&lt;/p&gt;
&lt;p&gt;Q: What is the core externality problem studied in this paper?
A: Firms use fire as a cheap land-clearance method, but once set, fires risk spreading beyond the igniter&amp;rsquo;s own concession onto land owned by others, creating an uncompensated externality. The decision to use fire rather than mechanical clearance is de facto a decision to impose this spread risk on third parties. The paper asks whether firms adjust this decision depending on the extent to which spread costs fall on themselves versus others, and whether government enforcement shapes that adjustment.&lt;/p&gt;
&lt;p&gt;Q: Why is Indonesia the empirical setting?
A: Indonesia holds a large share of the world&amp;rsquo;s tropical forests and is among the countries most affected by illegal land-clearing fires. The 2015 Indonesian fires alone released approximately 400 megatons of CO2 equivalent, at their peak emitting more daily greenhouse gases than all US economic activity, and caused an estimated 100,000 excess deaths across Indonesia, Malaysia, and Singapore. The palm oil industry in Indonesia and Malaysia, where fire is used extensively, accounted for 4.7% of global CO2 emissions from 1986 to 2016.&lt;/p&gt;
&lt;p&gt;Q: How are fire ignitions and spread identified in the data?
A: The paper starts from NASA MODIS daily hotspot data at 1 km resolution from October 2000 to January 2016. An iterative procedure assigns contiguous pixels burning on adjacent days to the same fire event, with a 1-pixel buffer allowing for spread detection. This yields 176,855 total fires across Indonesia, of which 107,334 remain after restricting to the major forested islands and the forest estate. The procedure may understate single-day spread since pixels burning on the same day are classified as part of the ignition area rather than spread.&lt;/p&gt;
&lt;p&gt;Q: What fraction of fires spread beyond their ignition area, and how much of the spread falls on outsiders?
A: 87% of fires burn for only one day and 89% do not spread beyond their initial ignition area. However, the largest fire in the data spread to cover 466 times its initial area, and the largest single fire burned 764 km2. Across all multi-day fires started inside concessions, 32% of the total land burned outside the initial ignition area is outside the concession where the fire began, quantifying the scale of the local externality.&lt;/p&gt;
&lt;p&gt;Q: How is wind speed used as an identification strategy?
A: Wind speed provides temporal and spatial variation in the probability that a fire will spread. A one-standard-deviation increase in wind speed (approximately 5 km/hr) increases the extent of fire spread by 287%. Because wind varies month to month and across space, while the composition of surrounding land types is fixed in the cross-section, the interaction of wind speed with surrounding land type identifies whether firms are more cautious about igniting fires when spread risk is high and spread costs would fall on their own land versus others&amp;rsquo; land.&lt;/p&gt;
&lt;p&gt;Q: What is the main result on firms&amp;rsquo; internalization of fire spread externalities?
A: Firms are significantly less likely to start fires on windy days when a larger share of the surrounding buffer zone belongs to their own concession. One additional buffer pixel in one&amp;rsquo;s own land decreases ignitions by 0.2–0.7%. A buffer zone entirely owned by the same concession holder reduces ignitions by 8–25% at mean wind speed, and by 22–61% at the 95th-percentile wind speed. This demonstrates that firms take fire spread risk into account when it threatens their own assets, but discount it when spread would damage others&amp;rsquo; land.&lt;/p&gt;
&lt;p&gt;Q: Do firms treat different types of neighboring land differently?
A: Yes. The benchmark category is unleased productive forest, which has the weakest property rights and receives the least de facto government protection. Relative to this benchmark, firms are more cautious about fire spread toward protected forest (national parks and watershed areas) and toward land outside the forest estate (typically villages and smallholders). One additional buffer pixel in protected forest versus unleased productive forest decreases ignitions by 0.9% at mean wind speed and 2.7% at the 95th-percentile wind speed; the deterrent for land outside the forest estate is even stronger at 1.6% and 4.6%, respectively. Firms treat other firms&amp;rsquo; concession land similarly to unleased productive forest, suggesting no effective private enforcement between concession holders.&lt;/p&gt;
&lt;p&gt;Q: What evidence shows fires are tied to intentional land clearance rather than natural ignition?
A: Fires are eight times more likely per hectare in oil palm and wood fiber concessions than in logging concessions, consistent with clear-cutting versus selective logging. Completely deforesting a 1 km pixel increases fire probability in that pixel in the subsequent year by 279%. Crucially, the effect reverses in the second year after deforestation — the pixel becomes less likely to burn than before — which rules out natural flammability as the mechanism and confirms deliberate slash-and-burn timing.&lt;/p&gt;
&lt;p&gt;Q: What does the electoral cycle evidence show about government enforcement?
A: Fires following deforestation fall by approximately 38% in oil palm concessions during district election years relative to the year prior to an election, and bounce back to pre-election levels in the year after. The decline is confined to productive forest zones where conversion is occurring; no electoral cycle appears in protected areas where conversion is already prohibited. This indicates that enforcement is tightened when political incentives are strong, and confirms that these fires are set intentionally and are responsive to government pressure.&lt;/p&gt;
&lt;p&gt;Q: How is the government&amp;rsquo;s de facto punishment function estimated?
A: The paper uses data on firms investigated by the Indonesian Ministry of Forestry following the 2015 fires, matching investigated firms (identified only by initials in the published list) to concession-holder names. A logistic regression of investigation probability on the land-type outcomes of a firm&amp;rsquo;s fires — conditional on total area burned — shows the government is substantially more likely to investigate firms whose fires burned protected areas or high-population-density land, but does not differentially investigate cases where fire damage is largely confined to other private concessions.&lt;/p&gt;
&lt;p&gt;Q: How closely do firm behavior and government enforcement weights align?
A: The relative weights across land types that the government applies in its investigation decisions correspond closely to the relative weights firms apply when deciding whether to ignite fires on windy days. Firms are most deterred by spread risk toward protected forest and populated areas outside the forest estate — the same categories the government prioritizes. Firms are least deterred by spread toward unleased productive forest and other private concessions — the categories the government largely ignores. This alignment is consistent with firms responding to Pigouvian-style implicit incentives generated by the government&amp;rsquo;s enforcement pattern.&lt;/p&gt;
&lt;p&gt;Q: What do the counterfactuals reveal about policy effectiveness?
A: Fully Coasian property-rights reform — where firms treat all surrounding land as their own — would reduce fires by only 14%. Tort reform enabling concession holders to recover damages from neighbors (treating neighboring concessions as own land) would reduce fires by only 6%. By contrast, uniform enforcement raising deterrence to the level currently applied to populated areas would reduce fires by 80%; applying the level currently applied to protected forest would reduce fires by 67%. An enforcement regime that perfectly prevented all fire spread outside the igniting concession would reduce area burned by only 23%; preventing spread into protected and populated areas alone would yield only a 2% reduction.&lt;/p&gt;
&lt;p&gt;Q: What do the benefit-cost ratios for fires look like?
A: The estimated external damages from the 1997/1998 Indonesian fires range from 1,286 to 6,074 USD per hectare burned (2020 USD). The average private benefit from using fire rather than mechanical clearance — accounting for fertilizers and other costs — averages approximately 52 USD per hectare (2020 USD). Benefit-cost ratios of 0.008 to 0.04 lie well below 1, indicating that the social damages from fires vastly exceed the private benefits, even though the government currently deters only the most costly categories of fire.&lt;/p&gt;
&lt;p&gt;Q: Why do Coasian private solutions perform poorly in this setting?
A: Coasian bargaining between concession holders would require them to reach agreements to bring fire use to a locally efficient level without government intervention. The evidence shows firms treat other concession holders&amp;rsquo; land essentially the same as unprotected unleased productive forest, implying that no such bargains are being struck. The counterfactual analysis confirms this: even a fully-Coasian outcome where every surrounding pixel is treated as own land would reduce fires by only 14%, because the bulk of fires occur when ignition costs to the firm&amp;rsquo;s own land are low regardless of wind speed.&lt;/p&gt;
&lt;p&gt;Q: What is the primary policy implication?
A: The most effective lever for reducing fires is not preventing spread after the fact, but rather deterring ignition in the first place by extending the enforcement regime uniformly across all land types. If firms were induced to treat all surrounding land with the same caution they currently apply toward populated areas — through broader and stronger penalties — fires would fall by 80%. This is substantially more effective than property-rights reforms, tort reforms, or targeted spread-prevention measures focused only on protected and populated areas.&lt;/p&gt;
&lt;p&gt;Externality (fire spread): In this paper&amp;rsquo;s usage, the cost imposed on third parties when a fire ignited inside one concession spreads to land owned by others. The externality is quantified as the share of area burned outside the igniting concession (32% of multi-day fire spread in the data) and the ratio of external damages (1,286–6,074 USD/ha) to private benefits (52 USD/ha) from using fire rather than mechanical clearance.&lt;/p&gt;
&lt;p&gt;Slash-and-burn (industrial scale): The two-stage land-clearance practice where valuable timber is first harvested (deforestation) and the remaining vegetation is then burned to prepare land for plantation crops. The paper establishes this cycle empirically: complete deforestation of a 1 km pixel increases fire ignitions by 279% in the following year, with the effect reversing in the second year, ruling out natural flammability.&lt;/p&gt;
&lt;p&gt;Pigouvian enforcement: Government-imposed penalties that alter private incentives to account for externalities. In this paper&amp;rsquo;s usage, the government&amp;rsquo;s de facto punishment function — which heavily weights fires spreading into protected areas and populated land — functions as an implicit Pigouvian tax, shaping which fires firms choose to avoid rather than uniformly deterring all illegal burning.&lt;/p&gt;
&lt;p&gt;Coasian bargaining failure: The absence of private negotiations between concession holders to internalize the externalities they impose on each other. The paper demonstrates this failure empirically by showing firms treat neighboring concession land no differently from unprotected unleased productive forest, indicating no effective private agreements are limiting cross-concession fire spread.&lt;/p&gt;
&lt;p&gt;Wind speed as spread risk shifter: Monthly average wind speed at each 1 km pixel, used as the time-varying component of fire spread risk. A one-standard-deviation increase (approximately 5 km/hr) increases fire spread area by 287%. The paper uses wind speed variation interacted with surrounding land type composition to identify whether firms adjust ignition decisions based on spread risk and who bears the cost.&lt;/p&gt;
&lt;p&gt;Unleased productive forest (benchmark): Land within the national forest estate that is neither in a designated concession nor in a protected zone, leaving ownership rights unclear and de facto unprotected. The paper uses firms&amp;rsquo; behavior toward this category as the baseline against which sensitivity to other land types is measured, because it attracts the least government attention and the weakest property rights.&lt;/p&gt;
&lt;p&gt;Government punishment function: The implicit weights the Indonesian government places on different types of fire damage when deciding whether to investigate a firm, estimated from logistic regression on the 2015 investigation data. The function heavily weights fires burning protected areas and high-population-density land, and places near-zero weight on damage to other private concessions, shaping which fire types firms strategically avoid.&lt;/p&gt;</description></item><item><title>Traditional Institutions in Modern Times: Dowries as Pensions When Sons Migrate</title><link>https://macropaperwarehouse.com/papers/traditional-institutions-in-modern-times-dowries-as-pensions-when-sons-migrate/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/traditional-institutions-in-modern-times-dowries-as-pensions-when-sons-migrate/</guid><description>&lt;p&gt;This paper asks whether dowry — a transfer from the bride&amp;rsquo;s family to the groom&amp;rsquo;s household upon marriage, prevalent throughout India — enables male migration by providing liquidity that compensates parents for the old-age support they would otherwise lose when sons leave the village. The core friction is that in patrilocal societies, sons traditionally co-reside with parents and share income in old age; migration disrupts this arrangement and introduces income-sharing frictions (limited commitment, information asymmetries, remittance costs). Dowry attenuates this friction by providing a liquid pool of resources at the time of marriage that the son can transfer to parents, lowering the net return to migration needed for a household to find migration optimal.&lt;/p&gt;
&lt;p&gt;The authors develop a collective household model in which parents and sons jointly maximize a Pareto-weighted utility function. The model yields six testable predictions: (1) net marriage transfers can flow in either direction; (2) parents are more likely to take from the dowry when sons migrate; (3) conditional on migration, the probability of parental taking increases in the son&amp;rsquo;s income and in parental bargaining power; (4) aggregate male migration rates are higher in districts with stronger historical dowry traditions; (5) migration responses to a reduction in migration costs are larger in dowry areas, provided migration rates are relatively low; and (6) parents who receive remittances from migrant sons are more likely to have also taken from the dowry.&lt;/p&gt;
&lt;p&gt;To test predictions 1–3 and the remittance auxiliary prediction, the authors collected two original datasets: a Destination Survey of 557 prime-age men in Gurugram (near Delhi) conducted in 2018, of whom 62% were migrants; and an Origin Survey of 2,541 households across 34 districts in six North Indian states conducted in 2020, covering 3,069 sons, 20% of whom were migrants. These are the first quantitative data on property rights over dowry in India. Across the Destination and Origin surveys, 45% and 27% of grooms&amp;rsquo; parents, respectively, took from the dowry on net. Parents of migrants are 27 percentage points (Destination) and 8 percentage points (Origin) more likely to take than parents of non-migrants. For migrant sons, a doubling of the son&amp;rsquo;s occupational score raises the likelihood of parental taking by 19 percentage points; no such relationship exists for non-migrants. When sons report that parents held veto power over the marriage — a proxy for parental Pareto weight — parents of migrant sons are 28 percentage points more likely to be net takers. Parents whose migrant son sends financial remittances are 17 percentage points more likely to have taken from the dowry (coefficient 0.168, SE 0.074).&lt;/p&gt;
&lt;p&gt;To test predictions 4 and 5, the authors use the Ancestral Characteristics data (Giuliano and Nunn 2018) to construct district-level measures of dowry tradition strength, validated against 1999 REDS and IHDS survey data, where a one-unit increase in the historical dowry measure is associated with 81–109% higher gross or net dowry payments. Using the NSS Round 64 migration module (2007–08), they find that the continuous dowry tradition measure is associated with a 2.7–3.7 percentage point increase in migration probability against a mean of 23.8%. For the highway construction identification strategy, the authors exploit the staggered rollout of the Golden Quadrilateral and North-South/East-West corridor (5,846+ km, $71 billion), using modern staggered-entry difference-in-differences estimators (Borusyak et al. 2021; Callaway and Sant&amp;rsquo;Anna 2020). Young men (ages 15–30) in dowry districts exhibit a large, significant increase in out-migration following highway construction with no pre-trends, while the effect for non-dowry males is indistinguishable from zero. Older males (ages 31–45) show no such effect in either group, consistent with the mechanism operating at marriage. The highway effects are concentrated in inter-district, employment-driven migration.&lt;/p&gt;
&lt;p&gt;Scope conditions: the migration-enabling mechanism operates through marriage-age liquidity and patrilocal support norms; results are specific to male migration in India. The model assumes parents and sons act collectively, matching is based on grooms&amp;rsquo; earning potential, and migration frictions cause income-sharing transfers to be infeasible when the son migrates.&lt;/p&gt;
&lt;p&gt;Q: What is the central hypothesis of the paper?
A: The hypothesis is that dowry, by providing a liquid transfer at the time of marriage, allows sons to compensate parents for the old-age support that would otherwise be lost when sons migrate. Because migration introduces frictions that prevent optimal post-migration income sharing between parents and sons, dowry lowers the minimum net return to migration required for the household to find migration optimal, thereby enabling more migration.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;Seeking&amp;rdquo; versus &amp;ldquo;Satisfied&amp;rdquo; distinction in the model, and why does it matter?
A: &amp;ldquo;Satisfied&amp;rdquo; parents are those whose own income plus the maximum feasible marriage transfer (bounded by the bride&amp;rsquo;s endowment dE when dowry is present) is at least as large as their consumption allocation under no migration; migration then Pareto-improves the household for any non-negative return R. &amp;ldquo;Seeking&amp;rdquo; parents have insufficient income plus endowment, so migration reduces their consumption unless the son&amp;rsquo;s return R exceeds a threshold B(d). Because dowry strictly increases the feasible transfer ceiling, B(d=1) ≤ B(d=0), meaning dowry converts some Seeking households into effectively Satisfied ones and lowers the migration threshold for the rest.&lt;/p&gt;
&lt;p&gt;Q: What share of grooms&amp;rsquo; parents actually take from the dowry, and how does migration status affect this?
A: In the Destination Survey (62% migrants), 45% of parents take from the dowry on net; in the Origin Survey (20% migrants), 27% do. Parents of migrants are 27 percentage points more likely to take in the Destination Survey and 8 percentage points more likely in the Origin Survey, consistent with the model prediction that migration increases net taking.&lt;/p&gt;
&lt;p&gt;Q: How does the son&amp;rsquo;s earnings level affect parental taking, and does this pattern hold for non-migrants?
A: For migrant sons, a 100% increase in the son&amp;rsquo;s occupational score increases the likelihood of parents taking by 19 percentage points. For non-migrant sons, the son&amp;rsquo;s occupational score has no meaningful association with taking. This asymmetry is consistent with prediction 3: when migration occurs and the alpha income-sharing channel is shut down, parents with higher-income migrant sons have a higher relative marginal return to consumption and thus take more of the dowry.&lt;/p&gt;
&lt;p&gt;Q: What is the remittance auxiliary prediction, and is it borne out in the data?
A: The model predicts that parents who receive remittances from migrant sons should also be more likely to have taken from the dowry, because households first exhaust the costless dowry transfer before making costly or risky remittances — so remittance-receiving parents are precisely those Seeking households where dowry was already taken. The data confirm this: parents whose migrant son sends financial remittances are 17 percentage points more likely to have taken from the dowry (coefficient 0.168, SE 0.074, significant at 5%) compared to parents of migrants who do not remit.&lt;/p&gt;
&lt;p&gt;Q: How is the district-level dowry tradition measure constructed and validated?
A: The measure merges the Giuliano and Nunn (2018) Ancestral Characteristics data — which uses ethnographic sources to estimate the share of each district&amp;rsquo;s current population belonging to historically dowry-practicing groups — with district-level demographic data. Validation against the 1999 REDS shows that a one-unit increase in the historical dowry measure is associated with 81% higher gross dowry payments and 109% higher net dowry payments without region fixed effects, with a still-significant 79% for net dowry including region fixed effects. Additional validation in the IHDS confirms the historical measure predicts gold payments at marriage (coefficient 0.152 without state fixed effects, 0.185 with state fixed effects).&lt;/p&gt;
&lt;p&gt;Q: What is the association between historical dowry traditions and migration in nationally representative data?
A: Using the NSS Round 64 migration module (2007–08) for males aged 15–45, against a mean migration rate of 23.8%, the continuous dowry measure is associated with a 2.66 percentage point increase in migration probability with no controls (significant at 1%), and 3.67 percentage points with full controls including state fixed effects, year-of-birth fixed effects, caste fixed effects, distance controls, and education controls (significant at 5%).&lt;/p&gt;
&lt;p&gt;Q: What is the highway construction identification strategy, and what does it show?
A: The authors exploit the staggered construction timing of the Golden Quadrilateral and NS-EW highway corridors (beginning 1999, 5,846+ km, $71 billion investment) across Indian districts, assembling new data on district-level construction timing from a complete capital projects database. Using staggered-entry event study estimators robust to heterogeneous treatment effects, they separately estimate highway effects in districts with and without strong dowry traditions. For young men aged 15–30, dowry districts show a large, significant increase in out-migration after highway construction with no pre-trends; non-dowry districts show an effect indistinguishable from zero. Older men (31–45) show no significant effect in either group.&lt;/p&gt;
&lt;p&gt;Q: Why is the age heterogeneity (15–30 vs. 31–45) in the highway results important for the mechanism?
A: The model predicts that dowry&amp;rsquo;s migration-enabling role operates at the time of marriage, when the liquid transfer is made. Men aged 31–45 at the time of highway construction would largely have already been married before the roads were built, so they cannot retroactively benefit from the new liquidity channel. Young men (15–30) are near or below marriage age and can time their marriages and migration decisions in response to reduced migration costs. The null result for older men and the strong result for younger men together confirm the marriage-time liquidity channel.&lt;/p&gt;
&lt;p&gt;Q: Why is the highway effect concentrated in inter-district rather than intra-district migration?
A: The Golden Quadrilateral connects districts to other districts, and the model&amp;rsquo;s mechanism relies on migration creating income-sharing frictions that are more severe at longer distances. Intra-district moves are shorter, less likely to disrupt co-residence and informal support arrangements, and less likely to require the dowry&amp;rsquo;s compensatory role. The concentration of effects in inter-district migration is directly consistent with the proposed channel.&lt;/p&gt;
&lt;p&gt;Q: How does the paper address concerns about pre-trends and robustness in the highway analysis?
A: The event study plots show no pre-trends in migration for either dowry or non-dowry districts prior to highway construction. Robustness checks include additional geographic controls, caste-by-year fixed effects, time-varying cultural controls, the alternative Callaway-Sant&amp;rsquo;Anna estimator, adjusted age distributions, and varying dowry tradition cutoffs at 1%, 10%, and 25% thresholds. Results are stable across these specifications.&lt;/p&gt;
&lt;p&gt;Q: What do the theory and evidence imply about the modern transformation of dowry&amp;rsquo;s function?
A: While dowry historically served as a pre-mortem bequest to the bride adapted to patrilocal society, the modern practice has evolved so that grooms&amp;rsquo; parents frequently capture the transfer. The evidence is consistent with this reallocation of property rights serving a new function: providing parents with a pension substitute when sons migrate and traditional co-residential support breaks down. The authors speculate this functional evolution may partly explain why dowry prevalence has grown despite legal bans, as declining patrilocality creates rising demand for this type of intergenerational transfer mechanism.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: The paper suggests that policies discouraging dowry — which has many well-documented negative consequences including intimate partner violence, female infant mortality, and adverse resource allocation — may be more effective if paired with expansions of formal pension programs or other mechanisms for old-age support. Without such alternatives, eliminating dowry could inadvertently reduce male migration and associated economic development benefits because the migration-enabling liquidity function of dowry would go unfilled.&lt;/p&gt;
&lt;p&gt;Q: Does the mechanism apply equally to households with both sons and daughters?
A: The theoretical appendix shows that in a household with a son and a daughter, the daughter&amp;rsquo;s dowry outflow partially offsets the son&amp;rsquo;s inflow, reducing but not eliminating the migration-enabling effect. However, the net aggregate effect on male migration remains positive because more sons live in households where sons outnumber daughters, so the dowry inflow for the son exceeds the outflow on average across the population.&lt;/p&gt;
&lt;p&gt;Dowry (in the paper&amp;rsquo;s sense): A transfer from the bride&amp;rsquo;s family accompanying marriage that in the modern Indian context is liquid at the time of the wedding and over which grooms&amp;rsquo; parents frequently exercise property rights — distinct from the traditional anthropological conception of dowry as a pre-mortem bequest to the bride.&lt;/p&gt;
&lt;p&gt;Net Taker: A groom&amp;rsquo;s parent who receives a positive net transfer from the son&amp;rsquo;s dowry (tau &amp;gt; 0 in the model), meaning the flow of dowry resources is from the son/bride&amp;rsquo;s side to the groom&amp;rsquo;s parents.&lt;/p&gt;
&lt;p&gt;Seeking vs. Satisfied parents: Model categories distinguishing parents whose consumption needs can be met from own income plus the maximum feasible marriage transfer (Satisfied, no migration distortion) from those whose needs cannot (Seeking, requiring a minimum migration return threshold B(d) &amp;gt; 0 for migration to be household-optimal).&lt;/p&gt;
&lt;p&gt;Migration friction (alpha = 0 under migration): The modeling assumption that income-sharing transfers between migrant sons and parents are infeasible or prohibitively costly due to limited commitment, information asymmetries, and remittance costs — the friction that dowry&amp;rsquo;s lump-sum transfer at marriage is designed to circumvent.&lt;/p&gt;
&lt;p&gt;Ancestral Characteristics dowry measure: The district-level variable from Giuliano and Nunn (2018) measuring the share of the current population belonging to historically dowry-practicing ethnic groups, used as a proxy for the strength of local dowry traditions.&lt;/p&gt;
&lt;p&gt;Patrilocality: The residential norm in which sons remain with or near their parents after marriage and provide old-age support — the norm whose breakdown via migration creates the income-sharing friction that dowry helps resolve.&lt;/p&gt;
&lt;p&gt;Pareto weight (theta): The weight assigned to parents&amp;rsquo; utility in the collective household problem, capturing parental bargaining power; empirically proxied by whether sons report that parents held veto power over the marriage choice.&lt;/p&gt;</description></item><item><title>Trust and Innovation Within the Firm</title><link>https://macropaperwarehouse.com/papers/trust-and-innovation-within-the-firm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/trust-and-innovation-within-the-firm/</guid><description>&lt;p&gt;This paper investigates whether and how a CEO&amp;rsquo;s inherited generalized trust enhances innovation within firms, offering a micro-foundation for the well-documented macro-level relationship between societal trust and economic growth. The author argues that trust — by inducing tolerance of failure — encourages researchers to undertake high-risk, explorative R&amp;amp;D rather than safe exploitation of known approaches.&lt;/p&gt;
&lt;p&gt;The empirical foundation is a matched CEO-firm-patent dataset covering 5,753 CEOs at 3,598 US public firms during 2000–2011, encompassing 700,000 patents and over one million inventors. CEO trust is measured as an inherited trait: each CEO&amp;rsquo;s ethnic origin is inferred probabilistically from their last name using de-anonymized US censuses from 1910–1940, and ethnic-origin-specific trust levels are drawn from the US General Social Survey (GSS), restricted to respondents in highly prestigious occupations. The resulting trust measure is the weighted average of ethnic-specific trust scores across a CEO&amp;rsquo;s likely ethnic composition.&lt;/p&gt;
&lt;p&gt;The main empirical strategy exploits within-firm variation across CEO transitions, using firm and year fixed effects to compare patenting before and after a CEO change. The identifying assumption — that the timing of CEO transitions and the new CEO&amp;rsquo;s trust level are not predicted by prior firm patenting trends — is supported by event-study tests showing flat pre-trends. A one-standard-deviation increase in CEO inherited generalized trust (equivalent to the difference between Greek and English averages) is associated with a 6.2–6.3% increase in patent filings, statistically significant at the 1% level. For the average firm, this equals approximately 1.1 additional patents annually, worth roughly $6.8 million. The effect is larger among exogenous transitions (CEO retirement or death): 8.5% in the restricted sample, and an IV estimate of 8.2%. The back-of-envelope calculation suggests this trust-innovation channel could account for approximately 37% (range: 16–58%) of the effect of trust on GDP per capita growth.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central mechanism — risk taking — is tested by examining the distribution of patent quality rather than the mean. Under the risk-taking mechanism, trust should increase the variance of R&amp;amp;D project quality, raising high-quality patents without necessarily increasing low-quality ones. Consistent with this, CEO trust raises only above-median quality patents (measured by forward citation decile), with effects increasing monotonically toward the top decile and no statistically significant effect on below-median patents. Average patent quality as measured by citation-weighted counts or patent value rises by 4–6%. Trust also disproportionately raises the share of explorative patents (those with at least 90% of backward citations outside the firm&amp;rsquo;s existing knowledge stock) by 1 percentage point over a base of 17%.&lt;/p&gt;
&lt;p&gt;The transmission channel is examined using BERT-based classification of nearly one million Glassdoor employee reviews. Under more trusting CEOs, firms exhibit stronger top-down trust sentiment (managers trusting workers), particularly among R&amp;amp;D workers and scientists. The effect materializes within the first two years of a CEO term. Director selection provides an additional transmission mechanism: under more trusting CEOs, newly appointed directors are more trusting and departing directors are less trusting.&lt;/p&gt;
&lt;p&gt;A within-CEO design using bilateral trust (toward researchers in specific countries) with CEO fixed effects addresses omitted CEO characteristics. A one-standard-deviation increase in CEO bilateral trust toward a country is associated with a 5% increase in patents by inventors in that country&amp;rsquo;s R&amp;amp;D lab, controlling for firm-by-year, CEO, and inventor-country fixed effects.&lt;/p&gt;
&lt;p&gt;The effect is strongest when CEO trust is matched to a high-quality researcher pool; in firms with mostly low-quality researchers, high trust may be counterproductive. Trust is also a substitute for R&amp;amp;D knowledge: the effect disappears when the CEO holds a non-MBA graduate degree or has prior R&amp;amp;D experience.&lt;/p&gt;
&lt;p&gt;Q: What is the main research question?
A: The paper asks whether a CEO&amp;rsquo;s generalized trust causes more and higher-quality innovation within the firm, and through what mechanism. It also asks how trust transmits from the CEO to researchers who rarely interact with the CEO directly.&lt;/p&gt;
&lt;p&gt;Q: How is CEO trust measured?
A: CEO trust is measured as an inherited trait using a two-step procedure. First, each CEO&amp;rsquo;s last name is probabilistically mapped to one or more ethnic origins using four de-anonymized US censuses (1910–1940). Second, ethnic-origin-specific trust is computed from GSS respondents in highly prestigious occupations. The CEO&amp;rsquo;s trust measure is the weighted average across ethnic compositions. This measure is shown to be more precise than an individual-level survey measure and approximately 80% as precise as a game-based measure, without introducing attenuation bias.&lt;/p&gt;
&lt;p&gt;Q: What is the baseline patent effect and how large is it economically?
A: A one-standard-deviation increase in CEO inherited trust is associated with a 6.2–6.3% increase in patent filings (statistically significant at 1%). For the average baseline firm, this is approximately 1.1 additional patents per year, valued at roughly $6.8 million. When patent quality is accounted for, the effect rises to 9.9% using citation-weighted patent count and 11.5% using patent value based on excess stock returns on grant dates.&lt;/p&gt;
&lt;p&gt;Q: Is the effect causal? What identification strategy is used?
A: The main strategy uses firm and year fixed effects, identifying the effect from within-firm variation around CEO transitions. Pre-trend tests confirm that neither the timing of CEO changes nor the new CEO&amp;rsquo;s trust level predicts prior firm patenting. Among exogenous transitions (CEO retirements and deaths), the effect is 8.5%, and an IV estimate using the predecessor&amp;rsquo;s trust as instrument yields 8.2% (significant at 10%), both comparable to the baseline.&lt;/p&gt;
&lt;p&gt;Q: What is the macroeconomic significance of the trust-innovation channel?
A: Combining the paper&amp;rsquo;s trust-to-patents estimate (0.042–0.062) with Akcigit et al.&amp;rsquo;s (2017) patents-to-GDP-growth estimate (0.026–0.066) and the cross-country trust-to-growth coefficient (0.007), the trust-innovation channel could explain approximately 37% of the effect of trust on growth, with a plausible range of 16–58%.&lt;/p&gt;
&lt;p&gt;Q: What is the mechanism linking CEO trust to innovation?
A: The conceptual mechanism is that a more trusting manager interprets researcher failure as bad luck rather than bad type, making her more likely to tolerate failure and continue employing the researcher. This increases the researcher&amp;rsquo;s incentive to pursue explorative, high-risk R&amp;amp;D over safe exploitation of known approaches. The mechanism implies a variance-increasing effect on the R&amp;amp;D quality distribution, rather than a mean shift.&lt;/p&gt;
&lt;p&gt;Q: How is the risk-taking mechanism tested against alternative mechanisms?
A: The paper examines the distribution of patent quality by citation decile. Under mean-shifting alternatives (delegation, cooperation, relational contracting), trust should raise all quality brackets. Under risk-taking, trust raises only high-quality patents. The results show CEO trust has monotonically increasing effects from low to high quality deciles, with no statistically significant effect on below-median patents, consistent only with the variance-increasing (risk-taking) mechanism.&lt;/p&gt;
&lt;p&gt;Q: What patent quality measures are used and what do they show?
A: Beyond forward citation deciles, the paper uses explorativeness (patents with at least 90% of backward citations outside the firm&amp;rsquo;s existing knowledge stock), disruptiveness (Funk and Owen-Smith, 2017), patent importance (Kelly et al., 2021), backward citations to scientific literature, and patent scope. Trust increases all these measures with statistically significant positive coefficients. The share of explorative patents rises by 1 percentage point over a base of 17%. Average citation count and patent value increase by 4–6%.&lt;/p&gt;
&lt;p&gt;Q: Does CEO trust raise R&amp;amp;D expenditure?
A: No. The coefficients from regressing R&amp;amp;D expenditure on CEO trust are neither statistically significant nor large enough to explain the innovation effect. The patent effect is also robust to controlling for R&amp;amp;D inputs, suggesting that trust affects the type of projects chosen (consistent with risk-taking) or their realized outcomes, rather than the scale of R&amp;amp;D.&lt;/p&gt;
&lt;p&gt;Q: How does CEO trust transmit to corporate culture?
A: Using BERT-based classification of nearly one million Glassdoor reviews covering 266 firms and 397 CEO terms between 2008 and 2017, the paper finds that CEO trust is associated with stronger top-down trust sentiment (managers trusting workers). The normalized effect of a one-standard-deviation increase in CEO trust on overall trust sentiment is 0.257, on top-down trust 0.531, and on bottom-up trust only 0.141 (statistically insignificant). The effect is strongest among reviewers who identify as scientists, researchers, or engineers, and materializes within the first two years of the CEO term.&lt;/p&gt;
&lt;p&gt;Q: What evidence exists for transmission via director selection?
A: Under more trusting CEOs, newly appointed directors — especially those who remain until the end of the CEO term — are more trusting, and departing directors are less trusting. The average director trust improves during the CEO&amp;rsquo;s term. Because 54% of director hirings and 46% of turnovers occur within the first two years, this change also materializes quickly, consistent with the dynamic pattern of trust culture change.&lt;/p&gt;
&lt;p&gt;Q: What is the within-CEO bilateral trust result and what does it add?
A: Using within-CEO variation in bilateral trust toward researchers from different countries (from Eurobarometer surveys), and controlling for CEO, inventor-country, and firm-by-year fixed effects, a one-standard-deviation increase in CEO bilateral trust toward a country is associated with a 5% increase in patents by inventors in that country&amp;rsquo;s R&amp;amp;D lab. This design allows CEO fixed effects, ruling out unobserved CEO-level confounders such as management style or R&amp;amp;D ability.&lt;/p&gt;
&lt;p&gt;Q: When is CEO trust counterproductive?
A: CEO trust is beneficial only when matched to a high-quality researcher environment. Using residual patent output (controlling for observable firm and CEO characteristics) as a proxy for researcher quality, the effect of CEO trust on patents, patent output per R&amp;amp;D dollar, and future sales/employment/TFP is significant only among firms in the top two quintiles of researcher quality. In firms with mostly low-quality researchers, high CEO trust may be counterproductive by failing to screen out bad researchers.&lt;/p&gt;
&lt;p&gt;Q: How does the trust effect vary by industry and CEO background?
A: The effect is ubiquitous across industries but especially pronounced in pharmaceutical and ICT firms. The timing varies: it manifests quickly in ICT (short R&amp;amp;D lag) and more slowly in pharma (long R&amp;amp;D horizon). The effect vanishes when the CEO holds a non-MBA graduate degree or has prior R&amp;amp;D experience, suggesting trust is a substitute for direct knowledge of R&amp;amp;D processes.&lt;/p&gt;
&lt;p&gt;Q: Are the results robust?
A: Yes. The paper reports 14 categories of robustness checks including alternative patent transformations, alternative trust measures (LASSO, World Value Survey, Global Preference Survey, alternative GSS questions), alternative standard error clustering, Poisson count models, restriction to granted patents, exogenous transition subsamples, modern difference-in-differences estimators (de Chaisemartin et al., 2024; Sun and Abraham, 2021; Callaway and Sant&amp;rsquo;Anna, 2021; Borusyak et al., 2024), and leave-one-ethnicity-out. The baseline result is stable across all these checks.&lt;/p&gt;
&lt;p&gt;Inherited generalized trust: The paper&amp;rsquo;s measure of a CEO&amp;rsquo;s trust disposition, defined as the probability-weighted average of ethnic-origin-specific trust levels (from the GSS) based on the CEO&amp;rsquo;s likely ethnic composition inferred from their last name and historical census records. It captures the culturally transmitted component of trust, distinct from individual-level noise.&lt;/p&gt;
&lt;p&gt;Explorative R&amp;amp;D: In the paper&amp;rsquo;s framework (building on March, 1991), research activities that involve testing untested paths, carrying high risk of failure but high potential for innovation, as opposed to exploitation of well-known approaches with low failure risk. The paper argues CEO trust encourages researchers to shift toward exploration.&lt;/p&gt;
&lt;p&gt;Tolerance of failure: A manager&amp;rsquo;s propensity to attribute a researcher&amp;rsquo;s failure to bad luck rather than bad type. Under the paper&amp;rsquo;s mechanism, a more trusting manager gives greater weight to bad luck, making her more likely to retain the researcher after failure, thereby incentivizing risk taking.&lt;/p&gt;
&lt;p&gt;Top-down trust: In the paper&amp;rsquo;s BERT-based classification of Glassdoor reviews, the direction of trust from managers toward workers (as opposed to bottom-up trust from workers toward managers). The paper finds CEO trust primarily raises top-down trust sentiment, especially among R&amp;amp;D workers.&lt;/p&gt;
&lt;p&gt;Patent explorativeness: A patent quality measure defined as the share of its backward citations that fall outside the firm&amp;rsquo;s existing knowledge stock; patents are classified as explorative if at least 90% of backward citations are outside that stock. The paper uses this as a direct measure of explorative R&amp;amp;D output.&lt;/p&gt;
&lt;p&gt;Bilateral trust: CEO d&amp;rsquo;s directed trust toward individuals from country c, computed analogously to inherited generalized trust but using Eurobarometer survey data on country-pair trust attitudes among European-origin populations. Used in the within-CEO design to control for CEO fixed effects.&lt;/p&gt;
&lt;p&gt;Variance-increasing mechanism: The paper&amp;rsquo;s characterization of the risk-taking channel, in which CEO trust raises the variance (not the mean) of the R&amp;amp;D project quality distribution by encouraging researchers to pursue high-risk, high-reward exploration. Empirically identified by the pattern that trust raises only above-median quality patents with monotonically increasing effects toward the top decile.&lt;/p&gt;</description></item><item><title>When Did Growth Begin? New Estimates of Productivity Growth in England from 1250 to 1870</title><link>https://macropaperwarehouse.com/papers/when-did-growth-begin-new-estimates-of-productivity-growth-in-england-from-1250-to-1870/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/when-did-growth-begin-new-estimates-of-productivity-growth-in-england-from-1250-to-1870/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; When did sustained productivity growth begin in England? This paper constructs new estimates of the evolution of productivity in England from 1250 to 1870, with the goal of both dating the onset of growth and using that dating to discriminate between competing theories of why growth began.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Methodological Innovation.&lt;/strong&gt; The core challenge is that real wages over this period were heavily distorted by Malthusian population dynamics. Plague-induced population collapses (most dramatically the Black Death of 1348, which killed roughly 25% of England&amp;rsquo;s population) drove enormous swings in real wages that reflect movements along a stable labor demand curve, not changes in productivity. A naive regression of wages on labor supply is therefore inconsistent, because in a Malthusian world productivity growth induces population growth, making labor supply endogenous to productivity.&lt;/p&gt;
&lt;p&gt;The authors address this by writing down and structurally estimating a full Malthusian model of the economy. Output is produced with fixed land and variable labor (and, in an extended model, capital) via a Cobb-Douglas production function. The labor demand curve equates the real wage to the marginal product of labor. Population growth is increasing in real per-capita income (the Malthus law of motion), capturing both preventive and positive checks. Productivity follows a random walk with drift, and the paper allows for two structural breaks in the average drift rate mu. Exogenous population shocks, modeled as infrequent, sizable plague draws from a beta distribution plus a Gaussian noise term, provide identification: plague shocks and productivity shocks generate observationally distinct dynamics &amp;ndash; plague shocks cause an immediate population drop that gradually reverts, while productivity shocks cause an immediate wage jump followed by a slow population rise to a new steady state. The model is estimated via Bayesian Hamiltonian Monte Carlo (Stan), and structural break dates for mu are chosen by maximizing the Bayes factor (marginal likelihood) over the observed data on real wages, population, and days worked per worker.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Key Data.&lt;/strong&gt; Real wages are from Clark (2010) unskilled building workers series. Post-1540 population is from Wrigley et al. (1997); pre-1540 population trends are from Clark (2007b) manorial records. Days worked per worker are from Humphries and Weisdorf (2019). All series are used as decadal averages.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings.&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Onset of growth: 1600.&lt;/strong&gt; Productivity growth was zero before 1600. The Bayes factor strongly favors a first structural break in mu at 1600; break dates before 1590 and after 1640 are clearly rejected.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Two-phase post-1600 growth.&lt;/strong&gt; Between 1600 and 1810, average productivity growth was 4% per decade (posterior mean; 95% credible interval approximately 2%-6%). After 1810, productivity growth accelerated sharply to 18% per decade (95% CI approximately 12%-23%). The second break date is estimated to 1810 (the only alternative not clearly rejected is 1800).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Magnitude of productivity change.&lt;/strong&gt; By the authors&amp;rsquo; estimates, productivity in England was approximately 540% higher in 1850 than in 1500. This contrasts sharply with Clark&amp;rsquo;s (2010) dual-approach TFP series, which implies essentially no change over this period. The authors attribute the discrepancy to mismeasurement in Clark&amp;rsquo;s land rent series.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Productivity growth preceded the Glorious Revolution.&lt;/strong&gt; Productivity rose by an estimated 48% between 1600 and 1680, well before the Glorious Revolution of 1688 and the English Civil War (1642-1651). This supports the view that economic change contributed to causing the bourgeois institutional reforms of the 17th century, consistent with the Marxist tradition (Hill, 1940, 1961), rather than that institutional change preceded and caused growth.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Weakness of Malthusian population force.&lt;/strong&gt; The elasticity of population growth with respect to real income (gamma) is estimated at 0.09. Combined with a slope of the labor demand curve (alpha) of 0.53, this implies a half-life of plague-induced population dynamics of approximately 150 years. A doubling of real per-capita income stimulated population growth by only 6 percentage points per decade &amp;ndash; indicating Malthusian forces were sufficiently weak to be overwhelmed by post-1800 productivity growth. The model implies that the post-1810 productivity growth rate would have produced a 28-fold long-run increase in steady-state real wages even without the Demographic Transition.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Capital extension.&lt;/strong&gt; When capital is explicitly incorporated, using rates of return on agricultural land and rent charges to infer the capital stock, results are broadly similar: productivity growth from 1600-1810 is 3% per decade and post-1810 is 14% per decade. Capital&amp;rsquo;s production function exponent is estimated at 0.18, confirming that capital accumulation explains only a modest share of growth.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; All estimates are for England specifically. The model assumes competitive factor markets, a Cobb-Douglas (or CES) production function, and a log-linear Malthusian population law of motion. Results are robust to alternative wage series (farm laborers, craftsmen, Allen&amp;rsquo;s series), alternative population sources (Broadberry et al., 2015), constant-days-worked assumption, and alternative prior distributions.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-cant-standard-ols-regression-of-wages-on-labor-supply-recover-productivity-in-this-setting"&gt;Q1. Why can&amp;rsquo;t standard OLS regression of wages on labor supply recover productivity in this setting?&lt;/h3&gt;
&lt;p&gt;In a Malthusian world, productivity growth causes population growth, which in turn raises labor supply. This means labor supply and productivity are positively correlated, biasing OLS estimates. The authors demonstrate this concretely: from 1300 to 1450 (plague era), wages and labor supply moved in opposite directions along a stable labor demand curve, while after 1630 the same data points begin shifting off that curve &amp;ndash; a pattern that OLS would confound with changes in the slope rather than shifts in the intercept.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-authors-distinguish-empirically-between-a-plague-shock-and-a-productivity-shock"&gt;Q2. How do the authors distinguish empirically between a plague shock and a productivity shock?&lt;/h3&gt;
&lt;p&gt;The two shocks generate fundamentally different dynamics. A plague shock causes an immediate, large drop in population and a corresponding spike in wages; over time, high wages induce population growth and both wages and population gradually return to their pre-plague levels. A permanent productivity shock, by contrast, causes an immediate rise in wages with no contemporaneous population change; population then slowly rises and wages partially revert until a new, higher steady-state population is reached. The model exploits these different impulse-response signatures in the joint data on wages and population to identify the two shocks separately.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-bayes-factor-evidence-for-the-1600-break-date"&gt;Q3. What is the Bayes factor evidence for the 1600 break date?&lt;/h3&gt;
&lt;p&gt;Figure 8 in the paper shows the Bayes factor for models with different first break dates (all holding the second break at 1810). The Bayes factor rises sharply from 1580 to 1600 and falls more gradually from 1600 to 1650. Break dates before 1590 and after 1640 are clearly rejected using the standard rule of thumb that a Bayes factor of 10 constitutes strong evidence. The 1600-1810 pair of break dates yields the highest marginal likelihood of any combination considered.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-papers-productivity-estimate-compare-to-clarks-2010-dual-approach-tfp-series"&gt;Q4. How does the paper&amp;rsquo;s productivity estimate compare to Clark&amp;rsquo;s (2010) dual-approach TFP series?&lt;/h3&gt;
&lt;p&gt;Clark&amp;rsquo;s series implies productivity in England was essentially unchanged between the 15th and mid-19th centuries &amp;ndash; a result the paper argues is implausible and inconsistent with Allen&amp;rsquo;s (2005) agricultural TFP estimates (which show a 162% increase in agricultural TFP between 1500 and 1850). The authors&amp;rsquo; baseline estimate implies productivity was approximately 540% higher in 1850 than in 1500. The authors conjecture that a key driver of the difference is mismeasurement in Clark&amp;rsquo;s land rent series, which appears essentially flat from 1250 to 1600 despite enormous plague-induced swings in the land-labor ratio over this period.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-malthusian-model-imply-about-engels-pause--the-apparent-stagnation-of-real-wages-during-early-industrialization"&gt;Q5. What does the Malthusian model imply about &amp;ldquo;Engel&amp;rsquo;s Pause&amp;rdquo; &amp;ndash; the apparent stagnation of real wages during early industrialization?&lt;/h3&gt;
&lt;p&gt;Between 1730 and 1800, real wages fell slightly despite what the model estimates to be substantial productivity growth. The conventional explanation is that the gains from early industrialization accrued to capitalists rather than workers. The authors offer an alternative Malthusian explanation: England&amp;rsquo;s population grew rapidly over this period, and in the Malthusian model this population growth depressed wages relative to productivity. The authors do not reject the distributional explanation but show that Malthusian forces alone are sufficient to explain the wage-productivity divergence.&lt;/p&gt;
&lt;h3 id="q6-how-quantitatively-important-are-days-worked-the-industrious-revolution-for-the-productivity-estimates"&gt;Q6. How quantitatively important are days worked (the Industrious Revolution) for the productivity estimates?&lt;/h3&gt;
&lt;p&gt;The authors find that their productivity estimates are largely insensitive to whether the Humphries-Weisdorf (2019) days-worked series or a constant-days assumption is used. The qualitative pattern &amp;ndash; zero growth before 1600, modest growth 1600-1810, rapid acceleration post-1810 &amp;ndash; and the quantitative magnitudes remain similar. What does change is the estimated slope of the labor demand curve alpha: assuming constant days makes the labor demand curve steeper. This robustness is reassuring given that the Industrious Revolution is a contested empirical phenomenon.&lt;/p&gt;
&lt;h3 id="q7-what-does-the-model-imply-about-the-speed-of-malthusian-population-dynamics-and-how-does-this-compare-to-prior-estimates"&gt;Q7. What does the model imply about the speed of Malthusian population dynamics, and how does this compare to prior estimates?&lt;/h3&gt;
&lt;p&gt;The estimated elasticity of population growth to real income gamma = 0.09, combined with alpha = 0.53, implies a half-life of population dynamics of approximately 150 years. This is consistent with but lies between prior structural estimates: Lee and Anderson (2002) find a half-life of 107 years, and Crafts and Mills (2009) find 431 years. All estimates agree that Malthusian dynamics in England were slow relative to the conceptual ideal of rapid subsistence convergence.&lt;/p&gt;
&lt;h3 id="q8-can-the-model-explain-the-post-1750-population-explosion-without-invoking-the-demographic-transition"&gt;Q8. Can the model explain the post-1750 population explosion without invoking the Demographic Transition?&lt;/h3&gt;
&lt;p&gt;Yes. The authors simulate predicted population paths from 1740 to 1860 taking real wages and days worked as given and using their estimated gamma and alpha. Despite the weak Malthusian population force, the model can explain the vast majority of the observed population growth from 6 million in 1740 to nearly 20 million in 1860 (10.4% per decade). The key mechanism is that days worked increased substantially over this period, raising per-capita income well above what real wages alone would suggest.&lt;/p&gt;
&lt;h3 id="q9-how-does-incorporating-capital-change-the-productivity-estimates"&gt;Q9. How does incorporating capital change the productivity estimates?&lt;/h3&gt;
&lt;p&gt;In the capital-augmented model, the capital stock is inferred from rates of return on agricultural land and rent charges (Clark 2002, 2010). The capital exponent beta is estimated at 0.18, indicating a modest role for capital in pre-industrial England. Average productivity growth from 1600-1810 falls from 4% to 3% per decade, and post-1810 growth falls from 18% to 14% per decade. The authors conclude that the vast majority of growth from 1600 to 1870 cannot be attributed to capital accumulation. From 1600 to 1860, the estimated capital stock grew by a factor of five (8% per decade).&lt;/p&gt;
&lt;h3 id="q10-what-theories-of-the-onset-of-growth-are-consistent-vs-inconsistent-with-the-authors-timing-evidence"&gt;Q10. What theories of the onset of growth are consistent vs. inconsistent with the authors&amp;rsquo; timing evidence?&lt;/h3&gt;
&lt;p&gt;Inconsistent: The North-Weingast (1989) view that the Glorious Revolution of 1688 was the key institutional trigger, since productivity had already risen 48% between 1600 and 1680. Also inconsistent: gradual-growth theories (Kremer 1993, Galor-Weil 2000) in which there is no discrete acceleration. Consistent: Marxist accounts (Hill 1940, 1961) that economic change drove 17th-century institutional change; Acemoglu-Johnson-Robinson (2005) accounts linking Atlantic trade enrichment to the demand for secure property rights (timing broadly consistent, though growth rates do not visibly accelerate after the Civil War or Glorious Revolution); cultural-change accounts (Mokyr, McCloskey) tracing the onset of growth to the spread of literacy and scientific rationalism around 1600; Allen&amp;rsquo;s (2009a) directed-technical-change theory linking 17th-century wage growth to the later profitability of labor-saving innovation in the Industrial Revolution.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-model-imply-about-the-long-run-real-wage-consequences-of-post-1810-productivity-growth-even-counterfactually-assuming-malthusian-forces-persisted"&gt;Q11. What does the model imply about the long-run real wage consequences of post-1810 productivity growth, even counterfactually assuming Malthusian forces persisted?&lt;/h3&gt;
&lt;p&gt;The steady-state real wage in the Malthusian model is w-bar = mu/(alpha*gamma) minus subsistence-related terms. For mu = 0.018 (the post-1810 estimate), this formula implies a long-run real wage 28 times higher than the steady state under zero productivity growth. In other words, even if the Demographic Transition had not occurred and birth and death rates had remained sensitive to income, post-1810 productivity growth was fast enough relative to the weak Malthusian force to generate substantial sustained rises in living standards.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Labor demand curve (in the paper&amp;rsquo;s sense).&lt;/strong&gt; The equilibrium relationship between real wages and labor supply derived from competitive profit maximization by landowners facing a fixed land endowment: w_t = phi - alpha*l_t + a_t. Productivity is identified as shifts in this curve across time periods. The slope alpha is not simply the land share under a CES production function but equals one minus the labor share divided by the elasticity of substitution between labor and land.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Malthusian population force.&lt;/strong&gt; The feedback mechanism by which higher real wages induce faster population growth, expanding labor supply and pushing wages back toward a steady state. Its speed is governed jointly by gamma (elasticity of population growth with respect to income) and alpha (slope of the labor demand curve); the half-life of wage/population dynamics after a shock equals log(0.5)/log(1 - alpha*gamma). In the paper&amp;rsquo;s estimates, this force was sufficiently weak (half-life approximately 150 years) that post-1800 productivity growth overwhelmed it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Plague shock (xi_1t).&lt;/strong&gt; An infrequent, large, exogenous negative population shock modeled as a draw from a beta distribution occurring with probability pi. Plagues are the primary source of identifying variation for the pre-1600 period: they generate movements along a stable labor demand curve and allow the slope alpha and the (lack of) productivity trend to be separately identified from labor demand shifts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural break in average productivity growth (mu).&lt;/strong&gt; The drift parameter in the random-walk model for the permanent component of productivity. The paper allows two breaks in mu, with break dates chosen to maximize the marginal likelihood (Bayes factor). The best-fitting breaks are at 1600 (zero to 4% per decade) and 1810 (4% to 18% per decade).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Permanent vs. transitory productivity component.&lt;/strong&gt; Productivity is decomposed into a permanent component a-tilde_t (random walk with drift, sigma_epsilon1) and a transitory component epsilon_2t (iid noise, sigma_epsilon2). The paper reports and interprets the permanent component as the meaningful measure of underlying technological change; transitory shocks are treated as measurement error and short-run fluctuations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Industrious Revolution.&lt;/strong&gt; The hypothesized long-run increase in days worked per worker in England, associated with de Vries (1994, 2008). The paper uses Humphries-Weisdorf (2019) estimates showing a sharp drop after the Black Death followed by a sustained rise from 1350 onward. A key robustness result is that the paper&amp;rsquo;s productivity estimates are insensitive to whether this Industrious Revolution is assumed to have occurred.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bayes factor (model selection).&lt;/strong&gt; The ratio of marginal likelihoods p(y|M_t)/p(y|M_t&amp;rsquo;) for two competing models, used here to select structural break dates for mu. A factor of 10 is treated as strong evidence. The bridge sampling method of Gronau, Singmann, and Wagenmakers (2020) is used to compute marginal likelihoods.&lt;/p&gt;</description></item></channel></rss>