<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Public-Finance | Macro Paper Warehouse</title><link>https://macropaperwarehouse.com/topics/public-finance/</link><atom:link href="https://macropaperwarehouse.com/topics/public-finance/index.xml" rel="self" type="application/rss+xml"/><description>Public-Finance</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><item><title>"Compensate the Losers?" Economic Policy and the Origins of U.S. Partisan Realignment</title><link>https://macropaperwarehouse.com/papers/compensate-the-losers-economic-policy-and-the-origins-of-u.s.-partisan-realignment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/compensate-the-losers-economic-policy-and-the-origins-of-u.s.-partisan-realignment/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Why have less-educated voters in the United States abandoned the Democratic Party over recent decades? The paper argues that the Democratic Party&amp;rsquo;s evolution on &lt;em&gt;economic policy&lt;/em&gt; — specifically its retreat from &amp;ldquo;predistribution&amp;rdquo; — is a central, previously understudied driver of partisan realignment by education.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conceptual Framework.&lt;/strong&gt; The authors distinguish between two categories of egalitarian economic policy: (1) &lt;em&gt;predistribution&lt;/em&gt; — policies that alter the pre-tax-and-transfer earnings distribution, including job guarantees, minimum wage increases, union support, and protectionist trade policies (following Hacker 2011); and (2) &lt;em&gt;redistribution&lt;/em&gt; — taxes and transfers. The paper&amp;rsquo;s central claim is that these two types of policy have sharply different educational gradients among voters, and that the Democratic Party moved away from predistribution beginning in the 1970s, triggering educational realignment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Methodology.&lt;/strong&gt; The authors harmonize over 1,000 surveys (N ≈ 2.2 million observations) spanning 1942–2020, drawn from Gallup, ANES, GSS, CCES, and historical survey archives housed at iPoll/Cornell. Education is translated into a common metric (adjusted years of schooling) using Census data, controlling for sex, race, year, and birth cohort to address the changing selectivity of educational categories over time. Congressional roll-call data come from the Comparative Agendas Project (CAP). Campaign finance data come from FEC filings, Congressional hearing records, and watchdog sources. DLC membership data are compiled from official Democratic Leadership Council records (available for 1985, 1986, 1991, 1993, and 1997 onward) and DLC-aligned Congressional caucus lists. House election returns are taken from King and Palmquist (1997) at the minor-civil-division-group (MCDG) level (~60 units per Congressional district), matched to 1980 Census demographic data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Voter preferences (demand side):&lt;/em&gt; The educational gradient for predistribution is large and negative: averaged across the four predistribution questions (job guarantee, minimum wage, union support, trade protection), each additional year of education reduces support by 0.044 standard deviations (p &amp;lt; 0.001). A college graduate relative to a high school graduate supports predistribution 0.176 standard deviations less — equivalent to roughly half the average Democrat-Republican gap in predistribution support (which is 0.34 standard deviations). This gradient has been stable since at least the 1940s. By contrast, the educational gradient for redistribution (higher taxes on the rich, views on own taxes, welfare spending) is close to zero (summary β = 0.004, not distinguishable from zero in the full sample). The difference between the two gradients is statistically significant (p &amp;lt; 0.001). These results replicate in white-only samples. Notably, the educational gradient on social issues — measured across nine questions on racial attitudes, gender roles, sexual norms — is positive (more education predicts more liberal positions) but has been largely &lt;em&gt;stable&lt;/em&gt; since the 1940s, not increasing, conditional on the long-run sample.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Party supply (supply side):&lt;/em&gt; Before 1976, predistribution topics accounted for roughly one-quarter of Democratic House roll-call votes when Democrats controlled the chamber. After 1976 (taking Jimmy Carter&amp;rsquo;s presidency as the start of the &amp;ldquo;New Democrat&amp;rdquo; era), this share falls by approximately nine to ten percentage points, while the redistribution share of votes holds steady. Between 1968 and 1980, the union share of total PAC donations to Democratic Congressional candidates falls from approximately 90 percent to 40 percent, coincident with 1970s campaign finance reforms that placed union and corporate PACs on equal legal footing and allowed corporations to exploit their naturally deeper pockets. Corporate PAC share of Democratic donations correspondingly rises from approximately 10 percent to 45 percent over the same period. In individual contributions to primary elections (data beginning in 1980), Democratic primaries rely on increasingly more-educated census tracts relative to Republican primaries; by 2018 Democratic primaries are financed from census tracts averaging 0.41 more years of education than Republican primaries (against a within-year standard deviation of 1.56 years).&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The New Democrat/DLC faction:&lt;/em&gt; The authors identify the anti-predistribution faction through official DLC membership records and aligned caucus lists. DLC membership as a share of Democratic House seats grows from near zero in the mid-1970s to approximately half by the early 2000s. Roll-call voting analysis (N = 3,428,405 vote-observations) shows DLC members are more conservative than other Democrats overall, and &lt;em&gt;especially&lt;/em&gt; so on predistribution: for a 10-percentage-point increase in the share of Republicans voting for a bill, the probability a DLC member votes in favor increases 36 percent more on predistribution bills than on other bills. DLC members show no differential conservatism on redistribution. They are also significantly more socially conservative — more likely than other Democrats to support the Defense of Marriage Act (by 16 pp), the Partial-Birth Abortion Ban (by 7 pp), and restrictive immigration bills (by 10 pp). DLC candidates receive significantly less from labor PACs and significantly more from corporate PACs, and draw their out-of-district individual donations from census tracts averaging more than 0.1 years more educated than non-DLC Democrats.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Voter reaction and the inflection point:&lt;/em&gt; Using the N ≈ 2.2 million partisan identification dataset, the authors estimate a structural break in the education-party identification gradient. From the 1940s through the mid-1970s, each additional year of education reduces the probability of identifying as a Democrat by approximately 3 percentage points. A Chow breakpoint test identifies 1976 as the inflection point. Since 1976, the gradient steadily rises; by 2000 it reaches zero; and today (as of the sample period end ~2020) each additional year of education &lt;em&gt;increases&lt;/em&gt; Democratic identification by approximately 3 percentage points — an almost exact reversal. The breakpoint for Republican identification occurs later, in 1992, consistent with the Democratic agenda changing first. A Gallup prosperity question (&amp;ldquo;which party will better keep the country prosperous?&amp;rdquo;) shows a parallel pattern: controlling for views on parties&amp;rsquo; economic performance explains approximately 44 percent of partisan realignment, interpreted as an upper bound on economic policy&amp;rsquo;s contribution.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Factional tests — hypothetical elections and actual results:&lt;/em&gt; In hypothetical general-election matchups from 1972–1992 Democratic primaries (in which most contests pitted a &amp;ldquo;New Democrat&amp;rdquo; against an &amp;ldquo;Old Democrat&amp;rdquo;), a voter with a college degree is roughly 3 percentage points &lt;em&gt;more&lt;/em&gt; likely to vote Democratic when the candidate is a New Democrat rather than an Old Democrat. In 1980s actual House elections using MCDG-level data, DLC candidates out-perform other Democrats in more educated neighborhoods by a magnitude large enough to erase approximately 90 percent of the general Democratic underperformance in highly educated areas. Combining these estimates, the party&amp;rsquo;s shift toward the DLC accounts for a lower bound of approximately 20 percent, and an upper bound (from the prosperity question) of approximately 50 percent, of educational realignment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; The analysis focuses on the United States, 1942–2015 (with some post-2015 discussion in the conclusion). The faction analysis focuses on the Democratic side; Republican faction changes are discussed but not the primary focus. The paper is explicit that between 20–50 percent of realignment is explained, leaving room for other factors, including social issues. The analysis ends mostly before 2016 to avoid complications from the closure of the DLC in 2011 and shifting post-2010 party dynamics.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-papers-central-conceptual-innovation-and-how-does-it-differ-from-prior-realignment-research"&gt;Q1. What is the paper&amp;rsquo;s central conceptual innovation, and how does it differ from prior realignment research?&lt;/h3&gt;
&lt;p&gt;The paper separates egalitarian economic policies into &amp;ldquo;predistribution&amp;rdquo; (pre-tax-and-transfer market interventions such as minimum wages, job guarantees, union support, and protectionism) and &amp;ldquo;redistribution&amp;rdquo; (taxes and transfers) and shows these two types have sharply different educational gradients. Prior work typically aggregated all economic policies into a single index, which the authors argue masks essential heterogeneity. By documenting that the educational gradient is large and negative for predistribution but close to zero for redistribution — a pattern stable since the 1940s — the paper reframes the &amp;ldquo;voting against economic interest&amp;rdquo; puzzle: less-educated voters leaving the Democratic Party may be responding rationally to changes in the supply of the type of economic policy they actually prefer.&lt;/p&gt;
&lt;h3 id="q2-how-large-and-stable-is-the-educational-gradient-on-predistribution-and-how-does-it-compare-to-social-issues"&gt;Q2. How large and stable is the educational gradient on predistribution, and how does it compare to social issues?&lt;/h3&gt;
&lt;p&gt;The average coefficient on adjusted years of schooling across the four predistribution questions is -0.044 (p &amp;lt; 0.001), stable over eight decades. A four-year difference in education (high school vs. college) shifts an individual&amp;rsquo;s support for predistribution by 0.176 standard deviations in the conservative direction — about half the average Democrat-Republican gap in predistribution support (0.34 standard deviations). For social issues, the summary gradient is positive (+0.028, p &amp;lt; 0.001 for the full sample), but this gradient has been largely &lt;em&gt;stable&lt;/em&gt; since the 1940s across nine social issue questions, not increasing over time. This stability undermines the interpretation that rising social liberalism among the educated is a new phenomenon driving realignment, at least through the supply of parties&amp;rsquo; social positions.&lt;/p&gt;
&lt;h3 id="q3-what-happened-to-predistribution-as-a-share-of-the-democratic-house-agenda-after-the-1970s"&gt;Q3. What happened to predistribution as a share of the Democratic House agenda after the 1970s?&lt;/h3&gt;
&lt;p&gt;Using the Comparative Agendas Project classification, predistribution topics (labor regulation, industrial policy, public works, trade) accounted for roughly one-quarter of all House roll-call votes during years Democrats controlled the Speakership before 1977. After 1977, this share falls by approximately 9–10 percentage points (a decline of nearly half from its pre-1977 share), and the decline is statistically significant (p &amp;lt; 0.001). The redistribution share of votes holds essentially constant. Party platform data from Hopkins et al. (2022) show a sharp decline in Democratic use of terms like &amp;ldquo;minimum wage,&amp;rdquo; &amp;ldquo;full employment,&amp;rdquo; and labor-relations language beginning in the 1970s and 1980s, while Republican platforms use these terms sparingly throughout.&lt;/p&gt;
&lt;h3 id="q4-how-did-1970s-campaign-finance-reforms-change-the-financial-composition-of-the-democratic-party"&gt;Q4. How did 1970s campaign finance reforms change the financial composition of the Democratic Party?&lt;/h3&gt;
&lt;p&gt;Before the early 1970s, unions enjoyed substantially more freedom than corporations under separate legal regimes governing PAC donations; mid-1970s reforms placed them on equal legal footing, enabling corporations to exploit their deeper pockets. The union share of total PAC donations to Democrats fell from approximately 90 percent in 1968 to approximately 40 percent by 1980, while the corporate share rose from approximately 10 percent to 45 percent. For Republicans, both series barely changed: unions had never donated substantially to the GOP, and the corporate share rose only modestly (from approximately 70 to 80 percent). The authors note the rapid decline cannot be attributed to falling union density in the economy, since both union and corporate PAC donations grew in absolute terms during this period; the relative shift was the result of the regulatory change.&lt;/p&gt;
&lt;h3 id="q5-who-are-the-new-democrats--dlc-and-when-did-they-emerge"&gt;Q5. Who are the &amp;ldquo;New Democrats&amp;rdquo; / DLC, and when did they emerge?&lt;/h3&gt;
&lt;p&gt;The DLC officially operated from 1985 to 2011, but members who would join it began entering Congress in large numbers in the 1970s (&amp;ldquo;Watergate Babies&amp;rdquo; of 1974, &amp;ldquo;Atari Democrats&amp;rdquo;). The DLC grew to approximately half of all Democratic House seats by the early 2000s. Members were drawn from suburban, affluent districts; their founder Al From explicitly criticized all four predistribution policies the paper studies (minimum wage, job guarantees, unions, and protectionism). The breakpoint test on DLC share in Congress identifies 1975 as the pivotal year — one year before the 1976 inflection point in partisan identification.&lt;/p&gt;
&lt;h3 id="q6-how-do-dlc-members-vote-differently-from-other-democrats-and-how-is-this-differential-conservatism-distributed-across-policy-types"&gt;Q6. How do DLC members vote differently from other Democrats, and how is this differential conservatism distributed across policy types?&lt;/h3&gt;
&lt;p&gt;In roll-call regressions (N = 3,428,405 observations, with roll-call fixed effects), a 10 pp increase in the Republican vote share for a bill increases the probability a DLC member votes in favor by 1.48 pp more than for other Democrats (baseline result for all bills). For predistribution-classified bills, this excess alignment with Republicans is 36 percent larger than for non-predistribution bills. Crucially, DLC members are no more conservative than other Democrats on redistribution-classified votes (the interaction with redistribution is near zero and insignificant). DLC members are also differentially more conservative on social issues, a result that proves useful in separating economic from social-issue explanations of realignment.&lt;/p&gt;
&lt;h3 id="q7-do-dlc-members-finance-differently-from-other-democrats"&gt;Q7. Do DLC members finance differently from other Democrats?&lt;/h3&gt;
&lt;p&gt;Yes. In primary elections, DLC candidates receive approximately 9.7 pp less of their PAC financing from labor unions and approximately 6.7 pp more from corporate PACs (with state fixed effects) relative to non-DLC Democrats. Out-of-district individual contributions to DLC primary candidates come from census tracts averaging more than 0.1 years more educated than those for non-DLC Democrats, while within-district contributions show no significant difference (0.060 years, insignificant). This pattern suggests educated out-of-district donors, rather than local constituency demands, drive DLC candidates&amp;rsquo; anti-predistribution orientation.&lt;/p&gt;
&lt;h3 id="q8-when-precisely-did-educational-realignment-in-democratic-party-identification-begin-and-what-does-the-inflection-point-analysis-show"&gt;Q8. When precisely did educational realignment in Democratic party identification begin, and what does the inflection-point analysis show?&lt;/h3&gt;
&lt;p&gt;Using N ≈ 2.2 million observations from 1,006 surveys, a Bai-Perron breakpoint test on the year-by-year education gradient in Democratic party identification identifies 1976 as the inflection point (with robustness to alternative specifications yielding breakpoints of 1978–1980 for white-only samples and unadjusted years of schooling). Before 1976, each additional year of education reduces the probability of Democratic identification by approximately 3 percentage points (a stable, significantly negative relationship since the 1940s). After 1976, the gradient steadily rises; it reaches zero around 2000 and today is approximately +3 percentage points per year of education — nearly an exact reversal of the baseline. The corresponding Republican inflection point occurs in 1992, about 16 years later, consistent with the Democratic Party&amp;rsquo;s agenda changing first.&lt;/p&gt;
&lt;h3 id="q9-how-do-hypothetical-presidential-matchup-surveys-test-the-dlc-mechanism"&gt;Q9. How do hypothetical presidential matchup surveys test the DLC mechanism?&lt;/h3&gt;
&lt;p&gt;The authors identify six Democratic primaries from 1972–1992 where a &amp;ldquo;New Democrat&amp;rdquo; and an &amp;ldquo;Old Democrat&amp;rdquo; were the top two contenders (e.g., Hart vs. Mondale in 1984, Clinton vs. Brown in 1992). Gallup and other surveys asked all respondents — regardless of party — whom they would vote for if either the New or the Old Democrat faced the eventual Republican nominee. A voter with a college BA is approximately 3 percentage points more likely to vote for the Democrat when the candidate is a New Democrat versus an Old Democrat (the &amp;ldquo;difference in differences&amp;rdquo; of hypothetical vote shares). This holds after controlling for state × election fixed effects and in five of the six election cycles studied (the 1976 exception is attributed to Mo Udall&amp;rsquo;s low name recognition, with 28 percent of respondents unfamiliar with him in a May 1976 poll). The result is attenuated but remains marginally significant when excluding non-white respondents, consistent with New Democrats&amp;rsquo; success with white voters due in part to their more conservative civil rights positioning.&lt;/p&gt;
&lt;h3 id="q10-what-do-actual-house-election-results-mcdg-level-data-show-about-dlc-electoral-performance-by-neighborhood-education"&gt;Q10. What do actual House election results (MCDG-level data) show about DLC electoral performance by neighborhood education?&lt;/h3&gt;
&lt;p&gt;Using 1980s House returns at the MCDG level (~60 neighborhoods per Congressional district), the authors regress Democratic vote share on neighborhood years of education interacted with a DLC candidate indicator, with Congressional district fixed effects. More-educated neighborhoods generally depress Democratic vote share (reflecting the still-negative overall educational gradient in the 1980s), but DLC candidates dramatically out-perform other Democrats in educated areas: the interaction coefficient is positive and significant, and its magnitude is large enough to erase approximately 90 percent of the general Democratic underperformance in highly educated neighborhoods. This result is robust to including District × Year fixed effects (so the identification comes from within-election, cross-neighborhood variation) and to adding controls for share white and share under age 35.&lt;/p&gt;
&lt;h3 id="q11-how-much-of-educational-realignment-can-the-papers-mechanism-account-for-and-how-is-this-calculated"&gt;Q11. How much of educational realignment can the paper&amp;rsquo;s mechanism account for, and how is this calculated?&lt;/h3&gt;
&lt;p&gt;Two bounding estimates are provided. Upper bound (~44–50%): controlling for a respondent&amp;rsquo;s view on which party is better for economic prosperity (from Gallup since 1950) explains approximately 44 percent of the change in the education-party identification gradient (specifically, the total difference in the unconditional gradient between the 1948–1967 baseline and 2001–2020 is 2.411 pp per year of schooling; after controlling for the prosperity question, the unexplained residual is 1.342 pp, leaving a share explained of 44.3 percent). Lower bound (~20%): the difference in the education gradient between matchups involving New versus Old Democrats in Table 4 (~0.75 pp) divided by the total realignment shift (~4 pp from pre-1976 to post-2008 for presidential voting) implies the faction shift accounts for at least approximately one-fifth of realignment. The authors interpret these as bounds because the prosperity question may partly capture party identification itself (upper bound concern), while the hypothetical matchup estimate misses the broader ideological shift not captured in a single election (lower bound).&lt;/p&gt;
&lt;h3 id="q12-can-social-issues-civil-rights-realignment-or-republican-changes-better-explain-the-1970s-inflection-point"&gt;Q12. Can social issues, Civil Rights realignment, or Republican changes better explain the 1970s inflection point?&lt;/h3&gt;
&lt;p&gt;Three alternative explanations are addressed. (1) &lt;em&gt;Civil Rights:&lt;/em&gt; Regional analysis shows that educated white Southerners &lt;em&gt;left&lt;/em&gt; the Democrats in the 1940s–1960s (not the 1970s), consistent with their realignment being driven by Democrats&amp;rsquo; liberal turn on civil rights rather than economic policy. After the 1960s, the South follows all other regions in the pace of educational realignment. (2) &lt;em&gt;Republican changes:&lt;/em&gt; The Republican party identification inflection point occurs in 1992, about 16 years after the Democratic inflection in 1976. Reagan elections in 1980 and 1984 do not appear to have differentially attracted less-educated voters (the &amp;ldquo;Reagan Democrats&amp;rdquo; were not differentially less educated). (3) &lt;em&gt;Social issues:&lt;/em&gt; The New Democrats were actually &lt;em&gt;more&lt;/em&gt; socially conservative than other Democrats (more likely to vote for DOMA, anti-abortion bills, restrictive immigration legislation), yet they disproportionately attracted educated voters. This internal inconsistency rules out a pure social-issues explanation for why educated voters preferred the DLC faction. (4) &lt;em&gt;Religion:&lt;/em&gt; Flexibly controlling for religious affiliation explains essentially none of partisan realignment (Appendix Figure A.24).&lt;/p&gt;
&lt;h3 id="q13-what-is-the-role-of-out-of-district-individual-donors-in-shifting-democratic-party-positions"&gt;Q13. What is the role of out-of-district individual donors in shifting Democratic Party positions?&lt;/h3&gt;
&lt;p&gt;Out-of-district primary donors are analytically important because they influence candidate supply without being able to vote in the election, isolating the &amp;ldquo;within-party&amp;rdquo; financial influence of educated supporters. By 1980, out-of-district primary donors to Democratic candidates already come from census tracts more educated than those for Republican candidates, even as local Democratic voters and within-district donors remain less educated than Republican counterparts. Democratic candidates also receive a substantially higher share of out-of-district contributions than Republican candidates — by almost 10 percentage points (Appendix Table A.7). Out-of-district donors thus represent a channel through which educated, anti-predistribution preferences are transmitted into the Democratic Party&amp;rsquo;s candidate supply before the electoral realignment is visible in vote totals.&lt;/p&gt;
&lt;h3 id="q14-are-predistribution-policies-becoming-less-popular-overall-which-might-independently-push-democrats-away-from-them"&gt;Q14. Are predistribution policies becoming less popular overall, which might independently push Democrats away from them?&lt;/h3&gt;
&lt;p&gt;The paper tests this alternative in Appendix Table A.9 and finds no evidence that predistribution has become less popular relative to redistribution over time. Predistribution appears on average more popular than redistribution across the sample period. If anything, support for predistribution has held steady or slightly risen relative to redistribution over time, conditional on the paper&amp;rsquo;s survey harmonization. The stability of the educational gradient (shown in Appendix Table A.10 to be unchanged even using educational rank within cohort rather than raw years of schooling) further suggests the negative education-predistribution relationship is a relative, not absolute, phenomenon — consistent with rising average education and stable preferences by education rank.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Predistribution:&lt;/strong&gt; Policies that aim to change the distribution of earnings or income &lt;em&gt;before&lt;/em&gt; taxes and transfers are applied. In this paper, this comprises government job guarantees, minimum wage increases, support for unions and collective bargaining, and protectionist trade policies. Distinguished from redistribution in that it operates on pre-tax market income rather than post-tax outcomes. The paper uses this term following Hacker (2011): &amp;ldquo;a focus on market reforms that encourage a more equal distribution of economic power and rewards even before government collects taxes or pays out benefits.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Redistribution:&lt;/strong&gt; Policies that change post-market income through the tax and transfer system, including higher taxes on the rich, views on own tax burden, prioritization of tax cuts, and transfers to the poor (welfare spending). In the paper&amp;rsquo;s usage, redistribution is analytically distinct from predistribution and has a near-zero educational gradient, in contrast to predistribution&amp;rsquo;s strongly negative gradient.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Educational Gradient:&lt;/strong&gt; The coefficient on adjusted years of schooling in a regression of an outcome variable (policy preference or partisan identification) on education, estimated separately by time period. The paper&amp;rsquo;s core finding is that the educational gradient for predistribution is stably negative (approximately -0.044 per year of schooling over the full sample), while the gradient for redistribution is close to zero, and the gradient for Democratic party identification shifts from approximately -0.03 to +0.03 per year of schooling between the 1940s and 2020.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;New Democrats / DLC (Democratic Leadership Council):&lt;/strong&gt; An explicitly anti-predistribution faction within the Democratic Party, identified through official DLC membership records and affiliated Congressional caucus lists. Founded formally in 1985 (operating through 2011), the DLC arose in part from the &amp;ldquo;Watergate Babies&amp;rdquo; cohort of 1974. DLC members were more conservative than other Democrats &lt;em&gt;especially&lt;/em&gt; on predistribution and social issues, relying differentially on corporate PACs and educated out-of-district donors. The paper treats DLC membership as a proxy for an anti-predistribution faction that gained bargaining power within the Democratic Party from the 1970s onward.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Adjusted Years of Schooling (AdjYearsEduc):&lt;/strong&gt; The paper&amp;rsquo;s harmonized education variable across more than 1,000 surveys spanning eight decades. Because raw educational categories change over time and represent different selectivity (e.g., in 1940 only one-quarter of adults had completed twelfth grade, versus nearly 90 percent today), the authors use Census microdata to predict years of schooling as a function of self-reported educational category, sex, race, year, and birth cohort in ten-year bins. This provides a common unit of measurement across surveys with incompatible category systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inflection Point (1976):&lt;/strong&gt; The structural break in the trend of the education-Democratic identification gradient, estimated using Bai-Perron (1998) methods on N ≈ 2.2 million observations. The data select 1976 as the year at which the previously stable negative gradient begins its upward trajectory. The corresponding Republican inflection point occurs in 1992. The paper argues that identification of this inflection point — not previously documented in the realignment literature — is made possible only by the large historical dataset assembled.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Minor Civil Division Group (MCDG):&lt;/strong&gt; The granular geographic unit used in the House election analysis for the 1980s, with approximately sixty MCDGs per Congressional district. Matched to 1980 Census demographic data to assign average years of education. Used to test whether DLC candidates out-perform other Democrats in more-educated neighborhoods, within the same Congressional district and election year, to address the concern that DLC candidates sort into more-educated districts.&lt;/p&gt;</description></item><item><title>(Not) Thinking About the Future: Financial Information and Maternal Labor Supply</title><link>https://macropaperwarehouse.com/papers/not-thinking-about-the-future-financial-information-and-maternal-labor-supply/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/not-thinking-about-the-future-financial-information-and-maternal-labor-supply/</guid><description>&lt;p&gt;This paper investigates whether information constraints — rather than fully forward-looking choices — contribute to mothers&amp;rsquo; reduced labor supply after childbirth, a key driver of gender inequality. The authors deploy two complementary methods in Switzerland: a representative descriptive survey of Swiss mothers aged 25–50, and a large-scale randomized controlled trial (RCT) among approximately 2,400 female public school teachers with children who work part-time.&lt;/p&gt;
&lt;p&gt;The descriptive survey first establishes that long-term financial factors are not top of mind for mothers making labor supply decisions: only about 11% of mothers spontaneously mention pensions or long-term career considerations when asked about their post-childbirth employment choices, compared to roughly half who mention child or own well-being. Beyond salience, the survey documents substantial misperceptions: 62% of women over-estimate pension receipt under part-time work by more than 10%, and a similar share believes wage growth under low part-time hours (40% FTE) is at least as high as under 80% employment. The authors label mothers with overly optimistic beliefs on both dimensions &amp;ldquo;cost-unaware&amp;rdquo;; 42% of the sample qualifies. Cost-unawareness is more prevalent among less-educated mothers and correlates with less financial interest and more gender-conservative attitudes.&lt;/p&gt;
&lt;p&gt;The RCT tests whether providing objective, individualized information shifts financial planning and labor supply. Teachers in treatment schools (two-thirds of all schools) were individually randomized into a treatment group viewing an informational video about the long-run earnings, pension, and life-event consequences of sustained part-time employment, plus access to a Future Calculator tool, or a placebo video on unrelated financial topics. The two-stage randomization (school-level first, then individual within treated schools) allows identification of both direct treatment effects and spillovers. Outcomes are measured in a Wave 1 post-video survey, a follow-up survey two months later, and linked administrative personnel records from the Department of Education one year post-intervention.&lt;/p&gt;
&lt;p&gt;Main findings: treated teachers are 31.26 percentage points (58% over the pure control mean) more likely to correctly rank the relative magnitude of long- versus short-term financial factors. Demand for financial planning tools rises by 0.39 standard deviations (SD) overall and by 0.31 SD among cost-unaware women specifically. In terms of stated labor supply plans, the treatment raises planned employment for the next academic year by 1.69 percentage points (ppt) in the full sample and by 4.95 ppt (9% over the pure control mean) among cost-unaware women. These plan effects persist two months later for cost-unaware women but fade for the full sample.&lt;/p&gt;
&lt;p&gt;Critically, stated plans translate into verified behavior: linked administrative data one year post-intervention show that cost-unaware teachers increase their contracted employment level by 3.87 ppt, or 7% over the pure control mean of 53.30% FTE. Cost-aware and overly pessimistic women do not reduce their labor supply upon learning they are better off than feared, an asymmetry consistent with agents responding more to perceived losses than gains. If the 3.87 ppt increase were sustained from age 40 onward, cost-unaware teachers would accumulate an additional 130,000 CHF in lifetime income and 40,000 CHF in pension wealth, shrinking the gender gap in lifetime income and pension receipt among teachers by approximately 18% each.&lt;/p&gt;
&lt;p&gt;The paper is scoped to Swiss female public school teachers — a population with linear pay scales, no part-time promotion penalty, and relatively low adjustment barriers — meaning the measured lifetime earnings and pension losses likely represent a lower bound relative to other occupations. Short-term RCT findings replicate among a sample of pregnant women in the general Swiss population, and the paper argues that similar labor supply adjustment magnitudes are feasible for a broader segment of part-time working mothers.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question and why does it matter?
A: The paper asks whether mothers&amp;rsquo; post-childbirth reduction in labor supply is partly driven by information constraints — specifically, whether mothers fail to account for the full long-term financial consequences of working reduced hours. This matters because if the child penalty partly reflects uninformed choices rather than deliberate tradeoffs, standard policy tools (parental leave, childcare subsidies) may underperform precisely because their long-term financial benefits are not internalized.&lt;/p&gt;
&lt;p&gt;Q: How prevalent is cost-unawareness among Swiss mothers?
A: 62% of mothers in the descriptive survey over-estimate pension receipt under part-time work by more than 10%, a similar share believes wage growth under low part-time (40% FTE) is at least as high as under 80% employment, and 42% are overly optimistic on both dimensions simultaneously. Cost-unawareness follows an education gradient: 77% of low-education women over-estimate pension receipt versus 51% of high-education women.&lt;/p&gt;
&lt;p&gt;Q: What share of mothers spontaneously considers long-term financial factors when deciding on their labor supply?
A: Only about 11% of mothers mention any long-term financial factor (pensions, financial independence, long-term career considerations) in open-ended responses; the share is similarly low across education groups (6% low, 12% mid, 13% high). About 50% mention child or own well-being; roughly 30% raise short-term financial factors such as current childcare costs.&lt;/p&gt;
&lt;p&gt;Q: What are the actual long-term financial stakes of the average female teacher&amp;rsquo;s part-time employment pattern in Switzerland?
A: Compared to full-time employment, the average female teacher&amp;rsquo;s employment trajectory produces a 35% reduction in potential lifetime earnings (approximately 3.34 million CHF versus 5.12 million CHF). Monthly pension receipt under the part-time scenario is 31% lower overall and 43% lower from the occupational second-pillar scheme specifically — a gap comparable to the average 47.5% gender pension gap observed in the second pillar in Switzerland in 2024.&lt;/p&gt;
&lt;p&gt;Q: How was the RCT designed and what populations were included?
A: The study recruited 2,359 part-time working mothers employed as public school teachers in a German-speaking Swiss canton. A two-stage randomization assigned two-thirds of schools to treatment schools (within which teachers were individually randomized 50/50 to treatment or spillover control) and one-third to pure control schools. This design allows estimation of direct treatment effects and spillover effects. The intervention was timed to precede December–January, the period when teachers communicate their preferred employment levels for the next school year.&lt;/p&gt;
&lt;p&gt;Q: What was the treatment intervention?
A: Treated teachers watched an informational video following a representative female teacher considering an employment-level increase, covering the impact of part-time work on lifetime earnings, monthly pension receipt, and financial exposure after adverse events such as divorce; it also benchmarked these magnitudes against childcare costs. Treated teachers additionally received individualized access to the Future Calculator, an online projection tool developed with a Swiss bank, calibrated to teachers&amp;rsquo; deterministic salary and pension schedules.&lt;/p&gt;
&lt;p&gt;Q: Did treated teachers understand and retain the treatment information?
A: Yes. Treated teachers were 31.26 ppt (58% over the pure control mean) more likely immediately after the intervention to correctly rank long- versus short-term financial factors in a vignette. Two months later, the treatment group remained significantly more likely to apply the information correctly (22.63 ppt higher), indicating the knowledge was not short-lived.&lt;/p&gt;
&lt;p&gt;Q: How did demand for financial planning tools respond to the treatment?
A: The treatment raised a financial information/tools index by 0.39 SD overall. For cost-unaware women specifically, demand for financial tools rose by 0.31 SD; cost-aware and pessimistic women showed no significant change. There was no significant average treatment effect on sign-up for an incentivized financial consultation.&lt;/p&gt;
&lt;p&gt;Q: How large were the labor supply plan effects in the survey, and did they persist?
A: For the full sample, treated teachers planned a 1.69 ppt higher employment level for the next school year immediately after the treatment, and 3.13 ppt higher in 10 years. For cost-unaware women, the short-run planned increase was 4.95 ppt (9% over the pure control mean of about 55%), and plans for 5 and 10 years into the future rose by approximately 4 ppt (6–7% over the mean). The short-run effects for cost-unaware women persisted to the two-month follow-up, while full-sample short-run effects faded.&lt;/p&gt;
&lt;p&gt;Q: What do the linked administrative data show about actual labor supply one year post-intervention?
A: Cost-unaware women in the treatment group increased their contracted employment level by 3.87 ppt relative to the pure control group (7% over the pure control mean of 53.30% FTE), closely matching the planned increase stated immediately after the treatment. Cost-aware women and the full sample showed no statistically significant shift in actual hours.&lt;/p&gt;
&lt;p&gt;Q: What asymmetry did the authors observe between cost-unaware and cost-aware women?
A: Cost-unaware (overly optimistic) women increased their labor supply upon learning the true financial costs; cost-aware and overly pessimistic women did not reduce their labor supply upon learning they were better off than expected. The authors interpret this as consistent with agents responding more to perceived losses (bad news for cost-unaware women) than to gains (good news for pessimistic women), and with cost-aware women already having incorporated the financial logic into their decisions even without precise estimates.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated lifetime impact of the observed labor supply adjustment?
A: If cost-unaware teachers maintain the 3.87 ppt employment increase from age 40 to retirement, they accumulate an additional 130,000 CHF in lifetime income and 40,000 CHF in pension wealth on average. This would reduce the gender gap in both lifetime income and pension receipt among teachers by approximately 18% each.&lt;/p&gt;
&lt;p&gt;Q: What emotional and social mechanisms did the paper document?
A: The treatment initially produced significantly negative emotional responses (−0.41 SD on an emotions index overall; −0.68 SD for cost-unaware women), consistent with cognitive dissonance from information conflicting with prior beliefs. Two months later, the treatment group reported feeling more in control and less stressed, and cost-unaware women returned to a neutral emotional baseline. Treated women were also 19.61 ppt more likely to have discussed the topic with anyone, with the largest effect on conversations with partners or family.&lt;/p&gt;
&lt;p&gt;Q: Did the treatment affect household-level labor supply — specifically, did partners reduce their hours?
A: No. The authors found no evidence that partners of cost-unaware women planned to work less in response to the treatment, and women did not plan to adjust future fertility. This suggests the observed hours increase by treated cost-unaware women was not offset by partner adjustments within the household.&lt;/p&gt;
&lt;p&gt;Q: Were there social spillover effects within schools?
A: Treated teachers were 11.59 ppt more likely to report having discussed the video with colleagues. Two months later, cost-unaware control teachers in treated schools (the spillover group) showed some evidence of absorbing the general treatment message and adjusting short-term labor supply plans upward, and a noisy increase in actual employment of roughly one-third the magnitude of the direct treatment effect, though these estimates were imprecise.&lt;/p&gt;
&lt;p&gt;Q: Why might cost-unaware women be uninformed in the first place?
A: In both the descriptive survey and the RCT sample, cost-unaware women lean more gender-conservative in their attitudes and report less interest in financial topics. The authors interpret this as suggesting a lack of information (rather than mere salience or forgetting) drives cost-unawareness, implying that passive information delivery through employers or pension funds could be effective.&lt;/p&gt;
&lt;p&gt;Q: What constraints to labor supply adjustment did the authors explore?
A: In a hypothetical scenario exercise, the scenario producing the largest desired employment increase for both treatment and control groups was if the partner were more engaged (roughly double the adjustment relative to a scenario of higher pay for additional hours). The treatment group adjusted their desired employment level by an additional 0.62–2.03 ppt relative to pure control across all scenarios except relaxing conservative gender norms.&lt;/p&gt;
&lt;p&gt;Q: How generalizable are the findings beyond the teacher sample?
A: The short-term RCT findings replicated among a sample of pregnant women in the general Swiss population. The authors also document that potential net gains from increasing labor supply — net of additional childcare costs — are large for the broader population of part-time working Swiss mothers, supporting feasibility of similar-magnitude adjustments outside teaching. The teaching context likely represents a lower bound for lifetime earnings and pension losses in other professions due to the absence of a part-time promotion penalty in teaching.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications?
A: The findings suggest that default exposure to individualized financial information about the long-term costs of part-time work — delivered by employers, pension funds, or the state — could improve decision quality and labor supply. More broadly, the results imply that policies designed to increase female labor supply (parental leave reforms, childcare subsidies) may underperform if mothers do not fully internalize the financial benefits of additional hours; ensuring that families solve the correct optimization problem is a precondition for unlocking the full potential of such policies.&lt;/p&gt;
&lt;p&gt;Child Penalty: The large and persistent reduction in women&amp;rsquo;s labor force participation and income following the birth of a first child, identified in the paper as the key driver of remaining gender inequality in the labor market in industrialized countries and a source of profound life-cycle financial consequences including reduced lifetime earnings and pension savings.&lt;/p&gt;
&lt;p&gt;Cost-Unaware: The authors&amp;rsquo; term for women who hold overly optimistic expectations about the financial consequences of part-time work — specifically, who over-estimate pension receipt under low part-time employment by more than 10% and who believe wage growth under low part-time is at least as high as under higher employment levels. In the descriptive survey 42% of mothers qualify on both dimensions.&lt;/p&gt;
&lt;p&gt;Future Calculator: An online individualized projection tool developed by the authors in cooperation with a Swiss bank, calibrated to teachers&amp;rsquo; deterministic salary and pension schedules, allowing users to estimate the long-term financial implications of different employment levels. Used both in the descriptive survey vignette and as part of the RCT treatment.&lt;/p&gt;
&lt;p&gt;Second Pillar (Occupational Pension Scheme, PP): Switzerland&amp;rsquo;s occupational pension scheme, the pillar most heavily affected by part-time work because contributions are directly proportional to earnings above a minimum annual earnings threshold. The paper documents an average gender pension gap of 47.5% in this pillar in 2024 and a 43% lower monthly pension receipt for the average female teacher&amp;rsquo;s part-time trajectory relative to full-time employment.&lt;/p&gt;
&lt;p&gt;Two-Stage Randomization: The experimental design used to separate direct treatment effects from spillover effects within schools. One-third of schools are assigned to a pure control group; in the remaining two-thirds, teachers are individually randomized into treatment or spillover control (untreated teachers in treated schools), enabling identification of both causal treatment impacts and social learning channels.&lt;/p&gt;
&lt;p&gt;Information Constraint: The paper&amp;rsquo;s central mechanism — mothers&amp;rsquo; failure to spontaneously account for the full long-term financial implications of reduced labor supply when making employment decisions, distinct from deliberate forward-looking tradeoffs. The authors document this both through the absence of long-term financial factors in open-ended decision narratives (only 11% of mothers mention them) and through systematic misperceptions of pension and wage outcomes.&lt;/p&gt;
&lt;p&gt;Cognitive Dissonance (as used in the paper): The authors use this term to describe the initial negative emotional response (−0.41 SD overall, −0.68 SD for cost-unaware women) when treated women learn that the true financial costs of part-time work are higher than they expected — information that conflicts with prior beliefs and prior choices, producing unpleasant emotions that subsequently reverse into lower stress levels two months later.&lt;/p&gt;</description></item><item><title>A Model of Multiple Hypothesis Testing</title><link>https://macropaperwarehouse.com/papers/a-model-of-multiple-hypothesis-testing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/a-model-of-multiple-hypothesis-testing/</guid><description>&lt;p&gt;This paper develops an economic framework for determining when and how much multiple hypothesis testing (MHT) adjustment is warranted in research settings. The research question is: under what conditions do MHT adjustments arise as an optimal solution to incentive misalignment between a researcher and a mechanism designer (social planner)?&lt;/p&gt;
&lt;p&gt;The model is a two-stage game. In the first stage, a benevolent social planner commits to a hypothesis testing protocol. In the second stage, a researcher decides whether to conduct a pre-specified experiment based on private costs and benefits. The planner&amp;rsquo;s utility function combines an ambiguity-averse (maximin) component—limiting harm from mistaken conclusions—with an expected-utility component capturing the generic benefits of research production. The framework focuses on multiplicity arising from testing multiple treatments or estimating effects within multiple subpopulations; multiple outcomes are treated as an economically distinct case covered in a companion paper.&lt;/p&gt;
&lt;p&gt;The main theoretical result is that separate t-tests are uniformly globally optimal under linearity of the researcher&amp;rsquo;s payoff and welfare functions and normality of test statistics. The optimal critical value takes the explicit form: t(J, Σ) = Φ⁻¹(1 − C(J, Σ) / (b · |J|)), where |J| is the number of hypotheses, C(J, Σ) is the experiment cost, and b is the researcher&amp;rsquo;s per-rejection benefit. This formula nests two limiting cases. When costs are fully fixed (invariant to |J|), the formula delivers a Bonferroni correction. When costs scale proportionally with the number of hypotheses, no MHT adjustment is warranted—because the researcher already faces sufficient deterrent from the incremental cost of each additional test.&lt;/p&gt;
&lt;p&gt;The key economic mechanism is as follows. In the worst states of the world (where all treatments are harmful relative to the status quo), a research study has only downside risk for society. The planner must keep the researcher&amp;rsquo;s expected payoff from false positives low enough that she chooses not to experiment. If critical values were invariant to |J|, for sufficiently many hypotheses the researcher&amp;rsquo;s expected payoff from false positives alone would exceed costs, inducing unwanted experimentation. Some upward adjustment to critical values (i.e., tighter thresholds) is therefore generically optimal. The same logic implies that critical values should also adjust for sample size, since larger samples raise costs.&lt;/p&gt;
&lt;p&gt;The framework is calibrated to two empirical applications. For FDA clinical trial approval, using Sertkaya et al. (2016) data on approximately 31,000 U.S. pharmaceutical trials (2004–2012), fixed costs constitute approximately 46% of average total trial cost. At a benchmark significance level of 5% and benchmark sample size, the optimal level is approximately 3.2% for two tests, 2.6% for three tests, and asymptotes to approximately 1.4% as |J| → ∞. Sidak&amp;rsquo;s correction yields 2.5% and 1.7% for two and three tests respectively, and tends to zero as |J| → ∞—more conservative than the model implies. Optimal adjustments must also be less conservative for larger samples to preserve researcher incentives to bear the correspondingly larger costs.&lt;/p&gt;
&lt;p&gt;For program evaluation in development economics, the paper uses a unique dataset of funding proposals submitted to J-PAL from 2009 to 2021. The estimated cost elasticity with respect to the number of treatment arms ranges from 0.13 to 0.22 (p &amp;lt; 0.05), indicating costs rise significantly but far less than proportionally. The implied optimal significance levels are slightly less conservative than Bonferroni/Sidak corrections but more conservative than unadjusted testing.&lt;/p&gt;
&lt;p&gt;Scope conditions: the framework assumes pre-specified experiments (no p-hacking), linear payoffs, normally distributed statistics, and a researcher whose preferences are common knowledge. The analysis focuses on multiple treatments and subpopulations, not multiple outcomes. Results extend to imperfectly informed researchers and heterogeneous variances.&lt;/p&gt;
&lt;p&gt;Q: What is the core mechanism by which MHT adjustments arise as optimal in this framework?
A: The planner must deter experimentation in the worst-case states—those where all treatments are harmful. If the testing protocol did not adjust for the number of hypotheses, a researcher testing sufficiently many hypotheses could earn enough expected payoff from false positives alone to justify experimentation, even when all treatments are truly harmful. Tighter critical values (higher thresholds) reduce the probability of false positives and thus cap the researcher&amp;rsquo;s expected payoff in the null space, deterring unwanted experimentation. This is the maximin optimality condition: the researcher&amp;rsquo;s expected payoff must be non-positive over the null space.&lt;/p&gt;
&lt;p&gt;Q: What are the two limiting cases of the optimal critical value formula, and what do they correspond to?
A: The optimal level of the separate t-tests is α(J, Σ) = C(J, Σ) / (b · |J|). When C(J, Σ) = ᾱ (costs are fixed, invariant to the number of hypotheses), this reduces to ᾱ/|J|, the Bonferroni correction. When C(J, Σ) = ᾱ · |J| (costs scale proportionally with the number of hypotheses), the optimal level equals ᾱ regardless of |J|—no MHT adjustment is warranted. The intuition for the second case is that proportional costs already deter excess testing; the researcher has no undue incentive to test many hypotheses because each additional test costs the same incremental amount.&lt;/p&gt;
&lt;p&gt;Q: Why do optimal critical values also depend on sample size, and what is the policy implication?
A: Since research costs C(J, Σ) increase with sample size (Σ captures design features including sample size), the optimal test level α(J, Σ) = C(J, Σ)/(b·|J|) rises with sample size. Equivalently, larger studies warrant less conservative significance thresholds. The policy implication is that a single uniform correction (e.g., Bonferroni at the 5% level) applied without regard to sample size is suboptimal: it is too conservative for large studies, which would over-deter valuable high-powered research.&lt;/p&gt;
&lt;p&gt;Q: What are the two optimality properties required of protocols in the paper&amp;rsquo;s main characterization?
A: The paper shows (Proposition 3.1) that a protocol is uniformly globally optimal—optimal for all values of the welfare weight λ and prior π—if and only if it is both maximin optimal and unbiased. Maximin optimality (Proposition 3.2) requires two conditions: the researcher&amp;rsquo;s expected payoff must be non-positive over the null space (deterring experimentation when all treatments are harmful), and expected welfare must be non-negative when some treatments are beneficial. Unbiasedness requires that the researcher&amp;rsquo;s maximum power strictly exceeds the test size, ensuring that experimentation is motivated when treatments are genuinely beneficial.&lt;/p&gt;
&lt;p&gt;Q: How does the paper rationalize conventional hypothesis testing asymmetry (type I vs. type II error weighting) without extreme restrictions?
A: In Tetenov (2012), justifying 5%-level testing with minimax regret in a single-agent model requires the decision-maker to place 102 times more weight on type I than type II regret—an extreme restriction. In this paper, the asymmetry arises naturally from the planner&amp;rsquo;s desire to prevent harmful treatment implementation: the planner is willing to forgo some power (probability of detecting beneficial treatments) to ensure that harmful treatments are not implemented. The researcher&amp;rsquo;s private incentives and the planner&amp;rsquo;s objective diverge in a way that makes tight size control endogenously optimal.&lt;/p&gt;
&lt;p&gt;Q: What does the FDA empirical calibration imply quantitatively about optimal versus standard adjustments?
A: Using Sertkaya et al. (2016) data showing that fixed costs are 46% of average total trial cost for U.S. pharmaceutical trials, and using Pocock et al. (2002) to set J̄ = 3 (average number of subgroups), the paper calculates that at a benchmark level of ᾱ = 0.05: the optimal level is approximately 3.2% for two tests, 2.6% for three tests, and asymptotes to approximately 1.4% as |J| → ∞. By contrast, Sidak&amp;rsquo;s correction yields 2.5%, 1.7%, and zero, respectively. Both the unadjusted 5% and the Sidak/Bonferroni levels are therefore suboptimal—the unadjusted level is too permissive while standard FWER corrections are too conservative.&lt;/p&gt;
&lt;p&gt;Q: What do the J-PAL data reveal about optimal MHT adjustment in program evaluation?
A: Using the universe of J-PAL funding proposals from 2009 to 2021, the paper estimates the cost elasticity with respect to the number of treatment arms to be 0.13–0.22, which is statistically significant (p &amp;lt; 0.05) but far below 1 (the proportional case). This means costs rise with arms but much less than proportionally. As a result, optimal significance levels for program evaluation studies are slightly less conservative than Sidak/Bonferroni corrections (e.g., approximately 3.8–4.5% versus 2.5% at a two-arm study with ᾱ = 5%) but more conservative than unadjusted testing. The testing thresholds also vary moderately with sample size, with larger samples implying less conservative procedures.&lt;/p&gt;
&lt;p&gt;Q: When are cross-study MHT adjustments warranted according to the framework?
A: Cross-study MHT adjustments are warranted only when there are cost complementarities across those studies. If studies are conducted independently with separate cost structures, each study&amp;rsquo;s costs do not depend on the number of hypotheses tested in other studies, so no cross-study adjustment is optimal. This provides a principled resolution to the disputed question of whether researchers should correct for tests performed in other papers.&lt;/p&gt;
&lt;p&gt;Q: When is FWER control (e.g., Bonferroni or Sidak) the appropriate form of MHT adjustment?
A: Appendix B.2 shows that FWER control is appropriate when the researcher&amp;rsquo;s payoff is nonlinear—specifically when the researcher requires at least one positive finding to receive any benefit (e.g., to publish). In the baseline linear payoff model, average size control (Bonferroni) is the correct adjustment only when all costs are fixed. The broader insight is that the form of compound error control—whether average error rate or FWER—is itself determined by economic fundamentals rather than being a statistical choice made in advance.&lt;/p&gt;
&lt;p&gt;Q: How does the paper extend to cases of heterogeneous variances across hypotheses?
A: Proposition 5.2 shows that under heterogeneous variances, the optimal protocol uses separate t-tests based on sample-equalizing allocations—dividing the sample equally across treatment arms—with critical values t*(J, n(J)) = Φ⁻¹(1 − C(J, n(J))/(b·|J|)), where n(J) is the total sample size. This protocol remains maximin optimal and unbiased, preserving the main qualitative results.&lt;/p&gt;
&lt;p&gt;Q: What does the paper contribute relative to Tetenov (2016) on single-hypothesis testing?
A: Tetenov (2016) showed that in the single-hypothesis case, separate t-tests are maximin optimal and uniformly most powerful (UMP) unbiased. This paper extends that result to multiple hypotheses, but two major complications arise: first, maximin optimality in the multi-hypothesis case requires verifying that welfare is non-negative even when treatment effects have opposite signs, which requires a non-trivial argument absent in the single-hypothesis case; second, no protocol is UMP unbiased in the multi-hypothesis case, so the paper develops a weaker notion of unbiasedness (power exceeding size) that is sufficient to motivate experimentation.&lt;/p&gt;
&lt;p&gt;Q: Why do multiple outcomes require different procedures than multiple treatments or subpopulations?
A: Multiple outcomes and multiple treatments are economically distinct types of multiplicity. For multiple outcomes that are noisy proxies for a common underlying quantity, the optimal rule tests an index formed using statistical weights (as in Anderson, 2008). When outcomes capture distinct components of the planner&amp;rsquo;s utility, economic weights are appropriate. In contrast, multiple treatments or subpopulations lead to separate t-tests with cost-adjusted critical values. Conflating these two forms of multiplicity leads to incorrect inferences about what procedures are appropriate.&lt;/p&gt;
&lt;p&gt;Maximin optimality: A hypothesis testing protocol is maximin optimal if it maximizes the planner&amp;rsquo;s worst-case welfare across all parameter values, equivalent to two conditions: deterring researcher experimentation over the null space (where all treatments are harmful), and ensuring non-negative expected welfare when some treatments are beneficial.&lt;/p&gt;
&lt;p&gt;Unbiasedness (in the paper&amp;rsquo;s sense): A protocol is unbiased if the researcher&amp;rsquo;s maximum achievable power strictly exceeds the test size, ensuring that experimentation is motivated when treatments are genuinely beneficial. This is a weaker condition than UMP unbiasedness, which does not exist in the multi-hypothesis case.&lt;/p&gt;
&lt;p&gt;Uniform global optimality: A protocol is uniformly globally optimal if it maximizes the planner&amp;rsquo;s objective for all values of the welfare weight λ ≥ 0 and all priors π over the parameter space, making it robust to uncertainty about the relative importance of deterrence versus research motivation.&lt;/p&gt;
&lt;p&gt;MHT correction factor: Defined as C(J, Σ) / (C̄ · |J|), this factor captures how the cost per test varies as the number of hypotheses grows. It equals 1/|J| (Bonferroni) when all costs are fixed, and equals 1 (no correction) when costs are proportional to the number of tests; the empirically appropriate correction lies strictly between these extremes.&lt;/p&gt;
&lt;p&gt;Cost function C(J, Σ): The private cost borne by the researcher for conducting the experiment, which depends on both the set of treatments J and the experimental design Σ (including sample size). The degree of optimal MHT adjustment is a direct function of how this cost varies with the number of hypotheses tested.&lt;/p&gt;
&lt;p&gt;Global null space Θ₀(J): The set of parameter vectors θ for which the welfare effect of implementing any combination of treatments is strictly negative—i.e., the status quo of no treatment dominates all interventions. Maximin optimality requires deterring researcher experimentation over this set.&lt;/p&gt;
&lt;p&gt;Cost complementarities across studies: Cost structures in which conducting multiple studies together is cheaper than conducting them separately. Cross-study MHT adjustments are warranted if and only if such complementarities exist; absent complementarities, each study&amp;rsquo;s optimal threshold is set independently of others.&lt;/p&gt;</description></item><item><title>A Welfare Analysis of Policies Impacting Climate Change</title><link>https://macropaperwarehouse.com/papers/a-welfare-analysis-of-policies-impacting-climate-change/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/a-welfare-analysis-of-policies-impacting-climate-change/</guid><description>&lt;p&gt;This paper extends and applies the marginal value of public funds (MVPF) framework to evaluate the welfare consequences of 96 climate-related tax and spending policies in the United States. The MVPF is a benefit-cost ratio in which the numerator captures all benefits to individuals (measured by their willingness to pay) and the denominator captures net government costs; policies with higher MVPFs are better spending policies, while those with lower MVPFs are more efficient revenue-raising instruments.&lt;/p&gt;
&lt;p&gt;The sample covers policies rigorously evaluated using quasi-experimental or experimental methods drawn from 18 major economics journals between January 1999 and December 2023. Policies fall into three primary categories: subsidies (wind production tax credits, residential solar, electric vehicles, hybrid vehicles, vehicle buybacks, appliance rebates, and weatherization), nudges and marketing, and revenue raisers (gasoline taxes, other fuel taxes, cap-and-trade). A selected set of international aid policies is also analyzed. The analysis applies a harmonized method for translating behavioral changes into emissions changes — using the EPA&amp;rsquo;s AVERT model for electricity-sector emissions — and a consistent set of externality valuations, including an EPA 2023 social cost of carbon (SCC) of $193 per ton of CO2 in 2020 (rising over time), with robustness checks at $76, $337, and $1,367.&lt;/p&gt;
&lt;p&gt;The primary methodological contribution is a new sufficient statistics approach to quantifying learning-by-doing (LBD) externalities. When marginal cost of production is an isoelastic function of cumulative production and demand is an isoelastic function of price, the time path of production satisfies a second-order ordinary differential equation whose solution yields society&amp;rsquo;s willingness to pay for LBD spillovers. LBD generates two types of externalities: a price externality (lower future consumer prices) and an environmental externality (increased future take-up of clean goods). The approach requires four inputs: price elasticity of demand, elasticity of marginal cost with respect to cumulative production, cumulative production at the time of the subsidy, and product cost at the time of the subsidy.&lt;/p&gt;
&lt;p&gt;The three main empirical findings are as follows. First, subsidies for production that directly displaces dirty electricity generation have the highest MVPFs. Wind production tax credits have an MVPF of 3.85 without LBD, rising to 5.87 with LBD. Residential solar subsidies have an MVPF of 1.45 without LBD, rising to 3.86 with LBD. EV subsidies have an MVPF of approximately 1.4 with LBD and approximately 1 without it. Consumer subsidies for appliances, weatherization, vehicle retirement, and hybrid vehicles have MVPFs around 1. Second, conservation nudges targeting electricity consumption can deliver MVPFs exceeding 5 in regions with relatively dirty electric grids, but fall below 1 in cleaner-grid regions such as California and the Northeast — and their effectiveness is expected to decline as grids decarbonize. Third, fuel taxes (gasoline, diesel, jet fuel) and cap-and-trade permit reductions are efficient revenue raisers, with nearly all having MVPFs below 1 and most below 0.7, reflecting the Pigouvian logic that current tax rates fall below the associated environmental externalities. Cap-and-trade permit reductions can produce MVPFs below zero, meaning revenue is raised while providing net positive welfare to individuals.&lt;/p&gt;
&lt;p&gt;The paper also constructs three cost-per-ton metrics — resource cost per ton, government cost per ton, and social cost per ton — and shows they can yield substantively different and sometimes opposite rankings relative to each other and to the MVPF. For example, EV subsidies carry a government cost per ton of $1,356 (among the highest in the sample) yet an MVPF above most consumer subsidies, because that metric omits non-CO2 benefits including LBD effects. The scope of the analysis is US historical policy, with the MVPF comparison most informative when social welfare weights across beneficiary groups are treated as roughly equal.&lt;/p&gt;
&lt;p&gt;Q: What is the MVPF framework and how does it differ from cost-per-ton analysis?
A: The MVPF equals benefits to individuals (sum of willingness to pay) divided by net cost to the government. It is designed for a decision-maker maximizing social welfare subject to a budget constraint, whereas cost-per-ton metrics serve a decision-maker minimizing cost subject to a fixed CO2 reduction target. A higher MVPF means more welfare gain per dollar spent; a lower MVPF means less welfare cost per dollar of revenue raised.&lt;/p&gt;
&lt;p&gt;Q: What are the three cost-per-ton definitions the paper distinguishes, and why do they differ?
A: Resource cost per ton measures the economic resources consumed per ton of CO2 abated, independent of subsidy incidence; government cost per ton measures net government outlays per ton, omitting all non-CO2 benefits; social cost per ton subtracts non-CO2 benefits from government costs. For appliance rebates, these three values are -$2, $474, and an intermediate figure — a range that reflects whether inframarginal transfers and non-CO2 co-benefits are counted.&lt;/p&gt;
&lt;p&gt;Q: What is the new methodological contribution regarding learning by doing?
A: The paper derives a sufficient statistics result showing that when marginal production cost is an isoelastic function of cumulative production and demand is isoelastic in price, the time path of production follows a second-order ordinary differential equation. Solving this equation yields society&amp;rsquo;s willingness to pay for LBD spillovers from four observable parameters: demand price elasticity, the LBD elasticity of marginal cost with respect to cumulative production, cumulative production at the subsidy date, and unit cost at that date. This allows LBD benefits to be incorporated into both MVPF and cost-per-ton calculations without requiring a fully calibrated dynamic model.&lt;/p&gt;
&lt;p&gt;Q: What LBD elasticities does the paper use, and where do they come from?
A: Drawing on Way et al. (2022), a 1% increase in cumulative solar production is associated with a 0.319% price reduction; for wind the elasticity is 0.194%, and for EV batteries it is 0.421%. These are treated as the isoelastic parameter in the sufficient statistics formula.&lt;/p&gt;
&lt;p&gt;Q: How does LBD affect the MVPF estimates for wind, solar, and EVs specifically?
A: For wind production tax credits, the MVPF rises from 3.85 to 5.87 when LBD is included. For residential solar, it rises from 1.45 to 3.86. For EV subsidies, the MVPF rises from approximately 1 to approximately 1.4. Without LBD, EV subsidies are in line with other consumer subsidies; LBD is the primary reason EVs outperform that group.&lt;/p&gt;
&lt;p&gt;Q: What is the baseline social cost of carbon used, and how sensitive are results to alternative values?
A: The baseline SCC is $193 per ton of CO2 in 2020, following EPA 2023 guidance at a 2% discount rate. Robustness checks use $76, $337, and $1,367. Higher SCC values raise the MVPF of all subsidies in the sample, but the relative ordering — with wind PTCs above all other consumer subsidies — remains consistent across the full range.&lt;/p&gt;
&lt;p&gt;Q: How are EV subsidies evaluated, and what accounts for their MVPF exceeding other consumer subsidies?
A: The analysis uses the California EFMP program studied by Muehlegger and Rapson (2022), which finds a price elasticity of demand of -2.1 and 85% pass-through to consumers (15% captured by dealers). A $1 subsidy generates $0.85 in consumer WTP, $0.15 in dealer WTP, $0.17 in CO2 co-benefits, $0.05 in local pollution and accident co-benefits, offset by $0.10 in damages from increased electricity generation. Most benefits are non-environmental (inframarginal transfers and LBD effects on future vehicle prices), which is why the government cost per ton of $1,356 appears high while the MVPF is approximately 1.4.&lt;/p&gt;
&lt;p&gt;Q: What drives the high MVPFs for nudges in dirty-grid regions, and what is the implication for the future?
A: Conservation nudges in dirty-grid areas have high MVPFs (exceeding 5) because each kilowatt-hour of reduced consumption displaces generation from high-emission sources, amplifying the environmental benefit per dollar of program cost. In cleaner-grid regions like California and the Northeast, the same nudge displaces lower-emission generation, pushing the MVPF below 1. As grids decarbonize nationwide, the paper notes that nudge MVPFs will decline over time.&lt;/p&gt;
&lt;p&gt;Q: How do cap-and-trade permit reductions compare to fuel taxes as revenue-raising instruments?
A: Nearly all fuel taxes (gasoline, diesel, jet fuel) have MVPFs below 1, with most below 0.7, meaning they impose a welfare cost of only $0.70 per dollar of revenue raised. Cap-and-trade permit reductions can have MVPFs below zero, meaning they can raise revenue while simultaneously providing net positive welfare gains to individuals because environmental benefits from reduced emissions outweigh the permit costs borne by emitters.&lt;/p&gt;
&lt;p&gt;Q: What do the international subsidy findings suggest, and what are their limitations?
A: Subsidies for efficient charcoal cookstoves in Kenya (Berkouwer and Dean 2022) generate US-specific gains from CO2 reductions that are 37 times the net cost of the subsidy; including global benefits raises the MVPF to 323. However, the paper flags substantial uncertainty: estimated policy impacts vary widely within similar international categories, and the US-specific MVPF is highly sensitive to assumptions about the incidence of the social cost of carbon on US residents and US government tax revenue.&lt;/p&gt;
&lt;p&gt;Q: Why does the social cost per ton metric give opposite rankings within wind, solar, and EVs relative to the MVPF?
A: EVs have a social cost per ton of -$415 versus -$32 for wind PTCs, making EVs appear superior on that metric — the reverse of the MVPF ordering. The paper explains that when SCPT values are negative (policies that abate CO2 while also yielding positive non-CO2 net benefits), the metric loses its Lagrange multiplier interpretation: increased non-CO2 benefits make SCPT more negative while increased abatement makes it less negative, preventing meaningful cross-policy comparisons.&lt;/p&gt;
&lt;p&gt;Q: What is the overall policy ranking implied by the MVPF analysis?
A: From highest to lowest MVPF: international clean energy subsidies &amp;gt; wind production tax credits &amp;gt; residential solar subsidies &amp;gt; energy conservation nudges (dirty grids) &amp;gt; EV subsidies &amp;gt; consumer appliance and weatherization subsidies &amp;gt; hybrid vehicle subsidies &amp;gt; vehicle buyback rebates &amp;gt; energy conservation nudges (clean grids) &amp;gt; revenue raisers (gas taxes, fuel taxes, cap-and-trade). The paper notes that shifting $1 of government revenue from gas taxes (MVPF ~0.67) to wind PTCs (MVPF ~5.87) generates $5.20 in net welfare benefits to individuals, assuming equal social welfare weights across groups.&lt;/p&gt;
&lt;p&gt;Marginal Value of Public Funds (MVPF): A benefit-cost ratio equal to the sum of individuals&amp;rsquo; willingness to pay for a policy divided by its net cost to the government. Policies with higher MVPFs deliver greater welfare gains per dollar spent; those with lower MVPFs impose lower welfare costs per dollar of revenue raised. Used to compare spending and revenue-raising policies on a common welfare-maximizing basis.&lt;/p&gt;
&lt;p&gt;Learning-by-Doing (LBD) Externality: The spillover by which current production of a technology lowers its future marginal cost, generating future consumer surplus (price externality) and additional future uptake with associated environmental benefits (environmental externality). Treated in this paper as an uninternalized external benefit of subsidizing current production.&lt;/p&gt;
&lt;p&gt;Sufficient Statistics Approach to LBD: The paper&amp;rsquo;s methodological contribution — showing that when marginal cost is an isoelastic function of cumulative production and demand is isoelastic in price, the LBD welfare benefit can be computed from four observables: the demand price elasticity, the LBD cost elasticity, cumulative production at subsidy date, and unit cost at subsidy date, without requiring a fully specified dynamic model.&lt;/p&gt;
&lt;p&gt;Resource Cost per Ton (RCPT): Economic resources consumed to produce and use a product, divided by tons of CO2 abated. Appropriate for private firms minimizing abatement cost; independent of subsidy take-up rates and inframarginal transfers.&lt;/p&gt;
&lt;p&gt;Government Cost per Ton (GCPT): Net government outlay per ton of CO2 abated. The correct metric for a government focused exclusively on CO2 reduction at minimum fiscal cost; omits all non-CO2 welfare impacts, including co-benefits and LBD effects.&lt;/p&gt;
&lt;p&gt;Social Cost per Ton (SCPT): Government cost net of all non-CO2 benefits, per ton of CO2 abated. Intended to capture the social cost of abatement, but loses its Lagrange multiplier interpretation when values are negative, preventing valid cross-policy comparisons in that region.&lt;/p&gt;
&lt;p&gt;Social Cost of Carbon (SCC): The monetized damage from one additional ton of CO2 emissions. Baseline value of $193 per ton in 2020 from EPA 2023 at a 2% discount rate, rising over time. A key parameter driving MVPF levels across all policy categories; robustness checked at $76, $337, and $1,367.&lt;/p&gt;
&lt;p&gt;Pigouvian Efficiency of Environmental Taxes: The paper quantifies that fuel taxes have MVPFs below 0.7 because current tax rates fall below the associated Pigouvian optimum — i.e., taxing polluting goods raises revenue while reducing a pre-existing negative externality, so the welfare cost of the revenue is less than one dollar per dollar raised.&lt;/p&gt;</description></item><item><title>Additionality and Asymmetric Information in Environmental Markets: Evidence from Conservation Auctions</title><link>https://macropaperwarehouse.com/papers/additionality-and-asymmetric-information-in-environmental-markets-evidence-from-conservation-auctions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/additionality-and-asymmetric-information-in-environmental-markets-evidence-from-conservation-auctions/</guid><description>&lt;p&gt;This paper investigates the problem of additionality — the likelihood that a conservation action is marginal to (i.e., caused by) an incentive — in the United States Department of Agriculture&amp;rsquo;s Conservation Reserve Program (CRP), one of the largest and most mature Payments for Ecosystem Services (PES) mechanisms in the world. The CRP pays landowners $1.6–$1.8 billion per year under 10-year contracts to retire cropland and plant grass mixes, trees, or wildlife habitats, using a discriminatory scoring auction in which landowners submit bids on a menu of heterogeneous contracts ranked by a scoring rule.&lt;/p&gt;
&lt;p&gt;The central argument is that additionality represents a form of asymmetric information. Landowners possess private knowledge about their counterfactual land use (whether they would have conserved anyway), while the auction screens only on their private cost of accepting the contract. Because lower-cost landowners are lower-cost partly because they expect to conserve regardless of the CRP, cost and additionality are positively correlated — generating adverse selection: the least costly participants to purchase are the least socially valuable. The status quo scoring rule implicitly assumes all landowners are fully additional (tau = 1), an assumption the paper tests and rejects.&lt;/p&gt;
&lt;p&gt;The authors construct a dataset linking confidential administrative CRP bid data across seven auctions from 2009 to 2021 to satellite-derived land use classifications from the Cropland Data Layer (30m resolution) and USDA administrative land use reports. They exploit a regression discontinuity (RD) in contract awards around the winning score threshold to estimate the causal effect of CRP contracts on land use at the margin. The first-stage is close to one. The key finding is that CRP contracts reduce cropping by approximately eight percentage points at the margin, but the 100%-additional benchmark predicts a reduction of roughly 33 percentage points (matching the share of land covered by a contract at the margin). Therefore, only approximately one quarter (22–29%) of marginal auction winners are additional — meaning three-quarters would have conserved without the CRP contract.&lt;/p&gt;
&lt;p&gt;To test for adverse selection, the authors use the 82% of rejected bidders in the 2016 auction (the most restrictive) for whom counterfactual land use is observed, constructing a landowner-specific additionality measure. They document a systematic positive correlation between bid rental rates (reflecting higher costs) and additionality, which persists conditional on rich observable characteristics including prior land use interacted with soil productivity. Contract choice further reveals additionality: tree-related contract bidders exhibit substantially lower additionality than base grassland contract bidders.&lt;/p&gt;
&lt;p&gt;To quantify welfare implications, the authors develop and estimate a joint structural model of bidding and additionality. Costs are inferred via revealed preferences in optimal bidding (following the empirical auctions literature), and additionality is estimated as a conditional expectation function of observable characteristics and unobserved costs, matched to observed land use among rejected bidders via Method of Simulated Moments. Social benefits are taken from the CRP literature and USDA revealed preferences.&lt;/p&gt;
&lt;p&gt;Key welfare findings: (1) Despite widespread non-additionality and adverse selection, a hypothetical uniform-price market for the base conservation contract generates social welfare gains of $14.37 per acre-year at the socially-optimal price. Setting price equal to the full social benefit B — ignoring counterfactual land use — causes welfare losses of $12.68 per acre-year, nearly eliminating the gains. (2) The status quo auction generates social welfare gains of approximately $120 million per auction relative to no market, but implements only 12% of the gains achievable under the efficient allocation. (3) Simple modifications to the scoring rule that incorporate expected additionality — via uniform adjustments and market-size reductions — close 37% of the gap between the status quo and the efficient allocation, increasing social welfare by over $300 million per auction. Nearly all gains arise from incorporating additionality into the scoring rule. These modifications are described as implementable by the USDA in practice.&lt;/p&gt;
&lt;p&gt;Q: What is additionality, and why does it matter for conservation markets?
A: Additionality is defined as the expected impact of contracting on a landowner&amp;rsquo;s conservation action — i.e., the probability that a landowner would not have conserved absent the incentive. Social surplus depends on both a landowner&amp;rsquo;s cost of accepting a contract and her additionality, but market mechanisms screen only on cost. When the lowest-cost participants are the least additional, standard procurement mechanisms fail to implement the efficient allocation, undermining the environmental and fiscal effectiveness of conservation programs.&lt;/p&gt;
&lt;p&gt;Q: What is the rate of additionality at the margin of CRP contract awards?
A: Approximately one quarter (22–29% depending on specification) of marginal auction winners are additional. The RD design shows contracts reduce cropping by about eight percentage points at the margin, compared to the 100%-additional benchmark of approximately 33 percentage points (the share of land covered by the contract at the margin). This implies three-quarters of marginal winners would have conserved without a CRP contract.&lt;/p&gt;
&lt;p&gt;Q: What is the empirical evidence for adverse selection?
A: Among rejected bidders in the 2016 auction — where additionality is directly observed for 82% of bidders — there is a systematic positive correlation between bid rental rates (reflecting higher costs of accepting the contract) and additionality. This correlation persists conditional on rich observable characteristics, including prior land use interacted with soil productivity estimates. Contract choice also reveals additionality: bidders selecting tree-related contracts have substantially lower additionality than those choosing base grassland contracts.&lt;/p&gt;
&lt;p&gt;Q: How does soil productivity relate to additionality?
A: USDA-constructed soil productivity estimates, which approximate the earning potential of a parcel, are predictive of additionality in practice, consistent with theory. Higher soil productivity is associated with lower additionality — landowners with less productive land are more likely to conserve regardless of the CRP. Soil productivity is not currently incorporated into the CRP scoring rule to rank bidders.&lt;/p&gt;
&lt;p&gt;Q: How is the RD design validated?
A: The histogram of normalized score distributions shows no bunching at the winning threshold, validating that bidders do not know the exact ex-post threshold realization. Pre-period RD coefficients are indistinguishable from zero in both the remote sensing and administrative land use data. The first stage (share of bidders with a CRP contract just above the threshold) is close to one. Treatment effect magnitudes are stable over the 10-year contract period with no evidence of attenuation, and there are no spillovers to non-bid fields.&lt;/p&gt;
&lt;p&gt;Q: What do the social welfare calculations show for a uniform-price market?
A: Despite widespread non-additionality and adverse selection, a hypothetical uniform-price market for the base conservation contract generates social welfare gains of $14.37 per acre-year at the socially-optimal uniform price. However, setting price equal to the full social benefit B — as the status quo implicitly does by assuming tau = 1 — causes welfare losses of $12.68 per acre-year, nearly eliminating all gains.&lt;/p&gt;
&lt;p&gt;Q: How does the status quo auction perform relative to the efficient benchmark?
A: The status quo auction generates social welfare gains of approximately $120 million per auction relative to no market. The efficient allocation, which awards contracts based on both landowner costs and expected social benefits (incorporating additionality), would be substantially larger. The status quo implements only 12% of the social welfare gains achievable under the efficient allocation.&lt;/p&gt;
&lt;p&gt;Q: Can the efficient allocation be implemented by any mechanism?
A: Not necessarily. Implementing the efficient allocation requires that the expected net social surplus function B·tau(c) - c be monotonically decreasing in cost, so that a standard incentive-compatible auction can rank bidders appropriately. If lower-cost landowners are sufficiently less additional that the allocation rule is non-monotone in cost, no incentive-compatible mechanism can implement the efficient allocation (per Myerson 1981). Empirically, the authors find that for the base contract the efficient allocation is in the implementable case (similar to their Figure 1a), but implementing it exactly via an incentive-compatible auction remains complex.&lt;/p&gt;
&lt;p&gt;Q: What alternative auction designs are proposed, and how much do they improve welfare?
A: The authors propose alternative scoring rules that incorporate expected additionality — through uniform adjustments to the scoring rule, reductions in market size, and differentiation among heterogeneously additional landowners based on observables such as soil productivity and contract choice. These simple modifications close 37% of the gap between the status quo and the efficient allocation, increasing social welfare by over $300 million per auction. Nearly all gains come from incorporating additionality into the scoring rule, with a large share accruing through simple uniform adjustments.&lt;/p&gt;
&lt;p&gt;Q: How is the structural model of bidding estimated?
A: Estimation proceeds in three steps. First, beliefs about the winning score threshold distribution are estimated by simulating auctions via resampling (following Hortacsu 2000). Second, landowner costs are estimated via Maximum Simulated Likelihood using revealed preference inequalities from optimal bidding in the scoring auction. Third, the additionality conditional expectation function is estimated via Method of Simulated Moments, matching observed additionality levels, its distribution across rejected bidders, its covariance with scores, and its distribution by contract choice.&lt;/p&gt;
&lt;p&gt;Q: What sources of scoring rule variation identify the model?
A: Three sources are used. A mid-mechanism policy change in the 2021 auction added carbon sequestration payments differentially across contracts, providing two bids from the same bidders under different scoring rules. A policy change around 2011 shifted Wildlife Priority Zone (WPZ) bonus points to be contract-specific. Air Quality Zone (AQZ) status shifts the level of the score. These sources provide variation in relative payments across contracts, though the authors note the variation is modest and rely also on parametric extrapolation.&lt;/p&gt;
&lt;p&gt;Q: What assumptions are required for identification and how robust are results?
A: Key assumptions include perfect compliance (validated by inspection of over 1,000 aerial photographs), no spillovers to non-bid fields (validated in Table 2), and stability of the additionality function tau(z,c,kappa) across auction years. The authors assess robustness to alternative functional forms of tau, conduct a non-parametric inversion exercise across cost quantiles, and construct alternative scoring rules using cross-auction and cross-tract variation to probe the stability assumption. Model-implied additionality at the RD margin (23%) closely matches the empirical RD estimate.&lt;/p&gt;
&lt;p&gt;Q: Are the adverse selection and additionality findings specific to the 2016 auction?
A: The 2016 auction provides the most complete view because bid fields are observed and 82% of bidders are rejected. But cross-auction evidence replicates the core patterns. RD estimates exploiting threshold variation across auctions show additionality ranging from 10–20% among lower bidders to 40–50% among higher bidders across auctions, consistent with adverse selection. Tree-contract null RD effects replicate across all auctions. Cross-tract cropping rates show similar observable heterogeneity across auctions.&lt;/p&gt;
&lt;p&gt;Q: What is the social welfare impact of the market for conservation existing at all?
A: Theoretically ambiguous because non-additional landowners may receive transfers without generating social value, and adverse selection may tilt the market toward low-additionality participants. Empirically, despite these concerns, there exist positive social welfare gains of $14.37 per acre-year at the socially-optimal uniform price for the base contract, indicating that conservation markets of this type can improve welfare even in the presence of substantial non-additionality and adverse selection.&lt;/p&gt;
&lt;p&gt;Additionality: The expected impact of contracting on a landowner&amp;rsquo;s conservation action — formally, tau(c) = E[1 - a_i0 | c = c_i], the probability that a landowner would not have conserved absent the incentive. A landowner is additional if she would have cropped without the CRP contract; the social benefit of contracting depends only on this incremental conservation impact.&lt;/p&gt;
&lt;p&gt;Adverse Selection: The positive correlation between landowner cost of accepting a contract and additionality. Because landowners with low costs are low-cost partly because they expected to conserve regardless of the program, lower-cost participants are less socially valuable. This upward-sloping contract value curve mirrors adverse selection in insurance markets as modeled by Einav, Finkelstein, and Cullen (2010).&lt;/p&gt;
&lt;p&gt;Contract Value Curve: The function B·tau(F^{-1}_C(q)) plotting the expected social value of contracting at each quantile q of the cost distribution. It lies below the social benefit B due to non-additionality and slopes upward due to adverse selection. The vertical distance between the contract value and marginal cost curves equals expected social surplus B·tau(c) - c.&lt;/p&gt;
&lt;p&gt;Efficient Allocation: The allocation that maximizes expected social surplus B·tau(c) - c by awarding contracts to landowners for whom this quantity is positive. Implementing this allocation via an incentive-compatible mechanism requires that B·tau(c) - c be monotonically decreasing in cost; if not, no standard mechanism can achieve it.&lt;/p&gt;
&lt;p&gt;Scoring Rule: The known function s(b_i, z^s_i) that converts a landowner&amp;rsquo;s multi-dimensional bid (rental rate and contract choice) and observed characteristics into a score, determining contract awards. The status quo scoring rule implicitly assumes full additionality (tau = 1), ranking bidders as if all conservation actions are marginal to the incentive.&lt;/p&gt;
&lt;p&gt;Source Text Origin: The classification of the text on which a summary is based — &amp;ldquo;pdf&amp;rdquo; or &amp;ldquo;oa-html&amp;rdquo; for full working paper text, or &amp;ldquo;abstract-only&amp;rdquo; which is blocked from summarization. Determines the validity and completeness of any summary produced.&lt;/p&gt;</description></item><item><title>Bridging micro and macro production functions: The fiscal multiplier of infrastructure investment</title><link>https://macropaperwarehouse.com/papers/bridging-micro-and-macro-production-functions-the-fiscal-multiplier-of-infrastructure-investment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/bridging-micro-and-macro-production-functions-the-fiscal-multiplier-of-infrastructure-investment/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper investigates the fiscal multiplier of infrastructure investment, specifically by incorporating firm-level investment decisions — a dimension absent from prior literature. The central analytical challenge is bridging the micro (firm-level) and macro (state-level) production functions for infrastructure, given that public capital is non-rivalrous: it can be used simultaneously by all firms without being depleted. The paper demonstrates that this non-rivalry generates a systematic discrepancy between firm-level and aggregate-level estimates of the elasticity of substitution between private and public capital, and it shows how this discrepancy shapes the magnitude of the fiscal multiplier.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors build and estimate a heterogeneous-firm general equilibrium model. Firms operate a constant-elasticity-of-substitution (CES) production function using private capital, non-rivalrous public capital (infrastructure), and labor. Firms are subject to idiosyncratic productivity shocks and make lumpy investment decisions subject to both fixed and convex capital adjustment costs, following Cooper and Haltiwanger (2006) and Winberry (2021). The economy has two regions — one with poor infrastructure and one with good infrastructure — motivated by the near-invariant cross-state distribution of infrastructure spending observed in U.S. data.&lt;/p&gt;
&lt;p&gt;The model is estimated via an extended Simulated Method of Moments (SMM) that treats market clearing prices as additional parameters estimated simultaneously with structural parameters, reducing computational cost relative to standard GE estimation. Estimation uses a multi-block Metropolis-Hastings algorithm. Target moments include lumpy investment fraction (0.14, from Zwick and Mahon 2017), average investment-to-capital ratio (0.10), standard deviation of i/k (0.16), private-to-infrastructure capital ratio (0.75, from BEA), high-infrastructure region&amp;rsquo;s private capital share (0.83, from Census BDS), and total working hours (0.33).&lt;/p&gt;
&lt;p&gt;The identification of the key parameter — the firm-level elasticity of substitution between private and public capital (λ) — comes from the relative size of private capital stocks across the two infrastructure groups: under greater complementarity, regions with more infrastructure should hold relatively more private capital.&lt;/p&gt;
&lt;p&gt;External validation is provided by estimating the state-level elasticity from the model&amp;rsquo;s simulated data using a nonlinear least squares method following An et al. (2019), and comparing it to empirical state-level estimates from actual U.S. state data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings with Quantitative Magnitudes&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Firm-level vs. aggregate-level elasticity gap.&lt;/strong&gt; The estimated firm-level elasticity of substitution is λ = 1.185, implying gross substitutability between private and public capital at the firm level. The state-level elasticity implied by the same model is 0.48 (or 0.35 in a decreasing-returns-to-scale specification), implying gross complementarity. The empirical state-level counterpart estimated from actual U.S. data is 0.445. The paper proves theoretically (Proposition 1) that, given non-rivalry and under mild conditions, firm-level gross substitutability implies aggregate-level gross complementarity. Proposition 2 further shows that this same mechanism micro-founds the increasing-returns-to-scale assumption in Baxter and King&amp;rsquo;s (1993) Cobb-Douglas aggregate production function.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Fiscal multiplier (baseline, 2-year horizon).&lt;/strong&gt; The aggregate output multiplier over a 2-year horizon in the heterogeneous-firm general equilibrium model is &lt;strong&gt;1.088&lt;/strong&gt; in response to a one-time unexpected infrastructure spending shock equal to 1% of steady-state GDP, financed by a lump-sum tax. The corresponding partial-equilibrium output multiplier (holding prices fixed at steady state) is 1.858; the gap reflects crowding out of private investment induced by the general equilibrium interest rate response. In the baseline, the interest rate rises by 0.39% after the shock; the investment multiplier is -0.043.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Comparison with representative-agent model.&lt;/strong&gt; When the same implied returns-to-scale parameters are used in a representative-agent model (following Baxter and King 1993), the output multiplier is 0.991 and the investment multiplier is -0.157, both substantially lower than the heterogeneous-firm baseline. The key mechanism: under convex adjustment costs, the Jensen&amp;rsquo;s inequality effect implies that heterogeneous firms face a greater average adjustment burden than the representative firm, making their investment less responsive to the general equilibrium crowding-out pressure.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Sensitivity to elasticity of substitution.&lt;/strong&gt; Across the heterogeneous-firm model: at λ = 3 (high substitutability), the output multiplier falls to 0.672; at λ = 0.5 (complementarity), it rises to 1.364. The multiplier is significantly more sensitive to λ in the heterogeneous-firm model than in the representative-agent model, because non-rivalry amplifies the effect of any given elasticity value through each firm&amp;rsquo;s production function.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Cross-state distribution of gains.&lt;/strong&gt; Under the baseline spending allocation (81% to Good states, 19% to Poor states), per $1 of infrastructure spending, Good states receive $1.072 of the $1.088 total output gain, while Poor states receive only $0.016. In a counterfactual with equal spending across states, the total output multiplier falls to 0.873, Good states&amp;rsquo; output multiplier falls to 0.810, and Poor states&amp;rsquo; output multiplier rises to approximately 0.062 (about four times the baseline level of 0.016). This quantifies a sharp efficiency-equality trade-off in the allocation of infrastructure investment.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Employment and earnings effects.&lt;/strong&gt; Compared to steady state, the baseline fiscal shock produces an average annual increase of 0.304% in employment and 0.389% in wages, yielding a $0.713 increase in earnings and a $0.148 increase in consumption per $1 of fiscal spending in general equilibrium. In partial equilibrium (no price changes), earnings increase by $1.294 and consumption by $0.605 per $1 spent.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Results are conditional on: (i) lump-sum tax financing of the fiscal shock; (ii) a one-time unexpected (MIT) shock with no persistence; (iii) a closed-economy framework with endogenous real interest rate; (iv) the estimated two-region structure calibrated to U.S. state-level infrastructure data; (v) firm-level investment dynamics calibrated to Compustat and BDS moments. The authors note that incorporating time-to-build assumptions (tested in an appendix) reduces the aggregate fiscal multiplier, consistent with Ramey (2020).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-theoretical-result-connecting-firm-level-and-aggregate-level-elasticities-and-what-is-the-intuition"&gt;Q1. What is the core theoretical result connecting firm-level and aggregate-level elasticities, and what is the intuition?&lt;/h3&gt;
&lt;p&gt;A: Proposition 1 proves that, given non-rivalrous public capital and mild data conditions (at least one firm has private capital below total infrastructure, and aggregate private capital exceeds total infrastructure), if the firm-level elasticity of substitution λ ≥ 1 (gross substitutes), then the aggregate-level elasticity ξ &amp;lt; 1 (gross complements). The intuition is that a marginal increase in public capital raises the marginal product of private capital for every firm simultaneously due to non-rivalry; the sum of these MPK gains across all firms exceeds any single firm&amp;rsquo;s gain. To represent this amplified benefit within an aggregate production function, a stronger complementarity is required than what any single firm faces. Put differently, non-rivalry means aggregate private and public capital &amp;ldquo;look&amp;rdquo; more complementary than they truly are at the firm level.&lt;/p&gt;
&lt;h3 id="q2-how-does-non-rivalry-micro-found-the-baxter-king-aggregate-production-function"&gt;Q2. How does non-rivalry micro-found the Baxter-King aggregate production function?&lt;/h3&gt;
&lt;p&gt;A: Proposition 2 shows that if firms use a CES production function with gross substitutability (λ ≥ 1) and non-rivalrous public capital, then fitting aggregate output with a Cobb-Douglas production function (as in Baxter and King 1993, H(K,N,L) = zK^α L^{1-α} N^ζ) yields ζ &amp;gt; 0, implying increasing returns to scale (IRS). This is the paper&amp;rsquo;s micro-foundation for a widely-used but previously ad hoc assumption in the macro-fiscal literature. The corollary states that both gross complementarity in the aggregate CES function and IRS in the aggregate Cobb-Douglas follow from the same non-rivalry mechanism at the firm level.&lt;/p&gt;
&lt;h3 id="q3-why-does-the-heterogeneous-firm-model-produce-a-higher-output-multiplier-than-the-representative-agent-model"&gt;Q3. Why does the heterogeneous-firm model produce a higher output multiplier than the representative-agent model?&lt;/h3&gt;
&lt;p&gt;A: Two mechanisms drive the difference. First, due to Jensen&amp;rsquo;s inequality and the convexity of adjustment costs, heterogeneous firms face a higher average adjustment burden than the representative (average) firm; this means heterogeneous firms are less responsive to interest rate changes that crowd out investment. The investment multiplier is -0.043 in the heterogeneous-agent baseline versus -0.157 in the representative-agent model. Second, the fixed adjustment cost (present in the baseline but absent from the representative-agent model) further dampens investment sensitivity via the extensive margin. Because less private investment is crowded out, more of the direct output boost from infrastructure spending survives into the aggregate multiplier, yielding 1.088 versus 0.991.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-novel-estimation-procedure-and-why-is-it-necessary"&gt;Q4. What is the novel estimation procedure and why is it necessary?&lt;/h3&gt;
&lt;p&gt;A: Standard SMM applied to GE models requires solving for market-clearing prices for every candidate parameter vector, creating a nested optimization loop that is computationally prohibitive. The authors extend SMM by treating market-clearing prices (wage w and marginal utility of consumption p) as additional parameters and appending market-clearing conditions as additional target moments — effectively requiring those moments to equal zero. A multi-block Metropolis-Hastings algorithm jointly draws from the price block and the parameter block. This approach generates posterior draws that simultaneously satisfy market clearing and fit empirical moments, without the inner loop. The resulting market-clearing accuracy is e^{-4} at the posterior mean.&lt;/p&gt;
&lt;h3 id="q5-how-is-the-firm-level-elasticity-of-substitution-λ-identified-from-the-data"&gt;Q5. How is the firm-level elasticity of substitution (λ) identified from the data?&lt;/h3&gt;
&lt;p&gt;A: λ is identified from the cross-state difference in private capital stocks between high- and low-infrastructure regions. Under the model, if private and public capital are more complementary (lower λ), high-infrastructure regions should attract relatively more private capital. The data moment used is the Good region&amp;rsquo;s share of aggregate private capital (0.83 from Census BDS data). This identification strategy is analogous to Bartik-instrument approaches in the empirical literature, where a parameter governing cross-state sensitivity to aggregate shocks is identified from cross-sectional variation.&lt;/p&gt;
&lt;h3 id="q6-how-is-the-model-validated-externally"&gt;Q6. How is the model validated externally?&lt;/h3&gt;
&lt;p&gt;A: The authors compute the state-level elasticity from the estimated model by fixing firm-level parameters and re-estimating only the elasticity and regional productivity from the model&amp;rsquo;s simulated state-level data, using the same NLLS estimator as An et al. (2019). The model-implied state-level elasticity is 0.349 (DRS specification) or 0.482 (CRS specification). The empirical estimate from actual U.S. state-level data following the same estimator is 0.445. Both indicate gross complementarity at the state level, consistent with the theoretical prediction. This external validation is not used in the estimation itself, providing an independent check.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-roles-of-extensive-vs-intensive-investment-margins-in-the-crowding-out-effect"&gt;Q7. What are the roles of extensive vs. intensive investment margins in the crowding-out effect?&lt;/h3&gt;
&lt;p&gt;A: Table 9 decomposes the investment multiplier of -0.043 by investment margin. When only the extensive margin (the discrete decision of whether to invest) is allowed to respond, the investment multiplier is -0.032 — approximately 74% of the baseline crowding-out effect. When only the intensive margin (investment size conditional on adjusting) responds, the multiplier is -0.011 — about 25% of the total. Thus the extensive margin is the dominant channel through which higher interest rates crowd out private investment. When both margins are held fixed, the output multiplier rises to 1.139, confirming that investment crowding-out reduces the output multiplier by about 0.05.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-elasticity-of-substitution-affect-the-fiscal-multiplier-quantitatively-and-why-does-this-matter-more-in-the-heterogeneous-firm-model"&gt;Q8. How does the elasticity of substitution affect the fiscal multiplier quantitatively, and why does this matter more in the heterogeneous-firm model?&lt;/h3&gt;
&lt;p&gt;A: In the heterogeneous-firm GE model: λ = 3 gives an output multiplier of 0.672, λ = 1.185 (baseline) gives 1.088, and λ = 0.5 gives 1.364 — a range of 0.692. In the representative-agent model, the comparable range across the implied ζ values is much narrower (0.970 to 0.998). The amplification in the heterogeneous-firm model occurs because non-rivalry means each firm&amp;rsquo;s production function directly incorporates the public capital stock, so the elasticity parameter has first-order consequences for every firm&amp;rsquo;s investment incentive response to a fiscal shock. This heightened sensitivity underscores why accurately estimating λ at the firm level — rather than importing a state-level estimate — is critical for quantifying infrastructure multipliers.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-efficiency-equality-trade-off-in-cross-state-infrastructure-allocation"&gt;Q9. What is the efficiency-equality trade-off in cross-state infrastructure allocation?&lt;/h3&gt;
&lt;p&gt;A: Under the baseline allocation (81% of infrastructure spending to Good states, 19% to Poor states), per $1 of infrastructure spending, the Good states receive $1.072 of output gains and Poor states receive only $0.016. In the equal-spending counterfactual, the total output multiplier falls from 1.088 to 0.873. The Poor states&amp;rsquo; output multiplier rises from $0.016 to $0.062 (approximately fourfold), while the Good states&amp;rsquo; falls from $1.072 to $0.810. The Poor states also see earnings multipliers more than double (from $0.017 to $0.042). This trade-off arises because Good states have both more private capital (benefiting from non-rivalry) and higher estimated TFP — so each dollar of infrastructure is more productive there. Equal allocation reduces aggregate efficiency while partially mitigating regional inequality.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-papers-multiplier-estimates-compare-to-the-existing-literature"&gt;Q10. How do the paper&amp;rsquo;s multiplier estimates compare to the existing literature?&lt;/h3&gt;
&lt;p&gt;A: In partial equilibrium (no GE adjustment), the authors find an output multiplier of 1.858, consistent with Chodorow-Reich&amp;rsquo;s (2019) cross-sectional multiplier of approximately 1.8. Once the general equilibrium interest rate effect is included, the multiplier falls to 1.09, which falls within the 0.6-1.2 range from Ramey (2011). Literature using representative-agent models without non-rivalry (e.g., Ramey 2020) typically reports multipliers of 0.3 to 0.8 using returns-to-scale parameters of 0.07-0.12; the paper shows these correspond to fiscal multipliers of 0.847-0.882 in the representative-agent framework. The heterogeneous-firm model, once it incorporates the non-rivalry-corrected elasticities, yields a meaningfully higher multiplier of 1.088.&lt;/p&gt;
&lt;h3 id="q11-what-role-does-time-to-build-play-and-how-does-the-paper-handle-it"&gt;Q11. What role does time-to-build play, and how does the paper handle it?&lt;/h3&gt;
&lt;p&gt;A: The baseline model assumes a time-to-build period s = 1 year (one-year lag before new infrastructure is productive). The paper notes in Appendix H that incorporating extended time-to-build reduces the aggregate fiscal multiplier, operating through two channels: a news effect (agents adjust behavior upon anticipating future infrastructure) and a general equilibrium effect endogenous to the news effect. This finding is consistent with Ramey (2020). The baseline results are therefore reported under the minimal one-year time-to-build assumption, with longer lags serving as a robustness check.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-role-of-region-specific-tfp-heterogeneity-in-the-model"&gt;Q12. What is the role of region-specific TFP heterogeneity in the model?&lt;/h3&gt;
&lt;p&gt;A: The model includes two regions that differ both in infrastructure levels and in region-specific productivity (TFP) levels. The TFP of the Good region is estimated to be approximately double that of the Poor region (x = 2.064 for Good vs. 1 for Poor). This productivity difference is estimated to partially capture heterogeneous congestion effects (which are not separately modeled) and is estimated jointly with the infrastructure elasticity. The productivity differential is identified from the Good region&amp;rsquo;s share of aggregate output (0.849 in the data). The large TFP gap is also the reason why equal spending on Poor states generates a much smaller output gain than spending on Good states: not only is infrastructure utilization lower (fewer firms), but underlying productivity is also lower.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Non-rivalry of public capital&lt;/strong&gt;: The property by which infrastructure stock (Nj,t) enters each firm&amp;rsquo;s production function at the full regional level, not divided among firms. Formally, a single marginal unit of public capital raises every firm&amp;rsquo;s marginal product of private capital simultaneously, so the aggregate marginal product gain summed across firms exceeds any single firm&amp;rsquo;s gain. This is the central mechanism driving the micro-macro elasticity discrepancy in the paper.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Firm-level elasticity of substitution (λ)&lt;/strong&gt;: The elasticity governing the degree of substitutability between private capital (k) and public infrastructure (N) in the firm&amp;rsquo;s CES production function. At λ = 1 the production function is Cobb-Douglas; λ &amp;gt; 1 is gross substitutability; λ &amp;lt; 1 is gross complementarity. In the paper&amp;rsquo;s estimation, λ = 1.185, meaning private and public capital are gross substitutes at the firm level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gross substitutability vs. gross complementarity&lt;/strong&gt;: Two inputs are gross substitutes (complements) if an increase in the quantity of one raises (lowers) the demand for the other, holding output price fixed. In the paper&amp;rsquo;s framework, private and public capital are gross substitutes at the firm level (λ = 1.185 &amp;gt; 1) but gross complements at the state level (ξ ≈ 0.48 &amp;lt; 1), with non-rivalry explaining the inversion upon aggregation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Convex adjustment cost&lt;/strong&gt;: A cost C(I,k) = (µ/2)(I/k)² · k that scales quadratically with the investment rate. In the heterogeneous-firm model, this cost plays a critical role: by Jensen&amp;rsquo;s inequality, heterogeneous firms&amp;rsquo; average adjustment burden under a convex cost exceeds that of the representative (average) firm, making aggregate investment less sensitive to interest rate changes and thereby dampening crowding out.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fixed adjustment cost (ξ)&lt;/strong&gt;: A one-time overhead cost drawn from a uniform distribution [0, ξ̄], paid only when a firm makes a large-scale investment outside the &amp;ldquo;inaction band&amp;rdquo; [−νk, νk]. This cost generates lumpy investment at the firm level, with about 14% of firms making lumpy investments in any given year. It also creates an extensive margin of investment adjustment that accounts for approximately 74% of the baseline crowding-out effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fiscal multiplier (as defined in this paper)&lt;/strong&gt;: The ratio of the present value of aggregate output deviations from steady state to the present value of the fiscal spending shock, both summed over a T-year horizon. For the short run, T = 2 years; for the long run, T = 5 years. This is computed as a perfect-foresight transition path response to a one-time MIT shock equal to 1% of steady-state GDP.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;MIT shock (one-time unexpected shock)&lt;/strong&gt;: An unanticipated, non-persistent one-period deviation in infrastructure spending. The term &amp;ldquo;MIT shock&amp;rdquo; refers to a deterministic transition experiment where agents have perfect foresight about all future values after the initial shock occurs. This contrasts with persistent policy rules and allows isolating the dynamic effects of a one-time fiscal impulse.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extended SMM with market-clearing moments&lt;/strong&gt;: The paper&amp;rsquo;s estimation innovation. Rather than solving for market-clearing prices at each parameter candidate (the standard costly inner loop), wages (w) and marginal utility of consumption (p) are treated as parameters with associated moments being the market-clearing conditions set to zero. A multi-block Metropolis-Hastings algorithm draws from the price block and the parameter block separately, generating posterior draws that jointly satisfy market clearing and empirical moment conditions.&lt;/p&gt;</description></item><item><title>Cap‐and‐Trade and Carbon Tax Meet Arrow–Debreu</title><link>https://macropaperwarehouse.com/papers/capandtrade-and-carbon-tax-meet-arrowdebreu/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/capandtrade-and-carbon-tax-meet-arrowdebreu/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Anderson and Duanmu (2025) ask how general equilibrium (GE) interactions — factor reallocation across sectors, capital misallocation under climate uncertainty, and the distributional incidence of damages — alter the social cost of carbon (SCC) relative to the partial equilibrium (PE) estimates embedded in standard integrated assessment models (IAMs). The paper also characterizes conditions for Pareto improvements through climate policy and derives the optimal carbon tax in second-best environments with pre-existing distortions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Framework&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors build a dynamic Arrow-Debreu economy with L goods, K capital stocks (including climate stocks), and T periods. The climate module specifies that the carbon stock evolves as S_{t+1} = S_t + sum_j e_j(q_j) − alpha·S_t, and climate damage functions D_j(S_t) = 1 − d_j·(S_t − S_0) reduce sector-specific production possibilities sets. Firms and households take the climate trajectory as given and do not internalize their own emissions&amp;rsquo; impact, generating the externality. Under standard regularity conditions, the authors prove existence of a competitive equilibrium and establish that it is inefficient: output is too high and climate-intensive sectors are too large relative to the social optimum.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;General Formula for the SCC&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper derives a general SCC formula — SCC_t = Sum_{tau &amp;gt;= t} beta^(tau−t) · [dW/dS_tau / (dW/dY_t)] — that decomposes into four components: (1) the standard direct productivity-loss term, (2) a GE factor-reallocation term capturing inefficient reallocation as damages shift relative prices, (3) a capital-misallocation term reflecting distortions in investment from climate uncertainty, and (4) a distribution term reflecting the welfare losses from the regressive incidence of climate damages. All three correction terms are positive under standard conditions, so the GE SCC exceeds the PE SCC. The paper shows that this formula nests existing IAM frameworks as special cases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Calibrating to three leading IAMs, the authors find that general equilibrium interactions raise the SCC by 15–40% above standard PE estimates:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;DICE-calibrated: GE correction of &lt;strong&gt;18%&lt;/strong&gt; above the PE estimate.&lt;/li&gt;
&lt;li&gt;FUND-calibrated: GE correction of &lt;strong&gt;15%&lt;/strong&gt; above the PE estimate.&lt;/li&gt;
&lt;li&gt;PAGE-calibrated: GE correction of &lt;strong&gt;40%&lt;/strong&gt; above the PE estimate, the largest correction owing to greater sector heterogeneity in that model.&lt;/li&gt;
&lt;li&gt;Median calibration: a PE SCC of &lt;strong&gt;$51/tCO₂&lt;/strong&gt; rises to a GE SCC of &lt;strong&gt;$62/tCO₂&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Decomposing the aggregate GE correction: factor reallocation across sectors accounts for &lt;strong&gt;55%&lt;/strong&gt;, capital misallocation due to climate uncertainty for &lt;strong&gt;30%&lt;/strong&gt;, and the distributional regressivity of damages for &lt;strong&gt;15%&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-Best Policy and Uncertainty&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In environments with pre-existing distortions, the optimal carbon tax deviates from the SCC: revenue recycling through labor tax cuts generates additional welfare gains of &lt;strong&gt;10–15%&lt;/strong&gt; of carbon tax revenue; undertaxed capital implies the optimal carbon tax should be set above the SCC (double dividend); and in monopolistically competitive sectors the optimal carbon tax is below the SCC because the carbon tax amplifies monopoly distortions. Under climate uncertainty, the SCC carries a risk premium proportional to the variance of damage estimates times the coefficient of relative risk aversion, estimated at &lt;strong&gt;+$8–15/tCO₂&lt;/strong&gt; (15–25% of the base SCC).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The quantitative corrections are calibrated to DICE, FUND, and PAGE and therefore inherit those models&amp;rsquo; parameterizations of damage functions and discount rates. The GE factor-reallocation and capital-misallocation channels are larger when sectors are more heterogeneous in damage exposure — as is explicit in the PAGE result. Second-best corrections depend on the sign and magnitude of pre-existing distortions (labor taxes, capital taxes, market structure).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-inefficiency-result-and-what-does-it-imply-about-the-competitive-equilibrium"&gt;Q1. What is the core inefficiency result, and what does it imply about the competitive equilibrium?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s efficiency theorem establishes that the competitive equilibrium is Pareto inefficient because firms and households take the climate trajectory as given and do not internalize the impact of their own emissions on the carbon stock. As a consequence, output is too high and climate-intensive sectors are too large relative to the social optimum. This externality is the fundamental justification for climate policy in the model.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-papers-general-scc-formula-extend-existing-approaches-and-what-are-the-novel-terms"&gt;Q2. How does the paper&amp;rsquo;s general SCC formula extend existing approaches, and what are the novel terms?&lt;/h3&gt;
&lt;p&gt;The general formula SCC_t = Sum_{tau &amp;gt;= t} beta^(tau−t) · [dW/dS_tau / (dW/dY_t)] nests standard IAM SCC formulas as special cases. The novel terms relative to partial equilibrium are: (i) a GE reallocation term capturing losses from inefficient factor reallocation as climate damages change relative prices across sectors; (ii) a capital-misallocation term capturing distortions in investment arising from climate uncertainty; and (iii) a distribution term capturing welfare losses from the regressive incidence of damages. All three terms are positive under standard conditions, implying GE SCC &amp;gt; PE SCC in all calibrations.&lt;/p&gt;
&lt;h3 id="q3-how-are-the-quantitative-ge-corrections-decomposed-and-which-channel-dominates"&gt;Q3. How are the quantitative GE corrections decomposed, and which channel dominates?&lt;/h3&gt;
&lt;p&gt;Of the total GE correction above the PE baseline, factor reallocation across sectors contributes 55%, capital misallocation due to climate uncertainty contributes 30%, and the distributional regressivity of damages contributes 15%. Factor reallocation is the dominant channel because, as climate damages alter relative prices, production shifts toward less-damaged sectors in ways that are distorted by the original carbon externality — generating second-order losses absent from PE damage functions.&lt;/p&gt;
&lt;h3 id="q4-why-does-the-page-calibration-produce-a-larger-ge-correction-40-than-dice-18-or-fund-15"&gt;Q4. Why does the PAGE calibration produce a larger GE correction (40%) than DICE (18%) or FUND (15%)?&lt;/h3&gt;
&lt;p&gt;The paper attributes PAGE&amp;rsquo;s larger GE correction to greater sector heterogeneity in that model&amp;rsquo;s parameterization. When damage exposure is more heterogeneous across sectors, the relative-price effects of marginal carbon are larger, amplifying the factor-reallocation channel. DICE and FUND, with more uniform sector-level damage structures, exhibit smaller reallocation corrections.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-median-calibration-implication-for-the-scc-in-dollar-terms"&gt;Q5. What is the median-calibration implication for the SCC in dollar terms?&lt;/h3&gt;
&lt;p&gt;In the median calibration, a PE SCC of $51/tCO₂ rises to a GE SCC of $62/tCO₂, an increase of roughly $11/tCO₂ or approximately 22%. This figure is directly computable from observable trade elasticities and sector-level damage estimates.&lt;/p&gt;
&lt;h3 id="q6-how-should-the-carbon-tax-be-adjusted-when-pre-existing-labor-market-distortions-are-present-and-what-is-the-magnitude-of-the-welfare-gain-from-revenue-recycling"&gt;Q6. How should the carbon tax be adjusted when pre-existing labor market distortions are present, and what is the magnitude of the welfare gain from revenue recycling?&lt;/h3&gt;
&lt;p&gt;When labor taxes create a pre-existing wedge, using carbon tax revenue to reduce labor taxes generates additional welfare gains of 10–15% of total carbon tax revenue — the double dividend in the labor market dimension. The optimal carbon tax in this case includes the SCC plus a correction term for the labor-market distortion.&lt;/p&gt;
&lt;h3 id="q7-how-do-capital-market-distortions-alter-the-optimal-carbon-tax-relative-to-the-scc"&gt;Q7. How do capital market distortions alter the optimal carbon tax relative to the SCC?&lt;/h3&gt;
&lt;p&gt;If capital is undertaxed (a pre-existing distortion in capital markets), the optimal carbon tax is set above the SCC. The intuition is that a higher carbon tax partially offsets the under-taxation of capital by raising the effective cost of carbon-intensive investment, capturing a double-dividend in the capital market.&lt;/p&gt;
&lt;h3 id="q8-how-does-monopolistic-competition-modify-the-optimal-carbon-tax"&gt;Q8. How does monopolistic competition modify the optimal carbon tax?&lt;/h3&gt;
&lt;p&gt;For monopolistically competitive sectors, the optimal carbon tax is below the SCC. The reasoning is that applying a carbon tax to these sectors amplifies existing monopoly markups and associated distortions, so the social cost of the carbon tax exceeds the raw SCC in those sectors. The optimal policy trades off carbon correction against monopoly amplification.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-risk-premium-in-the-scc-under-climate-uncertainty-and-how-is-it-estimated"&gt;Q9. What is the risk premium in the SCC under climate uncertainty, and how is it estimated?&lt;/h3&gt;
&lt;p&gt;The paper adds a term to the SCC proportional to the variance of damage estimates times the coefficient of relative risk aversion. Using empirical estimates of damage uncertainty, this risk premium is estimated at +$8–15/tCO₂, representing 15–25% of the base SCC. This term is absent from deterministic SCC calculations and constitutes a further reason standard PE estimates understate the true social cost.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-papers-claim-regarding-computability-of-the-ge-correction"&gt;Q10. What is the paper&amp;rsquo;s claim regarding computability of the GE correction?&lt;/h3&gt;
&lt;p&gt;The paper states that the novel GE terms are computable from observable trade elasticities and sector-level damage estimates, implying the GE correction is not merely a theoretical construct but can be implemented in quantitative policy analysis using data sources already available to researchers and policymakers.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Social Cost of Carbon (General Equilibrium Formula)&lt;/strong&gt;
Defined in the paper as SCC_t = Sum_{tau &amp;gt;= t} beta^(tau−t) · [dW/dS_tau / (dW/dY_t)], the present discounted value of the marginal welfare loss from an additional unit of carbon, expressed relative to the marginal utility of current output. The paper&amp;rsquo;s version adds GE reallocation, capital-misallocation, and distributional terms absent from standard PE formulations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GE Adjustment Factor&lt;/strong&gt;
The ratio of the general equilibrium SCC to the partial equilibrium SCC, expressed as GE/PE = 1 + phi_realloc + phi_capital + phi_distribution. Under standard conditions all three phi terms are positive, so the GE SCC strictly exceeds the PE SCC.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Climate Damage Function (Sector-Specific)&lt;/strong&gt;
Specified as D_j(S_t) = 1 − d_j·(S_t − S_0), a sector-specific multiplicative reduction in the production possibilities set as the carbon stock rises above the pre-industrial level S_0. Heterogeneity in d_j across sectors is the driver of the factor-reallocation GE correction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Carbon Stock Evolution&lt;/strong&gt;
S_{t+1} = S_t + sum_j e_j(q_j) − alpha·S_t, where alpha is the natural decay rate of atmospheric carbon and e_j(q_j) is sectoral emissions as a function of output. Firms and households treat S_t as exogenous, generating the externality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double Dividend&lt;/strong&gt;
In second-best environments, a carbon tax can generate two welfare gains simultaneously: correcting the carbon externality and reducing the deadweight loss from a pre-existing distortion (labor or capital tax). The paper finds revenue recycling via labor tax cuts yields 10–15% of carbon tax revenue as additional welfare gain; undertaxed capital implies the optimal carbon tax is set above the SCC.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Risk Premium in the SCC&lt;/strong&gt;
An additive term in the SCC under climate uncertainty, proportional to the variance of damage estimates times the coefficient of relative risk aversion. Empirically estimated at +$8–15/tCO₂, representing 15–25% of the base SCC.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-Best Optimal Carbon Tax&lt;/strong&gt;
Written as tau*_carbon = SCC + CORRECTION, where the correction depends on the sign and magnitude of pre-existing distortions. The correction is positive under undertaxed capital (raise above SCC), negative under monopolistic competition (lower below SCC), and augmented by revenue-recycling gains when labor taxes are present.&lt;/p&gt;</description></item><item><title>Catastrophes, Delays, and Learning</title><link>https://macropaperwarehouse.com/papers/catastrophes-delays-and-learning/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/catastrophes-delays-and-learning/</guid><description>&lt;p&gt;This paper develops a general model of experimentation under catastrophe risk in which the catastrophe is triggered when a stock variable exceeds an unknown threshold, but occurs only after a stochastic delay. The central contribution is the concept of the &amp;ldquo;legacy of the past&amp;rdquo;: at any planning date, past experiments may have already triggered a catastrophe that has not yet materialized, and the planner cannot observe whether triggering has occurred. The legacy is formally defined as the probability, conditional on survival, that a catastrophe was triggered in the past.&lt;/p&gt;
&lt;p&gt;The model unifies two canonical but previously incompatible approaches in the literature. In the hazard-rate approach, the catastrophe is bound to happen and the planner manages its timing and severity. In the unknown-threshold approach, learning is instantaneous and the catastrophe is certainly avoided if the stock has not yet exceeded the threshold. Neither approach captures the intermediate case where the planner remains uncertain about whether the catastrophe is already underway. By introducing a delay governed by an exponential distribution with parameter α, the authors show that both approaches are limiting special cases: as α → ∞ (no delay), the legacy vanishes and the unknown-threshold approach is recovered; when the legacy is set permanently to one (catastrophe triggered with certainty), the hazard-rate approach is recovered.&lt;/p&gt;
&lt;p&gt;Three benchmark stock levels anchor the analysis. QN is the long-run target absent any catastrophe risk. QD (&amp;ldquo;Damages&amp;rdquo;) is the optimal stabilization target when the planner knows a catastrophe was triggered in the past — it lies weakly below QN because the planner trades off current gains against the discounted marginal damage from raising the stock at the moment of eventual catastrophe occurrence. QE (&amp;ldquo;Experimentation&amp;rdquo;) is the stock level below which stabilization is suboptimal when the planner is certain no triggering has occurred — it also lies weakly below QN.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s two main theorems are distinguished by the ranking of QD and QE, which reflects whether mitigation strategies are effective.&lt;/p&gt;
&lt;p&gt;Theorem 1 (QE &amp;lt; QD): When damage is not highly sensitive to the stock level at catastrophe time — so mitigation is relatively ineffective — optimal paths are monotonically increasing and converge to a long-run stock level Q∞ ∈ [QE, QD]. The stopping condition equates the marginal benefit of experimentation to a weighted average of the expected cost under the unknown-threshold approach (weight 1 − π) and the marginal damage under the hazard-rate approach (weight π), where π is the legacy at stopping time. A higher legacy at the stopping time is associated with a higher long-run stock level. A higher initial legacy induces fatalism: since the catastrophe is more likely already triggered, the planner shifts priority toward current consumption rather than caution, leading to more total experimentation.&lt;/p&gt;
&lt;p&gt;Theorem 2 (QD &amp;lt; QE): When damage is highly sensitive to the stock level — so mitigation is valuable — the long-run target is uniquely QE regardless of the initial legacy. However, the short-run path is non-monotonic: for a sufficiently high initial legacy, the planner first reduces the stock sharply (lockdown, emissions cut) to mitigate pending catastrophe damages, then, as the legacy declines because no catastrophe occurs, gradually allows the stock to rise back toward QE. The direction of caution reverses relative to Theorem 1: a higher legacy now induces more caution, not less.&lt;/p&gt;
&lt;p&gt;Applications include pandemic management (stock = infected population, catastrophe = health system collapse) and climate change (stock = cumulative CO2 emissions or atmospheric pollution stock). In the disease control application, whether a planner prioritizes economic production or mortality reduction determines which theorem governs, with the key ratio being production losses relative to mortality increases. For pandemic policy, Theorem 2 produces a formal learning-based rationale for non-monotonic &amp;ldquo;hammer-and-dance&amp;rdquo; policies (strict early lockdown followed by relaxation) that differs from prior explanations in the literature. In the carbon budget application, Proposition 5 formally proves that higher initial legacy raises the optimal carbon budget under Theorem 1 conditions, and can imply unbounded consumption (certainty of catastrophe) above a critical legacy threshold π*. Under Theorem 2 conditions (Proposition 6), the optimal policy can involve first reducing then expanding the stock before stabilizing, with both transition dates increasing in the initial legacy.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;legacy of the past&amp;rdquo; and how is it computed?
A: The legacy πt is defined as the probability, conditional on survival to date t, that a catastrophe was already triggered by past experiments. Formally, πt = 1 − [1 − F(Qt)] / pt, where Qt is the highest stock level ever reached, F is the prior distribution over the threshold, and pt is the survival probability. A past experiment at time t&amp;rsquo; contributes to the current legacy with weight exp[−α(t − t&amp;rsquo;)], so recent experiments matter more than distant ones. As time passes without catastrophe, the legacy of any fixed past experiment declines geometrically at rate α.&lt;/p&gt;
&lt;p&gt;Q: How do the three benchmark stock levels QN, QD, and QE relate to each other?
A: QN is the optimal long-run stock without any catastrophe. QD is defined by the condition where the marginal net benefit of increasing the stock — ν(Q) − [α/(α+δ)]D&amp;rsquo;(Q) — equals zero, and satisfies QD ≤ QN. QE is defined by ν(Q) − [α/(α+δ)]ρ(Q)D(Q) = zero, and also satisfies QE ≤ QN. The ranking between QD and QE depends on whether damage is more sensitive to the marginal increase in stock at catastrophe time (which pushes QD below QE) or to the level of the stock at triggering (which pulls QD above QE).&lt;/p&gt;
&lt;p&gt;Q: What is the key optimality condition in Theorem 1 and how does it unify prior approaches?
A: The stopping condition (equation 15) states: ν(QT) = [α/(α+δ)] × [(1 − πT)ρ(QT)D(QT) + πT D&amp;rsquo;(QT)]. When πT = 0 (no legacy, unknown-threshold limit), this reduces to the experimentation stopping condition of Tsur and Zemel, governed by the hazard rate ρ(QT) times expected loss D(QT). When πT = 1 (full legacy, hazard-rate limit), it reduces to the damage-mitigation condition governed by marginal damage D&amp;rsquo;(QT). The legacy at stopping time thus serves as the mixing weight between the two canonical approaches, embedding both as special cases.&lt;/p&gt;
&lt;p&gt;Q: How does the initial legacy affect total experimentation under Theorem 1 versus Theorem 2?
A: Under Theorem 1 (QE &amp;lt; QD), a higher initial legacy π0 leads to more total experimentation (higher Q∞), because the planner becomes fatalistic — since the catastrophe is more likely already triggered and mitigation is relatively ineffective, current consumption is prioritized. Proposition 5 formally proves this for the carbon budget application: the optimal stopping date T and optimal budget QT are nondecreasing in π0. Under Theorem 2 (QD &amp;lt; QE), a higher legacy triggers more caution in the short run (larger reduction in the stock during the mitigation phase), but the long-run target QE remains the same regardless of π0.&lt;/p&gt;
&lt;p&gt;Q: What generates non-monotonic policies in Theorem 2, and what does this look like in the pandemic application?
A: Non-monotonicity arises because the optimal response to a high legacy is first to reduce the stock sharply to limit catastrophe damages (since damage is sensitive to the stock level), and then, as time passes without catastrophe and the legacy declines, to allow the stock to recover. In the disease control application with high mortality weight, a complete lockdown is optimal in the first phase whenever the legacy is strictly positive. As the legacy declines, the lockdown is gradually relaxed, and eventually the infection level returns to its pre-lockdown level. Figures 3 and 4 show that a higher initial legacy (π0 = 0.1, 0.5, or 0.9) leads to a longer lockdown and slower recovery, though all paths converge to the same long-run infection level.&lt;/p&gt;
&lt;p&gt;Q: How does the model&amp;rsquo;s disease control application determine which theorem governs?
A: Lemma 2 states that if 1 / [1 + (Y(r+d) − Y*) / (wµ&lt;em&gt;dI^D)] &amp;lt; ρ(I^D), then I^E &amp;lt; I^D and Theorem 1 applies; otherwise I^E &amp;gt; I^D and Theorem 2 applies. The key ratio is (Y(r+d) − Y&lt;/em&gt;) / (wµ*d), the production loss relative to mortality increase. A planner who weights economic activity heavily (large production loss ratio) falls under Theorem 1 and tolerates rising infections; a planner who weights mortality heavily falls under Theorem 2 and imposes an initial lockdown.&lt;/p&gt;
&lt;p&gt;Q: What is the carbon budget result under Theorem 1 (Proposition 5)?
A: Under the condition u1 &amp;gt; [α/(α+δ)]v0 (marginal consumption value exceeds discounted marginal damage), Theorem 1 applies and there exists a critical legacy threshold π* such that: below π*, the planner consumes maximally (qt = q-bar) until a finite date T and then stops, with QE &amp;lt; QT &amp;lt; QD; above π*, the planner consumes maximally forever, triggering the catastrophe with certainty. The stopping date T and the optimal budget QT are nondecreasing functions of initial legacy π0, formally proving that higher past emissions (captured through legacy) justify higher future carbon budgets in this model.&lt;/p&gt;
&lt;p&gt;Q: What is the carbon budget result under Theorem 2 (Proposition 6)?
A: Under condition u1 &amp;lt; [α/(α+δ)]v0, QD &amp;lt; QE and Theorem 2 applies. Starting from Q0 above QE, if π0 is small enough (specifically u1 &amp;gt; π0[α/(α+δ)]v0), the optimal policy is to stabilize the stock forever at Q0. Otherwise, there exist two finite dates t1 &amp;lt; t2, both increasing in π0, such that the planner first reduces the stock at maximum rate (qt = q-bar-negative) for t &amp;lt; t1, then expands at maximum rate for t1 &amp;lt; t &amp;lt; t2, then stabilizes at Q0 forever. The optimal carbon budget is Q0 in all cases, showing that the long-run target is independent of legacy under Theorem 2.&lt;/p&gt;
&lt;p&gt;Q: How does the model relate to the hazard-rate literature formally?
A: Papers such as Nordhaus and others that use an exogenous hazard rate h(Qt) for catastrophe — yielding survival probability pt = p0 exp(−∫h(Qτ)dτ) — are shown to be equivalent to the special case where the catastrophe was triggered in the past (legacy = 1 permanently). Their formulation corresponds to assuming α is constant and the legacy is identically one, which reduces the law of motion for pt to pt = p0 exp(−αt). The key difference is that in the hazard-rate approach the planner can reduce the arrival rate by lowering the stock (h is increasing in Q), whereas in the authors&amp;rsquo; model the delay parameter α is constant and policy affects only damages.&lt;/p&gt;
&lt;p&gt;Q: What is the role of the exponential delay distribution assumption?
A: The assumption that the delay τ follows an exponential distribution with parameter α is made for tractability. Under this assumption, the entire past trajectory of the stock (Qt)t≤0 can be summarized by just two state variables — the highest stock on record Q0-bar and the initial legacy π0 — because the exponential &amp;ldquo;memoryless&amp;rdquo; property means that the additional expected waiting time until catastrophe occurrence does not depend on how long the triggering has already been in effect. Without this assumption, the full chronicle of past experiments would be required as a state variable, making the problem intractable.&lt;/p&gt;
&lt;p&gt;Q: What happens when the delay parameter α approaches zero or infinity?
A: When α → ∞ (instantaneous catastrophe upon triggering), pt = 1 − F(Qt) and the legacy is identically zero, recovering the Tsur-Zemel unknown-threshold approach (Proposition 3). The optimal path converges to QE0 from below or stabilizes if already above QE0. When α → 0 (infinite delay, effectively no catastrophe), QE = QD = QN and the problem reduces to the simple stock-flow problem (Proposition 1), with the optimal path converging monotonically to QN.&lt;/p&gt;
&lt;p&gt;Q: Does the model allow for damage mitigation after triggering but before occurrence?
A: Yes, this is a key feature. The continuation payoff after catastrophe occurrence is V(QT) where QT is the stock level at the time of occurrence T, not at triggering time T(S). This means the planner can reduce the stock after triggering to lower damages — analogous to a skater turning back toward shore after the ice first cracks. The assumption that V depends on the stock at occurrence rather than at triggering or at the maximum historical level is what allows this mitigation channel and is explicitly noted as a modeling choice.&lt;/p&gt;
&lt;p&gt;Legacy of the past (πt): The probability, conditional on survival to date t, that past experiments have already triggered a catastrophe. Formally πt = 1 − [1 − F(Qt)] / pt. Recent experiments contribute more to the legacy than distant ones, with contribution decaying at rate α. The legacy is zero when α → ∞ and is the central state variable bridging the paper&amp;rsquo;s two canonical extremes.&lt;/p&gt;
&lt;p&gt;QE (&amp;ldquo;Experimentation&amp;rdquo; threshold): The stock level at which the net marginal gain from further experimentation, defined as ν(Q) − [α/(α+δ)]ρ(Q)D(Q), equals zero, under the assumption that no catastrophe has been triggered. Below QE, stabilization is suboptimal; above QE, the planner does not experiment further when the legacy is zero.&lt;/p&gt;
&lt;p&gt;QD (&amp;ldquo;Damages&amp;rdquo; threshold): The stock level at which the net marginal benefit from holding the stock, defined as ν(Q) − [α/(α+δ)]D&amp;rsquo;(Q), equals zero, under the assumption that the catastrophe is known to have been triggered. QD ≤ QN and represents the optimal long-run target when the hazard-rate approach applies.&lt;/p&gt;
&lt;p&gt;Marginal payoff ν(Q): Defined as uq(0, Q) + (1/δ)uQ(0, Q), it measures the net gain from marginally increasing the flow when the stock is stabilized at Q. It is strictly decreasing in Q under Assumption 1 and equals zero at QN.&lt;/p&gt;
&lt;p&gt;Damage function D(Q): Defined as (1/δ)u(0, Q) − V(Q), it measures the welfare loss from catastrophe occurrence when the stock is Q at occurrence time, relative to permanent stabilization at Q. Assumed weakly positive and weakly increasing in Q.&lt;/p&gt;
&lt;p&gt;Survival probability (pt): The probability, computed from prior beliefs F at the beginning of times, that the catastrophe has not yet occurred by date t. Its law of motion is ṗt = α[1 − F(Qt) − pt], driven solely by the catastrophe parameter α and the current maximum stock Qt.&lt;/p&gt;
&lt;p&gt;Fatalism (under Theorem 1): The policy implication that a higher legacy — meaning a higher probability the catastrophe is already triggered — leads the planner to increase the stock further and accept more experimentation, because mitigation is relatively ineffective (QE &amp;lt; QD) and current consumption must be enjoyed before the catastrophe arrives.&lt;/p&gt;</description></item><item><title>Collusion with Optimal Information Disclosure</title><link>https://macropaperwarehouse.com/papers/collusion-with-optimal-information-disclosure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/collusion-with-optimal-information-disclosure/</guid><description>&lt;p&gt;This paper asks how a third-party intermediary (an &amp;ldquo;algorithm&amp;rdquo;) that observes market demand or costs superior to competing firms should optimally disclose that information to maximize the firms&amp;rsquo; collusive profit in a repeated Bertrand competition setting. The motivation is the rise of algorithmic pricing intermediaries such as RealPage in apartment rentals, A2i Systems in retail gasoline, and Rainmaker in hotel rooms, as well as offline cartel facilitators like AC-Treuhand.&lt;/p&gt;
&lt;p&gt;The model extends the canonical Rotemberg–Saloner (1986) repeated Bertrand framework with stochastic demand. The key technical assumption is that firm profit is affine in the unknown state s, so expected profit depends only on the expected state. This holds for binary states, linear demand with unknown intercept (D(p,s) = s − p), and linear demand with unknown per-unit cost. The algorithm observes s and commits to a known disclosure policy mapping s to a public signal. The solution concept is pure-strategy subgame-perfect equilibrium, and the paper solves for the disclosure policy and equilibrium that jointly maximize collusive profit.&lt;/p&gt;
&lt;p&gt;The main result (Theorem 1) is that the unique optimal disclosure policy is upper censorship: there is a cutoff ŝ such that demand states s &amp;lt; ŝ are disclosed and result in the corresponding monopoly price p^m(s), while demand states s ≥ ŝ are pooled — only the event {s ≥ ŝ} is disclosed — and result in the monopoly price for the mean concealed state, p^m(s*), where s* = E[s | s ≥ ŝ]. The reduction to a static information design problem (Lemma 1) is the key technical step: optimal collusive profit equals V*, the greatest fixed point of V = max_{G ∈ MPC(F)} E_G[min{π^m(s), δV/((1−δ)(n−1))}]. The &amp;ldquo;capped monopoly profit&amp;rdquo; min{π^m(s), π^max} is convex-then-concave in s, and classical results from the static information design literature (Kolotilin 2018; Dworczak and Martini 2019) then imply upper censorship is uniquely optimal.&lt;/p&gt;
&lt;p&gt;Two features of the optimal equilibrium are notable. First, prices are rigid (constant at p^m(s*)) whenever s ≥ ŝ — the opposite of Rotemberg–Saloner&amp;rsquo;s &amp;ldquo;price wars during booms.&amp;rdquo; The logic is that pooling high demand states with a lower average state is more profitable than cutting prices, because pooling reduces the current-period deviation gain without sacrificing as much on-path profit. Second, for demand states s ∈ (ŝ, s*), the equilibrium price p^m(s*) exceeds the monopoly price p^m(s) — supra-monopoly pricing occurs for a range of intermediate states. Monopoly pricing is attainable at each such state in isolation, but recommending the higher price p^m(s*) is necessary to make the pooling incentive-compatible at states s &amp;gt; s*.&lt;/p&gt;
&lt;p&gt;Comparing to full disclosure, Proposition 1 shows that optimal disclosure leads to strictly higher prices at every demand state, and hence unambiguously lower consumer surplus. Proposition 3 shows that improving the algorithm&amp;rsquo;s accuracy (a mean-preserving spread of F) reduces expected consumer surplus whenever consumer surplus under monopoly pricing is concave in s — a natural condition. This result is more pessimistic than prior work (Sugaya–Wolitzky 2018; Miklos-Thal–Tucker 2019), which found ambiguous effects because those papers assumed full disclosure.&lt;/p&gt;
&lt;p&gt;Comparative statics (Proposition 2): fewer firms or a higher discount factor δ increases collusive profit V* and makes prices more flexible (raises ŝ). Collusion is impossible if and only if δ &amp;lt; (n−1)/n, the same threshold as under full disclosure.&lt;/p&gt;
&lt;p&gt;Extensions maintain the core results. With Markov (persistent) demand (Section 4 / Theorem 2), upper censorship remains optimal but the cutoff ŝ(s) depends on last-period demand s: under positive serial correlation, ŝ(s) is decreasing in s, so the algorithm discloses less information following high demand. With differentiated products under a symmetric linear demand system (Section 5 / Theorem 3), the optimal policy censors an intermediate interval [ŝ_L, ŝ_H] and discloses both the lowest and highest demand states, because at high states the absence of an upper bound on equilibrium profit makes disclosure with price-cutting optimal.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question and why is it policy-relevant?
A: The paper asks how an informed intermediary should optimally disclose demand or cost information to competing firms to maximize their collusive profit. It is directly motivated by antitrust cases against RealPage (sued by the US DOJ in August 2024), A2i Systems/Kalibrate, and Rainmaker, all of which gather market data from competing firms and recommend prices. The theory also applies to offline facilitators like AC-Treuhand, prosecuted by the European Commission for disclosing competitively sensitive information.&lt;/p&gt;
&lt;p&gt;Q: What is the affinity assumption and why does it matter?
A: The paper assumes that firm profit π(p, s) is affine (linearly increasing) in the demand or cost state s for each price p. This implies that expected profit for any distribution over states equals profit evaluated at the expected state: E[π(p,s)] = π(p, E[s]). As a consequence, any disclosure policy is equivalent, from a profit standpoint, to choosing a distribution G of the firms&amp;rsquo; posterior mean beliefs over s, and G must be a mean-preserving contraction of the prior F (by Blackwell 1953). The assumption is satisfied for binary states, linear demand with unknown intercept, and linear demand with unknown cost.&lt;/p&gt;
&lt;p&gt;Q: What is the key reduction result (Lemma 1) and what does it achieve?
A: Lemma 1 reduces the problem of finding an optimal repeated-game equilibrium to a static information design problem. Optimal collusive profit equals V*, the greatest fixed point of V = max_{G ∈ MPC(F)} E_G[min{π^m(s), δV/((1−δ)(n−1))}], and this is attained by a symmetric, stationary, grim-trigger equilibrium. The reduction works because, under Bertrand competition, static deviation gains are proportional to on-path payoffs, creating a one-to-one correspondence that allows the repeated-game constraint to be folded into a single-period objective.&lt;/p&gt;
&lt;p&gt;Q: Why is upper censorship the uniquely optimal disclosure policy?
A: The static information design problem has a &amp;ldquo;capped monopoly profit&amp;rdquo; objective: min{π^m(s), π^max}, where π^max = δV*/((1−δ)(n−1)) is the maximum per-period profit that satisfies incentive constraints. Because π^m(s) is convex (as the maximum of affine functions) and the cap π^max is constant, the overall objective is convex for s below the cap and constant (then concave) above it — i.e., convex-then-concave in s. Classical results for linear information design (Kolotilin 2018; Dworczak and Martini 2019) imply that the unique optimal policy for a convex-then-concave objective is upper censorship.&lt;/p&gt;
&lt;p&gt;Q: What is the supra-monopoly pricing result and why does it arise?
A: For demand states s ∈ (ŝ, s*), the equilibrium price is p^m(s*) &amp;gt; p^m(s), meaning firms charge above the monopoly price for the current state. This arises because the pooling policy must recommend a single price for all states s ≥ ŝ, and the recommended price is p^m(s*) where s* = E[s | s ≥ ŝ]. At intermediate states s ∈ (ŝ, s*), this price exceeds the local monopoly price. The algorithm accepts lower profit at these states because it is necessary to maintain the pooled recommendation at higher states where monopoly pricing would otherwise require a price cut.&lt;/p&gt;
&lt;p&gt;Q: How does optimal disclosure compare to full disclosure in terms of consumer surplus?
A: Proposition 1 shows that collusive prices under optimal disclosure are strictly higher at every demand state compared to full disclosure (Rotemberg–Saloner). In Rotemberg–Saloner, high demand states trigger price cuts (&amp;ldquo;price wars during booms&amp;rdquo;) to deter deviation; under optimal disclosure, high states are pooled and prices are instead rigid at p^m(s*). Because prices are higher at all states, consumer surplus is unambiguously lower under optimal disclosure.&lt;/p&gt;
&lt;p&gt;Q: What does Proposition 3 say about the effect of algorithmic accuracy on consumer surplus?
A: Proposition 3 states that if consumer surplus under monopoly pricing, CS(s), is concave in s, then a mean-preserving spread of F (i.e., improved algorithmic accuracy) reduces expected consumer surplus. This result is more pessimistic than prior work by Sugaya–Wolitzky (2018) and Miklos-Thal–Tucker (2019), which found ambiguous effects. The difference is that those papers assumed full disclosure, so better accuracy tightened incentive constraints and sometimes forced price cuts. Under optimal selective disclosure, a more accurate algorithm always raises average prices because the algorithm withholds information that would have forced price cuts.&lt;/p&gt;
&lt;p&gt;Q: What are the comparative statics with respect to the number of firms and the discount factor?
A: Proposition 2 establishes that a decrease in the number of firms n or an increase in the discount factor δ increases collusive profit V* and makes collusive prices more flexible (raises ŝ). The intuition for fewer firms making prices more flexible is that with fewer firms, incentive constraints bind for a narrower range of demand states, so less pooling is needed. Collusion is impossible if and only if δ &amp;lt; (n−1)/n, the same threshold as under full disclosure.&lt;/p&gt;
&lt;p&gt;Q: How does the model generate empirically testable predictions distinct from other collusion models?
A: The model predicts: (1) the equilibrium price distribution has support on an interval [p^m(s_bar), p^m(ŝ)] plus a single mass point at the higher price p^m(s*); (2) prices are pro-cyclical overall but rigidly fixed at p^m(s*) for all but the lowest demand states; (3) the gap p^m(s) − p(s) is non-monotone — zero at low states, negative (supra-monopoly) at intermediate states, and positive at high states; (4) prices are more flexible when firms are more patient or fewer. The rigid high price combined with a flexible interval of lower prices is described as a distinctive collusive marker not present in other models.&lt;/p&gt;
&lt;p&gt;Q: How does the model relate to the empirical literature testing Green–Porter versus Rotemberg–Saloner?
A: Rotemberg–Saloner predicts counter-cyclical prices (price wars during booms), while Green–Porter predicts pro-cyclical prices. Empirical tests (e.g., Porter 1983, Ellison 1994) have typically found pro-cyclical prices, favoring Green–Porter. The present model generates pro-cyclical prices through a different mechanism — perfect monitoring plus selectively disclosed demand information — showing that pro-cyclical prices are consistent with perfect monitoring when the information intermediary optimally pools high demand states. The paper suggests that distinguishing the theories requires estimating the gap between price and monopoly price over the cycle: under Green–Porter, collusion succeeds better in high demand states; under this model, collusion succeeds better in low demand states.&lt;/p&gt;
&lt;p&gt;Q: What narrative evidence from the RealPage case corroborates the model&amp;rsquo;s predictions?
A: The US DOJ complaint against RealPage states that &amp;ldquo;in down markets… [RealPage] instills pricing discipline in landlords, curbing normal fully independent competitive reactions by substituting them with interdependent decision-making,&amp;rdquo; and that RealPage advertised that its AI helps clients &amp;ldquo;avoid the race to the bottom in down markets.&amp;rdquo; This is consistent with the model&amp;rsquo;s prediction of flexible monopoly prices at low demand states and a rigid, supra-monopolistic price in normal times. The Kumatori Contractors Cooperative case (studied by Kawai, Nakabayashi, and Ortner 2024) corroborates the censorship result: that organization took drastic steps to limit bidders&amp;rsquo; information about costs on the largest projects — exactly the states where deviation is most tempting.&lt;/p&gt;
&lt;p&gt;Q: How do results change with persistent (Markov) demand?
A: Theorem 2 shows that upper censorship remains uniquely optimal with Markov demand, but the cutoff ŝ(s) now depends on last-period demand s. Under positive serial correlation, ŝ(s) is decreasing in s: the algorithm discloses less information after high demand because firms are more optimistic and thus more tempted to deviate. Under negative serial correlation, ŝ(s) is increasing. The optimal collusive price is no longer always equal to the monopoly price for the disclosed mean demand, and the expected price conditional on last-period demand can be countercyclical (similar to Rotemberg–Saloner), even though the current-period price is always monotone in current demand.&lt;/p&gt;
&lt;p&gt;Q: How does the optimal disclosure policy change with differentiated products?
A: With a symmetric linear demand system (Section 5, Theorem 3), the optimal policy censors an intermediate interval [ŝ_L, ŝ_H] and discloses both the lowest and the highest demand states. At high demand states s &amp;gt; ŝ_H, the algorithm discloses the state and recommends a price below monopoly (to satisfy incentive constraints), because with differentiated goods there is no upper bound on equilibrium profit and profit is convex in s at high states, making disclosure with price-cutting optimal. Mathematically, the capped monopoly profit is piecewise-convex rather than convex-then-concave, so the optimal policy is intermediate-interval censorship rather than upper censorship. The Appendix A version extends to general demand systems and capacity constraints with the same qualitative logic.&lt;/p&gt;
&lt;p&gt;Q: What are the main limitations and directions for future work acknowledged by the authors?
A: The paper identifies three main limitations. First, if profit is not affine in s (i.e., expected profit depends on more than the mean state), the information design problem becomes non-linear and upper censorship is typically suboptimal, though it remains approximately optimal when the problem is close to linear. Second, the model assumes the algorithm&amp;rsquo;s objective is to maximize industry profit; if the intermediary is a profit-maximizing seller of software (as in Harrington 2022), the objective may instead be to maximize the profit differential between adopters and non-adopters. Third, the model assumes all firms use the algorithm; allowing partial adoption would require modeling firms&amp;rsquo; incentives to subscribe. The paper notes that incorporating these considerations &amp;ldquo;could be an interesting direction for future research.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Upper Censorship (disclosure policy): A disclosure policy in which demand states below a cutoff ŝ are revealed to firms (along with the corresponding monopoly price recommendation), while states above ŝ are pooled — only the event {s ≥ ŝ} is disclosed — with a single monopoly price recommendation p^m(s*) for the mean concealed state s* = E[s | s ≥ ŝ]. This is the uniquely optimal disclosure policy in the baseline model.&lt;/p&gt;
&lt;p&gt;Capped Monopoly Profit: The per-period profit objective in the reduced static information design problem: min{π^m(s), π^max}, where π^max = δV*/((1−δ)(n−1)) is the maximum industry profit attainable in a single period without violating incentive constraints. This function is convex-then-concave in s, which drives the optimality of upper censorship.&lt;/p&gt;
&lt;p&gt;Supra-Monopoly Pricing: Equilibrium prices that exceed the monopoly price for the realized demand state. In the model, this occurs for states s ∈ (ŝ, s*), where the algorithm&amp;rsquo;s pooled recommendation p^m(s*) is above the local monopoly price p^m(s). It arises because the pooled recommendation must be incentive-compatible at the highest concealed states.&lt;/p&gt;
&lt;p&gt;Price Rigidity: The feature of the optimal equilibrium in which the collusive price is constant at p^m(s*) for all demand states s ≥ ŝ. The algorithm achieves this by withholding information about high demand states, preventing the &amp;ldquo;price wars during booms&amp;rdquo; predicted by Rotemberg–Saloner (1986) under full disclosure.&lt;/p&gt;
&lt;p&gt;Algorithmic Accuracy: In the paper&amp;rsquo;s terms, the informativeness of the algorithm&amp;rsquo;s signal about s, formalized as the precision of the distribution F. Improving accuracy corresponds to a mean-preserving spread of F (Blackwell 1953). A more accurate algorithm always increases collusive profit; under the concavity condition on consumer surplus, it also reduces expected consumer surplus.&lt;/p&gt;
&lt;p&gt;Mean-Preserving Contraction (MPC(F)): The set of distributions G of firms&amp;rsquo; posterior mean beliefs over s that are consistent with Bayesian updating of the prior F. By Blackwell (1953), a disclosure policy is feasible if and only if it induces a distribution G ∈ MPC(F). This is the feasibility constraint in the static information design problem.&lt;/p&gt;
&lt;p&gt;Affinity in the state: The assumption that π(p, s) is affine (linearly increasing) in s for each price p. This implies E[π(p,s)] = π(p, E[s]), so expected profit is determined entirely by the expected state, enabling the reduction of the disclosure problem to choosing a distribution of posterior means.&lt;/p&gt;</description></item><item><title>Community Engagement and Public Safety: Evidence from Crime Enforcement Targeting Immigrants</title><link>https://macropaperwarehouse.com/papers/community-engagement-and-public-safety-evidence-from-crime-enforcement-targeting-immigrants/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/community-engagement-and-public-safety-evidence-from-crime-enforcement-targeting-immigrants/</guid><description>&lt;p&gt;This paper studies how immigration enforcement affects public safety, asking two questions: (1) what is the effect of increased enforcement on criminal victimization, and (2) how does increased enforcement affect victims&amp;rsquo; willingness to report crimes to police? The authors exploit the staggered rollout of the U.S. Secure Communities (SC) program — the largest expansion of interior immigration enforcement in U.S. history — across counties between 2008 and 2013. SC expanded information sharing between local police and federal immigration authorities, causing ICE honored detainer requests to increase by over 50% following program activation.&lt;/p&gt;
&lt;p&gt;The primary data source is the restricted-access National Crime Victimization Survey (NCVS), which measures victimizations independently of whether they were reported to police and includes respondent ethnicity. This allows the authors to separately estimate effects on underlying crime incidence and on reporting behavior for Hispanic and non-Hispanic individuals. The empirical strategy uses a staggered difference-in-differences design following Sun and Abraham (2021), comparing earlier-treated counties to the last 25% of counties to activate SC, with estimates run separately by ethnicity.&lt;/p&gt;
&lt;p&gt;The main findings run contrary to the stated policy goal of improving public safety. Among Hispanic individuals, SC caused a statistically significant 0.15 percentage point increase in monthly victimization — a 16% increase relative to the pre-period baseline of 0.9 percentage points — implying approximately 1.3 million additional crimes against Hispanics in the two years following program activation. The increase is concentrated primarily in property crimes (a statistically significant 15% increase), with a similarly sized but imprecisely estimated 15% increase in violent crime victimizations. The victimization increase is larger for Hispanic females (0.23 percentage points, or 25%) and in counties with higher shares of non-citizen Hispanic residents.&lt;/p&gt;
&lt;p&gt;Simultaneously, SC caused a 9.5 percentage point decline in the likelihood that Hispanic victims report incidents to police — a 30% decline relative to the pre-period mean reporting rate of 33 percentage points. This reporting decline is primarily driven by a 34% decline in the reporting of property offenses. No changes in victimization or reporting are found for non-Hispanic individuals in the aggregate, though non-Hispanic individuals in neighborhoods with high Hispanic population shares do experience higher victimization rates after SC.&lt;/p&gt;
&lt;p&gt;Critically, reported crime rates (the product of victimization and reporting) are unchanged for both Hispanic and non-Hispanic individuals, explaining why prior studies using administrative reported-crime data found null effects of SC. The null effect on reported crime masks two large, opposing causal forces.&lt;/p&gt;
&lt;p&gt;The authors provide evidence that the decline in crime reporting is the primary driver of the increase in victimization. Cohorts with larger reporting declines experienced larger victimization increases, and a decomposition exercise shows the reporting decline is substantially more important than concurrent SC-induced changes in unemployment, wages, female-headed household shares, and the male immigrant share. Supporting data from 75 police departments confirm no change in 911 call volumes or total arrest volumes, while showing a decline in the Hispanic share of arrestees in both Hispanic and non-Hispanic neighborhoods — consistent with reduced reporting leading to reduced apprehension of offenders, with offending shifting toward non-Hispanic individuals.&lt;/p&gt;
&lt;p&gt;Scope conditions: results are estimated for the population residing in counties exceeding 100,000 residents (representing 61% of total U.S. population and 69% of the Hispanic population), excluding southern border counties and states that actively resisted SC implementation (Illinois, Massachusetts, New York). Effects apply to all Hispanic respondents — citizens and non-citizens — consistent with prior evidence that citizen Hispanics respond to immigration enforcement out of concern for non-citizen contacts.&lt;/p&gt;
&lt;p&gt;Q: What was the Secure Communities program and how was it implemented?
A: SC was a federal program launched in 2008 that required fingerprints of individuals booked into local jails to be forwarded not only to the FBI but also to the Department of Homeland Security, enabling automatic screening for immigration violations. Local authorities could not prevent federal officials from learning of an arrestee&amp;rsquo;s immigration status. The program rolled out county-by-county between October 2008 and January 2013 due to technological constraints and resource bottlenecks, generating the staggered variation used for identification.&lt;/p&gt;
&lt;p&gt;Q: How large was the first-stage effect on actual immigration enforcement?
A: County-level honored ICE detainer requests increased by over 50% following SC activation, with a similar 40% increase in all detainer requests. The number of honored detainers nationwide doubled between 2008 and 2012. Over 90% of detainers and removals in any given month were for individuals of Hispanic ethnicity.&lt;/p&gt;
&lt;p&gt;Q: What is the main finding on Hispanic victimization?
A: SC caused a 0.15 percentage point increase in monthly Hispanic victimization rates, a 16% increase relative to the pre-period baseline of 0.9 percentage points. This translates to approximately 1.3 million additional crimes against Hispanics over two years following program activation, calculated by multiplying the monthly effect by 24 months and the 35.3 million Hispanics in the sample counties.&lt;/p&gt;
&lt;p&gt;Q: What is the main finding on Hispanic crime reporting?
A: SC caused a 9.5 percentage point decline in the likelihood that Hispanic victims report incidents to police, a 30% decline relative to the pre-period mean reporting rate of 33 percentage points. This decline occurred relatively quickly after activation and was concentrated in property offenses, where reporting fell by 34%.&lt;/p&gt;
&lt;p&gt;Q: Why do reported crime rates show no change despite large shifts in victimization and reporting?
A: Reported crime rates — the probability of being victimized and reporting the crime — are unchanged because the 16% increase in victimization and the 30% decline in reporting are approximately offsetting in magnitude. This explains why prior work using administrative police data (Miles and Cox 2014; Treyger et al. 2014; Hines and Peri 2019) found null effects of SC on reported crime: those data sources cannot separately identify the two underlying changes.&lt;/p&gt;
&lt;p&gt;Q: Does SC affect non-Hispanic individuals?
A: In the aggregate, SC has no statistically significant effect on non-Hispanic victimization or reporting. However, non-Hispanic individuals living in neighborhoods with high Hispanic population shares do experience victimization increases, and in those neighborhoods their reporting rates also decline slightly. Re-weighting non-Hispanic respondents to match the county composition of Hispanic respondents yields an 8% increase in non-Hispanic victimization, suggesting spillover effects in Hispanic-dense areas.&lt;/p&gt;
&lt;p&gt;Q: What mechanism links the reporting decline to the victimization increase?
A: The authors argue that reduced victim reporting lowers the probability that offenders are apprehended, thereby reducing the cost of committing crimes. They demonstrate this through two analyses: first, cohorts of counties with larger reporting declines experienced larger victimization increases; second, a decomposition shows the reporting channel is substantially more important than concurrent SC-induced changes in unemployment, wages, female-headed household shares, and the male immigrant share of the population.&lt;/p&gt;
&lt;p&gt;Q: What do the police administrative data show about offender composition?
A: Data from 75 police departments show no change in 911 call volumes or total arrest volumes following SC — consistent with the NCVS finding of unchanged reported crime rates. However, the Hispanic share of arrestees declined after SC, with a 1.5 percentage point drop in Hispanic neighborhoods (off a base of 54%), suggesting the rise in offending was more concentrated among non-Hispanic offenders as reduced reporting lowered expected punishment probabilities.&lt;/p&gt;
&lt;p&gt;Q: How does the victimization effect vary by gender?
A: The victimization point estimate for Hispanic males is 0.085 percentage points and imprecisely estimated (SE = 0.088). For Hispanic females, the effect is over 2.5 times larger at 0.23 percentage points, a 25% increase. The decline in reporting is comparable in magnitude across male and female Hispanic victims, suggesting fear of enforcement is similar by gender but that females disproportionately bear the crime burden.&lt;/p&gt;
&lt;p&gt;Q: How does the victimization effect vary by neighborhood non-citizen Hispanic share?
A: Victimization effects for Hispanics are relatively constant across neighborhood types but are higher — around 25% — in neighborhoods with the highest shares of non-citizen Hispanics. Counties with higher non-citizen Hispanic shares also exhibit higher ICE removal rates, indicating greater total enforcement, and these counties have higher victimization effects. Reporting declines among Hispanics appear relatively uniform across neighborhood types.&lt;/p&gt;
&lt;p&gt;Q: Could survey attrition or compositional changes explain the results?
A: The authors rule this out through several tests. First, SC has no statistically significant effect on household survey response rates, even in Census tracts above the 90th percentile of Hispanic share. A worst-case bias calculation implies attrition could account for at most 26% of the victimization effect. Second, re-estimating using predicted victimization (based on pre-SC demographics) yields precise null effects, indicating the increase is not driven by compositional change. Third, results are stable when restricting to respondents present at all survey waves or using individual fixed effects.&lt;/p&gt;
&lt;p&gt;Q: Could the reporting decline be mechanical — reflecting a change in the types of crimes committed rather than behavioral change?
A: The authors test this by constructing predicted reporting rates using pre-SC incident characteristics. The largest alternative estimate is -1.45 percentage points, over six times smaller than the estimated main reporting effect of 9.5 percentage points, ruling out crime composition change as the primary explanation. Results also hold when focusing on always-respondents and using individual fixed effects, ruling out entry of low-reporting individuals into the survey.&lt;/p&gt;
&lt;p&gt;Q: How robust are the results to alternative empirical strategies?
A: Results are robust to including states that resisted SC (with somewhat smaller magnitudes as expected), alternative population cutoffs, TWFE specifications, the Borusyak et al. (2021) and Callaway and Sant&amp;rsquo;Anna (2021) estimators (which yield larger point estimates), a triple-differences specification using non-Hispanics as an additional control group, and the inclusion of time-varying unemployment rates. The dynamic event-study plots show parallel pre-trends across all specifications.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the null effect on aggregate victimization?
A: The authors estimate that the policy ruled out declines in aggregate victimization larger than 3.3%, indicating SC did not generate meaningful improvements in aggregate public safety. This contradicts the stated mission of immigration enforcement agencies. The findings imply that policies targeting immigrant communities can generate public safety costs through trust erosion that outweigh any deterrence or incapacitation benefits.&lt;/p&gt;
&lt;p&gt;Secure Communities (SC): A federal program launched in 2008 requiring automatic sharing of fingerprints from local jail bookings with the Department of Homeland Security, enabling identification of unauthorized immigrants among local arrestees and triggering ICE detainer requests; the largest expansion of interior immigration enforcement in U.S. history.&lt;/p&gt;
&lt;p&gt;Chilling effect: The mechanism by which immigration enforcement raises the perceived cost of contacting law enforcement for immigrant victims and witnesses — through fear that they, a family member, or neighbor will be detained or deported — thereby reducing willingness to report crimes independently of any change in underlying criminality.&lt;/p&gt;
&lt;p&gt;Victimization rate: The likelihood that an individual is the victim of a crime in a given period, measured via the NCVS independently of whether the crime was reported to police; the paper&amp;rsquo;s primary measure of public safety.&lt;/p&gt;
&lt;p&gt;Reporting rate: The likelihood that a criminal victimization is reported by the victim to the police, measured as a share of all crime incidents; distinct from victimization rate and central to the paper&amp;rsquo;s decomposition of reported crime into its two components.&lt;/p&gt;
&lt;p&gt;Reported crime rate: The joint probability of being victimized and reporting the crime, analogous to measures available in administrative police data such as the FBI UCR; this outcome masks the opposing effects of SC on victimization and reporting.&lt;/p&gt;
&lt;p&gt;Honored detainer: An ICE detainer request that results in a transfer of the arrested individual to ICE custody; the paper&amp;rsquo;s preferred measure of immigration enforcement intensity because it is available both before and after SC activation and is more directly linked to deportation actions than all detainer requests.&lt;/p&gt;
&lt;p&gt;Decomposition of victimization increase: The paper&amp;rsquo;s procedure for quantifying the relative importance of the reporting-channel (reduced probability of apprehension) versus other SC-induced social and economic changes (unemployment, wages, female-headed households, male immigrant share) in explaining the rise in Hispanic victimization.&lt;/p&gt;</description></item><item><title>Competing under Information Heterogeneity: Evidence from Auto Insurance</title><link>https://macropaperwarehouse.com/papers/competing-under-information-heterogeneity-evidence-from-auto-insurance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/competing-under-information-heterogeneity-evidence-from-auto-insurance/</guid><description>&lt;p&gt;This paper studies imperfect competition in selection markets where competing firms have heterogeneous information about consumers — a layer of asymmetry distinct from the classic buyer-seller information gap. The central questions are: how do inter-firm information asymmetries shape equilibrium pricing, consumer sorting, and market efficiency; and whether a centralized bureau that aggregates and equalizes firms&amp;rsquo; risk information can promote competition and improve welfare.&lt;/p&gt;
&lt;p&gt;The empirical setting is the Italian mandatory motor vehicle liability insurance market (Responsabilità Civile Auto). The authors use the IPER dataset from IVASS, a nationally representative panel of matched insurer-insuree contracts covering 124,428 liability insurance contracts for new customers in the province of Rome from 2013 to 2021. The panel tracks consumers across insurer switches, enabling construction of individual-specific risk estimates from ex-post claim records using Poisson regressions for claim frequency and log-normal regressions for claim severity. The analysis focuses on the top 10 largest firms plus a composite fringe firm.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s empirical strategy proceeds in three stages. First, individual risk types are estimated from multi-year claim panels. Second, demand parameters — price sensitivity and firm-level unobserved product attributes — are recovered using a novel fixed-point algorithm (extending Berry et al. 1995) that infers the full offered-price distribution from observed transaction prices alone, without parametric restrictions on price distributions across firms. Third, supply-side parameters — pricing coefficients, signal variances, and cost parameters — are identified by exploiting the monotone mapping between offered prices and private signals, borrowing from the nonparametric auction literature.&lt;/p&gt;
&lt;p&gt;The model features firms that each draw a private Gaussian signal about a consumer&amp;rsquo;s true risk type theta, with firm-specific signal standard deviation sigma_j. Lower sigma_j means higher information precision. Firms set prices as a linear function of their posterior risk rating: p_j = alpha_j + beta_j * E(theta | theta_j, D=j). Firms simultaneously choose pricing coefficients to maximize expected profits.&lt;/p&gt;
&lt;p&gt;Key empirical findings: (1) Firms differ substantially in how sensitively their premiums respond to realized consumer risk — a reduced-form measure of information precision — with Figure 2 showing wide cross-firm variation in premium-to-risk coefficients. (2) Structural estimation confirms substantial heterogeneity in signal standard deviations sigma_j across all 11 firms. Firms with less accurate risk-rating algorithms (higher sigma_j) tend to have more efficient cost structures (lower claim-processing cost parameter k_j), generating distinct comparative advantages. (3) Baseline pricing coefficients alpha_j and risk-sensitivity coefficients beta_j vary dramatically across firms. (4) Senior drivers are less price sensitive; urban drivers are more price sensitive. Lower-risk consumers show stronger preferences for Firms 3 and 5, while higher-risk consumers disproportionately choose Firm 8.&lt;/p&gt;
&lt;p&gt;Counterfactual simulations assess three information policies relative to the baseline. Under a centralized risk bureau — which collects each firm&amp;rsquo;s signal, aggregates them weighted by precision, and distributes the combined signal equally — average premiums fall by 21.6% and consumer surplus rises by 15.7%. The efficiency benchmark (firms observe true risk perfectly) yields a 25.7% premium reduction and a 16.9% consumer surplus gain, so the bureau recovers almost all the efficiency gap. The privacy benchmark (all firms restricted to the coarsest signal in the market) raises surplus for high-risk consumers by 6.9% but harms low-risk consumers.&lt;/p&gt;
&lt;p&gt;The bureau&amp;rsquo;s price reduction operates through two channels: it eliminates the market power that accrues to firms with superior private information, and it aligns firms&amp;rsquo; risk evaluations, enabling sharper undercutting. The bureau also reduces average costs by 12 euros per contract by enabling more efficient insurer-insuree matching — cost-efficient claim processors can better target the consumer types they have a comparative advantage in serving.&lt;/p&gt;
&lt;p&gt;The analysis is confined to new customers in Rome&amp;rsquo;s provincial market to avoid complications from dynamic pricing and consumer-firm learning. The model abstracts away from optional contract clauses (treated as observable characteristics) and does not model the specific mechanisms generating information heterogeneity.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s core research question?
A: The paper asks how information asymmetries between competing firms (not just between buyers and sellers) shape equilibrium pricing strategies, consumer sorting, and market efficiency in a selection market, and whether a centralized bureau that equalizes firms&amp;rsquo; access to aggregated risk information can improve competition and welfare. This extends the classic Akerlof-Rothschild-Stiglitz framework by introducing a second layer of asymmetry — across sellers themselves.&lt;/p&gt;
&lt;p&gt;Q: Why is the Italian auto insurance market well suited for this study?
A: Italy mandates liability insurance for all drivers and prohibits rejections, so the analysis focuses entirely on how consumers sort across insurers rather than on participation margins. The IPER dataset from IVASS is a nationally representative panel tracking policyholders even across insurer switches, providing both premium and ex-post claim records needed to construct individual risk types. The market has roughly 50 competing firms using demonstrably heterogeneous pricing algorithms, documented through a survey of major insurers and reduced-form regressions.&lt;/p&gt;
&lt;p&gt;Q: How do the authors measure firm-level information precision in the reduced-form analysis?
A: They estimate individual-specific risk types from a panel of claim records using Poisson regressions (claim frequency) and log-normal regressions (claim severity), then regress each firm&amp;rsquo;s premiums on those estimated risk measures. Firms whose premiums respond more sensitively to realized risk are inferred to have higher information precision. Figure 2 shows that these premium-to-risk coefficients vary significantly across firms — for example, Firm 7&amp;rsquo;s premiums are considerably more sensitive to risk than Firm 8&amp;rsquo;s — providing reduced-form evidence of heterogeneous information precision before any structural estimation.&lt;/p&gt;
&lt;p&gt;Q: What is the structural model&amp;rsquo;s signal structure?
A: Each firm j draws a private signal theta_j ~ N(theta, sigma_j^2) about a consumer&amp;rsquo;s true risk type theta, where sigma_j is the firm-specific signal standard deviation. A smaller sigma_j means higher precision. Signals are independent across firms conditional on theta, analogous to common-value auctions where firms receive noisy estimates of a shared unknown value (expected claim payouts). The parameter sigma_j is the key structural object the paper identifies and estimates.&lt;/p&gt;
&lt;p&gt;Q: What is novel about the demand estimation strategy?
A: Standard demand estimation assumes the same price is offered to all consumers or that the full price menu is observed. Here, only transaction prices are observed — the prices of unchosen insurers are not in the data. The authors apply the Wu and Xin (2024) fixed-point algorithm, which jointly estimates consumers&amp;rsquo; sorting probabilities, offered price distributions, and demand parameters by adding an outer loop over sorting propensities to the Berry (1994) contraction mapping. No parametric restrictions are imposed on the offered price distributions, and they are allowed to vary fully across firms.&lt;/p&gt;
&lt;p&gt;Q: How are firms&amp;rsquo; signal variances identified separately from pricing coefficients?
A: There is a one-to-one mapping between a firm&amp;rsquo;s offered price and its signal (prices increase monotonically in the signal, analogous to bids in auctions). After recovering the offered price distribution from the demand step, the authors observe price dispersion at a fixed risk level. By focusing on average prices conditional on each risk level, signal noise averages out, identifying the pricing coefficients beta_j. The residual price dispersion at fixed risk then identifies signal variance sigma_j^2.&lt;/p&gt;
&lt;p&gt;Q: What does structural estimation reveal about the relationship between information precision and cost efficiency?
A: Firms with higher signal standard deviations (less precise risk evaluation) tend to have lower claim-processing cost parameters k_j — they are more efficient at handling claims. This creates distinct comparative advantages: some firms excel at risk identification but face higher processing costs, while others process claims cheaply but evaluate risk less precisely. This heterogeneity means information-equalizing policies have differentiated firm-level impacts.&lt;/p&gt;
&lt;p&gt;Q: What are the quantitative effects of the centralized risk bureau on premiums and consumer surplus?
A: The bureau reduces average premiums by 21.6% relative to baseline and increases consumer surplus by 15.7%. The efficiency benchmark — where firms observe consumers&amp;rsquo; true risk perfectly — produces a 25.7% premium reduction and a 16.9% consumer surplus gain. The bureau therefore closes nearly all of the gap to the first-best allocation in surplus terms (15.7% vs. 16.9%).&lt;/p&gt;
&lt;p&gt;Q: Through what mechanisms does the bureau reduce prices?
A: Two distinct channels are identified. First, equalizing information precision eliminates the informational market power held by firms with superior signals, compelling them to compete more aggressively on price. Second, when all firms share the same risk evaluation of a consumer, they can undercut each other more precisely, which intensifies price competition further. Both channels operate simultaneously under the bureau.&lt;/p&gt;
&lt;p&gt;Q: How does the bureau affect consumer surplus distribution across risk types?
A: The bureau primarily benefits low-risk consumers because improved information allows firms to price discriminate more accurately on risk type, lowering prices for those who are low risk. High-risk consumers see smaller benefits and may face relatively higher premiums. This contrasts with the privacy benchmark, where restricting all firms to the coarsest signal in the market raises high-risk consumers&amp;rsquo; surplus by 6.9% — because it becomes harder for firms to distinguish them from low-risk consumers.&lt;/p&gt;
&lt;p&gt;Q: What is the cost efficiency effect of the bureau?
A: Under the centralized risk bureau, average costs per contract fall by 12 euros. This reflects more efficient insurer-insuree matching: when firms have equal and better information, those with cost advantages in claims processing can better identify and attract the consumer types they are relatively best equipped to serve. The authors note that given the scale of the Italian auto insurance market (approximately 31 million contracts annually), this per-contract saving implies a substantial aggregate impact.&lt;/p&gt;
&lt;p&gt;Q: What happens to firm profits under the bureau, and is the impact uniform?
A: Average profits decline overall due to lower prices. However, the impact is heterogeneous across firms. Firms that rely most heavily on superior information precision — often smaller, more specialized firms — experience greater profit losses, since the bureau most directly erodes their competitive advantage.&lt;/p&gt;
&lt;p&gt;Q: How does the privacy benchmark differ from the bureau scenario?
A: The privacy benchmark simulates a regulation that restricts all firms to using only basic consumer information, setting signal variance to the highest level observed in the market. Unlike the bureau (which improves and equalizes information), this benchmark degrades information uniformly. It produces opposite distributional effects: high-risk consumers gain 6.9% in surplus as cross-subsidization from low-risk to high-risk consumers increases, while low-risk consumers are worse off.&lt;/p&gt;
&lt;p&gt;Q: Why does the paper focus on new customers only?
A: Focusing on new customers avoids complications from dynamic pricing, where insurers update premiums based on accumulated claim history with a specific consumer, and from consumer-firm learning dynamics. This follows standard practice in the empirical asymmetric information literature, as cited in Chiappori and Salanie (2000) and Crawford et al. (2018).&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to and extend prior work on selection markets?
A: Prior empirical work on imperfect competition in selection markets — including Einav et al. (2010), Crawford et al. (2018), and related studies — assumes that competing firms have symmetric information about consumers. This paper is described as introducing the first tractable empirical framework for analyzing selection markets where firms have heterogeneous information. It also incorporates multidimensional cost heterogeneity on the supply side, adding to work by Salanié (2017) and Nelson (2025).&lt;/p&gt;
&lt;p&gt;Q: What do the reduced-form regressions reveal about pricing heterogeneity across insurers?
A: Firm-level regressions of premiums on observable risk factors show R-squared values ranging from 0.39 to 0.59. Estimated coefficients on key risk factors vary dramatically: being one year older reduces premiums by 0.25 to 1.68 euros depending on the firm; a higher bonus-malus class increases premiums by 12 to 32 euros; one additional accident in the previous five years raises premiums by 74 to 181 euros. These ranges reflect genuine differences in actuarial algorithms, not just sampling variation.&lt;/p&gt;
&lt;p&gt;Q: What is the bonus-malus system and why does its saturation matter for the paper&amp;rsquo;s setting?
A: Italy&amp;rsquo;s bonus-malus (BM) system assigns drivers to one of 18 risk classes based on accident history. Because approximately 80% of policyholders are in the best class (BM class 1), the public BM system provides limited granularity for risk evaluation. This saturation creates strong incentives for firms to develop proprietary risk-rating algorithms, which is the institutional basis for the substantial information heterogeneity that the paper documents and models.&lt;/p&gt;
&lt;p&gt;Information Precision (sigma_j): In the paper&amp;rsquo;s model, the firm-specific parameter measuring the dispersion of a firm&amp;rsquo;s private signal about a consumer&amp;rsquo;s true risk type. Firm j draws signal theta_j ~ N(theta, sigma_j^2); 1/sigma_j is information precision. A smaller sigma_j means the firm more accurately identifies consumer risk. This is not merely a theoretical construct — the paper identifies and estimates sigma_j structurally for each of the 11 firms.&lt;/p&gt;
&lt;p&gt;Heterogeneous Information: The condition where competing firms hold signals of different precision about the same consumer&amp;rsquo;s unobserved risk type, introducing asymmetry not just between buyers and sellers (as in Akerlof 1970) but among sellers themselves. This is the paper&amp;rsquo;s central departure from prior literature on selection markets, which assumed symmetric information among firms.&lt;/p&gt;
&lt;p&gt;Centralized Risk Bureau: A policy institution that collects each firm&amp;rsquo;s analyzed risk signal, aggregates them weighted by each firm&amp;rsquo;s information precision (producing a combined signal more precise than any individual firm&amp;rsquo;s signal), and makes the aggregated information equally accessible to all firms. The bureau is the paper&amp;rsquo;s primary policy counterfactual, and it is modeled as equalizing both the level and heterogeneity of information precision across competitors.&lt;/p&gt;
&lt;p&gt;Offered vs. Accepted Price Distribution: A distinction central to the paper&amp;rsquo;s identification strategy. The accepted price distribution is what is observed in transaction data — prices conditional on the consumer having chosen that firm. The offered price distribution is the full set of prices the firm would charge across all consumers, including those who did not select it. The paper recovers the offered distribution from the accepted distribution using a fixed-point algorithm, without imposing parametric restrictions.&lt;/p&gt;
&lt;p&gt;Selection Loop: The paper&amp;rsquo;s methodological extension of the Berry (1994) BLP contraction mapping for mean utilities. An outer loop iterates over consumers&amp;rsquo; sorting propensities to jointly recover offered price distributions, sorting probabilities, and demand parameters when only transaction prices are observed. This technique handles the endogeneity of which prices are accepted.&lt;/p&gt;
&lt;p&gt;Risk Rating: The firm&amp;rsquo;s posterior assessment of a consumer&amp;rsquo;s expected cost, computed as the posterior mean E(theta | theta_j, D=j) — the expected true risk type conditional on the firm&amp;rsquo;s private signal and the consumer selecting that firm. Firms set prices as a linear function of their risk rating: p_j = alpha_j + beta_j * E(theta | theta_j, D=j).&lt;/p&gt;
&lt;p&gt;Comparative Advantage (information vs. cost): The paper&amp;rsquo;s finding that firms with lower information precision (higher sigma_j) tend to have more efficient cost structures (lower k_j), and vice versa. This cross-sectional negative correlation between information advantage and cost advantage means that policy interventions that equalize information precision shift the basis of competition from information asymmetry to cost specialization.&lt;/p&gt;</description></item><item><title>Contextually Private Mechanisms</title><link>https://macropaperwarehouse.com/papers/contextually-private-mechanisms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/contextually-private-mechanisms/</guid><description>&lt;p&gt;Haupt and Hitzig introduce a framework for comparing the privacy properties of different mechanism protocols. The core research question is: when a designer commits to implementing a social choice rule, how much superfluous private information must they inevitably learn about agents, and how should they design the elicitation protocol to minimize that exposure?&lt;/p&gt;
&lt;p&gt;The setting is a finite-player extensive-form game in which a designer elicits agents&amp;rsquo; private types through a dynamic protocol to compute a social choice function. The authors explicitly exclude cryptographic tools and trusted mediators, working under the minimal assumption that the designer learns information if and only if an agent discloses it. This assumption is motivated by the historical prevalence of live dynamic auction formats — ascending formats at Sotheby&amp;rsquo;s, descending formats at Aalsmeer, oral ascending formats used by the U.S. Forest Service for timber, multi-round clock auctions for radio-spectrum allocation — and by settings where mediating technology is unavailable or costly.&lt;/p&gt;
&lt;p&gt;The central object is the contextual privacy violation. A protocol produces a contextual privacy violation for agent i at type profile θ if the designer can distinguish θ_i from some alternative type θ&amp;rsquo;_i while holding other agents&amp;rsquo; types fixed, yet the social choice rule assigns the same outcome at both profiles. Violations are defined at the level of individual agent–state pairs, not aggregated ex ante. A protocol is fully contextually private if it produces no violations; it is maximally contextually private if its set of violations is inclusion-minimal among all protocols that implement the same rule.&lt;/p&gt;
&lt;p&gt;The main characterization result (Theorem 1) connects privacy to pivotality: a social choice function admits a fully contextually private protocol if and only if, on every product subset of the type space where agents are collectively pivotal, at least one agent is individually pivotal. The contrapositive is what drives the paper&amp;rsquo;s impossibility results: whenever a rule contains a region where no single agent&amp;rsquo;s report changes the outcome but a group&amp;rsquo;s joint report does, any implementing protocol must produce contextual privacy violations.&lt;/p&gt;
&lt;p&gt;Using this characterization, the authors establish that the first-price auction rule (Proposition 2) and serial dictatorship (Proposition 3) admit fully contextually private protocols. Conversely, k-item Vickrey auction rules (Proposition 4) and any stable school-choice rule (Proposition 5) do not admit fully contextually private protocols, because these rules contain type-space regions where agents are only collectively — not individually — pivotal.&lt;/p&gt;
&lt;p&gt;For k-item Vickrey auctions, the authors study maximally contextually private protocols. They establish (Proposition 6) that, for a class of social choice rules on totally ordered type spaces that contains k-item Vickrey auctions, it is without loss to consider only protocols consisting of threshold queries that are monotonically increasing or decreasing after an initial guess. This reduction identifies two key design dimensions: the initial query posed to each agent, and the order in which agents are queried.&lt;/p&gt;
&lt;p&gt;The main constructive result (Theorem 2) proves that an ascending-join protocol is maximally contextually private for the k-item Vickrey auction. Proposition 7 formalizes the sense in which this protocol protects privacy by delaying queries to certain bidders — it repeatedly asks agents whether they can rule out a particular outcome, and postpones questioning agents whose privacy it is protecting.&lt;/p&gt;
&lt;p&gt;The authors also show (Proposition 19) that the ascending-join protocol is minimally relatively informative among protocols that are maximally contextually private. Extensions cover group contextual privacy (Proposition 11) and individual contextual privacy (Proposition 8), showing that individual contextual privacy violations equal the union of contextual privacy violations and nonbossiness violations.&lt;/p&gt;
&lt;p&gt;Q: What is a contextual privacy violation, precisely?
A: A protocol produces a contextual privacy violation for agent i at type profile θ if the designer can distinguish θ_i from some alternative type θ&amp;rsquo;_i — holding all other agents&amp;rsquo; types fixed — yet the social choice rule assigns the same outcome at both profiles. The violation is defined at the level of individual agent–state pairs. A single additional superfluous distinction at the same (i, θ) pair does not register as a second violation; the framework records whether any unnecessary disclosure occurs for that agent at that state, not the degree of overexposure.&lt;/p&gt;
&lt;p&gt;Q: How does contextual privacy differ from relative informativeness?
A: Relative informativeness compares two protocols by whether one distinguishes every pair of type profiles the other does, treating all disclosures as equally undesirable. Contextual privacy conditions the notion of a &amp;ldquo;violation&amp;rdquo; on the social choice rule: a distinction between θ_i and θ&amp;rsquo;_i counts as a violation only when the rule assigns the same outcome at both profiles. Relative informativeness thus penalizes the designer for learning information that is necessary to implement the rule, whereas contextual privacy imposes no penalty for learning pivotal information.&lt;/p&gt;
&lt;p&gt;Q: What is the pivotality characterization (Theorem 1)?
A: A social choice function admits a fully contextually private protocol if and only if, on every product subset of the type space where agents are collectively pivotal, at least one agent is individually pivotal. The necessity direction shows that if a collectively pivotal set exists where no agent is individually pivotal, any implementing iterative partition must contain an earliest node that distinguishes two type profiles leading to the same outcome. The sufficiency direction constructs a contextually private protocol inductively by always querying an individually pivotal agent, ensuring every distinction implies a different outcome.&lt;/p&gt;
&lt;p&gt;Q: Which social choice rules admit fully contextually private protocols?
A: The first-price auction rule (Proposition 2) and serial dictatorship (Proposition 3) admit fully contextually private protocols. The authors use Theorem 1 to show this: in both rules, any collectively pivotal region contains an individually pivotal agent. By contrast, k-item Vickrey auction rules (Proposition 4), any stable school-choice rule (Proposition 5), efficient allocations in housing assignment, and generalized median voting rules (Section B) do not admit fully contextually private protocols.&lt;/p&gt;
&lt;p&gt;Q: Why do k-item Vickrey auctions fail full contextual privacy?
A: Proposition 4 shows that k-item Vickrey auctions for k ≥ 1 do not admit fully contextually private protocols. The argument uses the necessary conditions from Theorem 1 (Corollaries 1 and 2): the Vickrey payment rule creates type-space regions where multiple agents together determine the price but no single agent is individually pivotal over the price, so any protocol implementing the Vickrey rule must produce violations for at least some agents at some type profiles.&lt;/p&gt;
&lt;p&gt;Q: What is the ascending-join protocol and what does Theorem 2 establish?
A: The ascending-join protocol is a specific dynamic elicitation protocol for k-item Vickrey auctions that repeatedly asks agents whether they can rule out a particular outcome, structured as threshold queries ascending from an initial guess. Theorem 2 proves that the ascending-join protocol is maximally contextually private for the k-item Vickrey auction. Proposition 7 formalizes the protection mechanism: the protocol delays queries to the bidders whose privacy it is protecting, querying them only when their responses become necessary for determining the outcome.&lt;/p&gt;
&lt;p&gt;Q: What does Proposition 6 establish about the structure of maximally contextually private protocols?
A: For a class of social choice rules on totally ordered type spaces that contains k-item Vickrey auctions, Proposition 6 shows it is without loss of generality to consider only protocols consisting of threshold queries that are monotonically increasing or decreasing in the threshold after an initial guess. This result serves as a theoretical reduction (enabling proofs that certain protocols are maximally private) and as a practical design principle (identifying the initial query and the ordering of agents as the two key design dimensions).&lt;/p&gt;
&lt;p&gt;Q: How does contextual privacy relate to obviously dominant strategies?
A: The paper treats privacy properties and incentive properties as largely orthogonal questions, to be analyzed separately. For the ascending-join protocol specifically, the authors verify obvious dominance — the most demanding incentive notion they consider — which requires that at every history, the worst-case payoff from the equilibrium action exceeds the best-case payoff from any deviation. This analysis proceeds after the contextual privacy properties of the protocol are established.&lt;/p&gt;
&lt;p&gt;Q: What is group contextual privacy and why do the authors focus on individual-level violations instead?
A: Group contextual privacy requires that whenever the designer learns any property of the joint type profile, that property must affect the outcome. The authors show (Proposition 11) that a protocol is fully group contextually private if and only if every query rules out at least one outcome. They argue this standard is extremely demanding and produces a very coarse partial order: improving in the group privacy order requires restructuring the entire protocol tree rather than making agent- or state-specific improvements. They also note that normative accounts of privacy, including Nissenbaum&amp;rsquo;s contextual integrity theory, center on individual rather than group information.&lt;/p&gt;
&lt;p&gt;Q: How does individual contextual privacy relate to nonbossiness?
A: Individual contextual privacy (Proposition 8) requires that if two type profiles differing only in agent i&amp;rsquo;s type are distinguished, they must lead to different allocations for agent i — presuming a private allocation domain. The paper shows that the set of individual contextual privacy violations equals the union of contextual privacy violations and nonbossiness violations: individual contextual privacy is violated precisely when either (a) agent i&amp;rsquo;s superfluous type information is revealed, or (b) agent i is &amp;ldquo;bossy&amp;rdquo; — able to change others&amp;rsquo; outcomes without changing their own.&lt;/p&gt;
&lt;p&gt;Q: What is the relationship between the ascending-join protocol and minimal relative informativeness?
A: Proposition 19 shows that the ascending-join protocol is not only maximally contextually private but also minimally relatively informative among protocols that are maximally contextually private. That is, among all maximally contextually private protocols, the ascending-join protocol reveals the smallest total amount of information about the type profile in the relative informativeness order. This establishes relative informativeness as a useful refinement for selecting among contextually privacy-equivalent protocols.&lt;/p&gt;
&lt;p&gt;Q: What motivates the exclusion of cryptographic tools and trusted mediators from the framework?
A: The authors work under the minimal assumption that the designer learns information if and only if an agent directly discloses it — no commitment to forget, anonymize, or cryptographically conceal. They motivate this on two grounds: first, many real-world auction formats are live and dynamic with no mediating technology; second, advanced cryptography is often costly in time, money, or computation, and studying the no-mediator benchmark can explain the historical prevalence of dynamic protocols and inform auction design in environments where cryptography may become unavailable (for example, due to quantum computing). The authors cite a Danish sugar-beet auction as a case where designers themselves questioned whether full multiparty computation was necessary.&lt;/p&gt;
&lt;p&gt;Contextual privacy violation: A protocol produces a contextual privacy violation for agent i at type profile θ if the designer can distinguish θ_i from some alternative type θ&amp;rsquo;_i — holding other agents&amp;rsquo; types fixed — yet the social choice rule assigns the same outcome at both profiles. The violation is assigned at the level of individual agent–state pairs.&lt;/p&gt;
&lt;p&gt;Maximally contextually private protocol: A protocol whose set of contextual privacy violations is inclusion-minimal among all protocols that implement the same social choice rule — equivalently, a protocol that lies on the Pareto frontier of implementation and contextual privacy, such that no other implementing protocol weakly reduces every violation and strictly reduces at least one.&lt;/p&gt;
&lt;p&gt;Iterative partition: A directed rooted tree whose nodes are subsets of the type space, where each non-leaf node is split into children by partitioning on a single agent&amp;rsquo;s type. Any protocol is equivalent (in terms of what the designer learns) to a partitional protocol induced by an iterative partition (Proposition 1).&lt;/p&gt;
&lt;p&gt;Individual pivotality: On a product set of type profiles, agent i is individually pivotal if there exist two subsets of agent i&amp;rsquo;s types such that every type profile from one subset leads to a different outcome than every type profile from the other subset, holding others&amp;rsquo; types fixed.&lt;/p&gt;
&lt;p&gt;Collective pivotality: Agents are collectively pivotal on a product set if there exist two type profiles in that set with different outcomes. Collective pivotality without any agent being individually pivotal is precisely the condition that forces contextual privacy violations (Theorem 1).&lt;/p&gt;
&lt;p&gt;Ascending-join protocol: A specific dynamic protocol for k-item Vickrey auctions that poses threshold queries in ascending order after an initial guess, repeatedly asking agents whether they can rule out a particular outcome. It is maximally contextually private (Theorem 2) and minimally relatively informative among maximally contextually private protocols (Proposition 19), and it achieves privacy protection by delaying queries to agents whose privacy it protects (Proposition 7).&lt;/p&gt;
&lt;p&gt;Relative informativeness: A partial order on protocols defined by: protocol P is less relatively informative than P&amp;rsquo; if every pair of type profiles P distinguishes is also distinguished by P&amp;rsquo;. Unlike contextual privacy, relative informativeness treats all disclosures as equally undesirable and does not condition on the social choice rule. The paper positions it as a useful refinement for selecting among contextually privacy-equivalent protocols.&lt;/p&gt;</description></item><item><title>Costly Multidimensional Screening</title><link>https://macropaperwarehouse.com/papers/costly-multidimensional-screening/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/costly-multidimensional-screening/</guid><description>&lt;p&gt;This paper studies when a principal can improve upon simple one-dimensional mechanisms by also deploying costly nonprice screening instruments — actions that are socially wasteful yet potentially informative about the agent&amp;rsquo;s private type.&lt;/p&gt;
&lt;p&gt;The model features a principal and an agent with quasilinear, additively separable preferences across two components: (i) a productive component, where allocations lie in a one-dimensional compact space X and generate genuine surplus, and (ii) a costly component, where any allocation y in an arbitrary measurable space Y satisfies sB(y, θB) ≤ 0 — it destroys or at best does not create social surplus. The agent&amp;rsquo;s private type is multidimensional, θ = (θA, θB), drawn from a commonly known distribution. Both components allow for nonlinear valuations and, on the principal&amp;rsquo;s side, interdependent preferences.&lt;/p&gt;
&lt;p&gt;The central result (Theorem 1) establishes that if the agent&amp;rsquo;s preferences between the productive and costly components are positively correlated — meaning that a higher θA implies a stochastically higher θB — then there exists an optimal mechanism that involves no costly screening. Moreover, if instruments are strictly costly, every optimal mechanism involves no costly screening almost everywhere. Positive correlation is defined in terms of stochastic dominance: θB | θA is stochastically nondecreasing in θA. A sufficient but not necessary condition is affiliation in the sense of Milgrom and Weber (1982).&lt;/p&gt;
&lt;p&gt;The intuition centers on two observations. First, under positive correlation, costly instruments can only help relax upward incentive constraints (deterring lower types from mimicking higher types). Second, under the surplus condition — a single-crossing condition on the surplus function sA(x, θA) requiring that if x generates more surplus than x&amp;rsquo; at some type, it continues to do so at all higher types — the principal can safely ignore upward incentive constraints at the optimum. The Downward Sufficiency Theorem (Theorem 2) formalizes the second observation: in any one-dimensional screening problem satisfying the surplus condition, there exists an optimal solution to the relaxed program (with only downward IC constraints) that also satisfies all upward IC constraints. Because monetary transfers fully substitute for costly instruments in relaxing downward constraints without destroying surplus, the costly instruments add no value under positive correlation.&lt;/p&gt;
&lt;p&gt;The proof proceeds via a monotone path decomposition of the multidimensional type space, exploiting a measurable monotone coupling (Lemma 1) to write θ = (θA, h(θA; ε)) where ε is independent of θA and h is nondecreasing. This reduces the problem to a family of one-dimensional paths, on each of which the Reconstruction Lemma (Lemma 2) shows that any costly mechanism can be weakly improved upon by one with no costly screening that satisfies all downward IC constraints.&lt;/p&gt;
&lt;p&gt;A partial converse (Proposition 1) shows that under negative correlation — when some dimension of θB is stochastically nonincreasing in θA — there exist utility functions satisfying the surplus condition for which any mechanism screening only the productive component is strictly dominated.&lt;/p&gt;
&lt;p&gt;The paper derives three applications. In monopoly pricing with costly signals (waiting in line, climbing stairs, collecting coupons), profit-maximizing mechanisms require no costly signals when higher-willingness-to-pay consumers also face weakly lower signal costs (Proposition 2). In monopsonistic labor market screening, the firm need not make offers contingent on costly credentials when higher-ability workers find credentialing easier — in contrast to the competitive Spence (1973) model where all screening must occur through costly effort because wages are pinned down by expected output (Proposition 3). In multiproduct pricing, the paper reinterprets bundle components as costly instruments for screening grand-bundle values, recovering Haghpanah and Hartline&amp;rsquo;s (2021) pure bundling optimality result and extending it to nested bundling (Proposition 4), under conditions that the incremental value of adding items to nested bundles is strictly increasing in type while the value of any non-nested bundle is nonincreasing relative to some nested superset.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s central research question?
A: The paper asks whether a principal can improve upon simple one-dimensional mechanisms by also deploying costly nonprice screening instruments when the agent has multidimensional private information. The goal is to characterize conditions under which augmenting a standard price menu with surplus-destroying actions — such as waiting in line, climbing stairs, or obtaining credentials — is or is not beneficial for the principal.&lt;/p&gt;
&lt;p&gt;Q: What does &amp;ldquo;positively correlated preferences&amp;rdquo; mean precisely in this model?
A: Positive correlation means that θB is stochastically nondecreasing in θA: for any θA &amp;lt; θ̂A, the conditional distribution of θB given θA first-order stochastically dominates that given θ̂A — i.e., θB | θA ≤_st θB | θ̂A. Observing a high θA conveys good news about θB in the stochastic dominance sense. A sufficient but not necessary condition is affiliation in the sense of Milgrom and Weber (1982). The condition is asymmetric and does not require full independence or monotone dependence in a deterministic sense.&lt;/p&gt;
&lt;p&gt;Q: What is the surplus condition and why does it matter?
A: The surplus condition is a single-crossing condition on the productive surplus function: for any x &amp;lt; x̂ and θA &amp;lt; θ̂A, if sA(x̂, θA) &amp;gt; sA(x, θA) then sA(x̂, θ̂A) &amp;gt; sA(x, θ̂A). It says that if a higher allocation generates more total surplus at some type, it continues to do so at all higher types. This condition ensures the existence of a monotone efficient allocation rule, and it is the key enabling condition for the Downward Sufficiency Theorem. It is automatically satisfied when the principal has no interdependent preferences and the agent satisfies increasing differences, and also when sA is strictly increasing in x or has nonnegative cross partial derivative.&lt;/p&gt;
&lt;p&gt;Q: What is the Downward Sufficiency Theorem and why is it the key technical result?
A: Theorem 2 states that in any one-dimensional screening problem satisfying the surplus condition, there exists an optimal solution to the relaxed program — which ignores all upward IC constraints — that also satisfies all upward IC constraints. This means the principal can solve the easier downward-IC-only problem and the solution is fully incentive compatible. The result is novel and uncovers a general property of one-dimensional screening problems beyond the standard monotone allocation rule setting. It is key because, combined with the observation that costly instruments under positive correlation can only relax upward constraints, it implies there is no benefit to using costly screening.&lt;/p&gt;
&lt;p&gt;Q: How does the proof handle the case of multidimensional types?
A: The proof uses a monotone path decomposition. By Lemma 1 (measurable monotone coupling), under positive correlation there exists a random variable ε independent of θA and a nondecreasing measurable function h such that θ =^d (θA, h(θA; ε)). This writes the joint type distribution as a family of monotone paths indexed by ε. On each path ε = e, the types are ordered by θA alone, reducing the problem to a one-dimensional screening problem. The Reconstruction Lemma (Lemma 2) then shows that on each such path, any mechanism involving costly screening can be replaced by one without costly screening that weakly improves principal payoff and satisfies all downward IC constraints.&lt;/p&gt;
&lt;p&gt;Q: What does the partial converse (Proposition 1) establish?
A: Proposition 1 shows that when some dimension i of the costly component satisfies that θi is stochastically nonincreasing in θA (negative correlation), and the type distribution has a density with |X| &amp;gt; 1 and |Y| &amp;gt; 1, then there exist utility functions satisfying the surplus condition for which any mechanism screening only the productive component is strictly dominated by one involving costly screening. This is not a full converse — it establishes existence of cases where costly screening is strictly beneficial, not that it is always beneficial under negative correlation.&lt;/p&gt;
&lt;p&gt;Q: How does the insurance example illustrate the two correlation cases?
A: In Example 1 (negative correlation), a low-risk type (θA = 0) values insurance at 2, a high-risk type (θA = 1) values it at 3; costs are 0 and 5/2 respectively; and the high-risk type also has higher disutility for the costly action. Without costly screening, the optimal mechanism sells full insurance at price 2 to both types for a profit of 3/4. With costly screening (e.g., requiring the agent to climb stairs to get full insurance), only the low-risk type purchases, yielding profit of 1 &amp;gt; 3/4. In Example 2 (positive correlation), the high-risk type has lower disutility for the costly action; any mechanism using the costly instrument is strictly dominated by simply selling full insurance at price 2 to both types.&lt;/p&gt;
&lt;p&gt;Q: How does the labor market application differ from Spence (1973)?
A: In Spence (1973), wages are competitive and pinned down by expected output, leaving no room to screen workers via monetary payments, so all screening must occur through costly credentials. In Yang&amp;rsquo;s model, the monopsonistic firm sets wages and all types face the same outside option, so monetary transfers can screen types. Proposition 3 says that when θB is stochastically nondecreasing in θA — higher-ability workers find credentials easier — no credential is needed in the optimal mechanism. The paper thus shows that costly screening is a feature of competitive, not monopsonistic, labor markets, under positive correlation of preferences.&lt;/p&gt;
&lt;p&gt;Q: What is the bundling application and what new results does it yield?
A: The paper reinterprets the multiproduct pricing problem by treating the grand bundle as the productive component and sub-bundles as costly instruments (since selling a sub-bundle instead of the grand bundle destroys social surplus relative to selling the grand bundle). Proposition 4 (nested bundling) establishes that a nested menu B of bundles is optimal among deterministic mechanisms if: (i) the incremental value of adding items to move from bundle b to b&amp;rsquo; ⊃ b in B is strictly increasing in θ, and (ii) for any bundle b not in B, there exists a nested superset b&amp;rsquo; ∈ B such that the value of b relative to b&amp;rsquo; is nonincreasing in θ. This extends and complements Haghpanah and Hartline (2021), which is recovered as the special case of pure bundling (Proposition 5).&lt;/p&gt;
&lt;p&gt;Q: What are the key scope conditions that delimit when Theorem 1 applies?
A: Theorem 1 requires: (i) additive separability of preferences across productive and costly components; (ii) the surplus condition on sA (single-crossing of total surplus in the productive component); (iii) the positive correlation condition (stochastic monotonicity of θB in θA); and (iv) the costly instruments satisfy sB(y, θB) ≤ 0 for all y, θB. The productive allocation space X must be compact and one-dimensional; Y can be any measurable space. The agent&amp;rsquo;s type space can be multidimensional. The result holds for both private values and interdependent valuations on the principal&amp;rsquo;s side.&lt;/p&gt;
&lt;p&gt;Q: Under what conditions does costly screening arise in practice, according to the model?
A: The model predicts that if costly screening instruments are observed in practice, the consumers or agents with higher willingness to pay (or ability) for the productive good must tend to face higher costs for the screening action. For instance, higher-willingness-to-pay consumers who find waiting in line more costly (positively correlated preferences) would not be subjected to waiting as a screening device. If a firm uses waiting in line, it must be because higher-willingness-to-pay consumers find waiting less costly — consistent with negative correlation.&lt;/p&gt;
&lt;p&gt;Costly Instruments: Allocations in the space Y such that the ex post social surplus sB(y, θB) = uB(y, θB) + vB(y, θB) ≤ 0 for all y and all θB. These include actions like waiting in line, collecting coupons, or obtaining credentials that destroy social surplus but may convey private information useful for screening.&lt;/p&gt;
&lt;p&gt;Productive Component: The one-dimensional allocation dimension X in which both principal and agent derive non-negative surplus, representing the intrinsically valuable output of the mechanism (e.g., insurance coverage, job placement, bundle of goods).&lt;/p&gt;
&lt;p&gt;Positive Correlation (Stochastic Monotonicity): The condition that θB is stochastically nondecreasing in θA: for any θA &amp;lt; θ̂A, the conditional distribution of θB given θA first-order stochastically dominates that given θ̂A. Equivalently, observing a higher θA conveys good news about θB. A sufficient condition is affiliation (Milgrom-Weber), but positive correlation is strictly weaker.&lt;/p&gt;
&lt;p&gt;Surplus Condition: A single-crossing condition on the total surplus function sA(x, θA) for the productive component: for any x &amp;lt; x̂ and θA &amp;lt; θ̂A, if x̂ generates strictly more surplus than x at type θA, it continues to do so at θ̂A. This ensures a monotone efficient allocation rule exists and is the enabling condition for the Downward Sufficiency Theorem.&lt;/p&gt;
&lt;p&gt;Downward Sufficiency Theorem (Theorem 2): The result that in any one-dimensional screening problem satisfying the surplus condition, there exists an optimal solution to the relaxed program (which ignores upward IC constraints) that also satisfies all upward IC constraints. This implies the principal need only enforce downward incentive constraints at the optimum.&lt;/p&gt;
&lt;p&gt;Monotone Path Decomposition: A proof technique that writes the multidimensional type distribution as θ =^d (θA, h(θA; ε)) where ε ⊥ θA and h is nondecreasing in θA. Borrowed from dynamic mechanism design (Eso-Szentes, Pavan-Segal-Toikka), it reduces multidimensional IC problems to families of one-dimensional paths indexed by the independent residual ε.&lt;/p&gt;
&lt;p&gt;Nested Bundling: A menu B of product bundles that can be totally ordered by set inclusion (b1 ⊂ b2 ⊂ &amp;hellip; ⊂ bK). The paper shows that nested bundling is optimal under conditions that the incremental value of nesting is strictly increasing in type for bundles within B, and nonincreasing relative to any nested superset for bundles outside B.&lt;/p&gt;</description></item><item><title>Debiasing and T-Tests for Synthetic Control Inference on Average Causal Effects</title><link>https://macropaperwarehouse.com/papers/debiasing-and-t-tests-for-synthetic-control-inference-on-average-causal-effects/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/debiasing-and-t-tests-for-synthetic-control-inference-on-average-causal-effects/</guid><description>&lt;p&gt;Chernozhukov, Wüthrich, and Zhu propose a debiased synthetic control (SC) estimator and an accompanying self-normalized t-test for making inferences on the average treatment effect on the treated (ATT) in aggregate panel data settings with one treated unit. The inferential target is the time-averaged treatment effect τ = (1/T1) Σ_{t=T0+1}^{T} (Y0t(1) − Y0t(0)), a one-number summary of the overall causal impact that admits standard-form confidence intervals, in contrast to per-period effects (which cannot be consistently estimated with one treated unit) and sharp null hypotheses (which do not inform effect magnitude).&lt;/p&gt;
&lt;p&gt;The method addresses two structural challenges in SC inference. First, the canonical SC estimator τ_SC is biased because the weights are estimated from high-dimensional pre-treatment data, and the bias can be substantial under misspecification. Second, even if true weights were known, constructing standard errors requires estimating the long-run variance (LRV), for which classical estimators such as Newey-West are unreliable in the small samples typical of SC applications.&lt;/p&gt;
&lt;p&gt;The debiasing procedure is a K-fold cross-fitting scheme applied to the pre-treatment period. The pre-treatment sample is split into K consecutive blocks. For each fold k, SC weights w_(k) are estimated on the leave-one-block-out pre-treatment data H_{(-k)}, and a component estimator τ_k is formed as the difference between the post-treatment SC residual (using w_(k)) and the in-block pre-treatment SC residual. The latter serves as an estimator of the bias, which under the model assumptions is stable across the pre- and post-treatment periods. The final estimator τ_hat is the average of τ_k across folds. A self-normalized t-statistic T_K = sqrt(K)(τ_hat − τ)/σ_τ is constructed using the cross-fold variance; its asymptotic distribution is t_{K-1}, so no LRV estimation is required and (1−α) confidence intervals take the textbook form τ_hat ± t_{K-1}(1−α/2) × σ_τ/sqrt(K).&lt;/p&gt;
&lt;p&gt;The t-test is proven valid with both stationary and non-stationary data. With stationary data (Theorem 2), it is valid under arbitrary misspecification. With non-stationary data, validity holds either when all units share a common nonstationarity (Theorem 3, also misspecification-robust) or when units deviate from a common nonstationarity under restrictions on the magnitude and heterogeneity of deviations but SC is correctly specified (Theorem 4). The latter covers heterogeneous deterministic time trends and certain cointegration structures. Researchers therefore need not pre-test for unit roots and select inference procedures accordingly.&lt;/p&gt;
&lt;p&gt;A formal efficiency result (Section 3.3) shows that the asymptotic variance of the debiased SC estimator is no larger than that of difference-in-differences (DID), because SC minimizes prediction error and w* dominates the equal-weight DID vector. The relative asymptotic efficiency (RAE) of the t-test versus DID rises with K: K=3 yields RAE of 63.56%; K=5 yields 82.08%; K=10 yields 92.25%.&lt;/p&gt;
&lt;p&gt;Simulations calibrated to Andersson&amp;rsquo;s (2019) Swedish carbon tax application — T0=30, T1=16, N=14, Gaussian AR(1) errors — show that the t-test at K=3 achieves coverage close to the nominal 90% level across correct-specification and misspecification DGPs, while Newey-West standard errors produce substantial undercoverage (coverage = 0.72–0.84) at moderate to high AR(1) coefficients. The method performs comparably to or better than subsampling (Li, 2020) and synthetic DID (Arkhangelsky et al., 2021), and avoids bandwidth selection.&lt;/p&gt;
&lt;p&gt;In the empirical application, the debiased SC t-test (K=3) applied to annual CO2 emissions from transport across Sweden (treated, 1990) and 14 OECD control countries over 1960–2005 yields a negative and statistically significant ATT, with a 90% confidence interval lying entirely below zero, implying approximately an 11% average reduction in per capita CO2 emissions from transport attributable to the Swedish carbon tax over 1990–2005. The pre-treatment AR(1) coefficient of SC residuals is approximately 0.31, supporting K=3 as appropriate. These findings corroborate and extend Andersson&amp;rsquo;s (2019) permutation-based results by providing a confidence interval for the magnitude of the average effect. The method is implemented in the R package scinference.&lt;/p&gt;
&lt;p&gt;Q: What is the primary inferential target and why is it preferred over per-period effects or sharp nulls?
A: The target is the ATT τ = (1/T1) Σ_{t=T0+1}^{T} (Y0t(1)−Y0t(0)), the time-averaged treatment effect on the treated unit over the post-treatment period. Per-period effects cannot be consistently estimated when there is only one treated unit, yielding wide and uninformative confidence intervals. Sharp nulls (e.g., of no effect whatsoever) are useful starting points but do not inform policy decisions about effect magnitude. The ATT provides an interpretable one-number summary and admits standard-form confidence intervals.&lt;/p&gt;
&lt;p&gt;Q: What are the two main inferential challenges that the paper addresses?
A: First, the canonical SC estimator τ_SC is biased due to estimation error in the high-dimensional weights, even under correct specification, and the bias can be substantial under misspecification. Second, even with known true weights, standard error estimation requires the long-run variance (LRV), for which classical estimators such as Newey-West (1987) and Andrews (1991) are not sufficiently accurate in the small samples typical of SC applications.&lt;/p&gt;
&lt;p&gt;Q: How does the K-fold cross-fitting procedure debias the SC estimator?
A: The pre-treatment period is divided into K consecutive blocks H1,&amp;hellip;,HK. For each fold k, SC weights w_(k) are estimated using leave-one-block-out pre-treatment data H_{(-k)}. The component estimator τ_k subtracts the in-block pre-treatment SC residual (an estimator of the bias in period Hk) from the post-treatment SC residual (using w_(k)). Because the bias is assumed stable across pre- and post-treatment periods, this subtraction removes it. The final estimator τ_hat averages τ_k across k=1,&amp;hellip;,K.&lt;/p&gt;
&lt;p&gt;Q: How does the self-normalized t-statistic avoid LRV estimation?
A: The statistic T_K = sqrt(K)(τ_hat − τ)/σ_τ uses σ_τ = sqrt(1 + Kr/T1) × sqrt[(1/(K−1)) Σ_k (τ_k − τ_hat)^2], which is the cross-fold standard deviation of the component estimators scaled by a factor reflecting the ratio of pre- to post-treatment block lengths. Under the asymptotic theory, T_K converges to a t_{K-1} distribution, which is pivotal and requires no bandwidth or kernel choice. The cross-fold structure acts as a self-normalizer analogous to the fixed-b approach in the LRV literature.&lt;/p&gt;
&lt;p&gt;Q: What does the paper prove about validity with non-stationary data?
A: Theorem 3 establishes that when all units share a common nonstationarity (Assumption 4: Yt(0) = Vt(0)+θt and Xt = Zt+1_N·θt where {Vt(0),Zt} is stationary and θt is unrestricted), T_K → t_{K-1} under arbitrary misspecification. Theorem 4 establishes validity when units deviate from common nonstationarity (Assumption 5) under restrictions on the magnitude and heterogeneity of deviations, but requires SC to be correctly specified. These results jointly imply that researchers need not pre-test for unit roots before applying the t-test.&lt;/p&gt;
&lt;p&gt;Q: How does the paper formally show that debiased SC is more efficient than DID?
A: The pseudo-true SC weights w* minimize mean squared prediction error over W_SC, so the residual variance σ^2_* = E(Yt(0)−Xt&amp;rsquo;w*)^2 ≤ E(Yt(0)−Xt&amp;rsquo;w_DID)^2 = σ^2_DID, where w_DID = (1/N,&amp;hellip;,1/N)&amp;rsquo; is the equal-weight DID vector. This inequality holds regardless of whether SC is correctly specified or not, so the efficiency gain over DID is unconditional. The t-test is also valid when the parallel trends assumption underlying DID is violated, making it more robust.&lt;/p&gt;
&lt;p&gt;Q: What is the trade-off in choosing K, and what does the paper recommend?
A: A larger K produces shorter confidence intervals (higher RAE: 63.56% at K=3 versus 92.25% at K=10) but may reduce coverage accuracy in finite samples because the t_{K-1} approximation improves with K while each block becomes smaller. The paper recommends K=3 as a starting point for typical SC applications where T0 is small, based on simulation evidence showing excellent 90% coverage at K=3. When T0 is moderate or large, K can be increased without loss of coverage accuracy.&lt;/p&gt;
&lt;p&gt;Q: What do the simulations show about the performance of Newey-West standard errors versus the t-test?
A: In simulations calibrated to the Swedish carbon tax application (T0=30, T1=16, N=14, AR(1) errors), the t-test at K=3 achieves coverage close to the nominal 90% level across both correct-specification and misspecification DGPs. Newey-West standard errors produce coverage of only 0.72–0.84 when the AR(1) coefficient of the error process is moderate to high. DID achieves nominal coverage when parallel trends hold but is biased and has poor coverage under violations of parallel trends.&lt;/p&gt;
&lt;p&gt;Q: How does the method compare with Li (2020) subsampling and synthetic DID (Arkhangelsky et al., 2021)?
A: Compared with Li (2020), the t-test allows N to grow with (T0,T1) rather than treating N as fixed, directly corrects for SC estimation bias via cross-fitting, avoids the need to pre-process data for stationarity, and does not require a subsampling bandwidth choice. Compared with SDID (Arkhangelsky et al., 2021), the t-test is simpler, does not require homoskedasticity across units as SDID&amp;rsquo;s placebo variance estimator does, and is developed under a linear prediction model rather than a factor model. Simulations show the t-test performs comparably to or better than both alternatives in the application-calibrated DGP.&lt;/p&gt;
&lt;p&gt;Q: What are the empirical findings for the Swedish carbon tax application?
A: Using annual CO2 emissions from transport for Sweden and 14 OECD control countries over 1960–2005, with T0=30 (1960–1989) and T1=16 (1990–2005), the debiased SC t-test at K=3 yields a negative and statistically significant ATT. The 90% confidence interval lies entirely below zero. The estimated average effect is approximately an 11% reduction in per capita CO2 emissions from transport attributable to the carbon tax over 1990–2005. The pre-treatment SC residuals show an estimated AR(1) coefficient of approximately 0.31, confirming moderate persistence and supporting the use of K=3.&lt;/p&gt;
&lt;p&gt;Q: When does the paper recommend against using the t-test?
A: The paper advises against the t-test when T1 is very small (T1 &amp;lt; 8–10), as asymptotic approximations may be inaccurate; when there are structural breaks shortly after T0 (making the ATT ill-defined); and when SC fit is poor because the treated unit is very different from controls. The method requires T0, T1, N → ∞ for asymptotic validity, and T1 ≥ 10–15 is suggested for reliable finite-sample performance.&lt;/p&gt;
&lt;p&gt;Q: How does the paper cover higher-order improvements in finite samples?
A: Appendix D formally establishes that the coverage error of the confidence interval I_K(1−α) is O(1/T) rather than O(1/sqrt(T)), analogous to the fixed-b approach in the LRV literature. This provides a formal justification for the excellent finite-sample coverage observed in the simulations and distinguishes the t-test from Gaussian approximations whose coverage error is of larger order.&lt;/p&gt;
&lt;p&gt;K-fold cross-fitting debiasing: A procedure that splits the pre-treatment period into K consecutive blocks, estimates SC weights on the leave-one-block-out pre-treatment data for each fold, and subtracts the in-block pre-treatment prediction error as an estimator of the bias. Under the model, the bias is assumed stable across pre- and post-treatment periods, so this subtraction removes it from the final estimator.&lt;/p&gt;
&lt;p&gt;Self-normalized t-statistic: A scale-free test statistic T_K = sqrt(K)(τ_hat − τ)/σ_τ whose denominator is the cross-fold standard deviation of the K component estimators, scaled to account for the ratio of pre-treatment block length to post-treatment period length. The statistic converges to a t_{K-1} distribution without requiring any LRV estimation.&lt;/p&gt;
&lt;p&gt;Average treatment effect on the treated (ATT): The target parameter τ = (1/T1) Σ_{t=T0+1}^{T} (Y0t(1)−Y0t(0)), representing the time-averaged causal effect of the treatment on the treated unit over the post-treatment period. It provides an interpretable one-number summary that admits standard-form confidence intervals, in contrast to per-period effects (not consistently estimable with one unit) and sharp null hypotheses (informative about presence but not magnitude of effect).&lt;/p&gt;
&lt;p&gt;Common nonstationarity: The condition (Assumption 4) that all units share the same nonstationary component θt — formally, Yt(0) = Vt(0)+θt and Xt = Zt+1_N·θt with {Vt(0),Zt} stationary and θt unrestricted. Under this condition, the t-test is valid under arbitrary misspecification of SC weights, without requiring the researcher to specify or pre-test the type of nonstationarity.&lt;/p&gt;
&lt;p&gt;Relative asymptotic efficiency (RAE): The ratio of the asymptotic expected confidence interval length of the debiased SC t-test to a benchmark (taken as K→∞), quantifying the cost in interval length from using a finite K. At K=3, RAE = 63.56%; at K=5, RAE = 82.08%; at K=10, RAE = 92.25%.&lt;/p&gt;
&lt;p&gt;Long-run variance (LRV): The quantity that governs the asymptotic variance of time-averaged quantities in settings with serially correlated data. The paper argues that classical LRV estimators (Newey-West, Andrews) are insufficiently accurate in the small samples typical of SC applications, motivating the self-normalization approach that avoids LRV estimation entirely.&lt;/p&gt;
&lt;p&gt;Pseudo-true SC weights: The population minimizer w* = argmin_{w ∈ W_SC} E(Yt(0)−Xt&amp;rsquo;w)^2, defined as the best linear predictor of the treated unit&amp;rsquo;s counterfactual outcome within the SC simplex constraint. These weights exist and satisfy the efficiency bound even under model misspecification, providing the foundation for the efficiency comparison with DID.&lt;/p&gt;</description></item><item><title>Decision Theory for Treatment Choice Problems with Partial Identification</title><link>https://macropaperwarehouse.com/papers/decision-theory-for-treatment-choice-problems-with-partial-identification/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/decision-theory-for-treatment-choice-problems-with-partial-identification/</guid><description>&lt;p&gt;This paper applies classical statistical decision theory (Wald 1950) to treatment choice problems where the data only partially identify payoff-relevant parameters. The policy maker chooses an action a in [0,1] — interpreted as the share of the population assigned to a new policy — to maximize welfare that is linear in the action. The data are Gaussian, and the key departure from prior literature is that the mean function mapping parameters to data need not be injective, so even infinite data may not reveal the optimal action.&lt;/p&gt;
&lt;p&gt;The paper evaluates decision rules under three classical criteria: admissibility, maximin welfare, and minimax regret (MMR).&lt;/p&gt;
&lt;p&gt;Admissibility result (Theorem 1): Under nontrivial partial identification, every decision rule — however exotic — is welfare-admissible. No rule is dominated. This is a sharp reversal from point-identified settings, where admissibility meaningfully restricts the rule class: in the scalar point-identified case (n=1, m(theta)=theta), Karlin and Rubin&amp;rsquo;s (1956) result implies that any non-threshold rule is dominated. The proof exploits completeness of the Gaussian statistical model: if a dominating rule d&amp;rsquo; existed, it would have to agree almost everywhere with d, yielding a contradiction. Theorem 5 generalizes this result beyond Gaussian likelihoods, tying it to bounded completeness of the statistical model.&lt;/p&gt;
&lt;p&gt;Maximin welfare result (Theorem 2): The maximin criterion selects the no-data rule d(y) = 0 — preserve the status quo regardless of data — whenever the status quo welfare is the infimum over states with non-positive welfare contrast. In the running example, maximin welfare equals zero and is achieved by never assigning the new policy. This echoes critiques from Savage (1951) and Manski (2004) about ultra-pessimism.&lt;/p&gt;
&lt;p&gt;Minimax regret result (Theorem 3): In point-identified problems, the MMR rule is essentially unique and nonrandomized (Canner 1970; Stoye 2009a; Tetenov 2012). Under partial identification, when the identified set is large enough — formally, when I(0) is large enough and there exists mu in the identified set with I(mu) &amp;gt; I(0) — there are infinitely many MMR optimal rules, and any symmetric, weakly increasing MMR rule depending only on the sufficient statistic (w*)^T Y must randomize for some data realizations. Moreover, if I(mu) is differentiable at zero, no linear threshold rule is MMR optimal.&lt;/p&gt;
&lt;p&gt;Least randomizing MMR rule (Theorem 4): Because policy randomization is difficult to implement in practice, the authors uniquely characterize the MMR optimal rule that randomizes least frequently. Among all symmetric, weakly increasing, unimodal MMR optimal rules depending on (w*)^T Y, the rule d*_linear has the smallest randomization region — every other distinct such rule has a strictly wider randomization region. This rule can be profiled-regret dominant over the Stoye (2012a)/Yata (2023) MMR rule (Proposition 2), and the uniformly randomizing rule is inadmissible under profiled regret (Proposition 3). Under some conditions, d*_linear can also be obtained as the MMR rule within a class that penalizes randomized assignments equally (Proposition 4).&lt;/p&gt;
&lt;p&gt;Three applications ground the theory. First, in Ishihara and Kitagawa&amp;rsquo;s (2021) evidence aggregation framework — extrapolating treatment effects from n source countries to a target country — the least randomizing rule randomizes only when estimated bounds on the target treatment effect straddle zero, linking decision rules directly to identified-set estimators. Second, in LATE extrapolation (Mogstad et al. 2018), all decision rules are admissible and IV-based threshold rules are not dominated. Third, in the omitted-variable-bias setting of Diegert et al. (2022), the decision-theoretic breakdown point — the largest confounding magnitude under which the seemingly better policy should be adopted without hedging — tolerates strictly more confounding than Diegert et al.&amp;rsquo;s breakdown point, where the threshold is k = sqrt(pi/2) * sigma.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question?
A: The paper asks how classical statistical decision theory — admissibility, maximin welfare, minimax regret — applies when the data only partially identify the payoff-relevant parameters governing a binary treatment choice. Prior literature had developed these criteria for point-identified settings; this paper characterizes how partial identification fundamentally changes the answers.&lt;/p&gt;
&lt;p&gt;Q: What is the formal framework?
A: The policy maker chooses a in [0,1] (population share assigned to the new policy) with welfare W(a,theta) = a*W(1,theta) + (1-a)*W(0,theta), linear in a. The data are Y ~ N(m(theta), Sigma) with known m and Sigma. Partial identification arises when m is not injective, so distinct parameter values theta and theta&amp;rsquo; with opposite-sign welfare contrasts U(theta) = W(1,theta) - W(0,theta) can produce the same data distribution.&lt;/p&gt;
&lt;p&gt;Q: Why does admissibility lose all refinement power under partial identification?
A: Theorem 1 shows that every decision rule is admissible when there is nontrivial partial identification. The mechanism is Gaussian completeness: if a dominating rule d&amp;rsquo; existed, then for every data distribution in the model, d and d&amp;rsquo; would have equal expected values, which by completeness implies d = d&amp;rsquo; almost everywhere — a contradiction. This relies on the fact that nontrivial partial identification ensures that each data distribution is compatible with both positive and negative welfare contrasts, preventing the construction of a uniformly dominating rule.&lt;/p&gt;
&lt;p&gt;Q: What is the contrast with point-identified settings?
A: In the scalar point-identified case (n=1, m(theta)=theta, W(1,theta)=theta, W(0,theta)=0), Karlin and Rubin&amp;rsquo;s (1956) theorem implies any non-threshold rule is dominated; admissibility restricts attention to threshold rules. Partial identification completely eliminates this refinement: even randomized or otherwise arbitrary rules are admissible.&lt;/p&gt;
&lt;p&gt;Q: What does the maximin welfare criterion recommend?
A: Theorem 2 shows that when the status quo welfare equals the infimum of welfare over states with non-positive welfare contrast, the maximin optimal rule is d(y) = 0 for all y — preserve the status quo regardless of the data. In the running evidence-aggregation example, maximin welfare equals zero and is achieved by never assigning the new policy. The criterion ignores all data because the worst case is always achieved at states where the new policy performs no better than the status quo.&lt;/p&gt;
&lt;p&gt;Q: What is the minimax regret criterion and why is it preferred?
A: Expected regret at state theta is R(d,theta) = U(theta)*{1{U(theta)&amp;gt;=0} - E[d(Y)]} — the expected welfare loss relative to the oracle who knows theta. A rule is MMR optimal if it minimizes worst-case expected regret. Unlike maximin welfare, MMR uses data and balances risks across states. In point-identified settings it yields essentially unique, nonrandomized rules.&lt;/p&gt;
&lt;p&gt;Q: How does partial identification change the MMR solution set?
A: Theorem 3 shows that when the identified set is large enough — I(0) is sufficiently large and there exists mu with I(mu) &amp;gt; I(0) — there are infinitely many MMR optimal rules, and every symmetric, weakly increasing MMR rule depending on the sufficient statistic (w*)^T Y must randomize for some data realizations. If I(mu) is differentiable at zero, no linear threshold rule is MMR optimal. Different MMR rules can recommend different policies for the same data, creating a nontrivial multiplicity problem.&lt;/p&gt;
&lt;p&gt;Q: How is the least randomizing MMR rule characterized?
A: Theorem 4 shows that among all symmetric, weakly increasing, unimodal MMR optimal rules that depend on data only through (w*)^T Y, the rule d*_linear has the smallest randomization region: every other distinct rule in this class has a strictly wider randomization region, V(d*_linear) ⊆ V(F∘w*) with strict inclusion when F ≠ d*_linear. This characterization is essentially unique and provides a pragmatic refinement of the MMR solution set.&lt;/p&gt;
&lt;p&gt;Q: What is profiled regret and why is it used?
A: Profiled regret reports worst-case expected regret at each fixed value of the point-identified parameters, rather than worst-case over all parameters jointly. Proposition 2 shows that the least randomizing rule d*_linear can profiled-regret dominate the Stoye (2012a)/Yata (2023) MMR rule in the running example. Proposition 3 shows that the uniformly randomizing rule is profiled-regret inadmissible when profiling over point-identified parameters. This concept provides an additional selection criterion within the MMR solution set.&lt;/p&gt;
&lt;p&gt;Q: Can the least randomizing rule be derived from an explicit welfare penalty?
A: Proposition 4 shows that, under some conditions, d*_linear is minimax regret optimal within the class of rules that penalize all randomized assignments equally. This connects the least randomizing criterion to a modified welfare function that treats randomization itself as costly, providing an interpretation for the refinement beyond mere pragmatics.&lt;/p&gt;
&lt;p&gt;Q: What does the evidence aggregation application show?
A: In the Ishihara-Kitagawa (2021) framework — extrapolating effects from n source countries to a target country using Lipschitz smoothness — the least randomizing rule randomizes only (though not always) when the estimated bounds on the target treatment effect contain both positive and negative values. When bounds are entirely positive or entirely negative, the rule recommends a deterministic action. This shows how identified-set estimators directly enter decision-theoretically optimal rules.&lt;/p&gt;
&lt;p&gt;Q: What does the LATE extrapolation application show?
A: In the Mogstad et al. (2018) setting with a binary instrument and no covariates, where the payoff-relevant parameter is a policy-relevant treatment effect corresponding to expanding the complier subpopulation, Theorem 1 applies: all decision rules are admissible. In particular, the IV threshold rule — implement the policy for large IV estimates — is not dominated, providing decision-theoretic grounding for a common empirical practice.&lt;/p&gt;
&lt;p&gt;Q: What does the omitted variable bias application show?
A: In the Diegert et al. (2022) setting where the identified set for the long regression coefficient given the medium regression coefficient is [beta_med - k, beta_med + k], the least randomizing MMR rule is d*_linear(beta_hat_med) when k &amp;gt; sqrt(pi/2) * sigma. The decision-theoretic breakdown point — the largest k under which the seemingly better policy should be adopted without randomization — is strictly larger than Diegert et al.&amp;rsquo;s sensitivity breakdown point, meaning the decision-theoretic approach tolerates more confounding before recommending hedging.&lt;/p&gt;
&lt;p&gt;Q: How does Theorem 5 generalize Theorem 1 beyond Gaussian likelihoods?
A: Theorem 5 extends the admissibility result by connecting it to bounded completeness of the statistical model rather than Gaussian-specific completeness. This shows that the collapse of admissibility&amp;rsquo;s refinement power is not an artifact of normality but a general consequence of partial identification combined with a sufficiently rich statistical model.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s broader implication for empirical practice?
A: The results show that under partial identification, two of the three classical decision-theoretic criteria (admissibility and maximin welfare) provide no useful guidance — the former because everything passes, the latter because it ignores data entirely. MMR remains the operative criterion but yields infinitely many rules, all requiring some randomization. The least randomizing refinement provides a unique, practically implementable rule that connects to estimated identified sets and tolerates more ambiguity than purely statistical sensitivity analyses.&lt;/p&gt;
&lt;p&gt;Partial identification: A setting where even infinite data cannot uniquely determine payoff-relevant parameters, because the mean function m mapping parameters to data distributions is not injective. Distinct parameter values with opposite-sign welfare contrasts may be observationally equivalent.&lt;/p&gt;
&lt;p&gt;Welfare contrast U(theta): The difference W(1,theta) - W(0,theta) between the welfare under the new policy and under the status quo at parameter theta. The oracle optimal action is 1{U(theta) &amp;gt;= 0}.&lt;/p&gt;
&lt;p&gt;Admissibility (welfare): A rule d is admissible if no rule d&amp;rsquo; weakly dominates it in expected welfare at every theta with strict improvement at some theta. Under partial identification with Gaussian likelihood, every rule is admissible — admissibility has no refinement power.&lt;/p&gt;
&lt;p&gt;Maximin welfare optimality: A rule is maximin optimal if it attains the highest worst-case expected welfare. Under partial identification, this criterion selects the no-data rule (always preserve status quo) whenever the status quo welfare equals the infimum over states with non-positive welfare contrast.&lt;/p&gt;
&lt;p&gt;Minimax regret (MMR) optimality: A rule minimizes the worst-case expected welfare loss relative to the oracle action. Under severe enough partial identification, MMR optimal rules are non-unique and all require randomizing policy recommendations for some data realizations.&lt;/p&gt;
&lt;p&gt;Least randomizing MMR rule (d*_linear): The unique MMR optimal rule with the smallest randomization region among all symmetric, weakly increasing, unimodal MMR rules depending on the sufficient statistic. Characterized in Theorem 4; randomizes only when estimated identified set bounds straddle zero in the running example.&lt;/p&gt;
&lt;p&gt;Profiled regret: The worst-case expected regret at each fixed value of the point-identified parameters, treating them as a parameter of interest and profiling out the partially identified parameters. Provides a finer ranking within the MMR solution set and renders the uniformly randomizing rule inadmissible.&lt;/p&gt;</description></item><item><title>Default Options and Retirement Saving Dynamics</title><link>https://macropaperwarehouse.com/papers/default-options-and-retirement-saving-dynamics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/default-options-and-retirement-saving-dynamics/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; Does automatic enrollment (auto-enrollment) in retirement savings plans increase lifetime wealth accumulation and welfare? The prior literature established large short-run participation effects but had not traced the policy&amp;rsquo;s consequences over a full working life.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The paper draws on two primary sources. First, a proprietary panel of 401(k) administrative records from nearly 600 U.S. firms, covering roughly 159,216 first-year employees across 86 firms (for the &amp;ldquo;increasing default&amp;rdquo; fact) and 6,415 employees across 34 firms (for structural estimation), observed between December 2006 and December 2017. Second, 12 successive waves (2006–2017) of the U.K. Annual Survey of Hours and Earnings (ASHE), a 1% nationally representative panel of approximately 200,000 private-sector employees per year, including 37,120 job-switchers, used to exploit the phased rollout of the U.K. Pension Act of 2008.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Methodology.&lt;/strong&gt; The paper proceeds in three steps. (1) Three empirical stylized facts are documented using quasi-experimental variation (comparing employees hired before versus after changes in the default contribution rate within the same firm, and exploiting the staggered employer-size-based rollout of U.K. auto-enrollment). (2) A structural lifecycle model is estimated via the Method of Simulated Moments, using three preference parameters—intertemporal discount factor (δ), elasticity of intertemporal substitution (σ), and opt-out cost (k)—identified from the within-firm default variation in 34 U.S. firms. (3) The estimated model is used for out-of-sample validation and counterfactual welfare analysis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three stylized facts.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Fact I — Increasing the default reduces participation.&lt;/em&gt; Among 159,216 first-year employees in 86 auto-enrollment firms, each percentage-point increase in the default contribution rate reduces 401(k) participation by approximately 1 percentage point and increases contributions strictly below the new default by 1 percentage point. When the default rose from 3% to 6%, workers were 3.2 percentage points more likely to contribute at 1% or 2% of salary. This &amp;ldquo;drop-out&amp;rdquo; pattern is consistent with an opt-out cost model but is inconsistent with loss-aversion and psychological-anchoring theories, both of which predict that raising the default should weakly increase low-end contributions.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Fact II — Non-autoenrolled workers catch up within three years.&lt;/em&gt; In the estimation sample of 34 U.S. firms offering a 50% match up to 6% and an auto-enrollment default of 3%, median cumulative employee 401(k) contributions of non-autoenrolled workers equal those of autoenrolled workers after three years of tenure. Because non-autoenrolled workers compensate for initial non-participation by contributing more later—earning similar cumulative employer match and tax benefits over the full three-year horizon—a modest opt-out cost suffices to explain the observed inertia. Previous studies (which examined only the first year of tenure and did not allow future contribution adjustment) inferred opt-out costs of $1,000–$2,200 or more; the dynamic model implies a cost of only approximately $250.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Fact III — Prior auto-enrollment reduces saving in the next job.&lt;/em&gt; Using the phased U.K. policy rollout, workers who were auto-enrolled in their previous job and then move to a new employer that has not yet implemented auto-enrollment participate 12.8 percentage points less and contribute 0.55% of salary less in the new plan relative to otherwise similar job-switchers from non-auto-enrollment employers. When the new employer also has auto-enrollment, no statistically significant difference is observed. Placebo rollout tests confirm the effect is not a pre-existing selection pattern. This negative spillover contradicts a &amp;ldquo;savings habit&amp;rdquo; hypothesis and suggests that auto-enrollment&amp;rsquo;s short-run boost overstates lifetime savings effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural estimation results.&lt;/strong&gt; The estimated quarterly discount factor is δ = 0.987 (approximately 0.949 annually), and the elasticity of intertemporal substitution is σ = 0.435, both standard in lifecycle models. The opt-out cost is estimated at &lt;strong&gt;$254&lt;/strong&gt; per contribution-rate change (standard error $11). Sensitivity exercises show that combining a short observation window (first year only), sticky contributions (no intra-job adjustment), no income uncertainty, immediate vesting, and penalty-free DC withdrawals yields an opt-out cost of $3,004—broadly matching the range in previous studies. The low baseline estimate is thus driven by the dynamic nature of decisions (ability to compensate later), the illiquidity of retirement accounts (which reduces their perceived value), and income uncertainty (which expands the inaction range).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Long-run wealth effects.&lt;/strong&gt; Simulating a universal 3% auto-enrollment policy, the model predicts that &lt;strong&gt;wealth at retirement changes by less than 2% for the top 7 income deciles&lt;/strong&gt;. For individuals in the top two deciles, total wealth at age 65 is actually reduced by less than 1% because many high earners who would voluntarily contribute above 3% are pulled down to the default. At the &lt;strong&gt;bottom decile&lt;/strong&gt;, however, auto-enrollment raises total retirement wealth by more than &lt;strong&gt;12%&lt;/strong&gt;; savings increases are concentrated in the first 20 years of working life and peak around age 45, where bottom-quintile workers hold an additional 20% of average annual lifetime earnings. Even at the bottom, approximately one-third of the early savings gains are offset by lower contributions after age 45, as the wealth effect dominates. Crowd-out of liquid savings is limited: for bottom-quintile individuals, &lt;strong&gt;89%&lt;/strong&gt; of the increase in retirement savings at age 65 passes through to total wealth; for middle-quintile individuals, &lt;strong&gt;62%&lt;/strong&gt; passes through.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Out-of-sample validation.&lt;/strong&gt; The U.S.-estimated model is not rejected (at the 10% level) in 8 of 11 response moments in the 86-firm sample where defaults were raised between two positive rates, covering over 85% of workers. Recalibrated to U.K. institutions (using δ and σ from the U.S. and k = £160 via the average USD/GBP exchange rate), the model replicates the roughly 30-percentage-point increase in both participation and contributions at the 1% U.K. default. The model also predicts a 9.6-percentage-point drop in participation when workers move from an auto-enrollment to an opt-in employer, close to the empirical 12.8 percentage points.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Welfare and optimal policy.&lt;/strong&gt; Under utilitarian preferences (policymaker shares individuals&amp;rsquo; discount rate, no redistributive motive), the opt-in regime is always preferred to auto-enrollment regardless of policy incidence, because matching and tax incentives already induce over-saving relative to individuals&amp;rsquo; revealed time preferences. Under &lt;strong&gt;paternalistic&lt;/strong&gt; preferences (social discount factor = 1) or &lt;strong&gt;inequality-averse&lt;/strong&gt; preferences (Pareto weights inversely proportional to income, with degree of inequality aversion ν = 1 following Saez 2002), an auto-enrollment default at or near the employer matching threshold (6% of income) maximizes social welfare. A 6% auto-enrollment default improves welfare by 0.3% in lifetime consumption-equivalent for the bottom decile even under a utilitarian policymaker when incidence is on employers. These optimal policy rankings are robust to whether the opt-out cost is treated as fully welfare-relevant (π = 1) or welfare-irrelevant (π = 0), and hold under three incidence scenarios (employer profit reduction, match-rate adjustment, wage adjustment).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-core-mechanism-by-which-non-autoenrolled-workers-catch-up-at-the-median-and-why-does-this-reduce-the-implied-opt-out-cost-relative-to-prior-estimates"&gt;Q1. What is the core mechanism by which non-autoenrolled workers &amp;ldquo;catch up&amp;rdquo; at the median, and why does this reduce the implied opt-out cost relative to prior estimates?&lt;/h3&gt;
&lt;p&gt;A: Non-autoenrolled workers who do not contribute in their first year are not permanently forgoing employer matching and tax benefits; they can contribute more later in the same job and earn similar cumulative benefits. The paper shows that at the median and 75th percentile, cumulative employee 401(k) contributions among opt-in workers equal those of autoenrolled workers after three years of tenure in 34 U.S. firms offering a 50%-up-to-6% match at a 3% default. This dynamic substitutability means the opportunity cost of initial non-participation is far smaller than one-period back-of-the-envelope calculations suggest. Previous studies, which implicitly or explicitly assumed static contribution decisions or examined only the first year, inferred opt-out costs of $1,000–$2,200; in a fully dynamic model the same inertia requires only ~$254.&lt;/p&gt;
&lt;h3 id="q2-why-does-fact-i-higher-default-reduces-participation-specifically-rule-out-loss-aversion-and-anchoring-as-the-primary-mechanism-and-what-does-it-support-instead"&gt;Q2. Why does Fact I (higher default reduces participation) specifically rule out loss aversion and anchoring as the primary mechanism, and what does it support instead?&lt;/h3&gt;
&lt;p&gt;A: Under loss aversion, contributions above the default feel like losses while contributions below the default feel like gains. Raising the default shifts some contributions from the loss domain into the gain domain, making low contributions relatively less attractive. Proposition 2 demonstrates formally that loss-averse preferences predict a weakly lower fraction contributing below the new (higher) default — the opposite of what is observed. Similarly, Proposition 3 shows that psychological anchoring shifts preferences toward the new default, also predicting more participation at low rates when the default rises. Only the opt-out cost model (Proposition 1) predicts that a higher default causes some workers to incur the cost to switch &lt;em&gt;away&lt;/em&gt; from the default and end up at lower contribution rates, matching the empirical finding that each 1-percentage-point rise in the default increases contributions strictly below the old default by approximately 1 percentage point.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-quantitative-magnitude-of-the-opt-out-cost-and-what-modeling-assumptions-are-responsible-for-it-being-much-smaller-than-prior-estimates"&gt;Q3. What is the quantitative magnitude of the opt-out cost, and what modeling assumptions are responsible for it being much smaller than prior estimates?&lt;/h3&gt;
&lt;p&gt;A: The baseline estimate is $254 per contribution-rate change (s.e. $11), roughly an order of magnitude smaller than prior estimates of $1,000–$3,000+. Table 4 decomposes the sources of the difference: using only first-year data changes the estimate only slightly (to $226). Assuming contributions cannot be changed within a job (&amp;ldquo;sticky contributions&amp;rdquo;) raises the cost to $308 with four years of data or $712 with one year of data. Eliminating income uncertainty raises the estimate to $465. Assuming immediate vesting raises it to $344. Assuming penalty-free DC withdrawals raises it to $609. Combining all these restrictions simultaneously yields $3,004 — closely matching the prior literature. The three key drivers are thus: (1) the ability to adjust contributions over time within a job; (2) the illiquidity of the DC account (early-withdrawal penalties); and (3) income uncertainty widening the inaction range.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-paper-validate-the-structural-model-out-of-sample-and-what-confidence-does-this-provide-in-the-long-run-predictions"&gt;Q4. How does the paper validate the structural model out of sample, and what confidence does this provide in the long-run predictions?&lt;/h3&gt;
&lt;p&gt;A: Two out-of-sample exercises are reported. First, the model estimated on 34 U.S. firms (introduction of auto-enrollment from 0% to 3% default) is used to predict workers&amp;rsquo; response when 86 other firms raised the default from one positive rate to a higher rate. The model prediction cannot be rejected at the 10% level in 8 of 11 response-moment cases, covering 71 of 86 firms and more than 85% of workers. Second, the model is re-calibrated to U.K. institutions (keeping U.S. preference estimates, setting k = £160 via exchange rate) and applied to the phased rollout of the U.K. Pension Act of 2008. The model replicates the roughly 30-percentage-point increase in both participation and contributions at the 1% default following the policy, and predicts a 9.6-percentage-point drop in participation when previously autoenrolled workers move to a new opt-in employer — compared with an empirical estimate of 12.8 percentage points (s.e. 5.5 pp).&lt;/p&gt;
&lt;h3 id="q5-what-are-the-distributional-implications-of-a-universal-3-auto-enrollment-policy-for-wealth-at-retirement"&gt;Q5. What are the distributional implications of a universal 3% auto-enrollment policy for wealth at retirement?&lt;/h3&gt;
&lt;p&gt;A: The effect is concentrated at the bottom. For the top 7 income deciles, retirement wealth at age 65 changes by less than 2% relative to the opt-in counterfactual. For the top two deciles, total wealth at age 65 is actually reduced by less than 1% because high-earning workers who would voluntarily contribute above 3% are pulled down to the default. For the bottom decile, the policy raises total retirement wealth by more than 12%. Even at the bottom, roughly one-third of the early savings gains are later offset by lower contributions after age 45 as the wealth effect dominates, so even 20-year empirical follow-ups may overstate the policy&amp;rsquo;s lifetime effect at the bottom.&lt;/p&gt;
&lt;h3 id="q6-how-large-is-crowd-out-of-liquid-savings-by-auto-enrollment-and-what-explains-the-limited-degree-of-substitution"&gt;Q6. How large is crowd-out of liquid savings by auto-enrollment, and what explains the limited degree of substitution?&lt;/h3&gt;
&lt;p&gt;A: Crowd-out is modest. For bottom-quintile workers, 89% of the increase in retirement savings at age 65 translates into higher total wealth; for middle-quintile workers, 62% passes through. The limited crowd-out arises because liquid assets serve a precautionary motive and DC accounts serve a lifecycle motive — the two assets are not close substitutes. Additionally, as in Kaplan and Violante (2014), the marginal propensity to consume out of liquid assets is high in the model, so autoenrolled workers reduce consumption rather than run down liquid balances. These predictions align with Beshears et al. (2021), who find no significant increase in unsecured debt after four years, and Chetty et al. (2014), who estimate an 80% pass-through to total savings in a different Danish policy.&lt;/p&gt;
&lt;h3 id="q7-why-do-previously-autoenrolled-workers-contribute-less-when-they-switch-to-an-opt-in-employer-and-how-is-this-consistent-with-the-model"&gt;Q7. Why do previously autoenrolled workers contribute less when they switch to an opt-in employer, and how is this consistent with the model?&lt;/h3&gt;
&lt;p&gt;A: The most plausible explanation, and the one consistent with the model&amp;rsquo;s out-of-sample predictions, is a standard wealth effect: workers auto-enrolled early accumulate more retirement wealth and therefore have less incentive to contribute in a new job. The model predicts a 9.6-percentage-point participation drop for AE-to-non-AE movers, close to the empirical 12.8 pp. An alternative explanation — that previously autoenrolled workers rationally expect their new employer to soon adopt auto-enrollment and thus delay active enrollment — is partially ruled out by the finding that the empirical estimate is closer to the model prediction for job-switchers whose new employer is not expected to adopt auto-enrollment in the next 12 months.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-welfare-implications-of-auto-enrollment-under-utilitarian-paternalistic-and-inequality-averse-policymakers-and-how-robust-are-these-to-the-incidence-assumption"&gt;Q8. What are the welfare implications of auto-enrollment under utilitarian, paternalistic, and inequality-averse policymakers, and how robust are these to the incidence assumption?&lt;/h3&gt;
&lt;p&gt;A: Under utilitarian preferences (policymaker shares individuals&amp;rsquo; discount factor, no extra redistributive weight), the opt-in regime is always preferred regardless of whether the policy&amp;rsquo;s cost falls on employer profits, the match rate, or wages. The negative welfare effect is largest when incidence falls on wages (approximately 50% larger than under match-rate reduction). Under paternalistic preferences (social discount factor = 1), a 6% default (equal to the employer matching threshold) is optimal under all three incidence scenarios. Under inequality-averse preferences (ν = 1 Pareto weights), a 6% default is optimal when incidence falls on employers, and a 5% default when incidence falls on workers. These results are identical whether the opt-out cost is treated as fully welfare-relevant (π = 1) or welfare-irrelevant (π = 0). A 6% auto-enrollment default increases welfare by 0.3% in lifetime consumption-equivalent for the bottom income decile even under a utilitarian planner when incidence is on employers.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-address-heterogeneity-in-default-effects-across-age-and-income-groups-within-a-parsimonious-homogeneous-preference-model"&gt;Q9. How does the paper address heterogeneity in default effects across age and income groups within a parsimonious homogeneous preference model?&lt;/h3&gt;
&lt;p&gt;A: The model has only three estimated preference parameters (δ, σ, k), yet it endogenously replicates empirical heterogeneity. Conditional on participating, workers in their 20s are approximately 20 percentage points more likely to stay at the 3% default than workers in their late 50s and early 60s; the model attributes this to the option value of waiting: young workers can compensate for current non-saving by contributing more later, so the cost of opting out is effectively smaller for them. The lowest-income workers are approximately 40 percentage points more likely to remain at the default than the highest-paid; the model explains this primarily because the fixed opt-out cost of $254 represents a larger share of earnings for low-income individuals (and secondarily because high-income workers have more to gain from active contribution decisions due to higher marginal tax rates and a lower Social Security replacement rate). All model-predicted coefficients fall within the 95% confidence intervals of the empirical estimates.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-paper-conclude-about-the-broader-relevance-of-the-dynamic-opt-out-cost-framework-beyond-retirement-saving"&gt;Q10. What does the paper conclude about the broader relevance of the &amp;ldquo;dynamic opt-out cost&amp;rdquo; framework beyond retirement saving?&lt;/h3&gt;
&lt;p&gt;A: The paper argues that wherever individuals can compensate for present inaction with future actions — as in retirement saving — the observed inertia at a default understates the freedom of choice preserved by the nudge, and short-run effects overstate long-term consequences. In contrast, in domains such as healthcare plan choice or school selection, future actions cannot easily offset present inertia; opt-out costs are likely to remain large; and the distinction between a nudge and a hard mandate collapses. The paper therefore argues that the appeal of &amp;ldquo;libertarian paternalism&amp;rdquo; (Thaler and Sunstein 2003) is domain-specific and is strongest precisely where intertemporal adjustment is possible.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Opt-out cost (k).&lt;/strong&gt; In this paper, a utility cost — estimated at $254 per contribution-rate change — that individuals must pay every time they choose a retirement contribution rate different from the current default. The cost is modeled as a consumption reduction and captures both real transaction costs (form-filling, adviser fees) and behavioral costs (cognitive cost of attention and optimal-choice search). It is fixed and homogeneous across individuals, and applies symmetrically in any direction of deviation from the default.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Auto-enrollment default contribution rate.&lt;/strong&gt; The positive contribution rate at which new hires are automatically enrolled in a defined-contribution plan, with the option to opt out by incurring the opt-out cost. In the paper&amp;rsquo;s estimation sample, this is 3% of salary. The default is exogenous at the start of each new job but endogenous thereafter: once established, the default for subsequent periods equals the worker&amp;rsquo;s contribution rate in the previous period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Default eﬀect.&lt;/strong&gt; The empirically observed tendency of workers to remain at the default contribution rate rather than actively choosing a different rate. In this paper, the default effect is explained by opt-out costs rather than loss aversion or psychological anchoring — a distinction identified through the novel prediction that raising the default from a positive rate to a higher positive rate reduces overall participation (the &amp;ldquo;drop-out&amp;rdquo; effect), a pattern consistent only with opt-out costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Drop-out eﬀect.&lt;/strong&gt; The paper&amp;rsquo;s term (following Caplin and Martin 2017) for the empirical finding that increasing the auto-enrollment default contribution rate causes some workers to stop contributing altogether or to contribute at rates strictly below the initial default. This effect is used as a discriminating test between competing theories of the default effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic opt-out cost framework.&lt;/strong&gt; The paper&amp;rsquo;s core modeling insight: that opt-out costs must be estimated in a fully dynamic lifecycle model that allows workers to adjust contributions over time, to hold liquid assets and unsecured debt, and to face labor market risk. In a static or short-horizon model, the opportunity cost of initial non-participation appears large (because the worker permanently forgoes match and tax benefits), requiring large opt-out costs. In the dynamic model, the ability to compensate later shrinks the implied opportunity cost and hence the opt-out cost required to rationalize observed inertia.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crowd-out of liquid savings.&lt;/strong&gt; The extent to which higher DC retirement contributions induced by auto-enrollment reduce liquid asset holdings (or increase unsecured borrowing), rather than increasing total wealth. The paper estimates limited crowd-out (89% pass-through to total wealth for bottom-quintile workers, 62% for middle-quintile workers), attributable to the different roles of liquid assets (precautionary motive) and DC accounts (lifecycle motive) in the model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy incidence.&lt;/strong&gt; The channel through which employers balance their budget in response to higher matching costs created by auto-enrollment. The paper considers three scenarios: employers absorb costs through reduced profits; employers reduce the match rate; employers reduce wages. Optimal policy rankings and welfare magnitudes differ across these scenarios, but the qualitative conclusions — utilitarian policymaker prefers opt-in; paternalistic or inequality-averse policymaker prefers AE at 6% — are robust across incidence assumptions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Consumption-equivalent variation (γ).&lt;/strong&gt; The welfare metric used in the paper: the proportional increase in consumption in every period and every state of the world that would make the policymaker indifferent between an auto-enrollment policy at default d and the opt-in regime. A 6% default increases welfare by 0.3% in consumption-equivalent for the bottom income decile under a utilitarian policymaker when incidence is on employers.&lt;/p&gt;</description></item><item><title>Defying Distance? The Provision of Medical Services in the Digital Age</title><link>https://macropaperwarehouse.com/papers/defying-distance-the-provision-of-medical-services-in-the-digital-age/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/defying-distance-the-provision-of-medical-services-in-the-digital-age/</guid><description>&lt;p&gt;This paper asks whether digital platforms can improve healthcare outcomes by enabling needs-based matching between patients and physicians unconstrained by geography. Amanda Dahlstrand studies digital primary care in Sweden during 2016-2018, exploiting nationwide conditional random assignment between approximately 200,000 patients and 143 doctors employed by Europe&amp;rsquo;s largest digital primary care provider. Patients who selected the &amp;ldquo;first available doctor&amp;rdquo; option (82% of first visits) were effectively randomized to a doctor within each 3-hour shift-by-date stratum, generating quasi-experimental variation free of the patient-doctor sorting that confounds identification in physical primary care.&lt;/p&gt;
&lt;p&gt;The paper defines three observable dimensions of primary care physician skill: (1) identifying risky patients and triaging them to higher levels of care, measured by whether patients subsequently have an avoidable hospitalization within 90 days; (2) providing guideline-consistent treatment, measured by counter-guideline antibiotic prescriptions; and (3) leaving patients sufficiently informed so they do not unnecessarily seek additional in-person care within the following week. Doctor skill in each dimension is estimated via a value-added framework in a hold-out sample (Sample 1, the first 600 randomized consultations per doctor), using empirical Bayes shrinkage to reduce noise. Complementarities between doctor skill and patient risk are then estimated in a disjoint main sample (Sample 2).&lt;/p&gt;
&lt;p&gt;A central finding is that doctor skill is task-specific rather than governed by a single latent ability: skills across the three tasks are not positively correlated, meaning doctors within general practice have individual &amp;ldquo;specializations.&amp;rdquo; A patient ranked in the top 1% of avoidable hospitalization risk who is matched to a doctor ranked in the top 10% at reducing avoidable hospitalizations experiences a 90% reduction in that adverse outcome, relative to a patient with the same risk profile matched to the worst-performing doctor. Patients not estimated as risky show effects indistinguishable from zero when matched to the same high-skilled doctors, establishing a strong complementarity between doctor type and patient risk.&lt;/p&gt;
&lt;p&gt;Using the Average Match Function framework of Graham, Imbens, and Ridder (2014, 2020), the paper evaluates counterfactual reallocation policies. Reallocating only 2% of patients — those in the top 1% of predicted avoidable hospitalization risk — to doctors in the top 10% of triage skill reduces aggregate avoidable hospitalizations by 20% relative to random assignment, without adversely affecting counter-guideline prescriptions or other measured outcomes. Doctor skills across outcomes are not positively correlated, so this reallocation does not generate meaningful trade-offs. The paper benchmarks this matching policy against a selective hiring/expansion policy in which doctors with above-median skill in three tasks expand their hours by up to 70% at the expense of below-median peers; that policy yields no significant reduction in avoidable hospitalizations and only a 4% reduction in counter-guideline prescriptions — smaller gains than matching and harder to implement.&lt;/p&gt;
&lt;p&gt;The paper also documents that physical primary care quality is worse in lower-income and more deprived areas of Sweden (a negative relationship between deprivation index and patient-reported experience is statistically significant at the 1% level in a cross-section of roughly 120-150 primary care centers in Region Skane). Because the estimated risk of avoidable hospitalization and prior avoidable hospitalizations are concentrated in the lower end of the income distribution, needs-based digital matching reallocates triage skill toward lower-income patients, severing the correlation between local area income and service quality. Simulating positive assortative matching on patient income and doctor skill — approximating existing healthcare inequalities — leads to more avoidable hospitalizations than random assignment, because the most vulnerable patients tend to be the poorest. Scope conditions: findings derive from a single digital primary care provider in Sweden, 2016-2018, pre-pandemic, covering conditions amenable to video consultation and a patient pool younger and somewhat more urban than the average Swedish citizen.&lt;/p&gt;
&lt;p&gt;Q: What is the key identification strategy, and why is it valid in this setting but not in physical primary care?
A: Patients who selected the &amp;ldquo;drop in&amp;rdquo; (first available doctor) option — 82% of first visits — were assigned to whichever certified doctor was next in the roster within a 3-hour shift-by-date stratum, a by-product of the first-come-first-served queue. Neither patients nor doctors could intervene in this digital process. The author validates the assumption by regressing doctor characteristics on patient characteristics controlling for shift-by-date fixed effects and finds characteristics are balanced. In physical primary care, endemic patient-doctor sorting means doctors do not meet a common support of patient types, preventing causal identification of doctor effects.&lt;/p&gt;
&lt;p&gt;Q: How are doctor skill estimates constructed and why does the split-sample matter?
A: Doctor skill in each task is estimated as an empirical Bayes-shrunk random effect from a value-added regression on Sample 1, each doctor&amp;rsquo;s first 600 randomized consultations (40% of the sample). Sample 2 (60%) is entirely disjoint and used to estimate complementarities between doctor skill and patient risk. The split-sample design prevents overfitting: doctor skill was estimated on different patients than those in Sample 2. The Durbin-Wu-Hausman test does not reject random effects (p = 0.16).&lt;/p&gt;
&lt;p&gt;Q: What is the main quantitative result on avoidable hospitalization matching?
A: A patient ranked in the top 1% of predicted avoidable hospitalization risk matched to a doctor ranked in the top 10% at reducing avoidable hospitalizations could reduce that patient&amp;rsquo;s avoidable hospitalizations by 90%, relative to the worst-performing doctor in that skill. At the aggregate level, reallocating only 2% of patients (those in the top 1% risk group) to high-triage-skill doctors reduces avoidable hospitalizations across the full patient population by 20% compared to random assignment.&lt;/p&gt;
&lt;p&gt;Q: Does the avoidable hospitalization reallocation harm other outcomes?
A: No. The paper explicitly evaluates the Average Reallocation Effect on counter-guideline prescriptions and additional in-person care seeking when optimizing for avoidable hospitalizations, and finds no significant adverse effects on these other outcomes. The author attributes this to the fact that doctor skills across tasks are not positively correlated, so reallocating triage-skilled doctors does not systematically remove skill from other dimensions.&lt;/p&gt;
&lt;p&gt;Q: How does matching compare to selective hiring and hour expansion as a policy?
A: Even expanding the working hours of doctors with above-median skill across three tasks by as much as 70% yields no significant reduction in avoidable hospitalizations and only a 4% reduction in counter-guideline prescriptions — both smaller gains than the matching policy. Matching outperforms hiring expansion because patients have heterogeneous needs that can be identified from prior healthcare records, and doctors have differentiated skill sets relevant to some patients but not others.&lt;/p&gt;
&lt;p&gt;Q: What is the evidence that doctor skills are task-specific rather than reflecting a single latent ability?
A: The estimated doctor effects across the three tasks — triaging to avoid hospitalizations, guideline-consistent antibiotic prescribing, and minimizing unnecessary follow-up care — are not positively correlated with one another. This means a doctor who is effective at one task is not systematically effective at others, indicating individual specializations within general practice that are not accounted for in standard primary care organization.&lt;/p&gt;
&lt;p&gt;Q: How is patient risk for avoidable hospitalizations measured?
A: A propensity score is estimated from pre-digital physical healthcare data (2013-2015), regressing past number of avoidable hospitalizations on demographic and healthcare utilization variables — including age, a disease index of chronic diagnoses, and previous hospitalizations — all variables already available in patient medical records. The top 1% of predicted risk scores are classified as &amp;ldquo;risky.&amp;rdquo; Patients in the risky group had on average 0.35 avoidable hospitalizations in the prior 3 years, versus 0.01 for non-risky patients.&lt;/p&gt;
&lt;p&gt;Q: What is the distributional (equity) implication of needs-based matching versus income-assortative matching?
A: Estimated risk of avoidable hospitalization and the count of prior avoidable hospitalizations are concentrated in the lower end of the income distribution. Needs-based matching therefore reallocates triage skill toward lower-income patients. Simulating positive assortative matching on patient income and doctor skill — approximating observed inequalities in physical care — produces more avoidable hospitalizations than random assignment, because the most vulnerable patients are often the poorest. Needs-based digital matching can sever the link between local area income and service quality.&lt;/p&gt;
&lt;p&gt;Q: How does digital care usage sort by income and demographics in the data?
A: At the extensive margin, the deprivation index (Care Need Index) is similar among digital users and non-users in Region Skane. However, at the intensive margin, individuals with a higher deprivation index who use the digital service have more appointments in it; similarly, lower-income users use the service more intensively. Digital care users are younger than non-users and are more likely to live in cities than the average Swedish citizen.&lt;/p&gt;
&lt;p&gt;Q: What are avoidable hospitalizations and why are they the primary outcome?
A: Avoidable hospitalizations (also called hospitalizations for ambulatory care sensitive conditions) are hospital admissions defined in the medical literature as preventable by adequate and timely primary care. They are coded using ICD-10 diagnosis codes listed in Page et al. (2007). The most common diagnoses in the 90-day post-consultation window are respiratory and genitourinary, conditions commonly treated in digital care. The outcome is rare (0.2% of patients in the sample), but high-stakes: an estimated 1.1 potential life years are lost per avoidable hospitalization, and in Sweden they cost an estimated SEK 7.1 billion (~$820 million) annually (7% of inpatient curative and rehabilitative care costs).&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the counter-guideline antibiotic prescription outcome?
A: Non-adherence is coded against 16 guidelines from Sweden&amp;rsquo;s strategic programme against antibiotic resistance (Strama 2017, 2019), all designed to limit or narrow antibiotic use. The measured rate of non-adherence is described as quite low by international standards; the CDC estimates 28% of US antibiotic prescriptions are unnecessary, while the author&amp;rsquo;s sample rate is 2%. The guidelines require doctors to sometimes refuse patients who request antibiotics, introducing a behavioral compliance dimension to this skill.&lt;/p&gt;
&lt;p&gt;Q: What are the costs and feasibility considerations for implementing needs-based digital matching?
A: The paper characterizes matching as a &amp;ldquo;resource-neutral&amp;rdquo; policy because it reallocates existing doctors without hiring or training. The primary costs are a small increase in waiting time for some patients and the costs of importing data and developing the matching algorithm. Because the algorithm handles patient-doctor allocation while doctors retain all clinical decision-making, the policy functions as a complement to human skill rather than a substitute, which the author argues makes it less subject to &amp;ldquo;algorithm aversion.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Q: Why does the paper restrict to each patient&amp;rsquo;s first digital consultation only?
A: The first visit is the one subject to conditional random assignment; subsequent visits could reflect endogenous selection by patients who preferred a particular doctor or outcome. Using only first visits eliminates this concern. The restriction reduces the sample from approximately 378,000 to 210,171 patients (56% of the original), paired with 143 doctors who each had at least 600 randomized consultations.&lt;/p&gt;
&lt;p&gt;Conditional random assignment: The allocation mechanism by which patients selecting the &amp;ldquo;first available doctor&amp;rdquo; option in digital primary care were assigned to whichever certified doctor was next in the shift roster, conditional on 3-hour shift-by-date strata — a by-product of the first-come-first-served queue rather than an intended experimental design.&lt;/p&gt;
&lt;p&gt;Average Match Function (AMF): The conditional mean of a patient outcome given observable doctor type and patient type under random assignment, β(x,w) = E[Y|X=x, W=w], which serves as the building block for evaluating counterfactual reallocation policies.&lt;/p&gt;
&lt;p&gt;Average Reallocation Effect (ARE): The difference in expected patient outcomes between a counterfactual doctor-patient assignment and the status quo random assignment, taking into account the externality on the patient from whom a high-skilled doctor is moved.&lt;/p&gt;
&lt;p&gt;Task-specific doctor skill: The paper&amp;rsquo;s finding that primary care physician effectiveness is not governed by a single latent ability but varies across distinct tasks — triage/risk prediction, guideline-consistent prescribing, and minimizing unnecessary follow-up care — with skills across tasks not positively correlated.&lt;/p&gt;
&lt;p&gt;Avoidable hospitalization: A hospital admission coded to a diagnosis (per Page et al. 2007 ICD-10 classification) defined in the medical literature as preventable by adequate and timely primary care, used as the primary high-stakes outcome measure (0.2% incidence in the sample within 90 days of a digital consultation).&lt;/p&gt;
&lt;p&gt;Counter-guideline prescription: A prescription of an antibiotic in violation of one of 16 guidelines from Sweden&amp;rsquo;s Strama antibiotic resistance programme, all of which are designed to limit use or require narrower-spectrum first-line antibiotics; used as the primary guideline-adherence outcome (2% incidence in the sample).&lt;/p&gt;
&lt;p&gt;Empirical Bayes shrinkage: A procedure applied to raw doctor value-added estimates in which the noisy estimate of doctor quality is multiplied by the ratio of signal variance to total (signal plus noise) variance, yielding a best linear predictor of the underlying doctor random effect and reducing noise from small-sample estimation.&lt;/p&gt;</description></item><item><title>Demand Analysis under Latent Choice Constraints</title><link>https://macropaperwarehouse.com/papers/demand-analysis-under-latent-choice-constraints/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/demand-analysis-under-latent-choice-constraints/</guid><description>&lt;p&gt;Agarwal and Somaini study demand estimation in markets where consumers face latent choice constraints — situations where a consumer&amp;rsquo;s effective choice set is determined not only by her preferences but also by supply-side rationing or information frictions that restrict which options are actually available to her. Standard discrete choice methods assume consumers pick freely from the full product set, but this assumption fails in school and college admissions, entry-level labor markets, healthcare with selective admissions, and consumer markets with incomplete consideration sets. The paper provides a unified non-parametric identification framework for this class of models, proves necessity of the identifying instruments, proposes a computationally tractable estimator, and applies the framework to the California kidney dialysis market.&lt;/p&gt;
&lt;p&gt;The model combines a general random utility specification — accommodating multi-dimensional unobserved heterogeneity and product-level unobservables correlated with observed characteristics as in Berry (1994) and BLP (1995) — with a reduced-form acceptance policy function that governs which products accept which consumers. The consumer&amp;rsquo;s latent choice set is the set of products that accept her, and she picks her most preferred option within that set. Crucially, the acceptance decision may be arbitrarily correlated with consumer preferences, ruling out the independence assumptions common in the consideration-set literature.&lt;/p&gt;
&lt;p&gt;Identification rests on two sets of instruments. The first is a preference shifter, a consumer-product observable that affects utility but is excluded from the acceptance policy — distance to facility in the application. The second is a choice-set shifter, an observable that affects the acceptance decision but is excluded from consumer utility — short-term deviation of a facility&amp;rsquo;s caseload from its estimated target in the application. The main result (Theorem 1) establishes non-parametric point identification of the joint distribution of indirect utilities and acceptance decisions given both instruments. Proposition 1 establishes that the model is not identified when the choice-set shifter is absent — even when the preference shifter has full support — making both instruments necessary rather than merely sufficient.&lt;/p&gt;
&lt;p&gt;The application uses USRDS data on 41,913 new dialysis patients treated at 552 California facilities between 2015 and 2018. Most facilities are owned by Fresenius or DaVita. The choice-set shifter is the facility&amp;rsquo;s caseload deviation from target when a patient enters the market; facility and quarter fixed effects are included so that only short-term caseload variation drives identification. A reduced-form regression shows that higher caseload deviation significantly reduces the inflow of new patients to a facility, consistent with supply-side rationing. Patients also choose more distant facilities when nearby facilities have above-normal caseloads, providing further reduced-form evidence that rationing shapes allocations.&lt;/p&gt;
&lt;p&gt;A Gibbs sampler with data augmentation — drawing alternately from the distribution of latent choice sets conditional on utilities and from utility parameters conditional on choice sets — circumvents the curse of dimensionality that makes direct likelihood maximization over all possible choice sets infeasible.&lt;/p&gt;
&lt;p&gt;Estimation results show that the probability a patient is accepted at her first-choice facility is only 73.0%, with variation across facilities. Standard discrete choice models that ignore rationing misestimate facility quality, systematically assigning high desirability to low-caseload facilities in a manner that conflates easy access with genuine patient preference. A naive correction that includes the caseload measure in the utility function mischaracterizes the diversion pattern: rationed patients are marginal for the facility but strictly prefer it, so they divert differently from patients who voluntarily switch because of quality changes. Fresenius and DaVita facilities are estimated to be more selective than independent facilities, consistent with chain networks enabling coordinated patient-flow management across locations.&lt;/p&gt;
&lt;p&gt;Q: What is the core empirical problem the paper addresses?
A: Standard demand estimation inverts market shares to recover preference parameters under the assumption that consumers choose freely from the full product set. When choice sets are constrained by supply-side rationing or information frictions, the largest market share product need not be the one most preferred — it may simply be the one that accepts the most consumers. This makes the standard inversion inapplicable, and ignoring constraints yields biased preference estimates.&lt;/p&gt;
&lt;p&gt;Q: What does the paper&amp;rsquo;s model consist of?
A: The model has two components: (1) a random utility model for consumer preferences with rich observed and unobserved heterogeneity, allowing product-level unobservables correlated with observed characteristics; and (2) a reduced-form acceptance policy function sigma_jt taking values in {0,1} that determines whether product j accepts consumer i. The consumer&amp;rsquo;s latent choice set is the set of products that accept her; she picks her most preferred option within it. Utilities and acceptance decisions may be arbitrarily correlated.&lt;/p&gt;
&lt;p&gt;Q: What examples of latent choice constraints are covered by the framework?
A: The reduced form encompasses: selective admissions in healthcare (facility accepts patient if profitability exceeds a caseload-dependent threshold); two-sided matching markets where a pairwise stable allocation is described by cutoff scores (school admissions, entry-level labor markets); consideration set models where brand awareness advertising or inattention determines which products a consumer sees; fixed-sample consumer search; and product stock-outs. Each of these implies an acceptance policy function of the form specified in the paper&amp;rsquo;s reduced-form model.&lt;/p&gt;
&lt;p&gt;Q: What are the two identifying instruments and the intuition behind each?
A: The preference shifter yij is a consumer-product observable that affects the consumer&amp;rsquo;s indirect utility for product j but is excluded from that product&amp;rsquo;s acceptance decision. In the application this is distance: dialysis requires multiple weekly visits, so distance affects patient utility, but a facility&amp;rsquo;s decision to accept a patient does not depend on how far the patient lives. The choice-set shifter zij is an observable that affects the acceptance decision but is excluded from consumer preferences. In the application this is the deviation of facility caseload from its estimated target: short-term caseload swings affect whether a facility can take a new patient but, conditional on facility fixed effects, do not reflect facility quality as perceived by patients.&lt;/p&gt;
&lt;p&gt;Q: What does Theorem 1 establish and under what conditions?
A: Theorem 1 establishes non-parametric point identification of (i) the function gj mapping the preference shifter to its utility contribution, and (ii) the joint distribution of indirect utilities and acceptance indicators, for every consumer attribute vector and every value in the interior of the joint support of the instruments. Conditions required include: monotonicity of the acceptance policy in the choice-set shifter (higher z makes acceptance weakly less likely, with sigma=1 as z approaches negative infinity and sigma=0 as z approaches positive infinity); conditional independence of unobservables from the instruments given observed consumer attributes; and at least two products available.&lt;/p&gt;
&lt;p&gt;Q: What does Proposition 1 establish about necessity of the choice-set shifter?
A: Proposition 1 shows that if the choice-set shifter z has singleton support (no variation), then even when the preference shifter g has full support on R^|J|, the distribution of preferences is not identified wherever a choice set strictly smaller than the full product set has positive probability. The non-identification result applies on any open set where a constrained choice set has positive probability — it is not a knife-edge case. This makes the choice-set shifter a necessary condition for identification, not merely a convenient one.&lt;/p&gt;
&lt;p&gt;Q: How does the paper handle endogeneity of product characteristics?
A: Corollary 2 extends the baseline identification result to allow product-level unobservables that may be correlated with observed product characteristics, as in Berry (1994) and BLP (1995). Identification in this case requires an additional instrument that shifts product characteristics but is excluded from both preferences and choice sets — analogous to BLP supply-side instruments — alongside the two shifters already required. This extends Berry and Haile (2010) to settings with constrained choice sets.&lt;/p&gt;
&lt;p&gt;Q: What is the Gibbs sampler estimator and why is it needed?
A: With J products per market, the number of possible choice sets is 2^J, making direct likelihood computation infeasible for even moderate J. The Gibbs sampler uses data augmentation to alternate between: (a) drawing latent choice sets conditional on current utility parameters and observed choices; and (b) drawing utility parameters conditional on the augmented choice sets. Each conditional draw reduces to a standard problem, avoiding the curse of dimensionality. The Bernstein-von Mises theorem implies that the posterior mean of the sampling chain is asymptotically equivalent to the maximum likelihood estimator.&lt;/p&gt;
&lt;p&gt;Q: What is the reduced-form evidence for supply-side rationing in dialysis?
A: The regression of log(1 + new patient inflows to facility j in quarter q) on facility fixed effects, quarter fixed effects, and the caseload deviation z_jq yields a statistically significant negative coefficient on caseload deviation: above-target caseloads reduce new patient admissions even after controlling for facility-level and time-level averages. Additionally, patients whose nearest facilities have above-normal caseloads travel to more distant facilities, providing complementary evidence that rationing displaces patients geographically.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated probability of acceptance at a first-choice facility?
A: The structural estimates imply that a patient is accepted at her first-choice facility with probability only 73.0%, with variation across facilities. The implied 27.0% rejection rate is economically substantial, meaning a large share of observed allocations do not reflect unconstrained patient preference.&lt;/p&gt;
&lt;p&gt;Q: How do estimates from the constrained model differ from a standard discrete choice model?
A: The standard model, which ignores selective admissions, assigns higher utility to facilities with lower caseloads — a bias that conflates easy access with genuine patient preference. The constrained model separately identifies the facility&amp;rsquo;s acceptance propensity from the patient&amp;rsquo;s underlying preference, yielding different facility quality rankings. The largest facilities are not necessarily the most desirable once selective admissions are accounted for.&lt;/p&gt;
&lt;p&gt;Q: Why is the naive correction — including caseload in the utility function — insufficient?
A: The naive correction treats caseload as a quality attribute, implying that a patient turned away because of high caseload and a patient who voluntarily avoids a high-caseload facility are pulled from the same margin. In the constrained model, a rationed patient is marginal for the facility but strictly prefers it, so she diverts to a different set of alternatives than a patient who voluntarily switches. Not capturing this distinction produces quantitatively different diversion ratios.&lt;/p&gt;
&lt;p&gt;Q: What do the estimates say about chain versus independent facilities?
A: Fresenius and DaVita facilities are estimated to be more selective in their admissions than independent facilities. The paper interprets this as consistent with large chains having better ability to coordinate patient flows across their network of facilities, potentially directing turned-away patients to other chain locations.&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the identification results?
A: Identification is established within each market, for consumer attribute vectors in the interior of support, and for utility-acceptance pairs in the interior of the joint support of the instruments. The results are non-parametric in that they do not restrict the functional form of preferences or acceptance policies beyond monotonicity and support conditions, and they allow unobservables affecting choice sets to be arbitrarily correlated with preference unobservables. The empirical application implements a parametric version for tractability.&lt;/p&gt;
&lt;p&gt;Latent choice constraint: A restriction on a consumer&amp;rsquo;s effective choice set arising from supply-side rationing or information frictions, such that the consumer can only choose among the products that accept her rather than freely among all products in the market. Distinct from price-based market clearing.&lt;/p&gt;
&lt;p&gt;Acceptance policy function: A reduced-form function mapping consumer attributes, consumer unobservables, and the choice-set shifter to a binary accept/reject decision by product j. Indexed by product and market, allowing arbitrary variation in selectivity across products and time. The consumer&amp;rsquo;s latent choice set is defined as the set of products whose acceptance policy equals 1.&lt;/p&gt;
&lt;p&gt;Choice-set shifter: A consumer-product observable that shifts the acceptance probability — making product j more or less likely to accept consumer i — while being excluded from consumer indirect utility. In the application: short-term deviation of facility caseload from its estimated target. Necessary (not merely sufficient) for non-parametric identification of the model.&lt;/p&gt;
&lt;p&gt;Preference shifter: A consumer-product observable that shifts consumer utility for product j and is separable from consumer-specific unobservables, but is excluded from that product&amp;rsquo;s acceptance policy function. In the application: distance from patient&amp;rsquo;s residence to the facility. Also necessary for identification.&lt;/p&gt;
&lt;p&gt;Curse of dimensionality in constrained choice: The computational problem that the number of possible latent choice sets grows as 2^J with the number of products J, making direct likelihood integration over choice sets infeasible for even moderate J. Resolved in this paper by a Gibbs sampler with data augmentation that conditions alternately on latent choice sets or utility parameters.&lt;/p&gt;
&lt;p&gt;Diversion ratio under selective admissions: The share of patients lost by a facility who are captured by each alternative facility. In a model with selective admissions, rationed patients (marginal for the facility) divert differently from patients who voluntarily switch (marginal for the consumer), because rationed patients strictly prefer the rejecting facility. The naive correction conflates these two margins, yielding quantitatively different and biased diversion ratio estimates.&lt;/p&gt;
&lt;p&gt;Non-parametric necessity of instruments: The property that both the preference shifter and the choice-set shifter are individually necessary conditions for point identification of the joint distribution of preferences and acceptance decisions, not merely convenient sufficient conditions. Absence of either instrument leaves the model non-identified on any open set where a constrained choice set has positive probability.&lt;/p&gt;</description></item><item><title>Designing Dynamic Reassignment Mechanisms: Evidence from GP Allocation</title><link>https://macropaperwarehouse.com/papers/designing-dynamic-reassignment-mechanisms-evidence-from-gp-allocation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/designing-dynamic-reassignment-mechanisms-evidence-from-gp-allocation/</guid><description>&lt;p&gt;This paper studies the design of dynamic reassignment mechanisms—centralized systems that must not only provide good initial matches but also accommodate changes in agents&amp;rsquo; preferences over time. The empirical setting is Norway&amp;rsquo;s system for allocating patients to general practitioners (GPs), where every individual is assigned a specific GP whose panel has a binding capacity cap. Since 2016, Norway has allowed patients to join waitlists for oversubscribed GPs while retaining their spot on their current GP&amp;rsquo;s panel, with reassignment proceeding strictly first-come, first-served (FCFS) as vacancies arise.&lt;/p&gt;
&lt;p&gt;The paper makes three contributions. First, it provides direct evidence of unrealized gains from trade: in December 2019, 15 percent of the 133,332 patients then standing on waitlists could have been immediately reassigned via a single run of the Top-Trading Cycles (TTC) algorithm, which identifies not only bilateral swaps but arbitrary cycles. A mechanical simulation holding patient choices fixed shows that running TTC monthly from November 2016 through December 2019 would have left 23 percent fewer patients on waitlists by end-2019, with average waiting times among reassigned patients 29 percent shorter.&lt;/p&gt;
&lt;p&gt;Second, the paper introduces a dynamic TTC mechanism and clarifies why static properties do not carry over. In the static case, TTC is both strategy-proof and Pareto-improving (Shapley and Scarf, 1974; Roth, 1982). In a dynamic setting, neither property holds. Repeated TTC is not strategy-proof because patients&amp;rsquo; GP choices affect how long they wait. More importantly, TTC may leave some patients worse off: a panel slot that would have gone to the first person on a waitlist under FCFS may instead go to a later-arriving patient who can form a trading cycle, effectively de-prioritizing patients whose GPs are undersubscribed. In the mechanical simulation, 4.5 percent of patients face longer waiting times under TTC.&lt;/p&gt;
&lt;p&gt;Third, the paper estimates a structural model of patient attention and GP choice using monthly Norwegian administrative data covering 4.78 million patients and 6,470 GP panels (2014–2019), restricting estimation to the Trondelag region (approximately 8 percent of the country). The model specifies: a Poisson attention process (patients consider switching only when an attention shock arrives); preferences over GPs as a function of travel time, GP fixed effects, and match characteristics; and a belief model mapping observed waitlist lengths into expected waiting times. Parameters are recovered via a Gibbs sampler with Metropolis-Hastings for the discount rate. Key estimates: the annual discount factor is approximately 0.91; a female patient under 45 would travel 7.3 minutes farther to see a female GP (6.3 minutes for a female patient over 45); GP fixed effects have a standard deviation of 31 minutes&amp;rsquo; travel-time equivalent; idiosyncratic taste shocks have a standard deviation of 12.6 minutes.&lt;/p&gt;
&lt;p&gt;The paper then simulates a stationary equilibrium for each counterfactual mechanism. Under the status quo in stationary equilibrium, 9.4 percent of patients are on a waitlist, 82.2 percent of GPs have a waitlist, and average expected waiting time is 16.7 months. Introducing TTC reduces average waiting time to 14.1 months and raises mean patient welfare by the equivalent of 0.75 minutes&amp;rsquo; travel time (more than 13 percent of the gain achievable under a no-capacity-constraints benchmark). Over half of this gain (0.4 minutes) comes directly from patients obtaining geographically closer GPs. Benefits are concentrated among younger patients, female patients, and recent movers; rural patients gain 2.1 minutes. However, patients with undersubscribed GPs face waiting times that rise from 16.7 to 22.8 months and are worse off by the perpetuity equivalent of 0.8 minutes.&lt;/p&gt;
&lt;p&gt;Two modified mechanisms are evaluated. Deferred Acceptance (DA), which strictly respects FCFS priority, achieves essentially no improvement over the status quo, illustrating a fundamental trade-off between eliminating envy and exploiting gains from trade. A &amp;ldquo;TTC with Priority&amp;rdquo; (TTCP) mechanism, which gives priority for panel vacancies to patients with undersubscribed GPs before running TTC, achieves 61 percent of TTC&amp;rsquo;s welfare gains (0.46 minutes flow payoff; 1.08 minutes NPV) while leaving patients with undersubscribed GPs no worse off than under the status quo. A benchmark simulation eliminating waitlists altogether raises mean welfare slightly (0.19 minutes) but lowers median welfare (−0.60 minutes), with gains concentrated among highly mismatched patients.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the core market failure the paper documents?&lt;/strong&gt;
A: Norway&amp;rsquo;s waitlist mechanism assigns panel vacancies strictly first-come, first-served without allowing patients to trade. This creates a &amp;ldquo;double coincidence of wants&amp;rdquo; problem: patients can simultaneously be on each other&amp;rsquo;s waitlists but cannot swap. In December 2019, 15 percent of 133,332 waiting patients could have been immediately reassigned via a single TTC run. A mechanical simulation shows that monthly TTC would have left 23 percent fewer patients on waitlists by end-2019 and reduced average realized waiting times among reassigned patients by 29 percent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does TTC fail to be strategy-proof in a dynamic setting?&lt;/strong&gt;
A: In the static case, TTC gives every agent an assignment at least as good as their endowment, making truthful reporting a dominant strategy. In a dynamic setting, a patient&amp;rsquo;s choice of GP determines not only which GP they receive but also how long they wait — patients who choose less-demanded GPs reach the front of the waitlist faster. This creates incentives to misreport preferences strategically, breaking strategy-proofness. The paper shows this formally and builds it into the equilibrium model by requiring patients to optimize over both GP choice and expected waiting time.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does dynamic TTC harm some patients relative to the status quo?&lt;/strong&gt;
A: Under FCFS, the first person on a waitlist is guaranteed the next available slot on the target GP&amp;rsquo;s panel. Under TTC, a patient who arrived later but whose current GP is oversubscribed can form a trading cycle that redirects that slot, effectively jumping the queue. Patients with undersubscribed GPs — whose panel endowment is not a scarce resource that others want — cannot form cycles and are systematically de-prioritized. In the stationary equilibrium, their expected waiting time rises from 16.7 to 22.8 months, and they are worse off by the perpetuity equivalent of 0.8 minutes&amp;rsquo; travel time.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the main parameter estimates and what do they imply?&lt;/strong&gt;
A: The annual discount factor is estimated at approximately 0.91 once GP fixed effects are included (rising to near 0.95 without them, because more desirable GPs have longer waitlists). Gender homophily is worth 6.3–7.3 minutes of travel time for female patients under 45. Age homophily is worth approximately 1 minute. The standard deviation of GP fixed effects is 31 minutes and idiosyncratic shocks are 12.6 minutes, both in travel-time equivalents, indicating substantial horizontal differentiation across GPs and across patients&amp;rsquo; idiosyncratic tastes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How important are moves as a driver of GP switching?&lt;/strong&gt;
A: Moves are the dominant driver. Among non-movers, older men consider switching just once every 25 years; temporary residents consider switching approximately once every 7.5 years (1.084 percent per month). Among patients who moved more than 30 minutes, a temporary resident has an 18.59 percent monthly probability of considering switching in the month of or month after the move. For a permanent resident making a long-distance move, the cumulative attention probability over the 8 months surrounding the move rises to 34 percent (versus 22 percent for a short-distance move). In the data, 26 percent of waitlist users moved municipality during 2017–2019, versus 6 percent of non-switchers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the stationary equilibrium under the status quo look like?&lt;/strong&gt;
A: In the long-run stationary equilibrium, 9.4 percent of patients are on a waitlist, 82.2 percent of GPs have a waitlist, and the average expected waiting time to switch GPs is 16.7 months. Each month, 2,299 patients on average draw attention shocks; 85.2 percent of these choose to join a waitlist, while the remainder either switch to an open GP or stay with their current GP. The average attentive patient expects to successfully obtain their chosen GP after 16.8 months.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the distributional consequences of TTC across patient subgroups?&lt;/strong&gt;
A: Female patients benefit especially because they are more likely to be attentive (and thus use waitlists) than males. Recent movers gain 2.3 minutes&amp;rsquo; travel-time equivalent. Patients who have never moved still gain 1.0 minutes. Rural patients gain 2.1 minutes (larger than average), reflecting their longer baseline travel times and greater geographic mismatch potential. Urban patients also benefit but less so. The one group that is harmed is patients with undersubscribed GPs, who face longer waits and a welfare loss of 0.8 minutes perpetuity equivalent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does the Deferred Acceptance mechanism fail to improve on the status quo?&lt;/strong&gt;
A: DA strictly respects FCFS waiting-time priority: no patient may be reassigned to a GP for whom another patient has been waiting longer. This means DA can only execute swaps in which all patients ahead of each participant on their respective waitlists are also reassigned in the same month. In practice, this virtually never occurs, so DA reassigns almost no patients earlier than the status quo Waitlists mechanism. The result illustrates a fundamental trade-off: fully respecting FCFS priority eliminates nearly all gains from trade.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does TTCP restore fairness while preserving most of the efficiency gains?&lt;/strong&gt;
A: TTCP modifies TTC by prioritizing patients with undersubscribed GPs over those with oversubscribed GPs when assigning panel vacancies, while still respecting the constraint that patients cannot be assigned a GP they prefer less than their current one. This gives patients with undersubscribed GPs a compensating advantage in the queue that offsets their inability to trade via cycles. TTCP achieves 0.46 minutes&amp;rsquo; mean flow payoff improvement versus 0.75 for TTC (61 percent of TTC&amp;rsquo;s gains), and an NPV measure of 1.08 minutes versus 1.25 for TTC. Patients with undersubscribed GPs are left no worse off than under the status quo.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What happens when waitlists are eliminated entirely?&lt;/strong&gt;
A: Under No Waitlists, attentive patients may only choose among GPs with open panels at the moment of attention. Mean welfare rises slightly (0.19 minutes) because patients spend less time mismatched while waiting, but median welfare falls by 0.60 minutes. The gains are concentrated among a minority of highly mismatched patients who prefer limited choice with no waiting over broader choice with long waits, while most patients prefer the option to wait for a more preferred GP. The authors note this may partly explain why formal waitlists are rare in other primary care systems.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the welfare benchmark and how large are the gains?&lt;/strong&gt;
A: The benchmark is a &amp;ldquo;No Caps&amp;rdquo; scenario in which all panel caps are removed, representing the maximum achievable improvement. The mean welfare gain from TTC (0.75 minutes) represents more than 13 percent of this upper bound. The &amp;ldquo;Truthful TTC&amp;rdquo; benchmark, where patients submit full preference lists, yields 1.04 minutes, but its gains are also concentrated: the median patient is no better off than under the status quo Waitlists mechanism.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the scope conditions for these findings?&lt;/strong&gt;
A: The demand model is estimated on the Trondelag region of Norway (approximately 8 percent of the national population) over 2017–2019, a period when waitlists were growing rapidly rather than in steady state. Counterfactual comparisons are made in a stationary equilibrium calibrated to Trondelag. The model excludes patients under 16 (whose enrollment is managed by parents). The partially capitated payment structure and fixed panel caps are institutional features specific to Norway, though similar systems exist in Canada, the UK, Italy, and Sweden. GP characteristics are held fixed in the model. The analysis abstracts from health outcomes, focusing on preference-based welfare from GP assignment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Top-Trading Cycles (TTC) algorithm&lt;/strong&gt;: A centralized reassignment algorithm that takes agents&amp;rsquo; preference lists and objects&amp;rsquo; priority lists as inputs, has each agent &amp;ldquo;point to&amp;rdquo; their preferred object and each object &amp;ldquo;point to&amp;rdquo; their highest-priority current or waiting agent, identifies cycles of mutual pointing, and executes the trades in those cycles simultaneously. In the paper&amp;rsquo;s static application, TTC is both Pareto-improving (every participant receives an assignment at least as good as their endowment) and strategy-proof. In the dynamic setting studied here, neither property holds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic TTC mechanism&lt;/strong&gt;: A mechanism that runs the TTC algorithm repeatedly at the end of each period after naturally arising vacancies have been filled from waitlists. Because patients&amp;rsquo; GP choices affect how long they wait — not only which GP they receive — this mechanism is not strategy-proof and may leave patients with undersubscribed GPs worse off than under strictly FCFS waitlists.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;TTC with Priority (TTCP)&lt;/strong&gt;: A modified version of dynamic TTC that changes the priority ordering so that patients with undersubscribed current GPs are prioritized above patients with oversubscribed GPs when panel vacancies are allocated. This modification preserves patients&amp;rsquo; endowment rights but compensates the group harmed by standard TTC. In the paper&amp;rsquo;s simulations, TTCP achieves 61 percent of TTC&amp;rsquo;s mean welfare gains while leaving patients with undersubscribed GPs no worse off than under the status quo.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Patient attention model&lt;/strong&gt;: A model in which patients consider switching GPs only when they receive a Poisson-distributed attention shock. Attention rates vary by observable characteristics (age, gender, temporary vs. permanent residency, whether and how far the patient recently moved). The model interprets any switch request as evidence of both an attention shock and a preference for the requested GP over the current one. Patients who do not request switches may be either inattentive or attentive but satisfied.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Horizontal differentiation (GP preference heterogeneity)&lt;/strong&gt;: The extent to which different patients prefer different GPs for reasons unrelated to overall GP quality — primarily driven by geographic proximity, gender homophily (worth 6.3–7.3 travel-time-equivalent minutes for young female patients), and age similarity (approximately 1 minute). Horizontal differentiation is the fundamental source of gains from trade: if all patients preferred the same GP, there would be no mutual-benefit swaps to find.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Deferred Acceptance (DA) algorithm&lt;/strong&gt;: The patient-proposing DA algorithm, which strictly respects FCFS waiting-time priority: no patient may be reassigned ahead of another patient who has been waiting longer for the same GP. In the dynamic context, DA achieves essentially no welfare improvement over the status quo because its strict respect for priority eliminates nearly all trading opportunities, illustrating the trade-off between envy-freeness and efficiency.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Double coincidence of wants&lt;/strong&gt;: The situation in which two (or more) patients are simultaneously on each other&amp;rsquo;s waitlists and would mutually benefit from trading GP assignments, but cannot do so under the current mechanism because there is no vacancy on either panel. The paper&amp;rsquo;s direct evidence of this phenomenon — 15 percent of waiters could be immediately reassigned via one TTC run — motivates the counterfactual analysis.&lt;/p&gt;</description></item><item><title>Disincentive effects of unemployment insurance benefits</title><link>https://macropaperwarehouse.com/papers/disincentive-effects-of-unemployment-insurance-benefits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/disincentive-effects-of-unemployment-insurance-benefits/</guid><description>&lt;p&gt;This paper isolates the disincentive effects of pandemic unemployment insurance (UI) benefits on employment recovery, separating them from the simultaneously operating stimulative (demand) effects that previous studies conflate. The authors study the largest UI expansion in U.S. history — the CARES Act of March 2020 — which introduced three simultaneous provisions: a $600 weekly income supplement (FPUC) through end of July 2020, a 13-week extension of maximum benefit duration (PEUC), and expanded eligibility to workers previously ineligible for UI (PUA), together raising the median replacement rate to 145% and more than doubling the number of UI recipients.&lt;/p&gt;
&lt;p&gt;The empirical strategy uses high-frequency establishment-level data from Homebase (HB), a scheduling and payroll provider covering approximately 140,000 small U.S. businesses — predominantly restaurants and retailers — matched to Yelp price-tier data and Safegraph foot-traffic and spending data. The final estimation sample is 4,595 businesses within 1,195 local-industry cells, observed at weekly frequency from January 2019 to December 2020.&lt;/p&gt;
&lt;p&gt;The identification rests on comparing employment recovery of low-wage versus high-wage businesses within the same narrow local labor market (four-digit zip code), industry (two-digit NAICS), and price tier. Because neighboring businesses largely share the local demand stimulus from UI, differencing within local-industry cells removes common demand effects. The key variation is the expiration of the $600 supplement, which differentially compresses the replacement-rate gap between low- and high-wage businesses depending on local average wages — labor markets where the gap falls more sharply are the treated group.&lt;/p&gt;
&lt;p&gt;The main empirical finding is that a 100 percentage point decline in the replacement rate gap is associated with a 5.7 percentage point rise in low-wage business employment recovery relative to high-wage business employment recovery at 12 weeks after the $600 expiration. For the average labor market, the expiration of the $600 supplement decreased the replacement rate gap by 46 percentage points, implying a 2.6 percentage point closing of the low-versus-high-wage employment gap within 12 weeks. Importantly, hours per employee and hourly wages grew faster in low-wage businesses over the same period, consistent with a labor supply rather than a demand mechanism. When the comparison is conducted at the U.S. state level rather than within local-industry cells — as in Finamor and Scott (2021) — the effect disappears and reverses sign, illustrating how local demand effects obscure disincentive effects at broader geographic aggregations.&lt;/p&gt;
&lt;p&gt;To quantify the aggregate employment impact, the authors build and calibrate a McCall-style labor search model with heterogeneous firm wages, a UI-eligible and non-UI unemployed pool, and equilibrium reservation wages. The model is extended to include a probability (calibrated at 16.5%) that workers lose UI eligibility upon refusing a job offer, which reconciles the model with the empirical estimates; without this feature the baseline model substantially overstates the differential employment effect of the $600 expiration.&lt;/p&gt;
&lt;p&gt;The full model-implied aggregate employment loss from all CARES Act UI provisions combined is 3.4 percentage points on average between April and December 2020, representing approximately 20% of the average employment shortfall in the Leisure and Hospitality sector over that period. When each provision is implemented in isolation, the effects are modest ($600 supplement: 0.2 pp; extended duration: 0.2 pp; expanded eligibility: 1.0 pp), but their interaction generates the large combined effect. Expanded eligibility is identified as the most disruptive provision, particularly for low-wage businesses, because it depletes the pool of non-UI unemployed who are the primary source of hires for these firms. The unemployment duration elasticities implied by the model are modest and in line with the low-to-middle range of pre-pandemic estimates.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s scope is restricted to the disincentive channel and deliberately excludes the stimulative effects of UI; it studies small, in-person service sector businesses and the April–December 2020 recovery period only.&lt;/p&gt;
&lt;p&gt;Q: What is the core identification challenge this paper addresses?
A: Prior empirical studies find only modest net effects of pandemic UI on employment, but it is unclear whether this reflects small disincentive effects or the near-cancellation of two opposing forces — UI suppressing labor supply while simultaneously stimulating local consumer demand. Identifying the disincentive effect alone requires a design that neutralizes the demand channel. The authors accomplish this by comparing low-wage and high-wage businesses within the same narrow local market, industry, and price tier, so that common local demand shifts from UI are differenced out.&lt;/p&gt;
&lt;p&gt;Q: What data does the empirical analysis use, and how is the sample constructed?
A: The primary data source is Homebase, covering approximately 140,000 small U.S. businesses with daily employment, hourly wages, and hours worked. The estimation sample is restricted to 4,595 businesses present throughout 2019, matched to Yelp price-tier classification and Safegraph weekly foot traffic and credit-card spending. Businesses are grouped into 1,195 local-industry cells defined by four-digit zip code, two-digit NAICS industry, and Yelp price tier (inexpensive vs. expensive). Within each cell, businesses are classified as low-wage or high-wage, with high-wage businesses paying on average $1.80 per hour more — about 8% above the average hourly wage of $10.90.&lt;/p&gt;
&lt;p&gt;Q: How is the replacement rate defined in the empirical framework?
A: The business-specific replacement rate is the ratio of average UI receipts (state benefit plus the pandemic supplement, converted to hourly units) to the pre-pandemic average hourly wage of that business. Because the supplement is uniform across workers, businesses with lower pre-pandemic wages face higher replacement rates; the replacement rate gap between low- and high-wage businesses within a local market is therefore a function of both state benefit levels and the local wage dispersion.&lt;/p&gt;
&lt;p&gt;Q: What does the event-study analysis around the $600 expiration show?
A: The event study exploits cross-labor-market variation in how much the replacement rate gap between low- and high-wage businesses declined when the $600 FPUC supplement expired at end of July 2020. Labor markets with a larger decline in the gap see faster relative recovery in low-wage business employment after expiration. A 100 percentage point decline in the replacement rate gap is associated with a 5.7 percentage point rise in the low-versus-high-wage employment recovery gap at 12 weeks post-expiration. For the average labor market, the $600 expiration reduced the replacement rate gap by 46 percentage points, implying a 2.6 percentage point narrowing of the employment recovery gap.&lt;/p&gt;
&lt;p&gt;Q: Why does the estimated effect disappear when broader geographic aggregations are used?
A: When businesses are compared within U.S. state borders rather than within local-industry cells, the estimated coefficient on the replacement rate gap turns positive and statistically insignificant. This occurs because at the state level, low-wage areas benefit disproportionately from the purchasing power increase that generous UI provides to local unemployed workers, so demand effects swamp and reverse the supply-side disincentive. This finding explains why Finamor and Scott (2021), using Homebase data with state fixed effects, find no negative association between replacement rates and labor market re-entry.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports a labor supply rather than demand interpretation of the differential recovery?
A: During the period of the $600 supplement, hours per employee and hourly wages grew faster in low-wage businesses than in high-wage businesses, even as low-wage businesses lagged in employment levels. If the differential recovery reflected demand deficiencies at low-wage businesses, hours per employee and wages should have grown faster at high-wage businesses instead. The observed pattern is consistent with labor supply shortfalls at low-wage firms.&lt;/p&gt;
&lt;p&gt;Q: What is the structure of the quantitative labor search model?
A: The model features a unit measure of workers and a fixed measure of firms, each posting a constant idiosyncratic wage drawn from an exogenous distribution. Unemployed workers receive job offers at a rate determined by labor market tightness and accept offers above their reservation wage. Reservation wages are equilibrium objects because UI benefits depend on the worker&amp;rsquo;s previous wage. The unemployed are split into UI-eligible and non-UI pools; the non-UI pool accepts jobs from lower in the wage distribution and is the primary supply source for low-wage firms. The model is calibrated to pre-pandemic U.S. service sector averages, with a pre-pandemic UI replacement rate of 0.51, a UI recipiency probability of 14%, and a non-UI replacement rate of 0.15.&lt;/p&gt;
&lt;p&gt;Q: Why does the baseline model overstate the empirical effect, and how is this reconciled?
A: The baseline model dramatically overstates the differential employment impact of the $600 expiration because the CARES Act&amp;rsquo;s expanded eligibility (modeled as a rise in the recipiency probability from 14% to 70%) nearly empties the non-UI unemployed pool, which is the dominant labor supply source for low-wage firms. In the data, the share of unemployed receiving UI nearly tripled for in-person leisure and hospitality workers, but not to the degree that the model&amp;rsquo;s implied employment collapse would require. The model is reconciled by introducing a 16.5% probability that a worker loses UI eligibility upon refusing a suitable job offer — consistent with UI law — which reduces the effective outside option and raises acceptance rates for low-wage firms.&lt;/p&gt;
&lt;p&gt;Q: What are the aggregate employment losses implied by the model?
A: When all three CARES Act provisions are implemented jointly, the model estimates that the disincentive effects held back aggregate employment recovery by 3.4 percentage points on average between April and December 2020 — approximately 20% of the average employment shortfall in the Leisure and Hospitality sector. Implemented in isolation, each provision generates only modest losses: the $600 supplement alone accounts for 0.2 percentage points, extended duration for 0.2 percentage points, and expanded eligibility for 1.0 percentage points. The large combined effect arises from the interaction of all three provisions, not from any single one.&lt;/p&gt;
&lt;p&gt;Q: What are the conditional (interaction) effects of each provision when the other two are in place?
A: Conditional on the other two provisions being active, the income supplement holds back employment recovery by 1.6 percentage points, the extended duration by 1.5 percentage points, and expanded eligibility by 2.9 percentage points. This interaction effect is the central quantitative finding: individually modest provisions combine to produce effects far exceeding their sum when implemented simultaneously.&lt;/p&gt;
&lt;p&gt;Q: What are the implied unemployment duration elasticities, and how do they compare to the literature?
A: The $600 supplement alone raises average unemployment duration by 8% against a 343% rise in the replacement rate, implying an elasticity of 0.02. Extended duration alone raises unemployment duration by 6% against a 150% increase in potential benefit duration, implying an elasticity of 0.03. Expanded eligibility alone raises unemployment duration by 19%, implying an elasticity of 0.04. When each provision is activated on top of the other two, the implied elasticities rise substantially: 0.24 for the $600 supplement, 0.43 for extended duration, and 0.28 for expanded eligibility. These are in the low-to-middle range of pre-pandemic estimates (Katz and Meyer, 1990: 0.3–0.5; Johnston and Mas, 2018: 0.4–0.8; Rothstein, 2011: 0.06; Farber and Valletta, 2015: 0.15).&lt;/p&gt;
&lt;p&gt;Q: What is the role of expanded eligibility specifically?
A: Expanded eligibility is identified as the most disruptive CARES Act provision, accounting for 1.0 percentage points of employment loss alone and 2.9 percentage points conditional on the other provisions. Mechanically, expanded eligibility converts non-UI unemployed workers into UI-eligible workers, draining the pool of workers willing to accept low-wage job offers. Because low-wage firms depend disproportionately on the non-UI pool for hiring, this provision disproportionately depresses their employment. Using CPS data, the authors document that the share of unemployed workers receiving UI in the in-person leisure and hospitality sector nearly tripled in 2020 relative to the pre-pandemic period.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions and limitations of the analysis?
A: The empirical analysis is restricted to small, in-person service sector businesses (restaurants and retailers) in the Homebase sample, which may not be representative of the broader labor market. The quantitative model is explicitly focused on disincentive effects only and does not capture the stimulative or demand effects of UI. The model also abstracts from re-opening restrictions and other pandemic-specific confounders. The analysis covers April to December 2020; the 2021 pandemic UI extensions are not studied. The job-refusal probability (chi = 16.5%) is a reduced-form calibration target rather than a structurally identified parameter.&lt;/p&gt;
&lt;p&gt;Replacement rate gap: The difference in business-specific UI replacement rates between low-wage and high-wage businesses within the same local labor market; defined as UI benefits (state benefit plus supplement) divided by the business&amp;rsquo;s pre-pandemic average hourly wage. Larger gaps indicate greater relative disincentive for workers to accept jobs at low-wage firms.&lt;/p&gt;
&lt;p&gt;Disincentive effect: The negative impact of higher UI replacement rates on workers&amp;rsquo; willingness to accept job offers and thus on business employment recovery, isolated from the simultaneous stimulative demand effect of UI spending.&lt;/p&gt;
&lt;p&gt;Non-UI unemployed pool: Workers who are ineligible for or have exhausted UI benefits and therefore receive only social benefits at a lower replacement rate (calibrated at 0.15 in the model). This group has a lower reservation wage and constitutes the primary labor supply source for low-wage firms.&lt;/p&gt;
&lt;p&gt;Local-industry cell: The paper&amp;rsquo;s unit of comparison — businesses sharing the same four-digit zip code (covering on average four neighboring zip codes), two-digit NAICS industry, and Yelp price tier. Within-cell differencing is the mechanism that removes common local demand effects.&lt;/p&gt;
&lt;p&gt;Benefit recipiency probability: The probability that a newly separated worker enters the UI-eligible unemployed pool, combining UI eligibility and takeup. Pre-pandemic this is calibrated at 14%; under the CARES Act it rises to 70%, targeting the observed near-tripling of UI recipients in the CPS data.&lt;/p&gt;
&lt;p&gt;Job-refusal eligibility loss: A probability (calibrated at 16.5%) that a UI-eligible worker who rejects a job offer loses UI status and transitions to the non-UI pool. Motivated by UI law prohibiting refusal of suitable work; reduces the effective outside option and reconciles the model&amp;rsquo;s predicted employment gap with the empirical estimate.&lt;/p&gt;
&lt;p&gt;Equilibrium residual wage dispersion: The wage dispersion observed in equilibrium conditional on worker observables. The model generates realistic dispersion by calibrating the non-UI replacement rate to match the lower half of the wage distribution and the firm wage offer variance to match the upper half; the presence of the non-UI state substantially increases residual dispersion relative to standard search models.&lt;/p&gt;</description></item><item><title>Diversifying Society's Leaders? Determinants and Causal Effects of Admission</title><link>https://macropaperwarehouse.com/papers/diversifying-societys-leaders-determinants-and-causal-effects-of-admission/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/diversifying-societys-leaders-determinants-and-causal-effects-of-admission/</guid><description>&lt;p&gt;This paper studies why children from high-income families are more likely to attend Ivy-Plus colleges (Ivy League, Stanford, MIT, Duke, Chicago — 12 colleges total) and whether attending these colleges causally improves post-college outcomes. The authors construct a de-identified panel dataset linking federal income tax records, Department of Education college attendance data, College Board and ACT test scores, and application and admissions records from several Ivy-Plus and flagship public colleges covering approximately 2.4 million students across entering classes from 1998–2015.&lt;/p&gt;
&lt;p&gt;The central finding on the input side is that students from families in the top 1% of the income distribution (income above $611,000) are 2.3 times more likely to attend an Ivy-Plus college than middle-class students (defined as the 70th–80th percentiles of the national parental income distribution, approximately $91,000–$114,000) with comparable SAT/ACT scores. Two-thirds of this gap is attributable to higher admissions rates at Ivy-Plus colleges for high-income applicants; conditional on SAT/ACT scores, top-1% applicants are 58% more likely to be admitted than middle-class applicants. The remaining third splits between differences in application rates (roughly 20% of the total attendance gap) and matriculation rates (roughly 12%). In contrast, admissions rates at flagship public colleges are essentially uncorrelated with parental income conditional on test scores.&lt;/p&gt;
&lt;p&gt;Three admissions practices drive the high-income admissions advantage at Ivy-Plus colleges. First, legacy preferences: legacy applicants from the top 1% are admitted at more than five times the rate of non-legacy applicants with comparable test scores, demographics, and admissions ratings; children of alumni of a given Ivy-Plus college are not more likely to be admitted to other Ivy-Plus colleges, confirming that legacy status is not merely a proxy for unobservable credentials. Legacy preferences account for 52 of the estimated 168 &amp;ldquo;extra&amp;rdquo; top-1% students per average Ivy-Plus class (enrollment ~1,650). Second, non-academic ratings: students from the top 1% have markedly stronger non-academic credentials (extracurricular activities, leadership ratings) partly because they disproportionately attend private high schools whose students receive higher non-academic ratings despite no higher academic ratings; this accounts for 35 additional extra top-1% students. Third, athletic recruitment: the share of recruited athletes rises from 5% among admitted students from the bottom 60% to 13% among those from the top 1%, accounting for 27 additional extra top-1% students.&lt;/p&gt;
&lt;p&gt;On the output side, the authors estimate causal effects of attending an Ivy-Plus college using a new research design based on waitlisted applicants. The key identification assumption is that idiosyncratic variation in admissions decisions across waitlisted applicants at one Ivy-Plus college is uncorrelated with admissions decisions at other Ivy-Plus colleges — which the authors verify empirically. Under this assumption, comparisons of admitted vs. rejected waitlisted applicants identify causal effects for marginal students. The marginal student who attends an Ivy-Plus college instead of the average flagship public is approximately 50% more likely to reach the top 1% of the earnings distribution at age 33, nearly twice as likely to attend a highly-ranked graduate school, and 2.5 times as likely to work at a prestigious firm. Attending an Ivy-Plus college increases mean earnings by $101,000 at age 33 relative to a counterfactual mean of $143,000 at state flagships. Effects are concentrated in the upper tail of earnings — the impact on reaching the top quartile is small and statistically insignificant, while impacts on reaching the top 1% far exceed what a constant percentage treatment effect would predict. Effects are larger for students with weaker fallback options (i.e., whose home-state colleges channel fewer students to the top 1%).&lt;/p&gt;
&lt;p&gt;Critically, the three credentials driving the high-income admissions advantage — legacy status, athletic recruitment, and high non-academic ratings — are uncorrelated with or negatively correlated with post-college success once the college attended is held constant. Academic credentials (SAT/ACT scores, academic ratings) remain highly predictive of outcomes.&lt;/p&gt;
&lt;p&gt;Counterfactual simulations show that eliminating all three high-income admissions preferences and replacing those slots with students having the same test score distribution would increase enrollment from the bottom 95% of the parental income distribution by 8.8 percentage points — comparable in magnitude to the effect of race-based affirmative action on Black and Hispanic enrollment shares. Such a policy would have small effects on monetary leadership outcomes (e.g., Fortune 500 CEO share from bottom-95% families rises by only 0.4 pp, because Ivy-Plus graduates are a small fraction of all top earners) but larger effects on non-monetary leadership positions: the share of senators from the bottom 95% would rise by 1.7 pp and the share of Supreme Court justices by 5.4 pp. With need-affirmative policies (giving low-income students preferences comparable to those currently given to legacy applicants), the share of Supreme Court justices from families in the bottom 60% would rise by 17.5 pp. These predictions assume that the causal share of Ivy-Plus attendance in explaining observational differences in leadership outcomes is the same as that estimated for early-career outcomes, and they ignore general equilibrium effects.&lt;/p&gt;
&lt;p&gt;Q: How much more likely are top-1% students to attend an Ivy-Plus college than middle-class students with the same test scores?
A: Students from families in the top 1% (income above $611,000) are 2.3 times more likely to attend an Ivy-Plus college than students from the 70th–80th percentile of the parental income distribution (approximately $91,000–$114,000) with comparable SAT/ACT scores. This &amp;ldquo;missing middle&amp;rdquo; pattern is stable across entering classes from 1998 to 2018 and persists after controlling for race and ethnicity.&lt;/p&gt;
&lt;p&gt;Q: How is the overall attendance gap decomposed into application, admissions, and matriculation?
A: Differences in admissions rates explain two-thirds of the gap in Ivy-Plus attendance between top-1% and middle-class students conditional on test scores. Of the estimated 168 &amp;ldquo;extra&amp;rdquo; top-1% students per average Ivy-Plus class, 87 come from higher admissions rates for non-recruited athletes, 27 from athletic recruitment, and the remaining slack from application rate differences (accounting for roughly 20% of the overall attendance gap) and matriculation differences (roughly 12%).&lt;/p&gt;
&lt;p&gt;Q: How large is the admissions advantage for top-1% applicants at Ivy-Plus colleges?
A: Conditional on SAT/ACT scores, applicants from the top 1% are 58% more likely to be admitted to Ivy-Plus colleges than middle-class applicants. Students from the top 0.1% are 2.5 times more likely to be admitted than middle-class applicants with comparable test scores. At flagship public colleges, admissions rates are essentially constant across the income distribution conditional on test scores.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of legacy preferences and how is it established that legacy is not just a proxy for other credentials?
A: Legacy applicants from the top 1% are admitted at more than five times the rate of otherwise comparable non-legacy applicants at the college their parents attended. The paper isolates the legacy effect by showing that children of alumni at a given Ivy-Plus college are only slightly more likely to be admitted at other Ivy-Plus colleges — and the predicted counterfactual admissions rate for legacy students at other colleges closely matches their actual admissions rate — confirming that legacy status is not merely a proxy for other unobservable credentials. Legacy applicants constitute 2.5% of the overall applicant pool but over 9% of top-1% applicants.&lt;/p&gt;
&lt;p&gt;Q: How do non-academic credentials differ by parental income, and what drives the difference?
A: Top-1% applicants have markedly stronger non-academic ratings (measuring extracurricular participation and leadership traits) compared with other applicants, while the share achieving high academic ratings is essentially constant across the income distribution. Students from the top 1% are much more likely to have attended private high schools, whose applicants receive substantially higher non-academic ratings than students from public high schools with the same SAT/ACT scores. Non-academic ratings account for 35 of the estimated 168 extra top-1% students per Ivy-Plus class.&lt;/p&gt;
&lt;p&gt;Q: What is the research design for estimating causal effects, and what is the key identification assumption?
A: The authors focus on applicants who are waitlisted at a given Ivy-Plus college and compare those ultimately admitted versus rejected from the waitlist. The key identification assumption is that if different colleges&amp;rsquo; admissions committees make correlated assessments of underlying student merit but uncorrelated idiosyncratic admissions errors, then residual variation in admissions outcomes for waitlisted applicants at one college is orthogonal to students&amp;rsquo; long-run potential. The authors validate this empirically by showing that waitlist admission at one Ivy-Plus college is uncorrelated with admissions decisions and internal ratings at other Ivy-Plus colleges.&lt;/p&gt;
&lt;p&gt;Q: What are the causal effects of attending an Ivy-Plus college on post-college outcomes?
A: For the marginal student (one who attends an Ivy-Plus college instead of the average flagship public), attending an Ivy-Plus college increases the probability of reaching the top 1% of the earnings distribution at age 33 by approximately 50%, nearly doubles the probability of attending an elite graduate school, and increases the probability of working at a prestigious firm by approximately 2.5 times. Mean earnings at age 33 increase by $101,000 (relative to a counterfactual mean of $143,000 at state flagships). Effects on reaching the top quartile of earnings are small and statistically insignificant, while effects at the very top tail are disproportionately large.&lt;/p&gt;
&lt;p&gt;Q: Why do the findings differ from Dale and Krueger (2002) and related studies finding little effect of selective college attendance on earnings?
A: The authors replicate the matriculation design of Dale and Krueger (comparing outcomes conditional on the set of colleges to which students were admitted) and obtain estimates statistically indistinguishable from their waitlist design — the research designs are not the source of disagreement. Instead, the differences arise because (1) the authors have direct college fixed effects rather than relying on average test scores as a proxy for college quality, and (2) the authors focus on upper-tail outcomes (top 1% earnings, elite graduate schools, prestigious firms) rather than log mean earnings, where Ivy-Plus colleges have their largest effects.&lt;/p&gt;
&lt;p&gt;Q: Are the credentials that drive the high-income admissions advantage — legacy, athlete status, high non-academic ratings — predictive of better post-college outcomes?
A: No. Recruited athletes, students with higher non-academic ratings, and legacy students have equivalent or lower chances of reaching the upper tail of the income distribution, attending an elite graduate school, or working at a prestigious firm than comparable Ivy-Plus applicants once the college attended is held constant. By contrast, SAT/ACT scores and academic ratings are highly positively predictive of all three post-college outcome measures.&lt;/p&gt;
&lt;p&gt;Q: How much could changing admissions practices diversify Ivy-Plus enrollment and subsequently society&amp;rsquo;s leadership?
A: Eliminating legacy preferences, non-academic rating weights, and the differential recruitment of high-income athletes — and filling those slots with students having the same test score distribution as the current class — would increase enrollment from families in the bottom 95% of the parental income distribution by 8.8 percentage points, a magnitude comparable to race-based affirmative action&amp;rsquo;s effect on Black and Hispanic enrollment shares. For leadership positions, predicted effects are small for monetary outcomes (Fortune 500 CEOs from the bottom 95% would increase by only 0.4 pp) but larger for positions where Ivy-Plus graduates are a larger share: senators from the bottom 95% would increase by 1.7 pp and Supreme Court justices by 5.4 pp. A stronger need-affirmative policy (giving low-income students preferences equivalent to current legacy preferences) would increase the share of Supreme Court justices from the bottom 60% by 17.5 pp.&lt;/p&gt;
&lt;p&gt;Q: How are &amp;ldquo;elite&amp;rdquo; and &amp;ldquo;prestigious&amp;rdquo; employers defined in this study?
A: Elite firms are defined as those that disproportionately employ Ivy-Plus graduates relative to flagship public graduates, pulling firms from the top of that ratio ranking until 25% of Ivy-Plus attendee employment is accounted for. Prestigious employers are defined by the residual of that ratio after controlling for the firm&amp;rsquo;s predicted top-1% income probability — they are firms that disproportionately employ Ivy-Plus graduates conditional on their salaries, capturing high-status jobs that do not necessarily lead to the highest earnings. The paper validates this algorithmic approach against external rankings (Vault.com for law and consulting firms; Scimagoir for hospitals), finding substantial overlap.&lt;/p&gt;
&lt;p&gt;Q: How are treatment effect estimates adjusted for heterogeneity in students&amp;rsquo; fallback options?
A: Causal effects of Ivy-Plus attendance are much larger for students with weaker fallback options — specifically, students whose home-state flagship colleges channel fewer students to the top 1% of earnings. The authors exploit this heterogeneity to estimate the treatment effect for the marginal student who actually switches from a flagship public to an Ivy-Plus college. This heterogeneity also implies that the average causal effect across all admitted students may differ from the effect for the marginal admitted student.&lt;/p&gt;
&lt;p&gt;Q: What share of the overrepresentation of top-1% families at Ivy-Plus colleges is attributable to pre-application factors versus admissions practices?
A: Of the 245 &amp;ldquo;extra&amp;rdquo; top-1% students in an average Ivy-Plus class relative to an unconditionally income-neutral benchmark, 77 (31%) are attributable to the higher test scores of top-1% students (a pre-application factor). The remaining 168 (69%) reflect higher attendance rates conditional on test scores, of which the large majority is attributable to admissions practices (legacy, non-academic ratings, athletic recruitment) rather than application or matriculation rate differences.&lt;/p&gt;
&lt;p&gt;Ivy-Plus colleges: The twelve highly selective private colleges comprising the eight Ivy League institutions plus Stanford, MIT, Duke, and the University of Chicago — the focus group of the study, which together account for more than 10% of Fortune 500 CEOs, a quarter of U.S. senators, and three-fourths of Supreme Court justices appointed in the last half century despite enrolling less than 0.5% of Americans.&lt;/p&gt;
&lt;p&gt;Missing middle: The pattern by which attendance rates at Ivy-Plus colleges conditional on SAT/ACT scores are lowest for students from the middle class (70th–80th percentile of the parental income distribution, approximately $91,000–$114,000) — lower than both the top 1% and, slightly, the bottom 40% — producing a non-monotone income gradient in attendance.&lt;/p&gt;
&lt;p&gt;Legacy preference: An admissions advantage given to applicants whose parent(s) obtained an undergraduate degree from the college to which the student is applying. In the paper&amp;rsquo;s data, legacy applicants from the top 1% are admitted at more than five times the rate of non-legacy applicants with comparable test scores, demographics, and admissions ratings; the preference is college-specific (children of alumni are only slightly more likely to be admitted at other Ivy-Plus colleges).&lt;/p&gt;
&lt;p&gt;Waitlist research design: The paper&amp;rsquo;s primary identification strategy for causal effects, which exploits idiosyncratic variation in admissions decisions among waitlisted applicants. The design&amp;rsquo;s validity rests on the empirical finding that waitlist admissions at one Ivy-Plus college are uncorrelated with admissions decisions and internal ratings at other Ivy-Plus colleges, implying that residual variation conditional on being on the waitlist is orthogonal to students&amp;rsquo; long-run potential outcomes.&lt;/p&gt;
&lt;p&gt;Prestigious employers: Firms defined by the paper&amp;rsquo;s algorithm as disproportionately employing Ivy-Plus graduates conditional on those firms&amp;rsquo; predicted top-1% income probability — capturing high-status employment that does not necessarily lead to the highest earnings (e.g., prominent law firms, consulting firms, elite hospitals). Validated against external rankings (Vault.com, Scimagoir).&lt;/p&gt;
&lt;p&gt;Non-academic ratings: Numerical scores assigned by admissions officers measuring aspects of an application outside academic achievement, such as extracurricular activities and leadership traits. In the paper&amp;rsquo;s data, non-academic ratings differ substantially by parental income — particularly because top-1% applicants disproportionately attend private high schools whose students receive higher non-academic ratings — while academic ratings do not differ across the income distribution.&lt;/p&gt;
&lt;p&gt;Surrogate index: A prediction of later earnings outcomes (specifically, probability of reaching the top 1% at age 33 and mean income rank) constructed from individuals&amp;rsquo; graduate school attendance and employer fixed effects at ages 22–25, used to extend the outcome window for cohorts observed only early in their careers. The approach follows the terminology and methodology of Athey et al. (2019).&lt;/p&gt;</description></item><item><title>Do The Effects of Nudges Persist? Theory and Evidence from 38 Natural Field Experiments</title><link>https://macropaperwarehouse.com/papers/do-the-effects-of-nudges-persist-theory-and-evidence-from-38-natural-field-experiments/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/do-the-effects-of-nudges-persist-theory-and-evidence-from-38-natural-field-experiments/</guid><description>&lt;p&gt;This paper asks why the Home Energy Report (HER) — a widely deployed social-comparison nudge that shows households how their electricity consumption compares to their neighbors — produces behavioral changes that persist long after the nudge is discontinued, while analogous nudges in other domains (charitable giving, financial savings, voter turnout, tax compliance) fade almost entirely within a year or two. The authors formalize a research design to decompose the HER&amp;rsquo;s long-run effectiveness into two channels: technology adoption (a change in the stock of energy-efficient capital in the home) and habit formation (a change in the stock of habits or skills in the resident).&lt;/p&gt;
&lt;p&gt;The identifying strategy exploits the administrative rule that when the initial resident in an HER experiment moves out, HER mailings stop immediately — but electricity consumption in the home continues to be observed as new residents occupy it. Under three assumptions — (1) treatment assignment did not influence the initial resident&amp;rsquo;s decision to move; (2) treatment assignment did not influence the type of resident who moved in; and (3) energy-efficient technology adopted in response to the HER remained in the home after the move — the post-move HER effect identifies the fraction of the long-run treatment effect attributable to technology adoption (ATK), and the remainder identifies the fraction attributable to habit formation (ATH).&lt;/p&gt;
&lt;p&gt;Data come from 38 natural field experiments administered by Opower between 2008 and 2013 across 21 U.S. residential energy providers, comprising 61,310,166 electricity bills for 1,810,096 homes. The mover sample, restricted to homes where the initial resident deactivated service at or after the receipt of their fourth HER, contains 5,890,855 bills for 139,908 homes. Treatment and control homes enter the mover sample at statistically indistinguishable rates and have similar baseline electricity consumption.&lt;/p&gt;
&lt;p&gt;The main findings: the HER reduced electricity consumption by 2.1 percent in the long run (the pre-move ATE). After the initial resident moved and the HER was discontinued, 1.1 percent of the reduction persisted in the home — attributable to technology. The habit channel accounts for the remaining 1.0 percent reduction. Normalizing by the ATE, 51.4 percent (s.e. = 13.1) of the long-run effectiveness is attributable to technology adoption and 48.6 percent to habit formation. The persistence of the post-move effect is robust across alternative specifications, different HER-receipt cutoffs, balanced panels, and exclusion of low-consumption move-period homes. A falsification test using rental homes — where tenants do not typically own appliances and the technology channel is therefore shut down — yields a null post-move effect, consistent with the balanced-habits assumption.&lt;/p&gt;
&lt;p&gt;The authors use these results to explain a broader empirical pattern: one year after discontinuation, social comparison nudges targeting compliance, charitable giving, savings, and voter turnout retain on average only 4 percent of their initial effect, while nudges targeting energy and water conservation retain 65 percent. The paper argues this divergence reflects the relative abundance of enabling technologies in conservation contexts versus their absence in compliance or voting contexts. The findings also have cost-benefit implications: ignoring HER-induced technology adoption overstates net benefits by as much as 65 percent, depending on assumed technology cost per kWh saved (ranging from $0.03 per kWh saved per Gillingham et al. 2018 to $0.12 per kWh saved per Billingsley et al. 2014).&lt;/p&gt;
&lt;p&gt;Scope conditions: results are specific to electricity-consumption nudges in the U.S. residential sector; the technology channel identification requires that adopted equipment stays in the home after a move; the decomposition rests on a linear production function for outcomes in habits and technology.&lt;/p&gt;
&lt;p&gt;Q: What is the Home Energy Report and how was it administered in these experiments?
A: The HER is a mailed social-comparison report that contrasts a household&amp;rsquo;s electricity consumption with that of similar neighbors. In each of the 38 waves, homes were observed for a 12-month baseline, then randomly assigned to treatment (receiving HERs) or control. HERs were mailed monthly, bimonthly, or quarterly; generation ceased when the initial resident deactivated electricity service.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s central identification strategy?
A: The authors exploit a discontinuity created when the initial treated resident moves out: HER mailings stop, but the home&amp;rsquo;s electricity consumption continues to be measured as new residents move in. Under three assumptions about non-interference of treatment with moving decisions, balanced habits of subsequent residents, and stability of adopted technology, the post-move HER effect point-identifies the technology-adoption component (ATK) of the long-run average treatment effect (ATE). The habit-formation component (ATH) is then inferred as ATE minus ATK.&lt;/p&gt;
&lt;p&gt;Q: What are the three identifying assumptions and how are they tested?
A: Assumption 1 (no effect of treatment on moving rates) and Assumption 2 (balanced habits of subsequent residents) are tested with the data; treatment and control homes enter the mover sample at statistically indistinguishable rates and have similar baseline consumption, supporting Assumption 1. The rental-home falsification test supports Assumption 2: rental homes show a null post-move effect, consistent with renters having balanced habits because the technology channel is inactive in rentals. Assumption 3 (stable technology after a move) is untestable from the data; the authors note that violation of this assumption would imply the post-move effect is a lower bound on ATK, making the technology-adoption estimate conservative.&lt;/p&gt;
&lt;p&gt;Q: What are the main quantitative estimates of the decomposition?
A: The pre-move (long-run) ATE is -2.1 percent of baseline electricity consumption. The post-move effect (ATK) is -1.1 percent, and the habit-formation component (ATH) is -1.0 percent. Normalizing by the ATE, 51.4 percent (s.e. = 13.1) is attributed to technology adoption and 48.6 percent to habits.&lt;/p&gt;
&lt;p&gt;Q: How large is the HER effect in absolute terms during the comparison period?
A: During the comparison period, the HER reduced average daily electricity consumption by approximately -1.8 to -2.3 percent in the first year and -1.5 to -2.0 percent in the second year, with 95 percent confidence intervals excluding zero. In levels, these correspond to roughly -0.6 to -0.9 kWh per day — equivalent to using 2 to 4 sixty-watt incandescent bulbs for 5 fewer hours per day.&lt;/p&gt;
&lt;p&gt;Q: How persistent is the HER effect during the move period?
A: In the first year of the move period the HER continues to produce reductions of -1.7 and -1.4 percent; more than a year after the initial resident&amp;rsquo;s departure the estimated effect is -1.2 percent. All move-period estimates are statistically significant at conventional levels.&lt;/p&gt;
&lt;p&gt;Q: How does the paper explain variation in persistence across social-comparison nudge contexts?
A: One year after discontinuation, nudges targeting compliance, charitable giving, savings, and voter turnout retain on average only 4 percent of their initial effect, while nudges targeting energy or water conservation retain 65 percent on average. The paper argues the divergence reflects the relative availability of enabling technologies: households can adopt long-lived, input-efficient technologies (appliances, fixtures) to reduce energy and water use, but analogous technologies to facilitate compliance, donations, or voting are largely unavailable or absent.&lt;/p&gt;
&lt;p&gt;Q: How does this paper&amp;rsquo;s finding about technology adoption compare to Allcott and Rogers (2014)?
A: Allcott and Rogers (2014) used participation in utility-sponsored energy-efficiency programs as a proxy for technology adoption and found it explained no more than 2 percent of the HER&amp;rsquo;s long-run effectiveness. The authors reject this conclusion: their decomposition attributes 51.4 percent to technology, which is estimated precisely enough to statistically reject the 2 percent figure from Allcott and Rogers (2014). They attribute the discrepancy to the imperfect proxy used by Allcott and Rogers and low statistical power in analogous analyses.&lt;/p&gt;
&lt;p&gt;Q: What are the cost-benefit implications of accounting for HER-induced technology adoption?
A: Assuming monthly HERs for one year, a household electricity price of $0.10/kWh, and benefits accruing over two years, the baseline net benefit (ignoring technology costs) is $32.38 per household (electricity savings of $44.38 minus $12 administration cost). Using a technology cost of $0.03/kWh saved (Gillingham et al. 2018), net benefits fall to $27.14. Using $0.12/kWh saved (Billingsley et al. 2014), net benefits drop to $11.43 — a reduction of up to 65 percent from the baseline estimate. The HER still passes cost-benefit analysis but prior evaluations that ignore technology costs overstate net benefits substantially.&lt;/p&gt;
&lt;p&gt;Q: How robust are the decomposition results to alternative sample definitions and specifications?
A: The qualitative findings are stable across: alternative sets of control variables (Table A1); mover samples defined by receiving as few as 1 or as many as 5 HERs before moving (Table A2, with pre-move effects of -2.08 and post-move effects of -0.93 to -1.04 across cutoffs); balanced panels requiring fixed observation windows in each period (Table A3); and exclusion of homes showing unusually low consumption in the move period (Table A4, post-move effects of -1.19 to -1.48).&lt;/p&gt;
&lt;p&gt;Q: What policy implications does the paper draw for nudge design?
A: Policymakers seeking persistent nudge effects should target behaviors that can be augmented by readily available technologies, or pair social-comparison nudges with opportunities to adopt new technologies. In voting contexts, combining social-comparison nudges with opt-in mail-in or online ballot defaults could produce more persistent effects. In savings and charitable giving, pairing social comparisons with automatic contribution-rate defaults (as in Madrian and Shea 2001; Thaler and Benartzi 2004) is predicted to produce longer-lived effects than the nudge alone.&lt;/p&gt;
&lt;p&gt;Q: What methodological contribution does the paper offer beyond the HER application?
A: The mover-based decomposition is a generalizable research design for separating human capital (habits, skills) from physical capital (technology, infrastructure) as channels of policy effectiveness. The authors suggest it can be applied using other natural separation events — such as student graduation or employee departure — to assess the extent to which nudges build human capital in both recipients and the organizations in which they are embedded.&lt;/p&gt;
&lt;p&gt;Technology adoption channel (ATK): The component of the HER&amp;rsquo;s long-run average treatment effect attributable to increases in the stock of energy-efficient technologies in the home — identified empirically as the post-move HER effect that persists after the treated resident departs and the HER is discontinued.&lt;/p&gt;
&lt;p&gt;Habit formation channel (ATH): The component of the HER&amp;rsquo;s long-run treatment effect attributable to changes in the habits or skills of the resident — inferred as the residual after netting the technology component (ATK) from the total long-run effect (ATE).&lt;/p&gt;
&lt;p&gt;Post-move effect: The estimated difference in electricity consumption between treatment and control homes after the initial resident has moved out, the HER has been discontinued, and a new resident has taken occupancy; under the paper&amp;rsquo;s identifying assumptions this equals ATK.&lt;/p&gt;
&lt;p&gt;Balanced-habits assumption: The identifying assumption that treatment assignment did not influence the characteristics or habits of residents who subsequently moved into homes in the experimental sample, so that the habits of incoming residents are comparable across treated and control homes.&lt;/p&gt;
&lt;p&gt;Stable-technology assumption: The identifying assumption that energy-efficient technologies adopted in response to the HER remain in the home after the initial resident moves; relaxing this assumption implies the post-move effect is a lower bound on ATK.&lt;/p&gt;
&lt;p&gt;Home Energy Report (HER): A mailed social-comparison report that contrasts a recipient household&amp;rsquo;s electricity consumption with that of similar neighboring households; the treatment studied across all 38 experiments in this paper.&lt;/p&gt;
&lt;p&gt;Enabling technologies: Long-lived, input-efficient capital goods (appliances, lighting, insulation) that reduce the marginal cost of conservation and thereby lock in behavioral changes induced by a nudge; their relative abundance in energy and water conservation contexts — versus their absence in voting, giving, or compliance contexts — is the paper&amp;rsquo;s proposed explanation for cross-context variation in nudge persistence.&lt;/p&gt;</description></item><item><title>Does Deposit Insurance Promote Deposit Stability? Evidence from the Postal Savings System during the 1920s</title><link>https://macropaperwarehouse.com/papers/does-deposit-insurance-promote-deposit-stability-evidence-from-the-postal-savings-system-during-the-1920s/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/does-deposit-insurance-promote-deposit-stability-evidence-from-the-postal-savings-system-during-the-1920s/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; Does deposit insurance promote financial depth by arresting the outflow of deposits from the banking system during periods of bank distress? The paper tests and quantifies the deposit-stabilizing effect of state-level deposit insurance schemes operating in the United States during the 1920s.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and identification.&lt;/strong&gt; Between 1908 and 1929, eight primarily Midwestern states adopted some form of deposit insurance. The paper exploits the discontinuity in deposit insurance coverage at state borders to identify the causal effect of insurance on depositor behavior. The identification strategy compares outcomes in contiguous city pairs straddling deposit-insurance (DI) and non-deposit-insurance (NDI) state borders — a quasi-experimental design that controls for observed and unobserved confounders by using narrow geographic areas where the only relevant policy difference is the presence or absence of deposit insurance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Proxy for &amp;ldquo;mattress money.&amp;rdquo;&lt;/strong&gt; The paper uses postal savings deposits as a proxy for money withdrawn from the banking system. The U.S. Postal Savings System (established 1911) was backed by the full faith and credit of the federal government, with a maximum individual account limit of $2,500, and was widely viewed as a far safer alternative to commercial bank deposits. The authors validate this proxy by demonstrating, via Johansen cointegration tests, that the nationwide ratio of postal savings balances to total bank deposits is cointegrated (rank 1) with the currency-deposit ratio — a well-established indicator of banking distress.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The empirical analysis covers 1921–1929. The main postal savings dataset is drawn from Annual Reports of the Postmaster General. Bank suspension data are drawn from FDIC manuscript lists compiled in the 1930s by FDIC economist Clark Warburton, providing location, charter type, and suspension/reopening dates. The sample includes 74 city pairs across 14 states (7 DI: North Dakota, South Dakota, Nebraska, Kansas, Oklahoma, Texas, Mississippi; 7 NDI: Minnesota, Iowa, Missouri, Arkansas, Louisiana, Tennessee, Alabama), with an average distance between paired cities of approximately 18 miles.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings — postal savings regressions (Table 4).&lt;/strong&gt; Using OLS with city-pair and year fixed effects and standard errors clustered at the NDI city level, the paper finds that following a bank suspension within a 10-mile radius, postal savings deposits in NDI cities grew 16 percent more than deposits in the corresponding DI city. The effect is positive and statistically significant at the 20-mile radius but smaller — approximately 9 percent — and is statistically indistinguishable from zero at the 30-mile radius. The localized decay with distance is consistent with a geographically contained flight-to-safety response. Critically, when the same specification is estimated for periods after deposit insurance was discontinued, the effect at all radii is statistically nil, providing a falsification test ruling out omitted unobserved factors as the driver.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Persistence of effects (Table 5).&lt;/strong&gt; Arellano-Bond GMM dynamic panel regressions confirm that the disintermediation effects are persistent. The lagged dependent variable enters with a negative and statistically significant coefficient (approximately −0.20 for the 10-mile regression), indicating mean reversion, but the bank suspension coefficients remain robust. Implied long-run effects for the 10-mile and 20-mile equations are approximately 0.151 and 0.100, respectively, suggesting sustained rather than transitory deposit diversion away from the banking system in the absence of deposit insurance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Banking capacity (Table 6).&lt;/strong&gt; Because the postal savings deposit limit constrained the intake of funds — particularly severely during distress episodes, as documented through narrative evidence from the 1915 Congressional Record — the postal savings regressions underestimate the true effect of deposit insurance. The paper therefore estimates an alternative specification at the county level, comparing deposits at state-chartered banks in paired DI and NDI border counties. The results indicate that deposit insurance is associated with approximately a 56 percent increase in county-level deposits at state-chartered banks (coefficient 0.574, significant at 5 percent, robust to inclusion or exclusion of year fixed effects). By contrast, the analogous coefficient for national banks — which were prohibited by the OCC from participating in state deposit insurance schemes — is positive but statistically insignificant, providing a placebo test consistent with the interpretation that deposit insurance, not unobserved county characteristics, drove the banking capacity difference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; All effects are estimated for state-chartered bank deposits in predominantly agricultural, Midwestern border counties during 1921–1929, a period characterized by an average annual bank suspension rate of 2.22 percent (versus 0.3 percent during 1911–1920). The paper acknowledges that state deposit insurance schemes of this era generated moral hazard (as established by prior literature), and frames the contribution as quantifying the stability-enhancing component rather than the net welfare effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy implication.&lt;/strong&gt; The 56 percent banking capacity differential implies that deposit runoffs in the absence of insurance are substantially higher than the 3–10 percent runoff rates assumed in the Basel III Liquidity Coverage Ratio (LCR) framework, and more consistent with the 25–50 percent runoffs observed in non-systemic institutions in Denmark following an exogenous reduction in deposit insurance limits (Iyer et al., 2016).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-is-the-postal-savings-system-a-valid-proxy-for-mattress-money-and-what-evidence-supports-this"&gt;Q1. Why is the Postal Savings System a valid proxy for &amp;ldquo;mattress money,&amp;rdquo; and what evidence supports this?&lt;/h3&gt;
&lt;p&gt;The postal savings system was backed by the full faith and credit of the United States, making it categorically safer than commercial bank deposits, and was explicitly designed to attract savings hidden in mattresses. The authors validate the proxy empirically by showing that the nationwide ratio of postal savings balances to total bank deposits is cointegrated (Johansen test, rank 1) with the currency-deposit ratio — a series that rises during banking distress as depositors convert bank funds to currency. Contemporary narrative accounts from the 1915 Congressional Record further confirm that postal savings offices experienced sharp deposit inflows during local banking distress, with deposit intake frequently constrained by the $2,500 individual account cap.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-identification-strategy-and-why-does-it-address-endogeneity-concerns"&gt;Q2. What is the identification strategy, and why does it address endogeneity concerns?&lt;/h3&gt;
&lt;p&gt;The strategy exploits the discontinuity in deposit insurance at state borders by comparing relative postal savings deposit growth in contiguous city pairs — one city in a DI state, one in an adjacent NDI state — conditioning on bank suspensions within 10, 20, or 30 miles. The authors argue that deposit insurance legislation was a statewide political decision driven largely by partisan composition (Democrats favored it, Republicans opposed it), making it implausible that interests concentrated at border cities systematically determined which states adopted it. Six of the seven NDI control states introduced deposit insurance legislation but failed to pass it, underscoring that the policy variation was not determined by border-specific characteristics. A falsification test using the same city pairs after deposit insurance was discontinued shows zero effects, ruling out time-invariant unobserved heterogeneity as the driver.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-quantitative-results-from-the-city-pair-postal-savings-regressions"&gt;Q3. What are the main quantitative results from the city-pair postal savings regressions?&lt;/h3&gt;
&lt;p&gt;Following a bank suspension within 10 miles, postal savings deposits in NDI cities grew 16 percent more than in DI cities (coefficient 0.162, significant at 5 percent). At the 20-mile radius the differential is approximately 9 percent (coefficient 0.0933, significant at 5 percent). At the 30-mile radius the coefficient is 0.0997 and statistically indistinguishable from zero. These results are estimated with OLS using city-pair and year fixed effects and standard errors clustered at the NDI city level, based on 524 observations for the 10- and 20-mile specifications and 66 observations for the post-discontinuation falsification regressions.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-paper-establish-that-distance-matters-for-the-flight-to-safety-effect"&gt;Q4. How does the paper establish that distance matters for the flight-to-safety effect?&lt;/h3&gt;
&lt;p&gt;The monotonic decline in the estimated coefficient from 0.162 (10 miles) to 0.093 (20 miles) to a statistically insignificant 0.100 (30 miles) indicates that the diversion of deposits into postal savings was geographically localized. This pattern is consistent with depositors responding primarily to nearby bank failures rather than to distant ones, and it supports the interpretation that the effect is driven by local banking distress rather than by state-level or regional macroeconomic shocks that would affect all pairs symmetrically.&lt;/p&gt;
&lt;h3 id="q5-are-the-disintermediation-effects-of-bank-suspensions-temporary-or-persistent"&gt;Q5. Are the disintermediation effects of bank suspensions temporary or persistent?&lt;/h3&gt;
&lt;p&gt;The Arellano-Bond GMM dynamic panel regressions (Table 5) show that the effects are persistent. The lagged dependent variable coefficient is approximately −0.205 (10-mile) and −0.188 to −0.201 (20-mile), indicating partial mean reversion but not full reversal. Year-1, Year-2, and implied long-run dynamic effects are all statistically significant and of similar magnitude (approximately 0.145–0.152 for the 10-mile equation and 0.096–0.100 for the 20-mile equation), indicating that once depositors shift funds to postal savings in response to bank suspensions, a substantial portion of the effect persists in subsequent years. This is consistent with prior literature showing that deposits leave the banking system quickly but return slowly.&lt;/p&gt;
&lt;h3 id="q6-why-are-the-postal-savings-coefficient-estimates-considered-a-lower-bound-on-the-true-effect-of-deposit-insurance"&gt;Q6. Why are the postal savings coefficient estimates considered a lower bound on the true effect of deposit insurance?&lt;/h3&gt;
&lt;p&gt;Two institutional features constrained the postal savings system from fully capturing flight-to-safety deposits. First, individual accounts were capped at $2,500, and narrative evidence shows that this limit was severely binding during distress — depositors attempted to place far more than the ceiling allowed. Second, the re-depositing rate of postal savings funds back into local banks was not 100 percent: during 1921–1923 only 32–47 percent of postal savings deposits were re-deposited in banks, compared to 72–82 percent in calmer years. Because the postal savings system could not absorb unlimited deposits and did not fully recycle absorbed funds into local banking, its level understates the true flight of deposits from the banking system in NDI states.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-county-level-banking-capacity-test-address-the-censoring-problem"&gt;Q7. How does the county-level banking capacity test address the censoring problem?&lt;/h3&gt;
&lt;p&gt;The paper estimates log-ratio regressions comparing county-level deposits at state-chartered banks in DI versus NDI border counties, using a &amp;ldquo;DI Active&amp;rdquo; indicator that switches on when deposit insurance is in effect in a given state-year and switches off when schemes are discontinued. Because different states discontinued their insurance at different times, there is sufficient within-county variation to identify the DI coefficient even with year fixed effects. The estimated coefficient of 0.574 (without year FE) and 0.557 (with year FE) translates to approximately a 56 percent higher deposit level in state-chartered bank counties with deposit insurance, with virtually identical estimates across specifications.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-placebo-test-for-national-banks-and-what-does-it-show"&gt;Q8. What is the placebo test for national banks, and what does it show?&lt;/h3&gt;
&lt;p&gt;National banks were prohibited by the Office of the Comptroller of the Currency from participating in state deposit insurance schemes. If deposit insurance — rather than unobserved county characteristics — is responsible for the 56 percent banking capacity premium, then county deposits at national banks in DI states should show no corresponding premium. The Table 6 results confirm this: the DI Active coefficient for national bank deposits is positive (0.165 to 0.267) but statistically insignificant, providing a falsification result consistent with the causal interpretation for state-chartered banks.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-situate-deposit-insurances-stabilizing-benefits-relative-to-its-moral-hazard-costs"&gt;Q9. How does the paper situate deposit insurance&amp;rsquo;s stabilizing benefits relative to its moral hazard costs?&lt;/h3&gt;
&lt;p&gt;The paper explicitly frames its contribution as quantifying the stability-enhancing component of deposit insurance separately from the moral hazard component. It cites extensive prior literature (Calomiris 1992, 1993; Wheelock 1992, 1993; Wheelock and Wilson 1994) establishing that the 1910s–1920s state schemes generated moral hazard: insured banks reduced capital-to-asset ratios, relaxed lending standards, and increased risk exposure. The paper does not contest those findings but argues that the two effects are analytically separable and that the stabilization benefit had significant quantitative magnitude — a benefit that should be accounted for when assessing the net welfare effects of deposit insurance design.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-implications-for-the-basel-iii-liquidity-coverage-ratio-framework"&gt;Q10. What are the implications for the Basel III Liquidity Coverage Ratio framework?&lt;/h3&gt;
&lt;p&gt;The Basel III LCR formula assumes that during distress 3 percent of &amp;ldquo;stable deposits&amp;rdquo; and 10 percent of &amp;ldquo;less stable deposits&amp;rdquo; run off. The paper&amp;rsquo;s finding that deposit insurance is associated with a 56 percent increase in banking capacity implies that in the absence of insurance, deposit runoffs are far higher than these Basel assumptions — substantially larger than 10 percent and more consistent with the 25–50 percent runoffs observed for non-systemic banks in Denmark following an insurance limit reduction (Iyer et al. 2016). The authors argue their results suggest that empirical grounding for the LCR runoff assumptions remains insufficient, consistent with critiques by Allen (2014) and Diamond and Kashyap (2016).&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Postal Savings System (as &amp;ldquo;mattress money&amp;rdquo; proxy).&lt;/strong&gt; The U.S. Postal Savings System (1911–) accepted deposits up to $2,500 per individual, backed by the full faith and credit of the United States. In this paper, postal savings deposits are used as a quantitative proxy for money withdrawn from the banking system during distress — &amp;ldquo;money under the mattress&amp;rdquo; — validated by cointegration with the currency-deposit ratio.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy discontinuity / border-pair design.&lt;/strong&gt; The identification strategy exploits the fact that deposit insurance was adopted at the state level, creating a sharp policy discontinuity at state borders. Contiguous city pairs straddling DI and NDI state borders are treated as quasi-experimental units, with the within-pair difference in postal savings deposit growth serving as the outcome, controlling for time-invariant city-level heterogeneity and common time effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Relative Postal Savings Deposit Growth (RPS).&lt;/strong&gt; The dependent variable defined as the log-ratio of postal savings deposits in the NDI city to postal savings deposits in the DI city within a pair, and then first-differenced over time. This construction controls for city-pair-level time-invariant characteristics and isolates the differential response to bank suspensions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bank suspension.&lt;/strong&gt; In this paper&amp;rsquo;s context, a bank suspension is any closure of a bank (state-chartered or national) at a specific geographic location, as recorded in FDIC manuscript lists compiled by Clark Warburton during the 1930s. The variable used in regressions is the change in the number of suspensions within R miles (R = 10, 20, 30) of the paired postal savings offices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Financial depth / local banking capacity.&lt;/strong&gt; The paper uses county-level deposits at state-chartered banks as a measure of local banking market size. Deposit insurance is hypothesized to increase financial depth by preventing the diversion of funds out of the banking system during distress, and the 56 percent estimated premium is the paper&amp;rsquo;s primary measure of the insurance&amp;rsquo;s capacity-enhancing effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DI Active indicator.&lt;/strong&gt; A time-varying binary variable equal to 1 when deposit insurance was legally in effect in a given state at a given time, and 0 otherwise (including after repeal). Because different states repealed their schemes at different times (Oklahoma 1923, Texas 1927, South Dakota 1927, North Dakota 1929, Kansas 1929, Nebraska 1930, Mississippi 1930), this variable provides within-county variation that identifies the banking capacity coefficient after controlling for county and year fixed effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Moral hazard vs. stability-enhancing components.&lt;/strong&gt; The paper distinguishes analytically between the moral hazard effect of deposit insurance (insured banks undertake riskier projects, reduce capital buffers, relax lending standards) and the stability-enhancing effect (depositors retain funds in the banking system, preventing runs). The paper&amp;rsquo;s contribution is to quantify the latter component in isolation, using a setting where the two effects can be separated by focusing on depositor — rather than banker — behavior.&lt;/p&gt;</description></item><item><title>Dynamic Regulation with Firm Linkages: Evidence from Texas</title><link>https://macropaperwarehouse.com/papers/dynamic-regulation-with-firm-linkages-evidence-from-texas/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/dynamic-regulation-with-firm-linkages-evidence-from-texas/</guid><description>&lt;p&gt;This paper evaluates the efficiency of linked environmental regulation, a targeting mechanism whereby inspectors who discover violations at one plant can increase enforcement pressure on other plants sharing the same owner. The central research question is whether linking inspection decisions across co-owned plants adds value over unlinked, plant-level targeting and over random enforcement. The paper develops a new empirical framework of dynamic moral hazard under linked regulation, applies it to Texas environmental enforcement data, and uses the estimated model to evaluate counterfactual regulatory designs.&lt;/p&gt;
&lt;p&gt;The empirical setting is the Texas Commission on Environmental Quality (TCEQ), which enforces the Resource Conservation and Recovery Act (RCRA, governing hazardous waste) and the Clean Water Act using a two-dimensional scoring system. A plant-level &amp;ldquo;site rating&amp;rdquo; score captures the individual plant&amp;rsquo;s compliance history, while a firm-wide &amp;ldquo;person rating&amp;rdquo; score aggregates the weighted average of plant scores across all plants under the same manager. Both scores feed into a multiplicative penalty escalation rule and a logit-form inspection probability function. The data are an unbalanced panel of 9,792 plants from 2012–2020, with detailed records of inspections, violations, penalties, scores, and ownership. The average plant is inspected with probability 0.289 per year and is linked with approximately 2 other plants through common ownership, though some firms own portfolios exceeding 50 plants.&lt;/p&gt;
&lt;p&gt;The model features firms endowed with private types (abatement cost parameters) that may be affiliated within a firm&amp;rsquo;s portfolio, choosing continuous pollution actions to maximize discounted payoffs net of expected penalties. The regulator observes only scores and minimizes social costs subject to a binding inspection budget. A key computational innovation is &amp;ldquo;continuation value sufficiency&amp;rdquo;: because fully solving the portfolio optimization over large plant sets is infeasible due to the curse of dimensionality, each plant&amp;rsquo;s decision is approximated using three state variables — its own plant score, the firm-wide score, and a scalar summarizing other co-owned plants&amp;rsquo; continuation values — governed by an AR(1) transition process. Estimation proceeds in three stages: OLS/logit for inspection and penalty parameters, simulated method of moments for type distribution and curvature parameters, and inversion of the regulator&amp;rsquo;s first-order conditions to recover sector-specific marginal social harms.&lt;/p&gt;
&lt;p&gt;Descriptive evidence confirms three preconditions for linked regulation to add value: violations are positively correlated within firm portfolios, inspections are targeted toward higher-scoring plants on both dimensions, and higher inspection probabilities (instrumented by scores) are associated with fewer violations conditional on plant fixed effects. The coefficient on predicted inspection probability in the deterrence regression (specification 3, plant fixed effects, inspected years only) is −3.920, and an increase in log scores from 0 to 1.5 (roughly the interquartile range) reduces expected violations by approximately 0.5.&lt;/p&gt;
&lt;p&gt;Structural estimates show that plant-level and firm-level type variance are similar (σ²_J = 0.209, σ²_F = 0.275), indicating moderate within-firm cost correlation. The curvature parameter y = 0.403 governs diminishing returns to negligence. In counterfactual experiments centered on a 30% budget increase (approximately 10 percentage point rise in per-plant inspection probability), unlinked plant-score-based escalations reduce social costs by 31.9% relative to random inspections. Linked firm-score-based escalations reduce social costs by 41.8% relative to random. The optimal mix — approximately 40% unlinked and 60% linked — reduces social costs by 42.2% relative to random. A back-of-the-envelope cost-benefit calculation calibrating utility-sector violation costs at $3,157 per violation and inspection costs at $740 finds a return of $11.77 in avoided social costs per additional dollar spent on inspections under the optimal mixed regime, versus $8.28 under random inspections.&lt;/p&gt;
&lt;p&gt;The scope conditions are specific: the framework applies to RCRA and Clean Water Act plants in Texas, which typically cannot reallocate production across facilities (unlike Clean Air Act firms), so the pollution-substitution channel documented for multi-plant Clean Air Act firms is not modeled. The penalty schedule is taken as fixed; only inspection allocation is treated as a policy choice.&lt;/p&gt;
&lt;p&gt;Q: What is linked regulation and why might it improve on unlinked enforcement?
A: Linked regulation allows the regulator to increase inspection and penalty pressure on all plants owned by a firm when any one plant accumulates violations. It is efficient when compliance costs (types) are correlated within firms — e.g., due to managerial practices — because a violation at one plant is informative about likely violations at co-owned plants. This correlation means the regulator can target scarce inspection resources toward portfolios that are likely to harbor multiple bad actors, rather than inspecting each plant independently.&lt;/p&gt;
&lt;p&gt;Q: How does Texas implement linked regulation in practice?
A: Texas uses a two-dimensional scoring system. The plant score (&amp;ldquo;site rating&amp;rdquo;) summarizes the individual plant&amp;rsquo;s violation history over the past five years, normalized by complexity points. The firm score (&amp;ldquo;person rating&amp;rdquo;) is the complexity-weighted average of plant scores across all plants under the same manager. Penalties are then multiplied by escalation factors based on both scores: a firm in the &amp;ldquo;unsatisfactory performer&amp;rdquo; tier (firm score ≥ 55) faces a 1.1× firm escalation, while a &amp;ldquo;high performer&amp;rdquo; (firm score &amp;lt; 0.1) faces a 0.9× multiplier. Because the firm escalation applies to all plants in the portfolio simultaneously, even a small change in firm score can produce large aggregate deterrence effects across a large portfolio.&lt;/p&gt;
&lt;p&gt;Q: What descriptive evidence supports the preconditions for linked regulation to add value?
A: Three pieces of evidence are presented. First, a scatterplot (Figure 1) shows a positive cross-sectional correlation between a plant&amp;rsquo;s average violations per inspection and the leave-one-out average violations per inspection of its co-owned plants, indicating within-firm cost correlation. Second, Table 2 logit regressions show that both plant score (coefficient 0.121) and firm score (coefficient 0.062) significantly predict inspection probability, conditional on year and NAICS fixed effects. Third, Table 3 shows that conditional on plant fixed effects, predicted inspection probability is negatively associated with violations (coefficient −3.246 in specification 2, rising to −3.920 in specification 3 restricted to inspected plant-years), confirming dynamic deterrence.&lt;/p&gt;
&lt;p&gt;Q: What is the curse of dimensionality problem and how is it resolved?
A: In a multi-plant firm, each plant&amp;rsquo;s optimal action depends on the scores of every other co-owned plant, producing a state space of dimension n_plants + 1. For firms with portfolios of 50+ plants this is computationally infeasible. The paper introduces &amp;ldquo;continuation value sufficiency&amp;rdquo;: each plant&amp;rsquo;s decision is reduced to three state variables — its own score s_j, the firm score s_f, and a scalar W_j aggregating other co-owned plants&amp;rsquo; continuation values. Transitions are approximated by plant-specific AR(1) processes. This reduces the portfolio problem from one high-dimensional value function to n_plant separate three-dimensional value functions, each solved independently within an inner fixed-point loop.&lt;/p&gt;
&lt;p&gt;Q: How are the type distribution parameters identified?
A: The mean type for each NAICS sector θ̄_g is identified by average violations per inspection within that sector — a higher mean type implies more violations conditional on inspection. The plant-level type variance σ²_J is identified by the share of total violation variance occurring across plants within the same firm. The firm-level type variance σ²_F is identified by the share of total violation variance occurring across firms. The curvature parameter y is identified by the responsiveness of violations to changes in predicted inspection probability (the coefficient from specification 3 of Table 3, which equals −3.920 empirically and −6.095 in simulation moments).&lt;/p&gt;
&lt;p&gt;Q: What are the main counterfactual results?
A: A 30% increase in the inspection budget (approximately +10 percentage points in per-plant inspection probability) is allocated under four regimes. Random inspections reduce violations per plant by 0.31 from a baseline of 0.98. Unlinked (plant-score) escalations reduce social costs by 31.9% more than random. Linked (firm-score) escalations reduce social costs by 41.8% more than random. The optimal mix (approximately 40% unlinked, 60% linked) reduces social costs by 42.2% more than random. In detected violations, all three targeted regimes perform similarly (+0.7% detected violations versus random), meaning the social cost advantage of linked regulation comes through greater undiscovered deterrence rather than through detection rates.&lt;/p&gt;
&lt;p&gt;Q: How does the decomposition into static, own-plant, and cross-plant effects clarify the mechanism?
A: For unlinked escalations: the static effect accounts for −5.4% of social cost relative to random, own-plant dynamic deterrence accounts for −30.6%, and the cross-plant effect is +4.1% (slightly adverse, because unlinked escalations do not account for portfolio-level incentives). For linked escalations: the static effect is −2.4%, own-plant deterrence is −24.5% (smaller than unlinked because linked escalations are less precisely targeted to individual plant histories), and cross-plant deterrence is −14.9% (large and beneficial). The dominance of cross-plant deterrence under linked escalations is the key mechanism explaining why linking outperforms unlinked targeting.&lt;/p&gt;
&lt;p&gt;Q: What does the cost-benefit calculation find?
A: Calibrating utility-sector violation social costs at $3,157 per violation (from Kang and Silveira 2021 for California water utilities post-2006) and inspection costs at $740, the paper finds a return of $11.77 in avoided social costs per additional dollar spent on inspections under the optimal linked/unlinked mix, versus $8.28 under random inspections. This suggests a large return to expanding enforcement budgets, with the gain amplified substantially by optimal targeting design.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions and limitations acknowledged?
A: The framework applies to RCRA and Clean Water Act plants in Texas, where firms (e.g., gas station chains) typically cannot reallocate production across facilities, so the pollution-substitution channel documented by Gibson (2019) for Clean Air Act firms is not modeled. The penalty schedule is taken as fixed — only inspection allocation is treated as a policy choice — because Texas&amp;rsquo;s bylaws are prescriptive about how violations translate into penalties while leaving inspection targeting largely to regulator discretion. Social harm parameters h_g are identified only up to a scale normalization. The paper also does not model why types are correlated within firms (bad managers versus specialization), as the counterfactual results depend only on the degree of correlation, not its source.&lt;/p&gt;
&lt;p&gt;Q: How well does the model fit the data?
A: The model matches the targeted moments well (Table 5). Mean violations by NAICS sector are closely reproduced (e.g., utility: 0.201 empirical vs. 0.184 simulated; trade: 0.252 vs. 0.236). Responsiveness of violations to inspection probability matches closely (−6.398 empirical vs. −6.095 simulated). A non-targeted fit statistic — the correlation between a plant&amp;rsquo;s own violation rate and its co-owned plants&amp;rsquo; violation rates — is 0.32 in simulation versus 0.26 in the data, which the authors characterize as a good out-of-sample fit given it was not directly targeted in estimation.&lt;/p&gt;
&lt;p&gt;Q: How do heterogeneous effects shed light on the distributional consequences of regulation?
A: The own-plant deterrence effect is positive for all plants including those with low types that are unlikely to be targeted, but is especially pronounced for high-type plants under unlinked escalations. Under linked escalations, high-type plants are deterred less to the extent they are co-owned with lower-type plants, because firm-score-based targeting aggregates across the portfolio. Cross-plant effects are predictably small under unlinked escalations and larger under linked escalations, especially for firms with high-type portfolios, since those are the firms whose firm scores respond most to individual violations.&lt;/p&gt;
&lt;p&gt;Linked regulation: An enforcement mechanism in which the discovery of violations at one plant triggers increased inspection and penalty pressure on all other plants under the same owner. It exploits within-firm correlation in compliance costs to target scarce regulatory resources more efficiently than plant-by-plant escalation alone.&lt;/p&gt;
&lt;p&gt;Escalation mechanism: A penalty and inspection design in which plants with worse compliance records — measured by accumulated compliance scores — face disproportionately greater scrutiny and higher penalties per additional violation. The TCEQ&amp;rsquo;s two-dimensional scoring system is an escalation mechanism operating simultaneously at the individual plant and firm portfolio level.&lt;/p&gt;
&lt;p&gt;Plant score / firm score: The plant score (&amp;ldquo;site rating&amp;rdquo;) is a normalized index of a single facility&amp;rsquo;s violation history over the past five years, divided by investigation count and complexity points; the firm score (&amp;ldquo;person rating&amp;rdquo;) is the complexity-weighted average of all plant scores across the firm&amp;rsquo;s portfolio. Higher scores indicate worse compliance records and trigger both higher penalties and higher inspection probabilities.&lt;/p&gt;
&lt;p&gt;Continuation value sufficiency: The paper&amp;rsquo;s solution to the curse of dimensionality in large plant portfolios. Rather than tracking the full joint score state across all co-owned plants, each plant&amp;rsquo;s optimal action is approximated using three variables — its own score, the aggregate firm score, and a scalar W_j summarizing co-owned plants&amp;rsquo; continuation values — with state transitions governed by a plant-specific AR(1) process.&lt;/p&gt;
&lt;p&gt;Dynamic moral hazard under linked regulation: The firm&amp;rsquo;s problem of choosing how much to invest in pollution mitigation at each plant over time, given that current actions affect future scores, future penalties, and — through the firm-wide score — future scrutiny of all co-owned plants. The moral hazard arises because abatement costs are private information not directly observable by the regulator.&lt;/p&gt;
&lt;p&gt;Complexity points: A normalization factor in the TCEQ scoring system that adjusts raw violation counts for plant size and sector, enabling comparable compliance histories across heterogeneous facilities. They were introduced in 2012 specifically to prevent mechanically larger facilities from appearing riskier simply due to their scale.&lt;/p&gt;
&lt;p&gt;Cross-plant deterrence effect: The reduction in pollution actions at co-owned plants induced by increases in the firm-wide score following a violation at one plant in the portfolio. In the counterfactual decomposition, this effect accounts for −14.9 percentage points of social cost reduction under linked escalations and is the primary mechanism by which linked regulation outperforms unlinked plant-level escalation.&lt;/p&gt;</description></item><item><title>Efficiency Criteria, Income Taxation, and Heterogeneous Elasticities</title><link>https://macropaperwarehouse.com/papers/efficiency-criteria-income-taxation-and-heterogeneous-elasticities/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/efficiency-criteria-income-taxation-and-heterogeneous-elasticities/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Can income tax schedules be justified as utilitarian-optimal without adopting extreme normative assumptions about how household welfare should be measured? The paper proposes a welfare criterion strictly stronger than Pareto efficiency—called &lt;em&gt;rationalizability with bounded curvature&lt;/em&gt;—and asks whether observed US income taxes satisfy it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Starting Point.&lt;/strong&gt; Any Pareto-efficient nonlinear income tax schedule can, in principle, be rationalized as utilitarian-optimal under &lt;em&gt;some&lt;/em&gt; cardinalization of household utilities (i.e., some choice of how to measure the cardinal scale of each household&amp;rsquo;s well-being). However, the paper shows that rationalizing Pareto-efficient taxes in this way often requires cardinalizations under which there is &lt;em&gt;no&lt;/em&gt; population upper bound on the curvature of utility with respect to consumption. Equivalently, a utilitarian planner&amp;rsquo;s marginal willingness to transfer resources to households must fall arbitrarily quickly with the size of those transfers—an extreme form of status quo bias violated by virtually all quantitative optimal-tax exercises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Proposed Criterion.&lt;/strong&gt; The authors restrict attention to cardinalizations with &lt;em&gt;locally bounded curvature&lt;/em&gt;: there exists a finite (though potentially arbitrarily large) upper bound on the coefficient of relative risk aversion across the population. This admits two interpretations: (i) ex post, it requires that the social value of transfers not change arbitrarily quickly with transfer size; (ii) ex ante, it corresponds to a decision-maker behind a veil of ignorance with bounded risk aversion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Theoretical Result.&lt;/strong&gt; Within a standard Mirrlees model of nonlinear income taxation with arbitrary preference heterogeneity and intensive-margin labor supply, the paper proves that a tax schedule can be rationalized with bounded curvature if and only if government revenues are both &lt;em&gt;decreasing and concave&lt;/em&gt; (not merely decreasing) with respect to a class of narrowly targeted &amp;ldquo;two-bracket&amp;rdquo; reforms—reforms that raise retention by $1 local to some income level $z$ and zero elsewhere. This contrasts with Pareto efficiency, which requires only that revenues be decreasing in these reforms (Bierbrauer, Boyer, and Hansen 2023). The additional requirement of revenue concavity is what distinguishes the bounded-curvature criterion from pure Pareto efficiency.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sufficient Statistics.&lt;/strong&gt; The paper derives explicit sufficient-statistics expressions for the first- and second-order derivatives of tax revenue with respect to these targeted reforms. The second derivative depends on higher moments of the elasticity distribution, specifically the &lt;em&gt;income-conditional variance&lt;/em&gt; of compensated elasticities of taxable income (ETIs). Revenue convexity—which causes the second-order condition to fail—arises when income-conditional ETI variance is sufficiently high, even holding the mean ETI fixed. The economic mechanism is a &amp;ldquo;sort-and-extort&amp;rdquo; dynamic: a small tax reform sorts higher-elasticity households into income brackets where marginal taxes fall and lower-elasticity households into brackets where marginal taxes rise; repeating the reform then exploits this sorting by differentially taxing households by elasticity, as if applying group-specific tax schedules within a uniform income tax.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Findings.&lt;/strong&gt; Using the NBER panel of US tax returns from 1979 to 1990, the paper estimates income-conditional mean ETIs of approximately 0.2–0.3 at most income levels. Crucially, it estimates a &lt;em&gt;lower bound&lt;/em&gt; on income-conditional ETI variance by comparing elasticities of light versus heavy itemizers (defined by whether a household claims above or below the mean value of deductions in its income bracket). The low-elasticity group has an ETI of approximately zero and the high-elasticity group has an ETI of approximately one, implying a lower bound on ETI variance of roughly 0.2 at most incomes and approximately 0.25 at the top of the distribution. This lower bound is close to—and under plausible assumptions above—the threshold required for the second-order condition to fail. The authors conclude that the US income tax schedule in 1990 was likely Pareto efficient but likely &lt;em&gt;not&lt;/em&gt; rationalizable with bounded curvature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Welfare Gains.&lt;/strong&gt; In a calibrated model with a 50% top marginal tax rate, Pareto-tail shape of 2.5, mean ETI of 0.3, and ETI standard deviation of 0.75 (50% above the estimated lower bound), the planner gains significant welfare from either raising or lowering top marginal taxes. The welfare-maximizing top rate below the baseline is 13.3%, generating social value equivalent to a transfer of $1,966 per top earner. The welfare-maximizing top rate above the baseline is 71.2%, generating social value equivalent to a transfer of $972 per top earner. The revenue-maximizing rate is 80.9% under the baseline calibration, ranging from 74.6% to 86.8% as ETI standard deviation varies by ±25% of the lower bound.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; The theoretical analysis is restricted to intensive-margin labor supply (abstracting from extensive-margin decisions); the empirical application focuses on top incomes where extensive-margin effects are likely small. The empirical period is 1979–1990, covering major federal and state tax reforms. Results concern local efficiency of the tax schedule, not global optimization.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-exactly-is-rationalizability-with-bounded-curvature-and-how-does-it-differ-from-pareto-efficiency"&gt;Q1. What exactly is &amp;ldquo;rationalizability with bounded curvature&amp;rdquo; and how does it differ from Pareto efficiency?&lt;/h3&gt;
&lt;p&gt;A: Pareto efficiency requires that no small reform makes someone better off without making anyone worse off. Rationalizability (with &lt;em&gt;any&lt;/em&gt; cardinalization) is equivalent to Pareto efficiency in this setting. Rationalizability with bounded curvature additionally restricts the cardinalization: there must exist a finite upper bound on the coefficient of relative risk aversion (or equivalently, on the curvature of utility with respect to consumption) across the population. This is a strictly stronger criterion than Pareto efficiency. A schedule can be Pareto efficient but not rationalizable with bounded curvature if the only cardinalizations that rationalize it require unbounded consumption utility curvature.&lt;/p&gt;
&lt;h3 id="q2-why-do-extreme-cardinalizations-with-unbounded-curvature-arise-when-rationalizing-pareto-efficient-taxes"&gt;Q2. Why do &amp;ldquo;extreme&amp;rdquo; cardinalizations with unbounded curvature arise when rationalizing Pareto-efficient taxes?&lt;/h3&gt;
&lt;p&gt;A: When a Pareto-efficient schedule is rationalized as utilitarian, the cardinalization must make the set of feasible, recardinalized utilities convex so it can be separated from the set of Pareto-improving allocations. The paper constructs such a cardinalization explicitly: it takes the form of a function whose second derivative approaches negative infinity as utility approaches its baseline value. This implies the planner&amp;rsquo;s marginal value of transfers to a household falls precipitously as the household is made even slightly better off—an extreme status quo bias. Theorem 2.b establishes that &lt;em&gt;all&lt;/em&gt; cardinalizations rationalizing a schedule with convex revenues must share this pathology.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-sort-and-extort-mechanism-and-how-does-it-generate-revenue-convexity"&gt;Q3. What is the &amp;ldquo;sort-and-extort&amp;rdquo; mechanism and how does it generate revenue convexity?&lt;/h3&gt;
&lt;p&gt;A: When elasticities of taxable income (ETIs) are heterogeneous within an income level and the income density is declining steeply, a reform that lowers marginal taxes around income $z$ brings more households into the local bracket (because there are more households just below $z$ than above). Crucially, it disproportionately attracts households with &lt;em&gt;higher&lt;/em&gt; ETIs, since they respond more strongly to the marginal tax cut and relocate from further away, where the density differs more. Repeating the reform therefore faces a higher-elasticity composition at $z$, generating larger positive behavioral effects—making revenues convex in the size of the reform. The second step (&amp;ldquo;extort&amp;rdquo;) involves raising taxes on the now-concentrated low-elasticity households at adjacent brackets, achieving as-if group-specific taxation within a single income tax schedule.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-precise-relationship-between-revenue-convexity-and-eti-variance"&gt;Q4. What is the precise relationship between revenue convexity and ETI variance?&lt;/h3&gt;
&lt;p&gt;A: The paper shows (Theorem 4) that the second-order revenue derivative with respect to a narrow two-bracket reform around income $z$ equals a positive function of the income density times the expression $-[1-R&amp;rsquo;_0(z)]\varepsilon(z) + [1-R&amp;rsquo;_0(z)]\alpha(z)[\varepsilon^2(z) + \text{var}_h[\varepsilon^h | z^h_0=z]]$. The first term is always negative (pushing toward revenue concavity). The second term, which includes the income-conditional variance of ETIs, can dominate and create revenue convexity when ETI variance is sufficiently large. In the benchmark case with a single household type at each income (no within-income heterogeneity), the variance term vanishes and revenues are always concave whenever decreasing.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-sufficient-statistics-test-for-rationalizability-at-the-top-of-the-income-distribution"&gt;Q5. What is the sufficient statistics test for rationalizability at the top of the income distribution?&lt;/h3&gt;
&lt;p&gt;A: At top incomes (assuming no income effects, no super-elasticities, and CES preferences), taxes are Pareto efficient if and only if $\tau_\text{top} &amp;lt; \frac{1}{1+\alpha_\text{top}\varepsilon_\text{top}}$, and they are rationalizable with bounded curvature if and only if additionally $\tau_\text{top} &amp;lt; \frac{2}{1+\alpha_\text{top}(\varepsilon_\text{top} + \sigma^2_\text{top}/\varepsilon_\text{top})}$, where $\tau_\text{top}$ is the top marginal tax rate, $\alpha_\text{top}$ is the Pareto tail shape, $\varepsilon_\text{top}$ is the mean ETI at the top, and $\sigma^2_\text{top}$ is the income-conditional ETI variance at the top.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-estimate-a-lower-bound-on-income-conditional-eti-variance"&gt;Q6. How does the paper estimate a lower bound on income-conditional ETI variance?&lt;/h3&gt;
&lt;p&gt;A: The authors divide households at each income level into &amp;ldquo;heavy&amp;rdquo; and &amp;ldquo;light&amp;rdquo; itemizers based on whether their total deductions exceed the local income-bracket mean. They then estimate group-specific ETIs using local polynomial regressions of log income changes on log marginal retention changes, interacting tax changes with heavy-itemizer indicators. The within-year difference in elasticities between groups provides a lower bound on within-income ETI variance, since the two-group decomposition captures only a fraction of true variance. The interaction coefficient is allowed to vary by year to isolate within-year, within-income variation in elasticities rather than between-year compositional changes.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-estimated-magnitudes-of-mean-and-variance-of-etis"&gt;Q7. What are the estimated magnitudes of mean and variance of ETIs?&lt;/h3&gt;
&lt;p&gt;A: Income-conditional average ETIs are estimated at between 0.2 and 0.3 at most income levels, consistent with but somewhat below prior literature estimates. The low-elasticity group (light itemizers) has an ETI of approximately zero, while the high-elasticity group (heavy itemizers) has an ETI of approximately one. Given roughly equal group sizes, this implies a lower bound on ETI variance of approximately 0.2 at most incomes and approximately 0.25 at the ninety-fifth percentile. Subdividing the high-elasticity group into two, three, and four subgroups yields a lower bound of approximately 0.25 for variance at the top.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-back-of-the-envelope-calculation-work-to-assess-whether-the-second-order-test-fails"&gt;Q8. How does the back-of-the-envelope calculation work to assess whether the second-order test fails?&lt;/h3&gt;
&lt;p&gt;A: With $\tau_\text{top} \approx 0.5$, $\alpha_\text{top} \approx 2.5$, and $\varepsilon_\text{top} \approx 0.3$ (from prior literature), the second-order condition fails if and only if ETI variance exceeds approximately 0.27. The authors&amp;rsquo; lower bound estimate of ETI variance is already approximately 0.25 (standard deviation approximately 0.5), just below this threshold. The authors note that if the true standard deviation exceeds the lower bound by more than 4%, the second-order condition fails, making it empirically likely that the 1990 US tax schedule was not rationalizable with bounded curvature.&lt;/p&gt;
&lt;h3 id="q9-why-does-the-paper-focus-on-the-top-of-the-income-distribution-for-the-empirical-test"&gt;Q9. Why does the paper focus on the top of the income distribution for the empirical test?&lt;/h3&gt;
&lt;p&gt;A: The second-order condition is most likely to fail at high incomes for three reasons simultaneously: (i) the marginal tax rate is highest, (ii) ETI means are somewhat higher there, and (iii) the Pareto parameter $\alpha(z)$ is largest (income density falls steeply), which amplifies the sort-and-extort mechanism. The authors also note that extensive-margin labor supply responses—which are abstracted away in the theory—are likely small at high incomes.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-calibrated-quantitative-application-reveal-about-optimal-top-tax-policy"&gt;Q10. What does the calibrated quantitative application reveal about optimal top tax policy?&lt;/h3&gt;
&lt;p&gt;A: Calibrated with a 50% initial top marginal tax rate, Pareto tail shape of 2.5, mean ETI of 0.3, and ETI standard deviation of 0.75 (50% above the estimated lower bound), the model finds welfare gains in both directions of reform. The welfare-maximizing rate &lt;em&gt;below&lt;/em&gt; the baseline is 13.3%, yielding equivalent welfare gains of $1,966 per top earner. The welfare-maximizing rate &lt;em&gt;above&lt;/em&gt; the baseline is 71.2%, yielding equivalent gains of $972 per top earner. The revenue-maximizing rate is 80.9%, ranging from 74.6% to 86.8% when ETI standard deviation varies by ±25% of the lower bound. This sensitivity highlights that the optimal direction and magnitude of reform depend substantially on the uncertain degree of ETI heterogeneity.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-relate-to-the-inverse-optimum-literature"&gt;Q11. How does the paper relate to the &amp;ldquo;inverse optimum&amp;rdquo; literature?&lt;/h3&gt;
&lt;p&gt;A: The inverse optimum approach (Bourguignon and Spadaro 2012; Hendren 2020) infers the first-order welfare trade-offs implicit in an observed tax schedule. This paper goes further by inferring from second-order empirical moments—specifically the income-conditional ETI variance—whether taxes are consistent with &lt;em&gt;minimal&lt;/em&gt; requirements on how sensitive the planner&amp;rsquo;s trade-offs are to household welfare levels. Rather than assuming a welfare function, it tests whether &lt;em&gt;any&lt;/em&gt; welfare function with bounded curvature can rationalize the observed schedule.&lt;/p&gt;
&lt;h3 id="q12-is-revenue-convexity-possible-without-within-income-heterogeneity-in-preferences"&gt;Q12. Is revenue convexity possible without within-income heterogeneity in preferences?&lt;/h3&gt;
&lt;p&gt;A: Yes, but only under more specific conditions. The paper provides two supplemental examples. In the first, all households have constant-elasticity labor disutility but differ in both productivity and elasticity across income levels; when lower-income households have higher elasticities, a reform reducing marginal taxes at $z$ attracts higher-elasticity households and raises the average elasticity, leading to convex revenues. In the second, all households have the same initial elasticity but individual elasticities change in response to reforms. However, with the standard additively separable CES preferences and no within-income heterogeneity, revenues are always concave when decreasing—consistent with Werning&amp;rsquo;s (2007) observation that the Pareto planner&amp;rsquo;s problem is convex in this case.&lt;/p&gt;
&lt;h3 id="q13-what-is-the-role-of-random-tax-reforms-in-the-papers-logic"&gt;Q13. What is the role of random tax reforms in the paper&amp;rsquo;s logic?&lt;/h3&gt;
&lt;p&gt;A: Random tax reforms serve as an expository bridge. The paper shows that if the second-order revenue effect of a two-bracket reform is positive at some income $z$, then a &amp;ldquo;randomized&amp;rdquo; reform that applies the reform with equal probability in positive and negative directions generates an expected Pareto improvement—because the convexity of revenues implies expected revenues rise, while for any household with bounded risk aversion the reform&amp;rsquo;s second-order utility effect is also positive when the reform is sufficiently narrow. This establishes that revenue convexity implies random Pareto inefficiency under bounded risk aversion, and then the paper shows the analogous deterministic result for rationalizability.&lt;/p&gt;
&lt;h3 id="q14-what-scope-conditions-attach-to-the-sufficient-conditions-for-rationalizability-theorem-3"&gt;Q14. What scope conditions attach to the sufficient conditions for rationalizability (Theorem 3)?&lt;/h3&gt;
&lt;p&gt;A: Theorem 3 requires Assumptions 1 and 3 plus two boundary conditions: the ratio $\delta\text{Rev}(z)/(zg(z))$ must remain bounded away from zero as income approaches 0 or infinity, and at all incomes there must exist households with low enough compensated elasticities. Assumption 1 requires that average and marginal taxes have upper bounds below one, that marginal taxes have a lower bound, and that $zg(z)$ converges to zero at the boundaries. Assumption 3 is a regularity condition on how conditional moments of the elasticity distribution vary with income. These conditions ensure that the narrow, self-financing reforms considered in the necessity proof cannot generate welfare improvements once revenues are both decreasing and concave.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Rationalizability with Bounded Curvature.&lt;/strong&gt; The property that a tax schedule is utilitarian-optimal under some cardinalization of household utilities in which there exists a finite (though potentially arbitrarily large) upper bound on the curvature of utility with respect to consumption across the population. Formally, there exists a continuous function $\bar{\rho}$ such that, for all households, the absolute value of $[w_h \circ u_h]_{cc} / [w_h \circ u_h]_c$ is bounded by $\bar{\rho}$ evaluated at the household&amp;rsquo;s income. This criterion is strictly stronger than Pareto efficiency and strictly weaker than utilitarian optimality under a fixed cardinalization.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Two-Bracket Reform.&lt;/strong&gt; A targeted tax reform that increases retention (post-tax income) by $1 at incomes local to some level $z$ over a small bracket of width $\ell$, and zero elsewhere (smoothed at the edges). As $\ell \to 0$, this becomes an infinitesimally narrow reform. The first- and second-order revenue effects of these reforms—denoted $\delta\text{Rev}(z)$ and $\delta^2\text{Rev}(z)$—are the paper&amp;rsquo;s key objects: Pareto efficiency requires $\delta\text{Rev}(z) &amp;lt; 0$ for all $z$, and rationalizability with bounded curvature additionally requires $\delta^2\text{Rev}(z) \leq 0$ for all $z$.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Income-Conditional ETI Variance.&lt;/strong&gt; The variance of compensated elasticities of taxable income (ETIs) among households with the same income level, $\text{var}_h[\varepsilon^h | z^h_0 = z]$. This is the paper&amp;rsquo;s primary empirical object of interest and the key determinant of whether revenues are convex or concave in the size of targeted reforms. Unlike the literature&amp;rsquo;s focus on mean ETIs by income bracket, this within-income variance captures heterogeneity among households sharing the same pre-reform income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sort-and-Extort Mechanism.&lt;/strong&gt; The two-step economic mechanism underlying revenue convexity from ETI heterogeneity. In the first step (&amp;ldquo;sort&amp;rdquo;), a marginal tax cut around income $z$ disproportionately attracts higher-ETI households from lower incomes (because they respond more strongly and relocate from further away), shifting the elasticity composition at $z$ upward. In the second step (&amp;ldquo;extort&amp;rdquo;), repeating the reform finds higher-elasticity households concentrated where marginal taxes fall and lower-elasticity households where taxes rise, effectively applying differential tax treatment by elasticity within a single income tax schedule.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Pareto Parameter $\alpha(z)$.&lt;/strong&gt; Defined as $-d\log(zg(z))/d\log z$, where $g(z)$ is the income density. This captures the rate at which the income density is falling in income locally at $z$, and governs the strength of the sort-and-extort mechanism. High $\alpha(z)$ at top incomes (reflecting a steeply declining Pareto-type density) amplifies revenue convexity from ETI heterogeneity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Super-Elasticity.&lt;/strong&gt; A concept that captures how a household&amp;rsquo;s compensated ETI would change if its income were different, holding preferences fixed. Formally, it is the derivative of the household&amp;rsquo;s elasticity with respect to its log income, decomposing into effects from changes in preference curvature and changes in the local curvature of the tax schedule. Super-elasticities are zero in the benchmark case of additively CES preferences and locally CES retention schedules but contribute additional terms to the second-order revenue expression in the general case.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cardinalizing Function.&lt;/strong&gt; A strictly increasing function $w_h$ that maps household $h$&amp;rsquo;s indirect utility $V_h$ to a cardinalized utility level $w_h(V_h)$. The social planner maximizes the expectation of cardinalized utilities. Different choices of ${w_h}_h$ correspond to different stances on interpersonal comparisons, including unbounded curvature (rationalizing any Pareto-efficient schedule) or bounded curvature (the paper&amp;rsquo;s proposed restriction). Rawlsian social welfare is a limit of utilitarian welfare with increasingly concave cardinalizing functions.&lt;/p&gt;</description></item><item><title>Energy Transitions in Regulated Markets</title><link>https://macropaperwarehouse.com/papers/energy-transitions-in-regulated-markets/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/energy-transitions-in-regulated-markets/</guid><description>&lt;p&gt;This paper asks how rate-of-return (RoR) regulation in U.S. electricity markets affects the speed and efficiency of energy transitions, specifically the transition from coal to combined-cycle natural gas (CCNG) generation driven by fracking-induced cost declines. The authors build and estimate a structural model of regulated utility behavior in which utilities optimize investment, retirement, and hourly operations decisions against an incentive structure set by state Public Utility Commissions (PUCs).&lt;/p&gt;
&lt;p&gt;The regulatory environment combines two instruments: (1) an allowable rate of return that is decreasing in consumer electricity rates (incentive regulation), parameterized as s = (r/r₀)^{-γ}, where higher γ penalizes high-cost outcomes more severely; and (2) a &amp;ldquo;used-and-useful&amp;rdquo; standard in which a coal plant&amp;rsquo;s contribution to the rate base depends on its capacity utilization via a logit function. These two instruments create a tension: utilities want to lower costs to earn a higher RoR, but also want to run existing coal plants—even when uneconomical—to prove they are &amp;ldquo;used and useful&amp;rdquo; and thus maximize their rate base and profits.&lt;/p&gt;
&lt;p&gt;The authors estimate the model using publicly available EIA and EPA CEMS data spanning 2006–2017, covering 39 unique regulated utilities in the Eastern Interconnection across more than 4 million utility-hour observations (459 utility-years). Structural parameters are recovered via a nested fixed-point indirect inference approach that matches simulated regression coefficients to actual data; investment and retirement costs are estimated with a GMM nested fixed-point approach.&lt;/p&gt;
&lt;p&gt;Key reduced-form findings confirm the model&amp;rsquo;s two core mechanisms. First, a 10% increase in total variable costs is associated with a 2.5% decrease in variable profits per MW of capacity (with utility fixed effects), consistent with incentive regulation. Second, regulated utilities reduce coal generation by only a statistically insignificant 4.2 percentage points when coal fuel costs exceed import prices, compared to 16.1 percentage points for restructured utilities—consistent with regulated utilities running coal out-of-dispatch order to preserve used-and-useful status.&lt;/p&gt;
&lt;p&gt;In counterfactual simulations that impose 2018–20 natural gas prices ($2.01/MMBtu versus the 2006 price of $7.24/MMBtu) on utilities with their 2006 capital stocks, regulated utilities retire only 53% of coal capacity over 30 years and increase CCNG capacity by 296%, whereas a cost minimizer would retire most coal capacity while increasing CCNG by only 58%. The Averch-Johnson over-investment effect dominates: regulated utilities over-invest in CCNG while simultaneously over-using legacy coal.&lt;/p&gt;
&lt;p&gt;Carbon taxes on regulated utilities reduce short-run coal generation only 48% as much as when imposed on a cost minimizer (because the used-and-useful incentive partially offsets the carbon price signal), but in the long run result in 68% lower coal capacity and 77% lower coal generation relative to baseline by year 30—larger effects than for the cost minimizer. Eliminating the coal usage incentive (μ₂ = 0) produces 82% lower coal capacity and 92% lower coal generation over 30 years but requires utility variable profits to fall by over $300 million, threatening reliability without compensating transfers.&lt;/p&gt;
&lt;p&gt;Scope conditions: Results apply to regulated (non-restructured) utilities in the Eastern Interconnection, 2006–2017. The model estimates the coal-to-CCNG transition only; it explicitly does not model the ongoing transition to renewables and storage due to insufficient data variation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-central-research-question"&gt;Q1. What is the central research question?&lt;/h3&gt;
&lt;p&gt;The paper asks whether and how rate-of-return regulation in U.S. electricity markets slows energy transitions, and what alternative regulatory structures or carbon tax policies could accelerate the transition away from coal. It addresses this both theoretically—through a structural model of regulated utility behavior—and empirically, through estimation and counterfactual simulation using data on 39 regulated utilities over 2006–2017.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-key-regulatory-instruments-in-the-model-and-what-distortions-do-they-create"&gt;Q2. What are the two key regulatory instruments in the model, and what distortions do they create?&lt;/h3&gt;
&lt;p&gt;The first instrument is incentive regulation: the allowable rate of return declines as consumer electricity rates rise (s = (r/r₀)^{-γ}), so utilities have an incentive to lower costs. The second is the used-and-useful standard: a coal plant&amp;rsquo;s contribution to the rate base depends on its capacity utilization via a logit function, creating an incentive to run coal plants even when their fuel costs exceed import prices. Together, these instruments generate a tension between cost-reduction incentives and legacy-capacity-preservation incentives, causing the regulated utility to both over-invest in new CCNG capacity (Averch-Johnson effect) and over-use existing coal capacity relative to the cost-minimizing benchmark.&lt;/p&gt;
&lt;h3 id="q3-what-does-the-reduced-form-evidence-show-about-uneconomical-coal-usage"&gt;Q3. What does the reduced-form evidence show about uneconomical coal usage?&lt;/h3&gt;
&lt;p&gt;In a triple-difference specification, regulated utilities reduce coal generation by only 4.2 percentage points (statistically insignificant) when coal fuel costs exceed import prices, compared to a 16.1 percentage point reduction for restructured utilities. CCNG generation responds similarly under both regulatory regimes (21.1 vs. 19.7 percentage points), confirming that the distortion is specific to legacy coal under RoR regulation and not a general feature of high-cost generation. The six states with the largest responsiveness of coal usage to low market prices are all restructured states; out-of-dispatch-order coal generation also correlates strongly with utility ownership share across states.&lt;/p&gt;
&lt;h3 id="q4-what-do-the-structural-parameter-estimates-reveal-about-the-rate-base"&gt;Q4. What do the structural parameter estimates reveal about the rate base?&lt;/h3&gt;
&lt;p&gt;Each MW of CCNG capacity increases the rate base by $229,000. When fully utilized, each MW of coal capacity contributes 1.144 times as much as CCNG. When coal is not fully used, unused coal capacity contributes only 40% as much to the rate base as CCNG. NGT capacity contributes 79% more to the rate base than CCNG per MW. Operations cost estimates include O&amp;amp;M costs of $12.89/MWh for coal, $8.82/MWh for CCNG, and $44.63/MWh for NGT; a 100 MW coal ramp in one hour costs $4,770 versus $3,860 for CCNG.&lt;/p&gt;
&lt;h3 id="q5-what-happens-in-the-30-year-long-run-counterfactual-under-the-baseline-regulated-utility"&gt;Q5. What happens in the 30-year long-run counterfactual under the baseline regulated utility?&lt;/h3&gt;
&lt;p&gt;Facing a sudden drop to 2018–20 natural gas prices ($2.01/MMBtu vs. $7.24/MMBtu in 2006), regulated utilities retire 53% of coal capacity and increase CCNG capacity by 296% over 30 years. The Averch-Johnson over-investment effect dominates: utilities invest heavily in CCNG while retaining and using legacy coal far longer than a cost minimizer would. The social planner effectively eliminates coal generation immediately (99% reduction in the first period) and retires almost all coal capacity over the horizon.&lt;/p&gt;
&lt;h3 id="q6-how-does-a-cost-minimizer-behave-relative-to-the-regulated-utility-in-the-same-long-run-counterfactual"&gt;Q6. How does a cost minimizer behave relative to the regulated utility in the same long-run counterfactual?&lt;/h3&gt;
&lt;p&gt;A cost minimizer immediately reduces coal generation by 50% in the first period and retires most coal capacity over 30 years while increasing CCNG capacity by only 58%—versus the regulated utility&amp;rsquo;s 296% CCNG increase. Thirty years after the shock, the cost minimizer has retired 71% more coal capacity than the regulated utility. The cost minimizer&amp;rsquo;s much smaller CCNG expansion reflects that it does not face Averch-Johnson incentives to over-invest in rate-base capital.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-short-run-vs-long-run-impact-of-carbon-taxes-on-regulated-utilities-compared-to-cost-minimizers"&gt;Q7. What is the short-run vs. long-run impact of carbon taxes on regulated utilities compared to cost minimizers?&lt;/h3&gt;
&lt;p&gt;In the short run, carbon taxes on regulated utilities reduce coal generation only 48% as much as when imposed on a cost minimizer (34% vs. ~100% in immediate generation drop), because the used-and-useful incentive counteracts the carbon price signal. In the long run (30-year horizon), however, carbon taxes on regulated utilities result in 68% lower coal capacity and 77% lower coal generation relative to baseline—larger percentage reductions than for a cost minimizer—because the regulatory structure amplifies the retirement incentive over time once carbon costs erode the economic rationale for keeping coal in the rate base.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-short-run-operations-counterfactual-finding-for-carbon-taxes-in-the-sample-period"&gt;Q8. What is the short-run operations counterfactual finding for carbon taxes in the sample period?&lt;/h3&gt;
&lt;p&gt;Using each utility-year in the analysis sample, imposing carbon taxes on regulated utilities reduces carbon costs by only about $500 million relative to baseline—41% of the $1.3 billion carbon cost savings from imposing the same carbon taxes on a cost minimizer. Despite this limited carbon reduction, electricity rates nearly triple from $77.58/MWh to $224.18/MWh under the regulated utility with carbon taxes, as the utility passes through most carbon costs to consumers; regulated utility variable profits also fall by over $500 million.&lt;/p&gt;
&lt;h3 id="q9-what-happens-when-the-coal-usage-incentive-is-eliminated-μ--0"&gt;Q9. What happens when the coal usage incentive is eliminated (μ₂ = 0)?&lt;/h3&gt;
&lt;p&gt;Setting the coal usage incentive parameter μ₂ = 0 (eliminating the logit slope on capacity utilization) causes coal capacity to fall 82% and coal generation to fall 92% relative to baseline over 30 years—a slightly larger generation decline than for the cost minimizer. However, this comes at the cost of more than twice the CCNG capacity due to the Averch-Johnson effect, and requires utility variable profits to fall by over $300 million, raising reliability concerns unless accompanied by compensating transfers.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-papers-mechanism-relate-to-observed-differences-in-coal-exit-rates-between-regulated-and-restructured-states"&gt;Q10. How does the paper&amp;rsquo;s mechanism relate to observed differences in coal exit rates between regulated and restructured states?&lt;/h3&gt;
&lt;p&gt;Between 2006 and 2018, 26.0% of coal capacity exited in restructured states versus only 17.2% in regulated states—a gap the authors attribute primarily to the used-and-useful incentive structure in RoR regulation. The structural model quantifies how this regulatory feature specifically distorts coal usage and retirement decisions; it is not explained by demand or cost differences across states, as confirmed by the triple-difference evidence showing the gap is specific to coal (not CCNG) and to regulated (not restructured) utilities.&lt;/p&gt;
&lt;h3 id="q11-why-does-the-paper-argue-that-alternative-regulatory-adjustments-are-insufficient-to-replicate-cost-minimizing-transitions"&gt;Q11. Why does the paper argue that alternative regulatory adjustments are insufficient to replicate cost-minimizing transitions?&lt;/h3&gt;
&lt;p&gt;Changing regulatory parameters—such as increasing the coal usage incentive or adjusting the electricity rate penalty—does not come close to replicating the speed of the energy transition under a cost minimizer in the long-run simulations. Regulatory adjustments that do approach cost-minimizing outcomes (such as eliminating μ₂) require large reductions in utility variable profits sufficient to risk reliability, consistent with why the 2022 Inflation Reduction Act relied on substantial investment transfers rather than carbon taxes as its primary clean energy instrument.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-papers-identification-strategy"&gt;Q12. What is the paper&amp;rsquo;s identification strategy?&lt;/h3&gt;
&lt;p&gt;Identification exploits the sharp, exogenous decline in natural gas fuel prices from fracking, which had heterogeneous implications across utilities depending on their initial capital mixes (coal-heavy vs. CCNG-heavy). By comparing investment, retirement, and operations decisions across utilities and over time—particularly between utilities that had CCNG exposure before the price decline and those that did not—the authors recover the structural regulatory and cost parameters. The IV specification for reduced-form evidence uses the current natural gas price interacted with the utility&amp;rsquo;s initial CCNG generation share as an instrument for fuel and import costs.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-papers-explicit-limitations"&gt;Q13. What are the paper&amp;rsquo;s explicit limitations?&lt;/h3&gt;
&lt;p&gt;The paper estimates the coal-to-CCNG transition only and cannot speak to the transition to renewables and storage, because there is insufficient variation in the data to identify how regulators would treat CCNG as a legacy technology subject to used-and-useful standards, or how renewables and storage would contribute to the rate base. The authors note that over-investment in CCNG capacity may create future stranded asset problems for ratepayers and that usage incentives for CCNG are likely to further hinder the transition to renewables—but these are conjectures rather than estimated findings.&lt;/p&gt;
&lt;p&gt;Rate-of-return (RoR) regulation: A regulatory structure in which the PUC sets electricity rates so that utility revenues cover total variable costs plus an allowable return on the utility&amp;rsquo;s rate base (capital stock), with the allowable return parameterized as s = (r/r₀)^{-γ}, declining as consumer electricity rates rise.&lt;/p&gt;
&lt;p&gt;Used-and-useful standard: A prudence criterion under which a capital asset&amp;rsquo;s contribution to the rate base depends on its capacity utilization, modeled as a logit function of the generation-to-capacity ratio; fully used coal capacity contributes 1.144 times as much as CCNG per MW, while unused coal contributes only 40% as much.&lt;/p&gt;
&lt;p&gt;Rate base: The capital stock on which the PUC grants the utility its allowable rate of return; adjusted by prudence and used-and-useful assessments and described in the paper as &amp;ldquo;at best an arduous task&amp;rdquo; to quantify precisely.&lt;/p&gt;
&lt;p&gt;Averch-Johnson (AJ) over-investment effect: The tendency of regulated utilities to over-invest in capital because profits are proportional to the rate base; in this paper&amp;rsquo;s setting, this causes regulated utilities to increase CCNG capacity by 296% over 30 years following the natural gas price shock, compared to 58% for a cost minimizer.&lt;/p&gt;
&lt;p&gt;Incentive regulation: A modification of cost-plus RoR regulation in which the allowable rate of return declines as electricity rates rise; it provides efficiency incentives for cost reduction but does not achieve first-best outcomes and is insufficient to overcome the used-and-useful distortion for legacy coal.&lt;/p&gt;
&lt;p&gt;Out-of-dispatch-order generation: Running a generation unit when its fuel costs exceed the market import price; regulated utilities engage in this behavior with coal plants to maintain used-and-useful status and rate base contribution, whereas restructured utilities do not face this incentive.&lt;/p&gt;
&lt;p&gt;Nested fixed-point indirect inference: The estimation approach used to recover structural regulatory and operations parameters by minimizing the distance between regression coefficients from actual data and those from model-simulated data via a non-linear parameter search.&lt;/p&gt;</description></item><item><title>Environmental Consequences of Hydrocarbon Infrastructure Policy</title><link>https://macropaperwarehouse.com/papers/environmental-consequences-of-hydrocarbon-infrastructure-policy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/environmental-consequences-of-hydrocarbon-infrastructure-policy/</guid><description>&lt;p&gt;Covert and Kellogg study policies that aim to &amp;ldquo;keep carbon in the ground&amp;rdquo; by blocking fossil fuel infrastructure investment, with the Dakota Access Pipeline (DAPL) as their empirical application. DAPL moves more than 500,000 barrels per day of oil from the Bakken Shale of North Dakota to the U.S. Gulf Coast and was completed in June 2017 amid substantial opposition. The central research question is whether blocking pipeline construction actually keeps oil in the ground or merely shifts transport to alternative modes — specifically crude-by-rail — and what the net environmental and economic consequences are.&lt;/p&gt;
&lt;p&gt;The paper develops a two-period model of crude oil production and transportation mode choice. In the model, oil shippers decide in period 1 whether to commit to pipeline capacity under ship-or-pay contracts, then in period 2 allocate flows between the committed pipeline and the more flexible but costlier railroad alternative. Pipeline construction is an irreversible sunk cost with zero ongoing marginal cost; rail involves no sunk cost but substantial ongoing marginal costs including quadratic adjustment costs that capture capital investment in rail cars and loading/unloading facilities. Equilibrium pipeline capacity is determined by a shippers&amp;rsquo; indifference condition: expected per-barrel returns from pipeline access equal the FERC-regulated tariff.&lt;/p&gt;
&lt;p&gt;The empirical model is estimated using monthly Bakken oil production and transportation data, price differentials across three coastal destinations (Gulf, East, West), and drilling productivity data. Crude-by-rail marginal costs are estimated via 2SLS, yielding static marginal cost intercepts of $9.49/bbl to the East Coast, $12.64/bbl to the Gulf Coast, and $8.69/bbl to the West Coast, plus a dynamic adjustment cost of $1.28/bbl per mbbl/d of flow change. The upstream supply model follows Anderson, Kellogg, and Salant (2018), with old-well production following exponential decline (estimated decay parameter β = 0.955) and new-well drilling responding to current and lagged prices with a total long-run elasticity of 1.32. Shippers&amp;rsquo; beliefs about future oil prices are calibrated to an AR(1) process fit to historical price volatility (persistence φ₁ = 0.9925, volatility σ_G = 0.098). Model validation confirms a predicted expected return to pipeline commitment of $6.17/bbl against DAPL&amp;rsquo;s actual tariff of $5.50–$6.25/bbl.&lt;/p&gt;
&lt;p&gt;The main counterfactual asks what would have happened had DAPL&amp;rsquo;s construction been enjoined. In expectation, blocking DAPL reduces pipeline flows by 306 mbbl/d. Expected crude-by-rail flows increase by 248 mbbl/d, offsetting 81% of the pipeline reduction. Bakken oil production falls by only 58 mbbl/d, a 4% reduction. The modal shift from pipeline to rail worsens local environmental outcomes: per-barrel local pollution damages from rail transport substantially exceed those from pipelines, dominated by locomotive NOx emissions in populated areas. Foreclosing DAPL increases net local pollution damages by $444,000 per day (the decrease in pipeline-related harm of $144,000/day is more than offset by the increase from rail of $588,000/day). The total cost of blocking DAPL is $45/tonne of CO2 abated — $28/tonne from lost producer surplus and $17/tonne from increased local pollution damages — a figure comparable to the contemporaneous U.S. government social cost of carbon estimate of $42/tonne.&lt;/p&gt;
&lt;p&gt;An upstream production tax achieving the same CO2 reduction costs only $1.01–$2.68/tonne CO2 abated, an order of magnitude less, because it does not induce the distortionary modal shift to rail. Two caveats apply: if 57% of Bakken production reductions leak to other basins, the cost of blocking DAPL rises from $45/tonne to $104/tonne; and if reductions represent production delays rather than permanent reductions, effective abatement is further diminished. The analysis is scoped to Bakken crude oil and land transportation alternatives. The finding that blocking infrastructure increases local pollution is atypical of CO2 abatement policies, which usually generate local pollution co-benefits.&lt;/p&gt;
&lt;p&gt;Q: What is the core economic mechanism by which blocking a pipeline can keep oil in the ground?
A: When a pipeline is foreclosed, crude oil can still move by railroad, but rail transport involves substantial ongoing marginal costs. These costs create a wedge between upstream (Bakken) and downstream (Gulf Coast) prices that depresses upstream supply. Only when downstream prices are high enough to cover both rail marginal cost and this wedge will rail fully substitute for the pipeline; at lower prices, some production is uneconomical and stays in the ground. In the model, this price-depressing wedge is the mechanism that reduces production — but it operates only partially, since rail can substitute for much of the pipeline&amp;rsquo;s flow.&lt;/p&gt;
&lt;p&gt;Q: How much of the blocked pipeline flow substitutes to rail versus stays in the ground?
A: In expectation, blocking DAPL reduces pipeline flows by 306 mbbl/d. Expected crude-by-rail flows increase by 248 mbbl/d, offsetting 81% of the pipeline reduction. Bakken oil production falls by only 58 mbbl/d, or approximately 4%. In a specific simulated month (December 2019), 348 mbbl/d (67%) of the 520 mbbl/d of foregone pipeline flows would still move by rail.&lt;/p&gt;
&lt;p&gt;Q: How are crude-by-rail costs estimated, and what is the role of adjustment costs?
A: The authors estimate a 2SLS model of rail flows on price differentials, allowing for quadratic adjustment costs to capture investments and disinvestments in rail cars and loading facilities. Static marginal costs are $9.49/bbl (East Coast), $12.64/bbl (Gulf Coast), and $8.69/bbl (West Coast). The adjustment cost parameter γ is estimated at $1.28/bbl per mbbl/d, meaning a 10 mbbl/d monthly increase in rail flows raises marginal shipping cost by $12.76/bbl — a substantial share of total rail costs. Adjustment costs are necessary to reconcile the model with the sluggish observed response of rail flows to price differentials.&lt;/p&gt;
&lt;p&gt;Q: What is the structure of the upstream oil supply model and what are its key parameter estimates?
A: The model distinguishes &amp;ldquo;old&amp;rdquo; production from pre-existing wells, which follows exponential decline with estimated decay parameter β = 0.955, and &amp;ldquo;new&amp;rdquo; production from newly drilled wells, which is price-responsive with a total long-run elasticity of 1.32 — comparable to the 1.1–1.2 estimated by Newell and Prest (2019) across major U.S. shale plays. This structure implies that total production is highly inelastic in the short run (dominated by old wells) but responds to persistent price shocks over the long run through changes in drilling rates.&lt;/p&gt;
&lt;p&gt;Q: How do the local pollution damages of rail compare to those of pipeline transport?
A: At a social cost of carbon of $100/tonne, local air pollution damages from rail transport to the Gulf Coast are $1.66/bbl (plus $0.73/bbl in spill/accident costs), versus only $0.35/bbl local pollution (plus $0.11/bbl spills) for pipelines. Locomotive NOx emissions are the dominant factor, both because locomotives have high NOx emission factors and because these emissions often occur in densely populated areas. CO2 damages at $100/tonne SCC are roughly similar across modes ($0.79–0.83/bbl), so local pollution is the key differentiator.&lt;/p&gt;
&lt;p&gt;Q: What is the net welfare impact of foreclosing DAPL, and how is it decomposed?
A: Foreclosing DAPL reduces producer surplus by $716,000/day, increases net local pollution damages by $444,000/day (the $588,000/day increase from rail more than offsets the $144,000/day decrease from pipeline), and reduces CO2 emissions by 25.2 mtonnes/day from the 58 mbbl/d production reduction. The cost per tonne of CO2 abated is $28/tonne from lost producer surplus and $17/tonne from increased local pollution damages, totaling $45/tonne — broadly comparable to the U.S. government&amp;rsquo;s contemporaneous SCC estimate of $42/tonne. This means the policy&amp;rsquo;s abatement cost is approximately equal to the social value of each tonne abated, leaving little or no net social gain even before accounting for leakage.&lt;/p&gt;
&lt;p&gt;Q: How does the model validate against observed data and institutional parameters?
A: The model predicts an expected return to committed DAPL pipeline shipment of $6.17/bbl, which closely matches the actual DAPL tariff for committed shippers of $5.50–$6.25/bbl. The authors also validate simulated crude-by-rail flows against actual flows across destinations. The close match on the tariff is particularly meaningful because it tests the model&amp;rsquo;s equilibrium condition for pipeline capacity investment rather than a within-sample fit.&lt;/p&gt;
&lt;p&gt;Q: How does an upstream production tax compare to blocking DAPL as a policy instrument?
A: A production tax normalized to achieve the same CO2 reduction requires only $3.68/bbl if imposed after shippers have committed to DAPL (holding capacity fixed), or $3.24/bbl if announced before commitments are made (reducing pipeline capacity to 443 mbbl/d). The production tax reduces combined producer surplus and government revenue by only $96,000–$109,000/day versus $716,000/day under the DAPL ban, and reduces local pollution damages by $82,000/day rather than increasing them. The resulting cost per tonne CO2 abated is $1.01–$2.68 — an order of magnitude smaller than the $44.63/tonne for blocking DAPL.&lt;/p&gt;
&lt;p&gt;Q: What is the production leakage caveat and how large is its effect?
A: If blocking DAPL causes Bakken production to fall, production from other U.S. or global oil basins may increase, partially or fully offsetting the CO2 reduction. Following Prest (2022) and Prest et al. (2023), the authors note that if 57% of the Bakken production reduction leaks to other basins, the cost of blocking DAPL rises from $45/tonne to $104/tonne. Leakage would increase the cost per tonne for the upstream tax as well, but the relative advantage of the tax over the pipeline ban is unaffected by this caveat.&lt;/p&gt;
&lt;p&gt;Q: What is the production delay caveat?
A: Even absent leakage, the paper cautions that production reductions from either policy may represent production delays rather than permanent reductions — oil not extracted today may be extracted later as prices rise or technology improves. To the extent that reductions are temporary, the effective carbon abatement is smaller than the authors compute, and the cost per tonne of CO2 abated is correspondingly higher. The paper does not quantify this effect but flags it as a material caveat.&lt;/p&gt;
&lt;p&gt;Q: What institutional features drive pipeline capacity investment and risk allocation?
A: Pipelines are irreversible investments subject to ex-post holdup, so construction financing requires firm ship-or-pay commitments from shippers before construction and before future prices are known, meaning oil price risk is borne primarily by shippers rather than the pipeline owner. Pipeline tariffs are regulated by FERC on a cost-of-service basis. In the DAPL case, shippers executed binding ten-year ship-or-pay contracts in June 2014, and shippers&amp;rsquo; beliefs about future oil prices at that date — calibrated to historical price volatility using an AR(1) process with estimated persistence φ₁ = 0.9925 and volatility σ_G = 0.098 — determine equilibrium capacity investment.&lt;/p&gt;
&lt;p&gt;Q: How does the paper&amp;rsquo;s finding relate to the typical co-benefit structure of climate policies?
A: Most CO2 abatement policies generate local pollution co-benefits (reduced NOx, SOx, particulates), so the abatement cost is partially offset by local pollution gains. Blocking DAPL reverses this: the pipeline-to-rail modal shift increases local pollution damages, making local pollution a cost rather than a co-benefit of the policy. The authors note this is atypical but not unprecedented — urban densification and post-combustion emissions controls in fossil fuel boilers also present CO2–local pollution trade-offs.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Infrastructure foreclosure policy: A &amp;ldquo;keep it in the ground&amp;rdquo; strategy that blocks construction of specialized fossil fuel transportation infrastructure (pipelines) with the aim of inhibiting production of the fuels that would have been transported, without requiring direct acquisition or buyout of mineral rights.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Ship-or-pay agreement: A firm, up-front capacity commitment in which a pipeline shipper agrees to pay for reserved pipeline capacity whether or not they ultimately use it, made before construction and before future prices are realized; the institutional mechanism by which oil price risk is transferred from pipeline owners to shippers.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Crude-by-rail adjustment costs: Quadratic costs modeled as linear in the period-to-period change in rail volumes to a given destination, capturing capital investments and disinvestments in rail cars, loading facilities, and unloading terminals needed to expand or contract crude-by-rail capacity; estimated at $1.28/bbl per mbbl/d of monthly flow change.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Production leakage: The partial or full offset of production reductions in one oil basin (Bakken) by production increases in other U.S. or global basins in response to the same price signals; at 57% leakage, the cost of blocking DAPL rises from $45/tonne to $104/tonne of CO2 abated.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Old-well vs. new-well production dynamics: The distinction between production from pre-existing wells (which follows an exponential decline path insensitive to current prices, β = 0.955) and production from newly drilled wells (which responds to current and lagged upstream prices with long-run elasticity 1.32); this structure makes total short-run supply highly inelastic while allowing substantial long-run price responsiveness through drilling adjustments.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Local pollution damages from NOx: The dominant component of environmental harm from crude-by-rail transport, arising from locomotive NOx emissions that are both large in magnitude and concentrated in densely populated areas along rail corridors; at $100/tonne SCC, monetized local pollution damages from rail exceed CO2 damages for all three coastal destinations, whereas for pipelines CO2 damages exceed local pollution costs.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Cost per tonne of CO2 abated: The authors&amp;rsquo; metric for comparing infrastructure foreclosure to alternative policies; computed as the sum of lost producer surplus and net change in local pollution damages divided by the quantity of CO2 emissions avoided from reduced oil production and consumption; equals $45/tonne for blocking DAPL versus $1.01–$2.68/tonne for an equivalent upstream production tax.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>Firm Accommodation After Workplace Disability: Labor Market Impacts and Implications for Subsidy Design</title><link>https://macropaperwarehouse.com/papers/firm-accommodation-after-workplace-disability-labor-market-impacts-and-implications-for-subsidy-design/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/firm-accommodation-after-workplace-disability-labor-market-impacts-and-implications-for-subsidy-design/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper studies (1) how firm accommodation decisions respond to financial incentives in the context of workplace disability under workers&amp;rsquo; compensation, (2) what the causal effect of accommodation is on workers&amp;rsquo; subsequent labor market outcomes, and (3) whether the equilibrium level of accommodation is socially efficient, and what the welfare implications of wage subsidies for accommodation are.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Context and Data&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The analysis uses the universe of Oregon workers&amp;rsquo; compensation claims from 2005 through 2017 — over 131,000 disabling claims — linked to longitudinal quarterly earnings records from the Oregon Employment Department. The setting exploits Oregon&amp;rsquo;s Employer at Injury Program (EAIP), which subsidizes employers who provide &amp;ldquo;transitional work&amp;rdquo; accommodations (primarily through wage subsidies) to workers with temporary workplace disabilities. EAIP accounts for roughly 25 percent of claims on average, with the wage subsidy component representing over 96 percent of EAIP expenses.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identification Strategy&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors exploit a policy change in July 2013 that reduced the EAIP wage subsidy rate from 50 percent to 45 percent. They construct a firm-level &amp;ldquo;exposure&amp;rdquo; measure — the fraction of a firm&amp;rsquo;s claims that used EAIP in a baseline period (2005–2009) — and estimate a continuous difference-in-differences specification in which the interaction of exposure and a post-2013 indicator instruments for accommodation. The identifying assumption is strong parallel trends: firms with low baseline exposure are unlikely to respond to the subsidy reduction, while high-exposure firms respond more, generating cross-firm variation in accommodation rates after 2013. An MTE framework (Heckman and Vytlacil 2005) is then used to explore heterogeneous treatment effects along an unobserved resistance-to-treatment dimension.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Empirical Findings&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The subsidy reduction from 50% to 45% decreased accommodation rates by &lt;strong&gt;2.9 percentage points&lt;/strong&gt; (9.3 percent) for claims in firms with average exposure, implying a subsidy elasticity of accommodation of 0.9.&lt;/li&gt;
&lt;li&gt;The policy change led to a &lt;strong&gt;0.95 percentage point decrease in employment&lt;/strong&gt; and a &lt;strong&gt;$120 decrease in quarterly earnings&lt;/strong&gt; four quarters after disability for claims in average-exposure firms (roughly 1.3–1.5 percent declines relative to means), with no significant effect on worker turnover to other firms.&lt;/li&gt;
&lt;li&gt;IV estimates of the effect of accommodation itself (using predicted EAIP as instrument) show &lt;strong&gt;accommodation increases the probability of employment four quarters after disability by 33 percentage points&lt;/strong&gt; and &lt;strong&gt;increases quarterly earnings by approximately $4,100&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The MTE analysis reveals &lt;strong&gt;negative selection on gains&lt;/strong&gt;: workers with workplace disabilities who are least likely to receive accommodation have the highest potential gains from it, driven largely by severe disabilities with high accommodation costs.&lt;/li&gt;
&lt;li&gt;Descriptive and IV evidence is consistent with accommodation operating primarily as &lt;strong&gt;general human capital investment&lt;/strong&gt;: accommodation has no statistically significant effect on the probability of moving to a new firm, and earnings gains are not systematically lower for workers who change employers after accommodation.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Structural Model and Counterfactual Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A two-period frictional labor market model with risk-averse workers, risk-neutral firms, Nash bargaining, imperfect experience rating in workers&amp;rsquo; compensation, and firm accommodation as human capital investment is developed and estimated. Two inefficiency sources are identified: (1) a human capital externality — because accommodation builds general human capital, firms cannot capture the full surplus when workers separate, reducing accommodation incentives; and (2) a fiscal externality — imperfectly experience-rated firms do not fully internalize the workers&amp;rsquo; compensation cost savings from accommodation, further depressing it below the efficient level. Counterfactual simulations show:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Eliminating wage subsidies (from 50% to 0%) reduces accommodation rates from &lt;strong&gt;33% to 11%&lt;/strong&gt;, leading to a &lt;strong&gt;7% decline in post-disability employment&lt;/strong&gt; and a &lt;strong&gt;15% decline in post-disability quarterly wages&lt;/strong&gt; (roughly $1,358).&lt;/li&gt;
&lt;li&gt;A revenue-neutral reform eliminating wage subsidies reduces average welfare and the welfare of &lt;strong&gt;more than 90% of workers&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Welfare gains from the subsidy are &lt;strong&gt;larger for low-skilled workers&lt;/strong&gt; than high-skilled workers.&lt;/li&gt;
&lt;li&gt;Conditional on experiencing disability, eliminating wage subsidies decreases welfare by about &lt;strong&gt;10%&lt;/strong&gt;, while increasing the subsidy to 100% raises welfare for disabled workers by around &lt;strong&gt;30%&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Firm profit is maximized at a subsidy rate around 80%, after which higher taxes offset accommodation gains.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-employer-at-injury-program-eaip-and-how-does-it-differ-from-standard-workers-compensation"&gt;Q1. What is the Employer at Injury Program (EAIP), and how does it differ from standard workers&amp;rsquo; compensation?&lt;/h3&gt;
&lt;p&gt;A1: EAIP is an optional component of Oregon&amp;rsquo;s workers&amp;rsquo; compensation system that subsidizes employers for the costs of accommodating workers with temporary disabilities during a transitional return-to-work period. Unlike standard workers&amp;rsquo; compensation premiums (which are experience-rated at the firm level), EAIP is funded through a flat payroll tax on all firms that is not experience-rated — meaning firms that use EAIP do not pay higher premiums. The wage subsidy component accounts for over 96 percent of EAIP expenses; other reimbursable costs (worksite modifications up to $5,000, retraining up to $1,000, clothing up to $400) are rarely used. Eligible employers must be the employer at which the disability occurred, and accommodation is limited to a transitional period during which workers cannot simultaneously receive time-loss benefits.&lt;/p&gt;
&lt;h3 id="q2-how-is-firm-level-exposure-constructed-and-what-is-the-rationale-for-using-it-as-an-instrument"&gt;Q2. How is firm-level &amp;ldquo;exposure&amp;rdquo; constructed, and what is the rationale for using it as an instrument?&lt;/h3&gt;
&lt;p&gt;A2: Exposure is the fraction of a firm&amp;rsquo;s workers&amp;rsquo; compensation claims that used EAIP during a five-year baseline period from 2005 to 2009 — a separate historical period chosen to reduce volatility and avoid mean-reversion. The rationale draws on prior work (Aizawa et al., 2022) showing that firm fixed effects account for nearly 25 percent of variation in accommodation, far more than worker or disability characteristics (1 and 3 percent, respectively), suggesting permanent firm-level heterogeneity in the relative benefits and costs of accommodation. Firms with zero historical exposure are unlikely to change accommodation behavior in response to a subsidy reduction, while high-exposure firms respond more, creating differential quasi-experimental variation in accommodation rates after July 2013.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-first-stage-and-reduced-form-results-from-the-did-specification"&gt;Q3. What are the first-stage and reduced-form results from the DID specification?&lt;/h3&gt;
&lt;p&gt;A3: The first-stage DID coefficient shows that a ten-percentage-point increase in exposure is associated with a one-percentage-point decrease in EAIP take-up after 2013, implying a 2.9 percentage point decrease for claims in firms with average exposure (mean 0.27). The corresponding reduced-form results show a 0.35 percentage point decrease in employment four quarters post-disability and a $45 decrease in quarterly earnings for every ten-percentage-point increase in exposure, scaling to 0.95 percentage points and $120 at average exposure. There is no statistically significant effect on the probability of moving to a new firm. Pre-trend tests show parallel accommodation trends across exposure terciles prior to 2013, supporting the identifying assumption.&lt;/p&gt;
&lt;h3 id="q4-what-do-the-iv-estimates-imply-about-the-causal-effect-of-accommodation-on-labor-market-outcomes"&gt;Q4. What do the IV estimates imply about the causal effect of accommodation on labor market outcomes?&lt;/h3&gt;
&lt;p&gt;A4: Under the exclusion restriction that the subsidy change affects labor market outcomes only through accommodation, the IV estimates imply that receipt of accommodation increases the probability of employment four quarters after disability by &lt;strong&gt;33 percentage points&lt;/strong&gt; (against a mean of 72 percent) and increases quarterly earnings by approximately &lt;strong&gt;$4,100&lt;/strong&gt; (against a mean of $7,807). There is no significant effect on the probability of working at a new firm four quarters later. The authors note these large estimates reflect local average treatment effects for compliers — workers whose accommodation status was changed by the instrument — who disproportionately have high unobserved resistance to treatment and high accommodation returns, explaining the magnitude.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-mte-framework-reveal-about-the-distribution-of-accommodation-effects-and-selection"&gt;Q5. What does the MTE framework reveal about the distribution of accommodation effects and selection?&lt;/h3&gt;
&lt;p&gt;A5: The MTE curves show that workers with the highest unobserved resistance to treatment (least likely to receive accommodation) have the highest potential employment and earnings gains from accommodation. This negative selection on gains arises because these workers tend to have worse employment outcomes in the untreated state, consistent with more severe disabilities commanding higher accommodation costs. IV weights are concentrated at high-resistance values, explaining the large IV estimates. Negative selection on gains is also found along observable dimensions: workers in self-insured firms, healthcare support occupations, women, and those with wounds/cuts/burns show larger gains but lower likelihood of receiving accommodation.&lt;/p&gt;
&lt;h3 id="q6-what-evidence-supports-characterizing-firm-accommodation-as-general-rather-than-firm-specific-human-capital-investment"&gt;Q6. What evidence supports characterizing firm accommodation as general rather than firm-specific human capital investment?&lt;/h3&gt;
&lt;p&gt;A6: Three pieces of evidence point toward general human capital. First, the IV estimate shows accommodation has no statistically significant effect on the probability of working at a new firm four quarters after disability. Second, a triple-interaction specification (DID interacted with new-firm indicator) yields suggestive evidence of even larger earnings gains for workers who move to a new firm post-accommodation, though this is not statistically significant — a pattern inconsistent with firm-specific human capital. Third, the subset of claims that receive non-wage EAIP benefits (worksite modifications, retraining) do show lower mobility, but this comprises fewer than 5 percent of the sample, meaning the predominant form of investment in the context is general in nature.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-two-sources-of-market-inefficiency-in-accommodation-identified-in-the-model"&gt;Q7. What are the two sources of market inefficiency in accommodation identified in the model?&lt;/h3&gt;
&lt;p&gt;A7: The first is a human capital externality operating through worker turnover. Because accommodation builds general human capital that workers carry to new employers, a firm accommodating a worker does not capture the portion of future surplus that accrues to future employers upon separation. In a Nash bargaining framework with lack of commitment, this dynamic inefficiency is larger when industry-wide turnover rates are higher — consistent with the descriptive finding that accommodation rates are strongly negatively associated with industry separation rates. The second is a fiscal externality from imperfect experience rating: firms whose workers&amp;rsquo; compensation premiums are not fully linked to their own claim costs do not fully internalize the cost-savings from accommodation (i.e., reduced time-loss benefit payments), leading them to accommodate at inefficiently low rates.&lt;/p&gt;
&lt;h3 id="q8-how-is-heterogeneity-incorporated-in-the-structural-estimation-and-what-do-the-estimated-parameters-show"&gt;Q8. How is heterogeneity incorporated in the structural estimation, and what do the estimated parameters show?&lt;/h3&gt;
&lt;p&gt;A8: The model incorporates observed heterogeneity (firm insurance status, worker skill type — measured by pre-disability wages — firm baseline exposure, and pre/post policy change) and unobserved heterogeneity mapped to the MTE framework&amp;rsquo;s unobserved resistance to treatment. Indirect inference matches cross-sectional accommodation rates, earnings by subgroup, and the DID coefficients. Key findings: net output during the disability period is negative (accommodation is a costly short-run investment), while post-disability output is higher for accommodated workers. Low-skilled workers experience larger productivity gains from accommodation than high-skilled workers. Accommodation cost shock variance is lower for higher unobserved types, meaning high-gain workers are also more sensitive to subsidy changes, consistent with the large IV estimates. The model fits the DID coefficients for accommodation, employment, and wages well.&lt;/p&gt;
&lt;h3 id="q9-what-do-the-counterfactual-simulations-show-about-the-welfare-effects-of-varying-the-subsidy-rate"&gt;Q9. What do the counterfactual simulations show about the welfare effects of varying the subsidy rate?&lt;/h3&gt;
&lt;p&gt;A9: Eliminating wage subsidies from the current 50% rate reduces the accommodation rate from 33% to 11% and lowers post-disability employment by 7 percentage points and post-disability quarterly wages by 15% ($1,358). From a welfare perspective, eliminating subsidies in a revenue-neutral reform reduces average ex-ante worker welfare and lowers welfare for more than 90% of workers. Conditional on experiencing disability, eliminating subsidies reduces welfare by about 10% while raising the subsidy to 100% increases welfare of disabled workers by around 30%. Firm profit is increasing in the subsidy rate up to about 80%, then decreases. Ex-ante worker welfare gains from the current 50% subsidy relative to no subsidy are modest in consumption-equivalent terms (at most 0.6% increase in consumption), partly because the disability probability is low (2.2%) and because unaccommodated workers still receive two-thirds wage replacement through time-loss benefits.&lt;/p&gt;
&lt;h3 id="q10-what-distributional-implications-do-wage-subsidies-have-across-worker-and-firm-types"&gt;Q10. What distributional implications do wage subsidies have across worker and firm types?&lt;/h3&gt;
&lt;p&gt;A10: Welfare gains from higher wage subsidies are larger for low-skilled workers than high-skilled workers, so the subsidy has a redistributive dimension beyond efficiency correction. Welfare gains are also larger for workers in imperfectly experience-rated firms, where the fiscal externality creates the greater wedge from the efficient level. Self-insured firms, which already internalize workers&amp;rsquo; compensation cost savings and thus accommodate closer to the optimal rate, benefit less from the subsidy and can even be made worse off if subsidies are set very high (since they bear higher flat payroll taxes with smaller marginal accommodation gains). The fraction of worker-firm matches experiencing welfare gains exceeds 90% under the benchmark subsidy level, indicating broad rather than narrowly concentrated gains.&lt;/p&gt;
&lt;h3 id="q11-how-do-the-experience-rating-channel-and-the-worker-turnover-channel-interact-in-comparative-statics"&gt;Q11. How do the experience-rating channel and the worker-turnover channel interact in comparative statics?&lt;/h3&gt;
&lt;p&gt;A11: Model comparative statics show that reducing the job-to-job transition rate of workers with disabilities to one-quarter of its estimated value substantially raises accommodation rates, and this effect is more pronounced for imperfectly experience-rated firms than for self-insured firms. This occurs because self-insured firms already have a strong incentive to accommodate (to reduce workers&amp;rsquo; compensation premiums), so turnover is less marginal for them. Forcing all firms to be self-insured (perfect experience rating) would substantially increase accommodation rates in currently imperfectly rated firms. Lowering the accommodation cost during the disability period (increasing net output during the disability period) also raises accommodation rates for both firm types.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Firm Accommodation (EAIP):&lt;/strong&gt; In this paper&amp;rsquo;s specific sense, accommodation refers to a firm&amp;rsquo;s decision to offer a worker with a temporary workplace disability &amp;ldquo;transitional work&amp;rdquo; — alternative tasks, modified duties, or flexible arrangements — during their recovery period, funded in part through Oregon&amp;rsquo;s Employer at Injury Program wage subsidy. Accommodation is distinct from simple early return to work; it functions as a form of human capital investment by potentially providing skill development opportunities and preventing human capital depreciation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exposure (Instrument):&lt;/strong&gt; A firm-level continuous measure defined as the fraction of a firm&amp;rsquo;s workers&amp;rsquo; compensation claims that used EAIP during a five-year baseline period (2005–2009). Exposure captures permanent, time-invariant firm-level propensity to accommodate, and is used to construct a difference-in-differences instrument for the causal effect of accommodation by interacting exposure with a post-2013 indicator (when the subsidy rate was cut from 50% to 45%).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Imperfect Experience Rating:&lt;/strong&gt; The degree to which a firm&amp;rsquo;s workers&amp;rsquo; compensation insurance premium adjusts to reflect that firm&amp;rsquo;s own claims costs, rather than being set at an industry average. Fully experience-rated (self-insured) firms internalize 100% of claim costs and thus have strong incentives to accommodate. Partially experience-rated firms face a fiscal externality: because their premiums do not fully reflect their own time-loss benefit expenditures, they do not capture all the cost savings from accommodating workers, leading to under-accommodation relative to the social optimum.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Human Capital Externality (Dynamic Inefficiency in Accommodation):&lt;/strong&gt; The mechanism — analogous to Acemoglu and Pischke (1999) and Fang and Gavazza (2011) — by which worker turnover reduces firms&amp;rsquo; incentives to invest in general human capital (here, accommodation). When accommodation raises workers&amp;rsquo; general productivity, part of the future surplus from this investment accrues to future employers upon job-to-job separation. With Nash bargaining and lack of commitment (re-bargaining in the second period), the accommodating firm cannot capture this surplus, creating a dynamic inefficiency that is more severe in high-turnover industries.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negative Selection on Gains:&lt;/strong&gt; The empirical finding, established via the MTE framework, that workers with workplace disabilities who are least likely to receive accommodation (highest unobserved resistance to treatment) have the largest potential employment and earnings gains from accommodation. This pattern arises because workers with more severe disabilities have high accommodation costs (making firms unwilling to accommodate them) but also face far worse counterfactual labor market outcomes without accommodation, creating large potential gains.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marginal Treatment Effect (MTE):&lt;/strong&gt; Following Heckman and Vytlacil (2005), the treatment effect of accommodation evaluated at a specific quantile of unobserved resistance to treatment — defined here as the propensity score value at which a worker is indifferent between treatment and non-treatment. The MTE curve maps out the full distribution of treatment effects and reveals who benefits (and by how much), how IV estimates are weighted averages over this distribution, and which compliers drive the large IV estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;General vs. Firm-Specific Human Capital (in Accommodation Context):&lt;/strong&gt; Accommodation is characterized as general human capital investment if the productivity and earnings gains it produces are transferable across employers — i.e., if accommodated workers who move to new firms retain their wage gains. It is firm-specific if gains are tied to the current match. In this paper, general human capital is supported by the null effect of accommodation on new-firm employment probability, suggestive evidence of non-lower (possibly larger) earnings gains for new-firm movers, and the observation that fewer than 5% of claims use non-wage EAIP benefits associated with firm-specific investment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Revenue-Neutral Counterfactual:&lt;/strong&gt; A counterfactual policy experiment in which the wage subsidy rate for accommodation is varied while imposing that both the time-loss benefit program and the EAIP wage subsidy program remain budget-balanced. Higher subsidy rates raise firm accommodation, reduce time-loss benefit payouts (lowering base premiums for imperfectly experience-rated firms), but require a higher flat EAIP payroll tax on all firms, some of which is passed through to workers via lower first-period wages.&lt;/p&gt;</description></item><item><title>Genetic Prediction and Adverse Selection</title><link>https://macropaperwarehouse.com/papers/genetic-prediction-and-adverse-selection/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/genetic-prediction-and-adverse-selection/</guid><description>&lt;p&gt;This paper asks how much adverse selection would arise in critical illness insurance (CII) markets if consumers can observe polygenic indexes (PGIs) — genetic risk scores derived from millions of genetic variants — while insurers are legally barred from using genetic information. The authors develop an econometric method that measures selection under current PGI technology, then extends identification to expected future PGI accuracy using heritability bounds, even though future PGIs are not yet observable in data.&lt;/p&gt;
&lt;p&gt;The primary dataset is the UK Biobank (UKB), comprising approximately 446,570 genotyped individuals of European-like ancestry linked to NHS electronic health records. The authors study seven single-disease CII contracts (Alzheimer&amp;rsquo;s disease, breast cancer, coronary artery disease, colorectal cancer, prostate cancer, schizophrenia, and type 2 diabetes) and multiple-disease bundled contracts paying a lump sum upon onset. The econometric model assumes a probit disease probability, Gaussian PGI structure, and identification relies on published heritability estimates to pin down future PGI predictive power. The key selection metric is the implicit tax proposed by Hendren (2013): the percentage markup a marginal consumer must pay above her actuarially fair price due to adverse selection. The authors use the minimum implicit tax up to the 80th percentile of risk (t80) as their summary statistic, with market unraveling benchmarked at t80 between 43% and 83% from prior literature.&lt;/p&gt;
&lt;p&gt;The paper reports three main findings, all scoped to a population of 35-year-olds in the standard insurer risk class (those whose predicted risk falls within 0.75–1.25 times the population mean).&lt;/p&gt;
&lt;p&gt;First, under current PGI technology with full consumer adoption, selection is noticeable but heterogeneous across diseases. t80 ranges from 17.9% for coronary artery disease to 117.9% for Alzheimer&amp;rsquo;s disease. Coronary artery disease and colorectal cancer fall in the middle of the no-unraveling range; breast cancer, schizophrenia, and type 2 diabetes fall between the no-unraveling and unraveling ranges; Alzheimer&amp;rsquo;s disease and prostate cancer (t80 = 59.8%) reach or exceed the unraveling range. The current prostate cancer PGI explains 9.9% of liability variance, adding 8.3 percentage points over the 22.9% explained by non-genetic covariates.&lt;/p&gt;
&lt;p&gt;Second, under expected future PGI accuracy — bounded below by SNP heritability and above by twin heritability — selection becomes potentially crippling. Under the lower bound (Scenario 3L), t80 ranges from 57.5% for breast cancer to above 1,000% for Alzheimer&amp;rsquo;s. Under the upper bound (Scenario 3U), t80 exceeds 100% for all seven single-disease contracts and exceeds 1,000% for three of them. For prostate cancer, the reference case, t80 reaches 86.8% under Scenario 3L and 426.9% under Scenario 3U — far above Hendren&amp;rsquo;s unraveling benchmarks. For multiple-disease male contracts, t80 = 30.8% under current technology, rising to 54.4% (Scenario 3L) and 243.9% (Scenario 3U).&lt;/p&gt;
&lt;p&gt;Third, variation in selection across contracts is driven primarily by: the predictive power of the future PGI, the incremental predictive power over non-genetic covariates, and disease prevalence. Alzheimer&amp;rsquo;s and schizophrenia — high heritability, low prevalence — display the highest implicit taxes; breast and colorectal cancer — lower SNP heritability, lower incremental R2 — display the lowest.&lt;/p&gt;
&lt;p&gt;These findings are corroborated by a calibrated Akerlof-Einav-Finkelstein equilibrium model using HRS data: current PGI availability reduces equilibrium market quantity from 30% to 21.4%; future PGI availability drives equilibrium quantity to zero in a full adverse selection death spiral. Partial take-up robustness checks show that even at 50% consumer adoption, selection remains problematically high under future PGI accuracy for most contracts. The analysis is restricted to individuals of European-like ancestry due to data availability constraints.&lt;/p&gt;
&lt;p&gt;Q: What is the core market failure the paper analyzes?
A: The paper analyzes adverse selection arising from an asymmetric information gap: consumers can observe PGI-based disease risk predictions from consumer genetic tests (e.g., 23andMe), while insurers in many jurisdictions are legally prohibited from requesting or using genetic information. This creates a situation where high-risk consumers have private information allowing them to sort into insurance, driving up average claims costs and potentially unraveling the market.&lt;/p&gt;
&lt;p&gt;Q: What is a polygenic index (PGI) and why does it differ from classical genetic testing?
A: A PGI is a weighted sum of millions of genetic variants (typically over one million) each with individually tiny effects, constructed using effect-size estimates from genome-wide association studies (GWASs). This contrasts with traditional genetic testing focused on rare single-gene mutations (e.g., BRCA for breast cancer or PKD for kidney disease), which are rare, explain small shares of population-level disease variance, and can largely be inferred from family history. PGIs target common polygenic diseases and are the primary driver of the adverse selection concern because they aggregate diffuse genetic signals into a meaningful risk prediction.&lt;/p&gt;
&lt;p&gt;Q: What are the current PGI R2 values for the seven diseases studied?
A: Estimated on the liability scale in the UKB, current PGI R2 values are: Alzheimer&amp;rsquo;s disease 7.1%, breast cancer 6.7%, coronary artery disease 2.5%, colorectal cancer 2.2%, prostate cancer 9.9%, schizophrenia 4.9%, and type 2 diabetes 7.4%. These represent the share of liability variance explained by each disease&amp;rsquo;s current PGI in the study sample.&lt;/p&gt;
&lt;p&gt;Q: How does the paper identify the degree of selection under future PGI technology that does not yet exist in the data?
A: The identification strategy combines three elements: the normality of PGI distributions, the relationship between current and future PGIs (the current PGI is modeled as a noisy version of the future PGI with an independent Gaussian error), and published heritability estimates that bound the future PGI&amp;rsquo;s predictive power. Theorem 1 establishes that under five stated assumptions — including a probit disease model and known future R2 from heritability studies — the full joint distribution of loss, current PGI, future PGI, and non-genetic covariates is identified from observed data.&lt;/p&gt;
&lt;p&gt;Q: What heritability bounds are used for the future PGI scenarios, and why two bounds?
A: Scenario 3L sets future PGI R2 equal to each disease&amp;rsquo;s SNP heritability (estimated from common genetic variants), which the authors treat as a conservative lower bound because future PGIs will also incorporate rarer variants with better effect-size precision. Scenario 3U sets future PGI R2 equal to twin heritability, treating it as an upper bound since the theoretical maximum predictive power of a PGI is the trait&amp;rsquo;s narrow-sense heritability. For prostate cancer, these bounds are 18.0% (SNP) and 57.0% (twin); for Alzheimer&amp;rsquo;s, SNP heritability is 33.1% and twin heritability is 58%.&lt;/p&gt;
&lt;p&gt;Q: What is the implicit tax and how is it used as a benchmark?
A: The implicit tax t(r) for a consumer with private risk r equals the percentage by which her insurance cost exceeds her own actuarially fair price when she must pool with all consumers of equal or higher risk. It measures how much the marginal buyer overpays due to adverse selection. The authors follow Hendren (2013) in reporting t80, the minimum implicit tax up to the 80th percentile. Hendren&amp;rsquo;s benchmarks: t80 between 7–35% for markets that did not unravel; t80 between 43–83% for markets that had unraveled.&lt;/p&gt;
&lt;p&gt;Q: What are the single-disease contract results under current PGI technology (Scenario 2)?
A: With full consumer adoption of current PGI technology, t80 ranges from 17.9% for coronary artery disease to 117.9% for Alzheimer&amp;rsquo;s disease. Coronary artery disease (17.9%) and colorectal cancer (26.5%) fall in the middle of Hendren&amp;rsquo;s no-unraveling range. Breast cancer (36.9%), schizophrenia (42.1%), and type 2 diabetes (37.0%) fall between the no-unraveling and unraveling ranges. Alzheimer&amp;rsquo;s disease (117.9%) and prostate cancer (59.8%) reach or exceed the unraveling range.&lt;/p&gt;
&lt;p&gt;Q: What are the single-disease contract results under future PGI technology?
A: Under the lower bound (Scenario 3L, R2 = SNP heritability), t80 ranges from 57.5% for breast cancer to above 1,000% for Alzheimer&amp;rsquo;s disease. Under the upper bound (Scenario 3U, R2 = twin heritability), t80 exceeds 100% for all seven contracts and exceeds 1,000% for three (Alzheimer&amp;rsquo;s, schizophrenia, and at least one other). These figures substantially exceed Hendren&amp;rsquo;s unraveled-market benchmarks for virtually all contracts.&lt;/p&gt;
&lt;p&gt;Q: What drives cross-disease variation in the implicit tax?
A: The authors identify three main drivers: the expected accuracy of future PGI (higher heritability → higher implicit tax), the incremental predictive power of the future PGI over non-genetic covariates observable by insurers (more incremental information → more adverse selection), and disease prevalence (lower prevalence concentrates risk heterogeneity, amplifying selection). Alzheimer&amp;rsquo;s disease and schizophrenia — high heritability and low prevalence — have the highest implicit taxes. Breast and colorectal cancers — lower SNP heritability and lower incremental R2 — have the lowest.&lt;/p&gt;
&lt;p&gt;Q: What do the multiple-disease bundled contract results show?
A: For the male multiple-disease contract under Scenario 2 (current PGI), t80 = 30.8%, comparable to Hendren&amp;rsquo;s no-unraveling range. Under Scenario 3L, t80 = 54.4%; under Scenario 3U, t80 = 243.9%, both in or above the unraveling range. The female contract yields qualitatively similar results. Implicit taxes in bundled contracts are generally lower than in single-disease contracts, suggesting some diversification of genetic risk across diseases.&lt;/p&gt;
&lt;p&gt;Q: What does the calibrated equilibrium model find?
A: Using an Akerlof (1970) / Einav-Finkelstein-Cullen (2010) supply-and-demand model calibrated to match a 30% market participation rate and a 50% loss ratio in the UK CII market, and using HRS data on individual risk aversion, the model finds that current PGI availability reduces equilibrium quantity from 30% to 21.4%. Future PGI availability (both Scenario 3L and 3U) drives equilibrium quantity to zero — a complete adverse selection death spiral with no trade.&lt;/p&gt;
&lt;p&gt;Q: How robust are results to partial consumer adoption of genetic testing?
A: At 10% consumer take-up, selection is low regardless of PGI accuracy. At 50% take-up, selection remains problematically high for all single-disease contracts under future PGI accuracy (Scenarios 3L and 3U). For multiple-disease contracts at 50% take-up, t80 falls just below Hendren&amp;rsquo;s unraveling threshold under Scenario 3L but enters the unraveling range under Scenario 3U. This suggests market problems would materialize once predictive power exceeds the SNP heritability bound and take-up exceeds roughly 50%.&lt;/p&gt;
&lt;p&gt;Q: What role do risk preferences play, and do they confound the results?
A: The authors test whether risk tolerance correlates with disease risk in the UKB using a self-reported general risk tolerance measure. They find extremely low correlations between risk tolerance and each disease. This is consistent with low correlation between relative risk aversion and disease risk in the HRS calibration, and supports the finding that correlation between risk and risk preferences is unlikely to meaningfully affect the main results.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s assessment of preventive treatment as a mitigating factor?
A: The authors acknowledge that genetic testing could enable personalized preventive medicine, which would reduce actual disease incidence among high-risk individuals. However, they argue this is unlikely to substantially affect their main findings because the most commonly covered diseases under CII are cancers, for which preventive behaviors have bounded effectiveness.&lt;/p&gt;
&lt;p&gt;Q: What are the paper&amp;rsquo;s policy implications?
A: The paper situates the genetic information problem within the standard regulatory framework for selection markets, distinguishing laissez-faire (allow genetic underwriting — efficient but potentially unfair to high-risk consumers), government provision (unattractive for non-essential CII), and managed competition (community rating combined with subsidies and risk adjustment). The authors argue that a full ban on genetic underwriting — the current policy in many countries — may become untenable as PGI accuracy improves, because it generates potentially crippling adverse selection. Some level of community rating may remain desirable for redistribution, but needs to be paired with subsidies or risk adjustment to prevent market collapse.&lt;/p&gt;
&lt;p&gt;Q: What are the main data and scope limitations?
A: The analysis is restricted to individuals of European-like ancestry because most large GWASs were conducted in European ancestry samples and PGIs perform poorly across ancestries. The UKB sample was aged 40–69 at recruitment and the analysis adjusts for age-dependent covariates; the HRS replication uses approximately 20,000 individuals. The equilibrium model ignores moral hazard and uses a parsimonious binary loss framework. The paper does not specify a timeline for when PGI accuracy will reach heritability bounds.&lt;/p&gt;
&lt;p&gt;Polygenic Index (PGI): A weighted sum of an individual&amp;rsquo;s genetic variants across the genome (typically over one million variants), constructed using effect-size estimates from a genome-wide association study (GWAS) conducted in an independent sample. It is a noisy proxy for the individual&amp;rsquo;s true additive genetic factor for a disease, and its predictive power is bounded above by the trait&amp;rsquo;s narrow-sense heritability.&lt;/p&gt;
&lt;p&gt;Implicit Tax: A measure of adverse selection defined by Hendren (2013) as the percentage by which a consumer with private risk r must overpay relative to her own actuarially fair price if she is pooled with all consumers of equal or higher risk. The minimum implicit tax up to the 80th percentile of risk (t80) serves as the paper&amp;rsquo;s primary summary statistic; t80 above roughly 43% is associated with market unraveling in prior literature.&lt;/p&gt;
&lt;p&gt;SNP Heritability: The share of variance in a disease&amp;rsquo;s liability attributable to the set of common genetic variants (SNPs) used in heritability estimation. Used in this paper as a conservative lower bound on the predictive power of future PGIs, because future PGIs will additionally capture rarer variants.&lt;/p&gt;
&lt;p&gt;Twin Heritability: An estimate of a trait&amp;rsquo;s narrow-sense (additive) heritability computed by comparing resemblance of monozygotic twins (sharing 100% of their genomes) to dizygotic twins (sharing ~50% on average). Used as an upper bound on future PGI predictive power, since heritability is the theoretical maximum R2 for a PGI.&lt;/p&gt;
&lt;p&gt;Standard Risk Class: The set of consumers whose predicted disease risk (based on non-genetic covariates observable to insurers) falls between 0.75 and 1.25 times the population-wide average risk, following standard insurance underwriting practice. Insurers charge the same premium to all consumers in this class; any variation in risk within the class due to private genetic information constitutes the source of adverse selection analyzed in this paper.&lt;/p&gt;
&lt;p&gt;Private Risk Function: The probability rho(g, w) of contracting the disease conditional on both the consumer&amp;rsquo;s observed PGI g and non-genetic factors w. Contrasted with the non-genetic private risk function pi(w), which conditions only on non-genetic covariates. The dispersion of the private risk distribution across consumers in the same risk class determines the degree of adverse selection.&lt;/p&gt;
&lt;p&gt;Adverse Selection Death Spiral: The Akerlof (1970) mechanism in which high-risk consumers disproportionately purchase insurance, causing insurers to raise premiums, which deters low-risk consumers, which further raises the average risk of purchasers, ultimately driving equilibrium quantity to zero. The paper&amp;rsquo;s calibrated equilibrium model finds this outcome under future PGI accuracy for the HRS CAD contract.&lt;/p&gt;</description></item><item><title>Germs in the Family: The Short- and Long-Term Consequences of Intra-Household Disease Spread</title><link>https://macropaperwarehouse.com/papers/germs-in-the-family-the-short-and-long-term-consequences-of-intra-household-disease-spread/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/germs-in-the-family-the-short-and-long-term-consequences-of-intra-household-disease-spread/</guid><description>&lt;p&gt;This paper studies the short- and long-term consequences of intra-household respiratory disease transmission from older to younger siblings in Danish families. The central research questions are: (1) how do respiratory illnesses spread from preschool-aged older siblings to younger infant siblings during the first year of life, and (2) how does respiratory disease exposure during infancy causally affect younger siblings&amp;rsquo; long-term economic, human capital, and health outcomes?&lt;/p&gt;
&lt;p&gt;The study uses population-level Danish administrative data covering 1,230,180 children from 37 birth cohorts (1981–2017), linking records from the National Patient Register, income and labor market registers, education registers, and psychiatric care registers. The identification strategy combines birth order variation in respiratory disease vulnerability with within-municipality variation in local respiratory disease prevalence among children aged 13–71 months. The authors construct a municipality-level disease exposure index—cumulative respiratory hospitalizations per 100 children aged 13–71 months in a child&amp;rsquo;s municipality over their first 12 months of life—and estimate the differential effect of this index on younger versus older siblings, controlling for municipality fixed effects, birth year-month fixed effects, and an extensive set of individual and family background characteristics.&lt;/p&gt;
&lt;p&gt;The descriptive findings are stark: younger siblings have 2–3 times higher rates of hospitalization for acute respiratory conditions during their first year of life compared to older siblings at the same age, with the gap largest at ages two and three months. The gap is larger for winter births, shorter birth spacing, and when older siblings attend childcare centers—all patterns consistent with the older sibling serving as a disease vector.&lt;/p&gt;
&lt;p&gt;On the causal estimates, moving from the 25th to the 75th percentile of the disease exposure index distribution increases the younger sibling&amp;rsquo;s acute respiratory hospitalizations in the first year of life by 0.023 (32.9 percent above the sample mean), with effects more than twice as large for exposure in the first six months compared to the second six months.&lt;/p&gt;
&lt;p&gt;In the long run, an interquartile increase in first-year respiratory disease exposure reduces younger siblings&amp;rsquo; wage earnings (conditional on employment) at ages 25–32 by 0.8 percent and total income by 0.8 percent, and reduces their income percentile rank by 0.3 percentage points. There is no significant effect on labor force participation at the extensive margin. Effects on earnings are approximately twice as large when exposure is measured in the first six months of life. These earnings effects are comparable in magnitude to those from a 10 percent reduction in birth weight or a 9 percent increase in ambient air pollution at birth, and correspond to roughly two-thirds of the adult earnings impact of in utero exposure to the 1918 Spanish Influenza. When the disease index interaction is included, the main birth order coefficient declines by approximately 70 percent, suggesting intra-household disease transmission is an important channel underlying the documented birth order earnings disadvantage.&lt;/p&gt;
&lt;p&gt;Additional findings include: a 0.5 percentage point reduction in high school graduation and a 0.6 percentage point reduction in college graduation (interquartile effects); a 0.01 standard deviation penalty in ninth grade Danish test scores; a 20 percent increase (0.016 per hundred per year) in chronic respiratory hospitalizations at ages 16–26; and a 6.1 percent increase (0.5 additional visits per hundred per year) in psychiatric clinic visits at ages 16–26. Breastfeeding mitigates short-term effects, with 15 months of breastfeeding sufficient to entirely offset the elevated hospitalization risk.&lt;/p&gt;
&lt;p&gt;Scope conditions: findings apply to second-born relative to first-born children in Danish sibling pairs with at least 11 months birth spacing; long-term estimates are net of parental compensatory responses and any immunity benefits, and thus represent lower bounds of the uncompensated biological impact of respiratory illness in infancy.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the birth order gap in acute respiratory hospitalizations during infancy, and what patterns support an intra-household transmission mechanism?
A: Younger siblings have 2–3 times higher hospitalization rates for acute respiratory conditions in the first year of life compared to older siblings at the same age, with the gap especially large at ages two and three months. The gap is larger for winter births (when respiratory viruses circulate more), for siblings with shorter birth spacing, and when the older sibling attends a childcare center. Hospitalizations for non-infectious digestive diseases and injuries show no analogous birth order differences, ruling out differential parental healthcare-seeking as an explanation.&lt;/p&gt;
&lt;p&gt;Q: How is the disease exposure index constructed and what variation does it exploit?
A: The index is the cumulative count of acute respiratory hospitalizations per 100 children aged 13–71 months in a child&amp;rsquo;s municipality over their first 12 months of life, with the older sibling excluded from the count when applicable. It exploits irregular spatial and temporal waves of respiratory viruses (such as RSV and influenza) across Danish municipalities. The interquartile range of this index captures meaningful variation in community disease burden faced by infants across different places and years.&lt;/p&gt;
&lt;p&gt;Q: What is the first-stage relationship between the disease index and infant hospitalizations?
A: Moving from the 25th to the 75th percentile of the disease index increases younger siblings&amp;rsquo; acute respiratory hospitalizations in the first year of life by 0.023 (a 32.9 percent increase relative to the sample mean), while the effect on older siblings is substantially smaller. The interaction coefficient in the preferred specification implies that one additional hospitalization per 100 community children aged 13–71 months raises the younger sibling&amp;rsquo;s hospitalization count by 0.012 more than the older sibling&amp;rsquo;s. Effects are more than twice as large for exposure in the first compared to the second six months of life.&lt;/p&gt;
&lt;p&gt;Q: What are the estimated long-term effects on adult earnings, and how do they compare to benchmarks in the literature?
A: An interquartile increase in first-year respiratory disease exposure reduces younger siblings&amp;rsquo; wage earnings at ages 25–32 by 0.8 percent and total income by 0.8 percent, with a 0.3 percentage point reduction in income percentile rank. These magnitudes are comparable to a 1 percent earnings reduction from a 10 percent birth weight reduction (Black et al., 2007), a 1 percent earnings reduction from a 9 percent increase in ambient air pollution (Isen et al., 2017b), and roughly two-thirds of the in utero Spanish Influenza effect (Almond, 2006).&lt;/p&gt;
&lt;p&gt;Q: Does the birth order earnings disadvantage reflect intra-household disease transmission?
A: When the interaction between birth order and the disease index is excluded, the regression finds a 1.9 percent birth order earnings disadvantage for second-born children (consistent with Black et al., 2005 range of 1.2–4.2 percent). When the interaction is included, the main birth order coefficient declines by approximately 70 percent, suggesting that disease transmission from older to younger siblings is an important channel driving the birth order earnings penalty.&lt;/p&gt;
&lt;p&gt;Q: Are effects larger for exposure in the first versus second six months of life?
A: Yes, consistently across all outcomes. The interaction coefficient for acute respiratory hospitalizations is more than twice as large when exposure is measured in the first versus second six months. Effects on wage earnings are approximately 60 percent larger for first-half exposure, and effects on income rank are two to three times larger. This is consistent with biomedical evidence that infants&amp;rsquo; immune systems mature around six months when solid food introduction begins.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on educational outcomes?
A: An interquartile increase in first-year respiratory disease exposure reduces the likelihood of high school graduation by 0.5 percentage points (0.6 percent at the sample mean) and college graduation by 0.6 percentage points (1.7 percent at the sample mean), with effects approximately 60 percent larger when measuring first-half exposure. A 0.01 standard deviation reduction in ninth grade Danish test scores is also found. A back-of-the-envelope calculation using Danish returns to schooling suggests the reduction in educational attainment can explain approximately half of the estimated earnings effect.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on chronic respiratory and mental health outcomes?
A: An interquartile increase in first-year exposure increases chronic respiratory hospitalizations (asthma, COPD) at ages 16–26 by 0.016 per hundred per year (20 percent above the sample mean), with significant increases also apparent at ages one to two. For mental health, the same exposure is associated with 0.5 additional psychiatric clinic visits per hundred per year at ages 16–26 (6.1 percent above the sample mean), with effects becoming more significant in the early twenties. Effects on mental health from this paper are smaller than those estimated for more extreme fetal and early childhood shocks such as Ramadan exposure or maternal bereavement.&lt;/p&gt;
&lt;p&gt;Q: What does the acute respiratory trajectory look like beyond infancy?
A: Elevated acute respiratory hospitalizations persist at age one, then there is a reduction at ages two to three consistent with an immunity formation hypothesis, but this protective effect disappears by age four. There is no significant increase or decrease in acute respiratory hospitalizations at older ages, in contrast to the persistent increase found for chronic respiratory conditions.&lt;/p&gt;
&lt;p&gt;Q: What heterogeneity is found in short-term effects?
A: Effects on infant respiratory hospitalizations are larger for low birth weight children, for male infants (consistent with the fragile male hypothesis), for siblings with shorter birth spacing, and for sibling pairs where the older child attends childcare. The monotonic decline in effect size with increasing birth spacing is the opposite of what would be predicted if differential parental time investment were the main mechanism, supporting intra-household disease spread as the operative channel.&lt;/p&gt;
&lt;p&gt;Q: What is the role of breastfeeding as a moderator?
A: Using supplementary data on breastfeeding duration (covering 2009–2016, matched to 7.6 percent of the sample), the authors find that the impact of disease exposure on younger siblings&amp;rsquo; infancy hospitalizations declines significantly with longer breastfeeding duration. A linear specification implies that 15 months of breastfeeding entirely offsets the elevated hospitalization risk from higher disease exposure. Second-born children breastfed for less than half a month are particularly vulnerable to acute respiratory infections.&lt;/p&gt;
&lt;p&gt;Q: How do the authors validate the identifying assumption?
A: Three validation exercises are used. First, results are robust to adding municipality-specific linear and quadratic trends and maternal fixed effects. Second, using family background characteristics as outcomes in the interaction regression, at most two of fourteen coefficients are significant in any specification, and all effect sizes are less than one percent of sample means. Third, using alternative disease indices based on non-infectious digestive diseases and injuries shows no differential effects for younger siblings, ruling out a parental healthcare-seeking confound.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications?
A: The authors highlight breastfeeding support policies (paid family leave, workplace lactation accommodations), RSV vaccination campaigns for pregnant women and monoclonal antibody prophylaxis for infants, sick pay regulations, and childcare attendance policies as levers to reduce infant respiratory disease burden. They argue that current cost-benefit evaluations of such policies likely undercount the long-term human capital and earnings benefits. The COVID-19 pandemic illustrates the mechanism: restrictions reduced RSV spread during 2020 potentially benefiting infants with older siblings, while the subsequent RSV surge in 2021–2022 may have exposed later cohorts to above-average disease burden.&lt;/p&gt;
&lt;p&gt;Respiratory Disease Exposure Index: A municipality-level cumulative measure of acute respiratory hospitalizations per 100 children aged 13–71 months assigned to each child over their first 12 months of life (or first and second six months separately), designed to proxy for community respiratory disease burden faced by infants from slightly older children, with the child&amp;rsquo;s own older sibling excluded from the count.&lt;/p&gt;
&lt;p&gt;Intra-Household Disease Transmission: The mechanism by which preschool-aged older siblings, exposed to respiratory viruses in group childcare settings, bring home those viruses and infect younger infant siblings who are in a vulnerable stage of immune and brain development, creating a within-family externality in health outcomes.&lt;/p&gt;
&lt;p&gt;Differential Birth Order Effect (Identification): The quasi-experimental design exploits the interaction between birth order (younger siblings are more exposed to older siblings&amp;rsquo; illnesses) and local disease prevalence variation to identify causal impacts, netting out the main effects of both birth order and local disease environment through municipality and birth year-month fixed effects.&lt;/p&gt;
&lt;p&gt;Immunity Formation Hypothesis: The conjecture that early respiratory disease exposure may have a protective effect on later acute respiratory illness through immune system training; supported in the data by reduced acute hospitalizations at ages two to three, though this protection disappears by age four and does not prevent chronic respiratory disease development.&lt;/p&gt;
&lt;p&gt;Dynamic Complementarities with Sibling Health Spillovers: An extension of the Cunha-Heckman framework: while standard models incorporate investment complementarities across time periods for a given child, this paper&amp;rsquo;s findings imply that sibling health spillovers create differential returns to early-life health investments by birth order, since disease asymmetries between older and younger siblings are not incorporated in existing theoretical models.&lt;/p&gt;
&lt;p&gt;Net Long-Term Effects: The estimated long-run impacts incorporate not only the direct biological effects of respiratory illness on the younger sibling but also any parental compensatory responses and immunity benefits; thus they represent lower bounds of the uncompensated biological impact, as parental compensation would attenuate the measured sibling difference.&lt;/p&gt;</description></item><item><title>Growth Experiences and Trust in Government</title><link>https://macropaperwarehouse.com/papers/growth-experiences-and-trust-in-government/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/growth-experiences-and-trust-in-government/</guid><description>&lt;p&gt;This paper investigates whether individuals who have experienced stronger GDP growth over their lifetimes are more likely to trust their national government. The authors — Besley, Dann, and Dray — assemble a newly harmonized global dataset comprising approximately 3.3 million respondents across 166 countries since 1990, drawn from 11 major opinion surveys (Afrobarometer, Americasbarometer, Arabarometer, Asiabarometer, European Social Survey, Gallup World Poll, Integrated Values Survey, Latinobarometer, Life in Transition Survey, South Asia Barometer, and World Justice Project). They supplement this with longer-run U.S. evidence from the American National Election Studies (ANES) going back to 1958, covering respondents born as early as the 1880s, and longitudinal Swiss evidence from the Swiss Household Panel (SHP) which allows individual fixed-effects estimation.&lt;/p&gt;
&lt;p&gt;The core methodological contribution is the exploitation of country-cohort variation in lifetime GDP growth experiences. Following Malmendier and Nagel (2011), the authors construct a weighted average of past growth realizations across an individual&amp;rsquo;s lifetime, with weights decaying linearly over time (lambda = 1), so that more recent growth receives greater weight. The baseline specification includes country fixed effects, cohort-by-subcontinent fixed effects, survey-by-survey-year fixed effects, controls for log GDP per capita at year of birth, and individual characteristics (sex, marital status, education, religious denomination). More demanding specifications add country-by-survey-year and country-by-age fixed effects. For Switzerland, individual fixed effects are included, fully absorbing time-invariant personal characteristics.&lt;/p&gt;
&lt;p&gt;The main finding is that a one standard deviation increase in lifetime GDP growth experience — corresponding to approximately 2 percentage points of additional growth — is associated with a 2.1 percentage point increase in the probability of trusting the national government, significant at the 1 percent level. This corresponds to roughly 0.042 standard deviations of the trust outcome and approximately 5 percent of the global mean trust in government. The effect is quantitatively meaningful: it approximates between one-quarter and one-half of the difference in average trust between older and younger cohorts in India and Italy, respectively. For the U.S. ANES sample, a one standard deviation increase in growth experience (about 0.2 percentage points) increases trust in the federal government by 2.4 percentage points, explaining more than two-thirds of the average trust gap between Baby Boomers (born 1946–1964) and Millennials (born 1981–1996).&lt;/p&gt;
&lt;p&gt;Several scope conditions and heterogeneity findings sharpen the interpretation. First, the growth-trust link is specific to government institutions: there is no statistically significant effect of growth experience on interpersonal trust or trust in religious organizations, indicating the channel runs through perceptions of state performance rather than generalized social capital. Second, a recency heuristic operates: the linearly decaying weighting function (lambda = 1) outperforms both an unweighted lifetime average (lambda = 0) and a formative-years weighting. Growth experienced during formative years (ages 18–25) or before birth has no detectable effect on trust in government; the pre-birth result serves as a placebo test. Third, the positive growth-trust relationship is stronger in democracies than in autocracies, which the authors interpret as democracies producing citizens more responsive to government performance signals. Fourth, a &amp;ldquo;trust paradox&amp;rdquo; emerges: unconditionally, average trust in government is lower in democracies than in autocracies, and longer democratic experience is associated with lower trust, which the authors attribute to democratic institutions generating greater citizen skepticism about government performance. Fifth, core results are robust to controlling for other lifetime politico-economic experiences including inflation, banking and currency crises, epidemics, political unrest, executive turnover, stock market returns, and income inequality. The Swiss evidence further shows that private income growth experience does not drive the result — only aggregate macroeconomic growth does.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s core quantitative finding on the growth-trust relationship?
A: Using the global harmonized dataset of 3.3 million respondents across 166 countries, a one standard deviation increase in lifetime GDP growth experience (corresponding to approximately 2 percentage points of additional growth) is associated with a 2.1 percentage point increase in the probability of trusting the national government, significant at the 1 percent level. Using only the Gallup World Poll subsample (roughly half the observations), the estimated effect is somewhat larger at 3.6 percentage points per standard deviation increase. These estimates remain statistically significant under more demanding specifications with country-by-survey-year and country-by-age fixed effects, though the magnitudes decrease as these interacted fixed effects absorb variation in recent growth experiences.&lt;/p&gt;
&lt;p&gt;Q: How do the authors measure individual lifetime growth experience?
A: The growth experience variable is a weighted average of all past annual GDP per capita growth rates since an individual&amp;rsquo;s birth, with weights that decay linearly over time (lambda = 1 in the Malmendier-Nagel framework). Under this parameterization, the measure simplifies to how much recent economic performance (in the year prior to the survey) exceeds the long-run mean over the respondent&amp;rsquo;s lifetime, scaled by the respondent&amp;rsquo;s midpoint of life. This implies younger individuals are more sensitive to recent growth outcomes because their shorter life histories give recent events relatively greater weight. The authors validate this lambda = 1 choice via a grid search over alternative weighting structures using minimum residual sum of squares as the criterion.&lt;/p&gt;
&lt;p&gt;Q: How is reverse causality addressed?
A: The empirical strategy identifies the relationship using past, cumulative growth experiences measured prior to the survey, so current trust in government cannot cause past growth. Survey-year fixed effects absorb all aggregate time trends simultaneously affecting trust and growth. The authors also conduct a placebo test showing that GDP growth occurring before an individual&amp;rsquo;s birth has a precisely estimated null effect on their trust in government, which would not be the case if unobserved societal trends were jointly driving both growth histories and political perceptions.&lt;/p&gt;
&lt;p&gt;Q: Does growth experience affect interpersonal trust or trust in non-state institutions?
A: No. The estimated coefficient on lifetime growth experience is statistically insignificant at conventional levels when interpersonal trust replaces trust in government as the dependent variable, with narrow confidence intervals indicating a precisely estimated null. Similarly, growth experience has no systematic effect on trust in religious organizations such as churches or mosques. The authors interpret these null results as evidence against the alternative explanation that broad modernizing social changes are jointly driving both growth experiences and political trust.&lt;/p&gt;
&lt;p&gt;Q: What do the U.S. ANES results add?
A: The ANES data, which extends back to 1958 and captures cohorts born as early as the 1880s, provide a within-country test controlling for state fixed effects, generation dummies, and rich individual characteristics including partisan affiliation and partisan strength. A one standard deviation increase in U.S. growth experience (approximately 0.2 percentage points) raises trust in the federal government by 2.4 percentage points, significant at the 1 percent level. This estimate is quantitatively large enough to explain more than two-thirds of the average trust gap between Baby Boomers and Millennials. Results are robust to adding state-by-survey-year fixed effects and birth-state-by-generation fixed effects, and hold for a broader &amp;ldquo;trust in government index&amp;rdquo; covering beliefs about waste, corruption, and responsiveness of the federal government.&lt;/p&gt;
&lt;p&gt;Q: What do the Swiss Household Panel results contribute?
A: The SHP allows individual fixed-effects estimation, exploiting within-person changes in growth experience and trust over time from 1999 onward, which absorbs all time-invariant individual characteristics that could confound the global and U.S. cross-cohort results. The growth experience coefficient remains positive and significant, with a one standard deviation increase yielding a 1.9 percentage point increase in trust in the Swiss federal government (significant at the 1 percent level). The Swiss data also uniquely allow the authors to test whether personal income growth experience drives the result; they find no significant effect of private income growth experience on trust in government, only aggregate macroeconomic growth matters.&lt;/p&gt;
&lt;p&gt;Q: Does the recency heuristic hold — does growth in formative years matter?
A: No. The authors find no detectable effect of growth experienced specifically during formative years (ages 18–25) on trust in government. Additionally, in a grid-search exercise assessing model fit across different lambda values, the linearly decaying weighting scheme (lambda = 1, giving more weight to recent growth) outperforms both equal-weighted lifetime averages (lambda = 0) and weighting schemes that emphasize earlier life experiences (lambda less than 0). The pre-birth placebo result (null effect) and the absence of a formative-years effect together indicate that the operative mechanism is about evaluating current government performance based on recent macroeconomic experience, not the imprinting of long-lasting political dispositions during youth.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;trust paradox&amp;rdquo; and how is it documented?
A: The trust paradox refers to the empirical finding that average trust in government is lower in democracies than in autocracies at the cross-country level, and that longer experience with democratic institutions within countries is associated with lower levels of trust in government in the micro data. This is counterintuitive given the standard view that good institutions should foster confidence in government. The authors suggest the paradox likely reflects democracies cultivating greater citizen skepticism and more critical judgment of government performance, rather than indicating that democratic governance actually performs worse. Importantly, the positive effect of growth experience on trust remains present in democracies, and the growth-trust relationship is actually stronger in democratic regimes, consistent with citizens in democracies being more responsive to government performance signals.&lt;/p&gt;
&lt;p&gt;Q: How is the growth-trust finding related to corruption perceptions and living standards?
A: Using the Gallup World Poll, the authors find that stronger lifetime growth experience is associated with lower perceived corruption in government, greater satisfaction with personal living standards, and higher likelihood of feeling one lives comfortably on one&amp;rsquo;s present income. These results are consistent with citizens attributing economic success to government competence and integrity, and with growth translating into perceptions of improved personal circumstances through both direct income effects and indirect public goods provision.&lt;/p&gt;
&lt;p&gt;Q: Are the results robust to controlling for other lifetime politico-economic experiences?
A: Yes. When the authors include lifetime experience measures for political unrest, executive turnover, epidemic exposure, banking crises, currency crises, and inflation (both levels and volatility) simultaneously in equation (3), the growth experience coefficient remains consistently positive, stable, and significant across all specifications. Among the other experience variables, only lifetime unrest and epidemic exposure are independently negative and statistically significant at conventional levels. F-tests reject the null hypothesis that the crisis and growth experience coefficients are equal in magnitude. The U.S. results are also robust to adding lifetime experiences with S&amp;amp;P 500 returns, unemployment, and top-income-share inequality measures.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: The authors note that sustained economic growth may itself be a mechanism for building political trust, with positive downstream effects for policy compliance — a connection they document has been relevant during the COVID-19 pandemic (where higher-trust societies showed lower mobility during lockdowns and higher vaccine acceptance). The growth-trust channel could have implications for increasing compliance across a range of policy domains including climate action and tax morale. Governments that deliver sustained economic growth can expect citizens to update their trust upward, particularly in democracies where citizens are more performance-responsive, while governments that preside over stagnation or contraction face predictable erosion of political legitimacy across cohorts.&lt;/p&gt;
&lt;p&gt;Growth experience: A weighted average of all past annual GDP per capita growth realizations since an individual&amp;rsquo;s birth, with weights that decay linearly over time following Malmendier and Nagel (2011), so that more recent growth receives greater weight. Under the paper&amp;rsquo;s preferred parameterization (lambda = 1), the measure equals how much last year&amp;rsquo;s GDP per capita exceeds the respondent&amp;rsquo;s lifetime mean, scaled by the respondent&amp;rsquo;s midpoint of life.&lt;/p&gt;
&lt;p&gt;Trust in government: A binary dummy variable equal to one if a survey respondent expresses &amp;ldquo;a great deal&amp;rdquo; or &amp;ldquo;quite a lot&amp;rdquo; of trust or confidence in the national government, constructed from harmonized responses across 11 major opinion surveys. The paper treats this as reflecting respondents&amp;rsquo; perceptions of government performance rather than a deep interpersonal trust relationship.&lt;/p&gt;
&lt;p&gt;Trust paradox: The empirical regularity documented in the paper whereby average trust in government is unconditionally lower in democracies than in autocracies at the cross-country level, and whereby longer democratic experience within countries is associated with lower individual trust in government. The authors attribute this to democratic institutions generating more critical citizen judgment of government performance.&lt;/p&gt;
&lt;p&gt;Recency heuristic: The finding that more recent growth experiences carry greater weight in forming trust in government, as captured by the linear decay weighting scheme (lambda = 1) outperforming equal-weighted or early-life-weighted alternatives. Growth before birth and growth during formative years (ages 18–25) have no detectable effect, while recent macroeconomic performance is the operative signal.&lt;/p&gt;
&lt;p&gt;Cohort-level variation: The within-country differences in lifetime growth experiences across birth cohorts that form the paper&amp;rsquo;s primary identification strategy. Because different cohorts in the same country have lived through different sequences of growth episodes, differences in trust across cohorts within a country can be attributed to differential growth exposure rather than time-invariant country characteristics.&lt;/p&gt;
&lt;p&gt;Formative years effect: The hypothesis, tested and rejected in the paper, that economic experiences during ages 18–25 have a lasting imprint on political attitudes analogous to formative-years effects found in other political behavior literatures. The paper finds no statistically significant association between growth experienced during these years and trust in government.&lt;/p&gt;
&lt;p&gt;Source text origin: In the pipeline context relevant to this paper&amp;rsquo;s acquisition, this refers to whether a summary was generated from full working paper text (&amp;ldquo;pdf&amp;rdquo; or &amp;ldquo;oa-html&amp;rdquo;) versus abstract only (which is hard-blocked). The working paper was obtained from LSE Research Online (eprint 129614), classified as published version under CC BY 4.0.&lt;/p&gt;</description></item><item><title>Health Shocks, Health Insurance, Human Capital, and the Dynamics of Earnings and Health</title><link>https://macropaperwarehouse.com/papers/health-shocks-health-insurance-human-capital-and-the-dynamics-of-earnings-and-health/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/health-shocks-health-insurance-human-capital-and-the-dynamics-of-earnings-and-health/</guid><description>&lt;p&gt;Capatina and Keane build and calibrate a life-cycle model of labor supply and savings for U.S. men that incorporates health shocks, endogenous human capital accumulation via learning-by-doing, employer-sponsored health insurance (ESHI), means-tested social insurance, and endogenous medical treatment decisions. The model is calibrated to White males using the Medical Expenditure Panel Survey (MEPS) for 2000–2013, supplemented by CPS, HRS, and PSID data; separate calibrations are presented for Black and Hispanic men with high school or less education.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central research question is how health shocks affect labor supply, earnings, and earnings inequality over the life cycle, and through which mechanisms. Four channels are identified and quantified: (1) the direct labor supply effect — sick days and reduced tastes for work caused by health shocks; (2) the human capital effect — reduced work experience from health-shock-induced employment exits, which deteriorates future job and wage offers in a snowball dynamic; (3) the health-productivity effect — reduced functional health directly lowering wage offers; and (4) the behavioral effect — anticipation of health risk induces low-skill workers lacking ESHI to curtail labor supply to maintain means-tested transfer eligibility.&lt;/p&gt;
&lt;p&gt;The key quantitative findings from eliminating serious health shocks for working-age men (ages 25–64) are: the expected present value of lifetime earnings (PVE) for White men rises by 11% on average, and inequality in PVE falls by 12% (coefficient of variation). For White men with high school or less education the increase in PVE is 17.9%. For the typical White male the four channels contribute 5.7%, 2.7%, 1.4%, and 0.8% respectively. For low-skill White high school men the same channels contribute 10.7%, 14.8%, 1.3%, and 9.8% — with the human capital and behavioral effects dramatically larger for the low-skill group. For comparison, a severe health shock at age 40 reduces the present value of remaining lifetime earnings by 5.6% (approximately $53.9k) for a typical college man and by 11.5% (approximately $55.0k) for a typical high school man.&lt;/p&gt;
&lt;p&gt;Human capital amplification operates through employment persistence: a major health shock causes full-time employment to drop by 12 percentage points one year after the shock for the average man, and by 20 percentage points for high school men, with recovery still incomplete eight years later (employment remains 7.8 pp and 10 pp below baseline, respectively). Holding human capital fixed as in the pre-shock baseline causes employment to recover quickly, confirming that persistent wage-offer deterioration is the mechanism.&lt;/p&gt;
&lt;p&gt;On health insurance policy, the model evaluates providing public insurance to all workers lacking ESHI. This substantially increases medical utilization, improves health and life expectancy (survival to age 65 rises from 82% to 87% when health shocks are eliminated, as a related benchmark), reduces Medicaid and free-care costs, and raises labor supply among low-skill workers by weakening means-tested transfer incentives. The net program cost in a balanced budget simulation is modest, and all agent types are ex ante better off. By contrast, expanding Medicaid access creates perverse labor supply disincentives — workers reduce labor supply to maintain eligibility — does little to improve health, and makes almost all agents worse off in a balanced budget scenario.&lt;/p&gt;
&lt;p&gt;Scope conditions: the primary calibration covers non-institutionalized civilian White males; results for Blacks and Hispanics are presented only for the high school or less education group due to small samples. The model period ends at 2013, before ACA implementation.&lt;/p&gt;
&lt;p&gt;Q: What is the model&amp;rsquo;s overall estimate of how much health shocks reduce lifetime earnings for White men?
A: Eliminating serious health shocks at working ages (25–64) would increase the expected present value of lifetime earnings (PVE) for the average White male by 11% and reduce inequality in PVE by 12% as measured by the coefficient of variation. For White men with high school or less education the PVE gain is larger at 17.9%.&lt;/p&gt;
&lt;p&gt;Q: What are the four channels through which health shocks affect earnings, and how large is each for the average White male versus a low-skill high school male?
A: The four channels are (1) direct labor supply via sick days and reduced tastes for work, (2) human capital deterioration from lost work experience worsening future job/wage offers, (3) reduced health productivity lowering wage offers, and (4) behavioral responses to health risk reducing labor supply to preserve transfer eligibility. For the average White male the contributions to PVE are 5.7%, 2.7%, 1.4%, and 0.8%, respectively. For low-skill White high school men the same channels contribute 10.7%, 14.8%, 1.3%, and 9.8% — the human capital and behavioral effects are roughly five to twelve times larger for the low-skill group.&lt;/p&gt;
&lt;p&gt;Q: Why is the human capital effect so much larger for low-skill high school men than for college men?
A: Low-skill high school men are much more likely to exit full-time employment following a major health shock and are slow to return. Lifetime work years decline by 1.89 for the typical high school man versus only 0.84 for the typical college man following a major shock at age 40. Because job offer probabilities depend on lagged employment, absence from the labor market creates a snowball effect that persistently depresses offer quality; human capital accounts for 42% of the earnings decline for high school men versus 34% for college men.&lt;/p&gt;
&lt;p&gt;Q: How does the paper characterize the persistent employment effects of a major health shock?
A: For the average man, full-time employment drops by 12 percentage points one year after a severe shock and remains 7.8 pp below baseline after eight years. For high school men the initial drop is 20 pp, still 10 pp below baseline after eight years; for college men the figures are 7 pp and 3 pp. When human capital is held fixed at the pre-shock baseline — so wage and job offers do not deteriorate due to lost experience — employment recovers quickly for workers of all skill levels, confirming the human capital mechanism drives the persistence.&lt;/p&gt;
&lt;p&gt;Q: How does the behavioral effect operate for low-skill workers?
A: Workers without ESHI who face health risk have an incentive to maintain sufficiently low income and assets to qualify for means-tested social insurance, which provides a consumption floor approximating Medicaid, Food Stamps, SSDI, and SSI. This perverse incentive leads low-skill workers to curtail labor supply preemptively. When health risk is eliminated, this incentive disappears and labor supply rises, generating the behavioral effect of 9.8% of PVE for low-skill high school men versus only 0.8% for the average White male.&lt;/p&gt;
&lt;p&gt;Q: How does the paper correct for under-reporting of health shocks among the uninsured?
A: The measurement model assumes health shocks are correctly measured for the treated, but uninsured workers who do not seek treatment only record a shock with a shock-specific probability less than one. A key identifying assumption is that, conditional on health status, risk factors, age, and education, the true frequency of health shocks does not differ by insurance status per se — ruling out ex ante moral hazard. The measurement model parameters are calibrated to match observed frequencies of health shocks and high risk in MEPS for the uninsured.&lt;/p&gt;
&lt;p&gt;Q: What does the model estimate regarding the effect of a severe health shock on cumulative earnings relative to existing reduced-form evidence?
A: The model predicts an average cumulative (non-discounted) earnings loss of $42.8k over ten years following a severe shock for men aged 50, compared with Smith&amp;rsquo;s (2004) estimate of $37k from the HRS. The paper argues Smith&amp;rsquo;s estimate identifies effects on workers who actually experience shocks, who are a selected sample with low baseline earnings (as untreated shocks are more likely to be severe, and non-treaters tend to have low earnings). The model&amp;rsquo;s &amp;ldquo;average effect&amp;rdquo; — comparing a world where everyone experiences the shock to one where no one does — yields a substantially higher loss of $59.8k.&lt;/p&gt;
&lt;p&gt;Q: What are the key findings from the public insurance experiment (providing insurance to the uninsured)?
A: Providing public insurance to all workers lacking ESHI substantially increases medical utilization among the previously uninsured, who are intrinsically less healthy. This improves health and life expectancy, raising Social Security costs. However, it also generates positive labor supply incentives for low-skill workers (reducing their reliance on means-tested transfers), substantially reduces Medicaid and free-care costs, and increases tax revenue. On balance, the net program cost in a balanced budget simulation is modest, and all types of workers are ex ante better off.&lt;/p&gt;
&lt;p&gt;Q: Why does expanding Medicaid access produce perverse results in contrast to providing public insurance?
A: Medicaid is means-tested, so expanded access requires workers to maintain sufficiently low income and assets to remain eligible. This creates disincentives to work and save — workers reduce labor supply to preserve eligibility. The result is reduced earnings, lower tax revenue, little improvement in health (as access to care depends on maintaining low income), and almost all agents being worse off in a balanced budget scenario.&lt;/p&gt;
&lt;p&gt;Q: What role does insurance play beyond consumption smoothing in this model?
A: Beyond lowering out-of-pocket (OOP) costs and smoothing consumption, insurance grants access to care: in the US system, proof of insurance is often required before treatment, so uninsured workers may not have the option to treat at all. The model captures three distinct option sets for the uninsured — all options available, treatment not available, or default not available — each motivated by different real-world contexts. Non-treatment worsens health transition probabilities, so the access-granting role of insurance independently affects health trajectories beyond its cost-reducing role.&lt;/p&gt;
&lt;p&gt;Q: What explains the observed positive association between education, income, insurance, and health transitions in the data, and how does the model generate this without education entering the health production function directly?
A: The association between education and health is largely driven by the positive correlation between education and latent health types; controlling for latent health type in a descriptive logit largely eliminates the education coefficient. The association between insurance and health transitions is driven by the fact that the insured are more likely to receive treatment; controlling for treatment and true shocks eliminates the insurance coefficient. Education affects health indirectly through its effects on treatment decisions — via wages, job offers with ESHI, and consumption capacity — without appearing as a direct argument in the health production function.&lt;/p&gt;
&lt;p&gt;Q: How large are the effects of health shocks on key population health statistics according to the model?
A: Eliminating serious health shocks at working ages would increase the fraction of working-age men in good health from 60% to 75% and raise the probability of survival to age 65 from 82% to 87%. Average annual sick days of 16.42 would be eliminated, implying a 6% increase in work days for employed workers and an employment rate increase from 88% to 91%. Average annual medical costs would fall from $4,618 to $1,132.&lt;/p&gt;
&lt;p&gt;Q: How do the results for Black and Hispanic men compare to White men?
A: The results are qualitatively similar, but the magnitudes for Black men are somewhat larger. Eliminating health shocks would raise PVE for Whites, Blacks, and Hispanics with high school or less education by 17.9%, 23.7%, and 17.7%, respectively. Separate access-to-care probabilities are calibrated for each group, reflecting racial disparities in access that explain part of the observed differences in health outcomes and treatment rates.&lt;/p&gt;
&lt;p&gt;Q: What is the role of the consumption floor (means-tested social insurance) in shaping equilibrium outcomes for low-skill workers?
A: The consumption floor guarantees a minimum household consumption level approximating Medicaid, Food Stamps, SSDI, and SSI. It shields low-skill workers from the full cost of health shocks, reducing both the consumption-smoothing value of ESHI and precautionary saving incentives. However, it also creates a powerful disincentive for low-skill workers without ESHI to work, as earning above the eligibility threshold would eliminate benefits. This mechanism amplifies earnings inequality by generating perverse labor supply behavior concentrated among low-skill, uninsured workers.&lt;/p&gt;
&lt;p&gt;Functional Health (H): A discrete stock variable (Poor, Fair, or Good) measuring aspects of health that directly affect worker productivity and tastes for work; distinguished from asymptomatic health risk. Transitions depend on lagged health, latent health type, age, persistent health shocks, and whether shocks are treated.&lt;/p&gt;
&lt;p&gt;Asymptomatic Health Risk (R): A binary state (low or high) capturing risk factors such as obesity, high cholesterol, and hypertension that increase the probability of future health shocks but do not affect current productivity.&lt;/p&gt;
&lt;p&gt;Human Capital Effect: The channel by which health shocks reduce lifetime earnings not directly but indirectly — by causing employment exits that slow work experience accumulation, which in turn deteriorates future job offer probabilities and wage offers in a persistent, self-reinforcing (snowball) dynamic.&lt;/p&gt;
&lt;p&gt;Behavioral Effect: The reduction in labor supply — and associated earnings loss — that occurs because workers facing health risk and lacking ESHI have an incentive to keep income and assets low enough to maintain eligibility for means-tested social insurance, even absent any contemporaneous health shock.&lt;/p&gt;
&lt;p&gt;Tied Wage-Hours-Insurance Offer: The model&amp;rsquo;s labor market structure in which employment offers jointly specify a wage rate, hours (no offer, part-time, or full-time), and whether the offer includes ESHI; workers accept or reject the bundle rather than choosing hours and insurance independently.&lt;/p&gt;
&lt;p&gt;Source Text Origin: The paper&amp;rsquo;s own term distinguishing how the full text of a paper was obtained (PDF, OA-HTML, or abstract-only); used in the summarization pipeline. [Note: this concept is from the summarization pipeline metadata, not from the paper itself — omitting.]&lt;/p&gt;
&lt;p&gt;Treatment/Payment Options: The set of decisions available to a worker after a health shock occurs — whether to seek treatment and, if treated, whether to pay the out-of-pocket cost or default on bills. The available choice set differs by insurance status and context: the uninsured may face denial of access (option to treat unavailable) or required prepayment (default unavailable), or may have all options including free care.&lt;/p&gt;
&lt;p&gt;Latent Health Type: An unobserved permanent individual characteristic capturing innate biological resilience and pre-age-25 health investments; determines baseline transition probabilities for functional health conditional on shocks. Positively correlated with latent skill type within education groups.&lt;/p&gt;</description></item><item><title>Ideas Have Consequences: The Impact of Law and Economics on American Justice</title><link>https://macropaperwarehouse.com/papers/ideas-have-consequences-the-impact-of-law-and-economics-on-american-justice/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/ideas-have-consequences-the-impact-of-law-and-economics-on-american-justice/</guid><description>&lt;p&gt;This paper quantifies the effect of the Manne Economics Institute for Federal Judges — an intensive two-week economics training program run by the Law and Economics Center from 1976 to 1998 — on the decision-making of U.S. federal judges. The research question is whether exposure to a coherent set of economic ideas can directly shift the policy decisions of sitting policymakers, as distinct from effects operating through partisan affiliation or formal legal rules.&lt;/p&gt;
&lt;p&gt;The program trained nearly half of all federal judges over its two decades of operation. By 1990, forty percent of federal judges had attended; by the late 1990s, roughly half of circuit court cases had a Manne-trained judge on the panel. Instructors included Milton Friedman, Armen Alchian, Harold Demsetz, Martin Feldstein, Paul Samuelson, and Orley Ashenfelter, covering supply-and-demand theory, the Coase Theorem, externalities, property rights, and criminal deterrence following Becker (1968). The program was funded by pro-business foundations and had a recognized conservative-leaning orientation, though it invited both Republican- and Democrat-appointed judges and was popular across party lines.&lt;/p&gt;
&lt;p&gt;The identification strategy is a differences-in-differences design exploiting staggered attendance timing. Because the program was oversubscribed and admitted judges on a first-come-first-served basis — with applicants bumped to later cohorts when capacity was reached — the timing of attendance within the ever-attending population has a quasi-random component. The preferred control group consists exclusively of other ever-attending judges who had not yet attended, rather than never-attenders, because never-attenders differ systematically on observables and show a pre-existing positive trend in economics language use, likely from ambient diffusion through clerks, law schools, and organizations such as the Federalist Society. Judge fixed effects and circuit-by-year (or courthouse-by-year) fixed effects absorb time-invariant judge characteristics and court-level time trends. Elastic-net-selected covariates predicting attendance timing, fully interacted with year fixed effects, are added as robustness controls. Standard errors are clustered by judge.&lt;/p&gt;
&lt;p&gt;The data cover approximately 200,000 published circuit court opinions (1970–2005) from Bloomberg Law, a 5% random sample of circuit cases hand-coded for ideological direction from the Songer-Auburn database, machine-coded regulatory agency outcomes, a newly collected antitrust case dataset, and approximately 1.03 million district court criminal sentencing records (1992–2003) from TRAC.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. First, after attending the Manne program, judges increase their use of economics language in written opinions by approximately one-third of a standard deviation, measured via word-embedding similarity to an economics lexicon; this effect is statistically significant in the short-run event-study window but does not persist over the full career. Second, Manne attendance raises conservative voting in economics-related cases (labor and regulation) by approximately one-quarter of a standard deviation — corresponding to judges deciding in the conservative direction about 20 percent more often relative to the mean — with no significant effect on non-economics cases; the interaction effect is robust across specifications including never-attenders. Third, post-Manne judges vote more frequently against federal labor and environmental regulatory agencies, a result that is statistically significant and economically meaningful with no detectable pre-trends. Fourth, post-Manne judges impose longer and more frequent prison sentences, with no increase in sentencing harshness for drug crimes — consistent with Manne instructors having explicitly advocated drug legalization — and with the harshness gap between Manne and non-Manne judges widening after the 2005 Booker decision expanded judicial sentencing discretion. Fifth, there is some evidence of increased voting against antitrust enforcement, though this result is more sensitive to specification. Persuasion rates computed following DellaVigna and Gentzkow (2010) are slightly larger than those estimated for partisan media interventions such as Fox News and are closest to the effect of a 10-week Washington Post subscription on Democratic governor vote share. Neither the legalist model (judges follow statutes mechanically) nor the attitudinal model (judges follow party affiliation) can explain these within-judge, within-party shifts.&lt;/p&gt;
&lt;p&gt;Q: What is the central identification challenge and how do the authors address it?
A: The key threat is that judges who chose to attend the Manne program — or who attended at a particular time — may differ systematically from non-attenders in ways correlated with their decision trajectories. The authors address this in two steps. First, they restrict the control group to other ever-attending judges who had not yet attended, exploiting the first-come-first-served oversubscription rule that created quasi-random variation in timing among applicants. Second, they use judge fixed effects plus circuit-by-year fixed effects, and add elastic-net-selected biographical covariates (e.g., birth cohort indicators) interacted with year fixed effects as a robustness check. Republican affiliation — the most salient ideological predictor of attendance — is not a statistically significant predictor of attendance timing, supporting the exclusion restriction.&lt;/p&gt;
&lt;p&gt;Q: Why are never-attenders excluded from the preferred control group?
A: Never-attenders differ from attenders on observables including political party and show a positively trending use of economics language in their opinions even before any treatment, suggesting ambient diffusion of economics ideas through law clerks, law school curricula, and organizations such as the Federalist Society. Including never-attenders in the control group produces a near-zero coefficient on the language outcome, which the authors interpret as reflecting spillovers rather than a true null effect; the coefficient on conservative voting in the interaction specification, however, remains positive and significant even when never-attenders are included.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the effect on economics language use?
A: The within-judge effect of Manne attendance on the word-embedding similarity between judicial opinions and an economics lexicon is approximately one-third of a standard deviation, statistically significant in the short-run event-study window (covering six years before and after attendance). The effect shrinks and becomes non-significant when the full career of Manne judges is examined (rather than just the event-study window), consistent with broad diffusion of economics language across the judiciary over time rather than a persistent individual-level treatment effect.&lt;/p&gt;
&lt;p&gt;Q: How large is the effect on conservative voting, and is it concentrated in particular case types?
A: Post-Manne attendance raises conservative voting in economics-related cases (labor and regulation) by approximately one-quarter of a standard deviation, corresponding to judges deciding in the conservative direction about 20 percent more often relative to the mean liberal-conservative decision rate. There is no statistically significant effect on non-economics cases. The interaction coefficient — the differential effect on economics versus non-economics cases — is positive and significant across all specifications including the full sample with never-attenders, making this the most robust directional result in the paper.&lt;/p&gt;
&lt;p&gt;Q: What is the effect on regulatory agency voting?
A: Post-Manne judges vote more frequently against federal labor agencies (National Labor Relations Board, OSHA, Department of Labor, Federal Labor Relations Authority, Office of Worker&amp;rsquo;s Compensation Programs) and the Environmental Protection Agency. The event study shows a positive and significant increase that persists across the event-study window with no detectable pre-trends. This result is robust to both the baseline specification and the elastic-net-controls specification.&lt;/p&gt;
&lt;p&gt;Q: What is the effect on criminal sentencing, and what heterogeneity is found?
A: Post-Manne judges impose both more frequent prison sentences and longer sentences, consistent with Becker&amp;rsquo;s deterrence framework taught in the program&amp;rsquo;s criminal law curriculum. The sentencing effects are absent for drug crimes, consistent with Manne instructors — including Milton Friedman — having explicitly advocated against the drug war and for drug legalization. The gap in sentencing harshness between Manne and non-Manne judges widens after the 2005 United States v. Booker decision, which made the Federal Sentencing Guidelines advisory rather than mandatory; this is consistent with the program having shaped latent judicial preferences that are expressed more fully when formal constraints are relaxed.&lt;/p&gt;
&lt;p&gt;Q: How do the persuasion rates compare to benchmark media studies?
A: The persuasion rates computed following DellaVigna and Gentzkow (2010) are slightly larger than those estimated for partisan media interventions such as Fox News (DellaVigna and Kaplan, 2007) and are closest to the persuasion rates implied by a 10-week subscription to the Washington Post on Democratic governor vote share (Gerber et al. 2009). The comparison contextualizes the Manne program as a moderately high-intensity ideational intervention relative to documented cases of political persuasion.&lt;/p&gt;
&lt;p&gt;Q: What do the results imply for theories of judicial behavior?
A: The findings are inconsistent with both the legalist/formalist model — under which judges apply statutes and precedent without regard to extra-legal factors, predicting zero effect — and the attitudinal model — under which judges simply follow partisan preferences, also predicting zero effect since the program attended judges of both parties. The within-judge, within-party shifts point to a third channel: judicial worldviews and economic ideas, independent of formal law and partisan affiliation, shape high-stakes precedent-setting decisions.&lt;/p&gt;
&lt;p&gt;Q: Can the authors distinguish between a pedagogical (informational) and an ideological persuasion mechanism?
A: They cannot definitively distinguish between the two. Both mechanisms predict increased economics language, more conservative rulings in economics cases, deregulatory voting, and harsher non-drug sentences. The drug-crime heterogeneity is somewhat more consistent with a nuanced pedagogical channel, since Manne instructors explicitly discussed drug legalization, but this pattern is also consistent with complex ideological effects. Evidence on decision quality (citation rates, judicial promotion) is mixed and not robust, providing no clean test of the informational mechanism.&lt;/p&gt;
&lt;p&gt;Q: What does the antitrust evidence show?
A: Post-Manne judges tend to vote against antitrust claimants (i.e., in favor of less antitrust enforcement), but this result is more sensitive to specification than the regulatory agency and sentencing results and is not always statistically significant across specifications. The authors treat it as suggestive rather than conclusive.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to the literature on economics education and normative beliefs?
A: Prior work finds that economics students are less redistributive (Selten and Ockenfels 1998), view surge prices more favorably (Frey and Meier 2005), favor profit maximization (Rubinstein 2006), and that economics professors are less ideologically liberal than other social scientists (Jelveh et al. 2018). The present paper extends this literature by studying established professionals (judges) making high-stakes real-world decisions, and by documenting a direct policy impact rather than a change in survey responses or experimental choices.&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the dataset and the program coverage?
A: The circuit court dataset covers approximately 200,000 published opinions from 1970 through 2005. The district court sentencing dataset covers approximately 1.03 million cases from 1992 through 2003 (event study sample). The Manne program ran from 1976 to 1998, with roughly twenty judges per cohort; by 1990 forty percent of federal judges had attended, and by the late 1990s roughly half of circuit court cases had a Manne-trained panelist. Biographical information comes from the Federal Judicial Center; program attendance lists come from Butler (1999) supplemented by FOIA-obtained annual reports.&lt;/p&gt;
&lt;p&gt;Manne Economics Institute for Federal Judges: An intensive two-week economics training program for sitting U.S. federal judges, run by the Law and Economics Center from 1976 to 1998, covering supply-and-demand theory, the Coase Theorem, externalities, property rights, deterrence theory, and related topics; funded by pro-business foundations; admitted judges on a first-come-first-served basis and trained nearly half of all federal judges over its operation.&lt;/p&gt;
&lt;p&gt;Word-embedding economics language measure: A continuous measure of how closely a judicial opinion&amp;rsquo;s vocabulary aligns with a lexicon of law-and-economics phrases, constructed using word2vec embeddings (Mikolov et al. 2013) trained on the corpus of judicial opinions; measures the semantic proximity of opinion text to the Ellickson (2000) economics lexicon in embedding space, capturing implicit and contextual use of economics reasoning rather than raw phrase counts.&lt;/p&gt;
&lt;p&gt;Deterrence theory (Becker model): The framework, drawn from Becker (1968), taught in the Manne program&amp;rsquo;s criminal law curriculum, which holds that optimal crime deterrence requires setting the expected penalty — the economic cost of punishment times the probability of detection — high enough to outweigh the expected benefits of crime; treated in the paper as the theoretical basis for predicting harsher sentencing among post-Manne judges, and contrasted with retribution- or rehabilitation-based sentencing rationales that dominated before its diffusion.&lt;/p&gt;
&lt;p&gt;Conservative judicial decision (economics cases): In the paper&amp;rsquo;s usage, a ruling against the liberal/pro-plaintiff position in a case involving labor or regulation, as hand-coded by the Songer-Auburn database; includes ruling against a labor agency, rejecting a regulatory claimant, or voting against antitrust enforcement; the paper finds Manne attendance shifts judges in this direction in economics cases but not in non-economics cases.&lt;/p&gt;
&lt;p&gt;First-come-first-served oversubscription: The admission rule of the Manne program during its oversubscribed heyday (from the second cohort in 1977 through the late 1980s), under which applicants who did not secure a spot were bumped to the next year&amp;rsquo;s cohort; the authors argue this rule generates quasi-random variation in the timing of attendance among ever-attending judges, conditional on applying, providing the identifying variation for the differences-in-differences design.&lt;/p&gt;
&lt;p&gt;Persuasion rate: A summary statistic, following DellaVigna and Gentzkow (2010), measuring the fraction of the &amp;ldquo;persuadable&amp;rdquo; population that is convinced by a treatment; used in the paper to benchmark the Manne program&amp;rsquo;s effect size against documented media persuasion interventions such as Fox News and Washington Post subscriptions.&lt;/p&gt;</description></item><item><title>Identification of Time-Inconsistent Models: The Case of Insecticide-Treated Nets</title><link>https://macropaperwarehouse.com/papers/identification-of-time-inconsistent-models-the-case-of-insecticide-treated-nets/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/identification-of-time-inconsistent-models-the-case-of-insecticide-treated-nets/</guid><description>&lt;p&gt;This paper addresses two related problems: the formal identification of time-inconsistent preferences in dynamic discrete choice models with unobserved heterogeneous types, and the structural estimation of those preferences using data from a health intervention in rural Orissa, India. The identification challenge is fundamental — even the standard exponential discount factor delta is generically not identified in dynamic choice models (Rust 1994; Magnac and Thesmar 2002), and this non-identification extends a fortiori to the hyperbolic (beta, delta) parameterization. The paper&amp;rsquo;s first contribution is constructing identification conditions that overcome these results through two exclusion restrictions: a variable z that affects utility only through the perceived value of future states (played in the application by elicited beliefs about state evolution), and a variable r that acts as an imperfect signal of agent type but is uninformative about choices conditional on type.&lt;/p&gt;
&lt;p&gt;The general model accommodates a finite but unknown number of agent types — time-consistent (beta=1), time-inconsistent naive (beta&amp;lt;1, unaware of future present-bias), and time-inconsistent sophisticated (beta&amp;lt;1, aware of future present-bias) — as well as sub-types within each class. The paper proceeds in four identification steps when types are unobserved: identifying the total number of types (via the rank of an observable matrix), recovering type-specific choice probabilities, assigning type identities, and recovering preference parameters. For time-consistent and sophisticated agents, both beta and delta are point-identified. For naive agents, the parameters are set-identified in general, with point identification available under a monotonicity condition (Assumption 14) or by imposing a common exponential discount factor across types (Assumption 15).&lt;/p&gt;
&lt;p&gt;The empirical application studies demand for insecticide-treated nets (ITNs) and their periodic retreatment — a health-protective technology with low up-front cost but substantial future benefits — among households in malarious areas of rural Orissa. A key design feature is that households were offered either a standard ITN contract (with the option to purchase retreatment later) or a commitment contract bundling two consecutive retreatments, allowing the commitment product choice to serve as a noisy type signal r. Elicited beliefs about future state variables serve as the excluded z variable.&lt;/p&gt;
&lt;p&gt;The main empirical findings are: approximately 21% of the population is time-consistent, 49% are naive time-inconsistent, and 30% are sophisticated time-inconsistent — so time-inconsistent agents account for approximately 79% of the sample. The preferred estimates of the hyperbolic parameter beta are 0.16 for naive agents and 0.08 for sophisticated agents, indicating substantial present-bias in both groups. These estimates of the population type distribution and type-specific beta parameters are described as new to the literature.&lt;/p&gt;
&lt;p&gt;A counterfactual exercise quantifies the welfare cost of present-bias: the median undiscounted additional expected total cost of malaria during the study period attributable to under-investment in ITNs exceeds the price of a treated net by a factor of approximately six. However, because time-inconsistent households heavily discount future malaria costs, the discounted total costs of malaria are low for many inconsistent agents relative to the ITN price, explaining low demand from the agents&amp;rsquo; own subjective perspective. The paper also finds that commitment products are not disproportionately chosen by sophisticated agents — take-up of the commitment contract is actually higher among naive households — contradicting the deterministic mapping from commitment product purchase to sophistication that is commonly assumed in the literature. Finally, differences in per-period utilities across agent types exist but are not substantively important in explaining differential outcomes in the sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the core identification problem the paper addresses, and why is it hard?&lt;/strong&gt;
A: Even the standard exponential discount factor delta is generically not identified in dynamic discrete choice models (Rust 1994; Magnac and Thesmar 2002). This non-identification extends a fortiori to both beta and delta in the hyperbolic (beta, delta) model. When agents are also heterogeneous in unobserved type, the additional problem of identifying the population distribution of types — itself a key policy parameter — must be solved jointly with preference identification.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What two exclusion restrictions provide the key identifying variation?&lt;/strong&gt;
A: The first restriction is a variable z that affects utility only via the perceived value of future states but not per-period utility (Assumption 3); in the application this is played by elicited subjective beliefs about future state evolution. The second is a variable r that predicts agent type but, conditional on type and observables, provides no additional information about choices (Assumption 16); in the application r includes elicited time-preference indicators and the choice of the commitment versus standard ITN contract.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does the paper require at least three periods?&lt;/strong&gt;
A: Three periods are the minimum required to capture the notions of time-inconsistency studied here: with only two periods, no time-inconsistency problem would arise. Three periods allow the researcher to separately observe how an agent plans in period 1, how the agent actually behaves in period 2 (potentially deviating from the period-1 plan), and how the agent behaves in the terminal period 3 where the problem reduces to a static discrete choice.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is point-identified versus set-identified across agent types?&lt;/strong&gt;
A: For time-consistent agents, all per-period utilities and the (single) discount factor delta are point-identified. For sophisticated agents, both beta and delta are separately point-identified under the rank conditions in Assumptions 10-11. For naive agents, the parameters are in general only set-identified (Lemma 4 provides sharp bounds); point identification holds under either a monotonicity condition (Assumption 14) or the assumption that naive and sophisticated agents share the same exponential discount factor (Assumption 15).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the paper identify the total number of types in the population?&lt;/strong&gt;
A: The number of types equals the rank of a directly identified matrix P formed from the joint distribution of actions and states in adjacent time periods (Proposition 1). The rank provides a lower bound in general and equals the true number of types when the state space is sufficiently rich and type-specific choice probabilities vary sufficiently across the state space (Assumptions 17 and 19).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the paper distinguish naive from sophisticated agents among the identified type-specific choice probabilities?&lt;/strong&gt;
A: A key diagnostic is the function delta_hat_tau(x2,z2), which compares an agent&amp;rsquo;s period-1 view of the future against what would be expected given period 2-3 choices. For time-consistent and sophisticated agents, this function is constant across the state space (x2,z2); for naive agents it varies across the state space (Lemma 7, Proposition 2). This variation arises because naive agents incorrectly anticipate their future behavior in period 1, generating a wedge between planned and actual continuation values that shifts with the state.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What fraction of the sample is time-inconsistent, and what are the estimated beta parameters?&lt;/strong&gt;
A: Approximately 79% of the sample is time-inconsistent: 49% are naive and 30% are sophisticated. The preferred estimates of the hyperbolic (present-bias) parameter beta are 0.16 for naive agents and 0.08 for sophisticated agents. Both estimates indicate substantial present-bias. The paper states that these estimates of the population type distribution and the type-specific beta values are new to the literature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the welfare cost of present-bias in terms of malaria risk?&lt;/strong&gt;
A: Present-bias leads to lower ITN purchases and fewer retreatments, which increases the likelihood of contracting malaria. The median undiscounted additional expected total cost of malaria during the study period attributable to under-investment in ITNs exceeds the price of a treated net by a factor of approximately six. However, because inconsistent agents heavily discount future health costs, the discounted total costs of malaria are low relative to the ITN price for many such agents, which explains low demand from the agents&amp;rsquo; own subjective perspective despite large social costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the paper find about commitment products and agent sophistication?&lt;/strong&gt;
A: The commitment contract — bundling two consecutive retreatments — was designed to appeal to sophisticated present-biased agents who anticipate their future self-control problems. Contrary to the deterministic mapping from commitment product purchase to agent sophistication commonly assumed in the literature, take-up of the commitment contract is actually higher among naive households than sophisticated ones. The paper argues this is possible because the model allows commitment product choice to only imperfectly predict type, enabling a richer analysis than prior work that rules out type heterogeneity by assumption.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Are differences in per-period utilities across types an important alternative explanation for observed behavior?&lt;/strong&gt;
A: Per-period utilities do vary across agent types, but the paper finds they are not substantively important in explaining differential outcomes in the sample. This finding supports the interpretation that time-inconsistent preferences — rather than heterogeneity in static preferences over states — are the primary driver of the behavioral differences observed across agent types in this context.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the role of elicited beliefs in the identification strategy?&lt;/strong&gt;
A: Elicited beliefs about the future evolution of state variables serve as the excluded variable z that shifts the forward-looking component of the value function while leaving per-period utility unchanged. The use of expectational data, as advocated by Manski (2004), provides a natural and interpretable source of identifying variation for the discount parameters. The paper argues that this plausible exclusion restriction contributes to the encouraging Monte Carlo simulation results relative to other work in the identification literature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What happens to identification under partial sophistication?&lt;/strong&gt;
A: When agents are partially sophisticated — aware of some but not all of their future present-bias, so that beta_tilde in [beta, 1] rather than exactly equal to beta or 1 — the three time-preference parameters (delta, beta, beta_tilde) are not point-identified in general (Proposition 4 provides a set identification result). Point identification requires that the exponential discount factor delta be identified separately. The paper shows that partial and complete sophistication can be distinguished from time-consistency by whether the function delta_hat varies across the state space, and partially sophisticated types can be distinguished from fully sophisticated types under an additional variability condition (Assumption 23, Proposition 3).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hyperbolic (beta-delta) discounting:&lt;/strong&gt; A model of time-inconsistent preferences in which future utility at time s discounted from time t carries the factor beta*delta^(s-t), where beta&amp;lt;1 introduces an additional present-bias relative to pure exponential discounting. The parameter beta governs the wedge between the discount rate applied to immediate versus purely future tradeoffs; delta governs the intertemporal rate of substitution between any two future periods.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sophisticated vs. naive agents:&lt;/strong&gt; Both types are time-inconsistent (beta&amp;lt;1) and both are aware of their current present-bias. Sophisticated agents (tau_S) also correctly anticipate the extent of their future present-bias (beta_tilde = beta), while naive agents (tau_N) incorrectly believe their future self will behave as if beta_tilde = 1. This difference in beliefs about future behavior drives distinct choice dynamics across the three periods, providing the key observable variation used to distinguish the two types.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exclusion restriction (z variable):&lt;/strong&gt; A state variable that enters the transition probabilities and thus the value of future states but does not enter the current per-period utility function (Assumption 3). Variation in z shifts the forward-looking component of the Bellman equation while holding current utility fixed, providing the identifying variation needed to separately recover discount parameters from per-period utility parameters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Type indicator / type proxy (r):&lt;/strong&gt; An observed variable that is informative about an agent&amp;rsquo;s time-preference type but, conditional on type and other observables, provides no additional information about choices (Assumption 16). In the application, r includes elicited time-preference indicators and whether the agent chose the commitment versus standard ITN contract. Critically, the mapping from r to type is imperfect, so r does not directly reveal type for each individual.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conditional choice probability (CCP) inversion:&lt;/strong&gt; Following Hotz and Miller (1993), the type-specific conditional choice probabilities P_tau(a_t|x_t, z_t) — directly identified from data given type — can be inverted to recover per-period utility differences and combinations of discount parameters without solving the full dynamic programming problem. This approach underpins the constructive identification arguments throughout the paper.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Commitment contract:&lt;/strong&gt; A product design in which two consecutive ITN retreatments are bundled at purchase, intended to mitigate the time-inconsistency problem by removing the future self-control decision about retreatment. The commitment contract is theoretically predicted to be preferred by sophisticated present-biased agents; the paper finds this prediction fails empirically, with naive households showing higher take-up.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Present-bias welfare cost:&lt;/strong&gt; The undiscounted additional expected total cost of malaria attributable to under-investment in ITNs driven by present-bias. The paper estimates this cost exceeds the price of a treated net by a factor of approximately six at the median, capturing the gap between the social planner&amp;rsquo;s valuation of ITN adoption and the discounted valuation of time-inconsistent agents.&lt;/p&gt;</description></item><item><title>Ideological Alignment and Evidence-Based Policy Adoption</title><link>https://macropaperwarehouse.com/papers/ideological-alignment-and-evidence-based-policy-adoption/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/ideological-alignment-and-evidence-based-policy-adoption/</guid><description>&lt;p&gt;This paper investigates how the ideological alignment between knowledge-disseminating institutions and policymakers affects the adoption of evidence-based policies. The core research question is whether, and through which mechanisms, the ideology of the messenger — rather than the content of the message — determines whether local policymakers act on rigorous research evidence.&lt;/p&gt;
&lt;p&gt;The authors conduct a country-wide randomized controlled trial (RCT) across 5,678 touristic Spanish municipalities. The policy recommendation derives from Hinnosaar et al. (2021), an RCT demonstrating that minor improvements to municipalities&amp;rsquo; Wikipedia pages (adding photographs, local festival information, touristic landmark details) increased overnight tourist stays by 9%. This policy was chosen because it is ideologically neutral, low cost, within local policymakers&amp;rsquo; remit, and its implementation is directly traceable via Wikipedia edit histories.&lt;/p&gt;
&lt;p&gt;Municipalities were randomized into five treatment arms and a control group (approximately 950 municipalities each), stratified by ruling party ideology, population, and touristic accommodation count. Three arms received the same policy brief endorsed by: (1) an ideologically aligned think tank (FAES for right-wing municipalities, Fundación Alternativas for left-wing), (2) the ideologically opposite think tank, or (3) an ideologically nonsalient researcher from the London School of Economics. Two further arms received links to newspaper articles covering the same research from either an ideologically aligned outlet (El Mundo for right, Eldiario.es for left) or an ideologically opposite outlet. The control group received no information. The experiment ran from May to December 2022, with multiple reminder emails sent across the period.&lt;/p&gt;
&lt;p&gt;The main outcome is a binary indicator for whether a municipality&amp;rsquo;s Wikipedia page was changed in line with the recommended guidelines during the study period, coded blind to treatment status by two independent coders.&lt;/p&gt;
&lt;p&gt;Key findings: Pooled across all treatment arms, information provision increased the probability of policy adoption by approximately 0.98 percentage points (a 38% relative increase over the control group baseline), but this effect is only marginally above conventional significance thresholds (p-value = 0.13). The aggregate effect masks sharp heterogeneity by ideological alignment. When the informing institution&amp;rsquo;s ideology aligns with the policymaker&amp;rsquo;s, policy adoption increases by 1.68 percentage points (think tank) and 1.67 percentage points (newspaper) relative to the control group — equivalent to a 66% and 65% relative increase, respectively, both statistically significant at the 5% level. By contrast, information from an ideologically opposite institution produces a coefficient that is negligible and statistically indistinguishable from zero, indicating that misaligned information is no more effective than receiving no information at all. The ideologically nonsalient LSE researcher arm produced an intermediate effect (0.94 percentage points, 37% relative increase), but the p-value (0.27) exceeds conventional thresholds, and the effect is not statistically distinguishable from either the aligned or the control condition. Policy briefs and newspaper articles are equally effective when ideologically aligned (difference of 0.1 percentage points, p-value = 0.82).&lt;/p&gt;
&lt;p&gt;To decompose mechanisms, the authors propose a three-stage framework: (1) selective exposure to information, (2) belief updating, and (3) policy implementation. Email click-through rates (access to the full policy brief or article once the informing institution is revealed) do not differ significantly across treatment arms, ruling out selective exposure as the operative mechanism. A post-intervention online survey experiment with 1,600 policymakers from 1,196 municipalities shows that those receiving information from an aligned or nonsalient institution updated their beliefs about policy effectiveness significantly more than those receiving information from an opposite institution, implicating belief updating as one operative channel. However, comparing the survey experiment (where nonsalient and aligned treatments produce similar belief updating) with the main experiment (where the aligned arm adopts at nearly twice the rate of the nonsalient arm, though not statistically distinguishable) suggests that ideological alignment also affects the third stage — policy implementation — beyond mere belief updating.&lt;/p&gt;
&lt;p&gt;The estimated monetary cost of ideological misalignment is 2,192 euros per municipality per year, calculated using the impact of Wikipedia changes on touristic revenues from Hinnosaar et al. (2021).&lt;/p&gt;
&lt;p&gt;Scope conditions: The context is Spanish local government, a policy that is explicitly non-ideological, low-cost, and easily implemented. Generalizability to ideologically charged or costly policies is not established. Left-wing municipalities show larger responses to aligned information, though this heterogeneity is not statistically significant at conventional levels.&lt;/p&gt;
&lt;p&gt;Q: What is the baseline rate of policy adoption in the control group, and what does the aligned-institution treatment achieve in absolute terms?&lt;/p&gt;
&lt;p&gt;A: The paper reports that ideologically aligned institutions increase the share of municipalities implementing recommended Wikipedia changes by 1.68 percentage points (think tank) and 1.67 percentage points (newspaper) relative to the control group. Working backward from the stated 66% and 65% relative increases, this implies a control group baseline of approximately 2.5 percentage points. The aligned effects are statistically significant at the 5% level.&lt;/p&gt;
&lt;p&gt;Q: Does information from an ideologically opposite institution have any effect on policy adoption?&lt;/p&gt;
&lt;p&gt;A: No. The coefficient for opposite-ideology treatment arms is negligible in magnitude, closely resembling the near-zero coefficients from the placebo analysis conducted for the same months in 2019 (pre-intervention). The authors conclude that receiving information from an ideologically opposite institution is statistically indistinguishable from receiving no information at all. This null result is consistent across heterogeneity analyses by mayor ideology, municipality population, Wikipedia page length, and party type.&lt;/p&gt;
&lt;p&gt;Q: How does the ideologically nonsalient (LSE researcher) treatment compare to aligned and opposite arms?&lt;/p&gt;
&lt;p&gt;A: The nonsalient arm increases policy adoption by 0.94 percentage points (a 37% relative increase), approximately half the effect of the aligned arm (1.68 percentage points). However, the p-value is 0.27, and the effect is not statistically different from either the aligned arm (p-value = 0.34) or the control group at conventional confidence levels. The result should therefore be interpreted with caution.&lt;/p&gt;
&lt;p&gt;Q: Are policy briefs or newspaper articles more effective in promoting policy adoption?&lt;/p&gt;
&lt;p&gt;A: Neither format is significantly more effective than the other. Conditional on ideological alignment, the difference between policy brief and newspaper article effects is 0.1 percentage points with a p-value of 0.82. Both are equally effective when ideologically aligned with the receiving policymaker, a finding the authors describe as a novel contribution to the policy communication literature.&lt;/p&gt;
&lt;p&gt;Q: Does ideological alignment affect whether policymakers choose to access the full information (selective exposure)?&lt;/p&gt;
&lt;p&gt;A: No. Click-through rates on the links to policy briefs or newspaper articles — measured after policymakers have seen the informing institution&amp;rsquo;s identity — do not differ significantly across treatment arms. The observed average click-through rate is 6.42%. This null result is consistent with the hypothesis that policymakers do not strategically filter information acquisition based on the messenger&amp;rsquo;s ideology, at least for non-ideological policies.&lt;/p&gt;
&lt;p&gt;Q: What does the survey experiment reveal about belief updating?&lt;/p&gt;
&lt;p&gt;A: In the post-intervention survey experiment with 1,600 policymakers, participants first reported beliefs about a purportedly beneficial (but actually harmful) policy, then were randomly assigned to receive information about its negative effects from an aligned, opposite, or nonsalient think tank. Those receiving information from an aligned or nonsalient institution updated their beliefs significantly more than those receiving information from an ideologically opposite institution. This implicates belief updating — not just selective exposure — as a channel through which ideological alignment affects policy adoption.&lt;/p&gt;
&lt;p&gt;Q: Why do the authors conclude that ideological alignment also affects the third stage (policy implementation) beyond belief updating?&lt;/p&gt;
&lt;p&gt;A: In the survey experiment, aligned and nonsalient institutions produce statistically similar belief updating. Yet in the main field experiment, the aligned arm adopts policy at nearly twice the rate of the nonsalient arm (1.68 vs. 0.94 percentage points), although this difference is not statistically significant. The authors interpret this gap as suggestive evidence that ideological alignment affects policy implementation through channels beyond belief updating — such as career concerns, party cues, or the political economy of implementation — though they acknowledge the evidence is indirect and the treatment difference is not statistically distinguishable.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated economic cost of ideological misalignment?&lt;/p&gt;
&lt;p&gt;A: The authors estimate a cost of 2,192 euros per municipality per year attributable to ideological misalignment between the informing institution and the receiving policymaker. This calculation uses the estimated impact of Wikipedia changes on touristic revenues from Hinnosaar et al. (2021) and reflects not the cost of not implementing the policy, but the marginal cost of using an ideologically opposite rather than aligned institution to disseminate the research evidence.&lt;/p&gt;
&lt;p&gt;Q: How did outside researchers&amp;rsquo; predictions compare to actual results?&lt;/p&gt;
&lt;p&gt;A: Researchers surveyed on the Social Science Prediction Platform correctly anticipated the rank ordering of treatment effectiveness (aligned &amp;gt; nonsalient &amp;gt; opposite &amp;gt; control) but substantially overestimated adoption rates in every arm. They predicted relative increases of 144%, 103%, and 48% for aligned, nonsalient, and opposite conditions respectively, compared to actual relative increases of roughly 65%, 37%, and ~0%. Email opening rates were the most accurately predicted (49% predicted vs. 38% actual). The results highlight the difficulty of translating evidence into policy even for simple, low-cost interventions.&lt;/p&gt;
&lt;p&gt;Q: What are the main threats to validity and how are they addressed?&lt;/p&gt;
&lt;p&gt;A: Three main threats are considered. First, differential email opening rates across treatment arms: addressed by showing the informing institution was revealed only after email opening, and confirmed by finding no significant differences in opening rates across groups. Second, spillovers between municipalities: the endline survey shows only 5 of 236 control-group respondents reported receiving any information from external sources; spillover distance analyses in Table D.II find no significant effect on control municipalities&amp;rsquo; adoption rates. Third, contamination bias in multi-arm RCTs with strata fixed effects: addressed by replicating main results using the Goldsmith-Pinkham et al. (2022) method, yielding nearly identical estimates.&lt;/p&gt;
&lt;p&gt;Q: What heterogeneity is observed across left- and right-wing municipalities?&lt;/p&gt;
&lt;p&gt;A: The positive effect of receiving information from an ideologically aligned institution appears larger for left-wing municipalities, with coefficients approximately three times larger than for right-wing municipalities, but this difference is not statistically significant at conventional confidence levels. The authors caution that the strength of ideological alignment may differ systematically between the partner think tanks on the left and right, making direct comparisons between left- and right-wing effects difficult to interpret cleanly.&lt;/p&gt;
&lt;p&gt;Q: How does the paper relate to prior work on evidence-based policymaking?&lt;/p&gt;
&lt;p&gt;A: The closest prior work is Hjort et al. (2021) and Mehmood et al. (2024), which examine the impact of scientific evidence access on actual policy adoption, and DellaVigna and Kim (2022), which identifies ideology as a factor in the diffusion of innovative policies across governments. The present paper&amp;rsquo;s main contribution is being the first to isolate the causal effect of ideological alignment on policy adoption using a large-scale field experiment with real, authoritative ideological institutions — rather than surveys or hypothetical scenarios — while using a non-ideological policy recommendation to avoid confounding messenger ideology with policy ideology.&lt;/p&gt;
&lt;p&gt;Ideological alignment: In this paper&amp;rsquo;s usage, the congruence between the political ideology of the institution disseminating research evidence (think tank or newspaper) and the political ideology of the local government receiving that information. Alignment is operationalized by matching right-wing municipalities with right-leaning institutions (FAES, El Mundo) and left-wing municipalities with left-leaning institutions (Fundación Alternativas, Eldiario.es).&lt;/p&gt;
&lt;p&gt;Evidence-based policy adoption: The actual implementation by local policymakers of a policy recommendation derived from published peer-reviewed research — measured here as whether a municipality&amp;rsquo;s Wikipedia page was edited in line with specific recommended guidelines during the study period, not merely expressed intention or stated support.&lt;/p&gt;
&lt;p&gt;Knowledge brokers: Institutions, such as think tanks, that serve as intermediaries between academic researchers and policymakers, translating and disseminating research findings in accessible formats (policy briefs) to bridge the gap between evidence and policy.&lt;/p&gt;
&lt;p&gt;Nonsalient ideology: A condition in which the informing institution carries no salient or recognizable partisan affiliation, operationalized here by a foreign research university professor (LSE) whose institutional identity does not carry a clear left-right signal in the Spanish political context.&lt;/p&gt;
&lt;p&gt;Three-stage policy adoption framework: The authors&amp;rsquo; conceptual structure positing that ideology can interfere at three sequential stages: (1) selective exposure — whether policymakers choose to access information once the messenger&amp;rsquo;s ideology is revealed; (2) belief updating — whether policymakers revise their assessment of a policy&amp;rsquo;s effectiveness upon receiving evidence; and (3) policy implementation — whether policymakers act on updated beliefs to adopt the policy.&lt;/p&gt;
&lt;p&gt;Selective exposure: The tendency of individuals to avoid information from sources whose ideology conflicts with their own prior beliefs; in this paper, operationalized as differential click-through rates on links to policy briefs or news articles after the informing institution&amp;rsquo;s identity is revealed.&lt;/p&gt;
&lt;p&gt;Motivated reasoning: A documented tendency, also observed in policymakers, to reject or discount evidence that contradicts ideologically held prior beliefs — the mechanism proposed to explain why opposite-ideology information fails to update beliefs as effectively as aligned-ideology information.&lt;/p&gt;</description></item><item><title>Insurer Risk and Public Risk-Sharing: Quantifying the Value of Reinsurance</title><link>https://macropaperwarehouse.com/papers/insurer-risk-and-public-risk-sharing-quantifying-the-value-of-reinsurance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/insurer-risk-and-public-risk-sharing-quantifying-the-value-of-reinsurance/</guid><description>&lt;p&gt;Kim and Li study how publicly provided reinsurance affects insurer behavior and market outcomes in health insurance markets where firms face substantial cost uncertainty. The central question is whether standard expected-profit models—which predict that reinsurance reducing only cost volatility (not expected cost) should leave prices unchanged—miss an important mechanism: insurers internalizing the implicit financial cost of bearing claims uncertainty through &amp;ldquo;risk charges.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The paper develops a stylized monopoly-insurer model in which the insurer&amp;rsquo;s objective includes both expected claims cost and a risk charge term L(S), where S is a risk measure (e.g., standard deviation of total claims). This yields a first-order condition in which effective marginal cost includes both standard expected claims cost and a marginal risk charge. The model predicts that public reinsurance acts through two distinct channels: (1) a cost subsidy—reimbursing a share of high-cost claims reduces expected cost; and (2) risk protection—reducing the variance of claims lowers the risk charge and thus effective marginal cost. When both channels operate, the model predicts pass-through of public reinsurance to premiums can exceed unity, in contrast to the standard less-than-one pass-through under market power.&lt;/p&gt;
&lt;p&gt;Empirically, the authors use three primary data sources for the U.S. individual health insurance exchange market. NAIC Schedule S filings (2014–2023) provide transaction-level private reinsurance contracts, including ceded premiums, realized claims, and financial solvency measures. CMS Public Use Files and MLR reports provide plan-level premiums, enrollment, and claims. The Colorado All Payer Claims Database (CO APCD, 2014–2022) and Connect for Health Colorado administrative records (2015–2021) provide individual-level claims and insurance choices for structural analysis.&lt;/p&gt;
&lt;p&gt;Descriptive evidence establishes that 62% of exchange insurers purchase private reinsurance despite average reinsurance markups of 1.54 (reinsurance margin of 0.54), and that smaller, less financially solvent insurers are disproportionate buyers—consistent with risk charges driving demand for risk protection even at above-actuarially-fair prices.&lt;/p&gt;
&lt;p&gt;An event study exploiting staggered adoption of state-level public reinsurance programs finds that public reinsurance reduces premiums by approximately 14.5% on average (27% in Colorado Tiers 1–2, 46% in Tier 3), with a pass-through rate of 1.3—significantly greater than one (p = 0.037 one-sided). Public reinsurance reduces the probability of purchasing private reinsurance by 26 percentage points (a 42% reduction from baseline) and per-member private reinsurance expenditures by $19.5 (a 68% reduction from baseline). Premium and private reinsurance effects are larger for financially constrained insurers (RBC ratio below 3). No significant effects are found on insurer entry/exit, total medical expenses (ruling out moral hazard), or private reinsurance markups.&lt;/p&gt;
&lt;p&gt;The structural model, estimated on the Colorado exchange for 2017–2020, finds that the risk charge coefficient for regional insurers averages rho = 0.25, implying regional insurers face 9.8% higher effective costs than national insurers due to risk charges and private reinsurance expenses. Risk charges account for at least half the premium-cost wedge for small regional insurers. Counterfactual decomposition of Colorado&amp;rsquo;s program shows the direct cost subsidy accounts for approximately 75% of equilibrium price reductions; risk protection and competition effects together account for the remaining 25%. In a bang-for-buck comparison, public reinsurance dominates premium subsidies of equal government expenditure by approximately 20–30%, because reinsurance uniquely reduces risk charges and enhances competition by reducing smaller regional insurers&amp;rsquo; cost disadvantage.&lt;/p&gt;
&lt;p&gt;Q: What is the core theoretical innovation of the paper?
A: The paper adds a risk charge term L(S) to the standard expected-profit objective, where S is a risk measure of the insurer&amp;rsquo;s cost distribution. This makes the insurer behave &amp;ldquo;as if risk averse,&amp;rdquo; with effective marginal cost including both expected claims cost and a marginal risk charge that decreases with insured pool size due to risk pooling. When rho = 0, the model collapses to the standard monopoly case; when rho &amp;gt; 0, cost uncertainty directly inflates prices and creates a novel role for reinsurance even when reinsurance is actuarially fair priced.&lt;/p&gt;
&lt;p&gt;Q: What are the two distinct mechanisms through which public reinsurance affects insurer pricing?
A: The first is a cost subsidy: by reimbursing a portion of high-cost claims without requiring an actuarially fair premium upfront, public reinsurance lowers the insurer&amp;rsquo;s net expected cost. The second is risk protection: by providing ex-post payments for extreme health shocks, reinsurance reduces the variance of claims costs, lowering the risk charge component of effective marginal cost. Together, these channels can produce pass-through exceeding unity even under imperfect competition, where standard cost-subsidy pass-through is typically below one.&lt;/p&gt;
&lt;p&gt;Q: What does Proposition 1 say about actuarially fair reinsurance (theta = 1)?
A: Proposition 1(i) states that actuarially fair reinsurance—which does not alter net expected cost—still lowers the insurer&amp;rsquo;s price if and only if the insurer faces a risk charge (rho &amp;gt; 0). An insurer without risk charges is entirely unaffected by actuarially fair reinsurance. This result isolates the risk-protection channel as theoretically distinct from cost subsidization and establishes that pass-through exceeding one requires risk charges to be operative.&lt;/p&gt;
&lt;p&gt;Q: Why would an insurer purchase costly private reinsurance (theta &amp;gt; 1)?
A: Proposition 1(iii) shows that an insurer with no risk charge would never purchase private reinsurance with theta &amp;gt; 1, since it increases net expected cost with no offsetting benefit. An insurer facing a risk charge (rho &amp;gt; 0) may purchase private reinsurance because the risk-protection benefit—the reduction in cost variance and thus the risk charge—can outweigh the net cost increase. The paper documents that 62% of exchange insurers buy private reinsurance at an average markup of 1.54 (reinsurance margin 0.54), with smaller and financially weaker insurers more likely to purchase, consistent with this mechanism.&lt;/p&gt;
&lt;p&gt;Q: How does the paper establish empirically that insurers face and internalize cost uncertainty?
A: Three lines of evidence are presented. First, the CO APCD shows the claims distribution has a long right tail: the top 5% (1%) of consumers account for 68% (38%) of total expenses, and 2.5% of consumers exceed the $30,000 reinsurance threshold. Second, simulations show that with 1,000 enrollees, the probability that realized claims exceed expected costs by 25% is approximately 7%; even at 10,000 enrollees there is a 17% probability of exceeding expected costs by 5%. Third, in over 24% of insurer-year observations premium revenue falls short of realized claims costs, and the within-firm standard deviation of the claims-to-premium ratio is 0.15.&lt;/p&gt;
&lt;p&gt;Q: What are the event study findings on premiums?
A: Using staggered introduction of state-level public reinsurance programs, the event study finds premiums fell by 14.5% on average following program adoption. In Colorado specifically, Tiers 1 and 2 experienced 27% decreases and Tier 3 (highest reinsurance generosity) experienced a 46% decrease. The implied pass-through rate for 2020 is 1.3, meaning for every dollar the government spent on reinsurance, health insurance premiums fell by $1.30. A one-sided t-test rejects pass-through equal to one at p = 0.037.&lt;/p&gt;
&lt;p&gt;Q: What are the event study findings on private reinsurance?
A: Public reinsurance reduces the probability that an insurer purchases private reinsurance by 26 percentage points, a 42% decline from the pre-program baseline. Average per-member private reinsurance expenditures fall by $19.5, a 68% reduction from baseline. The substitution away from private reinsurance is consistent with the model prediction that public reinsurance displaces the demand for risk protection previously met by private markets, and reinforces the interpretation that risk management is a key driver of private reinsurance demand.&lt;/p&gt;
&lt;p&gt;Q: Do financially constrained insurers respond differently to public reinsurance?
A: Yes. The premium-reduction effect is significantly larger for insurers with RBC ratios below 3 (an additional interaction effect of -0.161 log points on top of the baseline -0.135). The reduction in per-member private reinsurance expenditures is also significantly larger for insurers with significant prior private reinsurance purchases (-$108.8 vs. baseline of -$19.5). This heterogeneity supports the hypothesis that the risk protection channel is more valuable for financially constrained insurers who face higher implicit costs of bearing risk.&lt;/p&gt;
&lt;p&gt;Q: Does public reinsurance affect insurer entry/exit, moral hazard, or private reinsurance markups?
A: The event study finds no statistically significant effect on market entry, total monthly medical expenses per enrollee, the probability that individual expenses exceed the reinsurance threshold (ruling out insurer moral hazard), or private reinsurance markups paid by primary insurers. These null results support the interpretation that premium reductions reflect reduced cost uncertainty rather than cost containment distortions, and that the competitive structure of the private reinsurance market is not directly altered by public programs.&lt;/p&gt;
&lt;p&gt;Q: What are the structural estimates of risk charges?
A: The estimated risk charge coefficient for regional insurers averages rho = 0.25. This implies that regional insurers incur, on average, 9.8% higher effective costs than national insurers (who are assumed not to face risk charges due to scale and diversification), stemming from both direct risk charges and private reinsurance expenses required to manage risk. Risk charges account for at least half the observed wedge between premiums and marginal claims costs for small regional insurers.&lt;/p&gt;
&lt;p&gt;Q: How does the structural model decompose the impact of Colorado&amp;rsquo;s reinsurance program?
A: Counterfactual analysis decomposes the equilibrium price reduction into three channels. The direct cost subsidy effect—reimbursing a share of high-cost claims between the $30,000 attachment point and $400,000 cap—accounts for approximately 75% of the price reduction. The risk protection effect (reduction in risk charges from lower portfolio variance) and the competition effect (smaller regional insurers facing lower cost disadvantages and competing more aggressively with national insurers) together account for the remaining 25% of the equilibrium price reduction.&lt;/p&gt;
&lt;p&gt;Q: How does public reinsurance compare to premium subsidies in bang-for-buck terms?
A: For equal government expenditure, public reinsurance is estimated to be approximately 20–30% more cost-effective than premium subsidies at reducing premiums. The advantage stems from two sources: reinsurance reduces risk charges, shifting down the marginal cost curve for regional insurers in a way demand-side premium subsidies do not; and reinsurance enhances competition by reducing the cost disadvantage of smaller regional insurers relative to national ones. The dominant effect is risk reduction rather than markup inflation, making reinsurance the more efficient instrument when the degree of financial risk is considerable.&lt;/p&gt;
&lt;p&gt;Q: What is the role of market size in risk charges, and why does this create a competitive asymmetry?
A: The model shows that the marginal risk charge decreases as the insured population grows (risk pooling), with marginal standard deviation equal to sigma_0 / (2*sqrt(q)), which vanishes as q approaches infinity. This implies that larger national insurers, covering very large populations, effectively face no risk charges, while smaller regional insurers face meaningful marginal risk charges. This size-asymmetry is the fundamental reason why public reinsurance disproportionately benefits smaller insurers—by reducing their risk charges, it narrows the cost gap with national insurers and intensifies competition.&lt;/p&gt;
&lt;p&gt;Q: What scope conditions apply to the structural findings?
A: The structural estimates are based on the Colorado individual health insurance exchange, covering years 2017–2020, chosen to avoid unsatisfactory early data quality and to net out systematic pandemic effects. The model assumes national insurers do not face risk charges in the baseline specification, and that aggregate (correlated) risk is not the primary driver during the sample period. Results are robust to staggered-treatment corrections (Callaway-Sant&amp;rsquo;Anna 2021; Borusyak et al. 2024), alternative outcome measures (benchmark premiums, Silver plan averages), alternative aggregation levels, and sensitivity analyses allowing for insurer entry/exit, correlated risks, moral hazard, and alternative risk charge functional forms.&lt;/p&gt;
&lt;p&gt;Q: What are the broader policy implications of the framework?
A: The framework applies to any market where firms face substantial cost uncertainty and internalize financial risk, including property and casualty insurance, flood insurance, wildfire insurance, and government loan guarantee programs. The analysis suggests that ignoring the risk protection channel causes policymakers to underestimate the effectiveness of public reinsurance relative to demand-side subsidies. Supply-side risk-sharing policies are particularly important for markets with small, financially constrained firms, where cost uncertainty most severely distorts pricing and competition, and where the competitive benefits of risk reduction are largest.&lt;/p&gt;
&lt;p&gt;Risk Charge: An additional cost term in the insurer&amp;rsquo;s objective function representing the implicit financial cost of bearing claims uncertainty, formalized as L(S) where S is a risk measure of total cost. Risk charges make the insurer behave &amp;ldquo;as if risk averse,&amp;rdquo; raising effective marginal cost above expected claims cost. In the baseline model the risk charge equals rho times the standard deviation of total claims.&lt;/p&gt;
&lt;p&gt;Risk Charge Coefficient (rho): The parameter governing the insurer&amp;rsquo;s marginal cost of financial risk, estimated structurally at an average of 0.25 for regional insurers in Colorado. It can be interpreted as either a direct risk-aversion parameter, the marginal cost of regulatory capital, or a reduced-form representation of financial and regulatory frictions that make bearing cost uncertainty costly.&lt;/p&gt;
&lt;p&gt;Risk Protection Channel: The mechanism through which reinsurance (public or private) reduces claims cost variance and thereby lowers the insurer&amp;rsquo;s risk charge, distinct from the cost-subsidy channel. The risk protection channel is operative even for actuarially fair reinsurance (theta = 1) and is responsible for pass-through rates exceeding unity under public reinsurance programs.&lt;/p&gt;
&lt;p&gt;Cost Subsidy Channel: The mechanism through which subsidized public reinsurance (theta less than 1) lowers the insurer&amp;rsquo;s net expected claims cost by reimbursing a share of high-cost claims without charging an actuarially fair premium. This channel operates regardless of whether the insurer faces risk charges and is the primary channel in standard models.&lt;/p&gt;
&lt;p&gt;Pass-Through Rate: The ratio of premium reduction to government expenditure on reinsurance. In standard models with market power, pass-through of cost subsidies is typically below one; the paper documents a pass-through rate of 1.3 in Colorado (p = 0.037 for the null of pass-through equal to one), attributing the excess to the risk protection channel reducing both expected cost and cost uncertainty simultaneously.&lt;/p&gt;
&lt;p&gt;Stop-Loss Reinsurance: A contract structure in which the reinsurer reimburses the primary insurer for individual claims costs exceeding a deductible (attachment point) kappa up to a cap. In Colorado&amp;rsquo;s program the attachment point is $30,000 and the cap is $400,000, with government coinsurance rates of 40–80% depending on county tier. More generous reinsurance corresponds to lower kappa; full reinsurance is kappa = 0.&lt;/p&gt;
&lt;p&gt;Risk-Based Capital (RBC) Ratio: The ratio of capital surplus (assets minus liabilities) to required risk-based capital, used by NAIC as a measure of insurer solvency. NAIC scrutinizes companies with RBC ratios below 200%; the paper uses RBC ratio below 3 as a proxy for financial constraint in heterogeneity analysis, finding larger premium and private reinsurance responses among constrained insurers.&lt;/p&gt;
&lt;p&gt;Tail-End Risk: The risk arising from the possibility that a small fraction of enrollees incurs extremely high medical costs, concentrated in the right tail of the claims distribution. In Colorado, the top 5% of consumers account for 68% of total expenses; tail-end risk is especially severe for small insurers with fewer than 10,000–100,000 enrollees and is the primary motivation for private reinsurance purchases even at above-actuarially-fair prices.&lt;/p&gt;</description></item><item><title>Investing in Influence: Investors, Portfolio Firms, and Political Giving</title><link>https://macropaperwarehouse.com/papers/investing-in-influence-investors-portfolio-firms-and-political-giving/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/investing-in-influence-investors-portfolio-firms-and-political-giving/</guid><description>&lt;p&gt;This paper investigates whether institutional investors influence the political activities of their portfolio firms, using political action committee (PAC) giving as a window into the broader question of whether institutional investors can leverage their concentrated ownership to extract benefits from portfolio firms for their own interests rather than those of their clients.&lt;/p&gt;
&lt;p&gt;The sample covers 574 institutional investors (those with at least $100 million in assets under management, i.e., 13-F filers) matched to 2,456 portfolio firms that had PACs, over the period 1980–2018. The primary source of variation is the first acquisition by an institutional investor of at least one percent of a portfolio firm&amp;rsquo;s outstanding shares, yielding 68,387 large acquisition events. PAC giving data come from FEC records matched by name to investor and firm entities. The main regression specification examines how the relationship between investor and firm PAC contributions to the same congressional district changes after such an acquisition, using a saturated set of fixed effects including firm × investor, firm × congressional district, firm × election cycle, investor × congressional district, investor × election cycle, and district × election cycle.&lt;/p&gt;
&lt;p&gt;The central finding is that, following a large block purchase, a firm&amp;rsquo;s PAC giving mirrors more closely that of the acquiring investment management company. In the preferred specification (column 8 of Table 2), the probability that a portfolio firm gives to a politician supported by its investor&amp;rsquo;s PAC increases by 31 percent after an acquisition. Using a cosine similarity measure of investor-firm PAC giving, the mean similarity of 0.10 at the acquisition cycle rises by 0.02–0.03 (a 20–30 percent increase) by the fourth post-acquisition election cycle.&lt;/p&gt;
&lt;p&gt;A key identification concern is that acquisitions may be driven by shared political preferences rather than representing a causal effect. To address this, the authors exploit stock index inclusions as exogenous shifters of institutional investor block purchases: when a firm is added to an index for the first time, passive indexers are compelled to rebalance toward that firm regardless of political alignment. Restricting to 5,601 index-inclusion acquisitions by passive investors, the authors find near-identical effect sizes (beta1 = 0.0132 in column 8 versus 0.0135 in the full sample), and an event study shows no pre-trend in giving convergence for the index subsample, in contrast to a slight pre-trend in the full sample. Divestment events exhibit the symmetric negative pattern: the interaction of post-divestment and investor PAC giving falls by between -0.074 and -0.058 across specifications.&lt;/p&gt;
&lt;p&gt;The authors argue that investors drive the convergence rather than portfolio firms adjusting investor preferences. Around acquisition dates, firms exhibit a larger drop in between-election-cycle cosine similarity than investors do. In a difference-in-differences comparison of the acquisition period relative to the preceding period, the difference in stability between investors and firms is 0.075 (significant at the 1 percent level), indicating that firms shift their giving more than investors. Investors obtaining a board seat at the portfolio firm amplifies the effect: in the preferred specification, the board-seat interaction is more than twice as large as the acquisition-alone interaction.&lt;/p&gt;
&lt;p&gt;Heterogeneity analysis provides evidence that the convergence reflects investors&amp;rsquo; partisan tastes rather than coordinated profit-maximizing political strategy. Acquisitions by more partisan investors (those whose giving is more skewed toward one party) produce a convergence coefficient roughly twice as large (0.020) as less partisan investors (0.010). Private fund families show more than twice the convergence effect of publicly owned fund families. The partisan composition of firm giving also shifts: a firm acquired by an investor giving exclusively to Republicans sees its Republican share increase by 2.8 percentage points relative to a baseline of 47.4 percent (a 5.9 percent increase).&lt;/p&gt;
&lt;p&gt;Finally, higher overall institutional ownership is associated with an increase in total PAC giving at the firm level, and this expanded giving does not go disproportionately to politicians on committees overseeing issues the firm actively lobbies — suggesting the ownership-driven increment in political spending is non-strategic from the firm&amp;rsquo;s profit standpoint and likely serves investors&amp;rsquo; own interests.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the central research question and why does it matter?&lt;/strong&gt;
The paper asks whether institutional investors influence the political giving of portfolio firms, motivated by the broader concern that the rise of institutional ownership — from 6 percent of U.S. public equities in 1950 to 65 percent in 2017 — concentrates not only economic but also political power in the hands of a small number of asset managers. This matters because if investors shape firms&amp;rsquo; PAC giving to serve investors&amp;rsquo; own preferences rather than firms&amp;rsquo; profit interests, it represents a misuse of corporate resources and a potential amplification of a small group&amp;rsquo;s political voice.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What data are used and how is the sample constructed?&lt;/strong&gt;
The analysis draws on 13-F filings (investors with at least $100M AUM) from Thomson-Reuters, matched to FEC PAC records via fuzzy and manual name matching. The resulting sample contains 574 investors with PACs and 2,456 portfolio firms with PACs, spanning 1980–2018. The Cartesian product of investor-firm pairs is restricted to those connected by at least one large acquisition event (defined as first acquisition of at least 1 percent of outstanding shares), yielding 68,387 such events. PAC contributions are measured at the investor- and firm-congressional-district-election-cycle level, linked to House of Representatives winners using MIT Election Data files.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the baseline regression and what does it find?&lt;/strong&gt;
The baseline regression (equation 1) interacts Log Investor PAC with a Post indicator (equal to 1 after the first large acquisition and while the stake is maintained) at the investor-firm-congressional-district-election-cycle level, with a saturated set of fixed effects. The coefficient on the interaction (beta1) is positive and highly significant (p &amp;lt; 0.001) across all eight specifications, ranging from 0.013 to 0.032. In the preferred specification, the increase in giving similarity is 31 percent relative to the pre-acquisition baseline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do the authors establish causality and rule out endogenous acquisitions?&lt;/strong&gt;
The primary identification strategy uses first-time inclusions of firms in stock indices (approximately 1,000 indices tracked in the sample) as exogenous shifters: passive indexers must rebalance toward the included firm regardless of political alignment. This subsample of 5,601 index-inclusion acquisitions produces near-identical coefficient estimates (0.0132 versus 0.0135 in the full sample), and the event study for this subsample shows no pre-trend in giving convergence, unlike the slight pre-trend in the full sample. Equality of the two coefficients cannot be rejected at standard significance levels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What evidence shows it is firms adjusting to investors rather than the reverse?&lt;/strong&gt;
The authors compute between-election-cycle cosine similarity separately for investors and firms around acquisitions. On average, investors exhibit more stable giving than firms at acquisition dates (Cos(xi,t, xi,t+1) &amp;gt; Cos(xf,t, xf,t+1)). The difference-in-differences estimate — comparing the acquisition period to the preceding period — is 0.075 (significant at 1 percent), indicating a relatively larger break in firm giving. Over a two-cycle window, the difference-in-differences estimate is 0.083, again indicating convergence is driven by firms shifting toward investors rather than the reverse.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What role does board representation play?&lt;/strong&gt;
In approximately 5 percent of acquisitions in the sample, the investor obtains a board seat. In specifications that include both the acquisition effect (Post × Log Investor PAC) and a board-membership interaction (Board × Log Investor PAC), both terms are positive and significant at the 1 percent level. In the preferred specification, the board-seat interaction is more than twice as large as the acquisition-alone interaction, indicating that a direct governance channel — board representation — substantially amplifies the convergence in political giving.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the divestment analysis show?&lt;/strong&gt;
Symmetric to the acquisition results, divestment events (where an investor exits a stake of at least 1 percent held for at least one election cycle) are associated with a decline in investor-firm PAC giving correlation. Post-divestment interaction coefficients range from -0.074 to -0.058 across specifications, and an event study confirms the correlation falls sharply after the divestment cycle.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does investor partisanship affect the magnitude of influence?&lt;/strong&gt;
Yes. Classifying investors as &amp;ldquo;More Partisan&amp;rdquo; (above-mean absolute deviation from 50/50 party split) versus &amp;ldquo;Less Partisan,&amp;rdquo; the interaction coefficient for More Partisan investors (0.020) is roughly twice that of Less Partisan investors (0.010). After a large acquisition by a fully Republican-giving investor, the acquired firm&amp;rsquo;s giving to that politician increases by 23.5 percent; the comparable figure for a Less Partisan investor is 7.6 percent. This pattern holds in both the full sample and the index-inclusion subsample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do private versus public fund families differ in their influence?&lt;/strong&gt;
Private fund families (e.g., Vanguard, Fidelity) show more than twice the convergence coefficient of publicly owned fund families (e.g., BlackRock, State Street, Invesco). The authors attribute this to private fund managers facing less outside scrutiny, allowing their giving to more readily reflect the preferences of owners and managers. Private investors also show greater partisan polarization: the 10th–90th percentile Republican-giving range for private investors is 6.3–100 percent, versus 21.7–88.3 percent for public investors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does increased institutional ownership expand overall firm PAC spending?&lt;/strong&gt;
Yes. In firm-year level regressions, institutional ownership is a positive and significant predictor of total firm PAC giving (significant at at least the 5 percent level in both cross-sectional and firm-fixed-effects specifications). Total corporate political expenditure by sample firms increased by nearly a factor of six over 1980–2018. The authors note that while many factors contribute, increased institutional ownership may be at least partly responsible for this expansion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does the additional giving driven by institutional ownership go to strategically important politicians for the firm?&lt;/strong&gt;
No. Regressions relating institutional ownership to giving to politicians on congressional committees overseeing issues the firm actively lobbies (a standard measure of politicians&amp;rsquo; strategic importance to firms) yield near-zero and statistically weak point estimates. In the preferred firm-fixed-effects specification, the share of total PAC giving devoted to such strategically relevant politicians is negatively associated with institutional ownership at marginal significance (p &amp;lt; 0.10), consistent with the interpretation that ownership-driven incremental political spending is non-strategic from the firm&amp;rsquo;s own profit perspective and expands total giving rather than displacing strategic giving.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the policy and legal implications?&lt;/strong&gt;
The authors flag three concerns: (i) the ownership-driven increment in political spending may represent a misuse of corporate resources that does not serve portfolio firm shareholders; (ii) it may constitute an illegal activity, since using a firm&amp;rsquo;s PAC to reimburse or proxy for an investor&amp;rsquo;s own political preferences can run afoul of campaign finance law; and (iii) it is a channel through which unequal resources amplify the political voice of a small number of fund managers at the expense of dispersed ultimate investors who are likely unaware of and do not sanction these contributions. The findings challenge the Supreme Court&amp;rsquo;s premise in Citizens United that corporate political speech reflects shareholder profit maximization.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PAC comovement (investor-firm giving similarity):&lt;/strong&gt; The increase in the probability that a portfolio firm&amp;rsquo;s PAC donates to a politician also supported by an acquiring investor&amp;rsquo;s PAC, measured as the interaction coefficient between Log Investor PAC and a Post-acquisition indicator in the baseline regression. In the preferred specification this represents a 31 percent increase relative to the pre-acquisition baseline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cosine similarity (cross-time and cross-entity):&lt;/strong&gt; A measure defined as the Euclidean dot product between two vectors of PAC giving (either the same entity across adjacent election cycles, or investor versus firm in the same cycle), taking values between 0 and 1, where 1 indicates identical giving patterns. Used both to confirm convergence post-acquisition and to attribute that convergence to firm rather than investor adjustment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Index-inclusion acquisition:&lt;/strong&gt; A large block purchase that results from a firm being added for the first time to a stock index tracked by a passive institutional investor, used as an exogenous shifter of investor stakes that is orthogonal to investor-firm political alignment. There are 5,601 such events in the sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Partisanship (investor):&lt;/strong&gt; Classified as &amp;ldquo;More Partisan&amp;rdquo; if an investor&amp;rsquo;s absolute deviation from a 50/50 party split in PAC donations is above the sample mean. More partisan investors produce roughly twice the convergence effect on portfolio firm giving compared to less partisan investors, used as evidence that personal political preferences rather than profit-maximizing business strategy drive the convergence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Post indicator (Postift):&lt;/strong&gt; A binary variable equal to 1 for all election cycles following an investor&amp;rsquo;s first acquisition of at least 1 percent of a portfolio firm&amp;rsquo;s outstanding shares, and remaining 1 as long as the investor holds any stake in the firm. The key source of temporal variation in the baseline regression.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Strategically important politicians:&lt;/strong&gt; Members of Congress sitting on committees that oversee issues on which a firm actively lobbies, identified by crosswalking lobbying reports from the Senate Office of Public Records to relevant committee jurisdictions. Used to test whether ownership-driven political giving displaces or supplements firm-profit-motivated giving.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Board seat channel:&lt;/strong&gt; The mechanism through which investor influence on firm political giving is amplified when the investor obtains representation on the portfolio firm&amp;rsquo;s board of directors (present in approximately 5 percent of acquisitions). The board interaction coefficient is more than twice the acquisition-alone coefficient in the preferred specification.&lt;/p&gt;</description></item><item><title>Lives Versus Livelihoods: The Impact of the Great Recession on Mortality and Welfare</title><link>https://macropaperwarehouse.com/papers/lives-versus-livelihoods-the-impact-of-the-great-recession-on-mortality-and-welfare/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/lives-versus-livelihoods-the-impact-of-the-great-recession-on-mortality-and-welfare/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Does the Great Recession reduce or increase mortality, and what are the welfare implications of incorporating recession-induced mortality changes into standard macroeconomic welfare frameworks?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and Identification.&lt;/strong&gt; The authors exploit spatial variation in the severity of the 2007–2009 Great Recession across 741 U.S. Commuting Zones (CZs), following the empirical design of Yagan (2019). The primary shock variable is the percentage-point change in the CZ unemployment rate between 2007 and 2009. The key identifying assumption is that no concurrent shocks to mortality coincide with the timing and geographic pattern of the Great Recession shock. Pre-trend evidence supports this: CZs subsequently harder hit experienced a slight relative &lt;em&gt;increase&lt;/em&gt; in mortality before 2007, which is the opposite sign from the main effect, supporting the validity of the design.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; Mortality data come from CDC restricted-use death certificate microdata (2003–2016) covering the universe of U.S. deaths, combined with SEER population denominators. A 20 percent random sample of Medicare enrollees aged 65–99 provides an individual-level panel that directly addresses concerns about endogenous migration. The main outcome is the log age-adjusted CZ mortality rate; economic indicators come from BLS, BEA, and FHFA; air pollution data from the EPA AQS monitor network (PM2.5); morbidity from the BRFSS; nursing home characteristics from federal certification inspections.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Mortality Finding.&lt;/strong&gt; A one-percentage-point increase in the local unemployment rate between 2007 and 2009 is associated with a 0.50 percent decline (SE = 0.15) in the annual age-adjusted mortality rate in 2007–2009, and a 0.58 percent decline (SE = 0.34) in 2010–2016; the two periods are statistically indistinguishable (p = 0.78). Because the national average unemployment rate rose by 4.6 percentage points, the Great Recession on average reduced the annual age-adjusted mortality rate by approximately 2.3 percent, with effects persisting for at least 10 years. The authors note this is equivalent to approximately two years of secular mortality improvement at the pre-recession trend pace of 1.1 percent per year. For a 55-year-old, the estimates imply that 1 in 25 gained an extra year of life from a shock of this magnitude.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity by Cause of Death.&lt;/strong&gt; Mortality declines appear across most major causes. Cardiovascular disease (34 percent of 2006 deaths) declines by 0.65 percent per percentage-point unemployment increase (SE = 0.21) and accounts for approximately 48 percent of the total estimated mortality reduction. Motor vehicle mortality falls by 1.7 percent (SE = 0.56) and liver disease by 1.1 percent (SE = 0.43). Suicides show a statistically significant 1.7 percent decline (SE = 0.5) in the 2010–2016 period. The notable exception is cancer (the second-largest cause of death), for which the estimated effect is a precise null of 0.02 percent (SE = 0.11). The null cancer result is interpreted as a specification check: if mortality declines were spurious (e.g., driven by population mismeasurement), cancer mortality should also decline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity by Demographics.&lt;/strong&gt; Recession-induced mortality declines are similar in percentage terms across gender and race/ethnicity, and statistically equi-proportional across age groups (p-value for equality across 25–64 versus 65+: 0.76). Because mortality is heavily concentrated in the elderly, those aged 65 and over account for approximately 74.3 percent of averted deaths, roughly proportional to their 72.5 percent share of 2006 mortality. The most striking heterogeneity is by education: the entire mortality decline is concentrated among the approximately 52 percent of the population with a high school degree or less. The estimated 2007-2016 effect is −1.3 percent per percentage-point unemployment increase (SE = 0.56) for those with high school or less, compared to +0.34 percent (SE = 0.68) for those with more than high school (statistically distinguishable at p &amp;lt; 0.01).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanisms.&lt;/strong&gt; The authors distinguish internal effects (own reduced employment or consumption improving health) from external effects (externalities from reduced aggregate economic activity, holding own employment/consumption fixed). Evidence strongly favors external effects as the primary driver. Three-quarters of averted deaths accrue to the elderly, who experienced no direct income effects from the labor market shock. Moreover, the timing pattern—an immediate mortality drop that does not grow over time—is inconsistent with health-behavior channels (e.g., smoking cessation, improved diet) that would build up gradually. Direct tests find no statistically significant impact on self-reported health behaviors (smoking, drinking, exercise) and no impact on healthcare use among Medicare enrollees.&lt;/p&gt;
&lt;p&gt;Among external channels, neither reduced spread of infectious disease nor improved nursing home staffing receives empirical support. Reduced air pollution (PM2.5) is identified as a quantitatively important channel. A one-percentage-point increase in CZ unemployment is associated with a 0.16 µg/m³ decline in PM2.5 (SE = 0.04), a 1.3 percent decline relative to the 2006 national average of 12 µg/m³. A mediation analysis (controlling for the PM2.5 shock) attenuates the estimated mortality effect by 37 percent, from −0.52 percent to −0.33 percent per percentage-point unemployment increase. Back-of-the-envelope calculations combining the PM2.5 decline with external estimates of PM2.5-mortality elasticities suggest pollution can explain 17 to 35 percent of total recession-induced mortality declines.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Lag Structure.&lt;/strong&gt; Exploiting variation in the speed of post-recession labor market recovery (measured by 2010–2016 EPOP ratio changes) conditional on the initial shock, the authors find that mortality reductions persist in areas that have fully recovered economically by 2016, suggesting lagged mortality effects of the initial economic downturn beyond what contemporaneous economic conditions alone explain.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Welfare Analysis.&lt;/strong&gt; The authors extend the Krebs (2007) consumption-based welfare cost-of-recessions model to incorporate endogenous mortality. For a 45-year-old with γ = 2 and a value of a statistical life-year (VSLY) of $250k (five times annual consumption), accounting for endogenous mortality reduces the willingness to pay to avoid all future recessions from 2.00 percent of average annual consumption to 0.91 percent—a reduction of approximately 55 percent. Starting around age 55, recessions become welfare-improving on net. For the Great Recession specifically, at age 55 endogenous mortality reduces the welfare cost by approximately 25 percent (from 2.39 to 1.80 percent of average annual consumption). Because mortality declines are concentrated among those with high school or less, accounting for endogenous mortality also substantially mitigates—and at older ages reverses—the finding that the Great Recession was more costly for the less educated.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions and Caveats.&lt;/strong&gt; (i) The design captures only differential local effects, not nationwide impacts (e.g., stock market collapse, nationwide malaise). (ii) Mortality impacts may not generalize to milder recessions, though the relationship appears approximately linear in shock size. (iii) The analysis excludes morbidity, though limited evidence suggests morbidity is also pro-cyclical and roughly equi-proportional across ages. (iv) The welfare analysis begins at age 35 and does not account for longer-run mortality costs of recession entry for younger cohorts.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-baseline-empirical-specification-and-why-does-the-design-exploit-cross-sectional-variation-rather-than-time-series-panel-regressions"&gt;Q1. What is the baseline empirical specification, and why does the design exploit cross-sectional variation rather than time-series panel regressions?&lt;/h3&gt;
&lt;p&gt;The estimating equation regresses the log age-adjusted CZ mortality rate on an interaction of the CZ-level Great Recession shock (2007–2009 unemployment change) with year indicators, plus CZ and year fixed effects, weighted by 2006 CZ population. The authors prefer this to the standard two-way fixed effects panel approach (area and year FE with contemporaneous unemployment rate) for three reasons: (1) it directly identifies the full dynamic lag structure of the shock rather than imposing contemporaneity; (2) exploiting a single spatially differentiated shock reduces risk of confounding from other concurrent area-level shocks; (3) the panel can be linked to individual-level Medicare data, allowing explicit control for endogenous migration, which the existing literature cannot do.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-address-the-concern-that-mortality-rate-declines-might-simply-reflect-unmeasured-population-outflows-from-hard-hit-areas-rather-than-genuine-reductions-in-deaths"&gt;Q2. How does the paper address the concern that mortality rate declines might simply reflect unmeasured population outflows from hard-hit areas rather than genuine reductions in deaths?&lt;/h3&gt;
&lt;p&gt;The authors offer two main responses. First, cancer mortality shows a precise null effect despite being the second-leading cause of death; if unmeasured population losses were driving the results, cancer deaths should decline proportionally. Second, using the Medicare individual-level panel, they fix each enrollee&amp;rsquo;s location at their 2003 CZ and find a statistically significant mortality decline of 0.35 percent per percentage-point unemployment increase in the reduced-form (2007–2009 period). A control function approach that instruments current-year location with 2003 location yields an estimate of −0.37 percent (SE = 0.17), similar to the baseline −0.50 percent from the aggregate specification, confirming that migration bias is not the primary driver.&lt;/p&gt;
&lt;h3 id="q3-how-long-do-the-mortality-reductions-from-the-great-recession-persist-and-does-the-paper-identify-whether-these-are-contemporaneous-or-lagged-effects"&gt;Q3. How long do the mortality reductions from the Great Recession persist, and does the paper identify whether these are contemporaneous or lagged effects?&lt;/h3&gt;
&lt;p&gt;The 2007–2009 period estimate is −0.50 percent per percentage-point unemployment increase and the 2010–2016 period estimate is −0.58 percent, and these are statistically indistinguishable (p = 0.78). To identify whether persistence reflects ongoing economic effects or true lagged mortality effects, the authors compare CZs with above- vs. below-median 2010–2016 EPOP recovery (conditional on initial shock decile). Both groups show similar 2010–2016 mortality declines despite the above-median recovery CZs having returned to pre-recession employment levels by 2016. This finding is consistent with lagged mortality effects of the initial economic downturn that persist independently of current economic conditions.&lt;/p&gt;
&lt;h3 id="q4-are-mortality-reductions-concentrated-among-individuals-already-near-death-harvesting-or-do-they-represent-meaningful-longevity-gains"&gt;Q4. Are mortality reductions concentrated among individuals already near death (&amp;ldquo;harvesting&amp;rdquo;), or do they represent meaningful longevity gains?&lt;/h3&gt;
&lt;p&gt;The authors use a Medicare auxiliary model to predict counterfactual remaining life expectancy for each enrollee based on age, demographics, and chronic conditions. The marginal life saved has only about 6 percent lower counterfactual remaining life expectancy than a typical decedent of the same age, and this difference is statistically insignificant. Because effects persist over 10 years (not just days or weeks), short-run mortality displacement (harvesting) is not the operative concern. The 6 percent difference is also small enough that the authors do not adjust their welfare analysis for it.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-educational-gradient-in-mortality-impacts-and-is-it-explained-by-age-composition-or-other-confounders"&gt;Q5. What is the educational gradient in mortality impacts, and is it explained by age composition or other confounders?&lt;/h3&gt;
&lt;p&gt;Mortality declines are entirely concentrated among those with a high school degree or less: the 2007–2016 estimate is −1.3 percent per percentage-point unemployment increase (SE = 0.56) for this group versus +0.34 percent (SE = 0.68) for those with more than high school, distinguishable at p &amp;lt; 0.01. This gradient holds within age groups (confirmed in Appendix analysis), and further disaggregation shows no mortality declines for those with some college or college-or-more separately. In Medicare data, the elderly mortality effect is concentrated among the approximately 12 percent enrolled in Medicaid (a proxy for low income), reinforcing the socioeconomic concentration.&lt;/p&gt;
&lt;h3 id="q6-what-evidence-rules-out-improved-health-behaviors-increased-exercise-reduced-smoking-reduced-alcohol-as-the-main-mechanism"&gt;Q6. What evidence rules out improved health behaviors (increased exercise, reduced smoking, reduced alcohol) as the main mechanism?&lt;/h3&gt;
&lt;p&gt;Two types of evidence argue against this channel. First, three-quarters of averted deaths are among the elderly, who experienced no direct income or employment effects from the local labor market shock and would not plausibly change their health behaviors in response to someone else losing employment. Second, the mortality decline is immediate in 2007 and flat through 2016 rather than growing over time; smoking cessation, for example, takes 10–15 years to accumulate mortality effects. Direct tests of behavioral outcomes from BRFSS find no statistically significant impact on smoking, drinking, exercise, or flu vaccination rates, individually or pooled. The pooled average treatment effect on six morbidity measures is statistically significant and negative (suggesting morbidity improvements), but behavioral covariates show no movement.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-evidence-for-and-against-improved-nursing-home-care-as-a-mechanism"&gt;Q7. What is the evidence for and against improved nursing home care as a mechanism?&lt;/h3&gt;
&lt;p&gt;Prior literature (Stevens et al. 2015; Konetzka et al. 2018; Antwi and Bowblis 2018) documents that recessions increase nursing home staffing and reduce nursing home deaths in earlier decades. However, the authors find no evidence for this channel in the Great Recession context. Estimated mortality impacts are virtually identical (approximately 0.5 percent per percentage-point unemployment increase) for the 7 percent of the elderly in nursing home care and the 93 percent not in nursing home care. Direct measures of nursing home staffing (direct-care staff hours per resident-day, highly skilled nurses ratio) show no statistically significant change in harder-hit areas: the point estimate for direct-care hours is −0.11 percent (SE = 0.22) in 2007–2009. Nursing home occupancy rates and resident characteristics also show no significant changes.&lt;/p&gt;
&lt;h3 id="q8-how-is-the-quantitative-importance-of-the-air-pollution-channel-estimated-and-what-are-the-two-complementary-approaches-used"&gt;Q8. How is the quantitative importance of the air pollution channel estimated, and what are the two complementary approaches used?&lt;/h3&gt;
&lt;p&gt;Approach 1 (back-of-the-envelope): The authors combine their estimate that a one-percentage-point unemployment increase reduces PM2.5 by 0.16 µg/m³ with external estimates from Deryugina et al. (2019) of PM2.5&amp;rsquo;s effect on elderly daily mortality, rescaled to annual exposure. This calculation implies pollution explains 17–35 percent of total recession-induced mortality declines, depending on which Deryugina et al. mortality estimates are used. Approach 2 (mediation analysis): Adding the county-level PM2.5 shock as an additional control in the mortality regression attenuates the Great Recession mortality coefficient from −0.52 percent to −0.33 percent per percentage-point unemployment increase—a 37 percent attenuation. Both approaches are suggestive rather than definitive, as the mediation analysis requires the strong assumption that the recession shock and PM2.5 shock are conditionally independent of other unmeasured mediators.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-specific-calibration-parameters-in-the-welfare-model-and-how-does-the-paper-set-the-mortality-decline-parameter"&gt;Q9. What are the specific calibration parameters in the welfare model and how does the paper set the mortality decline parameter?&lt;/h3&gt;
&lt;p&gt;The authors extend Krebs (2007)&amp;rsquo;s income process calibration (pH = 0.03, pL = 0.05, dH = 0.09, dL = 0.21, g = 0.02, σ = 0.01, πH = 0.5) and use 2007 SSA life tables for age-specific mortality rates in normal times. The recession mortality parameter is set to dm = −0.015 for all ages, derived from a 3.1 percentage-point unemployment increase in a typical recession multiplied by the estimated 0.5 percent mortality decline per percentage-point. VSLY values are parameterized at two, five, or eight times annual consumption ($100k, $250k, or $400k at $50k annual consumption). Risk aversion γ takes values 1.5, 2, and 2.5. For the Great Recession-specific exercise, dmA = −0.023 (4.6 × 0.5 percent), dmHS = −0.037, and dmC = 0.0006.&lt;/p&gt;
&lt;h3 id="q10-how-does-accounting-for-endogenous-mortality-change-the-distributional-welfare-analysis-of-the-great-recession-by-education-group"&gt;Q10. How does accounting for endogenous mortality change the distributional welfare analysis of the Great Recession by education group?&lt;/h3&gt;
&lt;p&gt;Under exogenous mortality, the welfare cost of the Great Recession at age 35 is 2.89 percent of average annual consumption for those with high school or less versus 1.23 percent for those with more than high school—the less educated bear roughly twice the burden. Under endogenous mortality, the mortality declines are concentrated entirely among the less educated (dmHS = −0.037 vs. dmC ≈ 0), so accounting for mortality disproportionately offsets welfare losses for that group. By around age 65, the welfare costs of the Great Recession converge across education groups, and after age 65, the less educated bear &lt;em&gt;lower&lt;/em&gt; welfare costs than the more educated, reversing the exogenous-mortality ranking. This result depends on the same education differential in mortality impacts that drives the main empirical finding.&lt;/p&gt;
&lt;h3 id="q11-what-robustness-checks-demonstrate-that-the-baseline-mortality-estimates-are-not-driven-by-geographic-or-functional-form-choices"&gt;Q11. What robustness checks demonstrate that the baseline mortality estimates are not driven by geographic or functional-form choices?&lt;/h3&gt;
&lt;p&gt;The baseline CZ-level estimate of −0.50 percent (SE = 0.15) is replicated almost exactly at the state level (−0.62, SE = 0.25) and county level (−0.49, SE = 0.10). A Poisson regression yields −0.45 percent (SE = 0.14). Dropping the top/bottom decile of CZs by shock size yields −0.46 percent (SE = 0.16). Adding Census-division-by-year fixed effects attenuates the estimate slightly to −0.38 percent (SE = 0.14) but retains statistical significance. Dropping CZs with high fracking activity and dropping the ten most populous CZs both produce estimates similar to baseline. Quartile regressions show monotone mortality reductions across quartiles of the unemployment shock, consistent with approximate linearity.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-expert-survey-reveal-about-prior-beliefs-and-how-does-the-papers-finding-compare"&gt;Q12. What does the expert survey reveal about prior beliefs, and how does the paper&amp;rsquo;s finding compare?&lt;/h3&gt;
&lt;p&gt;In a spring 2023 survey of over 300 experts, 50 percent predicted the Great Recession would &lt;em&gt;increase&lt;/em&gt; mortality and only 27 percent predicted a decrease. Of those predicting a decrease, 93 percent gave a magnitude larger (in absolute value) than the paper&amp;rsquo;s negative point estimate of 0.50 percent per percentage-point unemployment increase, and 82 percent gave a prediction larger than the upper bound of the 95 percent confidence interval. This illustrates that the paper&amp;rsquo;s finding—mortality is meaningfully pro-cyclical during the Great Recession—was highly surprising to the empirical and policy economics community.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Pro-cyclical mortality&lt;/strong&gt;: The phenomenon whereby mortality rates fall during economic downturns and rise during expansions. The paper documents this for the Great Recession using a spatial identification strategy, in contrast to the time-series correlation that had weakened in the two decades before the Great Recession. The term &amp;ldquo;pro-cyclical&amp;rdquo; means mortality moves in the same direction as the business cycle (up in booms, down in recessions), implying recessions are associated with fewer deaths.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Internal vs. external effects (of recessions on mortality)&lt;/strong&gt;: The paper distinguishes internal effects—whereby an individual&amp;rsquo;s own reduced employment or consumption affects her own mortality—from external effects, which are changes in mortality from reduced aggregate economic activity that hold constant one&amp;rsquo;s own employment and consumption. This distinction has direct welfare implications: external effects (e.g., less pollution from lower industrial output) are genuine welfare improvements for people who did not lose income, while internal effects of behavioral change are mitigated by the envelope theorem if behavior is privately optimal.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Commuting Zone (CZ) shock&lt;/strong&gt;: The paper&amp;rsquo;s primary treatment variable, defined as the percentage-point change in the CZ unemployment rate between 2007 and 2009. CZs are aggregations of counties (741 total) designed to approximate local labor markets. The median CZ experienced a 4.6-percentage-point increase, with substantial variation ranging from roughly 2.9 points (bottom quartile) to 6.7 points (top quartile).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Value of a Statistical Life-Year (VSLY)&lt;/strong&gt;: The dollar value placed on one additional year of life in expectation, used in the welfare calibration. In the paper&amp;rsquo;s framework it equals VSLY = bcγ − c/(γ−1), where b is a preference parameter governing the marginal utility of life-years. Results are reported for VSLYs of $100k, $250k, and $400k corresponding to two, five, and eight times average annual consumption of $50k, following Hall and Jones (2007).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous mortality in welfare analysis&lt;/strong&gt;: The paper&amp;rsquo;s central theoretical contribution is augmenting the Krebs (2007) welfare cost-of-recessions framework to allow mortality to vary with the aggregate state of the economy. When mortality is endogenously lower in recessions, the willingness to pay to eliminate recession risk falls—and at high enough VSLY or old enough ages, recessions become welfare-improving because the mortality benefit outweighs the consumption cost.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mortality displacement (harvesting)&lt;/strong&gt;: The possibility that short-run mortality declines merely reflect the premature death of already-frail individuals being slightly delayed, without meaningful longevity gains. The paper argues this is not the operative concern given 10-year persistence and uses auxiliary Medicare models to show marginal lives saved have only 6 percent shorter counterfactual life expectancy than average decedents of the same age.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PM2.5 mediation analysis&lt;/strong&gt;: An empirical approach in which the county-level change in fine particulate matter (PM2.5, in µg/m³) between 2006 and 2010 is added as a covariate in the mortality regression. Under the assumption that the recession shock and the PM2.5 shock are conditionally independent of other unmeasured mediators, the attenuation in the recession-mortality coefficient when controlling for PM2.5 identifies the share of the mortality effect operating through the pollution channel. A 37 percent attenuation is found in the 2007–2009 period.&lt;/p&gt;</description></item><item><title>Manipulation-Robust Prediction</title><link>https://macropaperwarehouse.com/papers/manipulation-robust-prediction/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/manipulation-robust-prediction/</guid><description>&lt;p&gt;This paper addresses the problem of algorithmic manipulation: when consequential decisions are encoded in machine learning algorithms, individuals strategically alter their behavior to achieve desired outcomes, undermining the predictive validity of the algorithm. The authors develop a &amp;ldquo;strategy-robust&amp;rdquo; approach to training decision rules that explicitly models the incentives and costs of manipulation, producing rules that remain stable even when fully transparent. They then deploy and evaluate this approach in a large field experiment in Kenya — the first real-world implementation and evaluation of such a strategy-robust empirical decision rule.&lt;/p&gt;
&lt;p&gt;The theoretical framework considers a policymaker who observes training data with features x_i and optimal decisions y_i, and wishes to estimate a decision rule to apply to new instances where behavior may be manipulated. While the standard approach (OLS or LASSO) selects a rule optimal for the training distribution, the strategy-robust approach models how individuals will adjust behavior in response to the incentive structure implied by any given rule. Under linear decision rules and quadratic manipulation costs, each individual shifts behavior by C_i^{-1} * beta away from their &amp;ldquo;bliss level,&amp;rdquo; where C_i captures individual- and behavior-specific manipulation costs. The strategy-robust estimator finds the rule that minimizes prediction error in the counterfactual world where people manipulate — a &amp;ldquo;Stackelberg&amp;rdquo; solution that commits the policymaker to a rule while anticipating equilibrium behavioral responses. Unlike LASSO, which penalizes all features equally without regard to their manipulability, the strategy-robust approach attenuates the weight on features that are both easily manipulated and subject to manipulation noise.&lt;/p&gt;
&lt;p&gt;The empirical setting is a smartphone app (&amp;ldquo;Smart Sensing&amp;rdquo;) deployed to 1,557 participants in Nairobi, Kenya, in collaboration with the Busara Center. The app passively collected over 1,000 behavioral indicators (calls, texts, app usage, mobility, etc.) and delivered weekly financial &amp;ldquo;challenges&amp;rdquo; that rewarded participants based on decision rules randomly assigned to them. Average weekly payouts were calibrated to approximate typical digital credit loan amounts in Kenya at the time (approximately $4.80). The experiment has two phases: a training phase using control (beta = 0) and simple single-behavior incentive rules to estimate manipulation cost parameters via GMM, and an implementation phase using complex multi-feature decision rules to compare strategy-robust versus LASSO classifiers.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. First, participants demonstrably manipulate behavior: a joint F-test that incentive diagonals all equal zero is rejected with p &amp;lt; 0.001. The number of texts sent was 49 times more responsive to incentives than the number of people called during the workday. Outgoing communications are cheaper to manipulate than incoming, and simple behaviors (e.g., average talk time) more manipulable than complex ones (e.g., standard deviation of talk time). Individuals who self-report higher tech skills find manipulation 9% easier on average, and the 90th percentile of gaming ability finds manipulation twice as easy as the 10th percentile.&lt;/p&gt;
&lt;p&gt;Second, in the implementation phase, strategy-robust decision rules outperform LASSO when the decision rule is made transparent to participants. Across all pooled outcomes, strategy-robust rules reduce RMSE by 11% (p = 0.024) relative to LASSO under transparency. For the single income-prediction outcome alone, the improvement is 5% ($0.19 RMSE reduction) but not statistically significant (p = 0.507).&lt;/p&gt;
&lt;p&gt;Third, the framework enables estimation of the &amp;ldquo;cost of transparency.&amp;rdquo; Making naive LASSO rules transparent lowers performance by 23%. Switching to strategy-robust rules under full transparency reduces that performance decline to 9.2% — a 60% reduction in the cost of transparency. The model predicts this cost to be 9.8%, close to the implemented value of 11.3%.&lt;/p&gt;
&lt;p&gt;The scope of the findings is bounded by the linear model with quadratic manipulation costs, a particular population of Kenyan smartphone users, and financial incentive magnitudes comparable to small digital credit loans. The mechanism relies on experimentally estimating manipulation cost parameters, though the authors also show that expert elicitation provides a correlated but noisier substitute (correlation 0.30 with experimental estimates).&lt;/p&gt;
&lt;p&gt;Q: What is the core market failure the paper addresses, and why do standard fixes fail?&lt;/p&gt;
&lt;p&gt;A: Standard machine learning training assumes the relationship between observed features and outcomes is stable, but implementing a consequential decision rule creates incentives for individuals to manipulate the features on which the rule is based (Goodhart&amp;rsquo;s Law; Lucas critique). The two common industry responses — restricting to &amp;ldquo;stable&amp;rdquo; predictors and keeping rules secret — are inadequate: restricting predictors amounts to a dogmatic prior that manipulation costs are either infinite or zero, while secrecy is increasingly at odds with demands for algorithmic transparency and fails anyway when sophisticated actors reverse-engineer the rule. Periodic retraining treats manipulation as generic covariate shift, can produce non-converging oscillations, and requires observing mistakes before learning from them.&lt;/p&gt;
&lt;p&gt;Q: How does the strategy-robust estimator differ from OLS and LASSO?&lt;/p&gt;
&lt;p&gt;A: OLS maximizes fit within the unincentivized training sample but ignores that implementing beta will shift behavior; LASSO adds a regularization penalty but still assumes behavior remains fixed at bliss levels and so penalizes all features equally regardless of manipulability. The strategy-robust estimator replaces each individual&amp;rsquo;s observed behavior x_i with their anticipated counterfactual behavior x_tilde_i(beta) = x_i + C_i^{-1} * beta, and finds the beta that minimizes prediction error in this manipulated distribution — a Stackelberg equilibrium. It attenuates features that are easily manipulated or subject to high manipulation noise, shifting weight toward harder-to-manipulate features even when the latter are less predictive in the training data.&lt;/p&gt;
&lt;p&gt;Q: What are the three ways the strategy-robust estimator differs from standard estimators?&lt;/p&gt;
&lt;p&gt;A: First, it anticipates level shifts in behavior: behaviors respond to beta, so observed training behaviors are replaced by counterfactual manipulated behaviors. Second, it accounts for signaling and noise: when manipulation ability correlates with the outcome of interest, manipulation can be informative about type (as in Spence 1973), but unobserved heterogeneity in gaming ability that is unrelated to outcomes introduces noise that attenuates coefficients on manipulable behaviors. Third, it achieves subgame perfection by anticipating how behaviors would respond to off-path deviations in beta, rather than assuming behaviors are fixed when beta deviates — yielding a Stackelberg rather than a one-step best-response solution.&lt;/p&gt;
&lt;p&gt;Q: How were manipulation cost parameters estimated in the Kenya experiment?&lt;/p&gt;
&lt;p&gt;A: In the training phase, each participant was randomly assigned to simple single-behavior incentive rules (e.g., &amp;ldquo;earn 12 Ksh. per incoming call this week, up to 250 Ksh.&amp;rdquo;) or control rules (beta = 0). This random variation in per-behavior incentives identifies how sensitive each behavior vector is to incentives, enabling GMM estimation of individual and behavior-specific cost parameters C and the heterogeneity scaling parameter omega. Off-diagonal elements of C were regularized to zero due to noisy estimation; diagonal elements used LASSO penalization with lambda = 1.0 set by cross-validation. Observable heterogeneity was allowed to vary with self-reported tech skills, which explained the most variation in preliminary analysis.&lt;/p&gt;
&lt;p&gt;Q: What patterns were found in manipulation costs across behaviors?&lt;/p&gt;
&lt;p&gt;A: Outgoing communications are cheaper to manipulate than incoming communications. Text messages, being relatively cheap to send, are more manipulable than calls. Simple behaviors such as average call duration are more manipulable than complex behaviors such as the standard deviation of talk time. Cross-behavior elasticities exist but are mostly noisy: 94.5% of off-diagonal incentive effects are not statistically significant (p &amp;lt; 0.05), 3.6% are significantly positive, and 1.8% are significantly negative.&lt;/p&gt;
&lt;p&gt;Q: How large is heterogeneity in gaming ability, and what predicts it?&lt;/p&gt;
&lt;p&gt;A: Individuals who self-report advanced or higher tech skills find it on average 9% easier to manipulate behaviors. Including unobserved heterogeneity, the 90th percentile of gaming ability finds manipulation twice as easy as the 10th percentile. Much of the heterogeneity arises from unobservables not captured by observables in the model.&lt;/p&gt;
&lt;p&gt;Q: What happened when the naive LASSO rule was made transparent versus when the strategy-robust rule was made transparent?&lt;/p&gt;
&lt;p&gt;A: Under the transparent treatment, participants received the full coefficients of the decision rule plus access to an interactive earnings calculator. Making naive LASSO rules transparent lowered performance by 23% relative to the opaque naive rule (RMSE $3.780 versus $4.641 in pooled outcomes). Switching to strategy-robust rules under full transparency reduced the performance decline to 9.2% — corresponding to a 60% reduction in the cost of transparency. The model predicted this cost to be 9.8%, which is close to the implemented value of 11.3%.&lt;/p&gt;
&lt;p&gt;Q: What does the reduced-form evidence on behavior change under complex decision rules show?&lt;/p&gt;
&lt;p&gt;A: Under the opaque treatment, participant behavior responses to complex decision rules were largely statistically insignificant and often in the wrong direction — 38.5% of estimated behavioral effects are in the same direction as the incentivized behavior. Under the transparent treatment, 75.4% of point-estimated effects are in the same direction as the incentive, confirming that transparency is a prerequisite for meaningful manipulation in this setting.&lt;/p&gt;
&lt;p&gt;Q: How does the paper compare strategy-robust estimation to iterative retraining?&lt;/p&gt;
&lt;p&gt;A: Simulation results show that iterative retraining of a naive LASSO model approaches the performance of the strategy-robust method after approximately 4 iterations. However, simulated performance of iterative retraining then begins to deteriorate; for the intelligence outcome, performance eventually falls below baseline performance before any retraining began. This illustrates that myopic best responses can produce non-convergent or suboptimal dynamics, while the strategy-robust approach finds the equilibrium rule directly.&lt;/p&gt;
&lt;p&gt;Q: How does the paper compare strategy-robust estimation to the &amp;ldquo;intuitive&amp;rdquo; approach of simply excluding highly manipulable features?&lt;/p&gt;
&lt;p&gt;A: The intuitive approach of excluding features above a manipulability threshold reduces predicted manipulability but also discards useful predictors. In some cases, the exclusions leave LASSO with no behaviors predictive enough to include, reducing performance. The strategy-robust approach can extract signal even from manipulable behaviors by adjusting their weights to account for manipulation noise, and outperforms the intuitive exclusion approach in the simulations reported in the Supplemental Appendix.&lt;/p&gt;
&lt;p&gt;Q: Can manipulation costs be estimated without an experiment?&lt;/p&gt;
&lt;p&gt;A: The authors briefly explore expert elicitation as a nonexperimental alternative: 171 individuals were surveyed to predict how Kenyans would manipulate phone behaviors when incentivized. Experts generally predicted lower costs (more manipulability) than observed experimentally, but the correlation between expert predictions and experimental estimates is 0.30. Using expert-elicited costs to train the strategy-robust model improved simulated performance substantially for one focal outcome and had an inconsequential negative effect for the other. Costs can also potentially be estimated from market prices and first principles when a structural model of underlying manipulations is available.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s interpretation of its results through the lens of the Lucas critique?&lt;/p&gt;
&lt;p&gt;A: The paper frames its contribution as a machine learning interpretation of Lucas (1976): just as implementing an economic policy changes the behavioral relationships on which the policy was calibrated, implementing a predictive decision rule beta changes the distribution of the very features the rule is based on. The key insight is that this counterfactual world has predictable structure — including a feature in the model tends to induce manipulation in that feature of a magnitude directly related to beta — so counterfactual fit can be estimated and rules can be optimized to perform well in the equilibrium they induce.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications for algorithmic transparency?&lt;/p&gt;
&lt;p&gt;A: The framework allows a policymaker to quantify and reduce the performance cost of transparency. The estimated equilibrium cost of transparency is roughly 10% when using strategy-robust rules, substantially less than the approximately 23% cost of making naive rules transparent. This means that strategy-robust rules can be disclosed — satisfying demands for a &amp;ldquo;right to explanation&amp;rdquo; under regulations such as GDPR — while losing far less performance than opaque naive rules would lose if disclosed.&lt;/p&gt;
&lt;p&gt;Strategy-robust decision rule: A decision rule trained to anticipate that individuals will manipulate the features on which it is based, by replacing observed training behaviors with anticipated counterfactual manipulated behaviors in the loss function. It yields a Stackelberg equilibrium in which the policymaker commits to a rule while correctly forecasting the equilibrium behavioral response.&lt;/p&gt;
&lt;p&gt;Manipulation costs (C_i): Individual- and behavior-specific quadratic costs that determine how far an individual shifts behavior from their bliss level in response to the incentive implied by a decision rule&amp;rsquo;s coefficient vector beta. Higher costs imply less behavioral response; costs are parameterized to allow separable heterogeneity by person and by behavior.&lt;/p&gt;
&lt;p&gt;Bliss level (x_i): An individual&amp;rsquo;s unincentivized behavior — the behavior they would exhibit absent any decision rule (i.e., when beta = 0). Estimated from control periods in the experiment.&lt;/p&gt;
&lt;p&gt;Gaming ability (gamma_i): Individual-level scaling factor for manipulation costs; a higher value means lower costs and easier manipulation. Modeled as a function of observable characteristics (e.g., self-reported tech skills) and unobservable heterogeneity.&lt;/p&gt;
&lt;p&gt;Counterfactual fit: Predictive fit evaluated in the counterfactual state of the world where the decision rule is implemented and agents manipulate their features in response. The strategy-robust approach maximizes counterfactual fit, sacrificing within-sample fit (as measured on unmanipulated training data) to improve performance in deployment.&lt;/p&gt;
&lt;p&gt;Cost of transparency: The reduction in predictive performance of a decision rule when its coefficients are disclosed to the individuals being evaluated. In the experiment, disclosure reduces performance of naive LASSO rules by 23% and strategy-robust rules by 9.2%, implying strategy-robust rules reduce the cost of transparency by 60%.&lt;/p&gt;
&lt;p&gt;Stackelberg equilibrium: The solution concept in which the policymaker (leader) commits to a decision rule, correctly anticipating the best-response behavior of individuals (followers), rather than taking behavior as fixed or updating myopically. The strategy-robust estimator implements this equilibrium concept.&lt;/p&gt;
&lt;p&gt;Performative prediction: The broader phenomenon, drawing on Perdomo et al. (2020), whereby a decision rule changes the distribution of the data it is applied to. The paper&amp;rsquo;s strategy-robust approach is an empirically estimable solution within this framework.&lt;/p&gt;</description></item><item><title>Marginal Returns to Public Universities</title><link>https://macropaperwarehouse.com/papers/marginal-returns-to-public-universities/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/marginal-returns-to-public-universities/</guid><description>&lt;p&gt;This paper asks whether enrolling in an American public university generates positive net returns for marginal students — those who barely qualify for admission — and whether those returns justify public expenditures. The question is policy-relevant because marginal students have weak academic preparation, face high dropout risk, and the net returns to expanding admission margins are theoretically ambiguous.&lt;/p&gt;
&lt;p&gt;The author assembles administrative records spanning all 35 public universities in Texas, covering the universe of Texas public high school graduates from 2004–2014 (approximately 2.7 million students). Texas public universities collectively enroll over 10 percent of all American public university students. The data link high school records (test scores, demographics, coursework, attendance, disciplinary infractions) to college application and admission records, postsecondary enrollment and degree completion records, financial aid packages, institutional expenditure data from IPEDS, and quarterly earnings records from the Texas Workforce Commission unemployment insurance system.&lt;/p&gt;
&lt;p&gt;The identification strategy exploits hundreds of decentralized SAT/ACT score cutoffs in university admissions — varying across schools and application years — that generate sharp discontinuities in admission probability. A fuzzy regression discontinuity design compares applicants just above versus just below each cutoff. On average, crossing a cutoff raises the probability of admission by 27 percentage points and the probability of enrolling at the target university by 15 percentage points. Density tests and pre-college covariate balance validate the smoothness assumptions. The typical cutoff complier is more disadvantaged than the average college applicant but comparable to the average Texas high school graduate.&lt;/p&gt;
&lt;p&gt;Roughly half of cutoff compliers would fall back to another, typically less selective, four-year institution if rejected; 43 percent would fall back to a two-year community college; and only about 6 percent would forgo higher education entirely. The pooled estimates therefore blend intensive-margin effects (more selective versus less selective four-year college) with extensive-margin effects (four-year college versus community college or no college).&lt;/p&gt;
&lt;p&gt;Main causal findings for enrollment compliers: the typical marginally admitted student completes approximately one additional year of credits in the four-year sector and becomes 12 percentage points more likely to ever earn a bachelor&amp;rsquo;s degree from any institution. About half of the additional four-year credits are offset by 15 fewer credits in the two-year sector, and associate degree or certificate completion falls by 7 percentage points. All bachelor&amp;rsquo;s degree gains are in non-STEM fields; STEM degree completion shows no detectable increase. Compliers become about 3 percentage points more likely to hold a graduate degree by 10 years out.&lt;/p&gt;
&lt;p&gt;On earnings, admitted compliers earn less than rejected counterparts in the first five years due to continued enrollment. Year six is the crossover point; by years 8–12, compliers earn a stable 8.6 percent earnings premium in log terms (8.2 percent in dollar ratio terms, representing a LATE of $3,339 against an untreated complier mean of $40,829), with earnings ranks rising approximately 4 percentiles from a base near the 50th percentile.&lt;/p&gt;
&lt;p&gt;Marginally admitted students pay no additional net tuition on average: $4,600 in additional gross tuition is nearly fully offset by grant aid, though they take on $5,300 more in student loans. Society incurs approximately $10,000 in additional educational expenditures per complier. Internal rates of return are 26 percent for students, 16 percent for society, and 7 percent for the government budget. At a 3 percent discount rate, the lifetime net present value of enrolling the typical marginal applicant is approximately $80,000 — $70,000 accruing to the student and $10,000 to taxpayers.&lt;/p&gt;
&lt;p&gt;Earnings gains are similar across institutions of varying selectivity, but significantly smaller for low-income compliers, who spend more time enrolled, complete fewer degrees, and major in less lucrative fields. A bounding method shows that extensive-margin compliers (those who would otherwise not attend any four-year college) experience larger effects than intensive-margin compliers.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question and why is credible evidence scarce?
A: The paper asks whether enrolling marginal students in American public universities generates positive net returns — private, social, and fiscal — and what drives heterogeneity in those returns. Credible evidence is scarce because most existing work is correlational and fails to account for selection bias: individuals with more college education may have had pre-existing advantages, confounding college&amp;rsquo;s causal effect with systematic sorting into it. Even if average returns are positive, the policy-relevant question is whether the marginal student — who has weak preparation and high dropout risk — represents a good investment.&lt;/p&gt;
&lt;p&gt;Q: What is the regression discontinuity design, and what does the first stage look like?
A: The author infers hundreds of decentralized SAT/ACT score cutoffs across approximately 700 application cells (combinations of university, year, GPA quartile, and test type) by searching for the score value with the largest discontinuity in admission and enrollment within each cell. This procedure delivers a superconsistent estimator of each cell&amp;rsquo;s true cutoff. Pooled across all cells, crossing a cutoff raises the probability of admission by 27 percentage points and the probability of enrollment at the target university by a precisely estimated 15 percentage points. The density of applicants and a rich set of pre-college characteristics run smoothly through the cutoffs, supporting the exclusion restriction.&lt;/p&gt;
&lt;p&gt;Q: Who are the cutoff compliers, and are they representative of any broader population?
A: Compliers — applicants who enroll in the target university if and only if they barely cross its cutoff — comprise approximately 15 percent of marginal applicants. In observable characteristics, compliers are roughly representative of the broader population of marginal applicants at the cutoff. They are significantly more disadvantaged than the average public university applicant, but broadly comparable to the average Texas public high school graduate in terms of academic preparation and family income.&lt;/p&gt;
&lt;p&gt;Q: What are the next-best alternatives for marginal applicants who are rejected?
A: Approximately 47 percent of compliers would fall back to another Texas four-year college (mostly public), 43 percent to a two-year community college, and approximately 9 percent would not enroll in any Texas institution. National Student Clearinghouse data for the 2008–2014 cohorts confirm that only 4 percent of untreated compliers attend a college outside the THECB universe, meaning approximately 6 percent of all compliers truly forgo higher education altogether if rejected. The empirically relevant extensive margin is therefore between the four-year sector and the two-year sector, not between college and no college.&lt;/p&gt;
&lt;p&gt;Q: How does cutoff crossing change the institutional characteristics a complier experiences?
A: Compliers are propelled into substantially better-resourced environments: the average math test score of college peers rises by half a standard deviation; peers are 12 percentage points less likely to have been low-income; gross tuition rises by $2,400 (a 42 percent increase over the untreated complier mean of $5,700); educational spending per student rises by $3,200 (43 percent over the untreated mean); peers&amp;rsquo; 10-year BA completion rate rises by 28 percentage points; and peer mean earnings 8–12 years after college entry are $6,700 higher.&lt;/p&gt;
&lt;p&gt;Q: What are the educational attainment effects?
A: Cutoff crossing causes compliers to complete approximately 28 additional credits at any four-year institution (roughly one full year of a four-year program) and increases the probability of ever earning a bachelor&amp;rsquo;s degree by 12 percentage points, raising the completion rate from approximately 40 percent to just above 50 percent. About 15 fewer two-year sector credits are offset against the four-year gains, and associate degree or certificate completion falls by 7 percentage points. All bachelor&amp;rsquo;s degree gains are in non-STEM fields; there is no detectable increase in STEM degrees. Graduate degree completion rises by approximately 3 percentage points by 10 years out.&lt;/p&gt;
&lt;p&gt;Q: What is the earnings trajectory, and when does the premium materialize?
A: Admitted compliers earn less than rejected counterparts in the first five years after application because they remain enrolled longer. Year six is the crossover point. By years 8–12, the earnings premium stabilizes at approximately 8.6 percent in log terms and 8.2 percent in dollar ratio terms (a LATE of $3,339 against an untreated complier mean of $40,829). Earnings rank rises by approximately 4 percentiles from a base near the 50th percentile. These results are robust across sandwich earnings, all-quarters-with-earnings, and zero-imputed specifications.&lt;/p&gt;
&lt;p&gt;Q: What does the cost-benefit analysis show?
A: Marginally admitted students pay no additional net tuition on average: $4,600 in additional gross tuition is nearly fully offset by additional grant aid. They do borrow $5,300 more in student loans, likely financing higher room, board, and consumption costs at four-year colleges. From society&amp;rsquo;s perspective, compliers generate approximately $10,000 in additional educational expenditures. Cumulative undiscounted earnings benefits surpass costs after 8 years for students, 11 years for society, and 19 years for taxpayers. At a 3 percent discount rate, the lifetime net present value is approximately $80,000 total — $70,000 accruing to the student and $10,000 to taxpayers — with internal rates of return of 26 percent for students, 16 percent for society, and 7 percent for the government budget.&lt;/p&gt;
&lt;p&gt;Q: Does selectivity of the admitting institution predict larger earnings returns?
A: No. Compliers at more selective institutions experience substantially larger increases in peer quality than those at less selective institutions, but they are also less likely to be on the extensive margin of four-year enrollment and experience smaller BA attainment gains. These factors roughly offset, producing no systematic difference in earnings gains across institutions of varying selectivity. More selective institutions also impose no additional cumulative cost on society, while compliers actually pay slightly less in additional net tuition at more selective schools.&lt;/p&gt;
&lt;p&gt;Q: How does the commonly used measure of college value-added (mean peer earnings) compare to actual complier returns?
A: Mean peer earnings overpredicts actual value-added for marginal students by a factor of two: compliers attend an institution with $6,700 higher average peer earnings as a result of admission but gain only $3,300 themselves. The measure also overpredicts the earnings return to selectivity by a factor of three: a 100-SAT-point increase in target school selectivity predicts $3,000 higher peer earnings but only a statistically insignificant $900 higher gain in the complier&amp;rsquo;s own earnings.&lt;/p&gt;
&lt;p&gt;Q: How do earnings returns differ by family income?
A: Compliers from low-income families experience significantly smaller earnings gains compared to higher-income compliers. The gap is not explained by differential changes in college quality induced by admission. Instead, low-income compliers gain fewer degrees despite spending more time in college and major in less lucrative fields, consistent with related findings in the literature on family income gaps in degree completion and major choice.&lt;/p&gt;
&lt;p&gt;Q: How do earnings returns differ by gender and by race?
A: Female and male compliers eventually earn similar log earnings and earnings rank gains, but women reach their gains more quickly — likely because men take longer to finish college. White and Asian compliers experience similar earnings gains and BA completion improvements as Black and Hispanic compliers, despite white and Asian students experiencing larger increases in college selectivity and spending per student as a result of admission.&lt;/p&gt;
&lt;p&gt;Q: What is the method for separating intensive- and extensive-margin effects?
A: The two complier types are not directly distinguishable in the data. The author first uses an endogenous but strong stratification variable — having at least one other Texas public university admission offer — to identify some mean potential outcomes for each type. He then imposes an empirically-informed rank assumption to bound the remaining unknown mean potential outcomes, delivering tightly informative upper and lower bounds on each margin&amp;rsquo;s effects without requiring full nonparametric identification. The results show that pooled effects are driven by larger returns for extensive-margin compliers who would not have attended any four-year college, with smaller contributions from intensive-margin compliers shifting between four-year institutions.&lt;/p&gt;
&lt;p&gt;Q: How do this paper&amp;rsquo;s earnings estimates compare to prior studies, and what explains the differences?
A: This paper&amp;rsquo;s 8 percent earnings gain is smaller than the 17–26 percent reported in prior studies (Zimmerman 2014: 22%; Kozakowski 2023: 26%; Smith, Goodman, and Hurwitz 2025: 17%; Bleemer 2024: 21%; Hoekstra 2009: 20%). The differences are likely explained by the much larger educational attainment and institutional quality gains induced by those studies&amp;rsquo; natural experiments: in Zimmerman (2014), enrollment compliers gain roughly three additional years of four-year education versus one year in this paper; in Bleemer (2024), compliers experience roughly $30,000 more in institutional spending per student versus approximately $3,000 in this paper.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions for these results?
A: The results pertain to marginal applicants to Texas public universities (excluding UT-Austin, which uses holistic admission with no detectable SAT/ACT cutoffs) from the 2004–2014 high school graduation cohorts. The identified effects are local average treatment effects for compliers — applicants who would enroll in the target university if and only if they barely crossed its admission cutoff — and do not represent effects for always-takers or infra-marginal students. Earnings are measured only for Texas-based workers covered by the state unemployment insurance system, which captures an estimated 90 percent of the civilian labor force.&lt;/p&gt;
&lt;p&gt;Cutoff complier: An applicant who enrolls in their target university if and only if their SAT/ACT score barely exceeds that university&amp;rsquo;s admission cutoff. Compliers are the population whose behavior — and thus whose treatment effects — are identified by the fuzzy RD design. They comprise approximately 15 percent of marginal applicants and are more disadvantaged than the average public university applicant but broadly comparable to the average high school graduate.&lt;/p&gt;
&lt;p&gt;Extensive versus intensive margin: The extensive margin refers to the contrast between attending any four-year college versus falling back to a two-year community college or no college. The intensive margin refers to the contrast between attending a more selective versus a less selective four-year institution. Approximately half of cutoff compliers are on each margin; the paper treats them as economically distinct parameters requiring separate identification.&lt;/p&gt;
&lt;p&gt;Fuzzy regression discontinuity (RD) design: An identification strategy that uses the discontinuous jump in admission probability at a test score cutoff as an instrument for enrollment, recovering the LATE for compliers via the ratio of the reduced-form discontinuity in outcomes to the first-stage discontinuity in enrollment. &amp;ldquo;Fuzzy&amp;rdquo; refers to the fact that crossing the cutoff changes admission and enrollment probabilities with a discrete jump rather than with certainty.&lt;/p&gt;
&lt;p&gt;Internal rate of return (IRR): The discount rate at which the net present value of an investment equals zero — here, the discount rate equating the discounted stream of earnings benefits to the discounted stream of costs. The paper estimates IRRs separately for students (26 percent), society (16 percent), and the government budget (7 percent), reflecting different cost and benefit definitions from each perspective.&lt;/p&gt;
&lt;p&gt;Rank assumption (bounding method): An empirically-informed assumption about the ordering of mean potential outcomes across latent complier types (extensive vs. intensive margin) that, combined with partial identification from a strong endogenous stratification variable, yields tight upper and lower bounds on each margin&amp;rsquo;s causal effects without requiring full nonparametric identification.&lt;/p&gt;
&lt;p&gt;Net tuition: Gross tuition charges minus grant aid. For the typical marginal complier, gross tuition rises by $4,600 but is nearly fully offset by additional grant aid, yielding approximately zero additional net tuition cost — meaning the private financial cost of attending a public university for marginal students is effectively zero on net, though they take on $5,300 more in student loans to finance room, board, and consumption.&lt;/p&gt;
&lt;p&gt;Sandwich earnings measure: A procedure applied to quarterly state earnings data that retains only quarters with positive earnings sandwiched between other quarters with positive earnings, discarding high-variance transition quarters between employment spells. Annualized by multiplying the quarterly average by four; used to reduce noise from entry and exit transitions in administrative earnings records.&lt;/p&gt;</description></item><item><title>Market Segmentation through Information</title><link>https://macropaperwarehouse.com/papers/market-segmentation-through-information/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/market-segmentation-through-information/</guid><description>&lt;p&gt;This paper asks what market outcomes an information designer — modeled as an internet platform that knows consumers&amp;rsquo; preferences — can achieve by choosing what information to disclose to competing oligopolistic firms who then make personalized price offers. The model features n firms each producing a single differentiated product at zero cost, a continuum of consumers with unit demand and multidimensional valuations (one per product), and a designer who commits to a mapping from consumer types to joint distributions over messages sent to firms before they play a simultaneous pricing game. The designer&amp;rsquo;s objective spans the full range from maximizing producer surplus to maximizing consumer surplus.&lt;/p&gt;
&lt;p&gt;The paper establishes two main results. First, under a necessary and sufficient condition called Aggregate Incentive Compatibility (AIC), the designer can implement full surplus extraction by firms — the producer-optimal outcome — in which every consumer buys her most preferred product at a price exactly equal to her valuation for it, capturing 100% of available surplus for producers. The AIC condition requires, for each firm i and each candidate deviation price p_hat_i, that the infra-marginal losses firm i would bear on its natural customers (those in Ei who value i most) from lowering price to p_hat_i must be weakly greater than the maximum business-stealing profit available from consumers who prefer other products but have valuation for i above p_hat_i. The condition is easier to satisfy when consumer preferences are more polarized, i.e., when consumers have stronger relative preferences for their most-preferred product. When firms offer homogeneous products the condition fails everywhere and no information structure can generate any producer surplus — Bertrand competition drives all profits to zero under any signal structure.&lt;/p&gt;
&lt;p&gt;Second, the paper characterizes the consumer-optimal information structure, which achieves the maximum possible consumer surplus across all equilibria induced by any information structure. The upper bound on consumer surplus is CS* = (total surplus) minus sum_i Pi*_i, where Pi*_i is the profit firm i can guarantee itself by ignoring the designer&amp;rsquo;s signal and setting the best uniform price assuming all rivals price at zero. This bound is tight: the designer can implement it by publicly partitioning consumers into groups by most-preferred product, inducing rival firms to price at marginal cost (zero) for consumers who prefer another firm&amp;rsquo;s product, and then applying the Bergemann-Brooks-Morris (2015) extremal segmentation within each firm&amp;rsquo;s natural customer set to preserve each firm&amp;rsquo;s guarantee profit while achieving efficiency.&lt;/p&gt;
&lt;p&gt;The illustrative two-firm example shows the quantitative stakes concretely. With no information disclosure, firms charge 4/5 and total producer surplus is about 76% of total surplus S*, consumer surplus is just under 10% of S*, and some consumers are excluded. With full disclosure, producer surplus rises to about 81% of S* and consumer surplus to 19%. The producer-optimal information structure (Case 3) achieves 100% of S* as producer surplus by pooling consumers who prefer different products into the same message submarket, giving each firm an incentive to price for its highest-valuing customers and ignore the others. The consumer-optimal information structure (Case 4) brings producer surplus down to about 57% of S* — its guaranteed lower bound — and delivers roughly 43% of S* to consumers, an outcome unattainable by full disclosure alone.&lt;/p&gt;
&lt;p&gt;Both producer-optimal and consumer-optimal outcomes are efficient: all consumers buy their most-preferred product in both cases. The paper further characterizes the full efficient frontier between consumer- and producer-optimal outcomes, showing that mixing the consumer-optimal and full-information structures (or consumer-optimal, full-information, and producer-optimal structures when the latter is implementable) spans every point on the frontier.&lt;/p&gt;
&lt;p&gt;The model assumes firms will price-discriminate if they can, that the designer has full knowledge of consumer types, and that the game is played once. The core results extend to continuous type distributions as shown in Online Appendix B.2. The analysis is restricted to a monopoly platform; competition among platforms is left for future work.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question and why does the two-benchmark comparison used by antitrust authorities miss important possibilities?&lt;/p&gt;
&lt;p&gt;A: The paper asks what market outcomes — combinations of consumer and producer surplus — an information designer (a platform) can achieve by choosing among all possible information structures, not just the two benchmarks of no-information and full-information. Antitrust analysis that compares only those two cases misses a vast middle ground: an intermediary can package information in ways that, for instance, implement perfect collusion (extracting all surplus as producer surplus) while appearing to use privacy-protective technologies, or can intensify competition well beyond the full-information benchmark to benefit consumers.&lt;/p&gt;
&lt;p&gt;Q: What is the producer-optimal information structure and when does it exist?&lt;/p&gt;
&lt;p&gt;A: A producer-optimal information structure is one that induces an equilibrium in which every consumer buys her most-preferred product at a price exactly equal to her valuation — full surplus extraction. It exists if and only if, for every firm i and every candidate deviation price p_hat_i, the Aggregate Incentive Compatibility (AIC) condition holds: the aggregate infra-marginal losses firm i would suffer on its natural customers Ei from lowering price to p_hat_i must be at least as large as the maximum business-stealing profit from consumers outside Ei who have valuation for i weakly above p_hat_i. This is a condition on the distribution of consumer valuations, not on the information structure per se.&lt;/p&gt;
&lt;p&gt;Q: What is the economic mechanism behind the producer-optimal structure — how does pooling consumers implement full surplus extraction?&lt;/p&gt;
&lt;p&gt;A: The designer assigns consumers who prefer product A to the same message submarket as consumers who prefer another product but have a lower valuation for A. Firm A is then price-recommended its highest-valuing customers&amp;rsquo; willingness to pay. The presence of the &amp;ldquo;outside&amp;rdquo; consumers in the same message makes it unprofitable for firm A to deviate downward to capture them, because the infra-marginal loss on the natural customers exceeds the additional revenue. Simultaneously, the rival firm cannot identify and undercut for A&amp;rsquo;s natural customers because the messages do not allow it to distinguish them. The result is that each firm plays a niche strategy, setting price equal to the valuation of its highest-type natural customers and excluding the others from its offer.&lt;/p&gt;
&lt;p&gt;Q: When does polarization of consumer preferences help achieve the producer-optimal outcome?&lt;/p&gt;
&lt;p&gt;A: Proposition 1 states that if a producer-optimal information structure exists under distribution f, it also exists under any distribution f_tilde that is more polarized than f — where more polarized means the mass of consumers who prefer i and have valuation above any threshold for i increases, and the mass of consumers who prefer j but have valuation above that threshold for i decreases. Intuitively, polarization slackens the Firm IC constraints because it reduces the business-stealing temptation: fewer consumers with high cross-product valuations are available for firm i to capture by undercutting. Concrete continuous-distribution examples include: uniform over the unit square (producer-optimal always exists), Hotelling anti-correlated values (exists everywhere), and truncated normal with mean 1/2 — producer-optimal is feasible for all standard deviations sigma &amp;gt; 0.15.&lt;/p&gt;
&lt;p&gt;Q: Why does the producer-optimal outcome fail entirely when products are homogeneous?&lt;/p&gt;
&lt;p&gt;A: Proposition 2 states that when all consumer types have equal valuations across products (the support of f lies on the diagonal of V^n), then for any information structure and any induced equilibrium, every consumer buys at price zero and all firms earn zero profit. The logic extends the standard Bertrand undercutting argument: with homogeneous products, any positive price a firm charges is undercut by a rival who can always profitably steal demand, and this applies to any posterior distribution induced by any signal realization. Even private signals cannot prevent this outcome because no signal realization can give a firm a non-contestable position.&lt;/p&gt;
&lt;p&gt;Q: How is the consumer-optimal information structure constructed, and what is its key economic logic?&lt;/p&gt;
&lt;p&gt;A: Theorem 2 shows the consumer-optimal structure has three layers. First, consumers are partitioned into n groups by most-preferred product (Ei). Second, firms j not equal to i are induced — by publicly revealing which group a consumer belongs to — to set price zero for consumers outside their group, because competing for those consumers is hopeless when their preferred firm is identified. Third, within each Ei, consumers are further partitioned into submarkets using the Bergemann-Brooks-Morris (2015) extremal segmentation applied to residual valuations (theta_i minus the maximum of competing valuations), ensuring firm i earns exactly its guarantee profit Pi*_i. By holding each firm down to its guarantee profit, the residual goes to consumers, maximizing CS.&lt;/p&gt;
&lt;p&gt;Q: What is the guarantee profit Pi*_i and how does it bound consumer surplus?&lt;/p&gt;
&lt;p&gt;A: Pi*&lt;em&gt;i is the maximum profit firm i can achieve by ignoring all designer signals and setting a single uniform price to all consumers, against the worst-case scenario in which all other firms price at zero. Formally, Pi*&lt;em&gt;i = max&lt;/em&gt;{pi} sum&lt;/em&gt;{theta in Ei: theta_i - pi &amp;gt;= max_{j not equal i} theta_j} pi * f(theta). Since firm i can always achieve Pi*_i regardless of the information structure (by simply ignoring signals), no information structure can push firm i&amp;rsquo;s profit below Pi*_i. The sum of these guarantee profits across all firms provides a lower bound on total producer surplus — and therefore an upper bound on consumer surplus — achievable by any information structure.&lt;/p&gt;
&lt;p&gt;Q: In the two-firm numerical example, what is the quantitative comparison across the four cases?&lt;/p&gt;
&lt;p&gt;A: Total available surplus S* = 0.84. Under no information (Case 1): producer surplus approximately 76% of S*, consumer surplus just under 10% of S*, and consumers of types (3/5, 2/5) and (2/5, 3/5) do not trade. Under full disclosure (Case 2): producer surplus approximately 81% of S*, consumer surplus 19% of S*, efficient. Under the producer-optimal structure (Case 3): producer surplus = 100% of S* (all surplus extracted), consumer surplus = 0%, efficient. Under the consumer-optimal structure (Case 4): producer surplus approximately 57% of S*, consumer surplus approximately 43% of S*, efficient. All cases except Case 1 are efficient; the no-information case excludes some consumers from trading.&lt;/p&gt;
&lt;p&gt;Q: Is the full-information disclosure structure consumer-optimal?&lt;/p&gt;
&lt;p&gt;A: Not in general. Proposition 3 states that full information is consumer-optimal if and only if all consumers in Ei have identical residual valuations (theta_i minus their second-best alternative) — a condition that generically fails. When residual valuations within Ei are heterogeneous, the designer can do strictly better for consumers by applying the extremal segmentation within each Ei rather than revealing full information, which would allow firms to price-discriminate on individual residual valuations and extract more surplus.&lt;/p&gt;
&lt;p&gt;Q: Can the designer trace out the entire efficient frontier between consumer- and producer-optimal outcomes?&lt;/p&gt;
&lt;p&gt;A: Yes, under two conditions. First, by mixing the consumer-optimal structure (point A) with the full-information structure (point B) using fractions lambda and 1-lambda respectively, the designer can implement any point on the efficient frontier between A and B. Second, when the producer-optimal outcome (point C) is also implementable, mixing the full-information structure with the producer-optimal structure by applying them to fractions lambda and 1-lambda of the consumer population respectively spans every point between B and C. The key insight is that the AIC condition, if it holds for f, also holds for any rescaled sub-distribution of f (it is scale-invariant), so the producer-optimal sub-problem remains feasible.&lt;/p&gt;
&lt;p&gt;Q: What are the regulatory implications of the analysis?&lt;/p&gt;
&lt;p&gt;A: The paper identifies a fundamental tension: banning information use sacrifices efficiency (some consumers excluded, wrong products purchased), but unrestricted use permits platforms to implement perfect collusion through information design. Critically, the paper shows that privacy-enhancing technologies that pool consumers into cohorts — like Google&amp;rsquo;s Privacy Sandbox — are equally consistent with the producer-optimal (collusive) and consumer-optimal (competitive) structures; the two differ only in the principle by which consumers are grouped. The paper suggests regulators could mandate that consumers in the same cohort share the same most-preferred product and that information be disclosed symmetrically across firms — the defining features of the consumer-optimal structure. This would block the producer-optimal grouping (which mixes consumers with different most-preferred products) while preserving efficiency.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to and extend Bergemann, Brooks, and Morris (2015)?&lt;/p&gt;
&lt;p&gt;A: Bergemann, Brooks, and Morris (2015) characterize achievable consumer and producer surplus outcomes when a designer discloses information to a single monopolist who can price-discriminate. The present paper extends this to oligopoly, where competition between firms creates both additional constraints (firms may undercut each other) and additional instruments (the designer can play firms against each other). The consumer-optimal construction directly applies the BBM (2015) extremal segmentation within each firm&amp;rsquo;s natural customer set Ei, but the outer layer — using public revelation of group membership to induce rival firms to price at zero — is new and arises specifically from the oligopoly setting.&lt;/p&gt;
&lt;p&gt;Information designer: An entity (modeled as a platform) that observes the full joint distribution of consumer valuations over all products and commits, before firms price, to a mapping from consumer types to joint distributions over messages sent to competing firms; the designer can be interpreted as an internet intermediary choosing how to package and share consumer data.&lt;/p&gt;
&lt;p&gt;Aggregate Incentive Compatibility (AIC): The necessary and sufficient condition on the distribution of consumer valuations for the existence of a producer-optimal information structure; for each firm i and each candidate deviation price p_hat_i, the aggregate infra-marginal losses firm i would incur on its natural customers by lowering price to p_hat_i must weakly exceed the maximum revenue firm i could gain by attracting consumers who prefer rival products but have valuation for i above p_hat_i.&lt;/p&gt;
&lt;p&gt;Producer-optimal information structure: An information structure that induces an equilibrium in which every consumer buys her most-preferred product at a price exactly equal to her full valuation for it, extracting 100% of available surplus as producer surplus — the outcome equivalent to the firms&amp;rsquo; fully collusive joint surplus maximum.&lt;/p&gt;
&lt;p&gt;Consumer-optimal information structure: An information structure that achieves the maximum consumer surplus attainable across all equilibria induced by any information structure, holding each firm to its guarantee profit Pi*_i (the best uniform-price profit the firm can secure by ignoring all signals) and allocating all residual surplus to consumers while maintaining allocative efficiency.&lt;/p&gt;
&lt;p&gt;Guarantee profit (Pi*&lt;em&gt;i): The maximum profit firm i can secure unilaterally by ignoring the designer&amp;rsquo;s signal and setting an optimal uniform price, computed against the worst case in which all rival firms price at zero; it equals max&lt;/em&gt;{pi} times the sum of f(theta) over all types in Ei for which theta_i minus pi exceeds all rival valuations.&lt;/p&gt;
&lt;p&gt;Polarization of preferences: A stochastic dominance condition under which, relative to a baseline distribution, the mass of consumers who prefer product i and have high valuations for it increases while the mass of consumers who prefer rival products but have high valuations for i decreases; higher polarization weakens the Firm IC constraints and makes the producer-optimal outcome easier to implement (Proposition 1).&lt;/p&gt;
&lt;p&gt;Separation and Consistency: Two structural properties any producer-optimal information structure must satisfy: Separation requires that the messages firm i sends to different consumers in Ei who have distinct valuations for i are disjoint in support; Consistency requires that every message firm i can send to any consumer type is contained in the union of messages firm i sends to consumers in Ei, preventing firm i from ever inferring that a consumer prefers a rival&amp;rsquo;s product.&lt;/p&gt;</description></item><item><title>Markov-Perfect Equilibria in Differential Games—With an Application to Climate Policy</title><link>https://macropaperwarehouse.com/papers/markov-perfect-equilibria-in-differential-gameswith-an-application-to-climate-policy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/markov-perfect-equilibria-in-differential-gameswith-an-application-to-climate-policy/</guid><description>&lt;p&gt;This paper by Jaakkola and Wagener addresses a long-standing open problem in the theory of differential games: how to make Markov-perfect equilibria (MPE) well-defined when best-response policy functions are generically discontinuous in the state variable. The paper&amp;rsquo;s primary contribution is methodological — it introduces discontinuous Markovian strategies into differential games and proves that, under this extension, (i) payoffs can always be computed and (ii) unique best responses exist for almost all strategy profiles of opponents. The authors then apply this framework to derive the entire set of symmetric MPE in a canonical non-cooperative climate mitigation model (van der Ploeg and de Zeeuw, 1992), finding welfare results that are quantitatively large and policy-relevant.&lt;/p&gt;
&lt;p&gt;The technical difficulty the paper resolves is that discontinuous policy functions can cause the ordinary differential equation governing state dynamics to lack classical solutions, making payoffs undefined. Prior literature responded either by restricting strategies to continuous functions — which rules out many natural best responses and imposes an unjustified constraint on the strategy space — or by allowing discontinuities only in &amp;ldquo;admissible&amp;rdquo; profiles, which makes each player&amp;rsquo;s strategy set depend on opponents&amp;rsquo; choices and thus violates the basic structure of non-cooperative game theory. The authors&amp;rsquo; solution is to adopt Filippov solutions (differential inclusions that convexify dynamics at discontinuities), so that a well-defined state trajectory and payoff exist for every strategy profile, not just admissible ones.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s three main theorems cover existence (Theorem 1), characterization (Theorem 2), and symmetric equilibrium conditions (Theorem 3). Theorem 1 establishes that, given any fixed set of potential jump points, the best-response correspondence maps almost all opponent strategy profiles to a unique Markovian best response — &amp;ldquo;almost all&amp;rdquo; in the sense of prevalence on infinite-dimensional function spaces. Theorem 2 provides necessary and sufficient conditions for a strategy to be a best response: it must satisfy the maximum principle where the value function is differentiable, value discontinuities may only occur at jump points of opponents&amp;rsquo; strategies where the player cannot unilaterally push the state back to the low-stock side, and the value at any such interface must exceed the static optimum. Theorem 3 translates these into conditions for symmetric Nash equilibrium.&lt;/p&gt;
&lt;p&gt;Applied to the van der Ploeg–de Zeeuw climate model — N symmetric countries choosing emissions a_i, with carbon stock x evolving as x-dot = sum(a_i) - delta&lt;em&gt;x, and flow utility u(x, a_i) = a_i - (1/2)a_i^2 - dx — the paper characterizes the complete set of symmetric MPE. The unique continuous globally defined equilibrium (the linear MPE, previously established by Rowat 2007) is shown to be weakly Pareto-dominated by every other MPE with a continuous value function. The best equilibria feature discontinuous strategies that act like stock-conditioned trigger strategies: when the carbon stock falls below a target steady state x&lt;/em&gt;, players respond with a discrete upward jump in emissions to rapidly return the economy to x*; when carbon rises above x*, players increase emissions only gradually, creating a threat of drifting to a higher-pollution steady state that disciplines deviations. In a calibrated example with N=10, delta=0.02, rho=0.02, and damage parameter d=0.5, the linear equilibrium steady state is approximately 2.5 times the first-best level, while the best continuous-value MPE steady state is approximately 1.2 times the first-best level. Choosing the best equilibrium rather than the linear equilibrium closes between 50 and 100 percent of the welfare gap to the first-best outcome, depending on initial conditions. The paper also identifies particularly bad equilibria involving value-function discontinuities — coordination failures in which no single country can unilaterally stop the carbon stock from rising past a threshold — that can yield welfare outcomes worse than the linear equilibrium at high carbon levels.&lt;/p&gt;
&lt;p&gt;The scope of the methodological results covers differential games with a single state variable and strategies that are real-analytic except at finitely many points. Extension to multiple state variables is left for future work. The climate application is restricted to the symmetric linear-quadratic van der Ploeg–de Zeeuw framework, chosen to facilitate comparison with prior literature.&lt;/p&gt;
&lt;p&gt;Q: What is the fundamental technical problem with MPE in differential games that this paper resolves?&lt;/p&gt;
&lt;p&gt;A: In differential games with Markovian strategies, best-response policy functions are generically discontinuous in the state variable. Discontinuous right-hand sides in the state dynamics ODE can prevent existence or uniqueness of classical solutions, making payoffs undefined for some strategy profiles. Prior literature either restricted attention to continuous strategies (causing non-existence of best responses to many profiles) or defined &amp;ldquo;admissible&amp;rdquo; strategy sets that depend on opponents&amp;rsquo; choices (violating non-cooperative game theory structure). This paper resolves both problems for the single-state-variable case.&lt;/p&gt;
&lt;p&gt;Q: How does the paper make payoffs well-defined under discontinuous strategies?&lt;/p&gt;
&lt;p&gt;A: The paper adopts Filippov solutions — differential inclusions that replace the dynamics at a discontinuity point with a convex hull of the left and right limits. At a &amp;ldquo;push-push&amp;rdquo; discontinuity (where dynamics push the state toward the jump point from both sides), the Filippov solution remains at the jump point and flow payoffs are a weighted average of left and right actions. This ensures a well-defined trajectory and payoff for every strategy profile, not just &amp;ldquo;admissible&amp;rdquo; ones, restoring the standard non-cooperative game-theoretic structure.&lt;/p&gt;
&lt;p&gt;Q: What does Theorem 1 establish, and what does &amp;ldquo;almost all&amp;rdquo; mean in this context?&lt;/p&gt;
&lt;p&gt;A: Theorem 1 establishes that, for any fixed collection of jump points, each player has a unique Markovian best response to almost every profile of opponents&amp;rsquo; strategies. &amp;ldquo;Almost all&amp;rdquo; is in the sense of prevalence on infinite-dimensional function spaces (following Hunt, Sauer, and Yorke 1992): the set of profiles for which a unique best response fails to exist is shy (measure-zero analog in infinite dimensions) and nowhere dense. This resolves the long-standing open problem of making MPE well-founded in differential games.&lt;/p&gt;
&lt;p&gt;Q: What are the necessary and sufficient conditions for a best response given by Theorem 2?&lt;/p&gt;
&lt;p&gt;A: A strategy phi_i is the best response to opponents&amp;rsquo; profile if and only if: (i) at all points where the value function is differentiable, the strategy satisfies the maximum principle; (ii) the value function is decreasing in the state (monotonicity); (iii) value discontinuities may occur only at opponents&amp;rsquo; jump points where player i cannot unilaterally move the state back to the low-stock region; (iv) at any such interface, the value must be at least as large as the static optimum u(x, a_i)/rho; and (v) the value is differentiable at push-push steady states. These conditions extend the standard maximum principle with local requirements that restrict which discontinuities are possible.&lt;/p&gt;
&lt;p&gt;Q: What is the van der Ploeg–de Zeeuw model and why is it used here?&lt;/p&gt;
&lt;p&gt;A: The van der Ploeg–de Zeeuw (1992) model has N symmetric countries choosing emissions a_i, with carbon stock evolving as x-dot = sum(a_i) - delta*x, and flow utility u(x, a_i) = a_i - (1/2)a_i^2 - dx. It is linear-quadratic, so a linear MPE exists and is analytically tractable, and prior literature (Dockner and Long 1993; Rowat 2007; Dockner and Wagener 2014) has studied it extensively. The paper uses it as a benchmark to demonstrate that the new methods yield novel and economically important results for even well-understood models.&lt;/p&gt;
&lt;p&gt;Q: What is the linear equilibrium and why does it produce poor welfare outcomes?&lt;/p&gt;
&lt;p&gt;A: The linear equilibrium phi_L(x) = alpha + beta*x, with beta negative, is the unique continuous globally defined MPE (Rowat 2007). In it, emissions decrease with the carbon stock because each player anticipates that opponents will also reduce emissions when carbon is high. This strategic substitutability creates adverse dynamic free-riding: players try to exploit the fact that high carbon stock will cause opponents to cut back, so each has an incentive to emit more when carbon is low. In the calibrated example, the linear equilibrium steady state is approximately 2.5 times the first-best level.&lt;/p&gt;
&lt;p&gt;Q: What do the best equilibria look like, and why do they achieve high welfare?&lt;/p&gt;
&lt;p&gt;A: The best equilibria feature a target steady state x* near the first-best level and a discontinuous upward jump in emissions when carbon falls slightly below x*. This threat rapidly returns any carbon reduction back to x*, eliminating the strategic incentive to free-ride on others&amp;rsquo; reductions. When carbon rises above x*, emissions increase only slightly, causing the economy to drift slowly toward a higher-pollution steady state — the threat of this bad outcome disciplines overshooting. This mechanism is analogous to a trigger strategy but is conditioned on the stock level rather than on past actions, making it compatible with Markovian strategies.&lt;/p&gt;
&lt;p&gt;Q: How large are the welfare gains from the best equilibrium relative to the linear equilibrium?&lt;/p&gt;
&lt;p&gt;A: In the calibrated example with N=10, delta=0.02, rho=0.02, and d=0.5, the best continuous-value MPE steady state is approximately 1.2 times the first-best level, compared to 2.5 times for the linear equilibrium. Choosing the best equilibrium closes between 50 and 100 percent of the welfare gap between the linear equilibrium and the first-best outcome, depending on initial conditions. The paper characterizes this as a quantitatively large, first-order welfare improvement.&lt;/p&gt;
&lt;p&gt;Q: What are &amp;ldquo;coordination failure&amp;rdquo; equilibria and when do they arise?&lt;/p&gt;
&lt;p&gt;A: Coordination failure equilibria feature discontinuities not only in the strategy (emission rate) but also in the value function itself. They arise when no single country can unilaterally prevent the carbon stock from rising past a threshold — formally, when N * a_max &amp;lt; delta * x at the discontinuity point. In such cases, if opponents are emitting heavily, no individual country can stop atmospheric carbon from rising even if it emits nothing, making heavy emission a best response. All players following this logic simultaneously produce a self-fulfilling collapse to high emissions. At high carbon levels these equilibria can yield welfare outcomes worse than the linear equilibrium.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s main policy implication for climate negotiations?&lt;/p&gt;
&lt;p&gt;A: The paper argues that international climate negotiations should be understood as a coordination problem over which of many MPE is played, rather than as bargaining over a limited cooperative surplus in a dynamic prisoners&amp;rsquo; dilemma. Since the best equilibria are self-enforcing (they are Nash equilibria, not cooperative solutions), they do not require external enforcement. The paper suggests effective agreements may involve threshold-based commitments — sharp decarbonisation if a carbon target is met, but acceptance of a substantially higher stabilisation target (e.g., 2.5 degrees C rather than 2 degrees C) if the first target is missed — to create the discontinuous strategic incentives that support good equilibria.&lt;/p&gt;
&lt;p&gt;Q: How does the paper handle the previously identified &amp;ldquo;local MPE&amp;rdquo; that could not be extended to the entire state space?&lt;/p&gt;
&lt;p&gt;A: Prior work (Dockner and Long 1993; Rubio and Casino 2002; Dockner and Wagener 2014) constructed nonlinear equilibria that were only locally defined, and the validity of such equilibria was questioned (Rowat 2007; Bernhard 2024) because they were undefined on the full state space. The present paper&amp;rsquo;s framework allows discontinuous strategies, so these locally defined equilibria can be extended into globally defined, discontinuous MPE. Most previously discovered equilibria are shown to be nested within the larger set of all symmetric MPE identified here.&lt;/p&gt;
&lt;p&gt;Q: What mathematical tools are used to prove the main results?&lt;/p&gt;
&lt;p&gt;A: The proofs rely on the theory of viscosity solutions to Hamilton-Jacobi-Bellman equations (Bardi and Capuzzo-Dolcetta 2008), building on and extending results of Barles, Briani, and Chasseigne (2013, 2014) on optimal control with discontinuous dynamics. A key departure from Barles et al. is that the paper cannot assume controllability of the dynamics near discontinuities without imposing undue restrictions on opponents&amp;rsquo; strategies. The application of these results to a fixed-point condition of the best-response correspondence to construct MPE conditions is described as entirely novel.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions and limitations of the methodological results?&lt;/p&gt;
&lt;p&gt;A: The main results (Theorems 1–3) apply to differential games with a single state variable and strategies that are real-analytic except at finitely many points with one-sided derivatives everywhere. The climate application is further restricted to the symmetric linear-quadratic van der Ploeg–de Zeeuw framework. Extension to multiple state variables is acknowledged as future work. The welfare calibration results are specific to the parameter values N=10, delta=0.02, rho=0.02, d=0.5.&lt;/p&gt;
&lt;p&gt;Markov-perfect equilibrium (MPE): A Nash equilibrium in Markovian strategies, where each player&amp;rsquo;s strategy conditions only on the current state variable and not on the history of play. The paper makes this concept well-founded in differential games by allowing discontinuous strategies, ensuring payoffs can be computed for all strategy profiles and unique best responses exist almost everywhere.&lt;/p&gt;
&lt;p&gt;Filippov solution: A solution concept for ordinary differential equations with discontinuous right-hand sides, which replaces the dynamics at a discontinuity point with a convex hull of the left and right limits. Used in this paper to define well-specified state trajectories and payoffs even when players&amp;rsquo; strategies have jumps, eliminating the need to restrict strategy sets to &amp;ldquo;admissible&amp;rdquo; profiles.&lt;/p&gt;
&lt;p&gt;Discontinuous Markovian strategy: A policy function phi: X -&amp;gt; A that maps the state to an action and is real-analytic except at finitely many points, with well-defined one-sided derivatives everywhere. The key innovation of the paper — allowing such strategies makes differential games well-behaved as standard non-cooperative games while capturing the generically discontinuous nature of optimal policy functions.&lt;/p&gt;
&lt;p&gt;Push-push steady state: A steady state at a discontinuity point of a strategy where the dynamics push the state toward that point from both sides. Under Filippov solutions the state remains at such a point, with flow payoffs being a weighted average of left and right actions. Theorem 2 requires the value function to be differentiable at these points in equilibrium.&lt;/p&gt;
&lt;p&gt;Coordination failure equilibrium: An MPE featuring discontinuities in both the strategy and the value function, arising when no single player can unilaterally move the state across a threshold. At high carbon levels, if opponents emit heavily, individual emission cuts are ineffective; heavy emission becomes a best response for all, sustaining a self-fulfilling high-emission outcome. These equilibria can yield welfare outcomes worse than the linear equilibrium.&lt;/p&gt;
&lt;p&gt;Linear equilibrium: The unique continuous globally defined symmetric MPE in the van der Ploeg–de Zeeuw model, characterized by emissions decreasing linearly in the carbon stock. It involves adverse strategic substitutability — each player reduces emissions in response to high carbon because opponents do likewise — and is weakly Pareto-dominated by every MPE with a continuous value function.&lt;/p&gt;
&lt;p&gt;Skiba point: A state at which the optimal policy is discontinuous because the value function has distinct left and right derivatives, corresponding to the boundary between two basins of attraction with different long-run outcomes. In this paper, the steady state of a best equilibrium is a Skiba-type point: below it, emissions jump up to return rapidly to the target; above it, emissions increase only gradually.&lt;/p&gt;</description></item><item><title>Measuring and Mitigating Racial Disparities in Tax Audits</title><link>https://macropaperwarehouse.com/papers/measuring-and-mitigating-racial-disparities-in-tax-audits/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/measuring-and-mitigating-racial-disparities-in-tax-audits/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Do Black taxpayers face higher IRS audit rates than non-Black taxpayers, despite race-blind audit selection? And if so, why — and what would mitigation look like?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Methodology.&lt;/strong&gt; The authors use comprehensive administrative microdata covering approximately 148 million individual income tax returns and 780,627 operational audits for tax year 2014, supplemented with 71,878 research audits from the IRS National Research Program (NRP) pooled over 2010-2014. Because neither the researchers nor the IRS observe taxpayer race, the authors employ Bayesian Improved First Name Surname Geocoding (BIFSG), which imputes the probability that a taxpayer is Black from first name, surname, and Census Block Group. They develop a novel partial identification strategy: two estimators (a probabilistic estimator and a linear estimator) that, under conditions verified using a matched North Carolina voter-registration dataset containing self-reported race, asymptotically bound the true racial audit disparity from below and above respectively. To address the selective labels problem — underreporting is observable only for audited returns — the authors combine operational audit data with NRP random-sample audits to simulate counterfactual audit selection algorithms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Magnitude of the disparity.&lt;/em&gt; The probabilistic estimator implies a racial audit disparity of 0.81 percentage points; the linear estimator implies 1.34 percentage points. Against a base audit rate of 0.54% for the overall U.S. population in 2014, these bounds imply that Black taxpayers are audited at between 2.9 and 4.7 times the rate of non-Black taxpayers.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Role of the EITC.&lt;/em&gt; The disparity is concentrated among EITC claimants. The estimated disparity within the EITC population is 1.96 to 2.90 percentage points, compared to only 0.10 to 0.18 percentage points among non-EITC claimants. In relative terms, Black EITC claimants are audited at 2.9 to 4.4 times the rate of non-Black EITC claimants. A formal decomposition attributes 70-73% of the overall disparity to higher audit rates among Black EITC claimants, 20-21% to racial differences in EITC claiming rates, and 7-8% to differential audit rates among non-EITC filers. Within EITC claimants, 78.5% of the observed audit disparity is attributable to the Dependent Database (DDb) program.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Source of the disparity — algorithmic objective.&lt;/em&gt; Using counterfactual audit selection algorithms estimated on NRP data, the authors find that allocating EITC audits to maximize detected total underreporting (from any source) would produce audit rates of 0.74% for Black EITC claimants versus 1.63% for non-Black EITC claimants — reversing the disparity. In contrast, the status quo, which prioritizes detecting overclaimed refundable credits, yields 3.00% for Black claimants versus 1.04% for non-Black claimants. The primary driver is a difference in the types of noncompliance that are more prevalent by race: dependent-claiming errors are more common among Black EITC claimants (dependent error rate of 26.6% vs. 16.3% for non-Black), while the highest underreporting via business income underreporting is disproportionately concentrated among non-Black EITC claimants. An algorithm focused on refundable credit overclaims implicitly targets dependent errors and therefore selects Black taxpayers at higher rates.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Prediction model bias.&lt;/em&gt; Even conditional on the refundable-credit objective, the status quo disparity (1.96 p.p.) exceeds the disparity that would arise under an oracle that uses actual rather than predicted refundable credit overclaims (1.08 p.p.), suggesting that prediction errors are unevenly distributed by race. The refundable credit prediction algorithm generates a disparity of 1.75 p.p., approximately 60% larger than the oracle. The authors find suggestive evidence of missingness in birth certificate data (paternal information is disproportionately missing for children claimed on Black taxpayers&amp;rsquo; returns) and differential predictive accuracy in the DDb risk score across race.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Operational consequences.&lt;/em&gt; Switching the objective from refundable credit overclaims to total underreporting would shift the composition of audited returns from predominantly dependent-eligibility issues (80% of refundable credit oracle-selected returns contain a dependent error) toward business income (86% of total-underreporting oracle-selected returns have business income underreporting). EITC returns with substantial business income (gross receipts above $25,000) cost on average $369.70 to audit versus $23.09 for other EITC returns. Holding the audit rate fixed, the switch would raise average examination costs by nearly an order of magnitude, while also increasing detected underreporting (mean adjustment of $22,578 per return under the total underreporting oracle versus $9,595 under the refundable credit oracle).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; Results pertain primarily to tax year 2014. The paper finds similar patterns for tax years 2010, 2012, 2016, and 2018. The analysis covers Black versus non-Black taxpayers; disparities for other racial and ethnic groups are not the focus. The selective labels identification strategy relies on the NRP random-audit sample and the bounding conditions verified in the North Carolina matched data.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-cant-the-disparity-be-attributed-simply-to-black-taxpayers-being-more-likely-to-claim-the-eitc-combined-with-eitc-claimants-facing-higher-audit-rates-generally"&gt;Q1. Why can&amp;rsquo;t the disparity be attributed simply to Black taxpayers being more likely to claim the EITC, combined with EITC claimants facing higher audit rates generally?&lt;/h3&gt;
&lt;p&gt;The authors test this directly by estimating racial audit disparities separately within EITC claimants and non-claimants. If differential EITC claiming rates were the full explanation, the within-EITC disparity would be close to zero. Instead, the disparity among EITC claimants (1.96-2.90 p.p.) is larger in absolute terms than the overall disparity (0.81-1.34 p.p.), indicating that Black EITC claimants face substantially higher audit rates than non-Black EITC claimants even holding EITC claimant status fixed. The formal decomposition attributes 70-73% of the overall disparity to differential audit rates within the EITC claimant population, not to differential claiming rates across the population.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-partial-identification-strategy-work-and-what-are-its-key-identifying-assumptions"&gt;Q2. How does the partial identification strategy work, and what are its key identifying assumptions?&lt;/h3&gt;
&lt;p&gt;The authors derive two estimators of the racial audit disparity that use BIFSG-imputed race probabilities rather than observed race. The probabilistic estimator weights each taxpayer&amp;rsquo;s contribution by their estimated probability of being Black; it is downward-biased when there is a positive residual covariance between audits and true race after conditioning on imputed race (E[Cov(Y,B|b)] &amp;gt; 0). The linear estimator regresses audit status on imputed race probability; it is upward-biased when there is a positive residual covariance between audits and imputed race after conditioning on true race (E[Cov(Y,b|B)] &amp;gt; 0). When both covariance terms are positive, the probabilistic and linear estimates bound the true disparity from below and above. The authors verify both conditions are positive and statistically significant (p &amp;lt; 0.01) in the matched North Carolina dataset, for the full population and the EITC population specifically.&lt;/p&gt;
&lt;h3 id="q3-does-the-racial-audit-disparity-within-eitc-claimants-disappear-when-comparing-taxpayers-with-similar-levels-of-underreporting"&gt;Q3. Does the racial audit disparity within EITC claimants disappear when comparing taxpayers with similar levels of underreporting?&lt;/h3&gt;
&lt;p&gt;No. The authors use NRP data to estimate audit rates by race within each underreporting decile among EITC claimants. Within every decile of the underreporting distribution, the estimated audit rate for Black taxpayers exceeds that for non-Black taxpayers. An oracle algorithm that selects returns in descending order of actual underreporting produces an audit rate of 0.74% for Black EITC claimants and 1.63% for non-Black EITC claimants — the opposite of the status quo pattern (3.00% for Black, 1.04% for non-Black). This rules out total-dollar underreporting as the primary driver of the observed disparity.&lt;/p&gt;
&lt;h3 id="q4-why-does-focusing-audit-selection-on-refundable-credit-overclaims-specifically-lead-to-higher-audit-rates-for-black-taxpayers"&gt;Q4. Why does focusing audit selection on refundable credit overclaims specifically lead to higher audit rates for Black taxpayers?&lt;/h3&gt;
&lt;p&gt;Two mechanisms operate simultaneously. First, EITC eligibility is linked to children, so detecting erroneously claimed dependents generates large refundable credit adjustments. The dependent error rate is higher among Black EITC claimants than non-Black EITC claimants (26.6% vs. 16.3% in the probabilistic estimate, or 30.8% vs. 15.4% in the linear estimate). Second, the highest-dollar noncompliance via underreported business income is disproportionately concentrated among non-Black EITC claimants: among EITC claimants in the top 1% of business income underreporting, the probabilistic estimate shows 0.05% are Black compared to 0.21% non-Black. An algorithm aimed at refundable credit overclaims implicitly targets dependent errors and therefore selects Black taxpayers at higher rates; one aimed at total underreporting would prioritize business income underreporting instead and therefore select non-Black taxpayers at higher rates.&lt;/p&gt;
&lt;h3 id="q5-how-do-the-simulated-algorithms-compare-to-the-actual-irs-algorithms"&gt;Q5. How do the simulated algorithms compare to the actual IRS algorithms?&lt;/h3&gt;
&lt;p&gt;The authors cannot directly replicate the IRS&amp;rsquo;s confidential DDb algorithm, but they provide three pieces of evidence that their refundable credit prediction algorithm is a reasonable proxy. First, public governmental documents describe DDb&amp;rsquo;s stated goal as identifying taxpayers who do not meet refundable credit eligibility requirements. Second, when selecting audits based on predicted refundable credit overclaims using largely the same features available to IRS, the authors generate a disparity (1.75 p.p.) close to the status quo disparity (1.96 p.p.). Third, operational audits of EITC returns are strongly associated with their predicted refundable credit overclaims measure but show a much weaker association with predicted total underreporting.&lt;/p&gt;
&lt;h3 id="q6-what-does-the-status-quo-disparity-exceeding-the-refundable-credit-oracle-disparity-reveal-about-prediction-model-design"&gt;Q6. What does the status quo disparity exceeding the refundable credit oracle disparity reveal about prediction model design?&lt;/h3&gt;
&lt;p&gt;The status quo disparity (1.96 p.p.) is approximately 80% larger than the disparity that would arise if the IRS were perfectly informed about actual refundable credit overclaims and selected accordingly (oracle disparity: 1.08 p.p.). The refundable credit prediction algorithm generates a disparity of 1.75 p.p., approximately 60% larger than the oracle. This gap between the oracle and prediction disparity is consistent with prediction errors being distributed unevenly by race. The authors find that birth certificates of children claimed on Black taxpayers&amp;rsquo; returns are substantially more likely to be missing paternal identity information, which may reduce the predictive accuracy of the DDb model for this population. They provide suggestive evidence that modifying the predictive features used could reduce the disparity without substantially degrading credit overclaim detection.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-downstream-operational-consequences-of-switching-the-algorithmic-objective"&gt;Q7. What are the downstream operational consequences of switching the algorithmic objective?&lt;/h3&gt;
&lt;p&gt;Switching from refundable credit overclaims to total underreporting would shift audited issues from dependent eligibility (80% of refundable credit oracle-selected returns have a dependent error) toward business income (86% of total underreporting oracle-selected returns have business income underreporting). Auditing business income returns is substantially more resource-intensive: $369.70 per return on average for returns with gross receipts above $25,000, versus $23.09 for other EITC returns. Holding the current EITC audit rate fixed, the share of audited returns with substantial business income would rise from 3% to 93%, raising total examination costs by nearly an order of magnitude. However, because total detected underreporting per audited return would also rise substantially (mean of $22,578 vs. $9,595), the increase in detected noncompliance would exceed the increase in audit costs, and the qualitative pattern persists even when accounting for higher per-return costs.&lt;/p&gt;
&lt;h3 id="q8-is-the-disparity-consistent-across-years-and-is-it-driven-by-a-particular-audit-type"&gt;Q8. Is the disparity consistent across years, and is it driven by a particular audit type?&lt;/h3&gt;
&lt;p&gt;The authors find comparable audit disparities for tax years 2010, 2012, 2016, and 2018, confirming the 2014 results are not year-specific. The disparity is concentrated in correspondence audits: the estimated disparity in correspondence audit rates is 0.804-1.328 p.p. for the full population, while the disparity in field/office audit rates is only 0.010-0.016 p.p. The disparity is present in both pre-refund and post-refund audits, though pre-refund audits show a larger disparity even among correspondence audits alone. Among EITC claimants, the correspondence audit channel is nearly entirely responsible for the group-level disparity.&lt;/p&gt;
&lt;h3 id="q9-what-heterogeneity-exists-within-eitc-claimants"&gt;Q9. What heterogeneity exists within EITC claimants?&lt;/h3&gt;
&lt;p&gt;The disparity is especially pronounced among unmarried male EITC claimants with dependents: among this subgroup, the audit rate for Black men exceeds the audit rate for non-Black men by more than 4 percentage points, and both are an order of magnitude above the overall U.S. population audit rate. Disparities are smaller among joint filers, unmarried women, and unmarried men without dependents, though the ratio of Black to non-Black audit rates remains substantial across all subgroups. The concentration of the disparity among unmarried men with dependents is consistent with the role of dependent-claiming errors, which are more likely to arise in family structures characterized by nonmarital cohabitation — a pattern more prevalent among Black Americans due to lower marriage rates.&lt;/p&gt;
&lt;h3 id="q10-can-the-disparity-be-attributed-to-disparate-treatment--ie-race-conscious-selection"&gt;Q10. Can the disparity be attributed to disparate treatment — i.e., race-conscious selection?&lt;/h3&gt;
&lt;p&gt;The authors rule out disparate treatment for the EITC population. The DDb audit selection process for EITC returns is automated (no manual review), and IRS does not use race or geography as an input into audit selection. The disparity is therefore the product of disparate impact: race-neutral selection criteria interact with racially correlated patterns of tax return characteristics to produce differential audit rates. For higher-income non-EITC taxpayers, where audit selection may involve human classifiers, the authors cannot rule out disparate treatment.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Audit Disparity (D).&lt;/strong&gt; Defined in the paper as D = E[Y|B=1] - E[Y|B=0], the difference in audit rates between Black taxpayers (B=1) and non-Black taxpayers (B=0). This is a group-level difference in selection rates, not conditional on any other characteristic, and is the primary estimand throughout.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Probabilistic Disparity Estimator.&lt;/strong&gt; An estimator that calculates group-specific audit rates by weighting each taxpayer&amp;rsquo;s contribution by their BIFSG-imputed probability of being Black (or non-Black). It is shown to be downward-biased when E[Cov(Y,B|b)] &amp;gt; 0, i.e., when there is residual positive association between true race and audits after conditioning on imputed race.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Linear Disparity Estimator.&lt;/strong&gt; An estimator based on regressing audit status (Y) on BIFSG-imputed race probability (b). It is shown to be upward-biased when E[Cov(Y,b|B)] &amp;gt; 0, i.e., when imputed race probability predicts audits even after conditioning on true race. Together, the probabilistic and linear estimators form bounds on the true disparity under conditions verified empirically.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;BIFSG (Bayesian Improved First Name Surname Geocoding).&lt;/strong&gt; A probabilistic race imputation method that uses Bayes rule under a conditional independence assumption (first name, surname, and geography are independent given race) to compute Pr[Black | first name, surname, Census Block Group]. Applied here to all 148 million tax returns; calibrated and validated against matched North Carolina voter registration data with self-reported race.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Selective Labels Problem.&lt;/strong&gt; The problem that noncompliance (underreporting) is observed only for returns selected for audit, not for the full filing population. In this paper it means the IRS cannot directly observe the underreporting distribution for unaudited returns. The authors address this using NRP random-audit data, which allows estimation of the unaudited underreporting distribution and construction of counterfactual selection algorithms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Algorithmic Objective.&lt;/strong&gt; The paper distinguishes between (1) the prediction component of audit selection — which model to use to forecast noncompliance — and (2) the objective component — what type of noncompliance to predict and pursue (overclaimed refundable credits versus total underreporting from any source). The paper finds that the objective, not just prediction error, is an independent driver of the racial audit disparity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dependent Database (DDb) Program.&lt;/strong&gt; The IRS&amp;rsquo;s primary EITC audit selection program, responsible for approximately 75% of audited EITC returns in 2014. DDb flags returns based on rules, heuristics, and proprietary risk scores, with the stated goal of identifying taxpayers who do not meet refundable credit eligibility requirements. Selection through DDb is fully automated, without human classifier review.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;National Research Program (NRP).&lt;/strong&gt; A stratified random sample audit program through which the IRS conducts near-line-by-line examinations of a small fraction of the filing population each year (approximately 2% of audited returns in 2014). The paper pools 71,878 NRP audits from 2010-2014 to identify the distribution of underreporting in the full EITC filing population and to estimate counterfactual selection algorithms.&lt;/p&gt;</description></item><item><title>Merger Effects and Antitrust Enforcement: Evidence from US Consumer Packaged Goods</title><link>https://macropaperwarehouse.com/papers/merger-effects-and-antitrust-enforcement-evidence-from-us-consumer-packaged-goods/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/merger-effects-and-antitrust-enforcement-evidence-from-us-consumer-packaged-goods/</guid><description>&lt;p&gt;This paper by Bhattacharya, Illanes, and Stillerman makes two contributions to the debate over US antitrust enforcement stringency. First, it documents the price, quantity, and assortment effects of a comprehensive set of consummated mergers in US consumer packaged goods (CPG). Second, it develops and estimates a model of agency enforcement decisions to quantify antitrust stringency and simulate counterfactual outcomes under stricter regimes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and scope.&lt;/strong&gt; The analysis covers 129 product markets across 47 transactions in US CPG from 2006 to 2017, using the NielsenIQ Retail Scanner Dataset (covering 35,000–50,000 stores and 2.6–4.5 million UPCs). The sample is restricted to all deals valued at $280 million or more where both the acquirer and target sold products in at least one overlapping product market-DMA. Geographic markets are NielsenIQ designated market areas (DMAs). The sample is defined to avoid selection bias from studying only mergers that attracted press attention or were litigation targets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identification strategy.&lt;/strong&gt; The empirical approach is a before-after event study within geography and product. For each merger, a brand-specific linear time trend is estimated from the 36 months prior to the merger announcement, controlling for UPC-DMA fixed effects, month-of-year fixed effects, input cost indices, and log median household income. Post-merger outcomes (24 months after completion) are measured as deviations from the extrapolated pre-merger trend. The identifying assumption is that secular demand and cost trends are gradual and well-captured by a linear trend. Pre-trend placebo tests show no significant departures from trend in the pre-period, and randomized-date placebos confirm that the linear trend is a better predictor of post-period outcomes under random merger dates than under actual merger dates, supporting the interpretation that observed post-period departures reflect merger effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Price effects.&lt;/strong&gt; The average price effect of consummated CPG mergers is small: across specifications, estimates range from -0.6% to 1.0%, with a baseline mean of 0.3%. However, heterogeneity is substantial. The standard deviation of merger-level price effects is 4.0–7.5 percentage points. In the baseline specification, the first quartile of price effects is -2.1% and the third quartile is 3.7%. Merging and non-merging party price changes are positively correlated (correlation = 0.49), consistent with strategic complementarity. Thirty-six percent of mergers lead both groups to lower prices; 36% lead both groups to raise prices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantity and assortment effects.&lt;/strong&gt; Total quantities fall on average by 0.4–1.0% across specifications, with 60% of mergers producing quantity reductions. Merging parties exhibit a larger average quantity decline of 6.4%. Mergers also lead to a 2.7% average reduction in the number of stores served by merging parties, a 2.2% reduction in the number of brands sold in a DMA by merging parties, and a 3.2% reduction for non-merging parties. Brands with less than 5% of the merged entity&amp;rsquo;s sales are 6 percentage points more likely to be dropped post-merger.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Enforcement model.&lt;/strong&gt; To interpret these outcomes relative to enforcement, the authors develop a model in which the agency receives a noisy signal of a merger&amp;rsquo;s price effect and challenges the merger if the posterior mean exceeds a threshold that is decreasing in deal size. They estimate the model by maximum likelihood using data on enforcement actions (6 mergers receiving remedies, 4 withdrawn under antitrust pressure) and realized price changes. The estimated sales-weighted average threshold is 4.8–6.3%: agencies act as if they challenge CPG mergers only when they expect a price increase exceeding this level. The posterior standard deviation of the agency&amp;rsquo;s assessment is 2.5–3.2 pp (aggregate prices) to 4.1–4.8 pp (merging-party prices).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Counterfactual stringency.&lt;/strong&gt; Tightening the threshold from approximately 6.1% to 2.5% would roughly quadruple the challenge probability (from 0.075 to 0.30), reduce aggregate price changes of consummated mergers by approximately 1.4 pp, and lower the share of allowed anti-competitive mergers from roughly 50% to 35%. Critically, type I errors (blocking pro-competitive mergers) remain negligible at thresholds down to approximately 3%; at 0% threshold only 10% of blocked mergers would be type I errors. The primary cost of tighter enforcement is a significantly larger agency workload, not an increase in blocked pro-competitive mergers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; Results pertain specifically to large CPG mergers (deal size ≥ $280 million) sold through US retail outlets, 2006–2017. Findings on structural presumptions show DHHI and merging share have predictive value for price changes, but structural metrics alone explain less than 10% of the variance in price effects (adjusted R-squared never exceeds 10% even with third-order interactions).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the average price effect of consummated CPG mergers and how should it be interpreted?&lt;/strong&gt;
A: Across specifications, the average price effect is between -0.6% and 1.0%, with a baseline mean of 0.3%. This small average does not imply that enforcement is strict: Carlton (2009) shows that with perfect foresight, the largest observed price change — not the average — would indicate stringency. Because agencies face uncertainty, the distribution of realized price changes reflects both inframarginal approved mergers and the noise in agency forecasts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How large is the heterogeneity in merger price effects?&lt;/strong&gt;
A: The standard deviation of merger-level price effects is 4.0–7.5 percentage points across specifications. In the baseline, the first quartile of price effects is -2.1% and the third quartile is 3.7% for all parties combined. Merging parties specifically show a first quartile of -3.2% and third quartile of 3.7%, meaning a full quarter of mergers raise merging-party prices by more than 3.7%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do merging and non-merging party prices co-move?&lt;/strong&gt;
A: Price changes for merging and non-merging parties are positively correlated (correlation = 0.49, s.e. = 0.08), consistent with strategic complementarity in pricing. Thirty-six percent of mergers lead both groups to lower prices, 36% lead both to raise prices, 13% cause merging parties to lower while non-merging parties raise, and 15% cause the reverse. The timing evidence shows merging-party prices begin changing upon merger completion, with rivals following suit.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What happens to quantities following mergers?&lt;/strong&gt;
A: Total quantities fall on average between 0.4% and 1.0% across specifications, with 60% of mergers producing quantity reductions. Merging parties bear the bulk of quantity adjustment, with an average quantity decline of 6.4% and a standard deviation and interquartile range both around 30 pp. Non-merging party quantity changes are much less variable. The correlation between merging and non-merging party quantity changes is 0.36 (s.e. 0.08), which is positive — at odds with theoretical predictions from demand systems with the &amp;ldquo;type aggregation property&amp;rdquo; (Nocke and Schutz, 2018, 2024), where mergers should produce negatively correlated quantity changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What non-price competitive responses do mergers trigger?&lt;/strong&gt;
A: Merging parties reduce the number of stores they serve by 2.7% on average, though in 38% of mergers store networks expand. Both merging and non-merging parties reduce product portfolios: merging parties drop the number of brands in a DMA by 2.2% on average and non-merging parties by 3.2%. Brands most likely to be dropped are those with less than 5% of the merged entity&amp;rsquo;s sales (6 pp more likely to be dropped), brands in small DMAs, and brands with small DMA shares.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Do the Merger Guidelines&amp;rsquo; structural presumptions (HHI, DHHI, merging share) predict price effects?&lt;/strong&gt;
A: DHHI and merging share have statistically significant but quantitatively modest predictive power. A 100-point increase in average DHHI is associated with a 0.2 pp increase in merging-party price changes and 0.3 pp for non-merging parties. Price effects are significantly larger when merging share exceeds 30%. However, structural metrics alone explain very little variance: adjusted R-squared never exceeds 10% even with third-order interactions of HHI, DHHI, merging share, private label share, and market size. Within-merger, DHHI is positively correlated with local price changes, and markets with DHHI above 200 exhibit significantly higher price effects than those below.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do the authors model antitrust enforcement and identify its stringency?&lt;/strong&gt;
A: The agency observes a noisy signal of a merger&amp;rsquo;s price effect, forms a posterior distribution combining a normally distributed prior (mean X&amp;rsquo;beta, standard deviation sigma_p*) with a normally distributed signal error (standard deviation sigma_epsilon), and challenges the merger if the posterior mean exceeds a threshold that is decreasing in deal size. The model is estimated by maximum likelihood: for approved mergers, the realized price change is observed; for withdrawn/remedied mergers, the posterior mean must have exceeded the threshold. Six mergers (from four deals) received remedies for horizontal market power concerns and four mergers (from two deals) were withdrawn under antitrust pressure, forming the challenged set.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the estimated enforcement threshold and how does it vary across mergers?&lt;/strong&gt;
A: The sales-weighted average threshold is 4.8–6.3% using aggregate price changes and 6.6–7.8% using merging-party price changes. The threshold is lower for larger mergers: a 10% increase in merging-party sales is associated with an approximately 0.06 pp decrease in the threshold. The first quartile of thresholds across mergers is 4.5–5.6% and the third quartile is 5.6–6.9%, reflecting that the agencies apply stricter standards to larger deals.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How accurate are the agencies&amp;rsquo; forecasts of merger price effects?&lt;/strong&gt;
A: Using only the prior (structural characteristics), the agency&amp;rsquo;s accuracy in classifying mergers as anti-competitive versus pro-competitive is 56% (s.e. 3 pp). Adding the signal increases accuracy to 83% (s.e. 9 pp). The correlation between the prior mean and the true price change is 0.29 (s.e. 0.08); the correlation between the posterior mean and the true price change is 0.85 (s.e. 0.15). The posterior standard deviation is 2.5–3.2 pp for aggregate price changes and 4.1–4.8 pp for merging-party price changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What would happen under stricter antitrust enforcement?&lt;/strong&gt;
A: Tightening the average threshold from 6.1% to 2.5% would raise the challenge probability from approximately 0.075 to 0.30 — roughly quadrupling it — and would reduce aggregate price changes of consummated mergers by approximately 1.4 pp (from roughly 0.2% to -1.2%). Moving to a 0% threshold would result in challenges to 57% of mergers, with 60–70% of consummated mergers then causing price decreases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How large are type I and type II errors at the current and counterfactual thresholds?&lt;/strong&gt;
A: At the current threshold (~6.1%), approximately 50% of allowed mergers are type II errors (anti-competitive mergers that should have been challenged). Type I errors (pro-competitive mergers wrongly blocked) are negligible at the current threshold and only become non-trivial starting around a 3% threshold. At a 2.5% threshold, the type II error share falls to 35%; at a 0% threshold, to 16%, while type I errors reach 10% of blocked mergers. The primary trade-off of stricter enforcement is therefore a larger agency workload, not an increase in blocking pro-competitive mergers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What identification strategy is used and how is it validated?&lt;/strong&gt;
A: The strategy is a within-product, within-geography before-after comparison using a brand-specific linear pre-merger trend as the counterfactual. Validation proceeds through three checks: (1) coefficient plots from an extended event study show no significant pre-trends after controlling for the linear trend; (2) a plot of brand trends against estimated price effects shows little explanatory power (statistically significant negative correlation but small magnitude, not consistent with results being driven by trend extrapolation); (3) placebo tests randomizing merger dates within the same markets yield a distribution centered at zero, narrower than the true distribution, and a significantly higher mean squared prediction error in the post-period, confirming that the linear trend is a better predictor under randomly assigned merger dates than under true dates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why do the authors not use alternative control group approaches?&lt;/strong&gt;
A: Non-merging firms in the same market are rejected as controls because they may strategically respond to the merger. Synthetic controls using similar-industry untreated markets are rejected because deals often treat multiple similar markets (ruling out natural donors) and estimates prove sensitive to individual donors. Geographic controls (markets where merging parties have small shares) are rejected because they omit all 39 national mergers, untreated markets are not randomly selected, and regional pricing by non-merging parties could propagate effects into untreated regions, biasing estimates toward zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Merger retrospective.&lt;/strong&gt; In this paper&amp;rsquo;s usage, an ex-post empirical study of the price, quantity, and assortment effects of a consummated merger, using pre-merger trends as the counterfactual, as opposed to forward-looking merger simulation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Enforcement stringency.&lt;/strong&gt; The marginal price increase at which the antitrust agency would expect to challenge a merger. Measured here as the sales-weighted average posterior-mean threshold: the value above which the agency acts as if it would propose a remedy, estimated at 4.8–6.3% for US CPG mergers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Type I error (antitrust).&lt;/strong&gt; The mistake of challenging (blocking) a merger that would have reduced prices (a pro-competitive merger). In the model, this occurs when an adverse signal causes the agency to block a merger whose true price effect is below the threshold.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Type II error (antitrust).&lt;/strong&gt; The mistake of allowing a merger that increases prices (an anti-competitive merger). In the model, this occurs when a favorable signal causes the agency to approve a merger whose true price effect is above the threshold. Estimated at approximately 50% of allowed mergers at the current enforcement threshold.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural presumptions.&lt;/strong&gt; The HHI-based rules in the 2010 and 2023 Merger Guidelines that create a presumption of competitive harm when DHHI exceeds specified thresholds (e.g., DHHI &amp;gt; 200 and post-merger HHI &amp;gt; 2,500 for the &amp;ldquo;red zone&amp;rdquo;). The paper finds DHHI and merging share have statistically significant but low explanatory power (adjusted R-squared below 10%) for actual price changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prior and signal (in the enforcement model).&lt;/strong&gt; The agency&amp;rsquo;s prior is a normal distribution over the merger&amp;rsquo;s true price effect, parameterized by structural characteristics (HHI, DHHI). The signal is a noisy draw centered on the true price effect, capturing information gathered through due diligence (e.g., evidence of efficiencies). The posterior mean — combining prior and signal — determines whether the agency challenges the merger.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Product market-deal pair (merger).&lt;/strong&gt; The unit of observation in the empirical analysis: a specific NielsenIQ product module (e.g., soluble coffee) within a specific acquisition transaction (e.g., a food conglomerate merger). The sample contains 129 such pairs across 47 deals.&lt;/p&gt;</description></item><item><title>Minimum Wages, Efficiency, and Welfare</title><link>https://macropaperwarehouse.com/papers/minimum-wages-efficiency-and-welfare/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/minimum-wages-efficiency-and-welfare/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; Can minimum wages improve welfare through efficiency — by correcting monopsony-driven under-employment — and, if so, by how much? What is the optimal minimum wage, and how much of the welfare gain from a higher minimum wage comes from efficiency versus redistribution?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model and methodology.&lt;/strong&gt; The paper develops a tractable general equilibrium oligopsony model with heterogeneous workers (four types: non-high-school, high-school, college workers, and capital owners) and heterogeneous firms (varying in total factor productivity), embedded in a continuum of local labor markets where firms compete strategically in Cournot fashion. Firms face downward-sloping labor supply curves; their market power generates wages below the marginal revenue product of labor (markdowns). The model is calibrated to US data using the Census Longitudinal Business Database (LBD, 2014), the Bureau of Labor Statistics Current Population Survey (CPS, 2019), and the Survey of Consumer Finances (SCF). Key calibration targets include: average firm size of 22.83 workers (LBD), 29 percent of workers earning below $15/hr (CPS), labor and capital income shares, and household-level earnings and capital income ratios. The model is validated by quantitatively replicating four strands of empirical evidence: (i) reallocation effects of the German minimum wage introduction (Dustmann et al., 2021); (ii) employer spillover responses to Amazon&amp;rsquo;s voluntary $15 minimum wage (Derenoncourt et al., 2021); (iii) wage distribution compression evidence from Brazil (Engbom and Moser, 2021); and (iv) heterogeneous employment effects by market concentration (Azar et al., 2019).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three channels for efficiency gains.&lt;/strong&gt; The model identifies three mechanisms through which a minimum wage can improve efficiency under oligopsony: (1) a &lt;em&gt;direct effect&lt;/em&gt; in which constrained firms with monopsony markdowns increase wages and expand employment toward the competitive level (Region II firms); (2) a &lt;em&gt;spillover effect&lt;/em&gt; in which unconstrained competitor firms narrow their own markdowns in response to constrained firms&amp;rsquo; increased wages and market shares; (3) a &lt;em&gt;reallocation effect&lt;/em&gt; in which employment is shifted away from low-productivity firms (which enter Region III — constrained on labor demand) toward more productive firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings on efficiency versus redistribution.&lt;/strong&gt; Under the $15.12/hr minimum wage that maximizes social welfare under utilitarian weights (population-share weights), less than 5 percent of the welfare gains come from improved efficiency, while more than 95 percent come from redistribution. When the government is additionally given access to budget-neutral lump-sum transfers that fully address redistribution goals, the efficiency-maximizing minimum wage narrows to a range of approximately $7.50–$10.00 per hour, which is robust across social welfare weight specifications. The welfare gains attributable to efficiency alone are approximately 0.16–0.20 percent in consumption-equivalent terms, representing only about 1–2 percent of the welfare gains achievable in an economy with no labor market power at all (which would be 15.26 percent in consumption-equivalent terms under the same conditions with optimal transfers).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why efficiency gains are small.&lt;/strong&gt; Three structural reasons limit efficiency gains: (i) low-productivity firms — which are the firms most affected by a binding minimum wage in Region II — have endogenously narrow markdowns even absent a minimum wage, because they face more elastic labor supply and command small market shares; (ii) the calibrated production function has relatively flat marginal revenue product of labor schedules (decreasing returns parameter α = 0.940), so once firms enter Region III, employment rationing occurs rapidly; (iii) the large, high-productivity firms with the widest markdowns are not materially affected by the minimum wages of their small, low-wage competitors because those competitors have small market shares — making spillovers quantitatively negligible even though the model matches empirical cross-employer wage elasticities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Optimal minimum wages under alternative frameworks.&lt;/strong&gt; Without transfers and under utilitarian weights, the optimal minimum wage is $15.12. Without transfers but under Negishi weights (which rationalize the observed competitive equilibrium and load approximately 62 percent of weight on college workers and owners versus their 35 percent population share), the optimal is $6.97. Under a 97 percent weight on high-school graduates, the optimal rises to $18.32. With optimal lump-sum transfers, the optimal collapses to $7.76–$10.11 regardless of social welfare weights — a range robust across Frisch elasticity variants (ϕ ∈ {0.30, 0.62, 0.86}), regional decompositions (low, medium, and high income US states), short-run capital-fixed scenarios (where the optimum declines by approximately $1 under utilitarian weights), and the removal of household heterogeneity entirely (which yields an optimum of $7.74).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Distributional proxies versus welfare.&lt;/strong&gt; Wage inequality (college–non-college log wage premium, cross-sectional variance of log wages) and the labor income share are monotonically improving as the minimum wage rises, even as welfare is hump-shaped and eventually declining. A rise in the minimum wage from $7.50 to $15 reduces the college–non-college log wage premium from 0.53 to 0.43 (roughly one-fifth), reduces the cross-sectional variance of log wages by nearly half, and raises the aggregate labor income share by approximately 3 percentage points — all while welfare (under utilitarian weights with no transfers) reaches its maximum at $15.12 and then declines. These standard proxies therefore do not reliably indicate welfare.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; All results are long-run steady-state comparisons unless otherwise noted. Results assume no price passthrough and a unit elasticity of substitution between capital and labor. The paper abstracts from capital–labor substitution responses and occupational choice. The redistribution channel quantified here is specific to the utilitarian welfare criterion and to the existing distribution of capital and profit income, in which owners (6 percent of households) earn 92 percent of dividends.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-three-regions-of-firm-behavior-in-response-to-a-binding-minimum-wage-and-what-are-their-efficiency-implications"&gt;Q1. What are the three regions of firm behavior in response to a binding minimum wage, and what are their efficiency implications?&lt;/h3&gt;
&lt;p&gt;A: A firm can be in one of three regions. In Region I the minimum wage is not binding: the firm pays its optimal monopsony wage and employment is inelastically below the competitive level. In Region II the minimum wage binds and exceeds the firm&amp;rsquo;s optimal monopsony wage, but labor supply at the minimum wage still falls short of labor demand: employment and efficiency improve as the shadow markdown narrows. In Region III the minimum wage exceeds the competitive wage, so unconstrained labor supply would exceed demand: the firm rations employment and the rationing constraint binds, reducing efficiency. At the boundary of Region II and Region III, the shadow markdown equals one and the firm is at its efficient employment level. Only a firm-specific minimum wage targeting each firm&amp;rsquo;s competitive wage could deliver economy-wide efficiency.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-paper-define-and-use-shadow-wages-to-characterize-equilibrium"&gt;Q2. How does the paper define and use &amp;ldquo;shadow wages&amp;rdquo; to characterize equilibrium?&lt;/h3&gt;
&lt;p&gt;A: The shadow wage for a firm is the effective wage that rationalizes equilibrium employment given rationing constraints. Formally, when a firm rations employment (Region III), households act as if facing a shadow wage equal to the actual minimum wage multiplied by a rationing factor p &amp;lt; 1 (the Lagrange multiplier on the rationing constraint, normalized as a fraction). Shadow wages aggregate across firms into market- and type-level shadow wages via CES aggregation. The key insight is that shadow wages, not observed wages, are allocative: aggregate labor supply for each worker type is determined by the type-level shadow wage, not by the minimum wage that firms actually pay. This allows the paper to express aggregate efficiency via two wedges — the aggregate shadow markdown (capturing average market power) and a misallocation term — without tracking all firm-specific constraints individually.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-two-aggregate-efficiency-wedges-and-how-do-they-behave-as-the-minimum-wage-rises"&gt;Q3. What are the two aggregate efficiency wedges and how do they behave as the minimum wage rises?&lt;/h3&gt;
&lt;p&gt;A: The two wedges are: (i) the aggregate shadow markdown µ̃, which is a productivity-weighted average of firm-level shadow markdowns and measures the extent to which aggregate wages fall short of marginal revenue products; and (ii) the misallocation term ω, which measures whether employment is allocated toward more productive firms and equals one when all shadow markdowns are identical. As the minimum wage rises from zero, µ̃ initially narrows (improving efficiency) because firms in Region II expand toward their competitive employment level and constrained firms&amp;rsquo; market shares rise, tightening the residual labor supply of unconstrained competitors and narrowing their markdowns. But as the minimum wage rises further, Region III rationing causes shadow markdowns to widen rapidly — first for low-productivity firms and then progressively for more productive ones — so µ̃ turns back downward. The misallocation term ω first improves as low-productivity firms are pushed out, but then worsens because rationing at intermediate-productivity firms redirects employment from high- to medium-productivity firms.&lt;/p&gt;
&lt;h3 id="q4-what-does-the-model-validation-exercise-on-the-german-minimum-wage-dlsub-2021-show"&gt;Q4. What does the model validation exercise on the German minimum wage (DLSUB 2021) show?&lt;/h3&gt;
&lt;p&gt;A: The paper calibrates the model to the German context by setting a minimum wage of $8.95/hr equivalent to 48 percent of the pre-reform median wage — matching Germany&amp;rsquo;s 8.50 euro introduction in 2015, where 15 percent of workers earned below the threshold. The model produces employment effects that are slightly positive (consistent with empirical findings of no disemployment), average wage increases consistent with both constrained and unconstrained firms raising wages, a negative elasticity of the number of operating firms with respect to minimum wage exposure (correctly signed, moderately smaller than data), and a positive elasticity of average firm size with respect to exposure (slightly larger than the data). The reallocation direction — small unproductive firms shrinking and workers moving to larger, more productive firms — matches the data qualitatively and within the range of data estimates across specifications.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-amazon-spillover-replication-dnwt-2021-show-and-what-does-it-imply-about-the-minimum-wage-spillover-channel"&gt;Q5. What does the Amazon spillover replication (DNWT 2021) show, and what does it imply about the minimum wage spillover channel?&lt;/h3&gt;
&lt;p&gt;A: Derenoncourt et al. (2021) estimate a cross-employer wage elasticity of 0.26: when Amazon raised wages by approximately 18.1 percent, competitors raised wages by 4.7 percent on average. The model replicates this by treating Amazon as the largest (or second-largest) firm in each market, exogenously narrowing its markdown by a fraction ζ calibrated to deliver an 18.1 percent wage increase. Competitors in the model raise wages through the strategic interaction mechanism: Amazon&amp;rsquo;s higher wage and market share tightens competitors&amp;rsquo; residual supply curves, inducing them to narrow their own markdowns. The model matches the 0.26 cross-employer elasticity when Amazon is the largest firm in markets with at least 36 competitors, or the second-largest in markets with at least 12. Critically, the authors note that this empirical evidence concerns responses to a &lt;em&gt;large&lt;/em&gt; firm raising wages; for minimum wages the question is whether &lt;em&gt;large&lt;/em&gt; firms respond to their small wage competitors, which the model shows they do not substantially, because small firms have negligible market shares.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-separate-efficiency-from-redistribution-and-what-is-the-key-methodological-innovation"&gt;Q6. How does the paper separate efficiency from redistribution, and what is the key methodological innovation?&lt;/h3&gt;
&lt;p&gt;A: The paper gives the government access to budget-neutral, unrestricted lump-sum transfers across households in addition to the minimum wage. With transfers available, the government can use them to meet any redistributive objective encoded in arbitrary social welfare weights. Whatever is left for the minimum wage to do must be purely efficiency-improving. The paper shows (via aggregation theorems) that optimal lump-sum transfers can be computed in closed form for any social welfare weights, and that the social welfare maximizing allocation subject to transfers can be decentralized by transfers that sum to zero across households. Under this framework, the efficiency-maximizing minimum wage lies between $7.50 and $10.00 per hour regardless of whether utilitarian, Negishi, or 97 percent high-school-weighted social welfare functions are used — collapsing the original $0–$31 range to a tight interval.&lt;/p&gt;
&lt;h3 id="q7-how-are-negishi-weights-computed-and-why-are-they-important-for-interpreting-the-results"&gt;Q7. How are Negishi weights computed, and why are they important for interpreting the results?&lt;/h3&gt;
&lt;p&gt;A: The Negishi weights are the social welfare weights under which a planner would choose the observed competitive equilibrium with zero lump-sum transfers. They are computed by inverting the planner&amp;rsquo;s first-order conditions: for the competitive equilibrium to be optimal under some set of weights, the implied consumption ratios must match observed data. The calibrated Negishi weights assign a combined weight of approximately 62 percent to college workers and owners, who constitute only 35 percent of the population. This means the competitive equilibrium is disproportionately aligned with higher-income households. A utilitarian planner, which weights households by population shares, therefore sees large scope for redistribution toward non-college workers — which is exactly why the utilitarian-optimal minimum wage is $15.12 and why 94 percent of its welfare gains come from redistribution rather than efficiency.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-quantitative-welfare-gains-from-the-efficiency-maximizing-minimum-wage-and-how-small-are-they-relative-to-the-potential-gains-from-eliminating-monopsony"&gt;Q8. What are the quantitative welfare gains from the efficiency-maximizing minimum wage, and how small are they relative to the potential gains from eliminating monopsony?&lt;/h3&gt;
&lt;p&gt;A: With optimal lump-sum transfers, the welfare gains from the efficiency-maximizing minimum wage are approximately 0.16–0.20 percent in consumption-equivalent terms, robust across social welfare weight specifications, Frisch elasticity variations, and regional decompositions. The welfare gains associated with an economy in which all firms&amp;rsquo; markdowns are set to one (no labor market power at all), also evaluated with optimal transfers, are 15.26 percent in consumption-equivalent terms. The efficiency-maximizing minimum wage therefore recovers approximately 1–2 percent of the potential welfare gains from eliminating monopsony. Equivalently, the efficiency gains correspond to roughly a 0.1 percent increase in TFP. These gains are small despite the model matching all empirical evidence on the channels through which efficiency gains could occur.&lt;/p&gt;
&lt;h3 id="q9-how-do-employment-effects-of-minimum-wages-vary-by-market-concentration-and-why"&gt;Q9. How do employment effects of minimum wages vary by market concentration, and why?&lt;/h3&gt;
&lt;p&gt;A: In concentrated markets (upper tercile of HHI), firms have larger monopsony markdowns, so a binding minimum wage pushes them into Region II — where employment expands — over a wider range of minimum wage values before entering Region III. This produces large, positive employment effects in concentrated markets. In less concentrated markets, firms already have narrow markdowns (they are closer to competitive), so even small minimum wage increases push them into Region III, where employment contracts. The model replicates the statistically significant positive effects in high-concentration markets and negative effects in low-concentration markets documented by Azar et al. (2019), for initial minimum wages below approximately $8/hr. At higher initial minimum wages, however, even high-concentration markets exhibit negative employment effects as more firms enter Region III.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-robustness-exercise-for-mississippi-reveal"&gt;Q10. What does the robustness exercise for Mississippi reveal?&lt;/h3&gt;
&lt;p&gt;A: Mississippi has the lowest per capita income in the US, and a $15 minimum wage would bind for 41.3 percent of its workers (versus 29.4 percent nationally). Despite this, the model finds that Mississippi would benefit from a $15 federal minimum wage under utilitarian weights, and the Mississippi-specific optimal minimum wage is $14.89 — nearly identical to the national optimum. The reason is an offsetting compositional effect: while Mississippi has lower average wages (pushing toward a lower optimal), it has a larger share of high-school graduates (63 percent versus 52.8 percent nationally) who prefer higher minimum wages (around $17 in the model). These two forces wash out, producing a stable optimal close to the national figure.&lt;/p&gt;
&lt;h3 id="q11-what-happens-to-common-empirical-proxies-for-inequality-and-worker-power-as-the-minimum-wage-rises"&gt;Q11. What happens to common empirical proxies for inequality and worker power as the minimum wage rises?&lt;/h3&gt;
&lt;p&gt;A: The college–non-college log wage premium declines from 0.53 to 0.43 (a fall of roughly one-fifth) as the minimum wage rises from $7.50 to $15. The cross-sectional variance of log wages falls by nearly half over this range, driven equally by declining within- and between-type inequality. The aggregate labor income share rises by approximately 3 percentage points, and the share of output created in non-high-school jobs paid to non-high-school workers rises by 7 percentage points. All of these proxies are monotonically improving in the minimum wage throughout, even as aggregate welfare under the model&amp;rsquo;s social welfare function is hump-shaped and declining past the optimum. The paper concludes that observations of declining inequality or a rising labor share are consistent with falling welfare, so these proxies cannot serve as reliable welfare indicators.&lt;/p&gt;
&lt;h3 id="q12-how-does-the-short-run-fixed-capital-analysis-differ-from-the-long-run-baseline"&gt;Q12. How does the short-run (fixed-capital) analysis differ from the long-run baseline?&lt;/h3&gt;
&lt;p&gt;A: In the short run, capital at each firm is fixed at the type-specific level chosen under a zero minimum wage. This creates sharper decreasing returns in labor (parameter γα rather than α̃), overhead costs that can make operation unprofitable, and a narrower range of minimum wages over which firms remain in Region II. The result is that firms in the short run enter Region III at lower minimum wages than in the long run, limiting the range of efficiency gains. Quantitatively, the efficiency-maximizing optimal minimum wage declines by approximately $1 under utilitarian weights (from about $10 to about $9 in the short-run exercise) and by only about $0.20 under Negishi weights. The robustness conclusion is that the difference between short- and long-run optimal minimum wages is modest, and the main finding that efficiency gains are small is preserved.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Shadow wage (w̃ᵢⱼ):&lt;/strong&gt; The effective wage that rationalizes a firm&amp;rsquo;s equilibrium employment in the presence of a minimum wage. When labor is rationed at firm ij (Region III), the shadow wage equals the actual minimum wage multiplied by a rationing factor pᵢⱼ &amp;lt; 1, where pᵢⱼ is derived from the Lagrange multiplier on the household&amp;rsquo;s rationing constraint. The shadow wage is allocative — it determines labor supply decisions — while the observed minimum wage wage is not. When the rationing constraint is slack (Regions I and II), the shadow wage coincides with the observed wage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Shadow markdown (µ̃ᵢⱼ):&lt;/strong&gt; The ratio of a firm&amp;rsquo;s shadow wage to its marginal revenue product of labor. In Region I (unconstrained), this equals the standard monopsony markdown. In Region II (constrained, on the labor supply curve), the shadow markdown narrows as the minimum wage increases, moving the firm toward its efficient employment level. In Region III (constrained, on the labor demand curve), the shadow markdown equals the rationing multiplier pᵢⱼ and widens, reflecting efficiency losses from rationing. An aggregate shadow markdown µ̃ is computed as a productivity-weighted average of firm-level shadow markdowns across all firms in the economy.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Misallocation wedge (ω):&lt;/strong&gt; A productivity-weighted measure of how well employment is allocated across firms. In an efficient allocation with identical shadow markdowns, ω = 1. When high-productivity firms have wider markdowns than low-productivity firms (the baseline oligopsony outcome), ω &amp;lt; 1 because employment is directed away from productive firms. A minimum wage can improve ω by shrinking low-productivity firms but worsens it when high-productivity firms enter Region III and are over-rationed relative to medium-productivity firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Oligopsony with Cournot competition:&lt;/strong&gt; The specific form of labor market power in this model. In each local labor market (defined as a NAICS 3-digit industry × commuting zone cell), a finite number of firms compete strategically in employment quantities, taking their competitors&amp;rsquo; employment levels as given (Cournot assumption). Each firm has an upward-sloping labor supply curve derived from nested CES household preferences, and exercises a markdown on the marginal revenue product of labor. This differs from monopsony (one firm) or perfect competition (infinitely many firms), and generates both direct effects and spillover effects of minimum wages.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negishi weights:&lt;/strong&gt; The vector of social welfare weights under which the observed competitive equilibrium allocation would be the solution to a social planner&amp;rsquo;s problem with zero lump-sum transfers. In this model, the calibrated Negishi weights assign roughly 62 percent combined weight to college workers and owners (who constitute only 35 percent of the population), reflecting the fact that the market equilibrium allocates a disproportionate share of consumption to high-income households. The Negishi weights are used both to identify the gap between market outcomes and utilitarian objectives (motivating redistribution) and as one alternative normative benchmark.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Efficiency-maximizing minimum wage:&lt;/strong&gt; The minimum wage that maximizes social welfare when the government additionally has access to budget-neutral lump-sum transfers across households. Because transfers can be optimized to handle any redistributive objective encoded in any arbitrary social welfare weights, the minimum wage under this framework serves solely to improve productive efficiency. In the calibrated model, the efficiency-maximizing minimum wage is approximately $7.50–$10.00 per hour, robust to social welfare weight specifications, Frisch elasticity variations (ϕ ∈ {0.30, 0.86}), and regional income differences.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rationing constraint (n̄ᵢⱼₖ):&lt;/strong&gt; A firm-specific, type-specific upper bound on the labor a household may supply to a firm in equilibrium. These constraints are taken as given by households and determined in equilibrium by firms&amp;rsquo; labor demand decisions. When the minimum wage is above the firm&amp;rsquo;s competitive wage (Region III), the firm&amp;rsquo;s labor demand is less than what households would want to supply at that wage, so the rationing constraint binds. The binding rationing constraint generates the shadow wage discount (pᵢⱼ &amp;lt; 1) and is the mechanism by which high minimum wages reduce efficiency in the model.&lt;/p&gt;</description></item><item><title>Mis(sed) Diagnosis: Physician Decision Making and ADHD</title><link>https://macropaperwarehouse.com/papers/missed-diagnosis-physician-decision-making-and-adhd/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/missed-diagnosis-physician-decision-making-and-adhd/</guid><description>&lt;p&gt;This paper develops and estimates a structural model of ADHD diagnosis to decompose the mechanisms driving the observed 2.3:1 male-to-female diagnostic difference in the United States. The research question is: to what extent does the large gender gap in ADHD diagnosis reflect true differences in symptom prevalence, versus patient-side utilization costs, versus physician decision-making under uncertainty? The setting is particularly well-suited to this question because DSM-V diagnostic guidelines for ADHD are explicitly gender-neutral, making any gender difference in physician thresholds a detectable deviation from uniform clinical rules.&lt;/p&gt;
&lt;p&gt;The data come from de-identified electronic health records from a large Arizona healthcare system covering January 2014 through September 2017. The sample encompasses 36,193 unique encounters for approximately 11,070 pediatric patients. The raw male-to-female diagnostic ratio in the data is 2.32:1 (7.2% of males vs. 3.1% of females receive a clinical ADHD diagnosis). This gap persists after controlling for demographics, general healthcare utilization, and mental health utilization in reduced-form regressions, motivating the structural approach.&lt;/p&gt;
&lt;p&gt;Because two key variables — whether a patient received a behavioral assessment (Qi) and the ADHD match signal observed by the physician (xi) — are not directly recorded in the EHR, the author constructs them from clinical doctor note text. A random forest machine learning classifier trained on labeled appointments predicts behavioral assessment take-up for unlabeled encounters; approximately 20.8% of children are predicted to have received a behavioral assessment (23.2% of males vs. 18.3% of females). The ADHD match signal is constructed via an adjusted Bag-of-Words cosine similarity measure comparing each patient&amp;rsquo;s aggregated note text to the DSM-V symptom list, rescaled to [0,1]. The average signal is 0.319 overall, with males averaging 0.326 and females 0.311.&lt;/p&gt;
&lt;p&gt;The structural model has three stages. First, patients/caregivers decide whether to schedule a behavioral assessment, a function of underlying latent ADHD risk (vi) and mental healthcare utilization costs (ci). Second, conditional on assessment, the physician receives a noisy signal of vi and updates beliefs via Bayesian learning; signal quality ρ governs diagnostic uncertainty. Third, the physician diagnoses ADHD if posterior risk exceeds a gender-specific diagnostic threshold τ. Population mean ADHD risk (μ) is identified using regression-adjusted initial primary care provider referral rates as a quasi-exogenous cost-shifter — patients of high-referral-rate providers select into assessment less selectively, so their observed signals approach population mean risk. This extrapolation approach follows Arnold et al. (2022).&lt;/p&gt;
&lt;p&gt;The structural parameter estimates reveal that male and female children have similar but slightly different mean ADHD risk (μm = 0.290 vs. μf = 0.262) and similar mean utilization costs (cm = 0.116 vs. cf = 0.109). The most striking differences are in physician parameters: signal quality is lower for male patients (ρm = 0.479 vs. ρf = 0.552), indicating higher diagnostic uncertainty for boys; and diagnostic thresholds are substantially lower for male patients (τm = 0.257 vs. τf = 0.312), meaning physicians are willing to diagnose ADHD in boys with lower posterior risk.&lt;/p&gt;
&lt;p&gt;Counterfactual decomposition simulations attribute approximately 20–25% of the 2.32:1 diagnostic gap to underlying differences in ADHD risk, approximately 20% to differences in selection into behavioral assessments, and the remaining majority — approximately 55–60% — to physician decision-making. Within physician decision-making, differences in diagnostic thresholds alone account for roughly two-thirds of the overall diagnostic gap.&lt;/p&gt;
&lt;p&gt;The paper offers economic rationales for why gender-specific thresholds may be consistent with physician rationality despite uniform guidelines: higher diagnostic uncertainty for boys justifies lower thresholds under Bayesian updating; hyperactive/impulsive symptoms predominant in boys impose larger classroom externalities (Aizer, 2008); and female patients show higher rates of internalizing co-morbidities (anxiety, depression) that may reduce the marginal benefit of an additional ADHD diagnosis. A type-specific threshold extension finds that for male patients the threshold for hyperactive/impulsive symptoms is significantly lower than for inattentive symptoms, consistent with salience of externally disruptive behaviors. These rationalizations do not vindicate the gap as fully guideline-consistent, but suggest physicians may be responding to real heterogeneity in external costs and co-morbidity patterns.&lt;/p&gt;
&lt;p&gt;Q: What is the main research question and why is ADHD a useful setting?
A: The paper asks what mechanisms produce the 2.3:1 male-to-female ADHD diagnostic difference: true symptom prevalence, patient utilization costs, or physician decision-making. ADHD is well-suited because (1) clinical guidelines (DSM-V) are explicitly gender-neutral and require the same symptom count threshold regardless of sex; (2) diagnosis is based on subjective behavioral assessment rather than objective testing, creating substantial physician discretion; and (3) both missed and excess diagnosis carry meaningful costs — missed diagnosis limits educational accommodations; excess diagnosis exposes children to Schedule II controlled substances.&lt;/p&gt;
&lt;p&gt;Q: What data does the paper use and what are the key descriptive facts?
A: The data are de-identified electronic health records from a large Arizona healthcare system, 2014–2017, covering 36,193 encounters for 11,070 pediatric patients aged 5 and above. Overall ADHD diagnosis rate is 5.2%, with males at 7.2% and females at 3.1%, a 2.32:1 ratio that matches national levels. Approximately 49.5% of the sample is Hispanic, which the author notes contributes to a below-national-average overall diagnosis rate. The gender diagnostic gap persists even after controlling for demographics, general healthcare utilization, and mental health utilization in reduced-form regressions.&lt;/p&gt;
&lt;p&gt;Q: How does the paper construct the behavioral assessment indicator (Qi) and the ADHD match signal (xi)?
A: Qi is constructed using a random forest classifier trained on doctor notes from appointments where assessment status is known with near-certainty (ADHD diagnosis or DSM-V comorbid diagnosis = positive; non-mental-health diagnosis code for patients with no mental health history = negative). The classifier uses 41 features including note length and top-20 word frequencies for each label class. xi is constructed via an adjusted Bag-of-Words cosine similarity between each patient&amp;rsquo;s combined behavioral assessment notes and the DSM-V symptom list, separately for inattentive and hyperactive/impulsive sub-types, taking xi = max{xi1, xi2}. The average xi is 0.319 (males 0.326, females 0.311) in the behavioral assessment subsample.&lt;/p&gt;
&lt;p&gt;Q: What is the identification strategy for recovering population mean ADHD risk (μ)?
A: Because xi is observed only for endogenously selected patients, the observed sample mean overestimates population mean risk. The author uses regression-adjusted referral rates of each patient&amp;rsquo;s initial primary care provider (IPCP) as a quasi-exogenous cost-shifter satisfying (a) relevance — IPCP referral intensity lowers patient scheduling costs — and (b) independence from patient ADHD risk vi, since IPCPs are typically chosen before behavioral symptoms develop and only 28% of IPCPs in the sample ever diagnose ADHD themselves. Population mean risk is then recovered by extrapolating the relationship between IPCP referral propensity and average observed xi to propensity = 1, following Arnold et al. (2022). The maximum observed IPCP referral propensity is only about 0.75, so the estimate requires extrapolation beyond the observed support.&lt;/p&gt;
&lt;p&gt;Q: What are the estimated structural parameters and what do they imply?
A: Mean ADHD risk is μm = 0.290 vs. μf = 0.262 — males have modestly higher underlying risk. Mean utilization costs are cm = 0.116 vs. cf = 0.109 — nearly identical across genders. Signal quality (diagnostic certainty) is lower for males: ρm = 0.479 vs. ρf = 0.552, indicating physicians face more diagnostic uncertainty when assessing boys. Most importantly, diagnostic thresholds are lower for males: τm = 0.257 vs. τf = 0.312, meaning physicians diagnose ADHD in boys at a lower required posterior risk level, consistent with viewing missed diagnosis as relatively more costly for male patients.&lt;/p&gt;
&lt;p&gt;Q: How much of the 2.32:1 diagnostic gap can be attributed to each mechanism?
A: Counterfactual simulations decompose the gap as follows: differences in underlying ADHD risk distribution account for approximately 20–25% of the diagnostic difference; differences in selection into behavioral assessments (utilization costs operating through assessment rates) account for approximately 20%; and physician decision-making differences account for the remaining majority, approximately 55–60%. Within physician factors, differences in diagnostic thresholds (τm &amp;lt; τf) are the single largest contributor, explaining roughly two-thirds of the overall male/female diagnostic gap.&lt;/p&gt;
&lt;p&gt;Q: What do the type-specific threshold estimates reveal?
A: When the baseline model is extended to allow separate diagnostic thresholds for inattentive vs. hyperactive/impulsive symptom sub-types, male patients show significantly lower thresholds for hyperactive/impulsive symptoms relative to inattentive symptoms (τ^HI_m &amp;lt; τ^Inatt_m). This is consistent with the hypothesis that more externally salient and disruptive symptoms carry larger classroom externalities, which physicians may implicitly factor into diagnosis decisions (following Aizer, 2008). For female patients, the threshold differences across symptom types are smaller and less statistically significant.&lt;/p&gt;
&lt;p&gt;Q: What economic rationales does the paper offer for gender-specific diagnostic thresholds despite uniform guidelines?
A: Three mechanisms are identified. First, higher diagnostic uncertainty for males (lower ρm) implies that under symmetric costs, Bayesian-rational physicians should set lower thresholds when the signal is noisier — this alone partially rationalizes the threshold gap. Second, hyperactive/impulsive symptoms predominant in boys impose greater externalities on classroom peers (Aizer, 2008), increasing the social benefit of diagnosis for boys on the margin. Third, females show substantially higher rates of co-morbid internalizing conditions (anxiety, depression) whose treatment may mitigate ADHD-related behaviors or whose interaction with stimulant medication makes the marginal ADHD diagnosis less beneficial for girls (Currie et al., 2014). These factors together suggest physicians may be responding to genuine heterogeneity in net diagnosis benefits, even if their behavior deviates from gender-neutral clinical guidelines.&lt;/p&gt;
&lt;p&gt;Q: What share of the 2.3:1 national diagnostic gap is consistent with genuine symptom prevalence differences?
A: Simulations indicate that only about 20–25% of the 2.32:1 male/female diagnostic difference can be explained by the underlying difference in ADHD risk distributions. The majority — roughly 75–80% — reflects factors beyond true prevalence: selection into care and, most substantially, physician decision-making differences including both signal quality and diagnostic thresholds.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications?
A: The findings suggest that targeted interventions in physician awareness and clinical training are likely more effective than generic awareness campaigns, since the dominant driver of the diagnostic gap is physician threshold-setting rather than symptom prevalence. Structured decision support tools or updated training that make physicians aware of gender-specific diagnostic patterns could reduce medically unwarranted diagnostic differences. Policies targeting patient-side access barriers (the ~20% explained by selection) remain relevant but secondary. The roughly 20–25% of the gap attributable to genuine symptom prevalence differences is, by construction, guideline-consistent and should not be targeted for elimination.&lt;/p&gt;
&lt;p&gt;Q: What are the methodological contributions?
A: The paper makes three methodological contributions. First, it develops a structural model of mental health diagnosis that explicitly incorporates endogenous patient selection — a feature absent from standard physician decision-making models — which is shown empirically important. Second, it applies machine learning and NLP to clinical doctor note text to construct key unobserved clinical variables (behavioral assessment indicator and ADHD match signal) that are unavailable as structured data in EHRs. Third, the identification of population mean health risk uses a quasi-exogenous variation approach (IPCP referral rates) analogous to Arnold et al. (2022)&amp;rsquo;s method for measuring racial discrimination in bail decisions, adapted here to a continuous health risk setting with endogenous selection.&lt;/p&gt;
&lt;p&gt;Diagnostic threshold (τ_θ): The gender-specific posterior ADHD risk level above which a physician chooses to diagnose ADHD. Set ex-ante, it reflects the physician&amp;rsquo;s perceived tradeoff between the costs of over-diagnosis (misdiagnosis) and under-diagnosis (missed diagnosis). A lower threshold implies the physician views missed diagnosis as relatively more costly for that patient group. By construction, uniform clinical guidelines imply a single threshold independent of patient gender.&lt;/p&gt;
&lt;p&gt;ADHD match signal (x_i): A physician-observed, noisy signal of a patient&amp;rsquo;s true latent ADHD risk (v_i), observed only conditional on the patient receiving a behavioral assessment. In estimation, it is proxied via a cosine similarity measure between the patient&amp;rsquo;s aggregated clinical doctor note text and the DSM-V symptom list, constructed separately for inattentive and hyperactive/impulsive sub-types.&lt;/p&gt;
&lt;p&gt;Signal quality / diagnostic uncertainty (ρ_θ): The correlation between the physician&amp;rsquo;s observed ADHD match signal and the patient&amp;rsquo;s true ADHD risk. Higher ρ means the physician&amp;rsquo;s signal is more informative and diagnostic uncertainty is lower. In the Bayesian updating framework, higher ρ implies the physician places more weight on the observed signal relative to the prior.&lt;/p&gt;
&lt;p&gt;Mental healthcare utilization cost (c_i): The composite of all patient/caregiver factors that affect the decision to schedule a behavioral assessment net of child symptom level. Includes non-monetary barriers such as time constraints, distance, stigma, and information from primary care providers during wellness visits; does not include monetary out-of-pocket costs since insurance typically covers behavioral assessments.&lt;/p&gt;
&lt;p&gt;Initial Primary Care Provider (IPCP) referral rate: The regression-adjusted share of a given PCP&amp;rsquo;s patients who ultimately receive a behavioral assessment at some point in the sample. Used as a quasi-exogenous cost-shifter that influences patient scheduling costs without being correlated with patient ADHD risk, enabling identification of population mean ADHD risk via extrapolation.&lt;/p&gt;
&lt;p&gt;Latent ADHD risk (v_i): An unobserved continuous measure of a child&amp;rsquo;s underlying ADHD-related behavioral symptoms, drawn from a gender-specific normal distribution N(μ_θ, σ²_θ). A child&amp;rsquo;s true ADHD status is Si = 1(v_i &amp;gt; v̄), where v̄ is the DSM-V minimum symptom threshold, defined identically for boys and girls.&lt;/p&gt;
&lt;p&gt;Adjusted Bag-of-Words (BOW) cosine similarity: The NLP method used to construct the ADHD match signal proxy. Patient notes are tokenized into uni-grams and bi-grams after preprocessing (spell check, abbreviation replacement, part-of-speech tagging, synonym replacement), and tf-idf weighted. The cosine similarity between the resulting document vector and the DSM-V symptom text vector is computed separately for each ADHD sub-type and rescaled to [0,1].&lt;/p&gt;</description></item><item><title>On the Nature of Entrepreneurship</title><link>https://macropaperwarehouse.com/papers/on-the-nature-of-entrepreneurship/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/on-the-nature-of-entrepreneurship/</guid><description>&lt;p&gt;This paper uses a novel longitudinal administrative dataset drawn from U.S. Internal Revenue Service (IRS) and Social Security Administration (SSA) records to characterize income dynamics and the determinants of entrepreneurial entry for pass-through business owners — sole proprietors, partners, and S corporation owners — who collectively account for over 50 percent of all U.S. business net income. The sample covers 2000–2015 and includes up to 1.3 billion person-year observations for individuals aged 25–65. The authors construct balanced panels using birth cohorts 1950–1975, impute education (college attainment) and skill (cognitive, interpersonal, manual) via machine-learning classifiers trained on CPS and O*NET data, and estimate life-cycle income profiles using a three-component model that separates individual fixed effects, group-specific time effects, and group-cohort-specific age effects.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central departure from prior work is coverage of the full income distribution, including the high-earning right tail that household surveys such as the CPS misrepresent due to top-coding and small samples. When the IRS and CPS samples are compared on a consistent classification basis, median self-employment income is lower in the IRS data at all ages, consistent with the survey literature&amp;rsquo;s emphasis on the &amp;ldquo;typical&amp;rdquo; self-employed individual. However, mean incomes diverge sharply: the IRS shows mean self-employment income rising from $23 thousand at age 25 to $93 thousand at age 55, whereas the CPS (with incorporated owners reclassified) shows a rise from only $41 thousand to $73 thousand. Roughly 80 percent of self-employment income in the IRS data accrues to individuals above the $100 thousand threshold, compared to 42–53 percent in the CPS. The IRS-CPS gap is dominated by the right tail and concentrated in professional services and health care. For paid-employed individuals, the IRS and CPS medians and means are close at all ages, confirming the discrepancy is specific to self-employment.&lt;/p&gt;
&lt;p&gt;The life-cycle estimation finds that individuals who have &amp;ldquo;tried self-employment&amp;rdquo; — a group earning virtually all self-employment income — start at similar average incomes to primarily paid-employed peers at age 25 but reach $134 thousand by age 55, compared with $79 thousand for paid-employed peers with the same observable characteristics. Age effects for the self-employed are 63 percent higher than for the paid-employed at age 26 and remain elevated until age 55. Time effects show dramatically greater cyclical volatility for the self-employed: income growth declined by $9,655 (2008) and $8,785 (2009) for the self-employed versus $373 and $1,583 for paid-employed in the same years, concentrated in real estate and construction.&lt;/p&gt;
&lt;p&gt;On the determinants of entry, the paper finds: (i) no evidence that house-price appreciation raises entry rates, contra collateral-constraint hypotheses; (ii) most entrants have lower asset incomes than future entrants with the same characteristics, arguing against a liquid-wealth precondition; (iii) most entrants have higher prior labor income than future entrants, consistent with entry being driven by on-the-job experience rather than fallback from low-paid work; (iv) almost all founders report positive individual tax income in their first year of operation despite negative business net income and no external debt financing. Self-employed income growth exhibits greater dispersion — a 10th-to-90th percentile range roughly 2.5 times wider than for the paid-employed — and a Kelly skewness about 0.1 higher. A standard consumption-risk model calibrated with household-finance estimates of risk aversion rationalizes the patterns if individuals are insured against the most adverse downside shocks. Entry and exit rates are stable across the sample period, including the Great Recession, and the entrepreneurship share does not decline.&lt;/p&gt;
&lt;p&gt;The subgroup congruent with non-pecuniary motivation — primarily self-employed individuals earning less than paid-employed peers with matching characteristics — comprises roughly 57 percent of primarily self-employed by count but earns only 16 percent of total self-employment income.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-do-irs-and-cps-data-give-such-different-pictures-of-self-employment-income"&gt;Q1. Why do IRS and CPS data give such different pictures of self-employment income?&lt;/h3&gt;
&lt;p&gt;The CPS suffers from top-coding of high incomes and small samples that underrepresent high earners in key industries. The IRS-CPS mean income gap for the self-employed is dominated by the right tail: in the main IRS sample, individuals above the $100 thousand threshold earn roughly 80 percent of all self-employment income, versus 42 percent in the comparable CPS sample. The average income of top earners above $100 thousand is $355 thousand in the IRS versus $218 thousand in the CPS. The gap is concentrated in professional services and health care and persists across all income thresholds and sample definitions tested. No analogous discrepancy exists for paid-employed individuals, where IRS and CPS medians and means are close at all ages.&lt;/p&gt;
&lt;h3 id="q2-what-does-the-comparison-look-like-at-the-median-versus-the-mean"&gt;Q2. What does the comparison look like at the median versus the mean?&lt;/h3&gt;
&lt;p&gt;At the median, IRS self-employment income is lower than both CPS samples at all ages, with the gap largest for younger owners and those with incorporated businesses — a pattern consistent with the survey-based &amp;ldquo;self-employment discount&amp;rdquo; narrative. At the mean, the IRS shows much higher income at older ages: by age 55, IRS mean self-employment income is $93 thousand versus $73 thousand in the CPS sample that includes reclassified incorporated-owner wages. The divergence arises because the mean is sensitive to the right tail, which the CPS systematically underrepresents.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-estimate-life-cycle-income-profiles-while-separating-age-time-and-cohort-effects"&gt;Q3. How does the paper estimate life-cycle income profiles while separating age, time, and cohort effects?&lt;/h3&gt;
&lt;p&gt;Individual income is decomposed into an individual fixed effect (permanent latent ability and preferences), a group-specific time effect (business-cycle fluctuations common to a group), and a group-cohort-specific age effect (life-cycle income growth). Identification exploits the overlapping cohort structure of the 16-year panel: age effects are assumed equal across cohort bins of size at least two, allowing time and age effects to be separately identified. The model is estimated in levels rather than logs to accommodate business losses. Groups are defined as a Cartesian product of 32,256 subgroups based on education, three skill dimensions, industry (21 two-digit NAICS codes), demographics (gender, cohort, marital status, children), and employment-status history.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-headline-life-cycle-income-profile-findings-for-self--versus-paid-employed"&gt;Q4. What are the headline life-cycle income profile findings for self- versus paid-employed?&lt;/h3&gt;
&lt;p&gt;Among the &amp;ldquo;primarily employed&amp;rdquo; group, those who have tried self-employment and those who are primarily paid-employed have similar average incomes at age 25. By age 55 the self-employed reach an estimated $134 thousand (2012 dollars) versus $79 thousand for paid-employed peers with identical observable characteristics. The estimated age effect for the self-employed is 63 percent higher than for the paid-employed at age 26 and remains higher through age 55. These gaps would widen further if incomes were adjusted upward for the BEA-estimated net misreporting rates of 46 percent for unincorporated owners and 14 percent for S corporation owners.&lt;/p&gt;
&lt;h3 id="q5-how-large-is-the-group-consistent-with-non-pecuniary-motivation-and-how-much-income-does-it-earn"&gt;Q5. How large is the group consistent with non-pecuniary motivation, and how much income does it earn?&lt;/h3&gt;
&lt;p&gt;The non-pecuniary subgroup — primarily self-employed individuals (at least 12 years in self-employment) who earn less on average than primarily paid-employed peers matched on gender, education, skills, and other characteristics — is numerically larger, comprising approximately 57 percent of primarily self-employed by count. However, this group earns only 16 percent of total self-employment income. Adjusting for paid-employed fringe benefits and self-employed income misreporting can change the group&amp;rsquo;s size but does not alter the finding that it accounts for a small income share. The paper concludes that non-pecuniary motives may guide occupational choice for many individuals but are not the driver of the typical dollar earned in self-employment.&lt;/p&gt;
&lt;h3 id="q6-how-does-idiosyncratic-income-risk-compare-between-self--and-paid-employed"&gt;Q6. How does idiosyncratic income risk compare between self- and paid-employed?&lt;/h3&gt;
&lt;p&gt;Self-employed income changes are substantially more dispersed: the 10th-to-90th percentile range of income growth is roughly 2.5 times wider for the self-employed than for the paid-employed. Income changes for the self-employed are also more right-skewed, with a Kelly skewness difference of approximately 0.1. When a standard consumption-risk model — augmented with a lower bound on consumption growth to allow for external insurance — is parameterized with risk-aversion estimates from the household finance literature, the observed patterns are rationalized if individuals are insured against the most adverse downside shocks, i.e., the attractive aspect of self-employment is large potential upside with insured downside.&lt;/p&gt;
&lt;h3 id="q7-what-happened-to-self-employed-income-and-exit-rates-during-the-great-recession"&gt;Q7. What happened to self-employed income and exit rates during the Great Recession?&lt;/h3&gt;
&lt;p&gt;Time effects show steep income growth declines for the self-employed of -$9,655 in 2008 and -$8,785 in 2009, compared with much more modest declines of -$373 and -$1,583 for paid-employed peers. The aggregate income declines are concentrated in cyclically sensitive self-employed subgroups in real estate and construction, with their paid-employed counterparts experiencing only modest declines. Despite these large income shocks, exit rates from self-employment showed little change during the Great Recession, either in aggregate or in the cyclically sensitive sectors. Entry rates were likewise stable, and the share of entrepreneurs in the population did not decline over the full sample period.&lt;/p&gt;
&lt;h3 id="q8-does-the-evidence-support-collateral-constraints-as-a-binding-barrier-to-entrepreneurial-entry"&gt;Q8. Does the evidence support collateral constraints as a binding barrier to entrepreneurial entry?&lt;/h3&gt;
&lt;p&gt;No. The paper tests the hypothesis, standard in the liquidity-constraints literature, that entry rates should be higher for homeowners experiencing house-price appreciation (which raises collateral value). The IRS data do not support this prediction. Separately, comparing asset incomes (interest, dividends, capital gains) of current entrants and future entrants with the same characteristics, the paper finds that most current entrants have lower asset incomes and less liquid wealth than those who switch later, which also argues against a liquid-wealth precondition for entry.&lt;/p&gt;
&lt;h3 id="q9-what-does-prior-labor-income-reveal-about-why-people-enter-self-employment"&gt;Q9. What does prior labor income reveal about why people enter self-employment?&lt;/h3&gt;
&lt;p&gt;Current entrants have higher prior labor income than matched future entrants with the same characteristics, indicating they enter with accumulated on-the-job experience rather than being pushed into self-employment as a fallback after failure in paid work. This is consistent with self-employment being a deliberate, experience-driven career transition for most entrants rather than a last resort for low earners. The paper interprets this as positive evidence for the role of experience-based human capital in driving entrepreneurial choice.&lt;/p&gt;
&lt;h3 id="q10-how-do-founders-finance-startup-costs-if-most-have-negative-business-net-income-in-early-years"&gt;Q10. How do founders finance startup costs if most have negative business net income in early years?&lt;/h3&gt;
&lt;p&gt;Almost all founders in the sample report positive income on their personal (individual) tax form in the first year of operation, even though most report negative business net income and carry no external debt financing. This pattern suggests founders rely on personal income sources — prior savings, part-time paid employment, or spousal income — to cover startup costs rather than external debt, implying that formal credit-market financing constraints are not the primary barrier to entry for most entrants in the sample.&lt;/p&gt;
&lt;h3 id="q11-what-are-the-scope-conditions-and-key-limitations"&gt;Q11. What are the scope conditions and key limitations?&lt;/h3&gt;
&lt;p&gt;The sample covers pass-through owners (sole proprietors, partners, S corporation owners) and excludes C corporation shareholders, whose entrepreneurial income does not flow to individual returns until distributed. Income measures exclude most employer fringe benefits; capital gains are excluded from self-employment income, and the authors note their inclusion would strengthen the main findings. The analysis covers 2000–2015 for cohorts born 1950–1975, and income is reported before taxes and transfers. Baseline estimates are not adjusted for misreporting, though BEA-implied adjustments of 46 percent for unincorporated owners and 14 percent for S corporation owners would widen the income gaps further.&lt;/p&gt;
&lt;p&gt;Pass-through business owner: An individual who owns a sole proprietorship, partnership, or S corporation, such that business net income flows directly onto the owner&amp;rsquo;s personal tax return; excludes C corporation shareholders whose income appears only upon dividend or capital-gains distributions.&lt;/p&gt;
&lt;p&gt;Tried self-employment: The paper&amp;rsquo;s primary self-employed comparison group within the &amp;ldquo;primarily employed&amp;rdquo; category — individuals with any years in self-employment (including frequent switchers and those with most years in self-employment) — who collectively earn virtually all self-employment income.&lt;/p&gt;
&lt;p&gt;Group-specific age effect: The paper&amp;rsquo;s estimate of how individual income changes with age within a defined subgroup (determined by education, skill, industry, demographics, and employment history), identified by exploiting overlapping birth cohorts in the 16-year panel and separated from individual fixed effects and business-cycle time effects.&lt;/p&gt;
&lt;p&gt;Primarily employed: Individuals with at least 12 of 16 sample years in either self- or paid-employment, with at most one intermediate year of non-employment; the paper&amp;rsquo;s main analytical focus for life-cycle income comparisons.&lt;/p&gt;
&lt;p&gt;SOI Databank: The Statistics of Income Databank, a de-identified balanced panel combining SSA demographic records with IRS tax filing data for all living U.S. individuals with a Social Security number over 1996–2015; the paper&amp;rsquo;s primary data source providing Schedule C, K-1, W-2, and related filing information.&lt;/p&gt;
&lt;p&gt;Kelly skewness: A robust measure of distributional asymmetry used by the paper to characterize income growth; the paper reports that Kelly skewness of self-employed income changes exceeds that of paid-employed by approximately 0.1, indicating greater right-skewness in self-employment income dynamics.&lt;/p&gt;
&lt;p&gt;Non-pecuniary motivation subgroup: Primarily self-employed individuals who earn less on average than primarily paid-employed peers matched on observable characteristics, taken by the paper as consistent with non-wage job amenities (autonomy, flexibility) driving occupational choice; found to be 57 percent of primarily self-employed by count but earning only 16 percent of total self-employment income.&lt;/p&gt;</description></item><item><title>On the Optimal Design of a Financial Stability Fund</title><link>https://macropaperwarehouse.com/papers/on-the-optimal-design-of-a-financial-stability-fund/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/on-the-optimal-design-of-a-financial-stability-fund/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks how to optimally design a Financial Stability Fund (Fund) for a union of sovereign countries that must simultaneously (i) prevent sovereign default, (ii) provide risk-sharing and consumption smoothing, (iii) respect countries&amp;rsquo; sovereignty (limited enforcement on both sides), (iv) address moral hazard from governments&amp;rsquo; non-contractable policy reform effort, and (v) never impose permanent transfers or incur undesired expected losses. The paper develops the formal theory of such a Fund and evaluates it quantitatively against an incomplete-markets economy with sovereign default (IMD), calibrated to euro area &amp;ldquo;stressed countries&amp;rdquo; (Greece, Italy, Portugal, Spain — the GIPS).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model Setup and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The Fund is modeled as a long-term contract between a risk-neutral lender (the Fund) and a risk-averse, relatively impatient borrower (a small open-economy sovereign). The government maximizes lifetime utility over consumption, leisure, and effort, where effort is private information (non-contractable) and determines the distribution of future endogenous government expenditure shocks. Two-sided limited enforcement (LE) constraints govern the contract: the borrower&amp;rsquo;s constraint ensures the country never prefers autarky-with-default to staying in the Fund; the lender&amp;rsquo;s constraint ensures the Fund never prefers investing at the risk-free rate to continuing the contract. The lender&amp;rsquo;s constraint is set with Z = 0 in the benchmark, meaning the Fund never accepts any expected permanent transfers — no ex-ante or ex-post redistribution.&lt;/p&gt;
&lt;p&gt;Because LE and moral hazard (MH) constraints are forward-looking, standard dynamic programming cannot be applied directly. The paper uses recursive contracts (a Saddle-Point Functional Equation, SPFE) with a discounted relative Pareto weight x as the co-state variable. The SPFE characterizes the constrained-efficient allocation. The paper then proves two welfare theorems, providing a novel decentralization of the Fund contract as a recursive competitive equilibrium (RCE) with state-contingent long-term bonds, Pigouvian taxes on Arrow securities (budget-neutral in equilibrium), and endogenous borrowing limits.&lt;/p&gt;
&lt;p&gt;The benchmark (IMD) economy features long-term non-contingent defaultable debt modeled following Chatterjee–Eyigungor, with asymmetric default penalties and probabilistic market re-entry after default (λ = 0.264). Both economies are calibrated to GIPS data for 1980–2015 using a panel Markov regime-switching AR(1) productivity process with three regimes (crisis, intermediate, normal). Key parameters: β = 0.929, r = 2.48%, δ = 0.814, κ = 0.083, labor share α = 0.566.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings with Quantitative Magnitudes&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Borrowing capacity&lt;/strong&gt;: The Fund supports a long-run average debt-to-GDP ratio of 191 percent, compared with 78.6 percent in the IMD economy — more than double — while eliminating default episodes entirely. At the state-level, the maximum debt capacity of the Fund ranges from roughly 99–293 percent of GDP across states, versus 1.6–184 percent in the IMD economy; capacity in bad states (low θ, high g) under the IMD falls to under 2 percent, while the Fund can absorb close to 100 percent even in the worst state.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Consumption volatility&lt;/strong&gt;: The relative volatility of consumption to output falls from 139 percent in the IMD economy to 36 percent under the Fund, reflecting greatly improved risk sharing through state-contingent payments.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Primary surplus co-movement&lt;/strong&gt;: The cyclical correlation of the primary surplus with output rises from 0.23 (mildly procyclical — consistent with some consumption smoothing but limited by borrowing constraints and default risk) in the IMD to 0.94 under the Fund, enabling counter-cyclical primary deficits during crises.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Effort&lt;/strong&gt;: The long-run mean effort is 17 percent higher under the Fund than in the IMD economy in normal times, reflecting the Fund&amp;rsquo;s long-horizon incentive structure. However, during a crisis, effort is lower under the Fund than under the IMD — the Fund deems high effort in a crisis not part of the efficient allocation, in contrast to the IMD where spreads and borrowing constraints impose austerity-like discipline.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Welfare gains&lt;/strong&gt;: Starting from zero initial debt, the consumption-equivalent steady-state average welfare gain of the Fund is approximately 8.5 percent (ergodic mean-weighted), ranging from 7.0 percent in the best state (high θ, low g) to 10.3 percent in the worst state (low θ, high g). In a counterfactual crisis simulation initialized at pre-crisis GIPS levels (70 percent debt-to-GDP, 0.8 percent spread), the welfare gain rises to approximately 10.59 percent in consumption-equivalent terms, exceeding the zero-debt benchmark of 8.57 percent for the same shock state.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Welfare decomposition&lt;/strong&gt;: For the two worst-shock states examined, higher debt capacity (channel iii) and state-contingent insurance (channel iv) together account for more than 90 percent of total welfare gains — specifically, 63.65 percent and 28.10 percent for (θl, gh), and 51.92 percent and 41.39 percent for (θl, gl), respectively. The direct costs of default (output penalty and market exclusion) together contribute less than 10 percent of total gains.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Spreads&lt;/strong&gt;: The IMD economy generates positive spreads reflecting default risk. The Fund economy generates only non-positive spreads in equilibrium — negative spreads arise when the lender&amp;rsquo;s limited enforcement constraint is binding (i.e., when continuing to lend risks permanent Fund losses, so the Fund restrains the borrower). This negative spread is interpretable as a Debt Sustainability Analysis signal.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Calibration is to GIPS countries over 1980–2015. The Fund assumes full exclusivity (absorbs all sovereign debt). A follow-up paper by other authors shows similar welfare gains hold when only a minimal fraction of debt is absorbed. The benchmark sets Z = 0 (no solidarity transfers); relaxing Z &amp;lt; 0 would allow greater risk sharing. The borrower is strictly more impatient than the lender (η = β(1+r) = 0.9684 &amp;lt; 1).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-two-limited-enforcement-le-constraints-in-the-fund-contract-and-what-do-they-individually-prevent"&gt;Q1. What are the two limited enforcement (LE) constraints in the Fund contract, and what do they individually prevent?&lt;/h3&gt;
&lt;p&gt;A: The borrower&amp;rsquo;s LE constraint (constraint 1) ensures the country&amp;rsquo;s continuation value under the Fund always weakly exceeds its outside option V°(s) — the value of defaulting and entering incomplete markets as a defaulter. This prevents the borrower from reneging on the Fund contract. The lender&amp;rsquo;s LE constraint (constraint 3) ensures the Fund&amp;rsquo;s expected net present value of transfers never falls below Z (set to 0 in the benchmark), preventing the Fund from making permanent expected losses. Together, these two constraints define an interval [x(s), x̄(s)] for the relative Pareto weight within which both parties remain voluntarily in the contract.&lt;/p&gt;
&lt;h3 id="q2-how-does-moral-hazard-enter-the-model-and-what-is-the-key-assumption-enabling-the-first-order-condition-foc-approach"&gt;Q2. How does moral hazard enter the model, and what is the key assumption enabling the first-order-condition (FOC) approach?&lt;/h3&gt;
&lt;p&gt;A: Government effort e ∈ [0,1] is non-contractable; it shifts the distribution of future government expenditure shocks g in a first-order stochastically dominant direction (higher effort → lower expected g). The incentive compatibility constraint (ICC, constraint 2) imposes that the marginal cost of effort v′(e) equals the marginal benefit in terms of expected future utility changes. The FOC approach is validated by Assumption 1 (monotone likelihood ratio condition on the g-shock transition, and convexity of the CDF with respect to effort), which guarantees the ICC is sufficient as well as necessary. Without this assumption, the full optimization problem would need to replace the ICC, making the recursive formulation substantially more complex.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-achieve-a-recursive-formulation-despite-forward-looking-le-and-mh-constraints"&gt;Q3. How does the paper achieve a recursive formulation despite forward-looking LE and MH constraints?&lt;/h3&gt;
&lt;p&gt;A: The paper uses the saddle-point Lagrangian approach (following Marcet–Marimon). Rather than tracking the full history of constraints, it introduces a discounted relative Pareto weight x ≡ [β(1+r)]^t · (µ_b,t / µ_l,t) as the sufficient co-state variable. The law of motion for x adjusts at each state realization: the borrower&amp;rsquo;s LE multiplier ν_b raises x (rewards the borrower), the lender&amp;rsquo;s LE multiplier ν_l lowers x (restrains the borrower), and the MH multiplier ρ̺ shifts x up or down depending on whether the realized g provides a positive or negative signal about effort (monotone likelihood ratio). This collapses the problem to a stationary Saddle-Point Functional Equation (SPFE) in (x, s).&lt;/p&gt;
&lt;h3 id="q4-what-are-the-key-properties-of-the-optimal-fund-allocation-characterized-in-the-paper"&gt;Q4. What are the key properties of the optimal Fund allocation characterized in the paper?&lt;/h3&gt;
&lt;p&gt;A: (i) When neither LE constraint binds, consumption increases with x and is constant in s (perfect Pareto weight-determined risk sharing), labor supply is undistorted and increases in θ, and x declines over time due to borrower impatience (η &amp;lt; 1). (ii) When the borrower&amp;rsquo;s LE binds (x ≤ x̄(s)), consumption, labor, and x are pinned at x̄(s) and the borrower is prevented from receiving less. (iii) When the lender&amp;rsquo;s LE binds (x ≥ x̄(s)), the same constancy holds and the lender is prevented from being overexposed. Moral hazard introduces state-contingency in the inter-period evolution of x even when neither LE binds, via the likelihood ratio term. The paper shows that immiseration (consumption converging to zero) is prevented by the borrower&amp;rsquo;s LE constraint, even in the presence of moral hazard.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-modified-inverse-euler-equation-in-this-model-and-how-does-it-differ-from-standard-formulations"&gt;Q5. What is the modified inverse Euler equation in this model, and how does it differ from standard formulations?&lt;/h3&gt;
&lt;p&gt;A: In the standard pure moral hazard problem, the inverse of the marginal utility process is a positive supermartingale, leading to immiseration (consumption converging to zero) when the borrower is impatient. In this model with two-sided LE and MH, the inverse Euler equation (Lemma 4, equation 21) has the form: E_s[{1/u′(c(x′,s′))} · {(1+ν_l)/(1+ν_b)}] = η · {1/u′(c(x,s))}. The LE multipliers truncate the supermartingale whenever borrower or lender constraints bind, recurrently preventing both immiseration and permanent lender losses. The MH constraint introduces state-contingent perturbations to the path of consumption (via likelihood ratios) even between binding episodes.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-novel-decentralization-result-and-why-is-it-theoretically-significant"&gt;Q6. What is the novel decentralization result, and why is it theoretically significant?&lt;/h3&gt;
&lt;p&gt;A: The paper provides two welfare theorems (Propositions 1 and 2). The Second Welfare Theorem shows that any constrained-efficient Fund contract can be decentralized as a recursive competitive equilibrium with: (a) long-term state-contingent (Arrow security) assets, (b) Pigouvian state-contingent taxes τ^a(s′) on Arrow securities — which are budget-neutral in equilibrium — where 1/(1+τ^a(s′)) = 1 + χ(x,s)·u′(c(x,s))·[∂_e π(s′|s,e)/π(s′|s,e)], and (c) endogenous borrowing limits &amp;ldquo;not too tight&amp;rdquo; relative to outside options. The First Welfare Theorem shows the reverse. This decentralization is novel because it handles both limited commitment and dynamic moral hazard simultaneously — prior work handled each in isolation. The taxes internalize the full social value of effort by creating a wedge between the borrower&amp;rsquo;s and lender&amp;rsquo;s intertemporal rates of substitution, removing the need to impose the ICC directly as a constraint in the competitive equilibrium.&lt;/p&gt;
&lt;h3 id="q7-what-drives-the-negative-spreads-in-the-fund-economy-and-how-do-they-differ-from-the-positive-spreads-in-the-imd-economy"&gt;Q7. What drives the negative spreads in the Fund economy, and how do they differ from the positive spreads in the IMD economy?&lt;/h3&gt;
&lt;p&gt;A: In the IMD economy, positive spreads reflect the probability of default: the bond price embeds an expected default discount. In the Fund economy, default is eliminated by construction. Negative spreads arise when the lender&amp;rsquo;s LE constraint is binding in some future state s′ (i.e., ν_l(x′,s′) &amp;gt; 0): this means the borrower&amp;rsquo;s Pareto weight is so high that the Fund risks permanent losses by continuing to lend. The asset price equation (45) shows the Arrow security price equals the maximum of the borrower&amp;rsquo;s discounted marginal utility valuation and the risk-free discounted return — so when the lender&amp;rsquo;s constraint binds, the price is driven by the risk-free return (q(s′|s) = π(s′|s,e)·A(s′)/(1+r)), which generates a negative implicit spread. The negative spread acts as a DSA-like signal: the Fund is better off restraining lending in those states.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-calibration-match-the-gips-data-and-what-is-the-main-misfit"&gt;Q8. How does the calibration match the GIPS data, and what is the main misfit?&lt;/h3&gt;
&lt;p&gt;A: The IMD economy is calibrated to average GIPS moments over 1980–2015 using a panel Markov regime-switching AR(1) for productivity (three regimes: crisis, intermediate, normal) and a three-state government expenditure process. The model matches well: average debt/GDP of 78.57 percent (data: 78.33), average spread of 4.17 percent (data: 4.15), labor moments, relative volatility of spreads (1.74 vs. 1.67 in data), government-output correlation (0.38 matches data), and relative volatility of the primary surplus (0.97 vs. 1.00 in data). The main misfit is the average primary surplus/GDP: the model generates a positive value (consistent with stationarity and debt servicing), while the data shows a slight deficit over the sample, plausibly reflecting growth expectations. The paper notes this level misfit does not compromise its core welfare-comparison results, since what matters is the relative time-series behavior.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-fund-compare-to-the-imd-economy-in-the-crisis-simulation-initialized-at-pre-2008-gips-conditions"&gt;Q9. How does the Fund compare to the IMD economy in the crisis simulation initialized at pre-2008 GIPS conditions?&lt;/h3&gt;
&lt;p&gt;A: The economy is initialized at 70 percent debt-to-GDP and 0.8 percent spread (consistent with 2005–2007 GIPS averages), then hit with a negative productivity and high government expenditure shock. In the IMD economy, this shock generates a wave of defaults (Figure 6), sharp spread increases (spreads spike, consistent with GIPS experience of 2009–2010 where spreads reached 4.04 percent on average), and a required increase in labor supply despite low productivity. Under the Fund, no defaults occur: instead, the country runs a large primary deficit financed by the state-contingent component of the Fund contract (debt actually falls under the Fund while rising in the IMD), consumption is higher than in the IMD for approximately the first 10 periods of the crisis, and labor supply is allowed to fall (consistent with efficiency). The welfare gain in this counterfactual is approximately 10.59 percent in consumption-equivalent terms, exceeding the zero-debt-initial-condition gain of 8.57 percent for the same shock state, demonstrating that welfare gains are amplified when the Fund takes over pre-existing debt.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-fund-affect-effort-incentives-differently-in-normal-times-versus-crisis-times"&gt;Q10. How does the Fund affect effort incentives differently in normal times versus crisis times?&lt;/h3&gt;
&lt;p&gt;A: In normal times, the Fund provides better incentives for effort: long-run average effort is 17 percent higher under the Fund than in the IMD economy. The Fund&amp;rsquo;s long-term contract links future government expenditure outcomes directly to future lifetime utility via the law of motion for x (equation 5): low g realizations shift x upward (reward the borrower), creating forward-looking incentives. In crisis times, the Fund allows effort to fall relative to the IMD economy; the IMD imposes higher effort in bad states through spread increases and effective borrowing constraints that make budget relief through effort more valuable. The paper interprets this as the efficient outcome: &amp;ldquo;austerity&amp;rdquo; (high effort during a crisis) is not part of the constrained-efficient Fund allocation.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-welfare-decomposition-methodology-and-what-does-it-reveal-about-channels-of-welfare-gain"&gt;Q11. What is the welfare decomposition methodology, and what does it reveal about channels of welfare gain?&lt;/h3&gt;
&lt;p&gt;A: The authors construct a sequence of counterfactual IMD economies. Channel (i) removes the output penalty upon default, isolating its welfare cost: contributes 6.58 percent (θl, gh) and 5.31 percent (θl, gl) of total gain. Channel (ii) additionally removes market exclusion after default (immediate return): contributes 1.67 percent and 1.38 percent respectively. Channel (iii) solves counterfactual economies with the Fund&amp;rsquo;s state-specific endogenous borrowing limits but no default allowed, quantifying the value of greater debt capacity: contributes 63.65 percent and 51.92 percent. Channel (iv) is the residual attributable to state-contingent insurance payments: contributes 28.10 percent and 41.39 percent. The decomposition reveals that in the worst state (θl, gh), debt capacity dominates (63.65 percent), while in (θl, gl) — where the low government expenditure partially offsets low productivity — state-contingent insurance is relatively more important (41.39 percent). Together, channels (iii) and (iv) exceed 90 percent of total gains in both cases examined.&lt;/p&gt;
&lt;h3 id="q12-why-is-the-funds-decentralization-unlikely-to-emerge-from-private-international-capital-markets"&gt;Q12. Why is the Fund&amp;rsquo;s decentralization unlikely to emerge from private international capital markets?&lt;/h3&gt;
&lt;p&gt;A: Two reasons are given. First, private international lenders typically lack the legal authority to impose state-contingent taxes (τ^a(s′)) on domestic economies; these taxes are a necessary component of the decentralization to internalize the social value of effort. Second, even if such taxes were optimal from the joint perspective of borrower and lender, the borrower has no unilateral incentive to impose them given market conditions — the taxes are only individually rational within the Fund&amp;rsquo;s constrained-efficient contract. This provides a rationale for an institutional implementation of the Fund rather than reliance on decentralized sovereign debt markets.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Financial Stability Fund (Fund)&lt;/strong&gt;: A long-term partnership contract between a risk-neutral lender (the Fund) and a risk-averse sovereign borrower, designed to provide risk-sharing and consumption smoothing through state-contingent transfers subject to two-sided limited enforcement and moral hazard constraints, without ever incurring expected permanent losses. Distinguished from standard lending by its long-term contingent structure and dual role as risk-sharing mechanism and crisis-resolution tool.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Two-sided limited enforcement (LE) constraints&lt;/strong&gt;: Forward-looking constraints in the Fund contract that prevent either party from reneging. The borrower&amp;rsquo;s LE constraint ensures the contract always delivers at least as much lifetime utility as defaulting and entering incomplete debt markets. The lender&amp;rsquo;s LE constraint (with Z = 0 in the benchmark) ensures the Fund never accumulates a negative expected net present value from its contractual obligations — i.e., no permanent transfers occur. Both constraints are binding recurrently in the long-run ergodic set.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Moral hazard (MH) / incentive compatibility constraint (ICC)&lt;/strong&gt;: The constraint arising from the fact that government policy reform effort e is non-contractable (sovereign right). The ICC requires that the marginal cost of effort v′(e) equals the marginal lifetime benefit, which depends on the likelihood ratio of future shocks with respect to effort. The Fund contract provides long-horizon performance-based rewards and punishments (via the law of motion of the relative Pareto weight x) to induce efficient effort, without imposing ex-ante austerity conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Discounted relative Pareto weight (x)&lt;/strong&gt;: The key co-state variable in the recursive formulation, defined as x_t = [β(1+r)]^t · (µ_b,t / µ_l,t), where µ_b and µ_l are the time-varying Pareto weights of borrower and lender. It captures the entire history of binding constraints and serves as the state variable summarizing the borrower&amp;rsquo;s &amp;ldquo;entitlement&amp;rdquo; in the contract. Declines over time due to borrower impatience (η = β(1+r) &amp;lt; 1), but is upward-adjusted when the borrower&amp;rsquo;s LE constraint binds, and shifts state-contingently due to MH likelihood ratios.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Saddle-Point Functional Equation (SPFE)&lt;/strong&gt;: The recursive formulation of the Fund contracting problem (equation 6), analogous to Bellman&amp;rsquo;s equation but for saddle-point (min-max) problems. Required because standard dynamic programming fails when constraints are forward-looking; solved by the Marcet–Marimon recursive contract approach. The SPFE characterizes the constrained-efficient Fund allocation as a function of the co-state x and exogenous state s.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Incomplete markets with default (IMD) economy&lt;/strong&gt;: The benchmark comparison economy in which the sovereign borrows via non-contingent long-term defaultable bonds (parameterized by maturity δ and coupon κ), with asymmetric output penalties upon default and probabilistic market re-entry. Calibrated to GIPS countries 1980–2015. Generates positive spreads that reflect default risk; serves as both the status quo and the source of the borrower&amp;rsquo;s outside option V°(s) in the Fund contract.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pigouvian Arrow security taxes&lt;/strong&gt;: State-contingent taxes τ^a(s′) on Arrow security holdings, defined by 1/(1+τ^a(s′)) = 1 + χ(x,s)·u′(c)·[∂_e π/π], introduced in the decentralization of the Fund contract. These taxes create a wedge between the borrower&amp;rsquo;s and lender&amp;rsquo;s intertemporal rates of substitution to internalize the full social value of non-contractable effort. Budget-neutral in equilibrium: the government&amp;rsquo;s lump-sum transfer τ(s) exactly offsets expected tax revenue.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Debt Sustainability Analysis (DSA) interpretation&lt;/strong&gt;: The paper interprets the lender&amp;rsquo;s LE constraint (Z = 0) as a Fund-level DSA: it sets the boundary beyond which the contract would embed permanent transfers. A negative spread in the Fund economy signals that the lender&amp;rsquo;s LE constraint is binding in some future state — a DSA warning that the Fund is better off investing at the risk-free rate rather than extending more credit.&lt;/p&gt;</description></item><item><title>Online Business Models, Digital Ads, and User Welfare</title><link>https://macropaperwarehouse.com/papers/online-business-models-digital-ads-and-user-welfare/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/online-business-models-digital-ads-and-user-welfare/</guid><description>&lt;p&gt;Acemoglu, Huttenlocher, Ozdaglar, and Siderius develop a two-sided platform model to study the welfare consequences of digital advertising as an online business model. The platform intermediates between a firm selling a horizontally differentiated product and a continuum of users who derive utility from both entertaining content and informative signals about product quality embedded in ads. Users have a two-dimensional type: a sophistication dimension (sophisticated with probability lambda, naïve with probability 1-lambda) and a product-quality dimension (high quality with prior probability q). The central departure from the standard informational-advertising literature is that sophisticated users hold the correct model of the ad signal process, while naïve users underestimate the false-positive rate — the probability that a low-quality product generates a positive ad signal (phi_0). Naïve users perceive this false-positive rate to be phi_{0,N} = omega_N * omega_P * phi_0, where omega_N &amp;lt;= 1 captures inherent naïveté and omega_P &amp;lt;= 1 captures failure to understand personalized targeting, so phi_{0,N} &amp;lt; phi_0. The equilibrium concept is Berk-Nash equilibrium (Esponda and Pouzo 2016), meaning all agents are Bayesian given their subjective model.&lt;/p&gt;
&lt;p&gt;The platform chooses ad load alpha (Poisson rate of ad displays), subscription fees, and the monetary transfer from the firm; the firm sets product price p after observing the platform&amp;rsquo;s contract. The central finding (Proposition 2) is that when the objective false-positive rate phi_0 exceeds a threshold phi-hat_0(lambda, phi_1, phi_{0,N}) — which is increasing in lambda and phi_{0,N} and decreasing in the true-positive rate phi_1 — the unique equilibrium is an advertising-based plan that fully segments the market: naïve users receive an ad load that extracts all their surplus, while sophisticated users are excluded entirely. In this regime the firm charges a strictly higher price p-hat* &amp;gt; p-bar*, where p-bar* = (beta*q + c)/2 is the monopoly price without advertising. The ad-based equilibrium emerges precisely when ads are more misleading (larger gap between phi_0 and phi_{0,N}), not when they are more informative — a comparative static the authors describe as paradoxical.&lt;/p&gt;
&lt;p&gt;Welfare consequences (Proposition 4) are unambiguous in the advertising regime: both naïve and sophisticated users are strictly worse off than the baseline without any platform. Naïve users over-purchase due to inflated posteriors from misread signals; sophisticated users are harmed through the price channel — the firm&amp;rsquo;s higher profit-maximizing price p-hat* applies to all buyers. In the fully rational benchmark (phi_{0,N} = phi_0), the unique equilibrium is subscription-based and user welfare equals the no-platform baseline (Proposition 3).&lt;/p&gt;
&lt;p&gt;These results extend to richer menus (Proposition 5), mixed subscription-plus-advertising plans (Proposition 7), and to multi-firm and multi-platform competition (Propositions 9-12). Digital ads soften Bertrand competition by generating endogenous horizontal differentiation among otherwise identical firms, so equilibrium prices can exceed marginal cost even with two competing firms. Platform competition similarly fails to restore welfare: platforms compete away subscription fees but both adopt ad-based plans targeting naïfs when phi_1 exceeds a threshold, maintaining the welfare loss.&lt;/p&gt;
&lt;p&gt;On policy, the first best (planner observes types) cannot be decentralized because naïve users prefer more ads than is socially optimal, inverting the usual self-selection constraint. The second best (planner subject to incentive-compatibility constraints) is a single pooling plan with an intermediate ad load alpha^{SB} in [alpha^{FB}_N, alpha^{FB}_S] and yields average welfare above the no-platform baseline, though below first best (Proposition 13). This second best can be decentralized with a nonlinear digital ad tax, a per-unit product subsidy, and a platform subscription subsidy (Proposition 14). A simpler flat tax on digital ad revenues — above a threshold gamma-bar &amp;lt; 1 — also improves welfare relative to the ad-based equilibrium, though it does not restore the second best (Proposition 15).&lt;/p&gt;
&lt;p&gt;Four robustness extensions are developed: endogenous manipulation (platform always chooses the most manipulative environment, lowest phi_{0,N}); naïve learning dynamics (learning raises the sophisticate share in steady state, making ad-based models less profitable but not overturning the main results); imperfect price discrimination by the firm (naïfs are unambiguously worse off, threshold for advertising equilibrium shifts down); and an added price-sensitivity dimension (the platform runs a 2x2 menu separating by both sophistication and price sensitivity, preserving the result that naïve users tolerate and receive more ads than sophisticates in every stratum).&lt;/p&gt;
&lt;p&gt;Q: What is the key asymmetry between naïve and sophisticated users that drives the main results?
A: Sophisticated users hold the correct Bayesian model of the ad signal process and thus correctly account for the false-positive rate phi_0 when updating beliefs from positive ad signals. Naïve users perceive the false-positive rate as phi_{0,N} = omega_N * omega_P * phi_0 &amp;lt; phi_0, so they treat positive signals as stronger evidence of high product quality than they actually are. Because naïve users overestimate the informativeness of ads, their (interim) subjective valuation of an ad-based plan is higher, making them more tolerant of ad loads and more willing to join platforms with heavy advertising. This asymmetry is what makes it profitable to target naïfs with high ad loads while excluding or charging subscription fees to sophisticates.&lt;/p&gt;
&lt;p&gt;Q: Why does advertising to sophisticated users generate no additional firm profit, while advertising to naïve users does?
A: Lemma 1 establishes that with linear-quadratic utility the firm extracts no surplus from advertising to sophisticates: because sophisticated agents are fully Bayesian, their expected posterior equals the prior (E_S[pi_i] = q), so expected demand after advertising is identical to demand before advertising. By contrast, Lemma 2 shows that the firm&amp;rsquo;s profit from naïve agents is positive and strictly increasing in ad load alpha, because naïve users&amp;rsquo; average demand curve drifts upward as alpha rises — their inflated perceived informativeness of ads causes them to over-update on positive signals, systematically raising their willingness to pay. The platform captures this surplus from the firm via the advertising transfer m*.&lt;/p&gt;
&lt;p&gt;Q: What is the threshold condition determining whether the equilibrium is subscription-based or advertising-based?
A: Proposition 2 identifies a threshold phi-hat_0(lambda, phi_1, phi_{0,N}) that is increasing in the sophisticate share lambda and in the naïve false-positive perception phi_{0,N}, and decreasing in the true-positive rate phi_1. When the objective false-positive rate phi_0 is below this threshold, the profit-maximizing business model is subscription-based with price P* = T - v and product price p* = p-bar* = (beta&lt;em&gt;q + c)/2. When phi_0 exceeds the threshold, the advertising model dominates: the platform sets a high ad load alpha-hat&lt;/em&gt; that makes naïve users exactly indifferent between participating and their outside option v, excludes sophisticates, and the firm charges p-hat* &amp;gt; p-bar*. The threshold falls with phi_1, meaning more informative ads expand the range of phi_0 over which the advertising equilibrium obtains.&lt;/p&gt;
&lt;p&gt;Q: How does allowing the platform to offer menus change the results relative to the baseline two-plan case?
A: Proposition 5 shows that with menus the platform can simultaneously serve both user types: sophisticates receive a subscription plan at P* = T - v and naïve users receive an ad-based plan with the same high load alpha-hat* as in the baseline. The threshold for the advertising equilibrium shifts down to phi*&lt;em&gt;0(lambda, phi_1, phi&lt;/em&gt;{0,N}) &amp;lt; phi-hat_0, so advertising business models arise for a strictly larger set of parameters. Welfare consequences are unchanged (Corollary 1): when phi_0 &amp;gt; phi*_0, both types have welfare strictly below the no-platform baseline. Proposition 6 further shows consumer welfare is monotonically decreasing in both phi_0 and phi_1: higher phi_1 (more informative true-positive signals) also reduces welfare because any surplus from greater informativeness is fully captured by the platform.&lt;/p&gt;
&lt;p&gt;Q: What is the welfare ranking across the three regimes: no platform, advertising equilibrium, and subscription equilibrium?
A: In the subscription equilibrium (regime (a) of Proposition 2 or 4), user welfare for both types equals the no-platform base case W_base(tau) — the platform captures all surplus it creates and users are no better or worse off. In the advertising equilibrium (regime (b)), both naïve and sophisticated users are strictly worse off than with no platform: W-hat*(tau) &amp;lt; W_base(tau) for both tau in {S, N}. The first-best, where a planner controls ad loads separately by type, yields W^{FB}(tau) &amp;gt; W_base(tau) for both types because informative ads can genuinely improve sophisticated users&amp;rsquo; decisions and a constrained amount improves naïve users&amp;rsquo; decisions too.&lt;/p&gt;
&lt;p&gt;Q: How does firm-level competition interact with digital advertising to affect prices and welfare?
A: Without advertising, two ex ante identical firms compete à la Bertrand and price at marginal cost (p*_1 = p*_2 = c). Proposition 9 establishes that when phi_1 &amp;gt; phi^F_1 and phi_0 &amp;gt;= phi^F_0(phi_1), the platform offers an ad-based plan and equilibrium prices p-hat*_1 and p-hat*_2 are both strictly above p-bar* — the monopoly price without advertising. The mechanism is endogenous horizontal differentiation: users who see positive ad signals for one firm&amp;rsquo;s product form higher valuations for that product, so the two products become differentiated in the eyes of consumers even though they are ex ante identical, breaking Bertrand logic. Example 1 further illustrates that advertising can be more prevalent with competition than without: a second firm&amp;rsquo;s entry can push the equilibrium from no-advertising to separating.&lt;/p&gt;
&lt;p&gt;Q: Does platform competition protect users from the welfare losses associated with digital advertising?
A: Not fully. Proposition 11 shows that with two competing platforms (M=2, N=1) and no advertising, platforms compete away both subscription fees and ad loads, and welfare reaches the fully rational benchmark. However, when phi_1 exceeds threshold phi^P_1, both platforms adopt ad-based plans targeting naïve users, charge no subscription fees, and the product price rises to p-hat*_P &amp;gt; p-bar* (Proposition 12). Competition reduces subscription fees to zero but does not eliminate the incentive to target naïfs with heavy ads, because naïve users&amp;rsquo; over-valuation of ads means they remain willing to join ad-heavy plans. The fundamental inefficiency from naïve users&amp;rsquo; misspecified model persists under platform competition.&lt;/p&gt;
&lt;p&gt;Q: Why is the first-best allocation not implementable as a decentralized equilibrium?
A: Proposition 13 explains the obstacle: the social planner would ideally offer naïve users fewer ads (alpha^{FB}_N) than sophisticated users (alpha^{FB}_S), with alpha^{FB}_N &amp;lt;= alpha^{FB}_S. However, naïve users have a higher subjective valuation for ads than sophisticates because they believe ads are more informative. If offered a menu with both options, naïve users would self-select into the plan with the higher ad load alpha^{FB}_S — the exact opposite of what the planner wants. The incentive-compatibility constraints therefore force the planner toward a single pooling plan with an intermediate ad load alpha^{SB} in [alpha^{FB}_N, alpha^{FB}_S]. Average welfare under the second best exceeds the no-platform baseline, confirming that some advertising is socially valuable, but falls short of the first best whenever alpha^{FB}_N &amp;gt; 0.&lt;/p&gt;
&lt;p&gt;Q: How does a flat digital ad tax improve welfare, and what are its limitations?
A: Proposition 15 establishes that whenever the equilibrium features an ad-based plan, a flat tax on digital ad revenues at rate gamma &amp;gt; gamma-bar &amp;lt; 1 improves welfare by discouraging advertising-based business models and inducing the platform to shift toward subscription-based plans. The mechanism is that taxing ad revenue reduces the platform&amp;rsquo;s marginal gain from increasing ad load, making the subscription plan relatively more profitable. However, the flat tax does not achieve the second best because it operates linearly rather than targeting the nonlinear distortion: the optimal nonlinear tax-subsidy scheme (Proposition 14) requires a threshold-style ad tax at rate mu &amp;gt; mu-bar combined with a per-unit product subsidy delta* and a platform subscription subsidy eta &amp;gt; eta-bar.&lt;/p&gt;
&lt;p&gt;Q: What happens when the platform can endogenously choose how manipulative its ads are?
A: Proposition 16 shows that a profit-maximizing platform always chooses the lowest feasible phi_{0,N} = phi-bar — the most manipulative environment. Two reinforcing channels drive this: the pricing channel (lower phi_{0,N} amplifies naïve demand shifts per positive signal, so the downstream firm raises price and sales, increasing ad revenues extracted by the platform) and the participation channel (lower phi_{0,N} raises naïve users&amp;rsquo; perceived informational value of ads, relaxing their participation constraint and permitting a higher ad load alpha). Platform competition constrains the equilibrium ad load through tighter participation constraints but does not alter the choice of phi_{0,N} = phi-bar, so competition limits ad quantity but not ad manipulativeness.&lt;/p&gt;
&lt;p&gt;Q: How do naïve learning dynamics affect the main results?
A: Proposition 17 introduces a birth-death environment where exposure to disconfirming evidence gradually converts naïve agents to sophisticates. A unique steady-state sophisticate share lambda*(alpha_N, phi_0) exists; both higher ad load alpha_N and higher phi_0 accelerate the conversion of naïfs, raising future sophisticate share and reducing future ad revenues. This creates a new intertemporal trade-off that constrains the platform&amp;rsquo;s choice of ad loads relative to the static case. The key result (part ii) is that the main characterization of Proposition 7 carries through under a modified cutoff phi-tilde^{dynamic}&lt;em&gt;0 &amp;gt;= phi-tilde_0(lambda-tilde, phi_1, phi&lt;/em&gt;{0,N}), so learning dynamics make the ad-based business model less likely but do not overturn the fundamental welfare results.&lt;/p&gt;
&lt;p&gt;Q: How does imperfect price discrimination by the firm affect naïve users?
A: Proposition 18 considers a firm that observes a user&amp;rsquo;s sophistication type with probability kappa in [0,1]. With price discrimination, the firm sets type-specific prices satisfying p*_N &amp;gt;= p* &amp;gt;= p*_S, moving toward the type-specific monopoly levels. Naïfs are unambiguously worse off: when identified (with probability kappa), they face the higher price p*_N and a higher equilibrium ad load. The threshold for the advertising equilibrium also shifts down relative to the baseline, meaning advertising business models emerge for a larger parameter range when price discrimination is possible.&lt;/p&gt;
&lt;p&gt;Q: How does the paper define and measure user welfare, and why is ex post rather than interim welfare the relevant concept?
A: User welfare W(tau_i) is defined as ex post utility, which depends on the actual product quality theta_i realized after consumption, not on interim beliefs formed after viewing ads. Naïve users&amp;rsquo; interim assessment inflates expected product quality, but their ex post utility depends on whether the product is genuinely high quality for them (theta_i = 1 with probability q, theta_i = 0 with probability 1-q). Because naïve users over-purchase due to misread signals — consuming more than optimal when theta_i = 0 — their ex post utility is strictly lower than their interim expected utility, and strictly lower than the no-platform baseline in the advertising equilibrium. The ex post welfare concept is the relevant one precisely because it captures the actual material consequences of manipulation, not the subjectively perceived gains from ads.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Naïve vs. Sophisticated Users&lt;/strong&gt;: The paper&amp;rsquo;s primary user heterogeneity dimension. Sophisticated users hold the correct model of the ad signal process, setting phi_{0,S} = phi_0 (the true false-positive rate). Naïve users hold a misspecified model with phi_{0,N} = omega_N * omega_P * phi_0 &amp;lt; phi_0, underestimating the probability that a low-quality product generates a positive ad signal, due to inherent naïveté (omega_N) and failure to understand personalized targeting (omega_P).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ad Load (alpha)&lt;/strong&gt;: The Poisson rate at which ads are displayed to a user per unit time. Total ad displays follow a Poisson(alpha*T) distribution. Higher ad load means less time on entertaining content — expected entertainment time is (1-alpha)&lt;em&gt;T — and a higher probability (1 - exp(-alpha&lt;/em&gt;T)) that the user sees the ad at least once. The platform chooses alpha as its primary instrument for extracting surplus from naïve users.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;False-Positive Rate (phi_0)&lt;/strong&gt;: The objective probability that a low-quality product (theta_i = 0) generates a positive (&amp;ldquo;good&amp;rdquo;) ad signal. The gap between phi_0 (objective) and phi_{0,N} (naïve users&amp;rsquo; perceived rate) is the key parameter driving all welfare results: a larger gap implies greater de facto manipulation and a stronger incentive for the platform to adopt an advertising-based model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Berk-Nash Equilibrium&lt;/strong&gt;: The solution concept from Esponda and Pouzo (2016), used to model agents with misspecified subjective models. All agents are Bayesian conditional on their own subjective model. Sophisticates&amp;rsquo; subjective model equals the objective model (standard Bayesian), while naïfs update using the misspecified phi_{0,N}. Perfection requires sequential rationality at each information set given beliefs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;De Facto Manipulation&lt;/strong&gt;: The paper&amp;rsquo;s term for a situation in which the platform and firm exploit naïve users&amp;rsquo; misspecified model to boost demand and extract surplus, without requiring any outright deception in the formal sense. It arises because naïve users voluntarily choose high-ad-load plans (believing ads to be highly informative) and voluntarily over-purchase (having updated on what they mistakenly think are strong positive signals). The manipulation is &amp;ldquo;de facto&amp;rdquo; because it operates through the users&amp;rsquo; own rational (but misspecified) decision-making.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Separating Equilibrium&lt;/strong&gt;: An equilibrium in which naïve and sophisticated users self-select into distinct platform plans. In the advertising equilibrium, naïve users join an ad-heavy plan (extracting all their surplus via inflated willingness to pay for ads) while sophisticated users are either excluded or placed on a subscription plan. This separation is the vehicle through which the platform maximizes revenue from naïf manipulation while limiting the disciplining force of sophisticates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-Best Allocation&lt;/strong&gt;: The welfare-maximizing allocation subject to the incentive-compatibility constraints that users self-select into plans. Because naïve users prefer more ads than sophisticated users (the inverse of what the planner desires), the second best is a single pooling plan with an intermediate ad load alpha^{SB} in [alpha^{FB}_N, alpha^{FB}_S]. This is strictly worse than the first best but achieves average welfare above the no-platform baseline, and can be decentralized with a nonlinear ad tax, product subsidy, and platform subscription subsidy.&lt;/p&gt;</description></item><item><title>Open Rule Legislative Bargaining</title><link>https://macropaperwarehouse.com/papers/open-rule-legislative-bargaining/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/open-rule-legislative-bargaining/</guid><description>&lt;p&gt;This paper revisits the open rule legislative bargaining model of Baron and Ferejohn (1989) — the dominant workhorse model in political economy for analyzing how legislatures divide a surplus — and provides a more complete characterization of its stationary equilibria. The core research question is whether the equilibrium typically cited in the literature as the &amp;ldquo;open rule equilibrium&amp;rdquo; is actually the unique equilibrium, or whether it rests on implicit and unstated assumptions that, once relaxed, reveal a much richer equilibrium set.&lt;/p&gt;
&lt;p&gt;The model features n=3 negotiators dividing a surplus normalized to one, operating under simple majority rule (2 of 3 votes required). The common discount factor is Delta in (0,1). In each period, a proposer is selected uniformly at random; under the open rule, an amender is then selected uniformly at random from the two non-proposers and may either accept or counter-propose. Sincere voting determines the outcome. The authors analyze stationary subgame perfect equilibria (SSPE), in which strategies depend only on current role, not history.&lt;/p&gt;
&lt;p&gt;The existing literature implicitly adopted what the authors call the &amp;ldquo;standard assumption&amp;rdquo;: when given the opportunity to amend, the amender proposes the same allocation she would propose as a proposer in a closed rule game. Under this assumption, the unique SSPE has the proposer receiving share 1-Delta and each of the other two negotiators receiving Delta/2 (in the Pareto-efficient equilibrium). The literature treated this as the definitive open rule solution.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s first main result is that this standard-assumption equilibrium is indeed a valid SSPE, but it is not the only one. The key mechanism generating multiplicity is the treatment of off-path behavior: what the amender does when the proposer deviates to a non-equilibrium proposal. With n=3, a deviating proposer can exploit the structure so that the amender becomes a &amp;ldquo;free&amp;rdquo; coalition member — the proposer does not need to buy the amender&amp;rsquo;s vote separately, because the amender is already included in the majority once she counter-proposes. This expands the set of credible threats and supports a continuum of additional Pareto-undominated SSPEs.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s second main result characterizes the broader equilibrium set: all Pareto-undominated SSPEs belong to a class in which the proposer offers (1-Delta) to herself and equal shares to both other negotiators. In the non-standard equilibria, the amender always amends, generating equilibrium delay — agreements are not reached immediately, and payoffs are discounted by Delta^(t-1) for each period of delay.&lt;/p&gt;
&lt;p&gt;The third main result is that among all Pareto-undominated SSPEs, the unique Pareto-efficient one is the standard-assumption equilibrium (no delay). All other equilibria involve delay and are therefore Pareto-inferior in expectation.&lt;/p&gt;
&lt;p&gt;The institutional design implication reverses a widely held view: the open rule was thought to promote more egalitarian allocations relative to the closed rule. The authors show this is not the case for Pareto-efficient equilibria. The Pareto-efficient open rule equilibrium is actually a special case of the closed rule equilibrium — the proposer captures 1-Delta and offers Delta to the coalition. More broadly, open rule bargaining tends to generate longer equilibrium delays and less egalitarian surplus allocations than previously predicted by Baron and Ferejohn. Scope conditions: the formal analysis is restricted to n=3 negotiators; generalization to larger legislatures is noted as an open direction.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;standard assumption&amp;rdquo; and why does the existing literature rely on it?&lt;/p&gt;
&lt;p&gt;A: The standard assumption holds that when an amender gets the opportunity to counter-propose, she proposes the same allocation she would choose if she were the proposer in a closed rule game. The existing open rule literature — including Baron and Ferejohn (1989), Jackson and Morelli (2004), Baron (2012), van Weelden (2013), and Austen-Smith and Banks (1999) — accepted this assumption implicitly, treating the resulting equilibrium as the unique open rule equilibrium. The assumption sidesteps the question of off-path behavior: what happens when the proposer deviates to a non-equilibrium proposal that the amender would want to amend. Because deviations are resolved within the same bargaining session under the open rule, off-path specifications are consequential.&lt;/p&gt;
&lt;p&gt;Q: What is the unique SSPE under the standard assumption, and what are its payoff implications?&lt;/p&gt;
&lt;p&gt;A: Under the standard assumption with n=3 and discount factor Delta, the unique SSPE has the proposer receiving a share of 1-Delta of the surplus and each of the other two negotiators receiving Delta/2. There is no delay: the proposal passes immediately in the period it is made. This equilibrium is Pareto-efficient relative to all other stationary equilibria identified in the paper.&lt;/p&gt;
&lt;p&gt;Q: What is the mechanism by which the equilibrium set is larger than the standard assumption predicts?&lt;/p&gt;
&lt;p&gt;A: With n=3, when a proposer deviates to a non-equilibrium proposal, the amender — who responds by counter-proposing — automatically becomes part of the passing coalition without the proposer needing to separately compensate her. This makes the amender a &amp;ldquo;free&amp;rdquo; coalition member in the deviation subgame, which changes the cost structure of deviations and expands the range of proposals the proposer can credibly make. Consequently, a wider set of strategies by the amender can be sustained as equilibrium responses, yielding a continuum of additional Pareto-undominated SSPEs beyond the standard-assumption equilibrium.&lt;/p&gt;
&lt;p&gt;Q: What do the non-standard equilibria look like in terms of proposals, delay, and payoffs?&lt;/p&gt;
&lt;p&gt;A: In the non-standard Pareto-undominated SSPEs, the proposer offers (1-Delta) to herself and equal shares (Delta/2 each) to the other two negotiators — note the proposer&amp;rsquo;s own share is the same as in the standard equilibrium, but the off-path behavior differs — and the amender always chooses to amend rather than accept. The amendment triggers a vote in which the amendment fails (or the process repeats), pushing resolution to the next period. This generates equilibrium delay: agreements take multiple periods to reach, and all payoffs are discounted by Delta^(t-1) per period of delay, making these equilibria Pareto-inferior to the no-delay equilibrium.&lt;/p&gt;
&lt;p&gt;Q: Which equilibrium is Pareto-efficient among all Pareto-undominated SSPEs, and why?&lt;/p&gt;
&lt;p&gt;A: The unique Pareto-efficient SSPE is the standard-assumption equilibrium, because it is the only one that involves no delay. All other Pareto-undominated SSPEs involve at least one period of delay, which destroys surplus through discounting (payoffs shrink by a factor of Delta per period). Since delay is costly for all negotiators and generates no compensating redistribution, any equilibrium with delay is Pareto-dominated by the no-delay equilibrium.&lt;/p&gt;
&lt;p&gt;Q: What are the implications for the classic efficiency comparison between open and closed rules?&lt;/p&gt;
&lt;p&gt;A: The closed rule always generates an efficient outcome (no delay in SSPE). The open rule can also generate an efficient outcome — under the standard-assumption equilibrium — but uniquely admits a continuum of inefficient equilibria involving delay. Therefore the open rule is weakly dominated by the closed rule from an efficiency standpoint: at best it matches the closed rule (one efficient equilibrium), and at worst it generates costly delay. This reverses the common inference that open rule unambiguously improves outcomes.&lt;/p&gt;
&lt;p&gt;Q: What are the implications for the classic fairness comparison between open and closed rules?&lt;/p&gt;
&lt;p&gt;A: The open rule was commonly believed to promote more egalitarian surplus divisions relative to the closed rule, which allows the proposer to extract a large share. The paper shows this view is misleading. In the Pareto-efficient open rule equilibrium, the proposer still captures 1-Delta — the same as under the closed rule — and the result is no more egalitarian. In the delay equilibria, the proposer does offer equal shares to both other negotiators, but this comes at the cost of inefficiency (delay). There is no Pareto-undominated open rule equilibrium that is both efficient and more egalitarian than the closed rule.&lt;/p&gt;
&lt;p&gt;Q: What is the class of &amp;ldquo;Pareto-undominated stationary strategies&amp;rdquo; and why does the paper focus on it?&lt;/p&gt;
&lt;p&gt;A: A stationary strategy profile is Pareto-undominated if no other stationary strategy profile gives every negotiator at least as high an expected payoff with at least one strictly better off. The paper focuses on this class to provide a tractable but principled selection criterion within the large set of SSPEs: it eliminates equilibria that are dominated from every player&amp;rsquo;s perspective, retaining only those that could plausibly arise if players coordinate on mutually beneficial outcomes. The characterization of this class reveals that equilibrium multiplicity is already substantial even after imposing this selection.&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the formal results, and what is left open?&lt;/p&gt;
&lt;p&gt;A: The formal analysis is restricted to n=3 negotiators with simple majority rule (2 of 3 votes). The authors acknowledge that generalization to larger n is an important open question. The three-legislator case is the simplest non-trivial instance of the majority-rule bargaining problem, and the authors use it to isolate the mechanism cleanly. The model assumes sincere voting, a common discount factor Delta in (0,1), and stationary strategies.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to Baron and Ferejohn (1989)?&lt;/p&gt;
&lt;p&gt;A: Baron and Ferejohn (1989) originated both the closed rule and open rule bargaining frameworks and derived the standard-assumption equilibrium for the open rule. Subsequent literature (Eraslan 2002, Cho and Duggan 2003, 2009, Banks and Duggan 2000) extended various aspects of the B&amp;amp;F framework. The present paper takes the B&amp;amp;F open rule model as given but demonstrates that B&amp;amp;F&amp;rsquo;s open rule analysis was incomplete: it did not systematically address off-path behavior, and as a result the equilibrium it identified is not unique. The paper&amp;rsquo;s main contribution is to show that the B&amp;amp;F open rule predictions — more egalitarian allocations and prompt agreement — do not hold generally across the full equilibrium set.&lt;/p&gt;
&lt;p&gt;Open Rule: A bargaining protocol in which, after an initial proposal is made, a nominated amender may make a counter-proposal before a vote is taken; contrasted with the closed rule, under which the initial proposal is voted on without amendment.&lt;/p&gt;
&lt;p&gt;Closed Rule: A bargaining protocol in which a vote is taken directly on the first proposal, with no opportunity for amendment.&lt;/p&gt;
&lt;p&gt;Standard Assumption: The implicit assumption, used by Baron and Ferejohn (1989) and subsequent literature, that when the amender counter-proposes under the open rule, she proposes the same allocation she would choose as a proposer in a closed rule game; the paper shows this assumption is consequential for equilibrium uniqueness.&lt;/p&gt;
&lt;p&gt;Stationary Subgame Perfect Equilibrium (SSPE): An equilibrium concept in which each player&amp;rsquo;s strategy depends only on her current role (proposer, amender, or voter) and not on the history of play; the paper characterizes SSPEs of the open rule model.&lt;/p&gt;
&lt;p&gt;Pareto-Undominated Stationary Strategy Profile: A stationary strategy profile for which no other stationary strategy profile gives every negotiator weakly higher expected payoff with at least one strictly higher; used as a selection criterion to prune the large equilibrium set.&lt;/p&gt;
&lt;p&gt;Equilibrium Delay: The phenomenon in which agreement is not reached in the current period because the amender always counter-proposes and the counter-proposal also fails, pushing resolution to a future period and discounting payoffs; all non-standard-assumption Pareto-undominated SSPEs involve delay.&lt;/p&gt;
&lt;p&gt;Off-Path Behavior: The specification of what strategies players use following a deviation from equilibrium play; the paper shows that different specifications of off-path behavior by the amender support different equilibria, and that the existing literature was not systematic about this.&lt;/p&gt;</description></item><item><title>Optimal Decision Rules When Payoffs are Partially Identified</title><link>https://macropaperwarehouse.com/papers/optimal-decision-rules-when-payoffs-are-partially-identified/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-decision-rules-when-payoffs-are-partially-identified/</guid><description>&lt;p&gt;This paper derives asymptotically optimal statistical decision rules for discrete choice problems when the payoffs associated with some choices are only partially identified. The research question is: how should a decision maker who can bound but not point-identify a payoff-relevant parameter θ use data to make optimal policy choices?&lt;/p&gt;
&lt;p&gt;The framework separates two parameter types. The reduced-form parameter µ is point-identified and can be estimated from data. The structural parameter θ — such as the average treatment effect (ATE) in a target population — is set-identified, meaning only that θ ∈ Θ0(µ) can be established, where the identified set is indexed by µ. The decision maker confronts both ambiguity (arising from partial identification of θ given µ) and statistical uncertainty (µ must be estimated).&lt;/p&gt;
&lt;p&gt;The authors propose a hybrid optimality criterion that applies minimax reasoning to the partially-identified parameter θ — choosing actions that minimize maximum risk over Θ0(µ) — while applying average (integrated) risk minimization over µ, reflecting the asymmetric nature of the two identification problems. This asymmetric treatment follows the generalized Bayes-minimax principle of Hurwicz (1951).&lt;/p&gt;
&lt;p&gt;The optimal decision rule is implemented by computing, for each action, the maximum risk (or regret) over θ ∈ Θ0(µ) conditional on µ, then averaging this maximum risk across either (i) a bootstrap distribution for an efficient estimator µ̂, (ii) a posterior distribution for µ in parametric models, or (iii) a quasi-posterior based on a limited-information criterion in semiparametric models. The optimal action is whichever choice has the smallest average maximum risk.&lt;/p&gt;
&lt;p&gt;A central theoretical result (Theorems 1 and 4) establishes formal asymptotic optimality for both parametric and semiparametric settings: Bayes and quasi-Bayes decisions with any prior whose density is positive, bounded, and continuous are asymptotically equivalent and optimal. Critically, the optimality of these rules is asymptotically independent of the choice of prior for µ. The authors also establish a necessity result (Theorems 2 and 5): any decision rule not asymptotically equivalent to the Bayes or bootstrap rule is strictly sub-optimal.&lt;/p&gt;
&lt;p&gt;A key finding is that &amp;ldquo;plug-in&amp;rdquo; rules — which substitute an efficient point estimate µ̂ directly into the oracle decision rule — can be sub-optimal. This failure occurs generically under partial identification because the maximum risk function R(d,µ) is typically only directionally differentiable (not fully differentiable) in µ, owing to max and min operators in intersection bounds, linear program value functions, or other bound constructions. When full differentiability holds, Corollary 1 confirms plug-in rules are optimal; otherwise they are not. The empirical illustration demonstrates the practical consequence: for German male youths deciding whether to adopt a job-training program based on 14 RCT studies from Card, Kluve, and Weber (2017), the optimal rule recommends treatment (average quasi-posterior robust welfare contrast b̄n &amp;gt; 0) while the plug-in rule recommends against treatment (plug-in value b(µ̂) &amp;lt; 0). The lower bound maximum of µ̂k − C‖x0 − xk‖ is −0.3190 for the leading US study and −0.3298 for the second-best Brazilian study; because these two values are close relative to the average standard error of 0.034 across studies, the lower bound distribution is right-skewed (behaving like the maximum of two Gaussians), pushing b̄n positive even though b(µ̂) is negative.&lt;/p&gt;
&lt;p&gt;The paper extends optimality theory to semiparametric models via a least favorable parametric submodel, introduces the concept of σ-optimality for cases where the average maximum risk criterion is infinite (relevant when the dimension K of µ exceeds 1), and provides detailed implementation guides for treatment assignment under intersection bounds, IV-like estimands, and non-separable panel data, as well as for optimal pricing decisions where revealed-preference demand theory bounds counterfactual demand responses via linear programming.&lt;/p&gt;
&lt;p&gt;Scope conditions: optimality results apply to discrete action spaces, require efficient estimation of µ, require the identified set Θ0(µ) to be known as a set-valued mapping, and assume no &amp;ldquo;first-order ties&amp;rdquo; (the oracle decision is unique at µ0). The asymptotic framework is local, mimicking the finite-sample problem where µ is not known with certainty.&lt;/p&gt;
&lt;p&gt;Q: What is the core decision problem this paper addresses?&lt;/p&gt;
&lt;p&gt;A: A decision maker must choose from a finite set of actions D = {0, 1, &amp;hellip;, D}. Payoffs depend on a structural parameter θ that is only set-identified — the data can establish θ ∈ Θ0(µ) but not pin down θ exactly. The reduced-form parameter µ is point-identified and estimated from data. The decision maker faces both ambiguity (which θ in Θ0(µ) is true?) and sampling uncertainty (what is µ?). The paper asks how to construct decision rules that are optimal in large samples under this dual uncertainty.&lt;/p&gt;
&lt;p&gt;Q: What is the proposed optimality criterion, and why is it asymmetric across parameters?&lt;/p&gt;
&lt;p&gt;A: The criterion applies minimax reasoning to the partially-identified θ — the maximum risk over Θ0(µ) given µ is the relevant loss — and integrates this maximum risk over µ using Lebesgue measure on local perturbations h = √n(µ − µ0) of a fixed µ0. The asymmetry reflects the fact that θ is not updated by the data (the prior for θ is not identified), while µ can be learned efficiently from the data. Full minimax over both (θ, µ) is rarely tractable even for simple binary treatment problems; the asymmetric approach yields tractable optimal rules for a broad empirically relevant class of settings.&lt;/p&gt;
&lt;p&gt;Q: What are the Bayes, bootstrap, and quasi-Bayes implementations of the optimal rule?&lt;/p&gt;
&lt;p&gt;A: In all three cases, the decision maker computes R̄n(d) — the average maximum risk for action d — and chooses the action that minimizes it. The Bayes rule averages R(d, µ) over the posterior πn(µ|Xn) for µ using Bayes&amp;rsquo; theorem with a prior π on M. The bootstrap rule averages R(d, µ̂*) over bootstrap redraws µ̂* of the efficient estimator µ̂. The quasi-Bayes rule (for semiparametric models) uses a limited-information quasi-posterior N(µ̂, (nÎ)−1) combining a Gaussian quasi-likelihood with a prior for µ. All three implementations are asymptotically equivalent and optimal under the regularity conditions of Theorems 1 and 4.&lt;/p&gt;
&lt;p&gt;Q: What do Theorems 1 and 2 (and their semiparametric analogues Theorems 4 and 5) establish?&lt;/p&gt;
&lt;p&gt;A: Theorem 1 establishes sufficiency: Bayes decisions with any prior in the class Π are asymptotically equivalent to each other and are optimal; any rule asymptotically equivalent to such a Bayes decision is also optimal. Theorem 2 establishes necessity: any rule in the admissible class D that is not asymptotically equivalent to the Bayes rule has strictly higher average excess risk at any µ0 where asymptotic equivalence fails. Together, these theorems fully characterize the class of asymptotically optimal rules and show that the Bayes/bootstrap class is not merely sufficient but also necessary for optimality.&lt;/p&gt;
&lt;p&gt;Q: When are plug-in rules sub-optimal, and when are they optimal?&lt;/p&gt;
&lt;p&gt;A: Plug-in rules substitute an efficient point estimate µ̂ directly into the oracle decision δo(µ̂). If R(d, µ) is fully differentiable at µ0 for all oracle-optimal actions d, then the directional derivative is linear and plug-in and Bayes rules are asymptotically equivalent; Corollary 1 confirms plug-in rules are then optimal. However, under partial identification, max and min operators in bound constructions — intersection bounds, linear program value functions, revealed-preference bounds — generically induce only directional (non-linear) differentiability of R(d, µ). In these cases asymptotic equivalence can fail, and Theorem 2 implies plug-in rules are sub-optimal. Manski (2021, 2023) documents poor finite-sample performance of plug-in rules numerically; the authors&amp;rsquo; necessity result provides a general theoretical explanation under the asymptotic average risk criterion.&lt;/p&gt;
&lt;p&gt;Q: How does the treatment assignment empirical illustration demonstrate the difference between optimal and plug-in rules?&lt;/p&gt;
&lt;p&gt;A: Using data from Ishihara and Kitagawa (2021) with K = 14 RCT studies from Card, Kluve, and Weber (2017) and Lipschitz constant C = 0.25, the decision is whether to adopt a job-training program for German male youths or female youths in 2010 (GDP growth 3.48%, unemployment 9.45%). For male youths, the largest lower bound value µ̂k − C‖x0 − xk‖ is −0.3190 (US study) and the second-largest is −0.3298 (Brazilian study), separated by only 0.0108 against an average standard error of 0.034 across studies, so the lower bound distribution is right-skewed (maximum of two near-tied Gaussians). This right-skew pushes the quasi-posterior mean b̄n positive, yielding a treatment recommendation, while the plug-in value b(µ̂) is negative, yielding a non-treatment recommendation — a concrete reversal of the policy decision. For female youths, the minima and maxima are better separated, the distribution is near-Gaussian, and b̄n ≈ b(µ̂), so both rules agree on treatment.&lt;/p&gt;
&lt;p&gt;Q: What are intersection bounds and why do they generate directional differentiability?&lt;/p&gt;
&lt;p&gt;A: Intersection bounds arise when the ATE is bounded in K separate observational studies by lower bounds bL,k(µk) and upper bounds bU,k(µk). The combined identified set uses bL(µ) = max_{1≤k≤K} bL,k(µk) and bU(µ) = min_{1≤k≤K} bU,k(µk). Even if each component bound is smooth in µk, the max and min operators make bL and bU only directionally differentiable (not fully differentiable) in µ. The directional derivative is positively homogeneous of degree one but non-linear, which is the property that drives the wedge between Bayes and plug-in rules.&lt;/p&gt;
&lt;p&gt;Q: How does the paper extend to semiparametric models, and what technical tool does it use?&lt;/p&gt;
&lt;p&gt;A: In semiparametric models, the data distribution depends on both µ ∈ R^K and an infinite-dimensional nuisance parameter η. Integrating over local perturbations of η as well as µ raises measure-theoretic problems in infinite-dimensional spaces. The authors instead restrict attention to local perturbations of µ0 within a least favorable parametric submodel, which is the direction that makes the problem hardest. The quasi-posterior N(µ̂, (nÎ)−1) is then used as the averaging distribution, combining a Gaussian quasi-likelihood with a prior for µ. Theorem 4 establishes optimality and Theorem 5 establishes necessity under these semiparametric conditions, mirroring the parametric Theorems 1 and 2.&lt;/p&gt;
&lt;p&gt;Q: What is σ-optimality and why is it needed?&lt;/p&gt;
&lt;p&gt;A: When the dimension K of µ exceeds 1, the integrated average excess risk criterion R({δn}; µ0) — which integrates over Lebesgue measure on R^K — may be infinite for all decision sequences in D, making the criterion uninformative. σ-optimality approximates the improper Lebesgue prior on h by a sequence of proper priors indexed by σ, and requires that the decision rule minimize the resulting criterion for all σ. Theorem 3 shows that the limiting behavior of σ-optimal rules coincides with that of the Bayes rule δ*n(·; π), preserving the practical implementation.&lt;/p&gt;
&lt;p&gt;Q: How is the optimal pricing application structured and what role do revealed-preference bounds play?&lt;/p&gt;
&lt;p&gt;A: A monopolist observes repeated cross-sections of individual demands across B budget sets and must choose a price vector from D = O ∪ C, where O contains observed prices and C contains counterfactual prices. For observed prices, average demand is identified; for counterfactual prices, only bounds are available. Following Kitamura and Stoye (2019), the space of goods is partitioned into GARP-compatible regions, and sharp bounds on counterfactual demand are computed by solving linear programs over the mass allocated to each region subject to GARP consistency constraints. The reduced-form parameter µ collects empirical choice probabilities across observed budget-region cells, estimated consistently by sample frequencies. The optimal pricing decision averages the linear-program bound solutions across quasi-posterior draws of µ.&lt;/p&gt;
&lt;p&gt;Q: How does this approach relate to minimax and conditional Γ-minimax approaches?&lt;/p&gt;
&lt;p&gt;A: Full minimax over (θ, µ) requires strong distributional assumptions and tractable finite-sample distributions; the authors note that no minimax treatment rule exists even for binary treatment with binary outcomes and estimated bounds. Conditional Γ-minimax (DasGupta and Studden, 1989; Giacomini, Kitagawa, and Read, 2021) fixes a prior for µ and takes minimax over the set of priors for θ conditional on µ; this is closely related to the authors&amp;rsquo; approach but can be conservative when the marginal prior for µ varies. The authors&amp;rsquo; framework fixes the marginal prior for µ and takes minimax over θ ∈ Θ0(µ) conditional on µ, which is shown to arise as the equilibrium of a two-player zero-sum game where adversarial nature chooses a prior for θ ∈ Θ0(µ) conditional on µ and the available data for µ.&lt;/p&gt;
&lt;p&gt;Q: What is the technical contribution regarding directionally differentiable functions?&lt;/p&gt;
&lt;p&gt;A: Hirano and Porter (2009) derived asymptotic optimality for treatment rules under fully differentiable welfare contrasts. This paper extends that theory to settings with directional (but not full) differentiability — a generic feature whenever bounds involve max/min operators or linear program values. The key technical building block is the asymptotic distribution of the quasi-posterior mean of directionally differentiable functions (Propositions 2 and 3 in Appendix C). While Kitagawa, Montiel Olea, Payne, and Velez (2020) characterized the asymptotic behavior of the posterior distribution of such functions, this paper instead characterizes the frequentist distribution of the posterior mean — a distinct and novel contribution to the literature on asymptotics for non-smooth functions (Dümbgen, 1993; Fang and Santos, 2019).&lt;/p&gt;
&lt;p&gt;Q: What are the key scope conditions and limitations of the optimality results?&lt;/p&gt;
&lt;p&gt;A: The action space D must be finite and discrete (continuous pricing must be approximated by a grid of whole-currency units, as noted in the introduction). The identified set mapping Θ0(·) must be known. Efficient estimation of µ is required, along with a consistent estimator of its asymptotic variance for quasi-Bayes implementation. The optimality criterion assumes &amp;ldquo;no first-order ties&amp;rdquo; — the oracle decision must be unique at µ0. The framework is asymptotic (local perturbations around a fixed µ0), and the theory is designed for settings where deriving exact finite-sample optimal rules is intractable. The results do not cover the case where θ affects the data distribution (only payoffs are partially identified, not identification of µ itself).&lt;/p&gt;
&lt;p&gt;Partially-identified parameter (θ): A structural parameter — such as the ATE in a target population — about which the data can establish only set membership θ ∈ Θ0(µ), not a point value. The identified set Θ0(µ) is indexed by the point-identified reduced-form parameter µ.&lt;/p&gt;
&lt;p&gt;Oracle decision (δo(µ)): The infeasible first-best decision that minimizes maximum risk over the identified set Θ0(µ) for a known value of µ. It serves as the benchmark against which practical rules are evaluated; any data-dependent rule can only do weakly worse.&lt;/p&gt;
&lt;p&gt;Maximum risk (R(d, µ)): The supremum of risk r(d, θ, µ) = Eθ[l(d, Y, θ, µ)] over all θ ∈ Θ0(µ) conditional on µ. Under the regret criterion for binary treatment, R(0, µ) = (bU(µ))+ and R(1, µ) = −(bL(µ))−.&lt;/p&gt;
&lt;p&gt;Robust welfare contrast (b(µ)): In the treatment assignment application, b(µ) = (bU(µ))+ + (bL(µ))−, whose sign determines the oracle decision: treat if b(µ) ≥ 0. The optimal rule replaces b(µ) with its quasi-posterior mean b̄n.&lt;/p&gt;
&lt;p&gt;Directional differentiability: A function f : M → R^k is directionally differentiable at µ0 if limits of (f(µ0 + tn hn) − f(µ0))/tn exist for all sequences tn ↓ 0 and hn → h, yielding a directional derivative ḟµ0[·] that is positively homogeneous but not necessarily linear. Max/min operators and linear program value functions are generically only directionally differentiable, not fully differentiable. This property is what causes plug-in rules to fail.&lt;/p&gt;
&lt;p&gt;Quasi-posterior: In semiparametric models, a posterior-like distribution for µ formed by combining a limited-information Gaussian quasi-likelihood N(µ̂, (nÎ)−1) with a prior π, yielding πn(µ|Xn) ∝ exp(−½(µ − µ̂)T(nÎ)(µ − µ̂))π(µ). Used in place of a full Bayesian posterior when the exact likelihood of the data-generating process is unavailable.&lt;/p&gt;
&lt;p&gt;σ-optimality: An optimality concept that replaces the improper Lebesgue prior on local perturbations h ∈ R^K with a sequence of proper priors indexed by σ, used when the average excess risk criterion is infinite for K &amp;gt; 1. Theorem 3 establishes that the σ-optimal decision rule converges to the Bayes rule as σ → ∞.&lt;/p&gt;
&lt;p&gt;Plug-in rule (δplug_n): A decision rule formed by substituting an efficient point estimate µ̂ directly into the oracle decision: δplug_n = δo(µ̂). Optimal when R(d, µ) is fully differentiable (Corollary 1), but generically sub-optimal under partial identification because directional differentiability of R(d, µ) breaks the asymptotic equivalence between the plug-in and Bayes rules.&lt;/p&gt;</description></item><item><title>Optimal Public Transportation Networks: Evidence from the World's Largest Bus Rapid Transit System in Jakarta</title><link>https://macropaperwarehouse.com/papers/optimal-public-transportation-networks-evidence-from-the-worlds-largest-bus-rapid-transit-system-in-jakarta/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-public-transportation-networks-evidence-from-the-worlds-largest-bus-rapid-transit-system-in-jakarta/</guid><description>&lt;p&gt;This paper studies how commuter preferences over wait times, travel times, and transfers should shape the design of urban bus networks, using the world&amp;rsquo;s largest Bus Rapid Transit (BRT) system — TransJakarta in Jakarta, Indonesia — as the empirical laboratory. The setting provides unusually rich identification: between January 2016 and February 2020, TransJakarta launched 93 new BRT and non-BRT feeder routes in a staggered, city-wide expansion, during which the operating bus fleet more than doubled from roughly 700 to over 1,600 vehicles. The authors combine over 500 million smart-card tap records, GPS tracking of every bus at 5–10 second intervals, and anonymized smartphone location data covering 35 million weekday trips from 2.3 million devices.&lt;/p&gt;
&lt;p&gt;The paper proceeds in three steps. First, the authors classify new route launches into three event types and estimate their causal impact on ridership via difference-in-differences. Event 1: a new direct connection between an origin-destination pair already served by transfer only, with no travel-time improvement — raises BRT ridership by 0.16 log points. Event 2: a new direct connection that also reduces travel time (by 0.29 log points on average) — raises ridership by 0.27 log points. Event 3: additional buses on an already-directly-connected pair, which increases the bus arrival rate by 0.32 log points and reduces wait times — raises ridership by 0.09 log points, implying a ridership elasticity with respect to wait times of approximately −0.29 for BRT. For non-BRT routes the implied wait-time elasticity is −1.05, raising the possibility of multiple equilibria in service levels. Crucially, none of the three event types produce detectable increases in aggregate trip volumes measured by smartphone data, implying the ridership gains reflect modal substitution toward the bus rather than trip generation.&lt;/p&gt;
&lt;p&gt;Second, the authors estimate a structural demand model. At its core is a route-choice model in which bus arrivals follow independent Poisson processes, so wait times are exponentially distributed and idiosyncratic. This formulation avoids the red-bus/blue-bus aggregation problem endemic to logit models. Commuters are also allowed to be partially inattentive to routes whose travel time exceeds the fastest available option by more than an estimated threshold. Structural parameters are recovered by classical minimum distance, matching seven reduced-form moments. Key findings: wait time is valued 2.4 times more than time on the bus for BRT routes, and 4.2 times more for non-BRT routes. There is no additional transfer penalty beyond the wait time and travel time costs of the second leg. Commuters pay significantly less attention to options with travel time more than roughly 34–44 percent above the fastest option in their choice set.&lt;/p&gt;
&lt;p&gt;Third, the authors use the estimated preference parameters to characterize optimal bus networks. Because the optimization problem is high-dimensional (418 grid cells, 1,536 possible edges, yielding on the order of 10^500 configurations) and exhibits neither global convexity nor simple complementarity, they reformulate the social planner&amp;rsquo;s problem as a discrete choice over networks with additive logit shocks — effectively sampling from a multinomial logit distribution via simulated annealing. The result: optimal networks cover approximately 66 percent of grid cells versus 42 percent under the actual TransJakarta network, and would give 91 percent of Jakarta residents bus access versus 73 percent currently. Bus frequency in the city center is somewhat lower in the optimal network. Despite commuters&amp;rsquo; high sensitivity to wait times, the current network concentrates too many buses in the city center where wait times are already short, rather than extending reach to underserved areas. Comparative statics show that doubling the wait-time cost parameter produces much more concentrated optimal networks (23 percent of origin-destination pairs connected, 41 percent fewer than baseline), while increasing the transfer penalty by the equivalent of 15 minutes of wait time raises the direct-connection share of served pairs from 12 to 16 percent.&lt;/p&gt;
&lt;p&gt;Q: What are the three event types and why are they analytically distinct?&lt;/p&gt;
&lt;p&gt;A: Event 1 is the launch of the first direct route between an origin-destination pair already connected by transfer, where the direct route is not faster than the existing transfer option; it isolates the effect of directness absent a travel-time change. Event 2 is the same but with a faster direct route (average reduction of 0.29 log points in travel time), combining directness and speed improvements. Event 3 is the launch of a new route that overlaps an existing direct route, increasing bus frequency and cutting wait times (arrival rate up 0.32 log points) without substantially changing travel time or directness. The three events together provide variation across the key dimensions — directness, speed, and frequency — needed to separately identify commuter preference parameters.&lt;/p&gt;
&lt;p&gt;Q: What are the main ridership effects and how large are they in levels?&lt;/p&gt;
&lt;p&gt;A: For BRT routes, Event 1 raises ridership by 0.16 log points (approximately 19 additional riders per week for a treated origin-destination pair with a baseline of 111 weekly riders), Event 2 by 0.27 log points (approximately 24 additional riders per week), and Event 3 by 0.09 log points (approximately 20 additional riders per week). For non-BRT routes, proportional effects are larger but level effects are similar: Event 1 yields roughly 34 additional weekly riders, Event 2 roughly 21, and Event 3 roughly 15. Event-study graphs show clear, discrete jumps in ridership at route launch with no pre-trends, and some gradual adjustment in the months following.&lt;/p&gt;
&lt;p&gt;Q: What does the paper find about aggregate trip generation versus modal substitution?&lt;/p&gt;
&lt;p&gt;A: Using smartphone location data to measure all trips regardless of mode, the authors find no statistically significant increase in aggregate trip volumes for any of the three event types. For BRT Event 1, the estimated aggregate-trip coefficient is −0.008 with a standard error of 0.051, allowing rejection at the 95 percent level of any positive impact above roughly 0.091 log points — small relative to the precise 0.11 log-point bus ridership effect in the same sample. The authors interpret this as evidence that the ridership gains over the 10-month post-event window reflect substitution from private modes (motorcycles, cars, taxis) toward TransJakarta rather than trip generation, and they use this null result to justify holding destination choices fixed in the structural model.&lt;/p&gt;
&lt;p&gt;Q: How does the model avoid the red-bus/blue-bus aggregation problem?&lt;/p&gt;
&lt;p&gt;A: The paper&amp;rsquo;s route-choice model assumes bus arrivals follow independent Poisson processes, so wait times are exponentially distributed. A key proposition (Proposition 1) proves that splitting one route into two identical routes with half the buses each produces exactly the same choice probabilities and expected utility as the original single route — because the sum of two independent Poisson processes is itself Poisson with the summed rate. Standard logit models fail this invariance because splitting a route creates two options with independent error draws, artificially inflating expected utility. The invariance property is essential for the optimal network design exercise, where the planner freely reallocates buses across routes.&lt;/p&gt;
&lt;p&gt;Q: What are the estimated preference parameters and what do they imply about commuter behavior?&lt;/p&gt;
&lt;p&gt;A: The paper estimates that wait time is valued 2.4 times more than time on the bus for BRT routes and 4.2 times more for non-BRT routes. There is no additional transfer disutility beyond the wait time and travel time costs implied by the extra leg. Commuters become substantially inattentive to routes with travel time more than approximately 34 percent above the fastest available option (BRT threshold) or 44 percent (non-BRT). The high relative cost of waiting versus riding reflects both the discomfort of waiting at exposed non-BRT stops and the fact that TransJakarta runs without a published schedule, so commuters cannot minimize wait time by timing arrivals.&lt;/p&gt;
&lt;p&gt;Q: What explains the non-BRT wait-time elasticity exceeding −1?&lt;/p&gt;
&lt;p&gt;A: For non-BRT routes, Event 3 raises ridership by 0.450 log points while raising the bus arrival rate by 0.425 log points, yielding an implied elasticity of ridership with respect to wait times of −1.05. Because the baseline arrival rate for non-BRT treated pairs is 2–4 times lower than for BRT pairs, the absolute reduction in wait time per additional bus is much larger. An elasticity exceeding −1 in absolute value implies that adding buses on some non-BRT routes could increase ridership enough to maintain or even raise average ridership per bus — the extreme form of the Mohring effect — suggesting the possibility of a high-ridership/low-wait-time equilibrium distinct from the current low-ridership/high-wait-time one.&lt;/p&gt;
&lt;p&gt;Q: How is the optimal network characterized and what algorithm is used?&lt;/p&gt;
&lt;p&gt;A: The social planner chooses a network to maximize utilitarian welfare (average expected utility across all commuters) from the estimated demand model, plus a network-level logit shock capturing cost and other factors outside the model. This transforms the combinatorially explosive optimization into sampling from a multinomial logit distribution over networks, which the authors approximate using simulated annealing. They run the algorithm multiple times to obtain a sample of networks drawn asymptotically from the planner&amp;rsquo;s distribution, then estimate optimal network characteristics and comparative statics from sample analogs. The theoretical framework is general and, the authors note, applicable to other high-dimensional spatial planning problems where welfare differences can be computed for pairs of counterfactuals.&lt;/p&gt;
&lt;p&gt;Q: How does the optimal network differ from the current TransJakarta network?&lt;/p&gt;
&lt;p&gt;A: The typical optimal network covers approximately 66 percent of 2km grid cells versus 42 percent for the actual network, and 91 percent of Jakarta residents would have bus access versus 73 percent currently. The optimal network reduces bus frequency in the city center relative to the current network, accepting longer wait times there in order to extend reach to peripheral areas. The paper finds no tension between distributional and efficiency concerns in this setting — expanding coverage improves both aggregate welfare and access for underserved areas.&lt;/p&gt;
&lt;p&gt;Q: What do the comparative statics reveal about the sensitivity of optimal network design to preference parameters?&lt;/p&gt;
&lt;p&gt;A: Doubling the wait-time cost parameter leads to substantially more concentrated optimal networks: only 23 percent of origin-destination pairs are connected, 41 percent fewer than in the baseline optimal network. This is because higher wait-time costs make it more valuable to concentrate buses on fewer routes to achieve short headways. Increasing the transfer penalty by the equivalent of 15 minutes of wait time raises the share of connected location pairs with a direct (non-transfer) connection from 12 to 16 percent. These comparative statics link micro-level preference parameters to macro-level network topology, clarifying which parameters most influence design choices.&lt;/p&gt;
&lt;p&gt;Q: How does the paper validate the destination imputation from tap-in-only smart card data?&lt;/p&gt;
&lt;p&gt;A: For the subset of BRT stations where tap-out is enforced (36 percent of stations), the authors estimate bivariate regressions of imputed daily ridership shares against actual observed ridership shares, obtaining R-squared of 0.85. They also show robustness by varying the grid cell size from 500 meters to 2 kilometers, finding no systematic decline in treatment effect magnitudes, which rules out large displacement effects within the network as an explanation for the results.&lt;/p&gt;
&lt;p&gt;Q: Does the response to network improvements vary by local poverty rates?&lt;/p&gt;
&lt;p&gt;A: The authors interact all six event types with an indicator for above-median poverty rate at the origin grid cell (from SMERU 2014 data), controlling for population. They find no clear pattern of heterogeneity by income level — richer and poorer areas respond similarly to service improvements. The paper notes this absence of heterogeneity as relevant context for interpreting optimal network design: the case for extending reach is not offset by a differential preference for frequency among poorer commuters.&lt;/p&gt;
&lt;p&gt;Mohring Effect: The externality arising from ridership responsiveness to wait times — more riders justify more buses, which reduce wait times for all riders, further increasing ridership. The paper estimates a BRT wait-time elasticity of −0.29, confirming the effect operates in Jakarta; for non-BRT the elasticity of −1.05 suggests the possibility of multiple equilibria in service levels.&lt;/p&gt;
&lt;p&gt;Negative Exponential Distribution Model (Daganzo 1979): The route-choice model used in the paper, in which bus arrivals on each route follow independent Poisson processes and wait times are exponentially distributed. The model is invariant to aggregation of identical routes (avoids the red-bus/blue-bus problem) and yields tractable closed-form expressions for choice probabilities and expected utility.&lt;/p&gt;
&lt;p&gt;Partial Inattention: The model feature whereby commuters assign near-zero effective arrival rates to bus options whose travel time exceeds the fastest available option by more than an estimated threshold (34–44 percent depending on route type). Captures the empirical finding that commuters in a large, complex network do not appear to consider all available options.&lt;/p&gt;
&lt;p&gt;Event Types (1, 2, 3): The paper&amp;rsquo;s taxonomy of service improvements induced by new route launches. Event 1 isolates the value of directness (new direct route, no speed gain). Event 2 combines directness and speed (new direct route that is also faster). Event 3 isolates the value of frequency (additional buses on an already-direct route, reducing wait time without changing travel time).&lt;/p&gt;
&lt;p&gt;Optimal Network Characterization via Social Planner&amp;rsquo;s Logit: The paper&amp;rsquo;s approach to the combinatorially intractable network optimization problem. The planner is modeled as making a logit discrete choice over all possible networks, with welfare from the demand model plus a network-level idiosyncratic shock. Sampling via simulated annealing yields estimates of optimal network characteristics and comparative statics without requiring identification of a single globally optimal network.&lt;/p&gt;
&lt;p&gt;Network Concentration vs. Extensiveness Tradeoff: The core design tension the paper formalizes — for a fixed bus fleet, concentrating buses on fewer routes reduces wait times on served routes but leaves more areas without coverage, while spreading buses across more routes extends reach at the cost of longer headways. The estimated preference parameters (high wait-time sensitivity) make this tradeoff non-trivial; nonetheless, the paper finds the current network is too concentrated relative to the optimum.&lt;/p&gt;</description></item><item><title>Optimal Taxation and Market Power</title><link>https://macropaperwarehouse.com/papers/optimal-taxation-and-market-power/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-taxation-and-market-power/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;This paper asks whether and how optimal income taxation should change when firms have market power. The question is motivated by the documented rise in economy-wide markups since 1980, which has compressed the labor share, widened the gap between worker and entrepreneurial income, and generated allocative inefficiency through excessive pricing.&lt;/p&gt;
&lt;p&gt;The authors develop a Mirrleesian optimal taxation framework augmented with three features absent from the canonical literature: (i) oligopolistic intermediate goods markets with endogenous, variable markups, (ii) heterogeneous firm productivities, and (iii) two occupational groups—wage-earning workers and profit-earning entrepreneurs—whose abilities are private information. Entrepreneurs strategically set prices under Cournot competition, which means that the tax system affects profits both through a firm&amp;rsquo;s own behavior and through the responses of its competitors. This strategic interaction is the critical novelty relative to prior work that assumes monopolistic competition.&lt;/p&gt;
&lt;p&gt;The main theoretical contribution is the derivation of optimal tax formulas for both labor income and profit income that decompose into four named components: (i) the Mirrleesian incentive component, which reflects the standard trade-off between redistribution and labor supply distortions; (ii) the Pigouvian component, which corrects for the externality from market power by subsidizing labor and entrepreneurial effort to offset the output shortfall from high markups; (iii) the Reallocation Effect (RE), which shifts the profit tax to redirect labor inputs from low-markup firms to high-markup firms where labor is inefficiently scarce, and which emerges only under heterogeneous markups; and (iv) the Indirect Redistribution Effect (IRE), which uses changes in competitors&amp;rsquo; product prices—a channel present only under oligopolistic (not monopolistic) competition—to redistribute income between entrepreneurs.&lt;/p&gt;
&lt;p&gt;For the labor income tax, the dominant force is the Pigouvian component. As average markups rise, the Pigouvian subsidy to labor supply grows, mechanically reducing optimal labor income tax rates. The profit tax is shaped by all four components in opposing directions; the net quantitative effect is resolved empirically.&lt;/p&gt;
&lt;p&gt;The model is calibrated to match distributions of labor income (from the Current Population Survey), profits (from Compustat-based data in De Loecker, Eeckhout, and Unger 2020), and firm-level markups (also from De Loecker, Eeckhout, and Unger 2020, using the cost-minimization approach) for the US in 1980 and 2019. The cost-weighted average markup rose from 1.25 in 1980 to 1.33 in 2019, with the increase concentrated at the top of the markup distribution.&lt;/p&gt;
&lt;p&gt;The central quantitative prescription is that the optimal labor income tax rate should decline by 7.7 percentage points between 1980 and 2019 (average optimal rate falls from 22.0 percent to 14.3 percent), while the optimal profit tax rate should rise by 2.2 percentage points on average (from 58.4 percent to 60.5 percent) and by 29.1 percentage points at the top. The decline in the labor income tax is driven primarily by the rise in average markups reducing the Pigouvian component. The increase in the profit tax, especially at the top, is driven primarily by the Mirrleesian component operating through the skill gap, which rises because higher markups reduce profit elasticity. The Pigouvian and reallocation components push in the opposite direction on the profit tax, but the Mirrleesian effect dominates.&lt;/p&gt;
&lt;p&gt;The optimal profit tax structure is regressive for large, high-markup firms—reflecting the RE, which requires lower tax rates for high-markup firms to incentivize labor reallocation toward them—but less regressive in 2019 than in 1980, reflecting the distributional tightening from rising markup inequality.&lt;/p&gt;
&lt;p&gt;Robustness checks across parameter values for the social welfare curvature k, the span of control ξ, and the elasticity of substitution σ confirm that the directional results hold: labor income tax rates decrease and profit tax rates increase from 1980 to 2019 across all parameter configurations. Extensions to nonlinear sales taxes and conditioning on markups confirm that even when the planner can observe markups directly, the first-best is not achievable because markups are endogenous to entrepreneurs&amp;rsquo; unobservable decisions.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-fundamental-difference-between-this-papers-model-and-prior-work-on-optimal-taxation-with-market-power"&gt;Q1. What is the fundamental difference between this paper&amp;rsquo;s model and prior work on optimal taxation with market power?&lt;/h3&gt;
&lt;p&gt;Prior work using monopolistic competition (e.g., Gürer 2021; Boar and Midrigan 2019) assumes each entrepreneur holds monopoly power in its own market, so no strategic interaction exists between firms. Under monopolistic competition, entrepreneurs price to maximize utility given competitors&amp;rsquo; choices, and the envelope theorem implies that tax changes have no first-order effect on prices or utility through the pricing channel—the Indirect Redistribution Effect (IRE) disappears. In this paper, entrepreneurs compete in Cournot oligopolistic markets with a finite number of firms I, so each firm&amp;rsquo;s pricing depends on competitors&amp;rsquo; output. A change in one firm&amp;rsquo;s output (induced by taxation) shifts competitors&amp;rsquo; prices, opening a redistribution channel through product markets that is entirely absent in monopolistic competition. Additionally, the Reallocation Effect (RE) emerges only when firm-level markups are heterogeneous, which requires oligopolistic (not perfectly competitive) markets.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-four-components-of-the-optimal-tax-formula-and-how-does-each-relate-to-market-power"&gt;Q2. What are the four components of the optimal tax formula and how does each relate to market power?&lt;/h3&gt;
&lt;p&gt;The optimal tax wedge for both labor and profit income decomposes into four components. First, the Mirrleesian component reflects the standard trade-off between redistribution and the efficiency cost of taxation; in the presence of market power, it is modified because the skill gap for entrepreneurs depends on markups through the profit elasticity. Second, the Pigouvian component corrects the externality from market power, which causes prices to exceed marginal cost and output to be inefficiently low; it implies a subsidy to both worker and entrepreneurial effort, scaled by the reciprocal of the average markup (for the labor tax) or firm-level markup (for the profit tax). Third, the Reallocation Effect (RE) applies only to the profit tax and reflects that labor should be shifted toward high-markup firms where it is inefficiently underemployed; it reduces the tax rate for firms whose markup exceeds the average. Fourth, the Indirect Redistribution Effect (IRE) captures redistribution through competitor price changes under oligopolistic interaction; it can either raise or lower the profit tax rate depending on the distribution of social welfare weights and the cross-inverse demand elasticity.&lt;/p&gt;
&lt;h3 id="q3-what-happens-to-the-labor-income-tax-formula-as-average-markups-rise"&gt;Q3. What happens to the labor income tax formula as average markups rise?&lt;/h3&gt;
&lt;p&gt;The labor income tax formula contains a Pigouvian component equal to the reciprocal of the employment-weighted average markup. As average markups rise, this reciprocal falls, reducing the optimal labor income tax rate. Quantitatively, the optimal average labor income tax rate declines from 22.0 percent in 1980 to 14.3 percent in 2019, a decrease of 7.7 percentage points. In a purely competitive benchmark economy, the top labor income tax rate would be around 60 percent (consistent with Saez 2001); in the calibrated model with market power, it is 34.2 percent in 1980 and 28.7 percent in 2019. The Pigouvian component accounts for essentially the entire difference because the Mirrleesian component, when calibrated to the same labor income distribution, is unchanged.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-mirrleesian-component-cause-the-top-profit-tax-rate-to-rise-with-market-power"&gt;Q4. How does the Mirrleesian component cause the top profit tax rate to rise with market power?&lt;/h3&gt;
&lt;p&gt;The Mirrleesian component of the profit tax is driven by the skill gap, defined as the proportional rate of change in the composite entrepreneur ability measure. The skill gap depends on markups through the profit elasticity: as markups rise, profit elasticity falls (since profit elasticity is approximately the reciprocal of markup minus the span-of-control parameter minus the inverse of the labor supply elasticity term), which increases the skill gap. A higher skill gap amplifies the income divergence across entrepreneur types, increasing the Mirrleesian incentive to redistribute at the top. Quantitatively, Figure 5 shows that the rise in the skill gap from 1980 to 2019 tracks almost exactly the change in the inverse of profit elasticity, confirming that markup changes—not changes in the ability distribution—are the primary driver of increased Mirrleesian pressure on top profit taxes.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-reallocation-effect-influence-the-structure-progressivity-of-the-profit-tax"&gt;Q5. How does the Reallocation Effect influence the structure (progressivity) of the profit tax?&lt;/h3&gt;
&lt;p&gt;The RE term equals the ratio of the average markup to the firm-level markup minus one: RE(θe) = μ/μ(θe) − 1. For firms with markups above the average, RE is negative, reducing their optimal tax rate; for firms below the average, RE is positive, increasing it. This implies that the optimal profit tax should be regressive relative to markup (i.e., high-markup firms face lower marginal tax rates), even though the overall profit tax rises on average. This provides a novel rationale for why the profit tax schedule in practice is less progressive—or even regressive—for large firms. As markups rise across the distribution, the reallocation effect pushes down the top profit tax but does not offset the larger increase from the Mirrleesian component in the quantitative exercise.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-indirect-redistribution-effect-and-why-does-it-disappear-under-monopolistic-competition"&gt;Q6. What is the Indirect Redistribution Effect and why does it disappear under monopolistic competition?&lt;/h3&gt;
&lt;p&gt;The IRE captures the change in entrepreneurial utility that arises because a tax reduction for one entrepreneur increases their output, which reduces the prices of substitute goods produced by competitors, thereby lowering competitors&amp;rsquo; incomes. Under oligopolistic competition with I &amp;gt; 1 firms per market, the cross-inverse demand elasticity is nonzero, so competitor prices are sensitive to any one firm&amp;rsquo;s output decision, and this redistribution channel is open. Under monopolistic competition (I = 1), each entrepreneur is the sole producer in its market; competitors&amp;rsquo; prices do not depend on the firm&amp;rsquo;s output, the cross-inverse demand elasticity is zero, and the IRE vanishes by the envelope theorem. The IRE is also absent in perfectly competitive economies. Empirical evidence for the US suggests the hazard ratio of profits is sufficiently high that the IRE generally pushes toward a lower top profit tax rate, but the Mirrleesian effect dominates in the quantitative results.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-quantitative-effect-of-rising-markups-on-the-optimal-tax-rates-and-what-drives-the-net-change-in-the-profit-tax"&gt;Q7. What is the quantitative effect of rising markups on the optimal tax rates, and what drives the net change in the profit tax?&lt;/h3&gt;
&lt;p&gt;The model calibrated to 1980 and 2019 US data prescribes a decline in the optimal average labor income tax rate of 7.7 percentage points (from 22.0 to 14.3 percent) and an increase in the optimal average profit tax rate of 2.2 percentage points (from 58.4 to 60.5 percent). At the top of the profit distribution, the increase is 29.1 percentage points. The net profit tax increase results from four opposing forces: the Pigouvian component falls (pushing toward lower taxes) and the RE decreases for high-markup firms (also pushing down the top rate), while the IRE and especially the Mirrleesian component rise (pushing up top rates). The Mirrleesian effect is the dominant force, driven by rising markup inequality reducing profit elasticity and widening the skill gap for top entrepreneurs.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-counterfactual-analysis-isolate-the-role-of-markups-from-productivity-changes"&gt;Q8. How does the counterfactual analysis isolate the role of markups from productivity changes?&lt;/h3&gt;
&lt;p&gt;The counterfactual fixes the markup distribution at its 1980 level while holding the 2019 productivity distribution constant, then solves for optimal taxes. The result is that high-profit entrepreneurs would face lower optimal tax rates under 1980 markups than under 2019 markups, while low-profit entrepreneurs would face higher rates. Decomposing the difference, the Pigouvian component and the RE are larger for high incomes under 1980 (lower) markups, making the profit tax more regressive, while the IRE and the Mirrleesian component are smaller under 1980 markups, producing a lower top rate. The increase in the Mirrleesian component due to the markup increase from 1980 to 2019 is identified as the primary reason top profit taxes rise. This isolates the markup channel from the productivity channel in accounting for changes in optimal taxes.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-robustness-analysis-reveal-about-parameter-sensitivity"&gt;Q9. What does the robustness analysis reveal about parameter sensitivity?&lt;/h3&gt;
&lt;p&gt;The main qualitative result—labor income taxes decline and profit taxes rise from 1980 to 2019—holds across a broad parameter space. The optimal profit tax rate is largely insensitive to the social welfare curvature parameter k: across k ∈ {0.77, 1, 3}, the average optimal profit tax rate is approximately 58 percent in 1980 and 61 percent in 2019. The optimal average labor income tax rate is more sensitive to k: for k = 0.7, 1, and 3, the 1980 rates are 20.3, 26.7, and 44.6 percent, and the 2019 rates are 12.5, 19.4, and 39.1 percent, respectively. Changes in the span-of-control parameter ξ and the substitution elasticity σ do not affect the labor income tax wedge schedule directly but do influence it indirectly through the markup distribution. The directional results are confirmed for all tested parameter configurations.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-role-of-the-additivity-property-from-prior-externality-literature-and-why-does-it-fail-here"&gt;Q10. What is the role of the &amp;ldquo;additivity property&amp;rdquo; from prior externality literature, and why does it fail here?&lt;/h3&gt;
&lt;p&gt;The additivity property from the Pigouvian externality literature (see Kopczuk 2003; Sandmo 1975) states that the Pigouvian correction is separable from other components of the optimal tax formula, implying that rising markups would simply decrease the optimal tax rate (since 1/μ falls). This property holds under simplifying assumptions that abstract from the general equilibrium and incentive effects of market power. In the present model, the additivity property does not hold because markups enter all four components of the optimal tax formula—not just the Pigouvian term—through the skill gap (Mirrleesian component), the RE, and the IRE. As a result, rising markups can increase the optimal profit tax rate even though the Pigouvian component falls, because the skill gap and Mirrleesian force dominate.&lt;/p&gt;
&lt;h3 id="q11-can-the-government-attain-the-first-best-by-conditioning-taxes-on-markups"&gt;Q11. Can the government attain the first-best by conditioning taxes on markups?&lt;/h3&gt;
&lt;p&gt;No. The paper demonstrates that even if the planner can observe and condition taxes on firm-level markups, the first-best is not achievable. The reason is that markups are endogenous to the entrepreneurs&amp;rsquo; unobservable decisions: an entrepreneur&amp;rsquo;s markup depends on their privately known type and chosen output. When the planner designs a mechanism that conditions on markup, the incentive constraint facing entrepreneurs remains the same as in the benchmark model, because the promise-keeping constraints are independent of the entrepreneur&amp;rsquo;s true type when markups are observable. The optimal allocation with markup-conditioned taxes is shown to be equivalent to the second-best with nonlinear sales taxes, which still falls short of the first-best.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-policy-implications-for-the-design-of-the-profit-tax-schedule"&gt;Q12. What are the policy implications for the design of the profit tax schedule?&lt;/h3&gt;
&lt;p&gt;The model yields three concrete prescriptions for the joint design of labor and profit income taxes in the context of rising market power. First, labor income taxes should be reduced and top profit taxes should be increased as market power rises. Second, for large, high-productivity firms the profit tax should be designed to be appropriately regressive to enhance allocative efficiency through the Reallocation Effect—this provides a new normative justification for why profit tax schedules observed in practice are often less progressive than labor income taxes. Third, while profit taxes should be regressive for large firms, the degree of regressivity should decrease as market power rises, reflecting the trade-off between efficiency and equality: higher markups increase the Mirrleesian pressure for redistribution at the top, reducing the optimal regressivity.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Mirrleesian component (of the optimal tax formula):&lt;/strong&gt; The standard incentive component of the optimal tax, capturing the trade-off between direct redistribution and the efficiency cost of taxation. In the presence of market power, this component is modified because the skill gap for entrepreneurs depends on markups through the profit elasticity: higher markups reduce profit elasticity, widen the skill gap, and amplify the Mirrleesian force toward higher top profit taxes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pigouvian component:&lt;/strong&gt; The correction in the optimal tax formula for the externality from market power. Because oligopolistic pricing causes output to be inefficiently low, the optimal tax subsidizes both worker and entrepreneurial labor supply. In the labor income tax formula, the Pigouvian component is the reciprocal of the employment-weighted average markup; in the profit tax formula, it is the reciprocal of the firm-level markup. As average markups rise, the Pigouvian component reduces the optimal labor income tax rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reallocation Effect (RE):&lt;/strong&gt; A component of the optimal profit tax formula that captures the efficiency gain from reallocating labor inputs from low-markup firms (where labor&amp;rsquo;s marginal product is high relative to value) to high-markup firms (where labor demand is inefficiently low). It equals the ratio of the average markup to the firm-level markup minus one. It implies a lower optimal marginal tax rate for firms with markups above the average, producing a regressive structure in the profit tax for large firms. This effect is absent under monopolistic competition (uniform markups) and in competitive markets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Indirect Redistribution Effect (IRE):&lt;/strong&gt; A component of the optimal profit tax formula specific to oligopolistic competition, capturing redistribution through competitor prices. Lowering the marginal tax rate of a high-productivity entrepreneur raises their output, which reduces the prices of substitutable goods produced by their competitors, thereby lowering competitors&amp;rsquo; incomes and redistributing toward workers who benefit from lower prices. This effect is present only when the cross-inverse demand elasticity is nonzero—i.e., only under oligopolistic (Cournot) competition with multiple firms per market—and vanishes under monopolistic competition and in the limit as the number of firms grows to infinity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Skill gap (for entrepreneurs):&lt;/strong&gt; The proportional rate of change in the composite entrepreneur ability measure with respect to entrepreneur type, analogous to the Mirrleesian skill gap for workers. Under market power, the entrepreneur skill gap depends on the markup through the profit elasticity: as firm-level markups rise, profit elasticity falls, the skill gap increases, and the income dispersion across entrepreneurs widens, which amplifies the Mirrleesian incentive to redistribute at the top and raises the optimal top profit tax rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Symmetric Cournot Competitive Tax Equilibrium (SCCTE):&lt;/strong&gt; The equilibrium concept used in the paper. It is a combination of a tax system, symmetric allocation, and symmetric price system such that all agents (final goods producer, entrepreneurs of each type, workers) are optimizing, strategic interaction in the intermediate goods market is a Cournot Nash equilibrium within each granular market, and all commodity and labor markets clear. Strategic interaction is restricted to within each granular market (firms in the same market compete), so decisions across markets are taken as given.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Composite ability:&lt;/strong&gt; A combined measure of entrepreneur productivity that determines equilibrium allocations and optimal taxation in the nested-CES economy. It aggregates the entrepreneur&amp;rsquo;s raw ability (affecting output capacity) and the demand parameter (affecting the market-level markup). The markup-relevant component and the quantity-relevant component are not perfect substitutes in the composite, since equilibrium prices depend on their specific composition while equilibrium quantities depend only on their combined value.&lt;/p&gt;</description></item><item><title>Peer Effects and Rank Concerns in the Classroom</title><link>https://macropaperwarehouse.com/papers/peer-effects-and-rank-concerns-in-the-classroom/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/peer-effects-and-rank-concerns-in-the-classroom/</guid><description>&lt;p&gt;This paper investigates the mechanisms behind peer effects in the classroom using exogenous variation in study disruptions generated by the 2010 Maule mega-earthquake in Chile (magnitude 8.8, the seventh-largest ever instrumentally recorded). The central research question is why classroom peers can shape academic achievement — specifically, whether beyond production complementarities and a desire to conform, a desire to compete for classroom rank can drive peer influence on learning.&lt;/p&gt;
&lt;p&gt;The author constructs a novel dataset linking administrative and survey data from Chile&amp;rsquo;s Ministry of Education (SIMCE test scores, GPA, curriculum coverage, and school expenditure records) for two cohorts of roughly 150,000 eighth-grade students — one measured in 2009 before the earthquake, one measured in 2011 roughly 20–22 months after — to newly constructed measures of housing damage. Damage to each student&amp;rsquo;s home is built in three steps: (1) ground-shaking intensity using an established attenuation formula for the 2010 earthquake; (2) seismic vulnerability of each student&amp;rsquo;s home inferred from a latent-class-analysis model trained on census data linking housing construction materials to vulnerability classes; and (3) a combined expected &amp;ldquo;damage ratio&amp;rdquo; (fraction of home that needs to be rebuilt). Identification uses a difference-in-differences strategy that exploits the differential correlation between pre-existing seismic vulnerability and outcomes across the pre- and post-earthquake cohorts, controlling for socioeconomic composition.&lt;/p&gt;
&lt;p&gt;The main findings, holding fixed a student&amp;rsquo;s own earthquake exposure, are as follows. (1) Own home damage reduced test scores by 0.03 standard deviations (SD) per SD increase in damages (a 4.4 percentage-point increase in collapsed home fraction, approximately USD 3,600) and raised self-reported cost of study effort. GPA effects (–0.02 SD) are statistically insignificant. (2) A 1 SD increase in the mean damage among classroom peers raised test scores by 0.05 SD and GPA by 0.04 SD. School expenditure data (available for the 42% of schools in the preferential subsidy program) show schools responded by reallocating funds away from administrative activities toward educational and psychological support, accounting for this positive effect. (3) A 1 SD increase in the within-classroom standard deviation of peer damages lowered test scores and GPA by approximately 0.085 SD on average, but with sharply heterogeneous effects across the prior-achievement distribution: it lowered test scores and GPA of high-prior-achievement students by 0.08–0.11 SD and raised achievement of low-prior-achievement students, without corresponding changes in those students&amp;rsquo; GPA rank. Neither curriculum-coverage data nor school spending data show significant responses to damage dispersion, pointing to peer-to-peer interactions rather than school mediation.&lt;/p&gt;
&lt;p&gt;The null effect on GPA rank despite heterogeneous GPA effects is the pivotal empirical finding motivating the paper&amp;rsquo;s theory. The author argues that high-achieving students reduced effort in response to a less threatening competitive environment while maintaining their classroom standing — consistent with rank concerns driving effort decisions. Direct survey evidence shows a majority of students agreed they like to do better than classmates.&lt;/p&gt;
&lt;p&gt;Motivated by this evidence, the paper introduces a game-of-status model where each student chooses effort to maximize a utility function combining academic achievement and classroom GPA rank, with rank weighted by a preference parameter lambda &amp;gt; 0. The model admits a unique symmetric Bayesian Nash equilibrium. The model rationalizes all four main empirical patterns: positive mean-damage effects (school compensation); heterogeneous dispersion effects (rank competition changes the density of nearby competitors); null dispersion effects on GPA rank (simultaneous equilibrium adjustment preserves rank ordering); and the survey evidence on competitive preferences.&lt;/p&gt;
&lt;p&gt;The study is confined to Chilean public and subsidized private schools in earthquake-affected, non-coastal regions, with outcomes measured at the 8th grade. The pre/post cohort design removes schools that closed or received earthquake evacuees. Findings apply to a context where classroom rank is observable to peers (GPA) and where competitive preferences are prevalent among students.&lt;/p&gt;
&lt;p&gt;Q: What is the core identification strategy and why does it avoid the usual confounds in peer-effects research?
A: The paper uses a difference-in-differences estimator that exploits the differential relationship between pre-existing seismic vulnerability and outcomes across a pre-earthquake cohort (outcomes measured in 2009) and a post-earthquake cohort (outcomes measured in 2011). Because identification relies on variation in peer disruptions rather than in peer characteristics — and because students did not reallocate across classrooms or schools in response to the earthquake in the estimation sample — the strategy avoids the reflection problem and selection confounds that typically plague peer-effects identification. The identifying assumption is that the relationship between seismic vulnerability and outcomes would have been the same across cohorts absent the earthquake.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports the identifying assumption?
A: The paper provides three pieces of supporting evidence. First, the fraction of students switching schools or classrooms between grades 7 and 8 is identical across the pre- and post-earthquake cohorts in the estimation sample, indicating no earthquake-induced reallocation. Second, pre-trend tests show precise zero effects of own damage, mean peer damage, and SD of peer damage on lagged (4th-grade) test scores and GPA. Third, placebo tests using students in regions unaffected by the earthquake show no significant differential relationships between seismic vulnerability measures and outcomes across cohorts.&lt;/p&gt;
&lt;p&gt;Q: How was housing damage measured, and why does this matter for identification?
A: Damage is estimated in three steps: ground-shaking intensity at the student&amp;rsquo;s town is calculated from a validated attenuation formula; seismic vulnerability of the home is predicted using a latent-class-analysis model trained on pre-earthquake census housing data and then applied to student records; and the two are combined into a damage ratio (fraction of home to be rebuilt) using structural engineering damage-grade distributions. This constructed measure is not self-reported and is determined by physical and housing-quality factors largely predetermined before the earthquake, which supports exogeneity. Coastal towns are excluded because the accompanying tsunami caused damages not captured by the damage-ratio formula, and results are robust to different definitions of coastal proximity.&lt;/p&gt;
&lt;p&gt;Q: What were the effects of damage to a student&amp;rsquo;s own home on achievement?
A: A 1 SD increase in own home damages (corresponding to a 4.4 percentage-point increase in the collapsed fraction of the home, or roughly USD 3,600) reduced test scores by 0.03 SD. GPA fell by 0.02 SD but this was not statistically significant. Survey data show that own-home damages raised students&amp;rsquo; self-reported cost of study effort, suggesting this effort channel may mediate the achievement effects. These negative effects did not vary significantly across the baseline achievement distribution.&lt;/p&gt;
&lt;p&gt;Q: What were the effects of mean peer damage on own achievement, and what mechanism explains them?
A: A 1 SD increase in mean peer home damage raised own test scores by 0.05 SD and GPA by 0.04 SD. School spending data from SEP-program schools (42% of the sample) show that schools responded to higher average student damage by reallocating expenditures away from administrative activities (recruitment of non-teaching staff, equipment purchases) toward educational support and psychological support activities. This reallocation more than offset potential negative peer-environment effects, generating positive net achievement effects that were approximately uniform across the prior-achievement distribution.&lt;/p&gt;
&lt;p&gt;Q: What were the effects of within-classroom damage dispersion on achievement, and how do they vary across students?
A: A 1 SD increase in the within-classroom standard deviation of peer damages lowered average test scores and GPA by approximately 0.085 SD. These average effects mask sharp heterogeneity: high-prior-achievement students experienced losses of 0.08–0.11 SD in test scores and GPA, while low-prior-achievement students saw gains. For some students the dispersion effect was comparable to or larger than the effect of damage to their own home.&lt;/p&gt;
&lt;p&gt;Q: Why is the null effect of damage dispersion on GPA rank theoretically important?
A: Students with high prior achievement experienced drops in GPA in classrooms with more dispersed damages, but without an accompanying drop in their GPA rank. The paper argues this is inconsistent with students passively absorbing a changed study environment: instead, students appear to have adjusted effort precisely enough to maintain their classroom standing. This equilibrium pattern — GPA changes that leave rank ordering intact — is the paper&amp;rsquo;s key empirical signature of rank-motivated competition as a mechanism for peer influence.&lt;/p&gt;
&lt;p&gt;Q: What direct survey evidence is presented on rank concerns?
A: Survey data from the post-earthquake cohort show that a majority of students agreed with the statement &amp;ldquo;I like to do better than my classmates in school,&amp;rdquo; providing direct evidence that students value classroom rank. Additionally, students with higher initial achievement reported reductions in self-reported ability to engage with course content in classrooms with more dispersed damages, consistent with these students reducing effort when the competitive environment became less threatening to their rank.&lt;/p&gt;
&lt;p&gt;Q: Do schools mediate the damage-dispersion spillovers?
A: The available data on curriculum coverage and school spending do not show statistically significant responses to within-classroom damage dispersion (as distinct from mean damage). Emergency reconstruction funds were also allocated by schools based on overall damage severity, not its within-classroom dispersion. This absence of a detectable school-mediation channel for dispersion effects strengthens the interpretation that the heterogeneous achievement effects of dispersion reflect peer-to-peer interactions rather than differential school responses.&lt;/p&gt;
&lt;p&gt;Q: How does the game-of-status model rationalize the empirical findings?
A: In the model, each student maximizes a utility function over academic achievement and GPA rank, with rank weighted by lambda &amp;gt; 0. Students choose effort simultaneously, and their cost-of-effort type is shaped by prior test scores, socioeconomic characteristics, and earthquake damage. The model admits a unique symmetric Bayesian Nash equilibrium. In this equilibrium: schools&amp;rsquo; compensating inputs in response to mean damage raise achievement uniformly (rationalizing positive mean-damage effects); changes in damage dispersion alter the density of nearby types differently for high- and low-cost-effort students, changing the marginal benefit of exerting effort to overtake competitors (rationalizing heterogeneous GPA effects); and because all students adjust effort simultaneously, the rank ordering is approximately preserved (rationalizing null rank effects).&lt;/p&gt;
&lt;p&gt;Q: What is the mechanism by which damage dispersion produces heterogeneous effort incentives?
A: The key mechanism is that when students derive utility from rank, the marginal benefit of a unit of additional effort depends on how many competitors are &amp;ldquo;nearby&amp;rdquo; in the effort-cost distribution. When dispersion increases, the density of types just below a high-achiever (low-cost-effort student) decreases, reducing the gain from exerting more effort to maintain rank over nearby rivals; high-achievers therefore reduce effort and GPA falls. Conversely, when dispersion increases, low-achievers face a distribution where they can more effectively compete for higher ranks, raising their effort incentive and GPA.&lt;/p&gt;
&lt;p&gt;Q: How does this paper&amp;rsquo;s theory differ from prior theories of peer influence?
A: Prior theories have emphasized two mechanisms: production complementarities (peer ability directly improves own learning) and a desire to conform (students prefer to match their peers&amp;rsquo; effort or achievement). Both rationalize a linear-in-means model that captures only mean peer characteristics. This paper&amp;rsquo;s theory is the first in the peer-effects literature to rationalize why higher-order moments of the peer distribution (specifically dispersion) affect learning, through a competitive rank-concern mechanism that is parsimonious and does not require extensions to production technology or preferences beyond adding rank to the utility function.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the competitive-motive theory?
A: The theory implies that classroom composition policies affecting the dispersion of student ability — such as ability tracking, gifted programs, or reshuffling policies — can have heterogeneous and potentially perverse effects: policies that reduce ability dispersion may concentrate competitive incentives in ways that harm some students while benefiting others. Standard linear-in-means models of peer effects, which capture only mean peer characteristics, would not predict these distributional consequences. The author argues this means the competitive mechanism has been largely unexplored despite its intuitive appeal, and calls for structural estimation and policy analysis in future work.&lt;/p&gt;
&lt;p&gt;Q: What is the scope of the empirical findings?
A: The findings apply to 8th-grade students in Chilean public and private subsidized schools located in earthquake-affected, non-coastal regions, with outcomes observed approximately 20–22 months post-earthquake. The sample excludes schools that closed due to the earthquake and schools that received evacuees. The paper notes that while the theory is formulated around an earthquake shock, the competitive-motive mechanism applies whenever the dispersion of students&amp;rsquo; cost-of-effort types changes — including through classroom assignment policies or other shocks — and is not specific to the natural-disaster context.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Damage ratio&lt;/strong&gt;: The fraction of a student&amp;rsquo;s home that needs to be rebuilt, constructed by combining geocoded ground-shaking intensity (via the Astroza et al. attenuation formula for the 2010 Chilean earthquake) with the predicted seismic vulnerability class of the home (derived from a latent-class-analysis model trained on census housing data). Used as the paper&amp;rsquo;s measure of disruption to each student&amp;rsquo;s environment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exogenous peer effect&lt;/strong&gt; (in the sense of Manski 1993): The reduced-form impact on a student&amp;rsquo;s outcome of a change in the distribution of an exogenous characteristic — here, earthquake damage — among classroom peers, holding fixed the student&amp;rsquo;s own characteristics. Distinguished in the paper from endogenous peer effects (best-response functions).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rank concern&lt;/strong&gt;: Students&amp;rsquo; utility derived from their position (rank) in the classroom GPA distribution, irrespective of whether that rank is formally rewarded. The paper treats rank concern as a preference parameter (lambda &amp;gt; 0 in the utility function) and identifies it as a mechanism for peer influence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Game-of-status model&lt;/strong&gt;: The paper&amp;rsquo;s theoretical framework, in which students simultaneously choose study effort to maximize utility over own academic achievement and GPA rank. The model admits a unique symmetric Bayesian Nash equilibrium. The central insight is that the density of nearby competitors in the effort-cost distribution determines the marginal benefit of effort, generating heterogeneous incentives when peer cost-of-effort types become more dispersed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Effort-cost type&lt;/strong&gt;: Each student&amp;rsquo;s marginal cost of exerting study effort, shaped by prior test scores, socioeconomic characteristics, and earthquake damages to the student&amp;rsquo;s own home. The key primitive of the model that links individual disruptions to equilibrium effort choices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;SEP (Subvencion Escolar Preferencial)&lt;/strong&gt;: Chile&amp;rsquo;s preferential school subsidy program for disadvantaged students, which requires participating schools (42% of the sample) to submit detailed annual spending reports to the Ministry of Education. The paper uses these reports to identify school spending responses to mean and dispersed peer damages.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Seismic vulnerability class&lt;/strong&gt;: A classification of a home&amp;rsquo;s resistance to earthquake damage based on its construction materials (exterior walls, roof, floor), assigned using a logistic latent-class-analysis model estimated on census data. Found to align strongly with household socioeconomic status, enabling prediction of housing vulnerability from administrative student records.&lt;/p&gt;</description></item><item><title>Pigovian Transport Pricing in Practice</title><link>https://macropaperwarehouse.com/papers/pigovian-transport-pricing-in-practice/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/pigovian-transport-pricing-in-practice/</guid><description>&lt;p&gt;This paper reports on the MOBIS experiment, a large-scale randomized controlled trial (RCT) implementing a multi-modal Pigovian transport pricing scheme in urban areas of German- and French-speaking Switzerland. The central research question is whether a first-best transport pricing scheme — one that charges users the full marginal external costs of their travel choices, varying across time, space, and mode — generates meaningful behavioral responses, and how those responses compare to a pure information intervention.&lt;/p&gt;
&lt;p&gt;The study recruited participants from urban areas, requiring them to be between 18 and 65 years old and to use a car at least two days per week. After contacting over 90,000 individuals and an initial online screening of 21,800 respondents, 3,656 participants completed the RCT. Each participant agreed to have their daily travel tracked via a smartphone app (&amp;ldquo;Catch-My-Day&amp;rdquo;) for eight weeks: four weeks of observation followed by four weeks of treatment. Assignment to treatment and control groups was fully randomized without stratification.&lt;/p&gt;
&lt;p&gt;The pricing treatment gave participants a budget equal to their observed external costs during the observation period plus a 20% buffer, from which the external costs of their actual travel were deducted in real time; any remaining balance was theirs to keep. External costs were computed across all modes using official Swiss Federal Roads Office monetization factors, including congestion (via a MATSim-based average marginal cost approach), CO2 climate costs (CHF 136.08/ton), health costs from air pollution (PM10 and NOx), and accident and physical activity effects for active and public modes. Public transport also carried a peak-hour surcharge of CHF 0.10/km for congested zone-pairs. A second &amp;ldquo;information-only&amp;rdquo; treatment provided identical information about external costs but imposed no financial charge. A control group received only weekly summaries of kilometers traveled by mode.&lt;/p&gt;
&lt;p&gt;The regression framework is a difference-in-differences specification with person, calendar-day, and day-of-study fixed effects, estimated in levels for external-cost outcomes (due to negative values from walking&amp;rsquo;s net external benefit) and via Poisson Pseudo-Maximum Likelihood for non-negative outcomes.&lt;/p&gt;
&lt;p&gt;The pricing treatment reduced total external costs by CHF 0.215 per day (p &amp;lt; 0.01), a 5.1% reduction relative to the control group. The average private cost of transport for the control group during the treatment period was CHF 25.72 per day; the external cost was CHF 4.22 per day, implying that Pigovian pricing raised total transport costs by 16.4% on average. The implied price elasticity of external costs with respect to this price increase is -0.31. The reduction is attributable to mode substitution toward public transport and active modes and to departure time shifting away from peak hours, but not to a reduction in total distance traveled.&lt;/p&gt;
&lt;p&gt;The information-only treatment produced a coefficient of -0.087, which is not statistically significant at conventional levels for the full sample. The differential effect of adding pricing to information is -0.127 (marginally significant, p &amp;lt; 0.1), with the pricing increment particularly important for reducing congestion costs. Sensitivity analysis shows that removing the control group and time fixed effects inflates the before-vs.-after elasticity to between -0.57 and -0.71, substantially larger than the preferred estimate of -0.31, underscoring the importance of the experimental design.&lt;/p&gt;
&lt;p&gt;Heterogeneity analysis reveals that men respond more strongly than women, German speakers more than French speakers, participants under 30 more than older participants, and those with above-median altruistic values respond significantly even to information alone. Correct knowledge of the definition of external costs (present in 45% of the sample) is a key driver of the pricing treatment effect. These scope conditions — mode availability, urban Swiss context, short 4-week treatment window, mandatory car use eligibility, and the specific external cost monetization framework — bound the generalizability of the elasticity estimate.&lt;/p&gt;
&lt;p&gt;Q: What is the main treatment effect of the Pigovian pricing scheme on external transport costs?
A: The pricing treatment reduced total external costs by CHF 0.215 per day, which is a 5.1% reduction relative to the control group (p &amp;lt; 0.01). About half of the reduction came from health costs, with congestion and climate costs following in magnitude. The implied elasticity of external costs with respect to the Pigovian price increase is -0.31, meaning a 10% increase in total transport costs from Pigovian pricing would reduce external costs by approximately 3.1% in the short run.&lt;/p&gt;
&lt;p&gt;Q: How was the Pigovian price increase calculated, and what was its magnitude relative to private costs?
A: The average private cost of transport for the control group during the treatment period was CHF 25.72 per day, and the average external cost was CHF 4.22 per day. The external cost thus represents 16.4% of total (private plus external) transport costs, and dividing the 5.1% reduction in external costs by this 16.4% price increase yields the elasticity of -0.31.&lt;/p&gt;
&lt;p&gt;Q: What mechanisms drove the reduction in external costs?
A: The reduction resulted from a combination of mode substitution — a shift away from car use toward public transport and active modes — and departure time shifting away from peak hours. Critically, total distance traveled did not decline; the behavioral adjustment operated entirely through changes in how and when people traveled, not in how much.&lt;/p&gt;
&lt;p&gt;Q: What was the effect of the information-only treatment?
A: The information-only treatment produced a coefficient of -0.087 CHF per day, which was not statistically significant at conventional levels for the full sample. It was statistically significant only for subgroups, notably participants with above-median altruistic values. The differential effect of adding pricing to information (alpha_P minus alpha_I = -0.127) was marginally significant (p &amp;lt; 0.1) and was particularly concentrated in congestion cost reductions, suggesting that the monetary incentive is especially important for internalizing the congestion externality.&lt;/p&gt;
&lt;p&gt;Q: Why is the control group critical, and how does removing it affect the estimated elasticity?
A: The tracking data show a seasonal negative trend in external costs over the study period; without a control group, this trend would be incorrectly attributed to the treatment, inflating the estimated effect. When both day-of-study and calendar-day fixed effects are removed (approximating a before-vs.-after design without a control group), the estimated elasticity rises to between -0.57 and -0.71, roughly double the preferred estimate of -0.31. This highlights that most prior studies in the literature, which lack control groups, are likely to overestimate treatment effects.&lt;/p&gt;
&lt;p&gt;Q: What heterogeneity is observed in the treatment response?
A: Men respond more strongly than women to both treatments, with the gender gap particularly pronounced for congestion costs. German speakers respond more strongly than French speakers. Participants under age 30 show stronger responses than older participants. Those scoring above the median on an altruistic values index respond significantly not only to pricing but also to information alone. Participants who correctly defined external costs (45% of the sample) drive the pricing treatment effect; a causal forest analysis confirms knowledge of external costs, age below 30, and language region as key heterogeneity drivers.&lt;/p&gt;
&lt;p&gt;Q: How were external costs computed across modes, and what are the key monetization parameters?
A: For private road transport, GPS tracks were map-matched using Graphhopper and processed via MATSim modules; emission factors came from the HBEFA 3.3 database, and congestion was assessed via an average marginal cost approach incorporating spillback effects. Externalities were monetized at CHF 136.08/ton for CO2, CHF 515,497–1,358,461/ton for PM10 (rural vs. urban), CHF 7,109/ton for NOx (regional), and a value of travel time savings of CHF 25.77/hour. For other modes, per-km values from the Swiss Federal Roads Office were applied. Walking carries net external benefits (negative external costs), while cycling carries small net external costs because accident costs exceed physical activity benefits.&lt;/p&gt;
&lt;p&gt;Q: How was public transport priced in the experiment, and why was it simplified?
A: A second-best zonal peak-hour surcharge of CHF 0.10/km was applied to public transport stages between zone-pairs experiencing peak demand, with peak windows set at 7–9 am and 5–7 pm. Full first-best pricing of public transport crowding was deemed infeasible because crowding effects are highly heterogeneous spatially and temporally, often concentrated in very short windows on specific lines, making aggregate distribution unreasonable.&lt;/p&gt;
&lt;p&gt;Q: Was there evidence of gaming the mode detection system?
A: Because participants could manually correct the app&amp;rsquo;s algorithmic mode assignments — and the pricing group had an incentive to overclaim low-cost modes — the potential for strategic misreporting was examined. While the analysis could not rule out some gaming, the main results were shown to be robust to excluding potential gamers, suggesting that gaming did not materially distort the treatment effect estimates.&lt;/p&gt;
&lt;p&gt;Q: What does the study imply for transport pricing policy?
A: The elasticity of -0.31 provides a benchmark for policymakers: a full Pigovian pricing scheme that raises total transport costs by about 16% can be expected to reduce external costs by about 5% in the short run in an urban context. The finding that congestion costs respond more to pricing than to information alone suggests the monetary component is essential for this externality. Heterogeneous responses — particularly the weaker responses by women and French speakers — have distributional implications. The experiment is a proof of concept that first-best transport pricing can generate meaningful behavioral responses, but scaling it would require addressing privacy concerns from GPS tracking, technical infrastructure, and political economy challenges.&lt;/p&gt;
&lt;p&gt;Pigovian transport pricing: A pricing scheme that charges each user the marginal external costs of their transport choices — including health, climate, congestion, and noise costs — as they vary across time, space, and mode, intended to internalize the gap between private and social costs of travel.&lt;/p&gt;
&lt;p&gt;External costs of transport: Costs borne by society rather than the individual traveler, including congestion (delay imposed on others), climate damages (CO2 emissions), health costs (local air pollution, accidents), and noise; in this paper, computed in real time from tracked trips using official Swiss monetization values.&lt;/p&gt;
&lt;p&gt;Average treatment effect (ATE): The difference-in-differences estimate of the causal effect of the pricing or information treatment on outcomes, identified from the randomized assignment and controlling for person, calendar-day, and day-of-study fixed effects.&lt;/p&gt;
&lt;p&gt;Mode substitution: The behavioral response in which travelers shift from higher-external-cost modes (primarily car) to lower-external-cost modes (public transport, walking, cycling) in response to pricing, as distinct from reducing total travel distance.&lt;/p&gt;
&lt;p&gt;Departure time shifting: The behavioral response in which travelers adjust when they depart to avoid peak-hour congestion surcharges, contributing to reduced congestion externalities without reducing total distance traveled.&lt;/p&gt;
&lt;p&gt;Information-only treatment: An experimental arm receiving identical information about external costs as the pricing group but facing no financial charge, used to isolate the informational component of the pricing treatment from the monetary incentive component.&lt;/p&gt;
&lt;p&gt;Source text origin: pdf&lt;/p&gt;</description></item><item><title>Place-Based Redistribution</title><link>https://macropaperwarehouse.com/papers/place-based-redistribution/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/place-based-redistribution/</guid><description>&lt;h2 id="place-based-redistribution-overview"&gt;Place-Based Redistribution: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Should national governments redistribute income to residents of poor areas through place-based transfers, or should redistribution rely solely on place-blind (income-only) taxes? The longstanding view in urban economics—&amp;ldquo;help poor people, not poor places&amp;rdquo;—holds that place-based aid is inefficient because it channels activity to less productive locations. This paper challenges that view by formalizing the conditions under which place-based redistribution improves on purely income-based transfers, using tools from optimal tax theory embedded in a spatial equilibrium model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper develops a two-location model (&amp;ldquo;Distressed&amp;rdquo; and &amp;ldquo;Elsewhere&amp;rdquo;) with a unit mass of heterogeneous households who differ in skill level (θ) and idiosyncratic preference for living in Distressed (φ). Households choose where to live and how much to earn, facing competitive labor and housing markets in each location. Locations may differ in amenity levels, wage schedules (which may embody skill-specific comparative advantage), and housing costs. A utilitarian planner sets location-specific income tax schedules—observed earnings and location are the only signals of unobserved skill—maximizing a weighted average of household utilities and landlord profits subject to a budget constraint.&lt;/p&gt;
&lt;p&gt;The paper proceeds in three steps. First, it derives closed-form conditions for the optimality of a lump-sum place-based transfer under a fixed income tax. Second, it characterizes fully general optimal nonlinear, location-specific marginal tax rate (MTR) schedules (Proposition 2). Third, it calibrates the model numerically, anchoring to the U.S. Empowerment Zone (EZ) program.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three Sorting Mechanisms and Their Policy Implications&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper identifies three polar mechanisms that generate sorting of lower-skill households into Distressed:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;em&gt;Skill-taste correlation&lt;/em&gt;: higher-skill households have stronger tastes for Elsewhere, independent of wages or rents.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Comparative advantage&lt;/em&gt;: higher-skill workers are relatively more productive in Elsewhere.&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Income-based sorting&lt;/em&gt;: because Elsewhere is more expensive, lower-income households are priced into Distressed.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Under skill-taste correlation, place-based transfers to Distressed are unambiguously welfare-improving even when income taxes are already optimal, because high-skill households prefer Elsewhere for reasons that are orthogonal to income. Under comparative advantage, the direction of the optimal transfer depends on migration elasticities: low migration elasticities favor transfers to Distressed, while high migration elasticities can reverse the sign. Under pure income-based sorting (with homogeneous locational preferences), the conditions for superfluous commodity taxation (Atkinson-Stiglitz 1976) are satisfied, and optimal place-based transfers are zero—though idiosyncratic preference heterogeneity restores non-zero optimal transfers even in this case.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Numerical simulations use Census data and ACS moments calibrated to EZ areas. With high migration responsiveness (κ = 0.5, approximating urban EZs) and skill-taste correlation as the sole sorting driver, the optimal average place-based transfer to Distressed is &lt;strong&gt;$4,805&lt;/strong&gt;, with about 40% ($1,943) arising from lower MTRs rather than a higher demogrant. With low migration responsiveness (κ = 4, approximating rural EZs), the optimal transfer more than doubles to &lt;strong&gt;$10,918&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;When comparative advantage alone drives sorting and migration is low (κ = 4), the optimal transfer to Distressed is &lt;strong&gt;$7,091&lt;/strong&gt;, with a $3,740 larger demogrant. With high migration and comparative advantage, the transfer reverses to &lt;strong&gt;−$2,763&lt;/strong&gt; (i.e., Elsewhere receives the subsidy). For intermediate migration under comparative advantage (e.g., κ ≈ 1), the optimal policy is nonlinear: the poorest Distressed residents receive a place-based transfer of &lt;strong&gt;$1,254&lt;/strong&gt;, while high-skill Distressed residents face a place-based tax of &lt;strong&gt;$12,398&lt;/strong&gt; at the 99th percentile.&lt;/p&gt;
&lt;p&gt;In the empirically calibrated &lt;strong&gt;urban EZ baseline&lt;/strong&gt; (migration elasticity 0.82, rent ratio 0.86, sorting driven by skill-taste correlation and income effects), the optimal average place-based transfer is &lt;strong&gt;$3,143&lt;/strong&gt;, roughly matching the magnitude of actual EZ wage tax credits (~$3,000 for full-time eligible workers). The demogrant advantage for Distressed is &lt;strong&gt;$1,462&lt;/strong&gt;, with just over half of the transfer arising from lower MTRs.&lt;/p&gt;
&lt;p&gt;In the &lt;strong&gt;rural EZ baseline&lt;/strong&gt; (migration elasticity 0.20, rent ratio 0.54, comparative advantage and income effects), the optimal average transfer rises to &lt;strong&gt;$4,329&lt;/strong&gt;, concentrated in lower MTRs rather than a larger demogrant. Halving the migration elasticity from the rural baseline raises the optimal transfer to &lt;strong&gt;$6,906&lt;/strong&gt;, while doubling it reduces the transfer to near zero (&lt;strong&gt;$573&lt;/strong&gt;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;All results are derived under the assumption of &lt;em&gt;no market failures&lt;/em&gt;; the model deliberately excludes agglomeration spillovers or other Pigouvian motives, attributing the case for place-based redistribution purely to redistributive goals.&lt;/li&gt;
&lt;li&gt;The planner observes only earnings and location, not skill type directly.&lt;/li&gt;
&lt;li&gt;Household Pareto weights are set equal to one across types in the simulations, so redistribution is driven solely by diminishing marginal utility of consumption.&lt;/li&gt;
&lt;li&gt;The model abstracts from interactions with subnational governments, local public services, and endogenous amenities.&lt;/li&gt;
&lt;li&gt;Results on the desirability of transfers to Distressed hinge critically on the motive for sorting, not simply on the existence of spatial income inequality.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-equity-efficiency-tradeoff-formula-for-a-lump-sum-place-based-transfer-and-what-does-it-reveal"&gt;Q1. What is the equity-efficiency tradeoff formula for a lump-sum place-based transfer, and what does it reveal?&lt;/h3&gt;
&lt;p&gt;Lemma 1 shows that the first-order welfare effect of a small per-capita transfer from Elsewhere to Distressed starting from a place-blind tax system is dSWF/dt = (λ̄₁ − λ̄₀) + Eθ{m(0)·[T(z₁*) − T(z₀*)]}. The equity gain (λ̄₁ − λ̄₀) is positive when Distressed households have higher average social marginal welfare weights, which holds when their skill distribution is first-order stochastically dominated by Elsewhere&amp;rsquo;s. The fiscal cost equals the earnings-tax-revenue loss from movers: households induced to migrate to Distressed who earn less there generate lower tax payments. This formula identifies the earnings response to migration as a sufficient statistic for the efficiency cost of place-based policy.&lt;/p&gt;
&lt;h3 id="q2-what-characterizes-the-optimal-lump-sum-transfer-t-in-proposition-1"&gt;Q2. What characterizes the optimal lump-sum transfer t* in Proposition 1?&lt;/h3&gt;
&lt;p&gt;Proposition 1 shows t* = [λ̄₁(t*) − λ̄₀(t*) + Eθ{m(t*)·[T(z₁*) − T(z₀*)]}] / (Eθ[m(t*)] / [L₀(t*)L₁(t*)]). The optimal transfer is larger when (i) the average social marginal welfare weight gap between Distressed and Elsewhere is greater, (ii) migration responses m(t*) are small, and (iii) the earnings difference between locations for marginal movers is small. This formula holds regardless of whether the income tax schedule T(·) is itself set optimally.&lt;/p&gt;
&lt;h3 id="q3-under-skill-taste-correlation-why-are-place-based-transfers-always-welfare-improving-even-under-an-optimal-income-tax"&gt;Q3. Under skill-taste correlation, why are place-based transfers always welfare-improving even under an optimal income tax?&lt;/h3&gt;
&lt;p&gt;When sorting is driven by skill-taste correlation (high-skill households have stronger preferences for Elsewhere despite identical wages and rents), the equity gain λ̄₁ − λ̄₀ is positive because low-skill households concentrate in Distressed. A small positive transfer starting from t = 0 also incurs zero fiscal cost because movers between locations face identical wages and do not change their earnings. Thus, welfare unambiguously increases. The key insight is that skill-taste correlation violates the Atkinson-Stiglitz condition: high earners would still prefer Elsewhere even if forced to earn less, so location serves as a proxy for skill not captured by income taxes alone.&lt;/p&gt;
&lt;h3 id="q4-under-comparative-advantage-why-can-the-sign-of-the-optimal-transfer-reverse-with-migration-elasticity"&gt;Q4. Under comparative advantage, why can the sign of the optimal transfer reverse with migration elasticity?&lt;/h3&gt;
&lt;p&gt;When higher-skill workers are more productive in Elsewhere, movers to Distressed experience wage and earnings reductions, generating a fiscal externality. When migration elasticities are high (low κ), this fiscal cost is large and can dominate the equity gain, making transfers to Elsewhere optimal (simulated optimal transfer of −$2,763 at κ = 0.5). When migration elasticities are low (high κ), the fiscal cost is small and equity considerations dominate, yielding transfers to Distressed ($7,091 at κ = 4). At intermediate elasticities, the optimal policy is nonlinear, redistributing to poor Distressed residents while taxing rich Distressed residents more.&lt;/p&gt;
&lt;h3 id="q5-why-are-place-based-transfers-superfluous-under-pure-income-based-sorting-with-homogeneous-locational-preferences"&gt;Q5. Why are place-based transfers superfluous under pure income-based sorting with homogeneous locational preferences?&lt;/h3&gt;
&lt;p&gt;Example 6 (and its formal proof in Appendix B.3.5) demonstrates that when sorting arises solely from higher rents in Elsewhere and preferences over location are homogeneous (no idiosyncratic φ heterogeneity), the Atkinson-Stiglitz sufficient condition for commodity tax superfluousness is met: hypothetically forcing high earners to earn less would not change their preferred consumption bundle relative to low earners. Hence a place-blind income tax implements optimal redistribution without spatial supplements. As the variance of idiosyncratic location preferences κ shrinks toward zero, Figure 3 confirms that optimal place-based transfers tend toward zero across all three sorting motives.&lt;/p&gt;
&lt;h3 id="q6-what-new-terms-appear-in-the-optimal-location-specific-mtr-formulas-proposition-2-relative-to-a-standalone-economy-optimum"&gt;Q6. What new terms appear in the optimal location-specific MTR formulas (Proposition 2) relative to a standalone-economy optimum?&lt;/h3&gt;
&lt;p&gt;The optimal MTR schedules in Proposition 2 contain two new terms beyond the standard Mirrlees (1971)/Saez (2001) formula. The term Δτ+(θ) captures the fiscal externality from migration: raising Elsewhere&amp;rsquo;s MTR at skill level θ and above induces movers to Distressed who change their tax revenue by T₁(z₁*(s)) − T₀(z₀*(s)). The term (λ_L − 1)Δr+(θ) captures the equilibrium rent effect: MTR changes shift households between locations, altering rents in both communities and redistributing between renters and landlords. When λ_L &amp;lt; 1 (landlords are weighted less than average households), the rent term creates additional motives for spatial redistribution depending on the ratio of rents to housing supply elasticities across locations.&lt;/p&gt;
&lt;h3 id="q7-how-do-housing-supply-elasticities-affect-the-optimal-spatial-transfer-and-why-does-the-sign-differ-between-urban-and-rural-settings"&gt;Q7. How do housing supply elasticities affect the optimal spatial transfer, and why does the sign differ between urban and rural settings?&lt;/h3&gt;
&lt;p&gt;The rent redistribution term Δr+(θ) has sign determined by r₁/ϱ₁ − r₀/ϱ₀. For urban EZs, where Distressed has lower rents but also lower housing supply elasticity than Elsewhere (ϱ₁ = 0.24, ϱ₀ = 0.34 in the baseline), this ratio is positive, meaning transfers to Distressed shift households into relatively inelastic markets, raising rents there and generating landlord income. When λ_L &amp;lt; 1, this reduces the desirability of transfers to Distressed. For rural EZs, Distressed has higher housing supply elasticity (ϱ₁ = 0.60), so the ratio is negative: transfers shift households to more elastic markets where rents rise minimally. When λ_L &amp;lt; 1, this actually motivates more transfers to rural Distressed areas. In the 75%-landlord-weight sensitivity, optimal urban transfers fall by ~$1,000 while rural transfers rise by ~$1,000, illustrating this asymmetry.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-urban-ez-baseline-calibration-find-about-optimal-transfers-and-how-does-it-compare-to-actual-ez-policy"&gt;Q8. What does the urban EZ baseline calibration find about optimal transfers and how does it compare to actual EZ policy?&lt;/h3&gt;
&lt;p&gt;The urban baseline targets a migration elasticity of 0.82 (from Busso et al. 2013), a Distressed-to-Elsewhere rent ratio of 0.86, and 56% of Distressed residents earning under $50,000. The calibrated κ is 0.44. At the optimum, Distressed residents receive an average place-based transfer of $3,143, with $1,462 as a higher demogrant and the remainder from lower MTRs. By comparison, actual EZs provide a wage tax credit of approximately $3,000 per eligible full-time worker. The paper concludes that the magnitude—but not the capped, flat structure—of EZ transfers approximates the optimal level.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-rural-ez-calibration-find-and-how-sensitive-are-results-to-migration-assumptions"&gt;Q9. What does the rural EZ calibration find, and how sensitive are results to migration assumptions?&lt;/h3&gt;
&lt;p&gt;The rural baseline targets a migration elasticity of 0.20 (from Sprung-Keyser et al. 2022), a rent ratio of 0.54, and 60% of Distressed residents earning under $50,000, with sorting attributed to comparative advantage and income effects. The calibrated κ is 4.06. The optimal average transfer is $4,329, primarily arising from lower MTRs rather than a higher demogrant ($532). Doubling the migration elasticity reduces the optimal transfer to near zero ($573); halving it raises it to $6,906. The direction and magnitude of optimal transfers are therefore highly sensitive to the assumed level of migration responsiveness, highlighting the empirical importance of estimating migration elasticities—particularly heterogeneity in migration by income level and earnings changes for marginal movers.&lt;/p&gt;
&lt;h3 id="q10-do-within-income-transfers-arising-from-differences-in-marital-and-parental-status-across-communities-effectively-constitute-place-based-redistribution"&gt;Q10. Do within-income transfers arising from differences in marital and parental status across communities effectively constitute place-based redistribution?&lt;/h3&gt;
&lt;p&gt;Online Appendix A investigates this by estimating the implicit place-based transfer induced by marital and parental status differences between EZ communities and the rest of the country. Using ACS tract-level data merged with Piketty-Saez-Zucman distributional national accounts (DINA), the authors find that marital status and parental status have offsetting effects: marital status raises taxes on single households (common in Distressed), while parental status increases transfers to households with children (also common in Distressed). Across all preferred CPS-adjusted estimates, net within-earnings transfers are below $1,000 in magnitude, and the two factors essentially cancel. The authors conclude that marital and parental status differences do not yield substantial de facto place-based redistribution within income levels.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-mtr-decomposition-table-3-reveal-about-why-sorting-motives-generate-different-mtr-patterns"&gt;Q11. What does the MTR decomposition (Table 3) reveal about why sorting motives generate different MTR patterns?&lt;/h3&gt;
&lt;p&gt;The decomposition separates the optimal MTR into a within-community component (standard equity-efficiency tradeoff) and a between-community component (fiscal externality from migration). Under skill-taste correlation with high migration (κ = 0.5), both components contribute positively to the Distressed MTR (0.246 within + 0.234 between = 0.479), yielding lower MTRs in Distressed (0.479) than in Elsewhere (0.510). Under comparative advantage with high migration, the within-community component is negative (−0.111) because high MTRs at the optimum reduce the concentration of high-skill types in Distressed, depressing the standard revenue-raising benefit of MTRs. The large positive between-community component (0.655) reflects the large fiscal externality from movers and overcomes this, yielding higher Distressed MTRs (0.544 vs. 0.509 in Elsewhere). With low migration (κ = 4), between-community components shrink substantially, and MTRs in Distressed fall below Elsewhere in all sorting scenarios.&lt;/p&gt;
&lt;h3 id="q12-what-does-the-crosswalk-from-urban-to-rural-baseline-reveal-about-which-assumptions-drive-the-change-in-optimal-transfers"&gt;Q12. What does the crosswalk from urban to rural baseline reveal about which assumptions drive the change in optimal transfers?&lt;/h3&gt;
&lt;p&gt;Table 5 traces the urban-to-rural transition step by step. Starting from the urban baseline ($3,143 average transfer), replacing the migration elasticity target with the rural value of 0.20 triples the optimal transfer to $9,870. Subsequently replacing skill-taste correlation with comparative advantage as the sorting mechanism reduces the transfer by roughly half ($6,402). Adjusting rent to match the rural ratio (0.54) reduces it further to $2,780, as lower Distressed rent reduces the marginal utility of consumption at the bottom and increases income-based sorting. Targeting the rural income share (60% below $50K) raises it back to $4,140, and incorporating rural housing supply elasticities yields the rural baseline result of $4,329. This decomposition reveals that lower migration responsiveness is the single largest driver of higher optimal transfers in rural settings.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Place-based redistribution&lt;/strong&gt;: Transfer schemes in which economic benefits or tax burdens are conditioned on the geographic location of residence, as distinct from place-blind income taxes that condition only on earned income. In this paper, modeled as location-specific tax schedules T_j(z) that may differ across communities j.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Skill-taste correlation&lt;/strong&gt;: A source of spatial sorting in which households with higher skill levels (θ) have systematically stronger preferences for the &amp;ldquo;Elsewhere&amp;rdquo; location, independently of wage or rent differences. Formally, the conditional distribution G_θ(φ) of locational tastes given skill is weakly increasing in θ. This correlation breaks the Atkinson-Stiglitz sufficient condition for commodity tax superfluousness and generates unambiguously positive optimal transfers to Distressed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Comparative advantage (spatial)&lt;/strong&gt;: A sorting mechanism in which higher-skill workers are disproportionately more productive in Elsewhere than in Distressed, captured by the wage elasticity with respect to skill being higher in Elsewhere (γ₀(θ) &amp;gt; γ₁(θ)). Households with skill above a threshold sort into Elsewhere even with homogeneous locational preferences. The existence of spatial comparative advantage means that migrants to Distressed earn less, creating a fiscal externality for place-based transfers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Income-based sorting&lt;/strong&gt;: Sorting of lower-income, lower-skill households into Distressed arising purely from the higher cost of living in Elsewhere, without any systematic skill-taste correlation or comparative advantage. Because high-skill households are less sensitive to rent differences, they sort into Elsewhere when rents there are higher. When this is the sole sorting mechanism and locational preferences are homogeneous, the Atkinson-Stiglitz commodity tax superfluousness conditions are satisfied and optimal place-based transfers are zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fiscal externality (migration)&lt;/strong&gt;: The change in income tax revenue caused by migration responses to place-based policy changes, not by changes in incentives for stayers. When movers from Elsewhere to Distressed earn less in their new location, they generate lower tax payments, imposing a first-order cost on the government budget. This externality is measured by Δτ+(θ) in the optimal MTR formulas and equals the earnings-tax-revenue loss from movers across all skill levels above θ. This term is a &amp;ldquo;sufficient statistic&amp;rdquo; for the efficiency cost of place-based transfers in the sense of Chetty (2009).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Demogrant (∆₀)&lt;/strong&gt;: The difference in lump-sum transfers provided to zero-earners across the two locations (−T₀(0) − (−T₁(0)) = T₀(0) − T₁(0)). A positive ∆₀ means Distressed provides a larger transfer to non-earners. It represents the place-based redistribution that occurs at the bottom of the earnings distribution, independently of MTR differences. In the paper&amp;rsquo;s decomposition, total optimal place-based redistribution (∆_z) exceeds ∆₀ when Distressed also has lower MTRs, meaning redistribution grows with income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Income-constant average tax difference (∆_z)&lt;/strong&gt;: The paper&amp;rsquo;s preferred summary measure of the average place-based transfer, defined as an equally weighted average of two tax-difference indices: the tax difference evaluated at Elsewhere earnings levels and the tax difference evaluated at Distressed earnings levels. This measure isolates tax schedule differences from productivity differences across locations, avoiding conflation of tax policy and wage effects on measured income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Landlord welfare weight (λ_L)&lt;/strong&gt;: The social marginal welfare weight assigned to landlords relative to the multiplier on the government budget constraint. When λ_L &amp;lt; 1, the planner values a marginal dollar of public funds more than a marginal dollar to landlords, creating a motive to use place-based taxes to shift rent incidence. The rent redistribution effect on optimal MTRs operates through the term (λ_L − 1)Δr+(θ), which has opposite signs in urban (positive) and rural (negative) distressed areas because of their different housing supply elasticities.&lt;/p&gt;</description></item><item><title>Policy Diffusion and Polarization across U.S. States</title><link>https://macropaperwarehouse.com/papers/policy-diffusion-and-polarization-across-u.s.-states/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/policy-diffusion-and-polarization-across-u.s.-states/</guid><description>&lt;p&gt;DellaVigna and Kim study the innovation and diffusion of policies across U.S. states using a dataset of over 700 state laws spanning seven decades. The central question is what predicts whether a state adopts a policy — and how those predictors have changed over time. The paper draws on two primary data sources: the State Policy Innovation and Diffusion (SPID) Database (Boehmke et al., 2020), covering 676 policies, and a hand-collected sample of 57 policies from 91 NBER working papers (April 2012–September 2021) that feature state-level policy variation. The combined dataset covers 733 policies adopted from the 1950s onward across the contiguous 48 states.&lt;/p&gt;
&lt;p&gt;On policy innovation, the paper finds that state capacity plays only a small role: larger and richer states are only slightly more likely to introduce new policies, innovation originates from both Republican and Democratic states, and the patterns are largely idiosyncratic with respect to observable state characteristics. California is the most frequent innovator, but large states like Florida and Texas rank in the middle.&lt;/p&gt;
&lt;p&gt;For policy diffusion, the paper employs both a static Geary&amp;rsquo;s C clustering statistic (measuring whether the first 10 adopting states cluster geographically or politically relative to a random-diffusion benchmark) and a dynamic logit hazard model estimated separately by decade. The hazard model identifies three similarity channels — geographic, demographic, and political — and allows their coefficients to vary over time.&lt;/p&gt;
&lt;p&gt;The central finding is a structural break in diffusion patterns around 2000. From the 1950s to the 1990s, geographic proximity is the dominant predictor of policy adoption: the coefficient on geographic similarity is 0.34 in the 1970s and remains roughly constant at 0.33 in the most recent decade. Demographic similarity is consistently positive and stable (approximately 0.20 in the 1980s, 0.22 in the 2010s). Political similarity — measured by closeness in Republican vote-share from the most recent presidential election — is a modest predictor before 2000, with coefficients between 0.14 (1970s) and 0.17 (1990s). Since 2000, the political similarity coefficient triples: 0.46 in the 2000s and 0.52 in the 2010s, making it by far the strongest predictor. The overall pseudo R-squared rises from 0.13 in the 1970s to 0.19 in the 2010s.&lt;/p&gt;
&lt;p&gt;These patterns are more pronounced for policies studied by economists: in the NBER subsample, the political similarity coefficient reaches 0.66 (s.e.=0.09) in the most recent two decades, versus 0.42 (s.e.=0.04) in the SPID sample.&lt;/p&gt;
&lt;p&gt;The paper tests whether the increased role of political similarity reflects correlated voter preferences, learning, or competition versus party discipline. Against correlated-preferences explanations: adding cross-state migration flows as a similarity measure reduces geographic predictive power but leaves the political similarity coefficient entirely unchanged; and typical policy-outcome variables (poverty rate, opioid mortality, income) have not become more correlated among politically similar states over time. In favor of party discipline: similarity in unified state government has zero predictive power through the 1990s but a coefficient of 0.42 (s.e.=0.06) in the 2000s–2010s. An event study of switches to unified party control confirms this causally for 1991–2020: switching to unified government raises the probability of passing ideologically aligned laws by approximately 2 percentage points in the four years following the switch, with no pre-trends and no effect on neutral-leaning laws; the same event study for 1950–1990 yields no detectable effect.&lt;/p&gt;
&lt;p&gt;COVID policies (77 state laws since October 2019) show strong political similarity in adoption; historical vaccination mandate policies (28 laws since 1975) show no political similarity effect. The paper concludes that rising party polarization at the state level — detectable from the 2000s onward, lagging the Congressional trend by roughly four to five decades — is the primary driver of the shift in diffusion patterns. The authors additionally classify each of the 57 NBER-sample policies by type of diffusion as an input for difference-in-differences research design assessment.&lt;/p&gt;
&lt;p&gt;Q: What data do the authors use and what is its scope?
A: The main source is the SPID Database (Boehmke et al., 2020), covering 676 policies over seven decades. The authors supplement this with 57 policies hand-collected from 91 NBER working papers (2012–2021) that use state-level policy variation. The combined sample covers 733 policies adopted from the 1950s onward in the contiguous 48 states, with the SPID sample averaging 23 adopting states per policy and the NBER sample averaging 29.&lt;/p&gt;
&lt;p&gt;Q: Do states with more resources or larger populations systematically innovate more policies?
A: The evidence for a state-capacity hypothesis is weak. There is only suggestive evidence that higher per-capita income predicts being in the top-20% of innovators, and no clear difference in population between the top and bottom innovators. Innovations arise from both Republican and Democratic states. One consistent correlate is urban population share, but overall innovation is largely idiosyncratic with respect to observable characteristics.&lt;/p&gt;
&lt;p&gt;Q: What was the dominant predictor of policy diffusion before 2000?
A: Geographic proximity was the dominant predictor. The coefficient on geographic similarity in the hazard model is 0.34 in the 1970s and remains stable at approximately 0.33 in the 2010s. Demographic similarity contributes consistently at approximately 0.20. Political similarity before 2000 is modest, ranging from 0.14 in the 1970s to 0.17 in the 1990s — roughly one-third to one-half the magnitude of the geographic coefficient.&lt;/p&gt;
&lt;p&gt;Q: How dramatically does political similarity change after 2000, and is this finding robust?
A: The political similarity coefficient triples, rising from 0.17 in the 1990s to 0.46 in the 2000s and 0.52 in the 2010s, making it the largest single predictor in recent decades. This pattern is robust across linear probability models, alternative measures of political similarity, alternative thresholds for &amp;ldquo;closest&amp;rdquo; states (closest fifth, fourth, third, or half all yield comparable coefficients), and alternative ways of computing adoption counts.&lt;/p&gt;
&lt;p&gt;Q: Is the shift toward political diffusion stronger for policies economists study?
A: Yes. In the NBER subsample, the political similarity coefficient reaches 0.66 (s.e.=0.09) in the 2000s–2010s, compared to 0.42 (s.e.=0.04) in the SPID sample. Geographic similarity also has somewhat higher coefficients in the NBER sample throughout the period. This implies that the policies most studied for difference-in-differences evaluation are also those most subject to politically-driven diffusion.&lt;/p&gt;
&lt;p&gt;Q: What does the Medicaid case study illustrate about political polarization?
A: ACA Medicaid expansion spread almost exclusively along partisan lines, with Republican vote-share accurately predicting the year of adoption. Crucially, the states that delayed or declined adoption — higher Republican vote-share states — had a higher share of population that would benefit from the expansion and therefore face a worse policy-need match. By contrast, the original 1966 Medicaid rollout showed no relationship between state political leaning and timing of adoption, and neither did the 1960s–1970s food stamp program expansion.&lt;/p&gt;
&lt;p&gt;Q: How do the authors distinguish party discipline from correlated voter preferences as the mechanism?
A: Two tests point away from correlated preferences: (1) cross-state migration flows, when added as a similarity measure, absorb geographic predictive power but leave the political similarity coefficient entirely unaffected; (2) typical policy-outcome variables (opioid mortality, poverty rate, income, etc.) have not become more correlated among politically similar states over time, contradicting the hypothesis that local needs or environments have become politically correlated.&lt;/p&gt;
&lt;p&gt;Q: What is the direct evidence for party discipline as the operative mechanism?
A: The authors construct a measure of similarity based on unified party control (governor and both chambers of the same party). This variable has zero predictive power through the 1990s (point estimate near zero). In the 2000–2020s, the coefficient for unified-government similarity is 0.42 (s.e.=0.06), making it the strongest single predictor of adoption in those decades. States with divided governments show no predictive power of adoption by other divided-government states, further isolating the role of party control.&lt;/p&gt;
&lt;p&gt;Q: What does the event-study of switches to unified party control show?
A: Switches to unified party control in 1991–2020 produce a statistically significant increase of approximately 2 percentage points in the probability of adopting ideologically aligned laws within four years of the switch, relative to the year before. The effect emerges in year t+1 and is persistent, with no pre-trends, and the effect on neutral-leaning laws is zero, ruling out a simple reduced-gridlock story. The same event study for 1950–1990 detects no effect.&lt;/p&gt;
&lt;p&gt;Q: How do COVID state policies compare to historical vaccination policies in terms of political diffusion?
A: COVID policies (77 state laws, October 2019–August 2021) show significant political similarity in adoption, consistent with the recent-decade patterns. Vaccination mandate laws (28 policies since 1975) show no political similarity effect whatsoever, with demographic and modest geographic similarity being the relevant predictors. This contrast underscores that political polarization in policy adoption is a recent phenomenon that has spread even to policy areas without prior partisan patterning.&lt;/p&gt;
&lt;p&gt;Q: How does partisan polarization at the state level compare temporally to polarization in Congress?
A: Congressional polarization (measured by DW-NOMINATE) has been rising since the 1950s. State-level policy polarization, as documented here, does not emerge until the 2000s — a lag of roughly four to five decades. The paper notes it has risen rapidly and has already reached policy domains (such as COVID mandates) that showed no political patterning historically.&lt;/p&gt;
&lt;p&gt;Q: Does the diffusion pattern vary across policy types?
A: Yes. For economic policies, geography and demographics decline in importance over time with a smaller increase in political predictors. For non-economic (social) policies, geographic importance remains stable while political polarization is especially strong. Political polarization is strongest in Republican-leaning and Democratic-leaning states, and weaker among battleground states, consistent with a party-driven model where ideologically extreme states adopt from each other.&lt;/p&gt;
&lt;p&gt;Q: How do the authors classify individual NBER-sample policies by diffusion type?
A: Using Geary&amp;rsquo;s C statistics computed separately for geographic and political clustering for each of the 57 NBER policies, the authors identify three approximate clusters: (1) primarily politically-clustered (e.g., Medicaid expansion); (2) jointly geographically and politically clustered (e.g., ban on asking about past salary history); and (3) largely idiosyncratic, neither geographically nor politically clustered (e.g., anti-bullying laws). This classification has direct implications for assessing identification threats in difference-in-differences designs.&lt;/p&gt;
&lt;p&gt;Q: What does the overall predictability of policy adoption look like over time?
A: The pseudo R-squared from the logit hazard model rises from 0.13 in the 1970s to 0.19 in the 2010s. The increase in political similarity is large enough not only to surpass geographic similarity as a predictor but to make the overall process of state policy adoption more predictable over time.&lt;/p&gt;
&lt;p&gt;Policy diffusion: The process by which a policy adopted in one state subsequently spreads to other states; measured here along geographic, demographic, and political dimensions using a logit hazard model estimated by decade.&lt;/p&gt;
&lt;p&gt;Geary&amp;rsquo;s C statistic: A ratio of weighted to unweighted average pairwise squared differences in adoption status, adapted from spatial statistics (Geary, 1954). Values below 1 indicate clustering; values above 1 indicate anti-clustering. The paper reports 1−C so higher values mean more clustering among similar states.&lt;/p&gt;
&lt;p&gt;Policy innovation: First-year adoption of a law in any state; a state is an &amp;ldquo;innovator&amp;rdquo; if it adopts in the first year the policy appears anywhere. The paper distinguishes innovation (origination) from diffusion (spread).&lt;/p&gt;
&lt;p&gt;Logit hazard model: A discrete-time logit model estimated at the state-year-policy level for all states that have not yet adopted a given policy, with policy-decade fixed effects as a baseline hazard and three time-varying similarity measures (geographic, demographic, political) as key predictors.&lt;/p&gt;
&lt;p&gt;Political similarity: Closeness of two states&amp;rsquo; Republican vote-shares from the most recent presidential election; the closest third of states in this dimension are used to construct the diffusion measure. Shown to be independent of — and to have grown far more predictive than — geographic similarity since 2000.&lt;/p&gt;
&lt;p&gt;Unified party control: A state government in which the governor and both state legislative chambers belong to the same party. The paper shows this is the variable most predictive of politically-driven policy diffusion in the 2000s–2020s, with a coefficient of 0.42 where it was effectively zero before 2000.&lt;/p&gt;
&lt;p&gt;Party discipline / party polarization: The paper&amp;rsquo;s preferred explanation for post-2000 patterns: state politicians increasingly vote and adopt policies along party lines beyond what voter preferences alone would predict, with the effect detectable since the 2000s at the state level, lagging the Congressional polarization trend by roughly four decades.&lt;/p&gt;</description></item><item><title>Politics at Work</title><link>https://macropaperwarehouse.com/papers/politics-at-work/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/politics-at-work/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Do individual political views shape firm behavior and labor market outcomes in the private sector? Specifically, do business owners sort copartisan workers into their firms, and does employers&amp;rsquo; political discrimination drive this sorting?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Setting&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper studies the complete Brazilian formal labor market over 2002–2019, assembling a novel longitudinal worker-firm-owner-party matched dataset from three administrative sources: (1) RAIS (Relação Anual de Informações Sociais), the universe of formal-sector workers (87 million unique workers, 7.6 million unique firms); (2) the Receita Federal do Brazil (RFB) and Cadastro Nacional de Empresas (CNE), containing business ownership structures for all registered firms; and (3) the Tribunal Superior Eleitoral (TSE) registry of all party members (19.3 million individuals) over 2002–2019. Matching these sources yields political affiliation for 11.4% of all private-sector owners and 7.8% of all private-sector workers in the sample. Party affiliation in Brazil requires an active registration step and is interpreted as a signal of strong and visible political views, distinguishing affiliated from unaffiliated individuals who likely hold milder views. The 35 parties in the sample are highly fragmented; the top 7 account for nearly 70% of all party members.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Political assortative matching.&lt;/em&gt; Using a likelihood ratio index (Eika et al., 2019; Chiappori et al., 2020), the paper finds that workers and owners belonging to the same party are on average about twice as likely to match in the labor market relative to random matching. Once within-municipality geographical sorting is accounted for, this figure falls to approximately 55% excess probability of copartisan matching, and increases over time: from 1.41 in 2002–2006 to 1.67 in 2016–2019. A dyadic regression approach — constructing all worker-firm dyads within industry-municipality labor markets and controlling for shared gender, race, age, and education — confirms the result: across all years, a politically affiliated worker is between 41% and 75% more likely to be employed by a copartisan owner than by an owner affiliated with a different party. Political assortative matching is driven both by higher hiring probabilities (range: 32%–59% more likely for copartisans, hiring margin only) and by longer tenure: copartisan workers stay in the firm roughly 5.5% longer than otherwise comparable workers of a different party, even within the same firm and hire-year (column 3 of Table 2). In every year and by every method, the degree of political assortative matching exceeds that of gender (15%–31% excess probability under dyadic approach) and race (approximately 3.4%), which are themselves both positive and significant.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Mechanisms: political discrimination.&lt;/em&gt; Three sets of evidence point to employer political discrimination as a relevant driver. First, in the administrative micro-data: assortative matching decreases strongly with firm size — it is more than twice as large in firms with up to 10 employees than in medium firms and more than six times as large as in firms with more than 50 employees — and is stronger for higher occupational layers and for jobs requiring above-median social skills or interpersonal relationships. Political assortative matching is, if anything, larger for parties not in power locally, inconsistent with a patronage mechanism. An event study of 5,262 owners who switched party finds a sharp increase of about 0.2 standard deviations in hires from the new party and a corresponding drop in hires from the old party at the time of the switch, with the share of workers from the new party rising by roughly 5 percentage points persistently. Second, an incentivized resume rating (IRR) field experiment (150 business owners; nondeceptive design) shows that owners rate copartisan resumes 0.213 points higher on a 1–7 Likert scale (a 7.4% increase relative to the mean rating for different-party resumes, statistically significant at p &amp;lt; 0.05), with no significant effect on perceived candidate acceptance probability. Third, a representative survey of 891 owners and 1,003 workers finds that belief-based and taste-based discrimination are ranked as the leading explanations by both groups; 47% of owners and 58% of workers agree with the belief-based discrimination statement. Additionally, 29% of surveyed owners (22% say &amp;ldquo;Yes&amp;rdquo; and 7% &amp;ldquo;In some cases&amp;rdquo;) explicitly reveal that political views affect their hiring decisions.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Real consequences.&lt;/em&gt; Conditional on employment, copartisan workers are promoted faster: they are 0.448 percentage points more likely to be promoted from white-collar to managerial positions (against a base rate of 2.58%) and 0.44 percentage points more likely to be promoted from blue-collar to white-collar positions (base rate 2.98%). Workers from a different party than the owner face a promotion penalty of 0.104–0.180 percentage points for white-collar-to-manager promotions. On wages, copartisan workers earn 3.9% more than unaffiliated coworkers within the same firm and year (firm-year FE specification); the effect is 2.8% when restricting to the same occupation within the firm. Workers from a different party earn 1.6% less. Decomposing by tier: managers (copartisan premium 1.6%), white-collar workers (3.4%), blue-collar workers (1.5%). Despite better outcomes, copartisan workers are 2.1 percentage points (2.3% relative to the mean) less likely to be educationally qualified for their occupation, conditional on firm-year and controlling for a full set of demographics. Finally, a higher share of copartisan workers in the prior year is associated with lower firm employment growth (estimated β = −0.071), corresponding to approximately a 1 percentage point gap in annual growth rate for a one-standard-deviation difference in copartisan share — substantial relative to an average annual growth rate of 10%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;All findings pertain to the formal private sector in Brazil over 2002–2019. Political affiliation in the Brazilian system requires an active step and signals strong views; results apply to the approximately 7.8%–11.4% of workers and owners who are party-registered. The field experiment sample is limited to 150 business owners affiliated with major Brazilian parties who were actively seeking to hire. The firm growth result is explicitly characterized as suggestive, without a source of exogenous variation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-likelihood-ratio-index-and-what-does-it-show-for-political-matching-in-brazil"&gt;Q1. What is the likelihood ratio index and what does it show for political matching in Brazil?&lt;/h3&gt;
&lt;p&gt;The likelihood ratio index measures how many times more likely a match between a worker and owner of the same party is, relative to the expected frequency under random matching (conditional on the population shares of each party). Across 2002–2019, the unconditional index ranges from 1.56 to 1.85, implying workers and employers of the same party are on average about twice as likely to match as under random matching. After accounting for geographic sorting within municipalities, the index ranges from approximately 1.41 (2002–2006 average) to 1.67 (2016–2019 average), showing a clear increasing trend. The corresponding gender and race indexes average about 1.2 and 1.35, respectively, in the basic specification, both significantly lower than the party index in every year of the sample.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-dyadic-regression-estimates-control-for-omitted-characteristics-and-what-do-they-find"&gt;Q2. How do the dyadic regression estimates control for omitted characteristics, and what do they find?&lt;/h3&gt;
&lt;p&gt;The dyadic regression constructs all possible worker-firm pairs within each municipality-industry labor market in a given year. The dependent variable is an indicator for whether worker i is employed by firm f. The key coefficient of interest is the differential probability of employment for a copartisan pair relative to a different-party pair, controlling for indicators for shared gender, race, age bracket, and education level, as well as worker occupation fixed effects and experience. This controls for the concern that politically affiliated individuals share non-political traits that correlate with employment choices. After these controls, a politically affiliated worker is 41%–75% more likely (depending on year) to be employed by a copartisan owner than by a different-party owner. The effect stems primarily from copartisan workers being preferentially hired (not just from unaffiliated owners preferring any affiliated worker indiscriminately). The analogous dyadic estimate for shared gender is 15%–31% and for shared race is approximately 3.4%, both lower than the party estimate in all years.&lt;/p&gt;
&lt;h3 id="q3-how-is-political-assortative-matching-decomposed-into-hiring-versus-retention-margins"&gt;Q3. How is political assortative matching decomposed into hiring versus retention margins?&lt;/h3&gt;
&lt;p&gt;To isolate the hiring margin, the authors estimate the dyadic regression restricting to newly hired workers (not present in the firm in year t-1). They find that the probability of being hired by a copartisan owner is 32%–59% higher than by a different-party owner across years. The retention (tenure) margin is estimated by regressing the share of subsequent years a worker remains at the firm on partisan alignment at the time of hire. In the most stringent specification (year-of-hire × firm fixed effects), copartisan hires stay 5.5 percentage points longer (as a share of post-hire years) than different-party hires from the same firm and hire-year cohort. Both margins are significant, and both exhibit stronger political sorting than equivalent estimates for gender or race.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-evidence-against-political-patronage-as-the-primary-driver-of-political-assortative-matching"&gt;Q4. What is the evidence against political patronage as the primary driver of political assortative matching?&lt;/h3&gt;
&lt;p&gt;If political patronage (parties pressuring owners to hire copartisans) were the main driver, we would expect political assortative matching to be stronger when the owner&amp;rsquo;s party is in power locally, as those parties have greater leverage over business owners. The authors estimate a modified dyadic regression distinguishing between cases where the owner&amp;rsquo;s party is in the ruling coalition of the municipal mayor or state governor versus not in power. The results show that political assortative matching is, if anything, larger for parties not in power. This is inconsistent with patronage being the dominant mechanism and consistent with the discrimination channel being driven by owner preferences rather than external political pressure.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-event-study-of-owner-party-changes-show"&gt;Q5. What does the event study of owner party changes show?&lt;/h3&gt;
&lt;p&gt;The event study tracks 5,262 owners who switch party affiliation during 2002–2019, comparing their firms to control firms in the same market whose owners remain affiliated to the original party. At the time of the switch, there is a sharp increase of approximately 0.2 standard deviations in hires from the owner&amp;rsquo;s new party and a corresponding sharp decrease in hires from the old party. Hires from other parties and unaffiliated hires also decline modestly. The share of the workforce affiliated with the new party increases by roughly 5 percentage points and remains elevated in subsequent years. Because nonpolitical network ties (shared school, neighborhood, sports team) are unlikely to dissolve abruptly when an owner changes party, this design provides additional evidence that the change in hiring is driven by a direct change in the owner&amp;rsquo;s political preferences rather than by network overlap.&lt;/p&gt;
&lt;h3 id="q6-what-was-the-design-of-the-incentivized-resume-rating-experiment-and-why-does-it-identify-political-discrimination"&gt;Q6. What was the design of the incentivized resume rating experiment and why does it identify political discrimination?&lt;/h3&gt;
&lt;p&gt;The experiment was conducted with 150 Brazilian business owners recruited from the administrative data (who are already known to be affiliated with one of six major parties), targeting owners with active hiring interest through a leading job platform. Owners rated 20 synthetic resumes with fully randomized features (education, experience, training, skills, formatting). Sixteen resumes had no partisan cues; two contained cues signaling copartisanship with the rating owner; two signaled a party from the opposite side of the political spectrum. Incentives were provided by committing to send respondents real job-seeker profiles from the platform chosen by machine learning based on revealed preferences. Because all resume features other than the partisan cue were randomized, the experiment shuts down shared nonpolitical networks and patronage as explanations; the only channel is the employer&amp;rsquo;s direct preference for the candidate&amp;rsquo;s partisan affiliation. The response rate was 11% and the survey was conducted March–May 2022.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-quantitative-magnitude-of-the-field-experiment-result"&gt;Q7. What is the quantitative magnitude of the field experiment result?&lt;/h3&gt;
&lt;p&gt;Owners rate copartisan resumes 0.213 points higher on the 1–7 Likert scale relative to resumes from the opposite side of the political spectrum (statistically significant at p &amp;lt; 0.05), representing a 7.4% increase relative to the mean rating of different-party resumes (2.950). When resume-level controls (gender, high-skill experience flag, years of experience, programming skills, training) are added, the estimate is 0.254. There is no statistically significant effect on owners&amp;rsquo; perceived likelihood that a candidate would accept a job offer (coefficient 0.150–0.158, not significant), suggesting that the observed difference in interest ratings reflects a genuine direct preference for copartisans, not an expectation that copartisans are more likely to accept.&lt;/p&gt;
&lt;h3 id="q8-what-do-the-survey-findings-add-about-mechanisms-and-the-prevalence-of-political-discrimination"&gt;Q8. What do the survey findings add about mechanisms and the prevalence of political discrimination?&lt;/h3&gt;
&lt;p&gt;The survey of 891 owners and 1,003 workers (response rate 26.84%) presents five candidate mechanisms and asks respondents to evaluate each. Both groups rank belief-based discrimination (owners believe copartisans would be more productive) as the most likely explanation: 47% of owners and 58% of workers partially or strongly agree. Taste-based discrimination is second (36% owners, 52% workers agree), followed by networks (39% owners, 49% workers). Patronage and workers&amp;rsquo; preferences attract little agreement from either group. Among owners ranked by single strongest agreement, 29.7% most strongly agree with belief-based discrimination and 22.0% with taste-based, while 29% of all surveyed owners explicitly stated that political views do affect their hiring decisions. These patterns are broadly similar regardless of the respondent&amp;rsquo;s own political affiliation status.&lt;/p&gt;
&lt;h3 id="q9-how-large-are-the-political-promotion-and-wage-premia-and-how-do-they-compare-to-gender-and-race-effects"&gt;Q9. How large are the political promotion and wage premia, and how do they compare to gender and race effects?&lt;/h3&gt;
&lt;p&gt;For promotions, copartisan white-collar workers are 0.448 percentage points more likely to be promoted to manager (relative to unaffiliated co-workers hired in the same firm-year), against a base promotion rate of 2.58% — an effect of approximately 17% of the mean. For blue-collar-to-white-collar promotion, the copartisan premium is 0.44 percentage points against a base rate of 2.98%. For wages, copartisans earn 3.9% more than unaffiliated co-workers within the same firm and year; restricting to the same occupation within the firm, the premium is 2.8%. The political wage premium (3.9%) exceeds the gender wage premium (1.5%) and the race wage premium (1.0%) in the same specification. Workers from a different party than the owner earn 1.6% less than unaffiliated co-workers within the same firm-year.&lt;/p&gt;
&lt;h3 id="q10-are-copartisan-workers-better-qualified-than-those-they-displace-and-what-does-this-imply-for-firm-performance"&gt;Q10. Are copartisan workers better qualified than those they displace, and what does this imply for firm performance?&lt;/h3&gt;
&lt;p&gt;Copartisan workers are significantly less qualified in terms of education relative to their occupation: they are 2.1 percentage points less likely to be educationally qualified for their position than their unaffiliated co-workers within the same firm-year (2.3% relative to the mean qualification rate of 93.2%), with the largest effects for managers. Workers of a different party show only a small and economically negligible qualification gap. The fact that copartisans are paid more, promoted faster, and yet are less qualified is consistent with political discrimination substituting for competence in personnel decisions. The qualification shortfall is specifically attributed to copartisanship and not to shared gender, race, age, or education between owner and worker, as those coefficients are economically small.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-evidence-on-firm-growth-and-what-are-the-limitations-of-that-evidence"&gt;Q11. What is the evidence on firm growth and what are the limitations of that evidence?&lt;/h3&gt;
&lt;p&gt;Firms with a higher share of copartisan workers in the prior year grow less. The estimated coefficient β = −0.071, and a one-standard-deviation difference in the copartisan share is associated with approximately a 1 percentage point gap in annual employment growth, relative to a mean growth rate of 10%. The specification compares firms of the same size and with the same number of affiliated workers in the same year. The result is robust to adding municipality and municipality-industry fixed effects. The authors explicitly characterize this evidence as suggestive, noting the absence of an exogenous source of variation in political discrimination. The negative association is more consistent with taste-based discrimination (Becker, 1957) — in which politically homogeneous firms sacrifice productivity for the owners&amp;rsquo; amenity of employing copartisans — than with accurate belief-based discrimination.&lt;/p&gt;
&lt;h3 id="q12-how-is-political-assortative-matching-distributed-across-parties-and-does-it-depend-on-party-ideology"&gt;Q12. How is political assortative matching distributed across parties and does it depend on party ideology?&lt;/h3&gt;
&lt;p&gt;The likelihood ratio index shows large assortative matching across the entire political spectrum. For most years, relatively more ideologically extreme parties — on the left (PT, PDT) and on the right (PP, DEM) — display higher assortative matching than more centrist parties (PMDB, PSDB). This pattern is consistent with stronger partisan identity at the extremes leading to stronger preferences for copartisan workers, but the paper does not formally model the mechanism behind this heterogeneity.&lt;/p&gt;
&lt;h3 id="q13-what-is-the-role-of-workers-preferences-as-opposed-to-employers-discrimination-and-how-can-wages-distinguish-them"&gt;Q13. What is the role of workers&amp;rsquo; preferences as opposed to employers&amp;rsquo; discrimination, and how can wages distinguish them?&lt;/h3&gt;
&lt;p&gt;If workers have a preference for working with copartisan owners (treating this as a job amenity), compensating differentials theory would predict a negative wage premium for copartisan workers — they would accept lower wages in exchange for working with like-minded owners. The data show the opposite: copartisan workers earn significantly more, not less, than their unaffiliated co-workers. This evidence is inconsistent with workers&amp;rsquo; preferences being the primary driver of political assortative matching, and is instead consistent with employers&amp;rsquo; discrimination. The survey evidence corroborates this: both owners and workers assign low priority to the &amp;ldquo;workers&amp;rsquo; preferences&amp;rdquo; mechanism.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Political assortative matching&lt;/strong&gt;: The phenomenon by which workers and business owners belonging to the same political party are matched in the labor market at rates significantly exceeding what would occur under random matching within the local labor market. Measured via the likelihood ratio index and dyadic regressions that control for shared demographic characteristics. In this paper, political assortative matching is larger in magnitude than assortative matching along gender or racial lines.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Likelihood ratio index (S)&lt;/strong&gt;: A measure of assortative matching defined as the weighted sum of the ratios of observed same-party co-occurrence probabilities to their expected probabilities under random matching. S &amp;gt; 1 indicates positive assortative matching. The paper uses both a basic version and a geography-adjusted version that computes the index within municipalities to control for geographic concentration of party membership.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dyadic regression&lt;/strong&gt;: A regression approach that constructs all possible worker-firm pairs within a defined labor market (municipality × 2-digit industry) to estimate the differential probability that a worker is employed by a copartisan firm relative to a different-party firm. The key advantage is the ability to control simultaneously for multiple shared demographic characteristics between worker and owner, accounting for the correlation of assortative criteria.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Incentivized resume rating (IRR) experiment&lt;/strong&gt;: A nondeceptive field experiment design (following Kessler et al., 2019) in which business owners rate synthetic resumes with fully randomized characteristics. Truthful rating is incentivized because respondents are told that their revealed preferences will be used to select real job-seeker profiles sent to them by a partner platform via machine learning. This design allows direct identification of employer preference for copartisan candidates while ruling out alternative channels such as shared nonpolitical networks or patronage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Political wage premium&lt;/strong&gt;: The percentage wage difference earned by copartisan workers relative to unaffiliated co-workers within the same firm-year (and occupation), after controlling for a full set of socio-demographic characteristics. A positive political wage premium is the paper&amp;rsquo;s primary piece of evidence that workers&amp;rsquo; compensating differentials cannot explain political assortative matching, since amenity-based sorting would predict a negative premium.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Political promotion premium&lt;/strong&gt;: The differential probability that a copartisan worker is promoted to a higher organizational layer (blue-collar to white-collar, or white-collar to manager) relative to an unaffiliated co-worker hired in the same firm and year, net of demographic controls.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Educational mismatch (Qualified)&lt;/strong&gt;: An indicator variable equal to one if a worker&amp;rsquo;s educational level meets or exceeds the educational level required by their specific occupation in the CBO (Classificação Brasileira de Ocupações) classification. Used to assess whether politically favored (copartisan) workers are less competent along this observable dimension.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Belief-based discrimination vs. taste-based discrimination&lt;/strong&gt;: Two distinct theoretical channels for employer political discrimination. Belief-based discrimination (Phelps, 1972; Arrow, 1973) occurs when employers perceive copartisans to be more productive — e.g., because shared political views reduce intra-firm conflict. Taste-based discrimination (Becker, 1971) occurs when employers have a direct utility-affecting preference for copartisan workers, independent of productivity beliefs. The paper treats these as observationally distinct from patronage and network overlap, and uses the negative correlation between political homogeneity and firm growth as suggestive evidence favoring the taste-based channel.&lt;/p&gt;</description></item><item><title>Professional Motivations in the Public Sector: Evidence from Police Officers</title><link>https://macropaperwarehouse.com/papers/professional-motivations-in-the-public-sector-evidence-from-police-officers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/professional-motivations-in-the-public-sector-evidence-from-police-officers/</guid><description>&lt;p&gt;This paper studies how public sector workers balance professional motivations against private economic concerns, using arrest decisions by Dallas Police Department (DPD) officers as the empirical laboratory. The central institutional feature exploited is that arrests made near the end of an officer&amp;rsquo;s shift typically require the officer to stay and work overtime, generating private costs that must be weighed against the professional benefits of making an arrest (e.g., crime reduction or duty fulfillment). The paper further leverages variation from DPD&amp;rsquo;s &amp;ldquo;secondary employment&amp;rdquo; program: approximately 30% of officers held a registered second job at some point during 2019–2021, and on days when a second job is scheduled after the police shift, the opportunity cost of late-shift policing is higher.&lt;/p&gt;
&lt;p&gt;The data cover all DPD arrests from January 2015 to December 2021, linked to officer shift assignments, charge types, prosecutorial outcomes (whether the Dallas County Attorney chose to prosecute), and second-job schedules. The sample excludes traffic violations and arrests without shift information. The authors observe wide variation in prosecution rates by charge type: drug and gang offenses exceed 70%, property and violent crimes run 30–50%, and minor charges fall below 20%.&lt;/p&gt;
&lt;p&gt;Four main findings emerge. First, arrest rates fall sharply in the last 30–40 minutes of a shift, with the decline most pronounced for drug and gang charges (approximately 50% drop in arrest rate) and smallest for violent charges, consistent with officers having more discretion over the former. Second, arrests that do occur late in the shift are of higher quality: conditional on being made, they are approximately 1.5–2.5 percentage points more likely to result in prosecution than arrests made earlier, with the quality premium larger in more discretionary charge categories (drugs/gang &amp;gt; property &amp;gt; violent). Third, on days when an officer has a second job scheduled, arrest rates are lower by roughly 5–10% relative to baseline across the full shift, with effects concentrated in the second half; and the conditional probability of prosecution on those days is 1–2 percentage points higher than on non-second-job days. The second-job effect appears even earlier in the shift than the overtime effect alone, consistent with the second job magnifying the opportunity cost mechanism.&lt;/p&gt;
&lt;p&gt;Fourth, the authors estimate a dynamic structural model of the arrest decision. At each moment of the shift the officer chooses whether to arrest, trading off a professional benefit b_p against a private cost c(t, secondjob) that rises when overtime begins and rises further on second-job days. Structural estimates indicate the overtime cost is large enough to reduce the expected professional value of an arrest in the final 30 minutes of the shift by roughly 20–30%. The additional second-job cost reduces expected professional value by a further 10–20%. Counterfactual simulation implies that eliminating the overtime cost would increase overall arrests by approximately 5–8%, a magnitude the authors describe as economically significant. Welfare analysis shows that the desirability of high overtime costs depends on whether citizens weight quantity of arrests or quality: under quality-weighted preferences the current overtime-cost regime may be socially optimal because officers self-select toward arrests they perceive as likely to result in prosecution; under quantity preferences, reducing overtime costs would increase police activity.&lt;/p&gt;
&lt;p&gt;The identification strategy relies on within-officer variation in second-job scheduling, absorbing officer fixed effects (and officer-by-month fixed effects in robustness checks) and time fixed effects. The key identifying assumption is that second-job days are not systematically assigned to low-crime or low-patrol days. Supporting evidence includes balance tests showing second-job status is uncorrelated with local crime call patterns conditional on fixed effects, and the observation that officers who take second jobs do not exhibit a systematically different enforcement style (measured by arrest patterns across the shift) relative to officers who do not.&lt;/p&gt;
&lt;p&gt;Scope conditions: results are from a single medium-sized urban police department (approximately 3,000 officers) in Dallas, Texas, a city described as diverse by race, income, and political affiliation. The department is 29% Black, 43% Hispanic, 27% White, and 15% female. Generalizability to other jurisdictions or institutional structures is not established by this study.&lt;/p&gt;
&lt;p&gt;Q: What is the main research question?
A: The paper asks how public sector workers balance professional motivations (e.g., crime reduction, duty fulfillment) against private economic concerns (e.g., overtime costs, opportunity costs from second jobs). It uses police arrest decisions as the empirical setting because the shift-end timing of arrests generates a clear, observable private cost that varies within officer across days.&lt;/p&gt;
&lt;p&gt;Q: What is the key institutional feature that generates identification?
A: Arrests made near the end of a shift typically require the arresting officer to stay past the shift and work overtime. This creates a personal cost — more time, delayed transition to off-duty activities — that makes late-shift arrests more costly without changing their professional value. The DPD secondary employment program adds a second source of variation: on days when an officer has a registered second job scheduled after the police shift, the opportunity cost of any arrest (and especially a late-shift arrest) is higher.&lt;/p&gt;
&lt;p&gt;Q: How large is the drop in arrest rates near shift end?
A: The baseline arrest rate declines by approximately 0.12 percentage points per six-minute time bucket in the last 30 minutes of the shift, or about 5% relative to the mean arrest rate of 2.3 percentage points. The drop is most dramatic for drug and gang charges, where the arrest rate falls by approximately 50%, and smallest for violent charges, where officers appear to arrest regardless of shift timing.&lt;/p&gt;
&lt;p&gt;Q: How does arrest quality change near shift end?
A: Arrests made in the last 30 minutes of a shift are approximately 1.5–2.5 percentage points more likely to result in prosecution than arrests made earlier in the shift, after controlling for charge type composition and officer fixed effects. The quality premium is larger in more discretionary charge categories (drugs/gang, then property, then violent), consistent with officers becoming more selective to avoid overtime costs on arrests unlikely to result in prosecution.&lt;/p&gt;
&lt;p&gt;Q: Does the shift-end drop reflect officer fatigue or overtime cost?
A: The paper argues both pieces of evidence point to overtime cost rather than fatigue alone. First, arrest rates increase sharply after the official shift end when the officer is already earning overtime pay — if fatigue were the mechanism, arrests would also decline post-shift. Second, on second-job days arrest rates fall earlier in the shift and by more, consistent with higher opportunity costs rather than accumulated fatigue.&lt;/p&gt;
&lt;p&gt;Q: What is the effect of having a second job scheduled on arrest rates?
A: Having a second job scheduled reduces arrest rates by roughly 5–10% relative to the baseline across the full shift, with effects concentrated in the second half. The reduction is even larger in the final 30 minutes, consistent with the second job amplifying the overtime cost mechanism.&lt;/p&gt;
&lt;p&gt;Q: What is the effect of second-job days on arrest quality?
A: Arrests made on second-job days are 1–2 percentage points more likely to result in prosecution compared to arrests on non-second-job days, after controlling for time of day, charge type composition, and officer fixed effects. This parallels the shift-end quality effect and is consistent with officers applying higher selectivity thresholds when opportunity costs are elevated.&lt;/p&gt;
&lt;p&gt;Q: How is the second-job variation used for identification?
A: The main specification compares the same officer&amp;rsquo;s behavior on shifts where a second job is scheduled versus shifts where it is not, absorbing officer fixed effects and time fixed effects. The identifying assumption is that second-job scheduling is uncorrelated with unobservable determinants of enforcement intensity conditional on fixed effects. The authors support this with balance tests showing second-job status is not predicted by lagged activity measures or contemporaneous crime call patterns.&lt;/p&gt;
&lt;p&gt;Q: What does the dynamic structural model add?
A: The structural model formalizes the arrest decision as a dynamic problem where the officer compares the professional benefit b_p of an arrest to the private cost c(t, secondjob), which rises discontinuously when overtime begins and rises further on second-job days. Estimating the model by matching moments (baseline arrest rates, shift-timing patterns, quality changes, second-job effects) yields preference parameters. The model enables counterfactual and welfare analysis that the reduced-form estimates alone cannot provide.&lt;/p&gt;
&lt;p&gt;Q: What are the structural estimates of overtime and second-job costs?
A: The overtime cost c_ot is estimated to be large enough that arresting someone in the final 30 minutes of the shift reduces the expected professional value of that arrest by roughly 20–30%. The additional second-job cost c_sj reduces expected professional value by a further 10–20%. Both estimates are described as statistically precise.&lt;/p&gt;
&lt;p&gt;Q: What does the counterfactual removal of overtime costs imply for arrests?
A: Eliminating the overtime cost is estimated to increase overall arrests by approximately 5–8%, which the authors characterize as economically significant. This implies that officers&amp;rsquo; private costs have a first-order impact on the quantity of law enforcement activity.&lt;/p&gt;
&lt;p&gt;Q: What does the welfare analysis conclude about overtime costs?
A: The welfare effect of eliminating overtime costs depends on citizen preferences. Under quality-weighted preferences — where citizens value the probability that an arrest results in prosecution — the current overtime-cost regime may be socially optimal because it induces officers to self-select toward arrests they perceive as likely to stick. Under quantity preferences — where citizens value the total number of arrests per period — reducing overtime costs would increase police activity and benefit citizens.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions of the study?
A: The study is conducted entirely within the Dallas Police Department, a single medium-sized urban department with approximately 3,000 officers. Dallas is described as a diverse city by race, income, and political affiliation, and the department itself is relatively diverse (29% Black, 43% Hispanic, 27% White, 15% female). The findings may not generalize to departments with different overtime rules, labor contracts, or institutional cultures.&lt;/p&gt;
&lt;p&gt;professional motivations: The non-pecuniary benefits officers derive from making arrests, such as crime reduction, duty fulfillment, or the legitimacy of their work; modeled as a professional benefit b_p that motivates arrest independent of financial compensation.&lt;/p&gt;
&lt;p&gt;private costs of arrest: The personal costs borne by officers when making an arrest, chiefly the overtime cost when an arrest extends the shift past its scheduled end, and the opportunity cost on days when a second job is scheduled. These costs are distinct from professional motivations and respond to economic incentives.&lt;/p&gt;
&lt;p&gt;arrest quality: The conditional probability that an arrest results in prosecution by the Dallas County Attorney&amp;rsquo;s office; used as a revealed-preference measure of the officer&amp;rsquo;s assessment of arrest strength. Higher arrest quality near shift end reflects greater selectivity under elevated private costs.&lt;/p&gt;
&lt;p&gt;secondary employment (second job): A formal DPD program allowing officers to register as certified police officers for private security work after their primary shift. Approximately 30% of DPD officers held a second job at some point during 2019–2021. The scheduled second job raises the opportunity cost of late-shift primary-shift arrests and provides a second source of variation in private costs.&lt;/p&gt;
&lt;p&gt;overtime cost: The cost incurred when an arrest requires an officer to remain past the end of the scheduled shift to complete paperwork and processing. Modeled as c_ot per period spent in overtime, this cost is the primary mechanism reducing late-shift arrest rates and increasing arrest selectivity.&lt;/p&gt;
&lt;p&gt;dynamic model of arrest decisions: A structural model in which officers decide each moment whether to arrest, balancing professional benefit against private cost as a function of shift timing and second-job status. Estimated by minimum distance on moments from the data; used to recover preference parameters and conduct counterfactual welfare analysis.&lt;/p&gt;</description></item><item><title>Quantifying Supply-Side Climate Policies</title><link>https://macropaperwarehouse.com/papers/quantifying-supply-side-climate-policies/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/quantifying-supply-side-climate-policies/</guid><description>&lt;p&gt;This paper asks three questions about supply-side climate policies in the oil market: how do oil companies respond to production-based taxes; what are the aggregate effects of such taxes on global CO2 emissions; and what are the distributional consequences across consumers, producers, and governments? The study addresses a gap in empirical evidence at a time when supply-side restrictions on fossil fuel production are gaining policy traction but the quantitative literature remains limited.&lt;/p&gt;
&lt;p&gt;The authors use proprietary company-level data from Rystad Energy&amp;rsquo;s UCube database covering 49,023 oil assets across 84 countries representing 98.1% of global oil production from 2000 to 2019. They identify 84 production tax reforms (54 increases, 30 decreases) with an average magnitude of roughly 5–6 percentage points. The empirical strategy is a difference-in-differences design that compares a company&amp;rsquo;s activity in a treated tax regime before and after a reform to the same company&amp;rsquo;s activity in other regimes over the same period, absorbing company-tax regime fixed effects, company-year fixed effects, and region-year fixed effects. This within-company cross-border comparison is used to test for, and rule out, activity-shifting spillovers. Two-stage least squares instruments the after-tax oil price with production taxes to isolate tax-driven price variation.&lt;/p&gt;
&lt;p&gt;The primary behavioral margin is exploration: a one-percentage-point increase in the production tax rate reduces exploration expenditure by 2.6% on average over the study period, growing to 4.1% beyond five years. The elasticity of exploration with respect to the after-tax oil price is 1.96. Reduced exploration translates into fewer discoveries; a one-percentage-point tax increase reduces discovered oil amounts by 4.3% on average and by 8.9% beyond five years. The authors find no statistically significant effect of taxes on production from existing conventional fields, consistent with high adjustment costs for already-producing wells. Unconventional production (shale, oil sands, tar sands) exhibits a statistically significant intensive-margin production response to taxes. Taxes also have no detectable effect on the extraction cost of newly discovered deposits, indicating that firms do not redirect search toward lower- or higher-cost deposits at the margin.&lt;/p&gt;
&lt;p&gt;Translating these firm-level responses into market outcomes, the authors build a dynamic field-level model spanning 2020–2100, combining field-by-field production profiles calibrated from Rystad data with demand elasticities of −0.2 and −0.5 drawn from the literature. The existing average production-weighted royalty of 21% already implies an indirect carbon price of approximately $32/tCO2 at a reference oil price of $65/barrel, an order of magnitude above the current global average demand-side carbon price of $3.1/tCO2.&lt;/p&gt;
&lt;p&gt;Under a permanent global climate royalty surcharge of 20 percentage points, annual emissions from oil fall by 5–7% in the first five years and by 9–20% in the medium term (by year 2100). The cumulative reduction over 2020–2100 is 85–161 GtCO2, or 1.0–2.0 GtCO2 per year on average. The oil price rises by $8–14/bbl initially and by $23–27/bbl by year 2100. Tax revenue to oil-producing governments increases by $590–870 billion per year; consumer surplus falls by roughly $500–730 billion per year; producer surplus falls by $270–310 billion per year. The policy breaks even in direct economic terms at a social cost of carbon of $72–84/tCO2.&lt;/p&gt;
&lt;p&gt;When the surcharge is adopted only by OECD countries (30% of current production, 49% of global exploration), short-term carbon leakage is 16–37%, rising to 58–82% by year 2100 as non-OECD producers increase exploration and development in response to the higher oil price. Net cumulative global emission reductions under the OECD-only scenario are 54–107 GtCO2 (47–73% of what the OECD reduction alone would achieve), roughly two-thirds of the global scenario outcome.&lt;/p&gt;
&lt;p&gt;Q: What is the primary behavioral margin through which oil companies respond to production taxes?
A: The primary margin is exploration expenditure. A one-percentage-point increase in the production tax rate reduces exploration by 2.6% on average across the study period, growing to 4.1% in the period six to twenty years after the reform. The after-tax oil price elasticity of exploration is 1.96, meaning a 1% increase in the after-tax price raises exploration by approximately 2%. The Poisson regression, which accounts for firms with zero exploration in a regime, yields consistent results, indicating the finding is not driven by firm entry or exit.&lt;/p&gt;
&lt;p&gt;Q: Do production taxes affect output from existing oil wells?
A: For conventional oil fields, the production response is statistically indistinguishable from zero across all specifications and time horizons, consistent with high adjustment costs making already-producing conventional wells insensitive to tax-driven price changes. Unconventional production (shale oil, oil sands, tar sands, extra heavy oil) is the exception, exhibiting a statistically significant intensive-margin production response to taxes. This asymmetry aligns with Bjørnland et al. (2021), who find that unconventional production is more price-sensitive than conventional production.&lt;/p&gt;
&lt;p&gt;Q: Do taxes affect the cost profile of newly discovered deposits?
A: No. The paper finds no statistically significant effect of production tax changes on the extraction cost of newly discovered fields, across all specifications and time horizons. This implies that, at the margin, firms do not redirect exploration toward lower-cost or higher-cost deposits in response to taxes; the volume and cost distribution of new discoveries are therefore treated as invariant to the tax regime in the quantitative model.&lt;/p&gt;
&lt;p&gt;Q: How does the paper address potential activity-shifting spillovers across countries?
A: The paper directly tests for spillovers by including both the own-regime tax rate and the company&amp;rsquo;s exploration-weighted average tax rate abroad as regressors; the foreign average tax rate has no statistically significant effect on domestic exploration. The analysis is also repeated restricting to small companies operating in two or fewer countries, where spillovers would be most pronounced; the null result on spillovers holds. Dropping these small companies from the main sample leaves the primary estimates unchanged.&lt;/p&gt;
&lt;p&gt;Q: How does the paper address the potential endogeneity of tax reforms?
A: The event study plots show no statistically significant pre-trends before reforms, supporting the parallel trends assumption. The paper also finds no significant correlation between tax reforms and observable oil-sector or macroeconomic variables in the pre-period. Subsamples minimizing lobbying concerns — private (non-national) oil companies, small companies, companies without pre-existing production in the country, and non-OPEC countries — all yield similar estimates, suggesting that large incumbents&amp;rsquo; influence over tax-setting does not drive the findings.&lt;/p&gt;
&lt;p&gt;Q: How does the paper handle the staggered difference-in-differences design?
A: To address potential bias from heterogeneous and dynamic treatment effects in a two-way fixed effects framework, the paper implements a stacked regression following Cengiz et al. (2019), constructing 18 cohort-specific datasets using never-treated countries as controls. The stacked specification yields significant effects on exploration and discoveries and null results on production and extraction costs, consistent with the main estimates. The stacked event study shows no pre-trends.&lt;/p&gt;
&lt;p&gt;Q: What is the implicit carbon price of existing production-based oil taxes?
A: At the production-weighted average royalty rate of 21% and a reference oil price of $65/bbl, the existing taxes correspond to an indirect carbon price of approximately $32/tCO2, calculated using a CO2 content of 0.43 tCO2/bbl. This figure is an order of magnitude larger than the current global average demand-side carbon price of $3.1/tCO2 (a production-weighted average including zeros for unpriced emissions). This calculation pertains only to downstream combustion emissions and excludes upstream production emissions.&lt;/p&gt;
&lt;p&gt;Q: What are the quantified effects of a global 20-percentage-point climate royalty surcharge on emissions?
A: In the first five years, the surcharge reduces annual oil-embedded emissions by 0.7–1.0 GtCO2, a 5–7% reduction. By year 2100, annual reductions reach 1.2–2.6 GtCO2, a 9–20% reduction relative to baseline. The cumulative reduction over 2020–2100 is 85–161 GtCO2 (1.0–2.0 GtCO2 per year on average), representing 17–32% of the remaining carbon budget for 1.5°C warming or 7–14% of the budget for 2°C warming. All ranges span demand elasticities of −0.2 to −0.5.&lt;/p&gt;
&lt;p&gt;Q: What happens to the global oil price under a global supply-side surcharge?
A: The immediate contraction of unconventional oil production raises the oil price by $8–14/bbl in the short term. As new exploration and field development are suppressed over time, the price effect grows, reaching $23–27/bbl by year 2100. This price increase is roughly equivalent to a global carbon price of $53–63/tCO2 levied on oil consumers in the medium term.&lt;/p&gt;
&lt;p&gt;Q: How does the paper analyze distributional incidence under the global surcharge?
A: A 20-percentage-point surcharge reduces average annual consumer surplus by $500–730 billion and producer surplus by $270–310 billion per year. Tax revenue to oil-producing governments increases by $590–870 billion per year. The net present value of the aggregate economic loss is $1,000–1,400 billion; the policy breaks even in direct welfare terms at a social cost of carbon of $72–84/tCO2. Oil-producing governments are the primary beneficiaries; both consumers and oil companies lose surplus.&lt;/p&gt;
&lt;p&gt;Q: What is the carbon leakage rate under an OECD-only supply-side coalition?
A: In the short term, leakage is 16–37%, as non-OECD unconventional producers ramp up output in response to the higher oil price. By 2050 the leakage rate rises to 41–70%. By year 2100 the coalition has reduced annual production by 9,000–9,400 million barrels while non-OECD countries have increased theirs by 5,200–7,800 million barrels, implying a terminal leakage rate of 58–82%. The net cumulative global emission reduction of 54–107 GtCO2 represents 47–73% of what the OECD reduction alone achieves, and roughly two-thirds of the global scenario.&lt;/p&gt;
&lt;p&gt;Q: Why are the authors&amp;rsquo; supply elasticity estimates somewhat larger than the prior literature?
A: The authors offer two reasons. First, their approach captures elasticity through changes in exploration activity rather than only production or field development, a broader and more forward-looking margin. Second, they use tax-driven variation in prices rather than market-price variation; the event studies show that tax reforms produce persistent changes in tax rates and after-tax prices throughout the sample, so firms are likely responding to changes perceived as durable, which would naturally elicit larger responses than responses to short-run price fluctuations.&lt;/p&gt;
&lt;p&gt;Q: What are the key limitations and scope conditions of the model?
A: The quantification omits upstream (well-to-refinery) emissions and natural gas, meaning the estimated climate effects are conservative. The demand curve is held constant over time, abstracting from long-run substitution toward clean energy. The model does not account for depletion of low-cost reserves beyond 80 years. The empirical elasticities are estimated from tax reforms that may have been perceived as temporary, meaning permanent-policy elasticities could be larger, which would imply both larger emission reductions under a global policy and higher leakage rates under a partial coalition.&lt;/p&gt;
&lt;p&gt;Q: How do distributional consequences differ between the OECD-only and global scenarios?
A: Under the OECD-only surcharge, OECD consumers and OECD producers both lose surplus, while non-OECD producers and governments everywhere gain — non-OECD governments solely through the oil price increase without bearing any tax burden. The sum of OECD producer surplus losses and non-OECD producer surplus gains is slightly negative overall. The aggregate annual global economic loss under the OECD scenario is $120–170 billion, slightly lower than the global scenario ($130–220 billion), because the oil price increase and quantity reduction are both smaller in the OECD case.&lt;/p&gt;
&lt;p&gt;Production-based tax (royalty): A tax levied on gross oil production or gross income from oil, not on profit. Unlike profit-based taxes, these are not deductible against costs and therefore create incentives to curtail exploration and production. In the paper&amp;rsquo;s framework they are equivalent to a supply-side climate instrument because they reduce the after-tax price received by producers.&lt;/p&gt;
&lt;p&gt;Climate royalty surcharge: An additional production-based tax, layered on top of existing taxes, proposed as an explicit supply-side climate policy instrument. Following Prest and Stock (2023), the paper defines this as an ad valorem levy on oil production that implicitly prices downstream CO2 emissions through its effect on the after-tax oil price.&lt;/p&gt;
&lt;p&gt;Carbon leakage: The offsetting increase in oil production by non-coalition countries in response to an oil price rise caused by a supply-restricting policy adopted by a subset of producers. Measured as the ratio of the production increase in non-coalition countries to the production reduction in coalition countries, expressed as a percentage.&lt;/p&gt;
&lt;p&gt;After-tax oil price elasticity of exploration: The percentage change in exploration expenditure per one-percent change in the after-tax oil price, estimated via 2SLS instrumenting the after-tax price with production taxes. The preferred estimate is 1.96, implying elastic exploration responses to tax-driven price changes.&lt;/p&gt;
&lt;p&gt;Extraction cost (breakeven price): The constant oil price at which the net present value of developing a field equals zero, computed using a real discount rate of 7.5%. It is the minimum price at which a field is commercially viable absent profit taxes. In the quantitative model, fields are developed if and only if extraction cost falls below the after-tax oil price.&lt;/p&gt;
&lt;p&gt;Indirect carbon price: The implicit CO2 price embedded in a production-based oil tax, calculated as the ad valorem royalty rate multiplied by the oil price and divided by the CO2 content of oil. The paper calculates that the existing average 21% royalty at $65/bbl corresponds to an indirect carbon price of approximately $32/tCO2, applicable only to downstream combustion emissions.&lt;/p&gt;
&lt;p&gt;Stacked regression (staggered DiD): A robustness approach to two-way fixed effects with staggered treatment timing, constructing cohort-specific datasets for each treatment year using only never-treated units as controls, thereby avoiding contamination from using already-treated units as comparisons for later-treated units.&lt;/p&gt;</description></item><item><title>Quota Mechanisms: Finite-Sample Optimality and Robustness</title><link>https://macropaperwarehouse.com/papers/quota-mechanisms-finite-sample-optimality-and-robustness/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/quota-mechanisms-finite-sample-optimality-and-robustness/</guid><description>&lt;p&gt;Ball and Kattwinkel study quota mechanisms — linking mechanisms that impose aggregate constraints on agents&amp;rsquo; reports across multiple decision problems — and provide the first theoretical analysis under realistic finite-sample conditions with uncertainty about the type distribution. The canonical examples are mandatory grading curves, prescription drug monitoring programs, storable votes procedures, and lifetime assistance caps (TANF). Prior literature (Jackson and Sonnenschein 2007; Matsushima et al. 2010) established only asymptotic results under the assumption that the designer knows the exact population distribution, leaving the practical rationale for quotas incomplete.&lt;/p&gt;
&lt;p&gt;The paper works in the Jackson–Sonnenschein (2007) decision framework: a principal and n agents face K independent copies of a primitive collective decision problem with independent private values and additively separable utilities. A quota mechanism requires each agent&amp;rsquo;s K reported type distributions to average to a fixed quota; in each problem copy the social choice function is applied to independently sampled types from the submitted distributions. The key methodological innovation is a reformulation of each agent&amp;rsquo;s best-response as an optimal transport problem, enabling tight bounds.&lt;/p&gt;
&lt;p&gt;The central result (Theorem 1) is a tight ex-post decision error guarantee: for any q-cyclically monotone social choice function, the (x,q)-quota mechanism has a Bayes–Nash equilibrium in which the average frequency of incorrect decisions across K problems is bounded by the sum over agents of (|Θ_i| − 1) times the total variation distance between agent i&amp;rsquo;s quota and the empirical distribution of agent i&amp;rsquo;s realized type vector. The constants (|Θ_i| − 1) are tight — they cannot be reduced even by arbitrary linking mechanisms without transfers. The core technical challenge is a &amp;ldquo;cascade of lies&amp;rdquo;: when an agent&amp;rsquo;s realized type frequencies depart from his quota, he may misreport in a way that propagates errors across types. The optimal transport reformulation shows this cascade is bounded because, under a cyclically monotone social choice function, an optimal coupling of the empirical and quota distributions can always be chosen whose support contains no nontrivial cycles, so every transport path has length at most |Θ_i| − 1.&lt;/p&gt;
&lt;p&gt;Taking expectations (Theorem 2), with quotas set equal to the prior π, the expected decision error is at most (1/√(2K)) times the sum over agents of (|Θ_i| − 1)^(3/2), which is of order 1/√K and tight to within a factor of approximately 1.25. Applied concretely: with three treatment types and K = 200 patients, the expected share receiving the wrong treatment is at most 10%.&lt;/p&gt;
&lt;p&gt;Theorem 3 establishes implementation equivalence: a social choice function is (a) one-shot implementable with transfers, (b) π-cyclically monotone, (c) asymptotically implemented by quota mechanisms, and (d) asymptotically implementable by any linking mechanism with transfers, all if and only if each other holds. No linking mechanism, even with transfers, can asymptotically implement social choice functions that quota mechanisms cannot. A quota–transfer duality is identified: the transfer T_i(θ_i&amp;rsquo;) in the one-shot problem corresponds to the Lagrange multiplier on the quota constraint for type θ_i&amp;rsquo;, with the two implementations requiring dual pieces of information about the environment.&lt;/p&gt;
&lt;p&gt;Theorem 4 bounds the error from misspecified quotas: if the true distribution is π but the quota is set to q, the mechanisms asymptotically implement some social choice function x_π whose expected distance from the target is bounded by Σ_i (|Θ_i| − 1)||q_i − π_i||. With many patients and a quota that underestimates the need for one of three treatments by 1 percentage point, at most 2% of patients receive the wrong treatment. The constants are again tight.&lt;/p&gt;
&lt;p&gt;Theorem 5 addresses robustness to agents&amp;rsquo; beliefs: in the Bergemann–Morris (2005) rich type-space framework, for any type space satisfying exchangeability and independence, the (x,π)-quota mechanism admits a belief-free equilibrium in which each agent&amp;rsquo;s strategy depends only on his own payoff type, and the expected average decision error vanishes as K → ∞. The mechanism is belief-robust because each agent knows his opponents must respect the quota, which pins down the marginal distribution of their reports regardless of their beliefs. Extensions treat interdependent values and dynamic settings with sequentially arriving information.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-fundamental-practical-problem-with-quota-mechanisms-that-the-paper-addresses"&gt;Q1. What is the fundamental practical problem with quota mechanisms that the paper addresses?&lt;/h3&gt;
&lt;p&gt;The prior literature showed quota mechanisms work asymptotically when the designer knows the true type distribution and the number of linked decisions is large. In practice, both conditions fail: any finite sample produces an empirical type distribution that deviates from the quota due to sampling variation, and quotas are typically set using imperfect estimates of the population distribution. The paper is the first to quantify the decision errors arising from these two sources of discrepancy.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-decision-error-guarantee-in-theorem-1-and-why-are-the-constants-tight"&gt;Q2. What is the decision-error guarantee in Theorem 1 and why are the constants tight?&lt;/h3&gt;
&lt;p&gt;For a q-cyclically monotone social choice function x and any realization of agents&amp;rsquo; private information, the average fraction of incorrect decisions is bounded by the sum over agents i of (|Θ_i| − 1) times ||q_i − marg(θ_i)||. The constants |Θ_i| − 1 are exactly tight: if they were reduced even slightly, the bound would fail for some realization under some linking mechanism. Tightness is demonstrated via a lower bound (Remark 3) that, in the case of a single agent with two types, agrees exactly with the upper bound.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-cascade-of-lies-and-how-does-optimal-transport-resolve-it"&gt;Q3. What is the &amp;ldquo;cascade of lies&amp;rdquo; and how does optimal transport resolve it?&lt;/h3&gt;
&lt;p&gt;When an agent&amp;rsquo;s empirical type distribution differs from his quota, truthful reporting is infeasible; he must misreport some types, which can propagate further misreporting — a cascade. The key insight is that the agent&amp;rsquo;s best-response is equivalent to choosing a coupling (joint distribution) of his empirical distribution and his quota that maximizes a linear objective. Because the social choice function is cyclically monotone, Lemma 2 establishes that an optimal coupling exists whose support contains no nontrivial cycles; consequently transport paths visit each type at most once and have length at most |Θ_i| − 1, bounding the total probability moved at (|Θ_i| − 1) times the total variation distance.&lt;/p&gt;
&lt;h3 id="q4-what-does-the-expected-error-bound-theorem-2-say-quantitatively"&gt;Q4. What does the expected error bound (Theorem 2) say quantitatively?&lt;/h3&gt;
&lt;p&gt;With the quota set equal to the prior π and K problem copies, the expected average fraction of incorrect decisions is at most (1/√(2K)) × Σ_i (|Θ_i| − 1)^(3/2). For a single agent with |Θ| = 3 types and K = 200 problems, the bound evaluates to (1/√400) × (2)^(3/2) ≈ 0.10, so at most 10% of patients receive the wrong treatment. The bound is of order 1/√K and cannot be improved by more than a factor of approximately 1.25.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-implementation-equivalence-result-theorem-3-and-why-is-it-significant"&gt;Q5. What is the implementation equivalence result (Theorem 3) and why is it significant?&lt;/h3&gt;
&lt;p&gt;Theorem 3 shows that four conditions are mutually equivalent for any social choice function x: being one-shot implementable with transfers (Rochet 1987), being π-cyclically monotone, being asymptotically implemented by (x,π)-quota mechanisms, and being asymptotically implementable by any linking mechanism including those with transfers. The significance is that no richer linking mechanism — even one with monetary transfers — can asymptotically implement anything that quota mechanisms cannot, justifying the focus on quota mechanisms.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-quotatransfer-duality-identified-in-section-52"&gt;Q6. What is the quota–transfer duality identified in Section 5.2?&lt;/h3&gt;
&lt;p&gt;In the one-shot problem, the transfer T_i(θ_i&amp;rsquo;) for agent i reporting type θ_i&amp;rsquo; corresponds exactly to the Lagrange multiplier on the quota constraint for type θ_i&amp;rsquo;. The two implementations require dual pieces of information: quota implementation requires knowledge of the type distribution π_i (to set the quota) but not the utility function or cross-agent beliefs; transfer implementation requires knowledge of agent i&amp;rsquo;s utility function and interim beliefs but not the marginal distribution π_i. A concrete allocation example illustrates that transfers can implement the social choice function without knowing the type distribution, while quotas cannot.&lt;/p&gt;
&lt;h3 id="q7-how-does-theorem-4-bound-the-error-from-a-misspecified-quota"&gt;Q7. How does Theorem 4 bound the error from a misspecified quota?&lt;/h3&gt;
&lt;p&gt;If the quota q is set based on an incorrect estimate but the true distribution is π, the (x,q)-quota mechanisms asymptotically implement some social choice function x_π whose expected total variation distance from the target x is bounded by Σ_i (|Θ_i| − 1)||q_i − π_i||. The constants |Θ_i| − 1 are again tight. Applied to opioid prescription with |Θ| = 3 and a 1 percentage point underestimate (||q − π|| = 0.01) for one treatment, the long-run expected error is at most 2 × 0.01 = 0.02, so at most 2% of patients receive the wrong treatment.&lt;/p&gt;
&lt;h3 id="q8-how-is-belief-robustness-theorem-5-formalized-and-what-does-it-require"&gt;Q8. How is belief robustness (Theorem 5) formalized and what does it require?&lt;/h3&gt;
&lt;p&gt;The paper adopts the Bergemann–Morris (2005) rich type-space framework, in which each agent has a payoff type and a belief type. Theorem 5 requires the type space to satisfy exchangeability (joint distribution over payoff types is exchangeable across problem copies) and independence (payoff types are independent across agents). Under these conditions, the (x,π)-quota mechanism has a Bayes–Nash equilibrium in which each agent&amp;rsquo;s strategy depends only on his payoff type vector, not his belief type, and the expected average decision error converges to zero as K → ∞.&lt;/p&gt;
&lt;h3 id="q9-why-is-cyclical-monotonicity-the-key-structural-condition-and-what-is-its-relationship-to-rochet-1987"&gt;Q9. Why is cyclical monotonicity the key structural condition, and what is its relationship to Rochet (1987)?&lt;/h3&gt;
&lt;p&gt;Cyclical monotonicity requires that no cycle of types would strictly gain, on average, if each type received the allocation intended for the next type in the cycle. Rochet (1987) proved that a social choice function is one-shot implementable with transfers if and only if it is cyclically monotone. Ball and Kattwinkel&amp;rsquo;s Theorem 3 adds that this same condition characterizes asymptotic implementability by quota mechanisms and by any linking mechanism with transfers, establishing a deep equivalence between the transfer-based and quota-based approaches.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-new-quota-mechanism-formulation-differ-from-jackson-and-sonnenschein-2007-and-what-are-the-consequences"&gt;Q10. How does the new quota mechanism formulation differ from Jackson and Sonnenschein (2007) and what are the consequences?&lt;/h3&gt;
&lt;p&gt;Jackson and Sonnenschein require agents to report a K-vector of types with type frequencies matching the quota, which requires quotas whose components are integer multiples of 1/K and involves additional modifications for general quotas. Ball and Kattwinkel allow each agent to report a type distribution on each problem, with the average of the K distributions constrained to equal the quota. This enables direct application of optimal transport theory; every type gets weakly higher expected utility under the Theorem 1 equilibrium than under the JS equilibrium. Under JS&amp;rsquo;s definition, Theorem 1 still holds but with an additional error term of order 1/K.&lt;/p&gt;
&lt;h3 id="q11-does-the-optimality-result-in-theorem-1-extend-to-linking-mechanisms-with-transfers"&gt;Q11. Does the optimality result in Theorem 1 extend to linking mechanisms with transfers?&lt;/h3&gt;
&lt;p&gt;Yes. Theorem 1 states that the constants |Θ_i| − 1 cannot be reduced even using arbitrary linking mechanisms — and the text specifies this holds even for mechanisms without transfers. Theorem 3 further establishes that the class of social choice functions asymptotically implementable does not expand when transfers are added, reinforcing the conclusion that quota mechanisms are not dominated by richer mechanisms in the asymptotic sense.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key concepts&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Quota mechanism: A linking mechanism in which each agent&amp;rsquo;s K reported type distributions must average to a fixed quota profile q; the social choice function is then applied to types independently sampled from each reported distribution. Generalizes mandatory grading curves, prescription quotas, and storable votes procedures.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Cyclical monotonicity (q-cyclical monotonicity): A condition on a social choice function x requiring that no cycle of types would strictly gain, on average, if each type in the cycle received the allocation intended for the next type. With multiple agents, taken in expectation over co-agents&amp;rsquo; types drawn from q. Equivalent by Rochet (1987) to one-shot implementability with transfers.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Ex-post decision error: The average, over K problem copies, of the total variation distance between the implemented decision lottery and the socially desired decision lottery, evaluated at a particular realization of private information — not in expectation over types.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Cascade of lies: The phenomenon in which an agent whose empirical type distribution departs from the quota finds it optimal to propagate misreporting across multiple types, amplifying the decision error beyond the minimum necessary to satisfy the quota constraint. Bounded in magnitude by the optimal transport analysis.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Optimal transport reformulation: Each agent&amp;rsquo;s best-response choice of report vector is recast as selecting a coupling (joint distribution) of his empirical type distribution marg(θ_i) and his quota q_i to maximize a linear objective. The acyclic structure of optimal couplings under cyclical monotonicity yields the tight error bound (|Θ_i| − 1)||q_i − marg(θ_i)||.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Implementation equivalence: The result (Theorem 3) that one-shot implementability with transfers, π-cyclical monotonicity, asymptotic implementation by quota mechanisms, and asymptotic implementability by any linking mechanism with transfers are mutually equivalent conditions on a social choice function.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Belief-free equilibrium: An equilibrium of a quota mechanism in the Bergemann–Morris type-space framework in which each agent&amp;rsquo;s strategy depends only on his payoff type, not his belief type. Exists under exchangeability and independence, because the quota pins down the marginal distribution of opponents&amp;rsquo; reports regardless of beliefs.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Distributional robustness: The property that when the quota q_i is set based on an incorrect estimate of the true distribution π_i, the long-run decision error is bounded by (|Θ_i| − 1)||q_i − π_i||, proportional to the estimation error.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>Racial Disparities in Federal Sentencing: Evidence from Drug Mandatory Minimums</title><link>https://macropaperwarehouse.com/papers/racial-disparities-in-federal-sentencing-evidence-from-drug-mandatory-minimums/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/racial-disparities-in-federal-sentencing-evidence-from-drug-mandatory-minimums/</guid><description>&lt;p&gt;This paper studies racial disparities in federal criminal sentencing by analyzing abnormal bunching in the distribution of crack-cocaine amounts recorded at sentencing. The identifying variation comes from the Fair Sentencing Act (FSA) of 2010, which raised the 10-year mandatory minimum threshold for crack-cocaine from 50 grams to 280 grams. Because the new 280g threshold was set at a point with essentially zero pre-existing bunching, the author implements a difference-in-bunching design (following Kleven 2016) that compares the pre-2010 distribution of charged drug amounts — treated as the counterfactual — to the post-2010 distribution. The primary data are case-level records from the United States Sentencing Commission (USSC) covering all federal drug cases sentenced 1999–2015, restricted to crack-cocaine offenses (approximately 50,273 cases, of which 83.3% involve black defendants, 9.2% Hispanic, and 7.6% white).&lt;/p&gt;
&lt;p&gt;The main finding is that after 2010, the fraction of cases charged with amounts in the 280–290g range increases by 3.3 percentage points overall. This increase is disproportionately concentrated among minority defendants: black and Hispanic offenders are more than 2.5 times as likely as white offenders to be charged with 280–290g after the threshold shifts to that level. Approximately 80% of the excess mass at 280g is drawn from cases that had previously been charged in the 50–280g range, indicating that prosecutors are moving cases upward to cross the new threshold rather than negotiating downward from above it. For black and Hispanic offenders specifically, cases from the 50–280g range account for 88% of the increase at the new threshold.&lt;/p&gt;
&lt;p&gt;The author rules out differential drug involvement as an explanation. The pre-2010 distributions of charged amounts from 60–280g are nearly identical across racial groups; a Kolmogorov-Smirnov test fails to reject equality (p-value = 0.792). This implies the post-2010 racial disparity in bunching is a conditional disparity — arising not from differences in underlying drug involvement but from differential treatment of similarly situated defendants.&lt;/p&gt;
&lt;p&gt;The paper then traces the bunching to prosecutorial discretion specifically. Drug seizure records (NIBRS, DEA STRIDE), survey data on drug use and selling (NSDUH), and state-level conviction records from Florida all show no change in drug quantities or behaviors at the offender or law enforcement level coinciding with the FSA. Critically, there is no bunching at 280g in drug seizure data, pointing to decisions made after arrest. By contrast, case management files from the Executive Office of the US Attorney (EOUSA) show the fraction of cases recorded in the 280–290g range increases by 7.8 percentage points after 2010. Approximately 22–30% of prosecutors (depending on the detection method) are responsible for the rise in 280g cases. Bunching patterns persist across districts and mandatory minimum thresholds for the same prosecutors, indicating it reflects a prosecutor-level characteristic.&lt;/p&gt;
&lt;p&gt;The Supreme Court&amp;rsquo;s 5-4 decision in Alleyne v. United States (June 2013) raised the evidentiary standard for facts that trigger mandatory minimums and shifted that factual determination to juries. The share of EOUSA cases recorded in the 280–290g range fell from 9.1% (2011–2013) to 6.8% (2014–2016) after Alleyne, and a difference-in-discontinuities design confirms that bunching was partially reined in by this decision.&lt;/p&gt;
&lt;p&gt;On the question of discrimination, the racial disparity in bunching cannot be explained by observable defendant characteristics — education, sex, age, criminal history, seized drug amount, or other offense elements. Approximately 70% of the disparity persists after controlling for state-by-post fixed effects and 60% after district-by-post fixed effects. The disparity can be largely explained by a state-level measure of racial animus based on Google search data (Stephens-Davidowitz 2014): prosecutors operating in higher-animus states apply more disparate treatment, a pattern consistent with taste-based rather than statistical discrimination.&lt;/p&gt;
&lt;p&gt;Cases charged just above the 280g threshold receive longer sentences than those just below it in the post-2010 period, confirming that prosecutorial bunching has real consequences for sentence length.&lt;/p&gt;
&lt;p&gt;Q: What is the central empirical strategy of the paper?
A: The paper uses a difference-in-bunching design exploiting the Fair Sentencing Act of 2010, which shifted the 10-year mandatory minimum threshold for crack-cocaine from 50g to 280g. Because the 280g point had essentially zero bunching before 2010, the pre-2010 distribution of charged drug amounts serves as an empirical counterfactual for the post-2010 distribution absent the threshold change. The design allows the author to isolate bunching caused by the new threshold and to test whether that bunching is racially disparate.&lt;/p&gt;
&lt;p&gt;Q: What is the main quantitative finding on bunching?
A: After 2010, offenders sentenced for crack-cocaine are 3.3 percentage points more likely to be charged with amounts in the 280–290g range (Column 1, Table 2). Black and Hispanic offenders are more than 2.5 times as likely as white offenders to be charged with 280–290g after the threshold change (Column 2, Table 2). This racial gap is the central disparity the paper investigates.&lt;/p&gt;
&lt;p&gt;Q: Does the racial disparity in bunching reflect genuine differences in drug involvement?
A: No. The pre-2010 distributions of charged amounts from 60–280g are nearly identical across racial groups; a Kolmogorov-Smirnov test fails to reject equality with a p-value of 0.792. Because these pre-period distributions are taken as reflecting true drug involvement, their similarity by race implies the post-2010 disparity is a conditional racial disparity — arising from differential treatment of similarly situated defendants, not from differential drug involvement.&lt;/p&gt;
&lt;p&gt;Q: Where in the criminal justice process does the bunching originate?
A: The bunching originates in prosecutorial decisions, not at the arrest or law enforcement stage. Drug seizure records (NIBRS and DEA STRIDE) show no bunching at 280g, and survey data (NSDUH) show no post-FSA change in drug use or selling by minority defendants. Florida state-level records show no shift in the share of high drug-weight cases. By contrast, EOUSA case management files — which capture quantities recorded by prosecutors — show an increase of 7.8 percentage points in the fraction of cases in the 280–290g range after 2010.&lt;/p&gt;
&lt;p&gt;Q: What fraction of prosecutors engage in this bunching behavior?
A: Approximately 29.7% of prosecutors have a higher-than-normal percentage of cases at 280–290g after 2010 under a straightforward outlier criterion. Using the outlier detection procedure from Ridgeway and MacDonald (2009), approximately 22% are flagged as outliers. A Bayesian shrinkage method estimates approximately 30% (SE = 0.042) of prosecutors engage in this bunching. The behavior persists across districts and across multiple mandatory minimum thresholds for the same prosecutors, indicating it is a durable prosecutor-level characteristic.&lt;/p&gt;
&lt;p&gt;Q: What evidence links the bunching to upward manipulation rather than downward negotiation?
A: Approximately 80% of the excess mass at 280g is drawn from cases previously charged in the 50–280g range rather than from cases above 290g. For black and Hispanic offenders the share is 88%. This pattern indicates prosecutors are pushing amounts upward past the new threshold to secure longer sentences, not negotiating amounts downward from above the threshold — reversing the direction assumed in prior qualitative discussions.&lt;/p&gt;
&lt;p&gt;Q: What was the effect of Alleyne v. United States on bunching?
A: The Supreme Court&amp;rsquo;s 5-4 decision in Alleyne (June 2013) raised the evidentiary standard for facts triggering mandatory minimums and assigned those factual determinations to juries rather than judges. The share of EOUSA cases in the 280–290g range fell from 9.1% in 2011–2013 to 6.8% in 2014–2016. A difference-in-discontinuities design confirms that bunching expanded in the run-up to Alleyne and was partially curtailed afterward, providing additional evidence that the bunching reflects prosecutorial manipulation rather than genuine drug amounts.&lt;/p&gt;
&lt;p&gt;Q: Can observable defendant characteristics explain the racial disparity in bunching?
A: No. The racial disparity in bunching persists after controlling for education, sex, age, criminal history, seized drug amount, and other offense elements. Approximately 70% of the disparity remains after controlling for state-by-post fixed effects and 60% after controlling for district-by-post fixed effects. The disparity exists among observably similar defendants, ruling out the hypothesis that it is driven by correlated case characteristics.&lt;/p&gt;
&lt;p&gt;Q: What evidence distinguishes taste-based from statistical discrimination?
A: The racial disparity in bunching is largely explained by a state-level measure of racial animus constructed from Google search data (Stephens-Davidowitz 2014): prosecutors in higher-animus states apply more racially disparate treatment. Because statistical discrimination would predict disparate outcomes based on informative case characteristics rather than on the ambient racial attitudes of the jurisdiction, the correlation with racial animus is more consistent with taste-based discrimination than with statistical discrimination.&lt;/p&gt;
&lt;p&gt;Q: Does bunching at 280g have real consequences for sentence length?
A: Yes. Cases charged just above the 280g threshold receive longer sentences than those charged just below it in the post-2010 period, confirming that the mandatory minimum threshold is binding and that prosecutorial bunching translates into materially longer sentences for the affected defendants.&lt;/p&gt;
&lt;p&gt;Q: How does this paper contribute relative to Rehavi and Starr (2014)?
A: Rehavi and Starr (2014) linked arrest to sentencing records to show black offenders receive harsher sentences, driven by prosecutorial charging of mandatory minimums, but acknowledged that unobserved differences in criminal conduct within offense codes remained a concern. This paper addresses that concern by using the pre-2010 distribution of charged amounts as a counterfactual for drug involvement, documenting near-identical pre-period distributions by race, and tracing the post-FSA disparity through multiple data sources to isolate prosecutorial decisions specifically. The paper also quantifies the fraction of prosecutors involved and tests discrimination mechanisms.&lt;/p&gt;
&lt;p&gt;Q: What is the relationship between this paper&amp;rsquo;s findings and the policy goals of the Fair Sentencing Act?
A: The FSA achieved its stated goal of narrowing racial gaps attributable to the crack-powder disparity in mandatory minimum thresholds, and in line with prior work the author confirms a net decline in sentences after 2010. However, the increase in bunching at 280g by prosecutors — disproportionately applied to black and Hispanic defendants — dampened the FSA&amp;rsquo;s effectiveness. The paper thus documents a strategic response by a subset of prosecutors that partially offset the reform&amp;rsquo;s intended benefits for minority defendants.&lt;/p&gt;
&lt;p&gt;Q: How robust are the main bunching estimates?
A: The 3.3 percentage point overall increase and the 2.5x racial disparity are robust to various sample restrictions, inclusion of state fixed effects, time trends, state-specific time trends, offender-level controls, Logit/Probit/Poisson models, wider bunching range definitions (e.g., 280–380g), inclusion of cases with weights coded as a range, and alternative standard error calculations. Including range-coded cases actually exacerbates the estimated degree of bunching and the racial disparity.&lt;/p&gt;
&lt;p&gt;Bunching (in this paper&amp;rsquo;s sense): An excess mass of cases charged with a drug amount at or just above the mandatory minimum threshold, defined operationally as a disproportionate concentration of cases in the 280–290g range relative to the counterfactual distribution. Bunching reflects discretionary upward adjustment of charged amounts by prosecutors to trigger longer mandatory minimum sentences rather than true drug seizure quantities.&lt;/p&gt;
&lt;p&gt;Difference-in-bunching design: An empirical strategy adapted from Kleven (2016) that compares the actual post-2010 distribution of charged drug amounts to the pre-2010 distribution as a counterfactual for what the post-2010 distribution would have looked like absent the FSA threshold change. The method exploits the fact that the 280g threshold was a point of essentially zero bunching before 2010.&lt;/p&gt;
&lt;p&gt;Conditional racial disparity in bunching: A racial gap in the probability of being charged at 280–290g that remains after conditioning on similar underlying drug involvement, operationalized by the near-identical pre-2010 distributions of charged amounts from 60–280g across racial groups. The conditional disparity isolates differential treatment from differential conduct.&lt;/p&gt;
&lt;p&gt;Prosecutorial discretion (in this context): The legal authority of federal prosecutors to determine the drug quantity attributed to a defendant for sentencing purposes, which is not strictly bound to the amount physically seized at arrest. Prosecutors can rely on informant testimony, conspiracy attribution, or approximations to establish amounts above what was seized, giving them effective control over whether the mandatory minimum threshold is crossed.&lt;/p&gt;
&lt;p&gt;Taste-based discrimination: Racially disparate prosecutorial behavior that cannot be explained by observable case characteristics or informative statistical inference about defendant conduct, and that correlates instead with ambient state-level racial animus. In this paper&amp;rsquo;s framing, taste-based discrimination is distinguished from statistical discrimination by its correlation with the Stephens-Davidowitz racial animus measure rather than with defendant or offense characteristics.&lt;/p&gt;
&lt;p&gt;Mandatory minimum threshold (in federal crack-cocaine sentencing): A drug quantity cutoff — set at 50g before 2010 and 280g after the FSA — above which federal law mandates a sentence of at least 10 years unless specific departure conditions are met. The threshold creates a sharp discontinuity in expected sentence length that gives prosecutors an incentive to place cases just above it.&lt;/p&gt;
&lt;p&gt;State-level racial animus measure: A proxy for the prevalence of racially prejudiced attitudes in a state, constructed by Stephens-Davidowitz (2014) from Google Trends search volume data (2004–2007) for a specific racial slur and its plural, normalized by total search volume. Used here as a predictor of the size of the racial disparity in prosecutorial bunching across states.&lt;/p&gt;</description></item><item><title>Rationing by Race</title><link>https://macropaperwarehouse.com/papers/rationing-by-race/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/rationing-by-race/</guid><description>&lt;p&gt;Singh and Venkataramani ask whether resource scarcity causes discriminatory rationing of health care by patient race, with patient death as the starkest possible outcome of biased allocation decisions. They examine 107,221 inpatient admissions from 2015 to 2018 at two large urban academic teaching hospitals (each with over 500 beds) in a Southeastern U.S. city with a sizable Black population. Black patients accounted for 60% of admissions, were on average younger (52 vs. 59 years), more likely to be female (65% vs. 50%), and had similar comorbidity burdens and baseline in-hospital death rates (approximately 2% for both groups), but waited over two hours longer on average for an inpatient bed and were 27% less likely to be admitted to the ICU.&lt;/p&gt;
&lt;p&gt;The authors exploit quasi-exogenous hour-to-hour variation in hospital capacity strain — measured as the share of inpatient beds occupied at the hour of a patient&amp;rsquo;s arrival — which clinical and qualitative literature establishes is difficult to predict even day-to-day. Capacity strain is coded in hospital-specific deciles (beds filled ranged from 69–78% in decile 1 to 91–95% in decile 10). The core regression interacts patient race with strain decile, controlling for hospital-specific hour-of-day, day-of-week, month-of-year, and year fixed effects; physician-of-record fixed effects; and a rich vector of patient characteristics including Elixhauser comorbidity indices, insurance status, and vital signs. Identification rests on the assumption that strain at the hour of arrival is conditionally independent of unobserved patient characteristics correlated with race — an assumption validated through balance tests on demographics, comorbidities, vital signs, machine-learning-derived admission themes, and selective discharge patterns.&lt;/p&gt;
&lt;p&gt;The main finding is that in-hospital mortality rises for Black patients but not for White patients as hospitals approach capacity. At the tenth decile of strain, Black patients face a mortality rate 0.7 percentage points higher than White patients — a 47.6% relative increase over the 1.47% White mortality rate at the same decile. A pooled difference-in-differences estimate implies that approximately 15% of Black patient deaths at high strain (decile 10) would not have occurred had Black patients faced the same strain-mortality relationship as White patients (coefficient 0.0052, p = 0.025). This pattern is concentrated among patients with the greatest ex ante medical need as measured by above-median Elixhauser mortality index scores (a score with AUC of 0.92 for predicting in-hospital mortality) and, in qualitatively similar but less precisely estimated form, by abnormal vital signs at arrival.&lt;/p&gt;
&lt;p&gt;The authors identify wait time for an inpatient bed as the primary mechanism. At all levels of capacity strain, high-need Black patients wait longer than low-need White patients — a pattern the authors characterize as a striking inversion of any need-based allocation principle. Racial disparities in wait times widen further at the highest decile of strain, exactly mirroring the mortality pattern. As an additional, more suggestive mechanism, the authors analyze free-text clinical documentation (the Reason for Admission field) using descriptive text features (time to completion, character count, average word length), sentiment analysis (subjectivity and polarity scores via TextBlob), and adjective counts. Documentation for Black patients exhibits features consistent with lower provider effort at all strain levels — shorter notes, less time deferred to completion — and subjectivity of notes and adjective counts diverge further by race at the highest strain decile, with White patients receiving increasingly detailed and descriptive notes as strain rises.&lt;/p&gt;
&lt;p&gt;The findings are robust across sparse models (age, gender, hospital fixed effects only) through fully saturated specifications (DRG fixed effects, interactions of all controls with race and strain), and to replacing Elixhauser index composites with their 31 individual comorbidity components. The authors explicitly scope their findings to a pre-COVID-19 period (2015–2018), while noting that pandemic-era record capacity strain and racial disparities in health outcomes suggest de facto race-based rationing may have been far more severe during COVID-19.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question and why is the health care setting chosen?
A: The paper asks whether increasing resource scarcity causes discriminatory rationing on the basis of race in consequential, high-stakes real-world decisions. Health care is chosen because it is high-stakes (patient death is the outcome), has a long documented history of racial discrimination at both provider and system levels, and offers uniquely detailed time-stamped electronic health record data that enables identification from hour-to-hour variation in capacity strain — a finer temporal resolution than most prior work.&lt;/p&gt;
&lt;p&gt;Q: How is hospital capacity strain measured and what is the identifying variation?
A: Strain is measured as the total number of patients occupying inpatient beds at the specific hour of a patient&amp;rsquo;s arrival, converted into hospital-specific deciles. The first decile corresponds to 69–78% of beds filled and the tenth decile to 91–95%. The identifying variation is residual hour-to-hour fluctuation in this measure after removing hospital-specific hour-of-day, day-of-week, month-of-year, and year fixed effects, which absorbs all predictable capacity patterns. Clinical and qualitative evidence establishes that even day-to-day strain is difficult to anticipate, making hour-to-hour residual variation plausibly as-if random.&lt;/p&gt;
&lt;p&gt;Q: What are the main mortality findings, and how large are the racial disparities at peak strain?
A: At the tenth decile of capacity strain, Black patients face a mortality rate 0.7 percentage points higher than White patients, representing a 47.6% relative increase over the 1.47% White mortality rate at that decile. The pooled difference-in-differences estimate (comparing decile 10 to deciles 1–9) implies that approximately 15% of Black patient deaths at high strain would not have occurred if Black patients had the same strain-mortality relationship as White patients (coefficient 0.0052, p = 0.025). White patient mortality does not increase at high strain; if anything, small (imprecisely estimated) decreases appear at deciles 7–9.&lt;/p&gt;
&lt;p&gt;Q: Which patients drive the racial mortality disparity?
A: The disparity is concentrated among patients with above-median Elixhauser mortality index scores — the ex ante sickest patients. The Elixhauser Mortality Index has a predictive AUC of 0.92 for in-hospital mortality. At decile 10, high-need Black patients experience a sharp increase in mortality not seen for high-need White patients or for low-need Black patients. Qualitatively similar but less precisely estimated results appear when acute need is measured by abnormal vital signs at arrival, with the difference that the triple interaction (race × strain × high-need vitals) is not statistically significant, consistent with vital signs being noisier proxies for severity than the Elixhauser indices.&lt;/p&gt;
&lt;p&gt;Q: How do the authors validate the identifying assumption that strain is conditionally independent of patient composition by race?
A: They document five types of supporting evidence: (i) the distribution of Black and White patients across hours of arrival and across strain deciles is nearly identical; (ii) regressions of patient demographics, all five Elixhauser comorbidity measures, and five vital signs abnormalities on race × strain interactions show no significant differential selection by race at different strain levels; (iii) machine-learning (Latent Dirichlet Allocation) topic themes from free-text admission notes change similarly by strain for Black and White patients; (iv) there is no evidence of selective discharge to hospice care by race and strain, with point estimates running counter to the hypothesis; and (v) strain is computed at time of arrival to the hospital rather than time of admission to an inpatient bed, preserving exogeneity.&lt;/p&gt;
&lt;p&gt;Q: What is the primary identified mechanism for the mortality finding?
A: Wait time for an inpatient bed is the primary mechanism. Black patients experience greater increases in wait times as strain rises compared to White patients, with the clearest divergence at decile 10 — exactly mirroring the mortality pattern. More strikingly, at every decile of strain (including decile 1, when beds are most abundant), high-need Black patients wait longer for a bed than low-need White patients, implying that the disparity is not solely a product of logistical constraints but reflects ingrained factors in clinical protocols, likely including implicit or explicit provider bias.&lt;/p&gt;
&lt;p&gt;Q: What does the wait time evidence reveal about the role of medical need vs. race in allocation decisions?
A: At lower strain levels, low-need patients appropriately wait longer than high-need patients. However, at higher strain levels (deciles 8–10) this need-based gap almost entirely disappears, while the racial gap in wait times persists. The gap between high-need Black and low-need White patients is larger than the gap between high-need and low-need patients of the same race, meaning race is a stronger predictor of wait times than medical need. This pattern is consistent with the paper&amp;rsquo;s conceptual framework in which increasing strain reduces providers&amp;rsquo; ability to accurately assess medical need while increasing the weight assigned to racial identity.&lt;/p&gt;
&lt;p&gt;Q: How is provider effort measured and what are the findings?
A: Provider effort is inferred from features of free-text Reason for Admission documentation: time to completion, character count, average word length, TextBlob subjectivity and polarity scores, and adjective counts. Across all strain levels, Black patients&amp;rsquo; documentation exhibits features consistent with lower effort — shorter completion times (providers less likely to defer documentation for clinical tasks), shorter notes with fewer characters and shorter words. At the highest strain decile, subjectivity scores for Black patients&amp;rsquo; notes increase relative to White patients&amp;rsquo; (driven by both rising Black and falling White subjectivity), and White patients receive more adjectives as strain rises while Black patients&amp;rsquo; adjective counts do not increase. Polarity scores remain stable by race and strain.&lt;/p&gt;
&lt;p&gt;Q: What do the documentation patterns suggest about compensatory behavior by providers?
A: The authors speculate that providers may anticipate reduced care quality at high strain and compensate by becoming more conscientious with White patients — writing longer, more detailed, more descriptive notes as strain increases, and potentially exerting greater care effort correlated with these documentation improvements. This protective compensatory behavior appears substantially less pronounced or absent for Black patients, which the authors suggest may translate into the small imprecisely estimated decrease in White patient mortality at higher strain deciles. They explicitly characterize this interpretation as speculative and requiring further investigation.&lt;/p&gt;
&lt;p&gt;Q: How robust are the main mortality findings to specification choices?
A: The mortality findings hold across: (i) sparse models with only age, gender, and hospital/year fixed effects; (ii) linear probability and logistic models; (iii) models with DRG fixed effects to compare within-diagnosis; (iv) models interacting all control variables with patient race and strain; (v) models replacing the Elixhauser composite index with its 31 individual comorbidity components; and (vi) models additionally controlling for five individual abnormal vital sign indicators. Results are substantively unchanged across all these specifications.&lt;/p&gt;
&lt;p&gt;Q: What additional care intensity measures are examined and what do they show?
A: The authors also examine ICU admission, ICU length of stay, total inpatient length of stay, and inpatient charges. They find no strain-related racial disparities on these margins. However, they note that unconditionally (across all strain levels), Black patients receive fewer resources on average — they are 27% less likely to be admitted to the ICU. The authors treat these care intensity measures as harder to interpret because both over- and under-provision can harm patients, and thus view them as less informative for their research question.&lt;/p&gt;
&lt;p&gt;Q: What conceptual framework guides the empirical predictions?
A: The framework models providers as assessing perceived medical need N&lt;em&gt;ij(t) = Ni × exp(−γ × S(t)), where the parameter γ captures the diminishing ability to accurately assess true need as strain S(t) rises. Simultaneously, the racial weight R&lt;/em&gt;ij(t) = Ri × φ(S(t)) increases with strain through the parameter φ(S(t)). When γ = 0 and φ = 0, allocation is race-neutral and need-based. When both parameters are positive, increasing strain simultaneously degrades need assessment and amplifies reliance on racial identity in allocation decisions — the paper&amp;rsquo;s core prediction, which is confirmed empirically.&lt;/p&gt;
&lt;p&gt;Q: How do the findings relate to the COVID-19 pandemic?
A: The data predate COVID-19 (2015–2018). The authors argue that pandemic conditions — record hospital capacity strain (especially in hospitals serving Black patients), extreme provider burnout, and documented racial disparities in health access — suggest race-based rationing may have been considerably more severe during COVID-19. The paper also contextualizes its findings within the pandemic-era debate over whether explicit race-based triage protocols were ethical or legal, arguing that de facto rationing by race appears to occur in ordinary care settings under typical stressors irrespective of that normative debate.&lt;/p&gt;
&lt;p&gt;Q: What policy interventions do the authors suggest?
A: The authors propose: increasing provider awareness of implicit biases; developing new algorithms to improve triage decisions for high-mortality-risk patients who might otherwise be overlooked; correcting existing care algorithms with documented racial bias; building provider peer networks to reduce biased treatment decisions; supporting patient self-advocacy; improving capacity prediction systems (as spurred by COVID-19); and creating load-shifting protocols and inter-hospital transfer networks to prevent resources from being stretched beyond capacity during high-strain periods.&lt;/p&gt;
&lt;p&gt;Capacity strain: The state of a hospital when a high share of inpatient beds are occupied, measured here at the hour of patient arrival as hospital-specific deciles of bed occupancy (ranging from 69–78% full at decile 1 to 91–95% full at decile 10); the paper&amp;rsquo;s primary measure of resource scarcity.&lt;/p&gt;
&lt;p&gt;Rationing by race: The paper&amp;rsquo;s term for the phenomenon whereby, as resource scarcity deepens, allocation decisions increasingly reflect patient racial identity rather than medical need — a form of discriminatory rationing that the authors distinguish from explicit (de jure) race-based triage and document as de facto practice.&lt;/p&gt;
&lt;p&gt;Perceived need (N*): In the paper&amp;rsquo;s conceptual framework, the provider&amp;rsquo;s assessment of a patient&amp;rsquo;s medical need, which deviates from true need Ni by the factor exp(−γ × S(t)) as strain S(t) increases; captures the provider team&amp;rsquo;s diminishing ability or willingness to accurately assess true medical need under cognitive and resource constraints.&lt;/p&gt;
&lt;p&gt;Racial weight (R*): The weight assigned to a patient&amp;rsquo;s racial identity in allocation decisions, modeled as Ri × φ(S(t)), where the function φ is increasing in capacity strain; represents the potential for discrimination — from implicit bias, algorithmic bias, reduced patient advocacy, or provider-patient social distance — to intensify as strain rises.&lt;/p&gt;
&lt;p&gt;Wait time inversion: The condition, documented throughout the paper, where high-need Black patients wait longer for an inpatient bed than low-need White patients at every decile of capacity strain, including decile 1 when resources are most abundant — inverting the normative principle that greater medical need should yield faster access to care.&lt;/p&gt;
&lt;p&gt;Elixhauser Mortality Index: A widely validated composite score of patient comorbid conditions used to predict in-hospital mortality (AUC = 0.92); used in this paper as the primary measure of chronic medical need, with patients split at the median into relatively sick (above median) and relatively healthy (below median) groups.&lt;/p&gt;
&lt;p&gt;Provider effort (inferred): An unobserved construct inferred in this paper from features of free-text clinical documentation in the Reason for Admission field, including time to note completion, character count, average word length, TextBlob subjectivity and polarity scores, and adjective counts; features argued to reflect how much attention, detail, and care a provider invested in documenting — and by extension, in assessing — a patient&amp;rsquo;s condition.&lt;/p&gt;</description></item><item><title>Regulatory Competition in the US Life Insurance Industry</title><link>https://macropaperwarehouse.com/papers/regulatory-competition-in-the-us-life-insurance-industry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/regulatory-competition-in-the-us-life-insurance-industry/</guid><description>&lt;p&gt;This paper quantitatively assesses the consequences of jurisdictional competition in the US life insurance industry, an $8 trillion market. The central question is whether competition between state regulators over capital requirements for captive reinsurance subsidiaries — a form of regulatory competition — increases or decreases total surplus, and by how much.&lt;/p&gt;
&lt;p&gt;US life insurers are regulated at the state level. Since the early 2000s, states have competed to attract captive reinsurance subsidiaries (captives) by setting lower capital requirements on these entities. The externality structure is asymmetric: the captive state earns tax revenues on liabilities transferred to captives and sets their capital requirements, but bears default costs only for policyholders in its own state. Consumer states bear default costs for their own residents even when those policies have been transferred to an out-of-state captive. This mismatch between who sets capital requirements and who bears default costs creates the externality that drives the race-to-the-bottom dynamic studied in the paper.&lt;/p&gt;
&lt;p&gt;The empirical setting draws on a novel dataset covering 66 US life insurers from 2005 to 2020, with total liabilities of $1.9 trillion (approximately 25% of the sector). Data sources include NAIC filings via S&amp;amp;P, CompuLife pricing data, A.M. Best ratings, SEC filings, and state legislative records. The author assembles novel data on captives&amp;rsquo; capital levels from SEC filings, Iowa Insurance Department captive financial statements, and insurer reinsurance exhibits.&lt;/p&gt;
&lt;p&gt;Three motivating empirical findings ground the structural model. First, captives materially reduce insurers&amp;rsquo; capital: in 2019, risk-based capital ratios are 23% lower on average after accounting for captives, with the median insurer&amp;rsquo;s capital declining 24%, and this translates into an increase in 10-year default probability from 1.0% to 2.9%. Second, states&amp;rsquo; capital requirements are the primary determinant of where insurers locate captives: a 1 percentage point increase in a state&amp;rsquo;s captive capital rate is associated with a 1.6 percentage point decrease in the probability an insurer chooses that state (against a 1.1 percentage point unconditional probability), and this holds when insurers switch states over time as capital requirements change. Tax rates, geographic proximity, and amenities are not meaningfully correlated with captive location choice. Third, a difference-in-differences design exploiting Regulation XXX (effective January 1, 2000), which raised capital requirements differentially across product term lengths, shows that 30-year term products — which faced the largest capital requirement increases — experienced price increases averaging 10.3% relative to 10-year term products, with quantities declining monotonically for longer-term products, consistent with an inward supply shift.&lt;/p&gt;
&lt;p&gt;The paper develops a structural model of the insurance market with imperfectly competitive insurers, endogenous default following a Leland (1994) framework, discrete choice consumer demand (Berry, 1994), and state regulators who set captive capital rates to maximize a weighted objective over tax revenues, default costs, consumer surplus, and producer surplus. Regulators deviate from a utilitarian social planner in two ways: they are state-based (generating competition and default externalities) and face agency frictions (captured by welfare weights that differ from unity). The demand side implies an average price elasticity of 2.4. The regulator side reveals that state regulators are willing to trade $1 of default costs against $3.5 of tax revenues and $0.59 of consumer surplus — both diverging from the social planner&amp;rsquo;s equal weighting.&lt;/p&gt;
&lt;p&gt;The main counterfactual finding is that eliminating competition by federalizing insurance regulation would cause regulators to raise capital requirements by 19% (3 percentage points), reducing expected default costs by $2.4 billion while lowering consumer surplus by $880 million, for a net total surplus gain of $1.5 billion. Regulator utility would increase by $3.3 billion in equivalent tax revenues. Because regulators over-value consumer surplus relative to default costs, competition exacerbates rather than counteracts their agency frictions, making competition unambiguously welfare-reducing in the baseline. A social planner would set capital requirements even higher than a federal regulator. On distribution, large states such as California and New York gain most from federalization (they bear substantial default costs), while Vermont — the largest captive state by market share — loses because it would forfeit captive tax revenues. Unilateral bans are found to have limited equilibrium consequences: a New York ban on captive use by insurers selling in New York would achieve only 23% of the national default cost reduction that federalization achieves, and a ban on captives domiciled in Vermont would achieve only 10%, as insurers would redirect captives to other states.&lt;/p&gt;
&lt;p&gt;Q: What is a captive reinsurance subsidiary and why do states compete to attract them?
A: A captive is a wholly-owned subsidiary of a life insurance holding company that reinsures policies written by the operating company, moving liabilities off the operating company&amp;rsquo;s balance sheet. Captive states earn tax revenues on liabilities transferred to captives and can set their own capital requirements on those entities, which are lower than the uniform NAIC standards applied to operating companies. Because captives are taxed by the state where they are domiciled — not the consumer&amp;rsquo;s state — captive states can earn tax revenues on policies sold elsewhere, incentivizing competition through lower capital requirements to attract insurers.&lt;/p&gt;
&lt;p&gt;Q: What is the default externality at the core of this paper&amp;rsquo;s argument?
A: When an insurer defaults, the shortfall on policies sold to consumers in a given state is borne by that state&amp;rsquo;s guaranty fund and consumers, regardless of where the captive holding those liabilities is domiciled. So Vermont, as the captive state, sets the capital requirement on liabilities transferred from (for example) Massachusetts policyholders, but does not bear the default cost on those Massachusetts policies. This means Vermont internalizes only the default cost on its own consumers, leading it to set capital requirements lower than it would if it bore the full default cost — a classic externality.&lt;/p&gt;
&lt;p&gt;Q: How large is the effect of captives on insurers&amp;rsquo; capital levels?
A: Using novel data on captives&amp;rsquo; actual balance sheets, the author finds that in 2019, the size-weighted average risk-based capital ratio of sample insurers is 23% lower after consolidating captives into the operating company&amp;rsquo;s balance sheet. The median insurer&amp;rsquo;s capital ratio decreases by 24%. In terms of default risk, this adjustment corresponds to an increase in the 10-year default probability from 1.0% to 2.9% based on historical insurer default rates.&lt;/p&gt;
&lt;p&gt;Q: What is the state of competition among captive domiciles in the data?
A: Twenty-two states had passed laws allowing captives as of the sample period, with the set of competing states largely stabilizing after 2013. The market is moderately concentrated: the top five states (Vermont, Arizona, Delaware, Iowa, and South Carolina) account for 80% of all captive liabilities, and the Herfindahl-Hirschman Index is 0.20. Vermont has maintained its position as the largest captive state throughout the period.&lt;/p&gt;
&lt;p&gt;Q: What evidence shows that capital requirements — rather than taxes or other factors — drive captive location choice?
A: In a linear probability model of captive location with insurer-year fixed effects, a 1 percentage point increase in a state&amp;rsquo;s captive capital rate is associated with a 1.6 percentage point decrease in the probability that an insurer chooses that state (versus a 1.1 percentage point unconditional probability). Captive tax rates are not meaningfully correlated with location choice, consistent with federal tax laws prohibiting the use of reinsurance to reduce tax liabilities. A changes-on-changes specification confirms that insurers are more likely to shift their captives to states that lower their capital requirements over time.&lt;/p&gt;
&lt;p&gt;Q: How does the Regulation XXX natural experiment identify the supply-side effect of capital requirements on insurance prices?
A: Regulation XXX, effective January 1, 2000, increased reserve requirements for operating companies on a mechanical basis tied to policy term length, with longer-term products facing larger increases. Using a difference-in-differences design at the insurer-product-month level with insurer-product and month fixed effects, the paper finds that products with larger capital requirement increases experienced larger price increases immediately after the regulation took effect. Thirty-year term products experienced price increases averaging 10.3% relative to 10-year term products (the reference group) within three months. Quantities also declined monotonically for longer-term products, confirming an inward shift of the supply curve rather than a demand shift.&lt;/p&gt;
&lt;p&gt;Q: What are the estimated regulator welfare weights, and what do they imply about agency frictions?
A: Normalizing the weight on tax revenues to 1, the paper recovers that regulators value $1 of default costs as worth $0.29 (implying $3.5 of tax revenues trades off against $1 of default costs) and value consumer surplus at $0.59 per dollar. Because the social planner sets all weights equal to 1, these estimates show regulators over-weight tax revenues and consumer surplus relative to default costs. The higher weight on consumer surplus is consistent with political backlash from consumers facing high insurance prices.&lt;/p&gt;
&lt;p&gt;Q: What is the total surplus effect of eliminating regulatory competition through federalization?
A: Federalizing insurance regulation — modeled as a single federal regulator setting a uniform capital rate while holding fixed regulatory frictions — would lead regulators to raise captive capital requirements by 19% (3 percentage points) to internalize the default externality. Expected default costs would fall by $2.4 billion. However, higher capital requirements would raise insurance prices and reduce consumer surplus by $880 million. The net effect is a total surplus increase of $1.5 billion. Regulator utility (in equivalent tax revenues) would increase by $3.3 billion.&lt;/p&gt;
&lt;p&gt;Q: Would eliminating both competition and regulatory frictions (i.e., a social planner) produce a different outcome than just federalizing?
A: In the baseline estimates, a social planner would set capital requirements even higher than a federal regulator, because regulators&amp;rsquo; agency frictions lead them to under-weight default costs relative to consumer surplus, pushing capital requirements below the socially optimal level even absent competition. Competition further exacerbates these frictions by providing an additional incentive to lower capital rates. Thus, in the baseline, competition unambiguously decreases total surplus. The paper also reports results under alternative assumptions, providing a &amp;ldquo;menu&amp;rdquo; for policymakers that maps different assumptions about regulators&amp;rsquo; frictions to quantitative welfare statements.&lt;/p&gt;
&lt;p&gt;Q: What distributional consequences across states explain why federalization has not been adopted?
A: Federalization would benefit large states such as California and New York most, because those states bear substantial default costs on large volumes of policies sold to their consumers. States with large captive market shares, primarily Vermont, would be made worse off because they would lose captive tax revenues. These predicted gains and losses align with actual state policy positions: New York has called for a national ban on captives, California forbids insurers from setting up captives there, and Vermont has been the most aggressive state in attracting captive domiciles.&lt;/p&gt;
&lt;p&gt;Q: How effective are unilateral state bans as an alternative to federal coordination?
A: The paper estimates that a unilateral ban by New York on insurers selling in New York from using captives would achieve only 23% of the national default cost reduction that full federalization would achieve. A unilateral ban on captives domiciled in Vermont — the largest captive state — would achieve only 10% of federalization&amp;rsquo;s default cost reduction, because insurers would simply relocate their captives to other states that still allow them. This finding underscores the importance of cross-state coordination for meaningful regulatory reform.&lt;/p&gt;
&lt;p&gt;Q: What does the model&amp;rsquo;s demand estimation imply about consumer sensitivity to insurance prices?
A: The discrete choice demand model estimated on state-level sales, prices, and product characteristics implies an average price elasticity of demand of 2.4 for life insurance products. This elasticity disciplines the quantitative impact of capital requirements on product markets through their effect on insurance prices.&lt;/p&gt;
&lt;p&gt;Q: How does the paper recover regulators&amp;rsquo; objective functions?
A: The author uses the revealed preferences of state regulators, exploiting regulators&amp;rsquo; utility maximization first-order conditions and performing numerical perturbations around those conditions to calibrate the welfare weights (lambdas) on each component of the regulators&amp;rsquo; utility function. This approach recovers regulators&amp;rsquo; tradeoff weights from their observed policy choices — specifically their captive capital rate decisions — without directly observing regulators&amp;rsquo; preferences.&lt;/p&gt;
&lt;p&gt;Captive reinsurance subsidiary: A wholly-owned subsidiary of a life insurance holding company that reinsures liabilities from the operating company. Unlike operating companies, captives are regulated by the state in which they are domiciled (the captive state) under that state&amp;rsquo;s own capital requirements, which are typically lower than the uniform NAIC standards. Captives allow insurers to reduce their overall capital requirements by allocating liabilities to the captive.&lt;/p&gt;
&lt;p&gt;Default externality: The mismatch between who sets capital requirements for captives (the captive state) and who bears default costs when an insurer fails (the consumer&amp;rsquo;s state and its guaranty fund). Because the captive state bears default costs only for its own residents — not for residents of states where the insurer also sells — it has an incentive to set lower capital requirements than it would if it internalized the full default cost, leading to an externality on other states.&lt;/p&gt;
&lt;p&gt;Risk-based capital ratio (adjusted for captives): The author&amp;rsquo;s measure of insurer capitalization after consolidating the captive&amp;rsquo;s balance sheet with the operating company&amp;rsquo;s. This adjusted ratio is lower than the statutory risk-based capital ratio that ignores captives, by 23-24% in the 2019 sample, and translates into meaningfully higher default probabilities (from 1.0% to 2.9% over 10 years).&lt;/p&gt;
&lt;p&gt;Regulatory agency frictions: Deviations of state regulators&amp;rsquo; objective functions from a utilitarian social planner&amp;rsquo;s, captured by welfare weights (lambdas) on each component of the regulator&amp;rsquo;s utility. In the paper&amp;rsquo;s estimates, regulators over-weight tax revenues ($3.5 of tax revenues per $1 of default costs) and consumer surplus ($0.59 per $1 of default costs) relative to the social planner&amp;rsquo;s equal weighting, consistent with political economy pressures from consumers and revenue incentives.&lt;/p&gt;
&lt;p&gt;Captive capital rate: The state-level capital requirement on captives, defined empirically as the sum of capital divided by the sum of liabilities of all captives in the state each year. Higher values represent more stringent requirements. The mean in the sample is 4% with a standard deviation of 3%, and captive capital rates are lower on average than operating company capital rates.&lt;/p&gt;
&lt;p&gt;Race to the bottom: The dynamic under which competition between state regulators leads each state to set lower capital requirements than it would absent competition, in order to attract captive tax revenues, resulting in a collectively worse equilibrium with higher default risks. The paper finds this outcome in the baseline: competition lowers capital requirements by 19% (3 percentage points) relative to a federal regulator.&lt;/p&gt;
&lt;p&gt;External financing frictions: The costs insurers face in raising equity capital, modeled as a per-dollar cost theta on required capital. These frictions create the supply-side channel through which capital requirements affect insurance prices: higher capital requirements raise insurers&amp;rsquo; marginal costs, leading to higher prices and lower quantities, as documented in the Regulation XXX natural experiment.&lt;/p&gt;</description></item><item><title>Religion, Education, and the State</title><link>https://macropaperwarehouse.com/papers/religion-education-and-the-state/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/religion-education-and-the-state/</guid><description>&lt;p&gt;This paper studies how Indonesia&amp;rsquo;s Islamic education sector responded to one of the largest state-driven mass schooling expansions in history — the SD INPRES program (Sekolah Dasar Presidential Instruction) launched in 1973 — and whether that program achieved its secular nation-building objectives. The research question is three-part: Did Islamic schools enter or exit markets where the state built more primary schools? How did religious school choice shift across cohorts? And did the program advance secular identity formation among exposed individuals?&lt;/p&gt;
&lt;p&gt;The empirical setting is Indonesia in the 1970s onward. Under SD INPRES, the government used windfall oil revenues to build more than 61,000 primary schools between 1973 and 1980, allocating construction across districts proportional to the non-enrolled primary-school-age population. Because Islamic schools were historically more prevalent in underserved areas, this rule produced a strong positive correlation between SD INPRES intensity and pre-existing Islamic school density — the same markets where the state expanded were precisely those with the greatest Islamic education presence.&lt;/p&gt;
&lt;p&gt;The authors use several novel data sources: administrative registries covering nearly 220,000 secular and 160,000 Islamic schools with establishment dates; six rounds of the National Socioeconomic Survey (Susenas) from 2012–18; the Indonesia Family Life Survey (IFLS, 1993–2014); a 2018–19 curriculum timetable registry (SIAP) covering nearly 20% of madrasa; and a 2016 political/religious attitudes survey. Identification relies on difference-in-differences (DID) exploiting cross-district variation in SD INPRES intensity, the synthetic DID approach of Arkhangelsky et al. (2021) for robustness to violations of parallel trends, and a staggered village-level event study using the Borusyak et al. (2024) estimator.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. First, Islamic schools did not exit markets where the state expanded — they entered in greater numbers. A one standard deviation increase in SD INPRES construction led to 1.4 additional Islamic school entries per district above a mean of 1.9 per district in 1972. Entry was competitive at the primary level, where new madrasa (MI) entered at twice the baseline annual rate in the years immediately following INPRES construction, and strategic at the secondary level, where Islamic junior secondary schools (MTs) peaked 6–9 years after INPRES entry as graduates sought continued education. The Islamic sector financed this expansion through waqf (inalienable religious endowments), informal taxation (infaq, zakat), and revenues from a concurrent rice price spike; entry responses were stronger in villages with above-median waqf endowments and above-median potential rice yields.&lt;/p&gt;
&lt;p&gt;Second, rather than converging toward secular curricula, newly established Islamic schools in high-INPRES districts devoted more time to religious content. Each additional SD INPRES is associated with a 1.2 percentage point increase in the religious curriculum share among newly created madrasa, with increases of 1.3 and 2.4 percentage points at the primary and junior secondary levels respectively — the latter equaling 82% of the cross-school standard deviation. Some of this increase came at the expense of Pancasila/civic education and national language instruction.&lt;/p&gt;
&lt;p&gt;Third, while SD INPRES reduced Islamic primary school enrollment by roughly 7%, it increased overall Islamic school attendance: each additional SD INPRES increased the likelihood of attending any Islamic school by approximately 5%, as demand for secondary education outweighed substitution at the primary level. Female students exhibited stronger secondary-level demand effects, amplified in districts with a concurrent state ban on the Islamic veil in public schools.&lt;/p&gt;
&lt;p&gt;Fourth, SD INPRES did not advance its ideological objectives. In the 1977 and 1982 elections, Golkar&amp;rsquo;s vote share fell and the Islamic PPP&amp;rsquo;s rose by 0.5–1.0 percentage points per SD INPRES school in high-INPRES districts. Among exposed cohorts, SD INPRES did not increase Pancasila proficiency, national language use at home, or support for secular governance, but did increase Arabic literacy by approximately 3% per additional SD INPRES. Exposed cohorts also prayed more frequently, fasted more during Ramadan, gave more to charity, and expressed greater pilgrimage intentions. These religious patterns were transmitted to children of exposed cohorts, who were more likely to attend Islamic schools themselves.&lt;/p&gt;
&lt;p&gt;Q: What was the allocation rule for SD INPRES and why did it create confrontation with Islamic schools?
A: Presidential Instruction No. 10/1973 allocated school construction across districts proportional to the non-enrolled primary-school-age population in 1971. Because Islamic schools historically served underserved populations, this rule meant the state built more schools precisely where Islamic education was most prevalent. The paper shows graphically and in Table 1 that the number of SD INPRES schools built is strongly correlated with the pre-existing stock of Islamic schools, conditional on district population and enrollment.&lt;/p&gt;
&lt;p&gt;Q: How large was the Islamic sector&amp;rsquo;s entry response to SD INPRES at the district level?
A: In the standard DID specification (Table 2, panel a), a one standard deviation increase in SD INPRES schools led to 0.013 more Islamic schools per district-year per 1,000 children, equivalent to 1.4 additional Islamic school entries in the average district relative to a mean of 1.9 Islamic schools per district in 1972. The synthetic DID (panel b) delivers positive and slightly larger estimates, indicating the result is not an artifact of diverging pre-trends.&lt;/p&gt;
&lt;p&gt;Q: What was the timing of the Islamic sector entry response at the village level?
A: Using the Borusyak et al. (2024) estimator on a balanced panel from 1960 to 1999, the paper finds (Figure 4) that INPRES construction is followed by a jump in Islamic school entry. Primary madrasa (MI) entered at twice the baseline annual rate in the years immediately following INPRES construction and this elevated rate persisted for six years before reverting to baseline. Islamic junior secondary entry (MTs) peaked around years 6–9 after SD INPRES construction, consistent with newly graduated primary students seeking continued schooling.&lt;/p&gt;
&lt;p&gt;Q: How did the Islamic sector finance its expansion?
A: The sector relied on waqf endowments (inalienable religious land assets), informal faith-based contributions (infaq), and obligatory alms (zakat). Fortuitously, the initial year of SD INPRES coincided with a large spike in the global price of rice, Indonesia&amp;rsquo;s main agricultural commodity, boosting harvest revenues channeled through informal Islamic taxation. Table 3 shows that entry responses were significantly stronger in villages with above-median waqf endowments and above-median potential rice yields, and these heterogeneous effects did not arise in non-INPRES periods or for non-Islamic private schools. Survey data from 2007–13 further show higher rates of informal taxation in villages with Islamic schools built during this period.&lt;/p&gt;
&lt;p&gt;Q: Did Islamic schools converge toward secular curricula under competitive pressure from SD INPRES?
A: No. Table 4 shows that madrasa established in high-INPRES districts after 1972 devote more time to religious content, not less. Each additional SD INPRES is associated with a 1.2 percentage point increase in the share of classroom time devoted to religious subjects among newly created Islamic schools, with increases of 1.3 percentage points at the primary level and 2.4 percentage points at the junior secondary level — the latter equal to 82% of the cross-school standard deviation. Similar patterns hold for Arabic instruction, and the junior secondary increase comes partially at the expense of Pancasila/civic education and national language instruction.&lt;/p&gt;
&lt;p&gt;Q: Did curriculum differentiation responses vary with local religious ideology?
A: Yes. Appendix Table A.14 shows a stronger curriculum differentiation response in markets with greater historical support for conservative Islam, proxied by Islamic political party vote shares in the 1950s elections. The paper also constructs a school-name-based predicted ideology index using a ridge shrinkage estimator and finds (Appendix Table A.15) that madrasa entering high-INPRES districts after the program onset have a more religious ideology on this measure.&lt;/p&gt;
&lt;p&gt;Q: What happened to the formalization of the Islamic sector?
A: Figure 5 and Appendix Table A.6 show that formal madrasa entry increased as a share of all new school entry, while informal Islamic schools (pesantren, diniyah) declined as a share of all new schools and all new Islamic schools. This formalization mirrors the organizational structure of state schools (primary-to-secondary progression), facilitating switching between public and religious schools and providing option value to moderate but still religious families. Crucially, the newly entering formal madrasa introduced more religious curriculum than incumbent madrasa, so formalization did not reduce religious instruction.&lt;/p&gt;
&lt;p&gt;Q: What was the net effect of SD INPRES on Islamic school attendance?
A: Table 5 shows that SD INPRES reduced the likelihood of attending Islamic primary school by roughly 7% per additional SD INPRES school but increased Islamic secondary attendance, with the net effect being a roughly 5% increase in the likelihood of attending any Islamic school (column 4). This finding holds in both DID and synthetic DID. The IFLS validation (Appendix Table A.18) confirms decreased Islamic elementary attendance and increased Islamic junior secondary attendance, consistent with the Susenas results.&lt;/p&gt;
&lt;p&gt;Q: How does selection into secondary education affect the religious schooling results?
A: The authors address selection using parametric (Heckman 1976) and semiparametric (Newey 2009) selection-correction procedures, using exposure to a 1960s pilot compulsory schooling program as an exclusion restriction. Table 6, panels (c) and (d), show that selection-adjusted estimates are broadly consistent with unadjusted estimates, with similar signs and magnitudes. The selection-corrected estimates approximately identify a local average treatment effect among compliers: those induced to attend elementary school were less likely to attend Islamic elementary; those induced to continue to secondary were more likely to attend Islamic secondary.&lt;/p&gt;
&lt;p&gt;Q: How did gender shape the effects of SD INPRES on religious school choice?
A: Table 7 shows that SD INPRES had more limited impacts on total schooling for women than men (consistent with Duflo 2001) but that the secondary-level demand effect toward Islamic schools was stronger for women. Table 8 shows that within high-INPRES areas, the SD INPRES-induced increase in Islamic secondary education is three times larger for women in districts with greater exposure to the 1982 state ban on the Islamic veil in public schools, and this differential is specific to Islamic schooling rather than total schooling.&lt;/p&gt;
&lt;p&gt;Q: Did SD INPRES strengthen or weaken the secular ruling regime&amp;rsquo;s political standing?
A: It weakened it. Table 10 shows that in the 1977 and 1982 elections, Golkar&amp;rsquo;s vote share decreased and the Islamic PPP&amp;rsquo;s vote share increased in high-INPRES districts, in the range of 0.5–1.0 percentage points per SD INPRES school. This represents a 1.5–3.0% change in PPP vote share and a 0.5–1.0% change in Golkar vote share per standard deviation in SD INPRES intensity. The PPP gained most in areas where SD INPRES had the greatest potential to draw students away from Islamic schools.&lt;/p&gt;
&lt;p&gt;Q: Did SD INPRES produce a secular ideological shift among exposed cohorts?
A: No. Table 11 shows that SD INPRES did not increase self-reported Pancasila proficiency, national language use at home, national language literacy, or attitudes in favor of secular governance. By contrast, Arabic literacy increased by approximately 3% per additional SD INPRES among exposed cohorts, indicating that Islamic schooling exposure rather than secular schooling drove literacy gains in that language.&lt;/p&gt;
&lt;p&gt;Q: Did SD INPRES increase religiosity among exposed cohorts?
A: Yes. Table 12 shows that SD INPRES increased prayer frequency, fasting during Ramadan, charitable giving, and pilgrimage intentions among exposed cohorts. These effects on prayer and fasting are stronger among women, consistent with the stronger shift toward Islamic secondary schooling found in Table 7. These outcomes are consistent with greater exposure to Islamic education increasing religiosity rather than the secular curriculum reducing it.&lt;/p&gt;
&lt;p&gt;Q: Were the effects on religious identity and Arabic literacy transmitted to the next generation?
A: Yes. Table 13 shows that SD INPRES increased Arabic literacy among the children of exposed cohorts, and that children of exposed cohorts were more likely to attend Islamic schools themselves. These intergenerational results confirm that the preference for Islamic education instilled during the SD INPRES era persisted into the next generation rather than converging toward secular norms over time.&lt;/p&gt;
&lt;p&gt;Q: What are the key robustness checks for the school entry results?
A: Several checks support causal interpretation. The Roth and Rambachan (2022) procedure finds no systematic pre-trends in the standard DID. Historical Podes data from 1980, 1983, 1990, and 1993 confirm the post-1973 increase in Islamic school entry, addressing survival bias in the 2019 registry. Results are robust to allowing differential trends in waqf endowments, Muslim population share, Islamic party vote shares, historical Arab immigration, Islamist insurgency, and Transmigration resettlement. The heterogeneous entry responses by waqf and rice yield do not appear in non-INPRES periods or for non-Islamic schools.&lt;/p&gt;
&lt;p&gt;Q: What do the results imply for the political economy of education reform more broadly?
A: The paper argues that state capacity to homogenize culture through education is limited when strong non-state actors can mobilize their own resources and provide differentiated alternatives. Rather than crowding out religious schools, state expansion triggered competitive entry, curriculum differentiation, and formalization in the religious sector, producing an equilibrium where both sectors expanded simultaneously with distinct clienteles. The findings imply that the long-run cultural effects of education programs cannot be evaluated without accounting for equilibrium responses by competing non-state providers.&lt;/p&gt;
&lt;p&gt;SD INPRES (Sekolah Dasar Presidential Instruction): Indonesia&amp;rsquo;s 1973 mass public primary school construction program, financed by oil windfalls, which built more than 61,000 elementary schools between 1973 and 1980 by allocating schools to districts proportional to the non-enrolled child population; the program&amp;rsquo;s explicitly secular nation-building objectives brought it into direct confrontation with the Islamic education sector.&lt;/p&gt;
&lt;p&gt;Waqf: Inalienable Islamic religious endowments — of land, agricultural assets, or other property — that under Islamic law can only be used for religious or charitable purposes and cannot be seized or repurposed by the state; in this paper, the pre-existing waqf base in a village serves both as a long-run financing mechanism for Islamic school construction and as an index of Islamic sector organizational capacity.&lt;/p&gt;
&lt;p&gt;Madrasa: Formal day Islamic schools operating at the same primary-to-secondary grade levels as secular state schools, teaching standard academic subjects alongside a religious curriculum (including Islamic law, doctrine, ethics, Qur&amp;rsquo;an, Arabic, and history of the Prophets) that averages 26% of total instruction hours; distinct from the more informal pesantren (boarding schools) and madrasa diniyah (afternoon Qur&amp;rsquo;anic study schools).&lt;/p&gt;
&lt;p&gt;Curriculum differentiation: The strategy by which newly entering madrasa in high-INPRES districts increased the share of classroom time devoted to religious and Arabic instruction rather than converging toward the secular state curriculum; measured as classroom hours devoted to Islamic subjects, Arabic, Pancasila/civic education, and national language instruction from 2018–19 SIAP timetable data.&lt;/p&gt;
&lt;p&gt;Pancasila: The official secular nationalist ideology of the Indonesian state, consisting of five principles (monotheism, humanitarianism, national unity, democracy, and social justice) intended to transcend ethnic and religious divisions; SD INPRES sought to transmit Pancasila through civic education and national language instruction as part of its homogenizing nation-building agenda.&lt;/p&gt;
&lt;p&gt;Synthetic difference-in-differences (SDID): The Arkhangelsky et al. (2021) estimator used throughout the paper, which reweights and matches pre-INPRES trends in Islamic school construction across high- and low-INPRES exposure districts to deliver estimates more robust than standard DID to violations of parallel trends; applied with a binary treatment indicator (districts above the 51st percentile in INPRES intensity).&lt;/p&gt;
&lt;p&gt;Formalization: The documented shift in the composition of the Islamic sector after SD INPRES, whereby formal madrasa (organized along the same grade-level progression as state schools) increased as a share of all new Islamic school entry while informal pesantren and diniyah declined as a share; interpreted as a competitive response that expanded parental option value without sacrificing religious instruction intensity.&lt;/p&gt;</description></item><item><title>Rent Guarantee Insurance</title><link>https://macropaperwarehouse.com/papers/rent-guarantee-insurance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/rent-guarantee-insurance/</guid><description>&lt;p&gt;Abramson and Van Nieuwerburgh study Rent Guarantee Insurance (RGI), a product in which an insurer pays the landlord on behalf of a tenant who defaults on rent due to a negative income or health expenditure shock, in exchange for a monthly premium proportional to rent. The central question is whether RGI can be designed to be both welfare-improving and financially viable, given the frictions of moral hazard and adverse selection.&lt;/p&gt;
&lt;p&gt;The authors develop a dynamic overlapping-generations equilibrium model of the rental market that features endogenous rent default, security deposits, evictions, and homelessness. Households face idiosyncratic persistent and transitory income risk, idiosyncratic medical expenditure risk, and aggregate (cyclical) income risk. Rental contracts are non-contingent, households face borrowing constraints, and housing is indivisible with a minimum quality floor. Landlords set deposits to break even in expectation given observed tenant characteristics. An insurance agency can offer RGI and must also break even in the long run. The model is calibrated to the United States at monthly frequency. Income dynamics are estimated from CPS data (1994–2023) and incorporate transitions among employment, unemployment, out-of-labor-force, and retirement states along with transfer income (unemployment insurance, disability, food stamps) and a progressive tax system. Key moments targeted by Simulated Method of Moments include a delinquency rate of 12.15% (model: 12.69%), average security deposit of $984 (model: $992, from approximately 500,000 Craigslist listings across the 100 largest MSAs), homelessness rate of 1.43% (model: 1.42%), and home-ownership rate of 63.6% (model: 63.2%).&lt;/p&gt;
&lt;p&gt;The model&amp;rsquo;s pre-RGI analysis establishes that persistent income shocks — not transitory shocks or medical shocks — are the primary driver of rent defaults. Default risk remains elevated for 3–6 months following a persistent shock, implying that short-duration RGI coverage is insufficient to prevent eviction; coverage must span multiple months.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s main policy experiments introduce RGI under different access rules and provider types. Unrestricted RGI (available to all renters) generates large welfare gains through improved risk-sharing and lower security deposits — because insured tenants pose less default risk, landlords lower deposit requirements — but is not financially viable for either a public or private insurer due to moral hazard and adverse selection. Even a public insurer that internalizes the fiscal savings from reduced homelessness cannot break even under unrestricted access.&lt;/p&gt;
&lt;p&gt;Restricting access changes the viability calculus sharply. A publicly provided RGI targeted to households at the bottom of the wealth distribution can achieve financial viability: these households are precisely those most prone to homelessness, so the reduction in homelessness expenses — which the public insurer internalizes — offsets the insurance deficit. This restricted public RGI generates substantial welfare gains for the most vulnerable households.&lt;/p&gt;
&lt;p&gt;A privately provided RGI must instead target higher-wealth renters to break even, because these households have low default risk (limiting claim payouts) while remaining sufficiently risk averse to pay the premium. The intersection of financial viability and take-up is small, yielding a limited target audience. The private program has minimal impact on housing insecurity, and the most vulnerable households derive little benefit. This pattern matches observed private RGI markets, where providers restrict access to renters in good financial condition.&lt;/p&gt;
&lt;p&gt;An RGI mandate — requiring all renters to purchase coverage — mitigates adverse selection by improving the pool of insured tenants, dramatically increasing financial viability and allowing the insurer to reduce the premium substantially while still breaking even. Mandated RGI is highly effective at preventing housing insecurity and generates welfare gains concentrated among the most financially vulnerable households.&lt;/p&gt;
&lt;p&gt;Scope conditions: results are calibrated to U.S. income, medical, and housing market parameters as of 2019. The insurer&amp;rsquo;s borrowing cost matters: the public insurer faces lower, counter-cyclical municipal bond spreads, whereas private insurers face higher, pro-cyclical corporate spreads, which constrains the generosity of private contracts in recessions.&lt;/p&gt;
&lt;p&gt;Q: What is Rent Guarantee Insurance and how does it work mechanically in the model?
A: RGI is a contract under which a tenant pays a flat monthly premium equal to a fraction kappa of rent. When the insured tenant defaults, the insurer pays the landlord directly and deducts one period from the tenant&amp;rsquo;s stock of &amp;ldquo;insurance credit.&amp;rdquo; The tenant remains housed. Once insurance credit is exhausted, the insurer no longer covers defaults. The insurer sets the premium and the maximum coverage duration to break even in the long run.&lt;/p&gt;
&lt;p&gt;Q: Why do most rent defaults arise from persistent rather than transitory shocks?
A: The model shows that the renter population is disproportionately exposed to persistent unemployment and labor-force-exit spells, and that negative persistent income shocks are harder to smooth through savings than transitory ones. Default risk remains elevated for 3–6 months after a persistent shock but dissipates quickly after a transitory shock. This implies that RGI coverage periods of only a few months would fail to prevent eviction for the majority of defaulting tenants.&lt;/p&gt;
&lt;p&gt;Q: How does RGI affect security deposits in equilibrium?
A: Because landlords observe the tenant&amp;rsquo;s insurance status at lease signing and deposits are set to make landlords break even in expectation, insured tenants pose lower default risk and thus face lower upfront deposit requirements. This deposit reduction is a key welfare channel of RGI, as large deposits tie up a disproportionate share of poor households&amp;rsquo; wealth and price the most vulnerable out of housing entirely.&lt;/p&gt;
&lt;p&gt;Q: Why is unrestricted RGI financially non-viable even for the public insurer?
A: Unrestricted access induces both adverse selection — riskier households self-select into coverage — and moral hazard — insured households alter their default and savings behavior. These effects cause the insurer to run a persistent deficit. Even a public insurer that internalizes the fiscal cost savings from reduced homelessness cannot recoup enough to break even, implying that an unrestricted program would require an ongoing subsidy.&lt;/p&gt;
&lt;p&gt;Q: How does publicly provided restricted RGI achieve financial viability?
A: By targeting households at the bottom of the wealth distribution — precisely those most prone to homelessness — the public RGI program produces large reductions in homelessness. Because the public insurer internalizes the fiscal expenses associated with shelters, health services, and policing that accompany homelessness, these savings are passed through to the insurer and are sufficient to offset the insurance deficit. No such mechanism is available to a private insurer.&lt;/p&gt;
&lt;p&gt;Q: Why must private RGI target higher-wealth renters, and what are the consequences?
A: Private insurers must break even using only premium revenue, without access to homelessness cost savings. Higher-wealth renters have lower default probabilities, which limits claim payouts, while remaining sufficiently risk averse to demand coverage and pay the premium. The viable target audience is small given these competing requirements. As a result, private RGI covers few households, has minimal effect on housing insecurity, and provides essentially no benefit to the most vulnerable renters. This pattern is consistent with observed private RGI markets.&lt;/p&gt;
&lt;p&gt;Q: What are the two differences between public and private insurers in the model?
A: First, the public insurer internalizes the fiscal costs of homelessness (shelters, health services, policing), raising its net benefit from offering coverage. Second, the public insurer borrows at municipal bond spreads — which are lower than corporate spreads and counter-cyclical — whereas the private insurer faces higher, pro-cyclical corporate spreads. Counter-cyclical borrowing costs allow the public insurer to extend more generous coverage precisely when aggregate conditions deteriorate and claims rise.&lt;/p&gt;
&lt;p&gt;Q: How does an RGI mandate improve financial viability?
A: Mandatory enrollment forces all renters, including low-risk ones, into the insurance pool, which counteracts adverse selection. The expanded and higher-quality pool dramatically reduces per-insured expected claim costs, allowing the insurer to lower the premium substantially while still breaking even. The low-premium mandated policy is then both affordable and effective at preventing housing insecurity, with welfare gains concentrated among the most financially vulnerable renters.&lt;/p&gt;
&lt;p&gt;Q: What novel data does the paper use for calibration of security deposits?
A: The authors construct a dataset of approximately 500,000 Craigslist rental listings scraped across the 100 largest U.S. metropolitan statistical areas between November 2022 and March 2024 to measure the cross-sectional distribution of security deposits. The average deposit in this dataset is $984, which the model matches closely at $992. The data also reveal that the deposit-to-rent ratio is decreasing in house quality, reflecting the higher default risk of low-income renters in lower-quality units.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s definition of homelessness and what rate does the model match?
A: Homelessness is defined broadly to include sheltered homeless, unsheltered homeless (0.6% of households), and doubled-up families (0.83% of households), for a total of 1.43% of U.S. households. The model matches this rate closely at 1.42%.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s key implication for the design of housing policy?
A: The central implication is that financial viability and impact on housing insecurity are in tension for private insurers, and cannot both be achieved simultaneously. Only a publicly provided program that internalizes homelessness fiscal costs and faces counter-cyclical borrowing spreads can target the most vulnerable renters, break even, and materially reduce housing insecurity. Private RGI, while viable for a narrow segment, cannot substitute for public provision as a tool against homelessness.&lt;/p&gt;
&lt;p&gt;Q: How does RGI relate conceptually to rental assistance programs?
A: The paper distinguishes RGI from rental assistance on a structural basis: insurance contracts require tenants to pay premiums, making them potentially self-financing for private providers, whereas rental assistance is a net transfer that can never be self-financing. This conceptual distinction motivates studying whether RGI can be designed to eliminate the need for ongoing fiscal transfers, though the analysis ultimately shows that a public subsidy or mandate is required to serve the most vulnerable renters.&lt;/p&gt;
&lt;p&gt;Rent Guarantee Insurance (RGI): A contract under which an insured tenant pays a monthly premium equal to a flat percentage of rent; when the tenant defaults, the insurer pays the landlord directly, preserving tenancy, for a limited number of periods governed by the tenant&amp;rsquo;s stock of insurance credit.&lt;/p&gt;
&lt;p&gt;Insurance Credit: An endowment of periods of RGI coverage that households receive upon entry into the model; each time the insurer pays on behalf of a defaulting tenant, one unit of credit is consumed, and no further coverage is available once credit is exhausted.&lt;/p&gt;
&lt;p&gt;Housing Insecurity: In the paper&amp;rsquo;s framework, the set of outcomes — rent delinquency, eviction, and homelessness — arising from the combination of non-contingent rental contracts, borrowing constraints, and idiosyncratic or aggregate income and medical shocks.&lt;/p&gt;
&lt;p&gt;Security Deposit: An upfront payment from tenant to landlord, set by the competitive landlord to break even in expectation given the tenant&amp;rsquo;s characteristics and insurance status; a key channel through which RGI affects welfare by reducing the upfront cost barrier to obtaining housing.&lt;/p&gt;
&lt;p&gt;Moral Hazard (in RGI context): The change in a tenant&amp;rsquo;s default, savings, and housing choices induced by the presence of insurance coverage, which increases expected claim costs for the insurer relative to a world where behavior is held fixed.&lt;/p&gt;
&lt;p&gt;Adverse Selection (in RGI context): The tendency of renters with higher default risk to self-select into RGI when access is unrestricted, worsening the insurer&amp;rsquo;s risk pool and driving up expected payouts relative to premiums.&lt;/p&gt;
&lt;p&gt;Homelessness Externality: The fiscal costs borne by government — for shelters, health services, and policing — that accompany homelessness; the public insurer internalizes these costs, creating a net benefit from RGI that private insurers cannot capture.&lt;/p&gt;
&lt;p&gt;Counter-cyclical Borrowing Spread: The feature of public (municipal bond) financing whereby borrowing costs fall during recessions, allowing the public insurer to expand coverage when claims are highest; contrasted with private insurers&amp;rsquo; pro-cyclical corporate bond spreads that tighten precisely when aggregate conditions worsen.&lt;/p&gt;</description></item><item><title>Republican Support and Economic Hardship: The Enduring Effects of the Opioid Epidemic</title><link>https://macropaperwarehouse.com/papers/republican-support-and-economic-hardship-the-enduring-effects-of-the-opioid-epidemic/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/republican-support-and-economic-hardship-the-enduring-effects-of-the-opioid-epidemic/</guid><description>&lt;p&gt;This paper establishes a causal connection between the opioid epidemic and the political realignment toward the Republican Party in the United States from the mid-2000s through 2022. The authors—Carolina Arteaga and Victoria Barone—exploit rich geographic variation in Purdue Pharma&amp;rsquo;s initial marketing strategy for OxyContin, drawn from unsealed litigation records, to construct a quasi-exogenous measure of community-level exposure to the epidemic.&lt;/p&gt;
&lt;p&gt;The identification strategy rests on a documented feature of OxyContin&amp;rsquo;s 1996 launch: Purdue initially targeted the established cancer pain market—physicians and patients already using MS Contin—as an entry point into the much larger noncancer pain market. Areas with higher cancer mortality in 1996 received disproportionate pharmaceutical marketing, leading to outsized opioid prescription growth that spilled over from cancer patients to the broader population through shared physicians. The authors use 1996 commuting-zone (CZ) cancer mortality rates as a proxy for this initial targeting, interacted with year fixed effects in an event-study specification with CZ and state-year fixed effects. The sample covers 625 CZs across the continental United States from 1982 to 2022.&lt;/p&gt;
&lt;p&gt;The empirical chain runs through three stages. First, the instrument strongly predicts opioid supply: by 2012, a one-standard-deviation higher 1996 cancer mortality rate led to an additional 0.97 opioid doses prescribed per capita, 65% above the baseline mean. Second, the resulting epidemic caused measurable mortality and economic hardship. A one-standard-deviation increase in 1996 cancer mortality caused drug-induced deaths in 2017 to be 46% above the pre-epidemic average; by 2012 the same increase caused prescription opioid deaths to be 61% higher. Excess mortality was concentrated among individuals under age 55, with no significant effects for those aged 55 and older. The epidemic also raised disability applications: SSDI applications rose by 12% and SSI applications by 7.6% by 2012, effects that persisted through 2020. SNAP enrollment in exposed CZs was 8% higher by 2022, equivalent to a 0.14 standard deviation increase.&lt;/p&gt;
&lt;p&gt;Third, and centrally, the communities that endured these health and economic shocks shifted persistently toward the Republican Party. By the 2022 House elections, a one-standard-deviation increase in 1996 cancer mortality increased the Republican two-party vote share by 4.5 percentage points. Effects of similar magnitude appear in presidential elections (4.6 percentage points) and gubernatorial elections (4.3 percentage points). The vote-share shift is consistent across gender, age, race, and education, with no detectable change in voter turnout. The shift translates into actual seat gains: beginning in 2012, exposed areas consistently elected more Republican House members, moving the chamber&amp;rsquo;s roll-call voting in a more conservative direction. The effect is not driven by anti-incumbent sentiment—results hold regardless of which party held the seat at the time.&lt;/p&gt;
&lt;p&gt;The paper identifies three reinforcing mechanisms. First, the Republican Party repositioned itself during this period as the advocate of &amp;ldquo;forgotten America&amp;rdquo; and working-class economic hardship, a message that resonated acutely in opioid-devastated communities. Second, conservative-leaning newspapers covered the epidemic at higher rates, and their coverage tracked local mortality; liberal-leaning outlets showed no such correlation. Fox News covered opioid stories at 1.5 times the rate of CNN and 1.7 times the rate of MSNBC, emphasizing crime, trafficking, and cartels at twice the frequency of liberal outlets. Third, exposed communities expressed stronger preferences for Republican-favored policy responses: higher police presence, greater sense of safety around law enforcement, and lower support for marijuana legalization on state ballot initiatives.&lt;/p&gt;
&lt;p&gt;Pre-trend tests show no relationship between 1996 cancer mortality and outcomes before OxyContin&amp;rsquo;s launch. Out-of-sample exercises using 1976 cancer mortality find no analogous pattern in the pre-epidemic period (1982–1994). Placebo instruments based on unrelated causes of death yield null results. The baseline findings are robust to controlling for the China import shock, NAFTA, the 1994 Republican Revolution, the 2001 and 2008–2009 recessions, declining unionization, robot adoption, Fox News introduction, deaths of despair, and Southern and rural political realignment.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s central research question?
A: The paper asks whether the opioid epidemic causally increased Republican vote share in communities most severely affected by the crisis. It documents a causal chain from pharmaceutical marketing through drug mortality and economic hardship to political realignment, contributing the first causal estimate of a major public health crisis&amp;rsquo;s effect on partisan voting.&lt;/p&gt;
&lt;p&gt;Q: What is the identification strategy, and why is 1996 cancer mortality a valid instrument?
A: Purdue Pharma explicitly targeted physicians in the cancer pain market at OxyContin&amp;rsquo;s 1996 launch, then used those established relationships to expand into the noncancer pain market. CZs with higher cancer mortality in 1996 received disproportionate marketing, generating differential opioid prescription growth unrelated to pre-existing political or economic trends. Pre-trend tests confirm no differential patterns before 1996, out-of-sample tests using 1976 cancer mortality find no relationship with pre-epidemic outcomes, and placebos using unrelated causes of death yield null results.&lt;/p&gt;
&lt;p&gt;Q: How strong is the first stage linking 1996 cancer mortality to opioid prescriptions?
A: The relationship between 1996 cancer mortality and opioid prescriptions is positive and statistically significant from 1998 through 2020. By 2012—the year prescription rates peaked nationally at 81.3 per 100 persons—a one-standard-deviation higher cancer mortality rate led to an additional 0.97 morphine-equivalent doses prescribed per capita, 65% above the baseline mean. CZs in the highest cancer mortality quartile experienced a 1,800% increase in grams of oxycodone per capita between 1997 and 2010, compared to less than half that in the lowest quartile.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on drug-induced mortality?
A: Drug-induced mortality (a broad measure covering deaths from prescription opioids, heroin, and fentanyl) rose continuously in exposed CZs after 1996. By 2017, a one-standard-deviation increase in 1996 cancer mortality caused drug-induced deaths to be 46% above the pre-epidemic average. By 2012, the same increase caused prescription opioid deaths specifically to be 61% higher relative to the pre-epidemic average. Excess mortality was concentrated among individuals under age 55, with no statistically significant effects for those aged 55 and older.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on disability program take-up?
A: Applications for Social Security Disability Insurance (SSDI) rose by 12% and Supplemental Security Income (SSI) applications rose by 7.6% by 2012 for a one-standard-deviation increase in 1996 cancer mortality. These effects persisted: SSDI recipients grew by 15% and SSI recipients by 3.2% by 2020 in similarly exposed CZs. The increases in disability were concentrated among individuals under age 55, paralleling the mortality effects.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on SNAP enrollment?
A: Exposed CZs showed a continuous increase in SNAP enrollment over two decades following the epidemic&amp;rsquo;s onset. By 2020, a one-standard-deviation increase in 1996 cancer mortality corresponded to an 8% increase in the share of the population receiving SNAP benefits, equivalent to 0.14 standard deviations. By 2022, the corresponding figure remains 8%, indicating persistent economic strain in exposed communities.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the political effects in House elections?
A: A one-unit increase in the 1996 cancer mortality rate yielded a 7.9-percentage-point increase in the 2022 Republican House vote share relative to 1996. Scaled to one standard deviation (0.58 units), this corresponds to a 4.5-percentage-point increase in the Republican two-party vote share by the 2022 midterms. The vote-share shift became statistically significant beginning in 2006, but only translated into consistent seat-level Republican gains starting in 2012.&lt;/p&gt;
&lt;p&gt;Q: When did opioid exposure start winning Republicans additional House seats?
A: Although the Republican vote share in exposed areas began increasing around 2006, actual seat flips did not become consistent until 2012. The paper explains this lag by noting that initial vote-share gains were concentrated in communities with low baseline Republican support, where additional votes did not immediately cross the winning threshold. Starting in 2010, median-baseline-Republican CZs also began shifting, enabling additional seat changes.&lt;/p&gt;
&lt;p&gt;Q: How large are the presidential and gubernatorial election effects?
A: In presidential elections, a one-standard-deviation increase in 1996 cancer mortality raised the Republican vote share by 4.6 percentage points. In gubernatorial elections, the same increase raised the Republican vote share by 4.3 percentage points after approximately six election cycles (corresponding to 2017–2020). These effects are described as comparable in magnitude to the difference in Republican vote share between the top and bottom quartiles of NAFTA vulnerability.&lt;/p&gt;
&lt;p&gt;Q: Does the political shift reflect increased polarization toward extremist candidates?
A: No. The paper finds no increase in the probability of electing candidates at the extremes of the Nokken-Poole ideological scale in any given election year. The ideological shift in the House composition arises from changes in which party wins seats rather than from the election of more extreme Republicans. Campaign donations to Republican candidates did not increase; rather, donations to Democratic candidates declined (in 2016, a one-standard-deviation increase in cancer mortality widened the Republican-Democrat donation gap by 0.44 standard deviations). The shift is interpreted as a change in voting preferences in previously Democratic-leaning areas, not heightened polarization.&lt;/p&gt;
&lt;p&gt;Q: Is the shift driven by anti-incumbent sentiment?
A: The authors test this by splitting the sample by the incumbent&amp;rsquo;s party at the time of each election and by redefining the outcome as the incumbent&amp;rsquo;s vote share. Neither exercise produces evidence of a systematic anti-incumbent response. The changes in Republican vote share are not statistically distinguishable based on whether the incumbent was a Republican or Democrat. If anything, after 2016 there is a slight increase in the likelihood of incumbents retaining their seats.&lt;/p&gt;
&lt;p&gt;Q: Where geographically are the Republican gains largest?
A: Using state-level treatment effects estimated from an in-differences model interacting cancer mortality with state-year indicators, the paper finds a strong positive correlation between the magnitude of the epidemic&amp;rsquo;s effect on economic hardship (measured by SNAP participation) and the magnitude of the Republican vote-share increase. This correlation is strongest with a lag: SNAP effects measured in 2006 are most predictive of vote-share shifts in 2022, indicating that deterioration in community economic fabric preceded and predicted the political realignment.&lt;/p&gt;
&lt;p&gt;Q: How did conservative and liberal media differ in covering the opioid epidemic?
A: Republican-leaning local newspapers covered the opioid epidemic more extensively than Democratic-leaning papers throughout the epidemic period, and their coverage tracked local opioid mortality rates; Democratic-leaning coverage showed no such correlation with local incidence. Fox News covered opioid stories at 1.5 times the rate of CNN and 1.7 times the rate of MSNBC. In terms of content, Republican-leaning newspapers showed 23% higher frequency of economic hardship keywords, 19% higher frequency of illegal activity and crime keywords, and 22% higher frequency of rehabilitation and treatment keywords relative to Democratic-leaning papers. Fox News emphasized crime, drug trafficking, and cartels at double the frequency of more liberal outlets.&lt;/p&gt;
&lt;p&gt;Q: How did voter policy preferences align with Republican versus Democratic platforms?
A: Using 2020 CCES data, the authors find that higher 1996 cancer mortality predicts a greater expressed preference for increasing the number of police officers on the street and a greater reported sense of safety around law enforcement—both consistent with the Republican Party&amp;rsquo;s law enforcement approach. Conversely, exposure to the epidemic predicts lower support for marijuana legalization on state ballot initiatives across 18 states from 2012 to 2023, indicating opposition to a key Democratic harm-reduction policy.&lt;/p&gt;
&lt;p&gt;Q: What role did political actors themselves play in driving the realignment?
A: Relatively little. The opioid epidemic was largely absent from House floor speeches until 2015 and from campaign advertising until 2020. Neither party took a clear legislative lead on the issue during the first two decades of the crisis. The authors interpret the political realignment as driven primarily by the Republican Party&amp;rsquo;s broader repositioning as the champion of working-class economic hardship and by differential media framing, rather than by active legislative competition over opioid policy.&lt;/p&gt;
&lt;p&gt;Q: What major confounds are ruled out?
A: The authors control for exposure to the China import shock, NAFTA, the 1994 Republican Revolution, the 2001 and 2008–2009 recessions, declining unionization, robot adoption, Fox News entry, deaths of despair (which include but are not limited to opioid deaths), and the political realignment of the South, rural areas, evangelicals, and the population over 65. Results remain robust across all these specifications. Placebo instruments using unrelated causes of death yield null results.&lt;/p&gt;
&lt;p&gt;Q: Could the vote-share effects be mechanically driven by opioid-related deaths removing Democratic voters from the electorate?
A: The authors perform a back-of-the-envelope calculation and estimate that even if all opioid-related deaths would have been Democratic votes, the mechanical effect on the Republican vote share is at most 0.22 percentage points relative to the observed 2020 vote share—far smaller than the estimated 4.5-percentage-point shift by 2022. The result is also inconsistent with a turnout mechanism, as voter turnout shows no meaningful change with epidemic exposure.&lt;/p&gt;
&lt;p&gt;Opioid epidemic exposure instrument: The paper measures community-level exposure to the opioid epidemic using cancer mortality rates in 1996, the year OxyContin launched. This instrument is grounded in Purdue Pharma&amp;rsquo;s documented marketing strategy of targeting the cancer pain market first; areas with more cancer patients received disproportionate pharmaceutical marketing, generating differential opioid prescription growth that extended well beyond cancer patients to the broader noncancer population through shared physicians.&lt;/p&gt;
&lt;p&gt;Commuting zone (CZ): The paper&amp;rsquo;s unit of geographic analysis, defined to capture local economic markets. There are 720 CZs in the US, encompassing all metropolitan and nonmetropolitan areas. The authors use 625 CZs with more than 20,000 residents, which account for more than 99% of all opioid deaths and total population.&lt;/p&gt;
&lt;p&gt;Two-party Republican vote share: The ratio of votes for Republican candidates to the total votes for both Republican and Democratic candidates in a given election. The paper tracks this measure for House, presidential, and gubernatorial elections from 1976 or 1982 through 2020 or 2022, depending on data availability.&lt;/p&gt;
&lt;p&gt;Drug-induced mortality: The paper&amp;rsquo;s broadest mortality measure, covering deaths from poisoning and medical conditions caused by legal or illegal drugs, including prescription opioids, heroin, and synthetic opioids such as fentanyl. It is distinguished from the narrower measures of prescription opioid deaths and all opioid deaths.&lt;/p&gt;
&lt;p&gt;Issue ownership: The political science concept, used in the paper to describe how the Republican Party repositioned itself during the epidemic period as the voice of working-class economic hardship, &amp;ldquo;forgotten America,&amp;rdquo; and &amp;ldquo;America left behind.&amp;rdquo; The paper contrasts this with Democratic ownership of income inequality and argues that Republican ownership of the hardship narrative made the party&amp;rsquo;s message especially salient in heavily opioid-affected communities.&lt;/p&gt;
&lt;p&gt;Path dependency in pharmaceutical marketing: Purdue&amp;rsquo;s strategy of concentrating initial OxyContin promotion in cancer-market areas, then later focusing on top-prescribing physicians (the highest three deciles of the distribution), meant that areas receiving high initial cancer-market promotion continued to receive disproportionate promotion as the company expanded to the noncancer market. This created a persistent targeting advantage for high-cancer CZs throughout the epidemic&amp;rsquo;s first wave.&lt;/p&gt;
&lt;p&gt;Nokken-Poole ideological measure: A roll-call-based measure of elected House members&amp;rsquo; ideology along the liberal-conservative dimension. The paper uses this measure to show that the epidemic shifted the composition of the House toward more conservative members, not by electing more extreme candidates in any given election, but by changing which party won seats over time.&lt;/p&gt;</description></item><item><title>Running Primary Deficits Forever in a Dynamically Efficient Economy: Feasibility and Optimality</title><link>https://macropaperwarehouse.com/papers/running-primary-deficits-forever-in-a-dynamically-efficient-economy-feasibility-and-optimality/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/running-primary-deficits-forever-in-a-dynamically-efficient-economy-feasibility-and-optimality/</guid><description>&lt;h2 id="running-primary-deficits-forever-in-a-dynamically-efficient-economy-feasibility-and-optimality"&gt;Running Primary Deficits Forever in a Dynamically Efficient Economy: Feasibility and Optimality&lt;/h2&gt;
&lt;h3 id="research-question"&gt;Research Question&lt;/h3&gt;
&lt;p&gt;The paper addresses two questions about government debt rollover. First, a positive question: what is the maximum ratio of government bonds to capital that can be sustained forever without any primary budget surpluses? Second, a normative question: among sustainable bond-capital ratios along a balanced growth path, which one maximizes the welfare (steady-state utility) of consumers? The analysis is motivated by Blanchard&amp;rsquo;s (2019) AEA presidential address and the fiscal responses to the COVID-19 pandemic.&lt;/p&gt;
&lt;h3 id="setting-and-mechanism"&gt;Setting and Mechanism&lt;/h3&gt;
&lt;p&gt;The baseline environment is a standard two-generation (young and old) overlapping-generations model. Young consumers earn labor income and save; old consumers live off portfolio returns. The production function is Cobb-Douglas, Yt = (GtN)^(1−α) K^α, where G = 1+g is the gross growth rate of labor-augmenting productivity. Uncertainty enters exclusively through a stochastic i.i.d. durability shock ε_t to the depreciation rate of capital (δ − ε_t), so the rate of return on capital r = αk^(α−1) − δ + ε is stochastic even though the capital stock per unit of effective labor k is deterministic along a balanced growth path. Consumers have Epstein-Zin-Weil utility with an intertemporal elasticity of substitution equal to one. Because IES = 1 and labor income is earned only when young, aggregate saving of young consumers is a constant fraction β of their wage income, making total assets (capital plus bonds) non-stochastic.&lt;/p&gt;
&lt;p&gt;This structure creates a key wedge: the expected rate of return on capital R can exceed the growth rate g (dynamic efficiency) while the riskfree interest rate rf — determined by the portfolio equilibrium between risky capital and riskless bonds — can remain below g. In deterministic economies these two rates coincide, so dynamic efficiency and the infeasibility of permanent debt rollover always go together. In this stochastic model they can be decoupled.&lt;/p&gt;
&lt;h3 id="main-findings"&gt;Main Findings&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Positive finding.&lt;/strong&gt; The maximum sustainable bond-capital ratio, Bmax, is attained precisely when rf = g (equivalently, when the adjusted gross riskfree rate Rf = 1). Starting from a bond-less economy with rf &amp;lt; g (which may itself be dynamically efficient), introducing government bonds crowds out capital, raises the marginal product of capital and the constellation of returns, and drives rf upward toward g. Once rf = g is reached, any further increase in bonds would require rf &amp;gt; g, making rollover infeasible without primary surpluses. The maximum sustainable ratio Bmax is characterized as the unique root of f(Bmax, 1) = 0, and it is finite.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Normative finding.&lt;/strong&gt; The welfare-maximizing sustainable bond-capital ratio equals Bmax. Proposition 6 establishes that u′(B) ≥ 0 for all B ∈ [0, Bmax] whenever Rf ≤ 1, with strict inequality unless Rf = 1. Proposition 7 therefore concludes that the welfare-maximizing B is the corner solution Bmax. Intuitively, increasing B reduces capital and wages but raises the rate of return on capital. When rf ≤ g, the welfare gain from a higher return on capital in old age dominates the welfare loss from a lower wage when young (via the factor-price frontier and the intertemporal optimality condition E{uo′(co)} ≥ uy′(cy)). When rf = g (at Bmax), a marginal increase in bonds also provides no additional welfare improvement if all seignorage is transferred to young consumers (ζ = 1), but still raises welfare if some seignorage is wasted (ζ &amp;lt; 1). In either case, Bmax is the optimum. Critically, at the optimum the economy is dynamically efficient — even though the government is running permanent primary deficits.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dual role of bonds.&lt;/strong&gt; At the optimal bond-capital ratio, government bonds serve two purposes simultaneously: (1) they crowd out any dynamically inefficient overaccumulation of capital that might prevail without bonds, and (2) they supply riskfree assets to risk-averse consumers who would otherwise hold only risky capital, improving risk sharing.&lt;/p&gt;
&lt;h3 id="quantitative-illustration"&gt;Quantitative Illustration&lt;/h3&gt;
&lt;p&gt;The paper calibrates a 30-year-period OLG model with α = 0.33, β = 0.353 (annual discount rate 2%), annual productivity growth g = 1% (G = 1.35), and target mean return on unlevered equity m = 3% per year. Risk aversion γ ∈ {1, 3, 8, 10} and annualized standard deviation of capital returns s ∈ {0.02, …, 0.22}. Key results (ζ = 0): at γ = 10 and s = 0.22, Bmax = 0.478 and B∗ (the bond-capital ratio needed just to eliminate dynamic inefficiency) = 0.083, so there is a wide interval [0.083, 0.478] of dynamically efficient, permanently rollable bond-capital ratios. For a capital-output ratio of 2, the debt-GDP ratio corresponding to Bmax = 0.478 is approximately 0.956. Bmax is strictly increasing in both γ and s, and is invariant to ζ (the share of seignorage transferred rather than wasted).&lt;/p&gt;
&lt;h3 id="scope-conditions"&gt;Scope Conditions&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Results hold along balanced growth paths with constant g and constant rf; the sustainability characterization is more complex if either rate is stochastic.&lt;/li&gt;
&lt;li&gt;The key sufficient condition for Rf to be increasing in B (Proposition 1) is that risk aversion γ &amp;lt; Λ, a model-dependent upper bound that is always positive. All subsequent propositions assume R′f(B) &amp;gt; 0, which is satisfied for a potentially larger set of γ.&lt;/li&gt;
&lt;li&gt;The paper focuses on welfare along the balanced growth path; it does not study transition dynamics or welfare during convergence from an initial state.&lt;/li&gt;
&lt;li&gt;The No Ponzi Game (NPG) condition is violated by design in the feasible-rollover region (rf ≤ g); the value of government bonds is positive even though the present value of all future primary surpluses is non-positive.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-can-an-economy-be-both-dynamically-efficient-and-able-to-roll-over-government-bonds-forever-when-this-is-impossible-in-deterministic-models"&gt;Q1. Why can an economy be both dynamically efficient and able to roll over government bonds forever, when this is impossible in deterministic models?&lt;/h3&gt;
&lt;p&gt;In a deterministic economy, the riskfree rate rf and the rate of return on capital r are equal, so the conditions rf &amp;lt; g (feasibility of rollover) and r &amp;lt; g (dynamic inefficiency) are identical. In a stochastic economy, aggregate uncertainty drives a wedge between rf and the expected return on capital. Risk-averse consumers require a premium to hold risky capital over riskless bonds, so rf &amp;lt; E{r}. It is therefore possible that E{ln R} &amp;gt; 0 (the Zilcha sufficient condition for dynamic efficiency holds) while Rf &amp;lt; 1, i.e., rf &amp;lt; g. This decoupling is the central theoretical contribution of the paper.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-formal-criterion-the-paper-uses-for-dynamic-efficiency-and-how-does-it-relate-to-the-amsz-criterion"&gt;Q2. What is the formal criterion the paper uses for dynamic efficiency, and how does it relate to the AMSZ criterion?&lt;/h3&gt;
&lt;p&gt;Abel, Mankiw, Summers, and Zeckhauser (AMSZ, 1989) show that if the rate of return on capital exceeds g in all states (R &amp;gt; 1 always), the economy is dynamically efficient, and since rf &amp;lt; r, the economy has rf &amp;gt; g so rollover is infeasible; conversely if r &amp;lt; g always, the economy is dynamically inefficient. The AMSZ criteria are silent when R sometimes exceeds and sometimes falls short of one. Building on Zilcha (1991), the paper uses E{ln R} ≥ 0 as a sufficient condition for dynamic efficiency. In the five-region diagram (Figure 1), Region E satisfies E{ln R} &amp;gt; 0 (Zilcha-efficient) and Rf &amp;lt; 1 (rollover feasible simultaneously), which is the case of central interest.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-model-achieve-a-deterministic-capital-stock-despite-stochastic-capital-returns"&gt;Q3. How does the model achieve a deterministic capital stock despite stochastic capital returns?&lt;/h3&gt;
&lt;p&gt;The durability shock ε_t affects depreciation but is additively separable from the production function. Because (1) IES = 1 and (2) consumers earn income only when young, aggregate saving is the fixed fraction β of wage income, which depends only on capital k (itself non-stochastic). Total assets At+1 = Kt+1 + Bt+1 = St are thus non-stochastic. The stochastic shock to depreciation makes the rate of return on capital r = αkα−1 − δ + ε stochastic even though k is deterministic. Online Appendix B establishes that this model is isomorphic to a model with production function shocks, extending the scope of the results.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-financial-market-equilibrium-condition-that-pins-down-the-riskfree-rate"&gt;Q4. What is the financial market equilibrium condition that pins down the riskfree rate?&lt;/h3&gt;
&lt;p&gt;Young consumers optimally choose the portfolio share λ in riskfree bonds. The first-order condition for this portfolio problem along a balanced growth path is E{(λRf + (1−λ)R)^(−γ)(Rf − R)} = 0 (equation 20). In equilibrium, λ = B/(1+B) (the bond-capital ratio determines the portfolio share), so the equilibrium riskfree rate Rf satisfies the implicit equation f(B, Rf) = 0 (equation 21). Lemma 1 establishes that Rf = E{R^(1−γ)_a}/E{R^(−γ)_a}, a ratio-of-moments formula analogous to an Euler equation.&lt;/p&gt;
&lt;h3 id="q5-why-is-the-riskfree-rate-rf-an-increasing-function-of-the-bond-capital-ratio-b-and-what-is-the-sufficient-condition-for-this"&gt;Q5. Why is the riskfree rate Rf an increasing function of the bond-capital ratio B, and what is the sufficient condition for this?&lt;/h3&gt;
&lt;p&gt;Lemma 2 shows ∂f/∂B &amp;gt; 0; intuitively, more bonds reduce capital, raise the marginal product of capital, and raise R, inducing consumers to demand more capital and less bonds, pushing Rf up to restore equilibrium. Lemma 3 provides a sufficient condition for ∂f/∂Rf &amp;lt; 0, namely γ &amp;lt; Λ (where Λ is a positive parameter-dependent bound). Under this condition, the implicit function theorem implies Rf′(B) &amp;gt; 0 (Proposition 1). The condition γ &amp;lt; Λ is sufficient but not necessary, so the results of all downstream propositions hold potentially for a wider parameter range.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-maximum-sustainable-bond-capital-ratio-bmax-and-how-is-it-characterized"&gt;Q6. What is the maximum sustainable bond-capital ratio Bmax, and how is it characterized?&lt;/h3&gt;
&lt;p&gt;By definition, a bond-capital ratio B is sustainable if and only if Rf(B) ≤ 1. If Rf(0) ≥ 1, then Bmax = 0 (no positive amount of bonds is sustainable). If Rf(0) &amp;lt; 1, Bmax is the unique positive root of Rf(B) = 1, i.e., f(Bmax, 1) = 0 (Proposition 4). At Bmax, the riskfree rate exactly equals the growth rate: rf = g. The paper also shows Bmax ≤ (1−α)β/α − 1, an upper bound that depends only on production and preference parameters. Notably, Bmax is invariant to the parameter ζ (the share of seignorage transferred to young consumers rather than wasted), because at Bmax transfers are always zero regardless of ζ.&lt;/p&gt;
&lt;h3 id="q7-why-does-the-welfare-maximizing-sustainable-bond-capital-ratio-equal-bmax-rather-than-some-interior-value"&gt;Q7. Why does the welfare-maximizing sustainable bond-capital ratio equal Bmax rather than some interior value?&lt;/h3&gt;
&lt;p&gt;Proposition 6 shows that u′(B) ≥ 0 for all B ∈ [0, Bmax] whenever Rf ≤ 1, with strict inequality unless Rf = 1 and (1−ζ)B = 0. Since utility is weakly increasing throughout the feasible set, the optimum is the corner solution Bmax (Proposition 7). The mechanism: increasing B reduces k, lowering wages (bad for utility when young) but raising the marginal product of capital and hence the rates of return on capital and bonds (good for utility when old). The factor-price frontier ensures that the wage reduction equals the income gain accruing to initial capital, and the intertemporal optimality condition uy′(cy) = Rf E{uo′(co)} implies that when Rf ≤ 1 (so E{uo′(co)} ≥ uy′(cy)/Rf ≥ uy′(cy)), the welfare gain in old age dominates.&lt;/p&gt;
&lt;h3 id="q8-how-does-proposition-5-square-with-the-optimality-of-bmax-does-reducing-expected-consumption-not-reduce-welfare"&gt;Q8. How does Proposition 5 square with the optimality of Bmax? Does reducing expected consumption not reduce welfare?&lt;/h3&gt;
&lt;p&gt;Proposition 5 shows that when ζ = 1, a marginal increase in B at Bmax reduces expected aggregate consumption (dE{c}/dB &amp;lt; 0). However, welfare is not simply expected aggregate consumption: it also depends on the distribution of consumption across states. At Bmax, even though expected consumption falls, the increased risk sharing from holding more riskfree bonds — which smooth consumption between the high-return and low-return states of capital depreciation — is large enough to leave welfare unchanged (u′(Bmax) = 0 when ζ = 1) or to increase it (u′(Bmax) &amp;gt; 0 when ζ &amp;lt; 1). This illustrates that in stochastic economies, the welfare criterion diverges from the aggregate consumption criterion that characterizes dynamic inefficiency in deterministic economies.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-papers-welfare-analysis-relate-to-the-no-ponzi-game-npg-condition-and-the-fiscal-theory-of-the-price-level"&gt;Q9. How does the paper&amp;rsquo;s welfare analysis relate to the No Ponzi Game (NPG) condition and the fiscal theory of the price level?&lt;/h3&gt;
&lt;p&gt;The standard NPG condition requires that the value of government debt equals the present value of future primary surpluses. In the paper&amp;rsquo;s feasible-rollover region (rf ≤ g), the NPG condition is violated by design: the present value of future primary surpluses is non-positive (all primary balances are deficits or zero), yet the market value of outstanding bonds is strictly positive. This is possible because, as Santos and Woodford (1997) show, when the present value of aggregate consumption is infinite, the NPG can fail. The market value of the capital stock remains finite (it is the value of profits on a depreciating capital stock approaching zero), but the bubble value of government bonds is positive.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-quantitative-calibration-reveal-about-the-range-of-dynamically-efficient-permanently-rollable-bond-capital-ratios"&gt;Q10. What does the quantitative calibration reveal about the range of dynamically efficient, permanently rollable bond-capital ratios?&lt;/h3&gt;
&lt;p&gt;With α = 0.33, β = 0.353, g = 1% per year, G = 1.35, target mean equity return m = 3% per year, and risk aversion γ = 10 with annualized return standard deviation s = 0.22, the paper finds Bmax = 0.478 and B∗ = 0.083 (ζ = 0, Table 1). The interval [B∗, Bmax] = [0.083, 0.478] is the range of bond-capital ratios for which the economy is both dynamically efficient and able to roll over bonds permanently. For an economy with a capital-output ratio of 2, these bond-capital ratios correspond to debt-GDP ratios of up to 0.956. Both Bmax and B∗ are increasing in risk aversion γ and in the standard deviation of capital returns s; Bmax is independent of γ in any given column of the table for the ζ = 0 case (since R is independent of γ there), but rises substantially with γ in the ζ = 1 case.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-role-of-the-parameter-ζ-the-share-of-seignorage-transferred-vs-wasted"&gt;Q11. What is the role of the parameter ζ (the share of seignorage transferred vs. wasted)?&lt;/h3&gt;
&lt;p&gt;The parameter ζ governs what the government does with seignorage revenue: transfer it to young consumers (ζ = 1) or waste it (ζ = 0), or some mix. Corollary 1 shows that Bmax is completely invariant to ζ, because at Bmax, rf = g so seignorage (g − rf)Bt = 0 in any case. The value ζ does affect u′(Bmax): if ζ &amp;lt; 1, u′(Bmax) &amp;gt; 0; if ζ = 1, u′(Bmax) = 0. Both configurations yield Bmax as the welfare-maximizing level. The parameter ζ matters for welfare levels and for B∗ (only in the ζ = 1 case, where transfers are positive and boost saving capacity), but not for the main positive or normative results.&lt;/p&gt;
&lt;h3 id="q12-in-what-sense-is-the-model-tractable-and-what-are-its-key-limitations"&gt;Q12. In what sense is the model tractable, and what are its key limitations?&lt;/h3&gt;
&lt;p&gt;Tractability comes from three design choices: (i) the durability shock is additively separable from the production function, so labor income and aggregate saving are non-stochastic; (ii) IES = 1 with Epstein-Zin-Weil preferences, making saving a constant fraction of income; (iii) along balanced growth paths, g and rf are constant, so sustainability reduces to comparing two constants. Limitations acknowledged by the authors: the paper analyzes only balanced growth paths and does not characterize transition dynamics; the framework does not directly address economies where g or rf are stochastic; and the two-period OLG structure is stylized. The authors pose as an open question whether the result that optimal borrowing equals maximal borrowing generalizes to settings with random g.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Bond-capital ratio (B):&lt;/strong&gt; The ratio of outstanding government bonds to the capital stock, Bt/Kt. This is the paper&amp;rsquo;s central state variable and policy instrument. A value B is &amp;ldquo;sustainable&amp;rdquo; if the government can roll over its debt forever at the riskfree interest rate without any primary budget surpluses. The paper distinguishes B from the more commonly reported debt-GDP ratio (which equals B times the capital-output ratio).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Adjusted gross rate of return / riskfree rate (R, Rf):&lt;/strong&gt; R ≡ (1+r)/G and Rf ≡ (1+rf)/G, where r is the net return on capital, rf is the riskfree interest rate on bonds, and G = 1+g is the gross growth rate. Expressing returns in these &amp;ldquo;adjusted&amp;rdquo; gross units scales out balanced growth and simplifies the sustainability condition to Rf ≤ 1 (equivalently, rf ≤ g).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic efficiency (Zilcha criterion):&lt;/strong&gt; In the paper&amp;rsquo;s stochastic setting, the relevant criterion for dynamic efficiency is E{ln R} ≥ 0 (Zilcha 1991, as amended by Rangazas-Russell 2005 and Barbie-Kaul 2009), meaning the geometric mean of the adjusted gross return on capital is at least one. This differs from the deterministic condition r ≥ g. The paper&amp;rsquo;s Region E in Figure 1 is the key zone where E{ln R} &amp;gt; 0 (dynamically efficient) and Rf &amp;lt; 1 (rollover feasible) simultaneously.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bmax (maximum sustainable bond-capital ratio):&lt;/strong&gt; The largest value of B for which the bond-capital ratio is sustainable, defined as the unique root of Rf(B) = 1. At Bmax, the riskfree rate exactly equals the growth rate (rf = g). The paper proves Bmax is finite, invariant to ζ, and equals the welfare-maximizing sustainable bond-capital ratio.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;B∗ (dynamic efficiency threshold):&lt;/strong&gt; The bond-capital ratio at which the economy crosses from Zilcha-inefficiency into Zilcha-efficiency, defined by E{ln R} = 0. For B ∈ [B∗, Bmax], the economy is dynamically efficient and debt rollover is feasible. B∗ &amp;lt; Bmax when risk aversion γ or return volatility s is large enough, defining a non-trivial interval of dynamically efficient, permanently rollable bond levels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Durability shock (ε):&lt;/strong&gt; An i.i.d. random variable with mean zero that enters the capital depreciation rate as δ − ε_t. This shock makes the rate of return on capital r = αkα−1 − δ + ε stochastic while leaving the capital stock per unit of effective labor, aggregate wages, and aggregate saving non-stochastic. It is the only source of aggregate uncertainty in the model and is the mechanism that drives a wedge between rf and E{r}.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;No Ponzi Game (NPG) condition:&lt;/strong&gt; The condition that the present discounted value of government debt converges to zero (equivalently, debt equals the present value of future primary surpluses). Standard fiscal sustainability analyses assume this condition holds. The paper explicitly violates it: in the feasible-rollover region rf ≤ g, the present value of aggregate consumption is infinite and the NPG fails, yet government bond values are positive and debt rollover is sustainable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Seignorage (ζ):&lt;/strong&gt; The revenue the government obtains by issuing new bonds in excess of interest payments on existing bonds, equal to (g − rf)Bt when rf &amp;lt; g. The parameter ζ ∈ [0,1] governs the share transferred to young consumers (as lump-sum transfers τt) versus wasted (captured by the government but yielding no utility). A key finding is that Bmax is invariant to ζ, since seignorage is zero at rf = g regardless of ζ.&lt;/p&gt;</description></item><item><title>School Choice and the Housing Market</title><link>https://macropaperwarehouse.com/papers/school-choice-and-the-housing-market/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/school-choice-and-the-housing-market/</guid><description>&lt;p&gt;Grigoryan (2021) develops a unified general-equilibrium framework that jointly models school assignment mechanisms and the housing market to evaluate the welfare and distributional consequences of replacing traditional neighborhood assignment (NA) with the Deferred Acceptance (DA) mechanism. The paper fills a gap in the matching theory literature, where preferences and priorities are typically treated as exogenous, by making residential choices endogenous: families first observe which school assignment mechanism the district announces, then optimally select a neighborhood given market-clearing prices and other families&amp;rsquo; choices, and finally children are assigned to schools through the announced mechanism.&lt;/p&gt;
&lt;p&gt;The model features a continuum of families, each with a type defined by valuations over all neighborhood–school pairs, a finite set of neighborhoods and schools (one school per neighborhood), and competitive equilibrium prices. Three mechanisms are compared: NA (each child attends the neighborhood school), DA without neighborhood priority (DA), and DA with neighborhood priority (DN), where neighborhood residents receive priority at their local school.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s first major result (Theorem 3) is that DN unambiguously generates weakly higher aggregate welfare than NA. The proof exploits the fact that DN preserves NA&amp;rsquo;s option — families can still guarantee admission to the neighborhood school by living there — while additionally allowing families to access seats at other schools that go unclaimed by neighborhood residents. Although price effects under DN can make some individual families worse off relative to NA, aggregate welfare (inclusive of house sellers) is always weakly higher under DN. In simulations with 1,000 students, 10 neighborhoods, and 10 schools, DN yields average aggregate welfare gains of 2.40% relative to NA across the 18 parameter configurations studied.&lt;/p&gt;
&lt;p&gt;The welfare comparison between DA (without neighborhood priority) and NA is ambiguous in the general model: simulations show DA producing gains as large as +5.65% and losses as large as −18.26% relative to NA, depending on the degree of preference alignment across families (parameter α) and the variance in school capacities (parameter γ). DN also dominates DA in aggregate welfare under two sufficient conditions — identical ordinal preference rankings over neighborhoods and schools (Assumption 1 or 2) — though counterexamples exist when these assumptions fail.&lt;/p&gt;
&lt;p&gt;The second major result (Theorem 5, Corollaries 1–2) concerns the welfare of lowest-income families, defined as those with budget (maximum willingness to pay for housing) equal to zero or sufficiently close to zero. Under two jointly sufficient conditions — (1) neighborhoods that are underdemanded (zero-priced) under NA remain underdemanded under DA/DN, and (2) the schools in those underdemanded neighborhoods are themselves underdemanded — both DA and DN generate weakly higher welfare for the lowest-income families than NA. These conditions hold whenever families share common ordinal preference rankings (Corollary 1) and in the uniform economy where each valuation profile is equally likely (Corollary 2). The conditions are shown to be approximately necessary in a robustness sense (Theorem 6): for any economy violating them, an arbitrarily close economy exists in which a positive measure of zero-income families prefer NA. In simulations, DN raises lowest-income welfare by an average of 26.51% and DA by an average of 38.25% relative to NA.&lt;/p&gt;
&lt;p&gt;The paper also proves existence of a competitive equilibrium for the continuum economy under DA and DN via the Schauder-Tychonoff fixed-point theorem (Theorem 2), exploiting the continuity of school assignment probabilities in families&amp;rsquo; neighborhood choices. In discrete economies, assignment externalities can preclude equilibrium existence, but approximate equilibria exist in sufficiently large discrete markets and all welfare comparisons carry over approximately. The existence proof technique applies to general assignment games with externalities including peer preferences and complementarities.&lt;/p&gt;
&lt;p&gt;Scope conditions: results are derived for a model without direct peer externalities or endogenous school quality; a supplementary extension to local public financing finds that the aggregate welfare superiority of DA over NA may not survive when school spending is capitalized into housing prices, though the lowest-income welfare sufficiency conditions of Theorem 5 do extend to that environment.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question and why does the housing market matter for evaluating school choice?&lt;/p&gt;
&lt;p&gt;A: The paper asks how replacing neighborhood assignment with the Deferred Acceptance mechanism affects aggregate welfare and the welfare of the lowest-income families, accounting for the fact that families choose where to live in response to the school assignment mechanism. The housing market matters because under neighborhood assignment families can guarantee enrollment at a preferred school by purchasing a house in that school&amp;rsquo;s neighborhood; switching to DA changes these strategic incentives, alters equilibrium prices, and therefore changes who ends up in which neighborhood before any school assignment takes place. Ignoring residential choices would miss this feedback loop between assignment rules and housing demand.&lt;/p&gt;
&lt;p&gt;Q: What are the three mechanisms compared, and how do they differ?&lt;/p&gt;
&lt;p&gt;A: Neighborhood assignment (NA) assigns each child to the school in their neighborhood with certainty. DA without neighborhood priority allocates seats by student preference rankings and lottery numbers, with market-clearing cutoffs determined iteratively; no residential location confers a priority advantage. DN (DA with neighborhood priority) works like DA but grants neighborhood residents a priority of 1 at their local school and 0 at all other schools, effectively guaranteeing neighborhood families a seat at their local school while filling remaining seats by lottery among non-neighborhood applicants.&lt;/p&gt;
&lt;p&gt;Q: What does Theorem 3 establish, and what is the intuition for why DN dominates NA in aggregate welfare?&lt;/p&gt;
&lt;p&gt;A: Theorem 3 establishes that for any competitive equilibrium under DN and any competitive equilibrium under NA, aggregate welfare is weakly higher under DN. The intuition is that DN preserves all options available under NA — a family can always choose the neighborhood corresponding to its most-valued school and be guaranteed admission there — while additionally providing access to seats at other schools not claimed by their own neighborhood residents. The proof maps DN&amp;rsquo;s CE onto a Walrasian equilibrium of a continuum assignment game and invokes the welfare-maximization property of such equilibria from Gretsky, Ostroy, and Zame (1992).&lt;/p&gt;
&lt;p&gt;Q: Why is the welfare comparison between DA and NA ambiguous?&lt;/p&gt;
&lt;p&gt;A: Under NA, families with the highest cardinal valuations for a particular school can guarantee admission by purchasing a house in that neighborhood, and this targeted sorting can raise aggregate welfare when preferences over schools are strongly aligned. Under DA (without neighborhood priority), no location guarantees school admission, so families lose this signaling device; but DA allows families to live in preferred neighborhoods without sacrificing school quality, which raises welfare when preferences are heterogeneous. Neither effect dominates in general: in simulations, DA ranges from −18.26% to +5.65% relative to NA across the parameter space.&lt;/p&gt;
&lt;p&gt;Q: What role do neighborhood priorities play as a &amp;ldquo;signaling device,&amp;rdquo; and when does DN dominate DA?&lt;/p&gt;
&lt;p&gt;A: Neighborhood priorities allow families to credibly signal high valuations for a school by choosing to live in that school&amp;rsquo;s neighborhood, analogously to signaling devices in matching markets without money. When families have identical ordinal preference rankings over neighborhoods and schools (Assumptions 1 or 2), DN generates weakly higher aggregate welfare than DA because any DA assignment probability can be replicated under DN by mixing over neighborhoods, but the converse is not true. Counterexamples exist when preference rankings differ across families, so the DN-over-DA dominance is not universal.&lt;/p&gt;
&lt;p&gt;Q: What are the sufficient conditions for lowest-income families to prefer DA/DN to NA, and how tight are they?&lt;/p&gt;
&lt;p&gt;A: The two joint conditions are: (1) neighborhoods that have zero price (are underdemanded) under NA also have zero price under DA or DN after the mechanism switch; and (2) the schools located in those underdemanded neighborhoods are themselves underdemanded (have zero admission cutoffs) under DA/DN. Condition (1) reflects that the poorest neighborhoods are unlikely to become highly sought-after merely because the assignment mechanism changed. Condition (2) is consistent with the empirical finding of Owens and Candipan (2019) that in large US metropolitan areas the poorest neighborhoods typically have underperforming schools. Theorem 6 shows these conditions are approximately necessary: any economy violating them is arbitrarily close to one where a positive measure of zero-budget families prefer NA, so robustness requires them.&lt;/p&gt;
&lt;p&gt;Q: What do the simulations show about the magnitude of welfare effects for lowest-income families?&lt;/p&gt;
&lt;p&gt;A: In simulations with 10 lowest-income families (budgets of 0.05) among 1,000 total, DN raises lowest-income welfare by an average of 26.51% relative to NA and DA raises it by an average of 38.25% relative to NA, across the 18 parameter configurations. The gains are larger when preferences for neighborhoods and schools are less correlated (lower α) and when school capacities are more uniform (higher γ). DA consistently outperforms DN for lowest-income families in the simulations, even though DN dominates NA in aggregate welfare more reliably.&lt;/p&gt;
&lt;p&gt;Q: How does the paper handle equilibrium existence given the externalities created by residential choices?&lt;/p&gt;
&lt;p&gt;A: Because a family&amp;rsquo;s expected utility from a neighborhood depends on other families&amp;rsquo; neighborhood choices (through their effect on school assignment probabilities), standard existence results for assignment games do not directly apply. For the continuum economy, the author proves that school assignment probabilities under DA/DN are equicontinuous in families&amp;rsquo; neighborhood choices, which enables application of the Schauder-Tychonoff fixed-point theorem to guarantee the existence of a competitive equilibrium (Theorem 2). In finite discrete economies, assignment externalities can prevent equilibrium existence (illustrated by an example in Appendix B), but approximate equilibria exist for sufficiently large discrete markets, and all welfare comparisons hold approximately.&lt;/p&gt;
&lt;p&gt;Q: How does the paper&amp;rsquo;s model relate to and extend prior theoretical work on school choice and welfare?&lt;/p&gt;
&lt;p&gt;A: Prior theoretical work (e.g., Calsamiglia et al. 2015; Xu 2019; Avery and Pathak 2020) uses stylized models with single-parameter family types, identical ordinal school rankings, supermodular valuations, and no preferences over neighborhoods. This paper allows an unrestricted preference domain — families have arbitrary valuations over all neighborhood–school pairs — which generates novel findings: in the general model, lowest-income families do not necessarily benefit from DA (contrary to Calsamiglia et al. and Xu), aggregate welfare comparisons between DA and NA are ambiguous (whereas they are trivially resolved in the special cases of prior work), and neighborhood priorities can be welfare-improving even relative to DA without priorities.&lt;/p&gt;
&lt;p&gt;Q: Does the paper address the extension to endogenous school quality or local public financing?&lt;/p&gt;
&lt;p&gt;A: In Supplementary Appendix B, the model is extended to allow school spending to be financed by local property taxes, making school quality endogenous to neighborhood housing values. In that environment, the aggregate welfare superiority of DA/DN over NA may not hold: DA attracts non-neighborhood applicants to high-priced neighborhoods, and if those schools are a poor match for those applicants absent the spending, social welfare may fall — a result analogous to Barseghyan et al. (2013). However, the paper reports that the sufficiency conditions for lowest-income family welfare comparisons (Theorem 5) do extend to the local public financing environment, preserving the distributional results.&lt;/p&gt;
&lt;p&gt;Q: What does the paper say about alternative mechanisms such as Immediate Acceptance (Boston mechanism) and Top Trading Cycles?&lt;/p&gt;
&lt;p&gt;A: The Supplementary Appendix studies these alternatives. For Immediate Acceptance (IA), the paper shows that when there are neighborhood priorities, lowest-income families may prefer DA to IA, echoing the finding that IA is not strategyproof and may disproportionately hurt low-income families who are worse at gaming the system or have worse outside options (Pathak and Sonmez 2008; Calsamiglia et al. 2015). Top Trading Cycles and further extensions are also analyzed in the Supplementary Appendix, though detailed results are not developed in the main text.&lt;/p&gt;
&lt;p&gt;Neighborhood Assignment (NA): The baseline mechanism under which each family&amp;rsquo;s child is automatically enrolled in the school located in their chosen residential neighborhood, with no option to attend schools outside that neighborhood.&lt;/p&gt;
&lt;p&gt;Deferred Acceptance without Neighborhood Priority (DA): A strategyproof centralized assignment mechanism in which seats are allocated by families&amp;rsquo; stated preference rankings and lottery numbers via market-clearing admission cutoffs; residential location confers no priority advantage at any school.&lt;/p&gt;
&lt;p&gt;Deferred Acceptance with Neighborhood Priority (DN): A version of DA in which families residing in a neighborhood receive priority 1 at their neighborhood school and priority 0 at all other schools, guaranteeing neighborhood residents a seat at their local school before remaining seats are allocated by lottery to non-neighborhood applicants.&lt;/p&gt;
&lt;p&gt;Competitive Equilibrium (CE): A pair of neighborhood choices and a price vector such that (1) each family optimally selects the neighborhood maximizing expected utility net of price (subject to budget), (2) neighborhood capacities are not exceeded, and (3) neighborhoods with excess capacity are priced at zero.&lt;/p&gt;
&lt;p&gt;Underdemanded Neighborhood/School: A neighborhood whose equilibrium price is zero (excess housing supply) or a school whose admission cutoff is zero (excess capacity), meaning any applicant who lists it can gain admission.&lt;/p&gt;
&lt;p&gt;Assignment Externality: The indirect dependence of a family&amp;rsquo;s expected utility on other families&amp;rsquo; neighborhood choices, which operates through the effect of the population distribution across neighborhoods on the family&amp;rsquo;s school assignment probabilities under DA or DN. This externality can preclude competitive equilibrium existence in discrete economies.&lt;/p&gt;
&lt;p&gt;Aggregate Welfare: The utilitarian sum of all families&amp;rsquo; expected utilities from their neighborhood–school assignments, not netting out neighborhood prices (so it includes the welfare of house sellers as passive agents); the comparison criterion for Theorems 3 and 4.&lt;/p&gt;
&lt;p&gt;Signaling Device (neighborhood priority as): The interpretation that neighborhood priorities allow families to credibly reveal high valuations for a school by choosing to live in that school&amp;rsquo;s neighborhood, analogously to signaling instruments in matching markets without monetary transfers; the mechanism through which DN can improve welfare relative to DA.&lt;/p&gt;</description></item><item><title>Search Frictions and Product Design in the Municipal Bond Market</title><link>https://macropaperwarehouse.com/papers/search-frictions-and-product-design-in-the-municipal-bond-market/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/search-frictions-and-product-design-in-the-municipal-bond-market/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper investigates whether intermediaries in the U.S. municipal bond market strategically exploit product design to increase search frictions and, through that channel, capture rents. Specifically, it asks: do underwriters who negotiate bond design with local governments have an incentive to add nonstandard provisions that raise their own competitive advantage in subsequent secondary-market intermediation, even at the expense of issuing governments and their taxpayers?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and Data&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The study focuses on tax-exempt general obligation and revenue bonds issued via negotiated sales by local governments (counties, cities, school districts, and other special-purpose governments) from 2010 to 2013, tracking all secondary-market transactions through 2014. The final sample comprises 13,118 bond issues with a total face value of $266.9 billion. Bond attribute data come from Mergent; transaction data come from the Municipal Securities Rulemaking Board (MSRB). Issuer financial health, demographics, and economic conditions are drawn from the Census and American Community Survey; state revolving-door regulations are compiled from the National Conference of State Legislatures database. Structural estimation uses a subsample of 927 bonds concentrated in the five states that enacted revolving-door regulations during the study period and neighboring border counties.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identification Strategy&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A core empirical challenge is that unobserved factors may jointly determine bond complexity and market outcomes. The authors exploit panel variation in state-level revolving-door regulations — laws that restrict former public officials from taking employment at firms regulated by their former agencies for a &amp;ldquo;cool-off&amp;rdquo; period — as an instrument for bond complexity. Between 2010 and 2013, three states (Arkansas 2011, Indiana 2010, Maine 2013) enacted new legislation covering state officials, and two states (New Mexico 2011, Virginia 2011) extended existing regulations to cover local officials. A difference-in-differences regression, with county and year-month fixed effects, shows that adopting revolving-door regulations covering local officials reduces bond complexity by 6% on average (coefficient −0.064, p &amp;lt; 0.01). Regulations targeting only state officials, who are not directly involved in bond negotiations, yield smaller and statistically fragile effects. Placebo checks on auctioned bonds, where underwriters cannot influence design, show no effect, and there is no evidence of pre-existing trends in complexity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Flexibility vs. liquidity trade-off&lt;/strong&gt;: A 1% increase in the bond complexity index lowers the number of negative credit-watch events (a proxy for default risk) by 0.002, a 3% decrease relative to the mean of 0.074, confirming that nonstandard provisions provide genuine financial flexibility. However, increasing the complexity index from its mean (1.46) to the 75th percentile (1.69) raises the intermediation spread — the cost for an investor to buy and immediately sell a bond — by 17 basis points (a 14% increase over the average of 120 basis points), confirming that complexity raises trading frictions. For context, the average intermediation spread of 120 basis points is large relative to the 30–60 basis point bid-ask spread of corporate bonds in 2010–2013.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Underwriter incentive to complicate&lt;/strong&gt;: Increasing complexity from the mean to the 75th percentile raises the underwriter&amp;rsquo;s market share in secondary-market intermediation by 1.4 percentage points, an 11% increase over the average underwriter share of 12.2%. The underwriter&amp;rsquo;s gross profits from intermediation also increase with complexity.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Structural estimates — search costs&lt;/strong&gt;: For a median bond, average dealer search costs amount to 10% of monthly gross profits ($2,625 per month). The underwriter&amp;rsquo;s exclusive initial sales generate a client network that lowers its effective search costs by 21% relative to an average dealer, more than offsetting its initial geographical disadvantage (for 72% of bonds, the underwriter&amp;rsquo;s baseline search cost exceeds the median dealer&amp;rsquo;s). Nonstandard provisions increase both the initial search cost parameter (φ₀) and the network-effect parameter (φ₁): a 1% increase in the complexity index increases φ₀ by 3.79% and φ₁ by 1.66%, implying complex bonds raise search costs broadly but amplify the advantage of a large client network — a position the underwriter occupies via exclusive primary-market sales.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Investor demand&lt;/strong&gt;: Nonstandard provisions do not substantially change the average investor valuation but substantially increase the dispersion: the standard deviation of investor valuations is 0.003 for simple bonds and 0.013 for complex bonds, consistent with complex bonds being niche products that investors &amp;ldquo;either love or loathe.&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Government cost&lt;/strong&gt;: The marginal cost of paying debt obligations is convex in complexity, reaching a minimum at an interior level of provisions; the government&amp;rsquo;s marginal financial cost increases by 42% when a median bond is stripped of all nonstandard provisions, reflecting the value of payment flexibility.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Conflict of interest&lt;/strong&gt;: The estimated weight that government officials place on underwriter payoffs in the absence of revolving-door regulations (ψ₀) is 0.34, implying the underwriter&amp;rsquo;s value accounts for 6.7% of the government official&amp;rsquo;s payoff under the median unregulated issuer. With revolving-door regulations in place, ψ₁ is essentially zero.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Counterfactual Policies (on representative bond: face value $6.45 million, maturity 7.7 years)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Standardization mandate&lt;/strong&gt; (ban on all nonstandard provisions): The coupon rate falls from 2.81% to 2.16% (−23%), average dealer search costs fall 47%, and investor surplus rises 13.3%. However, the marginal financial cost (c₀) rises by 41% (from 0.615 to 0.871), so the issuer&amp;rsquo;s total debt payment cost — principal plus interest, weighted by c₀ — rises by 35%, from $5.13 million to $6.96 million. The standardization policy harms issuers even while saving 7.8% of raw principal-and-interest payments ($8,349K to $7,997K), because the loss of flexibility more than offsets the liquidity gain.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Issuer-driven design&lt;/strong&gt; (issuer sets complexity to minimize its own debt payment cost, then negotiates the coupon): Complexity falls 19% to 1.14, the interest rate falls to 2.37%, total issuer cost falls 1.5%, investor surplus rises 6%, and the underwriter&amp;rsquo;s secondary-market payoff falls 19.9%.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Underwriter intermediation ban&lt;/strong&gt; (underwriter excluded from trading after six months): Complexity falls 5.7% to 1.33, the coupon falls to 2.59%, issuer cost falls 1.5%, but investor surplus falls 1.84% and even other dealers are worse off by 3.97%, because the underwriter&amp;rsquo;s information on primary-market buyers is lost, offsetting the liquidity gains from lower complexity.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-five-nonstandard-bond-features-tracked-as-proxies-for-complexity-and-how-are-they-combined-into-a-single-index"&gt;Q1. What are the five nonstandard bond features tracked as proxies for complexity, and how are they combined into a single index?&lt;/h3&gt;
&lt;p&gt;Following Harris and Piwowar (2006), the paper focuses on five features that are particularly difficult for investors to price: (i) multiple or serial bonds per issue (as opposed to a single bond), (ii) call provisions allowing early redemption, (iii) sinking fund provisions requiring periodic debt retirement, (iv) nonstandard interest payment frequencies (other than semiannual), and (v) variable or floating interest rates. The complexity index is constructed as the simple average of the latter four provisions across bonds within an issue, plus a dummy for whether the issue contains multiple bonds.&lt;/p&gt;
&lt;h3 id="q2-why-do-revolving-door-regulations-that-target-local-officials-reduce-complexity-more-than-those-targeting-state-officials"&gt;Q2. Why do revolving-door regulations that target local officials reduce complexity more than those targeting state officials?&lt;/h3&gt;
&lt;p&gt;State officials are not directly involved in bond origination negotiations — they can only indirectly influence local governments through budget allocations. Local officials negotiate directly with underwriters and are thus the proximate counterparties whose incentives the regulations alter. Accordingly, revolving-door regulations covering local officials reduce complexity by 6% (coefficient −0.064, p &amp;lt; 0.01 with full controls), whereas regulations targeting only state officials produce a smaller effect (approximately 2%) that loses statistical significance once issuer financial health controls are added.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-paper-validate-that-revolving-door-regulations-are-a-valid-instrument-for-bond-complexity"&gt;Q3. How does the paper validate that revolving-door regulations are a valid instrument for bond complexity?&lt;/h3&gt;
&lt;p&gt;The paper provides three pieces of evidence. First, the regulations have no effect on the credit ratings of bonds issued prior to their enactment, on the annual amount of bond issuance, or on the maturity length and sale method conditional on issuance — confirming the regulations do not alter governments&amp;rsquo; risk management or underlying financing needs. Second, the regulations have no effect on complexity for competitively auctioned bonds, where underwriters cannot influence design — a direct placebo test. Third, a pre-trend analysis (Figure A1) finds no differential trend in complexity in states that subsequently adopted regulations.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-mechanism-by-which-underwriters-benefit-from-adding-nonstandard-provisions-and-why-does-this-advantage-not-diminish-over-time"&gt;Q4. What is the mechanism by which underwriters benefit from adding nonstandard provisions, and why does this advantage not diminish over time?&lt;/h3&gt;
&lt;p&gt;Underwriters purchase and distribute the entire bond issue at origination, giving them an exclusive network of investors who initially purchased the bonds. In the secondary market, knowing who owns a bond allows the underwriter to locate buyers and sellers with lower search effort. For complex bonds, this advantage is amplified: nonstandard provisions make investor education and persuasion more costly, increasing the value of pre-existing client relationships. The network-effect parameter φ₁ — which governs how rapidly search costs fall as a dealer&amp;rsquo;s cumulative trades grow — itself rises with complexity (by 1.66% per 1% increase in the complexity index), so the underwriter&amp;rsquo;s head start in client network accumulation translates into a persistently larger cost advantage precisely for the most complex bonds.&lt;/p&gt;
&lt;h3 id="q5-how-large-is-the-underwriters-search-cost-advantage-in-equilibrium-and-what-drives-it"&gt;Q5. How large is the underwriter&amp;rsquo;s search cost advantage in equilibrium, and what drives it?&lt;/h3&gt;
&lt;p&gt;At the equilibrium meeting rate, the underwriter&amp;rsquo;s effective search cost of maintaining a given meeting rate is 21% lower than that of an average dealer. This advantage arises despite the underwriter having a higher initial search cost type (φ₀ of $3,609 vs. $3,216 for the average dealer at λ = 1), because for 72% of bonds the underwriter has less local trading experience than the median dealer. The advantage is entirely driven by the underwriter&amp;rsquo;s network: its exp(−φ₁ log(b)) cost discount factor averages 0.34, 32% lower than the average dealer&amp;rsquo;s 0.50. The underwriter meets investors 20% more frequently than the average dealer (0.23 vs. 0.19 per month), despite higher absolute search expenditures ($3,045 vs. $2,625 per month).&lt;/p&gt;
&lt;h3 id="q6-how-does-bond-complexity-affect-investor-demand--mean-or-dispersion-of-valuations"&gt;Q6. How does bond complexity affect investor demand — mean or dispersion of valuations?&lt;/h3&gt;
&lt;p&gt;Structural estimates show that increasing the complexity index by 1% increases the standard deviation of investor valuations (γ₂) by 4.60% but has no statistically significant effect on the mean valuation (coefficient −0.085, standard error 0.561). This pattern is consistent with complex bonds being niche products — they attract a subset of investors with specific preferences for the embedded features (e.g., certain tax or cash-flow attributes), while being unappealing to most investors. The standard deviation of valuations is 0.003 for a low-complexity bond (25th percentile) and 0.013 for a high-complexity bond (75th percentile).&lt;/p&gt;
&lt;h3 id="q7-what-does-the-structural-estimate-of-ψ-imply-about-the-degree-of-collusion-between-government-officials-and-underwriters"&gt;Q7. What does the structural estimate of ψ₀ imply about the degree of collusion between government officials and underwriters?&lt;/h3&gt;
&lt;p&gt;The estimated collusion parameter without revolving-door regulations (ψ₀ = 0.34) implies that, for the median unregulated issuing government, the underwriter&amp;rsquo;s value from secondary-market trading accounts for 6.7% of the government official&amp;rsquo;s objective function. This is a substantial weight: it means officials act partly as agents for the underwriter rather than purely for taxpayers. With revolving-door regulations (ψ₁ ≈ 0), this collusive weight is essentially eliminated, explaining the empirical reduction in complexity found in Table 2.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-effects-of-a-full-standardization-mandate-on-each-class-of-market-participant-and-why-does-the-issuer-lose-overall-despite-paying-a-lower-coupon"&gt;Q8. What are the effects of a full standardization mandate on each class of market participant, and why does the issuer lose overall despite paying a lower coupon?&lt;/h3&gt;
&lt;p&gt;Under standardization, the coupon falls 23% (from 2.81% to 2.16%) and the raw principal-plus-interest payment falls 7.8% (from $8,349K to $7,997K). However, the marginal financial cost c₀ rises 41% (from 0.615 to 0.871), reflecting the loss of payment flexibility previously provided by call provisions and other features; the total issuer cost — c₀A(1 + rT) — rises by 35% (from $5.13 million to $6.96 million). Investors gain 13.3% in surplus because they value liquidity and, on average, do not value nonstandard features. The underwriter loses 36.6% of its secondary-market value while other dealers gain 36.1%, as standardization erodes the underwriter&amp;rsquo;s network advantage.&lt;/p&gt;
&lt;h3 id="q9-why-does-the-issuer-driven-design-scenario-outperform-standardization-in-terms-of-total-issuer-cost-even-though-complexity-does-not-fall-to-zero"&gt;Q9. Why does the issuer-driven design scenario outperform standardization in terms of total issuer cost, even though complexity does not fall to zero?&lt;/h3&gt;
&lt;p&gt;Under issuer-driven design, the government minimizes its total cost of debt payment c₀A(1 + rT), accounting for both the flexibility value of provisions and their effect on the negotiated coupon. The optimal complexity index is 1.14 — positive, but 19% below the current baseline of 1.41 — because some provisions genuinely lower c₀ by allowing flexible debt service. The cost of search frictions (and hence the liquidity premium embedded in the coupon) falls 32% and the negotiated coupon falls to 2.37%, sufficient to reduce total issuer cost by 1.5%. By contrast, full standardization imposes a complexity of zero, which overshoots: c₀ rises more than the coupon savings compensate, increasing total costs by 35%.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-net-welfare-effects-of-the-underwriter-intermediation-ban-and-why-is-investor-surplus-negative-despite-lower-complexity"&gt;Q10. What are the net welfare effects of the underwriter intermediation ban, and why is investor surplus negative despite lower complexity?&lt;/h3&gt;
&lt;p&gt;The ban reduces complexity by 5.7%, lowering the coupon to 2.59% and reducing issuer costs by 1.5%. However, the underwriter&amp;rsquo;s client network — built during exclusive initial sales — is a productive resource that improves match quality in the secondary market; banning the underwriter from trading after six months wastes this information. Average dealer search costs rise 1.2% and the meeting rate falls 1.7%, net of the complexity reduction. Investors face bonds with lower coupons and higher effective search frictions, so their surplus falls 1.84%. Non-underwriter dealers also lose 3.97% because lower coupons reduce the rents extractable from intermediation.&lt;/p&gt;
&lt;h3 id="q11-how-is-the-structural-model-estimated-and-what-role-do-revolving-door-regulations-play-in-the-estimation"&gt;Q11. How is the structural model estimated, and what role do revolving-door regulations play in the estimation?&lt;/h3&gt;
&lt;p&gt;Estimation proceeds in three steps. In Step 1, bond-specific trading market parameters (investor demand, dealer search costs, meeting rates, bargaining parameters) are recovered separately for each bond by minimizing squared differences between observed and simulated trading prices, quantities, and transaction timing. In Step 2, IV regressions using revolving-door regulations and their interactions with county/state attributes as instruments for endogenous complexity map Step 1 parameters to bond attributes, addressing the endogeneity of complexity in determining search costs and investor demand. In Step 3, GMM moment conditions derived from Nash bargaining first-order conditions for the equilibrium complexity and coupon rate identify government preference parameters (θ_c, ψ₀, ψ₁), using the orthogonality condition that unobserved financing cost shocks are mean-zero conditional on observed attributes, regulations, and bond supply from neighboring counties.&lt;/p&gt;
&lt;h3 id="q12-does-the-underwriting-market-show-signs-of-concentration-that-might-amplify-the-conflict-of-interest-problem"&gt;Q12. Does the underwriting market show signs of concentration that might amplify the conflict-of-interest problem?&lt;/h3&gt;
&lt;p&gt;Yes. The mean state-level Herfindahl-Hirschman Index (HHI) for underwriting is 0.12, with the top three firms covering 45% of the market on average. For smaller deals (under $10 million), concentration is markedly higher: mean HHI of 0.24 and top three firms covering 64% of the market. Repeat relationships are common — 41% of bonds issued in 2011–2017 were underwritten by a firm that had underwritten a prior bond for the same issuer within five years — reflecting both informational advantages of local presence and potentially entrenched relationships that may increase government officials&amp;rsquo; susceptibility to underwriter influence.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Complexity index (nonstandard provisions)&lt;/strong&gt;: A bond-level measure computed as the simple average, across bonds within an issue, of four nonstandard features — call provisions, sinking fund provisions, nonstandard interest payment frequency, and variable/floating interest rates — plus a dummy for whether the issue contains multiple bonds. Used as the primary measure of bond complexity in all regressions and the structural model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Revolving-door regulation&lt;/strong&gt;: A state-level law restricting former public officials or employees from engaging in lobbying or taking employment at regulated firms for a specified &amp;ldquo;cool-off&amp;rdquo; period (typically one to two years) after leaving office. The paper uses the presence and scope of such regulations (whether they cover state officials, local officials, or both) as a source of exogenous variation in government officials&amp;rsquo; incentives to align with underwriter interests.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intermediation spread&lt;/strong&gt;: The logarithm of the average dealer-to-investor sale price minus the logarithm of the average dealer-from-investor purchase price for a given bond. Used as the empirical measure of trading frictions; the sample average is 120 basis points.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Network effect in search (φ₁)&lt;/strong&gt;: The parameter governing how a dealer&amp;rsquo;s cumulative prior trades with investors in a given bond reduce its cost of meeting new investors for that bond. A higher φ₁ means a larger client network translates into steeper cost savings. The paper estimates that φ₁ itself increases with bond complexity, so complex bonds amplify the advantage of dealers (especially the underwriter) who accumulate large client networks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marginal cost of debt payment (c₀)&lt;/strong&gt;: A bond- and issuer-specific parameter capturing the effective cost to the government of repaying each dollar of principal and interest, net of the flexibility benefits provided by nonstandard provisions. Normalized to one for a bond with zero nonstandard provisions at average issuer characteristics; estimated to be convex in complexity with an interior minimum, implying some nonstandard provisions are beneficial from the government&amp;rsquo;s perspective.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Collusion weight (ψ)&lt;/strong&gt;: The weight a government official places on the underwriter&amp;rsquo;s secondary-market value from trading when negotiating bond design. Estimated at ψ₀ = 0.34 in the absence of revolving-door regulations (implying the underwriter&amp;rsquo;s interest accounts for 6.7% of the official&amp;rsquo;s objective) and at ψ₁ ≈ 0 when such regulations are present.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Underwriter dual role&lt;/strong&gt;: The institutional arrangement in which the same investment bank (i) negotiates and purchases the entire bond from the issuing government at origination, and (ii) subsequently acts as a dealer in the bond&amp;rsquo;s secondary market. This dual role creates an incentive to design complex bonds that strengthen the underwriter&amp;rsquo;s competitive advantage in secondary intermediation via network effects in search.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Issuer-driven design&lt;/strong&gt;: A counterfactual policy scenario in which the government sets the complexity level to minimize its total cost of debt payment — accounting for both the flexibility value of provisions and the anticipated effect on the negotiated coupon rate — before bargaining with the underwriter only over the coupon. This policy allows some nonstandard provisions (complexity index 1.14 vs. baseline 1.41) and reduces total issuer cost by 1.5% relative to the baseline.&lt;/p&gt;</description></item><item><title>Selection in Surveys: Using Randomized Incentives to Detect and Account for Nonresponse Bias</title><link>https://macropaperwarehouse.com/papers/selection-in-surveys-using-randomized-incentives-to-detect-and-account-for-nonresponse-bias/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/selection-in-surveys-using-randomized-incentives-to-detect-and-account-for-nonresponse-bias/</guid><description>&lt;p&gt;This paper addresses nonresponse bias in surveys — the distortion that arises when survey participants differ systematically from nonparticipants in ways that correlate with the survey&amp;rsquo;s outcomes of interest. The authors develop and apply methods to detect and correct for nonresponse bias using randomized financial incentives embedded in the survey design itself.&lt;/p&gt;
&lt;p&gt;The empirical application is the &amp;ldquo;Norge i Koronatid&amp;rdquo; (NiK) survey, conducted by Statistics Norway in April–May 2020 to study the immediate labor market consequences of Norway&amp;rsquo;s COVID-19 lockdown. The NiK survey has two features that make it unusually well-suited for studying nonresponse bias: (1) it is linked to full-population administrative data, providing a verifiable ground truth for the entire Norwegian adult population; and (2) survey invitees were randomly assigned to one of five financial incentive levels (0%, 1%, 5%, 7%, or 10% probability of receiving a 1,000 NOK prepaid card), generating exogenous variation in participation rates. The final sample of 10,000 randomly drawn adults achieved a 47.4% participation rate.&lt;/p&gt;
&lt;p&gt;The administrative data reveal large, statistically significant nonresponse bias across all six labor market outcomes examined. Participants in the high-incentive arm had on average roughly 930 USD (30%) higher monthly pre-lockdown earnings than the full population, and were 10.8 percentage points (19%) more likely to be employed. Standard corrections for selection on observable characteristics — including propensity-score reweighting on age, gender, immigration status, schooling, and municipality-level variables — fail to eliminate this bias. For the high-incentive arm, reweighting on individual characteristics more than doubles the nonresponse bias for earnings loss and employment loss measures relative to unweighted estimates, meaning that observable-based corrections can make things worse, not better.&lt;/p&gt;
&lt;p&gt;A key finding is that higher participation rates do not imply lower nonresponse bias. The high-incentive arm, with the highest response rate, exhibited larger nonresponse bias than the no-incentive arm. Marginal participants — those induced to respond by higher incentives — had much stronger pre-lockdown labor market attachment (average earnings of 6,806 USD/month vs. 3,666 USD/month for inframarginal participants) but suffered substantially greater lockdown impacts: 32.3% became furloughed or unemployed versus only 3.4% of inframarginal participants.&lt;/p&gt;
&lt;p&gt;Existing methods designed to handle selection on unobservables also perform poorly. Worst-case (Manski) bounds contain the truth but are very wide: employment before lockdown is bounded between 30% and 83% against a true value of 57%. Monotone response selection assumptions produce bounds that do not contain the population quantities for any of the six outcomes, because the marginal survey response function is empirically non-monotone. A Heckman parametric selection model produces point estimates inconsistent with the ground truth (e.g., estimating 51% pre-lockdown employment against the true 57%).&lt;/p&gt;
&lt;p&gt;Investigation of participation timing reveals that reminder emails attract a qualitatively different type of respondent than incentives do. This motivates the paper&amp;rsquo;s central methodological contribution: a two-dimensional participation model that distinguishes &amp;ldquo;active&amp;rdquo; nonparticipants (those who received the invitation and chose not to respond because the incentive was insufficient) from &amp;ldquo;passive&amp;rdquo; nonparticipants (those who never received or attended to the invitation but who may respond to reminders). These two groups have labor market outcomes that differ from participants in opposite directions, which is why single-dimensional monotone selection models fail. The two-dimensional model, exploiting both incentive randomization and the timing of responses, produces bounds that contain or are closer to the ground truth than all other methods examined — for example, bounding pre-lockdown employment at [48%, 63%] around the true value of 57%.&lt;/p&gt;
&lt;p&gt;The paper is scoped to a high-quality, randomly sampled, administrative-data-linked survey conducted during a period of acute economic disruption. The authors note the patterns observed may differ outside crisis periods, though the methods developed apply generally.&lt;/p&gt;
&lt;p&gt;Q: How prevalent is nonresponse bias discussion in economics research, and what methods do researchers currently use?
A: A systematic review of survey-based papers in top-five economics journals from January 2015 to August 2020 found that nearly half of studies omit any discussion of nonresponse bias despite often high nonresponse rates. Among studies using researcher-collected survey data, the average nonresponse rate is 50%; rates reach as high as 87%. When researchers do address nonresponse, 47% of own-survey papers compare sample means to a reference population and 16% apply reweighting on observables; virtually none use methods that address selection on unobservables.&lt;/p&gt;
&lt;p&gt;Q: How was the NiK survey designed to enable testing for nonresponse bias?
A: The 10,000-person random sample was assigned to five incentive groups with probabilities of receiving a 1,000 NOK credit card set at 0%, 1%, 5%, 7%, and 10%, yielding expected payoffs ranging from 1.1 USD to 11 USD. Because group assignment was random, the groups are probabilistically identical ex ante, so differences in average responses across groups — given an exclusion restriction that incentives do not directly affect answers — provide a direct test for nonresponse bias. Participation rates across the aggregated no/low/high incentive groups were 45.7%, approximately 47.6%, and approximately 51.7%, respectively; the joint test of equal participation across groups rejects with p-value &amp;lt; 0.01.&lt;/p&gt;
&lt;p&gt;Q: How large is nonresponse bias in the NiK survey as measured against the administrative ground truth?
A: Across all six administrative outcomes and all three incentive arms, joint tests of no nonresponse bias are rejected with p-values &amp;lt; 0.01. High-incentive arm participants had pre-lockdown monthly earnings roughly 930 USD (30%) above the population mean, and were 10.8 percentage points (19%) more likely to be employed. The high-incentive arm&amp;rsquo;s estimated post-lockdown employment rate of 58% overstates the true rate by 8 percentage points; a researcher comparing this to the true pre-lockdown rate of 57% would erroneously conclude employment was essentially unchanged, when in fact it dropped 7 percentage points.&lt;/p&gt;
&lt;p&gt;Q: Does correcting for observable characteristics remove nonresponse bias?
A: No. After reweighting by propensity scores constructed from age, gender, immigration status, schooling, and municipality or individual-level characteristics, joint tests of zero remaining nonresponse bias are rejected with p-values &amp;lt; 0.01 for each specification and incentive arm. In some cases, reweighting on individual characteristics more than doubles the nonresponse bias — for example, for earnings loss and employment loss measures in the high-incentive arm — meaning that standard observable-based corrections can amplify rather than reduce bias. Robustness checks using machine learning algorithms, class weights, imputation, and richer covariate sets including lagged outcomes yield the same conclusion.&lt;/p&gt;
&lt;p&gt;Q: Does nonresponse bias in survey responses (not just administrative outcomes) differ across incentive arms?
A: Yes. For survey-elicited outcomes, average responses differ significantly across incentive arms, with all joint equality tests rejected at p &amp;lt; 0.1. For example, 10.4% of high-incentive participants reported applying for UI benefits versus 7.5% in the no-incentive group. Estimated UI expenditure as a share of Norway&amp;rsquo;s 2020 social insurance budget varies from 13.2% (no-incentive arm) to 18.4% (high-incentive arm), illustrating the policy stakes.&lt;/p&gt;
&lt;p&gt;Q: Do higher response rates reduce nonresponse bias?
A: Not in this survey. The no-incentive arm, with the lowest participation rate (45.7%), exhibits smaller nonresponse bias than the high-incentive arm (51.7% participation). This finding contradicts standard guidance from the U.S. Office of Management and Budget and J-PAL research guidelines, which equate higher response rates with lower bias risk. The authors note that J-PAL has subsequently updated its guidance in response to this paper&amp;rsquo;s findings.&lt;/p&gt;
&lt;p&gt;Q: How do marginal participants (induced by higher incentives) differ from inframarginal participants?
A: Marginal participants — those who participate only under high incentives but not without them — had average pre-lockdown monthly earnings of 6,806 USD versus 3,666 USD for inframarginal participants (p-value 0.08), indicating much stronger pre-lockdown labor market attachment. Post-lockdown, both groups had similar earnings (approximately 3,600–3,800 USD/month). Consistent with this, 32.3% of marginal participants became furloughed or unemployed after the lockdown versus 3.4% of inframarginal participants. Notably, marginal and inframarginal participants do not differ significantly on observable background characteristics (age, gender, immigrant status, schooling; joint test p-value 0.70), confirming that selection is on unobservables.&lt;/p&gt;
&lt;p&gt;Q: Why do existing methods designed to handle selection on unobservables fail?
A: Worst-case (Manski) bounds contain the truth but are too wide to be informative — pre-lockdown employment is bounded at [30%, 83%] against a true value of 57%. Adding randomized incentives as instruments tightens bounds only modestly (8.5% width reduction for employment before lockdown). Monotone response selection assumptions fail because the empirically estimated marginal survey response function is non-monotone: for employment, the probability first decreases and then increases as a function of willingness-to-participate. The Heckman parametric selection model gives point estimates inconsistent with the ground truth for most outcomes (e.g., 51% estimated pre-lockdown employment vs. 57% true).&lt;/p&gt;
&lt;p&gt;Q: What motivates the two-dimensional participation model?
A: Analysis of participation timing shows that reminder emails attract a qualitatively different type of respondent than incentives alone. Reminders have a larger proportional effect on participation in the no-incentive group than in the high-incentive group, both in absolute and proportional terms. Early respondents (responding to initial contact) had lower pre-lockdown earnings and employment than late respondents (responding to reminders). This implies that the two types of unobservables — resistance to incentive and probability of receiving the invitation — are associated with outcomes that move in opposite directions, producing a non-monotone marginal survey response function that single-dimensional models cannot capture.&lt;/p&gt;
&lt;p&gt;Q: How does the two-dimensional model work and what are its results?
A: The model distinguishes active nonparticipants (saw the invitation, declined because the incentive was too low — more likely to be employed and higher earners) from passive nonparticipants (did not receive or attend to the invitation — more likely to have been adversely affected by the lockdown). By exploiting both the randomized incentive variation and the timing of responses (initial contact vs. reminder), the model partially identifies population mean outcomes under shape restrictions on the joint distribution of the two unobservables. For pre-lockdown employment, the model produces bounds of [48%, 63%] bracketing the true value of 57%, compared to worst-case bounds of [34%, 83%] and monotone selection bounds that do not contain the truth. Improvements are largest for pre-lockdown levels outcomes where the two types of nonparticipants differ most.&lt;/p&gt;
&lt;p&gt;Q: What are the practical recommendations for survey researchers?
A: Embedding randomized incentives in surveys at little or no additional cost enables an inexpensive test for nonresponse bias that does not require linked administrative data. When such a test detects bias, researchers should apply the two-dimensional model rather than relying on observable-based reweighting or conventional selection models. The question of who participates matters at least as much as how many participate; surveys should be designed to characterize and correct for selection, not merely to maximize response rates.&lt;/p&gt;
&lt;p&gt;Nonresponse bias: The difference between the mean response among survey participants and the true population mean, arising when the decision to participate is correlated with the outcome of interest. Distinct from sampling bias; it persists even with a randomly drawn sample.&lt;/p&gt;
&lt;p&gt;Selection on unobservables: Nonresponse bias that remains after conditioning on all observed characteristics. In the NiK survey, marginal and inframarginal participants are indistinguishable on observable demographics but differ dramatically in labor market outcomes, providing direct evidence that unobservables drive selection.&lt;/p&gt;
&lt;p&gt;Marginal vs. inframarginal participants: Under the Imbens-Angrist monotonicity condition, inframarginal participants would respond at any incentive level; marginal participants respond only at higher incentive levels. Their average responses are separately identified using an IV regression with the incentive as instrument.&lt;/p&gt;
&lt;p&gt;Marginal survey response (MSR): The function m(u) = E[Y*_i | U_i = u], giving the average outcome for individuals at the uth quantile of willingness to participate. The MSR is nonparametrically identified for u in [0, p(z_high)]; its empirically non-monotone shape in the NiK data explains why monotone selection assumptions produce bounds that miss the ground truth.&lt;/p&gt;
&lt;p&gt;Active vs. passive nonparticipants: Active nonparticipants received the survey invitation and declined because the incentive was insufficient; they tend to have higher labor market attachment. Passive nonparticipants never received or attended to the invitation but may respond to reminders; they tend to have been more adversely affected by the lockdown. This distinction motivates the two-dimensional model.&lt;/p&gt;
&lt;p&gt;Two-dimensional participation model: A model of survey participation with two unobservables — resistance to incentive (determining active nonresponse) and probability of receiving the invitation (determining passive nonresponse). By exploiting both incentive randomization and the timing of responses (initial contact vs. reminder), the model produces bounds or point estimates on population means that are narrower and closer to ground truth than single-dimensional alternatives.&lt;/p&gt;
&lt;p&gt;Exclusion restriction for incentives: The assumption that randomly assigned incentives affect participation rates but do not directly affect participants&amp;rsquo; answers to survey questions. This is required for incentives to serve as valid instruments for testing and correcting nonresponse bias; the authors test and find no evidence that it is violated.&lt;/p&gt;</description></item><item><title>Slum Upgrading and Long-Run Urban Development: Evidence from Indonesia</title><link>https://macropaperwarehouse.com/papers/slum-upgrading-and-long-run-urban-development-evidence-from-indonesia/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/slum-upgrading-and-long-run-urban-development-evidence-from-indonesia/</guid><description>&lt;p&gt;This paper estimates the long-term causal effects of the Kampung Improvement Program (KIP), one of the world&amp;rsquo;s largest slum upgrading programs, on urban development in Jakarta, Indonesia. KIP ran from 1969 to 1984 across three staggered waves (Pelita I-III), covered 110 square kilometers (25% of Jakarta&amp;rsquo;s area), and served approximately 5 million residents at a total cost of roughly $500 million (2015 USD). The program provided basic physical upgrades — paved roads and footpaths, sanitation and drainage, and community buildings such as schools and health clinics — along with a verbal non-eviction guarantee for 15 years. Residents were not relocated.&lt;/p&gt;
&lt;p&gt;The central research question is whether preserving slums through upgrading entails long-run dynamic inefficiency: as Jakarta formalizes, do KIP areas lag behind non-KIP areas in ways that generate opportunity costs from land misallocation?&lt;/p&gt;
&lt;p&gt;The authors assemble high-resolution data on KIP policy boundaries, current assessed land values (nearly 20,000 sub-blocks), building heights from a novel photographic survey of 19,518 pixels stratified across Jakarta, and multiple novel measures of informality — a rank-based photographic index (0 to 4), an attributes-based index across fifteen binary characteristics, and administrative data on unregistered land-parcel titles. They also use digitized historical maps from 1937 and 1959 to identify pre-KIP kampung boundaries.&lt;/p&gt;
&lt;p&gt;Two empirical strategies address program selection bias (KIP planners prioritized the worst-condition kampungs first). The first restricts the sample to historical kampungs that existed before KIP and includes locality fixed effects, comparing treated kampungs against nearby untreated ones within the same neighborhood. The second is a boundary discontinuity design (BDD) comparing observations within 200 meters of KIP boundaries. Both strategies include eighteen predetermined controls for historical landmarks, infrastructure, and topography including flood proneness.&lt;/p&gt;
&lt;p&gt;Average effects (robust across both strategies): KIP areas today have land values approximately 14-17 log points (roughly 15%) lower than observably equivalent non-KIP areas, and are about 8-12 percentage points less likely to contain buildings taller than three floors — half the control-group mean of 0.24. KIP areas are more informal across all three informality metrics: the rank-based index is higher by 0.29 standard deviations, the attributes-based index by 0.05 SD units, and the share of unregistered parcels is 3 percentage points higher. Building heights corroborate the land-value finding: imputing the hedonic value of missing tall buildings in KIP accounts for approximately 90% of the aggregate land-value impact ($2.2 billion of $2.4 billion).&lt;/p&gt;
&lt;p&gt;Heterogeneity by real estate potential is a central finding. The authors construct a predicted land index for 2,058 hamlets in Jakarta using non-KIP land values. In the lowest quintile (Q5), KIP areas show a positive and statistically significant effect of +10 log points on land values, consistent with direct capitalization of the upgrades. This effect reverses in higher-potential areas: the estimate reaches -28 log points in Q2 and -30 log points in Q1, as non-KIP neighborhoods formalize while KIP areas lag.&lt;/p&gt;
&lt;p&gt;Surplus calculations integrating land values, building heights, horizontal built-up coverage (35% for KIP vs. 18% for non-KIP), and demand and supply elasticities reveal that 90% of total surplus losses are concentrated in the top two quintiles (Q1 and Q2), which comprise 47% of KIP&amp;rsquo;s coverage area. In Q1, KIP surplus is lower by $2,369 per square meter; in Q2, the gap is $1,044 per square meter. In the bottom two quintiles, KIP delivers greater surplus (up to +$347 per square meter in Q5), covering an estimated 3 million residents across 57 square kilometers.&lt;/p&gt;
&lt;p&gt;Mechanisms consistent with delayed formalization include significantly higher population density in KIP areas (+33 log points, or 39%) and greater land fragmentation (+9 parcels per pixel relative to a non-KIP mean of 19), both of which raise relocation and land assembly costs. The original KIP investments show no differential effect by type or intensity after four decades, consistent with their 15-year projected useful life. Endogenous sorting is ruled out as a confounder: if anything, educational attainment is slightly higher in KIP areas.&lt;/p&gt;
&lt;p&gt;Q: What is the Kampung Improvement Program (KIP) and what did it provide?
A: KIP was a slum upgrading program implemented in Jakarta, Indonesia from 1969 to 1984 across three five-year plan waves (Pelita I, II, III). It covered 110 square kilometers and 5 million residents at a total cost of approximately $500 million (2015 USD). The program provided three categories of basic physical improvements — vehicular and pedestrian road access, sanitation and drainage infrastructure, and community buildings (schools, health clinics) — along with a verbal non-eviction guarantee for 15 years. Crucially, upgrades were designed to be basic, with a planned useful life of only 15 years, to avoid attracting higher-income groups.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question and theoretical concern motivating the paper?
A: The paper asks whether slum upgrading programs, while immediately beneficial to residents, entail dynamic inefficiency by delaying formalization as cities develop. The concern is that preserving slums through upgrades and non-eviction guarantees can create opportunity costs from land misallocation when surrounding areas formalize and redevelop into higher-value formal structures. This is framed as a trade-off between the direct welfare benefits of upgrading (affordable in-situ housing for millions) and the long-run costs to urban land productivity.&lt;/p&gt;
&lt;p&gt;Q: How does the paper address the selection bias problem — KIP targeted the worst-condition kampungs first?
A: Two complementary strategies are used. First, the historical kampung specification restricts the sample to areas that were kampungs before KIP (from 1937 and 1959 maps) and includes locality fixed effects, so treated and control units are compared within the same neighborhood and share the same real estate market by assumption. Second, a boundary discontinuity design (BDD) compares observations within 200 meters of KIP boundaries with boundary fixed effects and quadratic distance controls. A falsification test using sequential KIP waves confirms the approach: the raw data shows a monotonic pattern (Wave I worst: -0.40 log points, Wave II: -0.29, Wave III: -0.17) consistent with selection bias, but this pattern disappears in the historical kampung specification (Wave I: -0.13, Wave II: -0.11, Wave III: -0.14), supporting the identification assumption.&lt;/p&gt;
&lt;p&gt;Q: What are the average effects of KIP on land values and building heights?
A: In the historical kampung specification, KIP areas have land values 14 log points (approximately 15%) lower than non-KIP historical kampungs within the same locality. The BDD estimate is similar at -17 log points. For building heights, KIP areas are 12 percentage points less likely to contain a building taller than three floors in the historical kampung sample (8 percentage points in the BDD), relative to a non-KIP control mean of 0.24 — meaning KIP areas are roughly half as likely to have tall buildings. The average effect on floors is -1.6 floors, relative to a control mean of 5 floors.&lt;/p&gt;
&lt;p&gt;Q: How do the authors validate that land value estimates are not distorted by measurement error in informal areas?
A: The authors impute the hedonic value of missing tall buildings in KIP using a hedonic regression estimated solely on non-KIP historical kampungs. KIP areas have 145 fewer buildings with more than ten floors; combined with a 57% price premium for tall buildings (relative to a base price of 13.4 million Rupiahs per square meter), the implied land value loss from missing buildings above ten floors is approximately $1.3 billion, and from buildings between four and ten floors is $0.9 billion, for a total imputed effect of $2.2 billion. This accounts for approximately 90% of the aggregate land value impact from the historical kampung specification ($2.4 billion), assuaging concerns that lower measured land values in KIP reflect data quality differences rather than true price gaps.&lt;/p&gt;
&lt;p&gt;Q: How does the KIP effect vary across the distribution of real estate potential?
A: The authors construct a predicted land index for 2,058 Jakarta hamlets by regressing non-KIP log land values on hamlet fixed effects, then rank hamlets into quintiles. In Q5 (lowest predicted land values, least likely to formalize), KIP areas show a statistically significant positive effect of +10 log points on land values, consistent with direct capitalization of the upgrades. Moving to higher-potential areas, the effect attenuates and reverses: it is -28 log points in Q2 and -30 log points in Q1, where non-KIP areas have formalized. This cross-sectional pattern traces out the dynamic inefficiency predicted by theory.&lt;/p&gt;
&lt;p&gt;Q: What informality measures does the paper construct and what do they show?
A: The paper constructs three complementary informality metrics. First, a rank-based photographic index (0 = very formal, 4 = very informal) coded by two trained Jakarta-based research assistants from approximately 28,000 hand-coded photographs, with inter-rater correlation of 0.78. Second, an attributes-based index averaging fifteen binary characteristics across vehicular access, neighborhood appearance, and structural permanence, standardized to a z-score. Third, the area share of unregistered land parcels from the Indonesian National Land Agency&amp;rsquo;s 2020 digital land maps. KIP areas score higher on all three: the rank-based index is higher by 0.29 SD units, the attributes-based index by 0.05 SD units, and the unregistered parcel share is higher by 3 percentage points.&lt;/p&gt;
&lt;p&gt;Q: What mechanisms explain why KIP areas remain informal and have lower land values?
A: The paper identifies three mutually reinforcing mechanisms. First, KIP areas have significantly higher population density (+33 log points or 39% in the historical kampung sample, equivalent to 51 more people per pixel), which raises relocation costs. Second, KIP areas have greater land fragmentation, with 9 more parcels per pixel relative to a non-KIP mean of 19, exacerbating holdout problems during land assembly; a back-of-the-envelope calculation attributes a 9% land value effect (60% of the total 15% effect) to this channel. Third, the verbal non-eviction guarantees and improved conditions likely strengthened residents&amp;rsquo; tenure perceptions and encouraged them to stay, leading to sub-division of parcels over time. The original KIP investments show no differential effect by type after four decades, consistent with their designed 15-year useful life, and KIP areas have similar access to public amenities today.&lt;/p&gt;
&lt;p&gt;Q: How does the paper calculate surplus and what are the results?
A: The surplus framework compares KIP (informal, tends to stay informal) against non-KIP counterfactuals (more likely formal) on three dimensions: non-KIP areas have (i) higher land values, (ii) taller structures, but (iii) lower horizontal built-up coverage than slums (18% vs. 35% for KIP). Consumer surplus uses a linear demand approximation with elasticity of 0.2 for non-KIP and 0.16 for KIP (backed out from differences in housing budget shares). Producer surplus integrates a Cobb-Douglas supply curve with elasticities of 1.4 (formal) and 1.3 (informal). In Q1, KIP property value is $1,873 per square meter vs. $3,098 for non-KIP, a difference of $1,225 in value terms and $2,369 in surplus terms. The surplus gap falls to $1,044 in Q2, and halves again in Q3, becoming positive (+$347 per square meter) in Q5. Ninety percent of total surplus losses are concentrated in Q1 and Q2, which cover 47% of KIP&amp;rsquo;s area.&lt;/p&gt;
&lt;p&gt;Q: What do the case studies of kampung clearances illustrate?
A: Three Jakarta kampungs cleared in 2015-2016 are examined. Kampung Bukit Duri (Q5, lowest real estate potential) shows a surplus difference of +$572 per square meter in favor of KIP — meaning clearance there is socially inefficient. Kali Pessangrahan (Q3) shows a surplus difference of -$307. Kalijodo (Q2) shows -$910 per square meter, suggesting sizable societal gains from formalization. However, even in Kalijodo, residents were relocated 24 km away to Marunda (a Q5 area), where consumer surplus is only 46% of Kalijodo&amp;rsquo;s — illustrating that societal gains from formalization do not automatically translate into Pareto improvements for evicted residents.&lt;/p&gt;
&lt;p&gt;Q: What robustness checks address alternative explanations?
A: The paper runs several tests. A placebo BDD using 45 non-KIP historical kampung boundaries finds no significant discontinuity, ruling out the hypothesis that slums generically have persistently lower land values. Bandwidth robustness shows consistent BDD estimates from 150 to 500 meters. Tests for spatial spillovers find no spatial decay pattern in land values near KIP boundaries, consistent with the prevalence of gated communities in formal Jakarta minimizing neighborhood contamination. Endogenous sorting is examined using 2010 Census data on 10 million individuals: educational attainment is slightly higher in KIP, and in-migration is slightly lower (1-2 percentage points below mean) with migrants having slightly more years of schooling — both inconsistent with an explanation based on low-skill sorting into KIP. Direct congestion effects from population density are also ruled out by estimating spatial decay around 45 dense non-KIP informal hamlets, finding no decay large enough to explain the land-value effects.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications for slum upgrading in other developing countries?
A: The paper&amp;rsquo;s framework suggests that slum upgrading&amp;rsquo;s cost-benefit balance depends critically on where the upgraded area sits in the real estate potential distribution. In low-potential areas (bottom quintiles of the land index), upgrading delivers net surplus even decades later and implicitly provides affordable housing at scale to millions of residents. In high-potential areas (top quintiles), the opportunity costs from delayed formalization can be large — up to $2,369 per square meter in surplus terms — and the paper suggests that stronger land market institutions to share surplus with informal residents could partially mitigate these costs. The paper also notes that formalization involves complex institutional and political challenges: relocating millions of kampung residents is logistically difficult, compensation is frequently inadequate or absent, and land assembly faces severe holdout problems.&lt;/p&gt;
&lt;p&gt;Dynamic inefficiency in cities: The phenomenon, in the context of this paper, whereby preserving informal slum settlements through upgrading delays their formalization, generating opportunity costs from land misallocation as surrounding formal areas develop. Distinguished from static inefficiency: KIP may raise resident welfare while simultaneously reducing aggregate land productivity.&lt;/p&gt;
&lt;p&gt;Slum upgrading: A policy providing basic public goods improvements (roads, sanitation, community buildings) and tenure security (typically verbal non-eviction guarantees) to existing slum residents in situ, without relocating them. Contrasted with formalization (redevelopment) and sites-and-services programs.&lt;/p&gt;
&lt;p&gt;Boundary discontinuity design (BDD): The paper&amp;rsquo;s second identification strategy, comparing outcomes for observations within 200 meters on either side of KIP program boundaries, with boundary fixed effects and quadratic distance controls, under the assumption that absent KIP, unobserved real estate potential varies smoothly at program boundaries.&lt;/p&gt;
&lt;p&gt;Predicted land index: A hamlet-level index constructed by regressing non-KIP log land values on hamlet fixed effects across 2,058 Jakarta hamlets, used to proxy real estate market potential and rank neighborhoods into quintiles from highest (Q1) to lowest (Q5) development stage.&lt;/p&gt;
&lt;p&gt;Informal surplus: The surplus generated within the informal housing sector, including built-up volume from high horizontal coverage (35% for KIP kampungs) and low-cost informal structures, which is destroyed upon formalization and must be weighed against the gains from taller, higher-value formal developments.&lt;/p&gt;
&lt;p&gt;Land fragmentation: The number of distinct land parcels per unit area (pixel), measured from Jakarta&amp;rsquo;s 2011 cadastral maps. Higher fragmentation exacerbates holdout problems in land assembly, raising the cost of redevelopment and contributing to delayed formalization.&lt;/p&gt;
&lt;p&gt;Source text origin: A classification in the paper&amp;rsquo;s summarization pipeline indicating whether the paper text derives from a full PDF or open-access HTML (permitting summarization) versus abstract-only text (which blocks summarization). All claims in this summary derive from the full paper text.&lt;/p&gt;</description></item><item><title>State Capacity as an Organizational Problem</title><link>https://macropaperwarehouse.com/papers/state-capacity-as-an-organizational-problem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/state-capacity-as-an-organizational-problem/</guid><description>&lt;p&gt;Mastrorocco and Teso study how the internal organization of a state evolves during national development, framing state capacity as an organizational — specifically a principal-agent — problem. Using a new micro-database covering the U.S. federal bureaucracy from 1817 to 1905, they ask: once rulers have incentives to build a state apparatus, how do they organize it to perform its functions across a vast territory, and what drives transitions between organizational forms?&lt;/p&gt;
&lt;p&gt;The dataset is constructed from every issue of the Official Register of the United States published between 1817 and 1905 (44 biennial volumes, 15,801 pages digitized). It records full name, state of birth, state of appointment, occupation, salary, department, office, and location for 304,410 unique federal employees across 810,942 employee-year observations. The authors reconstruct the bureaucracy&amp;rsquo;s four-layer hierarchy (department → office/bureau → division → local office), link employees over time to track careers, categorize all 11,930 occupation codes into five tiers, and geo-code 9,651 places of employment to 1890 county boundaries.&lt;/p&gt;
&lt;p&gt;The paper first documents three sets of descriptive facts. On growth: the federal workforce expanded very slowly before the 1860s and then rapidly, with geographic expansion accounting for none of state growth before 1859 but roughly 29% after. On location: state presence responded positively to local manufacturing activity (a one standard deviation increase in manufacturing employment share raises presence probability by 1.3 percentage points), but distance from Washington DC significantly attenuated this relationship in 1817–1859 and not in 1861–1905. On organization: before the 1860s, employee turnover was high and spiked sharply at presidential transitions (reaching 72% of employees departing in 1861), supervisors&amp;rsquo; departures strongly predicted subordinates&amp;rsquo; departures (a one-for-one supervisor exit raised subordinate turnover probability by 37% pre-1841), and managerial delegation outside DC was stagnant or declining. After the 1860s, turnover trended down (35% at the 1897 transition), the supervisor-subordinate career link weakened materially, and field managers tripled relative to the 1850s.&lt;/p&gt;
&lt;p&gt;The authors argue that high monitoring costs in the early century made trust-based, personalistic organization the second-best solution to principal-agent problems. The limited supply of sufficiently trusted individuals constrained geographic expansion, delegation, and total size. As railroad and telegraph networks lowered communication and transportation costs, monitoring capacity increased, enabling a transition to a Weberian bureaucracy no longer constrained by trust supply.&lt;/p&gt;
&lt;p&gt;The causal identification strategy uses the staggered expansion of the railroad network. For each county and decade (1820–1900), the authors compute the minimum-travel-time route from the county centroid to DC using Donaldson and Hornbeck (2016) data on railroads, steamboat waterways, coastal routes, and land routes. The specification includes county fixed effects, state-by-decade fixed effects, and controls for local railroad presence in the county and for the county&amp;rsquo;s market access, so the identifying variation comes from distant changes in the network that altered travel time to DC without directly affecting the county&amp;rsquo;s local economy or trade access.&lt;/p&gt;
&lt;p&gt;Results: a one standard deviation decrease in travel time to DC raises the probability of federal state presence by approximately 3 percentage points (about 8% of the mean), raises log employment similarly, raises the probability of observing a local managerial layer by approximately 3 percentage points (about 8% of the mean), and reduces employee turnover by approximately 2 percentage points (about 4% of the mean turnover rate). Placebo tests confirm that travel time to other major economic centers does not predict state presence. Telegraph network data (1845–1852, Wang 2020) yield consistent results. An additional test using the post-Civil War decline in Southern-born employee shares shows that better railroad connection to DC narrowed the North-South employment gap, consistent with monitoring substituting for trust-based selection.&lt;/p&gt;
&lt;p&gt;Scope conditions: the paper covers the civilian executive branch of the federal government, excluding the Postal Office, navy yards, and the engineer department; results are robust to restricting to states already in the union at the start of the sample, ruling out frontier-specific dynamics.&lt;/p&gt;
&lt;p&gt;Q: What is the central theoretical claim of the paper?
A: The paper argues that state capacity is fundamentally an organizational problem shaped by principal-agent constraints. When communication and transportation costs are high, the government cannot effectively monitor distant agents, so the second-best solution is to staff the bureaucracy with trusted individuals connected through personal networks. This personalistic form limits size and delegation because the supply of sufficiently trusted individuals is inherently scarce. Technological reductions in monitoring costs allow a transition to a Weberian bureaucracy based on procedural oversight rather than trust, removing the supply constraint on organizational growth.&lt;/p&gt;
&lt;p&gt;Q: What data source does the study rely on, and what time period does it cover?
A: The study draws on the Official Register of the United States, a biennial government publication listing all federal employees, digitized for every issue from 1817 to 1905. The resulting dataset includes 304,410 unique employees and 810,942 employee-year observations, with each record carrying name, state of birth, state of appointment, occupation, salary, department, office, location, and — through hierarchical reconstruction — position in a four-layer chain of command.&lt;/p&gt;
&lt;p&gt;Q: How did the size of the U.S. federal bureaucracy evolve over the nineteenth century?
A: Growth was slow before the 1860s. The first Register for 1817 listed 1,056 employees across 33 pages; the 1905 volume listed over 120,000 employees across 1,254 pages. Geographic expansion contributed zero to state growth before 1859 — the share of counties with any federal employee hovered around 15% from 1817 to 1859 — but contributed approximately 29% of growth after 1859, when county presence rose to 24% by 1871, 38% by 1881, and 61% by 1905.&lt;/p&gt;
&lt;p&gt;Q: What were the three sources of state growth, and how did their relative importance change?
A: The authors decompose growth into: (1) functions (new offices/bureaus), (2) geographic expansion (new counties), and (3) intensity (more employees per county-office pair). Before 1859, growth was entirely driven by functions (~40%) and intensity (~60%), with zero contribution from geographic expansion. After 1859, geographic expansion accounted for ~29%, intensity for ~32%, and functions for ~39% of growth.&lt;/p&gt;
&lt;p&gt;Q: How did employee turnover behave across the century, and what pattern emerges at presidential transitions?
A: Turnover trended upward through the late 1850s and then declined. During presidential transitions, the rate rose from 52–53% in 1841 and 1845 to 60–63% in 1849 and 1853 and peaked at 72% in 1861; it then fell to 55% in 1869, 44–48% in 1885/1889/1893, and 35% in 1897. Turnover was consistently lower in DC than in the field: controlling for year-bureau-position fixed effects, being employed in DC was associated with a 40% reduction in turnover probability.&lt;/p&gt;
&lt;p&gt;Q: How tight was the link between supervisors&amp;rsquo; and subordinates&amp;rsquo; careers, and how did it change?
A: Before 1841, moving from none to all supervisors leaving an organizational unit increased subordinate turnover probability by 37 percentage points. The effect was similar between 1841 and 1859, then dropped substantially to 22 percentage points in the following twenty-year period, and remained roughly constant after 1881. This pattern is consistent with the early bureaucracy relying on chains of personal trust that broke when a supervisor departed.&lt;/p&gt;
&lt;p&gt;Q: What evidence describes the evolution of delegation outside DC?
A: The number of field managers did not grow between 1817 and 1859 — it actually declined in the 1820s and was flat through the mid-1850s — and then tripled by 1905 relative to the 1850s level. The probability that workers in a local office had an additional managerial layer between them and DC was unchanged between pre-1841 and 1841–1859, increased by 5 percentage points between 1861 and 1881, and by 6 percentage points post-1881.&lt;/p&gt;
&lt;p&gt;Q: How does the paper measure monitoring capacity for the causal analysis?
A: The primary measure is travel time in hours from each county centroid to Washington DC, computed decade by decade (1820–1900) as the minimum-cost route across the available railroad network, steamboat waterways, coastal routes, and land routes, using data from Donaldson and Hornbeck (2016). A second, complementary measure is the number of telegraph connections between a county and DC using data from Wang (2020) for 1845–1852.&lt;/p&gt;
&lt;p&gt;Q: What is the identification strategy for the railroad analysis, and why are controls for local railroads and market access important?
A: The specification includes county fixed effects, state-by-decade fixed effects, an indicator for whether the county itself has railroad (LocalRailroad), and the county&amp;rsquo;s market access. County fixed effects mean beta is identified within-county from changes over time. Controlling for local railroad removes the direct correlation between local construction and local economic growth. Controlling for market access removes the effect of distant rail expansion on trade flows that raised agricultural land values and manufacturing activity. The remaining variation in travel time to DC — coming from distant network changes that altered the DC-county connection without affecting local conditions or broader trade access — is the identifying source.&lt;/p&gt;
&lt;p&gt;Q: What are the main quantitative effects of reduced travel time to DC?
A: A one standard deviation decrease in travel time to DC is associated with: (1) approximately 3 percentage point increase in the probability of federal state presence (~8% of the mean); (2) a similar magnitude increase in log employment conditional on presence; (3) approximately 3 percentage point higher probability of an additional managerial layer (~8% of the mean); and (4) approximately 2 percentage point reduction in employee turnover (~4% of the mean turnover rate).&lt;/p&gt;
&lt;p&gt;Q: How do placebo tests support the monitoring interpretation?
A: The authors show that, conditional on the same controls, travel times from a county to a set of other major economic centers are not associated with larger federal state presence. Since these other cities had no role as monitoring headquarters, the absence of an effect for them and the presence of an effect specifically for DC is consistent with the channel operating through the government&amp;rsquo;s ability to supervise agents from the capital, rather than through generic economic connectivity.&lt;/p&gt;
&lt;p&gt;Q: What does the telegraph evidence add, and what is its limitation?
A: Telegraph data (1845–1852, Wang 2020) show that counties with more telegraph connections to DC have larger state presence, more managerial delegation, and lower turnover, consistent with the monitoring mechanism. The limitation is that the authors have limited ability to address the endogeneity of telegraph network timing — the telegraph analysis is treated as corroborating evidence rather than the primary causal identification.&lt;/p&gt;
&lt;p&gt;Q: How do the Southern-born employee results illuminate the trust mechanism?
A: After the Civil War, the share of Southern-born federal bureaucrats fell sharply, consistent with reduced trust toward individuals from former Confederate states. However, counties that became better connected to DC via railroad expansion experienced a relative increase in the share of Southern-born employees. This shows that when monitoring costs fell, the government was willing to hire individuals from groups with lower baseline trust — monitoring substituted for trust as the mechanism ensuring agent performance.&lt;/p&gt;
&lt;p&gt;Q: Does federal state presence crowd out state and local government?
A: No. The presence of federal bureaucrats is positively correlated with the presence of state and local government employees at the county level, suggesting complementarity rather than substitution across levels of government.&lt;/p&gt;
&lt;p&gt;Q: What alternative mechanisms do the authors consider and how do they address them?
A: Three alternatives are discussed. First, demand shocks (Civil War debt repayment, industrialization) could explain the post-1860s expansion; the empirical specifications control for year fixed effects to absorb aggregate time-varying incentives, and the identification relies on differential cross-county variation in DC connectivity. Second, patronage as an electoral tool is consistent with spoils-driven turnover spikes but cannot explain why better-connected counties show lower turnover before civil service reform. Third, cognitive models of the firm (lower communication costs complement managerial problem-solving even without agency problems) could also predict the positive delegation result; the authors note they cannot empirically distinguish the monitoring and cognitive channels, and both may contribute.&lt;/p&gt;
&lt;p&gt;Q: What are the implications for developing countries today?
A: The authors suggest that their findings from nineteenth-century U.S. history may apply to understanding why modern Weberian bureaucracies remain elusive in many developing countries. Where communication infrastructure is limited and monitoring costs remain high, personalistic organizational forms based on trust networks may persist as constrained optima — not failures of will or design, but rational responses to structural conditions. Infrastructure investment that lowers monitoring costs could be a precondition for bureaucratic modernization.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Personalistic state organization&lt;/strong&gt;: The paper&amp;rsquo;s term for the organizational form that prevails when monitoring costs are high. It is characterized by staffing decisions based on personal character, moral reputation, and relationships of trust between principals and agents — and between supervisors and subordinates — rather than on formal procedural monitoring of performance. Frequent turnover at leadership transitions and constrained delegation are defining features, because the supply of trusted individuals is limited.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Weberian bureaucracy&lt;/strong&gt;: In the paper&amp;rsquo;s usage (following Weber 1978), a modern state organization defined by a fixed hierarchy of officials monitored through procedural rules rather than personal trust, lower turnover, and effective delegation of managerial power to geographically dispersed units. The paper treats this as the organizational form enabled by low monitoring costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Monitoring capacity&lt;/strong&gt;: The principal&amp;rsquo;s (politicians in DC and their cabinets) ability to observe and evaluate the behavior of agents (federal employees) throughout the territory. In the paper&amp;rsquo;s operationalization, monitoring capacity is proxied inversely by travel time and communication cost between DC and the county: lower travel time and more telegraph connections mean higher monitoring capacity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Geographic expansion component&lt;/strong&gt;: One of three decomposed sources of state growth. Defined as the increase in state size attributable to the state becoming present in more county locations. This component contributed zero to federal growth before 1859 and approximately 29% of growth after 1859.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Employee turnover&lt;/strong&gt;: In the paper&amp;rsquo;s measurement, the share of employees who leave the federal bureaucracy in a given year. The paper distinguishes politically-driven spikes at presidential transitions — reaching 72% of employees in 1861 — from the secular trend, which rose through the late 1850s and then declined, reaching 35% by the 1897 transition.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Delegation of managerial power&lt;/strong&gt;: The probability that a local county office has an additional managerial layer between its workers and DC, rather than reporting directly to the bureau-level supervisor in Washington. The paper uses this as its measure of whether decision authority has been decentralized to the field.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Trust substitution&lt;/strong&gt;: The paper&amp;rsquo;s mechanism linking monitoring capacity to organizational form. In the absence of effective monitoring, principals substitute trust for oversight — selecting agents whose personal loyalty, moral character, or political alignment gives the principal confidence they will not shirk or defect. As monitoring costs fall, trust becomes less necessary as a screening device, and the trust-constrained supply limit on organizational growth is relaxed.&lt;/p&gt;</description></item><item><title>Taxes Depress Corporate Borrowing: Evidence from Private Firms</title><link>https://macropaperwarehouse.com/papers/taxes-depress-corporate-borrowing-evidence-from-private-firms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/taxes-depress-corporate-borrowing-evidence-from-private-firms/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Does corporate income taxation raise or lower corporate leverage? The canonical Modigliani-Miller (1963) view holds that the interest tax deduction makes debt more attractive, predicting a positive taxes-to-leverage relationship. Most prior empirical work using large public firms confirms this prediction. This paper re-examines the question using data on small private U.S. firms and finds the opposite: higher corporate taxes &lt;em&gt;depress&lt;/em&gt; leverage, at least for small, financially constrained private firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Identification&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The primary dataset is the Federal Reserve&amp;rsquo;s Y-14Q supervisory collection (2011–2017), which covers the loan portfolios of the 33 largest U.S. banks and includes firm-level income statements and balance sheets for privately held, bank-dependent borrowers. The sample is restricted to domestic private C-corporations with prior-year assets above $100 million (to screen for pass-through entities), yielding 39,363 non-singleton firm-year observations. The median firm has $288 million in book assets and total debt-to-assets of approximately 38%. A supplementary dataset from the Shared National Credit (SNC) Program (1993–2018, 50,203 firm-year observations) provides a longer time series on syndicated loan commitments. Public firm comparisons use CRSP-Compustat (91,314 observations, 1989–2017).&lt;/p&gt;
&lt;p&gt;The empirical strategy is a difference-in-differences event study using variation in state corporate income tax rates. A novel contribution is the manual collection of both &lt;em&gt;enactment&lt;/em&gt; dates (when legislation was signed into law) and &lt;em&gt;effective&lt;/em&gt; dates for each state tax change since 1975. Identification follows the narrative approach of Romer and Romer (2010) and Giroud and Rauh (2019) to exclude tax changes endogenous to local economic conditions. The specification includes firm and industry-by-year fixed effects, and the analysis uses heterogeneity-robust estimators (Borusyak et al. 2024; de Chaisemartin and D&amp;rsquo;Haultfoeuille 2020) to address staggered treatment timing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Empirical Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For small private firms (below-median total assets, i.e., below $288 million), long-term debt-to-assets rises by approximately 4% in the year of tax cut &lt;em&gt;enactment&lt;/em&gt; and remains elevated—at approximately 2%—four or more years later, indicating a permanent increase in leverage. This anticipation effect arises because firms respond to the law&amp;rsquo;s passage, not its effective date; results using effective dates are noisy and largely insignificant. The average tax cut during the sample period was 1.2 percentage points, representing approximately a 6% reduction in firms&amp;rsquo; tax bills (given an average private-firm tax rate of 21%), and the implied leverage change of about 6% at year four is correspondingly large, consistent with a low-interest-rate environment in which small changes in marginal q translate into large investment and borrowing responses.&lt;/p&gt;
&lt;p&gt;For large private firms (above-median assets), leverage shows no significant response to tax cuts in any event year. For public firms, evidence of any effect is scant, with at most transient significance and pre-trend issues that complicate interpretation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanism&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper argues two tax-sensitive costs of debt offset the standard interest tax shield. First, a higher tax rate reduces after-tax profits, raising default probabilities and credit spreads endogenously; a tax cut thus lowers credit spreads and incentivizes more borrowing. Second, because external equity finance is either unavailable or very costly for small private firms, debt and capital are complements in financing investment: a tax cut raises the marginal product of capital, inducing firms to invest and borrow more. For small firms with low capital adjustment costs, this capital-debt complementarity dominates the direct loss of interest tax shield value. For large firms with high capital adjustment costs (estimated at nine times the small-firm value), investment responds sluggishly to tax changes, the complementarity effect is muted, and the traditional tax shield effect becomes relatively more important—producing the standard, slightly positive taxes-to-leverage relationship.&lt;/p&gt;
&lt;p&gt;Bank-assessed default probabilities fall by 20–30 basis points (roughly a 10% decline from an average of approximately 2%) in the year of enactment or one year later for small borrowers, directly supporting the model&amp;rsquo;s credit spread mechanism.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Welfare Counterfactual&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Removing the interest tax deduction from the estimated model (while retaining profit taxation and restricted equity access) causes leverage to fall from 0.36 to −0.26. Firms substitute into cash holdings, shrinking the capital stock. In equilibrium, hours worked rise, the real wage falls, and consumer welfare drops by approximately 1.8%. The interest deduction thus raises welfare in a second-best sense by offsetting other frictions that impede optimal capital accumulation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-do-prior-studies-find-a-positive-taxes-to-leverage-relationship-and-how-does-this-paper-differ"&gt;Q1. Why do prior studies find a positive taxes-to-leverage relationship, and how does this paper differ?&lt;/h3&gt;
&lt;p&gt;Prior studies—including Titman and Wessels (1988), Heider and Ljungqvist (2015), and Faccio and Xu (2015)—predominantly use large public firms, for which the interest tax shield is the quantitatively dominant consideration. The present paper focuses on small private firms that face greater financial frictions (restricted equity access, higher default risk), in which two additional tax-sensitive costs of debt become quantitatively important. A further methodological difference from Heider and Ljungqvist (2015) is the use of firm fixed effects rather than first differences, which the authors argue is appropriate in a staggered DiD design.&lt;/p&gt;
&lt;h3 id="q2-why-use-enactment-dates-rather-than-effective-dates-as-the-event"&gt;Q2. Why use enactment dates rather than effective dates as the event?&lt;/h3&gt;
&lt;p&gt;Tax legislation is often signed into law one to two years before taking effect; in the sample of 125 tax packages since 1975, 33 became effective the following year and 13 became effective two or more years later. Firms that anticipate future tax changes will adjust leverage immediately upon enactment, not at the effective date. Results confirm this: event studies using enactment dates yield precise positive estimates for small firms (ranging from ~4% at year 0 to ~2% at year 4+), while results using effective dates are noisy and mostly insignificant. The paper therefore treats the enactment date as the economically relevant event and collects these dates as a novel contribution.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-economic-magnitude-of-the-leverage-response-for-small-private-firms"&gt;Q3. What is the economic magnitude of the leverage response for small private firms?&lt;/h3&gt;
&lt;p&gt;Small firms&amp;rsquo; long-term debt-to-assets rises by almost 4% in the enactment year and remains elevated at approximately 2% four or more years after enactment, consistent with a permanent adjustment. The average tax cut during the period was 1.2 percentage points, representing roughly a 6% reduction in the average tax bill (given an average effective rate of 21% for private firms, per Zwick et al. 2016). The estimated coefficient of 0.021 in year four also implies approximately a 6% change in leverage, a large response that the paper attributes to the low interest rate environment amplifying the marginal q effect of even modest tax changes.&lt;/p&gt;
&lt;h3 id="q4-do-large-private-firms-respond-differently-to-tax-cuts-and-why"&gt;Q4. Do large private firms respond differently to tax cuts, and why?&lt;/h3&gt;
&lt;p&gt;Large private firms (above the median of $288 million in total assets) show no statistically significant leverage response to tax cuts in any event year, and this null is not attributable to wider confidence intervals. The model estimation explains this via capital adjustment costs: the adjustment cost parameter for large firms is estimated to be nine times larger than for small firms. With high adjustment costs, investment responds sluggishly to a tax cut, so the complementarity channel (more investment requires more debt) is suppressed. The traditional tax shield effect then becomes relatively more important, producing a slightly positive (or zero net) taxes-to-leverage relationship consistent with the large-firm data moment.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-model-generate-a-negative-relationship-between-taxes-and-leverage-when-the-interest-tax-deduction-is-present"&gt;Q5. How does the model generate a negative relationship between taxes and leverage when the interest tax deduction is present?&lt;/h3&gt;
&lt;p&gt;Two mechanisms offset the tax shield. First, higher taxes reduce after-tax profits, pushing firms closer to the default threshold; this is capitalized into equilibrium credit spreads, raising the cost of debt. Specifically, for small firms, the model shows that once leverage exceeds approximately 0.47 of assets, the after-tax risky interest rate rises monotonically with the tax rate (rather than falling via the deduction effect). Second, capital and debt are complements in financing investment: because a tax cut raises the marginal product of capital, and because external equity is unavailable, firms substitute into capital by using more leverage. For small firms with low capital adjustment costs, both mechanisms outweigh the loss of interest tax shield value when taxes fall.&lt;/p&gt;
&lt;h3 id="q6-how-are-the-model-parameters-estimated-and-what-are-the-key-parameter-values"&gt;Q6. How are the model parameters estimated, and what are the key parameter values?&lt;/h3&gt;
&lt;p&gt;The model is estimated by simulated method of moments on the Y-14 small-firm sample, minimizing the distance between nine data moments and their model-simulated counterparts. The nine moments include the means and standard deviations of debt, investment, and operating income (all as ratios of assets), the serial correlations of investment and operating income, and the coefficient from a two-way fixed-effects regression of leverage on a tax-change dummy. The deadweight loss in default (ξ) is estimated at 0.6 for small firms and 0.32 for large firms, consistent with elevated financial frictions for small firms and in line with average recovery rates in Kermani and Ma (2023). Fixed operating costs (f) are approximately 0.15 for both samples, amounting to just under half of steady-state operating profits. The serial correlation of the tax process is estimated at 0.662, with innovation standard deviation of 0.022.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-models-welfare-counterfactual-and-what-does-it-imply"&gt;Q7. What is the model&amp;rsquo;s welfare counterfactual, and what does it imply?&lt;/h3&gt;
&lt;p&gt;The paper compares two economies both with profit taxation: one with the interest tax deduction and one without. Removing the deduction in the small-firm model causes leverage to fall from 0.36 to −0.26, as firms hold net cash rather than net debt. The capital stock shrinks, output falls, hours worked rise, and both the real wage and consumption decline. Consumer welfare drops by approximately 1.8%. Capital misallocation (measured following Hsieh and Klenow 2009) worsens from 0.89 to 0.88. The result has a second-best character: the interest deduction incentivizes debt-financed investment that partially offsets the distortion from restricted equity access.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-evidence-on-default-probabilities-add-to-the-empirical-case"&gt;Q8. What does the evidence on default probabilities add to the empirical case?&lt;/h3&gt;
&lt;p&gt;The Y-14 collection contains bank-assessed default probability estimates. In an event study covering Q1 2012–Q4 2018, the authors find that firms&amp;rsquo; assessed default probabilities decline significantly by 20–30 basis points in the year of enactment or one year later for small borrowers (those with total loan commitments of $10–$100 million), representing approximately a 10% decline from the sample average default rate of around 2%. This decline peaks two years after enactment and persists for three years. No comparable decline is observed for larger loan size buckets. Separately, in SNC data, the probability of a non-pass (i.e., below-investment-grade supervisory) rating falls by 1.7–2.2 percentage points following tax cut enactments, persisting roughly three years. Together, these findings directly validate the model mechanism by which tax cuts lower default risk and credit spreads.&lt;/p&gt;
&lt;h3 id="q9-are-the-results-robust-to-alternative-econometric-methods-that-address-heterogeneous-treatment-effects"&gt;Q9. Are the results robust to alternative econometric methods that address heterogeneous treatment effects?&lt;/h3&gt;
&lt;p&gt;Yes. The paper applies the Borusyak et al. (2024) imputation estimator, which imputes fixed effects from untreated observations onto treated observations to remove negative weighting bias; for small firms and event years 0–3, it finds significant positive estimates comparable to the baseline. The de Chaisemartin and D&amp;rsquo;Haultfoeuille (2020, 2021) estimator, based solely on first-time switchers to treatment, yields an effect of 0.036 on leverage for small firms in the enactment year and no effect for large firms, consistent with the baseline. Results using the narrative approach (excluding Connecticut 2011 and 2015, New York 2014, and Rhode Island 2014 as potentially endogenous) produce slightly larger leverage estimates.&lt;/p&gt;
&lt;h3 id="q10-are-tax-hike-effects-symmetric-to-tax-cut-effects"&gt;Q10. Are tax hike effects symmetric to tax cut effects?&lt;/h3&gt;
&lt;p&gt;Evidence on hikes is weaker because tax hikes are rare in the sample. In Y-14 data, hikes are associated with leverage declines for small firms in event year 4 and for large firms in event years 1, 2, and 4, but without sufficient pre-hike observations to identify pre-trends, these results are less credible than the cut results. In SNC data (which spans a longer period, 1992–2018), tax hikes are associated with large and significant reductions in total syndicated borrowing commitments of 6–7%, while cuts produce smaller and marginally significant increases. This asymmetry is consistent with the lower adjustment costs of reducing debt relative to increasing it.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-analysis-of-alternative-model-specifications-reveal-about-the-generality-of-the-mechanism"&gt;Q11. What does the analysis of alternative model specifications reveal about the generality of the mechanism?&lt;/h3&gt;
&lt;p&gt;Three model extensions are considered. In a collateral-constrained model (no endogenous default), the cost of debt is lost financial flexibility (the future shadow cost of the borrowing constraint), which remains tax-sensitive. In a model with costly equity issuance (linear cost λ = 0.11 following Hennessy and Whited 2007), equity issuance is rare, so the model behaves nearly identically to the baseline. In a solvency-based default model (default when firm value turns negative rather than when liquidity is insufficient), the negative taxes-to-leverage result is preserved. A news-shock extension (Jaimovich-Rebelo 2009) incorporating the anticipation of future tax changes also produces lower leverage in response to higher anticipated taxes, consistent with the empirical anticipation effects, though with smaller magnitudes because the news shock variance is smaller than the total tax-change variance.&lt;/p&gt;
&lt;h3 id="q12-why-do-contingent-claims-models-fischer-leland-goldstein-class-always-predict-a-positive-taxes-to-leverage-relationship"&gt;Q12. Why do contingent-claims models (Fischer-Leland-Goldstein class) always predict a positive taxes-to-leverage relationship?&lt;/h3&gt;
&lt;p&gt;In these models, shareholders have deep pockets, so negative cash flows can always be covered; this implies default is rare and the effect of taxes on the default put value is small relative to the direct interest tax deduction. Additionally, these models contain no capital stock, so there is no substitution mechanism between capital and a storage technology (i.e., cash/negative debt). Without endogenous investment, the only channel linking taxes to leverage is the tax shield, which necessarily implies a positive taxes-to-leverage relationship. This is why, as the paper notes, the result was &amp;ldquo;already hiding&amp;rdquo; in the Hennessy-Whited class of dynamic investment models but not visible in the contingent-claims literature.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Interest Tax Deduction (Tax Shield)&lt;/strong&gt;
The paper uses this in the standard corporate finance sense: the after-tax cost of debt is reduced because interest payments are deductible against corporate income. In the model, debt proceeds are discounted at the after-tax interest rate, and the deduction is taken at the time of debt issuance. The paper&amp;rsquo;s contribution is to show this benefit can be outweighed by two tax-sensitive costs of debt, reversing the sign of the taxes-to-leverage relationship for small, constrained firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tax-Sensitive Cost of Debt&lt;/strong&gt;
The paper defines two distinct tax-sensitive costs that offset the tax shield. First, taxes reduce after-tax profits, shifting the default threshold and raising equilibrium credit spreads; this is capitalized into the risky lending rate endogenously from the lender&amp;rsquo;s zero-profit condition. Second, taxes reduce the marginal product of capital, making debt-financed investment less attractive; because debt and capital are complements in a model without external equity, a higher tax rate lowers optimal capital and, with it, optimal debt.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital Adjustment Costs (ψ)&lt;/strong&gt;
Quadratic costs of changing the capital stock, parameterized as ψ(k&amp;rsquo; − (1−δ)k)² / (2k). The paper identifies this parameter as the key determinant of whether leverage responds positively or negatively to taxes: for small firms, ψ is estimated to be near zero (insignificantly different from zero), enabling free substitution between capital and the storage technology (negative debt), so the complementarity channel dominates. For large firms, ψ is estimated to be nine times larger, suppressing this substitution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Default Threshold&lt;/strong&gt;
In the model, default is triggered when the firm&amp;rsquo;s current after-tax profits plus recoverable capital are insufficient to repay debt: (1−τ)(y − wn − f) + (1−ξ)(1−δ)k &amp;lt; p. This threshold depends directly on the tax rate τ, so higher taxes move the threshold in the direction of default, raising credit spreads. The paper provides empirical support for this mechanism via the event study of bank-assessed default probabilities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Enactment Date vs. Effective Date&lt;/strong&gt;
The paper distinguishes between the date tax legislation is signed into law (enactment date) and the date it becomes operative (effective date), which can differ by one to two years. The paper collects novel data on enactment dates from state legislative records. The empirical finding that firms respond to enactment rather than effective dates constitutes evidence of anticipation effects: firms adjust leverage upon observing future expected tax changes, not when the changes actually take hold.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-Best Welfare Effect of the Tax Deduction&lt;/strong&gt;
The paper uses this term to characterize the welfare result from the counterfactual: in an economy already distorted by profit taxation and restricted equity access, the interest deduction raises consumer welfare by incentivizing debt-financed capital accumulation. Removing the deduction causes firms to substitute into cash, shrinking the capital stock and lowering wages and consumption. This is a second-best result because the deduction is welfare-improving only because it partially offsets the distortions created by other frictions; in a frictionless world, no such second-best rationale would apply.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Y-14Q Supervisory Data&lt;/strong&gt;
The Federal Reserve&amp;rsquo;s supervisory collection from the 33 largest U.S. banks, covering loan portfolios and associated borrower financial statements for firms with commercial and industrial loans exceeding $1 million in commitment. The paper uses this dataset because it covers private, bank-dependent firms—a population not previously studied in the tax-leverage literature—and contains firm-level balance sheets, credit ratings, and default probability estimates.&lt;/p&gt;</description></item><item><title>The crowding-in effects of local government debt in China</title><link>https://macropaperwarehouse.com/papers/the-crowding-in-effects-of-local-government-debt-in-china/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-crowding-in-effects-of-local-government-debt-in-china/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks how changes in the &lt;em&gt;composition&lt;/em&gt; (not the size) of Chinese local government debt influence bank risk-taking, credit allocation between privately owned enterprises (POEs) and state-owned enterprises (SOEs), and local total factor productivity. The focus is a 2015 debt-to-bond swap program in which local governments were required to convert outstanding implicit debt — primarily bank loans to local government financing vehicles (LGFVs) and LGFV-issued corporate bonds — into explicitly guaranteed local government bonds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Institutional Context&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Following China&amp;rsquo;s 2008–09 fiscal stimulus, local government debt outstanding rose from 5.8% of GDP in 2006 to 22% by 2013 and reached RMB 15.4 trillion (24% of GDP) by end-2014. The debt was largely held through LGFVs, which are nominally corporate firms but with implicit government backing. Under China&amp;rsquo;s amended budget law effective early 2015, all outstanding debt had to be converted to provincial government bonds through a three-year swap program. Before the swap, government bonds accounted for only 8% of outstanding local government debt; the remaining 92% (approximately RMB 14.17 trillion) needed to be swapped. Commercial banks hold on average 88% of newly issued local government bonds; the government bond share of commercial bank assets rose from 1.7% in 2014 to 14% in 2019.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanism&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Under Basel III capital adequacy ratio (CAR) regulations, Chinese commercial banks — specifically the Big Five systemically important banks using the internal-ratings-based (IRB) approach — assign risk weights above 80% on average to corporate loans, but only 20% (the regulatory approach) to local government bonds. Converting LGFV debt to government bonds therefore reduces banks&amp;rsquo; risk-weighted assets, loosening the binding CAR constraint. The paper formalizes this through a partial-equilibrium model of bank portfolio choice: a lower risk weight on government-bond assets (modeled as a fall in ξ_g) loosens an effective capital constraint, inducing banks to shift toward riskier (POE) lending and reducing the POE-SOE loan rate spread. The model predicts this effect is larger in provinces with higher initial outstanding government debt.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The empirical analysis uses: (1) confidential loan-level data from one of the Big Five Chinese commercial banks covering approximately 400,000 unique firm-loan pairs from 2008:Q1 to 2017:Q4 (regression sample 2013:Q1–2017:Q4); (2) province-level outstanding debt data at end-2014 for 25 provinces, constructed from prefectural-level data collected by Qu et al. (2023); and (3) firm-level balance sheet data from China&amp;rsquo;s Annual Survey of Industrial Firms (ASIF), covering above-scale manufacturing firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings with Quantitative Magnitudes&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Using a triple-difference (DDD) identification — interacting POE status, a post-2015 dummy, and provincial initial government debt — the paper finds:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;At the average level of provincial government debt, the debt swap program reduced the POE credit spread (loan rate deviation from benchmark rate, relative to SOEs) by approximately &lt;strong&gt;3.18 percentage points&lt;/strong&gt; (coefficient α = −3.182, significant at p &amp;lt; 0.01).&lt;/li&gt;
&lt;li&gt;For provinces with initial outstanding debt &lt;strong&gt;one standard deviation above the mean&lt;/strong&gt; (approximately 0.402 log units above mean), the swap reduced the POE credit spread by an additional &lt;strong&gt;1.15 percentage points&lt;/strong&gt; (= 0.402 × 2.849; coefficient β = −2.849, significant at p &amp;lt; 0.01), accounting for 10.1% of the standard deviation of loan rates in the sample.&lt;/li&gt;
&lt;li&gt;In terms of the raw loan rate gap between SOEs and POEs (averaging 42 basis points in the sample), the program narrowed this spread by approximately 6 basis points in high-debt provinces (one standard deviation above mean), accounting for about 1/7 of the average gap.&lt;/li&gt;
&lt;li&gt;On the extensive margin, in provinces with outstanding debt one standard deviation above the mean, the swap raised the &lt;strong&gt;probability of bank lending to POE firms&lt;/strong&gt; by approximately &lt;strong&gt;1.2 percentage points&lt;/strong&gt; (= 0.402 × 0.0292).&lt;/li&gt;
&lt;li&gt;2SLS estimates instrumenting swapped debt by initial outstanding debt interacted with the post-2015 dummy confirm: one standard deviation increase in swapped debt leads to an &lt;strong&gt;11.21% decline&lt;/strong&gt; in the POE loan rate deviation from benchmark relative to SOEs (= 3.723 × 3.013%), accounting for 0.98 standard deviations of the loan rate variable.&lt;/li&gt;
&lt;li&gt;For provincial total factor productivity (TFP), provinces with 1% higher outstanding government debt before the swap experienced a &lt;strong&gt;2.2% larger increase in TFP&lt;/strong&gt; after 2015. The debt swap amount itself (instrumented) has a positive and significant effect on provincial TFP.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions and Parallel-Trends Validation&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Pre-trend tests show that neither the average POE-SOE rate spread (α_τ) nor its interaction with provincial government debt (β_τ) is significantly different from zero in 2014 relative to the base year 2013. Both turn significantly negative only from 2015 onward, validating the parallel-trends assumption. Results are robust to: excluding LGFV firms, excluding large firms (top 10% by assets), restricting to central SOEs as controls (dropping local SOEs), controlling for local debt capacity, GDP growth, FDI/GDP, aged population, total loans, and bank branch fixed effects. A placebo test using the 2016 deleveraging policy shows no significant effect on bank risk-taking, distinguishing the debt-swap mechanism from contemporaneous policy changes.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-key-theoretical-channel-through-which-the-debt-to-bond-swap-affects-bank-lending-to-poes"&gt;Q1. What is the key theoretical channel through which the debt-to-bond swap affects bank lending to POEs?&lt;/h3&gt;
&lt;p&gt;The channel is the risk-weighting mechanism under Basel III capital adequacy ratio (CAR) regulations. Under the IRB approach used by Big Five banks, corporate loans carry average risk weights above 80%, while local government bonds carry a fixed regulatory weight of 20%. Converting LGFV corporate loans and bonds to local government bonds on the bank&amp;rsquo;s balance sheet reduces total risk-weighted assets, loosening the binding CAR constraint. The bank responds by adopting a riskier investment policy — lowering the cutoff ω̂ in the model — which increases lending to POE firms and reduces the POE-SOE credit spread.&lt;/p&gt;
&lt;h3 id="q2-why-is-the-effect-of-the-swap-predicted-to-be-larger-in-provinces-with-higher-initial-outstanding-government-debt"&gt;Q2. Why is the effect of the swap predicted to be larger in provinces with higher initial outstanding government debt?&lt;/h3&gt;
&lt;p&gt;Proposition 2 of the model shows that the sensitivity of the POE loan rate spread to the debt swap policy (∂²ΔR_loan / ∂ξ_g ∂g) is positive, meaning it increases with the amount of government debt g. Provinces with more outstanding debt at end-2014 have more LGFV loans to swap into lower-risk-weight bonds, implying a larger reduction in risk-weighted assets for banks operating in those provinces and hence a larger relaxation of the CAR constraint. Empirically, the correlation between province-level outstanding debt and the amount of swapped debt from 2015–2017 is 0.85 (p-value &amp;lt; 0.0001), confirming the mechanism.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-empirical-specification-identify-the-effect-of-the-debt-swap-rather-than-pre-existing-trends"&gt;Q3. How does the empirical specification identify the effect of the debt swap rather than pre-existing trends?&lt;/h3&gt;
&lt;p&gt;The authors use a triple-difference (DDD) design: the outcome (loan rate deviation from benchmark) is regressed on the interaction POE × Post × GovDebt, where GovDebt is the demeaned log of province-level outstanding debt at end-2014. Pre-trend analysis (Equation 16) estimates year-specific coefficients α_τ and β_τ using 2013 as the reference year. For 2014, both coefficients are statistically indistinguishable from zero. From 2015 onward, both turn significantly negative at the 95% confidence level, consistent with the debt-swap policy triggering the change and inconsistent with pre-existing differential trends by province debt level.&lt;/p&gt;
&lt;h3 id="q4-how-do-the-authors-establish-that-the-risk-taking-channel-rather-than-a-demand-side-story-drives-the-results"&gt;Q4. How do the authors establish that the risk-taking channel rather than a demand-side story drives the results?&lt;/h3&gt;
&lt;p&gt;Two complementary exercises address demand versus supply. First, the authors add firm × year-quarter fixed effects, which absorb all firm-level time-varying factors (including loan demand). After removing demand effects, the triple-difference coefficient on GovDebt × POE × Post becomes more negative (−23.66, significant at 5%) than the baseline (−2.849), suggesting demand-side movements are not the source of the finding. Second, adding bank-branch × year-quarter fixed effects to remove supply-side heterogeneity makes the triple-difference term insignificant while leaving the POE × Post coefficient at −2.196 (significant at 5%), implying the result is primarily supply-driven and province-specific supply factors captured by the triple interaction absorb into the branch-level controls.&lt;/p&gt;
&lt;h3 id="q5-what-heterogeneous-effects-across-firm-types-provide-additional-evidence-for-the-risk-taking-interpretation"&gt;Q5. What heterogeneous effects across firm types provide additional evidence for the risk-taking interpretation?&lt;/h3&gt;
&lt;p&gt;Three dimensions of heterogeneity all point toward bank risk-taking. (a) Size: the credit-easing effect (coefficient on GovDebt × POE × Post) is larger in magnitude for small POEs (by firm assets or by loan size) than for large POEs, consistent with small firms being riskier borrowers. (b) Credit rating: the effect is larger for low-rating POEs (below AA-) than for high-rating POEs, consistent with banks taking on more risk in response to a loosened CAR constraint. (c) Firm-bank distance: the effect is larger for firms located farther from the lending bank branch, where information asymmetry is more severe, consistent with increased bank risk-taking toward harder-to-monitor borrowers.&lt;/p&gt;
&lt;h3 id="q6-how-do-the-authors-confirm-that-the-debt-swap-program-is-the-operative-channel-rather-than-the-overall-regulation"&gt;Q6. How do the authors confirm that the debt swap program is the operative channel rather than the overall regulation?&lt;/h3&gt;
&lt;p&gt;Using the Bertrand-Mullainathan (2001) 2SLS approach, the authors treat the amount of swapped debt (ln(1 + Swap_jy)) as the channel variable, instrumented by GovDebt_j × Post_y (and its interaction with POE_i for the intensive-margin regression). The first-stage results are strong (F-statistics of 158–268), confirming that provinces with more initial outstanding debt swap more debt after 2015. The second-stage results show: (a) on the intensive margin, a one-standard-deviation increase in swapped debt leads to an 11.21% decline in the POE loan rate deviation from benchmark relative to SOEs; (b) on the extensive margin, provinces with more swapped debt show significantly higher probability of POE lending. Both second-stage estimates are significant, confirming the debt swap program as the transmission channel.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-effect-of-the-debt-swap-on-provincial-total-factor-productivity-and-through-what-channel"&gt;Q7. What is the effect of the debt swap on provincial total factor productivity, and through what channel?&lt;/h3&gt;
&lt;p&gt;Provinces with 1% higher outstanding government debt before the swap experienced a 2.2% larger increase in average provincial TFP after 2015 (column 2 of Table 13, coefficient = 0.0220, significant at p &amp;lt; 0.01), with the parallel-trend analysis showing no significant pre-2015 differential effect (the 2014 coefficient is 0.00346, insignificant). 2SLS estimates using swapped debt as the channel variable confirm a positive, significant effect of swapped debt on provincial TFP, with a coefficient of 0.0253 (p &amp;lt; 0.01) in the second stage. The mechanism is credit reallocation from less-productive SOEs to more-productive POEs, consistent with POEs having higher average productivity as documented in Hsieh and Klenow (2009).&lt;/p&gt;
&lt;h3 id="q8-how-do-the-authors-rule-out-that-the-deleveraging-policy-implemented-in-december-2015-drives-the-results"&gt;Q8. How do the authors rule out that the deleveraging policy (implemented in December 2015) drives the results?&lt;/h3&gt;
&lt;p&gt;A placebo test replaces the Post_y dummy (equal to 1 from 2015 onward) with DeLevy (equal to 1 from 2016 onward, coinciding with the deleveraging policy). Neither the coefficient on GovDebt × POE × DeLevy nor on POE × DeLevy is statistically significant in the placebo regressions (Table 11). This distinguishes the mechanism from the deleveraging policy and confirms that the debt swap program — not deleveraging — is the source of the credit reallocation to POEs.&lt;/p&gt;
&lt;h3 id="q9-how-do-the-authors-confirm-results-are-not-driven-by-the-debt-capacity-channel"&gt;Q9. How do the authors confirm results are not driven by the debt capacity channel?&lt;/h3&gt;
&lt;p&gt;The local government debt reform also regulated debt capacity (the ratio of outstanding debt to a centrally assigned debt limit) for each local government. The authors control for the province-level debt capacity measure (DebtCap_j, the average ratio of local government debt to the debt limit in 2016–2017) alongside the baseline interaction terms. Table 9 shows the baseline results remain valid and significant after including debt capacity controls: the coefficient on GovDebt × POE × Post is −2.210 (p &amp;lt; 0.05) and the POE probability of lending result (coefficient on GovDebt × Post = 0.0277, p &amp;lt; 0.01) both hold, ruling out the debt capacity channel as the driver.&lt;/p&gt;
&lt;h3 id="q10-what-does-the-model-predict-about-the-general-relationship-between-capital-adequacy-requirements-and-bank-risk-taking"&gt;Q10. What does the model predict about the general relationship between capital adequacy requirements and bank risk-taking?&lt;/h3&gt;
&lt;p&gt;Proposition 1 establishes that tightening the capital adequacy ratio requirement (increasing ψ) leads to a safer investment policy (ω̂ increases, meaning the bank sets a higher cutoff before taking risky projects) and a lower leverage ratio. This is the benchmark: the debt swap effectively softens the constraint by reducing risk-weighted assets, analogous to lowering the effective ψ̃, which induces the opposite effect — riskier investment policy (lower ω̂) and lower POE credit spreads. The IRB approach&amp;rsquo;s property that risk weights are higher and increasing in project riskiness (ξ&amp;rsquo;(ω) &amp;lt; 0 and ξ&amp;rsquo;&amp;rsquo;(ω) ≤ 0) is essential for these comparative statics to hold.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Debt-to-Bond Swap Program (2015):&lt;/strong&gt; China&amp;rsquo;s central government program requiring local governments to convert all outstanding non-government-bond debt (primarily bank loans to LGFVs and LGFV-issued corporate bonds) into explicitly guaranteed provincial government bonds over three years starting in 2015. The program covered RMB 15.4 trillion in outstanding debt, of which 92% needed to be converted; by end-2018, approximately 90% of non-government-bond debt had been swapped.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Risk-Weighting Channel:&lt;/strong&gt; The mechanism by which the change in debt composition affects bank lending. Under Basel III&amp;rsquo;s internal-ratings-based (IRB) approach, Chinese Big Five banks assign risk weights above 80% on average to corporate loans but only 20% (the regulatory approach) to local government bonds. Swapping LGFV debt for government bonds reduces the bank&amp;rsquo;s total risk-weighted assets without changing the size of assets, loosening the binding capital adequacy ratio constraint and enabling increased lending to riskier (POE) borrowers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;POE Credit Spread:&lt;/strong&gt; Defined in the paper as the difference between the loan rate for privately owned enterprises (POEs) and that for state-owned enterprises (SOEs), measured as the percentage deviation of each loan&amp;rsquo;s interest rate from the benchmark rate set by the central bank. SOEs are treated as effectively riskless borrowers due to implicit government guarantees; POEs are the riskier counterparts. The paper tracks the POE credit spread as the primary outcome variable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Government Financing Vehicles (LGFVs):&lt;/strong&gt; Nominally corporate firms established by Chinese local governments to raise funds for public investment — primarily through bank loans and LGFV-issued corporate bonds (&amp;ldquo;municipal corporate bonds&amp;rdquo;). LGFVs are implicitly backed by local governments but not explicitly guaranteed, so the bank loans and bonds they issue carry higher Basel III risk weights (treated as corporate exposures) than formal government bonds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital Adequacy Ratio (CAR) Constraint:&lt;/strong&gt; The Basel III requirement that a bank&amp;rsquo;s equity capital exceed a minimum fraction ψ of its risk-weighted assets. For systemically important Big Five banks in China, implemented via the IRB approach for corporate loans and the regulatory approach for government bonds since 2012. In the theoretical model, the CAR constraint is binding and determines the bank&amp;rsquo;s effective leverage; relaxing it (by reducing risk-weighted assets) permits the bank to shift toward riskier lending.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Internal Ratings-Based (IRB) Approach:&lt;/strong&gt; The Basel III methodology used by the Big Five Chinese banks to calculate risk-weighted assets for corporate loan portfolios. Under this approach, the risk weight is an increasing function of credit risk (higher-risk loans receive higher weights), so the average weight on corporate loans exceeds 80%, and even high-quality loans carry weights above 50%. This contrasts with the fixed 20% regulatory weight assigned to local government bonds.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Crowding-In Effect:&lt;/strong&gt; In this paper&amp;rsquo;s usage, the mechanism by which restructuring local government debt composition — specifically, replacing corporate-form LGFV debt with low-risk-weight government bonds — frees up bank capacity to extend credit to private firms (POEs) that would otherwise face higher credit spreads or loan denial. This is framed as the opposite of the standard crowding-out effect (where more government debt squeezes private credit), arising because it is the &lt;em&gt;composition&lt;/em&gt; rather than the &lt;em&gt;size&lt;/em&gt; of government debt that changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Total Factor Productivity (TFP) Reallocation Effect:&lt;/strong&gt; The paper measures provincial average TFP (using the Brandt et al. 2013 methodology) and documents that provinces with more government debt outstanding before the swap experienced larger TFP gains after 2015, attributing this to credit reallocation from less-productive SOEs to more-productive POEs. The effect is interpreted as a reduction in credit misallocation rather than within-firm productivity improvement.&lt;/p&gt;</description></item><item><title>The Earnings and Labor Supply of U.S. Physicians</title><link>https://macropaperwarehouse.com/papers/the-earnings-and-labor-supply-of-u.s.-physicians/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-earnings-and-labor-supply-of-u.s.-physicians/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; What do U.S. physicians earn, how is that earnings variation structured across geography and specialty, and how much does government healthcare payment policy shape those earnings and — through them — physicians&amp;rsquo; labor supply and long-run talent allocation?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The paper builds a novel administrative panel by merging the universe of U.S. federal individual income tax returns (2005–2017) with: the National Plan and Provider Enumeration System (NPPES) physician registry; Medicare billing records with procedure-level Relative Value Unit (RVU) rates (2012–2017); restricted-use American Community Survey responses; Social Security Administration demographic records; and medical school ranking and graduation data. The main sample covers 11.6 million physician-year observations for 965,000 unique physicians aged 20–70.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Earnings Facts.&lt;/strong&gt; In 2017, average physician total individual income was $350,000 (median $265,000); the distribution is right-skewed — the top 1% of age-40–55 physicians averages $4.0 million. Physicians in aggregate earned $297 billion in pre-tax dollars, equaling 8.6% of total U.S. healthcare spending. The age-earnings profile is steep: earnings are approximately $60,000 during residency, rise to roughly $185,000 by the early thirties, and peak near $425,000 at age 50. Business income — systematically underreported in survey data (ACS estimates are approximately $140,000 lower than tax data during peak career years, almost entirely due to non-reporting of business income) — accounts for nearly one-quarter of earnings at age 50. Earnings differ sharply across specialties: primary care physicians average $201,200 (ages 40–55), about half the sample mean, while surgeons earn roughly twice as much.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Geographic Pattern.&lt;/strong&gt; Contrary to the pattern for lawyers and workers broadly, physician earnings are not highest on the coasts. A movers-based event study (physicians who changed commuting zones once during 2005–2017) finds that roughly 70% of the cross-location income difference is driven by place rather than worker composition. A two-way fixed-effects variance decomposition reveals pronounced negative physician-location sorting: high-earning physicians tend to locate in lower-income commuting zones, while lower-earning physicians locate in higher-income areas — the opposite of the pattern for lawyers. Medicare&amp;rsquo;s relatively weak adjustment of reimbursement rates for local costs (the empirical elasticity of the Geographic Adjustment Factor to median household income is 0.09, versus 0.33 for a broader local price index) can, by the authors&amp;rsquo; estimates, account for approximately one-third of this unusual geographic earnings pattern.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Government Influence — Medicare Price Changes.&lt;/strong&gt; Using procedure-specific RVU changes as a simulated instrument for each physician&amp;rsquo;s Medicare price exposure, the authors find that a 10% increase in the Medicare price instrument leads to a 2.4% increase in professional earnings of physicians aged 40–55. The behavioral supply response is substantial: physicians bill 4.4% more RVUs (supply elasticity of 0.4 after netting out the mechanical component), of which 3.9% reflects more unique procedures and the rest a shift toward higher-paid procedures. Nearly all of the procedure-level supply increase (3.4 out of 3.8 percentage points) comes from treating additional patients rather than more frequent treatment of existing patients. Converting to pass-through: physicians retain $62 of each $100 in additional Medicare spending directly, or approximately $25 of each $100 of any insurance spending once Medicare&amp;rsquo;s documented spillover into private insurance rates is accounted for. For physicians aged 56–70, a 10% increase in earnings driven by reimbursement changes reduces retirement probability by 0.5 percentage points in that year.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Government Influence — ACA Insurance Expansion.&lt;/strong&gt; Using county-level variation in pre-ACA uninsurance rates (as of 2013) as a source of differential exposure to the ACA&amp;rsquo;s Medicaid expansions and Marketplace subsidies (in 24 states expanding Medicaid in 2014 or early 2015), the authors estimate that a 10 percentage point higher baseline uninsurance rate led to 3.9% higher physician earnings four years post-expansion. Scaling by the first stage (a 10 p.p. higher uninsurance rate translating to 4.96 p.p. higher insurance coverage post-expansion), the implied elasticity of physician earnings to the insurance rate is 0.41. The ACA expansion also reduced retirement probability — a 10 p.p. higher insurance coverage rate leads to a 1 p.p. decline in retirement probability — consistent with a medium-run retirement-to-income elasticity of approximately −1.1. In aggregate, 6% of the $110 billion in annual ACA insurance expansion spending accrued to physicians personally, slightly below their 8.6% baseline share of healthcare spending.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Talent Allocation.&lt;/strong&gt; Specialty choice is sticky and entry-restricted. The authors estimate a discrete-choice model of specialty choice using graduates of top-5 medical schools — physicians with effectively unconstrained specialty access — and an aggregate model using USMLE Step 1 score buckets as ability proxies. At the top of the ability distribution, higher specialty earnings strongly attract physicians: increasing primary care physicians&amp;rsquo; hourly income from $98 to $168 per hour (the level of medicine subspecialists) would raise the share of top-5 medical school graduates choosing primary care by approximately 20 percentage points (nearly doubling their representation in primary care). Moving down the USMLE score distribution, the earnings coefficient falls monotonically and turns negative for the lowest score groups — consistent with the model&amp;rsquo;s prediction that entry restrictions cause higher-paying specialties to displace lower-ability applicants as earnings rise, rather than simply attracting more entrants. A more modest counterfactual — raising internal medicine earnings to dermatology levels — raises the average USMLE score in internal medicine by 10 points (from 230.2 to 239.6).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; The earnings estimates are for the period 2005–2017. Pass-through estimates use a short-run price instrument; long-run pass-through may differ depending on private market spillovers and entry. The ACA analysis is restricted to 24 early-expanding states. The specialty-choice model is estimated on medical graduates entering the residency match; the extensive margin of entering medicine itself is not modeled. Health outcome effects of changing physician ability distributions are not estimated.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-level-and-composition-of-physician-earnings-in-the-tax-data-and-how-do-they-compare-to-survey-based-estimates"&gt;Q1. What is the level and composition of physician earnings in the tax data, and how do they compare to survey-based estimates?&lt;/h3&gt;
&lt;p&gt;In 2017, average physician total individual income was $350,000 and median was $265,000; the top 1% of age-40–55 physicians earned $4.0 million on average, more than twice the average of the top 5%. Business income constitutes nearly one-quarter of earnings at age 50 and is concentrated among top earners: 80% of physicians in the top 1% have business income exceeding $25,000, versus 35% overall. ACS survey data for the same physicians underestimate earnings by approximately $140,000 (roughly one-third of the administrative mean) during peak career years, driven entirely by non-reporting of business income on the extensive margin.&lt;/p&gt;
&lt;h3 id="q2-what-share-of-total-us-healthcare-spending-do-physician-earnings-represent-and-what-does-this-imply-for-policy"&gt;Q2. What share of total U.S. healthcare spending do physician earnings represent, and what does this imply for policy?&lt;/h3&gt;
&lt;p&gt;Physicians in aggregate earned $297 billion pre-tax in 2017, equaling 8.6% of total U.S. healthcare spending (approximately $913 of the average American&amp;rsquo;s $10,611 annual healthcare expenditure). After applying a 30% income tax rate, after-tax physician earnings equal approximately 6% of total healthcare spending, or roughly 1% of GDP. The authors note this provides an upper bound on the magnitude of savings available from policies aimed at reducing physician incomes as a strategy for lowering overall healthcare spending.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-age-earnings-profile-of-physicians-evolve-and-what-drives-growth-during-peak-years"&gt;Q3. How does the age-earnings profile of physicians evolve, and what drives growth during peak years?&lt;/h3&gt;
&lt;p&gt;Physician earnings average approximately $60,000 during residency, rise to roughly $185,000 by the early thirties, and peak near $425,000 at age 50, before declining gradually to approximately $270,000 in the late 60s. Growth during peak earning years (ages 40–55) is driven almost entirely by business income: average wages are approximately flat at $285,000 across this age range, while business income and the probability of filing Schedule C rise steadily.&lt;/p&gt;
&lt;h3 id="q4-how-large-and-unusual-is-the-geographic-pattern-of-physician-earnings-and-what-is-the-causal-role-of-location"&gt;Q4. How large and unusual is the geographic pattern of physician earnings, and what is the causal role of location?&lt;/h3&gt;
&lt;p&gt;Physician earnings are highest in lower-income states (not on the coasts), unlike lawyers and the broader workforce. A movers event study finds that approximately 70% of the cross-commuting-zone income difference is attributable to location rather than worker characteristics; within specialty the estimate rises to approximately 85%. A two-way fixed-effects variance decomposition (with limited-mobility-bias corrections following Andrews et al. 2008 and Kline et al. 2020) reveals pronounced negative physician-location sorting, with the corrected covariance between individual and location effects being 0.6–0.8 times the variance of location effects in magnitude but opposite in sign — a pattern that reverses to positive sorting when the same methods are applied to lawyers.&lt;/p&gt;
&lt;h3 id="q5-what-instrument-is-used-to-identify-the-causal-effect-of-medicare-price-changes-on-physician-earnings-and-why-is-it-valid"&gt;Q5. What instrument is used to identify the causal effect of Medicare price changes on physician earnings, and why is it valid?&lt;/h3&gt;
&lt;p&gt;The authors construct a physician-year &amp;ldquo;Medicare price instrument&amp;rdquo; by fixing each physician&amp;rsquo;s service mix at its 2012–2017 average and then multiplying those fixed quantities by annually-updated RVU rates, summing over services. Because the fixed quantity weights exclude behavioral responses, and because national RVU changes from CMS periodic reviews affect physicians differentially according to their pre-determined service mix, variation across physicians and over time is plausibly exogenous to individual physicians&amp;rsquo; income shocks. Year-by-specialty fixed effects absorb common specialty-level price trends.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-magnitudes-of-the-earnings-and-labor-supply-responses-to-medicare-price-changes"&gt;Q6. What are the magnitudes of the earnings and labor supply responses to Medicare price changes?&lt;/h3&gt;
&lt;p&gt;A 10% increase in the Medicare price instrument raises earnings of 40–55 year-old physicians by 2.4% (reduced-form), with a 2SLS elasticity of income to billed RVUs of 0.17. The total-RVU billing coefficient of 1.437 implies a supply elasticity of 0.437 (subtracting 1 for the mechanical component). At the procedure level, a 10% price increase for a specific code leads to 3.8% more billings for that code, of which 3.4 percentage points reflects treating additional patients. For physicians aged 56–70, a 10% earnings increase reduces that year&amp;rsquo;s retirement probability by 0.5 percentage points.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-aca-insurance-expansion-affect-physician-earnings-and-retirement-and-what-is-the-implied-pass-through"&gt;Q7. How does the ACA insurance expansion affect physician earnings and retirement, and what is the implied pass-through?&lt;/h3&gt;
&lt;p&gt;Counties with a 10 percentage point higher pre-ACA uninsurance rate saw 3.9% higher physician earnings by 2017 (four years post-expansion). Scaled by the first stage (4.96 p.p. higher coverage), the elasticity of physician earnings to insurance coverage is 0.41. A 10 p.p. higher insurance coverage rate leads to a 1 p.p. lower retirement probability post-expansion (medium-run elasticity of retirement to income of approximately −1.1). In aggregate, 6% of $110 billion in annual ACA expansion spending — roughly $7.1 billion, or about $8,400 per physician — accrued to physicians.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-earnings-specialty-choice-relationship-vary-across-the-physician-ability-distribution"&gt;Q8. How does the earnings-specialty choice relationship vary across the physician ability distribution?&lt;/h3&gt;
&lt;p&gt;In the individual-level discrete-choice model estimated on top-5 medical school graduates (likely unconstrained in specialty choice), the coefficient on hourly earnings is 0.014. In the aggregate score-group model, the implied earnings coefficient is 0.016 for USMLE scores above 260 and declines monotonically to −0.008 for scores at or below 190. This negative coefficient for low scorers is consistent with the theoretical prediction that higher earnings attract high-ability physicians, leaving fewer slots for lower-ability applicants due to binding entry restrictions — not a reversal of preferences.&lt;/p&gt;
&lt;h3 id="q9-what-are-the-quantitative-implications-for-specialty-choice-if-primary-care-incomes-were-raised-to-subspecialty-levels"&gt;Q9. What are the quantitative implications for specialty choice if primary care incomes were raised to subspecialty levels?&lt;/h3&gt;
&lt;p&gt;Raising primary care hourly income from $98 to $168 (the level of medicine subspecialists) would increase the share of top-5 medical school graduates choosing primary care by approximately 20 percentage points (about 48% would enter primary care, versus the current share), nearly doubling their representation. Nearly half of these reallocations would come from procedural specialties. An analogous exercise raising internal medicine earnings to dermatology levels shifts the average USMLE score in internal medicine from 230.2 to 239.6 — a 10-point increase — as higher-scoring applicants displace lower-scoring ones within a fixed slot constraint.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-pass-through-from-medicare-reimbursements-to-physician-earnings-and-how-does-it-compare-to-rent-sharing-elsewhere"&gt;Q10. What is the pass-through from Medicare reimbursements to physician earnings, and how does it compare to rent-sharing elsewhere?&lt;/h3&gt;
&lt;p&gt;Direct estimates imply physicians retain $62 of each $100 in additional Medicare spending. Accounting for Medicare&amp;rsquo;s documented spillover into private insurance rates (following Clemens and Gottlieb 2017), the pass-through drops to $25 per $100 of total insurance spending. The authors note this is substantially higher than the modest rent-sharing found for average workers in response to firm-level shocks (Card et al. 2018), but comparable to rent-sharing with high-skilled workers benefiting from patent rents (Kline et al. 2019).&lt;/p&gt;
&lt;h3 id="q11-can-medicares-geographic-pricing-policy-explain-the-unusual-geographic-earnings-pattern-for-physicians"&gt;Q11. Can Medicare&amp;rsquo;s geographic pricing policy explain the unusual geographic earnings pattern for physicians?&lt;/h3&gt;
&lt;p&gt;The elasticity of Medicare&amp;rsquo;s Geographic Adjustment Factor (GAF) to commuting zone median household income is 0.09, compared to 0.33 for a broader local price index. Using the authors&amp;rsquo; short-run estimate that a 10% increase in Medicare prices raises earnings by 2.4%, a counterfactual simulation shows that if the GAF-to-income elasticity rose to 0.33 (aligning Medicare rates with the general cost-of-living gradient), the geographic physician earnings pattern would more closely resemble that of lawyers. The authors estimate that the gap in Medicare&amp;rsquo;s local cost adjustment explains approximately one-third of the unusual physician earnings geography, conditional on the short-run pass-through estimate.&lt;/p&gt;
&lt;h3 id="q12-how-does-the-theoretical-model-of-specialty-choice-and-entry-restrictions-guide-the-empirical-predictions"&gt;Q12. How does the theoretical model of specialty choice and entry restrictions guide the empirical predictions?&lt;/h3&gt;
&lt;p&gt;The model features a unit mass of physicians with heterogeneous ability (Pareto-distributed) and idiosyncratic specialty preferences (exponentially distributed). Physicians choose whether to specialize in period 1; government sets reimbursement rates in period 2; physicians choose labor supply in period 3. With a fixed number of residency slots, higher specialty earnings raise the ability cutoff for entry (rationing by ability). This generates a key nonmonotonic empirical prediction: higher-ability physicians respond positively to earnings increases (choosing a specialty more frequently), while lower-ability physicians respond negatively (displaced by the shift upward in the ability cutoff). The model also implies that demand shocks are not moderated by contemporaneous entry, so incumbents capture the full rent — motivating the estimated pass-through.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Medicare Price Instrument (Simulated RVU Instrument).&lt;/strong&gt; A physician-year measure of Medicare payment exposure constructed by holding each physician&amp;rsquo;s service mix fixed at its 2012–2017 average and multiplying those fixed quantities by time-varying national RVU rates, then summing across services. This purges the instrument of behavioral responses, creating exogenous cross-physician variation in price exposure arising from the interaction of fixed service mix with national RVU policy changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Relative Value Unit (RVU).&lt;/strong&gt; The unit by which Medicare defines and reimburses each physician service in the Physician Fee Schedule. RVUs are intended to reflect the time, effort, and resources required to provide each service, but are subject to periodic review by CMS&amp;rsquo;s RVU Update Committee (RUC) and influenced by political factors. Changes in RVUs translate directly into changes in Medicare reimbursement rates for affected services.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pass-Through (Reimbursement to Earnings).&lt;/strong&gt; The share of an additional dollar of Medicare (or insurance) spending that accrues to physicians personally as earnings, after accounting for practice costs, intermediaries, and behavioral responses. The paper estimates $62 per $100 of direct Medicare spending or $25 per $100 of total insurance spending (the latter accounting for Medicare&amp;rsquo;s spillover into private rates).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negative Physician-Location Sorting.&lt;/strong&gt; The empirical finding — robust to limited-mobility-bias corrections — that higher-ability (higher-earning) physicians disproportionately locate in lower-income commuting zones, while lower-earning physicians concentrate in higher-income areas. This is the opposite of the pattern for lawyers and for worker-firm matching in the broader labor literature. The paper attributes part of this pattern to Medicare&amp;rsquo;s incomplete geographic adjustment of reimbursement rates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ability Cutoff (am) in Residency Matching.&lt;/strong&gt; In the paper&amp;rsquo;s theoretical model, the minimum ability level required to gain entry into a restricted-entry specialty. Because the number of residency slots is fixed, the cutoff rises when a specialty&amp;rsquo;s relative earnings increase (attracting more high-ability applicants), displacing lower-ability physicians who would otherwise have entered. This makes the earnings-specialty relationship nonmonotonic across the ability distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Business Income (Pass-Through Entity Income).&lt;/strong&gt; Income from physician-owned practices organized as sole proprietorships, S-corporations, or partnerships, reported on Schedule C or through pass-through entities rather than on Form W-2. In the tax data, business income accounts for nearly one-quarter of physician earnings at career peak and is the main source of earnings for top physicians, but is systematically underreported in survey data (ACS), leading to a roughly one-third underestimate of total earnings during peak years.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Geographic Adjustment Factor (GAF).&lt;/strong&gt; A Medicare policy parameter that multiplies the national RVU rate to adjust physician reimbursements for local input costs (specifically physicians&amp;rsquo; work, practice expenses, and malpractice). The paper documents that the GAF&amp;rsquo;s elasticity to local median household income is 0.09 — far below the 0.33 elasticity of the general local price index — constituting an effective subsidy to rural and lower-income markets relative to higher-income areas.&lt;/p&gt;</description></item><item><title>The Economics of Equilibrium with Indivisible Goods</title><link>https://macropaperwarehouse.com/papers/the-economics-of-equilibrium-with-indivisible-goods/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-economics-of-equilibrium-with-indivisible-goods/</guid><description>&lt;p&gt;This paper develops an economic theory of competitive equilibrium with indivisible goods that accommodates both complementarities and substitutabilities. The central research question is: what conditions on demand are sufficient, and essentially necessary, for the existence of competitive equilibrium prices when goods are indivisible?&lt;/p&gt;
&lt;p&gt;The classical answer — gross substitutes (Kelso and Crawford, 1982) — entirely rules out complementarities. Complementarities matter in practice, yet prior work showed that equilibrium does not generally exist when all goods are complements (Bikhchandani and Mamer, 1997), while certain patterns of complementarities are compatible with equilibrium (Greenberg and Weber, 1986; Danilov, Koshevoy, and Lang, 2013). The economic content of which patterns permit equilibrium has remained opaque, previously accessible only through combinatorial or tropical geometry.&lt;/p&gt;
&lt;p&gt;Jagadeesan and Teytelboym&amp;rsquo;s key conceptual move is to analyze complementarity and substitutability between bundles of goods, rather than between individual goods. They introduce a bundle consistency condition: each pair of relevant bundles — defined via the compensated price effects of agents — must be either consistently substitutable or consistently complementary across all agents. A bundle is relevant if it arises as a price effect (revealing either direct complementarity or hidden complementarity between a good and an opportunity to sell another good) or consists of a single good. Bundle consistency is formulated as: for each bundling composed only of relevant bundles, each pair of bundles within it must be consistent.&lt;/p&gt;
&lt;p&gt;The paper establishes three core results. First (Theorem 1), for economies in which each agent demands at most one unit of each good, bundle consistency is sufficient for competitive equilibrium existence. Second (Theorem 2), bundle consistency is essentially necessary: if competitive equilibria exist for all economies in which agents have valuations in an invariant domain, then those valuations are bundle-consistent. &amp;ldquo;Invariant&amp;rdquo; requires closure under addition of nonneg linear functions and inclusion of the zero valuation — a condition satisfied by all major prior domains including gross substitutes, consecutive games, substitutes-and-complements, and all classes of discrete convexity. Third, for the multiunit demand setting (Theorems 3 and 4), unit consistency is additionally required: units of the same good must be substitutes for each other. This rules out increasing returns to scale at the unit level, analogous to the absence of increasing returns in standard divisible-good theory.&lt;/p&gt;
&lt;p&gt;The sufficiency proof works by showing that unit- and bundle-consistent preferences lie within a class of discrete convexity (Danilov, Koshevoy, and Murota, 2001), with bundle consistency shown equivalent (Proposition 3) to total unimodularity of the matrix of all agents&amp;rsquo; price effects in {-1, 0, 1}^I. Equilibrium existence then follows from existing results for discrete convex economies.&lt;/p&gt;
&lt;p&gt;A testable characterization is provided: preferences are bundle-consistent if and only if the set of all agents&amp;rsquo; price effects in {-1, 0, 1}^I is totally unimodular (Proposition 3, under unit consistency). This gives a finite, computable test.&lt;/p&gt;
&lt;p&gt;The scope conditions are explicit: the full theorem applies to agents with continuous utility functions strictly increasing in money; income effects are permitted. The necessity results apply to invariant domains. The multiunit extension requires the additional unit consistency condition. The paper does not impose quasilinearity for the main theorems, though geometric appendices restrict to the quasilinear case for the connection to tropical geometry.&lt;/p&gt;
&lt;p&gt;The results unify all previously known sufficient conditions for equilibrium existence with indivisible goods — substitutes, consecutive games, substitutes-and-complements, and the geometric domains — as special cases of bundle consistency. Crucially, Example 3 (four goods, six agent types with additive and pairwise-complement valuations) demonstrates a case where equilibrium exists under bundle consistency even though no bundling makes all agents view bundles as substitutes, so the result cannot be derived from Kelso-Crawford by rebundling.&lt;/p&gt;
&lt;p&gt;Q: What is the fundamental obstruction to equilibrium existence with indivisible goods, according to this paper?&lt;/p&gt;
&lt;p&gt;A: The only essential obstruction is an inconsistency between substitutability and complementarity across a pair of relevant bundles — that is, one agent seeing two bundles as substitutes while another sees them as complements. With only two goods (or only two units), consistency between goods themselves suffices. With more goods, apparent consistency at the good level can mask bundle-level inconsistency, as shown in Example 1 (three goods, each pair complements, yet no equilibrium exists). Bundle consistency — requiring pairwise consistency for all relevant bundlings — captures the full obstruction.&lt;/p&gt;
&lt;p&gt;Q: What makes a bundle &amp;ldquo;relevant&amp;rdquo; for the purpose of bundle consistency?&lt;/p&gt;
&lt;p&gt;A: A bundle b in {-1, 0, 1}^I is relevant if it either arises as a compensated price effect for some agent (revealing which goods move together following a price decrease, including negative entries that reveal hidden complementarities between a good and the opportunity to sell another) or consists of a single good e_i. Bundles with negative components (sale opportunities) are included because sale opportunities can themselves be complementary to goods — the &amp;ldquo;hidden complementarity&amp;rdquo; concept from Ostrovsky (2008) and Hatfield et al. (2013, 2019).&lt;/p&gt;
&lt;p&gt;Q: Why does the three-cycle-of-complements example (Example 1) fail to have an equilibrium, and how does bundle consistency detect this?&lt;/p&gt;
&lt;p&gt;A: Three agents hold V^1 = 3 min{x_a, x_b}, V^2 = 3 min{x_b, x_c}, V^3 = 3 min{x_a, x_c}, with one unit of each good available. Every pair of goods is complementary for some agent, so no inconsistency appears at the goods level. However, under the bundling B = {(1,0,0), (1,1,0), (0,0,1)} (apples-and-bananas bundled, coconuts separate), a fall in the coconut price induces agent 2 to buy the apple-banana bundle and sell apple, making apple and coconut substitutes for agent 2 while they remain complements for agent 3 — a bundle inconsistency. Bundle consistency detects this whereas good-level consistency does not.&lt;/p&gt;
&lt;p&gt;Q: What distinguishes the consecutive-games pattern (Example 2) from the three-cycle pattern (Example 1), and why does equilibrium exist in the former?&lt;/p&gt;
&lt;p&gt;A: In Example 2, agent 3&amp;rsquo;s valuation is replaced by V^3 = 3 min{x_a, x_b, x_c}: coconuts are complementary to apples only in conjunction with bananas, not directly. Under the same bundling B, a fall in the coconut price again makes apple and coconut substitutes for agents 2 and 3, but now this substitutability is consistent — neither agent sees apple and coconut as direct complements independently of bananas. Bundle consistency holds, and Greenberg and Weber (1986) confirm equilibrium existence for all endowments. The difference between the two examples hinges entirely on whether coconuts are directly complementary to apples or only complementary to apples in combination with bananas.&lt;/p&gt;
&lt;p&gt;Q: How does bundle consistency relate to the prior geometric approaches (discrete convexity, tropical geometry)?&lt;/p&gt;
&lt;p&gt;A: Proposition 4 establishes that a family of utility functions belongs to a single class of discrete convexity (Danilov, Koshevoy, and Murota, 2001) if and only if the family is unit- and bundle-consistent. Proposition 3 establishes that (under unit consistency) preferences are bundle-consistent if and only if the set of all agents&amp;rsquo; price effects in {-1, 0, 1}^I is totally unimodular — the same mathematical condition underlying Baldwin and Klemperer&amp;rsquo;s (2019) totally unimodular demand types. The paper thus provides economic interpretations for the entire class of geometric domains, not just substitutes or specific named cases.&lt;/p&gt;
&lt;p&gt;Q: What does unit consistency require, and why is it needed in the multiunit setting?&lt;/p&gt;
&lt;p&gt;A: Unit consistency requires that for any good i and any two serial-number indices m &amp;lt; m&amp;rsquo;, the m-th and m&amp;rsquo;-th units of good i are substitutes for each other (Definition 6). This rules out increasing returns to scale in units of the same good: with one indivisible good, increasing returns arise if and only if units of that good are complements. Since units of the same good are mechanically substitutes in the divisible-good limit, complementarity between units creates an inconsistency between substitutability and complementarity at the unit level. Unit consistency is automatically satisfied when each agent demands at most one unit of each good.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;essentially necessary&amp;rdquo; sense of the necessity results (Theorems 2 and 4)?&lt;/p&gt;
&lt;p&gt;A: The results require that the domain be &amp;ldquo;invariant&amp;rdquo; — closed under addition of nonneg linear price functions and containing the zero valuation. This is satisfied by all major prior domains: substitutes, consecutive games, substitutes-and-complements, sign-consistent tree valuations, all classes of discrete convexity, and all totally unimodular demand types. For any such domain in which competitive equilibria are guaranteed to exist for all economies, the domain&amp;rsquo;s valuations must be bundle-consistent (Theorem 2) or unit- and bundle-consistent (Theorem 4). This is stronger than previous necessity results because it covers any invariant domain, not just specific named ones.&lt;/p&gt;
&lt;p&gt;Q: How can bundle consistency be tested computationally?&lt;/p&gt;
&lt;p&gt;A: Under unit consistency, Proposition 3 gives a finite test: collect all agents&amp;rsquo; compensated price effects that lie in {-1, 0, 1}^I and form a matrix with these vectors as columns. Preferences are bundle-consistent if and only if this matrix is totally unimodular. Total unimodularity of an integer matrix can be verified in polynomial time using standard results from combinatorial optimization (Schrijver, 1998). Example 3 demonstrates this explicitly: for four goods and six agent types (additive plus four pairwise-complement pairs plus one all-complement agent), the 4x9 price-effect matrix is verified to be totally unimodular, confirming bundle consistency and equilibrium existence.&lt;/p&gt;
&lt;p&gt;Q: Does bundle consistency imply that some rebundling of goods makes all agents treat bundles as substitutes?&lt;/p&gt;
&lt;p&gt;A: No — this is a key finding. Example 3 shows a case where bundle consistency holds and equilibrium exists, yet Danilov, Koshevoy, and Lang (2013) confirm that no bundling exists for which all agents view the bundles as substitutes. Thus, the paper&amp;rsquo;s equilibrium existence result is strictly stronger than what could be obtained by applying Kelso and Crawford (1982) after rebundling goods. Bundle consistency is a weaker condition than the existence of a substitute-making rebundling.&lt;/p&gt;
&lt;p&gt;Q: What are the implications of the results for auction design?&lt;/p&gt;
&lt;p&gt;A: The paper suggests that bidding languages for sealed-bid multi-item auctions can be extended beyond the quasilinear-substitutes case (where Milgrom&amp;rsquo;s (2009) assignment messages apply) by using the economic concepts of bundling and consumer theory. Since bundle consistency characterizes when market-clearing prices exist even with complementarities and income effects, auction formats that guarantee equilibrium existence could in principle be designed for the full bundle-consistent domain, accommodating richer preference structures including complementarities and income effects.&lt;/p&gt;
&lt;p&gt;Q: How do &amp;ldquo;hidden complementarities&amp;rdquo; enter the analysis and why must bundles with negative components be considered?&lt;/p&gt;
&lt;p&gt;A: When a good&amp;rsquo;s price falls and demand for another good decreases, this reveals a hidden complementarity between the first good and the opportunity to sell the second. Ostrovsky (2008) and Hatfield et al. (2013, 2019) identified this structure in trading networks. Ignoring these hidden complementarities would miss obstructions to equilibrium existence: Online Appendix E provides an example where the full set of obstructions is only revealed by including bundles with negative components (sale opportunities) among the relevant bundles. This is why relevant bundles are defined to include price effects with negative entries, and bundles in a bundling are allowed to have negative components.&lt;/p&gt;
&lt;p&gt;Bundle consistency: The condition that for each bundling composed solely of relevant bundles, each pair of bundles within it is either consistently substitutable or consistently complementary across all agents — meaning no two agents disagree on whether the bundles are substitutes or complements. This is the paper&amp;rsquo;s central sufficient and essentially necessary condition for equilibrium existence.&lt;/p&gt;
&lt;p&gt;Relevant bundle: A bundle b in {-1, 0, 1}^I that is either a compensated price effect for some agent (a vector describing how demand changes following a price decrease, including negative entries for goods whose demand falls) or the unit vector e_i for a single good i. Only relevant bundles determine the obstructions to equilibrium existence.&lt;/p&gt;
&lt;p&gt;Compensated price effect: A nonzero vector delta_x for which there exist a utility level u, a price vector p, and a lower price p&amp;rsquo;_i at which demand shifts from x to x + delta_x, with unique demand at both prices. Price effects identify which pairs of goods are strict complements (same-sign entries) and which involve hidden complementarities (opposite-sign entries).&lt;/p&gt;
&lt;p&gt;Hidden complementarity: A complementarity between a good and the opportunity to sell another good, revealed when a price effect has a negative entry — meaning demand for some good decreases following the price decrease of another. The concept unifies settings with substitutes and with complements by treating sale opportunities as analogous to goods.&lt;/p&gt;
&lt;p&gt;Unit consistency: The condition that for any good i and any two units m &amp;lt; m&amp;rsquo; of that good, the m-th and m&amp;rsquo;-th units are substitutes. This rules out increasing returns to scale at the unit level and is needed for equilibrium existence in the multiunit demand setting; it is automatically satisfied in the single-unit case.&lt;/p&gt;
&lt;p&gt;Total unimodularity (of price effects): The property, for the matrix formed by stacking all agents&amp;rsquo; price effects in {-1, 0, 1}^I as columns, that every square submatrix has determinant in {-1, 0, 1}. Proposition 3 establishes this is equivalent to bundle consistency under unit consistency, providing a computable test and linking the economic conditions to the geometric literature.&lt;/p&gt;
&lt;p&gt;Invariant domain: A domain V of valuations closed under addition of nonneg linear price functions (V(x) + p*x remains in V for all p &amp;gt;= 0) and containing the zero valuation. Invariance is the scope condition under which the necessity theorems apply; it is satisfied by all major prior equilibrium existence domains.&lt;/p&gt;</description></item><item><title>The Effect of Education Policy on Crime: An Intergenerational Perspective</title><link>https://macropaperwarehouse.com/papers/the-effect-of-education-policy-on-crime-an-intergenerational-perspective/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effect-of-education-policy-on-crime-an-intergenerational-perspective/</guid><description>&lt;p&gt;This paper studies the intergenerational effects of education policy on crime, asking whether a compulsory schooling reform that reduced crime among those directly exposed also reduced crime among their children. The authors exploit the staggered municipal rollout of Sweden&amp;rsquo;s comprehensive school reform, implemented gradually between 1949 and 1962 across more than 1,000 municipalities, which increased compulsory schooling by one to two years, abolished tracking into academic and vocational streams after 6th grade, and introduced a uniform national curriculum. The parent generation consists of all individuals born in Sweden between 1945 and 1955 (approximately 447,000 men and 450,000 women), and their children form the child generation (426,721 sons observed from age 15 to 29). Crime is measured by administrative conviction records from the Swedish National Council for Crime Prevention covering 1973–2010.&lt;/p&gt;
&lt;p&gt;The empirical strategy is difference-in-differences, comparing changes in conviction rates across cohorts in municipalities that implemented the reform at different times, with treatment assigned based on the parent&amp;rsquo;s birth municipality to avoid endogenous sorting bias. Standard errors are clustered at the municipality level. Parallel trends validity is supported by three tests: results are unchanged when municipality-specific linear trends are included, placebo tests using incorrect reform dates yield effects indistinguishable from zero, and residuals from crime regressions show no correlation with municipality-specific trends.&lt;/p&gt;
&lt;p&gt;The main finding is a significant 0.79 percentage point (pp) decline in conviction rates among sons of fathers exposed to the reform (p-value &amp;lt; 0.002), representing a 3.4 percent reduction relative to baseline. The decline spans multiple crime types: violent crime fell by 0.27 pp, traffic-related crime by 0.45 pp, fraud by 0.22 pp, and other offenses by 0.41 pp — percentage reductions of three to six percent across categories. Multiple convictions fell by 0.43 pp (5.8 percent). These second-generation effects are driven entirely by paternal exposure: the impact of maternal reform exposure is an order of magnitude smaller and statistically insignificant, and the difference between paternal and maternal effects is itself significant (p-value 0.048 for any conviction, 0.009 for multiple convictions). Effects on daughters in the child generation are much smaller, with only the residual &amp;ldquo;other crime&amp;rdquo; category showing a significant 0.129 pp (15.5 percent) decline.&lt;/p&gt;
&lt;p&gt;The asymmetry between paternal and maternal transmission is explained by the first-generation effects of the reform. For men, the reform increased schooling by 0.32 years, earnings by approximately 1 percent, the probability of white-collar employment by 1.2 percent, cognitive skills by 0.14 standard deviations, noncognitive skills by 0.17 standard deviations, spousal earnings by 1,022 SEK per year, and overall household income by approximately 1 percent. For women, the reform increased education by 0.21 years but did not raise earnings, household income, or white-collar employment, and did not reduce their already low crime rates. Only 13 percent of women in the 1945–55 cohorts were at or below the compulsory schooling threshold, versus 20 percent of men, substantially limiting the reform&amp;rsquo;s bite for women.&lt;/p&gt;
&lt;p&gt;A mediation analysis decomposes the intergenerational transmission through three channels: fathers&amp;rsquo; education accounts for 64.8 percent of the indirect effect, the decline in paternal crime accounts for 18.5 percent, and the increase in household disposable income accounts for 16.7 percent. The direct effect (unexplained by these mediators) accounts for 48 percent of the total effect. The paper also documents that children of treated fathers attended schools with lower peer crime rates and lived in neighborhoods with lower youth crime rates, supporting a neighborhood and peer effects channel alongside human capital and role-model channels.&lt;/p&gt;
&lt;p&gt;Scope conditions: the study covers male children observed to age 29 in Sweden; results apply to a context of near-universal administrative records, a specific postwar schooling reform, and cohorts born 1945–1955 in a Nordic welfare state.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the intergenerational crime reduction caused by the reform?&lt;/p&gt;
&lt;p&gt;A: Sons of fathers exposed to the reform experienced a 0.79 pp decline in conviction rates (p-value &amp;lt; 0.002), corresponding to a 3.4 percent reduction relative to the baseline conviction rate of approximately 24 percent for the child generation by age 29. Multiple convictions fell by 0.43 pp, a 5.8 percent reduction. These magnitudes are similar in percentage terms to the direct crime reduction the reform caused among fathers themselves.&lt;/p&gt;
&lt;p&gt;Q: Does the reform&amp;rsquo;s intergenerational effect on crime differ by the sex of the treated parent?&lt;/p&gt;
&lt;p&gt;A: Yes. The intergenerational effect is driven entirely by paternal exposure to the reform: the effect of maternal exposure is an order of magnitude smaller and insignificant at any conventional significance level. The difference between paternal and maternal effects is statistically significant, with p-values of 0.048 for any conviction and 0.009 for multiple convictions. The paper attributes this asymmetry to the much weaker first-generation effects of the reform on women&amp;rsquo;s earnings, household income, crime rates, and neighborhood sorting.&lt;/p&gt;
&lt;p&gt;Q: Which crime types declined significantly among sons of treated fathers?&lt;/p&gt;
&lt;p&gt;A: Significant declines were found in violent crime (−0.27 pp, Romano-Wolf p-value 0.09), traffic-related crime (−0.45 pp, RW p-value 0.057), fraud (−0.22 pp, RW p-value 0.09), and other offenses (−0.41 pp, RW p-value 0.047), each representing a three-to-six percent reduction relative to the mean incidence of that crime type. Property crime and drug-related crime did not show significant declines.&lt;/p&gt;
&lt;p&gt;Q: What were the direct effects of the reform on the parent generation&amp;rsquo;s human capital?&lt;/p&gt;
&lt;p&gt;A: For men, the reform increased schooling by 0.32 years, earnings by approximately 1 percent, the probability of white-collar employment by 1.2 percent, cognitive skills by 0.14 standard deviations, and noncognitive skills by 0.17 standard deviations, all measured at military enlistment. Spousal earnings increased by 1,022 SEK per year and overall household income rose by approximately 1 percent. For women, education increased by 0.21 years and marriage market matches improved, but earnings, household income, and white-collar employment probability did not increase significantly.&lt;/p&gt;
&lt;p&gt;Q: Why did the reform have stronger first-generation effects on men than on women?&lt;/p&gt;
&lt;p&gt;A: The average share of individuals at or below the compulsory schooling threshold — the margin at which the reform was binding — was 20 percent for men but only 13 percent for women in the 1945–55 cohorts. Because fewer women were constrained by the old compulsory schooling limit, the reform increased their education by less and produced smaller downstream effects on earnings and labor market outcomes.&lt;/p&gt;
&lt;p&gt;Q: What are the three channels through which the reform reduces child crime, and what is the relative contribution of each?&lt;/p&gt;
&lt;p&gt;A: The paper identifies three channels: (1) the human capital channel, whereby increased parental education raises household income and child human capital; (2) the role model channel, whereby reduced paternal crime participation directly reduces son&amp;rsquo;s crime; and (3) the neighborhood and peer effects channel, whereby higher income enables sorting into lower-crime neighborhoods and better schools. The mediation analysis attributes 64.8 percent of the indirect effect to fathers&amp;rsquo; increased education, 18.5 percent to the decline in paternal crime, and 16.7 percent to the increase in household disposable income. The direct effect unexplained by these three mediators accounts for 48 percent of the total effect.&lt;/p&gt;
&lt;p&gt;Q: What is the role model effect, and how strong is it in the parent generation?&lt;/p&gt;
&lt;p&gt;A: The role model channel operates through the strong intergenerational persistence in crime participation: sons are 2.06 times more likely to participate in crime if their fathers have been convicted (Hjalmarsson and Lindquist, 2012). The reform reduced the incidence of any conviction among treated men by 1.5 pp and repeat convictions by 1.5 pp — the latter representing an approximately 8 percent decline from a lower base. For women, the reform produced no reduction in crime, providing no analogous role model improvement through the maternal channel.&lt;/p&gt;
&lt;p&gt;Q: How does neighborhood and school peer quality change for children of treated fathers versus treated mothers?&lt;/p&gt;
&lt;p&gt;A: Sons of fathers exposed to the reform moved to neighborhoods with lower youth crime rates (−0.087 pp) and attended schools with lower peer crime rates (−0.077 pp). In contrast, sons of mothers exposed to the reform experienced higher neighborhood crime rates (p-value 0.06) and higher school peer crime rates (p-value 0.01), the opposite direction. This asymmetry helps explain why only paternal treatment generates significant second-generation crime reductions.&lt;/p&gt;
&lt;p&gt;Q: What happens to other outcomes for children of treated fathers beyond crime?&lt;/p&gt;
&lt;p&gt;A: Sons experienced a 1.2 percentile increase in school GPA (RW p-value 0.05), a 2.3 pp increase in employment (RW p-value 0.04), a matching 2.3 pp decline in unemployment benefit receipt, a reduction in hospitalization of 2.4 days (17 percent, RW p-value 0.02), and a decline in prescribed drugs of 31 doses (2.8 percent, RW p-value 0.09). The decline in prescribed drugs for sons is driven by nervous system drugs and painkillers, pointing to improved mental health. Daughters of treated fathers show a significant reduction in welfare dependency but no other significant improvements.&lt;/p&gt;
&lt;p&gt;Q: How does the paper validate the parallel trends assumption?&lt;/p&gt;
&lt;p&gt;A: Three tests are reported. First, including municipality-specific linear trends leaves the main coefficient unchanged (p-value 0.85 for the trend terms themselves). Second, placebo contrasts using incorrect reform implementation dates produce effects indistinguishable from zero for all tested dates. Third, graphical inspection of regression residuals shows no correlation with municipality-specific trends. Together these provide strong support for the identifying assumption.&lt;/p&gt;
&lt;p&gt;Q: Are the results sensitive to using a linear probability model instead of a nonlinear model?&lt;/p&gt;
&lt;p&gt;A: A Monte Carlo experiment was conducted replicating observed crime rates across municipalities and imposing the estimated average treatment effect. Assuming the true data-generating process is a probit model, the linear probability model biases the estimated average effect upward by only 5 percent — a difference that is statistically indistinguishable from zero in the actual data — validating the OLS approach.&lt;/p&gt;
&lt;p&gt;Q: What is the broader policy implication of the findings?&lt;/p&gt;
&lt;p&gt;A: The results show that well-designed education policies can reduce crime not only among the directly treated generation but also among their children, amplifying the social benefits of reform across generations. The authors interpret this as consistent with the theoretical framework of Becker and Tomes (1979) on intergenerational transmission of human capital, and suggest that education policy evaluations that focus only on the treated generation substantially understate total social returns.&lt;/p&gt;
&lt;p&gt;Intergenerational transmission of education reform effects: the phenomenon whereby an education policy that raises parental human capital produces improvements in children&amp;rsquo;s outcomes — including crime — through multiple channels including resource increases, parental role modeling, and neighborhood sorting, beyond any direct policy exposure of the child generation.&lt;/p&gt;
&lt;p&gt;Comprehensive school reform (Sweden, 1949–1962): a nationally mandated restructuring of compulsory schooling that extended required attendance by one to two years, abolished selection into academic and vocational tracks after 6th grade, and introduced a uniform national curriculum, rolled out staggered across 1,055 Swedish municipalities.&lt;/p&gt;
&lt;p&gt;Human capital channel: the mechanism by which increased parental education raises earnings and household income, enabling greater investments in children&amp;rsquo;s development and exploiting complementarity between parental and child human capital in the skill production function, thereby raising children&amp;rsquo;s opportunity cost of crime.&lt;/p&gt;
&lt;p&gt;Role model channel: the mechanism by which reduced parental crime participation directly reduces children&amp;rsquo;s crime, operating through the transmission of norms and information across generations; identified empirically by the strong intergenerational correlation in convictions (sons with convicted fathers are 2.06 times more likely to be convicted themselves).&lt;/p&gt;
&lt;p&gt;Neighborhood and peer effects channel: the mechanism by which increased parental income from the reform enables sorting into residential neighborhoods and schools with lower youth crime rates, exposing children to peers less involved in illegal activities and thereby reducing their own crime participation.&lt;/p&gt;
&lt;p&gt;Mediation analysis: a decomposition method following Heckman, Pinto, and Savelyev (2013) that quantifies the share of a total treatment effect accounted for by specific intermediate variables (here: fathers&amp;rsquo; education, fathers&amp;rsquo; crime participation, and household disposable income) versus the direct unexplained effect.&lt;/p&gt;
&lt;p&gt;Conviction rate: the proportion of individuals in a given generation and observation window who received at least one criminal conviction in Swedish administrative records; used as the primary outcome measure because it captures offenses that led to a court appearance, excluding minor infractions resolved by direct fine.&lt;/p&gt;</description></item><item><title>The Effect of Provider Diversity on Racial Health Disparities: Evidence from the Military</title><link>https://macropaperwarehouse.com/papers/the-effect-of-provider-diversity-on-racial-health-disparities-evidence-from-the-military/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effect-of-provider-diversity-on-racial-health-disparities-evidence-from-the-military/</guid><description>&lt;p&gt;This paper asks whether racial concordance between patients and medical providers — specifically, whether Black patients are treated by Black physicians — improves use of preventive care and reduces mortality among patients with chronic, manageable diseases. The authors argue that trust and communication deficits along racial lines cause Black patients to underuse low-cost, life-saving preventive care, and that increasing the share of Black providers addresses this deficit.&lt;/p&gt;
&lt;p&gt;The authors use data from the Military Health System (MHS) Data Repository covering fiscal years 2003–2013, encompassing roughly 9.6 million beneficiaries. A distinctive feature of the MHS is that active-duty providers are themselves MHS beneficiaries, so their race is observed in the same eligibility files used for patients — overcoming the typical absence of provider-race data in claims databases. The study focuses on four chronic, deadly but manageable conditions: diabetes, hypertension, hypercholesterolemia, and clinical atherosclerotic cardiovascular disease. Preventive care is measured by medication fill-days for condition-appropriate generic drugs, HEDIS-recommended Comprehensive Diabetes Care compliance, and (for a subset) blood pressure control. Mortality is tracked across the full sample period.&lt;/p&gt;
&lt;p&gt;The identification strategy exploits quasi-random variation in provider racial composition induced by across-base moves. The MHS setting generates abundant moves driven by DoD personnel management needs — not by patient health or preferences. Using a movers-only differences specification (analogous to Finkelstein et al. 2016), the authors compare differential changes in outcomes for Black versus non-Black patients who move to bases with larger versus smaller increases in the share of Black providers. This design includes fixed effects for both sending and receiving bases, controlling flexibly for regional quality differences. The estimand is an intent-to-treat effect among patients living within 10 miles of a base (who use on-base care 66% of the time).&lt;/p&gt;
&lt;p&gt;The findings are consistent across all four disease samples. For diabetes, a move-induced one-standard-deviation increase in the share of Black diabetes providers is associated with a roughly 6 additional metformin fill-days per year (approximately 16% relative to the mean) and a 3 percentage-point increase (roughly 8% relative to the mean) in Comprehensive Diabetes Care compliance for Black relative to non-Black patients. Mortality falls by 0.4 percentage points — a 33% relative decline — for Black relative to non-Black diabetes patients following such a move.&lt;/p&gt;
&lt;p&gt;Pooling across all four chronic-disease samples, a one-standard-deviation move-induced increase in the Black provider share is associated with approximately 3 additional fill-days of relevant preventive medication and a roughly 0.2 percentage-point reduction in mortality — approximately 15% relative to the mean mortality rate — for Black relative to non-Black patients.&lt;/p&gt;
&lt;p&gt;A decomposition analysis combining the paper&amp;rsquo;s estimates with medical-literature parameters on the mortality effects of preventive medications finds that between 55% and 69% of the concordance mortality effect across the four disease samples can be attributed to improved medication adherence alone, with the remainder attributed to other aspects of the provider-patient relationship (e.g., lifestyle effects, other preventive care).&lt;/p&gt;
&lt;p&gt;Scope conditions: results are local to MHS movers, who are on average slightly younger and healthier than non-movers, potentially understating concordance benefits for the full population. The MHS covers over 3% of all Black U.S. residents, but beneficiaries may differ from the general population. The paper measures Black patient / Black provider concordance specifically; it does not establish a symmetric concordance effect for non-Black patients. The concordance effect estimated is relative — it captures how much Black patients benefit more than non-Black patients from moving to a higher Black-provider-share base. A system-wide spillover mechanism (non-Black providers improving care for Black patients when working alongside more Black providers) cannot be ruled out and would also be consistent with the core concordance motivation.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question and why is the MHS an advantageous setting?
A: The paper asks whether racial concordance between providers and patients causes Black patients to use more preventive care and achieve better health outcomes, focusing on the trust and communication channel. The MHS is advantageous because active-duty providers are themselves MHS beneficiaries, making their race observable — a feature absent in most claims databases. Across-base moves are driven by DoD staffing needs rather than patient health or preferences, providing quasi-random variation in provider racial composition. The system offers complete claims data covering both on- and off-base care, allowing full mortality tracking.&lt;/p&gt;
&lt;p&gt;Q: How does the empirical strategy address selection concerns that plague prior concordance studies?
A: Prior studies face selection problems from Black patients choosing different doctors than white patients and from residential segregation concentrating Black patients and Black physicians in regions with distinct care quality. The movers-based differences specification directly addresses both problems: it uses only patients who move across bases, comparing how the same individual&amp;rsquo;s outcomes change relative to non-Black patients experiencing the same move, as a function of the move-induced change in the Black provider share. Inclusion of fixed effects for both sending and receiving bases accounts flexibly for regional quality differences. Balance tests on observable patient characteristics show no differential sorting of Black versus non-Black patients toward high-Black-provider-share bases.&lt;/p&gt;
&lt;p&gt;Q: What specific preventive care and outcome measures are used for each disease?
A: For diabetes, the primary measures are annual metformin fill-days and Comprehensive Diabetes Care (CDC) compliance — defined as receiving HbA1c testing, a retinal eye exam, and medical attention for nephropathy in the focal year — plus blood pressure control (available only from 2009 onward for on-base patients). For hypertension, the measures are annual fill-days of WHO-recommended antihypertensives (thiazides, ACEs/ARBs, or long-acting dihydropyridine CCBs) and blood pressure control. For hypercholesterolemia, the measure is fill-days of antilipemic agents, bile acid sequestrants, and statins. For atherosclerotic cardiovascular disease, the HEDIS statin therapy receipt indicator is used. Mortality is tracked across all four samples.&lt;/p&gt;
&lt;p&gt;Q: What are the main quantitative results for the diabetes sample?
A: A move-induced one-standard-deviation increase in the share of Black diabetes providers is associated with approximately 6 additional metformin fill-days annually for Black relative to non-Black patients (roughly 16% relative to the mean). Compliance with Comprehensive Diabetes Care increases by 3 percentage points for Black relative to non-Black patients (roughly 8% relative to the mean). Mortality falls by 0.4 percentage points for Black relative to non-Black patients — a 33% relative decline — in connection with the same one-standard-deviation increase in Black provider share.&lt;/p&gt;
&lt;p&gt;Q: What are the pooled results across all four chronic-disease samples?
A: Pooling across diabetes, hypertension, hypercholesterolemia, and atherosclerotic cardiovascular disease, a one-standard-deviation move-induced increase in the Black provider share is associated with approximately 3 additional preventive medication fill-days per year for Black relative to non-Black patients. The pooled mortality effect is a 0.2 percentage-point reduction — roughly 15% relative to the mean mortality rate — for Black relative to non-Black patients.&lt;/p&gt;
&lt;p&gt;Q: How much of the concordance mortality effect operates through medication adherence?
A: The decomposition combines the paper&amp;rsquo;s estimated concordance effects on medication fill-days with medical-literature estimates of the mortality impact of each additional fill-day. For the diabetes sample, increased metformin adherence (4.2 additional fill-days) explains approximately 58.8% of the 0.4 percentage-point concordance mortality effect, with the residual 41.2% attributed to other channels such as lifestyle changes or other preventive care. Across all four disease samples, the medication fill-day channel explains between 55% and 69% of the respective concordance mortality effects.&lt;/p&gt;
&lt;p&gt;Q: What specification checks do the authors conduct to validate causal identification?
A: The authors conduct five main checks. First, balance regressions show that move-induced changes in Black provider share are not differentially related to baseline patient characteristics for Black versus non-Black patients. Second, regressions of the probability of moving on initial Black provider share and its interaction with patient race yield a near-zero concordance coefficient (0.008, SE 0.023), indicating no differential sorting. Third, regressions of post-move on-base care share on the concordance interaction term yield a near-zero coefficient (0.002, SE 0.003), indicating no differential race-specific selection into on-base care. Fourth, a distance falsification test shows that concordance coefficients are near zero and statistically insignificant for patients living more than 10 miles from the base. Fifth, event-study dynamics show no pre-move divergence in preventive care adherence between Black and non-Black patients, with a positive divergence emerging only after the move to a higher Black-provider-share base.&lt;/p&gt;
&lt;p&gt;Q: How does the paper separate a concordance effect from a pure Black-physician-quality effect?
A: The paper estimates a &amp;ldquo;first stage&amp;rdquo; specification on the subsample receiving on-base care (where provider race is observed), regressing the change in the probability of visiting a Black provider on the move-induced change in Black provider density. The results show an approximately one-to-one relationship between higher Black provider availability and increased visits to Black providers for all patients, with only a modest differential by patient race. This confirms that non-Black patients also see more Black providers when Black provider density rises, allowing the interaction specification to isolate concordance from a pure physician-quality effect.&lt;/p&gt;
&lt;p&gt;Q: How do the authors assess the potential role of spillover effects?
A: The authors acknowledge they cannot rule out that some of the estimated concordance effect arises through system-wide spillovers — for instance, non-Black providers on bases with more Black colleagues may improve their care for Black patients through peer learning or information transmission. They note that even if such a spillover mechanism operates, it is still consistent with the paper&amp;rsquo;s core concordance motivation, because provider-knowledge deficiencies about treating Black patients are among the theorized channels of racial discordance.&lt;/p&gt;
&lt;p&gt;Q: What do the results imply for the overall racial mortality gap?
A: Among MHS beneficiaries aged 20–65, Black beneficiaries are roughly 38% more likely to have diabetes and die over the sample period than non-Black beneficiaries; this gap appears driven primarily by higher diabetes prevalence rather than a within-diabetes mortality gap. Applying the diabetes concordance mortality estimate (a 0.4 percentage-point reduction), the authors calculate that a one-standard-deviation increase in the Black provider share would reduce the overall diabetes mortality gap from 38% to approximately 21% — a substantial narrowing driven by the concordance effect operating through conditional-on-prevalence outcomes.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: The results imply that investments in increasing physician workforce diversity could meaningfully reduce racial mortality disparities in the United States, particularly for chronic diseases manageable through preventive medication. The paper notes the results are relevant to affirmative action policies in medical school admissions, specifically the pending Supreme Court cases Students for Fair Admissions v. University of North Carolina and Students for Fair Admissions v. Harvard at the time of writing. The MHS population covered in the study includes over 3% of all Black U.S. residents, so the policy stakes extend substantially beyond the military context.&lt;/p&gt;
&lt;p&gt;Q: What are the limitations of the study regarding generalizability?
A: Movers in the chronic-disease samples are on average about four years younger and 0.2 percentage points less likely to die than non-movers, suggesting the local average treatment effect for movers may understate concordance benefits for the full population. The MHS population may be healthier overall than the general population, though conditioning on chronic-disease patients mitigates this concern. The paper covers only Black-patient/Black-provider concordance; concordance effects for other racial and ethnic groups are not estimated. The estimate of the concordance coefficient technically captures how much the Black patient / Black provider concordance effect exceeds the non-Black patient / non-Black provider concordance effect, meaning the absolute magnitude of Black concordance benefits is understated if non-Black concordance effects are also positive.&lt;/p&gt;
&lt;p&gt;Racial concordance: In this paper&amp;rsquo;s usage, the match between the race of a patient and their treating physician — specifically Black patient / Black provider pairing — theorized to improve care through trust, communication, and reduced provider knowledge deficiencies about Black patients.&lt;/p&gt;
&lt;p&gt;Provider Black share: The fraction of outpatient office visits for a given chronic condition at a given military base that are attended by Black active-duty providers, used as the base-level treatment variable; varies across bases from zero to approximately 20 percentage points in the pooled sample.&lt;/p&gt;
&lt;p&gt;Movers-based differences specification: An identification strategy that restricts to patients who relocate across military bases exactly once during the sample period and estimates the differential change in outcomes for Black versus non-Black patients as a function of the move-induced change in the base&amp;rsquo;s Black provider share, including fixed effects for both the sending and receiving base.&lt;/p&gt;
&lt;p&gt;Intent-to-treat (ITT) effect: The concordance estimate as applied to all patients living within 10 miles of a base — regardless of whether they actually received on-base care — to avoid selection bias from differential race-specific decisions to seek care on versus off base.&lt;/p&gt;
&lt;p&gt;Comprehensive Diabetes Care (CDC): A HEDIS composite measure requiring receipt of all three of the following in the focal year: HbA1c testing, a retinal eye exam, and medical attention for nephropathy (via microalbumin exam, ACE/ARB therapy, or nephropathy treatment).&lt;/p&gt;
&lt;p&gt;Medication fill-days: Annual days of supply dispensed for condition-appropriate generic medications (metformin for diabetes; thiazides/ACEs/ARBs/CCBs for hypertension; antilipemic agents, bile acid sequestrants, and statins for hypercholesterolemia; statins for atherosclerotic cardiovascular disease), used as the primary preventive care adherence measure.&lt;/p&gt;
&lt;p&gt;Decomposition of concordance mortality effect: A calculation that uses the paper&amp;rsquo;s estimated concordance effect on medication fill-days, combined with medical-literature estimates of the mortality impact per fill-day, to determine what share of the total concordance mortality effect passes through medication adherence versus other channels (lifestyle, other preventive care).&lt;/p&gt;</description></item><item><title>The Effects of Gender Integration on Men</title><link>https://macropaperwarehouse.com/papers/the-effects-of-gender-integration-on-men/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effects-of-gender-integration-on-men/</guid><description>&lt;p&gt;Greenberg, Wasserman, and Weber (2024/2026) ask whether men negatively respond—in terms of job performance, behavior, and workplace perceptions—when women first enter an exclusively male occupation. They exploit the staggered 2017-onward integration of women into U.S. Army infantry and armor combat companies following the 2016 rescission of the Ground Combat Exclusion Policy. The setting offers unusually clean causal identification: integration timing within Brigade Combat Teams was neither systematic nor data-driven, the Army&amp;rsquo;s rigid pay scales meant integration posed no displacement or wage threat to incumbent men, and roughly 391 companies are observed over 2012–2020. The empirical strategy is a staggered difference-in-differences design with company fixed effects, BCT-by-year-of-arrival fixed effects, and month-of-year fixed effects, applied to an individual-level sample of newly arrived male soldiers. Outcomes come from monthly administrative personnel records (retention, misconduct separations, demotions, criminal investigations, drug tests, medical profiles, physical fitness scores) and the Defense Organizational Climate Survey (DEOCS), a congressionally mandated annual survey with response rates above 50% covering organizational effectiveness, equal opportunity, and sexual assault prevention and response. The main finding is that integrating women into previously all-male combat companies does not negatively affect men&amp;rsquo;s performance or behavioral outcomes. Estimates are precise enough to rule out small detrimental effects: two years post-integration, the authors can rule out a 3% increase in attrition, a 5% increase in demotions, and a 4% increase in criminal investigations relative to their respective means. One behavioral outcome shows a statistically significant improvement: integration reduces separations for misconduct by 1.3 percentage points (16% of the mean). Drug test positivity also declines. The sole potential negative administrative finding is a 1.8-point decline in physical fitness scores (0.7% of the mean, roughly 5% of a standard deviation), but this does not affect pass rates and becomes statistically insignificant when scores are imputed using observable covariates. An aggregate Performance and Behavior Index rules out reductions of 0.8% of a standard deviation; the No Adverse Outcomes measure rules out a 1.2 percentage point increase (3% of the mean). Despite these null-to-positive performance effects, survey data reveal that integration causes a 5% of a standard deviation decline in men&amp;rsquo;s overall perceptions of workplace quality. This perception decline is concentrated in companies that received a female officer shortly after integration. Among companies integrated only with female enlisted soldiers (no female officer), men&amp;rsquo;s workplace attitudes actually improve by 14.7% of a standard deviation. Two mechanisms are examined: increased male awareness of pre-existing workplace problems (supported by higher reported observations of bullying, hazing, and unwanted comments, especially among male officers in female-officer-integrated companies), and negative reactions to women in positions of authority (supported by broader declines in organizational effectiveness perceptions not confined to equal-opportunity items). Crucially, the perception decline does not translate into retaliatory behavior or performance deterioration; companies integrated with a female officer show some performance gains, and female enlisted soldiers in those companies report fewer workplace problems. Scope conditions: findings apply to a high-stakes, traditionally male-dominated, hierarchical occupational setting during 2017–2020, a period when U.S. deployment missions were primarily advise-and-assist rather than direct combat. Integration increased female representation by approximately 4.7 percentage points on average.&lt;/p&gt;
&lt;p&gt;Q: What was the policy change studied and why does it offer causal leverage?
A: In December 2015, Secretary of Defense Ashton Carter announced that all U.S. military occupations, including infantry and armor combat roles, would open to women starting in 2016. Women did not begin arriving at operational companies until 2017 due to training timelines. Within BCTs, the selection of which companies to integrate was neither systematic nor data-driven, and baseline characteristics of integrated and non-integrated companies are similar after conditioning on BCT and company-type fixed effects, supporting a parallel trends assumption.&lt;/p&gt;
&lt;p&gt;Q: What are the main administrative performance findings?
A: Integration has a positive but statistically insignificant effect on retention, and reduces misconduct separations by 1.3 percentage points (significant at the 5% level), representing a 16% reduction relative to the mean. Demotions, criminal investigations (including sex-related and domestic violence), and medical profiles show no significant negative effects, with precision sufficient to rule out 5% increases in demotions and 4% increases in criminal investigations. Physical fitness scores decline by 1.8 points (0.7% of mean, approximately 5% of a standard deviation), but pass rates are unaffected and the estimate becomes insignificant when scores are imputed with observable covariates.&lt;/p&gt;
&lt;p&gt;Q: What does the aggregate performance index show?
A: The Performance and Behavior Index—an equally weighted z-score average of retention, misconduct separations, demotions, criminal investigations, medical profiles, promotions to Sergeant, and physical fitness outcomes—shows a positive but insignificant effect of integration, ruling out reductions of 0.8% of a standard deviation. The No Adverse Outcomes measure rules out a 1.2 percentage point increase (3% of the mean incidence of adverse outcomes).&lt;/p&gt;
&lt;p&gt;Q: How do men&amp;rsquo;s workplace perceptions change after integration?
A: The overall workplace quality index constructed from all DEOCS Likert-scale items declines by 5% of a standard deviation following integration, spanning perceptions of organizational effectiveness, workplace inclusivity, and sexual assault prevention and response. This average effect masks critical heterogeneity by the rank composition of integrating women.&lt;/p&gt;
&lt;p&gt;Q: What is the key heterogeneity in survey responses?
A: The decline in men&amp;rsquo;s perceptions is entirely driven by companies that received a female officer shortly after integration. In companies integrated only with female enlisted soldiers (17% of integrating companies did not receive a female officer within a month), men&amp;rsquo;s perceptions improve by 14.7% of a standard deviation. Male officers show a larger negative shift than male enlisted soldiers in officer-integrated companies, and this difference is statistically significant.&lt;/p&gt;
&lt;p&gt;Q: What mechanisms explain the negative perception response to female officers?
A: Two mechanisms are investigated. First, increased awareness: male soldiers—especially male officers—report observing more bullying, hazing, and unwanted comments after a female officer is integrated but not after integration with only female enlisted, and the decline in perceptions of sexual assault prevention and response is significantly larger among male officers than enlisted men, consistent with shared leadership roles amplifying awareness of workplace problems. Second, negative reactions to female authority: declines in perceptions are more pronounced on organizational effectiveness questions than on equal-opportunity items and extend to issues unrelated to women, suggesting broader dissatisfaction with female leadership alongside heightened awareness.&lt;/p&gt;
&lt;p&gt;Q: Is the decline in perceptions related to actual differences in female officer qualifications or preferential treatment?
A: No. Female and male officers have similar baseline characteristics including educational background and experience. Companies integrated with female officers perform at least as well as non-integrated companies or those integrated only with enlisted women on administrative metrics. There is no evidence that male officers waited longer for leadership assignments relative to female colleagues, ruling out perceived preferential treatment as a driver.&lt;/p&gt;
&lt;p&gt;Q: Do men&amp;rsquo;s negative perceptions of female officers translate into retaliatory behavior toward women?
A: No. Administrative misconduct metrics show some improvements in male behavior when a female officer is present. Female enlisted soldiers in female-officer-integrated companies report fewer workplace problems on the climate survey than female enlisted soldiers in companies integrated without a female officer, indicating that the presence of a female officer generates benefits for female enlisted soldiers rather than backlash against them.&lt;/p&gt;
&lt;p&gt;Q: Does heterogeneity by integration intensity or women&amp;rsquo;s rank affect administrative outcomes for men?
A: Integration intensity (number of women initially integrated) and rank composition (female officers vs. only female enlisted) do not produce negative administrative outcomes in any subgroup. The aggregate Performance and Behavior Index shows a positive effect when a female officer is included. Effects also do not vary with male soldiers&amp;rsquo; rank (enlisted vs. officer) or their tenure in the company.&lt;/p&gt;
&lt;p&gt;Q: What happens in units that deploy to combat zones?
A: Approximately one in five integrated companies deployed to a combat zone within two years of integration. Integration does not negatively affect retention, behavior, or performance of men in deploying units. Declines in workplace perceptions are larger for deploying units and are most pronounced when integration occurs shortly after return from deployment, consistent with deployment strengthening in-group identity among male soldiers rather than women performing poorly during combat-zone service.&lt;/p&gt;
&lt;p&gt;Q: What do the findings imply for theories of identity economics and the pollution theory of discrimination?
A: The null-to-positive behavioral and performance responses to women&amp;rsquo;s entry contradict the predictions of Akerlof and Kranton&amp;rsquo;s (2000) identity economics model and Goldin&amp;rsquo;s (2014) pollution theory of discrimination, which predict retaliatory or otherwise unproductive behaviors when women enter a male-dominated occupation. The paper shows that, to the extent identity concerns shape male responses, these are confined to subjective perceptions and do not manifest in diminished performance, retention, or conduct.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications for employers considering gender integration?
A: The paper provides evidence against the argument that men will become less productive when women enter previously male-only occupations, a justification sometimes offered for excluding women from such jobs. The finding that performance and behavior are unaffected—and misconduct actually declines—allows policymakers and employers to weigh these results against concerns about operational or productivity costs of integration. The perception gap between men&amp;rsquo;s attitudes and actual outcomes points to a need for targeted leadership and organizational interventions, particularly around the introduction of female leaders.&lt;/p&gt;
&lt;p&gt;Ground Combat Exclusion Policy (GCEP): The U.S. military policy, rescinded in 2013 and fully eliminated by Secretary of Defense Carter in 2016, that precluded women from serving in infantry and armor positions; the policy whose removal is the source of the integration shock studied. | Staggered difference-in-differences: The empirical strategy exploiting the sequential, non-systematic integration of women into combat companies across years 2017–2023, using never-yet-treated companies as a comparison group with company fixed effects and BCT-by-year-of-arrival fixed effects. | Performance and Behavior Index: An equally weighted average of z-scored administrative outcomes (retention, no misconduct separations, no demotions, no criminal investigations, no medical profiles, promotion to Sergeant, physical fitness pass/fail and score), constructed for enlisted soldiers, oriented so higher values indicate better outcomes. | Leaders First policy: An Army requirement that a female officer be assigned to a combat company before or alongside female junior enlisted soldiers to ensure female leadership presence at integration; adherence was not universal, with 17% of integrating companies not following it within one month. | Defense Organizational Climate Survey (DEOCS): A congressionally mandated, annually administered, anonymous survey of military unit members covering organizational effectiveness, equal opportunity, and sexual assault prevention and response; the source of workplace perception outcomes. | Pollution theory of discrimination: Goldin&amp;rsquo;s (2014) theory that men may seek to exclude women from occupations because women&amp;rsquo;s presence is perceived to diminish the occupation&amp;rsquo;s prestige or status, potentially leading to retaliatory or unproductive behaviors among incumbent male workers. | Perception-performance wedge: The paper&amp;rsquo;s central finding that men&amp;rsquo;s subjective workplace quality perceptions decline with integration—especially when a female officer is present—even as objective administrative performance and behavior metrics show null to positive effects, a divergence between attitudes and measurable outcomes.&lt;/p&gt;</description></item><item><title>The Effects of Mandatory Profit-Sharing on Workers and Firms</title><link>https://macropaperwarehouse.com/papers/the-effects-of-mandatory-profit-sharing-on-workers-and-firms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effects-of-mandatory-profit-sharing-on-workers-and-firms/</guid><description>&lt;p&gt;This paper studies the causal effects of mandatory profit-sharing on workers and firms using a quasi-experimental design arising from a 1990 French reform that lowered the eligibility threshold for mandatory profit-sharing from 100 to 50 employees. The institutional setting is the French RSP (Réserve Spéciale de Participation), a profit-sharing scheme in place since 1967 that requires firms above the threshold to distribute a fraction of their excess profits — defined as net income above 5% of book equity — to employees according to a formula scaled by the firm&amp;rsquo;s labor share. For the median firm, this amounts to roughly 10.5% of pre-tax income transferred to workers.&lt;/p&gt;
&lt;p&gt;The authors employ two primary empirical strategies. First, a bunching analysis exploits the pre-reform distribution of firm employment around the 100-employee threshold as a revealed-preference test of whether firms perceive profit-sharing as a net cost. Second, a difference-in-differences design compares treated firms (55–85 employees in 1989–1990, who become newly subject to the regulation after 1991) against two control groups: small firms (35–45 employees, likely never subject) and large firms (120–300 employees, already subject). Data come from the universe of French corporate tax files (FICAS) and a linked employer-employee panel (DADS) covering approximately 4% of private-sector workers, spanning 1985–1997.&lt;/p&gt;
&lt;p&gt;The bunching analysis documents a 22.3% excess density in the 95–99 employee bin before the reform, which disappears after 1991. Three tests — comparing wage bills per employee across the threshold, cross-checking with DADS employment records, and examining profitability patterns — collectively support the conclusion that bunching reflects genuine employment reductions rather than under-reporting. The implied employment loss is approximately 1.67% of total employment among affected firms.&lt;/p&gt;
&lt;p&gt;The difference-in-differences results yield the following firm-level findings: (a) the total compensation share (wages plus profit-sharing divided by value added) rises by 1.8 percentage points for firms with positive excess profits; (b) 77% of this increase comes at the expense of firm owners — the profit share falls by 1.37 percentage points; (c) the remainder is borne by the government through a reduction in the corporate income tax share; (d) the wage share (base wages only) is unaffected, indicating that owners do not reduce wages to offset the cost of profit-sharing; (e) investment and total factor productivity show no statistically significant change — effects on productivity are bounded below ±1% for several TFP measures; and (f) the capital-labor ratio shows a small, mostly insignificant negative effect, consistent with a model-implied increase in the cost of capital of only 0.43 percentage points.&lt;/p&gt;
&lt;p&gt;Worker-level analysis using the linked employer-employee data confirms that average total compensation rises by approximately 3.5% for workers in treated firms, with no decline in base wages. Critically, this average conceals distributional heterogeneity across the skill spectrum. For low- and medium-skill workers (blue-collar workers, clerks, supervisors, skilled technicians), total compensation rises while base wages are unchanged — consistent with wage rigidity binding for these groups. For high-skill workers (managers, engineers, executives), base wages fall by enough to leave total compensation unchanged, consistent with more flexible wages at the upper end of the skill distribution. This pattern implies that mandatory profit-sharing is a progressive policy within firms, redistributing excess profits predominantly to lower-skill workers.&lt;/p&gt;
&lt;p&gt;The paper concludes that France&amp;rsquo;s mandatory profit-sharing scheme, as implemented, functions as a non-distortive redistributive tool: it transfers excess profits from shareholders to lower-skill workers without generating measurable productivity losses or large investment distortions. The fiscal cost is non-trivial: each dollar transferred to workers costs approximately 20 cents in foregone corporate income tax. The scheme also has an inherent inequality in its redistribution since it exclusively benefits workers in profitable firms, and firms&amp;rsquo; excess profits are highly persistent.&lt;/p&gt;
&lt;p&gt;Q: What is the French RSP and how does the formula work?
A: The RSP (Réserve Spéciale de Participation) is a mandatory profit-sharing fund established by executive order in 1967. The formula is RSP = 0.5 × (wage bill / value added) × max(net income − 5% × book equity, 0). The 5% deduction represents lawmakers&amp;rsquo; view of fair compensation to shareholders; any excess is split between shareholders and workers, with the split scaled by the firm&amp;rsquo;s labor share. For the median firm in the sample — ROE of 12%, labor share of 0.52, corporate tax rate of 37% — the formula yields roughly 9.5% of pre-tax income, and in post-1991 data the realized average is 10.5% of pre-tax income for firms with positive excess profits.&lt;/p&gt;
&lt;p&gt;Q: Why can&amp;rsquo;t a standard regression discontinuity be used at the 100-employee threshold?
A: Because firms strategically control their position relative to the threshold — the bunching analysis itself demonstrates this. When firms sort non-randomly around the cutoff, the local randomization assumption underlying RD is violated. The authors instead use a difference-in-differences design exploiting the time variation introduced by the 1990 reform.&lt;/p&gt;
&lt;p&gt;Q: How large is the pre-reform bunching and what does it imply?
A: The distribution of employment shows 22.3% excess density in the 95–99 employee bin relative to the post-reform counterfactual distribution. Interpreting this as real employment reduction (supported by three empirical tests), the implied employment loss is approximately 1.67% of total employment among firms in the 85–120 employee range. Dynamic bunching analysis shows this is persistent rather than temporary — the 100-employee threshold significantly constrained three-year employment growth for firms in the 85–99 range in the pre-reform period.&lt;/p&gt;
&lt;p&gt;Q: How do the authors establish that bunching is real rather than under-reporting of employment?
A: Three tests are conducted. First, wage bills per employee show no discontinuity around the 100-employee threshold in either period, ruling out systematic under-reporting of headcount while truthfully reporting wages. Second, employment from DADS payroll records — harder to manipulate — shows only a statistically insignificant gap of roughly 0.5 employees relative to tax-file employment just below the threshold, far too small to shift firms across the 100-employee bin. Third, profitability and value added per employee are significantly higher just below the threshold, consistent with more profitable firms having stronger incentives to bunch through genuine employment reductions.&lt;/p&gt;
&lt;p&gt;Q: What is the main identification strategy for the firm-level analysis?
A: A difference-in-differences design where treated firms have 55–85 employees in both 1989 and 1990 (newly subject to the mandate after 1991), compared to small control firms with 35–45 employees (likely never subject) and large control firms with 120–300 employees (likely always subject). Specifications include firm fixed effects and county-by-year and industry-by-year fixed effects. Parallel pre-trends are confirmed graphically and in event-study regressions. The design is intent-to-treat: by 1997, 26.7% of treated firms had shrunk below 50 employees and did not actually pay profit-sharing. LATE estimates are obtained via 2SLS.&lt;/p&gt;
&lt;p&gt;Q: What are the main firm-level findings on compensation and profit shares?
A: For treated firms with positive excess profits, the total compensation share rises by 1.8 percentage points. The wage share (base wages only, excluding profit-sharing) is precisely estimated at zero — owners do not reduce wages. The profit share falls by 1.37 percentage points, accounting for 77% of the increase in total compensation. The remaining approximately 23% is borne by the tax authority through a reduction in the corporate income tax share, since profit-sharing reduces the corporate income tax base. These findings are robust to balanced vs. unbalanced samples and to alternative control group definitions.&lt;/p&gt;
&lt;p&gt;Q: Does mandatory profit-sharing raise or lower firm productivity?
A: Across five different TFP estimators (Olley-Pakes, Olley-Pakes with Ackerberg-Caves-Frazer correction, Wooldridge, Levinsohn-Petrin, and Ackerberg-Caves-Frazer), the effect of mandatory profit-sharing on productivity is a precisely estimated zero. For several measures, effects larger than ±1% in magnitude can be rejected. Softer measures of effort — sick leave rates and the probability of working extra hours — also show no significant change. This null finding contrasts with the literature on voluntary profit-sharing adoption, which typically finds 3–5% productivity gains, likely reflecting selection bias in that literature.&lt;/p&gt;
&lt;p&gt;Q: Does mandatory profit-sharing distort investment?
A: The effect on investment is small and mostly statistically insignificant. The theoretical model shows why: the profit-sharing formula is based on excess profits (net income minus 5% of book equity), not total profits. When the firm&amp;rsquo;s actual cost of equity approximately equals the regulatory 5% benchmark, the distortion to the cost of capital is zero. The calibrated distortion to the user cost of capital is only 0.43 percentage points — approximately 1.9% of the standard user cost — implying an investment ratio reduction of about 0.84 percentage points using estimated elasticities from Chodorow-Reich et al. (2024). Empirically, capital-labor ratios show a small, largely insignificant negative effect.&lt;/p&gt;
&lt;p&gt;Q: How does profit-sharing incidence differ across the skill distribution?
A: The worker-level DADS analysis reveals that the average 3.5% increase in total compensation masks sharp heterogeneity. For low- and medium-skill workers (blue-collar workers, clerks, supervisors, skilled technicians), total compensation rises while base wages are unchanged. For high-skill workers (managers, engineers, executives), base wages decline sufficiently to leave their total compensation unchanged. The authors interpret this pattern as consistent with wage rigidity being more binding for lower-skill workers — due to the federal minimum wage and collective agreements — than for managers whose pay is more flexibly set.&lt;/p&gt;
&lt;p&gt;Q: Why does profit-sharing not affect base wages for low-skill workers?
A: Two candidate explanations are considered. The risk channel — that profit-sharing is risky and thus less valuable to risk-averse workers, who demand wage compensation — is rejected empirically because profit-sharing only marginally increases the variability of workers&amp;rsquo; total earnings. The wage rigidity channel is supported: France&amp;rsquo;s binding federal minimum wage and widespread collective agreements constrain downward adjustment in base wages for lower-skill workers, so firms cannot pass through profit-sharing costs as lower wages for this group.&lt;/p&gt;
&lt;p&gt;Q: What is the fiscal cost of the profit-sharing scheme?
A: Each dollar transferred to workers through mandatory profit-sharing costs approximately 20 cents in reduced corporate income tax receipts, since profit-sharing payments are deductible from taxable income. The paper notes this is a partial fiscal evaluation; a full assessment would also require analyzing personal income tax implications, which are left for future work.&lt;/p&gt;
&lt;p&gt;Q: How does this scheme compare to a corporate income tax as a redistributive tool?
A: Both instruments reduce firm profits and can benefit workers, but differ in three key respects. First, the tax base differs: profit-sharing targets excess profits above 5% of book equity whereas the corporate income tax applies to all corporate earnings, generating different distortions to investment. Second, profit-sharing goes directly to workers in the same firm, whereas corporate tax revenues are redistributed through general government spending — making the incidence more direct and more closely monitored by workers. Third, workers have stronger incentives to monitor firm compliance with profit-sharing (each euro of diverted excess profit reduces workers&amp;rsquo; collective income by roughly 10–15 cents) than with corporate taxes.&lt;/p&gt;
&lt;p&gt;Q: How does this paper compare to findings on mandatory profit-sharing in Peru?
A: Tolentino (2022) studies a mandatory profit-sharing scheme in Peru exploiting a 20-employee eligibility threshold and finds larger distortions — reductions in both investment and productivity. The authors attribute this difference to two features: the Peruvian scheme applies to the entirety of post-tax profits rather than excess profits above an equity deduction, creating a broader and more distortionary base; and there is pre-existing bunching at the Peruvian threshold even before the scheme was introduced, suggesting confounding pre-existing regulations.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions on the external validity of the findings?
A: The findings apply specifically to mandatory profit-sharing under the French RSP formula — which exempts a 5% equity return from the profit-sharing base, limiting distortions — during 1985–1997, for firms in the 55–300 employee range. The null productivity effect may not generalize to voluntary schemes, where selection on anticipated gains likely produces positive correlations. The redistributive finding (benefiting lower-skill workers) is specific to a context with binding minimum wages and collective agreements that constrain wage adjustment for that group. The fiscal cost calculation also excludes personal income tax effects.&lt;/p&gt;
&lt;p&gt;Excess profits: Defined in the paper as net income minus 5% of book equity — the amount above what lawmakers considered fair compensation to shareholders. Only excess profits (not total profits) are subject to the mandatory profit-sharing formula.&lt;/p&gt;
&lt;p&gt;RSP formula (Réserve Spéciale de Participation): The statutory formula RSP = 0.5 × (wage bill / value added) × max(net income − 5% × book equity, 0), scaled by the firm&amp;rsquo;s labor share to reflect labor&amp;rsquo;s contribution to production. Unchanged since 1967.&lt;/p&gt;
&lt;p&gt;Total compensation share: The ratio of (wage bill plus profit-sharing) to value added — the paper&amp;rsquo;s primary measure of workers&amp;rsquo; overall claim on firm output, as distinct from the wage share (wage bill alone divided by value added).&lt;/p&gt;
&lt;p&gt;Wage incidence parameter (λ): The fraction of profit-sharing that firms pass through to workers as lower base wages. λ = 1 means full incidence (workers&amp;rsquo; total compensation unchanged); λ = 0 means no incidence (workers fully benefit). The paper&amp;rsquo;s empirical findings are consistent with λ ≈ 0 for low-skill workers and λ ≈ 1 for high-skill workers.&lt;/p&gt;
&lt;p&gt;Bunching: The empirical phenomenon whereby firms cluster employment just below the 100-employee regulatory threshold to avoid mandatory profit-sharing. The paper uses the pre- vs. post-reform shift in the employment distribution as a revealed-preference test of whether firms perceive the scheme as a net cost.&lt;/p&gt;
&lt;p&gt;Intent-to-treat (ITT) design: The empirical design comparing firms that were in the newly eligible size range (55–85 employees) just before the 1990 reform against firms that were either always or never eligible, regardless of whether treated firms actually ended up paying profit-sharing post-reform. LATE estimates are obtained via 2SLS to recover effects on actual compliers.&lt;/p&gt;
&lt;p&gt;Distortion to user cost of capital: The additional cost of capital induced by profit-sharing, equal to ϕ × γ(1−λ) / [1 − γ(1−τ)] × (re − ρ), where ρ = 5% is the regulatory equity benchmark. When the firm&amp;rsquo;s actual cost of equity equals the 5% benchmark, this distortion is zero — a feature that distinguishes the French scheme from a standard corporate income tax.&lt;/p&gt;</description></item><item><title>The Effects of Medical Debt Relief: Evidence from Two Randomized Experiments</title><link>https://macropaperwarehouse.com/papers/the-effects-of-medical-debt-relief-evidence-from-two-randomized-experiments/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effects-of-medical-debt-relief-evidence-from-two-randomized-experiments/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks whether relieving downstream medical debt — debt that has been sold to third-party debt collectors — causes improvements in financial outcomes, mental and physical health, and healthcare utilization for recipients. The question is motivated by a large correlational literature documenting strong associations between medical debt and adverse outcomes, and by the rapid expansion of government and private debt relief programs that, as of mid-2024, had committed or planned over $14.6 billion in relief.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Design&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors partnered with RIP Medical Debt (a non-profit that purchases and forgives medical debt for government and private donors) to conduct two randomized controlled trials between March 2018 and October 2020. In total the experiments relieved medical debt with a face value of $169 million for 83,401 people.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Hospital debt experiment&lt;/strong&gt;: RIP purchased a random subset of debt from a large for-profit hospital system at the juncture when the hospital would normally sell accounts to a debt collector (approximately one year after the medical service). The purchase price was 5.5 cents per dollar of face value. The treatment group consisted of 14,377 people who received $19 million in face-value relief (average of $1,321 per person). The 61,496-person control group had their debt pursued by the collector under normal protocol.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Collector debt experiment&lt;/strong&gt;: RIP purchased a random subset of older debt already under collection on the secondary market for several years, at a price of less than one cent per dollar. The treatment group consisted of 69,024 people who received $150 million in face-value relief (average of $2,167 per person). The 68,014-person control group retained their debt.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Credit reporting sub-experiment&lt;/strong&gt;: Partway into the collector debt experiment, the debt collector ceased reporting medical debt to the credit bureaus, reflecting an industry-wide trend. The authors isolate 2,761 accounts (6.8% of wave 1) that were reported prior to treatment assignment to estimate the effects of debt relief when accounts would have been counterfactually reported, compared to the subsequent no-reporting environment.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Outcomes are tracked using quarterly depersonalized credit bureau data from TransUnion (spanning at least four quarters before to four quarters after treatment), collections account data on future bill accrual, and a multimodal survey of 2,888 hospital debt experiment respondents measuring mental and physical health, healthcare utilization, and financial wellness. The primary credit-bureau outcome is the number of accounts past due; the primary survey outcome is the share with at least moderate depression (PHQ-8).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Credit market outcomes (main experiments)&lt;/strong&gt;: In both the hospital and collector debt experiments — where there is no counterfactual credit bureau reporting — debt relief has no average effect on financial distress, credit access, or credit utilization. The effect on the number of accounts past due is -0.01 (statistically insignificant; 95% CI excludes effects smaller than -0.04, relative to a control mean of 1.20). Effects on credit card balances (95% CI: -$42 to $47 relative to a mean of $1,481) and auto loan balances (95% CI: -$235 to $148 relative to a mean of $8,020) are similarly precise nulls. These null effects hold for the hospital debt sample (younger debt, 1.3 years old on average) and the collector debt sample (older debt, 7.0 years old on average), and across all preregistered subgroups.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Credit reporting sub-experiment&lt;/strong&gt;: When control group accounts are counterfactually reported, debt relief immediately raises credit scores by an economically small average of 3.4 points (p-value 0.021), with a larger 13.8-point increase (p-value 0.008) for persons with no other debt in collections. Credit limits grow gradually, reaching $340 (15.3% of the post-reporting control mean of $2,231; p-value 0.010) after the no-reporting period begins, with larger effects for those with no other debt in collections. Once control group reporting ceases, both the credit score and credit limit effects converge to zero for those with other debts in collections. No effects on borrowing or financial distress measures are detected in this sub-experiment.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Collections account outcomes (bill repayment)&lt;/strong&gt;: Debt relief causes a statistically significant 1.1 percentage-point increase in the probability of having another unpaid bill sent to collections (6.6% of the control mean of 16.2%; p-value &amp;lt; 0.05) and a $15 increase in the dollar amount of future medical debt sent to collections (7.2% of the control mean of $208). The increase is almost entirely attributable to pre-relief medical services, indicating reduced repayment of existing bills rather than greater healthcare utilization.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Survey outcomes&lt;/strong&gt;: There are no detectable average effects on depression (primary outcome), anxiety, stress, subjective well-being, or general health. Debt relief raises the share with at least moderate depression by a statistically insignificant 3.2 percentage points (p-value 0.097; control mean 45.0%); a 95% CI rules out a reduction of more than 0.6 percentage points, well below the 7.0 percentage-point improvement predicted by the median expert respondent. There are similarly null effects on healthcare utilization and financial wellness as measured in the survey.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The study focuses specifically on downstream medical debt in collections — debt that has already been through the hospital billing cycle and sold to third-party collectors. Results do not necessarily apply to upstream debt relief (e.g., financial assistance programs applied closer to the time of the medical event), nor to populations with different baseline financial profiles. The credit reporting results are most relevant to the prior regime of widespread reporting; under the current environment in which most medical debt has been removed from credit reports, the credit-access channel is largely foreclosed.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-did-the-authors-focus-specifically-on-downstream-medical-debt-in-collections-and-how-does-this-define-the-scope-of-their-study"&gt;Q1. Why did the authors focus specifically on downstream medical debt in collections, and how does this define the scope of their study?&lt;/h3&gt;
&lt;p&gt;The authors focus on downstream medical debt because this is the target of essentially all large-scale government and private relief programs working with RIP Medical Debt, and because it is the category of debt that is most comprehensively observable. Downstream medical debt is defined as bills that have been or are about to be sold by the healthcare provider to a third-party debt collector. This focus excludes upstream unpaid bills still held by the hospital, bills being paid over time, and medical expenses charged to credit cards. The distinction matters because prior literature on hospital financial assistance programs finds substantial benefits from upstream interventions that relieve debt closer to the precipitating medical event; the authors&amp;rsquo; null results are explicitly scoped to the downstream, post-collection stage.&lt;/p&gt;
&lt;h3 id="q2-why-did-the-purchase-price-of-medical-debt-55-cents-per-dollar-for-hospital-debt-less-than-1-cent-per-dollar-for-collector-debt-suggest-caution-about-expected-financial-impacts-ex-ante"&gt;Q2. Why did the purchase price of medical debt (5.5 cents per dollar for hospital debt, less than 1 cent per dollar for collector debt) suggest caution about expected financial impacts ex ante?&lt;/h3&gt;
&lt;p&gt;The authors argue that in a competitive market, the purchase price of medical debt reflects the sum of expected recovery rates and collection costs. A price of 5.5 cents per dollar implies that actual recovery (what collectors expect to collect from patients) is very low. Even if all of the expected recovery is passed through to the patient as a financial benefit, the direct liquidity gain from debt forgiveness is a small fraction of the debt&amp;rsquo;s face value. For the collector debt experiment, where the purchase price is less than 1 cent per dollar, the expected direct financial benefit to recipients is even smaller. The authors note that survey respondents expected to pay 54% of their outstanding medical debt and thought it fair to pay 37%, suggesting that perceived (rather than actual) payment obligations may be what connects medical debt to financial behavior.&lt;/p&gt;
&lt;h3 id="q3-how-was-random-assignment-implemented-in-the-hospital-debt-experiment-and-what-design-features-ensure-the-validity-of-the-experiment"&gt;Q3. How was random assignment implemented in the hospital debt experiment, and what design features ensure the validity of the experiment?&lt;/h3&gt;
&lt;p&gt;Within each of 18 waves between August 2018 and October 2020, RIP received a portfolio of unpaid bills from the hospital system. Persons were grouped at the individual level and stratified by the amount of debt, state of residence, insurance status, and a collections score predicting repayment likelihood. Within strata, persons were randomly assigned to treatment or control, with approximately 20% treated per wave (varying with donor funding). The hospital was unaware of the intervention, eliminating scope for selection of particularly uncollectible accounts. Treatment notification occurred via two letters sent approximately three and six weeks post-purchase. Balance tests confirm successful randomization: all p-values on baseline characteristics are above 0.05, and F-tests fail to reject joint balance.&lt;/p&gt;
&lt;h3 id="q4-what-was-the-credit-reporting-sub-experiment-and-how-was-it-identified"&gt;Q4. What was the credit reporting sub-experiment and how was it identified?&lt;/h3&gt;
&lt;p&gt;The debt collector in the collector debt experiment historically reported medical debt to the credit bureaus but largely ceased doing so before the first intervention wave (March 2018), reflecting broader industry concerns about CFPB enforcement and data integrity risk. However, a subset of accounts — 2,761 accounts (6.8% of wave 1, with virtually identical match rates across treatment and control) — were still being reported until 2019 Q1 (three quarters after wave 1 and one quarter after wave 2). This created a natural sub-experiment: for this subset, treatment group accounts were removed from credit reports immediately upon debt relief, while control group accounts continued to be reported for three more quarters before also being removed. The authors identify reported accounts by matching dollar amounts in collections account data to credit bureau tradeline data in the four quarters prior to intervention, and use this variation to estimate effects separately for the &amp;ldquo;reporting&amp;rdquo; and &amp;ldquo;no-reporting&amp;rdquo; periods.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-exact-estimated-effects-on-credit-scores-and-credit-limits-in-the-credit-reporting-sub-experiment"&gt;Q5. What are the exact estimated effects on credit scores and credit limits in the credit reporting sub-experiment?&lt;/h3&gt;
&lt;p&gt;During the three quarters when control group accounts are still reported to credit bureaus, debt relief raises credit scores by an average of 3.4 points (p-value 0.021) for the full reporting subsample. The effect is concentrated among those with no other debt in collections: 13.8 points (p-value 0.008) versus 1.2 points (p-value 0.440) for those with other debt in collections. Credit limits increase gradually, reaching $340 (15.3% of the post-reporting control mean of $2,231; p-value 0.010) by the four quarters after control group reporting ceases. Among persons with no other debt in collections, this credit limit effect grows to $922 (23% of the control mean; p-value 0.070). Once control group reporting stops, both the credit score effect and the credit limit growth converge to zero for persons with other debts in collections. The event study coefficients show the credit limit effect growing approximately linearly over five quarters post-intervention before leveling out.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-rule-out-the-possibility-that-medical-debt-relief-increases-healthcare-utilization-thereby-causing-more-future-medical-bills"&gt;Q6. How does the paper rule out the possibility that medical debt relief increases healthcare utilization, thereby causing more future medical bills?&lt;/h3&gt;
&lt;p&gt;The collections account analysis separates future debt accrual into debt associated with pre-relief medical services (which can only result from reduced repayment of existing bills) and post-relief medical services (which could reflect either increased utilization or changed repayment of new bills). Panel B of Table VI shows that virtually all of the increased debt sent to collections — a $15 increase and 1.1 percentage-point increase in the probability of any future collection — is attributable to pre-relief services. Panel C shows statistically insignificant increases in future debt from post-relief services. The authors therefore attribute the effect to reduced payment of existing bills and conclude they &amp;ldquo;cannot rule in or rule out effects on healthcare utilization&amp;rdquo; for the post-relief services channel, but the dominant mechanism is behavioral change in repayment of already-incurred debt.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-three-mechanisms-proposed-to-explain-the-reduction-in-repayment-of-existing-medical-bills-and-which-mechanism-is-rejected"&gt;Q7. What are the three mechanisms proposed to explain the reduction in repayment of existing medical bills, and which mechanism is rejected?&lt;/h3&gt;
&lt;p&gt;The authors offer three candidate mechanisms for the 6.6% relative increase in the probability of future bill collections: (i) an expectations mechanism, in which beneficiaries reduce payments because they anticipate future debt relief from similar charitable programs; (ii) a targeting mechanism, drawing on Dobkin et al. (2018), in which patients tolerate a certain level of indebtedness — relieving some debt creates &amp;ldquo;room&amp;rdquo; in their debt budget, so they reduce payment of remaining bills to return to that target level; and (iii) a confusion mechanism, in which recipients mistakenly believe the relief applied to non-forgiven bills (the notification letter explicitly stated &amp;ldquo;the forgiveness is for this outstanding bill only&amp;rdquo; but patients may not have internalized this). The income effect or &amp;ldquo;flypaper&amp;rdquo; mechanism — the idea that financial relief of existing debt frees up mental-account resources for paying medical bills, thereby increasing repayment — is explicitly rejected by the data, as the effect goes in the direction of less repayment, not more.&lt;/p&gt;
&lt;h3 id="q8-what-did-the-expert-survey-predict-and-how-did-those-predictions-compare-to-the-experimental-estimates"&gt;Q8. What did the expert survey predict, and how did those predictions compare to the experimental estimates?&lt;/h3&gt;
&lt;p&gt;An expert survey conducted between April and May 2022 — after the interventions were completed but before results were released — asked academics, non-profit staff, hospital revenue-cycle practitioners, and policymakers to predict the impact of the hospital debt experiment. The median expert predicted a 7.0 percentage-point reduction in depression (8.0 points when weighted by confidence), a 10.2 percentage-point reduction in borrowing (13.7 points when confidence-weighted), and meaningful improvements in healthcare access. In total, 75.6% of respondents predicted medical debt relief is at least a moderately valuable use of charity resources, and 51.1% thought it very or extremely valuable. The authors estimate a statistically insignificant 3.2 percentage-point increase in depression (not a decrease), and a 95% confidence interval that rules out a reduction in depression of more than 0.6 percentage points — far below the 7.0 percentage-point expert prediction.&lt;/p&gt;
&lt;h3 id="q9-what-survey-methodology-was-used-and-what-response-rate-was-achieved"&gt;Q9. What survey methodology was used, and what response rate was achieved?&lt;/h3&gt;
&lt;p&gt;The survey, administered by NORC at the University of Chicago, targeted a random subset of 14,922 hospital debt experiment participants who entered the study after September 2019 (waves 6-18) and owed at least $500. The protocol spanned 13 weeks and included five postal mailings (including a $2 upfront incentive and a $5 incentive with the paper survey), twice-weekly email reminders, certified mail delivery of the full survey instrument, and telephone interviews by a US-based call center. Respondents received a $50 completion incentive. The protocol achieved a 19.4% response rate, with 68% responding via web, 10% via telephone, and 23% via mail. The survey was titled &amp;ldquo;Health and Financial Wellness Study&amp;rdquo; and made no reference to RIP Medical Debt to avoid priming respondents. Respondents were surveyed on average 13 months after treatment assignment (interquartile range 10 to 17 months).&lt;/p&gt;
&lt;h3 id="q10-what-heterogeneity-in-survey-outcomes-was-detected-and-how-do-the-authors-interpret-the-anomalous-depression-finding-for-high-debt-recipients"&gt;Q10. What heterogeneity in survey outcomes was detected, and how do the authors interpret the anomalous depression finding for high-debt recipients?&lt;/h3&gt;
&lt;p&gt;Across all four preregistered heterogeneity dimensions (medical debt amount, age of debt, age of person, amount of other debt in collections), null effects on survey outcomes were found in 15 of 16 subgroups. The exception is persons in the fourth quartile of medical debt eligible for relief, for whom debt relief caused a statistically significant 12.4 percentage-point increase in depression (p-value 0.002) relative to a control mean of 45.9%, with similar patterns for anxiety, stress, subjective well-being, and general health. The authors consider this may be a statistical fluke given the null results across all other 15 groups. They also note potential parallels with findings from unconditional cash transfer experiments, where the receipt of transfers raised the salience of financial deprivation without addressing its underlying causes. A charity-stigma mechanism (recipients did not request the assistance) is also considered. The authors caution against giving this result undue weight in the overall assessment.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-position-downstream-debt-relief-relative-to-upstream-interventions-and-what-does-prior-evidence-suggest-about-upstream-alternatives"&gt;Q11. How does the paper position downstream debt relief relative to upstream interventions, and what does prior evidence suggest about upstream alternatives?&lt;/h3&gt;
&lt;p&gt;The authors highlight that their null results do not extend to upstream medical debt relief. Adams et al. (2022), studying a hospital financial assistance program at Kaiser Permanente that bundled debt relief with reductions in cost-sharing close to the time of the medical event, found substantial increases in high-value healthcare utilization. The Oregon Health Insurance Experiment (Baicker et al. 2013) found that Medicaid reduced depression by 9 percentage points among low-income uninsured adults. The authors suggest several reasons why downstream relief may fail: the intervention occurs too late after the precipitating event (approximately 15 months after the medical service in the hospital debt experiment, and about 7 years in the collector debt experiment), patients may have habituated to the stress of debt collections, the relief amount may be too small relative to overall financial distress, and the direct financial benefit is inherently limited by the low market price of collections-stage debt.&lt;/p&gt;
&lt;h3 id="q12-how-do-the-authors-address-concerns-about-differential-survey-response-and-external-validity"&gt;Q12. How do the authors address concerns about differential survey response and external validity?&lt;/h3&gt;
&lt;p&gt;Treated persons were a statistically insignificant 1.3 percentage points more likely to respond to the survey (p-value 0.056). The authors address this in two ways. First, they estimate specifications that (i) add rich observable controls and (ii) use speed of survey response as a proxy for unobserved response propensity; neither exercise changes the estimates meaningfully. Second, to probe external validity, they test for heterogeneous effects by predicted response propensity (from a logistic regression of a response indicator on baseline characteristics) and by speed of response; neither yields evidence of differential effects for non-respondents. They also compare credit bureau treatment effects for the full hospital debt sample, the survey outreach sample, and the survey respondent sample and find similar estimates across all three groups.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Downstream medical debt&lt;/strong&gt;: Medical bills that have already been sent to third-party debt collectors by the healthcare provider after the initial billing cycle, as distinguished from upstream unpaid bills still held by the hospital at or near the time of the medical event. The paper studies debt at this late stage specifically because it is the target of most large-scale relief programs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Credit reporting sub-experiment&lt;/strong&gt;: An embedded quasi-experiment within the collector debt RCT, exploiting the fact that a subset of accounts (6.8% of wave 1) were still being reported to credit bureaus at the time of intervention while the debt collector had already ceased reporting for the remaining accounts. This allows separate estimation of debt relief effects with and without counterfactual credit bureau reporting, using the period until 2019 Q1 (when the collector stopped reporting entirely) as the &amp;ldquo;reporting&amp;rdquo; window.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Downstream bill repayment effect&lt;/strong&gt;: The paper&amp;rsquo;s finding that debt relief increases the probability of a subsequent unpaid medical bill being sent to collections. The paper attributes this primarily to reduced repayment of existing pre-relief medical bills rather than to increased healthcare utilization, consistent with an expectations, targeting, or confusion mechanism — and inconsistent with an income or flypaper effect that would increase repayment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Targeting a level of indebtedness&lt;/strong&gt;: A behavioral model (drawn from Dobkin et al. [2018]) in which patients implicitly target a certain level of indebtedness. Under this model, relieving some debt creates headroom in the patient&amp;rsquo;s implicit debt budget, leading to reduced repayment of remaining bills to restore the targeted level of total indebtedness.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Expert survey (pre-results)&lt;/strong&gt;: A structured elicitation of predicted treatment effects conducted between April and May 2022 — after the interventions were completed but before results were released — from academics, non-profit practitioners, hospital revenue-cycle managers, and policymakers. Used as a benchmark to quantify how far the causal estimates fall below prevailing beliefs, and to document that the null results were ex ante surprising to informed observers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PHQ-8 (Patient Health Questionnaire-8)&lt;/strong&gt;: An eight-item validated clinical screen for depression, used as the paper&amp;rsquo;s primary preregistered survey outcome. An indicator for &amp;ldquo;at least moderate depression&amp;rdquo; on the PHQ-8 is the main mental health measure against which the debt relief treatment effect is estimated.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Multimodal survey&lt;/strong&gt;: A survey protocol combining five postal mailings, twice-weekly email reminders, certified mail delivery of a paper survey instrument, and US-based call center telephone interviews, designed to maximize response rates in a hard-to-reach low-income population with medical debt in collections.&lt;/p&gt;</description></item><item><title>The Environmental Bias of Corporate Income Taxation</title><link>https://macropaperwarehouse.com/papers/the-environmental-bias-of-corporate-income-taxation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-environmental-bias-of-corporate-income-taxation/</guid><description>&lt;p&gt;This paper documents and quantifies an &amp;ldquo;environmental bias&amp;rdquo; embedded in the U.S. corporate income tax code: CO2-intensive (&amp;ldquo;dirty&amp;rdquo;) firms systematically face lower effective tax rates than clean firms, constituting an implicit subsidy on pollution. The authors — Iovino, Martin, and Sauvagnat — establish this cross-sectional fact, trace it to a specific mechanism, provide causal evidence using the 2017 Tax Cuts and Jobs Act (TCJA), and quantify aggregate emissions implications using a calibrated multi-sector general-equilibrium model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and sample.&lt;/strong&gt; The empirical analysis combines firm-level CO2 emissions from Trucost (scope 1 greenhouse gases) with financial data from Compustat North America for U.S. publicly listed firms, 2003–2021, yielding 11,223 firm-year observations with positive pretax and gross capital income. Effective tax rates are measured as income taxes paid divided by gross capital income (sales minus COGS minus SGA expenses, adding back R&amp;amp;D).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cross-sectional finding.&lt;/strong&gt; A one-standard-deviation increase in CO2 intensity is associated with a decrease in the effective tax rate equal to approximately 9% of its standard deviation (coefficient −0.021 to −0.022, significant at 1%). The negative relationship is entirely explained by the lower taxable fraction of gross capital income for dirty firms — that is, by larger interest expense deductions — rather than by differences in the statutory tax rate applied to pretax income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanism.&lt;/strong&gt; The chain of causation runs: CO2-intensive production requires tangible capital (primarily machinery and equipment) → tangible capital serves as collateral → higher collateral supports higher debt → higher debt generates larger interest deductions (the &amp;ldquo;tax shield of debt&amp;rdquo;) → lower effective tax rates. Once PPE-to-capital-income is controlled for, the coefficient on CO2 intensity in leverage, pretax income, and tax regressions becomes small and statistically insignificant. The relationship holds both across and within industries, including within the energy sector, though the dominant variation is cross-industry.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Causal evidence: TCJA 2017.&lt;/strong&gt; The paper exploits the federal corporate tax rate cut from 35% to 21% (effective January 2018) in a difference-in-differences design, comparing firms in the top quartile of 2017 CO2 intensity (&amp;ldquo;dirty&amp;rdquo;) to cleaner firms. Dirty firms experienced a relative increase in their federal effective tax rate of 2.4 percentage points post-reform. Correspondingly, dirty firms&amp;rsquo; total assets grew approximately 11% less than clean firms post-reform. This translates to a semi-elasticity of firm total assets to a one-percentage-point increase in the effective tax rate of approximately −4.8. Parallel pre-trends are confirmed visually and via Rambachan-Roth (2023) sensitivity analysis; a placebo using non-federal taxes shows no differential effect. Results survive controls for other TCJA provisions (interest deductibility limits, international tax changes, net operating loss restrictions), exposure to import tariffs and carbon taxes, leave-one-industry-out specifications, and a triple-difference using foreign firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;General-equilibrium model and counterfactuals.&lt;/strong&gt; A 375-sector model with input-output networks (both intermediate and investment networks), financial frictions linking equipment to debt capacity, and endogenous CO2 emissions through fossil fuel usage is calibrated to 2017 BEA and Compustat data. In the Cobb-Douglas benchmark, the 2017 tax cut raises output by 5.9% and emissions by only 4.5% — a less-than-proportional emissions response because clean sectors expand relatively more. A counterfactual eliminating the tax shield of debt while simultaneously cutting the tax rate from 35% to 30% (to hold GDP constant) reduces aggregate emissions by 1.3% with output declining only 0.1%. When equipment and fuel are treated as complements (elasticity of substitution below 1), the emissions reduction under the same policy rises to over 3.7%, implying an absolute reduction of 80–240 million metric tons of CO2 from 2017&amp;rsquo;s total of 6,457 million metric tons. Monetized at the social cost of carbon, this ranges from USD 8–24 billion (conservative, ~USD 100/ton) to USD 112–336 billion (USD 1,400/ton per Bilal and Kanzig 2024).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the central empirical finding of the paper?&lt;/strong&gt;
A: CO2-intensive firms in the U.S. face systematically lower effective corporate income tax rates than clean firms. A one-standard-deviation increase in CO2 intensity is associated with a roughly 9% of a standard deviation decrease in the ratio of taxes paid to gross capital income. This negative relationship is robust to alternative emissions measures (EPA data, scope 2 and 3 emissions), alternative tax scalings (taxes over sales or assets), log CO2 emissions, and leave-one-industry-out specifications.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the mechanism linking CO2 intensity to lower effective tax rates?&lt;/strong&gt;
A: Dirty firms rely on tangible capital — specifically machinery and equipment — to produce. Tangible capital is pledgeable as collateral, enabling higher debt. Higher debt generates larger interest expense deductions under the tax code (the &amp;ldquo;debt tax shield&amp;rdquo;), which reduces taxable income relative to gross capital income. Once PPE-to-capital-income is included as a control, the coefficient on CO2 intensity in regressions of leverage, pretax income, and taxes paid all become small and statistically insignificant, confirming that PPE fully mediates the relationship.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Which component of tangible capital drives the result?&lt;/strong&gt;
A: Machinery and equipment, not buildings, leases, land, natural resources, or construction in progress, explains virtually the entire positive relationship between PPE and CO2 intensity. This finding is based on the Compustat breakdown of PPE components available for roughly 70% of sample firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does the mechanism operate within industries or only across them?&lt;/strong&gt;
A: Both. Decomposing firm CO2 intensity into an implied industry component (sales-weighted from pure-play firms) and a firm residual, both components are significantly associated with higher tangible capital, leverage, lower taxable fraction of capital income, and lower taxes paid at the 1% level. However, the largest share of the total effect stems from cross-industry variation. Within the energy sector specifically, firms with greater fossil fuel production capacity (from EPA/EIA data) also have more tangible capital, higher debt, and lower effective tax rates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the 2017 TCJA cut affect clean versus dirty firms differently?&lt;/strong&gt;
A: Because dirty firms already shield a large fraction of their capital income from taxation via interest deductions, a uniform cut in the statutory rate benefits them less in proportional terms. The difference-in-differences estimates show that dirty firms (top quartile of 2017 CO2 intensity) experienced a relative increase in their federal effective tax rate of 2.4 percentage points post-reform compared to clean firms, and their total assets grew approximately 11% less than clean firms post-reform. The semi-elasticity of firm assets to a one-percentage-point increase in effective tax rate is approximately −4.8.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How is the parallel trends assumption supported?&lt;/strong&gt;
A: Event-study graphs show no pre-2018 divergence in federal effective tax rates or asset growth between dirty and clean firms. A placebo test using non-federal income taxes (which should be unaffected by the federal statutory rate change) shows no differential post-reform effect. The Rambachan-Roth (2023) sensitivity analysis confirms that the null of no differential effect can be rejected at the 1% level allowing for pre-trend deviations up to M = 0.5, and at the 10% level up to M = 1.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What robustness checks address other provisions of the TCJA and concurrent shocks?&lt;/strong&gt;
A: The authors exclude or control for firms affected by the TCJA&amp;rsquo;s interest deductibility limitation, multinational firms (more than 20% foreign sales), firms with large loss carryforwards, and manufacturing firms — results are unchanged. They also control for firm-level exposure to import tariff changes and carbon taxes (using the World Carbon Pricing Database), with coefficients of interest remaining virtually unchanged. Leave-one-industry-out specifications and a triple-difference using foreign firms (comparing U.S. dirty vs. clean firms pre/post-2018, against foreign equivalents in countries with stable tax rates) yield a semi-elasticity of −5.8, if anything larger than the baseline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the general-equilibrium model add that the difference-in-differences cannot?&lt;/strong&gt;
A: The DiD design identifies relative effects of the tax cut on dirty versus clean firms but cannot recover the absolute effect on aggregate output and emissions. The GE model, calibrated to 2017 data and validated against the untargeted DiD estimates, quantifies aggregate impacts: the 2017 tax cut raises steady-state output by 5.9% while emissions rise by only 4.5% — a less-than-proportional increase due to compositional reallocation toward clean sectors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the counterfactual removing the debt tax shield find?&lt;/strong&gt;
A: Eliminating the tax shield of debt while simultaneously lowering the corporate tax rate from 35% to 30% (to keep GDP constant) reduces aggregate emissions by 1.3% (Cobb-Douglas benchmark) while total output falls only 0.1% and GDP remains constant by design. The emissions reduction arises because clean sectors, which rely more on less-pledgeable capital, are made relatively cheaper once the tax advantage of debt is removed, redirecting demand away from CO2-intensive sectors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the complementarity assumption between equipment and fuel affect the results?&lt;/strong&gt;
A: When equipment and fuel are modeled as complements (elasticity of substitution below 1) rather than Cobb-Douglas substitutes, both policy counterfactuals yield larger emissions effects. For the tax shield removal policy, the predicted emissions reduction rises from 1.3% to over 3.7% as complementarity strengthens. This is because policies that raise the cost of equipment also induce firms to cut fuel consumption, amplifying the direct compositional effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the quantified absolute emissions impact of removing the tax shield?&lt;/strong&gt;
A: Given 2017 U.S. total emissions of 6,457 million metric tons, the model predicts an absolute reduction of 80–240 million metric tons of CO2, depending on the assumed complementarity between equipment and fuel. Monetized at conservative estimates (~USD 100/ton), the policy saves USD 8–24 billion; at USD 1,400/ton (Bilal and Kanzig 2024), the value rises to USD 112–336 billion. The authors note that the physical quantity measure is more reliable than the monetized figure given uncertainty in the social cost of carbon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does this paper relate to the ECB bond purchasing literature?&lt;/strong&gt;
A: Piazzesi et al. (2022) document that the ECB&amp;rsquo;s market-neutral bond purchases implicitly favor dirty firms because those firms issue more bonds due to higher tangible capital holdings. This paper identifies the same underlying mechanism — tangible capital → debt capacity — but on the tax side, showing that the corporate income tax code independently provides an implicit subsidy to dirty firms through the debt tax shield.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the policy implication for the debt tax shield specifically?&lt;/strong&gt;
A: The debt tax shield — the deductibility of interest payments but not dividends — has no clear economic rationale (both are returns to capital) and, per several policy proposals (CBO 1997, IMF 2016), is a candidate for elimination. This paper adds a new dimension: the tax shield indirectly subsidizes CO2 emissions by differentially benefiting capital-intensive, CO2-intensive sectors. A revenue-neutral reform eliminating the shield can reduce emissions without sacrificing GDP.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Effective tax rate (paper&amp;rsquo;s definition):&lt;/strong&gt; The ratio of corporate income taxes paid to gross capital income, where gross capital income equals sales minus cost of goods sold minus SGA expenses plus R&amp;amp;D spending. This differs from the tax-to-pretax-income ratio because it captures how much of total capital earnings — before any deductions — is remitted as tax.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Debt tax shield (tax advantage of debt):&lt;/strong&gt; The reduction in corporate tax liability arising from the deductibility of interest payments on corporate debt. Because dividends are not deductible, debt-financed capital faces a lower after-tax cost than equity-financed capital. The shield&amp;rsquo;s value is estimated at approximately 10% of firm value in prior literature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;CO2 intensity:&lt;/strong&gt; Metric tons of CO2 equivalent per USD 1,000 of output (tCO2/k$). The sample average is 0.1 tCO2/k$, with a heavily right-skewed distribution (median 0.02, 99th percentile 1.5).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Environmental bias of corporate taxation:&lt;/strong&gt; The paper&amp;rsquo;s central concept — the systematic difference in effective tax rates between dirty and clean firms that arises not from explicit environmental policy but from the interaction of the debt tax shield with the capital structure of CO2-intensive industries. This constitutes an implicit subsidy on pollution embedded in the corporate income tax.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Asset pledgeability (psi):&lt;/strong&gt; The fraction of a firm&amp;rsquo;s assets recoverable by creditors in the event of default. In the model, equipment has higher pledgeability than other capital (estimated b_psi = 0.23 additional pledgeability for equipment, a_psi = 0.35 base). Higher pledgeability allows firms to sustain more debt and thus benefit more from the tax shield.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;User cost of capital:&lt;/strong&gt; The total cost to a firm of using one unit of capital, combining depreciation, tax allowances from accelerated depreciation, and the financing cost advantage of debt over equity. The model formalizes how both the equity-financed component and the debt advantage component respond to tax rate changes, with the debt advantage term being larger for firms with more pledgeable (tangible) capital.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Investment network:&lt;/strong&gt; An input-output structure capturing which sectors&amp;rsquo; outputs are used to produce each type of capital good. The paper extends vom Lehn and Winberry (2021) by constructing separate equipment and non-equipment investment networks across 375 non-fuel BEA sectors, enabling emissions accounting that includes capital production alongside direct production inputs.&lt;/p&gt;</description></item><item><title>The Illiquidity of Water Markets</title><link>https://macropaperwarehouse.com/papers/the-illiquidity-of-water-markets/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-illiquidity-of-water-markets/</guid><description>&lt;p&gt;Donna and Espín-Sánchez investigate whether a market (sequential English auction) or a non-market institution (fixed quota) more efficiently allocates an intermediate good — irrigation water — when some buyers are liquidity constrained. The setting is Mula, a city in southeastern Spain, where farmers used an unregulated water auction continuously from 1244 until August 1, 1966, when the institution was replaced by a fixed quota system. This 700-year natural experiment, combined with the fact that water demand for a given crop is pinned down by the crop&amp;rsquo;s production function rather than by farmer wealth, allows the authors to separately identify liquidity constraints from unobserved heterogeneity in productivity.&lt;/p&gt;
&lt;p&gt;The empirical context has four features the authors exploit. First, the pre-1966 auction was entirely unregulated, so price differences directly reflect valuations without the confounds of regulatory changes. Second, water is an intermediate good for apricot production; conditional on plot area, tree count, and crop type, demand is determined by the apricot tree&amp;rsquo;s biological water requirements — not by the farmer&amp;rsquo;s wealth — so wealthy and poor farmers growing the same bulida apricot variety share the same underlying demand up to an idiosyncratic productivity shock. Third, farmers are classified as wealthy if they held positive urban real estate (non-agricultural wealth) in 1955 tax records; wealthy farmers&amp;rsquo; average annual urban rental income (5,702 pesetas) far exceeded their average annual water expenditure (500 pesetas, rising to 1,619 in the highest-expenditure year, 1963), supporting the assumption that wealthy farmers were never liquidity constrained. Fourth, the 1966 institutional shift to quotas — under which each farmer received a fixed water allotment (tanda) every three weeks proportional to plot size, paying only a small annual maintenance fee after the critical season — provides the counterfactual.&lt;/p&gt;
&lt;p&gt;The authors build a structural dynamic demand model with three key features: storability (irrigation raises soil moisture, creating intertemporal substitution between periods because water evaporates partially), liquidity constraints (poor farmers cannot always afford water during the critical season when prices peak), and weather seasonality (the critical season, corresponding to apricot fruit growth stages II–III and the Early Post-Harvest period, spans roughly weeks 18–32 and is when trees most need water). Farmers are forward-looking and form expectations about future prices and rainfall. The model&amp;rsquo;s production function, drawn from the agricultural engineering literature (Torrecillas et al., 2000; Allen et al., 2006), transforms soil moisture into apricot output via a transformation rate parameter gamma, a hydric stress coefficient, and a seasonal dummy.&lt;/p&gt;
&lt;p&gt;Demand parameters are estimated using a two-step conditional choice probability (CCP) estimator (Hotz et al., 1994) on wealthy farmers only, then projected onto poor farmers&amp;rsquo; welfare calculations. The sample consists of 24 single-crop apricot farmers observed in weekly auction records from January 1955 to July 1966, embedded in a market with over 500 total participants.&lt;/p&gt;
&lt;p&gt;The main finding is that the institutional change from auction to quota increased total efficiency. Welfare increased by 23.4 real pesetas per farmer per tree, a 6 percent increase in total apricot production relative to the market. This gain arises because: (1) farmers were relatively homogeneous in productivity (small idiosyncratic shocks), so the primary source of misallocation was not productivity heterogeneity but wealth heterogeneity; (2) liquidity constraints prevented poor farmers from purchasing water during the critical season when their valuation was high, causing them instead to buy earlier (at lower prices but with partial evaporation loss) or later (when their trees had already experienced hydric stress); and (3) the apricot production function is concave in water, so uniform quota allocation is more efficient than market allocation when farmers are approximately homogeneous. The paper provides the first empirical demonstration that liquidity constraints can reverse the standard efficiency ranking of markets over quotas.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question?
A: The paper asks whether a free market (water auction) or a non-market institution (fixed quota) more efficiently allocates an intermediate good when some buyers are liquidity constrained. The theoretical ranking is ambiguous when agents are heterogeneous in both productivity and wealth, making this an empirical question. The authors find that quotas dominated the auction in the specific Mula setting.&lt;/p&gt;
&lt;p&gt;Q: What was the historical water market in Mula and when did it end?
A: From 1244 to 1966 — over 700 years — Mula farmers used a sequential ascending-price (English) auction to allocate river water. The auctioneer sold water in discrete units called cuartas (each representing 3 hours of canal flow, or approximately 432,000 liters), holding 40 units per weekly Friday session. Farmers paid in cash on auction day. On August 1, 1966, the farmers&amp;rsquo; union (Sindicato de Regantes) replaced the auction with a fixed quota system, having secured a credit line to purchase water property rights share by share.&lt;/p&gt;
&lt;p&gt;Q: How did the quota system work, and how did it eliminate liquidity constraints?
A: Under the quota, each plot of land received a fixed water allotment (tanda) every three weeks, proportional to plot size. Farmers paid only a small annual maintenance fee to the Sindicato at year-end, after the critical season harvest. Because payment occurred after farmers collected harvest revenue, no farmer was liquidity constrained under the quota. The fee was substantially lower than the per-unit average price under the market.&lt;/p&gt;
&lt;p&gt;Q: How do the authors identify liquidity constraints separately from unobserved heterogeneity in productivity?
A: The key insight is that water is an intermediate good whose demand is determined by the apricot tree&amp;rsquo;s biological production function, not by farmer wealth. Two farmers growing the same bulida apricot variety with the same number of trees should have the same water demand up to an idiosyncratic shock. The authors use wealthy farmers (those with positive urban real estate in 1955 tax records) to estimate preferences, under the assumption that wealthy farmers are never liquidity constrained. They then verify that outside the critical season, wealthy and poor farmers purchase similar amounts of water; the purchasing divergence appears only during the high-price critical season, consistent with a cash constraint rather than a preference difference.&lt;/p&gt;
&lt;p&gt;Q: What empirical evidence shows poor farmers were liquidity constrained rather than simply less interested in water?
A: Poor farmers display a bimodal purchasing pattern inconsistent with the apricot tree&amp;rsquo;s biological water needs: they buy water before the critical season (when prices are low) in anticipation of not being able to afford it during the critical season, and again after the critical season (when prices fall) to prevent their trees from withering from dehydration. Wealthy farmers, by contrast, delay purchases strategically to the critical season when trees most need water (weeks 18–32). Regression analysis confirms that wealthy farmers purchase significantly more water per tree during the critical season than poor farmers growing identical bulida apricots, while the difference outside the critical season is not statistically significant.&lt;/p&gt;
&lt;p&gt;Q: How were wealthy farmers defined and why does their wealth validate the non-constrained assumption?
A: A farmer is defined as wealthy if the value of their urban real estate (from 1955 urban tax records) is positive, and as poor if it is zero. Urban real estate constitutes non-agricultural wealth uncorrelated with the apricot production function. Wealthy farmers&amp;rsquo; average annual urban rental income was 5,702 pesetas, while their average annual water expenditure was only 500 pesetas (rising to 1,619 pesetas in 1963, the highest-expenditure sample year). This large gap supports the assumption that wealthy farmers could always afford water purchases.&lt;/p&gt;
&lt;p&gt;Q: What is the model&amp;rsquo;s treatment of soil moisture dynamics and why does it matter?
A: Soil moisture (M_it) evolves according to an agricultural engineering formula: it increases with rainfall and irrigation purchases (each unit adding 432,000 liters divided by plot area) and decreases via evapotranspiration (ET), subject to a full-capacity ceiling (FC) and a permanent wilting point (PW) lower bound. This storage structure creates intertemporal substitution — water purchased early partially substitutes for future purchases, but at a cost (evaporative loss). The dynamics mean poor farmers who pre-buy water before the critical season lose some of that investment to evaporation, generating a real efficiency loss relative to the quota that delivers water closer to when it is biologically needed.&lt;/p&gt;
&lt;p&gt;Q: What are the two sources of potential inefficiency the authors identify?
A: The first is inefficiency due to heterogeneity: if farmers differ in ex-post productivity (captured by idiosyncratic shocks epsilon_it), allocating water to a less productive farmer at a given moment is wasteful. Markets correct this inefficiency (they direct water to highest-valuation buyers) while quotas do not. The second is inefficiency due to decreasing marginal returns (DMR): because the production function is concave in water, giving water to a farmer with already-high soil moisture is less productive than giving it to a farmer with low moisture. Quotas naturally avoid DMR inefficiency by allocating uniformly; markets with liquidity constraints exacerbate DMR inefficiency by directing scarce critical-season water to wealthy farmers who may have already accumulated moisture from prior purchases.&lt;/p&gt;
&lt;p&gt;Q: What is the main quantitative result of the welfare analysis?
A: Switching from the market auction to the fixed quota system increased welfare by 23.4 real pesetas per farmer per tree, representing a 6 percent increase in total apricot production relative to the market counterfactual. This is computed as the difference in yearly mean welfare per tree per farmer (net of irrigation costs, excluding water expenditures which are transfers) between the quota and market allocations using the estimated structural model.&lt;/p&gt;
&lt;p&gt;Q: Under what conditions is a quota more efficient than a market with liquidity constraints?
A: Quotas dominate markets when three conditions hold simultaneously: (1) farmers are relatively homogeneous in productivity (so the market&amp;rsquo;s advantage of directing water to high-valuation buyers is small), (2) liquidity constraints are significant (so the market misallocates water away from constrained high-valuation farmers), and (3) the production function is concave in water (so uniform allocation is efficient when farmers are homogeneous). The authors find all three conditions hold in Mula. Conversely, markets dominate quotas when heterogeneity in productivity is large relative to heterogeneity in wealth.&lt;/p&gt;
&lt;p&gt;Q: How is the transformation rate parameter gamma estimated and interpreted?
A: The transformation rate gamma measures how soil moisture above the permanent wilting point converts into apricot output (in pesetas) during the critical season, via the production function h() = gamma * (M_it - PW) * KS(M_it) * Z(w_t). It is identified from variation in purchasing patterns across seasons and variation in moisture across farmers within the same season. The preferred specification (column 3 of Table 3) yields gamma_L = 0.05. With average moisture per tree (accounting for the hydric stress coefficient) of 873.93 during the critical season, a farmer earns on average 29.09 pesetas per tree per week during the critical season, or 407.25 pesetas per tree per year.&lt;/p&gt;
&lt;p&gt;Q: How does ignoring liquidity constraints bias demand estimates?
A: If one estimates demand using the full sample (poor and wealthy farmers pooled), a decrease in demand during the critical season when prices rise conflates two effects: (1) the standard price effect (fewer farmers have valuations above the price) and (2) the liquidity constraint effect (some farmers with valuations above the price still cannot buy because they lack cash). Attributing the second effect to price sensitivity overstates the demand elasticity, biasing its absolute value upward.&lt;/p&gt;
&lt;p&gt;Q: What robustness checks do the authors provide against unobserved heterogeneity?
A: The authors provide four pieces of evidence that wealthy and poor farmers do not have systematically different underlying preferences: (1) wealthy and poor farmers are not geographically sorted into different locations (both groups appear in subareas 1, 2, 4, and 7); (2) wealthy and poor farmers grow the same bulida apricot variety; (3) outside the critical season, wealthy and poor farmers purchase statistically similar amounts of water; and (4) the purchasing divergence is significant only during the critical season when prices are high, precisely the pattern predicted by the liquidity constraint mechanism.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications for water allocation in developing countries?
A: The paper implies that before introducing water markets in regions where farmers may be liquidity constrained, policymakers should assess the magnitude of those constraints. If liquidity constraints are significant and farmers are relatively homogeneous in productivity, a quota system or a market supplemented with credit provision may deliver higher efficiency than a pure market. The standard presumption that markets outperform quotas can reverse when poor farmers cannot access credit to purchase water at the times they most need it.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to Che et al. (2013)?
A: Che, Gale, and Kim (2013) assume agents consume at most one unit with linear utility and find that markets always dominate quotas, though some non-market mechanisms with resale outperform markets. Donna and Espín-Sánchez extend this framework by allowing multiple discrete units, a concave utility function, and intertemporal dynamics. Under these extensions, the efficiency ranking between markets and quotas is theoretically indeterminate, and the authors show empirically that quotas can dominate markets. Both papers agree that non-market mechanisms with resale outperform both markets and simple quotas.&lt;/p&gt;
&lt;p&gt;Liquidity constraint (paper&amp;rsquo;s sense): A farmer is liquidity constrained when they lack sufficient cash to purchase water at the prevailing auction price, even if their valuation (marginal productivity of water) exceeds that price. In Mula, poor farmers without urban real estate income faced this constraint during the critical season when prices peaked, because they had already spent their harvest proceeds from the prior year and lacked access to credit markets.&lt;/p&gt;
&lt;p&gt;Soil moisture (M_it): The state variable measuring water accumulated in a farmer&amp;rsquo;s plot, computed using the agricultural engineering evapotranspiration formula. Moisture increases with rainfall and irrigation purchases (each auction unit contributing 432,000 liters divided by plot area) and decreases via evapotranspiration. It is bounded below by the permanent wilting point (PW) — below which trees die — and above by field capacity (FC). Moisture creates intertemporal substitution in demand.&lt;/p&gt;
&lt;p&gt;Critical season: The period corresponding to apricot fruit growth stages II and III and the Early Post-Harvest (EPH) period, spanning approximately weeks 18–32 (early May to early August). This is when the bulida apricot tree transforms water into fruit at the most rapid rate, when water demand peaks biologically, and when auction prices rise to their highest levels. It is the season during which liquidity constraints are binding.&lt;/p&gt;
&lt;p&gt;Transformation rate (gamma): The parameter in the apricot production function that measures the rate at which excess soil moisture (above the permanent wilting point) converts into apricot output (measured in real pesetas) during the critical season. Estimated at gamma_L = 0.05 in the preferred specification (column 3). It is identified from cross-seasonal variation in purchasing patterns and cross-farmer variation in moisture levels.&lt;/p&gt;
&lt;p&gt;Inefficiency due to decreasing marginal returns (DMR): One of two sources of allocation inefficiency identified in the paper. It arises when a farmer with already-high soil moisture receives water, yielding less additional output than if that water had gone to a farmer with lower moisture, given the concavity of the production function. Quotas avoid this inefficiency by allocating uniformly; markets with liquidity constraints exacerbate it by directing critical-season water to wealthy farmers who may have accumulated moisture from earlier purchases.&lt;/p&gt;
&lt;p&gt;Cuarta (quarter): The unit of water sold at Mula auctions, representing the right to use water flowing through the main channel for three hours. At approximately 40 liters per second of flow, each cuarta carried approximately 432,000 liters of water. Water rights and land rights were held independently; farmers who participated in auctions owned only land, while waterlords separately owned canal usage rights.&lt;/p&gt;
&lt;p&gt;Conditional choice probability (CCP) estimator: The two-step estimation procedure used to recover demand parameters from wealthy farmers&amp;rsquo; purchasing choices. In Step 1, transition probability matrices for observable state variables (moisture, week, price, rainfall) are computed and CCP is estimated via multinomial logit. In Step 2, the value function is forward-simulated using these transition matrices and parameters are estimated by GMM, following Hotz et al. (1994).&lt;/p&gt;</description></item><item><title>The Impact of Incarceration on Employment, Earnings, and Tax Filing</title><link>https://macropaperwarehouse.com/papers/the-impact-of-incarceration-on-employment-earnings-and-tax-filing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-impact-of-incarceration-on-employment-earnings-and-tax-filing/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;This paper estimates the causal effect of incarceration on employment, wage earnings, self-employment, and tax filing behavior using administrative criminal justice data linked to Internal Revenue Service (IRS) records for approximately half a million felony defendants in two U.S. states: North Carolina and Ohio. The study period covers cases filed from the early 2000s through 2014, with outcomes tracked through 2020 using IRS W-2 and 1040 records.&lt;/p&gt;
&lt;h3 id="research-question"&gt;Research Question&lt;/h3&gt;
&lt;p&gt;The central question is whether incarceration itself — as distinct from arrest, conviction, and other criminal justice interactions that precede or accompany it — causes lasting reductions in defendants&amp;rsquo; labor market outcomes. The paper explicitly holds fixed upstream interactions (conviction, arrest) to isolate the effect of the incarceration sentence.&lt;/p&gt;
&lt;h3 id="data-and-sample"&gt;Data and Sample&lt;/h3&gt;
&lt;p&gt;Criminal justice records from Ohio (Common Pleas courts in Franklin, Cuyahoga, and Hamilton counties, covering Columbus, Cleveland, and Cincinnati) and North Carolina (Administrative Office of the Courts and Department of Public Safety) are linked to de-identified IRS records via name, date of birth, sex, address, and partial Social Security Numbers. Match rates are 92% in Ohio and 95% in North Carolina. The sample is restricted to defendants aged 18–50 at time of offense with cases filed 2002–2014. IRS records include employer-reported W-2 wages (regardless of individual tax filing), self-employment income from Schedule C/SE, non-employee compensation (1099-MISC), and gig-economy earnings from 1099 returns. All dollar figures are adjusted to 2016 dollars using the PCE deflator.&lt;/p&gt;
&lt;h3 id="empirical-strategy"&gt;Empirical Strategy&lt;/h3&gt;
&lt;p&gt;Two independent quasi-experimental research designs are used:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;North Carolina — Sentencing guideline discontinuities&lt;/strong&gt;: North Carolina&amp;rsquo;s structured sentencing guidelines map offense class (E through I, the five least severe felony classes) and prior record points (a numerical criminal history score) into permissible punishment types (incarceration vs. probation) and sentence lengths. Allowable punishment types change discretely at five cell boundaries, generating discontinuities in incarceration sentences for otherwise similar defendants. The paper uses these five boundary discontinuities as excluded instruments in a parameterized regression discontinuity design stacked across offense classes. First-stage F-statistic = 115.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Ohio — Random assignment to judges&lt;/strong&gt;: Cases are randomly assigned by computer to judges at arraignment in the three counties studied. Judge leave-out mean sentence length is used as an instrument for individual sentence length. The design follows Norris et al. (2021) and yields F-statistic = 321. The instrument shifts sentences along both the extensive margin (any vs. no incarceration) and intensive margin (longer vs. shorter sentences).&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Both designs produce complier populations for whom at least 37–45% are shifted on the extensive margin (from no incarceration to some incarceration), based on partial identification bounds using linear programming.&lt;/p&gt;
&lt;h3 id="main-findings"&gt;Main Findings&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s central finding is that incarceration generates &lt;strong&gt;large short-run reductions&lt;/strong&gt; in labor market activity during the incapacitation period, but &lt;strong&gt;no detectable long-run reductions&lt;/strong&gt; in annual employment or earnings once defendants have been released and the incapacitation effects have dissipated.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;In the first year after case filing, when incarceration rates peak (roughly 75–100 additional days incarcerated for a 12-month sentence), employment falls by approximately &lt;strong&gt;10 percentage points&lt;/strong&gt; and total W-2 earnings contract commensurately.&lt;/li&gt;
&lt;li&gt;Within 3–4 years of filing, employment effects return to near zero and are statistically insignificant in both states.&lt;/li&gt;
&lt;li&gt;Five to nine years after filing, when effects on contemporaneous incarceration have dissipated, the estimated effect of a 12-month sentence on annual earnings is &lt;strong&gt;positive or near zero&lt;/strong&gt; in both states. The combined 95% confidence interval rules out reductions in annual wages greater than &lt;strong&gt;$231&lt;/strong&gt; (approximately 5% of the untreated complier mean) and rules out any adverse employment effects.&lt;/li&gt;
&lt;li&gt;Despite no long-run level effects, losses during incapacitation are never recouped. A one-year sentence reduces &lt;strong&gt;cumulative earnings over five years by approximately $2,914&lt;/strong&gt; — a 13% reduction relative to the complier mean.&lt;/li&gt;
&lt;li&gt;Effects on self-employment, independent contracting, 1040 filing, adjusted gross income, EITC take-up, and interstate migration are similarly null in the long run.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="incapacitation-vs-post-release-scarring"&gt;Incapacitation vs. Post-Release Scarring&lt;/h3&gt;
&lt;p&gt;The paper provides two tests for whether short-run earnings losses reflect incapacitation alone or also post-release scarring (e.g., human capital depreciation, employer discrimination, or discouragement effects):&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;A &amp;ldquo;visual IV&amp;rdquo; regression of year-t earnings effects on year-t days-incarcerated effects yields an R² of 0.83–0.85 across states, with the intercept near zero (positive and small), indicating that virtually all dynamic earnings impacts flow through contemporaneous incapacitation and not through a post-release channel.&lt;/li&gt;
&lt;li&gt;Constructed outcomes that impose the null of pure incapacitation (scaling pre-case average earnings or covariate-predicted earnings by the share of the year free from prison) closely track actual earnings effects in both states, further confirming that incapacitation is the dominant mechanism.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="pre-existing-labor-market-detachment"&gt;Pre-existing Labor Market Detachment&lt;/h3&gt;
&lt;p&gt;A key scope condition is defendants&amp;rsquo; severe labor market disadvantage prior to their case. Fewer than 50–60% of defendants are employed in the year before filing; average pre-case W-2 earnings (including zeros) are below $6,000. Among employed defendants, only 10% earn more than $22,000 per year. Untreated complier means for earnings in the year after case filing are below $4,000, with virtually no earnings or employment growth over the following nine years. The paper concludes that returning to pre-filing earnings levels is sufficient for incarcerated defendants to match their non-incarcerated peers — a low bar that is readily met.&lt;/p&gt;
&lt;h3 id="policy-implications"&gt;Policy Implications&lt;/h3&gt;
&lt;p&gt;Back-of-envelope aggregation implies incapacitation losses of approximately &lt;strong&gt;$6.16 billion per year&lt;/strong&gt; in foregone earnings for the U.S. prison population, concentrated in communities heavily affected by incarceration. However, a marginal reduction in incarceration rates would increase average earnings by only &lt;strong&gt;$51 for white men&lt;/strong&gt; and &lt;strong&gt;$213 for black men&lt;/strong&gt;, suggesting incarceration&amp;rsquo;s direct contribution to labor market inequality is modest relative to the $21,100 black-white earnings gap estimated by Bayer and Charles (2018). The paper concludes that upstream factors — other criminal justice interactions, human capital deficits, and broader socioeconomic disadvantage — are more plausibly responsible for low earnings among the formerly incarcerated.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-exact-treatment-variable-and-what-is-the-counterfactual"&gt;Q1. What is the exact treatment variable, and what is the counterfactual?&lt;/h3&gt;
&lt;p&gt;The treatment variable is months of incarceration sentenced in the focal case (a continuous, weakly positive ordered treatment). The counterfactual for non-incarcerated defendants in North Carolina is probation (all defendants are convicted by construction under structured sentencing guidelines). In Ohio, the authors cannot reject that all compliers who do not receive a prison sentence are still convicted, implying the counterfactual is also conviction and probation. All compliers therefore acquire a criminal record regardless of sentence. The treatment effect is thus the effect of incarceration conditional on conviction, holding fixed the criminal record.&lt;/p&gt;
&lt;h3 id="q2-how-are-effects-interpreted-given-multiple-instruments-and-continuous-treatment"&gt;Q2. How are effects interpreted given multiple instruments and continuous treatment?&lt;/h3&gt;
&lt;p&gt;Under a &amp;ldquo;weakly positive ordered treatment&amp;rdquo; assumption and standard LATE conditions, the 2SLS estimates can be interpreted as Average Causal Responses (ACRs) — weighted averages of the marginal dose effects (12 vs. 11 months, 6 vs. 5 months, 1 vs. 0 months, etc.) for complier subgroups shifted by each instrument. In North Carolina with five parameterized RD instruments, the estimate averages ACRs weighted by first-stage strength. In Ohio with a leave-out mean instrument, the estimate is a convex average of ACRs under the assumption that the linear first-stage model is a good approximation. Dosage weights for both states put mass on a wide range of sentence lengths including both extensive and intensive margins, though Ohio&amp;rsquo;s weights are more skewed toward shorter sentences.&lt;/p&gt;
&lt;h3 id="q3-how-large-are-the-first-stage-effects-and-how-strong-is-the-instrument"&gt;Q3. How large are the first-stage effects, and how strong is the instrument?&lt;/h3&gt;
&lt;p&gt;In North Carolina, sentences jump by 50% or more at sentencing guideline cell boundaries where allowable punishment types change to include incarceration. The first-stage F-statistic is 115. In Ohio, defendants assigned to the most severe judge receive incarceration sentences approximately six months longer than those assigned to the least severe judge (roughly 30% of the average non-zero sentence), with a slope of approximately 0.8 in the first-stage regression; F-statistic = 321. At least 37% of compliers in North Carolina and 45% in Ohio are shifted on the extensive margin (from no incarceration to some positive incarceration), with upper bounds as high as 95%.&lt;/p&gt;
&lt;h3 id="q4-what-evidence-supports-instrument-validity-exclusion-restriction-and-independence"&gt;Q4. What evidence supports instrument validity (exclusion restriction and independence)?&lt;/h3&gt;
&lt;p&gt;Instrument validity is tested by estimating 2SLS &amp;ldquo;effects&amp;rdquo; on pre-case outcomes measured 2–4 years before the focal case. In both states, the instruments show no relationship with pre-case employment, W-2 wages, total days previously incarcerated, or binary severe prior incarceration. The probability of being matched to IRS records and the quality of the match are also uncorrelated with the instruments. In Ohio, potential exclusion restriction violations from judges affecting conviction (not just sentence) are addressed empirically: nearly 90% of defendants are convicted, the most severe judge is only 0.7 p.p. more likely to convict than the least severe judge (t-stat = 1.53), and the estimated conviction rate among untreated compliers is 0.972 (s.e. 0.018), so one cannot reject that all non-incarcerated compliers are convicted.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-paper-test-for-the-incapacitation-mechanism-against-post-release-scarring"&gt;Q5. How does the paper test for the incapacitation mechanism against post-release scarring?&lt;/h3&gt;
&lt;p&gt;Two complementary exercises are conducted. First, a &amp;ldquo;visual IV&amp;rdquo; plot regresses year-t earnings effects on year-t days-incarcerated effects across all post-filing years. If incapacitation is the sole channel, all points should lie on a line through the origin. The R² is 0.83 in North Carolina and 0.85 in Ohio, the estimated intercept is near zero (positive and small) in both states, and the slope (earnings lost per day incarcerated) is approximately $12. This implies cumulative earnings losses of $12 × 268 days = $3,216, very close to the directly estimated $2,914. Second, constructed outcomes that scale pre-case earnings or covariate-predicted earnings by the share of the year not incarcerated closely track actual earnings effects throughout the post-filing period, and both converge to zero as incapacitation effects fade — consistent with pure incapacitation and no net scarring.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-long-run-59-years-earnings-and-employment-estimates-and-how-precisely-are-null-effects-ruled-out"&gt;Q6. What are the long-run (5–9 years) earnings and employment estimates, and how precisely are null effects ruled out?&lt;/h3&gt;
&lt;p&gt;Averaged across both states using inverse-variance weights, the estimated effect of a 12-month sentence on annual W-2 earnings five to nine years after filing is positive but statistically indistinguishable from zero. The 95% confidence interval rules out reductions in annual wages greater than $231 (approximately 5% of the untreated complier mean of roughly $4,500–$5,000). The 95% CI also rules out any adverse employment effects. The untreated complier mean for employment 5–9 years post-filing is approximately 40% in North Carolina and slightly above 40% in Ohio.&lt;/p&gt;
&lt;h3 id="q7-what-happens-to-cumulative-earnings-over-five-years-despite-null-long-run-level-effects"&gt;Q7. What happens to cumulative earnings over five years despite null long-run level effects?&lt;/h3&gt;
&lt;p&gt;Even though long-run annual earnings are unaffected, earnings losses during incapacitation are never made up. A one-year sentence reduces cumulative employment (measured as years with any W-2) and cumulative earnings over five years by approximately $2,914 — a 13% reduction relative to the complier mean. This reflects the mechanical loss of earnings during the period of physical incapacitation, without a subsequent compensating period of higher earnings after release.&lt;/p&gt;
&lt;h3 id="q8-do-defendants-with-stronger-pre-case-labor-market-attachment-show-different-long-run-patterns"&gt;Q8. Do defendants with stronger pre-case labor market attachment show different long-run patterns?&lt;/h3&gt;
&lt;p&gt;The sample is split between defendants employed in at least 2 of the 4 years prior to the case (53–57% of the sample across states) and those less attached. Both groups show zero long-run earnings and employment effects. Previously employed defendants experience much larger short-run earnings drops — more than three times larger in the first year post-filing — and their earnings recover more slowly, reaching zero effect approximately six years after filing (vs. three years for the previously unemployed). For a stricter cut (pre-case average earnings above $15,000, representing only 12–15% of the sample), the long-run earnings effect is −$1,426 (8% of the untreated complier mean), significant only at the 10% level, and partly attributable to residual incapacitation (19.6 additional days incarcerated 5–9 years post-filing). For defendants with pre-case earnings below $15,000, incarceration slightly increases long-run employment (2.4 pp, p = 0.01) and earnings ($400, p = 0.03), possibly reflecting rehabilitative benefits (GED or educational programs) for labor-market-detached individuals.&lt;/p&gt;
&lt;h3 id="q9-does-first-time-incarceration-extensive-margin-exposure-have-larger-long-run-effects-than-repeat-exposure"&gt;Q9. Does first-time incarceration (extensive-margin exposure) have larger long-run effects than repeat exposure?&lt;/h3&gt;
&lt;p&gt;The paper tests this by splitting the sample into defendants with and without prior incarceration history. Among defendants with no prior incarceration, the instruments generate large differences in lifetime exposure: a 12-month sentence increases the probability of ever being incarcerated over the next 5–9 years by 26 p.p. (North Carolina) and 41 p.p. (Ohio). Among those not receiving a sentence, 48% (North Carolina) and 19% (Ohio) are eventually incarcerated anyway, implying treatment causes a 52 and 81 p.p. increase in lifetime incarceration probability for extensive-margin compliers. Despite these large differences in lifetime exposure, long-run earnings and employment effects remain small and statistically insignificant in both subsamples. The difference in long-run effects between previously and never incarcerated defendants is not statistically significant (p = 0.29 for employment, p = 0.82 for earnings).&lt;/p&gt;
&lt;h3 id="q10-are-there-heterogeneous-effects-by-race-sex-or-criminal-history"&gt;Q10. Are there heterogeneous effects by race, sex, or criminal history?&lt;/h3&gt;
&lt;p&gt;There is no evidence of long-run scarring for any demographic or criminal history subgroup. Effects for black and non-black defendants are both positive for long-run earnings and employment. Non-black defendants show somewhat larger cumulative losses (consistent with marginally higher counterfactual earnings), but differences are not statistically significant. Estimates for women are imprecise due to small sample size. Among defendants with and without prior felony charges in the four years preceding the case, there are neither economically nor statistically significant long-run earnings or employment effects. Cumulative losses are somewhat larger for defendants without prior felony charges (p = 0.07), reflecting their higher pre-case earnings.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-handle-potential-migration-bias-in-outcomes"&gt;Q11. How does the paper handle potential migration bias in outcomes?&lt;/h3&gt;
&lt;p&gt;Tax filing and W-2 receipt in the state of sentencing are used to proxy for whether defendants remain in the same state. Among untreated compliers, 88% of those with a tax footprint maintain it in the state of sentencing. No statistically significant effects of incarceration on migration (measured as filing or receiving a W-2 in North Carolina or Ohio) are detected, suggesting prior studies of recidivism measured within-state are unlikely to be severely biased by migration responses.&lt;/p&gt;
&lt;h3 id="q12-what-effect-does-incarceration-have-on-mortality"&gt;Q12. What effect does incarceration have on mortality?&lt;/h3&gt;
&lt;p&gt;Incarceration reduces five-year mortality by approximately 0.8 percentage points (about 20% of the untreated mean). The authors note this is too small to explain the null long-run labor market effects: even if all defendants whose death was averted were employed, removing them from the employment count would reduce the employment effect of a 12-month sentence only to approximately zero.&lt;/p&gt;
&lt;h3 id="q13-how-do-the-papers-findings-compare-to-prior-studies-particularly-mueller-smith-2015"&gt;Q13. How do the paper&amp;rsquo;s findings compare to prior studies, particularly Mueller-Smith (2015)?&lt;/h3&gt;
&lt;p&gt;Mueller-Smith (2015) finds large and persistent negative incarceration effects on labor market outcomes in Texas using a structural decomposition and Lasso-based judge-covariate interactions as instruments. The paper argues methodological differences are the likely explanation: the Lasso-selected interacted instruments can be susceptible to many-weak instruments bias toward OLS. It notes that Mueller-Smith&amp;rsquo;s simpler 2SLS specifications (analogous to those used here) show no statistically significant earnings effects. North Carolina and Ohio are documented to be broadly similar to Texas (and the U.S. average) in rehabilitation program participation, recidivism rates, and incarceration rates, reducing the likelihood that genuine geographic heterogeneity explains the divergence.&lt;/p&gt;
&lt;h3 id="q14-what-is-the-papers-aggregate-extrapolation-of-incapacitation-earnings-losses"&gt;Q14. What is the paper&amp;rsquo;s aggregate extrapolation of incapacitation earnings losses?&lt;/h3&gt;
&lt;p&gt;Scaling the estimated $2,914 cumulative loss per 12-month sentence by the ratio of days exposed to total days in a year gives a per-day loss of approximately $12. Applied to the 1,435,500 people incarcerated in U.S. prisons on any given day in 2019 (excluding the more than 700,000 in jail), the implied aggregate yearly earnings loss from incapacitation is approximately $6.16 billion.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Incapacitation effect&lt;/strong&gt;: The mechanical reduction in earnings and employment that occurs while a defendant is physically confined in prison and unable to work, as distinct from any post-release scarring effect. The paper shows this is the dominant — and essentially sole — causal channel through which incarceration affects labor market outcomes in their sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Post-release scarring&lt;/strong&gt;: Persistent reductions in earnings or employment that persist after a defendant is released from prison, caused by mechanisms such as employer discrimination based on incarceration history, human capital depreciation, loss of job contacts, or psychological discouragement effects. The paper finds no evidence of scarring in either state.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Average Causal Response (ACR)&lt;/strong&gt;: The weighted average of the marginal dose effects of incarceration (e.g., effect of 12 vs. 11 months, 1 vs. 0 months) for groups of defendants whose sentence lengths are shifted by a given instrument. Contrasted with a binary LATE, the ACR averages across the full dosage distribution for compliers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Complier&lt;/strong&gt;: An individual whose incarceration sentence is shifted by the instrument — either from zero to some positive sentence (extensive margin) or from a shorter to a longer sentence (intensive margin). Counterfactual outcome means for compliers sentenced to zero months provide the baseline for evaluating effect magnitudes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sentencing guideline discontinuity&lt;/strong&gt;: The discrete jump in permissible punishment types and minimum sentence lengths at specific criminal history score thresholds within North Carolina&amp;rsquo;s structured sentencing grid. Defendants just above a threshold are more likely to be incarcerated than otherwise similar defendants just below, generating quasi-experimental variation exploited as an instrument.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Leave-out mean judge instrument&lt;/strong&gt;: In Ohio, each defendant&amp;rsquo;s assigned judge&amp;rsquo;s average incarceration sentence length computed over all other cases that judge handles (excluding the defendant&amp;rsquo;s own case), residualized on court-by-month fixed effects. Because judges are randomly assigned to cases, this measure is conditionally independent of defendant potential outcomes and serves as an instrument for sentence length.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Control complier mean&lt;/strong&gt;: The estimated mean potential outcome for compliers under the counterfactual of receiving zero months of incarceration. Used as a benchmark to evaluate the magnitude of treatment effects and to characterize how low the earnings baseline is for the population driving the causal estimates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extensive vs. intensive margin of incarceration&lt;/strong&gt;: The extensive margin refers to the binary shift from receiving no prison sentence to receiving any prison sentence; the intensive margin refers to increasing sentence length conditional on some incarceration. The paper argues that neither margin appears to produce long-run labor market scarring, and uses linear programming bounds to estimate that at least 37–45% of compliers in each state are shifted on the extensive margin.&lt;/p&gt;</description></item><item><title>The Long-Run Impacts of Public Industrial Investment on Local Development and Economic Mobility: Evidence from World War II</title><link>https://macropaperwarehouse.com/papers/the-long-run-impacts-of-public-industrial-investment-on-local-development-and-economic-mobility-evidence-from-world-war-ii/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-long-run-impacts-of-public-industrial-investment-on-local-development-and-economic-mobility-evidence-from-world-war-ii/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Does government-led construction of large manufacturing plants in previously under-industrialized regions generate long-run improvements in regional economic development and in the lifetime earnings of the incumbent residents who were already living there at the outset? And, if so, through what mechanism — developmental improvements during childhood or expanded adult labor market opportunities?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and Identification.&lt;/strong&gt; The paper exploits the United States industrial mobilization for World War II, specifically the construction of 90 large, government-financed, newly-built manufacturing plants (each costing $10 million or more in contemporary dollars, approximately $150 million in 2020 dollars) in dispersed locations outside the major prewar manufacturing hubs. Strategic and security considerations — not economic optimization — drove the military to insist these plants be sited away from congested industrial centers. Because private firms were unwilling to finance construction in isolated locations with uncertain postwar value, the government built them directly as government-owned, contractor-operated (GOCO) facilities through the Defense Plant Corporation. Site selection within the set of sufficiently populated regions was governed by idiosyncratic, short-run factors — the immediate availability of suitable parcels, informal connections to procurement officers, and expedience — rather than systematic economic characteristics of the receiving counties. The paper documents no systematic association between publicly-funded wartime plant construction and prewar county-level economic or demographic characteristics conditional on population size, and finds parallel prewar trends and balanced outcome levels across treatment and comparison counties in all decades leading up to WWII. A placebo test using 1910-to-1940 intergenerational mobility in matched Census records confirms no differential prewar upward mobility in treatment counties.&lt;/p&gt;
&lt;p&gt;The comparison group consists of 1,400 counties outside the 100 largest prewar manufacturing counties that did not receive large public plants. Treatment assignment for individuals is based on birth county, not adult county of residence, enabling the paper to track outcomes regardless of where individuals ultimately live.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The analysis draws on the 1945 War Production Board data book for plant-level investment; county-level panels from Decennial and Economic Censuses spanning 1900–2000; the SSA NUMIDENT file (birth county and date); IRS Form 1040 individual income tax returns in 1969, 1974, 1979, and 1984 (covering wage earnings and adjusted gross income); the full-count 1940 Census (parent earnings, demographics); the 2000 Census long form (educational attainment); and W-2 earnings histories from the SSA Detailed Earnings Record matched to a CPS-linked subsample, with employer information linked to the Business Register.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regional Effects.&lt;/strong&gt; By 1970, counties receiving large public wartime plants had approximately 30 percent higher manufacturing employment, 20 percent larger populations, and 7–8 percent higher median family income than comparison counties. Manufacturing employment as a share of total employment rose and remained elevated through the 1970s before converging toward parity with the comparison group by 1990. Treated counties were permanently larger — with population stabilizing at a new, persistently higher equilibrium roughly 20 percent above comparison counties by end of century — even after the manufacturing employment share converged, consistent with path dependence and multiple equilibria. Average production worker pay in manufacturing rose by approximately 10 percent, closely tracking value-added per worker, while average retail wages rose by only one-third as much and were not statistically significant in most years. In the 40 years after the war, treated counties saw median family earnings increase by 5–10 percent, concentrated in higher average wages and employment shares in manufacturing and semi-skilled blue-collar occupations, with limited effects on non-manufacturing, white-collar occupations, or female individual income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Individual Earnings Effects.&lt;/strong&gt; Men born in treatment counties in the 18 years before the war (birth cohorts 1922–1940) earned approximately $1,200–$1,300 more per year (2020 dollars) in average wage earnings reported on 1040 returns in 1969, 1974, 1979, and 1984 — an increase of 2.5–3 percent and roughly a one-percentile rise in the national earnings distribution. Effects were largest for children of parents at the bottom of the 1939 earnings distribution: children of the lowest-income parents saw adult wage earnings rise by approximately $1,800–$2,000 per year (3–4 percent), with effects declining linearly by parent rank and effectively vanishing for children of the highest-earning parents. Black men experienced larger average earnings effects (4–6 percent, or $1,500–$2,500 in 2020 dollars) than White men (2–3 percent, or $1,000–$1,500), with the racial earnings gap estimated to have narrowed by about 2 percent in the treatment group. When examining Form 1040 returns (tax-unit level), effects are comparable for men and women, but W-2 individual earnings data from the SSA-CPS subsample show no positive effect on women&amp;rsquo;s own earnings — the 1040 effects for women are entirely driven by their husbands&amp;rsquo; higher earnings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanism.&lt;/strong&gt; The balance of evidence points to access to higher-wage jobs in adulthood as the primary channel, rather than developmental human capital improvements accumulated during childhood. War plants modestly increased male educational attainment — children from the lowest-earning families completed approximately one-quarter of a year more schooling and were 3 percentage points more likely to graduate high school — but education effects are too small to account for the full earnings increase. Critically, there is no gradient in earnings effects by birth cohort: children who were younger at the start of the war and therefore had longer childhood exposure to improved regions did not benefit more, contradicting a childhood exposure-effect mechanism as in Chetty and Hendren (2018b). Adult earnings effects are entirely accounted for by adult location: conditioning on 1979 county of residence eliminates the treatment effect. Stayers in treatment counties show large earnings differences relative to stayers in comparison counties, while movers show none. Men born in treatment counties are also directly documented to have worked in industries with higher wage premiums as adults, with coarse industry classification alone accounting for approximately one-third of the estimated log wage increase.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy Scope Conditions.&lt;/strong&gt; The paper argues these effects are specific to the WWII postwar institutional context — high global demand for U.S. manufactured goods, limited international competition, labor-intensive production techniques, and strong union bargaining power — conditions that no longer hold. Reexamination of &amp;ldquo;million-dollar plant&amp;rdquo; openings in the 1980s and 1990s shows manufacturing employment expanded but average manufacturing wages did not increase, suggesting contemporary plant openings do not generate the same high-wage opportunities. The association between manufacturing employment density and upward mobility visible in 1950 has entirely vanished by the end of the twentieth century.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-exactly-defines-the-treatment-group-and-why-were-these-plants-built-by-the-government-rather-than-private-firms"&gt;Q1. What exactly defines the treatment group, and why were these plants built by the government rather than private firms?&lt;/h3&gt;
&lt;p&gt;A: The treatment group consists of 90 counties outside the 100 largest prewar manufacturing regions that received at least one new, fully publicly-financed manufacturing plant costing $10 million or more (approximately $150 million in 2020 dollars) under the WWII industrial mobilization. Private firms refused to finance construction in dispersed, isolated locations with highly uncertain postwar value; the Air Force historians recorded that &amp;ldquo;industrialists&amp;rsquo; reluctance to invest in dispersed plant facilities was at odds with the government&amp;rsquo;s hope that private capital could finance new inland construction.&amp;rdquo; The government built and owned these facilities as GOCO plants, operated by private firms under contract. The 353 plants meeting the cost threshold (including both large and smaller public plants) account for 70 percent of all spending on new plants during the war.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-authors-establish-that-plant-siting-was-quasi-random-conditional-on-population-size"&gt;Q2. How do the authors establish that plant siting was quasi-random conditional on population size?&lt;/h3&gt;
&lt;p&gt;A: Identification rests on three forms of evidence. First, historical documents show procurement decisions were driven by idiosyncratic factors — availability of a suitable parcel, informal connections to procurement officers, short-run expedience — rather than systematic economic characteristics. Members of Congress had little ability to influence siting, and Rhode et al. (2018) find little evidence that federal politics drove the geographic distribution of wartime spending. Second, balance tests (estimating prewar county characteristics as outcomes in Equation 1) show no significant differences between treatment and comparison counties in earnings levels, demographics, manufacturing development, or industrial composition after conditioning on 1940 population, with a joint p-value of 0.30 (0.36 when also conditioning on geography and infrastructure). Third, a placebo test using children in the 1910 Census matched to the 1940 Census finds no differential economic outcomes or upward mobility rates in counties that would eventually receive treatment plants, conditional on basic region size.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-county-level-effects-on-the-structure-of-the-labor-market-in-the-medium-run"&gt;Q3. What are the county-level effects on the structure of the labor market in the medium run?&lt;/h3&gt;
&lt;p&gt;A: By the 1960s–1970s, treated counties had higher predicted union coverage rates and a greater share of men in semi-skilled production occupations, driven primarily by movement away from farm work and supplemented by higher male labor force participation. Average wages in craftsperson and operator occupations rose by 8 percent in treated counties — more than double the increase in wages for high-skill professional and managerial occupations. Treated counties had 8 percent higher median male individual incomes by 1979. Effects on female median individual income were minimal, and there were no effects on female labor force participation rates.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-estimated-magnitude-of-the-individual-earnings-effects-and-how-do-they-vary-by-parent-income"&gt;Q4. What is the estimated magnitude of the individual earnings effects, and how do they vary by parent income?&lt;/h3&gt;
&lt;p&gt;A: Men born in treatment counties averaged $1,200–$1,300 more per year in real wage earnings (2020 dollars) on 1040 tax returns across the four observation years 1969, 1974, 1979, and 1984, a 2.5–3 percent increase equivalent to roughly one percentile in the national earnings distribution. Heterogeneity by parent rank is pronounced and monotone: children of parents at the very bottom of the 1939 earnings distribution gained approximately $2,000 per year (about 4 percent), while children of the highest-earning parents experienced no significant effect. When county weighting is equalized to eliminate the differential representation of rural (lower-income) counties, effects are roughly constant across the bottom six deciles of the parent earnings distribution and then drop steeply at the top, showing that the earnings gradient was not simply an artifact of plant openings in poorer, smaller counties.&lt;/p&gt;
&lt;h3 id="q5-how-did-effects-differ-by-race"&gt;Q5. How did effects differ by race?&lt;/h3&gt;
&lt;p&gt;A: Wartime plant construction increased annual adult earnings of Black men by 4–6 percent ($1,500–$2,500 in 2020 dollars) and of White men by 2–3 percent ($1,000–$1,500 in 2020 dollars). The racial earnings gap in the treatment group is estimated to have narrowed by about 2 percent. However, the pattern of heterogeneity by parent income differs by race: for White men, effects are largest for children of below-median parents and effectively zero for children of above-median parents. For Black men, the largest effects — 7–10 percent ($4,000–$5,000 in 2020 dollars) — accrue to children of parents with earnings above the pooled-race national median, while effects for lower-income Black families range from 3–6.5 percent, suggesting that Black workers from higher-income backgrounds particularly benefited from wartime anti-discrimination policies and the opening of previously restricted manufacturing occupations.&lt;/p&gt;
&lt;h3 id="q6-why-do-the-1040-returns-show-comparable-effects-for-men-and-women-while-w-2-data-show-no-effect-on-womens-individual-earnings"&gt;Q6. Why do the 1040 returns show comparable effects for men and women, while W-2 data show no effect on women&amp;rsquo;s individual earnings?&lt;/h3&gt;
&lt;p&gt;A: Form 1040 returns are filed at the tax-unit level — for married couples, they report the combined wages of both spouses. Because more than 80 percent of women in the sample are married, an increase in a husband&amp;rsquo;s earnings raises the joint 1040 figure for both spouses. The SSA-CPS subsample with individual W-2 records shows that the entire effect on men&amp;rsquo;s Form 1040 wages directly reflects increases in their own W-2 earnings, while women&amp;rsquo;s own W-2 earnings show no positive treatment effect. This finding is consistent with county-level evidence of no impact on female individual income or female labor force participation, and with Rose (2018) finding that women were almost universally excluded from manufacturing jobs after the war&amp;rsquo;s conclusion despite high wartime female manufacturing employment.&lt;/p&gt;
&lt;h3 id="q7-what-evidence-tests-the-developmental-effects-mechanism"&gt;Q7. What evidence tests the developmental-effects mechanism?&lt;/h3&gt;
&lt;p&gt;A: Three tests argue against childhood developmental effects as the primary driver. First, educational attainment effects — while statistically significant for children of the lowest-income parents (approximately one-quarter of a year more schooling, 3 percentage points more likely to graduate high school) — are too small to account for the earnings increase: a Mincer-equation calculation shows that the education effects can explain less than one-half of the estimated effect on 1979 wages. Second, there is no gradient in earnings effects by birth cohort — children younger at the war&amp;rsquo;s start, who had longer post-treatment childhood exposure, did not benefit more, in direct contrast to the Chetty-Hendren childhood-exposure framework. Third, postwar in-migrants into treatment counties were not drawn from better-educated or higher-income families and did not themselves have more education than in-migrants into comparison regions, ruling out peer effects from selective in-migration.&lt;/p&gt;
&lt;h3 id="q8-what-evidence-directly-implicates-adult-labor-market-access-as-the-operative-mechanism"&gt;Q8. What evidence directly implicates adult labor market access as the operative mechanism?&lt;/h3&gt;
&lt;p&gt;A: Four pieces of evidence point to contemporaneous adult labor market access. First, individuals born in treatment counties lived as adults in counties with 3–4 percent higher median male earnings and higher wages in semi-skilled blue-collar occupations but not in highly-skilled professional occupations — a pattern quantitatively consistent with the individual earnings effects. Second, the entire earnings effect is concentrated among those who remain in their birth counties: stayers in treatment counties show earnings differences of similar magnitude to county-level manufacturing wage effects, while movers show no difference compared to movers from comparison counties. Third, conditioning on 1979 county of residence eliminates the earnings effect entirely (1979 location fixed effects specification). Fourth, using W-2 data matched to the Business Register in the SSA-CPS sample, men born in treatment counties are directly shown to work in industries with higher wage premiums, with coarse industry classification alone accounting for approximately one-third of the log wage increase.&lt;/p&gt;
&lt;h3 id="q9-is-the-persistence-of-regional-effects-driven-by-continued-cold-war-military-spending-at-the-plants"&gt;Q9. Is the persistence of regional effects driven by continued Cold War military spending at the plants?&lt;/h3&gt;
&lt;p&gt;A: No. The paper separates ordnance and ammunition plants — which predominantly became GOCO facilities or Air Force Bases after WWII and received disproportionately more Vietnam War-era defense spending — from general manufacturing plants, which overwhelmingly transitioned to privatized civilian production. Both types of plants show similarly persistent effects on manufacturing employment and comparable impacts on the long-run earnings of local children. Moreover, general manufacturing plants — which did not generate increased postwar military spending — had large permanent effects on overall population growth, while ordnance plants had smaller population effects. The persistence therefore does not appear to reflect continued federal expenditure.&lt;/p&gt;
&lt;h3 id="q10-what-mechanism-explains-the-permanent-population-effect-even-after-manufacturing-employment-shares-converge"&gt;Q10. What mechanism explains the permanent population effect even after manufacturing employment shares converge?&lt;/h3&gt;
&lt;p&gt;A: The authors interpret the permanent population differential — treated counties remain roughly 20 percent larger than comparison counties even at the end of the 20th century, after manufacturing employment shares converge — as evidence of path dependence and multiple equilibria. Once a region reaches a new, larger equilibrium, self-sustaining forces (expanded non-tradable employment, public infrastructure investment) maintain it. Treatment counties are more likely to have been connected to the interstate highway system in subsequent decades and show positive effects on local government capital outlays for utilities. The medium-term persistence is attributed partly to the sunk costs of site establishment (surveying, local approvals, infrastructure connections), which make reinvestment at existing sites more attractive than greenfield construction elsewhere.&lt;/p&gt;
&lt;h3 id="q11-do-smaller-plant-openings-generate-comparable-effects"&gt;Q11. Do smaller plant openings generate comparable effects?&lt;/h3&gt;
&lt;p&gt;A: No. Counties receiving smaller publicly-financed plants costing between $1 and $10 million show no detectable effects on manufacturing employment, population, median family income, or individual adult earnings comparable to those from the large plants. The authors cannot rule out the presence of small effects, but the null results for smaller plants — combined with evidence that the largest effects are in counties with the highest investment intensity per 1940 resident — are consistent with threshold effects (&amp;ldquo;big push&amp;rdquo;) in regional development, though the wide confidence intervals do not allow the authors to conclusively distinguish threshold effects from a linear-in-investment model.&lt;/p&gt;
&lt;h3 id="q12-what-do-modern-million-dollar-plant-openings-reveal-about-the-contemporary-relevance-of-these-findings"&gt;Q12. What do modern &amp;ldquo;million-dollar plant&amp;rdquo; openings reveal about the contemporary relevance of these findings?&lt;/h3&gt;
&lt;p&gt;A: Reexamining plant openings from Greenstone et al. (2010) using an event-study design, the authors find that 1980s–1990s million-dollar plant openings expanded manufacturing employment (consistent with Greenstone et al.) but had no impact on average manufacturing wages — in sharp contrast to the WWII findings. Slattery and Zidar (2020) similarly find no impacts on county-level incomes for plant openings since 2000. The correlation between manufacturing employment density and upward mobility rates visible in 1950 had entirely vanished by the end of the 20th century. The authors attribute the divergent results to the changed institutional environment: contemporary production is highly automated, relies on interchangeable labor from staffing agencies, faces intense international competition, and is conducted under much weaker collective bargaining institutions.&lt;/p&gt;
&lt;h3 id="q13-what-is-the-papers-assessment-of-aggregate-welfare-implications"&gt;Q13. What is the paper&amp;rsquo;s assessment of aggregate welfare implications?&lt;/h3&gt;
&lt;p&gt;A: The paper is explicit that its local estimates do not allow clean conclusions about aggregate effects. Publicly-financed plant construction in peripheral locations may have crowded out private investment that would otherwise have occurred in major manufacturing hubs. If so, the documented regional gains represent geographic reallocation of manufacturing activity rather than a net increase in the aggregate plant stock. Aggregate gains from reallocation would require that the benefits in the selected dispersed locations exceeded what would have occurred in the counterfactual locations — a plausible conjecture given the paper&amp;rsquo;s evidence that effects are larger in counties with lower prewar manufacturing employment shares and lower initial market access, but one the authors cannot demonstrate decisively.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Government-Owned, Contractor-Operated (GOCO) Plants:&lt;/strong&gt; Manufacturing facilities built and owned by a U.S. government agency (typically the Defense Plant Corporation) during WWII but built and operated by private firms under cost-plus contracts. GOCO status meant the government bore full construction risk and that post-war disposition (sale to private buyers at a fraction of construction cost, or continued GOCO operation for ordnance production) was determined by public agencies, not by the constructing firm&amp;rsquo;s investment calculus.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Place-Based Predistribution:&lt;/strong&gt; The paper&amp;rsquo;s term for the mechanism by which wartime plant construction raised the incomes of existing residents — not through ex-post redistribution of income via taxes and transfers, but by expanding the set of high-wage employment opportunities available to incumbent workers in the region, thereby changing the pre-tax, pre-transfer wage structure facing those workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Adult Labor Market Access (vs. Childhood Developmental Exposure):&lt;/strong&gt; A distinction the paper draws in explaining why children born in treated counties had higher adult earnings. The &amp;ldquo;developmental exposure&amp;rdquo; mechanism (as in Chetty and Hendren 2018b) implies benefits scale with the amount of time spent in an improved childhood environment. The &amp;ldquo;adult labor market access&amp;rdquo; mechanism means children benefit irrespective of years of childhood exposure because they can access improved local labor market conditions when they reach working age as adults — what the paper operationalizes through the finding that earnings effects are entirely accounted for by 1979 county of residence and are concentrated among individuals who remain in their birth counties.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Upward Mobility (Absolute and Relative):&lt;/strong&gt; Following Chetty et al. (2014), the paper uses both concepts: absolute upward mobility means children from low-income backgrounds have higher lifetime earnings than comparable children in counterfactual regions; relative upward mobility means their outcomes converge toward those of children from affluent backgrounds. The paper documents both: large earnings effects for the lowest parent-income deciles, declining linearly to zero for the top deciles.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conditional Independence (Plant Siting as Quasi-Random):&lt;/strong&gt; The paper&amp;rsquo;s identification assumption — that among counties with observably similar population sizes and basic geographic/infrastructure characteristics, the specific choice of plant siting locations was driven by idiosyncratic, short-run factors uncorrelated with potential postwar outcomes. This is a level-balance assumption (not merely a parallel-trends assumption), required because individual outcomes are only observed in the post-period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Industry Wage Premium:&lt;/strong&gt; The paper uses Krueger and Summers (1988) estimates of inter-industry wage differentials (the portion of a sector&amp;rsquo;s average wage unexplained by worker characteristics) to classify adult employers of treated individuals. Finding that men born in treatment counties work at employers in higher-premium industries — with industry category alone explaining approximately one-third of the log wage increase — provides direct evidence of the adult labor market access mechanism operating through industry sorting.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Path Dependence / Multiple Equilibria in Regional Development:&lt;/strong&gt; The paper documents that treated counties remain permanently larger in population than comparison counties even after manufacturing employment shares converge and the original plants begin to close. This self-sustaining population differential, inconsistent with a unique spatial equilibrium, is interpreted as evidence that the temporary wartime shock shifted treated regions into a permanently higher equilibrium, sustained by subsequent infrastructure investment and non-tradable sector expansion proportional to the larger population base.&lt;/p&gt;</description></item><item><title>The Optimal Taxation of Couples</title><link>https://macropaperwarehouse.com/papers/the-optimal-taxation-of-couples/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-optimal-taxation-of-couples/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; What is the optimal joint nonlinear earnings tax schedule for married couples? How should one spouse&amp;rsquo;s marginal tax rate depend on the other&amp;rsquo;s earnings? When is individual earnings-based (separable) taxation optimal versus family-income-based taxation, and what determines the sign and magnitude of &amp;ldquo;jointness&amp;rdquo; — the dependence of one spouse&amp;rsquo;s marginal tax on the other&amp;rsquo;s earnings?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The paper studies a canonical unitary household model in which each couple consists of two spouses who jointly maximize utility subject to a joint budget constraint. Spousal productivities are drawn from a joint distribution F with arbitrary dependence structure. The planner maximizes a weighted sum of couples&amp;rsquo; utilities, with Pareto weights that are decreasing functions of productivities. Utility takes a quasi-linear form in consumption and labor disutility with constant labor supply elasticity parameter γ (implying earnings elasticity γ/(γ-1)). The tax problem is equivalent to a two-dimensional mechanism design problem in which the planner chooses allocations as functions of reported productivity types, subject to incentive compatibility and budget feasibility. Because spousal productivities are two-dimensional, the problem is a multi-dimensional screening problem whose properties are poorly understood in general.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Methodology.&lt;/strong&gt; The authors proceed in two directions. First, they establish conditions under which the first-order approach (FOA) — restricting attention to local incentive constraints — is valid in this bi-dimensional setting. They show, for the special case of the benchmark economy (symmetric, independent types, separable Pareto weights), that FOA validity is equivalent to convexity of a certain transformation of the value function, and derive necessary and sufficient conditions that are strictly weaker than their unidimensional analogs — so the FOA is more likely to hold in two dimensions than in one. For the general economy, they invoke an Implicit Function Theorem argument in Hölder space to show that the FOA holds for Pareto weights sufficiently close to utilitarian (i.e., when the planner is not &amp;ldquo;too redistributive&amp;rdquo;). Second, assuming FOA validity, they characterize optimal taxes via a second-order nonlinear PDE. Since this PDE cannot be solved analytically in general, they apply the Coarea Formula to derive closed-form expressions for conditional averages of optimal tax distortions over various subsets of the type space, expressed entirely in terms of structural primitives (labor supply elasticities, Pareto weights, and elasticities of the joint distribution of productivities).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings.&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Average distortions and assortativeness.&lt;/strong&gt; Average optimal distortions on married individuals are ranked by the degree of positive quadrant dependence (PQD) in spousal productivities: more assortative matching implies higher optimal tax rates. Optimal distortions on married individuals are always weakly lower than on single individuals with the same productivity, same elasticities, and same marginal productivity distribution — strictly so unless matching is perfectly positively assortative. The intuition is that when couples pool resources, intra-family redistribution already occurs, and distortionary taxation crowds this out; more random matching produces more within-family redistribution, reducing the marginal social value of public redistribution through taxation.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Optimality of separable (individual earnings-based) taxation.&lt;/strong&gt; In the benchmark economy with independent types, optimal taxes are exactly separable (individual earnings-based), and optimal distortions on married individuals equal precisely one-half of those on comparable single individuals. With separable Pareto weights and independent types more generally, taxes remain separable. Once types are positively dependent, however, the planner optimally introduces jointness even under separable social weights.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Jointness and tail (in)dependence.&lt;/strong&gt; Optimal jointness — whether one spouse&amp;rsquo;s marginal tax rate increases or decreases in the other&amp;rsquo;s earnings — depends critically on tail dependence of the joint productivity distribution, captured by the copula and survival copula elasticities. For right-tail dependent distributions (so that extremely productive individuals are likely to be matched with extremely productive partners), positive jointness is optimal at the top (raising taxes on high earners whose partners are also high earners) and negative at the bottom. For right-tail independent distributions (such as the Gaussian copula, which is tail-independent for any finite ρ), the distortion-reducing motive dominates: optimal jointness is negative at the top and positive at the bottom, conditional on standard convergence conditions.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Primary vs. secondary earners.&lt;/strong&gt; The secondary earner (lower-productivity spouse) faces on average higher optimal distortions than the primary earner when the planner values redistribution to couples with a very unproductive spouse (α(w,0) ≥ 1), because the phasing out of transfers targeted to such couples generates high marginal tax rates on secondary earners. Family earnings-based taxation is optimal only when total family productivity and relative spousal productivity are independent, and when social weights are measurable only with respect to total family output.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Restricted taxation.&lt;/strong&gt; Optimal distortions under any of the three restricted tax regimes (anonymous, separable, family earnings-based) exactly equal the relevant conditional average of unrestricted optimal distortions. This establishes that the welfare difference between the restricted and unrestricted optimum stems solely from the planner&amp;rsquo;s inability to tag taxes to individual productivity types within the restricted class.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Findings (calibrated to 2020 CPS data on U.S. married couples, ages 25-65, worked ≥ 20 weeks).&lt;/strong&gt; Spousal productivities are positively but not perfectly dependent, with Kendall&amp;rsquo;s tau = 0.21 and Pearson correlation = 0.25 for productivities (0.21 for earnings). The joint distribution is well approximated by a Gaussian copula (ρ = 0.33) with Pareto-lognormal marginals (a = 2.95, Gini = 0.31). The Gaussian copula is tail-independent, so consistent with analytical results, optimal jointness is positive for low earners and negative for high earners (the latter arising at earnings above approximately $8.5 million in the benchmark specification). The quantitative magnitude of optimal jointness is small — marginal taxes for one spouse change by at most several percentage points as a function of the other spouse&amp;rsquo;s earnings. Individual earnings-based taxation provides a good approximation to the unrestricted optimum. By contrast, family earnings-based (joint) taxation is a poor approximation in all specifications, with marginal taxes on family income varying substantially with the earnings share of the secondary earner, and this conclusion holds even when Pareto weights explicitly favor family earnings-based taxation (k = 0 case). The implied top marginal tax rate converges toward approximately 55 percent (corresponding to limiting distortion of ≈1.35 = 1/γa with γ = 0.25, a = 2.95) but the convergence is slow, so optimal marginal rates remain substantially below this limit even at earnings of $300,000.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-mechanism-design-formulation-and-why-is-foa-validity-a-key-concern-in-the-bi-dimensional-setting"&gt;Q1. What is the mechanism design formulation, and why is FOA validity a key concern in the bi-dimensional setting?&lt;/h3&gt;
&lt;p&gt;A: The planner&amp;rsquo;s problem is cast as a direct mechanism in which couples report their two-dimensional productivity type (w1, w2) and receive allocations (consumption, earnings). Incentive compatibility requires that no couple prefers to misreport. In one-dimensional models (Mirrlees 1971), restricting attention to local incentive constraints (the FOA) yields the standard ODE characterization of optimal taxes and is valid for a broad class of primitives. In two dimensions, solutions to multi-dimensional screening problems generically display &amp;ldquo;bunching&amp;rdquo; (Rochet-Choné 1998, Armstrong 1996), and the FOA may fail. The key difference exploited in this paper is the absence of participation constraints in the public finance setting, which eliminates the main force driving FOA failure in industrial organization models.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-necessary-and-sufficient-conditions-for-foa-validity-in-the-benchmark-economy-with-independent-types"&gt;Q2. What are the necessary and sufficient conditions for FOA validity in the benchmark economy with independent types?&lt;/h3&gt;
&lt;p&gt;A: (Proposition 1) In the benchmark economy (symmetric, independent types, separable Pareto weights), FOA validity is equivalent to the condition that x·(1 + λ̃(x^{-γ})/2) is increasing in x, where λ̃(t) = [∫_t^∞ (1-α̃(w))g(w)dw] / (γtg(t)). The unidimensional analog requires x·(1 + λ̃(x^{-γ})) to be increasing. Since the bi-dimensional condition multiplies λ̃ by 1/2 rather than 1, the set of primitives satisfying it is strictly larger: every (G, α̃, γ) for which the unidimensional FOA holds also satisfies the bi-dimensional condition, but not vice versa. Economically, the FOA holds as long as the planner is not &amp;ldquo;too redistributive&amp;rdquo; — i.e., Pareto weights on low types are not so high as to violate these monotonicity conditions.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-coarea-formula-result-equation-27-and-why-is-it-the-central-technical-tool"&gt;Q3. What is the Coarea Formula result (equation 27) and why is it the central technical tool?&lt;/h3&gt;
&lt;p&gt;A: Given that the optimality conditions form a PDE system that cannot generally be solved pointwise, the authors integrate the optimality condition (equation 20) over subsets of the type space defined by level sets of an arbitrary function Q(w1, w2). The Coarea Formula allows them to express the result as: E[Σ_i λ*_i γ_i (∂lnQ/∂lnw_i) | Q=t] = [1 − E[α|Q≥t]] / [−∂ln P(Q≥t)/∂ln t]. By choosing different Q functions (e.g., Q = w_i, Q = max{k_1 w_1, k_2 w_2}, Q = R(w) for total family productivity, Q = I(w) for relative productivity), the formula delivers closed-form expressions for distinct conditional averages of optimal distortions, all expressed in terms of exogenous primitives. This contrasts with variational approaches (Golosov et al. 2014, Spiritus et al. 2022) that express optimal taxes in terms of endogenous moments.&lt;/p&gt;
&lt;h3 id="q4-how-do-optimal-distortions-on-married-individuals-compare-to-those-on-single-individuals-and-what-is-the-exact-quantitative-relationship-in-the-independent-types-benchmark"&gt;Q4. How do optimal distortions on married individuals compare to those on single individuals, and what is the exact quantitative relationship in the independent-types benchmark?&lt;/h3&gt;
&lt;p&gt;A: (Proposition 4) In the benchmark economy with independent types, the optimal distortion on spouse i with productivity t equals exactly one-half of the optimal distortion λ^{sng,&lt;em&gt;}(t) in the corresponding unidimensional economy: λ&lt;/em&gt;&lt;em&gt;i(t, w&lt;/em&gt;{-i}) = (1/2)λ^{sng,*}(t), and this is independent of the partner&amp;rsquo;s productivity w_{-i}. The intuition: the deadweight cost of taxing any individual depends only on her own characteristics (elasticity, productivity, density), not on whom she is married to. However, the redistributive benefit of taxation depends on matching — when matching is random, every high-productivity individual is married on average to an average person, so the incremental social benefit of extracting tax revenue from her is exactly half of what it would be if she were single (since half the benefit goes to a partner who is already average). More generally (Proposition 5 and Corollary 2), average distortions are weakly lower for married individuals than for singles as long as matching is not perfectly positively assortative.&lt;/p&gt;
&lt;h3 id="q5-what-is-average-jointness-and-how-is-it-measured"&gt;Q5. What is average jointness and how is it measured?&lt;/h3&gt;
&lt;p&gt;A: Average jointness J_i(t) is defined as the ratio of average distortions on spouse i conditional on the partner having above-t productivity to average distortions conditional on the partner having below-t productivity, minus one. Jointness is positive if the marginal tax rate on spouse i is on average increasing in the partner&amp;rsquo;s productivity, negative if decreasing, and zero for separable (individual earnings-based) taxes. The paper characterizes jointness through auxiliary functions H_i(t) (conditional distortion relative to unconditional average), whose behavior is determined by the copula elasticities η_i and survival copula elasticities η̄_i — the percentage change in the conditional quantile of the partner&amp;rsquo;s productivity when one spouse&amp;rsquo;s productivity quantile increases by 1%.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-role-of-tail-dependence-in-determining-the-sign-of-optimal-jointness"&gt;Q6. What is the role of tail dependence in determining the sign of optimal jointness?&lt;/h3&gt;
&lt;p&gt;A: (Proposition 7, Lemma 4) For right-tail dependent distributions — where the probability that an extremely productive person is married to an extremely productive partner remains bounded away from zero as productivity → ∞ — the redistributive benefit of positive jointness (targeting taxes to the richest couples) dominates its distortionary cost, so optimal average jointness is positive at the top. For right-tail independent distributions (where this probability converges to zero), the distortionary cost of positive jointness dominates, and optimal jointness is negative at the top. Exactly symmetric logic applies at the bottom using the survival copula and left-tail dependence. The bivariate lognormal/Gaussian copula is right-tail independent for any finite correlation ρ, while a distribution with perfect assortative matching in the tails would be right-tail dependent. The speed of convergence to tail independence, measured by κ = lim_{u→0} ln(u)/ln(C(u,u)) ∈ [1/2, 1), also matters: slower convergence (κ closer to 1) implies smaller optimal jointness under tail independence.&lt;/p&gt;
&lt;h3 id="q7-when-is-individual-earnings-based-separable-taxation-optimal-and-when-is-family-earnings-based-taxation-optimal"&gt;Q7. When is individual earnings-based (separable) taxation optimal, and when is family earnings-based taxation optimal?&lt;/h3&gt;
&lt;p&gt;A: (Propositions 4, 8, Corollary 1) Individual earnings-based taxation is optimal when Pareto weights are separable and spousal productivities are independent. When types are positively dependent, the planner introduces jointness even with separable social weights, because conditioning taxes on both spouses&amp;rsquo; earnings facilitates redistribution across couple types. Family earnings-based taxation is optimal when: (i) social weights are measurable only with respect to total family productivity r (i.e., the planner cares only about total family output, not the identity or relative productivity of individual spouses), and (ii) total family productivity r and relative spousal productivity ι are statistically independent. When r and ι are not independent, even a planner with an intrinsic preference for family earnings-based taxation will find it optimal to depart from it.&lt;/p&gt;
&lt;h3 id="q8-what-does-proposition-9-corollary-7-establish-about-the-relationship-between-restricted-and-unrestricted-optimal-taxes"&gt;Q8. What does Proposition 9 (Corollary 7) establish about the relationship between restricted and unrestricted optimal taxes?&lt;/h3&gt;
&lt;p&gt;A: (Corollary 7) For each restricted tax regime (anonymous, individual earnings-based, family earnings-based), the optimal distortions under the restricted tax equal the corresponding conditional average of unrestricted optimal distortions. Specifically: optimal individual earnings-based distortions equal E[λ*_i | w_i = t] (the average unrestricted distortion at productivity t); optimal family earnings-based distortions equal E[weighted average of λ*_i | R(w) = r]. This reveals that the unrestricted and restricted planners solve the same tradeoff between redistribution benefits and distortionary costs, but the restricted planner must apply a single tax rate to groups of couples that cannot be distinguished under the restriction. The welfare loss from restriction comes entirely from this forced bunching, not from a different objective or a different first-order condition.&lt;/p&gt;
&lt;h3 id="q9-what-do-the-quantitative-results-say-about-the-goodness-of-approximation-of-separable-vs-family-earnings-based-taxation"&gt;Q9. What do the quantitative results say about the goodness of approximation of separable vs. family earnings-based taxation?&lt;/h3&gt;
&lt;p&gt;A: In the calibrated benchmark economy (Gaussian copula, ρ = 0.33, Pareto-lognormal marginals, γ = 0.25, m = 0.35), optimal jointness is quantitatively small — the marginal tax rate on one spouse changes by at most several percentage points as a function of the other spouse&amp;rsquo;s earnings over the plotted range. Individual earnings-based (separable) taxation therefore provides a good approximation to the unrestricted optimum across all specifications considered. By contrast, family earnings-based taxation is a poor approximation: the marginal tax rate on family income varies substantially with the earnings share of the secondary earner (the ratio min{y1,y2}/(y1+y2)), and the deviation from the optimal unrestricted tax is large. This finding is robust across different Pareto weight specifications (m ∈ {0.35, 1.5}, k ∈ {0, 1, 2}) and holds even when k = 0, i.e., when the planner&amp;rsquo;s social weights inherently prefer family earnings-based taxation.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-calibration-results-relate-to-the-analytical-comparative-statics-predictions"&gt;Q10. How do the calibration results relate to the analytical comparative statics predictions?&lt;/h3&gt;
&lt;p&gt;A: The calibration validates the analytical predictions quantitatively. The analytical result (Proposition 5) that optimal distortions in the U.S. lie between those under random matching (1/2 of single-individual rates) and perfect assortative matching (same as single-individual rates) is confirmed: optimal tax rates for married individuals in the calibrated economy lie between the independence and perfect-dependence gray-line benchmarks in Figure 6. The analytical prediction (Proposition 7) that the Gaussian copula implies positive jointness at the bottom and negative at the top is confirmed, with the switch to negative jointness occurring above approximately $8.5 million in earnings. The slow convergence of the Gaussian copula to tail independence (κ = (1+ρ)/2 ≈ 0.665) explains the small magnitude of optimal jointness relative to the FGM copula (which has κ = 1/2, faster convergence, and exhibits more pronounced jointness as shown in the appendix). The analytical limiting distortion of E[λ*_i | w_i = t] → 1/(γa) ≈ 1.35 as t → ∞ (corresponding to a top marginal tax rate of approximately 55 percent) is confirmed, though convergence is slow and rates remain substantially below this limit at $300,000 in earnings.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-relate-to-and-advance-beyond-kleven-kreiner-and-saez-20072009"&gt;Q11. How does the paper relate to and advance beyond Kleven, Kreiner, and Saez (2007/2009)?&lt;/h3&gt;
&lt;p&gt;A: Kleven et al. (2009) studied couples taxation but avoided the multi-dimensional screening complexity by restricting the secondary earner to binary labor supply. The working paper by Kleven et al. (2007) considered the continuous setting but noted the difficulty of the FOA and derived several special-case insights. The current paper extends KKS in several systematic ways: it provides the first formal proof that the FOA conditions are strictly weaker in bi-dimensional than unidimensional settings; generalizes the formula for average distortions to arbitrary joint distributions (not just independent types); characterizes optimal jointness under positive dependence (not just independence); establishes the role of tail (in)dependence in determining the sign of jointness; compares optimal taxes for married vs. single individuals; and derives conditions under which family earnings-based or individual earnings-based taxation is optimal. It also shows that the KKS result on jointness sign (determined by the third derivative of the SWF) applies only under independence and can be reversed even with arbitrarily small positive dependence, as demonstrated with the Gaussian copula example.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;First-Order Approach (FOA) in multi-dimensional taxation.&lt;/strong&gt; The restriction of the mechanism design problem to local incentive constraints only — dropping global (non-local) incentive compatibility conditions and solving a relaxed problem. In the paper&amp;rsquo;s context, FOA validity is equivalent to convexity of a specific transformation vx* of the optimal utility function in the &amp;ldquo;linearized&amp;rdquo; type space X. The paper shows that the condition for FOA validity is strictly weaker (i.e., a strictly larger set of primitives satisfies it) in the bi-dimensional couples setting than in the corresponding unidimensional model, because the absence of participation constraints eliminates the main force driving FOA failure in industrial organization multi-dimensional screening.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;Optimal tax distortion λ&lt;/em&gt;_i(w).&lt;/em&gt;* The monotone transformation of the marginal tax rate defined by λ_i(w) = [∇_i T(y(w))] / [1 − ∇_i T(y(w))], where ∇_i T is the partial derivative of the tax function with respect to spouse i&amp;rsquo;s earnings. This transformation maps [−∞, ∞] marginal tax rates to (−1, ∞) distortions. The optimal tax schedule is characterized by the function λ* satisfying a system of PDEs; the paper studies conditional averages of λ* rather than λ* pointwise.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Coarea Formula.&lt;/strong&gt; A mathematical result from geometric measure theory that, in this context, converts an integral of the PDE optimality condition over a two-dimensional domain into an integral over the level sets of an arbitrary function Q(w). Applied to equation (20), it yields: E[Σ_i λ*_i γ_i (∂lnQ/∂lnw_i) | Q=t] = [1 − E[α|Q≥t]] / [−∂ln P(Q≥t)/∂ln t]. By choosing different Q functions, the formula delivers conditional averages of optimal distortions over different subsets of the type space, all in terms of exogenous primitives. This is the paper&amp;rsquo;s principal analytical tool for characterizing optimal taxes without solving the PDE explicitly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Jointness (positive/negative).&lt;/strong&gt; The dependence of the optimal marginal tax rate on one spouse&amp;rsquo;s earnings on the other spouse&amp;rsquo;s earnings. Taxes are positively jointed at w if ∂²T/∂y_1∂y_2 &amp;gt; 0 (so raising one spouse&amp;rsquo;s earnings increases the marginal tax rate on the other); negatively jointed if this cross-partial is negative; disjointed (separable) if it is zero. Average jointness J_i(t) at productivity t is measured as the ratio of conditional average distortions above and below the partner&amp;rsquo;s productivity threshold, minus one. Optimal jointness is the paper&amp;rsquo;s primary policy object for understanding how taxes on one spouse should respond to the other&amp;rsquo;s earnings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Copula and survival copula elasticities (η_i, η̄_i).&lt;/strong&gt; Defined as η_i(t) = ∂ln C(u)/∂ln u_i and η̄_i(t) = ∂ln C̄(u)/∂ln ū_i, where C is the copula of the joint productivity distribution, C̄ is the survival copula, and u_i = G_i(t_i), ū_i = 1−G_i(t_i) are the corresponding quantiles. These elasticities measure the percentage change in the conditional quantile of the partner&amp;rsquo;s productivity when one spouse&amp;rsquo;s productivity quantile increases by 1%. They quantify the additional distortionary cost introduced by jointness relative to a separable tax schedule: smaller elasticities (stronger dependence) correspond to larger distortionary costs of jointness at the boundaries of probability mass.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tail (in)dependence.&lt;/strong&gt; A joint distribution F is right-tail dependent if lim_{t→∞} P(w_{-i}≥t | w_i≥t) &amp;gt; 0, i.e., extremely productive individuals have a positive probability of being matched with equally extreme partners. It is right-tail independent if this limit is zero. The speed of convergence to tail independence is measured by κ = lim_{u→0} ln(u)/ln(C(u,u)) ∈ [1/2, 1). Tail dependence determines the sign of optimal average jointness in the tails: right-tail dependence favors positive jointness at the top; right-tail independence favors negative jointness at the top. The Gaussian copula is right-tail independent for any finite ρ; a perfectly assortative matching distribution is right-tail dependent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Positive quadrant dependence (PQD) order.&lt;/strong&gt; A partial ordering on joint distributions with the same marginals: F^b ≥_{PQD} F^a if F^b(w) ≥ F^a(w) for all w, equivalently if Cov(φ_1(w_1), φ_2(w_2)) ≥ 0 for any two increasing functions. The paper uses this order to rank economies by the &amp;ldquo;assortativeness&amp;rdquo; of matching, and shows that optimal average distortions are monotone in this order (Proposition 5): more assortative matching implies weakly higher optimal tax distortions on each married individual.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pareto-lognormal (PLN) distribution.&lt;/strong&gt; Used in the calibration to model the marginal distribution of spousal productivities. Defined as G(t) = Φ((ln t − μ)/σ) − a·exp(aμ + a²σ²/2)·Φ((ln t − μ)/σ − aσ), parameterized by location μ, scale σ, and tail parameter a. The PLN family has a lognormal body and a Pareto tail with tail parameter a, making it suitable for capturing the empirical finding of a thin left tail (implying optimal marginal taxes approaching zero as earnings → 0) and a thick right tail (implying a positive limiting marginal tax rate of approximately 1/(1 + 1/(γa)) as earnings → ∞).&lt;/p&gt;</description></item><item><title>The Origins and Control of Forest Fires in the Tropics</title><link>https://macropaperwarehouse.com/papers/the-origins-and-control-of-forest-fires-in-the-tropics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-origins-and-control-of-forest-fires-in-the-tropics/</guid><description>&lt;p&gt;This paper studies the economics of illegal tropical forest fires in Indonesia, framed as a modern counterpart to Pigou&amp;rsquo;s canonical externality example of sparks from railway engines. The central research question is whether private firms adjust their fire-setting behavior depending on the degree to which the costs of fire spread fall on themselves versus others, and what enforcement architecture shapes that adjustment.&lt;/p&gt;
&lt;p&gt;The empirical setting is Indonesia&amp;rsquo;s national forest estate, where palm oil and wood fiber concession holders use fire as a cheap land-clearance method — burning primary forest costs 44–70% less than mechanical clearance — despite the practice being illegal. The paper assembles a novel dataset of 107,334 fires across Indonesia&amp;rsquo;s major forested islands from October 2000 to January 2016, constructed from NASA MODIS daily satellite hotspot data (1 km resolution, four flyovers per day). Fire ignitions and spread paths are traced by linking contiguous pixels burning on adjacent days. This fire data is merged with geocoded concession boundaries (logging, palm oil, wood fiber), land-use classifications (protected forest, unleased productive forest, areas outside the forest estate), annual deforestation data from Hansen et al. (2013) at 30 m resolution, daily wind speed data from NOAA NCEP-DOE Reanalysis 2 interpolated to each 1 km pixel, and data on firms investigated by the Indonesian government following the 2015 fires. The main analytical sample focuses on the 39,077 fires started inside wood fiber and palm oil concessions.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s identification strategy exploits two intersecting sources of variation: (1) temporal and spatial variation in monthly wind speed, which predicts the probability and extent of fire spread — a one-standard-deviation increase in wind speed (approximately 5 km/hr) increases fire spread area by 287%; and (2) cross-sectional variation in the land-type composition of the area surrounding each ignition pixel, which determines whether spread costs would fall on the fire-setter or on others. The interaction of these two factors identifies whether firms are more cautious about igniting fires on windy days when surrounding land is their own versus when it belongs to others.&lt;/p&gt;
&lt;p&gt;Three main findings emerge. First, fires are systematically human-caused and linked to industrial land clearance. Fires are eight times more likely per hectare in oil palm and wood fiber concessions than in logging concessions. Completely deforesting a 1 km pixel increases the probability of fire ignition in that pixel in the subsequent year by 279%, and this effect reverses in the year after (two years post-deforestation), ruling out natural flammability as the explanation and confirming a deliberate slash-and-burn cycle. Fire use following deforestation falls by approximately 38% in oil palm concessions during district election years, consistent with tighter enforcement when political incentives favor suppression.&lt;/p&gt;
&lt;p&gt;Second, firms partially internalize the externalities from fire-setting. They are significantly less likely to set fires on windy days when surrounding pixels belong to their own concession rather than to others. A buffer zone entirely owned by the same concession holder reduces ignitions by 8–25% at mean wind speed, and by 22–61% at the 95th-percentile wind speed. However, firms treat neighboring concession land and unleased productive forest similarly — suggesting Coasian bargaining between concession holders is not occurring.&lt;/p&gt;
&lt;p&gt;Third, the government&amp;rsquo;s enforcement pattern shapes firm behavior. Using data on firms investigated after the 2015 fires, the paper shows the government disproportionately investigates firms whose fires burned protected areas or high-population-density land, but not those whose fires damaged other private concessions. The relative weights firms place on different land types when deciding whether to ignite fires align closely with this government punishment function, consistent with firms responding to implicit Pigouvian incentives.&lt;/p&gt;
&lt;p&gt;Counterfactual simulations show that broadening enforcement to treat all land types as the government currently treats populated areas would reduce fires by 80%; treating all land like protected forest would reduce fires by 67%. By contrast, fully Coasian property-rights solutions yield only 14% reductions, and tort reform allowing concession holders to recover damages from neighbors yields only 6%.&lt;/p&gt;
&lt;p&gt;Q: What is the core externality problem studied in this paper?
A: Firms use fire as a cheap land-clearance method, but once set, fires risk spreading beyond the igniter&amp;rsquo;s own concession onto land owned by others, creating an uncompensated externality. The decision to use fire rather than mechanical clearance is de facto a decision to impose this spread risk on third parties. The paper asks whether firms adjust this decision depending on the extent to which spread costs fall on themselves versus others, and whether government enforcement shapes that adjustment.&lt;/p&gt;
&lt;p&gt;Q: Why is Indonesia the empirical setting?
A: Indonesia holds a large share of the world&amp;rsquo;s tropical forests and is among the countries most affected by illegal land-clearing fires. The 2015 Indonesian fires alone released approximately 400 megatons of CO2 equivalent, at their peak emitting more daily greenhouse gases than all US economic activity, and caused an estimated 100,000 excess deaths across Indonesia, Malaysia, and Singapore. The palm oil industry in Indonesia and Malaysia, where fire is used extensively, accounted for 4.7% of global CO2 emissions from 1986 to 2016.&lt;/p&gt;
&lt;p&gt;Q: How are fire ignitions and spread identified in the data?
A: The paper starts from NASA MODIS daily hotspot data at 1 km resolution from October 2000 to January 2016. An iterative procedure assigns contiguous pixels burning on adjacent days to the same fire event, with a 1-pixel buffer allowing for spread detection. This yields 176,855 total fires across Indonesia, of which 107,334 remain after restricting to the major forested islands and the forest estate. The procedure may understate single-day spread since pixels burning on the same day are classified as part of the ignition area rather than spread.&lt;/p&gt;
&lt;p&gt;Q: What fraction of fires spread beyond their ignition area, and how much of the spread falls on outsiders?
A: 87% of fires burn for only one day and 89% do not spread beyond their initial ignition area. However, the largest fire in the data spread to cover 466 times its initial area, and the largest single fire burned 764 km2. Across all multi-day fires started inside concessions, 32% of the total land burned outside the initial ignition area is outside the concession where the fire began, quantifying the scale of the local externality.&lt;/p&gt;
&lt;p&gt;Q: How is wind speed used as an identification strategy?
A: Wind speed provides temporal and spatial variation in the probability that a fire will spread. A one-standard-deviation increase in wind speed (approximately 5 km/hr) increases the extent of fire spread by 287%. Because wind varies month to month and across space, while the composition of surrounding land types is fixed in the cross-section, the interaction of wind speed with surrounding land type identifies whether firms are more cautious about igniting fires when spread risk is high and spread costs would fall on their own land versus others&amp;rsquo; land.&lt;/p&gt;
&lt;p&gt;Q: What is the main result on firms&amp;rsquo; internalization of fire spread externalities?
A: Firms are significantly less likely to start fires on windy days when a larger share of the surrounding buffer zone belongs to their own concession. One additional buffer pixel in one&amp;rsquo;s own land decreases ignitions by 0.2–0.7%. A buffer zone entirely owned by the same concession holder reduces ignitions by 8–25% at mean wind speed, and by 22–61% at the 95th-percentile wind speed. This demonstrates that firms take fire spread risk into account when it threatens their own assets, but discount it when spread would damage others&amp;rsquo; land.&lt;/p&gt;
&lt;p&gt;Q: Do firms treat different types of neighboring land differently?
A: Yes. The benchmark category is unleased productive forest, which has the weakest property rights and receives the least de facto government protection. Relative to this benchmark, firms are more cautious about fire spread toward protected forest (national parks and watershed areas) and toward land outside the forest estate (typically villages and smallholders). One additional buffer pixel in protected forest versus unleased productive forest decreases ignitions by 0.9% at mean wind speed and 2.7% at the 95th-percentile wind speed; the deterrent for land outside the forest estate is even stronger at 1.6% and 4.6%, respectively. Firms treat other firms&amp;rsquo; concession land similarly to unleased productive forest, suggesting no effective private enforcement between concession holders.&lt;/p&gt;
&lt;p&gt;Q: What evidence shows fires are tied to intentional land clearance rather than natural ignition?
A: Fires are eight times more likely per hectare in oil palm and wood fiber concessions than in logging concessions, consistent with clear-cutting versus selective logging. Completely deforesting a 1 km pixel increases fire probability in that pixel in the subsequent year by 279%. Crucially, the effect reverses in the second year after deforestation — the pixel becomes less likely to burn than before — which rules out natural flammability as the mechanism and confirms deliberate slash-and-burn timing.&lt;/p&gt;
&lt;p&gt;Q: What does the electoral cycle evidence show about government enforcement?
A: Fires following deforestation fall by approximately 38% in oil palm concessions during district election years relative to the year prior to an election, and bounce back to pre-election levels in the year after. The decline is confined to productive forest zones where conversion is occurring; no electoral cycle appears in protected areas where conversion is already prohibited. This indicates that enforcement is tightened when political incentives are strong, and confirms that these fires are set intentionally and are responsive to government pressure.&lt;/p&gt;
&lt;p&gt;Q: How is the government&amp;rsquo;s de facto punishment function estimated?
A: The paper uses data on firms investigated by the Indonesian Ministry of Forestry following the 2015 fires, matching investigated firms (identified only by initials in the published list) to concession-holder names. A logistic regression of investigation probability on the land-type outcomes of a firm&amp;rsquo;s fires — conditional on total area burned — shows the government is substantially more likely to investigate firms whose fires burned protected areas or high-population-density land, but does not differentially investigate cases where fire damage is largely confined to other private concessions.&lt;/p&gt;
&lt;p&gt;Q: How closely do firm behavior and government enforcement weights align?
A: The relative weights across land types that the government applies in its investigation decisions correspond closely to the relative weights firms apply when deciding whether to ignite fires on windy days. Firms are most deterred by spread risk toward protected forest and populated areas outside the forest estate — the same categories the government prioritizes. Firms are least deterred by spread toward unleased productive forest and other private concessions — the categories the government largely ignores. This alignment is consistent with firms responding to Pigouvian-style implicit incentives generated by the government&amp;rsquo;s enforcement pattern.&lt;/p&gt;
&lt;p&gt;Q: What do the counterfactuals reveal about policy effectiveness?
A: Fully Coasian property-rights reform — where firms treat all surrounding land as their own — would reduce fires by only 14%. Tort reform enabling concession holders to recover damages from neighbors (treating neighboring concessions as own land) would reduce fires by only 6%. By contrast, uniform enforcement raising deterrence to the level currently applied to populated areas would reduce fires by 80%; applying the level currently applied to protected forest would reduce fires by 67%. An enforcement regime that perfectly prevented all fire spread outside the igniting concession would reduce area burned by only 23%; preventing spread into protected and populated areas alone would yield only a 2% reduction.&lt;/p&gt;
&lt;p&gt;Q: What do the benefit-cost ratios for fires look like?
A: The estimated external damages from the 1997/1998 Indonesian fires range from 1,286 to 6,074 USD per hectare burned (2020 USD). The average private benefit from using fire rather than mechanical clearance — accounting for fertilizers and other costs — averages approximately 52 USD per hectare (2020 USD). Benefit-cost ratios of 0.008 to 0.04 lie well below 1, indicating that the social damages from fires vastly exceed the private benefits, even though the government currently deters only the most costly categories of fire.&lt;/p&gt;
&lt;p&gt;Q: Why do Coasian private solutions perform poorly in this setting?
A: Coasian bargaining between concession holders would require them to reach agreements to bring fire use to a locally efficient level without government intervention. The evidence shows firms treat other concession holders&amp;rsquo; land essentially the same as unprotected unleased productive forest, implying that no such bargains are being struck. The counterfactual analysis confirms this: even a fully-Coasian outcome where every surrounding pixel is treated as own land would reduce fires by only 14%, because the bulk of fires occur when ignition costs to the firm&amp;rsquo;s own land are low regardless of wind speed.&lt;/p&gt;
&lt;p&gt;Q: What is the primary policy implication?
A: The most effective lever for reducing fires is not preventing spread after the fact, but rather deterring ignition in the first place by extending the enforcement regime uniformly across all land types. If firms were induced to treat all surrounding land with the same caution they currently apply toward populated areas — through broader and stronger penalties — fires would fall by 80%. This is substantially more effective than property-rights reforms, tort reforms, or targeted spread-prevention measures focused only on protected and populated areas.&lt;/p&gt;
&lt;p&gt;Externality (fire spread): In this paper&amp;rsquo;s usage, the cost imposed on third parties when a fire ignited inside one concession spreads to land owned by others. The externality is quantified as the share of area burned outside the igniting concession (32% of multi-day fire spread in the data) and the ratio of external damages (1,286–6,074 USD/ha) to private benefits (52 USD/ha) from using fire rather than mechanical clearance.&lt;/p&gt;
&lt;p&gt;Slash-and-burn (industrial scale): The two-stage land-clearance practice where valuable timber is first harvested (deforestation) and the remaining vegetation is then burned to prepare land for plantation crops. The paper establishes this cycle empirically: complete deforestation of a 1 km pixel increases fire ignitions by 279% in the following year, with the effect reversing in the second year, ruling out natural flammability.&lt;/p&gt;
&lt;p&gt;Pigouvian enforcement: Government-imposed penalties that alter private incentives to account for externalities. In this paper&amp;rsquo;s usage, the government&amp;rsquo;s de facto punishment function — which heavily weights fires spreading into protected areas and populated land — functions as an implicit Pigouvian tax, shaping which fires firms choose to avoid rather than uniformly deterring all illegal burning.&lt;/p&gt;
&lt;p&gt;Coasian bargaining failure: The absence of private negotiations between concession holders to internalize the externalities they impose on each other. The paper demonstrates this failure empirically by showing firms treat neighboring concession land no differently from unprotected unleased productive forest, indicating no effective private agreements are limiting cross-concession fire spread.&lt;/p&gt;
&lt;p&gt;Wind speed as spread risk shifter: Monthly average wind speed at each 1 km pixel, used as the time-varying component of fire spread risk. A one-standard-deviation increase (approximately 5 km/hr) increases fire spread area by 287%. The paper uses wind speed variation interacted with surrounding land type composition to identify whether firms adjust ignition decisions based on spread risk and who bears the cost.&lt;/p&gt;
&lt;p&gt;Unleased productive forest (benchmark): Land within the national forest estate that is neither in a designated concession nor in a protected zone, leaving ownership rights unclear and de facto unprotected. The paper uses firms&amp;rsquo; behavior toward this category as the baseline against which sensitivity to other land types is measured, because it attracts the least government attention and the weakest property rights.&lt;/p&gt;
&lt;p&gt;Government punishment function: The implicit weights the Indonesian government places on different types of fire damage when deciding whether to investigate a firm, estimated from logistic regression on the 2015 investigation data. The function heavily weights fires burning protected areas and high-population-density land, and places near-zero weight on damage to other private concessions, shaping which fire types firms strategically avoid.&lt;/p&gt;</description></item><item><title>The Productivity of Professions: Evidence from the Emergency Department</title><link>https://macropaperwarehouse.com/papers/the-productivity-of-professions-evidence-from-the-emergency-department/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-productivity-of-professions-evidence-from-the-emergency-department/</guid><description>&lt;p&gt;This paper studies the productivity of nurse practitioners (NPs) versus physicians performing overlapping tasks in Veterans Health Administration (VHA) emergency departments (EDs), exploiting a quasi-experiment created by the VHA&amp;rsquo;s December 2016 grant of full practice authority to NPs. The identification strategy instruments patient assignment to NPs versus physicians using quasi-random variation in the number of NPs on duty on a given ED-day, conditional on ED-by-time-category fixed effects. The sample covers 1.1 million ED visits across 44 VHA EDs from January 2017 to January 2020, seen by 1,348 physicians and 156 NPs. The instrument is validated by demonstrating balance in patient observable characteristics across values of the instrument, stability of IV estimates across 256 combinations of patient covariate controls, and absence of spillover effects from NP presence onto physician performance.&lt;/p&gt;
&lt;p&gt;On average in the ED setting, NPs increase patient length of stay by 11 percent (approximately 18 additional minutes) and raise the cost of the ED visit by 7 percent (approximately $66 per visit). NPs raise the 30-day preventable hospitalization rate by 0.25 percentage points, a 20 percent increase relative to the mean. No statistically significant effect on 30-day mortality is detected (95 percent confidence interval: -0.34 to 0.11 percentage points). OLS estimates carry the opposite sign because NPs are assigned healthier patients in observational data; the IV design corrects for this selection.&lt;/p&gt;
&lt;p&gt;The average NP-physician performance gap varies systematically by case complexity and severity. For the highest-complexity quartile of cases (by Elixhauser comorbidities), NPs increase ED costs by 12 percent and length of stay by 28 percent. For cases at or above the 95th percentile of severity (based on 30-day mortality by diagnosis), NPs increase ED costs by 25 percent, length of stay by 99 percent, and admissions by 26 percentage points (42 percent relative to the mean), while reducing 30-day preventable hospitalization by 3 percentage points — suggesting that NPs&amp;rsquo; higher care intensity partially offsets worse intrinsic skill for the most severe cases. For lower-complexity cases, the cost and length-of-stay gaps are smaller, but NPs still significantly raise preventable hospitalizations.&lt;/p&gt;
&lt;p&gt;NPs exhibit clinical decision-making patterns consistent with lower diagnostic skill: they are more likely to order consults (2.6 percentage points, or 11 percent of the mean), CT scans (1.2 percentage points, or 8.3 percent), and X-rays (2.0 percentage points, or 6.9 percent). NPs lower opioid prescriptions by 1.8 percentage points (20 percent of the mean) and raise antibiotic prescriptions by 4.0 percentage points (6.3 percent of the mean), consistent with threshold adjustment under lower diagnostic skill with asymmetric error costs. Downstream, patients treated by NPs incur similar opioid use disorder rates despite lower opioid prescribing, and higher infection-related return visit rates despite higher antibiotic prescribing.&lt;/p&gt;
&lt;p&gt;Counterfactual analysis finds that allocating one quarter of ED patients to NPs increases net spending by $129 million per year to the VHA after accounting for NPs&amp;rsquo; lower wages (approximately half of physicians&amp;rsquo;). However, deploying NPs exclusively to the least-complex quarter of cases reduces net spending to approximately one-fifth of this amount.&lt;/p&gt;
&lt;p&gt;A distributional analysis deconvolving provider-specific IV estimates reveals that within-profession productivity variation substantially exceeds the average between-profession gap. The interquartile range in annual spending attributable to provider productivity within each profession is approximately $900,000, roughly three times the mean annual spending difference between the average NP and the average physician. A randomly chosen NP outperforms a randomly chosen physician in up to 38 percent of pairs. Within professions, individual provider productivity shows essentially no relationship with wages or case complexity assigned, whereas between professions, case assignment and wages are strongly sorted by professional class.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question?
A: The paper asks whether NPs and physicians, who perform overlapping tasks in the ED but differ sharply in training, selectivity, and pay, differ in productivity, and how that average between-profession difference compares to productivity variation within each profession. It also asks what mechanisms drive any observed gap and how case assignment responds to provider skill differences.&lt;/p&gt;
&lt;p&gt;Q: What is the identification strategy and why is it credible?
A: The authors instrument patient assignment to NPs with the number of NPs on duty on the ED-day, conditional on ED-by-year, ED-by-month, ED-by-day-of-week, and ED-by-hour fixed effects. Credibility rests on: provider schedules being set months in advance, decoupling NP availability from arriving patient characteristics; patient characteristics being well balanced across values of the instrument conditional on fixed effects; IV estimates being stable across all 256 covariate-control combinations; and on-duty physician and NP characteristics also being balanced across the instrument.&lt;/p&gt;
&lt;p&gt;Q: What are the main average effects of NPs on resource use?
A: IV estimates show NPs increase patient length of stay by 11 percent (approximately 18 minutes) and ED cost by 7 percent (approximately $66 per visit). There is no significant average effect on inpatient admissions in the overall sample, though NPs significantly raise admissions for high-severity cases.&lt;/p&gt;
&lt;p&gt;Q: What is the effect of NPs on patient health outcomes?
A: NPs raise 30-day preventable hospitalizations by 0.25 percentage points, a 20 percent increase relative to the mean. The 95 percent confidence interval for 30-day mortality is -0.34 to 0.11 percentage points, implying no statistically significant mortality effect in the overall sample.&lt;/p&gt;
&lt;p&gt;Q: Why do OLS and IV estimates have opposite signs?
A: In observational data, NPs treat healthier patients than physicians: NP patients are younger (60.7 versus 62.5 years), have fewer Elixhauser comorbidities (3.2 versus 3.7), and have fewer prior inpatient stays (0.4 versus 0.7). This selection causes OLS estimates of NP effects to be negative. The IV corrects for this by exploiting quasi-random variation in NP availability; IV estimates are stable across all combinations of patient controls, consistent with the instrument being orthogonal to unobservable patient health.&lt;/p&gt;
&lt;p&gt;Q: How does the NP-physician performance gap vary with case complexity and severity?
A: For the highest-complexity quartile, NPs increase length of stay by 28 percent and ED costs by 12 percent without a significant preventable hospitalization effect. For cases at or above the 95th severity percentile, NPs increase length of stay by 99 percent, ED costs by 25 percent, and admissions by 26 percentage points (42 percent relative to the mean), while reducing 30-day preventable hospitalization by 3 percentage points. For lower-complexity quartiles, NPs show smaller cost and length-of-stay effects but significantly raise preventable hospitalizations, suggesting the higher care intensity at high severity compensates for lower skill.&lt;/p&gt;
&lt;p&gt;Q: What does the heterogeneity by severity imply for optimal case assignment?
A: The pattern is consistent with skill-task matching: NPs have a comparative and absolute disadvantage in complex cases, so optimal assignment directs less complex cases to NPs and fewer patients to NPs when physicians are more available. Empirically, NPs are indeed assigned healthier patients from the available pool, and are assigned a modestly smaller share when the ED is less busy.&lt;/p&gt;
&lt;p&gt;Q: What mechanisms explain the average NP-physician gap?
A: Three mechanisms are examined. First, experience: a one-standard-deviation increase in specific experience is associated with a 5.8 percent decline in the NP-physician length-of-stay gap, and general experience with a 10 percent decline; however, experience does not significantly narrow the preventable hospitalization gap. Second, information acquisition: NPs order more consults, CT scans, and X-rays, consistent with compensating for lower diagnostic skill. Third, prescription thresholds: NPs reduce opioid prescribing by 20 percent and raise antibiotic prescribing by 6.3 percent, consistent with threshold adjustment under asymmetric error costs, but downstream outcomes are not improved correspondingly.&lt;/p&gt;
&lt;p&gt;Q: What do prescription patterns and downstream outcomes reveal about NP diagnostic skill?
A: NPs prescribe fewer opioids yet patients treated by NPs obtain similar downstream opioid use disorder rates; NPs prescribe more antibiotics yet patients treated by NPs have higher rates of return visits with infections. This pattern is consistent with NPs exhibiting higher rates of both false positives and false negatives, not merely adjusted thresholds, suggesting genuinely lower diagnostic skill rather than threshold differences alone.&lt;/p&gt;
&lt;p&gt;Q: What do counterfactual cost calculations show?
A: Allocating one quarter of ED patients to NPs raises non-wage spending by $197 million per year to the VHA; after accounting for NP wages being half of physician wages (approximately $120,000 versus $240,000 per year), net cost is still $129 million per year. Restricting NP deployment to the least-complex quarter of cases reduces net spending to approximately one-fifth of this amount, illustrating that targeted case assignment substantially improves NP cost-effectiveness.&lt;/p&gt;
&lt;p&gt;Q: How large is within-profession productivity variation relative to between-profession differences?
A: The interquartile range in annual spending attributable to provider productivity within each profession is approximately $900,000, roughly three times the mean annual spending difference between the average NP and the average physician. A randomly chosen NP outperforms a randomly chosen physician in up to 38 percent of random pairs. The authors conclude that, despite stark differences in training and selection between professions, within-profession variation dominates.&lt;/p&gt;
&lt;p&gt;Q: Is individual provider productivity reflected in wages or case assignment within professions?
A: Within each profession, provider productivity shows essentially no relationship with wages or with the complexity of assigned cases. This contrasts sharply with between-profession patterns, where professional class strongly predicts both wages (NPs earn approximately $120,000 per year versus $240,000 for physicians) and assigned case complexity. The authors interpret this as evidence of informational and organizational frictions in recognizing individual productivity within professional classes, and note that professional class is a far stronger predictor of pay and case assignment than is individual productivity.&lt;/p&gt;
&lt;p&gt;Q: How do complier characteristics relate to the broader patient population?
A: Compliers — cases whose provider type is determined by the instrument — are healthier than the average case: younger, with fewer comorbidities, fewer prior inpatient stays, and lower predicted mortality. Never-takers are riskier than the average case. There are no always-takers since patients cannot be assigned to NPs on days when no NPs are on duty.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to the literature on NP scope-of-practice laws?
A: The scope-of-practice literature estimates general-equilibrium effects of allowing NPs greater autonomy, including labor reallocation between professions. This paper instead estimates the partial-equilibrium causal effect of assigning a patient to an NP versus a physician, holding the broader labor market fixed. The two literatures are complementary: the heterogeneity findings here suggest that scope-of-practice expansions may be more beneficial in lower-complexity primary care settings where the NP-physician performance gap is smaller.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: Three implications are highlighted. First, the efficiency of using NPs depends critically on case assignment: deploying NPs on the least-complex cases reduces net costs to approximately one-fifth of indiscriminate deployment. Second, the substantial overlap between NP and physician productivity distributions provides support for NP use in less complex settings even within the ED context. Third, within-profession productivity variation far exceeding between-profession differences suggests that individual-level productivity assessment, rather than professional class, may be a more accurate guide to case assignment and compensation.&lt;/p&gt;
&lt;p&gt;Quasi-experimental variation in NP availability: The identification strategy exploits day-to-day variation in the number of NPs scheduled to work in a given VHA ED, conditional on ED-by-time-category fixed effects, as an instrument for whether a patient is assigned to an NP versus a physician. Schedules are set months in advance, rendering the NP count orthogonal to arriving patient characteristics conditional on those fixed effects.&lt;/p&gt;
&lt;p&gt;30-day preventable hospitalization: A standardized quality-of-care outcome defined by the Agency for Healthcare Research and Quality, measuring hospitalizations occurring within 30 days of ED discharge that are classified as preventable given adequate prior outpatient management. Used by the paper as the primary downstream health outcome beyond the ED visit itself.&lt;/p&gt;
&lt;p&gt;Elixhauser comorbidities: A set of 31 binary indicators for chronic conditions (e.g., cancer, diabetes) based on medical histories in the prior 365 days, used in this paper to measure and stratify case complexity into quartiles for heterogeneity analysis.&lt;/p&gt;
&lt;p&gt;Productivity distributions within professions: Provider-specific productivity estimates derived from a just-identified IV model that instruments assignment to individual providers by indicators for on-duty providers, then deconvolved into underlying distributions using the Efron (2016) and Kline-Rose-Walters (2022) method. These distributions characterize the spread of productivity within each professional class, separate from measurement error.&lt;/p&gt;
&lt;p&gt;Prescription threshold adjustment: The mechanism, formalized in Chan, Gentzkow, and Yu (2022), by which providers with lower diagnostic skill optimally adjust treatment thresholds in response to asymmetric costs of false-positive versus false-negative errors. In this paper&amp;rsquo;s application, NPs lower the opioid prescription rate (where false positives carry higher costs: addiction and overdose) and raise the antibiotic prescription rate (where false negatives carry higher costs: untreated infection), but downstream outcomes do not improve correspondingly.&lt;/p&gt;
&lt;p&gt;Skill-task matching: The organizational economics principle (Acemoglu and Autor 2011) that efficiency requires assigning more complex tasks to higher-skilled workers. The paper documents that between professions, case assignment broadly follows this principle (NPs receive less complex patients on average), but within professions, essentially no matching between individual provider productivity and case complexity is observed.&lt;/p&gt;
&lt;p&gt;Full practice authority (VHA, December 2016): The VHA policy that allowed NPs to treat patients independently without physician supervision at VHA facilities, superseding state-level restrictions. This policy change defines the start of the paper&amp;rsquo;s sample period and establishes the institutional context in which the quasi-experiment occurs, as it removed the requirement for physician oversight that previously constrained NP independence.&lt;/p&gt;</description></item><item><title>The Social Tax: Redistributive Pressure and Labor Supply</title><link>https://macropaperwarehouse.com/papers/the-social-tax-redistributive-pressure-and-labor-supply/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-social-tax-redistributive-pressure-and-labor-supply/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks whether informal redistributive pressure — the social obligation to share earned income with kin and social networks — distorts labor supply in low-income communities. The authors conceptualize such pressure as a &amp;ldquo;social tax&amp;rdquo; on earnings and develop the first direct causal test of whether it reduces labor supply, output, and earnings among full-time workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and Sample&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The study works with 474 full-time piece-rate factory workers (464 of whom are women) employed in cashew processing plants run by Olam in Côte d&amp;rsquo;Ivoire. Workers are paid biweekly in cash entirely through piece rates for individual nut-peeling output, creating a direct mapping between labor supply and income. At baseline, workers report transferring 25–35% of their income to individuals outside their household, with 77% having made at least one transfer in the previous 3 months. Workers also strongly believe that earning more triggers more transfer requests: 77% agree that if someone starts earning more by working harder, people will ask that person more often for financial help.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intervention&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors introduce a blocked savings account into which workers can deposit any earnings above a self-chosen threshold (set at least as high as their own baseline average earnings). Earnings above the threshold are automatically deposited by the factory directly into the account with the Banque Populaire de Côte d&amp;rsquo;Ivoire; the cash component of pay is unchanged. Funds cannot be withdrawn until the end of the blocked period (9 months in Phase 1; 3 months in Phase 2). The key design feature is that the account reduces the effective social tax rate only on earnings &lt;em&gt;increases&lt;/em&gt; above baseline, thereby eliminating income effects and generating only a pure substitution effect — an unambiguous positive prediction on labor supply if a social tax exists.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Experimental Design&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Workers are randomized into three conditions: (1) Control (no account); (2) Private account (existence unknown to anyone outside the worker); (3) Non-private account (existence and forthcoming unblock date revealed to network members via promotional text messages). The contrast between Private and Non-private isolates the role of redistributive pressure specifically — holding constant all other features of the blocked account product. The experiment runs in two cross-randomized phases conducted between 2018 and 2019.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Take-up of blocked accounts is dramatically higher when accounts are private: 60% in Phase 2 (Private) versus 14% (Non-private), a 77% decline (p&amp;lt;0.001). Among workers who declined Non-private accounts, 96% cite anticipated increases in transfer requests as an important factor.&lt;/p&gt;
&lt;p&gt;Being offered a Private account sharply raises labor supply. Pooling both phases, the Private arm increases average daily earnings by 175.9 FCFA, or &lt;strong&gt;11.4%&lt;/strong&gt; (p=0.012), relative to Control or Non-private arms. This is accompanied by a &lt;strong&gt;6.2 percentage point (9.7%)&lt;/strong&gt; increase in daily work attendance (p=0.023), with the entire attendance effect driven by reduced absenteeism rather than turnover. Effects in Phase 1 (Private vs. Control: +11.3%, p=0.032) and Phase 2 (Private vs. Non-private: +11.5%, p=0.043) are nearly identical in magnitude, indicating the results are not sensitive to cross-phase design. The treatment effect magnitude is equivalent to each worker working an additional 1.19 days in every two-week paycycle. Because 89% of workers have no income outside the factory, these constitute increases in total earned income.&lt;/p&gt;
&lt;p&gt;Heterogeneity is consistent with the hypothesized mechanism: among workers who report difficulty saving due to redistributive pressure, the Private treatment increases earnings by &lt;strong&gt;15.0%&lt;/strong&gt; (p=0.018); among those not reporting such difficulty, the estimated effect is near zero and insignificant (p=0.95). Among workers who report transfers to acquaintances (the most likely social-tax-motivated transfers), the effect is &lt;strong&gt;17.5%&lt;/strong&gt; (p=0.014). Workers without a partner — for whom intra-household redistribution is irrelevant — experience a &lt;strong&gt;15.8%&lt;/strong&gt; earnings increase (p=0.017), indicating that extra-household pressure drives the results.&lt;/p&gt;
&lt;p&gt;Outgoing transfers do not decline. The design leaves cash-on-hand unchanged by construction, and consistent with this, there is no significant change in the likelihood or amount of transfers from treated workers to their networks. Total outgoing transfers are if anything higher among Private account workers (p=0.049), suggesting no loss in redistribution to the network.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Social Tax Rate Estimation&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Combining the 11.4% treatment effect on output with a labor supply elasticity estimated from an end-of-experiment piece-rate randomization (intensive-margin elasticity of 0.17; total elasticity of approximately 1.11), the authors estimate the social tax rate for the average worker in the sample at &lt;strong&gt;9–14%&lt;/strong&gt;. For the subset who actually take up Private accounts, the implied social tax rate is &lt;strong&gt;19–23%&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Results pertain to full-time female piece-rate workers in formal cashew processing plants in Côte d&amp;rsquo;Ivoire, with average tenure of 1.7 years. Because the intervention lowers the tax only on earnings &lt;em&gt;above&lt;/em&gt; baseline (not on all earnings), the estimates do not directly capture the total distortion from eliminating all redistributive pressure. Alternative confounds — fairness/morale effects, self-control, privacy concerns, goal-setting — are each tested and ruled out as primary drivers.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-theoretical-basis-for-predicting-that-private-accounts-unambiguously-increase-labor-supply"&gt;Q1. What is the theoretical basis for predicting that Private accounts unambiguously increase labor supply?&lt;/h3&gt;
&lt;p&gt;The authors model redistributive pressure as a social tax rate τ₁ on gross earnings. The blocked account reduces this tax to τ₂ &amp;lt; τ₁ only on earnings &lt;em&gt;above&lt;/em&gt; baseline labor supply e₁, creating a kink in the budget constraint. Starting from e₁, the worker faces only a pure substitution effect (no income effect) when τ₂ falls, because her net earnings at e₁ are unchanged. Equation (2) in the paper shows formally that the income effect term drops out, and the derivative of labor supply with respect to τ₂ is unambiguously negative (i.e., reducing τ₂ increases effort). This &amp;ldquo;clean&amp;rdquo; prediction — no income effect, no ambiguity — is the central design advantage relative to simply shielding existing earnings.&lt;/p&gt;
&lt;h3 id="q2-how-do-take-up-rates-differ-between-private-and-non-private-accounts-and-what-do-workers-say-explains-the-difference"&gt;Q2. How do take-up rates differ between Private and Non-private accounts, and what do workers say explains the difference?&lt;/h3&gt;
&lt;p&gt;In Phase 2, take-up of Private accounts is 60% versus only 14% for Non-private accounts — a 77% reduction (p&amp;lt;0.001). Among workers who declined a Non-private account, 96% cite the anticipation of increased transfer requests from network members knowing about the account as an important factor in their decision. Only 5% cite any other reason. This pattern is strong direct evidence that the fear of redistribution — not other features of the accounts — drives take-up differences.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-treatment-effects-on-earnings-and-attendance-and-how-consistent-are-they-across-phases-and-subsamples"&gt;Q3. What are the treatment effects on earnings and attendance, and how consistent are they across phases and subsamples?&lt;/h3&gt;
&lt;p&gt;Pooled across both phases, the Private arm raises daily earnings by 175.9 FCFA (11.4%, p=0.012) and attendance by 6.2 percentage points (9.7%, p=0.023). In Phase 1 alone (Private vs. Control), earnings rise 11.3% (p=0.032). In Phase 2 alone (Private vs. Non-private), earnings rise 11.5% (p=0.043). Restricting to workers not previously treated in Phase 1, the effect is 12.8% (p=0.034); restricting further to workers new to the study in Phase 2 only, the effect is 17.3% (p=0.020). The authors cannot reject that effects across these three Phase 2 subsamples are statistically the same (p=0.427), ruling out sensitivity to the cross-randomized design.&lt;/p&gt;
&lt;h3 id="q4-how-does-treatment-effect-heterogeneity-support-the-redistributive-pressure-mechanism"&gt;Q4. How does treatment effect heterogeneity support the redistributive pressure mechanism?&lt;/h3&gt;
&lt;p&gt;Workers who report difficulty saving because &amp;ldquo;someone else will need it for something urgent&amp;rdquo; see earnings increase by 15.0% (p=0.018) from the Private treatment; those not reporting this difficulty see near-zero, insignificant effects (p=0.95). Workers who make transfers to acquaintances — transfers especially unlikely to reflect altruism — see earnings rise 17.5% (p=0.014). Workers with below-median baseline earnings, potentially those facing the strongest relative disincentive to work, see larger effects. Each of these heterogeneous patterns is in the direction predicted if the social tax is the operative mechanism.&lt;/p&gt;
&lt;h3 id="q5-do-the-treatment-effects-reflect-substitution-away-from-outside-earnings-or-genuine-total-income-gains"&gt;Q5. Do the treatment effects reflect substitution away from outside earnings or genuine total income gains?&lt;/h3&gt;
&lt;p&gt;No. The paper finds no treatment effects on earnings outside the factory. At baseline, 89% of workers report zero outside earnings, and on average 93% of total income comes from factory wages. Consequently, the 11.4% earnings increase represents a near-one-for-one increase in total earned income.&lt;/p&gt;
&lt;h3 id="q6-do-private-accounts-reduce-transfers-to-the-network"&gt;Q6. Do Private accounts reduce transfers to the network?&lt;/h3&gt;
&lt;p&gt;No. The design ensures that cash-on-hand is unchanged by construction — workers receive the same or slightly higher take-home cash pay (the difference is positive but insignificant). Consistent with this, neither the probability of making transfers (p=0.37) nor transfers to family (p=0.35) or non-family (p=0.93) change significantly. Total outgoing transfers in the endline survey are if anything higher in the Private arm (p=0.049, though this may partly reflect redistribution of unblocked savings). The net transfer amount is positive but insignificant (p=0.32). The authors conclude the intervention did not make others in workers&amp;rsquo; networks worse off.&lt;/p&gt;
&lt;h3 id="q7-how-do-the-authors-rule-out-morale-or-fairness-effects-as-an-explanation"&gt;Q7. How do the authors rule out morale or fairness effects as an explanation?&lt;/h3&gt;
&lt;p&gt;Treatment assignment was conducted by lottery with ID numbers drawn in front of workers, clearly dissociating it from employer favoritism. More directly, the authors test for morale effects using the 3–4 week &amp;ldquo;announcement period&amp;rdquo; between treatment disclosure and account activation. If disgruntlement among non-Private workers drove results, output should fall during this period — but estimated announcement effects are near zero (0.8% of control mean, p=0.859 in Phase 2). In contrast, effects arise immediately in the first active paycycle: earnings jump 11.4% (p=0.082) even before workers have seen any deposits occur. The fairness story also cannot explain why effects are concentrated precisely among workers who report more redistributive pressure.&lt;/p&gt;
&lt;h3 id="q8-how-do-the-authors-test-and-rule-out-self-control-as-the-primary-mechanism"&gt;Q8. How do the authors test and rule out self-control as the primary mechanism?&lt;/h3&gt;
&lt;p&gt;Self-control cannot explain why Non-private accounts — which offer the same commitment benefit — have dramatically lower take-up than Private accounts. Separately, the authors test a core prediction of time inconsistency models by surprising workers with an option to opt out of the next deposit, randomly varying whether the offer comes 4 days before payday or on payday itself. Under quasi-hyperbolic preferences, workers should be more likely to opt out on the payday itself. Counter to this prediction, 94% of workers keep their earnings in the account on payday, compared to 86% four days before — and these means are not statistically distinguishable, with the relative magnitudes actually running opposite to time inconsistency predictions.&lt;/p&gt;
&lt;h3 id="q9-how-do-the-authors-address-the-concern-that-non-private-accounts-may-raise-the-tax-rate-above-the-baseline-inflating-treatment-effect-estimates"&gt;Q9. How do the authors address the concern that Non-private accounts may raise the tax rate above the baseline, inflating treatment effect estimates?&lt;/h3&gt;
&lt;p&gt;The concern is that Non-private SMS alerts could make network members more aware of available cash than under the status quo, pushing the effective comparison above the Control level. The authors note that (a) paydays are already publicly known in this setting and workers regularly face transfer requests around them; (b) workers must physically withdraw savings from a bank after the unblock date, and can even re-block funds; and (c) the magnitude of effects when comparing Private to Control is nearly identical to the effect when comparing Private to Non-private (11.3% vs. 11.5%), suggesting the Non-private condition does not materially raise the tax above the status quo.&lt;/p&gt;
&lt;h3 id="q10-how-do-the-authors-rule-out-privacy-concerns-rather-than-redistributive-pressure-as-the-driver-of-low-non-private-take-up-and-treatment-effects"&gt;Q10. How do the authors rule out privacy concerns (rather than redistributive pressure) as the driver of low Non-private take-up and treatment effects?&lt;/h3&gt;
&lt;p&gt;Four arguments are provided. First, Phase 1 effects (Private vs. Control, no Non-private arm) are the same magnitude as Phase 2 effects, yet Phase 1 cannot be confounded by privacy concerns. Second, among workers who refused Non-private accounts, 96% cite transfer request anticipation; none volunteer generic privacy concerns. Third, heterogeneity effects — concentrated among high-redistributive-pressure workers — have no obvious connection to privacy preferences. Fourth, two placebo SMS exercises: 95% of Non-private workers grant permission to send generic bank promotional texts, and 88% of workers who had Phase 1 Private accounts grant permission for messages about their past (already-spent) savings — indicating no inherent aversion to having some financial information shared with networks. Since these workers forgo 11.5% of full-time earnings by refusing Non-private accounts, privacy concerns alone are implausible as a full explanation.&lt;/p&gt;
&lt;h3 id="q11-how-is-the-social-tax-rate-estimated-and-what-does-the-range-look-like"&gt;Q11. How is the social tax rate estimated and what does the range look like?&lt;/h3&gt;
&lt;p&gt;The authors combine the 11.4% ITT treatment effect (used as the ratio e₁/e₂) with a compensated labor supply elasticity ζ estimated from an end-of-experiment piece-rate randomization. The piece-rate experiment (varying piece rates over four values from −15% to +30% of baseline over 6 days) yields an intensive-margin elasticity of 0.17. Using the ratio of attendance to intensive-margin effects from Table 3, the implied extensive-margin elasticity is 0.94, giving ζ ≈ 1.11. With this elasticity and assuming τ₂ = 0 (most conservative), the ITT-implied social tax rate is 9%; assuming τ₂ = 5%, it is 14%. For compliers (workers who actually take up Private accounts), the estimated rate is 19–23%. If instead the lower elasticity estimate of 0.32 (comparable to Goldberg 2016) is used, the ITT tax rate would be at least 29%.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-broader-implications-discussed-by-the-authors"&gt;Q12. What are the broader implications discussed by the authors?&lt;/h3&gt;
&lt;p&gt;The authors propose that if redistributive pressure distorts work incentives, it may also distort other costly income-generating actions: technology adoption, human capital investment, and formal sector participation. They note that 74% of workers believe taking a formal job would increase transfer requests, even though network members could also access such jobs. A speculative but highlighted policy implication is that formal safety nets (health or unemployment insurance) could reduce social tax burdens on non-recipients by absorbing demand for redistribution, potentially generating positive productivity externalities.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Social Tax&lt;/strong&gt;: The paper&amp;rsquo;s central concept. Redistributive pressure from kin and social networks is modeled as a tax rate τ₁ on gross earnings — not altruistic transfers, but transfers made under social pressure that workers would prefer to avoid. The &amp;ldquo;tax&amp;rdquo; analogy captures that the obligation is proportional to visible income and reduces the private return to earning more. The paper explicitly does not take a stance on the underlying microfoundation (risk-sharing, cultural norms, or a mix).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Blocked Savings Account&lt;/strong&gt;: A date-based savings account (implemented with Banque Populaire de Côte d&amp;rsquo;Ivoire) into which any earnings above a worker-chosen threshold are automatically deposited by the factory. Funds are inaccessible until the blocked period ends (3–9 months). Workers cannot withdraw during the period, making deposited earnings unavailable to fulfill transfer requests and therefore effectively reducing the social tax rate on earnings increases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Private vs. Non-private Treatment&lt;/strong&gt;: The paper&amp;rsquo;s key experimental contrast. A Private account&amp;rsquo;s existence is unknown to anyone in the worker&amp;rsquo;s network. A Non-private account triggers SMS messages to network members disclosing that the worker is saving and announcing when the unblock date approaches. The contrast isolates whether the shielding of income from social visibility — not the commitment device per se — drives take-up and labor supply.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Substitution Effect without Income Effect&lt;/strong&gt;: The paper&amp;rsquo;s design deliberately places the tax reduction only on earnings &lt;em&gt;above&lt;/em&gt; baseline, creating a kink in the budget constraint. Starting from the existing labor supply level, there is no change in net earnings at the margin — eliminating the income effect of a tax reduction — so any labor supply response is a pure compensated (substitution) effect. This makes any observed increase in labor supply an unambiguous signal that a distortionary social tax exists.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intent to Treat (ITT) vs. Treatment on the Treated (ToT)&lt;/strong&gt;: The ITT estimate (11.4% earnings increase) reflects the effect of being &lt;em&gt;offered&lt;/em&gt; a Private account on all offered workers, including those who did not take up. The ToT estimate — relevant for workers who actually used the accounts — implies a higher social tax rate (19–23%) because only roughly half of offered workers take up the accounts and only those workers face a materially reduced effective tax rate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Compensated (Hicksian) Labor Supply Elasticity (ζ)&lt;/strong&gt;: The ratio used to infer the social tax rate from the observed treatment effect. The paper estimates ζ ≈ 1.11 (extensive margin ζₐ ≈ 0.94, intensive margin ζₑ ≈ 0.17) from an end-of-experiment piece-rate randomization. The social tax rate is recovered as τ₁ = 1 − (1−τ₂)(e₁/e₂)^(1/ζ) from Equation (5).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Piece Rate Setting&lt;/strong&gt;: Workers earn a linear piece rate for every kilogram of cashews peeled, with no fixed pay component. This setting ensures that every unit of additional effort by a worker translates directly into higher earnings, and that any observed earnings changes cleanly reflect labor supply responses rather than hour or schedule effects.&lt;/p&gt;</description></item><item><title>Vanguard: Black Veterans and Civil Rights After World War I</title><link>https://macropaperwarehouse.com/papers/vanguard-black-veterans-and-civil-rights-after-world-war-i/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/vanguard-black-veterans-and-civil-rights-after-world-war-i/</guid><description>&lt;p&gt;This paper provides the first causal evidence on how military service shaped Black civil rights activism in the aftermath of World War I. The research question is whether random induction into the segregated National Army caused Black men to join the nascent NAACP and become prominent community leaders during the New Negro era. The authors leverage the WWI draft lottery — in which each registrant&amp;rsquo;s unique serial number was drawn from a bowl to determine induction order — as an instrument for military service, a source of exogenous variation not previously exploited in the literature.&lt;/p&gt;
&lt;p&gt;To support this analysis, Ang and Chinoy construct an unusually rich dataset by digitizing nearly one million Black draft registration cards from the first registration (June 17, 1917), linking them through the 1930 full-count census to 233,517 NAACP member observations across 227 branches from 1912 to 1940, and supplementing with Veterans Administration records, Army Transport Service passenger lists, and biographical dictionaries of prominent African Americans. The instrument — serial number percentile within draft board and race (SNP%) — is validated against all observed pre-draft registrant characteristics and yields a first-stage F-statistic of 1,051 in the preferred specification.&lt;/p&gt;
&lt;p&gt;The main finding is that Black men randomly induced to serve in the military were nearly three times more likely to join the NAACP than observably similar registrants from the same draft board (TSLS coefficient 0.0219, se = 0.0049, against a sample mean NAACP participation rate of 0.8%). The authors estimate that the draft induced more than 10,000 Black men to join the NAACP in total. Military service also raised the probability of appearing in biographical dictionaries of historically prominent African Americans by a factor of roughly 1.6 (TSLS coefficient 0.0027, se = 0.0012, sample mean 0.17%). These results are robust to alternative instruments, flexible polynomial specifications of SNP%, state-year fixed effects, and alternative veteran-status measures from VAMI and ATS records. They are also not explained by differential residential mobility: adding controls for interstate and North-South migration leaves the main coefficient essentially unchanged (0.0217-0.0218).&lt;/p&gt;
&lt;p&gt;In contrast, TSLS estimates for all socioeconomic outcomes — literacy, home ownership, employment, census-predicted income, actual 1940 income, and educational attainment — are small and insignificant, ruling out human capital acquisition as a mechanism. Club involvement measured in the census is likewise unaffected, indicating that NAACP membership reflects specifically civil rights activism rather than generically greater social participation.&lt;/p&gt;
&lt;p&gt;The mechanism the paper identifies is experienced discrimination. Effects on NAACP participation increase monotonically with the racial gap in induction rates across draft boards (significant at p = 0.01). Effects are large and significant for men assigned to camps that restricted Black soldiers&amp;rsquo; access to military training (coefficient 0.0351, se = 0.0104) and to officer promotion (coefficient 0.0360, se = 0.0111), and are large for men in both restriction types simultaneously (coefficient 0.0367, se = 0.0114). In contrast, men attending less discriminatory camps show small and insignificant effects. Among the two all-Black combat divisions, NAACP participation is highest for veterans of the 92nd Division — subjected to constant racial abuse under U.S. command — and lower for the 93rd Division, which served under more hospitable French command. Previously unstudied veteran surveys from Virginia and Connecticut corroborate this narrative: respondents from camps with training and promotion restrictions were more than twice as likely to mention racial injustice, and mentions of injustice were more predictive of postwar civic engagement than any other survey theme.&lt;/p&gt;
&lt;p&gt;The scope of the paper is Black male registrants in the first WWI draft registration (men aged 21-30 as of June 17, 1917), linked to a sample of approximately 300,000 in the 1930 census. Effects are attenuated for men from counties with greater racial hostility — proxied by Confederate state status, Confederate monument density, and county lynching rates — consistent with the interpretation that activism was more feasible in less repressive environments.&lt;/p&gt;
&lt;p&gt;Q: What is the core identification strategy and why was it not feasible to use it before this paper?
A: The paper uses each Black registrant&amp;rsquo;s serial number percentile within his draft board and racial group (SNP%) as an instrument for WWI military service. Unlike the WWII and Vietnam drafts, which used birthday-based lotteries, the WWI lottery assigned induction order by drawing unique serial numbers from a bowl, making serial number rank the source of quasi-random variation. This source had never been exploited in the literature, partly because the serial numbers had to be hand-captured from digitized draft card images.&lt;/p&gt;
&lt;p&gt;Q: How strong is the first stage, and was the lottery truly random?
A: The first-stage F-statistic is 1,051, and a ten-percentile decrease in SNP% is associated with a 34.5 percentage point increase in the probability of serving. Bivariate serial numbers show some non-random patterns — nine of 13 pre-draft characteristics correlate with raw SN% — likely because some Southern boards inflated numbers for white registrants. Conditioning on board fixed effects and using SNP% within board-race cells eliminates these correlations; Panel B of Appendix Table A1 shows the largest standardized coefficient falls to 0.006.&lt;/p&gt;
&lt;p&gt;Q: What is the magnitude of the effect on NAACP membership and how does the causal estimate compare to a naive OLS?
A: The TSLS coefficient is 0.0219 (se = 0.0049) against a sample mean of 0.8%, implying roughly a threefold increase in NAACP membership. The OLS estimate of 0.0116 understates the causal effect, consistent with the marginal man induced by the lottery being observationally weaker than infra-marginal volunteers.&lt;/p&gt;
&lt;p&gt;Q: Does the effect reflect simply that veterans moved to Northern cities where NAACP branches were more accessible?
A: No. Adding indicators for interstate migration and North-South migration leaves the TSLS coefficient essentially unchanged at 0.0218 and 0.0217, respectively. The Great Migration channel is thus not the operative mechanism.&lt;/p&gt;
&lt;p&gt;Q: Did military service improve Black veterans&amp;rsquo; economic outcomes?
A: TSLS estimates for literacy, home ownership, employment, census-predicted income, actual 1940 income, and educational attainment are all small and statistically insignificant. This contrasts sharply with evidence on Black veterans of WWII and Korea (Greenberg et al., 2022) and is consistent with the documented absence of meaningful postwar benefits or training for Black WWI soldiers.&lt;/p&gt;
&lt;p&gt;Q: If it was not human capital or migration, what mechanism does the paper establish?
A: The primary mechanism is exposure to institutional discrimination during military service. Three distinct empirical patterns converge: (1) effects increase monotonically with draft board racial disparities in induction rates; (2) effects are large and significant for men at camps that denied training and promotion, and near zero for men at less discriminatory camps; (3) veteran survey mentions of racial injustice are more common among men from discriminatory camps and are more predictive of postwar NAACP membership than any other survey theme.&lt;/p&gt;
&lt;p&gt;Q: How do the two all-Black combat divisions differ in their postwar NAACP participation, and what does this reveal?
A: Veterans of the 92nd Division, who fought under U.S. command amid constant racial abuse, show the highest NAACP participation rates. Veterans of the 93rd Division, who fought under French command and were received with relative hospitality, show lower (though not statistically significantly lower) participation. Since both divisions received similar formal training and neither group shows socioeconomic gains, the differential reflects discrimination exposure rather than skill acquisition.&lt;/p&gt;
&lt;p&gt;Q: What is the quantitative scale of the effect for the most discriminatory camps?
A: For men assigned to camps with restrictions on both training and promotion, the TSLS coefficient on NAACP membership is 0.0367 (se = 0.0114) — more than 1.5 times the average estimate of 0.0219. Men at camps without restrictions show coefficients that are small and statistically insignificant.&lt;/p&gt;
&lt;p&gt;Q: How does county-level racial hostility moderate the effect?
A: The effects of military service on NAACP membership are larger — more positive — for men from counties with fewer Confederate monuments, lower lynching rates, and non-Confederate state status. This is interpreted as evidence that activism in response to discriminatory military experiences was more feasible in less racially hostile local environments, rather than as evidence that discrimination exposure was lower.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s aggregate policy implication regarding the scale of the draft&amp;rsquo;s effect on the civil rights movement?
A: The authors estimate that the WWI draft induced more than 10,000 Black men to join the NAACP. Veterans accounted for nearly 15% of all male NAACP members, against roughly 8% of Black male adults in the population, and were significantly more likely to appear in biographical dictionaries of prominent African Americans. The draft thus constituted a sizable and measurable contribution to the organizational vanguard of the early civil rights movement.&lt;/p&gt;
&lt;p&gt;Q: How does the paper contribute to the economics of discrimination beyond documenting discriminatory behavior by majority actors?
A: Most economics research on discrimination studies the conduct of white decision-makers (e.g., racial bias in hiring, lending, or bail). This paper examines how experiences of discrimination reshape the political behavior and aspirations of the minority group itself. The results show that institutional betrayal — systematic exclusion, degradation, and denial of training — generated deep discontent that translated into aggressive political mobilization, a dynamic the authors trace through subsequent episodes including the WWII Double V campaign and responses to police killings.&lt;/p&gt;
&lt;p&gt;Serial number percentile within draft board and race (SNP%): The instrument constructed by the authors. Each WWI registrant received a serial number from 1 to the size of his draft board; those numbers were drawn in random order to determine induction priority. SNP% measures where a registrant fell in that draw relative to others in his board and racial group, and serves as the source of quasi-random variation in veteran status.&lt;/p&gt;
&lt;p&gt;New Negro era: The period of invigorated Black political and cultural assertiveness following WWI, characterized by renewed racial pride, economic independence, and progressive politics. The movement spanned the Harlem Renaissance, the Universal Negro Improvement Association, the American Negro Press, and the Brotherhood of Sleeping Car Porters, and represented a rejection of the &amp;ldquo;conservatism, parochialism, and political accommodationism&amp;rdquo; of older Black leaders.&lt;/p&gt;
&lt;p&gt;Draft board racial gap: The authors&amp;rsquo; measure of draft board discrimination, defined as the difference in induction rates between Black and white registrants within a given draft board. The interquartile range spans roughly 0 to 20 percentage points, with a notable fraction of boards exhibiting gaps exceeding 30 percentage points.&lt;/p&gt;
&lt;p&gt;Camp discrimination: The denial of military training and officer promotion opportunities to Black soldiers, documented in War Department reports by military intelligence officers tasked with monitoring the treatment of Black soldiers. The paper classifies each camp as restricted or unrestricted on each dimension and uses this classification to estimate heterogeneous treatment effects.&lt;/p&gt;
&lt;p&gt;Institutional betrayal: The paper&amp;rsquo;s characterization of the U.S. government&amp;rsquo;s treatment of Black WWI soldiers — drafting them at higher rates than whites, denying them training and promotion, and assigning them to menial labor — as generating a profound sense of injustice that motivated postwar political activism rather than loyalty or accommodation.&lt;/p&gt;
&lt;p&gt;NAACP membership as civil rights activism proxy: The paper uses dues-paying membership in local NAACP branches as its primary quantitative measure of civil rights participation. Membership involved active financial cost (annual fees of $1 to $10 at a time when median Black family income was below $500), exposure to harassment and violence in the South, and participation in local protest and legal advocacy, distinguishing it from passive civic engagement.&lt;/p&gt;</description></item><item><title>What Do Policies Value?</title><link>https://macropaperwarehouse.com/papers/what-do-policies-value/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/what-do-policies-value/</guid><description>&lt;p&gt;This paper asks a fundamental question about policy design: when a program prioritizes one group over another, is that because the group benefits more from the intervention, or because the policy assigns them higher intrinsic welfare weight? Björkegren, Blumenstock, and Knight develop a two-stage method to decompose observed allocation decisions into their underlying components: (i) welfare weights assigned to different types of people, (ii) heterogeneous treatment effects of the intervention, and (iii) relative weights on different outcomes. The key insight is that the same allocation rule can be consistent with very different value systems depending on how much each group actually benefits.&lt;/p&gt;
&lt;p&gt;The method works as follows. In a first stage, the analyst estimates heterogeneous treatment effects — how much each individual benefits on each outcome dimension — using OLS or machine learning methods (e.g., causal forests). In a second stage, the analyst reconciles the observed ranking of beneficiaries with an implicit welfare function using an exploded logit likelihood, recovering welfare weights (who is valued), impact weights (how different outcomes are valued), and a base value for treatment independent of measured outcomes. Identification requires an exclusion restriction: the covariates used to estimate treatment effect heterogeneity must include variables excluded from the welfare weight specification, allowing the analyst to compare households with similar welfare weights but differential treatment effects. Variants of the method that impose known welfare weights or known impact weights can be used without the exclusion restriction.&lt;/p&gt;
&lt;p&gt;The paper demonstrates the method using PROGRESA, Mexico&amp;rsquo;s large conditional cash transfer program launched in 1997. PROGRESA ranked households by a proxy means test poverty score and transferred approximately 197 pesos per month (roughly $20 USD) to eligible poor households, conditional on school attendance and doctor visits. The analysis uses endline survey data on 7,767 households and focuses on three outcomes emphasized in program documents: log per-capita consumption, child sick days (ages 0-5), and school days missed (ages 6-16).&lt;/p&gt;
&lt;p&gt;The program&amp;rsquo;s average treatment effects were: a 0.149 log point increase in monthly consumption (SE=0.015), a 0.165 reduction in sick days per child (SE=0.051), and a near-zero effect on school days missed (-0.0053, SE=0.028). These effects were heterogeneous: indigenous households, for instance, benefited substantially more from the program.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central empirical finding inverts the naive interpretation of PROGRESA&amp;rsquo;s targeting. Indigenous households were ranked 60.6 log points higher in the program&amp;rsquo;s priority order. A simple regression suggests the program favored them. But after accounting for the fact that indigenous households benefit substantially more from treatment, the method finds that the program&amp;rsquo;s implied welfare weight on indigenous households is, if anything, lower by 17.4% relative to non-indigenous households — not higher. The program&amp;rsquo;s prioritization of indigenous households is thus explained by efficiency, not by preferential welfare weighting.&lt;/p&gt;
&lt;p&gt;Because PROGRESA cash transfers relax household budget constraints and outcomes like consumption reflect household choices, the impact weights capture the difference between how the policy values outcomes and how households value them. The estimates strongly reject non-paternalism: the policy implicitly values consumption and potentially health differently from household decision-makers. Of the total welfare impact, approximately 55% is attributed to the base value of the transfer itself (independent of measured outcomes), approximately 45% to consumption impacts, and less than 1% to health and schooling impacts combined. The implied value of providing the transfer independent of outcomes corresponds to 0.16 log points of consumption, or about 23.1 pesos per person per month — slightly below the average transfer of 33.9 pesos per person per month.&lt;/p&gt;
&lt;p&gt;The paper also runs counterfactual exercises showing how alternative preference structures would have changed the allocation. A policy maximizing only educational impacts would have prioritized richer, smaller households; one maximizing only consumption impacts would have further prioritized indigenous households. These counterfactuals are mapped onto a Pareto frontier across the three outcomes. The estimated welfare weights from the implemented policy align closely with preferences elicited in a 2023 survey of 429 Mexican residents, though residents placed higher value on child health relative to what the policy implied.&lt;/p&gt;
&lt;p&gt;Q: What is the core identification challenge the paper addresses?
A: When a policy prioritizes a group, it could be because the group benefits more (efficiency) or because the policy assigns them intrinsically higher value (preference). These two explanations are observationally equivalent from the allocation alone. The paper separates them by first estimating heterogeneous treatment effects and then inverting the allocation to recover residual welfare weights.&lt;/p&gt;
&lt;p&gt;Q: What is the exclusion restriction required for full identification?
A: The covariates used to estimate treatment effect heterogeneity (x-tilde) must include at least some variables excluded from the welfare weight specification (x). This allows the analyst to compare households with similar welfare weights but different predicted treatment effects, pinning down how much of the ranking reflects efficiency versus preference. Without this restriction, one can still recover conditional preferences by imposing known values for either welfare weights or impact weights.&lt;/p&gt;
&lt;p&gt;Q: How does the exploded logit likelihood work in this setting?
A: The analyst observes a single full ranking of all households, rather than partial orderings from multiple decision-makers. The welfare impact of treating household i is modeled as a linear function of predicted treatment effects scaled by welfare and impact weights, plus an extreme-value-distributed shock. The likelihood of observing household i ranked above household i-prime is the ratio of their exponentiated welfare scores, summed over all households ranked below i. Maximum likelihood recovers the welfare weights, impact weights, and base value simultaneously.&lt;/p&gt;
&lt;p&gt;Q: What were PROGRESA&amp;rsquo;s average treatment effects on the three focal outcomes?
A: Average treatment increased log monthly consumption by 0.149 (SE=0.015), reduced child sick days by 0.165 (SE=0.051), and had a near-zero effect on school days missed (-0.0053, SE=0.028). The consumption and health effects are statistically significant; the schooling effect is not distinguishable from zero.&lt;/p&gt;
&lt;p&gt;Q: What does the analysis find about the welfare weight assigned to indigenous households?
A: In the raw ranking regression, indigenous households are ranked 60.6 log points higher, suggesting the program favored them. After accounting for the fact that indigenous households benefit substantially more from treatment, the method finds the implied welfare weight on indigenous households is lower, not higher — specifically, about 17.4% lower than non-indigenous households. The program&amp;rsquo;s higher ranking of indigenous households is explained entirely by their larger treatment effects, not by preferential weighting.&lt;/p&gt;
&lt;p&gt;Q: How are the impact weights on consumption, health, and schooling interpreted given that outcomes reflect household choices?
A: Because PROGRESA relaxes household budget constraints and outcomes like consumption result from household optimization, the estimated impact weights capture the difference between how the policy values outcomes relative to how households value them (internalities), rather than the absolute policy valuation. A nonzero weight implies the policy disagrees with household preferences — paternalism. The positive coefficient on log consumption implies the policy values this outcome more than households do.&lt;/p&gt;
&lt;p&gt;Q: How much of PROGRESA&amp;rsquo;s welfare impact comes from the base transfer value versus measured outcomes?
A: The base value of the transfer (independent of measured impacts on consumption, health, and schooling) accounts for approximately 55% of total implied welfare impact. The impact on consumption accounts for approximately 45%. Impacts on health and schooling together account for less than 1%. The implied value of the base transfer corresponds to 0.16 log points of consumption per capita, or about 23.1 pesos per person per month — somewhat below the average transfer amount of 33.9 pesos per person per month.&lt;/p&gt;
&lt;p&gt;Q: Does the analysis reject egalitarian welfare weights and non-paternalism?
A: Yes, using Wald tests with bootstrapped covariance matrices. The hypothesis of egalitarian weights (all gamma equal to one) is rejected. Non-paternalism (all beta equal to zero) is strongly rejected. The joint hypothesis of egalitarianism and non-paternalism is also rejected across all specifications tested.&lt;/p&gt;
&lt;p&gt;Q: How do the estimated welfare weights compare to stated preferences of Mexican residents?
A: The 2023 survey of 429 Mexican residents elicited preferences using multiple price lists over how to prioritize different household types. The welfare weights implied by the implemented policy are broadly similar to resident preferences, but the policy places relatively higher welfare weight on indigenous households than the median survey respondent does. Survey respondents value child health impacts more than household decision-makers and more than the implemented policy does, consistent with support for paternalism.&lt;/p&gt;
&lt;p&gt;Q: What do counterfactual allocations reveal about the relationship between policy goals and targeting priorities?
A: A policy maximizing only consumption impacts would further prioritize indigenous households with lower income. A policy maximizing only educational impacts would instead prioritize richer, smaller households. A policy maximizing only health impacts would largely preserve indigenous household prioritization while placing less emphasis on lower-education households. These three extreme policies map to the corners of a Pareto frontier, and the implemented PROGRESA policy lies close to the allocation consistent with surveyed resident preferences.&lt;/p&gt;
&lt;p&gt;Q: What changed when Mexico reformed PROGRESA&amp;rsquo;s poverty score in 2003?
A: The 2003 reform increased the priority of older and smaller households. Applying the method to the new poverty score reveals that it implicitly switched to assigning a positive welfare weight to indigenous households (compared to the negative implied weight under the original score), and placed less welfare weight on lower-income and younger households relative to the original design.&lt;/p&gt;
&lt;p&gt;Q: What are the main limitations and scope conditions of the method?
A: Full identification requires an exclusion restriction (some treatment effect heterogeneity predictors excluded from welfare weights) and sufficient variation in treatment effects across household types. If treatment effects are homogeneous, welfare weights and impact weights cannot be separately identified. If correlated unobservables drive the ranking but are not modeled, the method recovers preferences consistent with included variables only, analogous to omitted variable bias in OLS. The method also requires a way to estimate treatment effect heterogeneity, which is most credible with a randomized pilot, though non-experimental methods are in principle applicable.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to the inverse optimum public finance literature?
A: The inverse optimum literature (Bourguignon and Spadaro 2012; Saez and Stantcheva 2016; Hendren 2020) recovers the redistribution preferences consistent with income tax schedules, conditioning on a single covariate (pre-tax income) affecting a single outcome (net-of-tax consumption). This paper generalizes that framework to arbitrary allocation policies conditioning on a vector of covariates and affecting a vector of outcomes, and extends it to settings beyond income taxation where heterogeneous treatment effects can be estimated.&lt;/p&gt;
&lt;p&gt;Q: Can the method be applied when only a binary allocation is observed rather than a full ranking?
A: Yes. A binary allocation corresponds to a ranking with only two levels, and the same exploded logit procedure applies, though with reduced statistical power. The paper provides an empirical illustration of this setting in Section 5.2.1.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Welfare weights (w(x_i)):&lt;/strong&gt; The policy&amp;rsquo;s differential valuation of one household&amp;rsquo;s utility relative to another, expressed as a multiplicative function of household characteristics. Distinct from how much a household benefits — two households may be ranked identically despite different benefits if their welfare weights differ proportionally.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Impact weights (beta_j):&lt;/strong&gt; The policy&amp;rsquo;s relative valuation of different outcome components (consumption, health, schooling). For outcomes that are household choices, impact weights capture the difference between how the policy values the outcome and how the household values it — an internality or paternalistic preference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Base value (alpha):&lt;/strong&gt; The value a policy assigns to providing a treatment independent of its measured impact on any specific outcome. Captures either a direct utility benefit of treatment or the value of relaxing household budget constraints when outcomes are choices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exclusion restriction:&lt;/strong&gt; The requirement that the set of covariates used to estimate treatment effect heterogeneity includes at least some variables excluded from the welfare weight specification. Enables separate identification of efficiency-based and preference-based components of a ranking by comparing households similar in welfare weight but different in predicted treatment effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Exploded logit likelihood:&lt;/strong&gt; The econometric procedure used in the second stage, adapted for a single complete ranking of all alternatives rather than partial orderings. Treats the observed ranking of household i as a choice from the set of all households ranked below it, with likelihood given by the softmax of welfare scores.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Value audit:&lt;/strong&gt; A retrospective application of the method that reads the implicit values encoded in an implemented policy&amp;rsquo;s allocation decisions, enabling comparison against stated policy objectives, constituent preferences, or normative benchmarks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Paternalism (in this paper&amp;rsquo;s sense):&lt;/strong&gt; A policy is paternalistic if it assigns nonzero impact weight (beta_j ≠ 0) to outcomes that are household choices — meaning the policy values those outcomes differently from the households making the choices. The envelope theorem implies a non-paternalistic policy would place zero weight on choice outcomes beyond the general constraint relaxation.&lt;/p&gt;</description></item><item><title>What Jobs Come to Mind? Stereotypes About Fields of Study</title><link>https://macropaperwarehouse.com/papers/what-jobs-come-to-mind-stereotypes-about-fields-of-study/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/what-jobs-come-to-mind-stereotypes-about-fields-of-study/</guid><description>&lt;p&gt;Conlon and Patel test whether students stereotype the link between college majors and occupations — that is, whether they exaggerate the likelihood that majors lead to their &amp;ldquo;representative&amp;rdquo; careers (those most overrepresented among a major&amp;rsquo;s graduates relative to other majors, as measured by a likelihood ratio in US census data). The representative career for each major is intuitive: doctors for biology/chemistry, lawyers for political science, counselors for psychology, journalists for communications, artists for art, and so forth.&lt;/p&gt;
&lt;p&gt;The authors draw on three bodies of evidence. First, surveys of first-year undecided undergraduates in Ohio State University&amp;rsquo;s Exploration program (primarily Fall 2020 and Fall 2021 cohorts, ~80% response rate), asking students their beliefs about the share of US graduates in various careers conditional on major, as well as their beliefs about their own likely career. Beliefs are benchmarked against true career shares computed from the 2017–2019 American Community Survey restricted to college graduates aged 30–50. Second, 40+ years (1975–2018) of the CIRP Freshman Survey from UCLA, covering more than nine million nationally representative US college freshmen, which records intended major and intended career. Third, a field experiment embedded in the 2021 OSU survey with an RD design, in which treated students were shown the true share of their top major&amp;rsquo;s representative career before reporting beliefs, intentions, and — via administrative records — actual course enrollments and major declarations up to three years later.&lt;/p&gt;
&lt;p&gt;The main finding is large, systematic overestimation of representative careers. In the OSU survey, students believe 53% of art majors work as artists (true: 17%), 47% of journalism majors work as journalists (true: 4%), 38% of political science majors work as lawyers (true: 16%), and 43% of psychology majors work as counselors (true: 21%). OLS regressions of beliefs on true career frequency and a representative-career indicator yield a stereotyping coefficient θ of 0.32 p.p. (p &amp;lt; 0.01) without career fixed effects and 0.28 p.p. (p &amp;lt; 0.01) with them, meaning students believe representative careers are roughly 28–32 percentage points more common than equally prevalent non-representative careers. These patterns are similar across gender, ethnicity, and first-generation status, replicate in an MTurk sample (θ = 0.30, p &amp;lt; 0.01) and a nationally representative US adult sample (θ = 0.33, p &amp;lt; 0.01).&lt;/p&gt;
&lt;p&gt;In the CIRP data, 63% of biology freshmen expect to become doctors (true: 23%), 62% of psychology freshmen expect to be counselors (true: 21%), 65% of art freshmen expect to be artists (true: 17%), and 42% of communications/journalism freshmen expect to be writers or journalists (true: 4%). The average gap between expected and actual representative-career attainment is 36 p.p., and this gap has been roughly stable since at least the 1970s.&lt;/p&gt;
&lt;p&gt;An implicit association test (IAT) administered to 434 OSU students shows that implicit associations between representative major–career pairs are 0.30–0.36 standard deviations stronger than for non-representative pairs (p &amp;lt; 0.01), and remain 0.24–0.28 SDs stronger (p &amp;lt; 0.01) after controlling for true career frequency. A one-SD increase in individual IAT scores predicts 2.8–4.1 p.p. greater stereotyped beliefs (p &amp;lt; 0.01). Knowing someone with a non-representative major–career combination predicts beliefs 16 p.p. lower for the representative career (p &amp;lt; 0.01) — more than half the stereotyping effect — and also predicts lower IAT scores, suggesting associations arise from personal experience.&lt;/p&gt;
&lt;p&gt;An equilibrium model shows that stereotyping causes students to infer that representative careers have unusually favorable unobservable attributes, and that this inflates enrollment in the representative major among marginal students who are poorly suited to it. Correlational evidence from the NSCG, SIPP, and SHED confirms that majors subject to greater stereotyping are associated with more job dissatisfaction (+6.0% per SD, p &amp;lt; 0.01), greater job-skill mismatch (+3.1%, p &amp;lt; 0.05), more major-career mismatch (+5.4%, p &amp;lt; 0.05), and more regret about field of study (+4.8%, p &amp;lt; 0.05).&lt;/p&gt;
&lt;p&gt;The field experiment shows that correcting beliefs reduces stereotyping and shifts major choices. A 10 p.p. reduction in beliefs about the top major&amp;rsquo;s representative career lowers intentions toward that major by 3.5 p.p. (p &amp;lt; 0.01), reduces enrollment in that major&amp;rsquo;s courses by 0.22 credits in the next semester (p &amp;lt; 0.05), and reduces the probability of declaring that major within one year by 6.1 p.p. (p = 0.23). The same information boosts intentions toward students&amp;rsquo; second-ranked major by 2.1 p.p. (p = 0.17), increases second-major course enrollment by 0.20 credits (p &amp;lt; 0.10), and raises the probability of declaring the second major within a year by 9.9 p.p. (p &amp;lt; 0.01). Treated students also spend on average 0.21 more semesters undecided before declaring a major (p &amp;lt; 0.05). Effects are concentrated in the first year and partially fade over the two-to-three-year follow-up window.&lt;/p&gt;
&lt;p&gt;Q: How do the authors define a major&amp;rsquo;s &amp;ldquo;representative career&amp;rdquo;?
A: The representative career of major M is the career c that maximizes the likelihood ratio R(c, M) = p_{c|M} / p_{c|not-M}, where p_{c|M} is the true share of major-M graduates working in career c and p_{c|not-M} is the share of graduates from all other majors working in c. This ratio captures how much more common a career is among one major&amp;rsquo;s graduates relative to all other graduates. For example, the representative career of communications/journalism is &amp;ldquo;writers and journalists,&amp;rdquo; whose graduates are between 155% and 1,751% more likely to hold their major&amp;rsquo;s representative career than graduates of other majors, even though the absolute frequency of such careers is often modest (ranging from 2% to 60% across fields).&lt;/p&gt;
&lt;p&gt;Q: What is the core model of stereotyped belief formation?
A: The model draws from Bordalo et al. (2016). Let p_{c|M} be the true career share and π_{c|M} the student&amp;rsquo;s belief. The model specifies π_{c|M} = (1 − θ) p_{c|M} + θ · 1[c = c*(M)], where c*(M) is the representative career and θ ∈ [0,1] measures the extent of stereotyping. When θ = 0 the student holds rational beliefs; when θ = 1 beliefs assign all probability mass to the representative career. This formulation implies that students overweight representative careers because those careers come to mind more easily, grounded in a representativeness heuristic based on likelihood ratios.&lt;/p&gt;
&lt;p&gt;Q: What does the regression test for stereotyping find in the OSU survey?
A: The authors regress individual beliefs π_{c|M} on the true frequency p_{c|M} and an indicator for c being the representative career of M, clustering standard errors at the individual and career-by-major level. The estimated θ is 0.32 (p &amp;lt; 0.01) without career fixed effects (Column 1 of Table 1) and 0.28 (p &amp;lt; 0.01) with career fixed effects (Column 2). For self-beliefs about students&amp;rsquo; top-ranked major, the estimates are 0.36–0.43 p.p. (p &amp;lt; 0.01 both with and without career fixed effects). These estimates imply that students regard a major&amp;rsquo;s representative career as 28–43 percentage points more common than an equally prevalent non-representative career for the same major.&lt;/p&gt;
&lt;p&gt;Q: Do the OSU results replicate in other samples?
A: Yes. An MTurk convenience sample of 430 current college students yields a stereotyping coefficient of 0.30 (p &amp;lt; 0.01). A nationally representative sample of US adults yields a coefficient of 0.33 (p &amp;lt; 0.01); this pattern holds separately for college-educated and non-college-educated respondents and for both younger respondents (aged 18–29) and older respondents (aged 30+). The authors also ran a pre-registered 2021 replication survey in a new OSU Exploration cohort and found similar results.&lt;/p&gt;
&lt;p&gt;Q: What does the CIRP Freshman Survey data show about the persistence and scale of stereotyping?
A: Pooling more than nine million US college freshmen surveyed from 1975 to 2018, the CIRP data show that students systematically intend to enter their major&amp;rsquo;s representative career far more often than graduates actually do. Among students who have decided on a major, 63% intend to have their major&amp;rsquo;s representative career while only 27% of college graduates actually attain it — a gap of 36 p.p. (p &amp;lt; 0.01). The specific examples include: 63% of biology freshmen intend to become doctors (true: 23%), 62% of psychology freshmen expect to be counselors (true: 21%), 65% of art freshmen expect to be artists (true: 17%), and 42% of communications/journalism freshmen expect to be writers or journalists (true: 4%). The gap has been stable over the full 40+ year window, with no sign of convergence, and amounts to 40,000–200,000 students per year expecting careers in representative fields that they will not attain.&lt;/p&gt;
&lt;p&gt;Q: Can alternative mechanisms such as overconfidence or motivated reasoning explain the results?
A: The authors argue no, for two reasons. First, students overestimate the prevalence of representative careers not only for majors they plan to pursue (where overconfidence or motivated reasoning might apply) but also for majors they do not plan to pursue — the pattern holds for the gray (population belief) bars across all ten majors in Figure 1. Second, a Shapley-Sharrocks decomposition reported in Table A.V shows that the stereotyping mechanism accounts for a larger share of variance in beliefs than any other mechanism tested. A pre-registered survey also rules out unawareness of non-representative occupations as a driver: students are aware of the overwhelming majority of the 100 most common non-representative occupations, and such unawareness as exists is uncorrelated with stereotyped beliefs.&lt;/p&gt;
&lt;p&gt;Q: What does the IAT reveal about the mechanism behind stereotyping?
A: The IAT was run on 434 OSU Exploration students in Fall 2021, measuring implicit associations between five major–career pairs (Humanities-Writers and Journalists, Sciences-Healthcare, STEM-Business, Social Science-Law, Social Science-Counseling/Education). Participants sorted stimuli faster in &amp;ldquo;matched&amp;rdquo; blocks (where the representative career shares a response key with its major) than in &amp;ldquo;unmatched&amp;rdquo; blocks, yielding DID-IAT effects of 0.30–0.36 SDs (p &amp;lt; 0.01) for all five pairs. After controlling for true career frequency with career and major fixed effects, the effect shrinks only slightly to 0.24–0.28 SDs (p &amp;lt; 0.01), confirming that associations are driven by representativeness beyond base rates. At the individual level, a one-SD increase in DID-IAT scores predicts 4.1 p.p. greater stereotyped beliefs (p &amp;lt; 0.01) without career-by-major fixed effects and 2.8 p.p. (p &amp;lt; 0.01) with them.&lt;/p&gt;
&lt;p&gt;Q: What does the role-model heterogeneity analysis show?
A: Students were asked which major–career combinations they knew personally. Controlling for career-by-major fixed effects, knowing someone with a non-representative major–career combination (i.e., a non-default path) predicts beliefs about the representative career that are 16 p.p. lower (p &amp;lt; 0.01). This is more than half the size of the baseline stereotyping effect (28–32 p.p.). Knowing such a person also predicts lower IAT scores (p &amp;lt; 0.01), implying that personal exposure can reduce both implicit associations and explicit stereotyped beliefs.&lt;/p&gt;
&lt;p&gt;Q: What does the equilibrium model predict about misallocation?
A: The model embeds stereotyped beliefs in a two-stage choice framework: students choose a major first, then choose a career after graduation. It shows two main results (Propositions 1 and 2 in Online Appendix A.1). First, students who perceive the representative career as more common than it is will infer — through a rational expectations mechanism — that the unobservable amenities of that career are particularly favorable, so they will be surprised upon graduation. Second, stereotyping raises misallocation because it draws in marginal students whose career preferences make them poorly matched to the major&amp;rsquo;s representative career, while the inframarginal students who would have chosen the major anyway are better matched. The misallocation effect increases in the extent of stereotyping.&lt;/p&gt;
&lt;p&gt;Q: What correlational evidence links stereotyping to post-graduation mismatch outcomes?
A: Using major-level stereotyping estimates from the OSU data merged with three nationally representative surveys (NSCG, SIPP, SHED), the authors find: a one-SD increase in major-level stereotyping is associated with 6.0% more job dissatisfaction (p &amp;lt; 0.01, NSCG), 3.1% more reports that the job does not fit the worker&amp;rsquo;s skills and experience (p &amp;lt; 0.05, NSCG), 5.4% more reports that the job is unrelated to the field of study (p &amp;lt; 0.05, SIPP), and 4.8% more regret about field of study choice (p &amp;lt; 0.05, SHED). The authors note these are correlational and cannot rule out confounders such as underlying complexity of the career mapping.&lt;/p&gt;
&lt;p&gt;Q: How does the field experiment work and what is its identifying strategy?
A: The experiment was embedded in the second 2021 OSU survey, with students in the treatment group shown the true share of their top major&amp;rsquo;s representative career before reporting beliefs and intentions; control students answered the same questions without receiving this information. The main regression relates outcomes to (True Share − Prior Belief), set to zero for controls. Because students with less accurate prior beliefs may be more likely to choose the relevant major, OLS is potentially inconsistent; the authors use an RD design where the running variable is the information shock (True Share − Prior Belief), with the threshold at zero. Students just above (who overestimated) receive negative news; students just below (who underestimated) receive positive news. The RD estimates are combined with a first-stage estimate of belief updating to produce IV estimates of the effect of a 10 p.p. change in beliefs. Balance tests on predetermined demographics confirm no discontinuities at the threshold.&lt;/p&gt;
&lt;p&gt;Q: What are the first-stage belief-updating results?
A: Students update their posterior beliefs in response to the treatment: in response to information that the representative career is 1 p.p. less likely, students update their posterior beliefs down by 0.37 p.p. (p &amp;lt; 0.01). This under-reaction is consistent with Bayesian updating when priors are informative (Mobius et al. 2022). Students also update beliefs about non-representative careers: a 1 p.p. reduction in the representative career&amp;rsquo;s stated likelihood increases the expected probability of other careers by 0.27 p.p. (p &amp;lt; 0.01).&lt;/p&gt;
&lt;p&gt;Q: What are the effects of the information intervention on major intentions?
A: A 10 p.p. reduction in beliefs about the top major&amp;rsquo;s representative career reduces intentions (stated probability of graduating with that major) by 3.5 p.p. (p &amp;lt; 0.01). This effect is similar across subgroups (Columns 2–4 of Table 2). For students&amp;rsquo; second-ranked major, a 10 p.p. reduction in stereotyping boosts intentions by 2.1 p.p. (p = 0.17), which is imprecisely estimated but consistent in sign with all other outcomes.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on actual course enrollments?
A: In the semester immediately following the experiment, learning that the representative career of the first major is 10 p.p. less likely causes students to enroll in 0.22 fewer credits in that major&amp;rsquo;s field (95% CI: [−0.41, −0.02], p &amp;lt; 0.05), relative to a mean of 0.85 credits. Learning that the representative career of the second major is 10 p.p. less likely causes students to enroll in 0.20 more credits in the second major&amp;rsquo;s field (95% CI: [0.004, 0.40], p &amp;lt; 0.10), relative to a mean of 0.36 credits.&lt;/p&gt;
&lt;p&gt;Q: What are the effects on official major declarations?
A: Within one year of the experiment, students who learned the representative career of their top major is 10 p.p. less likely are 6.1 p.p. less likely to have declared that major (95% CI: [−16.0, 3.8], p = 0.23) and 9.9 p.p. more likely to have declared their second major (95% CI: [2.5, 17.4], p &amp;lt; 0.01); the difference between these two effects is 16.0 p.p. (p &amp;lt; 0.01). By two years out, the effects are more attenuated. Treated students also spend on average 0.21 more semesters undecided before declaring a major (95% CI: [0.02, 0.40], p &amp;lt; 0.05). Effects do not appear to be driven by dropout: treated students are if anything slightly more likely to still be taking classes two to three years later.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Representativeness (likelihood ratio):&lt;/strong&gt; The representativeness R(c, M) of career c for major M is defined as the ratio p_{c|M} / p_{c|not-M} — how much more common career c is among major-M graduates than among graduates of all other majors. This is a relative, not absolute, frequency measure. The representative career (or exemplar) of a major is the career that maximizes this ratio.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stereotyping (as exaggeration of a kernel of truth):&lt;/strong&gt; In this paper&amp;rsquo;s framework, stereotyping means overweighting the representative career when forming beliefs about a major&amp;rsquo;s career distribution. The belief model is π_{c|M} = (1 − θ) p_{c|M} + θ · 1[c = c*(M)], where θ &amp;gt; 0 implies beliefs exaggerate how common the representative career is relative to equally prevalent non-representative careers. This is distinct from overconfidence, motivated reasoning, or simple noise.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DID-IAT score (difference-in-differences implicit association test):&lt;/strong&gt; The paper&amp;rsquo;s adaptation of the standard IAT to measure relative implicit associations between major and career groups. For a focal major–career pair, the DID-IAT score is the difference in the matched-vs-unmatched IAT D-score for the focal major (relative to a comparison major). A positive score indicates the focal major is more strongly associated with the focal career than the comparison major is. This measures implicit memory-based associations rather than deliberate beliefs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Misallocation (as used in the model):&lt;/strong&gt; The welfare loss arising because stereotyped beliefs draw marginal students — those on the margin between choosing the representative major and not — who have career preferences close to the average rather than being the students best suited to that major. These marginal students end up choosing careers other than the representative career after graduation at higher rates, producing major-career mismatch. Misallocation is shown (Proposition 2) to increase in the extent of stereotyping θ.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Information shock:&lt;/strong&gt; In the field experiment, the information shock for a given student and major is the difference between the true share of the major&amp;rsquo;s representative career and the student&amp;rsquo;s prior belief about that share. Positive shocks correspond to students who overestimated (and thus receive bad news); negative shocks correspond to students who underestimated (and receive good news). The RD design uses the threshold at shock = 0 to generate quasi-experimental variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Source text origin (implicit in the paper&amp;rsquo;s design):&lt;/strong&gt; The paper measures beliefs about career distributions benchmarked against American Community Survey data on actual career outcomes of college graduates aged 30–50, restricting to respondents born 1958–1997. This defines the objective ground truth against which stereotyping is measured throughout the paper.&lt;/p&gt;</description></item><item><title>What Works and for Whom? Effectiveness and Efficiency of School Capital Investments Across the U.S.</title><link>https://macropaperwarehouse.com/papers/what-works-and-for-whom-effectiveness-and-efficiency-of-school-capital-investments-across-the-u.s./</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/what-works-and-for-whom-effectiveness-and-efficiency-of-school-capital-investments-across-the-u.s./</guid><description>&lt;h2 id="what-works-and-for-whom-effectiveness-and-efficiency-of-school-capital-investments-across-the-us"&gt;What Works and for Whom? Effectiveness and Efficiency of School Capital Investments Across the U.S.&lt;/h2&gt;
&lt;h3 id="research-question"&gt;Research Question&lt;/h3&gt;
&lt;p&gt;This paper investigates which types of school facility investments benefit students (as measured by test scores) and are valued by homeowners (as measured by house prices), and for which student populations these investments are most effective. Prior state-level studies had reached conflicting conclusions about the returns to school capital spending, and no nationwide evidence had distinguished impacts across spending categories or student backgrounds.&lt;/p&gt;
&lt;h3 id="data-and-methodology"&gt;Data and Methodology&lt;/h3&gt;
&lt;p&gt;The authors assemble a novel panel dataset covering approximately 14,000 school bond referenda in 29 U.S. states and 10,146 districts enrolling 71% of all U.S. students, for the period 1990–2017. The dataset combines: (1) ballot-level bond election records including vote shares, proposed amounts, and ballot text; (2) district-level test scores from the Stanford Education Data Archive (SEDA) extended backward to 2003 for all states and as early as 1995 for some, normalized to a national scale via NAEP; (3) a Census-tract-level house price index (Contat and Larson, 2022) aggregated to school districts; and (4) NCES district finance and demographic data.&lt;/p&gt;
&lt;p&gt;Bond ballot texts are classified into eight spending categories using text-analysis: classroom construction/renovation; HVAC; other infrastructure (plumbing, roofs, furnaces); safety and health (pollutant removal, building safety); STEM equipment and labs; athletic facilities; land purchases; and transportation vehicles.&lt;/p&gt;
&lt;p&gt;The identification strategy exploits quasi-random variation from close bond elections, building on the dynamic regression discontinuity (DRD) framework of Cellini et al. (2010). A key methodological contribution is a stacked DRD design that addresses heterogeneous treatment effects correlated with timing: each treatment cohort (districts that narrowly authorize a bond in year c) is matched against &amp;ldquo;clean controls&amp;rdquo; — districts that also proposed a bond in the same cohort but narrowly failed to authorize it and did not authorize any bond in the following ten years. Cohorts are stacked, and a dynamic RD model is estimated controlling for cohort fixed effects and a district&amp;rsquo;s bond proposal history.&lt;/p&gt;
&lt;h3 id="main-findings-with-quantitative-magnitudes"&gt;Main Findings with Quantitative Magnitudes&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Average effects.&lt;/strong&gt; Bond authorization raises capital spending by approximately $1,650 per pupil cumulatively over five years. Test scores increase gradually, reaching 0.079 standard deviations (sd) higher five to eight years after authorization, and 0.073 sd higher nine to twelve years after. 2SLS estimates, amortizing spending over a 30-year project life at a 9% depreciation rate, imply that a $1,000 increase in the flow value of capital spending raises test scores by 0.048 sd. House prices rise by approximately 9% eight to nine years after authorization. When house price effects are estimated against only locally-financed capital spending (not state aid), the 2SLS estimate is 0.8% per $1,000 — roughly consistent with efficiency — suggesting that the larger reduced-form house price response is driven primarily by state aid that supplements local funds rather than by an inefficiently low ex ante spending level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity by spending category.&lt;/strong&gt; Category-specific estimates reveal that only certain project types raise test scores: HVAC (+0.20 sd, largest effect), safety and health (+0.15 sd), other infrastructure/plumbing/roofs (+0.15 sd), STEM equipment (+0.15 sd implied), and classroom space (+0.10 sd), all measured three to six years post-election. By contrast, bonds for athletic facilities, land purchases, and transportation produce no detectable effects on test scores. The pattern for house prices is the inverse: athletic facilities generate a 17% house price increase; classroom space generates 14%; STEM generates 11% — while HVAC and safety/health bonds produce no significant effect on house prices. The correlation between category-level test score and house price estimates is −0.07, indicating these are largely orthogonal outcomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity by student socioeconomic status.&lt;/strong&gt; Effects are concentrated in districts serving socioeconomically disadvantaged students (top tercile of the share of students eligible for free or reduced-price meals, denoted low-SES). In low-SES districts, bond authorization raises test scores by 0.13 sd after seven years and house prices by 15%; in high-SES districts, neither outcome shows a significant effect. 2SLS estimates confirm that a $1,000 increase in cumulative spending raises test scores by 0.08 sd in low-SES districts but produces no detectable change in high-SES districts. The SES gradient persists after conditioning on spending amounts, spending categories, and baseline capital stock, indicating that students in disadvantaged districts have higher marginal returns to capital improvements independent of these channels. High-minority districts (top tercile of Black and Hispanic share) similarly see a 0.12 sd test score gain and 15% house price gain after seven years, versus 0.04 sd and 3% in low-minority districts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Role of baseline capital stock.&lt;/strong&gt; Among districts with below-median capital stock, test score effects are 0.20 sd in low-SES districts seven years post-election. Even among above-median-stock districts, low-SES districts see house price effects exceeding 10% while high-SES districts see no effect. Differences by SES persist after conditioning on capital stock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Policy simulation.&lt;/strong&gt; Closing the spending gap between high- and low-SES districts (approximately $1,000 over 10 years) without changing the composition of spending would raise low-SES test scores by roughly 0.08 sd, closing about 8% of the roughly 1 sd achievement gap. Targeting that same additional spending toward HVAC and safety/health (the highest-impact categories) would generate test score increases approximately three times as large, potentially closing up to 25% of the observed achievement gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reconciling prior literature.&lt;/strong&gt; Replicating state-level estimates, the authors show that Ohio&amp;rsquo;s positive effects are explained by a high share of bonds in low-SES districts funding infrastructure, while Texas&amp;rsquo;s near-zero effects reflect a high share of bonds in higher-SES districts funding classrooms and athletic facilities.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-first-stage-effect-of-bond-authorization-on-capital-spending-and-does-it-contaminate-other-spending-categories"&gt;Q1. What is the first-stage effect of bond authorization on capital spending, and does it contaminate other spending categories?&lt;/h3&gt;
&lt;p&gt;A1: Bond authorization raises per-pupil capital spending by approximately $700 per year at two years post-election and $590 at three years, with cumulative spending $1,650 higher over five years in treated districts relative to districts that narrowly failed to authorize a bond. Bond revenues are legally restricted to capital uses, and the paper confirms that non-capital (current) spending and instructional spending are not affected following authorization. This establishes a clean first stage: bond authorization raises only capital outlays.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-standard-drd-estimator-of-cellini-et-al-2010-require-refinement-and-what-problem-does-the-stacked-drd-design-solve"&gt;Q2. Why does the standard DRD estimator of Cellini et al. (2010) require refinement, and what problem does the stacked DRD design solve?&lt;/h3&gt;
&lt;p&gt;A2: The original CFR estimator assumes treatment effects are uncorrelated with the timing of treatment — an assumption potentially violated when, for example, bonds financing HVAC (high-impact) versus athletic facilities (amenity-focused) have different propensities to be proposed at different points in time. The stacked DRD design avoids &amp;ldquo;forbidden comparisons&amp;rdquo; by comparing each treatment cohort only against clean controls that propose but fail to authorize a bond in the same year and do not authorize any bond in the subsequent ten years. This ensures consistency even when treatment effects are heterogeneous across cohorts and correlated with timing.&lt;/p&gt;
&lt;h3 id="q3-how-do-the-authors-validate-the-quasi-random-assignment-assumption-of-the-regression-discontinuity-design"&gt;Q3. How do the authors validate the quasi-random assignment assumption of the regression discontinuity design?&lt;/h3&gt;
&lt;p&gt;A3: Three tests are performed. First, a McCrary (2008) density test on the vote margin distribution shows no discontinuity at the cutoff in the pooled or stacked data (p-values of 0.59 and 0.24, respectively), though discontinuities are found in Arkansas, Missouri, and Oklahoma — those three states are excluded. Second, pre-election district covariates (income, education, SES shares, enrollment, revenues, expenditures) are smooth around the cutoff in both datasets. Third, pre-election trends in test scores and house prices are flat and parallel between marginally approved and marginally rejected districts.&lt;/p&gt;
&lt;h3 id="q4-how-are-the-eight-spending-categories-constructed-and-how-many-bonds-are-successfully-classified"&gt;Q4. How are the eight spending categories constructed, and how many bonds are successfully classified?&lt;/h3&gt;
&lt;p&gt;A4: Categories are drawn from the SchoolBondFinder.com classification produced by The Amos Group, then refined by splitting capital improvements into HVAC versus other infrastructure, splitting construction/renovation into classroom versus athletic facility projects, and adding land purchases as a separate category. Keyword-based text analysis of ballot language successfully assigns 75% of the approximately 14,000 bonds to at least one of the eight categories. More than two-thirds of classified bonds receive multiple category designations, with a mean of 2.9 categories per proposed bond and 3.2 per authorized bond.&lt;/p&gt;
&lt;h3 id="q5-why-do-hvac-bonds-raise-test-scores-but-not-house-prices-while-athletic-facility-bonds-raise-house-prices-but-not-test-scores"&gt;Q5. Why do HVAC bonds raise test scores but not house prices, while athletic facility bonds raise house prices but not test scores?&lt;/h3&gt;
&lt;p&gt;A5: The authors interpret this divergence as reflecting what different types of improvements offer to different stakeholders. HVAC improvements reduce excessive heat and air pollution exposure in classrooms, directly improving students&amp;rsquo; learning experiences — consistent with Park et al. (2020) on heat and Gilraine and Zheng (2022) on air pollution. These improvements are not visibly salient to homeowners without school-age children and carry no amenity value for the broader community. Athletic facilities, by contrast, are highly visible and provide a community amenity valued in the housing market regardless of their impact on academic instruction. The near-zero correlation (−0.07) between category-level test score and house price estimates confirms that the two outcomes respond to largely distinct features of capital investments.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-three-candidate-explanations-for-the-larger-effects-of-bond-authorization-in-low-ses-districts-and-which-explanations-survive-empirical-scrutiny"&gt;Q6. What are the three candidate explanations for the larger effects of bond authorization in low-SES districts, and which explanations survive empirical scrutiny?&lt;/h3&gt;
&lt;p&gt;A6: The three candidates are: (1) larger spending increases after authorization in low-SES districts; (2) a different composition of spending categories (more toward high-impact HVAC and safety); and (3) higher marginal returns per dollar for disadvantaged students, holding spending size and composition fixed. The data confirm all three operate, but the third is the residual: 2SLS estimates show a $1,000 increase raises test scores by 0.08 sd in low-SES districts versus a statistically zero effect in high-SES districts, and within-category estimates show HVAC bonds raise scores by 0.27 sd in low-SES districts but have no detectable effect in high-SES districts. Differences by SES also persist after conditioning on the estimated baseline capital stock, though low capital stock accounts for part of the gap.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-role-of-state-aid-alter-the-interpretation-of-the-house-price-effect-for-spending-efficiency"&gt;Q7. How does the role of state aid alter the interpretation of the house price effect for spending efficiency?&lt;/h3&gt;
&lt;p&gt;A7: A 9% house price increase after bond authorization, if taken at face value under Brueckner&amp;rsquo;s (1979) efficiency test, would suggest the ex ante level of school capital spending was inefficiently low. However, state grants that partly match local bond revenues raise actual spending without raising local property taxes proportionally. When the 2SLS house price effect is estimated against only locally financed capital spending (using proposed bond size as the relevant measure), the implied house price increase is just 0.8% per $1,000 — consistent with rough efficiency on average across the full sample. The authors conclude that the large reduced-form house price response is driven primarily by the capitalization of state aid, not by an undersupply of capital investments at the aggregate level.&lt;/p&gt;
&lt;h3 id="q8-does-household-sorting-account-for-the-observed-test-score-and-house-price-gains-following-bond-authorization"&gt;Q8. Does household sorting account for the observed test score and house price gains following bond authorization?&lt;/h3&gt;
&lt;p&gt;A8: Bond authorization produces small but detectable compositional changes: the share of high-SES students is approximately 3 percentage points higher seven years after an election (a roughly 4% increase relative to an average share of 0.73), while enrollment and the share of white students are largely unaffected. However, controlling for district-by-year shares of each sociodemographic group only slightly attenuates the test score and house price estimates, indicating that sorting accounts for a small share of the observed gains.&lt;/p&gt;
&lt;h3 id="q9-are-the-findings-robust-to-alternative-research-designs"&gt;Q9. Are the findings robust to alternative research designs?&lt;/h3&gt;
&lt;p&gt;A9: The results are robust to five alternative estimation approaches: (1) the original one-step TOT estimator of Cellini et al. (2010); (2) a version of the stacked DRD where clean controls are districts that do not approve any bonds in the full [c−5, c+10] window; (3) a version that matches treated and control districts in each cohort based on bond history; (4) a version not controlling for future bond history; and (5) the extended two-way fixed effects (ETWFE) estimator of Wooldridge (2021). Results are also robust to linear polynomials with different slopes and quadratic polynomials of the vote margin.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-capital-stock-measure-illuminate-mechanism-and-what-are-its-limitations"&gt;Q10. How does the capital stock measure illuminate mechanism, and what are its limitations?&lt;/h3&gt;
&lt;p&gt;A10: The authors construct a district-level capital stock as the 30-year depreciated sum of capital spending from Census of Governments data (1967–2017) at a 5% depreciation rate. This stock is negatively correlated with the share of low-SES students, confirming that more disadvantaged students attend schools in worse structural condition. Conditioning on this proxy, the SES gradient in bond impacts is reduced but remains. Among districts with below-median capital stock, low-SES districts see test score gains of 0.20 sd after seven years, while among above-median-stock districts the gap narrows to approximately 0.10 vs. 0.05 sd. A key limitation is that detailed school-condition data are unavailable nationally, so the capital stock is a proxy only.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-quantitative-policy-implication-of-the-targeting-exercise"&gt;Q11. What is the quantitative policy implication of the targeting exercise?&lt;/h3&gt;
&lt;p&gt;A11: On average, low-SES districts receive about $97 per pupil per year less in capital spending than high-SES districts, so closing this gap over ten years implies approximately $970 in additional cumulative spending. Without changing spending composition, this would raise test scores by roughly 0.08 sd in low-SES districts, closing about 8% of the approximately 1 sd achievement gap between high- and low-SES districts. Redirecting that same additional spending toward the highest-impact categories (HVAC and safety/health) would generate test score gains roughly three times larger, potentially closing up to 25% of the observed achievement gap.&lt;/p&gt;
&lt;h3 id="q12-how-do-the-cross-state-differences-documented-in-prior-literature-map-onto-the-papers-heterogeneity-findings"&gt;Q12. How do the cross-state differences documented in prior literature map onto the paper&amp;rsquo;s heterogeneity findings?&lt;/h3&gt;
&lt;p&gt;A12: The authors replicate earlier state-level estimates and show that Ohio&amp;rsquo;s relatively large positive effects — found by Conlin and Thompson (2017) — are explained by a high concentration of bonds in low-SES districts funding infrastructure, while Texas&amp;rsquo;s near-zero effects — found by Martorell et al. (2016) — reflect a high share of bonds in higher-SES districts funding classrooms and athletic facilities. Wisconsin and Michigan, which showed null effects in earlier studies, similarly have bond compositions and student demographics that predict small impacts under the paper&amp;rsquo;s heterogeneity framework.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Stacked Dynamic Regression Discontinuity (Stacked DRD).&lt;/strong&gt; The paper&amp;rsquo;s primary estimation strategy, which combines the dynamic RD framework of Cellini et al. (2010) with a stacked-cohort design adapted from the staggered difference-in-differences literature. For each treatment cohort (year in which a bond barely passes), &amp;ldquo;clean controls&amp;rdquo; are defined as districts that also proposed a bond in the same year but narrowly failed to authorize it and did not authorize any subsequent bond within ten years. Cohort-specific datasets are stacked and estimated jointly with cohort fixed effects, ensuring that estimates are robust to treatment effect heterogeneity correlated with timing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Clean Controls.&lt;/strong&gt; Districts used as the counterfactual for treated districts in a given cohort: those that propose a bond in the same year as the treated cohort, barely fail to authorize it, and remain untreated for ten subsequent years. Their &amp;ldquo;clean&amp;rdquo; status is quasi-random because their future non-authorization results from narrow electoral loss rather than any endogenous district choice.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Bond Spending Categories.&lt;/strong&gt; Eight mutually-non-exclusive classifications of bond spending derived from ballot text using keyword analysis: classroom space; HVAC; other infrastructure (plumbing, roofs, furnaces); safety and health (pollutant removal, compliance upgrades); STEM equipment and labs; athletic facilities; land purchases; and transportation. These categories are defined in the paper not by administrative accounting codes but by the stated intended use of funds in ballot language.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Treatment-on-the-Treated (TOT) Estimator.&lt;/strong&gt; The CFR estimator that captures the effect of bond authorization against the counterfactual of never authorizing a bond in the foreseeable future, achieved by including leads and lags of a district&amp;rsquo;s bond proposal history as controls. This addresses the problem that multiple elections over time make simple treated-vs-control comparisons confounded by past and future bond activity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital Stock (District-Level Proxy).&lt;/strong&gt; A measure of each district&amp;rsquo;s accumulated school facility capital at a given point in time, constructed as the depreciated 30-year running sum of capital expenditures from the Census of Governments, using a 5% annual depreciation rate. Used as a proxy for facility conditions in the absence of nationally available building-quality data, and confirmed to be negatively correlated with district share of low-SES students.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Brueckner Efficiency Test.&lt;/strong&gt; An application of the theoretical framework linking public good provision levels to house price responses. If a spending increase raises house prices, the initial spending level was below the efficient level; if it lowers house prices, spending was too high. In this paper, the test is refined to use only locally-financed capital spending as the explanatory variable, to strip out the capitalization of state aid and isolate the efficiency assessment for locally-determined spending.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Socio-Economic Status (SES) Terciles.&lt;/strong&gt; Districts are ranked by the share of students eligible for free or reduced-price school meals as of 1995. &amp;ldquo;Low-SES districts&amp;rdquo; refers to those in the top tercile of this share (most disadvantaged); &amp;ldquo;high-SES districts&amp;rdquo; refers to those in the bottom tercile (least disadvantaged). Effects are estimated separately for these subsamples throughout.&lt;/p&gt;</description></item><item><title>Who's Afraid of the Minimum Wage? Measuring the Impacts on Independent Businesses Using Matched U.S. Tax Returns</title><link>https://macropaperwarehouse.com/papers/whos-afraid-of-the-minimum-wage-measuring-the-impacts-on-independent-businesses-using-matched-u.s.-tax-returns/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/whos-afraid-of-the-minimum-wage-measuring-the-impacts-on-independent-businesses-using-matched-u.s.-tax-returns/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks how independent (pass-through) businesses in the United States accommodate minimum wage increases — specifically whether they reduce employment, compress profits, pass costs through to customers, or exit — and what happens to the low-earning workers and business owners affected by these adjustments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors construct a novel linked firm-worker-owner panel dataset from the universe of U.S. tax returns, covering approximately 235,000 pass-through firms (S-corporations, partnerships, and LLCs) per year in highly exposed industries over 2010–2019. &amp;ldquo;Highly exposed&amp;rdquo; industries are defined as those where at least 15% of workers earned below the full-time equivalent of the federal minimum wage ($15,080 per year) in 2013. The dataset links annual business income tax returns to the individual income tax returns and W-2 information reports of all workers and owners.&lt;/p&gt;
&lt;p&gt;The causal identification strategy exploits the six state minimum wage increases that took effect in 2014 (California, Connecticut, Delaware, Michigan, Minnesota, and New Jersey) relative to 24 states that did not change their wage floors at any point from 2012–2018. The empirical workhorse is a panel difference-in-differences event study (Equation 1), augmented by DFL re-weighting (DiNardo et al., 1996) to improve comparability of treatment and control firms on observables. The analysis covers cumulative effects through 2018, by which point the average minimum wage across treatment states had risen 30.6%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Employment:&lt;/strong&gt; The average exposed independent firm does not meaningfully reduce employment. The authors estimate an own-wage elasticity of -0.209 (s.e. = 0.0112). Employment adjustments manifest as moderately lower hiring rather than layoffs of existing workers. Reduced hiring is wholly concentrated among teenagers and very part-time jobs paying less than $3,900 annually (with 67% earning less than $1,000 per year).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Worker earnings:&lt;/strong&gt; Despite the hiring reduction, low-earning workers employed at exposed independent firms experience average earnings gains of approximately $2,000 per year by 2018, relative to comparable workers in untreated states. Young individuals aged 20–26 without a 2013 job earn roughly $4,000 more per year by 2018; teenagers without a 2013 job gain approximately $1,000 per year. Workers in these groups are no less likely — and in some cases slightly more likely — to be employed five years after the minimum wage increase.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Wage bills:&lt;/strong&gt; Average wage bills among surviving treated firms rose 7.03% (s.e. = 0.0153) by 2018. Earnings gains are concentrated among workers earning $15,600–$35,000 annually, with no evidence of reduced earnings for higher-paid workers. The 7% average wage bill increase amounts to only 1.4% of 2013 firm revenues, easing pass-through.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Revenue and profits:&lt;/strong&gt; Revenues of surviving treated firms grew approximately 2.1% more than control firms by 2018. On average, this revenue increase fully offsets the higher wage bill, yielding a small net profit increase of roughly $3,360 (s.e. = $1,123) per owner by 2018, or about 2.7% of mean 2013 owner income.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Firm exit:&lt;/strong&gt; On average across all highly exposed industries, minimum wages increased the five-year exit probability by 0.9 percentage points (s.e. = 0.0029), relative to a baseline raw exit rate of approximately 29%. Exit effects are driven entirely by restaurants: by 2018, restaurants in treated states were 1.85 percentage points (s.e. = 0.0039) more likely to have exited, while the exit response for non-restaurant exposed firms is a precisely estimated zero.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity by productivity within restaurants:&lt;/strong&gt; Exit is concentrated entirely in the bottom productivity quartile (coefficient = 0.0254, s.e. = 0.0079), with no significant effect in the upper three quartiles. Profits among surviving small restaurants rise by $5,941 (s.e. = $1,546) by 2018 relative to 2013. Among small restaurants, the profit gains are larger for firms in the higher productivity quartiles (Q3: +$7,915; Q4: +$9,161). Surviving restaurants also increase non-labor input spending by 2.53% (s.e. = 0.0101), consistent with expanded output following competitor exits.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Entrant characteristics:&lt;/strong&gt; Post-reform restaurant entrants in treatment states have higher wage bills (13.8% higher in logs), higher revenues (4.0% higher), higher value-added (8.4% higher), and higher productivity (net income/revenue ratio 2.24 percentage points higher) than entrants in control states, indicating the minimum wage raises the productivity floor for new entrants.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Owner outcomes after exit:&lt;/strong&gt; Owners of small restaurants forced out by the minimum wage are significantly less likely to own an independent business five years later, but earn no less on average in wages plus business income. Policy-induced exiters are significantly less likely to report negative incomes, suggesting substitution away from risky or marginally profitable business ownership.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Theoretical Framework&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors present a Cournot competition model with heterogeneous firm productivity and fixed production costs. A minimum wage cost shock raises marginal costs, narrowing margins for all firms. Firms whose cost increases exceed the market price increase cannot cover fixed costs and exit. Remaining firms gain higher markups and larger market shares as demand is reallocated from exiting firms. Selection on ex-ante productivity (the least productive firms exit) limits the distortion to market quantity and amplifies profit gains among productive survivors. The model predicts profit increases only in markets with firm exit, which matches the data: profits rise among restaurants (where exit occurs) but not among retailers (where exit is a precisely estimated zero).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Findings pertain to the short-to-medium run (up to five years post-legislation) of phased-in minimum wage increases averaging 30.6% in six U.S. states. The sample covers pass-through (independent) businesses in highly exposed industries. Longer-run effects may differ if entrants adopt production technologies that rely less on low-wage labor or incumbents reconfigure inputs. Border-county retailers appear to be less able to pass through costs than interior firms, suggesting product market competition is a key moderating factor.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-do-the-authors-focus-on-pass-through-businesses-rather-than-publicly-traded-corporations"&gt;Q1. Why do the authors focus on pass-through businesses rather than publicly traded corporations?&lt;/h3&gt;
&lt;p&gt;Pass-throughs (S-corporations, partnerships, and LLCs) comprise 78% of non-sole-proprietorship businesses and 79% of firms with fewer than 20 employees. They represent the majority organizational form for independent businesses in virtually all two-digit NAICS industry groups except utilities and enterprise management. Because minimum wage concerns are disproportionately raised on behalf of small independent businesses, and because most minimum wage workers in restaurants are employed at pass-throughs, studying pass-throughs directly addresses the policy debate. Additionally, pass-through tax returns link business income directly to the individual tax returns of each owner, enabling the authors to separately identify employee versus owner responses.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-authors-define-highly-exposed-industries-and-why-does-this-matter-for-identification"&gt;Q2. How do the authors define &amp;ldquo;highly exposed&amp;rdquo; industries and why does this matter for identification?&lt;/h3&gt;
&lt;p&gt;Highly exposed industries are defined as four-digit NAICS industries where at least 15% of workers earned below the full-time federal minimum wage equivalent ($15,080 per year) in 2013, using tax data to construct a proxy for minimum wage workers. The analysis focuses on these industries because minimum wage workers are extremely concentrated — the vast majority are in Leisure/Hospitality and Retail. Restricting to highly exposed industries allows the authors to estimate average effects within affected markets and conduct heterogeneity analysis across firm characteristics within those markets, including comparing firms with different baseline shares of low-earning workers that nonetheless all face the market-level cost shock.&lt;/p&gt;
&lt;h3 id="q3-how-do-the-employment-effects-decompose-into-hiring-versus-retention"&gt;Q3. How do the employment effects decompose into hiring versus retention?&lt;/h3&gt;
&lt;p&gt;The average firm subject to a higher wage floor does not lay off existing workers (the retention line is flat in event study estimates). By 2018, firms in treated states hire roughly one fewer worker on average than similar firms in control states, entirely through reduced hiring. This reduced hiring is wholly concentrated among teenagers in very part-time jobs: the missing hires consist entirely of workers who would have earned less than $3,900 annually, with 67% earning less than $1,000 per year. Simultaneously, workers already employed at exposed firms are 2 to 4 percentage points more likely to remain with their 2013 employer by 2016, with prime-age low-earning workers exhibiting the largest retention increases.&lt;/p&gt;
&lt;h3 id="q4-what-happens-to-low-earning-workers-and-young-people-in-individual-level-panels"&gt;Q4. What happens to low-earning workers and young people in individual-level panels?&lt;/h3&gt;
&lt;p&gt;Low-earners (those earning below $25,000 in each year from 2012–2014) at exposed independent firms experience average earnings gains of approximately $2,000 per year by 2018 relative to similar workers in untreated states, including teenage low-earners. Young individuals aged 20–26 with no job in 2013 experience a relative earnings increase of approximately $4,000 per year by 2018; teenagers without jobs in 2013 gain approximately $1,000 per year. These workers are no less likely — and often slightly more likely — to be employed relative to their counterparts in control states, so the earnings gains are not offset by employment losses at the individual level.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-magnitude-of-the-cost-shock-for-firms-and-how-does-it-compare-to-revenues"&gt;Q5. What is the magnitude of the cost shock for firms and how does it compare to revenues?&lt;/h3&gt;
&lt;p&gt;By 2018, the average wage bill among surviving firms in treated states was 7.03% (s.e. = 0.0153) higher than comparable firms in control states. This is consistent with a back-of-envelope calculation: low-earning workers account for about 21% of wage bills at these firms, and states raised minimum wages by 30.6% on average (0.21 × 0.306 = 0.064). However, the 7% wage bill increase amounts to only approximately 1.4% of 2013 firm revenues, making cost pass-through relatively modest. Higher minimum wages have no discernible impact on pension contributions but slightly reduce deductions for other benefits including health insurance.&lt;/p&gt;
&lt;h3 id="q6-how-do-surviving-firms-finance-the-increased-wage-bill-and-what-happens-to-profits"&gt;Q6. How do surviving firms finance the increased wage bill, and what happens to profits?&lt;/h3&gt;
&lt;p&gt;Surviving firms finance the wage increase primarily through higher revenues. By 2018, revenues of firms in treated states grew approximately 2.1% more than revenues of firms in control states. On average, this revenue increase outpaces the higher wage bill, resulting in a net profit increase of approximately $3,360 (s.e. = $1,123) per owner by 2018, representing about 2.7% of mean 2013 owner income. There is no evidence of redistribution from middle- or high-income workers within firms; wage bill increases are concentrated among workers earning $15,600–$35,000 annually, consistent with minimum wage spillovers to workers slightly above the statutory floor.&lt;/p&gt;
&lt;h3 id="q7-why-do-restaurants-experience-exit-effects-but-retailers-do-not"&gt;Q7. Why do restaurants experience exit effects but retailers do not?&lt;/h3&gt;
&lt;p&gt;The asymmetry stems from the intensity of low-wage labor in production. While low-earning workers account for a similar share of labor costs at restaurants (41.8%) and retailers (38.5%), labor costs overall are more than twice as large at restaurants relative to retailers. Wage bills account for 39% of variable costs and 27% of revenues at restaurants, but only 16% of variable costs and 13% of revenues at retailers. As a result, raising the minimum wage raises variable costs by 5.76% at restaurants. Non-restaurant exposed firms are able to fully pass through their smaller cost shock, yielding flat profits and neither employment nor exit impacts.&lt;/p&gt;
&lt;h3 id="q8-why-is-firm-exit-concentrated-in-the-lowest-productivity-quartile-of-restaurants-rather-than-among-the-most-exposed-firms"&gt;Q8. Why is firm exit concentrated in the lowest productivity quartile of restaurants rather than among the most exposed firms?&lt;/h3&gt;
&lt;p&gt;The Cournot framework predicts exits among firms with the lowest ex-ante productivity (highest marginal costs), the largest cost shock (highest share of low-wage labor per unit of output), or a combination. Empirically, productivity is the primary determinant: restaurants across all productivity quartiles use similar shares of low-earning workers (40–44% of wage bills for Q1 through Q4). Exit rises significantly only among restaurants in the bottom productivity quartile (coefficient = 0.0254, s.e. = 0.0079), with no significant effects in Q2–Q4. Among the lowest-productivity restaurants, those most dependent on low-earning labor face the largest exit rates.&lt;/p&gt;
&lt;h3 id="q9-how-do-the-models-predictions-about-profit-heterogeneity-match-the-data"&gt;Q9. How do the model&amp;rsquo;s predictions about profit heterogeneity match the data?&lt;/h3&gt;
&lt;p&gt;The Cournot model predicts profits should rise only in markets with firm exit (via increased margins and market share reallocation to survivors). This is exactly what the data show. Among restaurants, where exit is concentrated in the bottom productivity quartile, profits among surviving small restaurants rise by $5,941 (s.e. = $1,546) by 2018. Among small restaurants specifically, profit gains increase with productivity: Q3 restaurants gain $7,915 (s.e. = $3,326) and Q4 restaurants gain $9,161 (s.e. = $2,127), while Q1 and Q2 gains are statistically indistinguishable from zero. In non-restaurant exposed industries where the exit effect is a precise zero, profits are also flat — exactly as the model predicts.&lt;/p&gt;
&lt;h3 id="q10-what-happens-to-the-characteristics-of-new-restaurant-entrants-after-the-minimum-wage-increase"&gt;Q10. What happens to the characteristics of new restaurant entrants after the minimum wage increase?&lt;/h3&gt;
&lt;p&gt;Post-reform restaurant entrants in treatment states are systematically more productive than entrants in control states. They have wage bills 13.8% higher (in logs), revenues 4.0% higher, value-added 8.4% higher, and productivity ratios (net income/revenue) 2.24 percentage points higher than new entrants in control markets. This implies the minimum wage raises the minimum viable productivity threshold for entrant restaurants, consistent with Sorkin (2015)&amp;rsquo;s insight that minimum wages shape the capital and technology choices of entering firms. The restaurant industry thus becomes more productive on average through both the exit of the least productive incumbents and the entry of more productive new firms.&lt;/p&gt;
&lt;h3 id="q11-how-do-worker-transition-patterns-reflect-the-reallocation-of-output-to-surviving-firms"&gt;Q11. How do worker transition patterns reflect the reallocation of output to surviving firms?&lt;/h3&gt;
&lt;p&gt;Workers at large independent businesses (top revenue quartile) are 3.52 percentage points more likely to remain with their 2013 employer in 2018 and 2.36 percentage points less likely to switch to another large firm. The large firms that retain more of their existing workforce also reduce their hiring of very part-time teenagers the most — in the top revenue quartile, firms shed roughly 4.5 employment relationships on average, comprising higher retention of 4.15 existing workers offset by reduced hiring of 8.67 very part-time teenage workers. Workers originally at smaller exposed firms are more likely to be found working at larger firms five years out, consistent with demand reallocation from exiting and shrinking small firms toward larger, more productive survivors.&lt;/p&gt;
&lt;h3 id="q12-what-happens-to-owners-of-restaurants-that-exit-due-to-the-minimum-wage"&gt;Q12. What happens to owners of restaurants that exit due to the minimum wage?&lt;/h3&gt;
&lt;p&gt;Policy-induced exiters of small restaurants are significantly less likely to own an independent business five years later and less likely to receive all earnings from business ownership, relative to owners of restaurants that exited for other reasons in control states. However, their average incomes (wage income plus ordinary business income) are no lower. This income stability is partly explained by the fact that policy-induced exiters are significantly less likely to report negative incomes five years out, suggesting they substitute away from potentially risky or marginally profitable business ownership toward wage employment or other activities. The utility implications are ambiguous: these former owners may have preferred business ownership even if it did not yield higher income.&lt;/p&gt;
&lt;h3 id="q13-what-is-the-role-of-product-market-competition-in-mediating-pass-through-as-evidenced-by-border-county-analysis"&gt;Q13. What is the role of product market competition in mediating pass-through, as evidenced by border-county analysis?&lt;/h3&gt;
&lt;p&gt;The border county robustness analysis reveals that product market competition is central to pass-through success. Retailers near state borders, where consumers can cross-state-border shop, face more elastic demand and are less able to finance the wage cost shock with new revenues, exhibiting reduced profits and higher exit rates (though estimates are imprecise). Further from the border, where the cost shock is more commonly felt by all potential substitutes (making market demand elasticity rather than firm demand elasticity the relevant parameter), results are very similar to the full-sample aggregate findings. This confirms that the common nature of the minimum wage cost shock — shared by all competing firms in the market — is a key reason firms can pass through costs to consumers.&lt;/p&gt;
&lt;h3 id="q14-how-do-the-findings-address-the-divide-among-independent-business-owners-on-minimum-wage-policy"&gt;Q14. How do the findings address the divide among independent business owners on minimum wage policy?&lt;/h3&gt;
&lt;p&gt;The heterogeneous outcomes rationalize why surveys consistently find business owners divided. Among restaurants, some owners (those operating the least productive small restaurants) face exit and loss of business ownership, while surviving productive restaurateurs see higher profits of $5,941–$9,161 per year. Among non-restaurant exposed businesses, owners are broadly unaffected in terms of profits and viability. Uncertainty about whether a given firm&amp;rsquo;s demand is elastic enough to bear cost pass-through — given that owners may be more familiar with the elasticity of firm-level demand from prior unilateral price changes, rather than the relevant market-level demand elasticity applying to a common cost shock — may broaden opposition to include even owners who would ultimately benefit.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Pass-through businesses (independent businesses):&lt;/strong&gt; Privately owned firms organized as S-corporations, partnerships, or LLCs, taxed by passing income through to the individual returns of owners rather than at the entity level. In 2015, these comprised 78% of non-sole-proprietorship U.S. businesses and 46% of employment. The paper uses &amp;ldquo;pass-through&amp;rdquo; and &amp;ldquo;independent business&amp;rdquo; interchangeably as the unit of analysis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Highly exposed industries:&lt;/strong&gt; Four-digit NAICS industries where at least 15% of workers earned below the annual full-time equivalent of the federal minimum wage ($15,080) in 2013, as measured in the authors&amp;rsquo; administrative tax data. This threshold proxies the concentration of minimum-wage workers across industries and drives the sample selection for firm-level analysis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Own-wage elasticity of employment:&lt;/strong&gt; The estimated percentage change in employment at a firm associated with a given percentage change in the firm&amp;rsquo;s minimum wage. The authors estimate this as -0.209 (s.e. = 0.0112), reflecting the average effect across all exposed independent businesses, conditional on the firm&amp;rsquo;s industry, size, and local market characteristics.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DFL re-weighting (DiNardo-Fortin-Lemieux):&lt;/strong&gt; A non-parametric reweighting procedure that adjusts the distribution of control-group firms to match the distribution of treatment-group firms on observables (specifically, two-year lagged value-added within three-digit NAICS industries). Used to improve pre-reform comparability of treatment and control firm samples without parametric functional form assumptions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Firm productivity (in this paper&amp;rsquo;s sense):&lt;/strong&gt; Measured as the ratio of net profits to revenues (net income/revenue) at the firm level in the base year 2013, used to assign firms to productivity quartiles for heterogeneity analysis. This is a firm-level profitability measure constructed from pass-through tax returns, not a total factor productivity estimate requiring production function estimation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Firm exit:&lt;/strong&gt; An indicator for a firm that filed a tax return in 2013 but did not file a return in a subsequent year t. The average one-year exit rate for highly exposed independent businesses is 5.2%; the cumulative five-year raw exit rate is approximately 29% across treatment and control states.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cournot competition with heterogeneous productivity and fixed costs:&lt;/strong&gt; The paper&amp;rsquo;s conceptual framework, in which N firms compete in quantities with asymmetric marginal costs (reflecting heterogeneous productivity), a common output price, and a fixed cost of production. Under this framework, a minimum wage cost shock narrows margins unevenly, induces exit among firms that cannot cover fixed costs, and generates both demand reallocation and market share gains for productive survivors — rationalizing simultaneous exit and profit increases in the same industry.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Common cost shock:&lt;/strong&gt; The property that a minimum wage increase raises production costs for all firms employing low-wage workers in the same market simultaneously. Because all competing firms face higher costs, the relevant pass-through parameter is the elasticity of market demand rather than the (higher) elasticity of individual firm demand, facilitating cost pass-through to consumers and distinguishing minimum wages from unilateral price changes by a single firm.&lt;/p&gt;</description></item><item><title>Why Doesn't the United States Have National Health Insurance?</title><link>https://macropaperwarehouse.com/papers/why-doesnt-the-united-states-have-national-health-insurance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/why-doesnt-the-united-states-have-national-health-insurance/</guid><description>&lt;p&gt;This paper investigates a critical juncture in the development of national health insurance (NHI) in the United States: the post-World War II period when most peer nations moved to establish comprehensive public coverage while the U.S. did not. The authors examine the causal role of the American Medical Association (AMA), which in 1949 hired Whitaker &amp;amp; Baxter&amp;rsquo;s Campaigns, Inc. — the country&amp;rsquo;s first political public relations firm — to direct a nationwide campaign opposing NHI and promoting private (voluntary) health insurance (PHI).&lt;/p&gt;
&lt;p&gt;The Campaign had two main components. First, a physician outreach component in which AMA members distributed pamphlets to patients warning against &amp;ldquo;socialized medicine&amp;rdquo; and encouraging enrollment in private plans, and acted as liaisons to local civic organizations to solicit resolutions against NHI sent to elected officials (nearly 50 million pieces of material were sent to physicians). Second, a mass newspaper advertising component, in which a standard ad was placed across newspapers nationwide, with an additional $19 million (approximately $240 million in current dollars) in coordinated tie-in advertising from roughly 23,000 corporations and industry associations. The messaging framed NHI as &amp;ldquo;un-American&amp;rdquo; and associated private insurance with &amp;ldquo;freedom&amp;rdquo; and &amp;ldquo;the American way,&amp;rdquo; providing little substantive information about insurance products.&lt;/p&gt;
&lt;p&gt;The authors construct novel measures of Campaign exposure by combining (a) per capita pamphlets distributed by AMA physicians and (b) per capita advertising circulation scaled by local newspaper readership, using archival data from the Whitaker &amp;amp; Baxter Archives (Sacramento), the National Archives (Washington D.C.), digitized AMA Medical Directories, the N.W. Ayer &amp;amp; Son&amp;rsquo;s Newspaper Directory, and newly discovered Blue Shield enrollment data from AMA Council on Medical Service annual reports covering 1946–1954.&lt;/p&gt;
&lt;p&gt;The primary estimation strategy exploits spatial variation in Campaign intensity combined with its timing, using event studies with state and year fixed effects and design controls for income per capita and unionization. The identifying assumption — that Campaign intensity was conditionally as-good-as-randomly assigned — is supported by balance tests showing no pre-Campaign correlation between exposure and enrollment or sociodemographic characteristics (with the exception of Black population share), and by the historical record that the Campaign was organized hastily following Truman&amp;rsquo;s unexpected 1948 electoral victory.&lt;/p&gt;
&lt;p&gt;Main findings: A one standard deviation increase in Campaign exposure explains approximately 20% of the post-Campaign increase in PHI enrollment, corresponding to roughly 14 million additional enrollees — an effect comparable in magnitude to increasing average per capita income by approximately $100 (about 7 percent). On public opinion, a one standard deviation increase in Campaign exposure led to a six percentage point decline in popular support for NHI per Gallup survey wave, a reversal occurring against a backdrop of 69% pre-Campaign approval that was trending upward. For context, this six-point magnitude approximates the entire gap in NHI support between union and non-union households, or one-third the racial gap in support. Campaign intensity also predicts civic organizations passing resolutions favoring PHI, Republican legislators adopting speech semantically similar to Campaign propaganda, and — by 1952 — AMA members being five times more likely to donate to the Eisenhower-Nixon ticket than non-AMA physicians, with donation rates increasing in Campaign intensity.&lt;/p&gt;
&lt;p&gt;Scope conditions: The analysis covers 48 U.S. states from 1946 to 1954, ending at the 1954 IRS tax code change that expanded commercial insurers&amp;rsquo; market share. The enrollment data capture Blue Shield (physician-run) plans specifically; the paper explicitly notes that commercial insurer granular data are unavailable for the main Campaign period. The authors argue that multiple subsequent factors — middle-class acquisition of private coverage reducing demand for a public option, incumbent interests defending the status quo, and the persistent ideological linkage of private insurance with freedom — help explain why NHI was not adopted in subsequent decades, though these persistence mechanisms are outside the paper&amp;rsquo;s direct empirical scope.&lt;/p&gt;
&lt;p&gt;Q: What was the AMA&amp;rsquo;s Campaign, and what prompted it?
A: In response to Harry Truman&amp;rsquo;s unexpected 1948 presidential victory alongside a Democratic Congress — and with a majority of informed voters favoring NHI — the AMA hired Whitaker &amp;amp; Baxter&amp;rsquo;s Campaigns, Inc. to run the National Education Campaign (NEC). The Campaign had two components: physician outreach (pamphlet distribution to patients, liaison to civic organizations) and mass newspaper advertising. The AMA paid Whitaker &amp;amp; Baxter approximately $1.2 million per year in current terms, and coordinated an additional $19 million in 1950 dollars (roughly $240 million today) in tie-in advertising from allied corporations and trade groups.&lt;/p&gt;
&lt;p&gt;Q: How is Campaign exposure measured, and how is it validated as conditionally exogenous?
A: Campaign exposure combines two standardized components: per capita pamphlets distributed by AMA physicians (pamphlet quantity from W&amp;amp;B archives scaled by state AMA membership share) and per capita advertising circulation scaled by local newspaper readership (share of adults with more than five years of schooling). The two components are summed and standardized. Exogeneity is supported by balance tables showing no pre-Campaign correlation between exposure and enrollment or Gallup opinion, by the absence of discontinuous changes in income or unionization at Campaign onset, and by the historical fact that Campaign logistics relied on pre-existing networks assembled hastily in response to Truman&amp;rsquo;s unanticipated victory.&lt;/p&gt;
&lt;p&gt;Q: What is the main effect of the Campaign on private health insurance enrollment?
A: A one standard deviation increase in Campaign exposure is associated with a two percentage point increase in the share enrolled in PHI in the preferred specification (Column 4 of Table 1, which includes income, unionization, state fixed effects, and year fixed effects; coefficient 0.020, se 0.007, significant at 1%). This accounts for approximately 20% of the overall post-Campaign increase in PHI enrollment, corresponding to roughly 14 million new enrollees. The pre-Campaign coefficient is not statistically significant (coefficient 0.004, se 0.005), and the F-test p-value for pre-trends is 0.958.&lt;/p&gt;
&lt;p&gt;Q: What is the effect of the Campaign on public opinion toward NHI?
A: Using Gallup survey data, a one standard deviation increase in Campaign exposure led to an approximately six percentage point decline in individual support for NHI legislation per survey wave, against a pre-Campaign approval level of 69% that was trending upward. The F-test p-value for pre-trends in the Gallup event study is 0.179. This six-point effect is approximately equal to the gap in NHI support between union and non-union households, and approximately one-third the racial gap in support.&lt;/p&gt;
&lt;p&gt;Q: What evidence links the Campaign to civic organizations and the legislative process?
A: The Campaign&amp;rsquo;s archives document all civic organizations &amp;ldquo;on record against compulsory health insurance,&amp;rdquo; meaning they had passed resolutions in favor of PHI. The authors find a positive relationship between Campaign intensity and civic organizations passing such resolutions at the county level. Resolutions sent to elected officials were traced to the Congressional Record and to physical folders in the National Archives; their semantic similarity to AMA-WB propaganda is confirmed. Republican legislators&amp;rsquo; speech in the 81st Congress shows increased similarity to Campaign language in proportion to Campaign intensity in their district or state, while Democrat legislators do not show this pattern. NHI and the AMA experienced spikes in mention frequency in the Congressional Record during this period.&lt;/p&gt;
&lt;p&gt;Q: Did the Campaign affect physician political behavior beyond the clinic?
A: By 1952, when the Republican platform had fully adopted the AMA&amp;rsquo;s position, AMA members were approximately five times more likely to donate to the Eisenhower-Nixon ticket than non-AMA physicians, with donation probability increasing in Campaign intensity. The authors digitized the donor list from the National Professional Committee for Eisenhower (NPCE) — a separate lobbying entity created because the AMA legally could not endorse candidates — and linked approximately 80% of physician donors to the AMA Medical Directory.&lt;/p&gt;
&lt;p&gt;Q: What alternative explanations for PHI growth does the paper address, and how?
A: The standard literature attributes PHI growth to the 1942 Stabilization Act wage freeze (which left benefits unconstrained), collective bargaining rights clarified in the late 1940s, and the 1954 IRS tax exemption for employer-paid premiums. The authors include income per capita and unionization as core design controls and show that their Campaign exposure coefficient is stable across specifications with and without these controls (coefficients of 0.025 and 0.020 in Table 1 Columns 1–2 vs. 3–4, respectively). The analysis stops in 1954 before the tax change, and the authors note that by 1952 roughly 63% of households already had some form of medical expense insurance.&lt;/p&gt;
&lt;p&gt;Q: What is the conceptual mechanism through which the Campaign operated?
A: The authors adapt Sobbrio (2011)&amp;rsquo;s indirect lobbying model. Voters hold uniform priors over whether NHI enactment yields net positive or negative social surplus. The private-sector advocate (AMA-WB) sends messages that shift voters&amp;rsquo; posterior beliefs toward the negative-surplus state and, simultaneously, encourage PHI enrollment, which reduces voters&amp;rsquo; private valuation of a public option. Because citizens were likely unaware of the coordinated tie-in advertising across industries and the financial motivation behind physician messaging, the framing operated through naive belief updating. The public-sector advocate (Truman administration, Committee for the Nation&amp;rsquo;s Health) was vastly outresourced — the CNH raised only $104,000 in 1949 — and faced legal constraints on executive lobbying.&lt;/p&gt;
&lt;p&gt;Q: What advertising tactics specifically characterized the Campaign, and what do they imply about mechanisms?
A: Campaign pamphlets and ads provided little or no substantive information about insurance products (coverage, eligibility, cost) and instead tied health insurance to ideological symbols: &amp;ldquo;freedom,&amp;rdquo; &amp;ldquo;the American way,&amp;rdquo; &amp;ldquo;the voluntary way,&amp;rdquo; and warnings about &amp;ldquo;socialized medicine.&amp;rdquo; Word clouds from Campaign materials confirm &amp;ldquo;America&amp;rdquo; and &amp;ldquo;freedom&amp;rdquo; as dominant terms. The authors connect this to behavioral models of advertising (Mullainathan, Schwartzstein and Shleifer 2008) whereby advertisers create or exploit associations to influence product beliefs. The absence of informational content is consistent with effects operating through ideology and identity rather than rational product evaluation.&lt;/p&gt;
&lt;p&gt;Q: What explains why the U.S. did not adopt NHI in subsequent decades after the immediate Campaign period?
A: The authors offer three mechanisms (discussed outside their main empirical scope): First, as middle-class Americans obtained PHI through employers, demand for a public option diminished — the model formalizes this as reduced private valuation of NHI. Second, incumbents who benefit from the private status quo — Blue Cross Blue Shield, AMA, American Hospital Association, and pharmaceutical companies, which today comprise four of the top ten direct federal lobbyists — actively work to maintain it (Acemoglu, Egorov and Sonin 2021). Third, the Campaign&amp;rsquo;s ideological framing proved durable: ideologically similar rhetoric opposing &amp;ldquo;socialized medicine&amp;rdquo; appeared in campaigns against both Clinton-era and Obama-era reform efforts, and has been linked to increased adverse selection and preventable deaths (Bursztyn et al. 2022; Galvani et al. 2022).&lt;/p&gt;
&lt;p&gt;Q: What are the paper&amp;rsquo;s main contributions to the literature?
A: The paper provides the first causal evidence on the AMA&amp;rsquo;s political role in blocking NHI at the post-WWII juncture, contributing to the economic history of U.S. social insurance development. It contributes to the advertising literature by providing credible estimates of a sustained national campaign combining trusted field agents (physicians) with mass media, and to the lobbying literature by documenting indirect lobbying — persuasion of ordinary citizens — as a distinct and effective tool alongside direct lobbying. It also documents physician behavior outside the clinical setting, showing how rents from supply-side constraints were deployed to shape the market for medical services.&lt;/p&gt;
&lt;p&gt;Indirect lobbying: In the paper&amp;rsquo;s usage, persuasion of ordinary citizens via campaigns — as distinct from direct lobbying of policymakers — used to shift median voter beliefs and behavior to achieve legislative goals. Whitaker &amp;amp; Baxter are credited with creating this field through their work at Campaigns, Inc.&lt;/p&gt;
&lt;p&gt;Campaign exposure: The paper&amp;rsquo;s composite treatment variable, constructed as the sum of two standardized components: per capita pamphlets distributed by AMA physicians (physician outreach) and per capita advertising circulation scaled by local newspaper readership (mass communications), then re-standardized to mean 0, standard deviation 1.&lt;/p&gt;
&lt;p&gt;Tie-in advertising: Coordinated newspaper advertisements by third-party corporations and trade associations placed simultaneously with the main AMA-WB Campaign ad, sharing the &amp;ldquo;Voluntary Way is the American Way&amp;rdquo; slogan. Approximately 60% of newspapers with a main Campaign ad also had tie-in ads, averaging three per issue; third-party spending totaled approximately $19 million in 1950 dollars (~$240 million current).&lt;/p&gt;
&lt;p&gt;Voluntary (private) health insurance: In the paper&amp;rsquo;s framing, the AMA-promoted alternative to NHI — prepaid medical service plans run by state medical societies (Blue Shield) or nonprofit hospitals (Blue Cross) — deliberately labeled &amp;ldquo;voluntary&amp;rdquo; to contrast with &amp;ldquo;compulsory&amp;rdquo; NHI, embedding the product within an ideological frame of free choice.&lt;/p&gt;
&lt;p&gt;National Education Campaign (NEC): The AMA&amp;rsquo;s official name for the anti-NHI campaign directed by Whitaker &amp;amp; Baxter starting in 1949, characterized as &amp;ldquo;educational&amp;rdquo; to provide legal cover; the name itself illustrates the indirect lobbying strategy of framing political advocacy as public information.&lt;/p&gt;
&lt;p&gt;Source text origin / abstract-only block: Not a paper-defined concept; excluded.&lt;/p&gt;
&lt;p&gt;Naive voter updating: The paper&amp;rsquo;s modeling assumption (drawn from Sobbrio 2011) that voters held uniform priors on health insurance policy outcomes and updated beliefs via Bayesian message receipt, without awareness of coordination across industries or the financial motivation of physician messengers — making the ideological framing effective.&lt;/p&gt;
&lt;p&gt;Physician field agents: In the Campaign&amp;rsquo;s design, AMA member physicians served as credible, trusted intermediaries who distributed pamphlets to patients and solicited civic organization resolutions, leveraging their social status to amplify the Campaign&amp;rsquo;s reach into communities where mass advertising alone would be insufficient.&lt;/p&gt;</description></item></channel></rss>