<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Firm-Dynamics | Macro Paper Warehouse</title><link>https://macropaperwarehouse.com/topics/firm-dynamics/</link><atom:link href="https://macropaperwarehouse.com/topics/firm-dynamics/index.xml" rel="self" type="application/rss+xml"/><description>Firm-Dynamics</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><item><title>Anatomy of the Phillips Curve: Micro Evidence and Macro Implications</title><link>https://macropaperwarehouse.com/papers/anatomy-of-the-phillips-curve-micro-evidence-and-macro-implications/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/anatomy-of-the-phillips-curve-micro-evidence-and-macro-implications/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper addresses a fundamental puzzle in macroeconomics: why do estimates of the New Keynesian Phillips curve (NKPC) slope differ sharply depending on whether real marginal cost or the output gap is used as the real activity variable? The conventional, output gap-based NKPC yields very flat slope estimates (e.g., 0.006 to 0.024 in Hazell et al. 2022 and Rotemberg and Woodford 1997), which has led to the widespread view that the Phillips curve is &amp;ldquo;flat,&amp;rdquo; at least during the pre-pandemic period. The authors argue that this view conflates two distinct structural relationships: the elasticity of inflation with respect to real marginal cost, and the elasticity of marginal cost with respect to the output gap.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors assemble a unique quarterly micro-level dataset covering 4,598 manufacturing firms in Belgium over 84 quarters (1999:Q1–2019:Q4), totaling 132,915 observations. The dataset combines product-level domestic prices and quantities from the PRODCOM administrative database, customs data on foreign competitors&amp;rsquo; prices, and firms&amp;rsquo; variable production costs (labor costs from social security declarations plus intermediate input costs from VAT declarations). Intermediate inputs account for approximately 75 percent of total variable costs on average and are the most volatile cost component (within-firm coefficient of variation 1.77, versus 0.77 for labor costs).&lt;/p&gt;
&lt;p&gt;Their estimation strategy follows a &amp;ldquo;bottom-up&amp;rdquo; approach. Starting from a theoretical framework with heterogeneous firms subject to Calvo (1983) nominal rigidities and strategic complementarities in price setting (imperfect competition including dynamic oligopoly and Kimball demand), they derive a forward-looking dynamic pass-through regression linking a firm&amp;rsquo;s current price to discounted present values of its own marginal costs and competitors&amp;rsquo; prices, plus a lagged price level that serves as an error-correction term. This is Model A; robustness variants include Model B (absorbing competitor prices via industry-by-time fixed effects), Model C (imposing an AR(1) process for marginal cost), and Model A-U (unrestricted lagged-price coefficient).&lt;/p&gt;
&lt;p&gt;The structural parameters governing the NKPC slope — the degree of nominal rigidity (θ) and the strength of strategic complementarities (Ω) — are estimated jointly via GMM. Instruments for marginal cost are four-quarter-lagged firm-level total factor productivity (TFPQ), and instruments for competitors&amp;rsquo; prices exploit variation in EU-area export prices to third-country destinations and bilateral exchange rates between non-EU competitor currencies and the Euro. Sector-by-time fixed effects and firm fixed effects absorb confounding trends, shifting trend inflation, and permanent markups.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings with Quantitative Magnitudes&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The baseline estimate (Model A) yields θ = 0.711 (SE 0.014), implying that prices remain fixed for approximately three to four quarters on average, consistent with Nakamura and Steinsson (2008) Belgian PPI data (0.72). The strategic complementarity parameter is Ω = 0.570 (SE 0.059), indicating that competitor price dynamics reduce the pass-through of own marginal cost shocks by approximately half relative to the no-complementarities benchmark.&lt;/p&gt;
&lt;p&gt;These structural estimates imply a slope of the marginal cost-based NKPC of λ = 0.052 (SE 0.007), tightly estimated and robust across specifications: λ = 0.077 in Model B, λ = 0.069 in Model C, and λ = 0.056 in the unrestricted Model A-U. This slope is two to ten times larger than existing estimates of the conventional output gap-based NKPC slope (κ ≈ 0.024, Rotemberg and Woodford 1997; κ ≈ 0.006, Hazell et al. 2022).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reconciling the High Cost-Based Slope with the Flat Output-Based Slope&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper shows that the output-based slope κ equals the product of the cost-based slope λ and the output elasticity of marginal cost σ_y: κ = λ · σ_y. Using Bartik-style instruments based on high-frequency ECB monetary policy surprises interacted with industry-level sensitivities, the authors estimate σ_y using two models. Model D yields σ_y = 0.406 and κ = 0.021; Model E (directly regressing changes in marginal cost on changes in output) yields σ_y = 0.112 and κ = 0.006. These estimates are consistent with, and overlap with, Rotemberg and Woodford (1997) and Hazell et al. (2022) during the pre-pandemic sample period. The low elasticity of marginal cost to output is attributed to near-constant short-run returns to scale at the firm level and wage rigidity that mutes general equilibrium effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Aggregate Inflation Dynamics&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Feeding an aggregate marginal cost index (constructed as a Törnqvist-weighted average of firm-level marginal costs) into the model-implied inflation expression produces a series that tracks Belgian manufacturing PPI inflation well: marginal cost fluctuations alone account for approximately 70 percent of inflation variation (R² = 0.68, correlation 0.8), without appealing to unobservable cost-push shocks or inflation lags.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model Validation via Supply Shocks&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;A validation exercise using identified oil shocks (Känzig 2021 — measured as unexpected OPEC-day movements in oil futures prices) confirms the model. A one-standard-deviation shock to oil prices (a 15.7 percent increase in Brent crude) raises firms&amp;rsquo; real marginal costs by approximately 1.5 to 3 percent within the first three quarters, before reverting. The price response peaks at approximately 3 percent after six quarters, consistent with nominal rigidities generating a delayed but persistent response. Impulse-response matching yields λ_IRF = 0.042 (SE 0.005), within the confidence bands of the micro-level estimate λ = 0.052, validating the bottom-up approach.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;All estimates are drawn from Belgian manufacturing firms over 1999–2019, a period of moderate inflation during which Calvo pricing provides a good approximation of firm behavior. The authors note that the elasticity of marginal cost to output may be time-varying and nonlinear, and that during large aggregate shocks (such as the post-pandemic inflation surge), both the frequency of price adjustment and the sensitivity of marginal cost to output can rise substantially, requiring state-dependent pricing models (addressed in a companion paper, Gagliardone et al. 2025).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-primitive-formulation-of-the-nkpc-and-how-does-it-differ-from-the-conventional-formulation"&gt;Q1. What is the primitive formulation of the NKPC, and how does it differ from the conventional formulation?&lt;/h3&gt;
&lt;p&gt;A1: The primitive NKPC features real marginal cost (in log-deviation from its steady state) as the real activity variable: π_t = λ·mc_t + β·E_t{π_{t+1}} + u_t, where λ is the slope depending on nominal rigidities and strategic complementarities. The conventional formulation uses the output gap (or unemployment gap) as a proxy for marginal cost, which is valid only under specific conditions including perfectly flexible wages. When those conditions fail, the output gap is a poor proxy for marginal cost, typically leading to downward bias in slope estimates. Even when a proportionality holds, the output-based slope κ equals λ multiplied by σ_y (the output elasticity of marginal cost), so the two slopes carry different economic content.&lt;/p&gt;
&lt;h3 id="q2-what-structural-parameters-govern-the-slope-of-the-cost-based-nkpc-and-what-is-the-formula"&gt;Q2. What structural parameters govern the slope of the cost-based NKPC, and what is the formula?&lt;/h3&gt;
&lt;p&gt;A2: The slope is λ = &lt;a href="1%e2%88%92%ce%a9"&gt;(1−θ)(1−βθ)/θ&lt;/a&gt;, where θ is the Calvo probability of price non-adjustment (capturing nominal rigidity) and Ω = Γ/(1+Γ) is the strategic complementarities parameter derived from the markup elasticity Γ with respect to relative prices. High nominal rigidity (high θ) flattens the slope by making individual price adjustments less frequent; strong strategic complementarities (high Ω) flatten it further because firms mute their price response to marginal cost in order to avoid deviating from competitors. The discount factor β is calibrated at 0.99 for quarterly data.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-dynamic-pass-through-regression-differ-from-the-static-long-run-pass-through-regressions-used-in-prior-literature"&gt;Q3. How does the dynamic pass-through regression differ from the static (long-run) pass-through regressions used in prior literature?&lt;/h3&gt;
&lt;p&gt;A3: The dynamic pass-through regression (Model A) includes the firm&amp;rsquo;s lagged price as a regressor, which functions as an error-correction term controlling for persistent deviations between the price and the optimal reset price. Failing to include this term with quarterly data leads to omitted variable bias of magnitude −θ·Var(Δp_ft), since the cointegration error is autocorrelated with coefficient θ. Static pass-through regressions (as in Amiti, Itskhoki and Konings 2019 using annual data) are appropriate only when nominal rigidities can be ignored (θ ≈ 0); with quarterly data and θ ≈ 0.711, the orthogonality condition of the static model fails and the dynamic framework is necessary.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-baseline-estimates-of-the-structural-parameters-and-how-robust-are-they"&gt;Q4. What are the baseline estimates of the structural parameters, and how robust are they?&lt;/h3&gt;
&lt;p&gt;A4: The baseline Model A yields θ = 0.711 (SE 0.014) and Ω = 0.570 (SE 0.059), implying prices fixed for approximately three to four quarters and competitor-price influence roughly equal to own marginal cost influence. The implied NKPC slope is λ = 0.052 (SE 0.007). Robustness checks across six specifications (Models B, C, A-U, variable SR-RTS controls, Translog TFPQ, eight-quarter-lagged instrument) yield λ in the range 0.044 to 0.077, with all estimates statistically significant and within each other&amp;rsquo;s confidence bands. The unrestricted model (A-U) cannot reject the restriction Ϛ = θ on the lagged-price coefficient (p-value 0.90).&lt;/p&gt;
&lt;h3 id="q5-what-is-the-short-run-elasticity-of-a-firms-own-price-to-a-permanent-marginal-cost-shock-and-how-do-nominal-rigidities-and-strategic-complementarities-each-contribute"&gt;Q5. What is the short-run elasticity of a firm&amp;rsquo;s own price to a permanent marginal cost shock, and how do nominal rigidities and strategic complementarities each contribute?&lt;/h3&gt;
&lt;p&gt;A5: The short-run pass-through elasticity is (1−Ω)(1−θ) ≈ (1−0.570)(1−0.711) ≈ 0.125. This is substantially below one because both forces dampen price adjustment: nominal rigidity (1−θ ≈ 0.289) means most firms cannot adjust in any given quarter, and strategic complementarities (1−Ω ≈ 0.430) mean that adjusting firms reduce their pass-through to avoid deviating from competitors&amp;rsquo; prices. Without strategic complementarities (Ω = 0), the elasticity would be roughly 0.289; without nominal rigidities (θ = 0), it would be roughly 0.430; both together produce the observed 0.125.&lt;/p&gt;
&lt;h3 id="q6-how-is-marginal-cost-measured-in-the-data-and-why-is-the-inclusion-of-intermediate-input-costs-important"&gt;Q6. How is marginal cost measured in the data, and why is the inclusion of intermediate input costs important?&lt;/h3&gt;
&lt;p&gt;A6: Marginal cost is proxied by average variable cost per unit of output: the log-nominal marginal cost equals ln(TVC_ft/Y_ft) + ln(1+ν_ft), where TVC is the sum of intermediate input costs (from VAT declarations) and labor costs (wage bill from social security declarations), and Y_ft is a quantity index. Intermediate inputs account for approximately 75 percent of total variable costs on average and are the most volatile component (within-firm coefficient of variation 1.77 vs 0.77 for labor). The authors note that DSGE models typically feature only labor as a variable input, but accounting for intermediates is pivotal because intermediate goods price shocks were among the most important drivers of the post-pandemic inflation surge.&lt;/p&gt;
&lt;h3 id="q7-what-instruments-are-used-for-marginal-cost-and-competitors-prices-and-what-are-the-identifying-assumptions"&gt;Q7. What instruments are used for marginal cost and competitors&amp;rsquo; prices, and what are the identifying assumptions?&lt;/h3&gt;
&lt;p&gt;A7: The instrument for marginal cost is the four-quarter lagged firm-level TFPQ (physical total factor productivity), estimated as the residual from a gross-output production function. Its relevance depends on TFP persistence (confirmed); the exclusion restriction requires that persistent TFP variation is orthogonal to current and future demand shocks after removing permanent demand components (via firm fixed effects) and industry trends (via sector-by-time fixed effects). Two instruments for competitors&amp;rsquo; prices exploit international trade variation: (i) sales-weighted average export prices of EU-area competitors to non-Belgium, non-EU destinations (orthogonal to Belgian demand shocks by construction), and (ii) bilateral exchange rate movements between non-EU competitor currencies and the Euro. All instruments pass the Cragg-Donald and Kleibergen-Paap F-statistics (strongly rejecting weak instruments) and Hansen-Sargan over-identification tests (failing to reject validity).&lt;/p&gt;
&lt;h3 id="q8-what-evidence-supports-the-validity-of-the-tfpq-instrument-against-capacity-utilization-concerns"&gt;Q8. What evidence supports the validity of the TFPQ instrument against capacity utilization concerns?&lt;/h3&gt;
&lt;p&gt;A8: The authors run two empirical tests. First, regressing marginal cost on four-quarter-lagged capacity utilization yields a small, statistically insignificant elasticity (0.011, SE 0.052), suggesting the TFPQ instrument&amp;rsquo;s predictive power does not reflect capacity utilization variation. Second, re-estimating with &amp;ldquo;purified&amp;rdquo; TFPQ instruments adjusted for capital utilization (Column 4) and for both capital and labor utilization (Column 5) produces parameter estimates and NKPC slopes essentially unchanged from baseline. Additionally, regression residuals show only weak and short-lived autocorrelation (−0.09 at one-quarter lag, p=0.09; −0.01 at two-quarter lag, p=0.69), indicating demand shocks are highly transitory after conditioning on fixed effects.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-model-track-aggregate-belgian-manufacturing-ppi-inflation-and-what-does-this-imply-for-cost-push-shocks"&gt;Q9. How does the model track aggregate Belgian manufacturing PPI inflation, and what does this imply for cost-push shocks?&lt;/h3&gt;
&lt;p&gt;A9: Using the reduced-form expression π_t = λ̃(mc_t^n − p_{t-1}) + α + θu_t, where the reduced-form slope λ̃ = 0.22 is evaluated at baseline structural estimates, the model produces a model-implied inflation series that accounts for approximately 70 percent of variation in manufacturing PPI inflation (R² = 0.68, correlation 0.8), without including inflation lags or cost-push shocks. The model captures the inflation drop during the 2008 financial crisis, the run-up in 2016, and the subsequent decline. This contrasts with the quantitative DSGE literature in which cost-push shocks (variation in desired price and wage markups) account for approximately 70 percent of inflation volatility (e.g., Primiceri, Schaumburg and Tambalotti 2006).&lt;/p&gt;
&lt;h3 id="q10-how-do-the-authors-estimate-the-output-elasticity-of-marginal-cost-σ_y-and-what-do-they-find"&gt;Q10. How do the authors estimate the output elasticity of marginal cost σ_y, and what do they find?&lt;/h3&gt;
&lt;p&gt;A10: They use two approaches. Model D is a pricing equation directly relating firm-level prices and nominal output (value added), estimated via GMM, instrumented with Bartik-style shifters based on high-frequency ECB monetary policy surprises (Altavilla et al. 2019) interacted with industry-level sensitivities. Model E directly regresses changes in nominal marginal cost on changes in nominal output, also instrumented. Model D yields σ_y = 0.406 (SE 0.099) and implied κ = 0.021 (SE 0.005); Model E yields σ_y = 0.112 (SE 0.026) and κ = 0.006 (SE 0.001). The low σ_y is consistent with near-constant short-run returns to scale at the firm level and wage rigidity muting general equilibrium labor-market feedback, at least during the moderate-inflation pre-pandemic period.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-oil-shock-validation-exercise-confirm-the-cost-based-nkpc-slope-estimate"&gt;Q11. How does the oil shock validation exercise confirm the cost-based NKPC slope estimate?&lt;/h3&gt;
&lt;p&gt;A11: Following Känzig (2021), the authors identify oil shocks as unexpected movements in Brent crude oil futures around OPEC meeting days, normalizing to a one-standard-deviation shock (15.7 percent Brent increase). Local linear projection IRFs show that firms&amp;rsquo; real marginal costs rise 1.5 to 3 percent within three quarters and then revert, while prices peak at approximately 3 percent increase after six quarters (consistent with nominal rigidity delaying the price response). Impulse-response matching — minimizing the weighted distance between empirical and model-implied price IRFs — yields λ_IRF = 0.042 (SE 0.005), which is close to and within the confidence bands of the micro-level estimate λ = 0.052, validating the bottom-up estimation approach.&lt;/p&gt;
&lt;h3 id="q12-what-do-the-estimates-imply-about-why-the-conventional-nkpc-appears-flat-in-normal-times"&gt;Q12. What do the estimates imply about why the conventional NKPC appears flat in normal times?&lt;/h3&gt;
&lt;p&gt;A12: The flat conventional NKPC slope (κ ≈ 0.006–0.024) does not reflect limited transmission of marginal cost fluctuations to inflation — that transmission is high (λ ≈ 0.052–0.077). Rather, flatness reflects a weak link between the output gap and marginal cost during the pre-pandemic period (σ_y ≈ 0.112–0.406), attributable to near-constant short-run returns to scale in production and wage rigidity. This decomposition matters for policy: supply shocks that directly raise marginal cost will pass through strongly to inflation even when output does not move much, whereas demand shocks that operate through the output-cost channel face attenuated transmission.&lt;/p&gt;
&lt;h3 id="q13-under-what-conditions-does-the-cost-based-phillips-curve-decompose-cleanly-into-a-product-of-the-two-elasticities"&gt;Q13. Under what conditions does the cost-based Phillips curve decompose cleanly into a product of the two elasticities?&lt;/h3&gt;
&lt;p&gt;A13: The decomposition κ = λ · σ_y requires assuming that real wages are flexible and determined in general equilibrium at the industry level, with real wages increasing in industry output with elasticity σ_w; that the natural level of output is defined as the equilibrium under flexible prices and constant desired markups; and that the firm&amp;rsquo;s marginal product of labor depends on productivity and output with a common short-run returns-to-scale parameter ν (homogeneous across firms and time-invariant). Under these assumptions (which parallel those used to derive the conventional NKPC in the standard NK model), the output elasticity of marginal cost is σ_y = σ_w + ν, and the theoretical restriction κ = λ · σ_y holds exactly.&lt;/p&gt;
&lt;h3 id="q14-how-do-macroeconomic-complementarities-from-aggregate-decreasing-returns-to-scale-affect-the-nkpc-slope"&gt;Q14. How do macroeconomic complementarities from aggregate decreasing returns to scale affect the NKPC slope?&lt;/h3&gt;
&lt;p&gt;A14: If aggregate SR-RTS fall below unity, the NKPC slope formula gains an additional term Θ = 1/(1+γν(1−Ω)) &amp;lt; 1, where ν is inversely related to average SR-RTS and γ is the within-industry elasticity of substitution. However, empirical estimates of sectoral SR-RTS range from 0.93 to 0.98, with an aggregate estimate of approximately 0.965 (implying ν ≈ 0.036). Given this and calibrating γ = 4, Θ ≈ 0.941, so macroeconomic complementarities would reduce the NKPC slope by only about 6 percent — well within the confidence bounds of the baseline estimates. The authors conclude that the constant-returns assumption in their main framework is a good approximation.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Primitive (cost-based) NKPC slope (λ):&lt;/strong&gt; The coefficient linking inflation to real marginal cost in the underlying New Keynesian pricing equation, defined as λ = &lt;a href="1%e2%88%92%ce%a9"&gt;(1−θ)(1−βθ)/θ&lt;/a&gt;. It captures how strongly firms&amp;rsquo; aggregate price setting responds to movements in real marginal cost per unit of output, holding the discount factor, nominal rigidity, and strategic complementarities fixed. Estimated at 0.052 (tightly, range 0.044–0.077 across specifications) for Belgian manufacturing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Calvo probability of price non-adjustment (θ):&lt;/strong&gt; The parameter from Calvo (1983) staggered price setting capturing the share of firms that cannot change their price in a given period, equal to one minus the per-period probability of price adjustment. In this paper, θ is estimated directly from the dynamic pass-through regression coefficient on lagged prices, yielding θ ≈ 0.711, implying prices fixed approximately three to four quarters on average.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Strategic complementarities parameter (Ω):&lt;/strong&gt; Defined as Ω = Γ/(1+Γ), where Γ is the elasticity of a firm&amp;rsquo;s desired markup with respect to its own relative price. Captures the extent to which a firm weights competitors&amp;rsquo; prices (rather than its own marginal cost) when resetting its price. High Ω means firms strongly mute price responses to own cost changes to avoid relative price deviations from competitors. Estimated at Ω ≈ 0.570, implying competitor prices and own marginal cost enter the reset price with roughly equal weight.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic pass-through regression:&lt;/strong&gt; A forward-looking pricing equation (Model A) relating observed firm prices to the discounted present values of own marginal costs and competitors&amp;rsquo; prices, plus lagged own price as an error-correction term. The structural parameters θ and Ω are identified jointly from the regression coefficients, using GMM with instruments for the present values. The dynamic specification is necessary at quarterly frequency because the error-correction term (omitted in static pass-through models) is non-negligible when θ &amp;gt; 0.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Output elasticity of marginal cost (σ_y):&lt;/strong&gt; The elasticity of firm-level real marginal cost with respect to the firm-level output gap, defined under the assumptions that real wages are flexible and industry-level, equal to σ_y = σ_w + ν (wage elasticity with respect to industry output plus the short-run returns-to-scale parameter). This parameter bridges the cost-based and output-based Phillips curve slopes via κ = λ · σ_y. Estimated from micro data using monetary policy shock instruments at σ_y ≈ 0.112–0.406 in the pre-pandemic period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Short-run returns to scale (SR-RTS):&lt;/strong&gt; The extent to which a firm&amp;rsquo;s marginal cost rises with output scale in the short run, parameterized by ν in the cost function MC^n_ft = C_{it} · A_{ft} · Y_ft^ν. If ν = 0, marginal cost is independent of output scale (constant returns), which the authors assume in their baseline. Firm- and sector-level estimates from Translog production functions yield SR-RTS ≈ 0.93–0.98 across sectors (aggregate ≈ 0.965), broadly consistent with the constant-returns assumption and implying modest macroeconomic complementarities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reduced-form aggregate pass-through slope (λ̃):&lt;/strong&gt; A composite parameter capturing the contemporaneous pass-through of aggregate real marginal cost (defined as nominal marginal cost relative to the lagged price level) into quarterly inflation under the assumption that nominal marginal cost follows a random walk. Evaluated at θ ≈ 0.70 and Ω ≈ 0.52 (median across models), λ̃ = 0.22. This is distinct from the structural NKPC slope λ because it also captures the persistence of cost shocks.&lt;/p&gt;</description></item><item><title>Auctions with Frictions: Recruitment, Entry, and Limited Commitment</title><link>https://macropaperwarehouse.com/papers/auctions-with-frictions-recruitment-entry-and-limited-commitment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/auctions-with-frictions-recruitment-entry-and-limited-commitment/</guid><description>&lt;p&gt;This paper develops an auction model that jointly incorporates three frictions pervading informal price-formation processes: (1) costly recruitment by the seller, (2) costly participation by bidders, and (3) the seller&amp;rsquo;s inability to commit to a recruitment level or reserve price. The authors argue these frictions are especially prevalent in markets for idiosyncratic assets such as mergers and acquisitions, real estate, and home repair contracting, where auction houses like Christie&amp;rsquo;s and Sotheby&amp;rsquo;s command fees of 20–30% of revenues precisely because they reduce the underlying inefficiencies.&lt;/p&gt;
&lt;p&gt;The model features a single seller who exerts recruitment effort gamma at cost gamma*s, generating a Poisson-distributed number of contacted bidders with mean gamma. Each contacted bidder independently decides whether to pay entry cost c &amp;gt; 0 to learn their private value and participate in a first-price auction (FPA). The seller cannot commit to gamma (which is unobservable to bidders) or to a reserve price. Two scenarios are analyzed: PO (participation-observable, where bidders observe the number of entrants before bidding) and PU (participation-unobservable).&lt;/p&gt;
&lt;p&gt;The central tension is between the seller&amp;rsquo;s incentive to recruit more bidders to intensify competition and raise revenue, and bidders&amp;rsquo; rational concern that excessive recruitment makes entry unprofitable. Because the seller cannot commit, this tension generates several novel inefficiency results.&lt;/p&gt;
&lt;p&gt;In the PO scenario, the seller&amp;rsquo;s marginal revenue from recruitment Ro&amp;rsquo;(lambda) is single-peaked, meaning there is a minimum profitable participation scale lambda_o below which the seller will never recruit. Combined with a maximum participation level lambda-bar_c above which bidders will not enter (defined by U(lambda-bar_c) = c, where U is the bidder&amp;rsquo;s expected payoff), no-trade equilibrium is the unique outcome whenever lambda-bar_c &amp;lt; lambda_o — even for arbitrarily small recruitment cost s. This result holds because with unobservable effort, bidders correctly anticipate the seller will target participation above lambda_o, making entry unprofitable. When lambda-bar_c &amp;gt; lambda_o, three regimes arise: (i) no trade if s exceeds a threshold s-bar_o; (ii) an interior equilibrium with full entry (q* = 1) and lambda* = lambda_o(s) for intermediate s; and (iii) for small s, an equilibrium with lambda* = lambda-bar_c and partial entry q* = Ro&amp;rsquo;(lambda-bar_c)/s &amp;lt; 1. In regime (iii), total recruitment cost lambda*(s/q*) equals the constant lambda-bar_c * Ro&amp;rsquo;(lambda-bar_c) regardless of s — so even as s approaches zero, wasteful recruitment costs do not vanish, because they are determined by incentive constraints rather than by technology.&lt;/p&gt;
&lt;p&gt;In the PU scenario, a no-trade equilibrium always exists for all parameter values, because the seller cannot credibly disclose participation, creating self-reinforcing expectations of zero competition. The seller&amp;rsquo;s recruitment incentive xi(lambda) is strictly weaker than Ro&amp;rsquo;(lambda) in the PO scenario (proven via revenue equivalence: Ro&amp;rsquo;(lambda) = xi(lambda) + a positive term reflecting how greater participation induces more aggressive bidding). This yields ranking reversals: for intermediate s and small c, the PO scenario dominates PU; but for small s or large c, the PU scenario&amp;rsquo;s weaker recruitment incentive reduces wasteful over-recruitment, making PU preferable. These comparisons translate directly to a comparison of FPA and SPA with unobservable participation: the two formats are not equivalent in the presence of recruitment and entry frictions because they generate different recruitment incentives.&lt;/p&gt;
&lt;p&gt;A sampling-curse mechanism drives near-complete market unraveling when sellers have privately known recruitment costs drawn from a continuous uniform distribution on [0, s_o]. Because low-cost sellers recruit more, a contacted bidder believes the seller is more likely to have low costs — and hence to have recruited many other bidders — making entry unprofitable. Proposition 3 establishes a threshold c-hat such that for c in (c-hat, c-bar), as the lower bound of the cost distribution approaches zero, the fraction of seller types that remain inactive approaches one — near-complete unraveling — even though each type would be active if its cost were commonly known.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s main modeling innovation relative to the existing literature?
A: The paper&amp;rsquo;s central novelty is combining all three frictions — costly recruitment by the seller, costly participation by bidders, and limited seller commitment — in one model. The existing literature had studied entry and recruitment separately; Szech (2011) examined costly recruitment with costless entry; McAfee and McMillan (1987) and Levin and Smith (1994) studied costly entry with an exogenously given number of potential bidders; Milgrom (1987) and McAfee and Vincent (1997) studied limited commitment to a reserve price with a fixed bidder set. None combine all three.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;minimum profitable scale&amp;rdquo; result and why does it arise?
A: Because the seller cannot commit to a reserve price, the first few bidders are complementary — they stimulate competitive bidding, causing the seller&amp;rsquo;s marginal revenue Ro&amp;rsquo;(lambda) to be initially increasing, then decreasing (single-peaked). This means the seller&amp;rsquo;s profit Pi_o(lambda, q) is maximized either at zero or at a participation level above a minimum scale lambda_o, defined by Ro&amp;rsquo;(lambda_o) = s-bar_o. The seller will never choose a participation level between 0 and lambda_o.&lt;/p&gt;
&lt;p&gt;Q: Under what conditions does the market completely shut down in the PO scenario?
A: No-trade is the unique equilibrium outcome whenever lambda-bar_c &amp;lt; lambda_o, where lambda-bar_c is defined by U(lambda-bar_c) = c (the participation break-even level) and lambda_o is the seller&amp;rsquo;s minimum profitable scale. This condition arises when entry costs c are large enough relative to the competitive dynamics. Importantly, no trade occurs for every recruitment cost s &amp;gt; 0, including arbitrarily small s — commitment failure alone can cause complete market breakdown even when recruiting bidders is nearly costless.&lt;/p&gt;
&lt;p&gt;Q: What is the inefficiency in regime (iii) of Proposition 2 (small s, PO scenario)?
A: When s &amp;lt; Ro&amp;rsquo;(lambda-bar_c), equilibrium has lambda* = lambda-bar_c and q* = Ro&amp;rsquo;(lambda-bar_c)/s &amp;lt; 1. The total recruitment cost is lambda* * (s/q*) = lambda-bar_c * Ro&amp;rsquo;(lambda-bar_c), a strictly positive constant independent of s. As s approaches zero, total recruitment effort and its cost do not vanish — they are pinned by incentive constraints. This waste could be avoided if the seller could commit to an effort level below lambda-bar_c, illustrating that commitment failure creates persistent inefficiency even when the technology of recruitment is inexpensive.&lt;/p&gt;
&lt;p&gt;Q: Why does a no-trade equilibrium always exist in the PU scenario but not always in the PO scenario?
A: In the PU scenario, if bidders expect zero participation, they bid zero conditional on being contacted; the seller then has no incentive to recruit, validating the expectation. This equilibrium is self-sustaining for all parameter values (Claim 2). In the PO scenario, the equilibrium refinement (requiring that off-path beliefs not support negative seller payoff at lambda = 0 when trade equilibria exist) rules out no-trade equilibria when lambda-bar_c &amp;gt; lambda_o and s is not too large; specifically, Proposition 2 shows that no-trade equilibrium is unique only when s &amp;gt; s-bar_o or lambda-bar_c &amp;lt; lambda_o.&lt;/p&gt;
&lt;p&gt;Q: What drives the ranking reversal between PO and PU scenarios?
A: The core result is Claim 3: Ro&amp;rsquo;(lambda) &amp;gt; xi(lambda) for all lambda &amp;gt; 0, meaning the marginal incentive to recruit is strictly stronger under PO than PU. This follows from revenue equivalence: Ro&amp;rsquo;(lambda) = xi(lambda) + (d/d lambda-hat) Ru(lambda, beta_{lambda-hat})|_{lambda-hat=lambda}, and the second term is strictly positive because greater expected participation induces more aggressive bidding. For intermediate s and small c, stronger PO recruitment incentives support higher participation and revenue. For small s or large c, those same stronger incentives generate wasteful over-recruitment in PO, making PU preferable.&lt;/p&gt;
&lt;p&gt;Q: How does the paper connect its PO/PU comparison to a comparison of first- and second-price auctions?
A: In any standard auction where the highest-value bidder wins, payoff and revenue equivalence imply that the bidder payoff function U(lambda) and seller revenue Ro(lambda) are identical. In particular, the dominant-strategy equilibrium of the SPA (where bidders bid their true values regardless of participation) generates the same outcomes as the PO equilibrium, because with truthful bidding the observability of participation is irrelevant. Therefore, comparing PO and PU with an FPA is equivalent to comparing the SPA and FPA with unobservable participation. The two formats are not revenue-equivalent when recruitment and entry frictions are present: their ranking depends on s and c in exactly the way described for PO vs. PU.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;sampling curse&amp;rdquo; and how does it cause market unraveling?
A: The sampling curse arises when sellers have privately known recruitment costs. Because a lower-cost seller optimally recruits more bidders, the probability of any given bidder being contacted is higher when the seller has a lower cost. Conditional on being contacted, a bidder therefore believes the seller more likely has a low cost and thus has recruited many competitors, reducing the value of entry. In the binary-type case (Claim 8), if sL is sufficiently small relative to sH, the low-cost seller must recruit so many bidders that entry becomes unattractive; the resulting low q* makes the marginal recruitment cost sH/q* prohibitively high for the high-cost type, driving it out (lambda*_H = 0).&lt;/p&gt;
&lt;p&gt;Q: What does Proposition 3 establish about near-complete unraveling with a continuum of seller types?
A: With seller costs uniformly distributed on [s-bar, s_o], Proposition 3 establishes a threshold c-hat strictly between 0 and c-bar such that: (i) for c in (c-hat, c-bar), as s-bar approaches zero, the fraction of seller types with zero recruitment approaches one — near-complete market unraveling; (ii) for c &amp;lt; c-hat, all seller types remain active regardless of how small s-bar is. This is striking because for any commonly known s in (0, s_o), the PO scenario supports trade for all c &amp;lt; c-bar; unraveling arises purely from the interaction of private cost information and the sampling curse, not from any type&amp;rsquo;s cost being intrinsically too high.&lt;/p&gt;
&lt;p&gt;Q: What does the welfare analysis say about equilibrium efficiency?
A: The welfare-maximizing participation level lambda_w satisfies U(lambda_w) = c + s (equating the marginal bidder&amp;rsquo;s surplus to the full social cost of one more participant), with full entry q_w = 1. In equilibrium under PO, q* &amp;lt; 1 in some cases (wasted recruitment) and lambda* differs from lambda_w for almost all (s, c) pairs — both excessive participation (lambda* &amp;gt; lambda_w) and deficient participation (lambda* &amp;lt; lambda_w) can arise. Full efficiency requires Ro&amp;rsquo;(lambda*) = s and U(lambda*) = s + c simultaneously, but since both U and Ro&amp;rsquo; are independent of s and c as parameters, these equalities generically fail.&lt;/p&gt;
&lt;p&gt;Q: Does the seller benefit from being able to commit to recruitment effort?
A: Claim 10 shows that with observable effort in the PO scenario, the seller commits to gamma-hat = min{lambda-bar_c, lambda_o(s)} when lambda-bar_c &amp;gt;= lambda_o, and to lambda-bar_c (if profitable) when lambda-bar_c &amp;lt; lambda_o. Commitment strictly improves the seller&amp;rsquo;s profit whenever gamma-hat = lambda-bar_c: it enables positive trade when lambda-bar_c &amp;lt; lambda_o and Ro(lambda-bar_c) &amp;gt; lambda-bar_c * s (otherwise impossible without commitment), and it saves recruitment costs when lambda-bar_c &amp;gt; lambda_o and Ro&amp;rsquo;(lambda-bar_c) &amp;gt; s. However, the commitment outcome is always welfare-inefficient: lambda-bar_c &amp;gt; lambda_w whenever s &amp;gt; 0.&lt;/p&gt;
&lt;p&gt;Q: What anecdotal evidence do the authors cite for the model&amp;rsquo;s relevance?
A: Subramanian (2010) and Boone and Mulherin (2004, 2009) show that the majority of merger and acquisition auctions are &amp;ldquo;informal&amp;rdquo; — mixtures of auctions and negotiations rather than structured processes with rules laid out in advance — and that sellers are typically unable to credibly commit to participation levels. Milgrom (2003) states from consulting experience that marketing an auction is often more critical than clever mechanism design. Fees of 20–30% of revenues paid to intermediaries like Christie&amp;rsquo;s and Sotheby&amp;rsquo;s are offered as quantitative evidence of the magnitude of the inefficiencies that such intermediaries reduce. Home repair contracting is cited as a familiar informal-auction setting where both recruitment and entry costs are material.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Recruitment effort (gamma):&lt;/strong&gt; The seller&amp;rsquo;s costly action of contacting potential bidders, modeled as a Poisson process with mean gamma at cost gamma*s; unobservable to bidders in the baseline model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Participation-observable (PO) vs. participation-unobservable (PU) scenarios:&lt;/strong&gt; The two variants of the model; in PO, bidders observe the total number of entrants n before bidding; in PU, they do not observe n and the seller cannot credibly disclose it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Minimum profitable scale (lambda_o):&lt;/strong&gt; The smallest positive participation level the seller will ever choose in equilibrium, defined as the value where Ro&amp;rsquo;(lambda_o) equals the peak of the average revenue curve s-bar_o. The seller always recruits either zero bidders or at least lambda_o, due to the initial complementarity of bidders (they stimulate each other&amp;rsquo;s bids) under no-commitment-to-reserve-price.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Break-even participation level (lambda-bar_c):&lt;/strong&gt; The maximum participation level at which a bidder&amp;rsquo;s expected gross payoff U(lambda) equals the entry cost c; bidders will not enter if they expect participation above lambda-bar_c.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sampling curse:&lt;/strong&gt; The adverse-selection mechanism arising when sellers have privately known recruitment costs: because low-cost sellers recruit more, a contacted bidder infers the seller is more likely to have a low cost and thus to have recruited many competitors, making entry less attractive and potentially driving higher-cost seller types out of the market.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;xi(lambda):&lt;/strong&gt; The seller&amp;rsquo;s marginal revenue with respect to recruitment in the PU scenario, defined as the total derivative of Ru(lambda, beta_{lambda-hat}) evaluated where actual and expected participation coincide (lambda-hat = lambda). Strictly less than Ro&amp;rsquo;(lambda) for all lambda &amp;gt; 0, reflecting that in PU the seller loses the ability to leverage bidder aggression via observable competition.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wasteful recruitment:&lt;/strong&gt; The equilibrium phenomenon in which total recruitment cost lambda*(s/q*) remains at the positive constant lambda-bar_c * Ro&amp;rsquo;(lambda-bar_c) even as s approaches zero, because incentive constraints — not technology — pin the equilibrium effort level.&lt;/p&gt;</description></item><item><title>Bottom-Up Markup Fluctuations</title><link>https://macropaperwarehouse.com/papers/bottom-up-markup-fluctuations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/bottom-up-markup-fluctuations/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper asks how firm-level, sector-level, and aggregate markups comove with output at different levels of aggregation, and whether a single structural model can reconcile seemingly contradictory empirical findings about markup cyclicality that arise when researchers use different aggregation schemes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors build a granular macroeconomic model featuring oligopolistic competition with a nested constant-elasticity-of-substitution (CES) demand structure following Atkeson and Burstein (2008). The economy contains N sectors, each with a discrete number of firms competing under Cournot oligopoly with flexible prices. Firm-level markups are endogenously increasing in within-sector market shares: under Cournot, the sectoral markup is a simple function of the sector&amp;rsquo;s Herfindahl-Hirschman index (HHI), and the aggregate markup is a function of the expenditure-share-weighted average of sectoral HHIs. Firm-level productivity follows a discretized random growth (Gibrat&amp;rsquo;s law) process as in Carvalho and Grassi (2019), generating fat-tailed firm-size distributions and granular aggregate fluctuations. The baseline calibration features only idiosyncratic firm-level productivity shocks and abstracts from aggregate shocks, because—in the model—aggregate shocks that move all firms proportionately do not affect relative market shares and hence do not affect markups.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The empirical analysis uses French administrative firm-level data from the FICUS-FARE datasets covering the universe of French firms from 1994 to 2019, yielding approximately 9.38 million firm-year observations across 26 years, 22 two-digit sectors, and 275 five-digit NAF sectors. Firm-level markups are estimated following De Loecker and Warzynski (2012) using a translog production function estimated by GMM (following De Ridder et al. 2024) on a subsample of approximately 220,733 firm-year observations where physical output quantity is available from the Enquete Annuelle de Production survey (2009-2019). Using quantity rather than revenue as the output measure avoids the measurement biases documented in Bond et al. (2021).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings and Quantitative Magnitudes&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Markup-market-share relationship (firm level):&lt;/strong&gt; Regressions of the change in the inverse firm markup on the change in firm market share yield a negative and significant coefficient of approximately -0.268 to -0.293 (depending on fixed-effect specification), consistent with the model prediction that markups rise with market share. Sector-level analogues yield a slope of the change in inverse sector markup on the change in sector HHI of approximately -0.37, which is simultaneously a calibration target (implying sigma = 1.8 given epsilon = 5) and an empirical moment the model closely matches (model counterpart: -0.36).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Within-between decomposition of sector markup changes:&lt;/strong&gt; In the model under Cournot competition, changes in firm-level markups (the &amp;ldquo;within&amp;rdquo; term) account for exactly 50% of changes in sector-level markups, with between-firm reallocation accounting for the other 50%. In the French data, for the median sector, the within term accounts for 59% of changes in sector markups (interquartile range across sectors: 34%-81%).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Firm-level markup cyclicality with sector output (heterogeneous by size):&lt;/strong&gt; The average firm&amp;rsquo;s markup is countercyclical with respect to own-sector output (beta_1 approximately -0.073 in levels specification), but this relationship reverses for large firms: firms with market shares roughly above 10% (top 0.1% of the market-share distribution) have procyclical markups (interaction coefficient beta_2 approximately 0.574 in levels). The model qualitatively and roughly quantitatively reproduces this heterogeneity.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Sector-level markup cyclicality with sector output (procyclical):&lt;/strong&gt; Following Nekarda and Ramey (2013), sector markup changes comove positively and significantly with sector output changes: estimated coefficient of 0.160 (standard error 0.040) in first-differences. The calibrated model yields a median coefficient of 0.139 (std dev 0.057 across 5,000 simulated 25-year samples), close to the data. Consistently, sector concentration (HHI) is also procyclical with sector output (estimated coefficient 0.332, std error 0.067 in first-differences).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Sector-level markup cyclicality with aggregate output (acyclical to weakly countercyclical):&lt;/strong&gt; Following Bils et al. (2018), the comovement between sector markups and aggregate output is fragile in sign and significance: the French data yields a point estimate of -0.239 (std error 0.116) in first-differences, marginally significant (t-stat 2.06) and with sign sensitive to detrending method. The model without aggregate shocks predicts positive comovement (median coefficient 0.165) that is not statistically different from zero across samples. Adding aggregate productivity shocks (calibrated to match French aggregate output volatility) brings the model-implied coefficient close to zero (median 0.008), with 20-30% of 25-year simulated samples displaying countercyclical sectoral markups relative to GDP—consistent with the ambiguity in the data.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Aggregate output volatility:&lt;/strong&gt; The baseline calibration with only granular firm-level shocks generates a standard deviation of detrended aggregate output of 0.83%, equal to 26% of the 3.16% observed in the French data. (The comparable granular ratio from Carvalho and Grassi 2019 for a perfectly competitive US model is 30%.) Variable markups dampen granular aggregate volatility: the standard deviation of aggregate output under variable markups is 0.87 times that under heterogeneous-but-constant markups (95% CI: 0.82-0.97), because incomplete pass-through reduces the effective weight of large firms in the price index.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Aggregate markup volatility:&lt;/strong&gt; In the data, the relative standard deviation of aggregate markup to aggregate output is 0.40-0.50 (depending on detrending). The model generates a relative volatility of 0.36 (median across samples). The correlation between aggregate markup and output in the data is at most 0.06; the model without aggregate shocks implies a counterfactually large median correlation of 0.91, which falls to 0.27 when aggregate TFP shocks are superimposed (with 16% of 25-year samples displaying countercyclical aggregate markups).&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Results pertain to French private-sector firms (including formerly government-owned firms, most of which privatized during the sample period) across manufacturing and some non-manufacturing sectors at the national-market level. The analysis abstracts from import competition (market shares are computed relative to all French firms in the sector), local geographic markets (relevant for non-tradeable goods where national-level shares understate local concentration), and multi-product firm structure. Findings are for a flexible-price model driven by idiosyncratic productivity shocks; the paper explicitly discusses how nominal rigidities would further strengthen procyclicality at the sector level.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-central-mechanism-by-which-granular-firm-level-shocks-generate-markup-cyclicality"&gt;Q1. What is the central mechanism by which granular firm-level shocks generate markup cyclicality?&lt;/h3&gt;
&lt;p&gt;A: Because markups are endogenously increasing in within-sector market shares under oligopolistic competition, a firm that receives a positive productivity shock gains market share and therefore raises its markup, while its competitors lose market share and lower their markups. The net effect on the sectoral markup depends on the shocked firm&amp;rsquo;s initial size: a positive shock to a sufficiently large firm (above a threshold market share) raises the sectoral markup, while a positive shock to a small firm lowers it. Since sectoral expansions in a granular economy are disproportionately driven by large firms, sector output and sector markup tend to comove positively in the medium run.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-sign-of-markup-cyclicality-differ-depending-on-the-level-of-aggregation"&gt;Q2. Why does the sign of markup cyclicality differ depending on the level of aggregation?&lt;/h3&gt;
&lt;p&gt;A: Sector-level markups react only to within-sector idiosyncratic shocks, so sectors that happen to be driven by large-firm booms display positive comovement between sector markup and sector output. However, a given sector&amp;rsquo;s markup is uncorrelated with aggregate output movements coming from other sectors. In small samples (such as 25-year windows), whether a sector&amp;rsquo;s markup comoves positively or negatively with aggregate output depends on whether the sector happens to lead or lag the aggregate cycle. Over sufficiently long samples, the model implies positive comovement of sector markups with aggregate output, but in finite samples the relationship is indeterminate. This asymmetry across aggregation levels explains why researchers using different reduced-form specifications in the same dataset can reach opposing conclusions about procyclicality versus countercyclicality.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-within-between-decomposition-of-sectoral-markup-changes-and-what-does-it-imply-quantitatively"&gt;Q3. What is the within-between decomposition of sectoral markup changes and what does it imply quantitatively?&lt;/h3&gt;
&lt;p&gt;A: Changes in the inverse sectoral markup can be decomposed into (i) a within term—changes in firm-level markups holding market shares fixed—and (ii) a between term—changes in market shares holding firm-level markups fixed. Under Cournot competition, the within and between terms are analytically equal in every period, so each accounts for exactly 50% of the change in sectoral markups; this 50-50 split holds globally (not only to first order). In the French data, for the median sector, within-firm markup changes account for 59% of sector markup changes (interquartile range across sectors: 34%-81%), close to but slightly above the model&amp;rsquo;s 50% prediction.&lt;/p&gt;
&lt;h3 id="q4-how-do-variable-markups-affect-granular-aggregate-output-volatility-relative-to-a-model-with-constant-markups"&gt;Q4. How do variable markups affect granular aggregate output volatility relative to a model with constant markups?&lt;/h3&gt;
&lt;p&gt;A: Variable markups (endogenous pass-through that is decreasing in firm size) reduce granular aggregate output volatility relative to a model where markups are heterogeneous but fixed. The intuition is that larger firms have lower pass-through rates, so their productivity shocks translate into smaller price changes and therefore smaller output responses than they would under constant markups—effectively reducing the weight of large firms in the aggregate price index in a way similar to a decline in market concentration. Quantitatively, using first-order approximations around equilibrium distributions from the calibrated model, the standard deviation of aggregate output under variable markups is 0.87 times that under heterogeneous-but-constant markups (95% confidence interval: 0.82-0.97). The overall standard deviation under variable and heterogeneous markups is only 1.02 times that under homogeneous and constant markups (95% CI: 0.99-1.14), meaning markup heterogeneity and variability together have limited net effects on aggregate output volatility.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-model-predict-for-firm-level-markup-cyclicality-and-how-heterogeneous-is-this-across-firm-size"&gt;Q5. What does the model predict for firm-level markup cyclicality, and how heterogeneous is this across firm size?&lt;/h3&gt;
&lt;p&gt;A: Proposition 4 states that, in the asymptotic limit, firm-level markups comove positively with own-sector output for firms with market shares above a threshold, and negatively for firms below it. This occurs because large firms have a disproportionate impact on sector-level price and output (when the product of market share and pass-through rate is increasing in size), so large-firm shocks simultaneously drive sector expansions and raise large-firm markups while compressing small-firm markups. In the French data, the average firm&amp;rsquo;s markup is countercyclical with respect to sector output (beta_1 approximately -0.073 in log-levels with firm and year fixed effects), but firms with market shares above roughly 10% (top 0.1% of the distribution, since the average market share is only 0.07%) display procyclical markups (interaction coefficient beta_2 approximately 0.574). The model reproduces this qualitative pattern and the order of magnitude of these estimates.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-calibrate-the-key-demand-elasticities-and-what-are-the-resulting-pass-through-implications"&gt;Q6. How does the paper calibrate the key demand elasticities, and what are the resulting pass-through implications?&lt;/h3&gt;
&lt;p&gt;A: The within-sector substitution elasticity is set to epsilon = 5, a standard value. The cross-sector substitution elasticity sigma is calibrated to match the slope of the inverse sector markup on sector HHI in first-differences. The empirical slope is -0.37; under the model, the slope equals -(epsilon/sigma - 1)/(epsilon - 1), and given epsilon = 5, sigma = 1.8 delivers a model counterpart of -0.36. These parameter values imply own-cost pass-through rates that are decreasing in firm size; for large firms (with market share &amp;gt;= 57%, approximately the top 0.004% of the distribution), the implied pass-through rate is 0.63, within the confidence intervals reported in Amiti, Itskhoki, and Konings (2019) for large Belgian firms.&lt;/p&gt;
&lt;h3 id="q7-why-do-aggregate-productivity-shocks-not-affect-markups-in-the-model-and-what-are-the-implications-for-aggregate-markup-cyclicality"&gt;Q7. Why do aggregate productivity shocks not affect markups in the model, and what are the implications for aggregate markup cyclicality?&lt;/h3&gt;
&lt;p&gt;A: In the model, firm-level markups are functions of within-sector market shares, not the level of productivity. An aggregate shock that shifts all firms&amp;rsquo; productivity proportionately leaves relative market shares unchanged and therefore leaves all markups unchanged. This means aggregate shocks increase aggregate output volatility but leave markup volatility unchanged, reducing the correlation between aggregate markup and aggregate output. When aggregate TFP shocks are added to match French aggregate output volatility, the model-implied median correlation between aggregate markup and output falls from 0.91 (without aggregate shocks) to 0.27 (with aggregate shocks), while 16% of 25-year simulated samples display countercyclical aggregate markups—more consistent with the weak and fragile empirical relationship.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-paper-address-the-potential-measurement-error-bias-in-the-negative-correlation-between-markups-and-marginal-costs"&gt;Q8. How does the paper address the potential measurement-error bias in the negative correlation between markups and marginal costs?&lt;/h3&gt;
&lt;p&gt;A: Since marginal cost is computed as price divided by estimated markup, regressing market shares or markups on marginal costs risks spurious correlation via measurement error in the markup (which appears in both sides). The authors address this concern by constructing an instrumental variable for marginal cost based on firm-specific energy intensity interacted with energy price changes, following Ganapati, Shapiro, and Walker (2020). Table A10 confirms that instrumenting for marginal cost yields negative effects on both markup and market share with larger point estimates than the OLS specifications in Table 4, validating the baseline findings.&lt;/p&gt;
&lt;h3 id="q9-is-the-50-50-within-between-decomposition-of-sectoral-markup-changes-robust-to-the-choice-of-competition-mode"&gt;Q9. Is the 50-50 within-between decomposition of sectoral markup changes robust to the choice of competition mode?&lt;/h3&gt;
&lt;p&gt;A: No. The exact 50-50 split of within and between terms in sectoral markup changes is a specific property of Cournot competition and holds globally (not just as a first-order approximation). Under Bertrand competition, the within and between terms are generally not equal to each other. The paper derives analytic results under both competition modes and focuses on Cournot for quantitative work because it generates more markup variation and better matches the estimated pass-through rates and markup-size relationship.&lt;/p&gt;
&lt;h3 id="q10-what-do-model-simulations-imply-for-the-magnitude-and-cyclicality-of-aggregate-markups-versus-the-data-and-what-is-the-role-of-variable-versus-constant-markups"&gt;Q10. What do model simulations imply for the magnitude and cyclicality of aggregate markups versus the data, and what is the role of variable versus constant markups?&lt;/h3&gt;
&lt;p&gt;A: In the data (detrended), the standard deviation of aggregate markup is 1.27% with a relative volatility (to output) of 0.40 and a correlation with output of 0.03. The baseline model with only granular shocks yields a median markup standard deviation of 0.30%, relative volatility of 0.36, and correlation with output of 0.91. The model with aggregate shocks added yields median markup standard deviation of 0.30%, relative volatility of 0.09, and correlation of 0.27. Counterfactually fixing markups at their initial heterogeneous levels while keeping the same market shares and shock variance yields aggregate markup standard deviation approximately 0.93 times the variable-markup value (standard deviation of markups under variable markups is 1.08 times that under constant markups, with a 95% CI of 1.00-1.18), and a correlation with output of 0.92 versus 0.87 under variable markups. Overall, the magnitude and cyclicality of aggregate markups are not substantially different between variable and constant-markup specifications.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-reconcile-its-findings-with-prior-literature-on-markup-cyclicality-bils-et-al-2018-vs-nekarda-and-ramey-2013"&gt;Q11. How does the paper reconcile its findings with prior literature on markup cyclicality (Bils et al. 2018 vs. Nekarda and Ramey 2013)?&lt;/h3&gt;
&lt;p&gt;A: Nekarda and Ramey (2013) find procyclical sector markups with respect to sector output in US data—a result replicated in French data (beta approximately 0.160). Bils, Klenow, and Malin (2018) find countercyclical sector markups with respect to aggregate output in US data. Both results can be generated simultaneously in the model: sector markups are positively correlated with own-sector output because granular booms in a sector are driven by large-firm expansions that raise sector markups; however, a given sector&amp;rsquo;s markup is weakly and ambiguously correlated with aggregate output because aggregate fluctuations reflect shocks across many sectors, only some of which are in the same sector. The model can therefore simultaneously predict procyclicality with respect to sector output and an acyclical-to-weakly-countercyclical relationship with aggregate output—explaining why both empirical findings can be correct.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-data-limitations-and-how-do-they-affect-the-interpretation-of-results"&gt;Q12. What are the data limitations and how do they affect the interpretation of results?&lt;/h3&gt;
&lt;p&gt;A: Three limitations are noted. First, market shares are computed relative to total revenue of all French firms in the sector without accounting for imports, so foreign competition is ignored and domestic concentration may be overestimated. Second, revenues are reported at the national level, so for non-tradeable goods (whose relevant market is local) the paper underestimates true local market concentration, attenuating the markup-concentration relationship in those sectors. Third, the model abstracts from entry and exit (the number of firms per sector is held fixed at sector-year averages), though Appendix D demonstrates robustness of main empirical results to restricting the sample to continuing firms.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Granular macroeconomic model:&lt;/strong&gt; A model in which the economy consists of a finite (large but discrete) number of firms, so that idiosyncratic firm-level shocks to large firms do not average out and instead generate aggregate fluctuations. In the paper&amp;rsquo;s usage, granularity means that sectoral and aggregate business-cycle fluctuations are driven primarily by shocks to the largest firms, which also have the highest markups and market shares.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nested CES demand structure (Atkeson-Burstein):&lt;/strong&gt; A two-level constant-elasticity-of-substitution aggregation where the final good aggregates N sectors with cross-sector elasticity sigma, and each sector aggregates the output of its Nk firms with within-sector elasticity epsilon &amp;gt; sigma. This structure generates firm-level markups that are endogenously increasing in within-sector market shares (under both Cournot and Bertrand competition) and yields closed-form expressions for sector-level markups as a function of sector HHI and aggregate markups as a function of the expenditure-weighted average of sector HHIs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Markup elasticity with respect to market share (Gamma_ki):&lt;/strong&gt; Under Cournot competition, the semi-elasticity of firm i&amp;rsquo;s log markup with respect to its log market share, equal to (epsilon/sigma - 1)s_ki / (epsilon/(epsilon-1) - (epsilon/sigma - 1)s_ki). This is strictly positive for epsilon &amp;gt; sigma and increasing in market share, implying that larger firms have markups that are more responsive to changes in their competitive position.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pass-through rate (alpha_ki):&lt;/strong&gt; The fraction of an idiosyncratic cost shock that is passed into the firm&amp;rsquo;s price relative to the sectoral price index, given by 1/(1 + (epsilon-1)Gamma_ki). Pass-through is decreasing in market share (larger firms have lower pass-through), which dampens their price response to own shocks and mutes the impact of large-firm shocks on aggregate price volatility—acting like a reduction in market concentration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Within-between decomposition of sector markup changes:&lt;/strong&gt; The change in inverse sector markup decomposed into (i) a within term measuring changes in firm-level markups holding market shares fixed, and (ii) a between term measuring reallocation of market shares across firms with heterogeneous markups. Under Cournot competition, these two terms are exactly equal (each 50%) for any firm-level shocks—a result that holds globally (not merely as a first-order approximation)—because the forces that increase the within term (higher markup sensitivity) also raise heterogeneity between firms (increasing the between term).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sectoral markup (mu_kt):&lt;/strong&gt; Defined as the ratio of sectoral revenues to total wage payments in the sector, equal to the harmonic mean of firm-level markups weighted by market shares. Under Cournot competition, this is a simple increasing function of the sector&amp;rsquo;s HHI: mu_kt = (epsilon/(epsilon-1))[1 - (epsilon/sigma - 1)/(epsilon-1) x HHI_kt]^(-1). This mapping between concentration and the markup price-cost wedge gives the central empirical prediction tested at the sector level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Markup cyclicality (at different aggregation levels):&lt;/strong&gt; The comovement between markups and output, which the paper distinguishes sharply across three levels: (i) firm markup vs. own-sector output—countercyclical for small firms, procyclical for large firms; (ii) sector markup vs. own-sector output—procyclical (positive covariance) under conditions proven in Proposition 3; (iii) sector markup vs. aggregate output—theoretically positive over long samples but ambiguous and close to zero in short samples, because aggregate output also reflects shocks to other sectors whose markups are uncorrelated with the focal sector&amp;rsquo;s markups. The paper&amp;rsquo;s central insight is that the same underlying model generates all three empirical patterns simultaneously.&lt;/p&gt;</description></item><item><title>Collusion with Optimal Information Disclosure</title><link>https://macropaperwarehouse.com/papers/collusion-with-optimal-information-disclosure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/collusion-with-optimal-information-disclosure/</guid><description>&lt;p&gt;This paper asks how a third-party intermediary (an &amp;ldquo;algorithm&amp;rdquo;) that observes market demand or costs superior to competing firms should optimally disclose that information to maximize the firms&amp;rsquo; collusive profit in a repeated Bertrand competition setting. The motivation is the rise of algorithmic pricing intermediaries such as RealPage in apartment rentals, A2i Systems in retail gasoline, and Rainmaker in hotel rooms, as well as offline cartel facilitators like AC-Treuhand.&lt;/p&gt;
&lt;p&gt;The model extends the canonical Rotemberg–Saloner (1986) repeated Bertrand framework with stochastic demand. The key technical assumption is that firm profit is affine in the unknown state s, so expected profit depends only on the expected state. This holds for binary states, linear demand with unknown intercept (D(p,s) = s − p), and linear demand with unknown per-unit cost. The algorithm observes s and commits to a known disclosure policy mapping s to a public signal. The solution concept is pure-strategy subgame-perfect equilibrium, and the paper solves for the disclosure policy and equilibrium that jointly maximize collusive profit.&lt;/p&gt;
&lt;p&gt;The main result (Theorem 1) is that the unique optimal disclosure policy is upper censorship: there is a cutoff ŝ such that demand states s &amp;lt; ŝ are disclosed and result in the corresponding monopoly price p^m(s), while demand states s ≥ ŝ are pooled — only the event {s ≥ ŝ} is disclosed — and result in the monopoly price for the mean concealed state, p^m(s*), where s* = E[s | s ≥ ŝ]. The reduction to a static information design problem (Lemma 1) is the key technical step: optimal collusive profit equals V*, the greatest fixed point of V = max_{G ∈ MPC(F)} E_G[min{π^m(s), δV/((1−δ)(n−1))}]. The &amp;ldquo;capped monopoly profit&amp;rdquo; min{π^m(s), π^max} is convex-then-concave in s, and classical results from the static information design literature (Kolotilin 2018; Dworczak and Martini 2019) then imply upper censorship is uniquely optimal.&lt;/p&gt;
&lt;p&gt;Two features of the optimal equilibrium are notable. First, prices are rigid (constant at p^m(s*)) whenever s ≥ ŝ — the opposite of Rotemberg–Saloner&amp;rsquo;s &amp;ldquo;price wars during booms.&amp;rdquo; The logic is that pooling high demand states with a lower average state is more profitable than cutting prices, because pooling reduces the current-period deviation gain without sacrificing as much on-path profit. Second, for demand states s ∈ (ŝ, s*), the equilibrium price p^m(s*) exceeds the monopoly price p^m(s) — supra-monopoly pricing occurs for a range of intermediate states. Monopoly pricing is attainable at each such state in isolation, but recommending the higher price p^m(s*) is necessary to make the pooling incentive-compatible at states s &amp;gt; s*.&lt;/p&gt;
&lt;p&gt;Comparing to full disclosure, Proposition 1 shows that optimal disclosure leads to strictly higher prices at every demand state, and hence unambiguously lower consumer surplus. Proposition 3 shows that improving the algorithm&amp;rsquo;s accuracy (a mean-preserving spread of F) reduces expected consumer surplus whenever consumer surplus under monopoly pricing is concave in s — a natural condition. This result is more pessimistic than prior work (Sugaya–Wolitzky 2018; Miklos-Thal–Tucker 2019), which found ambiguous effects because those papers assumed full disclosure.&lt;/p&gt;
&lt;p&gt;Comparative statics (Proposition 2): fewer firms or a higher discount factor δ increases collusive profit V* and makes prices more flexible (raises ŝ). Collusion is impossible if and only if δ &amp;lt; (n−1)/n, the same threshold as under full disclosure.&lt;/p&gt;
&lt;p&gt;Extensions maintain the core results. With Markov (persistent) demand (Section 4 / Theorem 2), upper censorship remains optimal but the cutoff ŝ(s) depends on last-period demand s: under positive serial correlation, ŝ(s) is decreasing in s, so the algorithm discloses less information following high demand. With differentiated products under a symmetric linear demand system (Section 5 / Theorem 3), the optimal policy censors an intermediate interval [ŝ_L, ŝ_H] and discloses both the lowest and highest demand states, because at high states the absence of an upper bound on equilibrium profit makes disclosure with price-cutting optimal.&lt;/p&gt;
&lt;p&gt;Q: What is the core research question and why is it policy-relevant?
A: The paper asks how an informed intermediary should optimally disclose demand or cost information to competing firms to maximize their collusive profit. It is directly motivated by antitrust cases against RealPage (sued by the US DOJ in August 2024), A2i Systems/Kalibrate, and Rainmaker, all of which gather market data from competing firms and recommend prices. The theory also applies to offline facilitators like AC-Treuhand, prosecuted by the European Commission for disclosing competitively sensitive information.&lt;/p&gt;
&lt;p&gt;Q: What is the affinity assumption and why does it matter?
A: The paper assumes that firm profit π(p, s) is affine (linearly increasing) in the demand or cost state s for each price p. This implies that expected profit for any distribution over states equals profit evaluated at the expected state: E[π(p,s)] = π(p, E[s]). As a consequence, any disclosure policy is equivalent, from a profit standpoint, to choosing a distribution G of the firms&amp;rsquo; posterior mean beliefs over s, and G must be a mean-preserving contraction of the prior F (by Blackwell 1953). The assumption is satisfied for binary states, linear demand with unknown intercept, and linear demand with unknown cost.&lt;/p&gt;
&lt;p&gt;Q: What is the key reduction result (Lemma 1) and what does it achieve?
A: Lemma 1 reduces the problem of finding an optimal repeated-game equilibrium to a static information design problem. Optimal collusive profit equals V*, the greatest fixed point of V = max_{G ∈ MPC(F)} E_G[min{π^m(s), δV/((1−δ)(n−1))}], and this is attained by a symmetric, stationary, grim-trigger equilibrium. The reduction works because, under Bertrand competition, static deviation gains are proportional to on-path payoffs, creating a one-to-one correspondence that allows the repeated-game constraint to be folded into a single-period objective.&lt;/p&gt;
&lt;p&gt;Q: Why is upper censorship the uniquely optimal disclosure policy?
A: The static information design problem has a &amp;ldquo;capped monopoly profit&amp;rdquo; objective: min{π^m(s), π^max}, where π^max = δV*/((1−δ)(n−1)) is the maximum per-period profit that satisfies incentive constraints. Because π^m(s) is convex (as the maximum of affine functions) and the cap π^max is constant, the overall objective is convex for s below the cap and constant (then concave) above it — i.e., convex-then-concave in s. Classical results for linear information design (Kolotilin 2018; Dworczak and Martini 2019) imply that the unique optimal policy for a convex-then-concave objective is upper censorship.&lt;/p&gt;
&lt;p&gt;Q: What is the supra-monopoly pricing result and why does it arise?
A: For demand states s ∈ (ŝ, s*), the equilibrium price is p^m(s*) &amp;gt; p^m(s), meaning firms charge above the monopoly price for the current state. This arises because the pooling policy must recommend a single price for all states s ≥ ŝ, and the recommended price is p^m(s*) where s* = E[s | s ≥ ŝ]. At intermediate states s ∈ (ŝ, s*), this price exceeds the local monopoly price. The algorithm accepts lower profit at these states because it is necessary to maintain the pooled recommendation at higher states where monopoly pricing would otherwise require a price cut.&lt;/p&gt;
&lt;p&gt;Q: How does optimal disclosure compare to full disclosure in terms of consumer surplus?
A: Proposition 1 shows that collusive prices under optimal disclosure are strictly higher at every demand state compared to full disclosure (Rotemberg–Saloner). In Rotemberg–Saloner, high demand states trigger price cuts (&amp;ldquo;price wars during booms&amp;rdquo;) to deter deviation; under optimal disclosure, high states are pooled and prices are instead rigid at p^m(s*). Because prices are higher at all states, consumer surplus is unambiguously lower under optimal disclosure.&lt;/p&gt;
&lt;p&gt;Q: What does Proposition 3 say about the effect of algorithmic accuracy on consumer surplus?
A: Proposition 3 states that if consumer surplus under monopoly pricing, CS(s), is concave in s, then a mean-preserving spread of F (i.e., improved algorithmic accuracy) reduces expected consumer surplus. This result is more pessimistic than prior work by Sugaya–Wolitzky (2018) and Miklos-Thal–Tucker (2019), which found ambiguous effects. The difference is that those papers assumed full disclosure, so better accuracy tightened incentive constraints and sometimes forced price cuts. Under optimal selective disclosure, a more accurate algorithm always raises average prices because the algorithm withholds information that would have forced price cuts.&lt;/p&gt;
&lt;p&gt;Q: What are the comparative statics with respect to the number of firms and the discount factor?
A: Proposition 2 establishes that a decrease in the number of firms n or an increase in the discount factor δ increases collusive profit V* and makes collusive prices more flexible (raises ŝ). The intuition for fewer firms making prices more flexible is that with fewer firms, incentive constraints bind for a narrower range of demand states, so less pooling is needed. Collusion is impossible if and only if δ &amp;lt; (n−1)/n, the same threshold as under full disclosure.&lt;/p&gt;
&lt;p&gt;Q: How does the model generate empirically testable predictions distinct from other collusion models?
A: The model predicts: (1) the equilibrium price distribution has support on an interval [p^m(s_bar), p^m(ŝ)] plus a single mass point at the higher price p^m(s*); (2) prices are pro-cyclical overall but rigidly fixed at p^m(s*) for all but the lowest demand states; (3) the gap p^m(s) − p(s) is non-monotone — zero at low states, negative (supra-monopoly) at intermediate states, and positive at high states; (4) prices are more flexible when firms are more patient or fewer. The rigid high price combined with a flexible interval of lower prices is described as a distinctive collusive marker not present in other models.&lt;/p&gt;
&lt;p&gt;Q: How does the model relate to the empirical literature testing Green–Porter versus Rotemberg–Saloner?
A: Rotemberg–Saloner predicts counter-cyclical prices (price wars during booms), while Green–Porter predicts pro-cyclical prices. Empirical tests (e.g., Porter 1983, Ellison 1994) have typically found pro-cyclical prices, favoring Green–Porter. The present model generates pro-cyclical prices through a different mechanism — perfect monitoring plus selectively disclosed demand information — showing that pro-cyclical prices are consistent with perfect monitoring when the information intermediary optimally pools high demand states. The paper suggests that distinguishing the theories requires estimating the gap between price and monopoly price over the cycle: under Green–Porter, collusion succeeds better in high demand states; under this model, collusion succeeds better in low demand states.&lt;/p&gt;
&lt;p&gt;Q: What narrative evidence from the RealPage case corroborates the model&amp;rsquo;s predictions?
A: The US DOJ complaint against RealPage states that &amp;ldquo;in down markets… [RealPage] instills pricing discipline in landlords, curbing normal fully independent competitive reactions by substituting them with interdependent decision-making,&amp;rdquo; and that RealPage advertised that its AI helps clients &amp;ldquo;avoid the race to the bottom in down markets.&amp;rdquo; This is consistent with the model&amp;rsquo;s prediction of flexible monopoly prices at low demand states and a rigid, supra-monopolistic price in normal times. The Kumatori Contractors Cooperative case (studied by Kawai, Nakabayashi, and Ortner 2024) corroborates the censorship result: that organization took drastic steps to limit bidders&amp;rsquo; information about costs on the largest projects — exactly the states where deviation is most tempting.&lt;/p&gt;
&lt;p&gt;Q: How do results change with persistent (Markov) demand?
A: Theorem 2 shows that upper censorship remains uniquely optimal with Markov demand, but the cutoff ŝ(s) now depends on last-period demand s. Under positive serial correlation, ŝ(s) is decreasing in s: the algorithm discloses less information after high demand because firms are more optimistic and thus more tempted to deviate. Under negative serial correlation, ŝ(s) is increasing. The optimal collusive price is no longer always equal to the monopoly price for the disclosed mean demand, and the expected price conditional on last-period demand can be countercyclical (similar to Rotemberg–Saloner), even though the current-period price is always monotone in current demand.&lt;/p&gt;
&lt;p&gt;Q: How does the optimal disclosure policy change with differentiated products?
A: With a symmetric linear demand system (Section 5, Theorem 3), the optimal policy censors an intermediate interval [ŝ_L, ŝ_H] and discloses both the lowest and the highest demand states. At high demand states s &amp;gt; ŝ_H, the algorithm discloses the state and recommends a price below monopoly (to satisfy incentive constraints), because with differentiated goods there is no upper bound on equilibrium profit and profit is convex in s at high states, making disclosure with price-cutting optimal. Mathematically, the capped monopoly profit is piecewise-convex rather than convex-then-concave, so the optimal policy is intermediate-interval censorship rather than upper censorship. The Appendix A version extends to general demand systems and capacity constraints with the same qualitative logic.&lt;/p&gt;
&lt;p&gt;Q: What are the main limitations and directions for future work acknowledged by the authors?
A: The paper identifies three main limitations. First, if profit is not affine in s (i.e., expected profit depends on more than the mean state), the information design problem becomes non-linear and upper censorship is typically suboptimal, though it remains approximately optimal when the problem is close to linear. Second, the model assumes the algorithm&amp;rsquo;s objective is to maximize industry profit; if the intermediary is a profit-maximizing seller of software (as in Harrington 2022), the objective may instead be to maximize the profit differential between adopters and non-adopters. Third, the model assumes all firms use the algorithm; allowing partial adoption would require modeling firms&amp;rsquo; incentives to subscribe. The paper notes that incorporating these considerations &amp;ldquo;could be an interesting direction for future research.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Upper Censorship (disclosure policy): A disclosure policy in which demand states below a cutoff ŝ are revealed to firms (along with the corresponding monopoly price recommendation), while states above ŝ are pooled — only the event {s ≥ ŝ} is disclosed — with a single monopoly price recommendation p^m(s*) for the mean concealed state s* = E[s | s ≥ ŝ]. This is the uniquely optimal disclosure policy in the baseline model.&lt;/p&gt;
&lt;p&gt;Capped Monopoly Profit: The per-period profit objective in the reduced static information design problem: min{π^m(s), π^max}, where π^max = δV*/((1−δ)(n−1)) is the maximum industry profit attainable in a single period without violating incentive constraints. This function is convex-then-concave in s, which drives the optimality of upper censorship.&lt;/p&gt;
&lt;p&gt;Supra-Monopoly Pricing: Equilibrium prices that exceed the monopoly price for the realized demand state. In the model, this occurs for states s ∈ (ŝ, s*), where the algorithm&amp;rsquo;s pooled recommendation p^m(s*) is above the local monopoly price p^m(s). It arises because the pooled recommendation must be incentive-compatible at the highest concealed states.&lt;/p&gt;
&lt;p&gt;Price Rigidity: The feature of the optimal equilibrium in which the collusive price is constant at p^m(s*) for all demand states s ≥ ŝ. The algorithm achieves this by withholding information about high demand states, preventing the &amp;ldquo;price wars during booms&amp;rdquo; predicted by Rotemberg–Saloner (1986) under full disclosure.&lt;/p&gt;
&lt;p&gt;Algorithmic Accuracy: In the paper&amp;rsquo;s terms, the informativeness of the algorithm&amp;rsquo;s signal about s, formalized as the precision of the distribution F. Improving accuracy corresponds to a mean-preserving spread of F (Blackwell 1953). A more accurate algorithm always increases collusive profit; under the concavity condition on consumer surplus, it also reduces expected consumer surplus.&lt;/p&gt;
&lt;p&gt;Mean-Preserving Contraction (MPC(F)): The set of distributions G of firms&amp;rsquo; posterior mean beliefs over s that are consistent with Bayesian updating of the prior F. By Blackwell (1953), a disclosure policy is feasible if and only if it induces a distribution G ∈ MPC(F). This is the feasibility constraint in the static information design problem.&lt;/p&gt;
&lt;p&gt;Affinity in the state: The assumption that π(p, s) is affine (linearly increasing) in s for each price p. This implies E[π(p,s)] = π(p, E[s]), so expected profit is determined entirely by the expected state, enabling the reduction of the disclosure problem to choosing a distribution of posterior means.&lt;/p&gt;</description></item><item><title>Comment on: Is it AI or data that drives market power?</title><link>https://macropaperwarehouse.com/papers/comment-on-is-it-ai-or-data-that-drives-market-power/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/comment-on-is-it-ai-or-data-that-drives-market-power/</guid><description>&lt;p&gt;This paper is a published comment by Miao Ben Zhang (USC Marshall School of Business) on Mihet, Rishabh, and Gomes (2025), &amp;ldquo;Is It AI or Data That Drives Market Power?&amp;rdquo; Zhang identifies three contributions of the commented paper and benchmarks each against the existing literature, offering targeted suggestions for strengthening the analysis.&lt;/p&gt;
&lt;p&gt;The first contribution Zhang discusses is the commented paper&amp;rsquo;s distinction between raw data, AI capability, and processed data. Raw data is modeled as a by-product of production linearly related to firm size; processed data is modeled as the abundance of signals improving the precision of firms&amp;rsquo; next-period productivity predictions. The commented paper&amp;rsquo;s key modeling innovation is a formula linking raw data (n_{i,t}), firm-level AI capability (z_i), and processed data (n_{i,t}-tilde): processed data equals a weighted sum of an information entropy effect — e^(-z_i) * (-n_{i,t} * ln(n_{i,t})) — and an AI capability effect — (1 - e^(-z_i)) * n_{i,t} * e^(n_{i,t}). Zhang notes this formula implies that the marginal value of raw data can turn negative for firms with low AI capability, consistent with information-theoretic constraints from the rational inattention literature (Sims, 2003). Zhang requests more empirical support for this equation, specifically asking whether low-AI firms exhibit lower TFP than high-AI firms at similar data-intensity levels, and encouraging discussion of existing measures of data-processing ability such as human capital in data engineering and ML pipeline automation.&lt;/p&gt;
&lt;p&gt;The second contribution is the commented paper&amp;rsquo;s modeling of a secondary market for trading processed data among firms. Zhang notes that facilitating processed data markets — for example via APIs or structured knowledge sharing — can, per the commented paper&amp;rsquo;s simulation and empirical analysis, democratize innovation and reduce market concentration, enabling even low-AI firms to compete. Zhang flags that the paper is silent on firm acquisition as an alternative channel for accessing processed data, arguing this omission is significant given that processed data, unlike ideas or technologies, is less portable and cannot be obtained simply by poaching skilled employees.&lt;/p&gt;
&lt;p&gt;The third contribution is the commented paper&amp;rsquo;s empirical strategy. The commented paper constructs firm-level proxies for AI intensity and data intensity, then exploits two exogenous technological shocks — the advent of AWS cloud computing and transformer-based architectures — to identify causal effects of improvements in compute and processed data accessibility. The evidence shows that compute improvements disproportionately benefit data-rich firms, while processed data access disproportionately benefits low-AI firms. The central empirical message is that access to raw data tends to foster market concentration, whereas access to processed data tends to reduce market concentration. Zhang raises a measurement concern: the commented paper relies on firm-level Herfindahl-Hirschman Index (HHI) calculations based on time-varying, text-based industry definitions (Hoberg and Phillips, 2016). Zhang argues a positive effect on this HHI could reflect either genuine firm growth relative to competitors or reclassification of the firm into different, possibly more concentrated, sectors — making the HHI measure alone insufficient to support claims about product market concentration. Zhang recommends complementing this with industry-level concentration measures anchored to fixed baseline industry codes (FIC codes from Hoberg and Phillips, 2016), constructed at the FIC-year level, following the approach of Gutierrez and Philippon (2017) on industries&amp;rsquo; growth and median Q.&lt;/p&gt;
&lt;p&gt;No quantitative magnitudes from regressions or calibrations are reported in the comment itself, as this is a discussion piece rather than an original empirical paper. All claims above are drawn directly from the text.&lt;/p&gt;
&lt;p&gt;Q: What are the three contributions of Mihet, Rishabh, and Gomes (2025) that Zhang identifies?
A: First, the paper explicitly models the distinct roles of raw data, AI capability, and processed data, linking the information entropy literature to firm production. Second, it models a secondary market for trading processed data among firms, relevant for policy on data sharing platforms. Third, it empirically tests the model&amp;rsquo;s predictions using firm-level proxies and two exogenous technological shocks.&lt;/p&gt;
&lt;p&gt;Q: What is the core formula linking raw data, AI capability, and processed data in the commented paper?
A: Processed data (n_{i,t}-tilde) equals e^(-z_i) * (-n_{i,t} * ln(n_{i,t})) plus (1 - e^(-z_i)) * n_{i,t} * e^(n_{i,t}), where z_i is firm-level AI capability and n_{i,t} is raw data. The first term captures the information entropy effect (which can reduce or negate the value of raw data for low-AI firms) and the second captures the AI capability effect (where AI turns raw data into abundant useful signals).&lt;/p&gt;
&lt;p&gt;Q: Why can the marginal value of raw data turn negative, according to the framework?
A: Information-theoretic constraints — long studied through concepts like Shannon entropy and Sims&amp;rsquo;s rational inattention — imply that unprocessed raw data may harm rather than help firms that lack adequate processing capabilities. Zhang situates this in the broader macro-finance literature on information choice (Sims, 2003; Veldkamp, 2011).&lt;/p&gt;
&lt;p&gt;Q: What empirical suggestion does Zhang make regarding the raw data versus AI capability distinction?
A: Zhang asks whether, in the commented paper&amp;rsquo;s sample of publicly-traded firms with measures of data intensity and AI intensity, low-AI firms exhibit lower TFP (following Imrohoroglu and Tuzel, 2012) than high-AI firms when controlling for similar levels of data intensity. Zhang also encourages discussion of anecdotal evidence for negative information entropy effects and of existing measures of data processing ability such as human capital in data engineering, annotation, cleaning, or ML pipeline automation (Abis and Veldkamp, 2024).&lt;/p&gt;
&lt;p&gt;Q: What is the policy relevance of the secondary market for processed data?
A: The commented paper&amp;rsquo;s simulation and empirical analysis shows that facilitating processed data markets (e.g., via APIs or structured knowledge sharing) can democratize innovation and reduce market concentration, enabling even low-AI firms to compete. This aligns with recent literature on secondary markets for structured data and foundation model outputs (Gans, 2018, 2024; Conti et al., 2023, 2024; Athey, 2019). Platforms may have incentives to restrict processed data access, potentially reinforcing incumbent power (Carballa Smichowski et al., 2023).&lt;/p&gt;
&lt;p&gt;Q: What channel does Zhang argue the commented paper neglects in its analysis of market concentration?
A: Zhang argues the paper is silent on firm acquisition as an alternative means by which firms access processed data, noting that processed data is less portable than ideas or technologies — it cannot be obtained simply by poaching a skilled employee. Zhang contends this acquisition channel appears central to the paper&amp;rsquo;s focus on market concentration and encourages the authors to include a discussion of it.&lt;/p&gt;
&lt;p&gt;Q: What is the central empirical finding of the commented paper regarding raw versus processed data and market concentration?
A: Access to raw data tends to foster market concentration, while access to processed data tends to reduce market concentration. The evidence shows that compute improvements (proxied by the AWS shock) disproportionately benefit data-rich firms, while processed data accessibility (proxied by the transformer architecture shock) disproportionately benefits low-AI firms, consistent with theoretical predictions.&lt;/p&gt;
&lt;p&gt;Q: What is Zhang&amp;rsquo;s specific concern about the HHI measure used in the commented paper?
A: The commented paper constructs firm-level HHI using time-varying, text-based industry definitions (Hoberg and Phillips, 2016). Zhang argues a positive effect on this HHI is ambiguous: it could reflect genuine firm growth relative to competitors or reclassification of the firm into different, possibly more concentrated, sectors. Zhang concludes that the HHI measure alone is not strong enough to support claims about product market concentration.&lt;/p&gt;
&lt;p&gt;Q: What robustness check does Zhang recommend for the empirical analysis?
A: Zhang recommends constructing industry-level concentration measures at the FIC-year level using fixed baseline FIC codes from Hoberg and Phillips (2016), available at the Hoberg-Phillips Data Library. The authors could then analyze how industries with high versus low average or median AI intensity and data intensity respond to the two technological shocks in terms of concentration. Zhang cites Gutierrez and Philippon (2017) as an example of this approach and notes it would help distinguish within-industry dynamics from shifts in firm business focus, aligning with best practices from De Loecker, Eeckhout, and Unger (2020) on persistent market power.&lt;/p&gt;
&lt;p&gt;Raw data: A by-product of firms&amp;rsquo; production, modeled as linearly related to firm size; represents unprocessed observations that have not yet been transformed into useful signals. Distinguished from processed data, which is what actually improves productivity predictions.&lt;/p&gt;
&lt;p&gt;Processed data: Modeled as the abundance of signals that improves the precision of firms&amp;rsquo; predictions of their next-period productivity (following Farboodi and Veldkamp, 2022). Unlike ideas or technologies, processed data is less portable and cannot easily be transferred by poaching skilled employees.&lt;/p&gt;
&lt;p&gt;AI capability (z_i): Firm-level ability to transform raw data into processed data. Firms with low AI capability may receive negative marginal value from additional raw data due to information entropy effects; firms with high AI capability extract large gains from the same raw data.&lt;/p&gt;
&lt;p&gt;Information entropy effect: The component of the raw-to-processed-data transformation — e^(-z_i) * (-n_{i,t} * ln(n_{i,t})) — that captures the information-theoretic cost of possessing raw data without adequate processing capability. At low AI capability, this effect can reduce or negate the precision of signals.&lt;/p&gt;
&lt;p&gt;Secondary market for processed data: A market in which firms trade processed data, modeled in the commented paper as a platform or API-based exchange. The commented paper&amp;rsquo;s analysis shows this market can democratize innovation and reduce market concentration by enabling low-AI firms to access processed data they cannot produce internally.&lt;/p&gt;
&lt;p&gt;Firm-level HHI (text-based): Herfindahl-Hirschman Index calculated using time-varying, text-based industry definitions (Hoberg and Phillips, 2016). Zhang identifies a measurement ambiguity: a positive effect on this measure could reflect genuine competitive gains or reclassification into more concentrated sectors.&lt;/p&gt;</description></item><item><title>Competition and the Phillips curve</title><link>https://macropaperwarehouse.com/papers/competition-and-the-phillips-curve/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/competition-and-the-phillips-curve/</guid><description>&lt;p&gt;Fujiwara and Matsuyama ask whether the well-documented flattening of the New Keynesian Phillips curve (NKPC) and the concurrent rise in market concentration and markup rates are causally linked or merely coincidental. Under the canonical New Keynesian model with CES demand, competition is irrelevant to the Phillips curve regardless of whether entry is endogenous — concentration neither changes its slope nor affects inflation directly. This paper overturns that irrelevance result by extending the canonical model in two directions: (1) incorporating endogenous firm entry and exit following Bilbiie, Ghironi, and Melitz (2008) and Bilbiie, Fujiwara, and Ghironi (2014), and (2) replacing CES with the Homothetic Single Aggregator (HSA) demand system (Matsuyama and Ushchev 2017, 2020b), a flexible, tractable class of homothetic demand systems that nests CES and Translog as special cases.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s theoretical results depend on two of Marshall&amp;rsquo;s laws of demand. The Second law states that the price elasticity of demand rises with the firm&amp;rsquo;s own price; the Third law states that the rate of increase in that elasticity falls with price. Together these conditions imply that the markup rate and pass-through rate are endogenous to the competitive environment.&lt;/p&gt;
&lt;p&gt;The main findings, delivered under both Rotemberg (1982) and Calvo (1983) pricing, are that higher entry costs — leading to market concentration — cause Phillips curve flattening through two distinct, complementary channels:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Structural (steady-state) effect.&lt;/strong&gt; Under Rotemberg pricing, the slope of the NKPC is proportional to the price elasticity zeta(z); market concentration reduces z, hence reduces zeta(z) under the Second law, directly flattening the curve. Under Calvo pricing, the slope is proportional to the pass-through rate rho(z); the Third law implies that concentration reduces rho(z), again flattening the curve. The Calvo–Rotemberg equivalence, which holds under CES to first order (Roberts 1995), breaks down under HSA: each pricing mechanism highlights a different channel.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Observational (omitted variable bias) effect.&lt;/strong&gt; Endogenous entry generates an endogenous cost-push shock through strategic complementarity in price setting. Because the number of firms N_t is omitted from a naive regression of inflation on real marginal cost, and because N_t is positively correlated with the marginal cost under the Second law, the omitted variable bias is negative — the estimated slope is biased downward. This bias is amplified with greater concentration under the Third law (Rotemberg case) and under both the Second and Third laws (Calvo case).&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Quantitatively, the paper simulates under three parametric HSA families — CES, Translog, and Co-PaTh (Constant Pass-Through). De Loecker, Eeckhout, and Unger (2020) document that aggregate markups rose from 21% above marginal cost to 61% — a rise of approximately 40 percentage points. The authors&amp;rsquo; simulations imply this increase corresponds to an entry cost roughly 3.5 times higher under Translog and roughly 2.5 times higher under Co-PaTh with pass-through rate rho = 0.5. Under these parameterizations, the accompanying market concentration can halve the slope of the NKPC. Impulse responses confirm that the responses of inflation to both technology shocks and monetary policy shocks become smaller as market concentration deepens.&lt;/p&gt;
&lt;p&gt;Scope conditions: results require departure from CES (the Second and/or Third law must hold); endogenous entry is necessary for the dynamic cost-push channel; the structural flattening requires only the Second law under Rotemberg but additionally the Third law under Calvo; the omitted variable bias requires the Second law under Rotemberg and both laws under Calvo. The model is closed-economy, with symmetric monopolistic competition and Rotemberg or Calvo price adjustment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q1: What is the irrelevance result the paper overturns, and why does CES produce it?&lt;/strong&gt;
Under CES, the market share function takes the form s(z) = gamma * z^(1-theta), yielding a constant price elasticity zeta = theta and a pass-through rate rho = 1, regardless of the number of firms or entry costs. As a result, concentration neither alters the slope of the NKPC nor generates any endogenous cost-push shock; competition is simply irrelevant to inflation dynamics. This irrelevance holds even with endogenous entry under CES.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q2: What is the Homothetic Single Aggregator (HSA) and why is it used?&lt;/strong&gt;
HSA is a class of homothetic demand systems, originally proposed by Matsuyama and Ushchev (2017), in which the market share of each intermediate input variety depends solely on its own price normalized by a single price aggregator A_t. This single aggregator serves as a sufficient statistic summarizing all competitive pressure effects on pricing behavior, including the markup rate and pass-through rate. HSA nests CES and Translog as special cases, is analytically tractable (equilibrium existence and uniqueness are straightforward to ensure with endogenous entry), and is flexible enough to accommodate both the Second and Third laws of demand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q3: What are Marshall&amp;rsquo;s Second and Third laws as defined in the paper?&lt;/strong&gt;
The Second law states that the price elasticity of demand zeta(z) is increasing in the normalized price z (equivalently, increasing in the single price aggregator A_t, which rises with fewer firms). The Third law, as defined by Matsuyama and Ushchev (2023b), states that the rate of increase in the price elasticity is decreasing in z. Together they ensure that both markup rates and pass-through rates respond systematically to changes in competitive pressure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q4: How does market concentration structurally flatten the NKPC under Rotemberg pricing?&lt;/strong&gt;
Under Rotemberg pricing, the slope of the NKPC equals (zeta(z) - 1) / chi, where chi is the Rotemberg price adjustment cost parameter. Higher entry costs reduce the equilibrium number of firms, which reduces competitive pressure and lowers z. Under the Second law, lower z reduces zeta(z), directly shrinking the slope coefficient. This is the steady-state effect of concentration: the structural slope of the curve declines because the price elasticity falls.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q5: How does market concentration structurally flatten the NKPC under Calvo pricing?&lt;/strong&gt;
Under Calvo pricing, the slope of the NKPC is positively related to the pass-through rate rho(z) rather than the price elasticity. The Third law implies that lower z (more concentration) reduces rho(z). Market concentration therefore causes structural flattening through the pass-through channel under Calvo. This is why the Calvo–Rotemberg equivalence — which holds to first order under CES — breaks down under HSA: Rotemberg highlights the Second law / price elasticity channel and Calvo highlights the Third law / pass-through channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q6: What is the endogenous cost-push shock and how does it arise?&lt;/strong&gt;
When the number of operating firms N_t changes endogenously, it alters the single price aggregator A_t and therefore the competitive environment facing each firm. Under the Second law, firms exhibit strategic complementarity in price setting: a firm reduces its markup when other firms lower their prices (A_t falls with more entry). Consequently, movements in N_t directly enter the NKPC as an additional term — (1/chi) * (1 - rho(z)) / rho(z) * N_hat_t — acting as an endogenous cost-push shock. This channel is absent under CES because rho = 1 makes the coefficient zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q7: How does the endogenous cost-push shock create a negative omitted variable bias?&lt;/strong&gt;
A naive regression of inflation on real marginal cost omits the N_hat_t term. Under the Second law, N_t is positively correlated with the marginal cost (more entry drives markups down, consistent with marginal cost movements), so the omitted variable N_hat_t is positively correlated with the included regressor. Because the true coefficient on N_hat_t in the NKPC is negative, omitting it biases the estimated slope on marginal cost downward (negative omitted variable bias). The estimated relationship between inflation and marginal cost is therefore weaker than the true structural relationship.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q8: How is the omitted variable bias amplified by concentration?&lt;/strong&gt;
Under the Third law (Rotemberg case) and under both the Second and Third laws (Calvo case), greater market concentration amplifies the magnitude of this negative bias. The intuition is that higher concentration makes the pass-through rate rho(z) smaller, which increases the coefficient on N_hat_t in the NKPC and thereby raises the magnitude of the bias when N_hat_t is omitted. Greater concentration thus generates both more structural flattening and more observational flattening simultaneously.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q9: What are the quantitative magnitudes of Phillips curve flattening in the simulations?&lt;/strong&gt;
De Loecker, Eeckhout, and Unger (2020) document that aggregate markups rose from 21% above marginal cost to 61% — approximately 40 percentage points. The paper&amp;rsquo;s simulations imply this corresponds to an entry cost increase of roughly 3.5 times under Translog and roughly 2.5 times under Co-PaTh with rho = 0.5. According to Figure 2, the accompanying market concentration can halve the slope of the NKPC. The slope declines more steeply for demand systems with smaller pass-through rates (rho further from 1).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q10: How do impulse responses change with market concentration?&lt;/strong&gt;
As entry costs rise (deeper concentration), the responses of the inflation rate to both technology shocks and monetary policy shocks become smaller in magnitude. Under the Second law, a positive technology shock increases the number of firms through a wealth effect, but strategic complementarity in price setting reduces markups, muting the inflation response relative to CES. The dynamic effect of endogenous entry thus weakens the transmission of real economic shocks to inflation — a supply side effect of monetary policy that parallels Baqaee, Farhi, and Sangani (2021) but operates through firm entry rather than the misallocation channel.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q11: What is the cyclicality of the markup rate under HSA, and why is it ambiguous?&lt;/strong&gt;
Under CES with flexible prices, the markup is constant. Under CES with sticky prices, the markup is procyclical (marginal cost falls with a positive technology shock but the price is rigid in the short run). Under the Second law with flexible prices, a positive technology shock increases firm entry, which reduces markups, making the markup countercyclical. In a sticky price equilibrium under the Second and Third laws, the cyclicality is therefore ambiguous: it depends on the tension between nominal rigidities (pushing toward procyclicality) and the pass-through rate (pushing toward countercyclicality).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q12: Why do the three price indices in the model differ, and which is used for the NKPC?&lt;/strong&gt;
The model features three aggregate price measures: the final goods price (CPI) P_t, which captures productivity effects of entry; the single price aggregator A_t, which captures competitive effects of entry and is the reference price for firms; and the average price index (PPI) p_t, which is not affected by entry effects and is the measured price index. Because entry effects shift P_t and A_t in ways that are not directly observed, the paper evaluates NKPC responsiveness in terms of p_t (PPI inflation), the measurable index.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q13: How does this paper relate to Wang and Werning (2022) and Baqaee, Farhi, and Sangani (2021)?&lt;/strong&gt;
Wang and Werning (2022) use a dynamic oligopoly model with exogenous entry and CES/Kimball demand, showing that higher concentration amplifies real effects of monetary policy and generates inflation persistence and endogenous cost-push shocks. Baqaee, Farhi, and Sangani (2021) use monopolistic competition with exogenous entry and Kimball demand under Calvo pricing, showing flattening through real rigidities and a misallocation channel (supply side effects of monetary policy). This paper uses monopolistic competition with endogenous entry and HSA under both Rotemberg and Calvo pricing; it produces supply side effects through firm entry rather than misallocation, and uses HSA rather than Kimball because HSA more readily guarantees equilibrium uniqueness with endogenous entry.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q14: What parametric families of HSA are used in simulations and what are their properties?&lt;/strong&gt;
Three families are used: CES (constant price elasticity theta, pass-through rho = 1, benchmark); Translog (satisfies the Second law, variable markups and pass-through); and Co-PaTh or Constant Pass-Through (proposed by Matsuyama and Ushchev 2020a, constant pass-through rate rho in (0,1) under flexible prices, containing CES as a limit as rho approaches 1). For Calvo pricing, a fourth family — PEM (Power Elasticity of Markup, proposed by Matsuyama and Ushchev 2023b) — is used; PEM satisfies the Third law in its strong form and contains Co-PaTh as a limit case. Translog is noted to behave similarly to Co-PaTh with rho = 0.5.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q15: What are the policy implications for central banks?&lt;/strong&gt;
Rising market concentration, by flattening the NKPC both structurally and observationally, reduces the effectiveness of monetary policy in achieving price stability through real economic activity — consistent with the concerns expressed by Federal Reserve officials (Clarida, Daly, Williams) quoted in the paper. The results suggest that empirical estimates of the NKPC slope that omit endogenous entry dynamics will be systematically biased downward, potentially leading central banks to underestimate the true structural responsiveness of inflation to demand conditions. Competition policy and barriers to entry thus have macroeconomic consequences beyond standard allocative efficiency considerations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Homothetic Single Aggregator (HSA):&lt;/strong&gt; A class of homothetic demand systems in which the market share of each input variety depends solely on its own price normalized by a single price aggregator A_t, which serves as a sufficient statistic for all competitive pressure effects on firm pricing behavior including the markup rate and pass-through rate. Nests CES and Translog as special cases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marshall&amp;rsquo;s Second Law of Demand (as used in the paper):&lt;/strong&gt; The condition that the price elasticity of demand zeta(z) is strictly increasing in the firm&amp;rsquo;s normalized price z. Under this condition, markup rates and pass-through rates vary endogenously with competitive pressure, and strategic complementarity in price setting arises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marshall&amp;rsquo;s Third Law of Demand (as used in the paper):&lt;/strong&gt; The condition, defined by Matsuyama and Ushchev (2023b), that the rate of increase in the price elasticity is decreasing in z. This law determines how the pass-through rate responds to concentration changes and is the relevant condition for structural flattening under Calvo pricing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Pass-through rate rho(z):&lt;/strong&gt; The fraction of a cost change that a monopolistically competitive firm passes through to its price under flexible pricing, defined as rho(z) = [1 - d&lt;em&gt;ln(zeta/(zeta-1))/d&lt;/em&gt;ln(z)]^(-1). Under CES, rho = 1 (complete pass-through); under the Second law, rho &amp;lt; 1 (incomplete pass-through); it declines with concentration under the Third law.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous cost-push shock:&lt;/strong&gt; The direct effect of changes in the endogenous number of firms N_t on inflation in the NKPC, arising from strategic complementarity in price setting under HSA. This term is absent under CES (where the coefficient is zero) and generates an omitted variable bias in naive regressions of inflation on marginal cost.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Steady-state (structural) flattening:&lt;/strong&gt; The reduction in the true structural slope of the NKPC caused by market concentration operating through lower price elasticity (Rotemberg channel) or lower pass-through rate (Calvo channel). This is the first of the paper&amp;rsquo;s two reasons for observed Phillips curve flattening.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Observational (omitted variable bias) flattening:&lt;/strong&gt; The downward bias in empirically estimated NKPC slopes arising because naive regressions omit the endogenous cost-push shock term. The bias is negative and is amplified by greater market concentration under the Third law and/or Second law depending on the pricing mechanism.&lt;/p&gt;</description></item><item><title>Competition in a Spatially-Differentiated Product Market with Negotiated Prices</title><link>https://macropaperwarehouse.com/papers/competition-in-a-spatially-differentiated-product-market-with-negotiated-prices/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/competition-in-a-spatially-differentiated-product-market-with-negotiated-prices/</guid><description>&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;How does individually negotiated pricing — where buyers make discrete choices among differentiated products and negotiate transaction-specific prices — affect market power and merger effects in oligopoly markets, and how do these effects differ from the uniform-pricing benchmark?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Setting&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper estimates the model using 13,788 transactions between the four main UK brick manufacturers and national house-building firms over 2003–2006. For each transaction (defined as a unique buyer-variety-destination-year combination), the data record the chosen product, negotiated price, production and delivery locations, volume, transport costs, and brick characteristics. The market is highly concentrated: four manufacturers held an 85% share of brick sales, with a two-firm concentration ratio of 0.60 and an HHI of 2,113. Spatial differentiation is a central feature — transport costs vary substantially by project location, and prices for the same brick product vary across the different projects of the same buyer depending on local competitive conditions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper develops an empirical model that adapts the Berry, Levinsohn, and Pakes (1995) differentiated-products framework to individually negotiated pricing. In the model, each buyer negotiates simultaneously and bilaterally with the sellers of the first-best and runner-up products (defined by surplus — value minus cost). The equilibrium first-best markup equals the minimum of (i) the unconstrained Nash bargaining solution, bj(wj(1) − w0), and (ii) the first-best seller&amp;rsquo;s surplus advantage over the runner-up, (wj(1) − wj(2)). Runner-up and lower-ranked sellers earn zero markups in equilibrium. This outcome is shown to be consistent with a range of non-cooperative bargaining models (Binmore 1985, Bolton and Whinston 1993, Manea 2018) and lies in the core of the associated coalition game. The TIOLI posted-price model is nested as the special case where seller bargaining skill equals one. A tractable likelihood for the joint probability of observed product choice and negotiated price is derived under the assumption that idiosyncratic taste terms follow a Generalized Extreme Value (GEV) distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The estimated mean seller bargaining skill is b̄ = 0.41 (s.e. 0.03), and a likelihood ratio test rejects the TIOLI restriction with a chi-squared statistic of 847 (p &amp;lt; 0.001), confirming that buyer bargaining power is economically and statistically significant. The model-implied price-cost margins (Lerner index) are low on average — mean of 0.08 — but vary widely across transactions (coefficient of variation of 0.78). Project location matters: sellers extract higher margins from buyers that are relatively close, taking advantage of their transport-cost proximity. Multi-product ownership also affects markups, but its relevance varies by project.&lt;/p&gt;
&lt;p&gt;Switching from negotiated to uniform pricing raises average markups by 34% at the observed market structure. However, effects are heterogeneous: approximately 15% of transactions see markup decreases. Buyers who benefit from uniform pricing are those with relatively little runner-up competition — precisely the buyers who face weak bargaining positions under negotiated pricing, and for whom the seller&amp;rsquo;s ability to use that position is constrained under a uniform rule.&lt;/p&gt;
&lt;p&gt;Under negotiated pricing, a merger affects a transaction&amp;rsquo;s markup only if it brings the first-best and runner-up products for that transaction under joint ownership. A demerger to single-product manufacturers reduces total manufacturer surplus by 25%. The merger of the two largest firms increases total manufacturer surplus by 19%, but with highly unequal transaction-level effects. Comparing the same mergers across pricing regimes, negotiated pricing abates average markup-increasing merger effects but worsens them for a minority of transactions — those where the merger creates a first-best/runner-up pairing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The model applies to complete-information settings where prices are negotiated transaction-by-transaction, buyers single-source for each discrete purchase occasion, and sellers have multiple spatially differentiated products. It is most directly applicable to business-to-business markets where individual transaction values are large enough to justify project-level negotiation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the fundamental difference between negotiated pricing in this paper and the standard Nash-in-Nash (NiN) bargaining framework?&lt;/strong&gt;
A: In standard NiN (Horn and Wolinsky 1988), a buyer negotiates one price per product and trades positive quantities of all products with negotiated prices, so all negotiated prices are observed in transaction data. In this paper, buyers make discrete single-sourcing choices — each project uses exactly one product — so only the chosen product&amp;rsquo;s price appears in data; the runner-up product and its counterfactual price are unobserved. Additionally, under NiN, prices are set at the buyer level and apply uniformly to all the buyer&amp;rsquo;s needs, whereas here prices are negotiated separately for each project, generating intra-buyer cross-project price variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the equilibrium markup formula, and what determines whether the Nash bargaining solution or the TIOLI constraint binds?&lt;/strong&gt;
A: The equilibrium first-best markup is ρ*j(1) = min[bj(1)(wj(1) − w0), (wj(1) − wj(2))], the minimum of the unconstrained Nash bargaining solution and the first-best seller&amp;rsquo;s surplus advantage over the runner-up. The TIOLI constraint (surplus advantage) binds when the seller&amp;rsquo;s bargaining skill is sufficiently high that the unconstrained NBS would exceed the surplus advantage — that is, when bj(1)(wj(1) − w0) &amp;gt; (wj(1) − wj(2)). Runner-up and all lower-ranked sellers earn zero markups in equilibrium because competition from the first-best drives their outside-option constraint to bind.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why do third-best and lower-ranked sellers have no effect on equilibrium outcomes?&lt;/strong&gt;
A: Because the most attractive offer any seller below the runner-up could make is a zero markup, and the runner-up already offers a zero markup due to competition from the first-best. Since the runner-up at zero markup already offers the buyer at least as much utility as any third-best product, the third-best cannot improve the buyer&amp;rsquo;s position. Proposition 1 (part iii) shows that the equilibrium markup and choice are invariant to N for N in {2, &amp;hellip;, N̄}.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the paper address the econometric challenge that the runner-up product and its price are unobserved?&lt;/strong&gt;
A: The paper derives a tractable closed-form likelihood for the joint probability of the observed product choice and the observed negotiated price, integrating out the unobserved idiosyncratic taste terms along with their implications for the identity and surplus of the unobserved runner-up product. The GEV distributional assumption on taste terms is crucial: it ensures that (1) choice probabilities have a closed form, (2) the surplus advantage can be expressed in terms of observed surpluses and GEV terms, and (3) the probability that the NBS is constrained has a closed form. This reduces the full problem to a lower-dimensional numerical integral over the normally distributed random effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What empirical evidence motivates the negotiated pricing model over simpler alternatives?&lt;/strong&gt;
A: Four data patterns motivate the model. First, prices vary across projects even after controlling for product identity and buyer identity — intra-buyer cross-project variation that is inconsistent with standard NiN where prices are set at the buyer level. Second, prices are lower, other things equal, when there is greater local competition from manufacturers not chosen for a project — inconsistent with standard NiN where excluded products play no competitive role. Third, buyers have many projects and make a discrete single-sourcing choice for each. Fourth, sellers are multi-product firms with products differentiated spatially and in other dimensions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What do the price regressions reveal about price determinants?&lt;/strong&gt;
A: Adding year effects to a simple regression explains only a small share of price variation (R² rises from 0.000 to 0.118 for the full sample). Adding variety-year effects raises R² to 0.775 and adding buyer-variety-year effects to 0.918, but still leaves substantial unexplained variation. Panel B regressions show that prices decrease with quantity, increase with input prices (gas price coefficient 27.2, wage coefficient 8.3), decrease with buyer-to-seller size ratio (coefficient −2.51), and decrease with greater local competition (a distance advantage indicator raises price by about 0.48–2.20 and N(DST) count reduces price by about 1.49–1.53 depending on specification).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What do the parameter estimates imply about spatial differentiation and buyer preferences?&lt;/strong&gt;
A: Transport costs have a strongly negative effect on value (coefficient on distance is −1.27, s.e. 0.04), and the interaction of distance with fuel costs is also negative and significant. The nesting parameter σJ is estimated at 0.47, indicating substantial within-group taste correlation across products from the same firm. Product characteristics matter: red and wire-cut bricks are preferred, and there are significant interactions between weather conditions and technical brick characteristics (frost positively interacts with strength; rainfall negatively interacts with absorption), indicating that buyers value bricks whose technical performance is suited to their project&amp;rsquo;s climate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How is the mean seller bargaining skill estimated, and how is the TIOLI model rejected?&lt;/strong&gt;
A: The mean seller bargaining skill b̄ is estimated at 0.41 (s.e. 0.03), substantially below one. The TIOLI restriction corresponds to b̄ = 1 (all markup determined by surplus advantage). A likelihood ratio test rejects this restriction with a chi-squared statistic of 847 (p &amp;lt; 0.001), providing strong statistical evidence that buyer bargaining power — not just competitive pressure — constrains markups below the TIOLI level.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the main findings regarding the distribution of price-cost margins?&lt;/strong&gt;
A: Price-cost margins (Lerner index form) are low on average, with a mean of 0.08, but vary widely across transactions, with a coefficient of variation of 0.78. Sellers set higher margins to buyers located relatively close to them (lower transport costs make the seller more attractive to the buyer, strengthening the seller&amp;rsquo;s position). Multi-product manufacturer portfolios also affect markups, but the relevance of multi-product ownership varies across projects depending on whether different products from the same firm compete as first-best and runner-up for a given project.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the uniform pricing counterfactual show, and how does it differ from the Hotelling benchmark?&lt;/strong&gt;
A: Switching from individually negotiated to uniform pricing raises average markups by 34% at the observed market structure. However, effects are heterogeneous: approximately 15% of transactions see markup decreases. Buyers who benefit from the switch are those in transactions with relatively weak runner-up competition — who had weak bargaining positions under negotiated pricing — and who gain because uniform pricing prevents sellers from exploiting that weakness. This contrasts with the result from the simple Hotelling linear city model (Thisse and Vives 1988), where switching to uniform pricing raises all markups.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the demerger counterfactual quantify multi-product effects?&lt;/strong&gt;
A: Decomposing the observed market to single-product manufacturers reduces total manufacturer surplus by 25%. This large reduction reflects the role of multi-product ownership in determining who the runner-up is for each transaction: when a manufacturer owns multiple products, it can avoid internal competition between its own first-best and runner-up products, preserving its surplus advantage. The impact is highly unequal across individual transactions, however, because the relevance of multi-product effects depends on whether any of a manufacturer&amp;rsquo;s other products would have been the runner-up for a given project.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the merger of the two largest firms imply for markups and surplus?&lt;/strong&gt;
A: The merger of the two largest firms (by market share) increases total manufacturer surplus in the industry by 19%. Markup increases are very unequal across transactions: the merger affects only those transactions for which the merging firms jointly become the first-best and runner-up, which is the mechanism highlighted in the 2010 US Merger Guidelines for negotiated pricing markets. The heterogeneity of effects means that aggregate market-level concentration measures (such as HHI changes) can be poor proxies for merger effects in these markets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the pricing regime interact with merger effects?&lt;/strong&gt;
A: Comparing the same mergers under negotiated versus uniform pricing, negotiated pricing abates the average markup-increasing effects of mergers. However, for a minority of transactions — specifically those where the merger creates a first-best/runner-up pairing that did not exist pre-merger — negotiated pricing makes the merger&amp;rsquo;s markup effect worse than it would be under uniform pricing. This implies that the direction of the pricing-regime effect on merger harm is not uniform across buyers, and that transaction-level analysis is required for accurate antitrust assessment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the paper relate to the Competition Commission&amp;rsquo;s 2007 assessment of the Wienerberger/Baggeridge merger?&lt;/strong&gt;
A: The CC (2007) found the market highly concentrated (HHI 2,113, implied HHI increase of 390 from the merger, both exceeding guideline thresholds) but approved the merger, judging profitability to be at or below average for comparable industries and competition to be more intense than the concentration level alone would suggest. This paper&amp;rsquo;s model provides formal underpinning for that assessment: with negotiated pricing and buyer bargaining power, markups are constrained by the runner-up competitive threat at the transaction level, not by market-wide concentration, and the low mean Lerner index of 0.08 is consistent with the CC&amp;rsquo;s profitability finding.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What external validity evidence supports the model&amp;rsquo;s cost specification?&lt;/strong&gt;
A: The paper compares the marginal costs implied by the estimated model to plant-month level production cost data that were not used in estimation. A good match between the two provides external validation of the cost specification and supports the model&amp;rsquo;s structural interpretation of the markup decomposition.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;First-best and runner-up products&lt;/strong&gt;: Defined at the project level in terms of surplus (value minus cost). The first-best product j(i,1) is the inside good yielding the highest surplus for project i; the runner-up j(i,2) is the highest-surplus inside good not sold by the first-best seller. These two products — and only these two — determine the equilibrium markup and buyer choice; third-best and lower-ranked products are irrelevant.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Surplus advantage&lt;/strong&gt;: The difference wj(i,1) − wj(i,2) ≥ 0 between the first-best product&amp;rsquo;s surplus and the runner-up&amp;rsquo;s surplus for a given project. This is the competitive constraint on the first-best seller&amp;rsquo;s markup under TIOLI pricing and the binding ceiling on the negotiated markup whenever the unconstrained Nash bargaining solution would exceed it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Negotiated pricing&lt;/strong&gt;: A pricing arrangement in which buyers negotiate prices specific to the individual purchase occasion (here, each construction project), as opposed to uniform pricing where the pre-transport price is the same for all buyers. Prices are determined bilaterally between buyer and competing sellers, with the buyer&amp;rsquo;s outside option — buying the runner-up at its anticipated negotiated price — serving as the competitive constraint.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Outside option principle (Binmore et al. 1989)&lt;/strong&gt;: The principle that a rival offer (outside option) has no effect on a bilateral Nash bargaining problem unless it would leave the receiving party better off than the Nash bargaining solution — i.e., it constrains rather than shifts the disagreement point. In the paper&amp;rsquo;s model, the runner-up seller&amp;rsquo;s zero-markup offer serves as the first-best seller&amp;rsquo;s constraining outside option when seller bargaining skill is high.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GEV (Generalized Extreme Value) taste distribution&lt;/strong&gt;: The distributional assumption on project-product idiosyncratic match terms that makes the joint likelihood of observed product choice and negotiated price tractable. The GEV structure yields closed-form choice probabilities (nested logit) and allows the surplus advantage — which depends on unobserved runner-up surplus — to be expressed analytically, enabling joint estimation from transaction-level data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Price-cost margin (Lerner index)&lt;/strong&gt;: Markup (price minus cost) divided by price, used here at the transaction level. The estimated mean Lerner index is 0.08 with a coefficient of variation of 0.78, reflecting wide dispersion driven by spatial variation in local competition and first-best surplus advantage across transactions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nash-in-Nash (NiN) vs. single-sourcing bargaining&lt;/strong&gt;: NiN (Horn and Wolinsky 1988) applies when a buyer trades positive quantities of all products with negotiated prices (multi-sourcing); the paper&amp;rsquo;s model applies when a buyer makes a discrete single-sourcing choice per occasion, so only the chosen product&amp;rsquo;s price is observed. The distinction generates different data observability and different competitive mechanisms — in NiN, excluded products play no role; in this paper, the runner-up&amp;rsquo;s potential zero-markup offer disciplines the first-best seller&amp;rsquo;s markup.&lt;/p&gt;</description></item><item><title>Competitive Advertising and Pricing</title><link>https://macropaperwarehouse.com/papers/competitive-advertising-and-pricing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/competitive-advertising-and-pricing/</guid><description>&lt;p&gt;Hwang, Kim, and Boleslavsky study how firms in an oligopoly simultaneously choose prices and advertising strategies, where advertising is modeled as the choice of how much product information to disclose to consumers. The paper extends the canonical Perloff-Salop (1985) random-utility discrete-choice framework — in which n firms engage in Bertrand competition for a consumer whose value for each product is independently drawn from a common distribution F — by endogenizing the information environment: each firm may choose any mean-preserving contraction (MPC) of F as its advertising strategy, with no structural restriction on feasible content. This full flexibility, drawn from the information design literature, allows each firm to choose the consumer&amp;rsquo;s effective value distribution, ranging from full information (choosing F itself) to complete concealment (a degenerate distribution at the mean). The model is silent on advertising costs, which are assumed to be zero throughout.&lt;/p&gt;
&lt;p&gt;The central result is that intense competition forces firms to provide precise product information. Formally, the full information equilibrium — in which every firm chooses F — exists in the advertising game (the subgame in which prices are fixed symmetrically) if and only if F^(n-1) is convex over its support. Because F^(n-1) represents the distribution of the consumer&amp;rsquo;s best outside option, convexity means the consumer likely faces an attractive alternative, incentivizing each firm to maximize the chance of offering the highest possible value. Crucially, this convexity condition is guaranteed to hold when n is sufficiently large, regardless of the shape of F, because the power function x^(n-1) becomes more convex as n rises. This establishes that under sufficiently intense competition, full information disclosure is the unique symmetric equilibrium.&lt;/p&gt;
&lt;p&gt;The general equilibrium advertising strategy G* — which governs cases where full information is not an equilibrium — satisfies two necessary and sufficient conditions: (i) (G*)^(n-1) is convex over the support of G*, and (ii) for almost all values in the support, G* either coincides with F (where the MPC constraint binds, preventing further dispersion) or (G*)^(n-1) is locally linear (where the firm is locally risk-neutral and has no incentive to alter its distribution). The paper proves existence and uniqueness of G* for any F satisfying the stated regularity conditions (density positive, continuously differentiable, bounded, with finitely many peaks). When F has log-concave density, a unique symmetric pure-price equilibrium (p*, G*) exists in the full game.&lt;/p&gt;
&lt;p&gt;The paper demonstrates that strategic advertising has ambiguous implications for prices and consumer welfare. Strategic advertising necessarily reduces social surplus through information loss, since consumers select suboptimal products with positive probability when G* differs from F. However, it compresses the support of the value distribution relative to F, which — by a new result (Proposition 3) — tends to lower the equilibrium price. Offsetting this, strategic advertising also redistributes marginal consumers in ways that may raise or lower the price. In the duopoly case with power distributions F(v) = v^alpha on [0,1], strategic advertising lowers the market price if and only if alpha &amp;gt; 1/sqrt(2) (approximately 0.7071), and raises consumer surplus if and only if alpha &amp;gt; 0.7928.&lt;/p&gt;
&lt;p&gt;The paper examines three extensions: (1) a binding consumer outside option, (2) multi-unit (k-out-of-n) demand, and (3) asymmetric firms with two types. In all three cases, full information cannot be a strict equilibrium for any finite n under the relevant structural condition, yet the equilibrium distribution G* converges pointwise to F as n tends to infinity, preserving the paper&amp;rsquo;s core asymptotic insight.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the main research question?&lt;/strong&gt;
A: The paper asks how much product information firms will voluntarily disclose when they compete both on price and advertising content in an oligopoly. Unlike the monopoly literature, the oligopoly context creates strategic interdependencies — each firm&amp;rsquo;s optimal disclosure depends on rivals&amp;rsquo; disclosure choices — that the paper characterizes fully.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How is advertising modeled, and why use mean-preserving contractions?&lt;/strong&gt;
A: Each firm&amp;rsquo;s advertising strategy is modeled as a choice of any mean-preserving contraction (MPC) of the true value distribution F. An MPC preserves the expected value but reduces dispersion, capturing the idea that a firm can selectively conceal information (moving toward a degenerate distribution) but cannot fabricate value dispersion beyond what F allows. Because consumers are risk-neutral and buy based on expected values net of prices, this MPC formulation captures full flexibility in information design without loss of generality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the precise necessary and sufficient condition for the full information equilibrium in the advertising game?&lt;/strong&gt;
A: The full information equilibrium — in which every firm chooses F — exists if and only if F^(n-1) is convex over its support [v, v̄]. The &amp;ldquo;only if&amp;rdquo; direction follows from Lemma 1: in any equilibrium, (G*)^(n-1) must be convex, so if F^(n-1) is not convex, F is not an equilibrium. The &amp;ldquo;if&amp;rdquo; direction follows because a convex F^(n-1) makes each firm locally risk-loving, so no MPC of F yields a higher payoff than F itself.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does sufficiently intense competition force full information disclosure?&lt;/strong&gt;
A: For any distribution F with positive, continuously differentiable, bounded density f with bounded derivative f&amp;rsquo;, the second derivative of F^(n-1) satisfies F(v)^(n-1)&amp;rsquo;&amp;rsquo; &amp;gt;= (n-1)F(v)^(n-3)[(n-2)epsilon^2 - M], where epsilon = min f(v)^2 &amp;gt; 0 and M = max |f&amp;rsquo;(v)| &amp;lt; infinity. This expression is strictly positive for n sufficiently large, so F^(n-1) is convex and the full information equilibrium exists. Economically, with many competitors each firm wins the consumer only when it offers the highest possible value, so providing full information is optimal.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;Q: What are the two necessary and sufficient properties characterizing the general equilibrium advertising strategy G&lt;/em&gt;?&lt;/em&gt;*
A: First (Lemma 1), (G*)^(n-1) must be convex over the support of G* — this prevents any firm from profitably concentrating mass to reduce dispersion. Second (Lemma 2), for almost all values in the support, either G* = F locally (the MPC constraint binds, preventing further dispersion) or (G*)^(n-1) is locally linear (the firm is locally risk-neutral and indifferent over distributions with the same local mean). Theorem 1 proves these two conditions are both necessary and sufficient, and that G* is unique for any F satisfying the stated regularity conditions.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;Q: What structure does G&lt;/em&gt; take when F^(n-1) has strictly quasi-concave density?&lt;/em&gt;*
A: By Corollary 2(1), there exists a cutoff v* in [v, v̄] such that G*(v) = F(v) for v &amp;lt;= v* (full information below the cutoff) and (G*)^(n-1) is linear above v*. As n increases, v* rises, meaning the region of full disclosure expands, and G* increases in convex order — so consumers receive strictly more information. One immediate implication is that consumer surplus strictly increases in n: consumers benefit both from more options and from more accurate information about each product.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What happens when F^(n-1) is concave?&lt;/strong&gt;
A: By Corollary 3, when F^(n-1) is concave, (G*)^(n-1) is linear over the entire support, with lower bound v. In the illustrative Example 1 (truncated exponential with n=2), this yields G* = U[0, 2*mu_F] — a uniform distribution on an interval whose upper bound is twice the mean of F.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does strategic advertising raise or lower equilibrium prices, and consumer surplus?&lt;/strong&gt;
A: Both effects are ambiguous and depend on the shape of F. Strategic advertising compresses the support of the value distribution (since G* is an MPC of F), which by Proposition 3(1) tends to lower equilibrium prices. But it also reshapes the distribution of marginal consumers, which may raise or lower prices. In the power distribution example (n=2, F(v) = v^alpha on [0,1]), strategic advertising lowers the market price if and only if alpha &amp;gt; 1/sqrt(2) ≈ 0.7071, and raises consumer surplus if and only if alpha &amp;gt; 0.7928. Thus even with deadweight loss from information suppression, consumers can be better off under strategic advertising than under forced full disclosure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does Proposition 3 contribute about equilibrium prices in the Perloff-Salop model?&lt;/strong&gt;
A: Proposition 3 delivers two results about how the distribution of marginal consumers (integral (F^(n-1))&amp;rsquo; dF) determines equilibrium prices. First, the measure of marginal consumers decreases if F is proportionally stretched over a larger support, confirming that longer support raises equilibrium prices. Second — presented as novel — among all distributions with support in [v, v̄], the power distribution F(v) = ((v-v)/(v̄-v))^(2/n) minimizes the measure of marginal consumers, corresponding to the maximum equilibrium price. The key property is that marginal consumers are uniformly distributed under this power distribution, and any deviation from uniformity allows a &amp;ldquo;flattening&amp;rdquo; adjustment that increases the measure of marginal consumers and lowers the price.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Under what condition does the full game (price plus advertising) have a unique symmetric pure-price equilibrium?&lt;/strong&gt;
A: Theorem 2 states that log-concavity of the density f is sufficient for existence and uniqueness of a symmetric pure-price equilibrium (p*, G*) as characterized in Theorems 1 and 2. Log-concavity ensures that the equilibrium distribution G* has a convex-linear structure (as in Corollary 2), which preserves log-concavity of each firm&amp;rsquo;s profit function even under compound deviations (simultaneous changes to both price and advertising strategy), making the first-order conditions sufficient for global optimality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Can strategic advertising create or destroy pure-price equilibria relative to the Perloff-Salop benchmark?&lt;/strong&gt;
A: Yes, both directions are possible. When F^(n-1) is convex (so G* = F), equilibrium existence in the Perloff-Salop (PS) model is necessary but not sufficient for existence in the full model, because compound deviations (changing both price and advertising) may be profitable even when pure price deviations are not. Conversely, when G* differs from F, the changed distribution of marginal consumers can sustain an equilibrium in the full model even when none exists in PS. Appendix E of the paper provides a specific example of the latter phenomenon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What happens with a binding consumer outside option?&lt;/strong&gt;
A: Proposition 4 shows that a full information equilibrium never exists in the advertising game when the consumer has a binding outside option (p* in (v, v̄)). The firm&amp;rsquo;s value function acquires a discrete jump at p* due to the indicator 1_{v &amp;gt;= p*}, making it optimal to pool mass around p* rather than disclose fully. Nevertheless, Proposition 5 proves that G* converges pointwise to F as n tends to infinity, because the jump of size F(p*)^(n-1) vanishes exponentially fast as n grows.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does the full information result survive multi-unit demand?&lt;/strong&gt;
A: No. Proposition 6 shows that with k &amp;gt; 1 units demanded (out of n products), the full information equilibrium never exists for any finite n or F. The reason is that phi&amp;rsquo;(v; F) — the firm&amp;rsquo;s marginal value of offering value v — is zero at v̄ when k &amp;gt; 1, so the firm can profitably pool values near the top of the support. However, Proposition 7 shows that G* converges pointwise to F as n tends to infinity (with k fixed), preserving the asymptotic full information result.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What happens with asymmetric firms differing in their value distribution supports?&lt;/strong&gt;
A: Proposition 8 shows a sharp dichotomy. If both firm types share the same upper bound of their value supports (v̄_1 = v̄_2), the full information equilibrium exists whenever both F_1^(n1-1) and F_2^(n2-1) are convex. If the supports have different upper bounds (v̄_1 &amp;lt; v̄_2), the full information equilibrium never exists regardless of n_1 and n_2, because type-2 firms face a downward kink in their winning probability at v̄_1 and always have an incentive to pool mass there. The authors conjecture that G*_1 and G*_2 still converge to F_1 and F_2 asymptotically but do not prove this due to technical complexity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does this paper relate to Ivanov (2013)?&lt;/strong&gt;
A: Ivanov (2013) also uses the Perloff-Salop framework and shows that full information is an equilibrium when n is sufficiently large, but restricts advertising to rotation-ordered strategies (in the sense of Johnson and Myatt, 2006). The present paper imposes no structural restriction and strengthens Ivanov&amp;rsquo;s result by: (a) providing a necessary and sufficient condition for the full information equilibrium (not just a sufficient condition for large n); (b) fully characterizing G* when full information is not an equilibrium; and (c) demonstrating robustness across multiple model variants.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What policy implication does the ambiguity result carry?&lt;/strong&gt;
A: The paper warns against assuming that mandating full information disclosure is unambiguously consumer-beneficial. While strategic advertising creates deadweight loss through information suppression, it can simultaneously compress support and alter the marginal consumer distribution in ways that lower equilibrium prices significantly. The power distribution example (alpha &amp;gt; 0.7928) shows consumers can be strictly better off under strategic advertising than under forced full disclosure. This ambiguity is a cautionary tale for disclosure regulation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mean-Preserving Contraction (MPC):&lt;/strong&gt; A distribution G_i is an MPC of F if it has the same mean as F but less dispersion (in the sense of second-order stochastic dominance). In the paper, each firm&amp;rsquo;s feasible advertising strategies are exactly the set MPC(F) — this captures all informationally feasible disclosures without structural restriction on content.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Advertising Game:&lt;/strong&gt; A restricted subgame of the full market game in which firms choose their advertising strategies G_i taking the symmetric price as given. An equilibrium in the advertising game is a necessary condition for equilibrium in the full game. The advertising game&amp;rsquo;s equilibrium uniquely pins down G* independently of the price level (under the baseline model without binding outside option).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Full Information Equilibrium:&lt;/strong&gt; An equilibrium of the advertising game in which every firm chooses the true underlying distribution F as its advertising strategy. This corresponds to complete, unobstructed product disclosure. The paper&amp;rsquo;s central result is that this equilibrium exists if and only if F^(n-1) is convex over its support.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Convexity of F^(n-1):&lt;/strong&gt; The key distributional condition governing advertising equilibria. F^(n-1) is the distribution of the consumer&amp;rsquo;s best alternative among (n-1) rivals&amp;rsquo; products. Convexity of F^(n-1) means its density is increasing, signaling a likely attractive outside option, which makes each firm risk-loving and induces full disclosure. This convexity is guaranteed for n sufficiently large.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;Locally Linear (G&lt;/em&gt;)^(n-1):&lt;/em&gt;* A region of the equilibrium distribution where (G*)^(n-1) has constant slope, making the firm locally risk-neutral. Over such a region, the firm is indifferent among all distributions with the same local mean, and the equilibrium G* need not coincide with F — it is only required to be an MPC of F on that interval. This alternating structure (coinciding with F on strictly convex regions; linear elsewhere) fully characterizes G*.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Marginal Consumers:&lt;/strong&gt; In the Perloff-Salop pricing formula, the equilibrium price p* = (1/n) / integral [(G*(v)^(n-1))&amp;rsquo; dG*(v)]. The integrand (G*(v)^(n-1))&amp;rsquo; * g*(v) is the density of consumers who are indifferent between a given firm&amp;rsquo;s product and their best alternative at value v. A larger measure of marginal consumers implies lower equilibrium prices through greater competitive pressure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Compound Deviation:&lt;/strong&gt; In the full game, a deviation by a firm that changes both its price p_i and its advertising strategy G_i simultaneously, rather than varying only one dimension. The possibility of compound deviations is what distinguishes equilibrium existence conditions in the full model from those in the standard Perloff-Salop model, even when G* = F.&lt;/p&gt;</description></item><item><title>Consistent Evidence on Duration Dependence of Price Changes</title><link>https://macropaperwarehouse.com/papers/consistent-evidence-on-duration-dependence-of-price-changes/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/consistent-evidence-on-duration-dependence-of-price-changes/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; This paper asks two related questions. First, can one develop a robust, distribution-free estimator for the discrete-time mixed proportional hazard (MPH) model of duration with unobserved heterogeneity? Second, what does that estimator reveal about the shape of the hazard of price changes, the role of heterogeneity in shaping aggregate price dynamics, and the distinction between regular price changes and sales?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Methodology.&lt;/strong&gt; The authors develop a linear generalized method of moments (GMM) estimator for the discrete-time MPH model, building on identification results in Honoré (1993). The model specifies that the probability a price spell ends at duration t, conditional on surviving to t, equals the product of a product-specific frailty parameter θ (unobserved, fixed over time) and a common baseline hazard bt. The estimator exploits repeated price spells per product via moment conditions that are linear in bt, making estimation and inference straightforward. It accommodates right- and left-censored data, competing risks, and spell-specific observable characteristics, without requiring any parametric assumption on the frailty distribution. The estimator is consistent as the number of products grows, even with a short time dimension. A Hansen-Sargan J-test of overidentifying restrictions and a test of the monotone-average-type prediction are also developed.&lt;/p&gt;
&lt;p&gt;The estimator is applied to two datasets: (1) IRI weekly store data (2001–2011), covering 30 product categories and more than 21 million products, yielding 684,919,778 pairs of durations; and (2) Online Micro Price data from Cavallo (2018), comprising approximately 250,000 products at daily frequency.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings with Quantitative Magnitudes.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Baseline hazard and heterogeneity.&lt;/em&gt; In the pooled IRI data, the Kaplan-Meier hazard is steeply declining throughout the entire range from 2 to 60 weeks. In contrast, the estimated baseline hazard is roughly constant until week 4 and then declines only modestly, with a noticeable spike at week 52. The ratio of the Kaplan-Meier hazard to the baseline hazard — the average type, E[θ|t] — drops by approximately 60 percent within the first 20 weeks, and continues to decline, reaching roughly 0.3 of its initial value after one year. This decomposition reveals substantial unobserved heterogeneity that accounts for a large fraction of the observed decline in the Kaplan-Meier hazard.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Implications for structural models.&lt;/em&gt; The finding of a decreasing baseline hazard is inconsistent with canonical state-dependent pricing models (Golosov and Lucas, 2007), which predict an increasing hazard, conditional on a given firm&amp;rsquo;s type. The decreasing baseline hazard is instead broadly consistent with time-dependent pricing models, though not with a constant-hazard (Calvo, 1983) specification.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Monetary policy impulse response.&lt;/em&gt; In a calibrated time-dependent pricing model with strategic complementarity (α = 0, 0.5, 0.95), the aggregate price level dynamics in the estimated heterogeneous-firm MPH economy are close to those of a homogeneous-firm economy that uses the Kaplan-Meier hazard as the common price-change hazard. The homogeneous-firm approximation is substantially closer to the MPH economy than a Taylor (1979, 1980) staggered-contract economy with the same Kaplan-Meier hazard, particularly when strategic complementarity is strong (α = 0.95). The Calvo economy provides a poor approximation due to its exponential (constant-speed) price convergence structure.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Regular versus temporary price changes.&lt;/em&gt; Using the competing-risks extension with spell-specific observables — classifying spells by whether they start and end with a price increase (+) or decrease (−) — the authors separately estimate four baseline hazards. The baseline hazard for consecutive price increases (b++t) is relatively flat, especially for the first 6 weeks, then flat until week 45, with a spike near one year, consistent with price-plan models. The baseline hazard for reversals (particularly b−+t, price decreases followed by price increases, associated with sales) is steeply declining. The J-test statistics are substantially lower for price trends (J++ = 3,920; J−− = 3,401) than for reversals (J+− = 8,737; J−+ = 7,910), and markedly lower than the pooled-model J = 10,498, indicating that the MPH structure fits regular price changes considerably better than sales.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; Results are conditional on weekly store-level price data for mostly packaged consumer goods (30 IRI product categories). The analysis focuses on price spells of at least 2 weeks to avoid spurious duration-one spells from mid-week price changes. The maximum duration examined is 60 weeks. The comparison of estimation methods relies on the IRI data only; the Online Micro Price data confirm weekly decision-making through a spike in the daily hazard every 7 days. Comparisons with maximum likelihood estimates show that GMM recovers more heterogeneity (average type declines to 0.37 at 6 months by GMM versus 0.48 by continuous-time MLE), and that time aggregation explains most of the discrepancy between the two methods.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-mixed-proportional-hazard-mph-model-as-used-in-this-paper-and-what-does-the-estimator-identify"&gt;Q1. What is the mixed proportional hazard (MPH) model as used in this paper, and what does the estimator identify?&lt;/h3&gt;
&lt;p&gt;A1. The MPH model specifies that the hazard that a price spell ends at duration t, conditional on surviving to t, equals θ·bt, where θ is a product-specific frailty parameter drawn from an unknown distribution G and bt is a baseline hazard common to all products. The estimator, which is linear in bt, identifies the baseline hazard up to a multiplicative constant using moment conditions derived from repeated spell data, without restricting the shape of the frailty distribution. Identification relies on comparing the joint survival probabilities of two consecutive spells for the same product and exploits the symmetry implied by the MPH structure across spells.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-kaplan-meier-hazard-relate-to-the-baseline-hazard-and-what-does-this-relationship-imply-about-heterogeneity"&gt;Q2. How does the Kaplan-Meier hazard relate to the baseline hazard, and what does this relationship imply about heterogeneity?&lt;/h3&gt;
&lt;p&gt;A2. The paper proves that the Kaplan-Meier hazard Ht equals bt times E[θ|t], the mean frailty among spells surviving to duration t. Because higher-type products (those with a higher propensity to change prices) exit the pool of surviving spells earlier, E[θ|t] is strictly decreasing in t — a form of dynamic selection. The ratio Ht/bt, normalized to 1 at the start, falls to approximately 0.4 by week 20 in the pooled IRI data and to approximately 0.3 after one year, documenting that a large share of the decline in the Kaplan-Meier hazard reflects heterogeneity rather than structural negative duration dependence.&lt;/p&gt;
&lt;h3 id="q3-what-does-the-estimated-baseline-hazard-imply-about-structural-models-of-price-setting"&gt;Q3. What does the estimated baseline hazard imply about structural models of price setting?&lt;/h3&gt;
&lt;p&gt;A3. A decreasing baseline hazard is inconsistent with the canonical state-dependent model of Golosov and Lucas (2007), in which a firm&amp;rsquo;s hazard of price change is increasing in the time since the last change, because larger deviations from the desired price accumulate with duration. The decreasing baseline hazard is instead consistent with time-dependent pricing models and with price-plan models where within-plan switches are costless. The mild spike at week 52 in the baseline hazard is consistent with Taylor-type annual pricing rules.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-approximate-aggregation-result-for-monetary-policy-and-how-quantitatively-accurate-is-it"&gt;Q4. What is the approximate aggregation result for monetary policy, and how quantitatively accurate is it?&lt;/h3&gt;
&lt;p&gt;A4. In the time-dependent pricing model without strategic complementarity (α = 0), the impulse response of the aggregate price level to a monetary shock in a heterogeneous-firm economy is exactly the same as in a homogeneous-firm economy whose single firm uses the Kaplan-Meier survival function. This extends Carvalho and Schwartzman (2015) to an approximation in the case with strategic complementarity (α = 0.5 and α = 0.95). Numerically, the path of aggregate prices in the estimated MPH economy is close to that in the homogeneous-firm Kaplan-Meier economy, and substantially closer to it than to the Taylor-contract economy — the difference is most pronounced at horizons beyond about half a year when α = 0.95, where the Taylor economy shows notably slower initial convergence and faster later convergence relative to the MPH and homogeneous economies.&lt;/p&gt;
&lt;h3 id="q5-how-do-the-papers-results-differ-from-those-obtained-using-maximum-likelihood-estimation-of-the-continuous-time-mph-model"&gt;Q5. How do the paper&amp;rsquo;s results differ from those obtained using maximum likelihood estimation of the continuous-time MPH model?&lt;/h3&gt;
&lt;p&gt;A5. The GMM estimator recovers substantially more heterogeneity than maximum likelihood (MLE) applied to the continuous-time model with continuous records (assumed gamma frailty). The average type falls from 1 to 0.37 at six months under GMM, versus only 0.48 under MLE. The authors investigate two sources of this discrepancy: the assumed frailty distribution family (gamma) and time aggregation. They conclude that time aggregation is quantitatively more important in the IRI weekly data — that is, the continuous-time MLE approach fails to properly account for the discrete nature of the data-generating process, leading it to understate heterogeneity and recover a steeper baseline hazard.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-distinguish-regular-price-changes-from-sales-without-directly-observing-a-sales-flag"&gt;Q6. How does the paper distinguish regular price changes from sales without directly observing a sales flag?&lt;/h3&gt;
&lt;p&gt;A6. The competing-risks extension classifies each spell by whether it starts with a price increase or decrease (observable characteristic χ ∈ {+, −}) and by whether it ends with a price increase or decrease (competing risk ρ ∈ {+, −}). Price trends — spells where the direction is the same at both the start and end (++ or −−) — are interpreted as regular price changes; price reversals (especially −+, i.e., price decrease followed by increase) are associated with sales. This approach is consistent with the statistical model used for estimation, avoids the bias from simply dropping suspected sales spells before estimation, and allows the MPH structure to hold only for the risks of interest even if it fails for others.&lt;/p&gt;
&lt;h3 id="q7-how-well-does-the-mph-model-fit-regular-price-changes-versus-sales"&gt;Q7. How well does the MPH model fit regular price changes versus sales?&lt;/h3&gt;
&lt;p&gt;A7. The J-test of overidentifying restrictions yields test statistics of J++ = 3,920 for consecutive price increases and J−− = 3,401 for consecutive price decreases, compared with J = 10,498 for the pooled model and J+− = 8,737 and J−+ = 7,910 for the reversal hazards. All rejections are at conventional significance levels (critical value 1,749 at 5%), but the rejection is substantially milder for price trends than for price reversals. For individual product categories, the model cannot be rejected for 8 categories (out of 30) for b++ and 21 categories for b−−, suggesting the MPH structure is a much better description of regular price changes than of sales.&lt;/p&gt;
&lt;h3 id="q8-what-role-do-one-week-price-spells-play-in-the-data-and-why-are-they-excluded"&gt;Q8. What role do one-week price spells play in the data, and why are they excluded?&lt;/h3&gt;
&lt;p&gt;A8. In the IRI data, prices are measured as the ratio of weekly revenue to quantity, so a price change occurring mid-week generates a spurious price spell of duration one week. If all spells including one-week spells are retained, the autocorrelation of spell durations is only 0.029 in levels and even negative (−0.042) in logs, which is inconsistent with a mixture model. Once one-week spells are excluded, the autocorrelation rises to 0.235 in levels and 0.233 in logs, and is stable when two-week spells are also excluded (0.248 and 0.256). The paper therefore sets the lower duration bound at T̲ = 2 weeks.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-daily-online-micro-price-data-add-relative-to-the-weekly-iri-data"&gt;Q9. What does the daily Online Micro Price data add relative to the weekly IRI data?&lt;/h3&gt;
&lt;p&gt;A9. The daily data reveal a sharp spike in the price-change hazard every seven days, suggesting that even when prices are observed daily, the decision to change prices is made at the weekly frequency. This justifies the use of a discrete-time model with a one-week period. The estimates from daily and weekly aggregations of the same data are broadly similar, though weekly data recovers somewhat less heterogeneity than daily data. Aggregating IRI weekly data to monthly frequency understates heterogeneity even more, confirming that frequency matters for measuring heterogeneity.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-computational-advantages-of-the-gmm-estimator-relative-to-maximum-likelihood"&gt;Q10. What are the computational advantages of the GMM estimator relative to maximum likelihood?&lt;/h3&gt;
&lt;p&gt;A10. Because the moment conditions are linear in the baseline hazard bt, the GMM estimator is obtained in closed form, making estimation fast and inference straightforward. On the pooled IRI sample, GMM estimation (including standard errors) required 70 minutes on a machine with 60 GB memory, whereas the maximum likelihood estimator required 15 hours on a machine with 256 GB memory and failed entirely on the 60 GB machine. The GMM approach also avoids the need to specify the frailty distribution family and guarantees a global solution (proved by the identification result), whereas the likelihood function is non-linear in bt and may have multiple local maxima.&lt;/p&gt;
&lt;h3 id="q11-what-is-the-shape-of-the-b-baseline-hazard-for-regular-price-increases-and-what-models-does-it-support"&gt;Q11. What is the shape of the b++ baseline hazard for regular price increases, and what models does it support?&lt;/h3&gt;
&lt;p&gt;A11. The baseline hazard for spells starting and ending with a price increase (b++) is decreasing during the first 6 weeks — dropping by almost 50% — and then flat until approximately week 45, with a pronounced spike at around one year. This shape is consistent with price-plan models (Eichenbaum, Jaimovich, and Rebelo, 2011) with Calvo-type switching between plans, where within-plan changes are costless and the hazard of between-plan switching is approximately constant. The annual spike is consistent with Taylor-type pricing. Approximately 76.8% of complete spells starting after a price increase last at most 6 weeks.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Baseline hazard (bt).&lt;/strong&gt; The component of the MPH hazard that is common to all products and may vary arbitrarily with elapsed duration t. It represents structural duration dependence — the tendency for a given product to be more or less likely to change price as a function of how long its current spell has lasted — net of heterogeneity. It is identified only up to a multiplicative constant.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Frailty parameter (θ) / frailty distribution (G).&lt;/strong&gt; The product-specific scaling factor in the MPH model, fixed over all spells for a given product, that captures permanent unobserved differences in price-change frequency across products. The paper treats G as a nuisance parameter and does not require a parametric assumption on its shape. A higher θ means the product has a higher baseline propensity to change its price.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Average type (E[θ|t]).&lt;/strong&gt; The mean frailty parameter among spells that have survived to at least duration t. Because high-type products change price earlier and exit the pool of surviving spells first, the average type is provably strictly decreasing in t under the MPH model. It is measured as the ratio of the Kaplan-Meier hazard to the baseline hazard, and its rate of decline measures the importance of dynamic selection.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Kaplan-Meier hazard (Ht).&lt;/strong&gt; The probability that a randomly drawn spell ends at duration t, conditional on having lasted at least t periods. It mixes together structural duration dependence (captured by bt) and dynamic selection (captured by changes in the average type). It can be estimated without imposing the MPH structure, requiring only stationarity of the duration process.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Competing risks.&lt;/strong&gt; The framework in which a price spell can end for multiple distinct reasons — here, ending with a price increase or a price decrease — each with its own hazard function. The paper&amp;rsquo;s GMM approach allows the MPH structure to hold for only a subset of risks and observables, without imposing any structure on the remaining risks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Price trends vs. price reversals.&lt;/strong&gt; A classification of spells based on the direction of the surrounding price changes. Price trends are spells where the direction of the price change at the start and end of the spell is the same (++ or −−), interpreted as regular price changes. Price reversals are spells where the direction switches (e.g., −+, a price decrease followed by a price increase), associated with sales and other temporary price changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Strategic complementarity in pricing (α).&lt;/strong&gt; The degree to which a firm&amp;rsquo;s target price responds to the average price set by other firms. Parameterized by α ∈ [0, 1), where α = 0 yields the exact aggregation result (only the Kaplan-Meier hazard matters) and higher α increases aggregate price stickiness by making firms reluctant to deviate from the average price when few others are adjusting.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic selection.&lt;/strong&gt; The mechanism by which the composition of the pool of surviving price spells shifts toward lower-type (more price-sticky) products as duration increases, because higher-type products change price sooner and exit the pool. This is the source of the gap between the steeply declining Kaplan-Meier hazard and the more modestly declining baseline hazard.&lt;/p&gt;</description></item><item><title>Counterfactual Analysis for Structural Dynamic Discrete Choice Models</title><link>https://macropaperwarehouse.com/papers/counterfactual-analysis-for-structural-dynamic-discrete-choice-models/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/counterfactual-analysis-for-structural-dynamic-discrete-choice-models/</guid><description>&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Discrete choice data identify only &lt;em&gt;differences&lt;/em&gt; in agents&amp;rsquo; utilities, not utility levels. In dynamic discrete choice (DDC) models this means many policy-relevant counterfactuals — those requiring knowledge of utility in levels — are not point-identified. Kalouptsidi, Kitamura, Lima, and Souza-Rodrigues ask: how much can researchers learn about counterfactual outcomes under mild, verifiable restrictions, without imposing the strong normalizations that are standard in applied work but often hard to justify and potentially sign-reversing in their effects?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and Methodology.&lt;/strong&gt; The paper works within a canonical infinite-horizon DDC framework where an agent chooses among a finite action set each period, with additively separable per-period payoffs and i.i.d. unobservables. The econometrician observes conditional choice probabilities (CCPs) and state transition functions from panel data, but the payoff vector is underidentified by X free parameters (one per state), which is the source of non-identification of many counterfactuals. The authors characterize the &lt;em&gt;sharp identified set&lt;/em&gt; for counterfactual CCPs, for low-dimensional outcomes such as average welfare, and develop both identification theory and a feasible inference procedure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Identification Results.&lt;/strong&gt; The sharp identified set for counterfactual CCPs is a smooth, connected manifold whose dimension equals the rank of a specific matrix (CJ*QJ) that the econometrician can compute directly from the data. This rank is at most X minus the number of linearly independent equality restrictions imposed. Two classes of commonly used restrictions reduce the dimension further without requiring full point identification: (i) &lt;em&gt;local counterfactuals&lt;/em&gt; — experiments affecting only a subset of the state-action space — reduce the dimension to at most the number of eigenvalues of the relevant transformation matrix that differ from one; (ii) &lt;em&gt;parametric payoffs&lt;/em&gt; with ηγ free parameters reduce the dimension to at most ηγ. Combining both achieves the tightest bound. Point identification is the special case where the rank equals zero.&lt;/p&gt;
&lt;p&gt;For scalar low-dimensional outcomes (e.g., average welfare), the identified set is a compact interval whose endpoints are obtained by solving constrained optimization programs implementable in standard nonlinear solvers (e.g., Knitro), feasible even when the state space is large.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Illustration.&lt;/strong&gt; In the firm entry/exit Monte Carlo with state space X = 4 and a counterfactual entry subsidy removal: under Restriction 1 alone (outside option = 0, non-negative costs, known variable profits), the identified set for the change in the long-run probability of being active is [-0.1235, 0.0000], correctly signed and containing the true value of -0.0638. Adding shape restrictions (Restrictions 1–2) tightens the upper bound to -0.0341; adding the scrap-value exclusion restriction (Restrictions 1–3) tightens it to -0.0421. Analogous patterns hold for consumer surplus (true: -0.0875; bounds narrowing from [-0.1735, 0.0000] to [-0.1735, -0.0573]) and firm value (true: 0.9513; bounds from [0.0000, 1.8229] to [0.6388, 1.8229]). Critically, the authors show that setting scrap values to zero — the standard identifying assumption — is &lt;em&gt;rejected by the data&lt;/em&gt; under Restrictions 1 and 2, because that payoff vector does not lie in the identified set.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Application.&lt;/strong&gt; Revisiting Das, Roberts, and Tybout (2007) on Colombian exporters, the paper re-examines the horserace among export revenue, fixed cost, and entry cost subsidies. The DRT ranking (revenue subsidies dominate, entry cost subsidies rank last) survives under weaker restrictions than originally imposed, but hinges on the assumption that scrap values do not vary across states. Without that restriction, entry cost subsidies can potentially outperform the other types, reversing the original conclusion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inference.&lt;/strong&gt; The paper develops a subsampling-based inference procedure that is asymptotically uniformly valid (bootstrap fails here due to non-regularity of the set boundary). The confidence set is constructed by inverting a quadratic-form distance test statistic. The critical practical recommendation is subsample size hN = N^{2/3}. The procedure remains feasible in binary choice models with state spaces up to X = 240 (dimension of the optimization problem: 720), where standard moment-inequality approaches are computationally infeasible.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why are many counterfactuals not point-identified in DDC models, even after the model is estimated?&lt;/strong&gt;
A: Choice data identify only differences in value functions across actions, not utility levels. The identifying matrix M has rank AX, leaving X free payoff parameters undetermined. Counterfactuals that depend on utility levels — such as the welfare impact of an entry subsidy when scrap values are unknown — therefore cannot be recovered uniquely from the data, even with a fully estimated model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the key object the paper characterizes, and what does it look like geometrically?&lt;/strong&gt;
A: The paper characterizes the sharp identified set for the counterfactual CCP vector p̃. Proposition 1 establishes that this set is a smooth, connected manifold with boundary, whose interior dimension equals rank(CJ*QJ). Connectedness is important because it means the set has no gaps and boundary tracing is sufficient to characterize it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the dimension of the identified set depend on the type of model restrictions imposed?&lt;/strong&gt;
A: Equality restrictions (d of them) reduce the maximum possible dimension from X to X–d. Local counterfactuals (affecting L state-action pairs) reduce the dimension further to at most the number of eigenvalues of the payoff transformation H(L) that differ from one, which is at most L. Parametric payoffs with ηγ free parameters cap the dimension at ηγ. Combining local counterfactuals with parametric payoffs gives the tightest bound: at most the number of eigenvalues of a related matrix D that differ from one, which is at most min(L, ηγ).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Under what conditions does the identified set for counterfactual behavior collapse to a point?&lt;/strong&gt;
A: When rank(CJ*QJ) = 0, every payoff vector in the identified set PI maps to the same counterfactual CCP — that is, p̃ is point-identified even though the structural payoff π may not be. This can occur through a combination of equality restrictions and specific structure of the counterfactual experiment, without requiring full identification of all model parameters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What properties does the identified set for a scalar low-dimensional outcome have, and how is it computed?&lt;/strong&gt;
A: Under continuity of the outcome function φ and boundedness of the payoff identified set, the identified set for a scalar outcome θ is a compact interval [θL, θU]. The endpoints are computed as the minimum and maximum of a constrained optimization program over the joint space of counterfactual CCPs and payoff vectors, subject to the model&amp;rsquo;s Bellman equations, model restrictions, and equality constraints linking observed to counterfactual behavior. These programs can be solved with standard nonlinear solvers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What do the Monte Carlo results show about the informativeness of the bounds?&lt;/strong&gt;
A: In the firm entry/exit example with X = 4, the identified sets under only mild restrictions (non-negative costs, known variable profits, zero outside option) are already informative and correctly signed. For the change in the probability of being active (true value: -0.0638), the set under Restriction 1 alone is [-0.1235, 0.0000], establishing that the probability does not increase. Adding shape restrictions and exclusion restrictions progressively tightens the interval. All intervals contain the true parameter value, confirming sharpness.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the paper show about the assumption of zero scrap values, which is standard in the entry cost literature?&lt;/strong&gt;
A: The paper shows that setting scrap values to zero can be rejected by the data: in the firm entry/exit example, the payoff vector with s = 0 does not belong to the identified set PI under Restrictions 1 and 2. This is empirically important because Kalouptsidi, Scott, and Souza-Rodrigues (2021) had previously shown that mistakenly setting scrap values to zero not only biases estimated entry costs downward but can also reverse the sign of a subsidy&amp;rsquo;s predicted effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the main finding of the empirical application to export subsidies?&lt;/strong&gt;
A: Revisiting Das, Roberts, and Tybout (2007), the paper finds that the DRT ranking — export revenue subsidies dominate, entry cost subsidies rank last — can be confirmed under restrictions weaker than those DRT originally imposed. However, the ranking is not robust to allowing scrap values to vary across states: under that generalization, entry cost subsidies can potentially outperform the other subsidy types, reversing the original policy conclusion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does the bootstrap fail for inference in this setting, and why does subsampling work?&lt;/strong&gt;
A: The test statistic ĴN(θ0) involves the minimum of a quadratic form over a non-regular (kinked), random, and possibly nonconvex set. Bootstrap critical values are not asymptotically uniformly valid in this non-regular setting. Subsampling with subsample size hN → ∞, hN/N → 0 (the paper recommends hN = N^{2/3}) delivers asymptotically uniformly valid critical values under weak conditions, because it does not require regularity of the constraint set boundary.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the inference approach handle the high dimensionality of DDC settings?&lt;/strong&gt;
A: The paper develops a computational algorithm specifically tailored to the structure of DDC models, exploiting the linear Bellman equation constraints to reduce the effective dimensionality of the optimization problem. In a binary choice model with X = 90, the joint optimization is over a 270-dimensional space; with X = 240 (as in Blundell, Gowrisankaran, and Langer, 2020), the dimension is 720. Standard moment-inequality inference methods (Kaido, Molinari, Stoye, 2019; Bugni, Canay, Shi, 2017) are computationally infeasible at these scales; the authors&amp;rsquo; algorithm remains tractable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the paper relate to Norets and Tang (2014), the closest alternative approach?&lt;/strong&gt;
A: Norets and Tang (2014) partially identify structural parameters and high-dimensional counterfactual CCPs by relaxing the assumed distribution of idiosyncratic shocks, focusing on binary choice models and using a pointwise-valid Bayesian approach. The present paper instead targets low-dimensional policy outcomes (nonlinear functions of payoffs and counterfactual CCPs), accommodates multinomial choice, provides asymptotically uniformly valid frequentist inference via subsampling, and restricts the source of underidentification to the payoff function rather than the error distribution. The two contributions are non-nested and complementary.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the practical workflow the paper enables for applied researchers?&lt;/strong&gt;
A: A researcher can (i) select any combination of model restrictions (equality or inequality, parametric or shape), (ii) specify any counterfactual experiment via an affine payoff transformation (H, g), and (iii) define any low-dimensional outcome of interest φ, then directly compute the identified set and a valid confidence interval by solving two constrained optimization programs — without deriving new analytical identification results for each specification. The rank condition for checking the dimension of the identified set is computable from the data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dynamic Discrete Choice (DDC) Model.&lt;/strong&gt; A discrete-time infinite-horizon model where agents choose among a finite action set each period, with per-period utilities additively separable into an observed payoff function π and an i.i.d. unobservable shock, and agents maximize expected discounted lifetime utility. The model is parameterized by payoffs π, transition function F, discount factor β, and shock distribution G.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conditional Choice Probability (CCP).&lt;/strong&gt; The probability that an agent selects a given action in a given state, integrating out the unobservable shocks. CCPs and state transitions are directly identifiable from panel data and serve as the sufficient statistics for the identified set, in place of the unidentified payoff vector.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sharp Identified Set for Counterfactual CCPs.&lt;/strong&gt; The set PĨ(p, F) of all counterfactual CCP vectors p̃ that are consistent with the observed data (p, F) and the imposed model restrictions, given the specified counterfactual transformation. Characterized as a smooth connected manifold with dimension equal to rank(CJ*QJ).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Local Counterfactual.&lt;/strong&gt; A counterfactual experiment in which the payoff transformation H modifies only a subset L of the state-action pairs, leaving the rest unchanged. Local counterfactuals reduce the dimension of the identified set relative to global experiments, because only the payoffs in the affected subset matter for the unidentified component of the counterfactual response.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Partial Identification / Identified Set for Outcomes.&lt;/strong&gt; Rather than seeking a unique estimate of a counterfactual outcome θ, partial identification recovers the set ΘI of all values of θ consistent with the data and restrictions. For scalar outcomes this is a compact interval [θL, θU] whose endpoints solve constrained optimization problems over payoff and counterfactual CCP spaces.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Subsampling Inference.&lt;/strong&gt; A procedure for constructing asymptotically uniformly valid confidence sets by repeatedly computing the test statistic on subsamples of size hN &amp;lt; N, approximating the sampling distribution of ĴN(θ0) without requiring regularity (smoothness) of the boundary of the constraint set — a requirement that fails here due to the kinked, nonconvex nature of the identified set.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rank Condition for Dimension.&lt;/strong&gt; The dimension of the identified set for counterfactual CCPs is determined by the rank of the matrix CJ*QJ, which depends on the counterfactual transformation H, the model restrictions, and the observed data. The econometrician can compute this rank from observables to assess, before imposing any strong assumptions, how many dimensions of freedom remain in the identified set.&lt;/p&gt;</description></item><item><title>Customer accumulation, returns to scale, and secular trends</title><link>https://macropaperwarehouse.com/papers/customer-accumulation-returns-to-scale-and-secular-trends/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/customer-accumulation-returns-to-scale-and-secular-trends/</guid><description>&lt;p&gt;This paper asks how rising returns to scale in production contributed to three concurrent U.S. secular trends since 1980: declining business dynamism, rising markups, and growing firm expenditures on customer acquisition. The author constructs a firm dynamics model in the Hopenhayn (1992) tradition with endogenous entry and exit, heterogeneous markups, and customer accumulation grounded in directed search in the product market. Firms compete for customers through both prices and selling activities; larger firms gain a competitive edge when returns to scale rise because their marginal costs fall more than those of smaller firms—even though the technological shift is uniform across firms. This demand-based channel triggers winners-and-losers dynamics and the rise of superstar firms.&lt;/p&gt;
&lt;p&gt;The empirical foundation rests on Compustat data for U.S. publicly traded firms (1977–2014) and Business Dynamics Statistics (BDS) for aggregate and sector-level dynamism measures. Production-function estimation using Ackerberg, Caves, and Frazer (2015) augmented with sales-share controls documents that aggregate returns to scale rose from approximately 1.0 in 1980 to approximately 1.05 by 2014—a within-sector increase, not a reallocation effect. Over the same period, the cost-weighted markup rose by 42%, the firm entry rate fell by 33%, the excess reallocation rate fell by 29%, and selling costs relative to production costs rose by 60%–90% depending on the measure used.&lt;/p&gt;
&lt;p&gt;The model is calibrated to 1980 steady-state moments (firm life-cycle patterns, markups, entry and reallocation rates). A 5% increase in returns to scale—matching the empirical estimate—accounts for: a +15 percentage point rise in the average cost-weighted markup (vs. +42% in the data); a 33% decline in the entry rate (exactly matching the data); a 21% decline in the reallocation rate (vs. 29% in the data); and a 23% increase in selling costs relative to production costs (vs. 60%–90% in the data). The model also generates a 53% rise in the share of firms aged 11 years or older (vs. 50% in the data) and a 58% decline in the employment share of firms aged 5 years or younger (vs. 56% in the data), closely tracking the aging of the U.S. firm population. Firm-level responsiveness to productivity shocks declines by 0.08 in the model, versus about 0.01 in Compustat and 0.09 in Decker et al. (2020).&lt;/p&gt;
&lt;p&gt;Sector-level panel regressions with sector fixed effects confirm the model&amp;rsquo;s directional predictions: within-sector increases in returns to scale are associated with lower entry rates (coefficient −2.89, significant at 1%), lower reallocation rates (−1.16, significant at 1%), higher markups (+3.15, significant at 1%), and higher selling costs relative to production costs (+1.85 for the advertising-based measure; +8.52 for adjusted SG&amp;amp;A).&lt;/p&gt;
&lt;p&gt;A key scope condition is that the model yields a constrained-efficient allocation: directed search and full internalization of returns to scale imply decentralized equilibrium efficiency, making the paper a laboratory for assessing how far efficient firm responses to technological change can explain the secular trends without invoking market failures. The model fits the post-2000 transition dynamics better than the 1980s–1990s period, and explains a substantial but incomplete share of the trends, suggesting complementary—possibly inefficient—forces also contributed.&lt;/p&gt;
&lt;p&gt;Q: What is the core mechanism through which rising returns to scale generate winners-and-losers dynamics?&lt;/p&gt;
&lt;p&gt;A: The marginal cost of production under increasing returns to scale (alpha &amp;gt; 1) is MC(z,n) = l(n,z)^(1−alpha) × (1/alpha) × (W/e^z), which depends on firm size l(n,z). A uniform rise in alpha rotates the marginal cost schedule clockwise by firm size: larger firms see a proportionally larger cost reduction than smaller firms, even though the technological change is identical across all firms. Because firms compete for the same pool of customers, this asymmetric cost advantage allows large firms to offer lower prices while sustaining higher margins, attracting customers away from small firms. The result is a demand-based channel that generates winners-and-losers dynamics and increases market concentration.&lt;/p&gt;
&lt;p&gt;Q: How does the model capture customer accumulation, and why is it central to the paper&amp;rsquo;s argument?&lt;/p&gt;
&lt;p&gt;A: The model introduces directed search in the product market, where firms post advertisements and customers—including those already matched with a firm—choose which submarket to enter by trading off offered utility against matching probability. A constant-returns-to-scale matching function governs match creation; in submarket with tightness theta, customers match with probability m(theta) = theta(1+theta)^(−1) and firms attract customers with probability q(theta) = (1+theta)^(−1). The customer accumulation motive creates an investment-harvest trade-off: firms can either post high promised utility (low prices) to grow their customer base or extract surplus through high prices. Rising returns to scale amplify large firms&amp;rsquo; ability to resolve this trade-off favorably, linking the technological change directly to markup dynamics, entry incentives, and selling expenditures.&lt;/p&gt;
&lt;p&gt;Q: What is the directed search framework&amp;rsquo;s role in ensuring equilibrium uniqueness and efficiency?&lt;/p&gt;
&lt;p&gt;A: The author introduces firm-side commitment contracts—specifying price, separation probability, and continuation utility contingent on productivity realizations—combined with directed search. Because search is directed on both sides and firms fully internalize returns to scale, the decentralized equilibrium is constrained-efficient. This delivers uniquely determined heterogeneous prices in equilibrium (solving the indeterminacy problem common in customer-market models) and establishes the paper&amp;rsquo;s efficient-mechanism benchmark: it tests how far profit-maximizing firm responses to technological change—without any market failure—can account for the secular trends.&lt;/p&gt;
&lt;p&gt;Q: How are prices structured in the model, and what life-cycle pattern do they generate?&lt;/p&gt;
&lt;p&gt;A: Each firm charges two distinct prices in each period: one to incumbent customers (the same for all incumbents, since they are identical conditional on being attached to the same firm) and one to newly acquired customers (which varies based on the promised utility in the submarket searched). Firms that are expanding their customer base offer greater promised utility and therefore charge lower prices to attract customers; firms harvesting their existing base charge higher prices. Because firms enter small and grow, this dynamic generates a price life cycle: young firms invest via low prices and mature firms harvest through higher prices, which the model reproduces as a rising markup pattern over the firm life cycle—an untargeted moment the model fits well.&lt;/p&gt;
&lt;p&gt;Q: What does the calibration target and what untargeted moments does the model reproduce?&lt;/p&gt;
&lt;p&gt;A: The model is calibrated to 1980 using: the number of employees of entrant firms (pinning entry customer base n_e), employees of age-5 firms (pinning convex cost chi_1), share of firms aged 11+ years (pinning chi_2), average firm size (operating cost f), entry rate (entry cost kappa), excess reallocation rate (exit shock delta), and average cost-weighted markup (linear cost c). Untargeted moments reproduced include: a sales-weighted markup of 0.28 (vs. 0.25 in De Loecker et al. 2020), endogenous customer turnover of approximately 9% (vs. 15% in Gourio and Rudanko 2014), and an elasticity of customer base shrinkage to price of 0.08 (within the 0.01–0.16 range from Paciello et al. 2019). The model also matches markup and selling-cost life-cycle patterns that are typically overlooked.&lt;/p&gt;
&lt;p&gt;Q: How large is the quantitative contribution of the 5% rise in returns to scale to each secular trend?&lt;/p&gt;
&lt;p&gt;A: Comparing the 1980 steady state (alpha = 1) to the 2014 steady state (alpha = 1.05): the average cost-weighted markup rises by 15% in the model versus 42% in the data; the entry rate declines by 33% in the model, exactly matching the data; the reallocation rate declines by 21% in the model versus 29% in the data; and selling costs relative to production costs rise by 23% in the model versus 60%–90% in the data. The model thus explains a substantial share of each trend while leaving a residual requiring additional mechanisms.&lt;/p&gt;
&lt;p&gt;Q: How does the model explain the aging of U.S. firms, and how well does it match the data?&lt;/p&gt;
&lt;p&gt;A: The winners-and-losers mechanism shifts activity toward larger, older firms, which mechanically ages the firm population. The model generates a 53% increase in the share of firms aged 11 years or older (vs. 50% in the data) and a 58% decline in the employment share of firms aged 5 years or younger (vs. 56% in the data). This aging arises because rising returns to scale increase the cost of customer acquisition, acting as a barrier to entry that disproportionately hurts new, small firms while allowing large incumbents to remain viable at lower productivity thresholds.&lt;/p&gt;
&lt;p&gt;Q: What is the channel through which rising returns to scale reduce business dynamism specifically?&lt;/p&gt;
&lt;p&gt;A: The unequal reduction in marginal costs intensifies competition for customers and raises customer acquisition costs. This operates through two simultaneous effects on the exit threshold: (i) lower marginal costs allow large firms to remain viable at lower productivity levels despite higher customer acquisition costs; and (ii) heightened competition forces smaller firms to require higher productivity to survive in a market that has become increasingly costly to operate in. Higher customer acquisition costs therefore function as an endogenous barrier to entry, reducing the entry rate and the reallocation of resources across firms.&lt;/p&gt;
&lt;p&gt;Q: Does the model attribute the secular trends entirely to efficient firm behavior, and what does it conclude about residual explanations?&lt;/p&gt;
&lt;p&gt;A: No. The model is explicitly designed as a constrained-efficient benchmark, and the paper finds that while rising returns to scale account for a substantial share of the trends—particularly in magnitude—the transition dynamics show a less accurate fit before the 2000s. The author concludes that complementary mechanisms, likely involving inefficiencies (such as market power from horizontal product differentiation or barriers to entry beyond those captured by the model), played a significant role in the earlier evolution of these trends and in the portion of the trends not explained by the efficient channel.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports the rising returns to scale finding, and what are its limitations?&lt;/p&gt;
&lt;p&gt;A: Production-function estimation using the Ackerberg-Caves-Frazer method with sales-share controls on Compustat data shows returns to scale rising from approximately 1.0 in 1980 to approximately 1.05 by 2014, driven primarily by within-sector increases rather than reallocation toward high-returns sectors. A translog production function finds limited evidence of heterogeneous increases across firm sizes within Compustat. However, Compustat predominantly covers large publicly traded firms; smaller firms outside the sample may have experienced minimal or no increase in returns to scale. If technology adoption involves fixed costs, the aggregate impact could be larger than estimated, meaning the quantitative exercises likely represent a conservative lower bound.&lt;/p&gt;
&lt;p&gt;Q: How does the paper relate to and extend the directed search literature in product markets?&lt;/p&gt;
&lt;p&gt;A: The paper builds on Gourio and Rudanko (2014) and Roldan-Blanco and Gilbukh (2020), where customers are locked in once matched, by introducing labor-search tools from Schaal (2017) to allow: (i) incumbent customer switching between firms at rates of 10%–25% annually (Gourio and Rudanko 2014), and (ii) a non-zero price sensitivity of incumbent customers (Paciello et al. 2019). It also allows firms to invest in demand through selling expenditures, which prior directed search models in product markets typically abstracted from, making it possible to study how technological changes affect customer reallocation and firms&amp;rsquo; cost structures jointly.&lt;/p&gt;
&lt;p&gt;Customer capital: The stock of customers a firm has accumulated through prior selling and pricing decisions; treated as a state variable that firms invest in (by offering low prices and spending on advertisements) or harvest from (by charging high markups), with a customer turnover rate estimated at 10%–25% annually in the literature.&lt;/p&gt;
&lt;p&gt;Directed search in the product market: A market structure in which both firms and customers choose which submarket (indexed by the promised utility level) to enter, trading off match probability against terms; delivers constrained-efficient equilibrium and uniquely determined heterogeneous prices.&lt;/p&gt;
&lt;p&gt;Investment-harvest trade-off: The firm&amp;rsquo;s dynamic choice between offering high promised utility (low prices, low current markups) to grow the customer base versus extracting surplus through high prices from an existing customer base; shaped by the firm&amp;rsquo;s current size, productivity, and the cost structure implied by returns to scale.&lt;/p&gt;
&lt;p&gt;Returns to scale (alpha): The curvature of the production function y = e^z × l^alpha; equals 1.0 under constant returns and approximately 1.05 by 2014 in the empirical estimates; the paper&amp;rsquo;s central technological change parameter, whose rise disproportionately reduces marginal costs for larger firms.&lt;/p&gt;
&lt;p&gt;Winners-and-losers dynamics: The reallocation of customers and market share from small to large firms triggered by the asymmetric cost advantage large firms obtain when returns to scale rise; the demand-based channel through which superstar firms emerge.&lt;/p&gt;
&lt;p&gt;Cost-weighted markup: The average markup aggregated using each firm&amp;rsquo;s costs as weights, as opposed to sales-weighted markup; the primary measure of market power used in the paper, rising by 42% in the data between 1980 and 2014.&lt;/p&gt;
&lt;p&gt;Constrained-efficient allocation: An equilibrium outcome in which, given the frictions present (search-and-matching in the product market), no social planner operating under the same constraints could improve welfare; the paper uses this as a benchmark to assess how far efficient firm responses explain secular trends without invoking market failures.&lt;/p&gt;
&lt;p&gt;Selling costs relative to production costs: The ratio of customer acquisition expenditures (advertising or adjusted SG&amp;amp;A) to cost of goods sold; rose by 60%–90% in the data between 1980 and 2014 and by 23% in the model&amp;rsquo;s steady-state comparison.&lt;/p&gt;</description></item><item><title>Customer Acquisition, Business Dynamism and Aggregate Growth</title><link>https://macropaperwarehouse.com/papers/customer-acquisition-business-dynamism-and-aggregate-growth/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/customer-acquisition-business-dynamism-and-aggregate-growth/</guid><description>&lt;p&gt;This paper asks whether firm-level customer acquisition — distinct from productivity differences — is a quantitatively important driver of aggregate economic growth, and whether ignoring it distorts predictions about growth policy efficacy. The authors build a novel endogenous growth model in which innovating firms must first accumulate customers to sell their products, with two channels of customer acquisition operating simultaneously: costly sales-and-marketing expenditure and below-static-markup pricing (sales-driven accumulation). The model is estimated using indirect inference against a combination of aggregate data (U.S. real GDP per worker growth of 1.43% annually, 1979–2019), Business Dynamics Statistics (BDS) life-cycle profiles, and firm-level data from Compustat matched to Capital IQ&amp;rsquo;s sales-and-marketing expense records covering 1997–2019.&lt;/p&gt;
&lt;p&gt;The benchmark model yields four closed-form propositions. First, a &amp;ldquo;firm-level market size effect&amp;rdquo;: higher customer retention raises a firm&amp;rsquo;s future profit base, strengthening incentives to conduct R&amp;amp;D. Second, an endogenous feedback loop: more productive firms invest more in customer acquisition, which expands their customer base and further strengthens R&amp;amp;D incentives. Third, customer base accumulation raises aggregate growth, but only indirectly — by boosting firm-level innovation rates — since aggregate productivity is a customer-weighted average of firm productivity levels. Fourth, the sensitivity of innovation to R&amp;amp;D subsidies increases with customer base growth, because firms with faster-growing customer bases discount future profits less steeply.&lt;/p&gt;
&lt;p&gt;In the quantitatively estimated full model — which relaxes the benchmark&amp;rsquo;s perfect-scaling restrictions and endogenizes firm entry and exit — the authors conduct two decomposition exercises. In a counterfactual scenario where expected customer retention is reduced to make average customer base growth zero among continuing businesses, firm-level innovation rates fall by approximately 40% relative to the full model. Of this 40% decline, only about 6 percentage points are attributable to the direct firm-level market size effect alone; the vast majority is driven by the endogenous feedback loop between innovation and customer acquisition. In a second decomposition focused on aggregate growth, the firm-level market size effect and a reallocation effect — whereby the feedback loop concentrates customers among high-productivity firms — together account for 44% of aggregate growth in the full model.&lt;/p&gt;
&lt;p&gt;On policy, the authors compare R&amp;amp;D subsidies and operational subsidies in the full model against an otherwise identical model that ignores customer accumulation. R&amp;amp;D subsidies are approximately twice as effective at boosting aggregate growth in the full model as in the model without customer accumulation. Conversely, operational subsidies produce a stronger decline in aggregate growth in the full model than in the benchmark-without-customer-accumulation, because aggregate growth in the full model is a customer-weighted average of firms&amp;rsquo; productivity growth rates, making the joint distribution of productivity and customer bases the relevant object of study.&lt;/p&gt;
&lt;p&gt;Firm-level data support three empirical predictions. Marketing expenditure, R&amp;amp;D intensity, and markups co-move in model-consistent directions both contemporaneously and over the life cycle. The estimated relative weight of marketing versus pricing as channels of customer accumulation is γ = 0.745, indicating marketing is the dominant channel. A model-consistent proxy for the severity of customer-base frictions, estimated in the cross-section of industries, shows that stronger frictions correlate with lower R&amp;amp;D investment, as predicted. The customer-base depreciation rate is estimated at ζ = 0.375, R&amp;amp;D cost scaling at σx = 1.264, and marketing cost scaling at σa = 1.405.&lt;/p&gt;
&lt;p&gt;Q: What is the firm-level market size effect and why does it arise?
A: When a firm retains more customers, successful innovations apply to a larger market, raising the profitability of each unit reduction in production costs. This increases the marginal benefit of R&amp;amp;D investment. In the benchmark model, Proposition 2(a) shows formally that firm-level innovation increases with customer base growth: ∂x/∂(1−ζ) &amp;gt; 0, where ζ is the customer separation rate.&lt;/p&gt;
&lt;p&gt;Q: What is the endogenous feedback loop between innovation and customer accumulation?
A: More productive firms have lower production costs and can therefore afford greater investment in marketing and can set lower markups, both of which attract more customers. A larger customer base raises firm value and strengthens R&amp;amp;D incentives further (Proposition 2(b)). This bidirectional feedback means that productivity growth and customer accumulation are jointly determined in equilibrium, not independent processes.&lt;/p&gt;
&lt;p&gt;Q: How large is the quantitative effect of customer accumulation on firm-level innovation?
A: In the counterfactual where expected customer retention is reduced so that average customer base growth among continuing firms is zero, firm-level innovation rates are approximately 40% lower than in the full model. Of this, only about 6% (of the total drop) is attributable to the direct market size effect in isolation; the feedback loop accounts for the remaining roughly 34 percentage points.&lt;/p&gt;
&lt;p&gt;Q: How much of aggregate growth do customer-acquisition channels explain?
A: The firm-level market size effect and a customer reallocation effect together account for 44% of aggregate growth in the full model. The firm-level market size effect alone reduces aggregate growth by about one-fifth (20%) in the relevant counterfactual. The reallocation effect — by which productive firms accumulate disproportionate market share — contributes the remainder of the 44%.&lt;/p&gt;
&lt;p&gt;Q: What is the reallocation channel for aggregate growth?
A: Because highly productive firms can invest more in customer acquisition, the feedback loop endogenously concentrates customers (market shares) among high-productivity firms. Since aggregate productivity in the model is a customer-weighted average of firm productivity levels (equation 16), this reallocation raises aggregate productivity growth beyond what the firm-level R&amp;amp;D incentive effect alone would produce.&lt;/p&gt;
&lt;p&gt;Q: How does customer accumulation change the efficacy of R&amp;amp;D subsidies?
A: R&amp;amp;D subsidies are approximately twice as effective at raising aggregate growth in the full model (with customer accumulation) as in an otherwise identical model that ignores customer accumulation. The mechanism is Proposition 4(b): faster customer base growth makes firms weight future profits more heavily, increasing their sensitivity to any change in R&amp;amp;D costs, including that brought about by a government subsidy.&lt;/p&gt;
&lt;p&gt;Q: What happens to aggregate growth under operational subsidies in the two models?
A: Operational subsidies lead to a stronger decline in aggregate growth in the full model than in the model without customer accumulation. The reason is that aggregate growth in the full model depends on the joint distribution of firm productivity and customer bases; operational subsidies alter this distribution in ways that reduce the customer-weighted average of productivity growth rates, an effect absent when customer accumulation is ignored.&lt;/p&gt;
&lt;p&gt;Q: How are the two customer-acquisition channels (marketing and pricing) measured empirically?
A: Marketing is measured using sales-and-marketing expenses from Capital IQ, available for 48% of the Compustat sample (34% report directly; an additional 14% report advertising or marketing sub-components). Markups are measured following De Loecker et al. (2020) as the inverse share of variable costs in sales multiplied by the cost-output elasticity, with variation across firms identified from balance sheet data under the assumption that cost-output elasticities are constant within industry-year cells.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated relative strength of marketing versus pricing in customer accumulation?
A: The relative weight on marketing is γ = 0.745, estimated by targeting the coefficient βµ = 0.04 (standard error 0.01) from a reduced-form regression of firm-level sales growth on changes in markups (equation 29). This implies that marketing is the dominant channel, consistent with evidence in Afrouzi et al. (2021) and Fitzgerald et al. (forthcoming).&lt;/p&gt;
&lt;p&gt;Q: What is the estimated customer-base depreciation rate and how is it disciplined?
A: The depreciation rate ζ is estimated at 0.375, targeted to match average firm-level employment growth from the BDS. This falls toward the lower end of existing estimates, which range from about 0.3 to 0.7 across studies.&lt;/p&gt;
&lt;p&gt;Q: How do R&amp;amp;D costs scale with firm size in the estimated model?
A: The R&amp;amp;D cost scaling parameter is σx = 1.264, estimated by targeting the reduced-form coefficient of −0.01 from a regression of log R&amp;amp;D intensity on log sales with industry-time fixed effects (equation 28). This is close to the estimate in Akcigit and Kerr (2018).&lt;/p&gt;
&lt;p&gt;Q: How do marketing costs scale with firm size?
A: The marketing cost scaling parameter is σa = 1.405, estimated by targeting a reduced-form coefficient of −0.01 from a regression of log sales-and-marketing intensity on log sales with industry-time fixed effects (equation 30).&lt;/p&gt;
&lt;p&gt;Q: What empirical co-movement evidence supports the model&amp;rsquo;s predictions?
A: In the cross-section of firms, marketing expenditure, R&amp;amp;D intensity, and markups all co-move in model-predicted directions, for both static (contemporaneous) relationships and dynamic (life-cycle) patterns. Additionally, a model-consistent industry-level proxy for the severity of customer-base frictions shows that stronger frictions are associated with lower R&amp;amp;D investment, as the model predicts.&lt;/p&gt;
&lt;p&gt;Q: How does endogenous firm exit work in the full model and why does it differ from standard models?
A: Firms pay a stochastic per-period operational cost and exit when that cost exceeds a threshold κ*_j = v(q_j, b_j)/W. Unlike standard growth models where exit depends only on productivity, here the exit threshold depends on both productivity and accumulated customers, so customer loss can trigger exit even for relatively productive firms.&lt;/p&gt;
&lt;p&gt;Q: What data sources are used and what are their key limitations?
A: The three primary firm-level sources are the Census Bureau&amp;rsquo;s BDS (broad coverage, employment-focused), Compustat (rich financial data but limited to publicly traded firms and lacking direct customer-acquisition measures), and Capital IQ (sales-and-marketing expenses available from 1997, matched to 91% of the Compustat sample). To address Compustat&amp;rsquo;s non-representativeness, employment-based weights aligning Compustat and BDS firm-size distributions are applied when computing model moments against Compustat targets.&lt;/p&gt;
&lt;p&gt;Firm-level market size effect: The mechanism by which higher customer retention raises a firm&amp;rsquo;s future profit base — because lower production costs from successful innovation apply to a larger market — thereby strengthening incentives to conduct R&amp;amp;D. This is the primary channel linking customer accumulation to innovation.&lt;/p&gt;
&lt;p&gt;Customer base (b_j): The mass of household members consuming a firm&amp;rsquo;s product variety, which varies endogenously across firms. It enters demand directly (equation 4) and serves as a state variable in the firm&amp;rsquo;s value function alongside productivity.&lt;/p&gt;
&lt;p&gt;Endogenous feedback loop: The bidirectional reinforcement between productivity growth and customer accumulation. More productive firms invest more in customers; a larger customer base raises the value of innovation; higher innovation raises productivity further.&lt;/p&gt;
&lt;p&gt;Reallocation effect: The concentration of customers (market shares) toward high-productivity firms that arises endogenously from the feedback loop, contributing to aggregate growth because aggregate productivity is a customer-weighted average of firm-level productivity.&lt;/p&gt;
&lt;p&gt;Customer-base depreciation rate (ζ): The exogenous rate at which a firm loses its existing customers each period, estimated at 0.375 in the paper&amp;rsquo;s calibration. It governs the baseline speed of customer attrition and is the key parameter for the firm-level market size effect.&lt;/p&gt;
&lt;p&gt;Sales-and-marketing expenses: Expenditures on sales force, brand development, customer service, advertising, and customer data acquisition — measured from Capital IQ — that directly drive marketing-based customer accumulation (the dominant channel with estimated weight γ = 0.745).&lt;/p&gt;
&lt;p&gt;Perfect scaling (Assumption 1): The benchmark restriction that R&amp;amp;D and marketing costs, and the sales-driven customer accumulation benefit, all scale one-for-one with a composite of firm productivity and customer base. This assumption enables closed-form solutions and is relaxed in the full model using estimated scaling parameters.&lt;/p&gt;</description></item><item><title>Diversification, Market Entry, and the Global Internet Backbone</title><link>https://macropaperwarehouse.com/papers/diversification-market-entry-and-the-global-internet-backbone/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/diversification-market-entry-and-the-global-internet-backbone/</guid><description>&lt;p&gt;This paper investigates how buyer demand for supplier diversification shapes entry incentives and market structure, using the global undersea fiber-optic cable industry as the empirical setting. The research question has two parts: first, how much of observed cable entry and surplus generation is attributable to buyers&amp;rsquo; diversification motives rather than standard price competition; and second, whether market forces produce too much or too little diversification relative to the social optimum.&lt;/p&gt;
&lt;p&gt;The empirical setting spans 2005–2021 and covers the worldwide network of undersea cables that carries more than 98% of all international internet traffic. Cables fail frequently — hundreds of faults per year — and industry professionals confirm that &amp;ldquo;no customer would buy capacity on a single cable.&amp;rdquo; The median monthly price for a 10Gbps lease fell from $55,500 in 2005 to $2,200 in 2021, and the number of active cables roughly doubled over the sample period.&lt;/p&gt;
&lt;p&gt;The authors use proprietary data from TeleGeography covering cable characteristics (construction costs, capacity, landing points, entry dates), quarterly bandwidth prices at the city-pair level, annual used bandwidth at the country-pair level, and 168 documented cable faults. Markets are defined as country-pairs in calendar quarters.&lt;/p&gt;
&lt;p&gt;The theoretical model begins with a representative buyer who splits bandwidth purchases equally across n symmetric cable operators to minimize expected disruption costs. Because disruption shocks are i.i.d. across cables, adding suppliers reduces the variance of realized bandwidth delivery, lowering the required over-provisioning buffer. This generates a &amp;ldquo;market expansion&amp;rdquo; channel: entry increases aggregate demand holding prices fixed, not just through price competition. The aggregate demand equation takes log-linear form with cable count indicators alongside price and demand shifters.&lt;/p&gt;
&lt;p&gt;The structural model adds a dynamic oligopoly game where firms make entry and exit decisions as a non-stationary Markov Perfect Equilibrium, with Cournot competition in each period. The three-step estimation procedure recovers: (1) price elasticities and diversification parameters from an IV demand regression using electricity generation cost shares as instruments; (2) marginal costs from firms&amp;rsquo; first-order conditions; (3) entry and fixed costs from a nested pseudo-likelihood (NPL) estimator, supplemented by construction cost data to separately identify entry costs given the near-absence of observed exits.&lt;/p&gt;
&lt;p&gt;Key demand results: the IV price elasticity is −1.36. The market expansion effect is large and exhibits decreasing marginal returns — entry of a second cable expands demand by as much as a 28.3% price decrease; a third cable is equivalent to a 19.3% price decrease; an eighth cable is equivalent to a 7.5% price decrease. The demand model achieves R² = 95%.&lt;/p&gt;
&lt;p&gt;The first counterfactual removes the diversification channel entirely (entry raises competition only). Without diversification, cable investment falls by 12%. The net present value of total surplus per market over the sample period averages $1.11 billion under the observed equilibrium; supplier diversification accounts for 11% of total surplus and 27% of consumer surplus.&lt;/p&gt;
&lt;p&gt;The second counterfactual quantifies two opposing distortions relative to the social optimum. Business-stealing creates excessive entry (entrants reduce incumbents&amp;rsquo; output), while diversity effects create insufficient entry (marginal entrants generate surplus through diversification they cannot fully capture). At end-of-sample (2021-Q4), diversity distortions in terms of number of entrants range from 54% to 125% of the business-stealing distortion. Business-stealing tends to dominate for most markets, producing moderately excessive entry. Relative to the market outcome, total surplus under the social planner&amp;rsquo;s solution is on average 10% higher: 53% of this welfare gap is attributable to diversity effects and 47% to business-stealing effects. These findings hold across market heterogeneity in entry costs, market size, and demand growth.&lt;/p&gt;
&lt;p&gt;The paper concludes that profit-maximizing suppliers fail to fully internalize diversification-related social benefits, and that targeted entry subsidies would pass cost-benefit tests in settings where diversity distortions dominate.&lt;/p&gt;
&lt;p&gt;Q: What is the core mechanism by which supplier diversification expands demand?
A: When buyers split purchases across n cable operators whose disruption shocks are i.i.d., adding a supplier reduces the variance of realized delivered bandwidth. The buyer therefore needs to hold a smaller over-provisioning buffer to achieve the same expected level of used bandwidth B. This lowers the effective cost of a given quantity of used bandwidth, shifting the aggregate demand curve outward. As the number of suppliers grows to infinity, the expected disruption cost converges to zero.&lt;/p&gt;
&lt;p&gt;Q: How large is the market-expansion effect of diversification empirically?
A: The effect is large but exhibits decreasing marginal returns. Entry of a second cable expands demand by as much as a 28.3% price reduction holding prices fixed; the third cable is equivalent to a 19.3% price reduction; and the eighth cable is equivalent to a 7.5% price reduction. All cable-count coefficients are positive and statistically significant in the IV demand model.&lt;/p&gt;
&lt;p&gt;Q: How is price endogeneity addressed in the demand estimation?
A: Bandwidth prices are instrumented using the marginal cost of electricity generation — specifically, country-level electricity generation shares (coal, gas, oil) interacted with quarterly commodity price series for coal, gas, and oil (Brent crude, Australian coal price, EU natural gas price). The first-stage results indicate electricity costs are strong predictors of bandwidth prices. Accounting for endogeneity raises the price elasticity from an OLS level to −1.36 in absolute value, consistent with the expected direction of OLS bias.&lt;/p&gt;
&lt;p&gt;Q: What share of cable investment and surplus is attributable to diversification motives?
A: In the counterfactual where the diversification channel is eliminated — entry raises competition and lowers prices but provides no diversification benefit — cable investment falls by 12%. Under the observed equilibrium, the net present value of total surplus per market over 2005–2021 averages $1.11 billion; supplier diversification accounts for 11% of this total surplus and 27% of consumer surplus.&lt;/p&gt;
&lt;p&gt;Q: How are the two distortions — business-stealing and diversity — defined and separated?
A: Business-stealing distortion arises because entrants reduce incumbents&amp;rsquo; outputs and revenues, so private entry benefits exceed social benefits, leading to excessive entry. Diversity distortion arises because entrants create surplus for buyers through diversification but cannot fully capture it without perfect price discrimination (following Spence (1976) and Mankiw and Whinston (1986)), leading to insufficient entry. The authors disentangle these by comparing: (i) the social planner&amp;rsquo;s solution (eliminates both distortions), and (ii) a coordinated entry solution maximizing producer surplus (eliminates only business-stealing). The residual gap between the two identifies the diversity distortion.&lt;/p&gt;
&lt;p&gt;Q: What is the net direction and magnitude of distortion in equilibrium market structure?
A: At 2021-Q4, for most markets, business-stealing dominates, leading to moderately excessive entry. Diversity distortions in number of entrants range from 54% to 125% of the business-stealing distortion across markets. Relative to the market outcome, the social planner&amp;rsquo;s solution yields average total surplus that is 10% higher. Of that welfare gap, 53% is attributable to diversity effects and 47% to business-stealing effects.&lt;/p&gt;
&lt;p&gt;Q: How do market characteristics affect which distortion dominates?
A: The paper analyzes cross-market heterogeneity and identifies market features — including the size of entry costs, market size, and the rate of demand growth over time — as determinants of whether insufficient diversification or excessive entry is the binding distortion. Markets with higher entry costs or slower demand growth are more likely to exhibit insufficient diversification.&lt;/p&gt;
&lt;p&gt;Q: How are entry costs identified given the near-absence of cable exits in the data?
A: Because exit events are rare in a nascent industry — only a handful of exits observed, mostly after 2020 — entry and fixed costs cannot be separated by exit decisions alone. The authors address this by using cable-level construction cost data from TeleGeography to estimate entry costs outside the dynamic model. With entry costs in hand, firms&amp;rsquo; optimal entry decisions identify fixed costs. Scrap values are normalized to zero, consistent with industry reports that retired cables are typically abandoned on the seabed.&lt;/p&gt;
&lt;p&gt;Q: What role does the non-stationarity of the market environment play in the model?
A: The data covers the industry&amp;rsquo;s earliest growth phase, with demand growing by roughly three orders of magnitude (used bandwidth from 5 Tbps in 2005 to 2,886 Tbps in 2021) and prices falling by a factor of roughly 25. The authors use a non-stationary Markov Perfect Equilibrium concept in which strategies and transition functions are indexed by time, aligning with the treatment of high-tech commodities in Igami (2017).&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: Because profit-maximizing suppliers do not fully internalize the diversification-related social benefits of entry, entry rates can be sub-optimal from a welfare perspective when diversity distortions dominate. The authors suggest targeted entry subsidies would pass cost-benefit tests in such cases. For antitrust analysis, regulators who ignore the demand-expansion effect of incremental suppliers may incorrectly judge a market as sufficiently competitive. In merger review, authorities must account for firms&amp;rsquo; private incentives to provide diversification to reach accurate welfare conclusions.&lt;/p&gt;
&lt;p&gt;Q: How does the paper verify that diversification demand is not a spurious empirical artifact?
A: Several checks support the causal interpretation. The estimated demand parameters are consistent with the predictions of the consumer-level utility maximization problem derived analytically: decreasing marginal returns to diversification and a positive relationship between the number of suppliers and demand. The demand model achieves R² = 95%, suggesting limited unobserved confounders. Additionally, 78% of cable faults involve only a single cable, confirming that disruptions are geographically isolated and that cross-cable diversification provides genuine insurance value.&lt;/p&gt;
&lt;p&gt;Q: What are the main data limitations acknowledged by the authors?
A: The authors cannot observe cable-level revenue or market shares, nor contracts between buyers and sellers; only aggregate country-pair used bandwidth is observed. Price coverage is not comprehensive — TeleGeography collects prices on a voluntary basis from dozens of providers. The cable faults dataset (168 faults) represents only a subset of total faults, as collection focuses on publicly disclosed events. The demand model also does not explicitly account for substitution patterns across firms due to lack of firm-level market share data, though the high R² partly mitigates this concern.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Diversification (in this paper&amp;rsquo;s sense):&lt;/strong&gt; Buyers&amp;rsquo; practice of splitting bandwidth purchases across multiple cable operators to reduce exposure to idiosyncratic disruption risk. Diversification across n cables with i.i.d. disruption shocks reduces the variance of realized delivered bandwidth and lowers the required over-provisioning buffer, making the effective cost of a given usage level B a decreasing function of n.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Market Expansion Effect:&lt;/strong&gt; The channel through which entry of additional cable suppliers raises aggregate demand holding prices fixed. This occurs because each additional supplier reduces disruption risk, allowing buyers to demand more used bandwidth for the same price. It is distinct from the conventional competition channel (entry lowering prices).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Diversity Distortion:&lt;/strong&gt; The tendency toward insufficient entry arising because marginal entrants generate consumer surplus through diversification benefits but cannot fully capture this surplus absent price discrimination. Follows Spence (1976) and Mankiw and Whinston (1986).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Business-Stealing Distortion:&lt;/strong&gt; The tendency toward excessive entry arising because entrants reduce incumbents&amp;rsquo; output and revenues, creating a gap between private and social returns to entry.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Non-Stationary Markov Perfect Equilibrium:&lt;/strong&gt; The equilibrium concept used for the dynamic entry game, in which strategies and equilibrium selection rules are indexed by calendar time to accommodate substantial secular trends in demand and costs — as opposed to a stationary MPE which assumes a stable long-run distribution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Used Bandwidth vs. Purchased Bandwidth:&lt;/strong&gt; Used bandwidth B is the amount the buyer is committed to delivering (to downstream customers or for internal use). Purchased bandwidth Q is what the buyer actually contracts for across all cables; Q &amp;gt; B because the buyer holds an over-provisioning buffer against disruption risk. The ratio B/Q is a decreasing function of the disruption cost parameter gamma and an increasing function of the number of suppliers n.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nested Pseudo-Likelihood (NPL) Algorithm:&lt;/strong&gt; The baseline estimator for the dynamic game, following Aguirregabiria and Mira (2007). It iterates on the best-response mapping to impose equilibrium restrictions. The authors supplement NPL with two-step estimators (1-PML, 1-MD) and the spectral algorithm of Aguirregabiria and Marcoux (2021), which solves for the root of a nonlinear system using a quasi-Newton method and is robust to fixed-point instability.&lt;/p&gt;</description></item><item><title>Dynamic Regulation with Firm Linkages: Evidence from Texas</title><link>https://macropaperwarehouse.com/papers/dynamic-regulation-with-firm-linkages-evidence-from-texas/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/dynamic-regulation-with-firm-linkages-evidence-from-texas/</guid><description>&lt;p&gt;This paper evaluates the efficiency of linked environmental regulation, a targeting mechanism whereby inspectors who discover violations at one plant can increase enforcement pressure on other plants sharing the same owner. The central research question is whether linking inspection decisions across co-owned plants adds value over unlinked, plant-level targeting and over random enforcement. The paper develops a new empirical framework of dynamic moral hazard under linked regulation, applies it to Texas environmental enforcement data, and uses the estimated model to evaluate counterfactual regulatory designs.&lt;/p&gt;
&lt;p&gt;The empirical setting is the Texas Commission on Environmental Quality (TCEQ), which enforces the Resource Conservation and Recovery Act (RCRA, governing hazardous waste) and the Clean Water Act using a two-dimensional scoring system. A plant-level &amp;ldquo;site rating&amp;rdquo; score captures the individual plant&amp;rsquo;s compliance history, while a firm-wide &amp;ldquo;person rating&amp;rdquo; score aggregates the weighted average of plant scores across all plants under the same manager. Both scores feed into a multiplicative penalty escalation rule and a logit-form inspection probability function. The data are an unbalanced panel of 9,792 plants from 2012–2020, with detailed records of inspections, violations, penalties, scores, and ownership. The average plant is inspected with probability 0.289 per year and is linked with approximately 2 other plants through common ownership, though some firms own portfolios exceeding 50 plants.&lt;/p&gt;
&lt;p&gt;The model features firms endowed with private types (abatement cost parameters) that may be affiliated within a firm&amp;rsquo;s portfolio, choosing continuous pollution actions to maximize discounted payoffs net of expected penalties. The regulator observes only scores and minimizes social costs subject to a binding inspection budget. A key computational innovation is &amp;ldquo;continuation value sufficiency&amp;rdquo;: because fully solving the portfolio optimization over large plant sets is infeasible due to the curse of dimensionality, each plant&amp;rsquo;s decision is approximated using three state variables — its own plant score, the firm-wide score, and a scalar summarizing other co-owned plants&amp;rsquo; continuation values — governed by an AR(1) transition process. Estimation proceeds in three stages: OLS/logit for inspection and penalty parameters, simulated method of moments for type distribution and curvature parameters, and inversion of the regulator&amp;rsquo;s first-order conditions to recover sector-specific marginal social harms.&lt;/p&gt;
&lt;p&gt;Descriptive evidence confirms three preconditions for linked regulation to add value: violations are positively correlated within firm portfolios, inspections are targeted toward higher-scoring plants on both dimensions, and higher inspection probabilities (instrumented by scores) are associated with fewer violations conditional on plant fixed effects. The coefficient on predicted inspection probability in the deterrence regression (specification 3, plant fixed effects, inspected years only) is −3.920, and an increase in log scores from 0 to 1.5 (roughly the interquartile range) reduces expected violations by approximately 0.5.&lt;/p&gt;
&lt;p&gt;Structural estimates show that plant-level and firm-level type variance are similar (σ²_J = 0.209, σ²_F = 0.275), indicating moderate within-firm cost correlation. The curvature parameter y = 0.403 governs diminishing returns to negligence. In counterfactual experiments centered on a 30% budget increase (approximately 10 percentage point rise in per-plant inspection probability), unlinked plant-score-based escalations reduce social costs by 31.9% relative to random inspections. Linked firm-score-based escalations reduce social costs by 41.8% relative to random. The optimal mix — approximately 40% unlinked and 60% linked — reduces social costs by 42.2% relative to random. A back-of-the-envelope cost-benefit calculation calibrating utility-sector violation costs at $3,157 per violation and inspection costs at $740 finds a return of $11.77 in avoided social costs per additional dollar spent on inspections under the optimal mixed regime, versus $8.28 under random inspections.&lt;/p&gt;
&lt;p&gt;The scope conditions are specific: the framework applies to RCRA and Clean Water Act plants in Texas, which typically cannot reallocate production across facilities (unlike Clean Air Act firms), so the pollution-substitution channel documented for multi-plant Clean Air Act firms is not modeled. The penalty schedule is taken as fixed; only inspection allocation is treated as a policy choice.&lt;/p&gt;
&lt;p&gt;Q: What is linked regulation and why might it improve on unlinked enforcement?
A: Linked regulation allows the regulator to increase inspection and penalty pressure on all plants owned by a firm when any one plant accumulates violations. It is efficient when compliance costs (types) are correlated within firms — e.g., due to managerial practices — because a violation at one plant is informative about likely violations at co-owned plants. This correlation means the regulator can target scarce inspection resources toward portfolios that are likely to harbor multiple bad actors, rather than inspecting each plant independently.&lt;/p&gt;
&lt;p&gt;Q: How does Texas implement linked regulation in practice?
A: Texas uses a two-dimensional scoring system. The plant score (&amp;ldquo;site rating&amp;rdquo;) summarizes the individual plant&amp;rsquo;s violation history over the past five years, normalized by complexity points. The firm score (&amp;ldquo;person rating&amp;rdquo;) is the complexity-weighted average of plant scores across all plants under the same manager. Penalties are then multiplied by escalation factors based on both scores: a firm in the &amp;ldquo;unsatisfactory performer&amp;rdquo; tier (firm score ≥ 55) faces a 1.1× firm escalation, while a &amp;ldquo;high performer&amp;rdquo; (firm score &amp;lt; 0.1) faces a 0.9× multiplier. Because the firm escalation applies to all plants in the portfolio simultaneously, even a small change in firm score can produce large aggregate deterrence effects across a large portfolio.&lt;/p&gt;
&lt;p&gt;Q: What descriptive evidence supports the preconditions for linked regulation to add value?
A: Three pieces of evidence are presented. First, a scatterplot (Figure 1) shows a positive cross-sectional correlation between a plant&amp;rsquo;s average violations per inspection and the leave-one-out average violations per inspection of its co-owned plants, indicating within-firm cost correlation. Second, Table 2 logit regressions show that both plant score (coefficient 0.121) and firm score (coefficient 0.062) significantly predict inspection probability, conditional on year and NAICS fixed effects. Third, Table 3 shows that conditional on plant fixed effects, predicted inspection probability is negatively associated with violations (coefficient −3.246 in specification 2, rising to −3.920 in specification 3 restricted to inspected plant-years), confirming dynamic deterrence.&lt;/p&gt;
&lt;p&gt;Q: What is the curse of dimensionality problem and how is it resolved?
A: In a multi-plant firm, each plant&amp;rsquo;s optimal action depends on the scores of every other co-owned plant, producing a state space of dimension n_plants + 1. For firms with portfolios of 50+ plants this is computationally infeasible. The paper introduces &amp;ldquo;continuation value sufficiency&amp;rdquo;: each plant&amp;rsquo;s decision is reduced to three state variables — its own score s_j, the firm score s_f, and a scalar W_j aggregating other co-owned plants&amp;rsquo; continuation values. Transitions are approximated by plant-specific AR(1) processes. This reduces the portfolio problem from one high-dimensional value function to n_plant separate three-dimensional value functions, each solved independently within an inner fixed-point loop.&lt;/p&gt;
&lt;p&gt;Q: How are the type distribution parameters identified?
A: The mean type for each NAICS sector θ̄_g is identified by average violations per inspection within that sector — a higher mean type implies more violations conditional on inspection. The plant-level type variance σ²_J is identified by the share of total violation variance occurring across plants within the same firm. The firm-level type variance σ²_F is identified by the share of total violation variance occurring across firms. The curvature parameter y is identified by the responsiveness of violations to changes in predicted inspection probability (the coefficient from specification 3 of Table 3, which equals −3.920 empirically and −6.095 in simulation moments).&lt;/p&gt;
&lt;p&gt;Q: What are the main counterfactual results?
A: A 30% increase in the inspection budget (approximately +10 percentage points in per-plant inspection probability) is allocated under four regimes. Random inspections reduce violations per plant by 0.31 from a baseline of 0.98. Unlinked (plant-score) escalations reduce social costs by 31.9% more than random. Linked (firm-score) escalations reduce social costs by 41.8% more than random. The optimal mix (approximately 40% unlinked, 60% linked) reduces social costs by 42.2% more than random. In detected violations, all three targeted regimes perform similarly (+0.7% detected violations versus random), meaning the social cost advantage of linked regulation comes through greater undiscovered deterrence rather than through detection rates.&lt;/p&gt;
&lt;p&gt;Q: How does the decomposition into static, own-plant, and cross-plant effects clarify the mechanism?
A: For unlinked escalations: the static effect accounts for −5.4% of social cost relative to random, own-plant dynamic deterrence accounts for −30.6%, and the cross-plant effect is +4.1% (slightly adverse, because unlinked escalations do not account for portfolio-level incentives). For linked escalations: the static effect is −2.4%, own-plant deterrence is −24.5% (smaller than unlinked because linked escalations are less precisely targeted to individual plant histories), and cross-plant deterrence is −14.9% (large and beneficial). The dominance of cross-plant deterrence under linked escalations is the key mechanism explaining why linking outperforms unlinked targeting.&lt;/p&gt;
&lt;p&gt;Q: What does the cost-benefit calculation find?
A: Calibrating utility-sector violation social costs at $3,157 per violation (from Kang and Silveira 2021 for California water utilities post-2006) and inspection costs at $740, the paper finds a return of $11.77 in avoided social costs per additional dollar spent on inspections under the optimal linked/unlinked mix, versus $8.28 under random inspections. This suggests a large return to expanding enforcement budgets, with the gain amplified substantially by optimal targeting design.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions and limitations acknowledged?
A: The framework applies to RCRA and Clean Water Act plants in Texas, where firms (e.g., gas station chains) typically cannot reallocate production across facilities, so the pollution-substitution channel documented by Gibson (2019) for Clean Air Act firms is not modeled. The penalty schedule is taken as fixed — only inspection allocation is treated as a policy choice — because Texas&amp;rsquo;s bylaws are prescriptive about how violations translate into penalties while leaving inspection targeting largely to regulator discretion. Social harm parameters h_g are identified only up to a scale normalization. The paper also does not model why types are correlated within firms (bad managers versus specialization), as the counterfactual results depend only on the degree of correlation, not its source.&lt;/p&gt;
&lt;p&gt;Q: How well does the model fit the data?
A: The model matches the targeted moments well (Table 5). Mean violations by NAICS sector are closely reproduced (e.g., utility: 0.201 empirical vs. 0.184 simulated; trade: 0.252 vs. 0.236). Responsiveness of violations to inspection probability matches closely (−6.398 empirical vs. −6.095 simulated). A non-targeted fit statistic — the correlation between a plant&amp;rsquo;s own violation rate and its co-owned plants&amp;rsquo; violation rates — is 0.32 in simulation versus 0.26 in the data, which the authors characterize as a good out-of-sample fit given it was not directly targeted in estimation.&lt;/p&gt;
&lt;p&gt;Q: How do heterogeneous effects shed light on the distributional consequences of regulation?
A: The own-plant deterrence effect is positive for all plants including those with low types that are unlikely to be targeted, but is especially pronounced for high-type plants under unlinked escalations. Under linked escalations, high-type plants are deterred less to the extent they are co-owned with lower-type plants, because firm-score-based targeting aggregates across the portfolio. Cross-plant effects are predictably small under unlinked escalations and larger under linked escalations, especially for firms with high-type portfolios, since those are the firms whose firm scores respond most to individual violations.&lt;/p&gt;
&lt;p&gt;Linked regulation: An enforcement mechanism in which the discovery of violations at one plant triggers increased inspection and penalty pressure on all other plants under the same owner. It exploits within-firm correlation in compliance costs to target scarce regulatory resources more efficiently than plant-by-plant escalation alone.&lt;/p&gt;
&lt;p&gt;Escalation mechanism: A penalty and inspection design in which plants with worse compliance records — measured by accumulated compliance scores — face disproportionately greater scrutiny and higher penalties per additional violation. The TCEQ&amp;rsquo;s two-dimensional scoring system is an escalation mechanism operating simultaneously at the individual plant and firm portfolio level.&lt;/p&gt;
&lt;p&gt;Plant score / firm score: The plant score (&amp;ldquo;site rating&amp;rdquo;) is a normalized index of a single facility&amp;rsquo;s violation history over the past five years, divided by investigation count and complexity points; the firm score (&amp;ldquo;person rating&amp;rdquo;) is the complexity-weighted average of all plant scores across the firm&amp;rsquo;s portfolio. Higher scores indicate worse compliance records and trigger both higher penalties and higher inspection probabilities.&lt;/p&gt;
&lt;p&gt;Continuation value sufficiency: The paper&amp;rsquo;s solution to the curse of dimensionality in large plant portfolios. Rather than tracking the full joint score state across all co-owned plants, each plant&amp;rsquo;s optimal action is approximated using three variables — its own score, the aggregate firm score, and a scalar W_j summarizing co-owned plants&amp;rsquo; continuation values — with state transitions governed by a plant-specific AR(1) process.&lt;/p&gt;
&lt;p&gt;Dynamic moral hazard under linked regulation: The firm&amp;rsquo;s problem of choosing how much to invest in pollution mitigation at each plant over time, given that current actions affect future scores, future penalties, and — through the firm-wide score — future scrutiny of all co-owned plants. The moral hazard arises because abatement costs are private information not directly observable by the regulator.&lt;/p&gt;
&lt;p&gt;Complexity points: A normalization factor in the TCEQ scoring system that adjusts raw violation counts for plant size and sector, enabling comparable compliance histories across heterogeneous facilities. They were introduced in 2012 specifically to prevent mechanically larger facilities from appearing riskier simply due to their scale.&lt;/p&gt;
&lt;p&gt;Cross-plant deterrence effect: The reduction in pollution actions at co-owned plants induced by increases in the firm-wide score following a violation at one plant in the portfolio. In the counterfactual decomposition, this effect accounts for −14.9 percentage points of social cost reduction under linked escalations and is the primary mechanism by which linked regulation outperforms unlinked plant-level escalation.&lt;/p&gt;</description></item><item><title>Energy Transitions in Regulated Markets</title><link>https://macropaperwarehouse.com/papers/energy-transitions-in-regulated-markets/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/energy-transitions-in-regulated-markets/</guid><description>&lt;p&gt;This paper asks how rate-of-return (RoR) regulation in U.S. electricity markets affects the speed and efficiency of energy transitions, specifically the transition from coal to combined-cycle natural gas (CCNG) generation driven by fracking-induced cost declines. The authors build and estimate a structural model of regulated utility behavior in which utilities optimize investment, retirement, and hourly operations decisions against an incentive structure set by state Public Utility Commissions (PUCs).&lt;/p&gt;
&lt;p&gt;The regulatory environment combines two instruments: (1) an allowable rate of return that is decreasing in consumer electricity rates (incentive regulation), parameterized as s = (r/r₀)^{-γ}, where higher γ penalizes high-cost outcomes more severely; and (2) a &amp;ldquo;used-and-useful&amp;rdquo; standard in which a coal plant&amp;rsquo;s contribution to the rate base depends on its capacity utilization via a logit function. These two instruments create a tension: utilities want to lower costs to earn a higher RoR, but also want to run existing coal plants—even when uneconomical—to prove they are &amp;ldquo;used and useful&amp;rdquo; and thus maximize their rate base and profits.&lt;/p&gt;
&lt;p&gt;The authors estimate the model using publicly available EIA and EPA CEMS data spanning 2006–2017, covering 39 unique regulated utilities in the Eastern Interconnection across more than 4 million utility-hour observations (459 utility-years). Structural parameters are recovered via a nested fixed-point indirect inference approach that matches simulated regression coefficients to actual data; investment and retirement costs are estimated with a GMM nested fixed-point approach.&lt;/p&gt;
&lt;p&gt;Key reduced-form findings confirm the model&amp;rsquo;s two core mechanisms. First, a 10% increase in total variable costs is associated with a 2.5% decrease in variable profits per MW of capacity (with utility fixed effects), consistent with incentive regulation. Second, regulated utilities reduce coal generation by only a statistically insignificant 4.2 percentage points when coal fuel costs exceed import prices, compared to 16.1 percentage points for restructured utilities—consistent with regulated utilities running coal out-of-dispatch order to preserve used-and-useful status.&lt;/p&gt;
&lt;p&gt;In counterfactual simulations that impose 2018–20 natural gas prices ($2.01/MMBtu versus the 2006 price of $7.24/MMBtu) on utilities with their 2006 capital stocks, regulated utilities retire only 53% of coal capacity over 30 years and increase CCNG capacity by 296%, whereas a cost minimizer would retire most coal capacity while increasing CCNG by only 58%. The Averch-Johnson over-investment effect dominates: regulated utilities over-invest in CCNG while simultaneously over-using legacy coal.&lt;/p&gt;
&lt;p&gt;Carbon taxes on regulated utilities reduce short-run coal generation only 48% as much as when imposed on a cost minimizer (because the used-and-useful incentive partially offsets the carbon price signal), but in the long run result in 68% lower coal capacity and 77% lower coal generation relative to baseline by year 30—larger effects than for the cost minimizer. Eliminating the coal usage incentive (μ₂ = 0) produces 82% lower coal capacity and 92% lower coal generation over 30 years but requires utility variable profits to fall by over $300 million, threatening reliability without compensating transfers.&lt;/p&gt;
&lt;p&gt;Scope conditions: Results apply to regulated (non-restructured) utilities in the Eastern Interconnection, 2006–2017. The model estimates the coal-to-CCNG transition only; it explicitly does not model the ongoing transition to renewables and storage due to insufficient data variation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-central-research-question"&gt;Q1. What is the central research question?&lt;/h3&gt;
&lt;p&gt;The paper asks whether and how rate-of-return regulation in U.S. electricity markets slows energy transitions, and what alternative regulatory structures or carbon tax policies could accelerate the transition away from coal. It addresses this both theoretically—through a structural model of regulated utility behavior—and empirically, through estimation and counterfactual simulation using data on 39 regulated utilities over 2006–2017.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-key-regulatory-instruments-in-the-model-and-what-distortions-do-they-create"&gt;Q2. What are the two key regulatory instruments in the model, and what distortions do they create?&lt;/h3&gt;
&lt;p&gt;The first instrument is incentive regulation: the allowable rate of return declines as consumer electricity rates rise (s = (r/r₀)^{-γ}), so utilities have an incentive to lower costs. The second is the used-and-useful standard: a coal plant&amp;rsquo;s contribution to the rate base depends on its capacity utilization via a logit function, creating an incentive to run coal plants even when their fuel costs exceed import prices. Together, these instruments generate a tension between cost-reduction incentives and legacy-capacity-preservation incentives, causing the regulated utility to both over-invest in new CCNG capacity (Averch-Johnson effect) and over-use existing coal capacity relative to the cost-minimizing benchmark.&lt;/p&gt;
&lt;h3 id="q3-what-does-the-reduced-form-evidence-show-about-uneconomical-coal-usage"&gt;Q3. What does the reduced-form evidence show about uneconomical coal usage?&lt;/h3&gt;
&lt;p&gt;In a triple-difference specification, regulated utilities reduce coal generation by only 4.2 percentage points (statistically insignificant) when coal fuel costs exceed import prices, compared to a 16.1 percentage point reduction for restructured utilities. CCNG generation responds similarly under both regulatory regimes (21.1 vs. 19.7 percentage points), confirming that the distortion is specific to legacy coal under RoR regulation and not a general feature of high-cost generation. The six states with the largest responsiveness of coal usage to low market prices are all restructured states; out-of-dispatch-order coal generation also correlates strongly with utility ownership share across states.&lt;/p&gt;
&lt;h3 id="q4-what-do-the-structural-parameter-estimates-reveal-about-the-rate-base"&gt;Q4. What do the structural parameter estimates reveal about the rate base?&lt;/h3&gt;
&lt;p&gt;Each MW of CCNG capacity increases the rate base by $229,000. When fully utilized, each MW of coal capacity contributes 1.144 times as much as CCNG. When coal is not fully used, unused coal capacity contributes only 40% as much to the rate base as CCNG. NGT capacity contributes 79% more to the rate base than CCNG per MW. Operations cost estimates include O&amp;amp;M costs of $12.89/MWh for coal, $8.82/MWh for CCNG, and $44.63/MWh for NGT; a 100 MW coal ramp in one hour costs $4,770 versus $3,860 for CCNG.&lt;/p&gt;
&lt;h3 id="q5-what-happens-in-the-30-year-long-run-counterfactual-under-the-baseline-regulated-utility"&gt;Q5. What happens in the 30-year long-run counterfactual under the baseline regulated utility?&lt;/h3&gt;
&lt;p&gt;Facing a sudden drop to 2018–20 natural gas prices ($2.01/MMBtu vs. $7.24/MMBtu in 2006), regulated utilities retire 53% of coal capacity and increase CCNG capacity by 296% over 30 years. The Averch-Johnson over-investment effect dominates: utilities invest heavily in CCNG while retaining and using legacy coal far longer than a cost minimizer would. The social planner effectively eliminates coal generation immediately (99% reduction in the first period) and retires almost all coal capacity over the horizon.&lt;/p&gt;
&lt;h3 id="q6-how-does-a-cost-minimizer-behave-relative-to-the-regulated-utility-in-the-same-long-run-counterfactual"&gt;Q6. How does a cost minimizer behave relative to the regulated utility in the same long-run counterfactual?&lt;/h3&gt;
&lt;p&gt;A cost minimizer immediately reduces coal generation by 50% in the first period and retires most coal capacity over 30 years while increasing CCNG capacity by only 58%—versus the regulated utility&amp;rsquo;s 296% CCNG increase. Thirty years after the shock, the cost minimizer has retired 71% more coal capacity than the regulated utility. The cost minimizer&amp;rsquo;s much smaller CCNG expansion reflects that it does not face Averch-Johnson incentives to over-invest in rate-base capital.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-short-run-vs-long-run-impact-of-carbon-taxes-on-regulated-utilities-compared-to-cost-minimizers"&gt;Q7. What is the short-run vs. long-run impact of carbon taxes on regulated utilities compared to cost minimizers?&lt;/h3&gt;
&lt;p&gt;In the short run, carbon taxes on regulated utilities reduce coal generation only 48% as much as when imposed on a cost minimizer (34% vs. ~100% in immediate generation drop), because the used-and-useful incentive counteracts the carbon price signal. In the long run (30-year horizon), however, carbon taxes on regulated utilities result in 68% lower coal capacity and 77% lower coal generation relative to baseline—larger percentage reductions than for a cost minimizer—because the regulatory structure amplifies the retirement incentive over time once carbon costs erode the economic rationale for keeping coal in the rate base.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-short-run-operations-counterfactual-finding-for-carbon-taxes-in-the-sample-period"&gt;Q8. What is the short-run operations counterfactual finding for carbon taxes in the sample period?&lt;/h3&gt;
&lt;p&gt;Using each utility-year in the analysis sample, imposing carbon taxes on regulated utilities reduces carbon costs by only about $500 million relative to baseline—41% of the $1.3 billion carbon cost savings from imposing the same carbon taxes on a cost minimizer. Despite this limited carbon reduction, electricity rates nearly triple from $77.58/MWh to $224.18/MWh under the regulated utility with carbon taxes, as the utility passes through most carbon costs to consumers; regulated utility variable profits also fall by over $500 million.&lt;/p&gt;
&lt;h3 id="q9-what-happens-when-the-coal-usage-incentive-is-eliminated-μ--0"&gt;Q9. What happens when the coal usage incentive is eliminated (μ₂ = 0)?&lt;/h3&gt;
&lt;p&gt;Setting the coal usage incentive parameter μ₂ = 0 (eliminating the logit slope on capacity utilization) causes coal capacity to fall 82% and coal generation to fall 92% relative to baseline over 30 years—a slightly larger generation decline than for the cost minimizer. However, this comes at the cost of more than twice the CCNG capacity due to the Averch-Johnson effect, and requires utility variable profits to fall by over $300 million, raising reliability concerns unless accompanied by compensating transfers.&lt;/p&gt;
&lt;h3 id="q10-how-does-the-papers-mechanism-relate-to-observed-differences-in-coal-exit-rates-between-regulated-and-restructured-states"&gt;Q10. How does the paper&amp;rsquo;s mechanism relate to observed differences in coal exit rates between regulated and restructured states?&lt;/h3&gt;
&lt;p&gt;Between 2006 and 2018, 26.0% of coal capacity exited in restructured states versus only 17.2% in regulated states—a gap the authors attribute primarily to the used-and-useful incentive structure in RoR regulation. The structural model quantifies how this regulatory feature specifically distorts coal usage and retirement decisions; it is not explained by demand or cost differences across states, as confirmed by the triple-difference evidence showing the gap is specific to coal (not CCNG) and to regulated (not restructured) utilities.&lt;/p&gt;
&lt;h3 id="q11-why-does-the-paper-argue-that-alternative-regulatory-adjustments-are-insufficient-to-replicate-cost-minimizing-transitions"&gt;Q11. Why does the paper argue that alternative regulatory adjustments are insufficient to replicate cost-minimizing transitions?&lt;/h3&gt;
&lt;p&gt;Changing regulatory parameters—such as increasing the coal usage incentive or adjusting the electricity rate penalty—does not come close to replicating the speed of the energy transition under a cost minimizer in the long-run simulations. Regulatory adjustments that do approach cost-minimizing outcomes (such as eliminating μ₂) require large reductions in utility variable profits sufficient to risk reliability, consistent with why the 2022 Inflation Reduction Act relied on substantial investment transfers rather than carbon taxes as its primary clean energy instrument.&lt;/p&gt;
&lt;h3 id="q12-what-is-the-papers-identification-strategy"&gt;Q12. What is the paper&amp;rsquo;s identification strategy?&lt;/h3&gt;
&lt;p&gt;Identification exploits the sharp, exogenous decline in natural gas fuel prices from fracking, which had heterogeneous implications across utilities depending on their initial capital mixes (coal-heavy vs. CCNG-heavy). By comparing investment, retirement, and operations decisions across utilities and over time—particularly between utilities that had CCNG exposure before the price decline and those that did not—the authors recover the structural regulatory and cost parameters. The IV specification for reduced-form evidence uses the current natural gas price interacted with the utility&amp;rsquo;s initial CCNG generation share as an instrument for fuel and import costs.&lt;/p&gt;
&lt;h3 id="q13-what-are-the-papers-explicit-limitations"&gt;Q13. What are the paper&amp;rsquo;s explicit limitations?&lt;/h3&gt;
&lt;p&gt;The paper estimates the coal-to-CCNG transition only and cannot speak to the transition to renewables and storage, because there is insufficient variation in the data to identify how regulators would treat CCNG as a legacy technology subject to used-and-useful standards, or how renewables and storage would contribute to the rate base. The authors note that over-investment in CCNG capacity may create future stranded asset problems for ratepayers and that usage incentives for CCNG are likely to further hinder the transition to renewables—but these are conjectures rather than estimated findings.&lt;/p&gt;
&lt;p&gt;Rate-of-return (RoR) regulation: A regulatory structure in which the PUC sets electricity rates so that utility revenues cover total variable costs plus an allowable return on the utility&amp;rsquo;s rate base (capital stock), with the allowable return parameterized as s = (r/r₀)^{-γ}, declining as consumer electricity rates rise.&lt;/p&gt;
&lt;p&gt;Used-and-useful standard: A prudence criterion under which a capital asset&amp;rsquo;s contribution to the rate base depends on its capacity utilization, modeled as a logit function of the generation-to-capacity ratio; fully used coal capacity contributes 1.144 times as much as CCNG per MW, while unused coal contributes only 40% as much.&lt;/p&gt;
&lt;p&gt;Rate base: The capital stock on which the PUC grants the utility its allowable rate of return; adjusted by prudence and used-and-useful assessments and described in the paper as &amp;ldquo;at best an arduous task&amp;rdquo; to quantify precisely.&lt;/p&gt;
&lt;p&gt;Averch-Johnson (AJ) over-investment effect: The tendency of regulated utilities to over-invest in capital because profits are proportional to the rate base; in this paper&amp;rsquo;s setting, this causes regulated utilities to increase CCNG capacity by 296% over 30 years following the natural gas price shock, compared to 58% for a cost minimizer.&lt;/p&gt;
&lt;p&gt;Incentive regulation: A modification of cost-plus RoR regulation in which the allowable rate of return declines as electricity rates rise; it provides efficiency incentives for cost reduction but does not achieve first-best outcomes and is insufficient to overcome the used-and-useful distortion for legacy coal.&lt;/p&gt;
&lt;p&gt;Out-of-dispatch-order generation: Running a generation unit when its fuel costs exceed the market import price; regulated utilities engage in this behavior with coal plants to maintain used-and-useful status and rate base contribution, whereas restructured utilities do not face this incentive.&lt;/p&gt;
&lt;p&gt;Nested fixed-point indirect inference: The estimation approach used to recover structural regulatory and operations parameters by minimizing the distance between regression coefficients from actual data and those from model-simulated data via a non-linear parameter search.&lt;/p&gt;</description></item><item><title>Firm dynamics and random search over the business cycle</title><link>https://macropaperwarehouse.com/papers/firm-dynamics-and-random-search-over-the-business-cycle/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/firm-dynamics-and-random-search-over-the-business-cycle/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;How do aggregate economic fluctuations reallocate workers across the firm productivity distribution over the business cycle? In particular, to what extent do recessions impede workers&amp;rsquo; movement up the job ladder toward more productive firms?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper develops a tractable random search model combining three features that had not previously been integrated in a single quantitative framework: (i) firm dynamics driven by idiosyncratic productivity shocks, with endogenous entry and exit; (ii) on-the-job search, generating a job ladder in which workers gradually move toward more productive firms; and (iii) aggregate productivity shocks. Multi-worker firms post employment contracts, choose hiring rates, and decide whether to continue or exit. The key tractability result — called &amp;ldquo;size-independence&amp;rdquo; (Result 1) — shows that, under a constant-returns hiring cost technology, firms&amp;rsquo; optimal policies (contract value, hiring rate, exit decision) are all independent of firm size, so the relevant state space reduces from the full joint distribution of firm productivity and size to the employment-weighted distribution of firm productivity alone. A further result (&amp;ldquo;rank-monotonic equilibrium,&amp;rdquo; Result 2) guarantees, under a sufficient convexity condition on hiring costs (hc&amp;rsquo;&amp;rsquo;(h)/c&amp;rsquo;(h) ≥ 1), that the optimal employment contract is increasing in firm productivity, so the job ladder maps one-for-one onto the firm productivity ladder. The optimal wage contract then admits a closed-form solution.&lt;/p&gt;
&lt;p&gt;The model is calibrated to British data for 1997–2018. Worker-level transition rates (unemployment-to-employment, employment-to-unemployment, and job-to-job) are drawn from the British Household Panel Survey (BHPS). Firm-level data on labor productivity (value added per worker) and employment costs per worker come from the Annual Respondents Database (ARD) and Annual Business Survey (ABS), merged with the Business Structure Database (BSD). The numerical solution adapts ideas from Krusell and Smith (1998), approximating the employment-weighted productivity distribution by a small set of moments and parameterizing value functions as polynomials in the aggregate state; standard linearization methods are inapplicable because endogenous firm entry and exit introduces a discontinuity in value functions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Model validation via the OP decomposition.&lt;/em&gt; The paper&amp;rsquo;s central validation exercise uses the Olley-Pakes (OP) decomposition of a labor productivity index constructed from firm-level data. The aggregate employment-weighted labor productivity index is decomposed into (a) the unweighted average firm productivity and (b) an interaction term (the &amp;ldquo;OP term&amp;rdquo;), which captures the covariance between employment shares and productivity — i.e., how well workers are allocated to productive firms. In the British firm-level data, approximately 20 percent of the variance of the aggregate labor productivity index is accounted for by this interaction (OP) term, with the remaining ~80 percent attributable to the unweighted average of firm productivity. The baseline model, with this moment untargeted, successfully replicates this 80/20 split. By contrast, the leading benchmark model of Moscarini and Postel-Vinay (2016) (MPV2016), calibrated to the same British data, attributes nearly all of the variance of labor productivity to the OP/worker reallocation term, grossly overstating the importance of job-ladder dynamics.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Structural decomposition of labor productivity.&lt;/em&gt; Using the calibrated baseline model to decompose the variance of aggregate labor productivity over the post-war British business cycle (&amp;ldquo;GDP shocks&amp;rdquo; going back to 1955), the baseline model attributes approximately 30 percent to the direct effect of the aggregate productivity shock, approximately 50 percent to changes in the distribution of active firms (the &amp;ldquo;firm ladder&amp;rdquo; or firm selection component), and approximately 20 percent to the worker reallocation component (the OP interaction term). This result is robust to an alternative calibration with a lower curvature of the hiring cost function (c1 = 1).&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Persistence and mechanisms.&lt;/em&gt; The impact of recessions on the job ladder is persistent: while the aggregate productivity shock is typically close to its pre-recession value four years after a typical recession onset, the overall allocation of workers to firms remains clearly worse relative to the pre-recession level at that same horizon. The Great Recession, viewed through the lens of the model, is a large but not unusually large recession.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Firm selection with multiple aggregate shocks.&lt;/em&gt; An unexpected finding concerns the direction of firm selection. With a single aggregate productivity shock, the model generates a standard &amp;ldquo;cleansing&amp;rdquo; mechanism: negative shocks raise the firm exit threshold, so surviving firms are on average more productive. However, when additional shocks to the exogenous separation rate (δ) and hiring cost scale (c0) are included — as required to match the volatility of labor market flows — firm selection instead amplifies the decline in labor productivity. The mechanism is a general equilibrium one: a higher separation rate lowers the optimal wage contract (since greater separation risk is passed on to workers), which in turn lowers the entry-exit threshold. Less productive firms become viable because their employees face higher unemployment risk and therefore accept lower wages; moreover, a larger pool of unemployed workers makes it easier for low-productivity firms to recruit.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Wage flexibility tension.&lt;/em&gt; The model implies a pass-through elasticity of wages to productivity shocks of approximately 0.7, well above the 0.05–0.2 range typically found empirically.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;All calibration and quantitative results pertain to Britain for the period 1997–2018 (firm-level data) and 1955–2018 (GDP-based aggregate shocks). The model abstracts from decreasing returns to scale in production and from nominal rigidities. The tractability results rely on specific assumptions about the hiring cost function; the rank-monotonicity condition requires sufficient convexity (hc&amp;rsquo;&amp;rsquo;(h)/c&amp;rsquo;(h) ≥ 1).&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-central-tractability-result-and-why-does-it-matter-for-computational-feasibility"&gt;Q1. What is the central tractability result and why does it matter for computational feasibility?&lt;/h3&gt;
&lt;p&gt;A: Result 1 (&amp;ldquo;size-independence&amp;rdquo;) shows that, because both the production technology and the hiring cost function are constant returns to scale, the firm&amp;rsquo;s present discounted value of profits is linear in employment. As a result, per-worker profits are independent of firm size, and optimal firm policies — the hiring rate, the contract value offered to workers, and the continuation/exit decision — all depend only on the firm&amp;rsquo;s current productivity, not on its size. This collapses the state space from the full joint distribution of firm productivity and employment size to the employment-weighted measure of firm productivity Lt(p), a uni-dimensional object. Without this result, the model would require tracking the entire joint firm distribution, making it computationally intractable.&lt;/p&gt;
&lt;h3 id="q2-what-is-a-rank-monotonic-equilibrium-rme-and-what-conditions-guarantee-it"&gt;Q2. What is a rank-monotonic equilibrium (RME) and what conditions guarantee it?&lt;/h3&gt;
&lt;p&gt;A: An RME is a recursive equilibrium in which the optimal contract offered by a firm is weakly increasing in that firm&amp;rsquo;s current productivity realization, for all aggregate states. Result 2 provides sufficient conditions: (i) the Markov process for firm-specific productivity satisfies first-order stochastic dominance (more productive firms today are more likely to be more productive tomorrow), (ii) the distribution of offered contracts is everywhere differentiable (ruling out mass points), and (iii) the hiring cost function satisfies hc&amp;rsquo;&amp;rsquo;(h)/c&amp;rsquo;(h) ≥ 1 — a sufficient convexity condition. The economic interpretation of the convexity condition is that firms must find retention (offering higher wages) sufficiently costly relative to new hiring that more productive firms optimally choose to use the wage margin to limit quits. The baseline calibration yields c1 ≈ 5.9 (so costs are highly convex in the hiring rate), though results are also reported for the minimum permissible c1 = 1.&lt;/p&gt;
&lt;h3 id="q3-what-does-the-optimal-employment-contract-look-like-in-a-rank-monotonic-equilibrium-and-what-does-it-reveal-about-rent-extraction"&gt;Q3. What does the optimal employment contract look like in a rank-monotonic equilibrium, and what does it reveal about rent extraction?&lt;/h3&gt;
&lt;p&gt;A: In an RME, the optimal contract V(p,ω,L) is a weighted average of the value of unemployment U(ω,L) and the firm-workers&amp;rsquo; joint surplus S(p,ω,L), where the weights are determined endogenously by the employment-weighted measure of firm productivity L. Specifically, the contract integrates the surplus of all firms with productivity below p, weighted by the share of employed workers at those firms, and divided by the mass of job seekers willing to accept the contract. As the employed workers&amp;rsquo; relative search intensity s approaches zero, the contract converges to the value of unemployment — workers receive no rents. The endogenous bargaining weight evolves with the aggregate state over the business cycle, unlike standard Nash bargaining models with a fixed exogenous weight.&lt;/p&gt;
&lt;h3 id="q4-what-firm-level-moments-are-used-to-calibrate-the-steady-state-model-and-what-is-the-logic-behind-the-parameter-moment-mapping"&gt;Q4. What firm-level moments are used to calibrate the steady-state model, and what is the logic behind the parameter-moment mapping?&lt;/h3&gt;
&lt;p&gt;A: Eight moments are targeted. From the BHPS worker data: the average UE rate (0.058) pins down the scale of hiring costs c0; the average EU rate (0.003) pins down the exogenous separation rate δ; and the average EE (job-to-job) rate (0.016) pins down the relative search intensity s. From the firm-level ARD/BSD data: average firm size (12.1 employees) pins down the entry probability µ; the share of job destruction from firm exits (0.526) disciplines the flow value of unemployment b; the autocorrelation of firm employment ln(n) (0.949 annually) disciplines the persistence of idiosyncratic productivity ρp; the interquartile range of firm-level labor productivity (1.129 log points) disciplines the volatility of idiosyncratic shocks σp; and the regression coefficient of firm employment growth on lagged labor productivity (0.136) disciplines the curvature of hiring costs c1. The baseline calibration fits all eight moments closely.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-calibrated-model-match-non-targeted-moments-and-what-does-this-establish"&gt;Q5. How does the calibrated model match non-targeted moments, and what does this establish?&lt;/h3&gt;
&lt;p&gt;A: The model generates several realistic features not targeted in calibration. It produces a realistic Pareto tail for the employment-size distribution (Pareto tail exponent of 1.033 in the model vs. 1.066 in the data), which arises from the combination of size-independent growth rates and firm entry and exit — conditions identified in the literature as generating power law distributions. The model also matches the dispersion of employment costs per worker across firms (capturing about 70 percent of the interquartile range of ECi,t), the slope of a regression of employment costs on labor productivity (model: 0.685 vs. data: 0.704), and the slope of a regression of employment growth on employment costs (model: 0.162 vs. data: 0.131). These non-targeted matches provide independent validation of the model&amp;rsquo;s wage-determination mechanism.&lt;/p&gt;
&lt;h3 id="q6-why-is-a-single-aggregate-productivity-shock-insufficient-to-match-labor-market-fluctuations-and-what-additional-shocks-are-needed"&gt;Q6. Why is a single aggregate productivity shock insufficient to match labor market fluctuations, and what additional shocks are needed?&lt;/h3&gt;
&lt;p&gt;A: With a single aggregate productivity shock calibrated to match the autocorrelation and standard deviation of log GDP, the model generates labor market fluctuations that are roughly an order of magnitude smaller than in the data. For example, the standard deviation of the EU transition rate is 4.1×10⁻⁴ in the single-shock model versus 2.3×10⁻³ in the data. Adding a discount rate shock (ω,r) partially helps but still leaves the job-finding rate (UE) more than 50 percent too smooth. Adding a separation rate shock (ω,δ) substantially increases EU and UE volatility but generates insufficient EE (job-to-job) volatility. The combination (ω,δ,c0) — adding a shock to the scale of hiring costs c0 — brings the standard deviations of EU and UE close to the data (2.0×10⁻³ and 4.0×10⁻⁴ vs. data 2.3×10⁻³ and 2.7×10⁻⁴), though the model still generates slightly under half the observed volatility in EE rates. This combination is the baseline for the quantitative analysis.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-op-decomposition-how-is-it-computed-from-the-firm-level-data-and-what-does-it-measure-in-the-model"&gt;Q7. What is the OP decomposition, how is it computed from the firm-level data, and what does it measure in the model?&lt;/h3&gt;
&lt;p&gt;A: The aggregate labor productivity index LPt is constructed from firm-level data as the employment-share-weighted average of log value added per worker across firms. The OP decomposition writes this as LPt = LPt_bar + OPt, where LPt_bar is the unweighted (simple) average of firm-level productivity and OPt is the covariance between employment shares and labor productivity (the &amp;ldquo;interaction term&amp;rdquo;). In the data, OPt increases when workers are disproportionately employed at above-average-productivity firms. In the model, LPt_bar maps onto the average (log) productivity of active firms — the support of the job ladder — while OPt maps onto the difference between the employment-weighted and the unweighted averages of firm productivity, directly measuring how high up the ladder workers are located relative to the set of active firms. Around 20 percent of the variance of LPt in the British data is accounted for by OPt, and the model replicates this.&lt;/p&gt;
&lt;h3 id="q8-how-does-the-great-recession-appear-in-the-op-decomposition-and-does-the-model-fit-the-decomposition-during-this-episode"&gt;Q8. How does the Great Recession appear in the OP decomposition, and does the model fit the decomposition during this episode?&lt;/h3&gt;
&lt;p&gt;A: During the Great Recession (2008q2–2009q3 in the UK), around 20 percent of the overall fall in the labor productivity index is accounted for by the fall in the OP interaction term, with the remaining 80 percent coming from the fall in the unweighted average firm productivity. The model, even though it does not target this decomposition in calibration, successfully matches both the average firm productivity component and the interaction (OP) component during the Great Recession. This matching holds both in the baseline calibration (c1 ≈ 5.9) and in the alternative calibration with c1 = 1. The model also matches the analogous decomposition for employment costs per worker (ECt), an additional non-targeted validation.&lt;/p&gt;
&lt;h3 id="q9-why-does-firm-selection-amplify-rather-than-cleanse-in-the-baseline-multi-shock-calibration"&gt;Q9. Why does firm selection amplify rather than cleanse in the baseline multi-shock calibration?&lt;/h3&gt;
&lt;p&gt;A: In the single-shock (productivity ω only) model, a negative productivity shock lowers surplus at all firms, raising the exit threshold pE and thus selecting out low-productivity firms — the standard &amp;ldquo;cleansing&amp;rdquo; mechanism. In the multi-shock baseline, the additional separation rate shock (δ) generates a less intuitive mechanism. A higher δ lowers the optimal wage contract (since increased separation risk is passed on to workers: ∂V/∂δ ≤ 0), which reduces the value of continued employment. This lowers the joint firm-worker surplus threshold for exit, making it viable for low-productivity firms to remain active. Moreover, the larger pool of unemployed workers (generated by the δ shock) depresses the outside option of workers and makes it easier for low-productivity firms to recruit. As a result, the entry-exit threshold pE,t falls — the set of active firms becomes less productive on average — producing a negative firm selection contribution to labor productivity and a positive (amplifying rather than cleansing) contribution to the variance of LPt.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-structural-variance-decomposition-of-labor-productivity-in-the-baseline-model"&gt;Q10. What is the structural variance decomposition of labor productivity in the baseline model?&lt;/h3&gt;
&lt;p&gt;A: Simulating the baseline model over the post-war British business cycle (1955–2020, GDP shocks), the variance of aggregate labor productivity LPt decomposes into three structural terms: approximately 30 percent (0.296) from the direct effect of the aggregate productivity shock ln(ωt); approximately 50 percent (0.541) from changes in the average productivity of active firms E[KP bar_t(ln p)] — the &amp;ldquo;firm ladder&amp;rdquo; or firm selection component; and approximately 20 percent (0.163) from the worker reallocation component OPt = E[LP bar_t(ln p)] − E[KP bar_t(ln p)]. This decomposition implies that roughly 70 percent of fluctuations in labor productivity are driven by worker reallocation broadly defined (the firm ladder plus the interaction term), with the firm selection component being the largest single driver. The result is robust to the alternative c1 = 1 calibration (30/49/22 percent split).&lt;/p&gt;
&lt;h3 id="q11-how-does-the-baseline-model-compare-to-mpv2016-in-the-variance-decomposition"&gt;Q11. How does the baseline model compare to MPV2016 in the variance decomposition?&lt;/h3&gt;
&lt;p&gt;A: In the multi-shock calibration (ω,δ,c0), the MPV2016 model calibrated to the same British data attributes approximately 97.7 percent (0.977) of the variance of LPt to the worker reallocation (OP) term, with essentially none attributed to a firm selection term (since there is no firm entry and exit in MPV2016). This is nearly five times the 20 percent share attributed to worker reallocation in the data and in the baseline model. In the single-shock (ω) calibration, both models attribute a more modest share to worker reallocation (7.2 percent for the baseline model, 0.1 percent for MPV2016 with c1=5), and the difference narrows considerably. The contrast thus stems from the interaction of firm dynamics with multiple aggregate shocks: allowing for endogenous firm entry and exit is critical to prevent the model from overstating the role of the job ladder.&lt;/p&gt;
&lt;h3 id="q12-how-persistent-is-the-impact-of-recessions-on-the-job-ladder-based-on-the-model-simulations"&gt;Q12. How persistent is the impact of recessions on the job ladder, based on the model simulations?&lt;/h3&gt;
&lt;p&gt;A: The paper simulates the structural decomposition of labor productivity starting from each of seven post-war British recessions (defined by two consecutive quarters of negative GDP growth). On average across these recessions, the aggregate productivity shock ln(ωt) is close to its pre-recession level by four years after the recession onset. However, the overall employment-weighted average productivity E[LP bar_t(ln p)] — reflecting workers&amp;rsquo; position on the job ladder — remains clearly below its pre-recession value at the four-year horizon, indicating persistent misallocation. The OP interaction term accounts for approximately 20 percent of the total drop in the employment-weighted productivity measure three years after a typical recession onset. Through the model&amp;rsquo;s lens, the Great Recession is a large recession but not an outlier relative to the historical distribution.&lt;/p&gt;
&lt;h3 id="q13-what-does-the-counterfactual-with-countercyclical-unemployment-benefits-reveal-about-the-tradeoff-between-firm-selection-and-worker-reallocation"&gt;Q13. What does the counterfactual with countercyclical unemployment benefits reveal about the tradeoff between firm selection and worker reallocation?&lt;/h3&gt;
&lt;p&gt;A: When the flow value of unemployment is made countercyclical (falling in recessions, rising in expansions — mimicking US unemployment insurance extension programs), the model generates a sign reversal in the firm selection (&amp;ldquo;firm ladder&amp;rdquo;) component. With countercyclical b, the unemployment value rises in recessions, which raises the minimum wage firms must offer and raises the exit threshold pE,t: fewer low-productivity firms survive, improving the composition of active firms. However, countercyclical benefits also amplify the slowdown in job-to-job reallocation: the higher value of unemployment reduces workers&amp;rsquo; willingness to accept job offers, and all firms cut recruitment since optimal wage contracts must rise. The OP interaction term therefore falls more sharply than in the baseline model. The counterfactual with ϵb,ω ∈ {−100, −50} finds that the positive &amp;ldquo;firm ladder&amp;rdquo; effect dominates on net, so the overall allocation of workers to firms improves relative to the baseline after a typical recession under countercyclical unemployment benefits.&lt;/p&gt;
&lt;h3 id="q14-what-is-the-numerical-solution-method-and-why-are-standard-linearization-approaches-inapplicable"&gt;Q14. What is the numerical solution method, and why are standard linearization approaches inapplicable?&lt;/h3&gt;
&lt;p&gt;A: The model is solved in two steps. First, aggregate shocks are shut down and the steady-state rank-monotonic equilibrium is solved numerically by discretizing the firm productivity process (401 grid points via Tauchen&amp;rsquo;s method) and iterating on the value function and the employment-weighted productivity measure until convergence. Second, aggregate shocks are reintroduced using a simulation-based approach adapted from Krusell and Smith (1998): the employment-weighted distribution of productivity is summarized by Nm = 2 moments (plus the unemployment rate), and the value functions are parameterized as polynomials in the aggregate state, with coefficients updated by regression until convergence. Standard linearization methods (Reiter 2009) are inapplicable because the endogenous entry-exit decision creates a kink (discontinuity) in value functions at the productivity threshold pE, making first-order approximations around the steady state inaccurate. Accuracy tests based on den Haan (2010) show that the polynomial approximation generates errors of at most 0.065 percent for value functions and at most 1 percentage point for the unemployment rate across simulation paths.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;1. Rank-Monotonic Equilibrium (RME)&lt;/strong&gt;
A recursive equilibrium in which the optimal state-contingent employment contract V(p,ω,L) offered by a firm is weakly increasing in the firm&amp;rsquo;s current productivity realization p, for all aggregate states (ω,L). This property implies that the job ladder maps one-for-one onto the firm productivity ladder: workers always prefer to work at more productive firms. The paper shows this property holds under a sufficient convexity condition on hiring costs (hc&amp;rsquo;&amp;rsquo;(h)/c&amp;rsquo;(h) ≥ 1) and first-order stochastic dominance of the productivity process.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Size-Independence&lt;/strong&gt;
The property that a firm&amp;rsquo;s optimal policies — the hiring rate h(p), the employment contract V(p), and the entry/exit decision χ(p) — are all independent of the firm&amp;rsquo;s current employment size n. This follows from constant returns to scale in production and hiring, which implies that firm profits are linear in employment. Size-independence reduces the model&amp;rsquo;s relevant state space to the employment-weighted distribution of firm productivity, enabling tractability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Employment-Weighted Distribution of Firm Productivity (L_t(p))&lt;/strong&gt;
The measure recording, for each productivity level p, the total employment at firms with productivity at most p. This is the sufficient statistic for the state of the job ladder at any point in time: combined with the aggregate shock ω, it determines all equilibrium policy functions and value functions. In the model, it replaces the full joint distribution of firm productivity and employment size that would otherwise be required.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;4. OP Decomposition (Olley-Pakes Decomposition)&lt;/strong&gt;
The decomposition of the aggregate employment-weighted labor productivity index LPt into: (a) the unweighted average firm productivity LPt-bar, which summarizes the productivity of active firms (the support of the job ladder); and (b) an interaction term OPt, the covariance between employment shares and firm-level productivity, which measures how well workers are allocated across the productivity distribution (i.e., how high up the ladder workers sit given the set of active firms). In the model, (a) maps to E[KP bar_t(ln p)] and (b) maps to OPt = E[LP bar_t(ln p)] − E[KP bar_t(ln p)].&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;5. Contract Posting&lt;/strong&gt;
The wage-setting protocol in which each firm commits upon entry to a full state-contingent employment contract — a schedule mapping each future realization of aggregate and idiosyncratic productivity to a wage and continuation decision — and is bound by an equal treatment constraint to offer the same contract to all employees. Workers cannot renegotiate based on outside offers. This protocol produces a well-defined closed-form for the optimal contract in an RME and differs from alternating-offer bargaining (Nash bargaining) in that the bargaining weights are endogenous rather than fixed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;6. Firm-Workers&amp;rsquo; Joint Surplus (S_t(p))&lt;/strong&gt;
The total present discounted value accruing to the firm-worker pair: firm profits per worker plus the contract value promised to workers. Because utility is transferable (risk neutrality) and the firm fully commits to its contract, this surplus depends only on the firm&amp;rsquo;s current productivity and the aggregate state — not on the promised contract value V. The surplus S_t(p) is the key object determining firm entry/exit (the firm continues if and only if S_t(p) ≥ U_t) and optimal hiring (the marginal return to an additional hire equals S_t(p) − V(p)).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;7. Cleansing vs. Anti-Cleansing Firm Selection&lt;/strong&gt;
In models with endogenous firm entry and exit, a negative aggregate shock can either raise or lower the productivity threshold for firm survival. &amp;ldquo;Cleansing&amp;rdquo; refers to the standard mechanism where a negative productivity shock raises the exit threshold, selecting out low-productivity firms and improving the average quality of survivors. &amp;ldquo;Anti-cleansing&amp;rdquo; (as in the baseline multi-shock calibration) occurs when separation rate or hiring cost shocks lower the optimal wage contract and reduce the exit threshold, allowing less productive firms to survive and worsening average firm productivity.&lt;/p&gt;</description></item><item><title>Firm idiosyncratic risk and productivity investment: Macroeconomic implications</title><link>https://macropaperwarehouse.com/papers/firm-idiosyncratic-risk-and-productivity-investment-macroeconomic-implications/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/firm-idiosyncratic-risk-and-productivity-investment-macroeconomic-implications/</guid><description>&lt;p&gt;This paper quantifies how idiosyncratic firm-level risk affects aggregate output, TFP, and firm life-cycle growth in an environment where firm productivity evolves endogenously through risky investment. The paper embeds endogenous productivity investment into a Lucas span-of-control model with risk-averse firm owners and endogenous entry and exit, and studies the effects of mean-preserving increases in the variance of returns to productivity investment. A mean-preserving increase in the variance of firm productivity shocks that raises the firm exit rate by 10% (from 0.10 to 0.11) is estimated to cause a 0.73% decline in output, a 0.38% decline in measured TFP, and a 3.69% decline in firm productivity investment; these elasticities remain approximately constant in the empirically relevant range. The driving force is that risk-averse firm owners reduce their risky productivity investment as variance rises; if capital financing constraints are present—as is common in developing economies—these effects are amplified and the increase in uncertainty may also slow firm life-cycle growth. Previously circulated as &amp;ldquo;Uncertainty, Firm Lifecycle Growth, and Aggregate Productivity.&amp;rdquo;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Summary based on a working paper version, AI-assisted and human-reviewed. See the linked published article for the authoritative version.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-distinguishes-this-paper-from-standard-models-of-firm-misallocation"&gt;Q1. What distinguishes this paper from standard models of firm misallocation?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Unlike the bulk of firm misallocation literature (Hsieh-Klenow 2009; Gopinath et al. 2017; Sraer-Thesmar 2023), which takes firm productivity as exogenous, this paper models productivity as an endogenous outcome of risky investment, so that idiosyncratic uncertainty affects allocative efficiency not only through selection effects but also through its discouragement of productivity investment by risk-averse owners.&lt;/strong&gt; The paper incorporates endogenous productivity investment into a standard Lucas span-of-control model, allowing the model to capture how higher uncertainty reduces the incentive to invest in productivity, on top of any selection effects from the exit option.&lt;/p&gt;
&lt;h3 id="q2-what-are-the-two-opposing-effects-of-higher-idiosyncratic-risk"&gt;Q2. What are the two opposing effects of higher idiosyncratic risk?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Higher idiosyncratic firm-level risk has two opposing effects on aggregate productivity: (i) a selection effect—a mean-preserving increase in variance leads to stronger selection and raises the productivity of survivors while reallocating exiters to alternative productive uses—that tends to raise average productivity; and (ii) a productivity investment effect—risk-averse owners reduce risky productivity investment in response to higher variance—that tends to reduce aggregate productivity and firm life-cycle growth.&lt;/strong&gt; The paper shows quantitatively that the productivity investment effect dominates in the baseline calibration, so that higher idiosyncratic risk reduces output and TFP despite positive selection effects.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-main-quantitative-findings"&gt;Q3. What are the main quantitative findings?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;A mean-preserving increase in the variance of firm productivity shocks calibrated to raise the firm exit rate by 10% (from 0.10 to 0.11) results in a 0.73% decline in output, a 0.38% decline in measured TFP, and a 3.69% decline in firm productivity investment; these elasticities remain approximately constant in the empirically relevant range.&lt;/strong&gt; The exit-rate increase from 0.10 to 0.11 is also associated with a 7.5% increase in the job destruction rate and a 14.6% increase in the standard deviation of firm growth rates—the latter is less than one-fifth of the increases in these risk measures observed when comparing India or Mexico to the U.S.&lt;/p&gt;
&lt;h3 id="q4-how-do-capital-financing-constraints-interact-with-the-results"&gt;Q4. How do capital financing constraints interact with the results?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;When firms face capital financing constraints—as is common in developing economies—the negative effects of higher idiosyncratic risk are amplified and the increase in uncertainty may also slow firm life-cycle growth.&lt;/strong&gt; The mechanism is that constrained firms must rely more heavily on internal financing, making risk-averse owners even more sensitive to increases in variance. The paper implies that the macro-financial implications of idiosyncratic risk are more severe in developing economies where both idiosyncratic risk levels and financing constraints are greater—consistent with cross-country patterns of firm growth dynamics.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;productivity investment&lt;/strong&gt; : endogenous spending by firms on activities that shift their productivity process; in the model, this investment exposes firm owners to idiosyncratic risk via the innovation in the productivity process; the key margin through which higher uncertainty reduces aggregate productivity and output.
&lt;strong&gt;mean-preserving increase in variance&lt;/strong&gt; : a statistical experiment that increases the spread of the distribution of returns to productivity investment while leaving the mean unchanged; used here to isolate the pure risk effect on firm behavior and aggregate outcomes from any change in expected returns.
&lt;strong&gt;span-of-control model&lt;/strong&gt; : the Lucas (1978) model of firm size distribution with decreasing returns to scale in the entrepreneurial input; used as the production environment; extended here by adding endogenous productivity investment and endogenous entry and exit.&lt;/p&gt;</description></item><item><title>Firm Quality Dynamics and the Slippery Slope of Credit Intervention</title><link>https://macropaperwarehouse.com/papers/firm-quality-dynamics-and-the-slippery-slope-of-credit-intervention/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/firm-quality-dynamics-and-the-slippery-slope-of-credit-intervention/</guid><description>&lt;p&gt;Crises have cleansing effects—low-quality firms face greater financial shortfalls and invest less than high-quality firms—but public credit support dampens these effects by reducing financing cost differentials, distorting the firm quality distribution downward and reducing total productivity. This trade-off between preserving output capacity and distorting quality determines the optimal size of intervention. The distortionary effects are self-perpetuating: a downward bias in quality necessitates interventions of greater scale in future crises, implying further distortions—a &amp;ldquo;slippery slope.&amp;rdquo; The distortions are amplified by expectations: because low-quality firms expect underpriced government funding in future crises, their Tobin&amp;rsquo;s q is biased upward, leading them to overinvest even in normal times, while high-quality firms may underinvest. A low interest rate environment exacerbates the distortionary effects because the low yield on savings discourages firms from accumulating precautionary internal liquidity against crises.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Summary of a forthcoming paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-are-the-cleansing-effects-of-crises-and-how-does-credit-intervention-dampen-them"&gt;Q1. What are the cleansing effects of crises and how does credit intervention dampen them?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Crises have cleansing effects because low-quality firms face tighter financial constraints and have lower Tobin&amp;rsquo;s q, causing them to invest less than high-quality firms; public credit support reduces this differential, preserving overall production capacity but distorting the quality distribution downward.&lt;/strong&gt; The model follows the limited-commitment literature (Kehoe-Levine, Kiyotaki-Moore, Rampini-Viswanathan): firms differ in productive capital quality that also serves as collateral. Government intervention is valued because the government has superior enforcement ability compared to private investors, but its credit support cannot be perfectly priced by quality—due to informational limits or political constraints—so it pulls financing costs of high- and low-quality firms closer together, dampening the cleansing mechanism.&lt;/p&gt;
&lt;h3 id="q2-what-is-the-slippery-slope-mechanism"&gt;Q2. What is the &amp;ldquo;slippery slope&amp;rdquo; mechanism?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The slippery slope arises because the downward bias in the quality distribution induced by one intervention necessitates larger interventions in future crises, generating a ratchet toward ever-larger public credit support.&lt;/strong&gt; After intervention, high-quality firms accumulate capital less rapidly than they would absent intervention, while low-quality firms&amp;rsquo; capital shares remain higher than in the laissez-faire equilibrium. The resulting lower aggregate productivity means that future crises are more severe in terms of output loss, requiring a larger optimal intervention, which in turn further distorts the quality distribution.&lt;/p&gt;
&lt;h3 id="q3-how-do-expectations-of-future-intervention-amplify-the-distortions"&gt;Q3. How do expectations of future intervention amplify the distortions?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Because low-quality firms expect underpriced credit support in future crises, their Tobin&amp;rsquo;s q is biased upward, motivating them to overinvest even in normal times; simultaneously, high-quality firms may underinvest because their Tobin&amp;rsquo;s q may fall below the first-best level.&lt;/strong&gt; The self-perpetuating distortion thus operates through both the crisis-time reallocation channel and the pre-crisis investment channel, amplifying the divergence from the efficient allocation relative to a setting with no anticipation effects.&lt;/p&gt;
&lt;h3 id="q4-why-does-a-low-interest-rate-environment-exacerbate-the-distortionary-effects"&gt;Q4. Why does a low interest rate environment exacerbate the distortionary effects?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;A low interest rate environment exacerbates the distortionary effects of credit intervention because the low yield on savings discourages high-quality firms from accumulating precautionary internal liquidity against crises, causing them to invest less in crises and requiring a greater scale of credit support.&lt;/strong&gt; Low-quality firms, expecting underpriced government funding, have even less incentive to self-insure through savings when interest rates are low, further worsening the quality distribution. The paper&amp;rsquo;s findings echo cautions against ultra-low interest rates (Brunnermeier and Koby, 2018; Quadrini, 2020) by providing a distinct mechanism operating through firm quality dynamics.&lt;/p&gt;
&lt;h3 id="q5-can-intervention-be-welfare-improving-despite-the-distortions"&gt;Q5. Can intervention be welfare-improving despite the distortions?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The paper shows that when carefully designed, intervention can improve welfare even though it generates distortionary effects on the firm quality distribution—the trade-off between preserving production capacity and distorting quality determines the optimal size of intervention.&lt;/strong&gt; This framing does not suggest intervention should be avoided, but that its optimal scale requires balancing the quantity-preserving benefit against the quality-distorting cost. The paper previously circulated as &amp;ldquo;The Distortionary Effects of Central Bank Direct Lending on Firm Quality Dynamics.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;cleansing effect of crises&lt;/strong&gt; : the tendency for crises to reduce the investment of low-quality firms relative to high-quality firms through tighter financial constraints, reallocating capital toward higher-productivity uses; credit intervention dampens this by reducing the financing cost differential.
&lt;strong&gt;slippery slope of intervention&lt;/strong&gt; : the self-perpetuating dynamic in which intervention-induced downward distortion of the quality distribution necessitates larger interventions in future crises, generating a ratchet toward ever-larger public credit support.
&lt;strong&gt;credit mispricing&lt;/strong&gt; : the inability of public credit support to differentiate financing costs by firm quality, arising from informational limits or political constraints on discriminatory treatment; the proximate source of the quality-distribution distortion.&lt;/p&gt;</description></item><item><title>Firm Responses and Wage Effects of Foreign Demand Shocks with Fixed Labor Costs and Monopsony</title><link>https://macropaperwarehouse.com/papers/firm-responses-and-wage-effects-of-foreign-demand-shocks-with-fixed-labor-costs-and-monopsony/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/firm-responses-and-wage-effects-of-foreign-demand-shocks-with-fixed-labor-costs-and-monopsony/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; The paper asks three related questions in the context of Belgium, a small open economy: (1) What do firms&amp;rsquo; responses to demand shocks reveal about their cost structures? (2) What are the worker and wage impacts of foreign demand shocks? (3) How sensitive are the aggregate wage effects of foreign demand shifts to firms&amp;rsquo; cost structures and imperfect competition in the labor market?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The analysis combines administrative micro-data from Belgium for 2002–2014, provided by the National Bank of Belgium. The linked dataset covers 995,739 firm-year observations from private, non-financial firms with at least one FTE employee, and integrates: (a) a Business-to-Business (B2B) VAT transactions registry capturing all annual domestic firm-to-firm sales above €250; (b) customs records and intra-EU declarations for imports and exports at the 8-digit product level; (c) annual accounts containing data on sales, labor costs, intermediate inputs, capital, and firm characteristics; and (d) employer-employee matched data from the Belgian social security administration (BCSS) for a random sample of 500,000 workers in firms with 10 or more FTE employees, covering 2003–2014.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identification Strategy.&lt;/strong&gt; To isolate variation in firms&amp;rsquo; sales driven by foreign demand rather than supply-side factors, the authors construct a firm-specific foreign demand instrument following Hummels et al. (2014) and Dhyne et al. (2021). The instrument is the weighted average of changes in world import demand facing a firm, using lagged export shares as weights and excluding Belgian imports from the world import measure. Crucially, the instrument captures both direct foreign demand exposure (for exporters) and indirect exposure through the domestic production network — including the foreign demand shocks passing through to upstream domestic suppliers via buyer-supplier links. Firm and industry-year fixed effects control for time-invariant heterogeneity and industry-level trends.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Key Empirical Facts.&lt;/strong&gt; Within-firm analysis over four-year windows finds that intermediate input purchases respond nearly proportionally to changes in sales (slope coefficient 0.82), while labor costs respond less than proportionally (slope coefficient 0.57). The less-than-proportional response of labor costs — with the employment slope of 0.48 and the average wage slope of 0.09 — is consistent with sizable fixed overhead costs in labor inputs and upward-sloping labor supply curves. Output prices co-move more with input prices than with average wages, consistent with labor constituting a smaller share of variable costs than intermediate inputs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;IV Estimates of Firm Responses.&lt;/strong&gt; In response to a foreign demand shock inducing a 10 percent instantaneous increase in a firm&amp;rsquo;s sales, the firm&amp;rsquo;s cumulative sales over four years increase by approximately 7.6 percent (balanced panel). Over the same four-year horizon, total input purchases increase by about 7.0–7.8 percent, while labor costs increase by only 3.5–4.1 percent — a substantially less-than-proportional response. Roughly one-quarter of the labor cost change comes from changes in average wages rather than employment changes. Domestic input purchases increase by 5.3–6.0 percent, indicating that firms pass on a large share of foreign demand shocks to their domestic suppliers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural Parameters.&lt;/strong&gt; The implied IV estimate of the labor cost elasticity with respect to sales is 0.53 (standard error 0.08), statistically significantly below one. The implied elasticity of total input purchases is 1.05 (standard error 0.15), close to one, so the fixed share of intermediate inputs is approximately zero. The labor supply elasticity estimated from the ratio of wage and employment responses is approximately 3.9 in the full sample and 2.3 in the stayer subsample; the implied wage markdown is 21 percent and 30 percent respectively. Incorporating upward-sloping labor supply into equation (15), the estimated share of total labor inputs that is fixed overhead is approximately 53 percent. By comparison, the fixed share of total costs (labor and intermediate inputs combined) is approximately 29 percent in Belgium — higher than the 18–22 percent found in U.S. data (De Loecker et al. 2020) and the 20 percent found in U.S. manufacturing plants (Ederhof et al. 2021).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;General Equilibrium Counterfactuals.&lt;/strong&gt; The authors parameterize and solve a small open economy general equilibrium model with monopsonistic competition in labor markets, monopolistic competition in product markets, and fixed and variable labor and intermediate input costs. Using the Dekle-Eaton-Kortum (2007) &amp;ldquo;hat algebra&amp;rdquo; technique, they simulate a 5 percent increase in foreign tariffs on all Belgian exports and compare four counterfactual economies: (1) baseline Belgium with fixed costs and imperfect labor market competition (ε = 3.9); (2) fixed costs and perfectly elastic labor supply (ε = ∞); (3) no fixed costs with imperfect competition; (4) no fixed costs and perfectly competitive labor markets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings on Wages.&lt;/strong&gt; In the baseline Belgian economy, a 5 percent increase in foreign tariffs produces a 4.9 percent fall in the average real wage. With fixed costs but perfectly elastic labor supply, the real wage falls by 4.8 percent — nearly identical. With upward-sloping labor supply but no fixed costs, the real wage falls by only 3.0 percent; without fixed costs and with perfectly competitive labor supply, the fall is only 2.8 percent. The paper concludes that fixed overhead costs in labor substantially amplify real wage declines, while incorporating upward-sloping labor supply appears quantitatively less consequential for aggregate wage outcomes. Standard models that assume no fixed costs and perfectly elastic labor supply — the typical modeling choice in the trade literature — may substantially understate (by roughly 43–75 percent of the true effect) the aggregate wage decline from a negative foreign demand shock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanism.&lt;/strong&gt; Fixed overhead costs reduce labor&amp;rsquo;s share of variable costs. When labor is a smaller share of variable costs, output prices are less sensitive to changes in wages. With a fixed aggregate labor supply, the economy must lower prices through wage reductions to restore equilibrium after a negative demand shock; the required wage decline is larger when fixed labor costs are taken into account. The findings are robust to adjustment cost specifications, a nested logit extension of the labor market model, and controlling for location-year fixed effects and import price changes.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-two-motivating-empirical-facts-about-belgian-firms-does-the-paper-establish"&gt;Q1. What two motivating empirical facts about Belgian firms does the paper establish?&lt;/h3&gt;
&lt;p&gt;A1: First, within-firm four-year changes show that intermediate input purchases respond nearly proportionally to changes in sales (slope coefficient 0.82), while labor costs respond less than proportionally (slope coefficient 0.57). The labor cost response decomposes into an employment slope of 0.48 and a wage slope of 0.09. Second, output prices co-move more strongly with input (intermediate goods) prices than with average wages, consistent with labor constituting a smaller share of variable costs than intermediate inputs.&lt;/p&gt;
&lt;h3 id="q2-how-does-the-instrument-for-foreign-demand-shocks-capture-indirect-exposure-through-production-networks"&gt;Q2. How does the instrument for foreign demand shocks capture indirect exposure through production networks?&lt;/h3&gt;
&lt;p&gt;A2: The instrument for firm k is a weighted average of changes in world import demand, where the weights reflect both the firm&amp;rsquo;s own direct export shares across countries and products and the firm&amp;rsquo;s indirect export exposure through its domestic buyers&amp;rsquo; export shares. The term H̃_{kn,t-1} captures the share of firm k&amp;rsquo;s total sales purchased by firm n directly and indirectly through all upstream chains. This means even non-exporting firms receive a non-zero instrument through their sales to directly-exporting firms. In fact, non-directly-exporting firms sell on average nearly 10 percent of their output indirectly to foreign markets.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-estimated-magnitude-of-the-labor-supply-elasticity-facing-belgian-firms-and-what-does-it-imply-for-wage-markdowns"&gt;Q3. What is the estimated magnitude of the labor supply elasticity facing Belgian firms, and what does it imply for wage markdowns?&lt;/h3&gt;
&lt;p&gt;A3: In the full main estimation sample (balanced panel), the IV estimate of the firm-specific labor supply elasticity is approximately 3.9, implying a wage markdown of about 21 percent relative to the marginal revenue product of labor. In the stayer subsample (incumbent workers only, holding workforce composition fixed), the estimated labor supply elasticity is approximately 2.3, implying a markdown of about 30 percent. The paper can reject perfect competition (infinite elasticity, zero markdown) at a significance level of 0.06 in the full sample and 0.001 in the stayer sample using the closure method.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-estimated-labor-cost-elasticity-with-respect-to-demand-driven-sales-changes-and-what-does-it-imply-about-fixed-labor-costs"&gt;Q4. What is the estimated labor cost elasticity with respect to demand-driven sales changes, and what does it imply about fixed labor costs?&lt;/h3&gt;
&lt;p&gt;A4: The IV estimate of the labor cost elasticity with respect to sales is 0.528 (standard error 0.085), statistically significantly below one. If labor supply were perfectly elastic, this would directly imply a fixed labor cost share of approximately 47 percent. Incorporating the estimated upward-sloping labor supply curve through equation (15), the model implies that approximately 53 percent of total labor inputs are fixed overhead. For context, occupational data from Belgium&amp;rsquo;s 2014 Structure of Earnings Survey shows that clerical support workers and managers together account for 21 percent of total earnings, and adding professionals raises this to 51 percent — broadly consistent with the estimated fixed share.&lt;/p&gt;
&lt;h3 id="q5-what-does-the-estimated-elasticity-of-input-purchases-with-respect-to-sales-imply-about-fixed-intermediate-input-costs"&gt;Q5. What does the estimated elasticity of input purchases with respect to sales imply about fixed intermediate input costs?&lt;/h3&gt;
&lt;p&gt;A5: The IV estimate of the elasticity of total input purchases with respect to sales is 1.050 (standard error 0.150), close to one. The implied fixed share of total intermediate inputs is therefore approximately zero. However, there is substantial heterogeneity by input type: purchases from the manufacturing sector (roughly half of all input purchases) have an elasticity close to one, whereas service-sector inputs (roughly 30 percent of total input purchases) have an implied fixed cost share of approximately 36 percent, with a size-weighted average cumulative response of 4.3 percent against a total cumulative sales increase of 6.7 percent.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-paper-rule-out-alternative-explanations-for-the-less-than-proportional-response-of-labor-costs"&gt;Q6. How does the paper rule out alternative explanations for the less-than-proportional response of labor costs?&lt;/h3&gt;
&lt;p&gt;A6: The paper considers three main alternatives. First, adjustment costs: even in the presence of labor adjustment costs, under a homothetic constant-returns production function a permanent shock should eventually produce a proportional labor response. The paper focuses on four-year cumulative responses where firm responses change little after the first couple of years, and shows identification of fixed costs holds even in models with quadratic or Calvo-style adjustment costs. Second, a non-homothetic CES production function without fixed costs: Appendix B.3 shows that such a specification predicts that if the labor cost elasticity is below one, the input purchase elasticity must be above one — at odds with the data, which shows the input purchase elasticity is close to one while the labor cost elasticity is well below one. Third, variable markups: a uniform markup change would reduce both elasticities proportionally, not create the large gap between labor cost and input purchase elasticities observed.&lt;/p&gt;
&lt;h3 id="q7-why-are-firms-domestic-suppliers-affected-by-foreign-demand-shocks-and-how-large-are-the-pass-through-effects"&gt;Q7. Why are firms&amp;rsquo; domestic suppliers affected by foreign demand shocks, and how large are the pass-through effects?&lt;/h3&gt;
&lt;p&gt;A7: Firms pass on foreign demand shocks to their domestic suppliers through buyer-supplier production network links. When a foreign demand shock increases a firm&amp;rsquo;s sales by 10 percent instantaneously, its domestic input purchases increase cumulatively by approximately 5.3–6.0 percent over four years. Total input purchases increase by 7.0–7.8 percent over the same period; the difference between total and domestic input purchases reflects service inputs (which have smaller responses) and the composition of imported versus domestic inputs.&lt;/p&gt;
&lt;h3 id="q8-what-is-the-aggregate-real-wage-effect-of-a-5-percent-increase-in-foreign-tariffs-on-belgian-exports-in-the-baseline-model"&gt;Q8. What is the aggregate real wage effect of a 5 percent increase in foreign tariffs on Belgian exports in the baseline model?&lt;/h3&gt;
&lt;p&gt;A8: In the baseline counterfactual representing the actual Belgian economy (with fixed overhead costs and labor supply elasticity ε = 3.9), a uniform 5 percent increase in foreign tariffs on all Belgian exports produces a 4.9 percent fall in the average real wage. The median firm reduces output by 3.8 percent, marginal costs by 4.8 percent, and wages by 7.9 percent. The fall in wages is driven by a general equilibrium mechanism: since the foreign price is exogenous and trade balance must hold, wages are the key adjusting margin.&lt;/p&gt;
&lt;h3 id="q9-how-much-does-the-modeling-of-fixed-overhead-costs-versus-imperfect-labor-market-competition-matter-for-the-aggregate-wage-counterfactual"&gt;Q9. How much does the modeling of fixed overhead costs versus imperfect labor market competition matter for the aggregate wage counterfactual?&lt;/h3&gt;
&lt;p&gt;A9: Fixed overhead costs account for nearly all of the amplification relative to the standard model. With fixed costs but perfectly elastic labor supply, the real wage falls 4.8 percent — almost identical to the 4.9 percent in the baseline. Without fixed costs but with the estimated upward-sloping labor supply, the fall is only 3.0 percent. Without either, the fall is 2.8 percent. Thus, incorporating fixed overhead costs in labor raises the estimated wage decline by approximately 1.9 percentage points, while incorporating imperfect labor market competition adds only about 0.1 percentage points. The paper concludes that fixed overhead costs, not monopsony, are the essential feature for accurately predicting tariff impacts on wages.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-mechanism-by-which-fixed-overhead-costs-amplify-the-aggregate-wage-decline-from-a-negative-demand-shock"&gt;Q10. What is the mechanism by which fixed overhead costs amplify the aggregate wage decline from a negative demand shock?&lt;/h3&gt;
&lt;p&gt;A10: Fixed overhead costs reduce the share of labor in firms&amp;rsquo; total variable costs. When labor constitutes a smaller fraction of variable costs, output prices are less sensitive to changes in wages. With aggregate labor supply fixed, the economy restores equilibrium after a negative demand shock by reducing prices through wage cuts. To achieve the same magnitude of price reduction when labor is a smaller fraction of variable costs, wages must fall by a larger amount — amplifying the aggregate wage impact. Fixed overhead costs in labor also make foreign inputs relatively more important in variable costs, as shown empirically in Appendix D.1.&lt;/p&gt;
&lt;h3 id="q11-is-the-conclusion-about-the-relative-importance-of-fixed-costs-versus-labor-market-imperfections-robust-to-alternative-specifications-of-the-labor-market"&gt;Q11. Is the conclusion about the relative importance of fixed costs versus labor market imperfections robust to alternative specifications of the labor market?&lt;/h3&gt;
&lt;p&gt;A11: Yes. The paper extends the model to a nested logit structure for worker preferences (following Lamadon et al. 2022), which allows Belgium to contain multiple labor markets (defined as industry-region nests), permits heterogeneous markdowns across markets, and is still identified from the data. Empirically, incorporating multiple labor markets and heterogeneous markdowns does not quantitatively alter the aggregate counterfactual predictions for the wage effects of foreign demand shocks.&lt;/p&gt;
&lt;h3 id="q12-are-heterogeneous-responses-to-the-foreign-demand-shock-observed-across-exporters-importers-and-domestic-only-firms"&gt;Q12. Are heterogeneous responses to the foreign demand shock observed across exporters, importers, and domestic-only firms?&lt;/h3&gt;
&lt;p&gt;A12: The paper finds no systematic differences in the elasticities of labor cost and input purchases between firms that trade internationally and those that do not. This implies that exporters and importers have higher absolute fixed costs (consistent with fixed export and import costs) but comparable fixed cost shares — since these firms tend to be larger and thus spread higher absolute fixed costs over larger output volumes.&lt;/p&gt;
&lt;h3 id="q13-do-the-findings-about-fixed-overhead-costs-extend-beyond-foreign-demand-shocks"&gt;Q13. Do the findings about fixed overhead costs extend beyond foreign demand shocks?&lt;/h3&gt;
&lt;p&gt;A13: Yes. The paper shows in Appendix D.4 that a uniform 5 percent reduction in the productivity of all Belgian manufacturing firms generates qualitatively and quantitatively similar conclusions: fixed overhead costs amplify the predicted wage effects of domestic productivity shocks, while imperfect competition in the labor market matters to a lesser but still meaningful extent.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Fixed Overhead Costs (Fixed Labor Costs / Fixed Intermediate Input Costs):&lt;/strong&gt; In the paper&amp;rsquo;s model, each firm has firm-specific fixed overhead input requirements for labor (denoted ℓ̄_k^f) and intermediate inputs (denoted q̄_k^f) that must be satisfied regardless of the firm&amp;rsquo;s output level. These fixed requirements are separate from the variable inputs used in production. Fixed labor costs may reflect administration, worker management, facility maintenance, and other tasks that do not directly translate into output. Fixed intermediate input costs include waste management, accounting services, and electricity payments that occur irrespective of sales. The share of total labor inputs that is fixed is identified by how much less than proportionally labor costs respond to demand-driven changes in sales.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Monopsonistic Competition in the Labor Market:&lt;/strong&gt; The paper models each firm as facing an upward-sloping firm-specific labor supply curve arising from workers&amp;rsquo; heterogeneous idiosyncratic preferences over non-wage firm attributes (amenities). Because workers&amp;rsquo; idiosyncratic tastes are private information, firms cannot price-discriminate and thus face an increasing marginal cost of labor. Each firm is infinitesimal within the aggregate labor market but has wage-setting power at the firm level. This gives rise to a constant-elasticity firm-level labor supply curve ℓ_k = A_k w_k^ε, where ε is the labor supply elasticity facing the firm.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wage Markdown:&lt;/strong&gt; The firm&amp;rsquo;s equilibrium wage is marked down relative to the marginal revenue product of labor by the factor ε/(1+ε), which is less than one when ε is finite. With a labor supply elasticity of 3.9, the implied markdown is approximately 21 percent; with a supply elasticity of 2.3 (stayer sample), the markdown is approximately 30 percent. Perfect competition corresponds to ε = ∞ and a markdown of zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Labor Cost Elasticity:&lt;/strong&gt; The elasticity of a firm&amp;rsquo;s total labor cost with respect to a demand-driven change in the firm&amp;rsquo;s sales, as derived from the model&amp;rsquo;s comparative statics (equation 15). This elasticity depends on both the variable share of labor inputs (ℓ_k^v / ℓ_k) and the labor supply elasticity ε. It lies strictly between zero (all labor fixed) and one (all labor variable), and is declining in ε for a given variable share. The paper estimates this elasticity at 0.528 via IV, implying substantial fixed overhead in labor.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Total Foreign Demand Shock:&lt;/strong&gt; The firm-level measure of foreign demand used as an instrument, defined as the weighted average of changes in world import demand (excluding Belgium) across country-product pairs, where the weights reflect both the firm&amp;rsquo;s own lagged direct export shares and its indirect exposure through the domestic production network (via the Leontief inverse matrix H̃). This measure captures both direct exporter exposure and indirect upstream exposure for non-exporting firms that supply to exporters.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Indirect Export Exposure:&lt;/strong&gt; The share of a firm&amp;rsquo;s output that reaches foreign markets indirectly through sales to domestic buyers who subsequently export. Defined recursively: the total export share of firm k equals its direct export revenue share plus the sum over all domestic buyers of the product of k&amp;rsquo;s revenue share from that buyer and the buyer&amp;rsquo;s own total export share. Even non-direct-exporting firms sell on average approximately 10 percent of their output indirectly to foreign markets in the Belgian data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Dekle-Eaton-Kortum Hat Algebra:&lt;/strong&gt; A technique for solving general equilibrium counterfactuals in trade models by expressing all outcomes as proportional changes (&amp;ldquo;hats&amp;rdquo;) relative to the observed equilibrium, without needing to recover the underlying structural parameters. The paper uses this approach to compute counterfactual wages under alternative tariff scenarios, holding fixed the observed firm-level expenditure shares from the reference year (2012) while allowing parameters such as productivity and technology weights to vary across counterfactual economies to rationalize identical observed firm-level observables.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Worker Rents:&lt;/strong&gt; In the monopsony model, inframarginal workers earn rents defined as the excess return over what would be required to make them indifferent between employers. These rents arise because firms cannot price-discriminate across workers with heterogeneous amenity valuations. The additional rents accruing to workers from a demand-driven increase in firm sales decompose into: (1) wage increases for incumbent workers multiplied by current employment, (2) rents for new hires (the excess of their wage bill over the amount required to induce them to switch to the expanding firm), and (3) a correction term related to the fraction of the labor cost increase borne by expanding employment rather than wages.&lt;/p&gt;</description></item><item><title>From Doubt to Devotion: Trials and Learning-Based Pricing</title><link>https://macropaperwarehouse.com/papers/from-doubt-to-devotion-trials-and-learning-based-pricing/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/from-doubt-to-devotion-trials-and-learning-based-pricing/</guid><description>&lt;p&gt;This paper studies a dynamic mechanism design problem in which an informed seller sells an experience good to a skeptical buyer who learns about the product through consumption. The central question is: how does a seller leverage proprietary data about product-buyer match quality together with the buyer&amp;rsquo;s ability to learn, and what are the welfare implications in equilibrium?&lt;/p&gt;
&lt;p&gt;The model features a seller who privately observes a binary match quality (theta in {H, L}) between their service and the buyer. The buyer does not observe match quality and has an initially unknown private value v for the good, drawn from a Myerson-regular distribution F with support [v_low, v_high] and normalized mean E[v] = 1. If the match is high, the buyer receives instantaneous utility rewards according to a Poisson process with flow rate lambda*I, where I in [0,1] is the seller-controlled access level. Upon receiving the first reward, the buyer perfectly learns both match quality theta and their own value v. The seller commits to a dynamic mechanism over time horizon T = [0, T] specifying access and prices conditional on reported histories. Both parties are risk-neutral and there is no discounting in the baseline.&lt;/p&gt;
&lt;p&gt;Two benchmark cases show the first-best is attainable absent both key features simultaneously. If trade is static (prices set only at time 0) or if the seller is uninformed about theta, the seller achieves first-best revenue of lambda&lt;em&gt;mu_0&lt;/em&gt;T by selling the entire service upfront. Proposition 1 establishes both cases; this implies that consumer data on theta is not required for maximizing social welfare, and it is weakly dominant for a seller to never collect consumer data in static environments.&lt;/p&gt;
&lt;p&gt;The central result is that the combination of dynamic pricing and seller private information breaks the first-best. A high-type seller can deviate by offering a &amp;ldquo;Myersonian free trial&amp;rdquo;: provide full access up to time tM (defined as argmax_t {(1 - exp(-lambda&lt;em&gt;t))&lt;/em&gt;(T - t)}), then offer the remaining service at post-trial price lambda&lt;em&gt;vM&lt;/em&gt;(T - tM), where vM is the Myerson monopoly price. The buyer accepts the trial regardless of beliefs (participation is weakly dominant) and purchases the post-trial service if and only if v &amp;gt;= vM. This deviation yields payoff pi_F = (1 - exp(-lambda&lt;em&gt;tM))&lt;/em&gt;(1 - F(vM))&lt;em&gt;lambda&lt;/em&gt;vM*(T - tM). Proposition 2 states that the first-best cannot be implemented in any equilibrium if and only if pi_F &amp;gt; lambda&lt;em&gt;mu_0&lt;/em&gt;T. Corollary 1 shows this condition holds for sufficiently large T, since pi_F grows proportionally with T while the first-best also grows with T but the ratio converges to a constant less than 1 only for some parameter configurations and exceeds 1 for others.&lt;/p&gt;
&lt;p&gt;Theorem 1 (the main mechanism design result) characterizes the boundary of the IC-IR feasible payoff set: any mechanism on this boundary is outcome-uniquely implemented by a trial mechanism, defined by a triple (v0, t0, p0) — a trial length, a post-trial value threshold, and a trial price. During [0, t0] uninformed buyers receive full access; after t0 only buyers who received a reward with v &amp;gt;= v0 continue at a premium. Trial length t0 is weakly increasing in the weight placed on the low-type seller and in the prior mu_0; post-trial threshold v0 is weakly decreasing in the same objects (Proposition 3).&lt;/p&gt;
&lt;p&gt;Equilibrium payoffs (Proposition 5) are precisely the IC-IR feasible pairs satisfying pi_H &amp;gt;= pi_F, implemented by pooling trial mechanisms in which both seller types propose identical mechanisms and the buyer updates beliefs only through private consumption signals. Under the D1 refinement (Proposition 6), only mechanisms with trial length tM and post-trial threshold vM survive. These have the shortest trial and highest post-trial price of all equilibrium mechanisms, minimize social surplus, and may leave both seller types strictly worse off than in a world without private information — directly contrasting the static informed principal result of Koessler and Skreta (2016) where data always helps the seller.&lt;/p&gt;
&lt;p&gt;When the seller can control service quality q in addition to access I (Section 6), the relevant equilibrium mechanisms become dynamic tiered pricing rather than binary trials: a low-quality, high-ad-load free tier provides learning opportunities while reducing information rents; convinced buyers upgrade to a premium ad-free tier. Counterintuitively, enriching the seller&amp;rsquo;s screening technology can reduce both revenue and social efficiency in equilibrium because additional instruments create additional signaling opportunities that distort outcomes further.&lt;/p&gt;
&lt;p&gt;Q: What is the core tension that prevents the first-best from being an equilibrium?&lt;/p&gt;
&lt;p&gt;A: When the seller is privately informed and pricing is dynamic, the high-type seller anticipates a greater likelihood of the buyer receiving a utility shock than the buyer&amp;rsquo;s own prior implies. This belief gap makes it profitable for the high-type seller to deviate from a proposed first-best mechanism by offering a free trial that &amp;ldquo;proves&amp;rdquo; high match quality and then extracting rent from convinced buyers. Because this deviation is profitable — yielding pi_F &amp;gt; lambda&lt;em&gt;mu_0&lt;/em&gt;T under some parameters — the first-best pooling contract unravels. The interaction of both ingredients (dynamic pricing and informed seller) is necessary: either ingredient alone is insufficient to break the first-best (Proposition 1).&lt;/p&gt;
&lt;p&gt;Q: What exactly is the Myersonian free trial and why does the buyer always accept it?&lt;/p&gt;
&lt;p&gt;A: The Myersonian free trial provides full service access up to time tM = argmax_t {(1 - exp(-lambda&lt;em&gt;t))&lt;/em&gt;(T - t)} at (approximately) zero price, then offers the remaining service at price lambda&lt;em&gt;vM&lt;/em&gt;(T - tM) where vM is the Myerson monopoly price. The buyer accepts the trial regardless of their prior belief about match quality because the trial itself is free and provides non-negative payoff. After the trial, the buyer purchases the post-trial service if and only if they received a reward with v &amp;gt;= vM; otherwise they exit. The deviation payoff is pi_F = (1 - exp(-lambda&lt;em&gt;tM))&lt;/em&gt;(1 - F(vM))&lt;em&gt;lambda&lt;/em&gt;vM*(T - tM).&lt;/p&gt;
&lt;p&gt;Q: Under what parametric conditions can the first-best not be supported in equilibrium?&lt;/p&gt;
&lt;p&gt;A: By Proposition 2, the first-best cannot be implemented if and only if pi_F &amp;gt; lambda&lt;em&gt;mu_0&lt;/em&gt;T. Corollary 1 states that for sufficiently large T this always fails, since as T grows, pi_F grows proportionally (the post-trial term (T - tM) dominates) while tM converges to a finite value. More precisely, for large T, pi_F / (lambda&lt;em&gt;mu_0&lt;/em&gt;T) converges to (1 - exp(-lambda*tM)) * (1 - F(vM)) * vM / mu_0, which exceeds 1 under appropriate parameter configurations. Conversely, when mu_0 is high or the service horizon is short, the first-best may remain implementable.&lt;/p&gt;
&lt;p&gt;Q: What is a trial mechanism and how does Theorem 1 characterize it?&lt;/p&gt;
&lt;p&gt;A: A trial mechanism is defined by a triple (v0, t0, p0): uninformed buyers receive full access on [0, t0] and no access thereafter; a buyer who reports a reward of value v &amp;gt;= v0 at time t receives full service for the remainder [t, T] at a price increment of lambda&lt;em&gt;v0&lt;/em&gt;(T - t0); the trial itself is priced at p0. Theorem 1 states that any payoff pair on the boundary of the IC-IR feasible set is outcome-uniquely attained by such a trial mechanism with appropriately determined (v0, t0, p0). The proof uses a relaxed problem retaining only two key constraint families: local incentive constraints on value reporting (IC-V) and a global intertemporal constraint preventing buyers from hiding the arrival of rewards forever (IC-U).&lt;/p&gt;
&lt;p&gt;Q: How does the trial length respond to changes in prior belief mu_0 and distributional spread?&lt;/p&gt;
&lt;p&gt;A: Proposition 3 states that t0 is weakly increasing in mu_0: as market belief becomes more optimistic, both seller types extract higher revenue from the trial, so the mechanism designer extends the trial. Proposition 4 adds that for a uniform distribution on [1-delta, 1+delta], trial length t0 is weakly increasing in delta (greater spread). The post-trial threshold v0 is weakly decreasing in mu_0, meaning that a more optimistic prior leads to a less exclusive post-trial cutoff.&lt;/p&gt;
&lt;p&gt;Q: What are the equilibrium payoffs and how does the high-type seller&amp;rsquo;s free-trial option constrain them?&lt;/p&gt;
&lt;p&gt;A: Proposition 5 states that (pi_L, pi_H) is an equilibrium payoff if and only if it lies in the IC-IR feasible set and pi_H &amp;gt;= pi_F. The lower bound pi_H &amp;gt;= pi_F reflects the high-type seller&amp;rsquo;s outside option: they can always deviate to the Myersonian free trial. Corollary 4 then shows that all &amp;ldquo;reasonable&amp;rdquo; equilibrium payoffs (those with pi_H &amp;gt;= pi_L, surviving a mild off-path refinement) are implemented by trial mechanisms with complete pooling — both seller types propose the same mechanism and the buyer updates beliefs only through private consumption signals, not the mechanism&amp;rsquo;s structure.&lt;/p&gt;
&lt;p&gt;Q: What does the D1 refinement select and why do it lead to worse outcomes?&lt;/p&gt;
&lt;p&gt;A: Proposition 6 shows that the only equilibrium trial mechanisms surviving the D1 criterion have trial length tM and post-trial threshold vM — the Myersonian free trial parameters. These have the shortest trial and highest post-trial price among all equilibrium mechanisms, resulting in the minimum social surplus. The intuition is that the high-type seller signals credibly by proposing mechanisms that generate high revenue from post-trial price discrimination (which the low type cannot profit from), pushing toward maximum learning-based discrimination. All D1-surviving payoffs are Pareto dominated by the point H (the unconstrained IC-IR optimum) for any prior mu_0, and Pareto dominated by point B when mu_0 is small.&lt;/p&gt;
&lt;p&gt;Q: Can having consumer preference data hurt the seller, and under what conditions?&lt;/p&gt;
&lt;p&gt;A: Yes. The distortion from signaling incentives can be so large that both seller types earn strictly less in the D1-surviving equilibrium than they would if neither possessed private information (where the first-best is attained). This result holds when the condition of Proposition 2 is satisfied — i.e., when pi_F &amp;gt; lambda&lt;em&gt;mu_0&lt;/em&gt;T. This contrasts sharply with the static result of Koessler and Skreta (2016), in which the ex-ante profit-maximizing mechanism is always supportable in equilibrium and data always (weakly) helps sellers.&lt;/p&gt;
&lt;p&gt;Q: How do trial mechanisms differ from the prior literature on signaling through introductory prices?&lt;/p&gt;
&lt;p&gt;A: The earlier literature (Milgrom and Roberts 1986; Bagwell 1987; Bagwell and Riordan 1991; Judd and Riordan 1994) uses two-period models with no seller commitment, so all pricing behavior is necessarily trial-like by model restriction. The present model instead allows the seller full flexibility to design any dynamic mechanism — including selling everything ex-ante, which would prevent buyers from gaining information rent. Trials emerge endogenously as the equilibrium outcome rather than being imposed by the model structure, and the paper provides new economic content on what determines trial length and price thresholds.&lt;/p&gt;
&lt;p&gt;Q: What happens when the seller controls service quality in addition to access?&lt;/p&gt;
&lt;p&gt;A: Section 6 extends the baseline by allowing the seller to choose (I, q) from a subset of [0,1]^2, where I governs the Poisson arrival rate and q scales the reward value (utility from a reward is v*q). Theorem 2 shows that the relevant equilibrium mechanisms now take the form of dynamic tiered pricing: a low-quality tier (interpreted as high ad load) provides learning opportunities while reducing information rents; once convinced, buyers upgrade to a premium high-quality tier. Enriching the screening technology in this way can reduce both revenue and social efficiency in equilibrium, because additional instruments create additional signaling opportunities that distort outcomes further from the revenue-maximizing benchmark.&lt;/p&gt;
&lt;p&gt;Q: What are the two sources of welfare loss relative to the first-best in D1-surviving equilibria?&lt;/p&gt;
&lt;p&gt;A: The welfare analysis in Appendix F identifies two sources. First, exclusion inefficiency: buyers with values v in [v_low, vM) who would generate positive surplus are excluded from post-trial service. Second, service truncation inefficiency: service access is cut off after trial length tM for buyers who were never convinced (theta = L type realizations and high-type buyers with v &amp;lt; vM), reducing total surplus below the first-best of mu_0 * lambda * T. Both losses are minimized (welfare is maximized) among trial mechanisms by longer trials and lower post-trial cutoffs, precisely the opposite of what D1 selects.&lt;/p&gt;
&lt;p&gt;Q: Does the model extend to continuous seller types or multiple buyer types?&lt;/p&gt;
&lt;p&gt;A: Appendix K outlines an extension to continuous seller types theta drawn from a distribution G on [theta_low, theta_high], where rewards arrive at rate lambda&lt;em&gt;I&lt;/em&gt;theta. The main economic forces persist: higher seller types anticipate faster buyer learning and have stronger incentives to offer trials. The main results generalize: equilibrium mechanisms are trial mechanisms, and under D1, pooling equilibria with maximum post-trial discrimination are selected. Appendix G similarly notes that the multiple-buyer-type extension preserves complete pooling and the D1 selection result.&lt;/p&gt;
&lt;p&gt;Q: What is the role of the &amp;ldquo;global intertemporal constraint&amp;rdquo; (IC-U) in the proof of Theorem 1?&lt;/p&gt;
&lt;p&gt;A: The canonical approach to dynamic mechanism design (Eso and Szentes 2007; Pavan, Segal, and Toikka 2014) relaxes the problem to only local incentive constraints on the initial report. This fails here because the informed seller causes buyer and seller to disagree on the evolution of buyer beliefs, making the timing of trade matter and requiring tracking of incentive constraints at every point in time. The paper identifies two key binding constraints in the relaxed problem: (IC-V) the buyer does not misreport their reward value, and (IC-U) the buyer does not remain silent about the arrival of a reward forever. Retaining only these two constraint families yields a tractable bang-bang solution for the optimal access policy, which is then verified to satisfy all original IC-IR constraints.&lt;/p&gt;
&lt;p&gt;Q: What are the implications for platform design and data collection strategy?&lt;/p&gt;
&lt;p&gt;A: The results imply that the value of consumer data depends critically on market dynamics. In static markets, collecting data about consumer match quality is weakly beneficial for sellers (Proposition 1, first point). In dynamic markets with buyer learning and sufficiently long service horizons, the same data can strictly reduce seller revenue by enabling a deviation that unravels first-best pricing. This suggests platforms in dynamic digital markets should weigh whether possessing and acting on proprietary match data improves or worsens their equilibrium position, and that regulatory attention to consumer data collection in dynamic markets may have welfare-ambiguous effects.&lt;/p&gt;
&lt;p&gt;Trial mechanism: A dynamic mechanism parameterized by (v0, t0, p0) in which the seller provides full service access during [0, t0] for uninformed buyers, offers continued service after t0 only to buyers who received a reward with value v &amp;gt;= v0, and charges a post-trial price of p0 + lambda&lt;em&gt;v0&lt;/em&gt;(T - t0) for those who qualify. In the paper&amp;rsquo;s usage, this is the unique outcome-implementing mechanism on the boundary of the IC-IR feasible payoff set.&lt;/p&gt;
&lt;p&gt;Myersonian free trial: The limiting trial mechanism as the trial price epsilon approaches zero, with trial length tM = argmax_t {(1 - exp(-lambda&lt;em&gt;t))&lt;/em&gt;(T - t)} and post-trial threshold vM equal to the Myerson monopoly price. It yields payoff pi_F = (1 - exp(-lambda&lt;em&gt;tM))&lt;/em&gt;(1 - F(vM))&lt;em&gt;lambda&lt;/em&gt;vM*(T - tM) to the high-type seller, and constitutes the binding outside option constraining equilibrium payoffs.&lt;/p&gt;
&lt;p&gt;Belief gap: The divergence between the seller&amp;rsquo;s and buyer&amp;rsquo;s beliefs about the rate at which the buyer will receive Poisson rewards. Because the high-type seller knows theta = H, they anticipate a higher probability of reward arrival than the buyer&amp;rsquo;s prior implies. This gap makes the buyer&amp;rsquo;s belief process non-martingale from the seller&amp;rsquo;s perspective, breaking the standard dynamic mechanism design approach and creating profitable deviation incentives.&lt;/p&gt;
&lt;p&gt;IC-IR feasible payoff set: The set of seller payoff pairs (pi_L, pi_H) achievable by mechanisms satisfying both incentive compatibility (for seller type reports and buyer learning reports) and individual rationality (non-negative ex-ante payoffs for all parties). Theorem 1 establishes that the boundary of this set is uniquely implemented by trial mechanisms.&lt;/p&gt;
&lt;p&gt;Dynamic tiered pricing: The equilibrium mechanism form that emerges when the seller controls both access I and service quality q. It features a low-quality tier (high ad load) providing learning opportunities at reduced information rent, and a premium tier offering full quality to buyers convinced of high match quality. This generalizes trial mechanisms to settings with richer screening technology.&lt;/p&gt;
&lt;p&gt;Global intertemporal constraint (IC-U): The constraint requiring that, upon receiving a Poisson reward, the buyer finds it suboptimal to remain silent about its arrival forever. Together with the local value-reporting incentive constraint (IC-V), these two constraints constitute the binding restrictions in the paper&amp;rsquo;s relaxed mechanism design problem, replacing the full continuum of incentive constraints that would otherwise be intractable.&lt;/p&gt;
&lt;p&gt;D1 criterion: A standard equilibrium refinement from signaling games applied here to the space of mechanism proposals. Among all pooling equilibrium trial mechanisms, D1 selects only those with parameters (tM, vM) — the shortest trial length and highest post-trial threshold — because the high-type seller has a strictly larger set of buyer responses for which deviation to a high-discrimination mechanism is profitable. These surviving mechanisms Pareto dominate no other equilibrium mechanism and minimize social surplus.&lt;/p&gt;</description></item><item><title>Heterogeneous innovations and growth under imperfect technology spillovers</title><link>https://macropaperwarehouse.com/papers/heterogeneous-innovations-and-growth-under-imperfect-technology-spillovers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/heterogeneous-innovations-and-growth-under-imperfect-technology-spillovers/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Jo and Kim ask two related questions: (1) How do firms use different types of innovation when learning others&amp;rsquo; technology takes time? (2) How does this process alter the aggregate implications of firm innovation, particularly in the context of increasing competition?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The paper develops a discrete-time infinite-horizon endogenous growth model with multi-product firms pursuing two types of innovation — &amp;ldquo;own-innovation&amp;rdquo; (improving existing product quality) and &amp;ldquo;creative destruction&amp;rdquo; (entering new product markets by displacing incumbents) — subject to a novel friction called &amp;ldquo;imperfect technology spillovers.&amp;rdquo; The friction takes the specific form of lagged learning: creative destruction builds on the one-period-lagged technology of the target market&amp;rsquo;s incumbent, while only the incumbent can observe the current frontier technology level. This one-period lag creates a technology gap (Δ = q_t / q_{t−1}) between the incumbent&amp;rsquo;s frontier and the level available to rivals. Four possible technology gap values arise in equilibrium: Δ₁ = 1 (no gap), Δ₂ = λ (one successful own-innovation), Δ₃ = η (one successful creative destruction), and Δ₄ = η/λ. The step sizes satisfy λ² &amp;gt; η &amp;gt; λ, meaning a single creative destruction improves quality more than a single own-innovation, but two consecutive own-innovations dominate a single creative destruction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Key Mechanisms.&lt;/strong&gt; The learning friction generates two novel mechanisms. First, the &amp;ldquo;market-protection effect&amp;rdquo;: incumbents with a technology advantage (Δ &amp;gt; 1) intensify own-innovation to widen the gap and protect their product lines when competitive pressure rises. Formally, own-innovation probability is highest for Δ₂ products and declines monotonically (z₂ &amp;gt; z₃ &amp;gt; z₄ &amp;gt; z₁), and ∂z₂/∂x &amp;gt; ∂z₃/∂x &amp;gt; 0 while ∂z₁/∂x &amp;lt; 0, conditional on value coefficients. Second, the &amp;ldquo;technological barrier effect&amp;rdquo;: higher overall own-innovation and creative destruction intensity widens the average technology gap across products, reducing rivals&amp;rsquo; conditional probability of successfully taking over a product market. This is distinct from the standard Schumpeterian effect (lower expected future profits) and from the escape-competition effect in step-by-step models (which apply only to neck-and-neck, single-product firms).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Empirical Strategy.&lt;/strong&gt; The empirical analysis combines the USPTO PatentsView database, the Longitudinal Business Database (LBD), the Longitudinal Firm Trade Transactions Database (LFTTD), the Census of Manufactures (CMF), Compustat, and NBER-CES data, covering the universe of U.S. patenting firms from 1976 to 2016, with main analyses from 1982 to 2007. Own-innovation is proxied by the self-citation ratio of patents (the ratio of self-citations to total backward citations); creative destruction by new products added and low-self-citation patents. Exogenous competitive pressure comes from China&amp;rsquo;s WTO accession in 2001, instrumented by the industry-level NTR tariff gap (the gap between non-NTR and NTR rates in 1999) following Pierce and Schott (2016).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Findings.&lt;/strong&gt; Pre-shock (1982–1999): patents with lower self-citation ratios (closer to creative destruction) have significantly longer backward citation gaps (coefficient −2.29 to −2.59, p &amp;lt; 0.01 across specifications), confirming that learning others&amp;rsquo; technology takes more time. Creative-destruction-type patents also have higher market value (Kogan et al. stock return measure) and scientific value (forward citations), with self-citation ratio negatively associated with both (e.g., coefficient on self-citation for market value: −0.289 without firm FE; −0.110 with firm FE, p &amp;lt; 0.01). Conditional on patenting, higher self-citation ratios are negatively associated with employment growth (coefficient −0.256, p &amp;lt; 0.05), number of industries added (−0.158, p &amp;lt; 0.05), and products added (−0.274, p &amp;lt; 0.01).&lt;/p&gt;
&lt;p&gt;Post-shock (DID): foreign competition had no statistically significant effect on overall patent counts, but firms with above-average innovation intensity in industries with high NTR gaps significantly increased their self-citation ratio — indicating a shift toward own-innovation. The triple-interaction coefficient is 0.795 (p &amp;lt; 0.01) with baseline controls. For a firm with average lagged innovation intensity (0.18) in an industry with an average NTR gap (0.291), this corresponds to a 4.2 percentage point increase in the seven-year growth rate of the self-citation ratio, representing a 15.0% increase relative to the average growth rate of 28.2 percentage points. Consistent with the technological barrier effect, firm entry rates are lower in industries with higher TFPR-skewness-based technological barriers (coefficient −0.012 to −0.016, p &amp;lt; 0.05).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantitative Analysis.&lt;/strong&gt; Calibrated to the U.S. manufacturing sector in 1992, the model matches six target moments including average number of products (2.3), products added (0.3), firm entry rate (7.6%), average productivity growth (1.9%), high-growth-firm employment growth (22.5%), and import penetration (15.3%). Creative destruction contributes approximately 1.88 times more to growth per unit than own-innovation (step size ratio 0.075/0.04). The aggregate R&amp;amp;D-to-sales ratio (untargeted) is 4.6% in the model vs. 4.1% in data.&lt;/p&gt;
&lt;p&gt;A counterfactual increasing outside entrants by 83% (matching the rise in import penetration from 15.3% to 25.1% between 1992 and 2007) generates a 1.51% increase in aggregate creative destruction arrival rate x, but firm-level creative destruction probability falls 1.33% and startup creative destruction also falls 1.33%. The aggregate R&amp;amp;D-to-sales ratio falls 1.6% and creative destruction R&amp;amp;D intensity falls 1.2%. Average domestic productivity growth declines 11.0%, with growth from creative destruction falling 13.0% and growth from domestic startups falling 1.7%. The total mass of domestic firms falls 6.4%.&lt;/p&gt;
&lt;p&gt;In economies with creative destruction costs 80 times higher than the U.S. baseline, the same competitive pressure shock raises rather than lowers total R&amp;amp;D (by 1.0%), but domestic growth still falls 9.7%, because the marginal decline in creative destruction impedes the growth contribution and firm entry even when aggregate innovation spending rises.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-key-friction-that-distinguishes-this-model-from-the-existing-multi-product-firm-literature-eg-klette-and-kortum-2004-akcigit-and-kerr-2018"&gt;Q1. What is the key friction that distinguishes this model from the existing multi-product firm literature (e.g., Klette and Kortum 2004; Akcigit and Kerr 2018)?&lt;/h3&gt;
&lt;p&gt;A: The key friction is &amp;ldquo;imperfect technology spillovers,&amp;rdquo; modeled as lagged learning: creative destruction can only build on the one-period-lagged technology of the target product (q_{j,t−1}), while the product&amp;rsquo;s current owner observes the frontier technology (q_{j,t}). In models without this friction — such as Akcigit and Kerr (2018) — rivals can instantly learn and copy frontier technology, so firms have no technological advantage and cannot protect their markets. In the current model, own-innovation by the incumbent widens the gap between q_{j,t} and q_{j,t−1}, creating a barrier that a rival must overcome even after successful creative destruction. This makes own-innovation an endogenous function of the technology gap, a feature absent from existing multi-product firm frameworks.&lt;/p&gt;
&lt;h3 id="q2-why-does-the-model-predict-that-own-innovation-increases-with-the-technology-gap-up-to-a-point-then-decreases"&gt;Q2. Why does the model predict that own-innovation increases with the technology gap up to a point, then decreases?&lt;/h3&gt;
&lt;p&gt;A: From Corollary 1, the ordering z₂ &amp;gt; z₃ &amp;gt; z₄ &amp;gt; z₁ reflects competing forces. Products with gap Δ₂ = λ gain the most from additional own-innovation in terms of reducing the probability of losing the product line (equation 2), so own-innovation is highest there. Products with Δ₃ = η or Δ₄ = η/λ already have substantial technological advantages from prior creative destruction, so the marginal value of own-innovation in reducing market loss probability is lower. Products with Δ₁ = 1 have no advantage at all: if a rival succeeds in creative destruction, the incumbent loses the product regardless of own-innovation (equation 1), so z₁ is lowest. Beyond a certain gap level, the incumbent is sufficiently protected that additional own-innovation has diminishing returns in deterrence.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-market-protection-effect-formally-and-for-which-products-is-it-strongest"&gt;Q3. What is the market-protection effect formally, and for which products is it strongest?&lt;/h3&gt;
&lt;p&gt;A: The market-protection effect (Corollary 2) is the positive response of a firm&amp;rsquo;s own-innovation to an increase in the aggregate creative destruction arrival rate x, conditional on the value coefficients A₁ and A₂ being fixed. It is strongest for products with Δ₂ = λ (∂z₂/∂x is the largest and positive), positive but weaker for Δ₃ = η (∂z₃/∂x &amp;gt; 0), of ambiguous sign for Δ₄ = η/λ, and negative for Δ₁ = 1 (∂z₁/∂x &amp;lt; 0). The asymmetry reflects the asymmetric payoff to own-innovation across gap levels: for Δ₂ products, successful own-innovation can turn a losing situation into a winning one because it shifts the technology gap from Δ₁ to Δ₂ from the rival&amp;rsquo;s perspective, effectively defeating the rival&amp;rsquo;s creative destruction attempt. This mechanism provides a micro-foundation for why frontier firms (like Google or NVIDIA) keep innovating intensely despite their technological leads, a pattern the standard step-by-step model cannot explain.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-technological-barrier-effect-and-how-does-it-differ-from-the-schumpeterian-effect"&gt;Q4. What is the technological barrier effect and how does it differ from the Schumpeterian effect?&lt;/h3&gt;
&lt;p&gt;A: The technological barrier effect refers to the reduction in rivals&amp;rsquo; incentive for creative destruction caused by an increase in the average technology gap across product lines. When incumbents do more own-innovation or when outside firms do more creative destruction, the distribution of technology gaps shifts rightward (density at Δ₁ falls; density at Δ₂, Δ₃, Δ₄ rises). This raises the average technology barrier rivals must overcome to successfully take over a product market, reducing the conditional takeover probability x^{takeover} and the expected value of creative destruction B. In the U.S. counterfactual, the technological barrier effect accounts for 17.0% of the total change in the aggregate creative destruction rate x and 15.0% of the change in startup creative destruction x_e. In contrast, the Schumpeterian effect refers to the reduction in expected future profits from owning a product due to increased displacement risk (through the value coefficient A₂), a mechanism present in standard quality-ladder models. Both operate simultaneously but the technological barrier effect is a novel feature of this framework.&lt;/p&gt;
&lt;h3 id="q5-how-is-own-innovation-vs-creative-destruction-measured-empirically-and-what-validates-this-measure"&gt;Q5. How is own-innovation vs. creative destruction measured empirically, and what validates this measure?&lt;/h3&gt;
&lt;p&gt;A: The self-citation ratio (the share of a patent&amp;rsquo;s backward citations that cite the same assignee&amp;rsquo;s earlier patents) is used as the primary measure: a higher ratio indicates greater reliance on the firm&amp;rsquo;s own prior knowledge, hence a higher probability that the innovation improves an existing product line (own-innovation). This is validated empirically in three ways. First, patents with lower self-citation ratios have significantly larger backward citation gaps (coefficient −2.29 to −2.59 across fixed-effect specifications on 728,721 observations), consistent with creative destruction requiring more time to learn others&amp;rsquo; technology. Second, lower self-citation patents have higher market value and scientific value (forward citations), consistent with η &amp;gt; λ (creative destruction contributes more per event to quality). Third, firm-level regressions show that lower self-citation ratios are associated with higher employment growth, more products added, and more industries entered, consistent with creative destruction contributing more to firm expansion.&lt;/p&gt;
&lt;h3 id="q6-how-does-the-did-identification-strategy-work-and-what-are-the-main-results"&gt;Q6. How does the DID identification strategy work, and what are the main results?&lt;/h3&gt;
&lt;p&gt;A: The identification exploits the removal of trade policy uncertainty (TPU) after China&amp;rsquo;s WTO accession in 2001. The treatment variable is the industry-level NTR gap (the gap between non-NTR and NTR tariff rates in 1999): industries with larger gaps experienced a larger reduction in uncertainty and thus a greater increase in Chinese import competition. The DID compares patenting firms across periods (1992–1999 vs. 2000–2007) and across high- vs. low-NTR-gap industries, with a triple interaction for firm-level innovation intensity (lagged five-year average patents per employee, normalized within two-digit NAICS). The main finding (Table 4): the NTR gap × Post interaction has no significant effect on overall patent counts (coefficient 0.238 without controls, standard error 0.237), but the triple interaction (NTR gap × Post × innovation intensity) has a positive and significant effect on the growth rate of the self-citation ratio (0.732 without controls, p &amp;lt; 0.05; 0.795 with baseline controls, p &amp;lt; 0.01). This implies that innovation-intensive firms in high-competition industries shifted their composition toward own-innovation, while overall patenting was unchanged — consistent with an offsetting rise in own-innovation and fall in creative destruction.&lt;/p&gt;
&lt;h3 id="q7-what-are-the-aggregate-growth-effects-of-increasing-competitive-pressure-in-the-calibrated-model"&gt;Q7. What are the aggregate growth effects of increasing competitive pressure in the calibrated model?&lt;/h3&gt;
&lt;p&gt;A: Using an 83% increase in outside entrants (matching the 1992–2007 rise in import penetration from 15.3% to 25.1%), average domestic productivity growth falls 11.0%. Decomposing: growth from domestic own-innovation falls 11.4%, growth from domestic creative destruction falls 13.0%, and growth from domestic startups falls 1.7% (Table 9). The aggregate R&amp;amp;D-to-sales ratio falls 1.6% and the creative destruction R&amp;amp;D intensity falls 1.2%, indicating that the decline in creative destruction R&amp;amp;D outweighs the rise in own-innovation R&amp;amp;D. The total mass of domestic firms falls 6.4% and the average number of products per firm falls 5.5%.&lt;/p&gt;
&lt;h3 id="q8-how-do-results-differ-in-economies-with-high-creative-destruction-costs-vs-the-us"&gt;Q8. How do results differ in economies with high creative destruction costs vs. the U.S.?&lt;/h3&gt;
&lt;p&gt;A: When creative destruction costs (χ̃) are set 80 times higher than the U.S. baseline, the initial equilibrium has much lower creative destruction: R&amp;amp;D-to-sales ratio is 1.39% (vs. 4.58% in U.S.), creative destruction R&amp;amp;D intensity is 8.6% (vs. 63.9%), average number of products is 1.0 (vs. 2.3), and average domestic productivity growth is 1.4% (vs. 1.9%). Under the same competition shock, total R&amp;amp;D actually rises by 1.0% in this high-CD-cost economy (because own-innovation increases more than creative destruction falls, given the already low baseline of creative destruction), in contrast to the −1.6% in the U.S. However, domestic growth still falls 9.7% even in this economy, driven by reductions in creative destruction by incumbents and startups combined with a decline in the mass of domestic incumbents. This result holds even with a fixed firm mass (Table E5), confirming the mechanism is not solely due to entry/exit dynamics.&lt;/p&gt;
&lt;h3 id="q9-what-is-the-technological-barrier-effects-quantitative-contribution-to-the-decline-in-creative-destruction"&gt;Q9. What is the technological barrier effect&amp;rsquo;s quantitative contribution to the decline in creative destruction?&lt;/h3&gt;
&lt;p&gt;A: In the U.S. counterfactual (Table 8 and associated decomposition), 17.0% of the total change in the aggregate creative destruction arrival rate x and 15.0% of the total change in startup creative destruction x_e are attributable specifically to the technological barrier effect — that is, to the shift in the technology gap distribution µ(Δℓ) holding all else equal. The conditional takeover probability x^{takeover} declines from 73.2% to 73.0%. The density at Δ₁ (the easiest gap to overcome) falls 0.4%, while densities at Δ₃ and Δ₄ rise 1.1% and 1.4% respectively, driven by increased creative destruction by outside firms and intensified own-innovation by incumbents.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-policy-implications-the-paper-draws-from-its-framework"&gt;Q10. What are the policy implications the paper draws from its framework?&lt;/h3&gt;
&lt;p&gt;A: The paper argues that policies evaluating innovation should account for composition, not just aggregate R&amp;amp;D levels or patent counts. Increased overall innovation driven by defensive own-innovation contributes less to economic growth than creative destruction and restricts firm entry — so it is less beneficial than it appears. In low-creativity economies (e.g., European economies with high regulatory barriers to creative destruction), increased foreign competition may raise aggregate R&amp;amp;D while still lowering domestic growth, misleading policymakers who track only total innovation spending. The model also suggests that the mixed empirical findings in the competition-innovation literature (Aghion et al. 2005; Bloom et al. 2016; Autor et al. 2020) can be reconciled by accounting for compositional shifts: the net effect of competition on total innovation is ambiguous because it raises own-innovation for technologically advantaged firms while reducing creative destruction for all firms.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Imperfect Technology Spillovers:&lt;/strong&gt; The novel friction introduced in this paper, modeled as lagged learning: firms attempting creative destruction can only access the one-period-lagged technology of the target product market (q_{j,t−1}), while the incumbent product owner observes and can improve from the current frontier (q_{j,t}). This asymmetry creates a persistent technological advantage for incumbents and enables strategic defensive innovation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Own-Innovation:&lt;/strong&gt; R&amp;amp;D investment by a firm to improve the quality of its existing product lines. Successful own-innovation raises product quality by a step size λ &amp;gt; 1. Own-innovation does not require learning others&amp;rsquo; technology and, in the model, constitutes the incumbents&amp;rsquo; defensive margin against creative destruction. At the aggregate level, it contributes more to total growth than creative destruction because it succeeds more frequently, but per successful event it contributes less (λ &amp;lt; η).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Creative Destruction:&lt;/strong&gt; R&amp;amp;D investment enabling a firm to enter a new product market by displacing the incumbent. Successful creative destruction improves the lagged quality of the target product by a step size η &amp;gt; λ, where λ² &amp;gt; η &amp;gt; λ. It requires learning the incumbent&amp;rsquo;s one-period-lagged technology, takes longer to develop (evidenced empirically by longer backward citation gaps), and contributes more to firm growth and product expansion per event than own-innovation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Technology Gap (Δ):&lt;/strong&gt; The ratio of a product&amp;rsquo;s current-period technology to its previous-period technology (Δ_{j,t} = q_{j,t}/q_{j,t−1}). This gap summarizes the technological advantage the incumbent holds in a product market under imperfect spillovers. Four values are possible in equilibrium: Δ₁ = 1, Δ₂ = λ, Δ₃ = η, Δ₄ = η/λ. The gap determines both the incumbent&amp;rsquo;s own-innovation incentive and the rival&amp;rsquo;s probability of successfully completing a product takeover conditional on creative destruction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Market-Protection Effect:&lt;/strong&gt; The mechanism by which incumbents with a technological advantage (Δ &amp;gt; 1) increase own-innovation in response to heightened competitive pressure (an increase in the aggregate creative destruction arrival rate x). This effect is maximized for products with Δ₂ = λ and positive but diminishing for Δ₃. It is absent for Δ₁ = 1 products (where own-innovation cannot prevent displacement) and is formally distinct from the escape-competition effect in step-by-step innovation models, which applies only to neck-and-neck single-product firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Technological Barrier Effect:&lt;/strong&gt; The reduction in rivals&amp;rsquo; incentive for creative destruction caused by an increase in the average technology gap across the economy&amp;rsquo;s product lines. When incumbents intensify own-innovation and/or when outside creative destruction increases, the distribution of technology gaps shifts toward higher Δ values, reducing the conditional probability that a rival successfully takes over any given product market. This feedback mechanism endogenously suppresses creative destruction and firm entry beyond what the Schumpeterian effect alone would predict.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Self-Citation Ratio:&lt;/strong&gt; The share of a patent&amp;rsquo;s backward citations that cite patents previously owned by the same firm. Used in the paper as a continuous proxy for the likelihood that a patent represents own-innovation vs. creative destruction: a ratio of 1 (100% self-citations) implies 100% probability of own-innovation; a ratio of 0 implies 100% probability of creative destruction. This measure follows Akcigit and Kerr (2018) and is validated in the paper against learning time, quality, and firm growth outcomes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;NTR Gap (Trade Policy Uncertainty Shock):&lt;/strong&gt; The industry-level difference between non-NTR (column 2) and NTR (column 1) U.S. tariff rates in 1999, used as an instrument for the exogenous increase in Chinese competitive pressure following China&amp;rsquo;s WTO accession and the U.S. granting of Permanent Normal Trade Relations (PNTR) in 2002. Industries with larger NTR gaps experienced a greater reduction in trade policy uncertainty and thus a larger increase in competitive pressure from foreign firms.&lt;/p&gt;</description></item><item><title>How Do You Identify a Good Manager?</title><link>https://macropaperwarehouse.com/papers/how-do-you-identify-a-good-manager/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/how-do-you-identify-a-good-manager/</guid><description>&lt;p&gt;This paper develops a novel experimental method to identify the causal contribution of managers to team performance, and uses it to evaluate which characteristics predict managerial effectiveness and how manager selection mechanisms affect organizational outcomes.&lt;/p&gt;
&lt;p&gt;The core identification challenge is that managers are not randomly assigned to teams in the field, and field managers are a highly non-random sample, making it difficult to infer which traits genuinely predict managerial performance. The authors address this by repeatedly randomly assigning managers to multiple teams in a controlled laboratory experiment, then estimating each manager&amp;rsquo;s average causal contribution to group output after conditioning on group members&amp;rsquo; individual productive skills. The intuition is that a good manager is someone who consistently causes their team to produce more than the sum of their parts.&lt;/p&gt;
&lt;p&gt;The experiment was conducted at the University of Essex lab with 555 participants (46% female, mean age 25, ethnically diverse) forming 728 groups of three across four rounds. Each group consisted of one manager and two workers who performed a Collaborative Production Task requiring coordination across three problem-solving modules (numerical, spatial, and analytical reasoning). The team score was the minimum module score — a weakest-link structure making coordination essential. Prior to group testing, all participants completed individual assessments of task-specific skill, fluid intelligence (CFIT), emotional perceptiveness (Reading the Mind in the Eyes Test, RMET), economic decision-making skill (the Assignment Game, which measures resource allocation under comparative advantage), Big 5 personality, and demographic characteristics. Manager selection was randomly varied at the session level: in 20 sessions, the participant with the strongest preference for leadership became manager (self-promotion); in 19 sessions, managers were assigned by lottery.&lt;/p&gt;
&lt;p&gt;The main quantitative findings are as follows. First, there are large, stable, and statistically significant manager effects: a manager one standard deviation above average improves team performance by approximately 0.23 standard deviations (p = 0.04). This estimate is roughly 90% the size of the combined productive skill coefficient for the two workers (approximately 0.26 sd), indicating that a good manager is roughly twice as valuable as a good individual worker. Manager contributions predict out-of-sample group performance in a leave-one-out procedure (p &amp;lt; 0.01).&lt;/p&gt;
&lt;p&gt;Second, among randomly assigned managers, only two predictors significantly explain managerial performance: fluid intelligence (CFIT) and economic decision-making skill (Assignment Game scores), both significant at below the 1% level. Gender, age, and ethnicity do not predict managerial performance.&lt;/p&gt;
&lt;p&gt;Third, self-promoted managers perform substantially worse than lottery-assigned managers, by approximately 0.10 standard deviations — roughly equivalent to being assigned a manager with fluid intelligence one full standard deviation below average. The mechanism is overconfidence: people who strongly prefer management roles are significantly more overconfident (d = 0.41 sd, p &amp;lt; 0.01) and exhibit a strong negative correlation between self-reported social skills and actual emotional perceptiveness on the RMET (r = -0.37, p &amp;lt; 0.001). Among self-promoted managers, self-reported extraversion and political skill are negatively correlated with managerial performance (rho = -0.24 and -0.26, p &amp;lt; 0.05); no such negative relationship appears among lottery managers.&lt;/p&gt;
&lt;p&gt;Fourth, selecting managers on economic decision-making skill rather than self-promotion improves average manager quality by 0.6 standard deviations — equivalent to replacing an average worker in every group with a worker at the 99th percentile of individual productivity.&lt;/p&gt;
&lt;p&gt;The three mechanisms through which good managers improve performance are: (1) monitoring — good managers (1 sd above average) cut monitoring errors from 16% to 8%; (2) optimal task allocation according to comparative advantage — groups with optimally assigned workers score 0.52 sd higher (p &amp;lt; 0.01); (3) worker motivation in late-stage effort — teams led by a 1-sd-above-average manager solve 0.6 more problems in the final two minutes versus only 0.3 more in the first two minutes.&lt;/p&gt;
&lt;p&gt;The experiment was conducted in a university lab in the UK, and the sample skews toward graduate students with limited work experience. Generalizability to field settings is supported by prior evidence that peer productivity spillover experiments yield similar magnitudes in lab versus field settings, and that the estimated manager effects are similar to Lazear et al. (2015) estimates from a large employer dataset.&lt;/p&gt;
&lt;p&gt;Q: What is the core methodological innovation of this paper?
A: The paper requires repeated random assignment of managers to multiple teams, combined with controls for individual productive skill measured prior to group work. This allows identification of each manager&amp;rsquo;s average causal contribution to group output, rather than confounding management quality with team composition or individual worker ability. The key estimand is the standard deviation of individual manager effects (sigma_alpha), interpreted as the impact of having a manager one standard deviation above average.&lt;/p&gt;
&lt;p&gt;Q: How large is the estimated manager effect, and how does it compare to worker effects?
A: A manager one standard deviation above average improves team performance by approximately 0.23 standard deviations (p = 0.04 by randomization inference). This is roughly 90% the size of the combined productive skill effect of both workers together (approximately 0.26 sd), implying a good manager is nearly twice as valuable as a good individual worker. Without conditioning on production skills, the manager effect rises to 0.29 sd.&lt;/p&gt;
&lt;p&gt;Q: What characteristics predict managerial performance among randomly assigned managers?
A: Only two measures predict managerial performance in the lottery arm: fluid intelligence (CFIT) and economic decision-making skill (scores on the Assignment Game), both significant at below the 1% level. These predictors are robust to controls for demographics, education, work experience, emotional perceptiveness, and personality traits. Gender, age, and ethnicity do not predict managerial performance.&lt;/p&gt;
&lt;p&gt;Q: What is the &amp;ldquo;Assignment Game&amp;rdquo; and why is it a strong predictor?
A: The Assignment Game (Caplin et al., 2024) places participants in a simulated managerial role where they must assign fictional workers to tasks. Performing well requires understanding comparative advantage intuitively, managing an attentionally demanding numerical environment, and avoiding biases such as anchoring. The paper argues its strong predictive power reflects that good managers excel at allocating workers according to comparative advantage — which the experiment directly identifies as a key mechanism.&lt;/p&gt;
&lt;p&gt;Q: How do self-promoted managers perform relative to lottery-assigned managers?
A: Self-promoted managers perform approximately 0.10 standard deviations below lottery managers, and this gap is robust across model specifications. The performance deficit is roughly equivalent to being assigned a manager whose fluid intelligence is one full standard deviation below average. This finding implies that common organizational practice of selecting managers partly via self-nomination actively reduces team productivity.&lt;/p&gt;
&lt;p&gt;Q: Why do self-promoted managers underperform?
A: The paper attributes underperformance primarily to overconfidence. People strongly preferring management roles are significantly more overconfident than those without strong preferences (d = 0.41 sd, p &amp;lt; 0.01). Self-promoted managers specifically overestimate their social skills: among them, self-reported people skills are strongly negatively correlated with actual emotional perceptiveness on the RMET (r = -0.37, p &amp;lt; 0.001), and self-reported extraversion and political skill are negatively correlated with managerial performance (rho = -0.24 and -0.26, p &amp;lt; 0.05). None of these negative relationships appear among lottery managers.&lt;/p&gt;
&lt;p&gt;Q: Who wants to be a manager, and does it differ by gender?
A: The three variables most strongly correlated with wanting to be in charge are extraversion, risk appetite, and being male. The relationship between high extraversion and preference for management is driven largely by men. Women are much less likely to nominate themselves for leadership roles despite being equally or more effective on average — a finding consistent with broader experimental evidence on gender and leadership self-selection.&lt;/p&gt;
&lt;p&gt;Q: How large are the potential gains from skill-based manager selection?
A: Compared to self-promotion, selecting managers based on economic decision-making skill yields managers who are 0.6 standard deviations better in terms of estimated manager effects. In terms of group performance, this is equivalent to replacing an average worker in every group with a worker at the 99th percentile of individual productivity. Selecting on both economic decision-making and fluid intelligence outperforms random assignment, selection on social skills, or selection on worker task performance (the Peter Principle).&lt;/p&gt;
&lt;p&gt;Q: What are the three mechanisms through which good managers improve team performance?
A: First, monitoring: good managers (1 sd above average) reduce monitoring errors — defined as having a worker on a module substantially above the minimum score at task end — from 16% to 8% (bivariate correlation with manager performance = -0.40, p &amp;lt; 0.001). Second, optimal task allocation: the probability of finding the optimal comparative-advantage-based assignment is positively associated with manager performance (rho = 0.19, p &amp;lt; 0.01), and groups with always-optimal starting assignments score 0.52 sd higher than those with never-optimal assignments (p &amp;lt; 0.01). Third, worker motivation: team performance in the final two-minute period is about 50% more influential for overall outcomes than the first two minutes (p = 0.038), and 1-sd-above-average managers generate 0.6 more problems solved in the final period versus 0.3 in the first, consistent with differential motivational effects emerging over time.&lt;/p&gt;
&lt;p&gt;Q: What is the Peter Principle, and how does this paper relate to it?
A: The Peter Principle refers to the practice of promoting employees based on their performance as line workers rather than their suitability for management — promoting individuals to their level of incompetence. Benson et al. (2019) document this selection pattern empirically. This paper shows that selecting managers on worker task skill is inferior to selecting on economic decision-making skill or fluid intelligence, confirming that task skill is not the right criterion for manager selection even if it predicts individual worker output.&lt;/p&gt;
&lt;p&gt;Q: How does the paper validate that manager effects are real and not noise?
A: The paper uses randomization inference with 5,000 simulated allocations to compute p-values, obtaining p = 0.04 for the main manager effect. Robustness checks include controlling for pre-existing social relationships, manager risk appetite, variance of individual scores, and granular skill measures — all yielding estimates near 0.22 sd. A leave-one-out out-of-sample prediction test confirms manager contributions significantly predict held-out group performance (p &amp;lt; 0.01), while the analogous worker out-of-sample estimate is less than half the magnitude and not statistically significant.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions on the experimental results?
A: The experiment is conducted in a university lab in the UK with graduate students averaging 25 years of age and two years of work experience, limiting direct generalizability to experienced workers or senior management. The task lasts approximately 15 minutes, which may not capture longer-run managerial dynamics. Compensation equalized average earnings between managers and workers, which differs from most real-world settings. The authors note their effect-size estimates closely match Lazear et al. (2015) from a large employer, and that Herbst and Mas (2015) find lab peer-productivity experiments generalize to the field.&lt;/p&gt;
&lt;p&gt;Manager Effect (sigma_alpha): The standard deviation of individual managers&amp;rsquo; average causal contributions to group performance, estimated via repeated random assignment and conditioning on individual productive skill. Represents the impact of having a manager one standard deviation above average, estimated at approximately 0.23 standard deviations of group output.&lt;/p&gt;
&lt;p&gt;Collaborative Production Task: A novel lab group task in which a manager and two workers solve problems across three modules (numerical, spatial, analytical reasoning), with team score defined as the minimum module score (weakest-link structure). Managers are responsible for worker assignment, monitoring, and motivation; workers face no financial performance incentives.&lt;/p&gt;
&lt;p&gt;Economic Decision-Making Skill: Defined by Caplin et al. (2024) as the ability to make good resource allocation decisions, assessed via the Assignment Game in which participants must optimally assign workers to tasks under comparative advantage. The single strongest predictor of managerial performance in the lottery arm.&lt;/p&gt;
&lt;p&gt;Monitoring Failure: Defined in the paper as having any group member working on a module at task end whose score is substantially greater (e.g., 10 points higher) than the minimum module score — meaning the worker&amp;rsquo;s effort is not contributing to the group score. Occurs in 16% of groups overall; managers one sd above average reduce this to 8%.&lt;/p&gt;
&lt;p&gt;Self-Promotion (as selection mechanism): A treatment condition in which the participant with the strongest stated preference for being manager (on a 1-10 scale) is assigned the managerial role. Contrasted with lottery assignment; self-promoted managers perform approximately 0.10 sd worse than lottery managers.&lt;/p&gt;
&lt;p&gt;Overconfidence (in managerial context): The gap between self-assessed skill (particularly social/interpersonal skill) and objectively measured skill (e.g., RMET score). Self-promoters are significantly more overconfident (d = 0.41 sd), and overconfidence is strongly negatively correlated with actual emotional perceptiveness (r = -0.33, p &amp;lt; 0.001).&lt;/p&gt;
&lt;p&gt;Comparative Advantage Allocation: The practice of assigning each worker to the module in which they have the highest relative (not absolute) performance advantage. Captured via whether a manager selects the optimal one-to-one assignment given pre-measured individual module scores; groups with always-optimal allocation score 0.52 sd higher.&lt;/p&gt;</description></item><item><title>Input Sourcing under Climate Risk: Evidence from U.S. Manufacturing Firms</title><link>https://macropaperwarehouse.com/papers/input-sourcing-under-climate-risk-evidence-from-u.s.-manufacturing-firms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/input-sourcing-under-climate-risk-evidence-from-u.s.-manufacturing-firms/</guid><description>&lt;p&gt;Blaum, Esposito, and Heise study how supply chain risk — specifically, the risk of unexpected shipping delays caused by ocean weather conditions — affects U.S. manufacturing firms&amp;rsquo; import sourcing decisions. The paper asks three related questions: Do weather-induced shipping delays harm firm performance? Do firms adapt their sourcing strategies ex ante in response to shipping time risk? And what are the aggregate welfare costs of heightened supply chain risk from climate change, geopolitical tensions, and port congestion?&lt;/p&gt;
&lt;p&gt;The empirical foundation is the U.S. Census Bureau&amp;rsquo;s Longitudinal Firm Trade Transactions Database (LFTTD), covering the universe of U.S. import transactions from 1992 to 2016, merged with the Longitudinal Business Database and Annual Survey of Manufacturers for firm-level outcomes. For ocean shipments, the authors reconstruct vessel routes using vessel names, foreign port stops, and U.S. ports of entry, then map those routes to hourly wave height and direction data from NOAA&amp;rsquo;s WaveWatch III model at 0.5-degree resolution across more than 40,000 distinct maritime routes (period: 2011–2016 for weather data).&lt;/p&gt;
&lt;p&gt;The identification strategy proceeds in two steps. First, observed shipping times are regressed on a rich set of fixed effects — supplier, product, route-month, vessel, buyer, relationship status — plus controls for shipping charges and weight, to strip out anticipated determinants of delivery time. Second, the residuals are projected onto realized wave height and direction along the vessel&amp;rsquo;s route to isolate the weather-induced, unexpected component of shipping time variation. The identifying assumption is that realized wave conditions along the entire multi-week ocean crossing are not predictable by importers at the time orders are placed, beyond seasonal patterns absorbed by route-month fixed effects. This assumption is supported by the literature on weather forecasting, which finds accuracy degrades sharply beyond seven days.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s first empirical result concerns the consequences of weather-induced delays. Defining an extreme delay as a weather-induced shipping time above the 95th percentile for a given product-route, the authors estimate that a one standard deviation increase in the share of input costs that are weather-delayed (2.66 percentage points) reduces firm sales by 6.5%, profits by 3.5%, and employment by 1.0% within the same year. These effects are estimated from panel regressions for 2011–2016, with importer, product, and year fixed effects. The magnitudes indicate that firms are typically unable to fully hedge supply chain disruptions through insurance or financial instruments.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s second empirical result concerns ex ante adaptation. Risk exposure is measured as the standard deviation of weather-induced shipping times over three-year rolling windows for each supplier-route-product combination, then aggregated to the importer-product-year level using pre-determined import shares as weights (Bartik shift-share). Moving from the 25th to the 75th percentile of this shipping risk distribution increases the number of routes used by 7.7% and the number of foreign suppliers by 4.9%, while reducing total import value by 5.1%, route concentration (HHI) by 4.6%, and supplier concentration (HHI) by 3.2%. The risk effect on imports is estimated conditional on average shipping time, indicating that uncertainty exerts an additional, independent negative effect on import demand beyond the level of delays.&lt;/p&gt;
&lt;p&gt;To rationalize these findings, the authors build a quantitative general equilibrium model of importing with firm heterogeneity. Firms source domestic and foreign inputs; foreign input quality is reduced when delivery is late, and firms face uncertainty about shipping times when placing orders. Risk-neutral firms nonetheless face a concavity in expected revenues from monopolistic competition, so higher variance in input quality reduces expected profits. Firms can diversify by adding foreign suppliers (at a per-supplier fixed cost), and a key theoretical result is that a mean-preserving spread in supplier quality variance increases the optimal number of suppliers but, because the extensive-margin elasticity is less than one, total import value necessarily falls.&lt;/p&gt;
&lt;p&gt;The calibrated model is used to evaluate three counterfactual scenarios. Ocean wave height volatility increased by 0.34% per year on average between 2011 and 2023; projecting this trend forward 50 years generates a climate change scenario. The Houthi attacks in the Red Sea caused rerouting that raised both the mean and variance of navigation time. Post-Covid port congestion (2021–2022) increased the variance of port waiting times. Across all three scenarios, U.S. real income falls by 0.4% to 1.33%, driven by firms substituting toward more expensive domestic inputs as they reduce exposure to risky foreign sourcing.&lt;/p&gt;
&lt;p&gt;The sample scope is U.S. manufacturing importers using ocean shipping during 2011–2016 for the main empirical results (weather data period), with an extended robustness sample of 1992–2016 using residualized shipping time volatility. The study covers 43,080 origin-destination port pairs, 401,700 unique vessels, and approximately 35.8 million seaborne transactions.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s core research question?
A: The paper asks how supply chain risk — specifically, the risk of unexpected delays in ocean shipping caused by weather conditions — affects U.S. manufacturing firms&amp;rsquo; import sourcing decisions and aggregate welfare. It examines both the disruption effects of realized delays and the ex ante adaptation of sourcing strategies to risk exposure, then quantifies aggregate costs through a calibrated general equilibrium model.&lt;/p&gt;
&lt;p&gt;Q: What data sources underpin the empirical analysis?
A: The primary dataset is the LFTTD, which covers the universe of U.S. import transactions from 1992 to 2016, recording importer and exporter identities, HS-10 product codes, values, quantities, shipping dates, vessel names, and port pairs. This is merged with the Longitudinal Business Database for employment and industry, and with Census of Manufactures and Annual Survey of Manufacturers for sales, material costs, and payroll. Weather data come from NOAA&amp;rsquo;s WaveWatch III model at hourly, 0.5-degree resolution for 2011–2016. Ocean routes are constructed using Eurostat&amp;rsquo;s SeaRoute program, covering over 40,000 distinct routes across approximately 10,500 route segments.&lt;/p&gt;
&lt;p&gt;Q: How do the authors isolate the unexpected component of shipping time variation?
A: They use a two-step residualization. In step one, observed log shipping times are regressed on supplier, product, route-month, vessel, buyer, and relationship-status fixed effects, plus controls for log shipping charges and log weight; the residuals capture variation not explained by anticipated factors. In step two, these residuals are projected onto realized average wave height and relative wave direction along the vessel&amp;rsquo;s route to extract the weather-induced component. The identifying assumption is that importers cannot forecast realized wave conditions beyond seasonal patterns when placing orders that initiate multi-week ocean crossings, consistent with evidence that weather forecasts lose accuracy beyond seven days and that ocean wave height is particularly hard to predict.&lt;/p&gt;
&lt;p&gt;Q: What are the estimated effects of weather-induced shipping delays on firm performance?
A: A one standard deviation increase in the share of input costs that are weather-delayed (2.66 percentage points) reduces firm sales by 6.5%, profits by 3.5%, and employment by 1.0% within the same year. Using a broader measure of residualized shipping time delays (not restricted to the weather-induced component) produces similar results: a one standard deviation increase reduces sales by 6%, profits by 3.2%, and employment by 0.9%. These effects are estimated from panel regressions for 2011–2016 with importer, product, and year fixed effects.&lt;/p&gt;
&lt;p&gt;Q: How do firms adjust their sourcing strategies in response to higher shipping time risk?
A: Moving from the 25th to the 75th percentile of the shipping risk distribution (a 61 log-point increase) raises the number of routes used by 7.7% and the number of foreign suppliers by 4.9%, while reducing route HHI by 4.6%, supplier HHI by 3.2%, and total import value by 5.1%. The margin of route diversification is larger than supplier diversification, consistent with shipping risk being determined primarily at the route level. Higher risk also increases the likelihood of switching to air freight by 1.0% over the same interquartile range.&lt;/p&gt;
&lt;p&gt;Q: Does the risk effect on imports operate independently of the level of shipping times?
A: Yes. The regressions of total import demand on risk exposure control for average shipping time, and the coefficient on risk remains negative and significant after this control. This indicates that the variance of shipping times has an independent negative effect on import demand beyond the first-moment effect of longer average delays.&lt;/p&gt;
&lt;p&gt;Q: What is the theoretical mechanism through which shipping time risk reduces import demand?
A: In the model, firms are risk-neutral but face monopolistically competitive output markets, which introduces curvature in the revenue function. Higher variance in input quality (stemming from unpredictable shipping times) reduces expected revenues even for risk-neutral firms. Firms can diversify by adding foreign suppliers at a per-supplier fixed cost, which reduces variance in average input quality. However, the elasticity of the optimal number of suppliers with respect to quality variance is less than one, so total import expenditure necessarily falls as variance rises — diversification is incomplete and firms substitute toward domestic inputs.&lt;/p&gt;
&lt;p&gt;Q: What does Proposition 1 state about the extensive margin response to risk?
A: Proposition 1 establishes that, under the condition that shipping time risk is small relative to expected revenues, a mean-preserving spread in the variance of supplier quality increases the optimal number of foreign suppliers. However, the elasticity of the optimal number of suppliers with respect to quality variance is strictly less than one, which implies that total import value necessarily falls whenever quality variance increases, regardless of the extensive margin diversification response.&lt;/p&gt;
&lt;p&gt;Q: How is the calibration structured and what moments does it target?
A: The model features firm heterogeneity in both productivity and shipping time risk (variance of delivery times). The calibration targets three sets of moments: the estimated effect of shipping time risk on the extensive margin of importing (number of suppliers), the negative association between firm sales and average shipping times (which disciplines the timeliness elasticity parameter tau), and the joint distribution of firm size and risk observed in the data — specifically, the empirical finding that larger importers are matched with safer (lower-risk) foreign suppliers, with a correlation of -0.12. The calibrated model replicates the key moments of shipping time risk and import demand.&lt;/p&gt;
&lt;p&gt;Q: What are the three counterfactual scenarios and their aggregate welfare costs?
A: (1) Climate change: ocean wave height volatility increased by 0.34% per year on average between 2011 and 2023; projecting this trend forward 50 years and passing the resulting increase in shipping time variance through the model. (2) Red Sea/Houthi attacks: re-routing around the Suez Canal raises both the mean and variance of navigation time. (3) Post-Covid port congestion: greater variability in port waiting times during 2021–2022. Across all three scenarios, U.S. real income falls by 0.4% to 1.33%, driven by firms substituting from cheaper foreign inputs toward more expensive domestic production to reduce risk exposure.&lt;/p&gt;
&lt;p&gt;Q: What is the role of the shift-share (Bartik) instrument in the risk exposure measure?
A: The exposure measure aggregates supplier-route-product level risk (standard deviation of weather-induced shipping times over three-year rolling windows) to the importer-product-year level using pre-determined import shares from the prior three years as weights. Using lagged shares rather than contemporaneous shares ensures that the weights are not endogenous to current sourcing decisions. This construction is standard in the Bartik shift-share literature and helps isolate variation in risk that is plausibly exogenous to the firm&amp;rsquo;s current sourcing choices.&lt;/p&gt;
&lt;p&gt;Q: How do the authors handle the endogeneity concern that firms may select into riskier routes?
A: The weather-induced component of shipping time variation is by construction driven by realized ocean conditions that are unpredictable at the time orders are placed. The residualization removes all fixed-effect variation associated with route, season, vessel, supplier, and buyer characteristics. Additionally, the shift-share construction uses pre-determined weights, so risk exposure does not mechanically reflect current sourcing decisions. The authors also show robustness using the longer 1992–2016 sample with residualized (rather than weather-specific) shipping time volatility, obtaining qualitatively and quantitatively similar results.&lt;/p&gt;
&lt;p&gt;Q: What does the paper contribute relative to the literature on shipping times and trade?
A: Prior work by Evans and Harrigan (2005) and Hummels and Schaur (2010, 2013) focused on the level of shipping times (the first moment) as a trade cost. This paper is the first to systematically study the variance of shipping times (the second moment) as an independent determinant of import demand and sourcing structure, both empirically and theoretically. The authors show that uncertainty around delivery times has negative effects on trade that are separate from the effects of longer average delays.&lt;/p&gt;
&lt;p&gt;Q: What are the robustness checks reported for the main empirical results?
A: For the effects of risk on sourcing behavior, the authors show that using residualized shipping time volatility over the longer 1992–2016 sample (rather than the weather-induced measure over 2011–2016) produces similar results: moving from the 25th to the 75th percentile increases routes by 6.6%, suppliers by 3.7%, decreases route HHI by 3.9%, and supplier HHI by 2.5%, while reducing total imports by 10.5%. For the effects of delays on firm performance, applying the same specification with residualized (not weather-induced) delay shares yields coefficients on sales, profits, and employment that are very close to the baseline estimates.&lt;/p&gt;
&lt;p&gt;Q: What are the welfare implications for firms that cannot hedge through financial markets?
A: The large negative effects of weather-induced delays on sales, profits, and employment — and the finding that firms respond by ex ante restructuring their supply chains rather than relying on insurance — indicate that financial hedging instruments are largely unavailable or insufficient for managing input delivery risk. This motivates the model&amp;rsquo;s assumption that firms must manage risk through sourcing diversification, which is costly because of per-supplier fixed costs and because it ultimately requires substituting toward more expensive domestic inputs.&lt;/p&gt;
&lt;p&gt;Weather-induced unexpected shipping time: The component of shipping time variation explained by realized ocean wave height and direction along the vessel&amp;rsquo;s route, after removing all variation attributable to anticipated factors (route, season, vessel, supplier, buyer characteristics, shipping charges, weight). Interpreted as unexpected because multi-week ocean crossings begin before accurate weather forecasts are available.&lt;/p&gt;
&lt;p&gt;Shipping time risk: Measured as the standard deviation of weather-induced residualized shipping times over three-year rolling windows for each foreign supplier-route-product combination. This captures the second moment (variance) of delivery time uncertainty, distinct from the first moment (average shipping time level).&lt;/p&gt;
&lt;p&gt;Shift-share risk exposure: An importer-product-year level risk measure constructed as a weighted average of supplier-route-product level risk, using pre-determined import shares from the prior three years as weights. This Bartik-style construction ensures exposure weights are not endogenous to current sourcing decisions.&lt;/p&gt;
&lt;p&gt;Timeliness elasticity (tau): A structural parameter in the model governing how rapidly input quality degrades when delivery is later than expected. Specifically, when a shipment arrives di days late, quality is reduced by the factor exp(-tau*(di - E[di])). Calibrated to match the observed negative association between firm sales and average shipping times in the data.&lt;/p&gt;
&lt;p&gt;Extensive margin diversification: The response of firms to higher shipping time risk by increasing the number of foreign suppliers and shipping routes used for a given product, rather than increasing the volume sourced from existing suppliers. In the model and data, this margin is the primary channel through which firms hedge delivery risk.&lt;/p&gt;
&lt;p&gt;Mean-preserving spread condition: The theoretical condition (Proposition 1) under which higher variance in supplier quality increases the optimal number of foreign suppliers. The condition requires that shipping time risk be small relative to expected revenues, so that the diversification benefit of adding suppliers (reducing variance in average quality) dominates the revenue-reducing effect of higher variance.&lt;/p&gt;
&lt;p&gt;Per-supplier fixed cost: A fixed cost in the model that must be paid for each foreign supplier relationship maintained. This cost limits the extent of diversification, ensuring that firms cannot fully eliminate shipping time risk by adding arbitrarily many suppliers, and that higher risk raises (rather than eliminates) per-unit sourcing costs.&lt;/p&gt;</description></item><item><title>Insurer Risk and Public Risk-Sharing: Quantifying the Value of Reinsurance</title><link>https://macropaperwarehouse.com/papers/insurer-risk-and-public-risk-sharing-quantifying-the-value-of-reinsurance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/insurer-risk-and-public-risk-sharing-quantifying-the-value-of-reinsurance/</guid><description>&lt;p&gt;Kim and Li study how publicly provided reinsurance affects insurer behavior and market outcomes in health insurance markets where firms face substantial cost uncertainty. The central question is whether standard expected-profit models—which predict that reinsurance reducing only cost volatility (not expected cost) should leave prices unchanged—miss an important mechanism: insurers internalizing the implicit financial cost of bearing claims uncertainty through &amp;ldquo;risk charges.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The paper develops a stylized monopoly-insurer model in which the insurer&amp;rsquo;s objective includes both expected claims cost and a risk charge term L(S), where S is a risk measure (e.g., standard deviation of total claims). This yields a first-order condition in which effective marginal cost includes both standard expected claims cost and a marginal risk charge. The model predicts that public reinsurance acts through two distinct channels: (1) a cost subsidy—reimbursing a share of high-cost claims reduces expected cost; and (2) risk protection—reducing the variance of claims lowers the risk charge and thus effective marginal cost. When both channels operate, the model predicts pass-through of public reinsurance to premiums can exceed unity, in contrast to the standard less-than-one pass-through under market power.&lt;/p&gt;
&lt;p&gt;Empirically, the authors use three primary data sources for the U.S. individual health insurance exchange market. NAIC Schedule S filings (2014–2023) provide transaction-level private reinsurance contracts, including ceded premiums, realized claims, and financial solvency measures. CMS Public Use Files and MLR reports provide plan-level premiums, enrollment, and claims. The Colorado All Payer Claims Database (CO APCD, 2014–2022) and Connect for Health Colorado administrative records (2015–2021) provide individual-level claims and insurance choices for structural analysis.&lt;/p&gt;
&lt;p&gt;Descriptive evidence establishes that 62% of exchange insurers purchase private reinsurance despite average reinsurance markups of 1.54 (reinsurance margin of 0.54), and that smaller, less financially solvent insurers are disproportionate buyers—consistent with risk charges driving demand for risk protection even at above-actuarially-fair prices.&lt;/p&gt;
&lt;p&gt;An event study exploiting staggered adoption of state-level public reinsurance programs finds that public reinsurance reduces premiums by approximately 14.5% on average (27% in Colorado Tiers 1–2, 46% in Tier 3), with a pass-through rate of 1.3—significantly greater than one (p = 0.037 one-sided). Public reinsurance reduces the probability of purchasing private reinsurance by 26 percentage points (a 42% reduction from baseline) and per-member private reinsurance expenditures by $19.5 (a 68% reduction from baseline). Premium and private reinsurance effects are larger for financially constrained insurers (RBC ratio below 3). No significant effects are found on insurer entry/exit, total medical expenses (ruling out moral hazard), or private reinsurance markups.&lt;/p&gt;
&lt;p&gt;The structural model, estimated on the Colorado exchange for 2017–2020, finds that the risk charge coefficient for regional insurers averages rho = 0.25, implying regional insurers face 9.8% higher effective costs than national insurers due to risk charges and private reinsurance expenses. Risk charges account for at least half the premium-cost wedge for small regional insurers. Counterfactual decomposition of Colorado&amp;rsquo;s program shows the direct cost subsidy accounts for approximately 75% of equilibrium price reductions; risk protection and competition effects together account for the remaining 25%. In a bang-for-buck comparison, public reinsurance dominates premium subsidies of equal government expenditure by approximately 20–30%, because reinsurance uniquely reduces risk charges and enhances competition by reducing smaller regional insurers&amp;rsquo; cost disadvantage.&lt;/p&gt;
&lt;p&gt;Q: What is the core theoretical innovation of the paper?
A: The paper adds a risk charge term L(S) to the standard expected-profit objective, where S is a risk measure of the insurer&amp;rsquo;s cost distribution. This makes the insurer behave &amp;ldquo;as if risk averse,&amp;rdquo; with effective marginal cost including both expected claims cost and a marginal risk charge that decreases with insured pool size due to risk pooling. When rho = 0, the model collapses to the standard monopoly case; when rho &amp;gt; 0, cost uncertainty directly inflates prices and creates a novel role for reinsurance even when reinsurance is actuarially fair priced.&lt;/p&gt;
&lt;p&gt;Q: What are the two distinct mechanisms through which public reinsurance affects insurer pricing?
A: The first is a cost subsidy: by reimbursing a portion of high-cost claims without requiring an actuarially fair premium upfront, public reinsurance lowers the insurer&amp;rsquo;s net expected cost. The second is risk protection: by providing ex-post payments for extreme health shocks, reinsurance reduces the variance of claims costs, lowering the risk charge component of effective marginal cost. Together, these channels can produce pass-through exceeding unity even under imperfect competition, where standard cost-subsidy pass-through is typically below one.&lt;/p&gt;
&lt;p&gt;Q: What does Proposition 1 say about actuarially fair reinsurance (theta = 1)?
A: Proposition 1(i) states that actuarially fair reinsurance—which does not alter net expected cost—still lowers the insurer&amp;rsquo;s price if and only if the insurer faces a risk charge (rho &amp;gt; 0). An insurer without risk charges is entirely unaffected by actuarially fair reinsurance. This result isolates the risk-protection channel as theoretically distinct from cost subsidization and establishes that pass-through exceeding one requires risk charges to be operative.&lt;/p&gt;
&lt;p&gt;Q: Why would an insurer purchase costly private reinsurance (theta &amp;gt; 1)?
A: Proposition 1(iii) shows that an insurer with no risk charge would never purchase private reinsurance with theta &amp;gt; 1, since it increases net expected cost with no offsetting benefit. An insurer facing a risk charge (rho &amp;gt; 0) may purchase private reinsurance because the risk-protection benefit—the reduction in cost variance and thus the risk charge—can outweigh the net cost increase. The paper documents that 62% of exchange insurers buy private reinsurance at an average markup of 1.54 (reinsurance margin 0.54), with smaller and financially weaker insurers more likely to purchase, consistent with this mechanism.&lt;/p&gt;
&lt;p&gt;Q: How does the paper establish empirically that insurers face and internalize cost uncertainty?
A: Three lines of evidence are presented. First, the CO APCD shows the claims distribution has a long right tail: the top 5% (1%) of consumers account for 68% (38%) of total expenses, and 2.5% of consumers exceed the $30,000 reinsurance threshold. Second, simulations show that with 1,000 enrollees, the probability that realized claims exceed expected costs by 25% is approximately 7%; even at 10,000 enrollees there is a 17% probability of exceeding expected costs by 5%. Third, in over 24% of insurer-year observations premium revenue falls short of realized claims costs, and the within-firm standard deviation of the claims-to-premium ratio is 0.15.&lt;/p&gt;
&lt;p&gt;Q: What are the event study findings on premiums?
A: Using staggered introduction of state-level public reinsurance programs, the event study finds premiums fell by 14.5% on average following program adoption. In Colorado specifically, Tiers 1 and 2 experienced 27% decreases and Tier 3 (highest reinsurance generosity) experienced a 46% decrease. The implied pass-through rate for 2020 is 1.3, meaning for every dollar the government spent on reinsurance, health insurance premiums fell by $1.30. A one-sided t-test rejects pass-through equal to one at p = 0.037.&lt;/p&gt;
&lt;p&gt;Q: What are the event study findings on private reinsurance?
A: Public reinsurance reduces the probability that an insurer purchases private reinsurance by 26 percentage points, a 42% decline from the pre-program baseline. Average per-member private reinsurance expenditures fall by $19.5, a 68% reduction from baseline. The substitution away from private reinsurance is consistent with the model prediction that public reinsurance displaces the demand for risk protection previously met by private markets, and reinforces the interpretation that risk management is a key driver of private reinsurance demand.&lt;/p&gt;
&lt;p&gt;Q: Do financially constrained insurers respond differently to public reinsurance?
A: Yes. The premium-reduction effect is significantly larger for insurers with RBC ratios below 3 (an additional interaction effect of -0.161 log points on top of the baseline -0.135). The reduction in per-member private reinsurance expenditures is also significantly larger for insurers with significant prior private reinsurance purchases (-$108.8 vs. baseline of -$19.5). This heterogeneity supports the hypothesis that the risk protection channel is more valuable for financially constrained insurers who face higher implicit costs of bearing risk.&lt;/p&gt;
&lt;p&gt;Q: Does public reinsurance affect insurer entry/exit, moral hazard, or private reinsurance markups?
A: The event study finds no statistically significant effect on market entry, total monthly medical expenses per enrollee, the probability that individual expenses exceed the reinsurance threshold (ruling out insurer moral hazard), or private reinsurance markups paid by primary insurers. These null results support the interpretation that premium reductions reflect reduced cost uncertainty rather than cost containment distortions, and that the competitive structure of the private reinsurance market is not directly altered by public programs.&lt;/p&gt;
&lt;p&gt;Q: What are the structural estimates of risk charges?
A: The estimated risk charge coefficient for regional insurers averages rho = 0.25. This implies that regional insurers incur, on average, 9.8% higher effective costs than national insurers (who are assumed not to face risk charges due to scale and diversification), stemming from both direct risk charges and private reinsurance expenses required to manage risk. Risk charges account for at least half the observed wedge between premiums and marginal claims costs for small regional insurers.&lt;/p&gt;
&lt;p&gt;Q: How does the structural model decompose the impact of Colorado&amp;rsquo;s reinsurance program?
A: Counterfactual analysis decomposes the equilibrium price reduction into three channels. The direct cost subsidy effect—reimbursing a share of high-cost claims between the $30,000 attachment point and $400,000 cap—accounts for approximately 75% of the price reduction. The risk protection effect (reduction in risk charges from lower portfolio variance) and the competition effect (smaller regional insurers facing lower cost disadvantages and competing more aggressively with national insurers) together account for the remaining 25% of the equilibrium price reduction.&lt;/p&gt;
&lt;p&gt;Q: How does public reinsurance compare to premium subsidies in bang-for-buck terms?
A: For equal government expenditure, public reinsurance is estimated to be approximately 20–30% more cost-effective than premium subsidies at reducing premiums. The advantage stems from two sources: reinsurance reduces risk charges, shifting down the marginal cost curve for regional insurers in a way demand-side premium subsidies do not; and reinsurance enhances competition by reducing the cost disadvantage of smaller regional insurers relative to national ones. The dominant effect is risk reduction rather than markup inflation, making reinsurance the more efficient instrument when the degree of financial risk is considerable.&lt;/p&gt;
&lt;p&gt;Q: What is the role of market size in risk charges, and why does this create a competitive asymmetry?
A: The model shows that the marginal risk charge decreases as the insured population grows (risk pooling), with marginal standard deviation equal to sigma_0 / (2*sqrt(q)), which vanishes as q approaches infinity. This implies that larger national insurers, covering very large populations, effectively face no risk charges, while smaller regional insurers face meaningful marginal risk charges. This size-asymmetry is the fundamental reason why public reinsurance disproportionately benefits smaller insurers—by reducing their risk charges, it narrows the cost gap with national insurers and intensifies competition.&lt;/p&gt;
&lt;p&gt;Q: What scope conditions apply to the structural findings?
A: The structural estimates are based on the Colorado individual health insurance exchange, covering years 2017–2020, chosen to avoid unsatisfactory early data quality and to net out systematic pandemic effects. The model assumes national insurers do not face risk charges in the baseline specification, and that aggregate (correlated) risk is not the primary driver during the sample period. Results are robust to staggered-treatment corrections (Callaway-Sant&amp;rsquo;Anna 2021; Borusyak et al. 2024), alternative outcome measures (benchmark premiums, Silver plan averages), alternative aggregation levels, and sensitivity analyses allowing for insurer entry/exit, correlated risks, moral hazard, and alternative risk charge functional forms.&lt;/p&gt;
&lt;p&gt;Q: What are the broader policy implications of the framework?
A: The framework applies to any market where firms face substantial cost uncertainty and internalize financial risk, including property and casualty insurance, flood insurance, wildfire insurance, and government loan guarantee programs. The analysis suggests that ignoring the risk protection channel causes policymakers to underestimate the effectiveness of public reinsurance relative to demand-side subsidies. Supply-side risk-sharing policies are particularly important for markets with small, financially constrained firms, where cost uncertainty most severely distorts pricing and competition, and where the competitive benefits of risk reduction are largest.&lt;/p&gt;
&lt;p&gt;Risk Charge: An additional cost term in the insurer&amp;rsquo;s objective function representing the implicit financial cost of bearing claims uncertainty, formalized as L(S) where S is a risk measure of total cost. Risk charges make the insurer behave &amp;ldquo;as if risk averse,&amp;rdquo; raising effective marginal cost above expected claims cost. In the baseline model the risk charge equals rho times the standard deviation of total claims.&lt;/p&gt;
&lt;p&gt;Risk Charge Coefficient (rho): The parameter governing the insurer&amp;rsquo;s marginal cost of financial risk, estimated structurally at an average of 0.25 for regional insurers in Colorado. It can be interpreted as either a direct risk-aversion parameter, the marginal cost of regulatory capital, or a reduced-form representation of financial and regulatory frictions that make bearing cost uncertainty costly.&lt;/p&gt;
&lt;p&gt;Risk Protection Channel: The mechanism through which reinsurance (public or private) reduces claims cost variance and thereby lowers the insurer&amp;rsquo;s risk charge, distinct from the cost-subsidy channel. The risk protection channel is operative even for actuarially fair reinsurance (theta = 1) and is responsible for pass-through rates exceeding unity under public reinsurance programs.&lt;/p&gt;
&lt;p&gt;Cost Subsidy Channel: The mechanism through which subsidized public reinsurance (theta less than 1) lowers the insurer&amp;rsquo;s net expected claims cost by reimbursing a share of high-cost claims without charging an actuarially fair premium. This channel operates regardless of whether the insurer faces risk charges and is the primary channel in standard models.&lt;/p&gt;
&lt;p&gt;Pass-Through Rate: The ratio of premium reduction to government expenditure on reinsurance. In standard models with market power, pass-through of cost subsidies is typically below one; the paper documents a pass-through rate of 1.3 in Colorado (p = 0.037 for the null of pass-through equal to one), attributing the excess to the risk protection channel reducing both expected cost and cost uncertainty simultaneously.&lt;/p&gt;
&lt;p&gt;Stop-Loss Reinsurance: A contract structure in which the reinsurer reimburses the primary insurer for individual claims costs exceeding a deductible (attachment point) kappa up to a cap. In Colorado&amp;rsquo;s program the attachment point is $30,000 and the cap is $400,000, with government coinsurance rates of 40–80% depending on county tier. More generous reinsurance corresponds to lower kappa; full reinsurance is kappa = 0.&lt;/p&gt;
&lt;p&gt;Risk-Based Capital (RBC) Ratio: The ratio of capital surplus (assets minus liabilities) to required risk-based capital, used by NAIC as a measure of insurer solvency. NAIC scrutinizes companies with RBC ratios below 200%; the paper uses RBC ratio below 3 as a proxy for financial constraint in heterogeneity analysis, finding larger premium and private reinsurance responses among constrained insurers.&lt;/p&gt;
&lt;p&gt;Tail-End Risk: The risk arising from the possibility that a small fraction of enrollees incurs extremely high medical costs, concentrated in the right tail of the claims distribution. In Colorado, the top 5% of consumers account for 68% of total expenses; tail-end risk is especially severe for small insurers with fewer than 10,000–100,000 enrollees and is the primary motivation for private reinsurance purchases even at above-actuarially-fair prices.&lt;/p&gt;</description></item><item><title>Investing in Influence: Investors, Portfolio Firms, and Political Giving</title><link>https://macropaperwarehouse.com/papers/investing-in-influence-investors-portfolio-firms-and-political-giving/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/investing-in-influence-investors-portfolio-firms-and-political-giving/</guid><description>&lt;p&gt;This paper investigates whether institutional investors influence the political activities of their portfolio firms, using political action committee (PAC) giving as a window into the broader question of whether institutional investors can leverage their concentrated ownership to extract benefits from portfolio firms for their own interests rather than those of their clients.&lt;/p&gt;
&lt;p&gt;The sample covers 574 institutional investors (those with at least $100 million in assets under management, i.e., 13-F filers) matched to 2,456 portfolio firms that had PACs, over the period 1980–2018. The primary source of variation is the first acquisition by an institutional investor of at least one percent of a portfolio firm&amp;rsquo;s outstanding shares, yielding 68,387 large acquisition events. PAC giving data come from FEC records matched by name to investor and firm entities. The main regression specification examines how the relationship between investor and firm PAC contributions to the same congressional district changes after such an acquisition, using a saturated set of fixed effects including firm × investor, firm × congressional district, firm × election cycle, investor × congressional district, investor × election cycle, and district × election cycle.&lt;/p&gt;
&lt;p&gt;The central finding is that, following a large block purchase, a firm&amp;rsquo;s PAC giving mirrors more closely that of the acquiring investment management company. In the preferred specification (column 8 of Table 2), the probability that a portfolio firm gives to a politician supported by its investor&amp;rsquo;s PAC increases by 31 percent after an acquisition. Using a cosine similarity measure of investor-firm PAC giving, the mean similarity of 0.10 at the acquisition cycle rises by 0.02–0.03 (a 20–30 percent increase) by the fourth post-acquisition election cycle.&lt;/p&gt;
&lt;p&gt;A key identification concern is that acquisitions may be driven by shared political preferences rather than representing a causal effect. To address this, the authors exploit stock index inclusions as exogenous shifters of institutional investor block purchases: when a firm is added to an index for the first time, passive indexers are compelled to rebalance toward that firm regardless of political alignment. Restricting to 5,601 index-inclusion acquisitions by passive investors, the authors find near-identical effect sizes (beta1 = 0.0132 in column 8 versus 0.0135 in the full sample), and an event study shows no pre-trend in giving convergence for the index subsample, in contrast to a slight pre-trend in the full sample. Divestment events exhibit the symmetric negative pattern: the interaction of post-divestment and investor PAC giving falls by between -0.074 and -0.058 across specifications.&lt;/p&gt;
&lt;p&gt;The authors argue that investors drive the convergence rather than portfolio firms adjusting investor preferences. Around acquisition dates, firms exhibit a larger drop in between-election-cycle cosine similarity than investors do. In a difference-in-differences comparison of the acquisition period relative to the preceding period, the difference in stability between investors and firms is 0.075 (significant at the 1 percent level), indicating that firms shift their giving more than investors. Investors obtaining a board seat at the portfolio firm amplifies the effect: in the preferred specification, the board-seat interaction is more than twice as large as the acquisition-alone interaction.&lt;/p&gt;
&lt;p&gt;Heterogeneity analysis provides evidence that the convergence reflects investors&amp;rsquo; partisan tastes rather than coordinated profit-maximizing political strategy. Acquisitions by more partisan investors (those whose giving is more skewed toward one party) produce a convergence coefficient roughly twice as large (0.020) as less partisan investors (0.010). Private fund families show more than twice the convergence effect of publicly owned fund families. The partisan composition of firm giving also shifts: a firm acquired by an investor giving exclusively to Republicans sees its Republican share increase by 2.8 percentage points relative to a baseline of 47.4 percent (a 5.9 percent increase).&lt;/p&gt;
&lt;p&gt;Finally, higher overall institutional ownership is associated with an increase in total PAC giving at the firm level, and this expanded giving does not go disproportionately to politicians on committees overseeing issues the firm actively lobbies — suggesting the ownership-driven increment in political spending is non-strategic from the firm&amp;rsquo;s profit standpoint and likely serves investors&amp;rsquo; own interests.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the central research question and why does it matter?&lt;/strong&gt;
The paper asks whether institutional investors influence the political giving of portfolio firms, motivated by the broader concern that the rise of institutional ownership — from 6 percent of U.S. public equities in 1950 to 65 percent in 2017 — concentrates not only economic but also political power in the hands of a small number of asset managers. This matters because if investors shape firms&amp;rsquo; PAC giving to serve investors&amp;rsquo; own preferences rather than firms&amp;rsquo; profit interests, it represents a misuse of corporate resources and a potential amplification of a small group&amp;rsquo;s political voice.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What data are used and how is the sample constructed?&lt;/strong&gt;
The analysis draws on 13-F filings (investors with at least $100M AUM) from Thomson-Reuters, matched to FEC PAC records via fuzzy and manual name matching. The resulting sample contains 574 investors with PACs and 2,456 portfolio firms with PACs, spanning 1980–2018. The Cartesian product of investor-firm pairs is restricted to those connected by at least one large acquisition event (defined as first acquisition of at least 1 percent of outstanding shares), yielding 68,387 such events. PAC contributions are measured at the investor- and firm-congressional-district-election-cycle level, linked to House of Representatives winners using MIT Election Data files.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the baseline regression and what does it find?&lt;/strong&gt;
The baseline regression (equation 1) interacts Log Investor PAC with a Post indicator (equal to 1 after the first large acquisition and while the stake is maintained) at the investor-firm-congressional-district-election-cycle level, with a saturated set of fixed effects. The coefficient on the interaction (beta1) is positive and highly significant (p &amp;lt; 0.001) across all eight specifications, ranging from 0.013 to 0.032. In the preferred specification, the increase in giving similarity is 31 percent relative to the pre-acquisition baseline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do the authors establish causality and rule out endogenous acquisitions?&lt;/strong&gt;
The primary identification strategy uses first-time inclusions of firms in stock indices (approximately 1,000 indices tracked in the sample) as exogenous shifters: passive indexers must rebalance toward the included firm regardless of political alignment. This subsample of 5,601 index-inclusion acquisitions produces near-identical coefficient estimates (0.0132 versus 0.0135 in the full sample), and the event study for this subsample shows no pre-trend in giving convergence, unlike the slight pre-trend in the full sample. Equality of the two coefficients cannot be rejected at standard significance levels.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What evidence shows it is firms adjusting to investors rather than the reverse?&lt;/strong&gt;
The authors compute between-election-cycle cosine similarity separately for investors and firms around acquisitions. On average, investors exhibit more stable giving than firms at acquisition dates (Cos(xi,t, xi,t+1) &amp;gt; Cos(xf,t, xf,t+1)). The difference-in-differences estimate — comparing the acquisition period to the preceding period — is 0.075 (significant at 1 percent), indicating a relatively larger break in firm giving. Over a two-cycle window, the difference-in-differences estimate is 0.083, again indicating convergence is driven by firms shifting toward investors rather than the reverse.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What role does board representation play?&lt;/strong&gt;
In approximately 5 percent of acquisitions in the sample, the investor obtains a board seat. In specifications that include both the acquisition effect (Post × Log Investor PAC) and a board-membership interaction (Board × Log Investor PAC), both terms are positive and significant at the 1 percent level. In the preferred specification, the board-seat interaction is more than twice as large as the acquisition-alone interaction, indicating that a direct governance channel — board representation — substantially amplifies the convergence in political giving.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the divestment analysis show?&lt;/strong&gt;
Symmetric to the acquisition results, divestment events (where an investor exits a stake of at least 1 percent held for at least one election cycle) are associated with a decline in investor-firm PAC giving correlation. Post-divestment interaction coefficients range from -0.074 to -0.058 across specifications, and an event study confirms the correlation falls sharply after the divestment cycle.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does investor partisanship affect the magnitude of influence?&lt;/strong&gt;
Yes. Classifying investors as &amp;ldquo;More Partisan&amp;rdquo; (above-mean absolute deviation from 50/50 party split) versus &amp;ldquo;Less Partisan,&amp;rdquo; the interaction coefficient for More Partisan investors (0.020) is roughly twice that of Less Partisan investors (0.010). After a large acquisition by a fully Republican-giving investor, the acquired firm&amp;rsquo;s giving to that politician increases by 23.5 percent; the comparable figure for a Less Partisan investor is 7.6 percent. This pattern holds in both the full sample and the index-inclusion subsample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do private versus public fund families differ in their influence?&lt;/strong&gt;
Private fund families (e.g., Vanguard, Fidelity) show more than twice the convergence coefficient of publicly owned fund families (e.g., BlackRock, State Street, Invesco). The authors attribute this to private fund managers facing less outside scrutiny, allowing their giving to more readily reflect the preferences of owners and managers. Private investors also show greater partisan polarization: the 10th–90th percentile Republican-giving range for private investors is 6.3–100 percent, versus 21.7–88.3 percent for public investors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does increased institutional ownership expand overall firm PAC spending?&lt;/strong&gt;
Yes. In firm-year level regressions, institutional ownership is a positive and significant predictor of total firm PAC giving (significant at at least the 5 percent level in both cross-sectional and firm-fixed-effects specifications). Total corporate political expenditure by sample firms increased by nearly a factor of six over 1980–2018. The authors note that while many factors contribute, increased institutional ownership may be at least partly responsible for this expansion.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does the additional giving driven by institutional ownership go to strategically important politicians for the firm?&lt;/strong&gt;
No. Regressions relating institutional ownership to giving to politicians on congressional committees overseeing issues the firm actively lobbies (a standard measure of politicians&amp;rsquo; strategic importance to firms) yield near-zero and statistically weak point estimates. In the preferred firm-fixed-effects specification, the share of total PAC giving devoted to such strategically relevant politicians is negatively associated with institutional ownership at marginal significance (p &amp;lt; 0.10), consistent with the interpretation that ownership-driven incremental political spending is non-strategic from the firm&amp;rsquo;s own profit perspective and expands total giving rather than displacing strategic giving.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the policy and legal implications?&lt;/strong&gt;
The authors flag three concerns: (i) the ownership-driven increment in political spending may represent a misuse of corporate resources that does not serve portfolio firm shareholders; (ii) it may constitute an illegal activity, since using a firm&amp;rsquo;s PAC to reimburse or proxy for an investor&amp;rsquo;s own political preferences can run afoul of campaign finance law; and (iii) it is a channel through which unequal resources amplify the political voice of a small number of fund managers at the expense of dispersed ultimate investors who are likely unaware of and do not sanction these contributions. The findings challenge the Supreme Court&amp;rsquo;s premise in Citizens United that corporate political speech reflects shareholder profit maximization.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PAC comovement (investor-firm giving similarity):&lt;/strong&gt; The increase in the probability that a portfolio firm&amp;rsquo;s PAC donates to a politician also supported by an acquiring investor&amp;rsquo;s PAC, measured as the interaction coefficient between Log Investor PAC and a Post-acquisition indicator in the baseline regression. In the preferred specification this represents a 31 percent increase relative to the pre-acquisition baseline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cosine similarity (cross-time and cross-entity):&lt;/strong&gt; A measure defined as the Euclidean dot product between two vectors of PAC giving (either the same entity across adjacent election cycles, or investor versus firm in the same cycle), taking values between 0 and 1, where 1 indicates identical giving patterns. Used both to confirm convergence post-acquisition and to attribute that convergence to firm rather than investor adjustment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Index-inclusion acquisition:&lt;/strong&gt; A large block purchase that results from a firm being added for the first time to a stock index tracked by a passive institutional investor, used as an exogenous shifter of investor stakes that is orthogonal to investor-firm political alignment. There are 5,601 such events in the sample.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Partisanship (investor):&lt;/strong&gt; Classified as &amp;ldquo;More Partisan&amp;rdquo; if an investor&amp;rsquo;s absolute deviation from a 50/50 party split in PAC donations is above the sample mean. More partisan investors produce roughly twice the convergence effect on portfolio firm giving compared to less partisan investors, used as evidence that personal political preferences rather than profit-maximizing business strategy drive the convergence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Post indicator (Postift):&lt;/strong&gt; A binary variable equal to 1 for all election cycles following an investor&amp;rsquo;s first acquisition of at least 1 percent of a portfolio firm&amp;rsquo;s outstanding shares, and remaining 1 as long as the investor holds any stake in the firm. The key source of temporal variation in the baseline regression.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Strategically important politicians:&lt;/strong&gt; Members of Congress sitting on committees that oversee issues on which a firm actively lobbies, identified by crosswalking lobbying reports from the Senate Office of Public Records to relevant committee jurisdictions. Used to test whether ownership-driven political giving displaces or supplements firm-profit-motivated giving.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Board seat channel:&lt;/strong&gt; The mechanism through which investor influence on firm political giving is amplified when the investor obtains representation on the portfolio firm&amp;rsquo;s board of directors (present in approximately 5 percent of acquisitions). The board interaction coefficient is more than twice the acquisition-alone coefficient in the preferred specification.&lt;/p&gt;</description></item><item><title>Making the Invisible Hand Visible: Managers and Worker Allocation</title><link>https://macropaperwarehouse.com/papers/making-the-invisible-hand-visible-managers-and-worker-allocation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/making-the-invisible-hand-visible-managers-and-worker-allocation/</guid><description>&lt;p&gt;This paper asks why managers matter for firm performance, and specifically whether managers improve productivity by matching workers to better-suited jobs inside firms rather than through supervision, motivation, or selection out of the firm. The setting is the internal labor market of a large private consumer goods multinational enterprise (MNE) operating in more than 100 countries, with annual turnover exceeding EUR 50 billion. The data cover the universe of white-collar workers and managers at the firm — 200,000 workers and 30,000 managers observed monthly over 11 years (January 2011 to December 2021) — linked to payroll, performance ratings, organizational chart, digital platform activity, employee surveys, and an independent sales productivity series for field sales workers in 15 countries.&lt;/p&gt;
&lt;p&gt;The paper confronts two identification challenges. First, the author constructs a measure of manager quality — &amp;ldquo;high flyers&amp;rdquo; — defined as managers who were promoted to the first managerial work level (WL2) by age 30. This threshold yields 26.2% of managers classified as high flyers. The measure is defined entirely ex ante, before the manager ever supervises the worker under study, which addresses reverse causality. It is validated against ex post performance metrics including future salary growth, probability of promotion to WL3, performance ratings, and anonymous subordinate feedback. Second, to identify causal effects of manager quality on workers, the author exploits the firm&amp;rsquo;s long-standing policy of rotating WL2 managers laterally across teams as part of their career development, a practice implemented for several decades. Using an event-study design centered on the worker&amp;rsquo;s first manager transition, the author compares workers who transition from a low-flyer to a high-flyer manager (LtoH) against workers who transition from one low-flyer to a different low-flyer (LtoL), netting out the effect of the transition itself. Pre-event parallel trends are confirmed empirically.&lt;/p&gt;
&lt;p&gt;The main findings are as follows. Gaining a high-flyer manager causes substantial reallocation of workers within the firm through lateral job transfers: seven years after the manager transition event, cumulative lateral moves are 40% higher for workers who gained a high-flyer manager relative to those who gained another low-flyer. These lateral moves are not confined to a single organizational margin — transfers rise within-team, across teams in the same function, and across functions — and they involve meaningfully larger shifts in task content, as measured by angular separation across O*NET cognitive, routine, and social task intensity dimensions, with cumulative task distance becoming statistically distinguishable from zero approximately seven quarters post-transition. These gains in lateral mobility translate into persistent wage growth: seven years after the manager transition, workers supervised by a high-flyer earn salaries 13% higher than the comparison group, with divergence beginning only after the transition date. Using independent sales bonus data, three years after gaining a high-flyer manager workers&amp;rsquo; sales productivity increases by 0.347 standard deviations, ruling out the interpretation that wage gains merely reflect manager favoritism rather than genuine productivity improvement. Establishment-level data further show that sites with a higher share of workers under high-flyer managers display higher output per worker and lower operational costs per unit.&lt;/p&gt;
&lt;p&gt;Effects are asymmetric: gaining a good manager has large positive effects, but losing one (comparing HtoL with HtoH transitions) produces no corresponding negative effects, implying that a single exposure to a high-flyer manager generates durable benefits that survive a subsequent downgrade in manager quality. A mediation analysis finds that 64% of the salary gain is explained by lateral job changes, though the author notes this understates the full allocation channel because it excludes vertical transfers and the gains from remaining well-matched in the current role. These findings hold under multiple robustness checks including restricting to new hires, using the Sun and Abraham (2021) interaction-weighted estimator, varying the age threshold for high-flyer classification, using a tenure-based alternative, and placebo tests with randomly assigned manager types.&lt;/p&gt;
&lt;p&gt;The scope conditions are specific to white-collar workers at a large, organizationally homogeneous consumer goods multinational. All workers hold college degrees, mean firm tenure is 8.5 years, team sizes average five workers, and the firm has the same organizational structure across all countries, functions, and years.&lt;/p&gt;
&lt;p&gt;Q: How does the paper define &amp;ldquo;high flyer&amp;rdquo; managers and what share of managers receive this classification?
A: High flyers are managers who achieved the first managerial work level (WL2) by age 30, a threshold derived from continuous age estimates constructed from 10-year age bands in the personnel records. This definition yields 26.2% of managers classified as high flyers. The measure is time-invariant and defined ex ante relative to any interaction with the workers whose outcomes are studied.&lt;/p&gt;
&lt;p&gt;Q: What validates the high-flyer measure as capturing genuine managerial ability rather than noise?
A: The high-flyer classification is significantly positively correlated with multiple ex post performance metrics recorded after the manager&amp;rsquo;s own promotion: future salary growth, probability of subsequent promotion to WL3 (director level), annual performance ratings, and anonymous upward feedback scores from subordinates on leadership. High flyers are also 14.5 percentage points less likely to be mid-career recruits, suggesting they are internally developed talent rather than external hires.&lt;/p&gt;
&lt;p&gt;Q: What is the source of identifying variation and how does the event-study design address endogeneity?
A: The firm has operated a decades-long policy of rotating WL2 managers laterally across teams to broaden their experience and to screen candidates for promotion to WL3. These rotations are asserted by firm executives and HR representatives to be orthogonal to worker and team characteristics. The author verifies this empirically by showing that a wide range of team characteristics measured over the two years before a transition — including team performance, inequality, transfer rates, and team diversity — cannot predict the type of incoming manager. The event-study design compares workers who receive a high-flyer replacement (LtoH) against workers who receive another low-flyer replacement (LtoL), netting out any generic effect of a managerial change, and confirms parallel pre-trends.&lt;/p&gt;
&lt;p&gt;Q: What is the effect of gaining a high-flyer manager on lateral job mobility?
A: Seven years after the manager transition, workers assigned to a high-flyer manager exhibit lateral moves that are 40% higher relative to workers assigned to another low-flyer. These lateral moves occur across all organizational margins: within the same team, across teams within the same function (the largest contributor), and across functions. Beyond frequency, lateral moves under high-flyer managers also involve larger task-content shifts, with cumulative task distance (measured using O*NET cognitive, routine, and social task dimensions via angular separation) becoming statistically distinguishable from zero approximately seven quarters after the transition.&lt;/p&gt;
&lt;p&gt;Q: What is the wage effect of gaining a high-flyer manager and when does it materialize?
A: Workers who transition from a low-flyer to a high-flyer manager earn a salary 13% higher than workers who transition to another low-flyer, measured seven years after the transition event. The divergence begins only after the transition date, consistent with the pre-event parallel trends assumption, and accumulates gradually rather than appearing as an immediate jump.&lt;/p&gt;
&lt;p&gt;Q: Does the wage gain reflect genuine productivity improvement or simply managerial favoritism in pay decisions?
A: The author uses an independent sales bonus series — based on monthly targets set by supply chain demand planning teams, not by managers — for 5,604 field sales workers in 15 countries from 2018 to 2021. Three years after gaining a high-flyer manager, workers&amp;rsquo; sales productivity increases by 0.347 standard deviations. This confirms that pay gains correspond to actual productivity improvement rather than inflated ratings for unchanged performance.&lt;/p&gt;
&lt;p&gt;Q: How much of the wage gain is attributable to the lateral reallocation channel specifically?
A: A mediation analysis attributes 64% of the 13% salary gain to lateral job changes. The author cautions that this is a lower bound because the mediation excludes vertical transfers (which mechanically raise salary) and does not capture gains for workers who remain in their current job because it represents a good match rather than requiring reallocation.&lt;/p&gt;
&lt;p&gt;Q: Are the effects symmetric — does losing a high-flyer manager reverse the gains?
A: No. Comparing workers who transition from a high-flyer to a low-flyer manager (HtoL) against workers who transition from a high-flyer to another high-flyer (HtoH) reveals no corresponding negative effects. The gains from a single prior exposure to a high-flyer manager are persistent and are not undone by a subsequent low-quality manager. The author interprets this as evidence that a good match, once created, endures independently of the manager who created it.&lt;/p&gt;
&lt;p&gt;Q: Does gaining a high-flyer manager raise the rate of worker exit from the firm?
A: No. There is no statistically detectable effect on either voluntary exits (quits) or involuntary exits (layoffs), with null results that are not masked by heterogeneity across high- and low-performing workers. This rules out the interpretation that high-flyer managers improve measured outcomes of retained workers by selecting out underperformers.&lt;/p&gt;
&lt;p&gt;Q: Do workers move into roles connected to their high-flyer manager&amp;rsquo;s prior network or follow their manager when the manager moves?
A: No. There is no evidence that workers move into roles connected to the high-flyer manager&amp;rsquo;s prior colleagues; if anything, subordinates of high-flyer managers are less likely to make such moves. Workers also do not follow their high-flyer managers when those managers subsequently rotate to a different team. These findings rule out favoritism, social network access, and information-advantage explanations as primary drivers.&lt;/p&gt;
&lt;p&gt;Q: How does the paper rule out on-the-job teaching (human capital transmission) as the primary mechanism?
A: If high-flyer managers improved worker outcomes primarily by teaching workers to be more productive in their current job, the prediction would be reduced lateral mobility (workers become too productive to leave their current role). The observed pattern — substantially higher rates of lateral reallocation under high-flyer managers — is the opposite of this prediction, making teaching as the dominant channel unlikely.&lt;/p&gt;
&lt;p&gt;Q: What does the manager behavior evidence show about how high flyers spend their time?
A: Time-use data from a random sample of approximately 600 WL2 managers in 2019 show that high-flyer managers spend 19% more time in one-on-one meetings with subordinates and engage more in communication and multitasking activities relative to low-flyer managers. Their skill profiles also differ: high flyers are more likely to have strengths in strategy and talent management rather than project management, consistent with a more coordination-intensive and people-development-oriented style.&lt;/p&gt;
&lt;p&gt;Q: What heterogeneity is there in who benefits from high-flyer managers?
A: Effects are larger when managers and workers are in the same physical office (proximity facilitates talent assessment), when the organizational unit has a more diverse set of job roles (more matching opportunities), and for younger workers who are still discovering their comparative advantages. Critically, benefits are not concentrated among high-baseline performers: workers with low initial pay growth experience gains comparable to those of high performers, suggesting high-flyer managers uncover and deploy hidden talent broadly rather than accelerating only already-visible stars.&lt;/p&gt;
&lt;p&gt;Q: Does high-flyer management aggregate to establishment-level productivity?
A: Yes. Establishments where a higher share of workers are supervised by high-flyer managers show higher output per worker (tons per FTE) and lower operational costs per unit of output (operational costs per ton), measured using establishment-year data across approximately 150 sites globally over 2019-2021. This is consistent with the individual-level allocation mechanism producing aggregate productivity gains.&lt;/p&gt;
&lt;p&gt;Q: What are the organizational design implications of the asymmetric effects?
A: Because the gains from a single exposure to a high-flyer manager persist even after a subsequent manager downgrade, firms do not need each worker to be continuously supervised by a high-flyer. It is sufficient to rotate high-flyer managers across teams so that each worker receives at least one exposure. This makes the allocation mechanism resource-neutral relative to hiring, firing, or formal training programs.&lt;/p&gt;
&lt;p&gt;High flyer (paper&amp;rsquo;s definition): A manager who achieved the first managerial work level (WL2) at the firm by age 30 — a time-invariant, ex ante classification representing the firm&amp;rsquo;s revealed-preference assessment of leadership potential, validated against subsequent salary growth, promotion probability, performance ratings, and subordinate feedback. Constitutes 26.2% of managers in the sample.&lt;/p&gt;
&lt;p&gt;Internal labor market (paper&amp;rsquo;s usage): The system within the firm through which workers are allocated to jobs via lateral transfers and vertical promotions, mediated by managers rather than by external price mechanisms; the institutional context within which manager-worker matching produces wage growth and productivity gains.&lt;/p&gt;
&lt;p&gt;Lateral transfer (paper&amp;rsquo;s usage): A horizontal reallocation of a worker to a different job title, team, subfunction, or function at the same work level, as distinct from a vertical promotion. Captured monthly in personnel records; operationalized as moves involving changes in task content measured by O*NET task distances.&lt;/p&gt;
&lt;p&gt;Task distance (paper&amp;rsquo;s usage): The angular separation between origin and destination occupations across three O*NET task dimensions (cognitive, routine, and social intensity), ranging from zero (identical task profiles) to one (completely distinct profiles), used to characterize the substantive scope of lateral moves induced by high-flyer managers.&lt;/p&gt;
&lt;p&gt;Manager rotation (paper&amp;rsquo;s usage): The firm&amp;rsquo;s longstanding policy of reassigning WL2 managers laterally across teams within a subfunction, designed to broaden managerial experience and screen for promotion to WL3; treated in the empirical strategy as generating plausibly exogenous variation in the manager type each worker encounters.&lt;/p&gt;
&lt;p&gt;Allocation mechanism (paper&amp;rsquo;s usage): The process by which managers discover workers&amp;rsquo; specific skills and match them to specialized jobs inside the firm, operating through lateral reallocation rather than through hiring, firing, or on-the-job training; identified in the paper as the primary channel through which high-flyer managers generate persistent wage and productivity gains.&lt;/p&gt;
&lt;p&gt;Asymmetric persistence (paper&amp;rsquo;s usage): The empirical pattern in which the gains from gaining a high-flyer manager are large and durable, while losing a high-flyer manager (transitioning to a low-flyer) produces no corresponding negative effects on the outcomes of previously well-matched workers, implying that good matches, once formed, survive a change in manager quality.&lt;/p&gt;</description></item><item><title>Manager Pay Inequality and Market Power</title><link>https://macropaperwarehouse.com/papers/manager-pay-inequality-and-market-power/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/manager-pay-inequality-and-market-power/</guid><description>&lt;p&gt;This paper asks whether managers are paid for market power. Bao, De Loecker, and Eeckhout build a general equilibrium model in which firms compete oligopolistically in goods markets (following Atkeson and Burstein 2008) while managers are allocated to firms through a competitive matching market (following Gabaix and Landier 2008 and Tervio 2008). The model identifies two distinct channels through which market power and firm size jointly determine executive compensation: a market power channel, whereby a more productive firm charges a higher markup given its output level, and a firm size channel, whereby higher total factor productivity expands output given markups. Because manager ability and firm type are complementary inputs into TFP, assortative matching arises: high-ability managers sort into high-type firms, amplifying both productivity dispersion and markup dispersion across firms.&lt;/p&gt;
&lt;p&gt;The authors estimate the model year-by-year using Simulated Method of Moments on Compustat data covering 1994 to 2019, targeting ten moments including the average salary share, markup distribution, employment, and manager compensation levels. Firm-level markups are estimated using the production approach of De Loecker, Eeckhout, and Unger (2020). The ExecuComp variable TDC1 — encompassing salary, bonus, restricted stock grants, and option grant values — measures manager pay. Finance, insurance, and real estate sectors (SIC 6000–6799) are excluded.&lt;/p&gt;
&lt;p&gt;Main findings: market power accounts for on average 45.8% of total manager pay over the sample period, rising from 38.0% in 1994 to 48.8% in 2019. Over the full period, average CEO compensation (net of reservation utility) roughly doubled, from approximately $2.94 million to $6.43 million. Of the $3.49 million cumulative increase, $2.02 million (57.8%) is attributed to rising market power, with the remainder ($1.47 million) due to the firm size channel. The market power channel&amp;rsquo;s dominance is concentrated among top managers: for the highest-ranked managers in 2019, 80.3% of pay is attributable to market power, and nearly all of their pay growth since 1994 stems from the market power channel. For lower-ranked managers, pay is determined primarily by the firm size channel and has been roughly flat over the period.&lt;/p&gt;
&lt;p&gt;Within the market power channel, changes in technology — specifically increasing dispersion in firm-level TFP — are the dominant factor, contributing $1.33 million (65.9% of total market power channel growth). The increasing importance of manager ability (rising parameter alpha) contributes an additional $1.14 million through the market power channel. Within the firm size channel, TFP change accounts for 70.1% ($1.03 million) of growth, but the large effects from rising alpha and rising complementarity (gamma) are substantially offset by increasing dispersion in firm type. Structural estimates confirm that the average number of firms per market declines from 4.40 to 3.15, and firm-type dispersion (sigma_z) rises from 0.51 to 0.77, both consistent with rising market power over the period.&lt;/p&gt;
&lt;p&gt;A counterfactual economy with no market power — firms priced at marginal cost — would yield a social welfare gain of 58.4% on average. The welfare cost of market power in 1994 could be offset by a 33.8% TFP increase; by 2019 the required TFP offset had risen to 51.7%. Without any market power, even the most talented managers would earn only their reservation utility, because firms earn zero profits regardless of productivity, eliminating the complementarity-driven matching surplus that makes top managers valuable. This confirms that superstar manager pay is intrinsically tied to the existence of market power in goods markets, not solely to firm size.&lt;/p&gt;
&lt;p&gt;Scope conditions: the model applies to publicly listed US firms covered by Compustat and ExecuComp. The mechanism relies on Cournot competition within oligopolistic markets, assortative matching between managers and firms, and complementarity between manager ability and firm type (elasticity of substitution gamma estimated to be negative throughout the sample). The findings on market power share apply to CEOs specifically; the authors argue the same logic extends to all managerial positions with span-of-control over other workers, which encompasses roughly one-fifth of the workforce.&lt;/p&gt;
&lt;p&gt;Q: What are the two channels through which manager pay is determined in the model, and how do they differ mechanically?
A: The market power channel captures how a given level of TFP translates into higher markups — more productive firms charge more above marginal cost — thereby increasing profits per unit of output. The firm size channel captures how higher TFP expands the quantity of output a firm produces, increasing total profits through scale rather than through price-cost margin. Both channels raise profits and thus the marginal product of managers, but they operate through distinct economic mechanisms: one through pricing power and the other through productive scale.&lt;/p&gt;
&lt;p&gt;Q: What is the empirical magnitude of the market power channel&amp;rsquo;s contribution to manager pay levels and growth?
A: Market power accounts for an average of 45.8% of total manager pay over 1994–2019, rising monotonically from 38.0% in 1994 to 48.8% in 2019. For the total pay increase of $3.49 million over the period, $2.02 million (57.8%) is due to the increase in market power, with the remaining $1.47 million attributable to the firm size channel.&lt;/p&gt;
&lt;p&gt;Q: How does the market power channel&amp;rsquo;s importance vary across the manager ability distribution?
A: For the highest-ranked managers, 80.3% of total pay in 2019 is attributable to market power, and nearly all of their pay growth since 1994 runs through the market power channel. For the lowest-ranked managers, pay is almost entirely explained by the firm size channel and has been approximately flat over the period. This heterogeneity arises because top managers sort into high-markup firms through assortative matching, making their compensation disproportionately dependent on those firms&amp;rsquo; market power.&lt;/p&gt;
&lt;p&gt;Q: How does the model generate assortative matching between manager ability and firm type?
A: Manager ability and firm type are complementary inputs into TFP (the CES aggregator with elasticity of substitution gamma less than one), which makes the matching output supermodular. In a frictionless matching market with transferable utility, supermodularity guarantees that high-ability managers match with high-type firms in equilibrium (Proposition 1). This positive assortative matching then amplifies productivity and markup dispersion, since the most productive firms become even more productive and gain larger market shares.&lt;/p&gt;
&lt;p&gt;Q: What structural changes drive the rising importance of market power in manager pay over time?
A: The dominant factor within the market power channel is changes in technology, specifically increasing firm-type dispersion (sigma_z rising from 0.51 to 0.77), which contributes $1.33 million or 65.9% of market power channel growth. The rising importance of manager ability (alpha, the weight on manager ability relative to firm type in the TFP aggregator) contributes another $1.14 million. The number of firms per market declines from an average of 4.40 to 3.15, further reducing competitive pressure and amplifying the markup premium for high-productivity firms.&lt;/p&gt;
&lt;p&gt;Q: What does the counterfactual with no market power (first-best pricing) imply for manager pay and social welfare?
A: Without market power, firms price at marginal cost and earn zero profits regardless of productivity, which eliminates the surplus from manager-firm matching. All managers would earn only their reservation utility, which is negligible relative to actual compensation. Social welfare would increase by 58.4% on average. The efficiency cost of market power — measured as the TFP increase needed to offset welfare losses — rose from 33.8% in 1994 to 51.7% in 2019, indicating a worsening welfare distortion over the period.&lt;/p&gt;
&lt;p&gt;Q: How are markups measured, and what is their trend in the data?
A: Markups are not directly observable and are estimated using the production approach of De Loecker, Eeckhout, and Unger (2020), which recovers firm-level price-cost margins from production data without requiring price data. Average markups in the Compustat sample rose from 1.53 in 1994 to 1.78 in 2019. The reduced-form elasticity of manager pay with respect to markups (controlling for firm characteristics, year, and firm fixed effects) increased substantially: in 2019 a one-percent increase in firm-level markup raises manager pay by 0.41 percent, which is 70.1% larger than the effect estimated in 1994.&lt;/p&gt;
&lt;p&gt;Q: How does the paper handle the identification challenges inherent in regressing manager pay on markups?
A: The reduced-form regression (with firm fixed effects, year effects, and interactions of year dummies with markups) documents a robust positive correlation but cannot establish causality due to reverse causality and omitted-variable bias. The paper addresses this by embedding the markup-manager pay relationship in a structural model where both are jointly determined by primitives — technology, market structure, and manager ability — and estimating those primitives via Simulated Method of Moments. The quantitative decomposition into market power and firm size channels derives from the model structure rather than from identifying variation in an instrumental variables sense.&lt;/p&gt;
&lt;p&gt;Q: What do the matching model estimates reveal about manager-firm complementarity over time?
A: The estimated elasticity of substitution between manager ability and firm type (gamma) is negative throughout the sample, confirming complementarity. Gamma was relatively stable before declining sharply from -2.22 in 2014 to -3.55 in 2019, indicating that manager ability and firm type became substantially more complementary in the latter part of the sample. The importance-of-manager parameter alpha is small (consistent with Gabaix and Landier 2008) but generally increasing, suggesting managers play an expanding role in determining firm-level TFP over time.&lt;/p&gt;
&lt;p&gt;Q: What are the broader macroeconomic and distributional implications of the findings?
A: Because approximately one-fifth of workers supervise other workers, the market-power-driven premium in managerial pay has implications beyond CEO compensation for the shape of the earnings distribution. The rise in top-1-percent income is identified as an efficiency concern, not just an equity concern: the best managers are hired by high-markup firms where they generate profits for shareholders but disproportionately little additional social value. Assortative matching between top managers and top firms widens the productivity gap between competitors, increasing market power and deadweight loss — the social return to managerial talent is therefore below the private return in equilibrium.&lt;/p&gt;
&lt;p&gt;Market Power Channel: The component of manager pay attributable to how a firm&amp;rsquo;s TFP raises its markup — the ratio of output price to marginal cost — given the level of output. Distinct from the firm size channel; operates through pricing power rather than scale.&lt;/p&gt;
&lt;p&gt;Firm Size Channel: The component of manager pay attributable to how a firm&amp;rsquo;s TFP expands output quantity given markups. Increasing output scale raises total profits and thus the marginal product of the manager even absent any change in price-cost margins.&lt;/p&gt;
&lt;p&gt;Assortative Matching: The equilibrium allocation of high-ability managers to high-type firms, arising because manager ability and firm type are complementary inputs into TFP (supermodular matching output). Matching is determined in a frictionless market with transferable utility.&lt;/p&gt;
&lt;p&gt;Markup: The ratio of output price to marginal cost, equal to the inverse of the price elasticity of demand under the nested CES preference structure. Endogenously determined by the firm&amp;rsquo;s sales share within its oligopolistic market and the elasticities of substitution within markets (eta) and across markets (theta).&lt;/p&gt;
&lt;p&gt;Manager-Firm Complementarity: The property that manager ability and firm type are imperfect substitutes with elasticity of substitution gamma less than one in the TFP aggregator. Complementarity is the necessary condition for positive assortative matching and for the supermodularity of matching surplus.&lt;/p&gt;
&lt;p&gt;Span of Control (Lucas 1978): The mechanism by which a manager raises the productivity of all workers under supervision, so that a more able manager generates a proportionally larger productivity gain the larger the firm. Provides the microfoundation for why firm size amplifies the value of manager ability.&lt;/p&gt;
&lt;p&gt;Market Structure: The number of firms in each oligopolistic sub-market (Ij), which varies across markets and over time. Together with the distribution of firm-level TFP within a market, market structure determines how much competitive pressure limits markup extraction. Average firms per market declines from 4.40 to 3.15 over 1994–2019.&lt;/p&gt;</description></item><item><title>Market Segmentation through Information</title><link>https://macropaperwarehouse.com/papers/market-segmentation-through-information/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/market-segmentation-through-information/</guid><description>&lt;p&gt;This paper asks what market outcomes an information designer — modeled as an internet platform that knows consumers&amp;rsquo; preferences — can achieve by choosing what information to disclose to competing oligopolistic firms who then make personalized price offers. The model features n firms each producing a single differentiated product at zero cost, a continuum of consumers with unit demand and multidimensional valuations (one per product), and a designer who commits to a mapping from consumer types to joint distributions over messages sent to firms before they play a simultaneous pricing game. The designer&amp;rsquo;s objective spans the full range from maximizing producer surplus to maximizing consumer surplus.&lt;/p&gt;
&lt;p&gt;The paper establishes two main results. First, under a necessary and sufficient condition called Aggregate Incentive Compatibility (AIC), the designer can implement full surplus extraction by firms — the producer-optimal outcome — in which every consumer buys her most preferred product at a price exactly equal to her valuation for it, capturing 100% of available surplus for producers. The AIC condition requires, for each firm i and each candidate deviation price p_hat_i, that the infra-marginal losses firm i would bear on its natural customers (those in Ei who value i most) from lowering price to p_hat_i must be weakly greater than the maximum business-stealing profit available from consumers who prefer other products but have valuation for i above p_hat_i. The condition is easier to satisfy when consumer preferences are more polarized, i.e., when consumers have stronger relative preferences for their most-preferred product. When firms offer homogeneous products the condition fails everywhere and no information structure can generate any producer surplus — Bertrand competition drives all profits to zero under any signal structure.&lt;/p&gt;
&lt;p&gt;Second, the paper characterizes the consumer-optimal information structure, which achieves the maximum possible consumer surplus across all equilibria induced by any information structure. The upper bound on consumer surplus is CS* = (total surplus) minus sum_i Pi*_i, where Pi*_i is the profit firm i can guarantee itself by ignoring the designer&amp;rsquo;s signal and setting the best uniform price assuming all rivals price at zero. This bound is tight: the designer can implement it by publicly partitioning consumers into groups by most-preferred product, inducing rival firms to price at marginal cost (zero) for consumers who prefer another firm&amp;rsquo;s product, and then applying the Bergemann-Brooks-Morris (2015) extremal segmentation within each firm&amp;rsquo;s natural customer set to preserve each firm&amp;rsquo;s guarantee profit while achieving efficiency.&lt;/p&gt;
&lt;p&gt;The illustrative two-firm example shows the quantitative stakes concretely. With no information disclosure, firms charge 4/5 and total producer surplus is about 76% of total surplus S*, consumer surplus is just under 10% of S*, and some consumers are excluded. With full disclosure, producer surplus rises to about 81% of S* and consumer surplus to 19%. The producer-optimal information structure (Case 3) achieves 100% of S* as producer surplus by pooling consumers who prefer different products into the same message submarket, giving each firm an incentive to price for its highest-valuing customers and ignore the others. The consumer-optimal information structure (Case 4) brings producer surplus down to about 57% of S* — its guaranteed lower bound — and delivers roughly 43% of S* to consumers, an outcome unattainable by full disclosure alone.&lt;/p&gt;
&lt;p&gt;Both producer-optimal and consumer-optimal outcomes are efficient: all consumers buy their most-preferred product in both cases. The paper further characterizes the full efficient frontier between consumer- and producer-optimal outcomes, showing that mixing the consumer-optimal and full-information structures (or consumer-optimal, full-information, and producer-optimal structures when the latter is implementable) spans every point on the frontier.&lt;/p&gt;
&lt;p&gt;The model assumes firms will price-discriminate if they can, that the designer has full knowledge of consumer types, and that the game is played once. The core results extend to continuous type distributions as shown in Online Appendix B.2. The analysis is restricted to a monopoly platform; competition among platforms is left for future work.&lt;/p&gt;
&lt;p&gt;Q: What is the central research question and why does the two-benchmark comparison used by antitrust authorities miss important possibilities?&lt;/p&gt;
&lt;p&gt;A: The paper asks what market outcomes — combinations of consumer and producer surplus — an information designer (a platform) can achieve by choosing among all possible information structures, not just the two benchmarks of no-information and full-information. Antitrust analysis that compares only those two cases misses a vast middle ground: an intermediary can package information in ways that, for instance, implement perfect collusion (extracting all surplus as producer surplus) while appearing to use privacy-protective technologies, or can intensify competition well beyond the full-information benchmark to benefit consumers.&lt;/p&gt;
&lt;p&gt;Q: What is the producer-optimal information structure and when does it exist?&lt;/p&gt;
&lt;p&gt;A: A producer-optimal information structure is one that induces an equilibrium in which every consumer buys her most-preferred product at a price exactly equal to her valuation — full surplus extraction. It exists if and only if, for every firm i and every candidate deviation price p_hat_i, the Aggregate Incentive Compatibility (AIC) condition holds: the aggregate infra-marginal losses firm i would suffer on its natural customers Ei from lowering price to p_hat_i must be at least as large as the maximum business-stealing profit from consumers outside Ei who have valuation for i weakly above p_hat_i. This is a condition on the distribution of consumer valuations, not on the information structure per se.&lt;/p&gt;
&lt;p&gt;Q: What is the economic mechanism behind the producer-optimal structure — how does pooling consumers implement full surplus extraction?&lt;/p&gt;
&lt;p&gt;A: The designer assigns consumers who prefer product A to the same message submarket as consumers who prefer another product but have a lower valuation for A. Firm A is then price-recommended its highest-valuing customers&amp;rsquo; willingness to pay. The presence of the &amp;ldquo;outside&amp;rdquo; consumers in the same message makes it unprofitable for firm A to deviate downward to capture them, because the infra-marginal loss on the natural customers exceeds the additional revenue. Simultaneously, the rival firm cannot identify and undercut for A&amp;rsquo;s natural customers because the messages do not allow it to distinguish them. The result is that each firm plays a niche strategy, setting price equal to the valuation of its highest-type natural customers and excluding the others from its offer.&lt;/p&gt;
&lt;p&gt;Q: When does polarization of consumer preferences help achieve the producer-optimal outcome?&lt;/p&gt;
&lt;p&gt;A: Proposition 1 states that if a producer-optimal information structure exists under distribution f, it also exists under any distribution f_tilde that is more polarized than f — where more polarized means the mass of consumers who prefer i and have valuation above any threshold for i increases, and the mass of consumers who prefer j but have valuation above that threshold for i decreases. Intuitively, polarization slackens the Firm IC constraints because it reduces the business-stealing temptation: fewer consumers with high cross-product valuations are available for firm i to capture by undercutting. Concrete continuous-distribution examples include: uniform over the unit square (producer-optimal always exists), Hotelling anti-correlated values (exists everywhere), and truncated normal with mean 1/2 — producer-optimal is feasible for all standard deviations sigma &amp;gt; 0.15.&lt;/p&gt;
&lt;p&gt;Q: Why does the producer-optimal outcome fail entirely when products are homogeneous?&lt;/p&gt;
&lt;p&gt;A: Proposition 2 states that when all consumer types have equal valuations across products (the support of f lies on the diagonal of V^n), then for any information structure and any induced equilibrium, every consumer buys at price zero and all firms earn zero profit. The logic extends the standard Bertrand undercutting argument: with homogeneous products, any positive price a firm charges is undercut by a rival who can always profitably steal demand, and this applies to any posterior distribution induced by any signal realization. Even private signals cannot prevent this outcome because no signal realization can give a firm a non-contestable position.&lt;/p&gt;
&lt;p&gt;Q: How is the consumer-optimal information structure constructed, and what is its key economic logic?&lt;/p&gt;
&lt;p&gt;A: Theorem 2 shows the consumer-optimal structure has three layers. First, consumers are partitioned into n groups by most-preferred product (Ei). Second, firms j not equal to i are induced — by publicly revealing which group a consumer belongs to — to set price zero for consumers outside their group, because competing for those consumers is hopeless when their preferred firm is identified. Third, within each Ei, consumers are further partitioned into submarkets using the Bergemann-Brooks-Morris (2015) extremal segmentation applied to residual valuations (theta_i minus the maximum of competing valuations), ensuring firm i earns exactly its guarantee profit Pi*_i. By holding each firm down to its guarantee profit, the residual goes to consumers, maximizing CS.&lt;/p&gt;
&lt;p&gt;Q: What is the guarantee profit Pi*_i and how does it bound consumer surplus?&lt;/p&gt;
&lt;p&gt;A: Pi*&lt;em&gt;i is the maximum profit firm i can achieve by ignoring all designer signals and setting a single uniform price to all consumers, against the worst-case scenario in which all other firms price at zero. Formally, Pi*&lt;em&gt;i = max&lt;/em&gt;{pi} sum&lt;/em&gt;{theta in Ei: theta_i - pi &amp;gt;= max_{j not equal i} theta_j} pi * f(theta). Since firm i can always achieve Pi*_i regardless of the information structure (by simply ignoring signals), no information structure can push firm i&amp;rsquo;s profit below Pi*_i. The sum of these guarantee profits across all firms provides a lower bound on total producer surplus — and therefore an upper bound on consumer surplus — achievable by any information structure.&lt;/p&gt;
&lt;p&gt;Q: In the two-firm numerical example, what is the quantitative comparison across the four cases?&lt;/p&gt;
&lt;p&gt;A: Total available surplus S* = 0.84. Under no information (Case 1): producer surplus approximately 76% of S*, consumer surplus just under 10% of S*, and consumers of types (3/5, 2/5) and (2/5, 3/5) do not trade. Under full disclosure (Case 2): producer surplus approximately 81% of S*, consumer surplus 19% of S*, efficient. Under the producer-optimal structure (Case 3): producer surplus = 100% of S* (all surplus extracted), consumer surplus = 0%, efficient. Under the consumer-optimal structure (Case 4): producer surplus approximately 57% of S*, consumer surplus approximately 43% of S*, efficient. All cases except Case 1 are efficient; the no-information case excludes some consumers from trading.&lt;/p&gt;
&lt;p&gt;Q: Is the full-information disclosure structure consumer-optimal?&lt;/p&gt;
&lt;p&gt;A: Not in general. Proposition 3 states that full information is consumer-optimal if and only if all consumers in Ei have identical residual valuations (theta_i minus their second-best alternative) — a condition that generically fails. When residual valuations within Ei are heterogeneous, the designer can do strictly better for consumers by applying the extremal segmentation within each Ei rather than revealing full information, which would allow firms to price-discriminate on individual residual valuations and extract more surplus.&lt;/p&gt;
&lt;p&gt;Q: Can the designer trace out the entire efficient frontier between consumer- and producer-optimal outcomes?&lt;/p&gt;
&lt;p&gt;A: Yes, under two conditions. First, by mixing the consumer-optimal structure (point A) with the full-information structure (point B) using fractions lambda and 1-lambda respectively, the designer can implement any point on the efficient frontier between A and B. Second, when the producer-optimal outcome (point C) is also implementable, mixing the full-information structure with the producer-optimal structure by applying them to fractions lambda and 1-lambda of the consumer population respectively spans every point between B and C. The key insight is that the AIC condition, if it holds for f, also holds for any rescaled sub-distribution of f (it is scale-invariant), so the producer-optimal sub-problem remains feasible.&lt;/p&gt;
&lt;p&gt;Q: What are the regulatory implications of the analysis?&lt;/p&gt;
&lt;p&gt;A: The paper identifies a fundamental tension: banning information use sacrifices efficiency (some consumers excluded, wrong products purchased), but unrestricted use permits platforms to implement perfect collusion through information design. Critically, the paper shows that privacy-enhancing technologies that pool consumers into cohorts — like Google&amp;rsquo;s Privacy Sandbox — are equally consistent with the producer-optimal (collusive) and consumer-optimal (competitive) structures; the two differ only in the principle by which consumers are grouped. The paper suggests regulators could mandate that consumers in the same cohort share the same most-preferred product and that information be disclosed symmetrically across firms — the defining features of the consumer-optimal structure. This would block the producer-optimal grouping (which mixes consumers with different most-preferred products) while preserving efficiency.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to and extend Bergemann, Brooks, and Morris (2015)?&lt;/p&gt;
&lt;p&gt;A: Bergemann, Brooks, and Morris (2015) characterize achievable consumer and producer surplus outcomes when a designer discloses information to a single monopolist who can price-discriminate. The present paper extends this to oligopoly, where competition between firms creates both additional constraints (firms may undercut each other) and additional instruments (the designer can play firms against each other). The consumer-optimal construction directly applies the BBM (2015) extremal segmentation within each firm&amp;rsquo;s natural customer set Ei, but the outer layer — using public revelation of group membership to induce rival firms to price at zero — is new and arises specifically from the oligopoly setting.&lt;/p&gt;
&lt;p&gt;Information designer: An entity (modeled as a platform) that observes the full joint distribution of consumer valuations over all products and commits, before firms price, to a mapping from consumer types to joint distributions over messages sent to competing firms; the designer can be interpreted as an internet intermediary choosing how to package and share consumer data.&lt;/p&gt;
&lt;p&gt;Aggregate Incentive Compatibility (AIC): The necessary and sufficient condition on the distribution of consumer valuations for the existence of a producer-optimal information structure; for each firm i and each candidate deviation price p_hat_i, the aggregate infra-marginal losses firm i would incur on its natural customers by lowering price to p_hat_i must weakly exceed the maximum revenue firm i could gain by attracting consumers who prefer rival products but have valuation for i above p_hat_i.&lt;/p&gt;
&lt;p&gt;Producer-optimal information structure: An information structure that induces an equilibrium in which every consumer buys her most-preferred product at a price exactly equal to her full valuation for it, extracting 100% of available surplus as producer surplus — the outcome equivalent to the firms&amp;rsquo; fully collusive joint surplus maximum.&lt;/p&gt;
&lt;p&gt;Consumer-optimal information structure: An information structure that achieves the maximum consumer surplus attainable across all equilibria induced by any information structure, holding each firm to its guarantee profit Pi*_i (the best uniform-price profit the firm can secure by ignoring all signals) and allocating all residual surplus to consumers while maintaining allocative efficiency.&lt;/p&gt;
&lt;p&gt;Guarantee profit (Pi*&lt;em&gt;i): The maximum profit firm i can secure unilaterally by ignoring the designer&amp;rsquo;s signal and setting an optimal uniform price, computed against the worst case in which all rival firms price at zero; it equals max&lt;/em&gt;{pi} times the sum of f(theta) over all types in Ei for which theta_i minus pi exceeds all rival valuations.&lt;/p&gt;
&lt;p&gt;Polarization of preferences: A stochastic dominance condition under which, relative to a baseline distribution, the mass of consumers who prefer product i and have high valuations for it increases while the mass of consumers who prefer rival products but have high valuations for i decreases; higher polarization weakens the Firm IC constraints and makes the producer-optimal outcome easier to implement (Proposition 1).&lt;/p&gt;
&lt;p&gt;Separation and Consistency: Two structural properties any producer-optimal information structure must satisfy: Separation requires that the messages firm i sends to different consumers in Ei who have distinct valuations for i are disjoint in support; Consistency requires that every message firm i can send to any consumer type is contained in the union of messages firm i sends to consumers in Ei, preventing firm i from ever inferring that a consumer prefers a rival&amp;rsquo;s product.&lt;/p&gt;</description></item><item><title>Markups Across Space and Time</title><link>https://macropaperwarehouse.com/papers/markups-across-space-and-time/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/markups-across-space-and-time/</guid><description>&lt;p&gt;Anderson, Rebelo, and Wong study the behavior of markups in the retail sector across regions and over time, using a combination of firm-level Compustat data and product-level scanner data from two large retailers — one operating over 100 stores across U.S. states (quarterly data from 2006 Q1 to 2009 Q3, covering roughly 3.6 million SKU-store pairs across 79 product categories) and one operating hundreds of stores across Canadian provinces (quarterly data from 2016 Q1 to 2018 Q4, covering 15.6 million item-store pairs across 41 product groups). Markups are measured using gross margins — sales minus cost of goods sold as a fraction of sales — computed at the product level using the replacement cost for every item. This measurement approach is appropriate for retail because cost of goods sold accounts for over 80 percent of total retail firm costs, making it a reliable proxy for marginal cost. The replacement cost data, available at the store level, is the cost used by managers in actual pricing decisions, distinguishing these datasets from typical scanner data that contain only average costs.&lt;/p&gt;
&lt;p&gt;The paper documents five main facts. First, markups are remarkably stable over time and display a mild procyclical pattern. At the aggregate level, gross margins are roughly acyclical or mildly procyclical while sales and cost of goods sold are highly procyclical. The elasticity of gross margins with respect to real GDP is statistically insignificant at both the aggregate and firm level. The conditional response of gross margins to high-frequency monetary policy shocks and oil price shocks is also statistically insignificant, while net operating profit margins fall significantly in response to both shocks. Operating profit margins are 3.4 times more volatile than gross margins at a quarterly frequency, and sales and costs are roughly 2.6 times more volatile.&lt;/p&gt;
&lt;p&gt;Second, there is large regional dispersion in gross margins. A variance decomposition shows that the regional variance of gross margins (0.103) is substantially larger than the time-series variance (0.013), with a near-zero covariance between the two components. Third, regions with higher incomes and more expensive houses have higher markups — gross margins are positively correlated with log household income and log median house value in both the U.S. and Canadian data.&lt;/p&gt;
&lt;p&gt;Fourth, these higher regional markups do not result from less intense competition or regional differences in marginal costs. Gross margins are uncorrelated with the Herfindahl index (a measure of competition) and with a rural dummy (a proxy for higher transportation costs). The cyclicality of markups is acyclical or mildly procyclical regardless of whether the underlying product costs are themselves acyclical, procyclical, or countercyclical.&lt;/p&gt;
&lt;p&gt;Fifth, and most distinctively, regional variation in markups arises from differences in assortment composition across regions rather than from deviations from uniform pricing. A decomposition of regional gross margin variance confirms that the dominant component is the term capturing differences in product assortment across markets; the term capturing differences in gross margins for the same item — which would be nonzero under geographic price discrimination — accounts for very little of the regional variation. When the same item is available in different regions, the retailer charges a uniform price, consistent with Della Vigna and Gentzkow (2019).&lt;/p&gt;
&lt;p&gt;To rationalize these five facts, the authors propose a model with non-homothetic, quadratic preferences (following Melitz and Ottaviano 2008). In the model, higher-productivity regions choose higher-quality goods, which have less elastic demand and therefore higher markups. The markup is procyclical with respect to productivity shocks (A) but acyclical with respect to labor supply shocks (N), so a mixture of both types of shocks produces mildly procyclical markups. The model generates uniform pricing across regions for the homogeneous good, with regional markup differences arising through quality and assortment selection rather than price discrimination.&lt;/p&gt;
&lt;p&gt;Q: How do the authors measure markups, and why is this approach appropriate for retail?
A: Markups are measured as gross margins — (sales minus cost of goods sold) divided by sales — computed at the product level using the replacement cost for every item. This is appropriate for retail because cost of goods sold is the predominant variable cost, accounting for over 80 percent of total retail firm costs. The replacement cost is the marginal cost concept used by managers in pricing decisions and is available at the store level rather than as a national average.&lt;/p&gt;
&lt;p&gt;Q: What is the cyclical behavior of gross margins at the aggregate retail level?
A: Gross margins are roughly acyclical or mildly procyclical. Sales and cost of goods sold are highly procyclical, suggesting that the business cycle primarily affects quantities sold rather than markups. Operating profit margins are 3.4 times more volatile than gross margins at a quarterly frequency, while sales and costs are roughly 2.6 times more volatile.&lt;/p&gt;
&lt;p&gt;Q: What is the conditional response of gross margins to monetary policy and oil price shocks?
A: The response of gross margins to both high-frequency monetary policy shocks (identified from Federal Funds futures data) and oil price shocks (identified via the Ramey-Vine 2010 VAR approach) is statistically insignificant. In contrast, net operating profit margins fall in a statistically significant manner in response to both types of shocks, indicating that fixed cost absorption rather than markup adjustment drives profit volatility.&lt;/p&gt;
&lt;p&gt;Q: How large is the regional dispersion in gross margins relative to their time-series variation?
A: The variance decomposition shows that the regional variance of gross margins is 0.103, compared to a time-series variance of only 0.013, with a covariance term close to zero. The vast majority of gross margin variation is therefore cross-sectional rather than time-series.&lt;/p&gt;
&lt;p&gt;Q: What variables explain the regional variation in gross margins?
A: In the U.S. data, gross margins are positively correlated with log household income and log median house value. Gross margins are uncorrelated with the Herfindahl index (a competition measure) and with the rural county dummy (a transportation cost proxy). Canadian data confirms the positive correlation between gross margins and both log household income and log median house value.&lt;/p&gt;
&lt;p&gt;Q: What is the mechanism through which higher-income regions have higher markups?
A: Regional markup differences are driven by assortment composition differences, not price discrimination. When the same item is sold in multiple regions, it sells at a uniform price. Higher-income regions carry different (higher-quality, higher-margin) products. The correlation between unique items sold and regional household income is 0.42 for the Canadian retailer and 0.17 for the U.S. retailer.&lt;/p&gt;
&lt;p&gt;Q: How is the variance of regional gross margins decomposed into assortment versus pricing components?
A: The variance decomposition separates total regional gross margin variance into: (1) a term for differences in gross margins for the same item across regions (would be nonzero with geographic price discrimination), (2) a term for differences in assortment composition holding gross margins fixed, and (3) an interaction term plus covariance terms. The dominant term is the assortment composition component; the same-item price difference term accounts for very little of the regional variation.&lt;/p&gt;
&lt;p&gt;Q: Does the acyclicality of gross margins hold for products with procyclical costs?
A: Yes. The authors divide products into those with acyclical, procyclical, and countercyclical costs and show (Table 7) that gross margins are acyclical or mildly procyclical for all three groups in both the U.S. and Canadian data. This implies that retailer pricing behavior contributes to price inertia even for products whose wholesale costs move with the cycle.&lt;/p&gt;
&lt;p&gt;Q: What fraction of gross margin changes are active versus passive?
A: In the U.S. data, 91 percent of margin changes are active (resulting from price changes, regardless of whether replacement cost has changed); 9 percent are passive (replacement cost changes with no price change). In the Canadian data, 93 percent of changes are active. Both the probability of active margin changes and the size of margin changes are acyclical with respect to unemployment and local house prices.&lt;/p&gt;
&lt;p&gt;Q: How does the Hall approach compare to gross-margin-based markup estimates?
A: When the Hall approach is implemented using output elasticities (deflating sales by a product-level price deflator to obtain quantity), the resulting markup estimates are very close to those from gross margins — the ratio is 1.014 for the U.S. firm and 0.991 for the Canadian firm. However, when revenue elasticities are used instead of output elasticities (the common practice in the literature due to data limitations), the implied markup is 14 percent lower for the U.S. firm and 13 percent lower for the Canadian firm, confirming the bias documented by Bond et al. (2020).&lt;/p&gt;
&lt;p&gt;Q: What are the key features of the theoretical model and what facts does it explain?
A: The model uses non-homothetic quadratic preferences (Melitz-Ottaviano form) in which demand elasticity falls as consumption quality rises. Higher-productivity regions optimally consume higher-quality varieties, which face less elastic demand and hence carry higher markups. The markup is procyclical in productivity (A) with an elasticity less than one (incomplete cost passthrough) and acyclical in labor supply (N), so a mixture of shocks generates mild procyclicality. Uniform pricing across regions for the homogeneous good holds by construction, and regional markup differences arise through quality-assortment selection.&lt;/p&gt;
&lt;p&gt;Q: Which existing macroeconomic models are consistent with the time-series evidence, and which are not?
A: The evidence is inconsistent with models featuring countercyclical markups (Rotemberg-Woodford 1992 imperfect competition, Ravn-Schmitt-Grohe-Uribe deep habits, Jaimovich-Floetotto entry-exit, and standard New Keynesian models with sticky prices and procyclical marginal costs). The time-series evidence is consistent with models featuring sticky retail prices and acyclical marginal costs (Nakamura-Steinsson 2010, Coibion-Gorodnichenko-Hong 2015) and models with price and wage rigidities at the manufacturing level (Erceg-Henderson-Levin 2000, Christiano-Eichenbaum-Evans 2005). Mildly procyclical search models (Alessandria 2009) are also consistent when procyclicality is mild.&lt;/p&gt;
&lt;p&gt;Q: Which existing trade and regional models are consistent or inconsistent with the regional evidence?
A: The spatial price discrimination models of Greenhut-Greenhut (1975) and Thisse-Vives (1988), which predict higher markups in less competitive regions, are inconsistent with the data. The Bertoletti-Etro (2017) non-homothetic model predicts that regional markup variation is driven by deviations from uniform pricing, which is also inconsistent. The Fajgelbaum-Grossman-Helpman (2011) model predicts countercyclical markups when costs are procyclical, contradicting the time-series results. Most existing macroeconomic models rely on homothetic preferences, predicting markups independent of regional income, inconsistent with the regional facts.&lt;/p&gt;
&lt;p&gt;Q: What are the scope conditions on the measurement approach?
A: Gross margins are valid proxies for markups only in the retail sector, where cost of goods sold is the dominant variable cost (over 80 percent of total costs). In manufacturing, where labor and other costs represent a larger fraction of total variable costs, gross margins would not be a reliable markup measure. The product-level scanner data cover the 2006-2009 period for the U.S. and 2016-2018 for Canada; the U.S. sample includes a recession while the Canadian sample covers a moderate expansion.&lt;/p&gt;
&lt;p&gt;Gross margin as markup proxy: The ratio of (sales minus cost of goods sold) to sales, computed at the product level using the replacement cost for each item at each store and time period. Used as a proxy for the price-cost markup because cost of goods sold is the dominant variable cost in retail (over 80 percent of total costs), and the replacement cost is the marginal cost concept managers use in pricing decisions.&lt;/p&gt;
&lt;p&gt;Replacement cost: The cost at which the retailer would replenish a unit of inventory at current prices, available at the store level in the scanner datasets. Distinct from average historical cost and used here as a direct proxy for marginal cost, eliminating one of the main sources of markup mismeasurement in prior empirical work.&lt;/p&gt;
&lt;p&gt;Assortment composition: The set of products stocked and the expenditure weights of those products within a region. The paper&amp;rsquo;s central mechanism for regional markup variation — higher-income regions carry different (higher-quality, higher-margin) goods rather than charging different prices for the same goods.&lt;/p&gt;
&lt;p&gt;Uniform pricing: The practice of charging identical prices for the same item across different geographic regions. Confirmed empirically in both the U.S. and Canadian scanner datasets, and embedded structurally in the theoretical model for the homogeneous good.&lt;/p&gt;
&lt;p&gt;Active versus passive margin changes: A decomposition of gross margin changes into active changes (arising from retailer price decisions, irrespective of cost changes) and passive changes (arising when replacement cost changes but the retailer holds price fixed). Ninety-one percent of U.S. margin changes and 93 percent of Canadian changes are active.&lt;/p&gt;
&lt;p&gt;Non-homothetic quadratic preferences: The utility specification (following Melitz and Ottaviano 2008) in which the absolute value of the own-price demand elasticity falls as quality consumption rises. This property implies that higher-quality goods carry higher markups and that richer regions, which demand higher quality, have higher average markups — the key mechanism linking income to markups in the model.&lt;/p&gt;
&lt;p&gt;Hall approach to markup estimation: A production-function-based method in which the markup equals the output elasticity with respect to a variable input divided by that input&amp;rsquo;s cost share in revenue. The paper shows this yields estimates close to gross-margin estimates when implemented with true output quantities, but produces markups roughly 13-14 percent lower when revenue is substituted for output (a common approximation), confirming the Bond et al. 2020 bias.&lt;/p&gt;</description></item><item><title>Merger Effects and Antitrust Enforcement: Evidence from US Consumer Packaged Goods</title><link>https://macropaperwarehouse.com/papers/merger-effects-and-antitrust-enforcement-evidence-from-us-consumer-packaged-goods/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/merger-effects-and-antitrust-enforcement-evidence-from-us-consumer-packaged-goods/</guid><description>&lt;p&gt;This paper by Bhattacharya, Illanes, and Stillerman makes two contributions to the debate over US antitrust enforcement stringency. First, it documents the price, quantity, and assortment effects of a comprehensive set of consummated mergers in US consumer packaged goods (CPG). Second, it develops and estimates a model of agency enforcement decisions to quantify antitrust stringency and simulate counterfactual outcomes under stricter regimes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and scope.&lt;/strong&gt; The analysis covers 129 product markets across 47 transactions in US CPG from 2006 to 2017, using the NielsenIQ Retail Scanner Dataset (covering 35,000–50,000 stores and 2.6–4.5 million UPCs). The sample is restricted to all deals valued at $280 million or more where both the acquirer and target sold products in at least one overlapping product market-DMA. Geographic markets are NielsenIQ designated market areas (DMAs). The sample is defined to avoid selection bias from studying only mergers that attracted press attention or were litigation targets.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Identification strategy.&lt;/strong&gt; The empirical approach is a before-after event study within geography and product. For each merger, a brand-specific linear time trend is estimated from the 36 months prior to the merger announcement, controlling for UPC-DMA fixed effects, month-of-year fixed effects, input cost indices, and log median household income. Post-merger outcomes (24 months after completion) are measured as deviations from the extrapolated pre-merger trend. The identifying assumption is that secular demand and cost trends are gradual and well-captured by a linear trend. Pre-trend placebo tests show no significant departures from trend in the pre-period, and randomized-date placebos confirm that the linear trend is a better predictor of post-period outcomes under random merger dates than under actual merger dates, supporting the interpretation that observed post-period departures reflect merger effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Price effects.&lt;/strong&gt; The average price effect of consummated CPG mergers is small: across specifications, estimates range from -0.6% to 1.0%, with a baseline mean of 0.3%. However, heterogeneity is substantial. The standard deviation of merger-level price effects is 4.0–7.5 percentage points. In the baseline specification, the first quartile of price effects is -2.1% and the third quartile is 3.7%. Merging and non-merging party price changes are positively correlated (correlation = 0.49), consistent with strategic complementarity. Thirty-six percent of mergers lead both groups to lower prices; 36% lead both groups to raise prices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantity and assortment effects.&lt;/strong&gt; Total quantities fall on average by 0.4–1.0% across specifications, with 60% of mergers producing quantity reductions. Merging parties exhibit a larger average quantity decline of 6.4%. Mergers also lead to a 2.7% average reduction in the number of stores served by merging parties, a 2.2% reduction in the number of brands sold in a DMA by merging parties, and a 3.2% reduction for non-merging parties. Brands with less than 5% of the merged entity&amp;rsquo;s sales are 6 percentage points more likely to be dropped post-merger.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Enforcement model.&lt;/strong&gt; To interpret these outcomes relative to enforcement, the authors develop a model in which the agency receives a noisy signal of a merger&amp;rsquo;s price effect and challenges the merger if the posterior mean exceeds a threshold that is decreasing in deal size. They estimate the model by maximum likelihood using data on enforcement actions (6 mergers receiving remedies, 4 withdrawn under antitrust pressure) and realized price changes. The estimated sales-weighted average threshold is 4.8–6.3%: agencies act as if they challenge CPG mergers only when they expect a price increase exceeding this level. The posterior standard deviation of the agency&amp;rsquo;s assessment is 2.5–3.2 pp (aggregate prices) to 4.1–4.8 pp (merging-party prices).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Counterfactual stringency.&lt;/strong&gt; Tightening the threshold from approximately 6.1% to 2.5% would roughly quadruple the challenge probability (from 0.075 to 0.30), reduce aggregate price changes of consummated mergers by approximately 1.4 pp, and lower the share of allowed anti-competitive mergers from roughly 50% to 35%. Critically, type I errors (blocking pro-competitive mergers) remain negligible at thresholds down to approximately 3%; at 0% threshold only 10% of blocked mergers would be type I errors. The primary cost of tighter enforcement is a significantly larger agency workload, not an increase in blocked pro-competitive mergers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; Results pertain specifically to large CPG mergers (deal size ≥ $280 million) sold through US retail outlets, 2006–2017. Findings on structural presumptions show DHHI and merging share have predictive value for price changes, but structural metrics alone explain less than 10% of the variance in price effects (adjusted R-squared never exceeds 10% even with third-order interactions).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the average price effect of consummated CPG mergers and how should it be interpreted?&lt;/strong&gt;
A: Across specifications, the average price effect is between -0.6% and 1.0%, with a baseline mean of 0.3%. This small average does not imply that enforcement is strict: Carlton (2009) shows that with perfect foresight, the largest observed price change — not the average — would indicate stringency. Because agencies face uncertainty, the distribution of realized price changes reflects both inframarginal approved mergers and the noise in agency forecasts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How large is the heterogeneity in merger price effects?&lt;/strong&gt;
A: The standard deviation of merger-level price effects is 4.0–7.5 percentage points across specifications. In the baseline, the first quartile of price effects is -2.1% and the third quartile is 3.7% for all parties combined. Merging parties specifically show a first quartile of -3.2% and third quartile of 3.7%, meaning a full quarter of mergers raise merging-party prices by more than 3.7%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do merging and non-merging party prices co-move?&lt;/strong&gt;
A: Price changes for merging and non-merging parties are positively correlated (correlation = 0.49, s.e. = 0.08), consistent with strategic complementarity in pricing. Thirty-six percent of mergers lead both groups to lower prices, 36% lead both to raise prices, 13% cause merging parties to lower while non-merging parties raise, and 15% cause the reverse. The timing evidence shows merging-party prices begin changing upon merger completion, with rivals following suit.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What happens to quantities following mergers?&lt;/strong&gt;
A: Total quantities fall on average between 0.4% and 1.0% across specifications, with 60% of mergers producing quantity reductions. Merging parties bear the bulk of quantity adjustment, with an average quantity decline of 6.4% and a standard deviation and interquartile range both around 30 pp. Non-merging party quantity changes are much less variable. The correlation between merging and non-merging party quantity changes is 0.36 (s.e. 0.08), which is positive — at odds with theoretical predictions from demand systems with the &amp;ldquo;type aggregation property&amp;rdquo; (Nocke and Schutz, 2018, 2024), where mergers should produce negatively correlated quantity changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What non-price competitive responses do mergers trigger?&lt;/strong&gt;
A: Merging parties reduce the number of stores they serve by 2.7% on average, though in 38% of mergers store networks expand. Both merging and non-merging parties reduce product portfolios: merging parties drop the number of brands in a DMA by 2.2% on average and non-merging parties by 3.2%. Brands most likely to be dropped are those with less than 5% of the merged entity&amp;rsquo;s sales (6 pp more likely to be dropped), brands in small DMAs, and brands with small DMA shares.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Do the Merger Guidelines&amp;rsquo; structural presumptions (HHI, DHHI, merging share) predict price effects?&lt;/strong&gt;
A: DHHI and merging share have statistically significant but quantitatively modest predictive power. A 100-point increase in average DHHI is associated with a 0.2 pp increase in merging-party price changes and 0.3 pp for non-merging parties. Price effects are significantly larger when merging share exceeds 30%. However, structural metrics alone explain very little variance: adjusted R-squared never exceeds 10% even with third-order interactions of HHI, DHHI, merging share, private label share, and market size. Within-merger, DHHI is positively correlated with local price changes, and markets with DHHI above 200 exhibit significantly higher price effects than those below.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How do the authors model antitrust enforcement and identify its stringency?&lt;/strong&gt;
A: The agency observes a noisy signal of a merger&amp;rsquo;s price effect, forms a posterior distribution combining a normally distributed prior (mean X&amp;rsquo;beta, standard deviation sigma_p*) with a normally distributed signal error (standard deviation sigma_epsilon), and challenges the merger if the posterior mean exceeds a threshold that is decreasing in deal size. The model is estimated by maximum likelihood: for approved mergers, the realized price change is observed; for withdrawn/remedied mergers, the posterior mean must have exceeded the threshold. Six mergers (from four deals) received remedies for horizontal market power concerns and four mergers (from two deals) were withdrawn under antitrust pressure, forming the challenged set.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the estimated enforcement threshold and how does it vary across mergers?&lt;/strong&gt;
A: The sales-weighted average threshold is 4.8–6.3% using aggregate price changes and 6.6–7.8% using merging-party price changes. The threshold is lower for larger mergers: a 10% increase in merging-party sales is associated with an approximately 0.06 pp decrease in the threshold. The first quartile of thresholds across mergers is 4.5–5.6% and the third quartile is 5.6–6.9%, reflecting that the agencies apply stricter standards to larger deals.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How accurate are the agencies&amp;rsquo; forecasts of merger price effects?&lt;/strong&gt;
A: Using only the prior (structural characteristics), the agency&amp;rsquo;s accuracy in classifying mergers as anti-competitive versus pro-competitive is 56% (s.e. 3 pp). Adding the signal increases accuracy to 83% (s.e. 9 pp). The correlation between the prior mean and the true price change is 0.29 (s.e. 0.08); the correlation between the posterior mean and the true price change is 0.85 (s.e. 0.15). The posterior standard deviation is 2.5–3.2 pp for aggregate price changes and 4.1–4.8 pp for merging-party price changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What would happen under stricter antitrust enforcement?&lt;/strong&gt;
A: Tightening the average threshold from 6.1% to 2.5% would raise the challenge probability from approximately 0.075 to 0.30 — roughly quadrupling it — and would reduce aggregate price changes of consummated mergers by approximately 1.4 pp (from roughly 0.2% to -1.2%). Moving to a 0% threshold would result in challenges to 57% of mergers, with 60–70% of consummated mergers then causing price decreases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How large are type I and type II errors at the current and counterfactual thresholds?&lt;/strong&gt;
A: At the current threshold (~6.1%), approximately 50% of allowed mergers are type II errors (anti-competitive mergers that should have been challenged). Type I errors (pro-competitive mergers wrongly blocked) are negligible at the current threshold and only become non-trivial starting around a 3% threshold. At a 2.5% threshold, the type II error share falls to 35%; at a 0% threshold, to 16%, while type I errors reach 10% of blocked mergers. The primary trade-off of stricter enforcement is therefore a larger agency workload, not an increase in blocking pro-competitive mergers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What identification strategy is used and how is it validated?&lt;/strong&gt;
A: The strategy is a within-product, within-geography before-after comparison using a brand-specific linear pre-merger trend as the counterfactual. Validation proceeds through three checks: (1) coefficient plots from an extended event study show no significant pre-trends after controlling for the linear trend; (2) a plot of brand trends against estimated price effects shows little explanatory power (statistically significant negative correlation but small magnitude, not consistent with results being driven by trend extrapolation); (3) placebo tests randomizing merger dates within the same markets yield a distribution centered at zero, narrower than the true distribution, and a significantly higher mean squared prediction error in the post-period, confirming that the linear trend is a better predictor under randomly assigned merger dates than under true dates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why do the authors not use alternative control group approaches?&lt;/strong&gt;
A: Non-merging firms in the same market are rejected as controls because they may strategically respond to the merger. Synthetic controls using similar-industry untreated markets are rejected because deals often treat multiple similar markets (ruling out natural donors) and estimates prove sensitive to individual donors. Geographic controls (markets where merging parties have small shares) are rejected because they omit all 39 national mergers, untreated markets are not randomly selected, and regional pricing by non-merging parties could propagate effects into untreated regions, biasing estimates toward zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Merger retrospective.&lt;/strong&gt; In this paper&amp;rsquo;s usage, an ex-post empirical study of the price, quantity, and assortment effects of a consummated merger, using pre-merger trends as the counterfactual, as opposed to forward-looking merger simulation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Enforcement stringency.&lt;/strong&gt; The marginal price increase at which the antitrust agency would expect to challenge a merger. Measured here as the sales-weighted average posterior-mean threshold: the value above which the agency acts as if it would propose a remedy, estimated at 4.8–6.3% for US CPG mergers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Type I error (antitrust).&lt;/strong&gt; The mistake of challenging (blocking) a merger that would have reduced prices (a pro-competitive merger). In the model, this occurs when an adverse signal causes the agency to block a merger whose true price effect is below the threshold.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Type II error (antitrust).&lt;/strong&gt; The mistake of allowing a merger that increases prices (an anti-competitive merger). In the model, this occurs when a favorable signal causes the agency to approve a merger whose true price effect is above the threshold. Estimated at approximately 50% of allowed mergers at the current enforcement threshold.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Structural presumptions.&lt;/strong&gt; The HHI-based rules in the 2010 and 2023 Merger Guidelines that create a presumption of competitive harm when DHHI exceeds specified thresholds (e.g., DHHI &amp;gt; 200 and post-merger HHI &amp;gt; 2,500 for the &amp;ldquo;red zone&amp;rdquo;). The paper finds DHHI and merging share have statistically significant but low explanatory power (adjusted R-squared below 10%) for actual price changes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Prior and signal (in the enforcement model).&lt;/strong&gt; The agency&amp;rsquo;s prior is a normal distribution over the merger&amp;rsquo;s true price effect, parameterized by structural characteristics (HHI, DHHI). The signal is a noisy draw centered on the true price effect, capturing information gathered through due diligence (e.g., evidence of efficiencies). The posterior mean — combining prior and signal — determines whether the agency challenges the merger.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Product market-deal pair (merger).&lt;/strong&gt; The unit of observation in the empirical analysis: a specific NielsenIQ product module (e.g., soluble coffee) within a specific acquisition transaction (e.g., a food conglomerate merger). The sample contains 129 such pairs across 47 deals.&lt;/p&gt;</description></item><item><title>Normal Approximation in Large Network Models</title><link>https://macropaperwarehouse.com/papers/normal-approximation-in-large-network-models/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/normal-approximation-in-large-network-models/</guid><description>&lt;p&gt;This paper proves a central limit theorem (CLT) for network formation models with strategic interactions and homophilous agents, addressing a foundational inferential gap in the econometrics of large networks. The setting is one where the econometrician observes a single large network — the asymptotic framework sends network size n to infinity — which is the empirically relevant case for most network datasets. The network moments of interest are averages of node-level statistics (1/n) Σ ψ_i, where ψ_i can capture degree, clustering coefficients, or subnetwork counts (triangles, k-stars) that have been used for structural inference in network formation games.&lt;/p&gt;
&lt;p&gt;The model is a pairwise-stability network formation game augmented onto a latent-space/geometric-graph structure. Each node i has an i.i.d. type (X_i, Z_i), where X_i is a continuously distributed position vector capturing homophilous attributes. Two nodes i and j form a link if a joint-surplus function V(·) exceeds zero, where V depends on the scaled distance r_n^{-1}‖X_i − X_j‖ between positions, a vector of strategic interaction statistics S_{ij} (functions of neighboring links), node attributes Z_i, Z_j, and an i.i.d. utility shock ζ_{ij}. Homophily enters as a monotonicity requirement: V is decreasing in the distance component, so dissimilar nodes are less likely to link. Sparsity is ensured by setting r_n = (κ/n)^{1/d}, which keeps expected degree asymptotically bounded.&lt;/p&gt;
&lt;p&gt;Strategic interactions enter through S_{ij}, which depends on links involving neighbors of i or j (local externalities), generating chains of cross-sectional dependence that are the central obstacle to the CLT. The paper identifies two distinct sources of dependence: (1) link interdependencies from best-response chains, where the realization of one link influences neighboring links; and (2) global coordination in equilibrium selection, where agents may condition on a common signal.&lt;/p&gt;
&lt;p&gt;The main technical contribution is adapting &amp;ldquo;stabilization&amp;rdquo; conditions from the literature on geometric graphs (Penrose and Yukich 2003, 2008) to the strategic setting. Exponential stabilization (Assumption 5) requires that the radius of stabilization R_i — the smallest neighborhood of i such that ψ_i depends only on nodes within that neighborhood — has a distribution with exponential tails. This bounds the effective dependence neighborhood and provides the weak dependence structure needed for the CLT.&lt;/p&gt;
&lt;p&gt;To verify stabilization from primitive conditions, the paper employs branching process theory. The key construct is the &amp;ldquo;strategic neighborhood&amp;rdquo; C_i^+, the component of i in the network of non-robust links D (pairs where strategic interactions can change the link outcome). The paper bounds |C_i^+| by a subcritical Galton-Watson branching process: if the mean offspring is below 1 (subcriticality, Assumption 7, stated as ‖h*‖_m &amp;lt; 1), the process is non-explosive and its size has exponential tails, yielding the required stabilization. The subcriticality condition directly restricts the strength of strategic interactions and is the network analog of the condition ‖β‖ &amp;lt; 1 in linear autoregressive models. A second condition (Assumption 8, decentralized selection) requires that equilibrium selection operates independently across disjoint strategic neighborhoods, ruling out global coordination; this holds under myopic best-response dynamics.&lt;/p&gt;
&lt;p&gt;For inference, the paper proposes a network HAC variance estimator hat_Σ_n = (1/n) Σ_i Σ_j k(d_{ij}/b_n) hat_ψ_i hat_ψ_j^T, where k(·) is a kernel, d_{ij} is the path distance in A, and b_n is a bandwidth, and a network bootstrap that resamples nodes with replacement. Both are shown to be consistent (Theorem 3). Simulation results with n up to 500, varying strategic interaction strength θ_2 from 0 to 0.5, show that the network HAC estimator achieves nominal 5% rejection rates and 95% coverage for n ≥ 500, while the bootstrap slightly over-rejects in small samples and performance degrades as θ_2 increases.&lt;/p&gt;
&lt;p&gt;The scope conditions are explicit: the CLT applies to sparse networks (expected degree bounded), undirected networks with local externalities, models admitting a pairwise-stability equilibrium, and equilibrium selection satisfying decentralization. Extensions to directed or denser networks are left for future work.&lt;/p&gt;
&lt;p&gt;Q: What is the primary research question and why does it require new theory?
A: The paper asks when sample averages of network statistics — degree, clustering, subnetwork counts — satisfy a CLT in strategic network formation models observed as a single large network. Standard CLT proofs require weakly dependent observations, but strategic interactions generate chains of link dependence of a priori unbounded length, and multiple equilibria allow global coordination, both of which can destroy asymptotic normality. Prior work (Leung 2019b; Menzel 2024) established laws of large numbers but not CLTs, which require stronger conditions.&lt;/p&gt;
&lt;p&gt;Q: What is the stabilization condition and why is it the right formulation of weak dependence?
A: Exponential stabilization (Assumption 5) requires that the radius of stabilization R_i — the smallest K such that ψ_i depends only on the K-neighborhood of i in the network — has a distribution with exponential tails: lim sup_{w→∞} w^{-η} max{log τ_{b,ε}(w), log τ_p(w)} &amp;lt; 0 for some η ∈ (0,1]. This implies that each node&amp;rsquo;s statistic depends effectively only on a bounded fraction of the network, making {ψ_i} weakly dependent. The condition is a modification of stabilization conditions from the geometric graph literature (Penrose and Yukich 2003, 2008) adapted to allow strategic interactions.&lt;/p&gt;
&lt;p&gt;Q: How does the paper connect the abstract stabilization condition to primitive model conditions?
A: The paper defines the strategic neighborhood C_i^+ as the union of one-step network neighborhoods of nodes in i&amp;rsquo;s component in the non-robust link network D (where D_{ij} = 1 iff the link A_{ij} can be switched by strategic interactions). The size |C_i^+| controls the radius of stabilization. By mapping exploration of C_i via breadth-first search onto a Galton-Watson branching process, subcriticality (mean offspring &amp;lt; 1, i.e., ‖h*‖_m &amp;lt; 1) implies that |C_i^+| has exponential tails, which yields exponential stabilization with η = 1 (Theorem 2).&lt;/p&gt;
&lt;p&gt;Q: What is the subcriticality condition and what does it restrict?
A: Subcriticality (Assumption 7) requires that the mean interaction-strength measure satisfies ‖h*‖_m &amp;lt; 1, where h* bounds the probability that a given link is non-robust as a function of node attributes. This restricts how strongly the existence of one link influences the probability of neighboring links. The authors explicitly analogize this to the condition ‖β‖ &amp;lt; 1 in linear autoregressive models: both bound the magnitude of &amp;ldquo;autoregressive&amp;rdquo; dependence below one to prevent explosive propagation of dependence.&lt;/p&gt;
&lt;p&gt;Q: What is the decentralized selection condition and what does it rule out?
A: Assumption 8 (decentralized selection) requires that the equilibrium selection mechanism operates independently across disjoint strategic neighborhoods: A_{H_l} = λ_{|H_l|}(r^{-1}T_{H_l}, ζ_{H_l}) for each disjoint strategic neighborhood H_l. This rules out global coordination where agents condition on a common signal (such as the type of a particular node) to jointly select an equilibrium. The condition is satisfied by myopic best-response dynamics and is described as the single-network analog of requiring equilibrium selection to be independent across networks under many-network asymptotics.&lt;/p&gt;
&lt;p&gt;Q: What is the structure of the CLT proof?
A: The proof has two steps. Step 1 proves a CLT for the Poissonized model where the number of nodes N_n ~ Poisson(n), leveraging results from Penrose and Yukich (2008) for geometric graphs extended to the strategic setting. Step 2 is a de-Poissonization argument that transfers the Poissonized CLT back to the fixed-n model. The abstract CLT (Theorem 1) requires Assumptions 5 and 6, and Theorem 2 establishes that Assumptions 1–8 imply Assumption 5 with η = 1.&lt;/p&gt;
&lt;p&gt;Q: How does the network HAC estimator work and what are its consistency conditions?
A: The estimator is hat_Σ_n = (1/n) Σ_i Σ_j k(d_{ij}/b_n) hat_ψ_i hat_ψ_j^T, where d_{ij} is the path distance between i and j in the observed network A, k(·) is a kernel function, b_n is a bandwidth, and hat_ψ_i = ψ_i(N_n) − (1/n) Σ_j ψ_j(N_n) is the demeaned statistic. Consistency (hat_Σ_n →^p Σ_n) is established under appropriate conditions on the bandwidth b_n (Theorem 3). The bandwidth plays the same role as in time-series HAC estimation, controlling the window over which covariances are summed.&lt;/p&gt;
&lt;p&gt;Q: What do the simulations show about finite-sample performance?
A: Using a DGP with X_i ~ U([0,1]^2), ζ_{ij} ~ N(0,1), and θ_2 varying from 0 to 0.5 to control strategic interaction strength, the network HAC estimator achieves nominal 5% rejection rates and 95% coverage at n ≥ 500 across all settings. The bootstrap slightly over-rejects in small samples. Performance of all procedures degrades as θ_2 increases (stronger strategic interactions), consistent with the theoretical condition that subcriticality must hold. These results support practical use of the inference procedures based on Theorem 1.&lt;/p&gt;
&lt;p&gt;Q: How does this paper relate to prior work on CLTs for network data?
A: Kojevnikov et al. (2021) prove a CLT for node-level data conditional on the network, but this does not apply to network formation because the network is the outcome, not a conditioning variable. Leung (2019b) and Menzel (2024) prove laws of large numbers for strategic network formation but not CLTs. Kuersteiner (2019) takes a different approach using a conditional mixingale assumption. The paper&amp;rsquo;s abstract CLT extends Penrose and Yukich (2008) by modifying the stabilization condition to accommodate strategic interactions; the primitive conditions are new and use branching process tools that build on Leung (2019b).&lt;/p&gt;
&lt;p&gt;Q: What network moments can the CLT be applied to?
A: The CLT applies to any average of node statistics ψ_i that depends only on the K-neighborhood of i in the network (Assumption 4 with finite K). Explicit examples include average degree (ψ_i = Σ_j A_{ij}), average clustering coefficient, and counts of connected subnetworks such as triangles and k-stars. Subnetwork counts have been used as the basis for structural identification and estimation of network formation games (Sheng 2020), making the CLT directly applicable to inference in those models.&lt;/p&gt;
&lt;p&gt;Q: What are the scope limitations and directions for future work?
A: The CLT applies to sparse undirected networks with local externalities (Assumption 2), homophily in positions (Assumption 1), and equilibrium selection satisfying decentralization (Assumption 8). It does not cover directed networks, denser networks where expected degree grows with n, or models with global link externalities. The authors identify extending results to directed and denser networks and developing more powerful inference procedures exploiting network structure as priorities for future work.&lt;/p&gt;
&lt;p&gt;Stabilization (exponential): The condition that the radius of stabilization R_i — the smallest neighborhood of i beyond which ψ_i does not depend on further nodes — has a distribution with exponential tails (lim sup_{w→∞} w^{-η} log τ(w) &amp;lt; 0 for η ∈ (0,1]). This is the paper&amp;rsquo;s operative formulation of weak dependence for network statistics and is adapted from geometric graph theory to the strategic setting.&lt;/p&gt;
&lt;p&gt;Strategic neighborhood (C_i^+): The union of one-step neighborhoods of nodes in i&amp;rsquo;s component in the non-robust link network D. A link (i,j) is non-robust (D_{ij} = 1) if strategic interactions can change its realization — i.e., the surplus V can be positive under some interaction configurations and non-positive under others. The size of C_i^+ governs the radius of stabilization and hence the degree of cross-sectional dependence.&lt;/p&gt;
&lt;p&gt;Subcriticality (‖h*‖_m &amp;lt; 1): The condition that the mean-field interaction strength measure satisfies ‖h*‖_m &amp;lt; 1, where h* bounds the conditional probability that a link is non-robust. Subcriticality ensures that breadth-first search of the strategic neighborhood is dominated by a subcritical Galton-Watson process (mean offspring &amp;lt; 1), preventing explosive growth of the dependence neighborhood. The paper explicitly frames this as the network analog of ‖β‖ &amp;lt; 1 in autoregressive models.&lt;/p&gt;
&lt;p&gt;Decentralized selection (Assumption 8): The requirement that the equilibrium selection mechanism assigns outcomes independently across disjoint strategic neighborhoods: A_{H_l} = λ_{|H_l|}(r^{-1}T_{H_l}, ζ_{H_l}) for each disjoint H_l. This rules out global coordination — agents conditioning on a common signal to select among equilibria — while permitting local coordination within strategic neighborhoods. Satisfied by myopic best-response dynamics.&lt;/p&gt;
&lt;p&gt;Pairwise stability: The solution concept underlying the model. A network A satisfies pairwise stability under transferable utility if A_{ij} = 1{V_{ij} &amp;gt; 0}, meaning a link forms exactly when the joint surplus is positive. This is the equilibrium condition from which the strategic interaction statistics S_{ij} and non-robustness indicators D_{ij} are derived.&lt;/p&gt;
&lt;p&gt;Network HAC estimator: The variance estimator hat_Σ_n = (1/n) Σ_i Σ_j k(d_{ij}/b_n) hat_ψ_i hat_ψ_j^T, where d_{ij} is the path distance in the observed network, k(·) is a kernel, and b_n is a bandwidth. It is the network analog of heteroskedasticity- and autocorrelation-consistent (HAC) estimators in time series, using path distance in place of temporal lag distance.&lt;/p&gt;
&lt;p&gt;Homophily (in this paper&amp;rsquo;s sense): The property that the joint-surplus function V is decreasing in the first argument r_n^{-1}‖X_i − X_j‖ (scaled positional distance), so nodes that are more dissimilar in position are strictly less likely to form links. Combined with the sparsity scaling r_n = (κ/n)^{1/d}, this ensures that links decay with distance in social space and that the network remains sparse as n grows.&lt;/p&gt;</description></item><item><title>Online Business Models, Digital Ads, and User Welfare</title><link>https://macropaperwarehouse.com/papers/online-business-models-digital-ads-and-user-welfare/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/online-business-models-digital-ads-and-user-welfare/</guid><description>&lt;p&gt;Acemoglu, Huttenlocher, Ozdaglar, and Siderius develop a two-sided platform model to study the welfare consequences of digital advertising as an online business model. The platform intermediates between a firm selling a horizontally differentiated product and a continuum of users who derive utility from both entertaining content and informative signals about product quality embedded in ads. Users have a two-dimensional type: a sophistication dimension (sophisticated with probability lambda, naïve with probability 1-lambda) and a product-quality dimension (high quality with prior probability q). The central departure from the standard informational-advertising literature is that sophisticated users hold the correct model of the ad signal process, while naïve users underestimate the false-positive rate — the probability that a low-quality product generates a positive ad signal (phi_0). Naïve users perceive this false-positive rate to be phi_{0,N} = omega_N * omega_P * phi_0, where omega_N &amp;lt;= 1 captures inherent naïveté and omega_P &amp;lt;= 1 captures failure to understand personalized targeting, so phi_{0,N} &amp;lt; phi_0. The equilibrium concept is Berk-Nash equilibrium (Esponda and Pouzo 2016), meaning all agents are Bayesian given their subjective model.&lt;/p&gt;
&lt;p&gt;The platform chooses ad load alpha (Poisson rate of ad displays), subscription fees, and the monetary transfer from the firm; the firm sets product price p after observing the platform&amp;rsquo;s contract. The central finding (Proposition 2) is that when the objective false-positive rate phi_0 exceeds a threshold phi-hat_0(lambda, phi_1, phi_{0,N}) — which is increasing in lambda and phi_{0,N} and decreasing in the true-positive rate phi_1 — the unique equilibrium is an advertising-based plan that fully segments the market: naïve users receive an ad load that extracts all their surplus, while sophisticated users are excluded entirely. In this regime the firm charges a strictly higher price p-hat* &amp;gt; p-bar*, where p-bar* = (beta*q + c)/2 is the monopoly price without advertising. The ad-based equilibrium emerges precisely when ads are more misleading (larger gap between phi_0 and phi_{0,N}), not when they are more informative — a comparative static the authors describe as paradoxical.&lt;/p&gt;
&lt;p&gt;Welfare consequences (Proposition 4) are unambiguous in the advertising regime: both naïve and sophisticated users are strictly worse off than the baseline without any platform. Naïve users over-purchase due to inflated posteriors from misread signals; sophisticated users are harmed through the price channel — the firm&amp;rsquo;s higher profit-maximizing price p-hat* applies to all buyers. In the fully rational benchmark (phi_{0,N} = phi_0), the unique equilibrium is subscription-based and user welfare equals the no-platform baseline (Proposition 3).&lt;/p&gt;
&lt;p&gt;These results extend to richer menus (Proposition 5), mixed subscription-plus-advertising plans (Proposition 7), and to multi-firm and multi-platform competition (Propositions 9-12). Digital ads soften Bertrand competition by generating endogenous horizontal differentiation among otherwise identical firms, so equilibrium prices can exceed marginal cost even with two competing firms. Platform competition similarly fails to restore welfare: platforms compete away subscription fees but both adopt ad-based plans targeting naïfs when phi_1 exceeds a threshold, maintaining the welfare loss.&lt;/p&gt;
&lt;p&gt;On policy, the first best (planner observes types) cannot be decentralized because naïve users prefer more ads than is socially optimal, inverting the usual self-selection constraint. The second best (planner subject to incentive-compatibility constraints) is a single pooling plan with an intermediate ad load alpha^{SB} in [alpha^{FB}_N, alpha^{FB}_S] and yields average welfare above the no-platform baseline, though below first best (Proposition 13). This second best can be decentralized with a nonlinear digital ad tax, a per-unit product subsidy, and a platform subscription subsidy (Proposition 14). A simpler flat tax on digital ad revenues — above a threshold gamma-bar &amp;lt; 1 — also improves welfare relative to the ad-based equilibrium, though it does not restore the second best (Proposition 15).&lt;/p&gt;
&lt;p&gt;Four robustness extensions are developed: endogenous manipulation (platform always chooses the most manipulative environment, lowest phi_{0,N}); naïve learning dynamics (learning raises the sophisticate share in steady state, making ad-based models less profitable but not overturning the main results); imperfect price discrimination by the firm (naïfs are unambiguously worse off, threshold for advertising equilibrium shifts down); and an added price-sensitivity dimension (the platform runs a 2x2 menu separating by both sophistication and price sensitivity, preserving the result that naïve users tolerate and receive more ads than sophisticates in every stratum).&lt;/p&gt;
&lt;p&gt;Q: What is the key asymmetry between naïve and sophisticated users that drives the main results?
A: Sophisticated users hold the correct Bayesian model of the ad signal process and thus correctly account for the false-positive rate phi_0 when updating beliefs from positive ad signals. Naïve users perceive the false-positive rate as phi_{0,N} = omega_N * omega_P * phi_0 &amp;lt; phi_0, so they treat positive signals as stronger evidence of high product quality than they actually are. Because naïve users overestimate the informativeness of ads, their (interim) subjective valuation of an ad-based plan is higher, making them more tolerant of ad loads and more willing to join platforms with heavy advertising. This asymmetry is what makes it profitable to target naïfs with high ad loads while excluding or charging subscription fees to sophisticates.&lt;/p&gt;
&lt;p&gt;Q: Why does advertising to sophisticated users generate no additional firm profit, while advertising to naïve users does?
A: Lemma 1 establishes that with linear-quadratic utility the firm extracts no surplus from advertising to sophisticates: because sophisticated agents are fully Bayesian, their expected posterior equals the prior (E_S[pi_i] = q), so expected demand after advertising is identical to demand before advertising. By contrast, Lemma 2 shows that the firm&amp;rsquo;s profit from naïve agents is positive and strictly increasing in ad load alpha, because naïve users&amp;rsquo; average demand curve drifts upward as alpha rises — their inflated perceived informativeness of ads causes them to over-update on positive signals, systematically raising their willingness to pay. The platform captures this surplus from the firm via the advertising transfer m*.&lt;/p&gt;
&lt;p&gt;Q: What is the threshold condition determining whether the equilibrium is subscription-based or advertising-based?
A: Proposition 2 identifies a threshold phi-hat_0(lambda, phi_1, phi_{0,N}) that is increasing in the sophisticate share lambda and in the naïve false-positive perception phi_{0,N}, and decreasing in the true-positive rate phi_1. When the objective false-positive rate phi_0 is below this threshold, the profit-maximizing business model is subscription-based with price P* = T - v and product price p* = p-bar* = (beta&lt;em&gt;q + c)/2. When phi_0 exceeds the threshold, the advertising model dominates: the platform sets a high ad load alpha-hat&lt;/em&gt; that makes naïve users exactly indifferent between participating and their outside option v, excludes sophisticates, and the firm charges p-hat* &amp;gt; p-bar*. The threshold falls with phi_1, meaning more informative ads expand the range of phi_0 over which the advertising equilibrium obtains.&lt;/p&gt;
&lt;p&gt;Q: How does allowing the platform to offer menus change the results relative to the baseline two-plan case?
A: Proposition 5 shows that with menus the platform can simultaneously serve both user types: sophisticates receive a subscription plan at P* = T - v and naïve users receive an ad-based plan with the same high load alpha-hat* as in the baseline. The threshold for the advertising equilibrium shifts down to phi*&lt;em&gt;0(lambda, phi_1, phi&lt;/em&gt;{0,N}) &amp;lt; phi-hat_0, so advertising business models arise for a strictly larger set of parameters. Welfare consequences are unchanged (Corollary 1): when phi_0 &amp;gt; phi*_0, both types have welfare strictly below the no-platform baseline. Proposition 6 further shows consumer welfare is monotonically decreasing in both phi_0 and phi_1: higher phi_1 (more informative true-positive signals) also reduces welfare because any surplus from greater informativeness is fully captured by the platform.&lt;/p&gt;
&lt;p&gt;Q: What is the welfare ranking across the three regimes: no platform, advertising equilibrium, and subscription equilibrium?
A: In the subscription equilibrium (regime (a) of Proposition 2 or 4), user welfare for both types equals the no-platform base case W_base(tau) — the platform captures all surplus it creates and users are no better or worse off. In the advertising equilibrium (regime (b)), both naïve and sophisticated users are strictly worse off than with no platform: W-hat*(tau) &amp;lt; W_base(tau) for both tau in {S, N}. The first-best, where a planner controls ad loads separately by type, yields W^{FB}(tau) &amp;gt; W_base(tau) for both types because informative ads can genuinely improve sophisticated users&amp;rsquo; decisions and a constrained amount improves naïve users&amp;rsquo; decisions too.&lt;/p&gt;
&lt;p&gt;Q: How does firm-level competition interact with digital advertising to affect prices and welfare?
A: Without advertising, two ex ante identical firms compete à la Bertrand and price at marginal cost (p*_1 = p*_2 = c). Proposition 9 establishes that when phi_1 &amp;gt; phi^F_1 and phi_0 &amp;gt;= phi^F_0(phi_1), the platform offers an ad-based plan and equilibrium prices p-hat*_1 and p-hat*_2 are both strictly above p-bar* — the monopoly price without advertising. The mechanism is endogenous horizontal differentiation: users who see positive ad signals for one firm&amp;rsquo;s product form higher valuations for that product, so the two products become differentiated in the eyes of consumers even though they are ex ante identical, breaking Bertrand logic. Example 1 further illustrates that advertising can be more prevalent with competition than without: a second firm&amp;rsquo;s entry can push the equilibrium from no-advertising to separating.&lt;/p&gt;
&lt;p&gt;Q: Does platform competition protect users from the welfare losses associated with digital advertising?
A: Not fully. Proposition 11 shows that with two competing platforms (M=2, N=1) and no advertising, platforms compete away both subscription fees and ad loads, and welfare reaches the fully rational benchmark. However, when phi_1 exceeds threshold phi^P_1, both platforms adopt ad-based plans targeting naïve users, charge no subscription fees, and the product price rises to p-hat*_P &amp;gt; p-bar* (Proposition 12). Competition reduces subscription fees to zero but does not eliminate the incentive to target naïfs with heavy ads, because naïve users&amp;rsquo; over-valuation of ads means they remain willing to join ad-heavy plans. The fundamental inefficiency from naïve users&amp;rsquo; misspecified model persists under platform competition.&lt;/p&gt;
&lt;p&gt;Q: Why is the first-best allocation not implementable as a decentralized equilibrium?
A: Proposition 13 explains the obstacle: the social planner would ideally offer naïve users fewer ads (alpha^{FB}_N) than sophisticated users (alpha^{FB}_S), with alpha^{FB}_N &amp;lt;= alpha^{FB}_S. However, naïve users have a higher subjective valuation for ads than sophisticates because they believe ads are more informative. If offered a menu with both options, naïve users would self-select into the plan with the higher ad load alpha^{FB}_S — the exact opposite of what the planner wants. The incentive-compatibility constraints therefore force the planner toward a single pooling plan with an intermediate ad load alpha^{SB} in [alpha^{FB}_N, alpha^{FB}_S]. Average welfare under the second best exceeds the no-platform baseline, confirming that some advertising is socially valuable, but falls short of the first best whenever alpha^{FB}_N &amp;gt; 0.&lt;/p&gt;
&lt;p&gt;Q: How does a flat digital ad tax improve welfare, and what are its limitations?
A: Proposition 15 establishes that whenever the equilibrium features an ad-based plan, a flat tax on digital ad revenues at rate gamma &amp;gt; gamma-bar &amp;lt; 1 improves welfare by discouraging advertising-based business models and inducing the platform to shift toward subscription-based plans. The mechanism is that taxing ad revenue reduces the platform&amp;rsquo;s marginal gain from increasing ad load, making the subscription plan relatively more profitable. However, the flat tax does not achieve the second best because it operates linearly rather than targeting the nonlinear distortion: the optimal nonlinear tax-subsidy scheme (Proposition 14) requires a threshold-style ad tax at rate mu &amp;gt; mu-bar combined with a per-unit product subsidy delta* and a platform subscription subsidy eta &amp;gt; eta-bar.&lt;/p&gt;
&lt;p&gt;Q: What happens when the platform can endogenously choose how manipulative its ads are?
A: Proposition 16 shows that a profit-maximizing platform always chooses the lowest feasible phi_{0,N} = phi-bar — the most manipulative environment. Two reinforcing channels drive this: the pricing channel (lower phi_{0,N} amplifies naïve demand shifts per positive signal, so the downstream firm raises price and sales, increasing ad revenues extracted by the platform) and the participation channel (lower phi_{0,N} raises naïve users&amp;rsquo; perceived informational value of ads, relaxing their participation constraint and permitting a higher ad load alpha). Platform competition constrains the equilibrium ad load through tighter participation constraints but does not alter the choice of phi_{0,N} = phi-bar, so competition limits ad quantity but not ad manipulativeness.&lt;/p&gt;
&lt;p&gt;Q: How do naïve learning dynamics affect the main results?
A: Proposition 17 introduces a birth-death environment where exposure to disconfirming evidence gradually converts naïve agents to sophisticates. A unique steady-state sophisticate share lambda*(alpha_N, phi_0) exists; both higher ad load alpha_N and higher phi_0 accelerate the conversion of naïfs, raising future sophisticate share and reducing future ad revenues. This creates a new intertemporal trade-off that constrains the platform&amp;rsquo;s choice of ad loads relative to the static case. The key result (part ii) is that the main characterization of Proposition 7 carries through under a modified cutoff phi-tilde^{dynamic}&lt;em&gt;0 &amp;gt;= phi-tilde_0(lambda-tilde, phi_1, phi&lt;/em&gt;{0,N}), so learning dynamics make the ad-based business model less likely but do not overturn the fundamental welfare results.&lt;/p&gt;
&lt;p&gt;Q: How does imperfect price discrimination by the firm affect naïve users?
A: Proposition 18 considers a firm that observes a user&amp;rsquo;s sophistication type with probability kappa in [0,1]. With price discrimination, the firm sets type-specific prices satisfying p*_N &amp;gt;= p* &amp;gt;= p*_S, moving toward the type-specific monopoly levels. Naïfs are unambiguously worse off: when identified (with probability kappa), they face the higher price p*_N and a higher equilibrium ad load. The threshold for the advertising equilibrium also shifts down relative to the baseline, meaning advertising business models emerge for a larger parameter range when price discrimination is possible.&lt;/p&gt;
&lt;p&gt;Q: How does the paper define and measure user welfare, and why is ex post rather than interim welfare the relevant concept?
A: User welfare W(tau_i) is defined as ex post utility, which depends on the actual product quality theta_i realized after consumption, not on interim beliefs formed after viewing ads. Naïve users&amp;rsquo; interim assessment inflates expected product quality, but their ex post utility depends on whether the product is genuinely high quality for them (theta_i = 1 with probability q, theta_i = 0 with probability 1-q). Because naïve users over-purchase due to misread signals — consuming more than optimal when theta_i = 0 — their ex post utility is strictly lower than their interim expected utility, and strictly lower than the no-platform baseline in the advertising equilibrium. The ex post welfare concept is the relevant one precisely because it captures the actual material consequences of manipulation, not the subjectively perceived gains from ads.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Naïve vs. Sophisticated Users&lt;/strong&gt;: The paper&amp;rsquo;s primary user heterogeneity dimension. Sophisticated users hold the correct model of the ad signal process, setting phi_{0,S} = phi_0 (the true false-positive rate). Naïve users hold a misspecified model with phi_{0,N} = omega_N * omega_P * phi_0 &amp;lt; phi_0, underestimating the probability that a low-quality product generates a positive ad signal, due to inherent naïveté (omega_N) and failure to understand personalized targeting (omega_P).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Ad Load (alpha)&lt;/strong&gt;: The Poisson rate at which ads are displayed to a user per unit time. Total ad displays follow a Poisson(alpha*T) distribution. Higher ad load means less time on entertaining content — expected entertainment time is (1-alpha)&lt;em&gt;T — and a higher probability (1 - exp(-alpha&lt;/em&gt;T)) that the user sees the ad at least once. The platform chooses alpha as its primary instrument for extracting surplus from naïve users.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;False-Positive Rate (phi_0)&lt;/strong&gt;: The objective probability that a low-quality product (theta_i = 0) generates a positive (&amp;ldquo;good&amp;rdquo;) ad signal. The gap between phi_0 (objective) and phi_{0,N} (naïve users&amp;rsquo; perceived rate) is the key parameter driving all welfare results: a larger gap implies greater de facto manipulation and a stronger incentive for the platform to adopt an advertising-based model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Berk-Nash Equilibrium&lt;/strong&gt;: The solution concept from Esponda and Pouzo (2016), used to model agents with misspecified subjective models. All agents are Bayesian conditional on their own subjective model. Sophisticates&amp;rsquo; subjective model equals the objective model (standard Bayesian), while naïfs update using the misspecified phi_{0,N}. Perfection requires sequential rationality at each information set given beliefs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;De Facto Manipulation&lt;/strong&gt;: The paper&amp;rsquo;s term for a situation in which the platform and firm exploit naïve users&amp;rsquo; misspecified model to boost demand and extract surplus, without requiring any outright deception in the formal sense. It arises because naïve users voluntarily choose high-ad-load plans (believing ads to be highly informative) and voluntarily over-purchase (having updated on what they mistakenly think are strong positive signals). The manipulation is &amp;ldquo;de facto&amp;rdquo; because it operates through the users&amp;rsquo; own rational (but misspecified) decision-making.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Separating Equilibrium&lt;/strong&gt;: An equilibrium in which naïve and sophisticated users self-select into distinct platform plans. In the advertising equilibrium, naïve users join an ad-heavy plan (extracting all their surplus via inflated willingness to pay for ads) while sophisticated users are either excluded or placed on a subscription plan. This separation is the vehicle through which the platform maximizes revenue from naïf manipulation while limiting the disciplining force of sophisticates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-Best Allocation&lt;/strong&gt;: The welfare-maximizing allocation subject to the incentive-compatibility constraints that users self-select into plans. Because naïve users prefer more ads than sophisticated users (the inverse of what the planner desires), the second best is a single pooling plan with an intermediate ad load alpha^{SB} in [alpha^{FB}_N, alpha^{FB}_S]. This is strictly worse than the first best but achieves average welfare above the no-platform baseline, and can be decentralized with a nonlinear ad tax, product subsidy, and platform subscription subsidy.&lt;/p&gt;</description></item><item><title>Optimal Resilience in Multitier Supply Chains</title><link>https://macropaperwarehouse.com/papers/optimal-resilience-in-multitier-supply-chains/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/optimal-resilience-in-multitier-supply-chains/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Grossman, Helpman, and Sabal ask what market failures arise in vertical supply chains with multiple production tiers, limited (non-anonymous) supply networks, arms-length transactions, and recurrent risks of disruption at every node. They then ask what government policies would be required to implement the socially efficient (first-best) allocation as a decentralized equilibrium, and — in a second-best environment where subsidies to firm-to-firm transactions are politically infeasible — how optimal policies to promote resilience and network formation differ.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper develops a general-equilibrium model of a closed economy with an arbitrary number S+1 of vertical production tiers (tier 0 through tier S). A finite measure of &amp;ldquo;lead&amp;rdquo; firms in tier S produce differentiated consumer goods under monopolistic competition using labor and a CES bundle of intermediate inputs from tier S-1 suppliers. Firms in each intermediate tier combine labor and inputs from the tier above using a Cobb-Douglas production function. Tier 0 firms produce from labor alone.&lt;/p&gt;
&lt;p&gt;Every firm faces an independent, non-zero probability of a catastrophic disruption (complete inability to produce). Firms may invest labor up front to moderate this risk — endogenous &amp;ldquo;resilience&amp;rdquo; — or may invest to forge relationships with a larger fraction of potential suppliers in the next upstream tier — endogenous &amp;ldquo;network thickness.&amp;rdquo; Each formed relationship costs k units of labor.&lt;/p&gt;
&lt;p&gt;After disruption shocks are realized, surviving firms negotiate quantities and payments bilaterally. Bargaining is sequential (beginning with lead firms negotiating with tier S-1, then tier S-1 with tier S-2, and so on to tier 0), and within each round is governed by Nash-in-Nash equilibrium (Horn and Wolinsky, 1988): each firm takes as given the outcomes of its negotiations with all other partners. The Nash surplus is split with exogenous bargaining weight β_s for the downstream buyer in the s-to-s−1 negotiation.&lt;/p&gt;
&lt;p&gt;The paper solves the planner&amp;rsquo;s direct-control problem and then characterizes the three sets of policy instruments needed to decentralize the first best: subsidies to input transactions between adjacent tiers, subsidies to investments in resilience (agility), and subsidies to network formation (redundancy). It then solves the second-best problem in which transaction subsidies are constrained to zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Transaction subsidies.&lt;/em&gt; In the competitive bargaining equilibrium, each pair of firms undervalues input transactions because the upstream firm anticipates paying a marked-up price when it bargains with its own suppliers. This cascading distortion means the private marginal cost of producing a tier-s good exceeds the social marginal cost. The optimal first-best transaction subsidy on sales by tier s firms (τ*&lt;em&gt;s) equals [γ_s + (1−γ_s)μ&lt;/em&gt;{s−1}]^{−1}, where γ_s is the labor share in tier s production and μ_{s−1} is the endogenous markup factor from bargaining at the s-to-s−1 interface. This subsidy depends only on production function parameters and bargaining weights at the immediately adjacent tier. No subsidy is needed at tier 0 (the most upstream tier), and no subsidy is applied to final-good sales. Under Assumption 1 — inputs become weakly less substitutable as goods proceed downstream — the optimal purchase subsidies rise monotonically as one moves downstream along the supply chain.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Resilience subsidies (first best).&lt;/em&gt; Two offsetting forces govern the optimal subsidy to resilience investments θ*&lt;em&gt;s at intermediate tiers: (i) firms capture only the fraction (1−β&lt;/em&gt;{s+1}) of the joint surplus that their resilience creates for downstream customers, creating underinvestment; (ii) optimal transaction subsidies inflate private profitability, creating a countervailing overinvestment incentive. The net optimal first-best subsidy for intermediate-tier firms is θ*&lt;em&gt;s = (1−β&lt;/em&gt;{s+1}) / τ*_s. This formula depends only on technological and bargaining parameters of tier s and the tier immediately adjacent; it does not depend on conditions elsewhere in the chain. When production parameters and bargaining weights are uniform across tiers, the first-best resilience subsidy is the same at every interior tier. If goods become strictly less substitutable downstream, the first-best subsidy for resilience declines monotonically as one moves downstream, and may turn into an optimal tax for middle tiers where the transaction subsidy is large enough to over-incentivize resilience investment. The first-best resilience subsidy always applies at both extreme ends of the chain.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Network formation subsidies (first best).&lt;/em&gt; Despite firms&amp;rsquo; private incentive to manipulate their number of upstream suppliers to improve bargaining position, the net strategic effect of network formation in general equilibrium exactly cancels the off-equilibrium spillovers to non-partners. As a result, the optimal first-best policy toward network formation at every tier is identical to the optimal policy toward resilience investment.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Second-best policies.&lt;/em&gt; When transaction subsidies are unavailable, uncorrected markups downstream from tier s depress demand for tier-s output, reducing profitability and incentives to invest in resilience below the first-best level. Second-best optimal subsidies for resilience and network formation therefore reflect production function parameters and bargaining weights throughout the entire downstream supply chain, not just at the immediately adjacent tier. Specifically, when buyer bargaining weights are non-increasing along the chain (β_{s+1} ≤ β_s for all s), the second-best subsidy to resilience falls monotonically as one moves downstream. This is the opposite pattern from what might be inferred from the first-best analysis when transaction subsidies are available: with non-increasing bargaining weights, second-best subsidies are larger for upstream producers than for downstream producers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Results are derived for a closed economy. Welfare is measured by the CES utility of the representative consumer over differentiated final goods. The sequential bargaining structure assumes contracts are written after disruption shocks are realized. Assumption 1 (σ_1 ≥ σ_2 ≥ … ≥ σ_S &amp;gt; ε, where σ_s is the elasticity of substitution between inputs at tier s and ε is the demand elasticity for final goods) is maintained for sharper monotonicity results on the structure of optimal subsidies across tiers.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-precise-structure-of-the-supply-chain-in-the-model-and-why-does-the-bargaining-take-place-sequentially-rather-than-simultaneously-across-all-tiers"&gt;Q1. What is the precise structure of the supply chain in the model, and why does the bargaining take place sequentially rather than simultaneously across all tiers?&lt;/h3&gt;
&lt;p&gt;A: The economy has S+1 tiers. Tier 0 firms use only labor; tier s firms (s = 1,…,S−1) use labor and a CES bundle of tier s−1 inputs with elasticity of substitution σ_s &amp;gt; 1; tier S firms produce final differentiated goods using labor and tier S−1 inputs under Cobb-Douglas technology. Sequential bargaining is imposed because the vast number of simultaneous negotiations across all tiers makes a grand coalition impractical. The timing is that lead firms (tier S) first negotiate input quantities and payments with their tier S−1 suppliers; those suppliers, now contractually obligated to their downstream customers, then negotiate with tier S−2, and so on up the chain until tier 1 firms contract with tier 0 suppliers.&lt;/p&gt;
&lt;h3 id="q2-how-is-the-markup-factor-defined-and-what-parameters-determine-it"&gt;Q2. How is the markup factor defined, and what parameters determine it?&lt;/h3&gt;
&lt;p&gt;A: The markup factor μ_s is the ratio of the payment per unit made by tier s+1 firms to the production cost of tier s firms. It equals μ_s = (1−β_{s+1}) · [σ_{s+1}/(σ_{s+1}−1)] + β_{s+1}, where β_{s+1} is the exogenous bargaining weight of the downstream (tier s+1) buyer. When the downstream firm has all bargaining power (β_{s+1} = 1), the markup equals unity (competitive outcome). When the upstream firm has all bargaining power (β_{s+1} = 0), the markup equals the standard monopoly markup σ_{s+1}/(σ_{s+1}−1). For intermediate bargaining weights, the markup is a weighted average. The markup enters the optimal transaction subsidy formula by inflating the private marginal cost of producing tier-s inputs above the social marginal cost.&lt;/p&gt;
&lt;h3 id="q3-why-are-no-subsidies-needed-for-the-most-upstream-tier-0-transactions-or-for-final-good-sales"&gt;Q3. Why are no subsidies needed for the most upstream (tier 0) transactions or for final-good sales?&lt;/h3&gt;
&lt;p&gt;A: For tier 0 transactions: when tier 0 and tier 1 firms bargain, the negotiations occur last sequentially and so do not affect any prior agreements. There are no downstream cascading markup effects — tier 0 firms produce from labor alone, so their private marginal cost equals their social marginal cost. The joint surplus maximization by the pair thus aligns with the planner&amp;rsquo;s objective, yielding τ*_0 = 1 (no intervention needed). For final-good sales: final producers do mark up above marginal cost under monopolistic competition, but all varieties are symmetric, so the markup affects all goods equally and does not distort relative consumption choices. Hence τ*_S = 1.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-two-offsetting-forces-that-determine-the-optimal-first-best-subsidy-to-resilience-investments-at-an-intermediate-tier"&gt;Q4. What are the two offsetting forces that determine the optimal first-best subsidy to resilience investments at an intermediate tier?&lt;/h3&gt;
&lt;p&gt;A: First, a firm in tier s captures only the fraction (1−β_{s+1}) of the joint surplus that its survival creates for its downstream customers (the rest is appropriated through bargaining by those customers), leading to underinvestment relative to the social optimum. Second, the optimal transaction subsidy τ*_s &amp;lt; 1 raises the private profitability of firms in tier s above its social value, because public finances bear part of the cost of their input purchases. This inflated private profitability encourages resilience investment beyond what the planner desires. The net optimal policy is θ*&lt;em&gt;s = (1−β&lt;/em&gt;{s+1}) / τ*_s, which may be a subsidy (θ*_s &amp;lt; 1) or a tax (θ*_s &amp;gt; 1) depending on which force dominates.&lt;/p&gt;
&lt;h3 id="q5-why-does-the-first-best-subsidy-for-resilience-at-an-intermediate-tier-depend-only-on-local-parameters-at-tier-s-and-its-immediate-neighbors-even-though-resilience-investments-generate-spillovers-to-firms-throughout-the-network"&gt;Q5. Why does the first-best subsidy for resilience at an intermediate tier depend only on local parameters (at tier s and its immediate neighbors), even though resilience investments generate spillovers to firms throughout the network?&lt;/h3&gt;
&lt;p&gt;A: When optimal transaction subsidies are in place at all tiers, a firm&amp;rsquo;s value becomes independent of the joint surplus in sales that occur between firms in tiers other than its own. That is, the positive spillovers to all firms farther upstream and downstream in a firm&amp;rsquo;s own network are exactly offset by the negative spillovers to firms in rival networks (including rival firms in the same tier). What remains after this general-equilibrium cancellation is only the benefit to the firm&amp;rsquo;s immediate downstream customers and the wedge created by the transaction subsidy. This result implies that the formula θ*&lt;em&gt;s = (1−β&lt;/em&gt;{s+1}) / τ*_s does not involve conditions at tiers other than s and s−1.&lt;/p&gt;
&lt;h3 id="q6-why-does-the-optimal-policy-for-network-formation-supplier-link-investment-equal-the-optimal-policy-for-resilience-investment-despite-the-fact-that-network-formation-also-strategically-improves-a-firms-bargaining-position"&gt;Q6. Why does the optimal policy for network formation (supplier link investment) equal the optimal policy for resilience investment, despite the fact that network formation also strategically improves a firm&amp;rsquo;s bargaining position?&lt;/h3&gt;
&lt;p&gt;A: Firms in intermediate tiers do have a private incentive to form additional supplier links specifically to improve their bargaining position vis-à-vis their upstream suppliers (by improving their outside options) and vis-à-vis their downstream customers (by the same mechanism). However, the authors show by comparing the firm&amp;rsquo;s first-order condition for link formation with the planner&amp;rsquo;s first-order condition that this strategic motivation exactly balances the offsetting general-equilibrium effects from rival firms doing the same. After this cancellation, the residual wedge between private and social incentives for network formation is identical to that for resilience investment. Hence #&lt;em&gt;_s = θ&lt;/em&gt;_s for all tiers.&lt;/p&gt;
&lt;h3 id="q7-how-do-second-best-policies-differ-from-first-best-policies-in-terms-of-both-the-magnitude-of-subsidies-and-the-information-required-to-set-them"&gt;Q7. How do second-best policies differ from first-best policies in terms of both the magnitude of subsidies and the information required to set them?&lt;/h3&gt;
&lt;p&gt;A: In the first best, the subsidy for resilience at tier s depends only on the bargaining weight β_{s+1} and the markup factor μ_{s−1} — parameters relevant to tier s and its immediate neighbors. In the second best, when transaction subsidies are unavailable, the optimal resilience subsidy at tier s is θ†&lt;em&gt;s = J^{−1} · [1 − (cumulative distortion of all downstream tiers)] · (1−β&lt;/em&gt;{s+1}), where J captures aggregate labor-market effects of all markups throughout the chain. This formula requires knowledge of production function parameters (labor shares γ_j, markups μ_j, elasticities σ_j) for every tier j downstream from s. The second-best subsidy may be larger or smaller than the first-best subsidy; it is more likely to exceed the first-best subsidy for upstream tiers, where the cumulative downstream distortions (uncorrected markups contracting demand) produce a larger shortfall in private profitability and hence a larger underinvestment in resilience.&lt;/p&gt;
&lt;h3 id="q8-under-what-condition-do-second-best-subsidies-fall-monotonically-as-one-moves-downstream-and-how-does-this-compare-to-the-first-best-pattern"&gt;Q8. Under what condition do second-best subsidies fall monotonically as one moves downstream, and how does this compare to the first-best pattern?&lt;/h3&gt;
&lt;p&gt;A: The ratio of second-best subsidies at adjacent tiers (θ†_{s−1} / θ†&lt;em&gt;s) equals [(1−β_s) / (1−β&lt;/em&gt;{s+1})] · [τ*&lt;em&gt;s]^{−1}, where τ*&lt;em&gt;s is the first-best transaction subsidy. If buyer bargaining weights are non-increasing along the chain — β&lt;/em&gt;{s+1} ≤ β_s for all s — then (1−β_s) ≤ (1−β&lt;/em&gt;{s+1}) and, combined with τ*&lt;em&gt;s ≤ 1, the second-best subsidy is larger upstream than downstream (θ†&lt;/em&gt;{s−1} ≥ θ†_s). This contrasts with the first-best policy: when parameters are uniform across tiers, first-best resilience subsidies are the same at every interior tier, while second-best subsidies are strictly larger upstream than downstream.&lt;/p&gt;
&lt;h3 id="q9-what-role-does-assumption-1-elasticities-of-substitution-non-increasing-as-goods-move-downstream-play-in-the-results"&gt;Q9. What role does Assumption 1 (elasticities of substitution non-increasing as goods move downstream) play in the results?&lt;/h3&gt;
&lt;p&gt;A: Assumption 1 (σ_1 ≥ σ_2 ≥ … ≥ σ_S &amp;gt; ε) ensures that the operating profit function ~v_s(η) is concave in a firm&amp;rsquo;s network size η, which in turn ensures interior solutions to the network formation problem. It also delivers sharper monotonicity results: under this assumption, if other production parameters and bargaining weights are similar across tiers, the optimal purchase subsidies rise monotonically downstream, and the optimal first-best resilience subsidies decline monotonically downstream (potentially turning into taxes at some interior tiers). The assumption reflects the realistic view that inputs become more differentiated and specialized as they approach the final consumer good.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-limitations-the-authors-identify-regarding-their-model-and-what-extensions-do-they-suggest"&gt;Q10. What are the limitations the authors identify regarding their model, and what extensions do they suggest?&lt;/h3&gt;
&lt;p&gt;A: Three main limitations are identified. First, the model assumes bargaining occurs after disruption shocks are realized, ruling out contingent contracts. Pre-disruption bargaining with contingent payments could mitigate double-marginalization inefficiencies and help internalize resilience externalities, though complex network-wide contingent contracts would likely be needed for full efficiency even in the second-best environment. Second, the model assumes symmetric firms within each tier, so downstream firms cannot sort on upstream firms&amp;rsquo; observable resilience levels; if observable differences existed, downstream firms could seek out more reliable partners, partially internalizing the resilience externality. Third, the model covers only a closed economy with idiosyncratic (uncorrelated) shocks. Extensions to global supply chains, correlated (geographic) shocks, cross-country differences in wages and technologies, and optimal cooperative versus unilateral policy are identified as important directions for future research.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Resilience (agility):&lt;/strong&gt; In the paper&amp;rsquo;s usage, a firm&amp;rsquo;s endogenous investment in reducing the probability of a catastrophic disruption to its own operations. A firm in tier s hires r_s units of labor up front, which raises its survival probability φ_s(r_s), with φ&amp;rsquo;_s &amp;gt; 0 and φ&amp;rsquo;&amp;rsquo;_s &amp;lt; 0. Resilience is a relationship-specific investment in the sense that its payoff is realized only conditional on the firm surviving and then trading with its downstream customers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Network thickness (redundancy):&lt;/strong&gt; The fraction η_s of firms in the next upstream tier with whom a firm in tier s forms a supply relationship prior to the disruption shock. Forming k units of labor per link creates a thicker network that hedges against supplier disruption, increases input variety (and thus CES productivity), and improves bargaining positions vis-à-vis both upstream suppliers and downstream customers. Distinct from resilience: resilience reduces the firm&amp;rsquo;s own probability of disruption; network thickness provides substitutability across suppliers should some fail.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Markup factor (μ_s):&lt;/strong&gt; The ratio of the per-unit payment made by tier s+1 firms to the production cost of tier s firms, as determined by Nash bargaining. Specifically, μ_s = (1−β_{s+1}) · [σ_{s+1}/(σ_{s+1}−1)] + β_{s+1}. The markup distorts private marginal costs above social marginal costs, causing underinvestment in transactions between firms and, transitively, in resilience and network formation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nash-in-Nash equilibrium:&lt;/strong&gt; The bargaining solution concept used in the paper (following Horn and Wolinsky, 1988). Each pair of firms negotiates as if all other bilateral negotiations involving either party proceed at their equilibrium outcomes, both on and off the equilibrium path. This is the appropriate equilibrium concept when grand coalitions across all firms and all tiers are impractical.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sequential bargaining:&lt;/strong&gt; The specific timing structure in which negotiations proceed from the most downstream tier (lead firms bargaining with tier S−1 suppliers) sequentially upstream until tier 1 firms bargain with tier 0 suppliers. Each tier of firms, at the time they bargain with their own suppliers, are already contractually obligated to deliver specified quantities to their downstream customers. This obligation anchors the downstream firm&amp;rsquo;s outside option in any given bilateral negotiation.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;First-best transaction subsidy (τ&lt;/em&gt;_s):&lt;/em&gt;* The fraction of the cost of a tier-s input that, under the optimal policy, the downstream (tier s+1) buyer must pay. Equals [γ_s + (1−γ_s) · μ_{s−1}]^{−1} &amp;lt; 1 for all intermediate tiers, i.e., it is always a subsidy. Designed to align private marginal cost in the bilateral negotiation with the social marginal cost by offsetting the distortion introduced by anticipated markups on the upstream firm&amp;rsquo;s own inputs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-best subsidy:&lt;/strong&gt; The optimal policy toward resilience and network formation when subsidizing firm-to-firm transactions is infeasible (constrained to τ_s = 1 for all s). Unlike first-best subsidies — which depend only on local tier parameters — second-best subsidies depend on production function parameters and bargaining weights throughout the entire downstream supply chain due to the uncorrected cumulative markup distortions.&lt;/p&gt;</description></item><item><title>Organizational Change and Reference-Dependent Preferences</title><link>https://macropaperwarehouse.com/papers/organizational-change-and-reference-dependent-preferences/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/organizational-change-and-reference-dependent-preferences/</guid><description>&lt;p&gt;Schmidt and von Wangenheim develop a dynamic model of organizational change in which workers have reference-dependent preferences — specifically loss aversion and social comparisons — to explain several empirically observed patterns that standard models cannot easily account for: organizational inertia in normal times, sudden productivity jumps during crises, persistent total factor productivity (TFP) differences across firms in the same industry, and effort and wage compression within firms.&lt;/p&gt;
&lt;p&gt;The motivating empirical puzzle is the early-1980s collapse of the Great Lakes iron ore and steel industry, which had been geographically shielded from foreign competition for over 100 years. When Brazilian competitors undercut prices, the industry responded by roughly doubling labor productivity within a few years — not through new technology or capital investment, but through organizational improvements and more efficient use of existing capital (Schmitz 2007). The broader puzzle is Syverson&amp;rsquo;s (2004) finding that at the four-digit industry level, the 90th-percentile firm has TFP 1.9 times that of the 10th-percentile firm, a gap that cannot be explained by observable input differences.&lt;/p&gt;
&lt;p&gt;The model features a principal (firm owner) bargaining with loss-averse workers (represented by a union) over organizational change — represented as a worker effort level x that adapts the firm to the state of technology θ. Workers&amp;rsquo; reference point is a convex combination of the status quo contract and their rational expectations of the agreed contract, with weight α on the status quo. Loss aversion parameter λ &amp;gt; 0 means that losses relative to the reference point are weighted more heavily than gains.&lt;/p&gt;
&lt;p&gt;The core static result (Proposition 1) is that loss aversion drives a wedge of 1 + αλ between the workers&amp;rsquo; marginal cost and the firm&amp;rsquo;s marginal benefit of organizational change. Below a threshold θ defined by ∂v(x₀,θ)/∂x = 1 + αλ, there is complete inertia: the firm does not change the effort level at all. Above θ, the firm adjusts effort, but to x(θ) &amp;lt; x^ME(θ), undershooting the materially efficient level. Higher λ or higher α both widen the inertia range and reduce the amount of implemented change (Proposition 2).&lt;/p&gt;
&lt;p&gt;A crisis — modeled as a cost shock that makes the status quo contract generate negative profits, threatening firm closure — changes workers&amp;rsquo; outside option from their current utility U₀ to the unemployment utility of zero. Workers are now willing to accept either wage cuts or effort increases to keep their jobs. Crucially, because both concessions are perceived as losses of equal size by workers, the firm prefers to increase effort rather than cut wages, since increasing effort is more productive when x &amp;lt; x^ME. The model thus provides a microfoundation for downward nominal wage rigidity: in a recession, workers make concessions through harder work rather than wage cuts.&lt;/p&gt;
&lt;p&gt;In the infinite-horizon dynamic model, workers accumulate a quasi-rent over time equal to αλ(x_{t-1} − x₀), which represents compensation paid for past effort increases. This quasi-rent is what the firm expropriates during a crisis, allowing a discontinuous jump in effort toward the materially efficient level. Firms founded at different times or hitting different idiosyncratic shocks will therefore have different effort histories and different productivity levels, generating persistent TFP differences even among firms with identical technologies. When forward-looking players anticipate the possibility of crisis, inertia in normal times actually widens further (x̃(θ) ≤ x(θ)), because firms rationally delay effort adaptation knowing it will be cheaper to implement change during a crisis.&lt;/p&gt;
&lt;p&gt;The expectations-management extension (Section 4) introduces a moral-hazard problem with a manager who chooses the probability of successful change. Because a higher probability of change raises the workers&amp;rsquo; expectation-based reference point and reduces their perceived adaptation cost, the firm&amp;rsquo;s optimization problem becomes convex when the cost of effort for management is sufficiently low relative to (1−α)λΔx. This delivers a bang-bang result: the principal induces either full implementation (p = 1) or no change (p = 0), never an interior probability. This formalizes the management-consulting advice that commitment and urgency are essential to organizational change.&lt;/p&gt;
&lt;p&gt;The social-comparisons extension (Section 5) shows that when workers compare their wages and effort to colleagues, the firm optimally compresses effort differences across workers — inducing the less productive worker to work more than efficiency requires and the more productive worker to work less. If productivity differences between workers are sufficiently small, the firm sets identical effort levels. Wage compression follows from effort compression. To avoid the cost of social comparisons entirely, it may be optimal for the firm to split into separate legal entities whose workers no longer form a common reference group — a new explanation for organizational unbundling.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the core mechanism by which loss aversion generates organizational inertia in normal times?&lt;/strong&gt;
A: Workers have a reference point that is a convex combination (weight α on status quo, weight 1−α on rational expectations) of their current contract and the expected new contract. Because workers perceive an effort increase above their reference effort as a loss, the firm must pay a wage premium of αλ per unit of additional effort on top of the material effort cost of 1. This raises the effective marginal cost of implementing change from 1 to 1 + αλ, so the firm only implements change when the marginal revenue of effort strictly exceeds 1 + αλ. Below the threshold technology level θ (defined by ∂v(x₀,θ)/∂x = 1 + αλ), there is complete inertia and the firm keeps x* = x₀.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does a crisis break the inertia?&lt;/strong&gt;
A: A crisis is a cost shock large enough to make the firm&amp;rsquo;s profits negative under the status quo contract, so the firm would close unless workers make concessions. Workers&amp;rsquo; outside option shifts from their accumulated utility U₀ to the unemployment utility of zero. Because wage cuts and effort increases are both perceived as losses of equal magnitude, the firm prefers to demand effort increases (which raise revenue) over wage cuts (which do not). At the margin, when workers are at zero utility, the loss-aversion terms cancel from the marginal rate of substitution, and the firm can push effort up to the materially efficient level x^ME — a discontinuous jump.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why do wages not fall during a recession in this model?&lt;/strong&gt;
A: Workers perceive both wage cuts and effort increases as losses of equal per-unit utility cost. Since increasing effort by one unit and cutting wages by one unit impose the same utility cost on workers but effort increases raise firm revenue while wage cuts do not, it is always more efficient for the firm to extract concessions through higher effort rather than lower wages. The firm therefore first drives effort to x^ME before cutting wages, and cuts wages only if the zero-utility constraint still is not binding at x^ME. This provides a microfoundation for Bewley&amp;rsquo;s (1999) observation that wages do not fall during recessions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Where does the quasi-rent exploited during a crisis come from?&lt;/strong&gt;
A: Every time the firm implements an effort increase in normal times it must compensate workers with a permanent wage increase to cover both the permanent higher effort cost (x_{t}−x_{t-1}) and the one-time behavioral adaptation cost αλ(x_{t}−x_{t-1}). Because the compensation for the adaptation cost must be spread over all future periods as a permanent payment, workers accumulate a quasi-rent that by period t equals αλ(x_{t-1}−x₀) above their initial utility U₀ = w₀−x₀. This is the rent the firm expropriates in a crisis to fund the discontinuous effort increase.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the dynamic model generate persistent TFP differences across firms in the same industry?&lt;/strong&gt;
A: Firms founded at different times start with different initial status-quo effort levels relative to the current technology θ. Because each firm&amp;rsquo;s path of organizational adaptation is history-dependent — inertia regions, timing of crises, and accumulated quasi-rents all depend on when the firm was founded and what idiosyncratic shocks it experienced — firms that start later (or hit crises earlier) can remain more productive than older firms for extended periods. The numerical example with v(x,θ) = θ ln(x), α = 0.5, λ = 1, δ implied parameters shows that a firm founded when θ = 7 at the materially efficient point can maintain a substantial productivity advantage over a firm founded when θ = 4 that has accumulated inertia, even though both firms have access to the same technology.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does rational anticipation of a future crisis increase or decrease inertia in normal times?&lt;/strong&gt;
A: It strictly increases inertia. When players assign probability µ &amp;gt; 0 to a crisis each period, forward-looking workers demand higher compensation for effort increases in normal times — specifically, the per-period compensation for behavioral adaptation cost rises from (1−δ)αλ to γ = (1−δ(1−µ))αλ, which is increasing in µ. Simultaneously, the firm anticipates that effort adaptation will be cheaper to achieve in a crisis and therefore delays effort increases. The result is that the inertia threshold shifts from x(θ) to x̃(θ) ≤ x(θ), a strictly wider inertia region (Proposition 6).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the expectations-management result and what drives it?&lt;/strong&gt;
A: When a manager chooses the probability of successful change p at cost c(p) = (c/2)p², the wage the firm must pay workers is concave in p (equation 22): w = x₀ + p(1+λ)Δx − p²(1−α)λΔx + U₀. The concavity arises because a higher p raises the expectation-based component of the reference point, lowering workers&amp;rsquo; perceived adaptation cost. When c &amp;lt; (1−α)λΔx, this makes the principal&amp;rsquo;s profit function convex in p, so the optimum is at a corner: the principal induces either p = 1 (full implementation) or p = 0 (no change). Even when an interior solution obtains, a decrease in α (more weight on expectations) increases p. This formalizes the practitioner prescription that organizational change requires convincing everyone that change is certain and unavoidable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the effort and wage compression result under social comparisons?&lt;/strong&gt;
A: When each worker compares his situation to his colleague&amp;rsquo;s, with weight β on the peer&amp;rsquo;s wage and effort in forming the reference point, the firm must pay both workers a social-comparison premium of λβ(x₂−x₁) per unit of effort difference (Lemma 5). The firm therefore optimally compresses effort differences: it induces the less productive worker to exert effort above his efficient level and the more productive worker below his efficient level, at first-order conditions ∂v₁/∂x = 1 − 2λβ and ∂v₂/∂x = 1 + 2λβ respectively. If the productivity difference is small enough (specifically if ∂v₂(x*,θ)/∂x &amp;lt; 1 + 2λβ at the equal-effort point), the firm sets x₁* = x₂* = x*, eliminating wage inequality entirely (Proposition 8).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why might it be optimal for a firm to split into separate entities?&lt;/strong&gt;
A: Social comparisons impose costs on the firm by requiring higher wages for both workers (each receives a premium of λβ(x₂−x₁) regardless of their relative rank) and by distorting effort levels away from their efficient values. If workers employed by legally separate firms no longer treat each other as part of their reference group — because β falls to zero across firm boundaries — the firm can eliminate these comparison costs by spinning off activities into independent entities. This provides an efficiency rationale for organizational unbundling that does not rely on asset specificity or transaction costs, addressing what the authors call the &amp;ldquo;Williamson puzzle.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the implications for older workers and for social insurance policy?&lt;/strong&gt;
A: Older workers have two compounding reasons to be more resistant to organizational change: shorter remaining time horizons reduce the present value of permanent wage compensation for adaptation costs, and Gächter, Johnson, and Herrmann (2022) report that loss aversion λ increases with age, income, and wealth. Both factors raise the cost of implementing change with older workers. For social insurance, generous unemployment benefits or policies preventing layoffs (such as short-time work schemes) reduce workers&amp;rsquo; concession costs in a crisis, weakening the mechanism by which crises trigger change. The model suggests this may contribute to slower technology adoption in countries with stronger labor market protections.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What empirical facts from the existing literature does the model account for?&lt;/strong&gt;
A: The model accounts for: (1) Syverson&amp;rsquo;s (2004) finding of a 90th/10th percentile TFP ratio of 1.9 in four-digit US industries; (2) the iron ore and steel case study (Schmitz 2007) in which labor productivity doubled within a few years of a competitive shock with no new technology; (3) Bloom et al.&amp;rsquo;s (2014) correlation between more intense competition and higher TFP; (4) Holmes and Schmitz&amp;rsquo;s (2010) survey finding that competitive shocks raise industry productivity mainly through survival and improvement of existing firms; (5) Bewley&amp;rsquo;s (1999) downward nominal wage rigidity; and (6) Hjort, Li, and Sarsons (2022) on multinational firms using headquarters wages as reference points for wages in low-wage locations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Loss aversion (λ):&lt;/strong&gt; The parameter measuring the degree to which workers weight losses relative to their reference point more heavily than gains. A meta-analysis (Brown et al. 2023) across 607 empirical estimates finds an average loss aversion parameter of 1 + λ = 1.955. In this paper, λ &amp;gt; 0 means workers perceive a wage cut and an effort increase as losses, raising the effective marginal cost of organizational change by a factor of 1 + αλ.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reference point (w^r, x^r):&lt;/strong&gt; The benchmark wage and effort level against which workers evaluate outcomes. Defined as a convex combination of the status quo contract (w₀, x₀) with weight α and the rational expectation of the agreed contract (w^e, x^e) with weight 1−α. Losses occur when the realized wage falls below w^r or the realized effort exceeds x^r.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Organizational inertia:&lt;/strong&gt; The firm&amp;rsquo;s failure to implement materially efficient organizational change even when doing so would increase total surplus. In the model, inertia arises because the effective marginal cost of effort to the firm is 1 + αλ rather than 1, so the firm only implements change above a threshold technology level θ. The range of inertia widens with higher λ, higher α, and higher initial effort x₀.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quasi-rent:&lt;/strong&gt; The utility accumulated by workers above their initial utility U₀ = w₀−x₀ as compensation for past effort increases. By period t it equals αλ(x_{t-1}−x₀). This quasi-rent is the source of concessions the firm can extract in a crisis: workers accept higher effort (or lower wages) in exchange for keeping their jobs rather than losing this accumulated utility through unemployment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Behaviorally efficient effort x(θ):&lt;/strong&gt; The effort level that maximizes joint surplus taking behavioral adaptation costs into account, defined by ∂v(x,θ)/∂x = 1 + (1−δ)αλ in the dynamic model. This is strictly below the materially efficient effort x^ME(θ) (defined by ∂v/∂x = 1) and strictly above the firm&amp;rsquo;s privately optimal effort in normal times.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Effort compression:&lt;/strong&gt; The result under social comparisons that the principal optimally reduces the effort difference between workers relative to the efficient allocation — inducing the less productive worker to work more and the more productive worker to work less than efficiency requires. Driven by social-comparison costs λβ(x₂−x₁) that both workers receive as premiums regardless of relative rank.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Expectations management:&lt;/strong&gt; The strategic use of commitment to high probability of change in order to shift workers&amp;rsquo; expectation-based reference point and reduce the perceived adaptation cost. When α is small (rational expectations dominate the reference point), making change more certain lowers the wage cost of implementation, creating a complementarity between commitment and cost reduction that produces the bang-bang result: implement with certainty or not at all.&lt;/p&gt;</description></item><item><title>Patent Term, Innovation, and the Role of Technology Disclosure Externalities</title><link>https://macropaperwarehouse.com/papers/patent-term-innovation-and-the-role-of-technology-disclosure-externalities/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/patent-term-innovation-and-the-role-of-technology-disclosure-externalities/</guid><description>&lt;p&gt;This paper examines how anticipated changes in patent term affect R&amp;amp;D and innovation, using the U.S. ratification of the Trade-Related Aspects of Intellectual Property Rights (TRIPs) agreement in 1995 as a quasi-natural experiment. The central research question is whether and how policy anticipation shapes the short- and long-run dynamics of innovative activity, given ambiguous theoretical predictions: news of a patent term reduction could either deter innovation (by signaling lower future returns) or accelerate it (by inducing innovators to file under the more favorable existing regime before it expires).&lt;/p&gt;
&lt;p&gt;The identification strategy exploits a difference-in-differences (DiD) design using two sources of variation across 621 4-digit International Patent Classification (IPC) technological fields. The first is cross-sectional variation in field-specific pending periods — the time between patent application and grant during which monopoly rights are not fully enforceable — which determines whether TRIPs increased or reduced each field&amp;rsquo;s effective patent term (from 17 years post-grant to 20 years post-application minus the pending period). Fields with average pending periods exceeding three years faced expected reductions; those below faced extensions. On average across fields, TRIPs extended patent term by approximately 473 days (about 15 months), but approximately 45% of fields faced greater than 5% probability that individual patents would receive a term reduction. The second source is time variation from two events: a news event at the end of 1992 (when the Blair House Accord substantially reduced uncertainty about TRIPs adoption) and implementation in June 1995. The empirical sample spans 1985Q1–2000Q4 using PATSTAT patent data, augmented by firm-level R&amp;amp;D data from NBER-Compustat for 2,410 listed U.S. firms.&lt;/p&gt;
&lt;p&gt;Three main empirical facts emerge. First (Fact 1), innovation and R&amp;amp;D accelerate more during the anticipation phase (1992Q4–1995Q2) in fields with a higher probability of patent term reduction. A one-percentage-point higher reduction probability corresponds to a 1.4% larger increase in granted patent applications before implementation; a one-month shorter average patent term extension corresponds to a 2.9% larger increase. At the firm level, a one-percentage-point higher reduction probability is associated with a 1.9% increase in annual R&amp;amp;D expenditure (approximately $1.7 million), ruling out the interpretation that rising patent counts merely reflect strategic filing adjustments.&lt;/p&gt;
&lt;p&gt;Second (Fact 2), this heightened innovative activity persists for at least five years after implementation. Two years post-implementation, a one-percentage-point higher reduction probability corresponds to 1.44 additional quarterly patents (+2.7% in Poisson estimates), and a one-month shorter term extension corresponds to 3.3 more patents (+5.9%). This persistence is driven by indirect effects: the anticipation-induced burst in patenting generates additional follow-on innovation through technology disclosure externalities linked to cumulative knowledge creation. The elasticity of post-implementation innovation to news-phase innovation is estimated at approximately 2.1.&lt;/p&gt;
&lt;p&gt;Third (Fact 3), the direct effect of patent term on innovation — estimated by augmenting the DiD specification to control for field-specific innovation histories — is negative for shorter extensions and consistent with prior literature. A one-month shorter patent term extension reduces quarterly patents by 1.7%, and a one-year reduction reduces them by 20.9%. These estimates align with Budish, Roin, and Williams (2015, 2016), who find that a one-year extension of patent monopoly increases R&amp;amp;D by 7%–22% in pharmaceuticals. The identification is supported by the absence of pre-trends, by the finding that pre-news pending period distributions predict realized post-news variation with coefficients near one (0.957–1.104), and by extensive robustness checks.&lt;/p&gt;
&lt;p&gt;Q: What was the effective change in U.S. patent term under TRIPs, and why did it differ across fields?
A: TRIPs shifted patent expiry from 17 years after grant to 20 years after application date. Because monopoly rights are only fully enforceable after grant, the effective term became 20 years minus the pending period. Fields with average pending periods shorter than three years received net extensions; fields with longer average pending periods faced net reductions. Cross-field variation in pending periods arises because applications in different technical fields are reviewed by distinct USPTO technical units with different complexity and backlog levels.&lt;/p&gt;
&lt;p&gt;Q: What was the news event, and how was anticipation established?
A: The paper identifies November 1992 — when the Blair House Accord substantially reduced uncertainty about TRIPs adoption — as the news event, with formal ratification in December 1994 and implementation in June 1995. Documentary evidence confirms anticipation: U.S. business executives were involved in TRIPs negotiations from 1986; the patent term change appeared in a 1991 GATT draft; an Advisory Committee report co-signed by IBM, 3M, Motorola, and others referenced it in August 1992; and a New York Times article noted proposed changes in September 1992.&lt;/p&gt;
&lt;p&gt;Q: How is the probability of patent term reduction (PL_j) constructed, and what is its distribution?
A: PL_j is the fraction of patents in field j granted before the TRIPs news with a pending period exceeding three years, computed using PATSTAT data on U.S. patents granted between January 1990 and May 1992. Approximately 45% of fields faced a reduction probability exceeding 5%, and 15% faced a probability exceeding 10%. Even fields with an average term extension greater than one year had individual-patent reduction probabilities as high as 40%. A 10-percentage-point increase in PL_j corresponds to approximately a four-month shorter average term extension.&lt;/p&gt;
&lt;p&gt;Q: What is Fact 1 and what are its quantitative magnitudes?
A: Fact 1 states that during the news phase, innovation and R&amp;amp;D increase relatively more in fields with higher patent term reduction probability and shorter average term extension. One year after the news (two years before implementation), a one-percentage-point higher reduction probability generates 0.19 additional quarterly patents (+0.5% in Poisson estimates); a one-month shorter average extension generates 0.35 additional units (+0.8%). These effects approximately triple one year before implementation. At the firm level, a one-percentage-point higher probability is associated with a 1.9% increase in annual R&amp;amp;D (~$1.7 million) in 1993.&lt;/p&gt;
&lt;p&gt;Q: Why does news of a potential patent term reduction accelerate rather than deter innovation?
A: Innovators who anticipate a reduction in future patent protection under the new regime have strong incentives to file applications before implementation to secure the longer 17-years-from-grant term while it remains available. The acceleration is therefore consistent with innovators preferring longer protection: they rush to file under the more favorable old regime rather than curtailing innovation. Complementary analyses exploiting within-field dispersion in pending periods find that firms were particularly responsive to scenarios involving adverse policy changes, consistent with loss aversion. The dynamics of the news-phase acceleration are also consistent with an R&amp;amp;D gestation lag of approximately two years, as estimated by Pakes and Schankerman (1984).&lt;/p&gt;
&lt;p&gt;Q: What is Fact 2 and what drives the post-implementation persistence?
A: Fact 2 states that the heightened innovation in fields with higher reduction probability persists for at least five years after June 1995, even though the direct effect of a shorter patent term is innovation-reducing. Two years post-implementation, a one-percentage-point higher reduction probability corresponds to 1.44 additional quarterly patents (+2.7% Poisson) and a one-month shorter extension to 3.3 additional patents (+5.9% Poisson). The persistence is driven by technology disclosure externalities: the news-phase acceleration generates new patented knowledge that subsequent innovations build upon. Fields where new inventions rely more heavily on past innovations from the same field — proxied by backward citation intensity — display stronger post-implementation persistence.&lt;/p&gt;
&lt;p&gt;Q: How does the paper separate direct from indirect (externality-driven) post-implementation effects?
A: Following Angrist and Pischke (2009), the paper augments the baseline DiD specification to control for field-specific innovation histories via a lagged moving average of past outcomes and pre-determined field attributes interacted with quarterly fixed effects. The resulting coefficients capture the effect of patent term variation orthogonal to the news-induced innovation dynamics. The direct effect estimates are negative post-implementation (Fact 3), while the overall estimates are positive (Fact 2), confirming that the indirect externality channel outweighs the direct channel in the post-implementation period.&lt;/p&gt;
&lt;p&gt;Q: What is Fact 3 and how does its magnitude compare to prior literature?
A: Fact 3 states that, controlling for the news shock, a shorter patent term extension leads to a relative decline in innovation post-implementation. The estimated semi-elasticity is 1.7% per one-month increase in patent term and 20.9% per one-year increase. These estimates align with Budish, Roin, and Williams (2015, 2016), who find a 7%–22% increase in pharmaceutical R&amp;amp;D per one-year extension, and with Hemous et al. (2023), whose model implies a 1.2% innovation increase per one-month extension.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated elasticity of post-implementation innovation to news-phase innovation, and what does it imply?
A: Point estimates imply that one additional patent during the news phase generates approximately 5.1 additional patents post-implementation. Given average patent counts of 408.5 during the news phase and 1,000.3 post-implementation, this corresponds to a percent-to-percent elasticity of approximately 2.1. This elasticity captures the technology disclosure externality channel by which transitory accelerations in patenting generate persistent follow-on innovation.&lt;/p&gt;
&lt;p&gt;Q: Why is ignoring anticipation (as in Abrams 2009) a problem for DiD identification?
A: Anticipation inflates patenting in fields with higher reduction probability during the pre-implementation period, violating the DiD assumption that pre-implementation outcomes provide an unaffected baseline. For example, between April 1994 and March 1995, average monthly patents in field C12P (high reduction probability) were 15.1 units above pre-news levels, versus only 2.4 in field E05D (low reduction probability). Using this inflated pre-implementation level as the DiD reference baseline reverses the sign of the estimated implementation effect relative to the specification that uses the unaffected pre-news baseline.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports the technology disclosure externality mechanism over alternative explanations?
A: The paper proxies technological dependence by backward citation intensity at the field level and finds that the news-phase acceleration propagates more strongly into post-implementation innovation in fields where new inventions more heavily cite prior same-field patents. Time-varying measures of technological dependence identify this channel as the primary driver of indirect post-implementation effects. Two alternative mechanisms — changes in technological competition and adjustments in patenting strategies — lack comparable empirical support. The finding is consistent with Hegde, Herkenhoff, and Zhu (2023), who document that permanent increases in knowledge diffusion speed permanently raise follow-on innovation rates.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of jointly considering anticipation and knowledge spillovers?
A: Standard patent term analyses that abstract from anticipation effects and knowledge spillovers may substantially mischaracterize full welfare implications. The paper shows that innovation-policy interventions shape both short- and long-run outcomes, and that near-term variation in innovative activity can itself drive medium- to long-term effects through technological externalities. The estimated semi-elasticities of news, direct, and indirect effects provide empirical calibration targets for normative endogenous growth models used to derive optimal patent term, complementing prior normative recommendations ranging from zero protection (Boldrin and Levine, 2013) to infinite protection (Gilbert and Shapiro, 1990).&lt;/p&gt;
&lt;p&gt;Effective patent term: The duration of legally enforceable monopoly granted by a patent, equal to 17 years after grant under the pre-TRIPs U.S. regime and 20 years after application minus the pending period under the post-TRIPs regime. Because enforcement begins only at grant, the pending period directly erodes effective protection.&lt;/p&gt;
&lt;p&gt;Patent term reduction probability (PL_j): The field-specific fraction of pre-TRIPs patents with a pending period exceeding three years, representing the probability that individual patent applications in that field obtain a net reduction in patent term under the new 20-years-from-filing rule.&lt;/p&gt;
&lt;p&gt;News effect: The incremental change in innovation or R&amp;amp;D at the time of policy announcement, induced by future anticipated changes in patent term, before the new policy enters into force. In this paper&amp;rsquo;s setting, the news effect is positive: higher reduction probability accelerates patenting as innovators rush to file under the favorable existing regime.&lt;/p&gt;
&lt;p&gt;Direct implementation effect: The component of the post-implementation change in innovation attributable to the patent term change itself, isolated by controlling for field-specific innovation histories (i.e., abstracting from the indirect effects of anticipation-induced knowledge accumulation). It is negative for shorter patent term extensions, with a semi-elasticity of 1.7% per one-month increase.&lt;/p&gt;
&lt;p&gt;Technology disclosure externality: The mechanism by which newly patented knowledge, disclosed through the patent system, enables subsequent inventors to build on prior innovations, generating follow-on inventive activity. In this paper, the transitory news-phase burst in patenting generates a persistent externality, particularly in fields with high backward citation intensity.&lt;/p&gt;
&lt;p&gt;Policy anticipation: The phenomenon whereby forward-looking agents adjust behavior in response to credible news about future policy changes before those changes take effect. In this paper, anticipation induces a pre-implementation acceleration in patenting that temporarily pushes innovation in the opposite direction from the direct long-run effect and generates persistent indirect post-implementation effects through knowledge spillovers.&lt;/p&gt;
&lt;p&gt;Pending period: The time between patent application and grant during which USPTO examines the application and during which full monopoly rights are not enforceable. Field-level heterogeneity in pending periods — arising from differences in examination complexity and USPTO unit congestion — is the source of cross-sectional identification in the DiD design.&lt;/p&gt;</description></item><item><title>Peer Effects in Consideration and Preferences</title><link>https://macropaperwarehouse.com/papers/peer-effects-in-consideration-and-preferences/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/peer-effects-in-consideration-and-preferences/</guid><description>&lt;p&gt;This paper develops a general nonparametric model of discrete choice in which peers influence agents through two distinct channels: (1) the set of alternatives an agent considers (consideration set effects) and (2) the agent&amp;rsquo;s preferences over those alternatives (preference effects). The framework embeds these peer mechanisms in a continuous-time Markov process where agents revise choices at Poisson alarm-clock rates. A peer is classified as a consideration peer, a preference peer, or both, and the network is encoded as two directed edge sets rather than one.&lt;/p&gt;
&lt;p&gt;The central identification challenge is recovering network structure, consideration probabilities, and preferences simultaneously, without relying on exogenous variation in covariates or the menu of available options. The paper shows this is achievable using time-series variation in the choices made by connected agents. The key insight is that consideration peers who adopt alternative v change the probability that the focal agent considers v — entering only the &amp;ldquo;consideration&amp;rdquo; term of the conditional choice probability (CCP) — while preference peers who adopt alternatives other than v change only the &amp;ldquo;conditional-on-consideration&amp;rdquo; selection probability. These cross-alternative patterns in the CCPs allow the researcher to distinguish the two channels. Once consideration-only peers are isolated, their choices serve as exclusion restrictions that mimic artificial menu variation, enabling nonparametric recovery of preferences.&lt;/p&gt;
&lt;p&gt;Identification proceeds in stages: (i) recover the full reference group of each agent from changes in CCPs; (ii) separate consideration-only peers from preference-affecting peers using cross-order effects across alternatives; (iii) distinguish preference-only peers from consideration-and-preference peers under an exclusion restriction (Assumption 4) requiring that an agent with a dual-channel peer also has at least one single-channel peer; (iv) recover consideration ratios Q(v|n+1)/Q(v|n) and then the full choice rule. The results allow arbitrary heterogeneity across agents and do not require exogenous menu variation or covariate shifters.&lt;/p&gt;
&lt;p&gt;For continuous-time data (Dataset 1), the CCPs and Poisson rates are exactly identified from the observed revision history. For discrete-time panel data (Dataset 2), identification is generic under a mild eigenvalue condition on the transition rate matrix.&lt;/p&gt;
&lt;p&gt;The empirical application studies store-opening decisions by China&amp;rsquo;s two dominant high-end tea chains — Heytea and Nayuki — across prefecture-level cities from their founding through end-2020. By that date, Nayuki had 485 stores in 57 cities and Heytea had 729 stores in 46 cities, in an industry whose total revenue grew from 42.2 to 83.1 billion yuan between 2017 and 2020. Each firm-market pair is modeled as an agent deciding whether to open a new store. The key exclusion restriction is that the cumulative store count of either firm in geographically neighboring markets shifts consideration probabilities but does not enter marginal profitability directly.&lt;/p&gt;
&lt;p&gt;Estimation via maximum likelihood yields four substantive findings: (1) Firms exhibit limited consideration — consideration probabilities for markets with no prior presence by either firm are substantially below one. (2) Stores in neighboring markets significantly raise consideration probabilities for a given market, for both own-firm and rival stores; this peer effect in consideration is described as economically large. (3) Own-market store density raises marginal profitability (density economies) while rival presence lowers it (competitive effects). (4) A full-consideration model that omits the attention stage overestimates the negative competitive effect and underestimates positive density effects.&lt;/p&gt;
&lt;p&gt;Counterfactual simulations show that removing attention constraints (full consideration) accelerates market penetration substantially: firms enter new markets earlier and achieve broader geographic coverage. Removing peer effects in consideration only — while retaining attention constraints — slows the diffusion of store openings across neighboring markets, because peer effects in consideration function as an informational cascade. Limited consideration also reduces competition by delaying rival entry into high-profitability markets, explaining a significant share of the geographic concentration in first- and second-tier cities during the early expansion phase. The paper&amp;rsquo;s scope is limited to settings with repeated, non-durable choices; it does not model forward-looking behavior or multiple equilibria, which the authors note as directions for future research.&lt;/p&gt;
&lt;p&gt;Q: What are the two peer-effect channels in the model, and how do they differ structurally?
A: A consideration peer influences whether an alternative enters the agent&amp;rsquo;s consideration set — specifically, the probability Q_a(v | n) that alternative v is considered is a function of the number n of consideration peers currently adopting v. A preference peer influences the choice rule R_a(v | y, C) — the probability that v is selected conditional on it being in the consideration set. Importantly, the paper models the two channels as affecting logically separate stages of the decision process, so the observed CCP factors into a consideration term and a conditional-selection term that respond to distinct sets of peers.&lt;/p&gt;
&lt;p&gt;Q: Why does the standard identification approach of varying menus fail here, and how does the paper substitute for it?
A: Menu variation requires the researcher to observe the same agent facing different sets of available alternatives, which is unavailable in many empirical settings. The paper replaces exogenous menu variation with endogenous variation generated by consideration-only peers: when a consideration-only peer adopts alternative v, the focal agent&amp;rsquo;s probability of considering v rises, effectively mimicking the removal of other alternatives from her consideration set. This peer-induced variation in consideration is then used to trace out the choice rule R_a over counterfactual menus without any actual menu changes.&lt;/p&gt;
&lt;p&gt;Q: How does the paper separate consideration peers from preference peers in the data?
A: The decomposition exploits an asymmetry in how the two peer types appear in the log-CCP. When a consideration peer switches to alternative v, the term ln Q_a(v | .) changes but the conditional-selection term ln D_a(v | .) remains unchanged, because the agent already considers v. Conversely, when a preference peer adopts an alternative other than v, only the conditional-selection term shifts. The paper formalizes this via cross-order effects of peers across alternatives in the CCPs (Propositions 3.1–3.3) and invokes Assumption 4 — requiring at least one single-channel peer when a dual-channel peer exists — to complete the separation.&lt;/p&gt;
&lt;p&gt;Q: What is Assumption 4 and why is it necessary?
A: Assumption 4 states that if agent a has a peer in N_CR_a (a peer affecting both consideration and preferences), then a also has at least one additional peer affecting only consideration or only preferences. Without this exclusion restriction, the consideration and preference effects of a dual-channel peer are not separately identified from each other; the single-channel peer provides the variation needed to pin down each component separately.&lt;/p&gt;
&lt;p&gt;Q: What does Proposition 2.1 establish and what does it require?
A: Proposition 2.1 establishes existence and uniqueness of an invariant equilibrium distribution mu over choice configurations, with full support. It requires Assumptions 1 (independent consideration), 2(i) (strictly positive consideration probability for every alternative), and 3(i) (strictly positive probability of selecting any non-default alternative from some reachable consideration set). The continuous-time Poisson structure ensures zero probability of simultaneous revisions, which rules out multiple equilibria in the data-generating process.&lt;/p&gt;
&lt;p&gt;Q: How does the paper handle discrete-time panel data, where only periodic snapshots of choices are observed?
A: The paper invokes results from Blevins (2017, 2026) to show that the transition rate matrix W of the continuous-time process is generically identified from the discrete-time transition matrix observed at interval Delta, provided the eigenvalues of W do not differ by integer multiples of 2&lt;em&gt;pi&lt;/em&gt;i/Delta. Once W is identified, the CCPs P and Poisson rates lambda_a are recovered. This result is described as generic, meaning it holds except on a measure-zero set of parameter values.&lt;/p&gt;
&lt;p&gt;Q: What data does the empirical application use, and what are the key sample statistics?
A: The application uses city-level store registration data sourced from the National Enterprise Credit Information Publicity System (via CnOpenData, 2021), supplemented by regional statistics from the China City Statistical Yearbook (2016–2021). The sample ends in 2020 to avoid COVID-19 demand shifts. By end-2020, Nayuki had 485 stores across 57 cities and Heytea had 729 stores across 46 cities. The high-end tea industry&amp;rsquo;s total revenue grew from 42.2 to 83.1 billion yuan between 2017 and 2020.&lt;/p&gt;
&lt;p&gt;Q: What is the key exclusion restriction in the empirical specification, and why is it plausible?
A: Stores in geographically neighboring markets (parameterized by distance bins d(m,m&amp;rsquo;)) enter the attention index pi_tilde but are excluded from the marginal profit index pi_bar. The rationale is that nearby store counts are informative signals that draw managerial attention to a market (an informational spillover) but do not directly alter the profitability of operating in that market — profitability depends on local demand, competition within the market, and own firm density, not on activity in adjacent markets. This restriction identifies the consideration-only peer channel.&lt;/p&gt;
&lt;p&gt;Q: What does the paper find about biases from ignoring limited consideration?
A: When the two-stage model (consideration + choice) is replaced by a single-stage full-consideration model, the estimated payoff parameters differ substantially. Specifically, the full-consideration model overestimates the negative effect of competition (rival presence in the same market) and underestimates the positive effect of own-store density. The intuition is that correlated entry patterns driven by shared consideration spillovers are misattributed to payoff interactions when the consideration stage is omitted.&lt;/p&gt;
&lt;p&gt;Q: What do the counterfactual simulations show about the role of limited consideration in market dynamics?
A: Three counterfactuals are compared against the baseline. Under full consideration (no attention constraints), market penetration is substantially faster — firms enter new markets earlier and achieve broader geographic coverage. Removing peer effects in consideration while retaining attention constraints slows geographic diffusion because the informational cascade that propagates entry to neighboring markets is eliminated. Limited consideration also reduces competition by delaying rival entry into high-profitability markets; markets with high potential demand remain underserved for longer. Collectively, limited consideration explains a significant portion of the geographic concentration of tea chain stores in first- and second-tier cities during the early expansion period.&lt;/p&gt;
&lt;p&gt;Q: What forms of heterogeneity does the identification allow, and what does it not require?
A: The nonparametric identification results accommodate arbitrary heterogeneity across agents in consideration mechanisms Q_a, choice rules R_a, Poisson revision rates lambda_a, and network positions. The identification requires neither exogenous covariates that shift preferences or consideration, nor variation in the set of available alternatives across observations. It relies solely on time-series variation in the choices made by connected agents, which are endogenous to the model and are themselves identified in the first stage.&lt;/p&gt;
&lt;p&gt;Q: How does the paper model history dependence, and does it change the main identification results?
A: Section 4.1 extends the model to allow consideration probabilities and choice rules to depend on the agent&amp;rsquo;s own choice history h_t in addition to the current configuration y. Proposition 4.1 states that under Assumptions 1–4 applied conditional on both y_{at} and h_t, all identification propositions from Section 3.1 remain valid. The extension also allows consideration probabilities to equal one, enabling nontrivial dynamics in consideration sets driven by past choices.&lt;/p&gt;
&lt;p&gt;Q: How is the unobservable default handled in the empirical application?
A: When the default alternative (e.g., &amp;ldquo;do not open a store&amp;rdquo;) is unobserved, the Poisson revision rate lambda_a cannot be separately identified from the CCPs without normalization. The paper normalizes lambda_a = 1 for each agent in the empirical application, treating the revision opportunity rate as fixed and recovering all remaining primitives under this normalization.&lt;/p&gt;
&lt;p&gt;Consideration set: The subset C of the full menu Y that agent a actually attends to at the moment of revision; formed before the choice rule is applied. Alternative v enters C independently with probability Q_a(v | n), where n is the number of consideration peers currently adopting v. The default alternative is always in the consideration set.&lt;/p&gt;
&lt;p&gt;Conditional choice probability (CCP): P_a(v | y), the ex-ante probability that agent a selects alternative v given choice configuration y; equal to the product of the consideration probability Q_a(v | .) and the conditional-selection probability D_a(v | .), integrated over all possible consideration sets.&lt;/p&gt;
&lt;p&gt;Choice configuration: The vector y = (y_a)_{a in A} recording the current alternative selected by every agent in the network simultaneously; the state variable of the continuous-time Markov process.&lt;/p&gt;
&lt;p&gt;Consideration-only peer: A peer a&amp;rsquo; in N_C_a \ N_R_a whose choices enter the consideration probability Q_a but not the choice rule R_a. Variation in the choices of consideration-only peers serves as an exclusion restriction that mimics artificial menu variation for identifying preferences.&lt;/p&gt;
&lt;p&gt;Preference-only peer: A peer a&amp;rsquo; in N_R_a \ N_C_a whose choices enter the choice rule R_a but not the consideration probability Q_a.&lt;/p&gt;
&lt;p&gt;Cross-order peer effect: The pattern in the CCP by which a consideration peer&amp;rsquo;s adoption of alternative v changes ln P_a(v | .) but not the conditional-selection component, while a preference peer&amp;rsquo;s adoption of a different alternative v&amp;rsquo; changes the conditional-selection component but not the consideration component; this asymmetry is the key to separating the two channels.&lt;/p&gt;
&lt;p&gt;Limited consideration: The situation in which Q_a(v | n) is strictly less than one for at least some alternatives v and peer counts n, so that the agent does not evaluate all available options before choosing; distinct from full rationality in which all alternatives are always considered.&lt;/p&gt;
&lt;p&gt;Mean attention index (pi_tilde): The latent index governing the consideration probability in the empirical specification; it depends on own and rival store counts in the same and neighboring markets and on firm fixed effects, but is excluded from the marginal profit index — constituting the empirical exclusion restriction that separates the consideration and payoff channels.&lt;/p&gt;</description></item><item><title>Quantifying the allocative efficiency of capital: The role of capital utilization</title><link>https://macropaperwarehouse.com/papers/quantifying-the-allocative-efficiency-of-capital-the-role-of-capital-utilization/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/quantifying-the-allocative-efficiency-of-capital-the-role-of-capital-utilization/</guid><description>&lt;p&gt;Standard measures of capital allocative efficiency—based on the dispersion of the average revenue product of capital (ARPK)—are severely biased when capital utilization is endogenous. When utilization is flexible, firms can bypass physical adjustment constraints by varying intensity, so that the correct efficiency measure requires the dispersion of average revenue product of capital services (ARPKS), defined as the log difference between revenue and utilized capital, not of ARPK. Contrary to the standard view that higher ARPK dispersion signals lower allocative efficiency, the paper demonstrates that when efficiency improvements arise from greater utilization flexibility, ARPK dispersion can increase alongside efficiency gains. An application to India&amp;rsquo;s capital market liberalization reform shows that the standard approach (ignoring utilization) predicts allocative efficiency gains of 5.25% (statistically significant), while the corrected approach accounting for utilization finds gains of only 0.04% (not statistically significant).&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Summary based on a working paper version, AI-assisted and human-reviewed. See the linked published article for the authoritative version.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-wrong-with-arpk-dispersion-as-a-measure-of-allocative-efficiency"&gt;Q1. What is wrong with ARPK dispersion as a measure of allocative efficiency?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;ARPK dispersion conflates two conceptually distinct quantities—the dispersion of physical factor allocation and the dispersion of utilization intensity—so it is not a monotone proxy for allocative efficiency when utilization is endogenous.&lt;/strong&gt; In the Hsieh-Klenow (2009) framework, higher ARPK dispersion is interpreted as higher misallocation. But ARPK is simply revenue over capital inputs, so when firms with too little (too much) capital relative to their productivity simply utilize their capital more (less) intensely, the variance of ARPK rises even as the efficiency of factor services allocation improves. The standard interpretation therefore has the causality backwards in economies with flexible utilization.&lt;/p&gt;
&lt;h3 id="q2-what-is-arpks-and-why-does-it-correctly-measure-capital-allocative-efficiency"&gt;Q2. What is ARPKS and why does it correctly measure capital allocative efficiency?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;ARPKS—the average revenue product of capital services, defined as the log difference between revenue and utilized capital—is the theoretically correct sufficient statistic for capital allocative efficiency when utilization is endogenous, because it measures the dispersion in the productivity of factor services rather than the dispersion of physical factor inputs.&lt;/strong&gt; The paper embeds endogenous utilization into a neoclassical investment model and shows formally that the variance of log ARPKS is zero in the efficient equilibrium, while the variance of log ARPK is not. ARPK dispersion is a combination of ARPKS dispersion and utilization variation, and the latter is not a sign of misallocation.&lt;/p&gt;
&lt;h3 id="q3-can-higher-arpk-dispersion-accompany-higher-allocative-efficiency"&gt;Q3. Can higher ARPK dispersion accompany higher allocative efficiency?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Yes: when allocative efficiency improvements arise from greater flexibility in capital utilization, ARPK dispersion increases even though allocative efficiency improves.&lt;/strong&gt; Greater utilization flexibility generates more variation in how intensively different firms use their capital, raising the ratio of revenue to physical capital input for high-utilization firms. A researcher using ARPK dispersion would therefore mistakenly conclude that allocative efficiency fell when it actually rose. The paper provides counterfactual simulations illustrating this phenomenon.&lt;/p&gt;
&lt;h3 id="q4-what-is-the-empirical-application-and-what-do-the-results-show"&gt;Q4. What is the empirical application and what do the results show?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;An application to Bau and Matray (2023)&amp;rsquo;s Indian capital market liberalization reform finds that: the standard approach (ignoring utilization) predicts allocative efficiency gains of 5.25% (statistically significant); accounting for utilization, the corrected estimate is only 0.04% (not statistically significant).&lt;/strong&gt; The analysis uses firm-level panel data for India including capacity utilization rates and capital maintenance expenses. Maintenance expenses serve as a proxy for capital utilization rates in the model, bridging the gap between observed capacity and the theoretically relevant utilization measure.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;average revenue product of capital (ARPK)&lt;/strong&gt; : the ratio of revenue to physical capital input; commonly used as a proxy for firm-level distortions in the Hsieh-Klenow (2009) framework, but shown to be a biased measure of allocative efficiency when capital utilization is endogenous.
&lt;strong&gt;average revenue product of capital services (ARPKS)&lt;/strong&gt; : the ratio of revenue to utilized capital; the theoretically correct measure of capital allocative efficiency when utilization is endogenous; its dispersion is zero in the efficient equilibrium.
&lt;strong&gt;capital utilization&lt;/strong&gt; : the intensity with which a firm deploys its physical capital stock; endogenous in the model, allowing firms to partially bypass adjustment constraints; the omission of utilization is the source of the bias in standard ARPK-based efficiency estimates.&lt;/p&gt;</description></item><item><title>Regulatory Competition in the US Life Insurance Industry</title><link>https://macropaperwarehouse.com/papers/regulatory-competition-in-the-us-life-insurance-industry/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/regulatory-competition-in-the-us-life-insurance-industry/</guid><description>&lt;p&gt;This paper quantitatively assesses the consequences of jurisdictional competition in the US life insurance industry, an $8 trillion market. The central question is whether competition between state regulators over capital requirements for captive reinsurance subsidiaries — a form of regulatory competition — increases or decreases total surplus, and by how much.&lt;/p&gt;
&lt;p&gt;US life insurers are regulated at the state level. Since the early 2000s, states have competed to attract captive reinsurance subsidiaries (captives) by setting lower capital requirements on these entities. The externality structure is asymmetric: the captive state earns tax revenues on liabilities transferred to captives and sets their capital requirements, but bears default costs only for policyholders in its own state. Consumer states bear default costs for their own residents even when those policies have been transferred to an out-of-state captive. This mismatch between who sets capital requirements and who bears default costs creates the externality that drives the race-to-the-bottom dynamic studied in the paper.&lt;/p&gt;
&lt;p&gt;The empirical setting draws on a novel dataset covering 66 US life insurers from 2005 to 2020, with total liabilities of $1.9 trillion (approximately 25% of the sector). Data sources include NAIC filings via S&amp;amp;P, CompuLife pricing data, A.M. Best ratings, SEC filings, and state legislative records. The author assembles novel data on captives&amp;rsquo; capital levels from SEC filings, Iowa Insurance Department captive financial statements, and insurer reinsurance exhibits.&lt;/p&gt;
&lt;p&gt;Three motivating empirical findings ground the structural model. First, captives materially reduce insurers&amp;rsquo; capital: in 2019, risk-based capital ratios are 23% lower on average after accounting for captives, with the median insurer&amp;rsquo;s capital declining 24%, and this translates into an increase in 10-year default probability from 1.0% to 2.9%. Second, states&amp;rsquo; capital requirements are the primary determinant of where insurers locate captives: a 1 percentage point increase in a state&amp;rsquo;s captive capital rate is associated with a 1.6 percentage point decrease in the probability an insurer chooses that state (against a 1.1 percentage point unconditional probability), and this holds when insurers switch states over time as capital requirements change. Tax rates, geographic proximity, and amenities are not meaningfully correlated with captive location choice. Third, a difference-in-differences design exploiting Regulation XXX (effective January 1, 2000), which raised capital requirements differentially across product term lengths, shows that 30-year term products — which faced the largest capital requirement increases — experienced price increases averaging 10.3% relative to 10-year term products, with quantities declining monotonically for longer-term products, consistent with an inward supply shift.&lt;/p&gt;
&lt;p&gt;The paper develops a structural model of the insurance market with imperfectly competitive insurers, endogenous default following a Leland (1994) framework, discrete choice consumer demand (Berry, 1994), and state regulators who set captive capital rates to maximize a weighted objective over tax revenues, default costs, consumer surplus, and producer surplus. Regulators deviate from a utilitarian social planner in two ways: they are state-based (generating competition and default externalities) and face agency frictions (captured by welfare weights that differ from unity). The demand side implies an average price elasticity of 2.4. The regulator side reveals that state regulators are willing to trade $1 of default costs against $3.5 of tax revenues and $0.59 of consumer surplus — both diverging from the social planner&amp;rsquo;s equal weighting.&lt;/p&gt;
&lt;p&gt;The main counterfactual finding is that eliminating competition by federalizing insurance regulation would cause regulators to raise capital requirements by 19% (3 percentage points), reducing expected default costs by $2.4 billion while lowering consumer surplus by $880 million, for a net total surplus gain of $1.5 billion. Regulator utility would increase by $3.3 billion in equivalent tax revenues. Because regulators over-value consumer surplus relative to default costs, competition exacerbates rather than counteracts their agency frictions, making competition unambiguously welfare-reducing in the baseline. A social planner would set capital requirements even higher than a federal regulator. On distribution, large states such as California and New York gain most from federalization (they bear substantial default costs), while Vermont — the largest captive state by market share — loses because it would forfeit captive tax revenues. Unilateral bans are found to have limited equilibrium consequences: a New York ban on captive use by insurers selling in New York would achieve only 23% of the national default cost reduction that federalization achieves, and a ban on captives domiciled in Vermont would achieve only 10%, as insurers would redirect captives to other states.&lt;/p&gt;
&lt;p&gt;Q: What is a captive reinsurance subsidiary and why do states compete to attract them?
A: A captive is a wholly-owned subsidiary of a life insurance holding company that reinsures policies written by the operating company, moving liabilities off the operating company&amp;rsquo;s balance sheet. Captive states earn tax revenues on liabilities transferred to captives and can set their own capital requirements on those entities, which are lower than the uniform NAIC standards applied to operating companies. Because captives are taxed by the state where they are domiciled — not the consumer&amp;rsquo;s state — captive states can earn tax revenues on policies sold elsewhere, incentivizing competition through lower capital requirements to attract insurers.&lt;/p&gt;
&lt;p&gt;Q: What is the default externality at the core of this paper&amp;rsquo;s argument?
A: When an insurer defaults, the shortfall on policies sold to consumers in a given state is borne by that state&amp;rsquo;s guaranty fund and consumers, regardless of where the captive holding those liabilities is domiciled. So Vermont, as the captive state, sets the capital requirement on liabilities transferred from (for example) Massachusetts policyholders, but does not bear the default cost on those Massachusetts policies. This means Vermont internalizes only the default cost on its own consumers, leading it to set capital requirements lower than it would if it bore the full default cost — a classic externality.&lt;/p&gt;
&lt;p&gt;Q: How large is the effect of captives on insurers&amp;rsquo; capital levels?
A: Using novel data on captives&amp;rsquo; actual balance sheets, the author finds that in 2019, the size-weighted average risk-based capital ratio of sample insurers is 23% lower after consolidating captives into the operating company&amp;rsquo;s balance sheet. The median insurer&amp;rsquo;s capital ratio decreases by 24%. In terms of default risk, this adjustment corresponds to an increase in the 10-year default probability from 1.0% to 2.9% based on historical insurer default rates.&lt;/p&gt;
&lt;p&gt;Q: What is the state of competition among captive domiciles in the data?
A: Twenty-two states had passed laws allowing captives as of the sample period, with the set of competing states largely stabilizing after 2013. The market is moderately concentrated: the top five states (Vermont, Arizona, Delaware, Iowa, and South Carolina) account for 80% of all captive liabilities, and the Herfindahl-Hirschman Index is 0.20. Vermont has maintained its position as the largest captive state throughout the period.&lt;/p&gt;
&lt;p&gt;Q: What evidence shows that capital requirements — rather than taxes or other factors — drive captive location choice?
A: In a linear probability model of captive location with insurer-year fixed effects, a 1 percentage point increase in a state&amp;rsquo;s captive capital rate is associated with a 1.6 percentage point decrease in the probability that an insurer chooses that state (versus a 1.1 percentage point unconditional probability). Captive tax rates are not meaningfully correlated with location choice, consistent with federal tax laws prohibiting the use of reinsurance to reduce tax liabilities. A changes-on-changes specification confirms that insurers are more likely to shift their captives to states that lower their capital requirements over time.&lt;/p&gt;
&lt;p&gt;Q: How does the Regulation XXX natural experiment identify the supply-side effect of capital requirements on insurance prices?
A: Regulation XXX, effective January 1, 2000, increased reserve requirements for operating companies on a mechanical basis tied to policy term length, with longer-term products facing larger increases. Using a difference-in-differences design at the insurer-product-month level with insurer-product and month fixed effects, the paper finds that products with larger capital requirement increases experienced larger price increases immediately after the regulation took effect. Thirty-year term products experienced price increases averaging 10.3% relative to 10-year term products (the reference group) within three months. Quantities also declined monotonically for longer-term products, confirming an inward shift of the supply curve rather than a demand shift.&lt;/p&gt;
&lt;p&gt;Q: What are the estimated regulator welfare weights, and what do they imply about agency frictions?
A: Normalizing the weight on tax revenues to 1, the paper recovers that regulators value $1 of default costs as worth $0.29 (implying $3.5 of tax revenues trades off against $1 of default costs) and value consumer surplus at $0.59 per dollar. Because the social planner sets all weights equal to 1, these estimates show regulators over-weight tax revenues and consumer surplus relative to default costs. The higher weight on consumer surplus is consistent with political backlash from consumers facing high insurance prices.&lt;/p&gt;
&lt;p&gt;Q: What is the total surplus effect of eliminating regulatory competition through federalization?
A: Federalizing insurance regulation — modeled as a single federal regulator setting a uniform capital rate while holding fixed regulatory frictions — would lead regulators to raise captive capital requirements by 19% (3 percentage points) to internalize the default externality. Expected default costs would fall by $2.4 billion. However, higher capital requirements would raise insurance prices and reduce consumer surplus by $880 million. The net effect is a total surplus increase of $1.5 billion. Regulator utility (in equivalent tax revenues) would increase by $3.3 billion.&lt;/p&gt;
&lt;p&gt;Q: Would eliminating both competition and regulatory frictions (i.e., a social planner) produce a different outcome than just federalizing?
A: In the baseline estimates, a social planner would set capital requirements even higher than a federal regulator, because regulators&amp;rsquo; agency frictions lead them to under-weight default costs relative to consumer surplus, pushing capital requirements below the socially optimal level even absent competition. Competition further exacerbates these frictions by providing an additional incentive to lower capital rates. Thus, in the baseline, competition unambiguously decreases total surplus. The paper also reports results under alternative assumptions, providing a &amp;ldquo;menu&amp;rdquo; for policymakers that maps different assumptions about regulators&amp;rsquo; frictions to quantitative welfare statements.&lt;/p&gt;
&lt;p&gt;Q: What distributional consequences across states explain why federalization has not been adopted?
A: Federalization would benefit large states such as California and New York most, because those states bear substantial default costs on large volumes of policies sold to their consumers. States with large captive market shares, primarily Vermont, would be made worse off because they would lose captive tax revenues. These predicted gains and losses align with actual state policy positions: New York has called for a national ban on captives, California forbids insurers from setting up captives there, and Vermont has been the most aggressive state in attracting captive domiciles.&lt;/p&gt;
&lt;p&gt;Q: How effective are unilateral state bans as an alternative to federal coordination?
A: The paper estimates that a unilateral ban by New York on insurers selling in New York from using captives would achieve only 23% of the national default cost reduction that full federalization would achieve. A unilateral ban on captives domiciled in Vermont — the largest captive state — would achieve only 10% of federalization&amp;rsquo;s default cost reduction, because insurers would simply relocate their captives to other states that still allow them. This finding underscores the importance of cross-state coordination for meaningful regulatory reform.&lt;/p&gt;
&lt;p&gt;Q: What does the model&amp;rsquo;s demand estimation imply about consumer sensitivity to insurance prices?
A: The discrete choice demand model estimated on state-level sales, prices, and product characteristics implies an average price elasticity of demand of 2.4 for life insurance products. This elasticity disciplines the quantitative impact of capital requirements on product markets through their effect on insurance prices.&lt;/p&gt;
&lt;p&gt;Q: How does the paper recover regulators&amp;rsquo; objective functions?
A: The author uses the revealed preferences of state regulators, exploiting regulators&amp;rsquo; utility maximization first-order conditions and performing numerical perturbations around those conditions to calibrate the welfare weights (lambdas) on each component of the regulators&amp;rsquo; utility function. This approach recovers regulators&amp;rsquo; tradeoff weights from their observed policy choices — specifically their captive capital rate decisions — without directly observing regulators&amp;rsquo; preferences.&lt;/p&gt;
&lt;p&gt;Captive reinsurance subsidiary: A wholly-owned subsidiary of a life insurance holding company that reinsures liabilities from the operating company. Unlike operating companies, captives are regulated by the state in which they are domiciled (the captive state) under that state&amp;rsquo;s own capital requirements, which are typically lower than the uniform NAIC standards. Captives allow insurers to reduce their overall capital requirements by allocating liabilities to the captive.&lt;/p&gt;
&lt;p&gt;Default externality: The mismatch between who sets capital requirements for captives (the captive state) and who bears default costs when an insurer fails (the consumer&amp;rsquo;s state and its guaranty fund). Because the captive state bears default costs only for its own residents — not for residents of states where the insurer also sells — it has an incentive to set lower capital requirements than it would if it internalized the full default cost, leading to an externality on other states.&lt;/p&gt;
&lt;p&gt;Risk-based capital ratio (adjusted for captives): The author&amp;rsquo;s measure of insurer capitalization after consolidating the captive&amp;rsquo;s balance sheet with the operating company&amp;rsquo;s. This adjusted ratio is lower than the statutory risk-based capital ratio that ignores captives, by 23-24% in the 2019 sample, and translates into meaningfully higher default probabilities (from 1.0% to 2.9% over 10 years).&lt;/p&gt;
&lt;p&gt;Regulatory agency frictions: Deviations of state regulators&amp;rsquo; objective functions from a utilitarian social planner&amp;rsquo;s, captured by welfare weights (lambdas) on each component of the regulator&amp;rsquo;s utility. In the paper&amp;rsquo;s estimates, regulators over-weight tax revenues ($3.5 of tax revenues per $1 of default costs) and consumer surplus ($0.59 per $1 of default costs) relative to the social planner&amp;rsquo;s equal weighting, consistent with political economy pressures from consumers and revenue incentives.&lt;/p&gt;
&lt;p&gt;Captive capital rate: The state-level capital requirement on captives, defined empirically as the sum of capital divided by the sum of liabilities of all captives in the state each year. Higher values represent more stringent requirements. The mean in the sample is 4% with a standard deviation of 3%, and captive capital rates are lower on average than operating company capital rates.&lt;/p&gt;
&lt;p&gt;Race to the bottom: The dynamic under which competition between state regulators leads each state to set lower capital requirements than it would absent competition, in order to attract captive tax revenues, resulting in a collectively worse equilibrium with higher default risks. The paper finds this outcome in the baseline: competition lowers capital requirements by 19% (3 percentage points) relative to a federal regulator.&lt;/p&gt;
&lt;p&gt;External financing frictions: The costs insurers face in raising equity capital, modeled as a per-dollar cost theta on required capital. These frictions create the supply-side channel through which capital requirements affect insurance prices: higher capital requirements raise insurers&amp;rsquo; marginal costs, leading to higher prices and lower quantities, as documented in the Regulation XXX natural experiment.&lt;/p&gt;</description></item><item><title>Rural Migrants and Urban Informality: Evidence From Brazil</title><link>https://macropaperwarehouse.com/papers/rural-migrants-and-urban-informality-evidence-from-brazil/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/rural-migrants-and-urban-informality-evidence-from-brazil/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question.&lt;/strong&gt; Does rural-urban migration increase or decrease urban informality, and through what mechanisms — and does the answer depend on the time horizon?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Setting and Data.&lt;/strong&gt; The paper studies internal migration in Brazil over 2000–2010. The empirical analysis combines: (i) two waves of the Decennial Population Census (2000 and 2010) covering working-age adults (ages 15–64) across 3,548 Minimum Comparable Areas (MCAs); (ii) the universe of formal firms and workers from the matched employer-employee administrative dataset RAIS (1997–2018); (iii) the ECINF informal firm survey (2003); and (iv) the annual National Household Survey (PNAD, 2001–2009) for year-on-year short-run analysis in 700 identifiable municipalities. Internal immigration to the average urban destination was large: 17.6 percent overall over the decade, 7 percent for state-to-state migration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Empirical Design.&lt;/strong&gt; The authors use a shift-share instrumental variable (IV) design. The shares are pre-existing migration networks (migrant flows by origin-destination pair, 1995–2000). The shifts are drought shocks constructed from the Standardized Precipitation-Evapotranspiration Index (SPEI) interacted with agricultural crop calendars and the value share of each crop in each origin municipality — accumulated over the 2000–2010 decade. A second independent instrument uses international commodity price shocks as push factors (following a China-analogous construction); the two instruments are nearly uncorrelated across origins (0.007) and only weakly correlated across destinations (-0.3), providing an independent validation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Long-Run Findings (decadal changes, 2000–2010).&lt;/strong&gt; A one-percentage-point increase in the immigration rate (equal to 18.5 percent of a standard deviation):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Increases the share of workers in formal wage employment by &lt;strong&gt;0.27 percentage points&lt;/strong&gt; (a 1.2 percent increase from the mean of 23 percent).&lt;/li&gt;
&lt;li&gt;Decreases the share in informal wage employment by &lt;strong&gt;0.29 percentage points&lt;/strong&gt; (a 2.9 percent decrease from the mean of 10 percent).&lt;/li&gt;
&lt;li&gt;Has no effect on overall wage employment, unemployment, or self-employment — the formalization effect is a reallocation from informal to formal jobs, not net job creation.&lt;/li&gt;
&lt;li&gt;Reduces formal sector wages by &lt;strong&gt;0.6 percent&lt;/strong&gt;, with no effect on informal wages.&lt;/li&gt;
&lt;li&gt;Increases the number of formal establishments by &lt;strong&gt;1.6 percent&lt;/strong&gt; and the number of formal jobs by &lt;strong&gt;2 percent&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Raises gross firm entry by &lt;strong&gt;2.8 percent&lt;/strong&gt; and gross firm exit by &lt;strong&gt;3 percent&lt;/strong&gt; (higher churn), with effects stable or slightly increasing through 2017–18.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These firm-creation effects are not driven by migrants starting businesses: migrants are not more likely to be business owners in high-immigration municipalities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Short-Run Findings.&lt;/strong&gt; Using year-on-year specifications with the PNAD (2001–2009), the authors replicate the results in the prior literature: municipalities receiving more migrants experience a reduction in formal wage employment, with no change in informal employment or non-employment — so the share of informal jobs rises. These short-run informality-increasing effects coexist with the long-run formalization results, and are not a sample artifact (the long-run results are unchanged when restricted to the same 700 PNAD municipalities).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanism — Downward Nominal Wage Rigidity (DNWR).&lt;/strong&gt; DNWR in the formal sector is the key mechanism reconciling short- and long-run effects. In Brazil, nominal wage cuts were illegal, and the national minimum wage rose regularly during the 2000s. Two municipality-level DNWR proxies are used: (i) the Kaitz index (national minimum wage / municipality median wage in 2000); (ii) the share of workers with negative year-on-year nominal wage changes (from RAIS, 1997–2000). In municipalities with higher DNWR: the positive formalization effects of immigration are smaller or fully muted; non-employment increases; and formal wages decline less. These cross-sectional patterns echo the Harris-Todaro-Fields prediction, and are consistent with DNWR being more binding in the short run (when nominal rigidities bind) than in the long run (when inflation and worker turnover allow real wage adjustment).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The paper develops and estimates a dynamic model of firm dynamics and informality, extending the canonical Hopenhayn framework with (i) two margins of informality — the extensive margin (whether a firm registers) and the intensive margin (whether a registered formal firm hires workers formally) — and (ii) heterogeneous long-run productivity parameters (nu) that generate firm-specific life-cycle growth profiles. Formal firms cannot revert to informality; informal firms can formalize by paying the cost differential between formal and informal entry costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Counterfactuals.&lt;/strong&gt; A simulated once-and-for-all 10 percent labor supply shock (approximately the 80th percentile of observed immigration shocks) produces: a 4.1 percent decline in the share of informal workers (IV: 7.5 percent); a 16.1 percent increase in formal firms (IV: 21.1 percent); and a 3.4 percent wage decline (IV: 5 percent). Of the increase in formal firms, &lt;strong&gt;40 percent&lt;/strong&gt; is accounted for by formalization of previously informal firms, highlighting the stepping-stone role of informality that a static or dual-economy model would miss. Average firm productivity declines by 1.4 percent due to worsening firm composition (the share of formal firms in the lowest productivity quartile rises by more than 4 percentage points). A counterfactual that nearly eliminates the extensive margin of informality (via steep enforcement costs) raises total output by 8.6 percent vs. 7 percent in the baseline shock, and increases average firm productivity by 2.1 percent vs. a decline of 1.4 percent — at the cost of displacing the least productive informal firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions.&lt;/strong&gt; Results pertain to internal (not international) migration; drought-induced migrants do not change the skill composition of the labor force at destination, justifying a homogeneous worker assumption. The formalization effects hold for migrants and non-migrants separately, and for high- and low-skilled workers separately. The model is calibrated to the average urban destination in Brazil, not a spatial general equilibrium.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-identification-strategy-and-what-are-the-key-threats-to-validity-the-authors-address"&gt;Q1. What is the identification strategy, and what are the key threats to validity the authors address?&lt;/h3&gt;
&lt;p&gt;The authors use a shift-share IV where shifts are drought shocks at origin municipalities (constructed from SPEI x crop calendar x crop revenue share, accumulated over 2000–2010) and shares are pre-2000 migration networks. Threats addressed: (i) pre-trends — no evidence of differential pre-trends in firm outcomes between 1997–98 and 1999–2000; (ii) demand channel — controlling for local drought shocks and distance-weighted neighboring shocks leaves results unchanged; (iii) capital reallocation — adding a bank-network-based shift-share control (following prior literature) does not change results; (iv) agricultural processing linkages — results hold after excluding agricultural firms and food/beverage/tobacco manufacturers; (v) migration persistence — controlling for baseline log population and 1995–2000 migration rates leaves results unchanged. The commodity-price-shock instrument provides an independent validation, yielding similar results despite near-zero cross-origin correlation with drought shocks and only -0.3 correlation across destinations.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-authors-reconcile-the-long-run-formalization-result-with-the-short-run-informality-increasing-result-and-what-role-does-dnwr-play"&gt;Q2. How do the authors reconcile the long-run formalization result with the short-run informality-increasing result, and what role does DNWR play?&lt;/h3&gt;
&lt;p&gt;DNWR is the key mechanism. Nominal wage cuts are illegal in Brazil&amp;rsquo;s formal sector, and the minimum wage rose through the 2000s, making DNWR binding especially in the short run. In the year-on-year specification (PNAD, 2001–2009), immigration reduces formal wage employment with no change in informal employment, raising the informal share — consistent with prior literature. Over the decade, inflation and worker turnover permit real formal wage adjustment, enabling formal sector expansion. Cross-sectional heterogeneity confirms this: in municipalities with above-median Kaitz index or below-median share of negative wage changes, the formalization effect of immigration is smaller or zero, and non-employment rises — precisely the Harris-Todaro-Fields prediction for rigid-wage environments.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-exact-magnitude-of-the-firm-level-effects-and-how-persistent-are-they"&gt;Q3. What is the exact magnitude of the firm-level effects and how persistent are they?&lt;/h3&gt;
&lt;p&gt;A one-percentage-point increase in the immigration rate increases formal establishments by 1.6 percent, formal jobs by 2 percent, firm entry by 2.8 percent, and firm exit by 3 percent — all decadal effects (1999–2000 to 2011–12). Effects on firms, entry, exit, and jobs remain stable or slightly increasing through 2017–18 as estimated using RAIS panel data, with no evidence of pre-trends (effects near zero in 1997–98 to 1999–2000 period). The effect on firm-level average wages is negative (consistent with the worker-level wage effect) but not statistically significant.&lt;/p&gt;
&lt;h3 id="q4-are-migrants-themselves-the-source-of-new-formal-firm-creation"&gt;Q4. Are migrants themselves the source of new formal firm creation?&lt;/h3&gt;
&lt;p&gt;No. The authors directly test and reject this channel. Migrants are not more likely to be business owners — either of small firms (fewer than 5 employees) or larger firms (6 or more employees) — in municipalities that receive more immigration. The increase in formal firm entry is driven by non-migrants responding to cheaper labor.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-two-margins-of-informality-in-the-model-and-why-does-the-intensive-margin-matter-for-the-migration-formality-nexus"&gt;Q5. What are the two margins of informality in the model, and why does the intensive margin matter for the migration-formality nexus?&lt;/h3&gt;
&lt;p&gt;The extensive margin is whether a firm registers formally (firm-level binary). The intensive margin is whether a formally registered firm hires workers without formal labor contracts (worker-level, within formal firms). The intensive margin is crucial because it links formal firms to migrants: newly arrived migrants may take informal jobs within formal firms, allowing formal firm creation to respond to the immigration shock even before the labor market fully formalizes. In the transition dynamics after an immigration shock with DNWR, new formal firms tend to be small and lower-productivity, and hire a substantial fraction of their workforce informally — so labor informality hovers near its initial level for several years even as firm informality declines quickly.&lt;/p&gt;
&lt;h3 id="q6-what-fraction-of-the-increase-in-formal-firms-in-the-counterfactual-comes-from-stepping-stone-formalization-versus-new-formal-entry"&gt;Q6. What fraction of the increase in formal firms in the counterfactual comes from stepping-stone formalization versus new formal entry?&lt;/h3&gt;
&lt;p&gt;In the baseline 10 percent labor supply counterfactual, approximately &lt;strong&gt;40 percent&lt;/strong&gt; of the increase in the number of formal firms comes from formalization of previously informal firms across their life cycles. The remaining 60 percent comes from new formal firm creation. A static framework would miss the stepping-stone channel entirely and substantially underestimate total formalization.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-models-calibration-pin-down-the-cost-structure-of-informal-vs-formal-firms"&gt;Q7. How does the model&amp;rsquo;s calibration pin down the cost structure of informal vs. formal firms?&lt;/h3&gt;
&lt;p&gt;The model is calibrated using a two-step minimum distance procedure. First-step parameters include the persistence of formal firms&amp;rsquo; productivity process (estimated from RAIS: rho_f = 0.92), and statutory tax rates (payroll tax tau_w = 0.375; revenue VAT tau_y = 0.293). Second-step parameters (12 total, including entry costs, exogenous death rates, productivity dispersion, and cost-function curvatures for both margins of informality) are estimated by minimizing the distance between simulated and observed moments from RAIS (2003 cross-section for static moments; 2000–2011 panel for growth moments) and ECINF (informal firms with up to 5 employees, 2003). Key calibrated values: formal entry costs are more than twice informal entry costs and correspond to over 30 times the 2003 monthly national minimum wage; the informal sector exogenous death rate (delta_i = 0.148) is more than twice the formal rate; productivity variance and persistence are similar across sectors.&lt;/p&gt;
&lt;h3 id="q8-what-happens-to-firm-productivity-and-output-per-worker-in-the-long-run-counterfactual"&gt;Q8. What happens to firm productivity and output per worker in the long-run counterfactual?&lt;/h3&gt;
&lt;p&gt;Average firm productivity declines by 1.4 percent despite lower informality. The composition of formal firms worsens: the share of firms in the lowest productivity quartile rises by more than 4 percentage points, while the share in the top quartile falls by about 3 percentage points. Total output and tax revenues increase (7 and 8.6 percent, respectively), but both decline in per capita terms. The authors note these are likely lower bounds because the model assumes no technological differences between formal and informal sectors and no differential capital access.&lt;/p&gt;
&lt;h3 id="q9-what-does-the-enforcement-counterfactual-reveal-about-the-dual-role-of-informality"&gt;Q9. What does the enforcement counterfactual reveal about the dual role of informality?&lt;/h3&gt;
&lt;p&gt;When the extensive margin of informality is nearly shut down (by making the informal cost function very steep), a 10 percent labor supply shock produces: output increase of 8.6 percent (vs. 7 percent with informality present); average firm productivity increase of 2.1 percent (vs. decline of 1.4 percent); much higher tax revenues due to greater formality. However, this comes at the cost of a sizable reduction in total firm count as the least productive informal firms are displaced. This illustrates the dual role: in the short run, the informal sector acts as an employment buffer and stepping-stone, which is more important when formal wage rigidity is stronger; but in the long run, it dampens aggregate economic benefits from immigration by sheltering low-productivity firms.&lt;/p&gt;
&lt;h3 id="q10-do-the-results-hold-for-both-migrants-and-non-migrants-and-across-skill-levels"&gt;Q10. Do the results hold for both migrants and non-migrants, and across skill levels?&lt;/h3&gt;
&lt;p&gt;Yes. Appendix results show similar employment and wage effects for migrants and non-migrants separately, though formal wage declines are more pronounced for non-migrants. Results are also similar for high- and low-skilled workers — which the authors attribute to the fact that drought-induced migration does not change the skill composition of the workforce at destination (confirmed empirically). Price-shock-induced migrants differ: they are more likely to be young and male, and do change workforce composition, providing a different set of compliers that strengthens external validity.&lt;/p&gt;
&lt;h3 id="q11-how-does-the-paper-relate-to-the-startup-deficit-literature-on-demographic-decline"&gt;Q11. How does the paper relate to the &amp;ldquo;startup deficit&amp;rdquo; literature on demographic decline?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s findings are the mirror image of the US startup deficit literature, which argues that demographic slowdown reduced firm entry, labor reallocation, and employment growth. The magnitudes are comparable in scale: the US startup deficit corresponds to a 5-percentage-point decline in firm entry between 1980 and 2012, while the rural-urban migration shocks studied here produce first-order effects on firm entry of similar or larger magnitude (2.8 percent per percentage point of immigration rate), suggesting labor supply growth is a primary driver of formal firm dynamics in both directions.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Downward Nominal Wage Rigidity (DNWR).&lt;/strong&gt; In the paper&amp;rsquo;s usage, the binding constraint that formal sector wages cannot be cut in nominal terms — in Brazil, both legal prohibition of nominal wage cuts and a rising national minimum wage. DNWR is the paper&amp;rsquo;s central mechanism explaining why immigration increases informality in the short run (wages cannot adjust) but reduces it over the decade (inflation and turnover permit real adjustment). Measured empirically via the municipality-level Kaitz index (national minimum wage / local median wage) and via the share of workers with negative year-on-year nominal wage changes in RAIS.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extensive Margin of Informality.&lt;/strong&gt; Whether a firm is registered with the government (formal) or not (informal). In the model, informal firms can avoid taxes but face a size-increasing cost of informality and the option to formalize by paying the difference in entry costs. This margin captures the firm&amp;rsquo;s legal registration status.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Intensive Margin of Informality.&lt;/strong&gt; Whether a formally registered firm hires individual workers with or without formal labor contracts (signed work booklet, carteira de trabalho). Formal firms face increasing costs for informal hiring but exploit this margin for lower-cost labor, especially when small or young. This margin is critical because it links formal firms to migration-induced informal labor supply and allows formal firms to absorb migrants before full wage adjustment occurs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Stepping-Stone Role of Informality.&lt;/strong&gt; The paper&amp;rsquo;s term for the dynamic channel through which the informal sector facilitates transitions to formality for both firms and workers. Informal firms accumulate productivity experience and formalize when productivity crosses the formalization threshold; informal workers within formal firms transition to formal contracts as firms grow. In the counterfactuals, 40 percent of the increase in formal firms following a labor supply shock is attributable to this channel. The stepping-stone role is most valuable during the short-run period of formal wage rigidity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Shift-Share Instrumental Variable.&lt;/strong&gt; The identification design combining pre-existing migration network shares (fraction of prior migrants to destination d from each origin o, computed 1995–2000) with exogenous push shocks at origin (drought shocks or commodity price shocks). The instrument predicts which destination municipalities receive more migrants based purely on exogenous origin-level shocks, purging the endogeneity from migrants self-selecting into prosperous cities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Minimum Comparable Area (MCA).&lt;/strong&gt; The paper&amp;rsquo;s geographic unit of analysis: a harmonized aggregation of Brazilian municipalities whose administrative borders changed during the study period, yielding 3,548 stable units covering all urban destinations studied. The authors call these &amp;ldquo;municipalities&amp;rdquo; for convenience.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Harris-Todaro-Fields Framework.&lt;/strong&gt; The theoretical benchmark against which the paper&amp;rsquo;s results are compared — the view (from Harris and Todaro 1970 and Fields) that rural-urban migration increases urban unemployment or informality because DNWR prevents the formal sector from absorbing migrants, who instead queue for formal jobs or enter the informal sector. The paper shows this prediction holds in the short run and in high-DNWR municipalities, but not in the long run where real wage adjustment occurs.&lt;/p&gt;</description></item><item><title>Screening and Segmenting: A Consumer Surplus Perspective</title><link>https://macropaperwarehouse.com/papers/screening-and-segmenting-a-consumer-surplus-perspective/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/screening-and-segmenting-a-consumer-surplus-perspective/</guid><description>&lt;p&gt;Bergemann, Heumann, and Wang study consumer surplus when a monopolist simultaneously engages in second-degree price discrimination (screening consumers within each market segment through quality-differentiated menus) and third-degree price discrimination (offering different menus across segments). The central question is which market segmentation maximizes aggregate consumer surplus, and under what conditions any segmentation benefits consumers at all.&lt;/p&gt;
&lt;p&gt;The model features a monopolist selling vertically differentiated goods of quality q at strictly convex cost c(q) to a continuum of buyers with privately known values v drawn from an aggregate market m*. A segmentation is any decomposition of m* into submarkets, each receiving a profit-maximizing screening menu. The seller observes segment identity but not individual values. The problem of finding the consumer-optimal segmentation is, on its face, an optimization over distributions of distributions — an infinite-dimensional object.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central methodological contribution is a dramatic dimensional reduction. Theorem 1 establishes that the maximum consumer surplus achievable by any segmentation equals the maximum of the expected local information rent, u(v,h) = h·Q(v−h), over all inverse hazard rate functions h satisfying a majorization constraint h ≺ h* (where h* is the aggregate market&amp;rsquo;s inverse hazard rate). The local information rent captures both the extensive margin (h measures the mass of higher-value buyers per unit of value-v buyers who earn rent from v&amp;rsquo;s allocation) and the intensive margin (Q(v−h) is the quality allocated to value v, decreasing in h as distortion increases). The two margins trade off: raising h widens the base of rent-earning buyers but worsens allocative distortion, making u(v,h) hump-shaped in h with an interior maximizer h̄(v).&lt;/p&gt;
&lt;p&gt;The consumer-optimal segmentation has a striking structural property: every buyer of a given value v receives the same quality in every segment in which they appear, even though the monopolist could in principle offer different qualities across segments. Prices, however, differ across segments for identical buyers. This holds because the optimal segmentation is always a uniform segmentation — one in which the inverse hazard rate hm(v) is equalized across all segments containing value v.&lt;/p&gt;
&lt;p&gt;Under log-concavity of both aggregate demand (equivalently, a non-increasing aggregate inverse hazard rate h*(v), satisfied by uniform, normal, logistic, and exponential distributions) and the supply function Q(v) (equivalent to c&amp;rsquo;&amp;rsquo;&amp;rsquo;(q)q/c&amp;rsquo;&amp;rsquo;(q) ≥ −1, satisfied by all power cost functions), the optimal segmentation takes a transparent two-regime form (Proposition 3): for values below a threshold v̂ where h*(v̂) = h̄(v̂), the inverse hazard rate is reduced to h̄(v) by concentrating low-value buyers; for values above v̂, the aggregate market is left unchanged. The resulting segments are nested convex intervals [vm, v̄], all sharing the same upper bound v̄, with pricing differing across segments only by a quality-independent base price Tm that increases with vm (Theorem 2).&lt;/p&gt;
&lt;p&gt;Corollary 3 delivers the sharpest policy-relevant finding: under log-concave demand and supply, zero segmentation is optimal — any segmentation harms consumers — if and only if h*(v̲) ≤ h̄(v̲) at the lowest value v̲. For iso-elastic costs c(q) = q^γ/γ (γ &amp;gt; 1), this becomes η*(v̲) ≤ γ/(1−γ), where η*(v̲) is the aggregate demand elasticity at the bottom of the distribution. When demand is sufficiently elastic relative to supply, the monopolist&amp;rsquo;s screening already provides near-optimal consumer rents and no redistribution of buyers across segments can improve them. More elastic supply (lower γ) shrinks the set of markets where zero segmentation is optimal (Proposition 4, Zγ&amp;rsquo; ⊂ Zγ for γ&amp;rsquo; &amp;lt; γ); more inelastic supply (higher γ) expands it, and in the limit γ → ∞ zero segmentation is suboptimal only when the aggregate allocation itself is efficient.&lt;/p&gt;
&lt;p&gt;For iso-elastic costs, the optimal segmentation assigns each segment a Pareto distribution below v̂ with shape parameter α = γ/(γ−1), and the aggregate market above v̂ (Corollary 1). Each segment&amp;rsquo;s demand elasticity equals the constant γ/(1−γ) below v̂ and the aggregate elasticity above (Corollary 2): the supply elasticity 1/(γ−1) determines how elastic demand must be made within segments to counteract monopoly distortions. The paper also extends the framework to adverse selection (where seller cost rises with buyer type), with the full reduction to inverse hazard rate optimization preserved when the rate of increase in adverse selection satisfies τ&amp;rsquo;&amp;rsquo;(v)v/τ&amp;rsquo;(v) ∈ [0,1] (Proposition 5).&lt;/p&gt;
&lt;p&gt;Q: What is the local information rent and why is it central?
A: The local information rent is u(v,h) = h·Q(v−h), where h is the inverse hazard rate at value v and Q is the inverse marginal cost (supply) function (equation 9). The factor h captures the extensive margin — the mass of higher-value buyers per unit of value-v buyers who earn rent from v&amp;rsquo;s quality allocation — while Q(v−h) captures the intensive margin — the quality allocated to v via the virtual value v−h, which falls as h rises. Because u is hump-shaped in h, there is an interior rent-maximizing inverse hazard rate h̄(v) for each value. Lemma 2 establishes that in every regular market, total consumer surplus equals the integral of u(v,hm(v))dFm(v), so the entire segmentation problem reduces to choosing h.&lt;/p&gt;
&lt;p&gt;Q: What is the majorization constraint and what does it exactly characterize?
A: The majorization constraint h ≺ h* requires that for all v ∈ V, the integral from v̲ to v of [h*(t) − h(t)]dF*(t) ≥ 0 (equation 18). Proposition 1 shows that for any segmentation σ, the average inverse hazard rate hσ must satisfy hσ ≺ h*. A partial converse holds: given h ≺ h* under regularity conditions, a uniform segmentation implementing h exists. The constraint is strictly weaker than the pointwise bound h ≤ h* available in the binary case because it permits h to exceed h* at some values (dilution) provided it falls sufficiently below h* at higher values (concentration) to maintain the cumulative inequality.&lt;/p&gt;
&lt;p&gt;Q: What are concentration and dilution, and how do they interact?
A: Concentration gathers buyers of a given value into fewer segments, lowering their inverse hazard rate below h*(v). Dilution raises the inverse hazard rate of value v by placing v in segments where immediately higher values are missing — creating gaps in the support — thereby increasing the support increment Δm(v) and hence hm(v) (equation 12). Dilution at v requires that values just above v have already been concentrated elsewhere to create the gaps; concentration thus enables dilution, linking the two tools. With only binary values, only concentration is available; with a continuum, dilution can strictly expand achievable consumer surplus by permitting h to exceed h* at low values.&lt;/p&gt;
&lt;p&gt;Q: What does Theorem 1 establish and why is it a major simplification?
A: Theorem 1 states that the maximum consumer surplus over all segmentations of m* equals the maximum of ∫u(v,h(v))dF*(v) over all h satisfying the majorization constraint h ≺ h* (equation 25). The original problem maximizes over distributions on the infinite-dimensional space of probability measures on V; the reduced problem is a standard optimal control problem over a single real-valued function h: V → R+, amenable to Karush-Kuhn-Tucker methods and often yielding closed-form solutions. Furthermore, every optimal segmentation is a uniform segmentation implementing some h solving the reduced problem, so the reduction is exact. The optimal h always satisfies regularity (h&amp;rsquo;(v) ≤ 1), meaning v − h(v) is non-decreasing, which ensures segments in the optimal uniform segmentation are themselves regular.&lt;/p&gt;
&lt;p&gt;Q: What is the structural property of consumer-optimal segmentations regarding quality across segments?
A: In any consumer-optimal segmentation, every buyer of value v receives the same quality in every segment in which they appear (the uniform quality property following from Theorem 1). This holds because the optimal inverse hazard rate h(v) is equalized across segments (uniform segmentation), and quality in a regular market is qm(v) = Q(v − hm(v)), which depends on the market only through hm(v). Prices, however, differ across segments for identical buyers: the monopolist does not redesign its product line across segments but adjusts only quality-independent base prices. This is counterintuitive because nothing in the monopolist&amp;rsquo;s problem requires quality uniformity — it emerges purely from the consumer surplus maximization.&lt;/p&gt;
&lt;p&gt;Q: What conditions guarantee the simple two-regime convex segmentation structure?
A: Log-concavity of aggregate demand — equivalently, h*(v) non-increasing in v, satisfied by uniform, normal, logistic, and exponential families — and log-concavity of the supply function Q(v), equivalent to c&amp;rsquo;&amp;rsquo;&amp;rsquo;(q)q/c&amp;rsquo;&amp;rsquo;(q) ≥ −1, together guarantee the structure of Proposition 3 and Theorem 2. Under these conditions, h̄(v) is strictly increasing in v (log-concave supply) while h*(v) is decreasing (log-concave demand), so they cross exactly once at v̂. The optimal h equals h̄(v) below v̂ and h*(v) above. Only concentration (not dilution) is ever used because log-concave supply makes u concave in h and log-concave demand ensures monotone ordering of marginal local information rents across values, so the binding majorization constraint becomes the pointwise constraint at the bottom.&lt;/p&gt;
&lt;p&gt;Q: What is the structure of convex segmentations and their menus (Theorem 2)?
A: Under log-concave demand and supply, the consumer-optimal segmentation consists of segments m with absolutely continuous supports [vm, v̄] for varying lower bounds vm ≤ v̂, all sharing the same upper bound v̄ (Part 1 of Theorem 2). Pricing across these segments differs only by a quality-independent base price Tm that is increasing in vm — more concentrated segments (lower vm) face a lower base price and carry higher information rents — while the quality menu p(q) is uniform across segments (Part 2). Equivalently, the monopolist offers nested menus all sharing the same efficient upper bound quality Q(v̄), differing in how far down the menu is extended and in the price of the lowest offered quality.&lt;/p&gt;
&lt;p&gt;Q: What do Corollaries 1 and 2 say for iso-elastic cost functions?
A: With iso-elastic cost c(q) = q^γ/γ (γ &amp;gt; 1) and log-concave demand, the consumer-optimal segmentation assigns each segment a Pareto distribution with shape parameter α = γ/(γ−1) below the threshold v̂, and the aggregate distribution above v̂ (Corollary 1). This delivers a constant demand elasticity of γ/(1−γ) within each segment below v̂, matching the aggregate market&amp;rsquo;s elasticity above v̂ (Corollary 2). The Pareto shape — and thus the degree of demand manipulation — is determined entirely by the supply elasticity 1/(γ−1): more elastic supply (lower γ) mandates a higher shape parameter α and more elastic within-segment demand to counteract larger monopoly distortions.&lt;/p&gt;
&lt;p&gt;Q: When is zero segmentation optimal, and what is the precise elasticity condition?
A: Under log-concave demand and supply, zero segmentation is optimal if and only if h*(v̲) ≤ h̄(v̲) — the aggregate inverse hazard rate at the lowest value already lies at or below its rent-maximizing level (Corollary 3). Since h* is decreasing under log-concavity, this condition at v̲ implies it holds everywhere, so the designer cannot improve rents at any value. For iso-elastic cost, the condition becomes η*(v̲) ≤ γ/(1−γ): aggregate demand elasticity at the bottom must be at least as large in magnitude as one plus the supply elasticity. For a Pareto aggregate distribution with shape parameter α, zero segmentation is optimal when α ≥ γ/(γ−1).&lt;/p&gt;
&lt;p&gt;Q: How does supply elasticity govern the scope for beneficial segmentation (Proposition 4)?
A: Proposition 4 establishes that for iso-elastic cost, the set of markets Zγ where zero segmentation is optimal is strictly nested increasing in γ: for any γ&amp;rsquo; &amp;lt; γ, Zγ&amp;rsquo; ⊂ Zγ. More elastic supply (lower γ) amplifies monopoly distortions and enlarges the set of markets where segmentation benefits consumers; more inelastic supply (higher γ) makes quality provision rigid, reducing segmentation&amp;rsquo;s scope. In the limit γ → ∞ (approaching unit demand), zero segmentation is suboptimal only if the aggregate allocation is already efficient — but this limit also means very inelastic supply, so the potential benefits from segmentation have shrunk toward zero simultaneously.&lt;/p&gt;
&lt;p&gt;Q: How does this paper compare to and depart from Haghpanah and Siegel (2023)?
A: Haghpanah and Siegel (2023) showed that in generic markets with a finite number of goods, some segmentation always improves consumer surplus relative to the aggregate market. This paper shows that with a continuum of qualities, this universal improvement result fails: Corollary 3 identifies a large, non-degenerate class of markets satisfying Haghpanah and Siegel&amp;rsquo;s genericity conditions where zero segmentation is optimal for consumers. The discrepancy arises because the log-concave supply condition (equation 27) is violated in finite-good environments — Haghpanah and Siegel explicitly provide a counterexample showing their result fails with a continuum of goods. This paper characterizes exactly when the finite-good gains vanish as the quality space becomes continuous, providing the precise elasticity conditions.&lt;/p&gt;
&lt;p&gt;Q: What changes and what is preserved when extending to adverse selection?
A: In the adverse selection specification, buyer net value v is private and the seller&amp;rsquo;s cost per unit is τ(v) − v, increasing in v when τ&amp;rsquo;(v) &amp;gt; 1. The local information rent becomes w(v,h) = u(v, τ&amp;rsquo;(v)·h), where adverse selection enters by amplifying the effective inverse hazard rate by τ&amp;rsquo;(v) (equation 40). Proposition 5 confirms that the full reduction to majorization-constrained optimization over h goes through, and the optimal segmentation features more elastic within-segment demand when adverse selection is more severe. The reduction requires τ&amp;rsquo;&amp;rsquo;(v)v/τ&amp;rsquo;(v) ∈ [0,1] (equation 39), bounding the rate of increase of adverse selection severity; if this fails, the key inequality (35) driving the optimality of uniform segmentations may break down.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications for regulation of price discrimination?
A: The results imply that blanket restrictions on market segmentation may harm consumers by preventing welfare-enhancing price discrimination in markets where demand is sufficiently inelastic relative to supply (the region outside the zero-segmentation condition). In markets satisfying η*(v̲) ≤ γ/(1−γ), allowing segmentation yields no consumer benefit, so restrictions are harmless to consumers. The key policy-relevant primitives are demand and supply elasticities, which are in principle measurable. The findings also imply that the welfare effects of data-driven personalized pricing depend critically on the interaction between consumer heterogeneity (demand shape) and cost structure (supply elasticity), rather than on the degree of segmentation per se.&lt;/p&gt;
&lt;p&gt;Local information rent: u(v,h) = h·Q(v−h), the total consumer surplus generated per unit mass of buyers at value v as a function of the inverse hazard rate h. The factor h is the extensive margin (mass of higher-value buyers per unit of value-v buyers who earn rent) and Q(v−h) is the intensive margin (quality allocated to v via the virtual value v−h). It is hump-shaped in h with interior maximizer h̄(v), and the segmentation problem reduces entirely to maximizing its expectation.&lt;/p&gt;
&lt;p&gt;Inverse hazard rate hm(v): in a continuous market, (1−Fm(v))/fm(v); generalized to accommodate atoms and support gaps (equation 12). It simultaneously determines the virtual value ϕm(v) = v − hm(v) (governing allocative distortion) and the scaled mass of higher-value buyers per unit of value-v buyers (governing the extensive margin of rents). The dual role requires both a continuum of qualities and endogenous segmentation.&lt;/p&gt;
&lt;p&gt;Majorization constraint h ≺ h*: for all v, the cumulative integral of [h*(t)−h(t)]dF*(t) from v̲ to v is non-negative (equation 18). It is the exact characterization of inverse hazard rate functions achievable by some segmentation of m*, strictly weaker than the pointwise bound h ≤ h* of the binary case because it permits h to exceed h* at some values (dilution) provided it falls sufficiently below h* at higher values (concentration).&lt;/p&gt;
&lt;p&gt;Uniform segmentation: a segmentation in which every buyer of value v faces the same inverse hazard rate hm(v) = hσ(v) in every segment containing v (equation 22). Theorem 1 establishes that every consumer-optimal segmentation is uniform; this class converts the double integral over segments and values into a single integral against F*, enabling the dimensional reduction of Theorem 1.&lt;/p&gt;
&lt;p&gt;Concentration and dilution: the two tools by which segmentation modifies inverse hazard rates. Concentration gathers buyers of a given value into fewer segments, lowering hm(v) below h*(v). Dilution raises hm(v) above h*(v) by placing value v in segments where immediately higher values are absent, creating support gaps. Dilution requires prior concentration of adjacent higher values, so the two tools are linked; under log-concave demand and supply, only concentration is used in the optimal segmentation.&lt;/p&gt;
&lt;p&gt;Convex segmentation: a segmentation whose constituent segments have nested convex interval supports [vm, v̄] all sharing the same upper bound v̄, with varying lower bounds vm. This is the consumer-optimal structure under log-concave demand and supply (Theorem 2). For iso-elastic cost, each segment below the threshold v̂ corresponds to a Pareto distribution with shape parameter α = γ/(γ−1) determined by cost convexity γ.&lt;/p&gt;
&lt;p&gt;Zero-segmentation condition: the condition under which no segmentation can improve consumer surplus over the aggregate market. Under log-concave demand and supply with iso-elastic cost c(q) = q^γ/γ, it is η*(v̲) ≤ γ/(1−γ): aggregate demand elasticity at the lowest value must be at least as large in magnitude as one plus the supply elasticity (Corollary 3). When this holds, any redistribution of buyers across segments strictly reduces consumer surplus.&lt;/p&gt;</description></item><item><title>Spread too thin: The impact of lean inventories</title><link>https://macropaperwarehouse.com/papers/spread-too-thin-the-impact-of-lean-inventories/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/spread-too-thin-the-impact-of-lean-inventories/</guid><description>&lt;p&gt;This paper investigates the macroeconomic consequences of widespread just-in-time (JIT) inventory management, documenting a fundamental trade-off: JIT raises firm profitability and reduces micro-level volatility in normal times, but renders the economy significantly more vulnerable to large unanticipated shocks.&lt;/p&gt;
&lt;p&gt;The empirical analysis draws on a novel dataset of approximately 200 publicly listed U.S. manufacturing firms for which the author identifies JIT adoption years using narrative records from SEC filings and historical news archives. Firm-level balance sheet data come from Compustat Fundamentals Annual (1980–2018), merged with county-level weather event data from NOAA. Four empirical facts are documented. First, JIT adoption is associated with a 13% decrease in inventory-to-sales ratios and a 9% increase in sales. Second, JIT adopters experience a roughly 7% decline in sales and employment growth volatility. Third, JIT adopters are approximately 25–30% more cyclical than non-adopters: a 1% increase in GDP growth predicts an additional 0.47 percentage point increase in JIT firm sales growth above the non-adopter baseline of roughly 1.6%. Fourth, a weather disaster predicts an additional 3% decline in JIT firm sales and employment relative to non-JIT firms.&lt;/p&gt;
&lt;p&gt;To explain and quantify these facts, the author builds and structurally estimates a dynamic general equilibrium model with a distribution of heterogeneous final goods firms that differ in idiosyncratic productivity, inventory holdings, and JIT adoption status. Materials must be drawn from inventory stocks; new orders are subject to stochastic fixed order costs. JIT producers draw from a first-order stochastically dominated order cost distribution relative to non-JIT firms. Adopting JIT requires an upfront sunk cost and a smaller continuation cost thereafter. The model is estimated via simulated method of moments (SMM) targeting 11 moments (adoption frequency, inventory-to-sales ratios, covariances, and spike frequencies for both firm types), with nine parameters to be estimated.&lt;/p&gt;
&lt;p&gt;In the estimated model steady state, JIT adoption delivers a 9–10% increase in output, a 40% decline in the aggregate inventory-to-sales ratio (close to the observed 35% decline in nonfarm inventories-to-final-sales from 1980 to 2018), a 1.3% increase in firm value, a 1.3% increase in measured TFP, and a welfare gain of 1.43% in consumption equivalent terms. These gains arise because lower order costs allow firms to better align material input use with realized productivity, smoothing inventory cycles.&lt;/p&gt;
&lt;p&gt;The vulnerability side is quantified through an unanticipated supply disruption calibrated to match the 3.4% drop in real U.S. GDP between 2019 and 2020. In response, the JIT economy experiences an approximately 0.40 percentage point excess output contraction relative to the no-JIT counterfactual, amounting to roughly 13–15% more output lost. The mechanisms are stockouts — firms that fully exhaust their inventories and cannot produce — and hoarding behavior, whereby firms that retain some inventory draw stocks down more slowly to preserve buffers, reducing material input use. Both channels reduce production relative to the counterfactual. The excess output loss is estimated at approximately $100 billion, comparable to state and local government allocations under the CARES Act.&lt;/p&gt;
&lt;p&gt;JIT nevertheless remains welfare-improving even under this shock. For a social planner to prefer a no-JIT world, the negative productivity shock to the intermediate goods sector would need to be nearly 14% — an order of magnitude larger than the calibrated COVID-19 shock. The trade-off is robust across alternative order cost distributions, parameterizations, partial anticipation scenarios, and stockout cost specifications.&lt;/p&gt;
&lt;p&gt;Q: What is the central trade-off identified by the paper?
A: JIT adoption reduces fixed order costs, enabling firms to place smaller and more frequent orders, which raises sales, reduces micro-level volatility, and increases firm value and welfare in normal times. However, because JIT firms hold fewer inventories, an unexpected aggregate shock increases the likelihood of stockouts and hoarding behavior, producing a deeper aggregate output contraction relative to an economy without JIT. Firms do not internalize the prospect of large shocks when making their private adoption decisions, generating the externality at the heart of the trade-off.&lt;/p&gt;
&lt;p&gt;Q: How does the paper measure JIT adoption, and how large is the sample?
A: The author constructs an adoption dummy for approximately 200 publicly listed manufacturing firms by exhaustively reviewing SEC filings and historical news archives for keywords including &amp;ldquo;JIT,&amp;rdquo; &amp;ldquo;just-in-time,&amp;rdquo; &amp;ldquo;lean manufacturing,&amp;rdquo; &amp;ldquo;pull system,&amp;rdquo; and &amp;ldquo;zero inventory.&amp;rdquo; Each document is individually analyzed to confirm the adoption year and to ensure it refers to the firm itself rather than its suppliers. More than half of observed adopters in the sample adopt prior to 1990, and nearly all adopt before 2000. The final Compustat-linked sample covers about 5,017 unique manufacturing firms from 1980 to 2018.&lt;/p&gt;
&lt;p&gt;Q: What are the firm-level efficiency gains from JIT adoption?
A: JIT adoption is associated with a 13% decrease in inventory-to-sales ratios and a 9% increase in sales; the corresponding standard deviation changes are –16% and +4%, respectively. Adopters also experience a roughly 7% decline in both sales and employment growth volatility, and a 5% increase in sales per worker relative to non-JIT firms. JIT firms additionally show a roughly 20% standard deviation reduction in squared forecast errors, indicating improved predictability of profitability.&lt;/p&gt;
&lt;p&gt;Q: How much more cyclical are JIT firms relative to non-JIT firms?
A: A 1% increase in GDP growth is associated with approximately a 1.6% increase in sales growth for non-adopters; JIT adopters experience an additional 0.47 percentage point increase above this baseline, making them roughly 25–30% more cyclical. This elevated cyclicality is estimated from variation external to the firm and reflects the heightened sensitivity of lean producers to aggregate demand fluctuations.&lt;/p&gt;
&lt;p&gt;Q: How are JIT firms affected by local weather disasters?
A: On average, a weather disaster predicts an additional 3% decline in JIT firm sales and employment relative to non-JIT firms. Using upstream supply chain linkages from Compustat Segment files, a unit increase in the average number of disasters hitting a firm&amp;rsquo;s suppliers predicts a 7–8% decline in firm sales and employment, with a similar excess decline for JIT firms. These results parallel the strategy in Barrot and Sauvagnat (2016).&lt;/p&gt;
&lt;p&gt;Q: What is the model structure, and how does the JIT adoption decision work?
A: The model features a representative household, a representative intermediate goods firm producing materials with capital and labor, and a continuum of heterogeneous final goods firms that differ in idiosyncratic productivity (AR(1) in logs), inventory holdings, and JIT adoption status. Each period has three stages: adoption decision, order decision (conditional on stochastic fixed order cost draw), and production decision. JIT producers draw order costs from a distribution first-order stochastically dominated by the non-JIT distribution, meaning JIT firms face systematically lower expected order costs. Adoption requires an upfront sunk cost c_s; maintaining JIT requires a smaller continuation cost c_f (estimated at slightly more than one-third of c_s), generating hysteresis: conditional on being an adopter, the probability of remaining one is estimated at 94%.&lt;/p&gt;
&lt;p&gt;Q: What moments are targeted in the SMM estimation, and how well does the model fit?
A: Eleven moments are targeted to identify nine parameters: the empirical adoption frequency, plus five moments each for JIT and non-JIT firms (mean inventory-to-sales ratio, the covariance matrix of inventory-to-sales ratios and log sales delivering three moments, and the frequency of positive inventory-to-sales ratio spikes exceeding 0.20). The model successfully fits targeted moments; non-targeted regression coefficients reproduce a quantitatively similar reduction in inventory-to-sales ratios after adoption, a comparable increase in sales among adopters, and reductions in firm volatility of 4–5% versus 6–7% in the data.&lt;/p&gt;
&lt;p&gt;Q: What are the estimated key structural parameters?
A: The upper support of the order cost distribution among non-adopters is estimated to be an order of magnitude larger than that of adopters, implying JIT firms place orders about 45% smaller than non-JIT firms. The estimated carrying cost is about 20% of inventory value. The estimated share of non-adopters in the model&amp;rsquo;s steady state implies a mass of JIT establishments of approximately 0.40. The technology parameters for the idiosyncratic productivity process are consistent with prior estimates in the structural firm dynamics literature.&lt;/p&gt;
&lt;p&gt;Q: What are the steady-state aggregate gains from JIT adoption in the model?
A: Relative to a counterfactual economy with no JIT option, the estimated model delivers a 9–10% increase in output, a 40% decline in the aggregate inventory-to-sales ratio (close to the observed 35% decline from 1980 to 2018), a 1.3% increase in firm value, a 1.3% increase in measured TFP, and a welfare gain of 1.43% in consumption equivalent terms. The TFP gain arises because lower order costs reallocate resources toward high marginal product producers at the aggregate level.&lt;/p&gt;
&lt;p&gt;Q: How is the unanticipated disaster calibrated, and what are its effects in the JIT versus no-JIT economies?
A: The disaster is an unanticipated negative shock to aggregate productivity in the intermediate goods sector, calibrated to match the 3.4% drop in real U.S. GDP between 2019 and 2020. In response, the JIT economy experiences approximately a 0.40 percentage point excess output contraction relative to the no-JIT counterfactual, amounting to roughly 13–15% more output lost. This excess loss equals approximately $100 billion, comparable to CARES Act allocations to state and local governments.&lt;/p&gt;
&lt;p&gt;Q: What are the two mechanisms through which JIT amplifies the disaster shock?
A: The first mechanism is stockouts: because JIT firms hold fewer inventories, an unexpected spike in order costs makes them more likely to fully exhaust their existing stocks, leaving them with no material inputs and forcing them to forgo production entirely. The second mechanism is hoarding: firms that do not fully stock out face a higher shadow value of inventories and cut back on material input use to draw inventories down more slowly, reducing output even without a full stockout. Both mechanisms reduce material input utilization in the JIT economy, causing a sharper drop in sales relative to the counterfactual.&lt;/p&gt;
&lt;p&gt;Q: Is JIT still welfare-improving when the COVID-19 shock is accounted for?
A: Yes. A social planner comparing welfare across steady states would not prefer to eliminate JIT even accounting for the deeper crisis it generates. For the planner to prefer a no-JIT world, the negative productivity shock to the intermediate goods sector would need to be nearly 14% — an order of magnitude larger than the calibrated 3.4% shock. This implies that the welfare gains from JIT in normal times substantially outweigh the welfare costs of the deeper recession under a COVID-19-scale shock.&lt;/p&gt;
&lt;p&gt;Q: How does the paper relate to the Great Moderation literature?
A: JIT adoption is credited in prior work (McConnell and Perez-Quiros, 2000; Blanchard and Simon, 2001; Kahn et al., 2002) as contributing to the roughly 35% reduction in the aggregate inventory-to-sales ratio between 1980 and 2018 and to the broader decline in macroeconomic volatility. The estimated model is consistent with this: JIT adoption reduces firm-level volatility and, in the steady state, implies a reduction in aggregate inventory-to-sales ratios close to the observed magnitude. However, the paper documents that the same forces that smooth normal-times fluctuations amplify unanticipated large shocks.&lt;/p&gt;
&lt;p&gt;Q: What robustness checks does the paper conduct?
A: The paper considers alternate parameterizations (all robustly show the micro-macro trade-off), larger disaster sizes calibrated to UK and France 2020 contractions (JIT economy contracts ~10% vs. ~8.7%, a ~15% larger contraction), partial anticipation (a sizable excess output drop persists because the left tail of firm outcomes is truncated at zero profits), stockout costs (trade-off remains with ~1.2% firm value gain and ~10% excess contraction), and an alternative right-skewed beta order cost distribution (firm value gain rises to 1.8%, trade-off remains). An alternative CUSUM-based measure of JIT adoption identifying approximately 560 firms produces qualitatively similar empirical results.&lt;/p&gt;
&lt;p&gt;Q: What is the subsample estimation finding on adoption costs over time?
A: Comparing 1980–1989 and 1990–2018 subsamples, the upfront sunk cost of JIT adoption estimated from the 1980s sample is about 26% higher than in the later subsample, implying it has become easier to initiate JIT production over time. Steady-state output rises by about 3.4% in the 1990–2018 period relative to 1980–1989, and the excess output contraction under the disaster shock is about 15% relative to the 1980s counterfactual, close to the baseline estimate.&lt;/p&gt;
&lt;p&gt;Just-in-Time (JIT) Production: A lean inventory management philosophy that minimizes the time between orders by committing to smaller and more frequent orders from suppliers, reducing costs of managing large material purchases and storing idle stocks; in the model, JIT is operationalized as drawing order costs from a distribution first-order stochastically dominated by the non-JIT distribution.&lt;/p&gt;
&lt;p&gt;Stockout: The condition in which a final goods firm enters a period with no inventories (s = 0) and chooses not to place an order, leaving it without any material inputs and forcing it to forgo production entirely for that period.&lt;/p&gt;
&lt;p&gt;Hoarding (in the disaster context): The behavior of firms that, facing a higher shadow value of inventories during an unexpected shock, cut back on material input use in order to draw down existing inventory stocks more slowly, preserving buffers at the cost of reduced current production.&lt;/p&gt;
&lt;p&gt;Fixed Order Cost: A stochastic, labor-denominated cost that a firm must pay each period in which it places a materials order; JIT adopters face a systematically lower distribution of these costs, enabling more frequent ordering at smaller quantities.&lt;/p&gt;
&lt;p&gt;Adoption Sunk Cost: The one-time upfront cost c_s a non-adopter must pay to initiate JIT status, which exceeds the continuation cost c_f paid by existing JIT firms to maintain their status; the gap between these costs generates hysteresis in the adoption decision.&lt;/p&gt;
&lt;p&gt;Simulated Method of Moments (SMM): The structural estimation procedure used to identify model parameters by minimizing the weighted distance between model-simulated moments and their empirical counterparts; here applied with 11 targeted moments to identify 9 parameters in an overidentified system.&lt;/p&gt;
&lt;p&gt;Micro-Macro Trade-off: The paper&amp;rsquo;s central finding that individual firms rationally adopt JIT for private profitability gains (1.3% increase in firm value, 1.43% welfare gain), while the aggregate economy becomes more fragile to unanticipated shocks (roughly 13–15% deeper output contraction) because firms do not internalize the systemic vulnerability created by economy-wide lean inventories.&lt;/p&gt;</description></item><item><title>Staffing agencies and in-house bargaining</title><link>https://macropaperwarehouse.com/papers/staffing-agencies-and-in-house-bargaining/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/staffing-agencies-and-in-house-bargaining/</guid><description>&lt;p&gt;This paper asks whether a labor market with search-and-matching frictions and firms producing under decreasing returns to labor is better characterized by in-house hiring with intra-firm wage bargaining (Stole-Zwiebel) or by an alternative arrangement in which intermediaries — &amp;ldquo;staffing agencies&amp;rdquo; — search for and employ workers and then rent them to producing firms on a frictionless, perfectly competitive market. The paper&amp;rsquo;s second and central question is what happens when firms can choose their optimal combination of the two arrangements simultaneously.&lt;/p&gt;
&lt;p&gt;The model is static. There are Z homogeneous firms with production function F(n) satisfying F&amp;rsquo;&amp;rsquo;(n) &amp;lt; 0, N homogeneous workers, and a standard concave constant-returns-to-scale matching function M = m(V, N). Firms can post vacancies, workers search, and Nash bargaining with worker bargaining weight β determines wages. The analysis is conducted with fully general production and matching functions throughout, deviating to specific functional forms (Cobb-Douglas matching, power production function F(n) = An^α) only when needed to illustrate a particular efficiency result. All main results hold for both directed and random search.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Comparing the two polar arrangements (Theorem 3.1).&lt;/strong&gt; When all hiring is in-house (Stole-Zwiebel), equilibrium firm size n^SZ, aggregate employment Zn^SZ, labor market tightness θ^SZ, and the equilibrium wage w^SZ are all strictly higher than their counterparts under full staffing-agency employment (n^SA, Zn^SA, θ^SA, w^SA). The mechanism is that under in-house hiring with decreasing returns, a worker&amp;rsquo;s threat to leave raises the marginal product — and hence the wage — of remaining workers, giving workers additional bargaining leverage. Firms respond by over-employing in-house hires to dilute each worker&amp;rsquo;s marginal product and thus moderate wages. This over-employment raises vacancy posting and tightness, which in general equilibrium bids up wages despite each firm&amp;rsquo;s individual wage-moderation motive.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Efficiency (Theorem 3.2).&lt;/strong&gt; Under the standard Hosios condition — worker bargaining weight β equals the elasticity η(θ) of the matching function with respect to vacancies — the staffing-agency equilibrium achieves the social planner&amp;rsquo;s optimum (θ^SA = θ*), while the in-house equilibrium posts strictly too many vacancies (θ^SZ &amp;gt; θ^SA = θ*). The in-house arrangement can be optimal for some β &amp;gt; η when workers&amp;rsquo; bargaining power is sufficiently high (Theorem 3.3, proved for Cobb-Douglas matching and power production function), because the over-employment incentive then counteracts the externality from underprovision of vacancies.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The main result: staffing agencies dominate (Theorem 3.4).&lt;/strong&gt; When firms choose their profit-maximizing combination of in-house hires n^SZ and rented staffers n^SA, the unique equilibrium has n^SZ = 0 and n^SA &amp;gt; 0 — firms use only staffers. The key is Lemma 3.1: renting one additional staffer reduces the wage paid to in-house workers by more than does hiring one additional in-house worker (formally, ∂w^SZ/∂n^SA &amp;lt; ∂w^SZ/∂n^SZ). This asymmetry arises because staffers cannot leave during intra-firm bargaining breakdowns — they remain regardless — so each additional staffer tightens the firm&amp;rsquo;s fallback position more effectively than an additional in-house hire. With continuous labor, any positive mass of in-house workers leaves residual scope for further wage moderation through staffers, so the firm always finds it profitable to convert the last in-house hire to a staffer. The corner solution n^SZ = 0 is thus the unique equilibrium. With discrete labor, a firm would be indifferent between exactly one and zero in-house workers.&lt;/p&gt;
&lt;p&gt;The paper also notes that this staffing-agency arrangement is formally equivalent to the &amp;ldquo;labor packer&amp;rdquo; or intermediate-good setup widely used in applied macroeconomics (e.g., Gertler, Sala, and Trigari 2008) to avoid Stole-Zwiebel complications, providing a micro-foundation for that modeling convention.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does in-house hiring with decreasing returns to labor generate higher wages and employment than the staffing-agency arrangement?&lt;/strong&gt;
A: Under decreasing returns, if a worker&amp;rsquo;s wage negotiation breaks down and the worker leaves, the marginal product of the remaining n−1 workers rises. This gives each in-house worker additional bargaining leverage beyond the standard β parameter. To counteract this, firms over-employ in-house hires to keep the marginal product low. In general equilibrium this raises tightness θ^SZ &amp;gt; θ^SA, which in turn raises wages w^SZ &amp;gt; w^SA even though each individual firm&amp;rsquo;s motive was wage moderation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the formal basis for the in-house wage equation, and what does it depend on?&lt;/strong&gt;
A: Following Stole and Zwiebel&amp;rsquo;s stability condition, the continuous-labor wage for a firm with n workers is w(n) = (1−β)b + n^(−1/β) ∫₀ⁿ z^((1−β)/β) F&amp;rsquo;(z) dz. The wage depends on the entire distribution of marginal products over [0, n], not merely on the marginal product at n. In the special case of a power production function, the integral yields an explicit power function in n.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: When is the staffing-agency equilibrium socially efficient?&lt;/strong&gt;
A: Under the standard Hosios condition β = η(θ*), the staffing-agency equilibrium attains exactly the planner&amp;rsquo;s tightness (θ^SA = θ*), because bargaining in staffing agencies is standard — the worker&amp;rsquo;s outside option does not affect other workers&amp;rsquo; wages and so the usual efficiency characterization applies. The in-house equilibrium then strictly over-posts vacancies (θ^SZ &amp;gt; θ*).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Can the in-house equilibrium ever be socially optimal?&lt;/strong&gt;
A: Yes, but only under specific parameter conditions. Theorem 3.3 shows that with Cobb-Douglas matching (η constant) and a power production function F(n) = An^α, there exists a threshold β̂ ∈ (η, 1) at which θ^SZ = θ*. The intuition is that strong worker bargaining power creates a vacancy-underprovision problem; the over-employment incentive under in-house hiring then partially corrects it. The functional form restriction is made for expositional convenience; the core logic is general.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is Lemma 3.1 and why is it the key to the main result?&lt;/strong&gt;
A: Lemma 3.1 states that, given any positive number of in-house hires n^SZ &amp;gt; 0, renting one additional staffer reduces the wage paid to in-house workers by more than does hiring one additional in-house worker: ∂w^SZ/∂n^SA &amp;lt; ∂w^SZ/∂n^SZ. This is proved by showing the relevant integral in the difference (∂w^SZ/∂n^SA − ∂w^SZ/∂n^SZ) is negative for n^SZ &amp;gt; 0 given F&amp;rsquo;&amp;rsquo; &amp;lt; 0.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does renting an additional staffer moderate in-house wages more than hiring an in-house worker?&lt;/strong&gt;
A: An in-house worker who is hired can, in principle, leave during a bargaining breakdown, triggering renegotiation all the way down to zero in-house workers and driving the firm&amp;rsquo;s fallback to zero profit. A rented staffer cannot leave; at minimum, all rented staffers remain in production regardless of in-house bargaining outcomes. Each additional staffer thus raises the firm&amp;rsquo;s floor payoff in bargaining by more than an additional in-house hire does, generating stronger wage moderation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Why does Theorem 3.4 produce a corner solution rather than an interior mix?&lt;/strong&gt;
A: Because labor is treated as a continuous input, any strictly positive mass n^SZ &amp;gt; 0 of in-house workers leaves the marginal in-house worker with positive bargaining leverage through the threat-to-leave mechanism. The firm can always improve its bargaining position by converting that marginal in-house worker to a staffer. This margin is present no matter how small n^SZ is, so the only equilibrium is n^SZ = 0. In discrete labor the firm would be indifferent between exactly one and zero in-house hires.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What happens to labor market tightness when both arrangements coexist?&lt;/strong&gt;
A: In the mixed equilibrium the tightnesses for in-house and staffer jobs must be equal in equilibrium (θ^SZ = θ^SA). If one tightness were higher, workers would prefer that job type (higher wage and higher probability of finding it), but firms would reduce vacancy posting there (costlier to fill), automatically equalizing tightness. This equilibration occurs even though in equilibrium vacancy posting for in-house jobs goes to zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Do the results require directed search or hold under random search as well?&lt;/strong&gt;
A: The results hold under both directed and random search. Appendix 3.C establishes that with random search and a single pooled matching function M = m(V^SZ + V^SA, N), the unique equilibrium also features n^SZ = 0. The directed-search assumption is made without loss of generality.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does this paper imply for the applied macroeconomics literature&amp;rsquo;s &amp;ldquo;labor packer&amp;rdquo; modeling convention?&lt;/strong&gt;
A: The paper provides a formal micro-foundation for the labor-packer or intermediate-good approach used in New Keynesian DSGE models (e.g., Gertler, Sala, and Trigari 2008) to sidestep Stole-Zwiebel bargaining. In that literature, a &amp;ldquo;wholesale firm&amp;rdquo; or &amp;ldquo;packer&amp;rdquo; searches for workers and sells their services to final-goods firms under perfect competition — formally identical to the staffing-agency arrangement in this paper. Theorem 3.4 shows this arrangement is the unique equilibrium outcome of rational firm choice, so the shortcut is not merely convenient but theoretically grounded.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the empirical literature say about wage differentials between in-house and agency workers?&lt;/strong&gt;
A: Drenik et al. (2023), using Argentine administrative data linking temp agencies to user firms, estimate a significant wage premium for in-house hires relative to temp workers. This is consistent with the paper&amp;rsquo;s theoretical prediction that w^SZ &amp;gt; w^SA in the polar-case comparison (Theorem 3.1), though the paper itself presents no empirical estimation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What are the implications of the staffing-agency arrangement for measured labor shares?&lt;/strong&gt;
A: The paper notes that costs for staffers typically appear in firm accounts as intermediate input costs rather than labor costs. A shift from in-house hires to staffers therefore reduces measured labor costs and, because it also reduces value added (by more than the labor-cost reduction), lowers the measured labor share at the firm even when actual labor input and output are unchanged.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What future extensions do the authors identify as priorities?&lt;/strong&gt;
A: The authors flag three main extensions: (i) heterogeneous workers and firms, which could generate predictions about which firms use each hiring mode; (ii) worker effort/loyalty differences between in-house and agency workers that could make in-house hiring attractive ex post; and (iii) a frictional rental market for staffers or heterogeneous tasks within the firm, where insufficient staffer supply in certain sub-markets could restore a role for in-house hiring.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Staffing agency (in this paper&amp;rsquo;s sense):&lt;/strong&gt; An intermediary that posts vacancies on the frictional labor market, employs workers through standard Nash bargaining, and rents those workers one-for-one to producing firms on a frictionless, perfectly competitive market. The staffing agency is separated from the firm&amp;rsquo;s production decisions; its search activity has constant returns to scale.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;In-house hiring with Stole-Zwiebel bargaining:&lt;/strong&gt; A market arrangement in which the producing firm itself posts vacancies, employs workers, and conducts intra-firm Nash bargaining. Under decreasing returns to labor, the bargaining outcome for worker i depends on the firm&amp;rsquo;s payoff if that worker left, which in turn depends on wages paid to the remaining n−1 workers — generating a system of interdependent bargaining problems captured by the differential equation w(n) = (1−β)b + n^(−1/β) ∫₀ⁿ z^((1−β)/β) F&amp;rsquo;(z) dz.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wage-moderation incentive (over-employment):&lt;/strong&gt; Under in-house hiring, a firm has an incentive to hire more workers than a social planner would recommend, because additional workers reduce each worker&amp;rsquo;s marginal product and hence the wage the firm must pay. This incentive is present because decreasing returns mean a departing worker raises the marginal product of remaining workers, giving each in-house worker leverage over the firm.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Differential wage-moderation effect (Lemma 3.1):&lt;/strong&gt; The finding that, given any positive mass of in-house hires, renting one additional staffer reduces in-house wages by more than hiring one additional in-house worker (∂w^SZ/∂n^SA &amp;lt; ∂w^SZ/∂n^SZ). The asymmetry arises because staffers cannot leave during intra-firm bargaining breakdowns, so they provide a more effective floor to the firm&amp;rsquo;s fallback payoff.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hosios condition (as applied here):&lt;/strong&gt; The standard efficiency condition β = η(θ), where β is the worker&amp;rsquo;s Nash bargaining weight and η(θ) is the elasticity of the job-offer arrival rate with respect to tightness. When this condition holds, the staffing-agency equilibrium is socially optimal (θ^SA = θ*) and the in-house equilibrium is inefficient (θ^SZ &amp;gt; θ*).&lt;/p&gt;</description></item><item><title>State Capacity as an Organizational Problem</title><link>https://macropaperwarehouse.com/papers/state-capacity-as-an-organizational-problem/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/state-capacity-as-an-organizational-problem/</guid><description>&lt;p&gt;Mastrorocco and Teso study how the internal organization of a state evolves during national development, framing state capacity as an organizational — specifically a principal-agent — problem. Using a new micro-database covering the U.S. federal bureaucracy from 1817 to 1905, they ask: once rulers have incentives to build a state apparatus, how do they organize it to perform its functions across a vast territory, and what drives transitions between organizational forms?&lt;/p&gt;
&lt;p&gt;The dataset is constructed from every issue of the Official Register of the United States published between 1817 and 1905 (44 biennial volumes, 15,801 pages digitized). It records full name, state of birth, state of appointment, occupation, salary, department, office, and location for 304,410 unique federal employees across 810,942 employee-year observations. The authors reconstruct the bureaucracy&amp;rsquo;s four-layer hierarchy (department → office/bureau → division → local office), link employees over time to track careers, categorize all 11,930 occupation codes into five tiers, and geo-code 9,651 places of employment to 1890 county boundaries.&lt;/p&gt;
&lt;p&gt;The paper first documents three sets of descriptive facts. On growth: the federal workforce expanded very slowly before the 1860s and then rapidly, with geographic expansion accounting for none of state growth before 1859 but roughly 29% after. On location: state presence responded positively to local manufacturing activity (a one standard deviation increase in manufacturing employment share raises presence probability by 1.3 percentage points), but distance from Washington DC significantly attenuated this relationship in 1817–1859 and not in 1861–1905. On organization: before the 1860s, employee turnover was high and spiked sharply at presidential transitions (reaching 72% of employees departing in 1861), supervisors&amp;rsquo; departures strongly predicted subordinates&amp;rsquo; departures (a one-for-one supervisor exit raised subordinate turnover probability by 37% pre-1841), and managerial delegation outside DC was stagnant or declining. After the 1860s, turnover trended down (35% at the 1897 transition), the supervisor-subordinate career link weakened materially, and field managers tripled relative to the 1850s.&lt;/p&gt;
&lt;p&gt;The authors argue that high monitoring costs in the early century made trust-based, personalistic organization the second-best solution to principal-agent problems. The limited supply of sufficiently trusted individuals constrained geographic expansion, delegation, and total size. As railroad and telegraph networks lowered communication and transportation costs, monitoring capacity increased, enabling a transition to a Weberian bureaucracy no longer constrained by trust supply.&lt;/p&gt;
&lt;p&gt;The causal identification strategy uses the staggered expansion of the railroad network. For each county and decade (1820–1900), the authors compute the minimum-travel-time route from the county centroid to DC using Donaldson and Hornbeck (2016) data on railroads, steamboat waterways, coastal routes, and land routes. The specification includes county fixed effects, state-by-decade fixed effects, and controls for local railroad presence in the county and for the county&amp;rsquo;s market access, so the identifying variation comes from distant changes in the network that altered travel time to DC without directly affecting the county&amp;rsquo;s local economy or trade access.&lt;/p&gt;
&lt;p&gt;Results: a one standard deviation decrease in travel time to DC raises the probability of federal state presence by approximately 3 percentage points (about 8% of the mean), raises log employment similarly, raises the probability of observing a local managerial layer by approximately 3 percentage points (about 8% of the mean), and reduces employee turnover by approximately 2 percentage points (about 4% of the mean turnover rate). Placebo tests confirm that travel time to other major economic centers does not predict state presence. Telegraph network data (1845–1852, Wang 2020) yield consistent results. An additional test using the post-Civil War decline in Southern-born employee shares shows that better railroad connection to DC narrowed the North-South employment gap, consistent with monitoring substituting for trust-based selection.&lt;/p&gt;
&lt;p&gt;Scope conditions: the paper covers the civilian executive branch of the federal government, excluding the Postal Office, navy yards, and the engineer department; results are robust to restricting to states already in the union at the start of the sample, ruling out frontier-specific dynamics.&lt;/p&gt;
&lt;p&gt;Q: What is the central theoretical claim of the paper?
A: The paper argues that state capacity is fundamentally an organizational problem shaped by principal-agent constraints. When communication and transportation costs are high, the government cannot effectively monitor distant agents, so the second-best solution is to staff the bureaucracy with trusted individuals connected through personal networks. This personalistic form limits size and delegation because the supply of sufficiently trusted individuals is inherently scarce. Technological reductions in monitoring costs allow a transition to a Weberian bureaucracy based on procedural oversight rather than trust, removing the supply constraint on organizational growth.&lt;/p&gt;
&lt;p&gt;Q: What data source does the study rely on, and what time period does it cover?
A: The study draws on the Official Register of the United States, a biennial government publication listing all federal employees, digitized for every issue from 1817 to 1905. The resulting dataset includes 304,410 unique employees and 810,942 employee-year observations, with each record carrying name, state of birth, state of appointment, occupation, salary, department, office, location, and — through hierarchical reconstruction — position in a four-layer chain of command.&lt;/p&gt;
&lt;p&gt;Q: How did the size of the U.S. federal bureaucracy evolve over the nineteenth century?
A: Growth was slow before the 1860s. The first Register for 1817 listed 1,056 employees across 33 pages; the 1905 volume listed over 120,000 employees across 1,254 pages. Geographic expansion contributed zero to state growth before 1859 — the share of counties with any federal employee hovered around 15% from 1817 to 1859 — but contributed approximately 29% of growth after 1859, when county presence rose to 24% by 1871, 38% by 1881, and 61% by 1905.&lt;/p&gt;
&lt;p&gt;Q: What were the three sources of state growth, and how did their relative importance change?
A: The authors decompose growth into: (1) functions (new offices/bureaus), (2) geographic expansion (new counties), and (3) intensity (more employees per county-office pair). Before 1859, growth was entirely driven by functions (~40%) and intensity (~60%), with zero contribution from geographic expansion. After 1859, geographic expansion accounted for ~29%, intensity for ~32%, and functions for ~39% of growth.&lt;/p&gt;
&lt;p&gt;Q: How did employee turnover behave across the century, and what pattern emerges at presidential transitions?
A: Turnover trended upward through the late 1850s and then declined. During presidential transitions, the rate rose from 52–53% in 1841 and 1845 to 60–63% in 1849 and 1853 and peaked at 72% in 1861; it then fell to 55% in 1869, 44–48% in 1885/1889/1893, and 35% in 1897. Turnover was consistently lower in DC than in the field: controlling for year-bureau-position fixed effects, being employed in DC was associated with a 40% reduction in turnover probability.&lt;/p&gt;
&lt;p&gt;Q: How tight was the link between supervisors&amp;rsquo; and subordinates&amp;rsquo; careers, and how did it change?
A: Before 1841, moving from none to all supervisors leaving an organizational unit increased subordinate turnover probability by 37 percentage points. The effect was similar between 1841 and 1859, then dropped substantially to 22 percentage points in the following twenty-year period, and remained roughly constant after 1881. This pattern is consistent with the early bureaucracy relying on chains of personal trust that broke when a supervisor departed.&lt;/p&gt;
&lt;p&gt;Q: What evidence describes the evolution of delegation outside DC?
A: The number of field managers did not grow between 1817 and 1859 — it actually declined in the 1820s and was flat through the mid-1850s — and then tripled by 1905 relative to the 1850s level. The probability that workers in a local office had an additional managerial layer between them and DC was unchanged between pre-1841 and 1841–1859, increased by 5 percentage points between 1861 and 1881, and by 6 percentage points post-1881.&lt;/p&gt;
&lt;p&gt;Q: How does the paper measure monitoring capacity for the causal analysis?
A: The primary measure is travel time in hours from each county centroid to Washington DC, computed decade by decade (1820–1900) as the minimum-cost route across the available railroad network, steamboat waterways, coastal routes, and land routes, using data from Donaldson and Hornbeck (2016). A second, complementary measure is the number of telegraph connections between a county and DC using data from Wang (2020) for 1845–1852.&lt;/p&gt;
&lt;p&gt;Q: What is the identification strategy for the railroad analysis, and why are controls for local railroads and market access important?
A: The specification includes county fixed effects, state-by-decade fixed effects, an indicator for whether the county itself has railroad (LocalRailroad), and the county&amp;rsquo;s market access. County fixed effects mean beta is identified within-county from changes over time. Controlling for local railroad removes the direct correlation between local construction and local economic growth. Controlling for market access removes the effect of distant rail expansion on trade flows that raised agricultural land values and manufacturing activity. The remaining variation in travel time to DC — coming from distant network changes that altered the DC-county connection without affecting local conditions or broader trade access — is the identifying source.&lt;/p&gt;
&lt;p&gt;Q: What are the main quantitative effects of reduced travel time to DC?
A: A one standard deviation decrease in travel time to DC is associated with: (1) approximately 3 percentage point increase in the probability of federal state presence (~8% of the mean); (2) a similar magnitude increase in log employment conditional on presence; (3) approximately 3 percentage point higher probability of an additional managerial layer (~8% of the mean); and (4) approximately 2 percentage point reduction in employee turnover (~4% of the mean turnover rate).&lt;/p&gt;
&lt;p&gt;Q: How do placebo tests support the monitoring interpretation?
A: The authors show that, conditional on the same controls, travel times from a county to a set of other major economic centers are not associated with larger federal state presence. Since these other cities had no role as monitoring headquarters, the absence of an effect for them and the presence of an effect specifically for DC is consistent with the channel operating through the government&amp;rsquo;s ability to supervise agents from the capital, rather than through generic economic connectivity.&lt;/p&gt;
&lt;p&gt;Q: What does the telegraph evidence add, and what is its limitation?
A: Telegraph data (1845–1852, Wang 2020) show that counties with more telegraph connections to DC have larger state presence, more managerial delegation, and lower turnover, consistent with the monitoring mechanism. The limitation is that the authors have limited ability to address the endogeneity of telegraph network timing — the telegraph analysis is treated as corroborating evidence rather than the primary causal identification.&lt;/p&gt;
&lt;p&gt;Q: How do the Southern-born employee results illuminate the trust mechanism?
A: After the Civil War, the share of Southern-born federal bureaucrats fell sharply, consistent with reduced trust toward individuals from former Confederate states. However, counties that became better connected to DC via railroad expansion experienced a relative increase in the share of Southern-born employees. This shows that when monitoring costs fell, the government was willing to hire individuals from groups with lower baseline trust — monitoring substituted for trust as the mechanism ensuring agent performance.&lt;/p&gt;
&lt;p&gt;Q: Does federal state presence crowd out state and local government?
A: No. The presence of federal bureaucrats is positively correlated with the presence of state and local government employees at the county level, suggesting complementarity rather than substitution across levels of government.&lt;/p&gt;
&lt;p&gt;Q: What alternative mechanisms do the authors consider and how do they address them?
A: Three alternatives are discussed. First, demand shocks (Civil War debt repayment, industrialization) could explain the post-1860s expansion; the empirical specifications control for year fixed effects to absorb aggregate time-varying incentives, and the identification relies on differential cross-county variation in DC connectivity. Second, patronage as an electoral tool is consistent with spoils-driven turnover spikes but cannot explain why better-connected counties show lower turnover before civil service reform. Third, cognitive models of the firm (lower communication costs complement managerial problem-solving even without agency problems) could also predict the positive delegation result; the authors note they cannot empirically distinguish the monitoring and cognitive channels, and both may contribute.&lt;/p&gt;
&lt;p&gt;Q: What are the implications for developing countries today?
A: The authors suggest that their findings from nineteenth-century U.S. history may apply to understanding why modern Weberian bureaucracies remain elusive in many developing countries. Where communication infrastructure is limited and monitoring costs remain high, personalistic organizational forms based on trust networks may persist as constrained optima — not failures of will or design, but rational responses to structural conditions. Infrastructure investment that lowers monitoring costs could be a precondition for bureaucratic modernization.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Personalistic state organization&lt;/strong&gt;: The paper&amp;rsquo;s term for the organizational form that prevails when monitoring costs are high. It is characterized by staffing decisions based on personal character, moral reputation, and relationships of trust between principals and agents — and between supervisors and subordinates — rather than on formal procedural monitoring of performance. Frequent turnover at leadership transitions and constrained delegation are defining features, because the supply of trusted individuals is limited.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Weberian bureaucracy&lt;/strong&gt;: In the paper&amp;rsquo;s usage (following Weber 1978), a modern state organization defined by a fixed hierarchy of officials monitored through procedural rules rather than personal trust, lower turnover, and effective delegation of managerial power to geographically dispersed units. The paper treats this as the organizational form enabled by low monitoring costs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Monitoring capacity&lt;/strong&gt;: The principal&amp;rsquo;s (politicians in DC and their cabinets) ability to observe and evaluate the behavior of agents (federal employees) throughout the territory. In the paper&amp;rsquo;s operationalization, monitoring capacity is proxied inversely by travel time and communication cost between DC and the county: lower travel time and more telegraph connections mean higher monitoring capacity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Geographic expansion component&lt;/strong&gt;: One of three decomposed sources of state growth. Defined as the increase in state size attributable to the state becoming present in more county locations. This component contributed zero to federal growth before 1859 and approximately 29% of growth after 1859.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Employee turnover&lt;/strong&gt;: In the paper&amp;rsquo;s measurement, the share of employees who leave the federal bureaucracy in a given year. The paper distinguishes politically-driven spikes at presidential transitions — reaching 72% of employees in 1861 — from the secular trend, which rose through the late 1850s and then declined, reaching 35% by the 1897 transition.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Delegation of managerial power&lt;/strong&gt;: The probability that a local county office has an additional managerial layer between its workers and DC, rather than reporting directly to the bureau-level supervisor in Washington. The paper uses this as its measure of whether decision authority has been decentralized to the field.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Trust substitution&lt;/strong&gt;: The paper&amp;rsquo;s mechanism linking monitoring capacity to organizational form. In the absence of effective monitoring, principals substitute trust for oversight — selecting agents whose personal loyalty, moral character, or political alignment gives the principal confidence they will not shirk or defect. As monitoring costs fall, trust becomes less necessary as a screening device, and the trust-constrained supply limit on organizational growth is relaxed.&lt;/p&gt;</description></item><item><title>Take the Goods and Run: Contracting Frictions and Market Power in Supply Chains</title><link>https://macropaperwarehouse.com/papers/take-the-goods-and-run-contracting-frictions-and-market-power-in-supply-chains/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/take-the-goods-and-run-contracting-frictions-and-market-power-in-supply-chains/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;This paper studies the efficiency of self-enforced relational agreements in manufacturing supply chains when sellers have market power and contracts cannot be externally enforced. The setting is Ecuador, an upper-middle-income country with slow commercial courts (debt enforcement takes around two years even after a 2016 reform) and highly concentrated manufacturing markets (average Herfindahl-Hirschman Index of 0.6 for 6-digit economic codes, well above the 0.25 threshold used by the US Department of Justice to identify highly concentrated markets).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; How efficiently do long-term trade relationships operate, period by period, when the seller can price discriminate and the buyer can opportunistically default on trade-credit debt? Does seller market power worsen or mitigate enforcement-driven inefficiencies?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data.&lt;/strong&gt; The paper uses three Ecuadorian government administrative databases: (1) an electronic invoicing (EI) system covering all sales of 49 large manufacturing firms in textiles, pharmaceuticals, and cement products for 2016–2017, providing product-level unit prices, quantities, and payment method for each buyer-seller pair (median seller has 600 buyers); (2) the universe of firm-to-firm VAT transactions from 2008–2015, used to measure relationship age (censored at 9 years); and (3) annual financial statements providing variable costs to proxy marginal cost.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; The author develops a dynamic contracting model that embeds non-linear pricing with heterogeneous buyers (following Jullien 2000 and Attanasio-Pastorino 2020) into an infinitely repeated game with limited enforcement (following Martimort et al. 2017). The seller holds all bargaining power, commits to a long-term menu of prices and quantities, and finances every transaction through trade-credit. The buyer has a privately observed, fully persistent type (willingness to pay) and can opportunistically default after delivery — &amp;ldquo;take the goods and run&amp;rdquo; — at the cost of losing the future relationship. The seller uses the value of the ongoing relationship as the enforcement instrument. The paper solves the seller&amp;rsquo;s profit-maximization problem using a recursive Lagrangian approach, yielding a modified virtual-surplus condition that governs optimal quantity allocations as a function of current and past limited-enforcement Lagrange multipliers (LE multipliers).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Six motivating empirical facts&lt;/strong&gt; documented in the data: (1) New buyers are ~35% of pairs but account for only ~10% of total trade; relationships lasting nine or more years are less than 10% of pairs but generate over 30% of trade. (2) Trade-credit is used in approximately 65% of transactions in the first year and 70–75% in older relationships. (3) Quantities increase as relationships age. (4) A 10% increase in quantity purchased is associated on average with a 2% decrease in unit price (quantity discounts). (5) Conditional on quantity, older buyers pay up to 3% less; these price discounts appear only in trade-credit transactions, not in pay-in-advance transactions. (6) Approximately 40% of new relationships survive one additional year, 60% of relationships aged 1–3 years survive, and more than 75% of relationships aged four or more years survive.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Key structural finding.&lt;/strong&gt; Almost all new relationships have binding enforcement constraints. The estimated LE multiplier equals 1 (unconstrained) only for the top 1% of pairs at tenure 0. As relationships age, the constraint relaxes and quantities are backloaded — consistent with the seller making promises of higher future trade to incentivize current debt repayment.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Efficiency results.&lt;/strong&gt; New relationships operate at approximately 30% of the frictionless (first-best) surplus level. Efficiency rises to 60% at tenure 2, 75% at tenure 4, and over 80% at tenure 5. Aggregating across buyers weighted by efficient quantities: only 5% of sellers trade efficiently with new buyers, rising to 70% by tenure 2 and 84% in the long term. By sector, 68% of textiles, 88% of pharmaceutical, and 95% of cement-product sellers reach efficient aggregate output by tenure 5. Sellers capture approximately 80% of generated surplus; the median buyer captures around 25%, and the smallest buyers may capture less than 10%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Counterfactuals reveal a second-best interaction.&lt;/strong&gt; Fixing enforcement alone (Counterfactual a: non-linear pricing with perfect enforcement) raises surplus for 75% of buyers in the early tenures but reduces surplus for essentially all buyers in later tenures, because the threat of buyer default was the force compelling the seller to promise growing quantities over time. Fixing market power alone (Counterfactual b: uniform pricing with limited enforcement) collapses surplus to 0–40% of the baseline because the seller can no longer tailor dynamic incentives to each buyer&amp;rsquo;s enforcement constraint, causing a large share of buyers to be excluded from trade. Addressing both frictions simultaneously (Counterfactual c: uniform pricing with perfect enforcement) raises surplus for most buyers in early tenures but remains welfare-reducing for high types in later tenures; the aggregate effect depends critically on weighting: positive (~40% gain) when weighted by number of buyers, negative (surplus falls to ~58% of baseline) when weighted by quantities.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-central-theoretical-mechanism-by-which-limited-enforcement-leads-to-backloading-of-quantities-in-the-model"&gt;Q1. What is the central theoretical mechanism by which limited enforcement leads to backloading of quantities in the model?&lt;/h3&gt;
&lt;p&gt;The buyer can default after delivery because payment is post-delivery (trade-credit). To prevent this, the seller must ensure the buyer&amp;rsquo;s discounted future net returns exceed the current payment obligation. This creates a forward-looking enforcement constraint: the seller must credibly promise sufficiently large future quantities at lower prices. As a result, current quantities are distorted downward (the seller delays granting full trade volumes), but quantities increase over time as past promises become binding promise-keeping constraints. The optimal contract is therefore non-stationary: total surplus generated and the buyer&amp;rsquo;s net return both increase over time even without efficiency gains in production.&lt;/p&gt;
&lt;h3 id="q2-how-does-seller-market-power-interact-with-enforcement-frictions--does-it-worsen-or-improve-efficiency-relative-to-a-perfect-enforcement-benchmark"&gt;Q2. How does seller market power interact with enforcement frictions — does it worsen or improve efficiency relative to a perfect-enforcement benchmark?&lt;/h3&gt;
&lt;p&gt;The paper&amp;rsquo;s key finding is that market power and enforcement constraints act as partially offsetting frictions. Seller market power creates downward quantity distortions (the seller restricts supply to extract rents). Limited enforcement, however, compels the seller to promise growing quantities to prevent buyer default, which counteracts the market-power distortion. Thus, in older relationships, the enforcement constraint effectively disciplines the seller&amp;rsquo;s rent-extraction incentives, producing trade levels that approach the frictionless first-best. This is an instance of the theory of second-best: each friction partially offsets the other, so removing only one friction can reduce total welfare.&lt;/p&gt;
&lt;h3 id="q3-what-are-the-six-motivating-empirical-facts-and-why-do-they-rule-out-standard-alternative-explanations"&gt;Q3. What are the six motivating empirical facts and why do they rule out standard alternative explanations?&lt;/h3&gt;
&lt;p&gt;The six facts are: (1) heavy concentration of trade in long-established relationships; (2) widespread trade-credit even in new relationships; (3) quantities increase with relationship age; (4) quantity discounts within any age cohort; (5) older buyers pay lower prices conditional on quantity; (6) survival rates increase with quantity and relationship age. Alternative models — efficiency gains, learning, demand assurance, and supply-side enforcement issues — cannot jointly account for all six patterns under realistic assumptions. Critically, Fact 5 holds only in trade-credit transactions and not in pay-in-advance transactions, which supports limited enforcement (not learning or demand assurance) as the underlying mechanism.&lt;/p&gt;
&lt;h3 id="q4-how-is-the-model-identified-from-cross-sectional-data-on-prices-and-quantities-for-a-single-seller"&gt;Q4. How is the model identified from cross-sectional data on prices and quantities for a single seller?&lt;/h3&gt;
&lt;p&gt;Identification exploits two sources of variation. First, because the seller offers non-linear price menus that induce type revelation, cross-sectional variation in prices and quantities across buyers reveals their underlying private types. Second, for the highest-type buyer at tenure 0, the cumulative LE multiplier equals 1 by construction, so the gap between the observed marginal price and marginal cost directly reveals the current enforcement multiplier for that type; cross-sectional variation across high-type buyers then identifies the elasticity parameter β. Once β is pinned down, the multipliers for all types and tenures are recovered as unique solutions to ordinary differential equations, and buyer types are recovered semi-parametrically. The approach requires only cross-sectional data from one seller per year — no panel of individual buyers is needed.&lt;/p&gt;
&lt;h3 id="q5-what-are-the-estimated-magnitudes-of-the-marginal-product-of-capital-wedge-and-how-do-they-compare-to-related-studies"&gt;Q5. What are the estimated magnitudes of the marginal product of capital wedge, and how do they compare to related studies?&lt;/h3&gt;
&lt;p&gt;The paper finds a wedge between the buyer&amp;rsquo;s marginal product of capital (MPK) and the transaction price of 40% for the median new relationship and 34% for the median tenure-5 relationship. These wedges are smaller than the 80% gaps estimated for Indian firms by Banerjee and Duflo (2014), and larger than the average 6% gap calculated by Blouin and Macchiavello (2019) in the international coffee market. They are also much smaller than the 300–500% gaps estimated for Mexican micro-enterprises by McKenzie and Woodruff (2008), which is consistent with the buyers in this sample being substantially larger (median yearly sales of USD 200,000).&lt;/p&gt;
&lt;h3 id="q6-what-does-counterfactual-a--perfect-enforcement-with-non-linear-pricing--reveal-about-the-intertemporal-trade-off"&gt;Q6. What does Counterfactual (a) — perfect enforcement with non-linear pricing — reveal about the intertemporal trade-off?&lt;/h3&gt;
&lt;p&gt;Counterfactual (a) shows massive short-run gains for low and middle types: surplus at tenure 0 increases to 1,508% and 628% of baseline for the bottom 10th and median buyer percentile groups respectively. However, for higher types (top 25%), perfect enforcement is immediately welfare-reducing because these buyers are already trading near efficiently and the seller loses the incentive to grow quantities over time once default is not a threat. By tenure 3 and beyond, perfect enforcement reduces surplus for essentially all buyers. The aggregate effect is negative because high-type buyers, who trade larger volumes, bear larger losses in later tenures when those tenures are weighted by quantity.&lt;/p&gt;
&lt;h3 id="q7-why-does-uniform-pricing-with-limited-enforcement-counterfactual-b-perform-so-poorly"&gt;Q7. Why does uniform pricing with limited enforcement (Counterfactual b) perform so poorly?&lt;/h3&gt;
&lt;p&gt;Under uniform pricing, the seller cannot tailor the dynamic contract to each buyer&amp;rsquo;s individual enforcement constraint. Without individualized price-quantity menus, many buyers cannot credibly commit to repaying their debts — because the seller cannot offer a sufficiently personalized future stream of benefits — and are thus excluded from trade entirely. For instance, at tenure 0, 95.8% of the bottom-decile buyers and 64% of median buyers are excluded. The aggregate surplus under this regime reaches only 3–68% of baseline across different tenures and percentile groups. This implies that the seller&amp;rsquo;s price discrimination ability, while generating informational rents, also serves a second purpose: it allows each buyer&amp;rsquo;s specific enforcement constraint to be satisfied, enabling trade that would otherwise be infeasible.&lt;/p&gt;
&lt;h3 id="q8-what-do-the-sector-level-results-suggest-about-the-generalizability-of-the-main-findings"&gt;Q8. What do the sector-level results suggest about the generalizability of the main findings?&lt;/h3&gt;
&lt;p&gt;All six motivating empirical facts are consistent across the three industries studied (textiles, pharmaceuticals, and cement products). The efficiency patterns also appear in all three sectors, though with heterogeneous speeds of convergence. Pharmaceutical and cement-product sellers converge faster (88% and 95% efficient at tenure 5) than textiles sellers (68% efficient at tenure 5). The finding that relationships approach efficiency in the medium and long term holds in every industry analyzed, suggesting that the underlying mechanisms — limited enforcement and seller market power — are broadly operative rather than sector-specific.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-paper-establish-that-the-standard-non-linear-pricing-model-without-enforcement-constraints-does-not-explain-the-data"&gt;Q9. How does the paper establish that the standard non-linear pricing model without enforcement constraints does not explain the data?&lt;/h3&gt;
&lt;p&gt;The paper tests whether the LE multiplier at tenure 0 (G0) is statistically distinguishable from the null hypothesis of a standard non-linear pricing model (which would imply G0 = 1 for all buyers). Based on t-statistics from the estimated distribution of G0 across seller-year markets, the null of a standard model is rejected for 86% of the markets (seller-years) in the sample. Additionally, the dynamic price discounts conditional on quantity — which are the key signature of backloading — appear only in trade-credit transactions and not in pay-in-advance ones, ruling out alternative explanations such as learning about buyer quality or demand assurance.&lt;/p&gt;
&lt;h3 id="q10-what-are-the-models-main-limitations-and-how-do-they-affect-the-counterfactual-conclusions"&gt;Q10. What are the model&amp;rsquo;s main limitations and how do they affect the counterfactual conclusions?&lt;/h3&gt;
&lt;p&gt;The author flags three principal limitations. First, buyer types are assumed fully persistent due to data constraints (only two years of invoice-level data); a Markov type structure would require longer buyer-level panels. Second, the identification strategy relies on the seller&amp;rsquo;s first-order optimality conditions and cannot recover counterfactual dynamic quantities — the counterfactuals are therefore static comparisons of per-period surplus rather than full dynamic simulations. Third, if buyers have unobserved outside options, the counterfactual efficiency results may be biased, though the direction of the bias is uncertain and depends on the distribution of types and the curvature of the return function.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Limited enforcement constraint (LE-B).&lt;/strong&gt; The paper&amp;rsquo;s central friction: because payment is post-delivery, the buyer can default and keep the goods. In the model, the contract is &amp;ldquo;default-free&amp;rdquo; only if the buyer&amp;rsquo;s post-delivery payment is weakly less than the discounted value of all future truthful net returns. The constraint is binding when this condition is tight — the buyer is on the margin of defaulting. When binding, it forces the seller to reduce current tariffs and quantities (to lower the attractiveness of default) while promising higher future quantities (to raise the continuation value).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Limited enforcement Lagrange multiplier (LE multiplier), Gt(α).&lt;/strong&gt; The shadow price on the buyer&amp;rsquo;s enforcement constraint at tenure t for a buyer at quantile α. It takes values in [0,1], equals 1 only when the enforcement constraint is slack (unconstrained buyer), and equals zero for the lowest type at all tenures. In the paper&amp;rsquo;s framework, the entire trajectory of Gt(α) across tenures encodes the history of past enforcement promises and is the key object identified and estimated to recover the dynamic distortions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Backloading.&lt;/strong&gt; The equilibrium property whereby the total surplus generated by the relationship and the buyer&amp;rsquo;s net return both increase over time. The seller achieves this by initially restricting quantities and promising growing future allocations as an enforcement device. Formally, quantities increase over time if and only if enforcement constraints are relaxed (gt+1(q) ≤ gt(q)).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Modified virtual surplus.&lt;/strong&gt; The object that replaces ordinary virtual surplus (which appears in standard non-linear pricing models) in the seller&amp;rsquo;s first-order condition. It augments standard virtual surplus by adding shadow costs for current binding enforcement constraints and subtracting corrections for past enforcement promises. Optimal quantity allocations are determined by an inverse-markup rule applied to this modified virtual surplus.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Relational agreement / self-enforced relational contract.&lt;/strong&gt; An informal long-term agreement sustained purely through the repeated interaction between the parties, without access to third-party (court) enforcement. In this paper&amp;rsquo;s setting, the seller disciplines the buyer&amp;rsquo;s opportunism exclusively through the threat of relationship termination; no legal recourse is available or used in equilibrium.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Quantity discounts (non-linear pricing / wholesale quantity discounts).&lt;/strong&gt; Price schedules under which the unit price decreases with the quantity purchased, offered by a seller with market power. In the paper&amp;rsquo;s empirical setting, a 10% increase in quantity is associated with a 2% decrease in unit price, and these discounts appear at every relationship age. The model generates them as the incentive-compatibility requirement that ensures higher-type buyers truthfully reveal their demand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Trade-credit.&lt;/strong&gt; Seller financing of the transaction, in which goods are delivered before payment is received. In the Ecuadorian data, approximately 65% of first-year purchases and 70–75% of purchases in mature relationships are conducted via trade-credit. Because the seller bears the full cost of buyer default, trade-credit is the financial arrangement that gives rise to the limited enforcement constraint studied in the paper.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-best interaction of frictions.&lt;/strong&gt; The paper&amp;rsquo;s counterfactual finding that removing a single friction (either enforcement or market power) can reduce total welfare when both frictions are present simultaneously. This occurs because the two frictions partially offset each other: enforcement constraints discipline the seller&amp;rsquo;s monopoly distortions, and market power allows the seller to price-discriminate in ways that enable enforcement in the first place. Addressing both frictions simultaneously can improve welfare, consistent with the Lipsey-Lancaster theory of second-best.&lt;/p&gt;</description></item><item><title>Talent Hoarding in Organizations</title><link>https://macropaperwarehouse.com/papers/talent-hoarding-in-organizations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/talent-hoarding-in-organizations/</guid><description>&lt;p&gt;This paper provides the first empirical evidence of talent hoarding in organizations — the practice whereby managers deliberately suppress workers&amp;rsquo; internal mobility to retain productive team members, thereby serving their own performance-based compensation interests at the expense of firm-wide talent allocation. The research question is whether managers with misaligned incentives hoard talent, how this can be measured, and what consequences it have for worker career outcomes and organizational efficiency.&lt;/p&gt;
&lt;p&gt;The study uses personnel records from a large German manufacturing firm with over 200,000 employees worldwide, focused on more than 30,000 white-collar and management employees in Germany, covering over 300,000 employee-by-quarter observations from 2015 to 2018. This is supplemented by a manager survey (62% response rate, over 3,000 responses) and an employee survey (50% response rate, over 15,000 responses), plus the universe of internal job application and hiring data covering over 16,000 job openings and over 200,000 applicants.&lt;/p&gt;
&lt;p&gt;The conceptual framework formalizes talent hoarding as a moral hazard problem: managers observe worker productivity and are compensated based on team performance, but are tasked with identifying and developing talent for promotion. When a high-productivity worker leaves, team productivity falls. The framework predicts that hoarding intensity increases with worker productivity, team vulnerability to departures (smaller teams), and manager-level hoarding incentives (performance-related pay, low talent visibility).&lt;/p&gt;
&lt;p&gt;The key administrative measure of hoarding is the systematic gap between managers&amp;rsquo; private performance ratings (not shared outside the team) and public potential ratings (widely circulated within the firm). Managers who suppress potential ratings relative to what would be predicted given worker performance are interpreted as strategically reducing worker visibility. Managers with a 1 percentage point higher share of performance-related pay are 0.19 percentage points more likely to hoard talent; a one-person increase in team size reduces hoarding probability by 1.3 percentage points; and managers in low-visibility functional areas are 4.0 percentage points more likely to hoard. Survey-based hoarding measures yield directionally identical patterns.&lt;/p&gt;
&lt;p&gt;To identify causal effects on workers, the paper exploits quasi-random manager rotations. When a manager learns they will move to a different team — typically two to three quarters before the actual transition — their hoarding incentive ceases. This creates a temporary window of reduced hoarding. During this window, worker application rates increase by 2.3 percentage points, representing a 78% increase over the baseline application rate of 2.9%. An event study confirms flat pre-trends prior to the announcement period, supporting the identifying assumption.&lt;/p&gt;
&lt;p&gt;Using manager rotations as an instrument for worker applications, marginal applicants — those induced to apply only by the manager rotation — face a 49.1% likelihood of receiving a new position, compared to an average hiring likelihood of 27.6%. This positive selection implies that many deterred applicants would have been successful and that talent hoarding meaningfully degrades the quality of the internal applicant pool. Gender analysis reveals that women are 22% more likely to rely on manager career guidance and 26% more likely to prioritize preserving a good manager relationship. Marginal female applicants are more positively selected on education, past performance, and hiring probability for higher-level positions. The counterfactual reduction in the gender pay gap from eliminating talent hoarding is estimated at 86%.&lt;/p&gt;
&lt;p&gt;Scope conditions: the firm is a large European manufacturer with long average tenures (13 years), an application-based internal labor market, and centralized online job portal. Results apply most directly to white-collar and management employees in Germany. External validity is supported by comparisons to German workforce surveys and by the fact that 83% of top publicly listed German companies and half of 665 global organizations in industry surveys report talent hoarding as a significant organizational friction.&lt;/p&gt;
&lt;p&gt;Q: How is talent hoarding formally defined in this paper?
A: Talent hoarding is defined as actions taken by managers that lower the likelihood that a worker applies for and receives a promotion or any internal transfer outside the team. In the formal framework, a manager chooses hoarding intensity β ≥ 0, where β &amp;gt; 0 reduces the equilibrium probability that a worker gets promoted. The definition encompasses all forms of managerial action that reduce worker departure probability, including suppressing visibility, restricting access to trainings, explicit discouragement, and threats.&lt;/p&gt;
&lt;p&gt;Q: Why do managers have an incentive to hoard talent?
A: Managers are compensated based on team performance, so losing a high-productivity worker (whose replacement is a random draw from an outside distribution with expected productivity ᾱ) reduces team performance and thus manager compensation. The framework shows that when a worker&amp;rsquo;s productivity αi exceeds the expected productivity of an outside hire ᾱ, the manager optimally sets β* &amp;gt; 0. The cost of hoarding (parameterized as φm) is convex and varies across managers, capturing altruism, reputation risk, or detection probability.&lt;/p&gt;
&lt;p&gt;Q: What share of managers in the survey self-report talent hoarding?
A: 75% of managers reported that they sometimes find themselves in situations where they need to dissuade a team member from exploring opportunities in another department due to immediate team needs or performance goals. Additionally, 45% cite the risk of losing talent as a reason not to invest in employee career development, and 66% cite the need to prioritize short-term performance targets over long-term employee development.&lt;/p&gt;
&lt;p&gt;Q: How are misaligned incentives documented in the manager survey?
A: 55% of managers agree or strongly agree that talent development entails a conflict of interest because more developed workers are more likely to leave the team. While 96% believe their direct intervention has a large impact on workers&amp;rsquo; career development, only 36% perceive that impact to be valued by the firm as much as team performance impact. Similarly, 87% say talent development is a high-impact area for the firm, but only 40% believe a track record in talent development matters for their own compensation and promotion.&lt;/p&gt;
&lt;p&gt;Q: How is the administrative measure of talent hoarding constructed?
A: The measure is the residual from an OLS regression of a worker&amp;rsquo;s potential rating (a public signal of promotion readiness, widely circulated within the firm) on their performance rating (a private signal of current task performance, not shared outside the team) and worker characteristics including age, education, gender, and tenure. The manager-level measure is the average of these residuals across all workers and quarters under that manager. Managers in the top tercile (mean deviation above 0.1036) are classified as hoarding-prone.&lt;/p&gt;
&lt;p&gt;Q: Does the hoarding measure respond to the incentive proxies as predicted by the framework?
A: Yes. A 1 percentage point higher share of performance-related compensation is associated with a 0.19 percentage point increase in the probability of being classified as hoarding-prone (p = 0.000), corresponding to a 13 percentage point difference between the 90th and 10th percentiles of the financial incentive distribution. A one-person increase in team size reduces hoarding probability by 1.3 percentage points (p = 0.000), again a 13 percentage point difference across percentiles. Managers in low-visibility functional areas are 4.0 percentage points more likely to hoard (p = 0.002) relative to high-visibility areas.&lt;/p&gt;
&lt;p&gt;Q: Is the training-based hoarding measure consistent with the potential-rating measure?
A: Yes. A complementary measure based on managers restricting worker access to high-visibility in-person trainings yields nearly identical patterns: a 1 percentage point increase in performance-related pay increases hoarding probability by 0.20 percentage points (p = 0.000); a one-person increase in team size reduces it by 1.4 percentage points (p = 0.000); low-visibility areas increase hoarding by 2.98 percentage points (p = 0.021). The direction and economic magnitudes are highly similar across both administrative measures and the survey-based measures.&lt;/p&gt;
&lt;p&gt;Q: How are manager rotations used to identify causal effects on workers?
A: When a manager learns they will move to a different position — typically two to three quarters before the rotation — their incentive to hoard workers on their current team ceases. This creates a quasi-random window of reduced talent hoarding for workers on that team. An event study with worker and quarter fixed effects shows flat pre-trends in application rates beyond three quarters before the rotation, consistent with the identifying assumption that managers do not yet know about their rotation in that earlier window. Balance tests confirm workers exposed to rotations are observationally similar on demographics and past performance to non-exposed workers.&lt;/p&gt;
&lt;p&gt;Q: How large is the effect of manager rotations on worker applications?
A: Manager rotations increase worker application rates by 2.3 percentage points in the quarter of rotation, representing a 78% increase over the baseline application rate of 2.9%. The effect is transitory: application rates return to baseline within one quarter after the new manager settles in. The effect is not driven by managers taking subordinates with them (97% of applications are to positions outside both the current team and the manager&amp;rsquo;s new team).&lt;/p&gt;
&lt;p&gt;Q: Does the rotation effect vary with predicted hoarding intensity as the framework requires?
A: Yes. The rotation effect is larger for workers with higher productivity, those whose replacement would be costlier (consistent with the prediction that workers harder to replace face more hoarding), and those working under managers with lower utility costs of hoarding. The paper tests these cross-sectional predictions using continuous interactions between the rotation indicator and standardized proxies for hoarding intensity, and all patterns are consistent with the talent hoarding mechanism rather than alternative explanations.&lt;/p&gt;
&lt;p&gt;Q: How successful would the deterred applicants have been?
A: Marginal applicants — those induced to apply by the manager rotation who would not otherwise have applied, identified via IV assumptions — face a hiring probability of 49.1%, compared to the average hiring likelihood of 27.6% across all applicants. This large positive selection implies that a substantial share of deterred applicants would have been successful, and that talent hoarding meaningfully degrades the quality and quantity of the firm&amp;rsquo;s internal applicant pool and the firm&amp;rsquo;s ability to promote high-productivity workers.&lt;/p&gt;
&lt;p&gt;Q: Does talent hoarding have differential effects by gender?
A: Yes. Women are 22% more likely to place high value on preserving a good relationship with their manager and 26% more likely to rely on manager career guidance when making career decisions. Consistent with this, marginal female applicants are more positively selected on educational qualifications, past performance, and hiring probability for higher-level positions than marginal male applicants. When comparing potential earnings outcomes, both men and women would earn more in the absence of talent hoarding, but the larger earnings gains for women imply a counterfactual reduction in the gender pay gap of 86%.&lt;/p&gt;
&lt;p&gt;Q: What evidence supports external validity of the findings?
A: The firm&amp;rsquo;s employee demographics closely match those of large manufacturing firms in the German BiBB workforce survey across gender, age, citizenship, and marital status. The firm&amp;rsquo;s internal labor market design is standard for large German firms, where 83% of top publicly listed companies cite talent hoarding as a key organizational friction. Industry surveys also report that half of 665 global organizations report managers hoarding talent by discouraging worker mobility, and talent hoarding occurs through many of the same behaviors documented in this study.&lt;/p&gt;
&lt;p&gt;Q: How does the paper rule out confounding mechanisms for the rotation effect?
A: The paper tests and rules out several alternatives: worker-manager specific match effects (the effect does not depend on characteristics of the incoming or outgoing manager); finite project timelines driving a rush to apply; and workers being recruited by managers to their new teams (97% of applications are outside the current team and not to the manager&amp;rsquo;s new team). Balance tests show workers exposed to rotations are observationally similar to non-exposed workers, and event studies confirm absence of pre-trends in team-level outcomes including absenteeism.&lt;/p&gt;
&lt;p&gt;Q: What are the policy implications of the findings?
A: The findings suggest firms forgo productivity gains when hoarded workers are not allocated to positions where they would be most productive. Potential organizational responses include monitoring or rewarding managers for promoting talent, reducing performance-related pay tied to team composition, or structuring career development activities in ways that cannot easily be suppressed by individual managers. The paper notes that firms generally do not compensate managers for promoting workers, partly due to practical difficulties of such contracts, and that the misalignment between what managers believe benefits the firm and what is recognized in their own compensation is particularly pronounced for talent development relative to all other managerial responsibilities.&lt;/p&gt;
&lt;p&gt;Talent hoarding: Actions taken by managers that lower the likelihood that a worker applies for and receives a promotion or internal transfer outside the team, driven by managers&amp;rsquo; incentive to retain productive workers to protect team performance and manager compensation. Distinct from mere neglect — it is strategic and deliberate.&lt;/p&gt;
&lt;p&gt;Potential rating: A public signal of a worker&amp;rsquo;s future potential for higher-level positions, assigned by the direct supervisor and widely circulated within the firm (e.g., via HR lists of high-potential workers); distinguished from performance ratings by its visibility outside the worker&amp;rsquo;s current team, making it a lever for strategic manipulation by hoarding managers.&lt;/p&gt;
&lt;p&gt;Performance rating: A private, task-specific signal of a worker&amp;rsquo;s past performance in their current position, not shared with other units in the firm; used as the baseline against which potential ratings are compared in the paper&amp;rsquo;s administrative hoarding measure.&lt;/p&gt;
&lt;p&gt;Visibility suppression (hoarding measure): The manager-level average residual from a regression of workers&amp;rsquo; potential ratings on their performance ratings and worker characteristics; a positive average residual indicates the manager systematically assigns lower potential ratings than predicted, suppressing worker visibility outside the team in a manner consistent with strategic talent hoarding.&lt;/p&gt;
&lt;p&gt;Manager rotation: An event in which a manager leaves their current team for a different internal position within the firm, temporarily eliminating their hoarding incentive for current team workers and creating the paper&amp;rsquo;s quasi-experimental source of variation in hoarding exposure.&lt;/p&gt;
&lt;p&gt;Marginal applicant: In the IV framework, a worker who applies for an internal position only because their manager is rotating and would not have applied otherwise; estimated via complier analysis (Abadie 2003) and used to characterize the counterfactual quality and hiring probability of workers deterred by talent hoarding.&lt;/p&gt;
&lt;p&gt;Utility cost of hoarding (φm): A manager-level parameter capturing the convex private cost to a manager of engaging in talent hoarding; may reflect altruism, detection risk, or reputational consequences; managers with lower φm hoard more intensively, and variation in φm is proxied empirically by performance-related pay, team size, and functional-area talent visibility.&lt;/p&gt;</description></item><item><title>Taxes Depress Corporate Borrowing: Evidence from Private Firms</title><link>https://macropaperwarehouse.com/papers/taxes-depress-corporate-borrowing-evidence-from-private-firms/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/taxes-depress-corporate-borrowing-evidence-from-private-firms/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Does corporate income taxation raise or lower corporate leverage? The canonical Modigliani-Miller (1963) view holds that the interest tax deduction makes debt more attractive, predicting a positive taxes-to-leverage relationship. Most prior empirical work using large public firms confirms this prediction. This paper re-examines the question using data on small private U.S. firms and finds the opposite: higher corporate taxes &lt;em&gt;depress&lt;/em&gt; leverage, at least for small, financially constrained private firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Identification&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The primary dataset is the Federal Reserve&amp;rsquo;s Y-14Q supervisory collection (2011–2017), which covers the loan portfolios of the 33 largest U.S. banks and includes firm-level income statements and balance sheets for privately held, bank-dependent borrowers. The sample is restricted to domestic private C-corporations with prior-year assets above $100 million (to screen for pass-through entities), yielding 39,363 non-singleton firm-year observations. The median firm has $288 million in book assets and total debt-to-assets of approximately 38%. A supplementary dataset from the Shared National Credit (SNC) Program (1993–2018, 50,203 firm-year observations) provides a longer time series on syndicated loan commitments. Public firm comparisons use CRSP-Compustat (91,314 observations, 1989–2017).&lt;/p&gt;
&lt;p&gt;The empirical strategy is a difference-in-differences event study using variation in state corporate income tax rates. A novel contribution is the manual collection of both &lt;em&gt;enactment&lt;/em&gt; dates (when legislation was signed into law) and &lt;em&gt;effective&lt;/em&gt; dates for each state tax change since 1975. Identification follows the narrative approach of Romer and Romer (2010) and Giroud and Rauh (2019) to exclude tax changes endogenous to local economic conditions. The specification includes firm and industry-by-year fixed effects, and the analysis uses heterogeneity-robust estimators (Borusyak et al. 2024; de Chaisemartin and D&amp;rsquo;Haultfoeuille 2020) to address staggered treatment timing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Empirical Findings&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;For small private firms (below-median total assets, i.e., below $288 million), long-term debt-to-assets rises by approximately 4% in the year of tax cut &lt;em&gt;enactment&lt;/em&gt; and remains elevated—at approximately 2%—four or more years later, indicating a permanent increase in leverage. This anticipation effect arises because firms respond to the law&amp;rsquo;s passage, not its effective date; results using effective dates are noisy and largely insignificant. The average tax cut during the sample period was 1.2 percentage points, representing approximately a 6% reduction in firms&amp;rsquo; tax bills (given an average private-firm tax rate of 21%), and the implied leverage change of about 6% at year four is correspondingly large, consistent with a low-interest-rate environment in which small changes in marginal q translate into large investment and borrowing responses.&lt;/p&gt;
&lt;p&gt;For large private firms (above-median assets), leverage shows no significant response to tax cuts in any event year. For public firms, evidence of any effect is scant, with at most transient significance and pre-trend issues that complicate interpretation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanism&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The paper argues two tax-sensitive costs of debt offset the standard interest tax shield. First, a higher tax rate reduces after-tax profits, raising default probabilities and credit spreads endogenously; a tax cut thus lowers credit spreads and incentivizes more borrowing. Second, because external equity finance is either unavailable or very costly for small private firms, debt and capital are complements in financing investment: a tax cut raises the marginal product of capital, inducing firms to invest and borrow more. For small firms with low capital adjustment costs, this capital-debt complementarity dominates the direct loss of interest tax shield value. For large firms with high capital adjustment costs (estimated at nine times the small-firm value), investment responds sluggishly to tax changes, the complementarity effect is muted, and the traditional tax shield effect becomes relatively more important—producing the standard, slightly positive taxes-to-leverage relationship.&lt;/p&gt;
&lt;p&gt;Bank-assessed default probabilities fall by 20–30 basis points (roughly a 10% decline from an average of approximately 2%) in the year of enactment or one year later for small borrowers, directly supporting the model&amp;rsquo;s credit spread mechanism.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Welfare Counterfactual&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Removing the interest tax deduction from the estimated model (while retaining profit taxation and restricted equity access) causes leverage to fall from 0.36 to −0.26. Firms substitute into cash holdings, shrinking the capital stock. In equilibrium, hours worked rise, the real wage falls, and consumer welfare drops by approximately 1.8%. The interest deduction thus raises welfare in a second-best sense by offsetting other frictions that impede optimal capital accumulation.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-do-prior-studies-find-a-positive-taxes-to-leverage-relationship-and-how-does-this-paper-differ"&gt;Q1. Why do prior studies find a positive taxes-to-leverage relationship, and how does this paper differ?&lt;/h3&gt;
&lt;p&gt;Prior studies—including Titman and Wessels (1988), Heider and Ljungqvist (2015), and Faccio and Xu (2015)—predominantly use large public firms, for which the interest tax shield is the quantitatively dominant consideration. The present paper focuses on small private firms that face greater financial frictions (restricted equity access, higher default risk), in which two additional tax-sensitive costs of debt become quantitatively important. A further methodological difference from Heider and Ljungqvist (2015) is the use of firm fixed effects rather than first differences, which the authors argue is appropriate in a staggered DiD design.&lt;/p&gt;
&lt;h3 id="q2-why-use-enactment-dates-rather-than-effective-dates-as-the-event"&gt;Q2. Why use enactment dates rather than effective dates as the event?&lt;/h3&gt;
&lt;p&gt;Tax legislation is often signed into law one to two years before taking effect; in the sample of 125 tax packages since 1975, 33 became effective the following year and 13 became effective two or more years later. Firms that anticipate future tax changes will adjust leverage immediately upon enactment, not at the effective date. Results confirm this: event studies using enactment dates yield precise positive estimates for small firms (ranging from ~4% at year 0 to ~2% at year 4+), while results using effective dates are noisy and mostly insignificant. The paper therefore treats the enactment date as the economically relevant event and collects these dates as a novel contribution.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-economic-magnitude-of-the-leverage-response-for-small-private-firms"&gt;Q3. What is the economic magnitude of the leverage response for small private firms?&lt;/h3&gt;
&lt;p&gt;Small firms&amp;rsquo; long-term debt-to-assets rises by almost 4% in the enactment year and remains elevated at approximately 2% four or more years after enactment, consistent with a permanent adjustment. The average tax cut during the period was 1.2 percentage points, representing roughly a 6% reduction in the average tax bill (given an average effective rate of 21% for private firms, per Zwick et al. 2016). The estimated coefficient of 0.021 in year four also implies approximately a 6% change in leverage, a large response that the paper attributes to the low interest rate environment amplifying the marginal q effect of even modest tax changes.&lt;/p&gt;
&lt;h3 id="q4-do-large-private-firms-respond-differently-to-tax-cuts-and-why"&gt;Q4. Do large private firms respond differently to tax cuts, and why?&lt;/h3&gt;
&lt;p&gt;Large private firms (above the median of $288 million in total assets) show no statistically significant leverage response to tax cuts in any event year, and this null is not attributable to wider confidence intervals. The model estimation explains this via capital adjustment costs: the adjustment cost parameter for large firms is estimated to be nine times larger than for small firms. With high adjustment costs, investment responds sluggishly to a tax cut, so the complementarity channel (more investment requires more debt) is suppressed. The traditional tax shield effect then becomes relatively more important, producing a slightly positive (or zero net) taxes-to-leverage relationship consistent with the large-firm data moment.&lt;/p&gt;
&lt;h3 id="q5-how-does-the-model-generate-a-negative-relationship-between-taxes-and-leverage-when-the-interest-tax-deduction-is-present"&gt;Q5. How does the model generate a negative relationship between taxes and leverage when the interest tax deduction is present?&lt;/h3&gt;
&lt;p&gt;Two mechanisms offset the tax shield. First, higher taxes reduce after-tax profits, pushing firms closer to the default threshold; this is capitalized into equilibrium credit spreads, raising the cost of debt. Specifically, for small firms, the model shows that once leverage exceeds approximately 0.47 of assets, the after-tax risky interest rate rises monotonically with the tax rate (rather than falling via the deduction effect). Second, capital and debt are complements in financing investment: because a tax cut raises the marginal product of capital, and because external equity is unavailable, firms substitute into capital by using more leverage. For small firms with low capital adjustment costs, both mechanisms outweigh the loss of interest tax shield value when taxes fall.&lt;/p&gt;
&lt;h3 id="q6-how-are-the-model-parameters-estimated-and-what-are-the-key-parameter-values"&gt;Q6. How are the model parameters estimated, and what are the key parameter values?&lt;/h3&gt;
&lt;p&gt;The model is estimated by simulated method of moments on the Y-14 small-firm sample, minimizing the distance between nine data moments and their model-simulated counterparts. The nine moments include the means and standard deviations of debt, investment, and operating income (all as ratios of assets), the serial correlations of investment and operating income, and the coefficient from a two-way fixed-effects regression of leverage on a tax-change dummy. The deadweight loss in default (ξ) is estimated at 0.6 for small firms and 0.32 for large firms, consistent with elevated financial frictions for small firms and in line with average recovery rates in Kermani and Ma (2023). Fixed operating costs (f) are approximately 0.15 for both samples, amounting to just under half of steady-state operating profits. The serial correlation of the tax process is estimated at 0.662, with innovation standard deviation of 0.022.&lt;/p&gt;
&lt;h3 id="q7-what-is-the-models-welfare-counterfactual-and-what-does-it-imply"&gt;Q7. What is the model&amp;rsquo;s welfare counterfactual, and what does it imply?&lt;/h3&gt;
&lt;p&gt;The paper compares two economies both with profit taxation: one with the interest tax deduction and one without. Removing the deduction in the small-firm model causes leverage to fall from 0.36 to −0.26, as firms hold net cash rather than net debt. The capital stock shrinks, output falls, hours worked rise, and both the real wage and consumption decline. Consumer welfare drops by approximately 1.8%. Capital misallocation (measured following Hsieh and Klenow 2009) worsens from 0.89 to 0.88. The result has a second-best character: the interest deduction incentivizes debt-financed investment that partially offsets the distortion from restricted equity access.&lt;/p&gt;
&lt;h3 id="q8-what-does-the-evidence-on-default-probabilities-add-to-the-empirical-case"&gt;Q8. What does the evidence on default probabilities add to the empirical case?&lt;/h3&gt;
&lt;p&gt;The Y-14 collection contains bank-assessed default probability estimates. In an event study covering Q1 2012–Q4 2018, the authors find that firms&amp;rsquo; assessed default probabilities decline significantly by 20–30 basis points in the year of enactment or one year later for small borrowers (those with total loan commitments of $10–$100 million), representing approximately a 10% decline from the sample average default rate of around 2%. This decline peaks two years after enactment and persists for three years. No comparable decline is observed for larger loan size buckets. Separately, in SNC data, the probability of a non-pass (i.e., below-investment-grade supervisory) rating falls by 1.7–2.2 percentage points following tax cut enactments, persisting roughly three years. Together, these findings directly validate the model mechanism by which tax cuts lower default risk and credit spreads.&lt;/p&gt;
&lt;h3 id="q9-are-the-results-robust-to-alternative-econometric-methods-that-address-heterogeneous-treatment-effects"&gt;Q9. Are the results robust to alternative econometric methods that address heterogeneous treatment effects?&lt;/h3&gt;
&lt;p&gt;Yes. The paper applies the Borusyak et al. (2024) imputation estimator, which imputes fixed effects from untreated observations onto treated observations to remove negative weighting bias; for small firms and event years 0–3, it finds significant positive estimates comparable to the baseline. The de Chaisemartin and D&amp;rsquo;Haultfoeuille (2020, 2021) estimator, based solely on first-time switchers to treatment, yields an effect of 0.036 on leverage for small firms in the enactment year and no effect for large firms, consistent with the baseline. Results using the narrative approach (excluding Connecticut 2011 and 2015, New York 2014, and Rhode Island 2014 as potentially endogenous) produce slightly larger leverage estimates.&lt;/p&gt;
&lt;h3 id="q10-are-tax-hike-effects-symmetric-to-tax-cut-effects"&gt;Q10. Are tax hike effects symmetric to tax cut effects?&lt;/h3&gt;
&lt;p&gt;Evidence on hikes is weaker because tax hikes are rare in the sample. In Y-14 data, hikes are associated with leverage declines for small firms in event year 4 and for large firms in event years 1, 2, and 4, but without sufficient pre-hike observations to identify pre-trends, these results are less credible than the cut results. In SNC data (which spans a longer period, 1992–2018), tax hikes are associated with large and significant reductions in total syndicated borrowing commitments of 6–7%, while cuts produce smaller and marginally significant increases. This asymmetry is consistent with the lower adjustment costs of reducing debt relative to increasing it.&lt;/p&gt;
&lt;h3 id="q11-what-does-the-analysis-of-alternative-model-specifications-reveal-about-the-generality-of-the-mechanism"&gt;Q11. What does the analysis of alternative model specifications reveal about the generality of the mechanism?&lt;/h3&gt;
&lt;p&gt;Three model extensions are considered. In a collateral-constrained model (no endogenous default), the cost of debt is lost financial flexibility (the future shadow cost of the borrowing constraint), which remains tax-sensitive. In a model with costly equity issuance (linear cost λ = 0.11 following Hennessy and Whited 2007), equity issuance is rare, so the model behaves nearly identically to the baseline. In a solvency-based default model (default when firm value turns negative rather than when liquidity is insufficient), the negative taxes-to-leverage result is preserved. A news-shock extension (Jaimovich-Rebelo 2009) incorporating the anticipation of future tax changes also produces lower leverage in response to higher anticipated taxes, consistent with the empirical anticipation effects, though with smaller magnitudes because the news shock variance is smaller than the total tax-change variance.&lt;/p&gt;
&lt;h3 id="q12-why-do-contingent-claims-models-fischer-leland-goldstein-class-always-predict-a-positive-taxes-to-leverage-relationship"&gt;Q12. Why do contingent-claims models (Fischer-Leland-Goldstein class) always predict a positive taxes-to-leverage relationship?&lt;/h3&gt;
&lt;p&gt;In these models, shareholders have deep pockets, so negative cash flows can always be covered; this implies default is rare and the effect of taxes on the default put value is small relative to the direct interest tax deduction. Additionally, these models contain no capital stock, so there is no substitution mechanism between capital and a storage technology (i.e., cash/negative debt). Without endogenous investment, the only channel linking taxes to leverage is the tax shield, which necessarily implies a positive taxes-to-leverage relationship. This is why, as the paper notes, the result was &amp;ldquo;already hiding&amp;rdquo; in the Hennessy-Whited class of dynamic investment models but not visible in the contingent-claims literature.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Interest Tax Deduction (Tax Shield)&lt;/strong&gt;
The paper uses this in the standard corporate finance sense: the after-tax cost of debt is reduced because interest payments are deductible against corporate income. In the model, debt proceeds are discounted at the after-tax interest rate, and the deduction is taken at the time of debt issuance. The paper&amp;rsquo;s contribution is to show this benefit can be outweighed by two tax-sensitive costs of debt, reversing the sign of the taxes-to-leverage relationship for small, constrained firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tax-Sensitive Cost of Debt&lt;/strong&gt;
The paper defines two distinct tax-sensitive costs that offset the tax shield. First, taxes reduce after-tax profits, shifting the default threshold and raising equilibrium credit spreads; this is capitalized into the risky lending rate endogenously from the lender&amp;rsquo;s zero-profit condition. Second, taxes reduce the marginal product of capital, making debt-financed investment less attractive; because debt and capital are complements in a model without external equity, a higher tax rate lowers optimal capital and, with it, optimal debt.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Capital Adjustment Costs (ψ)&lt;/strong&gt;
Quadratic costs of changing the capital stock, parameterized as ψ(k&amp;rsquo; − (1−δ)k)² / (2k). The paper identifies this parameter as the key determinant of whether leverage responds positively or negatively to taxes: for small firms, ψ is estimated to be near zero (insignificantly different from zero), enabling free substitution between capital and the storage technology (negative debt), so the complementarity channel dominates. For large firms, ψ is estimated to be nine times larger, suppressing this substitution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Default Threshold&lt;/strong&gt;
In the model, default is triggered when the firm&amp;rsquo;s current after-tax profits plus recoverable capital are insufficient to repay debt: (1−τ)(y − wn − f) + (1−ξ)(1−δ)k &amp;lt; p. This threshold depends directly on the tax rate τ, so higher taxes move the threshold in the direction of default, raising credit spreads. The paper provides empirical support for this mechanism via the event study of bank-assessed default probabilities.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Enactment Date vs. Effective Date&lt;/strong&gt;
The paper distinguishes between the date tax legislation is signed into law (enactment date) and the date it becomes operative (effective date), which can differ by one to two years. The paper collects novel data on enactment dates from state legislative records. The empirical finding that firms respond to enactment rather than effective dates constitutes evidence of anticipation effects: firms adjust leverage upon observing future expected tax changes, not when the changes actually take hold.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Second-Best Welfare Effect of the Tax Deduction&lt;/strong&gt;
The paper uses this term to characterize the welfare result from the counterfactual: in an economy already distorted by profit taxation and restricted equity access, the interest deduction raises consumer welfare by incentivizing debt-financed capital accumulation. Removing the deduction causes firms to substitute into cash, shrinking the capital stock and lowering wages and consumption. This is a second-best result because the deduction is welfare-improving only because it partially offsets the distortions created by other frictions; in a frictionless world, no such second-best rationale would apply.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Y-14Q Supervisory Data&lt;/strong&gt;
The Federal Reserve&amp;rsquo;s supervisory collection from the 33 largest U.S. banks, covering loan portfolios and associated borrower financial statements for firms with commercial and industrial loans exceeding $1 million in commitment. The paper uses this dataset because it covers private, bank-dependent firms—a population not previously studied in the tax-leverage literature—and contains firm-level balance sheets, credit ratings, and default probability estimates.&lt;/p&gt;</description></item><item><title>Technology Transfer and Early Industrial Development: Evidence from the Sino-Soviet Alliance</title><link>https://macropaperwarehouse.com/papers/technology-transfer-and-early-industrial-development-evidence-from-the-sino-soviet-alliance/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/technology-transfer-and-early-industrial-development-evidence-from-the-sino-soviet-alliance/</guid><description>&lt;p&gt;This paper estimates the causal effect of technology and knowledge transfers on early industrial development using the Sino-Soviet Alliance of the 1950s as a natural experiment. Between 1950 and 1957, the Soviet Union supported the &amp;ldquo;156 Projects&amp;rdquo; — 139 approved civil projects for constructing technologically advanced, large-scale, capital-intensive industrial facilities in China. The intended program comprised two components: a &amp;ldquo;basic&amp;rdquo; transfer of Soviet state-of-the-art machinery and equipment (including blueprints, site surveys, and plant construction assistance), and an &amp;ldquo;advanced&amp;rdquo; know-how transfer involving Soviet experts residing in Chinese plants for roughly three years to train engineers and production supervisors in organizational, technological, and planning methods. Total investment amounted to approximately $80 billion in 2020 figures (45.7% of Chinese GDP in 1949).&lt;/p&gt;
&lt;p&gt;Identification exploits idiosyncratic delays in project completion caused by Soviet production capacity constraints, insufficient experts, translator shortages, and miscommunication — factors documented in historical records as unrelated to project-specific characteristics. When the Sino-Soviet Split in 1960 abruptly ended the program, all 139 plants had been built but differed in what transfers they had received: 46 received both machinery and know-how (advanced), 46 received only machinery (basic), and 47 received neither (comparison). The paper verifies, via ANOVA tests, multinomial logit models, balancing regressions on 26 plant characteristics, pre-trend tests, and Oster (2019) selection-on-unobservables bounds, that the three groups were statistically equivalent prior to receiving the Soviet transfers.&lt;/p&gt;
&lt;p&gt;The primary data source is plant-level annual reports from the Steel Association covering 94 steel firms (1,410 plants) from 1949 to 2000, matched to 304 steel plants across the 156 Projects. Supplementary sources include the declassified 1985 Second Industrial Survey (7,592 largest Chinese firms) and the China Industrial Enterprises database (1998–2013, over 1 million firms).&lt;/p&gt;
&lt;p&gt;Three main results emerge. First, receiving only the basic (machinery) transfer had positive but short-lived effects: output of basic plants peaked at 14.7 percent above comparison plants six years after receiving Soviet machinery, then declined monotonically and became statistically insignificant after 20 years — consistent with the estimated 15–20 year life cycle of Soviet capital. Second, the advanced transfer had large and persistent effects: advanced plants&amp;rsquo; output rose 8.4 percent relative to basic plants within two years, 19.7 percent within 20 years, and 49.5 percent cumulatively after 40 years. TFPQ of advanced plants reached 47.9 percent above basic plants after 40 years. These magnitudes held across industries in 1985 and 1998–2013 data, where value added of advanced firms was 41.4–52.0 percent higher and TFPR 39.5–49.3 percent higher than basic firms. Third, the program generated horizontal spillovers (12.9 percent higher output, 12.4 percent higher productivity for steel plants in counties hosting advanced plants) and vertical spillovers (16.4 percent productivity gain for supply-chain firms in counties of advanced nonsteel plants), with spillover effects conditional on post-1990s market liberalization to materialize in private firms.&lt;/p&gt;
&lt;p&gt;The mechanism driving persistence is the accumulation of organizational and human capital during the advanced transfer, which enabled advanced plants — uniquely — to develop new production processes endogenously, home-fabricate continuous casting furnaces to replace obsolete Soviet open-hearth equipment, and produce export-quality steel. Advanced plants employed more engineers and high-skilled technicians, established professional schools, and their counties had 10.4 percent higher STEM university degree rates and 16.8 percent more technical schools.&lt;/p&gt;
&lt;p&gt;Scope conditions: results apply to large-scale, capital-intensive state-planned industrial facilities in a country at an early stage of industrialization, under conditions of near-complete trade isolation (1960–1978) that prevented basic plants from compensating via imported foreign capital. The estimated aggregate contribution of the program is that, without both transfer types, Chinese real GDP per capita growth between 1953 and 1978 would have been 7 to 19 percent lower.&lt;/p&gt;
&lt;p&gt;Q: What distinguishes the &amp;ldquo;basic&amp;rdquo; from the &amp;ldquo;advanced&amp;rdquo; Soviet transfer?
A: The basic transfer involved duplication of whole Soviet plants through provision of state-of-the-art Soviet machinery, equipment, blueprints, geological surveys, and construction assistance. The advanced transfer added visits of Soviet experts — expected to stay approximately three years — to teach Chinese technicians how to operate the machinery and to provide within-firm training in engineering (math, physics, chemistry, organizational and planning methods) and supervisory management based on &amp;ldquo;scientific management&amp;rdquo; principles including quality-control methods.&lt;/p&gt;
&lt;p&gt;Q: What caused plants to receive different levels of transfer, and why is this variation credible for identification?
A: Delays arose from Soviet production capacity constraints (by 1955, one-third of annual Soviet steel-rolling output was destined for China), insufficient experts, translator shortages, and bilateral miscommunication — all documented in historical records as unrelated to project characteristics. When the 1960 Split ended the program, plants&amp;rsquo; treatment status was determined by where they happened to be in the delivery queue. ANOVA tests find no significant differences in approval year, investment, workforce, equipment value, project length, or capacity across the three groups, and a multinomial logit on province and industry fixed effects shows no group had higher ex-ante probability of receiving either transfer type.&lt;/p&gt;
&lt;p&gt;Q: What were the output effects of the basic transfer, and why did they fade?
A: Output of basic plants was not significantly above comparison plants for the first two years, peaked at 14.7 percent higher six years after receiving Soviet machinery, then declined monotonically and became statistically insignificant after 20 years. This timing corresponds to the estimated 15-year life cycle of Soviet capital goods. TFPQ of basic plants followed the same pattern, peaking at 14.5 percent above comparison plants. Without the know-how component, basic plants could not develop new processes or home-fabricate replacement capital, so productivity advantages disappeared as Soviet equipment became obsolete.&lt;/p&gt;
&lt;p&gt;Q: What were the output and productivity effects of the advanced transfer?
A: Advanced plants&amp;rsquo; output rose 8.4 percent relative to basic plants within two years of the Soviet transfer and 19.7 percent within 20 years, reaching a cumulative effect of 49.5 percent after 40 years. TFPQ of advanced plants increased from 8.3 percent above basic plants two years after the transfer to 47.9 percent after 40 years. These effects were driven by output growth rather than differential input use — the number of workers, coke, and iron were statistically indistinguishable across the three plant types — ruling out government input reallocation as an explanation.&lt;/p&gt;
&lt;p&gt;Q: Did the advanced transfer affect steel quality?
A: Advanced plants produced substantially more crude steel (higher quality, lower carbon content) and less pig iron than basic and comparison plants, and this quality advantage persisted well beyond the 20-year life cycle of Soviet capital. Basic plants also shifted toward crude steel initially but the quality advantage dissipated once Soviet machinery became obsolete, whereas advanced plants maintained the shift through adoption of the basic oxygen process and later continuous casting furnaces.&lt;/p&gt;
&lt;p&gt;Q: What is the main mechanism through which the advanced transfer generated persistent effects?
A: The advanced transfer equipped engineers and supervisors with organizational, technological, and planning knowledge, enabling advanced plants to develop and adopt the basic oxygen steelmaking process independently during China&amp;rsquo;s 1960–1978 period of trade isolation. Advanced plants had a 15.2 percent higher probability of using the basic oxygen process five years after the transfer and a 65.1 percent higher probability twenty years after, relative to basic plants. They also home-fabricated continuous casting furnaces, making them 26.7 to 78.4 percent more likely to use such furnaces 10 to 20 years after the transfer; basic plants showed no differential advantage over comparison plants on this measure.&lt;/p&gt;
&lt;p&gt;Q: What role did trade openness play in the divergence between basic and advanced plants?
A: Once China opened to international trade from 1978, advanced plants relied dramatically less on imported foreign capital than basic plants — likely because they had developed domestic production capabilities. At the same time, advanced plants exported 45.5 percent more steel and produced 51.1 percent more steel above international quality standards than basic plants. Basic plants showed no differential imports of foreign capital or differential exports relative to comparison plants, suggesting that once both types could access foreign machinery, basic plants lost any remaining productivity edge.&lt;/p&gt;
&lt;p&gt;Q: What were the human capital effects of the advanced transfer?
A: Over time, advanced plants opened training schools for high-skilled technicians and offered within-firm training programs for engineers. As a result, advanced plants employed more engineers and high-skilled technicians and fewer low-skilled workers than basic plants, while the human capital composition did not differentially change between basic and comparison plants. At the county level, universities hosting advanced plants were 10.4 percent more likely to offer STEM degrees, had 16.8 percent more technical schools, 14.3 percent more STEM college graduates, and 17.6 percent more high-skilled workers than counties hosting basic plants.&lt;/p&gt;
&lt;p&gt;Q: Did the government differentially favor basic or advanced plants after the Split?
A: The paper finds no evidence of special government favor. Government transfers and loans were not differentially allocated to basic or advanced plants in either the short or long run. Distance from railroads and roads did not change differentially across plant types. Measures of political connection and politician quality at the prefecture level showed no significant differences across the three groups in the 40 years after the Soviet transfer. County-level total investment and investments in related and unrelated industries were also statistically indistinguishable.&lt;/p&gt;
&lt;p&gt;Q: What were the intra-firm spillover effects?
A: Steel plants in the same firm as advanced plants increased their steel production by 24.9 percent and were 22.1 percent more productive relative to plants in the same firm as basic plants, after the Soviet transfer. Plants in the same firm as basic plants showed no differential performance relative to plants in the same firm as comparison plants. The within-firm spillovers appear driven by the transmission of new technologies and production methods through formal within-firm training programs, as supported by historical records.&lt;/p&gt;
&lt;p&gt;Q: What were the horizontal spillover effects across firms?
A: Steel plants in the same counties as advanced plants produced 12.9 percent higher output and were 12.4 percent more productive than those in counties hosting basic plants, after the transfer. They were more likely to adopt basic oxygen converters and continuous casting furnaces, and from 1978 they exported significantly more and produced more steel above international quality standards, mirroring the patterns of the advanced plants themselves.&lt;/p&gt;
&lt;p&gt;Q: What were the vertical spillover effects?
A: Steel plants in counties hosting nonsteel basic plants produced 14.2 percent more steel than those in counties hosting nonsteel comparison plants, suggesting some output spillover from basic machinery. However, only plants in counties of advanced nonsteel plants experienced a productivity increase — estimated at 16.4 percent — relative to plants in counties of basic nonsteel plants. These supply-chain firms were also the only ones to show increased adoption of basic oxygen and continuous casting furnace technology and differential engagement in trade.&lt;/p&gt;
&lt;p&gt;Q: How did market liberalization reforms interact with the spillover effects?
A: Starting in the late 1990s, privatized firms economically related to advanced plants outperformed their counterparts in terms of value added, TFPR, and exports, while state-owned firms in the same counties no longer showed a competitive advantage. New private firms locating in counties that had hosted advanced plants received an additional performance gain. At the county level, counties hosting advanced plants had on average 16.6 percent more private firms and 25.2 percent more privately-produced industrial output than counties hosting basic plants. The mechanism appears to be the stock of industry-specific human capital concentrated in those counties, which private firms could draw on once allowed to compete for workers.&lt;/p&gt;
&lt;p&gt;Q: What is the estimated aggregate contribution of the Soviet transfer to Chinese growth?
A: Province-level regressions show that each additional basic project increased province-level output by 1.1 percent per year on average, and each additional advanced project by 6.2 percent per year. A back-of-the-envelope calculation implies that without both transfer types, Chinese real GDP per capita growth between 1953 and 1978 would have been 7 to 19 percent lower.&lt;/p&gt;
&lt;p&gt;Q: How does the paper rule out selection on unobservable characteristics?
A: Using the Oster (2019) methodology, the paper finds that for the treatment effects to become statistically insignificant, selection on unobserved variables would need to be 8 to 19 times larger than selection on observed variables — a range the authors characterize as implausible given the strong balancing on observables and the historical documentation of delay causes.&lt;/p&gt;
&lt;p&gt;Q: How does this paper differ from Heblich et al. (2020), which also studies Sino-Soviet technology transfer?
A: Heblich et al. (2020) study long-run negative spillovers of the 156 Projects on counties that hosted them relative to counties that were geographically suitable but ultimately not selected, focusing on an outside-the-program comparison. This paper instead exploits within-program variation — differences across the three plant types — using plant-level data to assess short-, medium-, and long-run direct effects and spillover effects of different transfer intensities.&lt;/p&gt;
&lt;p&gt;Basic Transfer: The provision of Soviet state-of-the-art machinery, equipment, blueprints, geological surveys, and plant construction assistance — duplicating a whole Soviet plant — without accompanying human capital or organizational training.&lt;/p&gt;
&lt;p&gt;Advanced Transfer: The full Soviet technology and know-how package: basic machinery provision plus multi-year visits of Soviet experts who taught Chinese engineers and production supervisors organizational, technological, and planning methods based on &amp;ldquo;scientific management&amp;rdquo; principles.&lt;/p&gt;
&lt;p&gt;Comparison Plants: Plants approved under the 156 Projects that received neither Soviet machinery nor technical assistance due to delays compounded by the Split, and which continued operating with traditional domestic technology.&lt;/p&gt;
&lt;p&gt;156 Projects: An array of 139 approved, technologically advanced, large-scale, capital-intensive industrial facilities whose construction the Soviet Union agreed to support between 1950 and 1957 as part of the Sino-Soviet Alliance, representing 45.7% of Chinese GDP in 1949.&lt;/p&gt;
&lt;p&gt;Tacit Knowledge: Industry- and firm-specific knowledge embodied in workers and organizations — including operational methods, quality-control procedures, and process innovation capabilities — that cannot be transferred through capital goods alone and requires extensive on-the-job training from foreign experts.&lt;/p&gt;
&lt;p&gt;Basic Oxygen Process: A steelmaking process innovation that became predominant in the 1960s by blowing oxygen through molten pig iron to reduce carbon content; adopted by advanced plants through endogenous process development, while basic plants showed no differential adoption relative to comparison plants.&lt;/p&gt;
&lt;p&gt;Source Text Origin: The paper&amp;rsquo;s classification scheme for the grounding of evidence — in this case, full working paper text obtained from NBER WP 29455, enabling comprehensive summary of quantitative results, mechanisms, and robustness tests.&lt;/p&gt;</description></item><item><title>The Architecture of Social Networks and the Diffusion of Innovations</title><link>https://macropaperwarehouse.com/papers/the-architecture-of-social-networks-and-the-diffusion-of-innovations/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-architecture-of-social-networks-and-the-diffusion-of-innovations/</guid><description>&lt;p&gt;This paper examines how the architecture of social networks shapes the success or failure of technology diffusion when adoption decisions exhibit strategic complementarities. The research question is: which structural feature of a network determines whether a new technology spreads or fails, and in which direction does that feature work?&lt;/p&gt;
&lt;p&gt;The paper builds on the canonical threshold diffusion model of Morris (2000) and Granovetter (1978), in which an agent adopts a new technology if the share of his neighbors who have adopted exceeds a threshold Q in [0,1]. The key innovation is the addition of a second structural object — a set of decision-making units C — that captures the empirically common phenomenon that subsets of agents (friends, family, neighbors, colleagues) can coordinate and make joint adoption decisions. The model is purely theoretical; the paper derives characterizations and comparison theorems rather than estimating parameters from data.&lt;/p&gt;
&lt;p&gt;The central structural concept introduced is insularity: the extent to which agents concentrate their connections to a narrow set of other agents, rather than distributing connections broadly. A formal partial order over networks is defined: network {w̃} is less insular than network {w} if there is no local increase in insularity in {w̃} relative to {w}, where a local increase in insularity occurs when one agent&amp;rsquo;s proportionate connections to a narrow set S are strictly higher and another agent&amp;rsquo;s proportionate connections to a superset R are strictly lower (with the first agent&amp;rsquo;s share of S weakly exceeding the second agent&amp;rsquo;s share of R). Moving from a network toward a convex combination with the complete network strictly reduces insularity under this definition (Lemma 3).&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s main characterization result (Proposition 1) establishes that the set of non-adopters of technology Q is precisely SQ — the maximal (1−Q)-subgroup-cohesive set — defined as the largest set in which every decision-making unit C contained in SQ has at least one agent with at least fraction (1−Q) of his connections inside SQ. This extends Morris&amp;rsquo;s (2000) cohesion characterization to the joint-decision setting.&lt;/p&gt;
&lt;p&gt;The main theorem (Theorem 1) establishes that for any two societies sharing the same decision-making structure C but differing in network insularity, there exists a cutoff threshold mu in [0,1] such that: (i) for technologies with Q &amp;lt; mu, adoption is weakly higher in the less insular network; and (ii) for technologies with Q &amp;gt;= mu, adoption is weakly lower in the less insular network. The direction reversal at mu reflects two competing mechanisms. Insular connections hinder singleton diffusion: an agent over-connected to a narrow set will not adopt individually until others in that set adopt, blocking entry of the technology from outside. But insular connections facilitate joint adoption: the same over-connectedness makes it profitable for the group to adopt together if they can coordinate, because each member already has a high share of neighbors within the group. High-threshold technologies depend crucially on joint adoption cascades and so benefit from insularity; low-threshold technologies spread person-to-person and are impeded by insularity when agents cannot coordinate.&lt;/p&gt;
&lt;p&gt;Proposition 2 establishes a complementary monotonicity result: expanding the set of decision-making units (C subset of C&amp;rsquo;) weakly increases adoption for any technology and any network, because joint decision-making resolves local coordination failures.&lt;/p&gt;
&lt;p&gt;The main result is extended to heterogeneous thresholds (Section 7). Proposition 3 shows that Theorem 1 continues to hold when agent-specific idiosyncratic components theta_i are bounded within an interval [−gamma/2, gamma/2] for some gamma &amp;gt; 0. Proposition 4 characterizes the necessary conditions for the main result to break: the specification fails only if there exist two agents i and j with theta_i &amp;gt; theta_j + Q2 − Q1, meaning the idiosyncratic gap between them exceeds the difference between the two technology thresholds being compared.&lt;/p&gt;
&lt;p&gt;Q: What is the paper&amp;rsquo;s central research question?
A: The paper asks how the architecture of a social network — specifically the structure of agents&amp;rsquo; connections — determines whether a new technology spreads widely or fails to diffuse. It focuses on technologies with strategic complementarities, where an agent&amp;rsquo;s benefit from adopting depends on neighbors adopting and those neighbors&amp;rsquo; benefit depends on their neighbors, creating potential for both snowballing and coordination failure.&lt;/p&gt;
&lt;p&gt;Q: What is the key modeling innovation relative to the standard threshold model?
A: The paper adds a set of decision-making units C, a collection of subsets of agents each of which can make a joint adoption decision. In the standard Morris (2000) model, only individual agents decide; here, groups such as friends, family, or neighbors can collectively agree to adopt, resolving their local coordination problem. The set C is subject only to closure under subsets and inclusion of all singletons, making the framework highly flexible.&lt;/p&gt;
&lt;p&gt;Q: How does the diffusion process work formally?
A: At each period t &amp;gt;= 1, agent i adopts if either: (1) more than fraction Q of his neighbors adopted in period t−1 (singleton adoption), or (2) i belongs to a decision-making unit C not yet adopted, and for every j in C the fraction of j&amp;rsquo;s neighbors in A_{t−1} union C exceeds Q (joint adoption). Actions are irreversible, and Appendix C proves this irreversibility assumption is without loss of generality for the final adoption set under myopic best-response dynamics.&lt;/p&gt;
&lt;p&gt;Q: What is the characterization of non-adopters (Proposition 1)?
A: The set of agents who do not adopt technology Q equals SQ, the unique maximal (1−Q)-subgroup-cohesive set — the largest set S such that every decision-making unit C contained in S has at least one member i with Pi(S minus C) &amp;gt;= (1−Q), meaning at least fraction (1−Q) of i&amp;rsquo;s connections remain inside S outside of C. This extends Morris (2000)&amp;rsquo;s p-cohesion concept: when C contains only singletons, (1−Q)-subgroup cohesion collapses to (1−Q)-cohesion in Morris&amp;rsquo;s sense.&lt;/p&gt;
&lt;p&gt;Q: What does the simple eight-agent example illustrate?
A: With two four-clique subgraphs (agents 1-4 and 5-8), Network A has agents 1, 3, 5, 7 each holding 3/4 of their connections within their four-agent group; Network B reduces those within-group shares to 5/8 by weakening two within-group links from weight 1 to weight 1/2 and adding cross-group links of weight 1/2. For Q = 3/10: in Network B all eight agents adopt (group {1,2,3,4} adopts jointly at t=1, then agents 5 and 7 adopt as singletons at t=2, agents 6 and 8 at t=3), while in Network A only {1,2,3,4} adopt (agents 5-8 each have only 1/4 of neighbors adopted, below Q = 3/10). For Q = 7/10: in Network A group {1,2,3,4} adopts jointly (each has 3/4 &amp;gt; 7/10 of neighbors adopting), while in Network B there is zero adoption (agent 3 has only 5/8 &amp;lt; 7/10 of neighbors in the joint group). This is the concrete illustration of the threshold-dependent reversal in Theorem 1.&lt;/p&gt;
&lt;p&gt;Q: What is insularity and how is it formally defined?
A: Insularity is the extent to which agents concentrate their connections to a narrow set of others. A local increase in insularity in {w} relative to {w̃} occurs when, for some agents i and j and sets S subset of R: (1) Pi(S) is strictly higher in {w} and Pj(R) is strictly lower in {w}, and (2) Pi(S) &amp;gt;= Pj(R) in {w}. Network {w̃} is less insular than {w} if no local increase in insularity exists in {w̃} relative to {w}. Lemma 3 establishes that the lambda-convex combination of any non-complete network with the complete network is strictly less insular.&lt;/p&gt;
&lt;p&gt;Q: What is the main theorem (Theorem 1) and its precise statement?
A: For two societies sharing the same decision-making structure C but differing in network insularity — with {w̃} strictly less insular than {w} — there exists a cutoff mu in [0,1] such that: for Q &amp;lt; mu, adoption is weakly higher in the less insular network; and for Q &amp;gt;= mu, adoption is weakly lower in the less insular network. The cutoff mu depends on the specific networks and decision-making structure. The result is a clean reversal: less insular is better for low-threshold technologies and worse for high-threshold technologies.&lt;/p&gt;
&lt;p&gt;Q: What are the two competing mechanisms driving Theorem 1?
A: First, insular connections hinder individual diffusion: an agent with a high share of connections concentrated inside a set will not adopt as a singleton until others in that set adopt, blocking entry of the technology from outside via individual contagion. Second, insular connections facilitate joint adoption: precisely because an agent has a high share of connections to a narrow group, jointly adopting with that group is profitable — each member has enough neighbors already within the group to exceed the threshold when the group adopts together. For high-threshold technologies, joint adoption is the only viable mechanism, so the second effect dominates; for low-threshold technologies, singleton diffusion suffices and the first effect dominates.&lt;/p&gt;
&lt;p&gt;Q: How does joint decision-making affect adoption (Proposition 2)?
A: Expanding the set of decision-making units from C to any C&amp;rsquo; containing C weakly increases adoption of technology Q for any network and any Q. The proof shows that the non-adopter set SQ under C&amp;rsquo; is also (1−Q)-subgroup cohesive under C, making it a subset of non-adopters under C. The economic logic is that any group able to make a joint decision can solve its local coordination problem: agents who individually would not adopt because too few neighbors have adopted may collectively adopt if each would benefit from group adoption.&lt;/p&gt;
&lt;p&gt;Q: How robust is Theorem 1 to heterogeneous thresholds?
A: Proposition 3 shows that Theorem 1 extends with the same cutoff structure when each agent i has an idiosyncratic threshold component theta_i in [−gamma/2, gamma/2] for sufficiently small gamma &amp;gt; 0. Proposition 4 establishes the necessary condition for the result to break with unbounded heterogeneity: there must exist agents i and j with theta_i &amp;gt; theta_j + Q2 − Q1, meaning the idiosyncratic gap must strictly exceed the technology threshold gap being compared. The underlying intuition of Theorem 1 persists even when the precise specification fails.&lt;/p&gt;
&lt;p&gt;Q: What are the policy and managerial implications?
A: A firm with a low-threshold technology should target less insular societies to maximize uptake, while a firm with a high-threshold technology should target more insular societies; the paper cites Facebook&amp;rsquo;s initial launch within closed university networks as consistent with the high-threshold logic. Policymakers and firms can increase adoption by encouraging joint decision-making — sanitation campaigns that organize neighborhood workshops, family mobile-plan discounts, or online coordination platforms all work through this channel. Conversely, governments trying to suppress collective action such as protest can prohibit in-person gatherings or online communication to prevent joint decision-making. The paper notes results abstract from seeding, leaving optimal seeding under joint decision-making as a future research direction.&lt;/p&gt;
&lt;p&gt;Insularity: The extent to which agents concentrate their connections to a narrow set of other agents rather than distributing connections broadly; formally defined via a partial order based on local increases in agents&amp;rsquo; proportionate connections to nested sets S subset of R.&lt;/p&gt;
&lt;p&gt;Decision-making unit: A set C of agents who can make a joint decision to adopt together; the collection C of all decision-making units is closed under subsets and contains all singletons, capturing informal group coordination among friends, family, or neighbors.&lt;/p&gt;
&lt;p&gt;p-Subgroup cohesion: A set S is p-subgroup cohesive if every decision-making unit C contained in S (of any size, including singletons) is p-connected in S — meaning at least one agent in C has at least fraction p of his connections to S minus C; the paper&amp;rsquo;s generalization of Morris (2000)&amp;rsquo;s p-cohesion to settings with joint decision-making.&lt;/p&gt;
&lt;p&gt;Threshold of adoption (Q): A parameter Q in [0,1] summarizing a technology&amp;rsquo;s strategic complementarities, such that an agent is better off adopting if and only if more than fraction Q of his neighbors adopt; low Q means the technology is valuable even with few adopters, high Q means it requires near-universal neighborhood adoption.&lt;/p&gt;
&lt;p&gt;Local increase in insularity: A pairwise comparison between two networks: {w} exhibits a local increase in insularity relative to {w̃} when one agent&amp;rsquo;s proportionate connections to narrow set S are strictly higher and another agent&amp;rsquo;s proportionate connections to superset R are strictly lower in {w}, with the first agent&amp;rsquo;s share of S weakly exceeding the second agent&amp;rsquo;s share of R in {w}.&lt;/p&gt;
&lt;p&gt;SQ (maximal non-adopter set): The unique maximal (1−Q)-subgroup-cohesive set in a society, constituting exactly the agents who do not adopt technology Q in the final outcome; it is the union of all (1−Q)-subgroup-cohesive sets and is itself (1−Q)-subgroup-cohesive (Lemma 2, Proposition 1).&lt;/p&gt;</description></item><item><title>The Dynamics of Verification when Searching for Quality</title><link>https://macropaperwarehouse.com/papers/the-dynamics-of-verification-when-searching-for-quality/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-dynamics-of-verification-when-searching-for-quality/</guid><description>&lt;p&gt;This paper develops a dynamic principal-agent model in which a principal seeks to select exactly one project from a stream of possibilities emerging over time, while a biased agent (who wants any project selected, regardless of quality) reports project quality each period. The principal cannot observe quality directly but can pay a cost c to verify it. Monetary transfers are unavailable. The central question is how verification and selection rules should optimally evolve over time as new options arrive.&lt;/p&gt;
&lt;p&gt;The model is set in discrete time with an infinite horizon (extended to finite horizons in Section 6.1). Each period, a project of quality h with probability q = λΔ or quality l with probability 1 − q arrives i.i.d. The principal selects at most once; the agent receives utility 1 from any selection and 0 otherwise; the principal&amp;rsquo;s payoff equals project quality net of verification costs. Both parties share discount factor δ = e^{−ρΔ}.&lt;/p&gt;
&lt;p&gt;When verification costs are low (c ≤ h) and the horizon is effectively infinite, the optimal mechanism exhibits decreasing skepticism: verification of high-quality reports occurs with a probability that is strictly declining over time, hitting zero at an endogenous deadline T* = ⌈(1−q)(δr − l) / (qc(1−δ))⌉. At that deadline, the principal selects any project irrespective of quality. Before the deadline, the agent reports truthfully — proposing only high-quality projects — and is incentivized by the threat of verification catching a lie, which triggers permanent exclusion. As the deadline approaches, the agent&amp;rsquo;s continuation value rises (guaranteed allocation arrives sooner), so the loss from a detected lie grows, and less verification is needed to deter misreporting. The deadline length is weakly increasing in h and r and decreasing in l and c; as c → 0, T* → ∞ and the principal&amp;rsquo;s payoff converges to the first-best of qh/(1−δ(1−q)).&lt;/p&gt;
&lt;p&gt;When verification costs are high (h &amp;lt; c &amp;lt; c̄, where c̄ is an explicitly computed threshold), deterministic selection is suboptimal. The optimal mechanism has two sequential phases: a randomization phase (periods 1 through T_R = ⌊log(h/c)/log(1−q)⌋ + 1) in which the principal randomizes between selecting and never selecting after a high-quality report without any verification, and a subsequent verification phase matching the low-cost structure. Verification is strictly backloaded: the principal never uses both tools simultaneously in the same period, and randomization always precedes verification. The intuition is that verification acts as a reward to the agent (guaranteeing allocation when h is realized), so delaying it allows earlier periods to exploit the prospect of future verification to relax incentive constraints across more periods, accumulating gains that justify the high verification cost.&lt;/p&gt;
&lt;p&gt;When the horizon is short (T ≤ T̄ := ⌊−(1−q)l/(qc)⌋) and l &amp;lt; 0 (static bias), increasing skepticism emerges: verification probability rises toward 1 in the final period. This occurs because a shrinking horizon reduces the agent&amp;rsquo;s continuation value, weakening the punishment for a detected lie, so more verification is required to maintain incentive compatibility. The paper also establishes that under renegotiation-proofness (Ray 1994), the optimal mechanism takes the same qualitative form as the full-commitment case but with permanent exclusion replaced by a mechanism restart. The leading application is board oversight of CEO-proposed acquisitions, motivated by the Smith v. Van Gorkom Delaware Supreme Court ruling; Graham et al. (2020) is cited as broad empirical support for decreasing oversight of CEOs over time.&lt;/p&gt;
&lt;p&gt;Q: What is the core agency conflict in the model?
A: The agent receives utility 1 from any selection regardless of quality, while the principal&amp;rsquo;s payoff equals quality minus verification costs. The agent always prefers immediate selection, while the principal prefers waiting for high quality, formalized by the condition qh + (1−q)l &amp;lt; qh/(1−δ(1−q)). This is &amp;ldquo;dynamic bias.&amp;rdquo; &amp;ldquo;Static bias&amp;rdquo; additionally arises when l &amp;lt; 0, meaning the principal prefers not allocating to allocating a low-quality project; this second source of conflict is more common in static settings.&lt;/p&gt;
&lt;p&gt;Q: What is the endogenous deadline T* and what determines its length?
A: T* = ⌈(1−q)(δr − l)/(qc(1−δ))⌉. It is weakly increasing in h and r (higher upside makes waiting worthwhile), weakly decreasing in l (a less costly low type shortens the horizon), and decreasing in c (cheaper verification makes longer search feasible). The term δr − l reflects the value of an additional quality draw relative to selecting low quality. As c → 0, T* → ∞ and the principal&amp;rsquo;s payoff converges to the first-best.&lt;/p&gt;
&lt;p&gt;Q: Why does the verification probability decline over time under decreasing skepticism?
A: As the deadline T* approaches, the agent&amp;rsquo;s continuation value from truthful play rises because guaranteed allocation is nearer. The loss from having a lie detected — permanent exclusion — therefore grows in absolute expected terms. Since more severe punishment requires less verification to deter misreporting, the minimum verification probability that satisfies the low type&amp;rsquo;s incentive compatibility constraint falls strictly over time, reaching zero exactly at T*.&lt;/p&gt;
&lt;p&gt;Q: When is randomization of the selection rule optimal, and when is verification strictly better?
A: Randomization is optimal if and only if c &amp;gt; h — when verification would guarantee a negative ex-post payoff for the principal. When c ≤ h, replacing randomization probability (1 − p̂_h) with verification probability x_h = 1 − δu_{t+1} maintains incentive compatibility while yielding a net gain to the principal proportional to h − c &amp;gt; 0 per period. The condition c &amp;gt; h is both necessary and sufficient for the randomization-augmented mechanism to dominate.&lt;/p&gt;
&lt;p&gt;Q: Why is verification backloaded when c &amp;gt; h?
A: Verification guarantees allocation whenever h is realized, which is a valuable reward for the agent. Deploying this reward later allows earlier randomization-phase periods to exploit the prospect of future verification to relax incentive constraints across multiple periods, accumulating gains. Moving verification earlier yields the same static cost but foregoes these accumulated gains; thus backloading verification is optimal. The principal never simultaneously randomizes and verifies in the same period.&lt;/p&gt;
&lt;p&gt;Q: What are the two phases in Theorem 2 and how long does each last?
A: The randomization phase runs from period 1 through T_R = ⌊log(h/c)/log(1−q)⌋ + 1; during this phase the principal randomizes allocation after a high-quality report (with the outside-option probability declining toward 0) but never verifies. The verification phase runs from T_R + 1 through a deadline at T* or T* + 1, with verification probability declining over time exactly as in Theorem 1. The total deadline is T* = T_R + ⌊(h − c − (l − δr)/(1−δ))(1−q)/(qc)⌋.&lt;/p&gt;
&lt;p&gt;Q: Under what conditions does increasing skepticism emerge?
A: Increasing skepticism arises when the horizon is finite and short — specifically when T ≤ T̄ = ⌊−(1−q)l/(qc)⌋, which requires l &amp;lt; 0 (static bias present). In this regime, verification probability rises to 1 in the final period T. Before T, the agent&amp;rsquo;s continuation value shrinks as fewer drawing opportunities remain, weakening the punishment for detected lies, so verification must increase to maintain incentive compatibility. Decreasing skepticism necessarily emerges only given a horizon long enough to overcome static bias.&lt;/p&gt;
&lt;p&gt;Q: How does the renegotiation-proofness extension modify the optimal mechanism?
A: Under renegotiation-proofness following Ray (1994), the mechanism cannot indefinitely withhold allocation following a detected lie, because both parties would prefer to restart rather than receive zero forever. The optimal renegotiation-proof mechanism takes the same qualitative form as Theorems 1 and 2, but permanent exclusion is replaced by a restart to the first period whenever a lie is verified during the verification phase or allocation is withheld during the randomization phase after a high-quality report. Deadlines, verification dynamics, and the phase structure are otherwise unchanged.&lt;/p&gt;
&lt;p&gt;Q: What is the three-region form of the value function?
A: Lemma 4 identifies thresholds u_low &amp;lt; u_high such that: for promised utility u ∈ [0, u_low], x_h(u) = 0 (no verification; only randomization); for u ∈ [u_low, u_high], dV/du = h − c (verification is interior, slope equals net benefit of verification); and for u &amp;gt; u_high, x_h(u) + y(u) = 1 (verification is at maximum). The slope h − c is constant on the middle region because increasing verification by ε raises promised utility by qε and the objective by q(h−c)ε, yielding a constant marginal rate.&lt;/p&gt;
&lt;p&gt;Q: What revelation-principle simplifications reduce the problem?
A: Lemmas 1–3 establish: (i) only high-type reports are ever verified (x_l = 0), since verification of the low type cannot improve principal payoffs; (ii) following verified truthfulness, allocation occurs with probability 1 (p*_{hh} = 1); (iii) the high type&amp;rsquo;s incentive constraint never binds in the optimal solution; and (iv) only the low type&amp;rsquo;s incentive compatibility constraint binds. These reduce the optimization to four free variables — x_h, p̂_h, p̂_l, û_l — subject to two binding constraints.&lt;/p&gt;
&lt;p&gt;Q: How does the paper relate to Kovac et al. (2013)?
A: The model builds most directly on Kovac et al. (2013)&amp;rsquo;s principal-agent stopping problem, which lacks costly verification. The key addition is the verification technology; the paper shows that when c ≤ h, verification eliminates the need for randomized selection rules that arise in Kovac et al. (2013). Kovac et al.&amp;rsquo;s randomization logic resurfaces in the randomization phase when c &amp;gt; h, and the analysis applies and extends Kovac et al.&amp;rsquo;s innovations.&lt;/p&gt;
&lt;p&gt;Q: What empirical and institutional evidence motivates the model?
A: The Smith v. Van Gorkom Delaware Supreme Court ruling (1985) established that boards must make meaningful efforts to become informed — exercising verification — as part of their duty of care in acquisition approvals; the TransUnion board was found negligent after approving an acquisition following a twenty-minute presentation with no written materials. Graham et al. (2020) provides broad empirical support for decreasing board oversight of CEOs over time, consistent with the paper&amp;rsquo;s decreasing skepticism prediction. Gompers et al. (2020) on VC analysts&amp;rsquo; project evaluation processes also illustrates the general applicability.&lt;/p&gt;
&lt;p&gt;Decreasing skepticism: The property of the optimal mechanism whereby the principal verifies high-quality reports with a probability that strictly declines over time, reaching zero at the endogenous deadline. Reflects diminishing concern about misrepresentation as the agent&amp;rsquo;s continuation value — and thus the cost of a detected lie — rises as the deadline approaches.&lt;/p&gt;
&lt;p&gt;Endogenous deadline (T*): The period at which the principal allocates any project irrespective of quality, ending the mechanism. Determined by T* = ⌈(1−q)(δr − l)/(qc(1−δ))⌉, balancing the value of waiting for additional quality draws against verification costs; weakly increasing in h and r, decreasing in l and c.&lt;/p&gt;
&lt;p&gt;Static bias vs. dynamic bias: Dynamic bias denotes the conflict that the principal prefers waiting for high quality while the agent prefers immediate selection. Static bias is the additional conflict (arising when l &amp;lt; 0) that the principal prefers withholding allocation to selecting a low-quality project, mirroring the agent-prefers-higher-action conflict in standard static models. Decreasing skepticism necessarily obtains absent static bias; static bias may flip dynamics to increasing skepticism if the horizon is short.&lt;/p&gt;
&lt;p&gt;Backloaded verification: The property that when c &amp;gt; h, verification is deployed only after a complete randomization phase, never simultaneously with randomization. Arises because verification acts as a reward to the agent by guaranteeing allocation when high quality is realized, and delaying this reward allows its incentive-relaxation benefits to compound across more randomization-phase periods.&lt;/p&gt;
&lt;p&gt;Randomization phase: The initial phase (periods 1 to T_R) in the high-cost regime, in which the principal randomizes the allocation decision after a high-quality report (outside option selected with declining probability) without using the verification technology. The randomization probability is set to keep the low type indifferent between truthful reporting and misreporting.&lt;/p&gt;
&lt;p&gt;Increasing skepticism: The opposite verification dynamic from decreasing skepticism, arising when the horizon is short (T ≤ T̄) and l &amp;lt; 0 (static bias). Verification probability rises over time toward 1 in the final period, because the agent&amp;rsquo;s continuation value shrinks as drawing opportunities dwindle, weakening the deterrent effect of detection and requiring more frequent verification to maintain incentive compatibility.&lt;/p&gt;
&lt;p&gt;Incentive compatibility via verification: The mechanism through which the principal deters low-type misreporting: by verifying a reported high-quality project with probability x_h, and punishing detected lies with permanent exclusion (or restart under renegotiation-proofness). This strictly dominates selection randomization when c ≤ h because the net per-period gain equals h − c &amp;gt; 0 while maintaining the same incentive compatibility condition for the low type.&lt;/p&gt;</description></item><item><title>The Effect of High-Tech Clusters on the Productivity of Top Inventors: Comment</title><link>https://macropaperwarehouse.com/papers/the-effect-of-high-tech-clusters-on-the-productivity-of-top-inventors-comment/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-effect-of-high-tech-clusters-on-the-productivity-of-top-inventors-comment/</guid><description>&lt;p&gt;This paper is a comment on Moretti (2021b), which studied agglomeration effects for innovation by testing whether the size of technology clusters causes patenting. The original paper (M21) used US patent data from 1971 to 2007 (Zucker and Darby, 2014) and reported a baseline elasticity of patenting with respect to cluster size of 0.0676, along with event study and instrumental variables (IV) evidence supporting a causal interpretation.&lt;/p&gt;
&lt;p&gt;Wiebe identifies two major methodological problems that undermine M21&amp;rsquo;s causal claims.&lt;/p&gt;
&lt;p&gt;Problem 1 — Misspecified event study. M21&amp;rsquo;s event study (Figure 6) was designed to test for selection bias from &amp;ldquo;rising star&amp;rdquo; inventors sorting into large clusters. The event is inventors moving across cities exactly once. However, M21&amp;rsquo;s specification interacts pre-move average cluster size with pre-move event-time indicators and post-move average cluster size with post-move event-time indicators separately — it does not exploit the change in cluster size generated by the move itself. Following the standard &amp;ldquo;mover&amp;rdquo; design literature (Finkelstein et al., 2016; Molitor, 2018; Cantoni and Pons, 2022), the correct specification uses the change in average cluster size as the treatment variable, interacted with event-time indicators. Wiebe implements this corrected event study and finds no statistically significant pre-trend and no statistically significant treatment effect post-move. Notably, the baseline elasticity estimated on the mover sample using all observed variation is large and significant at 0.3145 (SE 0.0953), but no effect is detected when variation is restricted to that generated by moving. The null result could also partly reflect attenuation bias from misclassified moves, since the dataset does not distinguish inventors who share the same name.&lt;/p&gt;
&lt;p&gt;Problem 2 — Coding error in IV. M21&amp;rsquo;s Table 5 instruments cluster size using variation in the number of inventors in other cities employed by firms also active in the focal inventor&amp;rsquo;s city, with the instrument calculated via first-differencing. Due to a coding error, M21 sorts data by firm, field, and year but not by city before first-differencing, so the differencing is taken across cities rather than within cities. Because firm-field-year is not a unique sorting key, Stata&amp;rsquo;s sort command pseudo-randomly orders observations with tied values, making the results unreproducible across runs. When Wiebe corrects the code to sort by city and compute first-differences within city, the 2SLS estimates become unstable and nonsignificant, with the first-stage F-statistic falling to approximately 7. This means M21 provides no valid IV evidence against confounding from city-field-year shocks such as local subsidies.&lt;/p&gt;
&lt;p&gt;Beyond these two major problems, the Appendix documents seven additional issues. The positive effect of cluster size on patent quality (M21 Table 6) disappears and reverses when the log transformation is corrected from log(y + 0.00001) to log(y + 1) or Poisson regression — the corrected estimate is negative and significant, implying that cluster size reduces citations per patent along the intensive margin and the overall quality effect is negative. Heterogeneous elasticity estimates (M21 Table 8) contain a coding error; corrected estimates show substantial heterogeneity. The distributed lag model (M21 Figure 5) uses an incorrectly defined lag structure in an unbalanced panel; corrected estimates yield nonsignificant contemporaneous effects. Cluster quality estimates (M21 Table A.8) use a cluster size definition differing from the text, and corrected elasticities are approximately half as large. M21&amp;rsquo;s claimed extensive margin effect in Table A.7 is logically unsupported since no zeros are observed. The team size robustness check is conceptually flawed because it controls twice for per-coauthor adjustment. A gap-interpolation coding error in Table A.6 biases estimates downward. Broader computational reproducibility failures arise from many-to-many merges with non-unique sort orders. Wiebe explicitly notes that the null IV and event study results are not evidence against agglomeration effects per se.&lt;/p&gt;
&lt;p&gt;Q: What is the baseline finding in M21 that Wiebe contests?
A: M21 reports a baseline elasticity of patenting with respect to cluster size of 0.0676, estimated from linear regressions with extensive fixed effects including inventor fixed effects. M21 presents an event study and IV strategy as additional evidence supporting a causal interpretation of this elasticity.&lt;/p&gt;
&lt;p&gt;Q: What is wrong with M21&amp;rsquo;s event study specification?
A: M21&amp;rsquo;s event study interacts pre-move average cluster size with pre-move event-time indicators and post-move average cluster size with post-move event-time indicators, but never uses the change in cluster size associated with moving. The standard mover design (Finkelstein et al., 2016; Molitor, 2018) uses the change in average environment as a constant treatment variable interacted with all event-time indicators. Because M21&amp;rsquo;s specification does not exploit moving-induced variation, it would be identified even if moving induced no change in cluster size.&lt;/p&gt;
&lt;p&gt;Q: What does Wiebe&amp;rsquo;s corrected event study find?
A: Wiebe&amp;rsquo;s corrected mover event study shows no statistically significant pre-trend (consistent with no systematic sorting of rising-star inventors into large clusters) and no statistically significant post-move treatment effect. In contrast, the baseline fixed-effects elasticity on the mover sample using all observed variation is 0.3145 (SE 0.0953) — large and significant — indicating the null result is specific to the moving-generated variation.&lt;/p&gt;
&lt;p&gt;Q: What alternative explanation does Wiebe offer for the null event study result?
A: The null result could be partly explained by attenuation bias from misclassified moves. M21&amp;rsquo;s code creates inventor identifiers based on names, but the COMETS dataset does not distinguish inventors who share the same name, so an apparent cross-city move may simply be two different inventors with the same name living in different cities.&lt;/p&gt;
&lt;p&gt;Q: What is the coding error in M21&amp;rsquo;s IV strategy?
A: M21 constructs the instrument by first-differencing a variable measuring inventors in other cities working for firms also active in the focal city. The code sorts by firm, field, and year before differencing, but omits city from the sort key, so first-differencing is computed across cities rather than within cities, generating an instrument that does not match the definition in the text.&lt;/p&gt;
&lt;p&gt;Q: Why does the coding error also cause non-reproducibility?
A: Firm-field-year is not a unique sorting key because multiple cities can share the same firm-field-year values. Stata&amp;rsquo;s sort command pseudo-randomly orders observations with tied values, so each run produces a different city ordering within tied groups and therefore a different instrument and different estimates.&lt;/p&gt;
&lt;p&gt;Q: What do the corrected IV results show?
A: After correcting the sort order to include city and computing first-differences within city, the 2SLS estimates are unstable and nonsignificant. The first-stage F-statistic falls to approximately 7, indicating a weak instrument. This does not constitute evidence against agglomeration effects, but means M21&amp;rsquo;s IV strategy provides no valid evidence against confounding from city-field-level shocks such as local subsidies.&lt;/p&gt;
&lt;p&gt;Q: What happens to the patent quality results when the log transformation is corrected?
A: M21 uses log(citations + 0.00001), which assigns very large weight to the extensive margin. When Wiebe uses log(citations + 1) or Poisson regression instead, the estimated effect of cluster size on patent quality is negative and statistically significant, reversing M21&amp;rsquo;s finding. The corrected result implies that while cluster size may raise the probability of producing any cited patent, it reduces citations per patent for inventors who do produce cited patents, and the overall effect is negative.&lt;/p&gt;
&lt;p&gt;Q: What are the corrected aggregate agglomeration loss estimates?
A: Using the corrected constant elasticity, the estimated output reduction from equalizing cluster sizes is -9.15% (slightly smaller than M21). Using corrected heterogeneous elasticities based on within-field-year size quartiles, the output loss is -23.75% (about twice as large). Using elasticities based on global size quartiles, the loss is -35.11%.&lt;/p&gt;
&lt;p&gt;Q: What is wrong with M21&amp;rsquo;s distributed lag model (Figure 5)?
A: M21&amp;rsquo;s code defines lags and leads using sequential observations in the panel rather than calendar years. Because the inventor-year panel is unbalanced, a coded &amp;ldquo;one-year lag&amp;rdquo; can refer to any number of years prior. When Wiebe restricts to inventors with 11 consecutive years and correctly defines year-based lags, confidence intervals widen substantially and the contemporaneous effect estimate becomes nonsignificant.&lt;/p&gt;
&lt;p&gt;Q: What is the conceptual flaw in M21&amp;rsquo;s team-size robustness check?
A: M21&amp;rsquo;s Table A.8 controls for the number of coauthors on a patent, but the dependent variable is already measured as patents per coauthor. Controlling for team size after already dividing by team size effectively controls for the same variable twice.&lt;/p&gt;
&lt;p&gt;Q: What are the broader computational reproducibility problems in M21?
A: The cleaning code uses many-to-many merges with non-unique sort orders, generating slightly different datasets on each run. For example, when merging inventors with patent assignees, patent identifiers are not unique because multiple firms can be assigned to a single patent. Removing name suffixes also causes distinct inventors (e.g., Paul H. Hamisch Jr. and Sr.) to be assigned the same identifier. Additionally, using reghdfe with the keepsingletons option retains singleton groups explicitly warned against by the package due to biased standard errors.&lt;/p&gt;
&lt;p&gt;Agglomeration elasticity: The elasticity of an inventor&amp;rsquo;s patent output with respect to the size of the technology cluster (city-field-year cell) in which they work; reported as 0.0676 in M21&amp;rsquo;s baseline and 0.3145 on the mover sample with all observed variation.&lt;/p&gt;
&lt;p&gt;Mover event study design: An event study specification in which the treatment variable is the change in an individual&amp;rsquo;s average environment (here, cluster size) before and after a geographic move, interacted with event-time indicators — the standard design used in Finkelstein et al. (2016) and Molitor (2018), which M21&amp;rsquo;s specification does not follow.&lt;/p&gt;
&lt;p&gt;Cluster size: The number of inventors (or cluster density) active in the same city-field-year cell as the focal inventor, used as the key independent variable in M21&amp;rsquo;s regressions.&lt;/p&gt;
&lt;p&gt;First-stage F-statistic: A measure of instrument strength in 2SLS IV estimation; the corrected instrument yields F ≈ 7 (indicating weakness), whereas M21&amp;rsquo;s incorrectly constructed instrument produced a stronger first stage by exploiting spurious cross-city variation.&lt;/p&gt;
&lt;p&gt;Extensive vs. intensive margin (patent quality): The extensive margin captures whether an inventor produces any cited patent; the intensive margin captures citations per patent conditional on having any. M21&amp;rsquo;s log(y + 0.00001) transformation overweights the extensive margin, and the corrected intensive-margin effect of cluster size on quality is negative and significant.&lt;/p&gt;
&lt;p&gt;Computational reproducibility: The property that running code on the same data produces identical results across runs. M21&amp;rsquo;s code fails this standard due to non-unique sort orders in merges and first-differencing steps, causing the IV instrument to differ across runs.&lt;/p&gt;
&lt;p&gt;Rising star sorting: The hypothesized selection mechanism whereby inventors with increasing patent trajectories are preferentially hired into large clusters, which would bias OLS agglomeration elasticity estimates upward; M21&amp;rsquo;s event study was designed to test for this but is incorrectly specified and does not use moving-induced variation.&lt;/p&gt;</description></item><item><title>The Environmental Bias of Corporate Income Taxation</title><link>https://macropaperwarehouse.com/papers/the-environmental-bias-of-corporate-income-taxation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-environmental-bias-of-corporate-income-taxation/</guid><description>&lt;p&gt;This paper documents and quantifies an &amp;ldquo;environmental bias&amp;rdquo; embedded in the U.S. corporate income tax code: CO2-intensive (&amp;ldquo;dirty&amp;rdquo;) firms systematically face lower effective tax rates than clean firms, constituting an implicit subsidy on pollution. The authors — Iovino, Martin, and Sauvagnat — establish this cross-sectional fact, trace it to a specific mechanism, provide causal evidence using the 2017 Tax Cuts and Jobs Act (TCJA), and quantify aggregate emissions implications using a calibrated multi-sector general-equilibrium model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and sample.&lt;/strong&gt; The empirical analysis combines firm-level CO2 emissions from Trucost (scope 1 greenhouse gases) with financial data from Compustat North America for U.S. publicly listed firms, 2003–2021, yielding 11,223 firm-year observations with positive pretax and gross capital income. Effective tax rates are measured as income taxes paid divided by gross capital income (sales minus COGS minus SGA expenses, adding back R&amp;amp;D).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cross-sectional finding.&lt;/strong&gt; A one-standard-deviation increase in CO2 intensity is associated with a decrease in the effective tax rate equal to approximately 9% of its standard deviation (coefficient −0.021 to −0.022, significant at 1%). The negative relationship is entirely explained by the lower taxable fraction of gross capital income for dirty firms — that is, by larger interest expense deductions — rather than by differences in the statutory tax rate applied to pretax income.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Mechanism.&lt;/strong&gt; The chain of causation runs: CO2-intensive production requires tangible capital (primarily machinery and equipment) → tangible capital serves as collateral → higher collateral supports higher debt → higher debt generates larger interest deductions (the &amp;ldquo;tax shield of debt&amp;rdquo;) → lower effective tax rates. Once PPE-to-capital-income is controlled for, the coefficient on CO2 intensity in leverage, pretax income, and tax regressions becomes small and statistically insignificant. The relationship holds both across and within industries, including within the energy sector, though the dominant variation is cross-industry.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Causal evidence: TCJA 2017.&lt;/strong&gt; The paper exploits the federal corporate tax rate cut from 35% to 21% (effective January 2018) in a difference-in-differences design, comparing firms in the top quartile of 2017 CO2 intensity (&amp;ldquo;dirty&amp;rdquo;) to cleaner firms. Dirty firms experienced a relative increase in their federal effective tax rate of 2.4 percentage points post-reform. Correspondingly, dirty firms&amp;rsquo; total assets grew approximately 11% less than clean firms post-reform. This translates to a semi-elasticity of firm total assets to a one-percentage-point increase in the effective tax rate of approximately −4.8. Parallel pre-trends are confirmed visually and via Rambachan-Roth (2023) sensitivity analysis; a placebo using non-federal taxes shows no differential effect. Results survive controls for other TCJA provisions (interest deductibility limits, international tax changes, net operating loss restrictions), exposure to import tariffs and carbon taxes, leave-one-industry-out specifications, and a triple-difference using foreign firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;General-equilibrium model and counterfactuals.&lt;/strong&gt; A 375-sector model with input-output networks (both intermediate and investment networks), financial frictions linking equipment to debt capacity, and endogenous CO2 emissions through fossil fuel usage is calibrated to 2017 BEA and Compustat data. In the Cobb-Douglas benchmark, the 2017 tax cut raises output by 5.9% and emissions by only 4.5% — a less-than-proportional emissions response because clean sectors expand relatively more. A counterfactual eliminating the tax shield of debt while simultaneously cutting the tax rate from 35% to 30% (to hold GDP constant) reduces aggregate emissions by 1.3% with output declining only 0.1%. When equipment and fuel are treated as complements (elasticity of substitution below 1), the emissions reduction under the same policy rises to over 3.7%, implying an absolute reduction of 80–240 million metric tons of CO2 from 2017&amp;rsquo;s total of 6,457 million metric tons. Monetized at the social cost of carbon, this ranges from USD 8–24 billion (conservative, ~USD 100/ton) to USD 112–336 billion (USD 1,400/ton per Bilal and Kanzig 2024).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the central empirical finding of the paper?&lt;/strong&gt;
A: CO2-intensive firms in the U.S. face systematically lower effective corporate income tax rates than clean firms. A one-standard-deviation increase in CO2 intensity is associated with a roughly 9% of a standard deviation decrease in the ratio of taxes paid to gross capital income. This negative relationship is robust to alternative emissions measures (EPA data, scope 2 and 3 emissions), alternative tax scalings (taxes over sales or assets), log CO2 emissions, and leave-one-industry-out specifications.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the mechanism linking CO2 intensity to lower effective tax rates?&lt;/strong&gt;
A: Dirty firms rely on tangible capital — specifically machinery and equipment — to produce. Tangible capital is pledgeable as collateral, enabling higher debt. Higher debt generates larger interest expense deductions under the tax code (the &amp;ldquo;debt tax shield&amp;rdquo;), which reduces taxable income relative to gross capital income. Once PPE-to-capital-income is included as a control, the coefficient on CO2 intensity in regressions of leverage, pretax income, and taxes paid all become small and statistically insignificant, confirming that PPE fully mediates the relationship.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Which component of tangible capital drives the result?&lt;/strong&gt;
A: Machinery and equipment, not buildings, leases, land, natural resources, or construction in progress, explains virtually the entire positive relationship between PPE and CO2 intensity. This finding is based on the Compustat breakdown of PPE components available for roughly 70% of sample firms.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: Does the mechanism operate within industries or only across them?&lt;/strong&gt;
A: Both. Decomposing firm CO2 intensity into an implied industry component (sales-weighted from pure-play firms) and a firm residual, both components are significantly associated with higher tangible capital, leverage, lower taxable fraction of capital income, and lower taxes paid at the 1% level. However, the largest share of the total effect stems from cross-industry variation. Within the energy sector specifically, firms with greater fossil fuel production capacity (from EPA/EIA data) also have more tangible capital, higher debt, and lower effective tax rates.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the 2017 TCJA cut affect clean versus dirty firms differently?&lt;/strong&gt;
A: Because dirty firms already shield a large fraction of their capital income from taxation via interest deductions, a uniform cut in the statutory rate benefits them less in proportional terms. The difference-in-differences estimates show that dirty firms (top quartile of 2017 CO2 intensity) experienced a relative increase in their federal effective tax rate of 2.4 percentage points post-reform compared to clean firms, and their total assets grew approximately 11% less than clean firms post-reform. The semi-elasticity of firm assets to a one-percentage-point increase in effective tax rate is approximately −4.8.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How is the parallel trends assumption supported?&lt;/strong&gt;
A: Event-study graphs show no pre-2018 divergence in federal effective tax rates or asset growth between dirty and clean firms. A placebo test using non-federal income taxes (which should be unaffected by the federal statutory rate change) shows no differential post-reform effect. The Rambachan-Roth (2023) sensitivity analysis confirms that the null of no differential effect can be rejected at the 1% level allowing for pre-trend deviations up to M = 0.5, and at the 10% level up to M = 1.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What robustness checks address other provisions of the TCJA and concurrent shocks?&lt;/strong&gt;
A: The authors exclude or control for firms affected by the TCJA&amp;rsquo;s interest deductibility limitation, multinational firms (more than 20% foreign sales), firms with large loss carryforwards, and manufacturing firms — results are unchanged. They also control for firm-level exposure to import tariff changes and carbon taxes (using the World Carbon Pricing Database), with coefficients of interest remaining virtually unchanged. Leave-one-industry-out specifications and a triple-difference using foreign firms (comparing U.S. dirty vs. clean firms pre/post-2018, against foreign equivalents in countries with stable tax rates) yield a semi-elasticity of −5.8, if anything larger than the baseline.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the general-equilibrium model add that the difference-in-differences cannot?&lt;/strong&gt;
A: The DiD design identifies relative effects of the tax cut on dirty versus clean firms but cannot recover the absolute effect on aggregate output and emissions. The GE model, calibrated to 2017 data and validated against the untargeted DiD estimates, quantifies aggregate impacts: the 2017 tax cut raises steady-state output by 5.9% while emissions rise by only 4.5% — a less-than-proportional increase due to compositional reallocation toward clean sectors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What does the counterfactual removing the debt tax shield find?&lt;/strong&gt;
A: Eliminating the tax shield of debt while simultaneously lowering the corporate tax rate from 35% to 30% (to keep GDP constant) reduces aggregate emissions by 1.3% (Cobb-Douglas benchmark) while total output falls only 0.1% and GDP remains constant by design. The emissions reduction arises because clean sectors, which rely more on less-pledgeable capital, are made relatively cheaper once the tax advantage of debt is removed, redirecting demand away from CO2-intensive sectors.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does the complementarity assumption between equipment and fuel affect the results?&lt;/strong&gt;
A: When equipment and fuel are modeled as complements (elasticity of substitution below 1) rather than Cobb-Douglas substitutes, both policy counterfactuals yield larger emissions effects. For the tax shield removal policy, the predicted emissions reduction rises from 1.3% to over 3.7% as complementarity strengthens. This is because policies that raise the cost of equipment also induce firms to cut fuel consumption, amplifying the direct compositional effect.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the quantified absolute emissions impact of removing the tax shield?&lt;/strong&gt;
A: Given 2017 U.S. total emissions of 6,457 million metric tons, the model predicts an absolute reduction of 80–240 million metric tons of CO2, depending on the assumed complementarity between equipment and fuel. Monetized at conservative estimates (~USD 100/ton), the policy saves USD 8–24 billion; at USD 1,400/ton (Bilal and Kanzig 2024), the value rises to USD 112–336 billion. The authors note that the physical quantity measure is more reliable than the monetized figure given uncertainty in the social cost of carbon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: How does this paper relate to the ECB bond purchasing literature?&lt;/strong&gt;
A: Piazzesi et al. (2022) document that the ECB&amp;rsquo;s market-neutral bond purchases implicitly favor dirty firms because those firms issue more bonds due to higher tangible capital holdings. This paper identifies the same underlying mechanism — tangible capital → debt capacity — but on the tax side, showing that the corporate income tax code independently provides an implicit subsidy to dirty firms through the debt tax shield.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Q: What is the policy implication for the debt tax shield specifically?&lt;/strong&gt;
A: The debt tax shield — the deductibility of interest payments but not dividends — has no clear economic rationale (both are returns to capital) and, per several policy proposals (CBO 1997, IMF 2016), is a candidate for elimination. This paper adds a new dimension: the tax shield indirectly subsidizes CO2 emissions by differentially benefiting capital-intensive, CO2-intensive sectors. A revenue-neutral reform eliminating the shield can reduce emissions without sacrificing GDP.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Effective tax rate (paper&amp;rsquo;s definition):&lt;/strong&gt; The ratio of corporate income taxes paid to gross capital income, where gross capital income equals sales minus cost of goods sold minus SGA expenses plus R&amp;amp;D spending. This differs from the tax-to-pretax-income ratio because it captures how much of total capital earnings — before any deductions — is remitted as tax.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Debt tax shield (tax advantage of debt):&lt;/strong&gt; The reduction in corporate tax liability arising from the deductibility of interest payments on corporate debt. Because dividends are not deductible, debt-financed capital faces a lower after-tax cost than equity-financed capital. The shield&amp;rsquo;s value is estimated at approximately 10% of firm value in prior literature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;CO2 intensity:&lt;/strong&gt; Metric tons of CO2 equivalent per USD 1,000 of output (tCO2/k$). The sample average is 0.1 tCO2/k$, with a heavily right-skewed distribution (median 0.02, 99th percentile 1.5).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Environmental bias of corporate taxation:&lt;/strong&gt; The paper&amp;rsquo;s central concept — the systematic difference in effective tax rates between dirty and clean firms that arises not from explicit environmental policy but from the interaction of the debt tax shield with the capital structure of CO2-intensive industries. This constitutes an implicit subsidy on pollution embedded in the corporate income tax.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Asset pledgeability (psi):&lt;/strong&gt; The fraction of a firm&amp;rsquo;s assets recoverable by creditors in the event of default. In the model, equipment has higher pledgeability than other capital (estimated b_psi = 0.23 additional pledgeability for equipment, a_psi = 0.35 base). Higher pledgeability allows firms to sustain more debt and thus benefit more from the tax shield.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;User cost of capital:&lt;/strong&gt; The total cost to a firm of using one unit of capital, combining depreciation, tax allowances from accelerated depreciation, and the financing cost advantage of debt over equity. The model formalizes how both the equity-financed component and the debt advantage component respond to tax rate changes, with the debt advantage term being larger for firms with more pledgeable (tangible) capital.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Investment network:&lt;/strong&gt; An input-output structure capturing which sectors&amp;rsquo; outputs are used to produce each type of capital good. The paper extends vom Lehn and Winberry (2021) by constructing separate equipment and non-equipment investment networks across 375 non-fuel BEA sectors, enabling emissions accounting that includes capital production alongside direct production inputs.&lt;/p&gt;</description></item><item><title>The Macroeconomics of Irreversibility</title><link>https://macropaperwarehouse.com/papers/the-macroeconomics-of-irreversibility/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-macroeconomics-of-irreversibility/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research question.&lt;/strong&gt; How does partial capital irreversibility — arising from a wedge between the purchase price and the resale (discounted) price of capital — shape the persistence and amplitude of aggregate capital fluctuations? And what is the quantitative magnitude of the capital price wedge that is needed to simultaneously reconcile micro-level investment behavior with macroeconomic propagation?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Methodology.&lt;/strong&gt; Baley and Blanco build a continuous-time investment model for a continuum of firms facing (i) idiosyncratic productivity shocks (geometric Brownian motion), (ii) fixed capital adjustment costs proportional to productivity, and (iii) a capital price wedge ω, under which firms buy capital at price p and sell at p(1−ω). The key state variable is the log capital-productivity ratio k̂. The optimal policy takes the form of an inaction region with two distinct reset points — one for upsizing (k̂*₋) and one for downsizing (k̂*₊) — instead of the single reset point that arises without the wedge.&lt;/p&gt;
&lt;p&gt;Their central innovation is the Cumulative Impulse Response (CIR): the cumulative deviation of average capital-productivity ratios following a small, permanent, unanticipated aggregate productivity shock. They show the CIR can be expressed analytically through three sufficient statistics derived entirely from the steady-state cross-sectional distribution of k̂ and capital age a: (i) Var[k̂], (ii) Cov[k̂, a], and (iii) an &amp;ldquo;irreversibility term&amp;rdquo; reflecting how idiosyncratic shocks change the anticipated direction of the next adjustment. Because idiosyncratic and aggregate shocks enter the law of motion symmetrically, steady-state moments encode the aggregate propagation.&lt;/p&gt;
&lt;p&gt;To handle the path dependence introduced by the dual reset points, they condition all behavior on the previous reset (upsizing or downsizing) and characterize transitions across reset points via a Markov chain. They then derive explicit mappings from observable microdata — size and direction of investment adjustments, duration of inaction spells, and cross-spell transition probabilities — back to the unobservable capital-productivity distributions and sufficient statistics. These mappings require no revenue or productivity data; investment actions alone suffice.&lt;/p&gt;
&lt;p&gt;They extend the baseline model to a generalized hazard framework (stochastic, asymmetric fixed costs), enabling the model to match the full empirical investment-rate distribution, and apply everything to annual establishment-level manufacturing data from Chile (Encuesta Nacional Industrial Anual, 1980–2011), restricting to plants observed for at least ten years with more than ten workers.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main findings.&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Price wedge estimate.&lt;/strong&gt; A capital price wedge of ω = 0.12 (12%) is selected as the preferred value because it maximizes joint consistency between the model&amp;rsquo;s predicted CIR decomposition and the data, while also matching the distribution of investment rates. At ω = 0 the model generates a CIR of 0.92 and a negative covariance term, inconsistent with the data. At ω = 0.18 the aggregate CIR level (2.39) is close to data (2.33) but the decomposition diverges. At ω = 0.12, the CIR is 1.93 and the decomposition into sufficient statistics closely mirrors the data structure.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Irreversibility doubles persistence.&lt;/strong&gt; In the analytically tractable case of zero drift and only a price wedge (no fixed costs), the CIR equals exactly twice the ratio Var[k̂]/σ², compared to the single fixed-cost case. This means irreversibility doubles the persistence of aggregate capital fluctuations for a given cross-sectional dispersion. More generally, under the calibrated model, a 1% decrease in aggregate productivity generates a nearly 2% cumulative deviation of average capital-productivity ratios from steady state. Without irreversibility, the CIR collapses to approximately 1.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Decomposition of the CIR.&lt;/strong&gt; At ω = 0.12, the variance term Var[k̂]/σ² accounts for 72% of the CIR; the covariance term ν·Cov[k̂,a]/σ² accounts for 10%; and the irreversibility term accounts for 18%. The positive covariance (Cov[k̂,a] = 0.152 &amp;gt; 0) reflects that firms subject to downward rigidity accumulate older capital stocks above the economy&amp;rsquo;s average, amplifying persistence. This positive covariance arises because the price wedge&amp;rsquo;s downward-rigidity force dominates the drift&amp;rsquo;s negative effect.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Micro-level evidence.&lt;/strong&gt; In the Chilean data, the inaction rate is 40%. More than 96% of adjustments are positive (upsizing), fewer than 4% are negative. The probability of upsizing after a previous upsize is P⁻⁻ = 0.958; the probability of downsizing after a downsize is P⁺⁺ = 0.124. A logistic regression yields an odds ratio of 3.3, meaning a firm is more than three times as likely to purchase capital following a prior purchase than following a prior sale. The average duration of inaction conditional on a prior purchase is E⁻[τ] = 1.72 years; conditional on a prior sale it is E⁺[τ] = 1.98 years. These patterns are qualitatively consistent with the serial correlation in adjustment sign predicted by the model.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Comparison with existing wedge estimates.&lt;/strong&gt; The calibrated ω = 0.12 lies between micro-level studies based on liquidating firms (Ramey and Shapiro, 2001: ω ≈ 0.72; Kermani and Ma, 2023: ω ≈ 0.65) and structural models calibrated to static moments of investment distributions (Cooper and Haltiwanger, 2006; Khan and Thomas, 2013: ω = 0.025–0.07). The lower value relative to liquidation studies is attributed to selection effects (liquidating firms face fire-sale dynamics) and firm-internal capital reallocation that mitigates irreversibility for continuing firms.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Scope conditions.&lt;/strong&gt; The analysis is a partial equilibrium characterization of transitional dynamics, maintaining constant interest rates and steady-state investment policies throughout the transition (a general equilibrium extension delivering constant prices as an equilibrium outcome is provided in Appendix D). Results apply to small, permanent, unanticipated aggregate productivity shocks; nonlinearities for shocks below 5% are found to be tiny. The empirical application is specific to Chilean manufacturing establishments, 1980–2011.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-economic-mechanism-by-which-capital-irreversibility-generates-persistence-in-aggregate-capital-fluctuations"&gt;Q1. What is the economic mechanism by which capital irreversibility generates persistence in aggregate capital fluctuations?&lt;/h3&gt;
&lt;p&gt;Irreversibility creates two distinct reset points rather than one. When a negative aggregate productivity shock hits, it shifts more firms into the downsizing region. Downsizing firms, because they have been selling capital sequentially, maintain capital-productivity ratios persistently above the economy&amp;rsquo;s average and continue to do so for multiple periods. This increases the share of firms in a persistent &amp;ldquo;downsizing phase,&amp;rdquo; which prolongs the aggregate deviation from steady state. Two channels compound: first, the population tilts toward more downsizing firms; second, their mean deviations become larger and converge more slowly. Both channels increase the CIR. Crucially, without irreversibility, firms become identical after their first adjustment and there is no additional persistence beyond what fixed costs alone generate.&lt;/p&gt;
&lt;h3 id="q2-how-are-the-three-sufficient-statistics-derived-and-what-does-each-capture"&gt;Q2. How are the three sufficient statistics derived, and what does each capture?&lt;/h3&gt;
&lt;p&gt;The CIR is characterized as a steady-state cross-sectional average of a recursive function m(k̂). Integrating over firms first and then time, and splitting each firm&amp;rsquo;s horizon at its first adjustment, yields three steady-state terms (Proposition 4). The first statistic, Var[k̂]/σ², measures how far firms allow their capital-productivity ratio to drift from the frictionless optimum — the &amp;ldquo;insensitivity of incomplete spells&amp;rdquo; to idiosyncratic productivity shocks. The second statistic, ν·Cov[k̂,a]/σ², is a bias-correction term that removes drift effects from the variance, ensuring only Brownian-shock sensitivity is captured. The third statistic, unique to the irreversibility case, measures how much idiosyncratic shocks alter the anticipated direction of the next adjustment — the &amp;ldquo;insensitivity of complete spells&amp;rdquo; — and equals the difference in expected cumulative deviations between departing and ending points of an inaction spell, scaled by duration.&lt;/p&gt;
&lt;h3 id="q3-why-is-the-cir-exactly-twice-as-large-under-pure-irreversibility-no-fixed-costs-as-under-pure-fixed-costs-for-a-given-level-of-dispersion"&gt;Q3. Why is the CIR exactly twice as large under pure irreversibility (no fixed costs) as under pure fixed costs, for a given level of dispersion?&lt;/h3&gt;
&lt;p&gt;Proposition 5, case (ii) shows that with zero drift and only a price wedge, the CIR = 2 × Var[k̂]/σ², because the first and third sufficient statistics are identical and the covariance term is zero. In contrast, with only fixed costs (case (i)), the CIR = Var[k̂]/σ². The doubling arises because the price wedge generates history-dependence through the dual reset: after a firm adjusts, whether it upsized or downsized predicts its future adjustment direction. This &amp;ldquo;anticipated terminal condition&amp;rdquo; effect (captured by the third statistic) adds an equal contribution to the CIR as the pure inaction effect (the first statistic), doubling total persistence for the same cross-sectional dispersion.&lt;/p&gt;
&lt;h3 id="q4-how-does-the-empirical-strategy-recover-the-capital-price-wedge"&gt;Q4. How does the empirical strategy recover the capital price wedge?&lt;/h3&gt;
&lt;p&gt;The price wedge cannot be identified from the investment rate distribution alone: for any price wedge ω, the generalized hazard framework can find an adjustment hazard function Λ(k̂) such that the product Λ(k̂)·g(k̂) matches the observed investment density h(Δk̂). Instead, the authors use the CIR&amp;rsquo;s sufficient statistics — specifically the covariance term and the irreversibility term — as additional discriminating moments. At ω = 0, the model produces a negative covariance (inconsistent with the positive Cov[k̂,a] = 0.152 in the data) and no irreversibility term. At ω = 0.12, all three sufficient statistics simultaneously align with their data counterparts in relative importance (72%, 10%, 18%), selecting this wedge as preferred. The CIR level at ω = 0.12 is 1.93, somewhat below the data value of approximately 2.54–2.60, but the preferred criterion is mechanistic consistency, not just level matching.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-role-of-the-markov-chain-across-reset-points-in-handling-path-dependence"&gt;Q5. What is the role of the Markov chain across reset points in handling path dependence?&lt;/h3&gt;
&lt;p&gt;Because optimal investment features serial correlation in the sign of adjustment (P⁻⁻ = 0.958 and P⁺⁺ = 0.124 in the data), firms&amp;rsquo; future behavior depends on their most recent reset point. To maintain tractability, the authors condition all densities, durations, and expectations on the previous reset (upsizing g⁻(k̂) or downsizing g⁺(k̂)). The transition matrix P encoding probabilities P⁻⁻, P⁻⁺, P⁺⁻, P⁺⁺ determines the steady-state shares of upsizing and downsizing firms (as the eigenvector of P) and the renewal weights r⁻ and r⁺ that rescale conditional densities to account for observational bias (firms with longer inaction spells contribute more to the cross-section). This Markov structure is sufficient because one adjustment erases all heterogeneity except the direction of adjustment.&lt;/p&gt;
&lt;h3 id="q6-what-do-the-microdata-mappings-recover-and-how-are-the-reset-points-identified"&gt;Q6. What do the microdata mappings recover, and how are the reset points identified?&lt;/h3&gt;
&lt;p&gt;Stage I mappings (Propositions 6–9) recover: drift ν = E[Δk̂]/E[τ]; volatility σ² from cross-spell moment E[(k̂τ&amp;rsquo; + ντ&amp;rsquo;)² − (k̂*)²]/E[τ]; conditional means E±[k̂] as midpoints of inaction spells weighted by relative adjustment size; Var[k̂] from differences in cubed stopped values; Cov[k̂,a] from variance, average age, and the dynamic covariance E[(k̂τ&amp;rsquo; − E[k̂])²τ&amp;rsquo;]/E[τ]; and the irreversibility term from differences in expected deviations at departing vs. ending reset points. Stage II (Proposition 10) recovers the two reset points k̂*₋ and k̂*₊ from optimality conditions that equalize the investment price to the expected discounted marginal product of capital during inaction plus the expected value of undepreciated capital, conditioning on the prior reset. The inner inaction region width k̂*₊ − k̂*₋ = 0.813 in the Chilean data, of which 45% is attributed to the exogenous price wedge and 55% to the endogenous response to the wedge.&lt;/p&gt;
&lt;h3 id="q7-how-does-the-sign-of-covka-depend-on-the-price-wedge-vs-the-drift"&gt;Q7. How does the sign of Cov[k̂,a] depend on the price wedge vs. the drift?&lt;/h3&gt;
&lt;p&gt;With zero price wedge and negative drift ν &amp;lt; 0 (depreciation exceeding productivity growth), firms with older capital have capital-productivity ratios below average, yielding Cov[k̂,a] &amp;lt; 0. The drift makes old capital-productivity ratios negative. Introducing a price wedge creates downward rigidity: unproductive firms delay selling, so old firms accumulate capital-productivity ratios above average, pushing Cov[k̂,a] toward positive values. The covariance turns positive once ω &amp;gt; 0.08 (in the illustrative parametrization in Figure V). In the Chilean calibration at ω = 0.12, Cov[k̂,a] = 0.152 &amp;gt; 0, confirming that the price wedge&amp;rsquo;s effect dominates the drift&amp;rsquo;s negative effect. A positive covariance amplifies the CIR (through the second sufficient statistic with ν &amp;gt; 0).&lt;/p&gt;
&lt;h3 id="q8-what-is-the-generalized-hazard-extension-and-why-is-it-needed"&gt;Q8. What is the generalized hazard extension and why is it needed?&lt;/h3&gt;
&lt;p&gt;The baseline model with a single fixed cost θ generates an investment distribution concentrated at two mass points (purchases and sales of fixed size), which does not match the empirical distribution&amp;rsquo;s coexistence of large and small investment rates and its convex shape. The generalized hazard model replaces the deterministic fixed cost with a stochastic, state-dependent adjustment cost, parameterized by a hazard function Λ(k̂) giving the probability of adjusting per unit time at any capital-productivity ratio in the outer inaction region. This function is recovered non-parametrically from the data by fitting a Gamma distribution to the investment density and inverting the Kolmogorov Forward Equation. The generalized hazard model nests the baseline model, random fixed cost models (Thomas 2002, Khan and Thomas 2008), and asymmetric adjustment models, while preserving the sufficient statistics characterization.&lt;/p&gt;
&lt;h3 id="q9-how-does-the-model-handle-the-problem-with-reinjection-that-arises-from-path-dependence-after-the-first-adjustment"&gt;Q9. How does the model handle the &amp;ldquo;problem with reinjection&amp;rdquo; that arises from path dependence after the first adjustment?&lt;/h3&gt;
&lt;p&gt;Without irreversibility, a firm&amp;rsquo;s initial state k̂₀ does not affect behavior after the first adjustment, because there is a unique reset point; subsequent behavior is independent of the aggregate shock magnitude. With irreversibility, firms only partially absorb the aggregate shock at the first adjustment, since the initial state affects the probability of subsequently upsizing or downsizing. In principle, one must track firms through infinitely many adjustments. The paper&amp;rsquo;s resolution (Proposition 2) is to note that the first adjustment erases all heterogeneity except the direction (upsizing vs. downsizing), allowing subsequent behavior to be summarized by just two numbers m(k̂*₋) and m(k̂*₊), combined with the transition probabilities P⁻(k̂₀) and P⁺(k̂₀). This yields a recursive formulation for m(k̂) governed by an HJB equation with two boundary conditions at the reset points, making the problem tractable.&lt;/p&gt;
&lt;h3 id="q10-what-is-the-role-of-the-stationarity-condition-in-pinning-down-the-cir"&gt;Q10. What is the role of the stationarity condition in pinning down the CIR?&lt;/h3&gt;
&lt;p&gt;The HJB for m(k̂) has infinitely many solutions (m(k̂) + a for any constant a). The stationarity condition, requiring that the cross-sectional average of m(k̂) in steady state is zero (no fluctuations without shocks), pins down the unique solution. Economically, it says that average cumulative deviations from complete upsizing spells and complete downsizing spells must exactly balance the deviations from incomplete inaction spells. For upsizing firms, deviations are negative (they hold too little capital relative to average); for downsizing firms, deviations are positive (they hold too much capital). The stationarity condition imposes a linear relationship between m(k̂*₋) and m(k̂*₊) that together with the HJB uniquely determines the solution.&lt;/p&gt;
&lt;h3 id="q11-how-are-the-results-extended-to-assess-nonlinearities-and-robustness"&gt;Q11. How are the results extended to assess nonlinearities and robustness?&lt;/h3&gt;
&lt;p&gt;Appendix G studies nonlinearities numerically in the generalized hazard model for different signs and magnitudes of the aggregate productivity shock. The authors find tiny nonlinearities and asymmetries for productivity shocks below ε = 5%, validating the first-order approximation used throughout. Appendix E.7 provides comparative statics on the output-capital elasticity α. The model is estimated with an inaction threshold of ι = 0.01 (investment rates below 1% in absolute value are treated as inaction), consistent with Cooper and Haltiwanger (2006). The investment distribution is truncated at the 2nd and 98th percentiles to remove outliers.&lt;/p&gt;
&lt;h3 id="q12-what-broader-applicability-do-the-authors-claim-for-the-cir-sufficient-statistics-framework"&gt;Q12. What broader applicability do the authors claim for the CIR sufficient statistics framework?&lt;/h3&gt;
&lt;p&gt;The authors argue the framework applies wherever path-dependent lumpy adjustments occur, including: inventory management (with two types of ordering decisions), durable goods consumption, and labor markets with sticky wages. The key requirement is the existence of a finite number of reset points and sufficient microdata to discipline the transition probabilities across them. Future extensions noted in the paper include: analysis of other aggregate shocks (profitability, capital prices, interest rates); corporate tax reform; monetary policy interacting with investment frictions; time-varying and endogenous price wedges in secondary markets; and higher-order cross-sectional moment responses (variance, skewness of capital-productivity ratios) by choosing different functions f(k̂) for the generalized CIR.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Capital price wedge (ω).&lt;/strong&gt; The fractional discount between the purchase price of capital p and its resale price p(1−ω). In the model this creates two distinct reset points for investment (one for buying at price p, one for selling at the discounted price) and represents the core source of irreversibility. It reflects asset specificity, adverse selection, intermediary fees, and obsolescence. The preferred calibrated value for Chilean manufacturing is ω = 0.12.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cumulative Impulse Response (CIR).&lt;/strong&gt; The integral over all future dates of the impulse response function of the average capital-productivity ratio following a small, permanent, unanticipated aggregate productivity shock. It summarizes both the impact and persistence of aggregate capital fluctuations in a single scalar. Without investment frictions, the CIR is zero (firms adjust instantaneously); the calibrated CIR at ω = 0.12 is 1.93, meaning a 1% aggregate shock generates a 1.93% cumulative deviation.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;Dual reset points (k̂&lt;/em&gt;₋ and k̂&lt;/em&gt;₊).** The two levels to which firms reset their capital-productivity ratio upon adjustment: k̂*₋ after a capital purchase (upsizing) and k̂*₊ after a capital sale (downsizing). With a price wedge, k̂*₊ &amp;gt; k̂*₋, creating an &amp;ldquo;inner inaction region&amp;rdquo; [k̂*₋, k̂*₊] with path-dependent behavior. The inner inaction region width is 0.813 in the Chilean data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sufficient statistics for the CIR.&lt;/strong&gt; Three steady-state cross-sectional moments that together fully characterize the CIR up to first order: (i) Var[k̂]/σ², the scaled cross-sectional variance of capital-productivity ratios (captures insensitivity of incomplete spells to idiosyncratic shocks); (ii) ν·Cov[k̂,a]/σ², the scaled covariance of capital-productivity ratios with capital age (a drift-bias correction); (iii) the &amp;ldquo;irreversibility term&amp;rdquo; measuring how idiosyncratic shocks change the anticipated direction of the next adjustment (unique to the irreversibility case, zero without a price wedge).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Serial correlation in adjustment sign.&lt;/strong&gt; The property, implied by the dual-reset structure, that a firm is more likely to purchase capital following a prior purchase and more likely to sell following a prior sale. In the Chilean data, P⁻⁻ = 0.958 (probability of upsizing after a prior upsize) vs. P⁺⁺ = 0.124 (probability of downsizing after a prior downside), and a logistic regression yields an odds ratio of 3.3.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Generalized hazard function Λ(k̂).&lt;/strong&gt; A state-dependent adjustment probability per unit time, allowing for stochastic and asymmetric fixed costs, that generates the full empirical investment rate distribution. It replaces the single deterministic fixed cost of the baseline model. The hazard function is recovered non-parametrically from microdata by fitting a Gamma distribution to the investment density and inverting the Kolmogorov Forward Equation, conditional on the price wedge.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Renewal weights (r⁻, r⁺).&lt;/strong&gt; Weights used to construct the unconditional density of capital-productivity ratios from the two conditional densities (conditional on prior purchase g⁻(k̂) and prior sale g⁺(k̂)). They rescale adjustment shares by relative average duration, correcting for the observational bias that firms with longer inaction spells are over-represented in the cross-section: r± = (N±/N) × (E±[τ]/E[τ]).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Endogenous irreversibility.&lt;/strong&gt; The component of the inner inaction region width (k̂*₊ − k̂*₋) that arises not from the exogenous price wedge directly but from firms&amp;rsquo; endogenous responses to the wedge — specifically, the differences in expected marginal products and user costs across the two types of inaction spells. At ω = 0.12, 45% of the inner inaction region is attributed to the exogenous wedge and 55% to endogenous amplification.&lt;/p&gt;</description></item><item><title>The Origins and Control of Forest Fires in the Tropics</title><link>https://macropaperwarehouse.com/papers/the-origins-and-control-of-forest-fires-in-the-tropics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-origins-and-control-of-forest-fires-in-the-tropics/</guid><description>&lt;p&gt;This paper studies the economics of illegal tropical forest fires in Indonesia, framed as a modern counterpart to Pigou&amp;rsquo;s canonical externality example of sparks from railway engines. The central research question is whether private firms adjust their fire-setting behavior depending on the degree to which the costs of fire spread fall on themselves versus others, and what enforcement architecture shapes that adjustment.&lt;/p&gt;
&lt;p&gt;The empirical setting is Indonesia&amp;rsquo;s national forest estate, where palm oil and wood fiber concession holders use fire as a cheap land-clearance method — burning primary forest costs 44–70% less than mechanical clearance — despite the practice being illegal. The paper assembles a novel dataset of 107,334 fires across Indonesia&amp;rsquo;s major forested islands from October 2000 to January 2016, constructed from NASA MODIS daily satellite hotspot data (1 km resolution, four flyovers per day). Fire ignitions and spread paths are traced by linking contiguous pixels burning on adjacent days. This fire data is merged with geocoded concession boundaries (logging, palm oil, wood fiber), land-use classifications (protected forest, unleased productive forest, areas outside the forest estate), annual deforestation data from Hansen et al. (2013) at 30 m resolution, daily wind speed data from NOAA NCEP-DOE Reanalysis 2 interpolated to each 1 km pixel, and data on firms investigated by the Indonesian government following the 2015 fires. The main analytical sample focuses on the 39,077 fires started inside wood fiber and palm oil concessions.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s identification strategy exploits two intersecting sources of variation: (1) temporal and spatial variation in monthly wind speed, which predicts the probability and extent of fire spread — a one-standard-deviation increase in wind speed (approximately 5 km/hr) increases fire spread area by 287%; and (2) cross-sectional variation in the land-type composition of the area surrounding each ignition pixel, which determines whether spread costs would fall on the fire-setter or on others. The interaction of these two factors identifies whether firms are more cautious about igniting fires on windy days when surrounding land is their own versus when it belongs to others.&lt;/p&gt;
&lt;p&gt;Three main findings emerge. First, fires are systematically human-caused and linked to industrial land clearance. Fires are eight times more likely per hectare in oil palm and wood fiber concessions than in logging concessions. Completely deforesting a 1 km pixel increases the probability of fire ignition in that pixel in the subsequent year by 279%, and this effect reverses in the year after (two years post-deforestation), ruling out natural flammability as the explanation and confirming a deliberate slash-and-burn cycle. Fire use following deforestation falls by approximately 38% in oil palm concessions during district election years, consistent with tighter enforcement when political incentives favor suppression.&lt;/p&gt;
&lt;p&gt;Second, firms partially internalize the externalities from fire-setting. They are significantly less likely to set fires on windy days when surrounding pixels belong to their own concession rather than to others. A buffer zone entirely owned by the same concession holder reduces ignitions by 8–25% at mean wind speed, and by 22–61% at the 95th-percentile wind speed. However, firms treat neighboring concession land and unleased productive forest similarly — suggesting Coasian bargaining between concession holders is not occurring.&lt;/p&gt;
&lt;p&gt;Third, the government&amp;rsquo;s enforcement pattern shapes firm behavior. Using data on firms investigated after the 2015 fires, the paper shows the government disproportionately investigates firms whose fires burned protected areas or high-population-density land, but not those whose fires damaged other private concessions. The relative weights firms place on different land types when deciding whether to ignite fires align closely with this government punishment function, consistent with firms responding to implicit Pigouvian incentives.&lt;/p&gt;
&lt;p&gt;Counterfactual simulations show that broadening enforcement to treat all land types as the government currently treats populated areas would reduce fires by 80%; treating all land like protected forest would reduce fires by 67%. By contrast, fully Coasian property-rights solutions yield only 14% reductions, and tort reform allowing concession holders to recover damages from neighbors yields only 6%.&lt;/p&gt;
&lt;p&gt;Q: What is the core externality problem studied in this paper?
A: Firms use fire as a cheap land-clearance method, but once set, fires risk spreading beyond the igniter&amp;rsquo;s own concession onto land owned by others, creating an uncompensated externality. The decision to use fire rather than mechanical clearance is de facto a decision to impose this spread risk on third parties. The paper asks whether firms adjust this decision depending on the extent to which spread costs fall on themselves versus others, and whether government enforcement shapes that adjustment.&lt;/p&gt;
&lt;p&gt;Q: Why is Indonesia the empirical setting?
A: Indonesia holds a large share of the world&amp;rsquo;s tropical forests and is among the countries most affected by illegal land-clearing fires. The 2015 Indonesian fires alone released approximately 400 megatons of CO2 equivalent, at their peak emitting more daily greenhouse gases than all US economic activity, and caused an estimated 100,000 excess deaths across Indonesia, Malaysia, and Singapore. The palm oil industry in Indonesia and Malaysia, where fire is used extensively, accounted for 4.7% of global CO2 emissions from 1986 to 2016.&lt;/p&gt;
&lt;p&gt;Q: How are fire ignitions and spread identified in the data?
A: The paper starts from NASA MODIS daily hotspot data at 1 km resolution from October 2000 to January 2016. An iterative procedure assigns contiguous pixels burning on adjacent days to the same fire event, with a 1-pixel buffer allowing for spread detection. This yields 176,855 total fires across Indonesia, of which 107,334 remain after restricting to the major forested islands and the forest estate. The procedure may understate single-day spread since pixels burning on the same day are classified as part of the ignition area rather than spread.&lt;/p&gt;
&lt;p&gt;Q: What fraction of fires spread beyond their ignition area, and how much of the spread falls on outsiders?
A: 87% of fires burn for only one day and 89% do not spread beyond their initial ignition area. However, the largest fire in the data spread to cover 466 times its initial area, and the largest single fire burned 764 km2. Across all multi-day fires started inside concessions, 32% of the total land burned outside the initial ignition area is outside the concession where the fire began, quantifying the scale of the local externality.&lt;/p&gt;
&lt;p&gt;Q: How is wind speed used as an identification strategy?
A: Wind speed provides temporal and spatial variation in the probability that a fire will spread. A one-standard-deviation increase in wind speed (approximately 5 km/hr) increases the extent of fire spread by 287%. Because wind varies month to month and across space, while the composition of surrounding land types is fixed in the cross-section, the interaction of wind speed with surrounding land type identifies whether firms are more cautious about igniting fires when spread risk is high and spread costs would fall on their own land versus others&amp;rsquo; land.&lt;/p&gt;
&lt;p&gt;Q: What is the main result on firms&amp;rsquo; internalization of fire spread externalities?
A: Firms are significantly less likely to start fires on windy days when a larger share of the surrounding buffer zone belongs to their own concession. One additional buffer pixel in one&amp;rsquo;s own land decreases ignitions by 0.2–0.7%. A buffer zone entirely owned by the same concession holder reduces ignitions by 8–25% at mean wind speed, and by 22–61% at the 95th-percentile wind speed. This demonstrates that firms take fire spread risk into account when it threatens their own assets, but discount it when spread would damage others&amp;rsquo; land.&lt;/p&gt;
&lt;p&gt;Q: Do firms treat different types of neighboring land differently?
A: Yes. The benchmark category is unleased productive forest, which has the weakest property rights and receives the least de facto government protection. Relative to this benchmark, firms are more cautious about fire spread toward protected forest (national parks and watershed areas) and toward land outside the forest estate (typically villages and smallholders). One additional buffer pixel in protected forest versus unleased productive forest decreases ignitions by 0.9% at mean wind speed and 2.7% at the 95th-percentile wind speed; the deterrent for land outside the forest estate is even stronger at 1.6% and 4.6%, respectively. Firms treat other firms&amp;rsquo; concession land similarly to unleased productive forest, suggesting no effective private enforcement between concession holders.&lt;/p&gt;
&lt;p&gt;Q: What evidence shows fires are tied to intentional land clearance rather than natural ignition?
A: Fires are eight times more likely per hectare in oil palm and wood fiber concessions than in logging concessions, consistent with clear-cutting versus selective logging. Completely deforesting a 1 km pixel increases fire probability in that pixel in the subsequent year by 279%. Crucially, the effect reverses in the second year after deforestation — the pixel becomes less likely to burn than before — which rules out natural flammability as the mechanism and confirms deliberate slash-and-burn timing.&lt;/p&gt;
&lt;p&gt;Q: What does the electoral cycle evidence show about government enforcement?
A: Fires following deforestation fall by approximately 38% in oil palm concessions during district election years relative to the year prior to an election, and bounce back to pre-election levels in the year after. The decline is confined to productive forest zones where conversion is occurring; no electoral cycle appears in protected areas where conversion is already prohibited. This indicates that enforcement is tightened when political incentives are strong, and confirms that these fires are set intentionally and are responsive to government pressure.&lt;/p&gt;
&lt;p&gt;Q: How is the government&amp;rsquo;s de facto punishment function estimated?
A: The paper uses data on firms investigated by the Indonesian Ministry of Forestry following the 2015 fires, matching investigated firms (identified only by initials in the published list) to concession-holder names. A logistic regression of investigation probability on the land-type outcomes of a firm&amp;rsquo;s fires — conditional on total area burned — shows the government is substantially more likely to investigate firms whose fires burned protected areas or high-population-density land, but does not differentially investigate cases where fire damage is largely confined to other private concessions.&lt;/p&gt;
&lt;p&gt;Q: How closely do firm behavior and government enforcement weights align?
A: The relative weights across land types that the government applies in its investigation decisions correspond closely to the relative weights firms apply when deciding whether to ignite fires on windy days. Firms are most deterred by spread risk toward protected forest and populated areas outside the forest estate — the same categories the government prioritizes. Firms are least deterred by spread toward unleased productive forest and other private concessions — the categories the government largely ignores. This alignment is consistent with firms responding to Pigouvian-style implicit incentives generated by the government&amp;rsquo;s enforcement pattern.&lt;/p&gt;
&lt;p&gt;Q: What do the counterfactuals reveal about policy effectiveness?
A: Fully Coasian property-rights reform — where firms treat all surrounding land as their own — would reduce fires by only 14%. Tort reform enabling concession holders to recover damages from neighbors (treating neighboring concessions as own land) would reduce fires by only 6%. By contrast, uniform enforcement raising deterrence to the level currently applied to populated areas would reduce fires by 80%; applying the level currently applied to protected forest would reduce fires by 67%. An enforcement regime that perfectly prevented all fire spread outside the igniting concession would reduce area burned by only 23%; preventing spread into protected and populated areas alone would yield only a 2% reduction.&lt;/p&gt;
&lt;p&gt;Q: What do the benefit-cost ratios for fires look like?
A: The estimated external damages from the 1997/1998 Indonesian fires range from 1,286 to 6,074 USD per hectare burned (2020 USD). The average private benefit from using fire rather than mechanical clearance — accounting for fertilizers and other costs — averages approximately 52 USD per hectare (2020 USD). Benefit-cost ratios of 0.008 to 0.04 lie well below 1, indicating that the social damages from fires vastly exceed the private benefits, even though the government currently deters only the most costly categories of fire.&lt;/p&gt;
&lt;p&gt;Q: Why do Coasian private solutions perform poorly in this setting?
A: Coasian bargaining between concession holders would require them to reach agreements to bring fire use to a locally efficient level without government intervention. The evidence shows firms treat other concession holders&amp;rsquo; land essentially the same as unprotected unleased productive forest, implying that no such bargains are being struck. The counterfactual analysis confirms this: even a fully-Coasian outcome where every surrounding pixel is treated as own land would reduce fires by only 14%, because the bulk of fires occur when ignition costs to the firm&amp;rsquo;s own land are low regardless of wind speed.&lt;/p&gt;
&lt;p&gt;Q: What is the primary policy implication?
A: The most effective lever for reducing fires is not preventing spread after the fact, but rather deterring ignition in the first place by extending the enforcement regime uniformly across all land types. If firms were induced to treat all surrounding land with the same caution they currently apply toward populated areas — through broader and stronger penalties — fires would fall by 80%. This is substantially more effective than property-rights reforms, tort reforms, or targeted spread-prevention measures focused only on protected and populated areas.&lt;/p&gt;
&lt;p&gt;Externality (fire spread): In this paper&amp;rsquo;s usage, the cost imposed on third parties when a fire ignited inside one concession spreads to land owned by others. The externality is quantified as the share of area burned outside the igniting concession (32% of multi-day fire spread in the data) and the ratio of external damages (1,286–6,074 USD/ha) to private benefits (52 USD/ha) from using fire rather than mechanical clearance.&lt;/p&gt;
&lt;p&gt;Slash-and-burn (industrial scale): The two-stage land-clearance practice where valuable timber is first harvested (deforestation) and the remaining vegetation is then burned to prepare land for plantation crops. The paper establishes this cycle empirically: complete deforestation of a 1 km pixel increases fire ignitions by 279% in the following year, with the effect reversing in the second year, ruling out natural flammability.&lt;/p&gt;
&lt;p&gt;Pigouvian enforcement: Government-imposed penalties that alter private incentives to account for externalities. In this paper&amp;rsquo;s usage, the government&amp;rsquo;s de facto punishment function — which heavily weights fires spreading into protected areas and populated land — functions as an implicit Pigouvian tax, shaping which fires firms choose to avoid rather than uniformly deterring all illegal burning.&lt;/p&gt;
&lt;p&gt;Coasian bargaining failure: The absence of private negotiations between concession holders to internalize the externalities they impose on each other. The paper demonstrates this failure empirically by showing firms treat neighboring concession land no differently from unprotected unleased productive forest, indicating no effective private agreements are limiting cross-concession fire spread.&lt;/p&gt;
&lt;p&gt;Wind speed as spread risk shifter: Monthly average wind speed at each 1 km pixel, used as the time-varying component of fire spread risk. A one-standard-deviation increase (approximately 5 km/hr) increases fire spread area by 287%. The paper uses wind speed variation interacted with surrounding land type composition to identify whether firms adjust ignition decisions based on spread risk and who bears the cost.&lt;/p&gt;
&lt;p&gt;Unleased productive forest (benchmark): Land within the national forest estate that is neither in a designated concession nor in a protected zone, leaving ownership rights unclear and de facto unprotected. The paper uses firms&amp;rsquo; behavior toward this category as the baseline against which sensitivity to other land types is measured, because it attracts the least government attention and the weakest property rights.&lt;/p&gt;
&lt;p&gt;Government punishment function: The implicit weights the Indonesian government places on different types of fire damage when deciding whether to investigate a firm, estimated from logistic regression on the 2015 investigation data. The function heavily weights fires burning protected areas and high-population-density land, and places near-zero weight on damage to other private concessions, shaping which fire types firms strategically avoid.&lt;/p&gt;</description></item><item><title>The Power of Proximity to Coworkers</title><link>https://macropaperwarehouse.com/papers/the-power-of-proximity-to-coworkers/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/the-power-of-proximity-to-coworkers/</guid><description>&lt;p&gt;This paper studies how physical proximity to coworkers affects on-the-job training and productivity, using software engineers at a Fortune 500 online retailer observed from 2019 to 2024. The authors exploit two quasi-experimental shocks to proximity: the office closures of 2020, which eliminated proximity differentials that previously existed across team types, and the firm&amp;rsquo;s subsequent return-to-office (RTO) mandates in 2022 and 2023, which restored proximity for co-located teams while leaving geographically-distributed teams apart. The core identification strategy is a difference-in-differences design comparing engineers whose teams were co-located in a single headquarters building to those whose teams were split across two buildings a ten-minute walk apart — a distinction that became immaterial once offices closed.&lt;/p&gt;
&lt;p&gt;The central finding is that sitting near teammates substantially increases the digital feedback engineers receive on their code. Before the office closures, engineers on co-located teams received 23.9% (1.92 comments per program) more code review feedback than engineers on multi-building teams. Once offices closed, this advantage narrowed by 18.3% (1.47 comments per program, p-value = 0.0026). The lost comments were disproportionately those predicted by a machine-learning classifier to be helpful, actionable, well-reasoned, and impactful, with high-quality comments declining by 21–23% — exceeding the overall volume decline. Face-to-face and digital communication are complements, not substitutes: proximate engineers drew on a wider pool of reviewers and asked 48.4% more follow-up questions, a differential that vanished once offices closed.&lt;/p&gt;
&lt;p&gt;Proximity&amp;rsquo;s effects are highly heterogeneous. Gains in feedback are concentrated among less-tenured, younger, and female engineers — those with the most to learn. Junior engineers on co-located teams lost 2.03 more comments per program upon office closure than junior engineers already on distributed teams (p-value = 0.001); young engineers lost 2.47 more comments (p-value = 0.0001). Female engineers lost 38.9% more comments than their distributed female counterparts (p-value &amp;lt; 0.0001), partly because women stop asking as many people for feedback when they cannot do so in person.&lt;/p&gt;
&lt;p&gt;Proximity improves code quality for inexperienced engineers. Around the second RTO (three days per week), engineers on co-located teams became 2.2 percentage points less likely to add files subsequently deleted — a measure of churn — and 1.4 pp less likely to introduce bugs, relative to distributed teams (p-values of 0.041 and 0.022 respectively). These gains were roughly twice as large for less-tenured and younger engineers. The benefits persist: engineers who spent more pre-closure time on co-located teams continued to write higher-quality code during the fully remote period.&lt;/p&gt;
&lt;p&gt;However, mentorship is costly for those who provide it. Senior engineers on co-located teams wrote 0.76 fewer programs per month in the main codebase before closures (p-value = 0.0005), a gap that closed when offices did and widened again during the second RTO. The firm faces a fundamental tradeoff: proximity accelerates junior engineers&amp;rsquo; human capital development while reducing experienced engineers&amp;rsquo; immediate coding output.&lt;/p&gt;
&lt;p&gt;These dynamics shape hiring. The firm shifted toward hiring older, more experienced engineers during closures — buying talent it could no longer build in-house — and back toward younger hires once offices reopened. Nationally, young college graduates in remotable occupations (classified per Dingel and Neiman, 2020) experienced a 0.88 pp increase in unemployment between 2017–2019 and 2022–2024, while older graduates saw a marginal decline of 0.11 pp. A triple-difference estimate finds a 0.65 pp greater increase in young workers&amp;rsquo; unemployment in remotable versus non-remotable occupations (p-value = 0.029), a pattern that predates generative AI diffusion and is robust to controlling for AI exposure. Back-of-the-envelope, remote work accounts for an estimated 64% of the total unemployment increase among young college graduates over this period.&lt;/p&gt;
&lt;p&gt;The paper also documents that proximity is fragile: a ten-minute walk between two buildings reduces feedback as much as being multiple states away, and even a single distant teammate imposes negative externalities on those who remain co-located, reducing their feedback by 1.71 comments per program (p-value = 0.095) via a &amp;ldquo;one Zoom, all Zoom&amp;rdquo; norm.&lt;/p&gt;
&lt;p&gt;Q: What is the main identification strategy for the office-closure analysis, and what is the key parallel-trends evidence?&lt;/p&gt;
&lt;p&gt;A: The authors compare engineers on co-located teams (all members in one headquarters building) to those on multi-building teams (split across two buildings a ten-minute walk apart), before and after the March 2020 office closures. Co-located teams lost more proximity when offices closed, while multi-building teams experienced a smaller shock, enabling a difference-in-differences design. Pre-closure trends in feedback are parallel across the two team types (Figure I), supporting the identifying assumption. Standard errors are clustered by team, the unit of treatment assignment.&lt;/p&gt;
&lt;p&gt;Q: How large is the effect of proximity on total code review feedback, and how is it broken down by feedback source?&lt;/p&gt;
&lt;p&gt;A: Before closure, co-located engineers received 23.9% (1.92 comments per program) more feedback than multi-building engineers. The DiD estimate indicates that losing proximity reduced feedback by 18.3% (1.47 comments per program, p-value = 0.0026, Column 3 of Table II). This decline stems entirely from reduced feedback from teammates; there is no detectable effect on feedback from engineers on other teams — a placebo check that supports the identification strategy and rules out explanations based on differential project complexity.&lt;/p&gt;
&lt;p&gt;Q: How does proximity affect the quality — not just the quantity — of code review comments?&lt;/p&gt;
&lt;p&gt;A: Using a gradient-boosted decision tree trained on 5,377 human-labeled comments, the authors predict comment quality across all 174,014 comments. Losing proximity reduced comments predicted to be helpful, well-reasoned, actionable, and likely to change the code by 21–23% — exceeding the 18.3% overall volume decline. The residual comments were lower quality: 2.9 pp fewer were helpful (p-value = 0.039), 1.7 pp fewer explained their reasoning (p-value = 0.094), and 1.9 pp fewer were likely to change the code (p-value = 0.072).&lt;/p&gt;
&lt;p&gt;Q: What mechanisms drive the complementarity between face-to-face interaction and digital feedback?&lt;/p&gt;
&lt;p&gt;A: Proximity increases feedback on both the extensive and intensive margins. On the extensive margin, co-located engineers draw on a wider pool of reviewers, returning less frequently to the same commenter. On the intensive margin, losing proximity reduces follow-up questions by 48.4% (0.12 questions per program, p-value = 0.0083), accounting for roughly half of the total feedback decline. The other half comes from reduced initial reviewer feedback. References to other communication channels (e.g., Slack) within code reviews also decline when proximity is lost, confirming that face-to-face and digital communication are complements.&lt;/p&gt;
&lt;p&gt;Q: How small a physical barrier is sufficient to reduce feedback substantially?&lt;/p&gt;
&lt;p&gt;A: A ten-minute walk between two buildings on the same headquarters campus reduces feedback by as much as being multiple states away — both groups receive significantly less feedback than engineers whose entire team sits in the same building (Figure Ib). This finding aligns with research on academics showing that different floors or buildings reduce coauthorship, and extends it to daily teammates sharing projects.&lt;/p&gt;
&lt;p&gt;Q: What are the externality effects of a single distant teammate?&lt;/p&gt;
&lt;p&gt;A: Through the firm&amp;rsquo;s implicit &amp;ldquo;one Zoom, all Zoom&amp;rdquo; norm, even one teammate in a different location shifts all team meetings to video calls. Engineers in the same building exchange 14.5% less feedback when even one teammate is in another building versus when all teammates are co-located (p-value = 0.037). When a new hire transforms a co-located team into a multi-building one, feedback between the original co-located teammates drops by 1.71 comments per program (p-value = 0.095); adding a new co-located hire produces no such decline.&lt;/p&gt;
&lt;p&gt;Q: How does the effect of proximity on feedback differ by engineer tenure, age, and gender?&lt;/p&gt;
&lt;p&gt;A: Less-tenured engineers on co-located teams lost 2.03 more comments per program upon closure than less-tenured engineers on distributed teams (p-value = 0.001). Young engineers (under 29) on co-located teams lost 2.47 more comments per program than young distributed engineers (p-value = 0.0001). Female engineers on co-located teams lost 38.9% (3.71) more comments than female engineers on distributed teams (p-value &amp;lt; 0.0001), partly because women draw feedback from 14.7% fewer people when proximity is lost (p-value = 0.0078), compared to a negligible 2.6% decline for men. The extra feedback women receive in person is of higher quality, not rude or condescending.&lt;/p&gt;
&lt;p&gt;Q: How is the effect of proximity on code quality identified using the RTO design, and what are the magnitudes?&lt;/p&gt;
&lt;p&gt;A: The RTO design compares engineers on co-located (same-city) teams to geographically-distributed teams across three periods: full closure, first RTO (two days per week), and second RTO (three days per week). The authors predict γ_closed ≈ 0 (office assignment irrelevant when closed) and γ_2nd_RTO &amp;gt; γ_1st_RTO (more in-office days means more proximity). Both predictions are confirmed. During the second RTO, co-located engineers were 2.2 pp less likely to add files later deleted (p-value = 0.041) and 1.4 pp less likely to introduce bugs (p-value = 0.022), with effects roughly twice as large for less-tenured and younger engineers.&lt;/p&gt;
&lt;p&gt;Q: Does the benefit of co-location on code quality persist after remote work resumes?&lt;/p&gt;
&lt;p&gt;A: Yes. After all engineers returned to remote work, those who had been on co-located teams pre-closure were 2.37 pp less likely to write disposable code (p-value = 0.013) and 3.09 pp less likely to introduce bugs (p-value = 0.0012). Code quality improves monotonically with the number of pre-closure months spent on co-located teams (Figure A.5). These gaps persist when including current team fixed effects, meaning within the same post-closure team, the previously co-located engineer writes higher-quality code.&lt;/p&gt;
&lt;p&gt;Q: What is the cost of mentorship for senior engineers, and how does it manifest in coding output?&lt;/p&gt;
&lt;p&gt;A: Senior engineers on co-located teams wrote 0.76 fewer programs per month in the main codebase when offices were open (p-value = 0.0005). Once offices closed, this gap disappeared, and senior engineers who lost proximity to their teammates saw a relative increase in output of 0.58 programs per month (p-value = 0.0014). During the second RTO, engineers with more than sixteen months of tenure on co-located teams wrote fewer programs, while no significant difference emerged for less-tenured engineers. Overall, the DiD estimate indicates losing proximity to teammates increases immediate output by 0.48 programs per month (p-value = 0.0002).&lt;/p&gt;
&lt;p&gt;Q: How does the firm&amp;rsquo;s hiring age distribution respond to changes in proximity?&lt;/p&gt;
&lt;p&gt;A: When offices were closed, the firm shifted toward hiring older engineers: the share of hires under age 29 fell from over half pre-closure to less than a third during the closure. After the RTOs, the firm shifted back toward younger hires. Geographic variation reinforces this: headquarters-campus hires were 7–10 years younger than those hired into distributed roles when offices were open; this gap narrowed substantially during closures when everyone was far from teammates.&lt;/p&gt;
&lt;p&gt;Q: Does proximity affect which engineers are poached by other firms?&lt;/p&gt;
&lt;p&gt;A: Yes. During the office closures, 1.2% of co-located engineers were poached per month, compared to 0.9% of multi-building engineers of similar tenure, age, and engineering group (p-value = 0.044). By the end of the closure period, nearly a quarter of co-located engineers had been poached versus a sixth of multi-building engineers. There is a dose response: more pre-closure time on co-located teams predicts higher poaching rates. The effect is concentrated among younger and female engineers, consistent with their feedback building more transferable general human capital. Tenure does not moderate the poaching effect, consistent with less-tenured engineers&amp;rsquo; feedback being more firm-specific.&lt;/p&gt;
&lt;p&gt;Q: What does national unemployment data show about the scarring effects of remote work on young workers?&lt;/p&gt;
&lt;p&gt;A: Between 2017–2019 and 2022–2024, young college graduates (under 29) in remotable occupations experienced a 0.88 pp increase in unemployment (p-value &amp;lt; 0.00001), while older graduates in the same occupations saw a marginal decline of 0.11 pp (p-value = 0.053). A triple-difference regression finds a 0.65 pp greater increase in young workers&amp;rsquo; unemployment in remotable versus non-remotable occupations (p-value = 0.029). Back-of-the-envelope, scaling this estimate by the 61% share of young graduates in remotable jobs predicts a 0.4 pp increase in young college graduates&amp;rsquo; overall unemployment — equal to 64% of the realized 0.63 pp increase.&lt;/p&gt;
&lt;p&gt;Q: Is the unemployment increase among young workers in remotable jobs driven by generative AI rather than remote work?&lt;/p&gt;
&lt;p&gt;A: The authors argue against AI as the primary driver on two grounds. First, the uptick in young workers&amp;rsquo; unemployment in remotable occupations predates the rapid diffusion of generative AI. Second, the differential increase is not concentrated among occupations with the highest AI task exposure. The triple-difference estimate is robust to controlling for occupational AI exposure using the Eisfeldt, Schubert and Zhang (2023) index. The authors acknowledge that AI may become more important as it diffuses further.&lt;/p&gt;
&lt;p&gt;Q: How do young workers&amp;rsquo; own office attendance decisions reflect the value of proximity?&lt;/p&gt;
&lt;p&gt;A: At the partner firm, engineers under 29 were 8.8 pp (37.6%) more likely to come into the office during the RTOs than older engineers when on co-located teams (solid line in Figure VIIa). This difference was roughly halved on geographically-distributed teams (p-value of difference = 0.0085), indicating that the draw is specifically proximity to teammates. Co-located managers raised attendance by 2.6 pp, while co-located teammates raised it by 5.1 pp. Nationally, Stack Overflow survey data show nearly half of engineers under 25 are in the office each day, versus a quarter of older engineers (p-value &amp;lt; 0.00001).&lt;/p&gt;
&lt;p&gt;Q: What does the paper imply about why remote work was rare before the pandemic despite workers&amp;rsquo; stated preferences for it?&lt;/p&gt;
&lt;p&gt;A: The paper offers a resolution: firms may have recognized that the value of the office lies in training for tomorrow and improving the quality — not the quantity — of work today. Remote work boosts immediate output, especially for experienced workers, but it reduces mentorship and long-run skill development. The tradeoff between current and future productivity, and between individual and collective returns to human capital, explains why firms historically resisted remote work even when workers preferred it and short-run output was unaffected.&lt;/p&gt;
&lt;p&gt;Q: What are the implications for gender equity in remote work?&lt;/p&gt;
&lt;p&gt;A: The findings suggest remote work has ambiguous gender effects. While remote work may help working mothers remain in the workforce, it appears costly for young women&amp;rsquo;s professional development, which is especially sensitive to physical proximity. Women receive substantially more high-quality feedback when co-located, draw feedback from a wider network in person, and lose disproportionately more feedback when proximity is lost. Young female engineers on co-located teams were also disproportionately poached — suggesting their human capital gains from co-location are more general and transferable.&lt;/p&gt;
&lt;p&gt;Code review feedback: The digital comments engineers exchange when reviewing each other&amp;rsquo;s code before it is merged into the live codebase; the paper&amp;rsquo;s primary measure of on-the-job training and mentorship investment, distinct from mere volume because the authors also classify comments by helpfulness, reasoning, actionability, and expected impact using supervised machine learning.&lt;/p&gt;
&lt;p&gt;Co-located team: A team in which all members are assigned to the same office building; the treatment group in the difference-in-differences designs, distinguished from multi-building teams (split across two headquarters buildings, a ten-minute walk apart) and geographically-distributed teams (members in different cities or permanently remote).&lt;/p&gt;
&lt;p&gt;One Zoom, all Zoom norm: The implicit team practice of holding all meetings virtually if any single teammate cannot be physically present; the mechanism by which one distant colleague generates negative externalities for the remaining co-located teammates, reducing their in-person interaction and feedback.&lt;/p&gt;
&lt;p&gt;Proximity fragility: The finding that even small physical barriers — a ten-minute walk between buildings — reduce feedback as much as being multiple states away, implying that the relationship between physical distance and mentorship is highly nonlinear near zero.&lt;/p&gt;
&lt;p&gt;Churn (disposable code): Files that are added by an engineer but deleted within the subsequent six months, either because the code was poorly structured or because it introduced a feature later abandoned; used as one of two code quality proxies in the RTO analysis (occurring in 15% of programs).&lt;/p&gt;
&lt;p&gt;Bugs (immediate reversions): Programs that are immediately and fully reverted after being merged, typically indicating the engineer&amp;rsquo;s changes precipitated an emergency requiring rollback to an earlier version; used as the more serious of the two code quality proxies (occurring in 3.5% of programs).&lt;/p&gt;
&lt;p&gt;Scarring effects: The persistent adverse impact on young workers&amp;rsquo; human capital and labor market outcomes from reduced mentorship during the remote work period; manifested both as lower code quality at the individual level and higher unemployment rates nationally among young college graduates in remotable occupations.&lt;/p&gt;
&lt;p&gt;Remotable occupation: An occupation classified by Dingel and Neiman (2020) as feasibly performed from home; used to construct the national triple-difference analysis comparing age gaps in unemployment across remotable and non-remotable jobs before and after the pandemic.&lt;/p&gt;</description></item><item><title>Trust and Innovation Within the Firm</title><link>https://macropaperwarehouse.com/papers/trust-and-innovation-within-the-firm/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/trust-and-innovation-within-the-firm/</guid><description>&lt;p&gt;This paper investigates whether and how a CEO&amp;rsquo;s inherited generalized trust enhances innovation within firms, offering a micro-foundation for the well-documented macro-level relationship between societal trust and economic growth. The author argues that trust — by inducing tolerance of failure — encourages researchers to undertake high-risk, explorative R&amp;amp;D rather than safe exploitation of known approaches.&lt;/p&gt;
&lt;p&gt;The empirical foundation is a matched CEO-firm-patent dataset covering 5,753 CEOs at 3,598 US public firms during 2000–2011, encompassing 700,000 patents and over one million inventors. CEO trust is measured as an inherited trait: each CEO&amp;rsquo;s ethnic origin is inferred probabilistically from their last name using de-anonymized US censuses from 1910–1940, and ethnic-origin-specific trust levels are drawn from the US General Social Survey (GSS), restricted to respondents in highly prestigious occupations. The resulting trust measure is the weighted average of ethnic-specific trust scores across a CEO&amp;rsquo;s likely ethnic composition.&lt;/p&gt;
&lt;p&gt;The main empirical strategy exploits within-firm variation across CEO transitions, using firm and year fixed effects to compare patenting before and after a CEO change. The identifying assumption — that the timing of CEO transitions and the new CEO&amp;rsquo;s trust level are not predicted by prior firm patenting trends — is supported by event-study tests showing flat pre-trends. A one-standard-deviation increase in CEO inherited generalized trust (equivalent to the difference between Greek and English averages) is associated with a 6.2–6.3% increase in patent filings, statistically significant at the 1% level. For the average firm, this equals approximately 1.1 additional patents annually, worth roughly $6.8 million. The effect is larger among exogenous transitions (CEO retirement or death): 8.5% in the restricted sample, and an IV estimate of 8.2%. The back-of-envelope calculation suggests this trust-innovation channel could account for approximately 37% (range: 16–58%) of the effect of trust on GDP per capita growth.&lt;/p&gt;
&lt;p&gt;The paper&amp;rsquo;s central mechanism — risk taking — is tested by examining the distribution of patent quality rather than the mean. Under the risk-taking mechanism, trust should increase the variance of R&amp;amp;D project quality, raising high-quality patents without necessarily increasing low-quality ones. Consistent with this, CEO trust raises only above-median quality patents (measured by forward citation decile), with effects increasing monotonically toward the top decile and no statistically significant effect on below-median patents. Average patent quality as measured by citation-weighted counts or patent value rises by 4–6%. Trust also disproportionately raises the share of explorative patents (those with at least 90% of backward citations outside the firm&amp;rsquo;s existing knowledge stock) by 1 percentage point over a base of 17%.&lt;/p&gt;
&lt;p&gt;The transmission channel is examined using BERT-based classification of nearly one million Glassdoor employee reviews. Under more trusting CEOs, firms exhibit stronger top-down trust sentiment (managers trusting workers), particularly among R&amp;amp;D workers and scientists. The effect materializes within the first two years of a CEO term. Director selection provides an additional transmission mechanism: under more trusting CEOs, newly appointed directors are more trusting and departing directors are less trusting.&lt;/p&gt;
&lt;p&gt;A within-CEO design using bilateral trust (toward researchers in specific countries) with CEO fixed effects addresses omitted CEO characteristics. A one-standard-deviation increase in CEO bilateral trust toward a country is associated with a 5% increase in patents by inventors in that country&amp;rsquo;s R&amp;amp;D lab, controlling for firm-by-year, CEO, and inventor-country fixed effects.&lt;/p&gt;
&lt;p&gt;The effect is strongest when CEO trust is matched to a high-quality researcher pool; in firms with mostly low-quality researchers, high trust may be counterproductive. Trust is also a substitute for R&amp;amp;D knowledge: the effect disappears when the CEO holds a non-MBA graduate degree or has prior R&amp;amp;D experience.&lt;/p&gt;
&lt;p&gt;Q: What is the main research question?
A: The paper asks whether a CEO&amp;rsquo;s generalized trust causes more and higher-quality innovation within the firm, and through what mechanism. It also asks how trust transmits from the CEO to researchers who rarely interact with the CEO directly.&lt;/p&gt;
&lt;p&gt;Q: How is CEO trust measured?
A: CEO trust is measured as an inherited trait using a two-step procedure. First, each CEO&amp;rsquo;s last name is probabilistically mapped to one or more ethnic origins using four de-anonymized US censuses (1910–1940). Second, ethnic-origin-specific trust is computed from GSS respondents in highly prestigious occupations. The CEO&amp;rsquo;s trust measure is the weighted average across ethnic compositions. This measure is shown to be more precise than an individual-level survey measure and approximately 80% as precise as a game-based measure, without introducing attenuation bias.&lt;/p&gt;
&lt;p&gt;Q: What is the baseline patent effect and how large is it economically?
A: A one-standard-deviation increase in CEO inherited trust is associated with a 6.2–6.3% increase in patent filings (statistically significant at 1%). For the average baseline firm, this is approximately 1.1 additional patents per year, valued at roughly $6.8 million. When patent quality is accounted for, the effect rises to 9.9% using citation-weighted patent count and 11.5% using patent value based on excess stock returns on grant dates.&lt;/p&gt;
&lt;p&gt;Q: Is the effect causal? What identification strategy is used?
A: The main strategy uses firm and year fixed effects, identifying the effect from within-firm variation around CEO transitions. Pre-trend tests confirm that neither the timing of CEO changes nor the new CEO&amp;rsquo;s trust level predicts prior firm patenting. Among exogenous transitions (CEO retirements and deaths), the effect is 8.5%, and an IV estimate using the predecessor&amp;rsquo;s trust as instrument yields 8.2% (significant at 10%), both comparable to the baseline.&lt;/p&gt;
&lt;p&gt;Q: What is the macroeconomic significance of the trust-innovation channel?
A: Combining the paper&amp;rsquo;s trust-to-patents estimate (0.042–0.062) with Akcigit et al.&amp;rsquo;s (2017) patents-to-GDP-growth estimate (0.026–0.066) and the cross-country trust-to-growth coefficient (0.007), the trust-innovation channel could explain approximately 37% of the effect of trust on growth, with a plausible range of 16–58%.&lt;/p&gt;
&lt;p&gt;Q: What is the mechanism linking CEO trust to innovation?
A: The conceptual mechanism is that a more trusting manager interprets researcher failure as bad luck rather than bad type, making her more likely to tolerate failure and continue employing the researcher. This increases the researcher&amp;rsquo;s incentive to pursue explorative, high-risk R&amp;amp;D over safe exploitation of known approaches. The mechanism implies a variance-increasing effect on the R&amp;amp;D quality distribution, rather than a mean shift.&lt;/p&gt;
&lt;p&gt;Q: How is the risk-taking mechanism tested against alternative mechanisms?
A: The paper examines the distribution of patent quality by citation decile. Under mean-shifting alternatives (delegation, cooperation, relational contracting), trust should raise all quality brackets. Under risk-taking, trust raises only high-quality patents. The results show CEO trust has monotonically increasing effects from low to high quality deciles, with no statistically significant effect on below-median patents, consistent only with the variance-increasing (risk-taking) mechanism.&lt;/p&gt;
&lt;p&gt;Q: What patent quality measures are used and what do they show?
A: Beyond forward citation deciles, the paper uses explorativeness (patents with at least 90% of backward citations outside the firm&amp;rsquo;s existing knowledge stock), disruptiveness (Funk and Owen-Smith, 2017), patent importance (Kelly et al., 2021), backward citations to scientific literature, and patent scope. Trust increases all these measures with statistically significant positive coefficients. The share of explorative patents rises by 1 percentage point over a base of 17%. Average citation count and patent value increase by 4–6%.&lt;/p&gt;
&lt;p&gt;Q: Does CEO trust raise R&amp;amp;D expenditure?
A: No. The coefficients from regressing R&amp;amp;D expenditure on CEO trust are neither statistically significant nor large enough to explain the innovation effect. The patent effect is also robust to controlling for R&amp;amp;D inputs, suggesting that trust affects the type of projects chosen (consistent with risk-taking) or their realized outcomes, rather than the scale of R&amp;amp;D.&lt;/p&gt;
&lt;p&gt;Q: How does CEO trust transmit to corporate culture?
A: Using BERT-based classification of nearly one million Glassdoor reviews covering 266 firms and 397 CEO terms between 2008 and 2017, the paper finds that CEO trust is associated with stronger top-down trust sentiment (managers trusting workers). The normalized effect of a one-standard-deviation increase in CEO trust on overall trust sentiment is 0.257, on top-down trust 0.531, and on bottom-up trust only 0.141 (statistically insignificant). The effect is strongest among reviewers who identify as scientists, researchers, or engineers, and materializes within the first two years of the CEO term.&lt;/p&gt;
&lt;p&gt;Q: What evidence exists for transmission via director selection?
A: Under more trusting CEOs, newly appointed directors — especially those who remain until the end of the CEO term — are more trusting, and departing directors are less trusting. The average director trust improves during the CEO&amp;rsquo;s term. Because 54% of director hirings and 46% of turnovers occur within the first two years, this change also materializes quickly, consistent with the dynamic pattern of trust culture change.&lt;/p&gt;
&lt;p&gt;Q: What is the within-CEO bilateral trust result and what does it add?
A: Using within-CEO variation in bilateral trust toward researchers from different countries (from Eurobarometer surveys), and controlling for CEO, inventor-country, and firm-by-year fixed effects, a one-standard-deviation increase in CEO bilateral trust toward a country is associated with a 5% increase in patents by inventors in that country&amp;rsquo;s R&amp;amp;D lab. This design allows CEO fixed effects, ruling out unobserved CEO-level confounders such as management style or R&amp;amp;D ability.&lt;/p&gt;
&lt;p&gt;Q: When is CEO trust counterproductive?
A: CEO trust is beneficial only when matched to a high-quality researcher environment. Using residual patent output (controlling for observable firm and CEO characteristics) as a proxy for researcher quality, the effect of CEO trust on patents, patent output per R&amp;amp;D dollar, and future sales/employment/TFP is significant only among firms in the top two quintiles of researcher quality. In firms with mostly low-quality researchers, high CEO trust may be counterproductive by failing to screen out bad researchers.&lt;/p&gt;
&lt;p&gt;Q: How does the trust effect vary by industry and CEO background?
A: The effect is ubiquitous across industries but especially pronounced in pharmaceutical and ICT firms. The timing varies: it manifests quickly in ICT (short R&amp;amp;D lag) and more slowly in pharma (long R&amp;amp;D horizon). The effect vanishes when the CEO holds a non-MBA graduate degree or has prior R&amp;amp;D experience, suggesting trust is a substitute for direct knowledge of R&amp;amp;D processes.&lt;/p&gt;
&lt;p&gt;Q: Are the results robust?
A: Yes. The paper reports 14 categories of robustness checks including alternative patent transformations, alternative trust measures (LASSO, World Value Survey, Global Preference Survey, alternative GSS questions), alternative standard error clustering, Poisson count models, restriction to granted patents, exogenous transition subsamples, modern difference-in-differences estimators (de Chaisemartin et al., 2024; Sun and Abraham, 2021; Callaway and Sant&amp;rsquo;Anna, 2021; Borusyak et al., 2024), and leave-one-ethnicity-out. The baseline result is stable across all these checks.&lt;/p&gt;
&lt;p&gt;Inherited generalized trust: The paper&amp;rsquo;s measure of a CEO&amp;rsquo;s trust disposition, defined as the probability-weighted average of ethnic-origin-specific trust levels (from the GSS) based on the CEO&amp;rsquo;s likely ethnic composition inferred from their last name and historical census records. It captures the culturally transmitted component of trust, distinct from individual-level noise.&lt;/p&gt;
&lt;p&gt;Explorative R&amp;amp;D: In the paper&amp;rsquo;s framework (building on March, 1991), research activities that involve testing untested paths, carrying high risk of failure but high potential for innovation, as opposed to exploitation of well-known approaches with low failure risk. The paper argues CEO trust encourages researchers to shift toward exploration.&lt;/p&gt;
&lt;p&gt;Tolerance of failure: A manager&amp;rsquo;s propensity to attribute a researcher&amp;rsquo;s failure to bad luck rather than bad type. Under the paper&amp;rsquo;s mechanism, a more trusting manager gives greater weight to bad luck, making her more likely to retain the researcher after failure, thereby incentivizing risk taking.&lt;/p&gt;
&lt;p&gt;Top-down trust: In the paper&amp;rsquo;s BERT-based classification of Glassdoor reviews, the direction of trust from managers toward workers (as opposed to bottom-up trust from workers toward managers). The paper finds CEO trust primarily raises top-down trust sentiment, especially among R&amp;amp;D workers.&lt;/p&gt;
&lt;p&gt;Patent explorativeness: A patent quality measure defined as the share of its backward citations that fall outside the firm&amp;rsquo;s existing knowledge stock; patents are classified as explorative if at least 90% of backward citations are outside that stock. The paper uses this as a direct measure of explorative R&amp;amp;D output.&lt;/p&gt;
&lt;p&gt;Bilateral trust: CEO d&amp;rsquo;s directed trust toward individuals from country c, computed analogously to inherited generalized trust but using Eurobarometer survey data on country-pair trust attitudes among European-origin populations. Used in the within-CEO design to control for CEO fixed effects.&lt;/p&gt;
&lt;p&gt;Variance-increasing mechanism: The paper&amp;rsquo;s characterization of the risk-taking channel, in which CEO trust raises the variance (not the mean) of the R&amp;amp;D project quality distribution by encouraging researchers to pursue high-risk, high-reward exploration. Empirically identified by the pattern that trust raises only above-median quality patents with monotonically increasing effects toward the top decile.&lt;/p&gt;</description></item><item><title>Who's Afraid of the Minimum Wage? Measuring the Impacts on Independent Businesses Using Matched U.S. Tax Returns</title><link>https://macropaperwarehouse.com/papers/whos-afraid-of-the-minimum-wage-measuring-the-impacts-on-independent-businesses-using-matched-u.s.-tax-returns/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/whos-afraid-of-the-minimum-wage-measuring-the-impacts-on-independent-businesses-using-matched-u.s.-tax-returns/</guid><description>&lt;h2 id="layer-1--overview"&gt;Layer 1 — Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Research Question&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;This paper asks how independent (pass-through) businesses in the United States accommodate minimum wage increases — specifically whether they reduce employment, compress profits, pass costs through to customers, or exit — and what happens to the low-earning workers and business owners affected by these adjustments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data and Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors construct a novel linked firm-worker-owner panel dataset from the universe of U.S. tax returns, covering approximately 235,000 pass-through firms (S-corporations, partnerships, and LLCs) per year in highly exposed industries over 2010–2019. &amp;ldquo;Highly exposed&amp;rdquo; industries are defined as those where at least 15% of workers earned below the full-time equivalent of the federal minimum wage ($15,080 per year) in 2013. The dataset links annual business income tax returns to the individual income tax returns and W-2 information reports of all workers and owners.&lt;/p&gt;
&lt;p&gt;The causal identification strategy exploits the six state minimum wage increases that took effect in 2014 (California, Connecticut, Delaware, Michigan, Minnesota, and New Jersey) relative to 24 states that did not change their wage floors at any point from 2012–2018. The empirical workhorse is a panel difference-in-differences event study (Equation 1), augmented by DFL re-weighting (DiNardo et al., 1996) to improve comparability of treatment and control firms on observables. The analysis covers cumulative effects through 2018, by which point the average minimum wage across treatment states had risen 30.6%.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Main Findings&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Employment:&lt;/strong&gt; The average exposed independent firm does not meaningfully reduce employment. The authors estimate an own-wage elasticity of -0.209 (s.e. = 0.0112). Employment adjustments manifest as moderately lower hiring rather than layoffs of existing workers. Reduced hiring is wholly concentrated among teenagers and very part-time jobs paying less than $3,900 annually (with 67% earning less than $1,000 per year).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Worker earnings:&lt;/strong&gt; Despite the hiring reduction, low-earning workers employed at exposed independent firms experience average earnings gains of approximately $2,000 per year by 2018, relative to comparable workers in untreated states. Young individuals aged 20–26 without a 2013 job earn roughly $4,000 more per year by 2018; teenagers without a 2013 job gain approximately $1,000 per year. Workers in these groups are no less likely — and in some cases slightly more likely — to be employed five years after the minimum wage increase.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Wage bills:&lt;/strong&gt; Average wage bills among surviving treated firms rose 7.03% (s.e. = 0.0153) by 2018. Earnings gains are concentrated among workers earning $15,600–$35,000 annually, with no evidence of reduced earnings for higher-paid workers. The 7% average wage bill increase amounts to only 1.4% of 2013 firm revenues, easing pass-through.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Revenue and profits:&lt;/strong&gt; Revenues of surviving treated firms grew approximately 2.1% more than control firms by 2018. On average, this revenue increase fully offsets the higher wage bill, yielding a small net profit increase of roughly $3,360 (s.e. = $1,123) per owner by 2018, or about 2.7% of mean 2013 owner income.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Firm exit:&lt;/strong&gt; On average across all highly exposed industries, minimum wages increased the five-year exit probability by 0.9 percentage points (s.e. = 0.0029), relative to a baseline raw exit rate of approximately 29%. Exit effects are driven entirely by restaurants: by 2018, restaurants in treated states were 1.85 percentage points (s.e. = 0.0039) more likely to have exited, while the exit response for non-restaurant exposed firms is a precisely estimated zero.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity by productivity within restaurants:&lt;/strong&gt; Exit is concentrated entirely in the bottom productivity quartile (coefficient = 0.0254, s.e. = 0.0079), with no significant effect in the upper three quartiles. Profits among surviving small restaurants rise by $5,941 (s.e. = $1,546) by 2018 relative to 2013. Among small restaurants, the profit gains are larger for firms in the higher productivity quartiles (Q3: +$7,915; Q4: +$9,161). Surviving restaurants also increase non-labor input spending by 2.53% (s.e. = 0.0101), consistent with expanded output following competitor exits.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Entrant characteristics:&lt;/strong&gt; Post-reform restaurant entrants in treatment states have higher wage bills (13.8% higher in logs), higher revenues (4.0% higher), higher value-added (8.4% higher), and higher productivity (net income/revenue ratio 2.24 percentage points higher) than entrants in control states, indicating the minimum wage raises the productivity floor for new entrants.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Owner outcomes after exit:&lt;/strong&gt; Owners of small restaurants forced out by the minimum wage are significantly less likely to own an independent business five years later, but earn no less on average in wages plus business income. Policy-induced exiters are significantly less likely to report negative incomes, suggesting substitution away from risky or marginally profitable business ownership.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Theoretical Framework&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The authors present a Cournot competition model with heterogeneous firm productivity and fixed production costs. A minimum wage cost shock raises marginal costs, narrowing margins for all firms. Firms whose cost increases exceed the market price increase cannot cover fixed costs and exit. Remaining firms gain higher markups and larger market shares as demand is reallocated from exiting firms. Selection on ex-ante productivity (the least productive firms exit) limits the distortion to market quantity and amplifies profit gains among productive survivors. The model predicts profit increases only in markets with firm exit, which matches the data: profits rise among restaurants (where exit occurs) but not among retailers (where exit is a precisely estimated zero).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope Conditions&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Findings pertain to the short-to-medium run (up to five years post-legislation) of phased-in minimum wage increases averaging 30.6% in six U.S. states. The sample covers pass-through (independent) businesses in highly exposed industries. Longer-run effects may differ if entrants adopt production technologies that rely less on low-wage labor or incumbents reconfigure inputs. Border-county retailers appear to be less able to pass through costs than interior firms, suggesting product market competition is a key moderating factor.&lt;/p&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-why-do-the-authors-focus-on-pass-through-businesses-rather-than-publicly-traded-corporations"&gt;Q1. Why do the authors focus on pass-through businesses rather than publicly traded corporations?&lt;/h3&gt;
&lt;p&gt;Pass-throughs (S-corporations, partnerships, and LLCs) comprise 78% of non-sole-proprietorship businesses and 79% of firms with fewer than 20 employees. They represent the majority organizational form for independent businesses in virtually all two-digit NAICS industry groups except utilities and enterprise management. Because minimum wage concerns are disproportionately raised on behalf of small independent businesses, and because most minimum wage workers in restaurants are employed at pass-throughs, studying pass-throughs directly addresses the policy debate. Additionally, pass-through tax returns link business income directly to the individual tax returns of each owner, enabling the authors to separately identify employee versus owner responses.&lt;/p&gt;
&lt;h3 id="q2-how-do-the-authors-define-highly-exposed-industries-and-why-does-this-matter-for-identification"&gt;Q2. How do the authors define &amp;ldquo;highly exposed&amp;rdquo; industries and why does this matter for identification?&lt;/h3&gt;
&lt;p&gt;Highly exposed industries are defined as four-digit NAICS industries where at least 15% of workers earned below the full-time federal minimum wage equivalent ($15,080 per year) in 2013, using tax data to construct a proxy for minimum wage workers. The analysis focuses on these industries because minimum wage workers are extremely concentrated — the vast majority are in Leisure/Hospitality and Retail. Restricting to highly exposed industries allows the authors to estimate average effects within affected markets and conduct heterogeneity analysis across firm characteristics within those markets, including comparing firms with different baseline shares of low-earning workers that nonetheless all face the market-level cost shock.&lt;/p&gt;
&lt;h3 id="q3-how-do-the-employment-effects-decompose-into-hiring-versus-retention"&gt;Q3. How do the employment effects decompose into hiring versus retention?&lt;/h3&gt;
&lt;p&gt;The average firm subject to a higher wage floor does not lay off existing workers (the retention line is flat in event study estimates). By 2018, firms in treated states hire roughly one fewer worker on average than similar firms in control states, entirely through reduced hiring. This reduced hiring is wholly concentrated among teenagers in very part-time jobs: the missing hires consist entirely of workers who would have earned less than $3,900 annually, with 67% earning less than $1,000 per year. Simultaneously, workers already employed at exposed firms are 2 to 4 percentage points more likely to remain with their 2013 employer by 2016, with prime-age low-earning workers exhibiting the largest retention increases.&lt;/p&gt;
&lt;h3 id="q4-what-happens-to-low-earning-workers-and-young-people-in-individual-level-panels"&gt;Q4. What happens to low-earning workers and young people in individual-level panels?&lt;/h3&gt;
&lt;p&gt;Low-earners (those earning below $25,000 in each year from 2012–2014) at exposed independent firms experience average earnings gains of approximately $2,000 per year by 2018 relative to similar workers in untreated states, including teenage low-earners. Young individuals aged 20–26 with no job in 2013 experience a relative earnings increase of approximately $4,000 per year by 2018; teenagers without jobs in 2013 gain approximately $1,000 per year. These workers are no less likely — and often slightly more likely — to be employed relative to their counterparts in control states, so the earnings gains are not offset by employment losses at the individual level.&lt;/p&gt;
&lt;h3 id="q5-what-is-the-magnitude-of-the-cost-shock-for-firms-and-how-does-it-compare-to-revenues"&gt;Q5. What is the magnitude of the cost shock for firms and how does it compare to revenues?&lt;/h3&gt;
&lt;p&gt;By 2018, the average wage bill among surviving firms in treated states was 7.03% (s.e. = 0.0153) higher than comparable firms in control states. This is consistent with a back-of-envelope calculation: low-earning workers account for about 21% of wage bills at these firms, and states raised minimum wages by 30.6% on average (0.21 × 0.306 = 0.064). However, the 7% wage bill increase amounts to only approximately 1.4% of 2013 firm revenues, making cost pass-through relatively modest. Higher minimum wages have no discernible impact on pension contributions but slightly reduce deductions for other benefits including health insurance.&lt;/p&gt;
&lt;h3 id="q6-how-do-surviving-firms-finance-the-increased-wage-bill-and-what-happens-to-profits"&gt;Q6. How do surviving firms finance the increased wage bill, and what happens to profits?&lt;/h3&gt;
&lt;p&gt;Surviving firms finance the wage increase primarily through higher revenues. By 2018, revenues of firms in treated states grew approximately 2.1% more than revenues of firms in control states. On average, this revenue increase outpaces the higher wage bill, resulting in a net profit increase of approximately $3,360 (s.e. = $1,123) per owner by 2018, representing about 2.7% of mean 2013 owner income. There is no evidence of redistribution from middle- or high-income workers within firms; wage bill increases are concentrated among workers earning $15,600–$35,000 annually, consistent with minimum wage spillovers to workers slightly above the statutory floor.&lt;/p&gt;
&lt;h3 id="q7-why-do-restaurants-experience-exit-effects-but-retailers-do-not"&gt;Q7. Why do restaurants experience exit effects but retailers do not?&lt;/h3&gt;
&lt;p&gt;The asymmetry stems from the intensity of low-wage labor in production. While low-earning workers account for a similar share of labor costs at restaurants (41.8%) and retailers (38.5%), labor costs overall are more than twice as large at restaurants relative to retailers. Wage bills account for 39% of variable costs and 27% of revenues at restaurants, but only 16% of variable costs and 13% of revenues at retailers. As a result, raising the minimum wage raises variable costs by 5.76% at restaurants. Non-restaurant exposed firms are able to fully pass through their smaller cost shock, yielding flat profits and neither employment nor exit impacts.&lt;/p&gt;
&lt;h3 id="q8-why-is-firm-exit-concentrated-in-the-lowest-productivity-quartile-of-restaurants-rather-than-among-the-most-exposed-firms"&gt;Q8. Why is firm exit concentrated in the lowest productivity quartile of restaurants rather than among the most exposed firms?&lt;/h3&gt;
&lt;p&gt;The Cournot framework predicts exits among firms with the lowest ex-ante productivity (highest marginal costs), the largest cost shock (highest share of low-wage labor per unit of output), or a combination. Empirically, productivity is the primary determinant: restaurants across all productivity quartiles use similar shares of low-earning workers (40–44% of wage bills for Q1 through Q4). Exit rises significantly only among restaurants in the bottom productivity quartile (coefficient = 0.0254, s.e. = 0.0079), with no significant effects in Q2–Q4. Among the lowest-productivity restaurants, those most dependent on low-earning labor face the largest exit rates.&lt;/p&gt;
&lt;h3 id="q9-how-do-the-models-predictions-about-profit-heterogeneity-match-the-data"&gt;Q9. How do the model&amp;rsquo;s predictions about profit heterogeneity match the data?&lt;/h3&gt;
&lt;p&gt;The Cournot model predicts profits should rise only in markets with firm exit (via increased margins and market share reallocation to survivors). This is exactly what the data show. Among restaurants, where exit is concentrated in the bottom productivity quartile, profits among surviving small restaurants rise by $5,941 (s.e. = $1,546) by 2018. Among small restaurants specifically, profit gains increase with productivity: Q3 restaurants gain $7,915 (s.e. = $3,326) and Q4 restaurants gain $9,161 (s.e. = $2,127), while Q1 and Q2 gains are statistically indistinguishable from zero. In non-restaurant exposed industries where the exit effect is a precise zero, profits are also flat — exactly as the model predicts.&lt;/p&gt;
&lt;h3 id="q10-what-happens-to-the-characteristics-of-new-restaurant-entrants-after-the-minimum-wage-increase"&gt;Q10. What happens to the characteristics of new restaurant entrants after the minimum wage increase?&lt;/h3&gt;
&lt;p&gt;Post-reform restaurant entrants in treatment states are systematically more productive than entrants in control states. They have wage bills 13.8% higher (in logs), revenues 4.0% higher, value-added 8.4% higher, and productivity ratios (net income/revenue) 2.24 percentage points higher than new entrants in control markets. This implies the minimum wage raises the minimum viable productivity threshold for entrant restaurants, consistent with Sorkin (2015)&amp;rsquo;s insight that minimum wages shape the capital and technology choices of entering firms. The restaurant industry thus becomes more productive on average through both the exit of the least productive incumbents and the entry of more productive new firms.&lt;/p&gt;
&lt;h3 id="q11-how-do-worker-transition-patterns-reflect-the-reallocation-of-output-to-surviving-firms"&gt;Q11. How do worker transition patterns reflect the reallocation of output to surviving firms?&lt;/h3&gt;
&lt;p&gt;Workers at large independent businesses (top revenue quartile) are 3.52 percentage points more likely to remain with their 2013 employer in 2018 and 2.36 percentage points less likely to switch to another large firm. The large firms that retain more of their existing workforce also reduce their hiring of very part-time teenagers the most — in the top revenue quartile, firms shed roughly 4.5 employment relationships on average, comprising higher retention of 4.15 existing workers offset by reduced hiring of 8.67 very part-time teenage workers. Workers originally at smaller exposed firms are more likely to be found working at larger firms five years out, consistent with demand reallocation from exiting and shrinking small firms toward larger, more productive survivors.&lt;/p&gt;
&lt;h3 id="q12-what-happens-to-owners-of-restaurants-that-exit-due-to-the-minimum-wage"&gt;Q12. What happens to owners of restaurants that exit due to the minimum wage?&lt;/h3&gt;
&lt;p&gt;Policy-induced exiters of small restaurants are significantly less likely to own an independent business five years later and less likely to receive all earnings from business ownership, relative to owners of restaurants that exited for other reasons in control states. However, their average incomes (wage income plus ordinary business income) are no lower. This income stability is partly explained by the fact that policy-induced exiters are significantly less likely to report negative incomes five years out, suggesting they substitute away from potentially risky or marginally profitable business ownership toward wage employment or other activities. The utility implications are ambiguous: these former owners may have preferred business ownership even if it did not yield higher income.&lt;/p&gt;
&lt;h3 id="q13-what-is-the-role-of-product-market-competition-in-mediating-pass-through-as-evidenced-by-border-county-analysis"&gt;Q13. What is the role of product market competition in mediating pass-through, as evidenced by border-county analysis?&lt;/h3&gt;
&lt;p&gt;The border county robustness analysis reveals that product market competition is central to pass-through success. Retailers near state borders, where consumers can cross-state-border shop, face more elastic demand and are less able to finance the wage cost shock with new revenues, exhibiting reduced profits and higher exit rates (though estimates are imprecise). Further from the border, where the cost shock is more commonly felt by all potential substitutes (making market demand elasticity rather than firm demand elasticity the relevant parameter), results are very similar to the full-sample aggregate findings. This confirms that the common nature of the minimum wage cost shock — shared by all competing firms in the market — is a key reason firms can pass through costs to consumers.&lt;/p&gt;
&lt;h3 id="q14-how-do-the-findings-address-the-divide-among-independent-business-owners-on-minimum-wage-policy"&gt;Q14. How do the findings address the divide among independent business owners on minimum wage policy?&lt;/h3&gt;
&lt;p&gt;The heterogeneous outcomes rationalize why surveys consistently find business owners divided. Among restaurants, some owners (those operating the least productive small restaurants) face exit and loss of business ownership, while surviving productive restaurateurs see higher profits of $5,941–$9,161 per year. Among non-restaurant exposed businesses, owners are broadly unaffected in terms of profits and viability. Uncertainty about whether a given firm&amp;rsquo;s demand is elastic enough to bear cost pass-through — given that owners may be more familiar with the elasticity of firm-level demand from prior unilateral price changes, rather than the relevant market-level demand elasticity applying to a common cost shock — may broaden opposition to include even owners who would ultimately benefit.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Pass-through businesses (independent businesses):&lt;/strong&gt; Privately owned firms organized as S-corporations, partnerships, or LLCs, taxed by passing income through to the individual returns of owners rather than at the entity level. In 2015, these comprised 78% of non-sole-proprietorship U.S. businesses and 46% of employment. The paper uses &amp;ldquo;pass-through&amp;rdquo; and &amp;ldquo;independent business&amp;rdquo; interchangeably as the unit of analysis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Highly exposed industries:&lt;/strong&gt; Four-digit NAICS industries where at least 15% of workers earned below the annual full-time equivalent of the federal minimum wage ($15,080) in 2013, as measured in the authors&amp;rsquo; administrative tax data. This threshold proxies the concentration of minimum-wage workers across industries and drives the sample selection for firm-level analysis.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Own-wage elasticity of employment:&lt;/strong&gt; The estimated percentage change in employment at a firm associated with a given percentage change in the firm&amp;rsquo;s minimum wage. The authors estimate this as -0.209 (s.e. = 0.0112), reflecting the average effect across all exposed independent businesses, conditional on the firm&amp;rsquo;s industry, size, and local market characteristics.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DFL re-weighting (DiNardo-Fortin-Lemieux):&lt;/strong&gt; A non-parametric reweighting procedure that adjusts the distribution of control-group firms to match the distribution of treatment-group firms on observables (specifically, two-year lagged value-added within three-digit NAICS industries). Used to improve pre-reform comparability of treatment and control firm samples without parametric functional form assumptions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Firm productivity (in this paper&amp;rsquo;s sense):&lt;/strong&gt; Measured as the ratio of net profits to revenues (net income/revenue) at the firm level in the base year 2013, used to assign firms to productivity quartiles for heterogeneity analysis. This is a firm-level profitability measure constructed from pass-through tax returns, not a total factor productivity estimate requiring production function estimation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Firm exit:&lt;/strong&gt; An indicator for a firm that filed a tax return in 2013 but did not file a return in a subsequent year t. The average one-year exit rate for highly exposed independent businesses is 5.2%; the cumulative five-year raw exit rate is approximately 29% across treatment and control states.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cournot competition with heterogeneous productivity and fixed costs:&lt;/strong&gt; The paper&amp;rsquo;s conceptual framework, in which N firms compete in quantities with asymmetric marginal costs (reflecting heterogeneous productivity), a common output price, and a fixed cost of production. Under this framework, a minimum wage cost shock narrows margins unevenly, induces exit among firms that cannot cover fixed costs, and generates both demand reallocation and market share gains for productive survivors — rationalizing simultaneous exit and profit increases in the same industry.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Common cost shock:&lt;/strong&gt; The property that a minimum wage increase raises production costs for all firms employing low-wage workers in the same market simultaneously. Because all competing firms face higher costs, the relevant pass-through parameter is the elasticity of market demand rather than the (higher) elasticity of individual firm demand, facilitating cost pass-through to consumers and distinguishing minimum wages from unilateral price changes by a single firm.&lt;/p&gt;</description></item></channel></rss>