Macro Paper Warehouse
Online First [Journal of Political Economy] doi:10.1086/743901 Online 31 Aug 2026

Using Consumption Data to Derive Optimal Top Income and Capital Tax Rates

Christian Hellwig

Nicolas Werquin

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

How high should taxes on top incomes and on savings be? The usual answer is built from income data. This paper shows that two famous results - the formula for the top income tax rate, and the case for not taxing savings at all - are two sides of one coin, and that using consumption data instead changes the answer substantially. Consumption is far less concentrated at the top than income is, and putting the consumption figure through the same formula cuts the implied top income tax rate from 74% to 57%. The gap between the two numbers is itself informative about how savings should be taxed.

What this paper finds — and why it matters

Two of the most influential results in optimal tax theory — Saez’s (2001) formula for the top income tax rate and Atkinson and Stiglitz’s (1976) case for not taxing savings — are usually invoked separately, and this paper shows they are two sides of one coin: the revenue-maximising top income tax equals Saez’s τ if and only if uniform commodity taxation applies and the optimal savings tax is zero, so any departure from that joint benchmark forces a trade-off between the two instruments. The reason the split matters empirically is that consumption is much less concentrated than income at the top: the authors’ own estimates from the 2005–2021 waves of the PSID put the Pareto coefficient of income at 1.8 but of consumption at about 3.1, and plugging the consumption tail into the same formula would cut the implied top income tax from 74% to 57%. Working in a Mirrleesian model where agents work, consume and save for retirement, the paper derives representations of optimal top income and savings taxes in terms of three consumption-based sufficient statistics — the ratio of Pareto coefficients, the ratio of compensated elasticities, and a scaling parameter identified by the elasticity of intertemporal substitution — and shows that consumption data, not savings data, are what identify how the combined wedge splits. Calibrating to available empirical estimates, the paper finds it optimal across all its calibrations to shift part of the top earners’ burden from income onto savings, with an optimal savings tax between 12.5% and 30% and a top marginal income tax falling from 74.1% to between 63% and 70.4%; the paper is explicit that the size of the shift is sensitive to the calibration strategy and that identifying the pass-through elasticity separately from the ratio of Pareto tails is critical, since varying the two across their plausible ranges moves the optimal savings tax from 4.5% to 26.3%. It also draws a pointed implication for policy: because the two benchmarks stand or fall together, Diamond and Saez’s (2011) simultaneous case for high top income taxes and for positive capital taxes is, on the paper’s argument, internally inconsistent — and it is the second recommendation the paper’s evidence supports.

Summary of a published paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What is the puzzle the paper opens with?

That two independent data sources — income and consumption — should identify the same optimal top income tax, but do not. In the static model of Saez (2001), the revenue-maximising top income tax rate is 1/(1 − ζ_Y^I + ζ_Y,τ^H · ρ̄_Y), where ρ̄_Y is the Pareto coefficient of the upper tail of taxable income, ζ^H the compensated elasticity of taxable income, and ζ^I the income effect of a lump-sum levy. But the same static model admits an alternative, consumption-based representation of exactly the same object, obtained from three model-implied identities — that the Pareto coefficients of consumption and income are equal, and that the consumption elasticities map one-for-one into the income ones — all of which follow from the static budget constraint equating consumption to after-tax income. The paper’s observation is that these identities are testable over-identifying restrictions, and they fail in the data: consumption is significantly more evenly distributed than income among top earners. So the static model gives no guidance about which measure to use, and is itself inconsistent with the empirical discrepancy.

Q2. How much does the choice of measure matter numerically?

A great deal: substituting the consumption tail for the income tail in the same formula cuts the implied top rate from 74% to 57%. With the elasticities ζ^H = 0.33 and ζ^I = 0.25 and a Pareto coefficient of yearly income of 1.8, the Saez formula gives 74%. Using a Pareto coefficient for consumption of about 3 — the value the paper cites from Toda and Walsh (2015), Buda et al. (2022) and Gaillard et al. (2023) — the optimal top income tax rate falls to 57%, which the paper notes would increase the after-tax income of top earners by 66%. This is the gap the rest of the paper is built to resolve.

Q3. What is the theoretical set-up?

A Mirrleesian economy in which agents with heterogeneous labour productivities work, consume, and save for retirement. The additional consumption–savings margin is what allows consumption to be separated from after-tax income, and it introduces capital taxes as a second redistributive instrument. The paper describes the baseline model as kept deliberately as simple as possible, and shows in a later section that the results carry over to much more general settings — multi-good economies with arbitrary planner preferences, a life-cycle economy in which tax perturbations have additional inter-temporal spillovers, and models with multi-dimensional heterogeneity where the optimal tax structure is determined by the co-variation of average consumption and savings with income. It also discusses how the model can accommodate heterogeneous asset returns and dynamic income risks.

Q4. What is Theorem 1, and why is it the paper’s organising result?

It establishes that the Saez top income tax formula and uniform commodity taxation are equivalent — neither can hold without the other, except in the trivial case where savings vanish at the top. The revenue-maximising top income tax equals τ_Y^Saez if and only if uniform commodity taxation applies and the optimal savings tax is zero. Conversely, departures from this joint benchmark identify a trade-off: either it is optimal to tax savings and reduce the income tax below τ^Saez, or it is optimal to subsidise savings and raise the income tax above it. The intuition offered is that in a dynamic setting the optimal income tax continues to equal τ^Saez if and only if its design can be reduced to a static trade-off between labour supply and after-tax earnings, with no information needed about how earnings are split between consumption and savings — and that condition holds precisely when preferences for savings are independent of incentives to work, which is the Atkinson–Stiglitz condition.

Q5. What are the two optimality conditions the results are derived from?

A no-arbitrage condition and a revenue-spillover condition. The no-arbitrage condition states that a marginal shift from income to savings taxes (or vice versa) that leaves all agents’ utilities unchanged should not raise or lower tax revenues, since otherwise it would deliver a strict Pareto improvement; it shows that the departure from uniform commodity taxation depends on comparing the ratio of Pareto coefficients ρ̄_Y/ρ̄_C against the ratio of compensated elasticities ζ^H_{C,τ_Y}/ζ^H_{Y,τ_Y}. The revenue-spillover condition augments the static income-tax trade-off with a fiscal spillover from income taxes to savings-tax revenues, and implies that the optimal income tax is strictly lower than (respectively equal to or higher than) the static wedge τ^Saez if and only if the savings tax is positive (respectively zero or negative). Theorem 1 connects the two and provides a simple test for departure from the joint benchmark.

Q6. Why is consumption data essential, rather than savings data?

Because at the top, savings co-move one-for-one with income, so savings data are redundant given income data — while consumption is what pins down how the combined wedge splits. In the empirically relevant case where consumption has a thinner tail than income, the highest earners’ labour supply decision again rests on a static trade-off, but between labour supply and savings rather than consumption. So the static wedge τ^Saez, and hence income data, still matter — but they now determine the optimal combined wedge between income and savings. As the savings share of income converges to one, savings data become uninformative relative to income data, which is why formulas built on the distribution and behavioural responses of savings, in the tradition of Saez (2002), can identify only the combined wedge and not the two rates separately. The paper adds a technical corollary: perturbing one top tax rate at a time no longer suffices, because at the top both perturbations become colinear — labour supply responds only to the combined wedge — so the no-arbitrage condition, which perturbs both taxes simultaneously holding the combined wedge fixed, is essential.

Q7. Why does the model call for shifting the burden onto savings?

Because top earners have a vanishing marginal propensity to consume, and the empirical pass-through evidence is too high for non-homothetic preferences alone to account for the tail gap. If consumption is less concentrated than income, the marginal propensity to consume vanishes at the top. That could arise either because savings have a higher income elasticity than consumption for given preferences — agents view savings as a luxury good relative to consumption — or because preferences for savings correlate positively with labour productivity, or a combination. Empirical studies estimating consumption responses to permanent income changes find pass-through coefficients too high to rationalise the gap entirely through non-homothetic preferences. Hence, in the paper’s words, it rejects uniform commodity taxation in favour of a shift towards positive savings taxes and lower income taxes.

Q8. Where do the Pareto coefficients come from, and how robust are they?

From the 2005–2021 PSID waves, using a statistical procedure that tests the Pareto fit against alternatives: ρ̄_Y = 1.8 for income and ρ̄_C = 3.1 for consumption on average across years. The paper reports that evidence of stable Pareto coefficients already appears around the top 10%, so the result is not restricted to the very top. It also computes the coefficients on average consumption and net worth conditional on income rank — the object the multidimensional generalisation says is the relevant one — by regressing across the 71st, 73rd, …, 99th income quantiles: the estimated slope of ρ̄_Y/ρ̄_C = 0.626 implies ρ̄_C = 2.88 when ρ̄_Y = 1.8, only slightly below the unconditional value. The baseline sets ρ̄_Y = ρ̄_S = 1.8 and ρ̄_C = 3 (so ρ̄_Y/ρ̄_C = 0.6), falling between the two estimates, with robustness at 0.7 (described as a very conservative lower bound to account for potential under-reporting of consumption) and 0.5.

Q9. What scope condition does the paper attach to whom the analysis applies to?

It is explicit that the model is best interpreted as applying to the “rich” but not the “super-rich”. The model features a unique labour income source and is not designed to fully account for income and wealth concentration at the very top — business income, entrepreneurship, capital gains and the like. The paper therefore states it may be best interpreted as applying to high-income working professionals in the top 10% of the population, but not the top 1% or 0.1%, and that the characterisation of optimal income and savings taxes remains valid within the top 5–10% if the relevant sufficient statistics are stable within that range. It also flags a data limitation in the same direction: PSID consumption may not accurately measure the spending patterns of rich and super-rich households, with bequests — as concentrated as wealth — the one important missing category the authors could identify, which they suggest makes it reasonable to read the non-homothetic preference structure over savings as capturing a warm-glow bequest motive.

Q10. Does the model need implausible return heterogeneity to fit the wealth data?

The paper argues not, on a back-of-the-envelope calculation. The slope estimate for net worth of ρ̄_Y/ρ̄_W = 1.316 implies a Pareto coefficient of average net worth of 1.37, close to the unconditional 1.4 reported in Gaillard et al. (2023) and well below the 1.8 for income. Writing average wealth as savings times gross returns, the two can be reconciled by an elasticity of average returns to income of 0.316 — implying that as household income doubles, average compounded returns rise by 13% over a 25-year horizon and average annual returns by roughly 50 basis points. The paper notes it is not aware of studies estimating the elasticity of average returns to income, but argues the required variation appears well within the range of return variation by wealth documented in the empirical literature. The model-implied Pareto coefficient for capital income of 1.2 is described as almost exactly the value estimated by Gaillard et al. (2023).

Q11. What are the labour-supply elasticities set to?

A Hicksian elasticity of 1/3 and an income effect of 1/4, giving the 74.1% static benchmark. The paper cites Chetty’s (2012) meta-analysis for a preferred Hicksian elasticity of 0.33, and Gruber and Saez (2002) for 0.5 for top income earners. It describes evidence on income effects as mixed: Gruber and Saez find small income effects, while Golosov et al. (2021) estimate that $1 of additional unearned income reduces pre-tax income by 67 cents in the highest income quartile, which at a 50% top marginal rate translates into an income effect of 0.33; Vivalt et al. (2024) estimate an income elasticity of earnings of 0.28 for low-income households in a universal basic income field experiment. Robustness is checked at ζ^H = 1/2 (giving τ^Saez = 60.6%) and ζ^I = 1/3 (giving 78.9%).

Q12. What is the hardest identification problem the paper confronts?

Separating non-homotheticity of preferences from unobserved preference heterogeneity — which is precisely what determines whether savings should be taxed. The paper explains why the distinction is not a technicality: the economic rationale to tax savings is the hidden correlation of preferences for savings with permanent income, and many existing empirical studies lack the flexibility to identify both mechanisms separately, leading to biased estimates of the relevant pass-through elasticity. Structural estimates that impose homotheticity — the paper names Heathcote, Storesletten and Violante (2014) — force the pass-through from uninsurable after-tax income shocks to consumption to equal one regardless of risk preferences, so their overall pass-through estimates come out close to the ratio of Pareto coefficients while their conditional pass-through is biased towards homotheticity. Unobserved preference heterogeneity biases cross-sectional or panel estimates in the same direction. Two other complications are flagged: available estimates are typically uncompensated while the formulas need compensated elasticities (the paper derives bounds that hold with equality when preferences are weakly separable in earnings), and the elasticity ratio and the EIS cannot be chosen independently when preferences admit complementarities, yet few studies estimate them jointly.

Q13. What pass-through evidence does the paper use?

Estimates it selects for controlling preference heterogeneity, clustering around 0.7 to 0.9, above the 0.6 ratio of Pareto tails. From the UBI field experiments, Bartik et al. (2024) and Vivalt et al. (2024) imply income effects on consumption of at least 0.65 which, coupled with an income effect on earnings of 0.28, give a lower bound of 0.90 for the elasticity ratio; Bohmann et al. (2025) report German UBI results implying a pass-through of 0.75. From panel work explicitly controlling for preference heterogeneity, Straub (2019) estimates 0.7 instrumenting for permanent income, and Lins (2025) estimates 0.8 controlling for education and other proxies for time-preference heterogeneity. Aggregating Blundell, Pistaferri and Saporta-Eksten’s (2016) individual elasticities to the household level gives 0.79. The paper notes pointedly that the OLS estimates in Straub and Lins that do not control for unobserved preference heterogeneity are around 0.6 — very close to the ratio of Pareto coefficients — which is the bias it warns about. Baseline value: 0.75.

Q14. What EIS does the paper use, and why is it contested?

A baseline EIS of 2, taken from studies of wealthy households’ responses to tax reforms, and the paper is candid that this is near the upper end of existing estimates. Jakobsen et al. (2020), focusing specifically on the wealthiest households, find an EIS as large as 2 and higher for the very wealthy is necessary to replicate the effects of a large Danish wealth tax reform on wealth accumulation; Gruber (2013) and Holm et al. (2024) estimate about 2 and 1.6 respectively from spending responses to capital income tax changes and a dividend tax news shock. The paper acknowledges these estimates are near the upper end of the existing range, possibly because of confounding tax avoidance responses to the policy change, and that macroeconomic estimates from the business cycle literature tend to be closer to 1 while cross-sectional estimates of consumption responses to interest rates are often well below 1. It defends the choice on two grounds: values well above 1 are common in asset pricing (citing Bansal and Yaron’s preferred Epstein–Zin calibration of 1.5, and noting these focus on investors who tend to be among the wealthiest); and within the model the EIS of the richest agents exceeds the population average precisely when the elasticity ratio exceeds 1, so evidence that the EIS increases in wealth supports the calibration.

Q15. What are the headline optimal tax numbers?

A combined wedge of 74.1%, split into a savings tax between 12.5% and 30% and a top income tax between 63% and 70.4%, depending on the calibration. Under homothetic preferences (calibration A), the optimal top savings tax is 22.2% and the income tax falls to 66.7%. Calibrating the pass-through elasticity to 0.75 with an EIS of 1 (calibration B) gives a savings tax of 12.5% and an income tax of 70.4%. Calibrating the EIS to 2 (calibration C) gives a savings tax of 30% and an income tax of 63%, with the top earners’ after-tax income rising by 43 percent relative to the static benchmark. The paper translates the savings wedge into more familiar units: at a 4% annual return over 25 years from working life to retirement, a savings tax of 30% corresponds to a 1.4% annual tax on accumulated wealth, or a 37% capital income tax; read as a model of retirement savings, it means top earners receive a present value of $0.77 in additional pension payments for each additional dollar of social security contributions.

Q16. How sensitive are these numbers, and to what?

Very sensitive to the joint calibration of the Pareto ratio and the pass-through elasticity — the paper reports an almost six-fold range — and much less so once non-homotheticity and heterogeneity reinforce each other. As the ratio of Pareto tails varies from 0.5 to 0.7 and the pass-through from 0.75 to 1, the optimal savings tax moves from 4.5% to 26.3%. The paper reads this as illustrating the critical importance of identifying the pass-through elasticity separately from the ratio of Pareto tails and of controlling for unobserved preference heterogeneity, since it is the latter that determines the strength of the savings-tax motive. Under calibration C the two forces reinforce each other and the variation across the same rows is much smaller, from 25% to 32.2%. Raising the compensated taxable-income elasticity scales down both distortions while leaving the trade-off between them unchanged in calibrations A and B; raising the income effect on labour supply raises the combined wedge and shifts distortions towards savings, so the savings tax is unambiguously increasing in it. Setting the preference complementarity to its theoretical upper bound makes the savings taxes two to three times as large as in the baseline in all three calibrations.

Q17. What are the paper’s three conclusions from the calibration?

First, that some shift from income to savings is optimal across every calibration considered — anywhere between 12.5% and 30% on savings, with income tax reduced accordingly, for reasonable estimates of the Pareto coefficients and behavioural elasticities. Second, that the magnitude of the shift is sensitive to the calibration strategy: matching consumption responses to wealth taxes (calibration C) gives much larger savings taxes and lower income taxes than matching pass-through of permanent income to consumption (calibration B), with the homothetic case in between. The paper draws out the substantive difference between the two: the pass-through evidence, which is not specifically tailored to top earners, suggests the average household treats savings as a luxury relative to regular consumption, while the estimates based on wealth tax reforms suggest the richest households treat their consumption as a luxury.

Q18. What does this imply for policy debates?

That income and savings tax recommendations cannot be made independently — and that a prominent pair of recommendations is, on this argument, mutually inconsistent. The paper observes that the Saez formula and the Atkinson–Stiglitz theorem have both been highly influential but are typically invoked separately, the former in income tax design discussions and the latter in debates about savings taxes. It singles out Diamond and Saez (2011), which simultaneously makes a case for high top income taxes based on τ^Saez and for positive capital taxes by questioning the assumptions underlying uniform commodity taxation, and states that by Theorem 1 the two recommendations are mutually inconsistent. Its own characterisation then supplies the empirical guidance on resolving the trade-off: government revenue is maximised by shifting part of the burden from income to savings, so the paper provides empirical support for the second Diamond–Saez recommendation while invalidating the first.

Q19. How does the paper position itself relative to the savings-tax literature?

As offering a unified representation of both top rates where earlier work characterised one at a time using savings-based statistics. It notes that Scheuer and Slemrod (2021), Gerritsen et al. (2020), Schulz (2021), Ferey, Lockwood and Taubinsky (2024), and Lefebvre, Lehmann and Sicsic (2025) characterise or estimate optimal savings taxes using savings-based sufficient statistics. Its twofold contribution is, first, to combine the revenue-spillover and no-arbitrage conditions so as to identify the joint Saez/Atkinson–Stiglitz benchmark and offer a unified representation of both top rates in a model encompassing these studies; and second — the point it calls most important — to emphasise the role of consumption rather than savings data. It states its model is formally equivalent to Ferey, Lockwood and Taubinsky (2024), with its own characterisation complementing theirs by focusing specifically on top income earners.

Key terms in this paper

Definitions below follow the paper's own usage.

Saez top income tax formula (τ^Saez)
the revenue-maximising top income tax rate in a static economy, expressed in terms of the Pareto coefficient of the upper tail of taxable income, the compensated (Hicksian) elasticity of taxable income, and the income effect of a lump-sum levy on the highest earners.
Uniform commodity taxation (UCT)
the Atkinson–Stiglitz result that it is optimal to leave consumption choices undistorted — and hence not to tax savings — if preferences for savings are independent of labour productivity; savings taxes must therefore be rationalised by departures from this benchmark.
Combined wedge
the overall distortion between earnings and savings at the top, equal to τ^Saez and identified by income data alone; consumption data are what determine how it is decomposed into an income tax and a savings tax.
No-arbitrage condition
the requirement that a marginal shift between income and savings taxes leaving all agents' utilities unchanged should neither raise nor lower revenue, since otherwise it would be a strict Pareto improvement — the condition that perturbs both taxes at once and so identifies the two rates separately.
Revenue-spillover condition
the second optimality condition, which augments the static income-tax trade-off with the fiscal spillover from income taxes to savings-tax revenues, and which implies the optimal income tax lies below τ^Saez exactly when the savings tax is positive.
Non-homotheticity of preferences
the property that savings have a higher income elasticity than consumption for given preferences — agents treating savings as a luxury good relative to consumption — one of two candidate explanations for the gap between the income and consumption tails.
Preference heterogeneity (as used here)
the alternative candidate explanation for that gap — a positive correlation between labour productivity and the preference for savings relative to current consumption — which, unlike non-homotheticity, is what generates the rationale to tax savings, and which the paper argues most existing estimates fail to separate from non-homotheticity.
Pass-through elasticity
the elasticity of consumption to permanent income changes, whose value relative to the ratio of Pareto tail coefficients determines the departure from uniform commodity taxation.
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.