The equity premium: A puzzle
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Why have U.S. stocks paid investors so much more, on average, than safe government debt? Using ninety years of data from 1889 to 1978, Mehra and Prescott find that stocks earned about seven percent a year after inflation while safe short-term securities earned less than one percent -- a six-point gap. Testing a standard economic model of saving and risk-taking, calibrated to match how bumpy actual U.S. consumption growth has been, they find it can generate at most a few tenths of a point of that gap, even allowing realistic caution toward risk. The shortfall suggests something is missing from how economists model the tradeoff between risk, reward, and time.
What this paper finds — and why it matters
This 1985 Journal of Monetary Economics paper by Rajnish Mehra and Edward C. Prescott documents that over the ninety-year period 1889-1978, the average real annual return on the Standard and Poor’s 500 Composite Index was 6.98 percent while the average real return on short-term, virtually default-free securities was only 0.80 percent, an average equity premium of 6.18 percent (standard error 1.76 percent), and asks whether this gap can be accounted for by a class of general-equilibrium pure-exchange models that abstract from transactions costs, liquidity constraints, and other frictions absent from the Arrow-Debreu framework (Sections 1-2, pp. 145-147, 155-156). Using a variant of Lucas’s (1978) exchange-economy model – modified so that the growth RATE, rather than the level, of the (single) representative household’s consumption endowment follows a two-state Markov process, in order to accommodate the sustained, non-stationary growth actually observed in U.S. per capita consumption – the authors calibrate the consumption process to match the historical mean, standard deviation, and first-order serial correlation of U.S. consumption growth, then search over the whole empirically defensible range of risk-aversion (0 to 10) and time-discount parameters to see what combinations of average risk-free rate and average equity premium the model can produce (Sections 3-4, pp. 150-155). They find that the model can generate a maximum average equity premium of only 0.35 percent – compared with the 6.18 percent observed – and that this conclusion is robust to measurement error in inflation, to the level of temporal aggregation, to alternative specifications of the consumption process (including higher moments), and to introducing financial leverage on the modeled firm or extending the model to allow production and capital accumulation (Section 4, pp. 155-158). The intuition, as the authors explain it, is that in a growing economy the levels of risk aversion needed to generate a six-percent premium also imply implausibly high average real interest rates, since sufficiently risk-averse agents discount future (higher) consumption so heavily that the model-implied risk-free rate rises far above the 0.80 percent actually observed (Section 1, pp. 146-147; Section 5, p. 157). The paper concludes that the puzzle can equally well be framed as asking why the risk-free rate was so low rather than why the equity return was so high, and conjectures that resolving it will most likely require an equilibrium model with some friction or market incompleteness – rather than a complete-markets Arrow-Debreu economy – doubting that either investor heterogeneity or non-time-additive preferences alone will suffice (Section 5, pp. 157-159).
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What empirical fact motivates the paper, and over what period is it documented?
Over the ninety-year period 1889-1978, the average real annual return on the Standard and Poor’s 500 Composite Index was 6.98 percent, while the average real annual return on short-term, virtually default-free securities was only 0.80 percent, yielding an average equity risk premium of 6.18 percent (standard error 1.76 percent) (Section 2, Table 1, pp. 146-147, 155-156). The data are drawn from five series matching those used by Grossman and Shiller (1981): a real stock price index and real dividends for the S&P 500, per capita real consumption on nondurables and services (Kuznets-Kendrick-USNIA), a consumption deflator, and the nominal yield on relatively riskless short-term securities (Treasury Bills from 1931, Treasury Certificates 1920-1930, and Prime Commercial Paper before 1920) (Section 2, pp. 147-148).
Q2. What kind of model do Mehra and Prescott use to ask whether this premium is explicable, and how does it differ from Lucas’s (1978) framework?
The paper uses “a variation of Lucas’ (1978) pure exchange model,” with one important modification: because U.S. per capita consumption has grown persistently over time, the authors assume that the growth RATE of the endowment (rather than its level, as in Lucas’s original model) follows a Markov process, an assumption that “requires an extension of competitive equilibrium theory” established separately in Mehra and Prescott (1984) and lets the model capture the non-stationarity in the consumption series while keeping growth rates and asset returns themselves stationary (Section 3, pp. 150-151). A single representative “stand-in” household maximizes the expected discounted sum of constant-relative-risk-aversion (CRRA) utility over consumption of a single perishable good produced by one productive unit, whose output (and dividend) follows a two-state Markov chain on its growth rate (Section 3, pp. 151-152).
Q3. How do the authors calibrate and “test” the model, and how do they characterize this exercise methodologically?
The authors explicitly frame the exercise as “a quantitative theoretical exercise” rather than “an estimation exercise designed to obtain better estimates of key economic parameters” (Section 1, p. 146). They select the technology parameters of the two-state Markov chain (average growth rate, its standard deviation, and first-order serial correlation) so that the model’s stationary distribution matches the historical U.S. values (0.018, 0.036, and -0.14, respectively) exactly (Section 4, pp. 153-154). They then search over the full empirically defensible range of the risk-aversion parameter alpha (restricted to 0 to 10, based on a survey of estimates from Arrow, Friend and Blume, Kydland and Prescott, Altug, Kehoe, Hildreth and Knowles, and Tobin and Dolde, cited on p. 154) and the discount factor beta (0 to 1), asking whether ANY combination in this admissible parameter space can jointly reproduce the observed risk-free rate and equity premium (Section 4, pp. 154-156).
Q4. What is the model’s maximum attainable equity premium, and how does that compare with the historical value?
The maximum average equity premium the model can generate, for any risk-aversion parameter up to 10 and any discount factor between zero and one, is 0.35 percent – a “sharp contrast to the six percent premium observed” (Section 1, p. 145; Section 4, pp. 155-156). Figure 4 plots the entire “admissible region” of (average risk-free rate, average equity premium) pairs the model can produce while keeping the average risk-free rate between zero and four percent, and the observed U.S. pair of 0.80 percent and 6.18 percent lies well outside that region (Section 4, p. 156).
Q5. What is the economic intuition for why the model cannot generate a large premium without also generating an implausibly high risk-free rate?
With real per capita consumption growing at nearly two percent per year on average, the paper explains, “the elasticities of substitution between the year t and year t+1 consumption good that are sufficiently small to yield the six percent average equity premium also yield real rates of return far in excess of those observed.” In a growing economy, highly risk-averse (or, equivalently, highly consumption-smoothing) agents discount the future more heavily than less risk-averse agents would, because future consumption will on average exceed present consumption and its marginal utility is correspondingly lower – so real interest rates rise with the degree of risk aversion needed to generate a large premium (Section 1, pp. 146-147). For a curvature parameter of 2, for instance, the model’s average risk-free rate is at least 3.7 percent, far above the observed 0.80 percent, whose sample standard deviation is only 0.60 (Section 5, p. 157).
Q6. What robustness checks do the authors perform, and does any of them narrow the gap?
None of the robustness checks meaningfully narrows the gap. The authors check sensitivity to: (i) measurement error in inflation, which biases the real risk-free rate and equity return by the same amount and so does not affect the estimated risk premium itself (Section 4.1, p. 156); (ii) the length of the underlying time period, varied from 1/128 of a year to two years, which had “negligible effect” on the admissible region (Section 4.1, p. 156, Appendix pp. 159-160); (iii) the average consumption growth rate mu (varied between 1.4 and 2.2 percent), which barely changed the result; (iv) the standard deviation of consumption growth delta, to which the equity premium is roughly proportional in its square, and the persistence parameter phi, decreases in which reduced the premium (Section 4.1, pp. 156-157); and (v) higher moments of the consumption growth distribution, introduced via a four-state Markov chain while holding the first two moments fixed, which raised the maximum attainable premium only from 0.35 to 0.39 percent (Section 4.1, pp. 156-157, Appendix p. 160).
Q7. Does allowing for firm leverage, or introducing production and capital accumulation, change the conclusion?
No. To better approximate an actual leveraged corporation, the authors re-price a security whose dividend is output net of a fixed fraction (theta = 0.9, reflecting a roughly ten percent corporate profit share) of expected output committed in advance; this “increased the equity risk premium by less than one-tenth percent,” because financial arrangements do not affect resource allocation or the underlying Arrow-Debreu prices (Section 4.2, p. 158). The authors also argue, citing a companion result (Mehra, 1984), that admitting capital accumulation and production “cannot overturn our conclusion, because expanding the set of technologies in this way does not increase the set of joint equilibrium processes on consumption and asset prices” (Section 4.3, p. 158).
Q8. How does the paper reframe the puzzle in its conclusion, and what does that reframing depend on?
The authors suggest that “the equity premium puzzle may not be why was the average equity return so high but rather why was the average risk-free rate so low” – a reframing that follows if one accepts the Friend and Blume (1975) finding that the curvature parameter alpha significantly exceeds one, since at alpha = 2 the model’s average risk-free rate (at least 3.7 percent) is far above the 0.80 percent sample average (Section 5, p. 157). They note this is not an isolated anomaly: currency, for example, is dominated by Treasury bills with positive nominal yields, yet sizable amounts of currency are held, another case of an asset earning a lower return than Arrow-Debreu theory implies (Section 5, p. 157).
Q9. What candidate resolutions does the paper consider, and which does it doubt?
The authors doubt that investor heterogeneity by itself resolves the puzzle, citing Constantinides (1982)’s finding that heterogeneous-agent economies within the Debreu (1954) competitive framework impose the same tested restrictions; they also doubt that non-time-additive-separable preferences would resolve it, since that would require near-term consumptions to be POORER substitutes than widely separated ones (Section 5, pp. 157-158). Instead, they conjecture that some departure from the complete-markets Arrow-Debreu setting – a “friction” such as legal or borrowing constraints, non-enforceable contracts, or the infeasibility of contracting with as-yet-unborn generations – will be needed to rationalize the observed premium, and note that testing such theories would likely require consumption data disaggregated by income or age group (Section 5, pp. 158-159).
Q10. What are the model’s key simplifying assumptions and self-acknowledged limits?
The model has a single representative household and a single productive unit whose entire output is the dividend paid on one traded equity share, so “the security priced in our model does not correspond to the common stocks traded in the U.S. economy” – an actual economy has a near-continuum of capital types with differing risk characteristics, and the paper’s leverage exercise (Q7) is only a partial attempt to bridge this gap (Section 4.2, p. 158). The endowment process is exogenous, with neither capital accumulation nor production in the baseline specification (Section 4.3, p. 158). The authors explicitly caveat that their simple class of economies, while well suited to this particular question, “clearly is poorly suited for other issues, in particular issues such as the volatility of asset prices” (Section 1, p. 146). The risk-aversion parameter is restricted a priori to the 0-10 range based on the existing empirical literature, a restriction the authors say is important because with unrestricted, very large risk aversion “virtually any pair of average equity and risk-free returns can be obtained by making small changes in the process on consumption” (Section 4, pp. 154-155).
Key terms in this paper
Definitions below follow the paper's own usage.
- Equity premium
- the difference between the average real return on a risky, broadly diversified equity security (in this paper, the Standard and Poor's 500 Composite Index) and the average real return on a short-term, virtually default-free security; the paper documents an average of 6.98 percent on equity versus 0.80 percent on the riskless security over 1889-1978, an equity premium of 6.18 percent (standard error 1.76 percent) (Section 2, pp. 147, 155-156).
- Constant relative risk aversion (CRRA) preferences
- the class of period utility functions U(c) = (c^(1-alpha) - 1)/(1-alpha) used throughout the paper (with the logarithmic function as the limiting case at alpha = 1), where the single parameter alpha indexes the household's willingness to substitute consumption between adjacent time periods and its aversion to consumption risk; the paper restricts attention to 0 < alpha <= 10 on the grounds that this range is what the empirical literature on risk aversion (Arrow, Friend and Blume, Kydland and Prescott, and others cited on p. 154) supports (Section 3, pp. 151-152, 154).
- Admissible region (of risk-free rate/equity-premium pairs)
- the set of (average risk-free rate, average equity risk premium) pairs that the model can generate for any preference-parameter pair (alpha, beta) with 0 < alpha < 10 and 0 < beta < 1, given the consumption process calibrated to U.S. 1889-1978 data; plotted in fig. 4, the observed U.S. pair (0.80 percent, 6.18 percent) lies far outside this region, whose maximum attainable premium is 0.35 percent (Section 4, pp. 155-156, Appendix, pp. 159-160).
- Growth-rate (rather than endowment-level) Markov process
- a modification of Lucas's (1978) pure-exchange asset-pricing model in which the growth RATE of the aggregate endowment, rather than its LEVEL, is assumed to follow a Markov process; this extension, requiring an extension of competitive equilibrium theory established in Mehra and Prescott (1984), is needed to capture the sustained, non-stationary rise in U.S. per capita consumption over the ninety-year sample while keeping the growth-rate process itself, and hence asset returns, stationary (Section 3, pp. 150-151).
- Non-Arrow-Debreu resolution
- the paper's concluding conjecture that, since a complete-markets Arrow-Debreu exchange economy cannot rationalize the observed premium, "most likely some equilibrium model with a friction" -- some form of market incompleteness such as borrowing/legal constraints, non-enforceable contracts, or the infeasibility of trading with as-yet-unborn generations -- will be the one that does (Section 5, p. 158, and Section 1, p. 145).