Macro Paper Warehouse
Online First [Journal of Money, Credit and Banking] doi:10.1111/jmcb.70065 Online 4 Jun 2026

Another Look at the (Ir)Relevance of Long-Run Risks for Equity Risk Premia

Paulo Maio

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Can investors' fear of slow-moving shifts in how fast consumption grows -- and in how uncertain that growth is -- explain why some groups of U.S. stocks earn higher returns than others? Using quarterly data from 1963 to 2018 and portfolios built around seven well-known return patterns, this paper reports the answer is no: pricing errors stay large, the model's fit is usually worse than assuming every portfolio earns the average premium, and the key preference parameter comes out near zero or negative rather than above one as the literature assumes. That matters because much macro-finance work rests on those assumed values.

What this paper finds — and why it matters

This paper derives a three-factor consumption-based asset pricing model that extends the baseline Consumption CAPM by adding two long-run risk factors – the innovation in expected future consumption growth (“consumption growth news”) and the innovation in the expected future variance of consumption growth (“consumption variance news”) – and then asks whether those factors are priced in a reasonably demanding cross-section of U.S. equity risk premia. Because the model is derived from recursive (Epstein-Zin) preferences, its three factor risk prices map directly onto the coefficient of relative risk aversion and the elasticity of intertemporal substitution, so the cross-sectional estimates can be read as estimates of the preference parameters the long-run risks literature calibrates. Using quarterly data from 1963:III to 2018:IV and decile portfolios sorted on seven prominent CAPM anomalies – book-to-market, asset growth, price momentum, accruals, net stock issues, operating profitability, and residual return variance – plus an augmented test that prices 43 excess returns simultaneously, the author finds the model is largely rejected on both statistical and economic grounds: in no single estimation does it deliver small pricing errors together with economically plausible and statistically significant risk prices, the cross-sectional explanatory ratio is negative in most cases, and the estimated elasticity of intertemporal substitution is near zero, insignificant, or negative rather than above one. The conclusion the author draws is that long-run consumption risks do not rescue the Consumption CAPM, which he presents as a major challenge for the voluminous long-run risks literature; the finding is stated as an empirical rejection within this model class, this VAR-based factor construction, and this test-asset set, not as a general claim about all consumption-based models.

Summary of a forthcoming paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What exactly is the model, and how does it differ from earlier attempts to bring long-run risks into the cross-section?

The paper derives a three-factor model in which the factors are current consumption growth, consumption growth news, and consumption variance news, and in which the risk prices are tied to both recursive-utility preference parameters – relative risk aversion (gamma) and the elasticity of intertemporal substitution (psi). The baseline CCAPM of Lucas (1978) and Breeden (1979) and the two-factor model of Restoy and Weil (2011) are both nested as special cases: the CCAPM obtains when gamma = 1/psi. The author positions the model against two nearby papers. Relative to Bansal et al. (2014), there are “two major differences (among other discrepancies)”: all three factors here are directly related to aggregate consumption, and the risk prices depend on both gamma and psi rather than on gamma alone. Relative to Constantinides and Ghosh (2011), who substitute the two long-run consumption factors with observables (innovations in the price-to-dividend ratio and the risk-free rate), this paper keeps all three factors connected to consumption, which the author argues is “a sensible choice” if the object of interest is the original long-run risks framework.

Q2. What data and test assets are used, and over what period?

All data are quarterly; the VAR estimation covers 1963:III to 2018:IV, with the start date constrained by the availability of some of the portfolio return data. Consumption growth is log real per-capita consumption of non-durables and services, built from St. Louis Fed data using the “end-of-period” timing convention. The benchmark VAR state vector (VAR I) is the real short-term interest rate, the log dividend-to-price ratio for the S&P index (Shiller), the log consumption-to-wealth ratio cay (Lettau and Ludvigson), the log real stock market return (CRSP), the realized consumption variance, and consumption growth. The test assets are seven groups of decile portfolios from Kenneth French’s data library, corresponding to book-to-market (BM10), asset growth (AG10), price momentum (M10), accruals (ACC10), net stock issues (NSI10), operating profitability (OP10), and residual stock return variance (RVAR10) – portfolio sets used by Fama and French (1996; 2015; 2016) to test their own factor models. Alongside the seven single-anomaly estimations there is an augmented estimation that drops the four intermediate deciles in each group (to keep the cross-section from overwhelming the time-series dimension) and adds the market factor, giving 43 excess portfolio returns.

Q3. How poorly does the baseline CCAPM itself do, before long-run risks are added?

Very poorly: the cross-sectional explanatory ratio is negative in all seven single-anomaly tests, and the risk aversion estimates required are extremely large. Risk aversion estimates range between 153 (in the estimation with RVAR10) and 267 (with BM10), and only three of the seven (BM10, AG10, ACC10) are significant at the 5% level, the remaining four only at the 10% level. In the augmented estimation with the 42 extreme portfolios plus the market factor, the R-squared is -0.75, the average absolute pricing error is 0.70% per quarter, and the implied risk aversion estimate of 180 is only marginally significant (10% level). The author draws the moral explicitly: “unlike the univariate case of the aggregate equity premium, large risk aversion estimates are not sufficient to generate small or zero pricing errors,” because the model must price 10 or 43 excess returns rather than a single risk premium.

Q4. Does adding consumption growth news alone (the nested two-factor model) help?

No – the two-factor Restoy-Weil model is “qualitatively very similar” to the baseline, producing negative explanatory ratios for all seven portfolio groups. In the augmented 43-portfolio estimation its R-squared is -0.60 and its average pricing error 0.64% per quarter, only marginally below the single-factor model’s. Where consumption growth news does earn a significantly positive risk price (OP10, ACC10), the short-run consumption risk price turns significantly negative, so the implied gamma and psi are significantly negative – which the author describes as “entirely at odds with theory” and “a major economic rejection of the model.” Where the implied psi estimates are positive (BM10, AG10, NSI10), their magnitudes are “very close to zero and are clearly insignificant (even at the 10% level).”

Q5. What happens with the unrestricted three-factor model, LRCCAPM1?

LRCCAPM1 “does not outperform the two-factor model on a qualitative basis.” It yields a positive explanatory ratio in three single-anomaly estimations (BM10, AG10, RVAR10), but the levels of fit are between 4% and 17% and “far from statistically significant,” with bootstrap p-values so large that the null of a zero fit cannot be rejected. In the augmented estimation the R-squared is -0.36 and the average pricing error 0.56% per quarter – only slightly below the two-factor model – so “the addition of the variance news factor does not materially improve the model’s ability to explain risk premia.” The variance-news risk price is negative in six of the seven single-anomaly estimations (M10 the sole exception), the sign the theory predicts, but significant only with AG10 or RVAR10; the consumption-growth-news risk price is significantly positive only with OP10 and ACC10; and the short-run consumption risk price (hence gamma) is insignificant even at the 10% level across the single-anomaly estimations.

Q6. Does the restricted version that is closer to the theory do better?

No. LRCCAPM2, in which only two structural parameters are identified so the variance-news risk price is no longer freely estimated, “does not perform significantly better than LRCCAPM1,” and there is not a single cross-sectional estimation in which it is validated on both statistical and economic grounds. The cross-sectional R-squared is negative in all estimations, and because LRCCAPM2 is a restricted version of LRCCAPM1 it produces larger pricing errors than LRCCAPM1. Its one improvement is more credible estimates of gamma, which the author judges “far from enough to rescue the model” given the very large pricing errors and, more damagingly, negative estimates of psi in all cases – most of them statistically significant.

Q7. And the version built to match the theory’s stochastic-volatility structure?

LRCCAPM3, whose three factors come from a heteroskedastic VAR estimated by weighted least squares with all variables scaled by the inverse of realized consumption volatility, “does not do better than LRCCAPM2, being strongly rejected on both economic and statistical grounds.” The explanatory ratios are negative in all cases and the estimates of psi are largely insignificant even at the 10% level in nearly all estimations. Where gamma estimates are statistically significant they take extreme magnitudes, “around or above 200,” which the author counts as a further violation of the model on economic grounds. This version matters for the paper’s framing because the author shows (Section 6) that a restricted version of LRCCAPM3 is equivalent to the original long-run risks model of Bansal and Yaron (2004), so the rejection is connected to that model rather than only to a loose empirical cousin of it.

Q8. Why is the elasticity of intertemporal substitution the parameter that carries the argument?

Because the original long-run risks model’s performance for the aggregate equity premium and risk-free rate “critically depends on calibrations of psi above one,” and this paper’s cross-sectional estimates are nowhere near that. Across the many specifications estimated for the three versions of the model, only one case delivers an estimate of psi that is positive and statistically significant at the 5% level, and even there the magnitude is “around zero (0.01).” The author notes that low estimates of this parameter are not peculiar to his setting – other recursive-utility cross-sectional tests produce values very close to zero – and cites Havranek (2015), who examines 2,735 estimates of psi reported in 169 studies and concludes that “calibrations greater than 0.8 are inconsistent with the bulk of the empirical evidence.”

Q9. How is the model evaluated, and why does the paper distrust the formal specification test?

Evaluation rests on the cross-sectional explanatory ratio with bootstrapped empirical p-values, the average absolute pricing error, and pairwise R-squared comparisons against the nested models, with the chi-squared specification test treated as unreliable. In several estimations the model is not rejected by the chi-squared test (p-values above 5%) despite negative or very small explanatory ratios – for LRCCAPM1 this happens with M10, BM10, AG10, OP10, and the all-groups test. The author reads this as confirming “the severe limitations of the chi-squared statistic” rather than as evidence for the model. He also flags that the reported standard errors on risk prices, and hence on the implied preference parameters, “are likely to understate the ’true’ statistical uncertainty” because both long-run factors are estimated from a first-order VAR rather than observed.

Q10. What robustness checks are run?

The negative verdict survives imposing alternative values of the log-linearization parameter, estimating alternative VAR state vectors containing other macro variables, using a different proxy for the realized consumption variance, employing an alternative definition of consumption growth, and relying on a different definition of consumption variance news. The alternative realized-variance measure averages squared consumption growth over four consecutive quarters instead of two; the alternative variance-news definition sets the conditional variance equal to the current realized variance, which has the advantage of precluding negative fitted values for expected realized variance. The author reports that the general pattern is robust across these alternative specifications for each of the three model versions.

Q11. What does the paper claim, and what does it deliberately not claim?

The claim is that long-run consumption risks “do not seem to represent a valid solution to rescue the consumption asset pricing framework,” and that the results “question the ability” of long-run risks to solve the equity premium and risk-free rate puzzles – not that consumption-based pricing is hopeless in general. A footnote is explicit that “other, equally parsimonious, consumption-based linear factor models have some success when it comes to explaining cross-sectional equity risk premia,” citing Yogo (2006), Savov (2011), Lioui and Maio (2014), Kroencke (2017), and Chen and Lu (2018). The author also lists the openings his design leaves: extending the cross-sectional tests to other asset classes such as bonds or currencies; the possibility that the findings arise from a limited participation problem, in which aggregate consumption is not a good proxy for the consumption of the average stock market investor, making disaggregated-consumption long-run risks worth testing; alternative proxies for the conditional consumption volatility process, such as the multivariate GARCH approach of Ma (2013); and modelling long-run risks for consumption growth and inflation jointly, as in Creal and Wu (2020).

Q12. How does this evidence sit relative to earlier cross-sectional tests that reported success?

The paper argues the earlier tests were run against cross-sections too thin to discriminate. Constantinides and Ghosh (2011) use only four equity portfolios – two sorted on book-to-market and two on equity market capitalization – plus the market portfolio and the risk-free rate; since the U.S. size premium “has disappeared for several decades,” the author reads their cross-section as essentially accounting for the value premium alone, “a very modest hurdle to judge the capacity of an asset pricing model (containing three factors) for cross-sectional risk premia.” Bansal, Kiku, and Yaron (2016) employ the same four portfolios and claim “success,” but the author objects that the magnitudes of the pricing errors – either as a fraction of the raw risk premia to be explained or in comparison with a benchmark such as the Fama-French (1993) three-factor model – are not reported. He also notes that even on that very small cross-section Constantinides and Ghosh’s model is formally rejected by the J-statistic, and that their reported estimate of psi above one is “clearly insignificant.”

Key terms in this paper

Definitions below follow the paper's own usage.

Long-run risks (LRR)
in this paper, the two consumption-based risk factors that extend the baseline consumption model: the innovation in expected future consumption growth and the innovation in the expected future variance of consumption growth, both derived endogenously from a recursive-utility pricing kernel rather than imposed as empirical factors -- so that a test of whether they are priced in the cross-section is a test of the long-run risks framework itself.
Consumption growth news
the revision, between t and t+1, in the discounted sum of expected future consumption growth; estimated here as a linear combination of the shocks from a first-order VAR that includes the real short rate, the log dividend-to-price ratio, the consumption-to-wealth ratio, the real market return, the realized consumption variance, and consumption growth.
Consumption variance news
the revision in the discounted sum of the expected future conditional variance of consumption growth, where the unobserved conditional variance is proxied by the conditional expectation of next period's realized variance, itself a two-quarter moving average of squared consumption growth (a four-quarter version is used as a robustness check).
Elasticity of intertemporal substitution (psi)
the recursive-utility preference parameter separated from risk aversion under Epstein-Zin preferences; the paper treats it as the load-bearing quantity because the long-run risks literature's resolution of the equity premium and risk-free rate puzzles depends on calibrating it above one, whereas the cross-sectional estimates here are near zero, insignificant, or negative.
Explanatory ratio (cross-sectional R-squared)
the cross-sectional OLS R-squared of realized average excess returns on model-predicted premia, with empirical p-values from a bootstrap simulation; a negative value means the model explains cross-sectional dispersion in risk premia worse than a trivial model that assigns every test asset the cross-sectional average risk premium.
LRCCAPM1, LRCCAPM2, LRCCAPM3
the paper's three empirical versions of the same three-factor model: an unrestricted version whose three risk prices are all freely estimated (LRCCAPM1); a restricted version in which only two structural parameters are identified so the variance-news risk price is no longer free (LRCCAPM2); and a version whose factors come from a heteroskedastic VAR estimated by weighted least squares, with all shocks scaled by realized consumption volatility, a restricted case of which is equivalent to the original Bansal-Yaron (2004) long-run risks model (LRCCAPM3).
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.