Habit Formation in Consumption and Its Implications for Monetary-Policy Models
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Textbook consumption theory says spending should jump the moment news arrives, but U.S. data show it building gradually and peaking about a year later. This paper asks whether that gap can be closed by assuming people compare today's consumption with what they consumed recently, so they dislike sudden changes as well as low levels. Estimating the model on postwar U.S. data, the author finds habit formation is overwhelmingly supported and reproduces the hump-shaped response the standard model misses. His conclusion is a warning as much as a result: a model that cannot match these dynamics should not be trusted to rank monetary policies.
What this paper finds — and why it matters
The argument starts from a complaint about method: a model used to rank monetary policies must be trusted to represent how consumers and firms actually behave over the policy horizon, and the author contends that most optimisation-based sticky-price models of the late 1990s had not been validated in a way that would earn that trust. Matching first and second unconditional moments is not enough, and neither is matching a single impulse response – especially the response to a monetary policy shock, since the unanticipated component of policy accounts for only a small share of the variance of output, inflation or interest rates. Instead the paper advocates likelihood-based evaluation, comparing the full vector autocovariance function of a structural model against that of an unconstrained VAR in which the structural model is nested. Judged that way, the standard life-cycle consumption model fails in a specific and diagnosable manner: consumption behaves like a “jump variable,” front-loading its entire response to a shock, whereas identified VARs show a gradual hump-shaped response peaking around a year out. Adding Campbell-Mankiw rule-of-thumb consumers does not fix this. The paper’s proposed fix is habit formation in the Carroll-Overland-Weil form, in which utility depends on consumption relative to a reference level built from past consumption; the author’s own explanation of why it works is that this “mixes utility from the level of consumption with utility from the change in consumption,” so consumers acquire a motive to smooth changes as well as levels. He deliberately rules out the alternative repair – assuming serially correlated structural errors – on the grounds that a model with its dynamics hidden in the errors “becomes vulnerable to a Lucas critique of its errors.” Estimating the linearised consumption function by numerical maximum likelihood on U.S. quarterly data for 1966:1-1995:4, the habit parameter comes in at 0.80 with a standard error of 0.19, the rule-of-thumb income share at 0.26, the curvature parameter at 6.11 (implying a small intertemporal elasticity of substitution), and the forward-looking discount parameter at 0.99 per quarter. The restriction that habit formation is unimportant is rejected with a chi-squared statistic of 21.4 and a p-value of 4 times ten to the minus six; the restriction that rule-of-thumb behaviour is unimportant is rejected with a statistic of 12.6 and a p-value of 4 times ten to the minus four. Notably, the full set of cross-equation and zero restrictions implied by the structural model and rational expectations yields a test statistic of 32.8, “not significant at even the 10 percent level” – which the author calls “one of relatively few cases” where an optimisation-based rational-expectations model survives against the VAR that nests it. One estimate comes out lower than expected: the memory parameter in the habit reference level is essentially zero, implying the reference level is simply last quarter’s consumption, and the author defends this at length by showing that a single lag suffices to deliver the smoothing and that substituting a long memory changes the model’s disinflation dynamics only slightly. In a disinflation simulation that cuts the inflation target from about 5% to 2% unexpectedly, consumption in the habit model responds gradually with a peak at about a year and a full response over three to four years, and the author shows that the counterfactually fast real-side response in the no-habit model also significantly damps the persistence of inflation – so misspecifying the real side corrupts the nominal side. He declines to compute optimal policy in the model, giving three reasons: with empirically significant rule-of-thumb consumers it is unclear whose utility to maximise; the model contains no explicit cost of inflation, so the consumer-optimal policy degenerates to minimising consumption fluctuations; and the representative-agent structure is a poor vehicle for welfare costs that arguably fall on discrete employment shifts for a small fraction of the population. He is also explicit that the specification “might not be robust across shifts in monetary or other policy regimes,” and that only testing across regime shifts can settle that.
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What is the paper’s methodological complaint, and what does it propose instead?
That candidate monetary policy models had not been tested against the dynamic behaviour of the data, and that the right test is likelihood-based, using the full vector autocovariance function of an unconstrained VAR that nests the structural model. The author argues that “it is not enough to match first and second unconditional moments, or a subset of conditional moments implied by the model,” as in early equilibrium business cycle studies, because “the working assumption among most economists is that monetary policy has only short-run effects on real variables. If so, it would be a major omission not to fully evaluate the short-run dynamic effects of monetary policy in a candidate model.” He is equally sceptical of impulse-response matching: relying on the response to a federal funds rate shock “may be quite misleading,” since Leeper, Sims and Zha show “the fraction of the variance of output, inflation, or interest rates accounted for by the unanticipated component of monetary policy is generally quite small.” The likelihood is preferred because it “incorporates all of the dynamic covariances among observable variables, weighted according to their contribution to the likelihood,” with the autocovariance function serving as a graphical companion that “may highlight visually a behavioral deficiency in the model that is more difficult to interpret from statistical evidence.”
Q2. What exactly fails in the standard model?
Consumption and investment behave like jump variables, front-loading their entire response to shocks, against identified-VAR evidence of a gradual response peaking around a year. Summarising his earlier work, the author lists for consumption “extremely significant unexplained serial correlation in the consumption-income ratio, parameter estimates that indicate very little or no forward-looking behavior, and excessive sensitivity of consumption to current income arising from the rule-of-thumb behavior,” and for investment “very significant unexplained serial correlation in the investment-capital ratio, extreme sensitivity of model stability and uniqueness to small perturbations in parameter estimates, and a negative estimate of the capital share in income.” The dynamic failure is the one he pursues: disinflation simulations show “both consumption and investment act like ‘jump variables,’ completely front-loading or pulling forward in time their responses to shocks,” whereas identified VARs show “a gradual response over several years, with the peak response at one year or so.”
Q3. Why does the paper refuse to fix the problem with correlated errors?
Because a model whose dynamics live in its error processes cannot claim those errors are policy-invariant. The author acknowledges that “in principle, one can augment a structural model that is dynamically deficient with an arbitrary error structure so as to exactly replicate the dynamic structure in the data,” and names Rotemberg and Woodford as taking that route. His objection is that “the error processes so identified cannot be considered ‘structural’ in any meaningful sense. We have no idea in what way the errors are linked to underlying behavior, and thus we can have no more confidence about their policy invariance than we have for reduced-form VARs or 1960s structural models sans explicit expectations. In essence, by putting the dynamic structure in the errors, the model becomes vulnerable to a Lucas critique of its errors.” He therefore assumes throughout that structural innovations are uncorrelated across time, though possibly correlated across equations, and states that “none of the dynamics in the results reported below may be attributed to across-time correlation in the error terms.”
Q4. What form does habit formation take?
Period utility depends on consumption relative to a reference level that is a geometric average of past consumption, with one parameter indexing how much the reference matters and another indexing how far back its memory reaches. The specification follows Carroll, Overland and Weil, and is “related in spirit to the pioneering work of Duesenberry.” Utility is no longer time-separable “because the consumption choice today influences the future habit reference level in next period’s and all future periods’ utility.” The author flags an admissibility restriction: the habit-importance parameter cannot exceed one, “because it implies that steady-state utility is falling in consumption.” The consumer’s problem is solved with a standard budget constraint and a time-varying real rate, where the ex ante real rate is a discounted weighted average of model-consistent forecasts of short real rates – the duration of the implied long-term real bond being “set to ten years for this paper” – and a separate discount rate for future income, following Campbell and Mankiw, indexes how far forward consumers look.
Q5. How is the model estimated, and on what data?
Numerical maximum likelihood on a system, nested inside an unconstrained VAR, using U.S. quarterly data for 1966:1-1995:4. The estimator follows Fuhrer and Moore (1995); its advantages are that “it allows estimation to proceed naturally from an unrestricted linear vector autoregression that nests all of the linear models considered to successively more-restricted linear models,” and that its finite-sample properties may be better than method-of-moments estimators. The acknowledged drawback is that “to the extent that any equation in the system is mis-specified, estimates of all the parameters in the system will (in principle) be affected.” The VAR contains log per capita nondurables and services consumption, log per capita disposable personal income, the federal funds rate, the price level, and log per capita GDP other than nondurables and services. Consumption and income are chain-weighted, per capita and detrended with the trend segmented in 1974; the price measure is the CPI excluding food and energy; the short rate is the quarterly average effective federal funds rate. In the first stage only the consumption function’s parameters are estimated, with income, funds rate, price and other-GDP equations left as unconstrained VAR equations.
Q6. What do the estimates say?
Habit formation matters a great deal, rule-of-thumb behaviour accounts for about a quarter of income, intertemporal substitution is small, and consumers are highly forward-looking – but the reference level is only one quarter deep. The baseline estimates are 0.80 (standard error 0.19) for the habit-importance parameter, 0.0015 (0.0039) for the memory parameter, 0.26 (0.13) for the rule-of-thumb income share, 6.11 (1.81) for the curvature parameter, 0.99 (0.01) for the forward-looking discount parameter, and 28.49 (5.17) for the coefficient on the expected real interest rate term; the log-likelihood is 2366.4. The author reads these as: “(1) habit formation is an economically important determinant in the utility function; (2) the habit formation reference level is essentially last period’s consumption level; (3) rule-of-thumb behavior is important, with about one-fourth of income accruing to rule-of-thumb consumers; (4) the intertemporal elasticity of substitution is quite small; (5) for those who look forward, the horizon is long; the parameter takes the estimated value .996 on a quarterly basis, .984 on an annual basis; and (6) the model explains most, but not all, of the autocorrelation in the consumption data,” the last shown by a Ljung-Box Q(12) of 83.7 for consumption with a p-value of essentially zero. (The Q statistics for income, the funds rate and inflation, at 6.1, 13.5 and 8.2, are all far from significant.)
Q7. How strongly is habit formation supported statistically?
Overwhelmingly, and the same is true of rule-of-thumb behaviour – while the structural model’s full set of restrictions is not rejected at all. “The hypothesis that habit formation is unimportant in this model – that the exponent on the reference level of consumption is zero – is overwhelmingly rejected. The chi-squared likelihood ratio test for this single restriction takes the value 21.4, with p-value of 4 times ten to the minus six. Similarly, the hypothesis that rule-of-thumb behavior is unimportant is strongly rejected. The chi-squared likelihood ratio test for the restriction takes the value 12.6, with p-value = 4 times ten to the minus four.” The third result is the one the author treats as notable: the likelihood ratio test for the constrained baseline model, “which incorporates the many zero restrictions and cross-equation restrictions implied by the structure of the consumption model and by rational expectations, takes the value 32.8, not significant at even the 10 percent level. This is one of relatively few cases in which the joint restrictions imposed by an optimization-based model with rational expectations cannot be rejected relative to the unconstrained model in which the constrained model is nested.”
Q8. The memory parameter came out near zero. Does that undercut the habit story?
The author says it is “lower than expected” and defends it in two ways rather than explaining it away. First, algebraically: rewriting period utility and setting the memory parameter to zero “shows that the essence of habit formation is that it mixes utility from the level of consumption with utility from the change in consumption. That is, the habit formation model with any normally shaped utility function will imply smoothing of both the level of consumption and its changes (provided [the habit-importance parameter] is not zero). Larger values of [the memory parameter] simply define the changes relative to a longer distributed lag of past consumption.” So “a single lag of consumption in the reference level is sufficient to impart the smoothness to changes in consumption expenditures that is absent in the standard life-cycle model.” Second, numerically: he repeats the disinflation simulation substituting 0.9 for the memory parameter, and reports that “the model’s behavior is altered only slightly by the change from one-quarter memory to more persistent memory in the reference level.” He adds a candid aside in a footnote: in this light “the habit formation may provide a reasonable approximation to a model with a standard utility function and costs of adjustment” in the change in consumption.
Q9. How well does the estimated model match the data’s dynamics?
The consumption dynamics are recovered well, and the remaining discrepancies are inside the VAR’s own confidence bands. “The model has recovered the dynamic covariances for consumption expenditures quite well, capturing the persistence in the autocorrelation, as well as the persistent dynamic correlations between consumption and income, interest rates, and inflation.” Against 90% confidence intervals computed by Monte Carlo around the VAR’s autocovariance function, “the differences between the two autocorrelation functions are generally insignificant at the 10% level. Thus the correlations that the structural model cannot match are generally not precisely determined in the data.” Setting both the habit and rule-of-thumb parameters to zero, by contrast, leaves “the primary consumption dynamics in the model… almost totally missing. The simple PIH model cannot replicate the dynamics in the data.” When further restrictions are added – a simple Taylor rule for the funds rate, a Fuhrer-Moore contracting price specification, and an equation for non-consumption GDP – none “significantly deteriorate the likelihood,” though the author notes specific imperfections: the model makes the correlation of consumption with the lagged funds rate and lagged inflation “too strongly” negative, and makes the correlation between the funds rate and lagged consumption negative “while the VAR says it should be mildly positive.”
Q10. Could the failure instead lie in the price and policy-rule restrictions rather than in consumption?
The paper tests that interpretation directly and rejects it. “An alternative interpretation of the results of this paper and Fuhrer (1996) is that the restrictions imposed on the price specification and the funds rate reaction function are invalid, and are interfering with the real-side dynamics of consumption and output. To test this possibility, I estimate a model with reduced-form processes for consumption and income, so that only the restrictions from the price and interest rate specifications constrain the model.” The comparison “suggests that this interpretation is invalid. The Fuhrer-Moore price specification and the simple reaction function capture the dynamics in these variables without distorting their dynamic interactions with consumption and income (or vice versa).” The converse test, reported in a footnote, runs the other way: “the model with restrictions on consumption, but without restrictions on prices and interest rates, requires rule-of-thumb and habit formation behavior to match the moments in the data.”
Q11. What does the disinflation simulation show?
Habit formation produces the gradual, hump-shaped consumption response the data show, and the absence of it corrupts the model’s inflation dynamics as well. Starting from steady state, the long-run inflation target is cut unexpectedly “from about 5% (the estimate from the data) to 2%.” In the habit model, “inflation falls gradually from its old steady state to the new, lower equilibrium. Interestingly, consumption also responds gradually, with its peak response at a year or so; the full response takes three to four years.” Against this, the no-habit model gives “an example of a model in which the mis-specification of the real side compromises the behavior of the nominal side. The persistence of inflation in this model… is significantly decreased by the rapid (and counterfactual) response of real variables to a disinflationary shock.” The author is careful to note that the problem is not solved by rule-of-thumb consumers alone: “the specification including rule-of-thumb consumers still exhibited rapid response to shocks.”
Q12. Is the linear approximation good enough to carry these results?
The paper checks and finds it is. Solving the nonlinear model with the parameters estimated from the linear one, for the same disinflation simulation, gives “nearly identical results.” Substituting the linear model’s solutions for consumption, income and real rates into the nonlinear first-order conditions, “the maximum absolute error in the nonlinear Euler equations is about .01, compared to steady-state marginal utility of about -1.” And “the estimate of lifetime utility for the disinflation simulation is very similar whether computed using the solution paths from the linear model or from the nonlinear model.” The author does flag the method used: the nonlinear solutions use “a ‘certainty equivalence’ solution technique that does not compute the stochastic distribution of the endogenous variables via value-function programming,” because “the state dimension of the model would make the computation time for such a method prohibitive.”
Q13. Why doesn’t the paper compute optimal monetary policy?
Three reasons, all about what the model cannot support rather than about computation. First, “given the empirical significance of consumers who appear to follow a rule of thumb, it is difficult to know whose utility we should maximize in evaluating policy alternatives. One could maximize the utility of the (forward-looking, rational) habit-formation consumers, but one could not know the welfare implications for the rule-of-thumbers.” Second, “the model as it stands includes no explicit cost of inflation” – inflation reaches consumers only indirectly, through the real disruptions the Fed causes in bringing it back to target – so “without any direct cost of inflation, the optimal policy from the consumers’ point of view is one that minimizes fluctuations in consumption. This is not a satisfying or interesting policy conclusion.” Third, “because the bulk of the welfare cost arguably arises from discrete shifts in employment status for a small fraction of the population, the representative agent model may not provide an accurate measure of the relative welfare costs of pursuing different monetary policies.” He suggests the usual weighted output-and-inflation loss function “may be a reasonable approximation to use until the thornier issues… are tackled,” and “leave[s] the computation of ‘optimal’ policy responses in this model for future work.”
Q14. What does the paper claim, and what does it decline to claim?
It claims that only specifications imposing smoothness on the change in consumption will succeed empirically, and that the hump-shaped response is a statistically significant feature of the data; it declines to claim that habit formation is the unique or regime-invariant explanation. The conclusion states that habit formation “allows the model to replicate key dynamic correlations among consumption, output, interest rates, and inflation to a degree that standard models cannot,” specifically the hump-shaped response “to income, interest rate, and inflation shocks,” and that it does so “because it imparts a motive for consumers to smooth the change, as well as the level of consumption.” But immediately: “Other specifications may also afford improvements in the empirical performance of the standard model. This paper suggests, however, that only specifications that impose some smoothness on the change in consumption will be successful empirically.” And on stability: “The specification set forth in this paper might not be robust across shifts in monetary or other policy regimes. But only through rigorous econometric testing of this and alternative specifications across regime shifts can observational equivalence (or empirical dominance) of alternative specifications and stability of any one specification across policy shifts be determined.” The author’s own summary of the contribution is deliberately modest – “a small step,” “a modest improvement.”
Key terms in this paper
Definitions below follow the paper's own usage.
- Habit formation
- in this paper, the Carroll-Overland-Weil specification in which period utility depends on consumption relative to a reference level that is a geometric average of past consumption, governed by two parameters -- one indexing how much the reference level matters and one indexing how far back the "memory" reaches. The author's own gloss on why it works is that it "mixes utility from the level of consumption with utility from the change in consumption," so that any normally shaped utility function built this way implies smoothing of both, and he notes it "may provide a reasonable approximation to a model with a standard utility function and costs of adjustment" in the change in consumption.
- Hump-shaped response
- the gradual, delayed response of real spending to a shock -- building over time, peaking around a year, and taking three to four years to complete -- which identified VAR studies find in the data and which the standard life-cycle model, even augmented with rule-of-thumb consumers, cannot generate; the paper treats reproducing this shape as the central empirical test, and argues that "only specifications that impose some smoothness on the change in consumption will be successful empirically."
- Jump-variable behaviour
- the author's term for the behaviour of consumption and investment in standard optimising models, which "completely front-load or pull forward in time their responses to shocks" rather than adjusting gradually; he identifies this as the specific counterfactual feature that habit formation is introduced to fix, and shows in a disinflation simulation that it also corrupts the nominal side of the model by damping the persistence of inflation.
- Likelihood-based model evaluation
- the paper's proposed standard for judging monetary policy models -- comparing a structural model's full vector autocovariance function against that of an unconstrained VAR in which the structural model is nested, and testing the restrictions by likelihood ratio. The author argues explicitly against the alternatives: matching first and second unconditional moments is not enough, and matching a single impulse response (especially to a monetary policy shock, which accounts for little of the variance of output, inflation or interest rates) "can be quite misleading."
- Rule-of-thumb consumers
- Campbell and Mankiw's consumers whose current consumption equals current income; the paper allows a fraction of total income to accrue to them and estimates that fraction at about one-fourth, finding the restriction that it is zero strongly rejected. The author's point is that this behaviour is necessary but not sufficient -- the life-cycle model augmented with rule-of-thumbers "still exhibited rapid response to shocks," so it does not on its own deliver the hump.
- Lucas critique of the errors
- the author's argument that a structural model which repairs its dynamic failings by assuming serially correlated errors rather than by changing behaviour "becomes vulnerable to a Lucas critique of its errors" -- since nothing links those error processes to underlying behaviour, one can have no more confidence in their policy invariance than in a reduced-form VAR. This is why he restricts structural innovations to be uncorrelated across time throughout the paper.