<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Nowcasting-Forecasting | Macro Paper Warehouse</title><link>https://macropaperwarehouse.com/topics/nowcasting-forecasting/</link><atom:link href="https://macropaperwarehouse.com/topics/nowcasting-forecasting/index.xml" rel="self" type="application/rss+xml"/><description>Nowcasting-Forecasting</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><item><title>Misspecified Expectations among Professional Forecasters</title><link>https://macropaperwarehouse.com/papers/misspecified-expectations-among-professional-forecasters/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/misspecified-expectations-among-professional-forecasters/</guid><description>&lt;p&gt;Analyzing panel data from the U.S. Survey of Professional Forecasters (SPF, 1992Q1–2019Q4, 77 forecasters, 1,520 forecaster-quarter observations), Julio Ortiz finds that a &amp;ldquo;misspecified expectations&amp;rdquo; model — in which forecasters perceive an AR(2) data-generating process to be an AR(1), causing them to misperceive its underlying persistence — tends to outperform a noisy-information rational benchmark and two leading non-FIRE alternatives (overconfident and diagnostic expectations) when fit to forecast errors and revisions. The models are estimated by maximum likelihood and ranked using forecast-encompassing weights; for the baseline real GDP growth case, misspecified expectations earns the largest encompassing weight (0.539 vs. 0.462 for diagnostic, ~0 for rational and overconfident) and the highest log-likelihood. Across 14 macroeconomic variables, misspecified expectations provides the best fit for most series both in-sample and out-of-sample, though diagnostic expectations fits better for some (e.g., GDP deflator, industrial production, real residential investment) and rational expectations fits the unemployment rate best. The author argues misspecified expectations succeeds in part because its bias enters both the prediction and updating equations, producing overreaction to new information plus overextrapolation across horizons, which makes forecast errors longer-lived; he concludes it can serve as a &amp;ldquo;suitable approach&amp;rdquo; / useful benchmark to model professional-forecaster expectation formation, while emphasizing the results are specific to the context of professional forecasting and may not carry over to household or firm expectations.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Summary of a forthcoming paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-question-does-the-paper-address"&gt;Q1. What question does the paper address?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The paper undertakes a formal comparison of competing non-FIRE theories of expectation formation to move toward establishing a benchmark non-FIRE model in the context of professional forecasting.&lt;/strong&gt; Ortiz motivates this with the observation that survey forecast errors are predictably correlated with real-time information — a violation of full-information rational expectations (FIRE) — but that, as noted in Reis (2020), the literature &amp;ldquo;has not yet settled on a benchmark non-FIRE model.&amp;rdquo; The paper offers &amp;ldquo;a partial answer to this question.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="q2-what-models-are-compared"&gt;Q2. What models are compared?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Four models are estimated: a noisy-information rational expectations baseline plus three biased non-FIRE models — overconfident expectations (Daniel et al., 1998), diagnostic expectations (Bordalo et al., 2020), and misspecified expectations (in the spirit of Fuster et al., 2010).&lt;/strong&gt; All are embedded in a common noisy-information environment where the latent variable is unobservable and forecasters update via a Kalman filter from a noisy private signal. Overconfidence has forecasters misperceive their signal noise as smaller than it is; diagnostic expectations introduces a representativeness distortion ϕ &amp;gt; 0 generating overreaction to recent news; misspecified expectations has forecasters treat an AR(2) process as an AR(1).&lt;/p&gt;
&lt;h3 id="q3-what-exactly-is-misspecified-expectations-in-this-paper"&gt;Q3. What exactly is &amp;ldquo;misspecified expectations&amp;rdquo; in this paper?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Misspecified expectations is a model in which the underlying state follows an AR(2) process but forecasters treat it as an AR(1), so they misperceive the true persistence of the data-generating process.&lt;/strong&gt; The author notes this version is &amp;ldquo;closest to natural expectations as modeled in Fuster et al. (2010),&amp;rdquo; with forecasters neglecting longer lags. Importantly, forecasters still understand the information structure. If the perceived persistence loads excessively onto the first lag, forecasters overextrapolate. The author flags three technical differences from Fuster et al. (2010): he does not model an AR(2) in levels with AR(1)-in-growth-rates forecasting; the perceived persistence is estimated from the data rather than defined as a function of the true autocorrelation parameters; and he does not define expectations as a weighted average of rational and naive AR(1) expectations.&lt;/p&gt;
&lt;h3 id="q4-what-data-and-sample-are-used"&gt;Q4. What data and sample are used?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The estimation uses U.S. SPF panel data from 1992Q1 to 2019Q4, yielding 77 unique forecasters and 1,520 forecaster-quarter observations for the baseline.&lt;/strong&gt; The 1992 start is chosen to avoid spanning different regimes and because the survey redefined output from GNP to GDP in 1992. The procedure requires unbroken observation sequences, so only each forecaster&amp;rsquo;s longest spell is kept, with a minimum spell length of eight quarters (because entry/exit may be non-random, per Engelberg et al., 2011). Real GDP growth is the baseline variable; 13 other macroeconomic variables are also estimated. Real-time forecast errors (not errors based on revised figures) are used, following the literature.&lt;/p&gt;
&lt;h3 id="q5-how-are-the-models-estimated-and-compared"&gt;Q5. How are the models estimated and compared?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The models are estimated via a three-step maximum likelihood procedure, and their relative fit is compared using forecast-encompassing weights (West, 2001; Harvey et al., 1998; West, 2006), supplemented by AIC and a Vuong (1989) non-nested likelihood-ratio test.&lt;/strong&gt; Step 1 estimates the fundamental process parameters (ρ₁, ρ₂, σ_w) from the macro time series and fixes them across models; step 2 estimates the signal-noise dispersion σ_v from the rational model and calibrates it across the other three; step 3 estimates each bias parameter (α_v, ϕ, ρ̂) by MLE on SPF data. This keeps fundamental and information parameters consistent across biased models so they are evaluated solely on the biases they generate, and makes identification transparent (notably, σ_v and α_v cannot be jointly identified in the overconfidence model). Encompassing weights are obtained from a constrained linear regression of realizations on model-based one-quarter-ahead forecasts, with weights summing to 1.&lt;/p&gt;
&lt;h3 id="q6-what-are-the-baseline-real-gdp-growth-results"&gt;Q6. What are the baseline real GDP growth results?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;For real GDP growth, the misspecified expectations model produces the highest log-likelihood and the largest encompassing weight, 0.539, versus 0.462 for diagnostic expectations and approximately 0.000 for both rational and overconfident expectations.&lt;/strong&gt; The fundamental process estimates imply relatively low persistence (first-order autocorrelation ρ₁ ≈ 0.434, second-order ρ₂ ≈ −0.006). The estimated bias parameters are: overconfidence ≈ 0.72, diagnosticity ≈ 0.23, and perceived persistence ρ̂ ≈ 0.564. Because ρ̂ ≈ 0.56 exceeds the estimated ρ₁ ≈ 0.43, the misspecified model implies forecasters overestimate the first-order autocorrelation and neglect the partial reversal in the second lag, generating overreactions. The signal-to-noise ratio implied by the estimated private noise dispersion is σ_w/σ_v ≈ 1.09. AIC rankings (and BIC) do not change the ordering relative to the maximized likelihoods.&lt;/p&gt;
&lt;h3 id="q7-does-the-result-hold-across-other-macroeconomic-variables"&gt;Q7. Does the result hold across other macroeconomic variables?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Across the 14 SPF macroeconomic variables, misspecified expectations provides the best in-sample fit for most series, but not all.&lt;/strong&gt; Diagnostic expectations registers larger encompassing weights for certain series — the GDP deflator (0.771), industrial production (1.000), and real residential investment (0.624). Rational expectations provides the best fit for the unemployment rate (0.745) and housing starts (in-sample). For the bulk of the remaining variables (e.g., CPI 0.859, payroll employment 1.000, real consumption 0.777, real federal spending 1.000, real GDP 0.539, real nonresidential investment 1.000, real state/local spending 1.000, 3-month Treasury bill 0.713, 10-year bond 0.746), misspecified expectations carries the largest weight. Overconfident expectations &amp;ldquo;does not yield particularly large encompassing weights for any variable.&amp;rdquo;&lt;/p&gt;
&lt;h3 id="q8-why-does-misspecified-expectations-fit-better-and-for-which-variables-especially"&gt;Q8. Why does misspecified expectations fit better, and for which variables especially?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The author finds that, among variables exhibiting overreactions, misspecified expectations tends to offer a better fit for less persistent series, because the scope for it to generate overreaction (ρ̂ − ρ₁) is greater when ρ₁ is low.&lt;/strong&gt; Unlike the alternatives, the persistence bias ρ̂ − ρ₁ can be positive or negative, allowing the model to account for both overreacting and underreacting variables; the alternative models cannot generate forecaster-level underreaction. Figure 2 plots the encompassing weight on misspecified expectations against the sum of autoregressive coefficients and suggests (with some exceptions) that less persistent variables have higher weight on misspecified expectations.&lt;/p&gt;
&lt;h3 id="q9-does-the-model-perform-out-of-sample"&gt;Q9. Does the model perform out of sample?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The misspecified expectations model also provides a better out-of-sample fit for more of the variables, estimated on 1992Q1–2005Q4 and evaluated on the latter half of the sample.&lt;/strong&gt; However, out of sample diagnostic expectations now outperforms for the GDP deflator (0.987), industrial production (0.959), payroll employment (0.813), and real federal government expenditures (0.591); overconfident expectations outperforms for the 10-year government bond (0.653); and rational expectations outperforms for housing starts (0.502) and the unemployment rate (1.000). The author cautions that these results do not imply forecasters could improve their forecasts in real time, because the MLE observations include contemporaneous individual and consensus forecast errors that are not known to forecasters when they issue forecasts; for the same reason, the results are &amp;ldquo;not inconsistent with&amp;rdquo; Eva and Winkler (2023) on the poor out-of-sample performance of error-predictability regressions.&lt;/p&gt;
&lt;h3 id="q10-could-the-apparent-advantage-of-misspecified-expectations-just-reflect-learning"&gt;Q10. Could the apparent advantage of misspecified expectations just reflect learning?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The author argues that learning about the data-generating process does not appear to drive the relative model rankings in favor of misspecified expectations, based on two exercises.&lt;/strong&gt; First, using the full pre-COVID sample (1968Q4–2019Q4) over 25-year rolling windows (three-year roll), the misspecified model outperforms diagnostic expectations in six of ten sub-samples and all models in five of ten, while diagnostic expectations wins four of ten — patterns that &amp;ldquo;do not indicate that learning over time favors misspecified expectations.&amp;rdquo; Second, splitting forecasters by &amp;ldquo;age&amp;rdquo;/tenure (a proxy for experience), misspecified expectations outperforms the others among experienced (above-median age) forecasters (encompassing weight 0.766, with overconfidence 0.234) and is dominant among inexperienced ones (1.000). The author concedes learning &amp;ldquo;is likely reflected in professional forecasts&amp;rdquo; but does not appear to drive the rankings.&lt;/p&gt;
&lt;h3 id="q11-what-additional-moments-does-misspecified-expectations-match"&gt;Q11. What additional moments does misspecified expectations match?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Beyond overall fit, the author shows in the appendix that misspecified expectations matches five features of the data — overreaction, underreaction, overshooting, persistent disagreement, and updating behavior — and is the only model generating delayed overshooting.&lt;/strong&gt; All three non-rational models generate individual-level overreaction (Bordalo et al., 2020 errors-on-revisions regression) and aggregate underreaction (Coibion-Gorodnichenko, 2015 consensus regression). But when simulating impulse responses, &amp;ldquo;only the misspecified expectations model generates a sign switch in the forecast error,&amp;rdquo; indicating delayed overshooting (Angeletos et al., 2020). The author reports &amp;ldquo;stronger evidence&amp;rdquo; favoring misspecified expectations on two further moments: it better generates persistent disagreement across horizons, and it better matches the relative weights forecasters place on priors versus news — because its bias also enters the prediction equation (not just the update equation), producing longer-lived errors.&lt;/p&gt;
&lt;h3 id="q12-what-are-the-scope-conditions-and-limitations-the-author-stresses"&gt;Q12. What are the scope conditions and limitations the author stresses?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The author emphasizes that the results are specific to the context of professional forecasting and that the relative model rankings &amp;ldquo;may be different&amp;rdquo; for household or firm expectations, or for micro-level expectations rather than aggregate forecasts.&lt;/strong&gt; He notes professional forecasters are arguably the most well-informed agents, so the literature has treated their predictions as informative about a lower bound on economy-wide information frictions and biases. The paper abstracts away from learning in the model setup and from theories that generate only underreaction. Models excluded from the comparison (e.g., imperfect memory, multi-frequency forecasting, asymmetric attention, learning) are set aside mainly because they cannot be flexibly nested into the common setting and would introduce additional parameters posing identification challenges.&lt;/p&gt;
&lt;h3 id="q13-what-does-the-author-conclude-and-recommend"&gt;Q13. What does the author conclude and recommend?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Ortiz concludes that misspecified expectations &amp;ldquo;can serve as a suitable approach&amp;rdquo; / useful benchmark to model expectation formation among professional forecasters for a variety of macroeconomic aggregates, while framing this as only &amp;ldquo;a partial answer&amp;rdquo; to the search for a non-FIRE benchmark.&lt;/strong&gt; He highlights a practical advantage: embedding this form of misspecified expectations into a quantitative model &amp;ldquo;only requires introducing two parameters into an otherwise standard model.&amp;rdquo; He also notes misspecification can arise either from a behavioral bias or because adopting parsimonious forecasting models is optimal (Branch and Evans, 2006; Pfajfar, 2013). A promising avenue for future research is whether evidence favors misspecified expectations in other settings.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key concepts&lt;/h2&gt;
&lt;dl&gt;
&lt;dt&gt;&lt;strong&gt;Full-information rational expectations (FIRE)&lt;/strong&gt;&lt;/dt&gt;
&lt;dd&gt;The benchmark in which forecast errors are uncorrelated with any information in the forecaster&amp;rsquo;s time-t information set; the orthogonality conditions it implies &amp;ldquo;tend to be violated in the data,&amp;rdquo; motivating non-FIRE models.&lt;/dd&gt;
&lt;dt&gt;&lt;strong&gt;Misspecified expectations&lt;/strong&gt;&lt;/dt&gt;
&lt;dd&gt;The paper&amp;rsquo;s focal bias — the true state follows an AR(2) process, xₜ = ρ₁xₜ₋₁ + ρ₂xₜ₋₂ + wₜ, but forecasters treat it as an AR(1), xₜ = ρ̂xₜ₋₁ + uₜ, misperceiving its persistence; forecasters retain the correct information structure. The bias enters both the predict and update equations.&lt;/dd&gt;
&lt;dt&gt;&lt;strong&gt;Persistence bias (ρ̂ − ρ₁)&lt;/strong&gt;&lt;/dt&gt;
&lt;dd&gt;The gap between perceived AR(1) persistence and true first-order autocorrelation; positive values generate overextrapolation/overreaction, negative values generate underreaction, and its overreaction scope is larger when ρ₁ is low.&lt;/dd&gt;
&lt;dt&gt;&lt;strong&gt;Overconfident expectations&lt;/strong&gt;&lt;/dt&gt;
&lt;dd&gt;Forecasters misperceive their private signal noise as smaller (σ̃_v = α_v σ_v, α_v ∈ [0,1]) than it truly is, placing excessive weight on new private information.&lt;/dd&gt;
&lt;dt&gt;&lt;strong&gt;Diagnostic expectations&lt;/strong&gt;&lt;/dt&gt;
&lt;dd&gt;A representativeness-based distortion (Bordalo et al., 2020; Gennaioli-Shleifer, 2010) in which, with diagnosticity ϕ &amp;gt; 0, forecasters overweight outcomes representative relative to a &amp;ldquo;no news&amp;rdquo; reference scenario, generating overreaction to recent news.&lt;/dd&gt;
&lt;dt&gt;&lt;strong&gt;Encompassing weight&lt;/strong&gt;&lt;/dt&gt;
&lt;dd&gt;The model-comparison metric — a weight wₖ from a constrained linear regression of realized one-quarter-ahead values on competing models&amp;rsquo; forecasts, with weights summing to one; a larger weight indicates a better-fitting model.&lt;/dd&gt;
&lt;dt&gt;&lt;strong&gt;Delayed overshooting&lt;/strong&gt;&lt;/dt&gt;
&lt;dd&gt;The Angeletos et al. (2020) pattern of initial underreaction followed by later overreaction to a shock; in this paper, only misspecified expectations produces the sign switch in the forecast-error impulse response that signals it.&lt;/dd&gt;
&lt;dt&gt;&lt;strong&gt;Overreaction vs. underreaction&lt;/strong&gt;&lt;/dt&gt;
&lt;dd&gt;Individual-level overreaction is measured via the Bordalo et al. (2020) errors-on-revisions regression; aggregate/consensus-level underreaction via the Coibion-Gorodnichenko (2015) regression — the data exhibit both, and a successful non-FIRE model must reproduce both.&lt;/dd&gt;
&lt;/dl&gt;</description></item><item><title>Mixing It Up: Inflation at Risk</title><link>https://macropaperwarehouse.com/papers/mixing-it-up-inflation-at-risk/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/mixing-it-up-inflation-at-risk/</guid><description>&lt;p&gt;This paper introduces a Bayesian Gaussian mixture density regression framework that estimates the complete forecast distribution of inflation — not just selected quantiles — and decomposes the entire risk outlook into contributions from individual economic predictors. The methodology accommodates multimodality, skewness, and fat tails without parametric restrictions, and allows construction of risk measures calibrated to the central bank&amp;rsquo;s own loss function rather than generic percentile-based measures. Applied to the recent U.S. inflation surge, the framework finds that post-pandemic inflation risk was primarily driven by the recovery of the U.S. business cycle and surging commodity prices, while adjustments in monetary policy contributed negatively — partially mitigating the increase in right-tail inflation risk — and credit spreads also offset some risk. The Gaussian mixture structure enables fast MCMC estimation and produces well-calibrated density forecasts across a range of macroeconomic variables.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Summary of a forthcoming paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-is-the-key-methodological-contribution-relative-to-existing-inflation-at-risk-approaches"&gt;Q1. What is the key methodological contribution relative to existing inflation-at-risk approaches?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Existing approaches to macroeconomic at-risk measures focus on specific quantiles of the forecast distribution — typically the 5th or 25th percentile — discarding information contained in the rest of the distribution; this paper redirects attention to the full forecast distribution while retaining the nonparametric flexibility of quantile regression.&lt;/strong&gt; The Gaussian mixture density regression estimates a conditional distribution that is a weighted mixture of Gaussians, capturing multimodality, asymmetry, and fat tails simultaneously. The key innovation is decomposability: each predictor&amp;rsquo;s contribution to any region of the forecast distribution can be quantified, enabling a driver-level accounting of what generates tail risk in any given period.&lt;/p&gt;
&lt;h3 id="q2-what-does-the-us-application-reveal-about-the-inflation-surge"&gt;Q2. What does the U.S. application reveal about the inflation surge?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The framework attributes the increase in right-tail U.S. inflation risk during 2021–2023 primarily to surging commodity prices and the recovery of the domestic business cycle, while monetary policy tightening contributed negatively — its effect partially offset the upward pressure from commodity and cycle drivers.&lt;/strong&gt; Credit spreads also partially mitigated the risk. The decomposition implies that the dominant drivers of inflation risk were supply-side and aggregate-demand factors, and that monetary policy, when it tightened, reduced the right-tail risk as intended — providing quantitative support for the interpretation that policy was reactive but directionally correct.&lt;/p&gt;
&lt;h3 id="q3-how-does-the-framework-construct-policy-relevant-risk-measures"&gt;Q3. How does the framework construct policy-relevant risk measures?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The framework allows weighting probability mass over the forecast distribution by any user-specified loss function, including asymmetric central bank preferences, yielding risk measures that integrate the full distributional information in proportion to the policymaker&amp;rsquo;s actual valuation of different inflation outcomes.&lt;/strong&gt; A central bank that penalizes above-target inflation more heavily than below-target inflation (consistent with empirical evidence on CB loss functions) would weight the upper tail more, producing a risk statistic that is higher than a symmetric measure for the same distribution. This policy-preference-aligned risk measure could have provided a more accurate signal of the urgency of the 2021–2023 inflation risk than standard percentile measures.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;inflation at risk&lt;/strong&gt; : the quantile-based or distribution-based characterization of future inflation uncertainty; extended in this paper from a single quantile to the complete forecast distribution and its risk decomposition by driver.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;density regression&lt;/strong&gt; : a regression model in which the conditional distribution of the outcome — not just its mean or a specific quantile — is the object of estimation; the paper uses a Gaussian mixture density regression to capture non-standard distributional shapes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;risk decomposition&lt;/strong&gt; : the attribution of shifts in the full forecast distribution to individual predictor variables; the paper&amp;rsquo;s key tool for identifying which economic factors drive right-tail inflation risk in any period.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;CB-preference-aligned risk measure&lt;/strong&gt; : a summary statistic constructed by weighting probability mass over the forecast distribution by the central bank&amp;rsquo;s loss function; captures asymmetric preferences and goes beyond standard percentile measures.&lt;/p&gt;</description></item><item><title>Quality Adjustment at Scale: Hedonic versus Exact Demand-Based Price Indices</title><link>https://macropaperwarehouse.com/papers/quality-adjustment-at-scale-hedonic-versus-exact-demand-based-price-indices/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://macropaperwarehouse.com/papers/quality-adjustment-at-scale-hedonic-versus-exact-demand-based-price-indices/</guid><description>&lt;h2 id="layer-1-overview"&gt;Layer 1: Overview&lt;/h2&gt;
&lt;p&gt;This paper implements and evaluates methods for constructing quality-adjusted price indices from item-level retail scanner data at scale — across hundreds of product categories, heterogeneously encoded product attributes, and rapid product turnover. Using proprietary NPD Group data covering five general merchandise categories (2014–2018) and NielsenIQ scanner data covering 50+ food product groups (2006–2015), the paper compares hedonic superlative indices (using the Erickson-Pakes EP-TV methodology and a novel machine-learning extension) against exact demand-based indices (the Feenstra 1994 lambda-ratio adjustment and the Redding-Weinstein 2020 CES Unified Price Index, CUPI). The central finding is that quality adjustment is quantitatively large: the hedonic Tornqvist index shows roughly 2.5–2.9 percentage points per year faster price decline than the matched-model Tornqvist in high-tech categories (headphones, memory cards) and 4–5 percentage points of cumulative additional disinflation relative to matched-model indices for food. The Feenstra index agrees closely in magnitude with the hedonic approach, but the CUPI is highly sensitive to the choice of common-goods rule (CGR) and in some specifications shows 20–40 percentage points more cumulative disinflation than the Feenstra, a gap the paper attributes to the CUPI&amp;rsquo;s extreme sensitivity to low-share goods. The paper establishes that hedonic superlative price indices are feasible to implement at scale, including with machine learning on sparse text descriptions, and recommends them as the practical benchmark for re-engineering official price statistics.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Summary of a published paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="in-depth"&gt;In depth&lt;/h2&gt;
&lt;h3 id="q1-what-problem-in-price-index-construction-does-the-paper-address"&gt;Q1. What problem in price index construction does the paper address?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The paper addresses the long-standing challenge of simultaneously accounting for consumer substitution and quality change due to product entry and exit in official price statistics, and shows these two corrections are now feasible to implement in real time from item-level retail transactions data.&lt;/strong&gt; Standard official statistics (CPI, PCE) use an arithmetic Laspeyres index that holds spending weights fixed — the Boskin Commission documented this overstates the true cost of living due to substitution bias. Scanner data permit superlative indices (Tornqvist, Fisher) that correct for substitution at the item level, but further correcting for quality change requires dealing with the millions of products that enter and exit the market each quarter. The paper&amp;rsquo;s contribution is to show that both corrections can be combined at scale.&lt;/p&gt;
&lt;h3 id="q2-what-data-infrastructure-does-the-paper-use-and-what-does-it-reveal-about-product-turnover"&gt;Q2. What data infrastructure does the paper use, and what does it reveal about product turnover?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The NPD Group data cover five product groups with quarterly turnover rates of 4.5–13.5 percent per quarter (both entry and exit), and exhibit a characteristic life-cycle pattern: prices peak at entry and decline steadily thereafter while market shares follow a hump shape — rising as products gain distribution and then declining as newer products displace them.&lt;/strong&gt; Memory cards, for example, show approximately a 50 percent price decline over their life cycle and a 200 percent increase in market share in the first year after entry. These interrelated price-quantity dynamics mean that any matched-model index that ignores entering and exiting goods misses substantial quality improvement. The NielsenIQ data cover over 2.6 million UPCs across 40,000 stores, with nominal food sales closely tracking BEA PCE food expenditures, validating its representativeness.&lt;/p&gt;
&lt;h3 id="q3-what-is-the-ep-tv-hedonic-method-and-why-does-the-paper-prefer-it"&gt;Q3. What is the EP-TV hedonic method and why does the paper prefer it?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The Erickson-Pakes time-varying unobservables (EP-TV) method estimates hedonic price indices from item-level transactions data using a two-step procedure: first predict log price levels from observable characteristics to recover item-level residuals, then predict log price changes from characteristics plus the lagged residual, which allows the model to track changing valuations of unobservable attributes over time.&lt;/strong&gt; This approach outperforms the simpler log-level hedonic and the EP-F (fixed unobservables) approach in model fit for price changes: EP-TV achieves R² of 0.13–0.50 versus 0.05–0.24 for log-level models across product groups. The paper extends EP-TV to the NielsenIQ data using deep neural networks to decode sparse abbreviated product descriptions (e.g., &amp;ldquo;ZR DT LN/LM CF NBP CT&amp;rdquo; for a diet soft drink), achieving out-of-sample R² of roughly 75% for price-level predictions and above 50% in-sample for price changes.&lt;/p&gt;
&lt;h3 id="q4-what-are-the-main-quantitative-findings-for-the-hedonic-indices"&gt;Q4. What are the main quantitative findings for the hedonic indices?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The EP-TV hedonic Tornqvist index indicates price declines approximately 2.9 percentage points per year faster than the matched-model Tornqvist for memory cards, 2.5 pp/year for headphones, 1.3 pp/year for boys&amp;rsquo; jeans, 0.7 pp/year for coffee makers, and 0.4 pp/year for occupational footwear; for NielsenIQ food categories, the hedonic Tornqvist is approximately 4 percentage points lower cumulatively over 2006–2015 than the matched-model Tornqvist.&lt;/strong&gt; These gaps represent the quality improvement embedded in product turnover — the fact that new memory cards at a given price embody more storage than their predecessors, new headphones better sound quality, etc. The paper also shows that the hedonic Laspeyres, which only adjusts for exiting goods (following the standard Pakes 2003 bounding result), is substantially lower than the matched-model Laspeyres, confirming that the selection bias from ignoring exiting goods in official statistics is quantitatively important.&lt;/p&gt;
&lt;h3 id="q5-how-do-the-demand-based-exact-price-indices-compare-with-the-hedonic-approach"&gt;Q5. How do the demand-based exact price indices compare with the hedonic approach?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The Feenstra (1994) lambda-ratio index, which adjusts the Sato-Vartia index for product entry and exit via the ratio of entering to exiting expenditure shares scaled by (1/(σ−1)), shows cumulative disinflation approximately 2 percentage points beyond the matched-model Sato-Vartia across all five NPD product groups and approximately 5 percentage points for NielsenIQ food, broadly comparable in magnitude to the hedonic adjustment.&lt;/strong&gt; The estimated substitution elasticities (σ) range from about 5.2 to 7.8 across NPD product groups and have a median of about 6 across food product groups, consistent with the literature. The Feenstra-hedonic agreement is a reassuring finding: two methodologically distinct approaches based on different identifying assumptions yield similar magnitudes of quality correction.&lt;/p&gt;
&lt;h3 id="q6-what-is-the-cupi-and-why-is-it-problematic"&gt;Q6. What is the CUPI and why is it problematic?&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;em&gt;The Redding-Weinstein (2020) CES Unified Price Index (CUPI) generalizes the Feenstra index by adding a taste-shock correction (S&lt;/em&gt; ratio) and a Jevons index (P&lt;/em&gt; ratio), both of which are unweighted geometric means across common goods — making the CUPI extremely sensitive to products with tiny expenditure shares, because any product with a low share is inferred to have low appeal, which the model translates into a large quality-adjustment downward.** Without a common-goods rule (CGR), the CUPI shows 30–40 percent per year price declines for high-tech goods and boys&amp;rsquo; jeans, 10–30 percentage points below the Feenstra; with a 25th-percentile market-share CGR, the CUPI is still more than 40 percentage points below the Feenstra for food in 2015 on a cumulative basis. The key concern is that very low expenditure shares for entering or exiting products can reflect search frictions, limited distribution, or clearance-rack effects rather than genuinely low consumer appeal, so the CUPI&amp;rsquo;s unweighted components may conflate these factors with quality change. The paper concludes that more research on the appropriate CGR is needed before the CUPI can be recommended for official statistics.&lt;/p&gt;
&lt;h3 id="q7-is-the-hedonic-approach-robust-to-missing-or-omitted-product-attributes"&gt;Q7. Is the hedonic approach robust to missing or omitted product attributes?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Yes — omitting key observable characteristics (memory size for memory cards, major brand dummies for apparel and footwear) from the EP-TV estimation has only minimal effect on the resulting hedonic price index, in contrast to log-level hedonic models where such omissions produce much larger distortions.&lt;/strong&gt; For memory cards, the baseline EP-TV Tornqvist index produces an average annual cumulative chained price change of −20.12%; excluding entering products whose size or speed is outside the range of continuing products changes this to −20.09%, even though roughly 50% of entering products (accounting for 25% of entering-product sales) are excluded. This robustness reflects the EP-TV design: the first-stage residual absorbs time-varying unobservable characteristics that would otherwise confound the hedonic mapping.&lt;/p&gt;
&lt;h3 id="q8-what-are-the-implications-for-official-statistics"&gt;Q8. What are the implications for official statistics?&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;The paper argues that adopting hedonic superlative price indices from real-time scanner data would produce official CPI and PCE price measures that simultaneously correct for substitution bias and quality change, resulting in meaningfully lower measured inflation in categories with high product turnover — likely understating quality-adjusted price declines in current official statistics by several percentage points per year in high-tech consumer goods and by roughly half a percentage point per year in food.&lt;/strong&gt; The practical case for adoption is that the EP-TV approach can be implemented across heterogeneously encoded data (both structured NPD attributes and unstructured NielsenIQ text), is feasible in real time with transaction data, is relatively insensitive to chain drift (full-imputation hedonic indices avoid the transitory price volatility that contaminates matched-model chained indices), and satisfies bounding properties under general conditions established by Pakes (2003). The paper thus operationalizes long-standing recommendations of the Boskin Commission (1998) that called for exactly these corrections.&lt;/p&gt;
&lt;h2 id="key-concepts"&gt;Key Concepts&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;EP-TV hedonic index&lt;/strong&gt; : the Erickson-Pakes time-varying unobservables hedonic price index, which imputes quality-adjusted price changes for entering and exiting products using a two-step regression that includes lagged residuals from a log-level hedonic to control for products whose unobservable attributes change their market valuations over time; the paper&amp;rsquo;s preferred hedonic approach.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Feenstra (1994) lambda-ratio adjustment&lt;/strong&gt; : a correction to the Sato-Vartia CES price index that accounts for product entry and exit by multiplying by (λ_{t,t-1}/λ_{t-1,t})^{1/(σ-1)}, where the lambda terms are the expenditure shares of continuing goods relative to all goods in each period; larger entry shares relative to exit shares produce a downward adjustment reflecting quality improvement from new products.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;CES Unified Price Index (CUPI)&lt;/strong&gt; : the Redding-Weinstein (2020) extension of the Feenstra index that additionally incorporates time-varying product appeal shocks via an unweighted Jevons index (P* ratio) and an unweighted expenditure-share ratio (S* ratio); the paper finds this index is highly sensitive to the common-goods rule and may overstate quality adjustment in practice due to sensitivity to low-share goods.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;common-goods rule (CGR)&lt;/strong&gt; : a threshold rule that restricts the set of goods entering the CUPI&amp;rsquo;s unweighted components to those with sufficiently large or long-duration market shares, introduced by Redding and Weinstein (2020) to limit the influence of fringe products; the paper finds CUPI results are sensitive to the CGR specification in a way that varies across product groups.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;full-imputation hedonic index&lt;/strong&gt; : a hedonic price index that uses the hedonic mapping to impute price changes for all goods (including continuing goods), rather than only entering and exiting goods; reduces chain drift relative to partial-imputation approaches because imputed prices are less volatile than observed prices.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;product turnover&lt;/strong&gt; : the quarterly entry and exit of products from the market, ranging from 4.5% to 13.5% per quarter in NPD data; the primary mechanism through which quality change is embedded in item-level scanner data and the main source of mismeasurement in matched-model price indices.&lt;/p&gt;</description></item></channel></rss>