A Model of the Data Economy
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
If data is simply information a firm accumulates to predict a moving target, what kind of economy does that make? The paper's stated contribution is the framework rather than any single prediction. In the long run such an economy behaves like an ordinary one accumulating machines, with diminishing returns; in the short run it shows increasing returns from a feedback loop, S-shaped firm growth with possible traps, early losses, and data bartered for goods at no money price, so measured output can understate activity. Choices are efficient unless data is also used to poach rivals' customers. Why it matters: it gives statisticians and modellers a common starting point.
What this paper finds — and why it matters
The paper builds a tractable dynamic general-equilibrium model in which data is information that firms accumulate to forecast a moving target, is generated as a by-product of production, is tradeable and (partially) non-rival, and depreciates over time. Its primary stated contribution is the framework itself — a tool that maps to many observable macro and finance measures and can be “calibrated and estimated like its old-economy DSGE counterpart” — rather than any single prediction. The model implies that a data economy resembles a Solow capital-accumulation economy with diminishing returns in the long run, but exhibits new short-run behavior: increasing returns from a “data feedback loop,” S-shaped firm growth with possible growth traps, initial losses, and the barter of data for goods. Because data can be exchanged at a zero monetary price, the authors argue conventional GDP can understate economic activity, and the model delivers a data-specific depreciation formula and a recursive value function for valuing data. In the baseline economy equilibrium choices are socially efficient despite non-rivalry, increasing returns, and data being a by-product; but once data is also used for “business stealing,” the model generates over-investment in capital and excessive trade in data. These are properties of a stylized model under specific assumptions (competitive price-taking firms, normally distributed shocks, a quadratic quality loss, an aggregate state variable, and an exogenous rental rate of capital), not calibrated empirical estimates.
Summary of a forthcoming paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What is the paper’s core contribution?
The primary contribution is a tractable modeling tool for the data economy, not a specific empirical prediction. The authors write that “the primary contribution of the paper is not the particular predictions we explore, but our model as a tool” — a way “to value data, measure its effects and to think clearly about the aggregate economic consequences of data accumulation.” Because the resulting framework “maps to many observable macro and finance measures, it can be calibrated and estimated like its old-economy DSGE counterpart.” Several predictions, the authors note, are “unsurprising, given the model assumptions”; their value is that the realism of the predictions supports the framework’s usefulness.
Q2. How is “data” modeled, and why does that modeling choice matter?
Data is modeled as information — noisy signals about a persistent, drifting state that firms use to forecast the optimal production technique — rather than as a direct addition to productivity. Each firm’s good has quality that declines in the squared distance between its chosen technique and an optimal technique with a persistent AR(1) component and a transitory component. Data points are noisy signals about that state; the number of data points a firm obtains is a by-product of its physical output (n = z·kᵃ). The authors emphasize three differences from Solow’s capital: data is used for forecasting, data is a by-product of economic activity, and data is at least partially non-rival. This information-based formulation is what lets the model capture “the tension between diminishing and increasing returns that is central to data valuation,” a tension the authors argue models where data contributes directly to productivity cannot represent.
Q3. How does the model say data should be depreciated?
The model delivers an explicit data-depreciation formula, equal to 1 − 1/(ρ² + σ²_ε·Ω_t), which depends on the persistence and volatility of the forecasted environment and on how much the firm already knows. A larger fraction of the “stock of knowledge” is lost to depreciation when the state changes a lot from period to period (high σ²_ε), when there is already lots of knowledge (high Ω_t), and when persistence makes the state more variable. The authors contrast this with accounting rules, which “depreciate all data like software, by amortizing it over three years” — a flat 30% per year — and argue instead that the depreciation rate of data “may vary widely” depending on whether the data forecasts something static (e.g., consumer location or tastes) or ephemeral (e.g., equity order flow). The formula follows from applying the conditional-variance operator to the state’s AR(1) law of motion (a modified Kalman filter / Riccati structure).
Q4. How does the model value data, and what does non-rivalry do to data prices?
Data is valued through a recursive Bellman value function V(Ω) with a single state variable, the “stock of knowledge” Ω (the precision of a firm’s forecast); the marginal value ∂V/∂Ω is a firm’s willingness to pay for data. The authors show the firm’s problem collapses to a deterministic recursive problem in one state variable because, under the quadratic loss and normal shocks, the conditional variance (precision) is a sufficient statistic for expected quality and all future choices. Non-rivalry adds a kink that acts “like a negative bid-ask spread”: when a firm sells data it loses only a fraction ι of what it sells, so the seller effectively earns more per unit forfeited than the buyer pays — exchanging data results in more total data being owned. This lets the authors import market-equilibrium measurement tools from finance to value data that is retained rather than transacted.
Q5. What distinctive short-run dynamics does a data economy exhibit?
In the short run, when data is scarce, a “data feedback loop” can generate increasing returns, producing S-shaped knowledge accumulation, a bifurcated firm-size distribution, and possible growth (poverty) traps. More data raises quality and profit per unit, inducing more production and transactions, which generate still more data. The authors prove (Proposition 1) that a single low-data firm entering an economy of steady-state firms has net data flow that first increases and then decreases with its stock of knowledge, tracing an S-shaped path — so firms “do not remain mid-sized for long.” Increasing returns also create growth traps: a data-poor firm earns low profits, produces little, and therefore generates little data, keeping it data-poor. For other parameter values (abundant learnable risk) accumulation is instead concave; the appendix characterizes when the feedback loop is strong enough to make inflows convex.
Q6. Why can new data-economy firms lose money yet carry high valuations?
The model shows (Proposition 2) that a firm entering with zero data cannot make positive expected profit without first making negative expected profit — it optimally produces at a loss to generate data, a form of costly “active experimentation.” Production loses money early because initial expected quality is too low to be profitable, but producing generates information that raises future quality and enables future profits — “like a bandit problem.” Because accounting rules let book value include data only if it was purchased (not self-produced), such firms display high market value relative to book value, the hallmark used to measure intangible capital. The authors liken this to Amazon, which “during its first 17 quarters as a public company … lost $2.8 billion, before turning a profit.”
Q7. What is “data barter,” and how does it connect to “missing GDP”?
Data barter arises when goods are exchanged for customer data at a zero monetary price; the authors argue that, more generally, every transaction carries a data-barter element, so conventional GDP can understate economic activity. Proposition 3 shows a firm may optimally choose positive production even when the goods price is zero, because the marginal benefit is data that can be sold next period. “Free” apps given away in exchange for data are the clean case, but the authors note “every transaction, in principle, has a data barter element,” so every firm should charge slightly less because of the value of accompanying data. They suggest the value of the knowledge asset generated — V(Ω_{i,t}) − V(Ω_{i,t−1}) — could in principle be used to fill in this missing value in GDP measurement, while stressing that detailed measurement is beyond the paper’s scope.
Q8. In the long run, can data accumulation sustain growth on its own?
No: within the model, data has diminishing returns and generates no long-run growth without innovation, and the authors prove this outcome is hard to avoid. Conceptually, information has diminishing returns because its ability to reduce forecast-error variance shrinks as beliefs become more precise, and forecast errors “can, at best, be zero.” Proposition 4 shows that sustaining any positive growth rate forever would require the quality function to approach infinity as forecast error goes to zero — i.e., perfect one-period-ahead foresight yielding infinite output — which the authors argue no standard model delivers. Proposition 5 adds that even if one accepted infinite output under perfect foresight, sustained growth would also require no fundamental (unlearnable) randomness; if some future events are fundamentally random (σ²_a > 0), data must have diminishing returns. The long-run data economy therefore “looks just like a long-run capital economy, but for different reasons.”
Q9. When can data sustain growth?
If data is used as an input into research and development, the model becomes a data-driven endogenous-growth model that can sustain growth — which the authors use to draw a measurement lesson. Allowing data to raise the size of technology advances (an R&D extension) lets constant additions to the stock of ideas sustain growth. The take-away the authors stress is not that “anything is possible” but that measurement should “distinguish between data that is used for research and development and data that is not,” just as macroeconomists separate R&D investment from ordinary capital investment.
Q10. Is a data economy efficient, and when does it fail to be?
The baseline steady-state allocation is socially efficient (Proposition 6), but adding a “business-stealing” externality makes it inefficient, producing over-investment in capital and excessive trade in data (Proposition 7). In the baseline, equilibrium capital investment and data production are efficient because there are no externalities and the constraint that data can only be produced through goods production is faced by both firms and the planner; prices reflect marginal social value. The authors caution this “doesn’t mean that data cannot cause harm.” When data is used for business stealing (a quality externality indexed by b ∈ [0,1], with b = 1 meaning data has zero net social value), firms neglect that their data production and sales reduce rivals’ quality, so in equilibrium “too much output is produced and too much data is traded.” The authors highlight excessive data trade as a distinctive, testable margin of the data economy’s response to externalities.
Q11. What are the model’s scope conditions, and what does the paper not claim?
The results are properties of a deliberately streamlined model, and the authors repeatedly flag which assumptions are load-bearing. Key assumptions include competitive price-taking firms, normally distributed shocks with a quadratic quality loss (so precision is a sufficient statistic), an aggregate state θ (perfect cross-firm correlation of innovations, needed for data trade to exist), an exogenous world rental rate of capital (an industry or small-open-economy interpretation), and data productivity z held fixed. The paper does not provide calibrated empirical estimates of data’s value or GDP mismeasurement; it argues its value function and depreciation formula are a “first step” and “clear way to think about” data value, not “a cookbook recipe for assigning a dollar value to data.” Extensions the authors flag as outside the main model — imperfect competition, endogenous data-savviness, product-portfolio choice, and optimal data policy — are discussed in the conclusion as future work.
Key terms in this paper
Definitions below follow the paper's own usage.
- data economy
- an economy in which "transactions of goods and services generate information, which is stored, traded and depreciates" — data here specifically means transaction-generated information used by firms to forecast future outcomes and optimize business processes.
- stock of knowledge (Ω)
- the paper's single state variable, defined as the precision (inverse posterior variance) of a firm's forecast of the persistent state θ — a summary statistic for everything a firm has learned from its data.
- data feedback loop
- the mechanism by which more data makes a firm more productive, which raises production and transactions, which generate still more data — "more data begets more data," the source of short-run increasing returns.
- active experimentation
- producing (even at a loss) partly in order to generate data with future value — a bandit-like motive in which current negative-expected-value actions are taken because they produce information.
- data barter
- the exchange of goods for customer data at a zero (or reduced) monetary price; the authors argue every transaction has a barter element, which conventional price-weighted GDP misses.
- non-rivalry as a negative bid-ask spread
- because a seller loses only a fraction ι of the data it sells while the buyer gains the full amount, exchanging data increases total data owned — the seller effectively receives more per unit forfeited than the buyer pays.
- business-stealing externality
- a quality externality (index b ∈ [0,1]) in which one firm's data-driven marketing reduces other firms' ability to reach their preferred customers; at b = 1 data has zero net social value, and any b > 0 makes the equilibrium inefficient.