Macro Paper Warehouse
Published Classic [Brookings Papers on Economic Activity] doi:10.1353/eca.2012.0005 Vol. 2012, No. 1, pp. 81-135

Disentangling the Channels of the 2007-09 Recession

James H. Stock

Mark W. Watson

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Did the 2007-09 recession need a new explanation, or was it familiar shocks writ large? Fitting a small set of common factors to 200 American economic series back to 1959, this 2012 paper finds that a version estimated only on pre-crisis data tracks the downturn well once the realized factors are fed in, with no sign of a missing financial-crisis factor. The recession looks like unusually large draws of the same forces that drove earlier postwar recessions. It matters for how models are built, though the authors stress their individual named shocks are weakly identified and correlated, so the breakdown should not be read too literally.

What this paper finds — and why it matters

This 2012 Brookings Papers on Economic Activity paper by James Stock and Mark Watson asks whether the unusual severity of the 2007-09 recession and the weakness of the recovery that followed required some new economic mechanism – a “financial crisis factor” – or whether the episode can be understood as a larger, historically typical draw of familiar shocks propagating through the economy in its usual way. They address this with a high-dimensional dynamic factor model (DFM) fit to 200 quarterly U.S. macroeconomic series (1959Q1-2011Q2, 132 of them used to estimate the factors), extracting six common factors – a choice consistent with Bai-Ng (2002) information-criteria tests (which themselves gave mixed guidance, ranging from 3-4 to 12 factors depending on the criterion), visual inspection of a scree plot, and the number of distinct structural shocks examined – with factor loadings and dynamics estimated by principal components over the pre-crisis 1959Q1-2007Q3 subsample (the “old” factors) and then extended into the crisis period using a four-lag factor VAR fit over the full 1959Q1-2011Q2 sample. To identify structural shocks (oil, monetary policy, productivity, uncertainty, liquidity/financial risk, and fiscal), the paper develops what it treats as its most-cited contribution: a general “external-instruments” (proxy SVAR) identification strategy in which a shock is recovered as the population regression of an outside instrument – correlated with that shock and, under an exogeneity condition, uncorrelated with the other structural shocks – onto the reduced-form factor innovations, applied one instrument at a time across 18 candidate instruments spanning the six shock categories. The paper’s three headline findings are: (1) a DFM estimated only on pre-2007Q4 data, when fed the actual post-2007Q4 realizations of the “old” factors, tracks the 2007-09 downturn well, and formal tests find little evidence of a break in the factor loadings (rejected at the 5 percent level for only 15 percent of series against full-sample loadings, 12 percent against 1984Q1-2007Q3 loadings) or of a missing factor (a subsampling test gives p=0.59 at an 8-quarter post-2007Q3 horizon and p=0.90 at 15 quarters) – so the crisis reflects unusually large innovations to the same six factors that drove earlier postwar recessions, not a new factor or new dynamics; (2) those innovations were concentrated in financial and uncertainty-related series – the TED spread, VIX, and housing starts saw roughly 8-standard-deviation innovations in 2008Q4, while oil prices moved 1.7 standard deviations in 2007Q1 and 3.4 in 2008Q2 – and a composite uncertainty-liquidity shock (the first principal component of five estimated uncertainty and financial-risk shocks) is attributed roughly two-thirds of the 2007Q4-2009Q2 decline: 6.2 of a 9.2-percentage-point GDP shortfall and 4.5 of a 7.3-point employment shortfall, with oil and monetary-policy shocks contributing moderately and productivity and fiscal shocks contributing little; and (3) the slow recovery is mostly a story of secular trend decline rather than an unusually weak cyclical response – of the roughly 3-point shortfall in post-trough GDP growth relative to pre-1984 recoveries, about four-fifths (2.4 points) reflects slower trend growth rather than a weak cyclical rebound, and for employment the larger part of a roughly 6-point shortfall (3.3 points of trend versus 2.7 points cyclical) is trend-driven, traced to a decades-long decline in trend employment growth (trend GDP growth itself fell by roughly 1.2 percentage points from 1965 to 2005) attributed mainly to the plateauing of female labor-force participation and an aging-driven decline in male participation. Two scope conditions travel with these results and qualify how they should be used: the external-instrument shock estimates are frequently weakly identified (first-stage F-statistics below 5 in 10 of the 18 instrument cases, below 10 in all but 3) and are often substantially correlated with one another both within and across nominally distinct shock categories – most strikingly a -0.93 correlation between two ostensibly separate fiscal shocks whose underlying instruments correlate only -0.06 – which the authors say precludes treating the individual named shocks as a clean, mutually orthogonal decomposition; and because the DFM is linear and does not impose a zero lower bound, its monetary-policy shock estimates counterfactually permit negative interest rates and do not capture unconventional monetary policy, so the finding that monetary policy was “neutral or contractionary” during the crisis and recovery must be read subject to that caveat rather than as a claim about the effect of the actual, unconventional monetary-policy response.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What question does the paper ask, and what are its three headline answers?

The paper asks whether the 2007-09 recession’s unusual severity and the subsequent slow recovery required a genuinely new economic mechanism – a “financial crisis factor” – or whether the episode can be explained as a larger version of shocks the economy had experienced before, propagating in a historically predictable way. It reaches three substantive conclusions (stated in the abstract and Section V, pp. 81-82, 129): (1) the recession resulted from shocks that were larger versions of previously experienced shocks, to which the economy responded in a historically predictable way, so no new factor or new dynamics are needed to fit 2007Q4-2011; (2) these shocks emanated primarily, though not exclusively, from financial upheaval and heightened uncertainty; and (3) while the slow recovery partly reflects the nature and magnitude of the recession’s shocks, most of the slow recovery in employment – and nearly all of it in output – is due to a secular slowdown in trend labor-force growth, which also explains the “jobless recoveries” of 2001 and 2007-09 and implies future recoveries will likely be jobless too. A fourth, methodological conclusion is that ignoring such trend changes imparts low-frequency movements to structural-VAR/DFM errors that “seems likely to introduce subtle problems” into structural analysis (Section V, p. 129).

Q2. What data and modeling framework does the paper use?

The authors fit a dynamic factor model (DFM) to 200 quarterly U.S. macroeconomic series spanning 1959Q1-2011Q2 (data vintage November 2011), organized into 13 categories, with 132 disaggregated series used to estimate the factors and the remainder used as aggregates or outcomes (Section I). All series are transformed to stationarity (growth rates or differences) and standardized. The model has an observation equation Xt = ΛFt + et relating the panel Xt to a small vector of latent factors Ft via loadings Λ, plus idiosyncratic disturbances et, and a factor VAR Φ(L)Ft = ηt describing the factors’ own dynamics, with structural shocks εt related to the reduced-form factor innovations ηt by ηt = Hεt.

Q3. How many factors does the model use, and how are “old” (pre-crisis) factors constructed for comparison with the crisis period?

The DFM uses six factors, a choice the authors describe as consistent with Bai-Ng (2002) information-criterion tests, visual inspection of a scree plot, and the number of distinct structural shocks examined later in the paper – though the underlying statistical guidance was mixed: the Bai-Ng ICP1/ICP2 criteria selected three or four factors depending on the sample, ICP3 selected 12, and the scree plot dropped sharply to four or five factors before declining slowly (Section I, fn. 9). To compare pre- and post-crisis behavior, the factor loadings Λ and factor dynamics are estimated by principal components over the pre-crisis 1959Q1-2007Q3 subsample only, producing what the paper calls the “old” factors; these old factors are then extended through 2007Q4-2011Q2 by applying the pre-crisis loadings to post-2007Q4 data. Separately, the factor VAR used for the structural shock analysis has four lags and is estimated over the full 1959Q1-2011Q2 sample (Section III.C, fn. 7). The authors note their main results are not very sensitive to varying the number of factors over a reasonable range.

Q4. Was a new “financial crisis” factor needed to explain the recession, or new dynamics?

No: a DFM estimated purely on pre-2007Q4 data, when given the actual post-2007Q4 realizations of the “old” factors, tracks the 2007-09 downturn well, and two formal tests find little evidence of either a break in the factor loadings or a missing factor. First, Andrews (2003) end-of-sample instability tests on the factor loadings reject the null of stability at the 5 percent level for only 15 percent of all series when post-2007 loadings are compared against the full 1959Q1-2007Q3 sample, and for 12 percent when compared against the 1984Q1-2007Q3 subsample; the residual rejections concentrate in a handful of series (commodity/materials producer prices, duration of unemployment, monetary aggregates) rather than being widespread (Section II, Table 3). Second, a test for a missing factor – based on the ratio of the largest eigenvalue to the sum of all eigenvalues of the post-2007Q4 idiosyncratic disturbances’ second-moment matrix – finds this concentration ratio is actually smaller than in the 1960 and 1973 recessions; a subsampling test of whether the ratio matches its pre-2007Q4 mean gives p=0.59 over an 8-quarter window and p=0.90 over 15 quarters, providing no evidence of a missing factor (Section II, p. 100). What changed instead was the variance of the factor innovations, not the factor dynamics: Table 5 shows the factor component of oil prices moved 1.7 standard deviations in 2007Q1 and 3.4 in 2008Q2, while the TED spread, VIX, and housing starts moved to roughly 8-standard-deviation innovations in 2008Q4, with real-variable innovations comparatively moderate throughout 2007Q4-2009Q1.

Q5. What is the paper’s external-instruments (“proxy SVAR”) identification strategy, and why is it significant?

The paper identifies each structural shock as the predicted value from a population regression of an external instrument Zt on the reduced-form factor innovations ηt, a general identification strategy the authors present as their most influential contribution and that later work (Mertens-Ravn 2013; Gertler-Karadi 2015) builds on directly. For a single instrument identifying one shock ε1t, the strategy imposes three conditions (Section III.A, Eqs. 6-8, pp. 104-108): (i) relevance – E[ε1t Zt] = α ≠ 0, the instrument correlates with the shock of interest; (ii) exogeneity – E[εjt Zt] = 0 for all other structural shocks j; and (iii) the structural shocks are mutually uncorrelated (Σεε diagonal). An “external” instrument is a variable used for identification that is not itself one of the model’s factors (in a DFM) or included in the VAR (in a SVAR), distinguishing it from “internal” instruments built as linear combinations of variables already in the system. Each identified shock is normalized to have a unit impact on a chosen variable (e.g., an oil shock normalized to raise log oil prices by one unit). The paper gives the identification argument for the just-identified single-instrument case, since each shock here is estimated one instrument at a time; technical extensions to multiple instruments and weak/strong-instrument inference are developed in a companion paper, Montiel Olea, Stock, and Watson (2012). The authors note (fn. 15, p. 105) that this approach was first presented in Stock and Watson (2008) and developed independently by Mertens and Ravn (2012), that the broader idea of using constructed exogenous shocks as SVAR instruments dates at least to Hamilton (2003) (see also Kilian 2008a, 2008b), and that it generalizes the Romer-Romer (1989) practice of constructing exogenous shock components from outside the VAR.

Q6. Which six shocks does the paper identify, and with what instruments?

Six structural shocks are each identified one instrument at a time from a total of 18 candidate external instruments (Section III.B): oil price (Hamilton 2003 net oil price increase; Kilian 2008a OPEC production shortfall; Ramey-Vine 2010); monetary policy (Romer-Romer 2004; Smets-Wouters 2007 DSGE policy shock as recomputed by King-Watson 2012; Sims-Zha 2006; Gürkaynak-Sack-Swanson 2005 target factor); productivity (Galí 1999 long-run restriction, an internal instrument; Basu-Fernald-Kimball/Fernald TFP; Smets-Wouters 2007); uncertainty (Bloom 2009 VIX innovation; Baker-Bloom-Davis 2012 policy-uncertainty index); liquidity and financial risk (TED spread; Gilchrist-Zakrajšek excess bond premium; Bassett et al. 2011 bank loan supply); and fiscal policy (Ramey 2011a spending news; Fisher-Peters 2010 military-contractor stock returns; Romer-Romer 2010 tax shock).

Q7. How reliable are these instrument-identified shocks – are they weakly identified, and are they mutually uncorrelated as the theory assumes?

The instruments are frequently weak, and the resulting shocks are often substantially correlated with one another rather than mutually orthogonal, which the authors say undermines a clean decomposition. First-stage F-statistics are small: below 5 in 10 of the 18 instrument cases and below 10 in all but 3, echoing Kilian’s (2008b) observation that oil-shock instruments in particular tend to be weak; the paper notes this implies considerable sampling uncertainty that it does not attempt to quantify (p. 113). Second, because the external-instrument approach does not impose condition (iii) – mutual uncorrelatedness – in estimation, it can be checked empirically, and often fails: within a shock category, different instruments meant to identify “the same” shock are often only weakly correlated (the Hamilton- and Kilian-identified oil shocks correlate about 0.15, while Kilian and Ramey-Vine correlate about 0.60), and most strikingly, the Fisher-Peters (military spending) and Romer-Romer (tax) fiscal shocks correlate at -0.93 even though the underlying instruments themselves correlate at only -0.06 (pp. 113-116). Across categories, monetary and fiscal shocks have mean absolute correlation of about 0.51, and uncertainty and liquidity/financial-risk shocks about 0.73 – comparable in magnitude to within-category correlations, with a notably high 0.66 correlation between the Baker-Bloom-Davis policy-uncertainty shock and the Gilchrist-Zakrajšek excess-bond-premium shock (p. 116). The authors liken this to Rudebusch’s (1998) critique of SVAR-identified monetary shocks, conclude that the high correlations “preclude a compelling decomposition” among the five estimated uncertainty and liquidity/financial-risk shocks, and instead work with a composite – the first and second principal components of those five shocks (Section III.D, p. 117).

Q8. Quantitatively, which shocks explain the 2007-09 recession, and what qualification applies to the monetary-shock estimates?

By 2011Q2, GDP remained 8.2 percent below its trend value extrapolated from the 2007Q4 peak, of which 6.0 percentage points is attributed to the estimated factor shocks, with financial-risk and uncertainty shocks contributing the largest negative share. A composite uncertainty-liquidity shock (the first principal component of the five estimated uncertainty and liquidity shocks) accounts for roughly two-thirds of the recession’s decline over 2007Q4-2009Q2: 6.2 of 9.2 percentage points of the GDP decline and 4.5 of 7.3 points of the employment decline (Section III.D, pp. 116-119, Table 8). Oil and monetary-policy shocks make moderate negative contributions, with oil’s effect concentrated especially before the financial crisis, while productivity and fiscal shocks are estimated to have contributed little over this period (p. 119). A hard scope condition applies to the monetary-policy numbers: the DFM is linear and does not impose a zero lower bound, so it predicts a counterfactually negative federal funds rate over 2008-2011 and does not capture quantitative easing. The Romer-Romer (2004), Smets-Wouters (2007), and Gürkaynak-Sack-Swanson (2005) monetary shocks accordingly “indicate that monetary policy was neutral or contractionary during the recession and recovery, which is consistent with the model being linear and not incorporating a zero lower bound” (Section III.C, p. 119); the authors state plainly that “our identification scheme does not capture the unconventional monetary policy of the crisis and recovery,” so this finding should not be read as evidence about the actual effectiveness of crisis-era monetary policy.

Q9. Why was the recovery so much slower than past recoveries, and how much of that is cyclical versus a trend phenomenon?

Most of the slow recovery is a trend, not a cyclical, phenomenon: about four-fifths of the post-trough GDP growth shortfall and the larger part of the employment growth shortfall trace to declining trend growth rather than to a weaker-than-usual cyclical rebound. The recovery was unusually weak in absolute terms: 8 quarters after the 2009Q2 trough, GDP had grown only 5.0 percent (versus a 1960-2001 average of 9.2 percent) and employment only 0.6 percent (versus 4.0 percent), a starker contrast still with the robust 1960-82 recoveries (8-quarter GDP +11.0 percent, employment +5.9 percent) (Section IV, p. 121). Decomposing predicted post-2009Q2 growth into trend and cyclical components (Table 9), predicted GDP growth is about 3.0 percentage points below the pre-1984 average, of which 2.4 points (about four-fifths) is due to slower trend growth; predicted employment growth is 6.0 points below the pre-1984 average, of which 2.7 points is cyclical but the larger share, 3.3 points, is due to slower trend employment growth (Section IV.B, p. 124). That trend decline is itself long-running: trend GDP growth fell about 1.2 percentage points from 1965 (3.7 percent) to 2005 (2.5 percent) (Table 10), almost entirely due to declining trend employment growth, split roughly equally between a falling employment-population ratio and slower population growth; the falling employment-population ratio in turn traces mainly to the plateauing of female labor-force participation (Goldin 2006) and an aging-driven decline in male participation (Aaronson et al. 2006; Fallick-Pingle 2008) (Section IV.C). Notably, the DFM actually predicts a slower employment recovery than what occurred – the realized recovery was somewhat stronger than the model’s trough-conditioned forecast – and the authors say they cannot tell, with the limited post-crisis data available, whether this reflects the effectiveness of the extraordinary monetary and fiscal policy response or simple parameter instability (Section IV.A, p. 124; Section V, p. 129).

Key terms in this paper

Definitions below follow the paper's own usage.

Dynamic factor model (DFM)
the paper's core statistical framework, in which a large panel of 200 macroeconomic series Xt is modeled as driven by a small number (here, six) of common latent factors Ft plus series-specific idiosyncratic disturbances, Xt = ΛFt + et, with the factors' own dynamics governed by a VAR; used both to summarize broad macroeconomic comovement and, via the factor innovations, as the basis for structural shock identification.
External instrument / proxy SVAR identification
the identification strategy for recovering a structural shock from the reduced-form factor (or VAR) innovations by regressing an outside variable ("external instrument") -- one not itself a factor in the DFM or a variable in the VAR -- on those innovations; the instrument must be correlated with the shock of interest (relevance) and uncorrelated with all other structural shocks (exogeneity), and the resulting shock is normalized to have a unit impact on a chosen reference variable.
"Old" factors
the six DFM factors whose loadings and dynamics are estimated by principal components using only pre-crisis data (1959Q1-2007Q3), then extended forward through 2011Q2 by applying those fixed pre-crisis loadings to post-2007Q4 data; used throughout the paper as the basis for testing whether the crisis period is consistent with pre-crisis factor structure, as distinct from the four-lag factor VAR (estimated over the full sample) used for the structural shock analysis.
Relevance and exogeneity conditions
the two identifying assumptions underlying the paper's external-instruments approach -- relevance, that the instrument Zt is correlated with the shock of interest (E[ε1t Zt] ≠ 0), and exogeneity, that it is uncorrelated with every other structural shock (E[εjt Zt] = 0 for j ≠ 1) -- which the paper treats as testable in aggregate (via the correlation structure of the resulting estimated shocks) rather than simply assumed.
Composite uncertainty-liquidity shock
because the paper's separately estimated uncertainty and liquidity/financial-risk shocks turn out to be substantially cross-correlated rather than mutually orthogonal, the authors construct a composite measure -- the first (and second) principal component of the five estimated shocks in these two categories -- and use this composite, rather than the individual named shocks, to attribute roughly two-thirds of the 2007-09 GDP and employment decline to financial and uncertainty factors.
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.