Macro Paper Warehouse
Published Classic [International Economic Review] doi:10.1111/j.1468-2354.2011.00635.x Vol. 52, No. 2, pp. 461-487

Estimation and Inference by the Method of Projection Minimum Distance: An Application to the New Keynesian Hybrid Phillips Curve

Òscar Jordà

Sharon Kozicki

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Many macroeconomic relationships are awkward to estimate directly, even though their link to simple forecasting regressions is straightforward. This 2011 paper proposes doing it in two steps: first forecast how the economy responds horizon by horizon, then impose the theory's restrictions on those responses. It proves both steps are reliable in large samples, shows the approach coincides with a familiar alternative only when shocks are unrelated over time, and finds smaller bias in most simulation designs. Re-estimating a standard inflation equation on United States data from 1966 to 2001, it concludes the equation is misspecified. Why it matters: it can expose a bad model that other methods leave looking fine.

What this paper finds — and why it matters

This 2011 International Economic Review paper by Óscar Jordà and Sharon Kozicki proposes “projection minimum distance” (PMD), a two-step semiparametric estimator for macroeconomic models whose likelihood score function is nonlinear in the structural parameters but whose mapping from reduced-form Wold impulse-response coefficients to those parameters is linear – a class that includes forward-looking Euler equations and ARMA-type representations. In the first step, PMD estimates the Wold coefficients b_h semi-parametrically using Jordà’s (2005, 2009) local-projection regressions of y_{t+h} on lagged variables, run separately for each horizon; in the second step it exploits the model-implied linear restrictions on those coefficients via minimum distance (following Ferguson 1958) to estimate the structural parameters, with an information criterion following Hall, Inoue, Nason, and Rossi (2007) selecting the truncation horizon H by trading off identification gain from extra horizons against a many-weak-instruments cost. The paper establishes formally (Propositions 1-2 for the first-stage local-projection estimator; Lemmas 3-4 for the second-stage minimum-distance estimator) that both stages are consistent and asymptotically normal under stated rate conditions, and shows that in the simple case of i.i.d. shocks PMD coincides exactly with a GMM estimator that uses the lagged dependent variable as instrument – but the two estimators diverge once shocks are serially correlated, because the lagged dependent variable is then no longer a valid GMM instrument while the different (longer-lagged) instrument set implicitly used by local projections remains valid. Monte Carlo experiments (Section 4) show that in a univariate ARMA(1,1) design PMD converges to the true parameters even at T=50 and matches MLE’s standard errors closely by T=100-400, while MLE suffers numerical convergence difficulties in the pure-AR(1) and pure-MA(1) special cases that leave PMD unaffected; in a three-equation New Keynesian (Phillips-IS-Taylor) design calibrated from Lindé (2005) across 36 DGP variants, PMD has smaller bias than GMM in the majority of cases – both perform well when shocks are i.i.d. (though GMM’s output-gap estimates are somewhat downward biased), and both deteriorate once serial correlation and additional lagged terms are introduced, with GMM deteriorating more. Applying PMD to re-estimate Fuhrer and Olivei’s (2005) hybrid New Keynesian output and inflation Euler equations on 1966:Q1-2001:Q4 U.S. quarterly data, the paper finds PMD point estimates broadly similar to GMM, MLE, and OI-GMM but generally with similar or smaller standard errors, and uses the instability of the estimated coefficients across different choices of H together with frequent rejections of the overidentifying-restrictions test to argue that the hybrid Phillips curve is dynamically misspecified in ways that the other estimators’ point estimates alone do not reveal.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What estimation problem does PMD solve, and what are its two steps?

PMD is designed for models where the likelihood score function is nonlinear in the structural parameters but the mapping from the model’s reduced-form Wold impulse-response coefficients to those parameters is linear – a class that includes forward-looking Euler equations and ARMA representations (Abstract; Section 1, pp. 461-462). Using the prototype forward-looking Euler equation y_t = gammaE_t(y_{t+1}) + (1-gamma)y_{t-1} + eps_t, whose stable solution is an AR(1), y_t = betay_{t-1} + deltaeps_t with beta=(1-gamma)/gamma, the Wold coefficients satisfy b_h = beta^h and obey the linear moment condition b_h = gamma*b_{h+1} + (1-gamma)*b_{h-1} for h >= 1 (Eqs. 1-6, Section 2, p. 463). PMD’s two steps are: (1) estimate the b_h semi-parametrically using local projections (Jordà 2005, 2009), and (2) exploit the linear moment condition via minimum distance (Ferguson 1958) to recover gamma (p. 463).

Q2. In the simplest case, how does PMD relate to GMM, and what happens once shocks are serially correlated?

In the scalar case with i.i.d. errors, the PMD estimator gamma-hat_PMD = (b-hat_2 - 1)/(b-hat_1 - 1) is exactly equal to a GMM estimator that uses the lagged dependent variable y_{t-1} as instrument (Eq. 9, p. 464). This equivalence breaks once shocks are serially correlated (u_t = rho*u_{t-1} + eps_t): the local-projection-based estimator of b_h implicitly uses M_{t-2} as its effective instrument (Eq. 12, p. 465), which orthogonalizes the moment conditions against information omitted at t-2 and so remains valid, whereas y_{t-1} is no longer a valid GMM instrument because E(u_t | y_{t-1}) != 0 – so gamma-hat_PMD != gamma-hat_GMM (Eq. 12, p. 465). The authors describe this as the key practical advantage of PMD over GMM: local projections absorb serial correlation and dynamically misspecified model dynamics through the nuisance control X_{t,k} embedded in the projection regressions (Section 2, pp. 464-465).

Q3. Why work with the Wold representation rather than imposing structural identification directly, and what does that buy the estimator?

The paper gives two reasons for basing estimation on the reduced-form Wold representation rather than a structurally identified model: it avoids identification choices (recursive zero or sign restrictions) that are “often statistically untestable” and can introduce bias, and it remains valid when the true DGP is a VARMA rather than a VAR, since local projections consistently estimate Wold coefficients for the broad class of covariance-stationary processes including VARMA representations (pp. 465-466, citing Fernández-Villaverde et al. 2007 on structural identification bias). This is why PMD’s first stage never commits to a specific structural ordering of shocks: it estimates the b_h directly, and only the second-stage minimum-distance step maps those reduced-form coefficients into the structural parameters gamma via the model’s own restrictions.

Q4. What formal statistical theory does the paper establish for the two-stage estimator?

*For the first-stage local-projection estimator, Proposition 1 shows Â_1^h converges in probability to B_h under k²/T -> 0 and a truncation-remainder condition, and Proposition 2 establishes asymptotic normality, √(T-k-H)vec(B-hat_T - B_0) ->^d N(0, Omega_b), under the stronger rate k³/T -> 0 (Section 3.1, p. 468). For the second-stage minimum-distance estimator built from the vector generalization S_yB’ = S_x(I_r ⊗ B’)gamma (Eq. 17, p. 467), Lemma 3 establishes consistency of gamma-hat_T under an identification rank condition rank[WF_gamma] = dim(gamma), and Lemma 4 establishes asymptotic normality, √(T-H-k)(gamma-hat_T - gamma_0) ->^d N(0, Omega_gamma) with Omega_gamma = (F_gamma’WF_gamma)^{-1} (Section 3.2, pp. 469-470). The minimized distance criterion Q-hat_T, evaluated at the estimated (B-hat_T, gamma-hat_T), is asymptotically chi-square with degrees of freedom equal to the number of overidentifying moment conditions, giving a standard overidentification test (Section 3.2, p. 470).

Q5. How is the truncation horizon H (the number of local-projection horizons used) chosen, and why does this choice matter?

H is selected by minimizing an information criterion, Ĥ = argmin_h ln(|Omega-hat_gamma|) + h*ln(√T/k)/(√T/k), adapted from Hall, Inoue, Nason, and Rossi (2007), which balances the identification gain from including more horizons against a many-weak-instruments cost (Eq. 15, p. 466). The paper is explicit that including too many horizon-based moment conditions introduces the many-weak-instruments problem documented by Stock, Wright, and Yogo (2002), so despite the formal theory being derived for H fixed and given, the authors recommend parsimony in practice and let the information criterion pick H rather than fixing it arbitrarily (p. 466; Section 3.2, p. 470).

Q6. What do the Monte Carlo experiments show when comparing PMD to full-information maximum likelihood in a simple ARMA(1,1) design?

In a univariate ARMA(1,1) Monte Carlo design (y_t = rhoy_{t-1} + eps_t + thetaeps_{t-1}, four (rho,theta) combinations, T = 50, 100, 400, 1,000 replications), PMD converges quickly to the true parameter values even at the smallest sample size (T=50), and by T=100 and T=400 its standard errors are “virtually identical” to MLE’s (Table 1, Section 4.1, pp. 471-473). A notable asymmetry: MLE encountered numerical convergence difficulties in the pure-MA(1) (rho=0) and pure-AR(1) (theta=0) special cases, while PMD remained numerically stable throughout all four parameter combinations; adding extra overidentifying moment conditions (H=5 instead of H=2) did not appear to distort the PMD estimates (p. 473). The wiki source flags that the specific numerical parameter estimates and standard errors reported in Table 1 should be checked against the table directly before being cited in written work; only the qualitative findings above are reproduced here.

Q7. What do the Monte Carlo experiments show when comparing PMD to GMM in a three-equation New Keynesian model?

Across a three-equation New Keynesian model (Phillips curve, IS curve, Taylor rule) calibrated from Lindé (2005) with 36 DGP variants, T=200, and 1,000 replications, PMD estimates have smaller bias than GMM in the majority of cases (Section 4.2, pp. 471-477). The three benchmark calibrations range from mostly forward-looking (Case I: forward weights 0.7, backward weights 0.3) to mostly backward-looking (Case III: forward weights 0.3, backward weights 0.7) with increasing policy-rate persistence (beta_r = 0.09, 0.30, 1 across Cases I-III). When shocks are i.i.d. and there are no distortions to the model’s internal dynamics, both PMD and GMM produce good estimates, though GMM’s output-gap estimates tend to be somewhat downward biased; once richer dynamics are introduced (serially correlated shocks, a lagged Phillips-curve term, a lagged IS term), both estimators have difficulty, but the paper reports GMM’s distortions are larger. The authors are careful to note PMD “was not a universal panacea for every type of misspecification,” though in all cases examined it was less biased and more efficient than GMM when the model was correctly specified (pp. 475-477). As with Q6, the wiki source flags the specific median estimates and standard errors in Tables 2.1-2.3 as needing verification against the tables themselves; only the qualitative comparison is reproduced here.

Q8. What does the empirical re-estimation of Fuhrer and Olivei’s (2005) hybrid New Keynesian Phillips curve find?

Re-estimating Fuhrer and Olivei’s (2005) hybrid output and inflation Euler equations on 1966:Q1-2001:Q4 U.S. quarterly data with PMD at the information-criterion-selected horizon h=10, the paper finds PMD point estimates close to those from GMM, MLE, and OI-GMM but with standard errors that are generally similar to or smaller than the alternatives* (Table 3, Section 5, pp. 477-479). For the output Euler equation, where economic theory predicts gamma < 0, PMD’s forward-looking weight is estimated at roughly mu-hat ≈ 0.50 with the interest-rate coefficient gamma-hat of the theoretically expected negative sign at h*=10 (≈ -0.021 HP, -0.020 ST), but the paper cautions that this is not robust across H: Figure 1 shows gamma-hat is estimated with the “wrong” (positive) sign for H below about 7, and it is this H<7 instability – not a different numeric range – that the paper flags as evidence of dynamic misspecification (Section 5, p. 478). For the inflation Euler equation, where theory predicts gamma > 0, mu-hat is roughly 0.58-0.62 across the output-gap/RULC measures, and gamma-hat at h*=10 ranges from about -0.045 (HP) to -0.026 (ST) to +0.029 (RULC); the HP and segmented-trend specifications thus carry the theoretically “wrong” (negative) sign at h*=10, and Figure 1 shows this wrong sign is not confined to low H but holds “virtually… for any H” for HP and ST, while the RULC specification is instead mostly of the correct (positive) sign, turning negative only at H=3 and H=4. The paper interprets the instability of these estimates as H varies, together with the fact that overidentifying-restrictions tests reject most specifications, as diagnosing dynamic misspecification in the hybrid Phillips curve that the point estimates from GMM, MLE, or OI-GMM alone do not reveal (Section 5, p. 478). The wiki source flags all specific magnitudes in Table 3 (reproduced here with “≈” to signal approximation) as needing verification against the table before use in written work; see review_flags.

Q9. What are the paper’s own acknowledged scope limits?

The authors restrict PMD’s formal theory to models in which the Wold coefficients are linearly related to the structural parameters, note that the first-stage local-projection estimator requires an invertible Wold representation (det(B(z)) != 0 for |z| <= 1, ruling out noninvertible-MA VARMA representations), and assume homoskedastic innovations throughout (Section 1, p. 462; Section 3.1, pp. 467-468; Section 6, p. 480). The asymptotic theory also treats the truncation horizon H as fixed and given – allowing H to grow with the sample size is explicitly left to future research – and the paper flags the many-weak-instruments problem as a live practical risk whenever too many horizon-based moment conditions are used, mitigated but not eliminated by the Hall-Inoue-Nason-Rossi horizon-selection criterion (Section 3.2, p. 470; p. 466). Extension to nonlinear Wold-to-parameter mappings is explicitly left to indirect-inference methods rather than PMD itself.

Key terms in this paper

Definitions below follow the paper's own usage.

Projection minimum distance (PMD)
this paper's proposed two-step semiparametric estimator for models whose likelihood is nonlinear in the structural parameters but whose mapping from Wold impulse-response coefficients to those parameters is linear; step one estimates the Wold coefficients via local projections, step two recovers the structural parameters by minimum distance applied to the model's linear restrictions on those coefficients (Section 2, p. 463).
Wold representation / Wold coefficients (b_h, B_h)
the reduced-form moving-average representation y_t = mu + sum_{h=0}^infinity B_h*u_{t-h} of a covariance-stationary process in i.i.d. innovations u_t, used here as the object PMD estimates in its first stage rather than any structurally identified impulse response; the paper's key argument is that this reduced-form object can be linked to structural parameters without first taking a stand on structural identification (Eq. 16, Section 3, p. 466).
Local projections (as PMD's first stage)
the truncated single-horizon regressions y_{t+h} = A_1^h*y_t + ... + A_k^h*y_{t-k+1} + v_{k,t+h}, run separately for each horizon h following Jordà (2005), whose coefficient Â_1^h consistently estimates the Wold coefficient B_h (Proposition 1) and is asymptotically normal (Proposition 2) under stated rate conditions on the lag truncation k relative to the sample size T (Section 3.1, pp. 467-468).
Minimum distance (as PMD's second stage)
following Ferguson (1958), the estimator that chooses the structural parameter vector gamma to minimize the weighted quadratic distance Q-hat_T(B-hat_T; gamma) = f(B-hat_T; gamma)'*W-hat*f(B-hat_T; gamma) between the estimated Wold coefficients and the model's linear restrictions on them, with the minimized objective distributed chi-square under the null of correct overidentifying restrictions (Eq. 20, Section 3.2, pp. 468-470).
Optimal truncation horizon (H)
the number of local-projection horizons included in the second-stage minimum-distance system, selected in this paper by an information criterion adapted from Hall, Inoue, Nason, and Rossi (2007) that trades off the identification gain from additional horizons against the many-weak-instruments cost (Stock, Wright, and Yogo 2002) of including too many; treated as fixed in the paper's formal asymptotic theory even though it is chosen data-dependently in practice (Eq. 15, p. 466; Section 3.2, p. 470).
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.