Empirical exchange rate models of the seventies: Do they fit out of sample?
📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication
In brief
Can the exchange rate models economists had built during the 1970s actually predict currencies? This study runs them out of sample against the simplest possible rival, a random walk that just forecasts today's rate forever. None of them wins, at any horizon from one month to a year, for any of four dollar rates. The structural models are even handed the true future values of the variables they depend on, and still lose. The authors also stress that the random walk itself forecasts badly -- it is only that nothing else does better.
What this paper finds — and why it matters
This study compares the out-of-sample forecasting accuracy of the structural exchange rate models that had come to dominate the 1970s literature against simple time series alternatives, and finds that a random walk does at least as well as any of them. The competitors are three “asset” models – the flexible-price monetary (Frenkel-Bilson) model, the sticky-price monetary (Dornbusch-Frankel) model, and the Hooper-Morton model, which extends the latter to let the long-run real exchange rate move with unanticipated trade balance shocks – all nested in a single quasi-reduced form in relative money supplies, relative real income, the short-term interest differential, the expected long-run inflation differential and cumulated home and foreign trade balances. Against them stand six univariate time series techniques applied to raw and prefiltered data, a random walk with an estimated drift, an unconstrained vector autoregression in the same variables, the forward rate, and the spot rate itself. Estimation uses monthly, seasonally unadjusted data from March 1973, the start of the floating-rate period, through June 1981; forecasting begins in November 1976, and every model’s parameters – including its seasonal parameters – are re-estimated each period by rolling regression so that only information available at the time of each forecast is used. Horizons are one, three, six and twelve months, chosen to match the available forward rate maturities. The critical design choice is that the structural models are given the benefit of the doubt: their forecasts are built from the actual realized future values of their own explanatory variables, which “directly addresses one possible defense of these models: structural exchange rate models have explanatory power, but predict badly because their explanatory variables are themselves difficult to predict.” Even so, “none of the models achieves lower, much less significantly lower, RMSE than the random walk model at any horizon” for the dollar/mark, dollar/pound, dollar/yen or trade-weighted dollar. The result survives estimating the structural models by ordinary least squares, generalized least squares and Fair’s instrumental variables method, allowing lagged adjustment, freeing the domestic and foreign coefficients, swapping M1-B for M2 or the reserve-adjusted base, trying alternative inflation-expectations proxies, substituting price levels for monetary variables, running the models on cross-rates to sidestep unstable US money demand, starting the forecast period in November 1978, and ending it in November 1980. The authors are careful about what they can and cannot claim statistically: because formal tests of forecast-accuracy differences require restrictive assumptions, they assert only that “the other models do not perform significantly better than the random walk model,” not that the random walk is significantly better. They are equally careful that their result is not good news: “while the random walk model may be as good a predictor as any of major-country exchange rates, it does not predict well,” with root mean square errors of 1.99 percent at one month and 8.65 percent at twelve months even for the more predictable trade-weighted dollar, and 3.70 and 18.3 percent for the dollar/yen rate. Companion constrained-coefficient experiments lead them to conclude that “neither sampling error nor simultaneous equations bias can fully explain the results,” and they canvass – without settling among – structural instability from the oil shocks and policy-regime changes, inadequate modelling of expectations, failure to capture real disturbances, and misspecified money demand, describing the ranking of these explanations as “at this point speculative.”
Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.
Questions & answers
Q1. What question does the paper ask, and how does it differ from what the literature had been doing?
It asks how well the existing empirical exchange rate models fit out of sample, whereas the prior literature had evaluated them on in-sample fit. The framing is explicit: the paper “compares time series and structural models of exchange rates on the basis of their out-of-sample forecasting accuracy,” and the authors note that “when overall performance is measured by in-sample fit, regressions based on eq. (1) do reasonably well. See, for example, Frankel (1979) or Hooper and Morton (1982)” (Sections 1 and 2.1). The conclusion draws the methodological moral: “The results of our paper contrast with those of previous studies based on in-sample fit. Thus, from a methodological standpoint, our paper supports the view that out-of-sample fit is an important criterion to consider when evaluating empirical exchange rate models” (Section 6). They also locate the finding relative to earlier observations: Cornell, Mussa and Frenkel had noted that exchange rate changes are largely unpredictable, with Mussa writing that “the natural logarithm of the spot exchange rate follows approximately a random walk,” and “the present study systematically confirms this ‘stylized fact’” (fn. 1).
Q2. Which structural models are tested, and how are they nested?
Three asset models – Frenkel-Bilson, Dornbusch-Frankel and Hooper-Morton – all obtained as coefficient restrictions on a single quasi-reduced form. The general specification relates the log dollar price of foreign currency to the log ratio of US to foreign money supplies, the log ratio of US to foreign real income, the short-term interest differential, the expected long-run inflation differential, and cumulated US and foreign trade balances, plus a possibly serially correlated disturbance (Section 2.1, equation 1). “All of the models posit that, ceteris paribus, the exchange rate exhibits first-degree homogeneity in relative money supplies,” that is, a unit coefficient on relative money. The Frenkel-Bilson model “assumes purchasing power parity” and so zeroes the expected-inflation and both trade-balance coefficients; the Dornbusch-Frankel model, “which allows for slow domestic price adjustment and consequent deviations from purchasing power parity,” zeroes only the trade balances; and “none of the coefficients in eq. (1) is constrained to be zero in the Hooper-Morton model,” which “extends the Dornbusch-Frankel model to allow for changes in the long-run real exchange rate. These long-run real exchange rate changes are assumed to be correlated with unanticipated shocks to the trade balance.” The authors flag one shared restriction as a known risk: imposing that variables enter in differential form “implicitly assumes that the parameters of the domestic and foreign money demand and price adjustment equations are equal. While this parsimonious assumption is conventional in empirical applications, it is a potential source of misspecification” – though they report that freeing them “yields no gain in out-of-sample forecasting accuracy.”
Q3. How are the structural models estimated, and why does the choice of estimator matter?
By OLS, generalized least squares and Fair’s instrumental variables method, because the models’ right-hand-side variables are plausibly endogenous – and if a model were true and consistently estimated, the consistent estimator should win in a large sample. The concern is stated carefully: “Variables such as relative money supplies and relative incomes are typically treated as exogenous variables in the underlying theoretical models, but may be more realistically thought of as endogenous variables. Other variables, such as the short-term interest differential, are generally endogenous in these same theoretical models. Yet they are still treated as legitimate regressors in ordinary or generalized least squares regressions of eq. (1)” (Section 2.1). Evidence for endogeneity comes from the authors’ companion vector autoregression results and from block exogeneity tests in Glaessner. In Fair’s method, “money supplies, short-term interest rates, and expected long-term inflation rates are treated as endogenous variables.” The logical force of running all three estimators is spelled out: “if one of the models summarized by eq. (1) is true, and if its structural and serial correlation parameters can be consistently estimated with instrumental variables techniques, then such techniques will outperform inconsistent techniques in large enough samples. However, generalized least squares parameter estimates did not yield inferior forecasts in our experiments.” Specific comparisons are reported – fifth-order GLS performs worse than Cochrane-Orcutt, which also outperforms a stock-adjustment model, and Cochrane-Orcutt “does frequently yield marginally better results than Fair’s method. But this is not particularly encouraging since Cochrane-Orcutt estimates are less often of the theoretically correct sign.”
Q4. What time series models are the structural models up against?
Six univariate techniques applied to both raw and prefiltered data, a random walk with estimated drift, and an unconstrained vector autoregression. The prefilters are differencing, deseasonalizing and detrending. The six techniques are: a “long AR,” an unconstrained autoregression whose maximum lag is set by the deterministic rule M = N/log N; lag selection by the Schwartz criterion, which “provides a consistent estimate of lag length”; lag selection by the Akaike criterion, which “asymptotically produces minimum mean square prediction errors”; a weighted version of the long AR that weights observations “by powers of 0.95”; direct application of the Wiener-Kolmogorov prediction formula in the frequency domain; and a mean-absolute-deviations estimator, included because squared-error criteria “are inappropriate if, for example, exchange rates follow non-normal stable-Paretian distributions with infinite variance,” the MAD estimator being “more robust to fat-tailed distributions, and less sensitive to outlier observations” (Section 2.2). The random walk with drift estimates the drift “as the mean monthly (logarithmic) exchange rate change.” The VAR regresses each variable on lagged values of itself and all others, with a uniform lag length chosen by Parzen’s criterion (2 for the dollar/mark, 2 for the dollar/pound, 4 for the dollar/yen and 2 for the trade-weighted dollar), and a six-variable version constraining the two cumulated trade balances to enter in differenced form forecasts better than the seven-variable one. The VAR matters “since it does not restrict any variables to be exogenous a priori, and is therefore robust to some of the estimation problems that plague the structural models,” though the authors note they did not try shrinkage methods such as Litterman’s that “can sometimes markedly reduce the mean square prediction errors.”
Q5. What exactly is the forecasting design?
Monthly data from March 1973 to June 1981, with forecasts starting November 1976, all parameters re-estimated each period by rolling regression, at one, three, six and twelve month horizons. “All the competing models are estimated over a monthly data series which starts in March 1973, the beginning of the floating rate period, and extends through June 1981. Each model is initially estimated for each exchange rate using data up through the first forecasting period, November 1976. Forecasts are generated at horizons of one, three, six, and twelve months; these forecast horizons correspond to the available forward rate data. Then the data for December 1976 are added to the sample, and the parameters of each model, including the seasonal adjustment parameters, are re-estimated using rolling regressions” (Section 3). The start date is chosen for degrees of freedom, “especially [for] the profligately-parameterized vector autoregression.” Multiple horizons serve a purpose: “to see whether the structural models do better than time series models in the long run, when adjustment due to lags and/or a serially correlated error term has taken place,” and the authors note that “when lags and serial correlation are fully incorporated into the structural models, a consistently-estimated true structural model will outpredict a time series model at all horizons in a large sample.” Two robustness windows are used: a subperiod beginning November 1978, chosen partly because “that date marks a major change in U.S. intervention strategy,” and a sample truncated at November 1980, which “like November 1976, marks a U.S. presidential election.”
Q6. Why do the authors use realized rather than forecast values of the explanatory variables?
To close off the most natural defense of the structural models, at the price of making the comparison deliberately unfair in their favour. “The structural models require forecasts of their explanatory variables in order to generate forecasts of the exchange rate. To give these models the benefit of the doubt, we use actual realized values of their respective explanatory variables. This procedure directly addresses one possible defense of these models: structural exchange rate models have explanatory power, but predict badly because their explanatory variables are themselves difficult to predict” (Section 3). This is what gives the negative result its force, and the authors return to it repeatedly: “the structural models in particular fail to improve on the random walk model in spite of the fact their forecasts are based on realized values of the explanatory variables.” They also note a subtlety they do not exploit: “when the explanatory variables are endogenous they will in general be correlated with the error term in eq. (1). If available, information about this correlation could be used to construct better structural model forecasts” (fn. 10).
Q7. Why is the analysis run on logs, and what accuracy criteria are used?
Logs because the resulting statistics are unit-free and comparable across currencies and because they sidestep Jensen’s inequality; and three criteria – mean error, mean absolute error and root mean square error – because each guards against a different failure of the others. “Because we are looking at the logarithm of the exchange rate, these statistics are unit-free (they are approximately in percentage terms) and comparable across currencies. By comparing predictors on the basis of their ability to predict the logarithm of the exchange rate, we also avoid any problems arising from Jensen’s inequality. Because of Jensen’s inequality, the best predictor of the level of the dollar/mark rate might not be the best predictor of the mark/dollar rate” (Section 3). A long footnote argues that the standard dismissal of the Jensen problem rests on “an erroneous Taylor expansion,” gives a discrete counterexample where the approximation is off by two orders of magnitude, and observes that the term “is more likely to be large in data sets where an outside chance of a major intervention is incorporated into expectations” (fn. 12). Root mean square error is the principal criterion, mean absolute error is added because RMSE is inappropriate under fat tails, and mean error “provides another measure of robustness. By comparing MAE and ME we can ascertain whether a model systematically over- or underpredicts.”
Q8. What is the headline result?
No model beats the random walk in root mean square error at any horizon, for any of the four rates. “Ignoring for the present the fact that the spot rate does no worse than the forward rate, the striking feature of table 1 is that none of the models achieves lower, much less significantly lower, RMSE than the random walk model at any horizon. Although RMSE at three month horizons are not listed in table 1, they give the same result” (Section 4). Mean absolute errors, “which are generally 20-25 percent smaller than RMSE,” give “virtually the same rankings as RMSE” – and pointedly, “even the univariate technique designed to minimize mean absolute deviations fails to improve on the random walk model in out-of-sample MAE.” Mean forecast errors “are generally much smaller than the corresponding mean absolute errors, indicating that the models do not systematically over- or underpredict,” although “the structural models do tend to go systematically offtrack if no serial correlation is allowed for,” and “the random walk model is somewhat less dominant in ME than in RMSE and MAE,” particularly for the dollar/mark. “The dominance of the random walk model over the other models in RMSE remains when forecasting begins in November 1978, or alternatively if it ends in November 1980.”
Q9. Are there any exceptions, and how do the authors treat them?
There are several small ones, and the authors report each with the horizons and currencies where they do and do not hold, rather than suppressing them. Using the past twelve-month inflation differential as the expected-inflation proxy, “at one month horizons for the dollar/mark rate the Dornbusch-Frankel and Hooper-Morton models predict better than the random walk model in RMSE by 0.02 and 0.05 percent, respectively. They do worse at longer forecast horizons, though.” The Hooper-Morton model with the same proxy “also exhibits marginal improvement over the random walk model for the dollar/yen rate at six months (but not at one month or twelve months), and for the trade-weighted dollar at six months. At twelve months, the Hooper-Morton model improves by a more substantial 2 percent over the random walk model for the trade-weighted dollar” (Section 4). Among time series models, the random walk with drift “improves by about 0.5 percent at six and twelve month horizons for the dollar/pound rate, but is 0.1 percent worse at one month and does much worse at predicting the dollar/yen and trade-weighted dollar”; only the detrended long AR ever beats the random walk for the trade-weighted dollar, and “it is only 0.5 percent better at twelve months”; and only the Schwarz criterion with detrending helps for the dollar/yen, by “1.5 percent at six months and 4.8 percent at twelve months.” Running the Dornbusch-Frankel model on cross-rates, it fails for the pound/yen and pound/mark but does “0.6 percent better at six months (12.0 vs. 12.6 percent for the random walk) and 3.4 percent better at twelve months (16.0 vs. 19.4 percent) for the yen/mark cross-rate,” and the authors judge that “even this improvement is not so great as to provide a basis for asserting that money demand instability or misspecification is the main problem with the models” (Section 5). Forecast combination is also tried: “an estimated combination of all seven forecasts never improves upon the random walk model alone, but estimated linear combinations of the different forecasts taken two at a time do sometimes outperform the random walk model. However, the same combination never works for more than one exchange rate” (fn. 13).
Q10. What about the forward rate?
It does no better than the random walk except for the dollar/mark at twelve months, and the authors treat its performance as a side issue with a specific interpretation. “In table 1 the forward rate only improves on the random walk model in RMSE for the case of the dollar/mark rate at twelve month horizons. Given the joint assumptions of market efficiency and rational expectations, the relative performance of the forward rate may be interpreted as evidence on the existence of a risk premium. For example, the forward rate could predict worse than the random walk model when there is a time-varying risk premium, even if the risk premium is zero on average” (Section 4). They are explicit that this is not what the paper is for: “the interpretation of its relative performance is somewhat tangential to the main issue here, which is: How well do existing empirical exchange rate models fit out-of-sample?” (Section 1). They also note that among the many contemporaneous papers finding forward rates diverging from expected future spot rates, “Bilson (1981) … is the only author who uses an out-of-sample testing methodology. Although his model is not discussed, it too failed to outperform the random walk model at one month forecast horizons” (fn. 14).
Q11. What statistical claim do the authors actually make, given that they run no formal test?
Only the weaker, asymmetric one: that the other models are not significantly better than the random walk – not that the random walk is significantly better than they are. The limitation is disclosed before the results: “Although out-of-sample comparisons have considerable intuitive appeal, formal tests of whether these differences are statistically significant generally require restrictive assumptions” – specifically, the Granger-Newbold test “is applicable only when both forecast errors are independent and normally distributed with zero means and constant variances,” so it “can only be applied at forecast intervals greater than one month if overlapping multi-horizon forecasts are omitted” (Section 3 and fn. 11). The claim they then make is precisely bounded: “The results presented above do not answer the question of whether the random walk is significantly better than the other models in root mean square error, our primary criterion. However, given our finding that the random walk model almost invariably has the lowest root mean square error over all horizons and across all exchange rates, we can unambiguously assert that the other models do not perform significantly better than the random walk model” (Section 4).
Q12. Is the result good news for the random walk?
No – the authors are emphatic that the random walk itself forecasts poorly, and they give the numbers. “And while the random walk model may be as good a predictor as any of major-country exchange rates, it does not predict well. Even the RMSE in table 1 for the trade-weighted dollar – which as one might expect is more predictable than the bilateral rates – is 1.99 percent at one month and 8.65 percent at twelve months. The highest RMSE are for the dollar/yen rate: 3.70 percent at one month and 18.3 percent at twelve months” (Section 4). The follow-on sentence sets the agenda the next thirty years of literature took up: “One might hope to ultimately estimate a structural model which could perform substantially better than this, especially when forecasts are based on realized explanatory variable values.”
Q13. Why does the paper use seasonally unadjusted, point-sample data?
To estimate seasonal and structural parameters consistently, to avoid using information unavailable at forecast time, and to avoid the spurious serial correlation that monthly averaging induces. “All of the raw data used in this study are seasonally unadjusted, which makes it possible to estimate seasonal and structural parameters on a consistent basis. The use of seasonally adjusted data is especially likely to distort structural parameter estimates when the variables are not all adjusted by the same method.” And on real-time discipline: “Forecasts based on seasonally adjusted data, adjusted over the extended sample period or with a two-sided filter such as Census X-11, implicitly make use of information which would not have been available” (Section 2.3). Two adjustment procedures were tried – seasonal dummies and Sims’s expanding-parameterization method – with results “robust to the choice between these two techniques.” The bilateral rates are monthly point-sample data, which “have a decided advantage over monthly average data” because, per Working, “if the exchange rate follows a random walk on a mid-day to mid-day basis … a series consisting of monthly averages of mid-day rates will exhibit positive serial correlation.” The trade-weighted dollar, a weighted average of US dollar rates against the Group of Ten plus Switzerland, uses an average of daily rates for data availability and comparability with Hooper and Morton – which means, the authors note, one would expect estimated univariate models to beat the random walk for it in a large enough sample, “even if the latter model is true for point-sample data,” yet in practice almost none do.
Q14. Which explanations for the failure do the authors rule out?
Sampling error and simultaneous equations bias, on the basis of their companion constrained-coefficient experiments. The companion paper searches “a grid of coefficient constraints” built from theory – the homogeneity postulate for relative money supplies, the money demand literature for the income elasticity and interest semi-elasticity, and the purchasing power parity literature for the rate at which real exchange rate shocks are damped. The result: “Despite allowance for a serially correlated error term we find that no element of the grid yields a constrained-coefficient forecaster which improves on the random walk model for horizons under twelve months. While there is sometimes sporadic improvement at longer horizons, the overall results are similar to those presented here. … These further results appear to demonstrate that simultaneous equations bias and/or sampling error cannot be regarded as the primary rationalization of the evidence presented here” (Section 5). They also note that constraining coefficients to have the theoretically correct sign “does not, however, improve the structural model forecasts” (Section 4).
Q15. Does parameter instability explain it?
Not straightforwardly – the authors make the logical point that instability alone does not imply the random walk should win. They acknowledge the candidate: “their underlying parameters shifted over the course of the seventies due to the effects of the two oil shocks, changes in global trade patterns, or changes in policy regimes.” But then: “unless the structural model parameters themselves follow a random walk, it does not necessarily follow that parameter instability can explain why the random walk model outperforms the structural models.” They suggest Kalman filtering as a way to accommodate instability, and report that their one attempt along those lines – weighted least squares – “failed.” They also note an equivalence worth keeping in mind: “there is a sense in which parameter instability is equivalent to having omitted (perhaps binary) variables” (Section 5).
Q16. Which building blocks of the models do the authors suspect, and how confidently?
Uncovered interest parity, the inflation-expectations proxies, the goods market specification and money demand – discussed explicitly as speculation. They flag the epistemic status up front: “Any or all of the above may be a source of misspecification; the discussion below is speculative.” On uncovered interest parity: recent work “has strongly challenged” it, but “although the risk premia may be statistically significant, the evidence also suggests that the magnitudes are not large. Therefore, it is not evident that deviations from uncovered interest parity can explain the poor forecasting performance of the structural models.” On expectations: the sticky-price models “are potentially quite sensitive to the proxy used for the long-run expected inflation differential. Proxies such as long-term interest rates and past inflation rates may be grossly inadequate,” and while imposing rational-expectations cross-equation restrictions might help, “it is not clear why autoregressions should necessarily yield good expectations proxies for the exogenous variables during a period when autoregressions yield poor proxies for the endogenous variables.” On goods markets: “There is little question that purchasing power parity did not hold in the short run during the seventies,” while for long-run PPP “the evidence here is less clearcut”; the Hooper-Morton model tries to capture long-run real exchange rate movements “but it does not fit out-of-sample notably better than the other two models,” even though “temporary or permanent movements in the PPP level of the exchange rate to real shocks may be a major cause of exchange rate volatility.” On money demand: “The breakdown of empirical money demand relationships is widespread, and the phenomenon is particularly acute for U.S. money demand equations” (Section 5).
Q17. What did they do to test the money demand suspicion, and what happened?
Two direct tests – substituting price levels for monetary variables, and estimating on cross-rates to sidestep US money demand – both of which leave the conclusion intact. For the first, they use the money demand equation and its foreign counterpart to replace money supplies and real incomes with price levels. “For the Frenkel-Bilson model the theoretical values of the coefficients in eq. (1) are the same as the corresponding coefficients in eq. (4), so price levels alone remain as regressors after the substitution. The transformed model is thus a purchasing power parity equation,” while in the sticky-price models the substitution “only eliminates money supplies and real incomes.” The verdict: “The models still fail after the price levels substitution to improve on the random walk model in root mean square error,” and the accompanying table shows every modified model worse than the random walk at every horizon for every rate except the modified Hooper-Morton for the trade-weighted dollar at twelve months. For the second test, estimating Dornbusch-Frankel on cross-rates, the small yen/mark improvements described in Q9 are the only ones, and the authors judge them insufficient to make money demand the main culprit (Section 5).
Q18. What is the paper’s own summary of what remains unresolved?
That the structural models’ failure cannot be pinned on any one cause with the evidence in hand. The conclusion lists the candidates without ranking them: “Structural instability due to the oil price shocks and changes in macroeconomic policy regimes, as well as the failure of the models to adequately incorporate other real disturbances, may be important. Misspecification of the money demand functions which underpin the structural models is another likely problem, although it is true that the structural models do not predict better when price levels are substituted in for monetary variables, or when M2 or the reserve-adjusted base are used in place of M1-B. Difficulties in modeling expectations of the explanatory variables are yet another obvious source of trouble. But determining the relative importance of the possible problems listed above, or any of the others listed in section 5, is at this point speculative” (Section 6). They also note one avenue left untried: “we make no attempt to account for possible non-linearities in the underlying models” (Section 5).
Key terms in this paper
Definitions below follow the paper's own usage.
- Out-of-sample forecasting comparison
- the paper's test: each model's parameters are re-estimated every period by rolling regression using only information available at the time, forecasts are made at one, three, six and twelve month horizons, and accuracy is scored against realized rates by mean error, mean absolute error and root mean square error; the contrast with the prior literature is deliberate -- "when overall performance is measured by in-sample fit, regressions based on eq. (1) do reasonably well," and the authors conclude their results "support the view that out-of-sample fit is an important criterion to consider when evaluating empirical exchange rate models."
- Forecasts based on realized explanatory variables
- the authors' decision to feed the structural models the true future values of their own right-hand-side variables rather than forecasts of them, which "directly addresses one possible defense of these models: structural exchange rate models have explanatory power, but predict badly because their explanatory variables are themselves difficult to predict"; because the models still lose, this defense is closed off, and the paper's negative result is stronger than a like-for-like forecasting contest would have been.
- Frenkel-Bilson model
- the flexible-price monetary model, which assumes purchasing power parity holds and so drops the expected-inflation and cumulated-trade-balance terms from the paper's general quasi-reduced form; like the other two structural models it imposes first-degree homogeneity in relative money supplies.
- Dornbusch-Frankel model
- the sticky-price monetary model, which allows slow domestic price adjustment and hence short-run deviations from purchasing power parity, so the expected long-run inflation differential enters but the cumulated trade balances are still excluded; it assumes purchasing power parity only in the long run.
- Hooper-Morton model
- the sticky-price asset model that leaves every coefficient of the general specification free, extending Dornbusch-Frankel to allow changes in the long-run real exchange rate, which are assumed to be correlated with unanticipated shocks to the trade balance -- with cumulated deviations from trend entering because deviations from trend are treated as unanticipated.
- Random walk benchmark
- the benchmark that uses today's spot rate as the forecast of all future spot rates and requires no estimation at all; the authors also estimate a version with a drift parameter set to the mean monthly log change, and their conclusion is that "major-country exchange rates are well-approximated by a random walk model (without drift)" -- while noting that "as long as the exchange rate does not exactly follow a random walk, we would expect one of the estimated time series models to prevail in a large enough sample."
- The random walk does not predict well
- the authors' insistence that beating nothing is not the same as forecasting well -- "while the random walk model may be as good a predictor as any of major-country exchange rates, it does not predict well," with root mean square errors of 1.99 percent at one month and 8.65 percent at twelve months for the trade-weighted dollar and 3.70 and 18.3 percent for the dollar/yen rate.