Macro Paper Warehouse
Published Classic [Journal of Economic Literature] doi:10.1257/jel.51.4.1063 Vol. 51, No. 4, pp. 1063-1119

Exchange Rate Predictability

Barbara Rossi — ICREA-Universitat Pompeu Fabra, Barcelona GSE, and CREI

📄 Summarized from the full manuscript · Human-reviewed for faithfulness before publication

In brief

Thirty years after Meese and Rogoff showed that a random walk beats economic models at forecasting exchange rates, many papers claim to have solved the puzzle. This survey reviews them and then re-runs the horse race on up-to-date data for 19 countries. The honest answer is that it depends on almost every choice the forecaster makes. Taylor-rule fundamentals work at short horizons and panel monetary models at long ones, for some countries, sometimes. Predictive ability comes and goes over time, and no predictor beats the random walk everywhere.

What this paper finds — and why it matters

This article asks whether anything forecasts exchange rates and, if so, what – and answers that it depends, systematically, on five choices the forecaster makes. Since Meese and Rogoff (1983a,b, 1988) it has been known that a simple random walk usually beats economic models out of sample, and the survey opens by insisting on what that does not mean: the finding “should not be interpreted as a validation of the efficient market hypothesis,” which “does not mean that exchange rates are unrelated to economic fundamentals, nor that exchange rates should fluctuate randomly around their past values. Hence the puzzle.” The first half reviews the predictors, models, data choices and forecast evaluation methods proposed over the preceding decade, with deliberately limited scope: a reduced-form, out-of-sample focus at monthly and quarterly frequencies, on nominal rather than real exchange rates, because structural models “are too stylized to be literally taken to the data and successfully used for forecasting.” The second half re-runs the horse race on up-to-date data for a set of 19 countries against the US dollar, using a single common data transformation (seasonal adjustment), a rolling estimation window equal to half the sample, short (one month or quarter) and long (four-year) horizons, and a battery of tests: a Granger-causality test robust to instabilities, the Diebold-Mariano-West and Clark-West out-of-sample tests against a random walk without drift, and two tests that check robustness to the forecast sample and to the estimation window size. The literature’s consensus negatives are confirmed – purchasing power parity and monetary fundamentals do not forecast at horizons under about two to three years, and non-linear models are the least successful – and the one positive consensus is that Taylor-rule and net foreign asset fundamentals have more out-of-sample content than traditional interest rate, price, output and money differentials, with the disagreement being over how much of the puzzle they resolve. In the author’s own exercise most traditional predictors show in-sample predictive ability while almost none survives out of sample; Taylor-rule fundamentals are strongly significant under Clark-West for four countries and marginally for five more at the one-month horizon, and for none at four years, but are not significant at all under Diebold-Mariano-West – a discrepancy the author reads as the difference between judging a model “in population” and judging it “at the actual estimated parameter values.” A model loading up on real interest rates, the trade balance and the current account has the strongest in-sample fit of any considered and “extremely poor” out-of-sample performance, “possibly due to the large number of parameters to estimate.” A panel monetary model beats the random walk at long horizons for four countries and a Bayesian model averaging specification beats it nowhere. The sharpest result is about instability: the Fluctuation test finds predictive ability for monetary fundamentals only occasionally and briefly – never for Canada, for France in the late 2000s, Japan in the mid-2000s, the UK in 2009 – and window-size analysis shows that only small estimation windows detect German predictability (concentrated in the late 1980s) while only large ones detect Japan’s (emerging in the late 2000s). The author’s conclusion is therefore that although some predictors work somewhere sometimes, “none of the predictors, models, or tests systematically find empirical support for superior exchange rate forecasting ability of a predictor for all models, countries and time periods,” so “Meese and Rogoff’s (1983a,b) finding does not seem to be entirely and convincingly overturned” – and that the time-varying, occasional nature of predictability is itself a new puzzle, since existing time-varying-parameter models fail to capture it.

Summary of a classic paper, AI-assisted and human-reviewed. See the linked original for the authoritative claims and full conditions.


Questions & answers

Q1. What question does the article pose, and why is the existing literature hard to read as a whole?

It asks “does anything forecast exchange rates, and if so, which variables?”, and the literature is hard to read because papers differ simultaneously in predictors, models, estimation methods, samples, databases and evaluation tests, so results are not directly comparable. The author lists the choices a forecaster faces – “Which predictors to use? Which forecast horizon to predict? Which model to estimate? Which data frequency? Which sample?” – and states two distinct problems. First, coverage: one goal is “to provide guidance to researchers on navigating the existing literature as well as to provide a reliable overview of established findings.” Second, comparability: “existing papers rely on different predictors, tests, samples or databases; it is possible that such predictors might have lost their forecasting ability, or may not be robust to other databases or samples. In addition, while a predictor might be successful according to a metric/test, it may not be so according to a different one” (Section 1). That is the motivation for pairing the survey with a single unified empirical exercise rather than only tabulating other people’s results.

Q2. What is the paper’s answer, stated as precisely as the paper states it?

“It depends” – on predictor, horizon, sample period, model and evaluation method – with predictability most apparent under three specific conditions. The abstract’s formulation: “Predictability is most apparent when one or more of the following hold: the predictors are Taylor rule or net foreign assets, the model is linear, and a small number of parameters are estimated. The toughest benchmark is the random walk without drift.” The introduction adds two further characterizations: “There is some instability over samples for all models, and there is no systematic pattern across models in terms of which horizons or which sample periods the models predict best,” and “among the negative findings on which the literature has reached a consensus, typically, PPP and monetary models have no success at short (less than 2-3 years) horizons.”

Q3. What are the five general conclusions?

One on predictors, one on model specification, one on data handling, one on evaluation choices, and one on instability. (1) “The degree of success in forecasting exchange rates out-of-sample does depend on the choice of the predictor. Although there is disagreement in the literature, overall the empirical evidence is not favorable to traditional economic predictors (such as interest rates, prices, output and money). Instead, Taylor-rule fundamentals and net foreign asset positions have promising out-of-sample forecasting ability,” with the disagreement concerning “the degree to which they can resolve the Meese and Rogoff puzzle.” (2) “Among the model specifications considered in the literature, the most successful are linear ones,” and “typically, in single-equation linear models, the predictor choice matters more than the model specification itself.” (3) “Data transformations (such as de-trending, filtering and seasonal adjustment) may substantially affect predictive ability, and may explain differences in results across studies,” as can real-time versus revised data – “a concern for monetary fundamentals but less of a concern for Taylor-rule fundamentals” – while “the frequency of the data does not seem to affect predictability.” (4) “Empirical results vary with the benchmark model, the sample period, forecast evaluation method and the forecast horizon. The random walk consistently provides the toughest benchmark.” (5) The author’s own analysis confirms that “only Taylor-rules display consistently significant out-of-sample forecasting ability at short horizons; and panel monetary models display some forecasting ability at long horizons,” but also “reveals instabilities in the models’ forecasting performance: the predictability of fundamentals varies not only across countries, models and predictors, but also over time” (Section 1).

Q4. What are the article’s stated scope restrictions?

Reduced-form rather than structural, out-of-sample rather than in-sample, nominal rather than real rates, and monthly or quarterly rather than high-frequency data. The author gives reasons for each. On reduced form: “the majority of the empirical work in this area is done with a reduced-form approach,” and “while there are theoretical structural models of exchange rate determination, typically they are too stylized to be literally taken to the data and successfully used for forecasting exchange rates. Moreover, fully developed structural models typically do not fit exchange rate data well, not to mention forecast them.” Theoretical models are sketched only “to motivate the choice of economic predictors.” On frequency: “we focus on monthly and quarterly frequencies, as they are the ones of interest to economists; we will not consider very high frequency data analyses that are instead mostly of interest to risk management and finance.” On in-sample versus out-of-sample: “In-sample fit does not necessarily guarantee out-of-sample forecast success,” so the focus is out-of-sample, and since “typically, real exchange rates are fitted in-sample, while nominal ones are forecasted out-of-sample,” the object is the nominal rate (Section 1).

Q5. Why does this matter for policy, according to the article?

Because central bank policy is set on forecasts, and exchange rate projections feed directly into those forecasts. The author cites Wieland and Wolters for evidence that “Central Bank policies in the US and Europe are described by interest rate rules, where interest rates respond to forecasts of inflation and economic activity, rather than outcomes,” and quotes Greenspan: “implicit in any monetary policy action or inaction is an expectation of how the future will unfold, that is, a forecast.” Concretely, “prior to each Federal Open Market Committee meeting, the Federal Reserve staff produces forecasts of several macroeconomic aggregates at horizons up to two years,” and “the variables forecasted by the staff include exchange rates, which influence current account projections as well as US real GDP growth, eventually” – with policy scenarios that “may also include dollar depreciation/appreciation scenarios,” against a baseline in which the dollar is typically assumed constant “according to the random walk model.” Exchange rate projections matter especially “for Central Banks of countries that are heavy importers/exporters of commodities,” with the Bank of Canada’s use of the Amano and van Norden terms-of-trade model given as an example (Section 1).

Q6. Which traditional predictors does the survey cover, and what is the verdict on each?

Uncovered interest parity, purchasing power parity, flexible- and sticky-price monetary models, productivity differentials and portfolio balance – and the verdict is broadly negative out of sample. The monetary models are the Frenkel-Bilson flexible-price specification and the Dornbusch-Frankel sticky-price specification, in which “PPP holds in the long run but does not hold in the short run” (Section 3.1). The evidence is characterized as mixed in-sample and less positive out-of-sample: Meese and Rogoff’s original result “has been confirmed by Chinn and Meese (1995) for short horizon (one-month to one-year-ahead) forecasts, by Cheung, Chinn and Pascual (2005), who find that the monetary model does not predict well even at longer horizons (i.e. five years), and by Alquist and Chinn (2008).” The prominent dissent is Mark (1995), who “finds strong and statistically significant evidence in favor of the monetary model at very long horizons (i.e. three to four years),” but whose robustness “has been questioned by Berkowitz and Giorgianni (2001), Kilian (1999), Groen (1999), Faust et al. (2003) and Rossi (2005a).” On productivity differentials, in the Balassa-Samuelson spirit and measured by real GDP per employee, Cheung, Chinn and Pascual “find that the model with productivity differentials does not forecast better than the random walk.” On portfolio balance – the Frankel and Hooper-Morton specification with a stock-balance term proxied by cumulated trade or current account differentials or government debt – Meese and Rogoff found that augmenting the monetary model with a trade balance measure still does not beat the random walk, “a finding confirmed by Cheung, Chinn and Pascual.” The section’s own summary: “with a few exceptions, their predictive ability is not significantly better than that of a random walk at short horizons,” and “some predictors (i.e. interest rate differentials) show significant in-sample fit, although with coefficient signs that are inconsistent with economic theory.”

Q7. What are Taylor-rule fundamentals, and what is the evidence on them?

Predictors derived from assuming both countries set interest rates by a Taylor rule, so that under uncovered interest parity the bilateral exchange rate reflects their relative inflation and output gaps; the evidence is mostly but not uniformly favorable. The logic is stated compactly: “if one considers two economies, both of which set interest rates according to a Taylor rule, by UIRP their bilateral exchange rate will reflect their relative interest rates, and thus, as a consequence, their output gaps and their inflation levels” (Section 3.2). Following Molodtsova and Papell, the survey distinguishes an “asymmetric” specification, which augments the rule with the real exchange rate (on the Svensson argument that policy also tracks the PPP level of the rate) and with lagged interest rates (interest rate smoothing as in Clarida, Galí and Gertler), from a “symmetric” specification with common coefficients and neither the real exchange rate nor lagged rates. On evidence: “Molodtsova and Papell (2009) show that eq. (8) forecasts exchange rates out-of-sample significantly better than the random walk for several countries, although the performance depends on the exact specification” – typically “exchange rates of four to seven countries out of twelve are significantly predictable” across their specifications – and further support comes from Molodtsova et al., Giacomini and Rossi, and Inoue and Rossi. The dissent: “Rogoff and Stavrakeva (2008) find that the empirical evidence in favor of Taylor-rule fundamentals is not robust,” and in-sample, “Chinn (2008) estimates the Taylor model in-sample and finds that the coefficient signs are not consistent with theory, and that the choice of the gap measure is not innocuous.” A structural worry is also recorded: since the model builds on uncovered interest parity, “if the latter does not hold in the data, it is surprising that the former holds.” The survey notes Taylor rules “are generally deemed to be a good description of monetary policy in the past three decades, but monetary policy may have changed during the recent 2007 financial crisis,” which motivated work adding financial-stress measures (Libor-OIS and Euribor-OIS differentials, financial condition indices, the TED spread) and funding-liquidity aggregates, both of which “find positive evidence.”

Q8. What is the net foreign assets predictor and why should it work?

It is Gourinchas and Rey’s NXA measure, and it works because external adjustment can occur through a valuation-driven wealth transfer via currency depreciation rather than only through future trade surpluses. The contrast with the intertemporal current account view is explicit: that view “suggests that the country will need to run future trade surpluses to reduce this imbalance,” whereas “Gourinchas and Rey (2007) argue instead that part of the adjustment can take place through a wealth transfer between that country and the rest of the world occurring via a depreciation of the value of its currency.” The construction: “NXA is the deviation from trend of a weighted combination of gross assets, gross liabilities, gross exports and gross imports, and measures the approximate percentage increase in exports necessary to restore external balance, that is, to restore the long run equilibrium of net exports and net foreign asset ratios” (Section 3.3). The evidence is “overall favorable”: Gourinchas and Rey and Della Corte, Sarno and Sestieri “find that the net foreign asset model can predict (effective) exchange rates out-of-sample significantly better than the random walk at both long and short horizons,” while Alquist and Chinn find it works for bilateral rates “in some sub-sample … at short horizons for some countries,” with “less favorable” results at longer horizons.

Q9. What about commodity prices?

They are attractive because they can be treated as close to exogenous for small commodity exporters, but whether they forecast out of sample depends on the data frequency. Chen and Rogoff’s motivation is an identification one: “typically, exchange rates are endogenously determined in equilibrium together with other macroeconomic variables, so it is difficult to predict exchange rate changes based on reduced-form models. However, if it were possible to identify an exogenous shock to exchange rates, that would cleanly predict exchange rate fluctuations,” and commodity price changes “act as ’essentially exogenous’ shocks for small open economies” whose exports are commodity-heavy – Australia, Canada and New Zealand (Section 3.4). The results split by frequency: “Chen and Rogoff (2003) find in-sample empirical evidence in favor of commodity prices as predictors of exchange rates; Chen, Rogoff and Rossi (2010) find that commodity prices are not significant out-of-sample predictors of exchange rates in quarterly data, and Ferraro, Rogoff and Rossi (2011) find that they are in daily data.”

Q10. What does the article set up as the state of agreement and disagreement before running its own test?

Consensus on the negatives and on one positive; disagreement on the long-horizon monetary result and on short-horizon UIP. The consensus negatives: “the monetary and PPP fundamentals have no predictive ability at short horizons,” “non-linear models are the least successful models,” and “should monetary models have any predictive ability at long horizons, it only appears in single-equation ECM and panel ECM model specifications.” The disagreements: “whether monetary fundamentals do have predictive ability at long horizons and whether UIRP has predictive ability at short horizons,” with “some evidence in favor of BMA models, although not strong.” And the single positive consensus: “the only ‘positive’ finding the literature reached a consensus on seems the fact that Taylor-rules and net foreign assets models have significant forecasting ability at short horizons” (Section 7).

Q11. What data and design does the author’s own exercise use?

Monthly and quarterly data for 19 countries against the US dollar, one uniform data transformation, a rolling window of half the sample, and short and four-year horizons. The country set is Australia, Austria, Belgium, Canada, Denmark, Finland, France, Germany, Greece, Ireland, Italy, Japan, New Zealand, Portugal, Spain, Sweden, Switzerland, the UK and the US, with data on “overnight interest rates, 3-month Treasury Bills, 5-year Treasury Bonds, GDP/industrial production, CPI and the money stock,” plus current account, trade balance, public debt and deficit at annual or quarterly frequency, sourced from the IMF database via Datastream and Philip Lane’s website (Section 7.1). Sample selection is disclosed: a further fifteen emerging and Asian economies were collected but “the sample sizes of either the exchange rate or the fundamentals were severely limited, or potentially severely affected by measurement error. These countries were discarded from the analysis.” Because “countries’ geographical definitions have changed over time (for example, after the start of the euro), the sample size differs across countries,” and available fundamentals differ by country and frequency. To make studies comparable, “we will focus on a unifying framework where the only data transformation is seasonal adjustment, which we perform consistently for all series,” done with one-sided backward equal-weighted moving averages, with unadjusted results “qualitatively similar.” Two design limitations are stated: “data limitations prevent us from studying the importance of real-time versus revised data,” and “net foreign assets data are available only at the annual frequency, so an analysis with the techniques currently used in this paper is not possible due to the small sample sizes” – so one of the two most promising predictors cannot be tested here. Non-linear and time-varying-parameter models are also excluded, the former because they “fit well in-sample, but produce poor out-of-sample forecasts,” the latter because “the most recent studies report lack of empirical support” for them.

Q12. Which tests are used, and why so many?

Because in-sample and out-of-sample evidence diverge, out-of-sample tests differ in their power properties, and forecasting performance is unstable, so robustness has to be tested along the forecast sample and the window size too. The reported statistics are Rossi’s (2005b) Granger-causality test robust to instabilities for in-sample predictive ability, the ratio of the model’s root mean squared forecast error to the random walk’s, and p-values from both the Diebold-Mariano-West and Clark-West tests against a random walk without drift, all with Newey-West heteroskedasticity- and autocorrelation-robust covariance matrices (Section 7.3). Two further tests address implementation robustness: “since the literature review highlights the importance of the forecast sample, we formally investigate the robustness (or lack thereof) of the empirical findings to both the choice of the out-of-sample forecast period and the choice of the estimation window size by using the tests proposed by Giacomini and Rossi (2010) and Inoue and Rossi (2012), respectively” (Section 7.2). A technical point is flagged on using the Diebold-Mariano-West test with nested models: “in this article we interpret its critical values following Giacomini and White (2006), which show that the DMW test can be used to compare forecasts of nested models provided a rolling window is used for estimation.”

Q13. What does the author find for the traditional predictors?

In-sample predictive ability at one month for several countries, very little at four years, and out-of-sample ability that is “almost inexistent.” The panels “show some in-sample predictive ability at short (one-month-ahead) horizons for UIRP and the monetary ECM models, but very limited at long (four-year-ahead) horizons for any of the models,” with the counts given in a footnote: “seven countries are significant for UIRP, three for PPP, one for the monetary and nine for the monetary ECM models,” which “show that, for example, interest rates have been capable, at some point in time, to predict one-month-ahead exchange rates.” At four years, “PPP is significant for four countries, the other models/predictors perform much worse.” Out of sample: “typically, models’ MSFEs are larger than the random walk’s, and neither DMW nor CW find significant predictive ability” (Section 7.3). Notably, the author reports finding less long-horizon ability for the monetary error-correction model than other papers, and treats explaining that discrepancy as a task rather than an embarrassment.

Q14. What does she find for Taylor-rule fundamentals, and why do the two tests disagree?

Clark-West finds short-horizon predictive ability – strongly for four countries and marginally for five others at one month, for none at four years – while Diebold-Mariano-West finds none, and the gap is exactly about how parameter estimation error is treated. The result is stated plainly, along with its caveat: “Among the fundamentals that we consider, then, they are among the most successful out-of-sample at short horizons. Importantly, note however that the DMW test does not find predictive ability.” The interpretation is the most instructive methodological passage in the article: “The discrepancy between the CW and DMW tests emphasizes that the way we treat parameter estimation error matters: if we compare models and correct inference for the fact that the Taylor model estimates more parameters than the random walk, we conclude that the Taylor model, when evaluated ‘in population’, has better predictive ability; however, if we evaluate its forecasts ‘at the actual estimated parameter values’, their forecasting ability is not superior to that of the random walk” (Section 7.3). This is the concrete content of the abstract’s condition that predictability appears when “a small number of parameters are estimated.”

Q15. What happens to richer models?

More predictors buy in-sample fit and lose out-of-sample performance; the panel monetary model works at long horizons for a few countries; Bayesian model averaging works nowhere. For the quarterly specification with real interest rates, the trade balance and the current account: “among the models we consider, this is one with the strongest in-sample predictive ability; however, its out-of-sample performance is extremely poor (possibly due to the large number of parameters to estimate).” For the panel: “the monetary panel model displays some in-sample predictive ability (at some point in time) and almost inexistent out-of-sample forecasting ability at short horizons; interestingly, however, there is some evidence of out-of-sample forecasting ability at long horizons for four countries. Again, note the very different results for CW and DMW: once more, the latter never finds predictive ability.” For Bayesian model averaging: “the BMA is not significantly better than the random walk benchmark in forecasting out-of-sample at any horizon and for any test statistics” (Section 7.3). The summary scatterplots make the division of labour visual: Taylor-rule fundamentals cluster as significant at short but not long horizons, the monetary panel model as significant at long but not short horizons for some countries and nowhere for others, and “only the in-sample predictive ability of Taylor-rules survives out-of-sample.”

Q16. What do the instability tests reveal, and why is this the article’s most consequential finding?

That predictive ability appears in short, country-specific episodes, and that whether a study finds it can depend on the estimation window size alone – which means published positive results may not be robust to small changes in procedure. The Fluctuation test traces performance over time: “in Figure 1 the Fluctuation test for Canada is never above its critical value, thus the predictor never displays significant forecasting ability; on the other hand, the Fluctuation test does detect predictive ability for France in the late 2000s, Japan in mid-2000 and the UK in 2009. Overall, Figure 1 shows that monetary fundamentals have occasional and very short-lived predictive ability at the one-month horizon for some countries at some point in time.” The window-size test shows “that indeed the window size strongly affects predictability for some countries,” with predictability “enhanced by a large window size” for Switzerland and the opposite for Japan and the UK. Combining both dimensions makes the point sharpest: “Figure 5 shows that only small estimation window sizes can detect forecasting ability of monetary fundamentals in Germany, and their predictive power is concentrated in the late 1980s; Figure 6 shows instead that only large estimation window sizes can detect predictability of monetary fundamentals in Japan, and that the predictability emerges in the late 2000s.” The author draws the moral for the literature directly: “Clearly, for some window sizes, it is possible to find some evidence of predictive ability for monetary fundamentals, thus confirming previous results that have been reported in the literature for selected window sizes and/or out-of-sample evaluation periods. At the same time, our results highlight the lack of robustness of these analyses to small changes in the procedures for forecast estimation and evaluation” (Section 7.3).

Q17. So is the Meese-Rogoff puzzle solved?

No – the author’s verdict is that it “does not seem to be entirely and convincingly overturned,” and that the pattern of occasional predictability constitutes a second puzzle. The conclusion’s full statement: “although some predictors (Taylor-rule fundamentals and net foreign assets) do exhibit some predictive ability at short horizons, and others (monetary fundamentals, especially in panel models) reveal some predictive ability at long horizons, none of the predictors, models, or tests systematically find empirical support for superior exchange rate forecasting ability of a predictor for all models, countries and time periods: typically, when predictability appears, it does so occasionally for some countries and for short periods of time. Thus, Meese and Rogoff’s (1983a,b) finding does not seem to be entirely and convincingly overturned.” And the new puzzle: “the predictive ability of the fundamentals is time-varying and occasional, yet existing time-varying parameter models are not successful in capturing it” (Section 8). The author also notes her review finds “slightly less out-of-sample predictive ability in favor of some of the models and predictors used in the literature” than the original papers claimed, monetary fundamentals at long horizons being the example.

Q18. What research agenda does the article set?

Explaining why predictability moves around, and finding ways to exploit that movement – which the author ties back to Meese and Rogoff’s own conjectures. The closing questions: “why does the predictability of exchange rate models change over time? Is it possible to design ways to exploit instabilities and improve exchange rates’ forecasts? An answer to these questions would require an answer to another important question: why are exchange rates poorly forecasted by economic models?” The author’s own pointer is historical: “In their papers, Meese and Rogoff (1983a,b) conjectured that sampling error, model misspecification and instabilities could potentially explain the poor forecasting performance of the economic models. Insights on the latter may be particularly relevant to understand and resolve the Meese and Rogoff puzzle” (Section 8).

Q19. How does this survey differ from the earlier reviews it names?

By covering the newer predictors and, unusually for a survey, by running a unified empirical test of them on current data. Frankel and Rose “review the empirical literature on exchange rates up to 1995, whereas we focus on more recent contributions and include a thorough empirical analysis which includes recent data as well as predictors that have been identified in the last decade”; Engel, Mark and West “focus on explaining the fluctuations of exchange rates using selected models, countries and fundamentals”; Melvin, Prins and Shand take “an financial investor’s point of view, e.g. carry trades,” while this article “focus[es] instead on forecasting exchange rates using economic models and macroeconomic predictors”; Lewis surveys puzzles in international financial markets, “especially risk premia and home bias”; and Rogoff and Froot and Rogoff focus on purchasing power parity, which appears here only as one predictor among many (Section 1).

Key terms in this paper

Definitions below follow the paper's own usage.

Meese and Rogoff puzzle
the finding of Meese and Rogoff (1983a,b, 1988) that a simple, a-theoretical random walk generates better out-of-sample exchange rate forecasts than economic models; the author is careful to state what it is not -- "Meese and Rogoff's (1983a,b) finding that the random walk provides the best prediction of exchange rates should not be interpreted as a validation of the efficient market hypothesis," since that hypothesis "does not mean that exchange rates are unrelated to economic fundamentals, nor that exchange rates should fluctuate randomly around their past values. Hence the puzzle."
"It depends"
the article's headline answer to "are exchange rates predictable?" -- predictability depends on the choice of predictor, forecast horizon, sample period, model and forecast evaluation method, and is most apparent when the predictor is a Taylor rule or net foreign assets, the model is linear, and few parameters are estimated; the survey's contribution is to make the dependence systematic rather than to declare a winner.
Taylor rule fundamentals
exchange rate predictors built from the interest rate rule the monetary authority is assumed to follow: inflation and output-gap differentials (the "symmetric" version), optionally augmented with the real exchange rate and lagged interest rates (the "asymmetric" version of Molodtsova and Papell); in this survey they are the one predictor on which the literature has reached a positive consensus at short horizons, and they are the only predictor whose in-sample significance survives out of sample in the author's own exercise.
Net foreign assets (NXA)
Gourinchas and Rey's (2007) predictor, defined as the deviation from trend of a weighted combination of gross assets, gross liabilities, gross exports and gross imports, measuring "the approximate percentage increase in exports necessary to restore external balance"; its logic is that external adjustment can occur through a wealth transfer via currency depreciation rather than only through future trade surpluses, and the survey classes it with Taylor rules as one of the two promising predictors -- though annual-only data keep it out of the author's own exercise.
Random walk without drift as benchmark
in this survey, the reference model against which forecasts must be judged, and the one the author argues is the appropriate choice -- "the random walk without drift is the toughest benchmark to beat," so "choosing an inappropriate benchmark model may overstate the empirical evidence in favor of the economic model's predictive ability."
DMW versus CW tests
two out-of-sample tests that answer different questions and can disagree: the Diebold-Mariano-West test evaluates forecasts "at the actual estimated parameter values," while the Clark-West test corrects inference for the extra parameters the economic model estimates relative to the random walk and so evaluates predictive ability "in population"; in this article's own results Taylor-rule fundamentals are significant under Clark-West and not under Diebold-Mariano-West, which the author treats as the sharpest illustration that "the way we treat parameter estimation error matters."
Instability of predictive ability
the survey's central empirical theme, that a predictor's forecasting performance is not stable over time or across implementation choices; it is measured with Giacomini and Rossi's (2010) Fluctuation test, which traces predictive ability over the forecast sample, and Inoue and Rossi's (2012) test, which traces it over estimation window sizes -- and the finding is that "forecasting ability, when present, is an occasional and short-lived phenomenon."
How this summary was made. Bibliographic fields are pulled from Crossref and OpenAlex and are not model-generated. The summary was drafted from the open-access manuscript , checked by a claim-grounding and calibration review pass, and approved before publishing. Found an error or a misrepresentation? Flag it here — corrections are welcome, especially from the authors.