Predicting Trade: The Upgrade That Wasn’t (Part 2)

Seasonal decomposition was supposed to fix the moving average. It didn’t — and the reason is that the problem wasn’t what we thought.

trade
forecasting
time-series
evaluation
python
Author

Aaron Koenigsberg

Published

July 10, 2026

NoteA note on how this was made

Part of an ongoing series on IMF International Merchandise Trade Statistics (IMTS) data, built by directing an AI research assistant: I set the questions, the evaluation design, and the interpretation; Claude wrote and ran the code. The analysis is what emerged from that back-and-forth — including the parts where my hypothesis turned out to be wrong.

The issue we thought we had

Part 1 built a deliberately simple one-step-ahead forecaster — a 12-month moving average of partner export shares — and used its residuals to flag anomalies. It also diagnosed the moving average’s failure modes, one of which was calendar blindness: trade has seasonal cycles, and a trailing mean can’t see them. The promised fix was seasonal-trend decomposition (STL) — model the seasonality explicitly, and the forecasts should improve.

This post tests that fix. The result is short: the seasonal upgrade loses to the naive baseline on every country. The rest of the post is why — and the why says the problem was never seasonality in the first place.

Setup

Same data and universe as Part 1: monthly goods-export shares to each country’s top 6 bilateral partners, for Ukraine, Venezuela, Yemen, Mexico, and Denmark, January 2017 through December 2025.

The method we thought would fix it

Three one-step-ahead forecasters, all honest — at month \(t\) they only ever see data through \(t-1\):

  1. MA12 (the Part 1 baseline). Prediction = trailing 12-month mean, shifted one period.

  2. Out-of-sample STL (OOS-STL). At each month \(t\), fit STL to data through \(t-1\), then forecast \(\hat{s}_t\) = last trend value + the seasonal component from the same calendar month a year earlier. An expanding-window version of the decomposition idea.

  3. MA12 + seasonal. Deseasonalize with past-only STL, apply the trailing 12-month mean to the seasonally-adjusted series, re-add the seasonal. The hypothesis: keep the MA’s low-variance level estimate, fix only its calendar blindness.

The result: it doesn’t fit as well

One-step-ahead mean absolute error by country, averaged across each country’s partner series:

Table 1: One-step-ahead mean absolute error in percentage points of export share. Bold = best.
Country MA12 (pp) OOS-STL (pp) MA12+seasonal (pp)
Ukraine 1.38 1.79 1.72
Venezuela 5.33 8.52 8.34
Yemen 8.05 12.16 12.06
Mexico 0.34 0.44 0.42
Denmark 0.71 0.95 0.93

The naive baseline wins everywhere — not by a little. OOS-STL is 25–60% worse; even the conservative MA12+seasonal hybrid loses on every country. Two seemingly obvious explanations fail quickly: it isn’t the warm-up window (both seasonal methods get the same one), and it isn’t robustness settings (robust STL is already being used).

You can see the failure directly — the STL forecast is simply jumpier than the thing it’s trying to predict:

Figure 1: Actual export share vs. one-step-ahead forecasts for each country’s largest partner. The OOS-STL forecast (green) chases noise; the moving average (red) stays put.

Why: the problem isn’t what we thought

Part 1 asserted that “trade patterns often have seasonal cycles” and counted calendar blindness as an MA failure mode. That’s true for trade levels. It is much less true for partner shares — seasonality hits the numerator and the denominator together, and mostly cancels in the ratio.

We can quantify that with the seasonal-strength measure from Hyndman & Athanasopoulos: \(F_s = \max\left(0,\ 1 - \frac{\mathrm{Var}(R_t)}{\mathrm{Var}(S_t + R_t)}\right)\), where \(S\) and \(R\) are the STL seasonal and remainder components. 0 means no seasonality; 1 means the seasonal component explains everything the trend doesn’t.

Table 2: Seasonal strength: total export levels vs. partner shares (median across each country’s top-6 partner series).
Country Levels (total exports) Shares (median, top-6 partners)
Ukraine 0.22 0.26
Venezuela 0.09 0.01
Yemen 0.46 0.04
Mexico 0.37 0.24
Denmark 0.57 0.22

Denmark’s total exports are strongly seasonal (0.57); its partner shares are not (0.22). Yemen’s shares have essentially zero seasonal structure. When the true seasonal signal is this weak, STL doesn’t extract signal — it manufactures it from noise, and the estimation variance it adds exceeds the variance it removes. The moving average never had a seasonality problem on this data representation. The lesson isn’t “STL is bad”; it’s that the upgrade was aimed at a failure mode this data doesn’t have.

Move on

That’s the post. We diagnosed a problem, built the standard fix, and the fix lost to the baseline — because the diagnosis was wrong. Shares aren’t seasonal; levels are. The useful output isn’t a better model, it’s a corrected map of where the signal actually lives.

Where that points next:

  1. Model levels, not shares. Levels are more seasonal than shares for four of the five countries (up to 0.57 for Denmark). If a seasonal model has anything to work with here, that’s where it lives.
  2. Detectors matched to slow disruptions. Part 1’s residual rule catches abrupt breaks; the slow-bleed cases (Venezuela’s sanctions era) need methods built for sustained shifts, not single bad months. That’s a different post.

Data: IMF International Merchandise Trade Statistics, bulk flat file download.