Critical Materials Atlas
Method · rigor · out-of-sample validation

Does last year predict this year?

The atlas nowcasts the newest trade year before the official data lands, on the premise that supply structure is persistent. That premise is testable. Predicting each year 2019–2024 from prior years only, a naive persistence model recovers the top exporter 84% of the time. But read that honestly: 84% is the naive benchmark, not a measure of skill — and the verdict that matters is what persistence can’t do (below), plus the fact that whether the atlas’s engine beats this benchmark is not yet proven.

The 2025 figure on the slider is a nowcast, not measured BACI. Its defensibility rests on one empirical question: how much does the prior year actually tell you about the next? This page answers it the only honest way — by hiding the answer and scoring the prediction. No model here sees the year it predicts.

84%
the prior year names the same top exporter as the observed year (out-of-sample, 2019–2024, 192 material-years). 95% CI 77–92%, bootstrap clustered by material.
4.8 pp
mean absolute error on that top exporter’s share
306
mean absolute error on concentration (HHI, 0–10,000 scale)
42%
direction of the year-over-year change called correctly — ~a coin flip

What it means

Two things, and the atlas states both. The structure is highly persistent. Last year alone pins this year’s leading exporter 84% of the time and its share to within about 4.8 points — which is precisely why a nowcast that carries the prior year forward and reconciles the current year’s partial customs data is a defensible provisional estimate, not a guess. This persistence baseline is the floor the reconciliation engine improves on by adding real current-year data.

But 84% is the naive benchmark, not proof of skill — and that distinction is the whole point. In forecasting, the persistence (no-change) model is the mandatory baseline every method must beat, and it is famously hard to beat — Meese & Rogoff (1983) showed a random walk out-forecasts structural exchange-rate models. Our 84% is that persistence baseline, so it measures how stable trade structure is, not how clever the engine is. Comparing it to blind random guessing (which we did earlier) was a strawman, now dropped: nobody’s alternative is a random draw. Two honest caveats even on the 84% itself: the material-clustered 95% interval is 77–92%; and because one country (China) tops many materials at once, the material-years share shocks — the true effective sample is smaller than 32 and that interval is, if anything, optimistic. (Rewritten after an adversarial audit flagged the strawman baseline; see the changelog.)

The honest limit. Persistence is good at levels and poor at turning points. It calls the direction of a year-over-year move correctly only 42% of the time — no better than chance. So read the nowcast as “the structure of last year carried forward and re-measured,” not “a forecast of where shares are heading.” When a share genuinely turns, a persistence-grounded nowcast is the last thing to see it. The atlas nowcast mitigates this by reconciling actual current-year Comtrade rather than extrapolating — but the limit is real and named here rather than buried.

The harder test, pre-registered

Honesty requires naming what this backtest does not prove. Because top exporters rarely change, 84% is a high base rate — persistence is a strong baseline, not a low bar cleared. The open question is whether the reconciliation engine, which folds in the current year’s partial customs data, actually beats that baseline. It clearly moves: the engine’s 2025 nowcast departs substantially from simply carrying 2024 forward, so it is incorporating new information — but whether those moves are signal or noise cannot be known until the truth arrives.

Pre-registration. The atlas’s 2025 nowcast is frozen and public now. When CEPII releases official BACI 2025 (trade data lags ~1.5 years, so around 2027), we will score the frozen nowcast against it and against the naive-persistence baseline — with MASE (a scaled error where <1 means it beats naive, >1 means it doesn’t) and a Diebold–Mariano test for whether any edge is significant — and publish the result here, win or lose. Until then, the honest statement is that the engine’s skill over persistence is unproven. The prediction exists before the answer does; that is the difference between a forecast and a fit.

Can a smarter model beat persistence? We tried nine.

Rather than assume persistence is best, we ran a bake-off the standard methods for a short, persistent, compositional, low-N series. Each predicts every year 2019–2024 from prior years only, scored the same way:

Modeltop-exporter hitshare error
Persistence (naive)84.4%4.83pp
3-year moving average83.3%4.96pp
5-year moving average81.8%5.74pp
Exponential smoothing81.8%5.06pp
Shrink 70% → 5-yr mean82.3%4.66pp
Linear trend82.3%5.75pp
ETS damped (state-space)83.9%6.12pp
Panel-Ridge (borrows across materials)82.3%4.85pp
Compositional (CLR shrinkage)83.9%4.51pp
Bayesian Dirichlet (BDARMA-core)84.9%4.56pp
The verdict, after testing nine models across every relevant family. Read the table honestly: no model clearly dominates persistence. On who leads, most of the sophisticated methods (compositional CLR, ETS, panel-Ridge) are actually worse than persistence’s 84.4%; only the Bayesian Dirichlet model (the core of BDARMA) edges it, at 84.9% — half a point, on 32 materials, well inside the noise. On the leader’s share, the compositional and Bayesian models cut the error from 4.83 to ~4.5pp (~7%), a small, real-looking improvement. To stop reading noise as skill, we ran the paired block-bootstrap by material (2,000 resamples of the 32 materials, which respects the China-tops-many dependence; build_nowcast_bootstrap.py → out/nowcast_bootstrap.json). On who leads, no model’s paired interval separates from persistence — the Bayesian model’s 84.9% vs 84.4% is well inside noise (persistence’s own 95% CI is 77–91%). On the leader’s share, the compositional CLR-shrink model shows a nominal edge (−0.32pp, raw 95% CI −0.55 to −0.07). But a second reviewer pressed the obvious point: we tried several models, so a single interval excluding zero is not enough. Correcting for that (Holm–Bonferroni across the bootstrapped models), CLR-shrink’s edge does not survive (raw p=0.018, and it needed p<0.010); the only multiplicity-robust results are that the 5-year average and linear trend are significantly worse than persistence. So the honest, multiplicity-corrected verdict is stronger than before, not weaker: no model reliably beats naive persistence at all. Its unbeatability is not an artefact of model-shopping — it survives the correction. And because that correction conditions on the exact 2019–2024 window, we re-ran it as a two-way (material×year) cluster bootstrap, resampling test-years as well as materials: the verdict is unchanged — no model beats persistence there either. (Estimand: uncertainty across the 32-material atlas conditional on the 2019–2024 backtest window — not a claim about all future years. The heavier ETS/panel/Bayesian rows are single-run point estimates, outside this bootstrap.) The defensible reading is therefore modest: a series this persistent leaves little for a cleverer extrapolation to add — the same lesson as Meese–Rogoff (1983) for exchange rates, offered as an analogy, not as a significance test we have passed. The real lever is not the forecasting model at all; it is the reconciliation engine’s use of current-year data, which is the pre-registered test above. (Reproducible: build_nowcast_models.py; methods: Aitchison 1986 compositional data; Snyder/Ord state-space; Bayesian Dirichlet ARMA.)

Hit rate by test year

Top-exporter hit rate when each year is predicted from the year before, across all materials with data in both years.

Predicted yearcorrecthit rate
201927/3284%
202027/3284%
202127/3284%
202227/3284%
202326/3281%
202428/3288%

Method

For each test year T (2019–2024) and each material, the “prediction” is the observed structure of year T−1 (reconciled CEPII BACI), scored against the observed structure of year T. Metrics: whether the top exporter matches; absolute error on that exporter’s trade share; absolute error on the exporter-side HHI; and, using year T−2 as the basis for a trend, whether the sign of the year-over-year change is called correctly. A 3-year-mean baseline performs the same to within a point, so the persistence result is not an artifact of one model. This tests forward persistence of observed BACI — the premise under the nowcast — not the reconciliation engine’s accuracy, which is validated separately against official BACI (top-1 exporter 25/30, 3.5% share error) on the methodology page. Reproducible: python build_nowcast_backtest.py.