Critical Materials Atlas
Methodology & validation

Where do critical raw materials really come from?

Tracing 32 critical materials from mine to refinery to real trade — for any country. A method note, public data only.

For each material the tool answers three questions with real data: where is it mined, where is it refined, and who actually trades it with whom — globally, in both directions. Pick any material and any country. The European Union is available as one option among others — an aggregation you can look at, with a dedicated correction built for it (below) — but the tool is not EU-specific.

The layers

Why the refiner is not the source

Under customs rules, refining counts as "origin": when a country imports an ore, refines it and re-exports, the metal takes the refiner's nationality and the trade record stops there. One layer deeper, the picture changes — the apparent origin is often just the chokepoint, not the mine:

MaterialRefined origin (traded)Actually mined — USGS
CobaltChina 62%DR Congo 76%, Indonesia 10%
NickelNorway 33%Indonesia 67%, Philippines 11%
TantalumChina 37%DR Congo 40%, Rwanda 30%
LithiumChile 75%Australia 52%, Chile 22%

So an apparent China dependence on cobalt is, upstream, a Congo mining dependence funnelled through a Chinese refinery — China is the chokepoint, not the source. This mine-vs-refiner gap is well established in IEA and USGS supply-chain work; what this tool adds is to overlay all three layers — mine, refinery and customs trade — in one material-level interface, where mainstream dashboards (OEC, Resource Trade Earth, Comtrade) typically expose just one. It draws the implied mine → refiner → buyer chain, while being explicit (see Limitations) about what is measured versus referenced.

A method note: correcting the EU's import statistics

The same materials drive a classic correction specific to the European Union — the analysis this project grew out of, documented here (the live tool itself is global and country-agnostic). For extra-EU imports, the Eurostat Comext partner field is the country of origin. That single fact separates two views of the same data:

Across the thirty-two materials the naive and corrected panels disagree on the single most important fact — who the bloc depends on — and only the corrected one is true. China is the real origin for ten, far from a majority; the two sharpest dependencies are not China at all (beryllium is 100% United States, boron 98% Turkey), and the rest span South Africa, Brazil, Japan, Vietnam, Guinea, Gabon, Russia, Chile, Algeria/Qatar, Norway, Mexico, Kazakhstan and Tajikistan — almost none visible in the member-state view. A sample:

Material (CN8)Naive top — member stateCorrected top — originHHI
Rare-earth magnets (8505 11 10)Germany / Poland / NL spreadChina 93%, PH 3%, VN 2%0.86
Magnesium (8104 11 00)Netherlands 44%, DE 18%China 92%, IL 7%0.85
Beryllium, unwrought (8112 12 00)Spain 63%, FR 28%United States 100%1.00
Boron, natural borates (2528 00 00)scattered (none > 20%)Turkey 98%, BO 1%0.96
Ferro-niobium (7202 93 00)Netherlands 49%, DE 15%Brazil 84%, CA 16%0.73
Cobalt oxides (2822 00 00)Belgium 43%, DE 18%China 62%, GB 28%, BR 8%0.47
Lithium carbonate (2836 91 00)Netherlands 39%, DE 34%Chile 75%, US 13%, AR 6%0.58
Antimony (8110 10 00)France 37%, Belgium 37%Tajikistan 69%, VN 10%, CN 6%0.49
Manganese ore (2602 00 00)France 42%, ES 26%Gabon 51%, ZA 41%0.43
Bauxite (2606 00 00)Ireland 35%, DE 18%Guinea 59%, BR 18%0.39

Why the EU correction can't be faked from the raw download

Bloc aggregates double-count. The extract carries bloc reporters/partners (EU, EA, EU27_2020, WORLD) alongside the 27 member states. Summing without dropping them inflates origin totals ~4.6× — yet leaves the shares unchanged, so the error is invisible to anyone who validates on percentages alone.
The Rotterdam/Antwerp transit effect. Member-state figures attribute goods to the point of customs clearance, not consumption. Without the origin correction the largest "dependency" is a port.
Confidential & suppressed flows. For low-volume strategic materials, sensitive trade lines are withheld or masked in Comext; they must be detected and handled, not silently summed as zero.
CN8 code vintages. Eight-digit codes split and are reclassified over time (rare-earth magnets 8505 11 10 exists only from 2023). Series breaks must be respected, not stitched blindly.
Origin ≠ consignment, and units. The partner-as-origin identity holds only for extra-EU flows; quantities arrive as QUANTITY_IN_100KG and need conversion.

Because some codes (gallium, germanium) carry 15 years of data, dependency is a trajectory, not a snapshot: gallium's China-origin share runs 96.8% (2022) → 85% (2023) → 68% (2024) as Canada and Russia step in — China's July-2023 export controls visible directly in the customs record.

Limitations — what is measured vs referenced

This is an overlay of three different measures, not one observed supply chain. Read each row as three lenses laid over each other, not a single flow:

Provisional 2025 — a self-built nowcast

BACI lags ~1.5 years, so for 2025 (flagged 2025* on the slider) the atlas does not wait for CEPII — it uses a nowcast built here from raw UN Comtrade. A small Python pipeline (reconcile/) replicates BACI's method: it pulls both mirror reports of every flow (exporter FOB, importer CIF), estimates CIF/FOB markups, weights each reporter by a variance-components reliability model (E[discrepancy²] = vari + varj), and reconciles by inverse-variance averaging on logs. Validated against CEPII's official 2024 on what this atlas shows — exporter shares and concentration, not just a global correlation: the top exporter matches for 25 of 30 materials, exporter shares are within ~3.5% on average (median 3.2%), and the concentration ranking (HHI) correlates 0.92 (underlying flow structure: 0.975 log-value correlation across 21.7k flows). Absolute levels run ~1.8× higher (current Comtrade has absorbed revisions and late filers since BACI's early-2025 snapshot), so 2025 is calibrated to BACI's 2024 scale per material. A by-product finding: CIF/FOB rates are not identifiable on a 31-product slice (R²≈0.01) — BACI's gravity estimation needs the full ~5,000-product universe — so we fall back to a robust per-product median markup.

Caveat — partial coverage. Only about half of countries have filed 2025 annual data so far, so 2025 is genuinely provisional: stable monopolies (niobium, gallium) hold, but a material that leans on a country which has not yet reported (e.g. bauxite, whose dominant Guinea→China flow China has not filed) is distorted. Treat 2025 as indicative, not measured. Every earlier year (2018–2024) is reconciled CEPII BACI, untouched by this.

2026 (flagged 2026**) goes one step further — a directional scenario. The year is barely underway, so there is no annual data at all; only ~one quarter of monthly trade exists. Rather than invent a structure, it carries 2025's reconciled bilateral structure forward and scales each material by its Q1-2026 vs Q1-2025 export momentum (reporter-matched, from monthly Comtrade) blended with the change in its commodity price (World Bank Pink Sheet, for the ~7 base metals it covers). One quarter cannot credibly predict who-trades-with-whom, so shares are held at 2025 and only levels tilt: platinum rises (corroborated by a ~2× price move), antimony and graphite fall (export-control pressure). It is a trend signal — where each material is heading as of mid-2026 — not a measurement.

Validation & failure cases

The reconciliation engine (comtrade-reconcile, open + reproducible) is validated against official CEPII BACI on what this atlas shows — exporter shares and concentration, not a global level correlation.

Validation vs BACItop-1 exportertop-3 overlapshare MAEHHI corrlevel ratio
2024 (newest BACI year)25/302.57/33.5%0.92~1.8×
2022 (settled year)22/302.43/33.9%0.89~1.5×

The figures above are for exporters; importers validate similarly (2024: top-1 22/30, share MAE 4.2%, HHI corr 0.97) — both sides of the bilateral matrix hold. The validation is reproducible in CI from committed data (no API key) — see the engine repo.

Where it is weakest — surfaced, not hidden (2024):

Full per-material 2024 result (red rows = top-1 miss):

MaterialTop exporter (ours vs BACI)share MAEHHI oursHHI BACI
fluorsparZAF vs MEX9.0%0.250.25
graphiteTZA vs CHN6.7%0.140.21
arsenicJPN vs CHN4.5%0.180.18
phosphateMAR vs JOR4.0%0.160.15
tantalumCHN vs USA3.5%0.140.16
berylliumKAZ ✓5.5%0.530.80
cobaltCOD ✓5.2%0.220.17
vanadiumAUT ✓4.7%0.160.13
tungstenCHN ✓4.7%0.280.40
titaniumJPN ✓4.3%0.130.18
manganeseZAF ✓4.2%0.440.33
baryteIND ✓3.8%0.090.11
siliconCHN ✓3.5%0.140.20
strontiumDEU ✓3.4%0.430.44
feldsparTUR ✓3.4%0.320.21
Ga/Ge/HfCHN ✓2.9%0.130.14
boronTUR ✓2.8%0.470.38
heliumQAT ✓2.8%0.170.16
antimonyTJK ✓2.7%0.190.17
cokingcoalAUS ✓2.6%0.210.22
palladiumZAF ✓2.6%0.180.14
magnetsCHN ✓2.3%0.380.41
nickelNOR ✓2.3%0.110.09
niobiumBRA ✓2.2%0.540.59
bauxiteGIN ✓2.1%0.450.55
copperCOD ✓2.0%0.080.11
lithiumCHL ✓1.9%0.500.59
magnesiumCHN ✓1.9%0.480.54
platinumZAF ✓1.4%0.180.16
phosphorusVNM ✓1.1%0.370.35

Which findings are robust — a per-commodity status map

Three rounds of review added a lot of caveats (BACI is not truth; the engine understates some concentrations; thin codes swing wildly). Scattered, they bury the actual result. So we fuse them into one status per commodity (build_commodity_status.py, derived from the thinness screen, the engine-vs-BACI spread over 2002–2024, and the trend test). Draw headline findings from the “robust” rows only; read the rest as indicative. Of 31 tracked commodities: 21 robust, 4 spread-sensitive (a real but mild engine-vs-BACI gap), 2 shared-code (gallium/germanium can’t be separated in trade), 3 thin-fragile (graphite, beryllium, arsenic — too little trade to measure), and 1 engine-understated (cobalt, where the engine misses DRC’s real dominance).

Material2024 tradeConcentration statusTrend (2002–24)
cokingcoal$119402Mrobustsignificant rising
copper$86803Mrobustsignificant falling
platinum$14666Mrobustno significant trend
palladium$13483Mrobustsignificant falling
nickel$13010Mrobustsignificant falling
bauxite$11533Mrobust · single-country dominated (a finding, not a flag)significant rising
manganese$6405Mrobustsignificant rising
magnets$4963Mrobust · single-country dominated (a finding, not a flag)significant rising
silicon$3890Mrobustno significant trend
phosphate$3352Mrobustno significant trend
helium$3334Mrobustsignificant falling
titanium$1332Mrobustborderline
baryte$939Mrobustsignificant falling
magnesium$882Mrobust · single-country dominated (a finding, not a flag)significant rising
phosphorus$763Mrobustno significant trend
vanadium$672Mrobustsignificant falling
feldspar$662Mrobustno significant trend
boron$458Mrobustn/a
tantalum$340Mrobustsignificant falling
tungsten$130Mrobust · single-country dominated (a finding, not a flag)significant rising
strontium$101Mrobust · single-country dominated (a finding, not a flag)significant rising
lithium$3779Mspread-sensitive · max engine-BACI HHI gap 0.06significant rising
niobium$3463Mspread-sensitive · gap 0.08; single-country dominatedno significant trend
antimony$772Mspread-sensitive · max engine-BACI HHI gap 0.11significant falling
fluorspar$519Mspread-sensitive · max engine-BACI HHI gap 0.06significant falling
germanium$1106Mshared-code · shares its HS6 (gallium/germanium); not separableunreliable
gallium—shared-code · shares its HS6 (gallium/germanium); not separableunreliable
graphite$52Mthin-fragile · low trade valueunreliable
beryllium$27Mthin-fragile · low trade valueunreliable
arsenic$11Mthin-fragile · low trade valueunreliable
cobalt$777Mengine-understated · engine HHI up to 0.63 below BACI in some yearunreliable

What it is and isn't valid for. Valid: shares, ranks and top exporters. Concentration (HHI): reported as a spread, not a single number. The atlas engine is itself a full reconciliation — CIF/FOB deflation plus inverse-variance reliability weights on the two mirror reports (reconcile.py), the same family as CEPII BACI — not a plain average. An adversarial council pressed the right point: neither reconciliation is ground truth, so it is wrong to “correct toward BACI.” We now report concentration four ways per commodity (build_recon_envelope.py → out/recon_envelope.json): the two raw mirror reports — exporter-only (X/FOB) and importer-only (M/CIF) — and the two reconciliations (engine, BACI). A second review rightly rejected our first framing of this as the raw pair “bracketing the truth” with the reconciliations “inside”: two reports biased the same way (CIF vs FOB, confidentiality, re-exports) do not bound the truth, and because HHI is non-linear the aggregate need not fall between them — indeed BACI lands outside the raw pair for 24 of 30 commodities and the engine for 13 of 30 (2024). So the four views are a mirror-report spread — a measure of how uncertain a commodity’s concentration is — not a confidence interval, and no column is ground truth. The spread is wide: median ~0.07–0.09 HHI, up to 0.54 (manganese). The one directional thing we can say: the atlas engine tends to run low — often below both mirror reports (graphite 0.14 vs 0.23/0.48; beryllium 0.53 vs 0.62/0.78) — so its HHI is likely an underestimate for thin, dominated commodities. That is a direction, not a proven lower bound (an estimator can simply be wrong). Four ways, 2024 (exp / imp / engine / BACI): manganese 0.89 / 0.35 / 0.44 / 0.33, graphite 0.23 / 0.48 / 0.14 / 0.21, beryllium 0.62 / 0.78 / 0.53 / 0.80. The earlier one-number “beryllium 0.53→0.80” overstated false precision: read it as a wide, unresolved spread of roughly 0.5–0.8, the engine at the bottom and BACI at the top. Which commodities are too thin to carry a concentration number at all? Rather than judge that after the fact, we screen every code up front (build_thinness_screen.py): a commodity is flagged thin-fragile if its annual trade is under $100M or it has fewer than 8 active exporters. (High single-exporter leverage is not a flag on its own: in a deep market it is genuine one-country dominance — bauxite→Guinea, tungsten→China — a finding, not a data problem.) Only 3 of 31 are thin-fragile (graphite, beryllium, arsenic); their concentration is indicative, not measured. Building this ex-ante screen is also what caught our own cobalt error — see the status map above and the trend note below. Is the understatement fixable at source? We tried (reconcile/operator_test.py); simple operator substitutions do not remove it. The two-sided reconciliation’s geometric mean was the obvious suspect, but switching it to an arithmetic level-space mean changes the concentration almost nothing. The other diluter is importer-only flows (imports with no matching exporter report — often re-export/entrepôt) inflating the tail; but a full down-weight sweep shows no middle setting helps — partial weights barely move the BACI gap while worsening the fit, and only dropping them entirely improves raw-bracket coverage, at the cost of over-shooting past BACI. We also tested the principled candidate a reviewer named — the log-normal smearing correction (multiply the geometric mean by exp½σ² to retransform from log-median to level-mean). It fails badly here: with only two reports per flow the variance σ² is far too noisy, and the exponential over-inflates high-disagreement flows (a 100× mirror gap becomes a ~13× multiplier, pushing the value above both reports), so concentration over-shoots — raw-bracket coverage collapses from 17/30 to 6/30. We measure all of this two truth-agnostic ways: agreement with BACI (an external reconstruction, not ground truth) and coverage of the raw [exporter, importer] bracket; no operator we tested improves one without hurting the other. The one lever left untried is a gravity-fitted CIF like BACI’s, which reconcile.py documents as statistically unidentified on this 31-code slice (R² ≈ 0.01). So among the operators available to us, none improves the estimate — we report the spread as-is rather than hand-tune the engine toward BACI, which is not a target to hit. Does this band bias the trend claims? We had asserted “the dilution shifts the level not the trend” on only two years, which was not enough — so we did the full test (build_trend_robustness.py → out/trend_robustness.json): recompute the export-HHI series every year 2002–2024 on the engine reconciliation and independently on CEPII BACI, and run a Mann–Kendall trend test on each — and, because a reviewer noted annual HHI is serially correlated (which inflates plain-MK significance), using the Hamed–Rao autocorrelation-corrected variant. The correction barely moves the picture: it drops the engine’s significant trends from 21 to 20 of 29, so the series really are trending, not just persistent. A second honest point the reviewer forced: the raw “sign+significance agree for 26 of 29” headline is inflated because 8 of those agree only by both finding “no trend,” which is cheap. So we report the quantity that actually matters: among the 18 commodities where both series detect a significant trend, the direction agrees 18 of 18 — the reconciliation choice never flips a real trend’s sign, and all 18 survive a Benjamini–Hochberg multiple-testing correction. The three disagreements are all significance-level (one series clears the bar, the other just misses: arsenic, titanium), not direction reversals — except cobalt, the one substantive conflict (engine: a significant fall; BACI: a rise), which we ran down — and it exposes a real engine limitation, not a BACI one. (Correction, 31 Aug: our first pass called this a “thin-code artifact” on the strength of a buggy extraction that had dropped most exporters; cobalt’s code is not thin, and the story is the opposite.) Cobalt’s refined code (HS 2822.00, oxides/hydroxides) is a real $0.8–4.9 billion market, and BACI captures a genuine event: DR Congo’s share of those exports surged from 41% (2016) to ~85% (2020–2023) — the battery-era cobalt-hydroxide boom — then fell back in 2024 on the price crash (BACI HHI 0.25→0.71→0.17). The atlas engine missed it entirely: in 2020 it reconstructs only $0.7B of the trade (BACI: $3.5B) and puts DRC at 15%, not 84%. The reason is the engine’s own reliability weighting: DRC is a weak customs reporter, so its large but under-corroborated export claims get shrunk in reconciliation — exactly the “engine understates concentration” failure mode, here on a big, real dependency rather than a thin one. So the honest reading reverses: for cobalt, trust BACI and production (DRC ~70% of mine supply), not the engine’s smoothed series, whose “falling concentration” is an artefact of progressively shrinking DRC. This does not touch the 18/18 significant-trend agreement elsewhere, but it is a genuine scalp for the reconciliation engine, kept in view rather than buried. Not valid: absolute dollar levels — there is a known ~1.5–1.8× offset (current Comtrade runs above BACI's published values; diagnosed above), so arc and Sankey widths are relative, not dollar-precise. The earlier claim that the offset is a clean uniform multiplicative factor that “cancels in shares by construction” was too strong and is retracted: the two-year measurement above shows the offset falls unevenly on the leader, which is precisely why concentration needed correcting (reproducible: build_level_bias_audit.py, build_hhi_correction.py). 2025* is provisional (partial reporting, level-calibrated to BACI; not independently validated until BACI 2025 is released). 2026** is a directional scenario (Q1 momentum, shares held at 2025) — a trend signal, not a measurement.

Analytical layers & trend analysis

On top of the mine→refine→trade→reserves core, the atlas adds several screening lenses — each documented and caveated on its own page, and, with full citations, in the technical note:

The reconciled trade series runs 2002–2024, and the Trends view does not merely plot it: each material's export-concentration, China-share and origin-gap series is tested with the Mann–Kendall trend test, the Theil–Sen slope and the Pettitt structural-break test, Benjamini–Hochberg-corrected across the 32 materials. 9 of 32 materials show a statistically significant rising export-concentration trend, with structural breaks clustering in 2012–2016. The critical-minerals literature typically plots concentration; testing and dating it is the contribution. Full methods and references: technical note.

Data & reproducibility

Global bilateral trade (map / globe / flow): primary source UN Comtrade, used in its reconciled, mirror-harmonised form CEPII BACI (Gaulier, G. & Zignago, S., 2010, BACI: International Trade Database at the Product-Level, CEPII WP 2010-23) — releases HS02 + HS17 V202601, years 2002–2024 (HS02 vintage through 2016, HS17 from 2017). The year is selectable in the tool. Does the 2016/17 vintage join manufacture a trend? Tested: the mean year-over-year change in top-exporter share at the splice (4.6 pp) is if anything smaller than a typical year (5.1 pp average across all boundaries) — −0.5 standard deviations, well inside ordinary variation. The splice is invisible in the series, so trends across 2002–2024 are not artefacts of it (reproducible: build_splice_sensitivity.py).

EU import-origin lens (Table view): Eurostat Comext, dataset DS-045409, extra-EU imports, annual. Base-R pipeline, parameterised by CN code, refreshes from the live API. Mine production & reserves: USGS Mineral Commodity Summaries (annual editions 2020–2024, with a per-layer year slider). Refining shares: best-per-material — BGS World Mineral Statistics (OGC API, annual, for the metals it reports) else IEA Critical Minerals Outlook 2026 / EU CRM 2023 / USGS. All sources are public and free (no API key); figures are rounded.