The atlas is a thin presentation layer over committed, versioned data. Those files are the API: stable URLs, plain JSON, CC-friendly public sources. They update when the pipeline reruns (the same files power the interactive views and the per-material profiles).
The cube — one table behind everything
Most of what the atlas publishes now comes out of a single harmonised fact table. One row is one observation: a material, a country, a year, a measure, plus the identity of the source series it came from. 523,617 observations, 89 materials, 244 countries, 1900–2025, five public compilations in one place — BGS, USGS, CEPII BACI, World Mining Data and the IEA’s observed base-year columns. IEA projections stay out: a forecast must not sit where a query could difference it against measured history and call the result a trend.
| URL | What |
|---|---|
| /out/cube.parquet | The whole cube, 4.7 MB — for pandas, DuckDB, R, Polars |
| /out/cube.csv.gz | The same table as gzipped CSV, 5.8 MB — opens anywhere |
| /out/cube_summary.json | What is in it: coverage per source, per measure, and the join rules |
| /out/catalog.json | Coverage catalog — every series we hold or could reach, and whether it is ingested |
| /out/pairing.json | Cross-source comparison: which figures a second compilation agrees with, and why the rest differ |
| /out/sources.json | Source metadata — what each compilation measures, on what basis, and why two of them can report different numbers for the same thing without either being wrong |
The same table as SDMX
SDMX is the format statistical agencies exchange data in — the ECB, Eurostat, the IMF, the OECD and the UN all publish it. The cube is published that way too, so it can be loaded by anything that already reads SDMX without being told anything about us first.
The point of the standard is not the file format. It is that a data structure definition must declare which fields identify an observation and which merely describe it — and then nothing may be published that fails that key. Ours is FREQ . SOURCE . MATERIAL . MEASURE . STAGE . BASIS . NATIVE_CODE . REF_AREA over TIME_PERIOD. Two compilations counting the same year are two observations, not a conflict; ore and refined metal are different series, not one series to be summed. Asking whether our own data satisfied that key found 144 observations that did not — and all three causes were defects in the sources, not in our code. They are fixed and documented; today the answer is zero.
| URL | What |
|---|---|
| /out/sdmx/structure.json | The structure: 10 code lists, one data structure definition, one dataflow. Read this first — it says what may be pinned |
| /out/sdmx/mineral_flows.sdmx.csv.gz | 523,617 observations as SDMX-CSV, 3.8 MB |
| /out/cube_manifest.json | Every one of the 942 identities in the cube, with its years, country count and units — the menu to choose from before writing a filter |
| /out/sdmx.json | Summary: code list sizes, observation-status counts, and the licence of every source in the export |
Observation status uses the SDMX cross-domain code list, not our own words: A normal (520,685), E estimated by the source (1,263), Q missing because suppressed for confidentiality (1,222), N not significant — a real value that rounds to zero (447). Where a source publishes its own status codes, as BGS does, they are carried through rather than re-derived. One warning if you are mapping these yourself: USGS prints W for withheld, while SDMX’s W means “includes data from another category” — nearly the opposite claim.
Every source in the export is redistributable with attribution — BGS under the Open Government Licence, USGS as US public domain, CEPII BACI under Etalab 2.0 (cite Gaulier & Zignago 2010), World Mining Data free with attribution, and the IEA Critical Minerals Dataset under CC BY 4.0. The exporter refuses to write a dataflow containing a source whose terms are not recorded.
Why two sources report different numbers for the same thing
Each source is documented in /out/sources.json: what it measures, on what basis, its known thin spots, and one line on why it differs from the others. Coverage in that file is computed from the cube so it cannot go stale; the interpretation is written, because what a compilation means by “production” is not inferable from its rows. In short:
| Comparison | Where the difference comes from |
|---|---|
| BGS vs USGS | BGS sums what countries report; USGS estimates a world total. Where reporting is incomplete USGS is usually larger — a gap reads as coverage, not error |
| BGS vs IEA | BGS often reports ORE at gross weight where IEA reports CONTAINED metal. Lithium is the extreme: roughly a 40× ratio, which is a unit difference and not a disagreement |
| BGS vs World Mining Data | Genuinely comparable — both country-level mine production. But WMD estimates where a country is silent and BGS leaves the cell empty |
| Production vs BACI trade | Production is of a MATERIAL; trade is of a PRODUCT CLASS (an HS code). They can only be differenced where a crosswalk pairs the form to the code explicitly |
A different definition is a different dimension, not a different table
Sources disagree about what a word means. BGS counts manganese ore at gross weight; USGS counts contained manganese. Those are not rival numbers to be reconciled into one — they are two different quantities, and the cube stores both, distinguished by the columns that say so:
| Column | Why it exists |
|---|---|
| stage | mine / processed — where in the chain the quantity sits |
| basis | gross weight or contained metal. The single most common cause of two sources “disagreeing” |
| unit + conversion_factor | the native unit and exactly what was applied to reach tonnes. A tonnage is never stored without them |
| code_system + native_code + native_label | the source’s own identifier, kept forever. Never join on material alone — that is what turns a definition difference into a false contradiction |
| value_flag | nil (produced nothing), trace, withheld (confidential), estimated_by_source. A true zero is not an absence |
| source + series_id | which compilation said it, and which of its series |
So the answer to “can you merge sources that use different vocabularies?” is yes — by adding the dimension that the vocabularies differ on, never by picking a winner. What the cube refuses to do is aggregate across those dimensions silently: summing gross ore and contained metal produces a number that describes nothing. Historical states (USSR, Yugoslavia, Zaire) are kept under their own codes for the same reason — merging them into successors is a choice a query should make on purpose, not one the table makes for you.
Endpoints
| URL | What |
|---|---|
| /out/data.json | The 32 materials with their mine / refine / reserve layers + the EU import-origin lens |
| /out/flows_2018.json … flows_2024.json | Reconciled bilateral trade (CEPII BACI), one file per measured year |
| /out/flows_2025.json | Provisional nowcast (own reconciliation of partial Comtrade) |
| /out/flows_2026.json | Directional scenario (2025 structure, levels tilted by Q1 momentum × prices) |
Base: https://criticalmaterialsatlas.org. The CSV the atlas exports (the Download data button, top right of the table) is the per-material metric sheet for the selected year.
Schema — flows_<year>.json
{
"year": 2024,
"names": { "CD": "Congo [DRC]", "CN": "China", … }, // ISO-2 → display name
"materials": {
"cobalt": [
{ "from": "ZM", "to": "ZA", "value": 78998908 }, // exporter, importer, USD
…
],
"lithium": [ … ], …
},
"provisional": true // only on 2025/2026
}
Schema — data.json
{
"headlineYear": 2024, "dataUpdated": "15 Jun 2026",
"materials": [
{
"label": "cobalt", "title": "Cobalt oxides & hydroxides (CN 2822 00 00)",
"note": "…", "hhi": 0.474,
"mined": [ {"c":"CD","v":76}, {"c":"ID","v":10}, … ], // USGS, % world output
"refined": [ {"c":"CN","v":76} ], // IEA, % processing
"reserves": [ {"c":"CD","v":55}, {"c":"AU","v":15}, … ], // USGS, % reserves
"origins": [ {"c":"CN","v":62.25,"eur":12038563}, … ] // EU import-origin lens (Comext)
}, …
]
}
Use it
const f = await (await fetch( 'https://criticalmaterialsatlas.org/out/flows_2024.json')).json(); // top exporter of lithium, by value const o = {}; for (const {from,value} of f.materials.lithium) o[from]=(o[from]||0)+value; console.log(Object.entries(o).sort((a,b)=>b[1]-a[1])[0]); // → ["CL", …]