Critical Materials Atlas
Open data · static API

The data behind the atlas

Every figure on this site is read from a handful of plain JSON files, served straight from GitHub Pages. No key, no rate limit, no backend — fetch them directly, or download the table as CSV.

The atlas is a thin presentation layer over committed, versioned data. Those files are the API: stable URLs, plain JSON, CC-friendly public sources. They update when the pipeline reruns (the same files power the interactive views and the per-material profiles).

The cube — one table behind everything

Most of what the atlas publishes now comes out of a single harmonised fact table. One row is one observation: a material, a country, a year, a measure, plus the identity of the source series it came from. 523,617 observations, 89 materials, 244 countries, 1900–2025, five public compilations in one place — BGS, USGS, CEPII BACI, World Mining Data and the IEA’s observed base-year columns. IEA projections stay out: a forecast must not sit where a query could difference it against measured history and call the result a trend.

URLWhat
/out/cube.parquetThe whole cube, 4.7 MB — for pandas, DuckDB, R, Polars
/out/cube.csv.gzThe same table as gzipped CSV, 5.8 MB — opens anywhere
/out/cube_summary.jsonWhat is in it: coverage per source, per measure, and the join rules
/out/catalog.jsonCoverage catalog — every series we hold or could reach, and whether it is ingested
/out/pairing.jsonCross-source comparison: which figures a second compilation agrees with, and why the rest differ
/out/sources.jsonSource metadata — what each compilation measures, on what basis, and why two of them can report different numbers for the same thing without either being wrong

The same table as SDMX

SDMX is the format statistical agencies exchange data in — the ECB, Eurostat, the IMF, the OECD and the UN all publish it. The cube is published that way too, so it can be loaded by anything that already reads SDMX without being told anything about us first.

The point of the standard is not the file format. It is that a data structure definition must declare which fields identify an observation and which merely describe it — and then nothing may be published that fails that key. Ours is FREQ . SOURCE . MATERIAL . MEASURE . STAGE . BASIS . NATIVE_CODE . REF_AREA over TIME_PERIOD. Two compilations counting the same year are two observations, not a conflict; ore and refined metal are different series, not one series to be summed. Asking whether our own data satisfied that key found 144 observations that did not — and all three causes were defects in the sources, not in our code. They are fixed and documented; today the answer is zero.

URLWhat
/out/sdmx/structure.jsonThe structure: 10 code lists, one data structure definition, one dataflow. Read this first — it says what may be pinned
/out/sdmx/mineral_flows.sdmx.csv.gz523,617 observations as SDMX-CSV, 3.8 MB
/out/cube_manifest.jsonEvery one of the 942 identities in the cube, with its years, country count and units — the menu to choose from before writing a filter
/out/sdmx.jsonSummary: code list sizes, observation-status counts, and the licence of every source in the export

Observation status uses the SDMX cross-domain code list, not our own words: A normal (520,685), E estimated by the source (1,263), Q missing because suppressed for confidentiality (1,222), N not significant — a real value that rounds to zero (447). Where a source publishes its own status codes, as BGS does, they are carried through rather than re-derived. One warning if you are mapping these yourself: USGS prints W for withheld, while SDMX’s W means “includes data from another category” — nearly the opposite claim.

Every source in the export is redistributable with attribution — BGS under the Open Government Licence, USGS as US public domain, CEPII BACI under Etalab 2.0 (cite Gaulier & Zignago 2010), World Mining Data free with attribution, and the IEA Critical Minerals Dataset under CC BY 4.0. The exporter refuses to write a dataflow containing a source whose terms are not recorded.

Why two sources report different numbers for the same thing

Each source is documented in /out/sources.json: what it measures, on what basis, its known thin spots, and one line on why it differs from the others. Coverage in that file is computed from the cube so it cannot go stale; the interpretation is written, because what a compilation means by “production” is not inferable from its rows. In short:

ComparisonWhere the difference comes from
BGS vs USGSBGS sums what countries report; USGS estimates a world total. Where reporting is incomplete USGS is usually larger — a gap reads as coverage, not error
BGS vs IEABGS often reports ORE at gross weight where IEA reports CONTAINED metal. Lithium is the extreme: roughly a 40× ratio, which is a unit difference and not a disagreement
BGS vs World Mining DataGenuinely comparable — both country-level mine production. But WMD estimates where a country is silent and BGS leaves the cell empty
Production vs BACI tradeProduction is of a MATERIAL; trade is of a PRODUCT CLASS (an HS code). They can only be differenced where a crosswalk pairs the form to the code explicitly

A different definition is a different dimension, not a different table

Sources disagree about what a word means. BGS counts manganese ore at gross weight; USGS counts contained manganese. Those are not rival numbers to be reconciled into one — they are two different quantities, and the cube stores both, distinguished by the columns that say so:

ColumnWhy it exists
stagemine / processed — where in the chain the quantity sits
basisgross weight or contained metal. The single most common cause of two sources “disagreeing”
unit + conversion_factorthe native unit and exactly what was applied to reach tonnes. A tonnage is never stored without them
code_system + native_code + native_labelthe source’s own identifier, kept forever. Never join on material alone — that is what turns a definition difference into a false contradiction
value_flagnil (produced nothing), trace, withheld (confidential), estimated_by_source. A true zero is not an absence
source + series_idwhich compilation said it, and which of its series

So the answer to “can you merge sources that use different vocabularies?” is yes — by adding the dimension that the vocabularies differ on, never by picking a winner. What the cube refuses to do is aggregate across those dimensions silently: summing gross ore and contained metal produces a number that describes nothing. Historical states (USSR, Yugoslavia, Zaire) are kept under their own codes for the same reason — merging them into successors is a choice a query should make on purpose, not one the table makes for you.

Endpoints

URLWhat
/out/data.jsonThe 32 materials with their mine / refine / reserve layers + the EU import-origin lens
/out/flows_2018.json … flows_2024.jsonReconciled bilateral trade (CEPII BACI), one file per measured year
/out/flows_2025.jsonProvisional nowcast (own reconciliation of partial Comtrade)
/out/flows_2026.jsonDirectional scenario (2025 structure, levels tilted by Q1 momentum × prices)

Base: https://criticalmaterialsatlas.org. The CSV the atlas exports (the Download data button, top right of the table) is the per-material metric sheet for the selected year.

Schema — flows_<year>.json

{
  "year": 2024,
  "names":  { "CD": "Congo [DRC]", "CN": "China", … },   // ISO-2 → display name
  "materials": {
    "cobalt": [
      { "from": "ZM", "to": "ZA", "value": 78998908 },   // exporter, importer, USD
      …
    ],
    "lithium": [ … ], …
  },
  "provisional": true            // only on 2025/2026
}

Schema — data.json

{
  "headlineYear": 2024, "dataUpdated": "15 Jun 2026",
  "materials": [
    {
      "label": "cobalt", "title": "Cobalt oxides & hydroxides (CN 2822 00 00)",
      "note": "…", "hhi": 0.474,
      "mined":    [ {"c":"CD","v":76}, {"c":"ID","v":10}, … ],   // USGS, % world output
      "refined":  [ {"c":"CN","v":76} ],                       // IEA, % processing
      "reserves": [ {"c":"CD","v":55}, {"c":"AU","v":15}, … ],   // USGS, % reserves
      "origins":  [ {"c":"CN","v":62.25,"eur":12038563}, … ]      // EU import-origin lens (Comext)
    }, …
  ]
}

Use it

const f = await (await fetch(
  'https://criticalmaterialsatlas.org/out/flows_2024.json')).json();
// top exporter of lithium, by value
const o = {}; for (const {from,value} of f.materials.lithium) o[from]=(o[from]||0)+value;
console.log(Object.entries(o).sort((a,b)=>b[1]-a[1])[0]);   // → ["CL", …]
Attribution & licence. The data derives from public sources — UN Comtrade (reconciled via CEPII BACI), USGS Mineral Commodity Summaries, IEA Critical Minerals Outlook, Eurostat Comext, World Bank. Use freely; please cite the original source for serious work, and link this atlas as the derived/reconciled layer. The reconciliation method is documented in the technical note and reproducible from the engine repo.