Methodology & data notes

Everything needed to audit CO₂ Estimated vs Observed: where the numbers come from, what is computed, and what each published column means.

What this bundle contains

This directory is self-contained. It is the whole publishable artifact — app, data and documentation — and it has no dependency on the repository it was generated in.

FileRole
index.html, app.js, styles.cssThe app. All computation happens in the browser.
methodology.htmlThis page. The published methodology of record.
data/explorer_inputs.jsonThe input bundle the app fetches at load: metadata plus one object per year.
data/explorer_inputs.csvThe same rows as flat CSV, for spreadsheets and scripts. Not read by the app.
vendor/chart.umd.min.jsChart.js, vendored. No network requests at runtime.

To serve it, point any static file server at this directory — for example python3 -m http.server 8000 from here, then open http://localhost:8000/. There is no build step and no backend.

The bundle ships inputs only. No modelled series, deviation or metric is stored in the data files; all of them are recomputed in the browser from the coefficients on screen. That is deliberate — you cannot be handed a precomputed result that disagrees with the controls.

Observed CO₂ — measured, not ours

Three independent records are published, and the app scores the model against whichever one is selected. Each begins in a different year, so the baseline moves with the choice.

Reference seriesColumnSpanSource
NOAA global mean observed_noaa_global_ppm 1979– NOAA GML marine boundary layer global annual mean
Mauna Loa observed_noaa_mauna_loa_ppm 1959– NOAA GML Mauna Loa annual mean
Ice core + instrumental (spliced) observed_law_dome_instrumental_spliced_ppm 1900– Constructed here — see below

The spliced series is the only reference series we assemble rather than adopt. It takes the annually interpolated Law Dome ice-core reconstruction for 1900–1958, then the instrumental record from 1959: NOAA's global annual mean wherever it exists (1979 onwards), with Mauna Loa filling only the 1959–1978 gap. Because it is a construction, its citable form is the published file itself, data/explorer_inputs.csv, not a single upstream paper.

Two caveats travel with it. The ice-core part is smoothed by gas diffusion in the firn, so year-to-year variability before 1958 is damped and not comparable with the instrumental part. And Mauna Loa is one Northern-Hemisphere station reading above the global mean by a margin that grows over time (+0.75 ppm in 2000, +1.8 ppm in 2024), so the 1959–1978 window carries a small upward offset relative to the rest.

Emissions, and the two conversions

Global emissions come from the Global Carbon Budget via Our World in Data. The app offers three prediction series, assembled from the component columns rather than from the aggregate co2 column:

Prediction seriesOWID componentsPublished ppm field
Fossil fuels + flaringcoal_co2 + oil_co2 + gas_co2 + flaring_co2emissions_fossil_excl_cement_ppm
Fossil fuels + flaring + cementPrevious series + cement_co2emissions_fossil_cement_ppm
Fossil fuels + flaring + cement + change land-usePrevious series + land_use_change_co2emissions_fossil_cement_luc_ppm

The second series is the default. The third adds land-use change to the combined trajectory used previously. OWID component values are annual MtCO₂; blank component values are treated as zero before conversion.

MtCO₂ → GtC, factor 12/44. The sinks retain the carbon atom, not the whole molecule, so we keep the carbon fraction of CO₂ by molar mass. Published as emissions_gtc.

GtC → ppm, factor 2.124. The atmosphere weighs about 5.135·10¹⁸ kg, so one part per million by volume holds roughly 2.124 gigatonnes of carbon. This is the value used by the Global Carbon Budget and IPCC AR6. The selected prediction-series ppm field enters the model.

The model

For the selected observed series, baseline year B and active coefficients, the browser computes

C(T) = C_observed(B) + Σ E_ppm(t) · IRF(T − t), for B < t ≤ T

where the impulse-response function is a sum of one permanent term and three exponentials:

IRF(Δt) = a₀ + a₁·e^(−Δt/τ₁) + a₂·e^(−Δt/τ₂) + a₃·e^(−Δt/τ₃)

Each year's emissions are treated as a pulse; the IRF sets the fraction of that pulse still airborne in every later year. The baseline year contributes only its observed concentration — the model is anchored to measurement at one point and then runs free.

Sign convention: deviation = predicted − observed. Positive means the prediction overshoots reality; negative means it underestimates. This holds in the charts, the KPI cards and the period table, without exception.

The year-over-year changes shown for both series are comparison diagnostics computed after the fact. They are not inputs to the IRF.

Coefficients: two published sets, two exploratory ones

The IRF form above is shared by both IPCC parameter sets. What differs is the calibration — which is the point of offering both.

Preseta₀a₁ / τ₁a₂ / τ₂a₃ / τ₃Basis
Joos 2013 (AR5) 0.21730.2240 / 394.4 yr0.2824 / 36.54 yr0.2763 / 4.304 yr Fit to the multi-model mean of a 15-model carbon-cycle intercomparison. Adopted by IPCC AR5 WGI Ch.8 SM; published in Joos et al. 2013.
Forster 2007 (AR4) 0.2170.259 / 172.9 yr0.338 / 18.51 yr0.186 / 1.186 yr Response of the single Bern2.5CC reduced-complexity carbon-cycle-climate model on a 378 ppm background. The AR4 reference response for metrics — IPCC AR4 WGI Ch.2.
Short decay 0.170.21 / 120 yr0.27 / 25 yr0.35 / 3.5 yr Exploratory. Weights normalised to 1.0 but from no paper.
Long decay 0.270.19 / 600 yr0.27 / 60 yr0.27 / 8 yr Exploratory. Weights normalised to 1.0 but from no paper.

In all four sets the weights sum to 1.0 — a pulse starts out entirely in the atmosphere.

The two published sets differ in more than digits. AR4 puts a shorter leading timescale (173 yr against 394 yr) and more weight on the fast term against AR5's flatter distribution. AR5's own supplementary material describes its values as fitting parameters, not process coefficients with direct physical meaning, and states the IRF was updated from AR4. Switching between the two in the app is therefore a useful control: it shows how much of any deviation is a parameter choice inside the accepted IPCC range rather than a claim about the measurements.

Editing any coefficient by hand switches the preset to «Custom» and drops the source link, because at that point no publication stands behind the numbers.

Published column dictionary

ColumnUnitMeaning
yearcalendar year1900–2024, one row per year, sorted, no duplicates, no gaps.
emissions_gtcGtC/yrGlobal fossil-and-industry plus net land-use-change emissions, as carbon.
emissions_ppmppm/yrThe same emissions divided by 2.124. The model input.
emissions_fossil_excl_cement_gtc/ppmGtC/yr; ppm/yrCoal, oil, gas and flaring; excludes cement and land-use change.
emissions_fossil_cement_gtc/ppmGtC/yr; ppm/yrCoal, oil, gas, flaring and cement; excludes land-use change.
emissions_fossil_cement_luc_gtc/ppmGtC/yr; ppm/yrCoal, oil, gas, flaring, cement and land-use change.
observed_law_dome_instrumental_spliced_ppmppmSpliced observed series, 1900 onwards.
observed_noaa_global_ppmppmNOAA global annual mean. Empty before 1979.
observed_noaa_mauna_loa_ppmppmMauna Loa annual mean. Empty before 1959.

Empty cells are genuine absences of measurement, never zeros and never interpolated. The app starts each comparison at the first year the selected series actually has a value.

The JSON file carries the same rows under rows, plus a metadata block recording row count, the observed-series column names, and the relative path of the upstream CSV the bundle was generated from.

Validation applied at generation time, and the build fails if any of it breaks: rows sorted by year, no duplicate years, a continuous run from 1900, and units fixed as above.

What this model leaves out

Both IPCC coefficient sets are calibrated for a single pulse on a given background — 378 ppm for AR4's Bern2.5CC basis. Applying them to a whole emissions trajectory from a fixed baseline ignores that sink capacity depends on the state of the system, so the deviation cannot be read as a measured failure of the carbon cycle.

The baseline is anchored to a single observed year, which imports that year's measurement error into every later prediction. The spliced series mixes two measurement techniques with different smoothing. And the reported metrics reward different things: RMSE and mean bias are per-year averages, while cumulative deviation is a signed sum that grows with the number of years compared and cancels overshoots against underestimates — a value near zero there means the two directions balanced, not that the fit is good.

This is a model with declared assumptions, not a measurement.

Sources