About the Data

How the GRAND database is built — unifying science from around the globe.

Where it comes from

The GRAND database brings together temporally-resolved N2O flux measurements and their covariates from several major public sources — GRACEnet, Ag Data Commons, the NANORP N2O Network, and the Global N2O Database — plus roughly 45 additional peer-reviewed publications. Every record carries its original source citation, delivered with each download.

Getting that data into one consistent, queryable form takes several careful steps:

1

Extraction

Where data was published as tables, we ingest it directly. Where it lived only inside a figure, we digitize it point-by-point with WebPlotDigitizer, then add covariates and metadata by hand from the manuscript text.

Values read from figures carry the small reading uncertainty inherent to digitizing a plot.

2

Harmonization

Source datasets rarely agree on units, names, or structure. Python pipelines convert every measurement to common units and map inconsistent terminology onto a single vocabulary — crops as Scientific (Common) names, instruments as Static Chamber / Automatic Chamber / Eddy Covariance, and so on.

3

Treatment averaging

Replicate plots within a treatment are averaged to a single daily value, so each record represents a treatment's mean flux on a given date (with the replicate spread retained alongside it), rather than one individual chamber reading.

4

Covariate enrichment

Soil texture, pH, bulk density, organic carbon, and site climate (MAP/MAT) are taken from the source study wherever reported. Where a study didn't report a value, we fall back to gridded global products — SoilGrids for soil, NASA POWER for climate — and flag every such value as modeled rather than measured, so its origin is always clear.

Good to know

A few characteristics of the data worth keeping in mind as you use it:

Before you model

  • Measurement gaps. Most data comes from static chambers, which typically sample every 1–2 weeks. GRAND fills the calendar so dates are continuous, but it never invents flux values — there is no gap-filling.
  • Treatment-averaged. Records are treatment means, not raw replicate-level data.
  • Measured vs. modeled. Some soil and climate covariates are gridded estimates where the original study was silent — always flagged as such, so you can include or exclude them as your model requires.
  • Figure-digitized values carry a small uncertainty from reading points off a plot.

Full column definitions and complete per-record citations ship in the README bundled with every download. For the complete methodology, see Ackett et al. (2026).