Datasets · DS-01 – DS-04 · 12,307 rows · CC-BY-4.0
Datasets
Labelled corpora generated by the lab's engines and grounded in measured USGS spectra —
built for training against, benchmarking against, and breaking color pipelines with.
Each ships with a datasheet, grouped splits, and a published baseline.
Every row is produced by the released engines from a seeded build — the same commit
reproduces the same bytes — and each dataset's datasheet states where its labels come
from and what they do not establish. None of the labels are human judgements: observer
models are population averages, and the datasheets say so wherever it matters. Real
measured data enters through snapshot usgs-splib07a-1
(1,752 public-domain USGS reflectance spectra); everything derived is CC-BY-4.0.
Regenerate with node datasets/build/run.mjs.
The generator, the split hashing, the baseline, and the datasheet renderer are in the
repository under datasets/build/ — seeded throughout, so the
corpus is reproducible byte for byte from a commit hash. Splits are grouped by origin
(source record, anchor colour, palette, field), never split at random: a random row
split would let a model score by recognising the source rather than the property.
Browse the dataset sources →