What is a Croissant file?
Croissant is a metadata format for machine-learning datasets, standardized by MLCommons [1]. It is a layer of structured description that sits on top of the actual data files—CSVs, images, Parquet, TFRecords—and answers the questions a human or a data loader needs before touching the bytes: What is this dataset? Who made it, and under what license? What files does it contain, and how are they organized into records and fields?
Technically, a Croissant file is a JSON-LD document. It builds on schema.org vocabulary (so search engines and data catalogs already understand its top-level fields) and extends it with a cr: namespace that describes the dataset’s internal structure [2]. A Croissant file has three conceptual parts:
- Dataset-level metadata — name, description, license, citation, URL. The same fields you would expect on any dataset landing page.
- Distribution — the concrete resources:
FileObjects (individual files or archives) andFileSets (groups of files matching a pattern, such as “every.ome.tiffunder this directory”). - RecordSets — the logical structure: how the raw files map onto records and typed fields, so that a tool can load the data without bespoke parsing code.
The payoff is portability. A dataset shipped with a Croissant file can be loaded uniformly by any Croissant-aware tool—for example, mlcroissant or Hugging Face Datasets—instead of every consumer writing a one-off loader. It turns “here is a folder of files, good luck” into a machine-readable contract.
The dataset we’ll describe
For a concrete example, we’ll write a Croissant file for a real, published dataset from the Human BioMolecular Atlas Program (HuBMAP) [3]. HuBMAP builds open, spatially resolved maps of the human body at single-cell resolution, and every dataset is published with a stable HuBMAP ID, a UUID, and a DOI on the HuBMAP Data Portal.
The dataset is a CODEX (co-detection by indexing) multiplexed imaging run of human spleen tissue:
| Field | Value |
|---|---|
| Title | CODEX data from the spleen of a 18-year-old white male |
| HuBMAP ID | HBM548.TSMP.663 |
| UUID | afbd95090454842570908a76c80d1266 |
| DOI | 10.35079/HBM548.TSMP.663 |
| Organ | Spleen |
| Data provider | University of Florida TMC |
| Published | 2022-06-27 |
| Portal page | browse/dataset/afbd9…1266 |
After processing through HuBMAP’s codex-pipeline and SPRM, the dataset’s derived files follow a documented structure: a pipeline_output/expr/ directory of multi-channel OME-TIFF expression images, a pipeline_output/mask/ directory of segmentation masks, and a sprm_outputs/ directory of per-cell CSV feature tables [4].
A note on access. HuBMAP files are public but served through an authenticated assets gateway (and Globus), so the
contentUrls below name the real file locations but require a portal-issued token or a Globus transfer to download anonymously. Swap in your own access URL if you build a working loader.
The Croissant file
{
"@context": {
"@language": "en",
"@vocab": "https://schema.org/",
"citeAs": "cr:citeAs",
"column": "cr:column",
"conformsTo": "dct:conformsTo",
"cr": "http://mlcommons.org/croissant/",
"data": {
"@id": "cr:data",
"@type": "@json"
},
"dataType": {
"@id": "cr:dataType",
"@type": "@vocab"
},
"dct": "http://purl.org/dc/terms/",
"extract": "cr:extract",
"field": "cr:field",
"fileObject": "cr:fileObject",
"fileProperty": "cr:fileProperty",
"fileSet": "cr:fileSet",
"format": "cr:format",
"includes": "cr:includes",
"isLiveDataset": "cr:isLiveDataset",
"jsonPath": "cr:jsonPath",
"key": "cr:key",
"md5": "cr:md5",
"parentField": "cr:parentField",
"path": "cr:path",
"recordSet": "cr:recordSet",
"references": "cr:references",
"regex": "cr:regex",
"repeated": "cr:repeated",
"replace": "cr:replace",
"sc": "https://schema.org/",
"separator": "cr:separator",
"source": "cr:source",
"subField": "cr:subField",
"transform": "cr:transform"
},
"@type": "sc:Dataset",
"conformsTo": "http://mlcommons.org/croissant/1.0",
"name": "HuBMAP_CODEX_Spleen_HBM548.TSMP.663",
"description": "CODEX (co-detection by indexing) multiplexed imaging of human spleen tissue from an 18-year-old white male, published by the Human BioMolecular Atlas Program (HuBMAP). Processed with the HuBMAP codex-pipeline and SPRM, the dataset contains multi-channel OME-TIFF expression images, cell/nucleus segmentation masks, and per-cell feature tables with spatial coordinates.",
"url": "https://portal.hubmapconsortium.org/browse/dataset/afbd95090454842570908a76c80d1266",
"identifier": "HBM548.TSMP.663",
"sameAs": "https://doi.org/10.35079/HBM548.TSMP.663",
"license": "https://creativecommons.org/licenses/by/4.0/",
"creator": {
"@type": "sc:Organization",
"name": "University of Florida TMC (HuBMAP)",
"url": "https://hubmapconsortium.org/"
},
"publisher": {
"@type": "sc:Organization",
"name": "HuBMAP Consortium"
},
"datePublished": "2022-06-27",
"version": "1.0.0",
"keywords": [
"CODEX",
"multiplexed imaging",
"spleen",
"spatial biology",
"single cell",
"HuBMAP"
],
"citeAs": "HuBMAP Consortium. CODEX data from the spleen of a 18-year-old white male (HBM548.TSMP.663). HuBMAP Data Portal, 2022. https://doi.org/10.35079/HBM548.TSMP.663",
"isLiveDataset": false,
"distribution": [
{
"@type": "cr:FileObject",
"@id": "expr-image",
"name": "expr-image",
"description": "Multi-channel OME-TIFF expression image for region 1; one channel per antibody marker.",
"contentUrl": "https://assets.hubmapconsortium.org/afbd95090454842570908a76c80d1266/pipeline_output/expr/reg001_expr.ome.tiff",
"encodingFormat": "image/tiff"
},
{
"@type": "cr:FileObject",
"@id": "segmentation-mask",
"name": "segmentation-mask",
"description": "Cell and nucleus segmentation mask aligned to the region-1 expression image.",
"contentUrl": "https://assets.hubmapconsortium.org/afbd95090454842570908a76c80d1266/pipeline_output/mask/reg001_mask.ome.tiff",
"encodingFormat": "image/tiff"
},
{
"@type": "cr:FileSet",
"@id": "sprm-cell-tables",
"name": "sprm-cell-tables",
"description": "SPRM per-cell feature tables (mean/total marker intensity, cell centers, cluster assignments).",
"encodingFormat": "text/csv",
"includes": "sprm_outputs/reg001_expr.ome.tiff-*.csv"
},
{
"@type": "cr:FileObject",
"@id": "cell-mean-intensity",
"name": "cell-mean-intensity",
"description": "Per-cell mean marker intensity table produced by SPRM (one row per segmented cell).",
"containedIn": {
"@id": "sprm-cell-tables"
},
"encodingFormat": "text/csv",
"path": "sprm_outputs/reg001_expr.ome.tiff-cell_channel_mean.csv"
},
{
"@type": "cr:FileObject",
"@id": "cell-centers",
"name": "cell-centers",
"description": "Per-cell centroid coordinates produced by SPRM.",
"containedIn": {
"@id": "sprm-cell-tables"
},
"encodingFormat": "text/csv",
"path": "sprm_outputs/reg001_expr.ome.tiff-cell_centers.csv"
}
],
"recordSet": [
{
"@type": "cr:RecordSet",
"@id": "cell_centers",
"name": "cell_centers",
"description": "One record per segmented cell, with the cell's ID and centroid position in the image.",
"field": [
{
"@type": "cr:Field",
"@id": "cell_centers/cell_id",
"name": "cell_id",
"description": "Integer label matching the region of this cell in the segmentation mask.",
"dataType": "sc:Integer",
"source": {
"fileObject": {
"@id": "cell-centers"
},
"extract": {
"column": "ID"
}
}
},
{
"@type": "cr:Field",
"@id": "cell_centers/x",
"name": "x",
"description": "X coordinate of the cell centroid, in pixels.",
"dataType": "sc:Float",
"source": {
"fileObject": {
"@id": "cell-centers"
},
"extract": {
"column": "x"
}
}
},
{
"@type": "cr:Field",
"@id": "cell_centers/y",
"name": "y",
"description": "Y coordinate of the cell centroid, in pixels.",
"dataType": "sc:Float",
"source": {
"fileObject": {
"@id": "cell-centers"
},
"extract": {
"column": "y"
}
}
}
]
},
{
"@type": "cr:RecordSet",
"@id": "cell_mean_intensity",
"name": "cell_mean_intensity",
"description": "One record per segmented cell, with mean intensity for two representative markers. Add one field per channel present in the CSV header.",
"field": [
{
"@type": "cr:Field",
"@id": "cell_mean_intensity/cell_id",
"name": "cell_id",
"description": "Integer cell label; joins to cell_centers.cell_id.",
"dataType": "sc:Integer",
"source": {
"fileObject": {
"@id": "cell-mean-intensity"
},
"extract": {
"column": "ID"
}
}
},
{
"@type": "cr:Field",
"@id": "cell_mean_intensity/cd8",
"name": "CD8",
"description": "Mean CD8 marker intensity for the cell (T-cell marker).",
"dataType": "sc:Float",
"source": {
"fileObject": {
"@id": "cell-mean-intensity"
},
"extract": {
"column": "CD8"
}
}
},
{
"@type": "cr:Field",
"@id": "cell_mean_intensity/cd21",
"name": "CD21",
"description": "Mean CD21 marker intensity for the cell (follicular dendritic cell / B-cell zone marker).",
"dataType": "sc:Float",
"source": {
"fileObject": {
"@id": "cell-mean-intensity"
},
"extract": {
"column": "CD21"
}
}
}
]
}
]
}
Reading the file
A few things are worth pointing out in the example above.
-
The top block is just schema.org.
name,description,url,license,creator,datePublished, andciteAsare exactly what a data catalog or Google Dataset Search expects. This is why a Croissant file doubles as ordinary dataset metadata—and here every one of those values is the real, published HuBMAP record. -
distributiondescribes the physical files. The expression image and the segmentation mask are individualFileObjects addressed by their real pipeline paths (pipeline_output/expr/reg001_expr.ome.tiff,pipeline_output/mask/reg001_mask.ome.tiff). The SPRM per-cell CSVs are grouped as aFileSet(globsprm_outputs/reg001_expr.ome.tiff-*.csv), with two specific tables pulled out asFileObjects markedcontainedInthat set. -
recordSetdescribes the logical structure.cell_centersturns the centroid CSV into typedcell_id,x, andyfields;cell_mean_intensityexposes per-marker mean intensities (CD8,CD21, …). Each field’ssource/extractpoints back to a specific column, and the sharedcell_idlets a consumer join position to expression. -
Identifiers are stable and real. The HuBMAP ID (
HBM548.TSMP.663), the portal UUID, and the DOI all resolve to the same canonical dataset, which is what makes the metadata citable and reproducible.
To use it, save the JSON as croissant.json and load it with the reference implementation:
import mlcroissant as mlc
dataset = mlc.Dataset("croissant.json")
print(dataset.metadata.name)
for record in dataset.records(record_set="cell_centers"):
print(record["cell_centers/cell_id"], record["cell_centers/x"], record["cell_centers/y"])
break
The same file would load identically in any other Croissant-aware tool—which is the entire point of the format. (Because the HuBMAP assets gateway needs a token, point the contentUrls at a local copy or an authenticated URL before running the loader end to end.)
Summary
A Croissant file is a JSON-LD wrapper that makes a machine-learning dataset self-describing: schema.org metadata on the outside, an explicit map of files (distribution) and typed records (recordSet) on the inside. Wrapping the real HuBMAP CODEX spleen dataset HBM548.TSMP.663 in Croissant turns a tree of OME-TIFFs and SPRM CSVs into a portable, citable, tool-loadable object—readable by a human, a data catalog, and a training pipeline alike.
References
[1] MLCommons. “Croissant: A Metadata Format for ML-Ready Datasets.” Available at mlcommons.org/croissant.
[2] Mubashara Akhtar et al. “Croissant: A Metadata Format for ML-Ready Datasets.” Proceedings of the 2024 Workshop on Data Management for End-to-End Machine Learning (DEEM ‘24), 2024. doi:10.1145/3650203.3663326
[3] HuBMAP Consortium. “The human body at cellular resolution: the NIH Human Biomolecular Atlas Program.” Nature, 574:187–192, 2019. doi:10.1038/s41586-019-1629-x. Dataset: CODEX data from the spleen of a 18-year-old white male, HBM548.TSMP.663.
[4] HuBMAP Consortium. “codex-pipeline” and “SPRM” (Spatial Process & Relationship Modeling). GitHub: hubmapconsortium/codex-pipeline, hubmapconsortium/sprm.