icaoberg / An example of a Croissant file

Created Wed, 15 Jul 2026 00:00:00 +0000 Modified Mon, 10 Aug 2026 18:05:56 -0400

What is a Croissant file?

Croissant is a metadata format for machine-learning datasets, standardized by MLCommons [1]. It is a layer of structured description that sits on top of the actual data files—CSVs, images, Parquet, TFRecords—and answers the questions a human or a data loader needs before touching the bytes: What is this dataset? Who made it, and under what license? What files does it contain, and how are they organized into records and fields?

Technically, a Croissant file is a JSON-LD document. It builds on schema.org vocabulary (so search engines and data catalogs already understand its top-level fields) and extends it with a cr: namespace that describes the dataset’s internal structure [2]. A Croissant file has three conceptual parts:

  1. Dataset-level metadata — name, description, license, citation, URL. The same fields you would expect on any dataset landing page.
  2. Distribution — the concrete resources: FileObjects (individual files or archives) and FileSets (groups of files matching a pattern, such as “every .ome.tiff under this directory”).
  3. RecordSets — the logical structure: how the raw files map onto records and typed fields, so that a tool can load the data without bespoke parsing code.

The payoff is portability. A dataset shipped with a Croissant file can be loaded uniformly by any Croissant-aware tool—for example, mlcroissant or Hugging Face Datasets—instead of every consumer writing a one-off loader. It turns “here is a folder of files, good luck” into a machine-readable contract.


The dataset we’ll describe

For a concrete example, we’ll write a Croissant file for a real, published dataset from the Human BioMolecular Atlas Program (HuBMAP) [3]. HuBMAP builds open, spatially resolved maps of the human body at single-cell resolution, and every dataset is published with a stable HuBMAP ID, a UUID, and a DOI on the HuBMAP Data Portal.

The dataset is a CODEX (co-detection by indexing) multiplexed imaging run of human spleen tissue:

Field Value
Title CODEX data from the spleen of a 18-year-old white male
HuBMAP ID HBM548.TSMP.663
UUID afbd95090454842570908a76c80d1266
DOI 10.35079/HBM548.TSMP.663
Organ Spleen
Data provider University of Florida TMC
Published 2022-06-27
Portal page browse/dataset/afbd9…1266

After processing through HuBMAP’s codex-pipeline and SPRM, the dataset’s derived files follow a documented structure: a pipeline_output/expr/ directory of multi-channel OME-TIFF expression images, a pipeline_output/mask/ directory of segmentation masks, and a sprm_outputs/ directory of per-cell CSV feature tables [4].

A note on access. HuBMAP files are public but served through an authenticated assets gateway (and Globus), so the contentUrls below name the real file locations but require a portal-issued token or a Globus transfer to download anonymously. Swap in your own access URL if you build a working loader.


The Croissant file

{
  "@context": {
    "@language": "en",
    "@vocab": "https://schema.org/",
    "citeAs": "cr:citeAs",
    "column": "cr:column",
    "conformsTo": "dct:conformsTo",
    "cr": "http://mlcommons.org/croissant/",
    "data": {
      "@id": "cr:data",
      "@type": "@json"
    },
    "dataType": {
      "@id": "cr:dataType",
      "@type": "@vocab"
    },
    "dct": "http://purl.org/dc/terms/",
    "extract": "cr:extract",
    "field": "cr:field",
    "fileObject": "cr:fileObject",
    "fileProperty": "cr:fileProperty",
    "fileSet": "cr:fileSet",
    "format": "cr:format",
    "includes": "cr:includes",
    "isLiveDataset": "cr:isLiveDataset",
    "jsonPath": "cr:jsonPath",
    "key": "cr:key",
    "md5": "cr:md5",
    "parentField": "cr:parentField",
    "path": "cr:path",
    "recordSet": "cr:recordSet",
    "references": "cr:references",
    "regex": "cr:regex",
    "repeated": "cr:repeated",
    "replace": "cr:replace",
    "sc": "https://schema.org/",
    "separator": "cr:separator",
    "source": "cr:source",
    "subField": "cr:subField",
    "transform": "cr:transform"
  },
  "@type": "sc:Dataset",
  "conformsTo": "http://mlcommons.org/croissant/1.0",
  "name": "HuBMAP_CODEX_Spleen_HBM548.TSMP.663",
  "description": "CODEX (co-detection by indexing) multiplexed imaging of human spleen tissue from an 18-year-old white male, published by the Human BioMolecular Atlas Program (HuBMAP). Processed with the HuBMAP codex-pipeline and SPRM, the dataset contains multi-channel OME-TIFF expression images, cell/nucleus segmentation masks, and per-cell feature tables with spatial coordinates.",
  "url": "https://portal.hubmapconsortium.org/browse/dataset/afbd95090454842570908a76c80d1266",
  "identifier": "HBM548.TSMP.663",
  "sameAs": "https://doi.org/10.35079/HBM548.TSMP.663",
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "creator": {
    "@type": "sc:Organization",
    "name": "University of Florida TMC (HuBMAP)",
    "url": "https://hubmapconsortium.org/"
  },
  "publisher": {
    "@type": "sc:Organization",
    "name": "HuBMAP Consortium"
  },
  "datePublished": "2022-06-27",
  "version": "1.0.0",
  "keywords": [
    "CODEX",
    "multiplexed imaging",
    "spleen",
    "spatial biology",
    "single cell",
    "HuBMAP"
  ],
  "citeAs": "HuBMAP Consortium. CODEX data from the spleen of a 18-year-old white male (HBM548.TSMP.663). HuBMAP Data Portal, 2022. https://doi.org/10.35079/HBM548.TSMP.663",
  "isLiveDataset": false,
  "distribution": [
    {
      "@type": "cr:FileObject",
      "@id": "expr-image",
      "name": "expr-image",
      "description": "Multi-channel OME-TIFF expression image for region 1; one channel per antibody marker.",
      "contentUrl": "https://assets.hubmapconsortium.org/afbd95090454842570908a76c80d1266/pipeline_output/expr/reg001_expr.ome.tiff",
      "encodingFormat": "image/tiff"
    },
    {
      "@type": "cr:FileObject",
      "@id": "segmentation-mask",
      "name": "segmentation-mask",
      "description": "Cell and nucleus segmentation mask aligned to the region-1 expression image.",
      "contentUrl": "https://assets.hubmapconsortium.org/afbd95090454842570908a76c80d1266/pipeline_output/mask/reg001_mask.ome.tiff",
      "encodingFormat": "image/tiff"
    },
    {
      "@type": "cr:FileSet",
      "@id": "sprm-cell-tables",
      "name": "sprm-cell-tables",
      "description": "SPRM per-cell feature tables (mean/total marker intensity, cell centers, cluster assignments).",
      "encodingFormat": "text/csv",
      "includes": "sprm_outputs/reg001_expr.ome.tiff-*.csv"
    },
    {
      "@type": "cr:FileObject",
      "@id": "cell-mean-intensity",
      "name": "cell-mean-intensity",
      "description": "Per-cell mean marker intensity table produced by SPRM (one row per segmented cell).",
      "containedIn": {
        "@id": "sprm-cell-tables"
      },
      "encodingFormat": "text/csv",
      "path": "sprm_outputs/reg001_expr.ome.tiff-cell_channel_mean.csv"
    },
    {
      "@type": "cr:FileObject",
      "@id": "cell-centers",
      "name": "cell-centers",
      "description": "Per-cell centroid coordinates produced by SPRM.",
      "containedIn": {
        "@id": "sprm-cell-tables"
      },
      "encodingFormat": "text/csv",
      "path": "sprm_outputs/reg001_expr.ome.tiff-cell_centers.csv"
    }
  ],
  "recordSet": [
    {
      "@type": "cr:RecordSet",
      "@id": "cell_centers",
      "name": "cell_centers",
      "description": "One record per segmented cell, with the cell's ID and centroid position in the image.",
      "field": [
        {
          "@type": "cr:Field",
          "@id": "cell_centers/cell_id",
          "name": "cell_id",
          "description": "Integer label matching the region of this cell in the segmentation mask.",
          "dataType": "sc:Integer",
          "source": {
            "fileObject": {
              "@id": "cell-centers"
            },
            "extract": {
              "column": "ID"
            }
          }
        },
        {
          "@type": "cr:Field",
          "@id": "cell_centers/x",
          "name": "x",
          "description": "X coordinate of the cell centroid, in pixels.",
          "dataType": "sc:Float",
          "source": {
            "fileObject": {
              "@id": "cell-centers"
            },
            "extract": {
              "column": "x"
            }
          }
        },
        {
          "@type": "cr:Field",
          "@id": "cell_centers/y",
          "name": "y",
          "description": "Y coordinate of the cell centroid, in pixels.",
          "dataType": "sc:Float",
          "source": {
            "fileObject": {
              "@id": "cell-centers"
            },
            "extract": {
              "column": "y"
            }
          }
        }
      ]
    },
    {
      "@type": "cr:RecordSet",
      "@id": "cell_mean_intensity",
      "name": "cell_mean_intensity",
      "description": "One record per segmented cell, with mean intensity for two representative markers. Add one field per channel present in the CSV header.",
      "field": [
        {
          "@type": "cr:Field",
          "@id": "cell_mean_intensity/cell_id",
          "name": "cell_id",
          "description": "Integer cell label; joins to cell_centers.cell_id.",
          "dataType": "sc:Integer",
          "source": {
            "fileObject": {
              "@id": "cell-mean-intensity"
            },
            "extract": {
              "column": "ID"
            }
          }
        },
        {
          "@type": "cr:Field",
          "@id": "cell_mean_intensity/cd8",
          "name": "CD8",
          "description": "Mean CD8 marker intensity for the cell (T-cell marker).",
          "dataType": "sc:Float",
          "source": {
            "fileObject": {
              "@id": "cell-mean-intensity"
            },
            "extract": {
              "column": "CD8"
            }
          }
        },
        {
          "@type": "cr:Field",
          "@id": "cell_mean_intensity/cd21",
          "name": "CD21",
          "description": "Mean CD21 marker intensity for the cell (follicular dendritic cell / B-cell zone marker).",
          "dataType": "sc:Float",
          "source": {
            "fileObject": {
              "@id": "cell-mean-intensity"
            },
            "extract": {
              "column": "CD21"
            }
          }
        }
      ]
    }
  ]
}

Reading the file

A few things are worth pointing out in the example above.

  • The top block is just schema.org. name, description, url, license, creator, datePublished, and citeAs are exactly what a data catalog or Google Dataset Search expects. This is why a Croissant file doubles as ordinary dataset metadata—and here every one of those values is the real, published HuBMAP record.

  • distribution describes the physical files. The expression image and the segmentation mask are individual FileObjects addressed by their real pipeline paths (pipeline_output/expr/reg001_expr.ome.tiff, pipeline_output/mask/reg001_mask.ome.tiff). The SPRM per-cell CSVs are grouped as a FileSet (glob sprm_outputs/reg001_expr.ome.tiff-*.csv), with two specific tables pulled out as FileObjects marked containedIn that set.

  • recordSet describes the logical structure. cell_centers turns the centroid CSV into typed cell_id, x, and y fields; cell_mean_intensity exposes per-marker mean intensities (CD8, CD21, …). Each field’s source/extract points back to a specific column, and the shared cell_id lets a consumer join position to expression.

  • Identifiers are stable and real. The HuBMAP ID (HBM548.TSMP.663), the portal UUID, and the DOI all resolve to the same canonical dataset, which is what makes the metadata citable and reproducible.

To use it, save the JSON as croissant.json and load it with the reference implementation:

import mlcroissant as mlc

dataset = mlc.Dataset("croissant.json")
print(dataset.metadata.name)

for record in dataset.records(record_set="cell_centers"):
    print(record["cell_centers/cell_id"], record["cell_centers/x"], record["cell_centers/y"])
    break

The same file would load identically in any other Croissant-aware tool—which is the entire point of the format. (Because the HuBMAP assets gateway needs a token, point the contentUrls at a local copy or an authenticated URL before running the loader end to end.)


Summary

A Croissant file is a JSON-LD wrapper that makes a machine-learning dataset self-describing: schema.org metadata on the outside, an explicit map of files (distribution) and typed records (recordSet) on the inside. Wrapping the real HuBMAP CODEX spleen dataset HBM548.TSMP.663 in Croissant turns a tree of OME-TIFFs and SPRM CSVs into a portable, citable, tool-loadable object—readable by a human, a data catalog, and a training pipeline alike.


References

[1] MLCommons. “Croissant: A Metadata Format for ML-Ready Datasets.” Available at mlcommons.org/croissant.

[2] Mubashara Akhtar et al. “Croissant: A Metadata Format for ML-Ready Datasets.” Proceedings of the 2024 Workshop on Data Management for End-to-End Machine Learning (DEEM ‘24), 2024. doi:10.1145/3650203.3663326

[3] HuBMAP Consortium. “The human body at cellular resolution: the NIH Human Biomolecular Atlas Program.” Nature, 574:187–192, 2019. doi:10.1038/s41586-019-1629-x. Dataset: CODEX data from the spleen of a 18-year-old white male, HBM548.TSMP.663.

[4] HuBMAP Consortium. “codex-pipeline” and “SPRM” (Spatial Process & Relationship Modeling). GitHub: hubmapconsortium/codex-pipeline, hubmapconsortium/sprm.