icaoberg / MLCommons Croissant

Created Sun, 26 Apr 2026 00:00:00 +0000 Modified Sun, 26 Apr 2026 21:06:34 -0400

MLCommons Croissant

Anyone who has worked with multiple ML datasets knows the friction: every dataset comes with its own folder structure, its own README conventions, its own way of describing splits and features. You download a dataset, and before you can use it you have to read the docs, figure out the schema, and write custom loading code — then do it all over again for the next one.

Croissant is a metadata specification designed to fix that. Developed under the MLCommons umbrella, it gives datasets a shared vocabulary so that tools, search engines, and researchers can understand a dataset’s structure without custom per-dataset knowledge.


What Croissant is

Croissant is a metadata format — a structured, machine-readable way to describe a machine learning dataset. It defines what a dataset contains, how it’s organized, and how its fields map to ML concepts like splits, labels, and features.

It builds on schema.org, the widely-used vocabulary that powers structured web data, and extends it with ML-specific concepts. The result is a specification that:

  • Describes the dataset’s files, their formats, and how they relate to each other
  • Documents the ML-specific structure: training/validation/test splits, feature names and types, label definitions
  • Provides provenance and licensing information in a standard location
  • Is human-readable (JSON-LD) and machine-parseable by supporting tools

A Croissant file is typically a single JSON-LD document that travels with the dataset — or lives in a registry — and tells any Croissant-aware tool everything it needs to know to load and work with the data.


What it represents

Croissant represents a shift toward treating dataset documentation as infrastructure rather than an afterthought. The ML community has long had a reproducibility problem, and poor dataset documentation is a major contributor: unclear splits, undocumented preprocessing, ambiguous label definitions.

By standardizing how this information is expressed, Croissant makes datasets more FAIR — Findable, Accessible, Interoperable, and Reusable. A dataset with a Croissant metadata file can be indexed by search engines, automatically loaded by ML frameworks, and audited for bias and provenance — without any dataset-specific code.

It also reflects a broader effort within MLCommons to bring the kind of standardization to data that benchmarks like MLPerf brought to model evaluation.


What it’s used for

In practice, Croissant metadata serves several distinct purposes:

  • Discovery — search engines (including Google Dataset Search) can index Croissant files and surface datasets based on their features, tasks, and structure
  • Automatic loading — ML frameworks can read a Croissant file and load the dataset without custom preprocessing code
  • Documentation — researchers and auditors can inspect the file to understand splits, label definitions, feature types, and data provenance
  • Interoperability — a dataset described in Croissant can be consumed by any Croissant-aware tool, regardless of the framework the dataset creator originally had in mind
  • Reproducibility — fixed, versioned metadata makes it easier to reproduce experiments that depend on specific dataset configurations

Projects and platforms using Croissant

Croissant has seen adoption across major ML dataset platforms and tools:

  • Hugging Face — datasets on the Hub can export Croissant metadata; the Hub automatically generates it for supported datasets
  • Kaggle — Kaggle datasets support Croissant export, enabling cross-platform discovery
  • Google Dataset Search — indexes Croissant metadata for discovery across the open web
  • OpenML — integrating Croissant as part of its dataset metadata layer
  • TensorFlow Datasets (TFDS) — exploring Croissant as a loading interface alongside its existing catalog
  • MLCommons — the originating organization, uses Croissant for its own benchmark datasets

The specification is also supported by the mlcroissant Python library, which allows dataset creators to author and validate Croissant files programmatically and consumers to load Croissant-described datasets directly into Python.


Getting started

To explore or validate a Croissant file, the official Python library is the easiest entry point:

pip install mlcroissant

To load a dataset described by a Croissant file:

import mlcroissant as mlc

ds = mlc.Dataset("path/to/croissant.json")
for record in ds.records(record_set="train"):
    print(record)

Croissant files can also be authored by hand or generated from existing dataset metadata using the library’s builder API. The Croissant editor on Hugging Face Spaces provides a no-code interface for creating and validating metadata files.


Further reading