You may have heard FAIR mentioned in a grant proposal, a data management plan, or a funder’s requirements document. It stands for Findable, Accessible, Interoperable, and Reusable, and it describes a set of guiding principles for how scientific data — and increasingly, scientific software, code, and workflows — should be managed and shared.
If you are new to research, FAIR is worth understanding from the start. If you have been doing research for years, it may feel like a new label on old practices — but it demands more than that, and this post explains why.
Where FAIR Came From
The FAIR principles were introduced in a landmark 2016 paper published in Scientific Data:
Wilkinson MD, Dumontier M, Aalbersberg IJ, et al. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data. 2016;3:160018. DOI: 10.1038/sdata.2016.18
The paper identified a fundamental problem: scientific data was accumulating at a pace and scale that humans could no longer manage manually. Data was siloed in incompatible formats, locked in lab-specific systems, undocumented, and effectively lost when people left a project. Reproducibility suffered. Knowledge was wasted.
The FAIR principles were designed to solve this — not for humans alone, but explicitly for machines. The goal was data that computational systems could find, retrieve, interpret, and reuse automatically, without requiring a human intermediary at every step. That is a harder bar than it sounds.
The Four Principles
Findable (F)
Data and metadata must be discoverable — by both humans and automated systems.
- (F1) Data and metadata receive a globally unique, persistent identifier (such as a DOI)
- (F2) Data are described with rich metadata
- (F3) Metadata explicitly reference the identifier of the data they describe
- (F4) Data and metadata are registered or indexed in a searchable resource
A dataset that lives on a lab file server with no identifier and no metadata entry in any registry is not findable in any meaningful sense. Findability requires intentional action — not just saving a file somewhere.
Accessible (A)
Once found, data and metadata must be retrievable using open, standardized protocols.
- (A1) Data and metadata are retrievable by their identifier using a standard, open communications protocol
- (A1.1) The protocol is open, free, and universally implementable
- (A1.2) The protocol supports authentication and authorization where necessary
- (A2) Metadata remain accessible even if the data themselves are no longer available
This last point — A2 — is often overlooked. If a dataset is taken down or restricted, the metadata describing it should still exist so that others know it existed and can seek access through proper channels. Accessible does not mean everything must be open; it means access pathways must be clear and functional.
Interoperable (I)
Data must be able to integrate with other datasets and operate within computational workflows.
- (I1) Data and metadata use formal, accessible, shared, and broadly applicable languages for knowledge representation
- (I2) Data and metadata use vocabularies that themselves follow FAIR principles
- (I3) Data and metadata include qualified references to other data
Interoperability is where many scientists struggle. Using CSV files is not enough if the columns have non-standard names, the units are ambiguous, and the ontologies are lab-specific. Interoperability means using shared standards — community-agreed vocabularies, controlled terms, and formats that other tools can parse without custom conversion code.
Reusable (R)
The ultimate goal: data and metadata are described richly enough that they can be replicated and built upon in new contexts.
- (R1) Data and metadata are described with accurate, relevant, and comprehensive attributes
- (R1.1) Data are released with a clear and accessible usage license
- (R1.2) Data are associated with detailed provenance (where did they come from, how were they generated)
- (R1.3) Data meet domain-relevant community standards
Reusability depends on all the others. A dataset is only reusable if someone else can find it, access it, understand its format, and trust its provenance. That requires documentation that goes well beyond a README file.
Why FAIR Matters in Modern Science
Reproducibility. The reproducibility crisis in science is partly a data and methods problem. When researchers cannot access the data or code behind a published result, they cannot verify or build on it. FAIR infrastructure makes reproducibility structurally possible rather than dependent on the goodwill of individual authors.
Efficiency. Researchers spend enormous time redoing work that has already been done — often because the previous data or code were inaccessible or unusable. FAIR data reduces that waste.
Scale. Modern science generates data at volumes no human team can manage manually. Machine-actionable metadata and standardized formats allow automated pipelines to discover, integrate, and process data across thousands of datasets without human intervention at each step.
Equity. FAIR data levels the playing field. A well-documented, openly indexed dataset can be discovered and reused by a researcher at a small institution in a lower-income country just as easily as by a researcher at a large, well-funded university.
Funder requirements. Increasingly, FAIR compliance is not optional. NIH, NSF, and the European Union’s Horizon Europe program now require data management plans that align with FAIR principles. Funding and publication are increasingly contingent on demonstrating that your data are managed responsibly.
FAIR Is Not Just About Data
This is where FAIR becomes genuinely demanding — especially for researchers who have been publishing papers and datasets for years and consider themselves responsible sharers of scientific knowledge.
FAIR applies to software and code. A 2022 paper extended the FAIR principles specifically to research software:
Chue Hong NP, Katz DS, Barker M, et al. Introducing the FAIR Principles for research software. Scientific Data. 2022;9:622. DOI: 10.1038/s41597-022-01710-x
Known as FAIR4RS, this framework recognizes that software has characteristics data does not: it must be executed, it has dependencies, it evolves over time, and reusing it means more than reading it — it means running it. FAIR4RS requires that research software be:
- Deposited with a persistent identifier (e.g., via Zenodo or a software registry)
- Accompanied by machine-readable metadata describing its dependencies and execution environment
- Interoperable with other tools through standard APIs and data formats
- Released with a clear license and documented provenance
A script uploaded to a GitHub repository with no version tag, no license, and no environment specification is not FAIR software, even if it is publicly visible.
FAIR applies to workflows. The computational pipelines that connect data to results — the sequence of tools, parameters, and steps that transform raw inputs into published outputs — are themselves scientific objects that must be findable, accessible, interoperable, and reusable. Workflows should be:
- Stored in version-controlled repositories
- Registered in workflow registries such as WorkflowHub or Dockstore
- Described with machine-readable metadata specifying their inputs, outputs, and dependencies
- Packaged in portable environments (Docker, Singularity/Apptainer containers) that can run consistently across local machines, remote servers, and HPC clusters
FAIR applies to computational environments. On large-scale systems — HPC clusters, cloud platforms, institutional computing infrastructure — the environment in which code runs is as important as the code itself. A workflow that produced a result in 2022 may fail silently in 2026 if the underlying libraries have changed. Container-based environments freeze dependencies and make the execution context itself reproducible across time and systems.
This is not theoretical. Researchers who publish results from HPC workflows without documenting the software environment, module versions, and execution parameters are publishing results that cannot be independently reproduced — regardless of whether the code is available.
What This Means if You Have Been Doing This for a While
If you are an experienced researcher, FAIR may feel like renaming things you already do. You share your data, you post your code on GitHub, you describe your methods in your papers. That is a good start — but FAIR asks for more.
The honest challenges:
Persistent identifiers are not optional. A GitHub repository URL is not a persistent identifier. GitHub can rename repositories, change paths, or go away. A DOI minted through Zenodo, Figshare, or Dryad is. Every dataset, every software release, and every major workflow version should have one.
Metadata must be machine-readable. A well-written methods section is valuable, but it is not machine-readable metadata. FAIR requires structured, standardized metadata that software can parse — ideally in community-standard formats (schema.org, Dublin Core, domain-specific standards like MIMARKS for microbiology or MINSEQE for sequencing experiments).
Licenses must be explicit. “Available upon request” is not a license. “See README” is not a license. Every dataset and every software release must carry a clear, standard license (CC BY, MIT, Apache 2.0, etc.) attached to the data or code itself.
Provenance must be documented. Where did the raw data come from? What version of each tool was used? What parameters were passed? What was the hardware environment? This information needs to be recorded at the time of analysis, not reconstructed from memory at publication time.
FAIR is a culture, not a checklist. The largest barrier to FAIR adoption is not technical — it is cultural. Studies show that nearly 95% of researchers recognize FAIR’s value and 90% are willing to implement it — but only with the right resources and support. When implementation requires their own time and budget, motivation drops sharply. This reflects a real tension: FAIR demands sustained effort that current academic incentive structures do not adequately reward. Credit for data and software stewardship still lags far behind credit for novel publications.
Funders Are Paying Attention
- NSF requires a Data Management and Sharing Plan with all proposals and funds the FAIROS program specifically to advance FAIR open science practices
- NIH requires data management plans aligned with FAIR and mandates immediate public access to publications and supporting data
- Horizon Europe (EU) makes proper FAIR-aligned research data management mandatory for funded projects, with a Data Management Plan required by month six of every project
Non-compliance is increasingly a practical risk, not just an ethical lapse.
References
-
Wilkinson MD, Dumontier M, Aalbersberg IJ, et al. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data. 2016;3:160018. https://doi.org/10.1038/sdata.2016.18
-
Chue Hong NP, Katz DS, Barker M, et al. Introducing the FAIR Principles for research software. Scientific Data. 2022;9:622. https://doi.org/10.1038/s41597-022-01710-x
-
GO FAIR Initiative. FAIR Principles. https://www.go-fair.org/fair-principles/
-
The Turing Way Community. FAIR Principles. https://book.the-turing-way.org/reproducible-research/rdm/rdm-fair/
-
NSF. FAIROS — Findable, Accessible, Interoperable, Reusable Open Science. https://www.nsf.gov/funding/opportunities/fairos-findable-accessible-interoperable-reusable-open-science
-
OpenAIRE. How to comply with Horizon Europe mandates for Research Data Management. https://www.openaire.eu/how-to-comply-with-horizon-europe-mandate-for-rdm
-
WorkflowHub. https://workflowhub.eu
-
Dockstore. https://dockstore.org
-
Cornell University Library. FAIR Data and Reproducibility. https://data.research.cornell.edu/data-management/sharing/fair/
-
Digital Science, Figshare, Springer Nature. State of Open Data 2025. https://stateofopendata.com/