D Directive Publications Blog Guides for medical & scientific authors

FAIR Data Principles Explained for Researchers

Updated October 08, 2026

FAIR is not a badge or a licence. It is 15 short principles about identifiers, metadata, access rules and provenance that let people and machines find and reuse data.

Key point: FAIR does not mean open. A restricted clinical dataset can be highly FAIR if its metadata are public, its access route is clear and its reuse terms are stated.

Where the FAIR data principles come from

The FAIR data principles were published in 2016 as "The FAIR Guiding Principles for scientific data management and stewardship" by Mark Wilkinson and colleagues in Scientific Data. They grew out of a 2014 workshop in Leiden, Netherlands, called "Jointly Designing a Data Fairport," which brought together "a wide group of academic and private stakeholders."

The acronym stands for findable, accessible, interoperable and reusable. Behind those four words sit 15 short principles, each with an identifier such as F1 or R1.2. The original paper is open access and short enough to read in one sitting, and it is worth reading before you accept anyone's summary of it, including this one.

Two features of the paper explain most of what follows.

Machines are a primary audience. The authors state that, distinct from initiatives that focus on the human scholar, "the FAIR Principles put specific emphasis on enhancing the ability of machines to automatically find and use the data." A spreadsheet that a colleague can understand after an email exchange is not FAIR in this sense. A dataset that software can locate, retrieve and interpret with minimal human help moves closer.

They are principles, not a standard. The paper says they "precede implementation choices, and do not suggest any specific technology, standard, or implementation-solution." There is no single FAIR file format and no FAIR certificate. FAIRness is a matter of degree, and the authors apply the idea to all scholarly digital objects, "from data to analytical pipelines."

FAIR is not the same as open

This is the misunderstanding that causes the most trouble in medical research. Principle A1.2 says the retrieval protocol "allows for an authentication and authorization procedure, where necessary." Principle A2 says metadata should remain accessible "even when the data are no longer available."

The paper addresses sensitive data directly: for highly sensitive or personally identifiable data, publishing rich metadata to support discovery can provide a high degree of FAIRness even when the data themselves are not openly published. A controlled-access patient dataset with a public record, a persistent identifier, a documented application process and clear reuse conditions can be more FAIR than an open spreadsheet with no description posted on a lab website.

So the question for a clinical team is rarely "open or FAIR?" It is "what can we release, and how do we make the rest discoverable and requestable?" Our guide to sharing clinical data openly without breaching privacy covers the release decision itself.

The 15 principles in plain language

ID Principle (shortened) What it asks of you
F1 Globally unique, persistent identifier Deposit where the dataset gets a persistent identifier such as a DOI, not just a web address
F2 Rich metadata Describe the data well enough for a stranger to judge relevance
F3 Metadata include the data identifier The record and the files point to each other unambiguously
F4 Registered or indexed in a searchable resource The record is visible to search services, not buried in a folder
A1 Retrievable by identifier via a standard protocol Resolving the identifier leads to the data or to instructions for getting them
A1.1 Open, free, universally implementable protocol Use standard web protocols, not a proprietary download tool
A1.2 Authentication and authorisation where necessary Restricted access is allowed if the procedure is defined
A2 Metadata persist when data are gone Keep the record live after deletion or withdrawal
I1 Formal, shared language for knowledge representation Use machine-readable, documented formats and metadata schemas
I2 FAIR vocabularies Use established, documented terminologies
I3 Qualified references to other (meta)data Links say how objects relate, such as "is supplement to"
R1 Richly described with accurate, relevant attributes Document methods, variables, units and context
R1.1 Clear, accessible usage licence State reuse terms explicitly
R1.2 Detailed provenance Record where the data came from and how they were processed
R1.3 Domain-relevant community standards Follow the conventions of your field

The table paraphrases. For exact wording, use the paper's Box 2.

Findable and accessible in practice

Findability starts with the repository you choose. A general-purpose or institutional repository that mints DOIs handles F1, F3 and F4 for you, provided you fill in the record properly. If you are unsure where to deposit, ask your library or research data service which repositories your funder and your field recognise.

F2 is where most deposits fall short. Titles like "Dataset 1" and descriptions that say "data for the paper" defeat search. Write the description for someone who has not read your article: the population or sample, the measurement, the time period, the file contents and the variables.

For accessibility, the test is simple. Paste the identifier into a browser. Does it lead either to the files or to an unambiguous statement of how to request them, from whom and under what conditions? If the answer is "email the corresponding author," consider whether that address will still work in ten years and whether a data access committee or repository process would be more durable.

Plan for A2 now. If data must be deleted under a retention schedule or withdrawn after a consent change, the metadata record should stay public with a note explaining why the files are unavailable.

Interoperable and reusable in practice

Interoperability means your data can be combined with other data without manual guesswork. In practice:

  • Prefer open, documented file formats (CSV with a data dictionary, for example) over proprietary formats where you have a choice.
  • Use established terminologies where your field has them. Medical datasets commonly draw on vocabularies such as MeSH for subject indexing or LOINC for laboratory observations. Check licensing terms before redistributing any terminology content itself.
  • Give variables explicit names, units and code lists rather than relying on column headers only you understand.
  • Use typed links (I3) between the dataset, the article, the code and the protocol, so a reader or a machine knows which object supports which.

Reusability depends on three things that are easy to skip at the end of a project.

A licence or access statement (R1.1). Without stated terms, a careful reuser cannot tell what is permitted. Restricted data still need a statement of the conditions.

Provenance (R1.2). Record the source of the data, collection dates, processing steps, software versions and which version of the dataset the article used. Linking the code that produced the analysis is part of this. Our guide to citing data, software and code explains how to give each object its own citation.

Community standards (R1.3). Many fields have reporting conventions and data standards. The FAIRsharing registry catalogues data and metadata standards, databases and policies, and is a sensible first stop when you do not know what your field expects.

What funders and institutions expect

Policies refer to FAIR in different ways, so read the wording for each award.

The European Commission expects Horizon Europe beneficiaries to manage research data in line with the FAIR principles under the principle "as open as possible, as closed as necessary." The Commission's Horizon Europe Programme Guide sets out the obligations, including a data management plan due by month six of the project and deposit in a trusted repository.

In the United States, NIH's Data Management and Sharing Policy asks researchers to plan for data sharing, and its supplemental guidance on selecting a repository notes that deposit in a quality repository generally improves the FAIRness of data and sets out desirable repository characteristics consistent with FAIR.

Journals add their own layer through data availability statements. If you report a clinical trial, our guide to data sharing statements for clinical trials covers what editors check.

FAIR is also not the only framework that applies. For data about Indigenous Peoples, the CARE Principles for Indigenous Data Governance (collective benefit, authority to control, responsibility and ethics) were written to complement FAIR by addressing rights and power that a data-centric approach does not. Their authors summarise the goal as "Be FAIR and CARE."

For research offices that want to measure FAIRness consistently, the Research Data Alliance published the FAIR Data Maturity Model (version 1.0, 2020), a set of assessment indicators intended for reuse across evaluation tools.

A FAIR data checklist before you deposit

  1. The dataset is in a repository that assigns a persistent identifier
  2. The title and description would make sense to someone who never read the article
  3. Creators have ORCID iDs and affiliations are recorded consistently
  4. The metadata record will stay public even if files are restricted or later removed
  5. Access conditions and the request route are stated, if data are not open
  6. Files use open or widely supported formats, with a data dictionary for every variable
  7. Units, codes and terminologies are documented
  8. A licence or reuse statement is attached to this exact version
  9. Provenance is recorded: source, dates, processing steps and software versions
  10. The dataset, article, code and protocol link to each other with identifiers
  11. The relevant community standard for your data type has been checked

Common misreadings to correct

Misreading Better reading
"FAIR means we must make patient data public" FAIR allows authenticated, controlled access
"Uploading to a repository makes data FAIR" A repository helps, but poor metadata still limits findability and reuse
"Supplementary files on the journal site are enough" They may lack their own identifier, licence or rich metadata
"FAIR is a certification" It is a set of principles; FAIRness is a matter of degree
"FAIR is only about data" The authors apply it to all research objects, including workflows and code

The practical takeaway for institutions is to make FAIR a planning question, not a deposit-day task. Identifiers, metadata fields, consent wording and file formats are far cheaper to get right at the start of a study than to reconstruct after publication.

Frequently asked questions

What are the FAIR data principles?

They are 15 guiding principles, published in Scientific Data in 2016, describing how to make research data and metadata findable, accessible, interoperable and reusable. They cover persistent identifiers, rich metadata, standard retrieval protocols, shared vocabularies, licences, provenance and community standards. They put particular weight on making data usable by machines as well as by people.

Does FAIR data have to be open data?

No. The principles allow for authentication and authorisation where necessary, and the original paper notes that sensitive data can achieve a high degree of FAIRness through rich public metadata even when the data themselves are not openly published. Access can be restricted as long as the route to request it is clear.

Is FAIR a standard I can be certified against?

No. The authors describe FAIR as high-level principles that do not prescribe any specific technology, standard or implementation. Assessment frameworks such as the Research Data Alliance FAIR Data Maturity Model help evaluate how FAIR a dataset is, but FAIRness is a matter of degree rather than a pass or fail certificate.

Do funders require FAIR data?

Some do, in different ways. Horizon Europe expects beneficiaries to manage research data in line with the FAIR principles and to submit a data management plan, while NIH encourages deposit in repositories whose characteristics are consistent with FAIR. Requirements vary by funder and change over time, so check the current policy for each award.

What is the quickest way to make a dataset more FAIR?

Deposit it in a suitable repository that assigns a persistent identifier, then write metadata that let a stranger understand and reuse the files. Add a clear licence or access statement, use standard file formats and vocabularies, and link the dataset to the related article, code and protocol.

Ready to submit with confidence?

Use transparent peer review and clear author guidelines.

Submit your manuscript