D Directive Publications Blog Guides for medical & scientific authors

Sharing Clinical Data Openly Without Breaching Privacy

Updated October 08, 2026
A clinical dataset moving through stages of removing identifiers, generalising values and review by a data access committee before approved researchers can use it
Most clinical data are shared through controls, not simply posted online.

Funders and journals want clinical data shared, and participants were promised privacy. Institutions can honour both, but only if they stop calling coded data anonymous.

Key point: Pseudonymised is not anonymous: release each clinical dataset at the most open level that its consent, approvals and a documented re-identification assessment can support.

Openness is a spectrum, not a switch

Most of the trouble comes from treating sharing as a choice between posting everything online and locking everything away. Framed that way, one part of the institution promises a dataset it cannot legally release and another refuses everything to be safe.

Clinical data can be fully public, available to approved researchers under an agreement, analysable only inside a secure environment, or not shared with a stated reason. Place each dataset at the most open level that three things support at once:

  1. What participants consented to.
  2. What the ethics approval and contracts with sponsors or sites allow.
  3. A documented assessment of how identifiable the data would be to the recipient.

Keep the article and the data separate in your guidance. A paper can be open access under CC BY while its participant-level data sit behind an application process. Our piece on open access myths that mislead research offices covers what CC BY allows for articles; none of it decides what happens to patient records.

Many privacy failures start with a word: a consent form or data management plan calls a dataset "anonymised" when it has only been coded.

Under the EU and UK GDPR, pseudonymisation means processing personal data so they can no longer be attributed to a person without additional information, such as a key linking study IDs to record numbers, kept separately (Article 4(5)). Recital 26 says data that could be attributed to a person using that information remain personal data.

Anonymous information does not relate to an identifiable person, judged against all means reasonably likely to be used, by the controller or anyone else. Data protection law does not apply to it, although anonymising is itself processing.

In September 2025, in EDPS v SRB (C-413/23 P), the Court of Justice of the EU confirmed that pseudonymised data are not personal data in every case and for every person: for a recipient who cannot reasonably identify anyone, they may not be. The UK regulator's anonymisation guidance, under review at the time of writing, takes the same "whose hands?" approach. Neither lets you off: if your institution keeps the key, the data remain personal data for you. In July 2026 the European Data Protection Board adopted draft Guidelines 02/2026 on Anonymisation for consultation, built on three tests: no record isolation, no linkage and no inference. Check whether a final version exists.

The US uses different vocabulary. Under the HIPAA Privacy Rule, which applies to covered entities and their business associates, 45 CFR 164.514 sets out two de-identification methods:

  • Safe Harbor. Remove 18 types of identifiers, including all date elements except year, geographic units smaller than a state (the first three ZIP digits may stay only where that area holds more than 20,000 people), and ages over 89 unless grouped as 90 or older, with no actual knowledge that the rest could identify someone.
  • Expert Determination. A qualified expert applies statistical methods, finds the risk of identification by an anticipated recipient very small, and documents the analysis.
Term What has been done Status
Pseudonymised (GDPR) Identifiers replaced by codes; key kept separately Personal data for the key holder and anyone who can reasonably identify people
Anonymous (GDPR) Anonymous for everyone: no one can identify people by means reasonably likely to be used Outside data protection law
De-identified (HIPAA) Safe Harbor or Expert Determination met No longer protected health information
Limited data set (HIPAA) Direct identifiers removed; dates, town, state and ZIP code may stay Still protected; research, public health or operations only, under a data use agreement

Decision rule: in consent forms, data management plans and availability statements, write "anonymous" only when a documented assessment supports it. Otherwise write "pseudonymised" or "coded".

Why small clinical datasets re-identify

Removing names and record numbers is the easy part. The risk sits in combinations of indirect identifiers. The ICO illustrates this with a table of early COVID-19 admissions in part of South London showing only age band, hospital area and condition: the one patient aged 20 to 29 could be singled out and linked to news reports, because young patients were rare then.

The 2010 BMJ guidance on preparing raw clinical data for publication lists indirect identifiers including place of treatment, sex, rare disease or treatment, ethnicity, occupation, multiple pregnancies, very small numerators or denominators, age, and verbatim responses. Adopt its rule: a dataset with three or more indirect identifiers should be assessed by an independent researcher or ethics committee before publication.

Watch four more routes:

  • Prior knowledge. A relative or colleague may recognise a participant where a stranger could not. The ICO's motivated intruder test asks you to consider such people, alongside journalists.
  • Free text. Adverse event narratives and clinician comments carry names, places and dates that structured de-identification misses.
  • Scans. Image files carry metadata such as names and scan dates, and faces survive in the pixels: a 2019 New England Journal of Medicine letter reported face-recognition software matching faces reconstructed from head MRI scans to photographs of most volunteers tested.
  • Genomic data. A 2013 Science study traced participants in public sequencing projects from surnames inferred through genealogy databases plus age and state.

The EDPB's draft notes that re-identification becomes more likely over time as techniques improve and outside information accumulates, so give every assessment a review date.

Techniques that reduce risk, and what they cost

Start with the dataset the paper needs. The BMJ guidance defines the dataset for publication as the minimum detail needed to reproduce all the numbers reported in the paper, while encouraging sharing of more detailed underlying data where possible. Treat those as two releases, not a choice: the reproduction dataset at the most open level it supports, the richer version through controlled access. GDPR Article 89(1) points the same way, requiring research purposes that can be met with data that do not permit identification to be met that way.

Then combine techniques. The ICO describes two families, generalisation and randomisation, and warns that masking is not anonymisation on its own. Each step costs analytic value, so involve the study statistician.

Technique Clinical example Watch for
Remove direct identifiers Names, record numbers, contact details, device serial numbers Necessary, never sufficient
Generalise Age bands; region instead of postcode; year instead of full date Lost precision for age- or place-dependent outcomes
Suppress small groups Merge sparse categories; cap extreme values Rare-disease cohorts can lose most detail
Shift dates Move each participant's dates by the same random offset Intervals survive; calendar and seasonal analyses do not
Treat unstructured data Remove or code free text; strip image metadata; deface head scans Slow manual review; defacing can degrade some brain measurements

Set suppression thresholds before anyone sees the results. The ICO notes that the NHS standard for publishing health and social care data sets k at five, so every record shares its characteristics with at least four others. Follow your data provider's rule where one exists.

Controlled access when data cannot be made anonymous

Making many clinical datasets anonymous enough for public release strips out what makes them useful, so controlled access is normal, not a failure. The ICO puts it plainly: public release needs a very robust approach, because once data are out you cannot retract them. Releasing to defined groups lets you weigh what those recipients know, and keep more detail.

The Five Safes framework, used by the UK Data Service, is a practical design checklist: safe data, projects, people, settings and outputs. Common models:

  • A repository with a data access committee. At the European Genome-phenome Archive, requests go to the data controller's committee, not the archive, and that committee decides access.
  • Clinical research data-sharing platforms, such as Vivli, through which researchers share and request clinical trial data.
  • Trusted research environments, where approved analysts work inside a secure setting and only checked outputs leave.

Two points matter for institutions. First, the committee must outlast the grant: the archive asks committee contacts to name a replacement before changing institution, and can withdraw datasets whose committee stops responding. Make it a departmental role, not one investigator's inbox. Second, controls are not anonymisation. The EDPB's draft says technical, organisational and contractual measures restrict access but do not by themselves make data anonymous, so treat controlled-access data as personal data unless your assessment says otherwise.

What a data use agreement should cover

A data use agreement turns an access decision into enforceable terms. HIPAA sets a useful floor for limited data sets: the agreement must set permitted uses and users, and bind the recipient to safeguard the data, report unauthorised use, hold its agents to the same terms, and not identify the information or contact the individuals. Build on that:

  • Named purpose and users, with a route for adding team members
  • No linkage to other data without approval
  • Security requirements and no onward sharing
  • Prompt reporting of breaches or suspected re-identification
  • Output rules, such as suppressing small counts before publication
  • An end date, with deletion or return of the data
  • Citation of the dataset by its persistent identifier, and limits on cross-border transfers

Restricted data can still be citable: our guide to citing data, software, and code covers what to deposit openly around it.

What journals ask for in a trial report

Governance has to produce something an author can write down at submission. For trial reports this is a requirement, not a courtesy: the ICMJE recommendations state that manuscripts reporting results of clinical trials must contain a data sharing statement, and that trials beginning enrolment on or after 1 January 2019 must include a data sharing plan in the registration. Registration precedes publication by years, so the plan is due long before the paper.

The statement must say whether individual de-identified participant data will be shared, which data, what other documents come with them (protocol, statistical analysis plan), when and for how long, and who may access them for what analyses by what mechanism. "Undecided" is not acceptable for the first. The first and last are governance answers, so the data custodian and committee chair should supply them before submission. A statement naming the repository, the committee, the agreement and the expected turnaround is checkable; "available from the corresponding author on reasonable request" tells an editor nothing.

Before any plan or statement promises sharing, read the consent form: what did participants agree to about future use, by whom and for what purposes? Consent is not the only route:

  • EU and UK GDPR. Research with health data can rely on legal bases other than consent, using the research condition in Article 9(2)(j) and the safeguards in Article 89(1), with details set by EU or national law. Anonymising needs a lawful basis too, because it is processing.
  • US HIPAA. A covered entity can disclose protected health information for research with the individual's authorisation, under an IRB or privacy board waiver, as a limited data set under a data use agreement, or after de-identification. Other federal and state rules may also apply.

Legal permission is not ethical permission. An ethics committee can set tighter limits than the law, and a consent form promising use by the study team only limits what you can honestly offer. For new studies, write the sharing model into the protocol, participant information and consent form from the start, as the BMJ guidance recommends.

Before any clinical dataset leaves the institution

  1. Consent wording and ethics approval checked against the planned sharing
  2. Release level chosen: public, controlled, secure environment, or not shared with a reason
  3. Pseudonymised data not described as anonymous anywhere
  4. Key file held by a named custodian and never deposited
  5. Indirect identifiers listed; three or more sent for independent review
  6. Small groups, extreme values, dates, free text, scans and genomic data handled
  7. Motivated intruder test documented, with a review date
  8. Data access committee set up as a role with a deputy, and the agreement approved
  9. Data sharing statement matches the real access route

Give researchers one place to ask

The common failure is not a leak. It is a promise made by someone who never asked the data custodian. Name one contact who can answer "what can I offer for this dataset?", and file the answer with the study so the next request gets the same one.

Directive Publications' for-authors page links to its data availability and sharing policy, which authors should read alongside your institution's rules before drafting a statement. For questions about a specific manuscript, contact Directive Publications.

Frequently asked questions

Is de-identified patient data still considered personal data?

Often, yes. Under the EU and UK GDPR, data where identifiers have been replaced with codes and the key is kept separately are pseudonymised and remain personal data for anyone who holds the key or can reasonably obtain it. Whether data are anonymous can depend on whose hands they are in, but the safe working test is anonymous for everyone: no one can link the data to a person using means reasonably likely to be used. Under HIPAA, data de-identified by Safe Harbor or Expert Determination are no longer protected health information, so check which standard applies to your dataset.

Can clinical data be shared without patient consent?

Sometimes, depending on the jurisdiction and the route. HIPAA allows research use with a waiver approved by an IRB or privacy board, as a limited data set under a data use agreement, or after de-identification, and the GDPR allows research on legal bases other than consent, with safeguards and conditions set by EU or national law. Ethics approvals and the original consent form can set tighter limits, so check both before anyone promises to share.

What is the difference between anonymised and pseudonymised data?

Pseudonymised data can no longer be linked to a person without extra information, such as a key file, which is kept separately and protected. Anonymised data cannot be linked to a person using means reasonably likely to be used; regulators accept that this can vary between recipients, but the cautious test institutions should apply is that the data are anonymous for everyone. Pseudonymisation reduces risk but the data stay within data protection law, whereas truly anonymous data fall outside it.

What is a controlled access data repository?

It is a repository that holds data too sensitive or identifiable for public release and gives access only to approved users. Researchers apply, a data access committee checks the request against the consent terms and approvals, and approved users sign a data use agreement. Some services go further and let analysts work only inside a secure environment where outputs are checked before release.

How can small clinical datasets be re-identified?

A combination of ordinary details, such as age, sex, treatment dates, hospital and a rare diagnosis, can single out one person, especially to a relative, colleague or journalist who already knows something about them. Free text, medical images and genomic data add further routes. Generalise ages and dates, suppress small groups and extreme values, and have datasets with several indirect identifiers reviewed independently before release.

Ready to submit with confidence?

Use transparent peer review and clear author guidelines.

Submit your manuscript