Skip to content
PiBye

Research

Pseudonymized vs anonymized: which is your AI-ready copy?

If you can put the real names back, it is pseudonymized, not anonymized. Pseudonymization replaces identifiers with stand-ins and keeps a key; anonymization must be irreversible. The EDPB says pseudonymised data that can be attributed to a person with additional information is still information about an identifiable person, and the ICO says pseudonymous data is still personal data. So a copy with [PERSON_001] tokens is still personal information for you.

Key takeaway: A tokenized copy lowers what the AI provider learns. It does not take the document outside privacy law or professional confidentiality.

Last reviewed · 5 sources

What the regulators say

European Data Protection Board, Guidelines 01/2025: "Pseudonymised data, which could be attributed to a natural person by the use of additional information, is to be considered information on an identifiable natural person", and this holds "if pseudonymised data and additional information are not in the hands of the same person."

UK Information Commissioner's Office: "pseudonymous data is still personal data, as the data can be combined with additional information allowing the identification of people it relates to."

Québec's private-sector act sets the anonymization bar high: information is anonymized only if "it is, at all times, reasonably foreseeable in the circumstances that it irreversibly no longer allows the person to be identified directly or indirectly." A copy you can restore is, by design, reversible.

Side by side

Pseudonymization and anonymization compared
Pseudonymized copyAnonymized data
Can the real details be put back?Yes, by whoever holds the keyNo, by definition
Still personal information?Yes (EDPB, ICO)No, if the bar is truly met
Useful for drafting a client document?Yes: the draft comes back and names are restoredOnly for general analysis; the output cannot be tied back to the client
Main riskThe key, and details left in the textRe-identification from what remains

Why the key matters

NIST SP 800-188 describes the lookup table that maps identifiers to pseudonyms and warns that releasing it "will result in compromised identities. Thus, the lookup table or the information for the transformation must be highly protected." The EDPB adds that lookup tables "are personal data since they allow the identification of data subjects."

That is the practical line for AI work: keep the key where the AI provider cannot reach it, and do not derive tokens from the values they replace.

Which word to use

Call a reversible, tokenized copy pseudonymized or de-identified. NIST SP 800-188 goes further and recommends avoiding "anonymization" altogether because of inconsistent definitions. Searchers still type "anonymize a document for ChatGPT"; the accurate answer is that you are pseudonymizing it, which is what makes restoring the names possible.

Where PiBye fits

How PiBye handles this

PiBye pseudonymizes on purpose: names, ID numbers and addresses become stable tokens such as [PERSON_001], the key that reverses them stays on your Mac, and the real details go back into the returned document there. The AI sees the case, not the client. Because the copy is reversible for you, your confidentiality and consent obligations still apply to it.

PiBye restoring real names into a returned AI draft on the Mac
PiBye restoring real names into a returned AI draft on the Mac

Actual PiBye views rendered at 3× resolution with fictional sample data. Open full-resolution image ↗

1.0.1 · macOS 14.8.5 or later · Apple Silicon · 1.1 GB

Frequently asked questions

Can the AI provider reverse PiBye's tokens?

Not from the copy: the tokens are stand-ins, not encodings of the values, and the key stays on your Mac. Details you chose to keep, and facts in the text, are still visible to the provider.

Is de-identified the same as anonymized?

No. De-identification is the general process of removing identifying information; the result can still be reversible or re-identifiable. Anonymization is the irreversible end of that range.

Sources

  1. Guidelines 01/2025 on Pseudonymisation (version for public consultation), European Data Protection Board, January 16, 2025. Checked 25 September 2026.
  2. How do we ensure anonymisation is effective?, Information Commissioner's Office (UK). Checked 25 September 2026.
  3. Act respecting the protection of personal information in the private sector (CQLR c P-39.1), LégisQuébec, Government of Québec, up to date as of June 10, 2026. Checked 25 September 2026.
  4. NIST SP 800-188: De-Identifying Government Datasets, National Institute of Standards and Technology, September 2023. Checked 25 September 2026.
  5. Opinion 05/2014 on Anonymisation Techniques (WP216), Article 29 Data Protection Working Party, April 10, 2014. Checked 25 September 2026.