Skip to content
PiBye

Research

What we found redacting 50,000 real documents

We ran 50,119 real public documents through PiBye's detector and checked every redacted copy with regexes written independently of it. They found 84,037 number- and pattern-shaped identifiers; 24 were still readable afterwards, in 22 documents. Every survivor was number-shaped, mostly statute citations shaped like a SIN. The bigger defect was the opposite one: ordinary words treated as names, which we then cut by 52 percent.

Key takeaway: Automatic detection gets most identifiers, not all of them, and it can also hide too much. Both are reasons a person reviews every copy before it is shared.

Last reviewed · 2 sources

The numbers, from the published totals

All figures come from the battery's own totals file, last updated 2026-09-25. One document counts once: 618 batch reports holding 73,401 records were reconciled down to 50,119 processed documents, excluding 22,489 repeat runs.

PiBye test battery totals
MeasureValue
Documents successfully processed50,119
Documents attempted50,912
Failed attempts (reported separately)1,159
Missing-file attempts (reported separately)309
Identifiers found by the independent regexes84,037
Identifiers still readable after redaction24 (0.029% of those found)
Documents with a readable identifier22
Distinct readable values20
Flagged, then judged not to be identifiers5
Second-pass candidates in redacted text57,216

Source: website/data/battery-stats.json, also served at /api/battery-stats with its methodology.

What survived redaction

The 24 readable identifiers broke down as 17 SIN-shaped numbers, 6 phone numbers and 1 SSN-shaped number. All are counted as leaks. The SIN-shaped values are nine-digit runs in public filings, several beside firm telephone numbers, where text extraction can weld a number to the words beside it; the phone numbers were law-firm switchboards. Not every entry has been individually confirmed as harmless, and one is shaped like a US Social Security number.

What the check cannot see matters as much. These regexes look for number- and pattern-shaped identifiers. They do not count personal names that survive, so this battery says nothing about name recall. That gap is one reason PiBye shows every replacement for review instead of sending anything automatically.

The failure we went looking for: removing too much

A battery that only counts leaks rewards a tool for hiding everything, and a document with its ordinary words removed is useless to the AI reading it. So the full detector was run again over the redacted copies: it flagged 57,216 further candidates. These are not leaks, and reading a sample showed most were ordinary prose, with phrases like "almost all" and "typographical error" being treated as names.

On a sample of documents that contain no personal information at all (blank government forms and published statutes), PiBye had been replacing a median 108 name-like spans per thousand words. After the fix, the median is 51, a 52 percent drop, while the separate planted-identifier battery still caught every identifier it planted.

What this does not prove

  • Processing volume is not a leak-free rate. An independent regex finds only the identifier shapes it knows.
  • The corpus is public court records, government forms, statutes and financial filings. 978 documents from the original list were never processed (186 set aside on purpose); replacements are blank agency forms, which are not equivalent to the filings they replace.
  • No client documents are in the corpus. Your documents will differ, especially scans, handwriting and tables.
  • The battery does not test every import, sharing or restore workflow in the app.

Where PiBye fits

How PiBye handles this

PiBye publishes these results because the review step is the point: it proposes replacements, you check every one against the document, and nothing becomes an approved copy until you do. The first battery has finished and these totals are final for it; a second battery on harder material is still running and is reported on the test battery page, not here. The method is published with the totals.

PiBye review screen listing each detected identifier for approval
PiBye review screen listing each detected identifier for approval

Actual PiBye views rendered at 3× resolution with fictional sample data. Open full-resolution image ↗

1.0.1 · macOS 14.8.5 or later · Apple Silicon · 1.1 GB

Frequently asked questions

Were real client documents used?

No. The corpus is public documents: court records and filings, government forms, statutes and gazettes, and financial filings. No client or customer document is in it.

Why count over-redaction as a defect?

Because an AI can only reason about what it can read. Replacing ordinary words as if they were names makes the copy less useful without protecting anyone.

Does this mean PiBye catches every identifier?

No. 24 identifiers the independent check found were still readable, and the check does not measure names at all. Review every copy before sharing it.

Sources

  1. Inside the test battery, PiBye. Checked 25 September 2026.
  2. Battery totals and methodology (JSON), PiBye. Checked 25 September 2026.