Research
What we found redacting 50,000 real documents
We ran 50,119 real public documents through PiBye's detector and checked every redacted copy with regexes written independently of it. They found 84,037 number- and pattern-shaped identifiers; 24 were still readable afterwards, in 22 documents. Every survivor was number-shaped, mostly statute citations shaped like a SIN. The bigger defect was the opposite one: ordinary words treated as names, which we then cut by 52 percent.
Key takeaway: Automatic detection gets most identifiers, not all of them, and it can also hide too much. Both are reasons a person reviews every copy before it is shared.
Last reviewed · 2 sources
The numbers, from the published totals
All figures come from the battery's own totals file, last updated 2026-09-25. One document counts once: 618 batch reports holding 73,401 records were reconciled down to 50,119 processed documents, excluding 22,489 repeat runs.
| Measure | Value |
|---|---|
| Documents successfully processed | 50,119 |
| Documents attempted | 50,912 |
| Failed attempts (reported separately) | 1,159 |
| Missing-file attempts (reported separately) | 309 |
| Identifiers found by the independent regexes | 84,037 |
| Identifiers still readable after redaction | 24 (0.029% of those found) |
| Documents with a readable identifier | 22 |
| Distinct readable values | 20 |
| Flagged, then judged not to be identifiers | 5 |
| Second-pass candidates in redacted text | 57,216 |
Source: website/data/battery-stats.json, also served at /api/battery-stats with its methodology.
What survived redaction
The 24 readable identifiers broke down as 17 SIN-shaped numbers, 6 phone numbers and 1 SSN-shaped number. All are counted as leaks. The SIN-shaped values are nine-digit runs in public filings, several beside firm telephone numbers, where text extraction can weld a number to the words beside it; the phone numbers were law-firm switchboards. Not every entry has been individually confirmed as harmless, and one is shaped like a US Social Security number.
What the check cannot see matters as much. These regexes look for number- and pattern-shaped identifiers. They do not count personal names that survive, so this battery says nothing about name recall. That gap is one reason PiBye shows every replacement for review instead of sending anything automatically.
The failure we went looking for: removing too much
A battery that only counts leaks rewards a tool for hiding everything, and a document with its ordinary words removed is useless to the AI reading it. So the full detector was run again over the redacted copies: it flagged 57,216 further candidates. These are not leaks, and reading a sample showed most were ordinary prose, with phrases like "almost all" and "typographical error" being treated as names.
On a sample of documents that contain no personal information at all (blank government forms and published statutes), PiBye had been replacing a median 108 name-like spans per thousand words. After the fix, the median is 51, a 52 percent drop, while the separate planted-identifier battery still caught every identifier it planted.
What this does not prove
- Processing volume is not a leak-free rate. An independent regex finds only the identifier shapes it knows.
- The corpus is public court records, government forms, statutes and financial filings. 978 documents from the original list were never processed (186 set aside on purpose); replacements are blank agency forms, which are not equivalent to the filings they replace.
- No client documents are in the corpus. Your documents will differ, especially scans, handwriting and tables.
- The battery does not test every import, sharing or restore workflow in the app.
Where PiBye fits
How PiBye handles this
PiBye publishes these results because the review step is the point: it proposes replacements, you check every one against the document, and nothing becomes an approved copy until you do. The first battery has finished and these totals are final for it; a second battery on harder material is still running and is reported on the test battery page, not here. The method is published with the totals.
1.0.1 · macOS 14.8.5 or later · Apple Silicon · 1.1 GB
Frequently asked questions
Were real client documents used?
No. The corpus is public documents: court records and filings, government forms, statutes and gazettes, and financial filings. No client or customer document is in it.
Why count over-redaction as a defect?
Because an AI can only reason about what it can read. Replacing ordinary words as if they were names makes the copy less useful without protecting anyone.
Does this mean PiBye catches every identifier?
No. 24 identifiers the independent check found were still readable, and the check does not measure names at all. Review every copy before sharing it.
Sources
- Inside the test battery, PiBye. Checked 25 September 2026.
- Battery totals and methodology (JSON), PiBye. Checked 25 September 2026.