Skip to content
PiBye
← Back to PiBye

The test record

Inside the test battery.

We ran 50,119 real documents through PiBye, checked the results with regexes written independently of the detector, and published what survived. Including the parts that did not go our way.

Put through the battery

Real documents.
A growing test record.

Court records, public filings, and test documents. Every document in this count completed processing in the battery. Repeat runs count once.

50,119

documents through the test battery

Successfully processed50,000 document target

Verified against completed checkpoints through . This measures testing volume, not a leak-free success rate.

Count documents, not repeat runs.

The battery ran for three weeks across two Macs, and documents get re-run after a crash, a machine move or a rebalance. This snapshot reconciles 618 batch reports holding 73,401 records down to the 50,119 documents that were actually processed, excluding 22,489 repeat runs. It records 1,159 failed attempts and 309 missing-file attempts separately rather than folding them in.

An auditable snapshot.

The first corpus list held 50,000 documents and 978 of them were never successfully processed. 186 we set aside on purpose, single filings that pinned a parser for hours, which we quarantined rather than quietly drop. The rest failed extraction, mostly damaged or image-only PDFs. Rather than round up, we processed replacements drawn from corpus files the original list had not used, which is how the total passes 50,000. Those replacements are blank agency forms across nine authorities, and they are not equivalent to the large financial filings they stand in for. Your own documents are not in this corpus and never were.

Inspect the totals and methodology (JSON) ↗

What the battery found

84,037 identifiers found. 24 still readable.

A second set of regexes, written without reference to the detector, reads every document before and after redaction. Across 50,119 documents it found 84,037 identifiers. After redaction, 24 were still readable, across 22 documents and 20 distinct values.

phone
6
sin like
17
ssn like
1

We count all 24 as leaks. Every one is number-shaped and all come from the public corpus, mostly court filings. The SIN-shaped values are nine-digit runs, and several sit beside firm telephone numbers, where extracted text can weld a number to the words beside it; we have not individually confirmed that each of the 24 is harmless, and one is shaped like a US Social Security number. A further 5 were flagged and then judged not to be identifiers at all, such as a citation shaped like a number, and are not in the 24. These regexes check number- and pattern-shaped identifiers; they do not count personal names that survive, which is one reason you review every copy before sharing it.

The failure we went looking for

Removing too much is a defect too.

Every battery until now measured only what escapes. That rewards a tool for hiding everything, and a redacted document you cannot reason about is useless to the AI you are sending it to. So we built the opposite measurement: run the full detector again over the redacted copy and count what it still wants to hide. It returned 57,216 candidates. These are not leaks: they are what the detector would still replace if it ran again over text it had already redacted, and reading a sample showed most were ordinary prose, so the problem was ours. They were phrases like “almost all”, “examined whether” and “typographical error”. Ordinary words, being treated as names.

On a sample of documents that contain no personal information at all, blank government forms and published statutes, we had been replacing a median 108 name-like spans per thousand words. After the fix that number is 51, a drop of 52 percent, with no loss of protection: the adversarial battery still catches every planted identifier, the scanned identity-document battery reports zero leaks on fresh seeds, and the full test suite passes. In real court filings the only names that stopped being redacted were words like “Accordingly” and “officers”.

Second battery · in progress

Harder material, not finished.

The first fifty thousand were mostly blank forms and court filings. A second battery of up to 50,000 documents is aimed at what that one did not test: completed forms, handwriting, French and Spanish gazettes, accounting and audit material. It is in progress, and no conclusion should be drawn from it yet. So far 685 documents have been processed (220 more failed extraction, mostly large scans that hit a 15-minute limit, and are counted separately); none had an identifier still readable afterwards. The form-fill slice below is counted on its own. The count was last reconciled on .

Alongside it, a slice of about 1,360 government forms was filled with more than 6,000 planted identifiers and redacted. No planted name, address, phone number, email, date of birth, SIN, passport or account number was still readable afterwards. That slice did find a real leak, Canadian postal codes, and it is fixed in PiBye 0.4.3. Values PiBye keeps on purpose, such as province, monetary amounts and generic occupations, are not counted as leaks. Planted values are known to the test, so this shows the detector finds what it is told to look for in these layouts, not that every real document is clean.

What we test against.

Public court records and filings, government forms both blank and completed, statutes and gazettes, financial filings, and scanned pages that have to go through OCR. Alongside them we run two synthetic batteries: one that plants known identifiers in the layouts real files use, and one that renders identity documents the way immigration and licensing authorities print them, then degrades them the way a real scan does.

Where this is going

One million documents, scanned and probed for leaks.

The first fifty thousand were the start. We are working toward one million real documents, every one run through PiBye and checked afterwards by independent readers for anything that should have been replaced. Each slice is chosen to test something the last one did not: completed forms rather than blank ones, handwriting, French and Spanish, accounting and audit material, and documents dense with proper nouns that should never be redacted.

50,804 of 1,000,000 documents processed so far5.1%

The first fifty thousand plus the second battery so far. Counted from checkpoints, last reconciled 2026-10-05.

Every finding becomes an update. When a slice shows PiBye missing something, or removing something it should have kept, we fix it and release the fix. PiBye checks for updates each time it opens, and you choose when to install. Findings from each slice are published here whether or not they flatter us, so this page will keep changing, and the product will keep getting better.

Earlier detector report · 9 September 2026

An earlier, separate measurement.

An earlier report covered 21,631 documents, found 71,304 identifiers, and reported 147 documents with readable leftovers. It used a broader definition of a leftover than the figures above, so the two are not comparable. It is kept here rather than replaced.

Inspect the earlier detector report (JSON) ↗

What this does not prove. Processing volume is not a leak-free success rate, and an independent regex only finds the identifier shapes it knows. These tests do not validate every import, sharing or restoration workflow in the app. Read the approved copy before you send it. For the findings as a sourced summary, and what published standards say identifies a person, see our research.