GUIDE · VERIFICATION
How to Verify a PDF Redaction Actually Removed the Text
Short answer: a redaction worked only if the hidden text is gone from the file, not just covered. To check, open the redacted PDF and try to select or copy the text under each black box, search for the exact words you redacted, extract all of the text with a separate tool, and inspect the metadata, comments, attachments and bookmarks. If any redacted word can still be found, copied or extracted, the redaction failed.
Black boxes fail more often than people expect. A rectangle drawn over text only covers it; the characters underneath stay in the file and come right back with a copy-paste. That is how "redacted" court filings end up in the news. The checks below take about five minutes per document and work with free tools.
Six checks you can run on any redacted PDF
- Select and copy under the box. Drag your cursor across a black box, copy, and paste into a plain text editor. If words appear, the text was only covered.
- Search for what you redacted. Use the viewer's Find (Ctrl+F / Cmd+F) for each name, number or phrase you removed. A hit means it is still in the text layer.
- Extract all of the text with a second tool. Viewers can hide things. A text extractor such as
pdftotext(free, part of Poppler) dumps every character in the file. Search that output for the redacted terms. - Check the metadata. Open the document properties (title, author, subject, keywords). Redacted names often survive in a title like "Smith, J. – discipline file." Also check the XMP metadata if your tool shows it.
- Check comments, attachments, bookmarks and form fields. Sticky notes, embedded files, bookmark titles, links and filled-in form fields can each carry the information you removed from the page.
- Check scans and images by eye (or with OCR). On a scanned page the "text" is a picture, so a text search finds nothing even when a name is plainly visible. Look at every page at full zoom, or run it through OCR and search the result. If the PDF has layers, turn every layer on before you look.
If every check comes back clean, the redaction held. If any check finds a redacted term, don't release the file: redact it again with a tool that removes content (not one that draws shapes), then repeat the checks.
Why "it looks redacted" isn't proof
A PDF can have several layers of information on one page: the visible drawing, a separate text layer, annotations on top, and metadata that is never displayed. A black rectangle added in an ordinary editor is just one more drawing on top. The text layer underneath is untouched, which is why check #1 works on so many "redacted" files.
Real redaction deletes the characters under the box and then rewrites the file so that no earlier version of the page is left inside it. Verification means proving that deletion happened, instead of trusting that it did.
How RedactWorks runs these checks for you
RedactWorks does the removal and then automates most of the checks above on every document, before you download it. Here is exactly what it does, including what it doesn't check:
- Removes content, not just covers it. Approved items are deleted from the page with true redaction. Form fields and annotations are flattened into the page first so their text is redacted too. Embedded files, bookmarks, links and remaining annotations are removed, metadata and XMP are cleared, and the file is rewritten clean.
- Re-checks the finished file. After redaction it re-extracts the text and searches for every item you approved. It also checks the metadata, XMP, annotations and embedded files. For scanned documents and images, it re-reads each page with OCR to catch values that are still visible in the picture.
- Stops you before a bad download. If anything you approved can still be read, the download is held behind a red warning. You can only download that file by explicitly acknowledging the warning, and that override is recorded.
- Gives you a record. You can download a decision log PDF of what was found and what was redacted.
What the automatic check does not cover: boxes you draw by hand and whole-page redactions aren't text-checked, images inside a regular (non-scanned) PDF aren't read with OCR, and very short items (two characters or fewer) are skipped. For those, run checks #1 and #6 yourself. We'd rather tell you that here than have you find out later.
Frequently asked questions
How can I tell if a PDF redaction is real or just a black box?
Try to select and copy the text under the box, and search the document for the redacted words. If you can copy it or find it, it is only covered. Confirm with a separate text extractor, because some viewers hide text that is still in the file.
Can redacted text be recovered from a PDF?
Only if it wasn't really removed. Text hidden under a drawn shape, in metadata, in comments or in form fields can be recovered. Text deleted by true redaction, with the file rewritten afterward, can't be recovered from that file.
Does Adobe Acrobat verify redactions?
Acrobat's Redact tool removes content, and its Sanitize option removes hidden information such as metadata and comments. It doesn't search the finished file for the words you redacted, so run the checks above before you release anything.
Does RedactWorks guarantee a redaction is complete?
No tool can guarantee that, and we don't claim to. RedactWorks re-checks every document for the items you approved and holds the download behind a warning if any are still readable. Hand-drawn boxes, whole-page redactions and images inside regular PDFs still need a visual check.
Is there a free way to verify a redaction?
Yes. The six checks on this page use your PDF viewer plus free tools like pdftotext and any OCR app. RedactWorks also includes 3 free redactions, verification included, with no card required.