GhostXStudio tools

How to check that a PDF is properly redacted

Updated

Short answer

To check a PDF is properly redacted, try to recover what was removed: paste all its text into a plain-text editor, search for a redacted word, extract the text with a second tool, and inspect metadata, comments, attachments, and embedded images. If a redacted word turns up anywhere, the redaction failed.

Before you start

Work on a copy of the redacted file, never the only version. Write down the specific words and numbers that were supposed to be removed — a name, an account number, an address — because each check below is a search for those terms. If you only look at the page and see black boxes, you are checking what a screen shows, not what the file contains.

Step 1: Copy and paste every page

Open the PDF, select all text (Ctrl+A on Windows, Cmd+A on a Mac), copy it, and paste it into a plain-text editor. Repeat for each page if your viewer selects only one page at a time. Any redacted term that appears in the pasted text is still in the file. This single check catches the most common failure: boxes or highlights drawn on top of text that was never removed.

Step 2: Search for the redacted terms

Use the viewer’s find function to search for each term on your list, including partial forms: a surname alone, the last four digits of a number, an email domain. A search hit that lands on a black area is a failed redaction. Also search for terms you didn’t redact but that would identify the same person or case, such as an employer or a street name.

Step 3: Extract the text with a different tool

Viewers differ in which layers they expose. Open the file in a second program and extract everything. On the command line, pdftotext from the Poppler utilities prints a PDF’s full text layer (pdftotext -layout redacted.pdf out.txt). A browser’s built-in PDF viewer, a second PDF editor, or a text-extraction library all work too. What you are looking for is any text at all in the areas that were redacted.

Step 4: Look at the boxes and the images

  • Click on each black box. If it can be selected, moved, or deleted, it is an annotation lying on the page, not a redaction.
  • Zoom to 400% and inspect the edges of each box for the tops or tails of letters that stick out.
  • Remember that a box drawn over an embedded image doesn’t alter the image. Image-extraction tools such as pdfimages (also from Poppler) pull out the full original picture, including the part under the box.
  • On scanned documents, check for an invisible OCR text layer: if you can select words on a scanned page, that layer exists and must be redacted too.

Step 5: Check metadata, comments, and attachments

Open the document properties (File → Properties in most viewers) and read the title, author, subject, and keywords — titles often repeat a name or case number. Then open the comments, attachments, and bookmarks panels. Review comments can quote redacted text, a bookmark can be titled with a person’s name, and an attached spreadsheet can contain the full data set. Form fields keep their values even when a box is drawn over them.

Step 6: Look for earlier saved versions

Many PDF editors save changes by appending to the end of the file rather than rewriting it, a mechanism called an incremental update. The earlier version of the page can remain inside the file, and some tools can roll it back. One rough signal: open the PDF in a text editor and count occurrences of %%EOF. More than one can indicate incremental updates, although files saved for ‘fast web view’ (linearized) normally contain two. The reliable fix is to redact with a tool that writes a completely new file, then run these checks on that file.

Step 7: Check the file name and the delivery

The file name, the email subject line, and the message body travel with the document. Make sure none of them repeat what you removed, and confirm you are attaching the redacted copy rather than the original.

Faster checks with GhostX

For a file GhostX redacted, dropping it on the /verify page re-extracts its text layer in your browser and reports either ‘No Selectable Text Remains’ or ‘Text Was Found’ (a file redacted by rasterization should come back with no selectable text); /verify does not run that check on PDFs redacted by other tools. For any PDF, the X-ray tool reports a PDF’s document properties, XMP metadata, annotations, embedded files, form data, and scripts, and scans whatever text the file has for emails, phone numbers, ID numbers, and similar details. Both run locally; the file is not uploaded.

Two limits to keep in mind. Neither check can see pixels, so a box placed in the wrong spot on a rasterized page will pass — you still need to look at every page. And the X-ray does not currently detect incremental-update history, so use the %%EOF check above or re-save through a tool that rewrites the file.

Frequently asked questions

  • If copy-paste returns nothing, is the PDF safe to share?

    Not necessarily. Copy-paste only tests the text layer. Check the embedded images, document properties, comments, attachments, form fields, and earlier saved versions, then look at every page to confirm each box covers what it should.

  • What does ‘no selectable text remains’ mean?

    It means a text-extraction pass found no text in the file, which is expected for a rasterized (image-only) redaction. It does not mean every sensitive area is covered — that can only be confirmed by looking at the pages.

  • How do I check a redacted PDF that must stay searchable?

    Use the same steps. Copy-paste and search are especially important because the page still has a text layer; every redacted term should be absent from it while the rest of the text remains.

  • Can I tell if a PDF has hidden earlier versions?

    Count the %%EOF markers in the raw file: more than one suggests incremental updates, though linearized files normally have two. Acrobat also shows a ‘view signed version’ option for signed files with later changes. When in doubt, rebuild the file with a tool that writes a fresh copy.