GhostXStudio tools

Is covering text with a black box real redaction?

Updated

Short answer

No. A black box or highlight drawn over text in a PDF only adds a shape on top; the text underneath stays in the file and can be copied, searched, or extracted. Real redaction deletes the underlying content itself, so there is nothing left under the box to recover.

Why the text survives under the box

A PDF page is a list of drawing instructions: place these characters in this font at these coordinates, draw this image here, fill this rectangle there. The characters are stored as text, not as pixels, which is why you can select, copy, and search them. When you draw a black rectangle on top, the viewer paints the rectangle after the text, so your screen shows black. The text instructions are still in the file, unchanged.

The same is true for most of the ways people try to hide text in a PDF:

  • A rectangle drawn with a shape or comment tool is an annotation — a separate object layered over the page. It can often be selected, moved, or deleted, and the text underneath is untouched.
  • A highlighter set to black is also an annotation. Copy-paste and search go straight through it.
  • Changing the font colour to black-on-black or white-on-white hides the words visually but keeps them as ordinary text.
  • Drawing a black box in Word or Google Docs and then exporting to PDF usually produces the same result: text in the page, shape on top.
  • On a scanned page, a box over the image may cover the pixels, but many scanned PDFs also carry an invisible OCR text layer that still contains the words.

None of these change what a machine can read. Anyone who copies the page into a text editor, searches for a word, or runs a text-extraction tool sees the ‘redacted’ content.

Check a file in about a minute

You don’t need special software to spot a failed redaction. Work on a copy of the file and try to get the hidden words back:

  1. Select all the text on the page (Ctrl+A or Cmd+A), copy it, and paste it into a plain-text editor such as Notepad or TextEdit. If the redacted words appear, the redaction failed.
  2. Use the viewer’s search (Ctrl+F or Cmd+F) to look for a word you know was under a box. A match that highlights a black area means the text is still there.
  3. Click on the black box. If it can be selected, moved, or deleted, it is an annotation sitting on top of the page.
  4. Open the file in a second program — a different browser, or a command-line tool such as pdftotext — and extract all the text. Different readers expose different layers.

Passing these checks is necessary but not sufficient: metadata, comments, attachments, and earlier saved versions can hold the same information. Our guide to checking a redacted PDF walks through those as well.

What real redaction does

Real redaction removes information from the file, rather than hiding it on screen. There are two common ways to do it.

Content removal edits the page itself: the text operators, image pixels, and vector shapes inside the marked area are deleted, and a black box is drawn in their place. Professional editors such as Adobe Acrobat Pro work this way when you mark areas with their Redact tool and then apply the redactions. The rest of the page keeps its selectable text. The risk is in the details — content that overlaps the edge of the box, text in form fields or comments, or an image that extends under the mark all have to be handled correctly.

Rasterization turns each page into a picture, paints the black boxes into the pixels, and builds a new PDF from those pictures. Because the output contains images only, there is no text layer left anywhere for copy-paste or extraction to find, and the pixels under each box are replaced rather than covered. The trade-off is that text is no longer selectable or searchable, and files are usually larger.

Either way, a thorough redaction also deals with the parts of a PDF you can’t see on the page: document properties such as author and title, the XMP metadata packet, comments, form field values, embedded attachments, bookmarks, and earlier revisions saved into the same file.

How GhostX redacts a PDF

GhostX’s PDF redaction uses the rasterization approach, and it runs in your browser — the PDF is not uploaded. You draw boxes over the areas to remove, or use auto-detect to suggest them: it finds email addresses, phone numbers, US Social Security number patterns, card numbers (checked with the Luhn algorithm), IBANs (checked with the mod-97 rule), IP addresses, and dates, plus names and organizations when the on-device name model is loaded. Suggestions are never applied until you accept them.

When you apply, every page of the document is rendered to an image at roughly 144 DPI, the boxes are burned into the pixels as solid black, and a new image-only PDF is assembled. Because the file is rebuilt from page images, the original document properties, XMP metadata, annotations, form fields, attachments, and scripts are not carried over.

After redacting, GhostX re-extracts text from the output and re-runs its detectors. When nothing comes back, the result is labelled ‘Redacted · no text layer left’. That check reads text, not pixels: it cannot tell whether you missed an area or placed a box slightly off target, so look over every page of the result before you share it.

Mistakes that undo a good redaction

  • Sending the original file instead of the redacted copy — give the redacted file an unmistakable name.
  • Leaving identifying details in the file name, the email subject line, or the page headers and footers.
  • Redacting one instance of a name and missing others on later pages; search the original for every term first and make a list.
  • Redacting a Word document, then exporting to PDF with comments or tracked changes still in it.
  • Forgetting that the size and position of a box can hint at what was removed — a short box after ‘Patient:’ reveals a short name.

Frequently asked questions

  • Does a black highlight redact text in a PDF?

    No. A highlight is an annotation layered on top of the page. The text underneath stays in the file and can be copied, searched, and extracted. Use a redaction tool that removes the content, then verify the result.

  • Can a properly redacted PDF be un-redacted?

    Not from the file itself: when the content has been removed or the page has been rasterized with the box burned into the pixels, the original text is no longer present. Information can still leak through context — the length of a box, a name repeated elsewhere, or a copy of the original sent by mistake.

  • Is printing and scanning a safe way to redact?

    Printing a marked-up page and scanning it does produce an image with no original text layer, which is the same idea as rasterization. It is slow and error-prone, and some scanners add an OCR text layer that re-reads anything left visible, so check the scan the same way you would any redacted PDF.

  • Will my redacted PDF still be searchable?

    With a rasterizing redaction tool such as GhostX, no — the output is image-only, which is what guarantees no hidden text remains. Tools that remove content in place can keep the rest of the page searchable, but need more careful checking.