How to redact audio and video: spoken details, faces, and on-screen text
Updated
Short answer
Redacting audio or video means removing sensitive information from every channel it appears in: muting or bleeping spoken details in the soundtrack, covering faces and on-screen text in the frames, and removing caption tracks and container metadata such as location. Editing only one channel leaves the others exposed.
Where personal information hides in a recording
- Speech: names, phone and account numbers, addresses, and dates of birth said out loud — in a support call, an interview, a meeting, or a voicemail.
- The picture: faces, name badges, screens, documents on a desk, whiteboards, license plates, house numbers, and reflections.
- Caption and subtitle tracks, which repeat the spoken words as text.
- Container metadata: the recording date and time, the device model, and on many phones the location where the video was recorded.
- Separate transcripts or meeting notes generated from the recording.
Audio: muting versus bleeping
Muting sets the audio to silence for the time range; bleeping replaces the range with a tone. Both remove the speech when done properly. Bleeping makes the edit obvious to listeners, while muting is less distracting in long recordings. What matters is that the original sound is replaced. Mixing a tone over speech that is still present underneath, or just lowering the volume, can leave words recoverable with filtering or amplification.
To find what to remove, you can listen through and note time codes, or transcribe the recording with speech recognition and search the transcript. Transcription is faster for long recordings, but word timings are approximate, so pad each range slightly and listen to the result.
Video: covering faces and on-screen text
Visual redaction means drawing an opaque box, pixelating, or blurring an area of the frame for as long as the sensitive content is visible. Solid boxes are the safest choice for text: light blur or coarse pixelation of short strings, such as numbers, can sometimes be reversed or guessed by comparing candidates. Faces and objects move, so the covered area must follow them frame by frame. Visual edits also require re-encoding the video; the picture can’t be changed without decoding and re-compressing it.
Watch the whole result at normal speed, then scrub through slowly. A face that appears for a few frames, a document caught in a reflection, or a box that drifts off its target is easy to miss.
Metadata and extra tracks
Recordings from phones and cameras often include the capture time, the device, and a location. Video files can also carry subtitle tracks, chapter names, and data streams. Export the final file with metadata removed, and check it with a tool such as ffprobe or ExifTool, which list container tags and streams.
How GhostX redacts speech
GhostX’s audio and video redaction work on the soundtrack, in your browser. The recording is transcribed on your device with a Whisper speech-recognition model, and the transcript is scanned with the same detectors GhostX uses for documents: email addresses, phone numbers, Social Security number patterns, card numbers, IBANs, IP addresses, dates, your custom rules, and — when the on-device name model is loaded — names. You review the findings, then apply.
Each finding mutes the whole transcript segment it occurs in, not just the individual word. That over-redacts by design, because word-level timings from speech recognition are not precise enough to rely on. The tool shows the total time muted so you can judge the result, and you should listen through once before sharing.
Audio is saved as MP3; video is saved as MP4 with the original picture and the muted soundtrack. For MP4, M4V, and MOV sources the video stream is copied without re-encoding when possible; other formats are re-encoded. Container metadata, chapters, subtitle tracks, and data streams are removed from the output.
What GhostX doesn’t do
GhostX mutes speech; it does not bleep, and it does not blur or cover faces or on-screen text in video. For visual redaction, use a video editor with tracking masks and check every frame. Speech recognition can also miss or mishear words — accents, crosstalk, background noise, and spelled-out numbers are common trouble spots — so detection is a first pass, not a guarantee.
Frequently asked questions
-
Is bleeping better than muting?
Both remove the speech if the original audio is replaced rather than covered. Bleeping signals an edit to the listener; muting is less intrusive. GhostX mutes.
-
Can speech recognition find every piece of personal information?
No. It can mishear words, and some details — a name spelled out letter by letter, a number said in pieces — are hard to detect. Use it to find candidates quickly, then listen to the result.
-
Does muting change the length of the recording?
No. The muted range becomes silence, so the timing of everything else stays the same and the video stays in sync.
-
Does GhostX blur faces in video?
No. GhostX’s video redaction mutes spoken details and removes metadata. Faces and on-screen text need a video editor that supports masks or tracking.