New in Learn: a reproducible test of six free PDF redaction tools (GhostX, SafeRedact, PDF24, Smallpdf, Xodo, iLovePDF) run on one synthetic patient record. Every tool removed the text we marked, but four of them left the SSN from a form field in the file and kept the patient's name in the document metadata. The test found a bug in our own name detection too, which is fixed and disclosed in the write-up. The test document and the leak scanner are published so anyone can re-run it.
Auto-detect's on-device name detector was loading and running, but every name it found was thrown away before it reached the page, so redaction, the X-ray scan and speech muting only ever suggested structured data (emails, SSNs, cards, phone numbers). Names, organizations and places are now suggested again, a detected first name also covers the surname that follows it ("Jane Q. Testperson", not just "Jane Q"), and if name detection ever fails to run, the tool now says so instead of quietly reporting fewer results. We found this while testing redaction tools side by side.
GhostX is about keeping your files private, so the developer utilities (JSON formatter, base64, hashing, UUIDs, passwords, word and case tools) have been retired — their old links now lead home. In their place: dedicated pages for tools that were already built in — reorder PDF pages, PDF to PNG and to text, HEIC and WebP to PDF, Word to Markdown and HTML, watermark a photo, remove a photo's location, FLAC/OGG/AAC/Opus to MP3, AVI and WMV to MP4, extract audio from any video, auto-generate subtitles, lock any file with a passphrase, and check a file's SHA-256 fingerprint. New guided pages for redacting medical records and bank statements, comparisons with Adobe Acrobat's redaction, Otter.ai and remove.bg, and a new Learn section with ten plain-English explainers — whether a black box is real redaction, what hidden data a PDF carries, the 18 HIPAA identifiers, what a signature seal proves, and how to check any tool for uploads. The homepage gains a Redact & protect row. We also corrected a few pages that described things inaccurately: redaction flattens every page (not only the ones you mark), auto-detect reads the text layer and only OCRs pages without one, the site does use a service worker for offline caching, and the Security page now lists the signing seal among the features that contact our servers (it sends only a fingerprint).
The privacy tools now say what they do. 'Compliance X-ray' is 'Scan for private info', 'Redact' is 'Black out (redact)', 'Inspect & scrub' is 'Remove hidden data', 'Redact speech' is 'Mute private speech', 'Encrypt' is 'Lock with a passphrase' (a GhostX-only .gxenc — distinct from 'Password-protect PDF', which opens in any PDF reader), 'Timestamp' is 'Prove when it existed', and 'Checksum' is 'Check it’s unchanged' — each with a one-line description of what happens to your file. Page titles and URLs keep the familiar terms. Also: in Safari private windows, a file dropped on the homepage or a tool page now opens straight in the right tool — private browsing refuses to store files between pages, so the handoff now falls back to storing the raw bytes. The whole site was checked page by page in Chrome and Safari's engine (WebKit) on desktop, iPhone, iPad, and Android sizes: no layout overflow, no script or security-policy errors, and every privacy flow — scan, black out, verify, encrypt/decrypt, custom rules — passing end to end.
The privacy tools now read as one toolkit. After you drop a file, the studio's operations are split into 'Edit this file' and a new 'Privacy & proof' section, ordered the way a compliance check runs: Compliance X-ray, Redact, Inspect & scrub, Password-protect / Encrypt, Timestamp, Checksum — instead of being scattered through the general edit tools. And custom detection rules have a page of their own at /detection-rules, so you (or a team admin importing a shared rule file) can set rules up before dropping anything; a rule saved there is live in PDF Redact, image Redact, X-ray, and Redact speech on the next scan. Rules still live in your browser only. /security now documents that, and links to the page.
Two additions to the redaction toolkit. (1) Redact images (/image/redact, and a Redact op in the image studio): the PDF Redact experience pointed at a screenshot or photo — auto-detect reads the image's text on your device and suggests a box over each email, phone number, SSN, card or IBAN number, name, and custom-rule match; accept what you want, drag boxes over anything else (faces, signatures, plates), and apply. The boxes are burned into the pixels and the image is re-encoded, so what's underneath is destroyed and the hidden EXIF/GPS metadata is dropped too. Compliance X-ray's image fix now does this directly — 'Redact all N findings' instead of a metadata scrub that couldn't touch visible text. (2) Custom detection rules: add your own word lists (client names, project codenames) or patterns (employee IDs like EMP-12345, case numbers) and they apply everywhere at once — PDF Redact, image Redact, X-ray, and Redact speech. Rules are saved in your browser only (they can contain sensitive terms), with a live 'try it' box, validation that rejects patterns that could freeze the page, and JSON import/export so a team can share one rule set. Findings from your rules are counted as 'custom' in certificates and analytics — never by rule name. Also: emails that OCR splits at the @ ('jane.doe @example.com') are now detected, found while testing this in a real browser.
Every compliance flow was driven end to end in real Chrome against the production build and CSP: X-ray on PDF, image, and audio; /verify on redacted files and certificates; Redact speech on audio and a real VP8 WebM; encrypt → decrypt (including a renamed container); Markdown → PDF; zip and spreadsheet encryption; /security. That pass found and fixed three real problems. (1) Redact's Auto-detect OCR'd small page previews and reported "No sensitive text found" on a PDF containing an email and an SSN — text-layer pages are now detected and boxed from their own glyph geometry at any font size, with OCR only for scanned pages. (2) Redaction boxes placed word positions by character count, so a run of wide letters could leave the last glyph of an email visible — word extents now use real Helvetica proportions plus a safety margin (over-cover, never clip). (3) X-ray showed a green all-clear over a photo carrying GPS coordinates — the verdict now weighs hidden metadata, not just text. Also: redaction certificates withhold the source filename unless you opt in (filenames often carry the very PII redacted); spreadsheets can be encrypted straight from GhostSheet; muting a .mov that holds ProRes falls back to re-encoding instead of failing; watermarking a GIF says up front that animation isn't kept; and the wrong-passphrase message drops the jargon.
A five-reviewer pass over the compliance wave (correctness, security/privacy, testing, UX, adversarial) and every fix it found. X-ray on a text PDF now finds and boxes each finding in ONE pass from the text layer's own glyph geometry, on every page — previously it re-located findings by OCR-ing low-res thumbnails of the first 12 pages, so a finding could be reported but never covered; scan-only pages still go through OCR, and anything that can't be located makes the report partial. An encrypted PDF, an unreadable metadata tier, or a name model that silently failed now reads as partial/not-checked instead of clean, and images only offer the metadata scrub for metadata (it never touched visible text). Speech redaction: a finding in the last Whisper segment (which can report no end time) was silently left audible — it now mutes to the end of the media; the redacted transcript masked only the last of several findings in one segment — all are masked; mutes drop subtitle/data tracks, container metadata, and chapters; WebM/MKV/AVI sources are re-encoded instead of failing after transcription; names past ~64k characters of transcript are now scanned; cancelling stops the speech model; both models are freed before ffmpeg loads. Encryption: the .gxenc header is now authenticated (tampering fails), the original filename moved inside the ciphertext, the stored work factor is honored and bounded, passphrases need confirming, renamed containers still decrypt, and files with no studio (zip, txt, …) can be encrypted too. Markdown→PDF: multi-line code blocks and tabs no longer fail, unencodable characters are substituted instead of failing the document, underscores in URLs/identifiers survive, ~~~ fences close, and numbered lists keep their numbers. Copy: "verified clean" became "no text layer left" (the re-scan can't see pixels), the redaction certificate says plainly it's self-reported and unsigned, /verify gained a "couldn't re-check" state, and /security now describes the CSP, third-party scripts, telemetry, and offline behavior accurately. Also: /pdf/timestamp routes non-PDF drops to the right studio, watermarks keep transparency and shrink long text to fit, and checksum comparison can no longer hang or race.
A wave of privacy infrastructure built from engines GhostX already shipped, aimed at the compliance/redaction wedge. (1) Compliance X-ray (GhostXray, /x-ray): drop any file — PDF, image, audio, video, document — and get one severity-graded report of PII and hidden data: the GhostTrace metadata tier plus a content tier running the GhostRedact detectors over text layers, OCR (images and scanned PDFs), on-device Whisper transcripts, and document text. Previews are masked by default (screenshot-safe); an unreadable tier reports itself as partial, never clean. One-click remediation runs in-surface: PDFs get the destructive redact pipeline (boxes over detected findings, pages rasterized so no text layer remains — visually confirm the boxes), images get the metadata scrub (visible text in the picture is not removed), audio/video get the speech segments containing findings muted. (2) Redact speech (/audio/redact, /video/redact): Whisper transcribes on your device, the detectors find spoken PII (numbers, emails, cards, IBANs — names via the on-device model, auto-loaded on capable devices), you review each finding with its timestamp, and the mute pass silences the transcript segment containing each finding — segment-level, so the surrounding phrase goes too; the review shows each span and the total (MP4/MOV video is stream-copied untouched; other containers are re-encoded to H.264). A masked transcript ships alongside. The whisper worker is disposed before ffmpeg loads so the ~140 MB model and the ~32 MB engine never coexist. (3) Redaction certificates: the optional knob delivers a standalone PDF recording the SHA-256 of source and output, counts by kind (never the redacted text), and the result of a post-redaction re-scan that re-extracts the output's text layer and re-runs the structured-PII detectors. That re-scan shows no recoverable text layer remains — it can't see pixels, so it doesn't prove every sensitive area was boxed. The certificate is self-reported by the tool and not digitally signed; its hashes identify which files it describes. /verify learns both flavors — the certificate card, and a live re-check on any redacted file (a live scan beats any embedded claim, which is why the markers deliberately don't embed one). (4) GhostGuard: password-encrypt any file into the GXENC container (PBKDF2-SHA256 600k + AES-GCM 256, WebCrypto) — re-dropping a .gxenc anywhere opens a DecryptCard that restores the original through the normal studio load. Plus a SHA-256 checksum surface with byte-compare. (5) Image watermark (burned-in, 9-grid placement) and Markdown → PDF (.md now opens in the doc studio with a rendered preview and converts through a new parser; the resilience corpus caught a real dead-end before ship — un-encodable characters now map to specific copy instead of the generic fallback). (6) /security: the trust surface — architecture, the three stated exceptions (Send/Beam, multi-signer, GhostProof), how to verify every claim with the Network tab, and the honest limits. Also new: GhostRedact markers, verification, and X-ray in every hub's op rail; homepage verbs and footer links; detection-kind tagging from detector to certificate so counts are honest.
Eight CE review agents audited v3.0 in parallel. Found and fixed: (1) P0 self-await deadlock in finalize() — `writePromise = writePromise.then(() => finalize())` while finalize internally awaited writePromise → recipient hung at "Finalizing…" forever, never closing the sink. Dropped the redundant await; the chain itself serializes ordering. This was a complete success-path break — every transfer would have hung. (2) P1 security: recipient now enforces its own deriveCaps().recipientMax against sender-declared metadata.size, rejecting upfront with a clear browser-upgrade hint instead of OOM-crashing mid-transfer on the memory path. (3) High: handleCancel during cold-start now sets startupAbortedRef so a Cancel mid-await actually revokes the beam doc instead of leaking it. (4) High: public accept() now has an idempotency latch so React Strict Mode double-invoke / programmatic re-entry can't open two save pickers or leak the first FSA writable. The latch resets on save-picker cancel so the user can retry. (5) P1 reliability + security: new sweepOrphanedOpfsBeams() runs on every recipient page load — evicts beam-*.tmp files older than 30 minutes from OPFS so partial plaintext from a prior tab crash doesn't persist on shared devices. (6) parseMetadata now validates baseIV decodes to exactly 8 bytes (11 b64u chars) — catches malformed envelopes at the trust boundary. (7) BEAM_WIRE_FORMAT constant extracted from four duplicated string literals. (8) Sender + recipient chat dispatch converted from if-chains to exhaustive switch with never default — adding a new BeamChatFrame variant becomes a compile error instead of a silent drop. (9) chatChannel.onclose handlers now gated by terminated/aborted so successful transfers don't flip the chat pill to red "Disconnected" on the post-success teardown. All 258 tests still passing (209 frontend + 49 functions). Deferred to v3.1 with empirical groundwork: signal-based backpressure across the chat channel for slow-disk recipients, Wake Lock + visibilitychange for multi-GB mobile transfers, navigator.storage.estimate() quota probe at sink open, ciphertextSize ceiling unification across server + tryFetch + connectAndNegotiate, and the QR code on /beam (flagged as MVP gap for mobile-to-mobile pasting).
File caps tier up to 10 GB by detected browser capability — and the server bandwidth cost stays exactly zero because the bytes still flow direct browser-to-browser over WebRTC. Three recipient code paths: (1) Chrome/Edge/Opera get the File System Access path — the recipient picks the save location via showSaveFilePicker and we stream decrypted chunks straight to disk via a WritableStream, ~10 KB peak memory regardless of file size; (2) Firefox/Safari modern get the OPFS path — chunks land in an origin-private file then trigger a browser download, ~2× disk during transit but bounded RAM; (3) older browsers fall back to the current in-memory path with a capability-tightened cap (500 MB desktop, 250 MB mobile). The sender now reads via file.stream() instead of file.arrayBuffer(), so its memory peak is also a single chunk regardless of file size. Wire format upgrades to chunked-v1: each 64 KB plaintext chunk gets its own AES-GCM box with a deterministic IV (8-byte session-random base concatenated with a 32-bit big-endian chunk counter — 2^96 distinct IVs per key, well beyond any practical collision concern). New @ghostx/crypto primitives: generateChunkBaseIV / deriveChunkIV / encryptChunk / decryptChunk. Sender and recipient both use the WebRTC DataChannel's bufferedAmountLow event for backpressure-aware streaming. Capability detection module (./capabilities) returns the receivePath / cap / human-readable label plus an upgradeHint() — the sender page now displays "Up to {X}" derived from caps + a "Want larger files?" line that points users on suboptimal paths at Chrome/Edge/Opera on desktop. Server's MAX_FILE_PLAINTEXT_BYTES rises to 10 GB (was 500 MB) with a chunked-tag overhead buffer. Tests: 36 crypto tests now cover deriveChunkIV, encryptChunk/decryptChunk round-trips, IV uniqueness, wrong-index/wrong-base rejection, tampered-ciphertext rejection. parseMetadata tests rewritten against the chunked-v1 envelope shape with fmt+baseIV+chunkSize boundary cases. 209 frontend + 49 functions = 258 tests passing. Breaking wire format vs v2.27 — beams created on v3.0 use chunked-v1 envelopes and won't be readable by v2.x clients, but since beams expire in 15 min and require both peers online simultaneously, there's no real cross-version compatibility window in practice.
Previously chat only worked while the file was streaming. Now WebRTC auto-connects the moment the recipient clicks through the 'Continue to open the beam' gate, so the chat channel is alive before they see the file. Both sides can message each other while the recipient reviews the metadata + decides whether to accept. Hitting Accept now fires a 'go' control frame over the chat channel (instead of triggering the SDP handshake — that already happened in the background); the sender starts streaming on receipt. Chat keeps working through streaming, after the file lands, and until either tab closes or the 15-minute beam expires. Wire envelope is a tagged union {kind: 'msg'|'go'|'typing'}: text messages, the start-the-transfer control signal, and typing indicators all multiplex onto the single 'beam-chat' DataChannel, AES-GCM-256 encrypted with the same URL-fragment key. New presence: 'is typing…' indicator on both sides (debounced 1.5s idle), persistent connection pill that flips Connected → Disconnected on channel close so you know when the other side actually leaves vs is just quiet. Capability detection + chunked AEAD for 10GB+ files is queued for v3.0 — those two ship together so detection routes to a real code path instead of being informational UI.
Both sides of a beam now know when the other is alive. Sender sees 'Recipient online' the instant the recipient opens the link (via a new markBeamRecipientOpened Cloud Function that stamps a first-write-wins timestamp on the beam doc), and flips to 'Recipient connected' once the WebRTC DataChannel opens. Recipient sees the same connection status mirrored on their side. Alongside the file channel, GhostBeam now opens a second WebRTC DataChannel labelled 'beam-chat' — both sides can send short text messages back and forth during and after the transfer. Messages are AES-GCM-256 encrypted with the same URL-fragment key as the file payload and stream direct browser-to-browser; the chat content never touches our servers. Same hardening pass as the file channel: terminated/aborted latches, cleanup detaches handlers, length-capped at 2000 chars. Recipient surfaces (/send/:id, /beam/:id) also gained a discreet 'Powered by GhostX' footer linking to sibling tools — the recipient just had a great free experience, no signup required, and that's the warmest possible cross-promotion moment. Plus second-pass review fixes: timingSafeEqual on ownerToken compare with hex-shape precheck, strict typeof-string already-answered guard, submitBeamAnswer rate-limit consistent with createBeam, reapExpiredBeams retryCount=1, recipient ciphertextSize bounds at fetch, sender teardownTimer 5s→15s for low-end decrypt, getDoc 750ms retry on stale not-found, exhaustive mapWirePhase switches, mobile-OOM warning in AcceptCard, full unit coverage on parseMetadata + parseSessionDescription + CF helpers (BEAM_ID_RE tightened to match the alphabet exactly).
New product at /beam. Drop a file, share a one-time link, the recipient opens it, and the file streams direct browser-to-browser over WebRTC. The bytes never touch our servers — we mediate only an AES-GCM-encrypted SDP handshake. AES-GCM-256 key lives only in the URL fragment; both the file metadata and the SDP envelope are encrypted before they reach Firestore. Up to 500 MB per beam, 15-minute beam lifetime, both sender and recipient must be online simultaneously. NAT-traversal via Google STUN; clients behind symmetric NATs get a clear fallback message pointing at GhostSend. No TURN server (intentional — keeps the product free without ongoing bandwidth cost).
Removed the legacy ghostsign.app entries from the callable CORS allowlist, the Cloud Storage bucket CORS, and the Firebase Hosting `ghostsign-redirect` target. ghostx.tools is the only domain we serve from now on. Test fixtures and press-page copy updated to match. The legacy `ghostsign-redirect.web.app` site is still live in Firebase but no longer wired into this repo's deploys.
GhostSend now supports an optional password on every send. The link is one factor, the password (shared on a different channel) is the second — a leaked link alone won't open the payload. AES key is derived from the password via PBKDF2-SHA256 with 600,000 iterations (~250 ms on a laptop, ~1-3 s on mobile) so brute-forcing a leaked ciphertext is impractical. The salt rides in the URL fragment with a `p1:` version prefix; the password itself never crosses the wire. Wrong-password attempts get a fresh prompt without re-burning the share. Reaper schedule cut from hourly to every 5 minutes — unconsumed shares get deleted from our servers within ~5 min of TTL elapsing instead of up to 2 hours. GhostBeam queued as the next product: WebRTC peer-to-peer file transfer where the bytes stream directly browser-to-browser and never hit our servers.
Merged GhostNote and GhostShare into a single product at /send. One composer, two modes: paste a secret (text) or drop a file up to 100 MB. Same trust model — AES-GCM-256 in your browser, the decryption key never reaches our servers (it lives only in the URL fragment), atomic burn-after-read on first open, 1h / 24h / 7d expiry. Two SEO landing pages — /privnote-alternative and /wetransfer-alternative — render the same composer with competitor-targeted comparison tables and FAQs. Legacy /note and /share routes 301 to /send so existing links still resolve. Trust copy now explicitly contrasts GhostSend vs WhatsApp / Snapchat / iMessage / iCloud where the operator sees or retains the bytes.
Drop a file up to 100 MB. Get a self-destructing link. The recipient downloads it once, then it's gone forever. AES-GCM-256 encrypted in your browser; both the file bytes AND the filename are encrypted before they leave your machine. The decryption key lives only in the URL fragment so our servers never see it. Pick a 1h / 24h / 7d expiry — whichever comes first, the row + the ciphertext blob are deleted. No signup, no email, no logs.
Opening the Products menu now starts prefetching the JS chunks for sibling products in the background, so clicking through to GhostPDF / GhostQR / GhostNote / GhostSheet feels instant instead of waiting on a chunk fetch. Uses `requestIdleCallback` so it never competes with the page you're actually on, and only fires once per session per product.
Replaced the npm xlsx@0.18.5 dependency with the SheetJS CDN tarball v0.20.3, which closes CVE-2023-30533 (prototype pollution) and the known malformed-xlsx denial-of-service path. Bounded sheet parsing at 5,000,000 cells so a malicious workbook can't OOM the tab via a billion-row `!ref` header. Pasted HTML is now rejected outright (the paste box is for spreadsheet text, not markup). Formula cells are stripped before export so a re-shared file can't carry a smuggled =cmd|… formula. Cell edits now namespace by sheet (no more cross-sheet edit highlights), preserve leading-zero strings (ZIP codes / SSNs stay strings), guard against IME composition (Japanese / Chinese / Korean input no longer drops half-typed text), and force plain-text paste into cells. Extracted a shared `useDropdown` primitive into @ghostx/ui.
Drop an .xlsx, .xls, .csv, .ods, or .tsv (or paste tab-separated data) and see it instantly. Multi-sheet tabs, frozen header, full-text search across cells, click any cell to edit, then download as .xlsx / .csv / .tsv / .json. SheetJS does the parsing entirely in your browser — your file never leaves your machine. No Office install, no Google account, no upload.
Paste a password, an API key, or any secret. Get a self-destructing link. The recipient opens it once, then it's gone forever. AES-GCM encrypted in your browser; the key lives only in the URL fragment so our servers never see the plaintext. Pick a 1h / 24h / 7d expiry — whichever comes first, the row is deleted.
Generate styled QR codes for URLs, WiFi networks, vCards, email, SMS, plain text, or calendar events. Logo embed, color customization, PNG / SVG / PDF downloads — all in your browser. No watermark, no tracking pixel, no expiration.
Trimmed the GhostSign homepage by ~35% so the dropzone shows above the fold on more devices. The 'Why is GhostSign free?' section now matches the GhostPDF hub layout — same three cards, same vision. Verify link slid one spot left in the header so the most-used signing actions sit closer to the logo.
Refactored every PDF tool page to use a new `useObjectUrl` / `useObjectUrls` hook pair. Removes ~70 lines of repetitive boilerplate per page, eliminates a class of subtle memory-leak bugs around `URL.revokeObjectURL`, and clears the entire batch of React Compiler lint warnings.
GhostSign and GhostPDF now own their own chrome. The top-left logo, the cross-product link, and the donate popup all adapt to whichever surface you're on. Foundations for adding GhostNote / GhostShare etc. without touching layout code.
Public press kit at /press with brand assets, key facts, and pull-quotes. Added humans.txt, .well-known/security.txt, and the IndexNow API key so Bing, Yandex, and DuckDuckGo can be pinged on every deploy instead of waiting for their own crawl schedule.
Every /pdf/* tool now ships HowTo + FAQPage + BreadcrumbList JSON-LD into the initial HTML, plus a body section with 'How it works', 6 FAQs, and a related-tools rail. Tightened 16 oversized titles under Google's truncation budget.
Compress, merge, split, rotate, organize pages, add page numbers, watermark, crop, protect with a password, unlock, strip metadata, convert PDF↔Word, PDF↔Image, and extract text — every tool runs entirely in your browser. No upload endpoint. Free forever.
GhostSign code reorganized into a pnpm workspace with shared `@ghostx/*` packages (theme, brand, UI primitives, PDF core, signature, crypto, certificate). Sets up the family so each new product (GhostPDF, GhostNote, …) can share the same infrastructure without copy-paste.
GhostSign got the ability to send documents to other people for signature. The whole flow is AES-GCM-256 encrypted in your browser; the decryption key lives only in the magic-link URL fragment, which by web spec never reaches a server. No accounts for senders or recipients. Auto-deletes in 7 days.
Free in-browser PDF signing — no account, no upload, no watermark. Solo signing is 100% client-side: drop the PDF, place your signature, download the signed file with an audit-trail certificate. Built because every other free e-sign tool wanted our credit card.