GhostAudioStudio

Transcribe audio — on your device.

Speech to text, SRT, or VTT with OpenAI Whisper running in your browser. Dozens of languages, timestamps, no upload — the audio never leaves your device.

Turn any recording into text without uploading a single second of audio. GhostAudio runs OpenAI's Whisper speech-recognition model entirely in your browser — the audio is decoded locally, transcribed on your device, and never sent to a server. Export the result as plain text, SubRip subtitles (.srt), or WebVTT (.vtt) with timestamps. The multilingual model transcribes dozens of languages (and can auto-detect the spoken one); a faster English-only model is available when you want speed. Ideal for interviews, meetings, voice memos, lectures, and podcasts — including the sensitive recordings you can't hand to a cloud transcription service. Free, no signup, no minutes cap.

How it works

  1. 1

    Drop your audio

    Drag an MP3, WAV, M4A, or other recording onto the page or click to browse. It opens in the GhostAudio studio with Transcribe pre-selected.

  2. 2

    Pick format and language

    Choose plain text, SRT, or VTT, and either auto-detect the language or set it. The first run downloads the speech model once; after that it's cached.

  3. 3

    Download the transcript

    GhostAudio transcribes on your device and the transcript downloads straight from your browser.

Frequently asked questions

  • Is my audio uploaded to transcribe it?

    No. The Whisper model runs entirely in your browser via WebAssembly/WebGPU — the audio is decoded and transcribed on your device and never sent anywhere. You can confirm it in your browser's DevTools Network tab (only the model download appears, never your file).

  • What languages can it transcribe?

    The default multilingual model handles dozens of languages and can auto-detect the spoken one. For English audio you can switch to a smaller, faster English-only model.

  • Can I get subtitles with timestamps?

    Yes. Choose the SRT or VTT format and the transcript includes per-segment start/end times, ready to drop into a video editor or player.

  • Why does the first transcription take a while?

    The first run downloads the speech model (a one-time download that's then cached), and transcription runs on your own device, so speed depends on your hardware. Later runs skip the download and start immediately.

Every GhostAudio tool runs entirely in your browser. Your file is never uploaded — there's no upload endpoint to send it to.