Free tool

Free audio to SRT converter: transcribe MP3, podcasts, and voice notes to subtitles

Drop in a podcast episode, an interview, or a voice memo and get subtitles with real timings. Speech-to-text runs on your own device, and the SRT, VTT, or text transcript downloads straight to it.

  • No upload, runs in your browser
  • Private by design
  • No watermark, no sign-up
  • No size or length limit

What you get

1
00:00:00,000 --> 00:00:02,840
Welcome back to the show. Today

2
00:00:02,840 --> 00:00:05,200
we're talking about podcasts.
  • SRT, WebVTT, and plain-text transcript, all with the same timings
  • Fix a name or a word inline before you download
  • Detects the spoken language across the languages Whisper supports
  • Runs on your device: the audio is never uploaded

No sign-up, no watermark, no length limit.

How to transcribe audio to SRT for free

To convert audio to SRT for free, open the captionrich Free Audio to SRT Converter and drop in an MP3, WAV, M4A, OGG, or FLAC file. An open-source Whisper speech model runs on your device, detects the language, and returns a word-timed transcript in roughly a quarter of the audio's length on a laptop. Choose short, standard, or long subtitle lines, click any line to fix a word, then download the SRT or WebVTT file for YouTube, Spotify, Podbean, Premiere Pro, or any player, or grab the plain-text transcript. Nothing is uploaded, there is no sign-up, and there is no length limit.

How it works

  1. 1

    Drop in your audio

    Click the upload area or drag an MP3, WAV, M4A, OGG, FLAC, or a video onto it. The file is decoded locally; nothing is sent to a server.

  2. 2

    Let it transcribe on your device

    On the first visit the browser downloads the Whisper speech model once and saves it. Transcription then runs locally with a progress bar, and the language is detected from the first 30 seconds.

  3. 3

    Check the lines

    Play the audio next to the transcript; the current line highlights as it plays. Pick short, standard, or long lines, and click any line to fix a word.

  4. 4

    Download SRT, VTT, or text

    Click SRT for a SubRip file, VTT for WebVTT, or TXT for a paragraph transcript. Copy text puts the whole transcript on your clipboard.

What you get with the free audio to srt converter

  • Made for audio: MP3, WAV, M4A, AAC, OGG, Opus, FLAC, and AIFF, including iPhone voice memos and WhatsApp voice notes. Videos work too.

  • Speech-to-text runs on your device with an open-source Whisper model, so the recording never leaves your browser.

  • Three outputs with identical timings: SRT (SubRip), WebVTT, and a plain-text transcript split into paragraphs at pauses.

  • Pick short, standard, or long subtitle lines to match social clips, YouTube, or a podcast show-notes transcript.

  • Click any line to correct a name or a word before you download; click a timestamp to hear that moment.

  • Detects the spoken language automatically across the languages Whisper supports.

  • One-time model download, cached by your browser, so every later file transcribes offline.

  • No sign-up, no watermark, no length limit, no upload queue.

FAQ

Questions about the free audio to srt converter

Yes. There is no account, no credit card, no trial, and no per-minute charge. It is free because the speech recognition runs in your browser on your own hardware instead of on a server, so a three-hour podcast costs us nothing to transcribe.

Anything your browser can decode: MP3, WAV, M4A and AAC (iPhone voice memos), OGG and Opus (WhatsApp and Telegram voice notes), FLAC, AIFF, and WebM. Video files such as MP4 and MOV work too, since only the audio track is used. AMR files from some Android recorders are not supported by browsers; convert those to MP3 first.

No. The file is decoded by your browser and transcribed with an open-source Whisper model that runs on your device through WebGPU or WebAssembly. Nothing is transmitted, stored, or logged. The only download is the speech model itself, once, which the browser keeps for next time.

The on-device model is OpenAI's Whisper base, which handles clear speech well but can miss names, jargon, heavy accents, and crosstalk. Click any line to fix a word before you download. The captionrich editor uses a larger cloud speech model with up to 99% accuracy.

Roughly a quarter of the audio length on a laptop with WebGPU, so a one-hour episode takes about 15 minutes, and longer on a phone or in a browser without hardware acceleration. The first visit adds a one-time model download of 80 to 200 MB.

SRT and WebVTT are subtitle files with a timestamp on every line; upload them to YouTube, Spotify for Podcasters, or Podbean, or import them into Premiere Pro, DaVinci Resolve, or CapCut. The TXT download is the same words without timestamps, split into paragraphs where the speaker pauses, ready for show notes, a blog post, or a searchable archive.

Yes. Short keeps lines to about five words for social clips, Standard fits YouTube and most players at up to ten words, and Long makes roomier lines for a podcast transcript. Every setting keeps the same word timings, so the file stays in sync whichever you pick.

It detects the spoken language automatically from the first 30 seconds and transcribes in that language. Translation into other languages is a captionrich Pro feature in the editor, which keeps the word-level timing so translated subtitles stay in sync.

Drop the video into the captionrich editor instead: it transcribes the audio itself, styles the captions in 40+ animated styles, and lets you cut the video by deleting words from the transcript. The Free plan includes 3 videos a month with no watermark. To burn an SRT you already have into a video, use the Free Subtitle Burner.

No limit is set by us. Audio is transcribed in 30-second windows, so a multi-hour episode works; the practical ceiling is your device's memory and how long you are willing to wait. Very large files are decoded track by track rather than loaded whole.

Have the video too? Edit it like a document.

Upload the video to captionrich and it transcribes the speech, styles the captions in 40+ animated styles, and cuts the footage when you delete a word from the transcript. Remove silences and filler words in one click, then export an MP4 in your browser. No watermark on any plan, 3 videos a month free.

Open the free editor

More free tools