Voice Studio

Transcribe Mode

Speech-to-text with Whisper — editable transcripts and subtitles.

Transcribe mode turns an audio file into text. Drop a clip, and Voice Studio transcribes it locally with Whisper large-v3-turbo — 99 languages, fully offline once the weights are downloaded.

How it works

  1. Drop an audio file (WAV, MP3, FLAC, OGG, M4A, WebM).
  2. Optionally pick a language (or leave it on Auto-detect) and toggle timestamps.
  3. Transcribe — you get an editable transcript you can fix inline.

From the transcript you can Copy it, Send to Text-to-Voice to re-synthesize it in a chosen voice, or export .srt / .vtt subtitles.

First run

Whisper's weights (~1.6 GB) download on demand — the first time you enter Transcribe mode you'll see a Download Whisper button. It's the same download / delete flow as any other model (see Models & cache).

Subtitles for audio you generated

You can also subtitle audio you synthesized: the Subtitles (.srt) export in the toolbar transcribes the rendered podcast or single take and writes a SubRip file.

Whisper runs in-process and shares the GPU with synthesis, so transcription and generation never run at the same time.

On this page