Transcribe Mode
Speech-to-text with Whisper — editable transcripts and subtitles.
Transcribe mode turns an audio file into text. Drop a clip, and Voice Studio transcribes it locally with Whisper large-v3-turbo — 99 languages, fully offline once the weights are downloaded.
How it works
- Drop an audio file (WAV, MP3, FLAC, OGG, M4A, WebM).
- Optionally pick a language (or leave it on Auto-detect) and toggle timestamps.
- Transcribe — you get an editable transcript you can fix inline.
From the transcript you can Copy it, Send to Text-to-Voice to re-synthesize
it in a chosen voice, or export .srt / .vtt subtitles.
First run
Whisper's weights (~1.6 GB) download on demand — the first time you enter Transcribe mode you'll see a Download Whisper button. It's the same download / delete flow as any other model (see Models & cache).
Subtitles for audio you generated
You can also subtitle audio you synthesized: the Subtitles (.srt) export in the toolbar transcribes the rendered podcast or single take and writes a SubRip file.
Whisper runs in-process and shares the GPU with synthesis, so transcription and generation never run at the same time.