Quick Start
Your first synthesis in Podcast and Text-to-Voice modes.
Voice Studio has four project modes: Podcast, Text-to-Voice, Transcribe, and Dub. The two synthesis modes work on every engine — the backend splits multi-speaker scripts per line, so you can use either with any model.
Choose an engine
Use the engine selector in the right-hand Control Panel. Only one engine loads at a time. If an engine isn't installed yet, the selector offers Install (isolated engines) or Download (weights) with live progress. The first load of an in-process engine downloads its weights to the local Hugging Face cache.
Podcast mode
A multi-speaker segment editor:
- Add speakers in the Speaker Roster and assign each a voice.
- Write one segment per line as
Speaker N: .... - Hit synthesize — each line is generated as a separate call, then joined with silence gaps into one track.
Great for interviews, panels, and dialogue.
Text-to-Voice mode
A single textarea with one active voice, plus live character / word / duration counts. Pick the active voice in the library, type your text, and synthesize.
Great for narration, articles, and quick one-off clips.
Transcribe mode
Drop an audio file and Voice Studio transcribes it locally with Whisper — an
editable transcript you can copy, send to Text-to-Voice, or export as .srt /
.vtt subtitles. See Transcribe mode.
Dub mode
Drop a clip, transcribe it, pick a voice, and re-voice the whole track in that voice with the original timing preserved. See Dub mode.
Switching modes
A Mode Toggle switches between the four modes at any time — each mode keeps its own buffer, so you won't lose your work when you switch.