Voice Studio

Quick Start

Your first synthesis in Podcast and Text-to-Voice modes.

Voice Studio has four project modes: Podcast, Text-to-Voice, Transcribe, and Dub. The two synthesis modes work on every engine — the backend splits multi-speaker scripts per line, so you can use either with any model.

Choose an engine

Use the engine selector in the right-hand Control Panel. Only one engine loads at a time. If an engine isn't installed yet, the selector offers Install (isolated engines) or Download (weights) with live progress. The first load of an in-process engine downloads its weights to the local Hugging Face cache.

Podcast mode

A multi-speaker segment editor:

  1. Add speakers in the Speaker Roster and assign each a voice.
  2. Write one segment per line as Speaker N: ....
  3. Hit synthesize — each line is generated as a separate call, then joined with silence gaps into one track.

Great for interviews, panels, and dialogue.

Text-to-Voice mode

A single textarea with one active voice, plus live character / word / duration counts. Pick the active voice in the library, type your text, and synthesize.

Great for narration, articles, and quick one-off clips.

Transcribe mode

Drop an audio file and Voice Studio transcribes it locally with Whisper — an editable transcript you can copy, send to Text-to-Voice, or export as .srt / .vtt subtitles. See Transcribe mode.

Dub mode

Drop a clip, transcribe it, pick a voice, and re-voice the whole track in that voice with the original timing preserved. See Dub mode.

Switching modes

A Mode Toggle switches between the four modes at any time — each mode keeps its own buffer, so you won't lose your work when you switch.

Next steps

On this page