Voice Studio

Podcast Mode

The multi-speaker segment editor.

Podcast mode is a multi-speaker segment editor. You define a cast of speakers, assign each a voice, and write the conversation one segment at a time. Each segment is synthesized as its own call, then joined into a single track with short silence gaps — so multi-voice dialogue stays clean and each line is independently editable.

Build a cast

In the Speaker Roster at the top, add a speaker and assign it a voice from the library. Add as many as your engine supports — VibeVoice handles up to four distinct voices in one project. For engines with voice modes (OmniVoice, VoxCPM), each speaker also picks its own mode (clone / design / auto); see Engines Overview.

Write segments

Each segment is one line, attributed to a speaker. Under the hood the script is normalized to Speaker N: ... form, and every line is synthesized separately — which is exactly what keeps different voices from bleeding together.

  • Generate a single segment to preview it.
  • Generate All walks the project and synthesizes every uncached segment.
  • Regenerate forces a fresh take even when the text and voice are unchanged (useful for dialing in expressive engines).

Generated audio is cached per segment, so re-playing or re-exporting is instant unless you change the text, voice, or settings. See Models & Cache.

Play and export

The inline player plays segments in order (Play All), synthesizing any that aren't cached yet. When you're happy:

  • Export the joined track as a single WAV.
  • Export subtitles (.srt) — Voice Studio transcribes the rendered track with Whisper and writes timed captions.
  • Import / Export JSON saves the whole project (segments + speakers) so you can reload it later.

Prefer a single narrator? Use Text-to-Voice mode instead — both modes work on every engine.

On this page