Podcast Mode
The multi-speaker segment editor.
Podcast mode is a multi-speaker segment editor. You define a cast of speakers, assign each a voice, and write the conversation one segment at a time. Each segment is synthesized as its own call, then joined into a single track with short silence gaps — so multi-voice dialogue stays clean and each line is independently editable.
Build a cast
In the Speaker Roster at the top, add a speaker and assign it a voice from the library. Add as many as your engine supports — VibeVoice handles up to four distinct voices in one project. For engines with voice modes (OmniVoice, VoxCPM), each speaker also picks its own mode (clone / design / auto); see Engines Overview.
Write segments
Each segment is one line, attributed to a speaker. Under the hood the script is
normalized to Speaker N: ... form, and every line is synthesized separately —
which is exactly what keeps different voices from bleeding together.
- Generate a single segment to preview it.
- Generate All walks the project and synthesizes every uncached segment.
- Regenerate forces a fresh take even when the text and voice are unchanged (useful for dialing in expressive engines).
Generated audio is cached per segment, so re-playing or re-exporting is instant unless you change the text, voice, or settings. See Models & Cache.
Play and export
The inline player plays segments in order (Play All), synthesizing any that aren't cached yet. When you're happy:
- Export the joined track as a single WAV.
- Export subtitles (
.srt) — Voice Studio transcribes the rendered track with Whisper and writes timed captions. - Import / Export JSON saves the whole project (segments + speakers) so you can reload it later.
Prefer a single narrator? Use Text-to-Voice mode instead — both modes work on every engine.