Voice Studio

Dub Mode

Re-voice an audio clip in any voice — voice-to-voice dubbing.

Dub mode re-voices an audio clip in a voice you choose. Drop a clip, transcribe it, tidy the text, pick a voice, and Voice Studio regenerates the whole track in that voice — keeping the original timing. No extra model — it reuses the Transcribe (Whisper) and synthesis pipelines.

How it works

  1. Drop an audio clip. It's transcribed with Whisper into timed segments.
  2. Edit the per-segment text if the transcript needs a fix.
  3. Pick a target voice in the library (any voice the active engine supports).
  4. Generate Dub — each segment is re-synthesized in that voice, then laid back on the original timeline. It plays automatically, and you can Download the WAV.

Natural timing

Takes are placed back at their original positions: the leading silence matches the first segment's start, and the pause between segments matches the original gap. Each take plays at its natural length — audio is never time-stretched, so it always sounds clean.

Scope (v1)

Dub mode is same-language re-voicing: audio in, audio out, one voice for the whole clip. It reuses the same Download Whisper gate as Transcribe mode. Cross-language dubbing (translation), video, and speaker diarization are planned for later.

Tip

Any engine works, but built-in-voice engines like Kokoro or Kitten TTS Mini run on CPU and are a fast way to try dubbing without a GPU.

On this page