Voice Studio

Engines Overview

How engines load, and what each one does.

Voice Studio loads one engine at a time to keep memory low. Four engines run in their own isolated virtual environments as subprocesses, because their transformers / torch pins are mutually incompatible: Chatterbox, OmniVoice, VoxCPM, and Qwen3-TTS. VibeVoice, Kokoro, and Kitten TTS Mini run in the main environment (Kitten is ONNX-based and needs no GPU).

Engines are constructed at startup but loaded lazily on first use. Switching engines unloads the previous one, and your choice is remembered across restarts.

EngineSizeLanguagesNotes
VibeVoice-1.5B~5.4 GBEnglishExpressive multi-speaker
Kokoro-82M~350 MB4Fast, lightweight
Kitten TTS Mini~79 MBEnglishONNX, CPU-only, 8 voices — runs anywhere
Chatterbox V3~500 MB23Voice cloning
OmniVoice~3.3 GBMultilingualClone / design / auto modes
VoxCPM2~5 GB3048 kHz, ultimate cloning
Qwen3-TTS~1.7B params10Built-in voices + style prompt

Voice modes

Some engines support per-speaker voice modes:

  • OmniVoice — clone (reference clip), design (a free-text attribute prompt), or auto (no prompt).
  • VoxCPM — auto, design (inline style), clone, controllable clone (clone + inline style), and ultimate clone (reference + transcript).
  • Qwen3-TTS — built-in voices only (no cloning), with an always-available free-text style prompt and a Qwen-only advanced generation panel (temperature, top_p, top_k, repetition penalty, seed).

VibeVoice and Kokoro use their voice library directly without mode toggles.

Reference transcripts

Voices can carry an optional reference transcript (editable in the voice meta dialog) that drives VoxCPM's ultimate cloning. See Voice library & cloning.

On this page