Engines Overview
How engines load, and what each one does.
Voice Studio loads one engine at a time to keep memory low. Four engines run
in their own isolated virtual environments as subprocesses, because their
transformers / torch pins are mutually incompatible: Chatterbox,
OmniVoice, VoxCPM, and Qwen3-TTS. VibeVoice, Kokoro, and
Kitten TTS Mini run in the main environment (Kitten is ONNX-based and needs
no GPU).
Engines are constructed at startup but loaded lazily on first use. Switching engines unloads the previous one, and your choice is remembered across restarts.
| Engine | Size | Languages | Notes |
|---|---|---|---|
| VibeVoice-1.5B | ~5.4 GB | English | Expressive multi-speaker |
| Kokoro-82M | ~350 MB | 4 | Fast, lightweight |
| Kitten TTS Mini | ~79 MB | English | ONNX, CPU-only, 8 voices — runs anywhere |
| Chatterbox V3 | ~500 MB | 23 | Voice cloning |
| OmniVoice | ~3.3 GB | Multilingual | Clone / design / auto modes |
| VoxCPM2 | ~5 GB | 30 | 48 kHz, ultimate cloning |
| Qwen3-TTS | ~1.7B params | 10 | Built-in voices + style prompt |
Voice modes
Some engines support per-speaker voice modes:
- OmniVoice — clone (reference clip), design (a free-text attribute prompt), or auto (no prompt).
- VoxCPM — auto, design (inline style), clone, controllable clone (clone + inline style), and ultimate clone (reference + transcript).
- Qwen3-TTS — built-in voices only (no cloning), with an always-available free-text style prompt and a Qwen-only advanced generation panel (temperature, top_p, top_k, repetition penalty, seed).
VibeVoice and Kokoro use their voice library directly without mode toggles.
Reference transcripts
Voices can carry an optional reference transcript (editable in the voice meta dialog) that drives VoxCPM's ultimate cloning. See Voice library & cloning.