Voice Library & Cloning
Built-in voices, cloning, and voice metadata.
The voice library (left column) holds every voice the active engine can use — its built-in voices, plus any you've uploaded or cloned. The list is filtered to the active engine, so you only ever see voices that will actually work.
Built-in vs. cloned voices
- Built-in-voice engines — Kokoro, Kitten TTS Mini, and Qwen3-TTS ship a fixed catalog of designed voices and can't clone arbitrary clips. They expose only their own voices.
- Cloning engines — VibeVoice, Chatterbox, OmniVoice, and VoxCPM clone a voice from a short reference clip, so they show your uploaded/reference voices.
Add your own voice
Upload a reference clip (.wav, .mp3, .flac, or .ogg) from the library.
For a good clone:
- 10–30 seconds of clean speech is plenty.
- One speaker, minimal background noise, no music.
- A natural, consistent speaking tone — the clone mirrors what it hears.
You can also drop files straight into backend/voices/ to register them as
built-in voices (restart the backend afterward so they're scanned).
Voice metadata & reference transcripts
Each voice carries editable metadata — name, gender, language, and an optional reference transcript. The transcript drives VoxCPM's ultimate cloning (reference clip + exact transcript for the highest-fidelity clone). You can type it, or click Transcribe in the voice dialog to fill it automatically with Whisper, review it, and save.
Voice modes
Cloning engines with modes (OmniVoice, VoxCPM) let each speaker choose clone, design, or auto — so you don't always need a reference clip. See Engines Overview for the per-engine modes, and Dub mode to re-voice an existing recording in any of them.