Voice Studio

Models & Cache

Weight downloads and the synthesis cache.

Model weights

Weights aren't bundled — each engine downloads on demand into a local Hugging Face cache under backend/models/. When an engine's weights are missing, the selector shows a Download button instead of Switch, and a dialog streams live progress (percent, speed, ETA). Downloads run on their own thread, so they never block synthesis.

You can reclaim space at any time: Delete weights removes a model's files (re-downloadable later), and for isolated engines Uninstall removes the whole virtual environment. Both are in the engine menu. Whisper (speech-to-text) is managed the same way — its ~1.6 GB downloads and deletes like any other model.

Approximate sizes: Kitten ~79 MB · Kokoro ~350 MB · Chatterbox ~500 MB · Whisper ~1.6 GB · Qwen3-TTS ~3.5 GB · OmniVoice ~3.3 GB · VoxCPM2 ~5 GB · VibeVoice ~5.4 GB.

Synthesis cache

Every generated clip is cached to disk under backend/cache/, keyed by a content hash of the text + voice + engine + all settings (cfg, exaggeration, language, voice mode, style prompt, Qwen sampling params, and so on). Because the key folds in the engine and every knob, results never collide across engines or settings — and a repeat request returns instantly instead of hitting the GPU.

  • Per-segment takes live in the cache root; joined podcast downloads and dubbed tracks live in their own subfolders.
  • The cache is LRU-bounded (500 entries by default) so disk stays predictable.
  • It survives browser refreshes, model reloads, and server restarts.

The Recent generations panel in the right-hand Control Panel lists cached clips — play, download, or delete individual entries, Clear all, or open the cache folder in your file manager.

Disable it or change its size with CACHE_ENABLED and CACHE_MAX_ENTRIES, and relocate it with CACHE_DIR — see Configuration.

Where things live

DirectoryContents
backend/models/Downloaded model weights (Hugging Face cache)
backend/cache/Synthesized audio (per-segment, downloads/, dub/)
backend/voices/Built-in reference voices (scanned at startup)
backend/uploads/Voices you upload from the UI

Back these up (or mount them on a persistent volume) to keep your voices and downloaded weights across machines — see Deploy on a VPS.

On this page