Installation
Prerequisites, system dependencies, and setup.
Prerequisites
- Python 3.10+ and Node.js 18+
- PyTorch — CUDA on Windows/Linux, CPU-only (slower), or Apple Silicon (MPS, experimental)
- Disk for model weights (per engine): Kokoro ~350 MB · Chatterbox ~500 MB · VoxCPM2 ~5 GB · VibeVoice ~5.4 GB · OmniVoice ~3.3 GB
- VoxCPM requires Python 3.10–3.12
System dependencies
Two native tools are used at runtime and are not installed by pip.
python studio.py setup checks for them and prints the right command for your OS —
you can also install them yourself:
espeak-ng— required by Kokoro for text phonemization. Without it on yourPATH, Kokoro produces silent audio. The other engines don't need it.ffmpeg— used for some audio I/O.
| OS | espeak-ng | ffmpeg |
|---|---|---|
| Windows | winget install eSpeak-NG.eSpeak-NG | winget install Gyan.FFmpeg |
| macOS | brew install espeak-ng | brew install ffmpeg |
| Linux (Debian/Ubuntu) | sudo apt-get install espeak-ng | sudo apt-get install ffmpeg |
Restart the backend after installing so it picks them up on PATH.
Quick setup (recommended)
From the repo root, one command bootstraps everything — it creates the Python virtual environment, auto-detects your GPU and installs the matching PyTorch/CUDA build, installs dependencies, and lets you pick which models to download:
git clone https://github.com/msrbuilds/voice-studio.git
cd voice-studio
python studio.py setup
python studio.py startOpen http://localhost:5173 (dev) or http://localhost:8880 (prod).
On Windows, install a CUDA PyTorch wheel before the dependency install, or CUDA silently falls back to CPU.
studio.py setuphandles this for you.
Isolated engines
Chatterbox, OmniVoice, VoxCPM, and Qwen3-TTS each need a different (and mutually
incompatible) transformers / torch stack, so each runs in its own virtual
environment as a subprocess. Install them from the app — see
Installing engines.