Voice Studio

Installation

Prerequisites, system dependencies, and setup.

Prerequisites

  • Python 3.10+ and Node.js 18+
  • PyTorch — CUDA on Windows/Linux, CPU-only (slower), or Apple Silicon (MPS, experimental)
  • Disk for model weights (per engine): Kokoro ~350 MB · Chatterbox ~500 MB · VoxCPM2 ~5 GB · VibeVoice ~5.4 GB · OmniVoice ~3.3 GB
  • VoxCPM requires Python 3.10–3.12

System dependencies

Two native tools are used at runtime and are not installed by pip. python studio.py setup checks for them and prints the right command for your OS — you can also install them yourself:

  • espeak-ng — required by Kokoro for text phonemization. Without it on your PATH, Kokoro produces silent audio. The other engines don't need it.
  • ffmpeg — used for some audio I/O.
OSespeak-ngffmpeg
Windowswinget install eSpeak-NG.eSpeak-NGwinget install Gyan.FFmpeg
macOSbrew install espeak-ngbrew install ffmpeg
Linux (Debian/Ubuntu)sudo apt-get install espeak-ngsudo apt-get install ffmpeg

Restart the backend after installing so it picks them up on PATH.

From the repo root, one command bootstraps everything — it creates the Python virtual environment, auto-detects your GPU and installs the matching PyTorch/CUDA build, installs dependencies, and lets you pick which models to download:

git clone https://github.com/msrbuilds/voice-studio.git
cd voice-studio
python studio.py setup
python studio.py start

Open http://localhost:5173 (dev) or http://localhost:8880 (prod).

On Windows, install a CUDA PyTorch wheel before the dependency install, or CUDA silently falls back to CPU. studio.py setup handles this for you.

Isolated engines

Chatterbox, OmniVoice, VoxCPM, and Qwen3-TTS each need a different (and mutually incompatible) transformers / torch stack, so each runs in its own virtual environment as a subprocess. Install them from the app — see Installing engines.

On this page