Introduction
What Voice Studio is and why it exists.
Voice Studio by MSR is a fully-offline, local web UI for multiple open-source text-to-speech engines. Everything runs on your machine — no cloud, no API keys, and audio never leaves your computer.
Engines at a glance
- VibeVoice-1.5B — expressive multi-speaker synthesis
- Kokoro-82M — fast, lightweight, multilingual
- Kitten TTS Mini — ultra-light ~80M ONNX model, runs on CPU with no GPU
- Chatterbox Multilingual V3 — 23 languages with voice cloning
- OmniVoice — clone / design / auto voice modes
- VoxCPM2 — 2B model, 48 kHz, 30 languages
- Qwen3-TTS CustomVoice — 9 premium built-in voices, 10 languages
Only one engine is loaded at a time to keep memory low.
Four project modes
- Podcast — a multi-speaker segment editor.
- Text-to-Voice — single-voice narration from one textarea.
- Transcribe — speech-to-text with Whisper, plus subtitles.
- Dub — re-voice an audio clip in any voice you choose.