Voice Studio

Introduction

What Voice Studio is and why it exists.

Voice Studio by MSR is a fully-offline, local web UI for multiple open-source text-to-speech engines. Everything runs on your machine — no cloud, no API keys, and audio never leaves your computer.

Engines at a glance

  • VibeVoice-1.5B — expressive multi-speaker synthesis
  • Kokoro-82M — fast, lightweight, multilingual
  • Kitten TTS Mini — ultra-light ~80M ONNX model, runs on CPU with no GPU
  • Chatterbox Multilingual V3 — 23 languages with voice cloning
  • OmniVoice — clone / design / auto voice modes
  • VoxCPM2 — 2B model, 48 kHz, 30 languages
  • Qwen3-TTS CustomVoice — 9 premium built-in voices, 10 languages

Only one engine is loaded at a time to keep memory low.

Four project modes

  • Podcast — a multi-speaker segment editor.
  • Text-to-Voice — single-voice narration from one textarea.
  • Transcribe — speech-to-text with Whisper, plus subtitles.
  • Dub — re-voice an audio clip in any voice you choose.

On this page