Voice StudioVOICESTUDIO
ENGINE: LOADED//AUDIO STAYS LOCAL

Self-Hosted,OpenSource
alternativetoElevenLabs

Seven open-source TTS engines — voice cloning, multi-speaker podcasts, and 30+ languages. All on your machine. No cloud, no API keys, and audio never leaves your computer.

7 TTS ENGINES30+ LANGUAGESFULLY OFFLINE
Voice Studio — multi-speaker podcast editor with engine controls

// SEVEN OPEN-SOURCE TTS ENGINES · ONE LOCAL STUDIO

VibeVoice
Kokoro
Kitten TTS
Chatterbox
OmniVoice
VoxCPM2
Qwen3-TTS
FastAPI
PyTorch
Hugging Face
Voice Cloning
100% Offline
30+ Languages
VibeVoice
Kokoro
Kitten TTS
Chatterbox
OmniVoice
VoxCPM2
Qwen3-TTS
FastAPI
PyTorch
Hugging Face
Voice Cloning
100% Offline
30+ Languages

// CORE_CAPABILITIES

EVERYTHING, ON YOUR MACHINE

A full text-to-speech studio that runs locally — multi-engine, multi-speaker, and private by design.

READ THE DOCS
FIG. 01

PLUGGABLE ENGINES

Seven open-source engines — VibeVoice, Kokoro, Kitten TTS Mini, Chatterbox, OmniVoice, VoxCPM, and Qwen3-TTS — behind one clean interface. Only one loads at a time to keep memory low.

YOUR SCRIPTLOCAL AUDIO
FIG. 02

FULLY OFFLINE

No cloud, no telemetry, no API keys. Audio never leaves your machine.

FIG. 03

VOICE CLONING

Clone a voice from a short reference clip, or design one from a text prompt.

FIG. 04

VOICE MODES

Clone, design, or auto — per speaker.

FIG. 05

30+ LANGUAGES

Multilingual synthesis with native-script support: اردو, हिन्दी, 中文 and more.

FIG. 06

IN-APP INSTALLS

Install isolated engines and download model weights from the UI, with live progress.

FIG. 07

AUTO-UPDATES

Checks GitHub releases and applies updates in place, right from the app.

FIG. 08

FOUR PROJECT MODES

Podcast, Text-to-Voice, Transcribe (speech-to-text with Whisper), and Dub — re-voice any clip in a voice you choose.

FIG. 09

VOICE LIBRARY

Manage built-in and cloned voices with per-voice metadata and reference transcripts.

// HOW_IT_WORKS

FROM SCRIPT TO SPEECH

A pluggable engine pipeline: pick voices, load one engine, synthesize locally, and keep every byte of audio on your machine.

● 100% LOCALNO CLOUD
01_INPUT

SCRIPT & VOICES

Write a multi-speaker script or plain text, then assign a voice per speaker — built-in or cloned from a reference clip.

PODCASTMULTI
TEXT-TO-VOICESINGLE
VOICESCLONE / BUILT-IN
02_ENGINE [CORE]

PLUGGABLE ENGINE

The selected engine loads lazily on first use — only one at a time to keep memory low. Conflicting engines run in their own isolated venv.

ACTIVE1 ENGINE
ISOLATIONPER-VENV
LOADLAZY
03_SYNTHESIZE

LOCAL SYNTHESIS

Multi-speaker lines render as separate calls and join with silence gaps. Results are cached to disk so repeats are instant.

SPLITPER-LINE
CACHEON DISK
GPU/CPUAUTO
04_PRIVATE

STAYS ON DEVICE

Audio is written straight to your disk. Nothing is uploaded, streamed to a cloud, or tracked — ever.

NETWORKNONE
TELEMETRYOFF
API KEYS0

// TTS_ENGINES

SEVEN ENGINES, ONE STUDIO

Each engine has its own strengths. Only one loads at a time to keep memory low — install and download any of them right from the app.

MULTI-SPEAKER

VibeVoice-1.5B

Expressive multi-speaker synthesis for rich, natural dialogue.

~5.4 GBEnglish
LIGHTWEIGHT

Kokoro-82M

Fast and tiny, yet multilingual — great on modest hardware.

~350 MB4 languages
RUNS ANYWHERE

Kitten TTS Mini

An ~80M ONNX model that runs on CPU with no GPU — the low-end-hardware tier.

~79 MBEnglish
CLONING

Chatterbox V3

Voice cloning across a wide range of languages.

~500 MB23 languages
VOICE MODES

OmniVoice

Clone from a clip, design from a prompt, or go fully automatic.

~3.3 GBMultilingual
HI-FI

VoxCPM2

High-fidelity 48 kHz output with ultimate reference cloning.

~5 GB30 languages
BUILT-IN VOICES

Qwen3-TTS

Nine premium built-in voices with free-text style control.

~1.7B params10 languages
RUNS LOCALLYSWITCH ANYTIMEONE ACTIVE ENGINE

// UNDER_THE_HOOD

BUILT ON OPEN SOURCE

A FastAPI backend serving open-source models to a React editor — all running on your own hardware, no accounts required.

FastAPI0.115+

Backend API server

React + Vite18

Frontend editor UI

PyTorch2.4+

Model inference

Hugging Face

Model weights & cache

Transformers4.51

Engine runtimes

CUDA · MPS · CPU

Auto-detected acceleration

SYNTHESIS ENGINE
RUNNING LOCAL

ENGINE

1 ACTIVE

DEVICE

AUTO

CACHE

ON DISK

// LOG_STREAM

LIVE
00:01> load_engine(name="vibevoice")
00:01> device=cuda --dtype=fp16DONE
00:02> tokenize --speakers=3OK
00:02> synthesize --segment=1/4OK
00:03> cache_write --format=wavOK
00:03> join --silence-gap=250msOK
Local-only

No network calls

Uvicorn0.30+

ASGI server

Pydantic2.7+

Config & schemas

Python3.10+

Backend language

Tailwind CSS3.4

App styling

Disk cache

Synthesis cache

// BY_THE_NUMBERS

AT A GLANCE

LOCAL · PRIVATESTATUS: OFFLINE
0

TTS ENGINES

all open source

0+

LANGUAGES

multilingual synthesis

0kHz

SAMPLE RATE

VoxCPM2 hi-fi

0%

OFFLINE

no cloud, no API keys

// COMPLIANCE_LOGFULLY OFFLINE[VERIFIED]OPEN SOURCE[ACTIVE]NO API KEYS[COMPLIANT]

// GET_STARTED

UP AND RUNNING IN TWO COMMANDS

Python 3.10+ and Node 18+. Setup auto-detects your GPU and installs the matching PyTorch build, then lets you pick which models to download.

git clone https://github.com/msrbuilds/voice-studio.git
cd voice-studio
python studio.py setup     # venv, deps, CUDA auto-detect, model picker
python studio.py start     # backend + frontend
1Clone the repo2python studio.py setup3python studio.py start

100% OFFLINE • NO API KEYS • OPEN SOURCE