Voice Studio

FAQ

Frequently asked questions.

Is it really offline? Yes — after models are downloaded, no audio or text leaves your machine, and there are no API keys or telemetry.

Do I need a GPU? No. CUDA (NVIDIA) and MPS (Apple Silicon) are used when available; CPU works too, just slower. On a CPU-only machine, stick to Kokoro and Kitten TTS Mini (Kitten is CPU-native and runs anywhere).

How much VRAM do I need? About ~3 GB for VibeVoice in fp16, up to ~8 GB for VoxCPM2. Whisper adds ~1.6 GB while transcribing. Only one engine loads at a time, which keeps the footprint low.

Why only one engine at a time? Engines have large, mutually incompatible dependency stacks and heavy weights. Loading one at a time keeps memory predictable; switching engines unloads the previous one.

Can I use my own voice? Yes — VibeVoice, Chatterbox, OmniVoice, and VoxCPM clone from a short reference clip. See Voice Library.

Which engine should I pick? Start with Kokoro (fast, small) or VibeVoice (expressive multi-speaker); see Engines Overview.

Is there a character limit? Yes — 5,000 characters per request, counted across the whole script (all speakers combined). Longer text is rejected, so split it into lines, or use Podcast mode. Raise the cap with MAX_TEXT_CHARS (see Configuration).

Can I transcribe or add subtitles? Yes — Transcribe mode turns audio into text with Whisper (99 languages) and exports .srt / .vtt, and you can subtitle audio you generated too.

Can I re-voice an existing recording? Yes — Dub mode transcribes a clip and regenerates it in a voice you choose, same-language, with the original timing preserved.

Can I run it on a server? Yes — see Deploy on a VPS. Note it has no built-in authentication, so put it behind a reverse proxy with auth.

How do I update? Use the in-app updater (About → Check for updates) or git pull and re-run setup. See Updating.

Is it free? What license? Yes, it's free and open source — Voice Studio's own code is MIT-licensed. Each bundled model keeps its own license (MIT, Apache-2.0, etc.) and usage policy, so review those before redistributing generated audio.

Still stuck? Open an issue on GitHub.