FAQ
Frequently asked questions.
Is it really offline? Yes — after models are downloaded, no audio or text leaves your machine, and there are no API keys or telemetry.
Do I need a GPU? No. CUDA (NVIDIA) and MPS (Apple Silicon) are used when available; CPU works too, just slower. On a CPU-only machine, stick to Kokoro and Kitten TTS Mini (Kitten is CPU-native and runs anywhere).
How much VRAM do I need? About ~3 GB for VibeVoice in fp16, up to ~8 GB for VoxCPM2. Whisper adds ~1.6 GB while transcribing. Only one engine loads at a time, which keeps the footprint low.
Why only one engine at a time? Engines have large, mutually incompatible dependency stacks and heavy weights. Loading one at a time keeps memory predictable; switching engines unloads the previous one.
Can I use my own voice? Yes — VibeVoice, Chatterbox, OmniVoice, and VoxCPM clone from a short reference clip. See Voice Library.
Which engine should I pick? Start with Kokoro (fast, small) or VibeVoice (expressive multi-speaker); see Engines Overview.
Is there a character limit? Yes — 5,000 characters per request, counted
across the whole script (all speakers combined). Longer text is rejected, so
split it into lines, or use Podcast mode. Raise the cap with
MAX_TEXT_CHARS (see Configuration).
Can I transcribe or add subtitles? Yes — Transcribe mode
turns audio into text with Whisper (99 languages) and exports .srt / .vtt,
and you can subtitle audio you generated too.
Can I re-voice an existing recording? Yes — Dub mode transcribes a clip and regenerates it in a voice you choose, same-language, with the original timing preserved.
Can I run it on a server? Yes — see Deploy on a VPS. Note it has no built-in authentication, so put it behind a reverse proxy with auth.
How do I update? Use the in-app updater (About → Check for updates) or
git pull and re-run setup. See Updating.
Is it free? What license? Yes, it's free and open source — Voice Studio's own code is MIT-licensed. Each bundled model keeps its own license (MIT, Apache-2.0, etc.) and usage policy, so review those before redistributing generated audio.
Still stuck? Open an issue on GitHub.