Deploy on a VPS
Run Voice Studio on a remote Linux server (GPU or CPU), behind HTTPS.
Voice Studio is a local-first app, but it runs just as well on a remote Linux server — handy if you want a always-on studio, a shared team instance, or simply a machine with a bigger GPU than your laptop. This guide sets it up on Ubuntu 22.04, as a systemd service, behind an HTTPS reverse proxy.
Voice Studio has no built-in authentication. Anyone who can reach the port can use your GPU, upload audio, and read everything in your cache. Never expose port 8880 directly to the internet — always put it behind a reverse proxy with auth (or a VPN / firewall allowlist), as shown below.
Server requirements
| GPU server (recommended) | CPU-only server | |
|---|---|---|
| Use for | Every engine | Kokoro & Kitten TTS Mini |
| GPU | NVIDIA, CUDA 12.x driver | — |
| VRAM | ~3 GB (VibeVoice) → ~8 GB (VoxCPM2) | — |
| RAM | 8 GB+ | 4 GB+ |
| Disk | 40–80 GB (weights + venvs) | 15–20 GB |
| OS | Ubuntu 22.04 LTS | Ubuntu 22.04 LTS |
The heavier engines (VibeVoice, VoxCPM2, OmniVoice, Chatterbox) really want a GPU. On a CPU-only box, stick to Kokoro (~350 MB) and Kitten TTS Mini (~79 MB, ONNX, CPU-native) — both are fast enough to be usable. See Engines Overview for sizes.
1. Base packages
sudo apt-get update
sudo apt-get install -y git python3 python3-venv python3-pip ffmpeg espeak-ng curl
# Node.js 20 (for the frontend build)
curl -fsSL https://deb.nodesource.com/setup_20.x | sudo -E bash -
sudo apt-get install -y nodejsespeak-ng is required by Kokoro (without it Kokoro is silent); ffmpeg
handles audio I/O. See Installation for details.
2. NVIDIA driver + CUDA (GPU servers only)
Install a recent NVIDIA driver and confirm the GPU is visible:
sudo apt-get install -y nvidia-driver-550 # or your cloud image's recommended driver
nvidia-smi # should list your GPUYou don't need a system-wide CUDA toolkit — studio.py setup installs a
CUDA-matched PyTorch wheel into the app's virtual environment. Just make sure the
driver is new enough for CUDA 12.x.
3. Clone and set up
Run as a non-root user (e.g. ubuntu). studio.py setup is interactive (it
confirms the PyTorch build and lets you pick which models to download), so run it
in your SSH session:
cd ~
git clone https://github.com/msrbuilds/voice-studio.git
cd voice-studio
python3 studio.py setupSetup creates the virtual environment, auto-detects the GPU and installs the
right PyTorch build, installs dependencies, and builds the frontend. On a GPU box
it should report Detected NVIDIA GPU → CUDA wheel.
Verify it runs, then stop it with Ctrl+C:
python3 studio.py start --prod # serves UI + API on http://127.0.0.1:8880Once frontend/dist exists, the backend serves the built UI itself — so the
systemd service below can run backend.cli directly and still serve the full
app on one port.
4. Run it as a systemd service
Bind to 127.0.0.1 so only the reverse proxy can reach it. Create
/etc/systemd/system/voice-studio.service:
[Unit]
Description=Voice Studio by MSR
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=ubuntu
WorkingDirectory=/home/ubuntu/voice-studio
ExecStart=/home/ubuntu/voice-studio/backend/venv/bin/python -m backend.cli \
--host 127.0.0.1 --port 8880 --device cuda
Restart=on-failure
RestartSec=5
# First model download can be large; give it time before the timeout bites.
TimeoutStartSec=0
[Install]
WantedBy=multi-user.targetOn a CPU-only server use --device cpu. Then:
sudo systemctl daemon-reload
sudo systemctl enable --now voice-studio
sudo systemctl status voice-studio
journalctl -u voice-studio -f # follow logs (model load, requests)5. Reverse proxy + HTTPS + auth
Caddy is the shortest path — automatic HTTPS and one-line basic auth. Install
Caddy, generate a password hash with caddy hash-password, then use a
Caddyfile:
studio.example.com {
# Protect the whole app — Voice Studio has no auth of its own.
basic_auth {
youruser JDJhJDE0J...your-hash...
}
reverse_proxy 127.0.0.1:8880
request_body {
max_size 120MB # ASR uploads can be up to 100 MB
}
}Prefer nginx? The key extras are a large upload limit, long timeouts (some
generations take a while), and the WebSocket upgrade for /api/stream:
server {
server_name studio.example.com;
client_max_body_size 120M;
proxy_read_timeout 900s; # long synthesis / transcription jobs
location / {
proxy_pass http://127.0.0.1:8880;
proxy_set_header Host $host;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
auth_basic "Voice Studio";
auth_basic_user_file /etc/nginx/.htpasswd; # htpasswd -c ... youruser
}
}
# Run `certbot --nginx -d studio.example.com` for TLS.6. Firewall
Allow only HTTP/HTTPS and SSH; keep 8880 private:
sudo ufw allow OpenSSH
sudo ufw allow 80,443/tcp
sudo ufw enable
# Do NOT `ufw allow 8880` — it must stay reachable only from the proxy.7. Download model weights
Weights aren't bundled. Open the site, go to the engine selector, and hit Download on the engines you want — a live progress bar streams the download into the server's Hugging Face cache. Isolated engines (Chatterbox, OmniVoice, VoxCPM, Qwen3-TTS) also need a one-time Install to build their environment; see Installing engines. The first VibeVoice download is ~5.4 GB, so give it a minute.
Updating
git pull and re-run setup, or use the in-app updater (About → Check for
updates). See Updating — the in-app path guards for a clean git
checkout, then rebuilds the frontend for you.
Notes & limits
- One engine at a time. Switching engines unloads the previous one, so a single instance serves one engine at a time. Requests already serialize onto one GPU worker, so a shared instance queues rather than oversubscribes.
- Not multi-tenant. There are no per-user accounts or quotas — everyone behind the proxy shares the same voices, cache, and GPU.
- Keep the data dirs.
backend/voices,backend/uploads,backend/cache, andbackend/modelshold your voices and downloaded weights — back them up or mount them on a persistent volume if your VPS reprovisions.