Voice Studio

Deploy on a VPS

Run Voice Studio on a remote Linux server (GPU or CPU), behind HTTPS.

Voice Studio is a local-first app, but it runs just as well on a remote Linux server — handy if you want a always-on studio, a shared team instance, or simply a machine with a bigger GPU than your laptop. This guide sets it up on Ubuntu 22.04, as a systemd service, behind an HTTPS reverse proxy.

Voice Studio has no built-in authentication. Anyone who can reach the port can use your GPU, upload audio, and read everything in your cache. Never expose port 8880 directly to the internet — always put it behind a reverse proxy with auth (or a VPN / firewall allowlist), as shown below.

Server requirements

GPU server (recommended)CPU-only server
Use forEvery engineKokoro & Kitten TTS Mini
GPUNVIDIA, CUDA 12.x driver—
VRAM~3 GB (VibeVoice) → ~8 GB (VoxCPM2)—
RAM8 GB+4 GB+
Disk40–80 GB (weights + venvs)15–20 GB
OSUbuntu 22.04 LTSUbuntu 22.04 LTS

The heavier engines (VibeVoice, VoxCPM2, OmniVoice, Chatterbox) really want a GPU. On a CPU-only box, stick to Kokoro (~350 MB) and Kitten TTS Mini (~79 MB, ONNX, CPU-native) — both are fast enough to be usable. See Engines Overview for sizes.

1. Base packages

sudo apt-get update
sudo apt-get install -y git python3 python3-venv python3-pip ffmpeg espeak-ng curl
# Node.js 20 (for the frontend build)
curl -fsSL https://deb.nodesource.com/setup_20.x | sudo -E bash -
sudo apt-get install -y nodejs

espeak-ng is required by Kokoro (without it Kokoro is silent); ffmpeg handles audio I/O. See Installation for details.

2. NVIDIA driver + CUDA (GPU servers only)

Install a recent NVIDIA driver and confirm the GPU is visible:

sudo apt-get install -y nvidia-driver-550   # or your cloud image's recommended driver
nvidia-smi                                   # should list your GPU

You don't need a system-wide CUDA toolkit — studio.py setup installs a CUDA-matched PyTorch wheel into the app's virtual environment. Just make sure the driver is new enough for CUDA 12.x.

3. Clone and set up

Run as a non-root user (e.g. ubuntu). studio.py setup is interactive (it confirms the PyTorch build and lets you pick which models to download), so run it in your SSH session:

cd ~
git clone https://github.com/msrbuilds/voice-studio.git
cd voice-studio
python3 studio.py setup

Setup creates the virtual environment, auto-detects the GPU and installs the right PyTorch build, installs dependencies, and builds the frontend. On a GPU box it should report Detected NVIDIA GPU → CUDA wheel.

Verify it runs, then stop it with Ctrl+C:

python3 studio.py start --prod   # serves UI + API on http://127.0.0.1:8880

Once frontend/dist exists, the backend serves the built UI itself — so the systemd service below can run backend.cli directly and still serve the full app on one port.

4. Run it as a systemd service

Bind to 127.0.0.1 so only the reverse proxy can reach it. Create /etc/systemd/system/voice-studio.service:

[Unit]
Description=Voice Studio by MSR
After=network-online.target
Wants=network-online.target

[Service]
Type=simple
User=ubuntu
WorkingDirectory=/home/ubuntu/voice-studio
ExecStart=/home/ubuntu/voice-studio/backend/venv/bin/python -m backend.cli \
  --host 127.0.0.1 --port 8880 --device cuda
Restart=on-failure
RestartSec=5
# First model download can be large; give it time before the timeout bites.
TimeoutStartSec=0

[Install]
WantedBy=multi-user.target

On a CPU-only server use --device cpu. Then:

sudo systemctl daemon-reload
sudo systemctl enable --now voice-studio
sudo systemctl status voice-studio
journalctl -u voice-studio -f     # follow logs (model load, requests)

5. Reverse proxy + HTTPS + auth

Caddy is the shortest path — automatic HTTPS and one-line basic auth. Install Caddy, generate a password hash with caddy hash-password, then use a Caddyfile:

studio.example.com {
    # Protect the whole app — Voice Studio has no auth of its own.
    basic_auth {
        youruser JDJhJDE0J...your-hash...
    }
    reverse_proxy 127.0.0.1:8880
    request_body {
        max_size 120MB          # ASR uploads can be up to 100 MB
    }
}

Prefer nginx? The key extras are a large upload limit, long timeouts (some generations take a while), and the WebSocket upgrade for /api/stream:

server {
    server_name studio.example.com;

    client_max_body_size 120M;
    proxy_read_timeout 900s;      # long synthesis / transcription jobs

    location / {
        proxy_pass http://127.0.0.1:8880;
        proxy_set_header Host $host;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "upgrade";
        auth_basic "Voice Studio";
        auth_basic_user_file /etc/nginx/.htpasswd;   # htpasswd -c ... youruser
    }
}
# Run `certbot --nginx -d studio.example.com` for TLS.

6. Firewall

Allow only HTTP/HTTPS and SSH; keep 8880 private:

sudo ufw allow OpenSSH
sudo ufw allow 80,443/tcp
sudo ufw enable
# Do NOT `ufw allow 8880` — it must stay reachable only from the proxy.

7. Download model weights

Weights aren't bundled. Open the site, go to the engine selector, and hit Download on the engines you want — a live progress bar streams the download into the server's Hugging Face cache. Isolated engines (Chatterbox, OmniVoice, VoxCPM, Qwen3-TTS) also need a one-time Install to build their environment; see Installing engines. The first VibeVoice download is ~5.4 GB, so give it a minute.

Updating

git pull and re-run setup, or use the in-app updater (About → Check for updates). See Updating — the in-app path guards for a clean git checkout, then rebuilds the frontend for you.

Notes & limits

  • One engine at a time. Switching engines unloads the previous one, so a single instance serves one engine at a time. Requests already serialize onto one GPU worker, so a shared instance queues rather than oversubscribes.
  • Not multi-tenant. There are no per-user accounts or quotas — everyone behind the proxy shares the same voices, cache, and GPU.
  • Keep the data dirs. backend/voices, backend/uploads, backend/cache, and backend/models hold your voices and downloaded weights — back them up or mount them on a persistent volume if your VPS reprovisions.

On this page