Configuration
Settings, environment variables, and per-engine defaults.
Configuration is sourced, in order of precedence, from CLI flags, environment
variables, and a backend/.env file. Variable names are the setting name in
upper case (case-insensitive). Create backend/.env for persistent local config:
# backend/.env
DEVICE=cuda
DEFAULT_ENGINE=kokoro
MAX_TEXT_CHARS=8000
CACHE_MAX_ENTRIES=1000Common settings
| Variable | Default | Description |
|---|---|---|
DEVICE | auto | auto / cuda / cpu / mps |
DEFAULT_ENGINE | vibevoice | Engine activated on first run |
HOST | 0.0.0.0 | Bind host |
PORT | 8880 | Bind port |
LOG_LEVEL | info | debug / info / warning / error |
Directories
| Variable | Default | Description |
|---|---|---|
MODELS_DIR | backend/models | Hugging Face weight cache |
CACHE_DIR | backend/cache | Synthesized-audio cache |
VOICES_DIR | backend/voices | Built-in reference voices |
UPLOADS_DIR | backend/uploads | Uploaded voices |
Cache & limits
| Variable | Default | Description |
|---|---|---|
CACHE_ENABLED | true | Turn the synthesis cache on/off |
CACHE_MAX_ENTRIES | 500 | LRU bound before eviction |
MAX_TEXT_CHARS | 5000 | Max characters per request (see below) |
SYNTH_TIMEOUT_S | 600 | Per-request synthesis timeout |
Speech-to-text (Whisper)
| Variable | Default | Description |
|---|---|---|
ASR_MODEL_ID | openai/whisper-large-v3-turbo | ASR model |
ASR_MAX_UPLOAD_MB | 100 | Max upload size |
ASR_MAX_DURATION_SEC | 3600 | Max clip length |
Per-engine defaults
Each engine's model id and defaults are overridable too — e.g.
KOKORO_LANG_CODE, CHATTERBOX_MODEL_ID, CHATTERBOX_DEFAULT_CFG_WEIGHT,
CHATTERBOX_DEFAULT_EXAGGERATION, OMNIVOICE_MODEL_ID, VOXCPM_MODEL_ID,
VOXCPM_INFERENCE_TIMESTEPS, QWEN_MODEL_ID, and DEFAULT_CFG_SCALE.
Text length
MAX_TEXT_CHARS caps a single synthesis request at 5,000 characters by
default, counted across the whole script. It's enforced for every engine, and
over-limit requests return a 400:
MAX_TEXT_CHARS=10000This is an input guardrail, not a model context window — multi-line scripts are split per line before synthesis, so it bounds total input rather than what any single model processes at once. Very large values can risk out-of-memory errors on smaller GPUs.
Many of these map to CLI flags too.