Text-to-Voice Mode
Single-voice narration from one textarea.
Text-to-Voice mode is a single textarea with one active voice and live character, word, and duration counts. Pick the active voice in the library, type, and synthesize — ideal for narration, articles, and one-off clips.
Workflow
- Select the active voice in the library on the left.
- Type or paste your text. The counter shows characters, words, and an estimated duration as you go.
- Generate, then Play the result or Download the WAV.
The take is cached, so replaying it is instant. Editing the text, switching the voice, or changing an engine setting produces a fresh generation.
Voice modes & style
The controls adapt to the active engine:
- OmniVoice / VoxCPM — a Clone / Design / Auto toggle. Design takes a free-text attribute prompt (OmniVoice offers clickable chips like female, low pitch, british accent); Auto invents a voice; Clone uses the library voice.
- Qwen3-TTS — an always-available free-text style prompt (e.g. cheerful, slightly faster) plus an advanced panel (temperature, top_p, top_k, repetition penalty, seed).
See Engines Overview for the full per-engine breakdown.
Length limit & scripts
A request is capped at 5,000 characters (across the whole script); longer
text returns an error, so break large scripts into multiple lines or use
Podcast mode. Raise the cap with MAX_TEXT_CHARS — see
Configuration.
Right-to-left scripts (اردو, हिन्दी and more) are detected and laid out correctly as you type. Load a ready-made script from the Samples menu to try any engine instantly — see Samples & Export.