Voice Studio

Text-to-Voice Mode

Single-voice narration from one textarea.

Text-to-Voice mode is a single textarea with one active voice and live character, word, and duration counts. Pick the active voice in the library, type, and synthesize — ideal for narration, articles, and one-off clips.

Workflow

  1. Select the active voice in the library on the left.
  2. Type or paste your text. The counter shows characters, words, and an estimated duration as you go.
  3. Generate, then Play the result or Download the WAV.

The take is cached, so replaying it is instant. Editing the text, switching the voice, or changing an engine setting produces a fresh generation.

Voice modes & style

The controls adapt to the active engine:

  • OmniVoice / VoxCPM — a Clone / Design / Auto toggle. Design takes a free-text attribute prompt (OmniVoice offers clickable chips like female, low pitch, british accent); Auto invents a voice; Clone uses the library voice.
  • Qwen3-TTS — an always-available free-text style prompt (e.g. cheerful, slightly faster) plus an advanced panel (temperature, top_p, top_k, repetition penalty, seed).

See Engines Overview for the full per-engine breakdown.

Length limit & scripts

A request is capped at 5,000 characters (across the whole script); longer text returns an error, so break large scripts into multiple lines or use Podcast mode. Raise the cap with MAX_TEXT_CHARS — see Configuration.

Right-to-left scripts (اردو, हिन्दी and more) are detected and laid out correctly as you type. Load a ready-made script from the Samples menu to try any engine instantly — see Samples & Export.

On this page