POST
Continue an audio clip in the same voice — passes a short reference clip plus the new text and returns audio that sounds like the original speaker speaking the new content. Lower-friction than cloning + synthesizing in two calls when the goal is a single contiguous-feeling clip.

Authorization

string
required
Bearer token. Bearer API_key.

Request Body

string
required
Source audio URL providing voice + style. 5–30 seconds is the sweet spot.
string
required
New text to render in the source voice.
string
Output audio format. Options: wav, mp3. Default: wav.

When to use what

Tips

  • Reference length matters: longer reference clips capture more of the speaker’s prosody. 15+ seconds substantially improves long-form fidelity over 5-second references.
  • Prosody fidelity: voice extend preserves the cadence and emotional register of the reference — useful for podcast-style continuations where consistency matters more than literal word-for-word voice match.