Text to speech
Generate speech with the active engine, a cloned profile or a designed voice.
POST /v1/audio/speech takes the OpenAI speech shape and returns audio bytes.
tts-1 and tts-1-hd use the engine selected in VoiceStudio; an engine ID such
as omnivoice, voxcpm2, cosyvoice or kittentts targets that engine.
curl "$VOICESTUDIO_BASE_URL/audio/speech" \
-H "Authorization: Bearer $VOICESTUDIO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"voice": "default",
"input": "Every voice. Every language. On your machine.",
"response_format": "wav",
"speed": 1.0
}' \
--output speech.wavFields
| Field | Meaning |
|---|---|
input | Text to speak, up to 4,096 characters |
model | tts-1, tts-1-hd, or an engine ID |
voice | default, an OpenAI alias, an engine preset, or a profile ID |
response_format | mp3, opus, aac, flac, wav or raw pcm |
speed | 0.25 to 4.0 |
Engine controls
The request also takes engine-specific fields:
{
"model": "omnivoice",
"input": "Say this softly, then pause.",
"voice": "your-profile-id",
"language": "en",
"instruct": "quiet, intimate, unhurried",
"seed": 42,
"num_step": 32,
"guidance_scale": 2.0
}Support varies by engine. GET /v1/audio/voices lists installed engines and what
each can do: cloning, emotion, device and routing.
The response is binary
Write it to a file or stream it; don't parse it as JSON. Check Content-Type:
if an encoder is unavailable the server can fall back to WAV.
Long text
One request takes up to 4,096 characters. For books, podcasts or resumable jobs, use the audiobook and long-form routes in the API reference.