VoiceStudioDocs

Text to speech

Generate speech with the active engine, a cloned profile or a designed voice.

POST /v1/audio/speech takes the OpenAI speech shape and returns audio bytes. tts-1 and tts-1-hd use the engine selected in VoiceStudio; an engine ID such as omnivoice, voxcpm2, cosyvoice or kittentts targets that engine.

curl "$VOICESTUDIO_BASE_URL/audio/speech" \
  -H "Authorization: Bearer $VOICESTUDIO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "voice": "default",
    "input": "Every voice. Every language. On your machine.",
    "response_format": "wav",
    "speed": 1.0
  }' \
  --output speech.wav

Fields

FieldMeaning
inputText to speak, up to 4,096 characters
modeltts-1, tts-1-hd, or an engine ID
voicedefault, an OpenAI alias, an engine preset, or a profile ID
response_formatmp3, opus, aac, flac, wav or raw pcm
speed0.25 to 4.0

Engine controls

The request also takes engine-specific fields:

{
  "model": "omnivoice",
  "input": "Say this softly, then pause.",
  "voice": "your-profile-id",
  "language": "en",
  "instruct": "quiet, intimate, unhurried",
  "seed": 42,
  "num_step": 32,
  "guidance_scale": 2.0
}

Support varies by engine. GET /v1/audio/voices lists installed engines and what each can do: cloning, emotion, device and routing.

The response is binary

Write it to a file or stream it; don't parse it as JSON. Check Content-Type: if an encoder is unavailable the server can fall back to WAV.

Long text

One request takes up to 4,096 characters. For books, podcasts or resumable jobs, use the audiobook and long-form routes in the API reference.

On this page