VoiceStudioDocs

Transcription

Transcribe audio through the OpenAI-compatible endpoint.

POST /v1/audio/transcriptions takes multipart form data. whisper-1 maps to the ASR engine selected in VoiceStudio.

curl "$VOICESTUDIO_BASE_URL/audio/transcriptions" \
  -H "Authorization: Bearer $VOICESTUDIO_API_KEY" \
  -F "file=@interview.wav" \
  -F "model=whisper-1" \
  -F "response_format=verbose_json"

Output formats

FormatUse it for
jsonA simple { "text": "…" }
textPlain-text pipelines
verbose_jsonLanguage, duration, segments and timestamps
srtSubtitle editors and video tools
vttWeb video and browser captions

Optional fields: language (ISO 639-1), prompt (vocabulary, or the style of the previous segment) and temperature (not every engine uses it).

Models never download silently

With only a TTS model installed you get 409, naming the missing model and what to do. Install an ASR model in VoiceStudio and retry the request.

For live microphone streaming, use the dictation WebSocket in the API reference.

On this page