Transcription
Transcribe audio through the OpenAI-compatible endpoint.
POST /v1/audio/transcriptions takes multipart form data. whisper-1 maps to
the ASR engine selected in VoiceStudio.
curl "$VOICESTUDIO_BASE_URL/audio/transcriptions" \
-H "Authorization: Bearer $VOICESTUDIO_API_KEY" \
-F "file=@interview.wav" \
-F "model=whisper-1" \
-F "response_format=verbose_json"Output formats
| Format | Use it for |
|---|---|
json | A simple { "text": "…" } |
text | Plain-text pipelines |
verbose_json | Language, duration, segments and timestamps |
srt | Subtitle editors and video tools |
vtt | Web video and browser captions |
Optional fields: language (ISO 639-1), prompt (vocabulary, or the style of the
previous segment) and temperature (not every engine uses it).
Models never download silently
With only a TTS model installed you get 409, naming the missing model and
what to do. Install an ASR model in VoiceStudio and retry the request.
For live microphone streaming, use the dictation WebSocket in the API reference.