VoiceStudioDocs
Local reference

Audio

POST
/clean-audio

Accept a raw mic recording, run demucs vocal isolation, return clean WAV.

Request Body

multipart/form-data

audio*Audio

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/clean-audio" \
  -F audio="string"
json
null
POST
/generate

Request Body

multipart/form-data

text*Text
language?string|null
ref_audio?string|null
ref_text?string|null
instruct?string|null
duration?number|null
num_step?Num Step
Default16
guidance_scale?Guidance Scale
Default2
speed?Speed
Default1
t_shift?number|null
denoise?Denoise
Defaulttrue
postprocess_output?Postprocess Output
Defaulttrue
layer_penalty_factor?number|null
position_temperature?number|null
class_temperature?number|null
profile_id?string|null
seed?integer|null
effect_preset?Effect Preset
Default"broadcast"
engine?string|null
max_chunk_chars?Max Chunk Chars
Range0 <= value
Default800
crossfade_ms?Crossfade Ms
Range0 <= value <= 1000
Default50
pronounce?Pronounce
Defaulttrue
stream?Stream
Defaultfalse
hosted?Hosted
Defaultfalse

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/generate" \
  -F text="string"
json
null
POST
/transcribe

Transcribe an audio file to text.

Args: audio: The audio file to transcribe. language: Optional language hint (not currently used; auto-detected). model: Whisper model size (legacy; ignored in dual-mode architecture). mode: 'fast' (default) uses MLX Turbo for speed; 'accurate' uses WhisperX with forced alignment for word-level timing. refine: Opt-in local-LLM cleanup of the final text (disfluencies, self-corrections, punctuation) — same pipeline the live dictation socket uses. Off by default so MCP/CLI callers don't pay LLM latency unless they ask; honours the user's Settings → Dictation-refinement config and silently passes through when no LLM backend is configured. The raw text is always returned; refined_text is added only when the LLM actually changed something.

Returns: { "text": "full transcription", "refined_text": "cleaned text", # only when refine=true changed it "segments": [ {"start": 0.0, "end": 1.5, "text": "..."}, ... ], "language": "en", "duration_s": 4.2, "transcription_time_s": 0.8, "engine": "mlx-whisper" }

Request Body

multipart/form-data

audio*Audio
language?string|null
model?string|null
mode?string|null
refine?string|null

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/transcribe" \
  -F audio="string"
json
null
POST
/v1/audio/speech

Generate audio from text. Compatible with OpenAI's POST /v1/audio/speech.

Request Body

application/json

POST /v1/audio/speech — mirrors OpenAI's CreateSpeechRequest.

model?Model

TTS model to use. Maps to VoiceStudio engine IDs: 'omnivoice', 'voxcpm2', 'cosyvoice', 'mlx-audio', 'kittentts', 'moss-tts-nano'. Also accepts 'tts-1' and 'tts-1-hd' as aliases for the active engine.

Default"omnivoice"
input*Input

The text to synthesize. Max 4096 characters.

Lengthlength <= 4096
voice?Voice

Voice to use. For VoiceStudio: pass a voice profile ID, 'default', or a KittenTTS preset name. OpenAI voice names (alloy, echo, fable, onyx, nova, shimmer) are accepted but mapped to defaults.

Default"default"
response_format?Response Format

Audio output format.

Default"mp3"

Value in

  • "mp3"
  • "opus"
  • "aac"
  • "flac"
  • "wav"
  • "pcm"
speed?Speed

Speed of the generated audio (0.25 to 4.0).

Range0.25 <= value <= 4
Default1
language?|

Language code (ISO 639-1)

description?|

Voice description for voice design (VoxCPM2 only). E.g. 'young female, warm tone, slight British accent'.

instruct?|

Style instruction for the TTS engine.

duration?|

VoiceStudio extension: target output duration in seconds.

seed?|

VoiceStudio extension: deterministic sampling seed.

denoise?Denoise

VoiceStudio extension: prepend denoise control when supported.

Defaulttrue
preprocess_prompt?Preprocess Prompt

VoiceStudio extension: trim/preprocess reference prompt when supported.

Defaulttrue
chunk_duration?|

OmniVoice GGUF extension: long-form internal chunk duration.

chunk_threshold?|

OmniVoice GGUF extension: long-form internal chunk threshold.

num_step?|

VoiceStudio extension: iterative unmasking steps (app default 16; 32 = the model's documented quality preset).

guidance_scale?|

VoiceStudio extension: classifier-free guidance scale (app default 2.0).

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/v1/audio/speech" \
  -H "Content-Type: application/json" \
  -d '{
    "input": "string"
  }'
json
null
POST
/v1/audio/transcriptions

Transcribe audio to text. Compatible with OpenAI's POST /v1/audio/transcriptions.

Request Body

multipart/form-data

file*File

Audio file to transcribe

model?Model

ASR model. Accepts 'whisper-1' (maps to active engine), or an VoiceStudio engine ID: whisperx, faster-whisper, mlx-whisper, pytorch-whisper.

Default"whisper-1"
language?|

Language of the input audio (ISO 639-1). Optional.

prompt?|

Optional text to guide the model's style or continue a previous segment.

response_format?Response Format

Output format: json, text, verbose_json, srt, vtt.

Default"json"
temperature?|

Sampling temperature (0–1). Not used by all backends.

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/v1/audio/transcriptions" \
  -F file="string"
json
null
GET
/v1/audio/voices

List available voices. VoiceStudio extension to the OpenAI API.

Response Body

application/json

bash
curl -X GET "https://example.com/v1/audio/voices"
json
null