Audio
Accept a raw mic recording, run demucs vocal isolation, return clean WAV.
Request Body
multipart/form-data
Response Body
application/json
application/json
curl -X POST "https://example.com/clean-audio" \
-F audio="string"nullRequest Body
multipart/form-data
1621truetrue"broadcast"0 <= value8000 <= value <= 100050truefalsefalseResponse Body
application/json
application/json
curl -X POST "https://example.com/generate" \
-F text="string"nullTranscribe an audio file to text.
Args:
audio: The audio file to transcribe.
language: Optional language hint (not currently used; auto-detected).
model: Whisper model size (legacy; ignored in dual-mode architecture).
mode: 'fast' (default) uses MLX Turbo for speed; 'accurate' uses
WhisperX with forced alignment for word-level timing.
refine: Opt-in local-LLM cleanup of the final text (disfluencies,
self-corrections, punctuation) — same pipeline the live
dictation socket uses. Off by default so MCP/CLI callers don't
pay LLM latency unless they ask; honours the user's
Settings → Dictation-refinement config and silently passes
through when no LLM backend is configured. The raw text
is always returned; refined_text is added only when the
LLM actually changed something.
Returns: { "text": "full transcription", "refined_text": "cleaned text", # only when refine=true changed it "segments": [ {"start": 0.0, "end": 1.5, "text": "..."}, ... ], "language": "en", "duration_s": 4.2, "transcription_time_s": 0.8, "engine": "mlx-whisper" }
Request Body
multipart/form-data
Response Body
application/json
application/json
curl -X POST "https://example.com/transcribe" \
-F audio="string"nullGenerate audio from text. Compatible with OpenAI's POST /v1/audio/speech.
Request Body
application/json
POST /v1/audio/speech — mirrors OpenAI's CreateSpeechRequest.
TTS model to use. Maps to VoiceStudio engine IDs: 'omnivoice', 'voxcpm2', 'cosyvoice', 'mlx-audio', 'kittentts', 'moss-tts-nano'. Also accepts 'tts-1' and 'tts-1-hd' as aliases for the active engine.
"omnivoice"The text to synthesize. Max 4096 characters.
length <= 4096Voice to use. For VoiceStudio: pass a voice profile ID, 'default', or a KittenTTS preset name. OpenAI voice names (alloy, echo, fable, onyx, nova, shimmer) are accepted but mapped to defaults.
"default"Audio output format.
"mp3"Value in
- "mp3"
- "opus"
- "aac"
- "flac"
- "wav"
- "pcm"
Speed of the generated audio (0.25 to 4.0).
0.25 <= value <= 41Language code (ISO 639-1)
Voice description for voice design (VoxCPM2 only). E.g. 'young female, warm tone, slight British accent'.
Style instruction for the TTS engine.
VoiceStudio extension: target output duration in seconds.
VoiceStudio extension: deterministic sampling seed.
VoiceStudio extension: prepend denoise control when supported.
trueVoiceStudio extension: trim/preprocess reference prompt when supported.
trueOmniVoice GGUF extension: long-form internal chunk duration.
OmniVoice GGUF extension: long-form internal chunk threshold.
VoiceStudio extension: iterative unmasking steps (app default 16; 32 = the model's documented quality preset).
VoiceStudio extension: classifier-free guidance scale (app default 2.0).
Response Body
application/json
application/json
curl -X POST "https://example.com/v1/audio/speech" \
-H "Content-Type: application/json" \
-d '{
"input": "string"
}'nullTranscribe audio to text. Compatible with OpenAI's POST /v1/audio/transcriptions.
Request Body
multipart/form-data
Audio file to transcribe
ASR model. Accepts 'whisper-1' (maps to active engine), or an VoiceStudio engine ID: whisperx, faster-whisper, mlx-whisper, pytorch-whisper.
"whisper-1"Language of the input audio (ISO 639-1). Optional.
Optional text to guide the model's style or continue a previous segment.
Output format: json, text, verbose_json, srt, vtt.
"json"Sampling temperature (0–1). Not used by all backends.
Response Body
application/json
application/json
curl -X POST "https://example.com/v1/audio/transcriptions" \
-F file="string"nullcurl -X GET "https://example.com/v1/audio/voices"null