VoiceStudioDocs
Local reference

Dubbing

POST
/dub/parse-subtitle-text

Parse pasted subtitle text into timed cues. Stateless — no job, no I/O.

A thin wrapper over services.srt_parser.parse_srt so the client's "paste a translation" flow reuses the exact lenient parser the .srt import path uses (BOM / CRLF / .-vs-, ms / missing indices, plus de-overlapping). Unlike /dub/import-srt/{job_id} this mutates nothing: the caller maps these cues onto the segments it already has, keeping the existing timings and text_original.

Request Body

application/json

Raw pasted subtitle text (SRT/VTT-ish) to be parsed into timed cues.

Used by the "paste translation from an external source" flow: the user pastes what ChatGPT/DeepL/a human gave them and the client needs the SAME lenient cue parsing the .srt import path uses — parsing it here keeps services.srt_parser the single source of truth instead of growing a second, subtly-different implementation in JavaScript.

text*Text

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/parse-subtitle-text" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "string"
  }'
json
null
POST
/dub/import-srt/{job_id}

Replace job["segments"] with timestamps + text parsed from an SRT file. Used as a fallback when Whisper mis-transcribes — the user can point at their own pre-synced subtitles and skip ASR entirely.

Returns the new segment list plus counts of any cues we had to skip or re-time (overlap shifts). The caller surfaces these so the user knows if the import wasn't lossless.

Path Parameters

job_id*Job Id

Request Body

multipart/form-data

file*File

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/import-srt/string" \
  -F file="string"
json
null
POST
/dub/cleanup-segments/{job_id}

Re-run merge/stitch passes on a job's existing segments to drop fragments.

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/cleanup-segments/string"
json
null
POST
/dub/abort/{job_id}

Cancel in-flight upload/transcribe subprocesses for a job.

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/abort/string"
json
null
GET
/dub/history

Response Body

application/json

bash
curl -X GET "https://example.com/dub/history"
json
null
DELETE
/dub/history

Delete persisted dub rows and their on-disk dirs (scoped to known IDs).

Response Body

application/json

bash
curl -X DELETE "https://example.com/dub/history"
json
null
DELETE
/dub/history/{history_id}

Path Parameters

history_id*History Id

Response Body

application/json

application/json

bash
curl -X DELETE "https://example.com/dub/history/string"
json
null
POST
/dub/upload

Accept a media upload, write to disk, queue background prep task.

input_type is "video" (default) or "audio". Audio-only jobs (#119) skip scene detection, thumbnailing, and the final video mux — the transcribe → translate → TTS core is identical.

Returns 202 with {job_id, task_id, filename}. Client should open SSE on /tasks/stream/{task_id} to monitor extract/demucs stages and wait for the 'ready' event before starting transcription.

Request Body

multipart/form-data

video*Video
job_id?string|null
input_type?Input Type
Default"video"

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/upload" \
  -F video="string"
json
null
POST
/dub/ingest-url

Ingest a remote video URL via yt-dlp. Queues background prep task.

Returns 202 immediately with {job_id, task_id}. All work (download, audio extract, Demucs, scene detect, thumbnail) happens in the background task and progress is streamed via /tasks/stream/{task_id}.

Request Body

application/json

url*Url
job_id?string|null
fetch_subs?|
Defaultfalse
sub_langs?array<string>|null
cookie_file?string|null

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/ingest-url" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "string"
  }'
json
null
GET
/dub/transcribe-stream/{job_id}

Stream per-chunk segments via SSE, then emit diarized final pass.

Pre-flight checks (missing job, missing audio, ASR not loaded) are emitted as in-stream error events rather than HTTP errors, because EventSource on the client can't read non-2xx response bodies — a 503 there surfaces as an opaque "network error" instead of the actionable message we want.

num_speakers is an optional hint passed straight to pyannote. Left unset, pyannote auto-detects the count — but its auto-detect can collapse a multi-speaker clip to a single speaker (issue #274). When the user knows the exact count, supplying it forces pyannote to return that many speakers. On paths that can't honor the hint exactly (inline ASR turns, the silence-gap heuristic) it is never silently dropped: the heuristic cycles the requested count and a warning SSE event tells the user how far the labels can be trusted.

Path Parameters

job_id*Job Id

Query Parameters

num_speakers?integer|null
per_segment_refs?Per Segment Refs
Defaulttrue

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/transcribe-stream/string"
json
null
POST
/dub/transcribe/{job_id}

Legacy synchronous transcribe (kept for the headless CLI).

num_speakers mirrors the SSE endpoint's query param (same 1–20 clamp): an exact speaker count forwarded to pyannote, or cycled by the silence-gap heuristic when pyannote is unavailable. None → auto-detect.

Path Parameters

job_id*Job Id

Query Parameters

num_speakers?integer|null

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/transcribe/string"
json
null
POST
/dub/generate/{job_id}

Adds a dub generation job to the async batch task pool.

Path Parameters

job_id*Job Id

Request Body

application/json

segments*array<>
language?Language
Default"Auto"
language_code?Language Code
Default"und"
instruct?Instruct
Default""
num_step?Num Step
Default16
guidance_scale?Guidance Scale
Default2
speed?Speed
Default1
segment_ids?array<string>|null
regen_only?array<string>|null
preview?|
Defaultfalse
slot_fit?|
Default"time_stretch"
timing_strategy?|
Default"concise"
overflow_budget_s?|
Default0
fit_options?|null
voice_match?|
Default"per_line"

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/generate/string" \
  -H "Content-Type: application/json" \
  -d '{
    "segments": [
      {
        "start": 0,
        "end": 0,
        "text": "string"
      }
    ]
  }'
json
null
POST
/dub/preview-segment/{job_id}

Generate TTS for a single segment and return WAV bytes.

This is the fast path for interactive editing — 8 diffusion steps, no disk write, no mix. The preview IS synthetic audio leaving the app, so it carries the same invisible provenance mark as every other producer (#1169 — "no watermark" here used to be an exemption).

Path Parameters

job_id*Job Id

Request Body

application/json

segment_id?string|null
text*Text
language?Language
Default"Auto"
instruct?string|null
profile_id?string|null
speed?Speed
Default1
duration?number|null

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/preview-segment/string" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "string"
  }'
json
null
GET
/dub/tracks/{job_id}

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/tracks/string"
json
null
GET
/dub/segments-text/{job_id}

Per-segment texts for one generated track: {"texts": {segKey: text}}.

Backing store is job["segments_i18n"] (P1.2) — the authoritative per-language map every generate rebuilds. The Export preview tabs use it to hydrate segments whose in-browser translations[lang] entry is missing (tracks generated before per-language persistence, partial regens), so switching the preview language can't leave a mixed-language transcript. Empty map when the job predates segments_i18n or the track was never generated — the client keeps whatever it has.

Path Parameters

job_id*Job Id

Query Parameters

lang*Lang

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/segments-text/string?lang=string"
json
null
GET
/dub/download/{job_id}/{filename}

Path Parameters

job_id*Job Id

Query Parameters

preserve_bg?Preserve Bg

Mix background noise into dubbed tracks

Defaulttrue
default_track?Default Track
Default"original"
include_tracks?Include Tracks

Comma-separated list of tracks to include (e.g. 'original,de,es'). Empty = include all.

Default""
burn_subs?Burn Subs

Burn subtitles into the video stream (forces re-encode). Uses dual-subtitle layout when dual=1.

Defaultfalse
dual?Dual

When burn_subs=1, render translated on top of italicised original.

Defaultfalse
out_format?Out Format

Audio-only jobs (#119): output container — wav, m4a, mp3, or flac. Ignored for video jobs.

Default"m4a"

Header Parameters

X-VoiceStudio-Path-Authorization?X-Voicestudio-Path-Authorization
Default""

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/download/string/{filename}"
json
null
GET
/dub/download/{job_id}

Path Parameters

job_id*Job Id

Query Parameters

preserve_bg?Preserve Bg

Mix background noise into dubbed tracks

Defaulttrue
default_track?Default Track
Default"original"
include_tracks?Include Tracks

Comma-separated list of tracks to include (e.g. 'original,de,es'). Empty = include all.

Default""
burn_subs?Burn Subs

Burn subtitles into the video stream (forces re-encode). Uses dual-subtitle layout when dual=1.

Defaultfalse
dual?Dual

When burn_subs=1, render translated on top of italicised original.

Defaultfalse
out_format?Out Format

Audio-only jobs (#119): output container — wav, m4a, mp3, or flac. Ignored for video jobs.

Default"m4a"

Header Parameters

X-VoiceStudio-Path-Authorization?X-Voicestudio-Path-Authorization
Default""

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/download/string"
json
null
GET
/dub/media/{job_id}

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/media/string"
json
null
GET
/dub/preview-video/{job_id}

Return an inline-playable MP4 with the chosen dubbed track as sole audio.

Caches per lang+preserve_bg combination under exports/preview_{lang}_{bg}.mp4. Cache is invalidated when the underlying dubbed track mtime is newer than the cache.

Path Parameters

job_id*Job Id

Query Parameters

lang*Lang

Language code of the dubbed track to mux in

preserve_bg?Preserve Bg
Defaulttrue

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/preview-video/string?lang=string"
json
null
GET
/dub/onsets/{job_id}

Speech-onset times for the timeline editor's snap-to-onset ticks (#280).

Prefers the Demucs-isolated vocals track (clean speech energy); falls back to the mixed audio. Computed once per job and cached as onsets.json in the job directory; recomputed if the source audio is newer than the cache (e.g. re-ingest into the same job dir).

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/onsets/string"
json
null
GET
/dub/thumb/{job_id}

Serve the extracted dub video thumbnail (jpg). 404 if not generated.

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/thumb/string"
json
null
GET
/dub/audio/{job_id}

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/audio/string"
json
null
GET
/dub/preview/{job_id}/{segment_index}

Path Parameters

job_id*Job Id
segment_index*Segment Index

Query Parameters

lang?Lang

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/preview/string/0"
json
null
POST
/dub/qc/{job_id}

Re-recognize the dubbed audio and flag lines whose recognized text drifts from the target text. Opt-in, never fatal: the dub is untouched — this only annotates segments with a per-line drift score and a measured start/end, surfaced as "verify this line" markers feeding incremental re-dub. The generated text stays authoritative (design delta from pyvideotrans, which overwrites subtitles).

Path Parameters

job_id*Job Id

Query Parameters

lang?Lang
drift_threshold?Drift Threshold
Default0.5

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/qc/string"
json
null
GET
/dub/download-audio/{job_id}/{filename}

Path Parameters

job_id*Job Id

Query Parameters

lang?Lang
preserve_bg?Preserve Bg
Defaulttrue

Header Parameters

X-VoiceStudio-Path-Authorization?X-Voicestudio-Path-Authorization
Default""

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/download-audio/string/{filename}"
json
null
GET
/dub/download-audio/{job_id}

Path Parameters

job_id*Job Id

Query Parameters

lang?Lang
preserve_bg?Preserve Bg
Defaulttrue

Header Parameters

X-VoiceStudio-Path-Authorization?X-Voicestudio-Path-Authorization
Default""

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/download-audio/string"
json
null
GET
/dub/srt/{job_id}/{filename}

Path Parameters

job_id*Job Id

Query Parameters

dual?Dual
Defaultfalse
lang?Lang

Track language code. Emits that track's text (segments_i18n) when the job carries it; when that track was generated under Smart Fit or stretch_video, cue times come from the fitted timeline.

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/srt/string/{filename}"
json
null
GET
/dub/srt/{job_id}

Path Parameters

job_id*Job Id

Query Parameters

dual?Dual
Defaultfalse
lang?Lang

Track language code. Emits that track's text (segments_i18n) when the job carries it; when that track was generated under Smart Fit or stretch_video, cue times come from the fitted timeline.

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/srt/string"
json
null
GET
/dub/vtt/{job_id}/{filename}

Path Parameters

job_id*Job Id

Query Parameters

dual?Dual
Defaultfalse
lang?Lang

Track language code. Emits that track's text (segments_i18n) when the job carries it; when that track was generated under Smart Fit or stretch_video, cue times come from the fitted timeline.

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/vtt/string/{filename}"
json
null
GET
/dub/vtt/{job_id}

Path Parameters

job_id*Job Id

Query Parameters

dual?Dual
Defaultfalse
lang?Lang

Track language code. Emits that track's text (segments_i18n) when the job carries it; when that track was generated under Smart Fit or stretch_video, cue times come from the fitted timeline.

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/vtt/string"
json
null
GET
/dub/export-segments/{job_id}

Path Parameters

job_id*Job Id

Query Parameters

lang?Lang

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/export-segments/string"
json
null
GET
/dub/download-mp3/{job_id}/{filename}

Path Parameters

job_id*Job Id

Query Parameters

lang?Lang
preserve_bg?Preserve Bg
Defaulttrue
bitrate?Bitrate
Default"192k"

Header Parameters

X-VoiceStudio-Path-Authorization?X-Voicestudio-Path-Authorization
Default""

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/download-mp3/string/{filename}"
json
null
GET
/dub/download-mp3/{job_id}

Path Parameters

job_id*Job Id

Query Parameters

lang?Lang
preserve_bg?Preserve Bg
Defaulttrue
bitrate?Bitrate
Default"192k"

Header Parameters

X-VoiceStudio-Path-Authorization?X-Voicestudio-Path-Authorization
Default""

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/download-mp3/string"
json
null
GET
/dub/export-stems/{job_id}

Path Parameters

job_id*Job Id

Query Parameters

lang?Lang

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/dub/export-stems/string"
json
null
POST
/dub/translate

Request Body

application/json

segments*array<>
target_lang*Target Lang
provider?string|null
source_lang?string|null
job_id?string|null
quality?|
Default"fast"
glossary?array<>|null
dialect?string|null
auto_glossary?boolean|null
reflect?boolean|null
condense?|
Defaultfalse

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/dub/translate" \
  -H "Content-Type: application/json" \
  -d '{
    "segments": [
      {
        "id": "string",
        "text": "string"
      }
    ],
    "target_lang": "string"
  }'
json
null
GET
/glossary/{project_id}

Path Parameters

project_id*Project Id

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/glossary/string"
json
null
POST
/glossary/{project_id}

Path Parameters

project_id*Project Id

Request Body

application/json

source*Source
target*Target
note?Note
Default""

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/glossary/string" \
  -H "Content-Type: application/json" \
  -d '{
    "source": "string",
    "target": "string"
  }'
json
null
DELETE
/glossary/{project_id}

Clear every term. only_auto=true keeps user-added entries.

Path Parameters

project_id*Project Id

Query Parameters

only_auto?Only Auto
Defaultfalse

Response Body

application/json

application/json

bash
curl -X DELETE "https://example.com/glossary/string"
json
null
DELETE
/glossary/{project_id}/{term_id}

Path Parameters

project_id*Project Id
term_id*Term Id

Response Body

application/json

application/json

bash
curl -X DELETE "https://example.com/glossary/string/string"
json
null
PUT
/glossary/{project_id}/{term_id}

Path Parameters

project_id*Project Id
term_id*Term Id

Request Body

application/json

source?string|null
target?string|null
note?string|null

Response Body

application/json

application/json

bash
curl -X PUT "https://example.com/glossary/string/string" \
  -H "Content-Type: application/json" \
  -d '{}'
json
null
POST
/glossary/{project_id}/auto-extract

Ask the LLM to propose glossary entries from the project's source segments.

Writes them as auto=1 rows. Existing terms with the same (source,target) are NOT duplicated. Returns the full current glossary after the pass.

Path Parameters

project_id*Project Id

Request Body

application/json

source_lang?Source Lang
Default"en"
target_lang*Target Lang
segments?array<>

[{text: '...'}] — only text is read; other fields ignored.

max_terms?Max Terms
Default40

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/glossary/string/auto-extract" \
  -H "Content-Type: application/json" \
  -d '{
    "target_lang": "string"
  }'
json
null
POST
/tools/video-context/{job_id}

Analyse the source video's visual context for dubbing decisions.

Returns per-segment mood, brightness, and complexity cues that can be used as TTS instruct hints.

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/tools/video-context/string"
json
null
POST
/batch/enqueue

Enqueue a video for batch dubbing.

The video is saved to disk and a job is added to the queue. Returns the job ID for status polling.

Request Body

multipart/form-data

video*Video
langs?Langs
Default"es"
voice_id?string|null
preserve_bg?Preserve Bg
Defaulttrue

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/batch/enqueue" \
  -F video="string"
json
null
GET
/batch/jobs

List batch jobs, optionally filtered by status.

Query Parameters

status?string|null
limit?Limit
Default50

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/batch/jobs"
json
null
GET
/batch/jobs/{job_id}

Get the status of a specific batch job.

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/batch/jobs/string"
json
null
DELETE
/batch/jobs/{job_id}

Delete a batch job record and its video file.

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X DELETE "https://example.com/batch/jobs/string"
json
null
POST
/batch/jobs/{job_id}/cancel

Cancel a queued or running batch job.

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/batch/jobs/string/cancel"
json
null
GET
/batch/download/{job_id}/{lang}

Download a completed batch job's output video for a given language.

Path Parameters

job_id*Job Id
lang*Lang

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/batch/download/string/string"
json
null
GET
/engines/sonitranslate/status

Check SoniTranslate availability.

Response Body

application/json

bash
curl -X GET "https://example.com/engines/sonitranslate/status"
json
null
POST
/engines/sonitranslate/install

Clone and set up SoniTranslate (heavy — ~15GB with models).

Response Body

application/json

bash
curl -X POST "https://example.com/engines/sonitranslate/install"
json
null
POST
/engines/sonitranslate/start

Start the SoniTranslate Gradio server.

Response Body

application/json

bash
curl -X POST "https://example.com/engines/sonitranslate/start"
json
null
POST
/engines/sonitranslate/stop

Stop the SoniTranslate Gradio server.

Response Body

application/json

bash
curl -X POST "https://example.com/engines/sonitranslate/stop"
json
null
POST
/engines/sonitranslate/dub

Run full dubbing pipeline via SoniTranslate.

Transcribes, translates, generates TTS, and mixes audio. Returns the path to the dubbed output video.

KNOWN PROVENANCE GAP (#1169, documented — not silently ignored): the dubbed audio is synthesized and muxed entirely inside the external SoniTranslate sidecar (its own venv + gradio pipeline, Edge-TTS voices), which hands back a finished video file. VoiceStudio's tensor-stage mark_synthetic chokepoint never sees that audio; marking it would require a demux → embed → re-mux post-pass on the sidecar's output, which is a lossy re-encode of a pipeline we don't control. This opt-in engine (explicit install + start) is therefore NOT covered by the invisible AudioSeal provenance mark that every built-in synthesis path carries.

Request Body

application/json

video_authorization*Video Authorization
target_language?Target Language
Default"Spanish (es)"
source_language?Source Language
Default"Automatic detection"
tts_voice?Tts Voice
Default"es-ES-AlvaroNeural-Male"
max_speakers?Max Speakers
Default1
output_authorization?string|null

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/engines/sonitranslate/dub" \
  -H "Content-Type: application/json" \
  -d '{
    "video_authorization": "string"
  }'
json
null