Dubbing
Parse pasted subtitle text into timed cues. Stateless — no job, no I/O.
A thin wrapper over services.srt_parser.parse_srt so the client's
"paste a translation" flow reuses the exact lenient parser the .srt
import path uses (BOM / CRLF / .-vs-, ms / missing indices, plus
de-overlapping). Unlike /dub/import-srt/{job_id} this mutates
nothing: the caller maps these cues onto the segments it already has,
keeping the existing timings and text_original.
Request Body
application/json
Raw pasted subtitle text (SRT/VTT-ish) to be parsed into timed cues.
Used by the "paste translation from an external source" flow: the user
pastes what ChatGPT/DeepL/a human gave them and the client needs the
SAME lenient cue parsing the .srt import path uses — parsing it here
keeps services.srt_parser the single source of truth instead of
growing a second, subtly-different implementation in JavaScript.
Response Body
application/json
application/json
curl -X POST "https://example.com/dub/parse-subtitle-text" \
-H "Content-Type: application/json" \
-d '{
"text": "string"
}'nullReplace job["segments"] with timestamps + text parsed from an SRT
file. Used as a fallback when Whisper mis-transcribes — the user can
point at their own pre-synced subtitles and skip ASR entirely.
Returns the new segment list plus counts of any cues we had to skip or re-time (overlap shifts). The caller surfaces these so the user knows if the import wasn't lossless.
Path Parameters
Request Body
multipart/form-data
Response Body
application/json
application/json
curl -X POST "https://example.com/dub/import-srt/string" \
-F file="string"nullRe-run merge/stitch passes on a job's existing segments to drop fragments.
Path Parameters
Response Body
application/json
application/json
curl -X POST "https://example.com/dub/cleanup-segments/string"nullCancel in-flight upload/transcribe subprocesses for a job.
Path Parameters
Response Body
application/json
application/json
curl -X POST "https://example.com/dub/abort/string"nullcurl -X GET "https://example.com/dub/history"nullDelete persisted dub rows and their on-disk dirs (scoped to known IDs).
Response Body
application/json
curl -X DELETE "https://example.com/dub/history"nullcurl -X DELETE "https://example.com/dub/history/string"nullAccept a media upload, write to disk, queue background prep task.
input_type is "video" (default) or "audio". Audio-only jobs (#119) skip
scene detection, thumbnailing, and the final video mux — the transcribe →
translate → TTS core is identical.
Returns 202 with {job_id, task_id, filename}. Client should open SSE on /tasks/stream/{task_id} to monitor extract/demucs stages and wait for the 'ready' event before starting transcription.
Request Body
multipart/form-data
"video"Response Body
application/json
application/json
curl -X POST "https://example.com/dub/upload" \
-F video="string"nullIngest a remote video URL via yt-dlp. Queues background prep task.
Returns 202 immediately with {job_id, task_id}. All work (download, audio extract, Demucs, scene detect, thumbnail) happens in the background task and progress is streamed via /tasks/stream/{task_id}.
Request Body
application/json
falseResponse Body
application/json
application/json
curl -X POST "https://example.com/dub/ingest-url" \
-H "Content-Type: application/json" \
-d '{
"url": "string"
}'nullStream per-chunk segments via SSE, then emit diarized final pass.
Pre-flight checks (missing job, missing audio, ASR not loaded) are emitted
as in-stream error events rather than HTTP errors, because EventSource
on the client can't read non-2xx response bodies — a 503 there surfaces
as an opaque "network error" instead of the actionable message we want.
num_speakers is an optional hint passed straight to pyannote. Left unset,
pyannote auto-detects the count — but its auto-detect can collapse a
multi-speaker clip to a single speaker (issue #274). When the user knows
the exact count, supplying it forces pyannote to return that many speakers.
On paths that can't honor the hint exactly (inline ASR turns, the
silence-gap heuristic) it is never silently dropped: the heuristic cycles
the requested count and a warning SSE event tells the user how far the
labels can be trusted.
Path Parameters
Query Parameters
trueResponse Body
application/json
application/json
curl -X GET "https://example.com/dub/transcribe-stream/string"nullLegacy synchronous transcribe (kept for the headless CLI).
num_speakers mirrors the SSE endpoint's query param (same 1–20 clamp):
an exact speaker count forwarded to pyannote, or cycled by the silence-gap
heuristic when pyannote is unavailable. None → auto-detect.
Path Parameters
Query Parameters
Response Body
application/json
application/json
curl -X POST "https://example.com/dub/transcribe/string"nullAdds a dub generation job to the async batch task pool.
Path Parameters
Request Body
application/json
"Auto""und"""1621false"time_stretch""concise"0"per_line"Response Body
application/json
application/json
curl -X POST "https://example.com/dub/generate/string" \
-H "Content-Type: application/json" \
-d '{
"segments": [
{
"start": 0,
"end": 0,
"text": "string"
}
]
}'nullGenerate TTS for a single segment and return WAV bytes.
This is the fast path for interactive editing — 8 diffusion steps, no disk write, no mix. The preview IS synthetic audio leaving the app, so it carries the same invisible provenance mark as every other producer (#1169 — "no watermark" here used to be an exemption).
Path Parameters
Request Body
application/json
"Auto"1Response Body
application/json
application/json
curl -X POST "https://example.com/dub/preview-segment/string" \
-H "Content-Type: application/json" \
-d '{
"text": "string"
}'nullcurl -X GET "https://example.com/dub/tracks/string"nullPer-segment texts for one generated track: {"texts": {segKey: text}}.
Backing store is job["segments_i18n"] (P1.2) — the authoritative
per-language map every generate rebuilds. The Export preview tabs use it
to hydrate segments whose in-browser translations[lang] entry is
missing (tracks generated before per-language persistence, partial
regens), so switching the preview language can't leave a mixed-language
transcript. Empty map when the job predates segments_i18n or the track
was never generated — the client keeps whatever it has.
Path Parameters
Query Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/segments-text/string?lang=string"nullPath Parameters
Query Parameters
Mix background noise into dubbed tracks
true"original"Comma-separated list of tracks to include (e.g. 'original,de,es'). Empty = include all.
""Burn subtitles into the video stream (forces re-encode). Uses dual-subtitle layout when dual=1.
falseWhen burn_subs=1, render translated on top of italicised original.
falseAudio-only jobs (#119): output container — wav, m4a, mp3, or flac. Ignored for video jobs.
"m4a"Header Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/download/string/{filename}"nullPath Parameters
Query Parameters
Mix background noise into dubbed tracks
true"original"Comma-separated list of tracks to include (e.g. 'original,de,es'). Empty = include all.
""Burn subtitles into the video stream (forces re-encode). Uses dual-subtitle layout when dual=1.
falseWhen burn_subs=1, render translated on top of italicised original.
falseAudio-only jobs (#119): output container — wav, m4a, mp3, or flac. Ignored for video jobs.
"m4a"Header Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/download/string"nullcurl -X GET "https://example.com/dub/media/string"nullReturn an inline-playable MP4 with the chosen dubbed track as sole audio.
Caches per lang+preserve_bg combination under exports/preview_{lang}_{bg}.mp4. Cache is invalidated when the underlying dubbed track mtime is newer than the cache.
Path Parameters
Query Parameters
Language code of the dubbed track to mux in
trueResponse Body
application/json
application/json
curl -X GET "https://example.com/dub/preview-video/string?lang=string"nullSpeech-onset times for the timeline editor's snap-to-onset ticks (#280).
Prefers the Demucs-isolated vocals track (clean speech energy); falls
back to the mixed audio. Computed once per job and cached as
onsets.json in the job directory; recomputed if the source audio is
newer than the cache (e.g. re-ingest into the same job dir).
Path Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/onsets/string"nullServe the extracted dub video thumbnail (jpg). 404 if not generated.
Path Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/thumb/string"nullcurl -X GET "https://example.com/dub/audio/string"nullPath Parameters
Query Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/preview/string/0"nullRe-recognize the dubbed audio and flag lines whose recognized text drifts from the target text. Opt-in, never fatal: the dub is untouched — this only annotates segments with a per-line drift score and a measured start/end, surfaced as "verify this line" markers feeding incremental re-dub. The generated text stays authoritative (design delta from pyvideotrans, which overwrites subtitles).
Path Parameters
Query Parameters
0.5Response Body
application/json
application/json
curl -X POST "https://example.com/dub/qc/string"nullPath Parameters
Query Parameters
trueHeader Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/download-audio/string/{filename}"nullPath Parameters
Query Parameters
trueHeader Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/download-audio/string"nullPath Parameters
Query Parameters
falseTrack language code. Emits that track's text (segments_i18n) when the job carries it; when that track was generated under Smart Fit or stretch_video, cue times come from the fitted timeline.
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/srt/string/{filename}"nullPath Parameters
Query Parameters
falseTrack language code. Emits that track's text (segments_i18n) when the job carries it; when that track was generated under Smart Fit or stretch_video, cue times come from the fitted timeline.
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/srt/string"nullPath Parameters
Query Parameters
falseTrack language code. Emits that track's text (segments_i18n) when the job carries it; when that track was generated under Smart Fit or stretch_video, cue times come from the fitted timeline.
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/vtt/string/{filename}"nullPath Parameters
Query Parameters
falseTrack language code. Emits that track's text (segments_i18n) when the job carries it; when that track was generated under Smart Fit or stretch_video, cue times come from the fitted timeline.
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/vtt/string"nullcurl -X GET "https://example.com/dub/export-segments/string"nullPath Parameters
Query Parameters
true"192k"Header Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/download-mp3/string/{filename}"nullPath Parameters
Query Parameters
true"192k"Header Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/dub/download-mp3/string"nullcurl -X GET "https://example.com/dub/export-stems/string"nullRequest Body
application/json
"fast"falseResponse Body
application/json
application/json
curl -X POST "https://example.com/dub/translate" \
-H "Content-Type: application/json" \
-d '{
"segments": [
{
"id": "string",
"text": "string"
}
],
"target_lang": "string"
}'nullcurl -X GET "https://example.com/glossary/string"nullPath Parameters
Request Body
application/json
""Response Body
application/json
application/json
curl -X POST "https://example.com/glossary/string" \
-H "Content-Type: application/json" \
-d '{
"source": "string",
"target": "string"
}'nullClear every term. only_auto=true keeps user-added entries.
Path Parameters
Query Parameters
falseResponse Body
application/json
application/json
curl -X DELETE "https://example.com/glossary/string"nullcurl -X DELETE "https://example.com/glossary/string/string"nullPath Parameters
Request Body
application/json
Response Body
application/json
application/json
curl -X PUT "https://example.com/glossary/string/string" \
-H "Content-Type: application/json" \
-d '{}'nullAsk the LLM to propose glossary entries from the project's source segments.
Writes them as auto=1 rows. Existing terms with the same (source,target)
are NOT duplicated. Returns the full current glossary after the pass.
Path Parameters
Request Body
application/json
"en"[{text: '...'}] — only text is read; other fields ignored.
40Response Body
application/json
application/json
curl -X POST "https://example.com/glossary/string/auto-extract" \
-H "Content-Type: application/json" \
-d '{
"target_lang": "string"
}'nullAnalyse the source video's visual context for dubbing decisions.
Returns per-segment mood, brightness, and complexity cues that can be used as TTS instruct hints.
Path Parameters
Response Body
application/json
application/json
curl -X POST "https://example.com/tools/video-context/string"nullEnqueue a video for batch dubbing.
The video is saved to disk and a job is added to the queue. Returns the job ID for status polling.
Request Body
multipart/form-data
"es"trueResponse Body
application/json
application/json
curl -X POST "https://example.com/batch/enqueue" \
-F video="string"nullList batch jobs, optionally filtered by status.
Query Parameters
50Response Body
application/json
application/json
curl -X GET "https://example.com/batch/jobs"nullGet the status of a specific batch job.
Path Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/batch/jobs/string"nullDelete a batch job record and its video file.
Path Parameters
Response Body
application/json
application/json
curl -X DELETE "https://example.com/batch/jobs/string"nullCancel a queued or running batch job.
Path Parameters
Response Body
application/json
application/json
curl -X POST "https://example.com/batch/jobs/string/cancel"nullDownload a completed batch job's output video for a given language.
Path Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/batch/download/string/string"nullcurl -X GET "https://example.com/engines/sonitranslate/status"nullcurl -X POST "https://example.com/engines/sonitranslate/install"nullcurl -X POST "https://example.com/engines/sonitranslate/start"nullcurl -X POST "https://example.com/engines/sonitranslate/stop"nullRun full dubbing pipeline via SoniTranslate.
Transcribes, translates, generates TTS, and mixes audio. Returns the path to the dubbed output video.
KNOWN PROVENANCE GAP (#1169, documented — not silently ignored): the dubbed audio is synthesized and muxed entirely inside the external SoniTranslate sidecar (its own venv + gradio pipeline, Edge-TTS voices), which hands back a finished video file. VoiceStudio's tensor-stage mark_synthetic chokepoint never sees that audio; marking it would require a demux → embed → re-mux post-pass on the sidecar's output, which is a lossy re-encode of a pipeline we don't control. This opt-in engine (explicit install + start) is therefore NOT covered by the invisible AudioSeal provenance mark that every built-in synthesis path carries.
Request Body
application/json
"Spanish (es)""Automatic detection""es-ES-AlvaroNeural-Male"1Response Body
application/json
application/json
curl -X POST "https://example.com/engines/sonitranslate/dub" \
-H "Content-Type: application/json" \
-d '{
"video_authorization": "string"
}'null