VoiceStudioDocs
Local reference

Long form

POST
/stories/encode

Transcode an uploaded WAV to MP3/M4B/OGG and return the encoded bytes.

Provenance note (#1169): this endpoint is a pure TRANSCODER, not a synthesis producer — it never calls a TTS engine, so it must not call mark_synthetic (the upload may be arbitrary user audio, and marking human speech as synthetic would be wrong). Audio the Stories Editor stitched from VoiceStudio generations is already marked at its producing route, and the AudioSeal mark survives the lossy encode here.

Request Body

multipart/form-data

file*File
format?Format
Default"mp3"
bitrate?Bitrate
Default"192k"

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/stories/encode" \
  -F file="string"
json
null
POST
/audiobook/plan

Parse a script into a chapter/span plan (pure preview, no synthesis).

Request Body

application/json

text*Text
default_voice?string|null

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/audiobook/plan" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "string"
  }'
json
{}
POST
/audiobook/import

Import a .txt/.md/.epub/.pdf into a chapter-delimited script.

EPUB is parsed in spine order (stdlib only, local); PDF text is extracted with pypdf (pure-Python) then chapterized; plain text gets # headings inserted ahead of obvious chapter-title lines. Returns the script text (for the editor) + the resulting chapter count.

Request Body

multipart/form-data

file*File

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/audiobook/import" \
  -F file="string"
json
{}
POST
/audiobook/cover

Upload a cover image; returns a server-side path to pass back as cover_path in the synth request. Validated here (jpg/png + size cap) and again at render time.

Request Body

multipart/form-data

cover*Cover

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/audiobook/cover" \
  -F cover="string"
json
{}
POST
/audiobook/preview

Render a single chapter so the user can audition it before the full run.

Reuses the same content-addressed cache as the job, so a preview warms the cache (the later full render reuses it) and a re-preview is instant.

Request Body

application/json

num_step?|null
guidance_scale?|null
position_temperature?|null
class_temperature?|null
postprocess_output?boolean|null
seed?|null
emo_vector?array<number>|null
emo_text?|null
emo_alpha?|null
vary_repeats?Vary Repeats
Defaultfalse
text*Text
chapter_index?Chapter Index
Default0
default_voice?string|null
language?string|null
lexicon?|null
voice_map?|null

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/audiobook/preview" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "string"
  }'
json
{}
POST
/audiobook

Synthesize a chapterized audiobook from a script, streaming SSE progress.

Request Body

application/json

num_step?|null
guidance_scale?|null
position_temperature?|null
class_temperature?|null
postprocess_output?boolean|null
seed?|null
emo_vector?array<number>|null
emo_text?|null
emo_alpha?|null
vary_repeats?Vary Repeats
Defaultfalse
text*Text
default_voice?string|null
language?string|null
bitrate?Bitrate
Default"128k"
format?Format
Default"m4b"
loudness?string|null
cover_path?string|null
metadata?|null
lexicon?|null
voice_map?|null

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/audiobook" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "string"
  }'
json
null
POST
/longform/render

Render a pre-built chapter/span plan (the Stories Editor's compiled cast+lines) through the shared chapterized renderer — same resume, loudness, cover, metadata, and output formats as the Audiobook job.

Request Body

application/json

num_step?|null
guidance_scale?|null
position_temperature?|null
class_temperature?|null
postprocess_output?boolean|null
seed?|null
emo_vector?array<number>|null
emo_text?|null
emo_alpha?|null
vary_repeats?Vary Repeats
Defaultfalse
chapters?array<>
Default[]
default_voice?string|null
language?string|null
bitrate?Bitrate
Default"128k"
format?Format
Default"m4b"
loudness?string|null
cover_path?string|null
metadata?|null
lexicon?|null
voice_map?|null

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/longform/render" \
  -H "Content-Type: application/json" \
  -d '{}'
json
null
GET
/audiobook/jobs

List interrupted longform renders that can be resumed — a work dir that still holds a resume manifest (a job left mid-render by a crash/quit). The ids come from scanning the filesystem, so the UI can offer one-click resume.

Response Body

application/json

bash
curl -X GET "https://example.com/audiobook/jobs"
json
{}
POST
/audiobook/resume/{job_id}

Resume an interrupted longform render from its persisted manifest. The already-rendered chapters are content-addressed in the shared cache, so they return instantly — only the unrendered chapters synthesize again. Streams the same SSE event shape as the original render, under the original job_id.

Path Parameters

job_id*Job Id

Response Body

application/json

application/json

bash
curl -X POST "https://example.com/audiobook/resume/string"
json
null
GET
/longform/jobs

Finished Audiobook + Story renders, newest-first, ready to re-download.

Each item's output is served at /audio/<output>. Never 500s — on any backend hiccup it returns an empty list rather than an error.

job_store is bound at module import (top of file), NOT re-imported here at call time. A call-time from core import job_store re-resolves through sys.modules on every request — and several test suites purge and re-import the whole core/services namespace under a temporary OMNIVOICE_DATA_DIR (the isolated_db pattern, tests/smoke/test_boot_smoke.py, …) without restoring the old module tree. After one of those ran, the call-time import resolved a stale module world whose DB_PATH pointed at a different SQLite file than the one the test seeded through its collection-time bindings, so the library came back without the seeded jobs (the order-dependent test_route_handler_returns_jobs_envelope full-suite flake). The module-level binding keeps this route in the same world as whoever imported this module. Regression: tests/test_longform_jobs.py::test_route_survives_leaked_module_world_purge.

Query Parameters

limit?Limit
Range1 <= value <= 500
Default50

Response Body

application/json

application/json

bash
curl -X GET "https://example.com/longform/jobs"
json
{}