System and settings
curl -X GET "https://example.com/health"nullSettings page system info — model, tokens, data dir, timeout.
This endpoint MUST never throw — it's called on every Settings page load and a 500 here blocks the entire UI from rendering system details.
Response Body
application/json
curl -X GET "https://example.com/system/info"{
"app_version": "",
"data_dir": "string",
"outputs_dir": "string",
"crash_log_path": "string",
"idle_timeout_seconds": 0,
"model_checkpoint": "unknown",
"asr_model": "unknown",
"translate_provider": "unknown",
"has_hf_token": false,
"fast_download": {},
"device": "cpu",
"python": "",
"platform": "",
"arch": "",
"os_version": "",
"cpu_model": "",
"cpu_count": 0,
"ram_total_gb": 0,
"gpu_name": "",
"vram_total_gb": 0,
"disk_free_gb": 0,
"error": "string",
"ffmpeg_ok": false,
"ffmpeg_path": "",
"proxy_url": "",
"share_enabled": false,
"share_port": 0,
"lan_addresses": [],
"pin_required": false,
"backend_port": 3900,
"share_port_base": 3901,
"ui_port": 3901
}Tail the rolling runtime log — everything Python logged since last rotation.
Back-stop: if the rolling log doesn't exist yet (fresh install, disk error), fall back to the crash log so the UI always has something to show.
Query Parameters
200Response Body
application/json
application/json
curl -X GET "https://example.com/system/logs"nullTail the Tauri plugin log (or backend stdout redirect, whichever exists).
Query Parameters
200Response Body
application/json
application/json
curl -X GET "https://example.com/system/logs/tauri"nullServer-Sent Events stream of new log lines.
The client opens an EventSource connection and receives new lines as they are appended to the log file. This replaces the polling pattern used by the LogsFooter component.
Usage (frontend)::
const es = new EventSource('/system/logs/stream?source=backend');
es.onmessage = (e) => { const lines = JSON.parse(e.data); ... };Query Parameters
'backend' or 'tauri'
"backend"Poll interval in seconds
0.3 <= value <= 101Response Body
application/json
application/json
curl -X GET "https://example.com/system/logs/stream"nullTruncate the rolling runtime log and the crash log (what the Backend tab reads).
Response Body
application/json
curl -X POST "https://example.com/system/logs/clear"nullTruncate whichever Tauri-side log files we know about. OS-level rotation may recreate them.
Response Body
application/json
curl -X POST "https://example.com/system/logs/tauri/clear"nullcurl -X GET "https://example.com/sysinfo"{
"cpu": 0,
"ram": 0,
"total_ram": 0,
"vram": 0,
"gpu_active": false
}Aggressively release RAM/VRAM by clearing caches and running GC.
When unload_model=true, the TTS model is fully unloaded and will be re-loaded lazily on the next generation request.
Query Parameters
falseResponse Body
application/json
application/json
curl -X POST "https://example.com/system/flush-memory"nullReturn actionable notifications for the UI notification panel.
Each notification has:
- id: unique key (for dismiss tracking)
- level: "info" | "warn" | "error"
- title: short heading
- message: longer description
- action: optional {"label": str, "type": "navigate|link|api", "target": str}
Response Body
application/json
curl -X GET "https://example.com/system/notifications"nullNewest unclean-shutdown record from the previous backend run (#1164)
— the deployment-agnostic twin of the desktop shell's crash marker
(get_last_backend_crash), for browser/dev/Docker frontends that have no
shell to ask. Version-gated like the shell's markers: records from a
different release than the running build are ignored (kept on disk).
Response Body
application/json
curl -X GET "https://example.com/system/last-run-crash"nullMark the newest unclean-shutdown record as seen. Watermark semantics (like the shell's ack): the record itself is retained so bug reports can still attach the evidence; a NEWER death re-arms the notice.
Response Body
application/json
curl -X POST "https://example.com/system/last-run-crash/ack"nullMark the current crash log as seen — dismisses the 'crash-last-session' notification until the log changes again.
Response Body
application/json
curl -X POST "https://example.com/system/crash/ack"nullSet an environment variable at runtime, persisted across restarts.
Persistent keys (proxy, FFMPEG_PATH, translation provider keys, …) are
saved to prefs.json so they survive backend restarts (restored at
startup in main.py). HF_TOKEN is persisted via
huggingface_hub.login() (and cleared via logout()). Other keys
are set on os.environ for the running process.
The loopback-origin gate that previously lived inline here is now applied
at the router level via dependencies=[Depends(require_admin)] on
router — see the top of this file. Every route on this router is
gated, including this one. The 403 body and behavior are unchanged.
Request Body
application/json
Response Body
application/json
application/json
curl -X POST "https://example.com/system/set-env" \
-H "Content-Type: application/json" \
-d '{}'nullcurl -X GET "https://example.com/system/asr-backends"nullReturn the 3-source HF token cascade state for the Settings UI (Wave 2 React panel consumes this). Never returns the raw token — only a masked preview, whoami username, and per-source validity.
Response Body
application/json
curl -X GET "https://example.com/system/hf-token/state"nullRecent unhandled backend errors, newest first — structured, deduped (count per fingerprint), classified (error_class), pre-scrubbed. The bug-report pipeline reads this to auto-attach the most recent backend failure; Settings → Logs can render it as a triage view.
Query Parameters
1 <= value <= 5020Response Body
application/json
application/json
curl -X GET "https://example.com/system/errors/recent"nullBuild the drag-onto-a-GitHub-issue zip (core.diagnostic_bundle): self-check report, recent error journal, scrubbed log tails. Returns the local path so the UI can reveal it in the file manager. The path itself is NOT scrubbed — this response never leaves the machine; the zip's contents are scrubbed because the zip does.
Query Parameters
Include the hub reachability probe
falseResponse Body
application/json
application/json
curl -X POST "https://example.com/system/diagnostic-bundle"nullRun the self-check suite (core.diagnose) and return the structured report.
The hub probe can block up to ~5s (and deep=true far longer), so the
whole run goes through a threadpool; pass network=false for an
instant offline report. Output is pre-scrubbed (core.scrub) — safe to
paste into a GitHub issue.
Query Parameters
Include the HuggingFace hub reachability probe
trueAlso load the active engine and synthesize a short utterance (may cold-load the model — minutes on first run)
falseResponse Body
application/json
application/json
curl -X GET "https://example.com/system/diagnose"nullReport whether the running .app bundle has the macOS quarantine xattr.
On non-macOS platforms or dev runs (not inside a .app bundle), always
returns {"quarantined": false, "error_class": null}. The React
ErrorBoundary polls this endpoint on first load and renders the docs
deeplink when error_class is set (Plan 01-02 wired the deeplink).
Response Body
application/json
curl -X GET "https://example.com/system/quarantine-status"nullcurl -X GET "https://example.com/system/network/state"nullcurl -X POST "https://example.com/system/network/enable"nullcurl -X POST "https://example.com/system/network/disable"nullcurl -X GET "https://example.com/system/tailscale/status"nullcurl -X POST "https://example.com/system/tailscale/enable"nullcurl -X POST "https://example.com/system/tailscale/disable"nullReturn a curated model preset for the caller's device + architecture.
Data-driven from the curated_on field in models.yaml — adding or
retargeting a curated pick is a catalog edit, not a code change. Only the
TTS model is required; the ASR picks here are the optional "best for your
system" set the wizard and Settings surface for on-demand install.
Response Body
application/json
curl -X GET "https://example.com/setup/recommendations"nullcurl -X GET "https://example.com/setup/status"{
"models_ready": true,
"missing": [
{
"repo_id": "string",
"label": "string"
}
],
"hf_cache_dir": "string",
"disk_free_gb": 0,
"min_free_gb": 10,
"enough_disk": true
}curl -X GET "https://example.com/setup/preflight"{
"ok": true,
"has_warnings": false,
"checks": [
{
"id": "string",
"label": "string",
"status": "string",
"detail": "",
"fix": "string"
}
],
"device": {
"os": "string",
"arch": "string",
"gpu_vendor": "none",
"gpu_backend": "cpu",
"gpu_available": false,
"gpu_driver": "string",
"gpu_device_name": "string",
"gpu_family": "cpu",
"vram_gb": 0,
"ram_gb": 0,
"disk_free_gb": 0
},
"gpu_routing": {
"engine": "string",
"effective_device": "string",
"routing_status": "string",
"routing_reason": "string",
"host_family": "cpu",
"vram_gb": 0
},
"media_tools": {}
}Trigger a model load in the background so the first dub doesn't pay the cold-start tax.
Response Body
application/json
curl -X POST "https://example.com/setup/warmup"nullSSE: forward every HuggingFace download tqdm update as a JSON event.
Query Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/setup/download-stream"nullPersist a new HF token to the encrypted settings store + the HF canonical file (via huggingface_hub.login). Returns the updated cascade state.
Request Body
application/json
HuggingFace access token
1 <= lengthResponse Body
application/json
application/json
curl -X POST "https://example.com/api/settings/hf-token" \
-H "Content-Type: application/json" \
-d '{
"token": "string"
}'nullClear the App-source token. Optionally also call huggingface_hub.logout to clear the canonical HF file. Returns the updated cascade state.
Query Parameters
falseResponse Body
application/json
application/json
curl -X DELETE "https://example.com/api/settings/hf-token"null3-source HF token cascade state for the Settings UI.
fresh=1 drops the resolver's whoami validation cache first so the
response re-runs whoami for every source — this is what the panel's
"Test now" button sends. Plain GETs (panel mounts) keep the 300s cache
so repeat Settings visits don't hammer the HF API.
Query Parameters
falseResponse Body
application/json
application/json
curl -X GET "https://example.com/api/settings/hf-token/state"nullReturn the current torch.compile-disabled toggle + the runtime platform. UI uses the platform to render the toggle disabled (with an explainer) on non-Windows hosts, since the OOM is Windows-specific (issue #65).
Response Body
application/json
curl -X GET "https://example.com/api/settings/perf/torch-compile-disabled"nullPersist the toggle. Honoured by services.engine_env.build_engine_env()
which injects TORCH_COMPILE_DISABLE=1 on Windows when enabled.
Request Body
application/json
True to set TORCH_COMPILE_DISABLE=1 on engine subprocesses
Response Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/perf/torch-compile-disabled" \
-H "Content-Type: application/json" \
-d '{
"enabled": true
}'nullcurl -X GET "https://example.com/api/settings/compute-device"nullPersist the compute-device pick. Applied by the capability probe at
the next backend start (host caps are immutable per process — same
restart contract as the rest of the Performance tab). OMNIVOICE_DEVICE
always wins over this pick; the UI shows the pin instead of pretending.
Request Body
application/json
auto | cuda | rocm | xpu | mps | cpu
Response Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/compute-device" \
-H "Content-Type: application/json" \
-d '{
"value": "string"
}'nullcurl -X GET "https://example.com/api/settings/history-retention"nullPersist the retention cap. Enforced after every generation: the oldest unstarred takes over the cap are pruned (rows + their audio files); starred takes are never pruned. 0 disables pruning entirely.
Request Body
application/json
Max takes kept before the oldest UNstarred ones (rows + WAVs) are pruned; 0 = unlimited
0 <= value <= 100000Response Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/history-retention" \
-H "Content-Type: application/json" \
-d '{
"cap": 0
}'nullcurl -X GET "https://example.com/api/settings/dictation-refinement"nullRequest Body
application/json
Response Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/dictation-refinement" \
-H "Content-Type: application/json" \
-d '{}'nullcurl -X GET "https://example.com/api/settings/llm-endpoint"nullPersist base URL / model / API key for the OpenAI-compatible endpoint.
Reuses the env-var persistence path (prefs.json, restored at startup): base_url -> TRANSLATE_BASE_URL, model -> TRANSLATE_MODEL, api_key -> TRANSLATE_API_KEY. A None field is left unchanged; an empty string clears it. Ollama ignores the key; vLLM / LM Studio require it.
Request Body
application/json
Response Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/llm-endpoint" \
-H "Content-Type: application/json" \
-d '{}'nullAll providers with resolved base_url/model + whether a key is configured.
Never returns key material — only has_key/key_from_env booleans.
Response Body
application/json
curl -X GET "https://example.com/api/settings/llm-providers"nullSave a provider's key (encrypted) + optional base_url/model/account.
A None field is left unchanged; an empty api_key clears the stored key.
Path Parameters
Request Body
application/json
API key; '' clears it, None leaves unchanged
Cloudflare account id
falseResponse Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/llm-providers/string" \
-H "Content-Type: application/json" \
-d '{}'nullRequest Body
application/json
provider id to activate
Response Body
application/json
application/json
curl -X POST "https://example.com/api/settings/llm-providers/active" \
-H "Content-Type: application/json" \
-d '{
"provider": "string"
}'nullOne cheap round-trip against a provider to prove the key/URL work.
Temporarily activates the provider for the probe by resolving its config
directly (does not change the persisted active selection). Returns
latency_ms plus, on failure, a classified kind (config / auth /
not_found / rate_limit / network / error) so the UI shows an actionable,
localizable message instead of a raw exception string.
Path Parameters
Response Body
application/json
application/json
curl -X POST "https://example.com/api/settings/llm-providers/string/test"nullList model ids the provider's key can access (OpenAI-compat /models).
Powers the model-picker datalist in Settings → LLM Providers so users don't have to guess model names. Read-only; failures return the same classified shape as /test; capped so a huge catalog can't bloat the UI.
Path Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/api/settings/llm-providers/string/models"nullcurl -X GET "https://example.com/api/settings/llm-skills"nullToggle a skill and/or set its provider routing.
Field semantics match the providers PUT: an omitted field is left
unchanged; provider_override: ""/null clears the override.
404 for an unknown skill or an unknown provider id.
Path Parameters
Request Body
application/json
None leaves the toggle unchanged
provider id to route this skill to; '' or null clears it (skill follows the active provider). Omit to leave unchanged.
Response Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/llm-skills/string" \
-H "Content-Type: application/json" \
-d '{}'nullPersist a per-engine license-acceptance boolean.
Returns {"ok": True, "engine_id": ..., "accepted": ...} so the
caller can update its UI without a second round-trip. Validation:
engine_id must be in the in-tree allow-list ‑‑ refuses arbitrary
keys so the settings table can't be polluted via this route.
Request Body
application/json
1 <= length <= 64True to accept the license terms
Response Body
application/json
application/json
curl -X POST "https://example.com/api/settings/license" \
-H "Content-Type: application/json" \
-d '{
"engine_id": "string",
"accepted": true
}'{}Return {"engine_id": ..., "accepted": bool}.
Same allow-list as the POST handler so an unknown engine id is a
400 rather than a silent accepted=false for a non-existent
engine.
Path Parameters
Response Body
application/json
application/json
curl -X GET "https://example.com/api/settings/license/string"{}Current models directory: the persisted choice (from the durable env file — the same value main.py reads at startup), what's effective in this process, and the platform default.
Response Body
application/json
curl -X GET "https://example.com/api/settings/storage/models-dir"nullSet (or clear, with an empty path) the models download directory.
Validates the directory is writable, then writes OMNIVOICE_CACHE_DIR to the durable per-user env file so main.py applies it on the next launch. The env file is the only persisted store, so GET can never diverge from what was saved. Returns restart_required=True.
Request Body
application/json
One-shot native desktop authorization
Response Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/storage/models-dir" \
-H "Content-Type: application/json" \
-d '{
"authorization": "string"
}'nullDisk + per-category storage usage for the Settings → Storage panel.
refresh=1 bypasses the 5-minute cache and rescans. min_free_gb
reuses the setup wizard's constant so both surfaces warn at the same
threshold.
Query Parameters
falseResponse Body
application/json
application/json
curl -X GET "https://example.com/api/settings/storage"nullDelete VoiceStudio-owned temp files (Settings → Storage → Temporary files).
Removes only the omnivoice* entries in the OS temp dir — the exact
population the storage report's "temp" category counts — and invalidates
the cached report so the next scan reflects the reclaimed space. Partial
failures (files held open by a running job) are returned per entry.
Response Body
application/json
curl -X POST "https://example.com/api/settings/storage/temp/clear"nullcurl -X GET "https://example.com/api/settings/hf-mirror"nullRequest Body
application/json
HF_ENDPOINT URL; empty string clears it (official endpoint)
""'auto' switches to automatic endpoint selection (clears any explicit endpoint); 'manual' (or omitted — back-compat with older clients) pins the given url as an explicit choice.
Response Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/hf-mirror" \
-H "Content-Type: application/json" \
-d '{}'nullRe-run the endpoint race now (the Auto panel's "Test again").
Forces fresh probes and re-caches the decision. In manual mode this is a no-op (an explicit endpoint is never auto-switched) — the response simply reflects the current state.
Response Body
application/json
curl -X POST "https://example.com/api/settings/hf-mirror/test"nullcurl -X GET "https://example.com/api/settings/asr-openai-compat"nullRequest Body
application/json
'' clears it, None leaves unchanged
Response Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/asr-openai-compat" \
-H "Content-Type: application/json" \
-d '{}'nullCheap connectivity probe for the "Test connection" button.
GET {base_url}/models against the PERSISTED config — the panel saves first, then tests, same stale-config contract as /llm-providers/{id}/test. No audio leaves the machine, nothing is transcribed. Loopback-only via the router-level guard. Always 200 with a structured verdict ({ok, status, latency_ms, ...} — see services.asr_backend.probe_openai_compat_server) so the UI renders success/latency or the exact failure without a raw 500. The key is never logged or echoed back.
Response Body
application/json
curl -X POST "https://example.com/api/settings/asr-openai-compat/test"nullStructured release notes from the shipped CHANGELOG.md (newest first).
Bullets are raw markdown-lite (bold leads, code, (#NNN) refs) — the
frontend renders them safely without HTML. available: false when this
install has no changelog (never an error: the viewer just hides).
Query Parameters
1 <= value <= 505Response Body
application/json
application/json
curl -X GET "https://example.com/api/settings/changelog"nullNewest pre-migration database backup (or none yet). Feeds the "your data is backed up before every update" line in Settings → Updates.
Response Body
application/json
curl -X GET "https://example.com/api/settings/db-backup"nullcurl -X GET "https://example.com/api/settings/analytics"nullRequest Body
application/json
User's explicit choice. Default is OFF.
Response Body
application/json
application/json
curl -X PUT "https://example.com/api/settings/analytics" \
-H "Content-Type: application/json" \
-d '{
"enabled": true
}'nullcurl -X GET "https://example.com/media-tools/status"null(Re-)fetch the pinned, checksummed static ffmpeg/ffprobe build in the background. Idempotent; poll /media-tools/status for progress.
Response Body
application/json
curl -X POST "https://example.com/media-tools/acquire"nullFetch the newest yt-dlp wheel (sha256-verified against PyPI metadata) into the update-surviving overlay. Applies on next backend start.
Response Body
application/json
curl -X POST "https://example.com/media-tools/ytdlp/update"nullDrop the overlay — the app-tested, locked yt-dlp takes over on next start. Always safe (the locked install is never modified).
Response Body
application/json
curl -X POST "https://example.com/media-tools/ytdlp/restore"nullPath Parameters
Request Body
application/json
Response Body
application/json
application/json
curl -X POST "https://example.com/media-tools/string/custom-path" \
-H "Content-Type: application/json" \
-d '{
"authorization": "string"
}'nullcurl -X POST "https://example.com/media-tools/string/use-system"nullcurl -X POST "https://example.com/media-tools/string/restore"null