MCP server
Let Cursor, Claude Code, Codex and other agents speak, clone and transcribe with your local VoiceStudio.
VoiceStudio ships an MCP server so an AI agent can synthesize speech, clone a voice, transcribe audio and list your voices, locally, in a voice you choose per agent. It is mounted on the running app, so there is nothing extra to start once VoiceStudio is open.
Connect your client
Pick a client. Where it supports an install link you can add the server in one
click; otherwise copy the config. Every config points at
http://localhost:3900/mcp/ and carries no key, because loopback is never
challenged.
.cursor/mcp.jsonAdd to CursorCursor docs{
"mcpServers": {
"voicestudio": {
"url": "http://localhost:3900/mcp/",
"headers": {
"X-OmniVoice-Client-Id": "cursor"
}
}
}
}The server is mounted on the running app at http://localhost:3900/mcp/; open VoiceStudio first.
Keep the trailing slash
Use /mcp/. It works on every VoiceStudio version; bare /mcp only on
backends that serve it directly. If you moved the port with OMNIVOICE_PORT,
use that port.
Tools
| Tool | What it does |
|---|---|
generate_speech | Text to speech, WAV by default. Uses the agent's bound voice unless a profile_id is passed. Omit language to use the voice's saved language. |
clone_voice | Reference audio (base64, or a path under the base path) to a new voice profile. Returns a profile_id. |
describe_voice | A free-text description to the voice-design attributes it maps to, plus any words it could not map. Saves nothing. |
design_voice | A description to a new design voice profile with a stable identity. Returns a profile_id. |
transcribe | Audio (base64, or a path under the base path) to text, across 646 languages. |
list_voices, list_personalities, list_languages | Enumerate what is available. |
check_health | Backend status and the active GPU device. |
One voice per agent
Each client in the configs above sends an X-OmniVoice-Client-Id header
(cursor, claude-code, codex-cli, vscode). The app uses it to bind that
agent to a voice, so two agents can speak in two different voices. Change the
value to give an agent its own identity.
Keep audio out of the conversation
An agent pays for every byte it reads, and a WAV as base64 is a lot of bytes. Two variables on the backend's environment move audio onto disk and return a path instead:
| Variable | Values | Effect |
|---|---|---|
OMNIVOICE_MCP_OUTPUT_MODE | resources (default), files, both | files returns an audio_url and, with a base path, an output_path, instead of inline audio. |
OMNIVOICE_MCP_BASE_PATH | a directory | The security boundary for file traffic. transcribe and clone_voice read only from inside it, and files mode writes only into it. With none set, path arguments are refused. |
generate_speech also accepts format="ogg" or format="opus" (Opus in Ogg)
in files or both mode; that needs ffmpeg.
Other machines, and the limits
The MCP transport is not authenticated
Do not expose it to the open internet. To reach it from another machine, turn
on Settings → Sharing → Local network and send the shown access PIN as
X-OmniVoice-Pin. Plain HTTP sends that PIN unencrypted, so use it only on a
trusted LAN; otherwise put it behind an encrypted tunnel or HTTPS, with
authentication of its own.
In Docker, or behind a hostname, the MCP SDK rejects non-localhost Host
headers by default (a DNS-rebinding guard). Set OMNIVOICE_MCP_ALLOWED_HOSTS to
a comma-separated list of the hosts your agent connects from.
For clients that only speak stdio, the app bundles a shim
(python -m backend.mcp_shim) that proxies stdio to the mounted endpoint. It is
intended for source installs.
Where to go next
- Quickstart: the same app through the local HTTP API.
- Local API reference: every route, with an in-page client.
- Model library: which models, languages and engines power the tools above.