VoiceStudioDocs

MCP server

Let Cursor, Claude Code, Codex and other agents speak, clone and transcribe with your local VoiceStudio.

VoiceStudio ships an MCP server so an AI agent can synthesize speech, clone a voice, transcribe audio and list your voices, locally, in a voice you choose per agent. It is mounted on the running app, so there is nothing extra to start once VoiceStudio is open.

Connect your client

Pick a client. Where it supports an install link you can add the server in one click; otherwise copy the config. Every config points at http://localhost:3900/mcp/ and carries no key, because loopback is never challenged.

Add to .cursor/mcp.jsonAdd to CursorCursor docs
{
  "mcpServers": {
    "voicestudio": {
      "url": "http://localhost:3900/mcp/",
      "headers": {
        "X-OmniVoice-Client-Id": "cursor"
      }
    }
  }
}

The server is mounted on the running app at http://localhost:3900/mcp/; open VoiceStudio first.

Keep the trailing slash

Use /mcp/. It works on every VoiceStudio version; bare /mcp only on backends that serve it directly. If you moved the port with OMNIVOICE_PORT, use that port.

Tools

ToolWhat it does
generate_speechText to speech, WAV by default. Uses the agent's bound voice unless a profile_id is passed. Omit language to use the voice's saved language.
clone_voiceReference audio (base64, or a path under the base path) to a new voice profile. Returns a profile_id.
describe_voiceA free-text description to the voice-design attributes it maps to, plus any words it could not map. Saves nothing.
design_voiceA description to a new design voice profile with a stable identity. Returns a profile_id.
transcribeAudio (base64, or a path under the base path) to text, across 646 languages.
list_voices, list_personalities, list_languagesEnumerate what is available.
check_healthBackend status and the active GPU device.

One voice per agent

Each client in the configs above sends an X-OmniVoice-Client-Id header (cursor, claude-code, codex-cli, vscode). The app uses it to bind that agent to a voice, so two agents can speak in two different voices. Change the value to give an agent its own identity.

Keep audio out of the conversation

An agent pays for every byte it reads, and a WAV as base64 is a lot of bytes. Two variables on the backend's environment move audio onto disk and return a path instead:

VariableValuesEffect
OMNIVOICE_MCP_OUTPUT_MODEresources (default), files, bothfiles returns an audio_url and, with a base path, an output_path, instead of inline audio.
OMNIVOICE_MCP_BASE_PATHa directoryThe security boundary for file traffic. transcribe and clone_voice read only from inside it, and files mode writes only into it. With none set, path arguments are refused.

generate_speech also accepts format="ogg" or format="opus" (Opus in Ogg) in files or both mode; that needs ffmpeg.

Other machines, and the limits

The MCP transport is not authenticated

Do not expose it to the open internet. To reach it from another machine, turn on Settings → Sharing → Local network and send the shown access PIN as X-OmniVoice-Pin. Plain HTTP sends that PIN unencrypted, so use it only on a trusted LAN; otherwise put it behind an encrypted tunnel or HTTPS, with authentication of its own.

In Docker, or behind a hostname, the MCP SDK rejects non-localhost Host headers by default (a DNS-rebinding guard). Set OMNIVOICE_MCP_ALLOWED_HOSTS to a comma-separated list of the hosts your agent connects from.

For clients that only speak stdio, the app bundles a shim (python -m backend.mcp_shim) that proxies stdio to the mounted endpoint. It is intended for source installs.

Where to go next

On this page