Library
Every model. Every engine.
The 21 speech models and 6 transcription models VoiceStudio runs on your own machine, and exactly how each one runs: the engine, the languages, and the hardware.
OmniVoice
k2-fsa · Speech generation646 languagesNon-commercial termsVoiceStudio's default voice model: zero-shot cloning across 600+ languages at 24 kHz.
Voice cloningPyTorchllama.cpp / ggmlCUDAROCmApple SiliconCPUVulkan · buildCosyVoice 3
FunAudioLLM · Speech generation · Opt-in10 languages5.4 GBMultilingual zero-shot cloning with instructed speech, installed in its own Python environment.
Voice cloningPyTorchCUDACPUKittenTTS
KittenML · Speech generationEnglish0.08 GBA tiny English-only voice with 8 presets that runs in real time on any CPU.
Preset voicesONNX RuntimeCPUKokoro 82M
hexgrad · Speech generation8 languagesA small multilingual voice, and the default model of the MLX-Audio adapter.
Preset voicesMLXApple SiliconCSM 1B
Sesame · Speech generationEnglishThe one MLX-Audio model that accepts a reference clip, so the one that clones.
Voice cloningMLXApple SiliconQwen3-TTS VoiceDesign
Qwen · Speech generation10 languagesVoice design from a written description, 10 languages, on Apple Silicon.
Voice designMLXApple SiliconDia 1.6B
Nari Labs · Speech generationEnglishA dialogue model on Apple Silicon. English only.
Preset voicesMLXApple SiliconChatterbox
Resemble AI · Speech generationEnglishAn expressive English voice on Apple Silicon (the base model; Multilingual is a separate checkpoint).
Preset voicesMLXApple SiliconMeloTTS English
Speech generationEnglishA lightweight VITS voice for English on Apple Silicon.
Preset voicesMLXApple SiliconOuteTTS 1.0
OuteAI · Speech generation23 languagesA language-model-based voice covering 23 languages on Apple Silicon.
Preset voicesMLXApple SiliconVoxCPM2
OpenBMB · Speech generation30 languagesStudio-quality 48 kHz speech with cloning and voice design, in 30 languages.
Voice cloningVoice designPyTorchCUDAApple SiliconCPUMOSS-TTS-Nano
OpenMOSS · Speech generation20 languagesA 100M-parameter voice that runs in real time on a 4-core CPU, 20 languages, with cloning.
Voice cloningPyTorchCUDACPUMOSS-TTS v1.5
OpenMOSS · Speech generation · Opt-in31 languages16 GBAn 8B flagship: 31 languages, zero-shot cloning and duration control. Wants a 16 GB+ GPU.
Voice cloningPyTorchCUDAROCmIntel Arc · experimentalNPU · experimentalCPUGPT-SoVITS
RVC-Boss · Speech generation5 languagesConnects to your own GPT-SoVITS server for fast zero- and few-shot cloning in five languages.
Voice cloningBring your own serverRemote serverCUDACPUSherpa-ONNX TTS
k2-fsa · Speech generationLanguages varyA universal ONNX runtime for 20+ voice families (VITS, Piper, Kokoro and more). You supply one model.
Preset voicesONNX RuntimeCPUCUDAIndexTTS 2.5
Bilibili · Speech generation · Opt-in5 languagesNon-commercial termsMultilingual cloning with emotion control from a reference clip, a vector or text. Made for dubbing.
Voice cloningEmotion controlPyTorchCUDACPUSupertonic-3
Supertone · Speech generation31 languages0.4 GBNon-commercial termsA ~99M-parameter CPU-only voice with 31 languages and 7 presets at 44.1 kHz.
Preset voicesONNX RuntimeCPUdots.tts
rednote-hilab · Speech generation · Opt-in24 languages9 GBA 2B zero-shot cloning voice with 24 languages at 48 kHz. Linux and macOS only.
Voice cloningPyTorchCUDACPUConfucius4-TTS
NetEase Youdao · Speech generation · Opt-in14 languages5 GBCross-lingual zero-shot cloning in 14 languages at 22.05 kHz. CUDA recommended; CPU is slow.
Voice cloningPyTorchCUDAROCmIntel Arc · experimentalNPU · experimentalCPUPocketTTS
Kyutai · Speech generation6 languages0.1 GBA 100M-parameter CPU-only voice with cloning in six languages; the fastest CPU render.
Voice cloningPyTorchCPUBreeze-TTS-2
BreezeBlue · Speech generation · Opt-in2 languages4.73 GBNon-commercial termsA 3B model run by audio.cpp: English and Chinese cloning and voice design, with no Python.
Voice cloningVoice designaudio.cppCUDAVulkanApple SiliconCPUROCm · buildWhisper
OpenAI · Transcription90+ languagesThe transcription default, in four builds: WhisperX, Faster-Whisper, MLX Whisper and PyTorch Whisper.
Word timestampsCTranslate2MLXPyTorchCUDACPUApple SiliconParakeet TDT v3
NVIDIA · Transcription25 languages1.2 GB25 mostly European languages with auto-detect. Fast even on CPU.
Word timestampsNVIDIA NeMoPyTorchMLXCUDACPUApple SiliconMoonshine
Useful Sensors · Transcription · Opt-inEnglishEdge-optimized CPU transcription for quick notes on low-power machines. Plain text, no timestamps.
ONNX RuntimeCPUFunASR SenseVoice
Alibaba DAMO · Transcription · Opt-in50+ languagesAn all-in-one recognizer with 50+ languages, punctuation and built-in speaker diarization.
Speaker labelsPyTorchCUDACPUSherpa-ONNX Dictation
k2-fsa · Transcription90+ languages0.104 GBLive dictation on CPU with small int8 models, streaming or offline. The same on macOS, Windows and Linux.
Live dictationONNX RuntimeCPUOpenAI-compatible ASR
Transcription · Opt-inLanguages varyA client for any server that speaks /v1/audio/transcriptions, local or remote. Never active by default.
Bring your own serverRemote server
Everything on this page is read from the VoiceStudio app itself (its engine registries and engine docs at commit 0834c8b), not restated by hand. Which models install on your machine, and which engine is chosen, is decided by the app on your hardware. See Download to try it.