Speech generation · Opt-in
CosyVoice 3
Multilingual zero-shot cloning with instructed speech, installed in its own Python environment.
- FunAudioLLMBy
- 10Languages
- 24 kHzOutput
- ~5.4 GBDownload
Can do
- Voice cloningYes. YesSpeaks in a voice copied from a short reference clip.
- Voice designNo. NoBuilds a new voice from a written description.
- Emotion controlNo. NoTakes graded emotion from a clip, a vector or text.
- Preset voicesNo. NoShips fixed voices; no reference clip needed.
Languages
- Cantonese
yue - Chinese
zh - English
en - French
fr - German
de - Italian
it - Japanese
ja - Korean
ko - Russian
ru - Spanish
es
Runs on
PyTorch, isolated
isolatedPyTorchCUDACPUOne-click install into its own Python 3.10 environment.
Engine doc
Hardware
| Engine | CUDA | ROCm | Intel Arc | Apple Silicon | Vulkan | NPU | CPU |
|---|---|---|---|---|---|---|---|
| PyTorch, isolated | Supported | Not supported | Not supported | Not supported | Not supported | Not supported | Supported |
Supported Needs a build Experimental
Licence
- Code
- Apache-2.0
VoiceStudio is AGPL-3.0 and doesn't relicense model weights. Licence guide
Related
From the VoiceStudio app at 0834c8b. 27 models listed.