Speech generation
Supertonic-3
A ~99M-parameter CPU-only voice with 31 languages and 7 presets at 44.1 kHz.
- SupertoneBy
- 31Languages
- 44.1 kHzOutput
- ~0.4 GBDownload
Can do
- Voice cloningNo. NoSpeaks in a voice copied from a short reference clip.
- Voice designNo. NoBuilds a new voice from a written description.
- Emotion controlNo. NoTakes graded emotion from a clip, a vector or text.
- Preset voicesYes. YesShips fixed voices; no reference clip needed.
Languages
Runs on
ONNX, isolated
isolatedONNX RuntimeCPUEngine doc
Hardware
| Engine | CUDA | ROCm | Intel Arc | Apple Silicon | Vulkan | NPU | CPU |
|---|---|---|---|---|---|---|---|
| ONNX, isolated | Not supported | Not supported | Not supported | Not supported | Not supported | Not supported | Supported |
Supported Needs a build Experimental
Licence
- Code
- MIT (inference SDK)
- Weights
- OpenRAIL-M, with use restrictions you accept in the app
VoiceStudio is AGPL-3.0 and doesn't relicense model weights. Licence guide
Related
From the VoiceStudio app at 0834c8b. 27 models listed.