Speech generation
MOSS-TTS-Nano
A 100M-parameter voice that runs in real time on a 4-core CPU, 20 languages, with cloning.
- OpenMOSSBy
- 20Languages
- 48 kHzOutput
Can do
- Voice cloningYes. YesSpeaks in a voice copied from a short reference clip.
- Voice designNo. NoBuilds a new voice from a written description.
- Emotion controlNo. NoTakes graded emotion from a clip, a vector or text.
- Preset voicesNo. NoShips fixed voices; no reference clip needed.
Languages
- Arabic
ar - Chinese
zh - Czech
cs - Danish
da - English
en - French
fr - German
de - Greek
el - Hebrew
he - Hungarian
hu - Italian
it - Japanese
ja - Korean
ko - Persian
fa - Polish
pl - Portuguese
pt - Russian
ru - Spanish
es - Swedish
sv - Turkish
tr
20 languages.
Runs on
PyTorch, isolated
isolatedPyTorchCUDACPUEngine doc
Hardware
| Engine | CUDA | ROCm | Intel Arc | Apple Silicon | Vulkan | NPU | CPU |
|---|---|---|---|---|---|---|---|
| PyTorch, isolated | Supported | Not supported | Not supported | Not supported | Not supported | Not supported | Supported |
Supported Needs a build Experimental
Licence
- Code
- Apache-2.0
VoiceStudio is AGPL-3.0 and doesn't relicense model weights. Licence guide
Related
From the VoiceStudio app at 0834c8b. 27 models listed.