Speech generation

MOSS-TTS-Nano

A 100M-parameter voice that runs in real time on a 4-core CPU, 20 languages, with cloning.

  • OpenMOSSBy
  • 20Languages
  • 48 kHzOutput

Can do

  • Voice cloningYes. YesSpeaks in a voice copied from a short reference clip.
  • Voice designNo. NoBuilds a new voice from a written description.
  • Emotion controlNo. NoTakes graded emotion from a clip, a vector or text.
  • Preset voicesNo. NoShips fixed voices; no reference clip needed.

Languages

20languages
  • Arabicar
  • Chinesezh
  • Czechcs
  • Danishda
  • Englishen
  • Frenchfr
  • Germande
  • Greekel
  • Hebrewhe
  • Hungarianhu
  • Italianit
  • Japaneseja
  • Koreanko
  • Persianfa
  • Polishpl
  • Portuguesept
  • Russianru
  • Spanishes
  • Swedishsv
  • Turkishtr

20 languages.

Runs on

Hardware

EngineCUDAROCmIntel ArcApple SiliconVulkanNPUCPU
PyTorch, isolatedSupportedNot supportedNot supportedNot supportedNot supportedNot supportedSupported

Supported Needs a build Experimental

Licence

Code
Apache-2.0

VoiceStudio is AGPL-3.0 and doesn't relicense model weights. Licence guide

Related

From the VoiceStudio app at 0834c8b. 27 models listed.

Quick answers