Speech generation

Sherpa-ONNX TTS

A universal ONNX runtime for 20+ voice families (VITS, Piper, Kokoro and more). You supply one model.

  • k2-fsaBy
  • VariesLanguages
  • 22.05 kHzOutput

Can do

  • Voice cloningNo. NoSpeaks in a voice copied from a short reference clip.
  • Voice designNo. NoBuilds a new voice from a written description.
  • Emotion controlNo. NoTakes graded emotion from a clip, a vector or text.
  • Preset voicesYes. YesShips fixed voices; no reference clip needed.

Languages

Varieslanguages
Full list on the model card

Runs on

Hardware

EngineCUDAROCmIntel ArcApple SiliconVulkanNPUCPU
ONNX RuntimeSupportedNot supportedNot supportedNot supportedNot supportedNot supportedSupported

Supported Needs a build Experimental

Licence

Weights
Depends on the model you download

VoiceStudio is AGPL-3.0 and doesn't relicense model weights. Licence guide

Related

From the VoiceStudio app at 0834c8b. 27 models listed.

Quick answers