Speech generation

Supertonic-3

A ~99M-parameter CPU-only voice with 31 languages and 7 presets at 44.1 kHz.

  • SupertoneBy
  • 31Languages
  • 44.1 kHzOutput
  • ~0.4 GBDownload

Can do

  • Voice cloningNo. NoSpeaks in a voice copied from a short reference clip.
  • Voice designNo. NoBuilds a new voice from a written description.
  • Emotion controlNo. NoTakes graded emotion from a clip, a vector or text.
  • Preset voicesYes. YesShips fixed voices; no reference clip needed.

Languages

31languages
Full list on the model card

Runs on

Hardware

EngineCUDAROCmIntel ArcApple SiliconVulkanNPUCPU
ONNX, isolatedNot supportedNot supportedNot supportedNot supportedNot supportedNot supportedSupported

Supported Needs a build Experimental

Licence

Code
MIT (inference SDK)
Weights
OpenRAIL-M, with use restrictions you accept in the app

VoiceStudio is AGPL-3.0 and doesn't relicense model weights. Licence guide

Related

From the VoiceStudio app at 0834c8b. 27 models listed.

Quick answers