Speech generation · Opt-in

Confucius4-TTS

Cross-lingual zero-shot cloning in 14 languages at 22.05 kHz. CUDA recommended; CPU is slow.

  • NetEase YoudaoBy
  • 14Languages
  • 22.05 kHzOutput
  • ~5 GBDownload

Can do

  • Voice cloningYes. YesSpeaks in a voice copied from a short reference clip.
  • Voice designNo. NoBuilds a new voice from a written description.
  • Emotion controlNo. NoTakes graded emotion from a clip, a vector or text.
  • Preset voicesNo. NoShips fixed voices; no reference clip needed.

Languages

14languages
  • Chinesezh
  • Englishen
  • Frenchfr
  • Germande
  • Indonesianid
  • Italianit
  • Japaneseja
  • Koreanko
  • Malayms
  • Portuguesept
  • Russianru
  • Spanishes
  • Thaith
  • Vietnamesevi

14 languages.

Runs on

  • PyTorch, isolated

    isolated
    PyTorchCUDAROCmIntel Arc · experimentalNPU · experimentalCPU
    Engine doc

Hardware

EngineCUDAROCmIntel ArcApple SiliconVulkanNPUCPU
PyTorch, isolatedSupportedSupportedExperimentalNot supportedNot supportedExperimentalSupported

Supported Needs a build Experimental

Licence

Code
Apache-2.0
  • XPU and NPU selection is covered by mocked tests; synthesis is not certified on that hardware.

VoiceStudio is AGPL-3.0 and doesn't relicense model weights. Licence guide

Related

From the VoiceStudio app at 0834c8b. 27 models listed.

Quick answers