Transcription

Whisper

The transcription default, in four builds: WhisperX, Faster-Whisper, MLX Whisper and PyTorch Whisper.

  • OpenAIBy
  • 90+Languages

Can do

  • Word timestampsYes. YesTiming for each word, good enough to dub against.
  • Speaker labelsNo. NoSays who is speaking.
  • Live dictationNo. NoStreams text as you speak.

Languages

90+languages
Full list on the model card

Runs on

  • WhisperX

    CTranslate2CUDACPU

    faster-whisper plus wav2vec2 forced alignment, for word timing good enough to dub.

    Engine doc
  • Faster-Whisper

    CTranslate2CUDACPU

    Whisper on CTranslate2 without forced alignment.

    Engine doc
  • MLX Whisper

    MLXApple SiliconCPU

    Runs on the Apple Silicon GPU.

    Engine doc
  • PyTorch Whisper

    PyTorchCUDAApple SiliconCPU

    The Whisper build that actually uses AMD ROCm, and the rescue when CTranslate2 fails.

    Engine doc

Hardware

EngineCUDAROCmIntel ArcApple SiliconVulkanNPUCPU
WhisperXSupportedNot supportedNot supportedNot supportedNot supportedNot supportedSupported
Faster-WhisperSupportedNot supportedNot supportedNot supportedNot supportedNot supportedSupported
MLX WhisperNot supportedNot supportedNot supportedSupportedNot supportedNot supportedSupported
PyTorch WhisperSupportedNot supportedNot supportedSupportedNot supportedNot supportedSupported

Supported Needs a build Experimental

Licence

Terms
Not stated in the app's docs. Check the model card before commercial use.

VoiceStudio is AGPL-3.0 and doesn't relicense model weights. Licence guide

Related

From the VoiceStudio app at 0834c8b. 27 models listed.

Quick answers