SYSTRAN's Faster Whisper Medium Transcribes 99 Languages 2.4x Faster
A CTranslate2 conversion of OpenAI's Whisper medium delivers up to 4x faster transcription with lower memory, and it's crossed 830K downloads.
- SYSTRAN's faster-whisper-medium is a CTranslate2 conversion of OpenAI's Whisper medium checkpoint.
- Delivers up to 4x faster transcription than reference Whisper with identical FP16 accuracy.
- Supports FP16 on GPU and INT8 on CPU or GPU, cutting VRAM by 30-40%.
- Covers 99 languages, MIT licensed, ships with built-in Silero VAD filtering.
- Powers WhisperX and 26+ Hugging Face Spaces; over 830K downloads to date.
- Runs via the faster-whisper Python package with a minimal transcribe API.
Faster Whisper Medium speeds up multilingual transcription
SYSTRAN converted OpenAI’s Whisper medium checkpoint to CTranslate2, giving developers a faster, lower-memory way to run the same 769-million-parameter speech model. The converted checkpoint targets self-hosted transcription systems where latency, throughput, and GPU memory determine serving cost.
Whisper medium supports 99 languages, multilingual transcription, and speech translation into English. Its size places it between the cheaper small checkpoints and the more demanding large family, making it a practical candidate for subtitles, meetings, voice applications, and dataset labeling.
CTranslate2 reshapes inference
The repository stores the learned parameters from OpenAI’s upstream model in the format used by CTranslate2. That inference engine uses optimized kernels, decoder caching, kernel fusion, and low-precision computation to reduce runtime overhead.
FP16 execution preserves the original model parameters and generally provides comparable accuracy. Exact transcripts can still vary with decoding settings, numerical precision, hardware, and runtime versions. INT8 modes reduce memory use further, with a possible accuracy cost that should be measured on representative audio.
The benchmark gap, in numbers
The faster-whisper project’s published benchmark transcribes 13 minutes of audio with Whisper large-v2 on an NVIDIA GeForce RTX 3070 Ti 8 GB. Although the test uses a larger checkpoint, it shows how the CTranslate2 runtime behaves under comparable decoding settings.
| Implementation | Precision | Batch size | Time | Maximum VRAM |
|---|---|---|---|---|
| openai/whisper | FP16 | 1 | 2m 23s | 4,708 MB |
| whisper.cpp with Flash Attention | FP16 | 1 | 1m 05s | 4,127 MB |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.