Cactus Compute's Whistle Squeezes 7-Language Speech Recognition Into 16.9 MB
Cactus Compute's Whistle squeezes multilingual speech-to-text into a 16.9 MB CPU file that reportedly beats Whisper base with 9x less size and 6x speed.
- Whistle is a 16.9 MB on-device speech-to-text model from Cactus Compute, Apache 2.0 licensed.
- Supports English, German, French, Spanish, Italian, Dutch, Polish; up to 30 seconds of 16 kHz mono per pass.
- Claims it beats Whisper base on WER with roughly 9x smaller file and 6x speed on CPU.
- Shares the Needle C++ engine, so one binary turns audio into tool calls in a single call.
- Ships word timestamps, keyword biasing via Aho-Corasick, and speech embeddings from the encoder.
- Runs on 17 platforms including iOS, Android, WebAssembly, RISC-V and microcontrollers;
pip install cactus-needle.
Whistle puts seven-language speech recognition in a 16.9 MB file
Cactus Compute has released Whistle, an Apache 2.0 speech-recognition model designed to run entirely on a CPU. The 16.9 MB model transcribes English, German, French, Spanish, Italian, Dutch, and Polish, with automatic language detection and support for audio windows up to 30 seconds.
Because Whistle uses the same C++ engine and .cact container as Cactus Compute’s Needle language model, an application can add speech recognition without embedding another inference runtime. Cactus also reports lower word error rates than Whisper base on most of its listed tests, a model file roughly one-ninth the size, and six times the decoding speed in its Apple M4 Pro test. Those results are vendor-reported and use mixed runtimes and numerical precision.
Three outputs from one model
| Capability | Details |
|---|---|
| Transcription | Up to 30 seconds of 16 kHz mono audio per pass, with automatic or fixed language selection. |
| Word timestamps | Start time, end time, and probability for each word, derived from the decoder’s audio alignment. |
| Speech embeddings | One encoder vector every 80 milliseconds for matching, classification, or retrieval without generating a transcript. |
| Model format | A single 16.9 MB .cact file using 2-bit to 4-bit Cactus Quants. |
| Runtime | The CPU-focused C++ engine shared with Needle. |
Whistle also supports keyword biasing for names, places, brands, and domain-specific terms. During five-beam decoding, an Aho-Corasick matcher tracks requested keyword sequences and raises their scores. Candidate transcripts are ranked with length-normalized log probability over a vocabulary of 8,192 text pieces plus seven language tokens.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.