Cohere's Open-Source Transcribe Arabic Beats Whisper in 96% of Tests

Cohere's open-source 2B Arabic ASR model tops the Open Universal Arabic ASR Leaderboard, beating Whisper v3 Large by 11 WER points with native dialect and code-switching support.

·
·
Cohere's Open-Source Transcribe Arabic Beats Whisper in 96% of TestsPRO
Read2 min
TypeModel
  • New SOTA Arabic ASR: Cohere releases Cohere Transcribe Arabic, a 2B open-source model topping the Open Universal Arabic ASR Leaderboard with 25.87 average WER.
  • Beats Whisper by 11 WER points: Outperforms Whisper v3 Large (36.86 WER) and OmniASR-LLM-7B (28.32 WER), a 7B model more than 3x its size.
  • Dialect-native: Explicitly trained on Gulf, Najdi, Egyptian, Levantine, and North African Arabic, plus Arabic-English code-switching data.
  • 8x faster than OmniASR: RTFx of 525 vs. 66 for OmniASR-LLM-7B; optimized for high-throughput production via vLLM.
  • Human evals confirm it: Native Arabic-speaking reviewers preferred it over Whisper in 95.8% of head-to-head tests.
  • Apache 2.0, free to use: Weights on Hugging Face; also accessible via Cohere API with free rate-limited tier and paid Model Vault for production.

Arabic speech recognition has long been a graveyard of good intentions. The language has over 30 dialects, hundreds of millions of speakers mix Arabic and English mid-sentence, and most existing models just flatten everything into Modern Standard Arabic (MSA) -- the formal written register that many speakers rarely use in conversation. Cohere just shipped Cohere Transcribe Arabic, a 2B-parameter open-source ASR model that takes direct aim at this problem, and it's already sitting at the top of the Open Universal Arabic ASR Leaderboard.

The gap it's closing

The model targets the specific challenges of Arabic speech: dialect variety, bilingual Arabic-English conversations, code-switching, and specialized vocabulary. These aren't edge cases -- they're the norm for Arabic speakers in professional settings. Human reviewers preferred Cohere Transcribe Arabic to Whisper in 96% of tests. That's not a marginal improvement; it's a signal that existing tools were genuinely failing Arabic speakers at scale.

Where it lands on the benchmarks

Cohere says it outscores Whisper Large V3, the standard Cohere Transcribe model, and other systems in benchmarks. The numbers back that up. On the Open Universal Arabic ASR Leaderboard, it achieves an average WER (Word Error Rate -- the percentage of words transcribed incorrectly) of 25.87, compared to 28.32 for OmniASR-LLM-7B (a 7B model, more than 3x larger) and 36.86 for Whisper v3 Large.

ModelAvg WERCommon VoiceSADACasablanca
Cohere Transcribe Arabic25.875.8237.4749.71
OmniASR-LLM-7B28.329.7541.6156.46
Cohere Transcribe (03-2026)30.678.1760.1162.71
Whisper v3 Large36.8617.8355.9671.81

The benchmark covers six test sets spanning MSA, Egyptian, Gulf, Levantine, and Maghrebi dialects, making it a genuine multi-dialect stress test rather than a single-accent evaluation. The model ranks first on four of the six composite task sets.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Comments

avatar