StepFun's StepAudio 3 ASR Tops Speech Benchmark With 1.7% Error Rate

StepFun's new speech-to-text model ties for the top spot on the Artificial Analysis WER Index, dominating conversational audio at premium pricing.

·
·
StepFun's StepAudio 3 ASR Tops Speech Benchmark With 1.7% Error Rate
Read4 min
TypeNews
  • StepFun's StepAudio 3 ASR tops the AA-WER Index at 1.7%, down from 4.7% in v2.5.
  • Leads AA-AgentTalk at 1.4% WER but trails on long-form Earnings22 calls at 2.8%.
  • Priced at $6.67 per 1,000 minutes, the most expensive of the top-5 accurate models.
  • Transcribes at 88x real time, well behind MAI-Transcribe-2's 374x but ahead of Fun-Realtime-ASR-preview.
  • Built on an LLM backbone for better handling of names, jargon, and homophones across 20+ verticals.
  • Part of a five-model StepAudio 3 family covering realtime voice, TTS, audio generation and music.

StepFun’s StepAudio 3 ASR tops non-streaming speech benchmark

StepFun’s StepAudio 3 ASR now ranks first on the AA-WER leaderboard for non-streaming transcription. It recorded a 1.7% word error rate among 61 evaluated models, down from its predecessor’s 4.7%, a three-percentage-point reduction in one release.

Developers can access the model through the hosted StepFun API docs. The API accepts complete audio files after recording, which suits batch transcription and post-call processing. StepFun also released four other StepAudio 3 models covering real-time voice interaction, speech synthesis, general audio generation, and music generation.

Conversation drives the benchmark win

The AA-WER v2 index combines three datasets with different weights. Word error rate, or WER, divides inserted, deleted, and substituted words by the reference transcript’s word count, so lower scores indicate more accurate transcription.

AA-WER v2 component results
Dataset Weight StepAudio 3 ASR Comparison
AA-AgentTalk, voice-agent conversations 50% 1.4% ElevenLabs Scribe v2: 1.5%; Alibaba Fun-Realtime-ASR-preview: 2.0%
VoxPopuli-Cleaned-AA, European Parliament speech 25% 1.2% Alibaba Fun-Realtime-ASR-preview: 1.2%
Earnings22-Cleaned-AA, long earnings calls 25% 2.8% Alibaba Fun-Realtime-ASR-preview: 1.8%

StepAudio’s strongest result comes from AA-AgentTalk, which supplies half of the final score and reflects the turn-taking patterns found in voice-agent traffic. Its 2.8% result on Earnings22 shows a weaker fit for long, terminology-heavy financial calls. Recorded customer-support conversations resemble its strongest benchmark, while Alibaba’s model has the better published result for earnings calls.

88x speed at a premium price

Artificial Analysis reports throughput of 88 times real time, meaning the model processes one hour of audio in about 41 seconds under benchmark conditions. That result is close to ElevenLabs Scribe v2 at 84x and Gemini 3.5 Transcribe at 91x. MAI-Transcribe-2 reaches 374x, while Alibaba’s Fun-Realtime-ASR-preview reaches 20x.

Reported throughput and pricing
Model Throughput Price per 1,000 minutes
StepAudio 3 ASR 88x $6.67
ElevenLabs Scribe v2 84x $3.67
Gemini 3.5 Transcribe 91x Not provided
MAI-Transcribe-2 374x $1.67
Fun-Realtime-ASR-preview 20x Not provided

StepAudio 3 ASR costs $0.40 per audio hour, equivalent to $6.67 per 1,000 minutes. That makes it the most expensive of the leaderboard’s five most accurate models and about four times the listed price of MAI-Transcribe-2. StepFun lists the older StepAudio 2.5 ASR at $0.3667 per 1,000 minutes, giving cost-sensitive batch workloads a cheaper option with a higher 4.7% WER.

StepFun adds language-model context

StepFun describes StepAudio 3 ASR as a large-language-model-based recognizer that combines acoustic modeling with contextual reasoning. Acoustic modeling maps sound to candidate words, while language context can resolve ambiguous pronunciations and homophones. The approach targets difficult entities such as personal names, place names, medications, and technical terms.

StepFun’s model documentation lists support for Chinese, English, dialects, mixed Chinese-English speech, long recordings, and more than 20 specialized domains. Those domains include sports, pharmaceuticals, chemicals, automotive, legal, finance, and software development. The three AA-WER datasets do not independently verify that broader language and industry coverage.

Benchmark gaps to test in production

Non-streaming recognition supports enterprise transcription pipelines, documentation systems, post-call analytics, subtitles, and media archives. The AA-WER result provides a useful accuracy signal, but it does not cover several requirements that commonly determine production fit:

  • Workload quality: Test representative accents, background noise, overlapping speakers, abbreviations, names, and domain terminology.
  • Transcript features: Verify speaker diarization, timestamps, punctuation, number formatting, and confidence scores.
  • Operational limits: Check supported formats, maximum file duration, concurrency, rate limits, retries, and end-to-end latency.
  • Data governance: Confirm retention policies, regional processing, training-data use, security controls, and compliance terms.

The published results make StepAudio 3 ASR a strong candidate for recorded conversational audio when teams prioritize low WER and accept its higher API cost. Alibaba posts the better Earnings22 result, MAI offers much higher throughput, and StepAudio 2.5 remains StepFun’s lower-cost batch option. Production selection still depends on workload-specific accuracy, latency, features, and governance requirements.

Trending
  • No trending articles

Comments

avatar

Next Reads