xAI's Grok Voice Think Fast 2.0 Hits Sub-Second Response, Beating GPT-Realtime

SpaceXAI's new voice model hits #2 on the Speech-to-Speech Index with 0.70s latency, beating GPT-Realtime at half the price

·
·
xAI's Grok Voice Think Fast 2.0 Hits Sub-Second Response, Beating GPT-Realtime
  • Grok Voice Think Fast 2.0 debuts at #2 on the Artificial Analysis Speech-to-Speech Index with an 82.9% composite score.
  • Takes #1 on Tau Voice agentic benchmark at 56.5%, beating Qwen Audio 3.0 Realtime Plus (54.6%) and GPT-Realtime-2.1 High (45.7%).
  • Fastest model in the top 5 at 0.70s time-to-first-audio, vs. 4.02s for index leader Qwen Audio 3.0 Realtime Plus.
  • Full Duplex Bench score jumped from 77.8% (v1.0) to 95.1% — the biggest driver of the overall quality improvement.
  • Priced at $4.80/hr of input audio — more than GPT-Realtime-2 High ($4.14) but ~2.2x cheaper than GPT-Realtime-2.1 High ($10.75).
  • Available now via the xAI API, with OpenAI Realtime API compatibility for easy migration.

SpaceXAI just dropped Grok Voice Think Fast 2.0, the successor to its April 2026 flagship voice model, and the numbers are hard to ignore. On the Artificial Analysis Speech-to-Speech Index, the High variant debuts at #2 with an 82.9% composite score , up 7.3 percentage points from its predecessor , and takes the #1 spot on Tau Voice, the agentic benchmark that actually matters for production deployments.

The latency gap nobody else has closed

The headline number is 0.70 seconds time-to-first-audio. That makes Grok Voice Think Fast 2.0 the only model in the top five of the Speech-to-Speech Index under 1 second . For context: GPT-Realtime-2 High comes in at 1.14s, GPT-Realtime-2.1 High at 1.21s, and the overall index leader Qwen Audio 3.0 Realtime Plus at 4.02s . Sub-second response time is the threshold where a voice agent starts to feel like a real conversation rather than a walkie-talkie.

This is not just a speed story, though. The 1.0 model already had low latency , 1.25s , but its conversational dynamics score was a weak point. On the Full Duplex Bench subset, the model scores 95.1%, up from 77.8% for its predecessor , which was the largest single driver of the index gain. Full Duplex Bench (FDB) measures the things that make or break a real phone call: knowing when to stop talking, how to handle a user who cuts in mid-sentence, and whether the model correctly ignores filler words like "yeah" or "uh-huh" without treating them as turn-taking signals.

Where it leads , and where it doesn't

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves