Inception Labs' Mercury Voice Hits 320ms Response for Live AI Calls

Inception Labs' new diffusion-based model returns its first spoken token in 320ms median, finally fitting a reasoning model inside the 500ms voice budget.

·
·
Read5 min
TypeNews
TopicAudio · Llms
  • Inception Labs launched Mercury Voice, a diffusion LLM built for real-time agentic voice applications.
  • Hits 320 ms median time to first answer token, the only model under the 500 ms conversational budget.
  • Beats GPT-6 Luna, Gemma 4 31B, GLM-5.3-Flash on τ³-bench, IFBench, BFCL v4 composite scores.
  • 128K context, three reasoning-effort settings, up to 50K output tokens, OpenAI-compatible API.
  • Priced at $0.20/$0.75 per million tokens at launch, roughly $0.009 per minute of conversation.
  • Enterprise-only access via sales@inceptionlabs.ai; drops into LiveKit, Pipecat, Vapi, Retell stacks.

Mercury Voice targets subsecond reasoning for voice agents

Inception Labs has introduced Mercury Voice, a diffusion language model designed for latency-sensitive voice agents. The company reports a 320 millisecond median time to first answer token while retaining reasoning, tool calling, and long-prompt support. If those results hold across production workloads, developers could reduce the pauses that often force voice systems to use smaller models or more complex orchestration.

Parallel generation targets the pause

Conventional language models generate one token at a time, feeding each token back into the model before predicting the next. Diffusion language models refine multiple token positions in parallel, using GPU capacity more efficiently and increasing generation speed. Inception previously reported throughput above 1,000 tokens per second for Mercury 2; Mercury Voice applies the same architecture to conversational agents.

Voice applications place unusual pressure on startup latency because every turn includes endpoint detection, speech recognition, model processing, and speech synthesis. Mercury Voice handles the language-model stage. Its reported latency therefore represents one component of the full delay between a caller finishing a sentence and hearing the response.

  • Architecture: Diffusion language model
  • Primary workload: Tool-using voice agents
  • Context window: 128,000 tokens
  • Maximum output: 50,000 tokens
  • Reasoning controls: Low, medium, and high

A 320 ms median, with caveats

Inception measures time to first answer token (TTFAT), which includes internal reasoning before the model emits the first user-facing token. On the company’s set of customer-service prompts, Mercury Voice at low reasoning effort recorded a median TTFAT of about 320 ms and a p95 of 750 ms. A p95 of 750 ms means 95% of measured requests produced their first answer token within that time.

Inception Labs benchmark comparing time to first answer token across voice-agent models
Inception reports that Mercury Voice was the only tested model with a median TTFAT below its 500 ms conversational target. The results are vendor-reported.

TTFAT does not capture the caller’s complete mouth-to-ear delay. Network transit, turn detection, transcription, text-to-speech startup, and application code add latency around the model. Tail performance also matters: a system with a low median and a multi-second p95 will still produce frequent slow turns during longer calls.

Reasoning quality remains vendor-reported

Inception says Mercury Voice outperformed GPT-6 Luna, Gemma 4 31B, GLM-5.3-Flash, Gemini 3.5 Flash-Lite, and Qwen3.5-397B on a composite score covering τ³-bench Telecom, Retail, and Airline, IFBench, and BFCL v4. The company also says the model ran about twice as fast as the next-fastest model in that comparison.

Benchmark What it evaluates
τ³-bench Multi-turn customer-service agents that call tools and maintain task state
IFBench Compliance with detailed instructions and constraints
BFCL v4 Function selection, argument generation, and tool use

These tasks resemble common voice-agent workflows more closely than general knowledge tests, particularly when a call involves account lookup, order changes, or booking actions. Production evaluation still needs representative prompts, real tool schemas, long conversations, failure recovery, and latency measurements across each reasoning setting.

Pricing favors high-volume calls

Mercury Voice lists at $0.40 per million input tokens and $1.50 per million output tokens. A launch discount cuts those rates by half. Inception estimates language-model usage at about $0.009 per conversation minute and roughly one-fifth the cost of GPT-4.1 for its tested voice-agent profile.

Usage Launch price List price
1 million input tokens $0.20 $0.40
1 million output tokens $0.75 $1.50

Actual per-minute cost will vary with prompt length, conversation structure, tool output, and response length. Telephony, speech recognition, synthesis, observability, and infrastructure can add separate charges.

Integration follows familiar APIs

Access is currently limited to enterprise customers through the Inception API. Its OpenAI-compatible endpoint should let teams connect Mercury Voice through existing clients and place it in common voice orchestration stacks, subject to compatibility testing for streaming, tool schemas, retries, and usage reporting.

  • LiveKit
  • Pipecat
  • Vapi
  • Retell
  • Custom pipelines using OpenAI-compatible clients

Early deployments offer partial evidence

Inception cites several design partners in its announcement. Audivi AI uses Mercury Voice for automated drive-through ordering, including modifications and upsells. Altur deploys it for collections calls at financial institutions and reports that model inference is no longer the main bottleneck in its pipeline. OpenCall reports median model-response latency of about 170 ms on its production phone-agent workload.

Those figures come from different workloads and measurement setups, so the 170 ms production result cannot be compared directly with Inception’s 320 ms benchmark median. Independent tests would clarify performance under concurrent traffic, long prompts, repeated tool calls, and provider rate limits.

The architectural bet

Mercury Voice tests whether parallel token refinement can deliver reasoning and tool use within the latency budget of a live conversation. Google, Together AI, and other labs are also exploring diffusion-based language models, with voice agents, code completion, and iterative agent loops offering clear opportunities for faster generation.

For teams that currently route calls between separate planning and response models, one low-latency model could simplify orchestration and remove network hops. The practical decision will depend on end-to-end latency, task completion, tool-call accuracy, operating cost, and reliability under production load rather than the model benchmark alone.

Trending
  • No trending articles

Comments

avatar

Next Reads