Cerebras Beats Human Speed, Booking Restaurants 19x Faster Than Rivals
Cerebras ran Qwen3 27B at 1,500 tokens per second to book a dinner reservation in 22 seconds, 19x faster than Grok, Meta, and Claude assistants.
- Cerebras built an AI assistant that books a dinner reservation in 22 seconds, 19x faster than Grok, Meta Muse, and Claude Cowork.
- Setup uses Qwen3 27B on Cerebras at ~1,500 tokens per second with Pi as the agent harness.
- Competing assistants took between 4:36 and 7:40 minutes on the same restaurant booking task.
- Pre-cached site navigation procedures cut total tool calls by more than 80% in the run.
- Independent availability checks ran in parallel, with booking logic waiting only on combined results.
- Qwen3 27B endpoint is priced at $0.99 input / $1.49 output per million tokens with 128K context.
How Cerebras cut a booking agent to 22 seconds
Cerebras says it reduced a restaurant-booking agent’s median completion time to 22 seconds by combining faster inference, concurrent browser actions, and a cached navigation procedure. The published benchmark shows how model latency, serial tool use, and repeated site discovery compound during routine agent tasks.
Inside the 22-second booking
Cerebras asked four assistants to reserve a table at A Mano, Il Borgo, or Doppio Zero. It compared Grok Bot, Meta Muse, and Claude Cowork with an agent built on Qwen3.5 27B, Cerebras inference hardware, and Pi as the tool and browser harness. Cerebras also timed a human completing the task in 37 seconds.
| Assistant | Reported time | Conditions disclosed |
|---|---|---|
| Cerebras agent | 22 seconds | Reported median of two successful runs with a prepared workflow |
| Meta Muse | 4 minutes 36 seconds | Successful recorded run |
| Claude Cowork | 6 minutes 25 seconds | Successful recorded run |
| Grok Bot | 7 minutes 40 seconds | Successful recorded run |
The listed competitors took between 12.5 and 20.9 times as long as the reported Cerebras result. Their average time was about 6 minutes 14 seconds, or roughly 17 times the 22-second figure. Cerebras describes the overall gain as 19x, but reproducing that aggregate requires the complete run data and calculation method.
The agent loop, cut three ways
Browser agents repeatedly send page state to a model, wait for the next action, execute a tool, and return the result. Latency accumulates on every turn. Cerebras shortened that loop through three changes:
- Faster decoding: Cerebras reports roughly 1,500 output tokens per second for the Qwen3.5 27B endpoint under favorable conditions. Its comparison places typical GPU inference APIs near 150 tokens per second. Faster responses reduce the pause before each browser action.
- Concurrent availability checks: The agent queries independent restaurant options together, then evaluates the combined results. Serial checks force the model and browser to finish one path before starting another.
- A cached procedure: The team stored instructions for navigating the booking flow as an agent skill. Reusing that procedure reduced tool calls by more than 80% because the agent no longer had to rediscover each page and control.
The cached skill covers stable steps such as opening the reservation interface, selecting a party size, and locating available times. Live tools still retrieve changing information, including availability, prices, card requirements, and confirmation status. This division keeps reusable navigation separate from data that must be verified during each run.
Connecting to the endpoint
At publication, Cerebras listed Qwen3.5 27B on its OpenAI-compatible Inference Cloud API with a 64,000-token context window on the free tier and 128,000 tokens on paid plans. The quoted prices were about $0.99 per million input tokens and $1.49 per million output tokens.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.cerebras.ai/v1",
api_key=os.environ["CEREBRAS_API_KEY"],
)
response = client.chat.completions.create(
model="qwen-3.5-27b",
messages=[
{
"role": "user",
"content": "Book a table for two at 7 p.m."
}
],
)
print(response.choices[0].message.content)This client invokes the model through the chat-completions API. Browser control, reservation tools, parallel execution, retries, and final confirmation remain the agent harness’s responsibility. Pi supplied those capabilities in the Cerebras demonstration.
An independent Hacker News report measured median throughput near 890 tokens per second and encountered intermittent HTTP 429 rate-limit responses. The roughly 1,500-token figure therefore describes favorable conditions and carries no stated throughput floor. Production tests should track time to first token, sustained decode speed, concurrency limits, retry delays, and end-to-end completion time.
What the comparison leaves open
The benchmark combines a prepared Cerebras workflow with competitors performing cold discovery on the same sites. It also reports two successful Cerebras runs for a narrow task, leaving several questions unanswered:
- How often did each assistant fail, retry, or require human intervention?
- How do the systems compare when every agent receives the same site procedure?
- Does the advantage persist across unfamiliar sites, changed page layouts, logins, payment prompts, and unavailable time slots?
- What are the median and 95th-percentile completion times under concurrent load?
- How many model turns, tool calls, browser actions, and tokens did each run consume?
A stronger evaluation would randomize restaurants and requested times, separate warm and cold runs, use identical browser and network conditions, and publish success rates alongside latency percentiles. Those measurements would reveal how much of the gain comes from inference speed, workflow preparation, parallel execution, and service capacity.
A practical latency budget for agents
Cerebras’s experiment identifies three concrete controls for recurring web workflows: cache stable procedures, run independent reads concurrently, and reduce latency on every model turn. These techniques apply most directly to high-volume tasks with predictable interfaces. Agents facing unfamiliar sites or frequent layout changes will need more discovery, stronger recovery logic, and broader testing before they approach the reported 22-second result.