Parallel Web Systems' Turbo Cuts AI Search Costs 14x at 200ms
Parallel Web Systems launches Turbo mode for its Search API: 200ms median latency, $1 per 1,000 requests, and benchmark-leading accuracy against Exa, Tavily, and Brave.
- Parallel Search Turbo launches at $1/1,000 requests and 200ms median latency -- up to 14x cheaper than frontier model search. (Blog)
- Benchmark-leading accuracy: 51% on BrowseComp vs. Exa's 33.7%, Tavily's 19.3%, and Brave's 38.3%, all at faster latency.
- New architecture: Turbo is the first product on a re-engineered search stack spanning hardware, model training, and index design.
- Key use cases unlocked: voice AI, RL training with live web grounding, consumer chat with search-by-default, and small/local model augmentation.
- Company context: Parallel raised a $100M Series B at a $2B valuation led by Sequoia; total capital raised is $230M. (Series B announcement)
- Real-world traction: Nooks cut web search costs 70.5% switching to Parallel; Harvey, Notion, and Opendoor are also production customers.
Parallel Web Systems just launched Turbo, a new speed tier for its Search API that is gunning directly at the cost and latency assumptions that have defined AI web search. At $1 per 1,000 requests and a 200ms median latency, it is not a marginal improvement over the competition -- it is a different price point entirely.
The numbers that matter
The headline claims are specific and backed by benchmarks Parallel ran across four public datasets. On BrowseComp -- OpenAI's benchmark of 1,266 hard, multi-hop web questions -- Parallel Turbo achieved 51% accuracy at a 216ms median latency, compared to Exa Instant at 33.7% / 361ms, Tavily Ultra Fast at 19.3% / 357ms, Brave Search at 38.3% / 430ms, and SerpAPI at 23.3% / 999ms.
The cost story is equally sharp. With a median latency of 200ms and a price of $1 per 1,000 requests, Turbo is up to 14x cheaper than the default search in frontier models, while maintaining similar or better accuracy. For context, fast search APIs from competitors run $5--7 per 1,000 requests, frontier model search costs $10--14 per 1,000, and SERP APIs land around $1 per 1,000 but return raw results that require additional processing.
That last point matters. SERP APIs are cheap but dumb -- they hand you raw HTML and links. Turbo, like the rest of Parallel's Search API, returns dense, LLM-ready excerpts. By default, Turbo returns dense, relevant excerpts directly, so the model spends fewer input tokens on a better answer.
Why agents break the old search economics
The launch is timed to a structural shift in how software uses the web. A typical person runs a handful of searches a day, but an agent runs thousands: inside loops, across tools, and often multiple times per question. That volume makes the per-request price a first-class engineering constraint, not an afterthought.
Parallel is positioning itself as the only Search API built from the ground up for AI agents, meaning agents can specify declarative semantic objectives and Parallel returns URLs and compressed excerpts based on token relevancy. The distinction matters because traditional search ranks URLs for humans to click, while AI search needs something different: the right tokens in an agent's context window.
The practical consequence is that search costs have been a forcing function on product design. Teams building voice assistants, chat apps, or deep research pipelines have had to ration how often they call the web. At $1 per 1,000 requests, there is less need to ration search, and developers can run it in more scenarios than ever: from every chat request to every row in a batch job.
Use cases Turbo unlocks
Parallel is explicit about the workloads Turbo is designed for:
- Voice AI: Voice agents need to be low-latency to feel magical. Turbo can be used with real-time voice models to add web grounding without introducing awkward pauses.
- Deep research fan-outs: Applications that use deep research for decision-making can now perform significantly more searches, broadening the surface in comprehensive fan-outs.
- Consumer chat: Consumer chat applications can now use Turbo's low latency and cost to offer grounded search as the default user experience, rather than a heuristic-gated fallback.
- Reinforcement learning: Training models with web grounding depends on millions of live queries per rollout. Turbo cuts that latency so more of your compute budget goes to learning.
- Small, local models: Small models can answer harder questions with generous use of web search, making efficient local AI more viable.
The RL use case is the most interesting one to watch. Using live web search as a reward signal or grounding mechanism during training has been prohibitively expensive at scale. Turbo changes that math significantly.
A new architecture under the hood
Turbo is the first product powered by Parallel's new search architecture: a re-engineered search stack with innovation across hardware, model training, and index design to bring the cost and speed of search closer to zero. The company has not disclosed specifics of the architecture, but the latency numbers suggest meaningful changes at the infrastructure level, not just query optimization on top of existing systems.
Parallel maintains a large web index containing billions of pages, with crawling, retrieval, and ranking systems that add and update millions of pages daily to keep the index fresh. That proprietary index is the core moat -- Turbo's speed gains are meaningless without high-quality underlying data.
Who wins and who loses
The clearest winners are teams currently paying frontier model search rates. The Search API now supports turbo, basic, and advanced modes, with Turbo optimized for the fastest responses. Developers can now pick the right tier for the right workload rather than defaulting to a one-size-fits-all call.
The pressure lands squarely on Exa, Tavily, and Brave, all of which are undercut on both price and accuracy in Parallel's benchmarks. It also puts indirect pressure on OpenAI and Anthropic's built-in web search tools -- the Advanced mode provides higher quality with more advanced retrieval and compression, and defaults to advanced when omitted, meaning Parallel now spans from ultra-cheap to deep-research quality in a single API.
Real-world adoption is already moving. Nooks, the AI sales workspace, cut web grounding spend by 70.5% after moving from built-in LLM web search to Parallel's Search API, with better accuracy on time-sensitive facts. Harvey uses Parallel to ground legal reasoning across 60-plus jurisdictions. Notion's agents serve millions of users across knowledge work tasks.
The bigger picture
Parallel raised a $100 million Series B at a $2 billion valuation, led by Sequoia Capital, with Andrew Reed joining the board. Existing investors Kleiner Perkins, Index Ventures, Khosla Ventures, First Round Capital, Spark Capital, and Terrain Capital all increased their participation. The round more than doubled the company's valuation from five months prior and brought total capital raised to $230 million.
That capital is being deployed toward a specific thesis: the web's primary user is becoming AI agents, and the people who write, publish, and maintain the content those agents depend on need a direct stake in how it gets used. Turbo is the execution layer for that thesis -- make search so cheap and fast that every agent call becomes a web-grounded call by default.
To get started, Turbo is accessible via the Parallel Search API docs, the playground, the Parallel MCP server, and the CLI. Sign-up is free with no credit card required.