OpenAI Web Search Scores 74 but Loses on Cost to Cheaper Rivals
OpenAI's built-in web_search tool lands fifth on the Artificial Analysis Search Index at 74 points, trailing Perplexity, Octen, Parallel, and Brave.
- OpenAI Web Search debuts at 74 on the Artificial Analysis Search Index, ranking 5th among providers.
- Costs around $0.05 per task: $0.04 search ($10 per 1k calls) plus $0.009 model inference.
- Scores 72% on AA-Omniscience, nearly matching leader Firecrawl at 73%.
- Weakest on BrowseComp multi-hop browsing at 73.5%, 13th of 26 variants.
- Uses ~40k input tokens per task versus 125k for the leanest external Search API.
- Perplexity (medium) leads at 80, Octen beats OpenAI at half the price.
OpenAI Web Search scores 74 in its first Search API benchmark
Artificial Analysis has added OpenAI Web Search to its Search Index, giving the built-in tool its first direct comparison with standalone Search APIs. OpenAI scored 74, ranking fifth among providers and seventh among 26 product configurations. The result shows strong factual retrieval and efficient token use, alongside weaker performance on research that requires several linked searches.
One call, different plumbing
OpenAI Web Search is the index’s first integrated, first-party search tool. Standalone providers use Stirrup, an agent harness that lets a candidate model call an external search endpoint, inspect extracted pages, and repeat the process. OpenAI handles retrieval, reasoning, and answer generation inside one Responses API request, with the model choosing when and how often to search.
Artificial Analysis paired each provider with GPT-5.6 Luna at medium reasoning effort. The standard setup allowed up to 25 agent turns, returned as many as 10 results per search, and extracted page text with a 15-second timeout. The OpenAI run bypassed Stirrup and set search_context_size to medium.
The common model makes answer quality more comparable across providers, although the orchestration remains different. OpenAI’s score reflects its complete integrated workflow, while standalone API scores include Stirrup’s query planning, page extraction, and result handling.
Factual lookup leads its profile
The Search Index gives equal weight to three benchmarks. Its published score uses the underlying results, while the component figures below are displayed at lower precision.
| Benchmark | Task | Result | Relative position |
|---|---|---|---|
| AA-Omniscience | Factual question answering | 72% accuracy | Near Firecrawl’s leading 73% |
| DeepSearchQA | Broad research | 78 F1 | Mid-pack |
| BrowseComp | Multi-hop web browsing | 74% accuracy | 13th of 26 |
F1 combines precision and recall into one score, rewarding answers that recover relevant information without adding excessive incorrect material. BrowseComp measures a different capability: following clues across multiple searches and pages to reach an exact answer. Artificial Analysis used a 200-question subset drawn from BrowseComp’s 1,266-question evaluation pool.
OpenAI’s strongest relative result came from direct factual questions. Its 74% BrowseComp accuracy trailed Perplexity and Octen, which scored between 85% and 87%. That 11-to-13-point gap matters for agents that investigate obscure entities, reconcile details across sites, or build answers through several search steps.
GPT-5.6 Luna scored 33 without search, so the integrated tool added 41 index points. Perplexity’s medium configuration produced a 47-point gain with the same candidate model.
A score of 74 costs about five cents
OpenAI charges $10 per 1,000 web_search calls and bills retrieved search content as model input tokens. Under the benchmark workload, search averaged $40.19 per 1,000 tasks and model inference added $9.40. The combined cost was $49.59 per 1,000 tasks, or about $0.05 each.
| Provider | Index score | Search per 1,000 tasks | Model per 1,000 tasks | Total per 1,000 tasks | Time per task |
|---|---|---|---|---|---|
| Perplexity, medium | 80 | $62.30 | $8.04 | $70.34 | 27.5 seconds |
| Octen, highlights | 77 | $9.07 | $14.92 | $23.99 | 15.9 seconds |
| Parallel, advanced | 75 | $47.93 | $11.29 | $59.22 | 41.4 seconds |
| Brave, LLM context | 75 | $61.96 | $17.46 | $79.42 | 22.7 seconds |
| OpenAI Web Search, medium | 74 | $40.19 | $9.40 | $49.59 | 30.4 seconds |
| Model only | 33 | N/A | $2.92 | $2.92 | 21.0 seconds |
OpenAI was cheaper than 17 of the 25 standalone Search APIs in the index. Octen’s API delivered a higher score for less than half the total cost, placing OpenAI outside the quality-cost frontier in this run. Perplexity led on quality at a higher total price.
Integrated retrieval cuts the payload
OpenAI consumed roughly 40,000 input tokens per task, including retrieved content. The smallest measured footprint among the external Search API configurations was about 125,000 tokens. That difference kept OpenAI’s model-inference cost near $0.009 per task, even with GPT-5.6 Luna performing the reasoning.
Smaller context payloads can reduce input-processing latency, memory use for the key-value cache that stores prior token states, and exposure to context-based pricing. Those savings become more relevant when an application launches searches in parallel or runs repeated research loops.
The API surface stays small
A Responses API request gains web access through a tool declaration. The model can then decide whether to search and can issue several searches before returning an answer.
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-5.6-luna",
reasoning={"effort": "medium"},
tools=[{
"type": "web_search",
"search_context_size": "medium"
}],
input="Your question here"
)
print(response.output_text)One Responses request can trigger multiple billable search calls. The benchmark’s $40.19 search cost per 1,000 tasks, combined with OpenAI’s $10 rate per 1,000 calls, corresponds to roughly four search calls per task on average.
OpenAI does not expose search duration separately, so Artificial Analysis estimates it from streamed events. The reported search segment is an upper bound that includes some OpenAI processing time, which limits direct latency comparisons with standalone APIs. The benchmark also omits per-query metrics because an integrated request can generate several internal queries.
Where the built-in tool fits
Teams already using the Responses API can remove a separate search integration, including vendor authentication, request formatting, result parsing, and some retry handling. The benchmark supports the built-in tool for workloads with these characteristics:
- Questions center on direct factual lookup.
- Low token use matters for cost, latency, or context limits.
- A single API and billing relationship simplifies operations.
- End-to-end answer generation matters more than direct control over raw search results.
BrowseComp-style research provides a strong reason to test external providers. Perplexity Search and Octen led OpenAI by more than 10 points on that benchmark, and Octen paired its higher overall score with substantially lower cost.
Production evaluations should also measure source coverage, freshness, citation accuracy, regional behavior, rate limits, failure handling, and data-retention terms. The Artificial Analysis index captures answer quality, cost, token use, and latency under one controlled configuration; application-specific tests remain necessary before selecting a search stack.