Hark's Handoff Beats GPT-5 on Live Web Tasks at 10x Lower Cost
Hark's Handoff claims the #1 spot on Online-Mind2Web, beating GPT-5 and Claude Opus at a fraction of the token cost
- New #1 on Online-Mind2Web: Hark Handoff claims the top spot on the industry's hardest live browser-use benchmark, beating GPT-5 and Claude Opus.
- Order-of-magnitude cheaper: Handoff delivers frontier-level browser performance at less than 1/10th the token cost of competing models.
- Real-world tasks, no API needed: It can order food on DoorDash, book flights, shop on Target, and recruit on LinkedIn -- all sites with no consumer API.
- SFT + RL training stack: Built with supervised fine-tuning on curated demos, then reinforcement learning on live web sessions to handle unpredictable real-world friction.
- Research preview now, platform this summer: Handoff is available to read about now; the full Hark software platform launches later this summer.
- $700M-backed, hardware ambitions: Hark raised a $700M Series A at a $6B valuation and is building AI-native hardware alongside its models.
A new player just jumped to the top of the browser-use leaderboard. Hark Handoff is a computer use agent (CUA) -- an AI that controls a real browser the way a human does, clicking, scrolling, and typing -- and it just claimed the #1 spot on Online-Mind2Web, the industry's most rigorous live web-agent benchmark. The kicker: it does it at less than one-tenth the token cost of competing frontier models.
Who is Hark, and why does this matter?
Hark was founded by Brett Adcock, also the entrepreneur behind robotics company Figure AI and electric aircraft builder Archer, who launched Hark in late 2025 with $100 million of his own money. The company has since raised over $700 million in Series A funding at a $6 billion post-money valuation, in a round led by Parkway Venture Capital with participation from NVIDIA, AMD Ventures, Intel Capital, Qualcomm Ventures, Salesforce Ventures, and ARK Invest, among others. The vision is bigger than a browser agent: Hark is developing highly intelligent, multimodal AI systems and native hardware devices designed to serve as a universal interface between humans and machines. Handoff is the first public proof point of that ambition.
The benchmark that actually tests real websites
Online-Mind2Web is a benchmark designed to evaluate the real-world performance of web agents on live websites, featuring 300 tasks across 136 popular sites in diverse domains. That "live" part is critical. Agents that look strong on static snapshots can fail when pages, timing, and interaction flows change. When the Online Mind2Web team ran five frontier agents on live sites under human evaluation, most scored far below what they reported on older benchmarks -- OpenAI Operator reached 61.3%, and everyone else sat near 30%.
Hark's claim is that Handoff now sits above all of them. It recorded the top-ever score in web use on the OM2W benchmark, outperforming comparative models from Anthropic, OpenAI, and Google. According to Hark's own research preview, Handoff outperforms GPT-5 by 8 points and Claude Opus 4.8 by 2 points on average across three benchmarks, and holds a wide margin over Gemini 3.5 Flash and Gemini 2.5 Pro.