Hark's Handoff Beats GPT-5 on Live Web Tasks at 10x Lower Cost

Hark's Handoff claims the #1 spot on Online-Mind2Web, beating GPT-5 and Claude Opus at a fraction of the token cost

·
·
  • New #1 on Online-Mind2Web: Hark Handoff claims the top spot on the industry's hardest live browser-use benchmark, beating GPT-5 and Claude Opus.
  • Order-of-magnitude cheaper: Handoff delivers frontier-level browser performance at less than 1/10th the token cost of competing models.
  • Real-world tasks, no API needed: It can order food on DoorDash, book flights, shop on Target, and recruit on LinkedIn -- all sites with no consumer API.
  • SFT + RL training stack: Built with supervised fine-tuning on curated demos, then reinforcement learning on live web sessions to handle unpredictable real-world friction.
  • Research preview now, platform this summer: Handoff is available to read about now; the full Hark software platform launches later this summer.
  • $700M-backed, hardware ambitions: Hark raised a $700M Series A at a $6B valuation and is building AI-native hardware alongside its models.

A new agent just claimed the top spot on the browser-use leaderboard. Hark Handoff is a computer use agent (CUA) that controls a real browser the way a human does, clicking, scrolling, and typing, and it now holds the #1 position on Online-Mind2Web, the most rigorous live web-agent benchmark in the field. It does this at less than one-tenth the token cost of competing frontier models.

Who built it, and with what backing?

Hark was founded by Brett Adcock, the entrepreneur behind robotics company Figure AI and electric aircraft builder Archer. He launched Hark in late 2025 with $100 million of his own capital. The company has since raised over $700 million in Series A funding at a $6 billion post-money valuation, in a round led by Parkway Venture Capital with participation from NVIDIA, AMD Ventures, Intel Capital, Qualcomm Ventures, Salesforce Ventures, and ARK Invest. The broader vision extends well past a browser agent: Hark is building multimodal AI systems alongside purpose-built hardware, aiming to serve as a universal interface between humans and machines. Handoff is the first public proof of that direction.

The benchmark that tests live websites

Online-Mind2Web evaluates web agents on 300 tasks across 136 popular sites in real time. The live component matters. Agents that perform well on static page snapshots often collapse when timing, layouts, and interaction flows shift. When the Online-Mind2Web team ran five frontier agents on live sites under human evaluation, most scored far below their older benchmark numbers. OpenAI Operator reached 61.3%; everyone else landed near 30%.

Hark claims Handoff now sits above all of them. According to the company's research preview, Handoff outperforms GPT-5 by 8 points and Claude Opus 4.8 by 2 points on average across three benchmarks, with a wider margin over Gemini 3.5 Flash and Gemini 2.5 Pro.

What Handoff actually does

Handoff runs inside a dedicated virtual computer, clicking buttons and filling forms on behalf of the user across websites it has never seen before. Each session gets its own sandboxed environment with a browser, file system, and terminal. Users can connect existing accounts, and Handoff logs in using saved addresses, preferences, and payment details.

The tasks it targets are the ones that consume time without producing much:

  • Food ordering — place an end-to-end order on DoorDash or Uber Eats using your saved address and payment
  • Shopping — compare prices across Walmart, Target, or Costco, add to cart, and check out
  • Travel — book flights on United, Delta, or American, and hotels on Booking.com
  • Reservations — grab tables on OpenTable or Resy using your preferences
  • Recruiting — discover, message, and schedule interviews with candidates on LinkedIn
  • Research — cross-reference Reddit, review sites, and news sources for a synthesized answer

Hark's internal research puts 74.9% of all computer use inside a browser. That figure is the core thesis: most of the web has no API. Sites like DoorDash, Target, Walmart, OpenTable, and LinkedIn offer no consumer APIs, so the only way to automate them is to navigate them the way a human would.

How it was built

Hark's training follows a three-stage roadmap. The team started with post-training, the results shared today, has since moved into mid-training, and has pre-training planned for later this year. Starting at post-training lets the team iterate quickly, refining data pipelines, evaluations, and infrastructure on the shortest path to measurable results.

Two techniques build on each other:

  1. Supervised Fine-Tuning (SFT) — training on curated demonstrations of correct browser behavior. This has the most impact on harder task distributions where the base model has little signal to work from.
  2. Reinforcement Learning (RL) — letting the agent succeed and fail on real tasks, then updating from the outcome. On benchmarks where the base model is already strong, RL supplies most of the remaining headroom.

The model outputs a raw sequence of cursor and keyboard inputs: x/y coordinates, clicks, scrolls, and keystrokes that directly control a live browser. There is no intermediate abstraction layer translating intent into DOM manipulation. The agent sees the screen and acts on it.

The hostile nature of the web is a core design constraint. Bot blocking, pop-ups, CAPTCHAs, and ads exist specifically to stop automated agents. Hark's modeling and engineering teams co-evolve the model alongside the systems that handle this live friction, rather than training in a clean sandbox and hoping it transfers.

Cost is the sharpest edge

Hark can serve Handoff at less than one-tenth the token price of competing frontier models. For a CUA, cost compounds faster than it does for a chatbot: these agents run multi-step sessions with screenshots at every turn. An order-of-magnitude reduction in per-token cost, combined with lower per-step latency, is what separates an economically viable product from a demo.

What remains unknown

Handoff is a research preview, not a shipping product. Anyone can try it later this summer with the initial release of Hark's software platform. Exact pricing for end users has not been disclosed, and while the benchmark scores are independently verified on the OM2W leaderboard, some numbers for competing models are self-reported. Online-Mind2Web scores can also shift depending on whether evaluation used human judging, WebJudge, or a custom agentic judge, and methodology varies across submissions. That context is worth keeping in mind when reading the margins over competitors.

Where this fits in the broader field

A well-funded, vertically integrated lab has demonstrated that a purpose-built, domain-specific model can beat general-purpose frontier models at a specific agentic task and do it cheaper. That pattern tends to repeat as AI matures into distinct application layers. The open question is whether browser-use is a narrow enough domain to hold that advantage, or whether the next generation of general models closes the gap. Developers and researchers can sign up for early access on Hark's website.

Trending
  • No trending articles

Comments

avatar

Next Reads