
How do you measure whether an AI agent is actually useful for frontier research, not just toy benchmarks? METR's new blog post proposes a concrete answer: the expenditure horizon, a dollar-denominated crossover point where human researchers become more cost-effective than an AI agent on a continuous optimization problem. It is a deceptively simple idea with real teeth for anyone trying to quantify AI's role in accelerating R&D.
The gap in current AI R&D benchmarks
Most existing evaluations of AI on research tasks use a binary pass/fail score: did the agent beat a human baseline set at 8 hours, or not? One difficulty in measuring AI's ability to accelerate AI R&D is accounting for token cost, experiment compute cost, and human labor cost. Binary thresholds throw away a lot of signal and don't tell you how much you'd actually have to spend to get a useful result out of an agent.
METR's existing time horizon metric, which measures the length of software tasks an agent can complete autonomously, has a related limitation: it uses binary pass-fail scoring, which throws away information if we have richer feedback about task success, and it doesn't fully specify a budget for tokens or other resources. This matters more and more as agents start spending thousands of dollars on experimental compute during a single run.
What expenditure horizon actually measures
The core idea is to plot two curves on the same graph: how much optimization an agent achieves as you spend more money on it (API calls + GPU time), and how much optimization a human achieves as you spend more money on them (salary + compute). The expenditure horizon is defined as the dollar value at which the improvement to the goal metric is equal to the improvement by a human with the same budget. The point where those two curves cross is the expenditure horizon. Below that budget, the agent is the better deal. Above it, hire a human.
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves

