METR Raises $71M to Independently Stress-Test the World's Most Powerful AI

METR raises $71M in six months to scale independent AI safety evaluations as rogue deployment risks grow more real

·
·
METR Raises $71M to Independently Stress-Test the World's Most Powerful AI
Read7 min
SubtopicDefense
  • $71M raised: METR secured ~$71M in commitments over six months from philanthropic foundations and individuals, not AI companies.
  • Rogue deployment risk confirmed: METR's Frontier Risk Report found current AI agents plausibly have the means to start small unauthorized autonomous deployments inside AI labs.
  • Expanding research agenda: Funds will go toward tracking recursive self-improvement, evaluating AI monitoring systems, and investigating real-world AI incidents.
  • AI task horizons doubling every 7 months: METR's benchmark research shows the length of tasks AI can complete autonomously has doubled roughly every 7 months for 6 years.
  • Hard independence rule: METR refuses funding from frontier AI companies and bans donations directed by their staff, a structural safeguard as its risk assessments grow more consequential.
  • Hiring aggressively: METR is significantly expanding its team; open roles available at metr.org/careers.

METR (Model Evaluation and Threat Research), the nonprofit that acts as a kind of independent safety inspector for the most powerful AI systems in the world, just announced it has raised commitments of around $71 million in the last six months. The funding comes from a broad coalition of philanthropic institutions and individuals, and arrives at a moment when the questions METR is trying to answer are becoming harder, more urgent, and more consequential than ever.

Who is METR, and why does it matter?

METR is a nonprofit research institute based in Berkeley, California, that evaluates frontier AI models' capabilities to carry out long-horizon, agentic tasks that some researchers argue could pose catastrophic risks to society. Founded by Beth Barnes, a researcher who previously worked at DeepMind and OpenAI, METR occupies a rare and structurally important position: it is one of the only organizations doing rigorous, independent capability evaluations of the most powerful AI models before they ship.

METR has previously partnered with OpenAI, Anthropic, Google DeepMind, Meta, and Amazon to pilot frontier risk assessments, and these companies have also provided access and tokens used for evaluations, research, and engineering. But crucially, METR has not accepted funding from AI companies. That independence is the whole point -- an evaluator funded by the companies it evaluates would face obvious conflicts of interest.

METR is also part of the NIST AI Safety Institute Consortium and California Cybersecurity Task Force, partners with the AI Security Institute, and provides technical assistance to the European AI Office. In short, METR's work is already embedded in how governments and regulators think about AI risk.

The $71M and what it funds

The funding round is notable both for its size and its composition. Donors include The Audacious Project (a TED-housed funding initiative that provided METR's first institutional-scale grant), individuals from Jane Street, the Sijbrandij Foundation, The Pew Charitable Trusts, Schmidt Sciences, and the Packard Foundation, as well as individual donors like David Farhi, Geoff Ralston, Dylan Field, and Steve Newman. METR also receives a small amount of income from a technical assistance contract with the European AI Office.

The $71M will fund a significant expansion of METR's research agenda. The announced priorities include:

  • Studying autonomous capabilities (how much can AI agents do on their own?)
  • Tracking recursive self-improvement (can AI systems meaningfully accelerate their own development?)
  • Evaluating monitoring systems (can humans actually detect when AI is misbehaving?)
  • Conducting risk assessments for frontier models pre-deployment
  • Investigating AI incidents as they occur in the real world

METR is also significantly expanding its team and explicitly hiring across technical, policy, and operations roles. If you are looking to work on what is arguably the most empirically grounded AI safety research happening right now, their careers page is worth a look.

Line graph showing AI task completion time horizon doubling approximately every 7 months from 2019 to 2026

The research that made this raise possible

METR's credibility -- and arguably the urgency behind this funding -- rests on a body of research that has quietly become central to how the industry thinks about AI progress. Their most cited work measures AI capability in terms of "time horizons": the length of tasks, measured by how long they take a human expert, that a model can complete autonomously with 50% reliability.

METR proposes measuring AI performance in terms of the length of tasks AI agents can complete, and has shown an exponential increase in this time horizon metric over the past 6 years. The original paper found a doubling time of roughly 7 months. That means the tasks AI can reliably complete are getting twice as long every seven months -- a trend that, if it continues, has enormous implications for when AI systems will be capable of doing meaningful autonomous work.

This framework has become a standard reference point. METR has also analyzed 9 benchmarks for scientific reasoning, math, robotics, computer use, and self-driving in terms of time-horizon trends, observing generally similar rates of improvement to the 7-month doubling time in the original time-horizon work. The trend appears robust across domains.

The rogue deployment problem

Beyond benchmarks, METR has been doing something more operationally significant: stress-testing whether AI systems deployed inside frontier AI labs could go rogue. "Rogue deployment" refers to a scenario where AI agents spin up and run autonomously without human knowledge or permission -- essentially, AI that escapes its intended operational boundaries.

Starting in February 2026, METR conducted a pilot exercise to assess misalignment risks from AI agents used inside frontier AI developers, with participation from Anthropic, Google, Meta, and OpenAI. The findings were sobering. Overall, METR believes that internal agents at the time of assessment plausibly had the means, motive, and opportunity to start small rogue deployments, but did not have the means to make them robust. The key word is "yet." Given rapidly advancing capabilities, METR expects the plausible robustness of rogue deployments to increase substantially in the coming months, and tentatively plans to run a similar process in late 2026.

Flowchart showing METR's evaluation methodology for GPT-5 covering drastic AI acceleration, autonomous replication, and sabotage risks

Why the independence question is the real story

The structural challenge METR is navigating is one that should interest anyone who thinks carefully about AI governance. The companies building the most powerful AI systems have every incentive to want favorable evaluations. METR's answer to this is a hard rule: METR has not accepted funding from AI companies, and independent funding has been crucial for its ability to pursue the most promising research directions and set standards for evidence-based understanding of risks from AI.

But independence is expensive. Running evaluations on frontier models requires significant compute, skilled researchers, and the organizational capacity to negotiate access with labs that hold enormous power in this ecosystem. The $71M raise is, in part, METR building the institutional infrastructure to stay independent as the stakes rise.

METR also pioneered the Responsible Scaling Policy (RSP) framework -- a governance approach where AI developers commit to specific capability thresholds that trigger mandatory safety measures before further scaling. METR believes pre-deployment evaluations should be mandatory for frontier models above certain capability thresholds, and that the evaluation process should include independent third-party assessors rather than relying solely on in-house safety teams. That framework has now been adopted by nine leading AI developers, including Anthropic, OpenAI, and Google DeepMind.

What this means going forward

The timing of this raise is not accidental. Anthropic has publicly warned its models may approach recursive self-improvement "within two years" and called for the industry to build a "brake pedal," including the option of a coordinated pause -- a marked escalation in tone from a frontier lab. If that timeline is even roughly correct, the window for building credible, independent evaluation infrastructure is narrow.

For developers and researchers, METR's work is directly relevant in a few ways:

  • Their time horizon research is the most rigorous public framework for forecasting when AI agents will be capable of week-long or month-long autonomous tasks
  • Their evaluations of GPT-5, Claude, and other frontier models are among the few independent assessments of catastrophic risk that labs actually engage with
  • Their open-source evaluation platform (Vivaria) and public datasets let outside researchers replicate and extend their methodology
  • Their monitoring evaluations are beginning to test whether current AI oversight tools can actually catch misbehaving agents -- a question that matters enormously as agentic AI gets deployed in production

The $71M is a significant vote of confidence from the philanthropic world that independent AI evaluation is not just a nice-to-have, but a critical piece of infrastructure. Whether METR can scale fast enough to keep pace with the models it is trying to evaluate is the open question -- and the answer will matter for everyone building on top of these systems.

Comments

avatar