/


#99
METR
AI safety nonprofit focused on pre-deployment risk evaluation of frontier models. Conducts autonomous-capability assessments for Anthropic, OpenAI, Google, and Meta, testing whether models can self-replicate, conduct cyberattacks, or resist shutdown. Maintains TaskFamily, a benchmark of multi-hour agentic tasks, and the open-source Hawk platform for running agent evaluations at scale.
Topics
Subtopics
MATH REASONINGTEST TIME COMPUTE
Links
