Meta AI Sweeps Five STEM Olympiads With Perfect Scores and Zero Tools
Meta's AI models swept five STEM Olympiads — including perfect scores in physics — with no tools, no calculators, and no search allowed.
- Meta AI swept five STEM Olympiads, earning perfect scores in physics (APhO, IPhO) and gold medals in math and chemistry.
- All results were achieved with zero tool use — no search, no code, no calculator — testing pure model reasoning.
- The APhO perfect score of 30/30 cleared the gold medal threshold of 21-23 by a wide margin.
- Meta joins Huawei, Xiaohongshu, Google, and OpenAI in a fast-moving race to Olympiad-level AI reasoning.
- The results signal that LLMs can now handle genuinely novel, multi-step reasoning — not just pattern-matched recall.
- IMO is now considered a saturated benchmark; the field needs new, harder evaluation standards.
Meta AI just announced that its models competed in five of the world's most prestigious STEM Olympiad competitions and came away with gold across the board. The haul: perfect scores on the theory exams of both the Asian Physics Olympiad (APhO) and the International Physics Olympiad (IPhO), a gold medal at the International Mathematical Olympiad (IMO), and gold-medal-level performance at the International Chemistry Olympiad (IChO) and the Romanian Masters of Mathematics (RMM). All of this was done with zero tool use , no search, no code execution, no calculator.
What gold actually means here
These aren't multiple-choice tests. The IMO, held annually since 1959, brings together the brightest young mathematicians from over 100 countries to solve proof-based problems that emphasize creativity and logical reasoning rather than memorization or computation. Physics Olympiads are similarly brutal, requiring contestants to derive solutions from first principles under time pressure.
The scores Meta is claiming aren't just clearing the bar , they're vaulting over it. On the APhO theory exam, Meta's model scored a perfect 30/30. To put that in context, the gold medal threshold for APhO and related IPhO events typically falls in the range of 21 to 23 out of 30 , a perfect score doesn't just clear the bar, it clears it by a comfortable margin.
The no-tools constraint is what makes this particularly striking. The model had to produce full written proofs and derivations relying entirely on its internal reasoning , the same conditions human contestants face.
The competitive landscape this lands in
Meta isn't the only lab racing toward Olympiad-level reasoning, and the field has been moving fast. AI has advanced from silver-medal performance in 2024 to gold-level scores in 2025, before now reportedly achieving perfection at the IMO. The trajectory is steep.