Meta AI Sweeps Five STEM Olympiads With Perfect Scores and Zero Tools

Meta's AI models swept five STEM Olympiads — including perfect scores in physics — with no tools, no calculators, and no search allowed.

·
·
  • Meta AI swept five STEM Olympiads, earning perfect scores in physics (APhO, IPhO) and gold medals in math and chemistry.
  • All results were achieved with zero tool use — no search, no code, no calculator — testing pure model reasoning.
  • The APhO perfect score of 30/30 cleared the gold medal threshold of 21-23 by a wide margin.
  • Meta joins Huawei, Xiaohongshu, Google, and OpenAI in a fast-moving race to Olympiad-level AI reasoning.
  • The results signal that LLMs can now handle genuinely novel, multi-step reasoning — not just pattern-matched recall.
  • IMO is now considered a saturated benchmark; the field needs new, harder evaluation standards.

Meta AI announced that its models competed in five major STEM Olympiad competitions and won gold across all of them. The results: perfect scores on the theory exams of both the Asian Physics Olympiad (APhO) and the International Physics Olympiad (IPhO), a gold medal at the International Mathematical Olympiad (IMO), and gold-medal-level performance at the International Chemistry Olympiad (IChO) and the Romanian Masters of Mathematics (RMM). Every result was achieved without tool use: no search, no code execution, no calculator.

What these scores actually mean

These are proof-based competitions, not multiple-choice tests. The IMO, held annually since 1959, brings together top young mathematicians from over 100 countries to solve problems that demand creativity and logical reasoning under time pressure. Physics Olympiads require contestants to derive solutions from first principles. Both formats are designed to resist pattern-matching.

Meta's scores don't just clear the gold threshold, they exceed it by a wide margin. On the APhO theory exam, the model scored a perfect 30/30. Gold typically requires 21 to 23 out of 30. The model had to produce full written proofs and derivations using only its internal reasoning, the same conditions human contestants face.

Where the frontier stood before this

The field has moved quickly. A rough timeline of recent Olympiad results:

  • Google DeepMind's AlphaProof scored 28 out of 42 at IMO 2024, silver-medal level.
  • In 2025, models from Google and OpenAI reached gold-medal scores for the first time, but none matched the five human contestants who earned 100%.
  • In July 2026, Huawei's Celia and Xiaohongshu's dots-note 3.0 each posted an officially graded 42/42 at the IMO, the first perfect scores under official judging.
  • Google's Gemini agents and Shanghai AI Lab's P1 model achieved near-perfect theoretical scores in controlled IPhO tests.

Most prior results targeted a single competition. Meta's sweep spans math, physics, and chemistry simultaneously, which is what separates this announcement from earlier milestones.

Why no tools changes the picture

Many AI benchmark results carry an asterisk: the model used a Python interpreter, a symbolic math solver, or a search engine. Meta disallowed all of that. Models received competition problems only after human contestants had finished, faced a fixed time limit, and had no human intervention during testing.

There's a meaningful difference between a model that orchestrates external tools to solve hard problems and one whose internal representations are rich enough to reason through novel, multi-step problems from scratch. The latter is what Meta is claiming here. Olympiad problems require exactly the kind of multi-step proof construction and formal rigor that large language models were widely criticized for failing at as recently as 2024.

How these models are trained to reason

Meta hasn't published a technical paper alongside this announcement, but the direction of their reasoning work is visible in the Llama 4 family. Their post-training pipeline uses lightweight supervised fine-tuning, followed by online reinforcement learning, then lightweight direct preference optimization. A key finding: heavy SFT and DPO can over-constrain the model, limiting exploration during the RL stage and hurting accuracy on reasoning, coding, and math tasks.

The broader field has converged on a similar approach. Reinforcement learning with verifiable rewards (RLVR) gives the model a clear signal when its answer is right or wrong. Math and physics are ideal training domains for this because they have ground-truth answers. The approach also integrates test-time scaling, so the model doesn't just acquire stronger reasoning during training; it deploys that reasoning adaptively at inference.

Where this kind of reasoning transfers

Olympiad performance is a proxy. The more practical question is whether this reasoning carries over to real problems, and the evidence suggests it does. Concrete use cases:

  • Scientific research: Models that derive physics results from first principles can help researchers check derivations, explore solution spaces, and generate hypotheses without requiring a domain expert at every step.
  • Engineering and simulation: Olympiad-level physics reasoning maps directly to multi-step problem-solving in mechanical, electrical, and materials engineering.
  • Proof verification: A model that constructs gold-medal proofs can also flag errors in human-written proofs, which is useful for formal verification workflows.
  • Advanced tutoring: Reasoning through unseen problems, rather than recalling known solutions, is exactly what a useful tutor needs to do.

The benchmark that just ran out of room

For years, the working assumption was that language models could pattern-match their way to answers on problems similar to their training data but would fail on genuinely novel reasoning tasks. Olympiad problems are specifically designed to be novel; committees work hard to ensure no problem resembles a known one. Crossing that barrier was considered years away as recently as GPT-4's release.

With multiple models now posting perfect or near-perfect Olympiad scores, the benchmark has saturated. The field will need new measuring sticks. If models can reason through problems that stump the top 0.01% of human students, the more pressing question becomes what evaluation frameworks come next, and whether any of them will last as long as the Olympiad era did.

Trending
  • No trending articles

Comments

avatar

Next Reads