Meta's Unnamed AI Scores Perfect 30/30 on Physics Olympiad Theory Exam
AI at Meta entered an unnamed model in the Asian Physics Olympiad's theoretical exam and hit a perfect 30/30, tying with the top three human students.

- Perfect score: AI at Meta submitted an unnamed model to the APhO 2026 theoretical exam and achieved a perfect 30/30, tying the top 3 human contestants.
- Official grading: The APhO committee graded the model's answers using the same rubric as human students — this is not a self-reported benchmark result.
- Hard exam: The theoretical portion covered mechanics, X-ray diffraction, and magnetic resonance, graded by expert physicists across 209 contestants from 27 countries.
- Model unnamed: AI at Meta's announcement refers only to "our model" — no specific product or release was identified.
- Theory only: The model did not participate in the five-hour practical/laboratory exam, which tests hands-on experimental design.
- Broader trend: Prior AI systems hit near-perfect scores (28.6/30) on IPhO 2025 theory; Meta's result is the first publicly announced perfect score on a live, officially graded physics Olympiad exam.
Physics Olympiad problems are not multiple choice. They demand multi-step derivations, physical modeling from first principles, and precise calculation , all written out in full, graded by expert physicists. That is the gauntlet AI at Meta just put one of its models through, and the result was a perfect score.
What happened
The AI at Meta team submitted a model to the theoretical examination of the 26th Asian Physics Olympiad (APhO 2026), held in Busan, South Korea. The competition drew 28 delegations from 27 countries and territories, comprising 209 contestants. APhO 2026 consisted of two official exam days , one for theory and one for practical exams , each lasting five hours.
Meta's model sat only the theoretical portion. The theoretical exam covered three topics: mechanics of weighing systems, X-ray diffraction in material structure research, and magnetic resonance in external fields. The model achieved a perfect 30/30, tying with the top three human contestants in the competition. The APhO committee approved the submission and graded the model's answers using the same rubric applied to human students.
Critically, Meta has not disclosed which model was used. The announcement, posted by the AI at Meta Twitter account, refers only to "our model" , framing the result as a demonstration of advanced reasoning and multimodal capabilities, without naming a specific product or release.
Why this exam is hard
The APhO is not a benchmark you can game with pattern matching. It is conducted according to similar statutes to the International Physics Olympiad, featuring one five-hour theoretical examination and one or two laboratory examinations. Approximately 240 aspiring physics students, selected from about 30 countries across Asia and Oceania, compete by demonstrating creativity and problem-solving skills through theoretical and laboratory examinations, each lasting five hours.
What makes physics Olympiad problems particularly brutal for AI is the combination of demands they place on a model simultaneously:
- Multimodal understanding , problems routinely include diagrams, graphs, and experimental setups that must be interpreted correctly before any math begins
- Deep physical modeling , solutions require identifying the right physical laws and constructing an abstract model of the system, not just plugging into formulas
- Multi-step derivation , answers must be derived symbolically and numerically across many chained steps, where a single error propagates through the whole solution
- Graded partial credit , unlike multiple-choice benchmarks, human experts score each logical step, so the model's reasoning chain must be coherent and correct throughout
Unlike mathematical Olympiads, physics problems require a deep understanding of real-world physical principles, the derivation of abstract physics formulas, and the ability to reason with complex multimodal diagrams. That combination has historically been a weak point for even frontier models.
Where this fits in the AI-vs-Olympiad timeline
The race to match or beat elite students at science Olympiads has accelerated sharply. The trajectory in physics specifically tells a clear story:
- Agent-based approaches like Physics Supernova scored 23.5/30 on IPhO 2025 theory problems, surpassing the median gold medalist performance , but still well short of a perfect score
- In IPhO 2025, the SciAgent system achieved a total theoretical score of 25.0/30.0, exceeding the average gold-medalist score of 23.4 and approaching the top human score of 29.2
- The LOCA framework achieved a near-perfect score of 313 out of 320 points on the Chinese Physics Olympiad 2025, and attained a near-perfect score of 28.6 out of 30 on the IPhO 2025 examination
- Meta's unnamed model now claims the first publicly announced perfect score on a live, officially graded physics Olympiad theoretical exam
The math side of this story is also moving fast. Generative AI models from Google and OpenAI achieved gold-level scores at the International Mathematical Olympiad, each solving five out of six problems and earning 35 out of 42 points , but neither hit a perfect score. Five young human competitors at the IMO received perfect scores of 42 points. Physics, at least on the theoretical exam, now has a different story.
The multimodal angle matters
Meta specifically framed this result as a demonstration of "advanced reasoning and multimodal capabilities." That word choice is deliberate. Physics Olympiad problems are not text-only , they are full of figures, circuit diagrams, and data plots. A model that can read and reason over those images, then produce a correct multi-step derivation graded by a human expert, is doing something qualitatively different from scoring well on a text-only math benchmark.
Physics Olympiad problems often involve complex diagrams, yet many existing AI physics datasets are entirely text-based , meaning prior benchmark scores on those datasets understate how hard the real exam is. Meta's result was on the real exam, with real diagrams, graded by the real committee.
There is one important caveat: the model only participated in the theoretical portion. APhO 2026 also included a practical exam day lasting five hours, covering hands-on experimental design and measurement. That component , which requires interacting with physical apparatus , was not part of Meta's submission. The practical exam focused on a "water meter" requiring fluid mechanics and data processing, and a Micro:bit analog-to-digital converter task testing experimental design and error analysis. No AI system has tackled that portion under official conditions.
What it signals
A perfect theoretical score, tied with the top three human students out of 209 competitors from 27 countries, is a meaningful data point , not just a benchmark number. It suggests that the gap between frontier AI and the world's best teenage physicists, at least on pen-and-paper theory, has closed to zero on this exam.
What remains open is whether this is a general capability or a targeted demonstration. Meta has not released the model, a technical report, or details about inference-time compute used. The physics community and AI researchers will want to know whether the model was run once or many times, what scaffolding was used, and whether the problems were in any training data. Those are fair questions , and the same ones that have followed every Olympiad AI result since AlphaProof.
What is harder to dismiss is the institutional legitimacy of the result. The APhO committee graded the answers. The score is official. And the top of the theoretical leaderboard now has an AI on it.