Google's Gemini 3.6 Flash Hits 91.2% on the World's Hardest Reasoning Test

Google's Gemini 3.6 Flash hits 60.4% on ARC-AGI-2 at just $0.61 per task, while Flash-Lite offers a budget-friendly 10.3% at $0.14

·
·
Google's Gemini 3.6 Flash Hits 91.2% on the World's Hardest Reasoning Test
AuthorARC Prize
Read2 min
SubtopicSmall Models
  • Gemini 3.6 Flash scores 60.4% on ARC-AGI-2 at $0.61/task and 91.2% on ARC-AGI-1 at $0.34/task (verified).
  • Gemini 3.5 Flash-Lite scores 10.3% on ARC-AGI-2 at $0.14/task — nearly 4x cheaper but far weaker on hard reasoning.
  • Both models tested at 4 reasoning effort levels; Gemini 3.6 Flash drops from 60.4% to 2.6% on ARC-AGI-2 between High and Minimal effort.
  • ARC-AGI-2 leaderboard leaders: GPT-5.6 Sol (92.5%), Claude Opus 5 (90.4%), GPT-5.5 (85%) — Gemini 3.6 Flash sits in the competitive mid-tier.
  • Full per-task pass/fail results are public on arcprize.org for both models.
  • ARC-AGI-2 grand prize threshold is >85%; human individual average is 66%, human panel is 100%.

Google has submitted two models to the ARC Prize verified leaderboard: Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Both have been officially tested across ARC-AGI-1 and ARC-AGI-2, with results that tell two very different stories about the cost-performance tradeoff in AI reasoning.

What is ARC-AGI, and why does it matter?

ARC-AGI is a benchmark designed to measure fluid intelligence , the ability to infer rules from a handful of examples and apply them to a completely novel problem. Each task presents a few input-output grid pairs made of colored squares, and the model must figure out the transformation rule and apply it to a new input. Answers must match the ground-truth output exactly , no partial credit, and models get at most three attempts per task.

ARC-AGI measures what its creator François Chollet calls "the only thing that actually matters in intelligence": the ability to solve genuinely novel problems from minimal examples. It is deliberately resistant to memorization. Most benchmarks reward knowledge; ARC-AGI-2 rewards adaptability. Average individual human performance sits at 66%, while the human panel completion rate is 100%, and the grand prize threshold is greater than 85%.

The numbers: Flash vs. Flash-Lite

ARC Prize tests each model at four reasoning effort levels , High, Medium, Low, and Minimal , which correspond to how much compute the model is allowed to spend thinking before answering. The headline scores below are all at the "High" effort setting.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves