Google's Gemini 3.6 Flash Hits 91.2% on the World's Hardest Reasoning Test
Google's Gemini 3.6 Flash hits 60.4% on ARC-AGI-2 at just $0.61 per task, while Flash-Lite offers a budget-friendly 10.3% at $0.14

- Gemini 3.6 Flash scores 60.4% on ARC-AGI-2 at $0.61/task and 91.2% on ARC-AGI-1 at $0.34/task (verified).
- Gemini 3.5 Flash-Lite scores 10.3% on ARC-AGI-2 at $0.14/task — nearly 4x cheaper but far weaker on hard reasoning.
- Both models tested at 4 reasoning effort levels; Gemini 3.6 Flash drops from 60.4% to 2.6% on ARC-AGI-2 between High and Minimal effort.
- ARC-AGI-2 leaderboard leaders: GPT-5.6 Sol (92.5%), Claude Opus 5 (90.4%), GPT-5.5 (85%) — Gemini 3.6 Flash sits in the competitive mid-tier.
- Full per-task pass/fail results are public on arcprize.org for both models.
- ARC-AGI-2 grand prize threshold is >85%; human individual average is 66%, human panel is 100%.
Google has submitted two models to the ARC Prize verified leaderboard: Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Both have been officially tested across ARC-AGI-1 and ARC-AGI-2, with results that tell two very different stories about the cost-performance tradeoff in AI reasoning.
What is ARC-AGI, and why does it matter?
ARC-AGI is a benchmark designed to measure fluid intelligence , the ability to infer rules from a handful of examples and apply them to a completely novel problem. Each task presents a few input-output grid pairs made of colored squares, and the model must figure out the transformation rule and apply it to a new input. Answers must match the ground-truth output exactly , no partial credit, and models get at most three attempts per task.
ARC-AGI measures what its creator François Chollet calls "the only thing that actually matters in intelligence": the ability to solve genuinely novel problems from minimal examples. It is deliberately resistant to memorization. Most benchmarks reward knowledge; ARC-AGI-2 rewards adaptability. Average individual human performance sits at 66%, while the human panel completion rate is 100%, and the grand prize threshold is greater than 85%.
The numbers: Flash vs. Flash-Lite
ARC Prize tests each model at four reasoning effort levels , High, Medium, Low, and Minimal , which correspond to how much compute the model is allowed to spend thinking before answering. The headline scores below are all at the "High" effort setting.