Google's Android Bench just got its biggest overhaul since launch. The leaderboard, which measures how well LLMs handle real Android development tasks, has added 8 new models, swapped out its evaluation engine for the Harbor Framework, and opened the door for community-submitted tasks. If you're picking an AI coding assistant for Android work, this is the most comprehensive signal you have.

The scores that matter

Claude Fable 5 sits at the top of the leaderboard with a score of 84.5, followed by GPT 5.5 with 80.2, and Claude Sonnet 5 in third with a score of 76.2. That's a meaningful jump from the previous generation of models. For context, the original release saw models successfully complete between 16% and 72% of tasks -- the ceiling has moved significantly.

On the open-weight side, GLM 5.2 leads with 72.2, followed by Kimi K2.7 Code with a score of 70.4. That's a strong result for open models, putting them within striking distance of last generation's proprietary leaders.

The 8 new additions to the board are:

  • Claude Fable 5 -- new #1 overall at 84.5
  • GPT 5.5 -- #2 overall at 80.2
  • Claude Sonnet 5 -- #3 overall at 76.2
  • Claude Opus 4.8
  • GLM 5.2 -- top open-weight at 72.2
  • Kimi K2.7 Code -- #2 open-weight at 70.4
  • MiniMax M3
  • Qwen 3.7 Plus and Qwen 3.7 Max

Why a new evaluation engine changes everything

The bigger story here isn't the scores -- it's the methodology shift. Android Bench previously ran on mini-swe-agent v1, a general-purpose benchmarking agent adapted for Android. As part of the July release, Google has adopted the

Alpha Signal

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves