

Google's Android Bench just got its biggest overhaul since launch. The leaderboard, which measures how well LLMs handle real Android development tasks, has added 8 new models, swapped out its evaluation engine for the Harbor Framework, and opened the door for community-submitted tasks. If you're picking an AI coding assistant for Android work, this is the most comprehensive signal you have.
The scores that matter
Claude Fable 5 sits at the top of the leaderboard with a score of 84.5, followed by GPT 5.5 with 80.2, and Claude Sonnet 5 in third with a score of 76.2. That's a meaningful jump from the previous generation of models. For context, the original release saw models successfully complete between 16% and 72% of tasks -- the ceiling has moved significantly.
On the open-weight side, GLM 5.2 leads with 72.2, followed by Kimi K2.7 Code with a score of 70.4. That's a strong result for open models, putting them within striking distance of last generation's proprietary leaders.
The 8 new additions to the board are:
- Claude Fable 5 -- new #1 overall at 84.5
- GPT 5.5 -- #2 overall at 80.2
- Claude Sonnet 5 -- #3 overall at 76.2
- Claude Opus 4.8
- GLM 5.2 -- top open-weight at 72.2
- Kimi K2.7 Code -- #2 open-weight at 70.4
- MiniMax M3
- Qwen 3.7 Plus and Qwen 3.7 Max
Why a new evaluation engine changes everything
The bigger story here isn't the scores -- it's the methodology shift. Android Bench previously ran on mini-swe-agent v1, a general-purpose benchmarking agent adapted for Android. As part of the July release, Google has adopted the
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves
