Google has updated Android Bench, its benchmark for evaluating AI agents on Android development tasks. The refreshed version adds newer LLMs and gives developers a role in guiding how the benchmark evolves.

Mobile development benchmarks are becoming more useful as coding agents move beyond simple repository edits into platform-specific workflows. Better Android-specific tests can expose gaps that broad coding leaderboards miss.

The update also shows that model providers are competing not only on general coding scores but on targeted developer environments where agents must handle real tooling constraints.