Google has expanded Android Bench to evaluate coding agents on projects that may take a human engineer several days or a week. Version 2 adds long-horizon tasks such as upgrading dependencies, creating an application and moving a cross-platform project to native Android.

The benchmark also replaces an all-or-nothing result with continuous scoring. A submission receives credit for working functionality, visual fidelity and avoiding regressions, while violations of structural constraints carry defined penalties. This prevents one failed edge case from hiding progress on dozens of requirements.

Google reports that agents perform better when writing new code or making deterministic transformations, including Java-to-Kotlin conversion and replacing one networking library with another. They struggle more with runtime-only failures, breaking framework changes and unreleased libraries. In one cross-platform porting category, the best system reached 80 percent completion rather than finishing the task.

Full long-horizon pass rates remain low. At publication, Claude Opus 5.5 led with 32 percent, followed by GPT-6 Astra at 28 percent. Benchmark scores depend on the selected tasks, agent harness and evaluation rules, so they do not predict every repository. The update is valuable because it measures integration and regression control, not only whether a model can generate an isolated code snippet.