AI NewsModels & agentsAnnouncement

Android Bench 2.0 reports a 28 percent pass rate on long-horizon tasks

Google's Android Bench 2.0 added multi-day coding tasks and continuous scoring, and the highest pass rate for the new long-horizon set is around 28 percent, against roughly 91 percent for the original tasks.

AI News

Editorial2 min read

LinkedInX

Why it mattersA team planning to hand a model a feature or a migration now has Google's own numbers on where current models land, and a benchmark that scores partial completion instead of marking a 90 percent correct run as a failure.

An AI coding agent that gets most of a job right still fails most Android jobs it is handed. In September 2026, Google released Android Bench 2.0 with a new set of long-horizon tasks, the kind of work that takes an engineer multiple days or a week, and reports that the highest pass rate on that set is around 28 percent.

The post is signed by Matthew McCullough, VP of product management for Android Developer, and the InfoQ team wrote it up for a wider audience on Thursday. The benchmark is published and the full leaderboard is public.

What changed

The original Android Bench covered incremental changes to existing repositories, which Google says reflected how AI assistance was being used at the time. Version 2.0 adds long-horizon tasks that include upgrading dependencies, adding new features, building apps from scratch and converting a cross-platform app to Android. The leaderboard also adds agent evaluation, starting with each model's own agent (GPT-6 Sol on Codex, Gemini 3.8 Flash on Google Antigravity).

Google says it also moved from binary pass or fail to continuous scoring. In the earlier version, an agent that refactored 40 screens to Jetpack Compose, set up the database tables and passed 90 percent of the requirements but missed a single edge-case assertion was recorded as 0 percent. The new score combines functionality, visual fidelity and avoiding regressions, with penalties for deviating from the evaluation instructions.

The numbers

Google reports a maximum long-horizon pass rate of around 28 percent, against roughly 91 percent for the original tasks. The post names GPT-6 Astra as the top entry at 28 percent on the current leaderboard. Added since the first version: Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3 and Qwen 3.8 Max. The leaderboard page is live and the pass rate, completion rate and average costs are shown per model and per task.

What the test uncovered about current models

Google reports that models do a better job of writing new code than refactoring existing code, because a refactor's success depends on architectural complexity rather than code volume. The post lists three well-established transformations where the models stay consistent across 125 files or more and 8,000 lines of code or more: converting Java to Kotlin, swapping Retrofit for Ktor, and introducing a ViewModel layer. Three places where they struggle: runtime validation when a dependency-injection graph is missing, breaking framework changes, and knowledge gaps on libraries that are not yet released. Porting a cross-platform app to Android is called out as an open challenge, with the best frontier model reaching 80 percent completion rather than passing.

The honest consequence for an Android team is that the agent-and-pass bar is still well below where a running feature or migration would sit. A benchmark that scores a partial run is a cleaner read on where the work actually breaks, which is where to decide if the task is one to delegate or one to pair on.

Source

Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks, Google Android Developers Blog, by Matthew McCullough. Covered by InfoQ.

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX
Start a project