1 of 1
Story summary
- Google launched Android Bench 2.0 to evaluate AI models on complex Android tasks.
- Android Bench 2.0 adds long-horizon tasks that can require engineers several days to a week.
- Evaluation uses continuous scoring based on functionality, visual fidelity, and regression avoidance.
- OpenAI’s GPT-6 Astra achieved the highest pass rate of 28% on the benchmark.
