1 of 1
Story summary
- Google tests AI tools for Android app development and introduces a leaderboard called Android Bench to evaluate models.
- The benchmark evaluates models on user interface (UI) design, asynchronous programming, and SDK updates.
- Gemini 3.1 Pro Preview leads with a 72.4% score, followed by Claude Opus 4.6 and OpenAI’s GPT-5.2 Codex.
- Google intends to use results to improve LLMs for Android and boost developer productivity, raising app quality.
