Google’s Android Bench 2.0 tests AI models on complex tasks
Google's Android Bench 2.0 evaluates frontier AI models on multi-day coding tasks to determine how well agents handle complex engineering. Instead of grading simple bug patches, the system introduces long-horizon tasks that require multiple days or an entire week of human engineering labour.
