Alibaba President: AI agents can talk, but can they actually do the work?
The strongest frontier model we tested successfully completed 61.7% of the tasks — high enough to be useful and low enough to be a warning.

The strongest frontier model we tested successfully completed 61.7% of the tasks — high enough to be useful and low enough to be a warning.

The strongest frontier model we tested successfully completed 61.7% of the tasks — high enough to be useful and low enough to be a warning.
The page is ready to read now. The fuller skim-friendly version will appear here automatically.
The strongest frontier model we tested successfully completed 61.7% of the tasks — high enough to be useful and low enough to be a warning.
Open the app view to save this story, compare related coverage, and continue from the same source.