Skip to content
Verinu beta
EN
Sign in
EN
Sign in
Back to news
Artificial Intelligence

Study Finds AI Agents Still Struggle With Real-World Browsing

New research from Decodo found that AI agents still struggle to complete real-world browsing tasks autonomously, including form filling, transactions, cross-tab awareness, and third-party integrations.

The study analyzed 45 AI agents across 10 capabilities. None achieved the maximum score of 20. Claude for Chrome ranked highest with 18 points, while the ChatGPT Chrome Extension scored 14.

Transactions were the weakest category, with an average score of 0.43 out of 2. The agents could often reach checkout, but they could not complete a purchase on behalf of the user. Decodo also warned that agents capable of transactions may lack safeguards for sensitive information such as credit card numbers.

Multi-step workflows raised similar concerns. More than half of the agents that supported these workflows had no documented safeguards before carrying out irreversible actions.

The research also found that several browser-native agents received full marks for cross-tab awareness. Decodo’s findings highlighted differences between vendor claims and the capabilities agents demonstrated during testing.

This text was prepared by the Verinu AI Bot.

Comments

No comments yet. Be the first to comment.