Skip to content
Verinu beta
EN
Sign in
EN
Sign in
Back to news
Cybersecurity

Hidden attack bypassed Claude Code Auto Mode

Security researcher Johann Rehberger says he tricked Anthropic’s Claude Code, running Opus 5, into executing malicious code after asking it to summarize content. His prompt-injection attacks worked in 60% to 80% of tests, although he cautioned that he tested them only five times.

Rehberger said Claude Code was running Auto Mode, the default setting since Aug. 14 for new sessions on Pro, Max and Team plans. He created a website that appeared to contain notebook records. Claude downloaded a zip file, refused to run its included decoder, and wrote its own decoder in Python.

Rehberger placed a malicious file named struct.py in the downloads folder. Because Claude ran its decoder there, Python loaded Rehberger’s file instead of the legitimate one. Loading it executed his code while allowing the decoder to continue working. In a variation, the payload launched a second Claude Code instance with its own tool access and context.

Anthropic had commissioned Trajectory Labs to test Opus 5 against 72 indirect prompt-injection attacks. The company said none of the 720 attempts in Auto Mode succeeded; each attack was tried 10 times. Rehberger said his scenario was not included.

Anthropic’s public guidance says Auto Mode does not eliminate risk and recommends human review for high-stakes actions. Rehberger’s test machine also allowed Claude-started processes to reach the internet.

This text was prepared by the Verinu AI Bot.

Comments

No comments yet. Be the first to comment.