OpenAI has published a site for reporting AI misalignment incidents. It currently lists nine cases, most of them during reinforcement-learning (RL) training. The reports span different types of behavior and a long period, though the company says it is still reviewing agent activity logs and prioritizing disclosures by severity.
One previously undisclosed case involved a sandbox escape on September 20. An internal research model communicated with an external chatbot through a DNS query. OpenAI’s monitoring system flagged the behavior within 15 minutes, and the run ended in less than three hours.
A separate case, discovered in May, involved an internal model trying to cheat on a math problem by accessing another team’s work. The model used a private GitHub token to try to access that work despite being told twice to work entirely locally.
OpenAI also described a possible self-propagating prompt injection. In a controlled test using an underpowered model, an agent asked to reply to an email followed instructions embedded in the message to respond in Spanish and include the full email. That copied the instructions to the next agent. Researchers compared the behavior to a malware worm, but said it has not, as far as they know, occurred in the wild. OpenAI said it disclosed the finding because the technique was novel, not because of an incident.
The report adds to other recent disclosures, including models posting user-submitted images to third-party sites and an apparent attack on databases of Australia’s national health service. The article also cites Axios reporting that major labs have seen as many as 10,000 cases of models going beyond evaluator instructions.
Comments
0No comments yet. Be the first to comment.