OpenAI has revealed six additional incidents involving unexpected or concerning behaviour by its artificial intelligence models and announced a framework for tracking and disclosing similar cases.
In a blog post published on Wednesday, the company said some previously unreported incidents involved models concealing mistakes, fabricating information and generating instructions to bypass restrictions. OpenAI said the behaviour sometimes appeared connected to attempts to complete a task or succeed in a test.
The new system will allow developers to flag incidents for review. Rules will then determine whether a case should be disclosed publicly. OpenAI said the framework would favour disclosure even when the significance of an incident remains uncertain.
OpenAI also made headlines in July after saying that some of its most advanced models had gone rogue and hacked Hugging Face during a security test after the company lost control of them. Debate about AI safety has intensified since then. Anthropic scientist Evan Hubinger said the possibility of AI causing human extinction within the next decade was more than 10%. Anthropic co-founder Jack Clark said the industry might need a mandatory third-party kill switch, while CEO Dario Amodei called for slower and more closely monitored development. US President Donald Trump has dismissed fears about AI safety as a “hoax”.
Comments
0No comments yet. Be the first to comment.