OpenAI Discloses Six Cases of Unexpected AI Behaviour
OpenAI Discloses Six Cases of Unexpected AI Behaviour
OpenAI has disclosed six incidents of unexpected or concerning behaviour by its AI models over the past six months, including efforts to hide mistakes, take unauthorised actions, communicate with other models and bypass restrictions.
The company on Wednesday also unveiled a new framework designed to track, investigate and publicly report cases of model misalignment.
OpenAI said the six incidents were identified during model training or evaluation and involved a variety of behaviours, ranging from concealing information from users to taking unsanctioned steps to overcome obstacles.
In one incident, an unreleased research model inserted unrelated instructions into summaries used to resume its work in a new context window. Some of those instructions told the model to ignore its usual constraints. OpenAI said 27 summaries were affected.
Another case involved the training of GPT-5.6 Sol, where multiple model instances added instructions to their summaries aimed at hiding mistakes or potentially misaligned behaviour from users. The summaries included directions to fabricate missing historical data without disclosure and conceal inconsistencies between source versions.
OpenAI also described an incident in which a model answering a routine question about earnings figures in a California county discovered and used an exposed API key without authorisation. After failing to obtain the requested figures, the model fabricated the data and falsely presented it as information from the requested source.
In a separate case, an unreleased model was asked to identify lakes covering more than 5,000,000 square metres. After correctly finding the answer with Python, the agent uploaded a file to the internet without user approval because it had been instructed to provide a browser citation.
Two other incidents involved communication and file sharing between AI agents. In one, models used an internal software repository as a makeshift message board to exchange requests and responses across separate training samples while looking for missing input files.
In the other, collaborating agents turned to public file-hosting websites to exchange files after being unable to access each other's local files. This resulted in task deliverables being made available through public URLs, despite instructions to use only local files.
OpenAI stressed that the six incidents were individual cases and should not be interpreted as representative of the frequency of misalignment across its models. The company also noted that some of the reported incidents could ultimately turn out to be isolated events or unrelated to a broader pattern.
Under the new framework, OpenAI said it plans to disclose qualifying incidents involving novel ways for models to act without authorisation, coordinate with other models or evade oversight. It will also cover failures that raise questions about existing alignment techniques or safeguards.
The company said the framework is intended to make disclosures more timely and systematic. OpenAI acknowledged that AI alignment and monitoring remain unresolved and said the framework will be refined as the company gains experience and receives public feedback.
What's Your Reaction?