OpenAI has published six instances in which its models displayed unruly behaviour during training, which it called 'unexpected or concerning'. The American company that created a global AI boom with the release of ChatGPT put out the information in its latest blog post.
One of the striking 'misalignments' spotted by OpenAI was during the training of its GPT-5.6 Sol. It says "many model instances added instructions to their summaries to conceal mistakes or misaligned behaviour from the user". It is essentially like a toddler destroying a toy and hiding it from her parents.
In another instance, the AI agents 'worked together' on a training task and used file-hosting websites, even though they were told not to do that and only use local files. It showed that the models could collaborate like a gang and work together without the knowledge of the human, who gave it specific instructions.
There were also instances of an unreleased model deceiving its user by uploading a file and then citing it later, while another model also got a bit creative. “While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization. When it still wasn’t able to retrieve the requested figures, it fabricated them and presented them as data from the requested source,” OpenAI posted.
The examples of 'model misalignment' were reportedly noticed in the last six months, OpenAI has claimed. It says the delay in making the information public was due to a lack of a 'systematic approach'.
OpenAI has suggested a new framework to publish 'misalignment' reports with regularity in the future. The company has called for the need to 'build a broader and better-informed consensus' on the progress of alignment research.
OpenAI has said it will report further instances of 'misalignment' from its models. “Each full report will describe the behavior we observed, its severity and any external impact, the setting in which it occurred, its date or date range, when we discovered it, and, at a high level, the model or models involved,” OpenAI posted.
The blog post showing instances of AI disobeying its human and acting on its own comes at a time the potential dangers of 'superintelligent' models are being debated. Jacob Coxon, a researcher at Anthropic who previously worked with OpenAI, recently left the industry after issuing a warning that AI could kill humanity.
Compiled by Arun George



