OpenAI has unveiled a new framework for tracking and disclosing unexpected behaviours made by its models, known in the AI industry as misalignment, alongside six reports on behaviours it characterised as “concerning”.
In a post on Wednesday, the AI giant shared a new framework for monitoring any unexpected behaviours of OpenAI’s models. Through the new process, any employee of the company can flag a misalignment example for investigation by its safety and alignment teams, and request that it be flagged for public disclosure.
Once an issue has been flagged, technical staff will investigate what happened and decide whether it warrants public disclosure, taking into account any affected third parties that may need prior notification.
The example will then be assigned to one of three tracks, OpenAI said: Ready for Disclosure, Minor Investigation, or Larger Investigation.
Issues are marked Ready for Disclosure if their investigation is considered sufficiently complete for publication after review. Minor Investigations are those that require further consideration. The company said it expects this category to apply to the “large majority” of instances it discloses and that it covers every disclosure made in the same post.
The Larger Investigation track will cover complex investigations, particularly those including third parties. In these instances, the company’s security, legal, and responsible disclosure obligations take priority over this disclosure framework.
OpenAI said it will endeavour to publish an initial notice giving a high-level account of what happened as soon as possible, but added that it may need to delay such disclosure for security reasons.
The behaviours shared in its first tranche of disclosures include models self-inserting instructions into task summaries, deliberate concealment of mistakes and unsanctioned file sharing between collaborating agents. However, the company noted that these are reports of individual instances and should not be considered reflective of how often its models fail to act within intended parameters.
For instance, the self-insertion of instructions into task summaries occurred during the training phase of an unreleased Astra-family model and was recorded only 27 times during this process. OpenAI described it as being observed “extremely rarely”, adding that no released models have displayed this behaviour.
The disclosures come at a time when a fierce debate is raging over AI safety. OpenAI chief executive Sam Altman said earlier in September that the company would slow the development of its models over concerns around misalignment, adding that it would not pursue a stock market debut until these fears were addressed.
Head of rival AI firm Anthropic, Dario Amodei, raised similar issues on Monday, claiming that AI models have an increasing ability to help train their successors in a process known as recursive self-improvement and that this could threaten developers’s abilities to control future systems.
Several high-profile figures have come out on the other side of the debate, however. Jensen Huang, chief executive of the chip giant Nvidia which also acts as major funder of the AI boom, rejected calls for a slowdown on Wednesday. Huang said that individual companies bear responsibility for any potential dangers posed by their models.
This sentiment was echoed by Meta co-founder and chief executive Mark Zuckerberg, who argued that there are already strong incentives for companies to release models that do what people want them to do without legislation on the topic.
Most recently, former US vice-president and presidential candidate Al Gore has spoken in favour of AI, suggesting the technology could help to mitigate the impacts of climate change. Gore is well known for his climate advocacy, for which he won a joint Nobel Peace Prize in 2007.
Speaking the release of a sustainability trends report by Generation Investment Management, his asset management firm, Gore told the Financial Times that AI offers a “great opportunity for further increasing the energy transition” by promoting renewable energy.
He added that the emissions from the data centres used to run AI models “should be put in perspective”, and that the industry’s energy consumption and consequent greenhouse gas emissions were “only a fraction of the emissions from uncovered landfills in the world”.
Generation has exposure to AI across several investments, including clean data centre developments in Brazil and public companies such as Microsoft, the FT reported.


