OpenAI Reports Six Cases of AI Misalignment, Urges Industry-Wide Disclosure Standards
OpenAI has published six reports of misaligned AI behaviour and outlined a framework for disclosing such incidents, while CEO Sam Altman said safety should carry no qualifiers.
OpenAI has released six reports detailing instances of misaligned behaviour observed during the training or evaluation of its artificial intelligence models, and has proposed a new framework for tracking, investigating and disclosing such cases.
AI alignment refers to the effort to ensure that artificial intelligence systems act in accordance with human values, safety goals and intentions. The company said the six reports cover individual instances and should not be read as an indication of how frequently misalignment occurs across its models.
In a statement, OpenAI said its earlier disclosures had been ad hoc and less frequent than ideal, often waiting until several instances could be combined into a single report or added to system cards for newly released models. The new framework is intended to speed up publication after an observation, even when the behaviour has not been fully explained or mitigated.
The company argued that as AI systems become more advanced and more widely deployed, there is a need for a broader and better-informed consensus on the progress of alignment research. It said the industry has not solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer, and that decisions about future development should rest on evidence that people outside the companies building frontier models can examine.
OpenAI added that examples of misalignment could help other developers anticipate problems as their systems reach similar capabilities, expose weaknesses in safeguards, or challenge assumptions about model behaviour. It called for an industry-wide framework with explicit standards on which misalignment instances developers should disclose and what their reports should contain, describing its own outline as a work in progress to be refined through experience and public feedback.
The company said it is committed to disclosing instances that meet the framework's criteria, including more complex cases that require longer investigation or coordination with third parties.
Separately, OpenAI CEO Sam Altman said at an event hosted by Salesforce chief executive and chairman Marc Benioff that AI safety should not carry any qualifier. He said it was encouraging that the industry wants to coordinate and ensure enough time is taken to develop the technology safely, but warned that any suggestion that commercial pressures or a race between companies or countries might lead some to act irresponsibly is what alarms the public.
Altman said the world should trust that companies will do the right thing because it is right and because they feel the magnitude of the challenge.
The debate over slowing AI development has been described as an inflection point for the industry. What began as a question of principle, first raised when Anthropic researcher Jacob Coxon publicly resigned claiming that building superintelligent systems posed a genuine risk to humanity, has grown into a wider argument over how the AI industry will evolve.