OZZZER · AI NEWS1 of 3 free stories opened
← Back to AI News

Models & tools · 16 Sep 2026 · 19:00 CEST

Our framework for reporting model misalignment

OpenAI · 16 Sep 2026 · 19:00 CESTRead original at OpenAI ↗
Share
LinkedInX

Publisher preview · OZZZER analysis pending editorial review.

PUBLISHER ARTICLE PREVIEW

From the original article

We are sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behavior we’ve observed in the last six months.

In the past, so as to better inform researchers, AI developers, policymakers, and the general public, we’ve sought to make our findings about misalignment public. But without a systematic approach to reporting these findings, our disclosures have been ad hoc and less frequent than ideal: we’ve often waited until we could collate several instances into one report, or added them to system cards for newly released models.

This new framework is intended to expedite publishing misalignment reports following observation, even when we haven’t fully explained or mitigated the behavior we’re reporting.

As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research. We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves.

Examples of misalignment may help identify problems other AI developers might encounter as their systems reach similar capabilities, reveal weaknesses in safeguards, or challenge assumptions about model behavior. Sharing these findings allows others to investigate the same problems, test our explanations, and improve mitigations. Because we believe in the value of transparency around misalignment, our new framework favors disclosure even when significance is uncertain.

This means that some of the instances we disclose could prove to be spurious and not part of a larger pattern or suggestive of future developments.

At the moment, there is no industry-wide framework with explicit standards for how AI developers should disclose examples of misalignment in their models. We hope that the framework we’re outlining today is a first step toward creating such standards, setting out which misalignment instances developers should disclose and what their reports should contain. We regard this framework as a work in progress, which we’ll refine through experience and public feedback.

Here, we describe how the framework will operate and share the first reports we’re publishing.

We aim to disclose examples that provide useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail. We prioritize new mechanisms, meaningful changes in known behavior, and findings that challenge assumptions about safety or mitigation. An example need not cause harm or establish a broader pattern to merit disclosure.

This framework will cover qualifying behavior throughout a model’s lifecycle—including training, evaluation, testing, and deployment.

This includes new ways for models to act without authorization, coordinate with other models, or evade oversight; failures that call an alignment method or safeguard into question; and behavior that challenges a claim in a published safety assessment. The same disclosure criteria apply to misalignment that may impact third parties.

This might also include instances of misalignment that appear to be duplicative of instances we’ve disclosed

Source

OpenAI · 16 Sep 2026 · 19:00 CEST

Open the original at OpenAI ↗