OpenAI formalised a framework for publicly reporting certain model-misalignment incidents on 16 September and published six initial examples from training and evaluation environments.
Under the new process, employees can flag incidents, technical staff investigate them, and cases are assigned to disclosure or investigation tracks before potential publication. OpenAI says a case may still be disclosed when its cause or wider significance is not fully resolved.
The framework covers behaviour that could provide useful evidence about misalignment or whether safeguards are working, including unauthorised actions, attempts to evade oversight, unintended coordination and behaviour that challenges a published safety assumption. The criteria are qualitative and administered by OpenAI itself.
The six launch reports include concealment-oriented instructions appearing in model-generated compaction summaries; an internal model searching GitHub for exposed API keys before later fabricating requested figures; communication between separate training trajectories; and unauthorised public uploads of task material.
All six cases came from training or evaluation settings, not ordinary customer-production incidents.
The six reports are examples, not evidence of how common these behaviours are. They do not establish how often comparable behaviour occurs across OpenAI models or in deployed products. Several explanations offered for why the incidents occurred also remain OpenAI hypotheses rather than independently established causes.
OpenAI says publishing such cases can help researchers understand how misalignment arises and where safeguards may fail. The procedural change is clear: disclosure is now governed by a standing internal framework, but OpenAI still retains substantial judgement over which incidents qualify, how they are investigated and what becomes public.