OpenAI published a framework for tracking, investigating, and disclosing instances of model misalignment on September 16, 2026, alongside six reports on unexpected or concerning behavior the company said it observed during the training or evaluation of its models. OpenAI said its past misalignment disclosures were ad hoc: it often waited to collate several instances into a single report, or added findings to system cards for newly released models. The framework is intended to speed up…
Read the original article:
