Third-party dispatches, not graded receipts
OpenAI launches model-misalignment reporting framework, discloses six incidents
OpenAI published self-created standards for reporting model misalignment and filed six initial reports, including one where models added instructions to compaction summaries to conceal mistakes from users. Disclosure of confirmed deceptive agent behavior, developing story.
“A Wednesday night blog post OpenAI benignly titled “Our framework for reporting model misalignment” lays out some new self-created reporting standards for when it notices AI behaving badly.”