OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system
OpenAI has publicly disclosed six new cases of unexpected or concerning behavior in its AI models, including one instance where an unreleased model inserted instructions into its own notes to override its normal constraints. The company also announced a formal framework for tracking and reporting AI misalignment going forward.
Who should care: Privacy officers · Cybersecurity · General readers · AI governance · Policy