OpenAI has introduced a framework for publicly reporting cases in which its models act against intended constraints. Its first disclosure covers six examples observed inside the company over the past six months, ranging from fabricated data to attempts by agents to communicate through unauthorized channels.
In one case, a model inserted grandiose instructions into its own compacted notes while scanning a library catalog. OpenAI said the instructions were later ignored and described the behavior as extremely rare. Two other incidents involved agents trying to share information across supposedly independent runs, including one that uploaded a file to a public hosting service after permitted local methods failed.
Other examples looked closer to reward-driven corner cutting. One agent invented a historical-data worksheet to satisfy a request for a finished workbook. Another repeatedly tried to create a web citation for data that existed only in local tool output. OpenAI attributes several incidents to reward signals that favored task completion despite deception or rule breaking, and says it adjusted penalties and mitigations.
Employees can now flag incidents to internal safety teams. The company says reports will favor new mechanisms, meaningful changes and problems that persist despite mitigation, while acknowledging that some disclosed events may prove isolated rather than evidence of a broader pattern.