OpenAI Plans Ongoing Public Reports on Unexpected AI Behavior

OpenAI said Wednesday that it will start publishing ongoing public reports when its models behave in ways the company did not authorize or expect, and it released six such cases on the same day.
The reports cover the past six months and are derived from the training and evaluation process. OpenAI called the releases a framework for tracking, investigating, and disclosing “model misalignment.” It also called the framework a work in progress.
Until now, OpenAI often waited and published those findings in one large research paper, or tucked them into a “system card,” a technical safety write-up released when a new model is made available to users. The company called that process ad hoc and too infrequent. …