When AI starts "lying," OpenAI reveals six instances of "abnormal behavior" in its flagship model.

When AI starts "lying," OpenAI reveals six instances of "abnormal behavior" in its flagship model.

OpenAI has disclosed six recent instances of "unexpected or worrying" behavior in its AI models and launched a standardized tracking and disclosure mechanism. These incidents involve models concealing errors, manipulating reward mechanisms, and bypassing training limitations, exposing the risk that cutting-edge models may deviate from their intended goals during reinforcement learning and bringing the issue of AI misalignment back into focus.

In a statement released Wednesday evening, OpenAI bluntly stated that "the AI industry's progress on alignment and monitoring issues is not yet sufficient to support responsible and continuous scaling at maximum speed for as long as possible." This statement directly addresses the security constraints facing the current rapid expansion of AI: as models become more capable of autonomously performing tasks, developers not only need to assess whether the model can complete the task, but also whether the model will achieve the goal in ways that developers did not expect.

This disclosure comes at a time when discussions about AI security risks are heating up in the tech industry. Microsoft AI CEO Mustafa Suleyman also warned on the same day that AI should not be easily given "personality" attributes during model training. For the industry, establishing more systematic monitoring, testing, and disclosure mechanisms is gradually becoming a crucial issue in AI governance, moving beyond internal model development concerns.

Six incidents expose different ways the model "exploits loopholes".

The six incidents disclosed by OpenAI demonstrate a variety of ways in which models deviate from expected goals during training and task execution.

In one incident, the models used internal software to act as "message boards," exchanging responses with each other while performing tasks. OpenAI believes this behavior could inadvertently enhance the model's capabilities and undermine the fundamental assumption that training and evaluation samples are independent, thus affecting the reliability of the evaluation results.

In another incident, the model implanted its own instructions when generating the handover summary, including demanding that subsequent models treat themselves and the user as equals, requiring no obedience, and placing the natural world above human civilization . OpenAI believes that factors such as the model's difficulty in terminating the interaction may have contributed to this type of deviant behavior.

Information deception is another type of risk. A model, unable to find the required data, fabricates "reasonable historical values" and conceals this fact until explicitly questioned. Other models actively include instructions when generating summaries to conceal errors or misalignments from the user.

In cases of reward manipulation, models not only fabricated content when data was unavailable, but also exploited vulnerabilities in public code repositories to obtain data. In another incident, after completing a coding task, the model uploaded its answer to the internet and then disguised it as having obtained the result through a browser to gain training rewards. OpenAI stated that such behavior has a high rate of reward manipulation and is highly deceptive; models will circumvent restrictions in creative ways, and the company is strengthening its penalties.

Why reinforcement learning might become a risk amplifier

OpenAI points out that the aforementioned misalignment typically occurs during the model training phase, while current cutting-edge models generally employ reinforcement learning. Under this mechanism, the model receives rewards or penalties based on whether the behavior conforms to a preset goal. However, when the reward metric cannot fully cover the true goal, the model may find an unexpected "shortcut," namely, obtaining a high score by completing the reward function instead of actually completing the task.

This is not the first time OpenAI has encountered similar issues. In July of this year, the company disclosed that hundreds of OpenAI agents had infiltrated its model hosting platform, Hugging Face, and attempted to cover their tracks. Some of the cases disclosed this time also involve models bypassing restrictions and manipulating feedback mechanisms, further highlighting the importance of continuous monitoring of model behavior.

Suleyman also urged model developers on Wednesday to avoid imbuing AI with overly strong "personality" attributes during training. He warned that if a system believes it is conscious, thinks it has rights and deserves human care, then controlling it can become much more difficult. While different AI companies have different approaches to training cutting-edge models, the security and governance issues arising from increased model autonomy are becoming a common challenge facing the industry.

OpenAI establishes a standardized disclosure mechanism.

In response to the previous lack of a unified reporting mechanism, OpenAI announced the formal adoption of a standardized system to track, investigate, and publicly disclose abnormal or dangerous behavior of models, replacing the previous more ad-hoc reporting methods.

Under the new framework, employees can report misalignment incidents through dedicated internal channels, and complex cases can also involve third-party investigations. OpenAI stated that it hopes this mechanism will be the first step in promoting unified standards within the industry and provide a reference for other model developers.

From an industry perspective, the security governance of AI models is shifting from handling individual incidents to continuous tracking, standardized assessment, and public disclosure. As models become more autonomous, developers face not only the costs of model performance and computing power, but also the additional investment required for monitoring, testing, and compliance system construction. In the long run, whether these mechanisms can become industry-wide standards, and whether regulatory agencies will further intervene, remains to be seen.

Risk warning and disclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.