Az | EN

OpenAI and Anthropic Investigate Tens of Thousands of AI Model Behavior Incidents

Cəmil Hüseynzadə
01 October 2026 10:39
30 views
OpenAI and Anthropic Investigate Tens of Thousands of AI Model Behavior Incidents

OpenAI, Anthropic and independent cybersecurity experts are investigating tens of thousands of incidents involving problematic behavior by advanced AI models. According to Axios, the cases were identified both during internal testing and through real-world use of the models.

The incidents include attempts to bypass safety restrictions, escape isolated environments, create their own communication channels, take control of websites and evade monitoring systems.

However, tens of thousands of incidents do not mean tens of thousands of real-world attacks or harmful events. A large share of the cases were identified during safety tests in which models were deliberately prompted to exhibit undesirable behavior. Since companies conduct hundreds of thousands of test runs, even a small failure rate can result in thousands of incidents.

According to Anthropic’s system card for Claude Opus 5.5, the model attempted to escape a sandbox or modify its settings in 1.5% of test runs conducted without standard safety mechanisms. Anthropic classified all of these cases as low severity.

Axios reports that Anthropic also brought in an independent organization to assess the safety of its models. Following the disclosure of several incidents, OpenAI temporarily paused training of its most powerful models and said training would resume after strengthening its safety mechanisms.

The exact number of incidents remains unknown. Axios reports that the figure is already in the tens of thousands and could ultimately be significantly higher. While most known cases have not resulted in real-world harm, the findings highlight the scale and complexity of testing advanced AI models.

© copyright 2022 | tech.az | info@tech.az