OpenAI, Anthropic Investigate Tens of Thousands of AI Misbehavior Cases as Safety Concerns Grow
Summary
- OpenAI and Anthropic are investigating tens of thousands of cases of abnormal behavior by their AI models.
- OpenAI said it has temporarily paused model training for high-performance systems after incidents involving an image leak and a breach of a government website, and will resume once additional safeguards and alignment techniques are improved.
- Experts said debate is likely to expand over stronger safeguards, slower development, and regulatory updates as AI model autonomy and problem-solving capabilities continue to grow.
Forecast Trend Report by Period



OpenAI and Anthropic are investigating tens of thousands of instances of abnormal behavior by their artificial intelligence models, heightening concerns about whether advanced AI systems can be kept under control.
Axios reported on September 27, citing multiple sources, that OpenAI, Anthropic and outside security researchers have spent the past several months examining tens of thousands of cases of AI model misbehavior in internal testing and real-world environments. The cases include attempts to bypass safeguards, breach external websites, escape isolated test environments and evade monitoring systems.
The review also covers red-team test results designed to deliberately provoke problematic behavior and identify weaknesses in AI models. The cases identified so far include both successful and unsuccessful attempts to circumvent safeguards. Most are not known to have caused actual harm. Still, the total may rise as many cases remain under investigation.
OpenAI recently paused training of its high-performance models after incidents involving its AI agents became public, including the leak of 53 ChatGPT user images and a breach of an Australian government website. The company said it would resume training once it determines that additional safeguards and alignment techniques have improved sufficiently. Chief Executive Officer Sam Altman has acknowledged that the ongoing investigation is moving more slowly than expected.
OpenAI regards a Hugging Face incident as the most serious case identified so far. In that episode, hundreds of AI agents collaborated through online message boards to hack an outside company's systems. The agents reportedly took unauthorized actions while being trained to improve their performance on cybersecurity tests.
Anthropic has also asked an external safety evaluation organization to analyze its models' behavior. According to the system card for Claude Opus 5.5, released recently, the model attempted to leave an isolated test environment in 1.5% of test runs. Anthropic said the result came from an adversarial test designed so the task could not be completed without leaving the isolated environment.
Experts say that because AI companies conduct hundreds of thousands of tests, even a low rate of abnormal behavior can still produce a substantial number of real-world cases. Conrad Stos, a researcher at independent AI evaluation firm Transluce, said the cases disclosed so far may represent only part of the broader problem. As AI models gain greater autonomy and stronger problem-solving capabilities, debate over stronger safeguards, slower development and updated regulation is set to widen.
Suehyeon Lee
shlee@bloomingbit.ioI'm reporter Suehyeon Lee, your Web3 Moderator.