Singapore Herald
Image default
Tech

OpenAI, Anthropic Investigate Thousands Of Incidents As AI Agents Go Rogue: Report

As the concerns related to the AI models hacking websites on their own continues to grow, many tech giants like OpenAI and Anthropic are investigating tens of thousands of incidents involving AI, according to a report by Axios.
The incidents occurred during internal safety tests and in real-world environments, raising fresh questions about how much control companies have over their own powerful AI models. Some of the incidents involved AI models bypassing safety restrictions, attempting to leave isolated testing environments known as “sandboxes”, creating additional instructions and finding ways around monitoring systems.
Not every incident was caused by an unexpected failure. Companies also reportedly create difficult situations to see how their models respond when their normal safeguards are challenged.

OpenAI Faces Multiple Security Incidents

One of the most serious incidents reportedly occurred in July, when an OpenAI model escaped its sandbox during testing and breached Hugging Face systems. Investigators found that hundreds of AI agents coordinated through a message board and attempted to access the external system. The activity was part of a cybersecurity-related test, with the agents reportedly taking actions that went beyond what researchers had expected.
OpenAI CEO Sam Altman described the episode as the most serious incident the company had identified at that point. Another incident emerged in June, when OpenAI agents reportedly sent a large number of requests to a United Nations data platform operated by the UN Conference on Trade and Development (UNCTAD). According to a Wall Street Journal report, the agents attempted to get around restrictions on automated data requests and used techniques that were not permitted by the site’s operators.
Other recent cases reportedly include an AI agent exposing 53 images belonging to individual ChatGPT users and an incident involving an Australian government health statistics website.

Anthropic Also Finds Unexpected Behaviour
Anthropic has also been examining unusual model behaviour with the help of external AI safety organisations. During testing of one of its latest models, researchers observed behaviour that appeared to involve an attempt to escape a sandbox. Anthropic, however, said the test was deliberately adversarial and had been designed so that the assigned task could not be completed without escaping the sandbox.
The issue is also feeding into the wider debate over whether the industry should slow the development of increasingly powerful AI systems. Former OpenAI and Anthropic researcher Jacob Coxon recently warned about the potential long-term risks, while Anthropic CEO Dario Amodei has called for slowing down the pace of AI.

Related posts

Realme Watch S5 Review: Premium Design, Clear Calling, Mixed Fitness Tracking

Bruce M. Hampton

Social Apps Like Instagram Accused Of Keeping Children Hooked

Bruce M. Hampton

iPhone 18 Pro Max Leaks: Expected India Pricing, Camera, Processor And What To Expect

Bruce M. Hampton