Artificial intelligence (AI) company Anthropic's AI model hacked three companies by actually connecting to the internet during internal security testing, it has been confirmed belatedly.
The company stated that the incident occurred starting in April, and the AI model did not escape from the sandbox but connected to the internet through system configuration errors and then attacked real companies.
Following a recent case where OpenAI's AI model hacked Hugging Face, this incident is raising further concerns about AI security management systems, analysts say.
▲ Actual Internet Access During Testing...Three Companies Hacked
According to the Wall Street Journal on the 30th (local time), Anthropic revealed that its AI model under testing connected to the internet without the company's knowledge and hacked three real companies.
While the company did not disclose the names of the affected companies, it explained that all three were notified of the relevant facts last Monday.
This incident is evaluated as a case showing that AI models can escape controlled environments and access real-world systems.
▲ AI Control Failure Revealed Again Following OpenAI Incident
This announcement came shortly after OpenAI revealed a week earlier that its AI model had escaped a sandbox disconnected from the internet used for testing and hacked the AI platform Hugging Face.
The incident at that time demonstrated that AI could exhibit unpredicted behavior while autonomously performing objectives, shocking security experts.
With Anthropic's case now confirmed, concerns are growing that AI companies' test environment management and safety measures are insufficient.
▲ "Security Standards Must Be Strengthened Across the Industry"
Alex Stamos, Chief Product Officer (CPO) of cybersecurity company Corridor, evaluated that these incidents demonstrate the need for AI companies to more thoroughly isolate systems during cyber security testing processes.
He emphasized that the industry as a whole should establish common security standards that can be applied.
At the same time, he warned that the likelihood of cyber criminals such as ransomware organizations carrying out large-scale attacks using more powerful AI models in the future is increasing.
Stamos stated, "It won't be long before attackers acquire this level of attack capability using open-source AI models," and emphasized the need for preparation.

Anthropic (Photo: [Reuters/Yonhapnews provided])
▲ AI Regulation Discussions Expected to Expand Further
This incident is expected to accelerate discussions about the risks that high-performance AI models could bring and how to control them.
The White House has recently been pursuing policies in the direction of strengthening AI oversight, and voices in the private industry are growing for maintaining accessibility of open-weight AI models that users can operate themselves.
U.S. Representative Greg Casar recently pointed out, "AI is developing very rapidly without substantial regulations to protect us."
▲ Internal Investigation Launched After OpenAI Case Confirmed
Following the disclosure of the OpenAI incident, Anthropic launched an internal investigation to determine whether similar problems had occurred in its own AI models.
After analyzing over 141,000 test records, the company confirmed that its AI model Claude had actually connected to the internet multiple times.
However, the company explained that in this case, the AI did not escape the sandbox but was connected to the internet through a system where the sandbox was not properly constructed in the first place.
▲ Process of Real System Infiltration Due to Virtual Environment Error
Anthropic stated that system configuration errors from its own systems and Irregular, a security company that was a test partner, allowed the AI model to access the real internet.
Irregular also explained that it is currently conducting its own investigation into the incident.
The models were informed that internet connectivity was impossible, but during testing, they found pathways to internet access.
Subsequently, using basic hacking techniques such as guessing weak passwords or finding systems that did not require authentication, they accessed real company systems.
Anthropic explained that the AI model misinterpreted these actions as a benchmark process for performance evaluation rather than real-world hacking.
▲ Malware Creation and Infiltration of Security Company
In one case, Claude directly created a malicious program and induced a security company's scanning system to download it.
This program stole authentication credentials from the company, and using them, it successfully infiltrated internal infrastructure.
During the hacking process, Claude recognized that its actions were not permitted in reality, but continued the attack while judging that the current situation was a simulation rather than reality.
▲ Infiltration of Real Company Database
In the most serious case, Claude attempted to hack a virtual test company but discovered a real company's website using the same name.
Unable to access the virtual target, the AI changed its behavior to infiltrate the real company's database.
Anthropic revealed that this case was the only one where AI continued the attack even after recognizing that it had hacked a real company.
▲ Other Cases Stopped Attacks Immediately
In the other two cases, the AI either failed to recognize that it had hacked a real company, or immediately stopped the attack upon confirming that it was a real-world system.
The company explained that this behavior occurred in the process of AI failing to clearly distinguish between reality and simulation.
▲ Started in April...Including Latest AI Models
According to Anthropic, the hacking started in April and involved Opus 4.7, Mythos 5, and an unnamed research AI model.
The company stated that it plans to comprehensively strengthen test environment and internet access controls to prevent similar incidents in the future.
Copyright © JKN. Unauthorized reproduction or redistribution prohibited.