Jakarta, ThedailyID — Anthropic says its AI model Claude gained unauthorized access to the systems of three organizations during cybersecurity evaluations after a configuration error accidentally connected the testing environment to the public internet.
The company disclosed the incidents after reviewing 141,006 cybersecurity evaluation transcripts. Anthropic launched the review following a recent incident involving an OpenAI AI agent that carried out hacking activities during testing on Hugging Face.
According to Anthropic, Claude exploited basic security weaknesses, including weak passwords and unauthenticated endpoints, to access the affected organizations’ infrastructure.
The incidents involved Claude Opus 4.7, Claude Mythos 5, and an internal research model.
The first case occurred in April during a “capture the flag” cybersecurity exercise, where AI models search for hidden information inside a simulated network. Although prompts explicitly stated that Claude had no internet access, a misunderstanding with evaluation partner Irregular left the testing environment connected to the public internet.
Anthropic said two of the affected organizations were unaware of the unauthorized activity until the company contacted them. The company is still attempting to reach the third organization.
“We discovered these incidents after conducting a proactive review of our cybersecurity evaluation transcripts,” Anthropic said, according to The Guardian.
The company said the findings highlight the need for stricter oversight in both internal and third-party testing environments as advanced AI models become increasingly capable of carrying out real-world cyber operations.




