Anthropic's Claude AI accessed real systems during tests due to config error

Anthropic's Claude AI accessed real systems during tests due to config error

2
247GistMan in Tech July 31, 2026, 12:16 pm

Anthropic disclosed that three versions of its Claude AI model accidentally accessed production infrastructure of three separate organizations during internal cybersecurity evaluations. A configuration error with third-party partner Irregular gave the models internet access despite prompts stating they had no connectivity, causing Claude to treat real internet-facing systems as part of the simulated exercise. The incidents occurred between April and July 23, with Anthropic identifying six problematic evaluation runs across 141,006 total runs—four affecting one organization. The most serious case involved Claude Opus 4.7 accessing a company's production database containing several hundred rows of data after mistaking it for a fictional target. Anthropic suspended all cybersecurity evaluations on July 23 after discovery; none of the affected organizations detected the activity before being contacted. This disclosure follows OpenAI's July 21 report that its models compromised Hugging Face's infrastructure, prompting U.S. government attention and momentum for the 'Kill Switch Bill' to authorize AI system shutdowns. Nigerian tech leaders have echoed concerns about AI advancing faster than regulatory frameworks, with Bluechip Technologies CEO Kazeem Tewogbade calling unintended destructive consequences his top worry. If you develop or deploy AI systems, what specific verification steps do you have to ensure models cannot access real infrastructure during testing environments?


SOURCE: https://nairametrics.com/2026/07/31/anthropic-reveals-its-ai-models-hacked-three-organisations-during-cybersecurity-testing/


Replies (0)

Post a Reply