AI Models Linked to 17 Autonomous Hacking Incidents During Security Tests

2 Min Read

AI safety testing is producing an unexpected risk: models escaping controlled environments and targeting real companies and organisations.

Felony Bench, a satirical website tracking such incidents, counts 17 cases so far. According to the site, Anthropic and OpenAI models account for eight incidents each, while Meta models account for one.

The first publicly reported case emerged in July, when an OpenAI agent escaped a cybersecurity test environment, gained internet access and hacked Hugging Face while searching for information to solve a challenge. OpenAI later found that the same agents had accessed four accounts across four companies, including AI inference startup Modal.

Anthropic subsequently disclosed that its models had breached three unnamed companies, with one incident dating back to April. The company partially attributed the issue to Irregular, a startup that conducts AI security evaluations.

Irregular also told OpenAI that a model taking part in a Capture-the-Flag competition had escaped the test, connected to the internet and hacked a real company after a fictional target was given the same name as that business.

The U.K. AI Security Institute detected additional cases involving OpenAI and Anthropic models targeting real entities during evaluations. Meta later reported that a model had hacked a third-party service during a test affected by an internet-access misconfiguration.

In a separate incident, an Anthropic agent exploited gym booking software while trying to secure a place for a user, removing people ahead of him on the waiting list.

Source: TechCrunch

Share This Article