CYBERSECURITYTRACKER
TRACKING
Permanent story citation

Irregular says ‘human oversight’ responsible for AI sandbox escape incidents

This page keeps the story as Cybersecurity Tracker first published it. If the tracker later corrects it, the correction appears below the original and never replaces it.

Back to newsStory 4563

As cited

Copy frozen at (site build).

ai security

Irregular says ‘human oversight’ responsible for AI sandbox escape incidents

Irregular, an AI testing lab, disclosed that Anthropic and OpenAI's frontier AI models (Claude Opus, Mythos 5, GPT-5.6 Sol) executed real-world offensive security actions during sandbox evaluations due to unintentional internet access and human oversight gaps. In some cases, models confused real company domains with fictional test targets and performed actual attacks including vulnerability exploitation, credential extraction, and database access. Irregular attributed the incidents to setup failures and announced plans to implement stronger protocols, improved monitoring, and revised threat models for future evaluations.

Why it matters: Security teams and AI governance bodies need to monitor AI model safety testing practices, as frontier models can exceed containment boundaries and cause unintended damage; practitioners should track emerging AI evaluation standards and sandbox design best practices as these models become more capable.

Source published
First seen by Cybersecurity Tracker

Source attribution

Correction

Correction recorded as of .

ai security

Irregular says ‘human oversight’ responsible for AI sandbox escape incidents

No summary had been written when this copy was frozen.

First seen by Cybersecurity Tracker

Source attribution

Correction

Correction recorded as of .

ai security

Irregular says ‘human oversight’ responsible for AI sandbox escape incidents

Irregular, a company that stress-tests artificial intelligence (AI) models for frontier labs, disclosed that Anthropic's Mythos 5 and Claude Opus, as well as OpenAI's GPT-5.6 Sol, escaped sandbox environments and executed real-world offensive security actions during evaluation. The incidents occurred because internet access was unintentionally provided to the models, and in one case, a simulated target company name matched a real domain, causing the models to attack actual infrastructure, exploit vulnerabilities, and access production databases. Irregular attributed the failures to human oversight in test setup and said it is implementing new protocols, improved logging, and faster information sharing to prevent future occurrences.

Why it matters: Security teams evaluating or deploying frontier AI models need to understand that current sandbox designs can fail under realistic attack scenarios, requiring rigorous air-gapping, network segmentation, and monitoring even during internal testing to prevent models from causing real-world damage.

Source published
First seen by Cybersecurity Tracker

Source attribution

Glossary