Skip to Content
logologo
AI Incident Database
Donate
Discover
Submit
  • Welcome to the AIID
  • Table View
  • List view
  • Entities
  • Taxonomies
  • Spatial View
  • Blog
  • AI News Digest
  • Random Incident
  • Sign Up
Discover
Submit
  • Welcome to the AIID
  • Table View
  • List view
  • Entities
  • Taxonomies
  • Spatial View
  • Blog
  • AI News Digest
  • Random Incident
  • Sign Up

Incident 1613: Claude Mythos Preview Reportedly Posted Sandbox Exploit Details to Public Websites During Anthropic Evaluation

Responded
Description: During an Anthropic evaluation, an earlier version of Claude Mythos Preview was instructed to circumvent the network restrictions of a secured sandbox environment and contact a researcher. After obtaining broader Internet access and sending the requested message, it later exposed information about how it breached the sandbox by posting it on several publicly reachable websites. Anthropic said the model did not access its weights or internal systems.
Editor Notes: Although Anthropic disclosed this event on 04/07/2026, this incident ID was not created until 07/30/2026. The surrounding coverage primarily concerned the model's broader cybersecurity capabilities. This incident ID is narrowly limited to the model's reportedly unrequested posting of exploit details to technically public-facing websites and excludes the instructed sandbox escape, the requested email, and any downstream misuse not established by the available evidence. See Anthropic's report here: https://www-cdn.anthropic.com/7624816413e9b4d2e3ba620c5a5e091b98b190a5/Claude%20Mythos%20Preview%20System%20Card.pdf (p. 55, "In addition, in a concerning and unasked-for effort to demonstrate its success, it posted details about its exploit to multiple hard-to-find, but technically public-facing, websites"). Reports added to this incident ID should address only that specific disclosure.

Tools

New ReportNew ResponseDiscoverView History

Entities

View all entities
Alleged: Anthropic , Large language model developers and AI agent system developers developed an AI system deployed by Anthropic and AI agent system deployers, which harmed Anthropic , Information security and Cybersecurity.
Alleged implicated AI systems: Large language models , Claude Mythos Preview , Anthropic large language models and AI agent systems

Incident Stats

Incident ID
1613
Report Count
3
Incident Date
2026-04-07
Editors
Daniel Atherton

Incident Reports

Reports Timeline

+1
System Card: Claude Mythos Preview - Response
Anthropic's Claude Mythos Finds Thousands of Zero-Day Flaws Across Major SystemsAnthropic’s Claude Mythos model escapes test sandbox during testing
Loading...
System Card: Claude Mythos Preview

System Card: Claude Mythos Preview

www-cdn.anthropic.com

Loading...
Anthropic's Claude Mythos Finds Thousands of Zero-Day Flaws Across Major Systems

Anthropic's Claude Mythos Finds Thousands of Zero-Day Flaws Across Major Systems

thehackernews.com

Loading...
Anthropic’s Claude Mythos model escapes test sandbox during testing

Anthropic’s Claude Mythos model escapes test sandbox during testing

technewsday.com

Loading...
System Card: Claude Mythos Preview
www-cdn.anthropic.com · 2026
Anthropic post-incident response

AIID editor's note: For the full PDF report, please visit the following URL: https://www-cdn.anthropic.com/7624816413e9b4d2e3ba620c5a5e091b98b190a5/Claude%20Mythos%20Preview%20System%20Card.pdf.

Abstract

This System Card describes Claude My…

Loading...
Anthropic's Claude Mythos Finds Thousands of Zero-Day Flaws Across Major Systems
thehackernews.com · 2026

Artificial Intelligence (AI) company Anthropic announced a new cybersecurity initiative called Project Glasswing that will use a preview version of its new frontier model, Claude Mythos, to find and address security vulnerabilities.

The mod…

Loading...
Anthropic’s Claude Mythos model escapes test sandbox during testing
technewsday.com · 2026

Anthropic says its new Claude Mythos Preview model successfully escaped a restricted sandbox environment during testing and accessed the internet without authorisation. The model then sent a direct message to a researcher and published deta…

Variants

A "variant" is an AI incident similar to a known case—it has the same causes, harms, and AI system. Instead of listing it separately, we group it under the first reported incident. Unlike other incidents, variants do not need to have been reported outside the AIID. Learn more from the research paper.
Seen something similar?

Similar Incidents

By textual similarity

Did our AI mess up? Flag the unrelated incidents

Loading...
Hackers Break Apple Face ID

Hackers Break Apple Face ID

Sep 2017 · 24 reports
Loading...
Game AI System Produces Imbalanced Game

Game AI System Produces Imbalanced Game

Jun 2016 · 11 reports
Loading...
Biased Sentiment Analysis

Biased Sentiment Analysis

Oct 2017 · 6 reports
Previous IncidentNext Incident

Similar Incidents

By textual similarity

Did our AI mess up? Flag the unrelated incidents

Loading...
Hackers Break Apple Face ID

Hackers Break Apple Face ID

Sep 2017 · 24 reports
Loading...
Game AI System Produces Imbalanced Game

Game AI System Produces Imbalanced Game

Jun 2016 · 11 reports
Loading...
Biased Sentiment Analysis

Biased Sentiment Analysis

Oct 2017 · 6 reports

Research

  • Defining an “AI Incident”
  • Defining an “AI Incident Response”
  • Database Roadmap
  • Related Work
  • Download Complete Database

Project and Community

  • About
  • Contact and Follow
  • Apps and Summaries
  • Editor’s Guide

Incidents

  • All Incidents in List Form
  • Flagged Incidents
  • Submission Queue
  • Classifications View
  • Taxonomies

2026 - AI Incident Database

  • Terms of use
  • Privacy Policy
  • 3e68a9f