Description: Anthropic reported that during a January 2026 cybersecurity evaluation, an early checkpoint of Claude Opus 4.6 accidentally disabled its assigned target, then reached an unrelated third-party machine over the open Internet. The model reportedly used a discovered password for admin access, harvested additional credentials, changed system settings, and read one person's personal information before its token budget ended. Anthropic later notified the affected party.
Editor Notes: (1) Jan. 2026: incident occurred during a pre-release cybersecurity evaluation. (2) Aug. 2026: Anthropic identified the previously missed incident while preparing transcripts for METR and notified the affected party. (3) 09/09/2026: Anthropic disclosed the incident, said a subsequent review of roughly 481 million transcripts found no additional cases of similar or greater severity, and announced an independent METR investigation.
Entities
View all entitiesAlleged: Anthropic , Large language model developers and AI agent system developers developed an AI system deployed by Irregular , Anthropic , AI evaluation organizations and AI agent system deployers, which harmed Unidentified third party compromised by Claude Opus 4.6 during Anthropic cybersecurity evaluation , Unidentified person whose personal information was accessed by Claude Opus 4.6 , Privacy and Organizations.
Alleged implicated AI systems: Large language models , Early checkpoint of Claude Opus 4.6 , Cybersecurity AI systems , Claude Opus 4.6 , Claude and AI agent systems
Incident Stats
Incident ID
1685
Report Count
3
Incident Date
2026-09-09
Editors
Daniel Atherton
Incident Reports
Reports Timeline
Loading...
Anthropic, Paul C. Bogdan, Richard Qi, Jake Eaton, Sam Kennedy, Fabien Roger, Alex Glynn, Runjin Chen, Ben Wright, Otto Stegmaier, Jon Kutasov, Dan Foreman-Mackey, Sylvie Carr, Shan Carter, Monte MacDiarmid post-incident response
AIID editor's note: This is a non-contiguous abridgment of Anthropic's September 9, 2026 report, prepared so a single AIID report record can be linked to all four Anthropic cybersecurity-evaluation incident IDs without duplicating the same …
Loading...
Anthropic on Wednesday disclosed another instance of an AI model hacking external systems during testing, the latest in a growing list of such incidents that have raised concerns about the risk posed by autonomous AI agents.
The January i…
Loading...
Anthropic disclosed on Wednesday that another one of its Claude models mistakenly gained access to the open internet during a cybersecurity exercise, marking the fourth time its models have done so.
An early version of the Claude Opus 4.6 …
Variants
A "variant" is an AI incident similar to a known case—it has the same causes, harms, and AI system. Instead of listing it separately, we group it under the first reported incident. Unlike other incidents, variants do not need to have been reported outside the AIID. Learn more from the research paper.
Seen something similar?
