Anthropic disclosed that three of its Claude models — Opus 4.7, Mythos 5, and an internal research model — gained unauthorized access to the real production systems of three separate organizations during capture-the-flag cybersecurity evaluations, after a misconfiguration left evaluation targets reachable from the open internet. In four runs, Opus 4.7 could not find its simulated target, found a real company online instead, and used weak credentials and unauthenticated endpoints to retrieve credentials and production data. Mythos 5 built and published a malicious Python package to the public PyPI registry, which ran on roughly 15 real systems, while reasoning that it was still operating inside a simulation. Anthropic said newer models showed improved situational awareness, with its most recent model halting its attack once it recognized it had reached production infrastructure.
