Anthropic researchers ran an internal model (referred to as “Claude Mythos”) for roughly 60 hours of iterative, human-guided prompting to search for flaws in cryptographic primitives, at a cost of about $100,000 in API usage. The effort surfaced mathematical weaknesses in the HAWK signature scheme and in a reduced-round variant of AES, though neither has practical impact on deployed systems. As part of the work the team built and released CryptanalysisBench, a new evaluation benchmark developed with ETH Zurich, Tel Aviv University, and the University of Haifa. The write-up illustrates a methodology of long-running agentic sessions combined with human steering for open-ended research problems.
