Anthropic published its August 2026 Risk Report, raising its assessment of catastrophic misalignment risk in high-stakes settings from “very low” to “low,” citing added uncertainty following recent cybersecurity evaluation disclosures. The report also disclosed an unreleased internal model, Model 2, which the company says is somewhat more capable than its frontier Mythos 5 but has no current plans for external release. Anthropic said testing found no new or more concerning form of misalignment in Model 2 than behavior already documented for Mythos 5.
