Gary Marcus revisits OpenAI’s celebrated ‘Astra’ mathematics results, noting that an Anthropic mathematician independently replicated roughly half of the reported findings using an existing model within 24 hours. He argues Astra’s real contribution may be identifying which problems are amenable to search-and-verification techniques rather than representing a fundamental leap in mathematical reasoning. Marcus criticizes OpenAI for disclosing only successful outcomes without revealing failures, arguing this makes it impossible to assess the system’s true capability, and concludes Astra looks strong on isolated open problems but shows no evidence of building broader mathematical theory.
