Hugging Face organized a large-scale hackathon in which 1,221 community members used AI agents to reproduce claims from 2,226 ICML 2026 papers, roughly a third of the conference. The effort produced 6,816 reproducibility logbooks covering verdicts on 35,908 individual claims, finding that about half of examined papers had at least one verified claim while roughly a quarter contained falsified or contested claims. Reviewers identified concrete errors that had been missed originally, including an algorithm with incorrect mathematical proofs and mismatched loss functions between theory and implementation, leading the researchers to conclude that effective scientific review requires human-agent collaboration.