Researchers at the Allen Institute for AI introduced TutorMoments, a framework that pauses real one-on-one math tutoring transcripts at critical decision points and asks language models to continue the session. The dataset contains 462 de-identified, text-only transcripts with more than 1,500 annotated moments where human tutors chose between offering scaffolding or pushing students toward independent reasoning. Testing seven models found that generic “tutor well” instructions cause models to over-help, while explicit prompting about the scaffolding-versus-rigor trade-off improves performance, though with wide variation across models. The authors conclude a substantial gap remains between LLM tutoring behavior and human educator judgment at these decision points.
