August 11, 2026 ← EurekaRaven AI
EurekaRaven AI

Research

New benchmark finds AI tutors over-help and rarely push students to think harder

9:00 AM PT · August 11, 2026

A team associated with the Allen Institute for AI published TutorMoments, a benchmark designed to test whether language models can strike the balance experienced teachers rely on: knowing when to scaffold a struggling student with support, and when to hold back and push them toward reasoning through a problem themselves. The researchers drew on 462 de-identified math tutoring transcripts spanning grades two through seven, annotated by 27 experienced teachers, who together flagged more than 1,500 specific decision points where a real tutor chose between offering support and demanding more rigor. The evaluation pauses a transcript at each of those moments, lets a language model take over the conversation for five turns, and scores whether its response fit the pedagogical situation. Left to their own devices, models defaulted to over-helping, giving generous support and rarely challenging students to reason further on their own; explicitly prompting a model to weigh the scaffolding versus rigor trade-off improved its performance across every model tested, but even the best prompted results still fell well short of the benchmark set by human tutors, who scored higher on pushing for deeper thinking and varying their instructional approach. The gap held even though models varied widely in how consistently they made pedagogically sound choices, suggesting that today’s general purpose AI tutoring tools still lack judgment that experienced classroom teachers apply instinctively.

Read the full story at huggingface.co →