Ai2 has introduced TutorMoments, a benchmark for a subtle education problem: an AI tutor can be helpful and still teach badly if it gives away too much too soon. The framework replays real one-on-one math tutoring sessions and stops at moments where an experienced teacher had to choose between offering help and pushing the student to reason further.

At each decision point, a language model takes over as the tutor in a simulated session, with another model playing the student. The goal is not just to answer correctly, but to judge whether the tutor supports learning without removing the productive struggle.

Ai2 says early tests show current models tend to over-help when simply told to tutor well. Prompts that name the help-versus-hold-back trade-off improve behavior, but do not fully solve it. That makes TutorMoments useful less as a product claim than as a practical test for whether classroom AI systems can handle timing, not just knowledge.