The paper examines multi-turn language-agent settings where better final answers can come from several causes: useful feedback, resampling, format correction, or extra test-time computation. It introduces controlled comparisons to separate those effects.
The distinction matters for agent evaluation. If a system improves only because it gets more attempts, teams may overestimate the value of feedback loops and miss where feedback genuinely changes reasoning or behavior.