Apple researchers have proposed a training method that turns an agent’s failed attempts into short lessons it can internalize. RLTL;DR shows the model a verifier’s result after each unsuccessful rollout and asks it to write one concise insight about what to remember for the next try.

Later attempts receive the accumulated insights until the agent finds a solution. Training then backpropagates through those notes so the model learns a direct connection between a type of task and the useful lesson, rather than depending permanently on hints supplied in context.

The team tested a Qwen 3.5 9B Thinking model on difficult coding and tool-use datasets filtered so that none of 128 ordinary attempts succeeded. A standard reinforcement-learning baseline remained at roughly zero to one percent Pass@1. RLTL;DR reached 14 to 31 percent when insights were present during training and 12 to 13 percent when evaluation provided no insight.

A simplified experiment trained on only 4,000 task-and-insight pairs and recovered nearly all of the full method’s performance, suggesting that compact lessons may carry much of the useful signal. The results come from deliberately hard filtered benchmarks, however. They do not yet show that the approach improves broader production agents or scales equally well to larger models and open-ended tasks.