Spotify researchers have developed a training pipeline for conversational recommendation agents that can improve tool planning before a product has accumulated real user conversations. The work targets requests that require several steps, such as finding Italian indie artists a listener has not heard before, rather than simple one-shot searches.
The pipeline turns single-turn prompts into realistic multi-turn exchanges so teams can test an agent systematically during the cold-start period. A separate self-improvement loop compares variable outcomes, identifies planning and tool-use errors, and uses a coding agent to refine the instructions. The researchers report an 8% quality gain on top of an already optimized manual prompt.
Spotify says the system has been put into production and shortened the iteration cycle for launching a conversational recommendation agent. The paper also reports online A/B tests, but the result should be read as evidence from one company’s recommendation stack rather than proof that the method will transfer unchanged to other products. Its practical contribution is a way to exercise and repair an agent’s decision process before live interactions provide enough training evidence.