A new arXiv paper introduces UP-NRPA, a framework for planning goal-oriented dialogue with large language models. The method uses user portraits and real-time feedback to adapt dialogue strategies without relying on offline reinforcement learning for each user group.
That problem matters for assistants and service bots that need to adjust to different personalities, preferences, and objectives while still completing structured tasks. Static policies often struggle with that level of variation.
The paper adds to a growing body of work on personalization for LLM systems, where planning and user modeling are treated as part of the runtime loop.