The paper addresses the difficulty of auditing recommendation and personalization systems when researchers can only observe outputs from the outside. AI agents are used to simulate interactions and collect evidence across changing user histories.
Automated audits could make platform accountability studies broader and more repeatable, though they also raise questions about methodology, representativeness, and how platforms respond to synthetic users.