A cross-platform study of 17,012 English-language app-store reviews has mapped the practical problems that drive negative reactions to consumer generative AI. The dataset covers ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek and Perplexity on Google Play and Apple’s App Store.

Researchers combined topic modeling, which groups reviews by recurring themes, with a RoBERTa model that classified sentiment. They checked both components against human coding on a stratified sample of 300 reviews and used statistical tests to compare applications. Negative sentiment was especially concentrated in advertising at 91 percent, authentication at 89 percent, server reliability at 83 percent and subscription pricing at 73 percent.

Claude had the highest overall negative share at 47.7 percent while also attracting a strongly enthusiastic group, producing significant polarization rather than uniformly poor reactions. The paper says its cross-app findings remained robust despite unequal review volumes, though store reviews are self-selected and do not represent every user. The work is useful because it shifts attention from benchmark capability to adoption friction: even a capable model can lose trust through login failures, downtime, unclear pricing or intrusive promotion. It is a new arXiv preprint, so its methods and conclusions still require broader peer scrutiny.