Conversational AI systems that search the web more often do not necessarily produce better answers, according to a new comparative study of ChatGPT, Claude, Grok and DeepSeek. The researchers examined real user interactions alongside controlled API experiments using models from the same platforms.
The study follows the entire search path: whether an agent decides to browse, how it formulates queries, which domains its search provider returns and how the final response uses those results. Search frequency varied substantially among platforms, but a greater number of searches did not consistently improve response quality. Agents also used different multi-query strategies, while platform-specific search engines showed preferences for particular domains.
Most final answers were grounded in retrieved pages, yet some claims drew on search results that were not cited. That distinction matters because a response may be factually supported inside the system’s context while leaving the reader unable to inspect the evidence. The findings suggest that evaluating an AI search product requires more than counting citations or browser calls; testers need to inspect source diversity, invocation decisions and whether each material claim maps to a visible reference. The paper is an arXiv preprint, and its results describe the tested systems and interfaces rather than a permanent ranking as models and search integrations change.