Uber says it has cut end-to-end search latency in Uber Eats by 50% after redesigning work across retrieval, ranking, advertising, presentation and infrastructure. The team measured how long it took to render the first visible screen, rather than stopping the clock when the backend returned a response.

Pagination with server-side caching and asynchronous rendering improved that above-the-fold measure by more than 200 milliseconds. Engineers also found that the system enriched tens of thousands of candidates before ranking and then discarded many of them. Removing low-value retrieval paths saved about 120 milliseconds, while product-level embeddings reduced certain lookups more than 100-fold and saved another 50 milliseconds.

Separating ranking data from presentation data removed more than 100 milliseconds. Changes to dependency handling, request hedging and advertising storage produced further gains. Uber also used an agentic coding workflow to identify, benchmark and validate some optimizations, but the reported result reflects changes across the whole stack rather than AI coding alone.

The company is now testing product-based retrieval, microbatching, zero-pass ranking and HTTP multipart streaming so pipeline stages can overlap. Early product-search tests reportedly reduced 99th-percentile latency by more than 50%, though those experiments are not yet a promise of identical production results.