AI NewsProductivityReported
Uber Eats cut its search response time by 50 percent by doing less work and removing waits
Uber Eats rewrote its search pipeline across retrieval, feature fetching, ranking, ads and rendering, and reports a 50 percent reduction in end-to-end latency and a 50 percent reduction in p99 latency on an early product-based test, with a per-component ledger of where the milliseconds came from.

Image: InfoQ
Why it mattersAn engineering team sitting on years of search architecture decisions now has a step-by-step account of how one of the largest marketplaces found 500 milliseconds without faster hardware, which turns "we need to speed this up" into a plan with named line items.
The quickest way to speed up a search page is often to stop computing things nobody will see.
Uber Eats published a search-pipeline postmortem on its engineering blog on 10 September 2026, and InfoQ wrote it up on 2 October. The post is signed by five Uber engineers, led by Distinguished Engineer Nimish Sheth. It describes a rewrite across the whole pipeline, retrieval, feature fetching, ranking, ads and rendering, and reports a 50 percent reduction in end-to-end search latency, with a ledger that assigns milliseconds to the specific change that saved them.
Where the time went
The authors group the savings by pipeline stage. Shifting the primary metric from API response time to "Above-the-Fold" completion, and moving to asynchronous HTML template rendering with pagination, took more than 200 milliseconds off the time users actually see. Retrieval work came down by 120 milliseconds after low-value strategies were cut and product-level embeddings were introduced. Splitting ranking data from presentation data saved more than 100 milliseconds. Removing false dependencies and parallelising store ranking saved 55 milliseconds. Request hedging on hydration took another 40 milliseconds. Rebuilding the ads data path, from row-oriented to column-oriented bid data held in memory, saved around 140 milliseconds. Smaller changes to parallel encoding, embedding size and network connection management added up to roughly 150 milliseconds more.
On a separate test using product-based search catalogues instead of the older store-centric path, Uber reports a 50 percent reduction in p99 latency, with another 30 to 50 milliseconds expected from the next milestone.
The idea behind the rewrite
InfoQ attributes the central idea to Anubhooti Nagar at Uber: the gains came from doing less work and from cutting out unneeded waits, rather than from running the same work faster. The engineering authors frame the same point from the architectural side, writing that the pipeline had accumulated years of design decisions that each made sense when they were made, and compounded into waste when they ran together.
That is why the rewrite touches every stage rather than one place. Fetching features in parallel is pointless if the ranker still waits for presentation data it does not need to score. Faster retrieval does not help if the renderer blocks on everything at once instead of streaming the first screen.
The method the item points at
The post names its underlying systems: Apache Lucene for retrieval, Spark-based indexing, Kafka streaming updates, and a distributed serving layer. It does not claim a new technique. Request hedging, column-oriented data layouts, asynchronous rendering and parallel encoding are all textbook. What the post offers that is useful outside Uber is the accounting: how a team found each tenth of a second, what measurement convinced them it was worth chasing, and in which order the changes landed.
Two caveats are worth stating. The 50 percent end-to-end headline is Uber's own, measured on Uber's own traffic, and the component numbers do not add up to the headline because they come from different experiments. And the p99 result for product-based catalogues came out of an early test, not the whole user base.
Source
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


