AI NewsProductivityAnnouncement
An engineer at a retail company says shared LLM platform cut hallucinations from 15 percent to 1.5 percent in six months
Aditya Mulik, writing on InfoQ, reports that hallucinations on a retail inventory recommendation system fell from 15 percent of agent responses to 1.5 percent over six months, by moving retry logic, schema enforcement and tool authorisation out of each application and into a shared platform layer.

Image: InfoQ
Why it mattersSix platform primitives a team can build themselves moved the hallucination rate ten times further than a model swap would have, and the author published the full reference code so a second team can measure the same shift against its own baseline.
A retail company engineer named Aditya Mulik, writing in InfoQ today, says a multi-agent LLM system he worked on reduced its hallucination rate from roughly 15 percent of agent responses in its first month of production to 1.5 percent six months later, by moving six cross-cutting concerns out of each application and into a shared platform layer. The system is an inventory accuracy platform at a large retail organisation that routes through foundation models like Gemini and GPT, serves dozens of application teams, and analyses "discrepancy signals across millions of SKUs and billions of historical inventory records."
The author's own description of the change is direct. The lever that moved the number, in his words, was the decision to stop treating the LLM stack as an application concern and start treating it as platform infrastructure. The foundation model was the same model through the whole period.
Three failures that forced the platform
The article names three failures the team hit during beta that drove the shift. The first was API throttling under production traffic. The second was context. A bad recommendation almost always traced back to the data fed into the model rather than to the model's own choice. The third was hallucinations, which stayed individually rare but became a steady rate once production traffic was flowing.
The author identifies three properties that pushed those failures down into platform concerns: "A hallucinated response isn't a stack trace. It passes every traditional health check," there is no way to attribute token spend without per-request measurement, and latency dashboards show green while a model silently drifts into lower-quality but syntactically valid output.
What the platform actually does
The architecture is a single gateway, a root coordinator agent built on Google's Agent Development Kit that classifies intent on every turn, and specialist agents that each hold their own prompt and their own output schema. Tools live behind the Model Context Protocol in multiple independent MCP servers, each of which verifies the caller's token itself before any tool runs.
The author says the largest single piece of the 15 to 1.5 percent reduction came from the retry layer. The pattern is to classify each failure as a schema violation, a hallucination signal, or an infrastructure error, and route the three differently. Schema violations re-prompt with the error appended. Hallucination signals re-prompt with stronger grounding context. Infrastructure errors back off exponentially. Every request carries a hard retry cap, three in the author's case.
The grounded flag, and why it costs nothing
Every specialist output has a Pydantic schema with two scoring fields the author highlights: a grounded boolean that the model may set to true only when every claim is backed by what the tools returned, and an out_of_scope flag for a question outside its remit. The author writes: "A response that fails validation, has the wrong shape, missing fields, or unparseable output is the highest-precision hallucination signal a platform has."
The companion GitHub repository at adityamulik/llm-platform-primitives publishes the full reference implementation under Apache 2.0, with "no hidden abstractions." The repo runs a local Llama 3.1 model through LiteLLM instead of a hosted foundation model, so the whole stack runs end to end at zero API cost on a laptop. The 15 to 1.5 percent figure is one engineer's account of one platform at one retailer, and the author says so in the opening paragraph of the article.
Source
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


