CONTINUE TO SITE »
or wait 15 seconds

AI

How to keep your retail AI pilot out of the graveyard

To escape GenAI pilot purgatory, companies have to look at the four critical failure points where enterprises are getting the approach entirely wrong and what it actually takes to fix them.

Adobe Stock

September 3, 2026 by Sathya Narayanan Annamalai Geetha — Principal Architect, Google

Almost every retailer I talk to has an AI pilot initiative. A recommendation engine that lifted conversion in one specific area. A computer-vision demand-forecasting model that is poised to cut waste in a bunch of distribution centers. A chatbot that handled returns for a brand line without falling over.

But only a handful of them have scaled the pilots to production or have gotten the second one to stick, or the tenth. That's the part nobody puts in the case study. The MIT study titled "The GenAI Divide" finds that a staggering 95% of GenAI pilots are failing to deliver a measurable operational or P&L impact to the business.

The pilot is lying to you, just a little

A pilot or a proof-of-concept succeeds under conditions that don't exist anywhere else in the business. One team owns it, one data source feeds it and a store, in one region and one product category absorbs the mess. Everyone involved cares enough to babysit it through the rough patches.

Scale removes all of that protection at once. Now the model has to run against inventory data from three different systems of record, half of which update on different schedules and none of which agree on what 'in stock' means.

Now it has to survive a peak sales event without anyone standing by. What also tends to happen is that the team that built it has moved on to the next pilot, and the team that inherited it doesn't know why a specific threshold was set the way it was.

Almost none of that is caused by the frontier models; rather it's an operational problem. And enterprises that attempt scaling without ensuring the underlying architecture survives contact with reality are the ones who end up with a graveyard of proofs-of-concept. The pilot works but the scale-up stalls and the reason is almost never the model.

Where it actually breaks

I've spent a significant amount of time architecting and scaling enterprise-grade AI deployments, and the pattern is remarkably consistent.

To escape GenAI pilot purgatory, we have to look at the four critical failure points where enterprises are getting the approach entirely wrong — and what it actually takes to fix them.

1. Chasing the shiny object (technology over business value)
The fastest way to fail with AI is to start with a technology-first approach instead of a business problem. Right now, boardrooms are suffering from severe 'shiny object syndrome.' Executives see a powerful new Large Language Model (LLM) release and mandate their teams to "put GenAI in our app."
The result is a solution in search of a problem. When teams focus heavily on the AI buzz rather than core business friction, they build novelties; like a conversational sommelier for a grocery app that customers use once for fun and then abandon. Profound enterprise AI doesn't start with a model rather it starts with an operational bottleneck. If you aren't pointing the technology at a high-value, deeply painful business problem - like supply chain blind spots, massive customer service call volumes, or associate onboarding times - your pilot will never secure the operational buy-in required to scale.

2. The data void and the hallucination trap
The most dangerous assumption retailers make about Generative AI is treating it like a traditional data system. It is not. Generative AI is inherently non-deterministic, it predicts the next most likely probability of something based on its training, it does not "look up" facts. Because of this non-deterministic nature, if an LLM is not relentlessly grounded in your enterprise data, it will hallucinate. It will confidently offer a customer a discount that doesn't exist, or tell a store associate that an out-of-stock item is in aisle four. In retail, an AI hallucination isn't a funny quirk; it's a blown transaction and a permanent hit to brand trust. Retail data is notoriously fragmented across legacy point-of-sale, e-commerce, and supply chain systems. When an ungrounded GenAI pilot scales into this messy reality without a robust data foundation, it collapses into unreliability - causing more harm than good.

3. Treating integration as a "Phase 2" feature
Most Generative AI projects are built in a sandbox. They are launched as standalone projects with a slick UI, but absolutely zero backend integration into the systems that actually run the business.
An AI associate-assist tool is useless if it can't securely read the CRM to know the customer's loyalty tier, write to the ERP to update inventory, or trigger a return authorization in the order management system. When pilots exist as disconnected islands, they force employees and customers to swivel-chair between the AI and the actual systems of record. This friction inevitably leads to a complete lack of adoption. Integration is not a "Phase 2" feature; it is the entire ballgame.

4. The LLMOps and dynamic governance imperative
Governance designed for traditional software moves at a glacial pace. It assumes deterministic outputs: if I click "X", the system does "Y".
Generative AI broke that assumption. Because the model's outputs are non-deterministic and dynamic based on user prompts, you cannot simply QA test it once and deploy it. Most AI pilots are failing to scale because they lack AI governance and LLMOps (Large Language Model Operations). The same exact input fed to the same exact model can result in wildly different outputs. That is the whole idea of Generative AI. Now add to it the ever changing model landscape. It's chaos. Most latest and powerful models are not necessarily the most impactful to the use case in hand. Without a robust LLMOps, you have no way to monitor the model in real-time for data drift, biased outputs, toxicity, or prompt injection attacks. A pilot can survive with human babysitters; a production system across a 2,000-store network will become a massive liability if you do not have automated monitoring and active governance running 24/7.

What actually changes the outcome

  • The retailers who get past pilot purgatory tend to do a few unglamorous things early on, before they're forced to.
  • They invest in the data layer before they invest in the next model, because every future Generative AI capability depends entirely on it.
  • They architect for deep back-end integration from day one, treating AI as an active orchestration layer, because a model that cannot securely read and write to existing systems of record is destined to become an abandoned novelty.
  • They build governance as a fast, living process rather than a static document, with a real owner accountable for keeping it current as new models evolve.
  • Lastly, they staff for the middle of the AI lifecycle, i.e. the monitoring, retraining, and incident response, and not just the launch party.

None of this is as exciting as the pilot. There's no press release or announcements for 'we made our data pipelines consistent' or 'we can now move this workload between environments without a rebuild.' But it's the difference between a retailer with one good case study and a retailer with AI genuinely embedded across how it operates. To be honest, the proof-of-concept was never the hard part, it was just the part that photographs well.

The real work (and the real competitive advantage) is everything retailers have to get right after the pilot succeeds and before anyone notices.

About Sathya Narayanan Annamalai Geetha

Sathya AG is a Principal Architect with over 2 decades of experience specializing in cloud-native data and AI solutions for the retail industry. He is the author of "Enterprise-Grade Hybrid and Multi-Cloud Strategies", advisory board member for the CAIO Circle and a Fellow of the British Computer Society.

Connect with Sathya Narayanan:





©2026 Connect Media, All rights reserved.
b'S2-NEW'