Retail and Consumer Goods organisations are investing heavily in generative and agentic AI. While the capability is scaling fast, the measurable evidence of value has not kept pace. Our analysis of public disclosures from 10 leading organisations (between 2020–2025) shows AI capability narratives have grown steeply while quantified value narratives have seen a slower growth, and the mechanisms connecting the two (adoption, trust, utilisation, decision behaviour) are almost never disclosed.
We call this the AI Value Paradox - Organisations are becoming increasingly sophisticated at deploying AI, while the evidence needed to demonstrate, attribute, and track business value has evolved more slowly. The next competitive advantage will come from those who develop the discipline to connect AI investment to measurable outcomes.
This paper presents the evidence, diagnoses the reasons for the gap, and offers a practical framework for effectively closing it.
Per our review of the publicly available disclosures from 10 leading Retail and CPG organisations we found that between 2020 and 2025, AI language migrated from technology appendices into CEO letters and capital-allocation frameworks. Additionally, AI mentions in annual disclosures grew roughly eightfold. However, capability claims - platforms built, pilots launched, grew far faster than quantified, AI-attributed outcomes. Capability mentions represent about 76% of AI disclosure, outcome mentions about 24%, and quantified value with a specific attributed metric under 5%. The capability-to-value ratio widened from 2.3:1 in 2020 to 3.4:1 by 2025. Of the disclosed quantified value, cost-efficiency gains are concentrated in inventory management in Retail and selling, general and administrative expenses (SG&A) in CPG.
Most organisations measure AI inputs (technology, platforms, use cases) and business outcomes (revenue, margin, working capital). Almost entirely absent from public disclosure are the mechanisms connecting the two - adoption, utilisation, recommendation acceptance, override behaviour, trust, and decision quality. Inputs account for ~73% of AI disclosure, outcomes ~23%, and mechanisms under 4%. A retailer whose forecast AI improves accuracy but is never wired into replenishment decisions has a mechanism gap, not a technology gap. The same pattern appears when field teams for a consumer goods company override AI recommendations most of the time. A handful of companies buck this trend - Walmart's agentic assistant, Sparky, pairs a real outcome (order values ~35% higher) with an adoption mechanism (~50% of app users engaged); Mondelez links sales rep effectiveness (~80%) to a 2–4% topline lift; Tesco's AI-driven forecasting and Clubcard personalisation contributed ~GBP 500 million in productivity savings alongside ~82% Clubcard penetration; Unilever lifted product availability to 98% in a pilot while cutting manual planning effort ~30%. These are exceptions—most disclosures describe capability, some report outcomes, and almost none describe the mechanisms in between.
One organisation in our sample delivered the strongest cost-efficiency metrics of any company studied while carrying the lowest AI narrative - its performance came from structural discipline, not AI. Across all ten organisations, no consistent relationship exists between AI narrative intensity and margin change; high-narrative companies both improved and worsened. Many forces, alongside AI investment, shaped the results — data quality, leadership sponsorship, technology maturity, and market cycles — making AI's independent contribution difficult to isolate from disclosure alone. It is tempting to infer that mechanism-disclosing organisations also perform better financially, but we treat this as a question for further researchrather than a finding. The organisations, however, can control the building of an internal measurement infrastructure to see this connection directly.
The capability exists. The narrative is confident. The financial evidence is largely undisclosed. Something is failing in between — and it is not the technology.
Value realisation disciplines built for ERP, CRM, and cloud assume that adoption follows deployment, that outputs are predictable, and that benefits can be estimated upfront. Applied unchanged to AI, these assumptions break down.
Traditional systems are deterministic — same input, same output — so value is easy to calculate. Generative AI is probabilistic - the same question can produce different answers. Value is derived from judgment and decision support, not transactional accuracy. That means, organisations must evaluate AI on outcomes, not outputs. In addition, productivity gains vary widely across contexts and user skill levels, and adherence to AI recommendations is often incomplete, even when accurate. A single uniform productivity multiplier will overstate realiaable value.
Consider a grocery retailer’s category manager whose range-review prep time falls 40% using a GenAI assortment assistant. The real value realisation from this outcome will come from using this time to review more categories or improve assortment quality, rather than continuing the existing ways of working. In essence, the technology is identical, but the realised value depends on how the organisations capture the gain.
Traditional systems are mandatory. AI is not. Users can ignore recommendations and override outputs — trust governs behaviour, built through explainability, consistency, relevance, and track record. Organisations that skip trust-building see adoption plateau well before the required returns are achieved. Many AI initiatives get stuck in “pilot purgatory,” failing to scale due to weak supporting investment, misaligned decision rights, and middle-management resistance. This is only compounded by poor data quality. Users who learn not to trust a system stay disengaged even after data improves — trust damage outlasts the data problem.
E.g., in Retail store operations, the head office deploys AI for labour scheduling, shrinkage alerts, and planogram scores. Store managers — not involved in the design, not measured on compliance — routinely override or ignore it.
AI costs scale with usage — inference, monitoring, and governance rise with volume — making unit economics the essential discipline, not total platform spend. At the retail/CPG scale, a single replenishment system across thousands of SKUs and hundreds of stores can generate over a million decisions per week.AI also delivers value in stages - pilots yield learning, workflow integration yields operational gains, and operating-model change yields structural impact. Evaluating only early-stage results risks cancelling programs that would pay off later.
Agentic AI adds another layer - value depends not just on adoption of recommendations but on the quality of autonomous decisions, exception handling, and governance at execution. As AI moves from advising to acting, organisations must measure not just what it recommends, but what it does and how it’s supervised. Humans stay in the loop through supervision, approvals, and exception handling — but the locus of value shifts from what people choose to do with a recommendation to how well the system performs when it acts alone.
In short, traditional systems are deterministic and mandatory, measured by output. Generative AI is probabilistic and optional, measured by adoption. Agentic AI operates autonomously, and is measured by decision quality and governance.
These differences call for a fundamentally different measurement framework - one that tracks not just what AI delivers, but how it gets there. The Value Realisation Stack organises this across four layers - layer four is where most AI investment sits today, layers two and three - the connecting layers, are almost never disclosed, layer one can only be credibly attributed to AI once the layers beneath it are accounted for.
Closing the gap means treating value realisation as a continuous discipline across four stages, each requiring deliberate measurement — not the assumed adoption.
The table below summarises the shift each stage requires and the actions that operationalise it.
Stage |
Core Shift |
Key Actions |
Governing Questions |
1. Identify |
Start with value, not technology |
|
Where can AI create measurable value, and what is it worth? |
2. Justify |
Build cases that reflect AI economics |
|
Realised Value = Potential Value × Adoption × Economic Capture - are all three modelled? |
3. Realise |
Embed AI into decisions, not just systems |
|
Does the workflow generate adoption, override, and decision-quality data? |
4. Sustain |
Govern for continuous value |
|
Is value decaying (model drift, disengagement, cost creep, competitive response) unnoticed? |
We found through an AI readiness assessment that a mid-sized emerging market CPG organisation struggled with inconsistent data quality, no retraining of models, unclear ownership, and low utilisation despite classical ML models and early GenAI tools already in production. The company had AI, it did not have AI value realisation. The core issue was trust, not technology. Unreliable outputs had taught sales planners to keep manual workarounds. The company already tracked beat adherence, strike rate, and promo execution compliance - but these measured whether sales activities were executed as planned, not whether the right demand was proactively identified and captured before it was lost. The highest-value opportunity identified was a supervised virtual sales agent recommending reorder quantities and trade schemes. The business case was sized from the mechanism up, with adoption rate as the leading indicator and incremental sales per distributor the lagging one, making the mechanism gap trackable for the first time.
A review of one large US-based retailer’s AI governance found substantial delivery capability across platforms, plus a dedicated value realisation function. Yet the organisation could only report what was on track, not whether delivered technology was generating the business outcomes it was funded to produce. The gap was not capability; it was governance architecture. Quarterly business reviews lacked cross-platform consistency. The fix required unified KPI definitions and end-to-end traceability from feature delivery through to business outcome - proof that governance infrastructure does not automatically scale with AI ambition.
The evidence points to a consistent pattern: capability is scaling, but the measurement, governance, and behavioural disciplines needed to realise value from it have not kept pace. Closing the gap requires treating value realisation as a continuous management discipline - identify by value pool, justify with honest adoption and capture assumptions, realize by embedding AI into decisions, and sustain through governance across outcomes, behaviour, economics, and risk.
The first wave of AI rewarded experimentation, the second wave deployment. The next wave will reward value realisation. Access to AI technology will become widespread. The discipline to convert it into measurable business outcomes will remain rarer — and may prove to be one of the most important sources of competitive advantage in the age of Generative and agentic AI.
Research Support: Karthik Saravanathangam – data analysis, visualisation, and synthesis of public disclosure findings.