// Article · August 17, 2026 · 10 min read
The Model Was Right. The Store Manager Overruled It Anyway.
Retail supply chain AI has quietly stopped being a modelling problem. The 2026 evidence says the binding constraint is now inventory nobody counted, incentives that punish trusting the output, and a governance question nobody has answered.
// Contents
In May, ReFED — working with The Spoon — published The Food Operating System, a study of how AI is actually being deployed across the food system. Fresh grocery turns out to be where the evidence for AI is strongest anywhere in retail. Predictive ordering works. The waste reduction is real and measurable.
And the report's most interesting finding is that store managers keep overriding it.
Not occasionally. Routinely. Technically accurate, AI-generated ordering recommendations get overruled by the people on the floor — and ReFED's conclusion is not that the operators are irrational. It is that grocers need to restructure internal incentives so managers are rewarded rather than penalised for trusting the data. Corroborated
Think about what that implies. A manager who follows the model and ends up with an empty shelf takes the call from the regional director. A manager who over-orders and quietly wastes product does not. Under those incentives, overriding the AI is the rational move. The model is right and the human is right, and the system still produces the wrong outcome.
This is the single most useful thing to understand about AI in the retail supply chain in 2026, and it generalises far beyond produce: the bottleneck has moved. It is no longer the model. Almost nobody in the vendor conversation has caught up.
The modelling argument is over
It is worth being precise about how settled the forecasting question actually is, because vendors and sceptics both have reasons to keep it open.
The M5 competition ran on 42,840 real Walmart sales series. It was the first forecasting competition in which every top-performing method was a pure machine-learning approach, and every one of them beat every statistical benchmark and their combinations. That result was independently adjudicated and peer-reviewed. Nobody paid for it. Corroborated
The applied results line up. Afresh, which builds predictive ordering for fresh categories, is live in more than 12,500 departments across 40 states with Albertsons, Meijer and Wakefern, and raised $34m in April. It reports roughly 25% shrink reduction, a 3% sales lift and 7% better inventory turns — vendor-reported figures, and you should hold them loosely, but the direction is corroborated by ReFED's independent read of the sector. For context, US retail generated 3.98 million tons of surplus food in 2024, worth about $26.9bn.
Away from the shelf, the wins are even less glamorous. Modern Retail reports that C.H. Robinson built AI to read inbound shipper emails, classify intent and generate quotes — now booking around 3,000 appointments a day across 26,000 locations. ArcBest put AI on route optimisation and reports roughly $15m in savings. Target is building a digital twin for inventory. Corroborated
Read that list again. The highest-return AI deployment in freight right now is a machine that reads email. That is the actual state of the art, and it is considerably less exciting than the keynote.
So why is almost nobody running it?
Here is the number that should end most vendor meetings. Sage's 2026 State of the Supply Chain, a January survey of more than 200 retail and wholesale operators, found that 10% have AI live in supply chain workflows. Not piloting. Live. And adoption tracked data readiness and visibility maturity, not interest or experimentation. Corroborated
Gartner's supply chain surveys point the same way from a different angle: 67% of supply chain managers say current digital investment is going to AI (n=394, surveyed November 2025 to February 2026), while 55% of senior leaders say they are unclear what return it might yield (n=135, January to April 2026). Gartner also expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing cost, unclear value and inadequate risk controls.
Money is moving. Understanding is not. And when a majority of the people who bought the technology cannot tell you what it returned, every "42% accuracy improvement" in a vendor deck stops being a measurement and becomes a marketing asset.
The reason is not mysterious, and it is not really about AI at all.
Retail inventory record accuracy without RFID sits around 65%. RFID takes it above 95% — lululemon reports 98% across its store network. But per IHL's 2026 study of more than 400 retail brands, only about 20% of retailers use RFID at all, and just 9.1% have computer vision deployed in stores. Corroborated
Point a forecasting model at a stock file that is 65% accurate and it does not fail loudly. It produces confident, precise, wrong answers — faster than before, and with a much nicer dashboard. The highest-ROI "AI project" in most retailers is a tagging and counting programme with no AI in it whatsoever.
The same IHL study contains the statistic that should reframe the entire category: 89% of retail orders are fulfilled from a store. Every gleaming warehouse-robotics announcement describes a minority of the supply chain. The unsolved problem is the shop floor, where the inventory data is worst and the robots mostly aren't.
Which is why the most interesting deployment in grocery this year is not a forecasting engine. Schnuck Markets, a 111-store Midwestern chain, went chainwide with Simbe's Tally shelf-scanning robots. Tally covers the floor two to three times a day, scanning around 35,000 products per pass and roughly 4.2 million products daily, and a single robot identifies up to fourteen times as many out-of-stocks as a human worker. Corroborated Simbe reports up to 60% improvement in on-shelf availability across customers including BJ's and SPAR Austria — vendor-reported, but the mechanism is obvious enough that it barely needs defending.
The AI in that store is not recommending anything. It is counting. A regional grocer went chainwide before most national chains finished a pilot, which is its own quiet lesson about where the bureaucratic drag lives.
Three giants, three completely different fights
If you want to understand the shape of this market, stop reading vendor decks and look at where the three biggest AI companies have actually placed their bets. They looked at the same industry and picked three fights that barely overlap.
Anthropic is going after the operator's console. The strategy is not to build supply chain software but to become the interface to everyone else's, through MCP connectors. Earlier this month ShipBob shipped the first Anthropic-verified fulfilment connector — more on the mechanics below. The "verified" label is a real distinction: Anthropic reviews verified connectors for security, reliability and compatibility, where community connectors only clear automated checks, and enterprise admins can restrict which are permitted. The early-ROI sweet spot reported in logistics is deliberately unglamorous — classifying delivery exceptions and customs holds, drafting the customer comms, proposing a resolution. The highest-volume, highest-stress, most repetitive work in the building.
OpenAI is going after the shopfront — and is conspicuously absent from the back end. Search hard and there is no named retail supply chain deployment. What exists instead is the Deployment Company, launched with 19 global partners and an investment stake from Bain, which reaches supply chain through consultancies and private equity portfolios rather than through a product — "supply chain optimisation" appears as a named focus area for portfolio companies. Alongside that sits the Agentic Commerce Protocol with Stripe, live since late 2025 with Target, Instacart and DoorDash. That is checkout inside ChatGPT. Demand side, not supply side. Corroborated
The absence is the finding. OpenAI has the consumer surface area; warehouse management systems have the operational data. Those are different games, and OpenAI is playing the one it already leads.
Alibaba skipped the software argument and bought the physical network. In March, Cainiao launched a global robotic warehouse network across seven markets — Hong Kong, mainland China, the Netherlands, Spain, France, Germany and the United States — combining next-generation warehouse robots with an AI scheduling system coordinating robot fleets and equipment, aimed at expanding next-day and two-day coverage for Alibaba, AliExpress and Temu merchants.
The more striking half is Accio, Alibaba's AI sourcing agent: more than 10 million monthly active users as of March, trained on a billion product listings and 50 million supplier profiles, with a July plugin that finds suppliers, compares offers and keeps negotiations moving around the clock. Corroborated
Sit with that. Western retail has spent a year arguing about whether a shopping agent will buy your customer's trainers. Meanwhile ten million businesses are already using an AI agent to negotiate with factories. We are debating the shop window; the back door was automated while nobody was looking.
What "agentic" actually looks like when you open it up
ShipBob's launch on 5 August is the clearest available specimen of a real agentic supply chain deployment, and it is worth describing mechanically rather than in adjectives.
A merchant installs the ShipBob connector from Anthropic's directory in one click. Authentication is OAuth with their existing ShipBob account — no API tokens to copy, paste or rotate. Claude then has more than 70 read and write actions across ten operational areas, with every call hitting the live ShipBob developer API. An operations person types "which orders are at risk of missing their ship-by date, and reprioritise them," or "create an inbound receiving order from these three POs," and Claude queries live state and writes the change back. Every action is logged in an audit trail called MCP Insights inside a new AI Hub. ShipBob says merchants already submit thousands of MCP requests daily. Corroborated
The detail that makes this a story is the write access. Most connectors are read-only — AI that reports. This one acts on live warehouse state. That distinction is the whole difference between a dashboard and an operator.
Mike McPhail, senior director of distribution at Vuori, described the change as: "I can't even remember the last time I was on a call where someone asked me to decide which orders to prioritize."
But be clear about what this is. It is a natural-language console with write access and a logbook, sitting on top of conventional fulfilment software. The reordering and routing logic underneath is the same software it always was. The intelligence is in the interface, not in some new supply chain brain.
That is the honest shape of agentic supply chain in 2026 — and it is a much smaller, much more achievable promise than the autonomy being marketed. Kearney's assessment, reported by Modern Retail, is that the fully autonomous AI supply chain "remains aspirational," with only 17% of supply chain organisations pursuing immediate transformational redesign against 83% applying AI incrementally.
The question nobody has answered
Full read and write access to live fulfilment, through a chat window, is a genuinely new operational risk surface. The audit log exists precisely because you need one.
So: who approves a write?
When an AI can reorder stock and reroute shipments inside a production system, "who signed off on that?" stops being an IT question and becomes a financial-controls question. There is no established answer — no segregation-of-duties convention, no approval threshold, no standard for what a human must review before an agent commits an inventory change. Gartner's projection that 40% of agentic projects die by 2027 on "inadequate risk controls" is describing exactly this gap.
That is the governance fight of 2027, and it will be settled by finance directors and auditors, not by model quality.
A brief note on how fast fiction travels
While researching this piece, we repeatedly encountered a widely-circulated claim that Sainsbury's and Palantir deployed an AI fresh-produce spoilage forecasting system in July, cutting waste by a projected 15%. It is a good story, precisely detailed, and it appears in AI-industry news roundups.
We could not stand it up. It traces to a single conference news page. There is no Sainsbury's release, no Palantir announcement, and no coverage in the retail or grocery trade press. Unverified — which is not the same as false, but it is not something anyone should repeat on a stage.
That is worth flagging on its own terms. In a category where 55% of buyers cannot measure their own returns, unverifiable case studies circulate faster than verified ones, because nobody has a mechanism to check them.
What actually changed
Retail supply chain AI is not failing, and it is not delivering what was sold. Both are true, and the reason is that the industry is still arguing about the wrong layer.
The models work. That has been settled since M5. What has not been solved is a stock file that is 65% accurate, a store estate that fulfils 89% of orders with the worst data in the business, an incentive structure that punishes the manager who trusts the output, and a controls framework for AI that can write to production systems.
None of those are model problems. All of them are management problems wearing a technology costume.
Which is, incidentally, the most encouraging read available. Model quality is bought from three companies. Everything else on that list is within any retailer's own control — and the retailers deploying AI and edge infrastructure are, per IHL, planning to add store associates at nearly double the rate of those that are not. Whatever this is, it is not the automation story everyone rehearsed.
It is a counting story, an incentives story, and an approvals story. Considerably less cinematic. Considerably more actionable.
This post is also published on our Substack newsletter at edge-ai.forum. Subscribe for the weekly roundup direct to your inbox — fresh AI news, executive context, and devices + robotics every Friday morning.