In September 2025, Starbucks rolled out AI-powered inventory scanning to over 11,000 North American stores — iPads with computer vision, meant to turn an hour of manual counting into 10 minutes. Nine months later, Starbucks killed it. Baristas are back to counting by hand.
Somewhere in the rollout, “faster” quietly became the win condition. Reflective refrigerator doors double-counted milk. The system confused similar products. One employee summed it up to Fortune: it started off not particularly accurate and got less accurate over time. Fast and wrong is still wrong — it’s just wrong at scale, and wrong with confidence.
A win isn’t “the AI finished the task quicker.” A win is “the output was right, and it stayed right.” Those are different claims, and the second one is the only one that actually pays for the rollout. Speed without accuracy doesn’t save labor — it just moves the labor downstream, from counting the shelf to catching what the AI got wrong.
That’s exactly what a real pilot is for. It went straight to 11,300 stores with no visible test phase — no stretch where the AI count and a manual count ran side by side long enough to find out if the tool was actually right. A real pilot means deliberately doing the work twice for a while: let the AI count, and count it yourself too, across different lighting, different layouts, different shifts. That’s more labor during the trial, on purpose, in a handful of locations — instead of finding out for free, nine months later, in eleven thousand of them.
Run it once and call it a win is a guess wearing a results deck. Run it against a manual count until the two numbers agree — or you learn exactly where and why they don’t — and now you have a win you can actually scale.
So what — on your last AI rollout, was “it worked” based on a measured accuracy rate, or on how fast it felt? Now what — what’s the smallest pilot you could run this month that forces you to prove accuracy before you chase speed?
Source: Reuters (via Fortune, GeekWire, Fast Company)

Leave a comment