Every AI transformation has its moment of triumph—a successful pilot, enthusiastic users, and a compelling business case for scale. But a pilot proves only that the technology works under ideal conditions, not that an organisation is ready to depend on it. In this thought-provoking perspective, procurement transformation leader Palash Mukherjee challenges a common misconception, revealing why trust, governance, and human behaviour—not algorithms—usually determine whether AI programmes scale successfully or quietly unravel.
I’ve sat through a lot of AI demos in procurement over the last three years. Most of them are good. Genuinely good. The classification is sharp. The routing is fast. The dashboard looks like something out of a vendor’s fever dream of what procurement could be. And almost every one of those demos will disappoint someone, eighteen months later, in a steering committee meeting nobody wants to be in. Not because the technology failed. Because the demo was never what needed proving.
Palash Mukherjee
A demo is built on the vendor’s best data. It runs the vendor’s best scenario. It’s presented by the vendor’s best people, to an audience primed to be impressed. It is, by design, the easiest possible version of the problem. That’s not dishonest. It’s how demos work. Every experienced buyer in the room technically knows this. What they forget, under the pressure of a live decision and a budget cycle, is simple. The gap between “the pilot worked” and “the programme works” isn’t a scaling problem. It’s a completely different set of problems wearing the same trousers.
I was in the room for one of these. The pilot had run beautifully. Three months. One clean category. One cooperative business unit. Hand-picked data. The recommendation to scale was unanimous. Eleven months later, the same tool was live across the full spend base. And nobody trusted its output enough to act on it without checking it manually first.
Nothing about the AI had changed. Everything around it had.
That gap — between the thing that gets approved and the thing that gets delivered — is where most procurement AI programmes quietly go to die. And almost nobody names it before it happens.
THE PILOT DELUSION: WHAT WE ACTUALLY PROVE
Here’s the uncomfortable part. A successful pilot proves the model works. It doesn’t prove the organisation is ready for it to act.
Those are different claims. Procurement leaders routinely blur them — not out of carelessness, but because the pilot is designed to answer the first question and stays silent on the second. Nobody asks, “did we prove this will hold up outside a controlled environment?” The pilot was never built to test that. It was built to get funding for the next stage. So, the business case gets written on pilot evidence. The steering committee approves scale on pilot evidence. The vendor’s success story is the pilot. The organisation walks into full deployment having genuinely proven something true. Just not the claim everyone assumed they’d proven. This isn’t a criticism of pilots. It’s a criticism of what we let a pilot stand in for.
THE TRIPTYCH OF FAILURE: HETEROGENEITY, COOPERATION, AND TRUST
When a procurement AI programme stumbles after a strong pilot, the cause is almost always one of three things. And they rarely get discussed at the point where discussing them would help.
Data heterogeneity. The pilot ran on one clean category, hand-picked precisely because it was clean. Scale means every category. The messy ones. The ones with three overlapping taxonomies. The ones where “IT Services” and “Tech Consulting” and “Professional Fees” all describe the same spend. Nobody’s incentivised to volunteer the ugly category for the pilot. The vendor wants a win. The internal sponsor wants a win. So the pilot quietly avoids the exact conditions it will eventually have to survive.
Stakeholder cooperation collapses. The pilot business unit was chosen because someone there already believed in the initiative. A willing sponsor. An engaged team. People who wanted this to succeed and behaved accordingly. Scale means the business units with no stake in the outcome. The ones who never asked for this. The ones whose workarounds were already working fine, thank you. One cooperative team is not organisational buy-in. It’s a sample of one, dressed up as evidence.
Trust decays under real stakes. This is the one that gets talked about least. It’s also the one that actually killed the programme I watched. In a pilot, a wrong output is low-consequence. It’s a sandbox. Nobody’s supplier relationship gets damaged. Nobody’s budget gets misallocated. Nobody must explain a bad commercial decision to their CFO because the pilot’s output was wrong. At scale, a wrong output has weight. So users start manually checking everything — quietly, individually, without ever raising it as an issue. And the checking becomes the new normal. The automation is technically live. Nobody’s relying on it.
Of the three, trust decay is the most dangerous. It’s invisible until it’s structural. Data problems show up in error rates. Stakeholder resistance shows up in adoption metrics. Trust decay shows up nowhere — until you notice, eighteen months in, that the tool is running and everyone’s still doing the job manually alongside it.
CASE STUDY: THE MEETING WHERE IT ACTUALLY BROKE
The pilot had been a genuine success by every measure anyone had agreed to measure. One category. Three months. An engaged sourcing team that had wanted an AI-assisted classification tool for years and treated the pilot like the answer to a prayer. The output was accurate. The team trusted it within about two weeks, because they understood the category well enough to sanity-check the early recommendations. And the recommendations kept being right. The business case for scale wrote itself. Full spend base. All categories. Twelve-month rollout.
By month six of the rollout, the tool was live and technically functioning across the estate. Classification accuracy, measured the way the vendor measured it, hadn’t dropped. But in a steering committee update, someone from a different category team — not the pilot team, a team that had inherited the tool without ever asking for it — mentioned, almost in passing, that they’d started running a manual check on anything the AI classified with less than full confidence. Nobody challenged it. It sounded reasonable. Prudent, even.
Three months later, that “just to be safe” habit had spread. Not because anyone decided it should. Because nobody had ever explicitly said when the organisation was supposed to stop checking. The pilot team’s confidence had been earned through weeks of direct, hands-on validation in a category they knew intimately. Nobody had built an equivalent trust-building process for the teams who came after them. So each new team, arriving cold, defaulted to the only sensible response to a tool they didn’t yet trust: check it manually until proven otherwise. Except “until proven otherwise” never arrived. Because nobody owned that transition. The programme had a data migration plan, a change management plan, a training plan. It didn’t have a plan for the specific moment a team was meant to stop double-checking and start relying.
THE SILENT ACCUMULATION OF SHADOW WORK
Eighteen months in, the tool was processing the same volume it always had. The manual checking layered on top of it had become permanent, unbudgeted, invisible shadow work. Procurement doing the job twice — once by the AI and once by habit. Nobody in the governance structure had visibility into it, because nobody had ever asked the question: who owns the moment we start trusting this? The AI hadn’t failed. The organisation had simply never finished the job of deciding when to believe it.
Here’s something worth naming plainly, because it explains why this pattern repeats across so many organisations rather than being a one-off. Nobody in that steering committee was being dishonest. Nobody was hiding a problem. Each individual decision — hire more reviewers, add a manual check, keep the old spreadsheet running “just for now” — made sense in isolation. The failure wasn’t a decision. It was the absence of one.
That’s a harder problem to fix than a bad decision, because there’s no single moment to point to. No one signed off on shadow work. It simply accumulated, one reasonable choice at a time, until the programme was running two systems in parallel and calling it one.
This is why the fix can’t be a training session or a reminder email. It has to be a structural decision, made deliberately, before go-live — not discovered afterwards in a steering committee update that nobody thought to question.
WHAT “PROGRAMME-READY” ACTUALLY LOOKS LIKE
None of this means the pilot-to-scale gap is unmanageable. It means it needs to be treated as a distinct set of design problems, not an inevitable side effect of getting bigger. Three shifts matter most, and they map directly onto the three breaks above.
Include a messy category in the pilot, deliberately. Not as a stress test to be endured. As evidence that gets built into the business case. If the pilot only ever proves the tool works on easy data, the business case is only ever proven for conditions that won’t exist at scale. Choosing one unglamorous, genuinely difficult category for the pilot costs a few weeks of a harder result. It buys a business case that’s actually representative of what’s coming.
Put a resistant business unit in the pilot cohort. Not the enthusiast — the team that’s territorial, sceptical, or has something to lose. Not the team that’s been asking for this for years. Their reaction, not the enthusiast’s reaction, is the one that predicts what adoption looks like across the other eleven business units who didn’t volunteer.
Name someone as owner of the trust transition — before go-live, not after. This is the one almost nobody does. It’s also the one that would have changed the story above. Somebody needs explicit accountability for the specific moment a team is expected to move from “checking everything” to “trusting the output, checking exceptions.” That’s not a training deliverable or a change management workstream in the generic sense. It’s a named decision point, with a named owner, built into the programme plan with the same seriousness as a go-live date.
Get these right and the pilot stops being a demo of what’s possible in ideal conditions. It becomes a genuine rehearsal for the conditions the programme will actually face.
THE QUESTION NOBODY ASKS BEFORE SIGNING OFF SCALE
Every procurement AI business case I’ve reviewed answers the question “does the model work?” Very few answer the question that determines whether the programme succeeds: what happens the first time it’s wrong at scale — and whose job is it to notice?
Not whose job is it to fix the model. Whose job is it to notice that trust is quietly eroding. That a workaround is becoming a habit. That the automation everyone signed off on is technically running, while the organisation has, without ever deciding to, gone back to doing the work by hand.
That question won’t show up in a demo. It won’t show up in a pilot’s success metrics, because pilots are explicitly designed to be low-stakes. That’s what makes them safe to run. It’s also precisely what makes them unrepresentative of what comes next.
Before you scale your next pilot, ask the question the demo never had to answer: what happens the first time it’s wrong — and does anyone know whose job it is to notice?