Ninety-five percent of enterprise AI pilots produce no measurable financial impact. Three independent 2026 studies landed on that same number, which is unusual enough to take seriously. What the coverage does not say is where those pilots actually die. It is almost never the model. In the engagements I get called into, it is the security review, and by then the pilot has been dead for months because nobody scoped the work that leads to one.
The figure comes from MIT’s Project NANDA, from a 639-respondent field study by Domino Data Lab, and from an academic review of roughly 4.5 million production tests published this year. Different methods, different populations, the same answer. The supporting numbers point the same way. S&P Global found 42 percent of companies abandoned most of their AI projects in 2025. An IBM study of chief executives found only a quarter of initiatives delivered the expected return. Morgan Stanley reported that as of late 2025, only 21 percent of S&P 500 companies could name a measurable AI benefit. Set against something like 30 to 40 billion dollars of enterprise spend, that is an expensive amount of nothing.
The instinct is to read this as a verdict on the technology. It is not. It is a verdict on scoping.
Where the pilots actually die
I get called in at a specific moment. The demo worked, somebody senior liked it, and now it has to go into production. That is the moment the project meets security, identity, network and compliance for the first time, and it is usually the first time anybody has asked the questions those four functions ask.
Can this thing reach the systems it needs, from where it will actually run? What credential does it hold, what is that credential scoped to, and who rotates it? If it does something wrong at three in the morning, what stops it, and who gets paged? And six months from now, when someone asks what it did on a particular date to a particular record, can you answer with evidence rather than recollection?
None of those are AI questions. Every one of them is a question I have asked many times over about a firewall change, a new segment, or a service account nobody could account for. What has changed is that the answers now have to exist for a system that acts on its own initiative, and the pilot was never built to produce them. A practitioner put it well on X this month: the production bottleneck is governance, not intelligence, because most pilots fail at the transition when risk, security and legal cannot audit or bound the system’s behaviour. That matches what I see almost exactly.
The 80 percent nobody scoped
Roughly 80 percent of the work between a working pilot and a production system is data engineering, governance, workflow integration and measurement infrastructure. That estimate turns up repeatedly across the 2026 analyses, and it is consistent with how the work actually feels from the inside.
The trouble with that ratio is budgetary. The pilot budget covers the 20 percent, because the 20 percent is the part that produces a demo, and a demo is what gets the next tranche approved. The 80 percent has no line item, no sponsor and, most damagingly, no owner. So it becomes a slow argument between four departments who each believe it belongs to one of the other three. The project is rarely cancelled outright. It stalls, politely, and eighteen months later it shows up in a survey as a pilot with no measurable financial impact.
| The question | What the pilot answered | What production requires | Who owns it |
|---|---|---|---|
| Does it produce good output? | Yes, on curated inputs | Necessary, nowhere near sufficient | Data science |
| Can it reach the real systems? | Usually mocked or exported | Real integration, real credentials, real network path | Platform and network |
| Can you bound what it does? | Rarely asked | Scoped identity, authorisation boundary, kill switch | Security |
| Can you prove what it did? | No | An audit trail that survives an external question | Compliance |
| Did it move a number? | Not measured | A baseline captured before go-live | The business owner |
Ad hoc intelligence is not institutional intelligence
The sharpest explanation I have read this year came from a Smartsheet analysis published through Axios in July. Enterprises confused two different things: ad hoc intelligence, which makes one employee faster at one task, and institutional intelligence, which makes the organisation smarter over time. They built the first at scale and called it transformation. Their chief AI officer said it plainly, that individual task completion has got easier while working across systems and across teams has not improved at all.
That distinction resolves the apparent contradiction in the survey data, where staff report real enthusiasm and the finance function reports no impact. Both are true. Everyone genuinely is faster. None of it accumulates, because nothing crosses a system boundary, and crossing system boundaries is precisely the part that needs the integration, identity and governance work that nobody funded. It is the same shape as a control that reports green while protecting nothing, which I have written about in security controls that fail silently.
What is different in a German engagement
Anyone reading the American coverage and planning a rollout in Germany is missing three constraints that materially change the timeline.
The works council is a stakeholder, not a formality. German co-determination gives works councils rights over the introduction of technical systems capable of monitoring employee performance, and a tool that logs prompts, outputs and who submitted them sits squarely inside that description. This is not an obstacle to be routed around, it is a negotiation to be scheduled. Teams that discover it after the pilot lose a quarter. Teams that bring the council in while the pilot is still a pilot generally do not lose anything at all.
The compliance clock is already running. The German NIS2 implementation act came into force on 6 December 2025 and pulls roughly 29,500 companies into a documentation and reporting regime, with the registration deadline already behind us. Separately, the transparency obligations of the EU AI Act became enforceable on 2 August 2026. If your AI system touches an in-scope service, the evidence requirements are not a next-year problem. I have written before about how unprepared much of the Mittelstand is for the first of those, and AI projects are landing on top of that gap rather than beside it.
The documentation culture is an asset, if anyone points it at the project. German enterprises are usually better than their international peers at producing an audit trail, because regulation has required it of them for years. The failure I see is that this discipline gets applied to every system except the AI one, which is treated as an experiment exempt from the normal rules right up until the moment it very much is not.
What the successful minority did differently
The organisations getting a return are not distinguished by better models. Everyone has access to the same models. They are distinguished by having done four unglamorous things before go-live: they invested in the infrastructure first, they wrote the governance documentation before the pilot rather than after it, they captured baseline metrics before anything was switched on, and they gave the deployed system a named business owner who kept it after handover. That list comes from a practitioner who has run an autonomous agent in production for more than 300 days, and it is more useful than most vendor guidance because every item is something you can verify was or was not done.
The baseline metric is the one I would defend hardest. If you did not measure the process before you changed it, you cannot demonstrate improvement afterwards, and “no measurable financial impact” becomes literally true regardless of whether the system worked beautifully. The organisations reporting a return put it at roughly 3.70 dollars back for every dollar spent, and every one of them can answer the question “compared to what”.
What I would check before the next pilot
Five things, none of which require a strategy programme.
Name the owner of the 80 percent before the demo, not after it. One person, named, accountable for integration, identity, evidence and measurement. If nobody will take it, you have learned something important while it is still cheap to learn.
Put security in the pilot, not in the review. A security architect in the room during week one costs a few hours. The same person encountering the system for the first time at the production gate costs a quarter, and produces a worse answer, because by then the architecture has hardened around assumptions nobody challenged.
Capture the baseline in the same week you scope the pilot. Whatever the system is meant to improve, measure it now. This is the cheapest thing on the list and the most frequently skipped.
Write down what the system is allowed to do. Not what it is intended to do, what it is permitted to do, expressed as an identity with a scope. An AI system that acts is an account with standing access, and that has consequences I have set out in non-human identity sprawl. If the answer is a shared key with broad permissions, the security review will end the project and it will be right to.
Schedule the works council conversation on day one. Not as a compliance step near the end. As a design input, because it changes what you log and how, and retrofitting that is far more expensive than designing for it.
The part worth keeping
None of this says the technology does not work. It says the failure rate is measuring organisational readiness and reporting it as AI performance. That distinction matters, because the two have entirely different remedies. One requires a better model, which you cannot influence. The other requires somebody to own the boring 80 percent, which you can arrange this week.
The 95 percent is not evidence of a bubble. It is evidence that most organisations are running AI projects the way they ran network projects before change management became a discipline: impressive in the lab, undocumented in production, and impossible to prove anything about afterwards. We fixed that once, and the lessons transferred, as I set out in what enterprise security and compliance work actually teaches. The same fix is available here, and it is not a technology purchase.
If your pilot works and production keeps slipping, the blocker is usually integration, identity and evidence rather than the model. Request a review and I will give you an outside read on which of the 80 percent is missing.