9 minute read

Only 7% of enterprises say their data is ready for AI. The other 93% are running transformation programmes on a foundation they never built. A consultant’s guide to the layer nobody puts on a board slide, and why in Europe it has quietly become a legal precondition rather than an engineering nicety.

Every stalled AI programme gets the same autopsy. The model was not good enough. The use case was wrong. Adoption was poor. Change management failed. I have sat through enough of these post-mortems to notice a pattern: the diagnosis almost always stops one layer too high. Underneath the model, the use case, and the copilot nobody opens, there is usually a data foundation that was never built for what the organisation is now asking of it.

The number that should end the debate arrived in March 2026. A Cloudera and Harvard Business Review Analytic Services survey of 1,574 enterprise IT leaders found that only 7% say their data is completely ready for AI. Gartner expects enterprises to abandon 60% of AI projects through 2026 specifically because they lack AI-ready data foundations. The two figures describe the same problem from opposite ends. Almost nobody has built the foundation, and the projects built without it are the ones being quietly scrapped.

The previous posts in this series worked on the visible layers of enterprise AI: selecting use cases that matter, capturing the ROI that leaks, scaling agents, activating licences, redesigning the org chart, and most recently the verification bottleneck that caps how fast oversight can scale. This post goes underneath all of them. It is about the substrate every one of those layers silently depends on, and the reason it keeps failing is not technical. It is that almost no one owns it as an operating capability.

📊 The Numbers: Only 7% of enterprises say their data is completely AI-ready (Cloudera / HBR, March 2026), while Gartner expects 60% of AI projects to be abandoned through 2026 for lack of an AI-ready data foundation.

The failure everyone diagnoses one layer too high

Ask a leadership team why an AI initiative stalled and you get a familiar list. Wrong model. Immature vendor. Weak adoption. Insufficient training. Each of these is sometimes true. None of them is usually the binding constraint.

Post 2 in this series put a number on where the money actually goes. Data preparation alone consumes a large share of AI project budgets, and integration with legacy systems consumes more. Post 4 turned data readiness into a greenlight test: verify the data exists in usable form before committing to scale, because assuming it will be cleaned during the project is how the 95% end up underwater. What both posts pointed at, without naming it directly, is that the data layer is where most enterprise AI quietly dies. Gartner’s cost work is blunt about it. Up to 40% of AI project costs come from fixing data issues that were only discovered after deployment. That is not a modelling failure. It is a foundation nobody inspected.

The reason this keeps happening is structural, not careless. Data readiness is invisible in a demo. A copilot that answers three curated questions beautifully tells you nothing about whether the underlying estate can answer the fourth. So the foundation stays unexamined until scale exposes it, at which point the integration bill and the rework arrive together.

💡 Key Insight: A demo proves the model works. It proves nothing about whether your data can answer the next question, which is the only question that matters at scale.

“AI-ready” is not the data-quality problem you already know

Here is where most leaders get the scope wrong. They hear “AI-ready data” and reach for the familiar playbook: master data management, a data lake, a quality dashboard. Useful, and a decade out of date for what generative and agentic AI actually demand.

McKinsey’s technology practice, in work published only weeks ago, described the shift precisely. As AI leans on unstructured data (contracts, transcripts, emails, internal documents), that content no longer stays whole. It gets extracted, chunked, embedded, and recombined across systems, and each transformation alters the context of the original. The consequence is a new class of risk. When a single AI answer is assembled from fragments of many documents, the organisation often cannot reproduce which source, which version, or which transformation logic produced it. McKinsey’s phrase for the result is worth sitting with: the outputs become indefensible. In a regulatory audit or in legal discovery, that gap is not academic. It is material.

This is the real bar for 2026. AI-ready data is data that is discoverable, accessible in real time, governed by a single identity and policy model, high in quality with traceable lineage, and provisioned as reusable products that agents and copilots can consume across the estate. Most enterprise data estates fail at least three of those five tests. That is why the 7% figure is not an outlier. It is the honest number.

The five tests of AI-ready data

The five tests are a useful filter because they reframe the work. This is not a cleaning exercise you finish once. It is a set of properties your data has to hold continuously while agents pull it apart and put it back together thousands of times a day.

Why this is an operating-model problem, not a platform purchase

The instinct, once leaders accept the foundation is weak, is to buy a platform. Databricks, Snowflake, Microsoft Fabric, pick one and the problem is solved. It is not, and the reason is the most important point in this post.

The bars capture the trap. Almost no one has the foundation, a mature governance model, or a strategy anyone would honestly call mature, yet a clear majority of projects will be scrapped for exactly that reason. A platform licence does not move a single one of those bars on its own.

McKinsey’s point about platform maturity is the one to internalise. If two business units each build their own extraction logic, chunking strategy, embedding model, and retrieval configuration for the same underlying data, you do not get one AI-ready estate. You get two, and they disagree. The same question, asked in two parts of the company, returns two different answers, and neither is auditable. Duplication multiplies cost while destroying the one thing AI at scale requires, which is a consistent version of truth.

Platform maturity in an AI context therefore means something organisational: reusable extraction pipelines, shared embedding and indexing infrastructure, a common retrieval layer, and standardised guardrails that every application draws on. Tooling does not replace governance. It has to be governed centrally and consumed locally. This is the same shape Post 10 identified for the operating model as a whole. IBM’s 2025 research found that organisations running centralised or hub-and-spoke AI operating models see roughly 36% higher AI ROI than those where every team fends for itself. Data is the layer where that finding bites first.

In practice, the enterprises I see pulling ahead made one unglamorous decision early. They named an owner for AI-ready data as a governed product, with the authority to set shared standards, before they scaled a single agent. The ones still stuck treat data readiness as something each project fixes for itself, which guarantees they fix it five times and trust it zero.

⚠️ Watch Out: Buying a data platform without an owner and shared standards does not give you AI-ready data. It gives every team a faster way to build its own incompatible version of the truth.

Where the European lens changes the math

Everything above is true in every market. In Europe, it stops being optional.

From 2 August 2026, the EU AI Act’s obligations for high-risk AI systems are enforceable. Those obligations include logging, traceability, and a human-oversight architecture, all of which assume you can reconstruct how a given output was produced. Read that against McKinsey’s warning about indefensible outputs and the two collide directly. An AI system whose data lineage you cannot reconstruct is no longer only a quality problem or an ROI problem. It is a compliance exposure, with penalties that scale into the millions of euros or a share of global turnover.

This is the part US-centric AI strategy tends to skip, and it is a genuine advantage for European incumbents willing to use it. The 7% who have built lineage-traced, governed data are the ones who will be auditable when the deadline lands. Their competitors will be reconstructing provenance they never captured. For once, the regulatory environment rewards the organisations that did the boring foundational work first. In Frankfurt, I no longer treat data lineage as a data-team concern. It is a board-level readiness question with a date attached to it.

✅ Leadership Action: Treat data lineage as an EU AI Act deliverable, not a data-team preference. Any high-risk system whose outputs you cannot trace to source is a regulatory exposure before it is anything else.

The leader’s 30-day move

A concrete sequence, before the next platform invoice or the next stalled-pilot post-mortem.

  • Week 1. Trace one answer end to end. Take one high-value or one stalled use case and follow a single AI output back to its sources. If you cannot reproduce which document, which version, and which transformation produced it, you have found your real constraint, and it is not the model.
  • Week 2. Count the pipelines. Inventory how many separate extraction and embedding pipelines already exist for the same underlying data. Duplication is the clearest evidence that no operating model exists. Every duplicate is a divergent answer waiting to happen.
  • Week 3. Name one owner. Assign a single accountable owner for AI-ready data as a governed product, with authority to set shared standards across business units, and decide the shape (centralised or hub-and-spoke). This is one name on one slide, not another governance committee.
  • Week 4. Map lineage to the Act. Classify your AI systems against the EU AI Act high-risk criteria and check each one’s data lineage. Any high-risk system whose provenance you cannot reconstruct goes on the risk register now, not in August.

Do only these four things and you will already be operating ahead of the roughly nine in ten enterprises still treating data readiness as something a platform purchase will quietly resolve.

The consultant’s takeaway

Every layer of enterprise AI that leaders enjoy discussing (the models, the agents, the copilots, the org redesign) rests on a layer they prefer not to. The data foundation does not demo well, does not photograph well on a board slide, and does not generate a launch announcement. It simply determines whether everything above it works. The 7% who have built it will spend 2026 scaling. The 93% who have not will spend it running post-mortems that keep stopping one layer too high.

The uncomfortable reframe is that most AI strategies are data strategies that have not admitted it yet. The use-case portfolio, the agent architecture, the governance framework, all of it is downstream of whether the organisation can produce a consistent, traceable, reusable version of its own data. That is an operating-model decision about ownership and standards, not a procurement decision about platforms.

In Europe, the clock makes the point for me. When the EU AI Act’s high-risk obligations go live in August, the difference between the organisations that built a traceable data foundation and the ones that did not stops being a matter of AI performance and becomes a matter of who is auditable. In Frankfurt that is no longer a technical preference. It is the difference between a programme that scales and one that has to explain itself to a regulator it was never built to answer. The foundation was always the strategy. The only question is whether you build it deliberately, or discover it the hard way when the layer above it fails.

Updated: