AI Starts with your Data: Building an AI-Ready Foundation on AWS
In this article
Every enterprise AI roadmap eventually runs into the same wall, and it isn't the model.
Gartner projects that through 2026, organizations will abandon 60% of AI projects that aren't supported by AI-ready data, and a Gartner survey of 248 data management leaders found that 63% of organizations either lack the right data management practices for AI or aren't sure whether they have them. McKinsey's 2025 research adds the investment side of the picture: 92% of companies plan to increase AI investment over the next three years, while only 1% of leaders describe their organization as mature in how AI is deployed. And S&P Global Market Intelligence found the average organization scrapped 46% of its AI proof-of-concepts before they ever reached production.
Across these studies, the common constraint is the data foundation rather than the model.
That's the premise behind this post, before an organization can scale agents, copilots, or predictive models, it needs a data foundation that's actually ready for AI to use. Here's what that means, what tends to get in the way, how the layers fit together on AWS, and where to start.
Defining AI-ready data
AI-ready data isn't a data quality score or a single checklist. It's data that four things are simultaneously true of — and most enterprises have pockets of this, while very few have all four consistently across the data AI actually needs to touch, which is usually the messiest, most sensitive, most distributed data in the company.
Three structural constraints
The pattern recurs consistently across enterprise data programs WWT's research team has tracked: three constraints that compound rather than sit side by side.
Legacy architecture debt
Years of point-to-point integrations, siloed warehouses, and schema-on-write systems were built and tuned for one workload: scheduled batch reporting to a known audience, reviewed periodically. AI runs a completely different workload against the same infrastructure — on-demand, high-frequency queries; semantic as well as structural lookups; and access decisions evaluated dynamically, per query, rather than reviewed periodically at the report level. Legacy debt is not only aging systems; it is infrastructure tuned for a workload AI does not run.
Data gravity
The more data accumulates in each system, the harder and more expensive it becomes to move. Compute gets pulled toward where the data already lives, lock-in risk builds up passively well before a deliberate platform decision gets made, and "just copy it somewhere AI-friendly" becomes a recurring tax on every new use case. Left unmanaged, data gravity determines the platform decision by default.
Ungoverned AI data sprawl
As teams adopt AI tools independently — often with good intentions and real delivery pressure — copies of sensitive data get pulled into notebooks, vector stores, and third-party tools faster than governance can track them. The risk isn't hypothetical: access control, audit trail, and cost visibility degrade at once, usually discovered after the fact rather than designed around in advance.
The foundation, layer by layer
Strip away any specific vendor's product names, and an AI-ready data foundation has four layers that show up in almost every serious enterprise architecture. Understanding what each layer is responsible for — independent of tooling — makes it much easier to evaluate any platform's fit, AWS included.
The point of separating these layers conceptually first is that they fail independently. A strong storage layer with weak governance produces fast, ungoverned answers. Strong governance with poor connectivity produces a beautifully cataloged set of data nobody's systems can actually reach. AWS's 2026 data and AI investments — the Zero-ETL expansion, SageMaker Lakehouse, DataZone's generative-AI-assisted cataloging, and Bedrock's move to a managed knowledge base primitive — are best understood as AWS filling in all four layers concurrently, rather than optimizing any single one in isolation.
Approaches to organizing enterprise data
The four layers above are agnostic to a separate, equally important decision: who owns the data day to day, and how it's governed. No single architecture leads in 2026; the pattern is convergence. Data mesh and data fabric get discussed as if they're competing choices, but Gartner treats them as complementary, predicting that firms adopting one will adopt the other within two to three years. A data lakehouse, meanwhile, is increasingly just the shared technical foundation underneath whichever ownership model an organization picks.
| Approach | What it actually is | Where it fits |
| Centralized | One team owns data end to end and enforces standards directly. | Early-stage AI programs; regulated use cases needing tight, auditable control. |
| Federated / hub-and-spoke | A central function sets policy and shared infrastructure; domain teams execute within it. | The most common real-world pattern in 2026 — most enterprises land here. |
| Data mesh | Domain teams own their data as a product end to end, published into a shared catalog. | Organizations with real domain-level ownership readiness — not a default starting point. |
| Data fabric | An automation and metadata layer that connects distributed sources without moving the data. | Layered on top of almost any ownership model to cut integration work. |
| Data lakehouse | The underlying storage platform — open, queryable, combining lake flexibility with warehouse structure. | Increasingly the shared foundation underneath mesh, fabric, or centralized approaches alike. |
The survey data backs up the convergence story: across several 2026 enterprise governance studies, roughly 65% of data leaders now prefer federated or hybrid governance models over a purely centralized one — up sharply as data estates spread across multi-cloud, SaaS, and on-premises systems. Two real, public examples show how differently this plays out in practice. Uber's "Hive Federation" initiative decentralized 16,000 datasets off a single overloaded Hive warehouse using pointer-based federation, solving a scaling bottleneck rather than a governance problem. Zalando runs a hub-and-spoke model across roughly 400 teams, with about 200 datasets centrally stewarded, another 1,000 with domain input, and a long tail fully domain-owned — because, in their own words, "the central team is not the expert about everything."
The relevant question is not which architecture leads, but which combination of ownership model, connectivity approach, and storage foundation fits a given organization's constraints. That is a matter for evaluation rather than assumption.
Case study: a global automotive manufacturer
A global automotive manufacturer provides a publicly documented example of one point on that landscape, rather than a universal prescription. One of its plants was managing data across siloed systems with inconsistent quality and limited visibility into what existed or who could use it: a familiar starting point for any large manufacturer with decades of operational systems layered on top of each other.
For this specific problem, the plant chose a data mesh — a federated model where individual domains own and publish their own data products into a shared catalog, rather than funneling everything through one central team. On AWS, that meant Amazon DataZone as the governed marketplace for those data products, with domain owners registering their own assets and consumers discovering and requesting access through a self-service portal secured by federated identity. Generative AI was layered in specifically to generate business metadata automatically, which the organization expects to cut data access time from weeks to minutes.
Details and architecture vary by industry. what recurs is the general shape of the solution rather than the specific pattern — governed, federated data access tends to outperform both a fully centralized bottleneck and fully ungoverned sprawl.
Data maturity stages
The organizations that get this right rarely treat it as a single, big-bang re-architecture. They treat it as a maturity curve, moving one deliberate stage at a time — which mirrors how WWT's own research frames enterprise data maturity for AI: five stages, each unlocking a different class of AI use case rather than promising all of them at once.
The practical implication is that the scope of agentic AI deployment should match an organization's actual data maturity stage. A workload-first, platform-neutral roadmap — start from the specific business use cases that matter, then build only the foundation those use cases require — tends to outperform a foundation-first "boil the ocean" rebuild, both on time-to-value and on the odds the initiative survives its first budget review.
Enterprise AI use cases in production
The four-layer foundation matters more once it's clear which AI use cases enterprises are actually putting into production today — and at what rate. The ranking below, compiled from 2026 enterprise surveys across Gartner, McKinsey, IDC, Forrester and S&P Global Market Intelligence, orders common use cases by how widely they've been adopted into production so far.
Adoption tracks closely with how narrow and well-bounded the underlying data domain is. Customer service and code generation lead because the data they need — case history, product docs, source repositories — is comparatively contained and already reasonably governed. Legal, contract-heavy, and cross-domain use cases trail because they require pulling from more systems at once, with less consistent governance across all of them.
It's also why multi-agent orchestration is growing fastest inside the highest-adoption functions first: 22% of production agent deployments now coordinate three or more agents, and that share is projected to roughly double by 2027 — concentrated where the data foundation was already solid enough to support one agent reliably before a second and third were added.
Where to start
None of this requires having the whole five-stage journey mapped out before taking a first step. It requires an honest, evidence-based look at where an organization's data actually stands today — its governance posture, its connectivity gaps, its highest-value and highest-risk data — measured against the use cases it actually wants AI to handle, and against the full landscape of ownership and architecture approaches, not just the one that happens to be trending.
WWT's Cloud Practice works with clients on this assessment across three connected phases.
This is exactly the structure of WWT's Data Foundations of AI Briefing, available on AWS Marketplace, and it reflects lessons from WWT's own data strategy work across industries combined with hands-on AWS and generative AI expertise. WWT works as an independent partner: the evaluation is built to fit the client's environment rather than a particular hyperscaler's native tooling or a predetermined architecture pattern. For organizations earlier in scoping the question, a conversation with WWT's Cloud Practice is the more general starting point — what does "AI-ready" mean for this specific data estate, and what's the shortest credible path to it.