Digital & AI Series: The Data-AI Chicken & Egg Story – Do You Fix the Foundation First or Start AI first?

One question that comes up most frequently these days in all my recent conversations with business leaders, and it’s one that doesn’t have a clean, comfortable answer.

“Do we need to fix our data foundation before we can get real value from AI?

Or

Do we start with AI now and risk building on unstable ground? And if we wait to fix the data first, don’t we just fall further behind?”

It’s not a naive question. Rather, It’s one of the sharpest strategic tensions in enterprise AI today. And the fact that most frameworks don’t address it directly is part of why so many businesses are stuck.

Let’s try to decipher it honestly.

Why the question matters more than it used to?

In my previous article in this series, I wrote about the distinction between AI at the core, where AI is embedded into the fundamental operating model of the business and AI on top, where AI is layered over existing foundation to accelerate, assist, and augment without redesigning the underlying structure.

That distinction matters here because the data question isn’t the same question for both approaches.

If you are doing AI on top, your data requirements are often more contained. You are applying AI to specific tasks, specific documents, specific workflows. The model doesn’t need to understand your entire business; it needs to be useful in a defined context. The bar for “good enough data” is lower and often achievable without a full transformation of your data infrastructure.

If you’re doing AI at the core, rebuilding processes around AI, creating systems that make autonomous or semi-autonomous decisions, operating at enterprise scale, then the data foundation is not optional. Weak data pipelines, siloed records, inconsistent taxonomies, and poor data governance don’t just reduce the quality of AI outputs, they make the whole edifice structurally unsound.

The mistake is treating these two situations as the same problem.

The real cost of bad data in AI

Most organizations know, at some level, that their data isn’t perfect. What they underestimate is how differently that imperfection behaves inside an AI system versus a traditional reporting or analytics environment.

In a traditional BI environment, bad data produces bad reports. Humans look at those reports, notice anomalies, apply judgment, and compensate. The failure mode is visible and bounded.

In an AI system, especially one that’s integrated into a process or workflow, bad data produces plausible-looking outputs that are wrong in ways that are harder to detect. The model doesn’t flag its own uncertainty. It generates a confident answer. And that answer travels downstream, into decisions, into customer interactions, into operational processes before anyone has a chance to catch it.

MIT research in 2025 found that AI models are actually 34% more likely to use confident language when hallucinating than when giving factual answers. The system sounds most sure of itself precisely when it is making something up.

UnitedHealthcare: when AI confidence overrides clinical judgment. The US health insurer deployed an AI model called nH Predict to assist with post-acute care coverage decisions for elderly Medicare Advantage patients. A 2023 class-action lawsuit alleged that the model, despite a reported 90% error rate on appealed decisions, was being used to systematically deny claims that treating physicians had deemed medically necessary. The lawsuit alleged that UnitedHealth’s internal goal was to keep patient rehabilitation stays within 1% of the algorithm’s projection, effectively replacing clinical judgment with model output. What made this particularly dangerous was not that the AI was obviously wrong. It was that the outputs looked authoritative enough that only 0.2% of affected patients appealed. The errors travelled downstream unchallenged. The case remains active in federal court, with the company ordered in 2026 to produce broad discovery on the model’s implementation.

This is the compounding risk. It’s not just that AI amplifies bad data, it’s that it does so at a speed and scale that outpaces the human review mechanisms most organizations have in place. AI systems do not just consume data, they generate it continuously, with outputs feeding back into core systems, creating feedback loops that build over time. Unfortunately, if not governed and managed well, errors can spread much faster than they are caught.

So does that mean fix the data first?

Not exactly, and this is where most of the conventional advice go wrong!

“Fix the data first” sounds reasonable but is, in practice, a path to indefinite delay. Data is never fully clean. Data estates are never fully unified. If you set “complete data readiness” as the precondition for AI investment, you have set a condition that will never be met, and you will emerge in three years with a tidy data warehouse and a significant AI capability gap.

The right framing isn’t fixing data first, then do AI. It’s what data does this specific business opportunity need for AI to amplify outcomes, and is that data good enough for this specific purpose, right now?

That’s a very different question. And it’s one that has an answer, not a perfect answer, but a workable one.

A more useful way to think about sequencing

The three-level model below gives business leaders a practical way to sequence their AI investments against their actual data readiness, without waiting for perfection and without ignoring the structural risk of bad data.

Level 1: Contained context, bounded data

What it targets: knowledge management, internal search and retrieval, document processing, code assistance, structured workflow automation.

Data consideration: AI operates against a curated, defined corpus; a document library, a product catalogue, a transaction set. The model doesn’t need to understand your entire enterprise; it needs to perform reliably within a bounded scope. Morgan Stanley curated library of 350k research documents made queryable through an internal AI assistant. And as a result, document retrieval efficiency moved from 20% to 80%.

Data infrastructure required:

  • Document versioning and lineage so the AI retrieves current content
  • Sensitivity tagging and access controls at the document level
  • Quality thresholds on ingestion to prevent corrupted or incomplete files from entering the system

Level 2: Cross-functional context, integrated data

What it targets: customer-facing AI (service, personalization), AI-assisted operational decisions (credit, supply chain, workforce), and process automation that crosses system boundaries.

Data considerations: AI now draws from multiple source systems simultaneously – CRM, ERP, transaction records, customer history. This is where small inaccuracies in data become systemic: AI systems at this layer don’t follow fixed paths through clean datasets, they select and combine pieces of information on the fly. Different chunks of the same document, retrieved in different contexts, can produce different and contradictory outputs.

Data infrastructure required:

  • Entity resolution and master data management so that “customer X” means the same thing across CRM, billing, and support systems
  • Metadata that travels with content through transformation, so that a customer transcript carries its date, source, and sensitivity tags when it reaches the model
  • Governance controls that apply at the retrieval layer, not just at the storage layer. This one is critical: a document classified as confidential at rest can still surface sensitive content in model outputs if retrieval-layer controls aren’t in place.

Level 3: Enterprise-wide context, foundational data

What it targets: AI at the core, processes rebuilt around AI capability, agentic systems that take autonomous or semi-autonomous actions across multiple enterprise systems, and strategic decision intelligence at the board or C-suite level.

Data considerations: agentic AI systems trigger actions across multiple systems, work at machine speed, and interact with sensitive data. As per IDC analysis 2025, by 2027, organizations that have not prioritised high-quality, AI-ready data are projected to suffer a 15% productivity loss when attempting to scale at this level. The reason is structural – AI outputs at this layer feed back into core systems, as summaries in CRM, as decisions in ERP and those outputs then influence future actions. A small inaccuracy in the data foundation compounds through every cycle.

Data infrastructure required:

  • Unified data models with explicit lineage from source through every derived artifact
  • Observability that monitors not just pipeline health but retrieval behavior, answer quality, and content freshness over time
  • Governance that extends from storage through to runtime, controlling how information is assembled in prompts and generated in outputs

The sequencing logic isn’t “data first, then AI.” It’s “prove value in the narrow cases, build the data infrastructure progressively as you move into higher-stakes applications, and use earned credibility to justify the foundational investment where transformation is genuinely worth it.”

Who will get this right?

They won’t be the ones who waited for perfect data, nor the ones who launched broadly without governance. They will be the ones who diagnosed their data readiness honestly, started in the domains where readiness was sufficient, built progressively, measured rigorously, and invested in the foundation as they went, using evidence from early wins to justify the harder, slower foundational work.

As I wrote in my previous article, the gap between pilot and scale is where the real work is. The data-foundation question is one of the main reasons that gap exists. Closing it isn’t a precondition for starting, but it is a precondition for scaling.

The chicken-and-egg problem has a resolution. It just requires being more specific about which egg you are trying to hatch first.

Leave a comment