Data & AI Series: Everyone has an AI strategy, but how many can claim they have the outcomes?

Businesses are spending more on AI than ever before and a surprising number of them are struggling to explain what they’re getting back.

It’s not a fringe observation; it’s becoming one of the more uncomfortable conversations in boardrooms and leadership teams across industries. It’s an attempt to look honestly at where the equation currently stands, where it’s going, and what a sensible approach actually looks like for businesses trying to navigate it.

Spend on AI is growing!

Global enterprise AI spend is growing at a pace that would have seemed implausible five years ago. Cloud providers are reporting record demand. Model providers are raising at valuations that assume AI becomes infrastructure for everything. Boards are asking leadership teams to demonstrate an AI strategy.

The pressure to invest is enormous and the pressure to show results is equally enormous. And the gap between those two pressures is where a lot of confused spending is happening.

Some of it is producing real value. But a significant portion of enterprise AI spend today is going into pilots that don’t scale, tools that get adopted inconsistently, and capability-building that hasn’t yet connected to any measurable business outcome. The investment is running ahead of the operating model changes needed to make it land.

Although the value promise is strong, AI is not cheap!

There’s a seductive narrative in the market right now: compute costs are falling, models are commoditizing, and therefore AI is getting more affordable. But it’s a half-truth.

The cost of a token has dropped dramatically over the last two years. The frontier models that cost significant money to access in 2023 have been matched or surpassed by models that cost a fraction as much today. But token cost is not the same as the cost of AI in a business context.

The real costs, the ones that don’t appear in the per-token pricing, are the integration work, the data infrastructure required to make models useful against real business context, the governance and oversight frameworks, the change management needed for adoption, the ongoing human review that responsible deployment requires, and the accumulated technical debt when organizations move fast without thinking about how these systems will be maintained and monitored over time. Unfortunately, this everything else is a very large number compared to the cost of tokens!

If you still proceed, what happens when it goes wrong?

Air Canada deployed a customer service chatbot that told a grieving passenger he could buy a full-fare ticket and claim a bereavement discount retroactively within 90 days of travel. That wasn’t the airline’s policy. When Air Canada refused to honor what its own chatbot had promised, the passenger sued. The airline’s defense that the chatbot was a “separate legal entity responsible for its own actions” was dismissed by the British Columbia Civil Resolution Tribunal as a “remarkable submission.” Air Canada was found liable, ordered to pay damages, and the chatbot was quietly removed from its website shortly after. The case is good example of AI accountability.

McDonald’s spent three years and significant investment building an AI-powered voice ordering system, deployed across more than 100 drive-thru locations in the US. The technology struggled with accents, background noise, and basic corrections. Customers began filming the results, orders ballooning to 260 chicken nuggets, bacon appearing on ice cream and posting them on TikTok. The videos spread. In June 2024, McDonald’s pulled the plug entirely, confirming it was ending its partner and removing the technology from all testing locations.

Amazon spent years building an AI recruiting tool trained on a decade’s worth of historical résumés. Because most of those résumés came from men, reflecting the demographics of the tech industry at the time, the model taught itself to prefer male candidates. It began downgrading applications that included the word “women’s” and penalized graduates of all-women’s colleges. When the bias was identified, Amazon’s teams tried to fix it but couldn’t build sufficient confidence that bias wasn’t persisting in other, less visible ways. The project was scrapped. Nobody had intended to build a discriminatory hiring tool.

These aren’t edge cases from organizations that were careless. These are examples from companies with real resources, real governance teams, and real reputations, who moved faster than their control frameworks could keep up with.

The common thread isn’t the AI failing. It’s the absence of the right human oversight at the right points in the process. AI systems behave in ways their designers don’t fully anticipate. Without the governance structures to catch that early and the organizational clarity about where human judgment remains essential, the exposure compounds quietly until it surfaces publicly.

The reputational and legal costs of a visible AI failure are significant. The internal costs of undetected AI-driven errors in decisions, in outputs, in processes are harder to see and potentially much larger.

Where the value is actually showing up?

Despite all of this, AI is producing genuine, measurable value in specific domains. And there are enough real, named examples now to say something concrete about where it’s landing.

Morgan Stanley’s internal knowledge at scale: Morgan Stanley’s wealth management division manages a research library of over 350,000 documents, decades of institutional knowledge that financial advisors couldn’t practically navigate during client conversations. In 2023, the firm rolled out an internal AI assistant built on GPT-4 that made that entire knowledge base queryable in real time. Document retrieval efficiency went from 20% to 80%. Query times that previously took 30 minutes or more dropped to seconds. 98% of financial advisor teams actively use the tool, which in a firm of that size and that degree of regulatory caution is not a number that happens by accident. Morgan Stanley then extended this to “AI Debrief”, a tool that turns client meeting recordings into structured notes, action items, and draft follow-ups, automatically logged into their CRM.

GitHub Copilot, code generation and developer productivity case: Microsoft and Accenture ran a controlled study across 1,600+ developers. The group using GitHub Copilot completed tasks 55% faster than those who didn’t, a finding consistent with independent research from MIT and replicated across multiple enterprise deployments. The productivity gains are most pronounced for structured, repetitive coding tasks, and less dramatic for complex architectural work, which is actually the honest version of what AI-assisted development looks like in practice.

Walmart’s supply chain and operational optimization. Walmart’s AI-powered route optimization eliminated 30 million unnecessary delivery miles in a single fiscal year, avoiding 94 million pounds of CO2 emissions in the process, the work that won the Franz Edelman Award in 2023, widely regarded as operations research’s most rigorous recognition for real-world impact. The company has since commercialized the underlying technology as a SaaS product for other businesses. Separately, Walmart’s CEO Doug McMillon confirmed in the company’s August 2024 earnings call that generative AI had improved over 850 million product catalogue data points, a task he noted would have required 100 times the headcount to complete manually.

The pattern across all three isn’t AI operating autonomously. It’s AI absorbing the volume work, the retrieval, the drafting, the routing, the data processing, so that the humans in the loop can concentrate on the judgment, the relationship, and the decision. The organizations getting the best results have been deliberate about that distinction.

So, what is the right approach to generate more value while managing risk?

This is the single biggest question for businesses who are still on the boundary line trying to figure out the way forward! Should we go a full business reinvention with AI at the core or take an incremental journey with incremental gains by putting AI on the top of the current business?

AI at the core means embedding it into processes where it changes how the work is fundamentally done, where the operating model is rebuilt around the capability, not just augmented by it. This is higher risk, requires more infrastructure, demands real governance, and takes longer to get right. But when it works, it compounds. It produces structural advantage, not just efficiency – Telstra’s Data & AI JV with Accenture is a fantastic example of this approach.

AI on top means applying it as a layer over existing processes, accelerating, augmenting, assisting, without redesigning the underlying work. This is faster to deploy, lower risk, and easier to measure. It promises quick returns but doesn’t fundamentally change what the business is capable of. Most of the pilots, small-scale AI operationalization initiatives across businesses are of this nature. The biggest risk in this approach is if not governed or managed well, a heavily distributed, decentralized approach to AI application across the business will increase the very risks we discussed in the previous section.

The mistake most organizations are making is trying to do the first without the foundations or getting stuck in the second and calling it transformation.

A more honest approach acknowledges that most processes should start with AI on top, prove the value, build the operating capability, understand the failure modes, and then selectively, deliberately, move toward AI at the core in the domains where the structural advantage is worth the investment and the governance is genuinely in place.

Not everything needs to be rebuilt with or around AI, some things do. Knowing which is which and being disciplined about the sequencing is where the real strategic work lies.

How a cautious but committed approach could look like

We are all aware, there is no waiting game now. The businesses that disengage from AI investment now will face a genuine capability gap in three to five years that will be costly and disruptive to close. But there is a meaningful difference between being committed and being reckless.

Start with problems: The question isn’t “where can we apply AI?” It’s “what are the business outcomes we need, and is AI the most effective path to them?” or “what is the structural solution to it and can AI enhance the outcome both in terms of time to outcome and value of it?”

Build the governance infrastructure first without waiting for the first incident: Determine explicitly, in writing, where human oversight is non-negotiable. Where does an AI output require human review before it acts? Where does accountability sit when something goes wrong?

Be honest about what your data is ready for: AI models are only as useful as the data they operate against. Many organizations are investing in AI capability that their data infrastructure can’t yet support. Closing that gap is a higher-priority investment than the AI layer itself.

Sequence deliberately: Pick two or three domains where the value case is clear, the governance is manageable, and the risk of failure is contained. Get that right, measurably right. Use them to build internal confidence, capability, and credibility for the next wave of investment. Breadth without depth produces noise. Depth in the right places produces evidence.

In summary …

The commercial equation for AI doesn’t fully work yet for most businesses, not because the technology isn’t capable, but because the organizational leadership willingness, measured risk-taking decision-making mindset, and the infrastructure to deploy it responsibly and at scale is still evolving. Even if some small moves are made, the costs are front-loaded and often underestimated. The returns are real but concentrated and slower to materialize than the investment cycle that’s funding them.

The businesses that will look back on this period as having gotten it right are probably not the ones who moved fastest or spent most. They’re the ones who chose carefully, governed rigorously, measured honestly, and built the internal capability to make AI a sustained part of how they work, rather than a series of pilots that never quite scaled or a series of licenses given to workers that never quite impacted anything significantly.

That gap between pilot and scale is where the real work is. And it’s still largely unsolved.

Leave a comment