A wave of new industry research is underscoring a stark reality in the AI race: the top performers are pulling ahead not because they have better algorithms, but because they have better data foundations. Gartner’s latest analysis finds that companies successfully scaling AI spend up to four times more (as a percentage of revenue) on core data quality, governance, and infrastructure than those still struggling ([1]). This data-centric investment gap translates directly into results – organisations with mature data setups see up to 65% greater improvements in revenue growth and cost reductions from their AI initiatives compared to peers without strong data foundations ([2]).
These findings echo other recent studies indicating that only a small fraction of firms (around 5%) are realising significant AI ROI, while the majority remain mired in pilot projects that aren’t scaling ([3]). The key differentiator is that AI leaders treat data as a strategic asset. They break down silos, ensure data is clean and well-governed, and integrate their AI solutions deeply into business processes to drive value. Laggards, by contrast, often chase the latest AI tools or models without first solidifying their data foundation – a mistake that leads to underperforming models and disappointing results.
In short, having the most advanced model means little if it’s not fuelled by high-quality, well-integrated data. Enterprises that prioritise robust data architecture and data preparation are finding that AI can truly transform their operations and outcomes. Those that neglect these essentials risk falling behind, regardless of how impressive their algorithms might be.
Despite the excitement around generative AI, data issues are emerging as a primary bottleneck to success. A recent S&P Global survey of over 1,000 IT and business leaders revealed a sharp rise in AI project attrition: 42% of companies have now abandoned most of their AI initiatives before they ever hit production, up from just 17% the year prior ([1]). On average, these organisations report scrapping 46% of their AI proof-of-concepts as they struggle to move from pilot to scale. Experts have been frank about the reason – an MIT study found up to 95% of AI projects fail to meet expectations, and the chief culprit isn’t the algorithms at all. It’s the data: models too often are fed with poor-quality, incomplete or irrelevant information and thus never deliver useful results ([2]).
These data shortcomings take many forms. Corporate data is frequently fragmented across silos or locked in legacy systems, depriving AI of the consistent context it needs to generate accurate insights ([3]). Key data may be incomplete, inconsistent, or simply inaccurate, undermining model predictions and eroding trust ([4]). In many firms, data accessibility and integration remain challenges – data scientists still report spending the bulk of their time cleaning and organizing data instead of building AI solutions. All of this means that even the most powerful AI models will produce biased or unhelpful outputs if the input data is a mess.
Recognising this, industry leaders are urging a refocus on data fundamentals. An influential industry round-up this week put it bluntly: “The real challenge of enterprise AI is no longer the capabilities of the models themselves but the readiness of the organization to manage, govern, and integrate data effectively” ([5]). In fact, analysts predict some high-profile AI initiatives could falter in 2026, prompting a pause for reflection and new investments in “AI-specific” data governance to prevent disasters before they happen ([6]). In short, robust data governance and quality management have moved from a back-office IT concern to a frontline factor deciding AI success.
Amid these cautionary tales, companies with strong data foundations are demonstrating how proprietary data can become a powerful competitive moat in the AI era. Just this week, Thomson Reuters reported a 10% jump in first-quarter revenue, attributing the gain in large part to successful AI-driven products built on the firm’s vast troves of legal and financial data ([1]). Earlier in the year, the company’s stock price soared nearly 14% in a single day after it announced a new AI-powered CoCounsel legal assistant alongside a partnership with AI pioneer Anthropic ([2]). Investors saw this as proof that Thomson Reuters’ unique data assets and specialised AI tools are not just resilient in a turbulent market – they are fueling the company’s growth and outpacing competitors ([3]).
Notably, the nature of these "data moats" is evolving. As one industry expert observed, the advantage of proprietary data is shifting from sheer volume to a rich context layer built on metadata, governance, and domain-specific knowledge ([4]). In practice, this means leading organisations are investing in capabilities like detailed data catalogs, lineage tracking, domain ontologies and knowledge graphs, and even vector databases that enable retrieval-augmented generation (RAG) workflows ([5]). By layering context and oversight on top of their raw data, companies make their information far more usable for AI systems – which in turn allows AI to deliver deeper, real-time insights and more trustworthy decisions.
Concerns about data ownership and security are also shaping AI strategies. In sectors like finance, healthcare and defence, a growing number of firms now choose to fine-tune AI models on their own infrastructure or within virtual private clouds, to ensure sensitive customer information and intellectual property never leave their control ([6]). This trend reflects a broader realisation: in the age of generative AI, a company’s data is not just an operational resource, but a core piece of intellectual property. Keeping that data both protected and primed for AI use is becoming a strategic imperative.
The past week also saw major moves to modernise data architecture for AI. A prime example is SAP’s announcement of its acquisition of Dremio, an open lakehouse platform, which will help customers unify SAP and non-SAP data on a single platform using the open-source Apache Iceberg format ([1]). By integrating Dremio into SAP’s Business Data Cloud, the enterprise software giant aims to eliminate data fragmentation and give AI “agentic” systems high-performance access to governed, real-time data across the business ([2]). As SAP’s CTO put it, enterprise AI projects don’t stall for lack of sophisticated models, but because “the data isn’t ready for AI agents” ([3]) – this deal is about removing that obstacle.
Across the industry, data platform providers are similarly doubling down on "AI-first" architecture. Modern lakehouse designs, popularised by cloud data companies like Databricks and Snowflake, blend the flexibility of data lakes with the governance and speed of data warehouses – a formula now being adopted even by legacy vendors like SAP ([4]). Major cloud platforms are rolling out new capabilities to feed AI models with up-to-date information, moving from batch-processed data to continuous data pipelines and streaming analytics so that AI systems are never out of sync with the business’s latest context ([5]). In addition, the rise of vector databases (purpose-built for storing and retrieving the embeddings that power AI knowledge queries) is giving enterprises a way to provide “memory” to generative AI applications. One sign of this trend: in March, an open-source vector search firm secured $50 million in funding to scale its technology for enterprise AI deployments ([6]). It’s clear that tools like these – once seen as niche – are quickly becoming core parts of the AI infrastructure stack.
Common to all these innovations is an emphasis on openness and interoperability. Companies are increasingly favouring open standards and cloud-agnostic data architectures to avoid being locked into any one vendor’s ecosystem ([7]). By building flexible data foundations (for example, adopting open table formats like Iceberg and making data accessible across multiple cloud environments), enterprises can ensure their AI systems remain adaptable. This future-proofing means that as new AI models and techniques emerge, they can be integrated with the organisation’s entire corpus of trusted data, without having to re-engineer or migrate the underlying data layer.
All these developments are unfolding under intensifying regulatory scrutiny. Data protection authorities are making it clear that trustworthy AI requires strong data governance at every step. In Europe, regulators have already fined companies like Clearview AI for illegally scraping personal images to train face-recognition algorithms ([1]). Italy’s watchdog recently imposed a €5 million penalty on the American company behind the Replika chatbot for mishandling users’ personal data in training an AI system ([2]). These actions send a message: companies must embed privacy-by-design into their AI data pipelines, or risk severe consequences.
Meanwhile, governments are moving to harden data accountability through new laws. The European Union’s AI Act – along with its Digital Services Act, Data Act, and other regulations – is pushing organisations to adopt “governance-by-design,” building transparency, fairness, and security into data processes from the ground up. In 2026, with the EU’s AI regulatory framework coming into force, the cost of having disorganised or biased "messy data" is no longer just an efficiency issue – it’s a legal liability ([3]). Firms will need rigorous documentation of data provenance, quality metrics, and human oversight for high-risk AI systems, or they could face fines and reputational damage.
Another rising priority for CDOs and CTOs is ensuring data and AI sovereignty. Leaders across the globe are recognising that controlling their data – and the infrastructure their AI runs on – is now a strategic necessity, not an afterthought. In one global survey, only 11% to 27% of executives (depending on country) currently see sovereignty over AI and data as "mission-critical," but within three years a majority in major economies are expected to hold that view ([4]). Notably, the number-one driver behind this push (cited at twice the rate of any other factor) is the need to break data out of silos and maintain direct control, rather than geopolitical concerns ([5]). In practice, this means more organisations are considering localised cloud infrastructure, strict data residency controls, and open-source AI models that can be run on in-house data – all in an effort to balance AI innovation with compliance, security, and strategic control of data.