Finland has ordered Google to suspend construction of two large new data centres intended to power its artificial intelligence projects ([1]) ([2]). On 7 October the country’s environmental regulator (LVV) instructed Google’s local subsidiary to halt forest clearing and groundworks at the sites until a full environmental impact assessment is completed ([3]) ([4]). The move follows reports that Google cleared more than 300 hectares of boreal forest without the mandatory review ([5]). The data centres, part of a planned €13 billion AI infrastructure investment announced by Google in September, are now delayed at least until late October ([6]) ([7]).
The Finnish intervention underscores a new kind of compliance risk facing AI data strategy. European governments have welcomed data centre investments as digital infrastructure becomes critical to AI ambitions, but they are starting to weigh environmental and power grid impacts against economic benefits ([8]) ([9]). Google said it will work with Finnish authorities to address concerns and remains committed to its plans ([10]), but the halt highlights that building the physical side of “AI-ready” data - from cloud servers to networking - can be subject to costly regulatory delays. Tech leaders should factor these emerging constraints into their AI infrastructure strategies, especially in regions with stringent environmental rules.
Even where infrastructure is in place, many organisations are discovering that their data itself is not ready for AI. New research reveals a persistent gap between AI ambitions and data reality. Dun & Bradstreet’s latest global “AI Momentum” survey of 10,000 businesses found that while more than three-quarters of enterprises now report some ROI from AI, only 6% say their enterprise data is “fully ready” to support AI at scale ([1]) ([2]). The majority describe their data as only “partially” or “mostly” ready ([3]). Similarly, a joint study by data integrity firm Precisely and Drexel University found 43% of data and analytics leaders cite data readiness as their biggest barrier to aligning AI with business goals ([4]). In other words, most organisations remain far from the high-quality, well-governed, integrated data foundations needed to drive AI reliably across the enterprise.
Industry experts warn that this readiness shortfall is already stalling AI projects. Nearly half of respondents in an international poll said their AI initiatives had been indefinitely delayed or abandoned due to data problems, according to an IBM summary of the findings ([5]). In a pointed recommendation to companies, IBM’s data and AI team advises: “stop optimising the model and fix the data” ([6]). The clear implication is that, for AI to deliver scalable value, business leaders must first invest in improving data quality, governance and accessibility - or risk seeing promising AI pilots falter.
Another growing challenge is access to external data for AI applications. TechCrunch this week highlighted that many websites and content providers are now blocking automated AI web crawlers, effectively putting swathes of online data off-limits for training and “agentic AI” systems. Major social platforms like Reddit and X (formerly Twitter) have already restricted free data access through their APIs, and several news publishers - including the New York Times - have updated terms or robots.txt files to forbid scraping for AI model training ([1]). Even as OpenAI restored ChatGPT’s web browsing feature in late September, AI tools are entering a more closed internet where valuable data is either paywalled or protected by content owners.
These developments mean organisations can no longer assume that data from across the web will be readily available to fuel their AI models. Data licensing is becoming a strategic issue: companies may need to strike agreements or pay for high-quality external datasets, whether for training generative AI or powering real-time AI assistants. Leaders should evaluate their AI data supply chain and content partnerships to ensure they are prepared for a future where the “open” web is increasingly walled off. Crafting a robust data strategy - including first-party data governance and ethical data sourcing - is essential to mitigate these risks.