Data & Technology
Data is the key for AI: Why AI projects fail at the data layer
60–95% of AI projects deliver no impact. Why governance, lineage, and consent determine success.
•
acceleraid Editorial Team
5 min read
01
Acquire
Recognize signals
02
Onboard
Control activation
03
Grow
Next Best Action
04
Retain
Reduce churn
05
Reactivate
Reclaim potential

Part 1 of 2 of our series "Data is the Key for AI" on data architecture as a prerequisite for AI in banking. Part 2 shows the bridge from legacy data warehouses to Snowflake.
The Models Are Not the Problem
When AI projects in banks fail, it is almost never because of the model. It is because of the data layer underneath. Gartner quantifies the extent precisely: 63 percent of surveyed organizations have no or unclear data management practices for AI, based on a survey of 1,203 data management leaders in July 2024 (Gartner, Lack of AI-Ready Data Puts AI Projects at Risk). The consequence is a prediction that should change every budget discussion in IT management: By 2026, 60 percent of AI projects not supported by AI-ready data will be abandoned (Gartner).
The pattern runs through the entire investment chain. At least 30 percent of generative AI projects will be abandoned after the proof of concept — not due to poor model performance, but because of poor data quality, inadequate risk controls, and escalating costs with unclear business value. This is particularly painful because, according to the same survey of 822 business leaders, transformative GenAI deployments cost between $5 million and $20 million (Gartner Press Release, 2024). And even projects that survive the pilot phase often deliver no commercial impact: According to the MIT-NANDA report "The GenAI Divide," cited by Fortune, 95 percent of companies studied with GenAI achieve no measurable P&L effect — only about 5 percent of pilot programs deliver quick revenue acceleration. According to the report, the core cause is not a model problem, but flawed enterprise integration and a "learning gap" between pilot and production operations (Fortune on the MIT-NANDA Report).

Where It Fails in Practice
Two current industry surveys confirm that data quality is not a side issue, but the central bottleneck. At Informatica, 57 percent of data leaders cite data reliability as a key hurdle to moving AI projects from pilot to production. Half name data quality as the top challenge when rolling out agentic AI, and 76 percent say their AI governance is no longer keeping pace with actual use — among 600 global data leaders surveyed from the US, UK/EU, and APAC (Informatica, CDO Insights 2026). At dbt Labs, 56 percent of respondents cite data quality as a problem — despite broad AI adoption, trust in one's own data remains the top priority for data teams (dbt Labs, The state of analytics engineering in 2025).
For banks, this is not a new insight, but an old lesson in a new guise. Even the Basel principles for effective risk data aggregation, BCBS 239, demanded after the financial crisis that institutions must be able to aggregate risk exposures completely, quickly, and accurately — because many globally systemically important banks were simply not capable of doing so (Bank for International Settlements, BCBS 239). Anyone who has internalized this discipline in risk reporting immediately understands why it now applies just as much to AI applications in marketing and sales: prediction models, churn scores, and personalization are only as reliable as the data base on which they are trained and applied.
Regulation Is Catching Up
What was previously implicit best practice becomes an explicit obligation with the EU AI Act. Article 10 requires for high-risk AI systems that training, validation, and testing datasets be relevant, sufficiently representative, and as free of errors and complete as possible — in relation to the intended purpose. It also requires a documented data governance practice for data origin, preparation steps such as annotation, labeling, cleaning, updating, and aggregation, for checking bias, and for identifying data gaps (European Commission, AI Act Service Desk, Article 10). For banks that support creditworthiness assessments, fraud detection, or personalized financial products with AI, this is not an abstract compliance exercise, but directly relevant to audits.
The practical consequence of both developments — market pressure and regulation — is identical: Without a governed, documented, and consistent data base, neither reliable prediction nor legally secure personalization can be operated. The following overview summarizes where the requirements apply.
Requirement | Purpose | Source |
|---|---|---|
Data Governance for AI Use Cases | Alignment of data quality with specific use cases | |
Active Metadata Management & Lineage | Traceability of data origin and processing | |
Documented data preparation, bias checking | Mandatory for high-risk AI under Art. 10 EU AI Act | |
Complete, fast, accurate risk data aggregation | Historical foundation for data governance in banking |
What This Means for System Architecture
The mentioned requirements — governance, lineage, bias checking, completeness — cannot be retrofitted into a legacy, fragmented data landscape. They require a System of Record that consolidates customer data from CRM, core banking, and card processing in real time, manages consent, makes the origin of each data point traceable (lineage), and protects personally identifiable information (PII) end-to-end. This is precisely the function that a Customer Data Platform with an integrated data governance layer performs: not as an additional reporting tool, but as a reliable foundation on which prediction engines and personalization can be built in the first place.
The acceleraid platform maps this System of Record in the CDP & Data Governance module — with real-time connectivity to CRM, core banking, and card processing, consent management, lineage tracking, and PII protection, hosted in Germany and designed GDPR-by-design (acceleraid Platform). For banks in the DACH region, this is a double advantage: regulatory requirements from BCBS 239 and the EU AI Act are not treated as a downstream compliance task, but are anchored in the data architecture from the very beginning (acceleraid Banking).
This connection between data quality and prediction is particularly clear in use cases such as churn forecasting: A model designed to detect customer churn 90 days in advance is only as good as the transaction and interaction data from which it learns. How a governed data base is translated concretely into reliable churn scores has been described in our post on retention and churn forecasts (CLM Retail Banking: Churn Prediction 90 Days).
The Framework for the Next Steps
For marketing, digital, and data leaders in banks, a pragmatic assessment framework can be derived from this before starting the next AI project:
First: Governance before model. Before a prediction or GenAI project is set up, it must be clarified which data sources are considered the System of Record, how consent is captured, and how lineage is documented. Second: Metadata and quality are not a one-time project, but an ongoing process — the five steps to AI-ready data described by Gartner (use-case alignment, governance requirements, metadata management, pipelines, quality assurance) must be repeated iteratively. Third: Regulatory requirements such as Art. 10 of the EU AI Act should not be planned as a subsequent audit hurdle, but as an architectural principle from the very beginning. Institutions that follow this sequence avoid precisely that 60 percent cancellation rate predicted by Gartner for AI projects without an AI-ready data base.
In the second and final part of this series, we show how this governed data base can be implemented technically — specifically at the transition from legacy data warehouses to cloud platforms like Snowflake, without banks having to undergo a risky big-bang migration.
Illustration: AI-generated. AI-supported content: In creating our posts, we use AI technologies and automated agents, including those from Microsoft, Google, OpenAI, Anthropic, and other providers. Topics, professional direction, and final approval rest with our team.
Further Insights
We use cookies 🍪
Strictly necessary cookies (e.g. Pipedrive forms) remain active. With your consent, we also use Google Analytics (analytics) and Leadfeeder (visitor identification). Learn more in our Privacy Policy.