Data & Technology

Data is the key for AI: why projects fail at the data layer

60–95% of AI projects miss their goal. Why governance, lineage and consent decide success.

acceleraid Redaktion

6 min read

Customer Lifecycle Management

Customer Lifecycle Management

Customer Lifecycle Management

01

Acquire

Signale erkennen

02

Onboard

Aktivierung steuern

03

Grow

Next Best Action

04

Retain

Churn reduzieren

05

Reactivate

Potenziale zurückholen

Daten → KI-Score → Trigger → Kanal → Feedback

Daten → KI-Score → Trigger → Kanal → Feedback

Abstract illustration of a data layer as the foundation for AI systems in banking

Part 1 of 2 in our series "Data is the Key for AI" on why data architecture is the precondition for AI in banking. Part 2 covers the bridge from legacy data warehouses to Snowflake.

The model is rarely the problem

When AI projects in banking fail, the model is almost never the reason. The data layer underneath it is. Gartner puts a precise number on the scale of the issue: 63 percent of organizations surveyed have no or unclear data management practices for AI, based on a survey of 1,203 data management leaders conducted in July 2024 (Gartner, Lack of AI-Ready Data Puts AI Projects at Risk). The resulting forecast should reshape every budget conversation in IT leadership: by 2026, 60 percent of AI projects not backed by AI-ready data will be abandoned (Gartner).

The pattern repeats across the entire investment chain. At least 30 percent of generative AI projects are abandoned after the proof-of-concept stage — not because of weak model performance, but due to poor data quality, inadequate risk controls, and escalating costs against unclear business value. That is especially costly given that transformative GenAI deployments, according to the same survey of 822 business leaders, run between 5 and 20 million US dollars (Gartner Press Release, 2024). And even projects that survive the pilot phase often deliver no real economic effect: according to the MIT NANDA report "The GenAI Divide," as reported by Fortune, 95 percent of the companies studied see no measurable P&L impact from GenAI — only around 5 percent of pilot programs deliver rapid revenue acceleration. The report's core finding is that this is not a model problem but flawed enterprise integration and a "learning gap" between pilot and production (Fortune on the MIT NANDA report).


Five studies show AI projects fail at the data layer, not the model

Where it actually breaks down

Two recent industry surveys confirm that data quality is not a side issue but the central bottleneck. In Informatica's research, 57 percent of data leaders cite data reliability as the key barrier to moving AI projects from pilot to production. Half name data quality as their top challenge in rolling out agentic AI, and 76 percent say their AI governance is not keeping pace with actual usage — based on 600 global data leaders surveyed across the US, UK/EU, and APAC (Informatica, CDO Insights 2026). At dbt Labs, 56 percent of respondents cite data quality as a problem — despite widespread AI adoption, trust in data remains the top priority for data teams (dbt Labs, The state of analytics engineering in 2025).

For banks, none of this is new — it is an old lesson wearing new clothes. The Basel Committee's principles for effective risk data aggregation, BCBS 239, already required after the financial crisis that institutions be able to aggregate risk exposures completely, quickly, and accurately — precisely because many globally systemic banks could not (Bank for International Settlements, BCBS 239). Anyone who has internalized that discipline for risk reporting immediately understands why it now applies equally to AI applications in marketing and sales: prediction models, churn scores, and personalization are only as reliable as the data foundation they are trained and applied on.

Regulation is catching up

What used to be implicit best practice is becoming an explicit requirement under the EU AI Act. Article 10 requires that for high-risk AI systems, training, validation, and testing data sets be relevant, sufficiently representative, and, to the best extent possible, free of errors and complete for their intended purpose. It also demands documented data governance practices covering data provenance, preparation steps such as annotation, labeling, cleaning, updating, and aggregation, bias examination, and the identification of data gaps (European Commission, AI Act Service Desk, Article 10). For banks using AI to support creditworthiness assessments, fraud detection, or personalized financial products, this is not an abstract compliance exercise — it is directly subject to audit.

The practical takeaway from both developments — market pressure and regulation — points in the same direction: without a governed, documented, and consistent data foundation, neither reliable prediction nor legally sound personalization is possible. The table below summarizes where these requirements bite.

Requirement

Purpose

Source

Data governance for AI use cases

Aligning data quality with concrete applications

Gartner

Active metadata management & lineage

Traceability of data origin and processing

Gartner

Documented data preparation, bias checks

Mandatory for high-risk AI under Art. 10 EU AI Act

European Commission

Complete, fast, accurate risk data aggregation

Historical foundation for data governance in banking

BIS, BCBS 239

What this means for system architecture

These requirements — governance, lineage, bias checks, completeness — cannot be retrofitted onto a fragmented, organically grown data landscape after the fact. They require a system of record that consolidates customer data from CRM, core banking, and card processing in real time, manages consent, makes the origin of every data point traceable (lineage), and consistently protects personally identifiable information (PII). That is exactly the function a customer data platform with an integrated data governance layer performs: not as an additional reporting tool, but as the reliable foundation prediction engines and personalization need before they can operate at all.

The acceleraid platform implements this system of record through its CDP & Data Governance module — with real-time connections to CRM, core banking, and card processing, consent management, lineage tracking, and PII protection, hosted in Germany and designed GDPR-by-design (acceleraid Platform). For banks in the DACH region, this delivers a dual benefit: the regulatory requirements from BCBS 239 and the EU AI Act are not treated as a downstream compliance task but are embedded in the data architecture from the outset (acceleraid Banking).

This connection between data quality and prediction is especially visible in use cases like churn forecasting: a model designed to detect customer attrition 90 days in advance is only as good as the transaction and interaction data it learns from. We described how a governed data foundation translates into reliable churn scores in our article on retention and churn prediction (CLM Retail Banking: Churn Prediction 90 Days).

A framework for the next steps

For marketing, digital, and data leaders in banking, this translates into a pragmatic checklist to run through before launching the next AI project.

First: governance before model. Before a prediction or GenAI project is set up, it must be clear which data sources count as the system of record, how consent is captured, and how lineage is documented. Second: metadata and quality are not a one-time project but an ongoing process — Gartner's five steps toward AI-ready data (use-case alignment, governance requirements, metadata management, pipelines, quality assurance) need to be repeated iteratively. Third: regulatory requirements such as Article 10 of the EU AI Act should be built in as an architectural principle from day one, not treated as a late-stage compliance hurdle. Institutions that follow this sequence avoid exactly the 60 percent abandonment rate Gartner forecasts for AI projects without an AI-ready data foundation.

In the second and concluding part of this series, we show how this governed data foundation can be implemented technically — specifically at the transition from legacy data warehouses to cloud platforms like Snowflake, without banks having to take on a risky big-bang migration.

Illustration: AI-generated. AI-assisted content: We use AI technologies and automated agents in the creation of our articles, including from Microsoft, Google, OpenAI, Anthropic and other providers. Topics, editorial direction and final approval remain with our team.

We use Cookies 🍪

Strictly necessary cookies (e.g. Pipedrive forms) remain active. With your consent we also use Google Analytics (analytics) and Leadfeeder (visitor identification). More in our Privacy Policy.

Decline

Decline

Accept all

Accept all