Menu Schließen

Data Quality for AI: These 4 Areas will determine the success of your project

BI2run - Data Quality for AI

The AI pilot has been approved, the tool has been selected, and expectations are high. And then the system delivers numbers that no one in the controlling department trusts. The problem rarely lies with the model. According to industry reports, 70 to 85 percent of all AI projects fail to meet their goals, and the most common reason is data issues, not the technology.

Data quality describes how complete, consistent, and conceptually sound your data is. For AI applications, there’s a second layer to consider: data acquisition and preparation – that is, the process of moving data from disparate source systems to a foundation that an AI system can rely on. Checking this foundation before the project begins saves you from the most expensive form of rework: making corrections while the system is already in operation.

Why Does the Data Set Determine the Success of AI in Controlling?

Large language models have significantly lowered the barrier to entry for AI in controlling. Today, a model can understand natural language, explain discrepancies, and generate comments. What hasn’t changed: The model can only work with the data it is given. Different definitions of the same KPI, missing values, or inconsistent master data lead to unreliable results – even in AI applications.

The figures speak for themselves. Gartner estimates the cost of poor data quality at an average of $12.9 million per year per company. Experience from projects also shows that 70 to 80 percent of project time is spent on data preparation, not on models or prompts. Based on our project experience, data quality remains one of the biggest drivers of effort, even with LLMs.

That’s not bad news. It simply means that the key to successful AI lies not in comparing tools, but in the data strategy – and that can be systematically evaluated.

Bi2run - Nahaufnahme Mitarbeiterin am Laptop

What is the key question that must be addressed before every AI project in controlling?

The wrong first question to ask in an AI project is which tool to use. The right first question is: Which recurring process currently takes up the most manual time? Monthly commentary, variance analyses, data consolidation for reporting. That’s where AI creates the greatest value, because the benefit can be measured directly in hours saved.

Once the process has been identified, the second decision follows, and that concerns the order in which the work is done. We consistently work in two phases.

Where does your data foundation stand?

Is your data ready for an AI use case? We’ll give you an honest assessment – even if it means: data work first, AI second.

Schedule a call →

Phase 1: Data Preparation and Data Collection

First, the foundation. The data for the selected use case is checked, cleaned, and consolidated where it is missing or inconsistent.

Phase 2: The AI Model

Then there’s the technology. Once the foundation is solid, the AI model is built on top of it – not the other way around.

We never reverse this order. An AI model built on an untested data set delivers quick results, but not reliable ones.

Which 4 test areas are included in Phase 1?

The following four sections cover Phase 1 in its entirety. They serve as a self-assessment: Any question you can’t answer clearly indicates an area that needs attention before Phase 2 begins.

1. Data quality: Complete, consistent, and conceptually sound?

Check whether the same KPI is defined consistently across the board, whether any values are missing, and whether master data is consistent across systems. A revenue metric that is calculated differently in Sales than in Controlling will produce contradictory results in any AI application. Business transparency means that a person can explain how a number is derived. If they can’t, neither can the AI.

2. Domain Knowledge: Are KPIs and business rules documented?

In many companies, KPI definitions, technical terms, and business rules exist in people’s minds, not in documents. The test is simple: Would a new employee understand the meaning of the data without additional explanations? What applies to a new colleague applies just as much to any AI. LLMs come with a wealth of general knowledge; company-specific knowledge generally needs to be added in a targeted manner, for example through digital knowledge management.

3. Metadata: Is the data described in a way that is understandable to both humans and AI?

Metadata is the description behind the data. What does this field mean? What unit is this value in? Since when has it been recorded this way? Without this description, every AI query has to guess what a column represents. With a well-maintained metadata layer, a column of numbers becomes a dataset that an AI system can correctly classify. We explain why this is so crucial in the article “Why Metadata Is the Key to Successful AI Use.

4. Data Access: Who is allowed to use what, and how easy is it to access?

There are two sides to data access. The first is governance: Which data may be used for analytics and AI applications, who is responsible, and how are permissions and data protection regulated? The second is technology: Can the data be retrieved automatically via a defined interface, or does it require a manual export each time? Both aspects belong in the data strategy, not in the after-the-fact fixes.

The 4 fields translated to your company

Unsure about two or more fields? In the readiness check, we’ll go through them together and prioritize what needs to be done before phase 2.

Request readiness check →

How do you assess your AI readiness in practice?

The two-phase model and the four areas come together to form a concrete roadmap. Here’s how to proceed:

  1. First, select a use case. Choose a recurring process that involves a significant amount of manual work, such as the monthly review. The use case determines which data is relevant for Phase 1.
  2. Check data quality. Ensure completeness, consistency, and business-related traceability. Document gaps instead of silently ignoring them.
  3. Document domain knowledge. Write down the KPI definitions and business rules for the use case. This document will later serve as the business context for your AI system.
  4. Add metadata. Describe the fields, units, and source of the data so that they are understandable without further inquiry – for new employees as well as for the AI.
  5. Clarify data access. Specify in writing which data the AI system is permitted to use, and verify whether this data can be made available automatically. Phase 2 begins only once both of these are in place.

After following these five steps, you’ll know whether Phase 1 is complete or where there’s still data work to be done. Both outcomes are useful. It only gets expensive if you skip Phase 1 and start building the model right away.

BI2run- Business meeting

What do the numbers say about data quality in AI projects?

  • According to industry reports, 70 to 85 percent of AI projects fail to meet their goals, with data issues being the most common cause (Source: AIMultiple, 2026).
  • Poor data quality costs a company an average of $12.9 million per year (Source: Gartner).
  • Experience shows that 70 to 80 percent of project time is spent on data preparation rather than modeling (based on experience from customer projects, including quality.de, 2026).
  • Over 70 percent of companies cite data quality as the biggest hurdle for AI projects (Source: Industry studies, 2025).

Glossary: What Terms Should You Know?

TermDefinition
Data QualityThe degree to which data is complete, consistent, up to date, and understandable from a business perspective.
Data PreparationThe process of cleaning, standardizing, and making raw data from source systems usable for analytics or AI.
MetadataDescriptive information about a data field, such as meaning, unit, and origin, needed by both people and AI to interpret it correctly.
Data StrategyA company-wide framework defining how data is collected, managed, and used, including its use in AI.
AI ReadinessAn assessment of whether data, processes, and the organization are ready for productive AI use.
Data GovernanceRules and responsibilities for data access, data protection, and data maintenance within a company.
Two-Phase ModelAn approach in which phase 1 completes data preparation and sourcing before phase 2 sets up the AI model.

FAQ: Frequently Asked Questions About Data Quality for AI

Why isn’t a powerful LLM enough to compensate for data issues?

An LLM interprets the data it receives, but it cannot correct missing values, conflicting KPI definitions, or inconsistent master data. Based on poor data, it merely formulates incorrect answers that sound convincing. Data quality therefore remains one of the biggest cost drivers, even with LLMs.

How much data preparation is needed before the first AI use case?

Only as much as the specific use case requires – no more. If you just want to automate monthly commentary, you don’t need to clean up the entire data warehouse. Limiting the scope to the use case keeps Phase 1 predictable.

What are the key components of a data strategy for AI applications?

At least four key components: defined quality criteria, documented domain knowledge, a well-maintained metadata layer, and regulated data access, including data protection. Deciding which data may be used for AI should be part of the strategy, not the implementation.

How can you tell if your master data isn’t sufficient for AI?

Typical indicators include customers or products that are tracked under different keys in various systems, manually maintained mapping tables, and reports that show different totals depending on the source. Each of these indicators creates the same inconsistencies in AI applications as in traditional reporting.

What is metadata, and why isn’t data quality alone enough?

Data quality indicates whether a value is correct. Metadata explains what that value actually means, what unit it is in, and where it comes from. A value that is correct but lacks metadata is still difficult for AI to interpret. That is why both fields belong in Phase 1 – not just one of them.

Why is BI2run the right partner for data preparation and AI?

We’ve been building data platforms for planning and controlling for years – from IBM Planning Analytics to Power BI – and today we connect them directly to AI assistants. That’s why we understand both sides: the data work that precedes every AI project, and the use cases that justify the effort. That’s why we follow the same two-phase model for every client. Phase 1– data preparation and data acquisition – lays the foundation. Phase 2 – the AI model – brings in the technology. We never reverse this order.

And if our assessment shows that your data foundation isn’t yet sufficient for the desired use case, we’ll tell you so. In that case, we’ll provide you with a prioritized plan for Phase 1 instead of an AI project that’s bound to stall later on.

Free · Initial Consultation

AI Readiness Check for Your Controlling

Let’s find out together where your data foundation stands and which use case makes the most sense as a starting point. One initial conversation is enough for a clear assessment.

Book a meeting →

Share article:

LinkedIn
WhatsApp
Facebook
Email

More articles

Any questions? Our experts look forward to your call!