Menu Schließen

Introducing AI in the company: Why pilot projects fail – and what determines their path to productive operation

BI2run - AI in Business

The pilot is up and running, the demo was convincing, and yet no one on the steering committee is proposing a full rollout. This is exactly where most AI projects come to an end. According to an MIT study from the NANDA project, 95 percent of generative AI pilots in companies have no measurable impact on the bottom line. The reason isn’t the models themselves. It’s how the implementation is organized.

Implementing AI in a company means embedding a use case into the existing system and process landscape in such a way that it runs without special oversight. Results must be traceable, and the benefits must be quantifiable. A pilot proves feasibility. Production operation proves cost-effectiveness. These are two different tasks, and the second is almost always underestimated.

How far have German companies actually come with AI?

Adoption has doubled within a year. According to the 2026 Bitkom study, 41 percent of companies with 20 or more employees are actively using AI; twelve months earlier, the figure was 17 percent. Another 48 percent are planning or discussing its use, while only 11 percent reject it. The study is based on a survey of 604 companies.

More interesting is the second figure from the same survey: One-third of companies using AI find the technology more expensive than expected. Nearly one in five has cut jobs as a result. This aligns with the MIT findings. AI is being widely adopted, but poorly managed.

For you as a controller or CFO, this means two things. First: The difference no longer lies in whether to use AI at all. It lies in integrating it into routine operations. Second: If you don’t know the costs of an AI use case, you can’t evaluate its benefits. This isn’t an IT issue – it’s a controlling issue.

BI2run - hybrid Meeting

Why Do AI Projects Fail Between the Pilot Phase and Full-Scale Operation?

In projects, we always see the same five points of failure. They usually occur together, and none of them can be resolved by a better model.

1. The pilot ran on a data export, not on the system

For the demo, someone pulled an Excel file from the ERP system, cleaned it up, and fed it into the model. In production, the AI needs access to the source data – complete with permissions, history, and daily updates. The data cleanup – which took one person two days before the pilot – becomes an ongoing requirement. If the integration isn’t factored into the pilot, the project will have to start over from scratch once it’s approved.

2. There is no process manager, only a project manager

A pilot project has a project manager. A productive use case requires someone from the business unit who takes responsibility for the results and decides what to do if the AI produces an implausible output. If this role is missing, every query ends up with IT, and IT lacks the subject-matter expertise to make such decisions.

3. The benefit was never measured against a baseline

“It’s faster now” isn’t proof. Without the baseline figure, the post-implementation figure cannot be evaluated. If you didn’t record how many hours the monthly closing commentary used to take, you won’t be able to substantiate the savings afterward. And without documentation, there’s no budget for the rollout.

4. Authorization and traceability are not regulated

As soon as AI not only reads but also writes values or text back into a planning system, you need the same controls as you would for human intervention. Who approved it? What data was used to generate the value? Can the system be restored to its pre-change state? These questions will come from internal audit at the very latest – and they always come at the worst possible time.

5. Operating costs have not been calculated

Token or licensing costs scale with usage. A pilot project with five users costs very little. The same scenario with 200 users and daily runs looks quite different. On top of that, there are operational costs, model maintenance, and training. This is precisely where the Bitkom finding comes from – that one-third of companies find AI more expensive than expected.

How many of these five pain points apply to your pilot project?

Let’s take a quick look at them together – before the request to roll out the project gets stuck in the steering committee.

How many of the five failure points affect your pilot?

In a short conversation, we’ll take a look together – before the rollout request gets stuck in the steering committee.

Schedule a meeting →

What distinguishes a pilot operation from a production operation?

DimensionPilotProduction
Data SourceExport, prepared onceSystem integration with permissions
OwnershipProject leadProcess owner in the business department
Measuring ValueParticipant impressionsMetric measured against a documented baseline
Error HandlingPoint of contact within the project teamDefined escalation path and fallback
Write Accessusually noneApproval logic, logging, reversibility
CostProject budget, one-offOngoing cost item in the operating budget
User BasePilot group, trained and motivatedEntire business department, mixed prior knowledge

The table explains why a successful pilot is not a predictor of operational performance. The pilot tests the left column. The right column is what pays off.

BI2run - persönliches Gespräch

What steps make the introduction of AI robust?

This order has proven effective in our projects. It takes a bit more time before the pilot and saves months afterward.

  1. Choose the use case based on the process, not the technology. Start with a task that is currently time-consuming and recurring. Monthly closing commentary, document review, forecast creation, answering recurring technical questions.
  2. Measure the baseline before you start. Record: How many hours, how many runs, how many errors per month? Without these numbers, there will be no business case later.
  3. Build data access in the pilot as it should look in operation. Even if it takes two days longer. A pilot on an Excel copy proves nothing about operational capability.
  4. Appoint a process responsible person from the specialist department. She decides on plausibility and approval, not IT.
  5. Consider release and logging logic from the beginning. For every write access: Who releases it, what is logged, how is it reset?
  6. Calculate full costs for 12 months, scaled up to the target user group. Licenses, usage costs, operation, training, model maintenance.
  7. Define the termination criterion in advance. If the case does not reach the target value after three months of operation, it will be shut down. This protects against use cases that continue running out of habit.
Free Initial Consultation · AI in Controlling

From Pilot to Production

We’ll work with you on data integration, ownership, and baselines – so your use case doesn’t get stuck at one of the five failure points.

Schedule a consultation →

Why we are approaching the topic from the data side

We have been building planning and reporting systems on IBM Planning Analytics for years. That’s why we know where the data is located in the finance environment and how it is defined. That is exactly the part where AI implementations get stuck. Our AI use cases build on this data layer instead of creating a second truth alongside it.

This occasionally leads us to advise against a use case. If the metric definitions diverge between areas or data is only generated manually, AI brings nothing but faster incorrect numbers. Then data work is the first step, not the model.

Facts and Sources at a Glance
  • 41 percent of German companies with 20 or more employees actively use AI, up from 17 percent twelve months earlier. Another 48 percent are planning or discussing it, 11 percent reject it. Base: 604 companies surveyed. Source: Bitkom study, 2026.
  • One-third of companies using AI rate the technology as more expensive than expected. Nearly one in five has cut jobs as a result. Source: Bitkom study, 2026.
  • 95 percent of generative AI pilots at companies produce no measurable bottom-line effect. Base: 52 executive interviews, a survey of 153 leaders, and an analysis of 300 publicly documented deployments. Source: MIT, Project NANDA, 2025.
  • The MIT authors attribute the failures not to model quality but to lack of integration and a learning gap within organizations.
  • According to the same study, the highest documented returns occur in back-office functions, such as document review and support.

Glossary

TermMeaning
Pilot / Proof of ConceptA time-limited test that demonstrates the technical feasibility of a use case. Not proof of operational readiness.
ProductionPermanent, day-to-day operation of a use case with defined ownership, system integration, support, and a cost center.
BaselineA documented starting value before implementation, such as processing hours or error rate. The basis for any measurement of value.
Write-backWriting AI-generated values or text back into an operational system, such as a planning cube.
AI GovernanceThe set of rules for using AI: responsibilities, approvals, logging, and handling of errors and personal data.
Agentic Use CaseAI that doesn’t just respond but independently carries out multi-step tasks, such as retrieving data, calculating, and writing results.

Frequently Asked Questions about AI Implementation

How long should an AI pilot last?

Four to eight weeks are sufficient for most use cases in controlling. If it takes longer, it’s usually due to a lack of data connection or a decision. More important than the duration is that the pilot works with the same data source that will be used later in operation.

Which use case is suitable for getting started?

Recurring tasks with clear results and an existing data basis. In the finance environment, commenting on deviations, document and invoice verification, forecast suggestions, and answering technical questions on internal documents work well. Unsuitable are cases where the data foundation needs to be created first.

Why isn’t a good language model enough?

Because the model only responds as well as the context it receives. Without clean metric definitions, without authorization logic, and without up-to-date data, even a strong model delivers numbers that no one in controlling trusts. The work lies in the data layer, not in the model.

Who should be responsible for the introduction of AI in the company?

Professionally, the area that benefits, technically the IT, coordinating a designated person with a mandate. A purely IT responsibility leads to no one making decisions on technical plausibility questions. A purely technical responsibility without IT integration fails due to data connectivity.

How do we measure the benefits of an AI use case?

Through the comparison with the documented baseline: saved processing time, reduced error rate, shortened lead time, or avoided external effort. Offset all ongoing costs, scaled up to the target user group. Only in this way does a number emerge that holds up in the steering committee.

What to do if the pilot was successful but the budget is lacking?

Usually, it’s not the budget that’s missing, but the reliable figure. Calculate the case over twelve months, with full costs and measurable benefits against the baseline. A use case that pays for itself within a year usually finds a budget.

Share article:

LinkedIn
WhatsApp
Facebook
Email

More articles

Any questions? Our experts look forward to your call!