Menu Schließen

Machine Learning and Big Data in Controlling: What these Methods can do today

BI2run - Machine Learning & Big Data in Controlling

71 percent of companies that use AI employ it for text generation. For data analysis, the figure is 31 percent. These figures from the Bitkom survey conducted in April 2026 highlight where there is still room for improvement in controlling. The computational aspect of AI is used less frequently than the text-generation aspect, even though that is precisely where the core of controlling work lies.

Machine learning in controlling is no longer just a research topic. These methods are now built into planning platforms. The question that remains is which method answers which controlling question and what data is needed for that.

A quick Explanation

In management accounting, machine learning refers to methods that identify patterns in historical data and use them to generate forecasts or detect anomalies without anyone having to define the rules in advance. Typical applications include sales and cost forecasts, anomaly detection in financial statements, and driver analyses.

Big Data, on the other hand, describes the volume and diversity of data. In the controlling functions of small and medium-sized businesses, data volumes are usually manageable. A clean historical data set spanning 24 to 36 months is sufficient for reliable forecasts – not a sea of data.

What distinguishes machine learning from traditional updating?

Traditional forecasting applies a predetermined rule: revenue plus five percent, personnel costs based on the collective bargaining agreement. Machine learning derives the rule from the data, weighing multiple factors simultaneously.

The difference is particularly evident in two areas. Machine learning recognizes seasonality and correlations that are overlooked in a manual rule. However, it does not provide a rationale in a business context. A model predicts that a value is likely, not why it occurs.

For forecasting, we use these methods in AI forecasting and rolling forecasting, among other applications.

MethodWhat it calculatesUse in controllingData requirements

Time series forecasting

Projection based on trend and seasonality

Revenue and cost forecast per account

24 to 36 months in a consistent structure

Regression

Relationship between drivers and the target variable

Driver-based planning, sensitivities

History plus driver data

Classification

Assignment to categories

Account coding suggestions, risk rating of receivables

Examples with known outcomes

Anomaly detection

Deviation from the expected pattern

Review during closing, uncovering posting errors

Complete transaction data

Clustering

Groups with similar behavior

Customer and product segments for planning

Attributes per object

What controlling questions can this tool help answer?

The way a question is assigned to a specific process determines its usefulness. This overview makes that clear.

The right-hand column is the most important. Every process has a clearly definable boundary. Those who understand this use the results as a basis for decision-making rather than as the decision itself.

Question from controllingSuitable methodWhat you getWhat it does not do

How will revenue per region develop next quarter?

Time series forecasting

Suggested value per account and month

No statement on new markets without history

Which drivers explain the cost increase?

Regression

Weighting of influencing factors

Correlation is not causation

Which postings in the closing are unusual?

Anomaly detection

Prioritized review list

No assessment of whether a case is actually wrong

Which customers behave similarly?

Clustering

Segments for differentiated planning

No ready-made sales strategy

How likely is a payment default?

Classification

Risk rating per receivable

No decision on dunning levels

How much data does machine learning really need in controlling?

significantly less than the term “big data” suggests. For a monthly time-series forecast, 24 to 36 consecutive periods are generally sufficient. What matters is not the volume, but the consistency.

Three conditions are more important than any volume of data.

  • Consistent structure throughout the entire history. An accounting change in the previous year will invalidate any forecast if it is not reflected retroactively.
  • Unambiguous definitions. If revenue is defined differently in two systems, the model will reliably calculate the wrong figure. The article “Why Metadata Is the Key to Successful AI Use” explains why metadata is so crucial here.
  • Documented one-time effects. One-time events must be flagged; otherwise, the model will carry them forward.

For this reason, a well-maintained OLAP model is a better foundation than a large, unstructured dataset. Our article “What actually is TM1” explains the reasoning behind this.

BI2run - hybrid Meeting

As a controller, how do you evaluate an ML result?

Five steps that work even without statistical knowledge.

  1. Compare the forecast for the last six completed months with the actual figures. This retrospective analysis is the most important test.
  2. Examine the largest deviations individually. Is there a one-time effect that wasn’t accounted for in the model?
  3. Compare the forecast with a simple extrapolation. If the model isn’t better, use the simple extrapolation.
  4. Check the outliers. Negative revenue or costs close to zero indicate structural errors in the data.
  5. Document the approval with the date and the person responsible. Without this step, no one will take ownership of the numbers.

This verification routine is precisely the competency that, according to the EY AI Readiness Check 2026, only 23 percent of companies systematically practice. Information on how to organize this process can be found in the AI training course for Controlling.

Measure forecast accuracy properly, once

We run a backtest on your data over the last few periods and show whether a model is measurably better than your current roll-forward.

Have your forecast accuracy checked →

When are traditional methods the better choice?

There are three situations in which we advise against using machine learning.

When the data history is short or incomplete. With fewer than 24 consistent periods, a model cannot identify reliable patterns. In such cases, well-reasoned manual planning is a more reliable approach.

When values are set for political reasons. A sales target is a decision, not a forecast. Mixing the two leads to fruitless discussions.

When there are structural changes. If the business model, product lineup, or pricing logic is undergoing fundamental changes, the historical data describes a different company.

This brings us to the second most important point of this article: The greatest leverage currently lies not in a better algorithm, but in the underlying data structure. While 71 percent of AI users have the system generate text and only 31 percent analyze it, teams that keep their historical data clean gain a competitive edge in financial control. This advantage remains regardless of which model is current in two years.

Why BI2run starts with the data structure when making forecasts

We build planning models in which forecasting methods run directly on the production data structure. This ensures transparency regarding the values used to generate a forecast. In controlling, this transparency is more important than the last percentage point of forecast accuracy.

If a review shows that your current forecasting method is performing just as well, we’ll say so. A model that doesn’t improve anything only adds to the maintenance burden.

Free backtest · Forecast accuracy

Measure first, then build the model

We use your closed periods to check how well your current forecast performs. The result shows whether a model can add any value at all.

Book a backtest →
Fact box
  • Areas of application among AI-using companies: text creation and translation 71 percent, marketing and communication 53 percent, customer service 42 percent, data analysis 31 percent. Source: Bitkom, April 2026, 604 companies with 20 or more employees.
  • 54.4 percent of German companies actively used AI in May 2026, and 58.7 percent in industry. Source: ifo Institute, Business Survey.
  • Only 23 percent of companies systematically validate AI results. Source: EY AI Readiness Check 2026.
  • For monthly time series forecasts in controlling, 24 to 36 consistent periods are considered the practical minimum.

Glossary

TermDefinition

Machine learning

A method in which a system derives rules from example data instead of having them specified in advance.

Big data

Data sets whose volume, variety and velocity make them impossible to analyze with conventional tools.

Predictive analytics

The use of statistical methods to calculate probabilities for future values based on historical data.

Time series forecasting

A forecasting method that separates the trend, seasonality and random fluctuation of a variable over time.

Anomaly detection

A method that flags values that deviate from the expected pattern.

Forecast accuracy

A metric that measures how far the forecast is from the actual value.

Feature

An individual influencing variable used as input for a model, such as order intake or public holidays.

Frequently Asked Questions about Machine Learning and Big Data in Controlling

What prior knowledge is required for machine learning in controlling?

You don’t need a degree in statistics. What you need is an understanding of how a method works and what its limitations are, so that you can successfully evaluate the results.

How does predictive analytics differ from machine learning?

Predictive analytics describes the goal – that is, the prediction. Machine learning describes the path to achieving it. Many forecasts are generated using classical statistics and without a self-learning model.

Does machine learning in controlling require a cloud platform?

Not necessarily. The key is that the data and the computational logic are integrated. Many planning platforms already include these capabilities.

How do you determine whether a model is better than the previous plan?

A review of past periods. The forecast and the actual value are compared, and then the same calculation is performed for the previous method. The comparison of the two variances is decisive.

What data can be included in a model?

That depends on the type and source of the data, especially when it comes to personal data. The framework needs to be clarified before the first run, not after.

Share article:

LinkedIn
WhatsApp
Facebook
Email

More articles

Any questions? Our experts look forward to your call!