You’ve implemented an AI forecast, and the numbers look plausible. Then someone on the board asks whether the new forecast is better than the old one. And the room falls silent. Without a measured forecast accuracy, this question cannot be answered. And without an answer, any discussion of forecasting methods remains a matter of personal preference.
Forecast accuracy refers to the measurable difference between a forecast and the subsequent actual value. It is expressed using error metrics such as MAPE, WAPE, bias, and forecast value added. It is always calculated for a fixed forecast horizon, such as “forecast from month M for month M+2.” Without this context, the figure is not comparable.
Why isn’t comparing the forecast to the actual results enough on its own?
Because a single comparison says nothing about the quality of the process. A forecast that was off by 2 percent in June can still be systematically too optimistic. And a forecast with a 12 percent deviation can be accurate if business fluctuates widely.
Added to this is the aggregation effect. Accuracy almost always looks better at the aggregated level than at the product or location level. Overestimates and underestimates cancel each other out in the aggregate. Anyone who measures only the consolidated figure will consider a forecast to be good even if it’s unusable at the divisional level.
That’s why accuracy measurement requires three specifications before the first figure is generated:
- Aggregation level: Group, division, customer, product? Measure at the level where decisions are made.
- Forecast horizon: How many periods prior to the actual data was the forecast created? A forecast for the next month is not comparable to one for the next quarter.
- Freeze date: Which plan status applies? If values are overwritten retroactively, accuracy cannot be reconstructed. The plan status must be finalized and archived as of the cutoff date.
The third point is the most common stumbling block. In many systems, the forecast is constantly being overwritten. As a result, the number that needs to be measured no longer exists.

Which metrics indicate the quality of a forecast?
Four key metrics addressed the relevant questions in our projects. They answer different questions and are not interchangeable.
| Metric | What It Answers | Strength | Limitation |
|---|---|---|---|
| MAPE (Mean Absolute Percentage Error) | How large is the average percentage deviation? | easy to communicate, unit-free | distorted with small actual values, penalizes overestimation less than underestimation, undefined when actual = 0 |
| WAPE (Weighted Absolute Percentage Error) | How large is the deviation, weighted by volume? | robust with small and fluctuating values, useful for portfolios | individual small line items barely register |
| Bias (Forecast Bias) | Are we systematically forecasting too high or too low? | reveals behavioral patterns, such as safety buffers in sales | says nothing about the spread; a bias of zero can mask large errors |
| FVA (Forecast Value Added) | Is our method better than a simple benchmark? | checks whether the effort in the process pays off | requires a clean reference forecast |
Bias is the most uncomfortable number
MAPE tells you how far off the mark you are. Bias tells you in which direction, and that’s usually the more politically interesting piece of information. Here’s a real-world example: If a 12-month sales forecast is, on average, 8 percent too high, that’s not a methodological problem – it’s an incentive problem. You’ll only find patterns like this if you don’t adjust for the signs.
FVA answers the question that matters most in AI forecasting
Forecast Value Added measures the extent to which forecast quality improves as a result of an additional process step. The result is compared to a deliberately simple reference forecast, known as the “naive forecast.” Common examples include the most recent known value or the value for the same period in the previous year.
The principle was developed by Michael Gilliland and is well-established in demand planning. For controlling and AI forecasting, it is the most meaningful test. If an AI forecast does not outperform the prior-year value, it represents no progress – even if the dashboard looks impressive. The same applies to manual adjustments; they also incur costs and must demonstrate their value.
This is precisely where the value of the method lies for you. It evaluates not only models but every step in the process. The basic statistical forecast, adjustments by the business unit, consolidation in Controlling: every step is assigned a number.
Does your forecast beat the naive benchmark?
If nobody’s measured that yet, let’s take a look together – before the next steering committee meeting turns it into a matter of opinion.
How do humans and models work together to produce better forecasts?
The widespread expectation is that a model will replace manual forecasting. The evidence tells a different story. An article on predictive analytics in corporate planning published by Springer Nature describes an internal benchmark analysis. Result: The combined forecast was more accurate than both the purely machine-based forecast and the purely manual bottom-up forecast. The authors conclude that predictive analytics should involve greater integration with management accounting. This specific case does not allow for a company-wide generalization.
In practice, this means: The model provides the baseline, while the business unit contributes knowledge not contained in the data – such as tenders, customer losses, and price rounds. And the FVA measurement determines whether this contribution actually improves the forecast. Without this measurement, every adjustment is automatically approved, even those that are harmful.

What steps lead to a reliable accuracy measurement?
- Freeze and archive plan statuses. For each forecasting cycle, save an unchangeable version with a specific cutoff date. In IBM Planning Analytics, a version dimension in the cube is sufficient for this. Without this step, nothing else can proceed.
- Define reporting levels. Select a management level and an operational level, such as total revenue and revenue by product group. The corporate level alone is not sufficient.
- Define time horizons. Common horizons are M+1, M+3, and year-end. Measure each horizon separately.
- Include a naive reference forecast. Use the most recent actual value and the prior-year value for the same period. This requires two calculation steps in the model and forms the basis of every FVA statement.
- Embed key metrics in reporting. Include WAPE and bias for each level and horizon, plus FVA for each process step. As a fixed component of the monthly report, not as a special analysis.
- Set target values only after three months of measurement. Before that, you don’t know your baseline. Setting target values without a baseline measurement only leads to evasive behavior.
- Feed the results back into process design. A step with a consistently negative FVA is simplified or eliminated. That is the actual purpose of the measurement.
What changes as a result of measuring forecast quality?
| Situation | Without Accuracy Measurement | With Accuracy Measurement |
|---|---|---|
| Evaluating a new AI forecast | Discussion based on plausibility and gut feeling | Comparison against a naive benchmark, decision backed by a number |
| Systematic planning errors in sales | only surfaces at year-end results | becomes visible via bias after two to three periods |
| Effort in the forecasting process | grows because no step is ever questioned | steps that add no value get cut |
| Discussion in the steering committee | methodology dispute | metric-based comparison |
| Rolling forecast planning | more cycles, same uncertainty | visible at which horizon accuracy starts to drop |
How accurate is your forecast, really?
We’ll work with you to figure out how MAPE, WAPE, bias, and FVA can be mapped in your IBM Planning Analytics model – as a fixed part of planning, not a side calculation.
Schedule a consultation →- MAPE is the most common metric for measuring forecast accuracy. It expresses the percentage error relative to the actual value and isn’t meaningfully applicable when actual values are close to zero. Source: Imperia SCM, “MAPE in Forecasting: Formula, Good Values and Limitations.”
- Error metrics can be misleading when viewed in isolation. The reason is the aggregation effect: accuracy looks better at an aggregated level than at product or location level. Source: ToolsGroup, Forecast Accuracy Benchmarks.
- Forecast Value Added measures the percentage improvement in forecast performance contributed by each additional process step, measured against a naive reference forecast. Traced back to Michael Gilliland (2013, 2015). Sources: SAS white paper “Forecast Value Added Analysis: Step by Step”; Institute of Business Forecasting (IBF).
- The naive reference forecast is typically either the last observed value (random walk) or the actual value from the same period in the prior year (seasonal random walk). Source: SAS, white paper “Forecast Value Added Analysis: Step by Step.”
- Case finding from an internal company benchmark analysis: a combined forecast from model and controlling was more accurate than purely machine-generated or purely manual bottom-up forecasts. Not an industry-wide benchmark. Source: Springer Nature, “Predictive Analytics als digitaler Service.”
Glossary
| Term | Meaning |
|---|---|
| Forecast Accuracy | Measurable distance between forecast and actual value, expressed through error metrics for a defined level and horizon. |
| MAPE | Mean Absolute Percentage Error. Average of the absolute percentage deviations. Signs are removed. |
| WAPE | Weighted Absolute Percentage Error. Sum of absolute errors divided by the sum of actual values. More robust with small values. |
| Forecast Bias | Average deviation with sign. Reveals systematic over- or underestimation. |
| Forecast Value Added (FVA) | Improvement in forecast quality contributed by a process step, measured against a naive reference forecast. |
| Naive Forecast | Deliberately simple reference forecast, such as the last actual value or the prior-year value for the same period. |
| Forecast Horizon | Time gap between when the forecast is created and the forecasted period, e.g. M+1 or M+3. |
| Freeze | Immutably storing a plan state as of a specific cutoff date. A prerequisite for any later accuracy measurement. |
Frequently Asked Questions About Forecast Accuracy
What MAPE value is considered good?
There is no universally applicable target value. A MAPE of 5 percent may be considered poor for stable subscription revenue but very good for project-driven business. The only reliable comparison is with your own historical data and with the naive reference forecast. That is why the baseline measurement always comes before the target value.
Why should we evaluate MAPE and bias together?
Because they both highlight different types of errors. MAPE ignores the direction, while bias ignores the dispersion. A forecast with zero bias can contain large errors in opposite directions that cancel each other out mathematically. Only when considered together do these two metrics provide a useful picture.
How do we determine whether an AI forecast is better than our previous forecast?
About FVA. Calculate three values for the same time period, the same level, and the same horizon: a naive reference forecast, the previous forecast, and the AI forecast. Compare the error metrics. If the AI forecast does not outperform the naive reference forecast, it provides no added value. If it does not outperform the previous forecast, there is no justification for the switch.
What data do we need for a retrospective measurement?
Archived plan statuses with a reference date and the corresponding actual values in the same structure. If the plan history is missing because values were overwritten, accuracy cannot be reconstructed. In that case, the measurement begins with the next forecasting round, and the first step is to freeze the versions.
How often should the accuracy evaluation run?
Monthly, as a regular part of the reporting process. A one-time special analysis once a year sparks discussion but leads to no improvement. The value comes from repetition, because only then do patterns – such as systematic bias – become apparent.
What role does accuracy measurement play in rolling planning?
Especially there. Rolling planning increases the number of forecasting cycles. Without measuring accuracy, it simply increases the workload. By measuring accuracy, you can see at what point in the forecast horizon accuracy begins to decline, and you can make an informed decision about the number of cycles.
Why Accuracy Measurement Belongs in the Planning System and Not in a Separate Calculation
In practice, accuracy measurement almost always fails for one reason: the lack of a planning history. If you store forecast versions in separate Excel files, you won’t have a comparable data set after two years. In a planning system like IBM Planning Analytics, this is a modeling issue. A version dimension, a process for freezing data, and the key metrics are then calculated in the same cube used for planning.
That’s why we set up accuracy metrics as part of the planning model rather than as a separate report. The measurement uses the same definitions as the planning. This eliminates the debate over whether two numbers mean the same thing.
And if the analysis shows that your previous forecast significantly outperforms the naive benchmark, we’ll say so. Not every company needs an AI model for forecasting. Some just need proof that their process works.

























