Planning Intelligence Research
An ongoing research series exploring the future of planning, forecasting, drivers, and decision intelligence.
Planning intelligence should help a team learn from uncertainty, not disguise it behind a confident answer. This research note proposes an evaluation agenda for organizations exploring forecasting and decision support. It is an editorial research framework, not a report of Aarietech experiments, customer results, or validated product performance.
Why the evaluation question matters now
Oracle’s current EPM use-case guidance distinguishes bounded applications from more complex predictive models and cross-process orchestration. That distinction is useful context for research design: a successful narrow demonstration does not establish that a larger planning system will behave reliably. The guidance is vendor-authored and should be assessed alongside the organization’s own requirements.
Source: Oracle — Selecting the First EPM AI Use Case (Documentation accessed September 12, 2026).
Separate the questions being tested
A forecasting evaluation asks how predictions compare with later observations. A narrative evaluation asks whether an explanation is supported by the available evidence. A decision-support evaluation asks whether the output helps a person choose an appropriate action. These questions require different measures and should not be combined into a single “AI accuracy” score.
An illustrative experiment could compare a current planning approach, a simple baseline, and a proposed forecasting method on the same historical periods. Keep the information available at each forecast date consistent. Using later information during evaluation would make the comparison misleading.
Look beyond the average result
Evaluate performance at the business level where a decision will be made. A reasonable aggregate result can hide weak performance for a product group, entity, or planning horizon. Review unusually volatile periods and missing-data cases separately, and document where a method should not be used.
For generated commentary, reviewers can check whether numerical statements reconcile, whether claimed causes have supporting evidence, and whether uncertainty is stated appropriately. Record disagreement between reviewers rather than assuming there is always one unambiguous narrative. Useful research may reveal that better source information matters more than a different model.
Make the experiment repeatable
Record the data version, preparation steps, assumptions, model configuration, and acceptance criteria before comparing results. Keep a set of examples that were not used to tune the approach. When the method changes, rerun the comparison and examine regressions as well as improvements.
Finally, assess the human workflow. Does the proposed output reduce investigation effort, or merely shift that effort into checking AI text? Can a finance professional trace a conclusion and correct it? A planning assistant earns a larger role when it improves that decision process under realistic conditions, with its limitations still visible.