Forecasting · 10 min read · September 2026
Demand forecasting and optimization in 2026: beating the baseline, and the override that undoes it
By Sean Majidi, Founder, Thinklytics
Global inventory distortion is $1.7 trillion a year, two thirds of it stockouts. But a September 2026 study on half a million retail records found no model best on every metric, and a field study of 575,000 planner decisions found people override AI forecasts harder than identical statistical ones.
Global inventory distortion runs at $1.7 trillion a year, two thirds of it stockouts. And in a September 2026 study on roughly half a million retail records, no model was best on every metric: the deep learning model that won on percentage error lost to simple tree-based models on scaled error.
Both of those things are true, and holding them together is the whole job. The prize is real and large. The claim that a better model reliably captures it is not supported by what was published this year.
The size of the prize, stated accurately
Where the $1.7 trillion sits
- Stockouts, 66.9%. $1.14T. Revenue that never happened. Invisible in the ledger, so chronically underfunded.
- Overstock, 33.1%. $560B. Shows up as a write-off someone has to explain, which is why it gets the attention.
Equal to 6.2% of global retail sales in 2026, down from 10.4% in 2021 and from 6.5% in IHL's own September 2025 release. Distortion is improving, slowly. Source: IHL Group, 2026 Inventory Distortion Study, drawing on coverage of 4,500+ enterprise retailers and hospitality providers. No fieldwork window published; cite as published 2026.
Source: IHL Group, 2026 Inventory Distortion Study, drawing on coverage of 4,500+ enterprise retailers and hospitality providers. No fieldwork window published; cite as published 2026.
IHL Group's 2026 Inventory Distortion Study puts the global figure at $1.7 trillion annually, equal to 6.2% of retail sales, against 10.4% in 2021. The split is 66.9% stockouts, around $1.14 trillion, and 33.1% overstock, around $560 billion.
Two things about that number. IHL draws on coverage of more than 4,500 enterprise retailers and hospitality providers and does not publish a fieldwork window, so it belongs in a document as published in 2026 rather than fielded in 2026. And the trend in IHL's own series is downward: its September 2025 release put it at $1.73 trillion and 6.5%. Distortion is improving, slowly, which is a less exciting framing and a more credible one.
The stockout-heavy split matters for how the work gets scoped. Two thirds of the loss is revenue that never happened, which is invisible in the ledger and therefore chronically underfunded relative to overstock, which shows up as a write-off someone has to explain.
On the operator side there is one clean 2026-fielded measurement. inFlow surveyed 400 US inventory operators in March 2026 and found 44% experiencing stockouts at least monthly, 54% reporting holding costs above 10% of inventory value, and 81.2% wanting to implement AI against 11% currently using any AI tools at all. The sample skews to smaller businesses and construction, and the carrying-cost figures are self-reported rather than audited. It is still the only current, disclosed-method operator survey in the space.
What the models actually beat
The same dataset, graded four ways
Roughly 500,000 product, store and day records; 459,310 retained after processing. No model won on every metric.
| Finding | Result | What it means for a buying decision |
|---|---|---|
| No-skill baseline | Seasonal naive, MASE 1.00 | The threshold every vendor claim should be quoted against, rather than against no forecast at all |
| Best percentage error | LSTM, RMSE 0.3996, MAPE 21.63% | The deep learning model won here, which is the result that gets put in a deck |
| Best scaled error | Random forest and XGBoost | Two much simpler methods beat the LSTM on MASE. Which model is best depends on the metric you are graded on |
| Authors' own conclusion | No model consistently outperformed all others | Ask which metric the pilot will be scored on before the pilot starts, not after |
| Inventory risk classification | 61.1% Potential Risk or Critical, 16.3% Critical | The forecast is an input to a risk decision, not the deliverable |
| Published industry MAPE benchmark for 2026 | Does not exist | IBF stops at 2023 to 2024; APQC is paywalled with no fieldwork date; Gartner publishes none publicly |
Source: Gao Huan and Mohammad-Ali Sarvghadi, Information 17(9):909, published 17 September 2026. Figures quoted are those stated in the paper's retrievable text.
The benchmark that matters is a seasonal naive forecast, which is last season's number repeated. In scaled-error terms that defines the no-skill threshold at a MASE of 1.00, and every claim about forecasting improvement should be quoted against it rather than against no forecast at all.
A study published in Information in September 2026 ran that comparison on around 500,000 product, store and day records from a Chinese e-commerce company, of which 459,310 survived processing. The conclusion in the authors' own words is that forecasting performance varied by evaluation criterion, with no model consistently outperforming all others. The LSTM achieved the lowest testing RMSE at 0.3996 and MAPE at 21.63%. Random forest and XGBoost produced more favourable MASE values.
Read that again, because it is the finding most vendor material is designed to obscure. The model that won on percentage error lost on scaled error to two much simpler methods. Which model is best depends on which metric you are graded on, and nobody in the procurement process usually asks which one that will be.
The same study classified 61.1% of its SKU and store observations as Potential Risk or Critical, including 16.3% Critical, which is a reminder that the forecast is an input to an inventory risk decision rather than the deliverable.
What does not exist, and we looked hard: a 2026-fielded MAPE or WMAPE benchmark by industry or horizon. The Institute of Business Forecasting's published error benchmarking stops at 2023 to 2024. APQC's forecast accuracy figures are paywalled with no published fieldwork window. Gartner publishes no public accuracy benchmark. Any accuracy target has to be built from your own history, and any consultant quoting an industry MAPE at you should be asked for the source.
The finding that changes how you build it
The best new research of 2026 is not about models. It is about what people do with them.
A field study published in the Journal of Operations Management in April 2026 used around 575,000 observations from a quasi-natural experiment at a multinational retailer, plus two laboratory experiments, and found that users implement significantly greater adjustments for AI-based algorithms than for model-based ones. Their negative reaction to poor performance is amplified when the forecast is labelled as coming from AI. Adjustments reduced as the algorithms improved over time.
Be precise about what that does and does not show. It measures the magnitude of planner adjustment, not whether adjusting makes accuracy worse. The evidence that overrides destroy value is older forecast-value-added literature, not this paper.
What it does show is that the same forecast, labelled differently, gets treated differently. Which means the adoption problem in a forecasting programme is not a training problem to be solved after go-live. It is a design constraint from the first week: if the planners will override a number because of where it came from, the override path has to be instrumented, reviewed and fed back, or the model's measured accuracy and the company's actual forecast will diverge permanently.
Spend, and the absence of returns
Money committed, returns unknown
Four separate Gartner samples. Fieldwork windows differ and are given in the source line. Source: Gartner: n=394 fielded Nov 2025 to Feb 2026 (AI share of investment); n=135 fielded Jan to Apr 2026 (returns clarity); n=243 fielded 11 Nov to 18 Dec 2025 (spend and 2030 autonomy prediction); n=151 fielded 11 Nov to 21 Dec 2025 (approval revisits); n=140 fielded Oct to Nov 2025 (legacy integration).
- Supply chain digital investment allocated to AI
- Senior supply chain leaders unclear on AI returns
- Already spent $3M or more on planning automation
- Revisit final network-decision approvals at least once
- Say AI integration with legacy systems is a major challenge
- Will make 10%+ of planning decisions autonomously by 2030
Source: Gartner: n=394 fielded Nov 2025 to Feb 2026 (AI share of investment); n=135 fielded Jan to Apr 2026 (returns clarity); n=243 fielded 11 Nov to 18 Dec 2025 (spend and 2030 autonomy prediction); n=151 fielded 11 Nov to 21 Dec 2025 (approval revisits); n=140 fielded Oct to Nov 2025 (legacy integration).
Gartner published two surveys in August 2026 that are frequently merged into one sentence and should not be. Across 394 supply chain professionals fielded November 2025 to February 2026, 67% of digital investment is going to AI. Across a separate sample of 135 senior supply chain leaders fielded January to April 2026, 55% are unclear on the returns from those investments.
Then the autonomy forecast. In research fielded 11 November to 18 December 2025 across 243 organisations, Gartner predicts only 5% of organisations will make at least 10% of supply chain planning decisions autonomously by 2030, while 83% had already spent $3M or more on planning automation and 51% between $3M and $10M.
Money already committed, autonomy not expected within four years, and most leaders unable to state the return. That is not an argument against the work. It is an argument for scoping it around decisions a person still makes, because that is what the next four years are going to look like regardless of what the platform roadmap says.
The Hackett Group's 2026 Supply Chain Key Issues Study is worth one careful note. Its headline is that 83% have deployed or are piloting AI in supply chain intelligence and analytics, with 59% pursuing inventory optimization. Secondary coverage routinely drops the words "or piloting" and reports it as an adoption rate. It is not one. Hackett also does not publish a sample size or fieldwork window for that study.
Which numbers to stop using
Forecasting statistics in circulation, and whether they survive tracing
The first three are current. The last three are being presented as 2026 data and are not.
- Global inventory distortion $1.7 trillion, 6.2% of retail sales. IHL Group, published 2026. No fieldwork window disclosed, so cite the publication year.
- 55% of senior supply chain leaders unclear on AI returns. Gartner, n=135, fielded January to April 2026. The cleanest fully-2026 fieldwork in this space.
- 44% of operators hit stockouts at least monthly, 11% use any AI tool. inFlow, n=400, fielded March 2026, United States. Skews smaller businesses and construction.
- AI cuts forecasting errors 20 to 50%. McKinsey, published April 2017. Nine years old and still recycled in 2026 content with no date attached.
- 8.3% global out-of-stock rate. Gruen, Corsten and Bharadwaj for the Grocery Manufacturers Association, 2002. So are the associated shopper-behaviour figures.
- 83% of supply chains have adopted AI. Hackett Group 2026 says 83% have deployed or are piloting. Secondary coverage drops the piloting. No sample size or fieldwork window published.
Source: IHL Group 2026; Gartner, 5 August 2026; inFlow State of Inventory Management 2026; McKinsey, April 2017; GMA, 2002; Hackett Group 2026 Supply Chain Key Issues Study.
Source: IHL Group 2026; Gartner, 5 August 2026; inFlow State of Inventory Management 2026; McKinsey, April 2017; GMA, 2002; Hackett Group 2026 Supply Chain Key Issues Study.
Two figures dominate forecasting business cases and neither is current.
The claim that AI cuts forecasting errors by 20 to 50% and reduces lost sales from unavailability by up to 65% comes from a McKinsey report published in April 2017. It is nine years old and it is still being quoted in 2026 material without a date.
The 8.3% global out-of-stock rate traces to Gruen, Corsten and Bharadwaj research published in 2002 for the Grocery Manufacturers Association, based on 71,000 consumers across 29 countries. The associated figures about what shoppers do when an item is out of stock come from the same 2002 work. One 2026 report carrying those statistics states on its own page that no proprietary data of its own is presented in them.
Neither figure is fraudulent. Both are being presented as current when they are not, and a buyer who checks will conclude the rest of the document deserves the same treatment.
What we would do first
Three checks before a forecasting model is funded
All three can be done with data you already have, in under two weeks.
- Measure what your current process beats a seasonal naive forecast by. Take last year's actuals, repeat last season's number, compare. A small gap means more room than you think. A large gap means less.
- Instrument the planner overrides. How often the system number is changed, by how much, in which direction. Most organisations do not hold this data, and without it the forecast that ships is not the forecast that was built.
- Trace the decision the forecast feeds. 72% revisit final network-decision approvals at least once and more than half revisit three or more times. A better number entering a loop does not change the outcome.
- Expect autonomy to remove the human from the loop. Gartner predicts only 5% of organisations will make 10% or more of planning decisions autonomously by 2030. Design around the person who still decides.
- Set an accuracy target from an industry benchmark. No 2026-fielded MAPE or WMAPE benchmark by industry exists publicly. The target has to be built from your own history.
Source: Gartner, 14 July 2026 (n=151, fielded 11 Nov to 21 Dec 2025) and 24 September 2026 (n=243, fielded 11 Nov to 18 Dec 2025); Journal of Operations Management, April 2026.
Source: Gartner, 14 July 2026 (n=151, fielded 11 Nov to 21 Dec 2025) and 24 September 2026 (n=243, fielded 11 Nov to 18 Dec 2025); Journal of Operations Management, April 2026.
Three checks, in this order, before any model work is funded.
Establish your current baseline. Take last year's actuals, generate a seasonal naive forecast, and measure what your existing process beats it by. If the answer is "not much", the improvement available from a better model is larger than you think. If the answer is "a lot", the improvement available is smaller than the vendor is quoting.
Instrument the overrides. Count how often a planner changes the system number, by how much, and in which direction. That data does not exist in most organisations, and without it there is no way to tell whether the forecast that ships is the forecast that was built.
Trace the decision. Gartner research fielded 11 November to 21 December 2025 across 151 supply chain leaders found 72% revisit final approvals on network decisions at least once and more than half revisit three or more times. A more accurate number entering a decision process that loops three times does not change the outcome, and that is a process problem no model solves.
Our data foundation practice builds the demand history and master data layer these models depend on, and the AI readiness assessment is where we run the baseline and override checks above before anything is scoped. For where forecasting sits against the other supply chain use cases, see the 2026 logistics and supply chain AI readiness map. If the current dashboards are the problem rather than the forecast, start with why supply chain teams rebuild dashboards first.
One last note on 2026 conditions. KPMG surveyed 300 US C-suite leaders at organisations above $1B in February and March 2026 and found 26% in formal planning or active execution on reshoring, up from 10% six months earlier, with 55% planning price increases of up to 15% within six months. Tariffs are changing planning behaviour measurably. No published source quantifies a change in forecast error attributable to them, so if someone tells you tariffs degraded accuracy by a specific percentage, ask where the figure came from.
Frequently asked questions
How much is bad forecasting actually costing?
IHL Group's 2026 Inventory Distortion Study puts global inventory distortion at $1.7 trillion a year, equal to 6.2% of retail sales, down from 10.4% in 2021. The split is 66.9% stockouts, around $1.14 trillion, and 33.1% overstock, around $560 billion. IHL draws on coverage of more than 4,500 enterprise retailers and hospitality providers and does not publish a fieldwork window, so cite it as published in 2026. The direction in IHL's own series is worth noting: it is improving, from $1.73 trillion and 6.5% in its September 2025 release.
How much better is machine learning than a simple baseline?
Less than the marketing suggests, and the honest answer is that it varies by which error metric you pick. A September 2026 study in Information analysed roughly 500,000 product, store and day records from a Chinese e-commerce company and concluded that forecasting performance varied by evaluation criterion, with no model consistently outperforming all others. The deep learning model with the lowest RMSE and MAPE was beaten on scaled error by simple tree-based models. A seasonal naive forecast defines the no-skill threshold at a MASE of 1.00, and any vendor claim should be quoted against that baseline rather than against no forecast at all.
Do planners override AI forecasts more than statistical ones?
Yes, and this is the most useful new finding of 2026. A field study published in the Journal of Operations Management in April 2026, using around 575,000 observations from a quasi-natural experiment at a multinational retailer plus two laboratory experiments, found users implement significantly greater adjustments for AI-based algorithms than for model-based ones, and that their reaction to poor performance is amplified when the forecast came from AI. Adjustments fell as the algorithms improved over time. Note carefully what the study does not claim: it measures how much planners adjust, not whether adjusting makes accuracy worse.
Are supply chain leaders seeing returns on forecasting AI?
Mostly they cannot tell. Gartner surveyed 135 senior supply chain leaders between January and April 2026 and found 55% unclear on the returns from their AI investments. A separate Gartner survey of 394 supply chain professionals, fielded November 2025 to February 2026, found 67% of digital investment being allocated to AI. Those are two different samples and should not be presented as one survey. The combination is the problem: two thirds of the digital budget going to AI, and most leaders unable to say what it returned.
How much of planning will actually run autonomously?
Very little, on Gartner's own forecast. In research fielded 11 November to 18 December 2025 across 243 organisations, Gartner predicts only 5% of organisations will make at least 10% of supply chain planning decisions autonomously by 2030. In the same research 83% had already spent $3M or more on planning automation, with 51% spending between $3M and $10M. The gap between spend already committed and autonomy expected by 2030 is the single best argument for designing around human decisions rather than around removing them.
What actually blocks a forecasting project?
Data and integration, not modelling. The Hackett Group's 2026 Supply Chain Key Issues Study found 50% citing data quality, 47% integration and 45% AI talent. Gartner, surveying 140 supply chain organisations in October and November 2025, found 56% saying integration of AI with legacy systems is a major challenge and 50% reporting limited internal AI talent. On the operator side, inFlow's survey of 400 US inventory operators fielded in March 2026 found 44.8% naming inaccurate inventory data as a top challenge, second only to supplier reliability at 52%.
Which forecasting statistics should we stop quoting?
Two in particular. The claim that AI cuts forecasting errors by 20 to 50% comes from a McKinsey report published in April 2017, and it is still being recycled in 2026 content. The 8.3% global out-of-stock rate traces to Gruen, Corsten and Bharadwaj research published in 2002 for the Grocery Manufacturers Association. Both appear in 2026-dated material with no date attached. There is also no 2026-fielded MAPE or WMAPE benchmark by industry in the public record, which means any accuracy target you set has to be built from your own history.
Where should a forecasting programme start?
With the decision the forecast feeds, not with the forecast. Gartner research fielded 11 November to 21 December 2025 across 151 supply chain leaders found 72% revisit final approvals on network decisions at least once, and more than half revisit three or more times. A more accurate number arriving into a decision process that loops three times does not change the outcome. Fix the decision path first, establish what your current forecast beats a seasonal naive baseline by, and only then decide whether a better model is the constraint.
The work behind this
Ten engagements in the case library carry this capability, covering demand, inventory, staffing and revenue forecasting.
Every one names the client where we are permitted to and states the measured outcome: Forecasting and optimization, 10 engagements.
Topics covered
- demand forecasting
- inventory optimization
- forecast accuracy
- supply chain ai
- planner override
- seasonal naive baseline
- stockout cost
Frequently asked questions
How much is bad forecasting actually costing?
IHL Group's 2026 Inventory Distortion Study puts global inventory distortion at $1.7 trillion a year, equal to 6.2% of retail sales, down from 10.4% in 2021. The split is 66.9% stockouts, around $1.14 trillion, and 33.1% overstock, around $560 billion. IHL draws on coverage of more than 4,500 enterprise retailers and hospitality providers and does not publish a fieldwork window, so cite it as published in 2026. The direction in IHL's own series is worth noting: it is improving, from $1.73 trillion and 6.5% in its September 2025 release.
How much better is machine learning than a simple baseline?
Less than the marketing suggests, and the honest answer is that it varies by which error metric you pick. A September 2026 study in Information analysed roughly 500,000 product, store and day records from a Chinese e-commerce company and concluded that forecasting performance varied by evaluation criterion, with no model consistently outperforming all others. The deep learning model with the lowest RMSE and MAPE was beaten on scaled error by simple tree-based models. A seasonal naive forecast defines the no-skill threshold at a MASE of 1.00, and any vendor claim should be quoted against that baseline rather than against no forecast at all.
Do planners override AI forecasts more than statistical ones?
Yes, and this is the most useful new finding of 2026. A field study published in the Journal of Operations Management in April 2026, using around 575,000 observations from a quasi-natural experiment at a multinational retailer plus two laboratory experiments, found users implement significantly greater adjustments for AI-based algorithms than for model-based ones, and that their reaction to poor performance is amplified when the forecast came from AI. Adjustments fell as the algorithms improved over time. Note carefully what the study does not claim: it measures how much planners adjust, not whether adjusting makes accuracy worse.
Are supply chain leaders seeing returns on forecasting AI?
Mostly they cannot tell. Gartner surveyed 135 senior supply chain leaders between January and April 2026 and found 55% unclear on the returns from their AI investments. A separate Gartner survey of 394 supply chain professionals, fielded November 2025 to February 2026, found 67% of digital investment being allocated to AI. Those are two different samples and should not be presented as one survey. The combination is the problem: two thirds of the digital budget going to AI, and most leaders unable to say what it returned.
How much of planning will actually run autonomously?
Very little, on Gartner's own forecast. In research fielded 11 November to 18 December 2025 across 243 organisations, Gartner predicts only 5% of organisations will make at least 10% of supply chain planning decisions autonomously by 2030. In the same research 83% had already spent $3M or more on planning automation, with 51% spending between $3M and $10M. The gap between spend already committed and autonomy expected by 2030 is the single best argument for designing around human decisions rather than around removing them.
What actually blocks a forecasting project?
Data and integration, not modelling. The Hackett Group's 2026 Supply Chain Key Issues Study found 50% citing data quality, 47% integration and 45% AI talent. Gartner, surveying 140 supply chain organisations in October and November 2025, found 56% saying integration of AI with legacy systems is a major challenge and 50% reporting limited internal AI talent. On the operator side, inFlow's survey of 400 US inventory operators fielded in March 2026 found 44.8% naming inaccurate inventory data as a top challenge, second only to supplier reliability at 52%.
Which forecasting statistics should we stop quoting?
Two in particular. The claim that AI cuts forecasting errors by 20 to 50% comes from a McKinsey report published in April 2017, and it is still being recycled in 2026 content. The 8.3% global out-of-stock rate traces to Gruen, Corsten and Bharadwaj research published in 2002 for the Grocery Manufacturers Association. Both appear in 2026-dated material with no date attached. There is also no 2026-fielded MAPE or WMAPE benchmark by industry in the public record, which means any accuracy target you set has to be built from your own history.
Where should a forecasting programme start?
With the decision the forecast feeds, not with the forecast. Gartner research fielded 11 November to 21 December 2025 across 151 supply chain leaders found 72% revisit final approvals on network decisions at least once, and more than half revisit three or more times. A more accurate number arriving into a decision process that loops three times does not change the outcome. Fix the decision path first, establish what your current forecast beats a seasonal naive baseline by, and only then decide whether a better model is the constraint.