Most manufacturers we meet have a forecast that is accurate at the category level and useless at the level where decisions get made. You don't schedule production for "industrial fasteners, EMEA." You schedule it for SKU 48213-B, which sold 340 units one month, 12 the next, and zero for the six weeks after that.
The gap between aggregate accuracy and SKU-level accuracy is where safety stock bloats, expedites multiply, and planners quietly override the system with their own spreadsheets. Closing it is a solvable engineering problem — but not with the tools most companies are using.
Why SKU-level is a different problem
Three things break when you go from category to SKU.
Intermittency. In a typical industrial catalog, 40–70% of SKUs sell in fewer than half of all periods. Classical exponential smoothing degrades into forecasting the average of mostly zeros, which is precisely wrong for both the zero periods and the demand spikes.
Scale. Ten thousand SKUs across a dozen plants and warehouses is 100,000+ series. Nobody hand-tunes that. The system has to select methods, detect regime changes, and flag anomalies automatically.
Signal starvation. A single SKU's history is short and noisy. The information you need often lives elsewhere — in sibling SKUs, in the customer's order patterns, in the distributor's promotion calendar. Per-series models can't see any of it.
What actually works
Our standard architecture is a tiered one, and the tiers are chosen by data, not by preference.
Segment first. We split the catalog by demand pattern — smooth, erratic, intermittent, lumpy — using the standard ADI/CV² cut, plus a volume-value ABC layer. This determines both the modeling approach and, more importantly, the error metric that matters for each segment.
Global models for the head and middle. A single gradient-boosted model (LightGBM remains hard to beat) trained across all series, with SKU, plant, customer-mix, price, and calendar features, consistently outperforms per-series statistical models by 15–30% weighted MAPE in our engagements. The model learns cross-SKU structure: when a product family ramps, its variants follow. Where history is deep and the tail is long, we'll benchmark neural approaches, but we make the boosted global model the baseline to beat — often it isn't beaten by enough to justify the operational complexity.
Honest probabilistic treatment of the tail. For truly intermittent SKUs, we forecast a demand distribution, not a point. Croston variants are the floor; quantile forecasts from the global model are usually the answer. The output the planner needs is not "3.2 units next month" — it's "95% chance demand is under 9, set the reorder point accordingly."
Forecast value added (FVA) as the referee. Every model, and every human override, is measured against a naive baseline. In one build for a building-products manufacturer, we found planner overrides were destroying 6 points of accuracy on A-items — and adding 11 points on new products where the model had no history. Both findings changed the process: overrides became exception-based, gated to the segments where humans demonstrably add value.
The unglamorous 60%
The modeling is maybe 40% of the work. The rest is data engineering, and it decides whether the forecast survives contact with production.
Demand history has to be demand, not shipments — stockouts censor the signal, and a model trained on shipments learns to under-forecast exactly the SKUs you failed on. Order dates, not invoice dates. Promotions, price changes, and customer onboarding/offboarding as explicit features. SKU supersessions chained so that a renumbered part doesn't look like a death and a birth.
Then the pipeline itself: automated retraining on a fixed cadence, backtesting on rolling origins (never a single holdout), data-quality gates that block a forecast publish when input tables arrive late or anomalous, and lineage so a planner asking "why did the forecast for 48213-B jump 40%?" gets an answer with feature attributions, not a shrug.
Measuring what matters
We hold a hard line on two evaluation practices.
First, weighted metrics. Unweighted MAPE across 10,000 SKUs lets the model hide bad performance on the 500 SKUs that carry 80% of revenue. We report accuracy weighted by margin contribution and, separately, service-level impact for the tail.
Second, decision-level backtests. The forecast is an input to inventory and scheduling decisions, so the eval that matters is counterfactual: had you planned from these forecasts for the last 12 months, what would inventory, service level, and expedite spend have been? In a recent engagement that simulation showed a 14% inventory reduction at equal service level — a number a CFO can act on, unlike a MAPE delta.
The honest failure modes
Where does this go wrong? Forecasting demand that sales already knows about (integrate CRM opportunity data before blaming the model). Chasing accuracy on C-items that should simply carry a cheap two-bin policy. And shipping a model with no ownership plan — forecasts drift, catalogs churn, and an unmonitored pipeline is 12 months from being quietly replaced by the spreadsheet it was meant to retire.
A useful pilot scope: one product family, one plant, full pipeline from raw ERP extracts to a published forecast with FVA tracking, benchmarked against the incumbent process over an 8-week parallel run. That's a 4–6 week build plus the parallel run — enough to know, with numbers, whether to scale it. Production-or-nothing applies doubly to forecasting: a backtest in a notebook is an opinion. A forecast the planner scheduled from is a result.