Demand forecasting in apparel works predictably for replenishment basics and fails predictably for new fashion items: published retail forecasting literature has long reported mean absolute percentage errors in the thirty to fifty percent range for fashion items, versus far tighter bands for staple replenishment, and no method introduced through the mid-2020s closed that gap. The reason is structural. A forecast, which is a quantitative estimate of future sales by item, size, and location, needs history or a proxy for history, and a never-before-seen seasonal item has neither. The craft lies in knowing which model to trust for which product class.
Which forecasting models do apparel companies actually use?
Three families dominate practice, and each fits a different slice of the assortment.
- Time-series extrapolation: methods in the exponential smoothing family project each item's sales history forward, handling trend and seasonality from past patterns. They are cheap, transparent, and accurate for items with two or more seasons of stable history, meaning most basics and carry-over styles.
- Causal regression: models relate demand to explainable drivers such as price, weather, distribution, and promotions. They answer planning questions that extrapolation cannot, such as what a price change does to volume, at the cost of needing driver data that many brands keep badly.
- Analogous-item modeling: a new style's forecast borrows the sales curve of similar past styles, matched on attributes like silhouette, fabric weight, and price point. This is the standard answer for fashion items, and its accuracy equals the quality of the similarity judgment.
Machine learning entered this stack by refining attribute matching and blending the families, and by the mid-2020s most large vendors sold some learned component. Published comparisons through that period showed modest error reductions over well-tuned classical methods on real assortments, concentrated on mid-lifecycle items, per academic benchmark studies; no published method made new-item forecasts reliable in an absolute sense.
Why do fashion forecasts fail?
The failure mechanisms are specific, and naming them is more useful than generic humility.
- No history: a first-season item's demand is a hypothesis, and analogous matching transfers the past only as well as the item truly resembles its ancestors.
- Assortment interdependence: styles in a range compete for the same customer, so forecasting each style independently ignores cannibalization, the effect whereby one item's sales subtract from a similar item's, which can exceed the base forecast error itself.
- Demand creation: fashion sales respond to placement, influencer exposure, and store-level presentation, variables that sit outside most models' inputs and sometimes outside anyone's control.
- Aggregation traps: a colorway or size-level error of twenty percent can hide inside a total-style forecast that looks accurate, and inventory is bought at size level.
The last point deserves emphasis, because it is where planning careers are lost. A style forecast that hits within ten percent at the total can still strand inventory if the size curve skews, since nobody returns the excess mediums for smalls.
How does a forecast become a purchase order?
The planning pipeline below reflects standard mid-market practice, and the forecast's role shrinks as it moves down the chain.
- Financial planning sets a season-level sales budget by category, which is a target rather than a forecast in the strict sense.
- Merchandise planning allocates that budget across styles and options, blending model output with merchant judgment on the new items.
- Statistical forecasting produces item-location projections for carry-over and replenishment styles, where the models carry the load.
- Buy plans convert projections into orders, layering on minimums, lead times, and vendor capacity, which reshape the numbers materially.
- In-season re-forecasting adjusts both orders and prices as actuals arrive, and this loop, not the pre-season forecast, contains most of the controllable error.
Step five is where sophisticated retailers concentrate effort. A pre-season forecast for a fashion item is a guess with error bars; an in-season process that reads first-two-weeks signals and reallocates or marks down quickly converts a bad guess into a survivable one.
Related stories: PLM Software Tracks a Garment From Sketch to Invoice, Line by Auditable Line · Whole-Garment Knitting Produces a Finished Sweater in One Cycle, With Catches.
How should forecast accuracy be judged?
By error measured at the level where decisions are made, and by comparability across methods. Mean absolute percentage error, the average of absolute errors as a percentage of actuals, is the trade's default metric despite its distortion at low volumes, where one unsold unit wrecks the percentage. Better practice tracks MAPE at style-color-size level for buys, uses volume-weighted error for totals, and holds a rolling benchmark, because a model's accuracy next season is the only accuracy that matters.
| Item class | Preferred method | Realistic error expectation |
|---|---|---|
| Replenishment basics | Time-series with seasonality | Tight bands, improving with history depth |
| Carry-over fashion | Time-series plus causal adjustments | Moderate bands |
| New fashion items | Analogous matching plus judgment | Wide bands; thirty to fifty percent MAPE commonly reported |
| Size-level curves | Historical size curves by fit block | Moderate; skewed by fit issues and returns |
Who forecasts well, and what do they do differently?
The organizations with defensible forecasting records, as visible in published retail case discussions through 2025, share three habits rather than one algorithm. They buy in steps, committing early only to a fraction of the season and holding capacity for in-season reaction. They measure error at buy level and publish it internally, which keeps judgment honest and models calibrated. And they integrate returns data, since a fit-driven return spike reads as demand to a naive forecast and poisons the next cycle. Academic forecasting research, much of it published through operations journals with roots at engineering schools such as Cornell's, has argued for decades that process design beats model choice; the industry's leaders behave accordingly.
The practical takeaway for 2026 planning teams: spend modeling effort on basics where it pays, spend process design on fashion where models cannot go, and distrust any vendor quotation of forecast accuracy that does not state the item class and the error level.
What role do returns play in poisoning or fixing the forecast?
Returns are the quiet variable in apparel forecasting, because a sale that comes back was never demand, and most planning systems record it as one. When fit problems drive a high return rate on a style, the sales ledger shows strong demand through the return window, models extrapolate enthusiastically, and the next buy repeats the error at larger volume. Forecasting literature and retail post-mortems through the 2020s repeatedly identified unadjusted return rates as a systematic bias that no algorithm removes, because the model sees only what the ledger reports.
The fix is unglamorous data plumbing. Returns must be attributed to the original sale date, classified by reason, and netted from demand before models train on the series, with fit-driven returns separated from fashion-driven ones because their timing and remedies differ entirely. A fit problem is a pattern-making issue wearing a forecasting costume; feeding it to the forecasting team produces markdowns when it should produce a pattern correction.
Handled properly, returns become an input rather than a poison. Return reason data at style and size level tells planners where size curves mislead, which fits run small in customer experience despite matching spec, and which carry-over styles deserve a second season with adjusted curves. The retailers with the most credible forecasting operations treat the returns pipeline as part of the forecasting system, funded and maintained accordingly, while the rest discover each season that their models were right about the sales and wrong about the demand.
