STUDIES
Forecasting Returns Before They Happen: Can LLM-Augmented Models Cut Reverse-Logistics Waste?
By Dr. Kevin Rudolph
03 Aug 2026
The short version
Ask whether AI can forecast product returns, and the market already has an answer: yes, and it’s already running in production. Narvar’s prediction engine scores return likelihood before an order leaves the warehouse, across more than 1,500 retail brands (Narvar, 2025). McKinsey has published a named playbook for it - six levers, from AI-led disposition routing to structured feedback loops - built on a market where US consumers return close to a trillion dollars in merchandise a year (McKinsey, 2026). So the interesting question isn’t whether this works. It’s two narrower things.
First, the systems already at scale, including the market leader, run on transformer-derived machine learning, not generative AI. Narvar says directly that IRIS “is not using generative AI,” even though it borrows techniques from the same family (VentureBeat, 2025). Whether large-language- model embeddings of product text, images, and return-reason narratives add anything on top of that is a real question, and a surprisingly unexplored one. A 2026 systematic review of 335 studies on AI in reverse logistics, published between 2015 and 2024, sorts the field into three technology buckets - evolutionary algorithms, machine learning and deep learning, and fuzzy logic - and GenAI and LLMs don’t appear as an established category in any of them. The review’s own research agenda names “the impact as well as the effective and responsible utilization of Generative AI (GenAI) and Large Language Models (LLMs)” as a priority for future work, not a solved problem (Moosabeiki Dehabadi et al., 2026).
Second, and probably the bigger of the two: even where a proven playbook is public and a vendor will sell it to you today, most companies haven’t captured it. McKinsey’s broader research on enterprise AI adoption finds that while 88% of organizations use AI somewhere, fewer than a third have started scaling it across the enterprise, and only about one in five have redesigned a workflow around it instead of layering it onto what already existed (McKinsey State of AI, 2025). That gap has nothing to do with whether the model works.
Here’s the narrower bet: LLM-augmented signal should add a real edge in fashion and apparel and in cold-start SKUs, where returns come down to subjective, textual things like fit, color, or “it didn’t look like the photo,” and should add almost nothing in stable commodity categories current systems already handle well. Separately, a correct forecast only creates value if someone changes the workflow around it, which is the step most companies still haven’t taken.
Executive summary
- AI-driven reverse logistics is already scaling, not experimental. A single vendor’s prediction engine already covers 1,500+ brands (Narvar, 2025); McKinsey frames AI-led disposition routing as one of six practical levers against a market where US returns approach $1 trillion a year and cost roughly $200 billion annually to process (McKinsey, 2026).
- It’s one piece of a bigger shift. Circular business models more broadly have become a real financial lever, not just a sustainability talking point - Bain & Company found over 70% of manufacturing executives expect circular solutions to grow revenue by 2027, and companies like Philips, Hydro, and Siemens Mobility are already booking the results (Bain & Company, 2025).
- The open technical question is narrower than “does AI help”: does LLM/embedding signal add anything on top of what’s already deployed? A 2026 review of 335 reverse-logistics AI studies (2015-2024) doesn’t even give GenAI/LLMs their own technology category, and names it as a future research priority rather than settled ground (Moosabeiki Dehabadi et al., 2026). That’s close to true white space.
- The bigger gap is organizational, not technical. Across enterprise AI generally, fewer than a third of organizations have begun scaling AI enterprise-wide, and only about one in five have redesigned a workflow around it rather than layering it on top of an unchanged process (McKinsey State of AI, 2025). There’s no reason to expect reverse logistics is an exception.
- There’s a real, time-limited prize for moving first. Industry commentary frames 2026 as the year reverse logistics stops being an overflow process and becomes core capacity, with an estimated 18-24 months of competitive advantage available before AI-driven reverse logistics becomes table stakes rather than a differentiator.
- What we’re testing: the LLM-specific accuracy increment in high-signal categories, and what it structurally takes to turn a correct forecast into a changed routing decision. Whether “AI can help reverse logistics” is no longer one of the open questions.
Why reverse logistics, why now
The scale alone explains why this stopped being a niche topic. US consumers are expected to return close to $850 billion in merchandise in 2025, around 15.8% of total retail sales (National Retail Federation / Happy Returns, 2025); McKinsey separately puts 2024 US returns at nearly $1 trillion in merchandise value, with retailers spending roughly $200 billion a year processing and recovering value from it (McKinsey, 2026). The two figures come from different sources, years, and scopes, but they agree on the order of magnitude: this is a real cost line, not a rounding error. It’s concentrated unevenly, too - apparel and footwear see far higher return rates than electronics or home goods, driven by fit, color-as-photographed, and try-before-you-keep shopping behavior.
Regulation adds a second, less obvious push. The EU’s circular-economy agenda - the Ecodesign for Sustainable Products Regulation (ESPR) and its Digital Product Passport requirements, Extended Producer Responsibility (EPR) schemes expanding from packaging into textiles and electronics, and the Corporate Sustainability Reporting Directive (CSRD) - is pushing larger retailers toward far more granular tracking of product condition, return reason, and disposition outcome than reverse logistics has traditionally captured. The useful implication, if it holds up, is that compliance-driven data capture and forecasting-model training data are becoming the same dataset, without anyone having designed it that way on purpose.
Returns forecasting is one piece of a bigger shift
Reverse logistics is the operational slice of a much bigger story. Circular business models generally have moved from a sustainability talking point to a real line on the P&L. In a September 2025 survey, Bain & Company found that more than 70% of manufacturing executives expect circular solutions to increase revenue by 2027, and more than half expect them to cut costs despite the upfront investment required. Tellingly, 97% of companies already running circular programs say they’re doing it for reasons beyond sustainability (Bain & Company, 2025).
The examples span well beyond retail returns. Philips’s own 2024 annual report puts circular revenues at 24% of sales, largely from refurbishing and reselling medical equipment it takes back from hospitals (Philips, 2026). Aluminum producer Hydro is building toward 1,200 kilotons of recycled aluminum a year by 2030, a plan Bain estimates could add roughly $750 million in EBITDA, because recycled aluminum takes only about 5% of the energy that primary production does. Siemens Mobility uses predictive maintenance to cut maintenance costs by around 15%. Trane Technologies has grown a European HVAC rental business - equipment as a service instead of a one-time sale - by double digits annually (Bain & Company, 2025). None of these are returns-desk problems. They’re business-model decisions.
What connects these examples is Bain’s own read on what separates companies that capture circular value from companies that just talk about it: access, specifically preferential access to the physical material or the information needed to plan around it. A returns-forecasting model is a small instance of exactly that. Knowing what’s coming back, in what condition, and when, before it arrives, is the same kind of informational edge that lets Philips route equipment to refurbishment instead of scrap, or lets a rental business plan its fleet. The technical question here is small. The strategic category it sits inside is not.
Where the technology already works
Narvar’s IRIS engine, trained on more than 74 billion consumer interactions a year, already predicts which orders are likely to come back before they ship, and routes the returns process accordingly across its 1,500+-brand network (Narvar, 2025). McKinsey’s six levers - managing demand, improving data and insight, AI-led decisioning, operational optimization, re-commerce strategies, and structured feedback loops - read like an implementation checklist, not a research agenda (McKinsey, 2026). And individual retailers report real, measured results: footwear brand Flux, using ReturnGO’s returns platform, cut its return rate by roughly 7 points to 15.9% (ReturnGO, n.d.; U.S. Chamber of Commerce, 2026), and apparel retailer Piper & Scoot got 63.5% of returning customers to take an exchange instead of a refund, retaining about $55 per transaction that would otherwise have walked out the door (U.S. Chamber of Commerce, 2026).
None of that runs on generative AI, though. Narvar says directly that IRIS doesn’t use it, even though it draws on transformer-derived techniques from the same family (VentureBeat, 2025). A 2026 systematic review of 335 studies on AI in reverse logistics, published between 2015 and 2024, groups the field into three technology buckets - evolutionary algorithms and metaheuristics, machine learning and deep learning, and fuzzy logic - and generative AI isn’t one of them. Instead, the review’s authors list “the impact as well as the effective and responsible utilization of Generative AI (GenAI) and Large Language Models (LLMs) in human-AI and human-machine collaborative RL operations” as one of several open items on their own research agenda (Moosabeiki Dehabadi et al., 2026). So the honest way to describe the gap isn’t “can AI forecast returns” - that’s answered
- it’s whether the specific increment of LLM-embedded text and image signal improves on what’s already running in production.
Two generations of forecasting model
Current-generation production forecasting, the kind Narvar and its peers already run, combines classical time-series methods (ARIMA, Prophet, exponential smoothing) with tabular machine learning: gradient-boosted trees over structured features like category, price tier, promotional calendar, and transformer-derived behavioral sequences. It’s mature, well understood, and, per the results above, already earning its keep.
The next increment worth testing adds signal current systems still can’t use directly: embeddings of product descriptions and reviews (capturing fit-related language and quality complaints), embeddings of product images (capturing style and color ambiguity that correlates with returns), and embeddings of free-text return-reason narratives from similar past returns. In principle this should matter most exactly where current systems are weakest - a brand-new product with no sales history has nothing for a transactional model to learn from, but a text or image embedding can still borrow a reasonable prior from similar products that do have history.
The experiment: where the uplift shows up
Design (to be run for real in a later step). Benchmark a current-generation baseline (tabular ML plus classical seasonality, approximating what production systems already do) against an LLM-augmented hybrid (the same features plus text and image embeddings), segmented by category and by SKU maturity (established SKUs with 12+ months of history versus new SKUs under 8 weeks live). Evaluation metric: mean absolute percentage error (MAPE) on weekly return-volume forecasts, backtested across multiple seasons. Given the literature gap above, even a modest, honestly-reported version of this benchmark would be one of very few real-world data points to exist on the question.
Illustrative result (hypothesis, not measured):
| Segment | Typical return rate | Current-generation baseline (MAPE) | + LLM-augmented (MAPE) | Accuracy gain from LLM augmentation |
|---|---|---|---|---|
| Apparel & footwear | ~25-30% | 18% | 11% | -7 pts (large) |
| Beauty & personal care | ~10-15% | 13% | 9% | -4 pts (moderate) |
| Consumer electronics | ~8-10% | 9% | 8% | -1 pt (small) |
| Home & furniture (staple SKUs) | ~5-8% | 7.5% | 7.3% | ~0 pts (negligible) |
| New SKUs, any category, <8 weeks live | highly variable | 28% | 16% | -12 pts (largest gain) |
(Illustrative - hypothesis, not measured. The pattern we expect, and want to test: LLM augmentation earns its keep where current production systems are weakest, on subjective/textual return drivers and cold-start SKUs, and adds little where a current-generation system already forecasts well.)
If this pattern holds, it argues against retrofitting LLM signal across the whole catalog and for a targeted increment: apparel and new-SKU cold-start are where the remaining accuracy gap, and therefore the business case for going beyond what’s already deployed, is largest.
From forecast accuracy to network design
Disposition routing - deciding whether a returned item should be restocked, refurbished, resold, liquidated, or recycled, before it’s even back in the building - is already one of McKinsey’s named levers, not a novel proposal (McKinsey, 2026). The more specific question is whether a better, LLM-augmented pre-arrival forecast changes that routing decision often enough to matter, versus what a current-generation system would have routed anyway.
Illustrative economics (hypothesis, not measured). For a mid-size apparel-heavy retailer processing roughly 2 million returns a year at an average reverse-logistics handling and value-decay cost of $6-9 per unit, an LLM-augmented forecast that further cuts error-driven cost by 15-20% in the apparel-weighted portion of the catalog, the segment where the accuracy gain above is largest, would be worth a low-to-mid single-digit-million-dollar annual saving on top of whatever a current-generation system already captures. These figures are placeholders for a real unit-economics model, not a forecast.
Figure 1 (concept, to be turned into a network diagram). Two states worth drawing side by side: current-generation, where disposition is decided on arrival using a transactional-behavior model, and LLM-augmented, where a richer pre-arrival forecast, informed by product text and image embeddings, changes the routing decision for a defined slice of high-signal SKUs before the item ships back. The value isn’t a prettier diagram: it’s naming, precisely, which routing decisions flip.
The gap between capability and adoption
This is arguably the more important half of the story, and it’s the part most reverse-logistics AI commentary skips entirely. The tools exist. The playbook is public - McKinsey published it. A vendor will sell you a version of it today, already running across 1,500+ brands. And most companies still haven’t captured it, for reasons that have nothing to do with whether the model works.
Across enterprise AI generally, the gap between adoption and scale is stark and well documented: 88% of organizations use AI in at least one business function, and roughly 79% use generative AI, but fewer than a third have begun scaling AI across the enterprise, and only about 7% have scaled it enterprise-wide with measurable EBIT impact. Most tellingly, only about one in five organizations using generative AI have redesigned even some of their workflows around it. The rest are layering AI onto a process that hasn’t otherwise changed (McKinsey State of AI, 2025). Reverse logistics tells a related version of the same story. The Moosabeiki Dehabadi et al. (2026) review finds AI methods already applied across nearly every stage of the field - collection, recycling, reuse, repair, remanufacturing, refurbishing, disposal - drawn from 335 studies published between 2015 and 2024. The capability is documented and has been for years. The authors’ own agenda for future work is about extending and operationalizing those methods further, not proving they work in the first place - the same “we know it works, now what” position enterprise AI is in more broadly.
That combination is the strategic picture worth acting on: the firms that win the estimated 18-24 month window before AI-driven reverse logistics becomes table stakes won’t be the ones with access to the best model. Access is already commoditized - a vendor sells it. They’ll be the ones that change how the return workflow runs once the forecast exists: pre-approving disposition decisions, restructuring refurb-hub capacity planning, and giving the forecast the authority to change a routing decision instead of just producing a number a human reviews and routes the old way anyway.
Three bold calls
- The technology argument is over; the execution argument just started. Nobody serious is still asking whether AI can forecast returns. The competitive question is whether your organization will restructure the return workflow around that forecast before your competitors do, inside a window industry commentary already estimates at 18-24 months.
- The white space is narrower than “AI in reverse logistics”: it’s the LLM-specific increment. A 335-study review of the field doesn’t give GenAI or LLMs their own category and lists them as future work, not settled ground. That’s an unusually clean opportunity to be first with a real answer, but it’s a much smaller claim than “we brought AI to returns.”
- Redesigning the workflow beats adding the model. McKinsey’s own data says only about one in five organizations using generative AI have redesigned a workflow around it. In reverse logistics, that means giving the forecast authority to change a disposition decision, not just displaying it next to the decision a person was already going to make the old way.
What could go wrong
- The most common failure mode is conflating proof that the technology works with proof that you’re capturing value from it. Buying or piloting a model without restructuring the workflow around it reliably produces a dashboard, not a result.
- Cold start cuts both ways. Embeddings borrow signal from “similar” products, but similarity is itself a modeling choice; a novel product can get a confidently wrong prior from products that only look similar on the surface.
- Return-reason text is messy and inconsistent across customers and channels; a model trained on noisy free-text labels can learn the noise instead of the signal.
- Distribution shift is the norm in fashion in particular - a model tuned on last season’s return patterns may be a poor guide to a style cycle that changed underneath it.
- Compliance data collected for reporting isn’t automatically ML-ready. Data captured to satisfy a regulator’s checkbox and data structured well enough to train a model are not the same thing; closing that gap is real engineering work, not a side effect of compliance.
- Automated disposition routing raises its own governance question under emerging EU AI guidance for automated decision-making, worth scoping explicitly before a forecast gets the authority to route an item without human review.
The bottom line
The technology question is closed. What’s left splits into two: a narrower technical test, and a bigger organizational one - and the organizational point extends well past returns, into circular business models more broadly. Technically: test whether LLM-augmented signal beats what a current-generation system already does, on the highest-return, most textually-rich slice of the catalog plus cold-start SKUs, against an honest current-generation baseline, not a strawman classical model nobody would deploy in 2026. Organizationally: don’t stop at a better number. Give the forecast the authority to change a disposition decision before the item is back in the building, and do that inside the window that’s still open, because it won’t stay open long.
Sources and further reading
- National Retail Federation (2025) ‘Consumers expected to return nearly $850 billion in merchandise in 2025’, 24 October. Available at: https://nrf.com/media-center/press-releases/consumers-expected-to-return-nearly-850-billion-in-merchandise-in-2025 (Accessed: 22 August 2026). Figures from the 2025 Retail Returns Landscape report (NRF/Happy Returns): ~$850 billion in expected US merchandise returns for 2025, 15.8% of total retail sales.
- McKinsey & Company (2026) ‘From cost center to competitive advantage: Modernizing reverse logistics with AI’. Available at: https://www.mckinsey.com/industries/logistics/our-insights/from-cost-center-to-competitive-advantage-modernizing-reverse-logistics-with-ai (Accessed: 22 August 2026). ~$1 trillion in US merchandise returns (2024), ~$200 billion annual processing cost, six-lever modernization framework, AI-led disposition routing.
- McKinsey & Company (2025) ‘The state of AI in 2025: Agents, innovation, and transformation’. Available at: https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai (Accessed: 22 August 2026). Enterprise AI/gen-AI adoption versus scaling-gap statistics (88% using AI somewhere, ~1/3 scaling enterprise-wide, ~21% workflow redesign, ~7% scaled with EBIT impact).
- Bain & Company (2025) ‘Circular Business Models Unlock New Profit and Growth’ (CEO Sustainability Guide 2025), 15 September. Available at: https://www.bain.com/insights/circular-business-models-unlock-new-profit-and-growth-ceo-sustainability-guide-2025/ (Accessed: 22 August 2026). Executive survey statistics and the Hydro/Cisco/Siemens Mobility/Trane examples used above; fetched directly, high confidence.
- Philips (2026) Annual Report 2024 - Impact with Care, April. Available at: https://www.results.philips.com/app/uploads/2026/04/PhilipsFullAnnualReport2024-English.pdf (Accessed: 22 August 2026). Circular revenues at 24% of sales.
- Narvar (2025) ‘Narvar Introduces IRIS: The AI Engine Powering the Future of Post-Purchase’, PR Newswire, 9 January. Available at: https://www.prnewswire.com/news-releases/narvar-introduces-iris-the-ai-engine-powering-the-future-of-post-purchase-launches-narvar-assist-as-its-first-iris-powered-solution-302346999.html (Accessed: 22 August 2026). 1,500+ brand network, 74B+ annual consumer interactions; IRIS/Narvar Assist went generally available 15 January 2025.
- VentureBeat (2025) ‘How Narvar is using AI and data to enhance post-purchase customer experiences’, 22 December. Available at: https://venturebeat.com/ai/how-narvar-is-using-ai-and-data-to-enhance-post-purchase-customer-experiences (Accessed: 22 August 2026). Source for IRIS explicitly not being generative-AI-based, while drawing on transformer-derived techniques.
- Moosabeiki Dehabadi, F., Shahbazi Manshadi, S., Appolloni, A., Sun, X. and Yu, H. (2026) ‘How can artificial intelligence enhance reverse logistics practices: a systematic literature review and research agenda’, Cleaner Logistics and Supply Chain, 19, 100327. Available at: https://doi.org/10.1016/j.clscn.2026.100327 (Accessed: 22 August 2026). Review of 335 studies published 2015-2024; categorizes AI into evolutionary algorithms/metaheuristics, machine learning/deep learning, and fuzzy logic - GenAI/LLMs are not an established category and are named in the authors’ own future-research agenda instead.
- U.S. Chamber of Commerce (2026) ‘AI is revolutionizing product returns - and every business can benefit’, CO- by US Chamber of Commerce, 10 February. Available at: https://www.uschamber.com/co/good-company/launch-pad/ai-product-returns (Accessed: 22 August 2026). Source for the Flux Footwear (-7 points to 15.9%) and Piper & Scoot (63.5% exchange rate, ~$55 retained per transaction) case results used above.
- ReturnGO (n.d.) ‘Flux Footwear Case Study’, quoting Annie Phan, Head of Customer Experience, Flux Footwear. Available at: https://returngo.ai/case-study/flux-footwear-case-study/ (Accessed: 22 August 2026). Corroborating source naming ReturnGO as the specific returns platform behind the Flux Footwear result above; no publication date given on the page.
ABOUT THE AUTHOR(S)
Dr. Kevin Rudolph is a Managing Director in LoopSmart's Pegeia office (Cyprus).