ML vs. Value Sorting for Shipment Prioritization.

Jize Li· July 22, 2026 View original

Summary

This paper evaluates whether machine learning models for delay-risk prediction outperform a simple value-sorting baseline for prioritizing shipments. Across three supply-chain datasets, ML often fails to consistently beat sorting by highest-value shipments, highlighting the importance of severity learnability and calibration.

In supply chain management, models predicting delay risk are typically judged by their predictive accuracy. However, this research argues that the practical utility for managers, who can only review a limited number of shipments, lies in identifying which shipments to check first. The study investigates whether machine learning (ML) models can surpass a straightforward baseline: prioritizing shipments based solely on their highest monetary value. Using three real-world supply chain datasets—SCMS procurement, DataCo logistics, and Olist e-commerce—the researchers employed leakage-controlled rolling-origin evaluation and bootstrap confidence intervals. They found that ranking by predicted delay severity multiplied by known value (M1) generally outperformed severity-only ranking. However, M1 did not consistently beat the simple value-sorting approach across all datasets. The key differentiator appeared to be the "severity learnability" of the data. DataCo, with a higher R^2 and low calibration bias, showed ML outperforming value sorting. In contrast, SCMS and Olist, with poor R^2 and negative calibration bias, saw ML underperform. The study concludes that value sorting should remain a permanent benchmark, and ML models should only be deployed for prioritization after rigorous auditing of severity learnability and calibration under robust evaluation protocols.

Why it matters

Supply chain professionals and data scientists need to critically assess the real-world utility of ML models for prioritization, ensuring they genuinely add value beyond simpler, more transparent heuristics like value sorting.

How to implement this in your domain

  1. 1Establish value sorting as a mandatory baseline for any ML model designed for prioritization tasks.
  2. 2Implement leakage-controlled rolling-origin evaluation for all predictive models in production.
  3. 3Rigorously audit the severity learnability and calibration of ML models before deployment.
  4. 4Develop clear metrics for "practical utility" beyond just predictive accuracy for prioritization tasks.
  5. 5Educate stakeholders on the limitations of ML when simpler heuristics are highly effective.

Who benefits

Supply ChainLogisticsE-commerceManufacturingRetail

Key takeaways

  • ML models for prioritization must be benchmarked against simple value sorting.
  • Predictive accuracy alone is insufficient; practical utility matters more.
  • Severity learnability and model calibration are crucial for ML effectiveness.
  • Value sorting can often be a strong, no-model baseline.

Original post by Jize Li

"arXiv:2607.18573v1 Announce Type: new Abstract: Delay-risk models are usually judged by predictive accuracy. What matters in practice is narrower: with capacity to review only a few shipments, which ones should a manager check first? We evaluate whether machine learning clears a…"

View on X

Originally posted by Jize Li on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses