ERQDP Solves Long-Term Decision Making Under Risk

Irmaan (Mohammad), Mirzanejad, Nadjet Bourdache, Abdel-Illah Mouaddib· July 23, 2026 View original

Summary

This paper introduces ERQDP, an enumeration-free and sampling-free method for finite-horizon Markov Decision Process (MDP) planning under root-based risk objectives. ERQDP provides certified solutions or explicit gaps, enabling efficient risk-parameter sweeps.

This research addresses the challenge of long-term sequential decision-making in environments where risk is a critical factor. Specifically, it focuses on finite-horizon Markov Decision Process (MDP) planning under "root-based" risk objectives, which apply a rank-dependent functional to the distribution of total returns. These objectives are non-linear and typically break standard Bellman optimality, making direct optimization computationally intensive. The paper proposes ERQDP (Enumeration-free Rank-Quantile Dynamic Programming), a novel method that avoids scenario-tree enumeration and sampling. ERQDP solves a rank-quantile surrogate problem using exact dynamic programming, then evaluates candidate policies precisely by applying dynamic programming over return Probability Mass Functions (PMFs) on a discretized grid. It refines the surrogate in an anytime loop, providing explicit upper-lower bounds (certificates) for the target objective. Benchmarks show ERQDP delivers certified solutions or clear residual gaps, significantly speeds up risk-parameter sweeps, and supports both risk-averse and risk-seeking behaviors.

Why it matters

Professionals in fields requiring complex decision-making under uncertainty, such as finance or logistics, can leverage this method to optimize strategies with explicit risk considerations and performance guarantees.

How to implement this in your domain

  1. 1Evaluate ERQDP for optimizing sequential decision-making problems in your domain, especially those with non-linear risk objectives.
  2. 2Consult with quantitative analysts or researchers to understand how rank-dependent functionals apply to your specific risk profiles.
  3. 3Develop or adapt existing MDP planning tools to incorporate ERQDP's approach for risk-aware policy generation.
  4. 4Utilize the method's ability to provide explicit upper-lower gaps to certify solution quality in critical applications.

Who benefits

Financial ServicesLogisticsSupply ChainInsuranceEnergy

Key takeaways

  • ERQDP is a new method for MDP planning under complex, root-based risk objectives.
  • It avoids computationally intensive enumeration and sampling.
  • The method provides certified solutions or explicit performance gaps.
  • ERQDP enables efficient exploration of both risk-averse and risk-seeking strategies.

Original post by Irmaan (Mohammad), Mirzanejad, Nadjet Bourdache, Abdel-Illah Mouaddib

"arXiv:2607.19914v1 Announce Type: new Abstract: We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and gener…"

View on X

Originally posted by Irmaan (Mohammad), Mirzanejad, Nadjet Bourdache, Abdel-Illah Mouaddib on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses