New Method Optimizes Policies Using Nondeterministic Causal Models

Jessica Lally, Milad Kazemi, Nicola Paoletti, David Watson, Sander Beckers· August 5, 2026 View original

Key takeaways

  • Traditional counterfactual policy optimization often oversimplifies real-world stochasticity.
  • Nondeterministic causal models better separate latent confounding from irreducible randomness.
  • The new framework enables more robust policy optimization in complex, uncertain environments.
  • This approach has been validated in critical applications like sepsis treatment simulation.

Who benefits

HealthcareAutonomous VehiclesRoboticsFinanceLogistics

Summary

Researchers propose a novel framework for robust counterfactual policy optimization that accounts for irreducible stochasticity in real-world systems, moving beyond the typical assumption of deterministic causal models. This approach uses nondeterministic causal models to separate latent confounding from inherent randomness, validated on a sepsis treatment simulator.

A new research paper introduces a method for optimizing policies based on counterfactual reasoning, specifically designed for systems with inherent randomness. Traditional counterfactual approaches often assume that all uncertainty comes from hidden variables, implying a deterministic underlying causal structure. This new framework, however, acknowledges and models irreducible stochasticity, which is common in real-world scenarios like Markov Decision Processes. The proposed method formalizes counterfactual policy optimization using probabilistic nondeterministic causal models. This allows for a clearer distinction between latent confounding factors and the system's intrinsic randomness. The researchers also present a practical optimization problem for identifying robust counterfactual policies, incorporating a sensitivity analysis framework. The efficacy of this approach was demonstrated through simulations of sepsis treatment, where diabetes status served as a hidden confounder.

Why it matters

This advancement allows for more robust and realistic policy optimization in complex, stochastic environments, leading to better decision-making in critical applications where uncertainty is inherent.

How to implement this in your domain

  1. 1Assess existing decision-making systems for assumptions about determinism versus stochasticity in causal models.
  2. 2Investigate the potential for applying nondeterministic causal models to improve policy robustness in high-stakes environments.
  3. 3Explore sensitivity analysis frameworks to understand the impact of irreducible stochasticity on policy outcomes.
  4. 4Collaborate with research teams to adapt this methodology for specific domain challenges, such as healthcare or autonomous systems.
  5. 5Develop simulation environments that accurately reflect both latent confounding and inherent randomness to test new policies.

Original post by Jessica Lally, Milad Kazemi, Nicola Paoletti, David Watson, Sander Beckers

"arXiv:2608.02893v1 Announce Type: new Abstract: Counterfactual inference approaches for sequential decision-making typically assume deterministic causal models, where all randomness stems from latent variables. However, Markov Decision Processes (MDPs) are inherently stochastic.…"

View on X

Originally posted by Jessica Lally, Milad Kazemi, Nicola Paoletti, David Watson, Sander Beckers on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses