Causal Bandits Improve Decision-Making with Non-Manipulable Variables

Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi· July 20, 2026 View original

Summary

This research introduces new methods, Causal Thompson Sampling and Causal Information-Directed Sampling, for contextual causal bandits that incorporate non-manipulable variables. These approaches leverage known causal graphs to share information across interventions, leading to faster identification of optimal decisions and outperforming existing baselines.

The study addresses a critical challenge in causal bandits: situations where some influential variables cannot be directly controlled or manipulated, yet still provide valuable information for decision-making. Traditional causal bandit frameworks often struggle with these non-manipulable variables, limiting their effectiveness in complex real-world scenarios. Researchers developed novel algorithms, Causal Thompson Sampling and Causal Information-Directed Sampling (IDS), specifically designed for contextual causal bandits with such variables. These methods operate under the assumption of a known causal graph without latent confounding, using a Bayesian formulation where conditional probability tables represent unknown parameters. This allows observations from one intervention to inform reward estimates for others through their shared causal mechanisms. The proposed methods demonstrated superior performance over both causal and non-causal baselines in synthetic causal bandit tasks. This improvement stems from their ability to more effectively exploit shared information across different interventions, leading to faster and more accurate identification of high-reward decisions. The research also provides theoretical regret bounds for both algorithms, quantifying their efficiency.

Why it matters

For professionals dealing with complex decision-making in systems with interdependent variables, this research offers advanced methods to optimize outcomes even when some influencing factors are beyond direct control. It promises more efficient and accurate policy learning in dynamic environments.

How to implement this in your domain

  1. 1Map out the causal relationships within your system, identifying both manipulable and non-manipulable variables, to construct a causal graph.
  2. 2Implement Causal Thompson Sampling or Causal Information-Directed Sampling in A/B testing or recommendation systems where interventions have ripple effects.
  3. 3Collect data on contextual variables and post-intervention observations to feed into the Bayesian model for continuous learning.
  4. 4Evaluate the performance of these causal bandit algorithms against existing multi-armed bandit or reinforcement learning approaches in your specific domain.

Who benefits

HealthcareE-commerceMarketingFinanceLogistics

Key takeaways

  • New causal bandit algorithms can effectively handle non-manipulable variables in decision-making systems.
  • These methods leverage causal graphs to share information across interventions, improving learning efficiency.
  • Causal Thompson Sampling and Information-Directed Sampling outperform traditional baselines in identifying optimal decisions.
  • The approach is particularly valuable in complex systems where not all influencing factors can be directly controlled.

Original post by Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi

"arXiv:2607.15577v1 Announce Type: new Abstract: Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification of high-reward decisions. In many applications, however, some variables cannot be directly m…"

View on X

Originally posted by Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses