New Sampling Methods Improve Causal Bandits with Non-Manipulable Variables

Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi· July 20, 2026 View original

Summary

This research introduces causal variants of Thompson Sampling and Information-Directed Sampling (IDS) for contextual causal bandits with non-manipulable variables. These methods exploit causal graph structures to share information across interventions, leading to faster identification of high-reward decisions and improved regret bounds.

This paper explores contextual causal bandits in scenarios where some variables cannot be directly manipulated, yet they provide crucial information for decision-making and influence rewards. The research adopts a Bayesian framework, assuming a known causal graph without latent confounding, where the unknown parameters are the conditional probability tables of the observational distribution. This setup allows observations from one intervention to inform reward estimates for other interventions by leveraging shared causal mechanisms. The authors develop causal adaptations of two prominent bandit algorithms: Thompson Sampling and Information-Directed Sampling (IDS). For Thompson Sampling, they establish an entropy-dependent sublinear Bayesian regret bound. For IDS, they derive a similar entropy-dependent regret bound that explicitly accounts for errors introduced by Monte Carlo approximations. When exact quantities are available, this bound recovers the standard sublinear IDS rate. High-probability confidence bounds are also provided for the Monte Carlo estimates. Experimental results on several synthetic causal bandit tasks demonstrate that these proposed causal methods significantly outperform both causal and non-causal baselines. This superior performance is attributed to their more effective exploitation of information shared across different interventions, leading to accelerated identification of optimal decisions.

Why it matters

This research provides more efficient and robust algorithms for decision-making in complex systems where interventions have causal effects and some variables are beyond direct control. Professionals in fields like personalized medicine, marketing, and policy optimization can use these methods to make better, data-driven decisions with fewer experiments.

How to implement this in your domain

  1. 1Apply causal bandit algorithms to optimize sequential decision-making in complex systems.
  2. 2Incorporate known causal graph structures to improve information sharing across interventions.
  3. 3Utilize causal Thompson Sampling or Information-Directed Sampling for better exploration-exploitation trade-offs.
  4. 4Design experiments that account for non-manipulable variables and their informational value.

Who benefits

HealthcareMarketingPolicy MakingE-commerce

Key takeaways

  • Causal bandits improve decision-making by exploiting structural relationships.
  • New causal Thompson Sampling and IDS variants handle non-manipulable variables.
  • These methods share information across interventions more effectively.
  • They achieve superior performance and improved regret bounds in experiments.

Original post by Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi

"arXiv:2607.15577v1 Announce Type: cross Abstract: Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification of high-reward decisions. In many applications, however, some variables cannot be directly…"

View on X

Originally posted by Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses