Inference-Time Alignment Improves Fairness in Reinforcement Learning.

Umer Siddique, Peilang Li, Conor Wallace, Yongcan Cao· August 4, 2026 View original

Key takeaways

  • Fairness in RL can be achieved at inference time without costly policy retraining.
  • A multiplicative policy shaping framework adjusts action probabilities based on welfare scores.
  • This method significantly improves fairness objectives while preserving core task performance.
  • The approach is general and compatible with any deep RL agent.

Who benefits

HealthcareFinanceSocial MediaAutonomous SystemsHuman Resources

Summary

This research introduces a method to adjust pretrained reinforcement learning policies at inference time to achieve welfare-based fairness objectives without retraining the base policy. It proposes a multiplicative policy shaping framework that modifies action probabilities using welfare scores.

Deep reinforcement learning (RL) agents often struggle to adapt to new performance criteria like fairness once deployed, typically requiring costly retraining. This paper explores a novel approach inspired by large language model alignment, focusing on steering pretrained RL policies towards fairness objectives during inference. The proposed multiplicative policy shaping framework adjusts action probabilities based on action-dependent welfare scores, effectively modifying behavior without altering the underlying policy parameters. Extensive experiments across various domains demonstrate that this inference-time policy shaping significantly enhances welfare-based fairness while maintaining core task performance. This method offers a flexible and efficient way to incorporate new preferences into existing RL systems, addressing a key challenge in deploying adaptable and ethical AI.

Why it matters

Professionals can deploy more adaptable and ethically aligned AI systems by integrating fairness objectives post-training, reducing the need for expensive and time-consuming full model retraining.

How to implement this in your domain

  1. 1Evaluate existing RL models for potential fairness biases in deployment scenarios.
  2. 2Define specific welfare-based fairness objectives relevant to your application.
  3. 3Explore implementing inference-time policy shaping techniques to adjust model behavior.
  4. 4Monitor the impact of fairness adjustments on both fairness metrics and core task performance.
  5. 5Iteratively refine welfare scores and shaping parameters to optimize for desired outcomes.

Original post by Umer Siddique, Peilang Li, Conor Wallace, Yongcan Cao

"arXiv:2608.00175v1 Announce Type: new Abstract: Deep reinforcement learning (RL) agents achieve strong performance by optimizing scalar reward functions. However, once deployed, the policies of these RL agents are often rigid and costly to adapt to new performance criteria. For i…"

View on X

Originally posted by Umer Siddique, Peilang Li, Conor Wallace, Yongcan Cao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses