LEMUR Aligns Multi-Objective RL with Human Preference Feedback

Manith Adikari, Bei Peng, Samuele Vinanzi, Angelo Cangelosi· August 3, 2026 View original

Key takeaways

  • LEMUR enables multi-objective RL agents to learn from human preference feedback.
  • It jointly learns policies and objective-specific reward models.
  • The framework addresses challenges of specifying reward functions for competing objectives.
  • LEMUR demonstrates superior performance on various multi-objective tasks.

Who benefits

RoboticsAutonomous SystemsGamingHealthcareFinance

Summary

LEMUR is a novel framework that enables Multi-Objective Reinforcement Learning (MORL) agents to learn optimal policies by interactively learning from multiple human preferences, jointly learning policies and objective-specific reward models without predefined reward functions.

Traditional Reinforcement Learning (RL) typically relies on a single, well-defined scalar reward function, which is often difficult to specify for real-world tasks involving multiple, competing objectives like performance versus efficiency. While Multi-Objective RL (MORL) addresses these trade-offs by using vector rewards, it still assumes access to pre-specified reward functions for each objective. Preference-based RL (PbRL) has shown promise in learning without predefined reward functions by using human feedback, but it has largely been confined to single-objective settings. The LEMUR framework bridges this gap by introducing a novel approach that allows MORL agents to learn optimal multi-objective policies directly from the preferences of multiple humans. LEMUR jointly learns both the policies and multiple objective-specific reward models based on interactive human feedback. This enables agents to effectively balance competing objectives during the learning process, even when ground-truth reward functions are unavailable. Empirical results on various benchmark multi-objective tasks demonstrate LEMUR's superior performance compared to existing baseline methods, offering a promising direction for complex decision-making scenarios.

Why it matters

Professionals developing AI systems for complex real-world scenarios with conflicting objectives can use LEMUR to train agents more effectively by incorporating nuanced human preferences, leading to more aligned and adaptable AI behaviors.

How to implement this in your domain

  1. 1Explore integrating preference-based learning into multi-objective reinforcement learning projects.
  2. 2Design systems for collecting and interpreting human preferences for multiple, competing objectives.
  3. 3Develop methods to jointly learn policies and objective-specific reward models from human feedback.
  4. 4Evaluate LEMUR's approach for training agents in complex decision-making tasks without predefined reward functions.

Original post by Manith Adikari, Bei Peng, Samuele Vinanzi, Angelo Cangelosi

"arXiv:2607.29559v1 Announce Type: new Abstract: Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, competing objectives, such as performance versus effi…"

View on X

Originally posted by Manith Adikari, Bei Peng, Samuele Vinanzi, Angelo Cangelosi on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses