DMRL Optimizes Advertising Recommendations with LLM Skill Documents.

Wei Zhang, Hongji Li, Song Sun, Peng Yu, Xue Yang, Lei Zhao, Peng Jiang· September 3, 2026 View original

Key takeaways

  • DMRL automates the optimization of LLM skill documents for advertising recommendations.
  • It uses structured editing actions and attributes rewards to specific document changes.
  • Dual-Relative Policy Optimization (DRPO) and a Long-term Reward Predictor (LRP) are key components.
  • DMRL significantly outperforms baselines in optimizing advertising metrics.

Who benefits

Advertising TechnologyE-commerceSocial MediaDigital MarketingMedia & Entertainment

Summary

Document-Mediated Reinforcement Learning (DMRL) is a new framework that uses structured editing actions on LLM skill documents to optimize advertising recommendation systems. It employs Dual-Relative Policy Optimization and a Long-term Reward Predictor to attribute rewards and estimate long-term outcomes, outperforming baselines on a large-scale platform.

A novel framework called Document-Mediated Reinforcement Learning (DMRL) has been developed to optimize skill documents used by large language models (LLMs) in advertising recommendation systems. Traditionally, tuning complex system parameters in advertising to balance commercial returns and user experience is a labor-intensive, prompt-driven process. DMRL addresses this by modeling skill document optimization as a sequence of structured editing actions, allowing for a principled mechanism to attribute rewards to specific document changes. DMRL operates with an upper-level agent that performs controlled edits to skill documents, while a frozen lower-level task agent evaluates the impact of these edits through A/B testing. To tackle the challenges of credit assignment and long-term outcome prediction, the framework introduces two key components: Dual-Relative Policy Optimization (DRPO), a post-training method for robust and risk-aware advantage estimation, and a Long-term Reward Predictor (LRP), which models population heterogeneity using disentangled representation learning and cross-attention transfer. Deployed on a large-scale short-video ads platform, DMRL demonstrated superior performance over state-of-the-art baselines across critical advertising metrics.

Why it matters

For professionals in advertising, marketing, and product management, DMRL offers a powerful, automated way to continuously optimize recommendation systems. It moves beyond manual prompt engineering, enabling more efficient and effective tuning of LLM-driven advertising strategies to maximize commercial returns and user satisfaction.

How to implement this in your domain

  1. 1Assess current LLM-based advertising recommendation systems for integration with DMRL's skill document optimization.
  2. 2Pilot DMRL on a specific advertising campaign to evaluate its impact on key performance indicators (KPIs).
  3. 3Develop or adapt tools for structured editing of LLM skill documents to facilitate the upper-level agent's actions.
  4. 4Train data science and marketing teams on the principles of reinforcement learning for dynamic ad optimization.

Original post by Wei Zhang, Hongji Li, Song Sun, Peng Yu, Xue Yang, Lei Zhao, Peng Jiang

"arXiv:2609.02170v1 Announce Type: new Abstract: Advertising recommendation requires continuously tuning complex system parameters while balancing commercial returns and user experience. Recent work has introduced large language models (LLMs) with skill documents to assist this la…"

View on X

Originally posted by Wei Zhang, Hongji Li, Song Sun, Peng Yu, Xue Yang, Lei Zhao, Peng Jiang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses