Shopping Agents Learn from Online User Feedback

Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng Dou· August 13, 2026 View original

Key takeaways

  • Online user feedback is a rich, underutilized resource for improving shopping agents.
  • LOFA combines RL with feedback-aware distillation for learning from raw interaction logs.
  • The framework captures both behavioral patterns and user-specific preferences.
  • LOFA significantly enhances recommendation quality and user satisfaction in e-commerce.

Who benefits

E-commerceRetailCustomer ServiceMarketingSocial Commerce

Summary

LOFA is a framework that enables large language model-based shopping agents to learn directly from real online user interaction logs without human annotation. It combines reinforcement learning with feedback-aware on-policy distillation to capture both behavioral patterns and user-specific preferences, significantly improving recommendation quality and user satisfaction.

Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating vast amounts of user interaction data. This data contains valuable implicit and explicit feedback that can be used to improve agent performance. However, existing approaches primarily rely on offline training signals, often overlooking the rich, yet heterogeneous, sparse, and noisy conversational feedback from users. To address these challenges, researchers propose LOFA (Learning from Online Feedback for Agents), a framework that allows shopping agents to learn directly from real online interaction logs without requiring human annotation. LOFA integrates reinforcement learning, which leverages verifiable purchase outcomes, with feedback-aware on-policy distillation. This distillation process identifies in-dialogue user directives and converts them into dense, token-level supervision. By combining these complementary objectives, LOFA effectively captures both collaborative behavioral patterns and individual user preferences, leading to consistent improvements in recommendation quality, response helpfulness, and alignment with user satisfaction in real-world e-commerce experiments.

Why it matters

E-commerce professionals and product managers can use this framework to develop more intelligent and responsive shopping agents that continuously improve by learning from actual user interactions, leading to higher conversion rates and customer satisfaction.

How to implement this in your domain

  1. 1Analyze existing user interaction logs to identify patterns and potential feedback signals.
  2. 2Implement a reinforcement learning loop that uses verifiable purchase outcomes as rewards for shopping agents.
  3. 3Develop a feedback-aware on-policy distillation mechanism to extract dense supervision from conversational feedback.
  4. 4Integrate LOFA into current shopping agent architectures to enable continuous online learning.
  5. 5Monitor key metrics like recommendation quality, response helpfulness, and user satisfaction to quantify improvements.

Original post by Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng Dou

"arXiv:2608.11604v1 Announce Type: new Abstract: Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents. However, exis…"

View on X

Originally posted by Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng Dou on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses