New Benchmark for Multimodal AI Shopping Agents Launched

Zeying Hao, Hao Guo, Mengtao Xu, Yimin Hu, Yuheng Song, Zesheng Zhou, Jinsong Lan, Xiaoyong Zhu· August 3, 2026 View original

Key takeaways

  • MMShopBench is the first real-log benchmark for multimodal, multi-turn AI shopping agents.
  • It addresses the limitations of text-only or synthetic benchmarks in evaluating complex shopping interactions.
  • Agents must infer requirements from both user images and dialogue, then retrieve and verify products.
  • Fine-tuning with the companion training set significantly improves open-source model performance.

Who benefits

E-commerceRetailCustomer ServiceMarketingAI/ML Development

Summary

MMShopBench is introduced as the first real-log benchmark for multimodal, multi-turn AI shopping agents, addressing the limitations of text-only or synthetic benchmarks. It uses annotated shopping logs to evaluate agents' ability to infer purchase intent from images and dialogue, retrieve products, and verify requirements.

The increasing reliance of online shoppers on AI assistants, often involving both images and multi-turn dialogue to express complex product needs, highlights a gap in current evaluation methods. Existing benchmarks for shopping agents typically rely on text-only or synthetic requests, failing to capture the intricacies of real-world multimodal interactions. To bridge this gap, researchers have developed MMShopBench, the first benchmark derived from real shopping logs for multimodal, multi-turn AI shopping agents. MMShopBench is built from meticulously cleaned and manually annotated shopping logs, providing ground-truth annotations for purchase intent and mandatory product requirements within each request. Agents evaluated on this benchmark must infer these requirements from a combination of user images and conversational dialogue. They then need to retrieve suitable candidate products using both image and text search, and finally, verify that these candidates meet all specified requirements by analyzing product images and structured attributes. The benchmark evaluates both open-source and proprietary models using an evidence-grounded multimodal protocol. A companion training set is also provided for fine-tuning open-source models. Experiments show that fine-tuning significantly narrows the performance gap between open-source models and leading proprietary solutions, demonstrating the effectiveness of the training data and the benchmark's utility in fostering reproducible experimentation within an offline shopping sandbox.

Why it matters

This benchmark provides a realistic and robust tool for developers to create and test more sophisticated AI shopping assistants. Professionals in e-commerce and AI development can use MMShopBench to build agents that better understand and fulfill complex customer needs, leading to improved user experience and sales.

How to implement this in your domain

  1. 1Utilize MMShopBench to evaluate and improve the multimodal capabilities of your AI shopping assistants.
  2. 2Integrate multimodal input (images and text) processing into your agent development pipeline for richer customer understanding.
  3. 3Develop or fine-tune models specifically for inferring purchase intent and mandatory product requirements from combined visual and textual cues.
  4. 4Implement robust product retrieval and verification mechanisms that leverage both image analysis and structured product attributes.
  5. 5Set up an offline shopping sandbox environment to conduct reproducible experiments and iterate on agent performance using real-log data.

Original post by Zeying Hao, Hao Guo, Mengtao Xu, Yimin Hu, Yuheng Song, Zesheng Zhou, Jinsong Lan, Xiaoyong Zhu

"arXiv:2607.29002v1 Announce Type: new Abstract: Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-…"

View on X

Originally posted by Zeying Hao, Hao Guo, Mengtao Xu, Yimin Hu, Yuheng Song, Zesheng Zhou, Jinsong Lan, Xiaoyong Zhu on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses