New Benchmark Evaluates Personalized LLMs with User Behavior.

Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Jiang Feng, Min Yang· August 7, 2026 View original

Key takeaways

  • LUNAR benchmarks personalized LLMs using real-world, cross-domain user behavior logs.
  • Effective personalization requires intelligent evidence selection and cross-domain integration.
  • More context or larger models alone do not guarantee better personalization.
  • There is a critical trade-off between personalization strength and user privacy.

Who benefits

E-commerceMarketingSoftware DevelopmentCustomer ServiceAdTech

Summary

LUNAR is the first benchmark to evaluate personalized LLMs using longitudinal app interaction histories across diverse daily-life domains, revealing that effective personalization requires careful evidence selection and cross-domain integration, and that more context or larger models don't guarantee better performance. It also highlights the trade-off between personalization and privacy.

Current benchmarks for personalized Large Language Models (LLMs) often rely on static textual personas or isolated behavioral signals, which fall short in evaluating how LLMs personalize responses based on a user's heterogeneous, real-world daily activities across multiple domains. To address this critical gap, researchers have introduced LUNAR, a pioneering benchmark specifically designed to assess LLMs' ability to personalize responses using longitudinal app interaction histories. This includes diverse domains such as clothing, food, housing, and mobility. LUNAR's construction employs a multi-stage, coarse-to-fine synthesis pipeline, grounded in real-world behavioral patterns, to ensure scalability while mitigating data sparsity and privacy concerns. Fidelity analyses confirm that LUNAR's synthetic data closely aligns with actual behavioral distributions, making it a robust tool for evaluation. Experiments conducted on 19 mainstream LLMs using LUNAR yielded significant insights. The findings indicate that while access to behavioral logs is necessary for deep personalization, it is not sufficient on its own. Simply providing more context or using larger models does not automatically guarantee improved performance. Instead, effective personalization critically depends on the intelligent selection and integration of relevant evidence across different domains. Furthermore, the research highlights a crucial trade-off: stronger personalization often comes at the cost of privacy protection. Direct retrieval of fine-grained behavioral records consistently outperformed compressed memory approaches, underscoring the challenges in balancing utility and privacy for personalized LLMs.

Why it matters

This benchmark is vital for professionals developing personalized AI applications, as it provides a realistic framework to evaluate and improve LLMs' ability to understand and respond to individual user behaviors across diverse contexts, while also emphasizing the critical need to balance personalization with privacy.

How to implement this in your domain

  1. 1Re-evaluate personalization strategies for LLM-powered products, considering the need for cross-domain behavioral data integration.
  2. 2Investigate methods for intelligent evidence selection from user interaction logs to improve personalization effectiveness.
  3. 3Prioritize privacy-preserving techniques when designing systems that leverage extensive user behavioral data.
  4. 4Benchmark existing personalized LLM solutions against LUNAR-like scenarios to identify performance gaps and areas for improvement.

Original post by Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Jiang Feng, Min Yang

"arXiv:2608.05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily…"

View on X

Originally posted by Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Jiang Feng, Min Yang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses