Setoka Benchmark Evaluates Deeper User Understanding for Personalized Agents

Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou· July 31, 2026 View original

Key takeaways

  • Existing benchmarks for personalized agents lack deep user understanding evaluation.
  • Setoka introduces a hierarchical benchmark for semantic, episodic, behavioral, and personality understanding.
  • Current LLMs struggle with episodic memory and inferring behavior/personality from fragmented data.
  • More advanced memory mechanisms are needed for truly personalized agents.

Who benefits

Customer ServiceEdTechHealthcareE-commerceMarketing

Summary

Setoka is a new benchmark designed to assess personalized agents' hierarchical user understanding beyond explicit fact retrieval, covering semantic memory, episodic memory, behavior patterns, and personality traits. It uses a psychometrics-based pipeline to synthesize diverse, coherent user data and queries.

Personalized AI agents need more than just recalling explicit facts from past interactions; they require a deeper understanding of users, including abstract characteristics. Existing benchmarks primarily test simple information retrieval, failing to adequately assess this nuanced user comprehension. This paper introduces Setoka, a new benchmark specifically designed to evaluate memory-augmented personalized agents across multiple levels of user understanding. Setoka is grounded in cognitive and personality psychology, defining four hierarchical levels: semantic memory, episodic memory, behavior patterns, and personality traits. To enable realistic yet privacy-preserving evaluation, the benchmark employs a psychometrics-based pipeline to synthesize diverse and coherent heterogeneous user data and queries at scale. Initial evaluations using Setoka reveal that current language models combined with various memory systems perform well on semantic memory but struggle significantly with episodic memory, and even more so with tasks requiring integration of fragmented, long-term information to infer behavior patterns and personality traits. This highlights the need for more sophisticated memory mechanisms in personalized agents.

Why it matters

For product developers and researchers building personalized AI experiences, Setoka provides a crucial tool to rigorously evaluate and improve agents' ability to truly understand users, leading to more effective and empathetic interactions.

How to implement this in your domain

  1. 1Adopt benchmarks like Setoka to evaluate the depth of user understanding in your personalized agent development.
  2. 2Prioritize research and development into memory mechanisms that can integrate heterogeneous and fragmented user data over time.
  3. 3Design agent architectures that can infer abstract user characteristics beyond explicit fact retrieval.
  4. 4Conduct user studies to validate whether deeper user understanding translates to improved user satisfaction and task completion.
  5. 5Explore psychometrics-based data synthesis methods for creating diverse and privacy-preserving training/evaluation datasets.

Original post by Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou

"arXiv:2607.27056v1 Announce Type: new Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferri…"

View on X

Originally posted by Lingyang Zeng, Guangze Chen, Kaichen Yu, Zhicheng Pan, Siyang Weng, Zirui Hu, Xiangyun Du, Hailin He, Rong Zhang, Chengcheng Yang, Kai Huang, Xuan Zhou on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses