LLMs Struggle to Meet Real User Expectations

Miaomiao Li, Yang Wang, Bin Liang, Shudong Liu, Zhiwei Zhang, Kam-Fai Wong· July 24, 2026 View original

Summary

A new study reveals that despite strong benchmark performance, large language models often fail to align with diverse and subtle real-world user expectations. Researchers introduce ExpectBench, a benchmark grounded in actual user expectations, and LENS, a framework that helps models internalize these expectations for better-aligned responses.

While large language models (LLMs) consistently achieve impressive results on standard benchmarks, a significant gap exists between their perceived competence and their ability to truly satisfy real-world user expectations. Current evaluation methods, which often rely on model heuristics or expert rubrics, do not adequately capture the nuanced and varied expectations of human users, leading to a misalignment between model output and user intent.To address this, a systematic study of user expectations in actual LLM interactions was conducted. This research proposes a method for extracting semantically rich expectations and introduces ExpectBench, a novel benchmark built directly from real user data. Analysis of this benchmark highlights that contemporary LLMs frequently struggle to anticipate and fulfill what users genuinely hope to obtain from their interactions.Based on these findings, a lightweight framework called LENS (Latent Expectation-aware Response Generation) is proposed. LENS enables LLMs to internalize user expectations, leading to the generation of responses that are more closely aligned with human desires. Experiments show that LENS consistently improves expectation satisfaction, underscoring the critical importance of explicitly modeling user expectations for achieving genuine human-AI alignment.

Why it matters

This research is vital for professionals aiming to build truly user-centric AI products. It highlights that benchmark performance doesn't equate to user satisfaction and provides a path to developing LLMs that better understand and meet diverse human needs.

How to implement this in your domain

  1. 1Conduct user research to systematically identify and document real-world expectations for your LLM-powered applications.
  2. 2Integrate expectation-aware evaluation metrics into your LLM development and testing pipelines.
  3. 3Explore fine-tuning LLMs using datasets augmented with explicit user expectation signals, similar to the LENS framework.
  4. 4Design user feedback loops specifically to capture and analyze expectation misalignment for continuous model improvement.

Who benefits

Customer ServiceProduct ManagementUX DesignDigital MarketingAI Product Development

Key takeaways

  • LLMs often fail to meet diverse real-world user expectations despite strong benchmarks.
  • Existing evaluation methods don't fully capture subtle human expectations.
  • ExpectBench is a new benchmark grounded in actual user expectations.
  • LENS is a framework that improves LLM alignment by internalizing user expectations.

Original post by Miaomiao Li, Yang Wang, Bin Liang, Shudong Liu, Zhiwei Zhang, Kam-Fai Wong

"arXiv:2607.20485v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing evaluation approaches, relying on model heuristics,…"

View on X

Originally posted by Miaomiao Li, Yang Wang, Bin Liang, Shudong Liu, Zhiwei Zhang, Kam-Fai Wong on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses