New Benchmark Evaluates VLM Memory Beyond Simple Accuracy

Shmuel Berman, Jia Deng· September 2, 2026 View original

Key takeaways

  • Traditional VLM memory benchmarks are insufficient, focusing only on accuracy.
  • ECCBench introduces efficiency, compression, and calibration as key memory metrics.
  • VLMs struggle with video compression and show poor calibration across modalities.
  • Non-Transformer architectures may offer better memory trade-offs for long-horizon tasks.

Who benefits

AI/TechRoboticsAutonomous SystemsMedia & EntertainmentSurveillance

Summary

Researchers introduce ECCBench, a new benchmark and evaluation protocol that assesses the memory capabilities of Vision-Language Models (VLMs) across three dimensions: efficiency, compression, and calibration, moving beyond traditional accuracy-only metrics. The study reveals that current VLMs struggle with calibration and compression for video, and non-Transformer architectures show better trade-offs.

The memory capabilities of Large Language Models (LLMs) and Vision-Language Models (VLMs) are a recognized challenge, typically evaluated by their accuracy on long-context tasks. However, this narrow focus on accuracy overlooks crucial aspects vital for real-world, long-horizon applications. A new benchmark, ECCBench, has been developed to provide a more comprehensive evaluation of VLM memory. ECCBench measures memory along three axes: Efficiency (computational cost to retrieve information), Compression (ability to remember compressible inputs more effectively), and Calibration (the system's ability to express uncertainty and manage error costs). This protocol aims to reveal deeper insights into how VLMs handle and utilize information over extended periods. Initial findings using ECCBench indicate that while pre-trained VLMs can compress memory for text, they do not exhibit similar compression for video data. Furthermore, these models generally show poor calibration across both modalities. Interestingly, some non-Transformer architectures demonstrated superior compression-calibration trade-offs compared to RoPE Transformers, suggesting alternative architectural components might be more suitable for agents operating with long-term memory requirements.

Why it matters

This benchmark provides a more nuanced understanding of VLM memory, which is critical for developing robust AI agents capable of long-term reasoning and interaction, especially in complex, real-world scenarios where efficiency, data compression, and reliable uncertainty estimation are paramount.

How to implement this in your domain

  1. 1Adopt ECCBench or similar multi-faceted evaluation protocols for VLM development and selection.
  2. 2Prioritize VLM architectures that demonstrate better compression and calibration, not just raw accuracy.
  3. 3Investigate non-Transformer architectures for memory components in long-horizon AI applications.
  4. 4Develop strategies to improve VLM calibration, particularly for video-based tasks, to enhance reliability.

Original post by Shmuel Berman, Jia Deng

"arXiv:2609.00103v1 Announce Type: new Abstract: Memory is widely viewed as an important unsolved problem for LLMs and VLMs, and current benchmarks typically evaluate it by testing accuracy over long text or video. However, accuracy alone misses properties that matter for real lon…"

View on X

Originally posted by Shmuel Berman, Jia Deng on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses