ReMem Boosts MLLM Long Video Understanding with Adaptive Memory

Linghao Meng, Qiankun Li, Junyuan Mao, Pujin Liao, Zhicheng He, Enbo Zhang, Kun Wang, Yang Liu, Huazhu Fu, Yueming Jin· July 29, 2026 View original

Summary

ReMem is a new training-free framework that enhances Multimodal Large Language Models (MLLMs) for long video understanding. It uses a temporal granularity-adaptive keyframe selection mechanism with dual-level memory augmentation, achieving state-of-the-art zero-shot performance.

A novel framework named ReMem has been introduced to significantly improve Multimodal Large Language Models (MLLMs) in understanding long videos without requiring additional training. MLLMs typically struggle with long video contexts due to limited context windows, often relying on keyframe selection methods that can overlook crucial temporal information. ReMem addresses this by implementing a temporal granularity-adaptive keyframe selection framework with dual-level memory augmentation. At the query level, a Memory-Driven Question Parsing component uses LLM long-term memory to decode question temporal granularity and extract semantic entities. At the video level, Synergistic Dual-Semantic Frame Alignment leverages intrinsic structural memory to align frames with query semantics, guiding Structure-Aware Dynamic Frame Routing to optimally distribute sampling budgets. By explicitly preserving temporal information through these memory mechanisms, ReMem effectively suppresses redundancy and empowers MLLMs to perform robust multi-granular video reasoning. Evaluations across four popular LongVideoQA benchmarks, using three MLLMs, demonstrated highly efficient, state-of-the-art zero-shot performance, with notable gains on LVBench and LongVideoBench.

Why it matters

This advancement is critical for professionals developing AI applications involving video analysis, surveillance, content moderation, or educational platforms, as it enables MLLMs to process and understand extended video content more effectively and accurately.

How to implement this in your domain

  1. 1Evaluate existing MLLM-based video understanding pipelines for limitations in long video processing.
  2. 2Explore integrating ReMem or similar memory-augmented frameworks to enhance MLLM capabilities for long video QA.
  3. 3Benchmark the performance improvements of ReMem on specific long video analysis tasks relevant to your domain.
  4. 4Train AI development teams on advanced keyframe selection and temporal reasoning techniques for MLLMs.
  5. 5Develop new applications leveraging improved long video understanding for content summarization, event detection, or educational tools.

Who benefits

Media & EntertainmentSecurity & SurveillanceEdTechAutomotiveHealthcare

Key takeaways

  • ReMem is a training-free framework for MLLM long video understanding.
  • It uses a temporal granularity-adaptive keyframe selection with dual-level memory.
  • The framework improves MLLMs' ability to reason across multi-granular video events.
  • ReMem achieved state-of-the-art zero-shot performance on LongVideoQA benchmarks.

Original post by Linghao Meng, Qiankun Li, Junyuan Mao, Pujin Liao, Zhicheng He, Enbo Zhang, Kun Wang, Yang Liu, Huazhu Fu, Yueming Jin

"arXiv:2607.24794v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding. To accommodate this constraint, models typically resort…"

View on X

Originally posted by Linghao Meng, Qiankun Li, Junyuan Mao, Pujin Liao, Zhicheng He, Enbo Zhang, Kun Wang, Yang Liu, Huazhu Fu, Yueming Jin on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses