SERUM Extracts User Behavior Models from Unstructured Screen Video

Andy J. Phu, James Mooney, Karin de Langis, Khanh Chi Le, Dongyeop Kang· August 3, 2026 View original

Key takeaways

  • SERUM extracts structured user behavior models from unstructured screen video.
  • It uses multi-pass VLM annotation to refine labels and reduce errors.
  • The framework converges to stable behavioral vocabularies, called schematic equilibrium.
  • This enables scalable, interpretable user modeling for proactive AI assistants.

Who benefits

Software DevelopmentUX/UI DesignMarketingEdTechRobotics

Summary

Researchers developed SERUM, a multi-pass framework that extracts finite-state behavioral models from unstructured egocentric screen video using hierarchical VLM annotation. SERUM refines labels iteratively, converging to stable state vocabularies and enabling interpretable process models for user understanding.

This paper introduces SERUM (State Extraction and Refinement for User Modeling), a novel multi-pass framework designed to create structured models of user intent and workflow directly from raw, unstructured screen activity. The challenge lies in converting egocentric video recordings into interpretable behavioral models. SERUM addresses this by employing hierarchical Vision-Language Model (VLM) annotation. The framework processes screen recordings through a sliding window, alternating between activity-recognition and intent-inference passes. Each pass refines labels using accumulated prior context, which helps to reduce common issues like hallucination and temporal conflation often seen in single-pass annotation methods. Synonymous states are then merged using sentence embeddings and human-calibrated thresholds, resulting in a compact and coherent taxonomy of user behaviors. Evaluations across 61 egocentric videos in diverse domains (coding, cooking, physical activities, daily life) show that iterative label refinement converges to a stable "schematic equilibrium." Markov models built on these labels achieve significantly lower perplexity and higher action predictions compared to frequency baselines, especially in structured tasks like coding. SERUM is presented as the first system to automatically generate interpretable process models from unstructured egocentric screen video, offering a scalable pathway for user modeling and behavioral understanding in real-world scenarios.

Why it matters

Professionals in product design, UX research, and AI assistant development can leverage SERUM to gain deeper, scalable insights into user behavior and intent, enabling the creation of more proactive and personalized digital experiences.

How to implement this in your domain

  1. 1Explore using egocentric video data to understand complex user workflows in your product.
  2. 2Investigate integrating VLM annotation techniques for automated behavioral analysis.
  3. 3Pilot SERUM's multi-pass refinement framework to build structured user models from screen recordings.
  4. 4Apply the generated behavioral models to inform product feature development and AI assistant design.

Original post by Andy J. Phu, James Mooney, Karin de Langis, Khanh Chi Le, Dongyeop Kang

"arXiv:2607.29181v1 Announce Type: new Abstract: Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow. However, building these models from raw, unstructured screen activity remains an open challenge. We present SE…"

View on X

Originally posted by Andy J. Phu, James Mooney, Karin de Langis, Khanh Chi Le, Dongyeop Kang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses