Ex-Omni-2D Creates Expressive Omni-Modal Dialogue with Visual Avatars

Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo· August 12, 2026 View original

Key takeaways

  • Ex-Omni-2D generates coordinated text, personalized speech, and video for AI avatars.
  • It uses a Visual Thought Plan and a shared acoustic-temporal interface for learning.
  • The framework enables efficient incremental generation with a "Streaming Student" model.
  • This technology makes human-AI interactions more natural and visually engaging.

Who benefits

Customer ServiceEdTechMarketingEntertainmentVirtual Assistants

Summary

Ex-Omni-2D is an omni-modal dialogue framework that generates coordinated responses including text, personalized speech, and reference-conditioned video, giving AI avatars a native visual presence. It uses a Visual Thought Plan and a shared acoustic-temporal interface to learn from heterogeneous data and enables efficient incremental generation.

While omni-modal dialogue models can process multimodal inputs and generate spoken replies, their responses typically lack a visual embodiment. A new framework, Ex-Omni-2D, aims to bridge this gap by generating a fully coordinated response that includes text, personalized speech, and a reference-conditioned video, effectively giving AI avatars a native visual presence. Given a multimodal query, along with a reference image and audio, Ex-Omni-2D first predicts a structured "Visual Thought Plan" (VTP) that outlines scene, emotion, and motion. This is followed by the generation of response text and native multi-codebook speech units. These speech units form a crucial shared acoustic-temporal interface, allowing them to be decoded into speech and simultaneously aligned online with video frames. This innovative interface enables the model to learn both response and avatar pathways from diverse speech, dialogue, and avatar-video datasets, eliminating the need for extensive query-text-speech-video supervision. For efficient incremental generation, the framework distills a full-sequence Video Generator into a few-step block-causal "Streaming Student." This student model uses a Prefix Streaming mechanism to maintain a clean latent representation across consecutive chunks, minimizing cumulative degradation in later parts of the video. The complete four-GPU pipeline achieves an end-to-end real-time factor of 1.293, demonstrating a practical balance of quality and efficiency for generating expressive, visually embodied AI dialogue.

Why it matters

This research significantly advances human-AI interaction by enabling AI systems to communicate with expressive visual presence, making interactions more natural and engaging for users.

How to implement this in your domain

  1. 1Explore integrating expressive visual avatar generation into customer service or virtual assistant platforms.
  2. 2Pilot the use of omni-modal dialogue models for enhanced interactive educational content.
  3. 3Investigate the application of reference-conditioned video generation for personalized marketing campaigns.
  4. 4Develop internal guidelines for designing AI avatars with appropriate emotional and motion cues.
  5. 5Assess the computational resources required for deploying such real-time, visually rich AI interactions.

Original post by Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo

"arXiv:2608.10720v1 Announce Type: new Abstract: Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce \textbf{Ex-Omni-2D}, an omni-modal dialogue framework that generates a coordina…"

View on X

Originally posted by Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu, Yiwen Guo on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses