C$^2$MOE Enhances Multimodal Emotion Recognition with Incomplete Data

Yuntao Shou, Tao Meng, Wei Ai, Keqin Li· August 6, 2026 View original

Key takeaways

  • C$^2$MOE is a new framework for robust multimodal emotion recognition with incomplete data.
  • It unifies representation learning and missing modality imputation.
  • The framework uses consistency and complementarity-guided experts.
  • C$^2$MOE significantly outperforms state-of-the-art methods in various missing-modality settings.

Who benefits

Customer ServiceHealthcareEdTechAutomotiveSocial Media Analytics

Summary

Researchers propose C$^2$MOE, a novel Consistency and Complementarity-guided Mixture of Experts framework for multimodal emotion recognition in conversations (MERC) that effectively handles missing modalities. It unifies representation learning and imputation, significantly outperforming state-of-the-art methods across various missing-modality settings.

Multimodal Emotion Recognition in Conversations (MERC) systems typically rely on complete multimodal inputs, but real-world data often suffer from missing modalities due to transmission errors or user behavior. This incompleteness severely degrades model performance, and existing methods, while enhancing robustness through cross-modal consistency learning, frequently overlook modality complementarity, leading to biased data reconstructions. To overcome these limitations, a new framework called C$^2$MOE (Consistency and Complementarity-guided Mixture of Experts) has been introduced. This novel approach unifies representation learning and missing modality imputation within a principled information-theoretic framework. It factorizes multimodal knowledge into distinct consistency and complementarity components using interaction-aware experts. Consistency is achieved by maximizing cross-modal predictability, while complementarity is preserved by maximizing conditional entropy between modalities. C$^2$MOE features a dual-branch prediction mechanism for robust imputation: a consistency branch aligns imputed features with the joint distribution by minimizing uncertainty, and a complementarity branch exploits unique modality cues via entropy maximization. Finally, a learnable reweighting module dynamically assigns importance scores to each expert's output, resulting in a robust and adaptive fusion for imputation. Extensive experiments on multiple MERC benchmarks confirm that C$^2$MOE consistently surpasses state-of-the-art methods across various missing-modality scenarios, validating its robustness and generalization capabilities.

Why it matters

This framework significantly improves the reliability of multimodal emotion recognition systems in real-world scenarios where data incompleteness is common, making AI-driven emotional intelligence more practical and robust.

How to implement this in your domain

  1. 1Evaluate existing multimodal AI systems for their robustness to missing data and consider integrating C$^2$MOE's principles.
  2. 2Implement dual-branch prediction mechanisms for handling incomplete multimodal inputs in your AI models.
  3. 3Develop strategies for maximizing cross-modal predictability and conditional entropy in multimodal data fusion.
  4. 4Explore the use of learnable reweighting modules to dynamically adapt to varying data quality and completeness.
  5. 5Apply this framework to improve emotion recognition in customer service, mental health monitoring, or human-computer interaction applications.

Original post by Yuntao Shou, Tao Meng, Wei Ai, Keqin Li

"arXiv:2608.04013v1 Announce Type: cross Abstract: Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs. However, real-world data often suffer from missing modalities due to transmission errors or user behav…"

View on X

Originally posted by Yuntao Shou, Tao Meng, Wei Ai, Keqin Li on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses