New AI Framework Detects, Explains Harmful Humor in Memes

Shanhong Liu, Pai Chet Ng, De Wen Soh, Malika Meghjani, Konstantinos N. Plataniotis· July 20, 2026 View original

Summary

Researchers introduce MAR-12, a novel framework leveraging Vision Language Models to detect and explain harmful humor in internet memes by interpreting them through twelve structured perspectives. The system achieves high accuracy in humor and hate detection and provides coherent, context-grounded explanations.

Internet memes pose a significant challenge for AI due to their complex blend of visual, textual, and cultural elements, often containing humor, sarcasm, and potentially harmful intent. Existing multimodal AI systems struggle with providing reliable explanations for their interpretations. This new research introduces MAR-12, a framework designed to address these challenges. MAR-12 utilizes Vision Language Models (VLMs) to analyze memes from twelve distinct theoretical perspectives on humor and hate. It then employs a soft-gated attention mechanism to weigh the contribution of each perspective before classifying the meme. The system also generates transparent explanations based on these perspective-specific insights and learned attention weights. Evaluations on datasets like PrideMM and Memotion show MAR-12 outperforming current state-of-the-art methods in both humor and hate detection, achieving up to 80.3% and 75.9% accuracy respectively. Human and GPT-4 assessments confirm the quality of its explanations, especially for memes where humor and harm coexist.

Why it matters

Professionals in content moderation, brand safety, and social media management can leverage such AI to more effectively identify and address harmful content, improving platform integrity and user experience.

How to implement this in your domain

  1. 1Integrate advanced VLM-based detection systems into content moderation pipelines to identify nuanced harmful content.
  2. 2Develop internal guidelines for AI-assisted content review, incorporating multi-perspective analysis for complex cases like memes.
  3. 3Train human moderators on AI-generated explanations to enhance their understanding and decision-making processes.
  4. 4Pilot AI tools for proactive identification of potentially harmful trends in user-generated content before widespread dissemination.

Who benefits

Social MediaAdvertisingContent ModerationPublic Relations

Key takeaways

  • MAR-12 is a new VLM-based framework for detecting and explaining harmful humor in memes.
  • It analyzes memes from twelve structured perspectives derived from humor and hate theories.
  • The system outperforms state-of-the-art methods in both humor and hate detection.
  • It provides coherent and persuasive explanations, particularly for ambiguous memes.

Original post by Shanhong Liu, Pai Chet Ng, De Wen Soh, Malika Meghjani, Konstantinos N. Plataniotis

"arXiv:2607.15442v1 Announce Type: new Abstract: Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist. These complexities highlight the need for…"

View on X

Originally posted by Shanhong Liu, Pai Chet Ng, De Wen Soh, Malika Meghjani, Konstantinos N. Plataniotis on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses