Multimodal LLMs Inconsistent for Disaster Assistance, Fail Vulnerable Groups.

Anuridhi Gupta, Samara Mansoor, Hemant Purohit· August 18, 2026 View original

Key takeaways

  • Multimodal LLMs currently lack consistent performance across text and audio inputs.
  • Vulnerable populations experience heightened performance gaps with these AI systems.
  • Modality-dependent inequity undermines the humanitarian potential of MM-LLMs.
  • Rigorous testing and design for accessibility are crucial for AI in public services.

Who benefits

Emergency ServicesPublic SectorHealthcareNon-profit Organizations

Summary

A study found that current multimodal large language models (MM-LLMs) produce inconsistent responses across text and audio modalities, particularly failing to meet the needs of vulnerable populations in disaster communication scenarios. This inconsistency introduces modality-dependent inequity, hindering their humanitarian value.

Researchers investigated the reliability of open-weight multimodal large language models (MM-LLMs) when used for disaster risk communication, specifically focusing on their consistency across text and audio inputs. The study simulated real emergency scenarios for four vulnerable personas, including hard-of-hearing individuals and the elderly. The findings revealed that no tested MM-LLM achieved consistent outputs regardless of the input modality. Performance gaps were significantly wider for users with access needs, indicating that these systems introduce an inequity based on how users communicate. This undermines the potential humanitarian benefits of MM-LLMs in critical situations.

Why it matters

Professionals developing or deploying AI for public services, especially in critical sectors like emergency response, must understand the limitations of current MM-LLMs regarding modality consistency and equitable access.

How to implement this in your domain

  1. 1Prioritize robust testing of AI systems across all intended input modalities and user demographics.
  2. 2Implement human-in-the-loop verification for critical AI-generated communications, especially in high-stakes scenarios.
  3. 3Design fallback mechanisms or alternative communication channels for users who may be underserved by current AI capabilities.
  4. 4Advocate for and invest in research focused on improving modality consistency and accessibility features in AI models.

Original post by Anuridhi Gupta, Samara Mansoor, Hemant Purohit

"arXiv:2608.14651v1 Announce Type: new Abstract: Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pre…"

View on X

Originally posted by Anuridhi Gupta, Samara Mansoor, Hemant Purohit on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses