RobustMAD Benchmark Evaluates Small Multimodal LLM Robustness

Anushiya Arunan, Xin Li, Yan Qin, U-Xuan Tan, Nhu Khue Vuong, Xiaoli Li, Chau Yuen· July 21, 2026 View original

Summary

This paper introduces RobustMAD, a new benchmark for evaluating the real-world robustness of multimodal small language models (MSLMs) in industrial anomaly detection. It reveals critical robustness gaps in MSLMs, despite their promising capabilities, highlighting issues like fragile grounding and insufficient responses under challenging conditions.

Multimodal industrial anomaly inspection assistants are vital for smart factories, but large language models are often too computationally demanding or pose privacy risks for on-site deployment. Compact multimodal small language models (MSLMs) offer a deployable alternative, yet their real-world robustness has been poorly understood due to a lack of comprehensive benchmarks. To address this, researchers developed RobustMAD, the first benchmark specifically designed to evaluate MSLM robustness under diverse, open-ended industrial queries. This benchmark covers object understanding, anomaly detection, unanswerable problems, and visual quality degradations, reflecting realistic industrial conditions. Surprisingly, top-performing MSLMs showed promising capabilities, even outperforming larger models like GPT-5 Nano in some aspects. However, RobustMAD also exposed significant robustness gaps, identifying three recurring failure modes: fragile multimodal grounding under fine-grained distinctions or degraded visuals, insufficiently comprehensive responses, and weak logical grounding leading to hallucinations on ill-posed queries. These insights provide actionable guidance for designing more robust next-generation industrial inspection assistants.

Why it matters

For professionals in manufacturing, quality control, and AI deployment, RobustMAD provides a crucial tool and insights for selecting and developing robust, deployable multimodal AI assistants that can reliably perform anomaly detection in real-world industrial settings.

How to implement this in your domain

  1. 1Utilize benchmarks like RobustMAD to rigorously evaluate the robustness of MSLMs before deployment in critical industrial applications.
  2. 2Prioritize MSLM development efforts on improving multimodal grounding, especially under degraded visual conditions.
  3. 3Focus on enhancing the logical reasoning capabilities of MSLMs to prevent hallucinations on ambiguous queries.
  4. 4Design industrial AI assistants with mechanisms to handle unanswerable questions gracefully, rather than generating false information.

Who benefits

ManufacturingIndustrial AutomationQuality ControlLogisticsAutomotive

Key takeaways

  • MSLMs are promising for on-site industrial anomaly detection but lack real-world robustness.
  • RobustMAD benchmark reveals critical failure modes: fragile grounding, incomplete responses, and hallucinations.
  • Even top MSLMs fall short of safety-critical requirements despite outperforming some larger models.
  • Actionable insights are provided for designing more robust multimodal industrial assistants.

Original post by Anushiya Arunan, Xin Li, Yan Qin, U-Xuan Tan, Nhu Khue Vuong, Xiaoli Li, Chau Yuen

"arXiv:2607.16243v1 Announce Type: new Abstract: Multimodal industrial anomaly inspection assistants are a critical component of next-generation smart factories, enabling interactive vision-language-based querying. However, multimodal large language models remain impractical for o…"

View on X

Originally posted by Anushiya Arunan, Xin Li, Yan Qin, U-Xuan Tan, Nhu Khue Vuong, Xiaoli Li, Chau Yuen on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses