AgentPatch Repairs Merged Agentic MLLMs for Weak Task Recovery

Zibo Shao, Baochen Xiong, Chengdong Xu, Linhui Xiao, Kaichen Li, Haoran Gong, Yan Li, Yaguang Song, Xiaoshan Yang· August 10, 2026 View original

Key takeaways

  • Merging specialized agentic MLLMs is challenging due to capability preservation issues.
  • AgentPatch is a training-free framework to repair merged MLLMs.
  • It addresses weak-task degradation and behavior-critical forgetting.
  • The framework improves merged backbones and balances capability recovery.

Who benefits

AI/TechRoboticsSoftware DevelopmentAutonomous SystemsGaming

Summary

AgentPatch is a training-free, coarse-to-fine repair framework designed to address challenges in merging agentic multimodal large language models (MLLMs), specifically asymmetric capability preservation and behavior-critical forgetting. It restores diluted weak-task signals and recovers decisive behaviors without requiring additional training.

Merging specialized agentic multimodal large language models (MLLMs) into a single generalist model presents significant challenges. Key issues include the uneven preservation of capabilities with varying interaction complexity, leading to "weak tasks," and the forgetting of critical behaviors essential for long-horizon execution. This paper introduces AgentPatch, a novel, training-free framework designed to repair these merged MLLMs. AgentPatch employs a coarse-to-fine approach. It first selects a stable merged backbone, then restores diluted weak-task-specific signals through a process called Weak-Task Unique Residual Recovery. Finally, it applies an Agent-Guided Behavior-Critical Patch to recover decisive actions while explicitly protecting existing capabilities. This framework produces a single, static checkpoint without needing complex routing or ensembles. Experiments across multiple agentic and multimodal benchmarks demonstrate that AgentPatch effectively improves diverse merged backbones, mitigates weak-task degradation, and balances recovery with the preservation of complementary capabilities.

Why it matters

For professionals building complex AI agents, AgentPatch offers a practical solution to combine specialized MLLMs more effectively, leading to more versatile and robust generalist agents without extensive retraining.

How to implement this in your domain

  1. 1Evaluate AgentPatch for merging specialized MLLMs within your organization's AI development pipeline.
  2. 2Apply the Weak-Task Unique Residual Recovery technique to address performance degradation in merged models.
  3. 3Implement Agent-Guided Behavior-Critical Patching to safeguard crucial agentic behaviors.
  4. 4Explore the framework's applicability for consolidating various AI models into a unified system.

Original post by Zibo Shao, Baochen Xiong, Chengdong Xu, Linhui Xiao, Kaichen Li, Haoran Gong, Yan Li, Yaguang Song, Xiaoshan Yang

"arXiv:2608.06699v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are specialized for particular tools or environments, c…"

View on X

Originally posted by Zibo Shao, Baochen Xiong, Chengdong Xu, Linhui Xiao, Kaichen Li, Haoran Gong, Yan Li, Yaguang Song, Xiaoshan Yang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses