FedGAMMA Enables Federated Multimodal Graph Foundation Models

Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang· July 20, 2026 View original

Summary

FedGAMMA is a novel two-stage framework for federated multimodal graph foundation learning that addresses challenges from privacy-restricted data silos. It achieves semantic-structural alignment through pre-training with shared-private semantic enhancement and topology-aware graph fusion, followed by prompt-based fine-tuning, consistently outperforming baselines on diverse multimodal graph datasets.

Multimodal-attributed graphs (MAGs), which integrate various data modalities like images and text with topological structures, are increasingly prevalent but often fragmented across privacy-restricted data silos. This fragmentation poses a significant challenge for learning broadly transferable models. To address this, FedGAMMA is proposed as a two-stage framework for federated multimodal graph foundation learning, focusing on semantic-structural alignment. The first stage, federated pre-training, involves a shared-private semantic enhancer that disentangles common cross-modal information from modality-specific details, aligning them via optimal transport. A topology-aware graph fusion module then decouples semantic and structural views using semantic residual graphs and dual positional encodings. Client similarity is estimated through a dual-channel affinity-aware aggregation mechanism without exposing raw data. The second stage, prompt-based fine-tuning, adapts the pre-trained encoder using lightweight graph-aware prompts, a shared prompt pool with controlled exploration, and channel-wise prompt synchronization. Experiments across twelve multimodal graph datasets show FedGAMMA consistently outperforming a wide range of baselines, with gains up to 12.96%, and also demonstrating superior performance in few-shot learning scenarios across multi-domain datasets.

Why it matters

Professionals dealing with sensitive, distributed multimodal data can leverage FedGAMMA to build powerful, privacy-preserving AI models that unlock insights from fragmented datasets without compromising data confidentiality.

How to implement this in your domain

  1. 1Identify use cases involving multimodal graph data distributed across privacy-sensitive silos.
  2. 2Explore federated learning paradigms as a solution for collaborative model training without raw data sharing.
  3. 3Investigate FedGAMMA's two-stage semantic-structural alignment approach for multimodal graph foundation models.
  4. 4Consider implementing a proof-of-concept using FedGAMMA for a specific task requiring federated multimodal graph analysis.
  5. 5Assess the privacy implications and compliance requirements for federated learning deployments in your domain.

Who benefits

HealthcareSocial MediaE-commerceFinanceCybersecurity

Key takeaways

  • Multimodal graph data is often fragmented across privacy-restricted silos.
  • FedGAMMA enables federated learning for multimodal graph foundation models.
  • It uses a two-stage semantic-structural alignment for robust learning.
  • The framework consistently outperforms baselines, even in few-shot scenarios.

Original post by Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang

"arXiv:2607.15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer sem…"

View on X

Originally posted by Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses