Stealthy Backdoor Attacks Target Graph Foundation Models.

Minhua Lin, Zhicheng Gao, Yilong Wang, Hanqing Lu, Xiang Zhang, Suhang Wang· August 24, 2026 View original

Key takeaways

  • Graph Foundation Models (GFMs) are vulnerable to stealthy backdoor attacks like STAG.
  • STAG coordinates graph and text triggers to compromise GFMs on text-attributed graphs.
  • The attack is designed to be stealthy, making triggers readable and local graph structures natural.
  • Existing backdoor defenses are often ineffective against these aligned, multi-modal attacks.

Who benefits

CybersecuritySocial MediaFinancial ServicesGovernmentAI Development

Summary

Researchers propose STAG, a stealthy trojan attack framework designed to compromise Graph Foundation Models (GFMs) on text-attributed graphs (TAGs) by coordinating graph and text triggers. STAG ensures triggers are readable and local graph structures remain close to original, making detection difficult while effectively shifting model predictions towards a target class.

Graph Foundation Models (GFMs) operating on text-attributed graphs (TAGs) are designed to align graph representations with language semantics, enabling transferable graph learning. However, their vulnerability to backdoor attacks, especially under graph-language alignment, has been underexplored. Existing backdoor attacks typically target modalities independently, which is ineffective against GFMs where graph-only triggers can be constrained by clean text, and text-only triggers don't directly shift the aligned graph representation. Furthermore, TAGs pose a stealth challenge: triggers are exposed in both node text and local graph structure, making incoherent attributes or anomalous subgraphs easily detectable. To overcome these limitations, researchers introduce STAG (Stealthy Trojan Attack framework for GFMs). STAG coordinates a graph-trigger generator with a text-side soft prompt, ensuring that trigger-attached graph representations and triggered text representations converge towards the same target-class text region. STAG addresses stealthiness by realizing trigger nodes as readable text through candidate retrieval and regularizing the trigger-attached subgraph to maintain structural similarity to the original. Extensive experiments across multiple TAG datasets and GFMs confirm STAG's effectiveness and stealth, demonstrating its ability to subtly compromise these advanced models.

Why it matters

This research highlights a critical security vulnerability in emerging Graph Foundation Models, urging professionals to develop robust defense mechanisms against sophisticated, stealthy backdoor attacks that could compromise the integrity and trustworthiness of AI systems.

How to implement this in your domain

  1. 1Assess the security posture of Graph Foundation Models currently in use or under development within your organization.
  2. 2Develop and implement detection mechanisms for stealthy backdoor attacks like STAG, focusing on coordinated graph and text anomalies.
  3. 3Integrate adversarial training or robust fine-tuning techniques to enhance GFM resilience against such attacks.
  4. 4Conduct regular security audits and penetration testing on GFM deployments.
  5. 5Educate AI development teams on the risks of sophisticated backdoor attacks and best practices for model hardening.

Original post by Minhua Lin, Zhicheng Gao, Yilong Wang, Hanqing Lu, Xiang Zhang, Suhang Wang

"arXiv:2608.20991v1 Announce Type: new Abstract: Graph Foundation Models (GFMs) on text-attributed graphs (TAGs) align graph representations with language semantics to support transferable graph learning. Despite these advantages, the backdoor vulnerability of GFMs on TAGs remains…"

View on X

Originally posted by Minhua Lin, Zhicheng Gao, Yilong Wang, Hanqing Lu, Xiang Zhang, Suhang Wang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools