FUSE Framework Evaluates LLM Dangerous Capabilities and Safety.

Zhengyi Jin, Ru Zhang, Xiao Chen, Xinbo Liu, Jiaxuan Lin, Jia Huang, Jianyi Liu, Zhen Yang· September 3, 2026 View original

Key takeaways

  • FUSE provides a unified framework for evaluating LLM dangerous capabilities across three dimensions.
  • LLMs exhibit diverse safety profiles, with knowledge, defense, and harm often diverging.
  • Newer models do not always show monotonic decline in dangerous capabilities.
  • Systematic evaluation is crucial for responsible AI development and deployment.

Who benefits

Software DevelopmentCybersecurityGovernmentLegalHealthcare

Summary

This paper introduces FUSE, a modular framework for evaluating dangerous AI capabilities of large language models (LLMs) through three orthogonal pipelines: Knowledge (K), Defense (D), and Harm (H). It provides a standardized dangerous-capability profile, revealing divergent safety behaviors across models and showing that dangerous capabilities have not monotonically declined with newer models.

The fragmented nature of current safety evaluations for large language models (LLMs) hinders effective governance of their potentially dangerous capabilities. To address this, the FUSE framework offers a unified, modular approach to assess LLMs. It evaluates each model across three distinct, yet orthogonal, pipelines: Knowledge (K), which measures what the model knows about harmful topics; Defense (D), which assesses its refusal resilience to harmful requests; and Harm (H), which quantifies the harmfulness of its generated content when it does comply. The results are aggregated into a standardized dangerous-capability profile.The framework's modular design allows for pluggable components like scenario seeds and judge rubrics, while the core evaluation engine remains consistent. Instantiated with a chemical-biological (CB) module, FUSE was used to evaluate 12 commercial LLMs. This revealed significant differences in their safety profiles; models with similar knowledge levels could vary widely in their refusal rates, and strong defenders didn't necessarily produce less harmful content when they did comply. Family-level patterns also emerged, distinguishing models like Claude, DeepSeek, and GPT.Furthermore, a temporal analysis tracking K, D, and H against model release dates showed that dangerous capabilities have not consistently decreased over time. Newer models often exhibit deeper knowledge while only partially improving their defense mechanisms, indicating that advancements in scaling and alignment do not uniformly translate into enhanced safety across all dimensions. The framework's reliability was confirmed through high cross-judge consistency and pipeline orthogonality.

Why it matters

For professionals involved in AI development, deployment, and governance, FUSE provides a critical tool for systematically evaluating and understanding the safety risks of LLMs. It enables more informed decisions about model selection, risk mitigation, and responsible AI development.

How to implement this in your domain

  1. 1Adopt a structured framework like FUSE for comprehensive safety evaluations of LLMs before deployment.
  2. 2Evaluate LLMs across multiple dimensions: knowledge of harmful topics, refusal resilience, and actual harm generation.
  3. 3Develop internal protocols for assessing and profiling the dangerous capabilities of AI models.
  4. 4Monitor the evolution of dangerous capabilities in new model releases, recognizing that safety improvements may not be uniform.
  5. 5Use the generated dangerous-capability profiles to inform risk assessments and implement appropriate safeguards.

Original post by Zhengyi Jin, Ru Zhang, Xiao Chen, Xinbo Liu, Jiaxuan Lin, Jia Huang, Jianyi Liu, Zhen Yang

"arXiv:2609.02168v1 Announce Type: new Abstract: Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge ($K$), Defense ($D$), and Harm ($H$)---unde…"

View on X

Originally posted by Zhengyi Jin, Ru Zhang, Xiao Chen, Xinbo Liu, Jiaxuan Lin, Jia Huang, Jianyi Liu, Zhen Yang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses