SciHazard Benchmarks LLM Scientific Safety Risks.

Chunxiao Li, Yuan Xiong, Lijun Li, Tianyi Du, Wenlong Zhang, Lei Bai, Jing Shao· July 22, 2026 View original

Summary

SciHazard is a new benchmark and evaluation framework designed to measure the scientific safety risks of LLMs, particularly their ability to generate actionable misuse guidance. It uses real-world hazardous questions and a decomposed harm scoring method, revealing deep research agents as a critical blind spot.

Large language models (LLMs) are increasingly used to support scientific endeavors, but they also pose a risk by potentially converting hazardous scientific knowledge into actionable instructions for misuse. Existing safety benchmarks often rely on templated queries that don't reflect real-world dangers and use LLM-as-a-Judge paradigms without sufficient domain expertise. To address these shortcomings, researchers introduce SciHazard, a new benchmark grounded in real-world scientific risks. SciHazard comprises 2400 hazardous questions and 600 "oversafety" questions across 12 scientific disciplines. These queries are derived from regulated entities and documented failure scenarios, ensuring their relevance to actual hazards. The benchmark also introduces DeHarm-Score, a decomposed evaluation procedure that combines query hazard severity, the model's refusal behavior, and the risk level of its response. For responses that are not refused, DeHarm-Score further breaks down response-level harm into "Executability," quantified via dynamic checklists, and "Net-new risk," assessed through retrieval-augmented claim extraction and synthesis-barrier verification. Expert validation showed DeHarm-Score significantly improved agreement with human annotations compared to existing baselines. Benchmarking 31 frontier LLMs and deep research agents revealed that these autonomous agents yielded substantially higher mean DeHarm-Scores than standard LLMs, highlighting them as a critical, unaddressed safety concern.

Why it matters

Professionals developing or deploying LLMs, especially in scientific or sensitive domains, need robust tools like SciHazard to rigorously assess and mitigate the potential for misuse and ensure responsible AI development.

How to implement this in your domain

  1. 1Integrate the SciHazard benchmark into the safety evaluation pipeline for all LLMs used in scientific or critical applications.
  2. 2Adopt the DeHarm-Score methodology to gain a more nuanced understanding of potential harms from LLM outputs.
  3. 3Prioritize safety research and development specifically for "deep research agents" given their identified higher risk.
  4. 4Establish internal guidelines for responsible deployment of LLMs, particularly concerning their ability to generate actionable scientific guidance.
  5. 5Collaborate with domain experts to continuously update and expand hazard scenarios within safety benchmarks.

Who benefits

PharmaceuticalsBiotechnologyChemical ManufacturingDefenseCybersecurity

Key takeaways

  • SciHazard is a new, real-world-grounded benchmark for LLM scientific safety.
  • DeHarm-Score provides a decomposed, expert-validated method for measuring harm.
  • Deep research agents pose a higher scientific safety risk than standard LLMs.
  • Existing safety benchmarks may not adequately capture real-world hazards.

Original post by Chunxiao Li, Yuan Xiong, Lijun Li, Tianyi Du, Wenlong Zhang, Lei Bai, Jing Shao

"arXiv:2607.18665v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world ha…"

View on X

Originally posted by Chunxiao Li, Yuan Xiong, Lijun Li, Tianyi Du, Wenlong Zhang, Lei Bai, Jing Shao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses