CTIFoundry Boosts Cyber Threat Intelligence Agents with Structured Corpus.

Yutong Cheng, Changze Li, Qian Cui, Wei Ding, Lingzhi Wang, Yan Chen, Peng Gao· August 20, 2026 View original

Key takeaways

  • Unstructured CTI corpora bottleneck LLM agent performance in investigations.
  • CTIFoundry provides an agent-native, structured corpus scaffold for CTI.
  • Typed ontology graphs and procedural skills significantly boost agent accuracy and efficiency.
  • Structured data access enables smaller LLMs to outperform larger ones on flat data.

Who benefits

CybersecurityGovernmentDefenseFinancial ServicesIT Services

Summary

CTIFoundry is an agent-native corpus scaffold that materializes the latent structure of cyber threat intelligence (CTI) knowledge bases into a deterministic ontology graph. This structured approach, exposed via typed tools and procedural skills, significantly improves LLM agent performance in CTI investigations compared to traditional flat retrieval-augmented generation.

The way cyber threat intelligence (CTI) is consumed is shifting from human analysts to LLM agents that perform multi-step investigations. While agent harnesses have advanced, the underlying CTI corpora are still largely formatted for simple retrieval-augmented generation (RAG), treating threat reports as opaque chunks. This unstructured substrate is identified as a major bottleneck for effective agentic CTI investigation. To overcome this, CTIFoundry, an agent-native corpus scaffold, has been developed. At build time, CTIFoundry extracts and materializes the inherent structure of CTI knowledge bases, creating a deterministic ontology graph. This graph links authoritative sources like CVE, CWE, CAPEC, and ATT&CK through typed, traversable edges, and integrates a span-grounded report layer with alias-resolved entities. It also provides hybrid dense and lexical retrieval surfaces. At query time, this rich structure is exposed to LLM agents through seven typed tools and three procedural skills. Benchmarking on CTIConnect showed that simply swapping to CTIFoundry's action surface improved agent F1 scores by +0.19 to +0.28 across various models. A smaller model using CTIFoundry even outperformed a flagship model on the flat substrate, and the scaffolded agent was more accurate with roughly half the tool calls. This demonstrates that typed structure and procedural skills are key to enhancing agent performance in CTI.

Why it matters

Cybersecurity professionals and AI developers can leverage CTIFoundry's structured approach to build more effective and efficient LLM agents for cyber threat intelligence, leading to faster and more accurate investigations.

How to implement this in your domain

  1. 1Evaluate current CTI consumption methods and identify bottlenecks in LLM agent investigations.
  2. 2Explore integrating CTIFoundry or similar structured corpus scaffolds for CTI knowledge bases.
  3. 3Develop or adapt LLM agents to utilize typed tools and procedural skills that interact with the structured CTI corpus.
  4. 4Benchmark agent performance on CTI tasks using both flat RAG and structured corpus approaches.
  5. 5Train security analysts on how to interact with and interpret the outputs of these enhanced CTI agents.

Original post by Yutong Cheng, Changze Li, Qian Cui, Wei Ding, Lingzhi Wang, Yan Chen, Peng Gao

"arXiv:2608.18613v1 Announce Type: new Abstract: Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that compose multi-step investigations at query time. The harness side of this shift has matured rapidly (planning loops, tool protocol…"

View on X

Originally posted by Yutong Cheng, Changze Li, Qian Cui, Wei Ding, Lingzhi Wang, Yan Chen, Peng Gao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses