IFCMemoryBench Evaluates LLM Agent Long-Term Memory in BIM.

Changyu Du, Alexander Vosseler, Filippo Mazza, Andr\'e Borrmann· July 31, 2026 View original

Key takeaways

  • Existing LLM agent memory evaluations are insufficient for domain-specific professional tasks.
  • IFCMemoryBench evaluates long-term memory in BIM, requiring context reuse across sessions.
  • Current memory systems struggle with domain-transfer, achieving low accuracy in BIM tasks.
  • Reliable professional agents need domain-aware memory representations.

Who benefits

ArchitectureEngineeringConstruction (AEC)Software DevelopmentAI Research

Summary

IFCMemoryBench is a new benchmark designed to assess the long-term memory capabilities of LLM-based agents within the structured, domain-specific environment of Building Information Modeling (BIM). It tests agents' ability to recall and reuse information from prior sessions to answer complex queries requiring both remembered context and live model interaction.

While long-term memory is becoming a critical feature for LLM-based agents, current evaluation methods often focus on conversational recall in general settings. A new benchmark, IFCMemoryBench, proposes a more rigorous test by evaluating agents' ability to retain and apply information from past interactions within a live, structured, and domain-specific environment: Building Information Modeling (BIM). This professional engineering workflow requires agents to query large IFC models while also drawing upon project specifications, client decisions, and engineering conventions often discussed in conversations but not explicitly in the model. IFCMemoryBench comprises 143 multi-session tasks across 19 projects, incorporating 4,016 prior sessions. These tasks are derived from incomplete-information questions, where agents must combine remembered context from earlier conversations with live IFC queries to provide answers. The evaluation framework breaks down memory performance into ingestion, retrieval, and utilization, using expert-validated LLM judges for both answer and memory quality. Initial evaluations of vector-, graph-, and file-based memory systems show that even the strongest system achieves only 32.4% accuracy under realistic conditions, highlighting a significant gap in agent memory for domain-specific professional tasks.

Why it matters

For AI agents to be truly useful in professional domains like engineering, they must reliably retain and apply context over long periods and across multiple interactions, which current general-purpose memory systems struggle with.

How to implement this in your domain

  1. 1Explore the IFCMemoryBench framework to understand its methodology for evaluating long-term memory.
  2. 2Design and test domain-aware memory representations for LLM agents in specialized applications.
  3. 3Integrate multi-session task evaluation into your agent development and testing cycles.
  4. 4Collaborate with domain experts to validate memory quality and answer accuracy for specific use cases.

Original post by Changyu Du, Alexander Vosseler, Filippo Mazza, Andr\'e Borrmann

"arXiv:2607.26072v1 Announce Type: cross Abstract: Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or persona-grounded settings. We argue that a stronger test is whether an agent can reu…"

View on X

Originally posted by Changyu Du, Alexander Vosseler, Filippo Mazza, Andr\'e Borrmann on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses