AISE-Bench Challenges LLM Agents in Academic Information Seeking

Fanjin Zhang, Zhengyang Wang, Ruixuan Huang, Kefan Zhang, Amy Xin, Yuanchun Wang, Shu Zhao, Evgeny Kharlamov, Jie Tang, Juanzi Li· July 24, 2026 View original

Summary

AISE-Bench is a new, full-cycle benchmark for evaluating LLM agents' information-seeking capabilities on academic knowledge graphs, featuring realistic user intent, complex multi-step API planning, and grounded answers. It reveals that even frontier models struggle with API planning and execution, establishing a challenging testbed.

Large language models (LLMs) augmented with tools are increasingly used as autonomous agents for complex, long-horizon tasks, including information seeking. However, existing benchmarks for evaluating these agents on academic knowledge graphs often rely on simplified scenarios, synthetic templates, or narrow tasks, failing to capture the full complexity of real-world user intent, multi-step API planning, and grounded answer generation. To address these limitations, AISE-Bench has been introduced as a real-world, full-cycle annotated benchmark for information seeking on academic knowledge graphs. It comprises 1,133 question-answer pairs, complete with query taxonomies, detailed API execution trajectories, validated parameters, and source-grounded answers linked to references. A customized agent workflow was designed to facilitate high-quality annotation, enabling annotators to efficiently plan, execute, and revise complex API workflows. The benchmark includes a comprehensive evaluation protocol that measures answer quality, reference grounding, API-planning correctness, and execution success. In tests involving 14 different methods, even the most advanced model (PLAY2PROMPT with Gemini-3-Pro) achieved only moderate performance, frequently encountering difficulties with API planning and execution. AISE-Bench thus serves as a challenging new testbed for rigorously evaluating and improving the stepwise correctness, grounded summarization, and traceable reasoning abilities of multi-step API-using LLM agents.

Why it matters

For professionals developing and deploying LLM agents, AISE-Bench provides a robust tool to identify weaknesses in API planning, execution, and grounded reasoning, enabling the creation of more reliable and accurate information-seeking agents.

How to implement this in your domain

  1. 1Utilize AISE-Bench to evaluate the performance of existing or new LLM agents designed for information retrieval.
  2. 2Focus development efforts on improving LLM agents' multi-step API planning and execution capabilities.
  3. 3Prioritize grounding LLM agent answers with verifiable references, as measured by AISE-Bench.
  4. 4Design agent workflows that support iterative planning and revision, similar to the annotation process used for AISE-Bench.
  5. 5Explore fine-tuning or prompt engineering strategies specifically to address the challenges highlighted by AISE-Bench.

Who benefits

AI DevelopmentResearch & AcademiaInformation ServicesEdTechLegalTech

Key takeaways

  • AISE-Bench is a new benchmark for LLM agents on academic knowledge graphs.
  • It features realistic multi-step API planning and grounded answers.
  • Even frontier LLM agents struggle with API planning and execution.
  • The benchmark helps improve agent correctness, grounding, and reasoning.

Original post by Fanjin Zhang, Zhengyang Wang, Ruixuan Huang, Kefan Zhang, Amy Xin, Yuanchun Wang, Shu Zhao, Evgeny Kharlamov, Jie Tang, Juanzi Li

"arXiv:2607.20498v1 Announce Type: new Abstract: Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-horizon tasks. Current tool-using benchmarks for information seeking on academic…"

View on X

Primary sources

Originally posted by Fanjin Zhang, Zhengyang Wang, Ruixuan Huang, Kefan Zhang, Amy Xin, Yuanchun Wang, Shu Zhao, Evgeny Kharlamov, Jie Tang, Juanzi Li on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses