Telco-GAIA Benchmarks Bilingual AI Agents for Telecom

Dmitrii Khizbullin, Zaid Alyafeai, Abdelrahman Eldesokey, Nourah AlSultan, Raghad Alshalan, David R. Pugh, Bernard Ghanem· July 24, 2026 View original

Summary

Telco-GAIA is a new bilingual, multi-modal benchmark for evaluating tool-using AI agents in the telecommunications domain. It features 100 human-verified tasks in English and Arabic, requiring multi-hop reasoning over diverse data sources within a sandboxed environment.

Evaluating the capabilities of tool-using AI agents, especially in specialized domains, presents a significant challenge. To address this, a new benchmark called Telco-GAIA has been introduced, specifically designed for the telecommunications sector. This benchmark is bilingual, supporting both English and Arabic, and multi-modal, incorporating various data types. Telco-GAIA consists of 100 human-verified question-answering tasks, each demanding multi-hop reasoning, averaging 4.2 hops. These tasks require agents to process information from three heterogeneous sources: a static website snapshot (including HTML, images, and linked PDFs), a synthetic relational SQL database, and external web archives, spanning text, image, and tabular modalities. The benchmark is delivered as a sandboxed Docker environment, ensuring objective, deterministic, and reproducible evaluation through normalized exact string matching, without relying on LLM-as-a-Judge metrics. Initial evaluations of a purpose-built reference agent across twelve commercial and open LLMs reveal that Telco-GAIA is highly challenging. Even the strongest model solved only 71% of tasks, dropping to about 40% under a moderate cost budget. Visually grounded categories proved particularly difficult, with an average backend score below 30%, indicating substantial room for improvement in document and image understanding for AI agents.

Why it matters

For professionals developing or deploying AI agents in enterprise settings, particularly in telecommunications, Telco-GAIA provides a rigorous, reproducible standard for evaluating agent performance, highlighting current limitations and guiding future development.

How to implement this in your domain

  1. 1Utilize Telco-GAIA to benchmark existing or new AI agents for telecom-specific tasks.
  2. 2Identify weaknesses in current agent performance, especially in multi-modal and multi-hop reasoning.
  3. 3Focus development efforts on improving document and image understanding capabilities of agents.
  4. 4Adapt the Telco-GAIA template to create custom closed-domain benchmarks for other industries.
  5. 5Collaborate with research teams to contribute to the benchmark's evolution and agent improvement.

Who benefits

TelecommunicationsCustomer ServiceAI/ML DevelopmentEnterprise SoftwareData Analytics

Key takeaways

  • Telco-GAIA is a new bilingual, multi-modal benchmark for telecom AI agents.
  • It features 100 human-verified, multi-hop reasoning tasks over diverse data sources.
  • The benchmark is sandboxed, objective, and reproducible, not relying on LLM-as-a-Judge.
  • Current LLMs struggle significantly, especially with visually grounded tasks, showing much room for improvement.

Original post by Dmitrii Khizbullin, Zaid Alyafeai, Abdelrahman Eldesokey, Nourah AlSultan, Raghad Alshalan, David R. Pugh, Bernard Ghanem

"arXiv:2607.20510v1 Announce Type: new Abstract: We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Telco-GAIA comprises 100 human-verified question-answering tasks, in English and A…"

View on X

Originally posted by Dmitrii Khizbullin, Zaid Alyafeai, Abdelrahman Eldesokey, Nourah AlSultan, Raghad Alshalan, David R. Pugh, Bernard Ghanem on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses