HERO Library Benchmarks Federated Continual Learning for Diverse Data.

Thinh T. H. Nguyen, Le-Tuan Nguyen, Minh-Duong Nguyen, Nhi Trinh, Anh Tran Nam Nguyet, Dung D. Le, Kok-Seng Wong· July 13, 2026 View original

Key takeaways

  • HERO standardizes federated continual learning evaluation by isolating key heterogeneity factors.
  • Traditional average accuracy metrics can hide poor performance in specific client groups.
  • Task-order mismatch significantly impacts the effectiveness of FCL strategies.
  • The benchmark helps design more robust FCL systems for real-world distributed data.

Who benefits

HealthcareAutomotiveIoTFinanceTelecommunications

Summary

Researchers introduce HERO, a new benchmark library designed to standardize the evaluation of federated continual learning methods, accounting for varying data distributions and task sequences across clients. It helps compare different FCL approaches by disentangling key experimental variables.

Federated continual learning (FCL) involves distributed clients learning from evolving data streams while retaining past knowledge. A major challenge in this field has been the lack of standardized evaluation, making it difficult to compare different FCL methods due to simultaneous variations in datasets, task splits, client data distributions, and task orders. To address this, a new benchmark library called HERO has been developed. HERO allows researchers to systematically control and isolate three critical factors: task split, client data split, and client task sequence. This enables more rigorous and comparable evaluations of FCL algorithms. Initial evaluations using HERO on datasets like CIFAR-100 and TinyImageNet reveal that method performance can vary significantly across different settings, and average accuracy metrics can mask poor performance among individual clients. The benchmark also highlights how task-order mismatch influences strategy effectiveness and provides a framework for exploring domain-shift challenges beyond image data.

Why it matters

For professionals working with distributed AI systems, this benchmark offers a standardized way to evaluate and select robust continual learning algorithms, ensuring they perform reliably across diverse and evolving real-world data environments. It helps in understanding the true performance of FCL methods beyond simple average metrics.

How to implement this in your domain

  1. 1Adopt HERO for evaluating new federated learning models to ensure robust performance across heterogeneous client data.
  2. 2Analyze existing FCL deployments using HERO's metrics to identify potential weaknesses in client-specific performance.
  3. 3Design new FCL algorithms with specific attention to heterogeneity parameters like client data skew and task-order mismatch.
  4. 4Contribute to the HERO library by adding new datasets or FCL methods for broader community benefit.

Original post by Thinh T. H. Nguyen, Le-Tuan Nguyen, Minh-Duong Nguyen, Nhi Trinh, Anh Tran Nam Nguyet, Dung D. Le, Kok-Seng Wong

"arXiv:2607.08784v1 Announce Type: new Abstract: Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge. Existing evaluations are difficult to compare because they often change datasets, ta…"

View on X

Originally posted by Thinh T. H. Nguyen, Le-Tuan Nguyen, Minh-Duong Nguyen, Nhi Trinh, Anh Tran Nam Nguyet, Dung D. Le, Kok-Seng Wong on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI Engineering & DevToolsAI Research

Resilient Decentralized Federated Learning for Wireless IoT Networks

This paper introduces QEF-GT-AdamW, a communication-efficient and outage-resilient algorithm for decentralized federated learning over wireless IoT networks. It combines gradient tracking, AdamW optimization, and dual-stream biased quantization with error feedback to improve robustness and convergence under heterogeneous data and unreliable communication.

Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Vu Nguyen Ha, Symeon ChatzinotasAug 27, 2026
AI Engineering & DevToolsAI Research

FedQoS Predicts QoS Risk for Wireless Access Selection

This paper proposes FedQoS, a federated QoS-risk learning framework that predicts future QoS degradation for reliable access selection in heterogeneous indoor-outdoor wireless environments. It enables access nodes to locally learn from network logs and collaboratively train a global predictor without centralizing user data, significantly reducing QoS failure rates.

Nguyen Van Thieu, Ti Ti Nguyen, Ons Aouedi, Zerihun Huruy, Vu Nguyen Ha, Symeon ChatzinotasAug 27, 2026
AI ResearchAI Engineering & DevTools

Parametric Knowledge Graphs Show Storage-Retrieval Gap

This paper explores compiling knowledge graphs into LoRA adapters for parametric memory, finding that while adapters effectively store factual knowledge, retrieving it via semantic similarity or weight-space geometry is ineffective. This highlights a "storage-retrieval gap" and the need for new query-conditioned composition mechanisms.

Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov, Volker TrespAug 27, 2026