Flower Hub Platform Enhances Federated Learning Benchmarking

Yan Gao, Mohammad Naseri, Javier Fernandez-Marques, Dimitris Stripelis, Lorenzo Sani, Davide Eynard, Fan Zhang, Hong Jia, Ting Dang, D. B. Emerson, Fatemeh Tavakoli, Ole Werger, Lars Wulfert, Petros Demetrakopoulos, Sofia Tsekeridou, InSeo Song, KangYoon Lee, Honghao Li, Lingjuan Lyu, John P Dickerson, Daniel Janes Beutel, Nicholas D. Lane· August 27, 2026 View original

Key takeaways

  • Flower Hub standardizes federated learning benchmarking for reproducibility and comparability.
  • It allows the same benchmark application to run in both simulation and deployment.
  • The platform includes a multi-domain suite covering diverse FL tasks.
  • It supports system-aware reporting, including runtime and communication metrics.

Who benefits

HealthcareBFSILegalTechAI ResearchCybersecurity

Summary

Flower Hub is introduced as a new platform designed to improve the reproducibility, comparability, and extensibility of federated learning (FL) benchmarks. It packages benchmarks as executable, versioned applications, enabling unified evaluation across both simulation and real-world deployment environments.

Researchers have launched Flower Hub, a novel platform aimed at standardizing and streamlining the benchmarking process for federated learning (FL) applications. The platform addresses common challenges in FL evaluation, such as reproducibility issues, difficulties in comparing results across different studies, and limited extensibility of existing benchmarks. Flower Hub achieves this by packaging benchmarks as self-contained, executable, and versioned applications, complete with standardized metadata, pinned dependencies, and explicit evaluation workflows. The platform includes a multi-domain benchmark suite covering diverse settings like cross-silo and cross-device FL, with tasks ranging from medical imaging and financial tabular learning to legal instruction tuning and phishing URL detection. A key feature is its ability to run the same benchmarking application seamlessly across both simulated and deployed environments without requiring code changes, ensuring consistent evaluation. Beyond model quality, Flower Hub also supports system-aware reporting, capturing crucial metrics like runtime and communication overhead. This initiative marks a significant step towards more portable, executable, and reusable FL benchmarks.

Why it matters

For professionals working with federated learning, Flower Hub provides a critical tool for reliably evaluating and comparing FL models and systems, accelerating development and deployment of privacy-preserving AI.

How to implement this in your domain

  1. 1Explore Flower Hub to discover existing federated learning benchmarks relevant to your industry or use case.
  2. 2Utilize the platform to package your own FL models and datasets into reproducible benchmark applications.
  3. 3Run benchmarks on Flower Hub to compare your FL solutions against state-of-the-art methods in both simulation and real-world deployment.
  4. 4Contribute new benchmarks or extend existing ones to foster community collaboration and standardized evaluation.

Original post by Yan Gao, Mohammad Naseri, Javier Fernandez-Marques, Dimitris Stripelis, Lorenzo Sani, Davide Eynard, Fan Zhang, Hong Jia, Ting Dang, D. B. Emerson, Fatemeh Tavakoli, Ole Werger, Lars Wulfert, Petros Demetrakopoulos, Sofia Tsekeridou, InSeo Song, KangYoon Lee, Honghao Li, Lingjuan Lyu, John P Dickerson, Daniel Janes Beutel, Nicholas D. Lane

"arXiv:2608.25114v1 Announce Type: new Abstract: Federated learning (FL) has emerged as a key approach for training models across decentralized data, yet benchmarking in FL remains difficult to reproduce, compare, and extend. Existing evaluations are often tied to custom infrastru…"

View on X

Originally posted by Yan Gao, Mohammad Naseri, Javier Fernandez-Marques, Dimitris Stripelis, Lorenzo Sani, Davide Eynard, Fan Zhang, Hong Jia, Ting Dang, D. B. Emerson, Fatemeh Tavakoli, Ole Werger, Lars Wulfert, Petros Demetrakopoulos, Sofia Tsekeridou, InSeo Song, KangYoon Lee, Honghao Li, Lingjuan Lyu, John P Dickerson, Daniel Janes Beutel, Nicholas D. Lane on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools