Federated Learning Aggregation Robustness Under Poisoning and Backdoor Attacks

Soumya Mazumdar, Vineet Kumar Rakesh, Tapas Samanta· August 13, 2026 View original

Key takeaways

  • Trimmed Mean offers high accuracy in clean federated learning environments.
  • Krum is highly effective against sign-flipping and Gaussian model poisoning attacks.
  • Existing backdoor attack metrics and aggregation scaffolds can have critical implementation flaws.
  • Robust federated learning requires careful selection of aggregation methods and rigorous metric validation.

Who benefits

CybersecurityAI DevelopmentFinanceHealthcareTelecommunications

Summary

This research reconstructs and analyzes a benchmark for federated aggregation methods under various attacks, revealing that Trimmed Mean excels in clean accuracy while Krum performs best against sign-flipping and Gaussian attacks. It also identifies issues with existing metric implementations for backdoor attacks.

Federated learning, a distributed machine learning approach, faces significant security challenges, particularly from model poisoning and backdoor attacks. This research undertakes a comprehensive reconstruction and analysis of a 500-cell benchmark matrix to evaluate the robustness of different federated aggregation methods. The study considers five aggregation techniques, five datasets, five architectures, and four attack conditions: clean, sign-flipping, Gaussian, and BadNets. The analysis, based on 454 original and 36 repaired/rerun executions, revealed distinct performance profiles for the aggregation methods. Trimmed Mean demonstrated the highest macro-mean accuracy (76.02%) in clean conditions and the lowest mean within-task rank. Conversely, Krum proved most effective against both sign-flipping and Gaussian attacks, achieving the highest recorded accuracy under these adversarial configurations. These relative rankings remained consistent even when the analysis was restricted to tasks with complete original logs. Crucially, the researchers also identified critical issues with the supplied metric implementations. An audit of the BadNets metric revealed that it measures "Triggered Target-Label Rate" (TTLR) rather than a conventional attack success rate that excludes the target label. Furthermore, a potential discrepancy was found in the FedPARETO scaffold, where predictive summaries might reflect an uncorrupted local model while a separately corrupted update is applied for aggregation. The study emphasizes that its findings are descriptive comparisons within the recorded configurations due to the single seed per cell and incomplete attack lineage.

Why it matters

For professionals involved in deploying or securing federated learning systems, understanding the vulnerabilities to various attacks and the comparative robustness of aggregation methods is critical for building secure and reliable AI.

How to implement this in your domain

  1. 1Prioritize robust aggregation methods like Trimmed Mean for clean performance and Krum for resilience against specific poisoning attacks in federated learning deployments.
  2. 2Conduct thorough audits of evaluation metrics and implementation scaffolds to ensure they accurately reflect attack success and model behavior.
  3. 3Diversify testing scenarios to include various attack types (e.g., sign-flipping, Gaussian, backdoor) and dataset/architecture combinations.
  4. 4Investigate the use of multiple seeds and more comprehensive attack lineage tracking in internal benchmarks for statistical robustness.

Original post by Soumya Mazumdar, Vineet Kumar Rakesh, Tapas Samanta

"arXiv:2608.11423v1 Announce Type: new Abstract: Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. A 500-cell seed-1 evaluation matrix was reconstructed across…"

View on X

Originally posted by Soumya Mazumdar, Vineet Kumar Rakesh, Tapas Samanta on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses