CANN Bench: New Benchmark for AI-Generated Kernels on Huawei NPUs.

Xue-Jian Gao, Deng Pan, Yueming Su, Jiasheng Li, Bin Du, Fengming Zhu, Chengdi Ma, Junyi Fan, Qichen Liao, Chengqiu Hu, Xinxian Chen, Lingchao Zheng, Jun Li, Jiwei Yang, Yuwei Fan· July 24, 2026 View original

Summary

CANN Bench is an open benchmark for evaluating AI-generated operator code on Huawei's Ascend NPU, covering 53 operators and 1060 test cases across various precisions. It uses a three-dimensional weighted score for compilation, correctness, and performance against hardware limits.

This new research introduces CANN Bench, an open-source benchmark designed to evaluate the performance of AI-generated operator kernels specifically on Huawei's Ascend Neural Processing Units (NPUs). Unlike existing benchmarks that primarily focus on CUDA and Triton, CANN Bench addresses the need for a standardized evaluation framework in less-exposed hardware ecosystems. The benchmark includes a comprehensive set of 53 operators and 1060 test cases, ranging from basic element-wise operations to complex FlashAttention kernels, supporting multiple precision formats like FP16, BF16, FP32, and INT8. The evaluation methodology employs a sophisticated three-dimensional weighted composite score, which independently assesses compilation success, functional accuracy, and execution performance. Performance is measured against both a PyTorch-on-Ascend baseline and an analytical Hardware-Anchored Performance (HAP) limit, ensuring that reported scores reflect genuine optimization potential rather than measurement artifacts. The framework is built to prevent reward hacking and is intended for long-term community collaboration within the official CANN repository, providing a robust tool for advancing AI operator-authoring capabilities on Ascend NPUs.

Why it matters

Professionals developing AI models or hardware will find this crucial for evaluating and optimizing low-level AI operator performance on Huawei's Ascend NPUs, ensuring efficient deployment and competitive advantage in specific hardware ecosystems.

How to implement this in your domain

  1. 1Integrate CANN Bench into your CI/CD pipeline for automated performance testing of AI-generated kernels.
  2. 2Utilize the benchmark's scoring system to identify specific areas for optimization in your operator code (compilation, correctness, performance).
  3. 3Compare your AI agent's kernel generation capabilities against the provided baselines and hardware limits to gauge true optimization headroom.
  4. 4Contribute to the CANN Bench community to help expand its coverage and refine evaluation methodologies for Ascend NPUs.

Who benefits

SemiconductorAI/ML DevelopmentCloud ComputingAutomotive

Key takeaways

  • CANN Bench provides a critical evaluation tool for AI-generated operator kernels on Huawei Ascend NPUs.
  • Its three-dimensional scoring system offers a comprehensive assessment of compilation, correctness, and performance.
  • The benchmark helps identify genuine optimization headroom by comparing against hardware-anchored performance limits.
  • It fosters community collaboration for advancing AI operator development in specific hardware ecosystems.

Original post by Xue-Jian Gao, Deng Pan, Yueming Su, Jiasheng Li, Bin Du, Fengming Zhu, Chengdi Ma, Junyi Fan, Qichen Liao, Chengqiu Hu, Xinxian Chen, Lingchao Zheng, Jun Li, Jiwei Yang, Yuwei Fan

"arXiv:2607.20518v1 Announce Type: new Abstract: AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchmarks, however, focus almost exclusively on CUDA and Triton, leaving hardware ecos…"

View on X

Originally posted by Xue-Jian Gao, Deng Pan, Yueming Su, Jiasheng Li, Bin Du, Fengming Zhu, Chengdi Ma, Junyi Fan, Qichen Liao, Chengqiu Hu, Xinxian Chen, Lingchao Zheng, Jun Li, Jiwei Yang, Yuwei Fan on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses