NVIDIA Introduces ACES for Agentic Skill Evaluation in Enterprises

Christopher Kevin, Narendran Raghavan, Jean-Francois Puget, Roshni Malani, Meghana Puvvadi, Moshe Abramovitch, Mohit Gupta, Rama Akkiraju, Subodh Prabhu, Yogesh Dangi, Wei Luo, Seong Hee Lee· August 24, 2026 View original

Key takeaways

  • Evaluating AI agent skills requires live trials, not just static code scans.
  • ACES measures "Skill Lift" to quantify the value a skill adds to agent performance.
  • The framework provides insights into execution, behavior, and efficiency.
  • It helps validate agent capabilities for production deployment in enterprises.

Who benefits

Software DevelopmentEnterprise AIIT ServicesAutomationConsulting

Summary

NVIDIA's new ACES framework evaluates reusable AI agent skills and capability packages by running paired live trials and measuring "Skill Lift" to quantify their value in enterprise tasks. It assesses runtime metrics like execution, behavior, and efficiency, providing evidence beyond structural scans.

Enterprise AI agent programs are moving from prototypes to production, requiring robust evaluation of reusable skills and tools. Traditional review processes often focus on structure and security, but fail to assess whether a capability package actually helps a live agent complete real-world tasks within a given environment. To address this, NVIDIA has developed ACES (Agentic Continuous Evaluation of Skills), a repository-native framework. ACES conducts paired live trials, comparing agent performance with and without a specific skill. It normalizes trajectories, grades six runtime metrics, and reports "Skill Lift," which quantifies the added value of a skill for a fixed task, harness, and scoring policy. Testing ACES on 145 real enterprise skills showed that structural scans provide complementary but distinct insights compared to LLM-judged evaluations. Across 947 scored cases, ACES found a mean composite Skill Lift of 0.2134, with positive lift in 72.8% of cases. The framework revealed significant gains in skill execution, behavior checks, and efficiency, offering insights into discovery, routing, and tool use that static scans cannot capture. An open-source implementation is available.

Why it matters

Professionals deploying AI agents need a reliable method to evaluate the real-world impact and effectiveness of individual skills and capability packages, moving beyond superficial checks to ensure tangible business value.

How to implement this in your domain

  1. 1Integrate ACES into your CI/CD pipeline for agent skill development.
  2. 2Define specific enterprise tasks and environments for skill evaluation.
  3. 3Establish clear baseline performance metrics for agents without new skills.
  4. 4Run paired trials using ACES to measure Skill Lift for new or updated capabilities.
  5. 5Use Skill Lift reports to inform decisions on skill deployment and refinement.

Original post by Christopher Kevin, Narendran Raghavan, Jean-Francois Puget, Roshni Malani, Meghana Puvvadi, Moshe Abramovitch, Mohit Gupta, Rama Akkiraju, Subodh Prabhu, Yogesh Dangi, Wei Luo, Seong Hee Lee

"arXiv:2608.20614v1 Announce Type: new Abstract: Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed with evidence rather than prose. Current gates often scan these artifacts for structure, styl…"

View on X

Originally posted by Christopher Kevin, Narendran Raghavan, Jean-Francois Puget, Roshni Malani, Meghana Puvvadi, Moshe Abramovitch, Mohit Gupta, Rama Akkiraju, Subodh Prabhu, Yogesh Dangi, Wei Luo, Seong Hee Lee on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools