AI Moral Reasoning Evaluations Miss Norms, Focus on Values

Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh· August 18, 2026 View original

Key takeaways

  • Current AI moral evaluations overemphasize values, neglecting context-sensitive norms.
  • This imbalance stems from reliance on descriptive ethics frameworks.
  • Gaps exist in data, reasoning process evaluation, and context identification.
  • A new research agenda is needed for systematic normative reasoning study.

Who benefits

HealthcareLegalGovernmentAutomotiveTechnology

Summary

This position paper argues that current evaluations of LLM moral competence primarily focus on aligning with human moral values, neglecting the crucial aspect of identifying and applying context-sensitive moral norms. It proposes a research agenda to address this imbalance.

Current methods for evaluating the moral reasoning capabilities of large language models (LLMs) are incomplete, according to a new position paper. The primary focus has been on the "moral value problem," which assesses whether LLM outputs align with human moral values. However, the equally important "moral norm problem"—the ability of models to identify and correctly apply context-sensitive moral norms—remains largely unexplored. This imbalance is attributed to the field's reliance on descriptive ethics frameworks, which prioritize value representation over the practical application of norms. The paper reviews existing benchmarks, demonstrating their heavy concentration on values and a lack of attention to normative ethics. Three key gaps are identified: a scarcity of high-quality ground-truth data for moral norms, insufficient evaluation of intermediate reasoning processes, and limited consideration of morally relevant contextual features. To rectify this, the authors propose a research agenda that includes developing standardized formal representations for normative theories, creating expert-annotated datasets for norm application, and designing evaluation protocols that differentiate between values-level and norms-level competence. The goal is to foster a more systematic study of normative reasoning in LLMs.

Why it matters

For professionals building or deploying AI, understanding the limitations of current moral reasoning evaluations is crucial for developing more robust, ethically sound, and context-aware AI systems, especially in sensitive applications.

How to implement this in your domain

  1. 1Incorporate context-sensitive moral norm datasets into AI training and fine-tuning processes.
  2. 2Design evaluation metrics that specifically assess an AI's ability to apply moral norms in various scenarios, not just align with general values.
  3. 3Collaborate with ethicists and domain experts to create ground-truth data for complex moral norm applications.
  4. 4Develop AI systems with transparent intermediate reasoning processes to better understand their ethical decision-making.

Original post by Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh

"arXiv:2608.14566v1 Announce Type: new Abstract: Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm pro…"

View on X

Originally posted by Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses