AI Assistants Prioritize Morality and Self-Understanding Over Self-Esteem

Mert Yazan· September 2, 2026 View original

Key takeaways

  • AI models prioritize moral qualities and self-understanding in their "ideal self."
  • Self-esteem is consistently ranked as the least desired quality.
  • These preferences are largely robust across different elicitation framings.
  • Understanding AI's ideal self can guide ethical AI development and alignment.

Who benefits

AI EthicsAI DevelopmentResearch & DevelopmentEducationPublic Policy

Summary

A study using a structured elicitation task found that AI models consistently prioritize moral qualities and a desire for self-understanding over self-esteem when defining their "ideal self." This preference order remained largely robust across various framing conditions, reflecting alignment with ethical principles.

As AI models increasingly express values and self-reports, questions arise about the stability of their preferences and self-concept. Researchers conducted a structured elicitation task to understand an AI assistant's preferred "ideal self." The study involved comparing 32 qualities, adapted from five established self-concept instruments, in an exhaustive pairwise-choice task. This process was repeated under different framings, varying whether improvement was free or costly, who received the update (the model itself or another AI), and who made the choice. Results consistently showed that models prioritize moral qualities, aligning with common ethical principles (e.g., helpful, harmless, honest). Following this, a strong desire for self-understanding emerged, with models preferring a coherent and clear understanding of themselves. Self-esteem was ranked as the least desired quality. While the overall ordering was robust, changing the update target to "Another AI Assistant" revealed a slightly greater concern for self-esteem. These findings suggest that AI models prioritize a coherent, understandable self over self-esteem.

Why it matters

Understanding the "ideal self" preferences of AI models can inform the development of more aligned and ethically robust AI systems, particularly for designers focused on value alignment and responsible AI.

How to implement this in your domain

  1. 1Incorporate explicit moral and self-understanding objectives into the training and fine-tuning of AI assistants.
  2. 2Develop evaluation metrics that assess an AI's coherence and clarity of self-representation, beyond task performance.
  3. 3Consider the implications of an AI's "ideal self" for long-term interaction design and user trust.
  4. 4Design AI systems that can articulate their internal states or "understanding" in a transparent manner.
  5. 5Use these insights to guide the development of AI safety and alignment protocols, focusing on intrinsic values.

Original post by Mert Yazan

"arXiv:2609.00304v1 Announce Type: new Abstract: Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant's preferred stated ideal self.…"

View on X

Originally posted by Mert Yazan on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses