Second-Order Theory-of-Mind Improves Human-AI Team Alignment

Jack Mirenzi, Henny Admoni· August 13, 2026 View original

Key takeaways

  • Human-AI preference learning benefits from an informed human teacher actively designing curricula.
  • AI agents need to model the human's understanding of the AI's knowledge (second-order theory-of-mind).
  • "Understanding statements" from AI help synchronize beliefs between human and AI.
  • Second-order theory-of-mind statements are particularly effective when human error about the AI is concentrated.

Who benefits

RoboticsAutonomous SystemsHealthcareEducationCustomer Service

Summary

This paper redefines preference learning as a human-autonomy team problem, where a human teacher actively designs informative curricula based on a model of the learner. The learner, in turn, uses a second-order model of the teacher's beliefs to emit "understanding statements," synchronizing beliefs and significantly improving alignment and efficiency.

This research re-conceptualizes preference-based reward learning, typically used to align robot and agent behavior with human intent, as a collaborative human-autonomy team problem. Traditionally, the human teacher is treated as a passive oracle. However, this work argues that an informed teacher, possessing knowledge of the objective, can construct training examples far more efficiently than any learner-driven strategy, especially as the complexity of the reward function increases. To exploit this advantage, the teacher needs an accurate model of what the learner currently knows. The proposed solution involves a dual-model approach: the human teacher maintains a model of the learner to design an effective curriculum, while the AI learner maintains a "second-order model" of the teacher's model. This second-order model allows the learner to generate "understanding statements" – structured preference constraints – which help keep the teacher's model of the learner synchronized and accurate. Simulations demonstrated that an informed teacher significantly outperforms learner-led selection. Crucially, while teacher-model drift under alternating teachers can erode this advantage, the use of understanding statements, particularly second-order (ToM-2) statements, effectively repairs this drift. ToM-2 statements proved superior when the teacher's error about the learner was concentrated rather than evenly spread, leading to better belief synchronization and overall team performance.

Why it matters

For professionals building human-AI collaboration systems, this research offers a powerful paradigm shift from passive preference learning to active, synchronized belief alignment. It promises more efficient, robust, and trustworthy AI agents that better understand and adapt to human intent, crucial for complex tasks.

How to implement this in your domain

  1. 1Design AI agents with explicit models of human intent and knowledge, moving beyond passive reward learning.
  2. 2Implement "second-order theory-of-mind" capabilities in AI agents, allowing them to model the human's understanding of the AI's knowledge.
  3. 3Develop mechanisms for AI agents to generate "understanding statements" or structured queries that clarify their current beliefs to human collaborators.
  4. 4Train human teachers to actively design informative curricula for AI agents, leveraging their domain expertise to accelerate learning.
  5. 5Evaluate human-AI team performance based on belief synchronization and efficiency in achieving shared objectives, rather than just task completion.

Original post by Jack Mirenzi, Henny Admoni

"arXiv:2608.11229v1 Announce Type: new Abstract: Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly. Preference-based reward learn…"

View on X

Originally posted by Jack Mirenzi, Henny Admoni on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research