New Framework Diagnoses AI Agent Behavior Origins

Xichen Zhang, Yingjie Zhang, Tianshu Sun· July 21, 2026 View original

Summary

This perspective proposes "layer attribution," a diagnostic framework for AI agent behavior that distinguishes between foundational computational capabilities and behavioral modulation layers. It clarifies how behavior originates from architecture, memory, and objectives versus social interaction and governance, impacting evaluation and intervention strategies.

As AI agents become increasingly integrated into complex human systems like healthcare, politics, and science, understanding their behavior requires a robust diagnostic approach. This paper introduces "layer attribution," a framework designed to pinpoint the origins of AI agent behavior. The framework posits two distinct layers: the foundational computational layer, which dictates what behaviors are possible through the agent's architecture, memory, perception, and representation; and the behavioral modulation layer, which shapes how these capacities are expressed through identity, resources, objectives, social interaction, and governance rules. This distinction has three key implications: surrogate validity becomes a relationship between the model, task, and layer; human-AI divergence offers crucial diagnostic evidence; and effective governance necessitates attributing behavior to its source before intervention. The framework emphasizes that evaluating, explaining, and governing AI agents requires understanding where their behaviors truly originate.

Why it matters

Professionals involved in AI ethics, governance, and system design can use this framework to better understand, evaluate, and control AI agent behavior, ensuring safer and more aligned deployment in critical domains.

How to implement this in your domain

  1. 1Apply the layer attribution framework to analyze the behavior of AI agents in your organization.
  2. 2Distinguish between architectural limitations and behavioral modulation factors when diagnosing AI failures.
  3. 3Design AI systems with clear separation or traceability between computational and behavioral layers.
  4. 4Develop governance policies that target the appropriate layer of AI behavior for effective intervention.

Who benefits

AI Ethics & GovernancePublic PolicyHealthcareLegalResearch & Development

Key takeaways

  • AI agent behavior originates from distinct computational and modulation layers.
  • The layer attribution framework helps diagnose where AI behavior comes from.
  • Understanding these layers is crucial for valid evaluation and effective governance.
  • Interventions should target the specific layer responsible for observed behavior.

Original post by Xichen Zhang, Yingjie Zhang, Tianshu Sun

"arXiv:2607.17149v1 Announce Type: new Abstract: AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an…"

View on X

Originally posted by Xichen Zhang, Yingjie Zhang, Tianshu Sun on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses