TRACE Benchmark Diagnoses Human-AI Coordination Failures Under Drift

Joshua Zuniga, Srinivasan Subramanian, Ramya Madhuri Narapureddy, Md Abdullah Al Hafiz Khan· August 10, 2026 View original

Key takeaways

  • Trustworthiness in human-AI systems depends on understanding multi-layer interactions.
  • TRACE is a new benchmark for diagnosing drift and failures in these systems.
  • It provides time-aligned, multi-layer traces with detailed drift annotations.
  • Baseline studies confirm its utility in identifying and attributing drift.

Who benefits

RoboticsManufacturingAutonomous VehiclesHealthcareAerospace

Summary

TRACE is a new multi-layer benchmark designed to diagnose coordination breakdowns in human-AI systems under drift and failure conditions. It provides time-aligned traces across five execution layers, derived from household tasks, to help localize and attribute drift in cyber-physical and AI-assisted systems.

Modern AI-assisted and cyber-physical systems involve complex interactions between human operators, AI decision modules, and automated controllers. Ensuring trustworthiness requires understanding how drift and failures propagate across these layers, yet no standard benchmark exists to capture these multi-layer interactions. This paper introduces TRACE, a benchmark specifically designed to address this gap, focusing on drift. TRACE injects controlled drift into traces from the ALFRED benchmark, which covers everyday household tasks, creating 1,918 drifted traces. Each trace records step-by-step data across five layers (state, observation, decision, rules, control) and is labeled with drift type, affected layer, onset time, responsible actor, and causal mechanism. A baseline study using classical, recurrent, and attention-based models shows that drift is identifiable and attributable across all model families, indicating TRACE's effectiveness in diagnosing coordination issues.

Why it matters

Professionals developing or deploying human-AI systems can use this benchmark to rigorously test system resilience, identify failure points, and improve the robustness and trustworthiness of their integrated AI solutions.

How to implement this in your domain

  1. 1Utilize the TRACE benchmark to evaluate the robustness of existing human-AI control loops.
  2. 2Develop monitoring tools that capture multi-layer traces for drift detection and attribution.
  3. 3Integrate drift detection mechanisms into AI-assisted cyber-physical systems.
  4. 4Train engineering teams on diagnosing and mitigating drift propagation across system layers.

Original post by Joshua Zuniga, Srinivasan Subramanian, Ramya Madhuri Narapureddy, Md Abdullah Al Hafiz Khan

"arXiv:2608.06657v1 Announce Type: new Abstract: Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustworthiness depends on the whole loop, not any one model. Yet no standard benchmar…"

View on X

Originally posted by Joshua Zuniga, Srinivasan Subramanian, Ramya Madhuri Narapureddy, Md Abdullah Al Hafiz Khan on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses