Mechanistic Tomography for AI Interpretability and Control

Vijay Erramilli· August 21, 2026 View original

Key takeaways

  • Mechanistic tomography provides a unified framework for AI interpretability.
  • Designed measurements help recover internal model mechanisms and intervention effects.
  • A practical procedure involves iterative measurement selection and validation.
  • Control systems offer a demanding validation setting for interpretability estimates.

Who benefits

AI/ML DevelopmentAutonomous SystemsCybersecurityHealthcareFinance

Summary

This paper introduces mechanistic tomography, a framework for designed measurement to recover internal mechanisms and intervention effects in AI models. It formalizes how various interpretability techniques measure model internals, providing a practical procedure for selecting and expanding measurement families to improve understanding and control.

Mechanistic interpretability aims to uncover hidden states, component effects, and interactions within AI models that are not directly exposed. This research proposes "mechanistic tomography" as a unified framework to describe how different measurement techniques, such as patching, gradients, and subset interventions, contribute to understanding these internal mechanisms and their responses to interventions. The core idea is to treat these measurements as a system where observed responses (y) relate to target internal maps (x) through an intervention matrix (A) and noise (w). The framework offers a practical methodology: begin with the least costly measurements, validate them on held-out interventions, and refine simple mismatches. If structured residuals persist, expand the family of measurements. The paper highlights that control systems provide a rigorous validation context, as an estimate guiding an intervention acts as an observer, with control error linked to observer error. Experiments on various models, including HMMs and GPT-2-small, demonstrate the framework's utility. For instance, sparse aggregate measurements can recover finite-effect maps with fewer interventions than coordinate patching, and lifted measurements can reveal interactions missed by first-order maps. The study underscores that the appropriate measurement family depends on the chosen basis and the specific internal mechanisms being investigated.

Why it matters

As AI models become more complex and deployed in critical systems, understanding their internal workings and predicting intervention effects is paramount for safety, reliability, and ethical deployment. This framework provides a structured approach to model interpretability.

How to implement this in your domain

  1. 1Adopt a structured approach to interpretability by defining target internal mechanisms and intervention families.
  2. 2Start with low-cost interpretability measurements (e.g., simple gradients) and iteratively expand as needed.
  3. 3Validate interpretability insights by testing on held-out interventions and observing control system performance.
  4. 4Consider using tools like Tracr to explore different measurement families for specific model architectures.

Original post by Vijay Erramilli

"arXiv:2608.19338v1 Announce Type: new Abstract: Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions. Patching, gradients, Hessian-vector products, and subset interven…"

View on X

Originally posted by Vijay Erramilli on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses