Active SAE Features Show Less Holonomy in Gemma 2B

Larry Richards· July 24, 2026 View original

Summary

A preregistered study on Gemma 2 2B found that active sparse autoencoder (SAE) feature planes carry *less* holonomy than matched mixed-feature controls, reversing a prior prediction about semantic concentration. The cause remains open, with several alternative mechanisms proposed.

This research investigates the distribution of "holonomy" – a measure of how a local frame rotates when carried around small loops in the residual stream – across active sparse autoencoder (SAE) feature planes in the Gemma 2 2B model. The study aimed to test the prediction that holonomy would concentrate on active feature planes, which is a concrete operationalization of the broader idea that semantic information is concentrated there. The experimental design, analysis, and verdict rules were preregistered to ensure rigor. Contrary to the initial prediction, the study found a reversal: active-feature planes carried *less* holonomy than matched mixed-feature controls. The adjusted log contrast was significantly negative, indicating a clear difference. This result falsifies the hypothesis that holonomy concentrates on active SAE feature planes in this specific model and measurement context. The paper emphasizes that this is an auditable operational reversal, not a causal claim that meaning suppresses holonomy. The underlying cause for this observed phenomenon remains an open question, with several alternative mechanisms suggested, including activation-strength geometry, degree of feature engagement, dictionary geometry, and transport distortion. The study highlights the importance of rigorous, preregistered methodologies in mechanistic interpretability research.

Why it matters

For AI researchers and engineers working on mechanistic interpretability, this finding challenges assumptions about how semantic information is represented and processed within large language models, guiding future research into model internals.

How to implement this in your domain

  1. 1Re-evaluate existing hypotheses about feature representation and information flow within sparse autoencoders based on this new finding.
  2. 2Incorporate holonomy measurements into your mechanistic interpretability toolkit for analyzing model internals.
  3. 3Consider alternative explanations for feature behavior beyond simple semantic concentration, such as activation geometry.
  4. 4Adopt preregistration practices for interpretability research to enhance scientific rigor and reproducibility.

Who benefits

AI ResearchMachine Learning EngineeringAcademia

Key takeaways

  • Active sparse autoencoder (SAE) feature planes in Gemma 2 2B carry less holonomy than expected.
  • This finding reverses a prediction about semantic information concentration in active features.
  • The cause of this phenomenon is still unknown, with multiple potential mechanisms.
  • The study underscores the value of preregistered research in mechanistic interpretability.

Original post by Larry Richards

"arXiv:2607.20522v1 Announce Type: new Abstract: This paper tests whether holonomy concentrates on active sparse-autoencoder (SAE) feature planes in Gemma 2 2B, a concrete operationalization of the broader semantic-concentration prediction. Holonomy is measured at the final-token…"

View on X

Originally posted by Larry Richards on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses