New Framework Quantifies LLM Reasoning Effort in Chain-of-Thought Steps

Hui Wei, Junda Wu, Sheldon Yu, Sizhe Zhou, Yizhu Jiao, Ming Zhong, Bowen Jin, Tong Yu, Shijia Pan, Jiawei Han, Julian McAuley· August 3, 2026 View original

Key takeaways

  • SARE quantifies LLM reasoning effort at individual chain-of-thought steps.
  • Reasoning energy is non-uniform, with incorrect paths showing lower energy at critical points.
  • Internal geometric dynamics provide predictive information beyond output-level signals.
  • This framework offers a new lens for LLM interpretability and reliability.

Who benefits

AI DevelopmentSoftware EngineeringResearch & DevelopmentQuality Assurance

Summary

Researchers developed Step-Aware Reasoning Energy (SARE), a geometric framework using Centered Kernel Alignment (CKA) to measure computational effort at each step of an LLM's chain-of-thought reasoning. This method reveals non-uniform energy allocation and systematically lower energy in incorrect reasoning paths.

A new research paper introduces the Step-Aware Reasoning Energy (SARE) framework, designed to provide a granular understanding of how large language models (LLMs) expend computational effort during their chain-of-thought (CoT) reasoning processes. Unlike previous methods that offer only a high-level view, SARE quantifies effort at the individual step level. The SARE framework leverages Centered Kernel Alignment (CKA) to analyze the hidden states of tokens across transformer layers, capturing the intricate relational structure between tokens. This allows for the identification of "phase-like transitions" in reasoning energy that are otherwise invisible to broader, trajectory-level metrics. Experiments across various benchmarks and LLMs revealed that reasoning energy is highly uneven across different step types. Notably, incorrect reasoning trajectories consistently exhibited lower energy at crucial decision points. SARE-based features also proved effective in predicting outcomes, often outperforming output-based confidence measures, suggesting that internal geometric dynamics hold significant predictive information.

Why it matters

Understanding the internal reasoning dynamics of LLMs can lead to more reliable and interpretable AI systems, enabling developers to diagnose failures and improve model performance. This research offers a novel way to assess LLM "thinking" beyond just their final outputs.

How to implement this in your domain

  1. 1Integrate SARE-like metrics into LLM development pipelines to monitor internal reasoning quality.
  2. 2Develop debugging tools that visualize step-aware reasoning energy to identify problematic CoT steps.
  3. 3Utilize SARE features for real-time confidence scoring or early error detection in LLM applications.
  4. 4Design training strategies that encourage more robust energy allocation at critical reasoning junctions.

Original post by Hui Wei, Junda Wu, Sheldon Yu, Sizhe Zhou, Yizhu Jiao, Ming Zhong, Bowen Jin, Tong Yu, Shijia Pan, Jiawei Han, Julian McAuley

"arXiv:2607.28674v1 Announce Type: new Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods rely on output-level signals or collapse processing depth into…"

View on X

Originally posted by Hui Wei, Junda Wu, Sheldon Yu, Sizhe Zhou, Yizhu Jiao, Ming Zhong, Bowen Jin, Tong Yu, Shijia Pan, Jiawei Han, Julian McAuley on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses