LLMs Prioritize Context or Memory Through Activation Directions

Benjamin Shih, John Winnicki, Arianna Cao· September 2, 2026 View original

Key takeaways

  • LLMs use specific activation directions to choose between context and parametric memory.
  • These "authority directions" can causally influence the model's source choice.
  • The reusability of authority directions is limited across different tasks.
  • Understanding this mechanism can lead to more controllable and reliable LLMs.

Who benefits

AI DevelopmentContent GenerationInformation RetrievalCustomer ServiceLegalTech

Summary

Researchers investigated how language models decide between contextual information and their parametric knowledge, finding that specific activation directions can decode and steer this choice. Counterfactual experiments show these "authority directions" can reproduce a significant portion of the model's source choice, though their reusability across tasks is limited.

A recent study delves into the intriguing question of how large language models (LLMs) resolve conflicts between information provided in the input context and knowledge stored within their parameters (memory). The research demonstrates that by analyzing activation directions, it's possible to both decode and influence which source of information an LLM prioritizes. Specifically, "authority directions" were estimated from prompts where context and parametric knowledge aligned, and then these directions were manipulated in counterfactual experiments. The findings indicate that interchanging naturally occurring coordinates along these authority directions between matched prompts can reproduce a substantial portion (30-68%) of the model's decision to favor either context or memory across models like Qwen, Llama, and OLMo. However, the study also highlights limitations regarding cross-task reusability. Authority directions learned on one task showed only about 9% transferability to another task, compared to 57% when using a direction learned specifically for that task. This suggests that while authority representations exist and can causally influence source choice, the underlying computations may be highly task-dependent rather than universally reusable.

Why it matters

Understanding how LLMs balance context and internal knowledge is crucial for improving their reliability, reducing hallucinations, and developing more controllable AI systems for various applications.

How to implement this in your domain

  1. 1Analyze LLM behavior in scenarios where contextual information conflicts with parametric knowledge.
  2. 2Explore techniques to identify and manipulate "authority directions" to steer LLMs towards desired information sources.
  3. 3Develop strategies to enhance context-awareness or parametric knowledge recall based on task requirements.
  4. 4Design evaluation benchmarks that specifically test an LLM's ability to prioritize context over memory, or vice versa.

Original post by Benjamin Shih, John Winnicki, Arianna Cao

"arXiv:2609.00753v1 Announce Type: new Abstract: When contextual information conflicts with the knowledge stored in model parameters, activation directions can be used to decode and steer which source the model follows. However, steering along a direction does not establish causal…"

View on X

Originally posted by Benjamin Shih, John Winnicki, Arianna Cao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses