LLMs Reveal Internal Conflict Resolution in System vs. User Instructions

Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi, Sunishchal Dev, Callum Stuart McDougall, Anusha Mujumdar· September 1, 2026 View original

Key takeaways

  • LLMs exhibit varying behaviors when system and user instructions conflict.
  • Some models, like Llama-3.1-8B, show an "anti-hierarchy" preference for user instructions.
  • Internal conflict-resolution signals are decodable from LLM activations, even in anti-hierarchy models.
  • Steering these internal signals can significantly improve system compliance.

Who benefits

AI DevelopmentCybersecurityCustomer ServiceContent ModerationRobotics

Summary

Researchers investigated how instruction-tuned LLMs resolve conflicts between system and user instructions, identifying models as hierarchy-respecting, anti-hierarchy, or no-effect. They found that even anti-hierarchy models like Llama-3.1-8B possess decodable internal signals for conflict resolution, which can be steered to improve system compliance.

This research delves into the internal mechanisms by which instruction-tuned Large Language Models (LLMs) handle conflicting instructions from system prompts versus user inputs. A new benchmark with paired constraints and deterministic verifiers was used to evaluate eight models, categorizing their behavior into three types: hierarchy-respecting, anti-hierarchy (preferring user instructions), or showing no channel sensitivity. Notably, Llama-3.1-8B emerged as a strong anti-hierarchy model, frequently disregarding system instructions. Despite this behavioral preference, the study revealed that even in anti-hierarchy models like Llama-3.1-8B, an internal conflict-resolution signal is highly decodable from residual-stream activations. This signal can be leveraged for steering, significantly improving system compliance. The findings suggest that user-preferring arbitration does not imply an absence of internal conflict awareness, but rather a specific internal weighting, and that successful intervention depends on understanding the geometry of these internal representations.

Why it matters

Understanding and steering how LLMs prioritize instructions is critical for professionals developing reliable and controllable AI agents, especially in applications requiring strict adherence to safety guidelines or specific operational protocols.

How to implement this in your domain

  1. 1Develop internal benchmarks to test LLM behavior under conflicting system and user instructions.
  2. 2Analyze the "System Authority Delta" of chosen LLMs to understand their default instruction prioritization.
  3. 3Explore techniques for steering LLM internal representations to enforce desired instruction hierarchies.
  4. 4Implement robust prompt engineering strategies that account for potential instruction conflicts and model biases.
  5. 5Design AI agent workflows with explicit conflict resolution mechanisms, potentially informed by internal signal decoding.

Original post by Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi, Sunishchal Dev, Callum Stuart McDougall, Anusha Mujumdar

"arXiv:2608.28648v1 Announce Type: new Abstract: We study how instruction-tuned LLMs arbitrate direct conflicts between system and user instructions. We introduce a benchmark of 41 paired constraints with deterministic verifiers and evaluate eight models under matched baseline, co…"

View on X

Originally posted by Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi, Sunishchal Dev, Callum Stuart McDougall, Anusha Mujumdar on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses