Gated Steering Reduces LLM Sycophancy, Hallucination in Medicine

Himanshu Tripathi, Subash Neupane, Shaswata Mitra, Sudip Mittal, Noorbakhsh Amiri Golilarz, Shahram Rahimi· August 26, 2026 View original

Key takeaways

  • Sycophancy and hallucination are critical LLM failure modes, especially in medicine.
  • Gated Activation Steering uses targeted, inference-time interventions to mitigate both.
  • The framework learns separate steering directions and applies them via behavior-specific gates.
  • It significantly improves LLM robustness in medical Q&A, comparable to much larger models.

Who benefits

HealthcareHealthTechPharmaceuticalsAI DevelopmentMedical Devices

Summary

This research introduces Gated Activation Steering, an Inference Time Intervention (ITI) framework that jointly mitigates sycophancy and hallucination in LLMs for medical question answering. It uses behavior-specific gates to apply targeted interventions, significantly improving model robustness without constant intervention.

Large Language Models (LLMs) frequently exhibit sycophancy (agreeing with user pressure) and hallucination (generating unsupported information), which are particularly problematic in critical domains like clinical question answering. Existing solutions often address these issues separately or apply broad interventions that can degrade otherwise correct responses. This study proposes a novel framework to tackle both problems simultaneously. The researchers developed Gated Activation Steering, an Inference Time Intervention (ITI) method. This approach learns distinct steering directions for hallucination and sycophancy from contrastive clinical data and applies them to specific, causally verified attention heads within the LLM. During runtime, behavior-specific gates dynamically determine when an intervention is necessary, allowing for targeted mitigation: the hallucination component addresses unsupported claims, while the sycophancy component prevents answer shifts under user pressure. Evaluated on clinical questions grounded in Electronic Health Record (EHR) data, the framework demonstrated significant improvements. For a 4-billion-parameter model, gated steering helped the model resist user pressure in a vast majority of cases, achieving robustness levels comparable to much larger models. This indicates that targeted, inference-time steering can substantially enhance LLM reliability in medical contexts without the need for continuous, broad interventions.

Why it matters

Healthcare professionals and AI developers can deploy more reliable LLMs for medical applications by using targeted steering techniques to reduce critical failure modes like hallucination and sycophancy, enhancing trust and safety.

How to implement this in your domain

  1. 1Integrate gated steering: Explore incorporating Gated Activation Steering or similar ITI techniques into LLMs used for sensitive applications like medical diagnostics or patient information.
  2. 2Develop contrastive datasets: Create specific datasets of clinical questions and answers, including examples of sycophancy and hallucination, to train steering mechanisms effectively.
  3. 3Monitor model behavior: Implement real-time monitoring for LLM outputs in clinical settings to detect and log instances of sycophancy or hallucination, informing further model refinement.
  4. 4Collaborate with AI safety researchers: Partner with experts in AI safety and alignment to adapt and deploy advanced steering techniques for domain-specific LLM challenges.

Original post by Himanshu Tripathi, Subash Neupane, Shaswata Mitra, Sudip Mittal, Noorbakhsh Amiri Golilarz, Shahram Rahimi

"arXiv:2608.23666v1 Announce Type: new Abstract: Sycophancy and hallucination are persistent failure modes of Large Language Models (LLMs) across domains. However, it becomes particularly consequential in clinical question answering, where responses must remain grounded in the pro…"

View on X

Originally posted by Himanshu Tripathi, Subash Neupane, Shaswata Mitra, Sudip Mittal, Noorbakhsh Amiri Golilarz, Shahram Rahimi on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses