OPIUM Mitigates LLM Steering Externalities and Over-Refusal

Kavin Aravindan, Arihant Rastogi, Aadi Prasad, Krishak Aneja, Saiyam Jain, Vaishnavi Shivkumar, Ponnurangam Kumaraguru· July 23, 2026 View original

Summary

OPIUM (Optimizing Protected Injections via Utility Manifolds) is a training-free method that sanitizes steering vectors for large language models. It optimizes new steering vectors to preserve desired utility while matching safer reference behaviors, improving the safety-utility trade-off and mitigating unintended side effects like over-refusal.

This research introduces OPIUM, a novel training-free method designed to address the unintended side effects of activation steering in large language models (LLMs). While activation steering offers a lightweight way to control LLMs during inference, utility vectors can inadvertently weaken safety behaviors, and refusal vectors can lead to excessive refusal on benign prompts. These are termed "steering externalities" and "over-refusal." OPIUM tackles these issues by sanitizing steering vectors through representation matching. Given reference behaviors for both desired utility and safer responses, OPIUM optimizes a new steering vector. This optimized vector is crafted to maintain the downstream representations associated with the intended intervention while simultaneously aligning with a safer reference behavior in situations where the original vector might fail. Across various steering-externality and over-refusal scenarios, OPIUM consistently improves the safety-utility trade-off compared to vanilla steering and directional ablation, suggesting that many harmful side effects can be effectively mitigated directly within the activation space.

Why it matters

Professionals deploying and managing large language models can use OPIUM to fine-tune model behavior more safely and precisely, reducing unintended biases or over-refusal without costly retraining.

How to implement this in your domain

  1. 1Evaluate your current LLM steering mechanisms for potential externalities or over-refusal issues.
  2. 2Implement OPIUM's representation matching technique to sanitize existing steering vectors.
  3. 3Define clear reference behaviors for both desired utility and safety to guide the optimization process.
  4. 4Integrate OPIUM into your LLM inference pipeline to improve the safety-utility trade-off of steered models.

Who benefits

AI DevelopmentContent ModerationCustomer ServiceLegalTechHealthcare

Key takeaways

  • Activation steering in LLMs can cause unintended side effects like over-refusal or weakened safety.
  • OPIUM is a training-free method to sanitize steering vectors.
  • It optimizes vectors to preserve utility while matching safer behaviors.
  • OPIUM improves the safety-utility trade-off in steered LLMs.

Original post by Kavin Aravindan, Arihant Rastogi, Aadi Prasad, Krishak Aneja, Saiyam Jain, Vaishnavi Shivkumar, Ponnurangam Kumaraguru

"arXiv:2607.19806v1 Announce Type: new Abstract: Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refusal vectors…"

View on X

Originally posted by Kavin Aravindan, Arihant Rastogi, Aadi Prasad, Krishak Aneja, Saiyam Jain, Vaishnavi Shivkumar, Ponnurangam Kumaraguru on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses