Self-Distillation Improves LLM In-Context Watermarking Reliability

Yepeng Liu, Tianyi Chen, Xuandong Zhao, Dawn Song, Yuheng Bu· September 1, 2026 View original

Key takeaways

  • Current LLMs struggle to reliably follow in-context watermarking instructions while maintaining response quality.
  • A new two-stage self-distillation method significantly improves LLM adherence to ICW instructions.
  • This method requires no external teacher model or manual annotation, making it highly practical.
  • Reliable ICW is vital for content provenance, IP protection, and combating misinformation from LLMs.

Who benefits

Media & PublishingCybersecurityContent CreationAI DevelopmentLegal & Compliance

Summary

A new two-stage self-distillation method enables LLMs to reliably follow in-context watermarking instructions without degrading response quality, addressing a critical gap in current models. This technique, which requires no external teacher or manual annotation, significantly boosts watermarking detectability across various LLMs and instruction families.

In-context watermarking (ICW) offers a method for third parties to embed a statistically detectable signal into an LLM's response by simply prepending an instruction to a query, without needing access to the model's internal workings. However, the reliability of this technique hinges on the LLM's ability to consistently follow these instructions while maintaining high answer quality, a capability that current frontier models often lack. To address this, researchers introduced \mathsf{ICWBench}, a benchmark designed to measure both detectability and answer quality across various ICW instruction families. Their evaluation of 14 leading proprietary and open-source LLMs revealed that none could consistently achieve both objectives across all instruction types. To overcome this limitation, a novel two-stage training method was developed, utilizing self-distillation. The first stage, self-distillation with logits perturbation (SDLP), uses the base LLM itself as both teacher and student, where the teacher is guided by an instruction-equivalent decoding-time perturbation. The student then learns to match the teacher's output distribution. The second stage employs reinforcement learning with an automatic verifier providing rewards. This method significantly improved watermarking detectability (from 0.100 to 0.974 for Qwen3-14B and 0.337 to 0.968 for GPT-OSS-20B) while preserving response quality.

Why it matters

Reliable in-context watermarking is crucial for content provenance, intellectual property protection, and combating misinformation generated by LLMs. This breakthrough enables broader adoption of watermarking without requiring model retraining or internal access, enhancing trust and accountability in AI-generated content.

How to implement this in your domain

  1. 1Integrate self-distillation techniques into LLM training pipelines to enhance instruction following for security and provenance features.
  2. 2Adopt in-context watermarking as a standard practice for identifying AI-generated content in critical applications.
  3. 3Develop tools and APIs that leverage reliable ICW for content verification and intellectual property protection.
  4. 4Educate content creators and consumers about the existence and implications of AI watermarking for digital trust.

Original post by Yepeng Liu, Tianyi Chen, Xuandong Zhao, Dawn Song, Yuheng Bu

"arXiv:2608.29030v1 Announce Type: new Abstract: In-context watermarking (ICW) prepends an instruction to a query asking the model to embed a statistically detectable signal in its response. It thus equips LLMs with a watermarking interface that third parties can invoke without ac…"

View on X

Originally posted by Yepeng Liu, Tianyi Chen, Xuandong Zhao, Dawn Song, Yuheng Bu on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses