Rubrics Enhance Open-Ended AI Generation as Privileged Information

Deepika Bablani, Ajay Gupta, Wanming Chen· August 5, 2026 View original

Key takeaways

  • Rubrics used as "soft privileged information" significantly improve open-ended AI generation.
  • This method, RuPI, outperforms rubrics-as-reward RL and hard reference completion distillation.
  • Rubrics provide a richer, dense training signal by capturing preference structure.
  • The approach is effective across Qwen and Llama models on health and science benchmarks.

Who benefits

AI DevelopmentEducationHealthcareContent CreationCustomer Service

Summary

Researchers found that using rubrics as "soft privileged information" for on-policy self-distillation significantly improves open-ended AI generation, outperforming both rubrics-as-reward reinforcement learning and distillation with hard reference completions. This method provides richer training signals, leading to better quality responses across various LLM families and benchmarks.

This research explores a novel approach to improving open-ended text generation in large language models (LLMs) by leveraging rubrics as "privileged information" (PI) during on-policy self-distillation (OPSD). Unlike traditional methods that use rubrics as scalar rewards for reinforcement learning (RL) or rely on hard reference completions, this study demonstrates that rubrics provide a much richer, dense signal when used as soft PI. The findings indicate that distilling towards soft rubric PI is more effective than distilling towards hard reference completions, as rubrics capture the preference structure across a range of valid responses rather than over-constraining the model to a single correct answer. This method, termed RuPI, significantly outperforms rubric-as-reward RL and reference-PI distillation across Qwen and Llama model families on benchmarks like HealthBench and ResearchQA, leading to higher quality, more nuanced open-ended generations.

Why it matters

This advancement offers a more effective way to guide LLMs in generating high-quality, nuanced, and contextually appropriate open-ended responses, which is critical for applications requiring creativity, complex problem-solving, or adherence to specific quality criteria.

How to implement this in your domain

  1. 1Develop detailed rubrics for evaluating open-ended generation tasks to serve as privileged information for model training.
  2. 2Implement on-policy self-distillation (OPSD) techniques using these rubrics to fine-tune large language models for specific applications.
  3. 3Compare the performance of rubric-as-PI distillation against traditional RLHF methods and reference-based distillation for open-ended tasks.
  4. 4Apply this method to improve content generation, customer service responses, or educational feedback systems where quality and nuance are paramount.

Original post by Deepika Bablani, Ajay Gupta, Wanming Chen

"arXiv:2608.02948v1 Announce Type: new Abstract: On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in verifiable domains like math, where hard privileged information (PI) in the form of ground-truth…"

View on X

Originally posted by Deepika Bablani, Ajay Gupta, Wanming Chen on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses