APTER Improves LLM Reasoning with Expert-Grounded Rubrics

Xukai Wang, Liangqi Li, Zhiyue Xu, Jingang Zhou, Xiaoyu Shi, Jiansheng Cai, Bo Zhang, Zhe Li, Xu-Yao Zhang· August 17, 2026 View original

Key takeaways

  • LLMs in professional domains need structured, expert-grounded evaluation.
  • APTER uses expert criteria to create query-level rubrics for fine-grained feedback.
  • Rubric verdicts enable targeted diagnosis and optimization of LLM capabilities.
  • The framework significantly improves LLM performance in specialized reasoning tasks.

Who benefits

HealthcareLegalFinanceEngineeringEducation

Summary

APTER is a framework that enhances large language models' performance in professional domains by integrating structured domain knowledge through expert-grounded rubrics for fine-grained evaluation, optimization, and diagnosis. It uses rubric verdicts to identify and address persistent capability deficiencies through targeted supervised fine-tuning.

As large language models (LLMs) are increasingly deployed in professional settings, they must meet specific domain constraints, incorporate critical evidence, and provide comprehensive reasoning, moving beyond merely fluent responses. Current post-training methods often rely on broad preferences or outcome-level verification. While some rubric-based approaches exist, they typically generate rubrics independently for each query, which can lead to the omission of crucial requirements and inconsistency across samples, hindering the diagnosis and targeted correction of persistent capability gaps. To overcome these limitations, researchers propose APTER (Adaptive Post-Training with Expert-Grounded Rubrics). This framework systematically integrates structured domain knowledge into a fine-grained process for evaluation, optimization, and diagnosis, particularly for complex reasoning tasks in specialized domains. APTER's expert-grounded rubric construction begins with a framework of criteria established by domain experts, where each criterion represents a stable professional capability. For each query, APTER selects relevant criteria and instantiates them into query-level rubrics, linking them back to their source criteria. This transforms reusable expert criteria into executable, query-level supervision without needing reference answers. The adaptive post-training component uses these rubric verdicts as both optimization signals and criterion-level diagnostic indicators. By aggregating low-scoring verdicts based on criterion ID, persistent deficiencies are revealed, triggering targeted supervised fine-tuning updates during reinforcement learning. Experiments in mathematical reasoning and medical question answering demonstrated consistent performance gains across both domains, with APTER improving scores by up to 15.86 and 8.04 points, respectively, over base models across three model generations.

Why it matters

Professionals can leverage APTER to develop highly specialized and reliable LLMs for critical domain-specific applications, ensuring models adhere to professional standards and provide accurate, evidence-based reasoning.

How to implement this in your domain

  1. 1Collaborate with domain experts to define a comprehensive set of stable professional criteria for your LLM's target domain.
  2. 2Develop a system to dynamically select and instantiate relevant criteria into query-level rubrics for evaluation.
  3. 3Integrate rubric-based evaluation into your LLM post-training pipeline to generate fine-grained verdicts.
  4. 4Implement a diagnostic mechanism to aggregate low-scoring verdicts by criterion, identifying persistent model deficiencies.
  5. 5Apply targeted supervised fine-tuning or reinforcement learning updates based on these diagnostic signals to improve specific capabilities.

Original post by Xukai Wang, Liangqi Li, Zhiyue Xu, Jingang Zhou, Xiaoyu Shi, Jiansheng Cai, Bo Zhang, Zhe Li, Xu-Yao Zhang

"arXiv:2608.14212v1 Announce Type: new Abstract: As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post-training methods often r…"

View on X

Originally posted by Xukai Wang, Liangqi Li, Zhiyue Xu, Jingang Zhou, Xiaoyu Shi, Jiansheng Cai, Bo Zhang, Zhe Li, Xu-Yao Zhang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses