LLM-as-a-Judge Framework Validated for Agentic AI in Drug Discovery

Emma Granqvist, Roc\'io Mercado, Samuel Genheden· August 24, 2026 View original

Key takeaways

  • A new LLM-as-a-Judge framework evaluates agentic AI in drug discovery.
  • It defines four key output quality dimensions and includes tool call correctness checks.
  • Human expert alignment studies are crucial for validating and optimizing LLM judges.
  • The framework offers a scalable and reliable template for scientific AI agent evaluation.

Who benefits

PharmaceuticalsBiotechnologyHealthcareAI/ML DevelopmentResearch

Summary

Researchers developed and validated an LLM-as-a-Judge evaluation framework for agentic AI systems in drug discovery, specifically for AstraZeneca's ChatInvent. The framework defines four output quality dimensions and uses human expert alignment studies to optimize LLM judges, demonstrating improved agreement with human assessments.

Evaluating the open-ended outputs of agentic AI systems, particularly in complex scientific fields like drug discovery, presents a significant challenge. Traditional metrics often fail to capture semantic correctness, and human expert evaluation is not scalable. A new study introduces an "LLM-as-a-Judge" evaluation framework designed for AstraZeneca's agentic drug discovery assistant, ChatInvent, aiming to bridge this gap. The framework defines four key dimensions for evaluating output quality: Completeness, Relevancy, Structural Clarity, and Scope Adherence, alongside deterministic checks for tool call correctness. A crucial aspect of this work involved validating the LLM judge's alignment with human experts. Five expert annotators compared various LLMs (Gemini 3.1 Pro, Claude Opus 4.7, GPT-5, Llama 3.1 70B) as judges. The best-performing LLM judge was further optimized using few-shot demonstrations of human-annotated examples, which improved its alignment with the human majority vote from 0.80 to 0.86. Applying this optimized judge to held-out questions revealed specific limitations and surprisingly, found that informal phrasings did not degrade output quality, and even suggested that having the LLM rewrite the original question could be beneficial. This framework provides a robust, human-aligned template for evaluating agentic systems in scientific domains.

Why it matters

This framework provides a scalable and human-aligned method for evaluating complex AI agent outputs, crucial for accelerating development and ensuring reliability in high-stakes scientific applications like drug discovery.

How to implement this in your domain

  1. 1Adopt the proposed four output quality dimensions (Completeness, Relevancy, Structural Clarity, Scope Adherence) for evaluating your own agentic AI systems.
  2. 2Conduct a human alignment study with domain experts to validate and select the most suitable LLM judge for your specific application.
  3. 3Optimize your chosen LLM judge using few-shot examples derived from human-annotated data to improve its agreement with expert assessments.
  4. 4Integrate the validated LLM-as-a-Judge framework into your continuous integration/continuous deployment (CI/CD) pipeline for automated agent evaluation.

Original post by Emma Granqvist, Roc\'io Mercado, Samuel Genheden

"arXiv:2608.21057v1 Announce Type: new Abstract: Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. Reference-based metrics such as…"

View on X

Originally posted by Emma Granqvist, Roc\'io Mercado, Samuel Genheden on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools