New Benchmark Reveals AI Exam Scores Inflated by Standard Grading

Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh· August 20, 2026 View original

Key takeaways

  • Standard AI evaluation metrics can inflate scores on exams with non-additive grading.
  • Convex grading schemes penalize partial knowledge, impacting AI model rankings.
  • A model's error distribution, not just accuracy, affects its score under complex rubrics.
  • Accurate evaluation requires benchmarks that precisely replicate real-world scoring rules.

Who benefits

EdTechAI DevelopmentAssessment & CertificationGovernment

Summary

A new benchmark, THPT-Ladder, exposes how standard accuracy metrics inflate AI language model scores on exams with non-additive grading schemes, specifically Vietnam's 2025 National High School Graduation Examination. It shows that models can drop significantly in percentile rank when graded by the official convex rubric, which penalizes partial knowledge.

Traditional benchmarks for evaluating language models on human exams often simplify scoring by treating each response as either entirely right or entirely wrong. This approach assumes that partial knowledge is proportionally rewarded, which is not always true, especially with non-additive grading schemes. A new study introduces THPT-Ladder, a benchmark designed to accurately reflect Vietnam's 2025 National High School Graduation Examination grading. This exam uses a convex scoring system for certain sections, meaning correctly identifying three out of four true/false statements might earn 0.50 points, not the 0.75 points a proportional accuracy metric would suggest. This discrepancy can significantly inflate a model's apparent score. Evaluating eight different models, the research found that the official rubric consistently resulted in lower scores (0.020 to 0.159 points less per question) compared to proportional credit. For example, Qwen3.5-27B's standing on the History exam dropped from the 90th to the 77th percentile when graded by the official scheme. The study emphasizes that a model's overall accuracy does not reliably predict this penalty, as the distribution of errors greatly influences the final score under a convex system.

Why it matters

For professionals evaluating AI models, especially in high-stakes applications like education or certification, understanding how grading schemes impact reported performance is crucial to avoid overestimating AI capabilities.

How to implement this in your domain

  1. 1Scrutinize the grading rubrics of any benchmarks used to evaluate AI models, particularly for non-linear or convex scoring.
  2. 2Develop custom evaluation metrics that precisely mirror real-world grading schemes for critical applications.
  3. 3Conduct sensitivity analyses on AI model performance under different error distributions, not just overall accuracy.
  4. 4Advocate for transparent and detailed reporting of AI evaluation methodologies, including specific scoring rules.

Original post by Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh

"arXiv:2608.18336v1 Announce Type: new Abstract: When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption tha…"

View on X

Originally posted by Nguyen Quoc Hung, Nguyen Dang Minh, Le Nhu Quynh, Tran Khanh Linh, Nguyen Kieu Linh on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses