CHORUS Boosts LLM Code Generation for Hardware Verification

Hejia Zhang, Sheng Lu, Zhongming Yu, Chia-Tung Ho, Brucek Khailany, Jishen Zhao· August 12, 2026 View original

Key takeaways

  • CHORUS significantly improves LLM performance for hardware testbench stimulus generation.
  • It leverages complementary strengths from diverse SFT checkpoints and RL.
  • The framework consolidates specialists into a single, highly effective model.
  • CHORUS achieves substantial performance gains over much larger LLMs in hardware verification.

Who benefits

SemiconductorElectronics ManufacturingHardware EngineeringAutomotiveAerospace

Summary

Researchers introduce CHORUS, a post-training framework that significantly improves Large Language Model (LLM) performance in generating high-coverage testbench stimuli for hardware verification. It leverages complementary strengths of diverse SFT checkpoints and dense-reward RL to create a single, highly effective 4B model.

A new post-training framework called CHORUS has been developed to significantly enhance Large Language Model (LLM) performance in hardware verification, specifically for generating high-coverage testbench stimuli. This task is crucial in modern chip design, accounting for a substantial portion of the effort. CHORUS pushes performance beyond what typical supervised fine-tuning (SFT) to reinforcement learning (RL) pipelines achieve. The framework is built on two key observations. First, staged SFT produces behaviorally diverse checkpoints, which, when subjected to dense-reward RL, become strong experts with comparable overall performance but distinct strengths at the task level. Second, these complementary strengths can be effectively exploited through either training-free model merging or further post-training. By consolidating these specialized experts into a single 4B model, CHORUS achieved an 88.0% Pass@1 score on CVDP-ECov. This represents a substantial improvement of 13.5 percentage points over DeepSeek-R1 (671B), demonstrating its superior capability in generating effective testbench stimuli for hardware verification.

Why it matters

For professionals in hardware design and verification, CHORUS offers a breakthrough in automating a highly labor-intensive and critical task. It promises to accelerate chip development cycles, reduce verification costs, and improve the reliability of complex hardware designs by generating more comprehensive test cases.

How to implement this in your domain

  1. 1Investigate integrating CHORUS-like frameworks into hardware verification pipelines to automate testbench stimulus generation.
  2. 2Explore techniques for combining diverse AI models or checkpoints to leverage their complementary strengths for complex tasks.
  3. 3Pilot advanced LLM-based code generation tools for specialized engineering domains beyond general-purpose coding.
  4. 4Train hardware verification engineers on how to effectively use and validate AI-generated test stimuli.
  5. 5Assess the potential for AI-driven verification to reduce design cycles and improve chip quality.

Original post by Hejia Zhang, Sheng Lu, Zhongming Yu, Chia-Tung Ho, Brucek Khailany, Jishen Zhao

"arXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. Hardware verification is an important application of code generation and ac…"

View on X

Originally posted by Hejia Zhang, Sheng Lu, Zhongming Yu, Chia-Tung Ho, Brucek Khailany, Jishen Zhao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses