EMRB Benchmarks LLM Reasoning on Raw Electromagnetic Signals.

Mingxu Zhang, Ying Sun, Yuhan Li, Yang Ji, Dazhong Shen, Ke Zhang, Shan Huang· August 26, 2026 View original

Key takeaways

  • LLMs can be evaluated for raw electromagnetic signal analysis using EMRB.
  • Current LLMs struggle with complex system design tasks from raw data.
  • The ReconPilot method significantly improves LLM reasoning over raw signals.
  • This opens new possibilities for LLMs in physical-layer engineering and science.

Who benefits

TelecommunicationsAerospace & DefenseIoTAutomotiveScientific Research

Summary

This paper introduces EMRB, a multi-level benchmark designed to evaluate Large Language Models' ability to analyze raw I/Q electromagnetic signal data by writing and executing code. It reveals significant performance gaps across LLMs, particularly in complex system design tasks, and proposes ReconPilot, a structured method to improve reasoning.

While Large Language Models (LLMs) are increasingly used as code agents for scientific and engineering analysis, their capability to process and interpret raw physical-layer measurements, such as electromagnetic signals, has remained largely untested. The "EMRB" (Electromagnetic Reasoning Benchmark) addresses this gap by providing 200 problems across five difficulty levels and 27 question types, ranging from signal detection to OFDM design. Unlike benchmarks that rely on preprocessed features, EMRB presents only raw I/Q data, requiring LLMs to discover relevant quantities through code generation and execution. The evaluation of 14 diverse LLMs on EMRB showed a wide range of scores, from 24.1% to 78.9%, with a sharp decline in performance on more complex tasks like system design compared to basic measurements. To enhance LLM performance, the researchers propose "ReconPilot," a structured method that separates signal reconnaissance, targeted analysis, and self-verification steps. ReconPilot significantly improved overall scores across various LLM backbones, demonstrating its effectiveness in boosting reasoning capabilities for raw signal analysis.

Why it matters

Professionals in fields dealing with raw sensor data, signal processing, or complex engineering analysis can leverage LLMs more effectively if these models can directly interpret raw measurements. EMRB and ReconPilot offer a path to developing and evaluating such capabilities.

How to implement this in your domain

  1. 1Explore the EMRB benchmark to assess the raw signal analysis capabilities of LLMs.
  2. 2Integrate the ReconPilot structured reasoning method into LLM-based analysis workflows.
  3. 3Develop custom code agents that can interpret and process raw I/Q data using LLMs.
  4. 4Fine-tune LLMs on domain-specific raw signal datasets to improve performance.
  5. 5Collaborate with AI researchers to advance LLM capabilities in physical-layer analysis.

Original post by Mingxu Zhang, Ying Sun, Yuhan Li, Yang Ji, Dazhong Shen, Ke Zhang, Shan Huang

"arXiv:2608.24086v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as code agents for scientific and engineering analysis, but their ability to analyze raw physical-layer measurements remains untested. We introduce \textbf{EMRB} (\textbf{E}lectro\t…"

View on X

Originally posted by Mingxu Zhang, Ying Sun, Yuhan Li, Yang Ji, Dazhong Shen, Ke Zhang, Shan Huang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses