Multi-Agent AI Extracts Oncology Data with High Accuracy

Daniel Kang, Michelle Hu, Soorya Ram Shimgekar, Shayan Vassef, Yufan Wang, Anit Kumar Sahu, Munmun De Choudhury, Vedant Das Swain, Christian Poellabauer, Li Yan Khor, Koustuv Saha, Robert Wojciechowski, Elliot Kidd, Piyum Zonooz, Navin Kumar· September 1, 2026 View original

Key takeaways

  • nMAS is a multi-agent AI system for extracting oncology data from fragmented documents.
  • It extracts 328 clinician-defined attributes with high precision and recall.
  • The system significantly outperforms manual abstraction and comparator models.
  • It offers a scalable solution for converting unstructured clinical notes into structured data.

Who benefits

HealthcarePharmaceuticalsMedical ResearchHealthTech

Summary

The Nimblemind Multi-Agent System (nMAS) is a configurable AI workflow that extracts 328 clinically relevant oncology attributes from fragmented documentation. It achieved 85.0% F1 score, significantly outperforming a comparator, demonstrating feasibility for converting unstructured clinical notes into structured data.

Clinically relevant oncology information is often scattered across various heterogeneous and longitudinal documents, making manual abstraction a time-consuming and burdensome task. This manual process can take nearly half an hour per case, highlighting a critical need for scalable methods that can accurately extract structured data while preserving clinical context. Researchers evaluated the Nimblemind Multi-Agent System (nMAS), a configurable workflow designed for oncology information extraction. nMAS is built to extract 328 specific attributes, defined by clinicians, covering report metadata, diagnosis, staging, and cancer-type-specific details from fragmented oncology documentation. The system's design separates clinician-defined field specifications from the model's execution, combining complexity-aware extraction, report-level consolidation, and source-grounded validation. In a retrospective evaluation involving 230 de-identified oncology documents from 40 patients, nMAS achieved a rank-weighted value-level precision of 82.6%, recall of 87.5%, and an F1 score of 85.0%. This performance significantly surpassed an independently implemented comparator model, which achieved an F1 of 66.4%. These findings strongly support the viability of using a configurable, source-grounded extraction workflow to transform complex, fragmented oncology documentation into reusable structured data, which can then be used for analytics or presented in tumor boards.

Why it matters

This system dramatically reduces the manual burden of extracting critical oncology data, enabling faster, more accurate insights for research, clinical decision-making, and cancer registries.

How to implement this in your domain

  1. 1Pilot AI for data extraction: Explore implementing multi-agent AI systems for extracting structured data from unstructured clinical notes in specific medical domains.
  2. 2Define clear schemas: Collaborate with domain experts to define comprehensive and precise schemas for the data to be extracted, ensuring clinical relevance.
  3. 3Integrate validation steps: Design workflows that include source-grounded validation and clinician review to maintain accuracy and auditability of extracted data.
  4. 4Leverage structured data: Utilize the newly structured oncology data to enhance tumor boards, clinical trials, and population health analytics.

Original post by Daniel Kang, Michelle Hu, Soorya Ram Shimgekar, Shayan Vassef, Yufan Wang, Anit Kumar Sahu, Munmun De Choudhury, Vedant Das Swain, Christian Poellabauer, Li Yan Khor, Koustuv Saha, Robert Wojciechowski, Elliot Kidd, Piyum Zonooz, Navin Kumar

"arXiv:2608.28974v1 Announce Type: new Abstract: Clinically relevant oncology information is distributed across heterogeneous, longitudinal documentation, creating substantial abstraction burden and requiring accurate attribution across specimens, tumors, biomarkers, and time poin…"

View on X

Originally posted by Daniel Kang, Michelle Hu, Soorya Ram Shimgekar, Shayan Vassef, Yufan Wang, Anit Kumar Sahu, Munmun De Choudhury, Vedant Das Swain, Christian Poellabauer, Li Yan Khor, Koustuv Saha, Robert Wojciechowski, Elliot Kidd, Piyum Zonooz, Navin Kumar on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses