LLMs Benchmark Against Official Crash Data Coding

Sudhir Bharati, Rajendra K C Khatri, Sudip Bharati· August 3, 2026 View original

Key takeaways

  • Frontier LLMs show promise but also significant limitations in accurately coding crash data from narratives.
  • Performance varies greatly depending on the specific data attribute being extracted.
  • Simple baseline methods can sometimes outperform advanced LLMs for certain tasks.
  • Attribute-specific evaluation and human review are crucial for reliable LLM deployment in critical data extraction.

Who benefits

Public SafetyInsuranceAutomotiveGovernment

Summary

A study benchmarked six frontier LLMs against official crash databases using police narratives, finding that while LLMs can assist, their performance varies significantly by attribute and often falls short of simple baselines. GPT-5.5 High performed best among LLMs, but not always superior to basic rules.

This research evaluates the capability of six advanced large language models (LLMs) to extract and code crash attributes from police narratives, comparing their output against official, structured crash databases. The study linked over 5,500 fatal crash narratives from Arkansas with nearly 6,000 structured records, focusing on attributes like crash manner, intersection type, and light conditions.The LLMs were tested using a zero-shot prompt, and their performance was assessed using various metrics including agreement, F1 score, and Cohen's kappa. Results showed that GPT-5.5 High achieved the highest agreement among the LLMs. However, simple baselines, such as always-majority or keyword-rule approaches, often matched or even surpassed the LLMs in raw agreement or F1 scores.The study highlighted that performance varied more significantly across different crash attributes than across the LLM models themselves. Attributes like non-motorist relation and crash manner were coded with higher agreement, while light condition and roadway surface condition proved more challenging. This suggests that LLM deployment for such tasks requires attribute-specific evaluation and human oversight.

Why it matters

Professionals in data analysis, public safety, and insurance can understand the current limitations and potential of LLMs for extracting structured information from unstructured text, guiding realistic expectations and deployment strategies.

How to implement this in your domain

  1. 1Conduct pilot projects to evaluate LLM performance on specific data extraction tasks relevant to your domain, using clear baselines.
  2. 2Develop hybrid systems that combine LLM capabilities with rule-based or human-in-the-loop validation for critical data points.
  3. 3Prioritize LLM deployment for attributes where performance is demonstrably high and less prone to error.
  4. 4Establish transparent evaluation protocols and human review processes for any LLM-generated data to ensure accuracy and reliability.

Original post by Sudhir Bharati, Rajendra K C Khatri, Sudip Bharati

"arXiv:2607.29064v1 Announce Type: new Abstract: Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclear how well large language models (LLMs) reproduce official crash coding. This stud…"

View on X

Originally posted by Sudhir Bharati, Rajendra K C Khatri, Sudip Bharati on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses