Interpretable Anomaly Detection for Collider Physics Using Contrastive Learning

Haoyi Jia, Sagar Addepalli, Julia Gonski· August 17, 2026 View original

Key takeaways

  • Anomaly detection in complex systems often lacks interpretability.
  • ORCA uses contrastive learning to create an interpretable embedding space.
  • This framework improves sensitivity to anomalies and allows for their attribution.
  • Interpretable anomaly detection is crucial for scientific discovery and decision-making.

Who benefits

High-Energy PhysicsCybersecurityManufacturingHealthcareFinance

Summary

Researchers developed ORCA, a two-stage framework that uses supervised contrastive learning to create an interpretable embedding space for anomaly detection in collider physics. This method significantly improves sensitivity to new physics signals and allows for attributing anomalous events to known processes or characterizing unknown ones.

Anomaly detection in collider physics faces two significant challenges: interpreting the meaning of anomaly scores and preventing these scores from correlating with trivial factors like energy scale or object multiplicity. A new framework, Organized Representation via Contrastive learning for Anomaly detection (ORCA), addresses these issues by introducing a two-stage process. First, ORCA employs supervised contrastive learning to develop a highly organized embedding space. This space is learned across a diverse range of known physics processes, ensuring that distinct processes occupy separate, identifiable regions. In the second stage, a standard autoencoder is applied within this learned embedding space to generate event-level anomaly scores. This approach helps to disentangle true anomalies from background variations. Simulations mimicking conditions at the High-Luminosity Large Hadron Collider demonstrated that ORCA substantially improves both the breadth and depth of sensitivity to new physics signals compared to a baseline autoencoder. Beyond enhanced detection, the contrastive embedding makes the anomalous samples interpretable. By fitting templates of known processes to the embedding distributions, ORCA can attribute anomalous events to specific physics processes with quantified uncertainties, even recovering signals not included in the embedding's training. This capability establishes ORCA as a promising route for interpretable anomaly detection in collider experiments, providing higher-dimensional physics information for downstream statistical analysis.

Why it matters

For professionals in fields requiring robust anomaly detection and clear interpretability, ORCA offers a powerful framework. Its ability to not only detect anomalies but also explain their potential origins or characteristics is crucial for decision-making and scientific discovery, especially in high-stakes environments.

How to implement this in your domain

  1. 1Identify domains where anomaly detection is critical but interpretability is lacking.
  2. 2Explore applying supervised contrastive learning to create structured embedding spaces for your data.
  3. 3Integrate autoencoders or similar anomaly scoring models within the learned embedding space.
  4. 4Develop template fitting or clustering methods to interpret detected anomalies.
  5. 5Validate the framework's sensitivity and interpretability using simulated or real-world datasets.

Original post by Haoyi Jia, Sagar Addepalli, Julia Gonski

"arXiv:2608.13652v1 Announce Type: new Abstract: Generic event-level anomaly detection for collider physics has two recurring problems: anomaly scores are hard to interpret, and they correlate strongly with energy scale and object multiplicity. We present Organized Representation…"

View on X

Originally posted by Haoyi Jia, Sagar Addepalli, Julia Gonski on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses