New Method Detects Unsafe LLM Outputs Using Embedding Dynamics

Mohamed Akrout, Olivera Kotevska, Dan Wilson· August 21, 2026 View original

Key takeaways

  • A new dynamical systems framework classifies unsafe LLM outputs efficiently.
  • It analyzes prompt and response embedding dynamics using Koopman models.
  • Incorporating prompt embeddings improves detection of interaction-dependent violations.
  • This black-box method enhances LLM safety and compliance.

Who benefits

TechSocial MediaCustomer ServiceHealthcareFinance

Summary

This paper extends a dynamical systems framework to classify unsafe Large Language Model (LLM) outputs by analyzing prompt and response embedding dynamics. The method fits Koopman-based predictive models for safe and unsafe regimes, using a differential residual score to detect toxic or policy-violating content in a black-box manner.

The increasing deployment of Large Language Models (LLMs) in critical applications highlights the urgent need for robust methods to detect and prevent the generation of harmful, toxic, or policy-violating content. Current detection methods often struggle with efficiency and black-box applicability. Researchers propose an extension of a dynamical systems framework, previously used for hallucination detection, to classify LLM safety. This novel approach projects both the prompt and the LLM's response into high-dimensional embedding spaces. Separate Koopman-based predictive models are then fitted for what constitutes "safe" and "unsafe" content regimes. Safety classification is achieved by comparing the prediction errors of these safe and unsafe models using a new differential residual score. A key innovation is the incorporation of prompt and response embedding dynamics, allowing the fitted Koopman operators to capture crucial interaction patterns. Evaluations across three safety benchmarks and three embedding models show that including prompt embeddings consistently improves detection, especially for violations dependent on prompt-response interaction, while response-only violations benefit from dense semantic embeddings. This work opens new avenues for analyzing AI systems using dynamical systems principles.

Why it matters

For professionals deploying or managing LLMs, this research offers a promising black-box method to enhance safety and compliance by efficiently detecting harmful content, reducing reputational and operational risks associated with AI outputs.

How to implement this in your domain

  1. 1Explore integrating dynamic embedding analysis tools into LLM safety monitoring pipelines.
  2. 2Benchmark the proposed DMD-based classification method against existing safety filters for specific LLM applications.
  3. 3Develop internal datasets of safe and unsafe prompt-response pairs to train and validate dynamic safety models.
  4. 4Collaborate with AI safety researchers to understand and apply advanced dynamical systems techniques for LLM governance.

Original post by Mohamed Akrout, Olivera Kotevska, Dan Wilson

"arXiv:2608.19579v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or policy-violating content poses significant risks. Detecting these unsafe outputs efficiently in a…"

View on X

Originally posted by Mohamed Akrout, Olivera Kotevska, Dan Wilson on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses