Detecting LLM Deception: Lie Typology Impacts Performance

Amr Moustafa, Max Feser, Florian Mai· July 24, 2026 View original

Summary

This study systematically investigates factors influencing the detection of deceptive outputs from large language models, revealing that the type of lie in training data significantly impacts probe performance. It shows that probes trained on one lie type often fail on others, highlighting that deception detection is highly representation-dependent.

Detecting deceptive outputs from large language models (LLMs) remains a complex challenge, with previous research indicating that detection probes often fail when encountering out-of-domain scenarios. This means a probe trained on one type of lie may not effectively identify other forms of deception. This new work conducts a systematic analysis of various factors affecting detection performance. The study examines representation depth, probe expressivity, sparse feature representations, and crucially, the typology of lies within the training data. By augmenting standard benchmarks with a dataset containing diverse deception types—including fabrication, omission, and exaggeration—the researchers found that the choice of training data and the specific lie typology profoundly influence detectability. Optimal representation depth is dataset-dependent, more expressive probes offer only selective advantages, and sparse autoencoder features perform similarly to dense hidden states. Ultimately, the research underscores that effective deception detection is fundamentally tied to the specific data representations used.

Why it matters

Professionals building or deploying LLMs for sensitive applications (e.g., content moderation, legal tech, financial analysis) need to understand that current deception detection methods are highly sensitive to the type of lie, requiring diverse training data and careful validation.

How to implement this in your domain

  1. 1Review the types of deceptive content your LLM applications might encounter.
  2. 2Diversify training datasets for deception detection probes to include various lie typologies (fabrication, omission, exaggeration).
  3. 3Benchmark your LLM's deception detection capabilities across different lie types and representation depths.
  4. 4Implement human-in-the-loop verification for high-stakes scenarios where LLM deception detection is critical.

Who benefits

CybersecuritySocial MediaLegalTechFinanceContent Moderation

Key takeaways

  • Detecting LLM deception is challenging, especially out-of-domain.
  • Lie typology in training data significantly impacts detection probe performance.
  • Probes trained on one lie type often fail on others.
  • Effective deception detection is highly dependent on data representation.

Original post by Amr Moustafa, Max Feser, Florian Mai

"arXiv:2607.20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in out-of-domain scenarios -- training on one type of lie does not t…"

View on X

Originally posted by Amr Moustafa, Max Feser, Florian Mai on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses