Semalith v1.4: Compact Safety Classifier Excels in Prompt Injection.

Tejasvi C. Addagada· July 28, 2026 View original

Summary

Semalith v1.4 is a 184M-parameter DeBERTa-v3-base classifier that performs simultaneous three-axis safety classification, including prompt injection, general harm, and financial-services regulatory compliance, in a single pass. It outperforms Llama-Guard-3-8B on prompt injection benchmarks with 44x fewer parameters and achieves zero false positives on benign agentic prompts.

Deploying large language models (LLMs) in sensitive environments like financial services or agentic systems necessitates robust safety classifiers capable of handling prompt injection, regulatory compliance, and general harm simultaneously. Existing open-source guardrails typically do not offer this comprehensive, single-inference pass capability. Introducing Semalith v1.4, a compact 184-million-parameter DeBERTa-v3-base classifier designed to address this gap. This model performs a three-axis safety classification in one forward pass, covering prompt injection, general harm, and financial-services regulatory compliance. Its specialized 22-class output head, along with a 4-class auxiliary super-category head, was trained on a meticulously curated corpus of over 76,000 rows. In comparative evaluations against Llama-Guard-3-8B, Semalith v1.4 demonstrated superior performance in prompt injection detection, winning all seven benchmarks while using 44 times fewer parameters. It also achieved a perfect zero false positive rate on 208 benign agentic prompts, significantly outperforming Llama-Guard-3-8B's 0.063 FPR. While Llama-Guard-3 still leads on general harm benchmarks, Semalith v1.4 offers a compelling, efficient solution for specific high-stakes applications, particularly where BFSI compliance and prompt injection resistance are paramount.

Why it matters

For professionals deploying LLMs in production, especially in regulated or sensitive industries, a highly efficient and accurate safety classifier like Semalith v1.4 is crucial for mitigating risks such as prompt injection and ensuring compliance without excessive computational overhead.

How to implement this in your domain

  1. 1Assess your current LLM deployment's vulnerability to prompt injection and other safety risks.
  2. 2Evaluate Semalith v1.4 as a potential lightweight guardrail solution for your LLM applications.
  3. 3Integrate Semalith v1.4 into your LLM inference pipeline for real-time safety classification.
  4. 4Customize the safety classification rules to align with your organization's specific regulatory and compliance requirements.
  5. 5Monitor its performance in production, particularly for prompt injection detection and false positive rates on benign inputs.

Who benefits

BFSICybersecurityLegalAI/TechGovernment

Key takeaways

  • Semalith v1.4 is a compact (184M parameters) safety classifier for LLMs.
  • It simultaneously detects prompt injection, general harm, and financial compliance issues.
  • Semalith v1.4 outperforms Llama-Guard-3-8B on prompt injection with significantly fewer parameters.
  • It achieves zero false positives on benign agentic prompts, crucial for sensitive applications.

Original post by Tejasvi C. Addagada

"arXiv:2607.22545v1 Announce Type: new Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail ad…"

View on X

Originally posted by Tejasvi C. Addagada on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses