SLAPBench Benchmarks MLLMs for Fingerprint Verification

Bibesh Pyakurel, M. G. Sarwar Murshed· July 20, 2026 View original

Summary

SLAPBench is the first benchmark to evaluate multimodal large language models (MLLMs) for four-finger SLAP fingerprint verification, revealing that prompting strategies significantly impact performance and that proprietary models like Claude Opus 4.8 currently outperform open-source alternatives in discrimination.

Four-finger SLAP fingerprints are critical for identity verification in sensitive applications like border control and law enforcement. Despite the rise of multimodal large language models (MLLMs), their capability for this specific type of biometric verification had not been systematically assessed. To address this gap, researchers introduced SLAPBench, the first benchmark designed to evaluate MLLMs on four-finger SLAP fingerprint verification tasks. The benchmark, built using NIST SD302b data, tested several open-source MLLMs (InternVL3-8B, Qwen2.5-VL-7B, Qwen3-VL-8B, Gemma-3-12B) and the proprietary Claude Opus 4.8. A key finding was the profound impact of prompting strategies on MLLM behavior. Simple "task-description" prompts often led open-source models to collapse, resulting in near-100% False Accept Rates. However, "similarity-scoring" prompts mitigated this collapse, revealing significant capability disparities among the models. Claude Opus 4.8 demonstrated the best performance, resisting collapse and achieving high discrimination (AUC = 0.953). Among open-source models, Gemma-3-12B showed reasonable performance, while others struggled or even inverted their discrimination. The study also highlighted potential issues with the benchmark data, such as near-duplicate detection, and suggested that fairness disparities might increase as discrimination weakens. SLAPBench establishes a crucial baseline and underscores that effective prompting is paramount for MLLM performance in biometric verification.

Why it matters

For professionals in security, biometrics, and AI development, this benchmark highlights the current limitations and potential of MLLMs for critical identity verification tasks, emphasizing the importance of careful model selection and prompt engineering.

How to implement this in your domain

  1. 1When evaluating MLLMs for biometric tasks, prioritize "similarity-scoring" prompts over simple task descriptions to avoid performance collapse.
  2. 2Thoroughly benchmark MLLMs against specialized biometric systems for critical applications, as general-purpose MLLMs may not yet be robust enough.
  3. 3Investigate potential data shortcuts or biases in benchmark datasets when developing or evaluating MLLM-based biometric solutions.
  4. 4Consider proprietary MLLMs for higher-stakes biometric verification tasks, as they currently show superior discrimination.
  5. 5Develop robust prompt engineering strategies tailored to the specific nuances of biometric data and verification requirements.

Who benefits

BiometricsLaw EnforcementBorder SecurityCybersecurityAI/ML Development

Key takeaways

  • MLLMs can be applied to four-finger SLAP fingerprint verification, but performance varies widely.
  • Prompting strategy critically impacts MLLM performance, with similarity scoring outperforming task descriptions.
  • Proprietary MLLMs currently show superior discrimination capabilities compared to open-source models.
  • Benchmarks for biometrics must carefully consider data characteristics to avoid shortcuts and biases.

Original post by Bibesh Pyakurel, M. G. Sarwar Murshed

"arXiv:2607.15517v1 Announce Type: cross Abstract: Four-finger SLAP fingerprints are flat live-scan impressions of the index, middle, ring, and little fingers of one hand, used for identity verification in border control and law enforcement. No benchmark has evaluated whether mult…"

View on X

Originally posted by Bibesh Pyakurel, M. G. Sarwar Murshed on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses