Incomplete Prompts Jailbreak LLMs, Delaying Refusal

Yeonjea Kim, Bumjin Park, Jaesik Choi· July 24, 2026 View original

Summary

New research formalizes "incomplete prompt jailbreaks" (IPJ), where LLMs delay refusing harmful requests until sentence termination, making them vulnerable to partial harmful prompts. Training models to refuse these prompts is insufficient, but neuron-level interventions targeting "termination" and "continuation" neurons show promise for robust defenses.

A new study identifies and formalizes "incomplete prompt jailbreaks" (IPJ), a vulnerability in large language models where safeguards against harmful content can be bypassed by providing incomplete prompts. The research shows that LLMs systematically delay their refusal to generate harmful content until the sentence is fully terminated, making them susceptible to generating undesirable continuations from partial, harmful inputs. Empirical analysis reveals diverse "attractor types" associated with these incomplete continuations. Attempts to mitigate this through parameter tuning proved insufficient, as the models failed to generalize refusal across different content domains and attractor types. However, the study pinpointed specific "termination" and "continuation" neurons, suggesting that targeted interventions at this neuron level could offer a more precise and robust defense against IPJ.

Why it matters

For professionals deploying LLMs, understanding and mitigating incomplete prompt jailbreaks is crucial for maintaining model safety, preventing the generation of harmful content, and ensuring responsible AI use.

How to implement this in your domain

  1. 1Review current LLM safety protocols to identify potential vulnerabilities to incomplete prompt jailbreaks.
  2. 2Investigate the feasibility of implementing neuron-level interventions for enhanced safety in deployed LLMs.
  3. 3Develop robust input validation and sanitization techniques that account for partial or incomplete prompts.
  4. 4Stay updated on research into advanced jailbreak detection and prevention mechanisms.

Who benefits

AI DevelopmentCybersecurityContent ModerationSocial MediaPublic Sector

Key takeaways

  • LLMs are vulnerable to "incomplete prompt jailbreaks" where harmful content is generated from partial inputs.
  • Models often delay refusal until the prompt is fully terminated.
  • Standard fine-tuning for refusal is insufficient and lacks generalization.
  • Neuron-level interventions show promise for more robust and precise defenses.

Original post by Yeonjea Kim, Bumjin Park, Jaesik Choi

"arXiv:2607.20473v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion remains vulnerable to incomplete harmful prompts. In this work, we formalize th…"

View on X

Originally posted by Yeonjea Kim, Bumjin Park, Jaesik Choi on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses