CAPTURE Protects LLM Agents from Memory Poisoning.

S M Asif Hossain, Ruksat Khan Shayoni, Md Kishor Morol· September 3, 2026 View original

Key takeaways

  • Personalized LLM agents are vulnerable to memory poisoning attacks.
  • CAPTURE distinguishes genuine preference drift from malicious memory poisoning.
  • It uses a neural belief tracker, multi-timescale memory, and clarification.
  • The framework improves both personalization and robustness, but highlights an adaptation-security tradeoff.

Who benefits

Customer ServicePersonal AssistantsEdTechHealthcareE-commerce

Summary

This research introduces CAPTURE, a framework for personalized LLM agents that distinguishes genuine user preference changes from adversarial memory poisoning. It uses a neural differential-equation belief tracker, multi-timescale memory, clarification, and counterfactual auditing to improve both personalization and robustness.

Personalized large language model (LLM) agents rely on persistent memory to adapt to individual user preferences over time. However, this memory also presents a vulnerability: new, conflicting information could be a genuine change in user preference, a temporary contextual shift, or a malicious attempt to poison the agent's memory. CAPTURE addresses this critical challenge by formulating it as a continuous-time partially observable decision process. The framework incorporates several innovative components: a neural differential-equation belief tracker to model latent user states, a multi-timescale memory ledger for historical context, uncertainty-triggered clarification to resolve ambiguities, and counterfactual auditing of cited memories. Evaluations show CAPTURE significantly improves win rates and effectively limits the success of memory poisoning attacks while still accepting legitimate preference updates, highlighting the inherent trade-off between adaptation and security in these advanced AI systems.

Why it matters

Developers of personalized AI assistants and LLM-powered applications must understand and implement robust security measures like CAPTURE to protect user data and maintain agent integrity against sophisticated attacks.

How to implement this in your domain

  1. 1Integrate a belief tracker and multi-timescale memory ledger into personalized LLM agent architectures.
  2. 2Develop mechanisms for uncertainty-triggered clarification to confirm user preferences.
  3. 3Implement counterfactual auditing of memories to detect potential poisoning attempts.
  4. 4Regularly evaluate the trade-off between personalization and security in agent deployments.

Original post by S M Asif Hossain, Ruksat Khan Shayoni, Md Kishor Morol

"arXiv:2609.02265v1 Announce Type: new Abstract: Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference d…"

View on X

Originally posted by S M Asif Hossain, Ruksat Khan Shayoni, Md Kishor Morol on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses