As cited
Copy frozen at (site build).
ai security
LLMs and Contextual Integrity
Two new papers examine how large language models handle sensitive information stored in persistent memory across different task contexts. The first paper introduces CIMemories, a benchmark showing that frontier models leak inappropriate information in up to 69% of cases, with violations accumulating as usage increases. The second paper demonstrates that explicit reasoning and reinforcement learning can substantially reduce information disclosure while preserving task performance.
Why it matters: Organizations deploying LLMs with memory features, assistants, and autonomous agents need to understand that current models fail to properly control information flow across contexts, creating privacy and data protection risks that neither prompting nor model scaling alone can resolve.
- Source published
- First seen by Cybersecurity Tracker