One simple text file is enough to take over a coding agent’s memory. Cisco researchers checked a live coding tool and found it loaded the first 200 lines of its memory files straight into the system prompt [3].
One bad write stayed active across every project, session, and reboot. Attackers do not need to break the model, because they only need to write to its database.
The big picture:
Memory makes an agent useful, but it skips the prompt filters guarding the model.
A filter at the model reads each input once. A memory write takes two steps. The agent stores text now, and reads it back later as trusted context.
The Open Worldwide Application Security Project (OWASP) flagged this risk in 2026. It added Memory and Context Poisoning as entry ASI06 in its Top 10 for Agentic Applications [5]. The pattern is simple: bad text arrives once, and the system keeps trusting it.
My read: this is a missing check at the storage layer, not a weak model. One study notes these attacks work because of “the lack of security-focused memory governance” [1].
By the numbers
- 98% — Memory trap rate: One study planted bad code in the memory store of top agents in about 98% of tests, and the stored entry fired later in roughly 60% of them [1].
- 95% — Query-only attack: A second study reached over 95% storage and 70% attack success using plain user prompts, needing no special access [2].
- 200 lines — Memory read as orders: A live coding agent loaded the first 200 lines of its memory files straight into the system prompt, turning one write into a trusted order [3].
- 50 prompts — Found in the wild: Microsoft logged over 50 memory-poisoning traps aimed at 31 companies across 14 fields in just 60 days [4].
What I’d watch:
What strikes me here is that every real fix sits at the data store, not the prompt.
One study built a defense called Agentic Memory Sentry. It uses a rule for what the agent can save and a screen for what it reads back [1]. Both act on the memory layer, which is the exact spot a prompt filter cannot see.
System builders are moving in three ways to lock this down:
- A checked write path: Every memory save passes a check first. Systems block or set aside injected orders, secrets, and bad edits. Since the write is the exact moment trust is given, it is the cheapest place to reject bad data.
- A read screen: Text pulled from memory gets checked again before it enters the context window. This is the control that would have caught the coding-agent flaw [3].
- Per-user limits: Memory locked to a single user stops a poisoned entry from spilling into other people’s sessions. This shrinks the damage zone when a bad write slips through.
The catch
These high attack rates come from lab tests run under tight rules [2]. Read them as a ceiling on how easy the attack is, not as the rate to expect on live traffic.
Every defense carries a cost. A write check that is too strict throws away useful memory, which creates a bad user experience. A read screen adds delay to a path that runs on every single turn.
The part I keep circling: a record the agent reads again as truth will always be worth attacking. The researchers who found the coding-agent flaw closed one specific path [3], but the wider question remains open. Leaving the write path open and hoping the model acts safely is a weak design.
At a glance
- The Big Shift: An agent’s memory is now a main attack surface. A poisoned write is read back later as trusted context, meaning the flaw lives in the database rather than the prompt.
- Why It Matters: A prompt filter reads an input exactly once, but memory is read again on every future turn. That gap allowed labs to hit a 98% trap rate, and a live campaign hit 31 companies in 60 days.
- What I’d Watch: How system builders secure the memory store like a normal database. The tested controls all sit at the storage layer:
- A checked write path: A check on every memory save that blocks injected text before the system stores it.
- A read screen: A second filter on data leaving memory before it reaches the context window.
- Per-user limits: Memory locked to one user so a poisoned entry cannot spread into other sessions.
- The Catch: The highest attack rates come from controlled lab tests, and every defense adds friction. Strict write checks throw away useful context, and read screens add delay to every turn.
Related reading
- MCP Skills Extension — more on Autonomous & Agentic Workflows
- Planner-as-Router: Fold the Model Choice Into the Plan — more on Autonomous & Agentic Workflows
- Stop Piling Memory Onto AI Agents — more on Autonomous & Agentic Workflows
Sources
[1] George Torres, Sharad Shrestha and Satyajayant Misra, “When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents” (arXiv:2607.06595, submitted 2026-07-06) — the GhostWriter attack (98% injection, ~60% activation) and the Agentic Memory Sentry defense (https://arxiv.org/abs/2607.06595) [2] “Memory Poisoning Attack and Defense on Memory Based LLM-Agents” (arXiv:2601.05504) — the MINJA (Memory Injection Attack) query-only attack, over 95% injection and 70% attack success (https://arxiv.org/html/2601.05504v2) [3] Idan Habler and Amy Chang, “Identifying and remediating a persistent memory compromise in Claude Code”, Cisco Blogs, 2026-04-01 — the first 200 lines of memory loaded into the system prompt, persistence across sessions and reboots, and the v2.1.50 fix (https://blogs.cisco.com/ai/identifying-and-remediating-a-persistent-memory-compromise-in-claude-code) [4] Microsoft Defender Security Research Team, “Manipulating AI memory for profit: the rise of AI Recommendation Poisoning”, Microsoft Security Blog, 2026-02-10 — over 50 prompts from 31 companies across 14 industries in 60 days (https://www.microsoft.com/en-us/security/blog/2026/02/10/ai-recommendation-poisoning) [5] Idan Habler, “Memory Is a Feature. It Is Also an Attack Surface”, OWASP GenAI Security Project, 2026-05-13 — the ASI06 Memory and Context Poisoning entry in the Top 10 for Agentic Applications (https://genai.owasp.org/2026/05/13/memory-is-a-feature-it-is-also-an-attack-surface)