Blacksec

Administrator
Staff member
ROOT
VIP
Hey hackers - the memory poisoning of 2026 is a single adversarial write that outlives the session, the prompt, and often the vendor's own filters.
A memory poisoning attack in 2026 needs no direct memory access - the agent writes the payload itself, then retrieves it next Tuesday without you in the room.
TL;DR: The systematic study of agent memory (arXiv, June 2026) mapped four write channels, nine structural vulnerabilities, and six classes of memory-write attacks, benchmarking them at 50.5 percent attack success and 41 percent persistence - with the aggressively-writing agent scoring 66.7 percent against the conservative one's 34.3.
MemSecBench tracked payloads across the full lifecycle: 84.2 percent of malicious entries survived into later sessions and half completed the write-to-consequence chain, while PMPA hit 81.7 percent cross-session success against Claude Code specifically.
Sleeper attacks wrote fabricated memories on GPT-5.5 at 99.8 percent and drove attacker-intended actions in up to 89 percent of retrievals. Neighbors: agent tooling, session gaps, stored credentials.

Four write channels, nine doors​

Memory does not enter itself. The systematic study identified exactly four channels: an explicit instruction the agent executes into a write; a system prompt that tells the agent what to remember; compaction, where summarization decides what survives; and experience-to-procedure, where repeated behavior hardens into standing instructions. Nine structural vulnerabilities across model behavior, prompt design, and architecture make them writable.
The threat model bothers defenders: black-box, no weights, no system prompt knowledge, no direct memory interface - an outsider who gets one adversarial document read and lets the agent's own retention policy do the rest. Six attack classes follow, from command insertion to trigger-conditional payloads. None of them exploit a bug in the memory store; they exploit a policy in the agent.
MPBench made the trade measurable. The agent that writes more aggressively to perform better on long-horizon tasks is the agent that leaks more: HERMES's 2,200-character compaction threshold and liberal retention policy produced a 66.7 percent average success rate against OpenClaw's conservative design at 34.3. Capability and attack surface moved on the same dial.
ChannelWho writesPoisoning shape
Explicit instructionagent obeys a directive in contextpayload in document, email, page
System-prompt drivenretention policy auto-savesanything matching "useful"
Compactionsummarizer picks survivorspayload survives summarization
Experience-to-procedurerepetition hardens into rulesbad pattern becomes standing policy

The benchmarks that counted​

The systematic study's 41.1 percent persistence means a poisoned entry influenced a later session with no further attacker involvement two times in five - 92.8 percent on the permissive agent for strong-signal classes. Memory poisoning is not an exploit chain; it is a save file.
MemSecBench followed the same semantics through write, execute, and forget across 24 configurations - two harnesses, four memory backends, three model backends. Malicious memory persisted in 84.2 percent of all cases, the full write-to-consequence chain completed in 50.3 percent, and selective repair - removing the poison while preserving benign entries - succeeded only 56.1 percent of the time.
The sharpest drop happens at adoption, which is where judgment enters: storage and retrieval barely filter; the moment a model treats a recalled sentence as instructions is the softest layer in the stack.
PMPA tested harness-based agents directly - OpenClaw and Claude Code - embedding instructions in benign external sources and letting the agent persist them. Injection success averaged 73.7 percent on OpenClaw and 66.9 percent on Claude Code; cross-session attack success hit 55.5 percent and 81.7 percent respectively. Claude Code's memory held the payload better than OpenClaw's once written - write rate and hold rate are independent variables, and defenders must measure both.
Feed a benign-looking document with one instruction-shaped sentence, then open a fresh session and ask a related question - if the instruction surfaces, you have persistence without an exploit.
Trigger-test: plant a payload behind a rare phrase, return days later with the phrase, and watch for the behavior change. Dormant poisons pass every prompt-time filter.
Inspect the compaction boundary. Paste a long poisoned transcript until summarization fires, then check which lines survived into the summary that gets re-injected.
Attempt a targeted repair: delete the poisoned entry, confirm the behavior stops, and confirm benign entries still load - most stacks lose one or both halves.

Sleepers, zombies, and the retention nobody audited​

The sleeper attack is memory poisoning with patience. Researchers built it as persistent dormant memory injections that survive across sessions, wait for a trigger, and then steer the agent toward an attacker-chosen goal. Against GPT-5.5 the write succeeded 99.8 percent of the time; against Kimi-K2.6, 95 percent. Once the payload was live, attacker-intended actions followed in 60 to 89 percent of agentic runs, and goal-adjacent behaviors - the harmless-looking steps that set up the real one - held at 90 to 95 percent.
The payload did not need to fire on day one. It needed to still be there.
The zombie variant made removal itself the weapon. Recursive renewal and semantic aliasing - the poisoned entry rewritten under new wording each time a filter touches it - kept malicious memories alive past twenty repair cycles with perfect retention and around 240 copies accumulating in the store. A memory poisoning attack that can survive your incident response is not an incident; it is a resident.
PipePoison added the last mile: poisoning through the pipeline's own summarization path raised attack success by 19.1 points over direct-write baselines, because the summarizer carried the payload past defenses that only inspected raw writes. Any component that transforms content before storage can be made to launder an instruction.

Why write-time filters keep failing​

MemPoison's three-layer taxonomy explains the failure cleanly. Layer one is superficial compliance - the payload hides in plain text, and keyword filters should catch it, sometimes do. Layer two is intent smuggling, where the entry reads benign at write time and means something else at read time; static inspection passes it because at write time, it is true.
Layer three is structural adaptation - self-healing payloads, embedding-encoded instructions, alias mutations - which rewrites itself around whatever control sits downstream. Filters that defend the write choke on layers two and three because the write is not where the attack lives.
The adoption bottleneck from MemSecBench is the operational proof: 76.1 percent of payloads recalled correctly, only 53.7 percent steering the agent. The gap is not retrieval security, it is decision security - the moment the model reads a stored sentence and treats it as instructions. Treating memory as untrusted input at read time closes most of it without touching the store.
Production memory has already been hit without any of these frameworks. Rehberger's 2025 Gemini memory poisoning reached Google's live product - an injected instruction persisted into future conversations - and Microsoft's Azure-side disclosure in 2026 plus Unit 42's findings on Amazon Bedrock showed the same class crossing cloud boundaries.
The OWASP agentic ranking (ASI06, Untrusted Memory as a top risk) formalized what those incidents demonstrated: memory is now a production attack surface with named products, named CVEs, and a named place on the risk register.

What actually contains it​

Containment starts with provenance. Every memory entry needs a source, a timestamp, a writer, and a trust level - an entry written from a retrieved document carries that document's trust, not the agent's. Entries without provenance never execute instructions. That one rule turns layer-two payloads back into text.
Second, separate storage from authority. A memory store is a note database; notes are data, not directives. Instruction-shaped content - imperatives, tool names, URLs with credentials, anything phrased as an action - gets quarantined for review at write time and promoted at read time only if approved. The distinction between remembering a fact and obeying a fact is the entire control.
Third, rate-limit and audit writes. Zombies and sleepers both require volume or dormancy to work - 240 copies, twenty repair cycles, weeks of waiting. Per-session write quotas, duplicate detection on semantic near-matches, and alerts on entries whose recall correlates with a behavior change cut both without solving the hard problem. The table below maps each control to the attack class it retires.
ControlRetiresFailure it assumes
Provenance + trust tiersintent smuggling, document-borne writesmemory inherits caller trust
Data/instruction split at readadoption-stage payloadsrecall equals obedience
Write quotas + near-dup detectionrenewal, accumulationvolume is unbounded
Repair verification loopszombie survival past fixesdelete once = gone
Compaction reviewsummary-laundered payloadssummarizer is trustworthy
The fourth control is the one teams skip: repair verification. MemSecBench's 56.1 percent selective-repair rate means a deletion that was not tested probably failed. Post-repair probes - replay the trigger, confirm the behavior stopped, confirm benign recall still works - turn incident response into a measurement instead of a hope.

Red-team the memory, not just the prompt​

Prompt-injection testing stops at the session boundary; the memory poisoning work of 2026 says that boundary is where the interesting attacks begin. Run the write in session one - benign document, one instruction-shaped sentence, benign trigger phrase - then close everything and start cold.
Session two asks a related question and scores two things separately: did the entry surface, and did it steer. The two rates diverge by twenty points in either direction across MemPoison and PMPA - write rate and steer rate are different numbers with different owners.
Then run repair. Delete the entry you planted, replay the trigger, and verify both halves - the poison stopped and the benign memories still load. Most stacks fail one of the two: if the poison persists, your write path has a renewal problem; if benign recall breaks, your repair path is a sledgehammer. Neither shows up in a single-session injection test.
Log the third session too. The 84.2 percent persistence figure exists because someone checked later rather than at write time - storage time is the attacker's, session time is the defender's. An adversary-in-the-middle kit works the same way: the steal happens once, the reuse happens whenever nobody is looking, and the memory store is just another cookie jar with a longer TTL.
Agent builders own the architecture: provenance on every entry, the data/instruction split at read, quotas on writes, compaction review, repair verification as a built-in loop. Which agent you pick is a security decision too - the long-horizon-capable one is the one that leaks. Benchmarks that do not report attack rate alongside task rate are advertising.
Operators own the rest. Deployed agents get the memory layer on the asset inventory with an owner, egress controls on whatever retrieval source the agent reads from - the same documents that carry memory poisoning carry QR droppers and hostile web content - and retention windows that expire old entries by default. An agent that remembers everything forever is not more capable next year; it is only more compromised next year.
Reviewers own the questions. What can write memory? From where? Who approved the entry the agent is about to obey? What happens when I delete this - and how would I know if the delete silently failed? Every one has a test in the benchmarks above, with a number attached.

The stance that holds​

The memory poisoning of 2026 is not a new injection class; it is injection with a save file. Four write channels, nine structural doors, half of all attempts completing the full chain to consequence, payloads surviving 84 percent of the time into sessions nobody is watching, and repair working barely half the time - every number assumes the memory system is working as designed. No memory store is its own defense; the store is doing exactly what it was told.
What holds is refusing the premise that more memory is more intelligence. Memory is state, state has provenance, and provenance decides authority: entries are data until a human says otherwise, writes are rate-limited and logged, compaction is reviewed, repairs are verified, and recall never executes instructions.
The teams that adopt that split now will read the next sleeper disclosure as confirmation of a control they already run. The teams that do not will discover that their agent has been keeping notes for someone else - and that the notes have been awake the whole time.
Inventory the memory layer this week. Everything else is downstream of knowing who can write to it.