Prompt injection made stateful, and the memory layer that was built for it
2026. The 2026-09-06 audit found our fences and gates held. The second-pass audit, same trees, harder questions, found the seams those closures had, and we fixed them at fixpoint. This post is the market context for why we run these audits at all: 2026’s research says memory is where prompt injection goes to persist, and almost nobody ships the controls that survive it.
The threat got a name, a paper, and a benchmark
For two years, “prompt injection” meant a hostile turn: the model reads something malicious, maybe obeys it, and the conversation ends. 2026 made it stateful. The attack now writes itself into the one place your agent trusts most, its own memory, and replays every turn after.
The research landed in quick succession:
- OWASP’s Top 10 for Agentic Applications 2026 ranks Agent Goal Hijack #1 and defines ASI06 Memory & Context Poisoning as a top-tier risk, with the Gemini Memory Attack as their named example. OWASP now incubates a dedicated Agent Memory Guard project.
- Unit42 demonstrated indirect prompt injection persisting into long-term memory, the poisoned note waits in the store and fires in a later session (Palo Alto Networks).
- The MemPoison paper measured up to 0.95 attack success rate across memory mechanisms, and it generalizes across designs.
- The framing that stuck: “prompt injection made stateful”.
- And the tell that the market knows: MemGuard exists specifically to bolt trust scores and quarantine onto Mem0/Zep/Letta/LangMem after the fact.
Read those together and the conclusion is uncomfortable: if your memory layer auto-extracts and auto-writes what an agent says, you have given prompt injection a database.
What “answer it architecturally” means
Brain-server + its OpenClaw plugin were built with the assumption that everything trying to enter memory is hostile until a human says otherwise. The 2026 threat model describes our roadmap; here is the shipped answer, layer by layer:
- Screen at ingest. Every write passes a deterministic injection screen
(instruction-override blocklist with translation families, typoglycemia and
encoding tiers, optional local ONNX classifier). Suspect content is
quarantined, excluded from every retrieval leg: full-text, graph, and
vector.
Rejectpolicy never persists it at all. - The human promotion gate. By default, nothing an agent captures becomes
memory. It lands as a proposal in a review queue, scored deterministically,
carrying the exact capture context. The operator approves, and the approval
is digest-bound: it must carry the SHA-256 of the exact bytes the
reviewer saw (
400 digest_requiredwhen absent,409on drift). Rankers rank; they never promote. - Read-seam strips that survive re-assembly. A single-pass strip is not a
closure: our second-pass audit demonstrated
<scr<script>ipt>welding back into a live<script>after the element strip, and nested markdown constructs healing back into auto-fetch images after the dereference. The strips now iterate to a fixed point, pinned by tests with the exact adversarial vectors. - The fence, and the labels. Every injected hit is wrapped in an
unforgeable
UNTRUSTED_*fence, stripped of invisible-Unicode and bidi smuggling, dereferenced of image/link refs, taggeduntrusted: true, and hits that arrive without the flag are dropped, not injected. Captures from group/channel traffic carry a visible[memory | channel-capture]taint label for their whole life, and the host marks replayed labels in inbound text as untrusted. - Scoped principals, not shared gods. The agent authenticates as a scoped principal, recall/store/propose, no purge, no domains, no identity operations, with a probe-blind kill-switch wired into every auth door, and MCP tool scope capped by env.
The honest comparison
The memory-layer market is real and good at what it does: Mem0 for ecosystem and extraction pipelines, Zep/Graphiti for temporal knowledge graphs, Letta for self-editing agent memory. We don’t lead retrieval-quality benchmarks, our docs mark that pending rather than claiming it, and bi-temporal recall has better-published implementations. The 2026 comparisons are right that no system wins every dimension.
What none of them ship as native architecture is the axis above: ingestion screening with quarantine, a digest-bound human promotion gate, untrusted-fence rendering with fail-safe drop, provenance taint labels, tamper-evident audit with pinned heads, erasure certificates with tombstones and legal holds, and a local-first zero-token economy. The proof that the market treats this as a bolt-on is MemGuard’s existence, trust scores layered onto other people’s memory stores, and OWASP incubating a memory-guard project of its own.
If your deployment is a regulated or customer-facing one, ask your memory vendor the 2026 questions: what happens to content your screener flags? who can promote memory, and what binds that approval to the reviewed bytes? how do you prove an embedding was deleted? where does the audit chain’s key live? We published our answers as a control matrix and a live proof map, and a second-pass audit of our own closures, because the first pass is where the work starts, not where it ends.
Sources
- OWASP Top 10 for Agentic Applications 2026
- OWASP announcement, the benchmark for agentic security
- Unit42: Indirect Prompt Injection Poisons Long-Term Memory
- MemPoison (arXiv)
- Persistent Memory Poisoning in AI Agents
- MemGuard · OWASP Agent Memory Guard
- OWASP AI Agent Security Cheat Sheet
- AI Agent Memory Systems in 2026 compared
- Five systems, six dimensions, no winner