If your agent ingests email, scrapes web pages, reads pull request comments, or pulls from any source you do not fully control, you have a trust problem. Most engineers reach for the dual-LLM pattern to solve it. Then they wire up retrieval-augmented generation against a single shared vector store and quietly hand the trust problem back to themselves.
Simon Willison proposed the dual-LLM pattern in April 2023 (Willison, The Dual LLM pattern). The privileged LLM holds tools and never sees raw untrusted content. The quarantined LLM reads untrusted content and never holds tools. A controller passes opaque variable references between them. The pattern has held up well enough that a 2025 paper from IBM, Invariant Labs, ETH Zurich, Google, and Microsoft formalized it as one of six agent design patterns (Beurer-Kellner et al., Design Patterns for Securing LLM Agents). It also informed DeepMind's CaMeL pattern.
The pattern protects the live conversation, but memory sits outside it, and memory is where trust boundaries quietly collapse. Most agents being built in 2026 have memory.
The shape of the failure is straightforward. Workflow A reads an email, summarizes it on the quarantined side, and writes the summary to a shared memories table for future use. Workflow B, days later, runs on the privileged side and retrieves recent memories to inform a tool call. Workflow B's privileged LLM is now consuming quarantined-origin content with provenance washed off. The dual-LLM pattern has been silently inverted, and Unit 42 published a working proof of concept of exactly this against AWS Bedrock Agents in October 2025 (Unit 42, When AI Remembers Too Much). An indirectly injected web page poisoned the agent's long-term memory and persisted malicious instructions across sessions.
The fix has to live in the structure. Trust becomes a column on the memory store, retrievals declare which trust levels they accept, and privileged contexts require trusted-only retrieval. A retrieval audit table watches for violations. None of this depends on a particular vector database.
The pattern survives a vendor migration, a framework swap, or a pivot from Pinecone to pgvector to whatever ships next. The schema below is its simplest credible expression; adapt it to your own stack.
The pattern leaves three problems unsolved: bad trust assignment at write time, an attacker who poisons the embedding space itself (as PoisonedRAG demonstrated at USENIX Security '25 (Zou et al., PoisonedRAG)), and social engineering of the user. It reduces blast radius and forces explicit trust declarations rather than eliminating the threat. As Willison and others have written repeatedly, prompt injection has no clean solution; the goal is to make the failure modes bounded.
The dual-LLM pattern in one paragraph
The privileged LLM accepts input from trusted sources and has access to tools. The quarantined LLM processes any content that could carry an injection - emails, web pages, retrieved documents, third-party API responses - and has no tool access. A controller, written in regular code, passes opaque tokens between them. The privileged LLM sees variable references like $VAR1, never the underlying text. The privileged LLM acts only on data shapes it can verify. The quarantined LLM, even if fully compromised, has no tools to misuse. The original write-up by Willison (April 2023) remains the canonical reference. The Beurer-Kellner et al. June 2025 paper places it in a family of six patterns alongside Action-Selector, Plan-Then-Execute, Map-Reduce, Code-Then-Execute, and Context-Minimization (summary).
Why memory is where the boundary collapses
Long-term memory in modern agents typically lives in a vector store. The agent embeds chunks of conversation, summaries, retrieved documents, or task outputs and stores them with metadata. Future runs retrieve semantically relevant chunks and inject them into prompts. This works well for personalization, context retention, and grounding. It also works well for an attacker.
Three properties of memory turn it into the weakest link of the dual-LLM pattern.
Provenance is usually discarded at write time. Most production schemas store content, embedding, and a generic metadata JSONB. They do not record whether the content originated from a trusted source, an untrusted source, or a mixed-trust process. By the time it is retrieved, the question cannot be answered.
Retrieval is semantic, not authorial. A query for "recent decisions about vendor pricing" returns chunks similar in embedding space. The query treats user-written notes, trusted internal documents, and quarantined web extracts identically. The embedding does not carry trust.
Cross-workflow reads are the design intent. Memory exists precisely so that workflow A's outputs can inform workflow B's inputs. A defense that forbids this misses the point of memory; a defense that permits it without trust segmentation is the failure mode.
The Unit 42 demonstration on Amazon Bedrock Agents shows the failure end-to-end (Unit 42 PoC, Oct 2025). An attacker hosts a malicious web page. The user, in a normal session, asks the agent to look at the page. The page contains an indirect prompt injection that manipulates the agent's session-summarization process. The poisoned summary is written to long-term memory. In a future session, the agent's orchestration prompt includes that memory as part of system instructions. The malicious instructions now run with the agent's full privileges. No vulnerability in traditional code is required.
PoisonedRAG, presented at USENIX Security '25, demonstrates the embedding-space attack (Zou et al., 2024/2025). Injecting five malicious texts crafted to be retrieved for a target query achieves a 90% attack success rate against knowledge databases of millions of documents. The attack does not require write access to the database; it requires the database to ingest from any source the attacker can influence.
These attacks have a common structure: untrusted content reaches the memory layer, and the retrieval layer treats it the same as trusted content. The fix lives at the retrieval layer.
Trust as a column
The minimum viable schema adds two columns to whatever memory table the agent uses. The first records the trust level. The second records provenance - what process or source produced the content.
CREATE TABLE memories (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
content TEXT NOT NULL,
embedding vector(768) NOT NULL,
trust TEXT NOT NULL CHECK (trust IN ('trusted', 'untrusted', 'mixed')),
provenance TEXT NOT NULL,
workflow_id TEXT,
created_at TIMESTAMPTZ DEFAULT now(),
metadata JSONB DEFAULT '{}'
);
CREATE INDEX ON memories USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON memories (trust);
The trust column is constrained to three values. trusted means the content was produced by the user, by code, or by a process that operates entirely on trusted inputs. untrusted means the content was derived from anything the agent does not control - web pages, emails, scraped documents, API responses. mixed means the content combined trusted and untrusted sources during processing and provenance cannot be cleanly separated.
The provenance column is a free-text field that records the origin: user, human-note, quarantined:email-summary, quarantined:web-extract, mixed:research-pipeline, and so on. Provenance is for humans and for forensics. Trust is for the retrieval layer.
Trust is assigned at write time, by the workflow doing the writing, based on rules the workflow knows. A workflow that summarizes an email writes untrusted. A workflow that records a user's explicit instruction writes trusted. A workflow that combines an email and a user note writes mixed. The assignment is structural - encoded in the schema and enforced by the writer - not negotiated at retrieval time.
The schema generalizes. On Pinecone, Qdrant, Weaviate, or any vector database that supports metadata filtering, trust is a metadata field rather than a column, but the pattern is identical. The point is that trust must be queryable at retrieval time and stored at write time, never inferred and never derived from the embedding.
Retrieval as a required argument
Trust segmentation only works if retrievals declare what they accept. A default of "any" defeats the purpose. The pattern below makes trust a required keyword argument, raising a runtime error if the caller forgets.
from typing import Iterable, Literal, TypedDict
TrustLevel = Literal["trusted", "untrusted", "mixed"]
class Memory(TypedDict):
id: str
content: str
trust: TrustLevel
provenance: str
similarity: float
def retrieve(
*,
query_embedding: list[float],
allowed_trust_levels: Iterable[TrustLevel],
limit: int = 5,
threshold: float = 0.7,
) -> list[Memory]:
"""Retrieve memories matching the query embedding, filtered by trust.
`allowed_trust_levels` is required (keyword-only, no default).
Privileged contexts MUST pass ['trusted'].
Quarantined contexts may pass any subset.
Retrievals are logged to retrievals_audit.
"""
levels = list(allowed_trust_levels)
if not levels:
raise ValueError("allowed_trust_levels must not be empty")
rows = db.fetch(
"""
SELECT id, content, trust, provenance,
1 - (embedding <=> $1::vector) AS similarity
FROM memories
WHERE trust = ANY($2)
AND 1 - (embedding <=> $1::vector) >= $3
ORDER BY embedding <=> $1::vector
LIMIT $4
""",
query_embedding, levels, threshold, limit,
)
db.execute(
"INSERT INTO retrievals_audit (caller, levels, results_by_trust) VALUES ($1, $2, $3)",
current_workflow(), levels, count_by_trust(rows),
)
return [Memory(**row) for row in rows]
Three properties matter. First, allowed_trust_levels is keyword-only and has no default. Calling retrieve() without it raises a Python TypeError immediately. This is intentional. A default of ['trusted'] would silently pass the wrong levels for quarantined workflows; a default of ['trusted', 'untrusted', 'mixed'] would defeat the segmentation. The right default is no default.
Second, every retrieval is logged to a separate retrievals_audit table with the caller workflow, the trust levels accepted, and the count of results returned by trust level. A spike in untrusted retrievals into a workflow classified as privileged is a structural anomaly worth alerting on.
Third, the retrieval is a single SQL statement with trust as a WHERE clause, not a post-filter on a broader query. Filtering after the fact is correct in principle but invites bugs where a developer forgets the filter. Pushing trust into the query makes the boundary explicit at the one place every retrieval passes through; it is still the application choosing which trust levels to accept, not the database refusing to serve them, and the CI check below exists because of that.
A CI check that makes violations fail at merge time
The retrieval pattern catches mistakes at runtime. A static check catches them earlier. The check is straightforward: any workflow classified as dual-LLM in the workflow front-matter that retrieves from memories in a privileged node must pass ['trusted'], or carry an explicit override justification.
The implementation is a grep-shaped lint over workflow definitions. In the JonBot codebase, where workflows carry front-matter classification, the check looks like this:
# Pseudocode for a CI lint
for workflow in load_workflows():
if workflow.frontmatter.get("dual_llm") is not True:
continue
for node in workflow.privileged_nodes():
for call in node.find_calls("retrieve"):
levels = call.kwargs.get("allowed_trust_levels")
if levels != ["trusted"] and not call.has_override_comment():
fail(f"{workflow.path}:{call.line}: privileged retrieve must accept only 'trusted'")
The override comment exists because there are legitimate cases - a debug workflow inspecting raw quarantined memories, a reviewer dashboard surfacing recent untrusted summaries - where a privileged context legitimately reads untrusted content. Forcing those cases to be explicitly justified makes the exceptions visible during code review.
What this pattern does not solve
The pattern is a structural defense with known limits. Three failure modes remain.
Trust assignment errors at write time. If a workflow incorrectly writes trusted on content that should be untrusted, the retrieval layer cannot detect it. Mitigation lives at the write side: code review, type annotations on the writer functions, and ideally a writer pattern that derives trust from the source rather than letting the workflow assert it freely. The retrieval layer enforces a contract; it does not validate the contract was honored.
Embedding-space attacks. PoisonedRAG and similar attacks craft content that is both highly retrievable and instructionally malicious. If such content is written with trust='untrusted', the pattern correctly excludes it from privileged retrieval. If it is written with trust='trusted' (e.g., because the writer trusted the source incorrectly), no amount of retrieval-side filtering helps. Defense in depth here means treating any trusted content from an automated source as mixed until a human has reviewed it, and limiting bulk-ingest pipelines.
Social engineering of the user. The pattern defends the agent from the data. It does not defend the user from the agent. A privileged LLM that retrieves only trusted memories can still produce convincing language that manipulates the user into copying untrusted content into the trusted channel. Willison addressed this in the original 2023 post (Dual LLM pattern, You're still vulnerable to social engineering) and it remains true.
The Lethal Trifecta framing from June 2025 (HiddenLayer summary) names the conditions under which any of these failures becomes catastrophic: an agent with access to private data, exposure to untrusted content, and the ability to communicate externally. The trust-boundary pattern shrinks the second condition's blast radius without eliminating the trifecta; agents with all three capabilities need additional governance regardless.
Why this is the highest-leverage pattern to get right
Most security work in agent systems is about narrowing capability or filtering inputs. The trust-boundary pattern at the memory layer is different: it makes the dual-LLM separation persist across time, which is the dimension every other agent defense ignores. A workflow built today against an agent without memory looks safe. Add memory in a future release without trust segmentation, and the same workflow is compromised.
The pattern is also unusually portable. The schema works on any vector store with metadata filtering. The retrieval pattern works in any language. The CI check works against any workflow definition format that exposes call-site detail. A team that adopts it once carries it across stack migrations, framework swaps, and acquisition integrations; the pattern outlives any single implementation.
For builders deciding where to spend security effort on a new agent, this is the place. Detection-based defenses against prompt injection plateau around 97% effectiveness (Beurer-Kellner et al., 2025 paper summary) - attackers iterate faster than detectors. Structural defenses sidestep that race entirely: they do not depend on detecting the attack, only on giving it nowhere to land.
Adoption checklist
For a new agent build:
- Add
trustandprovenancecolumns to the memory schema before the first write - Make
allowed_trust_levelsa required keyword argument on retrieval functions - Default privileged contexts to
['trusted']only - Log every retrieval to a separate audit table with caller, levels, and count by trust
- Add a CI check that flags privileged retrievals without trust restriction or override
For an existing agent with memory:
- Audit current writes to determine the actual trust mix in the existing data
- Backfill
trustandprovenancecolumns where determinable; default tountrustedwhere not - Ship the retrieval pattern with the new schema
- Set a deadline for removing un-classified rows
- Run the CI check against existing workflow definitions to find violations
The migration is uncomfortable but bounded. The cost of not migrating is unbounded, because every memory ingestion path is a potential injection vector and every privileged retrieval is a potential exploitation point.
References
- Willison, S. The Dual LLM pattern for building AI assistants that can resist prompt injection (April 25, 2023): https://simonwillison.net/2023/Apr/25/dual-llm-pattern/
- Beurer-Kellner et al., Design Patterns for Securing LLM Agents against Prompt Injections, summary by Willison (June 13, 2025): https://simonwillison.net/2025/Jun/13/prompt-injection-design-patterns/
- Unit 42, When AI Remembers Too Much - Persistent Behaviors in Agents' Long-Term Memory (October 9, 2025): https://unit42.paloaltonetworks.com/indirect-prompt-injection-poisons-ai-longterm-memory/
- Zou et al., PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation, USENIX Security '25: https://www.usenix.org/system/files/usenixsecurity25-zou-poisonedrag.pdf
- HiddenLayer, The Lethal Trifecta and How to Defend Against It (November 25, 2025): https://www.hiddenlayer.com/research/the-lethal-trifecta-and-how-to-defend-against-it
- Agentic Patterns reference card, Dual LLM Pattern: https://agentic-patterns.com/patterns/dual-llm-pattern/