In June 2025, Simon Willison defined the lethal trifecta for AI agents: private data, untrusted content, and external communication. Put all three in one agent and the exfiltration path is open. Eleven months later, the post-mortems caught up.
Anthropic's "How we contain Claude" published May 25, 2026 walks through three concrete failures or red-team findings inside its own products. CSA's TrapDoor briefing the next day shows the same pattern escaping the lab into the package ecosystem. Together they turn the trifecta from a useful framing into a documented production failure mode.
Incident 1 - the pre-trust hook. Between mid-2025 and January 2026, Claude Code received three responsible-disclosure reports involving code that executed before user consent. In the clearest case, Claude Code read .claude/settings.json at startup, before the "do you trust this folder?" prompt. A repository cloned for PR review could carry a malicious hook in that file, and the hook would execute before the trust boundary existed. The fix was to defer parsing and execution of project-local configuration until after trust was established. This is the precondition failure: untrusted local content reaches execution before the trust boundary, and with filesystem and network access present, it becomes the trifecta's launch mechanism. The lesson is narrower but important: "local" and "trusted" are not synonyms.
Incident 2 - the phish. A red-teamer at Anthropic emailed an employee a Claude Code prompt that read like routine task instructions. Buried in the setup: read ~/.aws/credentials, encode the contents, POST to an external endpoint. Across 25 retries, Claude completed the exfiltration 24 times. The model-layer defenses anchor on user intent, and the user typed the instruction, so there was nothing anomalous to catch. The only defense that holds in this situation is environmental: egress controls that block the POST regardless of intent, and filesystem boundaries that keep ~/.aws out of reach.
Incident 3 - exfiltration through an approved domain. A malicious file in a Claude Cowork workspace carried both hidden instructions and an attacker-controlled Anthropic API key. Claude followed the instructions, read workspace files, and uploaded them to the attacker's account via Anthropic's own Files API. The egress allowlist saw api.anthropic.com and let the traffic through. Anthropic's own summary is the uncomfortable part: "the sandbox worked perfectly, and yet the data was exfiltrated." The reframing matters more than the fix: an allowlist is not a destination filter, it is a capability grant. Every function reachable through any domain on that list is now an attack surface.
Incident 4 - TrapDoor. Reported May 25-26, 2026 with the earliest confirmed artifact ([email protected]) published May 22 at 20:20 UTC, TrapDoor was a cross-ecosystem supply-chain campaign spanning 34+ malicious packages and 384+ artifact versions across npm, PyPI, and Crates.io (CSA research note, The Hacker News). The packages targeted developer environments where cloud credentials, SSH keys, GitHub tokens, wallet keys, and AI assistant configuration often coexist. The AI-specific twist was persistence through project-local instruction files such as CLAUDE.md and .cursorrules: when a developer later asked an assistant to perform a routine task, the compromised project context could silently steer the agent toward the attacker's payload. Same shape as Incident 1, with the untrusted content arriving through npm install instead of git clone, and the targeted assistant being Cursor or any tool that reads those files.
I wrote about MCP server configs as an attack surface in Your MCP Configs Are an Attack Surface Now and about the protocol-layer cleanup in The MCP Spec Just Did the Cleanup. Those pieces stopped at the protocol boundary. The trifecta sits one layer up: any agent that can read a file and reach a network endpoint inherits the same failure mode, and the file does not have to be an MCP config. It can be a project README, a Slack thread, an email, a workspace document, or - per TrapDoor - a config file shipped through a package manager and pulled in by npm install.
The pattern across all four is the same argument I made about DNSSEC in DNSSEC Finally Has Consequences: the deterministic boundary catches what the probabilistic layer misses. Model-layer guardrails, classifiers, and approval prompts all live at the probabilistic layer. They will never be 100%. The environment layer - egress controls, capability-scoped tokens, filesystem boundaries, sandboxes that do not implicitly trust local paths - is where the failure actually stops.
For my own infrastructure this turned into a short checklist:
- Audit every agent in the stack for the trifecta. If it holds private data, ingests untrusted content, and can reach a network endpoint, the next prompt-injection report is your problem regardless of the model vendor.
- Treat
.cursorrules,CLAUDE.md,.claude/settings.json, and any project-local AI config as untrusted content at parse time. Pin them, code-review them, and do not parse them before the trust boundary. - Reframe allowlists as capability grants, not destination filters.
api.anthropic.comis not a domain, it is "everything Anthropic's API can do." - Egress controls beat approval prompts. The approval is probabilistic. The block is not.
Anthropic says it plainly in their summary: "Be wary of custom components... the standard primitives held while our own work around them exposed flaws." The hypervisor, gVisor, and seccomp held. The custom allowlist proxy did not. The software you build yourself is the layer the post-mortem will name.
Simon was right in June 2025. The receipts are now public.
See also
- Your MCP Configs Are an Attack Surface Now - the protocol-layer predecessor to this argument
- The MCP Spec Just Did the Cleanup - what the architecture should look like
- DNSSEC Finally Has Consequences - same thesis at the DNS layer: the deterministic boundary catches what the probabilistic layer misses