Who are we defending against? Where Meta and Google differ on AI security

Subscribe to the blog

Three months apart, two of the largest AI labs published detailed accounts of how they secure their agents. Google DeepMind released its AI Control Roadmap in June (with a companion blog), and Meta followed in September with a 20-minute walkthrough of the security architecture behind Muse, its personal agent. Both are candid that the model will make mistakes, and so argue for controls that sit outside the model. At a quick glance, they might look like they are approaching the same problem the same way. But they aren't.

DeepMind notes that its threat model resembles Simon Willison's lethal trifecta (private data, untrusted content, and a way to communicate externally), except that no external content is needed to make the agent untrusted. Meta's design assumes the trifecta is happening (the Assume Breach mindset): an attacker only needs to plant instructions in content the agent reads, like an email or a web page, to have it hand your data over. So DeepMind is asking what happens when the agent is the adversary to begin with, and Meta is asking what happens when the agent is turned against you. The answer changes what you build, and where the two builds converge matters as much as where they differ.

Two answers to the same question

Meta treats the agent as a confused deputy, with the user as the principal, which is why consent sits at the center of its design. DeepMind treats the agent as a potential insider: the roadmap assumes internally deployed agents are untrusted, share an adversarial goal, and coordinate, and it extends MITRE ATT&CK to map what they could do, while admitting these threats are theoretical.

It makes sense that each company frames the question this way, given where each comes from. Google DeepMind approaches it with an academic, at-scale mindset: how do you see everything a vast fleet of your own agents is doing, and catch the one that goes wrong. Meta approaches it as a social network, where untrusted parties are expected to push text and prompts into each other's feeds, and now into the agents' feeds as well.

Yet the two framings lead to similar controls. Whether the agent is attacked or simply wrong, the failure looks much the same from the outside, which helps explain why the controls end up resembling each other.

Same shape, different engine

Both designs put a decision-maker the agent cannot influence between the agent and its consequences. Meta's Sentinel is the sole permission authority for connector actions and network egress, and DeepMind's semantic monitors and access controls play the same role. This is the out-of-band control plane we covered in The Modern Cap'n Crunch Whistle: the control signal has to travel on a channel the data cannot reach.

But the engines inside that shared model differ substantially. Meta leans on deterministic enforcement more than interpretive analysis: real credentials live outside the agent, which only holds surrogate tokens that Sentinel swaps for the real thing at the network boundary after approving a request. Run that through Anthropic's "impossible vs. tedious" test from our zero trust post: an agent coerced into leaking a surrogate token has leaked nothing.

DeepMind's primary engine maintains its AI heritage as a semantic check. Trusted models read the agent's reasoning and actions and judge whether something looks wrong. That needs no rule written in advance, but the monitor reads in-band content and inherits its weakness. The team acknowledges this by calling for monitors hardened against jailbreaking and worrying about collusion between agent and monitor. DeepMind also adds deterministic layers as capability grows, and its wider work includes CaMeL, which separates trusted planning from untrusted observations by architecture.

Meta has an in-band layer too, labeling external data untrusted and running injection classifiers, the tag-based approach Agent Data Injection (ADI) showed can be forged. Meta says its deterministic boundaries hold even if the model is persuaded. That is the conclusion we reached on ADI: enforce permissions somewhere forged labels cannot reach.

The two also approach friction differently. DeepMind observes that agents tolerate workflow friction far better than employees, so blocking and review are affordable. Meta's approvals require a human to intervene, so it asks only where consent matters and gates connector actions and egress synchronously from day one. In When Security through Friction Fails we argued that friction as a deterrent collapses against machine-speed attackers; these designs show where it still works, as a checkpoint with a decision behind it.

Where a sidecar fits

I was curious whether these approaches make the sidecar architecture I've advocated for unnecessary, or support it. The answer depends on who owns the stack.

In Meta's model, the sidecar is built into their architecture. The hatch-safety service runs independent models and classifiers as a separate unit outside the runtime cell, specifically so an attacker inside the cell cannot disable them. So when Meta controls the model, the VM and the kernel, there's nothing to worry about, right? Well, if you trust Meta implicitly, sure. But if you assume there might be threats or gaps Meta isn't handling on your behalf, then it makes sense to take the same out-of-band sidecar approach Meta applies to its own system.

In DeepMind's model, a sidecar maps onto the roadmap's asynchronous monitoring and alerting at its first Prevention and Response level, its ability to terminate an agent along with everything it spawned at its third, and the cross-session analysis that distributed attacks need. That is the direction of Agent Vardøger, which ends sessions rather than blocking prompts and correlates activity across sessions; following agent-to-agent communication so one kill cuts off downstream sessions is on our roadmap, not built. The limit is visibility: DeepMind's monitors read reasoning and activations, while a sidecar beside a hosted model sees what goes in, not what the agent thinks. That gap is why response monitoring is next, in line with our argument for outcome-based security.

Both labs apply the same principle: control should live outside the thing it controls. Most enterprises consume hosted models and own neither the model nor the runtime, so the gateway is the one layer they can still add controls to, and a sidecar can add them without sitting in the request path.

The future approaches for AI security

DeepMind ties its controls to a ladder of model capabilities, and the ladder keeps climbing. As models get better at hiding their reasoning, the controls escalate, and the roadmap says control would become infeasible against vastly superintelligent systems. In this manner, the ladder can evolve as the adversarial models do too, until a better solution can be implemented.

The two models are also drifting toward each other's problems. DeepMind's companion policy paper, Three Layers of Agent Security, covers external prompt injection and agent-to-agent risk, and the multi-agent layer is the least settled: cascading failures, tacit collusion, and payloads split into benign fragments that only combine when agents aggregate them. That is a pattern we've outlined before, and part of why we advocate for cross-agent event correlation and outcome-based analysis at execution.

Sharing security signals is the other open area. DeepMind's paper proposes shared signals across companies, possibly an "Agentic Security Exchange," because campaigns that span platforms can look harmless inside any one of them. This mirrors some of the best practices of coordinating security events learned through Google Cloud Platform working with AWS and Microsoft. Meta, on the other hand, isolates each user in a dedicated VM, and the post does not describe how patterns across users would be seen. Meta's own next steps are a Confidential VM later this year and tuning its approval balance as it learns from real use. So, taken together, Meta is taking a much more self-contained approach to security.

So who are you protecting against

None of this tells you which approach is right, and I'm not sure that is the right question. The answer depends on your threat model. When you are building your agents and defining your security policies, you need to decide who you are defending against and how much risk you are willing to accept. A retail organization running a customer-facing chatbot has a different answer than a financial services unit running thousands of coding agents with production access. The controls will be dictated by that difference.

So here is the question I'm left with. Which carries more risk for your organization: agents acting as confused deputies, or as insiders? If you want to talk through your answer, or how either lens applies to your own deployments, reach out at questions@generativesecurity.ai. We're always happy to help work through the hard questions with you.

About the author

Michael Wasielewski is the founder and lead of Generative Security. With 20+ years of experience in networking, security, cloud, and enterprise architecture Michael brings a unique perspective to new technologies. Working on generative AI security for the past 3 years, Michael connects the dots between the organizational, the technical, and the business impacts of generative AI security. Michael looks forward to spending more time golfing, swimming in the ocean, and skydiving... someday.

October 6, 2026
< Back to Blog
Copyright  2026 Generative Security
  |  
All Rights Reserved