The signature doesn't protect what it seems to
AP2 is a payment protocol that Google proposed so that LLM agents could pay on a person's behalf: receive a task, negotiate terms with a merchant, place an order, and transfer money. The trust framework here rests on two signed documents — the Checkout Mandate and the Payment Mandate. They lock in the terms of the deal, and once signed, they can't be swapped out unnoticed.
The catch is in the phrase "once signed." A signature guarantees that the data hasn't changed since the moment it was certified, and says nothing about how it came to be. And it's assembled from a stream of interactions: messages between agents over the A2A Protocol, tool calls via the Model Context Protocol, responses from external services, the contents of pages and emails. All of this lies outside the cryptographic protection. If someone slips a foreign instruction into the context before authorization, the signature will certify an already-distorted intent — and remain perfectly valid while doing so.
Hence the name of the study: the threat doesn't live inside the mandate, but "beyond the mandate."
What was found before, and why v0.2 required a fresh analysis
Weak spots in AP2 had been probed before: version v0.1 described replay attacks and prompt injection. In v0.2, some of these issues were closed — but the fixes brought new capabilities and new deployment assumptions along with them. Every such novelty is a potential attack surface, and the old analysis doesn't transfer mechanically to a new version.
That's what the team — Avital Aviv, Parth A. Gandh, Ron Bitton, and Asaf Shabtai — set out to do. The paper Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2) was posted to arXiv on August 24, 2026 under number 2608.23858, in the cs.CR and cs.AI categories.

Roles, phases, architectures, and trust boundaries
The analysis isn't built as a hunt for individual bugs, but as a mapping of the entire territory. The authors describe the roles of the participants and divide the transaction lifecycle into five sequential phases — from task assignment to settlement and any disputes that follow. In parallel, they identify five deployment architectures: they differ in where the agents are hosted, who holds the keys, and how trusted intermediaries are built into the scheme.
The key step is drawing the trust boundaries. And here the main asymmetry emerges: the mandate itself sits inside the protected perimeter, while the entire information flow from which it's forged is outside. Everything that happens between these zones is the main subject of the paper.

MAESTRO: actors, surfaces, goals
The formal part is built on MAESTRO — a threat modeling methodology for multi-agent environments. It's used to describe four threat actors, eleven attack surfaces, eighteen adversary capabilities, and six goals the attacker pursues.
This decomposition isn't for the sake of impressive numbers. It shows that the adversary isn't a single entity: it's not just an external attacker, but also a compromised tool, a dishonest merchant, a front service. And their motives vary — from directly siphoning off money to quietly nudging a choice in the desired direction.
A catalog of 48 threats
The output is a catalog of 48 threats, grouped into five attack families. They were assessed using AIVSS — a vulnerability scoring system adapted to the specifics of AI. Eight threats reach the High band in at least one architecture.
That "at least one" is telling in itself. The risk level depends not only on the code, but also on the deployment scheme: what's fatal for one integration option may be nearly harmless in another — and vice versa. There's no single answer to the question of how secure AP2 is.
Eight High-risk: the testbed and the demonstrations
There was no public deployment of the protocol at the time of the work, so the authors built a testbed themselves — and covered all five architectures in it. On it, they also built five proof-of-concept demonstrations: each covers its own group of threats, and together they span all eight High-risk scenarios along with the proposed countermeasures.
This is an important argument against the reproach of "it's all theory here." The threats aren't just listed — they're reproduced.

A scanner that knows about deployment
A separate practical result is a scanner that accounts for deployment specifics. Its logic is that what needs checking isn't the protocol in general, but a specific configuration. The scanner matches the threats applicable to a given scheme against three types of checks: static, cross-role consistency checks, and adversarial tests.
Cross-role checks are especially apt here. In a multi-agent scheme, the error often hides not inside an individual component, but at the seam between expectations: one agent is certain that another has already performed the necessary check.
What follows from this
The main conclusion sounds harsher than one would like: a valid mandate signature isn't enough to assert that a transaction reflects the user's intent. A signature confirms integrity, not meaningfulness. If the context was compromised before authorization, cryptography will neatly certify someone else's will.
Practical implications for those building agentic payments:
- Control the input context, not just the final transaction. The bulk of the risk lies before the moment of signing.
- Treat architecture as part of the threat model. The same protocol version yields a different risk profile under different deployment schemes.
- Narrow the agent's authority. The more specific the mandate, the less damage from a distorted intent.
- Check the seams between roles. That's precisely where checks fail most often.
In brief
Cryptography solves exactly the task it was given, and not an iota more. AP2 honestly protects a transaction from tampering after signing; everything that happens earlier remains the developer's responsibility. A study with 48 threats and eight High-risk scenarios isn't a verdict on the protocol, but a map of the terrain: it shows where the protection a signature provides ends, and where the work a signature won't do for you begins.



