Skip to content

Engineering blog

You don't need to trust the AI agent

INTYGA does not make agents trustworthy. It makes trusting them unnecessary — when your own service verifies a human's signed approval before it acts, and holds the only credential that can act.

Can we trust the agent to ask before it acts? The honest answer is no, and it is not a criticism of any particular model. A language model is a text predictor holding a tool. Prompt injection can redirect it, a long context can push the instruction out of reach, and a plain bug can skip a step. A control that depends on the agent remembering to use it has a hole the size of the agent.

INTYGA's position is narrower and more useful than making agents trustworthy: it makes trusting them unnecessary. That holds under one condition, stated up front because it is the whole argument. The check has to sit in the process that performs the action, and that process has to hold the only credential that can perform it. The first half is an integration. The second half is an inventory, and it is yours to do — this post comes back to it.

Asking the agent to check its own approval

The first integration most teams build looks reasonable: the agent calls a tool to request approval, polls until it sees APPROVED, and proceeds. The approval is real — a human signed it with a passkey, the signature is bound to the exact parameters, the ledger has the record. What is wrong is who is holding the result. The agent is. Whether it proceeds, and whether what it does matches what was approved, is a decision taken on the untrusted side of the line. Our own MCP post and use-case page describe this shape; read them as the evidence half, and this post as the enforcement half.

Concretely, an agent in that position can decline to ask at all. It can read APPROVED where the tool said DENIED. It can get approval for one database and drop another. It can present, a second time, a receipt a human genuinely signed ten minutes ago. It can skip the tool and call the production API directly. None of these require a malicious model. Prompt injection, state loss, parameter drift, tool-selection errors and ordinary software bugs are enough.

Evidence is not enforcement

An approval obtained over MCP is evidence. It is very good evidence: a signature over the canonical bytes of the action, verifiable by anyone with the approver's public key and no INTYGA secret, anchored in a witness ledger. But evidence describes what happened. Enforcement decides what happens next, and only the thing that performs the action can do that. The cryptography is identical in both cases. What differs is which process gets to say no.

For an AI agent, the signed v1 receipt also carries a configuration digest and an ordered session statement. The digest only helps when your executing service recomputes it from the live model, tools and prompt; it is not an attestation of the agent. Your service must keep a durable session head and budget so a chain of individually approved small actions cannot bypass an overall limit. The approval screen shows the agent identity and running amount, while the receipt keeps raw prompts out. Signing stays asynchronous and these runtime checks happen in your service.

The receipt travels with the action

The correct shape has the agent carry the receipt to the place the side effect happens. When it calls your service to perform the approved action, it sends the receipt and the challenge nonce alongside the request. Your service — not the agent, not INTYGA — then enforces five properties before it touches anything. The receipt is an authorization artefact, not authentication: your service still authenticates the caller exactly as it does today.

  Untrusted agent
        |
        |  action + receipt + nonce
        v
  Enforcing backend (your service)
        |
        |-- verify the signature(s) against approver keys it already holds
        |-- recompute the canonical payload from ITS OWN target, action and params;
        |   compare byte-for-byte with the signed bytes
        |-- check expiry and the signed quorum
        |-- pin the approval ceremony (origin, RP ID) and the expected requester
        |-- record the nonce before the side effect (single use)
        |
        v
  Production side effect
  1. Verify the signature against approver keys the service already trusts — loaded from its own configuration, never trusted from the receipt. A receipt that supplied the key that validates it would be vouching for itself. The keys come from the trust-anchor file your admins export from the console; ship it as configuration, load it once at startup, and rotate it the way you rotate any other trust root.
  2. Recompute the canonical payload from the request the service is about to execute — target, action type, parameters — and compare it with the signed bytes. The approval is over the canonical (RFC 8785) form of the parameters, so a one-byte difference in that form is a different action and the one-database-for-another swap fails. The target string has to be the same one the agent requested approval for: agree it once, per service.
  3. Check expiry, which the verifier enforces fail-closed by default against your clock, with a ±30 s tolerance by default, and the quorum your admins set in Approval Rules, which is signed into the requirement. Supply approvers as identities (the exported trust anchor does this), so a two-person approval is two people; with bare public keys the count is of credentials, and one approver with two passkeys is two. If your service wants its own floor, read requiredApprovals out of the canonical payload and refuse anything lower.
  4. Pin the approval ceremony and the requester. Every passkey approval is a WebAuthn assertion, and the verifier refuses one unless you supply the origin and RP ID of the console the human approved on — otherwise an assertion minted at any other site would verify. The requester is in the signed bytes too; assert the agent you expect (TypeScript and Python verifiers today; in Go, Rust or Java, pin the agent by your own authentication on the call).
  5. Consume the nonce. Record it as redeemed before the side effect — in the same transaction where the side effect is a database write; where it is an external call, insert first and accept that a failed call burns the receipt, which is the fail-closed direction. Keep the row until the signed expiry has passed, after which the receipt is dead anyway. A receipt is single-use because your service makes it so.

One detail that matters for the replay step. The gateway has its own redemption endpoint, POST /authorize/verify, for the case where the process that requested the approval is the one that acts; it byte-matches the parameters and burns the nonce server-side. It is bound to the principal that requested the challenge, so when the agent asked, only the agent can redeem there — deliberately, because it stops one agent redeeming another's approval. That is why a backend the agent calls keeps its own nonce ledger, which the DIV specification already names as a relying-party responsibility. Expiry bounds how long a receipt is valid; your ledger bounds how many times.

The credential the agent should not have

Everything above assumes the agent has exactly one way to reach production: through the service that verifies receipts. If the agent also holds a database password, a cloud role, a deploy key, or a direct URL to an unguarded admin endpoint, it has a second route, and the gate on the first route protects nothing. This is not a weakness in the receipt. It is the reason the receipt has to be checked by the process that holds the production credential, and why that process must be the only one that does.

Said plainly: the property this post describes holds only when INTYGA is placed at the execution boundary and the agent has no alternative credentials and no route around it. An inventory of what the agent's runtime can reach — environment variables, mounted secrets, IAM roles, network egress — is part of the integration, not an afterthought. The easiest gap to leave behind is an agent that calls a gated tool for the dangerous operation and still has the raw credential in its environment from the prototype.

Where the check can live

For the v1 AI-agent receipt, put the check at the tool or service that performs the side effect. That boundary owns the trusted action and session state; the agent only carries the challenge and receipt.

  • The relying party prepares agentContext from its own configuration and durable session state, sends it with the exact action to verify_human_authorization, and keeps a protected copy. The model must not choose the expected context.
  • On approval, the agent forwards the complete receipt and nonce to the backend or tool server that holds the execution credential. The backend recomputes the live configuration digest, verifies the receipt against its own action and approver keys, checks the locked session chain and aggregate, and atomically reserves the nonce, head and budget before acting.
  • For a tool server you control, @intyga/mcp-sdk performs these checks before calling its handler. For a stdio tool server, @intyga/mcp-proxy does so before forwarding the call. Both require an RP-owned AgentV1Runtime with trusted keys, live configuration and an atomic nonce/session/budget reservation; without it an AI_AGENT request fails closed.

With this boundary in place, the decision is taken by a process the agent does not control. The agent is a courier that carries a request in and a receipt back; its cooperation is not what authorizes the side effect.

Cooperative integration and enforced authorization

Cooperative integrationEnforced authorization
Who checks the approvalThe agent, by reading the tool resultThe process that performs the action
What a prompt injection can doSkip the check, or act on a different approvalOnly what the policy explicitly allows, or what a human signed — so its remaining target is the approver, who now sees the bound parameters
What the receipt is forA record after the factThe precondition for acting
ReplayThe gateway's consume step — if the agent calls itA nonce the backend already redeemed fails
Depends onThe model following instructionsThe agent having no other route to production

Cooperative integration is not worthless. It is how most teams start, it produces real evidence, and the gateway's MCP server publishes in-band instructions — in InitializeResult.instructions and in the descriptions of both authorization tools — telling the model to forward the receipt rather than act on a status. A cooperative model can follow them; nothing makes it. A control whose correctness depends on the model following instructions is the control this post is arguing against. Instructions make a cooperative agent do the right thing. The enforcing backend is what makes a hostile one irrelevant.