Skip to content

AI Agent Governance & MCP Control

Give AI agents capabilities. Keep humans in the loop for dangerous actions.

Autonomous AI agents run at machine speed, but when an agent requests a destructive MCP tool call—wiring money, deleting tables, or modifying cloud infrastructure—INTYGA holds the execution until a human signs the exact parameters. The signature is over the parameters the tool will actually receive, and it is checked by your runtime, so a persuaded model cannot talk its way past the gate any more than it can forge a passkey.

Why generic prompt rules fail

System prompts and LLM guardrails are soft controls easily bypassed by prompt injection or unexpected context. They live inside the same probabilistic system they are supposed to constrain. INTYGA enforces a hard cryptographic gate at the tool boundary — outside the model and outside the agent process, provided the agent cannot reach the policy file or spawn the server itself.

  1. The agent proposes an action via an MCP tool call or API request.
  2. INTYGA issues an out-of-band WebAuthn challenge to a human approver.
  3. The approver signs with passkey/Touch ID; the verifier checks the signature offline.
  4. Only then does the original tool handler run — with the parameters that were signed, not a later revision of them.

A protected action needs a valid tenant rule. New workspaces deny unknown action IDs by default; an administrator may explicitly route them through the tenant baseline. A malformed or conflicting rule refuses the request. None of these cases silently authorizes execution.

// The executing service prepares these values from trusted state.
// The model may carry them, but must not choose them.
{
  "action": "drop_database",
  "target": "prod-db-cluster-01",
  "actionDescription": "Drop the approved database",
  "params": { "database": "prod-1" },
  "agentContext": {
    "action": { "reversibility": "irreversible", "amount": null },
    "configDigest": "sha256:<digest-of-live-model-tools-prompt>",
    "delegatedBy": null,
    "session": {
      "id": "sha256:<digest-of-random-session-id>",
      "seq": "1", "prev": null, "aggregate": null
    }
  }
}
// Replace placeholders with actual SHA-256 digests before calling
// verify_human_authorization. The service keeps the same context.

Connect the agent and enforce the receipt

1. Connect to the hosted MCP server

Create an agent and API key in the console, exchange it for a 15-minute Bearer token, and connect a Streamable HTTP client to /mcp. The agent calls verify_human_authorization with the exact action and RP-prepared agentContext, then polls check_human_authorization.

2. Have the human sign

The owner reviews the agent identity, configuration digest, action and running amount in the browser and approves with a passkey. The agent forwards the complete receipt and nonce; an APPROVED status alone cannot unlock the action.

3. Verify at the execution boundary

Your service checks the receipt against trusted approvers, the actual operation, the live agent configuration and its locked session state. It then atomically reserves the nonce, advances the session head and enforces the budget before executing.

Follow the public AI setup checklist and the MCP reference. You can connect directly to hosted MCP, wrap your own server with @intyga/mcp-sdk, or place @intyga/mcp-proxy in front of a stdio server. Both adapters require an RP-owned v1 runtime for AI-agent keys so they can verify the receipt and reserve the session before the tool runs.

Where the security boundary actually is

Not in the prompt, and not in the agent's good intentions. The boundary is the tool handler that does not run until a receipt over its exact arguments verifies against keys you already trusted. Everything the model says is on the untrusted side of that line.

High-risk AI tool calls to gate

Financial transfers & API wires

Prevent agents from issuing unverified payouts or transferring funds without human sign-off.

Destructive DB & cloud ops

Hold schema wipes, table drops, and bulk data mutations for explicit passkey confirmation.

Credential & IAM modifications

Ensure prompt injections cannot escalate privileges, grant admin access, or rotate production keys.

Outbound messages that commit you

Emails to customers, support refunds, tickets closed as resolved — cheap to send, expensive to retract, and a favourite of injected instructions.

Code and infrastructure changes

An agent that can open a pull request is one merge policy away from an agent that can deploy. Gate the write, not the suggestion.

Bulk reads of sensitive data

Exfiltration rarely looks destructive. A full-table export or a mass document fetch deserves the same signature a delete does.

What the signature proves — and what it does not

The notification carries only an opaque reference; the approver's browser fetches the gateway-frozen action and parameters over TLS. The requester can still choose a misleading description or action ID, so the executing service must bind those fields to the operation it will actually perform. The agent's DID comes from its token rather than the model's text, and the executing service must check the expected DID independently.

The agent's policy is advisory. The rules are not.

The local policy manifest travels with the agent, so the gateway cannot authenticate it — a compromised agent could present a permissive one. The gateway therefore honours a client deny as an escalation and nothing else; an allow sent to the gateway never skips human approval. A separate service may allow its own low-risk operations locally only when its own policy and credentials make that decision, outside the model.

Whether a human must sign, how many, and with what class of key are server-side approval rules, evaluated by the gateway and frozen onto the challenge when it is created. There is no quorum argument on the wire, so an agent can request an authorization but cannot name a weaker one.

The agent supplies an action ID. The gateway matches it exactly and requires every specific rule to retain the tenant baseline. New workspaces deny unknown IDs unless an administrator explicitly chooses baseline approval. Older workspaces activate exact matching with a signed policy change. See approval rules for the full behaviour.

The manifest itself is zero-knowledge when you want it to be: encrypt it with your organization's public key, publish only ciphertext and a hash, and decrypt it in your own runtime. The gateway relays the blob and enforces version freshness. It never holds your private key and never sees the plaintext.

  • M-of-N quorum — distinct valid signatures before an agent's request is approved.
  • Four-eyes — the human who set the agent running cannot be the only one who ratifies what it asks for.
  • Hardware-key class — device-bound authenticators only, optionally narrowed by AAGUID to a specific model.
  • Attested requester — the agent's workload must have proved its identity by third-party attestation (OIDC / SPIFFE). A leaked client-credentials key does not satisfy it, and neither does key possession alone.
  • Escalation — a pending request widens to more approvers after a timeout. It never lowers the quorum or extends expiry.

Standing authority, granted by humans and witnessed

Some agents need a scope, not a prompt every time. An Agent Authority declares what one non-human identity may ask for — a target plus parameter patterns — and it grants nothing at all until a human quorum seals it through the ordinary approval ceremony. The seal is recorded in the anchored ledger, linking the agent to the named humans who granted it.

  • Sealing is the switch. An agent with no seals is unbounded; the declaration alone is inert.
  • An agent with at least one live seal is confined to the union of its sealed scopes at /authorize — out-of-scope requests are refused, and the refusal is itself witnessed.
  • A seal never approves anything. In-scope requests still take the ordinary human-approval path.
  • A seal ends at revocation or expiry; both are witnessed. Disabling the declaration merely stops it being re-sealed.

This is the narrowing mechanism, not a bypass: it makes an agent's permitted question set smaller and provably human-granted, rather than making its answers automatic.

What this breaks, and what it does not

Broken: injection that reaches a dangerous tool

A poisoned document, a hostile web page, or a malicious sub-agent can make the model ask for a destructive call. It cannot make the call execute. The request surfaces to a person as its real parameters, fetched from the gateway rather than narrated by the agent, and expires unsigned if nobody grants it. A leaked agent credential has the same ceiling: it buys the ability to request, never to approve.

Not broken: an agent that is wrong in a plausible way

If the parameters look reasonable and a human approves them, the action runs. The gate is a decision point, not a reviewer — it does not know your data model or your intent. Its value is that the decision exists, is attributable, and is bound to the exact arguments. Pair it with narrow tool scopes and reversible defaults.

The evidence an agent leaves

A receipt per approved call

Signed over the canonical payload — agent identity, target, tool, arguments, nonce, time. Verify it with @intyga/verify, offline, with no INTYGA secret.

Denials and refusals too

What an agent asked for and was refused is often the more interesting record. Out-of-scope requests under a sealed authority are witnessed as violations.

A tamper-evident chain

Entries append to a hash-linked ledger with periodic checkpoints, so committed records cannot be quietly rewritten after an incident.

Scoped honestly: the log proves committed entries were not altered. It cannot prove an event was never withheld from commitment. See the security architecture for how the chain is anchored, and the EU AI Act mapping for how a per-action signed record supports demonstrating human oversight.

This pattern extends beyond agents

The agent path is one adapter over a general primitive: no high-risk action without a signed human approval and a verifiable record. The same gate sits in front of deployments and Terraform applies, payments and treasury movements, privileged production operations, and package releases — humans and backend services call it the same way agents do.

The agent requests. The human approves.

Zero friction for read-only and benign tool calls; unyielding cryptographic proof for high-stakes operations. Approvals are not metered, so the Developer tier gates as many tool calls as your agents make, at no cost — what the paid tiers add is seats, a longer evidence-retention window, and quorum rules.

Explore security architecture →