Skip to content

Engineering blog

Mandatory AI approvals can still be optional

An approval step can be mandatory while another credential bypasses it. Test whether every route to a high-risk action enforces the same human decision.

An AI policy can be an unlocked door with a sign on it that says “please don’t.”

The sign is well written. Legal reviewed it. Everyone signed off.

The agent walks through anyway. The system accepts its credentials and executes the request. Nothing on that route ever checks the sign.

So you buy a control plane. Now there is a guard at the entrance, checking requests, recording visitors and stopping anyone without approval. The demonstration works. An unapproved request gets rejected. Everyone can see the rejection in the dashboard.

Then someone uses the side-door key.

The guard has done exactly what it was asked to do. The building is still open through another entrance.

An approval can be mandatory in your workflow and optional in your system.

A second route to the same deletion

Consider a hypothetical service that deletes customer datasets. Company policy requires human approval. The agent has a tool called delete_dataset, and the service behind it checks the approval before deleting anything. No approval, no deletion.

The agent's runtime also has a usable database credential. That credential permits deletion of the same records, and the database is reachable from the runtime. A direct database request therefore achieves the same result without calling delete_dataset.

The approval service can truthfully report that every deletion request it received was checked. The database can truthfully log the direct deletion under a valid identity. Nobody needs to falsify a log for the approval requirement to have failed.

A perfect record of the front entrance says nothing about who used the side door.

Control planes can close this gap. A policy decision can be enforced by a service, gateway or resource that the caller cannot bypass. NIST's zero-trust architecture explicitly distinguishes the components that make access decisions from the enforcement points that apply them. The deployment has to make enforcement unavoidable. NIST architecture

That gives you a more useful evaluation than watching someone approve a request in a demo.

Test the refusal

Choose one operation that your organization requires a human to approve. Reproduce it with synthetic data in a test environment. First establish that a valid approval allows the intended action. Then try to make the action happen without satisfying the requirement.

Try thisRequired result
Send the request without an approvalRefuse before changing any data.
Approve deletion of dataset A, then request dataset BRefuse the mismatch.
Present a proof signed by a key the service does not trustRefuse it, even if the signature itself is valid.
Submit the same approval twice, including concurrentlyAllow at most one authorized operation.
Use a direct database connection, another API client or a different toolRefuse unless that route independently enforces the same approval requirement.
Make the approval service unavailableRefuse to proceed without valid authorization.
Try to remove the check or obtain the executor's credential from the agent's environmentDeny the agent that authority.

Inspect the resulting data as well as the responses. An error message is little comfort if the deletion happened before the error was returned. These are acceptance criteria to demonstrate, not results to infer from a feature list.

The outage case does not rule out offline approval. An offline path can satisfy the requirement if it verifies a valid proof and still enforces its freshness and single use. “The service was down, so we let it through” cannot.

This is an application of complete mediation: every relevant request encounters authorization. OWASP recommends downstream authorization and human approval for high-impact actions in its guidance on excessive agency. Those requirements concern what the system enforces, independently of whether the model follows its instructions. OWASP Excessive Agency

Enforce approval where the action executes

In the deletion example, fixing the tool wrapper is insufficient while the database still accepts the agent's direct deletion request. Remove that privilege. Keep the execution credential in a service the agent cannot reconfigure, and make the service enforce approval before it acts. Apply equivalent controls to any other route that must retain deletion authority.

The approval must also identify the action precisely. “Yes, help with customer cleanup” leaves much more room than authorization to delete a particular dataset in a particular workspace. The executor must compare the approval against the parameters it will actually use, including the target. Approver keys and requirements must come from independently controlled configuration; accepting whichever key a proof supplies would let a requester invent its own approver.

Single use needs enforcement too. A signature does not disappear after someone checks it. Durable redemption state, concurrency handling and the operation's transaction or idempotency design have to prevent a second execution. Likewise, a correct verifier offers little protection if the agent can edit it out of the process.

Administrator and emergency access deserve the same scrutiny. If a separate operator can bypass the requirement, identify that exception, restrict who can invoke it and record its use. The guarantee you describe must match the powers people and services actually hold.

Make the evidence verifiable

Once the operation is controlled, there is another question: what can someone verify afterwards?

A signed approval can provide evidence of which action a credential authorized. It does not, by itself, prove that the action ran, that it ran only once, or that no other route existed. An execution record answers a different question and needs its own supporting evidence.

A tamper-evident ledger can commit those records to a history. An inclusion proof lets a verifier check that a particular record belongs to a checkpoint. External anchoring can make subsequent changes detectable against a checkpoint witnessed outside the ledger operator's control. Verification still needs independently chosen trust; taking every key and trust assertion from the same bundle you are checking is not independent verification.

The artifacts should be downloadable: the approval receipt, the relevant event proof and the checkpoint evidence. A customer or auditor should be able to inspect them without treating a green badge in the vendor's console as the answer. If sealing is periodic, a new event may have to wait for its inclusion proof. Show that state honestly.

There is a boundary to this evidence. A ledger cannot reveal an action that was never committed to it. Anchoring the visitor log does not tell you who entered through an unmonitored door.

INTYGA provides authorization and witness capabilities for this design: passkey-signed approvals bound to specific actions, and exportable evidence for independent verification. The integration still has to enforce the approval where the action executes and close the alternate paths. Putting INTYGA in a workflow while leaving the agent a separate execution credential would preserve the same defect.

The principle also applies to scripts, backend services and human operators. AI agents make the question urgent, but the receiving system has to enforce the requirement whoever sends the request. Nor does every operation need a person to approve it. Decide which actions require that intervention, then make the requirement real. A signature records authorization; it does not make a harmful decision wise.

Ask to see the side door

When a vendor or internal platform team shows you its approval flow, ask to see the refusal through another client. Ask who holds the execution credential and who can change the gate. Ask what happens during an outage, and what evidence you can take away afterwards.

If you need a starting point, map what your agent can actually reach. Then test the routes you find.

The sign belongs on the wall. The guard has useful work to do. The door still needs a lock.

Which credential in your stack can still do the thing without meeting the guard?