Agent PlaybookOpen the workbench
Generated footage

Measured, July 2026

When the procedure that governed the work was not already in front of them, frontier AI agents went and looked it up

3.5%of the time

106 real engineering tasks across 49 codebases, run against four frontier models including Claude and GPT-5. Compliance with refuse this was zero percent. Compliance with hand this to a human was zero percent.RepoComplianceBench, arXiv 2607.26819

Your agent has your rules.
It is not reading them.

This is not a prompt problem. The same research found that agents skipping the procedure is independent of how well they reason, so a smarter model does not fix it, and neither does telling it harder. Writing the rule into the prompt moves compliance from roughly 81% to 91%. One run in twelve still goes unguarded.

You have felt this without having a name for it. A job gets done a second time because nobody checked whether it had been done before. A procedure gets written, saved, and never opened again. The work looks finished. The dashboard says running. Nothing tells you otherwise for nine days.

What we built

A gate, not a reminder.

Before the agent can act, three questions are answered without it being asked: what is the procedure for this work, has this work been run before, and what went wrong last time. If there is no procedure and no precedent, the action is refused and a person is asked. Not logged. Refused.

  1. 01

    The procedure is loaded, never requested

    The agent is not asked to go and look. Looking is what fails 96.5% of the time. The governing procedure is resolved from the work itself and put in front of it before the first step.

  2. 02

    Loaded is not read, so the action waits

    Agents call retrieval and then decide before looking at what came back. The tool call is refused until the read has actually happened. This is an if statement in the harness, not an instruction in a prompt, which is why it holds.

  3. 03

    What it cannot do, it cannot be told to do

    Refusing and escalating score zero percent when they are rules. They score differently when the tool simply does not exist on that path. A prohibition an agent cannot violate is not a rule. It is a guarantee.

The standard underneath it

192 checks. Eight stages. Three tiers.

Every agent is graded against them, and the list filters itself to what your agent actually does. A simple internal routine needs about seventy. A client-facing receptionist needs far more.

114

Day 1

Mandatory before it does real work, even with you approving everything it sends. An agent with one open Day 1 check does not run.

41

Autopilot

Required before it acts without asking. Crash recovery, full tracing, shadow mode, 72 hours unattended.

37

Scale

Added later to make it smarter and cheaper. Never required to launch, and never before two clean weeks.

Build an agent against them

Where this sits

Everyone gates the action.
Nobody gates the knowledge.

We checked, against their own documentation. Several of these will stop an agent from calling a dangerous tool, and some do it well. Not one of them will stop an agent acting on work whose procedure it never opened.

  • Claude Agent SDKBlocks tool calls. Strongest of the set.No
  • OpenClawBlocks tool calls. The feature named policy does not enforce anything.No
  • Microsoft Agent GovernanceBlocks actions in under a millisecond. Free, and good.No
  • Meta MuseApproves what leaves. Approved a home address to a stranger three weeks in.No
  • OpenAI Agents SDKTripwires and approvals. The judgment is itself a model.No
  • Temporal, Inngest, RestateEnforce the workflow. No awareness of a model at all.No