Why "the model cannot do it" beats "the model must not do it" — Agents for Humans
Every safety constraint I wrote for a release agent came in two versions that look identical in code review and are worth about ten times different. A missing edge in a state machine, an action vocabulary with no word for a path, and a forged approval that fails not because it was detected but because there was nothing to look up.
I spent two weeks building a release agent — the kind that takes "please ship this safely" and turns it into an actual deployment — and the most useful thing I learned had nothing to do with prompts.
Every safety constraint I wrote came in two versions. One of them is worth roughly ten times the other, and they look almost identical in a code review.
The two versions
Take the simplest rule in a release pipeline: passing the tests is not permission to deploy.
Version one:
1
2
3
if not release.approved:
raise NotApproved("a human has to approve this first")
execute(plan)Version two: the state machine that governs a release has no edge from
GATED to EXECUTING.1
2
3
4
5
6
TRANSITIONS = {
...
# Deliberately excludes EXECUTING: gates passing is not authorisation.
ReleaseState.GATED: frozenset({ReleaseState.WAITING_APPROVAL, ReleaseState.BLOCKED}),
...
}Both stop an unapproved deploy today. They fail differently, and that is the whole thing.
Version one fails when someone forgets to write the check — in the new code path, the retry branch, the "quick fix" that calls the executor directly. Version two fails because the transition does not exist. There is nothing to forget, because there is nothing to remember.
The test I ended up using on every constraint:
Four of them
I built the same substitution four times.
The executor cannot approve its own work. Not "the executor shouldn't call approve" as a convention. Issuing and consuming are two different objects, and the executor is handed the consumer — which has no method that mints a grant.
The model cannot choose where a release lands. Not "validate parameters and reject
../". The action parameter vocabulary has no term for a path. Deployment actions are drawn from a seven-value enum, their parameters are schema-checked, and where is fixed in the release target before the agent object exists.The model cannot run commands. Not "check the command against an allowlist". No interface in the system accepts a command string. Gates are named, versioned definitions with a fixed
argv — you cannot ask for one that isn't in the registry, and you cannot describe one in prose.Unknown means blocked. Evidence tagged with a rule the policy table doesn't recognise blocks the release. Someone adding a rule upstream and forgetting to wire a decision stops a release, rather than silently passing one.
The one that convinced me
The sharpest version is the forged approval.
An approval grant in this system binds three digests: the plan, the repository content, and the target. So I wrote a test that constructs one by hand — all three digests correct, the shape perfect, expiry in the future — and tries to spend it.
It fails.
UnknownGrant.Not because a validator caught the forgery. Because consumption is a lookup in the ledger, and this grant was never recorded there. A perfect forgery fails for the same reason an empty string fails: there is nothing to look up.
That is the difference I want in a security property. Not hard to fake. Not a thing that can be faked, because the check isn't inspecting the object at all.
Where this meets an agent framework
I built the agent on the Strands Agents SDK, and the framework gave me one more place to put a structural boundary:
InterventionHandler, which intercepts a tool call before the tool body is entered.So the guard refuses tool names that must never resolve to anything —
approve_release, issue_approval, run_shell, execute_command — and every refusal lands in the evidence ledger.Which raises a fair question: the session already refuses out-of-order calls, so why guard again?
Because the two layers fail differently. The session refuses because a state transition does not exist. The handler refuses at a point the tool never reaches. A bug in one still meets the other. That is not belt-and-braces; it is two mechanisms with genuinely independent failure modes, which is the only kind of redundancy worth paying for.
The part that surprised me
I expected structural constraints to be more work. Mostly they were less.
Deleting an edge from a transition table is smaller than an
if statement placed correctly in every caller. An enum of seven action kinds is smaller than a path validator plus its tests plus the review comment explaining why os.path.normpath is not enough. And the tests get shorter, because you are asserting that something raises rather than enumerating everything it should reject.The expensive part is earlier, and it is not code. It is deciding what a component is not allowed to express — and then not adding the parameter later when it would be convenient.
I turned down an SDK integration on exactly that basis. Strands ships a
Sandbox abstraction, and adopting it would have been a nice thing to point at. Its single abstract method takes a shell string. My gate runner uses a fixed argv with cwd set to an isolated checkout and no shell anywhere in the path. Taking the integration would have meant introducing a string interface for commands into a system whose main claim is that no such interface exists.Three integration points that carry weight beat four where one is decoration.
Built for the AWS Agents for Humans hackathon with the Strands Agents SDK. The project is ReleaseSentinel: a release agent that audits scope, runs gates against an isolated copy, requires digest-bound human approval, verifies six dimensions after deploying, and writes an append-only receipt. Repository: https://github.com/yyswordsman-CN/release-sentinel
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article