
Agents for Humans: Why Seven Protected Effects Became Three Human Decisions
Explain Authority Cut Sets, semantic authority atoms, policy-defined bundles, exact minimum cover, readiness and the external-human authority boundary.
Designing Authority Cut Sets for professional agents with Strands Agents
When I started building Authority Cut for the Agents for Humans Hackathon, the obvious design was also the wrong one.
The workflow was vendor onboarding. An agent could collect documents, run checks, prepare an ERP record, configure payment settings and eventually prepare a first payment. Some steps were routine. Others clearly needed human authority.
The easy solution was to put a human approval in front of every protected action.
That looked safe at first. It also created seven separate approval moments in one small workflow.
I did not want an agent that asked, "May I run this tool?" seven times. I wanted it to ask the human the real business question behind those tools.
That became the starting point for Authority Cut.
The problem is not just whether to ask a human
Human-in-the-loop patterns already exist. Strands Agents supports intervention and approval patterns, and many agent systems can pause before a sensitive tool call.
The harder design question for me was different:
What is the smallest meaningful decision surface that should be shown to the human right now?
A professional operator does not naturally think in terms of tool calls.
A procurement manager does not want to approve:
activate vendor
sync ERP
enable purchasing
configure payment profile
set payment terms
prepare remittance
transmit first funds
Those are implementation effects.
The human thinks in semantic decisions such as:
Is this vendor risk exception acceptable?
May this vendor be enabled for payments?
May the first irreversible transfer actually be released?
I wanted the agent to keep those business decisions separate from its own execution tools.
Authority atoms instead of broad standing permission
In Authority Cut, protected effects declare the authority atoms they require.
A policy then defines valid decision bundles that may grant those atoms.
The runtime does not invent a broader approval just to reduce the number of prompts. It can only choose from the policy-defined bundles.
This distinction matters.
If a policy says "vendor exception" and "bank change" can be reviewed together as vendor-risk, the runtime may surface that bundle.
If the policy does not define a bundle that combines "vendor risk" with "first funds", the runtime cannot silently create one.
The goal is not fewer prompts at any cost.
The goal is the smallest valid semantic cover allowed by policy.
Computing the current Authority Cut
The core graph method is intentionally simple and inspectable.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
def minimum_authority_cut(self, state):
needed = self.unresolved_authorities(state)
if not needed:
return []
candidates = [
b for b in self.bundles.values()
if set(b.grants) & needed
]
for size in range(1, len(candidates) + 1):
valid = []
for combo in combinations(candidates, size):
covered = set().union(
*(set(b.grants) for b in combo)
)
if needed <= covered:
valid.append(combo)
if valid:
choice = min(
valid,
key=lambda c: tuple(
b.bundle_id for b in c
)
)
return list(choice)For the small hackathon workflow I used an exact search over the policy-defined bundles.
That gives me a useful property: minimality is not a model guess. It is deterministic and testable.
It is also deliberately bounded. I only claim exact minimality over the bundles supplied by the policy. I am not claiming that the policy itself is globally optimal.
A decision should not become actionable too early
Compression alone is not enough.
A human should not be asked to make a decision before the evidence needed for that decision exists.
Each semantic bundle can therefore define prerequisites.
For example, the
The
In the control plane, a human approval is rejected if its prerequisites have not executed:
That gives me a useful property: minimality is not a model guess. It is deterministic and testable.
It is also deliberately bounded. I only claim exact minimality over the bundles supplied by the policy. I am not claiming that the policy itself is globally optimal.
A decision should not become actionable too early
Compression alone is not enough.
A human should not be asked to make a decision before the evidence needed for that decision exists.
Each semantic bundle can therefore define prerequisites.
For example, the
vendor-risk decision is not ready until both the tax check and the bank check have executed.The
first-funds decision is not ready until the remittance preview exists.In the control plane, a human approval is rejected if its prerequisites have not executed:
1
2
3
4
if approved and not self._bundle_ready(bundle):
raise ValueError(
"decision bundle is not ready"
)When a decision is recorded, the runtime also binds it to the relevant prerequisite receipts.
This gave me a cleaner human interface:
execute safe work first
collect the evidence
surface only the semantic decision that is actually ready
keep future decisions hidden or not-ready
resume only the effects covered by the recorded authority
Keeping authority outside the model tool surface
The Strands agent in the public Authority Cut implementation has exactly three model-callable tools:
This gave me a cleaner human interface:
execute safe work first
collect the evidence
surface only the semantic decision that is actually ready
keep future decisions hidden or not-ready
resume only the effects covered by the recorded authority
Keeping authority outside the model tool surface
The Strands agent in the public Authority Cut implementation has exactly three model-callable tools:
1
2
3
execute_safe_vendor_work
get_authority_cut
execute_authorized_vendor_workThere is no approve tool.
There is no revoke tool.
The model can execute safe work, inspect the current decision surface and resume work that the control plane says is authorized.
The human principal changes authority through a separate control path.
This is an important boundary for me. The same agent that wants to execute a protected effect should not also be given the model-callable capability to mint the permission that makes the effect legal inside the workflow.
What happened in the fixed workflow
The controlled vendor-onboarding graph contains:
There is no revoke tool.
The model can execute safe work, inspect the current decision surface and resume work that the control plane says is authorized.
The human principal changes authority through a separate control path.
This is an important boundary for me. The same agent that wants to execute a protected effect should not also be given the model-callable capability to mint the permission that makes the effect legal inside the workflow.
What happened in the fixed workflow
The controlled vendor-onboarding graph contains:
1
2
3
5 safe actions
7 protected effects
3 semantic human authoritiesThe simple baseline is one approval per protected effect, which means seven human decisions.
Authority Cut surfaces three semantic decisions.
That is:
Authority Cut surfaces three semantic decisions.
That is:
1
(7 - 3) / 7 = 57.14%fewer approval decisions than that specific baseline in this fixed workflow.
I am deliberately not presenting 57.14% as measured productivity improvement. It is not a customer study and it is not a universal HITL benchmark.
It is a controlled mechanism result.
What matters more than the percentage is the shape of the interaction:
I am deliberately not presenting 57.14% as measured productivity improvement. It is not a customer study and it is not a universal HITL benchmark.
It is a controlled mechanism result.
What matters more than the percentage is the shape of the interaction:
1
2
3
4
5
6
7
8
routine work runs
-> evidence becomes available
-> vendor-risk becomes ready
-> human decides
-> authorized work resumes
-> payment-release becomes ready later
-> human decides
-> first-funds becomes ready only after remittance evidence existsThe operator sees business decisions instead of a stream of tool-level prompts.
Why Strands Agents was useful here
I wanted the control mechanism to sit around a real agent loop rather than a simulated orchestration script.
The public judge path constructs a real Strands Agent and runs the workflow through the published tool surface.
The deterministic custom Strands model provider makes that public proof reproducible and credential-free. It does not turn the control logic into model logic. The policy and state transitions remain explicit Python code.
That separation made testing much easier.
I can test:
the graph
the authority bundles
decision readiness
the model-callable tool boundary
execution state
later correction behavior
without treating model behavior as the source of truth for authorization.
What I learned
The main lesson from this part of the build is that "human in the loop" is too broad a design goal.
For professional agents, I think there are at least three separate questions:
What can the agent do autonomously?
What exact semantic authority does the human need to provide now?
What happens to downstream work if that authority changes later?
Authority Cut Sets address the second question.
The next part of the project addresses the third one, because an approval is not necessarily permanent.
A human can change their mind.
When that happens, logging the correction is not enough. The execution state has to change too.
Try the project
Live judge path:
https://evidencebound-authority-cut.vercel.app
Public source:
https://github.com/moneyparking/evidencebound-authority-cut
Run Run live Strands judge path to execute the four-phase vendor-onboarding proof.
The project is intentionally scoped. It does not claim to invent HITL, approval workflows or agent safety in general. The contribution I am testing is the combination of a policy-bounded minimum semantic authority surface with correction-aware downstream execution.
Why Strands Agents was useful here
I wanted the control mechanism to sit around a real agent loop rather than a simulated orchestration script.
The public judge path constructs a real Strands Agent and runs the workflow through the published tool surface.
The deterministic custom Strands model provider makes that public proof reproducible and credential-free. It does not turn the control logic into model logic. The policy and state transitions remain explicit Python code.
That separation made testing much easier.
I can test:
the graph
the authority bundles
decision readiness
the model-callable tool boundary
execution state
later correction behavior
without treating model behavior as the source of truth for authorization.
What I learned
The main lesson from this part of the build is that "human in the loop" is too broad a design goal.
For professional agents, I think there are at least three separate questions:
What can the agent do autonomously?
What exact semantic authority does the human need to provide now?
What happens to downstream work if that authority changes later?
Authority Cut Sets address the second question.
The next part of the project addresses the third one, because an approval is not necessarily permanent.
A human can change their mind.
When that happens, logging the correction is not enough. The execution state has to change too.
Try the project
Live judge path:
https://evidencebound-authority-cut.vercel.app
Public source:
https://github.com/moneyparking/evidencebound-authority-cut
Run Run live Strands judge path to execute the four-phase vendor-onboarding proof.
The project is intentionally scoped. It does not claim to invent HITL, approval workflows or agent safety in general. The contribution I am testing is the combination of a policy-bounded minimum semantic authority surface with correction-aware downstream execution.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article