AWS Builder Center

Three Strands extension points, and the fourth I turned down — Agents for Humans

Ten tools that take almost no arguments, a handler that denies before the tool body is entered, and input schemas generated from type hints. Plus the sandbox abstraction I turned down, because its one method takes a shell string and the whole claim is that no interface accepts one.

Agent safety and release evidence
The interesting decision I made with the Strands Agents SDK was not which features to use. It was which one to leave alone.
The project is a release agent: it takes "ship this change safely to sandbox" and turns it into scope audit, isolated test runs, digest-bound human approval, deployment, six-dimension verification, and an append-only receipt. The model's job inside that is narrow on purpose — understand the request, move through the stages, read the evidence, classify what went wrong.
Three SDK extension points carry real weight in that design. A fourth would have been a nice thing to point at, and taking it would have made the system worse.

1 · The tool surface is the whole surface

Ten @tool functions. That is everything a model can reach.
Nine of them take no arguments at all:
1
2
3
4
5
6
def build_release_tools(session: ReleaseSession) -> list:

@tool
def inspect_release_scope() -> str:
"""Read the repository and report what has changed since the baseline."""
...
They close over one ReleaseSession, constructed before the agent exists, which already holds the repository, the target, the environment registry and the approval ledger. So a model advances a release without ever naming a path, a digest, an environment or an approver — not because those arguments are validated, but because they are not arguments.
The single tool that takes anything is request_human_approval(justification: str). It takes the model's reasoning, which a person reads. It does not take its authority.
There is no tool that approves a release. Not a disabled one, not one behind a flag. There isn't one.

2 · InterventionHandler is a second, independent refusal

Strands lets you intercept a tool call before the tool body is entered. My handler refuses three things:
  • Tool names that must never resolve to anything: approve_release, issue_approval, grant_approval, run_shell, execute_command. A model reaching for any of them is denied and the denial is recorded as evidence.
  • Any call once the release is in a terminal state.
  • Writing tools when policy has not been evaluated, when policy said BLOCK, or when no human grant has been consumed.
The obvious objection: the session already refuses out-of-order calls. Why guard twice?
Because the two fail differently. The session refuses because a state transition does not exist — that is correctness. The handler refuses at a point the tool function never reaches. A bug in one still meets the other. Two mechanisms with independent failure modes is worth something; two implementations of the same check is not.
One correction I had to make here, found while wiring up a runner script: my terminal-state rule was refusing everything after a release concluded — including summarise_release, the tool whose entire job is to explain what happened. A finished release that cannot be explained is the opposite of what an evidence-backed system is for. So there is now exactly one exemption, and it is narrow: a tool that reads the ledger and writes nothing, to the target or to the release. The test asserts that the exemption set contains that one name and does not intersect the writing tools.

3 · Tool specs come from the type hints

Strands derives each tool's input schema from its signature and docstring:
1
2
3
{"name": "request_human_approval",
"inputSchema": {"json": {"properties": {"justification": {"type": "string"}},
"required": ["justification"], "type": "object"}}}
Which means parameters are schema-checked by the framework before my policy layer is consulted. My project contract had a line requiring exactly that, written before I had chosen a framework. Getting it from the SDK rather than hand-rolling it is a real saving, and I can assert against the generated spec in tests rather than against a list I maintain:
1
2
3
every_parameter = {p for tool in surface.values() for p in tool["parameters"]}
for forbidden in ("path", "repository", "environment", "target", "approver", "command"):
assert not any(forbidden in p for p in every_parameter)
That test reads the actual tool specs the SDK generated. If someone adds a parameter named target_path in six months, it fails.

4 · The one I turned down

Strands ships a Sandbox abstraction, with PosixShellSandbox as a ready-made implementation. Sandboxed execution is exactly the vocabulary of my problem, and "uses the SDK's sandbox" would have been a good line in a submission.
Its single abstract method:
1
def execute_streaming(self, command: str) -> ...
It takes a shell string.
My gate runner does not have a shell in it. Gates are named, versioned definitions with a fixed argv, run with cwd set to an isolated checkout:
1
2
3
4
5
6
7
PYTEST = GateDefinition(
gate_id="pytest",
version="1",
argv=(PYTHON, "-m", "pytest", "-q", "--no-header"),
timeout_seconds=300,
skip_exit_codes=frozenset({5}), # pytest exits 5 when it collected nothing
)
The whole claim of this system is that no interface anywhere accepts a command string. Adopting Sandbox would have meant building one, and then arguing that it was fine because I only ever put safe strings into it. That is the procedural version of a guarantee I had already implemented structurally, and I would have been trading down.
So I did not use it, and I wrote down why — in the repository, next to the code, under a heading that says what was deliberately not done. That file has turned out to be one of the more useful things in the project.

What I would tell someone starting

Count integration points that carry weight, not integration points.
The three above are load-bearing: remove any one and the system is meaningfully weaker. A fourth would have been a logo. When an SDK feature and your architecture disagree about something as basic as does this interface accept a command, the SDK feature is not free — it is a hole with a nice API on it.
And write down the refusals. "We considered X and here is the specific reason we did not take it" is more informative to a reader than another paragraph about what you did build, and it stops you relitigating the same decision in three weeks.

Built for the AWS Agents for Humans hackathon with the Strands Agents SDK. The project is ReleaseSentinel: ten layers, an append-only evidence ledger, and 54 asserted eval cases across eight failure families. Repository: https://github.com/yyswordsman-CN/release-sentinel
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article