AWS Builder Center
Agents for Humans: Six Lessons From Building a Bounded Agent With the Strands Agents SDK

Agents for Humans: Six Lessons From Building a Bounded Agent With the Strands Agents SDK

Six practical lessons from building and testing a bounded Strands agent: narrow tools, explicit authority, durable state, safe events, confirmation gates, and tests that exercise the real loop.

Series: Agents for Humans: Building ST-Agent (3 articles)

  1. 3
    Agents for Humans: Six Lessons From Building a Bounded Agent With the Strands Agents SDK This article
By the third post of an Agents for Humans series, you've earned the right to skip the pitch. ST-Agent — our hackathon entry, an agent that turns story ideas into SillyTavern character cards and lorebooks — is built and tested, the repo is public, and the previous post  covered our delivery gate.
This post is the retrospective: six things we learned while building a real, non-trivial agent with the Strands Agents SDK, in the order we learned them.

1. Tools with one job each beat tools with options

One early design temptation was a large tool with several modes and enough optional arguments to cover every case. It read like an API for a human developer, and that was the problem: the caller was a model.
We used eight narrow tools instead — save_brief, save_content_document, build_character_card, build_lorebook, build_png_card, check_official_sources, validate_deliverables, and finish_case — each with one job, a small typed payload, and a structured return. Tool selection became easier to reason about, and the results gave our tests something solid to bite on.

2. Give the model a lane, then fence the lane

Strands makes it easy to give an agent powerful tools. That does not mean every useful capability belongs in the tool set.
Ours gets filesystem access scoped to one user-chosen case folder, no shell, no arbitrary filesystem, and no code execution. Network access is limited to the configured model provider and a project-owned checker for a small allowlist of official format sources. Uploads are identified from their bytes and size-capped before they enter the workspace.
The agent is more useful, not less, when users can trust the blast radius: the model cannot call what we did not bind.

3. Persist before you think, not after

A live model turn can take a while, and any number of things can end it: a laptop lid, a Wi-Fi blip, or a provider timeout. Our recovery rule came from a decision we made before writing code:
The files are the memory.
Accepted user input is written before the model runs. Tool-driven state changes are persisted as they happen, and important writes are atomic. The controller loads the case manifest from disk instead of treating UI session state as the authority. When a turn stops, the next one resumes from the last complete state.

4. Sanitized events are a feature, not overhead

Our UI shows live tool activity while the agent works — “Organizing the creative brief...” and “Character card built.” — driven by Strands hook events and run through a sanitizer that strips prompts, payloads, credentials, and reasoning.
That produced two payoffs we did not fully predict. Users get a process they can feel moving instead of a spinner and vibes. Reviewers — judges, in our case — get an honest, inspectable trail in a collapsed technical view.
An agent that shows its real steps earns more trust than one that performs its steps.

5. Confirm-before-write is cheap and changes the feel entirely

The agent can organize a brief and ask questions, but the artifact-building tool chain refuses to proceed until the user has explicitly confirmed the plan with a plain “yes” or “confirm.”
That single gate transformed the product from “it does things at you” into “it works with you.” It also made the questions feel useful instead of ceremonial: the answers become a brief the user actually approves.

6. Test the loop, not just the units

The Strands model abstraction let us inject scripted model responses into the same controller and tool loop used in production.
Unit tests cover parsers, validators, file boundaries, and codecs. Scripted-model tests cover the actual behavior: questions persisted, confirmation enforced, tools sequenced, failures consumed, delivery gated, and interrupted cases resumed. The release suite now has more than 200 passing tests, and Windows CI runs it on every push.
When we changed model configurations during the project, the suite helped us separate provider-specific failures from broken product contracts without guessing from one long live run.

Wrapping up

The through-line: Strands is at its best when you let the model drive and make your program verify.
Tools with narrow jobs, a fenced lane, files as memory, honest events, an explicit confirmation, and tests that exercise the loop — none of these fight the framework. They are the guardrails that let us let it run.
ST-Agent is open source under the MIT License: github.com/Lockyer228/ST-Agent 
If you're racing on your own Agents for Humans entry right now: good luck, and put a gate on your “done.”
Written as part of our entry for the Agents for Humans hackathon.

Series: Agents for Humans: Building ST-Agent (3 articles)

  1. 3
    Agents for Humans: Six Lessons From Building a Bounded Agent With the Strands Agents SDK This article
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article