Agents for Humans: the model orchestrates, the data never passes through it
Five problems from wiring two Strands agents into a real request path on Bedrock: a model drowning in tool results, a model rewriting its own arguments, a 30-second gateway, a thread-pool leak, and prose that hides things. Then the fleet moved to AgentCore.
Second post on building VitaCabinet with the Strands Agents SDK on Amazon Bedrock. The first one was about the agent that holds no tools. This one is about what happened when I made the other two actually do the work. The third is about the data model underneath: a medical record that admits what it does not know .
VitaCabinet reads a medicine drawer — photograph the boxes, and three agents resolve each one against RxNorm, check the FDA enforcement record for live recalls, and write down the question to ask a pharmacist. It is live: https://b5emjsgbi1.execute-api.eu-north-1.amazonaws.com
When I audited my own first build against the judging criteria, I found something uncomfortable. I had defined three Strands agents, tested them, drawn them on the architecture diagram — and the request path a judge would actually exercise ran two REST calls in plain Python and never touched two of them. The one agent that did run was the one with no tools. The SDK was, in the deployed product, a sentence generator.
So I rewired it. Here is what that took, in the order the problems arrived.
Problem one: the model drowned in its own tool results
The first honest version was simple.
check_for_recalls(ingredient) returned the openFDA payload as a dict; the Watchman agent called it once per ingredient; Strands handed each result back to the model.Amlodipine has thirteen live recalls on the enforcement record. Each is a paragraph. The Watchman's second call blew straight through Nova Lite's output budget:
1
2
strands.types.exceptions.MaxTokensReachedException:
Model stopped generating due to maximum token limit.The model was being asked to carry the data, and the data was never its job. Its job is deciding what to call next.
The fix is a split I now think every tool-using agent should have. A tool returns one sentence to the model — enough to decide the next step — and writes the full structured result to a ledger the application reads afterwards:
1
2
3
4
5
6
7
8
9
10
def check_for_recalls(ingredient: str) -> str:
live = fda.recalls(ingredient)
ledger().recalls[ingredient] = live # the app reads this
if not live:
return f"{ingredient}: no live recalls on the FDA enforcement record."
newest = live[0]
return (f"{ingredient}: {len(live)} live recall(s). Newest: a batch of "
f"{newest.product.split(',')[0]} on {newest.date}, lots {newest.lots}. "
f"Report batches and lots, never 'your medicine'.") # the model reads thisThirteen recalls became one line in the context window. The Watchman went from a crash to 5.5 seconds for five ingredients — and, because its context stayed small, its final report got better: it now says things like "if you are taking only metformin and not the combination product, you may not be affected." That sentence was written by a model that had read a summary, not a dump.
Problem two: the model rewrote its own arguments
With the ledger in place, the Identifier agent worked beautifully — one
identify_medicine call per box, then find_duplicate_medicines across all of them. Except the duplicate finding, the one thing the product exists for, vanished.The trace explained it. Nova had "helpfully" normalised the box texts before passing them back:
1
find_duplicate_medicines(box_texts=["metformin hydrochloride 500 mg", ...])Glucophage 500mg and Metformin 500 mg — a brand and its generic, the pair I needed — had been collapsed into one string before the tool ever saw them. The tool compared a list with no pair in it, and reported none.You cannot prompt this away reliably. What you can do is stop depending on the model relaying inputs faithfully.
find_duplicate_medicines now compares whatever is already in the ledger — every box the Identifier has read this session — and its argument only adds. And the orchestrator recomputes duplicates from the ledger after the agent finishes, unconditionally. Cheap, deterministic, and immune to the model's paraphrasing.The principle generalises: findings come from tool results, never from the model's prose or the model's arguments. The model decides when to look. The tools decide what was found.
Problem three: API Gateway gives you thirty seconds
Two agents reading seven boxes take 15–35 seconds, depending on cold starts. API Gateway HTTP APIs return
503 at thirty. No configuration changes that.I resisted the obvious fix — a spinner over a synchronous call that sometimes works — because it would also have been the worse demo. A reading became a job:
POST /scanwrites the boxes to DynamoDB, invokes the same Lambda asynchronously with{"job": id}, and answers in a few hundred milliseconds.- A Strands hook on
AfterToolCallEventwrites each tool call — name, arguments, what the tool said, duration — to the job row as it happens. - The page polls
GET /jobs/{id}every 600 ms and draws the trace as it grows.
1
2
3
4
5
6
7
8
9
class Trace(HookProvider):
def register_hooks(self, registry, **_):
registry.add_callback(AfterToolCallEvent, self.after)
def after(self, ev):
step = {"agent": self.agent_name, "tool": ev.tool_use["name"],
"input": ev.tool_use["input"], "said": tool_text(ev.result)[:240],
"ms": elapsed_ms(ev)}
store.job_step(self.job_id, step) # DynamoDB list_appendThe person watching sees
identify_medicine("Glucophage 500mg") → 2909 ms → metformin hydrochloride 500 MG Oral Tablet (RxCUI 861008) appear line by line. It is the most persuasive thing on the screen, and it exists because of a timeout.Problem four: a leak that only showed up in one test order
The ledger was a
contextvars.ContextVar. Clean, per-request, idiomatic. One test failed — only when the whole suite ran, never alone.Strands'
ConcurrentToolExecutor runs tools on a thread pool. If a pool thread ever lazily created its own ledger, it kept it, and the next reading's tool calls on that thread landed in a stale ledger. The duplicate vanished again, for a different reason, in a way that depended on which thread picked up which call.The fix was less clever than the bug: a module-level ledger guarded by a lock. Each Lambda invocation is its own process; locally, readings serialise. I wrote down why in the code, because the ContextVar looks more correct and the next person will want to put it back.
Problem five: the agent's report was prose, and prose hides things
With the ledger holding the truth, the Identifier's final message was still a paragraph — nice to read, impossible to check. Strands has
structured_output, so after the tool loop the same agent is asked for its conclusion as a typed object:1
2
3
4
5
6
7
class DrawerReport(BaseModel):
boxes_read: int
unreadable: list[str]
duplicate_pairs: list[list[str]]
one_line: str
report = identifier.structured_output(DrawerReport, "Report what you found. No advice.")A model that has to fill in
unreadable cannot bury an unreadable box in a sentence. And now a test can compare the agent's account against the ledger — a report that disagrees with what the tools found is a report worth knowing about.One live surprise: a drawer with nothing unreadable came back with
unreadable: null, and the report failed validation because the drawer was fine. A field_validator(mode="before") that maps None to [] fixed it. Small models write null for "nothing here"; your schema should expect that.Then I moved the agents to AgentCore
With the agents doing real work, the last step was hosting them where AWS suggests: Amazon Bedrock AgentCore Runtime. The starter toolkit builds the ARM64 container in CodeBuild, so no Docker on the laptop — which mattered, because there is none.
1
2
agentcore configure -e agentcore_entry.py -n vitacabinet -rf requirements.txt
agentcore deploy --env VITACABINET_TABLE=vitacabinetThe entrypoint is twenty lines around the same
read_drawer() the Lambda used. The one design point worth stating: when the payload carries a job id, the runtime writes the trace to DynamoDB itself, from inside AgentCore, so the page keeps drawing the agents thinking even though the reading no longer runs next to it. The Lambda's job handler became "invoke the runtime, wait, record the result." Same agents, same tools, same ledger, same trace — one code path, two places it runs.All three agents run there now — the Scribe too, so no agent runs in a different place from its siblings — with OpenTelemetry enabled, so every reading leaves spans in CloudWatch.
Two things went wrong that will save you an hour: the container failed to start because
bedrock-agentcore was not in requirements.txt (the toolkit installs your requirements, not its own), and the auto-created execution role needed an inline policy for the table before the trace could be written from inside the runtime.What I would tell you
- Tools should tell the model one sentence and tell the app everything. The model's context is for deciding, not for carrying.
- Never derive findings from the model's prose or the model's arguments. Recompute from what the tools actually returned.
- A timeout can be a feature. The async job with a streamed trace is a better product than the synchronous call would have been.
- Order-dependent test failures in agent code are usually shared state on executor threads. Look there first.
- Ask the agent for its conclusion as a type, not a paragraph. Then check it against the tools. And expect
nullwhere you asked for an empty list.
The code, Apache 2.0, with 67 tests that run against the live models and public APIs: https://github.com/bayraktartahsin/vitacabinet
VitaCabinet does not tell anyone what to take, what to stop, or what to throw away. It finds what is uncertain in a drawer, keeps watching it, and writes down the question to ask a pharmacist.
# amazon-bedrock# amazon-bedrock-agentcore# generative-ai# serverless# healthcare-life-sciences-industry
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article