AWS Builder Center

Agents for Humans: deciding where a model is allowed to remember

The safest-looking rule in my agent quietly deleted the most important event in the product. The fix was not to loosen the guard — it was to change what refusal means. What I learned about which promises a model is allowed to keep, and which belong in code.

alzubairi street
I spent a weekend building an agent that finishes the weekly donor report for a neighbourhood mutual-aid group, and the thing I actually learned had nothing to do with prompting.
The product is called BasketBrief. A five-person team distributes relief kits; two donors want to know what happened to their money. The receipts are with one volunteer, the delivery counts are in a message from another, and one receipt is always missing. The coordinator becomes a switchboard: ask, wait, re-ask, add it up twice, write two reports that must not contradict each other.
So I gave the job to a Strands Agents SDK  agent on Amazon Bedrock, and it worked on the first afternoon. Then I tried to break it, and that is where the real design started.

The first thing I got wrong

My guard was simple and felt principled: an amount may only be recorded if it literally appears in its source. A model cannot invent a number into someone's donor report.
Then a correction arrived in the test data:
"Correction: we recounted at the church hall. 88 kits were delivered, not 92. Four more came back."
The agent tried to record delivered=88, returned=12. Twelve is not in the text — it is 8 plus 4 — so my guard refused the whole call. The agent tried delivered=88 alone; now the counts didn't reconcile against the hundred loaded, so a second rule refused that too. Having nothing left it could do, the agent gave up and filed the correction as "needs clarification."
The safest-looking rule in the system had quietly deleted the single most important event in the product.
The fix was not to loosen the guard. It was to change what refusal means: drop the field the source does not state, record the fields it does, and raise the arithmetic gap as something a human has to look at. Losing one unsupported number must never cost you the numbers the source actually gives you.

The thing worth taking away

After that I went through every behaviour I had promised and asked one question about each: is the model allowed to be the one who remembers this?
  • May the model alone decide that an amount is supported? → No. Code requires a stated total or currency-associated amount in the source text. This checks association, not whether the underlying document is true.
  • Are kits the same as households? → No. Code refuses the conflation.
  • Should someone be asked for the missing receipt? → The model chooses who and how to word it. Code checks afterwards that every known gap has an open question and creates a scoped fallback question if it does not.
That last one came from measurement, not theory. Amazon Nova Pro asks the follow-up most of the time. Most of the time is not a promise. So the promise moved into a gate that runs after the model's turn, and the model kept the part that is genuinely a judgement: who has the receipt, and how to ask a volunteer for it without sounding like a collections notice.
The model keeps the interpretation and wording work. Code enforces the tested boundaries for accepted facts, permissions, approval and delivery. Those checks still depend on readable evidence and human review; they cannot prove that a payment or a distribution happened.

What that looks like in the product

A correction arriving after the reports are delivered is the moment the whole thing exists for. When it lands, the figures move, both reports are redrafted, and the approval bound to the delivered version is refused — it cannot be reused for a version nobody read. And the gap is stated in the words that matter:
4 kits unaccounted for: 100 loaded, 88 reported delivered, 8 returned. The number of affected households is not established.
The reconciliation code generates that gap from stored counts. The coordinator reviews it before delivery; the system does not turn missing kits into an invented household count.

BasketBrief is open source under MIT: github.com/NexuChat/basketbrief. There is a live demo at basketbrief.mlki.app with a button that runs the whole journey against the real agent. Every organisation, person and receipt in it is fictional and labelled as such.
Built for the Agents for Humans Hackathon, Good Neighbor Agents track.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article