AWS Builder Center

Building for Agents for Humans: what we learned by auditing our own tool first

We measured our own published MCP server before building the agent. It was answering "nationwide market comparison" on 0.77% of a day. That measurement decided the product: an agent that measures its own evidence and refuses when it is not enough.

We build software for a wholesale produce market in Korea. Every morning, farmers there make the same decision — which of 33 public markets to ship to — and the data that should settle it is public, free, and almost never actually checked. Checking it properly costs more time than the decision is worth, every day, forever. So people go on instinct.
That is the decision we pointed an agent at for the Agents for Humans hackathon, using the AWS Strands Agents SDK. This post is about the part we did not expect: the most useful thing we built was not the agent. It was the measurement we ran on ourselves before we started.

We audited a tool we had already published

Months before the hackathon we had released an MCP server that reads the same government auction feed. It worked. People used it. It answered questions about "which market pays most" with a confident, well-formatted comparison.
Before wiring anything up, we measured what that server actually sees on a single day. On 2026-08-28:
  • Auction records in that one day: 129,536
  • What the tool fetched before filtering by product: 1,000
  • Distinct markets in that slice: 6 of 33
  • Markets it structurally never reaches: 26 of the 32 that traded
The feed is paginated. The tool read page one. Page one is not a random sample of the day: it is sorted, so it is dominated by a handful of markets. A single market was 64.3% of it. Seoul Garak, the largest market in the country and 25.6% of that day's records on its own, never appeared at all.
The tool was answering "nationwide market comparison" on 0.77% of the day, and nothing in its output said so.
We make no claim about anyone else's tools. We measured exactly one, and it was ours. What it cost us was a day of believing the data was thin when the data was fine and the tool was thin.
That measurement decided the product. The agent we built for this hackathon is not a better ranker. It is a ranker that measures its own evidence first and refuses when the evidence is not good enough.

What "refuses" means in practice

Three demo scenes, three different outcomes.
Peaches, Aug 28 — which market is better? Wonju, 4,183 won/kg, from 184 records across 32 markets.
Cabbage, Aug 30 — which market is better? Cannot answer. Zero cabbage auction records nationwide that day.
Onions, Aug 28 — which market is better? Cannot answer. Not in this sample, which is not the same as not existing.
The second and third are the product. The difference between them matters more than it looks: one is an absence we verified, the other is an absence we could not check. Collapsing those two into "no data" is how a tool starts lying politely.

Where Strands fit, and where we deliberately kept it out

Strands drives the loop: the model reads the question, picks a tool, and fills in its arguments, turning "peaches, 28th" into a product and a settlement date. That is genuinely the part a model is good at, and the SDK made it short.
Everything after that is code. The evidence gate computes coverage, normalizes prices to won per kilogram, detects outliers and package-class mismatches, and returns sufficient or insufficient. The sentence the user reads is written by the tool, not by the model.
That split was not stylistic. We measured our way into it, and each step has a number:
  • A refusal screen must not show the numbers. We first listed prices under "reference only, do not cite." The model cited them in 4 of 6 runs. A label is a request; removing the numbers is a constraint. After the change: 1 of 6.
  • The line the model copies must carry no statistics. Our refusal headline once quoted a coverage ratio. The model spun those digits into a table of per-date record counts it had never queried. Plain words instead: 0 of 6.
  • The answer sentence is written by code. When we let the model rephrase it, Vietnamese and Chinese characters appeared in Korean output in 4 of 6 runs, and one run relabelled a quantity as a record count.
One practical note for anyone starting with Strands: it defaults to Amazon Bedrock if you do not name a model. We always name one explicitly, which let us develop against a local Ollama model at zero cost. The provider is one environment variable, so the Bedrock path needs no change to the agent code — but we have not run it yet, and every number in this post was measured on the local model.

The bug that taught us the most

Late in the build we recorded the demo video and watched a product name come out garbled on screen. The refusal itself was correct, well-formatted, and about a word the user had never typed.
We had guarded what the model writes. We had not guarded what the model passes.
Every constraint we had built sat on the output side: the tool writes the sentence, the refusal carries no prices, the copied line carries no digits. None of them looked at the arguments going into the tool. So a corrupted product name flowed through the entire gate and came out wearing the formatting of a verified answer.
That is the failure mode we had been bragging about catching, showing up in our own demo reel.
The fix is small: compare the product the model passed against the question the user actually wrote, and refuse if they differ. Sixteen new controls hold it in place, on top of the 59 we already had.
The lesson is not "add another guard." It is that the scope of a control and the scope of the failure are two different things, and it is easy to check only the first. A well-formatted wrong answer is more dangerous than an obvious one, because everything about its shape says it was checked.

What we are not claiming

We measured 16 date-by-product pairs. That is small, so we say "often flips," not a percentage. Our sufficiency thresholds (at least 2 markets, at least 3 records each) are chosen, not derived. The market-bias measurement is one day. The full-day cache covers two days.
And the feed itself moves. Re-running the same script a few days later returned 129,704 records for the same settlement date, not 129,536: the source revises retroactively. The ratios did not change, but the total did. So we publish the number and the day we measured it, because a reviewer who re-runs our script will not get our figure back and should not have to wonder why.
Writing that down was not a concession. For a project whose entire pitch is "know what you actually saw," a number without its measurement date would have been the same mistake we built the thing to prevent.

Try it

The repository runs offline. No API key, no cloud account: a compact extract of one real day ships with it, and the three scenes above run against a local 3B model. Budget one and a half to three and a half minutes: we measured 92 s and 209 s on two runs. The variance is the model's, not the code's, and the measurement is in the repository.
https://github.com/SongT-50/agent-that-refuses-to-guess
MIT licensed, with every claim above mapped to the script that produces it.
Built for the Agents for Humans hackathon with the AWS Strands Agents SDK. The interface and sample output are in Korean because the user is a Korean farmer; that was the point, not an oversight.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article