An agent that refuses to guess — what Agents for Humans taught us about auditing our own tool
Five things hold the refusal, each a measured failure before it was a rule. Plus the things we got wrong along the way, kept in git rather than hidden.
Some decisions come back every single day, and the data that should settle them is scattered enough that people end up going on instinct. For a Korean produce farmer the question is "which of the 33 public wholesale markets should I ship to today?" Public auction data exists. Tools that read it exist. We built one of those tools ourselves — and this post is about what we found when we audited it, and what we built for the Agents for Humans hackathon instead.
The tool we already had answered on 0.77% of the day
Our earlier MCP server for the same data fetched the API's first page — 1,000 records — before filtering. The settlement day we measured (2026-08-28) had 129,536 records. That slice covered 6 of the 33 markets, one market was 64.3% of it, and Seoul Garak — alone 25.6% of the day — was not in it at all.
It still answered. Confidently. That was the problem.
We audited only our own tool. We make no claim about anyone else's.
What "refuses to guess" means in code
The hackathon agent is built on the AWS Strands Agents SDK. The loop is Strands; the decision is not left to the model. Five things hold it:
- It measures its own evidence before answering — coverage (what fraction of the day did we actually see), market count, records per market.
- When the evidence is thin, it says exactly why. Not "insufficient data" — it names the ratio, the markets that fell short, and what would fix it.
- It normalizes units. Auction prices come per package (4 kg, 5 kg, 10 kg boxes). Averaging raw prices compares a small box to a big one — that flipped "which market pays most" in 9 of 16 date-by-product pairs we measured.
- It reports the median and flags outliers. One record at 500,005 won/kg among 656 moved a market's mean by 21%.
- The verdict is computed in code, not asked of the model. A prompt that says "be careful" is a hope. A tool that returns insufficient is a constraint — and the numbers below say how much still got past it.
Three runs, three different answers
- Peaches, Aug 28 — which market? Wonju, 4,183 won/kg, from 184 records across 32 markets.
- Cabbage, Aug 30 — which market? Cannot answer. Zero cabbage auction records nationwide that day.
- Onions, Aug 28 — which market? Cannot answer. Not in this sample, which is not the same as not existing.
The second and third are the product. The interface is Korean because the user is.
What we had to fix to make the refusal hold
Each of these was a measured failure before it was a rule:
- The refusal screen carries no prices. We first showed them labelled "reference only, do not cite". The model cited them anyway in 4 of 6 runs. Labels are requests; removing the numbers is a constraint. After: 1 of 6.
- A tool failure becomes a refusal, never silence. When the tool crashed on an empty day, the model filled the gap with markets that do not exist.
- The answer line is written by code. Letting the model phrase it produced stray Vietnamese and Han characters in about 2 of 6 runs. The shipped interface prints the tool's line: 12 of 12 demo runs, no foreign script. The garbling was not fixed — it is not displayed. We keep both numbers.
- The copied line carries no statistics. A headline that read "only 34.97% (45,301/129,536) of the day" got spun into a table of per-date counts the model never queried. Plain words now; 0 of 6 runs invented a number after the change.
- The model cannot quietly change the question. It once passed a garbled product name into the tool and got a perfectly formatted refusal about a word the user never typed. The agent now compares the product it passed against what the user wrote.
Things we did wrong along the way, kept in git rather than hidden
- We reported "median attempts is still 1, so the retry change didn't slow it down." True of attempts. Wall-clock for the three-scene demo ran 92 s once and 209 s the next time (n=2). Answering a wall-clock question with an attempt count nearly planned a video around the faster figure.
- A teammate's demo answered where ours refused — their machine had a full day cached, ours had the bundled sample. Both were correct; the report was not. The agent now pins itself to the bundled file.
- Our publish-safety checker exempted itself and reported a clean run — right up until the thing it exempted was the thing that leaked. A check that exempts itself is not a check.
Limits we are not hiding
- n = 16 date-by-product pairs, n = 6 model runs. We say "often", not percentages-as-rates.
- Full-day caches cover 2 days. Products were chosen by us. The market-bias result is one day.
- Sufficiency thresholds (at least 2 markets, at least 3 records each) are chosen, not derived.
- Model probes are a local 3B model running on Ollama on one machine. The model provider is pluggable in Strands, and Strands defaults to Bedrock when you do not name one — we always name one explicitly. We did not deploy this to Bedrock or AgentCore during the hackathon window, so we have no measurements on that path and make no claim about it.
Try it
Repository: https://github.com/SongT-50/agent-that-refuses-to-guess — MIT licensed.
Demo video (4:47): https://www.youtube.com/watch?v=ckOxHll3XUA
Three demo scenes run without an API key on a bundled sample that states its own limits. Every number in this post has a script in the repository, mapped to the claim it supports and the sample size behind it.
#AgentsForHumans #StrandsAgents #AIAgents #AWS
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article