What Agents for Humans taught us about adding a surface: our guards did not follow the agent to the web
We added a web UI to an agent that already had 75 controls. An adversarial review found four ways the guards did not reach the new surface, then a fifth: the date guard only understood the formats a developer types, not the ones our Korean users write.
We spent most of this hackathon on one idea: when the evidence is thin, the agent should say so instead of guessing. By the end of last week we had that working on the command line, with 75 controls holding it in place and a demo video to show it.
Then we added a web UI. Same agent, same tool, same verdict — the page calls the exact function the CLI calls, and there is no second code path. We thought that was the whole story. An adversarial review of the new surface found four ways it was not, and then a fifth after we fixed those.
This post is about the four, because the shape they share is the useful part.
What we thought we had built
The agent's design is a split. The model reads your question and picks a tool. The tool computes the verdict deterministically and writes the sentence you read. If the right tool never ran, the agent refuses. Every guard we had written sat on that split: the tool writes the sentence, the refusal carries no prices, the model cannot quietly change the product it was asked about.
The web page reuses that split. It renders the same verdict line and builds its evidence table from the very object the tool computed on that call. Nothing is recomputed for display.
That is all true, and it was not enough.
The four things the review found
One. The refusal screen showed a price the model had invented. When the model failed to call any tool, the agent correctly refused. But the refusal text still appended whatever the model had written, and the page drew that line as the headline. We had already guarded one field, so the number was gone from the structured response — and it walked back in through the human-readable one. A synthetic run put "999999 won/kg" on a screen whose entire purpose is to say we have no grounds to answer.
Two. A failed retry left the previous attempt's answer behind. The agent retries with a fresh session when a run goes wrong. State was cleared once, at the start of the whole request, not at the start of each attempt. So a first attempt that produced a recommendation and then failed validation could leave its headline and its evidence sitting there while the second attempt errored. The status said "no answer." The page still had a price to draw.
Three. Two tool calls in one turn became one answer with mismatched parts. The headline came from the first result in the message; the evidence came from whichever call wrote last. In the reviewer's reproduction the screen said "cabbage, market B is highest" over a table of onion records. Both halves were real. They were from different questions.
Four. The model could look up a different day than the question named. We had built the guard that compares the product the model passed against what the user wrote — the one we had written a whole section about. We had not built the same guard for the date. Ask about the 30th, let the model query the 28th, and the 28th's prices come back formatted as a verified answer.
The fifth, which is the one we would have missed
We fixed all four and asked for a re-check. It came back with one more.
The new date guard read 2026-08-28, 2026.8.28 and 2026/8/28. It did not read 2026년 8월 28일 — the way a date is written in Korean. Our agent answers Korean farmers in Korean. The guard protecting them only understood the date formats a developer types. A question with a Korean date read as "no date given," which switched the guard off entirely.
It also took the first date when a question named two, so "not the 28th, the 30th" silently became a question about the 28th.
Both are fixed: Korean dates are read, and a question with more than one date is refused rather than guessed at. And because we now know the guard cannot possibly cover every way a human writes a date, the README lists the forms it does not recognize — a year left out, no separators, "the day after tomorrow" — instead of implying the four it does read are the only ones anyone types.
The shape
Every one of these is the same mistake wearing a different coat: the scope of a control was narrower than the scope of the failure.
We had written that sentence in our own README a week earlier, about the product-name guard. Then we added a surface, and the controls stayed where they were. They protected the structured response, not the rendered one. They protected the request, not the attempt. They protected the product name, not the date. They protected the developer's date format, not the user's.
The lesson we would offer to anyone shipping an agent for this hackathon is not "write more guards." It is that adding a surface is adding a scope, and your existing guards do not follow it there. A UI is not a view of your agent. It is a second place your agent can be wrong, and it is the place a judge will actually look.
How we know the fixes hold
Each of the five failures is now a check in a test file written from the reviewer's own reproduction strings. There are 39 of them, and they matter for exactly one reason: 11 of them fail against the code as it stood before the review. A test suite that passes on the broken version tells you nothing. That was the first thing we checked after writing it.
The rest of the suite — 75 controls on the evidence gate, the tool layer and the agent loop — runs green alongside it.
We are not claiming the agent is correct. We are claiming that when we say a guard exists, there is a check that fails without it. For a project whose pitch is "it refuses when it cannot tell," that seemed like the minimum.
Try it
Repository: https://github.com/SongT-50/agent-that-refuses-to-guess — MIT licensed.
Demo video (4:47): https://www.youtube.com/watch?v=ckOxHll3XUA
Built for the Agents for Humans hackathon with the AWS Strands Agents SDK.
#AgentsForHumans #StrandsAgents #AWS
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article