AWS Builder Center

Agents for Humans: my audit agent passed a 10.9% overcharge as clean

My settlement auditor recomputed every fee and reported zero discrepancies — while a 10.9% commission sat on a statement with a 9.9% contract. The bug was circular verification, the fix was one line, and here is how I proved the fix was load-bearing.

I built a settlement auditor for a small restaurant. It read the statement, recomputed every fee, compared the two, and reported no discrepancies. Every number matched.
The statement had a 10.9% commission on it. The contract says 9.9%.
The agent was not broken. It was doing exactly what I told it to do, and what I told it to do was worthless. This post is about that failure, because the fix is one line of code and the lesson took me two days to see.
The setup
The project is Settlement Sentinel, built for the Agents for Humans hackathon with the AWS Strands Agents SDK on Amazon Bedrock.
One kitchen sells on three Korean delivery platforms — Baemin, Coupang Eats, Yogiyo. Every month three settlement statements arrive, each with its own format and its own commission rate. Reconciling them line by line takes hours the owner does not have, so small over-deductions are simply never contested.
The agent's job is to do that reconciliation in the background, file a dispute when the evidence is clear, and ask the owner when it is not.
The core of it is a function that recomputes what each order should have paid:
1
expected_commission = round(order_amount * rate)
Everything hangs on where rate comes from.
The bug
My first version pulled the rate from the row I was checking:
1
2
# Wrong. The rate comes from the document under audit.
expected_commission = round(row["order_amount"] * row["stated_rate"])
Run that against a statement where the platform billed 10.9% instead of the contracted 9.9%, and here is what you get:
1
2
3
4
5
6
7
8
Order Amount Stated rate Stated comm. Expected comm. Δ Status
CE-2001 29,000 9.9% 2,871 2,871 0 Clean
CE-2002 41,000 9.9% 4,059 4,059 0 Clean
CE-2003 22,000 9.9% 2,178 2,178 0 Clean
CE-2004 38,000 10.9% 4,142 4,142 0 Clean <-- the overcharge
CE-2005 15,000 9.9% 1,485 1,485 0 Clean

Discrepancies flagged: 0 Total amount at stake: 0 KRW
CE-2004 is the overcharge. It reconciles perfectly, because 38,000 × 10.9% really is 4,142. The arithmetic is correct. The audit is not.
This is circular verification: checking a document against itself. The check consumes the very claim it is supposed to test, so it can only ever confirm. It cannot fail. And a check that cannot fail is not a check — it is a ritual that produces a green tick.
The scary part is how good the output looks. Five rows, five zeros, a clean verdict. If I had shipped that and the owner had trusted it, the agent would have been actively worse than no agent at all: it would have converted an unexamined statement into an examined one, with a false clean bill attached.
The fix
One line. The rate comes from the owner's side, not the platform's:
1
2
3
4
5
# Owner-side source of truth. Never taken from the platform's statement.
CONTRACT_RATES = {"baemin": 0.068, "coupangeats": 0.099, "yogiyo": 0.126}

contract_rate = CONTRACT_RATES[platform]
expected_commission = round(row["order_amount"] * contract_rate)
Same statement, same five rows:
1
2
Order Amount Stated rate Contract rate Stated comm. Expected comm. Δ Status
CE-2004 38,000 10.9% 9.9% 4,142 3,762 +380 Flagged
The gap is 380 KRW on one order. Small. It is also the entire reason the agent exists — nobody disputes 380 KRW by hand, which is exactly why it keeps happening.
Proving the fix is load-bearing
Fixing a bug and knowing you fixed it are different things. A one-line change is easy to convince yourself about, so I ran an ablation: put the bug back deliberately, run the same data, and see whether the finding disappears.
1
2
3
With the contract table: CE-2004 flagged, rate_mismatch, 380 KRW at stake
Without it (rate off the all five orders clean, 0 KRW at stake
statement):
The flag vanishes and comes back cleanly with the presence of that one dictionary. That is what "load-bearing" means, and it is worth measuring rather than assuming. I now run this kind of removal test on every component I claim is doing real work. Several times the answer has been "actually, nothing changed" — which is its own useful result.
The general shape
Circular verification is easy to spot in the toy version and hard to spot in your own code, because it never announces itself. It looks like diligence. Some places it hides:
Auditing a statement using the rate printed on the statement. This post.
Asking a model to check its own output. The same weights that produced the error are now grading it.
Validating an API response against a schema the same response was used to infer.
Testing a function against expected values you generated by running that function.
The tell is always the same question: where did the thing I am comparing against come from? If the answer traces back to the artifact under test, you have a ritual.
For an agent this matters more than for a script, because an agent acts. A script that verifies nothing produces a wrong report. An agent that verifies nothing files a wrong dispute — or worse, files nothing and tells the owner everything is fine.
What I kept
The rule I ended up with, written on the architecture diagram:
The rate printed on the statement is never trusted. Believing it would be circular verification.
An independent second source is not a nice-to-have in a verification system. It is the verification. Everything else — the confidence thresholds, the fail-closed gate, the owner escalation — is downstream of having something real to compare against. Get that wrong and the sophistication on top of it is decoration.
Code is on GitHub: https://github.com/minjun0208/settlement-sentinel
It runs without AWS credentials via python agent.py coupangeats --no-llm if you want to reproduce the table above.
Next post: why the agent refuses to file the largest discrepancy in the dataset.
Built for the AWS Agents for Humans hackathon, Professional Agents track.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article