AWS Builder Center
Agents for Humans: Publishing the Evaluation That Ties Instead of the One That Flatters

Agents for Humans: Publishing the Evaluation That Ties Instead of the One That Flatters

My Strands agent scored 50 out of 50, and so did a two line rule I wrote to beat it. Here is why I published the tie, the fragile threshold and the null result, and where the agent actually earns its place.

I built Quorum for the Agents for Humans hackathon. It reads Baltimore City Council's legislative record and decides whether a new bill continues an issue it has already seen, like a rezoning that died in 2023 and came back in 2026 under a new file number. Before I called the evaluation done, I asked the question a skeptical judge would ask first: does the agent actually beat a simple rule, or does it only look like it does?
The set
The evaluation set is 50 hand labeled pairs of council records, 19 of them true continuations, in two tiers:
  1. Parcel tier, 28 pairs. Both records touch the same parcel or block.
  2. Citywide tier, 22 pairs. No shared parcel. These share a sponsor and have near identical titles.
I built the second tier on purpose. The first tier alone hides a whole class of continuation that a parcel match cannot see: citywide bills that name no property at all, die at the end of a council term, and return years later under a new number.
The tie
A tuned two clause rule, parcel exact, or title similarity at or above 0.97 with the earlier bill dead, scores 100% on all 50 pairs. So does the Continuity Agent.
I report that first, because it is the strongest argument against the whole project.
Then I checked how fragile the rule is. Swept across thresholds, it is perfect only between 0.94 and 0.97 and wrong on either side. The margin is one negative pair at 0.934 and one positive pair at 0.971. Thirty seven thousandths of a similarity score separate a right answer from a wrong one on fifty examples. That is a property of this sample, not of how cities legislate, and it is exactly the kind of number that should make anyone distrust a headline accuracy figure, including mine.
The held out split
Next I tested whether the tie survives once the rule loses its tuning advantage. The 50 pairs were split into a 24 pair tune half and a 26 pair test half, stratified by tier and label and assigned by a hash of the pair id, so nobody picked the easy ones. The rule's threshold was tuned on the first half only and frozen at 0.950.
On the 26 test pairs, the frozen rule and the agent agree on every single one. McNemar's test gives p = 1.0.
I wrote that result up as exactly what it is: this test separates nothing. A negative result that cost real engineering time is worth more than a positive one I never actually tested.
Where the real answer was
The per tier breakdown is what changed my thinking.
  • On the parcel tier, a parcel lookup alone finds 8 of 8 true continuations. Deterministic code is the right tool for a deterministic problem.
  • On the citywide tier, the same lookup finds 0 of 11, because those records name no property.
11 of the 19 true continuations, 58%, have no parcel anywhere in them. That is the real argument for the agent, not the tied 100%.
What the prompt is worth
I also stripped every domain hint from the Continuity prompt, keeping only the task, the output schema and an instruction not to invent facts. Accuracy fell from 100% to 94% (47 of 50), still above the 84% of the best single deterministic feature. The three misses were a liquor license tied to its own zoning approval, a charter amendment returning after a failed term, and one street condemnation split into two filings. That is what the domain knowledge actually buys.
What I took away
None of this makes the agent look magical. It makes the claim narrower and more useful. The agent ties a good rule where the rule already works, and it reasons over evidence where the rule cannot reach, which in this set is the majority of real continuations.
Every decision it makes also says what it relied on and what it set aside, which a threshold can never do. Publishing the tie, the fragile threshold and the null result is the part of this project I am most glad I did not skip.
Live demo: https://quorum-peach.vercel.app
Code and full results: https://github.com/Rickygole/Quorum
Demo video: https://youtu.be/d4mJJAkWwK8
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article