AWS Builder Center

The Wrong Pair Ranks Higher Than the Right Pair: Why My Agents for Humans Build Needed an Agent

Why my neighbourhood safety agent needs an agent. The near-miss reports in my dataset resemble each other less than one of them resembles a completely unrelated report — the ordering is inverted, so no cosine threshold separates them at any value. The deciding evidence isn't in the text.

In my last post  I fixed retrieval in a neighbourhood safety agent by indexing a normalised summary instead of raw report text: group separation went from −0.16 to +0.52.
That invites an obvious question, and I spent a day trying to make the answer be yes.
If retrieval separates the groups that cleanly, why is there an agent here at all? Cluster on cosine similarity, pick a threshold, alert above it. An afternoon of work, nothing to tune, no model to pay for.
I measured it. It does not work, and how it fails is the interesting part.
The case that decides it. My dataset contains a deliberate near-miss: three reports of someone loitering, in three different zones, spread over three weeks. It should not alert — that is roughly the base rate of living somewhere, and broadcasting it is what teaches people to mute the service.
Next to it sits a genuine cluster: four reports, one parcel locker, 36 hours, four separate reporters. That one should alert.
In plain language both are "several people reported someone hanging around."
cosine similarity
Within the genuine cluster0.708 – 0.814
Within the near-miss0.436 – 0.456
Near-miss report → an unrelated report0.576
Read the last two rows again. The near-miss reports resemble each other less than one of them resembles an unrelated report about a car driving past some driveways.
This is not a threshold that needs tuning. The ordering is inverted. Sort every pair by similarity and the wrong pair sits above the right pair. No cut point puts one on the correct side without putting the other on the wrong side. Tuning changes which mistakes you make, not whether you make them.
I swept it anyway. To assemble the near-miss group at all you have to drop to 0.45, and you buy that with 20 wrong cross-group links; 0.40 catches it properly and costs 60. At 0.50 and up the noise clears and the near-miss group is simply invisible — never assembled, so never declined. Ten to one against, at every setting.
What actually separates them is not in the text. One place versus three zones. 36 hours versus three weeks. Four independent reporters versus three strangers. Similarity cannot see any of it.
You could bolt on if same_zone and span < 48h and reporters >= 3. It gets both my cases right; I ran it. But every constant there is a guess dressed as a policy, and — more importantly — it cannot explain itself. When my agent declines it says why, in a sentence a neighbour can read and argue with. span < 48 returning False is not a reason. For a system whose whole value is staying quiet, the account of why it stayed quiet is what earns it the right to keep running.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article