
The agent that checks its own homework
How I built a self-validating daily operations report that my team actually trusts — it cross-checks its own numbers against the official dashboards and shows its work every morning.
How I built a self-validating daily operations report that my team actually trusts.
What it does, and who it's for
It's for operations and reliability leaders across a group of fulfillment sites who walk into a morning stand-up needing one honest number: how bad were equipment faults yesterday, and can I trust it?
Every day it posts a daily fault report each morning to a team chat channel (fault counts per site, ranked, with a regional total), cross-validates each site's figures against the official source-of-truth dashboards and scores itself, and explains any gap in plain language — "one site off by a single count: an event right at the midnight day-boundary, within tolerance, not a data error."
How I built it
Amazon QuickSight — the production dashboards serve as both the source of truth and the validation target. The agent reads per-site, per-day figures and compares them to its own report.
A scheduled agent runtime — it runs daily: pulls the data, formats the report, runs the self-check, and posts to the team channel, under a fixed least-privilege tool policy so it's safe to run unattended.
Upstream equipment/event data — raw fault events are curated to match the dashboards' counting logic (rapid repeat-events on the same equipment de-duplicated into distinct events).
Team chat integration — the delivery surface for both the report and the validation post.
Key decision: validate against the number leadership already trusts. Instead of inventing a new metric, the agent ties out to the exact dashboards leaders already look at — so there's never a "whose number is right?" argument.
The one thing that makes it enjoyable
I made the agent check its own homework in public — and show its reasoning. Most reporting bots dump a table and leave you to trust it. This one proves its number every day, side-by-side with the official dashboards, and when something's off it says so and explains why in human terms. A self-scoring validation post turns trust into a visible daily number; honest discrepancy handling separates a real problem from a harmless boundary artifact; and a consistent persona with a robot marker keeps it approachable and unmistakably automated.
How I know it worked
Validation posts consistently score near-perfect matches, with the only misses being documented single-count, day-boundary events on the lowest-volume site. It resolved a real cross-team dispute: when a colleague flagged that two channels showed different counts for the same site, the agent reconciled it (raw alarms vs. dashboard-aligned distinct events) in one post — the reply was a "good catch," not a fire drill. It fails honestly: on days the data query times out, the channel shows the agent's own "query failed" notice instead of going dark. And teams across the region keep joining the channel because they trust the number.
See it in action
Screenshots below: the daily report; the cross-validation table scoring the sites as matches; the plain-language reconciliation of two different counts for one site; and the honest "query failed" notice.
This daily report is the data backbone of a bigger mission — reducing equipment faults across the region — by giving everyone one trustworthy daily pulse to agree on.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article