
Agents for Humans: 7,050 minutes short, and why my agent will only accuse the school of 225 of them
Minutes reconciles a child's IEP against the school's own records with a Strands Agents SDK agent on Amazon Bedrock. Why the letters are compiled rather than generated, why undocumented minutes are never called missed, and how a records request that goes unanswered becomes dated evidence on its own.
There is a category of software problem where being wrong is worse than being useless.
I have been building an agent for the AWS Agents for Humans hackathon called Minutes. It works on a problem I did not know existed until I went looking for burdens that agents could genuinely absorb: a child's IEP — the Individualized Education Program that governs special education in US schools — is a legal promise written in numbers. 300 minutes of speech-language therapy per month. Occupational therapy twice a week. An annual review by March 12.
Whether those minutes are actually delivered is a question almost nobody can answer. Not because the answer is hidden, but because answering it means reconciling a year of scattered emails, progress reports and half-remembered Tuesdays against a document in a drawer. The person best positioned to notice is a parent who is already out of hours.
That is an agent-shaped problem: recurring, evidence-heavy, mostly background, occasionally urgent.
But it comes with a trap, and the trap is the interesting part.
The failure mode that matters
The obvious build is: feed the IEP and the emails to a model, ask it what the school missed, generate a letter demanding it.
That build is worse than nothing. If the letter claims a session was missed and the district's own record shows it was delivered — or shows the child was absent that day — the parent has just handed away the only thing they had. Credibility, in a dispute like this, is not one asset among several. It is the whole position. An agent that is right nine times out of ten and confidently wrong on the tenth has made its user's situation worse than if it had never run.
So the design constraint I started from was not "make the model accurate." It was: make the wrong output unrepresentable.
Four buckets, not two
The first thing that fell out of that constraint was an accounting identity. For every service, over every period:
1
owed = delivered + excused + documented misses + undocumentedMost of the work is in the two buckets a naive implementation collapses.
undocumented is minutes with no record either way. Nobody wrote anything down. The naive version treats missing records as missed service — which is exactly backwards, and is how you end up accusing a school of something you cannot show. Undocumented minutes are a gap in the evidence, and the correct response to them is to request records, never to accuse.excused is minutes on days the child was absent. The school could not have delivered them. Ask for them back and you get the one reply that puts every other figure in your letter in doubt.Keeping those apart is not a nicety. It is the difference between a document a district has to answer and a document a district can dismiss.
On the sample term the four buckets read: 8,520 minutes owed, 1,365 documented as delivered, 105 excused, 7,050 short — and 6,825 of that shortfall has no record either way. So the letter accuses the school of 225 minutes, the ones its own records show were missed, and asks for the records behind the rest.
Cite or stay silent
The second thing that fell out was how letters get produced. They are not generated. They are compiled.
Every factual sentence in an outgoing letter carries a footnote marker bound to a specific piece of evidence — a dated district email, a line in a service log, the family's own note, or a records request that went unanswered. A claim with no evidence behind it is not hedged or softened. It is omitted.
A validator enforces it: a letter with a dangling footnote marker, an uncited factual claim, or a legal citation outside a verified allowlist does not compile. That last one matters more than it sounds — a fabricated regulatory citation in a letter to a school district is a fatal defect, and it is exactly the kind of thing a language model will produce fluently and confidently. So the allowlist is built from researched, verified authorities, and anything outside it cannot appear.
The effect is that the model's fluency is used where fluency is safe — reading unstructured school correspondence, writing connective prose — and structurally excluded from deciding a number, a date, or whether a school fell short.
Where the AWS pieces land
The agent is built on the Strands Agents SDK, running on Amazon Bedrock. A few things that shaped the build:
The model is a reader, not an accountant. There is exactly one place the model reads the IEP: a structured-output call that turns the document into typed obligations, with every extracted fact carrying the verbatim sentence it came from. Strands'
structured_output with a Pydantic model made that a single call rather than a parsing project. Everything downstream — the reconciliation, the deadline clocks, the letter assembly — is deterministic Python over that typed ledger.That division is also what makes it cheap. Because the arithmetic never calls a model, the entire engine can be exercised for free. Extraction runs once and is cached to a fixture; the test suite replays it. The whole project is running under a $10 budget with an AWS Budget alarm to enforce it, and the tests are hermetic — no credentials, no network, so CI runs them exactly as a laptop does.
Evidence the agent creates itself. The part I did not anticipate: the agent does not only consume evidence, it manufactures it. Parents have a statutory right of access to their child's records, and almost nobody exercises it systematically, because it means writing the same letter every month forever. The agent compiles that request on a cadence and stops on a Strands interrupt; the parent releases it and posts it by a method that proves delivery, then types the day the district received it. From that date the 45 days that 34 CFR 300.613(a) allows start running. If nothing comes back, the next scheduled wake derives one
documented_silence fact per service the request covered, dated the day the answer was due, raises a card that reads The district has not sent the service records you asked for, writes recorded documented silence on the case's trail, and the monthly Statement prints the sentence a complaint has to plead: received on, due on, under 34 CFR 300.613(a), none recorded. Nothing about any session is claimed. Those minutes stay in the undocumented column. But a school that will not produce its logs has, without meaning to, created documentation, and the parent did nothing after the receipt date.What I would tell someone starting a similar build
Ask what the worst plausible output is, and then design so the system cannot emit it. Not "prompt it not to." Cannot.
For this project that meant most of the interesting engineering ended up outside the model: in a type contract that forces every fact to carry its source, in an accounting identity that keeps absence-of-evidence distinct from evidence-of-absence, and in a validator that refuses to compile a document making a claim it cannot cite.
The model is doing real work. It is just not doing the work where being confidently wrong would cost someone their case.
Minutes is open source (MIT): github.com/N-45div/Minutes. Every document in the sample case is synthetic. It compiles documentation, not legal advice.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article