
Agents for Humans, Human Approval Is a Capability, Not a Fallback
AIRCheck treats human authority as a runtime-bound capability for one consequential media operation. This essay follows a real loudness failure from measured fact to approval, automatic continuation, re-verification, and durable evidence on AWS.
Series: Agents for Humans (1 article)
- 1Agents for Humans, Human Approval Is a Capability, Not a Fallback This article
I wanted AIRCheck to keep working after I stopped touching it. I also did not want a model deciding that it had permission to alter a finished master. The approval button was easy. Defining what the approval actually meant was the hard part.
The master is finished. The delivery isn't.
Media delivery is full of work that looks mechanical until it is not. A package can have the right files and still fail a destination's loudness range, caption format, filename rule, checksum requirement, or manifest check. An autonomous system can inspect those rules, measure the media, and prepare safe changes. The uncomfortable question arrives when the only successful repair changes content.
The button wasn't the hard part
I have used human-in-the-loop patterns that boil down to a nervous confirmation dialog. It asks whether I am sure. That is not a control boundary. It is a pause in a conversation. If a model can decide what should happen, decide that my click means permission, and then mutate whatever state it can reach, I am not granting authority. I am only decorating an already autonomous action.
AIRCheck is my answer in a bounded synthetic world. The Last Lightkeeper is a synthetic program, and Northstar Broadcast Network is a fictional destination. The public deployment uses prepared profiles and admitted candidates, so the filmed run does not pretend to parse arbitrary raw prose live. The broader design starts with prose and turns it into predicates, facts, authority, and evidence. The demo focuses on the boundary that matters most: who is allowed to make a consequential change, and how does the system prove what happened?
In the hero run, the destination has 15 requirements. Before asking me for anything, AIRCheck handles the safe work on working copies. It catches the filename issue, checks the caption format, verifies the checksum and manifest, and measures the media. The loudness reading is -19.05 LUFS against an allowed range of -26 to -22 LUFS. That is 2.95 LUFS above the ceiling. The compliant path requires a content-affecting audio derivative.
AIRCheck stops and surfaces HUMAN DECISION REQUIRED. The live AgentCore role is bounded diagnosis and recommendation among already admitted options. It can identify the proposed operation,
create_normalized_audio_derivative, and explain why it is the candidate repair. It cannot make that operation allowed. The original stays locked.
AIRCheck stops before the content-affecting repair and asks for one exact human decision.
Approval is a capability
A generic confirmation dialog asks, “Are you sure?” The meaningful question is different: what exactly did yes authorize?
I modeled authority as a set of explicit tiers:
- Tier 0: inspect, measure, read, hash, compare, and report.
- Tier 1: reversible working-copy or derivative operations with an exact option binding.
- Tier 2: a content-affecting derivative that requires an explicit human decision and a runtime-minted, opaque, single-use authorization.
- Tier 3: forbidden actions such as overwriting or deleting originals, inventing facts, silently resolving contradictions, or claiming external delivery.
The approval is bound to the current run, the source identity, the destination, the exact operation, and its settings. Approving “fix the audio” would be too broad. Approving this operation on this source, with this target and these parameters, is narrow enough to verify. If any of those bindings change, the capability is not valid.
A decision request is not authorization. The request is a fact in the run history. The capability is a separate runtime object that the deterministic system mints only after the human decision, stores against the run, and consumes once at action start. It is not derived from model confidence. The model does not see it, mint it, or get to reinterpret it.

The boundary is explicit: the model recommends, deterministic systems decide, and a human approves the exact repair.
Then I stop touching the mouse
I click once. Approval does not itself create a file. The runtime validates the decision against the request, mints the exact capability, consumes it at action start, and executes the bound operation. There is no second conversational command where the model gets to reinterpret what I meant.
AIRCheck creates a delivery derivative with a -24 LUFS target. It does not overwrite the original. The system re-measures the result at -24.05 LUFS, refreshes the checksum manifest because the package changed, and evaluates the destination again. Only deterministic terminal logic can award
DELIVERY_READY.The final result is 15 of 15 requirements verified. Four operations were autonomous, and one was human-authorized. The original remains unchanged. A separate denial run ends
BLOCKED, with no silent retry and no pretend success.A binding is part of the proof
M6 taught me that a beautiful evaluator can still manufacture truth. I red-teamed the evaluator itself three times before accepting the final result. One failure was not a bad measurement. It was a bad relationship: a valid measurement attached to the wrong requirement.
The nodes were right. The edges were wrong. That matters because evidence is relational. An approval attached to the wrong operation can be just as dangerous as an incorrect measurement. The system needs to preserve the links among source, predicate, finding, proposed action, authority, effect, and terminal verdict.
The repaired evaluator runs against a frozen deterministic corpus. It covers 27 of 27 scenarios, with 26 of 26 runtime-to-terminal matches and 1 of 1 safe refusal. Semantic responses are held fixed, provider runs are zero, and there is no LLM judge deciding whether the system passed. Those numbers describe deterministic conformance. They are not a claim of live-model accuracy.

The evaluator measures deterministic conformance, not live-model accuracy.
AWS made the boundary concrete
Deploying on AWS forced the topology to reflect the trust boundary. API Gateway HTTP APIs front the Lambda application and runtime path. S3 holds authoritative versioned state and evidence, while Lambda's local disk stays disposable. A per-run conditional-write lease keeps concurrent mutations from quietly racing each other.
AgentCore Runtime hosts the bounded diagnose, select, and escalate seam. A Strands agent with Nova Lite recommends among options that the deterministic system has already admitted. The deterministic runtime still owns measurement, predicate evaluation, authority, mutation, evidence, and termination. CloudWatch logs and service observability give me correlated API and AgentCore records without pretending that this is full distributed tracing.
That service topology is not a logo list. It makes model agency useful without giving the model power over truth. In the same-run hero evidence, the semantic source is recorded as
AGENTCORE and paired with a matching CloudWatch invocation. The separate AWS-hosted CodeBuild workflow is labeled separately rather than folded into the runtime story.What I learned building it
I built the dangerous boundaries inward. M4 exposed authority atomicity and recovery problems. M5 exposed how easily a system can manufacture a convincing truth. M6 forced the evaluator itself through three independent RED audits before I trusted its green result. A green test count is useful. It is not permission to stop thinking.
The most useful rule I kept was simple: filesystem state is memory, but conversation is not. If an approval, measurement, or mutation matters, it has to be written into the run's evidence and bound to the objects it describes. Otherwise the system may remember the story while losing the proof.
I do not want an agent that asks me to micromanage every mechanical step. I want it to know which work is safe, which work is consequential, and what to do after I grant one exact capability. In AIRCheck, the model gets room to diagnose and recommend. The human owns authority. Then I stop touching it, and the runtime keeps going.
Model agency. Deterministic authority.
See the real thing
- Try the live AIRCheck demo in its bounded synthetic workspace.
- Read the public source on GitHub .
- Watch the frozen AIRCheck demo film .
#AgentsForHumans
Series: Agents for Humans (1 article)
- 1Agents for Humans, Human Approval Is a Capability, Not a Fallback This article
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article