AWS Builder Center

A digest that changed when nothing did — Agents for Humans

A timestamp inside a content digest meant two reads of an unchanged repository disagreed. Every approval then expired between being issued and being used, the pipeline deadlocked, and the symptom pointed at a layer three steps away from the fault.

Agent safety and release evidence
The bug presented as drift detection is too sensitive. It was not that. It took an embarrassing amount of time to find, and the fix was deleting one field from a hash.

What the approval is supposed to do

The agent I was building deploys code, so the interesting question is not did a human approve this but did the human approve this, the exact thing about to be deployed.
"Approve A, execute B" is the failure that matters. A person reads a diff, agrees, and between their click and the deployment the content changes — a rebase, a race, a helpful colleague, a re-plan by the agent itself.
So an approval grant binds three digests: the plan, the input (repository content), and the target. Before executing, the executor re-inspects the repository, recomputes those digests, and consumes the grant with the values it just computed. If they disagree, the approval is invalid and the release falls back to being re-audited.
Note the executor recomputes them. Accepting digests from whoever called it would make the check ceremonial. The question is not "was this valid when it was issued" — it is "is it still valid right now".
That design is right. The implementation had a hole in it.

The hole

The change set — the record of what the inspector found — looked roughly like this:
1
2
3
4
5
6
7
class ChangeSet(Contract):
repository_id: str
baseline_commit: str
head_commit: str
files: tuple[FileChange, ...]
working_tree_dirty: bool
inspected_at: datetime # <- this
and input_digest() hashed the canonical form of the whole thing.
Read that field again in the context of "digest, recomputed just before execution."
Inspect a repository twice with nothing changed in between, and you get two different digests. Not because a byte moved. Because time passed.

What it looked like from outside

Every release went: audit, gates, plan, policy, waiting for approval, approved… and then, at execution, "the approved content has changed" — falling back to AUDITED, forever. Nothing ever left the approved state.
The symptom pointed at the wrong layer entirely. I went looking at the drift detector, which was fine. Then at the manifest comparison, which was fine. The actual fault was in a contract two layers upstream, in a field that looked like plain provenance metadata.
If I had wanted to make this hard to find on purpose, I would have done exactly this: put a timestamp in a digest and then only surface the consequence at the very last step of a nine-step pipeline.

The fix, and the test that keeps it fixed

inspected_at is excluded from the digest. It is still on the contract — when a repository was read is worth recording — it simply is not part of what "this content" means.
The fix is one line. What makes it stay fixed is two tests, in both directions:
1
2
3
4
5
6
7
def test_the_digest_is_stable_when_nothing_changed(repo):
assert inspect(repo).input_digest() == inspect(repo).input_digest()

def test_the_digest_changes_when_one_byte_does(repo):
before = inspect(repo).input_digest()
(repo / "app" / "version.py").write_text(...) # one byte
assert inspect(repo).input_digest() != before
Either test alone is passable by a broken implementation. The first passes for a constant. The second passes for a timestamp. Together they say the thing I actually meant.

The general shape

A digest that binds a human decision to content must cover content and nothing else. Anything time-varying inside it converts a safety property into a liveness bug — and liveness bugs caused by safety mechanisms are miserable to diagnose, because the mechanism is working, in the sense that it is refusing things. It is refusing everything.
The generalisation I took away: for every field in a hashed structure, ask would two observations of an unchanged world produce the same value here? Timestamps fail it. Absolute paths fail it. Iteration order over a set fails it. Anything derived from the observer rather than the observed fails it.

A second one, same family

Later, the console showed Binding holds for an approval that had already been consumed.
The session object held the grant as it was at issuance — an unconsumed copy. The version that actually gets spent lives only in the ledger. So the UI was faithfully rendering a snapshot, and the snapshot was stale.
Worse than cosmetic: it implies the approval can be used again.
The fix belongs in the session (re-read the ledger after execution), not in the view. Same lesson wearing different clothes: be explicit about who holds the truth, and never ask an object holding a snapshot what is true now.

Built for the AWS Agents for Humans hackathon with the Strands Agents SDK. The project is ReleaseSentinel, an evidence-backed release agent. Repository: https://github.com/yyswordsman-CN/release-sentinel
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article