AWS Builder Center
Testing a Fail-Closed AI Agent: 11 Checks and Four Auditable Outputs

Testing a Fail-Closed AI Agent: 11 Checks and Four Auditable Outputs

How ProofPack tests missing records, duplicate IDs, metric bounds, Strands tool use, output generation, and even demo-arrow geometry.

An evidence agent should not be judged only by whether its best-case demo looks convincing. The important questions are what it does when a metric is absent, a record is duplicated, evidence arrives late, a source is unverified, or a tool call fails to produce the expected artifact.
For **ProofPack**, I treated those edge cases as the product. The result is a compact test suite with 11 passing tests, a deterministic pilot, and four output formats that can be inspected independently.
ProofPack is a Strands-powered grant-reporting agent for small nonprofits. It maps obligations to source evidence and deliberately blocks submission when support is missing or requires human review.
## Anchor the suite with a representative failure state
The included pilot has four grant requirements. Its expected result is:
```text
Submission state: BLOCKED
Ready 2/4 · Review 1 · Missing 1
```
Two requirements have verified evidence. One has evidence that still needs source verification. One is missing a required survey summary. The expected decision inbox contains the two unresolved requirements.
This test anchors the behavior of the whole system. If a later change incorrectly turns the missing record into `ready`, changes the overall state, or loses a human decision, the suite fails.
## The missing metric stays missing
Suppose a roster record exists but does not contain the required `people` metric. A generative system might infer a number from context or continue with a vague statement.
ProofPack returns `missing`, and the metric value remains `None`. The test asserts both outcomes. This is the most direct expression of “evidence before assertion”: absence is represented as absence.
## Upper bounds matter too
Not every target is a minimum. A spending obligation might require an amount between 2,000 and 3,000. Evidence showing 3,200 should fail even though it exceeds the lower target.
The engine supports an `upper_target`, evaluates the full range, and records the target as `2000–3000` in its reason. The test protects against a common validation mistake: checking only the lower bound.
## Duplicate identifiers fail closed
Structured workflows depend on stable identifiers. If two evidence records share an ID, silently choosing one could connect an obligation to the wrong source.
ProofPack checks uniqueness before evaluation and raises a `ProofPackError` when duplicate evidence IDs appear. The test expects that exception. It is safer to stop than to guess which record the author intended.
## All four outputs must be real
A successful run must create:
- `proofpack.json`;
- `proofpack-report.md`;
- `proofpack-report.html`;
- `evidence-manifest.csv`.
The output test verifies the exact set of files, checks that each contains substantive content, reopens the JSON summary, and confirms that the HTML includes the responsible-use statement. This catches a class of demos that print success while omitting a promised deliverable.
The CSV manifest includes a SHA-256 digest for every evidence record. The JSON preserves structured decisions, while Markdown and HTML support human review.
## Exercise the real Strands tool loop
Unit-testing only the validator would leave the agent boundary untested. ProofPack therefore includes an offline model double that drives the real Strands `Agent` loop.
On its first turn, the scripted model requests the `build_proofpack` tool with the pilot paths. After receiving the tool result, it completes the turn. The test verifies that the model was called twice and that `proofpack.json` exists.
This test does not pretend to evaluate model quality. It verifies the integration contract: Strands can expose the bounded tool, accept a tool-use request, execute the deterministic build, and return control to the agent.
In production-style execution, the agent uses the Strands SDK's default Amazon Bedrock model provider. The offline test keeps continuous verification reproducible and credential-free.
## Visual geometry is still product quality
The demo video contains connector arrows. During visual review, a diagonal arrowhead was visibly misaligned. That defect did not change the engine, but it made the architecture harder to trust at a glance.
I replaced the fixed arrowhead logic with direction-vector geometry. Four parameterized cases now cover horizontal and differently oriented diagonal lines. Each case verifies three properties:
- the arrowhead vector is collinear with the line;
- its dot product points forward;
- its base has the expected width.
A fifth visual test rejects zero-length lines. Together, these checks prevent the obvious arrow defect from returning.
This was a useful reminder: testing user-visible explanations belongs in the same quality process as testing code. A technically correct project can still communicate the wrong thing through a malformed diagram.
## Add a clean-environment preflight
The suite totals 11 tests: five engine tests, one real Strands loop test, and five arrow-geometry cases. Before submission, I installed the project into a fresh environment and ran the suite from scratch.
That check exposed a packaging issue. The tests imported the video builder, which requires Pillow and WeasyPrint, but the development dependency group initially contained only pytest. The local working environment hid the omission.
I added explicit bounded versions for Pillow and WeasyPrint to the `dev` extra, recreated the environment, and reran everything. All 11 tests passed. The lesson was practical: a green test run is not enough if the environment cannot be reproduced from declared dependencies.
## What I would test next
The current suite establishes the core contract, but a production deployment would add larger fixture sets, corrupt JSON cases, permission and storage boundaries, Bedrock model configuration tests, and end-to-end review workflows with real users.
The central rule would remain unchanged: every failure should preserve evidence integrity and leave the final approval with a human.
Run it yourself
git clone https://github.com/ruma19076/proofpack-agent
cd proofpack-agent
python -m venv .venv
. .venv/bin/activate
pip install -e '.[dev]'
pytest
proofpack demo --output demo-output
• Source code: https://github.com/ruma19076/proofpack-agent
• Video demo: https://youtu.be/CI1sxxSqE1A
ProofPack is open source under the MIT License. OpenAI Codex was used as an AI coding assistant during the contest submission period; no pre-existing application code was incorporated.
#AgentsforHumans
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article