AWS Builder Center

Agents for Humans: Building Vouch and Learning Where an Agent Should Stop

Vouch — Agents for Humans (3-part series) • Part 1: Where agents stop — https://builder.aws.com/content/3JDKrxNPkMM7Z87gfpAGC1vurch • Part 2: Why readable evidence wasn’t safe — https://builder.aws.com/content/3JDLFi7qt9rVb1vE9sGSkbiZJyh • Part 3: When independent agents disagree — https://builder.aws.com/content/3JDLQMHwO6s3IXG9QaxAXPidiIW

I started Vouch with a manufacturing question.
Before an incoming lot of material can be used, how does a Quality team know that the supplier evidence in front of them is actually the right evidence for that lot, that the right specification is being used, and that any exceptions or equivalent test methods are valid?
At first, I thought the interesting part would be the pass/fail decision.
It wasn’t.
Once the right requirement and the right measurement are known, a lot of that work is ordinary deterministic software. A threshold like 462 MPa < 480 MPa does not need an agent.
The harder part is everything that has to be established before that comparison is meaningful.
That became the boundary Vouch is built around.

Give the agents the ambiguous work

Vouch uses two Strands agents running with Amazon Bedrock and Nova Pro.
The Investigator builds the case around the incoming material: the current specification, required tests, supplier context, exceptions, and approved equivalences.
The Independent Verifier reconstructs that basis separately. It does not receive the Investigator’s answer. Its job is to catch a conclusion that may look reasonable while resting on the wrong revision, the wrong-site exception, or an equivalence that does not actually cover the case.
Neither agent can release or quarantine material.
Neither has a tool for changing inventory or the production plan.
Their job ends when they have established enough context for ordinary software to take over.
That separation became more important as the build progressed.

Agreement is useful. Disagreement is useful too.

Early on, I thought of the Verifier as a quality check: Investigator reaches an answer, Verifier confirms it, move on.
That turned out to be the wrong mental model.
If an independent reviewer never disagrees, there is a fair question about whether it is actually independent.
One Vouch scenario contains two legitimate viscosity measurements taken using different methods. One points toward quarantine. The other is covered by an approved equivalence and points toward release.
Both agents can build defensible cases.
The plant data does not contain a rule saying which measurement should take precedence.
So Vouch stops.
It does not pick whichever agent sounds more confident. It does not average the answers. It does not ask a Quality Manager to approve an AI recommendation.
Instead, the Quality Manager establishes the missing fact for that lot: which measurement should be used for this decision.
The same DecisionRecord resumes with that new information. Both agents reassess, and only after they converge does deterministic evaluation run.
That was one of the biggest shifts in how I thought about the system.
Disagreement was not a failure mode to hide. It was a useful product state.

Make the things an agent should not do unreachable

Another lesson was that prompts are a weak place to put important boundaries.
Telling an agent “do not modify inventory” is not as strong as simply giving it no way to modify inventory.
Vouch’s model-backed roles have bounded read tools and structured outputs. RELEASE and QUARANTINE are not actions the agents can perform.
Once the case is ready, deterministic software evaluates the actual requirements. A scoped, single-use capability then permits the corresponding state change.
This became a general rule for the build:
When a behavior should never depend on model discretion, remove it from the model’s reachable surface.

Then test the idea against something simpler

I also built an evaluation comparing the agent path with a deterministic baseline.
The first result looked encouraging for the agents.
Then I improved the baseline.
The advantage disappeared.
That was uncomfortable, but useful.
It reminded me that the question is not, “Can I make an agent solve this?”
The more important question is:
After deterministic software has done everything it reasonably can, is there still meaningful judgment left?
For Vouch, I believe there is. Specifications reference other specifications. Exceptions have scope. Equivalent test methods have conditions. Supplier evidence comes in different forms. Sometimes two defensible interpretations really do exist.
But that belief is stronger because the system was forced to compete with a serious deterministic alternative rather than a deliberately weak one.

Where Vouch ended up

The final pipeline is simple to describe:
Admit → Assess → Reconcile → Decide → Act
Supplier evidence first has to be safe and attributable.
Strands agents establish the governing requirements and usable evidence.
Their assessments are reconciled.
Deterministic software makes the material decision.
A scoped capability permits the resulting state change.
The live workflow runs on Amazon Bedrock AgentCore Runtime, with Amazon Bedrock, Nova Pro, Guardrails, S3, Textract, and DynamoDB supporting different parts of that path.
But the AWS service list is not the thing I will remember most from building it.
I’ll remember the division of labor.
Agents are valuable where the work still requires interpretation. Deterministic software is better when the rules are known. And professionals should enter the workflow where their expertise supplies something the system genuinely cannot establish on its own.
For me, that is what Agents for Humans ended up meaning.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article