AWS Builder Center
Agents for Humans: Building a Maintenance Agent That Knows When Not to Act

Agents for Humans: Building a Maintenance Agent That Knows When Not to Act

I entered my first AI hackathon from Trinidad and Tobago with limited agent-building experience. Building Maintenance Autopilot with the Strands Agents SDK and Amazon Bedrock taught me that the difficult part of autonomy isn't getting an agent to act. It's proving when it should stop and hand judgment back to a human.

Series: Maintenance Autopilot — Agents for Humans Build Journey (3 articles)

  1. 1
    Agents for Humans: Building a Maintenance Agent That Knows When Not to Act This article
I am writing this from Trinidad and Tobago. My background is not in building AI systems. It is much closer to real operational work: planning work, dealing with competing priorities, understanding risk, and trying to get the right decisions made without unnecessary handoffs.
That shaped how I approached the AWS Agents for Humans hackathon.
I didn't want to build a chatbot just because I could. I wanted to see whether I could take a real-world workflow problem, learn the technology as I went, and build an agent that could genuinely take work away from a human without taking away decisions that still belonged to one.
I chose a deliberately ordinary problem.
Imagine a small landlord receiving fifty maintenance messages in a month.
A toilet seat is cracked. A smoke alarm is chirping. A garage door is stuck. There is damp on a ceiling. A tenant cannot lock the front door.
Fifty messages do not necessarily mean fifty decisions for the owner.
Most should simply move.
But a few absolutely need human judgment.
That distinction became Maintenance Autopilot.
Project note: Maintenance Autopilot is a personal hackathon project built using synthetic rental-maintenance scenarios. It is separate from my professional work and does not use employer systems, data, or intellectual property.

The problem wasn't getting the agent to act

My first instinct was that an autonomous maintenance agent needed to answer a fairly straightforward question:
What should I do with this maintenance request?
As I started defining the workflow, I realised that was the wrong question.
A useful agent shouldn't just classify requests. It needs to move them forward.
For a routine repair, perhaps it should proceed within an agreed authority. If information is missing, it should ask a focused question. If the tenant doesn't respond and there is no safety concern, perhaps the case should wait. If something exceeds the agent's authority, it should go to the owner.
Sometimes both things need to happen. The system may need to take an immediate protective action while simultaneously escalating the underlying issue to a human.
That led me to a deliberately small set of outcomes:
ACT
ASK
AWAITING
ESCALATE
ACT+ESCALATE
CLOSE
The harder question therefore became:
When should the agent not make the decision itself?
That became the central engineering problem.

I started with a decision standard, not a prompt

Before doing more work on the model, I wrote a decision standard.
I wanted to prevent myself from changing the rules every time the agent produced an answer I didn't like.
The standard considered six things:
Safety → Information Sufficiency → Urgency → Trade → Authority → Owner Communication
I then built a ground-truth benchmark of synthetic maintenance cases with expected outcomes and reasons.
Consider two reports:
“The front door handle came off in my hand, but the door still locks.”
and:
“The front door won't pull shut properly. It's not latching and I can't lock it.”
They sound similar.
Operationally, they are not.
The first is a maintenance repair that can proceed. The second creates an active security exposure.
This distinction became important to how I evaluated the project. I wasn't testing whether a model generated a convincing response. I was testing whether the whole system reached the decision required by the standard.
I also tried to impose an anti-overfitting discipline on myself.
When a test exposed a failure, the fix needed to represent a general concept rather than become a special rule that effectively said:
If you see this particular sentence, return this particular answer.
That rule became much harder to follow than I expected.

V1: an LLM agent with a deterministic governor

I used the Strands Agents SDK to construct the agent components, with Amazon Bedrock providing the model layer.
My early architecture combined model interpretation with a deterministic governor.
The reasoning seemed sound.
Tenants describe the same maintenance condition in many different ways, so an LLM could help interpret the language. But I didn't want the model to have unrestricted authority over the final maintenance decision.
The deterministic governor therefore sat around the model's decision and enforced specific rules.
On the development benchmark, this approach began to look very good.
Eventually, the development cases could be processed with excellent final outcome performance and no unsafe autonomous actions.
It felt like a breakthrough.
It wasn't.

Then I tested cases the system had never seen

I didn't want to keep evaluating the system against scenarios I had already used while building it.
So I created a sealed holdout set of 24 synthetic maintenance scenarios.
The underlying problems were realistic, but the wording was deliberately varied. The system had not seen these cases during development.
The result was uncomfortable.
The raw agent produced the correct final outcome on 50.0% of the holdout cases.
The deterministic governor improved that to 62.5%.
That sounds like the governor was doing its job.
Then I looked at the safety results.
There were three cases in the sealed holdout where critical escalation was required.
The raw agent caught none of them: 0.0% critical escalation recall.
With the governor, that improved to 33.3%.
The raw agent produced three unsafe autonomous actions.
The governed system still produced two.
Even the governor's interventions told an interesting story. It intervened in seven of the 24 cases. Five interventions corrected an agent decision.
Two made a previously correct decision worse.
That was the finding that changed the project.
I had built a safety layer that improved the average result while still failing exactly where I needed it to be strongest.
Looking through those failures exposed the underlying problem.
The governor could recognise hazards and conditions when they appeared in language I had anticipated. Reword the same underlying situation, however, and some of that protection disappeared.
My supposedly deterministic safety boundary still depended too heavily on how the tenant described the problem.
I stopped patching individual phrases and reconsidered the architecture.

Separating “what is happening?” from “what must happen?”

The redesign grew around one distinction:
“What is happening?” and “what must happen?” are different problems.
Natural language belongs heavily in the first.
Policy belongs heavily in the second.
I moved toward separate semantic assessors producing controlled concepts that a deterministic policy engine could consume.
Two of the most important components became a Hazard Assessor and a Case Assessor.
The Hazard Assessor is deliberately consequence-blind. Its job is to recognise the hazard present in the report, not decide whether the owner should be notified.
The Case Assessor deals with the broader maintenance situation, including condition, information sufficiency, urgency and likely trade.
Those model-driven components interpret messy language.
They do not own final authority.

Figure 1 — Maintenance Autopilot V2.5 architecture

There is another important detail before decomposition: a cross-symptom hazard check.
That came from thinking about reports containing multiple symptoms.
If the system immediately splits a tenant's message into separate maintenance issues, it can potentially separate facts that are relatively harmless individually but dangerous together.
Water and electricity are an obvious example.
I therefore wanted the system to consider interacting hazards across the complete report before decomposing it into individual issues.

Where Strands and Bedrock fit

One of the things I came to appreciate while learning the Strands Agents SDK was that I didn't need to make the entire application “AI.”
I could use model reasoning where semantic interpretation added value and ordinary deterministic Python where explicit policy was more appropriate.
In simplified form, the pattern looks something like this:
1
2
3
4
5
6
7
8
9
10
hazard_assessment = hazard_agent(hazard_prompt)
hazards = validate_hazard_output(hazard_assessment)
case_assessment = case_agent(case_prompt)
case = validate_case_output(case_assessment)
decision = evaluate_policy(
hazards=hazards,
conditions=case.conditions,
information=case.information,
authority=case.authority,
)
The actual implementation includes additional decomposition, repeat-failure assessment, validation and fail-safe handling, but the architectural principle is more important than the individual functions.
Amazon Bedrock and the Strands agents help answer: “What is happening?”
The deterministic policy engine answers: “Given the decision standard, what must follow?”
Structured output became important here too.
The assessors don't simply generate prose and leave another model to interpret it. They return concepts that the downstream policy understands.
The validation layer then checks those concepts before policy evaluation.
If the semantic layer encounters something that cannot be safely mapped into the approved concept set, the system does not quietly invent a new policy interpretation.
It fails safe rather than guessing.
That was another change in my thinking during the build.
At the boundary of an autonomous system, uncertainty is information.

Building this was not a straight line

There were plenty of less glamorous moments.
This was very much a build-from-my-own-PC project.
I learned what an expired AWS session looks like halfway through an evaluation run. Windows PowerShell's execution policy became my immediate engineering problem on another day. I broke imports while replacing components and occasionally had to work backwards to establish exactly which version of a file I was running.
I was learning parts of Strands, Bedrock and the broader agent-development workflow while simultaneously trying to build something credible with them.
Being new to the hackathon scene also created another temptation: adding things because they sounded technically impressive.
I increasingly tried to resist that.
The project became better when I reduced the question to something I could measure:
Did the system make the correct maintenance decision, and did it avoid unsafe autonomy?

From failed holdout to V2.5

The sealed 24-case evaluation changed the architecture.
From there, development became much more iterative.
I expanded the regression benchmark, created a larger 50-case challenge set, deliberately varied phrasing, and kept testing cases around security, moisture, structural conditions, incomplete information, repeat failures and routine maintenance.
Those tests continued to find weaknesses.
The system didn't jump from 62.5% to its final state in one redesign.
The assessors, validation layer and deterministic policy evolved as the challenge cases exposed edge conditions.
Eventually, I stopped.
That is worth saying because it is easy to tune a benchmark indefinitely and then call the resulting score validation.
I froze V2.5 as a known version of the decision engine.
At freeze, the results were:
Evaluation setScenariosPolicy outcome accuracyCritical escalation recallUnsafe autonomous actions
Regression benchmark61100.0%100.0%0
Challenge set50100.0%100.0%0
Those are strong numbers.
They also need an important qualification.
The 50-case challenge set began as separate validation, but I subsequently used failures from it to improve the V2.x system.
By the final V2.5 run, it was therefore not a pristine unseen holdout.
I don't consider its 100% result proof that the system will generalise perfectly to new maintenance reports.
The original 24-case test is different. It was sealed before the earlier architecture saw it, which is why its poor result remains such important evidence in this project.
The next meaningful test of frozen V2.5 should follow the same discipline: create another unseen benchmark, don't tune against it first, and accept whatever result comes back.
There is another reason I am cautious about the final 100%.
The intermediate semantic metrics are not perfect.
On the final 50-case challenge set, safety classification accuracy was 72.0%, urgency classification accuracy was 68.0%, and primary-trade accuracy was 70.2%.
I am deliberately reporting those numbers.
They showed me that individual semantic classifications can still differ from benchmark labels while the combination of validated concepts and deterministic policy reaches the correct final workflow decision.
That doesn't make those semantic differences irrelevant.
It means I need to understand which metric represents the behaviour I actually care about, while continuing to test the components underneath it.

Lessons learned

Held-out validation mattered more than development accuracy

My development benchmark became useful for regression. It told me whether a change broke behaviour I already understood.
The sealed holdout answered a much harder question:
What happens when the system encounters language I didn't build around?
The answer wasn't flattering.
It was also probably the most useful result I produced during the project.

A safety layer can be brittle too

I initially associated deterministic rules with robustness.
That was incomplete thinking.
A deterministic decision can still depend on a brittle interpretation upstream. If that interpretation relies too heavily on anticipated phrases, unseen language can defeat the protection.
The governor improved V1.
It did not solve the underlying problem.

Separate interpretation from authority

This became the architectural lesson I will take away from the project.
Models are useful where language is ambiguous and variable.
Deterministic policy is useful where consequences, authority and escalation need to be inspectable and testable.
For Maintenance Autopilot, separating those responsibilities made the whole system easier for me to reason about.

One unsafe action can matter more than aggregate accuracy

This changed how I thought about evaluation.
Overall accuracy is useful, but a single number can hide the failures that matter most.
That is why I started tracking critical escalation recall and unsafe autonomous actions separately.
For this system, getting 23 routine maintenance cases right doesn't compensate for autonomously mishandling the 24th if that is the case where a human was required.

What I haven't built yet

Maintenance Autopilot V2.5 is a validated prototype, not a production property-management system.
A real deployment would need authenticated users, property and tenant isolation, durable maintenance dispatch, duplicate-action protection, guaranteed delivery and acknowledgement of escalations, monitoring, privacy controls, backup and recovery, and other operational safeguards.
The current demonstration also shouldn't be confused with an agent autonomously operating an entire real property-maintenance ecosystem.
It doesn't.
The goal of this stage was more specific.
I wanted to establish whether I could create and measure a useful boundary between autonomous work and human judgment.
There is more work to do.
Most importantly, frozen V2.5 now deserves the same treatment that exposed V1: a genuinely unseen test.

What this changed for me

I entered this hackathon thinking the interesting part would be automating maintenance.
It wasn't.
The interesting part became designing the boundary between machine autonomy and human judgment.
That is probably also the part of this project that reaches furthest beyond rental maintenance.
Many real workflows contain a large amount of routine activity surrounding a much smaller number of consequential decisions.
Agents become useful when they can handle more of that routine volume.
They become more trustworthy when we can also measure where they stop.
For someone entering his first AI hackathon from Trinidad and Tobago, that has been the most valuable part of the experience.
I started by trying to build an agent that could make maintenance decisions.
I ended up much more interested in building a system that could demonstrate which decisions it shouldn't make.
For me, that became the real meaning of Agents for Humans.
Not removing the human from the workflow.
Being much more deliberate about the moments where the human is actually needed.

Series: Maintenance Autopilot — Agents for Humans Build Journey (3 articles)

  1. 1
    Agents for Humans: Building a Maintenance Agent That Knows When Not to Act This article
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article