
Agents for Humans: Autonomy Isn't the Same as Authority
Why I separated AI interpretation from decision control in Maintenance Autopilot
Series: Maintenance Autopilot — Agents for Humans Build Journey (3 articles)
- 2Agents for Humans: Autonomy Isn't the Same as Authority This article
When I started building Maintenance Autopilot, I thought the main challenge would be teaching an AI agent to understand maintenance problems.
I was wrong.
Understanding the problem turned out to be only half of it.
The harder question was:
Once an agent understands what is happening, who decides how far it is allowed to go?
That question changed the architecture of the project.
The problem wasn't classification
Maintenance requests are messy.
A tenant might write:
“The AC isn't cooling.”
Or:
“The second bedroom window doesn't latch.”
Or:
“The AC isn't cooling. A technician found a failed condensate pump and quoted $285 to replace it.”
A language model is useful here because people do not describe maintenance problems using neat database fields.
The agent needs to interpret what is happening, determine what information is present, identify possible hazards, recognise whether something is getting worse or recurring, and understand whether a diagnosis or repair estimate has already been provided.
But I realized I was mixing two fundamentally different problems:
- What is happening?
- What is the system authorised to do about it?
Those questions do not necessarily belong to the same decision-maker.
One report, two layers of reasoning
Maintenance Autopilot therefore evolved into two deliberately separate layers.
The first is semantic.
Using the Strands Agents SDK and Amazon Bedrock, the agent interprets the human report and produces structured, decision-relevant information.
The second is deterministic.
Those interpreted facts pass into an explicit policy engine that controls what the system is allowed to do next.
The principle became:
AI handles ambiguity. Policy handles authority.
That sounds simple now.
It was not where I started.

Strands/Bedrock interpretation → deterministic policy → governed outcome
Why not just put the rules in the prompt?
This became one of the most important design decisions in the project.
I could have told the model:
Repairs above $200 need landlord approval.
Critical hazards must escalate.
Ask for more information when necessary.
And then allowed the model to interpret the situation and decide whether those instructions applied.
But that would mean the same probabilistic system was doing two jobs:
interpreting the situation and deciding the limits of its own authority.
I wanted those responsibilities separated.
Consider a simple repair threshold.
If the property's delegated repair authority is $200, the relevant policy can be explicit:
1
2
if estimated_cost > repair_authority:
return ESCALATEThe interesting part is not the Python.
The interesting part is where the decision lives.
Once the relevant facts have been established, the model does not need to decide whether $205 is greater than $200 or whether it feels reasonable to proceed anyway.
The authority boundary is deterministic.
$200 is not a budget
This distinction became clearer after I showed the project to someone else.
They initially interpreted the $200 threshold as a maintenance budget.
It isn't.
It is a delegated decision right.
Suppose a repair is quoted at $195.
The repair may be appropriate, and the agent has been given authority to allow it to progress.
Now suppose the quote is $205.
The repair may be equally appropriate.
Nothing magical happened to the repair in those extra ten dollars.
What changed is:
The agent no longer has the authority to make that decision on its own.
So it returns the decision to the landlord.
That is a small example, but it captures what I became interested in during this project.
AI capability and AI authority are not the same thing.
Authority isn't only about money
The same idea applies to other maintenance decisions.
A bedroom window that “doesn't latch” may sound straightforward, but important information can still be missing.
Is this a security exposure?
Can the property be secured?
Is the window simply difficult to close?
Rather than confidently inventing the missing context, Maintenance Autopilot can return ASK.
A strong and worsening smell of gas is different again.
That takes the protective ACT+ESCALATE path.
And a resolved issue can return CLOSE rather than creating unnecessary work.
The policy therefore has six governed outcomes:
ACT · ASK · AWAITING · ESCALATE · ACT+ESCALATE · CLOSE
Not acting is not necessarily a failure.
Asking, waiting, escalating and stopping can all be successful agent behaviours.

Decision outcomes / four representative examples
This changes what “good automation” means
There is a natural temptation when building agents to maximise the amount of work they can perform autonomously.
More automation sounds better.
But for consequential workflows, I am increasingly unconvinced that autonomy percentage is the right objective.
Imagine two agents.
One autonomously processes 95% of cases but occasionally crosses a decision boundary it should not have crossed.
Another processes 70% autonomously and reliably returns the remaining decisions to a person.
Which is the better agent?
That depends on the domain, but the answer is not automatically the first one.
For Maintenance Autopilot, the objective became:
Automate where authority exists. Stop where it doesn't.
The human is part of the architecture
This also changed how I thought about the user experience.
I did not want the landlord supervising every decision the agent made.
If they have to read every report, approve every step and continuously check what the AI is doing, I have simply created another interface for the same administrative burden.
Instead, Maintenance Autopilot includes a Landlord View.
It answers four simple questions:
What came in?
What progressed?
What needs me?
What are we waiting on?
What progressed?
What needs me?
What are we waiting on?
The intention is not to remove the landlord.
It is to use their attention where it adds value.

Landlord View showing Evaluated / Progressed / Needs You / Waiting
Then testing exposed the other half of the problem
Separating authority from interpretation does not magically make the system reliable.
A deterministic policy can only govern the facts it receives.
That became particularly clear during testing.
After development and deployment testing, I froze the application and challenged it with 80 additional cases that had not been used to tune that version.
The frozen system scored 63/80 — 78.8%.
I kept the misses.
Several of them exposed the same architectural weakness: the deterministic policy behaved consistently with the information it received, but the semantic layer had failed to recognise an important signal such as repeat failure, progression or missing information.
That led to another principle:
Deterministic policy can control an agent's authority. It cannot compensate for a material fact the interpretation layer failed to recognise.
So bounded autonomy does not eliminate the need to improve the AI.
It makes the boundary between AI failure and control failure easier to see.
That distinction is useful engineering information.
One result mattered particularly
Before running that post-freeze evaluation, 15 cases were preregistered as requiring the protective ACT+ESCALATE behaviour.
The frozen system correctly governed:
15/15.
That is not a claim that the agent is universally safe.
It is a much narrower claim:
All 15 cases in that locked evaluation that were expected to require the protective path received that governed outcome.
For me, the imperfect 63/80 result and the 15/15 protective-path result belong together.
One shows where the interpretation still needs work.
The other shows why explicit decision boundaries are worth testing.
Maybe this architecture matters beyond maintenance
Maintenance Autopilot is currently a residential-maintenance prototype.
But the more I worked on it, the less I thought the interesting part was the maintenance interface itself.
The reusable pattern may be:
Unstructured human input
→ AI interpretation
→ structured decision facts
→ explicit authority policy
→ autonomous action or human decision
→ AI interpretation
→ structured decision facts
→ explicit authority policy
→ autonomous action or human decision
That could be useful anywhere an agent needs to understand ambiguous information but should not have unlimited authority over what follows.
The specific semantic model and policy would obviously change by domain.
The separation between understanding and authority may not.
Maybe the product isn't another application
Someone who reviewed the prototype recently asked what would trigger the process in a real implementation and whether it could automatically select maintenance resources.
Those questions led to an interesting product direction.
Today, the trigger is simply a report submitted through the Maintenance Autopilot web application.
But it does not have to remain that way.
An existing property-management system, tenant portal or messaging channel could send the report and relevant property context to the decision layer.
Maintenance Autopilot could return the governed decision.
If work is authorised to progress, a future resource layer could then consider:
- required trade
- location
- contractor availability
- workload and capacity
- previous performance
- service history
The existing system could continue managing the work order, contractor, payment and property record.
Conceptually:
Tenant report
↓
Existing system
↓
Maintenance Autopilot
Interpret → Apply authority → Govern
↓
Existing workflow / resource network
↓
Human where required
↓
Existing system
↓
Maintenance Autopilot
Interpret → Apply authority → Govern
↓
Existing workflow / resource network
↓
Human where required
That makes Maintenance Autopilot potentially less interesting as another property-management application — and more interesting as a decision-control layer inside systems that already exist.
Those downstream integrations are not part of the current prototype.
But they point toward a commercial question I want to explore after the hackathon:
Could bounded autonomy become an integration pattern for adding agentic capability to existing operational systems without handing an LLM unrestricted decision authority?

: Simple future decision-layer diagram
What I would build next
The evaluation gave me a fairly clear technical priority.
I would not start by adding more autonomy.
I would first strengthen the semantic layer's ability to recognise decision-critical context — particularly information sufficiency, progression, repeat failure and whether an actionable maintenance condition has actually been established.
Only then would I extend downstream execution.
That sequence matters.
If an agent is eventually going to select a contractor, create a work order or trigger a real action, I want the decision boundary to become stronger as capability increases, not weaker.
The question I am leaving this project with
I started Maintenance Autopilot asking:
Can an AI agent help a landlord manage maintenance?
I ended up with a question I think is more important:
As agents become capable of doing more, how do we decide what they are actually allowed to do?
My answer in this prototype is deliberately simple:
Use AI where ambiguity requires intelligence.
Use explicit policy where authority requires control.
Bring the human back when the boundary is reached.
Use explicit policy where authority requires control.
Bring the human back when the boundary is reached.
That is what bounded autonomy means to me.
And it may turn out to be the most reusable thing I built during this hackathon.
Maintenance Autopilot was built for the AWS Agents for Humans hackathon using the Strands Agents SDK, Amazon Bedrock and Amazon Bedrock AgentCore.
Series: Maintenance Autopilot — Agents for Humans Build Journey (3 articles)
- 2Agents for Humans: Autonomy Isn't the Same as Authority This article
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article