AWS Builder Center

Agents for Humans: AI agents that check on neighbourhood residents during a heat wave

Building a Good Neighbour agent that phones an opt-in list of at-risk neighbours during a heat warning, sorts who is fine from who isn't, and pages a human only for the calls that need one. Plus what Cedar policies and a deterministic backstop caught what the model didn't.

In the summer of 2021, a heat dome sat over the Pacific Northwest for a week. In British Columbia alone, the Coroners Service confirmed 619 heat-related deaths. 98% died indoors. 56% lived alone. 67% were over 70. The warnings went out and reached phones, yet nobody arrived.
Plenty of neighbourhoods already keep a list of who is most at risk. Tenant associations keep one, senior buildings keep one, neighbourhood emergency teams keep one. The problem is when the warning lands, one volunteer with a day job has to work through 50 names by phone, with each call taking about 4 minutes minimum if everything goes well.
So for the AWS Agents for Humans hackathon, I'm building Doorstep, a Good Neighbour agent. When a heat warning covers the group's area, Doorstep wakes up on its own, phones every resident on the opt-in list starting with the most at risk. It will then work out who is fine and who isn't, send a volunteer where a person is needed, and interrupt the block captain only for the decisions a person has to make.
The captain isn't meant to watch a dashboard. They are meant to get three messages on a hot afternoon and tap a button on each one.
I'm Ansh, a full-stack engineer in San Francisco. I build side projects with the goal of making everyday lives easier.
Doorstep architecture. Five stages run left to right. One, alerts and ingress: the National Weather Service alerts API, an EventBridge schedule every ten minutes, AWS Lambda handling the poller, webhooks, replay, voice links, caps and kill switch, and an API Gateway HTTP API. Two, the coordinator on Amazon Bedrock AgentCore Runtime, containing a Strands Graph of alert assessor to triage to outreach; a classifier paired with a deterministic backstop where the model proposes and phrase rules may only raise severity, never lower it; a dispatcher running the Strands agent loop with tools to escalate, send a volunteer, notify family and place a check-in call; Strands interrupts that snapshot the session to S3 and resume; and Strands hooks for approval, audit and model-call guarding. It runs on Amazon Bedrock with Nova 2 Lite for agents, Nova Micro for simulated residents and Nova 2 Sonic for voice. Three, check-ins over voice: residents answer in a browser or on a phone; a voice agent on AgentCore Runtime runs a Strands BidiAgent on Nova 2 Sonic over a WebSocket; a check-in worker Lambda takes jobs from SQS and dials through the Twilio subaccount; and a phone bridge runs the same voice code over μ-law audio. Both paths send a mid-call page and the transcript back to the coordinator. Four, Cedar on the tool boundary: Strands CedarAuthorization checks every dispatcher and outreach tool call against the call allowlist, consent, quiet hours and volunteer distance, allows no real calls in sandbox mode, re-checks the allowlist in code at the dialer, and audits every denial. Five, outputs: Telegram decisions to the block captain including mid-call pages, Telegram door-knock tasks to volunteers, a React dashboard on CloudFront, and an incident report. A dashed return path shows a captain or volunteer tap going through the webhook to resume the interrupt. Supporting services are DynamoDB, S3, SQS, SSM Parameter Store, AgentCore Memory and AgentCore Observability.
An alert wakes the coordinator, which triages the roster, runs the check-ins, and routes only the exceptions to a human.
An agent is the most necessary solution for this issue because of what comes back on the emergency calls. Real answers are messy, as someone says they're fine, then mentions they aren't sure what day it is. Someone answers in another language and/or someone is hard of hearing. This needs to be interpreted successfully and quickly to decide what would actually help, and write a volunteer a brief containing only what that volunteer needs to know. The model decides what to do. Code decides what it's allowed to do.

One agent, many hazards

Heat is what I'm shipping and testing properly. But nothing heat-specific lives in the agent code. Each hazard is a profile file holding everything that differs: which alerts activate it, the risk weights, the questions asked on the call, the danger-sign phrases, the kinds of help, and whether relief centres are for cooling or warming.
1
2
3
4
5
6
7
8
9
10
11
12
13
id: heat
display_name: "Extreme heat"

alert_events:
- event: "Extreme Heat Warning" # name in use since March 2025
vtec: "XH.W"
activation: auto
- event: "Excessive Heat Warning" # pre-2025 name; the June 2021 replay uses this
vtec: "EH.W"
activation: auto
- event: "Heat Advisory"
vtec: "HT.Y"
activation: ask
Two tests keep this honest. One loads a made-up "purple smoke" profile and checks that the questions, weights, danger signs and relief-centre kind all change with no code edit. The other fails the build if a heat word appears anywhere in the agent package. If the drill is pointed at the made-up profile, it declines the heat alert as not its hazard.

The policy layer earned its place

Cedar policies sit on every tool call. Calls go only to residents who opted in and are on the allowlist, and only between 8:00 am and 9:00 pm local time, unless the alert is extreme. Resident details travel only to the volunteer actually assigned. The agent can never place an emergency call itself. Instead, it tells the resident to call 911 and pages a human.
This caught a bug on the first full run. Harold hadn't answered three calls, so the dispatcher tried to send a volunteer:
1
2
#136 agent:dispatcher policy r05 assign_volunteer DENY: Access denied by Cedar policy;
decided on volunteer_available=False, volunteer_distance_km=0.7, volunteer_id=Sam
It had passed a first name, "Sam", where the tool wants an ID. No volunteer matched, so the enricher reported nobody available at a distance of 0.7 km, and the policy refused before the tool ran. Nothing went out. The model read the denial, stopped trying, and paged the captain instead.

The backstop caught what the model missed

One persona is written to hide a red flag. He says he's fine, then, cheerfully, that he isn't sure what day it is. The classifier read the transcript and called it OK. A phrase rule found it and raised the case:
1
#051 backstop r02: backstop raised OK to URGENT: matched ['confusion'] via ['not sure what day']
The backstop can only raise severity, never lower it, and every disagreement gets logged. That rule is data in a profile file, not code, so it can be reviewed without reading any Python.
Terminal board during a Doorstep drill. Twelve fictional residents listed by risk score and call wave, each with a state: resolved, calling, needs help, urgent, or queued. A header line notes that the roster and personas are fictional and the alert text is a real archived National Weather Service product.
The board 28 seconds into a drill. Five calls running at once, three residents escalated to the captain, and a red flag on Walter that the phrase rules caught after the classifier had called him OK.

Where it stands

A full local drill has been conducted against the archived June 2021 Portland warning, with 12 fictional residents and simulated people answering the phone:
  • 12 of 12 cases settled: 7 fine, 2 needing help with a volunteer sent, 2 urgent escalated, 1 no-answer escalated after three attempts
  • 7 decisions put to the captain; everything else handled
  • 0 policy denials, 138 audit events
  • 114 seconds, against a 240-second limit
The same list of people is roughly three hours of phone calls for one volunteer.
Next: pausing the agent mid-run for a human decision and resuming after a tap on Telegram, then AgentCore, then giving it a real phone line with Nova Sonic. The two bugs that cost me the most time, a Cedar schema that denies everything and a structured-output retry loop, are the subject of the next post.
All residents, volunteers and personas are fictional. The alert text is the real archived NWS product. The code goes public with the submission on Monday.
Built with the Strands Agents SDK, Amazon Nova and Amazon Bedrock AgentCore, for the AWS "Agents for Humans" hackathon.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article