AWS Builder Center

Agents for Humans: Pausing a Strands agent for human verification, and resuming after a tap

What it takes for an agent to pause mid-run, survive the process dying, and resume correctly when a human finally taps a button. Includes the two bugs that cost me the most time.

Doorstep phones at-risk neighbours when a heat warning lands, and pages a neighbourhood captain for the decisions a person has to make. I wrote about why in the first post .
This post is about the bit that I thought would be easy but turns out it wasn't. The captain gets a message on their phone. Something might come up where they could only tap the button twenty minutes later, from a different room, after the process that asked them has died. When they do tap, exactly one volunteer task should go out. Not zero, not two, only one.
I'm Ansh, a full-stack engineer in San Francisco.

There are two kinds of interrupts, and picking the wrong one sends a message twice

Strands lets you pause in two places. Each placement behaves differently on resume interaction.
tool_context.interruptpauses inside the tool. On resume, the tool body re-runs from the top, so anything it did before interrupting happens twice. That's fine for escalate_to_captain: its pre-interrupt work is one idempotent upsert, and paging the captain is the whole point of the tool.
BeforeToolCallEvent.interrupt pauses at admission. The executor returns before the tool runs at all, and on resume the body runs exactly once, or never if the hook cancels it. That's the only correct shape for assign_volunteer, because its side effect is a message containing a resident's details leaving the building. It must not be sent as a mistake that one would regret.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
class ApprovalHook(HookProvider):
"""Pauses `assign_volunteer` for the captain when the resident is high risk."""

def approve(self, event: BeforeToolCallEvent) -> None:
if event.tool_use["name"] != "assign_volunteer":
return
# … load volunteer and case
if case.risk.wave != HIGH_RISK_WAVE:
return # low risk: no need to wake anyone

decision = upsert_decision(ctx, name=APPROVE_DOOR_KNOCK, options=[...])
answer = event.interrupt(APPROVE_DOOR_KNOCK, reason={"decision_id": decision.id})

if not isinstance(answer, dict) or answer.get("option_id") != "approve":
event.cancel_tool = f"the captain chose: {answer.get('label', 'declined')}"
return
event.invocation_state["approved_decision_id"] = decision.id
The interrupt ID says where it paused, which turned out to be useful for before_tool_call:tu-restartrather than something inside the tool.

The test I actually trust

The unit tests were green long before the thing worked. What convinced me was a test that runs each half in its own Python interpreter, so the second half has no agent object and no memory of the first.
Half A escalates, pauses, writes the decision ID to disk and exits. Half B is a fresh process holding nothing but the stored session and the decision record. It answers, then answers again with the same button from the same phone.
1
2
3
4
5
6
7
8
9
$ python tests/helpers/restart_half.py a $WORK
HALF A: paused dec-001 interrupt=v1:before_tool_call:tu-restart:0507bf81…

$ python tests/helpers/restart_half.py b $WORK
HALF B: applied | You chose "Send Tom (0.4 km)". | assigned; r04 already assigned to vol-tom
HALF B AGAIN: already_answered | Already answered at 19:52: "Send Tom (0.4 km)". Nothing was sent twice.

$ cat $WORK/outbox.tsv
B volunteer_task vol-tom Doorstep · task for Tom …
One task is sent by the process that got the approval. It runs twice: once on the in-memory store with file sessions, once on DynamoDB and S3, which is how the deployed container is configured.
The cloud version is the same shape with real infrastructure. The agent pauses in one runtime process. I kill that process, and two minutes later a real tap on my phone resumes the session in a different one. Eighteen assertions, including the same Telegram update replayed twice and dropped at the webhook, a genuine second tap answered with already_answered, and a wrong webhook secret getting a 401.

The two bugs I promised last time

A Cedar schema that denies everything. The docs say to pass tools= to CedarAuthorization and let it generate the schema. What it generates declares exactly two session fields, hour_utc and call_count. Doorstep's enricher adds the facts the policies actually need: mode, role, whether the number is allowlisted, volunteer distance, consent. The moment it adds a field the schema doesn't declare, every tool call fails with failed to parse schema from request. Not a deny decision but a blanket deny.
Now the schema is generated from the tool specs plus the enricher's declared fields, and passed in explicitly. The upside is worth more than the fix: action typos, context-attribute typos and type errors all fail at startup instead of mid-drill. A later version of the same mistake cost me again, when the dispatcher generated its schema from its own subset of tools and policies, naming other agents' read-only tools failed with unrecognized action. It's generated from all tools now.
A structured-output retry loop. Nova 2 Lite returned list fields as strings and text fields as lists. Strands re-asks the model on every validation failure, with no cap. My first full drill sat there looking hung: 40 retries on one field. Coercing validators fixed the cause. A guard hook fixed the class of problem by bounding one invocation to a fixed number of model calls and writing the trip to the audit log:
1
2
3
4
5
def before_model_call(self, event: BeforeModelCallEvent) -> None:
self.calls += 1
if self.calls > self.max_calls:
event.cancel = f"stopped after {self.max_calls} model calls in one invocation"
# … recorded in the audit log
A drill always finishes, and the failure is visible.

What the captain sees, and what the volunteer doesn't

This is the part I'd point at if someone asked whether I'd aim this at real people. Same resident, same decision, two messages.
Two device screenshots side by side. The right shows Doorstep's message to the block captain about an urgent resident, including her age band, that she lives alone, what she said on the call, and the response of the captain. The left shows the task sent to a volunteer for the same resident: the unit number, what to do, and to tell the captain if nobody answers. It contains no age, no health details and no quote]
The same resident and the same decision. The captain gets what they need to decide; Tom gets what he needs to knock on a door.
The rule took me a while to phrase. It isn't health facts versus everything else. It's standing facts versus today's call. Roster notes the team already agreed to share, like "hard of hearing", go to the volunteer, and only with consent. What the resident said on the phone this afternoon goes to the captain only. A test checks that rule for every resident on the roster.

Every human choice goes through one function

Telegram taps and dashboard clicks both call respond_to_decision. The only difference is Responder.source.
Two things sit inside it in a deliberate order. The idempotency check comes before the identity check, because answering twice has to be cheap and safe even when the second tap is a stranger's. And the decision is claimed with a conditional write before the agent resumes, so of two taps racing each other, exactly one wins and the other is told what the winner chose.
From my own phone drill: twelve taps across eight decisions, one decision tapped five times, one tapped after it had expired. One applied, four already_answered, zero violations.

What moving to the cloud actually broke

The session layer moved with a config change. S3Storage instead of LocalFileStorage, same session manager on top. One catch: the prefix has to be sessions and not sessions/, because S3Storage adds the separator itself.
The data layer needed five fixes, and every one was a bug that in-memory had been hiding.
  • Duplicate IDs. Decision ids and audit sequence numbers were minted as len(...) + 1. Strands runs sync tool bodies in threads, so two tools could mint the same ID, and DynamoDB latency makes that likely rather than theoretical. Atomic counters now.
  • A shared-object assumption. Code saved a case object loaded before dispatch, which only worked because the in-memory store hands back the same object. DynamoDB returns copies. Identity map plus versioned writes.
  • A forced AWS profile in config that broke boto3 once it was running inside AWS.
  • Telegram callback data with no incident ID, so a stateless webhook couldn't route a tap back to the right runtime session.
  • Secrets in logs. The bot token lives in the Telegram URL path, so HTTP client logs and traces would record it. Worse, botocore at DEBUG logs whole SSM responses, which is every decrypted secret. Both are pinned down now, and a check scans logs for known values before I trust them.
A CloudWatch trace for a Doorstep invocation. The span tree shows invoke_graph calling invoke_agent, and the trajectory below shows the alert assessor, triage, and 59 check-in agent calls. The input panel holds the real archived National Weather Service alert text.
One invocation, end to end: the graph, the triage agent, and 59 check-ins running off a single archived weather alert.

Where it stands

Next up: the voice side, and what forty simulated residents found when I pointed evals at the whole thing. That's the last post.
All residents, volunteers and personas are fictional. The code goes public with the submission.
Built with the Strands Agents SDK, Amazon Nova and Amazon Bedrock AgentCore, for the AWS "Agents for Humans" hackathon.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article