AWS Builder Center

Durable Handoff Multi-Agent Workflow

Durable handoff multi-agent workflow with Amazon Simple Storage Service (S3)

AI Engineer

Introduction

In the previous post , we wrapped the non deterministic agent execution inside a deterministic outer loop using Step Functions and AgentCore Runtime. That post promised a closer look at what happens inside the inner loop of our agent workflow, where we demonstrate a production pattern that pairs prompt level data contracts with Amazon S3 as a durable, inspectable, replayable handoff layer between agents. Hence solving two problems: how to get a non deterministic LLM to produce reliable, structured output within a specific domain, and how to stop downstream agent failures from destroying expensive upstream work. A single analysis agent in our workflow run queries to multiple AWS services and can take minutes, when the downstream agent crashed, that output was gone and we had to start over. The full implementation is available in the sample-agentic-cost-optimizer repository .

The problem: ephemeral data passing in multi-agent workflows

Multi-agent AI systems split complex tasks across specialized agents. In the agentic-cost-optimizer , a cost analysis agent queries multiple AWS services to collect spending data and optimization recommendations. A report agent transforms those raw findings into a customer-facing document. Strands Agents framework's GraphBuilder  wires both agents as nodes in a directed graph and controls execution order.
The data-passing mechanism in most multi-agent frameworks is ephemeral. Strands Agents graph passes output from one node as input to connected nodes, but this orchestration state is not durable by default . If an agent crashes mid-execution, the graph does not resume from the failure point. All completed node work within that execution is lost.
If a report agent crashes after the analysis agent completes a 10-minute data collection run, the analysis results vanish. The team must re-run the entire workflow. Developers cannot inspect what the analysis agent produced. Teams cannot test prompt changes on the report agent without paying the full analysis cost again. The repository  addresses each of these gaps with a single architectural decision: route the handoff through Amazon S3.

Why durable handoffs matter in production

In-memory passing works fine for prototypes where iteration cycles are short and outputs are disposable. Production workflows are different. A single analysis run can take minutes and cost real money in API calls. That output has independent value, and losing it to a process crash wastes both.This pattern builds on AWS prescriptive guidance . AWS lists Amazon S3 as a shared memory option for multi-agent architectures, and the serverless multi-stage AI workflow  pattern uses S3 as both an event trigger and an output layer for chained AI processing stages. The sample-agentic-cost-optimizer  applies that guidance with session-scoped S3 paths and tool-level isolation between agents.

S3 as the durable handoff layer

The sample-agentic-cost-optimizer implements a two-agent workflow using the Strands Agents GraphBuilder. The analysis agent runs first, queries multiple AWS services, and writes its findings to Amazon S3. The report agent runs second, reads those findings from S3, and generates a customer-facing report. The S3 object, not an in-memory variable, is the handoff mechanism.
S3 Handoff Architecture Between Analysis and Report Agents
handoff
Session-Scoped Paths provide isolation between concurrent runs. Each workflow execution writes to s3://{bucket}/{session_id}/analysis.txt. The session_id originates from the caller's invocation of the agent. The Strands Agents @tool  decorator passes a tool_context  dictionary into the storage tool at runtime . This dictionary contains invocation_state, the caller-provided keyword arguments, which includes the session_id . The storage tool extracts the session identifier from the invocation state and constructs the S3 path automatically. Agents never build S3 paths themselves. They call the storage tool with a filename, and the tool handles the rest. The full storage tool implementation is available at the repository .
Tool-Level Isolation enforces a boundary between agents. Each agent receives a distinct tools  list at instantiation. The analysis agent's tools list includes the storage , journal tools and AWS service query tools. The report agent's tools list includes only storage and journal tools. This separation means the report agent cannot accidentally trigger a new analysis. The following snippet illustrates how the graph wires each agent with its own tool set:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
# See repository for full agent configuration
# Analysis agent: storage + journal + AWS service tools
analysis_agent = Agent(
system_prompt=analysis_prompt,
tools=[use_aws,
calculator,
storage,
journal,
current_time_unix_utc,
convert_time_unix_to_iso,]
)

# Report agent: storage and journal only
report_agent = Agent(
system_prompt=report_prompt,
tools=[storage,journal]
)

# Wire agents into a graph with explicit execution order
graph = GraphBuilder()
graph.add_node(analysis_agent, "analysis")
graph.add_node(report_agent, "report")
graph.add_edge("analysis", "report")
graph.set_entry_point("analysis")
graph.set_execution_timeout(900)
workflow = graph.build()
Both agents share a common schema. The analysis prompt defines the exact output structure of analysis.txt. The report prompt defines the exact input structure it expects. The S3 filename ties them together without either agent knowing the other's implementation.

Four benefits: durable, inspectable, replayable, bounded

DimensionIn-Memory EphemeralS3 Durable Handoff
DurabilityUpstream output lost if downstream agent crashes after receiving in-memory resultPersists in S3, survives downstream restarts
InspectabilityNot accessible after executionRead directly via S3 console or CLI
ReplayabilityMust re-run entire workflowRe-run downstream agent only
Crash recoveryComplete restart required, upstream work repeatedResume from last persisted artifact
Agent isolationNo enforcement boundaryTool-level separation, S3 as boundary
LatencyMinimal (in-process)Additional S3 read/write latency
InfrastructureNoneS3 bucket + IAM configuration

The S3 handoff pattern delivers four properties that in-memory passing lacks.

Durability means the analysis output persists beyond process failures. If the report agent crashes, the analysis results remain in S3. The team can restart the report agent and point it at the same session_id. This preserves completed work from upstream agents, reducing the cost of downstream failures.
Inspectability means developers can read analysis.txt directly in the S3 console or via the AWS Command Line Interface (CLI). When a report contains an unexpected recommendation, the developer can open the analysis file and check whether the analysis agent produced that recommendation or the report agent misinterpreted it. This separates "analysis bug" from "report bug" cleanly.
Replayability means developers can re-run the report agent against the same analysis.txt to test prompt changes. A developer can modify report_prompt.md, trigger the report agent with the same session_id, and compare the new output against the previous version. This cuts iteration time because only the report agent re-runs, not the full workflow.
Bounded isolation means each agent receives only the tools it needs at instantiation. The analysis agent receives the storage tool plus AWS service query tools. The report agent receives only storage and journal tool. This prevents the report agent from accidentally triggering a new analysis, keeping each agent focused on its designated role.

How the prompt contract defines what flows through S3

The S3 handoff layer solves the transport problem: how data moves durably between agents. The prompt contract solves the schema problem: what data the handoff file contains and how the downstream agent interprets it. Both analysis_prompt.md and report_prompt.md in the sample-agentic-cost-optimizer use a phased approach. Each prompt defines distinct phases that grant the agent autonomy within its domain while both prompts serve one shared goal: a complete, accurate cost optimization report.
How the Prompt Contract Ensures Valid S3 Handoff Data
Tiered Data Requirements gate which recommendations the agent can make. The analysis prompt defines three tiers: configuration-only, metrics-required, and logs-required. Each recommendation type maps to a specific tier. If the agent lacks metrics access, it skips metrics-dependent recommendations and documents the gap. This ensures analysis.txt contains only recommendations backed by available data, reducing the risk of hallucinated recommendations.
Graceful Degradation Chains handle failures without stopping the workflow. If the agent receives an AccessDenied error when querying logs, it skips log-based analysis, continues with metrics-based analysis, and documents the gap in a "Gaps & Limitations" section. Every failure gets documented, not hidden. The report prompt mirrors this approach: if the analysis file contains gaps, the report agent documents those limitations in the customer-facing output rather than inventing data to fill them.
Tool Fallback Windows prevent retry loops. The analysis prompt specifies a 5-step fallback chain for time-window queries: 30 days, then 15, then 7, then 3, then 1. If a 30-day metrics query fails, the agent tries 15 days. The prompt includes an explicit rule: "NEVER reuse timestamps from failed attempts." This breaks a common failure mode where the model retries the same failing request indefinitely.
Mandatory Partial Results ensure the downstream agent always receives input. The analysis prompt instructs the agent to produce analysis.txt regardless of how many individual steps encounter errors. The workflow continues through journaling failures and documents storage issues along the way. The report agent receives a file in the expected format, and this pattern maximizes successful end-to-end completion.
The phased structure in each prompt creates a division of labor. The analysis prompt's phases focus on data collection, validation, and gap documentation. The report prompt's phases focus on interpretation, formatting, and customer communication. Each prompt operates independently within its phases, but the output contract of one matches the input contract of the other. analysis.txt is the shared boundary.

When to use S3 handoffs

Not every multi-agent workflow needs durable handoffs. The added latency of S3 read/write operations and the infrastructure setup (S3 bucket + IAM configuration) add complexity. Three scenarios justify the pattern.
Use S3 handoffs when upstream agents perform expensive work. The sample-agentic-cost-optimizer's analysis agent queries multiple AWS services across different time windows and metric types. Losing that output to a crash wastes real money.
Use S3 handoffs when you need to iterate on downstream agents independently. Prompt engineering on the report agent means running it repeatedly against the same input. Without S3, each iteration re-runs the full workflow. With S3, developers can modify report_prompt.md and re-run only the report agent against the persisted analysis.txt.
Use S3 handoffs when developers need to debug agent-to-agent communication. When a customer report contains an unexpected recommendation, the developer needs to determine whether the analysis agent produced it or the report agent misinterpreted it. S3 makes the intermediate artifact directly readable.
In-memory passing suits simple, fast workflows. When upstream agents complete quickly and produce small outputs, in-memory propagation through the Strands Agents graph  keeps the architecture simple.

Conclusion

The sample-agentic-cost-optimizer repository demonstrates one production pattern with two reinforcing components: prompt-level data contracts that encode domain-specific decision logic, paired with Amazon S3 as a durable, inspectable, replayable handoff layer between agents.Take a complex prompt and restructure it as a contract: define tiered data requirements, add graceful degradation chains, and specify the exact output schema the downstream agent expects. Then route the handoff through S3 with session-scoped paths. The Strands Agents tool_context  mechanism passes the session_id via invocation_state, so agents never manage S3 paths themselves.Source code 
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article