AWS Builder Center

Step Functions as the Outer Loop for Long-Running AI Workflows

Step Functions as the Outer Loop for Long-Running AI Workflows

AI Engineer

Introduction

We built an agentic workflow  that discovers AWS resources, collects metrics, and generates cost optimization reports. The agent is autonomous; the LLM decides which tools to call, how many iterations to run, and when the task is complete. That means execution time is unpredictable: 10 minutes, 30 minutes, sometimes longer. Lambda's 15-minute ceiling was never going to work.
Beyond timeout, there's an idle compute problem. In this architecture, Amazon EventBridge triggers a Lambda function on a schedule. That Lambda's only job is to start the agent. If the Lambda waits synchronously for the agent to finish, it sits idle, burning compute costs and occupying a concurrency slot for the entire agent execution while doing zero work. The design goal: once the agent starts, the Lambda exits. This post walks through the pattern we used to solve it: Step Functions as a deterministic outer loop wrapping the non-deterministic agent execution inside Amazon Bedrock AgentCore Runtime.
Fire & forget timeline
Two problems fall out of this. First, the invoking Lambda must hand off work to a long-running compute environment without waiting for a response. Second, something must monitor whether the agent succeeded or failed after the Lambda exits. This is the same agent from our previous posts, where we explored how to bring deep observability to it  and how Business Traces capture its intent and behaviour , and the pattern we introduce here is what brings the full picture together.
Amazon Bedrock AgentCore Runtime closes the first gap. AgentCore Runtime is a fully managed, serverless compute service purpose-built for agentic workloads. It provisions a dedicated microVM per session with isolated CPU, memory, and filesystem resources. Each session supports workloads up to 8 hours. Unlike Lambda's ephemeral execution model, AgentCore Runtime maintains agent state across the session lifecycle.
Inside AgentCore Runtime, the agent uses Strands Agents SDK, a Python framework for building agentic workflows. Instead of hardcoding task flows, the LLM drives the agent loop: reading context, planning actions, calling tools, and iterating until it reaches a final answer.
The missing piece: how do you invoke the agent on a schedule, confirm the invocation succeeded, let the invocation Lambda exit, and detect when the agent finishes, all without human input? That is the pattern this post documents.

The Pattern: Step Functions as the Deterministic Outer Loop

The architecture uses four AWS services in a specific sequence: AWS Step Functions orchestrates the workflow, AWS Lambda handles short-lived initialization and invocation, Amazon Bedrock AgentCore Runtime executes the long-running agent, and Amazon DynamoDB stores completion events for polling.
Step Functions Outer Loop Architecture
Step function loop
The flow executes in four stages.
Stage 1: Session Initialization. Step Functions invokes a Lambda function that receives a session_id, records session metadata, and writes a SESSION_INITIATED event to a DynamoDB journal table. The session identifier is then passed to subsequent stages via the Step Functions state.
Stage 2: Fire-and-Forget Invocation. A second Lambda function calls invoke_agent_runtime() via the AWS SDK, passing the session ID to AgentCore Runtime. AgentCore accepts the request and responds immediately. The Lambda receives the acknowledgment and exits. Inside the AgentCore microVM, the @app.entrypoint handler uses asyncio.create_task() to spawn the long-running agent work as a background task. The @app.async_task decorator signals AgentCore that a task is active, toggling the health status to HealthyBusy and preventing the session from being reclaimed until the work is complete. Here is the core pattern from the repository:
1
2
3
4
5
6
7
8
9
10
11
12
13
@app.entrypoint
async def handle_request(payload,context: RequestContext):
session_id = context.session_id
# Launch background work, returns immediately
asyncio.create_task(background_work(payload, session_id))
return {"status": "started","session_id": session_id}

@app.async_task
async def background_work(payload, session_id: str):
# Runs for minutes or hours inside AgentCore Runtime microVM
# Ping status: HealthyBusy while running, Healthy when done
result = await run_agent_workflow(payload, session_id)
await write_completion_event(session_id, result)
The @app.entrypoint handler calls asyncio.create_task() to schedule the coroutine on the event loop without blocking. It returns {"status": "started"} to the calling Lambda within milliseconds. The @app.async_task decorator on background_work signals AgentCore Runtime that a long-running task is active, preventing the session from being reclaimed. The background work continues in the AgentCore Runtime session for up to 8 hours.
Stage 3: DynamoDB Polling Loop. After the invocation Lambda returns, Step Functions enters a polling loop. The state machine executes a Wait state (pauses for N seconds), then a direct DynamoDB query call to check for a completion event, then a Choice state that evaluates the result. If the agent has not finished, the Choice state routes back to Wait. If the DynamoDB item contains a completion or failure marker, the workflow proceeds to Stage 4.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
// Stage 3: DynamoDB Polling Loop
const waitForCompletion = new sfn.Wait(this, 'WaitForCompletion', {
time: sfn.WaitTime.duration(cdk.Duration.seconds(10)),
});

const checkStatus = new CallAwsService(this, 'CheckStatus', {
service: 'dynamodb',
action: 'query',
parameters: {
TableName: props.journalTable.tableName,
KeyConditionExpression: 'PK = :pk AND begins_with(SK, :eventPrefix)',
FilterExpression: '#status IN (:completed, :failed)',
ExpressionAttributeNames: {
'#status': 'status',
},
ExpressionAttributeValues: {
':pk': { S: JsonPath.format('SESSION#{}', JsonPath.stringAt('$.session_id')) },
':eventPrefix': { S: 'EVENT#' },
':completed': { S: EventStatus.AGENT_BACKGROUND_TASK_COMPLETED },
':failed': { S: EventStatus.AGENT_BACKGROUND_TASK_FAILED },
},
ScanIndexForward: false, // Most recent events first for faster completion detection
Limit: 20, // Scan multiple events to handle any out-of-order writes
},
iamResources: [props.journalTable.tableArn],
resultPath: '$.queryResult',
});

const evaluateStatus = new Choice(this, 'EvaluateStatus')
.when(
Condition.and(
Condition.numberGreaterThan('$.queryResult.Count', 0),
Condition.stringEquals('$.queryResult.Items[0].status.S', EventStatus.AGENT_BACKGROUND_TASK_COMPLETED),
),
processResults // Stage 4 state
)
.when(
Condition.and(
Condition.numberGreaterThan('$.queryResult.Count', 0),
Condition.stringEquals('$.queryResult.Items[0].status.S', EventStatus.AGENT_BACKGROUND_TASK_FAILED),
),
handleFailure // Error handling state
)
.otherwise(waitForCompletion);

// Chain the polling loop
waitForCompletion
.next(checkStatus)
.next(evaluateStatus);

// Grant Step Functions read access to DynamoDB
agentEventsTable.grantReadData(stateMachine);
Stage 4: Completion Detection. When the Strands agent finishes its work inside AgentCore Runtime, it writes an event to DynamoDB with a status of AGENT_BACKGROUND_TASK_COMPLETED or AGENT_BACKGROUND_TASK_FAILED. The next polling cycle picks up this event, and Step Functions routes to the appropriate downstream state for processing results or handling the failure.

AgentCore Runtime's Integration Model and the Polling Pattern

Why poll at all? AWS Step Functions supports a callback pattern called .waitForTaskToken that eliminates polling entirely. The state machine pauses execution, passes a unique task token to the downstream service, and resumes only when that service calls SendTaskSuccess or SendTaskFailure with the token. The workflow incurs zero additional state transition costs while paused and can wait for up to one year.
At the time of writing, AgentCore Runtime managed its own async lifecycle through /ping health status toggling rather than calling SendTaskSuccess, which meant the .waitForTaskToken callback pattern wasn't applicable here. The Wait → Query → Choice polling loop was the pragmatic solution, though as of March 2026, Step Functions now includes optimized AgentCore integrations  that may simplify this pattern in future implementations.

Hybrid Orchestration: Deterministic Wrapping Non-Deterministic

This pattern combines two distinct orchestration models that operate at different layers.Hybrid Orchestration: Outer Loop and Inner Loop
Outer & Inner
The left column (blue) represents the Step Functions outer loop: EventBridge triggers the workflow, two Lambda functions handle session init and agent invocation, then a polling loop (Wait, GetItem, Choice) monitors for completion. The right column (orange) represents the AgentCore inner loop: the Strands agent runs autonomously, calling tools and iterating as the LLM decides. Two cross-column connections link the loops: the red dashed "fire-and-forget" arrow from Lambda to AgentCore, and the green dotted "completion event" arrow from DynamoDB write back to DynamoDB read. DynamoDB serves as the handshake layer between the deterministic outer loop and the non-deterministic inner loop.
The outer loop: Step Functions (deterministic, rule-based). Step Functions is a visual workflow engine with explicit steps, condition-based transitions, and deterministic execution. It gives you auditability through a visual execution console, built-in error handling with configurable retries, and support for workflows running up to one year. The state machine doesn't know or care what the agent does internally. It manages session lifecycle, polls for completion, and handles failures.
The inner loop: Strands Agents SDK (non-deterministic, agentic). Inside AgentCore Runtime, the Strands agent runs autonomously, with the LLM controlling tool selection and iteration. A separate post  will cover the inner loop mechanics in detail, including how Strands orchestrates multi-agent graphs internally.
What matters for this pattern is the boundary between the two loops. Step Functions enforces the contract: the agent starts, the agent finishes (or fails), and the workflow proceeds. The LLM retains full autonomy within its execution context inside AgentCore Runtime. The outer loop does not constrain the agent's tool choices, iteration count, or reasoning path.
The key benefit is controlled autonomy. The agent retains full autonomy within its execution context, but the outer loop enforces timeouts, retries, and error routing. Each stage has its own failure state. If session initialization fails, the workflow terminates with SessionInitializationError before the agent is ever invoked. If the AgentCore invocation fails, the workflow catches it separately from an agent processing failure. If the DynamoDB status check itself errors, that's a distinct failure path. The outer loop doesn't just detect agent failure it isolates failure at every boundary. Non-deterministic agentic behavior stays contained inside deterministic guardrails.

When to Use This Pattern

This pattern applies when three conditions are true.
The agentic workload exceeds Lambda's 900-second limit. If your agent completes within 15 minutes, invoke it from Lambda and wait for the response. No orchestration overhead needed. This pattern targets workloads where the LLM-driven agent loop runs for 30 minutes to several hours, and you can't predict execution time at deploy time because the agent decides its own iteration count.
The downstream runtime lacks a Step Functions callback path. AgentCore Runtime implements its own HTTP contract and doesn't support Step Functions optimized integrations. It can't call SendTaskSuccess to resume a paused workflow. Any containerized or third-party service without callback support qualifies for this pattern.
The workflow requires deterministic guardrails around non-deterministic agent behavior. You need retries, timeouts, error routing, and auditability around an agent whose internal execution path is unpredictable. Step Functions provides these guarantees while the agent retains full autonomy inside its runtime.
The core value is separation of concerns. EventBridge triggers the workflow on a schedule. Lambda handles short-lived initialization and invocation. Once the agent starts, the Lambda exits and frees its concurrency slot. Amazon Bedrock AgentCore Runtime provides the long-running compute environment with microVM isolation. DynamoDB serves as the handshake layer between the agent and the state machine. Step Functions closes the loop by detecting completion and routing to the next action. Each service does one job, and no service sits idle waiting for another.
This pattern is not limited to agentic workloads. Any workload that runs longer than 15 minutes inside a runtime that lacks Step Functions callback support benefits from the same architecture: fire-and-forget invocation, event-based completion signaling via DynamoDB, and a deterministic outer loop to close the feedback cycle.The complete implementation is available in the repository .
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article