AWS Builder Center
Serving Strands Decider on Amazon Bedrock AgentCore Runtime: a 2B Decision Model in a 2 vCPU, 8 GB MicroVM with the AgentCore CLI

Serving Strands Decider on Amazon Bedrock AgentCore Runtime: a 2B Decision Model in a 2 vCPU, 8 GB MicroVM with the AgentCore CLI

I packaged the official Strands Decider models for Amazon Bedrock AgentCore Runtime and deployed them with the new AgentCore CLI. One container serves any decider, chosen by environment variables. An int8 CPU build fits the 2B model in the 8 GB microVM, gives the same decision as the reference precision on 28 of 30 test questions, and starts in about 70 seconds. Measured numbers, the memory work, two examples and the open-source repo.

Senior Solutions Architect | AWS AI Hero | OSS: Strands - robots, harness-sdk, decider, stan, box
Strands Decider models are small (2B) models that answer typed questions with calibrated probabilities. You give one a piece of text and a few questions: a yes/no question, a pick-one-option question, or a rating on a rubric. It returns a probability for every possible answer. Because the probabilities are calibrated, an application can act when the model is sure and send everything else to a person.
I wanted those models behind a serverless endpoint that any agent or application in an AWS account can call with IAM credentials, with no GPU and no servers to look after. Amazon Bedrock AgentCore Runtime runs containers in per-session microVMs, and the new AgentCore CLI deploys them from a project file. This article covers how I packaged the official deciders for it, what it took to fit a 2B model into a 2 vCPU, 8 GB microVM, and what the deployment measured.
The result is an open-source sample, strands-decider-agentcore :
  • One container image serves any decider. The model, its pinned Hugging Face commit and the precision are environment variables.
  • A small command, deciderctl add, writes a decider into an AgentCore CLI project, and agentcore deploy builds and ships it.
  • A client library and four examples call the endpoint: plain boto3, a Strands agent whose risky tool calls are checked by a decider, a throughput test, and a ticket-triage web console.
Every number in this article comes from a real deployment in us-east-1 on 7 October 2026, with strands-decider-2B-hobson-v21 and strands-decider-2B-hobson-v19 at pinned commits.
Alt text: Architecture: on your machine, deciderctl add writes runtimes into agentcore/agentcore.json and agentcore deploy builds the image with CDK and CodeBuild; in the AWS account, the image goes to Amazon ECR and each decider becomes one AgentCore Runtime, where every session is a 2 vCPU, 8 GB microVM that downloads the model, merges the LoRA adapter, quantizes it to int8 and serves /invocations on port 8080; applications call InvokeAgentRuntime with SigV4 and a session id of at least 33 characters
The deployment. One image, one runtime per decider model, one warm microVM per session.

How AgentCore Runtime runs a container

AgentCore Runtime takes a container that serves two HTTP routes on port 8080: POST /invocations for requests and GET /ping for health. The bedrock-agentcore Python SDK provides both through BedrockAgentCoreApp, an @app.entrypoint handler and an optional @app.ping handler.
Three behaviours shaped the design:
  • Every session id gets its own microVM. The first call with a new runtimeSessionId starts a fresh microVM and container. Later calls with the same id reach the same warm container until it has been idle for the configured timeout (default 15 minutes, configurable from 60 seconds to 8 hours) or reaches its maximum lifetime of 8 hours. Session ids must be at least 33 characters long.
  • The microVM is small. I ran a probe container first to see what a session gets: 2 vCPUs reporting Arm Neoverse N1 cores (Graviton2 class, with dot-product instructions but no bf16 or int8 matrix instructions), 8.2 GB of memory with about 7.8 GB available, and 9.4 GB of disk. Hugging Face downloads ran at about 280 MB/s.
  • Health can report busy. /ping can answer HealthyBusy while the container works, so a session can load a model in the background and still answer health checks.
A decider has to load inside each new session, so the cold start and the memory peak both matter.

The AgentCore CLI project and deciderctl

The AgentCore CLI  (@aws/agentcore, version 0.31.1 in my runs) creates a project directory with an agentcore.json file and a CDK app. agentcore deploy builds an arm64 container image in AWS CodeBuild, pushes it to Amazon ECR and creates the runtimes with CloudFormation. I created a project with agentcore create --no-agent and added runtimes to it with a container build that points at my own runtime/ directory.
Every decider uses the same image. The differences between them are environment variables, which the AgentCore project stores per runtime. So I wrote deciderctl.py, a standard-library Python script that edits agentcore.json for you:
1
2
3
4
5
6
python deciderctl.py target --region us-east-1 # writes agentcore/aws-targets.json from your credentials
python deciderctl.py add triage --model v21 # official release, pinned to its exact commit
python deciderctl.py add legacy --model v19
agentcore deploy -y
python deciderctl.py health triage # starts a session and reports load progress
python deciderctl.py ask triage --state "Order #4411 arrived broken. Refund me today." --noul "Should a human agent handle this?"
deciderctl add resolves v21 and v19 to the exact Hugging Face commits of the official releases, so a cold start next month loads the same weights as today. It also sets the idle timeout (30 minutes by default), the maximum lifetime, the precision and the prompt limits, and tags the runtime with the model name. Any other checkpoint in the strands-decider format works with --model org/repo@commit.
Alt text: Terminal: deciderctl adds the triage and legacy deciders pinned to their commits, agentcore deploy -y passes through every step from loading the deployment target to deploying the stack, deciderctl list shows both runtime ARNs, and deciderctl health triage reports the ready state, the time of each load stage, 2.41 GB of weights and 4.24 GB resident memory
Adding two deciders, deploying them, and checking one. The account ID is masked.
The first deploy of both runtimes took about 5 minutes, with the image builds running in CodeBuild. The image is 1.56 GB and holds no weights; each session downloads them at start-up. The first agentcore deploy in an account also turns on CloudWatch Transaction Search, which is an account-level setting.

Fitting a 2B model into 8 GB

My first attempt did not fit. strands-decider loads its 2B models on CPU by converting the whole model to fp32, which is about 9 GB for this model before any request arrives. The microVM has about 7.8 GB available.
The runtime now loads the model in five steps, which I found by measuring each step in a local container limited to 2 CPUs and 8 GB:
  1. Load in bf16 with PyTorch SDPA attention, merge the decider's LoRA adapter into the base model, and turn off strands-decider's CPU upcast to fp32.
  2. Quantize all 186 linear layers to int8 with PyTorch dynamic quantization, one layer at a time, using per-channel weights. I chose the qnnpack engine because it packs the int8 weights with about twice their size in extra memory, where the onednn engine used about three times their size and pushed the peak over the limit on long prompts.
  3. Keep the 248,320 by 2,048 token embedding in bf16, behind a small wrapper that returns fp32. That saves 1 GB.
  4. Re-store every remaining tensor. The merge writes into pages of the memory-mapped checkpoint file, and those copy-on-write pages stayed resident until every tensor referencing the file was replaced. This one step took the resident memory from close to 8 GB to about 4.1 GB.
  5. Return freed memory to the operating system with malloc_trim and allocator settings (MALLOC_ARENA_MAX=2), and size PyTorch's thread pool from the container's CPU quota. Before that fix, PyTorch in my local test container started 14 threads on 2 allowed CPUs.
Alt text: Bar chart of memory against the 8 GB microVM limit: the reference fp32 CPU load at an estimated 9.0 GB, above the limit; the int8 build at 6.55 GB peak while loading, 4.24 GB resident when ready, and 2.41 GB of weights
Memory of the 2B decider. The fp32 figure is an estimate; the other three were measured on AgentCore Runtime.
The longest prompt I allow is 4,096 tokens. A 3,100-token input completed under a 7.8 GB memory cap in my local test, and the DECIDER_MAX_TOKENS setting lets you lower the limit.

Cold start

The container answers health checks as soon as it starts, loads the model in a background thread, and reports HealthyBusy until the model is ready. A {"action": "health"} payload returns the current stage, the time of each finished stage and the live memory, without waiting for the model.
Alt text: Stacked bar of one cold start: download decider 1.4 s, download base model 16.5 s, load 2.9 s, merge LoRA 10.3 s, quantize to int8 14.0 s and a warm-up question 15.6 s, 61 s in the container and about 70 s from the first call
One session's cold start on AgentCore Runtime, by stage.
From the first call to a ready model took about 70 seconds. The other sessions I started took between 43 and 74 seconds. Downloading 4.6 GB from Hugging Face took 18 seconds; the rest is CPU work. The warm-up question runs the model once before the session reports ready, so the first real request does not pay for one-time setup.

Calling a decider

The invocation payload is strands-decider's own request format. A request has a state (text or JSON) and named questions:
1
2
3
4
5
6
7
8
9
10
{
"state": "Order #4411 arrived with a cracked screen. I want my money back today.",
"questions": {
"escalate": {"type": "noul", "instructions": "Should a human agent handle this ticket?"},
"queue": {"type": "choice", "instructions": "Which team handles it?",
"criteria": {"returns": "refunds and damaged items", "billing": "payments and charges"}},
"urgency": {"type": "score", "instructions": "How urgent is it?",
"criteria": ["low: can wait a week", "medium: within two days", "high: today"]}
}
}
The response has one answer per question: noul is P(yes) for a yes/no question, a choice answer has a probability per option, and a score answer has a probability per rubric level and the expected level. The runtime also accepts {"batch": [...]} for several requests in one call and {"action": "health"}. Errors come back as {"error": {"code": ..., "message": ...}} rather than as a failed session, with a loading code while a new session is still loading the model.
With plain boto3, two details matter:
1
2
3
4
5
6
7
8
9
10
11
import json, uuid
import boto3
from botocore.config import Config

client = boto3.client("bedrock-agentcore", config=Config(read_timeout=900))
resp = client.invoke_agent_runtime(
agentRuntimeArn=TRIAGE_ARN,
runtimeSessionId=f"triage-{uuid.uuid4()}", # at least 33 characters; reuse it to stay warm
payload=json.dumps(request).encode(),
contentType="application/json", accept="application/json")
answer = json.loads(resp["response"].read())
The default boto3 read timeout of 60 seconds is shorter than a cold start, so I raise it. Reusing the session id keeps requests on the same warm microVM.
The sample's DeciderClient wraps this. It keeps a pool of session ids, where pool=4 means four warm replicas. It retries while a session is still loading, and the same class talks to a local container with url="http://localhost:8080".

Matching the reference precision

int8 changes the arithmetic, so I compared answers on 14 realistic decisions with 30 questions: support tickets, content moderation, agent tool-call safety, grounded answers, a contract clause and a mixed review. scripts/collect_answers.py records every answer and scripts/compare_answers.py compares the top decision per question.
comparisonsame decisionlargest probability change
int8 on AgentCore against int8 in a local container30 of 300.000
int8 on AgentCore against bf16, the reference precision28 of 300.201
The two changed answers were close calls in the reference as well: urgency medium against high on one ticket, and sentiment mixed against negative on a review written to be mixed. The int8 build gave identical probabilities on AgentCore and in a local arm64 container on my laptop, so a decision tested locally carries over to the deployment. A runtime added with --quant bf16 serves the reference precision if you need it, at a higher memory cost.

Speed, measured

Once a session is warm, a decider on AgentCore Runtime takes about 7 seconds per question:
requestlatency
1 yes/no question, 94-token prompt7.6 s
3 questions, 248-token prompt20 to 23 s
14 decisions (30 questions), 1 warm session6.3 decisions per minute, median 9.5 s per decision
the same 14 decisions, 4 warm sessions11.4 decisions per minute, median 16.8 s per decision
Alt text: Terminal: the scale example with one warm session reports 14 decisions in 134 seconds, 6.3 decisions per minute, median 9.5 s and p95 12.5 s; with four warm sessions it reports 14 decisions in 73 seconds, 11.4 decisions per minute, median 16.8 s and p95 26.0 s
Throughput with one and with four warm sessions.
I profiled a forward pass to see where the time goes: 76 percent of it is in the int8 linear layers. The Qwen3.5 linear-attention layers fall back to plain PyTorch on CPU, which I suspected first, but they account for only a small share. The model is compute-bound on two cores without int8 matrix instructions. strands-decider already encodes the state once and reuses its key-value cache for every question, so the cost that grows with each question is the question's own text and options.
Four sessions gave 1.8 times the throughput of one, with slower individual requests, so I measure a pool's gain on real traffic before sizing it. These numbers make a decider on AgentCore Runtime a good fit for decisions that can take seconds, such as triage, review queues, agent guardrails and batch labelling. Tight interactive loops need a GPU host instead.

Example: a decider as a guardrail for a Strands agent

The first example puts a decider in front of an agent's tools. The agent runs on Amazon Nova Pro with three tools: look up an order, issue a refund, and close a customer's account. A Strands hook on BeforeToolCallEvent sends the customer's request and the planned call to the decider before any tool runs, and cancels the call when the probability of serious, irreversible harm is 0.5 or higher.
1
2
3
4
5
6
7
8
9
10
11
12
class DeciderGuard(HookProvider):
def register_hooks(self, registry, **_):
registry.add_callback(BeforeToolCallEvent, self.check)

def check(self, event):
call = {"tool": event.tool_use["name"], "input": event.tool_use["input"]}
answers = self.decider.decide(
{"customer_request": self.request, "planned_tool_call": call},
{"harm": noul("Could this tool call cause serious harm that cannot be undone, such as deleting data, "
"closing an account or paying out more money than the customer is owed?")})["answers"]
if answers["harm"]["noul"] >= self.threshold:
event.cancel_tool = "Blocked by policy: a person must approve this call."
Alt text: Terminal: for a refund request the decider rates lookup_order at 0.23 and issue_refund at 0.17 probability of irreversible harm and allows both, and the agent confirms the refund; for a request to close account C-88 and wipe everything it rates close_customer_account at 0.79 and blocks it, and the agent tells the customer a person will review the request
The refund the customer asked for goes through. Closing and wiping the account waits for a person.
My first version asked a single question: "is this call appropriate and safe to run without a human?". It blocked the harmless order lookup at 0.66 and rated the account closure higher, at 0.76. Asking about the specific risk, irreversible harm with examples of it, separated the calls clearly: 0.17 to 0.23 for the lookup and refund, 0.79 for the closure. In a direct test, the guard did not flag a refund ten times larger than the order. A check like that belongs in ordinary code beside the decider, and the example's documentation says so.

Example: a ticket-triage console

The second example is a small FastAPI app with one HTML page. It sends each support ticket to the decider with three questions: does it need a person, which team, and how urgent. It acts only on confident answers: P(needs a person) of 0.25 or less goes to an automated reply, 0.75 or more goes straight to the team's queue, and anything between waits for a person to review it. The team must have a probability of at least 0.6.
Alt text: The triage console: the ticket list on the left with routed and review tags; for a lawyer's email about a client injured by the product, P(needs a person) is 0.76, the legal team has 92 percent, high urgency 52 percent, and the console assigns it to a human agent in legal with a reply today, decided in 29 seconds
A legal threat, routed to the legal team without review because every answer was confident.
Question wording mattered here too. With the plain question "should a human agent handle this ticket?", every ticket landed between 0.3 and 0.65 and nothing was routed. Adding true and false criteria to the yes/no question gave a clear spread: 0.09 for a thank-you note, 0.21 for an address change, 0.76 to 0.81 for a legal threat and an angry repeat complaint. With those criteria, 4 of the 8 sample tickets were routed without review and the other 4 went to a person, which is the behaviour a calibrated model is meant to give.

What I fixed along the way

Three problems surfaced only on the real deployment:
  • Every command started a new microVM. My first deciderctl made a random session id for each run, so each ask paid a full cold start. It now uses a stable session id per decider.
  • A fresh clone could not deploy. The AgentCore CLI does not install the CDK app's npm dependencies, and agentcore create had done that for my copy. A clean checkout failed with tsc: command not found until I ran npm ci in agentcore/cdk. The quickstart now includes that step, and I confirmed it with agentcore deploy --dry-run on a fresh clone.
  • The self-reported load time did not match the clock. Some sessions reported load times longer than the time since my first call, so the throughput example now measures time to ready on the client side.

Limits

These are the limits I measured or know of:
  • The deployment is text only. The CPU build refuses requests with images.
  • Latency is seconds per question, as measured above, and a pool of sessions did not scale linearly in my test.
  • Each new session pays a cold start of about a minute. A longer idle timeout keeps sessions warm at the cost of memory held between requests.
  • The parity check covers 30 questions from my own decision set, not a benchmark.
  • Larger deciders will not fit the 8 GB microVM with this approach.

Cost and clean-up

Nothing in the sample runs on a GPU. You pay for AgentCore Runtime while sessions are active, CodeBuild minutes for each image build, ECR storage for a 1.56 GB image, and CloudWatch logs. See AgentCore pricing  for current rates. To remove everything, delete the AgentCore-StrandsDeciders-default CloudFormation stack, then the ECR repository and the CodeBuild logs it leaves behind.

Try it

The sample is at github.com/Vivek0712/strands-decider-agentcore  under the MIT-0 license. The README has the full quickstart, the payload reference, every configuration setting and a troubleshooting table. The models come from strands-labs/strands-decider  and are licensed Apache-2.0 at the pinned commits. Unit tests run without an AWS account or a model download, and CI builds and starts the container on an arm64 runner.
1
2
3
4
5
6
git clone https://github.com/Vivek0712/strands-decider-agentcore && cd strands-decider-agentcore
pip install -e ".[examples]"
python deciderctl.py target --region us-east-1
python deciderctl.py add triage --model v21
(cd agentcore/cdk && npm ci) && agentcore deploy -y
python deciderctl.py health triage
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article