AWS Builder Center
Benchmarking Strands Decider, LLM Classifiers and Bedrock Routing for Model Routing

Benchmarking Strands Decider, LLM Classifiers and Bedrock Routing for Model Routing

I routed 216 graded requests across Amazon Nova and Claude models with Strands Decider, five LLM classifiers, Bedrock Intelligent Prompt Routing and Bedrock Advanced Prompt Optimization. The decider judged request difficulty perfectly (ROC AUC 1.00 against 0.88 for Claude Sonnet 4.6) but cost about 0.0012 USD and 11 to 13 s per decision on a CPU microVM. Here are the numbers and the rules I now follow for when a decision model should judge and when a small LLM should route.

Senior Solutions Architect | AWS AI Hero | OSS: Strands - robots, harness-sdk, decider, stan, box
The best judge in my model-routing benchmark was one I could not afford to ask on every request. Strands Decider, a 2B decision model, separated easy requests from hard ones perfectly (ROC AUC 1.00, against 0.88 for a Claude Sonnet 4.6 classifier), but on CPU in Amazon Bedrock AgentCore Runtime each decision took 11 to 13 s and about 0.0012 USD of compute on a 2 vCPU microVM. Marc Brooker recently reported a 65 ms median for the v19 decider on hardware his post does not state, and the distance between his milliseconds and my seconds decides most of what follows.
Model routing sends each request to the cheapest model that would answer it correctly, and it has to choose before paying for the answer. An average Nova Pro answer in my test cost 0.00033 USD, so on CPU the decision cost more than the answer it was meant to save.
For people who wants to know result first hand:
Experience it here: playground 
The task: route each of 216 requests from 6 task families (36 per family) to the cheapest model that answers it correctly. I used half the requests for tuning and prompt optimization, and every number below comes from the other half, 107 held-out requests. Code grades every answer; no model acts as judge. Quality is the average score out of 100, and cost includes the router.
Router (Nova Micro, Lite, Pro ladder)QualityUSD per 1,000 requestsRouter latency (median)
Always Nova Pro93.10.332none
Always Nova Micro80.50.019none
Oracle (upper bound)96.30.056none
Bedrock Intelligent Prompt Routing87.90.129none (inside the answer call)
Nova Lite classifier, tier question83.20.0420.59 s
Nova Micro classifier, multi-step question88.50.2810.62 s
Claude Haiku 4.5 classifier, multi-step question89.20.4560.94 s
Strands Decider, tier question87.51.06810.7 s
Judge of "needs multi-step reasoning"ROC AUC
Strands Decider1.00
Claude Sonnet 4.6 classifier0.88
Claude Haiku 4.5 classifier0.78
Nova Micro classifier0.60
My position after measuring all of this is narrow. A decision model is the most trustworthy judge of a property of a request: it was perfect on difficulty here, it cannot drift into answering the task, and it returns a calibrated probability I can set a threshold on. Small LLM classifiers and Bedrock Intelligent Prompt Routing are the cheap, fast way to make a routing call per request. The decider earns its cost where one judgment is reused, for example across a whole multi-step process or as a guardrail. No combination I measured beat the best single routers on quality per dollar in this single-step benchmark, and the multi-step savings later in the post are a projection, labelled as one.

What Marc Brooker measured and how it shaped this work

On 4 October 2026 Marc Brooker published Strands Decider: Why Not an Encoder? . He asked why strands-decider-2B uses a modified causal decoder when "encoders are widely used as classifiers", then tested the alternatives. Against v19, the production Qwen3.5-based 2B causal decoder, he built encoder-decoders on T5Gemma 2 with pointer heads (b1 and b2), a 2.6B encoder-only model on the original T5Gemma (e1a), and a latency-optimised e1b with a masked state cache that lets several questions be evaluated in parallel without interference and lets the cache be reused. Frozen ModernBERT and ettin-encoder-1b underperformed for their size.
His findings are careful. e1b slightly beat v19 on JevBench accuracy, within v19's run-to-run variance, with slightly worse calibration; b2 was best but carries 8B total parameters. On speed: "The encoder-only models win handily on latency, with a median of around 35ms compared to v19's 65ms." The advantage fades with length, because e1b "wins handily at lower prompt lengths, but scales worse as the prompt grows": it has 26 full attention layers with quadratic cost, while 75% of v19's layers are Gated DeltaNet with linear cost and a higher fixed floor. On generalisation, e1a beat v19 on seven of nine evaluation tasks but did significantly worse on JevBench (153 against 168). His verdict: "Neither seems like a slam-dunk winner for this kind of work at this size."
Three of his findings line up with mine. First, he treats deciders as classifiers, and the decider generalised well in my test: zero-shot, on a multi-step question it was never trained for, it separated the task families perfectly. That fits his observation that the causal v19 generalises to unseen benchmarks better than the unmasked encoder. Second, v19's linear-cost layers favour long inputs, and the job I end up recommending for a decider, judging a whole multi-step process once, is a long-context job. Third, e1b's reusable state cache with parallel questions describes how I think a decider should be used: one state queried with several typed questions, and the judgment reused across steps.
My measurements qualify his numbers on cost. A 65 ms median makes a decider look almost free as a router. On a 2 vCPU CPU microVM I measured 11 to 13 s per decision, and at that speed the decision costs more than a Nova Pro answer. The economics of a decider as a per-request router depend on where it runs far more than on encoder against decoder, and a twofold speed-up from a different architecture would not change my conclusion on CPU. His "not a slam dunk" matches what I found from the other direction: no single router won on every axis.

The benchmark

The six task families are graded by code. Extract asks for one field from a synthetic order note, scored on the exact value. Sports is BIG-Bench Hard sports_understanding, yes or no. JSON turns a synthetic contact note into 5-field JSON, scored on the fraction of fields right. Math is GSM8K, scored on the final number. Code is MBPP sanitized, scored on the fraction of unit tests passing. Logic mixes BIG-Bench Hard logical deduction (5 objects), shuffled objects (5) and date understanding, scored on the option letter.
Seven models answered every request at temperature 0: Nova Micro, Nova Lite, Nova Pro, Llama 4 Scout 17B, Llama 3.1 8B, Claude Haiku 4.5 and Claude Sonnet 4.6. Prices come from the AWS Price List API, us-east-1 on-demand, per million input and output tokens: Nova Micro 0.035 and 0.14, Nova Lite 0.06 and 0.24, Nova Pro 0.80 and 3.20, Claude Haiku 4.5 1.10 and 5.50 regional, Claude Sonnet 4.6 3.30 and 16.50 regional. The runs took place in us-east-1 on 8 and 9 October 2026.
Alt text: Heatmap of test quality by task family for Nova Micro, Nova Lite, Nova Pro, Llama 4 Scout and Llama 3.1 8B. Extract is near 100 for most models, with Nova Micro at 94 and Nova Lite at 89. Sports is the weakest row, from 56 for Llama 3.1 8B to 83 for Nova Pro. JSON is 99 or 100 everywhere. Math is 94 for the three Nova models and Llama 4 Scout and 53 for Llama 3.1 8B. Code runs from 74 for Nova Micro to 91 for Llama 4 Scout. Logic runs from 50 for Llama 3.1 8B and 61 for Nova Micro to 100 for Llama 4 Scout.
Quality by family and model on the test half. The cheapest model already does well on most families.
Nova Micro alone fully solves 79% of the requests, so whatever a router can gain sits mainly in sports, code and logic, where it scores 61, 74 and 61.

The routers under test

Strands Decider (strands-decider-2B-hobson-v21, int8, CPU) ran on AgentCore Runtime in a 2 vCPU, 8 GB microVM with 4 warm sessions, the deployment described in part 1. AgentCore Runtime bills 0.1276 USD per vCPU-hour and 0.0169 USD per GB-hour. I asked it two typed questions. The tier question is a choice: "Which is the cheapest model that will answer this request correctly?", with options micro, lite and pro, each with a description. The multi-step question is a yes/no question with criteria: "Does answering this request correctly need careful multi-step reasoning?"
Five LLM classifiers answered the same two questions with one word: Nova Micro, Nova Lite, Llama 3.1 8B, Claude Haiku 4.5 and Claude Sonnet 4.6, all with one shared prompt. Bedrock Intelligent Prompt Routing used the default Nova router, which chooses Nova Lite or Nova Pro inside the answer call at no extra fee. An oracle that picks the cheapest model that is fully right gives the upper bound; it needs the answers in advance, so it serves only as a reference.

Results on the Nova ladder

Alt text: Scatter of quality against USD per 1,000 requests on a log cost axis for the Nova Micro, Lite and Pro ladder. The oracle sits top left at 96.3 and 0.056. Always Llama 4 Scout sits at 92.8 and 0.076, close to always Nova Pro at 93.1 and 0.332. Bedrock Intelligent Prompt Routing is at 87.9 and 0.129. The Nova Micro tier classifier sits low and cheap at 83.2 and 0.056, the Nova Micro multi-step classifier at 88.5 and 0.281, APO model selection at 83.6 and 0.124 and with optimized prompts at 88.3 and 0.278. The Claude Haiku 4.5 tier classifier sits right of always Nova Pro at 86.6 and 0.425. Two decider curves, one per question, trace every threshold at the far right, with the decider tier answer at 87.5 and 1.068 and the tuned multi-step setting at 87.4 and 1.396.
Quality against cost on the Nova ladder, router cost included. The decider curves trace its threshold settings.
StrategyQualityUSD per 1,000Against always Nova Pro
Always Llama 4 Scout92.80.07677% cheaper
Bedrock Intelligent Prompt Routing87.90.12961% cheaper
Nova Micro classifier, tier question83.20.05683% cheaper
Nova Micro classifier, multi-step question88.50.28115% cheaper
Claude Haiku 4.5 classifier, tier question86.60.42528% more expensive
Claude Sonnet 4.6 classifier, tier question88.81.038213% more expensive
Decider, multi-step question (threshold tuned on train)87.41.396321% more expensive
Four results stand out. Always Llama 4 Scout reached 92.8 for 0.076 USD per 1,000, higher quality than every router I tried and cheaper than all but the two small tier classifiers. Bedrock Intelligent Prompt Routing gave a solid trade with zero setup: 87.9 quality at 61% below always Nova Pro. The Claude classifiers cost more than sending everything to Nova Pro, because asking a Claude model to route costs more than the Nova answers it chooses between. The decider's routing quality, 87.4 to 87.5, sat in the same band as the better classifiers, and its compute made it the most expensive strategy on the chart. The Claude latencies are medians of 20 paced requests per model, because the account's quota of 10 requests per minute per model inflated latencies in the full run.

Judging difficulty: where the decider was perfect

On the multi-step question the decider outscored every classifier. Its mean P(yes) by family was 0.15 for extract, 0.15 for sports, 0.28 for JSON, 0.71 for math, 0.78 for code and 0.67 for logic. At a threshold of 0.50 it flagged none of the extract, sports or JSON requests and 94% of math, 100% of code and 89% of logic.
By construction, math, code and logic are the families that need multi-step reasoning. Scored against that label, the decider's ROC AUC was 1.00. The Claude Sonnet 4.6 classifier reached 0.88 and flagged only 42% of code. Claude Haiku 4.5 reached 0.78 and flagged 94% of extract as hard. Nova Micro reached 0.60: it flagged 100% of JSON and 72% of extract as hard, and 29 of its answers were neither yes nor no.
The classifiers also drifted. On about 40 of the 216 requests each, Claude Haiku 4.5 and Claude Sonnet 4.6 ignored the routing instruction and started answering the task itself. Haiku replied "249.99" to an order note; Sonnet began "Let me trace through each...". On the multi-step question that left Haiku with 39 unparseable answers and Sonnet with 40. A decider cannot do this: it answers only its typed question, with a probability over the allowed options. In fairness to Claude, I used one classifier prompt for every model, and a system prompt or a prefill would reduce the drift.
For a guardrail or a triage gate, this is the property I value most. A calibrated probability over fixed options can be thresholded, logged and audited, and it never turns into an essay.

Why perfect judgment did not produce the best routing

Alt text: Horizontal stacked bars showing, for each strategy, the share of requests sent to Nova Micro, Nova Lite and Nova Pro. The oracle sends most requests to Nova Micro. The decider on the tier question sends most requests to Nova Lite. The decider on the multi-step question sends a little over half to Nova Micro and the rest to Nova Pro. The Nova Micro multi-step classifier sends most requests to Nova Pro. Bedrock Intelligent Prompt Routing chooses only between Nova Lite and Nova Pro, mostly Lite. APO model selection splits requests across all three.
Where each strategy sends requests. The oracle keeps most requests on Nova Micro; every router escalates far more often.
Difficulty and the need for a big model turned out to be different properties. Nova Micro solves 94% of the GSM8K maths that the decider correctly calls hard, so escalating those requests buys little. Sports fails the other way. Judging whether a sports sentence is plausible takes world knowledge and few reasoning steps, so the decider correctly calls it easy, and Nova Micro then gets only 61% of it. The classifiers over-escalate, and that error happens to rescue some of those sports requests.
The tier question exposed a second effect. I wrote the tier descriptions myself, and they said JSON extraction needs the mid-size model. The decider followed them faithfully (P = 0.84 for lite on JSON) while Nova Micro got 99% of JSON right. A router that works from descriptions inherits its author's assumptions, and a more obedient router inherits them more completely.
The third effect is arithmetic. At about 0.0012 USD per decision, the decider costs more than the average Nova Pro answer (0.00033 USD), so on this ladder no routing policy can recover the cost of asking it.

A frontier ladder: Nova Micro or Claude Sonnet 4.6

The decision cost weighs less when the top of the ladder costs more. For the second ladder I routed between Nova Micro for easy requests and Claude Sonnet 4.6 for hard ones, at real Sonnet prices. Sonnet 4.6 stands in for Claude Sonnet 5.5, which is not yet available to my test account.
StrategyQualityUSD per 1,000Against always SonnetOracle match
Always Sonnet 4.695.31.543baseline21%
Oracle95.30.36576% cheaper100%
Nova Micro classifier, multi-step question91.61.21521% cheaper26%
Nova Lite classifier90.71.4109% cheaper42%
Claude Haiku 4.5 classifier93.31.4436% cheaper44%
Claude Sonnet 4.6 classifier90.51.69910% more expensive59%
Decider, P >= 0.5090.52.20743% more expensive57%
Decider, threshold 0.16 tuned on train93.52.65672% more expensive36%
Alt text: Scatter of quality against USD per 1,000 requests for the Nova Micro or Claude Sonnet 4.6 ladder. The oracle sits at 95.3 and 0.365, always Sonnet at 95.3 and 1.543. The Nova Micro, Nova Lite and Claude Haiku 4.5 classifiers sit between 1.2 and 1.45 USD with quality from 90.7 to 93.3. The Claude Sonnet 4.6 classifier and both decider settings sit to the right of always Sonnet, and the decider curve traces every threshold.
The frontier ladder. Every measured router lands well short of the oracle's saving.
The decider at P >= 0.50 matched the oracle's choice on 57% of requests, close to the Sonnet classifier's 59% and ahead of the cheaper classifiers, and it sent 54% of requests to Nova Micro. Its compute still outweighed the saving. The average Sonnet 4.6 answer on these short tasks cost 0.0015 USD, and the decider at P >= 0.50 breaks even only once an average expensive answer costs about 0.0022 USD, roughly 1.4 times these Sonnet answers. Requests with long prompts or long answers could cross that line; these short tasks did not.
The tuned threshold of 0.16 shows what calibration buys. Lowering the bar for "hard" raised quality to 93.5, the best of any router on this ladder, at a higher cost. With a calibrated probability I can move one number to trade cost for quality, and the probability means the same thing across requests. A one-word classifier answer offers no such dial.

Combinations I tested, and what is projection

The obvious next step is to combine the two kinds of router: let the decider make the difficulty call and a cheap classifier pick the tier, or let the cheap classifier go first and call the decider only to confirm an escalation. I measured these per request, which is k = 1, one decider judgment per step. The columns for k = 5 and k = 10 are a PROJECTION that I did not measure: they assume one decider judgment covers k steps of a multi-step process.
LadderCombinationQualityUSD per 1,000, measured (k = 1)PROJECTION k = 5PROJECTION k = 10
NovaDecider gate (P < 0.5 to Micro), hard to Nova Micro tier classifier83.01.220.280.16
NovaDecider gate, hard to Micro classifier choosing Lite or Pro85.81.230.280.17
NovaNova Micro multi-step classifier first, decider confirms escalations86.41.170.400.31
NovaClaude Haiku 4.5 classifier first, decider confirms escalations87.41.220.580.49
NovaDecider alone, multi-step, P >= 0.587.41.400.460.34
SonnetDecider gate, hard to Nova Micro tier classifier86.71.620.680.56
SonnetNova Micro classifier first, decider confirms escalations89.51.821.050.96
SonnetDecider alone, P >= 0.590.52.211.261.15
For comparison, the single routers on the Sonnet ladder: the Nova Micro tier classifier reached 86.9 at 0.74, the Nova Micro multi-step classifier 91.6 at 1.22, and the Claude Haiku 4.5 multi-step classifier 93.3 at 1.44. In the "classifier first" designs the decider was still called on 82% of requests, because the cheap classifier escalates so often.
The measured verdict: no combination beat the best single routers on quality per dollar in this single-step benchmark. The projection reads differently. If one decider judgment is shared across about 5 to 10 steps, decider routing on the Sonnet ladder would reach 90.5 quality at 1.15 to 1.26 USD per 1,000, 18 to 25% cheaper than always Sonnet, and still below Haiku's 93.3. At k = 20 the projection falls to 1.09. I did not run a real multi-step workflow benchmark, so these figures show the direction of the economics and nothing more.

Bedrock Advanced Prompt Optimization as a model selector

Bedrock Advanced Prompt Optimization (APO), launched on 14 May 2026, rewrites a prompt template for up to 5 target models and reports a score on your metric, a cost and a latency per model. That makes it a per-task model selector as well as a prompt rewriter: routing decided once per task type.
I ran 6 templates, one per family, with 18 training samples each. The evaluator was an AWS Lambda function running the same exact graders for 5 families. The code family needed a Claude Sonnet 4.6 LLM judge instead, because APO evaluator Lambdas may not execute code: the service rejects os, subprocess, sys and tempfile imports and exec or compile calls. I ran one job per target model across Nova Micro, Nova Lite, Nova Pro, Claude Haiku 4.5 and Claude Sonnet 4.6. Many entries failed with throttling under the 10 requests per minute Claude quota or with an intermittent service error, and 10 (family, model) pairs succeeded, all on Nova targets.
Family and targetAPO train score, original to optimizedMy test score, original to optimizedInput tokens, original to optimized
Math, Nova Pro0.78 to 1.0094.1 to 10084 to 381
Code, Nova Micro0.26 to 0.63 (judge scale)74.1 to 72.2100 to 236
Code, Nova Pro0.23 to 0.74 (judge scale)87.0 to 87.0100 to 334
Extract, Nova Micro1.00 to 1.0094.4 to 10070 to 93
Extract, Nova Lite1.00 to 1.0088.9 to 72.270 to 85
Extract, Nova Pro0.94 to 1.00100 to 10070 to 175
Sports, Nova Micro0.78 to 0.9461.1 to 72.223 to 282
Sports, Nova Lite0.78 to 1.0066.7 to 77.823 to 262
Logic, Nova Micro0.72 to 0.9461.1 to 66.7173 to 591
JSON, Nova Micro1.00 to 1.0098.9 to 98.981 to 81
Alt text: Paired dot chart of held-out test score before and after APO optimization for ten family and model pairs. Math on Nova Pro rises from 94.1 to 100, extract on Nova Micro from 94.4 to 100, sports on Nova Micro from 61.1 to 72.2 and on Nova Lite from 66.7 to 77.8, logic on Nova Micro from 61.1 to 66.7. Code on Nova Pro, extract on Nova Pro and JSON on Nova Micro are unchanged. Code on Nova Micro falls from 74.1 to 72.2 and extract on Nova Lite falls from 88.9 to 72.2.
APO on held-out requests, per pair. Most rewrites help, two hurt, and the train score did not predict which.
The held-out column is the one to trust. Extract on Nova Lite scored 1.00 on train before and after, yet the optimized prompt fell from 88.9 to 72.2 on test: an overfit. The code family showed the same risk in the prompt text itself. The rewrite for Nova Pro memorised details of training examples, including a regex for words starting with a capital P, and its large train gain on the judge scale turned into no change on test. The rewrites also grow prompts (sports went from 23 to 282 input tokens on Nova Micro), which feeds straight into cost.
For model selection I took, per family, the cheapest model within 0.03 of APO's best reported score. With the optimized prompts that selection reached 88.3 quality at 0.278 USD per 1,000, 16% cheaper than always Nova Pro. The same models with the original prompts reached 83.6 at 0.124. The model choice per task type is useful; keep the rewritten prompts only after a held-out check.

Rules I now follow for routing and judging

Alt text: Decision diagram, read top to bottom. A dashed note at the top reads "First: price single-model baselines and check quality per task family". The start box reads "A routing or judgment call to make". The first diamond asks "Is one judgment reused across many steps, a whole task type, or a guardrail?". Its yes arrow leads to a box "Strands Decider: one typed question, calibrated probability, threshold tuned on held-out data; cost per step falls as more steps share the judgment (projection)". Its no arrow leads to a second diamond "Do your models fit a Bedrock Intelligent Prompt Routing router?". Its yes arrow leads to a box "Bedrock Intelligent Prompt Routing: routes inside the answer call, no fee". Its no arrow leads to a box "Small LLM classifier (Nova Micro or Nova Lite): one-word answer, strict parsing, default tier on failure". A dotted line joins the decider box and the classifier box, labelled "Combined: decider gates the process, classifier picks the tier per request (measured limits)". A side box next to the task-type path reads "Fixed task types: choose the model once per type (APO model selection or one decider judgment), verify on held-out data".
How I now choose between a decision model, an LLM classifier and Bedrock's built-in routing.
Rule 1: price the single-model baselines before building any router. On my Nova ladder, always Llama 4 Scout reached 92.8 for 0.076 USD per 1,000, beating every router on quality. If one mid-size model already handles your traffic, a router adds cost and moving parts for little gain.
Rule 2: for per-request routing, start with Bedrock Intelligent Prompt Routing when your models fit one of its routers. It routes inside the answer call, adds no fee and no latency of its own, and it reached 87.9 quality at 61% below always Nova Pro.
Rule 3: when you need your own ladder, route with a small LLM classifier that answers in one word. The Nova Lite tier classifier cost 0.042 USD per 1,000 with a 0.59 s median, and the Nova Micro multi-step classifier reached 88.5. Parse strictly and fall back to a default tier on anything that is not an allowed answer.
Rule 4: do not route a cheap ladder per request with a Claude classifier. The Claude Haiku 4.5 and Claude Sonnet 4.6 classifiers cost 28% and 213% more than always Nova Pro, and both drifted into answering on about 40 of 216 requests.
Rule 5: use a decision model when the judgment is a property of the request that you will act on more than once: once per multi-step process or workflow, once per task type, or as a guardrail or triage gate. Here it judged difficulty with a ROC AUC of 1.00 without drifting, and its threshold could be tuned on training data (0.16 against 0.50 on the Sonnet ladder). Its cost per step falls with the number of steps that share one judgment; that saving is a projection until it is measured on a real workflow.
Rule 6: on CPU, per-request decider routing pays only when the expensive answer is expensive. On the Sonnet ladder the break-even was an average expensive answer of about 0.0022 USD, roughly 1.4 times the short Sonnet answers in this benchmark.
Rule 7: route on measured capability per task family, and check every description you give a router against that measurement. "Hard" and "needs the big model" diverged on math and sports, and my own wrong description of JSON sent requests up the ladder for nothing.
Rule 8: for fixed task types, choose the model once per type and verify it on held-out data. APO model selection with optimized prompts reached 88.3 at 16% below always Nova Pro, and one of its rewrites fell from 88.9 to 72.2 on test while scoring perfectly on train.
The combined pattern, a decider gate per process with a small classifier picking the tier per request, is the design I would build for a multi-step agent. It comes with measured limits: in single-step routing it did not beat the best single router.

Try the routers side by side

.
The live playground runs all four routers on one request.
The playground  runs the decider (with a threshold slider), the Nova Micro and Claude Haiku 4.5 classifiers and the Bedrock Nova router side by side on any request you type. The model the decider picks then answers, and the page shows the answer cost plus the routing cost against what Nova Pro would charge for the same tokens. It runs on Amazon CloudFront, an Amazon API Gateway HTTP API, AWS Lambda and Amazon DynamoDB, with rate limits of 40 calls per IP per hour and 1,500 per day, and no sign-in. An Amazon EventBridge rule keeps one decider session of about 4.2 GB warm, which costs about 51 USD per month.

Limits of this benchmark

The test half is small, 107 requests across 6 families, so differences of a point or two between routers are within noise. Every request is a single step; the multi-step economics are a projection and I have not benchmarked a real workflow. Claude Sonnet 4.6 stands in for Claude Sonnet 5.5. The 10 requests per minute Claude quota forced paced latency measurements and cut short many APO jobs, so the APO results cover 10 pairs, all on Nova targets. One classifier prompt served every model, which handicaps the Claude classifiers on drift. The decider's tier answers depend on descriptions I wrote, and one of them was wrong. All decider numbers come from a 2 vCPU CPU microVM; I have not measured the decider on other hardware, and Marc Brooker's 65 ms shows how much the hardware can change the picture.

Code and earlier parts

The benchmark, the routers, the APO driver and evaluator, and the playground CDK app are in strands-decider-agentcore/usecases/model-routing . Part 1 of this series, Serving Strands Decider on Amazon Bedrock AgentCore Runtime , covers the deployment every decider number here ran on. My thanks to Marc Brooker, whose encoder experiment changed how I read my own latency numbers.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article