Strands Decider on Lambda...
Strands Decider in an open weights decision making model. I wanted to see if it would be worth hosting it on Lambda for cheap decisions done serverlessly.
Can a tiny decision model run on Lambda?
I put Strands Decider on a CPU-only Lambda function. It works, it's cheap, and it's slow.
I spend a lot of time trying to make Lambda faster, so it seemed fair to see what happens when I ask it to do something it wasn't designed for. Strands Decider is a small open model that normally wants a GPU or a decent laptop. I wanted to know whether it could run on Lambda running a Graviton2 CPU and no GPU.
What is a decision model?
An LLM generates text. A decision model does something narrower: you give it some state, a question and a fixed set of options, and it picks one or scores it on a scale. The answer comes with a calibrated confidence.
"Is this message urgent?" "Which team should handle this: billing, sales or retail?"
The trade is flexibility. A decision model can't write code or hold a conversation, and it's weaker than a reasoning model on hard, multi-step problems. In return it always answers with one of your options, gives you a confidence you don't get from a frontier model API, and makes it cheap to ask several questions about the same input. That suits the small decisions inside an agent, like routing, tool call checks, triage and guardrails.
Strands Decider
Strands Decider 2B takes a pre-trained Qwen3.5-2B, removes the language model head so it can't generate text, and replaces it with a small pointer head that scores each option. It's fine-tuned with a LoRA adapter and released under Apache-2.0 with the weights, data and training scripts. And importantly for this experiment, it can run on a CPU.
It speaks the same request and response shape as Jev, the decision model from TypeSafe AI. The launch post reports a median of around 115ms per decision on an RTX 3090 and around 153ms on an M3 MacBook for small tasks.
Hosting it today
The model is on Hugging Face. The model card shows
pip install strands-decider for the CLI and Python, and strands-decider serve for an HTTP endpoint. The server binds to localhost with no authentication, so it's for experiments only.For a managed option, Hugging Face has Inference Endpoints, which run a model on a dedicated instance you pay for by the hour. I haven't deployed to one. The model is a LoRA adapter plus a custom head, so I'd expect to need a custom container.
Why self host, and why Lambda?
Self hosting keeps your data in your own account, puts access behind IAM, and gives you no shared rate limits. Lambda adds scale to zero and pay per request, with no servers. A Function URL with IAM auth gives you an authenticated endpoint without API Gateway, and container images can carry gigabytes of weights.
The catches are no GPU, and a memory ceiling of 10,240 MB, which comes with about 6 vCPUs.
What I built
The repo is ryancormack/strands-decider-lambda : a CDK stack in TypeScript and a container image Lambda in Python behind an IAM-authenticated Function URL. You POST a request:
1
2
3
4
5
6
7
8
9
10
11
12
{
"state": "Help! My payouts have been failing for 3 days.",
"model": "strands-decider-2B-hobson-v19",
"questions": {
"is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"},
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": null, "technical": null, "sales": null}
}
}
}and get a response:
1
2
3
4
5
6
7
8
9
10
11
12
13
{
"model": "strands-decider-2B-hobson-v19",
"answers": {
"is_urgent": {"type": "noul", "noul": 0.8091},
"team": {
"type": "choice",
"choice": "technical",
"probabilities": {"billing": 0.2023, "technical": 0.7833, "sales": 0.0144},
"confidence": 0.6749
}
},
"usage": {"input_tokens": 133, "output_tokens": 2}
}The Lambda configurations
Two facts about the model shape everything else. It runs in fp32 on the CPU, because the Decider upcasts the bf16 weights (bf16 kernels on CPU are slower), and that needs about 7 GiB of memory on its own. It's also big: the container image is about 5.2 GB, of which roughly 3.8 GB is weights.
So there was little to tune on memory. The fp32 model needs more than 8 GiB, and 10,240 MB is both Lambda's maximum and the setting that gets the most CPU. The things I could compare were thread count and request size.
Threads
I ran the same 3-question request on the live function with 1, 3 and 6 threads, three warm runs each:
| Threads | Server time | Speed-up | Cost per 1,000 requests |
|---|---|---|---|
| 1 | 27.5 s | 1.0x | about $3.67 |
| 3 | 10.7 s | 2.6x | about $1.43 |
| 6 | 4.4 s | 6.3x | about $0.59 |
It scales close to linearly, so inference uses all the cores. Because Lambda bills GB-seconds at a fixed memory size, more threads also means a lower cost per request. Six threads is the ceiling here, and going faster means a platform with more cores.
Request size
| Request | Server time |
|---|---|
| 1 question, short state (~130 tokens) | 1.7 s |
| 3 questions, short state | 3.5 s |
| 3 questions, ~370 token state | 5.5 s |
A single question is roughly ten times slower than the laptop figure from the launch post.
Cold starts
A cold start takes 18 to 30 seconds once the image is cached in the region, and over three minutes for the first request after a new image is pushed, because Lambda fetches large images lazily. Reading the weights in parallel before the model loads fixed the worst of it. I didn't try SnapStart, which I'd expect to help cold starts but not warm requests .
What it costs, and how that compares to Hugging Face
Lambda arm64 in eu-west-1 costs about $0.000133 per second at 10 GB, with nothing to pay while idle. At the 3-question rate that works out at $0.47 per 1,000 requests.
Inference Endpoints bill by the hour while the endpoint is running. These are the AWS rates from Hugging Face's pricing page, with an always-on month at 730 hours:
| Option | Spec | Hourly | Always on, per month |
|---|---|---|---|
| HF Endpoint, Intel SPR x8 | 8 vCPU, 16 GB | $0.268 | about $196 |
| HF Endpoint, Intel SPR x16 | 16 vCPU, 32 GB | $0.536 | about $391 |
| HF Endpoint, NVIDIA T4 | 1 GPU, 14 GB | $0.50 | $365 |
| HF Endpoint, NVIDIA L4 | 1 GPU, 24 GB | $0.80 | about $584 |
| Lambda, 10 GB arm64 | 6 vCPU | per request | $0 idle, $4.70 at 10k requests, $47 at 100k, $467 at 1M |
Against the 8 vCPU endpoint, Lambda is cheaper up to roughly 415,000 requests a month, and against the T4 up to roughly 775,000. Both figures assume the always-on endpoint and the 3-question rate.
An endpoint on a GPU should be far faster than my function. If you only need the model for a batch an hour a day, an endpoint that's only running then costs far less than always-on. I haven't tested how Hugging Face's scale to zero behaves for this model.
Should you?
Probably not for anything a user is waiting on. A warm single-question request takes nearly two seconds, most realistic requests take three to six, and a cold start can take minutes. If you want a decision model in a real-time path, Lambda probably isn't for you.
It could make sense when latency doesn't matter: batch triage, overnight classification, evaluating a pile of agent outputs. There, paying nothing while idle and a fraction of a cent per decision is attractive, especially if it means a bigger model handles fewer of the easy calls.
I'd love to see GPUs on Lambda for exactly this sort of thing, or even this model hosted on Bedrock .
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article