
Snap out of Provisioned Concurrency for a really fast Lambda
SnapStarting containers are probably the fast containers around
Cold starts are one of those Lambda topics that never quite goes away. I've even written about this before , a lot . I'm still not sure the actual user impact is that large, and they're unlikely to be the slowest part of your stack. For most workloads they're a tiny fraction of invocations, but they can hurt, especially if you're running something like NestJS on Node, where booting the framework alone can take over a second, and they're compounded with a Lambda Authorizer, or auto instrumented with OTEL tracers. The go-to answer for years has been Provisioned Concurrency: pay to keep some execution environments initialised and waiting. It works, but you're paying for idle compute every second of every day, whether anyone calls your function or not.
SnapStart has been the other answer, but only if you were writing Java, Python or .NET. Node was never on the list, and it still isn't for the managed
nodejs runtimes. What has changed is that SnapStart now works with container images, as long as the runtime client in the image implements the SnapStart restore API. The Lambda Web Adapter does exactly that. So you can take a normal Node HTTP app, put it in a container with LWA, and turn SnapStart on.I wanted to see how well that works, and whether it's a reasonable replacement for Provisioned Concurrency. I think for a lot of APIs it is.
How SnapStart works
When you publish a version of a function with SnapStart enabled, Lambda runs your init phase once, then takes a Firecracker microVM snapshot of the memory and disk of that initialised environment. It encrypts and caches that snapshot. From then on, every new execution environment for that version is restored from the snapshot instead of booting Node, loading your modules and starting your framework from scratch.
That means anything you do at init is frozen into the snapshot. That's the whole point for expensive setup, but it's a footgun for anything that should be unique (IDs, random seeds), anything time-sensitive (credentials, cached timestamps) and network connections. You get two hooks to deal with that: one just before the snapshot is taken and one just after a restore.
Running Node on SnapStart with the Lambda Web Adapter
The container is the AWS
lambda/nodejs:24 base image with LWA copied into the extensions directory. The only gotcha is the entrypoint. The base image starts its own Node runtime client, which would fight with LWA, so you override it and just start your web server:1
2
3
4
5
6
7
FROM public.ecr.aws/lambda/nodejs:24
COPY --from=public.ecr.aws/awsguru/aws-lambda-adapter:1.1.0 /lambda-adapter /opt/extensions/lambda-adapter
ENV NODE_ENV=production PORT=8080
WORKDIR ${LAMBDA_TASK_ROOT}
COPY --from=build /app/node_modules ./node_modules
COPY --from=build /app/dist ./dist
ENTRYPOINT ["node", "dist/main.js"]LWA waits for your app to answer a readiness check, then tells Lambda init is done. Because of that, everything before
app.listen() is part of the init phase and ends up in the snapshot.The SnapStart hooks are just HTTP endpoints on your app. You tell LWA where they are with environment variables:
1
2
3
4
5
6
7
8
9
10
11
environment {
variables = {
AWS_LWA_READINESS_CHECK_PATH = "/"
AWS_LWA_SNAPSTART_BEFORE_CHECKPOINT_PATH = "/snapstart/before"
AWS_LWA_SNAPSTART_AFTER_RESTORE_PATH = "/snapstart/after"
}
}
snap_start {
apply_on = "PublishedVersions"
}and implement them like any other route. LWA rejects requests to these paths from outside, so they can't be called by a user:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
('snapstart')
export class LifecycleController {
constructor(
private readonly repository: DetailsRepository,
private readonly runtime: RuntimeState,
) {}
('before')
(200)
beforeCheckpoint() {
this.repository.close();
return { ok: true };
}
('after')
(200)
afterRestore() {
this.repository.reconnect();
this.runtime.markRestored();
return { ok: true };
}
}That's it. No Lambda handler, no Lambda-specific code in the app itself, and locally it's just a Nest app you can run with
node.The experiment
I deployed the same NestJS 12 image three times: once with SnapStart, once with nothing (a regular on-demand function) and once with Provisioned Concurrency. All arm64, all in
eu-west-1. I then used the Lambda Power Tuner with onlyColdStarts turned on, which publishes a fresh version for every invocation so every single request is a cold start. That's 10 cold starts per memory size. The tuner handles SnapStart properly: it waits for each version's snapshot to be ready before invoking, and it counts restore time in the duration, so it's a fair comparison against a normal init.Cold starts on a slim Node image
To start with I used
node:24-slim as the base image.
SnapStart took the cold start from 1.2-1.4s down to 0.56-0.76s. Roughly twice as fast, at every memory size. (Interactive comparison )
Cold starts on the AWS Node base image
Then I switched to the AWS provided
lambda/nodejs:24 base image, which is what I'd use for anything real.
The AWS image was 30-100ms slower than the slim one for both variants, probably because it's a bigger image, but SnapStart was still around 2x faster. (Interactive comparison )
Cold starts with expensive init
A Nest app that does nothing at startup isn't very realistic. Most real apps load config, warm caches, build clients or parse something large. So I added about a second of pure CPU work at init (a
pbkdf2Sync loop that runs at module load) and ran it again.
Here, SnapStart barely moved: 0.4-0.55s at every memory size, because that second of work is already done and sitting in the snapshot. The regular function went up to 2.0-2.3s. The gap grows with however much work you do at init. (Interactive comparison )
I expected the regular function to get faster as I gave it more memory, since memory is how you buy CPU on Lambda. It didn't. 256MB was as quick as 3008MB. Lambda gives the init phase a burst of CPU for up to 10 seconds regardless of how much memory you configure, so CPU-bound init runs at close to full speed anyway. Throwing memory at a slow cold start doesn't really help. Size your memory for the work your handler does, not your init.
Provisioned Concurrency
The tuner can't measure Provisioned Concurrency (the versions it publishes don't carry the config), so I invoked that one directly with a provisioned concurrency of 1. Ten requests one after another took 2-8ms each, with no init at all. That's the thing Provisioned Concurrency is genuinely great at.
Then I sent 10 requests at the same time. One got the provisioned environment. The other 9 spilled over to on-demand and each paid a 1.75-3.0s cold start. Provisioned Concurrency only protects you up to the number you've provisioned. Past that you're back to regular cold starts, and because a function can't have both, you can't fall back to SnapStart either.
What does it cost?
These are
eu-west-1, arm64 prices, ignoring the free tier. SnapStart pricing for container images is the same as for Python and .NET (it's only free for Java).| Regular | SnapStart | Provisioned Concurrency | |
|---|---|---|---|
| Duration | $0.0000133334 per GB-s | $0.0000133334 per GB-s | $0.0000086726 per GB-s |
| Always-on charge | None | $0.0000015046 per GB-s cached, per published version, minimum 3 hours | $0.0000037168 per GB-s, per provisioned environment |
| Per cold start | Init duration at the normal rate (billed since August 2025) | $0.0001397998 per GB restored, plus restore duration | None, up to your provisioned amount |
For a 1GB function over 30 days, keeping one SnapStart snapshot cached costs $3.90. Keeping one provisioned environment costs $9.63. The key difference is what they scale with. The SnapStart charge is per version, and covers however many environments get restored from it. The Provisioned Concurrency charge is per environment, so it goes up with the concurrency you want to protect.
To put that in context, here's a made-up API: 1GB, 5 million requests a month at 100ms each, 20,000 cold starts a month, each with a 2 second init.
| Monthly cost | Notes | |
|---|---|---|
| Regular | $8.20 | $0.53 of that is init |
| SnapStart | $14.48 | $3.90 cache, $2.80 restores |
| Provisioned Concurrency of 5 | $53.51 | $48.17 of that is paying for idle, and any burst over 5 still cold starts |
A couple of things fall out of this.
SnapStart is not cheaper than doing nothing. It costs to store and to restore. It's a small amount of money for a big latency win, though.
Provisioned Concurrency does get cheaper per request the busier it is, because its duration price is lower. In
eu-west-1 it breaks even with on-demand at roughly 80% utilisation of the provisioned environments. If you have steady, predictable traffic that keeps it that busy, it can make financial sense. Most APIs I've worked on have spiky traffic and quiet nights, and there it's mostly a bill for idle. Even if you were hitting 80% utilisation in a given period, Lambda may not be the most cost efficient option (but I think it's certainly still the easiest).The SnapStart cost to watch is versions. Every published version keeps its snapshot cached (and billed) while it's active. 20 functions across 3 environments with a few old versions lying around each adds up quickly. Delete old versions as part of your deployment. The example repo has a small step that does this after every deploy.
What stays the same
Pretty much everything you like about Lambda. You're still using the same programming model, event sources, IAM, aliases, CloudWatch logs and metrics. Scaling behaviour is the same, except every new environment now starts from a snapshot. You still pay per request and per millisecond. If you already deploy Lambda from a container image, the only real change is the SnapStart config and the hooks.
The trade-offs
The biggest one: you own the container now. With a managed runtime, AWS patches Node and the OS for you, and for SnapStart on managed runtimes it even patches the cached snapshots. With a container image, AWS can't patch what's inside your image. Even if you use the AWS base image, you only get the patched version when you rebuild and redeploy. Treat it like any other container you run:
- Rebuild and redeploy on a schedule, not just when your code changes
- Keep base image and dependency updates automated with something like Renovate or Dependabot
- Turn on ECR image scanning so you know when you're behind
You're also tied to versions and aliases. SnapStart doesn't apply to
$LATEST, so everything needs to call an alias or a version. Each publish takes around 45 seconds for the snapshot to be ready, which slows your deployments down a bit.You need to think about what's in the snapshot. Anything unique, random, time-sensitive or connection-based has to be recreated after restore.
There are also some hard limits. SnapStart doesn't work with EFS, S3 Files, ephemeral storage above 512MB, or alongside Provisioned Concurrency on the same function.
Finally, it's fast, not instant. 400-600ms is a lot better than 2 seconds, but Provisioned Concurrency is single-digit milliseconds. If you have a strict latency requirement on every single request, Provisioned Concurrency is still the right tool. AWS also notes that functions invoked very infrequently may not see the same improvement.
So should you switch?
If you're running a Node API on Lambda and reaching for Provisioned Concurrency to hide slow cold starts, I think SnapStart on a container image is worth trying first. It covers every environment, not just the first one. You don't need to predict your concurrency or schedule scaling. It costs a few dollars per version rather than a few dollars per environment, and it gets relatively better the more you do at init. The price is owning a container and keeping it patched, which if you've run anything on ECS or EKS is a well-trodden path.
All the code is on GitHub: ryancormack/snapstart-vs-provisioned-vs-regular . It has the Nest app, the Dockerfile, OpenTofu for all three variants and the raw Power Tuner results.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article