
Welcome to LazyTown: How Being Lazy Made My Python Lambda 96.65% Faster
I benchmarked four Python Lambda cold-start optimizations on AWS: ARM64, lazy imports, memory tuning, and package size, each tested in isolation. The result: a 96.65% drop in Cold Total P50, without switching languages.
How I cut Python Lambda cold starts by 96.65% by actually measuring four optimizations (ARM64, lazy imports, memory tuning, and package size) instead of assuming "just use Go" was the only fix.
In LazyTown, Sportacus spends every episode chasing people away from the couch. In my benchmark, being lazy on purpose, in exactly one place, turned out to be the best decision I made all week.
1
2
3
4
Python baseline Cold Total P50 2457.86 ms
Optimized Python Cold Total P50 82.34 ms
Improvement 96.65%
Same Python. Same runtime. A very different setup.
This article is the write-up behind a talk I'm giving at Python Pizza, and it's about that setup: what I changed, what I measured, and which of my assumptions turned out to be wrong.
Why I Ran This Experiment
Every AWS Lambda cold-start conversation eventually arrives at the same sentence:
"Just use Go."
Go is genuinely a strong fit for a lot of serverless workloads: fast start, small binary, thin runtime. Nobody needs convincing of that, and this article isn't trying to un-convince anyone.
That wasn't the question I cared about, though. I wanted to know: if I optimize the architecture, the dependency loading, the memory configuration, and the deployment package, how much of Python's cold-start problem is actually the language, and how much is just how I was running it?
So I built a small benchmark rig, picked one workload, and changed one thing at a time until I had an answer.
First, What Am I Actually Measuring?
A Lambda invocation, simplified, looks like this:

During
INIT, Lambda may need to prepare the execution environment, boot the runtime, and run whatever happens at module import time. That's where "expensive top-level imports" quietly become a cold-start tax.For every variant in this benchmark, I collected:
Init P50Cold Total P50Warm P50Cold P90
Cold Total P50 is the number this whole article is built around, and it's worth being precise about it, because it isn't a metric AWS hands you. It's one I defined:1
Cold Total = Init Duration + Duration
I ran 30 cold invocations and 30 warm invocations per variant, and I only accepted a sample as "cold" when the Lambda
REPORT line actually included an Init Duration. My benchmark script forces a fresh execution environment before every cold sample by updating the function's Description with a unique nonce and waiting for the update to finish. It's a best-effort way to force a cold start, not an official AWS API for it:1
2
3
4
5
6
7
8
9
10
11
12
def force_refresh(function_name: str) -> None:
nonce = f"lazytown-bench-{uuid.uuid4().hex[:16]}"
run([
"aws", "lambda", "update-function-configuration",
"--function-name", function_name,
"--description", nonce,
"--output", "json",
])
run([
"aws", "lambda", "wait", "function-updated",
"--function-name", function_name,
])
And here's the part of the parser that decides whether a sample even counts as cold:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
def parse_report(log_text: str) -> dict:
m = REPORT_RE.search(log_text)
if not m:
raise RuntimeError(f"REPORT line not found:\n{log_text}")
duration = float(m.group("duration"))
init = float(m.group("init")) if m.group("init") else None
return {
"duration_ms": duration,
"init_ms": init,
"cold_total_ms": (duration + init) if init is not None else None,
"is_cold": init is not None,
}
One more honesty note before the results: these numbers describe this experiment, this workload, this region, and this runtime version. They're not a universal Lambda performance guarantee, and I'll keep saying "in this experiment" throughout, not as a disclaimer reflex, but because it's genuinely the only claim the data supports.
Round 1: The Comparison That Almost Fooled Me
I started small, on purpose: a trivial handler, tiny JSON in, tiny JSON out. Both Python and Go, same region, same memory.
1
2
3
4
Python Cold Total P50 65.62 ms
Go Cold Total P50 65.60 ms
Difference 0.02 ms
Basically a tie.
For this tiny workload, Lambda's own initialization overhead dominated the request so heavily that the runtime itself barely showed up in the number. If I'd stopped here, I would have walked away with a comforting and wrong conclusion: "runtime doesn't matter." Don't assume the runtime will dominate every workload. Measure it.
Important caveat: don't compare this ~65 ms Round 1 number to anything later in this article. The final race uses a much heavier workload, and the two aren't apples-to-apples.
Optimization 1: ARM64, for Free
The first Python-only change touched exactly one setting:
1
x86_64 -> arm64
Everything else (Python version, 512 MB memory, region, payload, 30 cold + 30 warm samples) stayed identical.
1
2
3
4
x86_64 Cold Total P50 85.09 ms
arm64 Cold Total P50 65.22 ms
Improvement 23.36%
A 23% drop from flipping one field. On Lambda,
arm64 runs on AWS Graviton processors, and the SAM change is genuinely this small:1
2
3
4
5
6
7
8
9
Resources:
Slide06ARM:
Type: AWS::Serverless::Function
Properties:
Runtime: python3.13
Handler: app.lambda_handler
Architectures:
- arm64
MemorySize: 512
If any of your dependencies ship native binaries, build for the target architecture (
sam build --use-container before sam deploy --guided).Lesson:
Architectures is a performance decision, not an implementation detail you set once and forget. (AWS Lambda instruction set architecture docs )Optimization 2: The Best Kind of Lazy
This is the one I'd put at the top of the talk slide, and the reason "LazyTown" is the name on the repo.
Picture a Lambda function with two request paths:
simple: doesn't touch pandasreport: genuinely needs pandas
The eager version
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
import json
import pandas as pd
def lambda_handler(event, context):
request_type = event.get("type", "simple")
if request_type == "simple":
return {
"statusCode": 200,
"body": json.dumps({"mode": "simple", "ok": True}),
}
rows = event.get("rows", [])
frame = pd.DataFrame(rows)
total = int(frame["value"].sum()) if not frame.empty else 0
return {
"statusCode": 200,
"body": json.dumps({"mode": "report", "total": total}),
}
Every single
simple request pays for import pandas as pd, even though simple never touches a DataFrame.The lazy version
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
import json
def lambda_handler(event, context):
request_type = event.get("type", "simple")
if request_type == "simple":
return {
"statusCode": 200,
"body": json.dumps({"mode": "simple", "ok": True}),
}
import pandas as pd
rows = event.get("rows", [])
frame = pd.DataFrame(rows)
total = int(frame["value"].sum()) if not frame.empty else 0
return {
"statusCode": 200,
"body": json.dumps({"mode": "report", "total": total}),
}
Same handler.
pandas just moved from the top of the file to the one branch that actually needs it.For a
simple request:1
2
3
4
Eager Cold Total P50 1796.69 ms
Lazy Cold Total P50 66.88 ms
Improvement 96.28%
That's the number that gets applause. But there's a second half to it that matters just as much, and I'd be doing you a disservice to leave it out.
For a
report request, the one that actually needs pandas:1
2
Eager report Cold Total P50 1815.13 ms
Lazy report Cold Total P50 6800.56 ms
Lazy loading didn't delete the cost. It relocated it, and in this run, paying for that
import pandas inside the first report invocation was almost four times slower than eating it eagerly at module load. Be lazy when you can skip the work entirely, not when you're only postponing it to a worse moment.Lazy imports are a genuine win when a dependency is expensive and a meaningful share of requests never touch it. They're not a free performance toggle you sprinkle over every import statement.
Optimization 3: Memory Is Also CPU
Next I froze everything except
MemorySize and stepped through four values on the same compute-heavy handler:1
2
3
4
5
6
Memory Cold Total P50 Warm P50
128 MB 2129.00 ms 2081.74 ms
512 MB 577.55 ms 508.58 ms
1024 MB 321.41 ms 256.29 ms
1769 MB 236.62 ms 167.03 ms
128 MB to 1769 MB:
1
2
2129.00 ms -> 236.62 ms
Improvement: 88.89%
Here's the part that made me stop and actually think:
Init P50 barely moved across all four. It stayed pinned around 69-70 ms. What got dramatically faster was the handler, not initialization. That tracks with how Lambda works: memory isn't just a ceiling, it's a proportional CPU allocation, and AWS documents that at 1769 MB a function gets the equivalent of a full vCPU.1
2
3
4
5
6
7
8
9
Resources:
PythonFunction:
Type: AWS::Serverless::Function
Properties:
Runtime: python3.13
Handler: app.lambda_handler
Architectures:
- arm64
MemorySize: 1769
I'm deliberately not turning this into a cost claim. The billed-duration numbers in my raw data are aggregated across cold and warm rows together, so I can't responsibly say "and it's cheaper too" without recomputing that split first. What I can say: for a compute-bound handler like this one,
MemorySize is a latency knob, not just a memory limit. (AWS Lambda memory configuration docs )Optimization 4: The Diet That Almost Wasn't
This is my favorite result in the whole experiment, precisely because it didn't confirm what I expected.
I built two deployment packages with the identical handler:
Fat:
pandas, numpy, requests, Pillow sitting in the SAM build directory, at roughly 125 MB, but never imported by the handler.Slim: no third-party dependencies at all.
1
2
3
4
5
6
7
8
9
10
FAT
Init P50 63.74 ms
Cold Total P50 65.56 ms
SLIM
Init P50 62.91 ms
Cold Total P50 64.75 ms
Difference 0.81 ms
Improvement 1.24%
A 125 MB build directory versus practically nothing, and the cold start barely blinked. (Worth being precise: that 125 MB figure is the size of the SAM build directory measured with
du -sh, not the deployed ZIP size, and I'm not going to pretend it is.) Shipping more and importing more are not the same cost.Smaller packages are still worth having, for build times, security surface, and general sanity, but "make the ZIP smaller" was not the highest-leverage lever in this test. If I'd started my optimization pass here instead of at lazy loading, I would have spent a lot of effort for 1% of the win.
The Final Race: Stacking It All Up
After running each optimization in isolation, I combined the ones that actually moved the needle into a final Python configuration:
1
2
3
4
5
Python 3.13
+ ARM64
+ 1769 MB
+ Lazy / minimized dependency loading
+ Slim package
I want to be upfront about one thing here: the final "optimized" handler doesn't literally contain a conditional
import pandas. For its actual workload (repeated SHA-256 hashing) it simply has no unnecessary eager dependencies to begin with. It applies the lesson from the lazy-loading experiment rather than reproducing that exact code.I ran this final configuration against the original baseline and against Go, same workload, same 30+30 sampling:
1
2
3
4
5
6
Variant Init P50 Cold Total P50
Python baseline 2407.15 ms 2457.86 ms
Python + ARM64 1949.68 ms 1990.65 ms
Python optimized 69.60 ms 82.34 ms
Go 49.91 ms 54.35 ms
The number I actually care about:
1
2
2457.86 ms -> 82.34 ms
96.65% lower Cold Total P50
Go still wins the final line, and I'm not going to pretend otherwise:
1
2
Go 54.35 ms
Optimized Python 82.34 ms
But that was never the bet I was making. This was never "Python beats Go." Python went from 2.46 seconds to 82 milliseconds, without changing the language.
One honest caveat before I move on: the baseline-to-optimized jump changes multiple variables at once (architecture, memory, dependency strategy). It's a combined configuration result, not proof that any single lever alone is worth 96.65%. Each lever's isolated contribution is exactly what the sections above this one measured separately.
What Actually Moved the Needle
Laid out side by side, this is the part I'd put on a slide:
1
2
3
4
5
ARM64 85.09 -> 65.22 ms (23.36% lower)
Lazy loading 1796.69 -> 66.88 ms (96.28% lower) <- simple path only
Memory tuning 2129.00 -> 236.62 ms (88.89% lower)
Package diet 65.56 -> 64.75 ms (1.24% lower)
Final Python 2457.86 -> 82.34 ms (96.65% lower)
Two of these are worth staring at together:
1
2
Lazy loading -> 96.28% improvement
Package diet -> 1.24% improvement
Both sound like reasonable serverless advice on their own. Only actually measuring told me which one mattered for this workload. That's exactly why I'm not turning this into a universal checklist ("always use ARM64," "always allocate 1769 MB," "always lazy-load everything"). That would be the wrong lesson from a right result.
The lesson I'm actually taking is a loop, not a list:
1
MEASURE -> FIND THE BOTTLENECK -> CHANGE ONE THING -> MEASURE AGAIN
What I'd Actually Do in Production
- Look at distributions, not single runs. P50, P90, P99: one lucky invocation tells you nothing.
- Keep cold and warm behavior separate. They're different execution paths with different costs.
- Audit your module-level imports. An expensive top-level import is an invisible tax on every cold start.
- Test ARM64 compatibility early, especially if native dependencies are involved.
- Tune memory against both latency and cost, but compute the cost split properly; don't eyeball a combined aggregate.
- Keep packages intentional, without assuming size alone is your bottleneck.
- Optimize for the request paths that actually exist. Lazy loading pays off exactly where a real fraction of traffic skips the expensive path, not everywhere by default.
Reproducing This
The full repository will follow a structure like this:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
lazytown-lambda-cold-starts/
├── README.md
├── benchmarks/
│ ├── benchmark.py
│ ├── summarize.py
│ └── result_card.py
├── experiments/
│ ├── 01-baseline/
│ ├── 02-python-vs-go/
│ ├── 03-arm64/
│ ├── 04-lazy-loading/
│ ├── 05-memory-tuning/
│ ├── 06-package-diet/
│ └── 07-final-race/
├── results/
│ ├── raw.csv
│ └── summary.csv
└── slides/
Command-line shape, for anyone who wants to point this at their own function:
1
2
3
4
5
6
7
8
9
10
11
12
python scripts/benchmark.py \
--function-name your-function-name \
--label optimized-python \
--payload payload.json \
--cold 30 \
--warm 30 \
--output results/raw.csv
python scripts/summarize.py \
results/raw.csv \
--csv results/summary.csv \
--md results/summary.md
The exact CLI matters less than what it enforces: the benchmark stays repeatable, the raw samples get preserved (not just the final chart), and P50/P90 get calculated from the actual data every time, not eyeballed off one lucky run.
Full repository:
lazytown-lambda-cold-starts (replace with final GitHub URL) Measure First
Go was fast because that's just what Go is. Python got fast because I made it earn it.
- An expensive top-level import is a tax every cold invocation pays, whether it needs the dependency or not.
- Architecture and memory aren't deployment trivia. In this test, they were two of the biggest levers I had.
- A big dependency you never import barely costs you anything at cold start; the same dependency imported lazily on the wrong path can cost you a lot more than eating it eagerly.
- None of this generalizes automatically to your workload. Rerun the measurement before you trust the conclusion.
Go still finished ahead: 54.35 ms to 82.34 ms. But Python recovered 96.65% of its own baseline cold-start time without a single line of it being rewritten in another language, and that's the number I actually walked away caring about.
Measure first. Optimize what actually hurts.
References
- AWS Lambda: Configure Lambda function memory
- AWS Lambda: Selecting an instruction set architecture
- AWS SAM: Build your application
- AWS SAM CLI:
sam buildcommand reference
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article