
What Happens to Your Application When Amazon DocumentDB Serverless Scales
Part 3 of 5 diving deep in Amazon DocumentDB Serverless through empirical testings
Series: Amazon DocumentDB Serverless: An Empirical Study (5 articles)
- …
- 3What Happens to Your Application When Amazon DocumentDB Serverless Scales This article
Scaling is transparent from a correctness perspective - queries succeed, connections hold. But scaling is not free at the latency layer. When Amazon DocumentDB Serverless scales in, the buffer cache is evicted along with the compute capacity. When your workload comes back after an idle window, the first reads hit storage instead of RAM, and individual page-fetches can take 100-180 ms. When a long analytical query is already running while the cluster scales out, the extra capacity doesn't accelerate the query that's already in flight. This article measures both effects with precision, and gives you the numbers to decide whether they matter for your workload.
Who this is for
You run an application on Amazon DocumentDB Serverless and are wondering: does my P99 read latency spike after quiet periods? Will my nightly analytics batch complete faster if I raise MinCapacity ahead of time? Is the cluster's buffer cache preserved during scale-in? This article addresses each question with measurement.
Executive summary
Three findings from cache-and-query experiments on DocumentDB 8.0.1 Serverless:
- Scale-in evicts the buffer cache. After a natural scale-in from ~10 DCU to floor (0.5 DCU), the same 1,000 random reads that ran at P50 0.68 ms warm ran at P50 3.14 ms cold - a 4.6× degradation at P50 and 8.7× at P99.
- The first reads on a cold cache are noticeably slow. Individual page-fetch latencies of 100-180 ms are visible for the first ~300 milliseconds of workload. After 30 seconds of sustained activity on a small hot working set, latency stabilizes.
- Scale-out does not accelerate long-running queries already in flight. A single heavy query ran in 28.2s starting at MinCap 0.5 vs. 22.0s starting at pre-warmed MinCap 8 - a 1.28× speedup despite 16× more capacity available. Scale-out opens capacity for subsequent work; it doesn't rewrite the pace of work already running.
How we measured this
Same lab as Parts 1 and 2: single-instance DocumentDB 8.0.1 Serverless in
us-east-1, PyMongo 4.9.2 workload from an EC2 client in the same VPC. The full methodology and reproducer scripts are in Part 0 of this series.Two additional pieces of instrumentation for this article:
- Client-side latency histograms for read operations, computed on 1,000-read batches at both warm and cold cluster states.
- A 5 GB collection (5M documents) for the heavy-query test, enabling an I/O-bound query that takes long enough to observe in-flight capacity changes.
Experiment 1: Buffer cache eviction on scale-in
Setup.
- Populate the cluster with 2M documents (~500 MB total data).
- Warm the cluster with 5 minutes of workload (100 conns × 30 ops/sec) so the buffer cache fully caches the dataset. Cluster scales to ~10 DCU during this window.
- Snapshot A (WARM): Run 1,000 seeded random reads. Record per-read latency.
- Stop workload and close connections. Wait 20 minutes for natural scale-in to floor (0.5 DCU).
- Snapshot B (COLD): Run the same 1,000 seeded reads. Same seed, same document IDs, same code path. Only the cluster state differs.
Result.
| Percentile | WARM (10 DCU, cache full) | COLD (0.5 DCU, cache empty) | Delta |
|---|---|---|---|
| P50 | 0.68 ms | 3.14 ms | +365% (4.6×) |
| P95 | 0.80 ms | 4.63 ms | +476% (5.8×) |
| P99 | 0.88 ms | 7.65 ms | +769% (8.7×) |
| MAX | 2.08 ms | 29.78 ms | +1332% (14.3×) |
Zero errors in either snapshot. The COLD reads all completed - just slower.
Two effects are compounded here. The buffer cache is empty (so reads fault in from storage), AND compute capacity is at floor (so the read-processing pipeline has less CPU headroom). The latency degradation is the sum of both.
Interpretation. The buffer cache lives in instance memory, and when the instance scales down, that memory shrinks and the cache is evicted. This is not surprising in principle - it's how caching in any elastic compute layer works - but the magnitude matters. If your application does bursty reads separated by 15-20 minute quiet windows, expect the first burst after each quiet window to see 5-9× worse read latency than the steady-state warm case.
A CloudWatch observation. During both snapshots, CloudWatch reported
BufferCacheHitRatio at 100%. The metric aggregates over 60-second windows and smoothed over the cache-miss burst at the start of the COLD snapshot. Client-side latency histograms were the only way to see it. If your monitoring depends on BufferCacheHitRatio alone, you may miss real cache-miss events.Experiment 2: Cold-cache re-warming rate
Setup. Continued from Experiment 1's end state: cluster at 0.5 DCU, cache empty. Run 20 concurrent PyMongo clients × 5 ops/sec against a small hot working set of 100 documents (25 KB total, easily fits in a 0.5 DCU buffer cache). Record per-op latency over 10 minutes, bucket into 30-second batches.
Result.
| Batch | Time window | P50 | Mean | Notes |
|---|---|---|---|---|
| 0 | 0-30s | 1.25 ms | 2.34 ms | Cache warming |
| 1-19 | 30s-10min | 1.12-1.21 ms | 1.20-1.30 ms | Stable |
But the first 300 ms tells a different story.
1
2
3
4
t=0.145s: 106.6 ms ← first read of first page
t=0.226s: 135.3 ms
t=0.286s: 181.4 ms ← MAX for entire 10-minute run
t=0.288s: 1.2 ms ← page loaded, subsequent read fastFor approximately 300 ms of wall clock, individual first-page-fetches produced 100-180 ms latencies - 265× the warm baseline for a brief window. These are not batch-averaged numbers; they are individual per-read latencies. Once the pages fault in, subsequent reads against the same pages complete at normal speed.
DCU response. The cluster scaled from 0.5 to 1.6 DCU over the 2-minute window in response to sustained storage-read pressure at just 100 ops/sec. This confirms IOPS pressure is an independent scale-out signal (see Part 1, Experiment 4). The compute-bound latency floor at 1.6 DCU stayed at P50 1.16 ms - did not return to the 10-DCU warm baseline of 0.68 ms.
Interpretation. The 30-second batch view says "cache warms in 30 seconds", but a user-facing request handler sees the 100-180 ms spikes during the first 300 ms. If your workload is highly latency-sensitive during startup (first request after quiet period), these spikes are visible.
Practical implication. If your application must serve consistent sub-2 ms P99 reads even on the first request after quiet, you need either:
- Pre-warm the cache before opening to traffic. Drive workload against the expected working set 1-2 seconds before real traffic arrives.
- Set MinCapacity high enough that the cache never fully evicts. This is what "reserved capacity for latency SLO" means in serverless: pay for the DCU that holds your working set resident, always.
- Accept the first-300ms spike. Fine for many workloads, unacceptable for real-time systems.
Part 5 covers the "when this matters" call more explicitly.
Experiment 3: In-flight queries during scale-out
Setup. Populate a 5 GB collection (5M documents, 1% match rate for a target
find query - the query must full-scan to find matches).Test A: Cold start. Cluster at MinCap 0.5. Launch a single find query that will scan the collection.
Test B: Pre-warmed. Same query on a cluster pre-warmed to MinCap 8 DCU.
Test B: Pre-warmed. Same query on a cluster pre-warmed to MinCap 8 DCU.
Measure end-to-end query time and per-batch throughput within the query.
Result.
| Test | Starting DCU | End DCU | Query time |
|---|---|---|---|
| A (cold start) | 0.5 | ~9 (scaled during query) | 28.2 s |
| B (pre-warmed) | 8 | 8 | 22.0 s |
Pre-warming to MinCap 8 (16× more capacity than 0.5) produced a 1.28× speedup. Not 16×, not 8×, not 3×. 1.28×.
Per-batch throughput analysis of Test A (does the query accelerate mid-flight when the cluster scales up?):
- First 10 batches: throughput X ops/sec
- Middle 10 batches: throughput X - 2%
- Last 10 batches: throughput X - 1.9%
The query completed at essentially the same rate throughout, even though DCU rose from ~2 to ~9 during execution. Scale-out opens capacity for subsequent queries; it does not rewrite the pace of work already running.
Interpretation. A query's execution plan is committed at parse time, using the resources available at that moment. Adding capacity mid-flight adds room for other queries in parallel, but the query in progress continues on its original resource allocation. This is architecturally consistent with how most database engines work; it is a fact worth knowing before you plan a "just scale up when the big report runs" strategy.
Practical implication. For workloads dominated by a few long-running analytical queries, the capacity available at query START determines the query time. If you have a nightly report that takes 30 minutes at MinCap 0.5, pre-warming to a higher MinCap 1-2 seconds before the query launches will help. Pre-warming AFTER the query starts will not.
For OLTP workloads with many small queries, scale-out helps the aggregate throughput even mid-workload - because the "aggregate throughput" is the parallel sum of many independent queries, each starting fresh at the current DCU.
Synthesis: latency behavior across scaling events
| Event | Correctness impact | Latency impact | Duration |
|---|---|---|---|
| Scale-out (rising load) | None | Transient throughput dips in some 1-second windows | ~5 seconds |
| Scale-in (falling load) | None | Buffer cache evicted with compute | Cache reload on next burst |
| First reads after quiet | None (all reads complete) | 100-180 ms individual page-fetches for first ~300 ms | Cache warms in ~30 seconds |
| Long query with scale-out during execution | None | Query continues at starting-DCU pace, does not accelerate | Full query duration |
What this means for your application
Rule 1: Read latency after quiet windows is a real user-facing concern for latency-sensitive workloads. If P99 read latency of ~4 ms warm vs. ~8 ms cold is meaningful for you, the mitigation is either pre-warm the cache or size MinCapacity so the working set stays resident. Provisioned instances don't have this problem because they don't scale down.
Rule 2: Long analytical queries need capacity provisioned at query START. If your report takes 30 minutes at MinCap 0.5, scaling up 5 minutes into it won't help. Pre-warm before the query launches, or use provisioned for analytical workloads.
Rule 3: The
BufferCacheHitRatio CloudWatch metric can smooth over transient cache misses. Trust client-side latency histograms for real-time performance investigation. CloudWatch is fine for trend monitoring; it is not fine for sub-minute anomaly detection.What we didn't measure
- The exact working-set-to-DCU ratio for cache retention. As a rough guide, each DCU provides ~2 GB of memory available for caching. Whether the entire working set must fit, or just the hottest N% for good hit ratios, is workload-dependent and not something we isolated.
- Cache retention across scale-in events on multi-instance serverless. All experiments are single-instance.
- Which specific storage tier serves the cold reads. DocumentDB storage is a shared distributed layer; whether "cold reads" hit the storage volume directly or an intermediate cache layer is not something the customer surface exposes.
Coming next
Part 4 - Sizing MinCapacity: Three Rules. Connection ceiling, working-set cache size, and steady-state throughput as inputs.
Part 5 - When to Choose Provisioned Over Amazon DocumentDB Serverless. The trade-offs across latency SLOs, sustained throughput, and cost predictability.
Previously:
- Part 1: What Triggers Amazon DocumentDB Serverless to Scale (And What Doesn't)
- Part 2: How Fast Does Amazon DocumentDB Serverless Scale?
Reproducibility. All 15 experiments, raw JSON results, and Python driver scripts are in the companion GitHub repository .
Series: Amazon DocumentDB Serverless: An Empirical Study (5 articles)
- …
- 3What Happens to Your Application When Amazon DocumentDB Serverless Scales This article
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article