
What Triggers Amazon DocumentDB Serverless to Scale (And What Doesn't)
Part 1 of 5 diving deep in Amazon DocumentDB Serverless through empirical testings
Series: Amazon DocumentDB Serverless: An Empirical Study (5 articles)
- 1What Triggers Amazon DocumentDB Serverless to Scale (And What Doesn't) This article
Part 1 of 5 · Amazon DocumentDB Serverless: An Empirical Study
Amazon DocumentDB Serverless scales capacity in response to workload pressure. The obvious question is: which pressure? This article measures four candidate signals - client connection count, idle connection presence, keep-alive ping frequency, and I/O pressure from cache misses - and reports which of them actually move the DCU meter. If you have ever wondered whether your connection pool is masking a scale-out signal, or whether closing idle clients at end-of-shift will help the cluster shrink back to its floor, this piece answers both questions with measurement rather than guess.
Who this is for
You are a developer or architect running (or evaluating) Amazon DocumentDB Serverless. You have read the product page and know that "scaling is driven by workload", but the word "workload" hides more than it reveals. You want to know which knobs actually turn the dial before you commit to a MinCapacity value or design a connection pool.
Executive summary
Four findings from controlled experiments on DocumentDB 8.0.1 Serverless:
- Connection count is not a scale-out signal. We opened 300 client connections against a cluster at MinCapacity 1.0 (ceiling 250). The extra 50 clients were refused for 8.5 continuous minutes. DCU never moved.
- Idle connections extend scale-in, but only above a threshold. 10 to 100 idle clients delay time-to-floor by ~6 minutes vs. all-closed. 500 idle clients pin DCU at ~2 indefinitely - the cluster holds exactly the DCU required to accommodate the open sockets and stops shrinking.
- Ping frequency doesn't matter. Whether idle clients pinged every 5s, every 30s, or not at all (TCP-only), the scale-in trajectory was the same. The autoscaler observes TCP presence, not application-layer activity.
- I/O pressure from cache misses DOES trigger scale-out. Sustained 100 ops/sec against a cold buffer cache drove DCU from 0.5 to 1.6 over ~2 minutes. It is not just CPU that matters.
Combined: DocumentDB Serverless reacts to resource pressure (CPU, IOPS, memory), not to counters that look correlated with load (connections, pings). Your connection pool is not masking a signal.
How we measured this
Every finding is drawn from a dedicated experiment on a single-instance DocumentDB 8.0.1 Serverless cluster in
us-east-1, driven by a PyMongo 4.9.2 workload from an EC2 client in the same VPC. Each experiment isolates one variable at a time and captures both server-side (CloudWatch, 30s cadence) and client-side (1 Hz latency probe, 60× resolution) observations. The full methodology, phase-by-phase design, and raw JSON results are in Part 0 of this series.Experiment 1: Does connection count trigger scale-out?
Setup. Cluster at MinCapacity 1.0 (measured client ceiling: 250 PyMongo clients). Open 300 clients in a controlled ramp. The 51st through 300th attempts should hit the ceiling and be refused. Observe DCU for 8.5 minutes with the refusal condition sustained.
Result.
| Time | Client attempts | Server-side sockets | Refused clients | DCU |
|---|---|---|---|---|
| t+0s | 300 opened | 500 (ceiling) | 50 | 1.0 |
| t+2min | 300 held | 500 | 50 | 1.0 |
| t+5min | 300 held | 500 | 50 | 1.0 |
| t+8.5min | 300 held | 500 | 50 | 1.0 |
Zero DCU response. Zero scale-out event. The refused clients hit
ConnectionFailure on the client side; the cluster served the 250 accepted clients at low CPU and moved on. Connection refusal is not among the signals the autoscaler observes.The mechanism is straightforward: DocumentDB Serverless watches for resource pressure (CPU utilization, IOPS, memory usage) produced by actual work happening inside the engine. A client that fails to connect at the socket layer never enters the engine, so it produces no pressure. From the autoscaler's perspective, a refused connection and a connection that was never attempted are indistinguishable.
Practical implication. If you want to accommodate a connection burst, you must set MinCapacity high enough to hold the peak socket count. Autoscaling will not rescue you. See Part 4 of this series for the sizing rule.
Experiment 2: Do idle connections trigger (or prevent) scale-in?
Setup. Warm the cluster with 100 clients × 30 ops/sec for 3 minutes (DCU stabilizes near 14). Stop the workload but keep N clients open, pinging every 30 seconds. Vary N across five values. Measure time to reach MinCapacity floor (DCU < 1.0).
Result.
| Idle clients held | Time to MinCap floor |
|---|---|
| 0 (all closed) | 10 minutes |
| 10 | 15 minutes |
| 50 | 16 minutes |
| 100 | 16 minutes |
| 500 | Never (DCU stabilized at ~2.0 through 20 min observation) |
Two things stand out.
Below 500, the effect is modest. Whether you hold 10, 50, or 100 idle clients, the cluster reaches its floor in 15-17 minutes vs. 10 minutes with everything closed. Six extra minutes of elevated capacity is real but small.
At 500 idle clients, the cluster stopped shrinking. DCU settled at ~2.0 and held there. This exactly matches the connection ceiling formula: 500 clients ÷ 250 clients-per-DCU = 2 DCU. The autoscaler will not scale below the level required to serve the open sockets, even if those sockets are transmitting no application traffic.

Figure: DCU trajectory over the 20 minutes after workload ends. With connections closed, the cluster reaches MinCap floor (0.5 DCU) in 10 minutes. With 100 idle connections held, the same drop takes 16 minutes. With 500 idle connections held, DCU stalls at ~2 (500 clients ÷ 250 clients-per-DCU) and never returns to floor.
Per-minute data (minutes after workload stops):
| t (min) | 0 conns | 100 conns idle | 500 conns idle |
|---|---|---|---|
| 0 | 13.12 | 14.50 | 12.50 |
| 2 | 12.50 | 14.50 | 11.00 |
| 4 | 12.00 | 14.50 | 8.00 |
| 6 | 5.99 | 14.45 | 6.33 |
| 8 | 5.00 | 11.04 | 5.00 |
| 10 | 0.57 | 6.50 | 2.15 |
| 12 | 0.51 | 2.31 | 2.00 |
| 14 | 0.51 | 1.58 | 1.80 |
| 16 | 0.51 | 0.96 | 1.64 |
| 18 | 0.51 | 0.53 | 1.62 |
| 20 | 0.51 | 0.53 | 1.60 |
Interpretation. Idle connections are not a scale-out signal (they never pushed DCU up), but they are a scale-in constraint. The autoscaler keeps enough capacity to honor the client ceiling for connections you have already opened.
Practical implication. If your application uses long-lived connection pools that don't close after a workload burst, expect the post-burst capacity to shrink slowly, and to floor out at
open_connections / 250 DCU rather than at MinCapacity. Configure driver-side idle timeout (maxIdleTimeMS in PyMongo, equivalents in other drivers) if you want the cluster to reach true floor between bursts.Experiment 3: Does ping frequency matter?
Setup. Same as Experiment 2 with 100 idle clients held. Vary the keep-alive ping interval across three values.
Result.
| Ping interval | Time to MinCap floor |
|---|---|
| 0s (no pings; TCP-only, PyMongo heartbeat off) | 17 minutes |
| 5s (aggressive pings) | 16 minutes |
| 30s (default-ish) | 16 minutes |
The three trajectories are within one minute of each other across a 20-minute window. Ping frequency does not measurably affect scale-in.
Interpretation. The autoscaler observes connection presence at the TCP layer, not activity at the application layer. A client that pings every 5 seconds looks the same as a client that never sends anything, as long as the socket stays open. This is consistent with Experiment 1: the autoscaler cares about actual engine-side pressure, and neither pings nor idle sockets create pressure.
Practical implication. You cannot game scale-in by lowering ping frequency. If you want faster scale-in, close the connections; don't just quiet them.
Experiment 4: Does I/O pressure from cache misses trigger scale-out?
This one asks the opposite question: can the cluster tell that reads are hitting storage instead of cache, and scale out to give the buffer more room?
Setup. After a natural scale-in cycle (cluster settled at MinCapacity 0.5, buffer cache eviced), run a workload of 20 clients × 5 ops/sec against a small hot set of 100 documents (25 KB total, easily fits in a 0.5 DCU cache once warm). Track DCU trajectory over 10 minutes.
Result.
- t+0s: workload starts, cluster at 0.5 DCU, cache empty. First reads hit storage at 100-180 ms individual latencies.
- t+30s: cache warm for the hot set, latency stabilizes at P50 ~1.2 ms.
- t+2min: DCU rises from 0.5 to 1.6. Sustained IOPS pressure during the cold-cache period (~100 ops/sec, all storage-side) triggered a scale-out response.
- t+2-10min: DCU holds at 1.6, cache stays warm, latencies stable.
Only 100 ops/sec was enough - but only because the reads were all hitting storage. The same 100 ops/sec against a warm cache would not have triggered scale-out because the pages would have served from RAM at near-zero I/O cost.
Interpretation. IOPS pressure is an independent scale-out signal from CPU. A read-heavy workload against a cold cache produces the same autoscaler response as a CPU-heavy compute workload of comparable size. This is consistent with the general model: the autoscaler watches for resource pressure at the engine level, and hitting storage is a form of pressure that CPU alone would not reveal.
Practical implication. If your workload has a large working set and starts cold, expect a scale-out cycle during the first minute or two as pages fault in. If you want to avoid this - because the initial storage-latency window is user-facing - pre-warm the cluster (drive workload before opening it to production traffic) or size MinCapacity high enough that the working set stays resident. Part 3 of this series covers cold-cache behavior in more detail.
Synthesis: what the autoscaler actually watches
Across four experiments, a consistent picture:
| Signal | Scale-out trigger? | Scale-in constraint? |
|---|---|---|
| CPU utilization | Yes | Yes (low CPU is a prerequisite for scale-in) |
| IOPS / storage reads | Yes | Yes (low IOPS is a prerequisite for scale-in) |
| Memory pressure | Yes (inferred; not isolated in these experiments) | Yes |
| Client connection count (attempted) | No | No |
| Client connections (established, idle) | No | Yes (holds capacity above sockets/500 per DCU) |
| Application-layer pings | No | No |
Scale-out is triggered by workload actually reaching the engine. Scale-in is gated by that same workload dropping AND by having room to reduce socket capacity below the current open-connection count.
What this means for your application
Three practical rules from this article. (The full 5-rule guidance set lives in Part 4.)
Rule 1: Don't design connection-pool behavior expecting it to influence autoscaling. A pool that holds 200 idle clients will not make the cluster scale down slower in any operationally meaningful way (below 500 conns, the delta is ~6 minutes). A pool that holds 1000 clients WILL pin DCU at
1000/250 = 4 DCU indefinitely. If your pool sizing crosses the ceiling-per-current-DCU line, understand you have effectively set a lower bound on capacity.Rule 2: If you want fast scale-in between bursts, close idle connections. Configure
maxIdleTimeMS (PyMongo) or the driver equivalent for a value shorter than your inter-burst quiet window. This keeps the socket count low enough that the autoscaler is free to return to floor.Rule 3: A cold-cache workload will scale the cluster out even at modest ops/sec. This is generally desirable (more capacity when needed), but it means the DCU trajectory during warm-up is not the same as the steady-state trajectory afterward. Don't size MinCapacity from a benchmark's peak-DCU number without checking whether the peak was warm-cache or cold-cache driven.
What we didn't measure
- Which of CPU / IOPS / memory dominates. All three appeared as scale-out signals; we did not run isolation experiments (e.g. IOPS-heavy but CPU-idle vs. CPU-heavy but IOPS-idle) that would let us weight them. Practically this doesn't matter much - the autoscaler treats them as combined pressure - but it's an open question for future work.
- Multi-instance cluster behavior. All experiments ran against single-instance serverless. Multi-instance serverless (writer + serverless replicas) may have different scale-in coordination.
- Non-PyMongo drivers. The connection-ceiling behavior (250 clients ↔ 500 sockets) is PyMongo-specific because of PyMongo's 2-socket-per-client model. Other drivers should exhibit the same scaling-signal logic; the ceiling numbers translate through their own socket-per-client ratio. See Part 4 for how to verify empirically.
Coming next
Part 2 - How fast does DocumentDB Serverless actually scale? Time-to-90%-throughput measurements across MinCapacity floors from 0.5 to 8 DCU, twelve iterations, and what the CloudWatch aggregation window hides.
Part 3 - Buffer cache eviction and cold-read latency. What actually happens to your P99 read latency after a scale-in cycle, and how long the cache takes to re-warm.
Part 4 - Sizing MinCapacity: three rules. Connection ceiling, working-set size, and steady-state throughput as inputs.
Part 5 - When NOT to use Serverless. Where provisioned still wins.
Reproducibility. All 15 experiments, raw JSON results, and Python driver scripts are in the companion GitHub repository . If you want to verify a specific behavior against your own driver or workload pattern, clone the repo and modify the relevant
phase_*.py script.Feedback. Comments, corrections, and questions welcome via the GitHub issues page on the reproducer repo, or reach out to me internally.
Series: Amazon DocumentDB Serverless: An Empirical Study (5 articles)
- 1What Triggers Amazon DocumentDB Serverless to Scale (And What Doesn't) This article
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article