AWS Builder Center
What Triggers Amazon DocumentDB Serverless to Scale (And What Doesn't)

What Triggers Amazon DocumentDB Serverless to Scale (And What Doesn't)

Part 1 of 5 diving deep in Amazon DocumentDB Serverless through empirical testings

Series: Amazon DocumentDB Serverless: An Empirical Study (5 articles)

  1. 1
    What Triggers Amazon DocumentDB Serverless to Scale (And What Doesn't) This article
Part 1 of 5 · Amazon DocumentDB Serverless: An Empirical Study

Amazon DocumentDB Serverless scales capacity in response to workload pressure. The obvious question is: which pressure? This article measures four candidate signals - client connection count, idle connection presence, keep-alive ping frequency, and I/O pressure from cache misses - and reports which of them actually move the DCU meter. If you have ever wondered whether your connection pool is masking a scale-out signal, or whether closing idle clients at end-of-shift will help the cluster shrink back to its floor, this piece answers both questions with measurement rather than guess.

Who this is for

You are a developer or architect running (or evaluating) Amazon DocumentDB Serverless. You have read the product page  and know that "scaling is driven by workload", but the word "workload" hides more than it reveals. You want to know which knobs actually turn the dial before you commit to a MinCapacity value or design a connection pool.

Executive summary

Four findings from controlled experiments on DocumentDB 8.0.1 Serverless:
  1. Connection count is not a scale-out signal. We opened 300 client connections against a cluster at MinCapacity 1.0 (ceiling 250). The extra 50 clients were refused for 8.5 continuous minutes. DCU never moved.
  2. Idle connections extend scale-in, but only above a threshold. 10 to 100 idle clients delay time-to-floor by ~6 minutes vs. all-closed. 500 idle clients pin DCU at ~2 indefinitely - the cluster holds exactly the DCU required to accommodate the open sockets and stops shrinking.
  3. Ping frequency doesn't matter. Whether idle clients pinged every 5s, every 30s, or not at all (TCP-only), the scale-in trajectory was the same. The autoscaler observes TCP presence, not application-layer activity.
  4. I/O pressure from cache misses DOES trigger scale-out. Sustained 100 ops/sec against a cold buffer cache drove DCU from 0.5 to 1.6 over ~2 minutes. It is not just CPU that matters.
Combined: DocumentDB Serverless reacts to resource pressure (CPU, IOPS, memory), not to counters that look correlated with load (connections, pings). Your connection pool is not masking a signal.

How we measured this

Every finding is drawn from a dedicated experiment on a single-instance DocumentDB 8.0.1 Serverless cluster in us-east-1, driven by a PyMongo 4.9.2 workload from an EC2 client in the same VPC. Each experiment isolates one variable at a time and captures both server-side (CloudWatch, 30s cadence) and client-side (1 Hz latency probe, 60× resolution) observations. The full methodology, phase-by-phase design, and raw JSON results are in Part 0 of this series.

Experiment 1: Does connection count trigger scale-out?

Setup. Cluster at MinCapacity 1.0 (measured client ceiling: 250 PyMongo clients). Open 300 clients in a controlled ramp. The 51st through 300th attempts should hit the ceiling and be refused. Observe DCU for 8.5 minutes with the refusal condition sustained.
Result.
TimeClient attemptsServer-side socketsRefused clientsDCU
t+0s300 opened500 (ceiling)501.0
t+2min300 held500501.0
t+5min300 held500501.0
t+8.5min300 held500501.0
Zero DCU response. Zero scale-out event. The refused clients hit ConnectionFailure on the client side; the cluster served the 250 accepted clients at low CPU and moved on. Connection refusal is not among the signals the autoscaler observes.
The mechanism is straightforward: DocumentDB Serverless watches for resource pressure (CPU utilization, IOPS, memory usage) produced by actual work happening inside the engine. A client that fails to connect at the socket layer never enters the engine, so it produces no pressure. From the autoscaler's perspective, a refused connection and a connection that was never attempted are indistinguishable.
Practical implication. If you want to accommodate a connection burst, you must set MinCapacity high enough to hold the peak socket count. Autoscaling will not rescue you. See Part 4 of this series for the sizing rule.

Experiment 2: Do idle connections trigger (or prevent) scale-in?

Setup. Warm the cluster with 100 clients × 30 ops/sec for 3 minutes (DCU stabilizes near 14). Stop the workload but keep N clients open, pinging every 30 seconds. Vary N across five values. Measure time to reach MinCapacity floor (DCU < 1.0).
Result.
Idle clients heldTime to MinCap floor
0 (all closed)10 minutes
1015 minutes
5016 minutes
10016 minutes
500Never (DCU stabilized at ~2.0 through 20 min observation)
Two things stand out.
Below 500, the effect is modest. Whether you hold 10, 50, or 100 idle clients, the cluster reaches its floor in 15-17 minutes vs. 10 minutes with everything closed. Six extra minutes of elevated capacity is real but small.
At 500 idle clients, the cluster stopped shrinking. DCU settled at ~2.0 and held there. This exactly matches the connection ceiling formula: 500 clients ÷ 250 clients-per-DCU = 2 DCU. The autoscaler will not scale below the level required to serve the open sockets, even if those sockets are transmitting no application traffic.
Figure: DCU trajectory over the 20 minutes after workload ends. With connections closed, the cluster reaches MinCap floor (0.5 DCU) in 10 minutes. With 100 idle connections held, the same drop takes 16 minutes. With 500 idle connections held, DCU stalls at ~2 (500 clients ÷ 250 clients-per-DCU) and never returns to floor.
Per-minute data (minutes after workload stops):
t (min)0 conns100 conns idle500 conns idle
013.1214.5012.50
212.5014.5011.00
412.0014.508.00
65.9914.456.33
85.0011.045.00
100.576.502.15
120.512.312.00
140.511.581.80
160.510.961.64
180.510.531.62
200.510.531.60
Interpretation. Idle connections are not a scale-out signal (they never pushed DCU up), but they are a scale-in constraint. The autoscaler keeps enough capacity to honor the client ceiling for connections you have already opened.
Practical implication. If your application uses long-lived connection pools that don't close after a workload burst, expect the post-burst capacity to shrink slowly, and to floor out at open_connections / 250 DCU rather than at MinCapacity. Configure driver-side idle timeout (maxIdleTimeMS in PyMongo, equivalents in other drivers) if you want the cluster to reach true floor between bursts.

Experiment 3: Does ping frequency matter?

Setup. Same as Experiment 2 with 100 idle clients held. Vary the keep-alive ping interval across three values.
Result.
Ping intervalTime to MinCap floor
0s (no pings; TCP-only, PyMongo heartbeat off)17 minutes
5s (aggressive pings)16 minutes
30s (default-ish)16 minutes
The three trajectories are within one minute of each other across a 20-minute window. Ping frequency does not measurably affect scale-in.
Interpretation. The autoscaler observes connection presence at the TCP layer, not activity at the application layer. A client that pings every 5 seconds looks the same as a client that never sends anything, as long as the socket stays open. This is consistent with Experiment 1: the autoscaler cares about actual engine-side pressure, and neither pings nor idle sockets create pressure.
Practical implication. You cannot game scale-in by lowering ping frequency. If you want faster scale-in, close the connections; don't just quiet them.

Experiment 4: Does I/O pressure from cache misses trigger scale-out?

This one asks the opposite question: can the cluster tell that reads are hitting storage instead of cache, and scale out to give the buffer more room?
Setup. After a natural scale-in cycle (cluster settled at MinCapacity 0.5, buffer cache eviced), run a workload of 20 clients × 5 ops/sec against a small hot set of 100 documents (25 KB total, easily fits in a 0.5 DCU cache once warm). Track DCU trajectory over 10 minutes.
Result.
  • t+0s: workload starts, cluster at 0.5 DCU, cache empty. First reads hit storage at 100-180 ms individual latencies.
  • t+30s: cache warm for the hot set, latency stabilizes at P50 ~1.2 ms.
  • t+2min: DCU rises from 0.5 to 1.6. Sustained IOPS pressure during the cold-cache period (~100 ops/sec, all storage-side) triggered a scale-out response.
  • t+2-10min: DCU holds at 1.6, cache stays warm, latencies stable.
Only 100 ops/sec was enough - but only because the reads were all hitting storage. The same 100 ops/sec against a warm cache would not have triggered scale-out because the pages would have served from RAM at near-zero I/O cost.
Interpretation. IOPS pressure is an independent scale-out signal from CPU. A read-heavy workload against a cold cache produces the same autoscaler response as a CPU-heavy compute workload of comparable size. This is consistent with the general model: the autoscaler watches for resource pressure at the engine level, and hitting storage is a form of pressure that CPU alone would not reveal.
Practical implication. If your workload has a large working set and starts cold, expect a scale-out cycle during the first minute or two as pages fault in. If you want to avoid this - because the initial storage-latency window is user-facing - pre-warm the cluster (drive workload before opening it to production traffic) or size MinCapacity high enough that the working set stays resident. Part 3 of this series covers cold-cache behavior in more detail.

Synthesis: what the autoscaler actually watches

Across four experiments, a consistent picture:
SignalScale-out trigger?Scale-in constraint?
CPU utilizationYesYes (low CPU is a prerequisite for scale-in)
IOPS / storage readsYesYes (low IOPS is a prerequisite for scale-in)
Memory pressureYes (inferred; not isolated in these experiments)Yes
Client connection count (attempted)NoNo
Client connections (established, idle)NoYes (holds capacity above sockets/500 per DCU)
Application-layer pingsNoNo
Scale-out is triggered by workload actually reaching the engine. Scale-in is gated by that same workload dropping AND by having room to reduce socket capacity below the current open-connection count.

What this means for your application

Three practical rules from this article. (The full 5-rule guidance set lives in Part 4.)
Rule 1: Don't design connection-pool behavior expecting it to influence autoscaling. A pool that holds 200 idle clients will not make the cluster scale down slower in any operationally meaningful way (below 500 conns, the delta is ~6 minutes). A pool that holds 1000 clients WILL pin DCU at 1000/250 = 4 DCU indefinitely. If your pool sizing crosses the ceiling-per-current-DCU line, understand you have effectively set a lower bound on capacity.
Rule 2: If you want fast scale-in between bursts, close idle connections. Configure maxIdleTimeMS (PyMongo) or the driver equivalent for a value shorter than your inter-burst quiet window. This keeps the socket count low enough that the autoscaler is free to return to floor.
Rule 3: A cold-cache workload will scale the cluster out even at modest ops/sec. This is generally desirable (more capacity when needed), but it means the DCU trajectory during warm-up is not the same as the steady-state trajectory afterward. Don't size MinCapacity from a benchmark's peak-DCU number without checking whether the peak was warm-cache or cold-cache driven.

What we didn't measure

  • Which of CPU / IOPS / memory dominates. All three appeared as scale-out signals; we did not run isolation experiments (e.g. IOPS-heavy but CPU-idle vs. CPU-heavy but IOPS-idle) that would let us weight them. Practically this doesn't matter much - the autoscaler treats them as combined pressure - but it's an open question for future work.
  • Multi-instance cluster behavior. All experiments ran against single-instance serverless. Multi-instance serverless (writer + serverless replicas) may have different scale-in coordination.
  • Non-PyMongo drivers. The connection-ceiling behavior (250 clients ↔ 500 sockets) is PyMongo-specific because of PyMongo's 2-socket-per-client model. Other drivers should exhibit the same scaling-signal logic; the ceiling numbers translate through their own socket-per-client ratio. See Part 4 for how to verify empirically.

Coming next

Part 2 - How fast does DocumentDB Serverless actually scale? Time-to-90%-throughput measurements across MinCapacity floors from 0.5 to 8 DCU, twelve iterations, and what the CloudWatch aggregation window hides.
Part 3 - Buffer cache eviction and cold-read latency. What actually happens to your P99 read latency after a scale-in cycle, and how long the cache takes to re-warm.
Part 4 - Sizing MinCapacity: three rules. Connection ceiling, working-set size, and steady-state throughput as inputs.
Part 5 - When NOT to use Serverless. Where provisioned still wins.

Reproducibility. All 15 experiments, raw JSON results, and Python driver scripts are in the companion GitHub repository . If you want to verify a specific behavior against your own driver or workload pattern, clone the repo and modify the relevant phase_*.py script.
Feedback. Comments, corrections, and questions welcome via the GitHub issues page on the reproducer repo, or reach out to me internally.

Series: Amazon DocumentDB Serverless: An Empirical Study (5 articles)

  1. 1
    What Triggers Amazon DocumentDB Serverless to Scale (And What Doesn't) This article
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article