AWS Builder Center
Sizing MinCapacity for Amazon DocumentDB Serverless: Three Rules

Sizing MinCapacity for Amazon DocumentDB Serverless: Three Rules

Part 4 of 5 diving deep in Amazon DocumentDB Serverless through empirical testings

Series: Amazon DocumentDB Serverless: An Empirical Study (5 articles)

  1. …
  2. 4
    Sizing MinCapacity for Amazon DocumentDB Serverless: Three Rules This article
MinCapacity is the most consequential configuration decision for Amazon DocumentDB Serverless. It's the floor DCU below which the cluster will not scale in, which makes it simultaneously your peak-connection ceiling, your working-set cache size, and your latency-SLO floor. Set it too low and you refuse client connections or evict working sets during quiet windows. Set it too high and you pay for capacity you don't need. This article gives three concrete inputs to size MinCapacity from - peak concurrent connections, working-set memory footprint, and steady-state throughput - with the empirical numbers behind each rule.

Who this is for

You are about to deploy an application on Amazon DocumentDB Serverless, or you have deployed and now need to right-size MinCapacity based on observed behavior. You want a decision framework, not vague guidance.

Executive summary

Three rules from measurement:
  1. Connection ceiling: 250 PyMongo clients per DCU (or 500 raw sockets per DCU). Verified from both server-side and client-side independently, with an exact 2.00 sockets-per-client ratio across the entire test range. Other drivers may have different ratios; verify empirically.
  2. Working-set cache: approximately 2 GB per DCU. As a rule of thumb, the buffer cache holds ~2 GB of pages per DCU. Size MinCap so your hot working set stays resident.
  3. Steady-state throughput: DCU settles at approximately throughput / 200. 300 ops/sec settled at ~7 DCU. 1,000 ops/sec settled at ~14 DCU. On our specific test setup (t4g.xlarge client running 300 PyMongo threads on a small-document insert workload), aggregate throughput plateaued at ~1,750-2,000 ops/sec around DCU 15-18. We did not isolate whether that plateau lived on the cluster side or the client side; different test setups may plateau at very different points. DocumentDB Serverless itself supports MaxCapacity up to 256 DCU.
Each rule gives an independent MinCapacity floor. Take the maximum of the three. Choosing MinCap = max(conn_floor, cache_floor, throughput_floor) is a reasonable default; scaling down from there requires understanding which floor gets crossed first.

How we measured this

Same lab as Parts 1-3: single-instance DocumentDB 8.0.1 Serverless in us-east-1, PyMongo 4.9.2 workload from an EC2 client in the same VPC. The full methodology and reproducer scripts are in Part 0 of this series.

Rule 1: Size for peak concurrent client connections

Setup. Ramp PyMongo MongoClient instances from 100 to 4,000 at MinCapacity 0.5, 1, 2, 4, 8. Record the connection count at which new clients start being refused.
Result.
MinCapacityPyMongo clients acceptedRaw sockets accepted
0.5124248
1250500
25001,000
41,0002,000
82,0004,000
The relationship is exactly linear: 250 PyMongo clients per DCU, or 500 raw sockets per DCU. These two formulas describe the same ceiling from different perspectives.

Why two numbers for the same ceiling

AWS documents the ceiling as 500 connections per DCU. We measured 250 client connections per DCU using PyMongo. The gap is fully explained by PyMongo's socket model:
Setup for the verification test. Progressively add MongoClient instances (1, +3 more, +10 more, +50 more). At each step, capture both server-side serverStatus.connections.current and client-side ss -tn state established | grep :27017. Compute the ratio of raw sockets to MongoClients added.
Result.
MongoClients addedServer-side deltaClient-side netstat deltaSockets per client
+1+2+22.00
+3+6+62.00
+10+20+202.00
+50+100+1002.00
Each PyMongo MongoClient opens exactly 2 TCP sockets to the target server. One for application operations, one for topology heartbeat monitoring. The 2.00 ratio held constant across a 50× range in client count. Server-side and client-side observations agreed exactly.
So: 250 clients × 2 sockets = 500 sockets = AWS-documented ceiling. Reconciled.
Practical implication. Your application-visible connection pool is not the same unit as the AWS-documented ceiling. Size against the client-level ceiling for your specific driver. Other drivers (Java, Node.js, Go) likely have different sockets-per-client ratios depending on their connection-pool + monitor architecture. Run the verification test with your driver if you're operating close to the ceiling.

Applying Rule 1

1
conn_floor_DCU = peak_concurrent_clients / clients_per_DCU_your_driver
For PyMongo with clients_per_DCU = 250:
Peak concurrent clientsMinimum DCU (Rule 1 only)
1000.5
2501
5002
1,0004
2,0008
Add a headroom margin. Sizing for exactly 100% ceiling means the next client is refused. A common practice is to size for 80% utilization at peak: conn_floor_DCU = (peak_clients / clients_per_DCU) / 0.8.

Rule 2: Size for working-set cache retention

Setup. Test whether a working set of size X fits in cache at DCU level Y. See Part 3, Experiment 2 for the full test. The rough conclusion: each DCU provides approximately 2 GB of memory for buffer caching.
Some of that 2 GB is used for other purposes (query buffers, connection state, internal data structures), but for order-of-magnitude sizing, 2 GB/DCU is the working number.
Result of applying this to Part 3's dataset.
The 500 MB dataset (2M docs) fit comfortably in 10 DCU (20 GB cache). At 0.5 DCU (1 GB cache), the same dataset did NOT fit - hence the observed 4.6× P50 latency degradation when the cluster scaled down and the cache was evicted.

Applying Rule 2

1
cache_floor_DCU = working_set_size_GB / 2
Working set (compressed)Minimum DCU (Rule 2 only)
500 MB0.5
2 GB1
10 GB5
50 GB25
"Working set" is not the same as "total data size". It is the portion of data actively read by queries within a window (typically a minute or less). A 1 TB dataset with only 10 GB of hot pages has a 10 GB working set, not 1 TB.
How to measure your working set. Watch BufferCacheHitRatio in CloudWatch:
  • Sustained >95% at steady state = current cache is large enough for your working set.
  • Sustained <95% = working set exceeds cache, MinCapacity should be higher.
  • Occasional dips during scale-in and recovery = expected. Persistent low ratio = undersized.
Note the CloudWatch caveat from Part 3: the metric aggregates over 60 seconds and can smooth over sub-minute cache-miss bursts. Trust client-side latency histograms for real-time verification.

Rule 3: Size for steady-state throughput

Setup. Test the DCU level at which the cluster settles under sustained partial workload. See Part 2, Experiment 6 for the full test.
Result.
Sustained throughputSettled DCU
300 ops/sec~7 DCU
1,000 ops/sec~14 DCU
The relationship is roughly linear: DCU ≈ throughput / 200. This coefficient depends on operation mix, document size, and index efficiency, so treat it as a starting estimate. Also note: our two validated data points are 300 ops/sec and 1,000 ops/sec — extrapolating to higher throughputs (say, 3,000+ ops/sec) assumes the coefficient stays linear, which we did not test. Load-test if your target exceeds ~1,000 ops/sec.
Observed throughput plateau in our test setup — bottleneck location not isolated. In Part 2's higher-workload tests (300 PyMongo clients × 30 ops/sec target from a t4g.xlarge EC2 client), aggregate throughput plateaued at ~1,750-2,000 ops/sec around DCU 15-18. Adding more client connections beyond that point did not increase throughput (per-connection rate dropped instead) AND did not push DCU higher (the autoscaler saw no additional CPU/IOPS/memory pressure).
Important caveat: we did not isolate where the ceiling actually lived. The plateau could have been on:
  • The cluster side — writer-node serialization, per-instance storage IOPS, wire-protocol contention, query engine bottlenecks
  • The client side — t4g.xlarge has 4 vCPUs and 300 concurrent PyMongo threads competing for the Python GIL is a plausible plateau point; TCP ephemeral port pressure, driver connection-pool internal locking, or network throughput on the EC2 client are also candidates
A test with a bigger client instance, a different driver, larger documents, or a different concurrency model (multiprocess instead of threading) could produce a different plateau. We did not run those variations.
DocumentDB Serverless supports MaxCapacity up to 256 DCU per AWS documentation. Our tests capped MaxCap at 64, which was never reached during the plateau. Real production workloads with different profiles may sustain significantly higher throughput on single-instance serverless.
Practical guidance: if your production workload will exceed the throughput levels we tested, don't extrapolate from our numbers — load-test with your own client sizing, driver, workload pattern, and document size, then observe where your actual plateau sits. That's the number you should size against.

Applying Rule 3

1
throughput_floor_DCU = expected_steady_state_ops_per_sec / 200
If MinCap is set below the natural steady-state DCU, the cluster will scale up on every workload cycle - which is fine functionally, but it means MinCap is not doing its job as a "warm baseline". If MinCap is set above the steady-state DCU, you are paying for capacity above the natural settling point.
The common pattern: set MinCap slightly BELOW the natural steady-state (say, 80% of it), so the cluster runs at MinCap during genuine idle windows and briefly scales up when the baseline workload arrives.

Rule 4 (implicit): Set MaxCapacity as a cost safety net

Not measured empirically in this series (it doesn't need measurement), but worth stating explicitly. MaxCapacity caps the ceiling of scale-out. Without a cap, a runaway query or bad deployment could drive the cluster to DCU 256 (the AWS-documented maximum), producing a large one-time cost surprise.
  • Set MaxCapacity to ~2-3× your expected steady-state DCU. This gives room for legitimate bursts while backpressuring runaway growth.
  • CloudWatch alarm on ServerlessDatabaseCapacity at MaxCapacity. If the cluster regularly hits MaxCap, either raise the cap (real growth) or investigate (bug).

Putting the rules together

For a typical OLTP workload:
  • Peak concurrent PyMongo clients: 800 → conn_floor = 800/250 = 3.2 DCU → round up to 4 DCU.
  • Working set (compressed): 12 GB → cache_floor = 12/2 = 6 DCU.
  • Steady-state throughput: 500 ops/sec → throughput_floor = 500/200 = 2.5 DCU.
MinCap = max(4, 6, 2.5) = 6 DCU. The working set was the binding constraint.
For a batch analytics workload:
  • Peak concurrent clients: 20 → conn_floor = 0.5 DCU.
  • Working set: 40 GB → cache_floor = 20 DCU.
  • Peak throughput during batch: 3,000 ops/sec → throughput_floor ≈ 15 DCU (extrapolated using /200 beyond our tested range — actual settling point may differ; load-test to confirm before committing).
MinCap = max(0.5, 20, 15) = 20 DCU. The working set was the binding constraint, so the extrapolated throughput floor doesn't affect the final sizing decision here — but note that if working set were smaller, throughput would drive the answer and the extrapolation uncertainty would matter more. Consider whether this workload is actually better on a provisioned instance - see Part 5.

What we didn't measure

  • The exact 2 GB/DCU cache ratio. This is a working number from observed behavior, not a documented AWS specification. It may vary by DocumentDB version or workload pattern.
  • Sizing for write-heavy workloads. All experiments were dominated by inserts and reads; sustained update-heavy or delete-heavy workloads may have different DCU-per-ops-sec coefficients.
  • The behavior of multi-instance serverless clusters. All measurements are single-instance. Multi-instance scaling may distribute cache and connections differently.

Coming next

Part 5 - When to Choose Provisioned Over Amazon DocumentDB Serverless. Trade-offs across latency SLOs, sustained throughput, and cost predictability.
Previously:

Reproducibility. All 15 experiments, raw JSON results, and Python driver scripts are in the companion GitHub repository .

Series: Amazon DocumentDB Serverless: An Empirical Study (5 articles)

  1. …
  2. 4
    Sizing MinCapacity for Amazon DocumentDB Serverless: Three Rules This article
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article