AWS Builder Center

The Hidden Cost of Calling PutMetricData on Every Request

Why PutMetricData on a hot path silently becomes your largest CloudWatch cost and the batching fix that solves it

Software Engineer come Cloud&Devops
CloudWatch bills feel invisible right up until they aren't. Compute, storage, and data transfer get watched closely because they're the "obvious" cost centers. Observability rarely gets the same scrutiny until a bill shows up with a line item nobody expected to be the biggest one.
This post walks through a specific, easy-to-reproduce cost bug which i faced while streaming data and ended up with $200: calling PutMetricData synchronously, once or twice per request, on a hot code path. It's a small, innocent-looking design choice that scales into real money far faster than most people expect and it's one of the more common patterns I've run into while auditing production systems.

The mechanism

PutMetricData is billed per API request, not per metric value: $0.01 per 1,000 requests, after the first 1,000,000 requests each month (AWS's own free tier). That sounds cheap. It is cheap per call.
The problem shows up when this call sits inside a hot path: a request handler, a per-tool wrapper, a middleware layer anything that runs on every single request your service handles. If your code calls put_metric_data even once per request, your request volume is your CloudWatch API request volume. Call it twice per request (one for latency, one for a status counter a very common pattern), and you've doubled it again.
Here's where it stops being a rounding error. A service handling a modest 50 requests/second, calling put_metric_data twice per request, generates:
1
50 req/sec × 2 calls × 86,400 sec/day × 30 days ≈ 259.2 million requests/month
At $0.01 per 1,000 requests (minus the 1M free tier), that's roughly $2,580/month for a service that most teams would describe as "medium traffic," not "at scale." I've seen this exact shape of bug account for the majority of a service's entire CloudWatch spend, dwarfing the actual custom-metric storage charges it was supposedly there to support.

Reproducing it

Here's a minimal, runnable demo of the naive pattern the kind you'd find inside a finally block wrapping a request handler:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
import boto3
import time

cloudwatch = boto3.client("cloudwatch")

def record_call_naive(tool_name: str, duration_ms: float):
"""Called on every single request. This is the bug."""
cloudwatch.put_metric_data(
Namespace="MyService",
MetricData=[
{
"MetricName": "CallDuration",
"Dimensions": [{"Name": "Tool", "Value": tool_name}],
"Value": duration_ms,
"Unit": "Milliseconds",
},
{
"MetricName": "CallCount",
"Dimensions": [{"Name": "Tool", "Value": tool_name}],
"Value": 1,
"Unit": "Count",
},
],
)

# Simulating 1,000 requests to show the call count, not to actually run against real CloudWatch:
for i in range(1000):
start = time.time()
# ... your actual request handling here ...
duration = (time.time() - start) * 1000
record_call_naive("get_data", duration)
# 1,000 requests => 1,000 live API calls, right here.
Run that shape of code at real production traffic, and the request count above is exactly what you get: one live, network-bound AWS API call per request, per metric group you're recording.

The fix: buffer, then flush on an interval

CloudWatch already gives you the tool to fix this: PutMetricData accepts pre-aggregated StatisticValues (SampleCount, Sum, Minimum, Maximum) instead of one data point per call. So instead of calling the API per request, buffer values in memory and flush the aggregate on a fixed interval 30 seconds is a reasonable default for most dashboards.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
import threading
import time
import boto3
from collections import defaultdict

cloudwatch = boto3.client("cloudwatch")
_lock = threading.Lock()
_buffer = defaultdict(list) # key: (metric_name, tool_name) -> list of values

def record_call_buffered(tool_name: str, duration_ms: float):
"""Called on every request, but does zero network calls."""
with _lock:
_buffer[("CallDuration", tool_name)].append(duration_ms)
_buffer[("CallCount", tool_name)].append(1)

def flush_metrics():
"""Runs on a timer, e.g. every 30 seconds this is the only place
that actually talks to CloudWatch."""

with _lock:
items = list(_buffer.items())
_buffer.clear()

metric_data = []
for (metric_name, tool_name), values in items:
metric_data.append({
"MetricName": metric_name,
"Dimensions": [{"Name": "Tool", "Value": tool_name}],
"Timestamp": time.time(),
"StatisticValues": {
"SampleCount": len(values),
"Sum": sum(values),
"Minimum": min(values),
"Maximum": max(values),
},
"Unit": "Milliseconds" if metric_name == "CallDuration" else "Count",
})

# CloudWatch allows up to 1,000 metric data points per PutMetricData call,
# but batch conservatively (20 here) to keep individual payloads small.
for i in range(0, len(metric_data), 20):
chunk = metric_data[i:i + 20]
if chunk:
cloudwatch.put_metric_data(Namespace="MyService", MetricData=chunk)

def flush_metrics_loop(interval_seconds: int = 30):
while True:
time.sleep(interval_seconds)
try:
flush_metrics()
except Exception as e:
print(f"metric flush failed, will retry next interval: {e}")
Wire flush_metrics_loop into your app's startup (a background thread, or an async task if you're on an asyncio framework), and every request now does zero blocking network calls for metrics it just appends to an in-memory list.

The before/after math

Same 50 req/sec, 2 metrics per request, 30-second flush interval, ~10 distinct tool names:
Naive (per-request)Buffered (30s flush)
API calls/month~259.2 million~2,880 flushes/day × 30 days ≈ 86,400 calls (upper bound, assuming every flush needs multiple batched calls)
Cost/month (at $0.01/1,000, beyond free tier)~$2,580Comfortably inside the 1M free-tier requests effectively $0
Latency added per requestOne blocking network callZero an in-memory list append
That last row matters as much as the cost row. This bug is never just a billing issue every one of those synchronous calls was also sitting in the response path of a real user's request, adding real latency for zero benefit to that specific request.

What to actually check in your own service

  1. Search your codebase for put_metric_data calls. Look at what's calling them a request handler, middleware, a finally block wrapping every operation? That's your hot-path candidate.
  2. Check whether it's per-request or batched already. Some frameworks and observability libraries already batch for you (the CloudWatch agent's StatsD/collectd integration, for instance) if you're calling the raw boto3 client directly and synchronously, you're probably not batched.
  3. Estimate your real request volume and run the math above with your own numbers. It scales linearly, so even a rough estimate tells you whether this is worth fixing today or not urgent yet.
  4. If you fix it, keep one path un-batched deliberately genuinely rare, high-signal events (a hard failure, a security-relevant action) are usually still worth sending live rather than waiting for the next flush interval. Batching is for high-frequency, per-request metrics, not for everything.

Takeaway

PutMetricData's per-request pricing is cheap in isolation and expensive in aggregate the exact shape of cost bug that's easy to miss because no single call looks wrong. If your service calls it directly on a hot path, it's worth five minutes to check whether you're paying for millions of individually-billed API calls that a simple in-memory buffer and a periodic flush would have batched into a few thousand, for free.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article