AWS Builder Center

How We Modernized RDS and ElastiCache—and Cut ~$450/Day in AWS Costs

How we turned AWS cost optimization into an engineering modernization initiative—upgrading RDS, moving Redis OSS workloads to Valkey, adopting newer instance generations, and right-sizing based on real usage. The result: ~$450/day in recurring savings without compromising reliability.

Cloud cost optimization is often treated as a finance exercise.
For us, it became an engineering modernization project.
We started with a relatively simple question:
How much are we paying for infrastructure that made sense when it was created—but no longer reflects our workloads today?
The answer led us to review a large part of our Amazon RDS and Amazon ElastiCache estate.
The result?
Approximately $450/day in recurring AWS savings — while simultaneously moving to newer infrastructure, supported database versions, and better-sized workloads.
That is roughly $13,500/month or $164,000/year if the savings remain consistent.
But the most interesting part wasn’t the number.
It was how we got there.
Step 1: Build the Inventory Before Optimizing Anything
Our first step wasn’t resizing instances.
It was visibility.
Across a large AWS environment, infrastructure naturally becomes heterogeneous over time. Different teams create databases at different moments, workloads evolve, traffic patterns change, and an instance that was correctly sized two years ago may be significantly oversized today.
We therefore started by identifying:
  • RDS and Aurora database instances
  • ElastiCache clusters and replication groups
  • Instance families and generations
  • Engine versions
  • CPU and memory utilization
  • Storage and I/O characteristics
  • High-availability requirements
  • End-of-support exposure
  • Potential right-sizing opportunities
This gave us something extremely important:
a prioritized modernization backlog instead of a random list of cost-saving ideas.
Step 2: Modernize the Compute Generation
One of the clearest opportunities was moving workloads away from older instance generations.
Where workload compatibility allowed it, we moved RDS and ElastiCache workloads from older generations toward newer AWS Graviton-based families.
Instead of looking only at:
“Can we use a smaller instance?”
we asked:
“Can we run this workload more efficiently?”
That distinction matters.
A newer instance generation can provide better price/performance without necessarily reducing the resources available to the application.
We therefore combined two operations wherever practical:
generation upgrade + right-sizing
For example, rather than simply replacing an older large instance with the equivalent size in a newer family, we looked at actual utilization and determined whether the workload could safely move down a size as well.
Step 3: Treat Database Upgrades as Part of FinOps
Cost optimization and lifecycle management are often managed as separate projects.
We decided to combine them.
Some RDS workloads were approaching—or already using—database versions where Extended Support had become relevant.
Instead of paying indefinitely for technical debt, we used the modernization initiative to push database upgrades as well.
For MySQL workloads, this meant moving applications away from older versions and toward supported releases.
For higher-risk databases, we used RDS Blue/Green Deployments where appropriate.
The pattern was roughly:
Current production → Blue/Green environment → Upgrade → Validate → Switchover
This gave application teams an opportunity to validate compatibility before the final cutover and reduced the operational risk compared with treating every major database upgrade as an in-place change.
The financial benefit here is easy to underestimate.
Avoiding unnecessary Extended Support costs is itself a FinOps optimization.
But you also remove technical debt at the same time.
Step 4: Redis OSS → Valkey
ElastiCache became another major part of the initiative.
We started migrating suitable Amazon ElastiCache for Redis OSS workloads to Valkey.
AWS supports cross-engine upgrades from Redis OSS to Valkey, and Valkey is designed as a drop-in replacement for Redis OSS 7. For supported versions, AWS performs the upgrade while retaining the application endpoint; node IP addresses can change during the operation. (AWS Documentation ⁠)
That last detail leads to an important lesson.
Test client reconnection behavior.
Infrastructure may handle an upgrade gracefully while the application does not.
During cache maintenance, node replacement, failover, or engine upgrades, existing client connections can be terminated. AWS explicitly recommends that Redis/Valkey clients implement retry behavior and exponential backoff. (AWS Documentation ⁠)
We found this particularly important for some application stacks where connection handling wasn’t behaving as expected after node replacement.
So before aggressively upgrading or right-sizing caches, verify:
  • connection retry behavior
  • DNS resolution behavior
  • connection pool recovery
  • failover handling
  • application timeouts
A cache engine upgrade should not become an application incident because the client assumes that a connection lives forever.
Step 5: Right-Size Using Real Workload Data
This was probably the most important principle of the entire project:
Do not right-size from instance specifications. Right-size from workload behavior.
For each candidate, we looked at historical utilization rather than a short snapshot.
Typical signals included:
  • CPU utilization
  • memory pressure
  • database connections
  • freeable memory
  • read/write IOPS
  • throughput
  • replica behavior
  • cache memory utilization
  • evictions
  • latency
  • traffic patterns
Peak events matter too.
An instance running at 20% CPU today might be intentionally sized for a known traffic event.
Our environment, for example, has workloads where Black Friday / Cyber Monday traffic is far more important than an average Tuesday.
So the question wasn’t:
“Can this instance survive today on a smaller size?”
It was:
“Can this instance safely support the workload we expect it to handle?”
That’s a very different FinOps conversation.
Step 6: Don’t Forget Architecture-Level Optimization
Some of the best savings didn’t come from changing an instance type.
They came from asking whether the architecture still required the same amount of infrastructure.
For clustered caches, for example, application optimizations can reduce memory requirements enough to reconsider the number of shards.
That creates another optimization dimension:
Instance size × instance generation × number of nodes
The same principle applies to databases.
Sometimes the right answer isn’t a smaller database.
It might be:
  • reducing unnecessary replicas
  • changing storage configuration
  • tuning expensive I/O behavior
  • removing unused resources
  • consolidating workloads
  • changing purchasing commitments
FinOps becomes much more powerful when engineers stop treating infrastructure configuration as immutable.
Step 7: Make Savings Measurable
We tracked the impact of each modernization activity.
Over time, the cumulative savings grew.
$50/day doesn’t sound transformational.
Neither does another $30/day.
Or another $20/day.
But infrastructure optimization compounds.
Eventually, our modernization work reached approximately:
$450/day in recurring AWS savings.
That’s approximately:
$13,500/month
and
$164,000/year
assuming similar utilization.
More importantly, these weren’t savings created by simply shutting down useful infrastructure.
We were simultaneously improving the platform:
Older infrastructure → newer generations
Oversized workloads → right-sized workloads
Older database engines → supported versions
Redis OSS → Valkey
Technical debt → modernization
That is the kind of cost optimization I like most.
What I Learned
After working through this initiative across multiple workloads and teams, a few lessons stand out.
1. FinOps should be an engineering activity
Finance can identify where money is being spent.
Engineers understand why it is being spent.
The biggest opportunities appear when those two perspectives meet.
2. Combine modernization with optimization
If you’re already touching an RDS instance, ask:
Can we upgrade the engine?
Can we move to a newer instance generation?
Can we right-size it?
Can we improve the storage configuration?
Can we remove Extended Support exposure?
One maintenance activity can potentially eliminate several pieces of technical debt.
3. Optimize continuously
Cloud infrastructure changes constantly.
Today’s perfectly sized instance can become tomorrow’s oversized instance.
Right-sizing shouldn’t be a yearly project.
It should be part of normal platform operations.
4. Measure savings per change
A small optimization becomes much easier to justify when you can say:
This change saves approximately $X/day.
It also makes the cumulative impact of engineering work visible to the organization.
5. Never optimize reliability away
The cheapest infrastructure is useless if it can’t handle production traffic.
Historical utilization, expected growth, failure scenarios, seasonal traffic, and major events must remain part of the decision.
Our objective wasn’t:
“Make AWS cheaper.”
It was:
“Run our workloads more efficiently without compromising reliability.”
The Bigger Picture
The ~$450/day number is great.
But I think the more important outcome was changing how we looked at infrastructure.
Cost optimization stopped being:
“Find something expensive and make it smaller.”
Instead, it became:
Observe → Modernize → Right-size → Validate → Measure → Repeat
And that is something any organization running AWS infrastructure at scale can start doing.
You don’t need a huge FinOps program to begin.
Pick ten RDS instances.
Pick ten ElastiCache clusters.
Look at their generation, engine version, utilization, architecture, and support lifecycle.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article