AWS Builder Center
AWS DynamoDB Deep Dive: NoSQL Fundamentals, Capacity Math, and Everything In Between

AWS DynamoDB Deep Dive: NoSQL Fundamentals, Capacity Math, and Everything In Between

NoSQL databases were built to solve exactly this. The full form is "Not Only SQL," meaning the world does not revolve around relational databases alone. NoSQL databases are non-relational, support flexible schemas, and scale horizontally by adding more nodes rather than buying bigger hardware.

SQL vs NoSQL at a Glance

FeatureSQLNoSQL
Data structureStructuredFlexible / Unstructured
SchemaStaticDynamic
ScalingVerticalHorizontal
JoinsSupportedNot supported
Use caseComplex queriesHigh performance at scale
ExamplesMySQL, PostgreSQL, OracleMongoDB, Redis, Cassandra
SQL databases shine when you need complex queries with joins, aggregations, and subqueries. NoSQL databases win when performance and scalability are the priority and your data structure may evolve over time.

Types of NoSQL Databases

NoSQL is not a single category. There are four main types:
Key-Value stores data as simple key-value pairs. Redis is the most well-known example.
Document stores data as JSON-like documents with flexible fields. MongoDB is the classic example.
Columnar databases store data by columns rather than rows, which makes them extremely efficient for analytics. Apache Cassandra, HBase, and Amazon Redshift fall into this category.
Graph databases model data as nodes and relationships, useful for social networks or anything where connections matter. Neo4j is the primary example.
AWS recognized this landscape and built its own managed NoSQL service: DynamoDB. It supports both document and key-value data models.

DynamoDB: No Instances, No Database, Just Tables

Here is the most important thing to understand about DynamoDB before anything else: it is serverless. There is no database engine to set up, no instance to manage, no operating system to patch.
In RDS, you create an instance first, then create a database on that instance, and then create tables inside that database. Three layers.
In DynamoDB, you skip straight to the table. No instance. No database. You go to the console, create a table, and you are done. That is why developers love it.
The performance also speaks for itself. DynamoDB delivers single-digit millisecond latency. Not one second. Not 100 milliseconds. Between 1 and 9 milliseconds, consistently, even at massive scale. Trillions of records, terabytes of data, and DynamoDB handles it without breaking a sweat.

DynamoDB Terminology: What Everything Is Called

DynamoDB uses different names for things you already know from SQL. Here is the mapping:
SQL TermDynamoDB Term
TableTable
ColumnAttribute
RowItem
Primary KeyPartition Key
A table holds all the data. Each row is an item. Each column is an attribute. And the primary key is called the partition key.

The Partition Key

The partition key is the one column you must define when creating a DynamoDB table. It must be unique and it cannot be null. DynamoDB also refers to it as the hash attribute.
When you create a table in DynamoDB, you start with just one attribute: the partition key. You do not need to define every column upfront. Other attributes can be added at runtime, by the application or manually. This is the dynamic schema in practice.

Sort Key and Composite Keys

Sometimes you need to allow duplicate partition key values in a table. By default, that is not possible since the partition key must be unique. But with a sort key, you can do it.
A composite key is the combination of partition key and sort key. With a composite key, the partition key alone does not need to be unique. What needs to be unique is the combination of partition key and sort key together.
Here is a practical example. Consider a games table:
PlayerName (Partition Key)GameName (Sort Key)Result
AlexPUBGWon
AlexCODLost
Alex appears twice. Normally that would violate the uniqueness constraint on the partition key. But because GameName is the sort key, the combination of Alex + PUBG and Alex + COD are both unique. This is valid.
The sort key is also called the range attribute in DynamoDB terminology. So when developers talk about hash attribute and range attribute, they mean partition key and sort key respectively.
The sort key is optional. If you have a simple table where partition key values are always unique, you do not need it.

DynamoDB Streams: Tracking Changes to Your Data

Imagine you update a record, changing a name from "Rashmita" to "Rashmi." Run a SELECT * query now and you only see "Rashmi." The previous value is gone unless you have transaction logs configured.
DynamoDB Streams solves this by capturing a change log of every insert, update, and delete on a table. Each stream record contains the old and new values of the item that was modified. This lets you track what changed and when.
DynamoDB Streams can work alongside Amazon Kinesis Data Streams for more advanced real-time processing pipelines, though they are separate features with different capabilities.

Backups and Point-In-Time Recovery

DynamoDB supports two backup approaches.
Manual backups can be created on demand at any time through the console or API.
PITR (Point-In-Time Recovery) is a continuous backup feature managed entirely by AWS. When enabled, it allows you to restore your table to any point in time within the retention window, down to the second. No scripts, no cron jobs, no manual effort.

Consistency Models: ECR and SCR

When you write data to DynamoDB, it is replicated across multiple nodes. The question is: when you read immediately after a write, do you get the latest data?
DynamoDB gives you two read consistency options.
Eventually Consistent Read (ECR) is the default. When you read data, you might get a slightly older version for a brief moment as replication completes. It is faster and cheaper.
Strongly Consistent Read (SCR) always returns the most up-to-date data. It waits for all replicas to be consistent before returning. Slightly more expensive in terms of capacity units.
The choice matters when you are calculating costs and designing read-heavy applications.

Capacity Modes

DynamoDB offers two modes for handling throughput.
Provisioned Capacity Mode requires you to specify how many Read Capacity Units (RCU) and Write Capacity Units (WCU) your table needs. AWS provisions that capacity and your table performs within those limits. Good for predictable workloads.
On-Demand Capacity Mode lets AWS handle scaling automatically. You do not specify RCU or WCU. AWS adjusts capacity based on actual traffic. Better for unpredictable or spiky workloads.

RCU and WCU: Formulas and Calculations

This is where the math comes in and it is simpler than it looks.

The Formulas

Read Capacity Units (RCU):
  • Eventually Consistent Read (ECR): 1 RCU = 2 reads per second for items up to 4 KB
  • Strongly Consistent Read (SCR): 1 RCU = 1 read per second for items up to 4 KB
Write Capacity Units (WCU):
  • 1 WCU = 1 write per second for items up to 1 KB

How to Calculate

For reads, the formula is:
1
RCU required = (Item size in KB / 4 KB) × reads per second / consistency divisor
Where the consistency divisor is 1 for SCR and 2 for ECR.
For writes, the formula is:
1
WCU required = (Item size in KB / 1 KB) × writes per second

Worked Examples

Example 1: Read Capacity (SCR)
You need to read 100 items per second. Each item is 8 KB. You are using Strongly Consistent Reads.
1
2
3
Step 1: 8 KB / 4 KB = 2
Step 2: 2 × 100 reads/sec = 200
Step 3: 200 / 1 (SCR divisor) = 200 RCUs required
Example 1 continued: Same scenario with ECR
1
2
3
Step 1: 8 KB / 4 KB = 2
Step 2: 2 × 100 reads/sec = 200
Step 3: 200 / 2 (ECR divisor) = 100 RCUs required
ECR costs half the RCUs of SCR for the same read volume.
Example 2: Write Capacity
You need to write 50 items per second. Each item is 4 KB.
1
2
Step 1: 4 KB / 1 KB = 4
Step 2: 4 × 50 writes/sec = 200 WCUs required

What Is the Default Capacity?

When you create a DynamoDB table in provisioned mode, AWS sets the default capacity at 5 RCUs and 5 WCUs. To put that in perspective, 1 RCU gives you roughly 5.2 million eventually consistent reads per month. AWS gives you 5 of those by default, which is more than enough for getting started.
If you need more, you increase the RCU and WCU numbers. That is what push-button scaling means. No instance restart, no downtime. Just update the values and DynamoDB scales accordingly.

Indexes: LSI and GSI

Indexes in DynamoDB work on the same principle as an index in a textbook. Without an index, finding a specific chapter means flipping through every single page. With an index, you jump straight to the right page. In database terms, an index means faster reads.
DynamoDB provides two types of indexes.

Local Secondary Index (LSI)

An LSI uses the same partition key as the base table but allows a different sort key. This lets you query data in a different order or by a different secondary attribute while still scoping to the same partition.
Critical rule: LSIs must be created at the time of table creation. You cannot add an LSI to an existing table. Once created, they cannot be modified or deleted.

Global Secondary Index (GSI)

A GSI allows an entirely different partition key and sort key combination. It is essentially a separate view of your table data organized around different keys.
GSIs are flexible: you can create them at any time, modify them, and delete them. There is no restriction on when they are created.

LSI vs GSI Summary

FeatureLSIGSI
Partition KeySame as base tableAny attribute
Sort KeyDifferent attributeAny attribute
Creation timeOnly at table creationAnytime
ModifiableNoYes

Scan vs Query: Always Prefer Query

Scan reads every single item in a table and filters afterward. On a table with a million rows, a scan reads all million rows, consumes massive RCUs, and returns slowly. Avoid it in production wherever possible.
Query uses the partition key (and optionally the sort key or indexes) to retrieve only the specific items you need. It is targeted, fast, and consumes far fewer RCUs.
Think of scan as SELECT * FROM table and query as SELECT * FROM table WHERE employee_id = 1000. One returns one record. The other returns everything. The difference in cost and performance is significant.

TTL: Automatic Item Expiration

TTL stands for Time to Live. It lets you set an expiration timestamp on individual items. When the timestamp is reached, DynamoDB automatically deletes that item from the table. No manual deletion required, no scheduled jobs.
A classic real-world example is a subscription service. When a user pays for a monthly subscription, a record is added to the table with a TTL set to 30 days from now. When the 30 days are up, the record is automatically deleted. When the user tries to access the service, their record no longer exists, so access is denied. The entire expiration is handled by DynamoDB with zero application code needed for cleanup.

DAX: When DynamoDB Needs Even More Speed

DynamoDB already delivers single-digit millisecond latency. If that is not fast enough for your use case, DynamoDB Accelerator (DAX) is an in-memory caching layer that sits in front of DynamoDB and delivers microsecond response times for read-heavy workloads. You attach it to your DynamoDB table and reads that hit the cache return almost instantly.

Global Tables

DynamoDB is a regional service. A table created in Mumbai lives in Mumbai. If you want that same table available in another region, for disaster recovery or for users in different geographies, you use Global Tables.
Global Tables provide multi-region, multi-active replication. Data written in one region is automatically replicated to all other configured regions.

Summary

DynamoDB is serverless, schema-flexible, and built for performance at scale. Here is what to keep in your head:
You go straight to creating a table with one mandatory attribute: the partition key. Sort key is optional but enables duplicate partition key values when the composite combination remains unique.
Reads have two consistency modes: ECR by default (faster, cheaper) and SCR (always fresh, costs more). Capacity is measured in RCU and WCU. The formulas are straightforward once you know that reads are based on 4 KB chunks and writes on 1 KB chunks.
LSIs must be created when the table is created. GSIs can be created anytime. Always prefer query over scan. Use TTL to automatically expire items without any cleanup code.
And when DynamoDB's millisecond latency is not fast enough, DAX brings it down to microseconds.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article