
The Race Condition Hiding Inside My Perfect Architecture
A single Lambda function, a working database, and no errors in the logs — yet two users still claimed the same last item. Here's what a race condition actually looks like in production, and how a conditional write fixes it.
The Race Condition Hiding Inside My Perfect Architecture
Imagine there is only one item left. Two users click "Claim" at almost exactly the same time The API is working, Lambda is working, and the database is working but somehow, both users get the item.
Nothing crashed. There is no obvious error in the logs. Every individual component is doing what it was designed to do.
So what went wrong?
This is how I came across one of those problems that is easy to miss when building cloud applications: a race condition.
It Looked Simple at First
The architecture itself was straight forward:
User
↓
API Gateway
↓
Lambda
↓
DynamoDB
When a request arrived, Lambda would read the current stock, check whether something was available, decrease the count, and save the update value. For a single request, everything worked exactly as expected.
The problem appeared when multiple requests started arriving at the same time.
Suppose the database contains just one available item. User A sends a request, and Lambda reads stock = 1. Before that request finishes updating the database, User B sends another request. Lambda B also reads stock = 1.
Both requests now believe they are allowed to continue.
The original flow looked like this:
Read stock
↓
Check stock > 0
↓
Decrease stock
↓
Save stock
↓
Check stock > 0
↓
Decrease stock
↓
Save stock
The issue is that the check and the update are separate operations. There is a small window between them another request can read the same value.
That small window is enough to create a race condition.
The Race was not Lambda
My first instinct was to look at the Lambda code. But the deeper problem was not that Lambda was doing something incorrectly. The problem was where the decision about stock availability was being made.
The application was effectively saying:
"I checked the stock and there is one available, so I will update it not."
But another request could make the same decision before the first update completed.
What we actually need is for the condition and the update to happen together:
1
2
Update stock
only if stock > 0This is where a conditional write in DynamoDB becomes useful. Instead of Lambda checking the value first and then performing a separate update, the database enforces the condition as part of the update itself:
1
2
3
4
UpdateItem
Key: { itemId }
UpdateExpression: SET stock = stock - :one
ConditionExpression: stock > :zeroNow, if two requests compete for the final item, only the request whose condition is satisfied at the time of the time of the write can succeed. DynamoDB rejects the other one outright the request fails cleanly instead of incorrectly claiming an item that no longer exists.
The important change was not adding another AWS service or making the Lambda more complicated. It was moving the correctness rule to the place that controls the data.
Why This Problem Is Bigger Than Inventory.
Once I understood the issue, I started noticing the same pattern in many other systems.
Think about the last available concert ticket, a limited-use coupon, a seat reservation, an account balance, or a job that should only be processed once. In all of these cases, multiple requests may be trying to change the same piece of state.
The question is always similar:
What happens if two requests try to change the same thing at exactly the same time?
That's a question worth asking early, because normal testing often won't reveal the problem.
Testing the System Under Pressure
A single successful API request doesn't tell us much about concurrency.
If I send requests one after another, each request gets a chance to finish before the next one starts. The race condition may never appear.
The interesting test is what happens when many requests arrive together. If the initial stock is 100 and 1,000 requests arrive concurrently, the system should never allow more than 100 successful claims. The final stock should never become negative, and rejected requests should not be treated as successful claims.
When I ran this under concurrent load, the conditional write held up exactly as expected: exactly 100 requests succeeded, the remaining 900 failed cleanly with a condition-check error, and stock never dropped below zero. No double-claims, no silent corruption just a clean split between requests that got in before the count hit zero and requests that didn't.
This is where load testing becomes more than a performance exercise. It's also a way of testing correctness under concurrency.
What I'd Ask Before Shipping
Now, whenever I build something that modifies shared data, I ask a few simple questions:
What happens if two requests arrive at the same time? If the answer relies on "check, then update" as two separate steps, that gap is where the race lives. The condition needs to live inside the write, not before it.
What happens if the same request is repeated? This is where idempotency keys earn their keep. Without one, a retried request looks identical to a brand-new one, and the system has no way to tell them apart.
What happens if a request succeeds but the client never receives the response? The client will likely retry. If the server-side operation isn't idempotent, that retry can double-charge, double-claim, or double process something that already happened.
What happens if the operation is retried by infrastructure, not the user? Lambda, queues, and API Gateway all retry under the hood during throttling or timeouts. The system needs to behave correctly even when nobody clicked twice.
These questions often reveal problems that aren't visible when everything happens in the expected order.
The Bigger Lesson
The interesting thing about this race condition is that nothing was actually "broken." The API worked. Lambda worked. DynamoDB worked. The application worked when tested normally.
The problem appeared because the system had to deal with something that wasn't obvious in the original design: multiple things happening at once. An architecture diagram doesn't show timing on paper it's just
API → Lambda → DynamoDB, but in reality hundreds of requests can move through that same path simultaneously, competing for the exact same database record while looking completely independent in the diagram.If there's a rule that must never be violated
stock >= 0, in this case that rule has to be enforced as close to the data as possible, not somewhere upstream where another request can slip through the gap.A system shouldn't only work when everything happens one request at a time. It should also know what to do when everything happens at once.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article