
π Reliable SaaS Workflows with Amazon SQS
Build reliable SaaS workflows with Amazon SQS using idempotency, transactional outboxes, partial batch responses, and recovery strategies.
Series: Building SaaS on AWS: From Architecture to Operations (3 articles)
- 3π Reliable SaaS Workflows with Amazon SQS This article
Reliable SaaS Workflows with Amazon SQS: Design for the Second Attempt
A customer clicks βGenerate document.β The request times out. They click again. Meanwhile, a worker has already generated the first file but has not recorded its success.
The difficult question is not whether the system can retry. It is whether the second attempt produces the correct business outcome.
This article uses a fictional document portal to propose a reliable workflow built around explicit state, stable identifiers, and recoverable failures.
Separate acceptance from completion
An HTTP response should describe what actually happened. A
202 Accepted response can mean the system has durably accepted work; it should not imply that the document already exists.Return a stable job identifier and a status location. The interface can then show progress and let a returning customer find the same operation. Define terminal outcomes such as
COMPLETED and REQUIRES_ATTENTION, as well as nonterminal states such as QUEUED and PROCESSING.For this portal, the job record is the source of truth for the user experience. Queue delivery is an implementation mechanism, not a user-visible status model.
Close the database-to-queue gap
Writing a job record and then sending an SQS message creates a failure window: the first write may succeed while the second fails. A transactional outbox addresses this by committing the job and an event record in one database transaction. A relay later publishes committed events to the queue. The relay may publish duplicates, so consumers still need idempotency. AWS guidance: transactional outboxΒ
For our design, the outbox event contains an event ID, tenant ID, job ID, schema version, and creation time. Keep large documents out of the message; reference controlled storage instead.
Monitor the oldest unpublished outbox entry. A healthy queue cannot tell you that the relay stopped publishing yesterday.
Make the business operation idempotent
Lambda's SQS integration can process records more than once. A successful path must therefore remain safe when repeated. AWS documentation: Lambda with SQSΒ
Choose an operation key based on business intent, such as tenant plus document request ID. An SQS message ID alone is insufficient when two messages represent the same request.
For a database-only effect, a unique constraint and transaction may be enough. For an external document provider, the problem is harder: the provider can complete the operation while your worker loses the response. Use the provider's idempotency key or lookup mechanism when available. Otherwise, mark the outcome as uncertain and reconcile it before attempting another irreversible action.
An
IN_PROGRESS flag also needs recovery rules. Record an owner or lease and an expiry, and define how a later worker determines whether it may resume. Never make a crashed worker's flag a permanent lock.Retry individual failures
With SQS batch processing, configure
ReportBatchItemFailures and return only the failed message identifiers. An unhandled function-level exception can still fail the entire batch. FIFO processing needs special care: stop after a failure and report failed and unprocessed records to preserve ordering. AWS documentation: partial batch responsesΒ Set the visibility timeout around the function's execution and batching configuration. AWS recommends at least six times the function timeout, plus the batching window when used. A redrive policy moves repeatedly unsuccessful messages to a dead-letter queue after the configured receive count. AWS documentation: SQS event source configurationΒ
Those settings support recovery; they do not decide whether a business action is safe to repeat.
Operate the failure path
For this portal, I would track end-to-end completion time, the oldest queued job, duplicate suppression, outbox delay, and dead-letter volume. Include tenant, job, and correlation identifiers in structured logs without logging customer document contents.
A dead-letter queue needs an owner and a replay procedure. Investigate the failure, fix its cause, verify that replay is safe, and monitor the replay. Moving messages back blindly can repeat a harmful side effect or recreate the same backlog.
Test four disruptions before calling the workflow reliable: duplicate delivery, failure after a durable write, temporary provider unavailability, and a malformed message mixed into a valid batch. Check the resulting business state, not just whether the function eventually returned success.
The target is one correct outcome per logical request, even when execution takes more than one attempt. The next article turns a small part of this model into a deployable AWS SAM exercise.
Series: Building SaaS on AWS: From Architecture to Operations (3 articles)
- 3π Reliable SaaS Workflows with Amazon SQS This article
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article