
Stop Building Dumb RAG: Multi-Agent Validation on AWS
Stop building RAG that hallucinates. Multi-agent validation patterns catch mistakes before production, plus cut Bedrock costs by 90%.
Stop Building Dumb RAG: Multi-Agent Validation on AWS
Most RAG implementations are exactly that—dumb retrieval + generation. You throw a question at your vector database, grab the top 3 results, throw them at an LLM, and hope the answer makes sense. It usually doesn't.
The issue: retrieval-based systems are brittle. Bad chunk boundaries, incomplete context, or noisy embeddings mean your LLM gets fed garbage and produces garbage. You're not building intelligence — you're amplifying whatever mistakes your vector database made.
Multi-agent RAG fixes this. Instead of one pipeline doing all the work, you build specialized agents that work together: one agent validates if the retrieval is even relevant, another agent reformulates the question if the first pass fails, another agent cross-checks facts. The final answer is better because multiple checks caught problems along the way.
This is the piece most RAG tutorials skip entirely. Here's how to actually build it on AWS.
What's Actually Happening in Your RAG Pipeline
When you build RAG on AWS, you're wiring together:
- Amazon Bedrock — LLM inference, no model management
- Knowledge Bases for Bedrock — Managed vector database + retrieval (sits on top of Aurora PostgreSQL with pgvector, OpenSearch Serverless, or Pinecone)
- Lambda — Orchestration and agent logic
- EventBridge or Step Functions — Workflow coordination
The flow looks simple from a distance: user question → retrieve relevant docs → augment prompt → LLM response. But production RAG has invisible layers. Your retrieval might miss the right document. Your chunks might be too small and lack context. Your embeddings might not capture the semantic meaning of the question.
A single agent can't fix all three. Multi-agent patterns distribute the problem.
The Three-Agent Pattern That Works
Agent 1: Query Reformulation
Takes the user's question and rewrites it three ways — once literal, once rephrased for embedding similarity, once expanded with domain context. If the original question fails, one of these reformulations often succeeds. Bedrock handles the LLM call; you handle the logic that decides which reformulation to use.
Takes the user's question and rewrites it three ways — once literal, once rephrased for embedding similarity, once expanded with domain context. If the original question fails, one of these reformulations often succeeds. Bedrock handles the LLM call; you handle the logic that decides which reformulation to use.
Agent 2: Retrieval Validator
Takes the documents returned from Bedrock Knowledge Bases and scores them. Not just "did the vector database say these are relevant" — actually asks the LLM "is this document useful for answering the user's question?" Rejects low-confidence retrievals and triggers fallback retrieval with a different query. This catches the case where your embeddings lied.
Takes the documents returned from Bedrock Knowledge Bases and scores them. Not just "did the vector database say these are relevant" — actually asks the LLM "is this document useful for answering the user's question?" Rejects low-confidence retrievals and triggers fallback retrieval with a different query. This catches the case where your embeddings lied.
Agent 3: Fact Checker
Takes the LLM's answer and cross-checks facts against the source documents. Catches hallucinations. Flags statements that the source docs don't actually support. In regulated domains (finance, healthcare), this isn't optional — it's required before showing the answer to a user.
Takes the LLM's answer and cross-checks facts against the source documents. Catches hallucinations. Flags statements that the source docs don't actually support. In regulated domains (finance, healthcare), this isn't optional — it's required before showing the answer to a user.
All three agents run on Lambda with Bedrock API calls. EventBridge orchestrates the flow: if validator rejects, trigger reformulation agent; if fact checker flags hallucination, return "I'm not confident" instead of a made-up answer.
The Bedrock Knowledge Base Cost Gotcha
Here's where most people get surprised. Bedrock Knowledge Bases manages your vector database for you, but the cost depends on which backend you pick:
OpenSearch Serverless — AWS's default recommendation. You get automatic scaling. You also get a bill surprise: $2.40 per Ingestion Data Unit (1 million tokens) and $0.30 per Query Data Unit. For a knowledge base with 100 documents (roughly 500K tokens ingested), querying once costs $0.15. At scale — hundreds of queries a day — this adds up fast.
Aurora Serverless v2 with pgvector — Postgres extension that handles vectors natively. Same retrieval quality, fraction of the cost. A small Aurora cluster (db.serverless) costs ~$0.50/hour when paused, scales up on traffic. For the same 100-document setup, you pay for database compute time, not per-query fees. One analysis showed a 90% cost reduction switching from OpenSearch to Aurora.
If you're building a prototype, OpenSearch Serverless is fine. If you're building for production or scaling, Aurora with pgvector is the move.
Why Multi-Agent Isn't Optional at Scale
Single-stage RAG works fine on toy datasets — when you have 10 well-curated documents. Production datasets are messy. Documents overlap, contain contradictions, or use domain jargon your embeddings don't understand. A multi-agent system catches these issues before they reach your users.
Also, LLMs are expensive on Bedrock. The more bad retrievals you feed to your LLM, the more token-wasting you do. Multi-agent validation cuts bad retrievals before they hit the LLM, saving money and improving latency.
Setting One Up
High-level flow:
- Set up Bedrock Knowledge Base with Aurora Serverless v2 + pgvector as the vector store (Quick Create option)
- Create three Lambda functions — one per agent (reformulator, validator, fact-checker)
- Wire them with EventBridge or Step Functions — reformulator runs first, on failure triggers validator, fact-checker runs last
- Add a simple API Gateway endpoint that users call with their question
Console path for Knowledge Base setup:
- Bedrock → Knowledge bases → Create knowledge base
- Select "Aurora PostgreSQL with pgvector" as vector store
- Upload your documents (S3 bucket or direct upload)
- Configure chunking: 512 token chunks, 100 token overlap (prevents context loss at chunk boundaries)
- Enable data synchronization for live doc updates
For the agents, use Bedrock API with function calling. Bedrock can call your Lambda functions directly if you set up the tool definitions right. Reformulator agent: invoke Bedrock with a system prompt asking for three rephrased versions. Validator agent: invoke Bedrock with "Rate the relevance of these retrieved docs from 1-10" and reject anything below 7. Fact-checker agent: invoke Bedrock with "List facts from the answer and mark which ones are supported by these source docs."
CLI example for querying Bedrock Knowledge Base:
1
2
3
4
aws bedrock-agent-runtime retrieve-and-generate \
--knowledge-base-id <kb-id> \
--input text="What's the company's refund policy?" \
--generation-configuration generationConfig={temperature=0.5,maxTokens=1024}
For multi-agent orchestration, Step Functions is cleaner than EventBridge. Define a state machine:
1
2
3
Start → Reformulate Query → Retrieve Docs → Validate Retrieval
├─ Valid? → Generate Answer → Fact-Check → Return
└─ Invalid? → Try Reformulated Query
Put Together
Article one covered CloudWatch alarms. Article two covered billing. Article three covered proactive monitoring. This is the fourth piece: building intelligent document systems that actually scale.
Single-agent RAG is a starting point. Multi-agent RAG is production-ready. The difference is catching mistakes before they become problems — exact same principle as alarms catching runaway Lambda costs before they hit your bill.
Build this on your AWS account with one document set first, measure latency and cost, then scale. The architecture holds.
Links & Sources
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article