
Building a RAG Pipeline on AWS with Amazon Bedrock, pgvector, and FastAPI
Learn how to build a retrieval-augmented generation (RAG) pipeline on AWS from scratch. This step-by-step guide uses Amazon S3 for documents, Amazon Bedrock for embeddings and answers, PostgreSQL with pgvector as the vector store, and FastAPI to expose it as an API, with code you can run and adapt.
Large language models are good at language, but they don't know your documents. Retrieval-Augmented Generation (RAG) fixes this by fetching relevant passages from your own data at query time and handing them to the model as context. In this post, I'll build a small but complete RAG backend on AWS using:
- Amazon S3 for document storage
- Amazon Bedrock for embeddings and answer generation
- PostgreSQL with pgvector (Amazon RDS) as the vector store
- FastAPI as the API layer
Architecture
The pipeline has two flows.
Ingestion (offline): PDF in S3 → extract text → split into chunks → embed each chunk with Bedrock → store text and vectors in Postgres.
Query (online): User question → embed the question → vector similarity search in Postgres → build a prompt from the top chunks → generate the answer with a Bedrock model → return the answer with its sources.
I chose pgvector because many teams already run Postgres, it keeps vectors next to relational metadata, and it avoids operating another service. If you outgrow it, the vector store is a small, replaceable layer (OpenSearch Serverless and S3 Vectors are alternatives worth evaluating).
Prerequisites
- An AWS account with Bedrock model access enabled for an embedding model and a text generation model in your region. Model availability varies by region, so check the Bedrock console first.
- An RDS for PostgreSQL instance with the
vectorextension available. - Python 3.11+ and these packages:
1
pip install fastapi uvicorn boto3 psycopg[binary] pypdf python-multipart
- IAM permissions for
bedrock:InvokeModel, plus S3 read access to your bucket. Use an IAM role rather than hardcoded keys.
Step 1: Set up the vector table
1
2
3
4
5
6
7
8
9
10
11
12
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE IF NOT EXISTS chunks (
id BIGSERIAL PRIMARY KEY,
source TEXT NOT NULL,
chunk_idx INT NOT NULL,
content TEXT NOT NULL,
embedding VECTOR(1024)
);
CREATE INDEX IF NOT EXISTS chunks_embedding_idx
ON chunks USING hnsw (embedding vector_cosine_ops);
The vector dimension must match your embedding model's output. I'm using 1024 here, which matches Amazon Titan Text Embeddings V2's default.
Step 2: Embeddings with Bedrock
1
2
3
4
5
6
7
8
9
10
11
12
13
# bedrock_client.py
import json
import boto3
REGION = "us-east-1" # use a region where your models are enabled
EMBED_MODEL = "amazon.titan-embed-text-v2:0"
bedrock = boto3.client("bedrock-runtime", region_name=REGION)
def embed(text: str) -> list[float]:
body = json.dumps({"inputText": text, "dimensions": 1024, "normalize": True})
resp = bedrock.invoke_model(modelId=EMBED_MODEL, body=body)
return json.loads(resp["body"].read())["embedding"]
Step 3: Ingest and chunk documents
Chunking quality has an outsized effect on answer quality. A simple, reliable start is fixed-size chunks with overlap, so sentences cut at a boundary still appear whole in a neighboring chunk.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
# ingest.py
import io
import boto3
import psycopg
from pypdf import PdfReader
from bedrock_client import embed
s3 = boto3.client("s3")
DB_URL = "postgresql://user:password@your-rds-endpoint:5432/ragdb" # load from Secrets Manager in production
def chunk_text(text: str, size: int = 800, overlap: int = 150) -> list[str]:
chunks, start = [], 0
while start < len(text):
chunks.append(text[start:start + size])
start += size - overlap
return chunks
def ingest_pdf(bucket: str, key: str) -> int:
obj = s3.get_object(Bucket=bucket, Key=key)
reader = PdfReader(io.BytesIO(obj["Body"].read()))
text = "\n".join(page.extract_text() or "" for page in reader.pages)
chunks = chunk_text(text)
with psycopg.connect(DB_URL) as conn:
for i, content in enumerate(chunks):
conn.execute(
"INSERT INTO chunks (source, chunk_idx, content, embedding) "
"VALUES (%s, %s, %s, %s::vector)",
(key, i, content, str(embed(content))),
)
return len(chunks)
Step 4: Retrieve and generate
For generation, I use the Bedrock Converse API, which gives one consistent request format across supported models, so swapping models is a one-line change.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
# rag.py
import psycopg
from bedrock_client import bedrock, embed
from ingest import DB_URL
GEN_MODEL = "your-bedrock-model-id" # a text model enabled in your account
def retrieve(question: str, k: int = 4) -> list[dict]:
q_vec = str(embed(question))
with psycopg.connect(DB_URL) as conn:
rows = conn.execute(
"SELECT source, chunk_idx, content, "
"1 - (embedding <=> %s::vector) AS score "
"FROM chunks ORDER BY embedding <=> %s::vector LIMIT %s",
(q_vec, q_vec, k),
).fetchall()
return [{"source": r[0], "idx": r[1], "content": r[2], "score": float(r[3])} for r in rows]
def answer(question: str) -> dict:
hits = retrieve(question)
context = "\n\n".join(f"[{i+1}] {h['content']}" for i, h in enumerate(hits))
system = [{"text": (
"Answer using only the provided context. "
"If the context doesn't contain the answer, say you don't know. "
"Cite sources as [1], [2], etc."
)}]
messages = [{
"role": "user",
"content": [{"text": f"Context:\n{context}\n\nQuestion: {question}"}],
}]
resp = bedrock.converse(
modelId=GEN_MODEL,
system=system,
messages=messages,
inferenceConfig={"maxTokens": 600, "temperature": 0.2},
)
text = resp["output"]["message"]["content"][0]["text"]
return {"answer": text, "sources": hits}
The system prompt does the most important job here: it tells the model to stay grounded in the retrieved context and admit when the answer isn't there. That single instruction cuts down on hallucinated answers more than most tuning does.
Step 5: Expose it with FastAPI
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
# main.py
from fastapi import FastAPI
from pydantic import BaseModel
from ingest import ingest_pdf
from rag import answer
app = FastAPI(title="RAG on AWS")
class IngestRequest(BaseModel):
bucket: str
key: str
class AskRequest(BaseModel):
question: str
def ingest(req: IngestRequest):
return {"chunks_indexed": ingest_pdf(req.bucket, req.key)}
def ask(req: AskRequest):
return answer(req.question)
Run it locally with
uvicorn main:app --reload, then try:1
2
3
curl -X POST localhost:8000/ask \
-H "Content-Type: application/json" \
-d '{"question": "What is the refund policy?"}'
Deploying
For a first deployment, containerize the app and run it on AWS App Runner or ECS Fargate, attach an IAM role for Bedrock and S3 access, and place RDS in private subnets reachable only from the service. Store the database credentials in AWS Secrets Manager rather than environment files. For spiky, low-traffic workloads, Lambda behind API Gateway also works well, though you'll want to manage database connections carefully (RDS Proxy helps).
Lessons learned and next steps
- Retrieval quality is the ceiling. If the right chunk isn't retrieved, no model can answer correctly. Log the retrieved chunks and scores while you develop.
- Tune chunk size and overlap on your own documents. There is no universal best value.
- Add metadata filtering. Storing fields like document type or date lets you narrow searches before the vector comparison.
- Evaluate with a test set. Build 20 to 30 real questions with known answers and rerun them after every change.
- Consider managed options. Amazon Bedrock Knowledge Bases can handle ingestion, chunking, and retrieval for you if you'd rather not run this plumbing yourself. Building it by hand first, as in this post, helps you understand what the managed service does.
- Add guardrails. Bedrock Guardrails can filter inputs and outputs for sensitive content.
Conclusion
With about 100 lines of Python, you have a working RAG backend: documents in S3, embeddings from Bedrock, similarity search in Postgres, and a clean FastAPI interface. From here, the most valuable investments are better chunking, evaluation, and observability, not a bigger model.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article