RAG Is Becoming Infrastructure: Will Amazon Bedrock Managed Knowledge Base Replace Custom RAG?
From custom RAG and Amazon Bedrock Knowledge Bases to Amazon Bedrock Managed Knowledge Base, a fresh look at enterprise RAG selection, governance, and the path to agents.
Over the past two years, RAG (Retrieval-Augmented Generation) has become the default architecture for enterprise generative AI. Customer-support knowledge bases, internal policy Q&A, contract review, sales enablement — most of these projects start with the same question: can it search our own knowledge? Building an agent is no different.
The RAG pattern sounds simple:
Chunk the documents → embed → load into a vector store → retrieve relevant passages → hand them to the LLM.
But once it reaches production, it behaves more like a system you have to operate for the long haul: connectors need maintenance, permissions have to be inherited, documents have to be parsed, chunking gets tuned again and again, embedding models get upgraded, rerankers need evaluation, answers must be traceable, data updates have to sync, costs have to be allocated, and agent calls need to stay under control.
So I've increasingly come to a view: RAG is shifting from "a module inside an application" to "a piece of infrastructure — a primitive — in the cloud."
Amazon Bedrock Managed Knowledge Base puts that shift out in the open. AWS's own framing is blunt about it: take the components you used to assemble and maintain yourself — storage, retrieval, embeddings, reranking, model selection — and fold them into a single managed primitive. What it targets is exactly the part of enterprise RAG that is most common, most repetitive, and least differentiating for the business.
So: will Managed Knowledge Base replace custom RAG?
My answer: it will replace a great deal of the "building RAG for the sake of building RAG" work — but not all custom RAG. Concretely:
- If you just want enterprise knowledge to reach a generative AI application reliably, quickly, and securely, Managed Knowledge Base should probably be your default.
- If you've already built mature retrieval capabilities around Zilliz / Milvus, Amazon OpenSearch Service, or LanceDB — or you're building a highly customized retrieval platform — custom RAG still earns its place.
- If you're moving from single-turn Q&A toward agentic AI, the Managed Knowledge Base + AgentCore Gateway combination is worth evaluating before you go write a pile of retrieval tool wrappers yourself.
Let me lay out the reasoning.

1. Why was enterprise RAG so hard before?
RAG demos tend to look great: a few dozen lines of code, read a handful of PDFs, drop them into a vector store, and you can ask questions. The hard part of enterprise RAG isn't in those few dozen lines — it's everything after the demo.
1.1 Connectors: enterprise knowledge never lives in one place
Enterprise knowledge won't sit neatly in a single S3 bucket. It's scattered across SharePoint, Confluence, Google Drive, OneDrive, internal wikis, ticketing systems, CRMs, databases, and email archives, in formats ranging from PDF, PPT, and Excel to web pages, images, and scans. Each system has its own authentication, pagination, incremental sync, permission model, and rate limits.
So the first layer of RAG complexity isn't vector search — it's: how do I get this data in, stably, securely, and continuously? Writing one connector is easy. Keeping it running in production without breaking is the hard part, and many teams underestimate it at the start.
1.2 Parsing: documents aren't plain text
RAG input is rarely clean Markdown. Real documents have tables, headers and footers, footnotes, tables of contents, multi-column layouts, scans, flowcharts, numbered contract clauses, approval records. Get the parsing strategy wrong and you poison everything downstream — embeddings and retrieval included. Common traps:
- Tables flattened into meaningless text
- Images and captions lost
- PDF reading order scrambled
- Headers and footers repeatedly landing in chunks and polluting retrieval
- Contract clause numbers separated from their text
- A single concept spread across pages, tables, and attachments, leaving retrieval with incomplete context
A lot of RAG failures aren't the model being too dumb — it's that the context handed to it was already broken.
1.3 Chunking: too big hurts recall, too small fragments context
Chunking is one of the most folklore-driven and most underestimated parts of RAG engineering. Too big, and the embedding's semantics blur, so retrieved passages carry more noise; too small, and the business meaning gets cut apart, so the model never sees the full context. Worse, different documents need different strategies: product manuals by section, FAQs by question-answer pair, contracts by clause, tables preserving their rows and columns, code docs by function and module, audit reports keeping the evidence chain and page numbers intact. One fixed chunk size that handles every document is basically a fantasy.
1.4 Retrieval and rerank: finding something isn't finding the right thing
Teams doing RAG for the first time usually only care whether something came back. Production cares about a different set of questions: does recall cover the key facts? Are the top-ranked passages actually useful? Should the user's question be decomposed into sub-questions? How do you route across multiple knowledge bases? When nothing relevant is found, should the system refuse to answer? Does the answer carry traceable citations?
Plain top-k vector similarity usually isn't enough. Enterprise RAG gradually pulls in hybrid search, metadata filters, rerank, query rewriting, multi-hop retrieval, plus an evaluation set and LLM-as-a-judge. By that point, RAG isn't a retrieval module anymore — it's a small search platform.
1.5 Permissions: easiest to ignore in a demo, hardest to get past in production
An enterprise knowledge base can't just answer "does this document exist." It has to answer: is this user allowed to see this content?
If an employee uses AI Q&A to pull HR, finance, contract, or security-audit content they were never supposed to see, the system stops being a productivity tool and becomes a data-leak entry point. Permissions span several layers: how source permissions sync, how retrieval filters per user, whose identity an agent acts under when it calls the knowledge base, whether generated answers leak sensitive context through citations, and whether logs and evaluation data retain sensitive content. This is why enterprise RAG always ends up at governance, rather than stopping at the vector store.
1.6 Evaluation: no eval set, no production-grade RAG
The most common bad smell in a RAG project is the sentence: "we tried a few questions, and it felt pretty good."
That's not evaluation. That's a gut feeling. Production-grade RAG needs at least a set of golden questions: high-frequency business questions, edge cases, permission-sensitive questions, multi-document questions, questions that should be refused, questions where old and new policies conflict, questions where citations must be exact. Every time you change the parser, chunking, embedding, rerank, prompt, or model, you should run a regression. Otherwise RAG drifts into a dangerous state: it looks better today and quietly gets worse tomorrow in a different scenario, and you don't notice.
2. What does Managed Knowledge Base add?
Amazon Bedrock Knowledge Bases already helps developers bring proprietary data into generative AI: retrieving relevant information, generating answers with citations, supporting reranking, and fitting into Bedrock Agents workflows.
Managed Knowledge Base, which became generally available at AWS Summit New York in June 2026, pushes more of the infrastructure work down a layer. Its capabilities cluster into three areas — native data connectors, Smart Parsing, and an Agentic Retriever — and it can be exposed as a prebuilt target for AgentCore Gateway, discoverable and callable by agent frameworks over an MCP-compatible interface.
2.1 Native data connectors: write fewer connectors, do more business
At launch, Managed Knowledge Base ships six native connectors for enterprise data sources: Amazon S3, SharePoint, Confluence, Google Drive, OneDrive, and Web Crawler.
The value isn't just "a few fewer sync scripts." Connectors hide a lot of the implicit complexity of enterprise RAG: source authentication, permission inheritance, incremental sync, content-type detection, document hierarchy, propagation of deletes and updates, source-specific formats. When those become managed, teams can put their time back into "which knowledge is worth bringing into the AI application" and "how does this business question get answered correctly" — instead of repeatedly patching data-moving scripts.
One permission detail is worth calling out: Managed Knowledge Base supports document-level permission filtering at retrieval time based on access control lists (ACLs), for every source except Web Crawler. That maps directly onto the easily-missed problem from section 1.5 — it brings the "filter by document permission at retrieval time" layer into the managed scope, though you still design the overall permission model yourself.
2.2 Smart Parsing: turning parsing from a tuning exercise into a managed default
The core of Smart Parsing is picking an appropriate parsing strategy automatically, per content type and per source. That matters, because a lot of retrieval failures don't happen at vector search — they happen at ingestion: structure breaks, tables get lost, the text-image relationship disappears, chunks lose their semantic boundaries, and after that even the strongest model can only answer from broken context.
It pushes that complexity down: a more suitable data model per source, handling of complex structures with images and tables, chunks generated according to document structure, while still leaving room for customization beyond the defaults.
A boundary worth keeping in mind: managed defaults cover most ordinary documents, but highly structured contracts, multi-column scans, and chart-heavy reports still warrant spot-checking the parsed output before you go live. In other words, the center of gravity shifts — where you used to spend a lot of time tuning parsers and chunking, you should now spend it on defining business eval sets, permission boundaries, and the knowledge lifecycle.
2.3 Agentic Retriever: from a single retrieval to multi-step retrieval
Traditional RAG is mostly single-turn: the user asks, the system retrieves top-k, the model answers. But real enterprise questions often aren't like that. For example:
"What's our ML platform team's cloud budget this year? And if we want to prepay an annual commitment, does finance policy allow it?"
That spans at least two knowledge domains: team budget and finance policy. A single retrieval might find one of them, but not necessarily connect the two. The Agentic Retriever aims to give retrieval a bit of planning ability: understand the user's intent, break a complex question into multiple retrieval steps, do multi-hop retrieval across one or more knowledge bases, then assemble better-grounded context.
This matters for agentic AI, because an agent's knowledge retrieval is no longer "find a passage" — it's more like repeatedly confirming facts, hunting for evidence, comparing rules, and filling in context in order to complete a task. That's exactly why the integration between Managed Knowledge Base and AgentCore Gateway is worth attention: the knowledge base stops being just a retrieval API on the application backend and becomes a tool an agent can discover, call, authorize, and observe.

3. Will it replace custom RAG? First, be clear about what you're building
I'd advise against framing this as "Managed Knowledge Base vs. Zilliz / OpenSearch / LanceDB — who wins." It isn't a database selection problem. A more useful question is:
Are you building a "business AI application," or a "reusable retrieval platform"?
If the goal is a business application, Managed Knowledge Base will likely turn a lot of the low-level work into managed defaults. If the goal is the retrieval platform itself — or you need deep control over indexing, recall, ranking, storage, or cross-cloud deployment — custom RAG still has a strong case.
The table below works as a quick reference.
| Dimension | Custom RAG: Zilliz / Milvus / OpenSearch / LanceDB, etc. | Amazon Bedrock Knowledge Bases | Amazon Bedrock Managed Knowledge Base |
|---|---|---|---|
| Core positioning | You own the retrieval stack; suited to platformization and deep customization | Build RAG apps inside Bedrock, reducing some integration complexity | Further manages the enterprise RAG pipeline, aimed at faster delivery and agent integration |
| Data ingestion | Build or integrate connectors yourself | Connect data sources, then build the knowledge base | Native connectors that reduce enterprise data-source onboarding work |
| Document parsing | Choose your own parser, OCR, table handling, multimodal strategy | Use Bedrock KB capabilities and configuration | Smart Parsing automatically handles multi-format data prep, lowering tuning cost |
| Vector storage | Fully your choice; deep optimization of index, storage, filtering, cost | Use a supported vector store, or create an OpenSearch Serverless vector store from the console | More of the underlying storage, embedding, and retrieval orchestration is managed |
| Retrieval strategy | Maximum customization: complex hybrid search, GraphRAG, domain rerank | Supports retrieval, generation, citations, reranking, and other RAG capabilities | Agentic Retriever for complex, multi-step, multi-knowledge-base retrieval |
| Model selection | Integrate models and embedding services yourself | Choose models within Bedrock | Retains Bedrock's model-selection flexibility while managing more default components |
| Permissions and governance | Design permission inheritance, auditing, logging, and evaluation yourself | Can fit into AWS IAM and Bedrock governance | Integrates with AgentCore Gateway for agent-tool permissions, observability, and evaluation; supports document-level ACL filtering at retrieval time |
| Agent integration | Usually requires wrapping your own MCP server / tool wrapper | Integrates with Bedrock Agents and others | Can act as a prebuilt AgentCore Gateway target, called over an MCP-compatible interface |
| Engineering effort | High; suited to organizations with a platform team | Medium | Low to medium; suited to quickly standardizing enterprise RAG |
| Best-fit scenarios | Retrieval platforms, very-large-scale customization, special indexing, cross-cloud/on-prem, algorithmic differentiation | Bedrock-native RAG apps that want less infrastructure assembly | Enterprise knowledge Q&A, internal agents, cross-team knowledge bases, wanting less RAG plumbing to maintain |
So Managed Knowledge Base isn't out to make vector databases disappear. What it actually replaces is repetitive labor: every team writing its own connector, every project re-tuning the parser and chunking, every application wrapping the knowledge base as a tool again, every agent implementing its own retrieval orchestration, every business line rediscovering evaluation from scratch. That work is important, but it isn't differentiating.
Real differentiation lives elsewhere: which knowledge gets included, how permission boundaries are defined, how business questions are modeled, how retrieval quality is evaluated, how agents use knowledge to complete tasks, and how results flow back into business processes. That's what "RAG becoming infrastructure" really means.
4. Three choices, and who each one fits
Turning the judgment into selection, there are roughly three paths. What I most want to make clear here is when not to choose each one.
4.1 Stay with custom RAG: when retrieval itself is your product
Three situations where I'd still seriously consider building your own, or at least keeping custom components.
First, you already have a mature retrieval platform: an indexing service shared across business lines, a stable ingestion pipeline, a custom metadata schema, hybrid retrieval and rerank strategies, complete permission filtering, retrieval evaluation and monitoring, cost and capacity governance. In that case you shouldn't refactor just because a new service shipped. The more realistic move is to add Managed Knowledge Base as a new standard option in your architecture, not migrate wholesale.
Second, you need very fine-grained low-level control: custom ANN index parameters, special multi-route recall, complex metadata filters, a domain-specific reranker, fusing graph retrieval with vector retrieval, cost optimization under multi-tenancy, extreme low latency — or data that must stay within a specific network or deployment form.
Third, retrieval quality is your competitive edge: legal search, medical knowledge, financial research, code search, patent search, scientific literature — domains where teams often need to keep innovating on the retrieval pipeline. Here, general enterprise knowledge can go to a managed service, while the core retrieval engine may still need to be your own.
The flip side: if you're only building internal Q&A but stand up an entire retrieval stack first in the name of "control," that's usually over-engineering.
4.2 Use Bedrock Knowledge Bases: when you want Bedrock-native RAG
If you're already building on Amazon Bedrock, Bedrock Knowledge Bases is a natural fit: your application mainly runs on AWS, you need proprietary data connected to Bedrock models, you want answers with citations for easy verification, you want to use models / Agents / guardrails / evaluation within Bedrock, you want to write less RAG glue code while keeping some configuration room, and your data sources, vector storage, and retrieval approach are already well defined.
A way to think about it: it connects your RAG application into Bedrock, but data sources, storage, sync, and retrieval configuration still need engineering attention from you. For many enterprises, that's already much lighter than going fully custom.
4.3 Go with Managed Knowledge Base: when you want enterprise RAG infrastructure
Three signals say you should evaluate Managed Knowledge Base first.
One: you want enterprise knowledge in an AI application fast, not to build a retrieval platform — especially when the sources are common systems like SharePoint, Confluence, Google Drive, OneDrive, S3, and web pages, where native connectors save a lot of up-front engineering.
Two: you don't want to keep feeding RAG plumbing. Connector upgrades, API changes, parsing-strategy adjustments, embedding-model swaps, rerank tuning, ingestion retries, index scaling, multi-team knowledge-base management — these look simple early and get tedious over time. If they aren't your core competency, managed is usually the better deal.
Three: you're building agents. What sets an agent apart from an ordinary chat application is that it actively calls tools, breaks down tasks, queries knowledge, takes actions — and has to be governed. When a knowledge base can be exposed as an MCP-compatible tool through AgentCore Gateway, the architecture shifts from "every agent writing its own knowledge-base calling logic" to "the enterprise knowledge base as a governed agent tool, exposed, authorized, observed, and evaluated in one place."
That's a big deal for enterprise architecture. Once you have a lot of agents, the hard thing to manage isn't the model — it's: which agent can call which tools, on whose behalf, which knowledge it accessed, whether it overstepped, how well it performed, who the cost belongs to, and how to trace it when it fails. Bringing knowledge access into the agent's tool governance is exactly the point of this combination.

5. A more realistic migration path: don't start with a "big migration"
Enterprise architecture migrations fear two extremes: refactoring everything the moment a new service appears, or never touching the old system because it still runs. I'd suggest three steps.
Stage one: PoC RAG — first prove the business problem is worth solving
The goal isn't to prove "RAG works," but to prove "this business problem is worth solving with RAG." In practice: pick a business domain with clear boundaries (internal IT Q&A, sales material, policies — any of these work); gather a set of high-quality documents instead of dumping every company document in at once; build 50 to 100 golden questions and record each one's expected answer, source document, permission requirement, and refusal condition; then use that same question set to compare basic RAG, Bedrock Knowledge Bases, and Managed Knowledge Base on both results and engineering effort.
The most important output of this stage isn't a demo — it's the eval set. Without it, every later architectural choice turns into a subjective call.
Stage two: production-grade RAG — fill in governance and operations
Once the PoC proves value, the next step isn't to rush every data source in — it's to fill in production capabilities: data-source owners, ingestion frequency, delete-and-update sync strategy, IAM and application permission model, sensitive-data handling, answer citations and traceability, retrieval-quality metrics, cost tags and chargeback, retry/alerting/monitoring, and regression tests for model upgrades and prompt changes.
Managed Knowledge Base's value shows up more clearly here than at the PoC stage. In a PoC you can get by with hand-written scripts; in production, every "get by" turns into operational debt.
Stage three: Agentic KB — turn the knowledge base into an agent tool
As RAG moves from Q&A toward agents, the knowledge base's role changes again: it's no longer the "user asks, system answers" backend module, but a trusted source of knowledge an agent uses to complete a task. An IT agent looks up the incident-handling runbook and then opens a ticket; a FinOps agent looks up cost policy and then drafts optimization recommendations; a sales agent looks up product material and pricing policy and then drafts a customer email; a security agent looks up the security baseline and exception approvals and then judges the risk of a change.
At this point, expose Managed Knowledge Base as a governed tool through AgentCore Gateway, and design a few things deliberately: which agents can access which knowledge bases, whether the user's identity is passed on each call, whether to filter by department/role/tenant, the Agentic Retriever's maximum iteration count and cost ceiling, how tool-call logs feed observability, and how results feed evaluation and feedback. Once you're here, RAG has genuinely moved from "application component" to "enterprise agent infrastructure."
6. Conclusion: the competitive edge of RAG is shifting
Pull back and look at RAG's evolution, and the focus of competition keeps moving. In the first phase, it was "can you get the document out?" In the second, "can you retrieve precisely, answer reliably, and cite your sources?" In the phase we're in now, it's "can you bring knowledge into enterprise agents, with clear permissions, evaluable quality, governable cost, and observable calls?"
So, back to the original question — will Managed Knowledge Base replace custom RAG? My judgment: yes, but what it replaces is the generic, repetitive, low-differentiation RAG plumbing. Where retrieval itself is the business moat, custom RAG will be around for a long time.
For a mature enterprise, the sensible posture isn't either/or — it's layering: general knowledge Q&A and ordinary business agents default to a managed Knowledge Base; businesses with special retrieval needs use custom or hybrid architectures; genuinely high-value retrieval capability gets consolidated into an enterprise platform and exposed to agents through a standard interface. Application teams shouldn't have to pick the parser, chunking, embedding, vector store, rerank, evaluation, and permission model all over again each time — these should sink down to the platform layer, the way logging, monitoring, identity, and networking did.
The significance of Managed Knowledge Base isn't to announce that "vector databases don't matter anymore." It's to make one thing clear:
Enterprises shouldn't rebuild a nearly identical RAG pipeline for every single AI application.
For most enterprise applications, the competitive focus of RAG will move from "who can build a pipeline" to who understands the business's knowledge boundaries better, who can build a high-quality eval set, who can do permissions and auditing properly, who can connect the knowledge base into agent workflows, and who can keep improving answer quality and cost. In a sentence:
The competitive edge of RAG is shifting from "can retrieve" to "can govern, can evaluate, can connect to agents."
That's also why I think architects should give Amazon Bedrock Managed Knowledge Base a serious look. It isn't yet another RAG demo tool. It's more like AWS making a statement: enterprise RAG is becoming infrastructure.
References
- AWS News Blog: Introducing Amazon Bedrock Managed Knowledge Base for faster, more accurate enterprise AI applications
https://aws.amazon.com/blogs/aws/introducing-amazon-bedrock-managed-knowledge-base-for-faster-more-accurate-enterprise-ai-applications/ - Amazon Bedrock User Guide: Retrieve data and generate AI responses with Amazon Bedrock Knowledge Bases
https://docs.aws.amazon.com/bedrock/latest/userguide/knowledge-base.html - Amazon Bedrock AgentCore Developer Guide: Amazon Bedrock AgentCore Gateway
https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway.html - AWS News Blog: Top announcements of the AWS Summit in New York, 2026
https://aws.amazon.com/blogs/aws/top-announcements-of-the-aws-summit-in-new-york-2026/
Authors

Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article