
Securely shipping GenAI systems in Amazon Bedrock with data-aware governance
By Pranava Adduri (Co-founder and CTO of Bedrock Data), Ishara Premadasa (Solutions Architect, AWS), Jean Malha (Senior Solutions Architect -Bedrock, AWS)
As organizations move generative AI applications from prototype to production using Amazon Bedrock, they face a new kind of governance challenge. Unlike traditional software where dependencies are static code libraries, AI models have dependencies on data. This includes training sets, fine-tuning data, and Retrieval-Augmented Generation (RAG) sources.
To ship these products securely, builders need more than just an inventory of data sources. They need a Data Bill of Materials (DBOM) to map an agent back to its underlying data, and a way to verify that safety controls match the sensitivity of that data.
In this post, we demonstrate how to use Bedrock Data ArgusAI to automatically generate a DBOM for your Amazon Bedrock Agents and perform a Guardrail Gap Analysis. We will walk through a technical scenario involving a PTO Approval Agent to show how you can detect and remediate data leakage risks before they reach production.
Solution Overview
This solution uses ArgusAI (by AWS Partner Bedrock Data) to scan your AWS account for Amazon Bedrock usage. This is done via a pre-deployed role in a target AWS that ArgusAI assumes to map the lineage between your Agents, Knowledge Bases, and underlying data sources to create a DBOM. It then compares this data against your active Amazon Bedrock Guardrails to identify gaps.
The workflow consists of four stages:
- Discovery: Mapping Agents to their Knowledge Bases and establishing lineage to the origin data.
- Classification: Scanning the data to identify sensitive information.
- Gap Analysis: Validating if the attached Guardrail handles the specific data types found.
- Remediation: Updating the Amazon Bedrock Guardrail configuration to close the gap.
Constructing the Data Bill of Materials
Before diving into the specific agent example, it is important to understand how the DBOM is constructed across different Amazon Bedrock resources.
ArgusAI traces dependencies for various workloads:
- Custom Models: It identifies training and fine-tuning datasets.
- Knowledge Bases: It maps vector stores back to their source S3 buckets or data connectors.
- Agents: It establishes lineage from the agent through the Knowledge Base to the origin data.
This comprehensive mapping ensures that whether you are fine-tuning a model or building a RAG agent, you have visibility into the data dependencies that drive model behavior.
Walkthrough: Securing a PTO Approval Agent
To demonstrate this workflow, we will look at a simplified, real-world scenario employed by a Bedrock Data customer: building an internal "PTO Approval Agent" using Amazon Bedrock.

In this example, the agent helps employees check their leave balance and request time off. To do this, the agent retrieves documents from two data sources:
- acme_pto_export: An HR export of PTO taken thus far by employees.
- Acme/HR Directory: The internal HR directory.
We have applied a PTO Agent Guardrail configured to block common threats like PII (Personally Identifiable Information), PHI (Protected Health Information), and PCI (Payment Card Information).
The question is simple: Is that guardrail enough to ship this agent safely?
Step 1: Automated Discovery and DBOM Generation
First, ArgusAI uses the pre-deployed role in a target AWS environment to introspect the Agent and Knowledge Base workloads present. ArgusAI does this using a combination of ListKnowledgeBases, ListAgents, GetKnowledgeBase and GetAgent API calls. It then traces the lineage from the PTO Approval Agent back to objects in the bucket acme_pto_export and pages in the SharePoint site Acme/HR Directory.
This lineage forms the Data Bill of Materials (DBOM). The goal of the DBOM is to map the model or agent back to its data so you understand the context of that data. For example, it helps you determine if your public-facing agent is inadvertently connected to a production data lake containing sensitive employee records.
Step 2: Data Classification and Risk Mapping
ArgusAI is built on Bedrock Data's Metadata Lake, a graph knowledge base of enterprise data and its attributes including the data's classification and lineage.
In this case, Bedrock Data's Metadata Lake is already aware that:
- acme_pto_export: Contains PII.
- Acme/HR Directory: Contains PII and Compensation Data.
ArgusAI builds on Metadata Lake to assess that the PTO Approval Agent can access both PII and Compensation Data.
Step 3: Guardrail Gap Analysis
This is the critical step where we validate our safety controls. ArgusAI retrieves the configuration of the PTO Agent Guardrail and compares the detected data types from Step 2 against the active blocking policies.
As shown in the architecture diagram, the analysis reveals two distinct results36:
- Blocked By Guardrail (Green): The system detects PII in the source data. Because the guardrail is configured to Deny PII, this risk is effectively mitigated.
- Exposed To Model (Red): The system detects Compensation data (salary details) in the HR Directory. The current guardrail blocks PII, PHI, and PCI, but it has no rule for financial compensation data.
A Critical Gap: The PTO Approval Agent can send salary information to the underlying model as well as return it to the caller. If a user asks "How much does the engineering lead make?", the agent might retrieve and generate an answer using that exposed data.
Step 4: Remediation via Amazon Bedrock Guardrails
To fix this, we must update the Guardrail configuration. ArgusAI provides the recommendation to close this specific gap.
We can update the UpdateGuardrail API parameters to explicitly redact or block the "Compensation" data type discovered by the DBOM. We can achieve this by adding a Topic Deny rule or a custom regex filter for salary formats.
Here is the example JSON configuration we used to close the gap using a Topic policy:

By applying this configuration:
- Ingestion: The model can still access the HR directory for authorized tasks.
- Inference: If a user attempts to access the exposed Compensation data, the Guardrail intercepts the request.
- Result: The gap is closed. The agent is now safe to ship because the specific data risks in the DBOM are covered by the guardrail.
Step 5: Continuous Monitoring
Because the DBOM is dynamic, if a data engineer adds a new folder containing "Performance Reviews" to the data source next week, ArgusAI will detect the new file type and re-run the Gap Analysis. It will alert you that your current guardrail lacks a filter for performance data, allowing you to update the policy before the risk is exploited.
Conclusion
By combining Amazon Bedrock for powerful model capabilities and Bedrock Data ArgusAI for governance, builders can deploy GenAI agents with confidence.
The Data Bill of Materials provides the visibility needed to understand what your models know. The Guardrail Gap Analysis ensures your safety controls are actually effective against your specific data risks. This allows you to ship AI products securely without slowing down innovation.
Call to Action
- Read the Docs: Explore Amazon Bedrock Guardrails to understand available policies
- Start Building: Implement a Guardrail for your own RAG agent and test it using the ApplyGuardrail API
- Try the Partner Solution: Learn more about automating this workflow at Bedrock Data
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article