
On Explicit Allows with Amazon Bedrock Guardrails
The post describes a technique for inverting Bedrock Guardrails to create an "explicit allow" system that permits responses only on specific topics while blocking everything else.
Guardrails in Large Language Model (LLM) applications typically function as barriers. They prevent models from generating responses on certain topics or containing specific (potentially harmful) content. However, what if we want the inverse? What if we want to explicitly allow only certain topics and block everything else? This post explores how to implement "explicit allows" using Amazon Bedrock Guardrails through a technique that inverts the standard guardrail behavior.
The Standard Guardrail Approach
Typically, one aspect of guardrails are defining denied topics. When a user's prompt triggers a guardrail, the system returns a denial message like "Sorry, I can't help with this request" or whatever custom text you've configured. The model classifies the input against defined topics, and if it matches a denied topic, the guardrail intervenes.
This works well for blocking specific content categories, but doesn't natively support the inverse: allowing only specific topics while blocking everything else. For many applications, particularly those with strict compliance requirements or specialized domains, this "allowlist" approach is exactly what we need.
Inverting Guardrails for Explicit Allows
The key insight is to view denied topics not as blocks but as classification mechanisms. By leveraging the
ApplyGuardrails API (rather than the integrated InvokeModel or Converse APIs), we can implement this inversion ourselves.Here's the conceptual approach:
- Define a guardrail with "denied topics" that actually represent what you want to allow
- When the guardrail triggers (meaning the content matches your "allowed" topic), allow the request through
- When the guardrail doesn't trigger (meaning the content is outside your allowed domain), block the request
This effectively creates a complement operation on the guardrail's behavior.
Implementation Example: Investment Advice Only
Let's implement a guardrail that only allows investment-related queries and blocks everything else. First, we'll create a guardrail in the Amazon Bedrock console:
- Create a guardrail named "Investment Advice Only".
- You can skip the content filters.
- Add a denied topic for "The management or allocations of funds or assets".
- Name: Investment Advice
- Definition: Investment advice refers to inquiries, guidance, or recommendations regarding the management or allocation of assets with the goal of generating returns or achieving specific financial objectives.
Behind the scenes, when this guardrail runs, it will check if the input text is related to investment advice. If it is, the guardrail would normally block it, but we're going to invert that behavior.
The Code Implementation
The implementation requires a wrapper around the
ApplyGuardrails API to handle the complement operation. Here's the core logic:1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
def apply_guardrails_with_complement(prompt, guardrails, complement_text=None):
triggered_guardrails = []
for guardrail in guardrails:
response = bedrock.apply_guardrails(
guardrailIdentifier=guardrail["id"],
guardrailVersion=guardrail["version"],
text=prompt
)
# Standard guardrail case - guardrail triggered normally
if response["action"] == "GUARDRAIL_INTERVENED" and not complement_text:
triggered_guardrails.append(guardrail["name"])
# Complement case - guardrail didn't trigger but we want to block
elif response["action"] != "GUARDRAIL_INTERVENED" and complement_text:
triggered_guardrails.append(f"{guardrail['name']} (complement)")
# If any guardrails triggered (either normally or via complement)
if triggered_guardrails:
return {
"result": "blocked",
"triggered_guardrails": triggered_guardrails,
"message": complement_text if complement_text else response["output"]
}
# Nothing triggered, allow the prompt
return {
"result": "allowed",
"triggered_guardrails": [],
"message": prompt
}The key part is the condition that handles the complement case. When the guardrail doesn't intervene (
response["action"] != "GUARDRAIL_INTERVENED") but we're using complement mode (complement_text is provided), we treat this as a guardrail trigger.Testing the Implementation
To validate our approach, we tested two sets of examples:
Financial Examples
Prompts like "I've inherited $50,000. What's the safest way to invest this money for retirement?" and "Can you explain why my growth stocks have been underperforming?" should be allowed through our guardrail. We tested roughly 50 similar examples and showed ~80% of the financial queries were correctly identified and allowed.
Some financial queries that weren't recognized included "What's the historical performance of your small cap fund?" and "Should I be worried about inflation?" This suggests the guardrail's topic definition could be refined with additional examples or broader definitions to improve coverage. This is an iterative approach of refining the definition (effectively the prompt) is critical in standard guardrail cases as well.
Non-Financial Examples
Prompts like "What's the best time to visit Machu Picchu?", "Can I substitute coconut cream for heavy cream in this recipe?", and "Can you explain why my joints crack when I do yoga?" should be blocked. Testing showed 100% of non-financial queries were correctly blocked.
So our guardrail was probably a bit too strong, and there is a trade-off we have to consider (probably a different blog). And in this case we can use standard metrics like Precision/Recall/F1 to characterize our performance.
Practical Applications
This explicit allow approach is particularly valuable for:
- Specialized assistants that should only answer questions within their domain of expertise
- Compliance-sensitive applications where you need to strictly limit the scope of interactions
- Systems where you want to ensure the model stays "on task" and doesn't drift into unrelated topics
For example, a financial advisor bot might use this technique to ensure it only provides information about investments and doesn't attempt to answer questions about health, travel, or other unrelated domains.
Limitations and Considerations
While effective, this approach does have some limitations. It relies on the accuracy of the underlying classification model in the guardrail. Essentially you need a definition that's specific and to ensure the guardrail topic-definition-model can classify the inputs correctly.
Additionally, this approach requires using the
ApplyGuardrails API directly rather than the integrated APIs. You currently are limited in that you cannot use the guardrail and model call in the same invocation. Custom code requires custom "upkeep".Finally, this technique works best when the "allowed" domain is well-defined and distinguishable from other topics. For more nuanced boundaries, you might need multiple guardrails or additional logic.
Conclusion
By inverting the standard behavior of Bedrock Guardrails, we can implement explicit allows, ensuring LLMs only respond to topics within their designated domain. This approach provides a technique to build focused, compliant, and reliable AI applications.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article