Hybrid Generative AI Architectures: Uniting Amazon Bedrock and Amazon SageMaker
Discover how to architect a hybrid Generative AI stack by pairing the serverless reasoning of Amazon Bedrock with the precision and custom control of Amazon SageMaker. Learn how to route requests, integrate tools via Lambda, and optimize inference costs.
When teams adopt Generative AI on AWS, architectural debates often reduce to a binary choice: Amazon Bedrock or Amazon SageMaker.
- Amazon Bedrock provides serverless, API-driven access to frontier Foundation Models (Claude, Llama, Mistral, Titan) with managed agents, knowledge bases, and guardrails.
- Amazon SageMaker provides full-lifecycle ML control—from custom distributed training on Trainium/Accelerated compute instances to custom inference runtimes, Hugging Face model deployment, and granular network isolation.
In production enterprise workloads, the most resilient, cost-effective architectures do not pick one over the other. They adopt a hybrid pattern: Bedrock functions as the cognitive router and zero-ops reasoning engine, while SageMaker handles domain-specialized, latency-critical, or custom fine-
Why Combine Them?
- Token Cost Optimization: Routing high-volume, narrow tasks (e.g., entity extraction, fraud classification, or tabular embeddings) to Bedrock frontier models burns tokens quickly. A smaller, quantized model (such as a fine-tuned 8B model or RoBERTa variant) hosted on a SageMaker endpoint delivers sub-15ms latency at a predictable, fixed hourly instance cost.
- Deterministic Domain Execution: While frontier LLMs excel at zero-shot reasoning, highly regulated compliance scoring or proprietary internal taxonomies often require proprietary fine-tuning and strict weight governance that SageMaker enables.
- Managed Orchestration: Bedrock provides managed conversational state, RAG integration via Bedrock Knowledge Bases, and native multi-agent coordination without requiring you to run custom LangGraph or Celery workers.
Technical Walkthrough: Calling SageMaker as a Bedrock Action Tool
In an agentic workflow, an Amazon Bedrock Agent can treat an Amazon SageMaker real-time endpoint as an OpenAPI action group. When user input requires deep domain analysis, the agent orchestrates the invocation through an AWS Lambda function that wraps the SageMaker runtime.
- The Orchestration Lambda Function (Python 3.12)
The following Lambda handler receives an action group event from Amazon Bedrock, packages the payload, queries the target SageMaker endpoint, and returns the structured response back into the Bedrock reasoning loop:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
import json
import os
import boto3
from botocore.exceptions import ClientError
sagemaker_runtime = boto3.client("sagemaker-runtime")
ENDPOINT_NAME = os.environ.get("SAGEMAKER_ENDPOINT_NAME", "domain-scorer-v1")
def lambda_handler(event, context):
agent = event.get("agent")
action_group = event.get("actionGroup")
api_path = event.get("apiPath")
parameters = event.get("parameters", [])
# Extract input parameter passed by Bedrock's planner
input_text = ""
for param in parameters:
if param.get("name") == "document_text":
input_text = param.get("value")
break
if not input_text:
return format_response(action_group, api_path, 400, {"error": "Missing 'document_text'"})
try:
# Invoke SageMaker Endpoint
payload = json.dumps({"inputs": input_text})
response = sagemaker_runtime.invoke_endpoint(
EndpointName=ENDPOINT_NAME,
ContentType="application/json",
Accept="application/json",
Body=payload
)
model_output = json.loads(response["Body"].read().decode())
response_body = {
"application/json": {
"body": json.dumps(model_output)
}
}
return format_response(action_group, api_path, 200, response_body)
except ClientError as e:
return format_response(
action_group,
api_path,
500,
{"error": f"SageMaker Invocation Error: {e.response['Error']['Message']}"}
)
def format_response(action_group, api_path, http_code, body):
"""Formats response schema required by Amazon Bedrock Agents."""
return {
"messageVersion": "1.0",
"response": {
"actionGroup": action_group,
"apiPath": api_path,
"httpStatusCode": http_code,
"responseBody": body
}
}- OpenAPI Schema for the Bedrock Agent
Save this OpenAPI 3.0 schema to Amazon S3 to define the tool contract for the Bedrock Agent:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
{
"openapi": "3.0.0",
"info": {
"title": "SageMaker Domain Analyzer Tool",
"version": "1.0.0",
"description": "Analyzes input text using a specialized SageMaker endpoint."
},
"paths": {
"/analyze": {
"post": {
"summary": "Run domain-specific inference on SageMaker",
"description": "Routes complex domain verification, anomaly classification, or fine-tuned scoring to the SageMaker model endpoint.",
"operationId": "analyzeDomainText",
"parameters": [
{
"name": "document_text",
"in": "query",
"description": "The raw text snippet or payload to evaluate",
"required": true,
"schema": {
"type": "string"
}
}
],
"responses": {
"200": {
"description": "Inference results from SageMaker",
"content": {
"application/json": {
"schema": {
"type": "object"
}
}
}
}
}
}
}
}
}Security, IAM, and Networking Guardrails
When bridging Bedrock and SageMaker in enterprise production, adhere to these key practices:
- IAM Principle of Least Privilege: Your Lambda function needs only
sagemaker:InvokeEndpointrestricted to the exact ARN of the target endpoint. Never grant wildcard access across SageMaker endpoints. - VPC Endpoints (PrivateLink): Ensure communication between Lambda and the SageMaker runtime passes entirely through AWS PrivateLink VPC endpoints, avoiding public internet traversal.
- Bedrock Guardrails: Wrap the entry point of your Bedrock Agent with Guardrails for Amazon Bedrock to filter PII, prevent prompt injections, and enforce safety boundaries before data ever reaches your custom downstream SageMaker models.
Conclusion
Building generative applications at scale does not require choosing between the convenience of managed APIs and the flexibility of custom ML infrastructure. By leveraging Amazon Bedrock as the conversational orchestrator and Amazon SageMaker as the specialized execution engine, you balance developer velocity, operational autonomy, and infrastructure costs across the entire lifecycle.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article