
The Missing Layer: Why Observability Alone Won't Save Your Agent
Observability tells you what your agent did, but not whether it was correct. This post walks through building dev-time and production evaluations for AI agents using DeepEval and Amazon Bedrock AgentCore Evaluations.
In our previous post, we explored how Business Traces for Agents helps you understand what your agent is trying to do. It captures intent and phase transitions that OTEL traces miss. Journaling solved a real problem, when our cost optimization agent timed out, we could finally see where it got stuck instead of hunting through 100+ spans.
But journaling answers "what was the agent thinking?" It doesn't answer "was the agent correct?"
We built an agentic cost optimizer that analyzes AWS Lambda functions and generates savings recommendations. The agent worked, completed phases and produced reports. But were the recommendations accurate?
When we started building agents, a combination of manual testing and intuition got us quite far. We'd run the agent, read the output, read the OTEL, journal... But this approach doesn't scale, manually verifying tool calls, catching prompt regressions, and reviewing production traffic becomes hard as complexity grows.
This post shares what we learned building a two-layer evaluation system: DeepEval for fast feedback during development and Amazon Bedrock AgentCore Evaluations for continuous monitoring in production.
The Problem: Observability Shows Behavior, Not Correctness
Our analysis agent executes an autonomous task with multiple phases: discover resources, collect metrics, fetch pricing, and generate recommendations. Each phase depends on accurate data from the previous one.
When we manually reviewed the output reports, we discovered problems. The agent wasn't always calling the pricing API. Instead, it sometimes guessed prices based on its training data. The numbers looked good, but they were wrong. This is a known limitation, LLMs excel at reasoning but are unreliable at precise math operations. Without explicit tool use for pricing lookups and time calculations, the LLM defaults to guessing. Similarly, time calculations were inconsistent, with the agent sometimes fetching metrics from invalid date ranges.
This is the gap that observability alone cannot fill. You can see that a tool was called, but not whether it was the *right* tool for the task. You can see that a phase completed, but not whether the output was *correct*. You can see that the agent finished, but not whether it achieved the *goal*.
We needed a way to answer three questions:
- Tool correctness: Did the agent use the right tools with the correct parameters?
- Task completion: Did the agent actually accomplish what we asked?
- Continuous validation: Are these checks running automatically against real production traffic, building meaningful averages over time?
Manual review answered these questions, but it was slow, inconsistent and difficult to reproduce across prompt or model changes. We needed automation.
Evaluations During Development
We have chosen DeepEval , an open-source LLM evaluation framework, for development-time evaluations. It provides out-of-the-box metrics for tool correctness and task completion, and integrates with Amazon Bedrock models for LLM-as-judge evaluations.
What We Measure?
We use two DeepEval metrics, both judged by an LLM:
ToolCorrectnessMetric validates that the agent called the expected tools with the right parameters. We define what tools we expect for a given task:
1
2
3
4
5
6
7
expected_tools = [
ToolCall(name="journal", input_parameters={"action": "start_task", "phase_name": "Discovery"}),
ToolCall(name="use_aws", input_parameters={"service": "lambda", "action": "list_functions"}),
ToolCall(name="current_time_unix_utc", input_parameters={}),
ToolCall(name="use_aws", input_parameters={"service": "pricing", "action": "get_products"}),
ToolCall(name="storage", input_parameters={"action": "write", "filename": "analysis.txt"}),
]TaskCompletionMetric evaluates whether the agent accomplished the goal, not just whether it called the right tools:
1
2
3
4
5
task_metric = TaskCompletionMetric(
task="Analyze Lambda functions, collect CloudWatch metrics, identify cost optimization opportunities, and save analysis results with specific recommendations and savings estimates",
model=judge_model,
threshold=0.7,
)Note: DeepEval metrics return scores from 0.0 to 1.0, with a default passing threshold of 0.5. We set ours to 0.7 as a higher bar, aiming for "good enough" rather than accepting borderline passes. You can tune this based on how strict you want your evals to be.
Mocking for Consistency
To make these evals repeatable, we mock the AWS API responses. This isn't an end-to-end test where we validate the full integration with AWS. It's somewhere between unit and integration testing where we run the real agent with real LLM calls, but mock the external dependencies. There's also a security benefit to not expose actual account data in tests that could run in CI/CD pipelines.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
def create_use_aws_tool(capture: ToolCapture):
"""Create use_aws tool with mock responses."""
def use_aws(service: str, action: str, **kwargs) -> dict:
"""Call AWS APIs."""
capture.record("use_aws", service=service, action=action, **params)
# Lambda API
if service == "lambda" and action == "list_functions":
return MOCK_LAMBDA_FUNCTIONS
# CloudWatch Metrics API
if service == "cloudwatch":
if action == "get_metric_data":
return MOCK_CLOUDWATCH_GET_METRIC_DATA
# Pricing API
if service == "pricing" and action == "get_products":
return MOCK_PRICING_LAMBDA_COMPUTE
return {"error": f"Not implemented: {service}.{action}"}
return use_awsThis way we're testing agent behavior, not whether CloudWatch returned the same metrics as yesterday. The agent runs against a controlled environment where we know exactly what data it should receive and what tools it should call.
Different Judge, Different Family
To get unbiased evaluation results, use a judge model from a different family than your agent. Our agent interacts with Claude, so we use Amazon Nova as the judge:
1
2
3
4
# Agent model - runs the actual workflow
AGENT_MODEL_ID = "us.anthropic.claude-sonnet-4-20250514-v1:0"
# Judge model - different family to avoid bias
JUDGE_MODEL_ID = "us.amazon.nova-pro-v1:0"Development Flow
Running evals during development helped us iterate faster. Change a prompt, swap a model, run the eval, and see if something regressed.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
----------------------------------------
TOOL CORRECTNESS
----------------------------------------
Score: 0.75
Reason: All expected tools were called. The agent selected storage
to read and write files, and journal to track workflow phases.
However, the agent could have potentially omitted some of the
journal calls or used a more direct approach, resulting in a
minor over-selection of tools.
----------------------------------------
TASK COMPLETION
----------------------------------------
Score: 0.9
Reason: Generated a comprehensive cost optimization report and
saved it to S3 along with evidence files, identifying significant
monthly savings. However, it did not explicitly mention loading
analysis results from storage as part of its process.
============================================================
PASSED
============================================================
1 passed in 95.75s (0:01:35)Evaluations in Production
Development evals with mocked data tells you the agent *can* work. But how do you know it's *still* working once deployed? Real traffic brings edge cases, unexpected inputs, and model drift that mocked tests won't catch.
There's also the non-determinism factor. Running an eval once doesn't tell you much because LLM outputs vary between runs. You need continuous evaluation across many sessions to get meaningful averages and spot patterns.
This is where AgentCore Evaluations comes in. It provides continuous, online evaluation of your agent against real production traffic.
Built-in Evaluators
AgentCore offers built-in evaluators that assess different dimensions of agent performance. For our cost optimization agent, we focus on three:
- ToolSelectionAccuracy: Did the agent select the appropriate tool for the task?
- GoalSuccessRate: Did the conversation successfully meet the user's goals?
- Correctness: Is the information in the agent's response factually accurate?
These evaluators use LLM-as-judge under the hood, scoring each session and surfacing patterns over time.
How It Works
AgentCore Evaluations integrates with your agent through OpenTelemetry traces. You configure which evaluators to run and set sampling rules for how much traffic to evaluate. The service then:
- Captures agent traces from production sessions
- Runs the configured evaluators against sampled sessions
- Stores scores and provides dashboards to investigate low-scoring session

Here's an example of what the evaluator feedback looks like for a passing production session:
1
2
3
4
5
6
7
8
9
10
11
12
GoalSuccessRate: "The user's goal was to analyze AWS costs and generate optimization
report. The AI assistant successfully discovered all 5 Lambda functions,
collected 30 days of CloudWatch metrics, analyzed memory usage, retrieved pricing
data, and generated comprehensive reports with specific recommendations..."
ToolSelectionAccuracy: "The agent systematically collected metrics for Lambda
functions discovered in the initial phase. The action is properly parameterized with
the correct function name, CloudWatch namespace, metric name, time range, and
statistics."
Correctness: "The analysis is data-driven, includes proper evidence, addresses gaps
and limitations, and provides actionable next steps."This is a happy path example. When sessions score lower, you get the same level of detail explaining what went wrong, helping you identify patterns and prioritize fixes.

AgentCore Evaluations Dashboard
Conclusion
You can introduce evals at any stage of agent development. In our case, we would have benefited from starting earlier. When you're iterating on prompts, swapping models, or refining tool definitions, evals give you fast feedback on whether those changes improve or break something.
For production, continuous evaluation becomes essential. LLM outputs are non-deterministic, so you need to evaluate across many sessions to get meaningful results and spot patterns. AgentCore Evaluations gives you that ongoing visibility into real agent behavior.
Built with Amazon Bedrock AgentCore and Strands. See the agentic cost optimizer for the full implementation, and AgentCore Evaluations for the managed evaluation service.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article