
Business Traces for Agents: Capturing Intent, Not Just Behaviour
OTEL tells you what your system did. Business traces tell you what your agent was trying to do. In this article, we show how to capture intent, not just behavior.
Tracing is critical for agent development. But when your agent fails silently, runs to completion with no errors, and produces wrong output, what do 100+ OTEL spans actually tell you?
We spent hours reading agent traces produced by OpenTelemetry while debugging an agentic cost optimizer . The spans were all green. The agent completed successfully. But the output was wrong. Somewhere in those 100+ spans, something went sideways. Finding where was the problem.
We needed a faster way to see which phase the agent was in when things went wrong. OTEL gives you technical traces: what the system did. We needed business traces: what the agent was trying to do.
Why Agents fail differently
Traditional services fail loudly. Exceptions, error codes, stack traces. Agents fail quietly. They complete successfully while producing wrong output.
This happens because you've handed control of execution flow to the agent. It decides which tools to call, how to interpret results, when to retry. These decisions are invisible to traditional infrastructure tracing. You're debugging a black box.
Technical traces answer "what did the system do?" with spans you have to hunt through. Business traces answer "what was the agent trying to do?" with entries you can read as a narrative.
Multi-agent orchestration
Single-agent systems can often get by with conversation memory as their audit trail. Multi-agent systems can't.
When something goes wrong, conversation memory fragments across agents. OTEL spans multiply. The orchestrator needs to understand global state: which agents have completed, which are stuck, what needs to be retried if something fails.
Infrastructure-level event journaling helps somewhat. Step Functions record events like
SESSION_INITIATED, AGENT_INVOCATION_STARTED, AGENT_INVOCATION_COMPLETED. This tells us the workflow executed and agents ran. But when an agent produces wrong output, these events don't help. We know the agents completed. We don't know what they were thinking.Infrastructure events (what we had):
1
2
3
4
5
6
7
[] SESSION_INITIATED
[] DISCOVERY_AGENT_INVOCATION_STARTED
[] DISCOVERY_AGENT_INVOCATION_COMPLETED
[] ANALYSIS_AGENT_INVOCATION_STARTED
[] ANALYSIS_AGENT_INVOCATION_COMPLETED
[] REPORTING_AGENT_INVOCATION_STARTED
[] REPORTING_AGENT_INVOCATION_COMPLETED Business traces (what we added):
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
[] TASK_DISCOVERY_STARTED
[] TASK_DISCOVERY_COMPLETED
[] TASK_USAGE_AND_METRICS_COLLECTION_STARTED
[] TASK_USAGE_AND_METRICS_COLLECTION_COMPLETED
[] TASK_ANALYSIS_AND_DECISION_RULES_STARTED
[] TASK_ANALYSIS_AND_DECISION_RULES_COMPLETED
[] TASK_RECOMMENDATION_FORMAT_STARTED
[] TASK_RECOMMENDATION_FORMAT_COMPLETED
[] TASK_COST_ESTIMATION_METHOD_STARTED
[] TASK_COST_ESTIMATION_METHOD_COMPLETED
[] TASK_REPORT_GENERATION_STARTED
[] TASK_OUTPUT_CONTRACT_STARTED
[] TASK_OUTPUT_CONTRACT_COMPLETED
[] TASK_S3_WRITE_REQUIREMENTS_STARTED
[] TASK_S3_WRITE_REQUIREMENTS_COMPLETED
[] TASK_REPORT_GENERATION_COMPLETEDInfrastructure events tell us which agents ran and for how long. Business traces tell us what each agent was doing and why. When the Analysis Agent's trace showed
date_range: "-0001-12-20 to 2025-01-19", we knew exactly where to look.Case study: how the pattern evolved
We developed this pattern while building the cost optimizer. The workflow uses two agents coordinated by Step Functions: a Discovery Agent finds resources across the AWS account, an Analysis Agent collects metrics and identifies optimization opportunities, and a Reporting Agent generates recommendations with estimated savings.
The workflow is triggered by EventBridge on a schedule or manually, with results saved to S3 and events tracked in DynamoDB. Here's how our journaling approach evolved through three stages.
Stage 1: simple status tracking
Our agent was timing out with "token limit reached" errors, but only sometimes. OTEL traces showed hundreds of spans per session. We knew that it failed, but not why some sessions succeeded and others didn't.
We broke the system prompt into meaningful phases and gave the agent a tool to record which phase it was in:
1
2
3
4
Phase: discovery → in_progress
Phase: discovery → completed
Phase: analysis → in_progress
...Immediately, we saw sessions getting stuck at "reporting → in_progress" and never completing. The journal pointed us to exactly where to look in OTEL logs.
Stage 2: immutable event stream
Initially, the journal tool updated a status field. This meant we only saw the final state. When sessions had inconsistent completion times, we couldn't tell which phase was slow.
We changed to append-only writes. Every status transition creates a new record. Now we could see:
1
2
3
4
5
6
[] discovery | in_progress
[] discovery | completed ← 45 seconds
[] analysis | in_progress
[] analysis | completed ← 16 seconds
[] reporting | in_progress
[] reporting | completed ← 4+ minutes!The reporting phase was the bottleneck. OTEL confirmed: the agent was making repeated failed attempts to call S3 APIs, each failure adding to the context window until it hit the token limit.
Stage 3: multi-layer journaling
When we moved to a serverless architecture with Step Functions, Lambda, and a background agent, we extended journaling beyond just the agent. Now the orchestrator, invoker, and agent all write to the same journal:
1
2
3
4
5
[] orchestrator | workflow_start | in_progress | "Starting cost analysis workflow"
[] invoker | agent_invoke | in_progress | "Triggering agent via Lambda"
[] invoker | agent_invoke | completed | {invocation_id: "..."}
[] agent | discovery | in_progress | "Discovering Lambda functions"
...This enabled the Step Function to query the journal and detect stuck agents, terminating early instead of waiting for hard timeouts.
Once we knew the reporting phase was the bottleneck, we could filter OTEL to just those spans and find the root cause in minutes.
Implementation: background capture
The obvious approach is giving agents a tool to record their own traces. We tried this first. It has problems.
Tools are on the critical path. Each write adds latency. Worse, agents can skip the tool entirely. Over long context windows, they "forget" to call it, especially when distracted by complex reasoning. You can prompt engineer around this, but you're fighting the model's attention.
We switched to background capture. Instead of asking the agent to trace itself, we intercept the agent's output stream and extract phase transitions automatically. The agent doesn't need to do anything. Capture is guaranteed.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
class BusinessTraceCapture:
def __init__(self, session_id: str, trace_store: TraceStore):
self.session_id = session_id
self.trace_store = trace_store
self.current_phase = None
# Keywords that indicate phase transitions
self.phase_patterns = {
"discovery": ["discovering", "scanning", "finding resources"],
"metrics_collection": ["fetching metrics", "collecting usage", "retrieving logs"],
"analysis": ["analyzing", "identifying opportunities", "evaluating"],
"report_generation": ["generating report", "creating recommendations", "writing output"],
}
async def wrap_stream(self, agent_stream):
"""Intercept agent stream, extract business events, pass through."""
buffer = ""
async for chunk in agent_stream:
buffer += chunk.content if hasattr(chunk, 'content') else str(chunk)
if phase := self._detect_phase_transition(buffer):
await self.trace_store.append_async(
session_id=self.session_id,
phase=phase.name,
intent=phase.intent,
status=phase.status,
timestamp=datetime.now()
)
self.current_phase = phase.name
buffer = ""
yield chunk
def _detect_phase_transition(self, text: str) -> Optional[PhaseTransition]:
"""Match keywords in agent output to detect phase changes."""
text_lower = text.lower()
for phase_name, keywords in self.phase_patterns.items():
if phase_name == self.current_phase:
continue # Already in this phase
for keyword in keywords:
if keyword in text_lower:
return PhaseTransition(
name=phase_name,
intent=text.strip(),
status="in_progress"
)
return NoneThe orchestrator wraps each agent call:
1
2
3
4
5
6
7
8
9
10
11
async def invoke_agent(self, task: str):
capture = BusinessTraceCapture(self.session_id, self.trace_store)
raw_stream = self.agent.stream(task)
traced_stream = capture.wrap_stream(raw_stream)
result = ""
async for chunk in traced_stream:
result += str(chunk)
return resultWhy observe externally instead of letting agents self-report?
This feels like traditional monitoring. You're parsing output from the outside rather than letting agents report their own state. Doesn't that defeat the agentic approach? We thought about this a lot. Agency matters for decisions: figuring out which tools to call, how to interpret results, when to retry. But recording "I'm now in the collecting phase" isn't a decision. It's bookkeeping. It doesn't benefit from reasoning.
When you give agents a tracing tool, you're asking them to do two jobs: solve the problem and maintain an audit trail. Those jobs compete for attention. Over long contexts, the model focuses on the actual task and the logging falls off. We saw this repeatedly. Agents would trace diligently for the first few phases, then get absorbed in complex reasoning and stop calling the tool. Background capture gives you less detail but more consistency. You won't have sessions where half the trace is missing because the agent got distracted.
When would you choose tool-based instead? When you need agents to capture why they made a decision, not just what they did. Background capture can see
evaluating | completed | {findings: 12} but can't capture "I filtered out 3 findings because they were below the $100/month threshold." If that reasoning needs to be in your audit trail, tool-based is worth the extra work.If you go the tool-based route, be explicit in your system prompt:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
JOURNALING REQUIREMENT:
You MUST use the update_journal tool to record your progress at each phase.
Journal at these points:
- Before starting each phase: status="in_progress"
- After completing each phase: status="completed" or "failed"
CRITICAL: Your intent should explain WHY you're making decisions, not just WHAT.
Bad: "Collecting data"
Good: "Fetching 30-day Lambda invocation metrics to identify underutilized functions"
Include decision rationale in context:
- selection_reason: Why you chose this approach
- alternatives_considered: Other options you evaluated
- error_details: What went wrong (for failures)
What you get
Faster debugging. Without business traces, you must scan spans, and cross reference timestamps to identify the problem. With them, you read a handful of entries and see the problem.
Stalled detection. The orchestrator queries the trace store and finds phases that have been "in_progress" past their timeout:
1
2
3
4
5
6
7
def check_stuck(session_id: str, agent_type: str, timeout_seconds: int) -> bool:
entries = trace_store.get_by_agent(session_id, agent_type)
in_progress = [e for e in entries if e['status'] == 'in_progress']
if not in_progress:
return False
started = datetime.fromisoformat(in_progress[-1]['timestamp'])
return (datetime.now() - started).total_seconds() > timeout_secondsPhase visibility. The orchestrator can show progress across all agents:
{"discovery/scanning": "completed", "analysis/collecting": "in_progress", ...}.Audit trails. Every session has a human-readable narrative. You can audit what data the agent accessed, what decisions it made, why it recommended one optimization over another.
When to use
This pattern fits multi-agent orchestration, long-running workflows, compliance requirements, and situations where you're spending too much time correlating OTEL spans.
It's overkill for simple single-turn agents, sub-second responses, or fully deterministic workflows where traditional logging works fine.
Relationship to other patterns
Business tracing is one point on a spectrum of agent observability patterns.
Event sourcing goes further. Instead of capturing phase transitions, you capture every streaming event from the agent. This enables replay and time-travel debugging, but generates more data and requires more infrastructure.
Durable execution goes further still. You checkpoint full agent state, not just intent. If the process crashes, you can resume from the last checkpoint. This is what you need for crash recovery, but it's significantly more complex to implement.
The saga pattern uses business traces for coordination. When a multi-agent workflow fails partway through, the orchestrator reads the trace store to understand what succeeded and what needs compensation. Business traces tell you what to roll back.
Conclusion
Traditional observability tells you what your system did. Conversation memory tells you what was said. Business traces tell you what your agent was trying to do.
That distinction matters when debugging autonomous systems that fail silently, get stuck in loops, or make decisions you don't understand.
The pattern is simple. Intercept agent output to extract phase transitions, or give agents a tool to record their narrative. Design meaningful phases. Store entries immutably. Query them for debugging and control.
Your agents are no longer black boxes. They're authors of their own story, and you can read it.
What's next
Traces tell you what happened. They don't tell you if it was correct. Part two covers how to build evaluations that do: The Missing Layer: Why Observability Alone Won't Save Your Agent
Built with Amazon Bedrock AgentCore and Strands. See the agentic cost optimizer for an example of this pattern.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article