AWS Builder Center
Agents for Humans and Making AI Financial Decisions Trustworthy: Cedar, Evidence Redaction, and HITL in Recoup

Agents for Humans and Making AI Financial Decisions Trustworthy: Cedar, Evidence Redaction, and HITL in Recoup

When an agent touches cloud spend, trust is architecture not marketing. Building Recoup: Agents for Humans and How We Designed a Safe AWS Recovery Workflow with Strands Graph covers our Strands SLA graph; this post is what the live hackathon demo enforces: sanitizer, Cedar (not LLM policy), and claim-bound HITL on Start Recovery.

When an agent touches cloud spend, “trust” is not a vibe, it’s architecture. Recoup treats evidence, policy, and approval as first-class gates before any recovery action.
Building Recoup: Agents for Humans and How We Designed a Safe AWS Recovery Workflow with Strands Graph  describes our 11-node Strands graph on Bedrock for the SLA credit path in repo and CI. This post covers what the public judge demo (J-FULL) actually enforces when an operator clicks Start Recovery and reaches the approval card aligned with operator-journey.md 

Live demo path (J-FULL)

On Start Recovery, Recoup runs a deterministic recovery assessment pipeline (evidence graph, confidence, safety checks) through the same policy semantics as production Cedar, then REQUIRE_APPROVAL before ledger or SNS. That path prioritizes reliable latency on App Runner (RECOVERY_LLM_ON_PROMOTE=false by default); optional Bedrock on promote exists for depth demos. Approve closes the case in the Recovery Ledger with claim-bound fields, not unattended remediation on every resource in the public UI.

Evidence sanitizer

Before model context sees customer data, we run an 8-pattern sanitizer (backend/src/recoup/evidence/sanitizer.py): auth tokens, JWTs, API keys, cookies, emails, AWS account IDs, private IPs, and AWS secret patterns. Goal: zero PII in agent reasoning context.

Cedar policy (deterministic gate)

The policy engine is not an LLM. Cedar rules in-repo define whether a proposed action is permitted, requires approval, or is forbidden including hard forbid on high-risk actions like terminate_ec2_instance. On the App Runner demo, we evaluate the same semantics deterministically in Python (backend/src/recoup/safety/cedar.py); CDK/IAM remain AgentCore-oriented for production hardening.

HITL with cryptographic binding

The approval card is the demo “wow moment”: operator sees why Recoup believes spend is unintended, safety checks, risk tier, exact action, and rollback, then Approve, Investigate further, or Decline.
Backend enforces binding:
  • claim_hash- digest of the recovery claim
  • amount- exact USD/month at approval time
  • state_version -optimistic concurrency
Any mismatch → 409 Conflict (covered by Playwright SEC tests).

Try the live demo

No AWS keys required: https://pdkeexzwxr.us-east-1.awsapprunner.com/scan  → Demo Scan → Start Recovery on three services → on opportunity detail, use Approve, Investigate further, or Decline and watch binding on tamper (Playwright SEC suite). Guest sessions: X-Demo-Session header + sidebar Reset Demo Data.
Step-by-step operator guide: docs/operator-journey.md 
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article