
Agents for Humans and Making AI Financial Decisions Trustworthy: Cedar, Evidence Redaction, and HITL in Recoup
When an agent touches cloud spend, trust is architecture not marketing. Building Recoup: Agents for Humans and How We Designed a Safe AWS Recovery Workflow with Strands Graph covers our Strands SLA graph; this post is what the live hackathon demo enforces: sanitizer, Cedar (not LLM policy), and claim-bound HITL on Start Recovery.
When an agent touches cloud spend, “trust” is not a vibe, it’s architecture. Recoup treats evidence, policy, and approval as first-class gates before any recovery action.
Building Recoup: Agents for Humans and How We Designed a Safe AWS Recovery Workflow with Strands Graph describes our 11-node Strands graph on Bedrock for the SLA credit path in repo and CI. This post covers what the public judge demo (J-FULL) actually enforces when an operator clicks Start Recovery and reaches the approval card aligned with operator-journey.md
Live demo path (J-FULL)
On Start Recovery, Recoup runs a deterministic recovery assessment pipeline (evidence graph, confidence, safety checks) through the same policy semantics as production Cedar, then
REQUIRE_APPROVAL before ledger or SNS. That path prioritizes reliable latency on App Runner (RECOVERY_LLM_ON_PROMOTE=false by default); optional Bedrock on promote exists for depth demos. Approve closes the case in the Recovery Ledger with claim-bound fields, not unattended remediation on every resource in the public UI.Evidence sanitizer
Before model context sees customer data, we run an 8-pattern sanitizer (
backend/src/recoup/evidence/sanitizer.py): auth tokens, JWTs, API keys, cookies, emails, AWS account IDs, private IPs, and AWS secret patterns. Goal: zero PII in agent reasoning context.Cedar policy (deterministic gate)
The policy engine is not an LLM. Cedar rules in-repo define whether a proposed action is permitted, requires approval, or is forbidden including hard forbid on high-risk actions like
terminate_ec2_instance. On the App Runner demo, we evaluate the same semantics deterministically in Python (backend/src/recoup/safety/cedar.py); CDK/IAM remain AgentCore-oriented for production hardening.Policy source: infra/policy/recoup-policy.cedar
HITL with cryptographic binding
The approval card is the demo “wow moment”: operator sees why Recoup believes spend is unintended, safety checks, risk tier, exact action, and rollback, then Approve, Investigate further, or Decline.
Backend enforces binding:
claim_hash- digest of the recovery claimamount- exact USD/month at approval timestate_version -optimistic concurrency
Any mismatch → 409 Conflict (covered by Playwright SEC tests).
Try the live demo
No AWS keys required: https://pdkeexzwxr.us-east-1.awsapprunner.com/scan → Demo Scan → Start Recovery on three services → on opportunity detail, use Approve, Investigate further, or Decline and watch binding on tamper (Playwright SEC suite). Guest sessions:
X-Demo-Session header + sidebar Reset Demo Data.Step-by-step operator guide: docs/operator-journey.md
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article