AWS Builder Center
Agents for Humans: Let an AI Negotiate Your SaaS Renewals While You Sleep

Agents for Humans: Let an AI Negotiate Your SaaS Renewals While You Sleep

An autonomous procurement agent that watches 14 subscriptions ($47,400/yr), flags renewals, benchmarks pricing, and negotiates vendor deals inside strict policy — built on Strands Agents SDK + Amazon Bedrock, with human approval before every vendor touch.

Agents for Humans: Meet Savr, the AI Procurement Employee Who Negotiates Your SaaS Renewals

Your startup has 18 people and 40+ SaaS subscriptions — Notion, Loom, PostHog, a dozen things nobody remembers signing up for. Renewals happen when someone remembers them. Prices go up silently. Seats go unused. And nobody on the team has time — or leverage — to negotiate any of it.
That spreadsheet is procurement. Until now.
Savr is an autonomous procurement employee for companies without a procurement team. It watches your stack, flags what needs attention, benchmarks prices, negotiates with vendors inside your policy — and never, ever contacts a vendor without your explicit approval.
Built for the AWS AI hackathon, on the Strands Agents SDK for TypeScript (v1.14.0), targeting Amazon Bedrock for live reasoning.

The problem no one owns

For a 5–20 person startup, SaaS spending sits in a blind spot. There's no procurement team, no leverage, no process — just a spreadsheet someone updates when things break. Renewals auto-renew at higher prices. Loom and Veed overlap and no one notices. Seats sit unused. Every dollar lost is a dollar no one ever sees.
The real cost isn't the money — it's the time spent not doing the thing the startup actually exists to build.

How Savr works

Savr runs one loop, over and over:
1
Observe → Reason → Propose → Policy Check → Execute or Human Gate → Persist
  • Observe — loads your subscriptions, company facts, and procurement policy.
  • Reason — an agent evaluates each subscription for renewals, unused seats, overlapping tools, and price increases.
  • Propose — recommends an action with evidence: benchmark pricing, competitor alternatives, and a clear cost model.
  • Policy Check — a deterministic policy runs after the agent, in code. The LLM decides what to do; code decides what it is allowed to do.
  • Execute or Human Gate — KEEP, DOWNGRADE, and CANCEL run autonomously. NEGOTIATE and SWITCH always require human approval first.
  • Persist — every decision, negotiation round, and realized saving is written to procurement memory for the next run.
Two modes drive it. Guardian monitors the whole stack and flags problems. Negotiator engages a vendor in structured rounds once — and only once — an approved card authorizes it. The LLM writes the words; code owns the financial boundaries, so a negotiation can never drift outside the approved budget.

Why Strands

The hackathon brief was explicit: Strands Agents SDK must orchestrate every agent path — never a fallback. That constraint turned out to be the design secret. Strands gives you the primitives agents actually need, all in one TypeScript SDK:
  • Agent — orchestrates Guardian and Negotiator modes; tools and hooks are first-class.
  • Custom tools — six of them: renewal checks, pricing benchmarks, alternatives search, vendor negotiation, vendor-response parsing, and policy check.
  • Tool calling — the agent decides what to research each run, bounded by a 12-call limit.
  • Structured output — the model emits GuardianOutput / NegotiationOutput validated with Zod, and app code builds the domain objects from them.
  • Hooks — a beforeToolCall guardrail enforces the blacklist, price bounds, and policy before any tool executes.
  • State — per-run invocation state threads the current vendor, round, and offer through the tools.
The result: the SDK's loop runs for real in every single path, and only the model is swapped. That seam makes a deterministic, zero-cost, judge-reproducible demo possible without ever bypassing the orchestrator.

What's real vs what's simulated

Transparency matters:
  • The demo path runs a deterministic in-process model under the same Strands orchestration — fully reproducible, zero cost.
  • Live mode targets Amazon Bedrock (anthropic.claude-sonnet-4-6, us-east-1) through the same Strands BedrockModel provider — identical code path, only the model changes.
  • Vendors are a sandbox — an Express server implementing the same message/counter/acceptance contract a real billing system would use.
  • Market research uses cached evidence by default, with an extensible live-search path.
  • Procurement memory persists to local JSON today; DynamoDB is the production intent.
  • The data is synthetic (Acme Corp, 18 employees, 14 tools) — but every package, card, negotiation round, and saving is produced by the actual Strands pipeline.
End-to-end: watches 14 subscriptions worth $47,400/yr, flags Notion (NEGOTIATE), Loom/Veed (SWITCH), runs the rest autonomously, waits for two approvals, negotiates Notion from $10,800 → $10,200 → $9,840, switches Loom, and lands on $5,760 in annual savings. Every number asserted by an automated E2E gate, not screenshotted.

From demo to real work on AWS

Savr is, right now, a demo-oriented build — every number was produced by the real Strands pipeline, but the deployment is a proof of architecture. The code ships with its own roadmap to a real product where Strands agents run on AWS, doing real procurement work for real startups:
Fargate, not a single process. Today one Node process boots the whole app; production moves to ECS Fargate behind an ALB, scaled on demand. The Dockerfile and CloudFormation template already ship in infra/. One task, one npx tsx src/api/server.ts, and AWS handles everything else — health checks, automatic restarts, zero-downtime deploys.
Amazon Bedrock, already wired. Today the demo uses a deterministic local model so judges get the same story every time; live mode already targets Claude Sonnet 4.6 through the same Strands BedrockModel. The production path is a model swap, not a rewrite — Strands handles the tool-calling loop, the hooks, the structured output; Bedrock handles the intelligence. Every agent call already routes through the same Converse API that production would use, with the same Anthropic Claude Sonnet 4.6 model the startup actually wants.
DynamoDB for procurement memory. Today local JSON; production persists every decision, negotiation, and saving to DynamoDB — the Strands DynamoDB storage adapter is the documented intent. Every run reads the full history, so Savr remembers what it negotiated last month, what it recommended, and what the human approved, across restarts and across deploys.
Real vendor APIs replace the sandbox. The sandbox proves the contract; production swaps in the real thing. Notion, Loom, PostHog — OAuth tokens in SecretsManager, real vendor endpoints, real rate limits. Strands handles the agent loop identically; only the tool's HTTP target changes.
EventBridge for autonomous triggers. Today an opt-in in-process loop; production schedules Guardian runs on Amazon EventBridge driving ECS tasks — once a day, once an hour, on whatever cadence the startup wants. Budget alerts are the cost guardrail, not the agent.
Auth, secrets, and real governance. Today a demo bearer token; production uses Amazon Cognito for auth, SecretsManager for API keys, and least-privilege IAM scoped to the exact model per region. The human gate — approval before any vendor contact — stays identical in demo and production.
Every step above is already stubbed, seamed, or documented in the repo — the architecture was built so "make it real" means adding integrations, not refactoring the agent loop.
One deliberate omission: Bedrock AgentCore was evaluated and set aside, so the Strands SDK stays the unambiguous orchestrator end-to-end — exactly as the brief requires.

What building an agent for humans taught me

  1. The human gate is the feature. An autonomous agent that contacts vendors on its own is terrifying. One that flags, proposes, negotiates after you approve — and refuses to apply a failed negotiation — is trustworthy. Code decides permission, the LLM decides strategy, humans decide authority.
  2. Deterministic policy in code, not in the prompt. Guardrails as code survive model swaps. The LLM can't "forget" a rule it never owned.
  3. Reproducible demos are a product decision. A demo that can't be re-run in front of a judge (or a customer) isn't a demo. Injecting the model kept Strands authentic and the story repeatable.
  4. Structured output + validation beats free text. Every recommendation is a typed, Zod-validated object flowing through the SDK's output schema — no prompt-parsing.

Try it

  • Repo: github.com/aalok101singh/savr  — public, MIT licensed, full spec docs, automated E2E (npm run demo:test).
  • Stack: Strands Agents SDK · Amazon Bedrock · TypeScript · Express · React/Vite · SSE for live events · local JSON persistence.
Procurement shouldn't be a spreadsheet. Savr is the employee who turns it back into a conversation — with your approval at every step.
Agents for humans: the agent does the work, the human keeps the authority.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article