Agents for Humans: Deterministic Guardrails in Strands
A synthetic, reproducible scoring example shows why fixed arithmetic, source evidence, and calibrated claims matter in a Strands agent.
This is an implementation note from RewardRadar for Agents for Humans. The project and article were prepared with substantial OpenAI Codex assistance.
A model can explain a tradeoff fluently while still using bad arithmetic or unsupported evidence. RewardRadar separates those failure modes: Strands orchestrates specialists and tools, while ordinary Python computes the decision. Neither layer should pretend that a heuristic score is proof of payment.
A small deterministic boundary
The Candidate dataclass contains explicit fields for advertised payout, effort, canonical status, competitors, lock state, escrow, sponsor verification, acceptance clarity, repository activity, payout readiness, deadline, and an optional base probability. The assess_candidate function returns probability, expected value, expected hourly value, confidence, verdict, and reasons.
With ordinary finite inputs, the implemented arithmetic is payment probability multiplied by advertised payout, then divided by estimated effort with a 0.5-hour denominator floor. A non-open canonical status sets probability to zero. Lock state, inactive repositories, competitors, and a short deadline reduce it. The result is bounded between zero and 0.92. These constants express a policy choice; they are not learned estimates from a validated payment dataset.
A reproducible synthetic example
On September 14, the current scoring implementation was rerun with a synthetic $200 opportunity, four hours of effort, and an explicitly assumed base probability of 0.10. Escrow, sponsor verification, clear acceptance, and payout readiness were set true as hypothetical inputs. This is not a real earning opportunity or a payment forecast.
- Open, zero competitors: probability 0.10; expected value $20.00; expected hourly value $5.00; verdict pursue.
- Open, two competitors: reported probability 0.0676; expected value $13.51; expected hourly value $3.38; verdict watch.
- Closed, zero competitors: probability 0; expected value $0; hourly value $0; verdict avoid.
The competitor adjustment is 1 / (1 + 0.24 × claimers). Monetary outputs are rounded for display after computation. When base_probability is explicitly supplied, optimistic bonuses from sponsor, acceptance, and payout-ready flags do not inflate that starting assumption; penalties can still reduce it.
A useful distinction: confidence is not payment probability
All three synthetic cases returned confidence 0.92, including the closed case. That is not a contradiction in the implementation: confidence is a separate, rule-based indicator derived from the presence of certain evidence fields. It is also not a statistically calibrated confidence interval. A high confidence display must never be read as a 92% chance of receiving money. This example is a practical reason to label every metric and expose its inputs.
What Strands contributes
The risk specialist can call the scoring tool, and the ranking specialist can compare its outputs. A model cannot rewrite the arithmetic inside that Python function through a prompt. It can still supply incorrect Candidate fields, omit contrary evidence, or produce misleading prose afterward. Source verification, constrained inputs, and human review remain necessary.
The credential-free demonstration runs the actual four-node Strands graph with a deterministic DemoModel and prescribed fixture tool calls. This tests graph construction, handoffs, tools, and outputs without cloud credentials. It does not prove open-ended reasoning or a hosted AWS deployment. The separate Bedrock runner is included, but no completed Bedrock invocation or AgentCore deployment is claimed.
Tests should target failure modes
The repository includes tests for closed items, increasing competition, ordering by adjusted hourly return, explicit probabilities, and graph tool execution. The tested source baseline passed 22 Python tests. Those checks support specific behavior in those cases, not a universal safety or accuracy guarantee.
Future hardening should test invalid numeric inputs, corrupted evidence, stale timestamps, and adversarial tool responses. The current dataclass is not a comprehensive schema validator, and deterministic calculations do not make untrusted inputs trustworthy. A deployment should validate finite, nonnegative numeric values and canonical evidence before ranking.
Reproduce and inspect
The scoring module is independent of cloud access. Inspect Candidate and assess_candidate in agent/core.py, then run the repository tests in a virtual environment using the pinned requirements. The public UI remains a dated snapshot; the synthetic example above was a direct local calculation, not a live marketplace recommendation.
- Exact scoring implementation
- Tests at the submitted source commit
- Companion article: Strands architecture and execution boundaries
- RewardRadar submission
Disclosure: substantial OpenAI Codex assistance was used for the project, example verification, and article. All monetary example values are synthetic assumptions. No award, received payment, or completed AWS runtime is claimed.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article