Agents for Humans: Building a Strands Agent That Distrusts Its Marketplace
How RewardRadar separates discovery, canonical verification, deterministic scoring, and a reproducible Strands graph without mistaking a static demo for an AWS deployment.
RewardRadar was built for the Agents for Humans hackathon with substantial OpenAI Codex assistance. This article explains the implemented Strands graph, not an unverified AWS deployment.
The fastest way to make a reward-finding agent look impressive is to display a large number. The fastest way to make it useful is to challenge that number. Independent developers need to know whether a task deserves their limited implementation time, not merely whether somebody advertised a reward.
Separate discovery from verification
RewardRadar uses four specialist nodes: Scout, Verifier, Risk Analyst, and ROI Ranker. Scout normalizes marketplace inventories. Verifier checks canonical issue or competition records. Risk Analyst identifies payment friction, crowded claims, unclear acceptance criteria, and costly participation requirements. ROI Ranker uses deterministic scoring to produce a pursue, watch, or avoid recommendation.
These are separate responsibilities because a marketplace row can outlive the issue it references. In the committed September 11, 2026 capture, 30 advertised rows yielded only eight canonical GitHub issues that were still open. The other 22 were closed, deleted, or unverifiable. One apparently unclaimed $1,500 row led to an HTTP 410 response. That dated observation is a regression case, not a claim about the marketplace today.
What actually runs
The implementation constructs four Strands Agent instances and connects them through GraphBuilder. Edges run from scout to verifier to risk to roi. The graph sets a maximum of eight node executions and an execution timeout of 120 seconds. Those SDK limits are configured safeguards, not a claim that every external operation can be interrupted at an exact wall-clock boundary.
There are two explicit paths. The credential-free DemoModel supplies one prescribed fixture tool call per specialist. This runs the real Strands graph and records tool results, but it is a deterministic reproduction adapter, not an open-ended foundation model. A separate Bedrock runner accepts an authorized AWS credential chain and only reports completion after Strands returns Status.COMPLETED. The project has no recorded completed Bedrock invocation or AgentCore deployment.
The public dashboard is another boundary: it replays a dated evidence snapshot and lets readers inspect rows and filters. It does not launch the Python graph or refresh marketplaces. The repository's separate capture scripts perform live-source collection when an operator explicitly runs them. Keeping these modes separate prevents an attractive UI from being mistaken for a cloud-runtime demonstration.
Tools provide evidence, not authority by assertion
The live verifier tools query canonical public APIs and return source state and provenance. Specialist prompts instruct the model not to fabricate evidence, but a prompt is not a security boundary: a model can still supply bad input fields or synthesize an unsupported conclusion. Source URLs and deterministic inputs therefore remain reviewable.
The scoring function is ordinary Python. A non-open status yields zero payment probability; crowding and other risks apply disclosed adjustments. The formulas are reproducible, but their probabilities are heuristics rather than calibrated predictions. No displayed amount, recommendation, submission, or possible prize is counted as received money.
Reproduce the boundary before adding a model
Clone the public repository, create a Python virtual environment, install requirements.txt, and run
python -m agent.demo. Then run python -m unittest discover -s tests -v. These commands exercise the credential-free path. Read the README before running live capture scripts, because current external evidence will naturally differ from the saved snapshot.A hosted-model run requires separately activated AWS access and authorized credentials and can incur charges. AWS account creation and an AWS Builder ID are not proof that Bedrock is enabled. This project's incomplete Bedrock activation remains explicitly disclosed rather than being hidden behind a successful local test.
The design lesson
A professional agent should absorb repetitive evidence gathering without taking responsibility away from the human for identity, legal agreements, and payment decisions. An evidence-backed rejection can save more time than another enticing lead. The next useful improvement would be a documented hosted-model run and labeled evaluation cases, followed by stronger sponsor-payment evidence—not a larger unverified reward total.
Code and evidence
- Four-node Strands implementation
- Dated source capture
- Public snapshot demo
- Submitted project and video
Disclosure: the project and this article were prepared with substantial OpenAI Codex assistance. The project uses the open-source Strands Agents SDK. No completed Bedrock run, AgentCore deployment, award, or payment is claimed.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article