AWS Builder Center
Agents for Humans: How I Built LusiScan, a DevSecOps Agent That Knows When to Step Back

Agents for Humans: How I Built LusiScan, a DevSecOps Agent That Knows When to Step Back

Follow the full build journey of LusiScan, an autonomous DevSecOps agent that knows when to step back. Built with the Strands Agents SDK, Amazon Nova, and Amazon Bedrock AgentCore Runtime, this post covers the agent loop, honest production failures, and the philosophy of building AI that respects human judgment.

DevOps Engineer | AI/ML Engineer
Every Python developer knows the quiet dread of a dependency upgrade. A bot opens a PR bumping pydantic from 1.x to 2.x, and now you have to read a migration guide, rewrite a dozen class Config blocks, and hope the tests catch what you missed. The tedious part isn't the typing it's the judgment. Which upgrades are mechanical? Which ones need a human to actually think?
That question became LusiScan, my entry for the Agents for Humans hackathon: an autonomous agent that does the repetitive, judgment-heavy work of Python dependency upgrades and surfaces a decision to a human only when one is genuinely needed. It's built on the Strands Agents SDK , reasons with Amazon Nova  on Bedrock, and runs on Amazon Bedrock AgentCore Runtime. This is the story of building it — what worked, what broke in production, and the principle that kept it honest: an agent for humans should know the limits of its own confidence.

The shape of the problem

I scoped LusiScan around two contrasting upgrades against a controlled demo repo:
  • requests 2.31.0 → 2.34.2 — a safe, mechanical bump. The agent should just do it.
  • pydantic 1.10.13 → 2.x — a breaking major upgrade with class Config → model_config restructuring. The agent should not touch the code; it should flag the change and ask a human.

The Agent Architecture

LusiScan runs one cycle of four stages per invocation:
Monitor parses pyproject.toml and asks PyPI what's newer. Planner fetches the changelog and asks Nova for a structured plan. Executor applies safe, AST-aware refactors and opens a PR. Validator polls GitHub Actions for the real test result. Everything lands in DynamoDB as pending_review, and then this is the "for humans" part the Notifier reaches out to a person. A human makes the final call in a Streamlit panel; on the next cycle the agent reads that decision and acts on it merge, leave open, or close.

Reaching the human: the notifier

An agent that quietly writes a row to DynamoDB and waits isn't really asking anyone anything. So the last thing LusiScan does each cycle is send one plain, human-facing notification and the channel it uses says a lot about the design philosophy.
It tries Slack first: if a webhook is configured, it posts a short message ("high confidence and tests pass a PR is ready for a quick approval"). If there's no webhook, it falls back to a PR comment via the GitHub API. If neither is available, it degrades to a no-op rather than crashing the loop. The delivery is best-effort by design a notification failure must never bring down an autonomous agent:
1
2
3
4
5
6
7
8
9
def notify(migration, *, repo_name=None, slack_webhook=None, ...):
text = build_notice(migration) # message tailored to the tier
if slack_webhook: # 1) Slack incoming webhook
result = _send_slack(slack_webhook, text)
if result["sent"]:
return result
if repo_name and migration.get("pr_number"): # 2) PR-comment fallback
return _comment_on_pr(repo_name, migration["pr_number"], text)
return {"channel": "none", "sent": False} # 3) no-op, never crash
The message itself changes with the situation: ready-to-approve when confidence is high and tests pass, guided review when confidence is low, and "this one needs an architectural decision, so I left the code untouched" for the human-required case. During the demo I ran without a webhook and watched LusiScan drop a comment straight onto the PR the human got pinged exactly where they already work.

Why Strands SDK

I evaluated the usual heavyweight orchestration frameworks and kept bouncing off the same wall: they wanted to own my control flow. For a pipeline where order and determinism matter, you must scan before you plan, plan before you refactor, refactor before you validate I didn't want an opaque planner deciding what runs next.
Strands' model fit perfectly. A tool is just a plain function with a decorator, and you invoke it by calling it:
1
2
3
4
@tool
def scan_packages(repo_path: str) -> dict:
"""Monitor stage: detect outdated packages."""
return package_tools.scan_packages(repo_path)
No AgentExecutor, no hidden graph my orchestrator sequences the stages with ordinary Python calls. The tools return structured dicts, so the pipeline acts on data, not on free-form model text. That decision paid for itself many times over, every stage is unit-testable in isolation, and the LLM never gets to "decide" whether to skip validation.

Why Amazon Nova

Dependency planning is a high-volume, low-glamour reasoning task: for every outdated package, summarize a changelog and emit a small JSON plan (confidence, strategy, estimated_risk, breaking_changes). I didn't need a frontier model to write poetry I needed fast, cheap, structured reasoning that runs on a schedule without a scary bill. Nova's tiers mapped cleanly onto the work: Nova Pro for the migration plan (the one call where reasoning quality matters), Nova Lite for high-throughput changelog summarization. Because it's native to Bedrock, one converse wrapper talks to all tiers, and low temperature keeps the JSON deterministic enough to parse reliably.

How AgentCore simplified deployment

Here's where a hackathon usually turns into a DevOps slog. With Bedrock AgentCore Runtime, I wrote the loop, wrapped it in a BedrockAgentCoreApp, and pointed the CLI at it:
1
2
3
4
5
app = BedrockAgentCoreApp()

@app.entrypoint
def invoke(payload: dict) -> dict:
return run_for_repo(payload) # {"repo_name": "owner/repo"} -> one full cycle
1
2
3
agentcore configure --entrypoint src/agentcore_app.py --name lusiscan ...
agentcore launch --env AWS_REGION=us-east-1
agentcore invoke '{"repo_name": "hamdani2020/lusiscan-demo-repo"}'
No API Gateway, no container orchestration, no server to babysit. And crucially for an autonomous agent, secrets stay out of the image: the GitHub token and Slack webhook are read from Secrets Manager at invoke time, and an EventBridge Scheduler (provisioned in Terraform, least-privilege role) can fire the runtime on a cadence. The agent genuinely runs on its own.

What broke (the honest part)

The demo worked locally on the first try. Then I deployed, and reality arrived.
1. The @tool decorator crashed the container on startup. Locally, strands wasn't installed, so my @tool fell back to an identity decorator everything imported fine. In the container, the real Strands decorator ran at import time and tried to build a Pydantic schema from every function signature. Two of my stage functions had dependency-injection parameters typed as Protocol classes, and Pydantic v2 can't schematize those:
1
2
pydantic.errors.PydanticSchemaGenerationError:
Unable to generate pydantic-core schema for <class 'ChangelogModel'>
The container never started; agentcore invoke just returned a RuntimeClientError. The fix was to keep the injected model as a plain Any in the tool signature, those params were never meant to be LLM-facing inputs anyway.
Lesson: the environment where you test and the environment where you deploy can disagree about what your decorators do.
2. Decimal is not JSON serializable. This one only appeared after I added DynamoDB persistence. DynamoDB returns every number as a Decimal. The moment the agent read a previously-stored migration (with a pr_number) back into its JSON response, the runtime blew up with a 500. Stateless runs never hit it; the first stateful re-invoke did. I added a recursive deserializer at the store's read boundary:
1
2
3
4
5
6
7
8
def _deserialize(value):
if isinstance(value, Decimal):
return int(value) if value == value.to_integral_value() else float(value)
if isinstance(value, list):
return [_deserialize(v) for v in value]
if isinstance(value, dict):
return {k: _deserialize(v) for k, v in value.items()}
return value
3. A proxy that truncated one Docker layer. agentcore launch's ECR push kept dying with write: broken pipe on one large layer through Docker Desktop's proxy. After two failed retries I switched approaches entirely docker buildx build --push, whose BuildKit uploader handled the proxy cleanly. Lesson: if an approach fails twice, stop patching and change the approach.

The payoff

After the fixes, a single live invocation drove both scenarios end-to-end:
1
2
3
4
5
6
7
8
9
{
"status": "completed", "errors": [],
"migrations": [
{ "package": "requests", "target": "2.34.2",
"strategy": "guided_pr", "test_summary": {"status": "passed"}, "pr_number": 1 },
{ "package": "pydantic", "target": "2.13.5",
"strategy": "human_required", "pr_url": null }
]
}
requests got a PR with green CI. pydantic got flagged for a human, with no code touched exactly the judgment call I wanted the agent to make. Both triggered a notification the moment they landed in pending_review, so the human didn't have to go looking. They approve in Streamlit; on the next cycle the agent merges the PR after an explicit recorded approval, and never before.
That last guarantee never merge without a human's yes, is the soul of the project. LusiScan isn't trying to be the developer; it hands the developer only the decisions that are actually theirs. That's what building for humans means to me.
In the next post, I dig into the hardest technical piece: making AST-aware refactoring safe, and teaching the agent to know exactly when to keep its hands off the code.
Built with Strands Agents SDK, Amazon Nova, and Amazon Bedrock AgentCore Runtime. #Agents for Humans
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article