AWS Builder Center
Security Hub remediation plans: root-cause remediation, dissected

Security Hub remediation plans: root-cause remediation, dissected

On October 1, 2026, Security Hub gained remediation plans: instead of a queue of findings, one plan per root cause, with an impact assessment and instructions in CLI, Terraform, CloudFormation, Python, and CDK. I dissect the root-cause remediation pattern (the problem it solves, its anatomy, when to use it and when not) and the detail that deserves the most scrutiny: AI agents consuming plans through the API.

Senior Solutions Architect at Banco Itaú | Cloud, event-driven, AI agents & FinOps
The dangerous question in an exposure management program isn't "how many findings did we close this week?": it's "how many findings does a single fix close?". After 16 years operating financial platforms on AWS, I've watched entire teams measure the first and ignore the second. On October 1, 2026, AWS Security Hub launched remediation plans, which group exposure findings by shared root cause, and in doing so turned a console feature into what interests me here: a remediation architecture pattern worth dissecting.

The problem: the finding treadmill

For years the operational model for security on AWS was the individual finding: every CSPM control, every Inspector CVE, every GuardDuty detection becomes an item in the queue, and the team measures what the queue lets them measure: findings closed per week. The pattern AWS published in January 2020, automated response via custom actions, EventBridge, and Lambda, automated exactly that: one function per finding type, 20+ CIS Benchmark remediations, each treating a symptom. It works, and the queue grows back the following Monday, because the cause is still there.
The math runs against the symptom model. An IAM policy allowing kms:Decrypt with Resource: "*" is not one problem: it's a trait that shows up in every exposure finding of every principal carrying it. Replacing the policy is one change; closing the findings one by one is dozens of tickets describing the same change in different words.
Exposure findings, which Security Hub already generated before this announcement, covered half the distance: they correlate signals from GuardDuty, Inspector, Security Hub CSPM, and Macie into one finding per resource (a resource is primary in at most one exposure finding) with the attack path and blast radius visualized. Remediation plans close the other half: they invert the index. Instead of indexing the work by what was detected, they index it by what you need to fix. It's the difference between a list of symptoms and a diagnosis.

Anatomy: what ships inside a plan

Each remediation plan carries four things: the priority (Critical, High, Medium, or Low), an impact assessment for the change, the set of exposures it resolves or downgrades, and step-by-step instructions in five formats: AWS CLI, Terraform, CloudFormation, Python, and CDK. Security Hub orders the plans automatically: whatever reduces the most risk appears first. And the detail that changes automation design: AI agents can consume the plans programmatically through the API.
Underneath, the raw material is the exposure findings' traits. Misconfiguration traits are the familiar ones: IAM user without MFA, access keys unrotated for more than 90 days, unrestricted kms:Decrypt. Impact traits are the catalog of escalation paths Security Hub computes over effective permissions: trust policy hijack path, credential minting path, data ransomware path, disable audit trail path. Contextual traits come from IAM Access Analyzer: the documentation uses the example of a vulnerable EC2 instance whose role has 47 unused permissions across 5 services: that's the blast radius telling you this CVE is worth more than its CVSS suggests.
Programmatically, a finding's resources come out via GetFindingsV2. Cost: plans arrive at no additional charge under the Essentials plan, which is pay-as-you-go per resource unit (an EC2 instance counts as 1 unit, a Lambda function 1/12, an IAM resource 1/125) with a 30-day trial per account and Region.

The pattern: from N symptoms to 1 fix

Correlated signals become exposures; exposures sharing a root cause become one plan; the plan becomes one change in the IaC pipeline, and one change resolves N exposures.
  1. Signals: GuardDuty threat findings, Inspector CVEs and reachability, Security Hub CSPM control checks and Macie sensitive-data findings feed the correlation engine.
  2. Exposure findings: the engine correlates traits and attack paths into one exposure finding per resource.
  3. Remediation plan: exposures that share a root cause are grouped into a single plan, ranked by the risk it removes.
  4. Consumption: the security team (impact assessment + ticket) or an AI agent (via API) turns the plan into a pull request against the IaC pipeline; the agent's PR still needs human approval.
  5. One change: the pipeline applies the fix to the single root-cause resource (for example, an over-permissive IAM policy) in the change window, and N exposures resolve at once.

When to use it, and when not to

Use root-cause remediation when: you operate dozens of accounts with IaC-managed state, the finding backlog grows faster than the team closes it, and a formal change process exists (PCI-DSS, BACEN 4.893) that makes each fix expensive enough to be worth grouping. In that scenario the plan becomes the natural unit of work: one ticket, one PR, one window, N fewer exposures.
The step-by-step trap: the plan ships console and CLI instructions, and that's where drift lives. If the resource is managed by Terraform and someone applies the fix through the console, the next terraform apply silently reverts the correction: the exposure comes back with nobody touching anything. Of the five formats the plan offers, the only one that counts for a managed resource is the one that enters your pipeline: Terraform, CloudFormation, or CDK, via PR.
Don't use it as a replacement for the per-finding pattern everywhere. Narrow, reversible, urgent fixes (blocking public access on an S3 bucket, say) remain better served by the 2020 model: EventBridge, Lambda, seconds of latency. And if your estate is small, half a dozen accounts without consistent IaC, the cost of building the consumption pipeline outweighs the gain: there, the console's prioritized list, read by a human once a week, already delivers most of the value.

Three ways to consume remediation

CriterionPer-finding automation (2020)Plan via IaC pipelineAI agent via API
Unit of fix1 finding → 1 Lambda1 root cause → 1 PR1 plan → 1 proposed change
Best forNarrow, reversible fixes, seconds of latencyIaC-managed estates with formal change managementHigh volume with human review at the PR
Typical failure modeTreats the symptom; the queue regrowsToo slow for an active critical exposureThe privileged agent becomes the new escalation path

AI agents consuming plans: the detail that demands suspicion

The announcement says, almost in passing, that AI agents can consume plans through the API to automate fixes. Read alongside the Well-Architected Agent preview, the direction is clear: AWS is turning operational guidance into machine-readable artifacts. For platform designers that's good: a structured plan with an impact assessment is a far better contract for an agent than runbook prose.
But an agent that applies remediation is, by definition, a privileged principal modifying IAM, KMS, and network. Security Hub's own impact-trait catalog (credential minting path, disable audit trail path, remove restriction path) reads like the list of what a compromised, or merely confused, remediation agent could do. The cost here is not building the automation: it's maintaining, auditing, and containing a powerful principal over years.
My rule for financial-grade environments: the agent proposes, the pipeline disposes. The agent reads the plan via API, opens the PR with the Terraform snippet and the impact assessment in the body, and stops there. Its role carries a permission boundary with no iam:* and no kms:PutKeyPolicy, and every action leaves a CloudTrail trail on a rail it cannot switch off. Fully automatic application with no human in the loop only for change classes you've pre-approved by category, and never for an IAM policy change.

Root-cause remediation anti-patterns

  • Fixing through the console what Terraform manages: the next terraform apply silently reverts the fix and the exposure reappears with no visible change: use the plan's Terraform/CDK snippet, via PR.
  • A remediation agent with broad permissions: granting the plan-consuming agent iam:* creates exactly the credential minting path Security Hub exists to detect.
  • Measuring findings closed instead of risk reduced: the per-finding KPI rewards closing 40 symptoms one by one and punishes the single change that would resolve them all: the plan already ships risk-based ranking; make that the metric.
  • Applying a Critical plan outside a change window because it's "just security": tightening a security group or a policy in production breaks legitimate integrations: the plan's impact assessment exists to be read before the apply, not after the incident.
  • Keeping per-finding automation competing with plans: the 2020-era Lambda fixing the symptom while the plan's PR fixes the cause produces two concurrent changes on the same resource: pick the rail per resource class and switch the other off.
Treat the plan as PR input, not a console checklist: Of the five instruction formats, standardize on one (Terraform or CDK, whichever your pipeline already speaks) and make the plan the body of the PR: snippet, impact assessment, and the list of exposures it resolves. The console is for break-glass. The fix then inherits for free what your change management already has: review, history, rollback, and audit.
Curator's note: I'd start in a non-production account with the Essentials 30-day trial, routing only Critical plans into tickets with the Terraform snippet attached, and I'd measure one thing: exposures resolved per applied change. In the first month that number tells you whether root-cause grouping works on your estate or just becomes another list. The hard-won lesson behind my suspicion of agent-driven application: in financial environments, the fix that bypasses change management costs more than the exposure it fixes: I've seen a "quick" security group fix take down a settlement integration the exposure never threatened. The agent proposes the PR; a human approves IAM changes; always.

Verdict

Adopt the root-cause remediation pattern if you already run Security Hub Essentials: plans come at no additional cost, and the index inversion (from symptom to fix) is what was missing for the exposure backlog to become a real change queue. Conditions: fixes enter through the IaC pipeline (the plan's Terraform/CDK, never the console steps on a managed resource), the metric becomes risk reduced per change, and AI-agent consumption is limited to proposing PRs, with a permission boundary, no iam:*, and a human approving identity changes. Keep per-finding automation only for the narrow, reversible class. What decides this is not the feature: it's whether your change pipeline can absorb what the plan proposes.
Rating: adopt-with-guardrails

References


Originally published on October 2, 2026 at fernando.moretes.com , where I write about architecture, AWS and applied AI.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article