AWS Builder Center
AI Found the Root Cause. Can It Fix It Safely?

AI Found the Root Cause. Can It Fix It Safely?

AWS DevOps Agent can investigate incidents. Discover how approved, pre-validated remediation workflows could make AI-powered operations safer and faster.

Cloud & DevOps Engineer

AI Found the Root Cause. Can It Fix It Safely?

Imagine receiving a production alert at 2 AM.
Your monitoring system detects an issue. An AI agent investigates logs, metrics, and traces. It identifies the likely root cause and recommends a fix.
Sounds like a major improvement for DevOps teams.
But one important question remains:
Should an AI agent that can diagnose a problem also be allowed to fix production automatically?

From investigation to remediation

AWS DevOps Agent helps investigate production incidents and generate findings. AWS has also published a technical approach for connecting those findings to remediation workflows using AWS Lambda Durable Functions, Amazon EventBridge, and Amazon Bedrock.
The important design principle is that investigation and remediation don't have to be the same operation.
An agent can investigate and recommend a fix while a separate workflow validates the proposed action and requests approval before execution.
That separation can help teams preserve human control without losing the speed of automation.

What a safer workflow could look like

Imagine this production incident workflow:
1. Detect
Amazon CloudWatch or another monitoring system identifies an incident.
2. Investigate
AWS DevOps Agent gathers evidence and identifies a probable root cause.
3. Prepare a fix
A controlled workflow translates the investigation findings into a proposed remediation.
4. Validate
Automated checks confirm that the action is allowed, the required conditions are met, and the proposed change is appropriate.
5. Request approval
The on-call engineer reviews the evidence, expected impact, and rollback plan.
6. Execute and verify
The approved action runs, and monitoring checks whether the incident has actually been resolved.
The goal isn't to remove engineers from incident response. It's to reduce repetitive investigation work while keeping consequential changes controlled.

Three safeguards I would prioritize

Least privilege: The remediation workflow should have only the permissions needed for its approved actions.
Idempotency and recovery: Repeated events should not accidentally apply the same change multiple times, and failed workflows should have a recovery strategy.
Verification and rollback: A successful API response doesn't necessarily mean the production incident is fixed. The workflow should verify system health and define what happens if the remediation fails.

Should AI be allowed to fix production?

I believe the right approach is graduated autonomy.
Low-risk, reversible actions with strong validation may eventually be automated. Changes affecting critical services, security boundaries, or customer data deserve stricter approval requirements.
The level of autonomy should depend on the potential impact—not simply on how confident the AI sounds.
AI-powered operations will be most useful when they combine fast investigation with controlled execution and measurable results.
My question for the AWS community:
Would you allow an AI agent to automatically remediate a low-risk production incident, or would you require human approval for every change?
Where would you draw the line?
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article