
SWE-InfraBench + Kiro: What Happens When a Coding Agent Tackles IaC?
Can a coding agent solve cloud infrastructure problems that stump even the best large language models? We ran Kiro on the SWE-InfraBench benchmark to find out, and the results tell an interesting story about where agentic approaches help, and where they don't.
SWE-InfraBench + Kiro: What Happens When a Coding Agent Tackles IaC?
Authors: Natalia Tarasova , Enrique Balp-Straffon , Aleksei Iancheruk , Yevhenii Sielskyi , Nikita Kozodoi , Liam H. Byrne , Jack Butler, Dayuan Jiang , Marcin Czelej , Andrew Ang , Yash Shah , Roi Blanco , Sergei Ivanov
Can a coding agent solve cloud infrastructure problems that stump even the best large language models? We ran Kiro on the SWE-InfraBench benchmark to find out, and the results tell an interesting story about where agentic approaches help, and where they don't.
The problem: IaC is hard for LLMs
Infrastructure-as-Code (IaC) is how modern teams define cloud resources (servers, databases, networks) as software artifacts. AWS Cloud Development Kit (CDK) lets developers write infrastructure in Python, TypeScript, and other languages, bringing the full power of programming to cloud provisioning.
But generating correct IaC is surprisingly difficult for language models. Unlike typical coding tasks, it requires:
- Understanding complex resource dependencies (a Lambda function needs an IAM role which needs a policy which references an S3 bucket...)
- Knowing exact API shapes for specific library versions
- Integrating new code into existing codebases without breaking anything
SWE-InfraBench is a benchmark designed to test exactly this. It contains 100 tasks drawn from real-world CDK repositories, each requiring a model to implement a specific modification based on natural language instructions. Success is determined by passing unit tests that verify the generated CloudFormation output.
When the benchmark was published, even the best models struggled. Claude 3.7 Sonnet, the top performer, solved only 34% of tasks in a single attempt. A two-turn approach with error feedback pushed the best result to 65%, but that still meant over a third of infrastructure tasks remained unsolved.
We noted that agent-like approaches allowing models to iteratively attempt, analyze errors, and refine solutions could substantially improve performance. With the rise of coding agents like Kiro and Claude Code since the benchmark was published, we decided to put this to the test.
Evaluation setup
We ran Kiro , powered by Claude Sonnet 5, on all 100 SWE-InfraBench tasks in two configurations:
Two-turn (comparable to paper): The agent gets one attempt, receives test error output, then gets a second attempt. Exactly matches our original two-turn VH evaluation protocol.
Iterative (full agent loop): The agent can invoke a sandboxed test runner as many times as it wants within a 600-second timeout. Each call returns pytest output showing which tests pass and fail. The agent iterates until all tests pass or it decides to stop.
In both configurations, the agent cannot read test source code. It only sees pytest output (pass/fail counts and error tracebacks). This provides the same information as our high-verbosity (VH) two-turn configuration.
The agent has access to file read/write tools, shell access (for
pip install, py_compile, etc.), web search, and the AWS IaC MCP server for CDK documentation lookup.Results

| Model | Correctness |
|---|---|
| Kiro (Sonnet 5, iterative agent) | 82% |
| Kiro (Sonnet 5, two-turn agent) | 69% |
| Claude 3.5 V2 + Two-Turn + RAG (previous best) | 65% |
| Claude 3.7 + Two-Turn VH | 64% |
| GPT-4.1 + Two-Turn VH | 48% |
| Claude 3.7 Sonnet (one-turn) | 34% |
| DeepSeek R1 (one-turn) | 24% |
The iterative agent achieves 82% correctness, a 17 percentage point improvement over our previous best and nearly 2.5x the best single-turn model.
Where iterative testing helps
Looking beyond the headline correctness number, the shift in error distribution reveals what's actually happening.
We categorize failures into: syntax errors (CDK construct errors causing all tests to fail), logical errors (partially correct), and format errors (no output generated / timeout).
| No Error | Logical | Syntax | Format | |
|---|---|---|---|---|
| Kiro (iterative) | 82% | 3% | 12% | 3% |
| Kiro (two-turn) | 69% | 17% | 14% | 0% |
| Claude 3.7 Two-Turn VH | 64% | 24% | 11% | ~1% |
| Claude 3.7 One-Turn | 34% | 28% | 38% | 0% |
The key difference between two-turn and iterative: logical errors drop from 17% to 3%. When the agent can run tests repeatedly, it almost always converges on a fully correct solution from a partially correct starting point. The tight edit-test-fix loop is extremely effective for resolving remaining configuration issues.
What this means
The results confirm two key findings:
- Iterative self-verification is the single biggest lever for IaC generation. The progression tells the story clearly: 34% (one-turn) to 65% (two-turn with feedback) to 82% (iterative agent). Each step of added iteration improves performance substantially. The ability to try, fail, learn from errors, and try again mirrors how human developers actually work with cloud infrastructure.
- Coding agents are ready for real IaC workflows. The jump from 34% to 82% demonstrates that the core challenge in IaC generation wasn't model capability per se, but the lack of a feedback loop. When agents can verify their work iteratively against real test infrastructure, they reliably produce correct CDK code across a wide range of cloud services and patterns.
Resources
- SWE-InfraBench dataset: Kaggle
- Paper: Amazon Science
- Kiro: kiro.dev
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article