AWS Builder Center
Introducing Kimi K3 on Amazon Bedrock: Long-Context AI for Coding and Knowledge Work

Introducing Kimi K3 on Amazon Bedrock: Long-Context AI for Coding and Knowledge Work

Amazon Web Services has announced the availability of Kimi K3 from Moonshot AI on Amazon Bedrock, giving developers another open-weight model option for coding and knowledge-work workloads.

Announced on September 18, 2026, Kimi K3 combines native vision capabilities, a 1-million-token context window, and explicit prompt caching. These capabilities are designed for workflows that need to work with large repositories, documents, images, and long-running coding or knowledge tasks.
For developers, AI engineers, DevOps teams, and platform engineers, this is particularly interesting because Amazon Bedrock provides a managed way to access the model while integrating it into existing AWS applications and AI workflows.

πŸš€ What Is Kimi K3?

Kimi K3 is an open-weight model from Moonshot AI that is now available through Amazon Bedrock.
AWS describes Kimi K3 as Moonshot AI's most capable open-weight model, combining native vision with a 1-million-token context window for long-running coding and knowledge workflows.
The model supports:
  • 🧠 Long-context reasoning
  • πŸ‘οΈ Native vision
  • πŸ’» Coding workflows
  • πŸ“š Large-document analysis
  • ⚑ Explicit prompt caching
  • πŸ› οΈ Tool calling
  • πŸ“¦ Structured outputs
Kimi K3 is available through the Amazon Bedrock runtime APIs, including the OpenAI-compatible Responses and Chat Completions APIs, as well as Converse and Invoke. AWS recommends the OpenAI-compatible APIs for new Kimi K3 applications.

🧠 Why a 1-Million-Token Context Window Matters

One of the biggest challenges when building AI applications is handling large amounts of context.
Consider a large software project:
1
2
3
4
5
6
7
8
9
10
11
12
13
GitHub Repository
β”‚
β”œβ”€β”€ Frontend
β”œβ”€β”€ Backend
β”œβ”€β”€ Terraform
β”œβ”€β”€ Kubernetes
β”œβ”€β”€ Docker
β”œβ”€β”€ CI/CD
β”œβ”€β”€ Documentation
└── Configuration
β”‚
β–Ό
AI Model
With smaller context windows, developers often need to:
  • Split information into smaller sections
  • Summarize documents
  • Retrieve only selected files
  • Build complex RAG pipelines
  • Maintain external memory
Kimi K3 provides a 1-million-token context window, allowing applications to maintain significantly larger amounts of context during a workflow.
This is particularly useful for:

πŸ’» Large Codebases

An AI coding assistant can potentially reason over a much larger portion of a repository rather than relying exclusively on small snippets.

πŸ“š Large Documents

Teams can provide extensive documentation, specifications, and technical references as part of an AI workflow.

πŸ—οΈ Infrastructure Projects

DevOps engineers can combine:
1
2
3
4
5
6
7
8
9
10
11
Terraform
+
Kubernetes YAML
+
Dockerfiles
+
CI/CD
+
Logs
+
Architecture Documentation
and ask the model to analyze relationships between them.

πŸ‘οΈ Native Vision Capabilities

Kimi K3 isn't limited to text.
The model supports image input, allowing applications to combine visual and textual information.
For example:
1
2
3
4
5
6
7
8
9
10
11
AWS Architecture Diagram
β”‚
β–Ό
Kimi K3
β”‚
β–Ό
Architecture Analysis
β”‚
β”Œβ”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”
β–Ό β–Ό β–Ό
Security Cost Reliability
Potential use cases include:
  • AWS architecture diagrams
  • Kubernetes diagrams
  • Monitoring dashboards
  • Application screenshots
  • CI/CD pipeline screenshots
  • Error screenshots
  • Technical diagrams
  • UI analysis
Imagine providing an architecture diagram together with Terraform code and asking:
"Identify inconsistencies between this architecture and the Terraform configuration."
That creates an interesting multimodal DevOps workflow.

⚑ Explicit Prompt Caching

Another important feature is explicit prompt caching.
AWS identifies Kimi K3 as an open-weight model on Amazon Bedrock that supports explicit prompt caching.
Why is this important?
Imagine an AI coding assistant that repeatedly sends the same information:
1
2
3
4
5
6
7
8
9
System Instructions
+
Coding Standards
+
Repository Documentation
+
Tool Definitions
+
Architecture Context
Only the user's question changes:
1
2
3
4
Request 1 β†’ Fix authentication
Request 2 β†’ Explain database connection
Request 3 β†’ Optimize API
Request 4 β†’ Fix Kubernetes deployment
Without caching, the application may repeatedly process the same large prompt context.
With prompt caching:
1
2
3
4
5
6
7
8
Large Static Context
β”‚
β–Ό
Cache
β”‚
β”Œβ”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”¬β”€β”€β”€β”€β”
β–Ό β–Ό β–Ό β–Ό
Req1 Req2 Req3 Req4
This can help reduce latency and input costs when the same context is reused.
AWS documentation states that Kimi K3 supports explicit prompt caching with a minimum cache checkpoint of 1,024 tokens and a cache retention period of at least 30 minutes.

πŸ—οΈ Kimi K3 on Amazon Bedrock Architecture

A simple architecture could look like this:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
Developer
β”‚
β–Ό
Application / API
β”‚
β–Ό
Amazon Bedrock
β”‚
β–Ό
Kimi K3
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό β–Ό
Text Images Documents
β”‚ β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β–Ό
AI Response
β”‚
β–Ό
Application / Agent
For an enterprise AI application, this could be extended with:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
Users
β”‚
β–Ό
Web Application
β”‚
β–Ό
API Gateway
β”‚
β–Ό
Lambda
β”‚
β–Ό
Amazon Bedrock
β”‚
β–Ό
Kimi K3
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό β–Ό
S3 Knowledge Base Tools
β”‚ β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β–Ό
AI Response
This allows Kimi K3 to become one component of a larger AWS-native AI architecture.

πŸ€– Kimi K3 for AI Agents

The combination of:
1
2
3
4
5
6
7
8
9
Large Context
+
Vision
+
Tool Calling
+
Structured Outputs
+
Prompt Caching
makes Kimi K3 interesting for agentic applications.
A simplified AI agent architecture could look like:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
User
β”‚
β–Ό
AI Agent
β”‚
β–Ό
Kimi K3
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό β–Ό
Search AWS Database
Tool Tool Tool
β”‚ β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β–Ό
Result
β”‚
β–Ό
User
For example, an AI DevOps assistant could receive:
1
2
3
4
5
6
7
Application Logs
+
Terraform
+
Kubernetes Manifests
+
Architecture Diagram
and then use tools to inspect AWS resources.
The model handles reasoning while AWS services and external tools provide the actual data and actions.

πŸ’» Example: Calling Kimi K3 from Amazon Bedrock

AWS documentation provides an OpenAI-compatible approach for Kimi K3.
A simplified Python example looks like this:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
from aws_bedrock_token_generator import provide_token
from openai import OpenAI

region = "us-west-2"

client = OpenAI(
api_key=provide_token(region=region),
base_url=f"https://bedrock-runtime.{region}.amazonaws.com/openai/v1",
)

response = client.responses.create(
input="Explain how Amazon EKS works.",
model="global.moonshotai.kimi-k3",
)

print(response.output_text)
The Kimi K3 model ID for the global inference profile is:
1
global.moonshotai.kimi-k3
AWS also provides the corresponding Bedrock runtime model ID:
1
moonshotai.kimi-k3
for programmatic access.

🌍 Regional Availability

Kimi K3 is available through Amazon Bedrock using US Geo cross-Region inference and Global cross-Region inference.
The current AWS documentation lists global availability through supported commercial AWS Regions, including Mumbai (ap-south-1) and Hyderabad (ap-south-2).
This is particularly interesting for developers building applications in India because Kimi K3 can be accessed through Amazon Bedrock's global inference profile from AWS infrastructure in India.
However, organizations should always check the latest AWS regional availability and data-routing requirements before deploying production workloads.

πŸ” Security and Governance

One of the advantages of using Kimi K3 through Amazon Bedrock is that the model becomes part of an AWS-managed application architecture.
A production environment could integrate:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
IAM
β”‚
β”œβ”€β”€ Authentication
β”œβ”€β”€ Authorization
└── Least Privilege
β”‚
β–Ό
Amazon Bedrock
β”‚
β–Ό
Kimi K3
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”
β–Ό β–Ό β–Ό
CloudWatch CloudTrail Application Logs
Organizations can combine Bedrock with existing AWS security and governance controls.
For production AI applications, important considerations include:
  • IAM permissions
  • Data access controls
  • Logging
  • Monitoring
  • Cost controls
  • Prompt security
  • Sensitive-data handling
  • Model access governance

πŸ’° Understanding Prompt Caching and Cost

Large-context AI applications can generate significant input-token usage.
This is where prompt caching becomes especially interesting.
Consider a workload where the application repeatedly sends:
1
2
3
4
5
6
7
500K tokens
+
500K tokens
+
500K tokens
+
500K tokens
because the same repository or documentation is included in every request.
Caching can allow reusable context to be handled more efficiently.
AWS currently lists separate pricing for:
  • Input tokens
  • Output tokens
  • Cache reads
  • Cache writes
for Kimi K3.
The actual economics depend on your workload, cache hit rate, request pattern, and service tier.
Therefore, developers should measure:
1
2
3
4
5
6
7
8
9
Context Size
+
Cache Hit Rate
+
Request Frequency
+
Output Tokens
=
Actual AI Cost

πŸ”§ Practical DevOps Use Case

Let's build a conceptual AI DevOps Assistant.

Input

1
2
3
4
5
6
7
8
9
10
11
Terraform Repository
+
Kubernetes Manifests
+
Dockerfiles
+
CI/CD Pipeline
+
CloudWatch Logs
+
AWS Architecture Diagram

AI workflow

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
Developer
β”‚
β–Ό
AI DevOps Assistant
β”‚
β–Ό
Amazon Bedrock
β”‚
β–Ό
Kimi K3
β”‚
β”œβ”€β”€ Analyze Code
β”œβ”€β”€ Analyze Diagram
β”œβ”€β”€ Analyze Logs
└── Reason Across Context
β”‚
β–Ό
Tool Calling
β”‚
β”Œβ”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”
β–Ό β–Ό β–Ό
AWS GitHub Kubernetes
β”‚ β”‚ β”‚
β””β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”˜
β–Ό
Response
The assistant could answer questions such as:
1
2
3
4
5
6
7
8
9
Why is my Kubernetes deployment failing?

Which AWS resource is causing the error?

Does my Terraform configuration match the architecture?

What security issue exists in this deployment?

What could be optimized in this infrastructure?
This is where long-context multimodal models can become useful beyond simple chatbot interactions.

🧩 Kimi K3 vs Traditional RAG Workflows

A large context window does not mean RAG becomes unnecessary.
Instead, the two approaches can complement each other.

Traditional RAG

1
2
3
4
5
6
7
8
9
10
11
12
13
Documents
β”‚
β–Ό
Embeddings
β”‚
β–Ό
Vector Database
β”‚
β–Ό
Retrieve Relevant Content
β”‚
β–Ό
LLM

Large-context workflow

1
2
3
4
5
6
7
Large Repository / Documents
β”‚
β–Ό
Kimi K3
β”‚
β–Ό
Analysis

Hybrid architecture

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
Documents
β”‚
β–Ό
Vector Store
β”‚
β–Ό
Relevant Context
β”‚
β–Ό
Kimi K3
β–²
β”‚
Additional Context
β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ β”‚
Images Repository
For very large knowledge bases, retrieval can still be useful for controlling context size, relevance, latency, and cost.

⚠️ Important Considerations

Kimi K3 has some current limitations that developers should understand.
According to AWS documentation:
  • Video inputs are not supported.
  • AWS recommends OpenAI-compatible APIs for new Kimi K3 applications.
  • Converse has known limitations for some multi-turn reasoning workflows.
  • Knowledge Bases integration is not currently supported for Kimi K3.
  • For combined image and text inputs, AWS recommends placing images before text for potentially better results.
  • Explicit prompt caching is supported through the Responses and Chat Completions APIs.
These limitations are important when deciding how to integrate the model into an existing AI application.

πŸš€ A Practical Project Idea

Build an AI-Powered DevOps Repository Analyzer

Create a project that combines:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
GitHub
β”‚
β–Ό
Repository
β”‚
β”œβ”€β”€ Terraform
β”œβ”€β”€ Kubernetes
β”œβ”€β”€ Docker
β”œβ”€β”€ GitHub Actions
└── Documentation
β”‚
β–Ό
AI Analyzer
β”‚
β–Ό
Amazon Bedrock
β”‚
β–Ό
Kimi K3
β”‚
β”Œβ”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”
β–Ό β–Ό β–Ό
Code Images Logs
β”‚ β”‚ β”‚
β””β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”˜
β–Ό
AI Report
The final report could include:

Infrastructure Analysis

1
2
3
4
5
6
VPC
IAM
EKS
RDS
S3
Load Balancer

DevOps Analysis

1
2
3
4
5
CI/CD
Docker
Terraform
Kubernetes
GitHub Actions

Security Analysis

1
2
3
4
IAM permissions
Secrets
Network exposure
Container configuration

Reliability Analysis

1
2
3
4
High availability
Scaling
Monitoring
Failure recovery
This would be an excellent practical project for demonstrating how Generative AI and DevOps can work together on AWS.

🎯 Key Takeaways

Kimi K3 on Amazon Bedrock brings several interesting capabilities to AWS AI developers:
1
2
3
4
5
6
7
8
Kimi K3
β”‚
β”œβ”€β”€ 1M-token context
β”œβ”€β”€ Native vision
β”œβ”€β”€ Prompt caching
β”œβ”€β”€ Tool calling
β”œβ”€β”€ Structured outputs
└── Long-running coding & knowledge workflows
The bigger story isn't simply another model becoming available on Amazon Bedrock.
The more interesting development is the combination of:
Long context + multimodal input + caching + managed AWS infrastructure.
For developers, this creates new possibilities for:
  • AI coding assistants
  • Repository analysis
  • Enterprise knowledge applications
  • Multimodal AI assistants
  • AI agents
  • Large-document analysis
  • DevOps automation
  • Infrastructure analysis

🌟 Final Thought

The evolution of AI applications is moving beyond simple question-and-answer systems.
We're moving toward AI systems that can understand:
1
2
3
4
5
6
7
8
9
10
11
Code
+
Documentation
+
Images
+
Infrastructure
+
Logs
+
Tools
and reason across all of them.
With Kimi K3 now available through Amazon Bedrock, developers have another model option for building these long-context and multimodal applications while using AWS infrastructure and services around it.
For DevOps and cloud engineers, this creates an exciting direction:
What if your AI assistant could understand your entire infrastructureβ€”not just a single file?
That is where long-context AI can become a powerful part of the modern cloud-native engineering workflow.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article