
From Prompt Engineering to Context Engineering: Building Agents That Actually Work at Scale
Building an AI agent that demos well is easy. Keeping it reliable in production is not. When agents break at scale, the model is rarely the problem — what the model receives is. Part 1 of this series introduces context engineering: the discipline of managing what information reaches your agent, when, and in what form — and why it changes everything
Section 1: The Problem — Why Agents Break in the Real World
Consider this scenario: a team builds an AI-powered customer service agent using Amazon Bedrock. In development, it performs beautifully. It retrieves product information, checks order status through API calls, follows company policies, and responds with the right tone. The demo impresses stakeholders. The agent moves to production.
Within weeks, the cracks appear.
A customer reaches out mid-conversation about a return, references an issue they raised two days ago, and the agent has no memory of it. A different customer asks a multi-part question — by the time the agent processes the third tool call, the earlier results have been pushed out of view, and it contradicts its own previous answer. During peak hours, costs spike dramatically — every agent invocation stuffs the full conversation history, complete tool schemas, and retrieved documents into the context window, burning through tokens whether or not the information is relevant.
The model is not the problem. The information the model receives — and does not receive — is.
This is the hidden challenge of building agentic systems. Unlike a simple chatbot handling single-turn questions, agents operate across multiple turns, invoke tools, maintain state, coordinate with other agents, and make decisions based on accumulated context. Every additional capability adds more information competing for limited space in the context window. And that window, regardless of whether it is 128K or 1M tokens, is always finite — and always expensive to fill.
The consequences show up across three dimensions simultaneously:
Quality degrades. When critical information gets lost, truncated, or buried among irrelevant details, agents make poor decisions. They hallucinate. They forget. They lose coherence across a conversation. The user experience suffers in ways that are difficult to debug because the root cause is not in the logic — it is in what the model could not see.
Performance slows. Larger context windows mean longer processing times. Stuffing everything into the prompt — "just in case" — creates latency that compounds in multi-step agentic workflows where one agent's output feeds into another. What felt instant in a single-turn demo becomes sluggish in a ten-step orchestration.
Costs escalate. Token consumption is directly tied to context size. In agentic architectures, where a single user request might trigger dozens of LLM calls — each carrying its own context payload — the math becomes unforgiving. Multiply that across concurrent users, and teams find themselves optimizing infrastructure costs before they have optimized the product itself.
These three forces — quality, performance, and cost — are not independent trade-offs. They are symptoms of a single underlying problem: the absence of a deliberate strategy for managing what information flows into and out of an agent's context, when, and in what form.
This is the problem that context engineering exists to solve.

Everything competes for the same finite window.
Section 2: Anatomy of Agent Context — What's Actually in the Window
Before engineering a solution, it helps to understand what exactly occupies an agent's context window. In a simple chatbot, context is straightforward — a system prompt and a user message. In an agentic system, the picture is far more complex. Six distinct types of information compete for space, each with different characteristics, lifespans, and levels of importance.
System Instructions sit at the foundation. These define who the agent is and how it should behave — its persona, guardrails, response format, escalation rules, and safety boundaries. In most architectures, system instructions are static and present in every single invocation. They are the tax that every context window pays before any real work begins. For a well-configured production agent, system instructions alone can consume thousands of tokens.
User Input includes the current query and, critically, the conversation history. A user asking "What about the blue one?" on turn twelve of a conversation is meaningless without the prior eleven turns. Conversation history grows linearly with every exchange, and in agentic workflows — where the agent itself generates intermediate reasoning steps — it grows even faster than the user expects.
Retrieved Knowledge is the information pulled from external sources at query time — the "R" in Retrieval-Augmented Generation. This might come from Amazon Bedrock Knowledge Bases, vector stores, enterprise search systems, or structured databases. RAG is powerful, but every retrieved chunk occupies context space. Retrieve too little, and the agent lacks the information it needs. Retrieve too much, and relevant details get buried in noise — a problem researchers call "lost in the middle," where models demonstrably pay less attention to information positioned in the center of long contexts.
Tool Definitions and Results represent the agent's ability to act. Tool definitions — the schemas that describe what each tool does, its parameters, and expected outputs — must be present for the model to know what actions are available. In a system with twenty or thirty available tools, definitions alone can consume a significant portion of the window. Then come the results: raw API responses, database query outputs, file contents. A single tool call to a search API might return kilobytes of JSON. Three tool calls in sequence, and the context window is carrying the weight of an entire conversation in tool output alone.
Agent Memory is what the agent knows beyond the current conversation. Short-term memory holds session-level state — the user selected a blue shirt, they prefer express shipping, they already rejected the first recommendation. Long-term memory spans sessions — this customer is a premium member, they had an unresolved complaint last month, they always ask about sustainability. Without memory, every conversation starts from zero. With unmanaged memory, the agent drowns in stale or irrelevant recollections.
Orchestration State emerges in multi-agent and multi-step architectures. When a planning agent delegates tasks to specialist agents, or when a workflow involves sequential reasoning steps, the system must track where it is in the process, what has been completed, what failed, and what remains. In frameworks built on Amazon Bedrock, where agents can invoke sub-agents or chain multiple actions, orchestration state becomes yet another consumer of context real estate.
Here is what makes this challenging: all six components are present simultaneously in a single context window, and their relative importance shifts with every turn.
On the first turn of a conversation, conversation history is negligible but retrieved knowledge is critical. By turn fifteen, conversation history dominates but most of it may be irrelevant to the current question. When the agent is about to invoke a tool, tool definitions matter enormously — but the moment the tool returns its result, the definitions for the other twenty-nine tools are dead weight. Long-term memory about the user's preferences is essential when making a recommendation, but adds nothing when the user asks for a shipping policy.
Static approaches — fixed prompt templates, hard-coded retrieval limits, blanket inclusion of all tool schemas — cannot adapt to this reality. What agentic systems need is an engineered approach to context: one that understands these components, their relationships, and their dynamic importance.
That approach has a name.

Six types of information, one finite window — their importance shifts with every turn.
Section 3: What Is Context Engineering?
The term "prompt engineering" has dominated the AI conversation for the past two years. It described a real and valuable skill — crafting the right instructions to get a useful response from a language model. But prompt engineering, as commonly practiced, is largely a static, single-turn discipline. It focuses on writing better text that goes into the prompt. For agentic systems, this is no longer sufficient.
Context engineering is the next evolution. Where prompt engineering asks "what should I write in the prompt?", context engineering asks "what system should I build to ensure the right information reaches the model at the right time, in the right form, at the right cost?"
The distinction matters. Prompt engineering is authoring. Context engineering is architecture.
More precisely, context engineering is the discipline of designing and implementing the systems, strategies, and logic that govern how information flows into and out of an agent's context window across turns, sessions, and agent boundaries. It treats context not as a static text block but as a dynamic, managed resource — one that must be curated, compressed, persisted, and routed with the same rigor that teams apply to data pipelines or network traffic.
Anthropic's research on building effective agents has consistently emphasized that the orchestration layer — what information reaches the model and when — is often more consequential than the model's raw capability. When two teams use the same foundation model on Amazon Bedrock but one builds a sophisticated context management system and the other stuffs everything into the prompt, the difference in agent quality is dramatic — and it has nothing to do with the model.
To make this discipline actionable, it helps to think of context management as a lifecycle with four stages:
Select — deciding what information deserves to be in the context window for this specific invocation. Not everything that could be included should be included. Selection involves choosing which conversation turns are relevant, which retrieved documents are worth the token cost, which tool definitions are needed right now, and which memories apply to the current task. Selection is the most impactful stage — information excluded here never gets a chance to influence the agent's behavior, for better or worse.
Compress — reducing the token footprint of selected information without losing its essential meaning. A fifteen-turn conversation history might be summarized into a concise paragraph. A raw API response with fifty fields might be distilled to the three fields that matter. Compression allows agents to carry more knowledge in fewer tokens — expanding effective context without expanding actual cost.
Persist — storing information beyond the boundaries of a single context window or a single session. Not everything can or should live in the active context at all times. Persistence strategies determine what gets written to memory, at what granularity, and with what metadata — so it can be retrieved later when it becomes relevant again. This is what gives agents the ability to remember, learn, and maintain continuity.
Route — directing the right context to the right agent or process in multi-agent architectures. When a planning agent hands off a task to a specialist agent, what context travels with that handoff? When two agents collaborate, do they share full context or scoped summaries? Routing decisions affect not only quality — giving each agent precisely the context it needs — but also security and cost, since over-sharing context across agent boundaries multiplies token consumption unnecessarily.
These four stages are not sequential in a strict sense. In a running agentic system, they operate continuously and often simultaneously. Every new turn triggers selection logic. Compression may happen in the background between turns. Persistence writes and reads happen throughout. Routing decisions execute every time work crosses an agent boundary.
Together, they form a context management system — the invisible but critical infrastructure layer that separates agents that work in demos from agents that work in production.

The Context Engineering Lifecycle — four stages that run continuously across every agent interaction.
The next question is practical: what specific techniques power each stage? that's what we are going to cover in part 2 of this article.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article