AWS Builder Center
I tried to make a coding agent play a text adventure. The bakery beat me, but I learnt so much

I tried to make a coding agent play a text adventure. The bakery beat me, but I learnt so much

I spent the Claude Code Camp trying to get an AI agent to do something that sounds trivial and turns out to be brutal: connect to a 1990s multiplayer text game (a MUD running on `localhost:4000`), log in, walk to the bakery, and read the menu.

All views are my own.

I tried to make a coding agent play a text adventure. The bakery beat me, and that's the point.

I spent the Claude Code Camp  from Andrew Brown, trying to get an AI agent to do something that sounds trivial and turns out to be brutal: connect to a 1990s multiplayer text game (a MUD running on localhost:4000), log in, walk to the bakery, and read the menu. No graphics, no API, just telnet and room descriptions like "You are in the southern end of the temple hall." A human does this in a minute. My agent burned about 65,000 tokens wandering and still didn't get there.
I want to write why, because the failure taught me more about building AI systems than any success would have, and almost all of it transfers to the cloud and agent work I actually get paid to do.

The first mistake was thinking it was a model problem

In the pre-week I did what most people do: I let the agent write throwaway Python and PowerShell scripts to open a socket, guess the login sequence, send commands, and parse output. It confused the CircleMUD login dance (name, then an immediate password prompt, then RETURN, then a menu), sent the wrong thing, dropped the connection after every command, and started over. I switched to smarter models. I tried DeepSeek and Qwen to save credits. The reasoning got a little better, but the core problem didn't get fixed magically.
The problem wasn't intelligence, it was architecture. I was asking the language model to regenerate deterministic plumbing, connection, authentication, framing on every single attempt, and that plumbing has fiddly state the model has no business rediscovering each time.
So I did the boring thing to get it working: put the deterministic parts behind a boundary. I built a small daemon that owns the socket, the telnet negotiation, the login dance, and speaks the Model Context Protocol (MCP) over stdio. The agent no longer knows what a MUD is. It calls look, move, rest like any other tool, and a separate process handles the telnet reality. The moment I did that, "can it log in?" stopped being a question.

Then it was a memory problem

With a reliable connection, the agent could move around and immediately revealed the real bottleneck. Every turn it re-read the full room description, re-reasoned about which way to go, forgot where it had already been, and walked north then straight back south. That's where the 65K tokens went: not thinking hard, but thinking the same thing over and over.
The fix was to give the agent memory instead of a bigger context window. I added a SQLite knowledge store: rooms identified by a stable fingerprint, exits and where they lead, a "frontier" of exits seen but not yet walked, visit counts, and player vitals. Before each model call I inject a compact [here] line for the current room, destination-aware exits, how much is unexplored, instead of dumping raw text. Route planning became plain breadth-first search over walked exits, no model involved.
One detail worth sharing: an open field where every room shares the same name, description, and exits. My fingerprinting collapsed them all into one node, so the map broke and exploration stalled. The fix is what real MUD mappers do, dead-reckoning: track an (x, y, z) coordinate from your movements and use it as part of a room's identity. Two rooms that look identical but sit at different coordinates are correctly treated as different places.
And the last honest twist: after all that, the agent still didn't reach the bakery. My character was parked in a high-level maze with zero movement points. But this time the memory told me why the stored vitals showed 0MV, and the log showed each exit correctly marked blocked. "The agent got confused" became "the agent is out of movement points in room #1 with three exits it can't use." That's the whole payoff of memory and observability: failures stop being mysterious.

What I'd tell future me or another builder to do

  • Put deterministic I/O behind an interface and let the model reason. Connection, auth, retries, and protocol framing should live in code (an MCP server, an SDK, a daemon), not be regenerated by the LLM each run. This single change fixed more than any model upgrade.
  • Give agents structured memory, not a bigger prompt. A tiny SQLite table plus a compact injected state block beat re-feeding raw output. Store what's known; inject a summary; let the model orchestrate.
  • Measure before you optimize. I only found the token sink because I logged moves, rooms, and token usage to JSON and rendered it. "It feels slow" is not a diagnosis.
  • You can test an agent loop with no API spend. I validated the entire tool-use loop with a scripted, deterministic "brain" standing in for the model — zero credits, real tools, real MUD. Save the paid model runs for when the mechanics already work.
  • Default-deny your tools. Scope which tools each task may call, down to parameter values. It keeps behavior predictable and your bill bounded.
  • Prefer boring, portable foundations. I built the whole thing on the standard library, hand-rolled MCP-over-stdio, SQLite, a self-contained HTML report instead of a hosted dashboard. Fewer dependencies meant fewer things breaking on Windows and a system I could actually reason about.
If you map this to AWS, the shape is familiar: the reasoning layer is a swappable model endpoint (like Bedrock), the deterministic capabilities are MCP servers or Lambda-backed tools behind a clear contract, memory is a managed store, and observability is non-negotiable. The pattern that made a text game tractable is the same one that makes production agents tractable, keep the model doing judgment, and everything deterministic in code around it, it'll also save you money and tokens.

Should you do the bootcamp?

Do it, especially if you're a cloud or backend engineer who suspects "AI agents" is mostly hype, they also have some great grants if cost would be a barrier, it was $100 full price and $50 as a community builder. The Claude Code Camp  is valuable because it hands you a goal that looks easy and then lets reality push back, you learn agent architecture by hitting the walls yourself, not by watching a demo where everything works, plus the Discord community and weekly session are really great. I finished understanding tool boundaries, context economics, structured memory, and how to instrument an agent so its failures are legible.
Budget more time than you think to go through the content, build and test, lean on the cheaper models (DeepSeek and Qwen carried a lot of my experimentation for a fraction of the cost), keep your own tests and folders tidy because the agent won't, and treat every failed run as the actual curriculum. My agent did eventually read that bakery menu but didn't defeat the Minotaur (though I did manually), so now I know exactly what it'll take, and that clarity is worth more than a lucky success.
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article