AWS Builder Center

ByteGeist Troubleshooting Buddy: one next step when you're stuck

A guided troubleshooting partner for beginning IT learners, built with Amazon Bedrock, Lambda, DynamoDB, and Amplify. One supported check at a time, with real deployment and testing evidence.

AWS Learner | Software Engineering Student | Cloud & Infrastructure
The browser walkthrough reaches the existing 100/100 incident review.
The browser walkthrough reaches the existing 100/100 incident review.
Getting a wall of ten commands when you're already stuck feels like being handed more homework. I wanted a troubleshooting agent that helps a beginner slow down, collect one useful result, and understand what it means before moving on.
I'm a software engineering student transitioning from physical work into tech. That shaped the audience for ByteGeist Troubleshooting Buddy: people studying IT fundamentals, preparing for their first support role, or practicing the reasoning behind Windows, Linux, and AWS troubleshooting. You can memorize a command and still struggle to decide when to use it. Buddy focuses on that decision and the evidence that comes afterward.

What Buddy does

Buddy is integrated into my existing ByteGeist Support Lab. The lab already had a simulated terminal, five incidents, adaptive coaching, deterministic scoring, and a learner performance report. I extended that foundation with a guided investigation flow.
Open an incident and choose Start guided troubleshooting. Buddy offers one supported check, explains the purpose, and asks you to notice something in the result. Load command into terminal places the exact command in the input field. You press Enter to run it. Then Review result + next step helps interpret the evidence and introduces the next check.
The five incidents cover a Windows DNS problem, file-share authorization, an S3 permission error, a Linux service that will not start, and an application outage on an otherwise running EC2 instance. These are training scenarios. The terminal returns simulated results; it does not execute commands on your computer or modify an AWS workload.
The coaching, however, uses real Amazon Bedrock responses. When you ask what an unfamiliar term means, the model can explain it in context. During the browser walkthrough, asking about DNS produced a plain-English explanation comparing it to a phone book that translates names into IP addresses.

How I built it on AWS

AWS Amplify Hosting serves the React and Vite frontend. The browser sends requests through an Amazon API Gateway HTTP API to Node.js Lambda functions in us-east-1. DynamoDB stores the scenario definitions and tutoring sessions. Lambda calls Amazon Nova Micro through the Amazon Bedrock Converse API, using structured tool output for the teaching response. AWS SAM and CloudFormation manage the backend deployment.
The guided loop lives in the application: read the server session, choose the next unfinished supported check, request an explanation from Bedrock, validate the response, save the tutoring state, and wait for the learner.
The browser does not decide which checks count as completed. The backend derives progress from supported command results recorded in that session. Help output and rejected commands do not advance the investigation. Repeating the same check does not inflate progress. After the planned checks are collected, Buddy changes to reflection and asks the learner to connect the evidence to a diagnosis and proposed fix.
I kept the existing rule: AI teaches. Code grades. The numeric score comes from deterministic application logic. The model supplies coaching and observations about the learner's reasoning. Collecting four checks does not automatically mean the learner understands them or has repaired a real system.

The small detail I cared about

The delightful detail is the command handoff. A beginner gets one exact, supported command and a button that loads it into the terminal, while retaining control over execution. That removes a copy-and-paste hurdle without removing the moment when the learner chooses to act.
I verified that behavior in the deployed browser. Clicking the load button filled the terminal input and left the output history unchanged. The next-step button stayed disabled until the pending check was actually run. After review, progress advanced and the card showed the next command. At the end of the DNS incident, it showed four of four checks collected and changed to a reflection prompt.
I also deliberately simulated a coach outage in the browser. The error explained that the terminal evidence was still available, and the transcript remained intact. A subsequent real request recovered normally. That matters because losing your work after an error is a quick way to make a learning tool frustrating.
These are observed functional walkthrough results. I still want feedback from beginning learners about whether the pacing and explanations help them learn.

Testing and proof

The local suite has 13 passing tests, including all five guided command sequences, rejection of fabricated selected evidence, session checks, model-response repair, and compatibility with the existing hint ladder. The frontend production build and lint checks passed.
The final deployed regression passed 120 API requests, including 70 successful coach responses from real Bedrock. It covered all five guided incidents and the existing questions, four hint levels, output explanations, and reasoning checks. Each correct seeded investigation and resolution received 100/100. Fresh sessions restarted with zero collected checks. I also exercised the browser's diagnosis form, result report, retry flow, and a 390-pixel-wide mobile layout. The mobile page fit without horizontal scrolling.
Testing caught two useful mistakes. A Windows ZIP archive used backslashes in asset names, so Amplify reported deployment success while the browser could not load the assets. Correcting the archive paths restored the app. Nova also omitted a required teaching-message field in a structured response. I reduced the guided schema and added one explicit repair attempt through another real model call. Invalid responses are rejected rather than presented as successful coaching.
Try the deployed Buddy at ByteGeist Support Lab . Open The Missing File Server and choose Start guided troubleshooting.
The source and repeatable checks are in the ByteGeist Support Lab repository .
The deployed Buddy offers the next DNS check after two recorded results.
The deployed Buddy offers the next DNS check after two recorded results.
My next step is to watch beginning learners use it and improve the explanations where they hesitate. The goal is a partner that helps someone reason through the next useful check and gradually need less help.
#agents
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article