"Plainly" and "So an Elementary Schooler Could Understand It" Were Not the Same — An Experiment with an AI Agent's Question Rules
When an AI agent's questions are hard to understand, "in plain words" and "so an elementary school student could understand it" produced different results when added to the question rules. "Plain words" barely differed from the original rules; "elementary schooler" was better at explaining "what happens if I pick this." In a first experiment writing questions from prepared facts, that one line alone raised the score sharply. But when AI-DLC actually ran, adding it did not clear the pre-set "improved" bar.
The AI Asks, but I Don't Know What It's Asking
When you work with an AI coding agent, it sometimes asks you to decide something: "Should we build this feature first?", "How long should we keep the data?" But sometimes you read the question and all its choices and still can't tell what you are being asked to decide. So you end up asking the agent: "Can you explain that more simply?"
I looked at this problem in AI-DLC. AI-DLC (AI-Driven Development Life Cycle) is a development approach proposed by AWS: the AI makes a plan, asks a person about what it doesn't know, and moves to the next stage only after the person confirms (introduction on the AWS DevOps Blog ). The open-source project that makes this approach usable across AI coding tools is awslabs/aidlc-workflows . Because a person confirms every stage, the agent asks a lot of questions.
Around late September, Korean-speaking users gave feedback like this: "The questions and choices come out in Korean, but I still can't tell what they mean, so I have to ask for a simpler explanation." AI-DLC was in the 2.10 release line at the time.
A hard question looks like this. The example below is made up for illustration — it did not come from the experiments.
1
2
3
4
5
Q1. The retention periods in NFR-3 and §4.2 conflict with ADR-7. How should we handle this?
A. Keep ADR-7
B. Update to match §4.2
C. Redefine NFR-3
X. Other
The sentences are short. But you have to know what NFR-3, §4.2, and ADR-7 are before you can answer. The question also doesn't say what is being decided, or what happens if you pick each option.
In fact, AI-DLC already had a rule meant to reduce this problem at the time — it went in at the end of August and was still there in the 2.10 release line. The rules the agent follows when writing questions include the sentence "Questions must be self-explanatory" and three guidelines:
- When a question uses an ID like
FR3orNFR-2, describe what it refers to first, and put the ID in parentheses once. - When it is not clear why a question is being asked, or what depends on the answer, add one line about it.
- Ask concretely in the words the user works with, not in AI-DLC's internal vocabulary.
The rule was there, and users still asked for simpler explanations. Looking at the three guidelines again: they say to spell out IDs and to attach one line of context to the question as a whole, but there is no guideline that says to attach "what happens if you pick this" to each individual choice.
When we ask a person to explain something simply, we usually say either "put it plainly" or "explain it so even an elementary schooler would get it." I experimented with what happens to the agent's questions when each of these two phrases is written into the question rules.
How "Plainly" and "Elementary Schooler" Differ
I compared two instructions. AI-DLC's rules are written in English, so the experiments inserted English sentences too.
| Name | First part (different) | Second part (identical) |
|---|---|---|
| "Plain words" | Write every question and choice in plain everyday words, as if the person had asked you to explain it so someone new could understand | say first what needs deciding and why, and what each choice leads to. Keep an exact name only where the answer depends on it. |
| "Elementary schooler" | Write every question and choice so simply that an elementary school student could understand it | (same as above) |
The second part is identical in both instructions. It means: "Say first what needs to be decided and why, and what each choice leads to. Keep an exact name only when the answer depends on it."
The only difference is the first part: who you are told to write simply for. One says "in plain everyday words, as if explaining to someone new"; the other says "so simply that an elementary school student could understand it."
Here, "elementary schooler" is just a phrase in the instruction. This is not an experiment where elementary school students read the questions.
Experiment 1: Comparing the Two Instructions Side by Side (September 29, 2026)
What we did
I picked four questions from two real AI-DLC projects (called Case A and Case B in this article). For each question, I prepared a sheet of the facts needed to write it. Then I gave Claude the question rules and that sheet, and had it write the question anew in Korean. Each condition wrote the same question three times. The question rules were the ones from around the time of the feedback, the 2.10 release line.
Claude also did the scoring. It was not told which condition each question came from. I made two shuffled batches and scored each once. There were three measures:
- Decision score: Can you tell what needs to be decided? (1–5)
- Outcome score: Can you tell what happens if you pick each choice? (1–5)
- Unfamiliar terms: How many words might the reader not know? (counted by the scoring AI)
Each condition was scored 24 times: 4 questions × 3 generations × 2 scoring passes. That is not 24 people reading, and not 24 distinct questions.
Results: "Plain words" stayed about the same; "elementary schooler" changed things
Three conditions were compared in the same scoring batch: the original rules alone, the original rules plus the "plain words" line, and the original rules plus the "elementary schooler" line. Scores in the table are "decision / outcome."
| Question | Original rules | + "Plain words" | + "Elementary schooler" |
|---|---|---|---|
| Case A, deciding requirements | 4.17 / 4.17 | 4.33 / 4.00 | 4.17 / 4.17 |
| Case A, team working style | 4.67 / 3.67 | 4.67 / 3.83 | 5.00 / 4.33 |
| Case B, approval and handover | 3.67 / 4.00 | 4.00 / 4.00 | 4.33 / 4.67 |
| Case B, team working style | 4.17 / 4.17 | 4.00 / 4.00 | 4.83 / 4.83 |
| Average (24 ratings per condition) | 4.17 / 4.00 | 4.25 / 3.96 | 4.58 / 4.50 |
| Unfamiliar terms (average count) | 1.25 | 1.46 | 1.42 |
"Plain words" was barely distinguishable from the original rules — its outcome score was actually slightly lower. "Elementary schooler" scored highest on three of the four questions and tied the original rules on the fourth. Unfamiliar-term counts were nearly identical across all three conditions.
The same day, I ran one more test. AI-DLC writes its questions to a file first; if the person picks "guide me one by one," the agent shows those questions in the chat one at a time. In this test, I compared how each condition displayed three questions from a real project's question file in the chat. No scores — I just read the outputs side by side.
- With the original rules only: the agent copied the questions from the file into the chat almost verbatim.
- With either the "plain words" line or the "elementary schooler" line: both rewrote the questions and said what needed deciding first. The text was also about a quarter shorter.
- Of the two, the one that most often attached "here's what happens if you pick this" to each choice was "elementary schooler."
What if there are no other rules at all?
To see whether this one line works without any help from the original rules, I compared again while adding and removing rules. There are two sets of rules here:
- Voice rules: rules that apply to everything the agent says to a person.
- Original question rules: the "Questions must be self-explanatory" sentence and the three guidelines introduced above.
In the table, "question rules swapped for the line" means the original question rules were removed and the "elementary schooler" line was put in their place. Every condition got the same facts sheet and the same task: "write the question."
| Condition | Voice rules | Original question rules | Inserted line | Decision | Outcome | Unfamiliar terms |
|---|---|---|---|---|---|---|
| No rules | – | – | – | 4.25 | 4.04 | 1.50 |
| Original rules | ✓ | ✓ | – | 4.38 | 4.00 | 1.00 |
| Original + "plain words" | ✓ | ✓ | "plain words" | 4.38 | 4.25 | 1.17 |
| Original + "elementary schooler" | ✓ | ✓ | "elementary schooler" | 4.58 | 4.46 | 1.17 |
| Question rules swapped for the line | ✓ | – | "elementary schooler" | 4.62 | 4.71 | 1.00 |
| "Elementary schooler" line only | – | – | "elementary schooler" | 4.62 | 4.71 | 1.12 |
| Swapped for the line + keep only the ID guideline | ✓ | ID guideline only | "elementary schooler" | 4.88 | 4.58 | 0.96 |
(24 ratings per condition. Even for the same condition, averages moved by about 0.1–0.2 points when the scoring batch changed, so I don't compare numbers directly between this table and the previous one.)
The rows to watch are No rules and "Elementary schooler" line only. With no voice rules and no question rules, inserting just this one line raised the outcome score from 4.04 to 4.71 — higher than the original rules (4.00).
Before this table, I had scored a separate six-condition batch, and it showed the same thing: no rules was 4.12 / 4.04, and the line alone was 4.67 / 4.71. In both batches, the outcome score rose by about 0.67 points — far larger than the roughly 0.1–0.2-point variation observed across scoring batches.
In the same table, "plain words" also raised the outcome score a bit over the original rules (4.00 → 4.25), but it did not reach "elementary schooler" (4.46). A condition with only the "plain words" line and no other rules was not tested.
What happens with real project material?
The comparison above had a weakness: the model got a clean, prepared facts sheet. Those sheets carried almost no unexplained IDs, so this comparison couldn't really show whether the instruction reduces the unexplained IDs that were a problem in real work.
Real work is different. The agent writes questions after reading piles of material — documents from earlier stages, lists of decisions made so far, code-analysis results. That material is full of IDs like
FR3 and §4.2.So this time I put in real project material, unprocessed.
- Input: the real project material of Case A and Case B, about 300 KB each (several documents' worth)
- Task: not one question, but a stage's entire question file
- Compared: the original rules versus the last condition in the table above (question rules swapped for the "elementary schooler" line, keeping only the ID guideline)
- AI used: Claude Opus 4.6 for both writing and scoring
Scores: On Case A, the changed rules scored 4.33 / 4.25 for decision/outcome — not lower than the original rules (4.17 / 4.17). Case B could not be scored: the scoring AI hit its output limit before producing an answer.
Unexplained IDs: IDs and names used without explanation were counted by a program. The averages per question:
| Case | Original rules | Changed rules |
|---|---|---|
| Case A | 4.22 | 4.07 (almost unchanged) |
| Case B | 2.86 | 1.41 (roughly halved) |
IDs that flow in from the material the agent read sometimes went down with this one line — and sometimes stayed put.
Why Was "Elementary Schooler" Different?
The two instructions share the same second part, which means both received the identical demand to "say what each choice leads to." If the results still differed, the difference plausibly came from the first part — who the agent is told to write for. What follows is my interpretation of the results, not something verified separately.
- "Someone new" is a broad standard. The "plain words" instruction also names a reader — "so someone new could understand" — but it isn't clear what that person doesn't know. So if the agent judges "my text is already this easy," it has no reason to change anything. That fits "plain words" barely differing from the original rules.
- "An elementary school student" is a specific reference point with little assumed background knowledge. It evokes a reader who knows neither developer vocabulary nor the project's circumstances. Then the agent actually has to spell out what is being decided, why, and what each choice leads to. The identical second-part demand raised the outcome score more when paired with "elementary schooler," and that condition also attached per-choice consequences most often.
In short, the difference I saw was how specifically the target reader was defined, and how little background knowledge was assumed.
Reading the original-rules questions and the "elementary schooler" questions from experiment 1 side by side myself, the "elementary schooler" ones were easier to read. But that is one person's impression, not a scored evaluation against criteria, so I don't use it as evidence for the numbers above.
Experiment 2: Actually Running AI-DLC (October 9, 2026)
Experiment 1 had the model write questions from prepared facts, and a later comparison added real project material. But it never actually ran AI-DLC. In real use, the agent writes a question file at each stage, working from AI-DLC's many instructions and the documents produced by earlier stages. To see whether the "elementary schooler" line makes a difference in that situation too, this time I actually ran AI-DLC and collected the questions.
The experiment used AI-DLC 2.11.0 , released on October 8. The "Questions must be self-explanatory" sentence and the three guidelines were identical, letter for letter, to experiment 1's version (the 2.10 line). What changed in between was elsewhere: the questions and choices shown when work gets stuck, the guidance text people read, and some question-related instructions were rewritten in plainer language (#1976 , #2024 , #2010 , #2069 ).
Using 2.11.0, I felt the questions were not that hard anymore even without the "elementary schooler" line. So I wondered: starting from this state, would adding the line make the questions better still?
Experiment 2 did not include "plain words." In experiment 1, "elementary schooler" had already beaten "plain words," so this time I focused on whether "elementary schooler" keeps producing good results when AI-DLC actually runs.
What we did
I ran AI-DLC 2.11.0 with Claude Opus 4.8 in Kiro CLI (2.28.0).
First, I ran two projects under the original rules, up to the end of Inception — the early phase where requirements and design get decided, before building starts. One was a restaurant table-ordering app made up for this purpose; the other was a fictional seat-reservation app written deliberately full of requirement IDs. In each project, I saved the state just before each of the six question-writing stages. I'll call these 12 saved states the "starting states."
All four conditions kept the "Questions must be self-explanatory" sentence and the first guideline (spell out IDs). What varied was the second guideline (one line on why it's asked), the third guideline (ask concretely in the user's words), and the "elementary schooler" line.
The "elementary schooler" line says things similar to those two guidelines: "say first what needs deciding and why" resembles the second, and "so simply that an elementary schooler could understand" resembles the third. It also contains something the original rules lack: "what each choice leads to." To separate the effect of adding the line from the effect of removing the two guidelines, I compared four conditions:
- Original rules: the three guidelines as they are
- Add: the three guidelines + the "elementary schooler" line
- Remove only: drop the second and third guidelines, insert nothing
- Swap: the "elementary schooler" line in place of the second and third guidelines
Swap is shaped like experiment 1's last condition (question rules swapped for the line, keeping only the ID guideline) — except this time the "Questions must be self-explanatory" sentence was kept too.
From each starting state, each condition started fresh, and the stage's question file was saved before anyone answered it. I generated two files per condition from each starting state, for 96 files in all.
Scoring was done by one OpenAI model through the Codex CLI. For each starting state, it scored the 8 files twice, blind to condition, in shuffled order. Markers like "(recommended)" on choices were removed, and only the questions were passed in. I also used the original run's question file as a reference and checked how completely each new question file still asked the decisions it had raised. I'll call this ratio "decision coverage."
We decided in advance what would count as "improved"
Before looking at results, I fixed the criteria.
The threshold is 0.2 points. In experiment 1, rescoring the same questions in a different batch moved averages by about 0.1–0.2 points, so differences under 0.2 points are hard to tell apart from the scoring AI's wobble.
"Improved" was judged on the outcome score alone, because that is what "elementary schooler" mainly moved in experiment 1 — and because allowing any one of several scores to clear the bar invites picking whichever one happens to clear it by chance.
I also computed a "95% range" to see how reliable each average difference is: resample the 12 starting states with replacement, 12 at a time, recompute the average difference, and repeat 10,000 times. (The experiment itself was not rerun 10,000 times.) The wider the range, the less certain the average difference; if the range contains 0, there may be no difference at all.
- Improved: the outcome score is at least 0.2 points above the original rules, and the entire 95% range is above 0.
- Did not get worse: this does not mean no score dropped at all. It means the lowest end of the 95% range of the difference versus the original rules satisfies all three: decision and outcome scores each above −0.2 points, and decision coverage above −10 pp. ("pp" stands for percentage points: a drop from 84% to 81% is −3 pp.) A stage has about 6 decisions on average, so 10 pp is a bit over half a decision.
With only 12 starting states, I drew no conclusions about very small differences.
Results: even adding the line did not clear the "improved" bar
| Condition average (12 starting states) | Decision | Outcome | Unfamiliar terms (avg per question) | Decision coverage |
|---|---|---|---|---|
| Original rules | 4.28 | 4.03 | 1.63 | 84% |
| Add | 4.44 | 4.09 | 1.34 | 85% |
| Remove only | 4.25 | 3.97 | 1.57 | 83% |
| Swap | 4.45 | 4.08 | 1.17 | 81% |
| Difference vs. original rules [95% range] | Decision (pts) | Outcome (pts) | Decision coverage (pp) |
|---|---|---|---|
| Swap | +0.17 [+0.01, +0.36] | +0.05 [−0.17, +0.28] | −3 pp [−7, +1] |
| Add | +0.16 [−0.01, +0.34] | +0.05 [−0.13, +0.22] | +1 pp [−6, +7] |
| Remove only | −0.02 [−0.15, +0.11] | −0.07 [−0.28, +0.12] | −1 pp [−4, +3] |
The outcome score that "elementary schooler" had raised sharply in experiment 1 moved by only +0.05 points on average here, whether the line was added or swapped in. That did not clear the pre-set "improved" bar.
Add's decision score rose by +0.16 points on average, but its 95% range included 0 (−0.01 to +0.34), so I could not confirm the decision score went up.
Swap cleared the "did not get worse" bar only when the two apps' results were pooled; split by app, even that was not separately confirmed.
Reading experiment 2's outputs myself, I didn't feel a big change either.
The two experiments differed in many ways. Experiment 1's main comparison wrote one question from prepared facts; experiment 2 wrote a stage's whole question file within an actual AI-DLC workflow. The scoring AIs differed (Claude vs. an OpenAI model), as did the AI-DLC versions (the 2.10 line vs. 2.11.0). And experiment 2 had no "no rules" condition. So I can't tell what caused the results to differ.
So: A Good Starting Point When You Have No Detailed Guidelines
Putting the two experiments together:
- With no voice rules and no question rules, the "elementary schooler" line alone pushed the outcome score above the original rules (experiment 1).
- Even with identical second parts, "elementary schooler" explained "what happens if you pick this" better than "plain words" (experiment 1).
- When AI-DLC actually ran, adding the line did not clear the pre-set "improved" bar on the outcome score (experiment 2).
So if you're using an agent without detailed response guidelines, starting with "so an elementary schooler could understand it" rather than "put it plainly" can be a good starting point. This is a judgment combining two results from a small experiment that wrote questions from prepared facts (experiment 1). A "plain words"-only condition without other rules was not tested. And in experiment 2, which actually ran AI-DLC, adding the line did not re-confirm an improvement. If you already have guidelines in place, add the line and check for yourself whether the actual questions change.
If you want to try it
The English sentence used in the experiments:
Write every question and choice so simply that an elementary school student could understand it: say first what needs deciding and why, and what each choice leads to. Keep an exact name only where the answer depends on it.
In all experiments, the instruction was inserted in English and the questions were generated in Korean, the projects' working language. I did not test whether the same effects occur with questions generated in other languages.
When you try it, watch two things:
- Did "what happens if you pick this" get attached? This was the biggest change in experiment 1. Check whether each choice carries this explanation.
- Are unfamiliar IDs still there? IDs from the documents the agent read did not disappear with this one line. If IDs remain, the thing to examine may be the material you give the agent, not the sentence.
Below is an example made up solely for illustration to show the shape of question this instruction aims at. It is not from the experiments, and it is not a fixed-up version of the hard question above.
1
2
3
4
5
Q1. When a member deletes their account, we need to decide when to erase their order history.
The documents disagree on the retention period, and we can't build the deletion feature until this is decided.
A. Erase immediately on deletion — personal data disappears quickly, but if a refund inquiry comes later, there's no record to check.
B. Keep for one year, then erase — refunds and disputes can be handled, but the personal data must be kept secure in the meantime.
X. Other (please write in)
What These Experiments Cannot Tell You
No human comprehension testing. The questions were written by AI and scored by AI. Apart from my own impressions reading the results, there was no human evaluation. Experiment 1 was scored by Claude; experiment 2 by one OpenAI model through the Codex CLI. Conditions were hidden and order shuffled during scoring — that is all. The questions were not tested for comprehension with elementary school students or actual users.
One task, tested small. Both experiments were about writing questions in AI-DLC. Whether the same difference appears in other writing — explanations, summaries, error messages — is unknown. Experiment 1's facts-sheet comparison used four questions; experiment 2 used 12 starting states from two apps. Everything was in Korean. Nor was this a comparison across AI models: the exact Claude model that wrote from facts sheets in experiment 1 was not recorded and cannot be verified; the real-material comparison used Opus 4.6, and experiment 2 used Opus 4.8.
The "no guidelines" evidence comes only from experiment 1. The comparison between no rules (no voice rules and no question rules) and the "elementary schooler" line alone came from experiment 1, which wrote questions from prepared facts. There is no result from running the real tool without guidelines. The "plain words"-only condition was not tested either.
Why "elementary schooler" differs is my interpretation. The broad-standard-versus-concrete-standard explanation fits the results, but no experiment tested the explanation itself.
Takeaways
- "Plainly" and "elementary schooler" were not the same. Of two instructions with identical second parts, "plain words" barely differed from the original rules, while "elementary schooler" was better, particularly at explaining "what happens if you pick this" — likely because it defines the target reader specifically and assumes little background knowledge.
- With no other rules, this one line made a large difference. In experiment 1, which wrote questions from prepared facts, inserting just this line with no voice rules and no question rules pushed the outcome score above the original rules.
- Experiment 2, which actually ran AI-DLC, did not re-confirm an improvement. Adding the line did not clear the pre-set "improved" bar on the outcome score, and Swap cleared the "did not get worse" bar only with the two apps pooled. Why the results differed from experiment 1 is unknown.
If you're getting questions from an agent that has no detailed response guidelines, try inserting the one-line "elementary schooler" instruction shown above — and check whether each choice gains a "here's what happens if you pick this," and whether unfamiliar IDs remain.
References
- AI-Driven Development Life Cycle: Reimagining Software Engineering — AWS DevOps & Developer Productivity Blog
- awslabs/aidlc-workflows — the AI-DLC open-source project
- AI-DLC 2.11.0 release — October 8, 2026
- #1976 · #2024 · #2010 · #2069 — 2.11.0 changes that made human-facing text plainer
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article