# Learning to Exercise Judgment

## Directing AI in the High School Classroom

Companion edition of *The Irreducible Officer*, for high school teachers, instructional leaders, and curriculum designers. September 2026. The original argues that the National War College has to teach and certify AI-enabled judgment. This edition argues that high schools should start building the same capacity, and uses a US History unit to show how. The classroom design is a proposal for teachers to test with their own students.

## I. Two Essays on the New Deal

A US History class gets a familiar prompt: was the New Deal a success?

One student works from the textbook and the teacher's document packet. She writes a careful essay arguing that the New Deal succeeded because unemployment fell and programs like Social Security lasted. It is organized, accurate, and conventional.

A second student uses AI differently. She has it speak as a Detroit autoworker in 1936, a Mississippi sharecropper, and a business owner who opposed the new regulations. She checks what each voice claims against the documents in the packet and throws out what she cannot confirm. Her essay argues that whether the New Deal succeeded depends on whose success you count, and she defends the standard she chose. It is the stronger essay.

Now ask which student understands the history. The essays will not tell you by themselves. The second student may have done harder thinking than anyone expects of a sixteen-year-old. Or the model may have supplied the "it depends on perspective" move, and she arranged its output well.

The first student is the harder case. She did everything the assignment asked, without help. She looks finished. She is also going to college and into jobs where her peers will pair similar reasoning with far more skill at directing AI. A school that rewards her essay and never teaches her the second student's method has taught to the old standard.

The goal for high school is students who can direct AI toward a question they own, decide when its output deserves weight, and explain their reasoning when someone pushes back. That takes subject knowledge. You cannot judge what a model says about sharecroppers if you do not know what sharecropping was. The knowledge gets built inside the work, while students practice directing the tool, rather than as a prerequisite that postpones the tool until senior year.

Adults still make the access decisions. Schools decide which tools are approved and whether students have individual accounts. The services set their own limits too; OpenAI, for example, requires users to be at least 13 and to have a parent's permission under 18 ([OpenAI Terms of Use](https://openai.com/policies/terms-of-use/)). None of that stops students from directing AI. A student can write the instructions, the criteria, and the check, and the teacher can run them on a shared screen.

## II. The Essay AI Writes

Most students start by typing the prompt as given: "Was the New Deal a success? Write an essay."

The model writes a good one. By standard estimates, it reports, unemployment fell from about 25 percent in 1933 to about 14 percent in 1937 before rising to 19 percent in the 1938 recession. It credits programs like the Civilian Conservation Corps, the Works Progress Administration, and Social Security with putting people to work and building protections that still exist. It notes that critics said the New Deal expanded federal power too far. It concludes that the New Deal was mixed but largely successful.

The facts are right, and the essay answers a narrower question than the one students should be asking. It judges success by economic recovery and by how long the programs lasted. It never asks success for whom.

That question changes the verdict. The old-age insurance program in the Social Security Act of 1935 did not cover agricultural or domestic workers, and those jobs employed about two-thirds of Black workers ([Dubin, 2024](https://law.stanford.edu/wp-content/uploads/2024/02/Dubin-Publication-Ready-1.pdf)). Congress added them in two steps, in 1950 and 1954.

Why they were left out is still argued, and the argument is good material for students. The committee that drafted the bill meant it to cover nearly every worker. The exclusion came from Treasury Secretary Henry Morgenthau, who said the Treasury could not collect payroll taxes from farms and households. Southern Democrats, who dominated the House Ways and Means Committee and had fought federal oversight elsewhere in the bill, backed the change, and the NAACP warned Congress that it would shut out most Black workers. Some historians read that record as racial politics working through an administrative argument (Lieberman, 1998; Katznelson, 2005). Others, including the Social Security Administration's own historian, find that the administrative concerns explain the exclusion (Davies & Derthick, 1997; [DeWitt, 2010](https://www.ssa.gov/policy/docs/ssb/v70n4/v70n4p49.html)), while granting that Southern support "no doubt reflected racial factors."

A student judging the New Deal by whether it protected the workers the Depression hurt most has to deal with that fact and that argument. A student judging it by national unemployment figures can leave both out.

Four of the original essay's failure modes are visible in this one exchange.

**Frame capture.** The model's standard of success becomes the student's. Her revisions polish the essay inside a boundary she never chose.

**Fluency substitution.** "Mixed but largely successful" sounds like a historian weighing evidence. It weighs everything and commits to nothing a reader could argue with.

**Invisible delegation.** "Write an essay" handed over the most important decision in the assignment, which is what success should mean. The student thought she was asking for help writing.

**Institutional monoculture.** Thirty students with similar tools will turn in many versions of "mixed but largely successful." The essays will differ in their examples and share one standard. A class discussion meant to surface competing arguments will mostly compare wording.

## III. Who Decides What Success Means

The teacher sets the lesson's purpose. She chose the unit, the prompt, and the documents, and she decided that students should practice evaluating a historical claim. That is her job, and a tightly bounded prompt is often good teaching.

The student sets the argument's standard. Whether success means recovery, durability, protection of the most vulnerable, or something else is the consequential choice the prompt leaves open. Two strong students can choose differently and both earn full credit if they defend the choice with evidence. That choice is the judgment the assignment exists to build, and it is exactly the choice the model makes silently when the student lets it.

AI can suggest standards too. A student who asks "what are different ways historians judge whether a policy succeeded?" will get a useful list. Accepting one from that list is fine. What matters is that she can say why it fits the question and what it leaves out.

## IV. Directing AI at Sixteen

The original essay describes a progression of how much judgment a person puts into an AI workflow. High school students can work at the middle of it.

At the bottom is the prompt students start with, "Write an essay on whether the New Deal was a success." It inherits the model's frame entirely.

One level up, the student states her standard and asks for help applying it: "I'm judging the New Deal by whether it protected the workers the Depression hurt most. Which programs in my packet support that case and which work against it?" Now the model is working on her question.

Above that is the evaluator loop, and it is where high school students can do something close to what the second officer does in the original essay. The student assigns the model several voices, such as a factory worker who gained union rights under the 1935 Wagner Act, a sharecropper, and a business owner, and asks each to argue whether the New Deal helped people like them. Then she weighs the disagreement and decides what she believes.

That exercise is also the best reliance lesson in the unit. Role-played voices invent things. A sharecropper character may cite a program that did not exist or quote a speech with the wrong date. The student has to check every factual claim against the packet and the textbook, and keep only what survives. Some claims will survive. The model's account of how the Agricultural Adjustment Act paid landowners to plant fewer acres is the kind of thing she should find confirmed in her documents and use. Accept what checks out. Throw out what does not. Asking a second chatbot to confirm the first is not a check.

Teachers should model this before students try it. Walk through one voice on a shared screen, check a claim against a document aloud, and show one you would keep and one you would throw out. Teaching students to plan, monitor, and evaluate their own thinking works best inside subject content with the teacher modeling it first ([Education Endowment Foundation](https://educationendowmentfoundation.org.uk/education-evidence/teaching-learning-toolkit/metacognition-and-self-regulation)).

Where students do not have their own accounts, they write the voices, the questions, and the checks, and the teacher runs them for the class. The students are still the ones directing.

## V. Knowledge Comes Inside the Loop

High school students often lack the background to judge an AI answer. The honest response is to teach that background while they work, not to keep AI out of the room until they have it.

Some effort is the learning. Reading a primary source and asking who wrote it, when, and why is the core practice of the history classroom, and it is what lets a student catch a role-played voice saying something no sharecropper in 1935 would have said. AI should not do that reading for her. Other effort is not the point of this unit. Formatting citations, finding page numbers, and fixing sentence mechanics are fine places for help.

Bastani and colleagues show what is at stake. High school mathematics students given unrestricted GPT-4 access improved during practice, then scored 17 percent worse than peers without access once the tool was removed. A version designed with teacher input to give hints rather than answers largely avoided that loss ([Bastani et al., 2025](https://doi.org/10.1073/pnas.2422633122)). Tool design and task design decide whether assistance builds skill or replaces it.

So start with the student's own attempt. Before AI enters, ask her to read two documents and say what standard of success each author seems to use. If she cannot, teach sourcing right there and try again. Then bring in the tool.

Reading supports, extra time, and other access accommodations stay in place throughout. Help with reading or expressing an answer is different from help supplying the reasoning, and the teacher should know which kind a student received.

## VI. Students Answer for Their Standard

A student who judges the New Deal only by national unemployment figures answers for leaving out the workers Social Security excluded. "The AI said it was mostly successful" does not tell a reader what success meant or who chose that meaning.

That responsibility fits a student's role. She is answerable for the claims in her essay and the standard behind them. She should be able to say which AI claims she kept, which she threw out, and why she stands behind her argument.

Adults answer for their parts. Teachers answer for the documents they chose, the prompt they wrote, and how they grade. School leaders answer for which tools are approved and how they are configured. The same structure runs through every role: whoever sets the standard answers for what follows from it.

## VII. Grading a Defended Argument

A polished essay no longer shows by itself what the student understood. Teachers need a little more evidence, gathered in a way that fits a class of thirty.

History prompts like this one have no single right answer, so grade the defense, not the verdict. Give credit when the student states her standard of success and why it fits, names a standard she considered and rejected, uses evidence from the documents that bears on her standard, identifies an AI claim she kept and how she checked it, and says what evidence would change her mind. Do not give credit for "there are many perspectives" without a position, and do not reward suspicion of AI for its own sake.

Add a short follow-up, spoken or written. Ask why she chose her standard, or ask about one AI claim she rejected. A student who owns the argument can answer quickly. A student who assembled the model's argument usually cannot. A rehearsed answer can sound strong and a nervous one can hide good reasoning, so no single response should settle a grade.

Then change the case. Ask students to judge whether their school's phone policy is a success. The move is the same one. Success by test scores, by what teachers report about attention, or by students who relied on their phones to coordinate rides and after-school jobs? A student who asks "success by what standard, and for whom?" without being prompted has carried the discipline to a new problem. A student who writes "it was mixed but largely successful" has not.

## VIII. A Pilot in One Unit

The original essay's five-step pilot fits inside an existing New Deal unit.

1. **Frame unaided.** Students read two or three documents and write a few sentences: what standard of success would you use, and why? This is also the readiness check. Teach sourcing to anyone who needs it.
2. **Direct AI against the frame.** Students write voices and questions, run them or have the teacher run them, and record which claims they checked and kept.
3. **Meet the misframed essay.** Students read the "mixed but largely successful" essay from Section II.
4. **Diagnose and revise.** They name the standard the essay used, what it left out, and whether that matters for their own argument, then revise.
5. **Defend.** A short follow-up on their standard and their reliance decisions, then the phone-policy case.

Try it with one class first. Review a handful of responses with a colleague who teaches the same course. Which step showed you something the essay alone would not have? Which response could you not interpret? Fix those before running it again.

## IX. What a Teacher Team Can Build

Teachers bring subject knowledge and knowledge of their students. They need enough practice with AI to direct it at the level they ask of students, see where it helps and where it invents, and model both in class. The quickest way to get there is for teachers to do the assignment themselves before students do.

A department can build this together. Have teachers run the same misframed essay, compare what they notice, and agree on what a strong defense looks like. One teacher may catch a missing perspective; another may notice that the documents are too hard for half the class to read. Both improve the task.

Keep a shared library: the prompt, the misframed AI answer and its frame flaw, one AI contribution worth accepting and the check behind it, the changed case, and notes from the first run. Include cases where the model was right. A library of nothing but traps teaches students to reject whatever a model says.

The same approach should reach other subjects and younger students, but it needs to be redesigned for them rather than simplified. A middle school science class might judge whether adding shade at a bus stop would help, which is a more concrete question with fewer frames in play. Each grade band and subject needs its own design, reviewed by teachers who teach it.

## X. Asking the Second Student

The second student's essay is better, and her teacher still has to find out whether its standard of success was hers. The way to find out is to ask her why she judged the New Deal by the workers it left out, and what evidence would change her mind. The first student is owed a lesson too, in how to put the tool to work on a question she has already thought hard about.

The next time you assign a prompt that asks students to evaluate something, write down the answer AI will most likely give and the standard hiding inside it. Find one claim in that answer worth keeping and a changed case your students care about. Teach it once, then read what students could explain.

## References

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. *Proceedings of the National Academy of Sciences, 122*(26), e2422633122. <https://doi.org/10.1073/pnas.2422633122>

Davies, G., & Derthick, M. (1997). Race and social welfare policy: The Social Security Act of 1935. *Political Science Quarterly, 112*(2), 217–235.

DeWitt, L. (2010). The decision to exclude agricultural and domestic workers from the 1935 Social Security Act. *Social Security Bulletin, 70*(4). <https://www.ssa.gov/policy/docs/ssb/v70n4/v70n4p49.html>

Dubin, J. C. (2024). The color of Social Security: Race and unequal protection in the crown jewel of the American welfare state. *Stanford Law & Policy Review, 35*, 104. <https://law.stanford.edu/wp-content/uploads/2024/02/Dubin-Publication-Ready-1.pdf>

Education Endowment Foundation. Metacognition and self-regulation. Teaching and Learning Toolkit. <https://educationendowmentfoundation.org.uk/education-evidence/teaching-learning-toolkit/metacognition-and-self-regulation>

Katznelson, I. (2005). *When affirmative action was white: An untold history of racial inequality in twentieth-century America*. W. W. Norton.

Lieberman, R. C. (1998). *Shifting the color line: Race and the American welfare state*. Harvard University Press.

OpenAI. Terms of use. <https://openai.com/policies/terms-of-use/>

Unemployment figures follow the standard historical estimates reported by the Bureau of Labor Statistics and Lebergott (24.9 percent in 1933, 14.3 percent in 1937, 19.0 percent in 1938). Other series that count work-relief employees as employed report lower figures, which is itself a question of standard worth raising with advanced students.
