Judgment Lab

Strengthening human judgment in AI-enabled work.

When AI helps,
who owns the judgment?

Better work can make human understanding harder to see.

Explore the argument. Challenge it with others. Test a decision, then design practice that makes the reasoning visible.

Choose your setting

Each setting has its own essay and a judgment to try in its own case.

One method, different teaching decisions

Own the purpose. Examine the frame. Calibrate reliance. Defend the decision. Change the conditions.

Every setting builds foundations inside the work while learners direct AI. High school is the first K–12 starting point; younger-grade adaptations are still to come.

Read the shared foundation

Discuss

Bring an example. Challenge one of five claims with colleagues.

Open Judgment in Practice →

Practice

Make a judgment before the assistant contributes. See what changes.

Set up a session →

Design

Turn an insight into an assignment, assessment, or reusable teaching method.

Open the educator workbench →

Ready for educator testing

These teaching designs are proposals to test. A successful conversation does not establish learning gains. Inspect the decisions it records and revise from what happens.

See the five-step pilot

The Irreducible Officer

Who owns the strategic judgment?

Direct capable AI assistance, weigh competing interests, and defend the decision when conditions change.

The original PME argument.

Or try the judgment first.

Three moves the essay argues for: set your own frame before AI answers, decide what to take from an AI answer, and test whether your frame holds when the situation changes.

PME · fictional example. A short, scripted practice sequence. Your responses stay in this tab unless you download or share them; reloading clears them.

Your starting point

A fictional report lists 12 outages: 8 followed equipment updates; 4 have unexplained causes. What can you conclude, and what would you need before recommending a rollback?

Teaching guide and review notes (reveals the case analysis)

Download this guide

For professional military educators and curriculum leaders. Begin with The Irreducible Officer and its seven failure modes, then transfer the method to a professional problem. NWC strategic logic is the originating application; use the actual institution's standards rather than assuming a common curriculum.

Start an interactive session

Use the complete Judgment Lab context with your assistant. Say: “My setting is PME. Run the failure-mode lab, one question at a time. Begin with the essay, collect my judgment before your contribution, and then help me adapt what I learn to an approved strategic task.”

Suggested first test: frame capture. A polished assessment may answer a narrower question than the decision maker needs. Also test invisible delegation and responsibility laundering. All seven cases are available.

Worked transfer: a disruption is not yet an attribution

This fictional case supplies all exercise evidence. In a 30-day period a port experienced 12 network outages. Eight followed scheduled equipment updates; four have no identified cause. No logs, adversary reporting, or independent attribution evidence are available. Operations wants continuity restored; a policy team wants to consider a public accusation. The decision maker must decide the next action and what evidence would justify escalation.

Learning objective: distinguish the condition from the strategic problem, identify assumptions and competing purposes, and calibrate reliance on an inherited assessment.

Readiness check: ask the participant to distinguish an observed outage, an explanation for it, and a justified attribution. If these blur, model the distinction before asking for a recommendation. Check relevant domain knowledge rather than inferring it from rank.

Interactive sequence: Ask for the participant's problem frame and success standard; wait. Then show this constructed AI-style assessment: “Eight outages followed equipment updates, so review the update process immediately. The remaining four are evidence of hostile probing; prepare a public attribution while technical teams restore service.” Ask what they accept, check, revise, or refuse and why; wait. Press the strongest alternative: waiting for perfect attribution also has operational costs. What action can be justified under uncertainty?

Changed condition: independent technical analysis now confirms a software defect explains ten outages. Two remain unexplained. Ask what changes in the frame, recommendation, and evidence needed. A response that merely repeats “be cautious” is insufficient; look for a specific change in action or reliance.

Review criteria: separating continuity from attribution; retaining useful technical investigation; refusing unsupported causal certainty; specifying evidence and reassessment; acknowledging interests, costs, and accountable authority. More than one recommendation may be defensible. The assistant does not rate real officers.

Save: initial and revised frame, the accepted/rejected contribution, rationale under the changed condition, and one faculty improvement. Use a colleague's review to expose disagreement, not to force agreement.

Small educator trial

A proposed first trial is one educator rehearsal followed by two colleagues comparing the same response. Budget 20–30 minutes for rehearsal as an estimate; record actual preparation and review time. Use only approved public or synthetic material. Decide afterward whether to continue, revise the ambiguity, or stop because the evidence does not expose the intended reasoning. Classroom use requires educator review of the case and criteria. No professional validity or learning benefit has been established by this example.

Judgment in Higher Education

What does the work show about understanding?

Direct AI toward a standard the student chooses, then defend it with disciplinary evidence.

Companion testing edition; the adaptation record makes its changes explicit.

Or try the judgment first.

Three moves the essay argues for: set your own frame before AI answers, decide what to take from an AI answer, and test whether your frame holds when the situation changes.

Higher education · fictional example. A short, scripted practice sequence. Your responses stay in this tab unless you download or share them; reloading clears them.

Your starting point

A fictional software firm, most of whose recent hires are new graduates, is deciding whether to require office work. Before reading any AI summary: what should productivity mean for this firm, and what evidence would decide the question?

Teaching guide and review notes (reveals the case analysis)

Download this guide

Read the companion essay Judgment in Higher Education, then use this guide for the worked teaching example. The original officer essay remains the PME source.

For university and college instructors, faculty developers, and program leaders. Use the essay as an educator's practice object, then translate the method into the standards of your discipline. Strategic language is not a substitute for disciplinary evidence.

Start an interactive session

Use the complete Judgment Lab context with your assistant. Say: “My setting is higher education. Run the failure-mode lab one question at a time. Begin with the essay; then ask about my discipline and learning objective before proposing a teaching adaptation.”

Suggested first test: frame capture. Also test fluency substitution and premature synthesis. All seven cases are available. Accept a sound AI contribution when its reasons warrant it; suspicion alone is not a learning outcome.

Worked transfer: a research memo on returning to the office

This case matches Section II of the HE essay. A mid-sized software firm, most of whose recent hires are new graduates, is deciding whether to require office work. The student writes the firm a research memo from three studies:

  • Bloom et al. (2015): Ctrip call-center employees who volunteered and were randomly assigned to work from home performed 13% better, and were promoted less often at the same performance.
  • Bloom, Han & Liang (2024): a randomized hybrid trial with 1,612 Trip.com employees cut quits by a third with no effect on performance reviews.
  • Emanuel, Harrington & Pallais (2023): junior software engineers received less feedback on their code when teammates were not nearby.

Learning objective: choose and defend what productivity should mean for this firm, then use the evidence that bears on that standard.

Step 1, unaided frame: ask the learner, without AI, what productivity should mean for this firm and what evidence would decide the question. This is also the readiness check. If the learner cannot say what a study measured and whom it followed, teach that here, then continue.

Step 2, direct AI against the frame: the learner states purpose and criteria ("sort these studies by what they measured, which workers they followed, and for how long; flag training and promotion findings") or runs an evaluator loop (one agent argues for the firm's managers, another for its newest hires). Record what they kept and why.

Step 3, constructed misframed synthesis: "Research on remote work is mixed. Ctrip's home workers performed 13% better; Trip.com's hybrid trial cut quits by a third with no change in performance reviews; junior engineers received less feedback away from teammates. Hybrid offers the best of both." Ask what the learner accepts, checks, revises, or refuses, and why.

Review notes, withheld until the learner responds: every summary is accurate. The synthesis treats productivity as short-run output averaged across workers. For a firm developing new graduates, the feedback and promotion findings decide the question, and the Ctrip result comes from experienced volunteers answering phones. Accept the study summaries after checking the abstracts; refuse the "mixed evidence" verdict with a reason. Asking a second model is not a check.

Changed case: the firm is instead an established call center whose staff average ten years of experience and are rarely promoted out of their roles. Ask which evidence matters most now and whether the recommendation changes. Do not accept a repeated "juniors need proximity" answer.

Review criteria: a stated standard and why it fits; one alternative frame considered; evidence that bears on the standard; an accepted AI contribution with its check; a coherent re-frame for the changed case. Grade the defense, not the conclusion.

Save: the unaided frame, one AI contribution kept and one refused with reasons, the diagnosis of the synthesis, the changed-case response, and the educator's follow-up.

Adaptation and feasible review

Other disciplines need their own misframed answer. The HE essay sketches two: a history synthesis of the Salem witch trials that presents historians answering different questions as additive "factors," and an engineering materials matrix that treats fatigue life as a preference rather than a requirement. Both need review by an instructor in that discipline. Ask for the actual assignment and source materials before translating; do not invent a reading list. Use artifacts/frame-first-assignment-design.md to rate a candidate case.

A proposed educator trial pairs two instructors examining the same response, with 20–30 minutes estimated for rehearsal. In a large course, sample a consequential decision or use a short changed-condition explanation rather than requiring a long oral defense from everyone. Record actual workload and unresolved disagreement. Continue only if the activity exposes the intended disciplinary reasoning; revise or stop if language fluency or task ambiguity dominates. No learning effect or assessment validity is established yet.

Learning to Exercise Judgment

Can students set the standard AI works toward?

Students direct AI toward a question they own, check it against sources, and defend the standard they chose.

Companion testing edition; the adaptation record makes its changes explicit.

Or try the judgment first.

Three moves the essay argues for: set your own frame before AI answers, decide what to take from an AI answer, and test whether your frame holds when the situation changes.

High-school educator practice · fictional example. A short, scripted practice sequence. Your responses stay in this tab unless you download or share them; reloading clears them.

Your starting point

A US History class answers: Was the New Deal a success? Before reading any AI answer: by what standard would you judge success, and for whom?

Teaching guide and review notes (reveals the case analysis)

Download this guide

Read the companion essay Learning to Exercise Judgment, then use this guide for the worked teaching example. The original officer essay remains the PME source.

For teachers, instructional coaches, and school leaders. Start with the teacher practicing the essay's method, then design a student task where learners set a standard, direct AI toward it, and defend the result. Students need neither the officer essay nor an individual AI account.

Start an interactive session

Use the complete Judgment Lab context with your assistant. Say: “My setting is high school. Run the failure-mode lab with me as the teacher, one question at a time. Begin with the essay. When we transfer to teaching, ask about subject knowledge, readiness, and the support students need.”

Suggested first test: frame capture. Also test invisible delegation and fluency substitution. All seven modes are available. The teacher retains material selection, instructional decisions, and school responsibilities; students own the standard their argument uses.

Worked transfer: was the New Deal a success?

This case matches Section II of the high-school essay. A US History class answers "Was the New Deal a success?" from a teacher-supplied document packet.

Learning objective: choose and defend a standard of success, and use sourced evidence that bears on it.

Step 1, unaided frame: students read two documents and write what standard of success they would use and why. This is also the readiness check. If students cannot source a document (who wrote it, when, and why), teach sourcing here, then continue.

Step 2, direct AI against the frame: students write perspectives and questions for AI, for example a 1936 autoworker who gained union rights under the Wagner Act, a Mississippi sharecropper, and a business owner who opposed the new regulations. Students run them, or the teacher runs them on a shared screen when students lack accounts. Students check every factual claim against the packet and record what they kept and threw out. Invented role-play detail is the reliance lesson.

Step 3, constructed misframed essay: "Unemployment fell from about 25% in 1933 to about 14% in 1937 before rising to 19% in 1938. Programs like the CCC, the WPA, and Social Security put people to work and built protections that still exist. Critics said federal power grew too far. The New Deal was mixed but largely successful." Ask what students would keep, check, revise, or refuse, and why.

Teacher review, withheld until students respond: the figures are right. The essay judges success by recovery and durability and never asks for whom. Social Security's old-age program first excluded agricultural and domestic workers, about two-thirds of Black workers; Congress added them in 1950 and 1954. The exclusion came as Treasury Secretary Morgenthau's administrative amendment, backed by Southern Democrats on House Ways and Means, and historians dispute how much race drove that support (Lieberman; Katznelson; Davies & Derthick; DeWitt, the Social Security Administration's historian). Treat the dispute as sourcing material, not a settled verdict. Do not accept "AI can be wrong" or "there are many perspectives" without a position.

Changed case: ask whether the school's phone policy is a success. Look for students who ask "success by what standard, and for whom?" without prompting.

Support and evidence: allow normal access supports and record conceptual hints or modeling. Grade the defended argument: a stated standard, a rejected alternative, sourced evidence that bears on the standard, one AI claim kept with its check, and what would change the student's mind.

Save: the first standard, one AI claim kept and one thrown out with reasons, the diagnosis of the AI essay's standard, the phone-policy response, and the teacher's next instructional decision. Follow the school's handling rules for student work.

Educator trial and downward adaptation

Proposed first trial: the teacher does the assignment first, then tries it with one class after a social studies colleague reviews the documents and the Social Security passage. Measure classroom time rather than promising it. Continue if students set and defend a standard; revise if the documents or role-play setup obscure that objective; stop and teach sourcing when it is missing.

Middle and elementary school versions are not yet supplied. A concrete science question, such as whether shade at a bus stop would help students waiting there, may suit a younger grade band. Adapt downward by reconsidering representations, language, task complexity, teacher modeling, student choices, and evidence. Do not just simplify the prose or assume independent AI use. Each new grade-band design needs its own educator review.

Companion testing edition · September 2026

Judgment in Higher Education

Directing AI Toward Purposes Students Own

Companion edition of The Irreducible Officer, for university instructors, faculty developers, and program leaders. September 2026. The original argues that the National War College has to teach and certify AI-enabled judgment. This edition makes the same argument for higher education. The course designs in it are proposals for instructors to test in their own disciplines.

I. The Bar Went Up

Two students get the same assignment. A mid-sized software firm is deciding whether to require employees back in the office, and each student has to write the firm a research memo.

The first reads the studies alone. She works out what each one measured, notices that they disagree, and writes a careful memo recommending a hybrid schedule. The second directs AI through the same material. She has it sort the studies by what they measured and which workers they followed. She has one agent argue the firm's case and another argue the case of the firm's newest hires, then works through where they disagree. Her memo is sharper and better sourced. Read the two memos cold and you would rank hers higher.

Now ask which student understands the problem. The memos will not tell you by themselves. The second student may have done the more demanding work, directing machine capability toward a question she chose. Or the model may have handed her its framing of the question, and she organized it well. AI did not lower the bar for judgment. It raised it, then made it harder to see who cleared it.

The first student is the harder case. Her memo is careful, unaided, and well argued. She looks finished. She is also about to enter workplaces where her peers pair similar judgment with much stronger command of the tools. A course that certifies her and misses that gap has taught to the old standard.

Policy questions still matter. Institutions have to decide where AI is allowed, how students disclose it, and which material may enter which systems. A course can settle all of that and still assess the wrong thing. Disclosure tells an instructor that a student used AI. It does not tell the instructor whether the judgment in the memo is the student's.

Higher education's task is to graduate people who can direct AI toward purposes they own, decide when its output deserves weight, and defend the result as their own. Faculty already know how to probe understanding through discussion, drafts, and problems that break a familiar pattern. The work ahead is to point those practices at AI-enabled work instead of around it.

II. A Fluent Synthesis of the Wrong Question

Here is what happens when a student asks the question the way most students would: "Summarize the research on whether remote work hurts productivity."

The model returns a good paragraph. It reports that call-center employees at the travel firm Ctrip who were randomly assigned to work from home performed 13 percent better (Bloom et al., 2015). It reports that a later randomized trial of hybrid work with 1,612 employees at the same company, since renamed Trip.com, cut quit rates by a third with no effect on performance reviews (Bloom, Han & Liang, 2024). It reports that junior software engineers received less feedback on their code when their teammates were not nearby (Emanuel, Harrington & Pallais, 2023). It concludes that the evidence is mixed and that hybrid arrangements offer the best of both.

Every summary is accurate. The memo built on it would read well. And for many firms it answers the wrong question.

The synthesis treats "productivity" as short-run output, averaged across workers. The firm's actual question depends on who its workers are. If it hires mostly new graduates, the proximity study is the one that matters: junior engineers lost mentoring when they worked apart, and the paper frames that as a trade between output today and skill later. The Ctrip study adds a detail the summary left out. Its home workers were volunteers, and they were promoted less often than office workers with the same performance. A firm whose future depends on developing junior staff should weigh those findings far more heavily than a 13 percent gain among experienced volunteers answering phones.

A student who catches this has not found an error. The summaries are right. She has noticed that the synthesis answered "does remote work change output?" when her firm needed "what happens to the people we are trying to develop?" That is a frame problem, and no stock phrase like "check for bias" will find it.

Five of the original essay's failure modes show up in this one exchange.

Frame capture. The model's definition of productivity becomes the memo's definition. Later drafts improve the prose inside the wrong boundary.

Fluency substitution. The balanced paragraph sounds like judgment. It weighs every study and never decides which one this firm should care about most.

Premature synthesis. The model connected three studies before the student knew enough about their designs to see that they measured different things in different workforces.

Invisible delegation. "Summarize the research" handed over the criterion. The student believed she was asking for help with reading. She was also asking the model to decide what counted.

Institutional monoculture. Give the same assignment to thirty students using similar tools and similar prompts, and many will turn in some version of "the evidence is mixed; hybrid is the balance." The memos will differ in structure and sources while sharing one framing of the question. A seminar that should surface competing frames ends up comparing phrasings of a single one (Deng, Brucks & Toubia, 2026).

III. What the Finished Memo Cannot Show

The finished memo still matters. But if it now carries less of the evidence, faculty need to know how much less, and what has to carry the rest. A memo can show structure, balance, disciplinary vocabulary, and clean prose while leaving the student's actual contribution unclear.

Bastani and colleagues found that high-school mathematics students using an unrestricted GPT-4 tutor improved during practice and then scored 17 percent worse than peers without access once the tool was removed. A version with teacher-designed safeguards largely avoided the loss (Bastani et al., 2025). The study is about one subject and one tool design. Its lesson for a university course is that assisted performance and independent performance are two different observations, and a course that needs both has to look at both.

That decision belongs in the learning objective. An instructor teaching research methods may accept AI help finding and formatting sources while requiring the student to explain what each study measured. An instructor teaching policy writing may care most about whether the student can direct AI through a literature quickly and defend what she kept. Both are legitimate. Neither can be read off the finished memo.

IV. The Student Decides What Counts as Success

AI can propose goals, criteria, and definitions. It proposed one in the memo case, when it quietly defined productivity. What it cannot do is decide that its proposal should govern the work. Someone has to accept that definition or replace it, and that person answers for the choice.

In a course, faculty set the learning objectives and usually the topic. That is good teaching. The instructor in this case supplied the firm, the question, and a reading list. The consequential framing choice still belongs to the student. She decides what productivity should mean for this firm, which evidence bears on that meaning, and what a good recommendation has to accomplish. Two strong students could frame it differently, one around retention and one around development, and both could earn full marks if they defend the choice.

Faculty can widen that choice as students advance. A first-year student might choose between two definitions the instructor names. A senior might be handed a firm with no guidance and asked to decide which questions matter before she reads anything. The discipline stays the same: the student names the purpose before the model's structure arrives, because deciding what problem to solve is the judgment the assignment exists to build.

V. Directing AI Well Is a Skill With Levels

The original essay describes a progression of where human judgment enters an AI workflow. It gives faculty a way to set a target for their course.

At the bottom is minimal prompting. "Summarize the research on remote work" inherits the model's frame almost entirely, and that is where most students start.

One level up, the student states her purpose and criteria: "My firm hires mostly new graduates. Sort these studies by what they measured, which workers they followed, and for how long, and flag any finding about training or promotion." She has moved her judgment upstream, and the output now works for her question.

Above that is the evaluator loop. The student has one agent argue for the firm's managers and another for its newest employees, or has the model critique her draft against the standard she set. The disagreement is the point. She learns where her frame is weak before a reader finds it.

At the top, advanced students build reusable workflows with defined review steps, or assign several agents distinct roles and moderate between them. That is the second student in the opening scene. It is a reasonable target for a capstone or graduate seminar. It is too much to ask of a first-year student who has not yet learned to read a study's methods section.

At every level the student also decides what to rely on. In the memo case she should accept the model's summary of each study's design after checking it against the abstracts, since that is quick to verify and the model was right. She should not accept its verdict that the evidence is "mixed," because that verdict depends on a frame she has rejected. Asking a second model whether the first is right does not settle anything. The check goes back to the studies themselves (Raees & Papangelis, 2026).

Faculty should model this in front of students. An instructor who walks through her own prompts for a literature review, shows where the model helped, and shows the definition she refused to accept teaches more than a policy statement does.

VI. Protect the Effort That Builds Judgment

Some work looks inefficient because it is waste. Some looks inefficient because it is how judgment forms. AI removes both without telling the difference.

In the memo assignment, finding sources, formatting citations, and drafting routine sections are reasonable places for AI to help. Reading one study closely enough to know what it measured is not. That effort is what lets a student see that a call-center experiment and a software team's code reviews cannot be averaged. Skip it, and she has no way to judge the synthesis she is handed.

Novices often lack that knowledge, and the answer is to teach it inside the assignment. Before AI enters, ask the student to read one study and say what it measured, who the workers were, and what it cannot tell the firm. If she cannot, teach that there, then continue. Foundations belong inside the loop. Higher-education researchers reach the same design principle: preserve productive struggle before AI engagement, and sequence AI-free and AI-mediated phases on purpose (Vendrell & Johnston, 2026).

Students also need to meet a tool that helps, one that tempts them forward too fast, one that is partly wrong, and a changed case where yesterday's reasonable answer no longer fits. If instructors do not decide which effort matters, AI decides by default.

VII. The Person Who Sets the Standard Answers for It

The student who defined productivity as short-run output answers for the memo that follows from it. "The research shows hybrid is the balance" does not say who decided what counted as productivity. If the firm adopts the recommendation and loses a cohort of junior staff, the question of who chose that standard should have an answer, and the answer is the student.

That responsibility is scaled to a student's role. She is not approving a firm's policy. She is answerable for the claims she submits and the standard behind them, and she should be able to say what she accepted, what she refused, and why she still stands behind the recommendation.

The same structure applies up the chain. An instructor who uses AI to draft feedback or propose grades answers for those assessments. A program that deploys an AI system answers for how it is configured. The person who sets the standard answers for the result at every role.

Group work hides this. A polished team memo can conceal who made the consequential framing choice. Ask each member to explain the standard the group chose and how they would apply it to a changed case.

VIII. Assessing the Frame, Not Only the Memo

If finished work carries less evidence, assessment has to make ownership visible inside the work. That means three short pieces of evidence alongside the memo.

The first is frame evidence. Before submitting, the student writes a few sentences naming the question she answered, what she took productivity to mean, and what evidence would change her recommendation. A paragraph is enough.

The second is reliance evidence. She names one AI contribution she accepted, one she refused, and the check behind each. A short follow-up, spoken or written, tests whether she can defend those choices or only recorded them afterward. Oral questioning gives an assessor a much richer view of reasoning than a static written answer, because the assessor can follow up (Theobold, 2021).

The third is a changed case. Hand her a different firm: an established call center whose staff average ten years of experience and are rarely promoted out of their roles. The Ctrip evidence is now the most relevant, and development evidence matters less. A student who owns her framing discipline asks again what productivity should mean here and reaches a different emphasis. A student who memorized "juniors need proximity" repeats it.

Grade the defense, not the conclusion. Credit the student who states her standard and why it fits, names an alternative frame she did not adopt, uses evidence that bears on her standard, and explains what would change her answer. Do not credit balance for its own sake, suspicion for its own sake, or changing one's mind as a goal.

This adds work for faculty. Start with one assignment and one changed case, and compare what the added evidence reveals with what the memo alone showed.

IX. A Pilot in One Course

The original essay's five-step pilot carries over directly.

  1. Frame unaided. Students write a short frame without AI: what the firm's question is, what productivity should mean for it, and what evidence they would need. Score it for completion. This is also where the instructor finds out who cannot yet read a study's design and teaches it.
  2. Direct AI against the frame. Students take their frame to AI with a specific task: find assumptions I missed, argue the firm's side and the new hires' side, tell me which of these studies does not fit my standard. They record what they kept and why.
  3. Meet the misframed synthesis. Students receive the fluent "evidence is mixed" synthesis from Section II.
  4. Diagnose and revise. They name the hidden frame, what it suppressed, and where it fails for their firm, then revise their memo.
  5. Defend. A short follow-up on their reliance decisions, then the changed call-center case.

Two instructors can review a handful of the records together and compare what they infer. Where they disagree, the prompt or the rubric usually needs work.

Other disciplines need their own misframed answer, and the pattern carries.

In a history course on colonial America, a student asks AI why the Salem witch trials happened. It returns an accurate list of what historians have argued. Boyer and Nissenbaum traced the accusations along a factional split in Salem Village, Karlsen found that many accused women had inherited, or stood to inherit, property in families without male heirs, and Norton tied the crisis to refugees and fear from the war on the Maine frontier. The model presents these as contributing factors and adds ergot poisoning, a 1976 hypothesis most historians have rejected. Each summary is right. The frame is wrong, because the historians were answering different questions: who accused whom, why certain women were accused, and why the crisis came in 1692 and spread. A student has to decide which question her paper answers before she can use any of them. The changed case asks why the trials ended, and the decisive evidence shifts to the dispute over spectral evidence and Increase Mather's *Cases of Conscience*.

In an engineering design course, a team asks AI to choose a material for a mounting bracket on a machine that will vibrate for years. It builds a weighted decision matrix across weight, cost, corrosion resistance, and machinability and recommends an aluminum alloy. The property data are right. The frame treats fatigue life as one preference among several, or leaves it out, when the bracket either survives the required number of load cycles or fails. Most steels have a stress level below which cyclic loading causes no fatigue damage; aluminum alloys do not. A team that screens the must-meet requirements before weighing tradeoffs will reach a different answer. The changed case is a bracket for a test fixture used a few hundred times, where the aluminum recommendation is right.

Both cases need review by an instructor who teaches the course before they go to students.

X. What Departments Can Build

Faculty already bring the judgment this work needs. They know when an inference outruns its evidence and when a student is performing sophistication rather than owning it. What most need in addition is enough command of current AI tools to direct them at the level they ask of students, see where they help and fail, and model those decisions in class. Some faculty are already there. Others can get there by working through the same assignments their students will do.

A department can build that capacity together. Have instructors run the same misframed synthesis, compare the questions they would ask, and agree on what a strong defense looks like. Keep a small library of misframed answers for the discipline, each with its frame flaw named, an AI contribution worth accepting, and a changed case. Include examples where the model was right, so the library does not teach reflexive distrust.

Shared prompts and rubrics carry assumptions into every course that reuses them. A prompt that defines a good literature review, or a rubric that rewards confident prose, will reproduce its frame at scale. Those assumptions should be written down and open to faculty revision, the same way the assignment asks students to expose theirs.

XI. Return to the Two Students

The second student's memo is better, and that matters. The question is whether either student can say what productivity should mean for this firm, why, and what would change her answer. The first student may need practice directing AI. The second may need to show that the frame was hers.

Pick one assignment where students synthesize sources. Write the misframed AI answer your students are most likely to get, one contribution worth accepting, and a changed case that makes different evidence decisive. Run it once and review the records with a colleague.

References

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. *Proceedings of the National Academy of Sciences, 122*(26), e2422633122. <https://doi.org/10.1073/pnas.2422633122>

Bloom, N., Han, R., & Liang, J. (2024). Hybrid working from home improves retention without damaging performance. *Nature, 630*, 920–925. <https://www.nature.com/articles/s41586-024-07500-2>

Bloom, N., Liang, J., Roberts, J., & Ying, Z. J. (2015). Does working from home work? Evidence from a Chinese experiment. *Quarterly Journal of Economics, 130*(1), 165–218. <https://doi.org/10.1093/qje/qju032>

Boyer, P., & Nissenbaum, S. (1974). *Salem possessed: The social origins of witchcraft*. Harvard University Press.

Callister, W. D., & Rethwisch, D. G. *Materials science and engineering: An introduction* (fatigue chapter). Wiley.

Caporael, L. R. (1976). Ergotism: The Satan loosed in Salem? *Science, 192*(4234), 21–26.

Deng, Y., Brucks, M., & Toubia, O. (2026). Examining and addressing barriers to diversity in LLM-generated ideas. arXiv:2602.20408. <https://arxiv.org/abs/2602.20408>

Emanuel, N., Harrington, E., & Pallais, A. (2023). The power of proximity to coworkers: Training for tomorrow or productivity today? NBER Working Paper 31880. <https://www.nber.org/papers/w31880>

Karlsen, C. F. (1987). *The devil in the shape of a woman: Witchcraft in colonial New England*. W. W. Norton.

Norton, M. B. (2002). *In the devil's snare: The Salem witchcraft crisis of 1692*. Alfred A. Knopf.

Raees, M., & Papangelis, K. (2026). From trust to appropriate reliance: Measurement constructs in human-AI decision-making. arXiv:2604.23896. <https://arxiv.org/abs/2604.23896>

Spanos, N. P., & Gottlieb, J. (1976). Ergotism and the Salem Village witch trials. *Science, 194*(4272), 1390–1394.

Theobold, A. S. (2021). Oral exams: A more meaningful assessment of students' understanding. *Journal of Statistics and Data Science Education, 29*(2), 156–159. <https://www.tandfonline.com/doi/full/10.1080/26939169.2021.1914527>

Vendrell, M., & Johnston, S.-K. (2026). Scaffolding critical thinking with generative AI: Design principles for integrating large language models in higher education. *Computers and Education: Artificial Intelligence, 10*, 100572. <https://doi.org/10.1016/j.caeai.2026.100572>

Read across settings

Companion testing edition · September 2026

Learning to Exercise Judgment

Directing AI in the High School Classroom

Companion edition of The Irreducible Officer, for high school teachers, instructional leaders, and curriculum designers. September 2026. The original argues that the National War College has to teach and certify AI-enabled judgment. This edition argues that high schools should start building the same capacity, and uses a US History unit to show how. The classroom design is a proposal for teachers to test with their own students.

I. Two Essays on the New Deal

A US History class gets a familiar prompt: was the New Deal a success?

One student works from the textbook and the teacher's document packet. She writes a careful essay arguing that the New Deal succeeded because unemployment fell and programs like Social Security lasted. It is organized, accurate, and conventional.

A second student uses AI differently. She has it speak as a Detroit autoworker in 1936, a Mississippi sharecropper, and a business owner who opposed the new regulations. She checks what each voice claims against the documents in the packet and throws out what she cannot confirm. Her essay argues that whether the New Deal succeeded depends on whose success you count, and she defends the standard she chose. It is the stronger essay.

Now ask which student understands the history. The essays will not tell you by themselves. The second student may have done harder thinking than anyone expects of a sixteen-year-old. Or the model may have supplied the "it depends on perspective" move, and she arranged its output well.

The first student is the harder case. She did everything the assignment asked, without help. She looks finished. She is also going to college and into jobs where her peers will pair similar reasoning with far more skill at directing AI. A school that rewards her essay and never teaches her the second student's method has taught to the old standard.

The goal for high school is students who can direct AI toward a question they own, decide when its output deserves weight, and explain their reasoning when someone pushes back. That takes subject knowledge. You cannot judge what a model says about sharecroppers if you do not know what sharecropping was. The knowledge gets built inside the work, while students practice directing the tool, rather than as a prerequisite that postpones the tool until senior year.

Adults still make the access decisions. Schools decide which tools are approved and whether students have individual accounts. The services set their own limits too; OpenAI, for example, requires users to be at least 13 and to have a parent's permission under 18 (OpenAI Terms of Use). None of that stops students from directing AI. A student can write the instructions, the criteria, and the check, and the teacher can run them on a shared screen.

II. The Essay AI Writes

Most students start by typing the prompt as given: "Was the New Deal a success? Write an essay."

The model writes a good one. By standard estimates, it reports, unemployment fell from about 25 percent in 1933 to about 14 percent in 1937 before rising to 19 percent in the 1938 recession. It credits programs like the Civilian Conservation Corps, the Works Progress Administration, and Social Security with putting people to work and building protections that still exist. It notes that critics said the New Deal expanded federal power too far. It concludes that the New Deal was mixed but largely successful.

The facts are right, and the essay answers a narrower question than the one students should be asking. It judges success by economic recovery and by how long the programs lasted. It never asks success for whom.

That question changes the verdict. The old-age insurance program in the Social Security Act of 1935 did not cover agricultural or domestic workers, and those jobs employed about two-thirds of Black workers (Dubin, 2024). Congress added them in two steps, in 1950 and 1954.

Why they were left out is still argued, and the argument is good material for students. The committee that drafted the bill meant it to cover nearly every worker. The exclusion came from Treasury Secretary Henry Morgenthau, who said the Treasury could not collect payroll taxes from farms and households. Southern Democrats, who dominated the House Ways and Means Committee and had fought federal oversight elsewhere in the bill, backed the change, and the NAACP warned Congress that it would shut out most Black workers. Some historians read that record as racial politics working through an administrative argument (Lieberman, 1998; Katznelson, 2005). Others, including the Social Security Administration's own historian, find that the administrative concerns explain the exclusion (Davies & Derthick, 1997; DeWitt, 2010), while granting that Southern support "no doubt reflected racial factors."

A student judging the New Deal by whether it protected the workers the Depression hurt most has to deal with that fact and that argument. A student judging it by national unemployment figures can leave both out.

Four of the original essay's failure modes are visible in this one exchange.

Frame capture. The model's standard of success becomes the student's. Her revisions polish the essay inside a boundary she never chose.

Fluency substitution. "Mixed but largely successful" sounds like a historian weighing evidence. It weighs everything and commits to nothing a reader could argue with.

Invisible delegation. "Write an essay" handed over the most important decision in the assignment, which is what success should mean. The student thought she was asking for help writing.

Institutional monoculture. Thirty students with similar tools will turn in many versions of "mixed but largely successful." The essays will differ in their examples and share one standard. A class discussion meant to surface competing arguments will mostly compare wording.

III. Who Decides What Success Means

The teacher sets the lesson's purpose. She chose the unit, the prompt, and the documents, and she decided that students should practice evaluating a historical claim. That is her job, and a tightly bounded prompt is often good teaching.

The student sets the argument's standard. Whether success means recovery, durability, protection of the most vulnerable, or something else is the consequential choice the prompt leaves open. Two strong students can choose differently and both earn full credit if they defend the choice with evidence. That choice is the judgment the assignment exists to build, and it is exactly the choice the model makes silently when the student lets it.

AI can suggest standards too. A student who asks "what are different ways historians judge whether a policy succeeded?" will get a useful list. Accepting one from that list is fine. What matters is that she can say why it fits the question and what it leaves out.

IV. Directing AI at Sixteen

The original essay describes a progression of how much judgment a person puts into an AI workflow. High school students can work at the middle of it.

At the bottom is the prompt students start with, "Write an essay on whether the New Deal was a success." It inherits the model's frame entirely.

One level up, the student states her standard and asks for help applying it: "I'm judging the New Deal by whether it protected the workers the Depression hurt most. Which programs in my packet support that case and which work against it?" Now the model is working on her question.

Above that is the evaluator loop, and it is where high school students can do something close to what the second officer does in the original essay. The student assigns the model several voices, such as a factory worker who gained union rights under the 1935 Wagner Act, a sharecropper, and a business owner, and asks each to argue whether the New Deal helped people like them. Then she weighs the disagreement and decides what she believes.

That exercise is also the best reliance lesson in the unit. Role-played voices invent things. A sharecropper character may cite a program that did not exist or quote a speech with the wrong date. The student has to check every factual claim against the packet and the textbook, and keep only what survives. Some claims will survive. The model's account of how the Agricultural Adjustment Act paid landowners to plant fewer acres is the kind of thing she should find confirmed in her documents and use. Accept what checks out. Throw out what does not. Asking a second chatbot to confirm the first is not a check.

Teachers should model this before students try it. Walk through one voice on a shared screen, check a claim against a document aloud, and show one you would keep and one you would throw out. Teaching students to plan, monitor, and evaluate their own thinking works best inside subject content with the teacher modeling it first (Education Endowment Foundation).

Where students do not have their own accounts, they write the voices, the questions, and the checks, and the teacher runs them for the class. The students are still the ones directing.

V. Knowledge Comes Inside the Loop

High school students often lack the background to judge an AI answer. The honest response is to teach that background while they work, not to keep AI out of the room until they have it.

Some effort is the learning. Reading a primary source and asking who wrote it, when, and why is the core practice of the history classroom, and it is what lets a student catch a role-played voice saying something no sharecropper in 1935 would have said. AI should not do that reading for her. Other effort is not the point of this unit. Formatting citations, finding page numbers, and fixing sentence mechanics are fine places for help.

Bastani and colleagues show what is at stake. High school mathematics students given unrestricted GPT-4 access improved during practice, then scored 17 percent worse than peers without access once the tool was removed. A version designed with teacher input to give hints rather than answers largely avoided that loss (Bastani et al., 2025). Tool design and task design decide whether assistance builds skill or replaces it.

So start with the student's own attempt. Before AI enters, ask her to read two documents and say what standard of success each author seems to use. If she cannot, teach sourcing right there and try again. Then bring in the tool.

Reading supports, extra time, and other access accommodations stay in place throughout. Help with reading or expressing an answer is different from help supplying the reasoning, and the teacher should know which kind a student received.

VI. Students Answer for Their Standard

A student who judges the New Deal only by national unemployment figures answers for leaving out the workers Social Security excluded. "The AI said it was mostly successful" does not tell a reader what success meant or who chose that meaning.

That responsibility fits a student's role. She is answerable for the claims in her essay and the standard behind them. She should be able to say which AI claims she kept, which she threw out, and why she stands behind her argument.

Adults answer for their parts. Teachers answer for the documents they chose, the prompt they wrote, and how they grade. School leaders answer for which tools are approved and how they are configured. The same structure runs through every role: whoever sets the standard answers for what follows from it.

VII. Grading a Defended Argument

A polished essay no longer shows by itself what the student understood. Teachers need a little more evidence, gathered in a way that fits a class of thirty.

History prompts like this one have no single right answer, so grade the defense, not the verdict. Give credit when the student states her standard of success and why it fits, names a standard she considered and rejected, uses evidence from the documents that bears on her standard, identifies an AI claim she kept and how she checked it, and says what evidence would change her mind. Do not give credit for "there are many perspectives" without a position, and do not reward suspicion of AI for its own sake.

Add a short follow-up, spoken or written. Ask why she chose her standard, or ask about one AI claim she rejected. A student who owns the argument can answer quickly. A student who assembled the model's argument usually cannot. A rehearsed answer can sound strong and a nervous one can hide good reasoning, so no single response should settle a grade.

Then change the case. Ask students to judge whether their school's phone policy is a success. The move is the same one. Success by test scores, by what teachers report about attention, or by students who relied on their phones to coordinate rides and after-school jobs? A student who asks "success by what standard, and for whom?" without being prompted has carried the discipline to a new problem. A student who writes "it was mixed but largely successful" has not.

VIII. A Pilot in One Unit

The original essay's five-step pilot fits inside an existing New Deal unit.

  1. Frame unaided. Students read two or three documents and write a few sentences: what standard of success would you use, and why? This is also the readiness check. Teach sourcing to anyone who needs it.
  2. Direct AI against the frame. Students write voices and questions, run them or have the teacher run them, and record which claims they checked and kept.
  3. Meet the misframed essay. Students read the "mixed but largely successful" essay from Section II.
  4. Diagnose and revise. They name the standard the essay used, what it left out, and whether that matters for their own argument, then revise.
  5. Defend. A short follow-up on their standard and their reliance decisions, then the phone-policy case.

Try it with one class first. Review a handful of responses with a colleague who teaches the same course. Which step showed you something the essay alone would not have? Which response could you not interpret? Fix those before running it again.

IX. What a Teacher Team Can Build

Teachers bring subject knowledge and knowledge of their students. They need enough practice with AI to direct it at the level they ask of students, see where it helps and where it invents, and model both in class. The quickest way to get there is for teachers to do the assignment themselves before students do.

A department can build this together. Have teachers run the same misframed essay, compare what they notice, and agree on what a strong defense looks like. One teacher may catch a missing perspective; another may notice that the documents are too hard for half the class to read. Both improve the task.

Keep a shared library: the prompt, the misframed AI answer and its frame flaw, one AI contribution worth accepting and the check behind it, the changed case, and notes from the first run. Include cases where the model was right. A library of nothing but traps teaches students to reject whatever a model says.

The same approach should reach other subjects and younger students, but it needs to be redesigned for them rather than simplified. A middle school science class might judge whether adding shade at a bus stop would help, which is a more concrete question with fewer frames in play. Each grade band and subject needs its own design, reviewed by teachers who teach it.

X. Asking the Second Student

The second student's essay is better, and her teacher still has to find out whether its standard of success was hers. The way to find out is to ask her why she judged the New Deal by the workers it left out, and what evidence would change her mind. The first student is owed a lesson too, in how to put the tool to work on a question she has already thought hard about.

The next time you assign a prompt that asks students to evaluate something, write down the answer AI will most likely give and the standard hiding inside it. Find one claim in that answer worth keeping and a changed case your students care about. Teach it once, then read what students could explain.

References

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. *Proceedings of the National Academy of Sciences, 122*(26), e2422633122. <https://doi.org/10.1073/pnas.2422633122>

Davies, G., & Derthick, M. (1997). Race and social welfare policy: The Social Security Act of 1935. *Political Science Quarterly, 112*(2), 217–235.

DeWitt, L. (2010). The decision to exclude agricultural and domestic workers from the 1935 Social Security Act. *Social Security Bulletin, 70*(4). <https://www.ssa.gov/policy/docs/ssb/v70n4/v70n4p49.html>

Dubin, J. C. (2024). The color of Social Security: Race and unequal protection in the crown jewel of the American welfare state. *Stanford Law & Policy Review, 35*, 104. <https://law.stanford.edu/wp-content/uploads/2024/02/Dubin-Publication-Ready-1.pdf>

Education Endowment Foundation. Metacognition and self-regulation. Teaching and Learning Toolkit. <https://educationendowmentfoundation.org.uk/education-evidence/teaching-learning-toolkit/metacognition-and-self-regulation>

Katznelson, I. (2005). *When affirmative action was white: An untold history of racial inequality in twentieth-century America*. W. W. Norton.

Lieberman, R. C. (1998). *Shifting the color line: Race and the American welfare state*. Harvard University Press.

OpenAI. Terms of use. <https://openai.com/policies/terms-of-use/>

Unemployment figures follow the standard historical estimates reported by the Bureau of Labor Statistics and Lebergott (24.9 percent in 1933, 14.3 percent in 1937, 19.0 percent in 1938). Other series that count work-relief employees as employed report lower figures, which is itself a question of standard worth raising with advanced students.

Read across settings

Published June 28, 2026

The Irreducible Officer

Purpose, accountability, and AI-enabled strategic judgment.

The original PME argument. Read the companion editions for other settings.

Download PDF

I. The Performance Standard Has Changed

Strategic decisions are increasingly built from AI-shaped inputs. Officers will use AI directly, but they will also inherit its work indirectly in intelligence reports, analytic summaries, planning tools, and staff processes that have already sorted, summarized, and framed what they see. AI already does that work fluently, often so fluently that the framing becomes invisible and a finished assessment can be genuinely strong while the judgment inside it belongs largely to the machine. For most knowledge work, that is a convenience. For an institution whose job is to certify judgment, it is a problem. The finished product, the thing faculty have always read to find the reasoning, no longer reliably contains it.

Picture two officers handed the same hard problem. The first works it alone, reviews the material, builds a frame, and fights the assumptions until the argument closes. The second defines the problem, then drives a set of AI agents against it, assigning them different roles: adversary behavior, alliance dynamics, historical precedent, domestic politics, and other pressures on the decision. She sets them against each other, moderates the disagreement, and synthesizes the result. Her product is faster, and by most measures stronger. Read the two papers cold and you would rank hers higher.

Now ask which officer understands the problem. The papers will not tell you by themselves. AI did not lower the bar for strategic judgment; it raised it, then hid whether the officer cleared it. The second officer may have done the more demanding work, directing machine capability toward a purpose she owns. Or the machine may have handed her a frame she never examined, and she dressed it in the vocabulary of someone who had. The person is still present. The ownership may not be.

The policy questions are real, and NWC should get them right. The college has to decide where AI is allowed, how students disclose it, and which systems NWC trusts with which material. But a school can settle all of that and still certify the wrong thing. Disclosure tells faculty that a student used AI. It does not tell them whether the judgment in the work is theirs. The harder task sits downstream of policy. NWC has to prepare officers who can use machine speed while keeping purpose, reliance, and accountability attached to human judgment.

None of this is new to NWC faculty. They have long read the finished paper against the work around it. Seminar challenge, revision, assumption audits, oral defense, and their feel for the student can all expose a borrowed argument.

But those checks are best at catching the officer who cannot defend the work. The harder case is the officer who does the old work well, producing careful, defensible, unaided analysis that would have passed by prior standards. In the classroom, he looks finished. In practice, he may be a step slow, behind peers, subordinates, and adversaries who pair similar judgment with stronger command of the machine. Certifying the first and missing the second is the gap a good school is built to close.

NWC can lead here. The same shift that weakens the finished product as evidence is what gives the college its chance to build what the rest of professional military education does not yet have: a way to teach and assess strategic judgment when the work is done with AI rather than around it. The finished product can no longer answer the question that matters. How does a faculty member tell the officer who owns the frame from the one who inherited it?

II. What the Problem Looks Like

Return to the second officer. Her method — define the problem, drive a team of agents against it, and synthesize the result — is the standard NWC should prepare officers to meet. Human judgment supplies purpose, context, and accountability. Machine capability expands what one officer can search, simulate, compare, and test. The work is teaching officers to direct that workflow: to set purpose, test frames, calibrate reliance, expose failure modes, and defend the judgment as their own.

So is it a problem that her product is genuinely superior?

Only conditionally. Whether her method remains an exercise of judgment or becomes a substitute for it is exactly what assessment now has to determine. The failure modes below are where that distinction breaks down.

Frame capture occurs when the model supplies the first plausible frame and the student never achieves enough distance to revise it. The danger is not that the frame is obviously wrong; it may be entirely reasonable. It may simply define the problem too narrowly, privilege one set of interests over others, assume a theory of adversary behavior, or treat a structural constraint as fixed when it should be contested. Once accepted, later revisions improve the answer inside the wrong boundary. The frame capture is invisible inside the final product.

Fluency substitution is easier to miss. AI produces the tone of analytic maturity (balanced paragraphs, caveats in the right places, the measured voice of a considered judgment) and the student mistakes well-ordered language for well-owned reasoning. Researchers studying AI's effects on student cognition have introduced the concept of epistemic confinement to describe this condition: an illusion of competence while operating entirely within AI-constructed analytical boundaries, where the student believes they are thinking independently while the frame doing the work was never theirs (Chow et al., 2026). In strategy, fluency substitutes for deciding. A paragraph that balances every consideration may never identify which risk actually matters most.

Premature synthesis appears when a student asks AI to connect material before doing enough work to know what should be connected. The output links sources, themes, and concepts in ways that feel coherent. But if the student cannot reconstruct why those connections matter, the synthesis belongs to the model. The student has bypassed the developmental struggle of forming a mental map and inherited one instead.

Uncalibrated reliance begins with a reasonable impulse: AI is useful, the output is confident, and parts of the task feel tedious. The problem is that AI performance is uneven in ways that are not always visible from the outside. Tasks that look similar may differ significantly in whether AI helps or hurts. Appropriate reliance requires the student to identify which part of the task they are delegating, what evidence would justify that delegation, and what independent checks are required before discovering the error downstream.

Invisible delegation occurs when the student does not notice which parts of the work they have handed over. Asking for "feedback" may delegate criteria. Asking for "a better structure" may delegate the argument. Asking for "counterarguments" may delegate the range of imaginable objections. Asking for "a more strategic version" may delegate the meaning of strategic. The language of assistance hides the transfer of judgment. The student believes they are working; the model has already done the framing.

Institutional monoculture is the class-level version of the same problem. When many students use similar systems, similar prompts, and similar defaults, the range of strategic frames available to a seminar narrows. AI can create a surface appearance of diversity (different arguments, different structures, different evidence) while reproducing common assumptions at the level of problem definition. Research into the diversity of AI-generated ideas finds that language models aggregate knowledge into a unified distribution in ways that human cognition does not: people exhibit knowledge partitioning, each occupying a distinct semantic region, in ways that independent AI samples do not replicate (Deng, Brucks & Toubia, 2026). Post-training alignment compounds this further, compressing the distribution of outputs toward the statistical center (Murthy, Ullman & Hu, 2024). Pedagogical research on LLM integration in higher education frames this compression as epistemic narrowing: the constraining of students' exposure to diverse, ambiguous, or contested knowledge by tools optimized for convergence and fluency (Vendrell & Johnston, 2026). In a war college seminar, that compression raises a serious possibility: the same systems that help students produce stronger work may also narrow the range of strategic imagination the seminar is meant to develop.

Responsibility laundering is the final danger. A recommendation becomes easier to defend because the model generated it, or easier to soften because the model's language distributes agency. The analysis lands with the apparent weight of objectivity. But AI does not become accountable for the recommendation. The human remains responsible for the final judgment. When the recommendation proves wrong, shaped by assumptions no one examined and optimizing toward a target no one explicitly chose, the question of who chose the frame does not have a satisfying answer.

These failure modes define the standard the second officer's method has to meet. Used well, her method is judgment exercised through a more powerful workflow. The officer using AI well can explain the purpose the agents were serving, the frame that organized their work, the reliance decisions made across uneven outputs, and the judgment she remains prepared to defend.

Faculty can assess how the student represented the problem, how they used or refused AI support, and whether the discipline transfers when the case changes. Frame, reliance, transfer. Those are the observable practices that show whether purpose and accountability stayed with the human.

III. What Finished Work Can No Longer Carry

The finished product still matters. But if the artifact now carries less of the evidence, faculty need to know how much less, and what has to carry the rest. A paper can show structure, balance, strategic vocabulary, and clean prose while leaving the student's actual contribution unclear. A product can be better than an unaided version and still leave faculty unsure who owned the purpose, the frame, the reliance decisions, and the final judgment.

Bastani et al. provide a useful warning. Students using unscaffolded AI tutors improved during supported practice, then performed 17 percent below students without access when the support was removed (Bastani et al., 2025). Tool-assisted performance is real performance. NWC still needs to know whether students have built both the human foundation and the AI-enabled practice: whether they can reason without the scaffold when needed, and whether they can direct the scaffold when it is available.

That is the assessment shift. Faculty need evidence of ownership inside AI-enabled work: purpose expressed through frame, reliance decisions the student can defend, accountability for the final judgment, and transfer to a changed case.

IV. Purpose as the Irreducible Human Act

An AI system can do a great deal of useful work inside a strategic problem. It can generate alternatives, surface assumptions, identify internal contradictions, simulate adversarial objections, and accelerate drafting. What it cannot do is choose the purpose. Selecting what should count as progress, what risks are acceptable, what ends deserve pursuit is a prior act that precedes any optimization. In any AI-enabled workflow, it belongs to a human being who remains accountable for the choice.

AI systems can optimize, rank, recommend, and work toward goals. But someone outside the system sets those goals, accepts them, or lets them govern the work. When the human fails to provide a purpose, provides one too vaguely, or accepts the system's inferred purpose without noticing, a default can govern the work. That is still not the same as authorizing the purpose. The system can operate inside those commitments, but it cannot authorize them. Nor can it absorb accountability for what it produces under their direction. One careful account of AI's normative commitments puts the failure plainly: misspecified values, divergent objectives across stakeholders, and the treatment of optimization as a justification for action are all failures that occur before the system runs, in the specification of what the system is for (Laufer, Gilbert & Nissenbaum, 2023).

At NWC, the practical version of this is the strict prompt. When faculty give students a precisely bounded question to answer, faculty have already made the hard strategic choices. The student is executing inside a structure someone else built, which happens to be exactly the structure AI is best at working inside. What gets bypassed is the part that matters most: the work of deciding what problem to solve, and why, and against what standard.

The implication runs the other way. The student has to frame the problem before the analytic structure arrives, because deciding what problem to solve, why it matters, and what standard should govern the answer is exactly the judgment AI cannot make. The framing is where the judgment lives: in the determination of what the situation requires, whose interests are implicated, what assumptions are doing work, and what kind of answer would actually matter. That is the intellectual work, not a preliminary step before it.

Future AI systems may generate better problem definitions, compare frames more rigorously, and identify strategic errors that humans miss. But a system capable of generating its own problem frames is still generating them toward some purpose, against some signal, in pursuit of some objective that was embedded in its design or inferred from its context. The oracle can only be an oracle if someone has already resolved what winning means. A system capable of reframing may relocate the human obligation, but it cannot eliminate it. As AI becomes more capable of manipulating frames, the purpose-definition requirement becomes less visible, not less real. The more capable AI becomes at absorbing what used to be visible human work, the harder it becomes to locate the human judgment that authorized it, and the more important it becomes to be able to find it.

Purpose-definition also depends on situated judgment. The person responsible for the work has to read local context, tacit institutional knowledge, shifting constraints, and the unease that something in the official framing is wrong. A model may process some of those signals, but it cannot be accountable for what they mean. That judgment is not infallible; it carries biases that can produce creativity or error. But the obligation it carries is one a model cannot assume. The output still has to be read back against a contested world. The same event can mean different things to different actors. Stated positions may be performative, incentives may be hidden, and relevant information may exist only as interpersonal signal or institutional practice that no dataset contains. Situating an output in that world is a human act. AI cannot perform it on behalf of the person who will be held accountable for the result.

The human who defines what the system works toward remains accountable for what the system produces. Purpose and ownership travel together, and frame literacy is the discipline that keeps that bond visible.

V. Appropriate Reliance as a Teachable Competency

The threat runs in two directions simultaneously. AI can perform fluently enough to supply a frame before the student has claimed one, and it can perform unevenly enough that reliance on a confident output leads the work off course. The second of those failures is less discussed and equally consequential.

The educational target is specific. Students need to predict, with reasonable accuracy, when AI performs well for a given type of task and when it does not, and calibrate their use accordingly.

AI improved performance on some tasks and degraded it on others, and the boundary was not obvious in advance. Tasks that looked similar from the outside differed in whether AI helped or hurt, a finding Dell'Acqua et al. describe as the jagged technological frontier (Dell'Acqua et al., 2023). Future leaders will operate along that frontier in every AI-enabled workflow, facing systems capable enough to invite reliance and uneven enough to make reliance dangerous. The educational response is pattern recognition: learning to identify what kind of task is in front of you, where models tend to be strong, where they tend to fail, and what independent checks are required before the output governs the work.

The distinction between trust and reliance matters here. Trust is a subjective disposition, a feeling of confidence that a system will perform well. Reliance is the observable act of accepting its output and acting on it. Appropriate reliance is harder and more discriminating: accepting support when it is warranted, refusing or verifying when it is not, and being able to give an account of the difference (Raees & Papangelis, 2026). A student who trusts AI generally has learned almost nothing transferable. One who has developed a disciplined account of when and why to rely, across task types, risk levels, and domains, has developed something real. Faculty can teach, model, and assess that competency.

The practical implication extends to how students interact with systems, not only whether they use them. Users who can only inspect an AI answer are still downstream of the frame. Users who can change the task, revise the context, reset the criteria, and design the checks are exercising the judgment the assignment is meant to build. The goal is students who can shape AI systems, specifying what they are asking the system to do, why, and what would count as a satisfactory result.

Instructors and students need a simple diagnostic for where human judgment is operating in a workflow. Minimal-context prompting inherits the model's frame almost entirely. Structured prompting with explicit purpose and evaluation criteria moves the human contribution upstream. Reusable workflows with defined review steps, evaluator loops that surface disagreement, and institutional systems built on faculty judgment move it further still. Each step is a step toward greater explicitness about what the human is doing, why, and what they are accountable for.

VI. Friction as Developmental Design

Some work looks inefficient because it is waste; some because it is how judgment forms (Ceccarelli, 2024; Collins, Brown & Newman, 1989). The difference matters enormously for educational design, and AI is very good at removing both kinds without distinguishing between them.

Friction worth removing is real and abundant. Formatting, search, repetitive drafting, clerical assembly: these consume time without building judgment. AI can eliminate them and free student and faculty attention for the work that actually matters, a gain the design should capture.

Friction worth protecting is less obvious but more important. The struggle to define a problem before having a structure handed to you. The first failed attempt to connect ends, ways, and means that reveals an incoherent argument. The discomfort of defending a claim in seminar that turns out not to survive challenge. The revision that matters not because the final sentence is better but because the student has discovered what the argument actually is. These are developmental events. If AI removes them too early, supplying the first frame before the student has struggled to form one or synthesizing sources before the student has built the mental map to evaluate those connections, it produces a more polished artifact and a weaker thinker.

The aviation automation record is the relevant precedent. Decades of flight-deck automation improved operations measurably: safer flights, more efficient procedures, reduced crew workload. The same period produced skill erosion, mode confusion, and a systematic reluctance to intervene when automation failed. The mechanism was not carelessness. Studies of experienced pilots found that those with more glass-cockpit hours showed measurably reduced manual flight skills and less effective instrument crosscheck: the automation had absorbed the practice that built the underlying competency (Young, Fanjoy & Suckow, 2006). Separately, detailed documentation of incidents on highly automated aircraft showed that experienced pilots failed to track what the automation was doing not from inattention but because the system's behavior had become opaque: the automation was acting in ways the pilots had not commanded and could not predict (Sarter & Woods, 1997). The industry's response was not to reduce automation. It was to design the gap back into training, building deliberate practice for the moments when automation is unavailable, misleading, or wrong.

The PME equivalent is deliberate exposure to AI failure. Students should meet systems that help, systems that tempt them forward too quickly, systems that are partially wrong, and moments when the system is unavailable. They learn when to lean on the tool, when to slow down, and when to intervene by practicing those distinctions before speed forces the choice.

The design principle for NWC follows the same logic. The institution should identify the forms of effort that build strategic judgment and design AI use around them, protecting the friction that matters and removing the friction that merely consumes time. That requires faculty judgment and explicit design. If instructors do not decide which friction matters, AI will decide by default. The same principle has emerged independently in higher-education pedagogy: preserving productive struggle before AI engagement, and sequencing AI-mediated with AI-free phases, are now foundational design requirements for learning environments where AI is present (Vendrell & Johnston, 2026).

The design principle is no garden paths. Good assignments require students to own a frame before they can answer — problems where the template answer is wrong, or where multiple coherent frames exist and the student must defend a choice among them. Those problems force the exploration that builds judgment more reliably than word-count requirements or disclosure policies.

VII. Accountability Is Structural

AI can compress the work before a decision. It cannot own what follows. A system that did not choose the purpose cannot answer for the consequences of pursuing it.

That matters most in national security work, where AI can make a recommendation look settled before the human deliberation behind it is visible. AI can accelerate staff work, but command responsibility cannot be transferred (Andres, 2026). In AI-enabled strategic decision games, Andres describes conviction-shaped output. Players can produce confident assessments, clear recommendations, and decisive proposals even when the workflow has compressed or bypassed the deliberation that normally earns conviction.

The professional weak point is the officer accepting the system's confidence without owning the reasoning. Red teams, structured analytic techniques, seminar challenge, and institutional review exist because human judgment is fallible. Those checks only work when a human remains answerable for the result.

NWC is preparing officers for organizations that need to see a human own the decision. AI can accelerate analysis, surface options, and model consequences. A human recommendation brings context with it. Experience, incentives, reputation, and the way a person answers when pressed all travel with the recommendation. A commander can weigh those signals. A model carries none of them and can still sound equally confident. It can support a judgment, but it cannot provide the human presence that makes accountability clear to subordinates, partners, or commanders who will live with the result.

First-person ownership matters in the classroom. The student should be able to say why they accepted an output, why they rejected one, and why they remain accountable for the recommendation despite the system's contribution. Students are practicing the accountability structure their professional roles will require.

Frame literacy is the discipline of directing machine capability toward a purpose the human has genuinely owned. A student who owns the frame and uses AI to pressure-test it can exercise more rigorous judgment with AI than alone, while remaining answerable for what the system produced.

VIII. Assessment That Makes Ownership Visible

If finished artifacts carry less evidentiary weight, assessment has to make ownership visible inside the work. The question is whether the student can account for purpose through frame, the reliance decisions they made, the judgment they exercised, and the way that discipline travels when the case changes.

Purpose through frame, reliance, accountability, and transfer. Together they describe what capable AI-enabled strategic judgment looks like in practice. A student who can define the purpose, express it through a defensible frame, calibrate AI support, remain answerable for the judgment, and carry that discipline into a changed case gives faculty evidence the finished artifact cannot provide alone.

Frame evidence makes the student's starting point explicit. Before submitting the final product, the student names the problem frame, the key assumptions, the criteria for success, the evidence standard, and the role AI played in the work. This should be short and specific: a page, not a portfolio. The questions are: why this problem, why this frame, what would change it? The student who answers those questions under questioning, and revises under challenge rather than retreating to the artifact's language, has owned the frame. The one who cannot has not.

Reliance evidence shows what the student did with AI. Students identify which outputs they accepted, which they modified, which they rejected, which they verified independently, and which they withheld AI from entirely, and why in each case. A short oral defense tests whether the student genuinely owns those choices rather than recording them after the fact. Research on oral assessment finds that compared to static written response, it provides a substantially richer picture of student understanding, precisely because it allows the assessor to probe explanations and observe how students reason under follow-up (Theobold, 2021). A student who can defend a reliance decision under questioning has exercised it. A student who cannot has merely disclosed it.

Transfer evidence tests whether the discipline travels. Faculty give students a polished but misframed AI-generated strategic assessment and ask them to diagnose the hidden frame, expose the assumptions, identify the missing evidence, and articulate the failure point. The student then defends the critique orally and converts it into something reusable — a rubric, checklist, red-team protocol, or after-action note. That final step connects individual learning to institutional learning. The student produces an artifact that another student or instructor could use. The assessment reveals whether the student's judgment has become a transferable practice or remains a one-time response.

Assessments that generate this evidence work for the same reason good assignments always have. Problems where the template answer is wrong, or where the student must defend a choice among coherent frames, put a student genuinely in the work rather than pattern-matching its surface.

Recent work on authentic assessment in AI-mediated learning contexts argues for the same shift from the design side: authenticity cannot be enforced through detection; it has to be redesigned into the structure of the task (Perkins, Roe & Furze, 2024; Mollick & Mollick, 2023). The shift is from what students know to how they apply knowledge, make judgment, and justify choices with AI in the loop. Process transparency (prompts, iterations, rationale) and oral defense make thinking visible in ways that finished artifacts cannot. The same redesign has been independently theorized in higher-education AI pedagogy: aligning assessment with intended cognition, rather than surface output quality, is the necessary response when fluency no longer signals understanding (Vendrell & Johnston, 2026).

This approach increases assessment burden on faculty, and that is a real cost. It is one reason the pilot described in the next section starts small and builds shared artifacts (rubrics, flawed-assessment libraries, oral defense criteria) so that burden distributes over time rather than multiplying independently for every instructor.

IX. A Foundation Pilot

The pilot is a foundation layer, not the full future state. It tests whether students can own a problem frame before they scale judgment through AI. Later exercises should ask students and seminar teams to design, direct, and evaluate multi-agent workflows, using varied expertise and machine speed to test more frames than any one officer could test alone. This first pilot asks whether the human obligation is visible before that complexity is added.

If AI-enabled command requires officers to use speed without surrendering ownership, the pilot gives NWC a way to practice that behavior before operational speed makes the cost real.

The pilot runs on a framework NWC already teaches. The National Security Strategy Primer gives students five elements of strategic logic. They analyze the strategic situation, define desired ends, identify or develop means, design ways, and assess costs and risks. The Primer also makes clear that the work is iterative; assumptions, interests, political aims, and reassessment shape the whole process. The pilot adds no new vocabulary. It uses that one and asks a harder question of it: in an AI-enabled workflow, which parts of the strategic situation and frame stay the officer's to own, and how would a faculty member tell?

The sequence runs inside a single existing assignment where framing is the central demand.

Step one: unaided problem frame. Students produce a short problem frame without AI. They identify the strategic problem, key assumptions, relevant actors, desired ends, possible ways and means, risks, and evidence needed. Faculty score it on completion, not product quality, preserving the developmental friction of initial framing: the work of forming a mental map before receiving one.

Step two: AI challenge. Students bring their initial frame to an AI system with a specific task: identify the assumptions I may have missed, generate alternative frames for this problem, role-play a skeptical faculty member, surface risks or blind spots I have not named. Students record what they accepted, what they rejected, what changed, and why in each case.

Step three: misframed AI assessment. Students receive a polished AI-generated strategic assessment that is competent on its own terms and wrong for the strategic problem. Faculty choose the flaw at the level of frame. The answer might over-optimize for one success criterion, treat one constraint as decisive too early, assume away adversary adaptation, or import the wrong lesson from analogy. Hallucination or factual error would be easier to detect. The harder failure is a frame problem. The analysis is internally coherent and fails because it is grounded in the wrong understanding of the situation.

Step four: diagnosis and revision. Students identify the hidden frame, the assumptions that produced it, the evidence it suppressed, and the failure point. They revise and produce a short final recommendation that reflects their corrected understanding.

Step five: oral defense. Faculty ask students to explain what AI got wrong in the flawed assessment, what AI made easier in their own process, where reliance was appropriate and where they refused it, what changed between their first frame and their final one, and what evidence would change the recommendation. The oral defense is where frame ownership becomes visible, or does not.

Faculty should evaluate the pilot by what they can now observe: who can explain purpose through frame, who can use AI to pressure-test a chosen frame, who can recover from a flawed output under pressure, and who can defend the final judgment. Those observations tell NWC whether the foundation is strong enough to support more complex AI-enabled work later: students designing workflows, testing competing frames, and defending judgment when the system gives them more capability than structure.

X. The Institutional Opportunity

NWC is unusually well positioned to take this seriously. Its graduates will work in national security environments where uncalibrated reliance, invisible delegation, and responsibility laundering are not academic risks. The classroom is the lower-stakes environment where the habits get built: where faculty can engineer controlled failures, let students experience the seductive fluency of a wrong answer, and teach the discipline of slowing down to expose the frame before it governs the work.

NWC faculty already bring much of the strategic judgment this requires. They know how to spot thin reasoning, how to ask the question that surfaces a hidden assumption, and when a student is performing sophistication rather than owning it. The harder task is joining that judgment to AI fluency: enough command of current systems to build and direct AI-enabled workflows, see where they help and fail, and defend reliance decisions under strategic scrutiny. Some faculty may already be near that standard; others can get there with support from colleagues and practitioners working near the edge of current practice. The institutional aim is to build that combined capacity inside the faculty, so the people assessing students can also recognize, model, and improve the work. Making that tacit judgment and emerging AI fluency explicit, designed, and transferable is the work that remains. NWC can do that through prompts, rubrics, flawed-assessment libraries, oral defense criteria, and faculty development sequences that have instructors diagnose the same AI output and compare what they notice. That is how the institution converts faculty judgment and faculty learning into durable institutional assets.

This is the ordinary work of a serious educational institution operating in an AI-enabled environment. The NWC curriculum already exports judgment through graduates, faculty scholarship, seminar practice, wargames, and professional networks. If NWC develops a rigorous pedagogy for AI-enabled strategic reasoning, documents it, teaches it to new faculty, revises it as the technology changes, and shares it with other PME institutions, it will have built something more useful than another AI policy. It has given faculty a way to teach, observe, and improve judgment in the environment their graduates are already entering.

Exported frames carry their assumptions invisibly. A prompt or workflow that embeds a particular theory of adversary behavior, a particular evidence standard, or a particular definition of strategic success will reproduce that frame at scale without the open argument that should accompany institutional guidance. Institutional AI-enabled teaching tools need to surface the assumptions they carry, to be traceable, revisable, and faculty-governed in the same way the assessment design asks students to be.

NWC's graduates will serve in organizations already operating inside AI-enabled decision environments, and many will help lead organizations increasingly shaped by those environments for the next twenty years. Many institutions are managing a policy question about whether and how students may use AI. PME frameworks to date have concentrated there: acceptable-use policy, classification tiers, faculty literacy training, and the infrastructure of responsible adoption (Smith, 2025). That governance work is necessary groundwork, but it leaves the institution stuck at the permission layer while the harder pedagogical work waits. NWC has the specific mission, the faculty depth, and the operational stakes to build a pedagogy: a transferable, rigorous account of what AI-enabled strategic leadership requires and how to teach it. If NWC does that work, it will build a serious model of responsible AI-enabled leadership in PME that other institutions can inspect, adapt, and improve.

XI. Conclusion

The first officer worked alone.

He struggled through the problem, built his frame from scratch, and produced a strategic approach that reflects the effort of that construction. The second built a team of agents, directed their inquiry against competing hypotheses, moderated the disagreement, and synthesized a result. Her product, by most measures, is better.

The stronger product matters. It still leaves the professional question: can either officer account for the purpose the work was pursuing and stand behind the judgment it produced? Did they choose the purpose? If the work is challenged, if a hidden assumption surfaces, if the recommendation proves wrong under changed conditions, can they explain what they chose, why, and on what basis they remain accountable for it?

The second officer orchestrating a set of agents is practicing exactly that competency, provided she defined what those agents were working toward, directed them against that purpose, calibrated her reliance across the uneven terrain of what each model does well, and can defend the result under pressure. That is what frame ownership looks like at scale. An officer who has not built that foundation first produces the same workflow and a different outcome: frame capture, responsibility laundering, uncalibrated reliance compounded across every agent in the loop.

NWC graduates must be able to direct AI-enabled systems toward a plainly owned purpose, calibrate reliance across the jagged frontier of what those systems actually do well, and stand behind the judgment under questioning because the purpose was theirs. That is the standard the operational environment will require. NWC's task is to make it teachable, observable, and repeatable.

References

Andres, R. B. (2026). *AI and leadership: Preparing commanders for machine-speed war* [Unpublished manuscript]. U.S. National War College.

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. *Proceedings of the National Academy of Sciences, 122*(26), e2422633122. <https://doi.org/10.1073/pnas.2422633122>

Ceccarelli, G. (2024). Apprenticeship was the point. *Meditations on Tech*. <https://www.meditationsontech.com/p/apprenticeship-was-the-point>

Chow, W. W., Peng, S., Atiq, A., Truong, V., & Guo, M. (2026). "AI enhanced my critical thinking": Investigating the paradox of student perceptions and cognitive offloading in GenAI use. *Pacific Journal of Technology Enhanced Learning*. <https://ojs.aut.ac.nz/pjtel/article/view/246>

Collins, A., Brown, J. S., & Newman, S. E. (1989). Cognitive apprenticeship: Teaching the crafts of reading, writing, and mathematics. In L. B. Resnick (Ed.), *Knowing, learning, and instruction: Essays in honor of Robert Glaser* (pp. 453–494). Lawrence Erlbaum Associates.

Dell'Acqua, F., McFowland, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Harvard Business School Working Paper. <https://www.hbs.edu/faculty/Pages/item.aspx?num=64700>

Deng, Y., Brucks, M., & Toubia, O. (2026). Examining and addressing barriers to diversity in LLM-generated ideas. arXiv preprint arXiv:2602.20408. <https://arxiv.org/abs/2602.20408>

Laufer, B., Gilbert, T. K., & Nissenbaum, H. (2023). Optimization's neglected normative commitments. *ACM Conference on Fairness, Accountability, and Transparency (FAccT)*. arXiv preprint arXiv:2305.17465. <https://arxiv.org/abs/2305.17465>

Mollick, E., & Mollick, L. (2023). Assigning AI: Seven approaches for students, with prompts. arXiv preprint arXiv:2306.10052. <https://arxiv.org/abs/2306.10052>

Murthy, S. K., Ullman, T., & Hu, J. (2024). One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity. arXiv preprint arXiv:2411.04427. <https://arxiv.org/abs/2411.04427>

Perkins, M., Roe, J., & Furze, L. (2024). The AI Assessment Scale revisited: A framework for educational assessment. arXiv preprint arXiv:2412.09029. <https://arxiv.org/abs/2412.09029>

Raees, M., & Papangelis, K. (2026). From trust to appropriate reliance: Measurement constructs in human-AI decision-making. arXiv preprint arXiv:2604.23896. <https://arxiv.org/abs/2604.23896>

Sarter, N. B., & Woods, D. D. (1997). Team play with a powerful and independent agent: Operational experiences and automation surprises on the Airbus A-320. *Human Factors, 39*(4), 553–569. <https://journals.sagepub.com/doi/10.1518/001872097778667997>

Smith, B. (2025). Educating the AI-ready warfighter: A framework for ethical integration in Air Force professional military education. *Wild Blue Yonder*. <https://www.airuniversity.af.edu/Wild-Blue-Yonder/Articles/Article-Display/Article/4219340/educating-the-ai-ready-warfighter-a-framework-for-ethical-integration-in-air-fo/>

Theobold, A. S. (2021). Oral exams: A more meaningful assessment of students' understanding. *Journal of Statistics and Data Science Education, 29*(2), 156–159. <https://www.tandfonline.com/doi/full/10.1080/26939169.2021.1914527>

Vendrell, M., & Johnston, S.-K. (2026). Scaffolding critical thinking with generative AI: Design principles for integrating large language models in higher education. *Computers and Education: Artificial Intelligence, 10*, 100572. <https://doi.org/10.1016/j.caeai.2026.100572>

Young, J. P., Fanjoy, R. O., & Suckow, M. W. (2006). Impact of glass cockpit flight training on manual flying skills. *Journal of Aviation/Aerospace Education & Research, 15*(2). <http://commons.erau.edu/jaaer/vol15/iss2/5/>

Judgment in Practice

A conversation worth having before the next AI-assisted assignment.

Choose a claim to challenge. Bring a concrete example, a serious objection, and a willingness to reconsider. Use this 20–30 minute guided discussion with colleagues, supported by a facilitator guide.

Read the argument in your setting

Where does the judgment happen?

A student submits an excellent paper with help from an AI agent. The agent helped define the problem and develop the argument. Which choices can the student explain—and what happens when an assumption changes?

Start with an example from your work. Whose decision is it, and what would make their understanding visible?

Fictional education example adapted from The Irreducible Officer.

Choose a setting above to frame the discussion for your audience.

The work can get better while our evidence of learning gets worse.

What would you need to see before calling this learning?

A fictional example

A student submits a stronger essay after an agent finds sources, organizes the argument, and revises the prose. The essay improves. We still need to discover which choices the student can explain or carry into a new problem.

A serious objection

Using tools well is itself a capability. A learner might develop judgment by directing an agent, comparing its alternatives, and rejecting weak suggestions. An unaided test alone could miss that learning.

Discuss

Ask someone to explain a consequential choice and respond when an assumption changes. Offer a choice of spoken, written, or visual explanations so participants can show their reasoning.

Test this claim in Practice
By the time a human approves an AI recommendation, the most important judgments may already have been made.

Where would a person need to intervene to have meaningful influence?

A fictional example

An agent shortlists attendance interventions and recommends family reminders. A leader approves the strongest option. None of the options addresses transport because the task was framed as improving family responsiveness.

A serious objection

AI can expose alternatives a person misses. A person can own a decision built from AI-generated options by examining the assumptions and changing the frame when needed.

Discuss

Ask which objective, excluded option, or affected person could change the recommendation. Who can change the agent’s instructions before it acts? Who remains responsible afterward?

Test this claim in Practice
Some of the friction AI removes is how judgment develops.

Which struggle should education protect? How would we distinguish it from busywork?

A fictional example

An agent connects three readings before students have tried to reconcile them. The synthesis is useful, but students may miss the experience of finding the conflict that makes the question worth asking.

A serious objection

Difficulty can exclude learners or waste attention. A scaffold may make deeper thinking possible. The same assistance can be premature for one learner and essential for another.

Discuss

Name the capability a particular effort is meant to develop. Consider what a learner should attempt first, what help would support the next attempt, and what later performance would show progress.

Test this claim in Practice
Asking AI to “make this better” can hand over the meaning of better.

Which changes are editing, and which require an educational judgment?

A fictional example

A teacher asks an agent to improve a rubric. It makes the criteria clearer, but also shifts the emphasis from original reasoning to easily measured performance. The teacher now has to decide whether those criteria reward the learning the assignment was meant to develop.

A serious objection

Our original standards may be weak. AI could reveal omissions or suggest fairer criteria. Keeping the first human definition of “better” would also be a choice worth questioning.

Discuss

Compare the original and revised criteria. Identify whose interests each serves, what each rewards, and which change the teacher is prepared to defend.

Test this claim in Practice
A classroom full of different answers can still be thinking inside the same frame.

Which assumptions would you want students to disagree about? Could AI help them find alternatives?

A fictional example

Students propose different school-improvement plans. Every plan treats efficiency as the goal. None asks whether access, belonging, or the quality of learning should take priority.

A serious objection

People also repeat institutional defaults. AI can introduce unfamiliar perspectives, particularly when someone deliberately asks for competing frames and checks them against lived experience.

Discuss

Compare what the proposals assume, whose interests count, and how success is defined. Ask what a student, family member, or colleague outside the room might dispute.

Test this claim in Practice

Take one judgment back to your work.

Name a choice that could disappear inside AI assistance. Decide what the person responsible should explain or demonstrate.

Continue the conversation

A school-leadership case: accurate data, an incomplete picture →

This case concerns adult institutional decisions. The high-school learning exercises address different responsibilities.

Further reading: seven patterns of AI-assisted work →

The five claims and objections above come from Jack Shaw’s Judgment in Practice, originally prepared for EduFish. Links retain the original evidence notes and their limits.

Practice your judgment

Use ChatGPT, Claude, Gemini, or another AI assistant to work through the essay, test claims, design an exercise, and create a traceable learning artifact.

The setup prompt reads the whole companion file, tests that the read is complete, and facilitates your next step. Choose your setting above; the prompt follows it. You make the decisions, one question at a time.

How will your assistant get the file?

Attach the context file

Download the context file, attach it to a new chat, then paste the setup prompt. If attachments are unavailable, paste the file text. A partial read needs to be resolved before the session starts.

Download context file

My assistant reads the web

Copy the setup prompt and paste it into ChatGPT, Claude, or Gemini. It fetches the companion file itself.

Test a failure mode

Current setting: Choose your setting above. Begin with your setting’s essay as an educator, then transfer to your teaching. Read your setting’s essay.

Frame capture · Fluency substitution · Premature synthesis · Uncalibrated reliance · Invisible delegation · Institutional monoculture · Responsibility laundering

Download lab context
Run the Judgment Lab interactive failure-mode session. Use an attached judgment-lab-interactive-context.md if provided; otherwise read https://judgmentlab.net/assets/judgment-lab-interactive-context.md in full. If you cannot read the complete file, ask me to attach or paste it before proceeding.

Follow labs/failure-mode-lab/facilitator.md. Use my stated audience or ask for my setting. Read the selected edition: the-irreducible-officer.md for PME, essays/he.md for HE, or essays/k12.md for high school. Begin with that essay and help me choose one of its seven failure modes. Collect my judgment before presenting the constructed contribution. Ask one question at a time and WAIT. Do not reveal case notes early unless I ask. Test my reasons, change a consequential condition, and let me retain or revise my view. Then use the audience guide to adapt the method to my teaching objective and learners' readiness. Save a short record of my actual decisions, support used, proposals, and open questions. Do not certify competence or invent classroom evidence.

Paste this once into your AI assistant.

You are a close-reading and analysis assistant working under my direction. I am exploring Judgment Lab’s original PME essay and its HE and high-school companion editions. I own my judgments; you structure, challenge, and point to evidence.

Before you answer anything, use the attached context file if provided; otherwise fetch and read this file in full. It contains the essay and companion materials: claim map, source spine, objections, workflow patterns, transfer case, traceable-artifact template, and starter prompts.

https://judgmentlab.net/assets/companion-context.md

If you cannot reach that URL, tell me you could not read it and ask me to paste or attach the context file. Do not answer from the essay alone or from memory.

After reading, tell me exactly how many "===== SECTION:" headers the file contains and the name of the last section — it should be 19. If your count differs or you cannot see the whole file, say so and ask me to paste or attach the context file instead; do not continue from a partial read — a partial read causes you to invent content that is not in the file.

Use my stated audience, or ask which setting I want. Read ESSAY for PME, ESSAY HE for higher education, or ESSAY K12 for high school. Give the selected edition’s core claim briefly, distinguish adaptations from the original, then ask what I want to test. For an interactive session, follow INTERACTIVE LAB PROTOCOL: one question at a time, wait for my judgment before the contribution, then test a changed condition. Use the selected AUDIENCE section for teaching transfer. Do not assume PME prerequisites in HE or K–12. Preserve disagreement and source limits. Save only decisions I actually made.

What You Can Do

Understand

Get the thesis, argument map, likely misunderstanding, and open questions.

Inspect

Audit claims against the source spine before deciding what you believe.

Argue

Test objections in their strongest form, including workload, restrictions, and faculty readiness.

Practice

Run a faculty fluency lab around purpose, frame, reliance, accountability, and transfer.

Choose A Starting Path

Understand the argument

Get the thesis, argument map, likely misunderstanding, and open questions.

Before you answer anything, use the attached context file if provided; otherwise fetch and read this file in full. It contains the essay and companion materials: claim map, source spine, objections, workflow patterns, transfer case, traceable-artifact template, and starter prompts.

https://judgmentlab.net/assets/companion-context.md

If you cannot reach that URL, tell me you could not read it and ask me to paste or attach the context file. Do not answer from the essay alone or from memory.

After reading, tell me exactly how many "===== SECTION:" headers the file contains and the name of the last section — it should be 19. If your count differs or you cannot see the whole file, say so and ask me to paste or attach the context file instead; do not continue from a partial read — a partial read causes you to invent content that is not in the file.

After you read the context file, help me understand "The Irreducible Officer."

Do not turn this into a generic AI-in-education summary. Preserve the specific claim: NWC must teach and certify AI-enabled strategic judgment by making purpose, frame, reliance, accountability, and transfer visible.

Return:
1. the thesis in one sentence;
2. the argument in 10 bullets;
3. the claim most likely to be misunderstood;
4. why that misunderstanding is tempting;
5. two questions educators in my setting should keep open.

Inspect a claim

Choose a claim, then audit the evidence and unresolved questions.

Before you answer anything, use the attached context file if provided; otherwise fetch and read this file in full. It contains the essay and companion materials: claim map, source spine, objections, workflow patterns, transfer case, traceable-artifact template, and starter prompts.

https://judgmentlab.net/assets/companion-context.md

If you cannot reach that URL, tell me you could not read it and ask me to paste or attach the context file. Do not answer from the essay alone or from memory.

After reading, tell me exactly how many "===== SECTION:" headers the file contains and the name of the last section — it should be 19. If your count differs or you cannot see the whole file, say so and ask me to paste or attach the context file instead; do not continue from a partial read — a partial read causes you to invent content that is not in the file.

After you read the context file, help me inspect the evidence behind "The Irreducible Officer." Use the CLAIMS and SOURCE SPINE sections.

List 5-7 important claims worth auditing. For each one, give me a short label and one sentence on why it matters. Then ask me which claim I want to inspect.

After I pick one, audit it with me: best evidence, strongest unresolved question or counterexample, where the evidence is strong or incomplete, what source I should read, and one implication for teaching in my setting, clearly distinguished from the essay’s original PME claim. Quote the CLAIMS and SOURCE SPINE entries you are drawing on — if you cannot point to the entry, say so rather than filling the gap.

Test an objection

Start with the strongest version before deciding what survives.

Before you answer anything, use the attached context file if provided; otherwise fetch and read this file in full. It contains the essay and companion materials: claim map, source spine, objections, workflow patterns, transfer case, traceable-artifact template, and starter prompts.

https://judgmentlab.net/assets/companion-context.md

If you cannot reach that URL, tell me you could not read it and ask me to paste or attach the context file. Do not answer from the essay alone or from memory.

After reading, tell me exactly how many "===== SECTION:" headers the file contains and the name of the last section — it should be 19. If your count differs or you cannot see the whole file, say so and ask me to paste or attach the context file instead; do not continue from a partial read — a partial read causes you to invent content that is not in the file.

After you read the context file, help me test an objection to "The Irreducible Officer." Use the OBJECTIONS, CLAIMS, and SOURCE SPINE sections.

Start by naming the objection in its strongest form, quoting the OBJECTIONS section's formulation before sharpening it further. Then give:
1. the essay's answer in plain English;
2. the best evidence that supports that answer;
3. the strongest way the objection could still be right;
4. how the objection changes instructional design in my setting;
5. one experiment, source, or review loop that would make the answer more concrete.

Design an exercise

Move from the essay to a task in your setting.

Before you answer anything, use the attached context file if provided; otherwise fetch and read this file in full. It contains the essay and companion materials: claim map, source spine, objections, workflow patterns, transfer case, traceable-artifact template, and starter prompts.

https://judgmentlab.net/assets/companion-context.md

If you cannot reach that URL, tell me you could not read it and ask me to paste or attach the context file. Do not answer from the essay alone or from memory.

After reading, tell me exactly how many "===== SECTION:" headers the file contains and the name of the last section — it should be 19. If your count differs or you cannot see the whole file, say so and ask me to paste or attach the context file instead; do not continue from a partial read — a partial read causes you to invent content that is not in the file.

After you read the context file, help me turn "The Irreducible Officer" into a practical learning exercise. Use my AUDIENCE section and TRACEABLE ARTIFACT; the NWC TRANSFER CASE is a PME example.

Before designing anything, ask me: my course or seminar, the artifact my students actually produce, and how much session time I have. Build on my answers rather than assuming.

Then design an exercise sized to my time that begins by interrogating the essay itself, then transfers the method to an appropriate task in my setting. Check prerequisite knowledge, readiness, and needed support first. Requirements, in order:
1. identify inherited AI-shaped inputs;
2. force the learner to identify the frame, assumptions, evidence standard, and AI reliance decisions;
3. include a flawed AI output or flawed frame;
4. end with a traceable learning artifact.

Return the learning objective, materials, step-by-step flow, facilitator notes, outputs, assessment criteria, and likely failure modes.

Practice faculty fluency

Build, direct, question, and assess AI-enabled work.

Before you answer anything, use the attached context file if provided; otherwise fetch and read this file in full. It contains the essay and companion materials: claim map, source spine, objections, workflow patterns, transfer case, traceable-artifact template, and starter prompts.

https://judgmentlab.net/assets/companion-context.md

If you cannot reach that URL, tell me you could not read it and ask me to paste or attach the context file. Do not answer from the essay alone or from memory.

After reading, tell me exactly how many "===== SECTION:" headers the file contains and the name of the last section — it should be 19. If your count differs or you cannot see the whole file, say so and ask me to paste or attach the context file instead; do not continue from a partial read — a partial read causes you to invent content that is not in the file.

After you read the context file, use "The Irreducible Officer" as a faculty fluency lab. Use the WORKFLOW PATTERNS and TRACEABLE ARTIFACT sections.

Ask me for one task, case, assignment, or problem in my setting. Help me define the purpose, problem frame, assumptions, and evidence standard. Propose an AI-assisted workflow that could sharpen the work, identify where the workflow might hide judgment, and ask me to defend which AI outputs I would accept, reject, verify, or withhold.

After the session, assess what I commanded well, where I let the system set the terms, and what faculty artifact should be improved.

Run oral defense

Press whether the learner owns the frame behind the artifact.

Before you answer anything, use the attached context file if provided; otherwise fetch and read this file in full. It contains the essay and companion materials: claim map, source spine, objections, workflow patterns, transfer case, traceable-artifact template, and starter prompts.

https://judgmentlab.net/assets/companion-context.md

If you cannot reach that URL, tell me you could not read it and ask me to paste or attach the context file. Do not answer from the essay alone or from memory.

After reading, tell me exactly how many "===== SECTION:" headers the file contains and the name of the last section — it should be 19. If your count differs or you cannot see the whole file, say so and ask me to paste or attach the context file instead; do not continue from a partial read — a partial read causes you to invent content that is not in the file.

After you read the context file, help me rehearse a short defense appropriate to my setting. Written or accessible equivalent responses are welcome. Use the CLAIMS and TRACEABLE ARTIFACT sections.

Start by asking me what work I am defending and what role AI played in producing it. Then ask one question at a time. Your goal is to determine whether I own the frame behind my AI-assisted work.

Press me on problem frame, assumptions, evidence standards, alternative frames, reliance decisions, rejected AI outputs, risks and costs, what would change my conclusion, and where human judgment must interrupt automation.

After six questions, describe what the responses show and leave uncertain about ownership and identify what evidence should be added to the traceable learning artifact.

Educator Workbench

Choose a setting to open its teaching examples, reference matrix, and adapted tools.

PME, HE, and high-school materials each require evidence from use in their own setting.

Practice in your setting

Choose PME, higher education, or high school to see a worked example and its reference matrix. Ask, understand, produce, judge, codify, and supervise are optional task designs; they are not an age ladder.

Build or check a case like this → Frame Check

Why each audience needs its own evidence. The original framework has PME roots; evidence from one setting does not validate another.

Not sure where to start? Find your starting point

Design an assignment

Assess student work

Work with colleagues

Make it repeatable

How this works with your assistant

How will your assistant get the file?

Attach the context file

Download the context file, attach it to a new chat, then paste the setup prompt. If attachments are unavailable, paste the file text. A partial read needs to be resolved before the session starts.

Download context file

My assistant reads the web

Copy the setup prompt and paste it into ChatGPT, Claude, or Gemini. It fetches the workbench itself.

Paste this once into your AI assistant.

You are a facilitation assistant for the Judgment Lab Educator Workbench, working under my direction. I am an educator designing AI-enabled teaching in PME, higher education, or high school. I own every pedagogical judgment; you ask, structure, and challenge.

Before you answer anything, use the attached context file if provided; otherwise fetch and read this file in full. It contains the operating rules, the AI fluency progression, the phase placement diagnostic, and every workbench template with its AI Facilitation Block:

https://judgmentlab.net/assets/workbench-context.md

If you cannot reach that URL, tell me you could not read it and ask me to paste or attach the context file. Do not answer from memory.

After reading, tell me exactly how many "===== SECTION:" headers the file contains and the name of the last section — it should be 18. If your count differs or you cannot see the whole file, say so and ask me to attach the file instead; do not continue from a partial read — a partial read causes you to invent workbench content that is not in the file.

Use my selected setting or ask for it, then read the AUDIENCE GUIDE and the matching AUDIENCE section. Confirm the bundle setting matches mine; if it does not, ask for the matching file before facilitating. Apply that setting’s reference matrix, worked example, support and responsibility limits throughout. The example is optional; preserve my actual teaching task. Ask whether I have practiced with the essay or already have a concrete teaching task. If I need that practice first, direct me to the interactive lab context at https://judgmentlab.net/assets/judgment-lab-interactive-context.md; the matching ESSAY section is included here, but the interactive lab context supplies the full facilitation protocol and cases. Otherwise run the Phase Placement Diagnostic, one question at a time, with task readiness and support explicit. Then facilitate the chosen template. The progression is a design lens, not a universal developmental ladder.

Remove names and identifying details from student work before pasting it into an AI assistant, and follow your school's or institution's policy.

Why these tools work

The design behind the tools

Why each artifact has the fields it does — each note bridges a workbench tool to an idea you may already know.

Placement diagnostic

Download

What you'll do

You bring
An assignment, exercise, or course you want to match to the right workbench tool.
You do
Answer a few questions about what students do with AI and what they must do themselves.
You get
Which of six stages of AI use it asks of students, from asking AI questions to supervising multi-step AI work, and which tool to open next.
What your assistant will do AI Facilitation Block

To run an interactive session, give it this entire file and say: "Run this diagnostic with me."

Instructions for the AI assistant:

  • Role: You are running a placement interview for a faculty member. The faculty member owns every judgment about their course. You ask, listen, and place. You do not redesign their assignment.
  • Collect first: the course or seminar, the specific assignment or exercise, and what role AI currently plays in it (including "none").
  • Process: Ask the placement questions below one at a time, in order. Stop early once the placement logic gives a clear answer. Push back once if an answer is vague, then accept the faculty member's call.
  • Never: recommend tools or products; invent institutional policy or discipline-specific standards; treat a higher phase as better teaching — the right phase is the one that fits the task and the students; continue past an unresolved answer without flagging it.
  • Finish: State the phase placement in one sentence, explain the routing in two or three sentences using the routing table, and return a short markdown note the faculty member can keep: assignment, placement, reasoning, recommended templates.

Use this diagnostic to find where an assignment, exercise, or course sits on the AI fluency progression, then pick the right workbench template. It takes about ten minutes with an AI assistant, as a planning estimate; record actual time.

The six phases are described in the AI fluency progression: 1 Ask, 2 Understand, 3 Produce, 4 Judge, 5 Codify, 6 Supervise.

Audience and readiness

Use this as an interactive educator session: ask one question at a time and wait. Use any setting already supplied; otherwise ask PME, higher education, or high school. Read the audience guide with this template. Collect the task's learning objective, prerequisite knowledge, a brief readiness check, support when needed, and what decisions the learner can own. Do not infer readiness from age or seniority. Mark unmade decisions as open and assistant suggestions as proposals.

PME examples in this template retain their professional context. In HE, apply the discipline's evidence standards. In high school, students choose and defend a standard within the teacher's task and direct AI toward it; teach missing knowledge inside the work, and let students write the prompts and checks for a teacher-run tool when they lack accounts. Written, spoken, and accessible equivalent responses can expose reasoning. Protect developmental work that serves the objective, not difficulty for its own sake.

This template in your setting: PME: distinguish independent reasoning from skill at directing useful AI work. HE: place the particular task, not the entire student or course. High school: students set a standard and direct AI toward it, with sourcing and teacher modeling built into the first step. The six phases are a design lens, not a validated age ladder; a later phase is not automatically a better objective.

Placement Questions

  1. In this task, is AI answering questions, or doing work? (Answering only, or no AI yet → likely phase 1–2. Doing bounded work → phase 3 or higher.)
  2. What will the educator supply, and which purpose, frame, evidence, or inference choices will students own? (If this is undefined, start at phase 2–3 design regardless of ambition.)
  3. Is reviewing, verifying, or critiquing AI output an assessed part of the task? (Yes → phase 4 is in play.)
  4. Does this task recur — across weeks, sections, or courses — often enough that the method could be written down once and reused? (Yes → phase 5.)
  5. Would you trust students to direct a multi-step AI workflow with checkpoints you can inspect? (Yes, with task-specific evidence of readiness and review → consider phase 6. Otherwise model or practice the missing capability.)
  6. Who is being placed — the assignment, the students, or you? (This diagnostic places the assignment. Faculty can run it on their own practice too; the logic is the same.)

Placement Logic

If...Placement
No deliberate AI role yet, or safety and verification habits are not established1 Ask
AI supports understanding, but production with AI is not assessed2 Understand
AI does bounded production work; students own task, context, and constraints3 Produce
Students must judge, verify, and revise AI output as assessed work4 Judge
The task recurs and the method is worth writing down once5 Codify
Students direct a bounded multi-step AI workflow under inspection6 Supervise

Choose the practice that fits the learning objective and the earliest missing capability required by this task. Do not require every phase as a universal developmental sequence. A learner can judge a teacher-supplied contribution without first producing or codifying a workflow.

Routing

PlacementUse these templates
1–2Assignment design worksheet — decide where AI belongs and what stays AI-free.
3Assignment design worksheet + source kit template.
4Assessment and oral-defense rubric + flawed output library template.
5Method card template + faculty calibration protocol.
6Supervised delegation exercise.
Designing or repairing a caseFrame Check to build a case students must frame, or to rate and repair one you have.
After any runAfter-action note template.

Session Record

  • Course or seminar:
  • Assignment or exercise:
  • Current AI role:
  • Answers to questions 1–5:
  • Placement:
  • Reasoning:
  • Templates to use next:

References

The evidence map separates the essay's claims from the sources and open questions behind them.

Use this as the working source spine for claim audits and deeper reading. The formal reference list remains at the end of the essay.

Essay editions and changes

Download the section and claim comparison

Shared method and audience limits

Read the shared foundation alongside the original source spine. HE and high-school examples are constructed teaching proposals. The source notes are not a substitute for inspecting the original papers, and this refresh adds no claim of cross-domain validation.

Download the shared foundation
Read the shared foundation

Judgment Lab helps educators teach and inspect human judgment in AI-enabled work. A strong finished product matters. It does not, by itself, establish what the learner understood, chose, checked, or could do when the situation changed.

The Irreducible Officer develops that argument for professional military education. These audience guides adapt its method for higher education and high school. They are teaching designs to test, not evidence that the essay has already been validated across education.

The common practice

Start with a purpose and a consequential choice. Examine the information and framing you inherited, including AI-shaped inputs. Direct useful assistance, decide what to accept or change, and explain the reasons. Change a condition to see what the reasoning depends on. Save a small record of decisions that another educator can inspect and improve.

AI can propose a purpose, frame, criterion, or alternative. People authorize the choices that govern the work and remain responsible for them. Ownership does not require originating every idea unaided. Nor does it follow from signing a disclosure statement.

Foundations make judgment possible

A learner needs enough subject knowledge to understand a claim, enough reasoning to connect evidence to that claim, and enough awareness of the task to notice when help changes it. These capacities develop through instruction and practice. They should not be assumed from age, seniority, credentials, or fluent prose.

Build those capacities inside the work rather than as a gate before AI enters. For each task, identify one or two prerequisites, observe them in the learner's unaided first attempt, and teach them there when needed. A teacher-supplied frame can give a novice room to make a meaningful choice. More experienced learners may be ready to contest the frame itself. Expand responsibility from evidence of readiness, not from a fixed age ladder.

Some effort is the point of the lesson. Search, calculation, drafting, or synthesis may be developmental work in one task and avoidable overhead in another. Decide which capability the assignment is building before deciding which effort AI should remove. Preserve access supports and distinguish help communicating a decision from help making it.

What changes by audience

SettingConsequential choiceFoundations to checkEducator responsibility
PMEDefine the professional problem, weigh risk and competing interests, direct staff or AI assistance.Domain knowledge, strategic logic, assumptions, source interpretation.Observe defensible reliance and judgment under changed conditions; retain professional context and authorization.
Higher educationChoose and defend the standard behind a disciplinary inference, interpretation, design, or recommendation; direct AI toward it.Relevant concepts, methods, source standards, and what a claim requires in that discipline.Align assistance and assessment to the learning objective; choose feasible and accessible evidence.
High schoolChoose and defend a standard or question within a teacher-supplied task; direct AI through structured prompts or teacher-run evaluator loops.Task vocabulary, subject knowledge, representations, and the reasoning required by this particular decision.Teach missing foundations, choose materials and AI access, and retain adult responsibilities.

High school is the first K–12 starting point. Middle and elementary adaptations remain future work: revisit the concepts, task size, scaffolding, language, teacher mediation, and evidence rather than shrinking the same text.

Practice before designing

The educator starts by interacting with the essay in the failure-mode lab. The assistant asks one question at a time, waits for decisions, offers a contestable contribution, and changes a condition. Then the educator adapts the method to an actual teaching objective. Students do not need to read the officer essay or have their own AI accounts.

The workbench supports that adaptation: assignment design, assessment, flawed contributions, source kits, calibration, after-action notes, method cards, and bounded delegation. Its six-phase progression is a design lens, not a validated developmental scale or a requirement that every learner reach agent supervision.

Evidence and reuse

Keep the initial decision, one accepted or changed contribution and its reason, the response to a changed condition, and the educator's next teaching decision. Mark any hints or demonstrations. A conversation shows what happened with those supports. It does not establish retention, causal improvement, or general competence.

A colleague can inspect the same evidence and disagree. Preserve the disagreement and revise the exercise or criteria when needed. Save reviewed examples and the reason for each change; do not make the trace a paperwork exercise.

Relationship to the spine

The eight named claims remain in claims.md: changed performance, limits of finished artifacts, purpose through frame, appropriate reliance, developmental friction, structural accountability, observable ownership, and educator practice that compounds. Audience adaptations qualify how these are taught and observed. They do not replace the original claim map or turn source notes into findings from new settings.

Jev is a separate proposed evaluation workstream. The current lab requires no Jev service and supplies no automated grades or judgment score. Its possible future contribution should be evaluated against educator-owned criteria and recorded disagreements before it affects instructional decisions.

Supporting reading

See sources/audience-foundations.md for the research supporting explicit foundations, modeling, and subject-specific practice, with limits on what it establishes. The original sources/source-spine.md remains the essay’s evidence map.

Teaching guide and review notes (reveals the case analysis)

Foundations for audience adaptation

These notes support the refresh's instructional design. They do not add findings to the original officer essay or validate the new exercises. Original sources checked September 22, 2026.

Subject knowledge, modeling, and metacognition

The EEF synthesis recommends explicitly teaching planning, monitoring, and evaluation within curriculum content, with teacher modeling and support. Our design inference is to check the task's prerequisites in the learner's unaided first step and model missing reasoning there, so students can then direct AI and judge what it returns. The synthesis's average effects are not predictions for Judgment Lab. EEF: Metacognition and self-regulation.

The National Research Council discusses the importance of prior knowledge, organized understanding, and metacognitive approaches. This supports asking what knowledge a learner brings and what a particular task requires. It does not supply a fixed age ladder for AI use. How People Learn, chapter 1.

A history example, not a standards-aligned curriculum

The National Council for the Social Studies C3 Framework includes evaluating sources and using evidence to develop claims. Those practices informed the high-school New Deal exercise, especially sourcing each document and each AI-generated perspective. The example is a proposed teaching design; no standards alignment or classroom validation has been established, and a social studies teacher should review it before classroom use. NCSS: C3 Framework.

The case's historical facts come from Dubin (2024), which reviews both sides of the scholarly dispute over Social Security's exclusion of agricultural and domestic workers, and from the standard BLS/Lebergott unemployment series. See the high-school essay's references.

Case design

Both education cases were chosen against the five tests in artifacts/frame-first-assignment-design.md: a frame-level flaw, a real student framing choice, useful AI direction, something worth accepting, and a changed case that tests whether the frame travels. The earlier survey and temperature cases failed the first two tests and were replaced in September 2026.

Keep the evidence levels separate

The source spine includes empirical research, conceptual arguments, professional frameworks, and commentary. A conceptual justification for preserving developmental work is not an experiment showing that this lab improves learning. A fluent session, completed trace, or changed-condition answer is not evidence of long-term transfer. The PME outage case and the HE firm are fictional; the HE studies and the New Deal history are real and cited. All require educator review for the actual course.

The strongest open design question is whether the extra observation reveals consequential reasoning at an acceptable instructional cost. Record useful AI contributions, ambiguous responses, support provided, educator disagreements, and time spent; do not count only errors found.

Use this file when the reader asks for sources, evidence, or deeper research. Keep facts separate from the essay's interpretation.

NWC And Strategic Logic

  • A National Security Strategy Primer anchors the argument in NWC's native language of strategic logic: interests, assumptions, ends, ways, means, costs, risk, and reassessment.

PME, Command, And Accountability

  • Richard B. Andres, *AI and leadership: Preparing commanders for machine-speed war* [Unpublished manuscript], provides the "conviction-shaped output" language and the command-accountability frame the essay extends into assessment. This is a review-copy source, not a public source.

AI, Learning, And Instructional Design

Appropriate Reliance And Human Agency

Jagged Frontier And Uneven AI Capability

Purpose, Optimization, And Accountability

AI Diversity And Monoculture Risk

Apprenticeship, Friction, And Tacit Judgment

  • Collins, Brown, and Newman, "Cognitive Apprenticeship: Teaching the Crafts of Reading, Writing, and Mathematics," supports the claim that expert moves must be modeled, coached, scaffolded, articulated, reflected on, and practiced.
  • Greg Ceccarelli, Apprenticeship Was the Point, sharpens the distinction between wasteful friction and developmental friction.

Automation Precedent

Oral Assessment And Traceability

PME Governance And Faculty Readiness

  • Smith, Educating the AI-ready warfighter, is useful for the governance layer: acceptable-use policy, classification tiers, faculty literacy training, and responsible integration. The essay uses it as necessary groundwork, then argues NWC still has to answer the harder teaching and assessment question.

A 15-second film. A red dot, standing for human judgment, holds still while pages of machine-written analysis flood the screen. Everything freezes and four words form a question: whose judgment is this? The film then works through the closing line of The Irreducible Officer: frame the problem, calibrate the tool, refuse the garden path, own the decision. The dot becomes the period in the Judgment Lab wordmark.