# Judgment in Higher Education

## Directing AI Toward Purposes Students Own

Companion edition of *The Irreducible Officer*, for university instructors, faculty developers, and program leaders. September 2026. The original argues that the National War College has to teach and certify AI-enabled judgment. This edition makes the same argument for higher education. The course designs in it are proposals for instructors to test in their own disciplines.

## I. The Bar Went Up

Two students get the same assignment. A mid-sized software firm is deciding whether to require employees back in the office, and each student has to write the firm a research memo.

The first reads the studies alone. She works out what each one measured, notices that they disagree, and writes a careful memo recommending a hybrid schedule. The second directs AI through the same material. She has it sort the studies by what they measured and which workers they followed. She has one agent argue the firm's case and another argue the case of the firm's newest hires, then works through where they disagree. Her memo is sharper and better sourced. Read the two memos cold and you would rank hers higher.

Now ask which student understands the problem. The memos will not tell you by themselves. The second student may have done the more demanding work, directing machine capability toward a question she chose. Or the model may have handed her its framing of the question, and she organized it well. AI did not lower the bar for judgment. It raised it, then made it harder to see who cleared it.

The first student is the harder case. Her memo is careful, unaided, and well argued. She looks finished. She is also about to enter workplaces where her peers pair similar judgment with much stronger command of the tools. A course that certifies her and misses that gap has taught to the old standard.

Policy questions still matter. Institutions have to decide where AI is allowed, how students disclose it, and which material may enter which systems. A course can settle all of that and still assess the wrong thing. Disclosure tells an instructor that a student used AI. It does not tell the instructor whether the judgment in the memo is the student's.

Higher education's task is to graduate people who can direct AI toward purposes they own, decide when its output deserves weight, and defend the result as their own. Faculty already know how to probe understanding through discussion, drafts, and problems that break a familiar pattern. The work ahead is to point those practices at AI-enabled work instead of around it.

## II. A Fluent Synthesis of the Wrong Question

Here is what happens when a student asks the question the way most students would: "Summarize the research on whether remote work hurts productivity."

The model returns a good paragraph. It reports that call-center employees at the travel firm Ctrip who were randomly assigned to work from home performed 13 percent better ([Bloom et al., 2015](https://doi.org/10.1093/qje/qju032)). It reports that a later randomized trial of hybrid work with 1,612 employees at the same company, since renamed Trip.com, cut quit rates by a third with no effect on performance reviews ([Bloom, Han & Liang, 2024](https://www.nature.com/articles/s41586-024-07500-2)). It reports that junior software engineers received less feedback on their code when their teammates were not nearby ([Emanuel, Harrington & Pallais, 2023](https://www.nber.org/papers/w31880)). It concludes that the evidence is mixed and that hybrid arrangements offer the best of both.

Every summary is accurate. The memo built on it would read well. And for many firms it answers the wrong question.

The synthesis treats "productivity" as short-run output, averaged across workers. The firm's actual question depends on who its workers are. If it hires mostly new graduates, the proximity study is the one that matters: junior engineers lost mentoring when they worked apart, and the paper frames that as a trade between output today and skill later. The Ctrip study adds a detail the summary left out. Its home workers were volunteers, and they were promoted less often than office workers with the same performance. A firm whose future depends on developing junior staff should weigh those findings far more heavily than a 13 percent gain among experienced volunteers answering phones.

A student who catches this has not found an error. The summaries are right. She has noticed that the synthesis answered "does remote work change output?" when her firm needed "what happens to the people we are trying to develop?" That is a frame problem, and no stock phrase like "check for bias" will find it.

Five of the original essay's failure modes show up in this one exchange.

**Frame capture.** The model's definition of productivity becomes the memo's definition. Later drafts improve the prose inside the wrong boundary.

**Fluency substitution.** The balanced paragraph sounds like judgment. It weighs every study and never decides which one this firm should care about most.

**Premature synthesis.** The model connected three studies before the student knew enough about their designs to see that they measured different things in different workforces.

**Invisible delegation.** "Summarize the research" handed over the criterion. The student believed she was asking for help with reading. She was also asking the model to decide what counted.

**Institutional monoculture.** Give the same assignment to thirty students using similar tools and similar prompts, and many will turn in some version of "the evidence is mixed; hybrid is the balance." The memos will differ in structure and sources while sharing one framing of the question. A seminar that should surface competing frames ends up comparing phrasings of a single one ([Deng, Brucks & Toubia, 2026](https://arxiv.org/abs/2602.20408)).

## III. What the Finished Memo Cannot Show

The finished memo still matters. But if it now carries less of the evidence, faculty need to know how much less, and what has to carry the rest. A memo can show structure, balance, disciplinary vocabulary, and clean prose while leaving the student's actual contribution unclear.

Bastani and colleagues found that high-school mathematics students using an unrestricted GPT-4 tutor improved during practice and then scored 17 percent worse than peers without access once the tool was removed. A version with teacher-designed safeguards largely avoided the loss ([Bastani et al., 2025](https://doi.org/10.1073/pnas.2422633122)). The study is about one subject and one tool design. Its lesson for a university course is that assisted performance and independent performance are two different observations, and a course that needs both has to look at both.

That decision belongs in the learning objective. An instructor teaching research methods may accept AI help finding and formatting sources while requiring the student to explain what each study measured. An instructor teaching policy writing may care most about whether the student can direct AI through a literature quickly and defend what she kept. Both are legitimate. Neither can be read off the finished memo.

## IV. The Student Decides What Counts as Success

AI can propose goals, criteria, and definitions. It proposed one in the memo case, when it quietly defined productivity. What it cannot do is decide that its proposal should govern the work. Someone has to accept that definition or replace it, and that person answers for the choice.

In a course, faculty set the learning objectives and usually the topic. That is good teaching. The instructor in this case supplied the firm, the question, and a reading list. The consequential framing choice still belongs to the student. She decides what productivity should mean for this firm, which evidence bears on that meaning, and what a good recommendation has to accomplish. Two strong students could frame it differently, one around retention and one around development, and both could earn full marks if they defend the choice.

Faculty can widen that choice as students advance. A first-year student might choose between two definitions the instructor names. A senior might be handed a firm with no guidance and asked to decide which questions matter before she reads anything. The discipline stays the same: the student names the purpose before the model's structure arrives, because deciding what problem to solve is the judgment the assignment exists to build.

## V. Directing AI Well Is a Skill With Levels

The original essay describes a progression of where human judgment enters an AI workflow. It gives faculty a way to set a target for their course.

At the bottom is minimal prompting. "Summarize the research on remote work" inherits the model's frame almost entirely, and that is where most students start.

One level up, the student states her purpose and criteria: "My firm hires mostly new graduates. Sort these studies by what they measured, which workers they followed, and for how long, and flag any finding about training or promotion." She has moved her judgment upstream, and the output now works for her question.

Above that is the evaluator loop. The student has one agent argue for the firm's managers and another for its newest employees, or has the model critique her draft against the standard she set. The disagreement is the point. She learns where her frame is weak before a reader finds it.

At the top, advanced students build reusable workflows with defined review steps, or assign several agents distinct roles and moderate between them. That is the second student in the opening scene. It is a reasonable target for a capstone or graduate seminar. It is too much to ask of a first-year student who has not yet learned to read a study's methods section.

At every level the student also decides what to rely on. In the memo case she should accept the model's summary of each study's design after checking it against the abstracts, since that is quick to verify and the model was right. She should not accept its verdict that the evidence is "mixed," because that verdict depends on a frame she has rejected. Asking a second model whether the first is right does not settle anything. The check goes back to the studies themselves ([Raees & Papangelis, 2026](https://arxiv.org/abs/2604.23896)).

Faculty should model this in front of students. An instructor who walks through her own prompts for a literature review, shows where the model helped, and shows the definition she refused to accept teaches more than a policy statement does.

## VI. Protect the Effort That Builds Judgment

Some work looks inefficient because it is waste. Some looks inefficient because it is how judgment forms. AI removes both without telling the difference.

In the memo assignment, finding sources, formatting citations, and drafting routine sections are reasonable places for AI to help. Reading one study closely enough to know what it measured is not. That effort is what lets a student see that a call-center experiment and a software team's code reviews cannot be averaged. Skip it, and she has no way to judge the synthesis she is handed.

Novices often lack that knowledge, and the answer is to teach it inside the assignment. Before AI enters, ask the student to read one study and say what it measured, who the workers were, and what it cannot tell the firm. If she cannot, teach that there, then continue. Foundations belong inside the loop. Higher-education researchers reach the same design principle: preserve productive struggle before AI engagement, and sequence AI-free and AI-mediated phases on purpose ([Vendrell & Johnston, 2026](https://doi.org/10.1016/j.caeai.2026.100572)).

Students also need to meet a tool that helps, one that tempts them forward too fast, one that is partly wrong, and a changed case where yesterday's reasonable answer no longer fits. If instructors do not decide which effort matters, AI decides by default.

## VII. The Person Who Sets the Standard Answers for It

The student who defined productivity as short-run output answers for the memo that follows from it. "The research shows hybrid is the balance" does not say who decided what counted as productivity. If the firm adopts the recommendation and loses a cohort of junior staff, the question of who chose that standard should have an answer, and the answer is the student.

That responsibility is scaled to a student's role. She is not approving a firm's policy. She is answerable for the claims she submits and the standard behind them, and she should be able to say what she accepted, what she refused, and why she still stands behind the recommendation.

The same structure applies up the chain. An instructor who uses AI to draft feedback or propose grades answers for those assessments. A program that deploys an AI system answers for how it is configured. The person who sets the standard answers for the result at every role.

Group work hides this. A polished team memo can conceal who made the consequential framing choice. Ask each member to explain the standard the group chose and how they would apply it to a changed case.

## VIII. Assessing the Frame, Not Only the Memo

If finished work carries less evidence, assessment has to make ownership visible inside the work. That means three short pieces of evidence alongside the memo.

The first is frame evidence. Before submitting, the student writes a few sentences naming the question she answered, what she took productivity to mean, and what evidence would change her recommendation. A paragraph is enough.

The second is reliance evidence. She names one AI contribution she accepted, one she refused, and the check behind each. A short follow-up, spoken or written, tests whether she can defend those choices or only recorded them afterward. Oral questioning gives an assessor a much richer view of reasoning than a static written answer, because the assessor can follow up ([Theobold, 2021](https://www.tandfonline.com/doi/full/10.1080/26939169.2021.1914527)).

The third is a changed case. Hand her a different firm: an established call center whose staff average ten years of experience and are rarely promoted out of their roles. The Ctrip evidence is now the most relevant, and development evidence matters less. A student who owns her framing discipline asks again what productivity should mean here and reaches a different emphasis. A student who memorized "juniors need proximity" repeats it.

Grade the defense, not the conclusion. Credit the student who states her standard and why it fits, names an alternative frame she did not adopt, uses evidence that bears on her standard, and explains what would change her answer. Do not credit balance for its own sake, suspicion for its own sake, or changing one's mind as a goal.

This adds work for faculty. Start with one assignment and one changed case, and compare what the added evidence reveals with what the memo alone showed.

## IX. A Pilot in One Course

The original essay's five-step pilot carries over directly.

1. **Frame unaided.** Students write a short frame without AI: what the firm's question is, what productivity should mean for it, and what evidence they would need. Score it for completion. This is also where the instructor finds out who cannot yet read a study's design and teaches it.
2. **Direct AI against the frame.** Students take their frame to AI with a specific task: find assumptions I missed, argue the firm's side and the new hires' side, tell me which of these studies does not fit my standard. They record what they kept and why.
3. **Meet the misframed synthesis.** Students receive the fluent "evidence is mixed" synthesis from Section II.
4. **Diagnose and revise.** They name the hidden frame, what it suppressed, and where it fails for their firm, then revise their memo.
5. **Defend.** A short follow-up on their reliance decisions, then the changed call-center case.

Two instructors can review a handful of the records together and compare what they infer. Where they disagree, the prompt or the rubric usually needs work.

Other disciplines need their own misframed answer, and the pattern carries.

In a history course on colonial America, a student asks AI why the Salem witch trials happened. It returns an accurate list of what historians have argued. Boyer and Nissenbaum traced the accusations along a factional split in Salem Village, Karlsen found that many accused women had inherited, or stood to inherit, property in families without male heirs, and Norton tied the crisis to refugees and fear from the war on the Maine frontier. The model presents these as contributing factors and adds ergot poisoning, a 1976 hypothesis most historians have rejected. Each summary is right. The frame is wrong, because the historians were answering different questions: who accused whom, why certain women were accused, and why the crisis came in 1692 and spread. A student has to decide which question her paper answers before she can use any of them. The changed case asks why the trials ended, and the decisive evidence shifts to the dispute over spectral evidence and Increase Mather's *Cases of Conscience*.

In an engineering design course, a team asks AI to choose a material for a mounting bracket on a machine that will vibrate for years. It builds a weighted decision matrix across weight, cost, corrosion resistance, and machinability and recommends an aluminum alloy. The property data are right. The frame treats fatigue life as one preference among several, or leaves it out, when the bracket either survives the required number of load cycles or fails. Most steels have a stress level below which cyclic loading causes no fatigue damage; aluminum alloys do not. A team that screens the must-meet requirements before weighing tradeoffs will reach a different answer. The changed case is a bracket for a test fixture used a few hundred times, where the aluminum recommendation is right.

Both cases need review by an instructor who teaches the course before they go to students.

## X. What Departments Can Build

Faculty already bring the judgment this work needs. They know when an inference outruns its evidence and when a student is performing sophistication rather than owning it. What most need in addition is enough command of current AI tools to direct them at the level they ask of students, see where they help and fail, and model those decisions in class. Some faculty are already there. Others can get there by working through the same assignments their students will do.

A department can build that capacity together. Have instructors run the same misframed synthesis, compare the questions they would ask, and agree on what a strong defense looks like. Keep a small library of misframed answers for the discipline, each with its frame flaw named, an AI contribution worth accepting, and a changed case. Include examples where the model was right, so the library does not teach reflexive distrust.

Shared prompts and rubrics carry assumptions into every course that reuses them. A prompt that defines a good literature review, or a rubric that rewards confident prose, will reproduce its frame at scale. Those assumptions should be written down and open to faculty revision, the same way the assignment asks students to expose theirs.

## XI. Return to the Two Students

The second student's memo is better, and that matters. The question is whether either student can say what productivity should mean for this firm, why, and what would change her answer. The first student may need practice directing AI. The second may need to show that the frame was hers.

Pick one assignment where students synthesize sources. Write the misframed AI answer your students are most likely to get, one contribution worth accepting, and a changed case that makes different evidence decisive. Run it once and review the records with a colleague.

## References

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. *Proceedings of the National Academy of Sciences, 122*(26), e2422633122. <https://doi.org/10.1073/pnas.2422633122>

Bloom, N., Han, R., & Liang, J. (2024). Hybrid working from home improves retention without damaging performance. *Nature, 630*, 920–925. <https://www.nature.com/articles/s41586-024-07500-2>

Bloom, N., Liang, J., Roberts, J., & Ying, Z. J. (2015). Does working from home work? Evidence from a Chinese experiment. *Quarterly Journal of Economics, 130*(1), 165–218. <https://doi.org/10.1093/qje/qju032>

Boyer, P., & Nissenbaum, S. (1974). *Salem possessed: The social origins of witchcraft*. Harvard University Press.

Callister, W. D., & Rethwisch, D. G. *Materials science and engineering: An introduction* (fatigue chapter). Wiley.

Caporael, L. R. (1976). Ergotism: The Satan loosed in Salem? *Science, 192*(4234), 21–26.

Deng, Y., Brucks, M., & Toubia, O. (2026). Examining and addressing barriers to diversity in LLM-generated ideas. arXiv:2602.20408. <https://arxiv.org/abs/2602.20408>

Emanuel, N., Harrington, E., & Pallais, A. (2023). The power of proximity to coworkers: Training for tomorrow or productivity today? NBER Working Paper 31880. <https://www.nber.org/papers/w31880>

Karlsen, C. F. (1987). *The devil in the shape of a woman: Witchcraft in colonial New England*. W. W. Norton.

Norton, M. B. (2002). *In the devil's snare: The Salem witchcraft crisis of 1692*. Alfred A. Knopf.

Raees, M., & Papangelis, K. (2026). From trust to appropriate reliance: Measurement constructs in human-AI decision-making. arXiv:2604.23896. <https://arxiv.org/abs/2604.23896>

Spanos, N. P., & Gottlieb, J. (1976). Ergotism and the Salem Village witch trials. *Science, 194*(4272), 1390–1394.

Theobold, A. S. (2021). Oral exams: A more meaningful assessment of students' understanding. *Journal of Statistics and Data Science Education, 29*(2), 156–159. <https://www.tandfonline.com/doi/full/10.1080/26939169.2021.1914527>

Vendrell, M., & Johnston, S.-K. (2026). Scaffolding critical thinking with generative AI: Design principles for integrating large language models in higher education. *Computers and Education: Artificial Intelligence, 10*, 100572. <https://doi.org/10.1016/j.caeai.2026.100572>
