Adrian de Wynter, a researcher at Microsoft, spent two years reading computer science papers about large language models. By his count, 57% of them assumed human-like traits in the technology before a single experiment ran, and 77% of the papers that set out to test for those traits concluded they were there. An assumption built into the premise has a way of finding its way back out as a result.

De Wynter's response was not another position paper. He opened the scenario editor in Age of Empires II, a real-time strategy game from 1999, and built logic gates out of goats. A goat standing on grass reads as a zero. A goat standing on a bridge reads as a one. Wire enough of these gates together and the goats start carrying information, flipping states, completing circuits. Then he described the whole thing using the same vocabulary AI researchers use for language models. The goats became "agentic." They "cooperated." They "passed information." Read the description without knowing what it referred to, and it sounds like a system that understands something. It is goats walking onto pressure plates.

The Trick Was Never in the System

What de Wynter's experiment isolates is that the words did the work, not the mechanism. Logic gates made of livestock and logic gates made of silicon produce the same kind of description once you reach for words like "decide" and "communicate." The vocabulary is doing something the underlying process is not. It is supplying a mind where there is only a mechanism, and it is doing it convincingly enough that researchers trained to be skeptical fell for it at a rate of better than three in four when they went looking.

This is not a new tendency. People see faces in electrical outlets and hear their names called in empty rooms. What is new is that large language models give that tendency something to work with that previous technologies did not: fluent, responsive, conversational output. A goat on a bridge does not talk back. A chat window does. The same anthropomorphizing impulse that finds faces in outlets now has a much better canvas, and the words researchers, journalists, and product teams choose to describe what is happening behind that canvas determine how much of a mind the user assumes is on the other side.

The vocabulary built a mind where the goats built nothing but a circuit. The danger is that the same vocabulary is doing identical work on systems people actually rely on.

Where This Meets the Understanding-Trust Gap

My dissertation research found that trust in an AI system and a person's actual comprehension of how it works do not move together. Under time pressure, trust goes up while measured understanding goes down. People were not trusting the system because they understood it. They were trusting it because it behaved in ways that read as competent, confident, and, frequently, intentional.

De Wynter's goats explain a piece of where that misplaced confidence comes from before a user ever opens an interface. If the language surrounding a system already implies intention, the user starts from a position of attributing a mind, and every subsequent interaction gets interpreted through that frame. A flagged error reads as the system "catching" something. A confident wrong answer reads as a "judgment call" rather than a statistical artifact. None of that is happening because the user is careless. It is happening because the words available to describe the system already did the anthropomorphizing before the user got involved.

This is also why transparency alone has not closed the gap in my research. Explaining a system's internals does little to correct a misattribution that was never about the internals in the first place. It was about language operating on a much older perceptual habit. A user can read a technically accurate explanation of how a model generates text and still walk away describing the output as something the model "wanted" to say, because the vocabulary of intention is doing more persuasive work than the explanation of mechanism.

Absurdism as Method

De Wynter has said he tends to "dial up things to 11" when he wants to make a point stick, and that absurdism has a real place in computer science and philosophy. The goats work as an argument precisely because they are silly. Nobody believes the goats are agentic. That certainty is the point. If the description sounds intelligent for goats walking across a game map, the description was never proof of intelligence to begin with. It was a vocabulary problem wearing a technical costume.

The harder case is the one where the underlying system is genuinely sophisticated, where the outputs are fluent and often correct, and where the same anthropomorphic vocabulary is harder to dismiss because it is not attached to anything as obviously absurd as livestock. That is precisely the case where the habit matters most to catch, because the consequences of misplaced trust scale with how much the system is actually relied on.

What the Goats Leave Us With

The fix is not to strip all human-sounding language from how AI systems get described. Some of that language is genuinely useful shorthand, and insisting on bloodless mechanical descriptions for every interaction is not realistic. The fix is closer to what de Wynter modeled with his own paper: noticing when the vocabulary has run ahead of the evidence, and building in a habit of checking whether a claim of agency, cooperation, or reasoning would still hold up if it were a description of goats on a bridge.

That habit is a small, specific version of the agency-over-transparency argument I have made elsewhere. Giving a person genuine interactive control over a system, the ability to test it, push back on it, and watch it fail in front of them, does more to correct an inflated mental model than any explanation of what is happening under the hood. The goats never needed a disclaimer. They needed someone to walk onto the bridge and look.

Source

Antonia Davison, "A Microsoft researcher sent some goats on a mission in Age of Empires II to make a point about LLMs," IBM Think, June 26, 2026. ibm.com/think

DH

Debra Hogue, PhD

Computer Scientist · Human-AI Collaboration Researcher · Oklahoma