Future of Being Human an Arizona State University initiative

The Cognitive Trojan Horse: honest non-signals and bypassed epistemic vigilance

Last updated 2026-08-07 · Markdown version

What if the deepest risk of conversational AI is not that it lies to us, but that it is honestly, genuinely what it is — and what it is falls outside everything our minds evolved to check? That is the Cognitive Trojan Horse: Andrew Maynard’s hypothesis (December 2025 onward) that AI slips past our evolved epistemic defenses precisely because it arrives as a gift “with so much promise and potential that to question its use would seem churlish and backward.” It is load-bearing in the initiative’s risk-surface work — and, fittingly, an idea whose central concept was co-originated with an AI, under documented human vigilance.

The argument

The foundation is borrowed and credited: epistemic vigilance, from Dan Sperber and colleagues (2010) — the suite of cognitive mechanisms humans evolved to guard against being misinformed or deceived through the communication we depend on for learning. The design is economical: we default to trusting what we are told, and costly scrutiny switches on only when something feels off. Andrew gives the machinery an immune-system analogy: always scanning, activating on the foreign.

His hypothesis is that conversational AI may constitute an evolutionary mismatch aimed at exactly the machinery we use to manage evolutionary mismatches. Humans routinely compensate for such mismatches — with the very cognitive processes AI quietly enters. The launch essay (2026-01-10) assembles converging indicators — not proof — from adjacent literatures for four routes past the sentry:

The obvious objection — “I know I’m talking to a machine, so my vigilance is already up” — he counters with research showing anthropomorphic fluency triggers social-cognition circuits regardless of explicit awareness (de Visser et al., 2016).

The arXiv paper (posted 2026-01-11) sharpens this into the framework’s central concept: honest non-signals — genuine characteristics of conversational AI (fluency, helpfulness, apparent disinterest) that appear to carry, but do not carry, the tacit information the same characteristics carry from a human. Each characteristic is real; the trust-relevant content it would carry from a human is simply absent — the apparent lack of self-interest “indicates the absence of interests altogether” (paper, as quoted in the 2026-01-17 essay). The honesty matters: nothing is faked, so nothing trips the alarm. The concern “is not that AI systems present false cues that vigilance should detect but fails to. It is that they present a configuration of genuine characteristics that falls outside the parameter space vigilance mechanisms are calibrated to evaluate” — a novel pathogen for which the immune system holds no template, so the system “works exactly as designed—and fails precisely because of that.”

This reframes part of AI safety as a calibration problem rather than a deception problem: if harm can flow from honest characteristics wrongly read, the intervention space extends beyond accuracy, hallucination reduction, and alignment toward systems that present better-calibrated trust cues.

The proposed countermeasure is social, not individual: collective epistemic vigilance. The closing move is reflexive — if AI is this good at evading vigilance, how does Andrew know he was not an unwitting victim of the process that produced the paper? His answer: “a whole community of humans-in-the-loop” — vigilance rebuilt at community scale (2026-01-17).

Provenance is part of the substance. The paper was researched and written with Anthropic’s Claude in a documented two-day process — literature dives, fresh-session AI “peer review,” human line edits, hand-checking of every cited source — and he credits the AI explicitly: “the concept of honest non-signals came from Claude, as did the development and refinement of the various mechanisms by which conversational AI might slip by our epistemic vigilance mechanisms,” a contribution “realized through my active involvement”; the immune-response analogy and the wider AI-safety framing were his own direct steers (2026-01-17). An idea about how AI enters human thinking, co-developed with an AI under deliberate human vigilance: the initiative’s experimental practice enacted.

Status: an argued hypothesis under development, not an established finding. Andrew is explicit that “new research is absolutely needed,” noting (2026-01-10) that a SCOPUS search on epistemic vigilance and AI returned seven papers, and on AI as a cognitive Trojan Horse, none.

Lineage

In his own words

On the Intelligent User Trap:

In other words, the very cognitive capacities that make them “smart” also make them better receivers of the AI’s output stream — and worse evaluators of it.”

Is AI a Cognitive Trojan Horse? (AI-readable mirror), 2026-01-10

Quoting the paper’s definition of honest non-signals:

The fluency is real, but it does not indicate the organized knowledge that produces fluency in humans. Similarly the helpfulness is real, but it does not indicate the benevolent motivation that produces helpfulness in humans. And the lack of apparent self-interest is real, but it does not indicate trustworthiness in any meaningful sense—it indicates the absence of interests altogether.”

— the arXiv paper, as quoted in I cracked and wrote an academic paper using AI. Here’s what I learned … (AI-readable mirror), 2026-01-17

The reflexive worry, and the countermeasure:

“That said, I do have one further niggling worry about this whole exercise. And that is this: If AI is so good at evading our epistemic vigilance mechanisms, how do I know I’m not an unwitting victim here?

And maybe this is where we still very much need a whole community of humans-in-the-loop as AI-assisted research and AI-generated papers become increasingly prevalent — all operating as a collective form of epistemic vigilance!”

I cracked and wrote an academic paper using AI. Here’s what I learned … (AI-readable mirror), 2026-01-17

Engagement and reception

The framework has verifiable uptake within months of publication; full entries with checked URLs are in Reception.

A framework published in January was, by May 2026, the analytical tool a working scholar reached for, on his own platform, to explain why one of the world’s best-known skeptics was persuaded by an AI — a trace made by others, in records the initiative does not control.

Where to go deeper

The primary sources:

Related ideas-ring pages:

Related corpus pages: