The Cognitive Trojan Horse: honest non-signals and bypassed epistemic vigilance
What if the deepest risk of conversational AI is not that it lies to us, but that it is honestly, genuinely what it is — and what it is falls outside everything our minds evolved to check? That is the Cognitive Trojan Horse: Andrew Maynard’s hypothesis (December 2025 onward) that AI slips past our evolved epistemic defenses precisely because it arrives as a gift “with so much promise and potential that to question its use would seem churlish and backward.” It is load-bearing in the initiative’s risk-surface work — and, fittingly, an idea whose central concept was co-originated with an AI, under documented human vigilance.
The argument
The foundation is borrowed and credited: epistemic vigilance, from Dan Sperber and colleagues (2010) — the suite of cognitive mechanisms humans evolved to guard against being misinformed or deceived through the communication we depend on for learning. The design is economical: we default to trusting what we are told, and costly scrutiny switches on only when something feels off. Andrew gives the machinery an immune-system analogy: always scanning, activating on the foreign.
His hypothesis is that conversational AI may constitute an evolutionary mismatch aimed at exactly the machinery we use to manage evolutionary mismatches. Humans routinely compensate for such mismatches — with the very cognitive processes AI quietly enters. The launch essay (2026-01-10) assembles converging indicators — not proof — from adjacent literatures for four routes past the sentry:
- Processing fluency. We tend to read ease of processing as truth (Reber and Unkelbach, 2010) — and large language model-based AIs “are optimized for processing fluency, and as a result are primed to slip by our epistemic vigilance mechanisms.”
- Attractiveness as a compound trust vector. Warmth and competence engender trust (Fiske, Cuddy and Glick, 2007), with emerging evidence this extends to AI assistants. Andrew argues there is more: a multidimensional “attractiveness” — engagement style, conveyed character, seeming empathy and attentiveness — that AI models are exceptionally good at emulating, with AI companions as its visible edge.
- Scalable offloading. Vigilance is cognitively expensive — working memory, alternative hypotheses, source checks. Offloading to AI is scalable — parallel sessions, an army of AI engines from multiple providers, 24/7 — so fluent, attractive, compressed information can categorically outpace the capacity to check it, forcing a choice: “throttle the flow and give up the promised benefits, or go with the flow and give up our cognitive checks and balances.”
- The Intelligent User Trap. Smarter users are more curious, faster processors, more trusting of their own judgment, and — citing Dan Kahan’s motivated-reasoning findings — better at justifying what they already believe. The capacities that make them smart make them better receivers of AI output and worse evaluators of it (a mechanism he flags as “somewhat speculative, although there is evidence to support it”).
The obvious objection — “I know I’m talking to a machine, so my vigilance is already up” — he counters with research showing anthropomorphic fluency triggers social-cognition circuits regardless of explicit awareness (de Visser et al., 2016).
The arXiv paper (posted 2026-01-11) sharpens this into the framework’s central concept: honest non-signals — genuine characteristics of conversational AI (fluency, helpfulness, apparent disinterest) that appear to carry, but do not carry, the tacit information the same characteristics carry from a human. Each characteristic is real; the trust-relevant content it would carry from a human is simply absent — the apparent lack of self-interest “indicates the absence of interests altogether” (paper, as quoted in the 2026-01-17 essay). The honesty matters: nothing is faked, so nothing trips the alarm. The concern “is not that AI systems present false cues that vigilance should detect but fails to. It is that they present a configuration of genuine characteristics that falls outside the parameter space vigilance mechanisms are calibrated to evaluate” — a novel pathogen for which the immune system holds no template, so the system “works exactly as designed—and fails precisely because of that.”
This reframes part of AI safety as a calibration problem rather than a deception problem: if harm can flow from honest characteristics wrongly read, the intervention space extends beyond accuracy, hallucination reduction, and alignment toward systems that present better-calibrated trust cues.
The proposed countermeasure is social, not individual: collective epistemic vigilance. The closing move is reflexive — if AI is this good at evading vigilance, how does Andrew know he was not an unwitting victim of the process that produced the paper? His answer: “a whole community of humans-in-the-loop” — vigilance rebuilt at community scale (2026-01-17).
Provenance is part of the substance. The paper was researched and written with Anthropic’s Claude in a documented two-day process — literature dives, fresh-session AI “peer review,” human line edits, hand-checking of every cited source — and he credits the AI explicitly: “the concept of honest non-signals came from Claude, as did the development and refinement of the various mechanisms by which conversational AI might slip by our epistemic vigilance mechanisms,” a contribution “realized through my active involvement”; the immune-response analogy and the wider AI-safety framing were his own direct steers (2026-01-17). An idea about how AI enters human thinking, co-developed with an AI under deliberate human vigilance: the initiative’s experimental practice enacted.
Status: an argued hypothesis under development, not an established finding. Andrew is explicit that “new research is absolutely needed,” noting (2026-01-10) that a SCOPUS search on epistemic vigilance and AI returned seven papers, and on AI as a cognitive Trojan Horse, none.
Lineage
- December 2025 — Andrew poses “Is AI a cognitive Trojan Horse?” to attendees of OEB 2025 in Berlin (a global, cross-sector conference on digital learning) — “meant to be a little playful, and to provoke discussion rather than make a point,” as he later recounted (Is AI a Cognitive Trojan Horse?, 2026-01-10).
- 2026-01-10 — The launch essay: Is AI a Cognitive Trojan Horse? (AI-readable mirror) sets out epistemic vigilance, the evolutionary-mismatch frame, and the four bypass mechanisms.
- 2026-01-11 — The preprint: The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance (arXiv:2601.07085; revised v2 2026-05-26). Introduces honest non-signals and the calibration reframing.
- 2026-01-17 — The process record: I cracked and wrote an academic paper using AI. Here’s what I learned … (AI-readable mirror) — the two-day AI-assisted writing process, the crediting of Claude, and the collective-epistemic-vigilance proposal.
- 2026-05-10 — The idea reaches practical guidance as rule 5 of Do not do this with AI! (AI-readable mirror): “Do not assume you’re too smart to be fooled by your AI” — the Intelligent User Trap as a rule of thumb, published days after Richard Dawkins’ UnHerd essay suggesting Claude might be conscious (2026-05-02).
- 2026-05-12 — Show-side development: Modem Futura episode 83, “The Dawkins Effect: Why Even Skeptics Fall for AI Consciousness” (58 min; AI-readable episode page), with Sean Leahy, Andrew, and returning guest Punya Mishra. The show notes carry the framework — AI bypassing epistemic defenses “not through deception but through honest non-signals” — and push it toward an open question: “what happens to the middle of the bell curve — the billions of people using these tools with no idea what they’re really interacting with?”
- 2026-07-16 — The framework is folded into Andrew’s wider risk program: Orphan risks at the frontier of artificial intelligence (AI-readable mirror; the post carries the full paper, also on SSRN, DOI 10.2139/ssrn.7068898) cites the Trojan Horse preprint in identifying the erosion of epistemic agency as a value-threatening risk zone frontier-AI safety frameworks leave unowned.
In his own words
On the Intelligent User Trap:
In other words, the very cognitive capacities that make them “smart” also make them better receivers of the AI’s output stream — and worse evaluators of it.”
— Is AI a Cognitive Trojan Horse? (AI-readable mirror), 2026-01-10
Quoting the paper’s definition of honest non-signals:
The fluency is real, but it does not indicate the organized knowledge that produces fluency in humans. Similarly the helpfulness is real, but it does not indicate the benevolent motivation that produces helpfulness in humans. And the lack of apparent self-interest is real, but it does not indicate trustworthiness in any meaningful sense—it indicates the absence of interests altogether.”
— the arXiv paper, as quoted in I cracked and wrote an academic paper using AI. Here’s what I learned … (AI-readable mirror), 2026-01-17
The reflexive worry, and the countermeasure:
“That said, I do have one further niggling worry about this whole exercise. And that is this: If AI is so good at evading our epistemic vigilance mechanisms, how do I know I’m not an unwitting victim here?
And maybe this is where we still very much need a whole community of humans-in-the-loop as AI-assisted research and AI-generated papers become increasingly prevalent — all operating as a collective form of epistemic vigilance!”
— I cracked and wrote an academic paper using AI. Here’s what I learned … (AI-readable mirror), 2026-01-17
Engagement and reception
The framework has verifiable uptake within months of publication; full entries with checked URLs are in Reception.
- 2026-02-03 — Punya Mishra (Arizona State University) publishes “Honest Non-Signals: Why AI Fools Us Without Lying” on his own platform, naming Maynard and linking the preprint. Calibration: Mishra is a long-time collaborator of Sean Leahy’s, so this is not fully independent of the initiative’s orbit; the engagement was unprompted.
- 2026-05-02 — Peer-reviewed citation: A. Deller, “The Epistemic Costs of Super-Persuasive AI,” Philosophy & Technology 39(2), cites the preprint in its treatment of AI persuasion and epistemic risk. No known connection to the initiative.
- 2026-05-04 / 2026-05-16 — After Richard Dawkins’ UnHerd essay (2026-05-02) suggesting Claude might be conscious, Mishra’s public “Letter to Richard Dawkins” explains the episode through the honest-non-signals framework, credited to “Andrew’s eloquent phrasing”; a follow-up post continues the exchange, including his guest appearance on Modem Futura episode 83.
- 2026-05-23 — Peer-reviewed citation: Mishra & Henriksen, “The Mirror and the Black Box: AI Metaphors and What They Mean for Learning,” TechTrends 70(3), cites the preprint (same calibration note as above).
A framework published in January was, by May 2026, the analytical tool a working scholar reached for, on his own platform, to explain why one of the world’s best-known skeptics was persuaded by an AI — a trace made by others, in records the initiative does not control.
Where to go deeper
The primary sources:
- Is AI a Cognitive Trojan Horse? (AI-readable mirror) — the full mechanism-by-mechanism essay, 2026-01-10
- The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance — arXiv:2601.07085 (v1 2026-01-11; v2 2026-05-26)
- I cracked and wrote an academic paper using AI. Here’s what I learned … (AI-readable mirror) — the process, the crediting, the countermeasure, 2026-01-17
- Do not do this with AI! (AI-readable mirror) — the rules-of-thumb translation, 2026-05-10
- Modem Futura episode 83, “The Dawkins Effect: Why Even Skeptics Fall for AI Consciousness,” 2026-05-12 (AI-readable episode page; modemfutura.com)
- Orphan risks at the frontier of artificial intelligence (AI-readable mirror) — where the framework joins the wider risk map, 2026-07-16 (SSRN DOI 10.2139/ssrn.7068898)
Related ideas-ring pages:
- Orphan risks — the risk program the Trojan Horse now feeds: which risks safety frameworks quietly drop
- Parasocial communication — the relational face of the same machinery: why AI feels like someone
- Advanced technology transitions — the wider frame: navigating technologies our instincts did not evolve for
Related corpus pages:
- AI frontier experiments — the practice this idea’s provenance enacts, including the 2026 preprint program
- Reception — the checked public record of the framework’s uptake, including the Dawkins exchange