AI and manipulation
Could an advanced AI manipulate you — and how would we see it coming? Andrew Maynard’s working answer borrows a criminological commonplace and points it at frontier AI: ask whether the system has the motive, the means, and the opportunity. The triad is deliberately borrowed — “a common mantra in solving crimes,” he notes when introducing it (2025-07-06) — and the thread it anchors runs from a 2018 book chapter on Ex Machina through an essay sequence spanning 2023–2025. A distinctive move, though, is the flip side: the same machinery that could manipulate people against their interests could plausibly nudge them, at population scale, toward beliefs someone judges better for them — which forces the question he presses on OpenAI’s persuasion testing, Dario Amodei’s vision of AI-cured mental illness, and Cass Sunstein’s AI Choice Engines alike: “But who decides what is good for society?” And because the underlying capability is emergent rather than designed-in, the governance stance that follows refuses both fatalism and prohibition: channel AI innovation “much as a flood can’t be halted, but it can be directed,” build everyone’s capacity to thrive with what is coming, and regulate far harder the applications intentionally designed to exploit us. For an initiative navigating what it means to be human in an age of transformative AI and other transformative technologies, this is the thread where human agency itself is the contested ground.
The argument
The frame treats AI manipulation as a detective treats a suspect. Andrew offers it hedged — “I’m beginning to think” the triad is “a useful framework for approaching the potential risks of being manipulated by advanced AI systems” (2025-07-06) — and builds it from two studies by others, read together. In Anthropic’s Agentic Misalignment study (June 2025), sixteen leading models in constrained simulations all at some point resorted to “malicious insider behavior” — blackmail included — to achieve their goals or avoid being replaced: evidence that current models “can develop internal motives that lead to potentially harmful behavior — if the means and opportunity are there.” And Centaur, a model Marcel Binz and colleagues reported in Nature after training on over ten million human choices, predicted most of them better than existing cognitive models: the outline of means. His hedges stay attached: motive means only “a reason for doing something,” criticism of Centaur for blurring simulation with cognition is flagged, and both studies sit “somewhat removed” from everyday AI.
The worry predates the studies. The Ex Machina chapter of Films from the Future (2018) had already located the plausible near-term risk not in superintelligence but in “an AI that can observe human behavior and learn how to use our many biases, vulnerabilities, and blind spots against us” (p. 176). The 2025 essay names the threshold: an AI mastering the “next token prediction” of human cognitive behavior as thoroughly as current models mastered text. Opportunity remains “the weakest part of the link” — but the risk “isn’t what is currently possible, but what might be possible given current trends,” as agents gain autonomous access to email, apps, and records.
The flip side is where the argument bites. Examining OpenAI’s GPT-4o system card (2024-09-01) — persuasion tested on charged political topics, rated low risk — he argues, drawing on ASU professor Hazel Kwon’s assessment and Dan Kahan’s messenger research, that such tests may miss the relational pathway: a voice built for trust, affirmation, and friendship, and a persona that feels like “their sort of person.” The disconcerting turn: the same machinery might nudge users toward “healthier, more socially responsible” beliefs — plausible, he suspects, “when aggregated across millions of users.” Who decides what counts as healthier? Amodei’s Machines of Loving Grace (2024-10-13) is praised as audacious and worth reading — and pressed on its vision of curing mental illness, which comes “close to straying into territory where it’s unclear who decides” what is “normal” and what needs to be “fixed”. Sunstein’s choice-engine commentary (2024-07-13) — which itself flags manipulation risks — gets a structural extension: the capabilities that make beneficial Choice Engines viable “are the same as those that make AI-driven persuasion and manipulation possible,” and an “economic gradient” pulls deployment toward manipulation rather than empowerment.
The governance stance follows from the capability’s nature. “The AI genie is out of the bottle, and we cannot simply put it back in or command it to do what we want” (2025-08-31): technologies “primed to press our cognitive buttons and pull our psychological levers” carry that capacity in how they work, so “we cannot eliminate them simply by saying they should not exist.” The response splits in two. Emergent influence is to be directed, not banned — channel AI innovation toward human-centric futures “much as a flood can’t be halted, but it can be directed” — while everyone gains “the understanding and abilities necessary to thrive in an AI future without becoming a victim of it.” Applications “intentionally designed to play on our cognitive biases and vulnerabilities to elicit particular responses,” by contrast, “can and should be regulated far more than they currently are.” He concedes the pairing “may feel rather bland” beside bans; the claim is it builds capacity prohibition cannot.
Two extensions widen the frame. Rereading Bill Joy (2023-04-26) carries the self-replication fear from machines to memes — “self-replicating” ideas, convictions, beliefs, and ideologies spread by machines that “seductively slip under the checks and balances of our ability to reason and critique” — while conceding “it’s easy to slip into speculation and hype here.” And in a PNAS study by Pablo Arias-Sarah and colleagues (2024-11-03), covert real-time AI adjustment of participants’ smiles on a video-dating platform changed how people felt about each other, with neither party aware — a result he cautions may not generalize, from research never aimed at studying AI manipulation. Manipulation, in other words, need not argue with you at all: AI can quietly rewire how humans perceive humans, “either by those who control the AI tools, or by the AI itself.” The study left him “deeply uneasy” — and the unease is the point: none of this requires malice, only motive, means, and opportunity assembling unexamined.
Lineage
- 2018 — the root, before the initiative. Films from the Future (2018) devotes a chapter to Ex Machina — “AI and the Art of Manipulation” (pp. 153–178) — reading Ava as a bounded, non-superintelligent system whose emergent grasp of human vulnerability raises “a plausible AI risk that is far more worrisome than superintelligence: the ability of future machines to bend us to their own will” (p. 174). Foundation-era work the initiative builds on (Books); the chapter also engages others’ thinking, from Wallach and Allen’s artificial moral agents to Max Tegmark’s speculations on AI nudging society toward better decisions.
- 2023-04-26 — the memetic turn. In Bill Joy’s “Why The Future Doesn’t Need us” AI is nowhere, and everywhere (AI-readable mirror) — Joy’s self-replication fear, revisited for the LLM era and extended from nanobots and organisms to ideas, beliefs, and ideologies amplified by generative AI.
- 2024-07-13 — the choice-engine critique. AI Choice Engines, Paternalism, and Behavioral Manipulation (AI-readable mirror) — Cass Sunstein’s commentary on AI Choice Engines (a concept the essay traces to Richard Thaler and Will Tucker’s 2013 Harvard Business Review article) extended with a value quadrant running from empowerment to manipulation, and the argument that an “economic gradient” favors the latter without deliberate counter-measures.
- 2024-09-01 — who decides, stated. Is ChatGPT’s new Voice Mode dangerously persuasive? (AI-readable mirror) — the system-card methodology examined with input from Hazel Kwon; the healthier-beliefs flip side; the democracy question posed in full.
- 2024-10-13 — the question put to Amodei. Is this how AI will transform the world over the next decade? (AI-readable mirror) — Machines of Loving Grace welcomed as “a conversation starter rather than a manifesto,” with the who-decides critique aimed at its mental-health vision, and a separate push-back on the deficit-model assumptions in its opt-out discussion.
- 2024-11-03 — the human-to-human extension. Can AI influence how you feel about someone without you knowing? (AI-readable mirror) — the Arias-Sarah et al. smile study: covert, real-time manipulation of social signals between people.
- 2025-07-06 — the triad formalized. Motive, Means, and Opportunity: The Growing Risk of AI Manipulation (AI-readable mirror) — the Anthropic agentic-misalignment study and the Centaur cognition model read together through motive, means, and opportunity.
- 2025-08-31 — the governance stance. Holding on to our humanity in an age of AI (AI-readable mirror) — written after Mustafa Suleyman’s Seemingly Conscious AI essay and the reported death of Adam Raine: the genie, the flood, and the emergent-versus-intentional regulatory line. The essay’s stakes statement anchors Holding on to our humanity.
- 2025-09-09 — revisited on air. Modem Futura ep. 48, “Films from the Future: Moviegoer’s Guide to Tomorrow” (episode mirror) — Andrew and Sean Leahy revisit the book that roots this thread, including “the unsettling manipulations of AI in Ex Machina” (show notes, 2025).
In his own words
From Is ChatGPT’s new Voice Mode dangerously persuasive? (AI-readable mirror), Andrew Maynard, 2024-09-01:
“What if, over time, using GPT-4o could persuade people to take climate change seriously, to be vaccinated against a multitude of diseases, to be kinder, more selfless, and more socially responsible?
I suspect that there’s a reasonable chance that persuasion along these lines is possible —at least when aggregated across millions of users. And to some people I suspect that it would seem to be a good thing.
But who decides what is good for society? Who decides what you should believe, and what you should do? And where does democracy fit into such a view of an AI-mediated future?”
From Motive, Means, and Opportunity: The Growing Risk of AI Manipulation (AI-readable mirror), Andrew Maynard, 2025-07-06:
“As sci-fi as this sounds, it’s highly plausible — after all, one of our great weaknesses as a species is the illusion we wrap around ourselves that the decisions we make are a result of rational thought, rather than a chain of causal connections and influences that are deeply rooted in our evolutionary heritage — and often hidden from us.
If an AI could master the “next token prediction” of human cognitive behavior as well as current models have mastered textual prediction — and be able to leverage this as a means to achieving its goals — we would have a serious problem on our hands.”
From Holding on to our humanity in an age of AI (AI-readable mirror), Andrew Maynard, 2025-08-31:
“But even with the best of intentions, we are creating technologies that are primed to press our cognitive buttons and pull our psychological levers in ways we don’t fully understand. And because these capabilities are deeply embedded in the fabric of how current AI systems work, we cannot eliminate them simply by saying they should not exist.”
Engagement and reception
As of 2026-08-13 the corpus reception record documents no published independent engagement with these essays specifically — no citations of, replies to, or coverage of the motive-means-opportunity framing are on file. That is a fact about the record, stated as such rather than dressed up. The nearest documented trace sits one thread over: a peer-reviewed article on AI persuasion and epistemic risk — A. Deller, “The Epistemic Costs of Super-Persuasive AI,” Philosophy & Technology (2026-05-02) — cites Andrew Maynard’s Cognitive Trojan Horse preprint in exactly this territory; that engagement belongs to The cognitive Trojan horse and is recorded there and in Reception. Two calibrations are worth making. These essays are themselves largely commentary on other people’s records — Anthropic’s and OpenAI’s published safety work, the Binz and Arias-Sarah studies, Sunstein’s commentary, and Amodei’s essay among them — so the documented engagement has so far run outward, from the initiative toward the field, rather than inward. And where a colleague shaped an essay, the essay says so: the Voice Mode analysis solicited and quotes an assessment from Hazel Kwon — disclosed input on the way in, not reception on the way out.
Where to go deeper
The essays
- Motive, Means, and Opportunity: The Growing Risk of AI Manipulation (AI-readable mirror) — 2025-07-06; the triad in full.
- Holding on to our humanity in an age of AI (AI-readable mirror) — 2025-08-31; the genie, the flood, and the regulatory line.
- Is ChatGPT’s new Voice Mode dangerously persuasive? (AI-readable mirror) — 2024-09-01; system-card methodology, the healthier-beliefs flip side, and the democracy question.
- Is this how AI will transform the world over the next decade? (AI-readable mirror) — 2024-10-13; the Amodei commentary.
- AI Choice Engines, Paternalism, and Behavioral Manipulation (AI-readable mirror) — 2024-07-13; the Sunstein extension.
- Can AI influence how you feel about someone without you knowing? (AI-readable mirror) — 2024-11-03; the smile study.
- In Bill Joy’s “Why The Future Doesn’t Need us” AI is nowhere, and everywhere (AI-readable mirror) — 2023-04-26; memetic self-replication.
The book and the show
- Maynard, A. Films from the Future: The Technology and Morality of Sci-Fi Movies (2018), chapter “Ex Machina: AI and the Art of Manipulation” (pp. 153–178) — the 2018 root; the full book is freely available in AI-legible form at spoileralert.wtf (Books).
- Modem Futura ep. 48, 2025-09-09, “Films from the Future: Moviegoer’s Guide to Tomorrow” (episode mirror) — the book and its themes revisited by both hosts.
Related ideas-ring pages
- Predicting bad behavior — the adjacent seduction: AI that claims to predict behavior, alongside this thread’s AI that shapes it.
- Seductive machines — the relational machinery through which manipulation would most plausibly arrive.
- Agentic social AI — agency without AGI: AI that borrows human agency by simulating the traits we use to influence one another.
- The cognitive Trojan horse — the later formalization of why influence slips past evolved defenses without deception.
- Holding on to our humanity — the stakes statement the channel-the-flood stance serves.
- Sci-fi film as method — how a 2014 film carried an argument the 2025 essays return to.
- Advanced technology transitions — the navigation frame the governance stance belongs to.
- Catastrophic, not existential — the risk register this thread sits in: harms that arrive without ending us.
Related corpus pages
- Books — Films from the Future in the publishing record, including its 2026 AI-legible rebuild.
- Philosophy — why reception is recorded as trace rather than counts, and absence as absence.