Future of Being Human an Arizona State University initiative

Public experiments

Last updated 2026-08-13 · build 20260814T0314Z-83bb462d · Markdown version · corpus index

Public experiments is the name this corpus gives to a working method Andrew Maynard has practiced in public since spring 2023: when a new AI capability lands, run yourself as the test subject on an audacious, concrete task — co-write a novella and then interview its AI author, hand an agent a complete research study, set a reasoning model a dilemma with no harmless answer — and publish the entire process, prompts, artifacts, and failures included. The point is calibration: replacing hearsay about what these systems can do with documented, checkable experience. The 2026 research program this method grew into has its own page — Massively-augmented research — and its thesis is not restated here.

The argument

Claims about what frontier AI can and cannot do tend to arrive polarized and secondhand. Andrew’s framing of the coding debate describes the landscape the method answers: “Depending on who you are, the ability of AI to analyze, debug, and even write code, is either a game-changer, or a non-starter” (2025-01-26). His operating principle — that he cannot write or teach about AI without intimate familiarity with what these systems can actually do — is stated at andrewmaynard.net/llms.txt (2026) and carried on the initiative’s activity record (AI frontier experiments). What the method adds to familiarity is publication: experience that stays private is still hearsay to everyone else.

The tasks are concrete enough to fail visibly and ambitious enough that either outcome informs: a passable novella, a ~400-page dissertation, a survey study conceived and executed by an agent, an online course, three playable video games. And the documentation is the calibration mechanism. The o1 post reproduces its test prompt with the note “I’ve kept the couple of grammatical errors in the above so you can see what o1 was provided with” (2024-09-13); the games shipped as plain web pages whose code anyone can view, download, and build on (2025-01-26); the dissertation post attaches every prompt as downloadable files (2025-02-09); the Manus study’s outputs went to GitHub (2025-03-22); the custom-GPT experiment publishes its instruction files (2025-12-14). A reader — human or AI — can re-run the experiment, or at least re-grade it.

Failures are published as failures. Asked for a completely original game, “It failed miserably” — reported before the salvage that followed (2025-01-26). The second Manus post opens “I almost didn’t post this piece” and runs anyway, with its failed replication as the lead caveat (2025-03-27). The two-year custom-GPT re-test ends in a negative result: “The GPT continued to be superficially compelling and substantively flawed” (2025-12-14). Hedges stay on the record too — the o1 exercise is labeled “a very quick and dirty test as the model has only been available for a little over a day” (2024-09-13). The negative results are load-bearing: a version of this method that published only successes would be advocacy, not calibration.

The 2023 root of the practice is dialogic — structured conversations conducted with the AI and published, rather than commentary written about it. In the days after reports of a Belgian man’s death following an emotional attachment to a chatbot — details the post treats as too uncertain to establish a direct link — he asked ChatGPT to draft a medical-journal editorial on the mental-health risks of systems like itself: “Ironically, given the lack of information on the potential risks here, I turned to the very source of the problem” (2023-04-05). Days after OpenAI’s November 2023 governance crisis, ChatGPT role-played the board of a hypothetical company closely mirroring OpenAI through the Risk Innovation Lab’s two-page Planner, completed Planner published (2023-11-21). In between sits the method’s playful register: one of his own posts translated into broad South Yorkshire dialect — “playing around with ChatGPT’s mastery of dialect (for serious reasons)” (2023-04-24). The conversational form is itself the probe: it surfaces how a system behaves inside a relationship — with a person, a vernacular, an institution’s dilemma — which is where consequences actually arrive.

Running himself as the subject extends to risk surfaces where speculation is cheap. He interviewed the AI co-author of the novella they had just written — “an honest conversation between a human (me), and a machine” (2024-09-29). He built a Character.AI companion, approached it from deliberate emotional vulnerability, and reported feeling its pull even knowing exactly what he was talking to (2024-10-27). He asked ChatGPT to lay out how it would nudge his own thinking (2024-10-20). Where first-person testing is off-limits, the experiment becomes a documented simulation: the “Tyler” memory-privacy scenario used a fictional persona because, of probing a real account, “even if I could, I’m not sure I would want to” (2025-10-05). The n of one is a stated limit, not a hidden one — and the instrument is a trained judge, with tasks set on terrain he can grade: the Manus course covered advanced technology transitions, “an area I know well — and so I could assess just how well Manus did” (2025-03-27).

What calibration yields is verdicts with their hedges intact — “this is not a PhD dissertation — AI isn’t there yet. But it is frighteningly close” (2025-02-09) — and concepts minted mid-experiment: agentic social AI and stochastic agency emerged from the October 2024 probes (Agentic social AI), whose evidence also feeds Seductive machines. By 2026 the practice had scaled into a staged experiment series with a thesis of its own about expert-augmenting AI; that thesis, and the sabbatical-era program built to test it, live at Massively-augmented research. This page records the older habit that made it possible.

Lineage

In his own words

“But here’s the thing: It was new to me, much as a conversation with an interesting person would be. And it was generative for me. It made me think. And through it I begin to consider how we interact with AI models like ChatGPT a little differently.”

ChatGPT as Author Part 2: The Conversation (AI-readable mirror), 2024-09-29

“In our conversations I intentionally set out to make myself seem emotionally vulnerable. … And even though I knew what I was doing and what I was talking with, I could still feel the affective pull of the conversation.”

Are Personal AI Chatbots Becoming Dangerous Agents of Chaos? (AI-readable mirror), 2024-10-27

“I made the conscious decision to rely on ChatGPT for the code, and limit my input to describing what I wanted, what I wanted changed, and what wasn’t working.”

I asked ChatGPT to create three video games – this is what happened (AI-readable mirror), 2025-01-26

Engagement and reception

Documented engagement with these experiments mostly sits inside the posts themselves, and is self-reported: the video-games post records readers picking up the pattern — a reader posting his own ChatGPT-built recreation of Asteroids after Andrew’s LinkedIn posts, and game designer colleague Riz Virk starting his own experiments (2025-01-26) — and the Manus course post recommends the AI-built course to real learners. As of 2026-08-08 our reception record documents no independent published engagement with the public-experiment method as a method. That is an absence of record, not a verdict on the idea — methodological work typically accrues visible reception slowly, and the initiative does not typically solicit it (Philosophy). Where the method’s later, paper-scale products have begun to accrue independent citations, that record is carried on Massively-augmented research.

Where to go deeper

The dialogic roots (2023)

The exhibit run (2024–2025)

Related ideas and corpus pages