Massively-augmented research
Massively-augmented research is Andrew Maynard’s name — coined in July 2026, early in a sabbatical devoted largely to public experimentation with frontier AI systems — for where he suspects frontier AI’s near-term significance for scholarship actually lies: not in autonomous discovery, but in the massive augmentation of an established expert who keeps epistemic control. The idea arrives with working parts: a replicable method, staged experiments that carry their own controls, and published accountability instruments — AI use statements and AI self-audits — with every stage documented in public as it happened. This is active work, a longitudinal experiment in progress rather than a settled finding: AI frontier experiments is the initiative’s activity record; this page is the idea.
The argument
The thesis was stated on 2026-07-04, after a month of staged public experiments with Anthropic’s Fable 5 model: as frontier models grow more capable, the place to watch, Andrew suspects, is “not so much in autonomous research and development … but in massively-augmented research and development.” The construction is deliberate: not a claim that autonomous AI scholarship is out of reach, but that the near-term significance runs through a frontier model working under a domain expert who keeps control of the judgments. The series’ opening post had already fixed the criterion: working with Claude “vastly accelerated the rate at which I could work, without me feeling as if I’d lost intellectual control” (2026-01-17).
And it is more than testimony, because the experiments came with their own baselines. One-shot prompting — something he “would usually never do” because it “tends to lead to outputs that are superficially OK and substantially poor” — became a deliberate diagnostic of the unassisted ceiling: the 2026-06-10 test asked Fable 5 for a complete research paper from a single prompt, judged in retrospect “superficially impressive, but substantially shallow” (2026-07-04). Two days later the same task went to Fable running agentically in Claude Code (2026-06-12). The output was better — and surfaced a different problem: evaluating such work demands expert labor most readers will not spend — which suggests, he wrote, that “any quick responses to AI-generated work like this are either coming from genius-class humans, are themselves the product of AI, or are not based on knowledgeable assessment.” His closing hedge: “we are potentially at the edge of a precipice where AI systems are capable of generating new knowledge and insights faster than we are currently capable of validating and even understanding them — or their consequences.”
The third stage put the model where the thesis says it belongs: beside an expert, on the expert’s own terrain. The 2026-07-04 experiment set Fable to work on Andrew’s risk-innovation scholarship of the past several years — ground where he was “uniquely positioned” to judge the output. The result became the orphan-risks paper (see Orphan risks): framing and analysis he calls “insightful — and genuinely novel”, produced in roughly two days rather than months — its value, on his own reading, “largely arises because of my previous work and my hands-on editing. This is not AI acting as an independent researcher.”
The method is documented for others to replicate (2026-01-17, refined through July 2026): iterative drafting with the model; adversarial “peer review” of drafts by the same model in fresh sessions — the circularity objection met in a footnote: a new session, in his experience, “has sufficient independence when augmented by human expert insight to provide valuable critical feedback”; line editing “very much in line with what I would have provided an accomplished grad student co-author”; manual download of cited works and verification of every cited source; and publication with a formal AI use statement plus, later, a published AI self-audit. The orphan-risks paper’s statement is the template: the research question, argument architecture, key concepts and “all editorial judgments” are the author’s; the model worked “under the author’s close direction”; “The author takes full responsibility for all content, claims and citations” (2026-07-16).
Companion practices extend this. Simulated-user evaluation routes assessment through firewalled agents: the 2026-03-29 degree-proposal experiment sent a draft to a second, separate Claude Code project for four independent reviews — pedagogy, employers, students, parents; the Hyperbubble build (2026-07-10) ran a concept competition scored by three AI judges, then an adversarial review where every claimed bug went to an independent verifier instructed to refute it. The habit predates Fable and crosses platforms: a 2025-05-25 footnote discloses “Using ChatGPT (model o3) as my “reviewer 2”” — and argues back against the objection it raised, in the same footnote.
The normative line is explicit. Using AI as “an academic profile-padder” he finds “distasteful”; “AI-assisted discovery and insights as a public good feels like something we should be embracing” (2026-01-17). The same post names the failure mode — academic literature threatened by “a tsunami of pseudo-intellectual AI slop” — and its alternative: a process that worked “because of how I used AI — not as a “slop prop,” but as a powerful research tool that extended what I was able to do, without supplanting my own intellectual contributions” (2026-01-17). The boundary runs through purpose and disclosed process, not through whether AI touched the work — and through honesty about the labor: “something like a 10:1 ratio of my time to Claude Code’s” (2026-03-29), and being left “deeply suspicious of anyone who claims they can get AI to churn out publishable papers in a matter of hours” (2026-07-04).
The thesis carries its own limits: the style ceiling — Fable’s published audit concedes “the machine has a style, not just a vocabulary,” requiring “a discerning human editor to keep catching it” (2026-07-04); the evaluation bottleneck; and a reflexive worry about the method itself — “If AI is so good at evading our epistemic vigilance mechanisms, how do I know I’m not an unwitting victim here?” — answered, provisionally, by community-scale human vigilance (2026-01-17; The Cognitive Trojan Horse). And it runs, deliberately, beside its opposite: AI-free intellectual craft, practiced by the same scholar in the same year (The artisanal intellectual).
Lineage
- 2025-02-04 / 2025-02-09 — the provocation. Does OpenAI’s Deep Research signal the end of human-only scholarship? (AI-readable mirror) and Can AI write your PhD dissertation for you? (AI-readable mirror) — four days with OpenAI’s Deep Research yield a 400-page AI-generated dissertation; the exercise the 2026-01-17 post cites as its starting point.
- 2025-05-25 — the “reviewer 2” habit, disclosed. Why parasocial communication around complex ideas is important – and why we need more of it (AI-readable mirror) carries, in a footnote, ChatGPT o3 as “reviewer 2” — adversarial AI self-review disclosed and argued with in print, pre-Fable and cross-platform.
- 2026-01-17 — the method, end to end. I cracked and wrote an academic paper using AI. Here’s what I learned … (AI-readable mirror) — the Cognitive Trojan Horse paper (arXiv:2601.07085, posted 2026-01-11) written in around two days; drafting, fresh-session review, line edits, citation checks and the profile-padding line all documented.
- 2026-03-29 — the orchestration test. Can AI create a comprehensive degree program proposal in the time it takes to grab a coffee? (AI-readable mirror) — a 223-page degree proposal via multi-agent Claude Code; the firewalled four-perspective review; the 10:1 editing ratio.
- 2026-06-10 / 2026-06-12 — the Fable diagnostic pair. Is Anthropic’s new AI model poised to change the AI higher education landscape … again? (AI-readable mirror) — the one-shot ceiling test — and A quick update on using Claude Fable 5 for research (AI-readable mirror) — the agentic re-run and the evaluation problem. Both posts also record the 2026-06-12 U.S. export-control suspension of Fable access, lifted July 1.
- 2026-07-04 — the thesis named. Just how good is Anthropic’s Fable at researching and writing an academic paper? (AI-readable mirror) — partner mode on his own terrain; “massively-augmented research and development” coined; Fable’s process audit published with the post.
- 2026-07-10 — the division of labor, audited. I asked Anthropic’s Fable 5 to create a video game inspired by my work. It’s mad! (AI-readable mirror) — the Hyperbubble game, with a self-audit written by Claude Fable 5 at Andrew’s request, which he endorsed and published.
- 2026-07-16 / 2026-07-19 — the accountability instruments in use. Orphan risks at the frontier of artificial intelligence (AI-readable mirror) publishes the paper with its AI use statement; Publish or Perish: AI vs Human (AI-readable mirror) puts Andrew’s rewrite and Claude’s draft side by side for readers to judge.
In their own words
“… as a research partner, these models are capable of seriously augmenting what an established expert/researcher is capable of achieving. And this, I suspect, is the where [sic] we need to be paying attention in the near future as these models only get more powerful—not so much in autonomous research and development (although I have no doubt that this is coming), but in massively-augmented research and development.”
— Andrew Maynard, Just how good is Anthropic’s Fable at researching and writing an academic paper? (AI-readable mirror), 2026-07-04
“Using AI as an academic profile-padder is something I still find distasteful — even though it’s never been easier to churn out new papers by the dozen using artificial intelligence. And yet, AI-assisted discovery and insights as a public good feels like something we should be embracing … as long as we can work out how to ensure the latter without the hollow self-aggrandizement of the former.”
— Andrew Maynard, I cracked and wrote an academic paper using AI. Here’s what I learned … (AI-readable mirror), 2026-01-17
“This document was written by the system it describes, which is a limitation worth stating plainly. … the corrections that mattered most (the game was too fast, flourishing was too easy, the panics were too crude, a whole feature deserved cutting) all came from the one human playing it. That division of labor is probably the honest headline: the machinery generated, tested, and repaired at scale; the judgment about what was worth keeping stayed human.”
— from “Making HYPERBUBBLE: a self-audit”, written by Claude Fable 5 at Andrew’s request — an audit he endorsed and published in I asked Anthropic’s Fable 5 to create a video game inspired by my work. It’s mad! (AI-readable mirror), 2026-07-10
Engagement and reception
The record on the thesis itself is thin, and we prefer to say so plainly: the posts ran between January and July 2026, and as of 2026-08-06 our reception record documents no independent published engagement with the massively-augmented-research argument as such. That is an absence of record, not a verdict on the idea — conceptual work of this kind typically accrues visible reception slowly, and the initiative does not typically solicit it (Philosophy).
What is checkable is the reception of the method’s first product. The AI Cognitive Trojan Horse preprint (arXiv:2601.07085, posted 2026-01-11) — the paper whose AI-assisted production the 2026-01-17 post documents — was cited in two peer-reviewed journals within months: Deller, Philosophy & Technology 39(2), published 2026-05-02, and Mishra & Henriksen, TechTrends 70(3), published 2026-05-23 (the latter not fully independent of the initiative’s orbit — Punya Mishra is a long-time collaborator of Sean Leahy’s; details and caveats in Reception). Those citations engage the paper’s argument, not the workflow that produced it — a distinction we keep explicit.
Where to go deeper
The experiment, post by post
- I cracked and wrote an academic paper using AI. Here’s what I learned … (AI-readable mirror) — 2026-01-17; the method stated end to end.
- Can AI create a comprehensive degree program proposal in the time it takes to grab a coffee? (AI-readable mirror) — 2026-03-29.
- Is Anthropic’s new AI model poised to change the AI higher education landscape … again? (AI-readable mirror) — 2026-06-10 — and A quick update on using Claude Fable 5 for research (AI-readable mirror) — 2026-06-12.
- Just how good is Anthropic’s Fable at researching and writing an academic paper? (AI-readable mirror) — 2026-07-04; the thesis post.
- I asked Anthropic’s Fable 5 to create a video game inspired by my work. It’s mad! (AI-readable mirror) — 2026-07-10; the Hyperbubble build and self-audit.
- Orphan risks at the frontier of artificial intelligence (AI-readable mirror) — 2026-07-16; the paper and its AI use statement.
- Publish or Perish: AI vs Human (AI-readable mirror) — 2026-07-19; the reader-judged AI-vs-human comparison.
Papers and episodes
- The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance — arXiv:2601.07085 (2026-01-11); the method’s first paper.
- Orphan Risks at the Frontier of Artificial Intelligence — SSRN preprint, DOI 10.2139/ssrn.7068898 (July 2026); the partner-mode paper.
- Modem Futura episode 92, 2026-07-14, “AI as Critical Infrastructure: Inside the Anthropic Fable Export Ban” (AI-readable episode page) — Sean Leahy and Andrew on the export-control suspension that interrupted the experiments.
Related ideas and corpus pages
- The Cognitive Trojan Horse — the idea the method first produced, and the vigilance worry the method turns on itself.
- Orphan risks — the paper written in partner mode.
- The artisanal intellectual — the deliberate counter-practice: AI-free intellectual craft, held in the same year by the same scholar.
- Writing for AI — the adjacent practice of making work legible to AI readers.
- AI frontier experiments — the initiative’s activity record of this work; Scholarship — where the 2026 preprints sit in the wider arc; Philosophy — why public experimentation without KPIs is the method.