B. AI Across the Research Lifecycle
Purpose: AI use in research is currently dominated by programmers and data analysts — tasks where it demonstrably works because success criteria are clear and outputs are verifiable. This document makes the case that AI can be used across the entire research project lifecycle, not just its technical phases. It maps where AI is productive, where it is fragile, and what the difference comes down to. Two moments are fragile in a specific and dangerous way — AI actively misleads, the misleading is hard to detect, and the damage compounds downstream: creative problem formation and literature finding. These are treated in depth in companion documents.
The foundational principle running through all of this: research is not a waterfall. It is a series of micro-discoveries, each one capable of redirecting the work. The design principle is not "try not to fail" but "fail quickly and let it be documented" — discover problems as early as possible, when correction is cheap, and treat every redirection as information rather than deviation. AI accelerates the feedback loops that make this possible. → B.lifecycle.0.iteration
The five phases: → B.lifecycle.1.creativity — research question formation in depth → B.lifecycle.2.literature — literature engagement in depth → B.lifecycle.3.datacapture — data collection and processing in depth → B.lifecycle.4.dataanalysis — analysis: execution vs. interpretation → B.lifecycle.5.manuscript — manuscript writing as thinking; the 5↔1 feedback loop
The programmer model and its limits
The current dominant model of AI-assisted work comes from software development: you have a task, you give it to AI, you check the output. This works because the task is well-defined and success criteria are external — the code runs or it does not. The feedback loop is tight and objective.
This model has spread into research through data science: run this regression, clean this dataset, generate this visualisation. These tasks are also relatively well-defined and their outputs are verifiable.
The claim of this document is that the model can be extended to the whole research lifecycle — but only if the extension is done with a clear understanding of where the analogy holds and where it breaks down.
The organizing principle: engineering-like vs. interpretive
Research is not an engineering problem. Its hardest and most important moments — forming a question, situating it in a field, deciding what evidence means, building an argument — operate on standards that are internal to a scholarly community, partly contested, and not reducible to external verification. These are not defects in research; they are features of genuine inquiry into hard questions.
But research contains a wide range of subproblems that are genuinely mechanical: finding and formatting sources, processing documents, transforming data, organising notes, managing files. These have clear success criteria and are straightforwardly delegable.
Between these poles is a gradient:
More engineering-like More interpretive
| |
Format → Search → Synthesize → Analyse → Interpret → Form question
AI is productive in proportion to how far left on this gradient the current task sits. Not because AI is a mechanical tool — it is not — but because the further right you go, the more the quality of the output depends on things Claude does not have: your position in the field, your intellectual history, your judgment about what matters to your community right now.
The practical implication: the question is not "should I use AI for my research?" but "which part of my research am I currently in, and what role for AI is appropriate here?"
Research is not a waterfall
The gradient above describes the character of individual tasks — how engineering-like or interpretive they are. It does not describe the order in which they happen.
The naive model of the research lifecycle is sequential: form a question, design a study, collect data, analyse it, write it up. This is how finished research gets reported. It is not how research gets done.
Research is a journey into the unknown. The question you start with is rarely exactly the question you end up answering. The design that seemed sound before collection meets material that does not fit it. The analysis produces findings that the original question did not anticipate. The manuscript reveals that what you thought you had established, you have not quite established — and that what you actually found is more interesting than what you were looking for.
Every phase generates feedback signals that can redirect earlier phases:
1.creativity ←──────────────────────────────────────┐
↓ │
2.literature ←─────────────────────────────────┐ │
↓ │ │
3.datacapture ←──────────────────────────┐ │ │
↓ │ │ │
4.dataanalysis ←──────────────────┐ │ │ │
↓ │ │ │ │
5.manuscript ─────────────────────┴──────┴─────┴────┘
The most powerful and most unexpected feedback loop is from manuscript writing back to everything — including back to the research question. Writing is the phase where all other phases become simultaneously visible and accountable to each other. What you cannot write clearly, you do not yet know clearly.
What AI adds to iterative research: Faster processing between phases tightens the feedback loops. When you can process a day's archival finds the same evening, the discovery that redirects tomorrow's search arrives earlier — when redirection is cheap — rather than after months of collection in the wrong direction. AI does not just make research faster; it enables more agile research, more responsive to its own discoveries.
The lifecycle, mapped
Project setup and organisation — Engineering-like. Writing CLAUDE.md, setting up folder structure, drafting timelines, organising notes. AI is highly productive here and the risks are low. This is also where the investment in longitudinal setup pays off most: a well-written CLAUDE.md becomes the context that makes every subsequent session better.
Literature finding — Engineering-like in principle, fragile in practice. Training cutoff, no paywall access, hallucination risk, and field-specific thinness make Claude unreliable for finding current and specialist literature. Use it for initial orientation in an unfamiliar field; use database searches (Scopus, Web of Science, JSTOR) for the actual work. → Detail in B.lifecycle.2.literature.
Literature synthesizing — Engineering-like, and genuinely productive once you have the sources. Claude can identify recurring themes, map tensions, surface connections, and draft structured summaries from material you provide. The limit: synthesis is not interpretation. Claude can tell you what the literature says; it cannot tell you what it means for your specific question. → Detail in B.lifecycle.2.literature.
Research question formation — Interpretive. The first fragile moment. The gap your research addresses is not in the training data; Claude's picture of a field is the centroid, not the frontier. The productive role for AI here is Socratic — stress-testing a question you have already formed, not generating one. → Full treatment in B.lifecycle.1.creativity.
Research design — Mixed. Methodology mechanics (standard methods for your data type, how comparable studies have been designed) lean engineering-like and Claude handles them well. Whether the design actually fits your specific question is interpretive. A useful pattern: describe your design and ask where its assumptions are most exposed, then evaluate Claude's response critically.
Data collection and empirical work — Mostly human; AI in a support role for preparation and post-collection processing. The micro-discovery loop — processing incoming material close to collection so that findings redirect further collection — is where AI most changes the character of this phase, enabling more agile research rather than just faster research. → Detail in B.lifecycle.3.datacapture.
Analysis — Data processing is engineering-like; interpretation is not. The execution vs. interpretation boundary is the organising principle: AI executes; you interpret. Qualitative coding has specific reliability, validity, and disclosure complications. Analysis generates feedback signals — the category that does not fit, the finding that contradicts the hypothesis — that return the project to earlier phases. → Detail in B.lifecycle.4.dataanalysis.
Manuscript writing — Mechanics are engineering-like; argument-building is interpretive. But more fundamentally: writing is where research gets its final shape. You do not write what you think — you think through writing. Every part of the manuscript that resists being written is a diagnostic signal pointing back to something unresolved. The manuscript is also connected to research question formation (phase 1) through the practice of proposal writing — even internal proposals for your own use — as a tool for clarifying thinking before and during the project. → Detail in B.lifecycle.5.manuscript.
Revision and peer review response — Mixed. Parsing reviewer comments, identifying patterns, drafting response letters — engineering-like and productive. Deciding whether a reviewer is right and how your argument should actually change — interpretive. Claude can help you explore options; it cannot make the judgment call.
The two fragile moments
Most lifecycle phases are fragile in the sense that interpretation stays human and AI's contribution is limited. These two are fragile in a different and more dangerous sense: AI actively misleads, the misleading is hard to detect, and the damage compounds into every phase that follows.
Creativity (question formation): Claude's picture of any field is the statistical centre of published literature — the centroid, not the frontier. When asked to generate research questions, it produces the most-already-answered ones: plausible, well-formed, and pointing away from where genuine contribution lies. The researcher cannot easily see this, because the questions look reasonable. Everything downstream — design, data, argument — is built on the question. A centroid question that goes undetected until manuscript writing has cost months. Use Claude as a Socratic stress-tester after you have a direction, not as a question generator before you have one. Note that the research question is not fixed at the start — micro-discoveries during data capture, analysis, and manuscript writing all feed back to refine or redefine it. → B.lifecycle.1.creativity
Literature finding: Claude hallucinates citations confidently. It also misses recent, paywalled, and field-specialist literature structurally — not randomly, but as a function of training cutoff and access. The researcher cannot know what they did not find. The synthesis, the argument, and the claim about the state of the field all follow from what was found. Undetected gaps and fabricated sources compound into every subsequent phase. Distinguish finding from synthesizing: Claude cannot reliably do the former; it can do the latter well, once you have the sources. Verification of citations is non-negotiable. → B.lifecycle.2.literature
The other phases have their own caution zones — but the risk is different in kind: the researcher knows they are interpreting, or the fragility is visible (if you cannot write it, you know you cannot write it). Brief pointers:
-
Data capture — The micro-discovery loop: process incoming material close to collection so findings redirect further collection. → B.lifecycle.3.datacapture
-
Analysis — Execution vs. interpretation is the line. Qualitative coding has specific complications. → B.lifecycle.4.dataanalysis
-
Manuscript — Writing is thinking. The manuscript is the most powerful feedback loop in the lifecycle. Proposal writing connects this phase back to phase 1. → B.lifecycle.5.manuscript
Managing AI across the full project arc
The practical challenge of beginning-to-end AI use is continuity. A research project runs for months or years; individual Claude sessions are ephemeral. The infrastructure that bridges them:
CLAUDE.md as project memory. A maintained CLAUDE.md describing your research question, key sources, methodological commitments, and current state allows each new session to begin with context rather than explanation. Update it when the project's direction changes.
A project history file. Tracking decisions, pivots, and reasoning (see A.markdown.history) means you can orient a new session to where you actually are, not where you were when you last wrote the CLAUDE.md.
The session as a unit. Sessions work best when focused on a specific subproblem — literature synthesis for section 2, argument structure for the introduction, a data processing step. The more scoped the session, the more engineering-like it becomes, and the more productive AI involvement is.
The full project arc belongs to you. The sessions within it can be collaborative. That is the beginning-to-end claim: not that AI manages the project, but that with the right infrastructure, AI can be a genuine partner in every phase — with involvement calibrated to the engineering-like or interpretive character of the current task.

Related
-
B.lifecycle.0.iteration — the foundational principle: fail quickly and let it be documented; the cost curve of failure; scaffolding the micro-discovery loops; the pre-mortem; documentation infrastructure
-
B.lifecycle.1.creativity — research question formation: the centroid problem, the Socratic reframe, good practice
-
B.lifecycle.2.literature — literature engagement: finding vs. synthesizing, Zotero, verification
-
B.lifecycle.3.datacapture — data collection and processing: the micro-discovery loop, what stays human
-
B.lifecycle.4.dataanalysis — analysis: execution vs. interpretation, qualitative coding, feedback signals
-
B.lifecycle.5.manuscript — manuscript writing as thinking: diagnostics, proposal writing, the 5↔1 loop
-
B.autonomy — agenda drift as the long-term risk of getting the lifecycle calibration wrong
-
B.epistemics — sycophancy, confirmation loops, and tunnel effect: the epistemic risks most dangerous in phases 1–2; counter-practices including cross-model checking
-
B.centroid-periphery — the center-periphery navigation model: when AI is genuinely useful at the center, when it misleads at the frontier, and how to navigate deliberately between them
-
B.usecases — specific scenario sketches; the lifecycle view is the frame these cases sit within
-
A.critical.limitations — execution vs. interpretation; hallucination in citations
-
A9.markdown-project-memory — CLAUDE.md as project memory; the longitudinal setup
-
A.markdown.history — project history file as continuity mechanism
-
A.markdown.meta-docs — the full meta-document system: _history, _ownership, progress, _index; the infrastructure that makes the full project arc manageable
-
A.issue.bounding — scoping individual sessions: making each phase tractable