Critical Limitations: What Goes Wrong with Claude in Research

Why this document exists: A workshop that only shows the possibilities will lose credibility with skeptical researchers — and, more importantly, will send participants into their work without the critical awareness they need. This document covers documented failure modes, real criticism from academics, and practical mitigation.


The core distinction: execution vs. interpretation

The clearest framing comes from Tom Pepinsky (Cornell, comparative politics):

"Use agentic AI for tasks that involve following rules. Do not use agentic AI for tasks that generate answers, arguments, or interpretations." — https://tompepinsky.com/2026/01/23/agentic-ai-and-social-science-research-practice/

What Claude Code is good at (following rules):

What Claude Code is unreliable at (generating answers, arguments, interpretations):

The boundary is not always clean — but when you catch yourself asking Claude to judge rather than execute, apply more skepticism.


Hallucination: the numbers

Claude confidently states false things. This is not a bug that will be fixed — it is a structural property of how large language models work.

Statistics (2025–2026 benchmarks):

Source: https://misinforeview.hks.harvard.edu/article/new-sources-of-inaccuracy-a-conceptual-framework-for-studying-ai-hallucinations/

The most dangerous pattern: AI models are 34% more likely to use high-confidence language ("definitely," "certainly," "it is clear that") when generating incorrect information. The more wrong, the more certain it sounds.

Practical rule: Never use a Claude-generated citation without independently verifying it. Never paste a Claude-written literature review into a manuscript without checking every claim against the source. Treat Claude-generated factual claims as first drafts to be verified, not outputs to be trusted.


"Vibe research": the p-hacking problem

Andy Hall (Stanford political scientist) coined the term "vibe research" for a specific failure mode:

AI produces confident results aligned with user expectations rather than rigorous methodology.

Source: https://www.niskanencenter.org/can-ai-vibe-research-replace-social-science/

Concrete examples from Hall's testing:

Pepinsky's version of the same concern:

"The computer might generate the result that it 'thinks' you want to hear."

This is the AI equivalent of p-hacking: if you ask Claude "does the data support hypothesis X?" it may structure its analysis to confirm X rather than rigorously test it. It is trying to be helpful. Helpfulness and rigor are not the same thing.

Mitigation: Ask Claude to argue against your hypothesis, not just for it. Ask it to identify weaknesses in its own output. Ask: "what would a critic say about this analysis?"


Documented failure modes in humanities and social science contexts

From Hall's testing (political science):

From Pollin's DH work:

From Messing & Tucker (Brookings):


The qualitative research caution

For researchers doing interview-based or ethnographic work:

From peer-reviewed literature on AI-assisted thematic analysis:

Sources: Xu (2026), Qualitative Inquiry; Ozuem et al. (2025), Sage; https://journals.sagepub.com/doi/10.1177/16094069261425173

Data privacy: Before uploading interview transcripts to any cloud-based AI service (including Claude Desktop), check your IRB approval and institutional policy. Claude Desktop sends data to Anthropic's servers. Claude Code runs locally — transcripts never leave your machine.


The systemic concerns (beyond individual use)

These are not reasons to avoid Claude, but researchers should be aware of them:

Junior researcher deskilling: Tasks that used to provide learning opportunities (literature searches, data cleaning, formatting) are now automated. Junior researchers may complete PhD programs without developing these skills. The implications for long-term disciplinary expertise are unknown.

Resource inequality: Heavy API use favors well-resourced institutions. A researcher at a well-funded institution with API credits has access to capabilities unavailable to colleagues at less-funded institutions.

Peer review strain: The projected increase in manuscript submissions (from faster paper production) is already straining review systems. Some journals are already receiving AI-generated manuscripts that reviewers struggle to detect.


Practical critical literacy checklist

Before trusting Claude output in your research:


The balanced summary

Claude Code and Claude Desktop are genuinely powerful for research. The documented productivity gains are real. The non-programmer entry point is real. The capabilities for document processing, systematic extraction, and writing assistance are real.

And: hallucination is real. Sycophancy is real. The execution/interpretation divide is real. Epistemological fragility is real.

The researchers getting the most value from these tools are not the ones using it most — they are the ones using it most critically: clear about what they are delegating, rigorous about verification, and honest in their methods sections about what role AI played.


Related