DO THE RESEARCH YOURSELF OR DELEGATE IT.
Deep research agents promise to turn a question into a referenced report in twenty minutes. They work, and sometimes surprisingly well. But getting oriented is not the same as publishing, and that difference still hinges on something no agent comes with built in: information expertise.
IN THIS ARTICLE
By María García-Puente · AI · July 2026
Two ways to answer the same question
Picture a third-year resident whose attending has asked her to review the evidence on an intervention. Two paths lie ahead of her.
The first is the familiar one: frame the question in PICO format, design a search strategy with its descriptors and synonyms, run it in PubMed and Embase, screen titles and abstracts, read the full texts, extract the data, and synthesize. Two weeks of work if it goes smoothly, with every decision documented and reproducible.
The second path is barely two years old: open a research agent (Claude Science, OpenAI's Deep Research, Gemini's version), type the question in plain language, and wait twenty minutes for it to come back with a structured report, complete with sections and references. It looks like the same task, finished seventy times faster.
The temptation to conclude that the first path is now obsolete is understandable, and it's also a mistake. But the opposite mistake, dismissing agents as tools that hallucinate citations, shouldn't be made either. We use these tools daily, and we've learned, through both wins and scares, that the useful question isn't which of the two paths wins, but what phase of the work you're in and how much rigor the outcome demands. That's what this article is about.
It's worth naming the two approaches before going further. We'll call approach A building the research step by step, using AI as a one-off assistant at each stage (drafting the search string, suggesting synonyms, helping with screening) while the researcher stays in charge of every decision. And approach B the end-to-end agent: you hand it the question and it returns the report, without explicit stage-by-stage supervision.
The map as of July 2026: what's out there, and how mature it is
The landscape shifts every few months, so it's worth dating the snapshot: this is what exists as of July 2026, distinguishing what you can use today from what is still an announcement or a research project.
Claude Science (Anthropic) is the most recent arrival: a beta launched on June 30, 2026, available to paying Claude users. It isn't a new model but a workbench that brings together, in one place, more than 60 skills and connectors to scientific databases. When you configure it, you choose which sources it can query, grouped by category: NCBI/NIH (PubMed, Entrez), genomics and biology (Ensembl, Reactome, KEGG), proteomics (UniProt, STRING, PDB), literature and citations (Semantic Scholar, arXiv, bioRxiv, Crossref, OpenAlex), and clinical and pharmacological data (ClinicalTrials, Open Targets, ClinGen), among others, with the ability to render protein structures or figures alongside the code that generates them. We tested it with a real search and can confirm two things that matter in our line of work. First, that PubMed is indeed among the native sources: we asked it to build and run the strategy for a systematic review, and it fired queries against NCBI's E-utilities, returning real counts step by step. Second, that the auditable trail isn't just a promise: every step is logged, and a reviewer agent later reproduced the counts one by one to confirm none had been invented. It's some of the most useful tooling we've seen for our work, with the usual caveat: a tool that documents and verifies its own steps doesn't exempt the information specialist from validating the strategy — it just makes that validation a lot easier.
OpenAI has two distinct offerings that are often confused. Deep Research is the actual agent, available since 2025: it searches, analyzes, and synthesizes hundreds of web sources to produce a report, and since February 2026 it can connect to external sources and restrict searches to trusted sites. "OpenAI for Science," by contrast, is an institutional initiative for scientific collaboration, with agreements such as the one with the U.S. Department of Energy — not a product you can open and use as an equivalent to Claude Science.
Google also has two pieces that shouldn't be conflated. Gemini Deep Research (and its 2026 evolution, Deep Research Max) is the functional equivalent of the above: an agent that searches, reads, and synthesizes existing literature. The AI Co-Scientist is something else entirely: a multi-agent system aimed at generating and debating new scientific hypotheses, published as research in Nature and accessible experimentally. It's interesting, but it belongs to the realm of ideation, not literature review.
Then there's the ecosystem of specialized academic tools, each built for a specific phase: Elicit searches across millions of papers and extracts data into structured tables (PICO, methods, results); Consensus measures the degree of agreement between studies and, in 2026, launched a medical mode focused on clinical guidelines and top-tier journals; Scite classifies each citation according to whether it supports, contradicts, or merely mentions the cited work, and keeps traditional Boolean search, a detail anyone who needs to reproduce a strategy will appreciate; Perplexity stands out for source traceability in general web search; and FutureHouse, a nonprofit lab, has taken the agentic approach all the way to the finish line: its Robin system identified an existing drug as a candidate for a new indication, going from idea to manuscript in two and a half months, with publication in Nature in May 2026. The coverage figures these tools advertise (138 million papers here, 220 million there) are vendor numbers; the Robin case, by contrast, went through peer review.
And there are many more tools than fit here (Undermind, ResearchRabbit, or Paperpal, to name three). You don't need to know them all, but you should know they exist and be able to place each one in the phase of the process it fits, along with its price and terms of use.
Where the end-to-end agent shines
We would be unfair, and not very credible, if we painted approach B as a trap. There are at least four situations where it works very well.
The first is orientation in unfamiliar territory. When you're assigned a topic you barely know anything about, a deep research report gives you, in half an hour, the map that used to take days to build: the key concepts, the authors who keep coming up, the open controversies. That report doesn't replace the review, but it works as the compass to start it with.
The second is generating search strategies, and here we can speak from our own experience. When we asked Claude Science for a strategy for a systematic review, it returned a search string that was more sensitive and comprehensive than the ones our own working templates produced: more synonyms, better handling of acronyms within the Boolean intersection. It surprised us, and we say so without reservation. With a caveat that's half the story: that was a draft strategy we then validated line by line before running it. The tool widened our starting point; the decision of what went in and what didn't remained ours.
The third is fast reading of a large corpus: once you already have the documents located and verified, asking an agent to go through them looking for a specific data point or comparing methodologies saves hours of skim reading.
And the fourth, early syntheses for internal use: a draft state-of-the-art overview to discuss with the team, one nobody is going to cite or publish, is a use case where speed more than makes up for the residual risk of error.
The common pattern across all four is that the agent's output is a starting point someone reviews, not a final product someone signs off on.
Where it falls short, with numbers
This is where enthusiasm needs data, and by 2026 we have it. Two recent studies, both peer-reviewed, put numbers to the two failures that matter most in our field.
The first is fabricated citations. A research letter published in The Lancet on May 7, 2026, by Maxim Topaz's team at the Columbia School of Nursing, analyzed roughly 2.5 million PubMed Central papers as part of the CITADEL project, looking for references that don't exist. The result traces a rising curve: in 2023, one in every 2,828 papers contained invented references; in 2025, one in every 458; and in the first weeks of 2026, one in every 277. The jump occurs in mid-2024, right as the use of generative chatbots for drafting became widespread. The authors attribute the phenomenon to a combination of AI-assisted writing without verification and publish-or-perish pressure. In other words, hallucinated citations aren't an anecdote about a careless user, but a measurable, growing problem in published biomedical literature — literature that already passed editorial screening and peer review.
The second figure concerns search sensitivity, the metric that decides whether a systematic review is worth anything. An independent evaluation published by Lau and Golder in Cochrane Evidence Synthesis and Methods (2025) compared Elicit against the traditional searches of four real systematic reviews. AI-assisted search achieved an average sensitivity of ~38% (ranging from 25.5% to 69.2%), against ~94% for the original traditional strategies. Put plainly: the tool left out more than half of the studies that should have been included in the review. Its precision was higher — it returned less noise. But in a systematic review, precision doesn't make up for sensitivity: an eligible study that doesn't turn up means a biased review.
This case also holds a lesson on how to read vendor figures. Elicit reports a 96.9% sensitivity on its own blog, and it isn't lying — it's measuring something else. That 96.9% evaluates the screening of abstracts already retrieved (deciding whether a document already in front of you is relevant), while the independent evaluation's ~38% measures the full search: finding all the studies that exist. Screening well what you've already found is not the same as finding all of it, and these tools' performance changes radically depending on which phase of the process you measure. Whenever you see a spectacular percentage on a product's website, the first question is always: exactly what task is it measuring?
One more caveat about these figures: both studies were done with the models available a few months earlier, and the pace at which these change is fast. Just in these first weeks of July 2026, Claude Fable 5 (within Claude) and OpenAI's GPT-5.6 family (Sol, Terra, and Luna, within ChatGPT) have become generally available. It's reasonable to think part of these results would shift if the evaluations were repeated today, so they're best read as a snapshot of a specific moment rather than a constant. What doesn't change with the model version is the underlying framework: exactly what each evaluation measures and which phase of the process it puts to the test.
Added to these two data points is a more structural problem: reproducibility. A Boolean search in PubMed can be written into the review's annex and anyone can rerun it; an agent's retrieval (based on semantic similarity, an opaque ranking and, in configurable environments, the specific tools each user happens to have connected) can produce different results across two runs, or between two users asking the same question. Specialist librarians such as Aaron Tay have flagged this as a structural problem, not an anecdotal one, for generic agents compared with fixed-workflow tools.
The blind spot: evidence quality
There's a gap no agent on the market closes today, and in health sciences it's the one that matters most: assessing the quality of the evidence. A deep research report can cite, in the same paragraph, a randomized controlled trial with thousands of patients and a case series with twelve, with nothing in the text warning you they don't carry the same weight. The hierarchy of study designs, each study's risk of bias, authors' conflicts of interest, grading certainty with systems like GRADE: all of that remains human analysis, with or without AI in the loop. It's also reasonable to think that agents drawing on the open web tend to see what's open-access and indexed more clearly than what sits behind a paywall, though, as far as we know, no one has yet measured that bias with a dedicated study in the biomedical domain, so we leave it as a reasonable hypothesis, not a fact.
The methodological community has taken a position on all of this. Cochrane's Rapid Reviews Methods Group published a 2025 position statement that explicitly recommends against using AI to fully automate a review or any of its methodological steps, because doing so risks introducing errors, bias, and a lack of transparency; sustained human oversight, they say, must remain the central principle of any AI-assisted evidence synthesis. In the same vein, a 2026 viewpoint published in JMIR reviewed deep research agents applied to medicine and concluded that they represent an incremental evolution, not a paradigm shift: they produce reports that look exhaustive and well-referenced, yet coexist with unresolved, clinically significant limitations. When the two most methodologically grounded voices in the field agree on the same nuance — useful, yes; autonomous, no — that should guide you more than any demo.
IN ONE SENTENCE
A deep research report is a starting point someone reviews, not a final product someone signs off on.
What's still your job
If agents already search and summarize quickly, what's left for the researcher or the information specialist to do? Quite a bit more than it seems, and, as it happens, the hardest part to automate.
You still have to frame the question well: an agent answers whatever you ask it, and a poorly scoped question produces a flawless report on the wrong problem. The PICO format hasn't lost its relevance; if anything, it's the best vaccine against generic reports. You can lean on a conversational AI to sharpen the wording, but that only works if you already know what you want to look for: the tool tidies up your question, it doesn't guess it for you. And, beyond being well framed, the question has to be useful — meaning its answer changes something you're about to decide, write, or recommend.
You still have to curate the corpus: decide which sources should be there, spot which ones are missing, and know which databases the tool you're using covers and which it doesn't. The agent doesn't know what it hasn't seen, and it won't tell you.
You still have to verify every citation, one by one, before any text leaves your computer for a journal, a report, or a guideline. After the Lancet finding, this has stopped being advice for perfectionists and become basic publishing hygiene. AI helps here too, if you know how to ask: you can have it cross-check each reference against PubMed, Crossref, or arXiv and flag the ones that don't add up. But the instruction has to be explicit, because checking isn't the same as drafting, and the agent that wrote the text won't review its own citations unless you tell it to.
And you still have to know what phase you're in, which is the judgment call that governs everything above: exploring isn't the same as screening, screening isn't the same as synthesizing, and synthesizing for an internal meeting isn't the same as synthesizing for a systematic review aiming to meet PRISMA. Each phase allows for a different level of delegation, and calls for different tools — the same judgment we apply when deciding when it makes sense to use NotebookLM to work over a closed corpus instead of sending an agent out onto the open web.
This division of labor isn't a defeat for AI, nor a guild's turf claim. It's simply what the 2026 evidence shows: these tools dramatically accelerate the phases where a mistake is cheap, and they still need close supervision in the phases where a mistake gets published.
Neither yes nor no: know what phase you're in
Let's go back to the resident from the beginning. Which path should she take? It depends on what for. If she needs to get oriented on the topic before Thursday's session, the deep research agent is probably the best tool that has ever existed for that, and denying her that in the name of rigor would be as absurd as banning the calculator. If what she's working on will end up in a manuscript, a guideline, or a clinical decision, the agent's report can be the scaffolding, but the building itself goes up with the disciplined method: a documented strategy, a reproducible search, screening with explicit criteria, verified citations, and evidence quality assessed by someone who knows how to do it.
The question worth asking in 2026 isn't "AI, yes or no?", which by now is an empty question, but three smaller, more useful questions to ask of each specific task.
THREE QUESTIONS
- What phase of the work am I in?
- What level of rigor does what I'm going to do with the result demand?
- What can I verify myself before signing off on it?
Anyone who can answer those three questions well can put any tool on this list to good use. Anyone who can't will get extremely fast, beautifully formatted reports that shouldn't support anything important. Learning to tell one case from the other has become as basic a skill as searching PubMed, which is why we devote growing space to it in our training.
And because we believe this judgment should be within everyone's reach, we've built a free, open tool, "Should I use this AI?". We built it following what the methodological community has published, the RAISE framework (Responsible use of AI in evidence SynthEsis), and the joint position of Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence, and we're making it available to everyone.
Free, open tool
Should I use this AI?
A short questionnaire that, based on a set of warning signs, helps you decide whether it makes sense to use an AI tool at a specific phase of your review, and with what precautions.
Open the tool (in Spanish)Sources
- Anthropic: Claude Science, an AI workbench for scientists (June 2026). Launch coverage in TechCrunch, MIT Technology Review, and STAT News.
- OpenAI: Introducing deep research and MIT Technology Review — Inside OpenAI's big play for science (January 2026).
- Google: Deep Research Max and Co-Scientist (published in Nature, May 19, 2026).
- FutureHouse: Demonstrating end-to-end scientific discovery with Robin, with publication in Nature (May 2026).
- Topaz et al.: research letter on fabricated references, The Lancet (May 7, 2026). Coverage in STAT News and phys.org.
- Lau and Golder: Comparison of Elicit AI and traditional literature searching, Cochrane Evidence Synthesis and Methods (2025). Vendor figure for contrast: Elicit blog — Evaluating Elicit's SLR capabilities.
- Cochrane Rapid Reviews Methods Group: Responsible Integration of AI in Rapid Reviews, position statement (2025).
- Wong, Ong, Merle, and Keane: Deep Research Agents: Major Breakthrough or Incremental Progress for Medical AI?, JMIR (2026).
- Aaron Tay: The agentic researcher.
- Recency note (July 2026 releases): OpenAI — GPT-5.6 Sol, Terra, and Luna (general rollout on July 9, 2026) and Anthropic — Claude Fable 5 (general availability in July 2026).
- RAISE (Responsible use of AI in evidence SynthEsis), a framework from Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence (2025): joint position statement and Cochrane's explainer, "Right tool for the right job: deciding when NOT to use an AI tool" (EN · ES). Methodological basis for our tool "Should I use this AI?".
KEEP READING
MORE INSIGHTS
INSIGHT · GOVERNANCE
PEER REVIEW
How peer review works step by step, and what you can do as an author.
READ MORE