Deep research tools, compared honestly: the report is where your work starts
Deep research agents from ChatGPT, Perplexity, and Gemini will read dozens of web sources and hand you a structured, cited report in minutes. They differ in speed, depth, and citation reliability, and they share one property the marketing skips: the report is a draft of the evidence, not the evidence. In the audit that measured these systems' ancestors, even the best cited sources incorrectly 37% of the time. Use the agents for the sweep. The judgment, and the checking, still belong to you.
Updated

Twenty minutes later, a beautiful report
You typed one paragraph, went for coffee, and came back to twelve pages: headings, findings, a confident executive summary, and footnotes all the way down. It looks like the output of a diligent analyst who worked through the night.
The feeling that report produces is the thing to be careful with. Length and structure read as diligence, and footnotes read as proof. But every claim in those twelve pages was assembled by the same kind of model that writes everything else, from sources it selected, read, and summarized without you. The report is not wrong because of that. It is unverified because of that, and those are different states.
How the big three differ
As of early 2026, the practical differences are speed, depth, and ecosystem. Perplexity's deep research is the fastest, returning reports in a few minutes with citations on every claim, and its underlying engine had the best citation record in Columbia's Tow Center audit, at 37% incorrect on news sourcing. ChatGPT's deep research runs much longer per task, produces the longest and most structured reports, and rations runs monthly depending on plan. Gemini's leans on Google's index and scholarly integration.
Pick by the job: quick landscape scans favor the fast one, single deep dives favor the long-running one, academic sweeps favor the scholarly one. And notice what did not vary: every one of them writes fluent, confident prose whose reliability you cannot judge from the prose itself.
Where deep research quietly fails
Three failure modes recur. Source quality laundering: the agent reads what the open web offers, and a well-formatted content-farm page weighs more than it should. Synthesis drift: ten accurate sources summarized into one paragraph can still yield a claim none of them makes, the same mischaracterization pattern Stanford measured in grounded legal tools at 17% to 33%. And coverage illusion: a report that reads exhaustive covered what the agent happened to retrieve, not what exists.
None of these show up in the report's tone. The twelve pages read identically whether the pipeline worked or failed, which by now you will recognize as the recurring theme of this whole series.
The workflow: sweep with agents, verify at the passage, decide yourself
The honest setup treats a deep research report as a map of candidate evidence. First, let the agent do the sweep; breadth is what it is genuinely good at. Second, extract the claims that will actually carry weight in your decision, usually three to five, and verify each at the passage level: open the cited source, compare the claim, note what it does not prove. Third, discard or flag everything consequential you could not verify.
For the claims that survive, you now have something better than a report: a small set of checked facts with open sources behind them. That is the material decisions deserve, and it took the 90-second ritual per claim, not another twenty minutes of agent time.
Your own documents deserve better than the open web treatment
One boundary matters more than any tool choice: deep research agents are built for the open web, and feeding them your contracts, client files, or internal reports means researching your confidential material with a tool shaped for public sources and public retention policies.
Questions about your own corpus belong in a grounded library: private by default, answering only from your documents and curated collections, every claim linked to its exact passage. That is Tatsulok's shape, and it means the verification step that costs real effort on a web report costs one click on your own material. Use the web agents for the world. Use a grounded library for what is yours.
The drill this week, no signup needed: take one deep research report you have received, pick its three most consequential claims, and run the passage check on each. Count how many survive intact. That number, not the page count, is the report's real value.
FAQ
- Which deep research tool is best?
- By job: Perplexity's is fastest with per-claim citations, ChatGPT's produces the deepest and most structured reports on longer runs, Gemini's integrates Google's scholarly index. All three produce confident prose whose reliability must be checked at the source, so the verification workflow matters more than the pick.
- Are deep research reports accurate?
- They are drafts of evidence, not verified evidence. The Tow Center audit found even the best engine citing sources incorrectly 37% of the time, and synthesis can produce claims none of the underlying sources makes. Verify the consequential claims at the passage level before relying on them.
- How do I verify a deep research report efficiently?
- Do not verify all of it. Extract the three to five claims that will actually carry your decision, open each cited source, compare claim to passage, and note what the passage does not prove. Flag or discard consequential claims that fail the check.
- Can I use deep research on my own documents?
- The web agents are shaped for public sources, public retention, and open-web retrieval. Confidential material belongs in a grounded, private library where answers come only from your corpus with passage-level citations, which is the architecture Tatsulok provides.
- Why do cited reports still contain wrong claims?
- Because citation and synthesis are separate steps: an agent can read accurate sources and still summarize them into a claim none of them makes, and a resolving link does not prove the passage supports the sentence. Stanford measured this mischaracterization pattern at 17% to 33% even in grounded legal tools.
Sources
Related guides

Prompts that force citations
Five prompt elements that force AI to show its evidence: scope the sources, demand a citation per claim, allow honest gaps, set a cutoff, fix the output shape. With a template you can copy today.
Read the article
Context rot
Context rot is the accuracy drop language models suffer as input grows, long before the window fills. At 32,000 tokens, 11 of 13 models scored under half.
Read the article
Chat with your documents
Chatting with your documents means uploading your files and asking questions in plain language. Tatsulok answers with citations to the exact source, privately.
Read the article