Every AI tool will confidently hand you a citation. The problem is that some of those citations don’t exist. A real-sounding article title, a plausible author, a URL that leads nowhere or somewhere else entirely. This is the biggest risk in using AI for research, and most comparisons skip it in favour of talking about writing quality.
Writing quality is easy to judge by reading. Citation accuracy takes verification work, which is exactly why it gets skipped and exactly why it matters more.
If you’re picking an assistant on general capability, our ChatGPT vs Claude vs Gemini comparison covers that broader question. This one is narrower and, for anything you’ll put your name on, more important.

Claude, one of the four tools covered here. A citation-shaped sentence reads the same whether it’s accurate or not.
How we compared
Two published benchmarks do the heavy lifting in this article, and they’re named inline wherever their numbers appear: Vectara’s hallucination leaderboard and Stanford HAI’s 2026 AI Index. Both publish their methodology. They also disagree with each other by a factor of more than ten, which turns out to be the most useful thing about them.
Pricing comes from each provider’s own pages, checked in July 2026. Screenshots are our own captures of publicly reachable pages.
We haven’t run a private fabrication test. Doing that properly means hundreds of prompts scored by more than one person, which is what the benchmarks below already do at a scale no article could match in a week.
The benchmarks disagree, and that’s the finding
Ask “how often does AI make things up” and you get two wildly different answers depending on what you measure.
Vectara’s hallucination leaderboard scores models on whether a summary stays faithful to a document it was given. It uses an evaluation model called HHEM against a private set of more than 7,700 articles spanning news, medicine, law, finance, and education. On that test, the best rows sit at roughly 2% to 3% hallucination. Vectara is careful about what this means, stating plainly that they are “not evaluating the quality of the summaries, only the factual consistency of them.” A model that just copied sentences out of the source would score perfectly and be useless.
Stanford HAI’s 2026 AI Index measures something closer to open recall, and the picture inverts. Across 26 leading models, the report puts hallucination rates between 22% and 94%. It documents GPT-4o’s accuracy dropping from 98.2% to 64.4% on the harder framing, and DeepSeek R1’s falling from over 90% to 14.4%.
Both numbers are real. They describe different jobs. Hand a model a document and ask for a summary, and fabrication is rare. Ask it to recall a fact it wasn’t given, with a source, and fabrication becomes common. That gap is the single most useful thing to know before trusting an AI citation, and it explains why “Deep Research” modes that browse before answering are meaningfully safer than a model answering from memory.
The Stanford finding with the sharpest practical edge is about framing. The report found that models handled a false claim reasonably when it was attributed to a third party, but that “when the same false statement is presented as something a user believes, performance collapses.” Phrase your question as “I read that X, can you confirm?” and you make agreement more likely, whether or not X is true. Ask neutrally instead.
For a broader look at how these same four tools stack up outside the research question, see our review of best AI for coding, which applies the same documented-evidence approach to a different job.
Why citations break even when the writing looks solid
These models generate text by predicting what a plausible answer looks like, not by looking up a fact and reporting it verbatim. Even with live web search, the model is writing a summary of what it found, and summaries drift. A citation-shaped sentence with a real-looking URL is exactly as easy to produce as an accurate one.
That’s why a confident tone tells you nothing. The failure is quiet, too. A wrong pasta recipe announces itself at dinner. A wrong citation in a research summary looks identical to a right one until somebody clicks through, which in practice is often nobody.

Gemini’s app. Its Deep Research mode browses the live web before answering, which per the benchmark gap above should help.
Perplexity is the one built for this job
Perplexity is the outlier worth understanding. It was designed from the start as a search-and-cite tool rather than a chat assistant with research bolted on, and it shows sources inline by default rather than when asked.

Perplexity’s Apple App Store listing, showing the vendor’s own published app screenshots. We used the store listing because perplexity.ai blocks automated capture.
Inline sourcing doesn’t make the underlying model more truthful. It makes checking faster, which in practice matters more. A tool that shows you four links you can click in ten seconds gets verified. A tool that buries its reasoning in a wall of prose does not.
What “good” actually looks like here
The best outcome isn’t a tool that never gets anything wrong. None of them manage that. It’s a tool that tells you when it isn’t sure.
Watch for hedging that tracks reality: “I don’t have a reliable source for this” on a genuinely obscure question, versus a flat, unhedged claim on something the model couldn’t plausibly know. A tool that hedges appropriately beats one that sounds more polished and is silently wrong at the same rate.
This is also the hardest quality to judge from one response, and it’s where the Stanford framing result is worth remembering. Confidence responds to how you asked, not just to what the model knows.
How to verify a citation in under a minute
You don’t need a research background to catch most fabrications.
- Open the link. A surprising share of fabricated citations use a URL that 404s or redirects to a homepage instead of the specific article.
- Check the byline and date. Real articles have both, consistently. Fabricated ones often have neither, or ones that don’t survive a quick search.
- Search the exact quoted phrase. Paste it in quotation marks. Nothing matching means treat the quote as unverified.
- Confirm the claim, not just the link. A real, working link can still be attached to a claim the source never made. This is the step people skip, and it’s the one that catches the errors that matter.
What each tool costs
| Tool | Free research features | Paid plan | Price |
|---|---|---|---|
| ChatGPT | Limited Deep Research | Plus | $20/mo |
| Claude | Web search on the free tier | Pro (adds Research) | $20/mo, or $17/mo annually |
| Gemini | Capped Deep Research reports | Google AI Pro | $19.99/mo |
| Perplexity | Limited daily Pro searches | Pro | $20/mo, or $200/year |
Perplexity Pro at $200 a year is the same effective rate as paying monthly, unlike Claude’s annual discount. Worth knowing before you commit twelve months up front for nothing in return.
Worth it for
- Tools with live web search cite real, checkable sources far more often than models answering from memory, which the benchmark gap above supports directly
- Perplexity shows sources inline by default, which makes verification fast enough that people actually do it
- All four will give you sources if you ask, which at minimum gives you something concrete to check
Skip it if
- No tool here is reliable enough to skip verification, and confident wording is not evidence of accuracy
- Fabrication rates vary enormously by task: roughly 2% on grounded summarisation, 22% to 94% on open recall
- How you phrase a question changes the answer, since presenting a false claim as your own belief makes agreement more likely
One habit worth keeping whichever tool you use: never paste an AI’s summary of a source into something you’re publishing or submitting without opening the original yourself. It costs a minute and it’s the most effective single defence against repeating a fabrication.
Alternatives worth a look
- Google Scholar or your library’s database. For academic or scientific claims, going straight to a primary-source search skips the summarisation step where drift happens.
- Wikipedia, read for its citations rather than its prose. Often faster than an AI chat for a quick check, since the sources are already listed and have survived some public scrutiny.
- A second AI tool as a cross-check. Running the same question through two tools catches fabrications one alone would miss. Agreement plus a shared real source raises your confidence; disagreement is your signal to verify by hand.
The verdict
WAIT. None of these four should be trusted to hand you a citation you can use unchecked. The published evidence is clear that fabrication is rare when a model summarises a document you gave it and common when it recalls facts on its own.
So use the shape of the task to decide how much to worry. Feeding a tool your own documents is relatively safe. Asking it what the research says is not, and needs every source opened. Use whichever tool’s search mode fits your budget, and build verification into the process every time rather than only when something looks off.
If you’re deciding whether a paid research tier is worth anything at all, our ChatGPT free vs Plus breakdown and the general AI subscription guide both cover the budget question. For how these four compare outside research, see ChatGPT vs Claude vs Gemini.
Common questions
Which AI doesn’t make up sources?
None of them are immune, but the risk depends heavily on the task. Vectara’s hallucination leaderboard puts the best models at roughly 2% to 3% fabrication when summarizing a document you gave them. Stanford HAI’s 2026 AI Index found rates between 22% and 94% when a model recalls a fact from memory and attaches a source itself. Perplexity is built specifically to cite sources inline, which speeds up verification, but it doesn’t make the underlying model more truthful.
Is ChatGPT Deep Research worth it vs Perplexity?
Both use live web search, which the benchmark gap above suggests is meaningfully safer than a model answering from memory alone. ChatGPT’s Deep Research is limited on the free tier and fuller on Plus at $20 a month. Perplexity was designed around search-and-cite from day one and shows sources inline by default rather than only when asked, which makes it the faster tool to actually verify.
Is Perplexity Pro worth it for research?
At $20 a month or $200 a year, Perplexity Pro raises your daily Pro-search allowance and adds Deep Research. It earns its cost if you’re checking sourced facts often enough to hit the free tier’s small daily allowance; if your research is occasional, the free tier’s unlimited basic search already covers it. Our full Perplexity Pro breakdown covers the model picker and exact pricing in more detail.