Ilustración editorial para ChatGPT Deep Research, Gemini Deep Research, Perplexity y Claude Research: cómo compararlos
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

What This Comparison Covers—and What It Cannot Conclude in Advance

A web research assistant is more than a search engine that returns a list of pages. You ask it to find information, organize it, and write a synthesis with references that a person can inspect. So an evaluation should not stop at confirming that the report contains links. You also need to check whether each source supports the claim beside it, whether relevant sources were found, and whether the report represents disagreements faithfully.

This guide proposes comparing ChatGPT Deep Research, Gemini Deep Research, Perplexity Pro Search, and Claude Research using equivalent assignments. It does not declare a winner or claim that one tool is more accurate, complete, or faster than another. The available provider documentation describes product features, but does not offer a shared comparative test of all four products. Moreover, the sources verified here do not include documentation for Claude Research that would confirm its options, availability, or operation. Its features should therefore be checked directly before an evaluation begins; they should not be inferred by analogy.

The distinction between an announced feature and a proven result is central. A product’s description of multistep research, web search, source selection, or links does not prove that it found the best sources for a particular query. To find out, you need to run the same tasks, preserve the test conditions, and manually review a sample of the main claims.

02

Design a Fair, Reproducible Test

A useful comparison starts with a record of the conditions, not an informal collection of screenshots. Note the date and time, the name and visible version of the feature used, the account type or plan, the controls enabled, the exact wording of each assignment, and any documents uploaded. Keep the full output too, including references, limit notices, and follow-up questions. Plans, quotas, and features change; without a record, a difference in results might be due to access rather than system behavior.

Use four assignments with different goals. For recent information, ask for a dated fact and a primary source. For a question with several parts, list each subquestion and ask the answer to indicate what evidence covers each one. For a restricted search, define the permitted domains in advance and check whether the final references comply. For work with user documents, use the same file in tools that allow uploads and require the report to distinguish documentary evidence from web evidence. If a feature is unavailable, record it as not evaluable under those conditions; do not substitute a different mode without saying so.

Using the same wording reduces differences in interpretation, though it does not eliminate them. It is useful to run a second round with equivalent phrasing to detect fragile results. Do not change the assignments midway through the test to favor a tool, and do not treat modes or plans as equivalent when they are not. When settings cannot be matched, describe that limitation rather than concealing it.

Minimum testing protocol

  1. 01Set four assignments: a recent fact, a multistage investigation, a domain-restricted search, and analysis of user documents.
  2. 02For each product, record the date, plan, feature, search settings, files, and visible limits.
  3. 03Enter the same assignment, save the complete response, and do not change settings during the test round.
  4. 04Check the central claims against the original pages and record citations that are correct, weak, irrelevant, or broken.
  5. 05Repeat important assignments and distinguish stable differences from variation between runs.
03

What to Measure in Sources, Citations, and Disagreements

Assess reference quality claim by claim. For each verifiable sentence in the report, identify the associated citation and open the original page. Check whether the page contains the information, whether it refers to the same population, period, and context, and whether the report overstates what the source allows you to conclude. A page may be relevant to the general topic yet fail to support the specific figure or causal relationship attributed to it by the assistant.

Record separately citations that open and are relevant, those that support only part of a claim, those that are not sufficiently related, and those that cannot be accessed. If a claim combines several facts, break it down before rating its support. An aggregate rate without this review can be misleading: many correct references for secondary details do not make up for an unsupported central claim.

Evidence diversity matters too. A primary source—such as a regulation, an original dataset, or a document from the institution responsible for a figure—does not always serve the same purpose as a journalistic analysis or an academic review. You do not need to require a primary source for every question, but you should identify the type of evidence used and whether independent sources are present. If two sources disagree, the report should explain what differs, why, and what cannot be resolved with the available evidence, rather than presenting one account as consensus.

Audit time completes the evaluation. Measure how long it takes a person to check the main claims, find the relevant passage, and decide whether the citation is sufficient. This is not an absolute measure of quality: it depends on the topic, source accessibility, and the reviewer’s experience. It is, however, a practical variable for teams that need to deliver verifiable reports against specific deadlines.

A scoring matrix for each report

CriterionWhat to look forRecommended record
CoverageWhether the response addresses every part of the assignment and identifies what remains unresolved.Parts covered, omitted, and not verifiable.
Citation supportWhether the original source directly supports the associated claim, with the correct scope and period.Direct, partial, insufficient, irrelevant, or inaccessible.
Quality and diversityPresence of relevant primary sources, independent sources, and conflicting perspectives.Source type, provenance, and disagreements identified.
TraceabilityHow easily a reader can open the source, find the passage, and understand which sentence it supports.Working link, findable passage, and ambiguities.
Review effortTime needed to check the most important findings.Minutes per report and difficulties encountered.
04

Four Tasks That Reveal Practical Differences

For a recent fact, the assignment should specify what counts as recent and what kind of support is needed. Ask for the date of the figure, the organization that publishes it, and a link to the original source. Then check whether the result confused the publication date with the measurement date, treated a provisional figure as final, or cited a news article that does not lead to the primary data. This test assesses recency and traceability, not general research ability.

For a multistage question, write a list of subquestions and require a synthesis that keeps the differences between them visible. A fluent response can conceal that one part rests on strong evidence while another relies on a weak source. Also check whether the conclusion follows from the sources or adds an inference that is not labeled as such. To make the audit easier, ask for each important point to have its own specific supporting reference.

For a search restricted to particular sites, first decide whether you want to allow only a list of domains or prioritize those domains without excluding others. These are different instructions. Then inspect the response’s sources: a system may offer a site-selection setting without that, by itself, guaranteeing that all final references come from those sites. OpenAI’s documentation describes a Deep Research search option restricted to trusted sites, and Gemini’s help documentation describes source controls and selection. These are claims about stated options, not verification that every run respects them.

For user documents, use a file without sensitive data that contains verifiable claims. Ask the assistant to explicitly separate information drawn from the document from information found on the web. Check whether citations lead to the file, web pages, or both, and whether the report attributes each piece of evidence correctly. Gemini’s official documentation describes file uploads and source selection as options; OpenAI also describes using uploaded files in Deep Research. Confirm availability and behavior in the account being tested.

Perplexity presents Pro Search as a mode with model selection and search modes, including Web, Academic, Finance, and Files, and says its answers link to original sources. These descriptions can help you design a test that records which mode was used. They do not replace checking the original pages or let you predict the proportion of correct citations. A valid comparison must record the mode, plan, and conditions, then evaluate each reference using the same criteria applied to the other products.

05

How to Interpret Features and Access Limits

Official pages are suitable for confirming what providers say about their products, but they are not independent evaluations. OpenAI describes Deep Research as a multistep research workflow with citations and distinguishes it from faster search. Its OpenAI Academy guide also explains that workflow. Google says Google Search is included by default as a source in Gemini Deep Research and documents source options, editing the research plan, and file uploads. Perplexity describes Pro Search, its modes, and links to original sources. These statements can help configure a test; they do not justify ranking the tools by accuracy.

Access limits should be part of the results. Record whether an account can use the feature, which plan the provider requires, what quotas are shown, and whether file uploads or source controls are available. Do not generalize the results from one account to all users. Plan pages can change, and a quota shown today should not be presented as permanent. Recheck the conditions immediately before publication and state the date of verification.

The academic evidence among the supplied sources also has limited scope. An arXiv study addresses source credibility and response support in ChatGPT and Perplexity, but does not include Claude or Gemini. Its results apply to the products and conditions studied in that paper; they should not automatically be carried over to current versions or used to proclaim a winner in this comparison. Earlier research can help formulate criteria, but testing current products requires original data and an explicit method.

In particular, the absence of a verified source about Claude Research means its capabilities or availability cannot be described here as facts. The tool can be added to the test if its official documentation and current access are confirmed. Until then, any claims about its controls, sources, or citations remain unverified. This caution avoids filling gaps with assumptions.

Decision matrix by task

NeedWhat to prioritizeWhat to check before choosing
Find a recent factUpdate date, primary source, and match between the data and its citation.Reference period, methodological changes, and source accessibility.
Research a complex questionCoverage of subquestions, handling of disagreements, and separation of evidence from inference.Whether each important conclusion has direct, sufficient support.
Restrict a searchClear controls for limiting or prioritizing domains and traceable final references.Whether the system actually followed the agreed list.
Combine web research with user documentsClear evidence attribution and references that lead back to the original document.Upload availability, file handling, and distinction between sources.
Deliver an auditable reportReadable citations, working links, and low verification effort.Actual review time and proportion of supported claims.
06

Choose Without Declaring an Absolute Winner

The choice depends on the task, the risk of error, and how much time is available for review. For a low-risk internal note, a report whose main claims are easy to check may be enough. For publication, regulatory analysis, or a professional decision, it may be necessary to compare primary sources, document disagreements, and keep an audit record. A tool with more links is not necessarily more traceable, and a longer response does not mean broader coverage.

Before adopting a solution, set an acceptance threshold. For example, define what proportion of central claims must have direct support, which sources are mandatory for a given kind of data, and which errors require rerunning the search. The threshold should depend on the intended use and apply equally to all products. If a tool cannot meet an essential condition—because of access, controls, or inadequate citations—the appropriate conclusion is that it does not fit this workflow under the tested conditions, not that it is universally inferior.

Keep the results, instructions, and verification notes so you can repeat the evaluation when products or plans change. The matrix is for comparing observed evidence, not for turning a small test into a universal ranking. If results vary considerably between runs, report that variation; if data are missing for a tool, mark the comparison as incomplete. The most defensible decision explains what was tested, what could be checked, and what remains uncertain.

Open questions

  • No verified documentation for Claude Research was provided; its features, sources, controls, and availability remain to be checked.
  • The supplied sources do not contain a current, shared test of all four products or observed results from equivalent assignments.
  • Plans, quotas, and access options can change; verify them for each account and again at editorial close.
  • Product documentation describes stated options but cannot by itself determine citation accuracy or report completeness.
07

Keep exploring

07

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction