A mixture of models, cross-checked.
Every frontier model (Claude, GPT, Gemini, Mistral, Grok) alongside our own, routed per sub-task and made to check each other. No single-model bias.
Claude Science is a research workbench: it runs your analysis, makes your figures, and drafts alongside them. Explore Science judges the finished manuscript, puts a Calibre score on it against a published rubric, and tests every reference in the bibliography.
Choose Explore Science when the question is how good the manuscript actually is: eight categories, a score out of 100 against the rubric we publish at /scoring, every reference checked against live databases, and each revision scored so you can see which categories moved. Claude Science is built for the work upstream of that, the analysis and the figures and the compute, and it is genuinely good at it. What it will not give you is a position: nothing in it is calibrated to tell you where this paper sits or whether it is ready to go out. Bring the draft here once it is written.
Compared July 2026 · from Claude's own published materials
| Capability | ![]() | Claude |
|---|---|---|
| What it is for | Judging a written manuscript and tracking it across drafts | Doing the research: analysis, figures, compute, and drafting alongside them |
| Running the analysisData, code, compute | None. We read the manuscript you bring us | Persistent Python and R kernels, SSH and Modal job submission, laptop to HPC cluster to on-demand GPUs |
| TurnaroundPer manuscript | One pass, up to 2 hrs, run to completion with nobody at the keyboard; what arrives is a report, not a session you steer | Interactive sessions; long analyses run as batch jobs you monitor |
| Review engine | Claude runs here too, alongside GPT, Gemini, Mistral, Grok and our in-house Explorer One, each checking what the others concluded | Anthropic's own models throughout, with coordinator and reviewer agents over 60+ curated skills |
| Quality score | Calibre: one score out of 100, eight fixed categories, formula published | The reviewer agent reports findings; no manuscript-level score or rubric is described in its published materials |
| Reference checking | Every reference in your list resolved against live databases, and checked against the claim it supports | Reviewer agent flags incorrect citations in the work Claude produced; a whole-bibliography pass is not described |
| Literature search | Novelty Check queries Semantic Scholar and OpenAlex to test whether the contribution has already been made | PubMed and OpenAlex sit among the connected databases and are searched inside a session; OpenAlex full text needs your own free API key |
| Field coverage | One rubric, any discipline | Preconfigured for genomics, single-cell, proteomics, structural biology and cheminformatics |
| Draft to draft | Every revision scored, with movement shown per category | Artifacts are versioned and can be diffed against a chosen version; no score is attached to a version |
| Paper-grounded chat | Rosa opens with the manuscript, its Calibre report and your back catalogue already in hand, and stays on that paper | Projects and memory hold context per project; grounding is whatever you have loaded into that project |
| Pricing | Free to start; $49 buys one full-depth review outright with nothing to cancel; plans from $99 a month | Pro $17/mo billed annually or $20 monthly; Max from $100/mo; Team $20 per seat annually, with discounted academic seats |
90% rate a review on par with or better than the human peer review they have had.
Two in three rate it better. From 304 head-to-head ratings by researchers who had received peer review on the same paper. These are our own users, self-reported, which is worth knowing when you weigh them.
Most tools optimize for speed. Explore Science runs the engine we built to conduct original research autonomously, and points it at your manuscript: it reads recent literature in your field, checks every reference, and works through the paper the way a careful reviewer does.
The time is the method, not an overhead on it.
How a paper is scoredTime spent thinking about your paper
How long a review actually takes
Industry observation; self-reported vendor figures; Publons peer-review time survey, 2018.
Three pieces of the architecture do the work behind every review.
Every frontier model (Claude, GPT, Gemini, Mistral, Grok) alongside our own, routed per sub-task and made to check each other. No single-model bias.
Dozens of specialised agents work each review in parallel, each tuned to one dimension of analysis, spawning more when a section needs them.
Memory that holds your paper in context for the whole review, live verification of every citation, and novelty assessment. Built over years of autonomous science.
Claude Science ships a reviewer agent, and it is a real thing: it flags incorrect citations and figures that do not match the data behind them. The scope is the work produced in the session. It is quality control on Claude's output, which is exactly what you want when an agent has just written eight paragraphs of related work.
Our pass runs the other way, over the bibliography you wrote. Every entry is resolved against live scholarly databases, and we check whether the source supports the claim it has been attached to. The dangerous reference is rarely the one an agent invented five minutes ago; it is the one you carried forward from a grant application two years back, half-remembered and never reopened.
Worth saying plainly: Claude is one of the models we run. A review here goes through Claude, GPT, Gemini, Mistral and Grok plus our in-house Explorer One, and their job is partly to disagree with each other. Where they diverge on a methods criticism or a claim of novelty, that divergence is the signal, and it is the reason no verdict rests on one model reading your paper once.
So the comparison is not Claude against something better than Claude. It is one vendor's models answering in a session, against several models adjudicated against each other and against a rubric that does not move between papers.
Claude Science holds a project together well: sessions keep context, artifacts are versioned and diffable, projects archive without losing anything, and memories stay scoped to the project they came from. All of that is state, and state is not the same as measurement. Nothing in it tells you that version four is better than version two, or by how much, or where.
Explore Science scores each revision on the same eight categories and shows you the movement. That is what turns a rewrite from a feeling into a decision: submit, or fix the two categories that did not move.
On Free, Pro and Max it depends on your Model Improvement setting. With it on, chats can be used to improve Claude and may be retained in de-identified form for up to five years in training pipelines; with it off, the 30-day retention window applies, and Incognito chats are excluded either way. Commercial plans (Team, Enterprise, Education and the API) sit under separate terms and are not covered by that consumer setting. Explore Science does not train on your manuscript at all, on any plan.
It is a desktop application, so it runs on the machine in front of you: co-authors on Windows are not covered, and on Team and Enterprise plans an administrator has to enable it before anyone can install it. Explore Science runs in a browser, and a paper's workspace is shared by sending a link, which matters when the person who needs to see the review is a supervisor in another department rather than someone with your setup.
Its general capabilities do, and the Python and R environments are field-agnostic. The connected databases and the 60-plus curated skills are weighted to genomics, single-cell analysis, proteomics, structural biology and cheminformatics, so a materials chemist or a health economist gets less of the preconfiguration and more of a general assistant with kernels attached. Calibre applies the same eight categories whatever the discipline, because it scores the argument and the evidence rather than the assay.
There are specialised passes for protocols and proposals, for editorial work, and for novelty, each run against the same manuscript rather than as a separate upload. Plagiarism, AI-text and AI-figure detection are available as plug-ins on top. These are the checks a journal office runs after you submit, which is a poor time to find out.
Claims about Claude reflect its own published materials; everything about Explore Science reflects what the platform does today. Sources: Claude Science (product page), Claude Science, an AI workbench for scientists (Anthropic, 30 June 2026), Claude Science changelog, Claude pricing, How large is the context window on paid Claude plans? (Claude Help Center), What are projects? (Claude Help Center), Is my data used for model training? (Anthropic Privacy Center), Updates to Consumer Terms and Privacy Policy (Anthropic). Something out of date? and we will fix it.
Get started free