Skip to content

Explore Science vs Penelope.ai

Penelope.ai checks in seconds whether a Word manuscript meets a journal's submission requirements, and journals wire it into their own workflow. Explore Science reviews the research: the whole paper read by a panel of models, scored out of 100, with every reference checked against live databases.

The short answer

Choose Explore Science when what you need to know is whether the work will survive review: the full manuscript read and cross-checked, a Calibre score out of 100 across eight fixed categories, every reference tested against the live literature, and each revision scored against the one before it. Penelope is a good desk check, free to the author on the journal pages that carry it, and it will find a missing funding statement faster than you will. Its checks ask whether a section is present, a count is inside the limit, a citation appears in the reference list. None of them asks whether the methods carry the conclusion, and that is the question a reviewer opens with.

Compared July 2026 · from Penelope.ai's own published materials

Choose Explore Science when

  • Reads the methods, results and argument, and reports where the design limits what the results are allowed to claim.
  • Verifies every reference against live scholarly databases, including whether the cited source actually supports the claim attached to it.
  • Calibre puts a score out of 100 across eight fixed categories on the paper, with the method published at /scoring, so two drafts can be compared on the same axes.
  • Claude, GPT, Gemini, Mistral and Grok run alongside our in-house Explorer One and cross-check each other's reading, rather than one rule firing per check.
  • Revisions stack in a workspace co-authors can join, with per-dimension movement shown draft to draft and Rosa answering questions on the paper, its review and your wider body of work.

Choose Penelope.ai when

  • Results come back in seconds, and the checklist report costs the author nothing on a participating journal's page.
  • Checks are configured against your target journal's own author guidelines, down to the wording of a required heading, which no general review tool does for you.
  • Runs inside the submission workflow for journals on Manuscript Manager, and through a public API, so the editor sees the same result the author does.
  • Publishes pass and fail criteria for every check and links each result to the marked-up text it was drawn from, so a failed check can be argued with.
  • Built by the team behind goodreports.org, whose EQUATOR Network collaboration won the 2018 Cochrane-REWARD prize; reporting checklists and trial registry IDs are surfaced where they apply.

Side by side

Capabilityexplore sciencePenelope.ai
What it examinesThe methods, results, argument and reference list, in any field30+ checks that a submission meets a journal's requirements: required sections and their titles, declarations, ethics statements, counts, referencing style, file cleanliness
Journal fitMatching a target journal and its submission systemJournal-independent; the review runs on any draft at any stageChecks configured per journal by its editor and served inside the submission workflow for journals on Manuscript Manager, plus a public API
TurnaroundUp to 2 hours, because the models read the paper end to end and check each other before anything is writtenSeconds; its BMJ Open case study describes authors emailed a link to their feedback within two minutes
Quality scoreCalibre: eight scored categories totalling 100, rubric publishedPass or fail per check, criteria published; no score for the work itself
ReferencesEvery reference verified against live scholarly databases, including whether the source supports the claim attached to itCitations matched to the reference list and back, plus style and ordering
NoveltyNovelty Check searches the live literature through Semantic Scholar and OpenAlexNot among its published checks
Review engineA panel of frontier models (Claude, GPT, Gemini, Mistral, Grok) plus our in-house Explorer One, each testing the others' reading of the paperA natural language processing engine of probabilistic checks, each with its pass and fail criteria published; no language model named
RevisionsEvery revision scored, with per-dimension movement shown draft to draft and co-authors able to join the workspaceEach upload checked afresh; no version-to-version comparison described in its published materials
Follow-up questionsRosa answers questions on the paper, its review and your earlier workAn annotated report you read; no channel to put a question back to it
Other passesProtocol and proposal, editorial and novelty passes, with plagiarism, AI-text and AI-figure detection as plug-insReporting checklists from the EQUATOR library surfaced when relevant, and clinical trial registry IDs checked
PricingFree to start, then $49 buys one full-depth review outright; monthly plans from $99 if you are revising continuouslyFree to the author for the checklist report on a participating journal's page, $9.50 for a track-changes Word file of suggestions on its demonstration journal; journals pay £750 to £4,850 a year by submission volume
Data policyWe do not train on your manuscriptPrivacy policy covers cookies and usage data and does not describe what happens to an uploaded manuscript (last updated January 2019)

90% rate a review on par with or better than the human peer review they have had.

Two in three rate it better. From 304 head-to-head ratings by researchers who had received peer review on the same paper. These are our own users, self-reported, which is worth knowing when you weigh them.

Depth

Why a thorough review takes the time it takes.

Most tools optimize for speed. Explore Science runs the engine we built to conduct original research autonomously, and points it at your manuscript: it reads recent literature in your field, checks every reference, and works through the paper the way a careful reviewer does.

The time is the method, not an overhead on it.

How a paper is scored

Time spent thinking about your paper

How long a review actually takes

Single-pass LLM
~2 min
Other AI reviewers
5 to 20 min
Human peer review
(active reading time)
~30 min
Explore Science
up to 2 hrs

Industry observation; self-reported vendor figures; Publons peer-review time survey, 2018.

The engine

What actually reads your paper.

Three pieces of the architecture do the work behind every review.

Models

A mixture of models, cross-checked.

Every frontier model (Claude, GPT, Gemini, Mistral, Grok) alongside our own, routed per sub-task and made to check each other. No single-model bias.

Agents

Scientific agents that reason like the field.

Dozens of specialised agents work each review in parallel, each tuned to one dimension of analysis, spawning more when a section needs them.

Infrastructure

An in-house science stack.

Memory that holds your paper in context for the whole review, live verification of every citation, and novelty assessment. Built over years of autonomous science.

Thirty checks, none of them about the argument

Penelope takes a Word file, reads it against a specific journal's author guidelines, and returns an annotated report: is there a dedicated ethical approval section with the title the journal wants, is a funder named, is the abstract structured with the subheadings PRISMA or CONSORT expects, are the figures numbered in ascending order, is the word count inside the limit, are there track changes or Endnote field codes still in the file. Every check links to the passage it fired on, and the pass and fail criteria are public. As a way of not losing three weeks to an editorial office bounce, it works, and on a journal page that carries it the report costs the author nothing.

What the list does not contain is a check on the research. No check reads the design, asks whether the sample supports the inference, or notices that the discussion claims more than the results earned. Explore Science runs at that level: methods, results and argument read end to end by several models that then check each other, and a Calibre score out of 100 across eight fixed categories, so the second draft can be measured against the first on the same axes rather than described as better.

The reference checks stop at the reference list

Penelope's referencing checks are internal ones. Does every citation in the text appear in the reference list, does every reference get cited, is the style numerical or named as the journal requires, is the ordering right. Those are the errors a copy editor would otherwise chase, and catching them before submission is worth something.

None of them leaves the document. Explore Science resolves each reference against live scholarly databases and reports back whether the work exists, whether it is attributed correctly, and whether it says what the sentence citing it needs it to say. A reference list can be perfectly consistent and still rest on a paper that found the opposite result.

Penelope answers to the journal

The product is sold to journals: £750 a year under 500 submissions, down to £0.80 a submission at scale, with the editor choosing which checks run, which are critical, and how the feedback is worded. Its natural home is the submit button, and journal-branded pages exist for titles including Access Microbiology and SPIE's Journal of Electronic Imaging. If your journal is on Manuscript Manager you may already be using it without having chosen it; Editorial Manager and ScholarOne are on a waiting list.

That placement is its strength and its ceiling. A check tuned to one journal's guidelines is useful at the moment you submit to that journal and inert before then, when the paper is still being argued into shape and the target may change. Explore Science works on the draft that has not been assigned a venue yet, and on the protocol before the study runs.

Common questions

Does Penelope check my statistics?

Its checks ask whether randomisation, blinding and sample size are addressed for experimental studies, and it points you to a reporting checklist from the EQUATOR library when one applies. That tests whether you said something, not whether what you said holds. Explore Science reads the analysis and reports where the design constrains the conclusions you can draw from it.

Which journals and submission systems can I use Penelope with?

Journals on Manuscript Manager, journals that have set up a branded check page (Access Microbiology and SPIE's Journal of Electronic Imaging among them), and anyone integrating its public API, as Peerwith did. Editorial Manager and ScholarOne are listed as a waiting list. If your journal has not added it, its demonstration journal page still runs a generic check against common requirements. Explore Science needs nothing from your publisher.

Penelope says its checks are 90% accurate. What does that figure cover?

Each check is tested against a set of manually labelled manuscripts and held to 90% or better, and the criteria are published per check, which is more openness than most tools in this space offer. It is a claim about a check firing correctly on the presence of a funding statement or a running head. It is not a claim about judgement, because none of the checks makes one.

I write in LaTeX. Can Penelope read my manuscript?

Not currently. It checks Word files, prefers .docx, and will attempt to convert a .doc before processing; its FAQ lists LaTeX and PDF as planned formats. Explore Science does not require a Word file.

Already using Penelope.ai? Put the same paper through a scored review.

Get started free

Explore Science vs Penelope.ai (2026) · Explore Science