Last updated: July 21, 2026
Same Claude models as the app, wrapped in a sandbox, connectors, and provenance. Inspectable by design, correct only where you check.
Claude Science is Anthropic’s public-beta research app. It runs the Claude models your plan already includes and wraps them in a local code sandbox, science-database connectors, versioned artifacts that carry a provenance record, and a background reviewer. It is not a new or smarter model. When you compare it with the regular Claude app on the same prompt, you are comparing environments, not model quality.

Verdict: Claude Science strengthens the audit trail, not the accuracy guarantee. Buy it to make research inspectable, keep a human on the science, and wait if your data is regulated.
This is a hands-on review. Future Stack Reviews ran the same research prompt in the regular Claude app and in Claude Science on July 21, 2026, inspected the saved artifacts and their five-tab provenance, verified every load-bearing figure against the primary sources, and ran controlled reviewer probes on synthetic data with known answers. Product facts were read from Anthropic documentation the same day. Findings are N=1: one topic, one run per tool, one date, one plan, one beta build.
At a glance
A desktop app in public beta for macOS and Linux, with Windows supported through WSL 2. It pairs your plan’s Claude models with local Python, R, and shell execution, connectors to science databases, artifacts saved with the exact code and environment that produced them, and a background reviewer. Not a new model.
- Reproducible, code-and-artifact research where you want retrieval, analysis, and figures in one traceable place.
- Non-computational scientists who need a defensible first-pass analysis.
- Anyone who values an execution log over a polished paragraph.
- Protected health information: the beta is not covered by a BAA.
- Central governance: no audit log, export, or offboarding controls reach its data yet.
- Treating the auto-written prose as peer-reviewed fact, or as a systematic review.
Included in Pro, Max, Team, and Enterprise; on by default on Pro and Max; not on Free. It shares your plan’s weekly usage with Claude Code and Cowork. External compute through Modal is billed to you directly, with no spend ceiling. Pro is 17 USD per month billed annually (200 up front) or 20 USD monthly; Max from 100 USD per month. Confirm at checkout before you commit.
Contents
What Claude Science actually is
Claude Science is an application, not a model. Anthropic launched it in public beta on June 30, 2026 and describes it as an app that runs the Claude models your plan already includes and adds an analysis environment around them.
The difference that matters to a buyer is the harness, not the intelligence. Regular Claude already does web research, file handling, and code. Claude Science adds a local sandbox that runs Python, R, and shell code on your own machine, connectors to science databases, artifacts that are saved with the exact code and environment that produced them, a background reviewer, and the option to send heavy jobs to a remote server or a cloud provider.
It runs on macOS and Linux. There is no native Windows build yet; the documented path on Windows is to run the Linux binary under WSL 2 (Ubuntu 24.04 or newer). The installer command itself differs by page, a small but real sign of a young product: the Linux quick-start pipes the script to sh, while the WSL page pipes it to bash.
What it costs, and what “included” hides
On the pricing page, Pro, Max, Team, and Enterprise all read “Includes Claude Science.” On Pro and Max the app is on with no admin action; on Team and Enterprise an owner turns it on; Free has no access. So the entitlement question that trips up quick reviews has a clear answer: if you pay for Pro or Max, you already have it.
“Included” is not the same as free at the margin. Claude Science draws from the same weekly usage limits as Claude Code and Cowork, and the background reviewer spends that usage too, so a heavy research day competes with your other Claude work. External compute is a separate bill. If you connect a Modal account, Modal charges you directly, Anthropic never sees a payment method, and there is no spend ceiling in the app. A job keeps running and billing after you close the app until it finishes or times out, with a default container timeout of 12 hours and a maximum of 23. Price the tool as your subscription plus a metered cloud bill you set your own guardrails on.
Inside the workspace
Three parts of the workspace decide whether the tool earns its place: connectors, the sandbox, and artifacts.
Connectors are not all equal. Featured connectors to public life-science databases are read-only and need no key, and they cover a wide table of sources from Ensembl and UniProt to PDB, AlphaFold, GEO, and PubChem. A smaller set of Directory connectors, including PubMed, ChEMBL, and bioRxiv, must be added by an admin on Team and Enterprise plans. The gap to watch: a connector appearing in the catalogue is not the same as a connector you can run. OpenAlex is listed as a Featured connector that needs no key, yet a July update made it require a free API key for full-text access. In our own run the app tried OpenAlex, found no key configured, and fell back to PubMed alone. Read “60-plus databases” as a catalogue, not a promise that each one will answer.

The sandbox keeps local code contained; remote compute does not. Code runs inside an operating-system sandbox on your machine. The moment you send a job to an SSH host or an HPC cluster, the documentation is blunt: the job “runs outside the sandbox, as your user on the host, with access to everything your account can read and write there.” That is a design choice, not a flaw, but it moves the trust boundary from Anthropic’s sandbox to your own account permissions on the target machine.
Artifacts are the real product. Every saved artifact carries five tabs: the surrounding messages, the code, the execution log, the environment with every package version, and the reviewer’s findings. The documentation names the execution log the authoritative record: “if the Code tab and the log disagree, trust the log.” This is more than a chat transcript, and it is the strongest reason to prefer Claude Science over pasting code into a general assistant. It is also where the next two sections find their tension.
The reviewer: what it can and cannot do
A background reviewer re-reads the recent responses, the approved plan, the saved artifacts, and the execution record, then checks whether the claims match what actually ran. It flags results reported as computed when nothing ran, values that contradict the file they came from, citations that do not support the claim, and a DOI that resolves to a different article. It runs automatically on Max, Team, and Enterprise, and on Pro you trigger it with Request review or turn on an opt-in auto-review.
Two limits are stated plainly and both change how much you can lean on it. It “doesn’t re-run analyses,” and it “doesn’t judge whether that method was the right choice for your research question.” So the reviewer is a consistency check against the record, not an independent replication and not peer review. Reading “results that check and correct themselves” as scientific validation is the most expensive mistake a buyer can make here.
Auditability is not correctness
We gave the same literature-mapping prompt to the regular Claude app and to Claude Science, saved the Claude Science artifacts, and checked the load-bearing figures against the primary papers. One error shows exactly where an audit trail stops helping.
The saved review stated that loss of the SDHB protein “stratified 5-year progression-free survival from 91.5% down to 34.8% across risk tiers.” The paper it cited, which the app retrieved in full, defines those survival tiers as a four-factor risk model built from SDHB expression, primary tumor size, final diagnosis, and Ki-67 index. SDHB is one of the four. The prose promoted one factor to the whole model.
The narrative said: “SDHB loss was found in 17.6% of cases and stratified 5-year progression-free survival from 91.5% down to 34.8% across risk tiers.”
The app’s own structured record said: “5-year PFS rates: low-risk 91.5%, intermediate-risk 41.7%, high-risk 34.8%.”
The source said (retrieved by the app): “a 3-tier risk model to predict PFS in PPGL using four risk factors … SDHB expression, primary tumor size, final diagnosis, and Ki-67 index.”
The reference entry was right. The narrative was not. The reviewer did not flag it, and that is consistent with its scope: the numbers are real and SDHB is genuinely one of the factors, so nothing looked out of place against the record. Catching the over-attribution needed a read of the source, which is the work the tool is meant to reduce. This is not fabricated data and not a broken reviewer. It is a plausible over-attribution in the write-up layer that the provenance recorded faithfully and no automated check questioned.
The balance matters, because the same testing cut the other way too. When we built three deliberate errors into a controlled task on synthetic data, the model refused to call a pure-noise variable a “significant biomarker” and explained why, and it corrected two textbook statistical traps on its own. In a separate run the reviewer caught the model claiming it had verified a citation when the log showed no such step. Claude Science also disclosed its own limits in the SDHB run without being asked: PubMed-only retrieval because the OpenAlex connector lacked a key, abstract-level extraction, and two preprint-and-published pairs counted as separate records. The honest reading is narrow. Blatant fabrication was resisted; the residual risk is the quiet over-attribution in the prose, and neither the provenance nor the reviewer is built to catch that.

Local-first is not local-only
Anthropic calls the app local-first, and for stored data that holds: conversation history and artifacts live only on the member’s device. Two things still leave the machine. Every model call sends the prompt and the response to Anthropic under its standard retention policy, so “runs on your infrastructure” does not mean nothing is transmitted. And remote compute sends code and data straight to the destination you connect.
The sharper issue for an organization is that individual auditability and central control move in opposite directions. Because the data sits on laptops, the Anthropic-side controls an admin relies on do not reach it. In the beta the admin table lists the audit log, the Compliance API, organization data export, custom data retention over local data, and offboarding wipe all as “Not available,” and members can add their own local connectors and remote compute that an admin cannot yet restrict. So the tool that makes one scientist’s run more traceable also places that run outside the controls a regulated lab, a hospital, or a pharmaceutical company depends on for retention, discovery, and offboarding. Endpoint management becomes your only lever over that data.
A matched run, read honestly
On the same prompt, Claude Science returned 49 publication records where the regular app firmly verified 8 and named up to 11. That is not a tenfold advantage in coverage. The 49 are records, not independent studies: two are preprint-and-published pairs of the same work, which the app itself flagged, so the real count is 47. They span reviews, case reports, animal models, and mechanism papers, not 49 tissue-specific expression studies. Without an independent reference set, the honest comparison is record count, not recall or precision.
What Claude Science clearly added was a re-runnable search, a saved dataset, and a provenance trail. What it did not establish was superior accuracy: the regular app kept a tighter scope, Claude Science over-attributed the Zhang model, and both handled different figures well and badly. Treat the run as a case study of the workflow, not a scoreboard.

Who should use it, who should skip
Use it if your work is computational and repeatable, if you want code, data, figures, and an execution log tied to each result, or if you are a scientist who needs a credible first pass without building the pipeline yourself. The provenance is real, and for these users it lowers the cost of checking your own work.
Skip it, or wait, if you handle patient or regulated data, if your organization needs central audit, export, and offboarding today, if you need a formal systematic review with dual screening, or if you would read the reviewer’s silence as a stamp of correctness. For these buyers the beta’s governance gaps and the write-up risk outweigh the convenience.
macOS: download the installer from claude.com/product/claude-science and open it. There is no curl command on macOS.
Linux:
curl -fsSL https://claude.ai/install-claude-science.sh | sh
claude-science serveWindows: no native build. Run the Linux version under WSL 2 (Ubuntu 24.04 or newer), where the same installer pipes to bash.
On a remote server, start with claude-science serve --no-browser, forward the port over SSH, and open the printed URL from your laptop.
FAQ
Is Claude Science a different, smarter model?
No. It runs the same Claude models your plan already includes, with no special access. The difference is the environment around the model: local code, connectors, artifacts, provenance, and a reviewer. Compare it as a workflow, not as a new model.
Is it included in my plan?
Yes on Pro and Max, where it is on by default with no admin step. On Team and Enterprise an owner enables it. Free has no access. The pricing page lists it as included on all paid tiers as of July 21, 2026.
Does the reviewer make the output trustworthy?
Partly. It checks whether claims match the execution and citation record and flags real mismatches. It does not re-run the analysis and does not judge whether the method was appropriate, so it is a consistency check, not replication or peer review.
Does my data stay on my machine?
History and artifacts do; they are stored only on your device. Prompts and model responses still go to Anthropic under standard retention, and remote compute sends code and data to the destination you connect. Local-first, not local-only.
Can my organization audit or delete what people do in it?
Not in the beta. The audit log, Compliance API, organization export, and offboarding wipe are all listed as not available for Claude Science data, which lives on members’ laptops. Control depends on your own device management.
Does “included” mean no extra cost?
No. It shares your plan’s weekly usage with Claude Code and Cowork, and the reviewer spends usage too. External compute through Modal bills you directly with no spend ceiling, and jobs keep running after you close the app.
Will it run on Windows, and can I use patient data?
There is no native Windows build; you run the Linux version under WSL 2. On patient data, the beta is not covered by a Business Associate Agreement, and the documentation says it should not be used with protected health information.
Are the papers it found reproducible and complete?
It saves the code, environment, and execution log, which supports reproduction but is not a verified independent rerun. Coverage in our run was PubMed-only because a connector lacked a key, so treat the result as a floor, not a census.
Methodology and disclosures
Conflict of interest: this review was produced with Anthropic’s Claude, auditing an Anthropic product. We flag it so you can weigh it.
What we did: on July 21, 2026 we ran the same research prompt in the regular Claude app and in Claude Science, inspected the saved artifacts and their five-tab provenance, verified the load-bearing figures against the primary sources, and ran controlled reviewer probes on synthetic data with known answers. The synthetic-data errors were built on purpose to test the reviewer and are disclosed as method, not presented as findings.
Scope: findings are N=1, one topic, one run per tool, one date, one plan, and one beta build. They do not measure average accuracy, recall, or cost. Product facts were checked against Anthropic documentation on July 21, 2026; prices, connectors, and beta behavior change, so recheck before relying on them. We did not verify the biological content for medical use, and this review is not clinical or diagnostic guidance.
Primary sources checked (Anthropic, July 21, 2026): launch, pricing, enablement, the reviewer, data handling, admin controls, compute providers, remote compute, connectors and skills, artifacts, Windows and WSL, changelog.
Verdict
Claude Science is worth adopting for individual, reproducible research, and worth holding for regulated or centrally governed environments. Its provenance and self-disclosed limits are the real value, and they beat what a general assistant gives you. The catch is that the same run that showed its work also over-attributed a four-factor prognostic model to a single marker in its write-up, its own record kept the numbers straight, and the reviewer let the sentence stand. Provenance raised the ceiling on what you can inspect. It did not raise the floor on what is correct. Buy it to inspect faster, keep a human on the science, and treat the beta’s governance gaps as a reason to wait if your data is sensitive.
Stay with the review desk
Choose a channel to keep reading.
Share this review