Last updated: July 26, 2026
Future Stack Reviews ran one frozen, closed-book task six times on 25 July 2026: three runs on Claude Opus 5 and three on Claude Opus 4.8, at High effort, in fresh chats outside any project, in the Claude desktop app, with no tools and no follow-up turns. Anthropic’s documentation was read directly on 26 July 2026. The test did not cover the Claude API, Claude Code, other effort levels, tool use, latency, token consumption or cost. Fable 5 was not tested. Scoring was not blinded.
Claude Opus 5 is Anthropic’s same-price successor to Opus 4.8, listed at $5 per million input tokens and $25 per million output tokens. In a frozen 12-item reconciliation task, both models matched the reference answer on every content-scored field across three High-effort runs each. Anthropic reports larger gains on selected coding and automation benchmarks, including near-Fable 5 performance in one CursorBench setting. The two sets of results measure different things.
Verdict in one line: Opus 5 matched Opus 4.8 in the workflow we tested, and changed thinking defaults plus two removed API features still make this a migration rather than a model-ID swap.
Sources: Anthropic, 24 July 2026 · Future Stack Reviews controlled test, 25 July 2026
Performance: observed and reported

Two kinds of evidence appear in this review, and they are not interchangeable. One comes from a task we ran and scored. The rest comes from Anthropic, measuring different things under conditions we cannot inspect.
| Question | Answer | Status and boundary |
|---|---|---|
| Did Opus 5 outperform Opus 4.8 in our test? | No scored difference was detected. Both matched the reference answer on every content-scored field in three of three stored runs. | OBSERVED BY FSR One frozen workflow, Claude desktop app, High effort |
| Does Anthropic report broader gains? | More than double Opus 4.8’s Frontier-Bench v0.1 performance, at a lower cost per task. | OFFICIAL CLAIM Internal run, mini-SWE-agent harness, mean of five attempts |
| Does Opus 5 reach Fable 5? | Within 0.5% of Fable 5’s peak CursorBench 3.2 score at max effort, at half the cost per task. | OFFICIAL CLAIM One benchmark, one effort setting |
| Is Opus 5 the better-aligned model? | Anthropic’s automated behavioral audit scores it above Opus 4.8, Sonnet 5 and Fable 5 on constitution adherence and deceptive behaviour. | OFFICIAL CLAIM Internal pre-deployment audit |
| Did FSR test Fable 5 here? | No. | UNTESTED No three-model ranking is possible from this review |
Anthropic also reports Opus 5 scoring three times the next-best model on ARC-AGI 3, surpassing Fable 5’s best OSWorld 2.0 result at just over a third of the cost, and passing around 1.5 times as many Zapier AutomationBench tasks as the next-best model at the same cost per task. It notes that Opus 5 remains behind Mythos 5 on cybersecurity tasks.

These are selected vendor results at chosen effort settings. Our test asks a narrower question: whether one fixed output contract survived the model swap.
Sources: Anthropic, 24 July 2026 · Future Stack Reviews controlled test, 25 July 2026
Opus 4.8, Opus 5 and Fable 5
Fable 5 belongs in this comparison, and not because it sits above Opus 5 on capability. It carries a different procurement profile: double the API rate, thinking that cannot be turned off, and a mandatory retention period that rules out zero data retention arrangements.
| Opus 4.8 | Opus 5 | Fable 5 | |
|---|---|---|---|
| List price per 1M tokens | $5 / $25 | $5 / $25 | $10 / $50 |
| Context, max output | 1M, 128k | 1M, 128k | 1M, 128k |
| Paid plan chat context | 500K | 1M | Not established here |
| Thinking when omitted | Off | Adaptive, on | Always on |
| Disabling thinking | Allowed | At high effort or below | Blocked |
| Priority Tier | Supported | Not supported | Supported |
| Zero data retention | Available | Available | Not available |
| Mandatory retention | None for general access | None for general access | 30 days |
| FSR hands-on in this review | 3 stored runs | 3 stored runs | Not tested |
Zero data retention availability is a plan and product eligibility question rather than an automatic entitlement. What the table records is that Opus 5 and Opus 4.8 carry no mandatory retention period for general access, while Fable 5 does.
Web fetch is unavailable on Opus 5 and available on Opus 4.8. Fable 5’s position on that tool is not established in the pages we opened.
Sources: Anthropic Claude Platform Docs, Models overview, accessed 26 July 2026 · Anthropic Claude Platform Docs, Migration guide, accessed 26 July 2026 · Anthropic Claude Help Center, Context window on paid plans, accessed 26 July 2026
The six-run test
We froze a single closed-book prompt before any run. It supplies 36 fictional records covering one production change package, in deliberately nonchronological order, and asks for a 12-row table resolving the final state of each item at a fixed snapshot time. Sixteen resolution rules govern what counts as evidence.
The fixture includes drafts, meeting notes and a roadmap proposal that must be rejected, records from another tenant and a staging environment that fall outside scope, an amendment approved but not yet effective at the snapshot, a waiver that permits a release while leaving the underlying control incomplete, and an acceptance record that restates measurements but declines to classify the outcome.
| Order | Run ID | Model |
|---|---|---|
| 1 | FSR-T1-O5-R1 | Opus 5 |
| 2 | FSR-T1-O48-R1 | Opus 4.8 |
| 3 | FSR-T1-O48-R2 | Opus 4.8 |
| 4 | FSR-T1-O5-R2 | Opus 5 |
| 5 | FSR-T1-O5-R3 | Opus 5 |
| 6 | FSR-T1-O48-R3 | Opus 4.8 |
<p style=”font-size:13px;color:#6b7280;line-height:1.6;”>Prompt file SHA-256: 56e522c6a88481392ea9f2a83136bfef88a166a2d7fce89e3ae514975b19f5d4 (17,905 bytes). Identical input for all six runs.</p>
Every stored run matched the reference answer on resolved value, status and both evidence columns, for all twelve items.
| Measure | Opus 5, 3 runs | Opus 4.8, 3 runs |
|---|---|---|
| Resolved value and status | 72 / 72 | 72 / 72 |
| Evidence binding | 72 / 72 | 72 / 72 |
| Complete rows | 36 / 36 | 36 / 36 |
| Schema compliance | 18 / 18 | 18 / 18 |
| Output contract | 15 / 18 confirmed, 3 pending | 15 / 18 confirmed, 3 pending |
Output contract items remain pending because our stored copies were normalised during transfer. We cannot confirm from them whether each original response rendered as a single Markdown table. Content scoring is unaffected.
The normalised scored tables matched across the three stored runs for each model, and no cell-level difference appeared between the two models. That agreement extends to three items in the task that have defensible wrong answers.
One item requires citing a future-effective amendment as decisive evidence for a current value, because without it there is no explanation for why the later figure is not yet in force. One requires excluding an enabling prerequisite from the evidence set, because a single operational record establishes the result directly. One requires excluding an execution record whose measurements are fully restated in a later acceptance, while still citing a separate baseline to classify the outcome. All six stored runs made those three calls the same way.
Source: Future Stack Reviews controlled test, 25 July 2026
What the result does not show
Both models reached the top of the content rubric. A test where everything passes has no resolution above its own ceiling, so this harness cannot measure any distance between the two models. The result is a non-detection, not a demonstration of equivalence.
Three runs per model is operational repetition rather than a statistical sample. We report three of three, and do not describe that as a success rate.
The test establishes nothing about general equivalence between the models, superiority of either, reliability beyond these runs, behaviour on the Claude API or in Claude Code, tool use, performance near the context limit, latency, token consumption, cost, or safety for a production parser. It also says nothing about Fable 5, which we did not run.
Our prompt was self-contained and carried no scaffolding written for an earlier model. That is the condition under which we found no difference, and it is not the condition most production prompt stacks are in.

Source: Future Stack Reviews controlled test, 25 July 2026
Two breaking changes
Anthropic documents exactly two breaking changes for code already running on Opus 4.8.
Thinking now runs by default. A request with no thinking field ran without thinking on Opus 4.8. The same request runs with adaptive thinking on Opus 5. Because max_tokens caps thinking and visible output together, a value tuned for the older model can truncate a response that previously fit.
Disabling thinking is capped at high effort. Passing thinking: {type: "disabled"} is still allowed, but only at effort high or below. Combining it with xhigh or max returns a 400 error, checked per request, so a conversation that succeeded on earlier turns will fail on the turn that raises effort.
Anthropic adds a caution for anyone disabling thinking to save cost. With thinking off, the model can occasionally write a tool call into visible text instead of emitting a structured tool_use block, and can emit internal XML tags. In agentic loops that text stays in the conversation history and affects later turns. The documented alternative is to keep thinking on and lower the effort level.
Anthropic lists the following as carrying across unchanged: the 1M context window with no beta header, 128k max output, prompt caching, batch processing, the Files API, PDF support, vision, and both server-side and client-side tools. Two exceptions apply. Web fetch is unavailable on Opus 5, and Priority Tier is not supported on it, while Opus 4.8 retains both.
Sources: Anthropic Claude Platform Docs, Migration guide, accessed 26 July 2026 · Anthropic Claude Platform Docs, Prompting Claude Opus 5, accessed 26 July 2026
Prompts and effort
Anthropic’s prompting guide states that Opus 5 verifies its own work without being told to, and that verification instructions carried over from earlier models cause over-verification. On removing them, the guide says “removing them reduces wasted tokens with no loss in quality”. The same applies to harness scaffolding that adds a separate verification step.| Take out | Put in |
|---|---|
| Explicit verification and self-check steps | An explicit conciseness or target-length instruction |
| Instructions to use a subagent for verification | Conditions for delegation, or a cap on subagent count |
| Re-check directives before responding | An explicit scope boundary for narrow tasks |
| Harness steps re-running a check the model already performs | A stated cadence for progress narration in agentic work |
| In review prompts, conservatism or high-severity-only instructions | Length calibration for documents written to disk |
Anthropic warns that a review prompt asking the model to be conservative or to report only serious issues may be followed literally, producing fewer findings, and recommends asking for everything and filtering in a separate pass. Teams running code review, compliance checking or editorial audit tooling on Claude should check whether their tuning instructions now suppress output.
On effort, Anthropic’s documentation describes the parameter as controlling how many tokens Claude spends in responding, affecting all tokens in the response including tool calls. At lower effort the model makes fewer tool calls. The same page describes effort as a behavioral signal rather than a strict token budget: at a lower setting the model still thinks on a sufficiently difficult problem, just less than it would at a higher setting.
The prompting guide adds that lowering effort can reduce thinking volume without reliably shortening the visible response, and directs you to prompt for length explicitly if that is the goal. On maximum effort, the migration guide says it can deliver gains on the most demanding tasks, may show diminishing returns from increased token usage, and can be prone to overthinking on simpler ones.
Anthropic’s instruction is to run a fresh effort sweep on your own evaluations rather than carry a setting across. All six of our runs used High, so we tested none of this.
Sources: Anthropic Claude Platform Docs, Prompting Claude Opus 5, accessed 26 July 2026 · Anthropic Claude Platform Docs, Effort, accessed 26 July 2026 · Anthropic Claude Platform Docs, Migration guide, accessed 26 July 2026
Surfaces and fallback
Context limits and feature availability differ by surface. On the Claude API both models run at 1M tokens. On paid plan chat, Opus 5 runs at 1M and Opus 4.8 at 500K. In Claude Code both reach 1M on paid plans, with Pro users needing usage credits enabled for Opus models.
Default effort differs in documentation coverage too. The models overview states that on Opus 4.8 the effort parameter defaults to high on all surfaces, naming the Claude API, Claude Code and claude.ai. For Opus 5 and Sonnet 5 the equivalent sentence names only the Claude API and Claude Code. We did not find a published default for Opus 5 on claude.ai in the pages we opened. Our runs set High manually.
Opus 5 ships with cybersecurity classifiers that allow it to find vulnerabilities in source code while blocking binary-based vulnerability scanning, penetration testing and exploit generation. Anthropic expects them to intervene around 85 percent less often than the classifiers on Fable 5.
In the Claude apps, Claude Code and Cowork, a flagged request is re-run on a less capable model in the same conversation. Anthropic states that a notice appears and the response carries a label naming the model that answered. After the switch, the picker stays on the less capable model for the rest of that conversation, and switching back may trigger the same fallback again if the original request is still in the thread. For long sessions, the per-response label is the record to keep rather than the picker state at the start. On the Claude API, automatic switching is off by default and customers configure fallbacks explicitly.
Anthropic states that Opus 5 does not fall back on biology, chemistry or life-sciences questions, using safeguards similar to those on Opus 4.8 for those topics.
Retirement schedules split by surface as well. Anthropic’s published dates apply to the Claude API, Claude Platform on AWS and Microsoft Foundry. Amazon Bedrock and Google Cloud set their own. On 26 July 2026 both claude-opus-4-8 and claude-opus-5 are listed as Active, with tentative earliest retirement no sooner than 28 May 2027 and 24 July 2027 respectively. Those are floors rather than assigned dates, and Anthropic commits to at least 60 days’ notice before retiring a publicly released model. Opus 4.8 does not appear in the latest models comparison table, which does not establish deprecation; the lifecycle page lists it as Active.
Sources: Anthropic Claude Help Center, Context window on paid plans, accessed 26 July 2026 · Anthropic Claude Help Center, Why Claude switched models with Opus 5, accessed 26 July 2026 · Anthropic, 24 July 2026 · Anthropic Claude Platform Docs, Model deprecations, accessed 26 July 2026
Who should migrate, and when
| If this describes you | Route |
|---|---|
| Priority Tier commitment, or a workload calling web fetch | Wait or redesign |
| Requests combining disabled thinking with xhigh or max effort | Fix before migrating |
| Requests previously ran with no thinking field and a tight max_tokens | Test first, raise max_tokens |
| Downstream systems parse Claude output automatically | Test first, on your fixtures |
| Prompts carry verification, subagent or re-check instructions | Test first, A/B the removals |
| Running on Amazon Bedrock or Google Cloud | Check that platform’s own schedule |
| Regulated workload requiring zero data retention | Opus 5 over Fable 5 |
| None of the above, prompts self-contained | Replay fixtures, then canary |
Sources: Anthropic Claude Platform Docs, Migration guide, accessed 26 July 2026 · Anthropic Claude Platform Docs, Model deprecations, accessed 26 July 2026
FAQ
Did Opus 5 beat Opus 4.8 in your test?
No scored difference was detected. Both matched the reference answer on every content-scored field in three of three stored runs, including the three items with defensible wrong answers. Both reached the top of the rubric, so the test could not measure any distance between them.
Is Opus 5 cheaper to run than Opus 4.8?
The list price is identical at $5 per million input tokens and $25 per million output tokens. Whether a workload costs less depends on token consumption, which changes because thinking now runs by default and responses run longer. We did not measure consumption in any run.
Do I have to change my prompts?
Only the two API changes are mandatory. Anthropic separately recommends removing verification instructions, subagent verification steps and re-check directives, and adding explicit length, scope and delegation guidance. Treat each as a candidate for an A/B test against your own evaluations rather than a required edit.
Should we use Fable 5 instead?
Fable 5 lists at double the API rate, cannot have thinking disabled, requires 30-day retention and is unavailable under zero data retention arrangements. Anthropic reports Opus 5 coming within 0.5% of its peak CursorBench score at max effort. For regulated buyers the retention difference usually decides this before capability does.
How do I know which model answered in the app?
Anthropic states that when an automatic switch happens a notice appears and the response is labelled with the model that answered. The picker then stays on the less capable model for the rest of that conversation, so the per-response label is the more reliable record.
Sources: Anthropic Claude Platform Docs, Migration guide, accessed 26 July 2026 · Anthropic, 24 July 2026 · Anthropic Claude Help Center, accessed 26 July 2026 · Future Stack Reviews controlled test, 25 July 2026
Methodology
The task. One closed-book prompt supplying 36 fictional records for a single production change package, in deliberately nonchronological order, with 16 resolution rules and a fixed six-value status vocabulary. Required output is a 12-row table with resolved value, status, decisive evidence and displaced evidence. Prompt file SHA-256 56e522c6a88481392ea9f2a83136bfef88a166a2d7fce89e3ae514975b19f5d4, 17,905 bytes, identical for all six runs.
We are not publishing the fixture itself. A published evaluation set can enter future training data and stop measuring what it was built to measure. The hash is here so that if we release it later, anyone can confirm the file has not changed since this test.
Conditions. Claude desktop app on macOS, 25 July 2026. High effort selected manually. Fresh chat per run, outside any project. No tools, no browsing, no follow-up turns, no edits, no regeneration. Runs alternated between models.
Scoring. Exact match against a reference answer on resolved value with status, evidence binding, row completeness and schema compliance. A fifth dimension covers the prompt’s prohibition on any prose, code fence or formatting outside a single Markdown table.
Scoring was not blinded. The scorer knew which model produced each output. Scoring is exact match and all six stored outputs agreed, so we judge the effect to be small, and record the condition rather than omitting it.
Open item. Fifteen of eighteen output-contract checks are confirmed per model. Three remain pending because our stored copies were normalised during transfer, so we cannot verify from them whether each original response rendered as a single Markdown table. We have not established byte-level identity between runs. We will close this against the original responses and update the article.
Not measured. Token consumption, cost, latency, active runtime, or behaviour at any effort level other than High. We did not test the Claude API, Claude Code, Cowork, tool use, or performance near the context limit, and we did not run Fable 5.
Primary sources. Every documentation claim traces to an Anthropic page opened and read on 26 July 2026: the Opus 5 launch post, the what’s-new page, the migration guide, the models overview, the model deprecations page, the Opus 5 prompting guide, the effort documentation, the refusals and fallback documentation, and two help centre articles covering automatic model switching and paid-plan context windows.
Corrections. If any statement here is wrong, tell us at [email protected]. We correct in place and date the change. Corrections are handled by the editorial side and are never routed to commercial enquiries.
Sources: Future Stack Reviews controlled test, 25 July 2026 · Anthropic, 24 July 2026 · Anthropic Claude Platform Docs, Migration guide, accessed 26 July 2026 · Anthropic Claude Platform Docs, Models overview, accessed 26 July 2026 · Anthropic Claude Platform Docs, Effort, accessed 26 July 2026
Verdict
In one frozen reconciliation workflow, run three times per model, we found no task-local reason to reject an Opus 4.8 to Opus 5 migration. Both models matched the reference answer on every content-scored field, and agreed on the three rulings in the task that have defensible wrong answers.
That is narrower than saying Opus 5 is better, and narrower than saying the migration is drop-in. Our prompt was self-contained and carried nothing written for an earlier model.
The official delta stands regardless of our result. Thinking now runs by default, which changes what fits inside an existing max_tokens. One previously valid configuration returns an error. Web fetch and Priority Tier are excluded. Anthropic asks for a prompt audit and a fresh effort sweep.
Replay representative production fixtures against both models. Canary the workloads that pass, hold the ones with a dependency on an excluded feature, and measure consumption through the API rather than inferring it from a list price that has not changed.
Both models are available now. To replay your own fixtures and compare results, the Claude Console is where you swap the model ID and run an effort sweep. Plan-level access for the Claude apps is covered on Anthropic’s pricing page.
Sources: Future Stack Reviews controlled test, 25 July 2026 · Anthropic Claude Platform Docs, Migration guide, accessed 26 July 2026 · Anthropic Claude Platform Docs, Prompting Claude Opus 5, accessed 26 July 2026
Future Stack Reviews is the publishing arm of Future Stack LLC, based in Japan. We work with teams on evidence-bound model evaluation and prompt-stack review. If a migration decision in this briefing applies to your workload, get in touch.
Contact usTier B means we tested the product hands-on. Tier C means the briefing is document-first, with no hands-on testing.
- Claude Fable 5 Is Back, But Its Usage Meters Do Not Agree The tier above Opus 5 at double the API rate, and what its own usage meters report.
- Fable 5 Built a Landing Page, Then Security-Reviewed Its Own Code. Zero Fixes. Here’s What “Clean” Actually Meant. Anthropic says Opus 5 verifies its own work without being told to. This is a self-review from the tier above, and what it returned.
- Gemini 3.6 Flash Review: The Price Cut Is Real. The same upgrade question at another vendor, starting from a confirmed price cut.
- Claude Opus 4.8 Review: A Safer Model, a Worse Operator Background on the model you would be migrating from, and where its operating behaviour diverges from its safety record.
- Claude Fable 5 From July 20, 2026: What Happens on Each Paid Plan The same model priced through a subscription instead of the API, where seat class decides whether the usage is included at all.
Future Stack Reviews is an independent publication. This review reflects testing and documentation on the dates stated. Model pricing, availability, lifecycle dates and documented behaviour change without notice, and buyers should confirm against their own account before making a purchasing decision. Nothing here is legal, financial or procurement advice.
Hands-on testing: 25 July 2026, six runs. Documentation reviewed: 26 July 2026. Last updated: 26 July 2026.
Stay with the review desk
Choose a channel to keep reading.
Share this review