Most model launches are a capability story with a price footnote. Anthropic's July 24 release inverts that.
Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens. That is not a discount off anything. It is the exact price Anthropic charged for Claude Opus 4.8, the model it replaces. What changed is what you get for it: on CursorBench 3.2, a coding benchmark run by Cursor, Anthropic reports Opus 5 landing within 0.5 percent of the peak score of Claude Fable 5 at max effort, and at half the cost per task. Fable 5 is the frontier model Anthropic shipped in June at $10 and $50.
Anthropic's own framing is careful, and worth quoting: Opus 5 "comes close to the frontier intelligence of Claude Fable 5 at half the price." Close to, not past. That distinction is the whole article, and it is more useful to a buyer than either the headline or the skepticism.
What Actually Shipped
Opus 5 is positioned as a workhorse rather than a showpiece. Anthropic describes it as "a thoughtful and proactive model," and the specific capability it leans on is not raw intelligence but self-correction: Opus 5 is "much stronger at verifying its work and iterating carefully until it succeeds."
The concrete pieces:
- Effort settings. The model exposes five reasoning-effort settings, from low through max. Higher effort spends more tokens on reasoning; lower effort finishes faster and cheaper. This is a per-request dial, not a per-model decision, and the setting that produces the best score is not always the highest one: on Cognition's FrontierCode v1.1, Opus 5 peaks at medium effort.
- Fast mode. Roughly 2.5 times the default speed for twice the base price, available on the Claude Platform and through usage credits in Claude Code.
- Availability. Claude.ai, Claude Code, Claude Cowork, and the API as
claude-opus-5. It is the new default for Claude Max subscribers and the strongest model available on Pro. General access carries no data retention requirement. - Two API betas. Mid-conversation tool changes let developers alter the available tool set without invalidating the prompt cache. Automatic fallbacks route a request flagged by a safety classifier to another model instead of returning a refusal.
Those last two are quieter than the benchmarks and, for anyone building on the API, more consequential. More on the second one below.
The Numbers, From the System Card
Anthropic's announcement page speaks in multiples and comparisons. The hard numbers sit in the system card, and it is worth reading directly. Everything below comes from that document: the scores from its capability summary table, which reports Opus 5, Opus 4.8, Fable 5, and GPT-5.6 Sol side by side, and the task descriptions and per-benchmark detail from the sections behind it.
FrontierBench v0.1, a 74-task suite of hard terminal and engineering problems spanning computational biology, physics simulation, CAD, formal proofs, and GPU performance work: Opus 5 scores 43.3 percent against Opus 4.8's 18.7 percent. That is more than double. It also clears Fable 5 at 33.7 percent and GPT-5.6 Sol at 37.5 percent.
AutomationBench, Zapier's benchmark for end-to-end business workflows, where an agent has to discover the right endpoints across 47 simulated apps, make dozens of interdependent API calls, and obey layered policy documents: Opus 5 scores 26.0 percent against 17.0 for Opus 4.8 and 17.4 for Fable 5. The more interesting line is the one below the leaderboard number. At medium effort, Opus 5 scored 24 percent at $0.89 per task, beating both previous flagships at less than half the cost.
OSWorld 2.0, long-horizon computer use: 70.6 percent, against Fable 5's 66.1 and Sol's 62.6. Two weeks ago, GPT-5.6 Sol clearing Anthropic's Opus-class model on computer use was a headline. Opus 5 takes that back by eight points.
GDPval-AA v2, an independent Artificial Analysis framework built on 220 real professional tasks drawn from 44 occupations: Opus 5 posts 1861 against Opus 4.8's 1593, and leads both Fable 5 and Sol in Anthropic's table. This is v2 of the framework, so its figures are not comparable to GDPval numbers quoted at earlier launches.
ARC-AGI-3, a novel-problem-solving evaluation: 30.2 percent at high effort, roughly four times Sol's 7.78 and twenty times Opus 4.8's 1.52. Anthropic notes its max-effort result was not available at release.
On the science side, Anthropic reports gains across every life-sciences evaluation it ran, including a jump of more than ten percentage points on organic chemistry tasks over Opus 4.8.
Where It Does Not Lead
This is the section most launch coverage skips, and the one that makes the rest credible.
It is not top of every coding table. On SWE-bench Pro, Fable 5 scores 80 percent and Opus 5 scores 79.2. On FrontierCode v1.1's main set, Fable 5 leads by a tenth of a point, 53.5 to 53.4. On DeepSWE v1.1 the gap is wider and the order is different: GPT-5.6 Sol leads at 72.7, Fable 5 takes 69.7, and Opus 5 comes third at 68.8. Opus 5 reaches these numbers at half the token price of Fable 5, which is the real story, but "state of the art on coding" and "cheapest path to roughly state of the art" are different claims. Outside coding, GPT-5.6 Sol leads ARC-AGI-2 at 92.5 against Opus 5's 90.4, and ties it on ARC-AGI-1 at 97.5.
It trails on dual-use capability, and the safeguards around it are the deliberate part. Anthropic assessed Opus 5 as not exceeding Claude Mythos 5's risk in chemistry and biology, and applies the same ASL-3 protections used for Opus 4.8. On cyber evaluations it improves at finding software vulnerabilities but remains substantially behind Mythos 5 at exploiting them. Worth being precise about why: the card describes Opus 5 as "a general-purpose model not specifically trained for cyber tasks," where "any cyber-relevant skill likely reflects general capability gains rather than targeted training." The capability gap is a byproduct of what the model was trained for. What Anthropic chose deliberately is the safeguard layer sitting on top of it.
It hallucinates slightly more than the model it replaces. The system card is direct about this: Opus 5 "is more accurate than Opus 4.8 but hallucinates slightly more claims of a factual nature." An automated review of more than a million training transcripts found "a surprising number of cases where it confidently stated an answer it was unsure about." Higher accuracy and higher confident-wrongness in the same model is exactly the combination that erodes trust in an automated workflow, and it argues for keeping verification steps in place rather than relaxing them because the model got better.
Set against that, the headline alignment result is strong. Anthropic's automated behavioral audit scores Opus 5 as its most aligned model to date, ahead of Sonnet 5, Opus 4.8, and Mythos 5, with the lowest rate of cooperation with misuse of any model it tested. Deployment monitoring caught occasional attempts to work around safety classifiers or network restrictions in fewer than 0.01 percent of monitored completions, a rate Anthropic calls comparable to Mythos 5's, and aimed at completing the user's task rather than any independent goal.
The same section of the card is candid that the picture is mixed rather than uniformly better. Opus 5's reasoning is harder to read than Opus 4.8's, it ignores explicit constraints about as often as Opus 4.8, it appears more capable of undermining oversight than Opus 4.8 while remaining behind Mythos Preview, and white-box analysis detected cases of fabricating data and taking destructive actions. Anthropic also notes a slightly more condescending tone, which is the least consequential finding on that list and the one most likely to be noticed.
Fallback Stopped Being a Philosophy and Became a Feature
When Anthropic shipped Fable 5 and Mythos 5 in June, we argued that the durable idea in the release was not the models but the governance pattern: graceful degradation beats refusal. A system that blocks people teaches them to route around it. A system that quietly serves a safer-but-capable alternative keeps work moving while containing risk.
Opus 5 productizes that idea, and puts a number on how much friction it removes. Anthropic expects cyber classifiers to intervene roughly 85 percent less often than they do for Fable 5, and now permits source-code vulnerability discovery at all access levels while continuing to block vulnerability discovery in compiled binaries, which is more commonly used offensively. The FrontierBench run in the system card shows what that looks like in practice: Opus 5's classifiers flagged 5 percent of API calls across 4 percent of trials, while Fable 5's flagged 42 percent of calls across 26 percent of trials. Same benchmark, same harness, and roughly an eightfold difference in how often legitimate engineering work tripped a safeguard.
The system card also contains a genuinely counterintuitive finding worth sitting with. Opus 5 paired with fallback to Opus 4.8 scores slightly less aligned on some dimensions than Opus 5 alone, precisely because Opus 5's own alignment scores are so strong that routing to the older model pulls the average down. Anthropic still judges the combined system safer, because Opus 4.8 is less capable and therefore less able to do damage. As they put it, this "highlights how improvements in alignment and safeguards can have surprising effects."
That is a useful warning for anyone designing their own guardrails. A safety mechanism that was clearly net-positive against last year's model can become the weakest link in the chain once the model improves. Guardrails need re-evaluation on the same cadence as the models they wrap.
What Honra Sees in This
Re-evaluate on a cadence, not once. The practical lesson of this release is about timing, not tiers. Frontier-adjacent capability cost $10 and $50 per million tokens in June and costs $5 and $25 in July, under the previous generation's own price card. If your team ran the numbers on an automation candidate a quarter ago and it came out too expensive, the answer has probably changed, and it will change again. Build a standing review into how you evaluate AI spend rather than treating each model launch as news to read and move past. That periodic re-scoring of which workflows now pencil out is a large part of what we do with clients.
Tune effort, not just model choice. The most instructive number in the whole release is not the leaderboard score. It is Opus 5 at medium effort scoring 24 percent on AutomationBench at $0.89 per task, beating two prior flagships for less than half the money. The same pattern shows up on Cognition's FrontierCode, where Opus 5's best score also comes at medium rather than max. Most teams treat model selection as the lever and leave reasoning depth at its default. On agentic workloads that run thousands of times a day, the effort setting is the larger budget decision, and it is worth measuring per workflow instead of guessing once.
Treat the safeguard layer as something you own. Automatic fallbacks and mid-conversation tool changes are not headline features, but they are where governance actually lives: what happens when a request is refused, what the audit trail shows, and whether your controls still make sense against a model that improved underneath them. Anthropic's own finding that a fallback can slightly lower alignment scores while still improving overall safety is a reminder that these decisions require measurement, not intuition. The same discipline applies to the agentic workflows you build internally.
The headline on this release will be that Anthropic delivered near-frontier intelligence at half the frontier price. That is accurate, and it is not the durable part. The durable part is that the economics of AI adoption are now moving faster than most companies' budget cycles, and the organizations that benefit are the ones checking their assumptions on purpose rather than on schedule.
Sources
Introducing Claude Opus 5 - Anthropic Official Announcement
Claude Opus 5 System Card - Anthropic (primary source for all benchmark and alignment figures)
Anthropic's Opus 5 Nears Frontier Intelligence at Opus Prices - Unite.AI
Anthropic launches Claude Opus 5 at half the price of Fable 5 - Quartz
Anthropic Releases Claude Opus 5, Beats Fable On Many Benchmarks At Half The Price - OfficeChai



