Anthropic shipped Claude Opus 5 on July 24, and the headline number isn't a benchmark — it's the invoice. The company is claiming intelligence near its frontier Fable 5 model for half of what Fable costs, at the same $5-in, $25-out per-million-token pricing its predecessor Opus 4.8 charged. The model is live everywhere already: it's the new default on Claude Max and the strongest option on Claude Pro.
The numbers Anthropic is showing
On Frontier-Bench v0.1, a hard coding-and-agents test, Anthropic says Opus 5 now leads every model outright and more than doubles Opus 4.8's score — while costing less per task. CursorBench 3.2 tells a similar story: within half a percent of Fable 5's peak, at half the spend.
It's not just code, either. The company reports triple the next-best score on ARC-AGI 3, which is built around problems a model hasn't seen before, and about 1.5 times the nearest pass rate on Zapier's AutomationBench, a test of finishing real business tasks rather than starting them. On the OSWorld 2.0 computer-use benchmark, Anthropic claims Opus 5 beats everything at any price point and passes Fable 5's best result at roughly a third of the cost. There's also a new effort dial — low, medium, high — so you choose per request whether you're paying for deep reasoning or a quick answer.
Worth remembering: these are launch-week numbers from the company that built the model. Independent testing is what turns claims into a verdict.
The part that's actually interesting
The anecdotes say more than the charts. Asked to rebuild a machine part as a 3D FreeCAD model from a drawing it literally couldn't view, Opus 5 wrote its own computer-vision pipeline, pulled the geometry out of the raw pixels, and built the part. Anthropic says no rival model solved that one in five tries. Handed a real bug in a widely used open-source package manager, it found the root cause and caught an edge case the community's own patch had missed. One trading-firm engineer built a full market data feed in a single session — and when the model found no live feed to test against, it built its own test harness first.
Customers in the early-access program back this up with specifics: Zapier says the model went 100% on its automation benchmark, running a churn-prevention workflow end to end where earlier models fell over. Box clocked an 8% gain over Opus 4.8 on enterprise document work. Cursor, Cognition, Harvey, and JetBrains all describe the same thing in different words — steadier judgment across long tasks.
Safety, in plain terms
Anthropic's own behavioral audit calls this its most aligned model yet: a 2.3 on overall misaligned behavior, the lowest of its recent lineup, with the least deception and the most resistance to being talked into misuse.
The dual-use picture is deliberate. Opus 5 stays behind Mythos 5 on biology research and offensive cyber work, and Anthropic says it never trained the model on cyber tasks at all. It got better at them anyway — that's what general capability does — so on OSS-Fuzz it nearly matches Mythos 5 at spotting vulnerabilities while staying far behind at turning them into exploits. Safety classifiers should interrupt about 85% less often than they do on Fable 5, with flagged requests quietly falling back to Opus 4.8. And since blocked biology requests on Fable 5 now route here, Opus 5 becomes the most capable science model Anthropic offers to the general public.
Why this matters
Cheap frontier capability changes the deployment math. The question stops being "can we afford the best model?" and becomes "what effort level does this specific job need?" — which is exactly the routing logic engineering teams used to build by hand. A Fast mode runs about 2.5x quicker for double the price, and two API betas (mid-conversation tool swaps, automatic fallbacks) round out the release. What to watch next: independent benchmarks, and whether competitors answer on capability — or on price.
Image: Markus Spiske, via Pexels




