A glass-scale reader on a small mill shows whether a model is careful. The scale sends a short serial string, the sign lives in one character, and a wrong minus sends the tool the wrong way. On July 24, 2026, Anthropic released Claude Opus 5 for that kind of daily work. It is available the same day, at $5 per million input tokens and $25 per million output tokens, the same list price as Claude Opus 4.8. A token is a chunk of text the model counts.
Anthropic says Opus 5 comes close to the frontier intelligence of Claude Fable 5 at half the price, and that it is a large step up at Opus 4.8's cost. It is the new default on Claude Max and the strongest model on Claude Pro. The post does not put Fable 5 on Pro. I have not run Opus 5 for this article. The scores below are Anthropic's.
What actually changed?
Opus 5 is the new Opus, priced like the old one. Anthropic says it works more efficiently than other models. The control is effort: a setting for how hard the model works. Higher effort spends more tokens. Lower effort conserves tokens for a faster, cheaper result. The prose names max, high, xhigh, and the lowest setting.
On CursorBench 3.2, Anthropic says max-effort Opus 5 lands within 0.5 percent of Fable 5's peak at half the cost per task. The post does not print Fable 5's token price, so half the price in the opening line and half the cost per task are both their wording, not a price I derived. On Frontier-Bench and GDPval-AA, Anthropic calls Opus 5 the new state of the art, and says it remains behind Mythos 5 on cybersecurity. An agent here is a setup that lets the model work for many steps, use tools, and check its own result.
How does the new piece work?
Anthropic stresses checking. One Frontier-Bench task gave the model a machine-part drawing and asked for a FreeCAD rebuild, with no direct way to view the drawing. FreeCAD is a free mechanical CAD program. Anthropic says Opus 5 wrote a vision pipeline, pulled geometry from the pixels, and rebuilt the part repeatedly. No competing model with that setup solved it in five attempts, Anthropic says. A Cursor note on the page says, "On CursorBench it's just under Fable 5 and has many of the same behaviors."
Even at the lowest effort, Anthropic says Opus 5 passes more Zapier AutomationBench tasks than any other model. That bench checks whether a model finishes a business task end to end. On a shop job that is unsafe if wrong, I would set effort high. On a changelog, I would set it low, and I would still read the diff.

What does this look like on a real project?
The mill reader is the project I would use. I did not run it on publication day. A small Python script opens a serial port, parses a scale packet, and prints millimeters. The bug I already understand is an off-by-one on the sign when the reading is negative. A "fix" that takes the absolute value looks clean and will crash a tool.
I would open Claude Code on Opus 5. The API name is claude-opus-5. Max defaults to it. Pro offers it as the strongest model on that plan, which is all the post claims. I would set effort high and order the work: write a test from a captured log with a positive, a negative, and a zero; then change the parser; then show the test failing on the old code. I would read the patch before it touched the machine. A README edit would get low effort. Fast mode stays off for a bug this small.

How does it compare with the previous version?
Opus 4.8 is the predecessor, at the same $5 and $25 prices. Anthropic says Opus 5 more than doubles Opus 4.8 on Frontier-Bench v0.1 at a lower cost per task. The footnote says that run used the mini-SWE-agent harness on Google Kubernetes Engine, with mean reward over five tries. A harness is the wrapper of tools and instructions around the model. If a safety classifier refused Opus 5 or Fable 5, Opus 4.8 answered instead, so the double is not a pure single-model number on every task.
On ARC-AGI 3, a test of novel problems, Anthropic says Opus 5's score is three times the next-best model. On Zapier AutomationBench, the pass rate is around 1.5 times the next-best model at the same cost per task. On OSWorld 2.0, a computer-use benchmark, Opus 5 beats every other model at a given cost and passes Fable 5's best result at just over a third of the cost. Computer use means the model drives software on a screen.
For science, Anthropic says Opus 5 beats Opus 4.8 on every life-sciences evaluation they list: 10.2 percentage points on an internal spectroscopy-structure benchmark, and 7.7 points on protein-sequence tasks about how changes affect function. Those are scores, not a procedure.
Where does it sit next to other tools a maker already uses?
On Claude Max, Opus 5 is the default. On Claude Pro it is the strongest model on the plan. The API name is claude-opus-5. Claude Code is how I would use it on a repo, and I would still run the tests myself. Claude Cowork uses the same safety fallback.
Mythos 5 is ahead on cybersecurity, by Anthropic's account. Fable 5 is the closer intelligence comparison, at the higher price they describe. Opus 4.8 is the price twin, and it is the fallback when a classifier flags a request in Claude.ai, Claude Code, or Claude Cowork. A firmware question can trip that. This article does not try to get around it.
Two API updates are in beta. You can change tools mid-conversation without clearing the prompt cache, the stored prefix that keeps you from paying full price to resend it. Flagged Opus 5 or Fable 5 calls can fall back to another model, which Anthropic says is the best available one by default. Read what that means on your account before a release.
What does it cost, and who can use it today?
Base price is $5 per million input tokens and $25 per million output tokens. Fast mode is about 2.5 times the speed at twice that base price, on the Claude Platform and through Claude Code usage credits. The post has no separate Fast price card.
On July 24 it is on all platforms. Max gets the default. Pro gets the strongest model on that plan. Anthropic says general access has no data-retention requirements, consistent with earlier Opus models. Effort has no separate sticker. It changes how many tokens you buy.
What is still unproven?
I did not reproduce these benchmarks.
Anthropic's behavioral audit scores Opus 5 at 2.3 on overall misaligned behavior, lowest among its recent models. It follows Claude's Constitution more closely than Opus 4.8, Sonnet 5, or Fable 5, with less deception, and it is the least likely of that set to take reckless hard-to-reverse actions.
Anthropic says Opus 5 does not move the dual-use frontier. It stays behind Mythos 5 on biology research and offensive cybersecurity. It was not trained on cyber tasks. Broader gains still improved that skill. Finding vulnerabilities is close to Mythos 5. Developing exploits is substantially behind. On OSS-Fuzz, identification is similar and exploit development is far behind. Safeguards allow finding vulnerabilities in source code, and they block binary vulnerability scanning, penetration testing, and exploit generation. Classifiers should fire about 85 percent less often than on Fable 5. Flagged chats fall back to Opus 4.8. The Cyber Verification Program gives members a less restricted copy. Biology requests blocked on Fable 5 now route to Opus 5. Long autonomous research is still weaker here than on Mythos 5.
Disclosure: The author is a paying subscriber to ChatGPT Plus, Claude Pro, and SuperGrok and uses all three services on a daily basis. The Makers Workbench is not affiliated with OpenAI, Anthropic, xAI, Google, or any of the other major AI companies covered in our reporting. No company receives favorable editorial treatment based on the author's personal subscriptions.
Sources and image credits
- Anthropic announcement (www.anthropic.com)
- Images: published by Anthropic with the announcement.
