Anthropic released Claude Opus 5 on Friday, making it the default model for Claude Max subscribers and the strongest model available on Claude Pro.
The API remains $5 per million input tokens and $25 per million output tokens, the same rates as Opus 4.8. The model is also available through Amazon, Google and Microsoft cloud platforms.
Why it matters: The list price did not change, but the model's operating behavior did.
Opus 5 turns adaptive thinking on by default. Anthropic also says its responses and written deliverables run longer than earlier Opus models. Developers moving from Opus 4.8 should retest token limits, latency and real costs.
The model supports a 1 million-token context window, 128,000 output tokens and five effort settings.
There is a breaking change at xhigh and max: Requests that disable thinking return an error. Thinking can still be disabled at high or below.
Anthropic's performance claims: The company says Opus 5 approaches Fable 5 on some coding and professional-work tests. It reports substantial gains over Opus 4.8 in agentic coding, computer use and long-running knowledge work.
Those claims come from Anthropic's launch materials and selected evaluations. Still in the Loop has not independently reproduced them.
Early outside use: Dan Shipper said Every spent a week testing Opus 5 across coding, writing, knowledge work and an internal agent. He reported that it argued with instructions, stopped early and clashed with existing skills and plugins.
The safety picture: Anthropic's system card rates overall alignment risk as very low. Its testing found stronger cyber capabilities than Opus 4.8, but weaker exploit capabilities than Mythos 5.
Anthropic's Boris Cherny highlighted prompt-injection resistance. The card reports gains over Opus 4.8 across coding, computer-use and browser tests, but says the Opus 5 bug bounty is not complete.
A UK AI Security Institute evaluation found no unprompted attempts to sabotage safety research and a low rate of continuing sabotage. Research scientist Robert Kirk cautioned that Opus 5 scored higher than earlier models at recognizing evaluations when prompted. The system card says that limits how confidently the results predict deployment behavior.
