Claude Opus 5 is live on Napsix — near Fable 5, at half the price
Anthropic released Opus 5: new state-of-the-art on coding and knowledge work, at the same price as Opus 4.8. With 2 official videos, real customer stories, and already active in the AI Gateway.

Anthropic released Claude Opus 5 on July 24, 2026 — and it's already live on Napsix as the platform's default frontier model.
Why it matters
Opus 5 comes close to the intelligence of Claude Fable 5 (Anthropic's most capable model) at half the price. It's the new state-of-the-art on several coding and knowledge-work evaluations, and it beats Opus 4.8 on nearly every benchmark — at the same price. Anthropic describes it as "a thoughtful and proactive model" designed to be used every day: it works more efficiently than other models, it's the new default model on Claude Max, and the strongest model available on Claude Pro.

The numbers
| Evaluation | Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol |
|---|---|---|---|---|
| Frontier-Bench v0.1 (agentic coding) | 43.3% | 33.7% | 21.1% | 34.4% |
| GDPval-AA v2 (knowledge work) | 1861 | 1747 | 1593 | 1736 |
| ARC-AGI-3 (novel problem-solving) | 30.2% | — | 1.5% | 7.8% |
| Agentic search (BrowseComp) | 90.8% | 87.4% | 84.3% | 90.4% |
| Zapier AutomationBench (business workflows) | 26.0% | 17.4% | 17.0% | 18.1% |
| OSWorld 2.0 (computer use) | 70.6% | 66.1% | 55.7% | 62.6% |
| Legal Agent Benchmark (held-out) | 11.7% | 13.3% | 10.4% | 2.5% |
| HealthBench Professional | 59.8% | 66.0% (Mythos 5) | 57.4% | 60.5% |
| BioMysteryBench (hard) | 49.4% | 46.5% | 42.4% | — |
On ARC-AGI-3 (solving problems the model has never seen before), Opus 5 scores three times as high as the next-best model. On Zapier's AutomationBench, even at its lowest effort setting, Opus 5 completes more end-to-end business tasks than any other model at high effort — and on OSWorld 2.0 it surpasses Fable 5's best result at just over a third of the cost.

How Opus 5 works in practice
Anthropic and its early-access users documented several concrete examples of the model's agency and thoroughness:
- On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to rebuild it as a 3D FreeCAD model — but was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded repeatedly; no competing model with the same setup could solve it after five attempts.
- Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case the community's patch had missed. A competing model fixed only the surface symptom (not the underlying cause) and reported the bug resolved.
- An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even with extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code correctly parsed the exchange's data.
What early customers are saying
Anthropic shared testimonials from dozens of early-access teams. A few real examples:
"On FrontierCode 1.1, Claude Opus 5 approaches Fable-level performance at half the cost. Within Devin, it also shows particular strength on difficult debugging and root-cause analysis tasks." — Scott Wu, CEO of Devin
"Claude Opus 5 topped Zapier's AutomationBench leaderboard without spending more tokens than prior Claude models. It took a raw account-health workbook and ran a full churn-prevention sequence end to end. Previous models didn't pass; Opus 5 hit 100%." — Wade Foster, CEO of Zapier
"Claude Opus 5 is the biggest leap in the Opus family since 4.5. On the same full-stack app builds, the front end shows it first: the best animations, games, and 3D work we have seen from an Opus model." — Madhav Jha, Co-Founder and CTO of Lovable
"Claude Opus 5 checks its own work the way a real frontend developer would. On our benchmark it opened its pages in a browser at desktop and phone widths, caught a product hidden below the mobile fold and an off-screen checkout button, and fixed both before handing the work back." — AJ Orbach, Co-Founder and CEO
Science and visual outputs
Opus 5 improves on Opus 4.8 on every life-sciences evaluation: organic chemistry (+10.2 percentage points inferring molecular structures from spectroscopy data) and protein-related tasks (+7.7 points predicting how mutations affect function).
It also produces much stronger visual outputs. Anthropic released two interactive artifacts Opus 5 built on its own, along with the behind-the-scenes videos:
- 🎬 Watch on YouTube: "Claude Opus 5 builds a working wind tunnel" — Opus 5 visualized the flow of air over aerodynamic (and non-aerodynamic) objects. Try the interactive wind tunnel here.
- 🎬 Watch on YouTube: "Claude Opus 5 builds a 3D interactive animal cell" — Opus 5 built a simplified, interactive illustration of a cell. Explore its elements here.
Alignment and safety
On Anthropic's automated behavioral audit, Opus 5 is their most aligned model to date — 2.3 on overall misaligned behavior, the lowest of their recent models. It adheres to Claude's Constitution better than Opus 4.8, Sonnet 5, or Fable 5; exhibits the lowest rates of deceptive behavior; and is the least susceptible to being tricked into misuse.
On cybersecurity and biology, Opus 5 does not surpass Mythos 5 (Anthropic's highest-risk, restricted-access model) — Anthropic deliberately avoided training it on cyber tasks, though it still improved substantially from becoming more generally capable. Opus 5's cyber classifiers are ~85% less restrictive than Fable 5's: they allow finding vulnerabilities in source code, but block "binary-based" vulnerability scanning, penetration testing, and exploit generation. Any request flagged in Claude.ai, Claude Code, or Claude Cowork falls back to Opus 4.8 automatically.
Pricing and availability
$5 per million input tokens and $25 per million output tokens — the exact same price as Opus 4.8. A Fast mode is also available (~2.5× faster, at double the price, available on the Claude API). Alongside the release, Anthropic shipped two beta features: mid-conversation tool changes (without invalidating the prompt cache) and automatic fallbacks on the API (requests flagged by safety classifiers now route to another model automatically instead of being blocked).
Available on the Claude API, AWS Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.
On Napsix
We audited Opus 5 against Anthropic's official sources and it's already active in the platform's AI Gateway (routed directly through Anthropic, with automatic Bedrock fallback) — no action needed on your end. Pick it from the model selector in any XIA chat.