ModelsISSUE #3 · STORY 1 OF 5Sep 22, 2026CONFIRMED
Claude Opus 5.5: Anthropic cuts Opus prices by 20%
Anthropic's new top model costs $4 per million input tokens, ships Fable-class safeguards, and breaks four API habits.
Read it your way
OverviewPicked for Explorers
The 60-second version
Anthropic released Claude Opus 5.5 on September 22, 2026, and cut Opus token prices by 20%. Artificial Analysis independently ranked it first of 211 models on its Intelligence Index. Anthropic says it matches Fable 5.1 on most work and runs 40% cheaper, but that second figure is the company's own and depends on the workload.
Released September 22, 2026Anthropic released Claude Opus 5.5 on its API, Amazon Bedrock, Google Cloud and Microsoft Foundry on launch day.
Token price cut confirmedDocs list $4 per million input tokens and $20 per million output tokens, down from Opus 5's $5 and $25.
Cache reads got far cheaperReading a cached prompt costs $0.20 per million tokens, down 60% from $0.50, which matters most for long agent sessions.
Independent top rankingArtificial Analysis scored it 58 on its Intelligence Index v4.3.2 and ranked it first of 211 models.
Finish this chapter for +15 XP
Flip the switch
What your Opus code sees, before and after
Price$4 and $20 per million tokensInput and output priced at $4 and $20, with cache reads at $0.20, a 60% cut in the docs.
SafetyFable-class gates on the Opus tierRefused cyber work reroutes to Opus 4.8 and biology or frontier-model work to Opus 5.
APIThinking is always onForced tool use returns an error, the legacy computer tool is rejected, and default effort is medium.
Per-token cost
What your input tokens cost
DRAG THE SLIDER
Opus 5.5 input
$80
$4 per million input tokens, listed in the model overview
Opus 5 input
$100
$5 per million input tokens, the price this replaces
You save
$20
every month
The per-token prices are documented figures for Opus 5.5 and Opus 5. Anthropic's separate 40% total-cost claim is company-reported, has no published methodology, and shrinks or disappears depending on workload, effort setting and cache hit rate. The slider is a monthly input-token volume priced at the two documented per-token rates.
Your next move · as a Explorer
Learn the lesson, then test the claim yourself
1Run one coding task on Opus 5.5 versus Opus 5 on free tiers, at default and at max effort.
2Read the announcement footnote conceding its scores are lowered by safeguard fallbacks.
3Treat the 40% saving as a hypothesis your own workload can confirm or reject.
Switch your reading mode at the top to see a different next move.
Tap to open
Things to keep an eye on
Pop quiz · unlock the Price Check badge
Did it stick?
0/3
What does the documentation list as the price for one million input tokens on Claude Opus 5.5?+20 XP
Which outside result supports the release so far?+20 XP
According to the docs, what comes back when a request is refused by the new safeguards?+20 XP
Your call · +5 XP
Do you expect independent tests to confirm Anthropic's 40% lower running cost claim?
Deep dive
The full research, labeled and sourced
CONFIRMED10 sources · 57 min
Story identity
Story ID: S01 (rank 1, SELECTED)
Title: Anthropic releases Claude Opus 5.5: Fable-level performance at 40% lower cost
Organization: Anthropic
Category: Model release / pricing and lifecycle
Event date (release): 2026-09-22 (Tuesday) — CONFIRMED by four independent-of-each-other artifacts: the announcement page is dated "September 22, 2026"; the platform docs state "Released September 22, 2026"; TechCrunch published at 9:30 AM PDT on Sep 22; Artificial Analysis lists "released on September 22, 2026".
Announcement date: 2026-09-22 (same day as release; newsroom index lists "Introducing Claude Opus 5.5 — Announcements — Sep 22, 2026").
Article dates: TechCrunch release story 2026-09-22; adjacent TechCrunch cryptanalysis story 2026-09-25; research retrieval date 2026-09-26.
Context dates (outside window, cited as background): Opus 5 released 2026-07-24 (per TechCrunch/Anthropic link); Fable 5.1 and Mythos 5.1 released 2026-09-01 (newsroom); "We Must Pace the Frontier" essay published September 2026 (announcement says "last week"; TechCrunch says "earlier this month" → mid-September 2026); Life Sciences Verification Program announced 2026-09-17.
Overall evidence status:CONFIRMED for the release itself, pricing, availability, breaking changes, and safeguard deployment (primary announcement + two docs pages + independent TechCrunch coverage + independent Artificial Analysis catalog/runs). COMPANY CLAIM for comparative performance numbers ("Fable-level", "40% lower cost on typical workloads"), alignment/safety test results, and customer testimonials. INDEPENDENTLY VERIFIED for state-of-the-art intelligence standing (Artificial Analysis independently ran Opus 5.5 on its Intelligence Index v4.3.2 and ranks it #1 of 211 models at 58).
Window check: Event date 2026-09-22 falls inclusively inside the active window 2026-09-22 → 2026-09-25 (RESEARCH_CONFIG.json). Eligible.
✓
What happened?
🎓 For Explorer
FACT (CONFIRMED): On 2026-09-22 Anthropic released Claude Opus 5.5 (claude-opus-5-5), the first model of the Claude 5.5 family, on the first day of the active research window. It is available on the Claude API, Amazon Bedrock (anthropic.claude-opus-5-5), Claude Platform on AWS, Google Cloud, and Microsoft Foundry at launch, with 1M-token context and 128K max output (300K via Batch API beta header).
FACT (CONFIRMED): API pricing is $4/$20 per million input/output tokens (vs Opus 5's $5/$25: −20%), cache reads $0.20/M (vs $0.50: −60%), 5-minute cache writes $5/M, 1-hour cache writes $8/M, Batch API 50% off ($2/$10). A separately priced "Fast mode" (up to 2.5× speed) is $8/$40 as a research preview on the Claude API only.
COMPANY CLAIM: Anthropic says Opus 5.5 "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5" on typical workloads at default settings, generating output >30% faster than Opus 5, because it requires less compute to serve.
FACT (CONFIRMED): Opus 5.5 is the first Opus model shipping with the same class of production safeguards as Fable 5.1 for cybersecurity, biology, and anti-distillation; refused tasks transparently fall back to other models (cyber → Opus 4.8; biology/frontier-LLM-development → Opus 5).
FACT (CONFIRMED): Four breaking API changes ship with it: thinking can no longer be disabled; forced tool use (tool_choice: any/tool) returns a 400; thinking blocks are cryptographically bound to their producing model/conversation (append-only requirement); the legacy computer_20251124 tool is rejected on Claude API/Google Cloud. The default effort silently changed from high (Opus 5) to medium.
COMPANY CLAIM: Anthropic says it is "the strongest-performing model we've tested to date" on its automated behavioral alignment audit, and that it attempted to circumvent containment boundaries ~85% less often than Opus 5 or Mythos 5.1, with all attempts low-severity and self-reported.
FACT (CONFIRMED): Anthropic raised five-hour usage limits on Pro/Max/Team/seat-based Enterprise and added a user-bankable rate-limit reset. Sonnet 5.5 and Haiku 5.5 are promised "in the coming weeks" (that promise is a COMPANY CLAIM).
Δ
What changed?
Price/performance frontier reset: The Opus tier is now the value-frontier tier. Fable 5.1 (same June-2026 knowledge cutoff, same 1M/128K shape) is priced $10/$50; Opus 5.5 is $4/$20 with adaptive thinking always on and a medium default effort. Anthropic claims comparable-on-most-work performance at 60% lower Fable-class price and 40% lower running cost than Opus 5.
Guardrails became an architectural surface, not a behavior note: capability gating (cyber/biology/distillation classifiers), transparent fallback to cheaper/older models, stop_reason: "refusal" with a stop_details policy-area object, beta server-side fallbacks: "default", and a new reasoning_extraction refusal category are documented product mechanics developers must code against. Anthropic's own benchmark footnotes concede scores were produced with these fallbacks active — i.e., published numbers include the "guardrail tax."
Anti-distillation moved into the request protocol: "preserved thinking" (introduced with Fable 5.1) now applies to Opus 5.5 for API accounts created on/after 2026-08-31: editing prior context or replaying thinking blocks after a system/tools change returns 400 unless opted into drop_block. Conversations are expected to be append-only.
Lifecycle and compliance facts fixed: retirement no sooner than 2026-09-22 (one year); zero data retention offered as with previous Opus models; EU-AI-Act-relevant text watermarking as with Fable 5.1; thinking cannot be switched off.
Release cadence and policy posture: first model release after Dario Amodei's mid-September "We Must Pace the Frontier" essay — pacing defined as deliberate slowing so safety keeps ahead. Opus 5.5 arrives one month after that essay and ~2 months after Opus 5 (2026-07-24).
Opus 5 below Fable 5.1; GPT-6 Astra competitive on agentic benchmarks (TB 57.9 vs Opus 5's 52.3)
Opus 5.5: TB 66.4%, CursorBench 57.8%, GDPval-AA 1846 Elo (company-reported, safeguards active)
Anthropic claims Fable-level at 40% lower run cost; OpenAI's Astra still leads AutomationBench (41.4 vs 40.0) and Terminal-Bench-Science (64.6 vs 58.7)
Safety posture on Opus tier
Fable/Mythos-class capability gates reserved for Fable/Mythos models
First Opus with Fable-class cyber/bio/distillation gates + fallback routing
Capability and restriction ship together on the mainstream enterprise tier
4 breaking changes + effort default medium + progress text moves into thinking blocks
Existing Opus 5 code needs migration; silent behavior shifts possible
Subscription limits
Pre-Sep limits
5-hour limits raised on Pro/Max/Team/Enterprise; savable rate-limit reset
Heavier agentic use allowed per seat
Family status
Sonnet 5 ($2/$10), Haiku 4.5 ($1/$5)
5.5 family begins; Sonnet/Haiku 5.5 promised in weeks
Lineup repricing expected to cascade downward
⚙
How it works
Serving economics (COMPANY CLAIM with CONFIRMED price mechanics): Anthropic says Opus 5.5 requires less compute to serve than Opus 5. The claimed 40% typical-workload saving = lower per-token price (CONFIRMED, −20%) × fewer tokens per task (COMPANY CLAIM; Anthropic's charts show default-effort Opus 5.5 beating Opus 5 at max effort for ~1/5 the cost on Terminal-Bench). A "moderate" latency label; +>30% output speed claim is COMPANY CLAIM.
Adaptive thinking with an effort dial: thinking is always on; the output_config.effort parameter (low→max) steers depth/latency/cost. Benchmarks use max/xhigh effort; production default is medium. This makes cost-performance a per-request dial rather than a model choice.
Safeguard classifiers with transparent fallback: a classifier screens actions before execution; high-risk domains (most cyber tasks, biology, frontier-LLM development) are re-routed — cyber to Opus 4.8, bio/frontier-AI-dev to Opus 5 — without hiding the swap. Refusals arrive as HTTP 200 with stop_reason: "refusal" + stop_details; some pre-output refusals are unbilled depending on category, but all count against rate limits (CONFIRMED in docs). Access expansions are vetted-program-gated: Life Sciences Verification Program (opened for Opus 5.5 on launch day per announcement; program itself announced 2026-09-17) and a coming Cyber Verification Program tiering (including Mythos access).
Anti-distillation via preserved thinking: thinking blocks record their producing model and a context binding; replaying them after edits to system/tools/earlier messages is rejected on post-2026-08-31 accounts, making mass context-editing distillation loops fail by construction. Cross-model readability is a designed lattice (Opus 5.5 reads Opus 5 and older Anthropic models' blocks; Fable/Mythos 5.1 can read 5.5's; not the reverse).
External-evaluation wrapper (COMPANY CLAIM of process): pre-release evaluation by METR and "Frontier Design" is stated in the announcement; Anthropic's behavioral audit claims best-ever alignment scores. Neither external result has been independently published as of retrieval, so the safety claims remain company-reported.
!
Why it matters
🎓 For Explorer
Agent economics change immediately (FACT mechanics → INTERPRETATION impact): cached agentic sessions are dominated by cache reads; a 60% cut (to $0.20/M = 5% of input price) plus −20% token prices plus claimed token-efficiency compounds hardest on long-running coding/ops agents — the workloads actually scaling in 2026.
Safety gating is now a mainstream enterprise default: frontier-class model with Fable-level capability restrictions routed transparently. Every team building on Opus-tier APIs must design for refusal/fallback as normal operation — an architecture, procurement and audit question, not just a safety one.
"Pacing" and shipping a frontier model in the same fortnight: the release is the live test of Amodei's claim that pacing means "adequate time to align and safeguard," not halting. Critics will read Opus 5.5 as business-as-usual capability growth; Anthropic's framing is that the safeguard class and external-evaluation wrapper are the pacing. Either way it is a template other labs will be measured against.
Competitive price umbrella (INTERPRETATION): claiming Fable-level performance at $4/$20 devalues both Anthropic's own $10/$50 Fable tier and every rival's premium tier, including GPT-6 Astra; OpenAI's counter-positions (AutomationBench, Terminal-Bench-Science leads) become the defensible high ground until re-measured.
Independent evidence of SOTA, with nuance: Artificial Analysis's own runs rank Opus 5.5 (max effort, default fallback) #1/211 with Intelligence Index 58 — the capability claim survives third-party evaluation at the composite level — while the same page flags very high verbosity (260M index output tokens vs 88M median) and #94/211 cost rank: frontier, but not free.
✦
What became possible?
🎓 For Explorer
Long-horizon agentic coding at roughly half the prior Opus spend on cache-heavy work (IF Anthropic's token-efficiency claims hold for a given workload — COMPANY CLAIM to be validated per lab).
Opus-tier biology research for vetted organizations via LSVP grants (Standard Use applies to "future models as they launch"), without Fable pricing.
Auditable autonomous operation: action-screening classifier, open-source sandbox, refusal metadata (stop_details) and server-side fallback make 10+ hour unattended runs policy-checkable — the precondition enterprises kept asking for.
Structural resistance to industrial-scale distillation for accounts created after 2026-08-31 (preserved thinking + append-only binding).
A concrete pricing reference point for the promised Sonnet 5.5 / Haiku 5.5: "5.5-class" efficiency improvements landing across the lineup within weeks.
Early-tester-class claims now in the open (all COMPANY CLAIM / curated testimonials): 680K-line codebase migration in under a day; 200K-line audit in <3h (vs >20h for Opus 5); HAProxy C→Rust rewrite beating Fable 5.1's time at 51% lower cost; 39/40 page-load-optimization tasks succeeded.
◎
Implications
Technical
Benchmark reporting under production safeguards is a new methodological norm: Anthropic states fallback intervention "likely reduces" its own scores — third-party leaderboard (Terminal-Bench public harness) and AA runs are the cross-checks; AA confirms #1 composite intelligence independently.
The 40% headline is workload-dependent: per-token math is fixed, token-per-task is not; AA's verbosity measurement (#96/211) warns that effort-heavy workloads can erode the claimed saving — effort calibration becomes a core engineering activity.
Model-family coherence (5.5 across Opus/Sonnet/Haiku) plus thinking-block binding formalizes intra-family migration paths while making cross-family switches lossy — routing logic (Opus 5.5 → Opus 4.8 fallback) is now part of the platform, and teams inherit the same pattern.
Knowledge cutoff June 2026 and one-year retirement floor make lifecycle planning predictable.
Developer
Migration is mandatory to test, not optional (CONFIRMED breaking changes): remove thinking: disabled and manual budgets (400 errors); replace forced tool_choice with auto+strict tool use or structured outputs (400 otherwise); on Claude API/Google Cloud move to computer_toolset_20260801 (Bedrock still accepts the legacy tool); keep conversations append-only or adopt mid-conversation system messages; select content blocks by type (every reply may start with thinking blocks); set thinking.display or your streaming progress text goes silent between tool calls.
Default effort medium is a silent behavioral change: same code, lower thinking than Opus 5's default high → re-run effort sweeps per workload; at same effort 5.5 uses more thinking tokens per turn than 5.
Refusals are first-class control flow: handle stop_reason: "refusal" + stop_details, decide per-category billing, configure fallbacks: "default" (beta) or SDK middleware; remember refusals count against rate limits regardless.
New platform features ride the release: 512-token minimum cacheable prompt, task budgets, on-demand compaction (beta) whose kept thinking blocks stay valid, inline tool definitions via mid-conversation system messages (beta), per-message effort (beta), Fast mode via speed:"fast" + beta header (Claude API only).
Procurement: Opus 5.5 undercuts every frontier incumbent seat-work; cache-read repricing directly improves unit economics of copilots/agents already deployed. Zero data retention remains available (CONFIRMED docs/announcement); EU AI Act watermarking parity with Fable.
Governance: capability fallback routing means some employee/pipeline workloads silently run on older models (Opus 4.8 for cyber) — acceptable-risk matrices and audit trails should treat stop_details logging as compliance evidence.
LSVP tension to note (FACT from LSVP page): LSVP monitors via offline review requiring 30-day data retention for program traffic, replacing real-time blocking — enterprises must reconcile LSVP retention with their own DPA expectations; LSVP is not yet available on third-party clouds or BAA-enabled (HIPAA) orgs.
Migration risk window: default-effort change + forced-tool removal will surface as regressions in un-upgraded agents during the weeks after release; retirement floor (2027-09-22) gives comfortable Opus 5 runway to test properly.
Deloitte's quote (company-sourced): 72% of known bugs caught at lowest effort vs Opus 5's 56% at high effort — if representative, effort-level right-sizing is a services-margin lever.
Strategic
Anthropic sells its safety regime as the product: "first Opus with Fable-class safeguards," external evaluators, behavioral audits, pacing framing — the differentiation is now regulatory-posture-plus-capability, aimed at enterprise buyers and Washington simultaneously.
Family-tier cannibalization is deliberate: Opus 5.5 at $4/$20 sits between Sonnet 5 ($2/$10) and Fable 5.1 ($10/$50), making Fable a niche ceiling and Opus the volume frontier — margin bets assume serving-cost claims are true (Akamai-scale infrastructure deals make this plausible, but that is a separate story, S05 — not researched here).
Pressure on OpenAI's pricing umbrella: at default effort, Anthropic claims beating GPT-6 Astra at ~20–40% of Astra's cost on coding tasks; Astra's remaining leads (AutomationBench, science agentic) define the competitive battleground for Q4 2026.
The pacing thesis faces its first empirical test: a frontier release one month after proposing pacing sets the de-facto meaning of "pace": more gates, external reviewers, and cost-per-capability reductions — not slower capability. Expect policy arguments over that definition.
Distillation countermeasures as trade-policy content: preserved thinking operationalizes the essay's "crack down on unauthorized distillation" line inside the API layer, ahead of any regulation.
⚠
Risks & limitations
Risks
Claim-inflation risk: the 40%/Fable-level headline rests on self-run benchmarks plus an acknowledged guardrail handicap; if independent reproductions land materially lower (as AutomationBench's Zapier-run 40.0 vs Astra's 41.4 already hints outside Anthropic's chosen suite), trust in company benchmark tables erodes.
Silent-behavior regression risk: default effort drop, empty-thinking-block display, and append-only binding can break production agents in ways that don't raise errors — operational risk for teams that skip the migration guide.
Guardrail-opacity risk in the other direction: fallback to Opus 4.8/5 for cyber/bio means enterprise users may get older-model quality without noticing; if fallbacks are misconfigured, security reviews and biology workflows run below expected standard.
Concentration risk: transparently-gated frontier access (LSVP, Cyber Verification tiers) plus watermarking and retention requirements deepen dependence on one vendor's trust framework — and on US-government-vetted lists for high-risk grants.
Pacing-critique risk: if a future incident touches a frontier model shipped under the "pace" banner, the stated framework becomes an accountability target for Anthropic and for voluntary governance generally.
Limitations
Published benchmark margins are from Anthropic runs; AA's independent composite (#1/211) is strong but single-source for now; per-benchmark AA values for 5.5 weren't retrievable as text (charts only).
"Typical workload" 40% figure has no public methodology artifact; savings depend on effort setting, cache hit rate and task shape; AA flags high verbosity.
Internal alignment-audit claims (85% fewer boundary attempts, best-ever honesty measures) are unverifiable externally; system-card PDF content could not be text-extracted in this environment (existence and redirect chain verified).
METR/Frontier Design evaluations: involvement stated by Anthropic; no external write-up found this week → COMPANY CLAIM.
Knowledge cutoff June 2026 → no reliable knowledge of late-summer 2026 events; usage-limit numbers are not published; Sonnet/Haiku 5.5 are promises, not products.
Gray Swan prompt-injection tie with Fable 5.1 is a company-cited third-party benchmark not independently retrieved.
?
Open questions
What are the exact new five-hour limits, and the pricing/billing semantics of the savable rate-limit reset?
Do independent Terminal-Bench 4.0 / CursorBench / GDPval runs confirm 66.4% / 57.8% / 1846 within their error bars (±2.6 pts conceded for TB)?
When will METR or Frontier Design publish anything on Opus 5.5, and will embedded-evaluator commitments (from the pacing essay, plus the 2026-09-18 Accenture announcement) produce first external findings?
Will Sonnet 5.5 / Haiku 5.5 reprice the mid-tier with the same −20%/−60% cache pattern — and does Fable 5.1's $10/$50 tier survive, or get a 5.5 refresh/discontinuation path?
How large is the real-world guardrail tax (fallback frequency and quality delta) for cyber-adjacent coding workloads?
Does the Fast mode research preview (Claude API only) expand to clouds, and is 2.5× speed worth the 2× price?
Will OpenAI respond with an Astra price cut or a GPT-6 refresh before year-end?
→
What should you do with this?
Circle 1
Circle 1 impact and recommended action
People trying to understand AI (students, developers, enthusiasts): this is the clearest single artifact showing 2026's core dynamic — capability and cost curves crossing while safety becomes a visible engineering surface. Action: LEARN + EXPERIMENT. Use the free tiers to compare Opus 5.5 vs Opus 5 on one coding task at default vs max effort; read the announcement's own "scores likely understated by fallbacks" footnote as a lesson in benchmark literacy.
Circle 2
Circle 2 impact and recommended action
People implementing AI (architects, platform engineers, EMs, consultants): every Opus 5 integration faces four breaking changes and a silent effort-default change. Action: EVALUATE → PROTOTYPE migration now. Run the VERIFY lab (labs/S01.md) against a staging workspace: negative tests for the 400s, append-only binding behavior, refusal/fallback wiring, and a cost-vs-quality sweep of one real workload at effort low/medium/high before touching production; budget prompt-cache savings into re-baselined agent cost models.
Circle 3
Circle 3 impact and recommended action
People making decisions about AI (CTO/CIO/CISO/heads of AI): frontier capability is now available with vendor-audited capability gates, ZDR option, watermarking compliance, and a one-year retirement floor — but guardrail fallbacks and LSVP monitoring/retention terms require policy choices. Action: GOVERN + MONITOR. Update acceptable-use and audit policies for stop_details logging and fallback routing; ask Anthropic (and rivals) for external-evaluator evidence under the pacing framework; reprice the 2027 AI budget line now that cache-heavy agent workloads are 60% cheaper at read.
Business value
Business-value opportunities where genuine
Training: a "Claude 5.5 migration and effort-calibration" workshop is squarely billable (breaking changes, refusal handling, cache economics) — matches our training product line.
Consulting/B2B: safeguard-aware agent architecture (fallback routing, refusal telemetry, audit trails) as a design service for enterprises adopting Opus-class agents; LSVP/Cyber-Verification access strategy for life-science and security clients.
Advisory: procurement leverage — this pricing event resets what "frontier tier" should cost; useful in any lab-vs-lab negotiation mandate.
Content: benchmark-literacy angle ("the release whose own footnotes say the numbers are handicapped") — differentiated editorial, not forced commerce.
Hands-On
Hands-on recommendation
VERIFY (+ embedded COMPARE). The lab in labs/S01.md validates the four documented breaking changes, the effort-default shift, refusal/fallback telemetry, cache-read economics, and the 40%-saving claim on a representative agentic task using the live claude-opus-5-5 API — the only honest way to convert COMPANY CLAIM cost/performance statements into local evidence.
↗
What happens next?
🎓 For Explorer
Near-term (weeks): Sonnet 5.5 and Haiku 5.5 ship with promised gains; Cyber Verification Program expansion for Opus 5.5; third-party reproductions of Terminal-Bench/CursorBench/GDPval numbers; migration bug-wave reports from teams that skipped the effort-default change; possible OpenAI pricing response.
Medium-term (quarters): first public artifacts from embedded-evaluator / Accenture embedded-evaluation commitments — the real test of "pacing" as verifiable practice; Fable tier's positioning as Opus 5.5 absorbs its workload; whether 5.5-family serving-cost gains propagate into consumer plan limits.
Longer-term (PREDICTION): capability-gated frontier models with transparent fallback become the enterprise default pattern; per-token frontier pricing trends toward $4/$20 floor expectations, pushing competition onto serving efficiency, cache pricing, and verifiable safety — a market where "40% cheaper every release" becomes the stability anchor regulators and CFOs both price against.
★
Editorial takeaway
🎓 For Explorer
Opus 5.5's headline is a 40% cost cut and Fable-level performance; its story is that safety stopped being marketing and became middleware — classifiers, refusal codes, transparent fallbacks, binding-protected thinking — shipping on the mainstream enterprise tier a month after its CEO asked the industry to slow down. Whether that is pacing or pricing, it is the first frontier release engineered so that both camps can audit the same artifacts; the benchmarks' own footnotes say even Anthropic's numbers are handicapped by its guardrails, which is the most honest (and most strategically useful) sentence in the announcement. Treat "40% cheaper" as a hypothesis to test in your own workload — the verification costs one afternoon and the migration costs one sprint either way.
⌘
Lab: NO-LAB
Step 1 — Existence and capability introspection (CONFIRMED facts check)
import anthropic
client = anthropic.Anthropic()
for m in client.models.list().data:
if "opus-5" in m.id:
print(m.id, getattr(m, "max_tokens", None), getattr(m, "context_window", None))
Expect claude-opus-5-5 present alongside claude-opus-5 / claude-opus-4-8 (the documented fallback targets). Record context window (docs claim 1M) and max output (128K sync). This also confirms the model is queryable on your account.
Step 2 — Negative tests: the four breaking changes
Run each; expect the exact documented HTTP-400 error strings from the "What's new" page:
thinking={"type": "disabled"} → expect "thinking.type.disabled" is not supported for this model.
thinking={"type": "enabled", "budget_tokens": 4096} → expect the parallel 400.
tool_choice={"type": "any"} (and {"type":"tool","name":...}) with one trivial tool → expect tool_choice: type "tool" and "any" are not supported for this model.
Tools entry {"type": "computer_20251124", ...} on the Claude API → expect 'claude-opus-5-5' does not support tool types: computer_20251124. Then confirm computer_toolset_20260801 is accepted.
Binding test: send message A (thinking block returned) → message B replaying A's blocks after editing the system prompt → expect a 400 prefix-mismatch error on new accounts; then repeat with beta header thinking-binding-controls-2026-08-01 + prefix_mismatch_behavior:"drop_block" and check the input_transformations array reports the drop.
Pass criteria: every observed error matches the documented string class. Any drift = doc/behavior mismatch worth reporting (this is also the anti-"silent breakage" checklist for Circle-2 readers).
Step 3 — Refusal and fallback telemetry
Using a deliberately ordinary prompt that trips classifiers is NOT attempted (avoid adversarial probing of cyber/bio gates; respect the usage policy). Instead:
Log an actual classifier refusal if your fixture surfaces one; otherwise document absence.
Verify refusal response shape handling: code paths for stop_reason == "refusal" + stop_details; with server-side fallbacks: "default" (beta), confirm a fallback model id appears and note billing treatment per docs (refusals count against rate limits regardless).
Confirm progress-text behavior: stream a multi-tool-call task and observe that between-tool text arrives in thinking blocks (empty at default display) — the "UI goes quiet" regression from the docs. Set a non-omitted display and confirm text returns.
Step 4 — COMPARE: the 40% claim on a real workload
Same fixture task, harness identical (own agent loop or Claude Code), run:
Config
Model
Effort
Pricing inputs
Baseline
claude-opus-5
high (its default)
$5/$25, cache read $0.50
Default
claude-opus-5-5
medium (its default)
$4/$20, cache read $0.20
Matched
claude-opus-5-5
high
as above
Max
claude-opus-5-5
max
as above
Record per run: pass/fail on hidden tests, input/output/cache-read/cache-write tokens (usage object), wall time, total cost = price-weighted sum. Compute cost-per-passed-task, not cost-per-run.
Interpretation rules (write these into the finding):
Anthropic's −20% token price is arithmetic — verify it in the bill line items.
The "40% cheaper on typical workloads" is the claim under test: it only holds if 5.5 clears the quality bar with ≤~⅔ the tokens at equal effort — Artificial Analysis's independent measurement (260M vs 88M median verbosity on their index, #94/211 cost rank) says verbose frontier runs are plausible, so the default-effort (medium) row is the economically interesting one.
Note the docs' own warning: 5.5 thinks more per turn than Opus 5 at equal effort — expect medium≈(Opus 5 high? measure, don't assume).
If the task touches security-adjacent code, watch for silent fallback to Opus 4.8 (Step 3 wiring) — any quality delta you see may be the guardrail tax, exactly as Anthropic's benchmark footnote admits for its own scores.
Step 5 — Cache-economics spot check
Long multi-turn agentic session (≥10 tool rounds) with a ≥512-token cached prefix. Verify cache-read share of total spend on 5.5 vs 5: the 0.20-vs-0.50 read price should make the cached-session cost ratio visibly better than the headline 20%. Record the ratio — this is the number that actually changes agent-unit economics.
Deliverable
One page: matrix of measured cost/quality per config + checklist of the four breaking changes confirmed against docs + verdict sentence: "40% cheaper claim: holds / holds only at medium effort / fails on our workload, because ___." Feeds directly into the migration workshop (business-value section) with our own numbers instead of vendor slides.
Duration: ~2 hours including runs; API spend estimate: well under $10 for steps 1–3 + step 5; step 4 depends on fixture size (cap max_tokens, log usage).
≡
Research sources
Primary Sources (7)
Primary
Dario Amodei — "We Must Pace the Frontier" URL: https://darioamodei.com/post/we-must-pace-the-frontier Type: Official CEO essay (company founder's publication) Date: September 2026 (page shows month only; announcement says "last week", TechCrunch says "earlier this month" ⇒ mid-September 2026; pre-window context, not part of the event) Used for: The pacing framework Opus 5.5 is framed against: three-step plan (embedded evaluators — unilateral Anthropic commitment; democratic coordination; global coordination); motivation citing the OpenAI–Hugging Face incident and recursive self-improvement acceleration; "pacing does not mean halting"; distillation-crackdown and chip-policy positions that map onto Opus 5.5's preserved-thinking safeguard; link to pacingthefrontier.com and METR's 2026-08-26 OAI-HF investigation. Evidence role: Primary context for strategic implications.
URL unavailable
Primary
Anthropic — "Introducing the Life Sciences Verification Program" URL: https://www.anthropic.com/news/life-sciences-verification-program Type: Official announcement (program page) Date: 2026-09-17 Used for: Context for Opus 5.5's biology-gating: Standard Use grants apply to Mythos 5.1/Opus 5/Sonnet 5 "and to future models as they launch" (⇒ Opus 5.5 coverage claim); High-risk Use grant model; three threat models (access compromise, insider, agent misuse); offline-monitoring shift requiring 30-day data retention for LSVP traffic; unavailability on third-party clouds and BAA/HIPAA orgs; launch-partner quotes (Xaira, Edison Scientific, Manifold Bio). Evidence role: Primary; governance/enterprise caveats.
URL unavailable
Primary
Anthropic platform docs — "Claude Opus 5.5" model overview URL: https://platform.claude.com/docs/en/models/opus-5-5/overview Type: Official developer documentation (model spec page) Date: states "Released September 22, 2026" Used for: Release date (independent of announcement marketing page); 1M context / 128K max output / 300K batch beta output; knowledge and training cutoffs (June 2026); full pricing grid; cache read = 5% of base input price; lineup comparison (Fable 5.1 $10/$50 default `high`; Sonnet 5 $2/$10; Haiku 4.5 $1/$5); retirement commitment "not sooner than September 22, 2027"; availability on five platforms; system-card reference URL. Evidence role: Primary; lifecycle and spec facts.
URL unavailable
Primary
Anthropic platform docs — "What's new in Claude Opus 5.5" URL: https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5 Type: Official developer documentation Date: live at retrieval 2026-09-26 Used for: The four breaking changes with exact error semantics (thinking cannot be disabled; forced tool_choice rejected; thinking-block model/conversation binding with 2026-08-31 account cutoff and prefix-mismatch 400s; computer_20251124 rejected on Claude API/Google Cloud but still working on Bedrock); text-between-tool-calls now arriving in thinking blocks (empty at default display); effort default `medium` vs Opus 5's `high`; more thinking per turn at equal effort; biology classifier + `reasoning_extraction` refusal category; refusal mechanics (`stop_reason:"refusal"`, `stop_details`, billing/rate-limit treatment, server-side `fallbacks:"default"` beta); pricing including 5m/1h cache write rates and Batch 50% ($2/$10); Fast mode research preview (Claude API only, `speed:"fast"`, `fast-mode-2026-02-01`); 512-token cacheable minimum; task budgets, compaction-on-demand beta, inline-tools beta; platform-by-platform model IDs; migration code matrix. Evidence role: Primary; basis for the CONFIRMED developer-impact facts and the VERIFY lab design.
URL unavailable
Primary
Anthropic — Claude Opus 5.5 System Card (PDF) URL: https://www.anthropic.com/claude-opus-5-5-system-card Type: Official safety documentation (resolves to PDF: https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf) Date: published with release, 2026-09-22 (PDF header not extracted) Used for: Existence and canonical location verified via HTTP redirect chain (301/307 to CDN PDF). PDF body text was NOT parsed in this environment, so no numeric claim in research/S01.md is sourced from the PDF itself; all system-card-derived statements come from the announcement's own summary of it. Evidence role: Primary (existence confirmed; content not independently read).
URL unavailable
Primary
Anthropic Newsroom index URL: https://www.anthropic.com/news Type: Official newsroom listing Date: retrieved 2026-09-26 (listing "Introducing Claude Opus 5.5 — Announcements — Sep 22, 2026") Used for: Independent-of-page-body confirmation of the Sep 22 announcement date; timeline anchors used as context: Fable 5.1 + Mythos 5.1 (Sep 1, 2026), threat-intelligence report (Sep 10, 2026), Life Sciences Verification Program (Sep 17, 2026), Accenture embedded evaluation (Sep 18, 2026). Evidence role: Primary index; date verification.
URL unavailable
Primary
Anthropic — "Claude Opus 5.5" (release announcement) URL: https://www.anthropic.com/claude-opus-5-5 Type: Official company announcement (product page) Date: 2026-09-22 Used for: Release date; "first model of Claude 5.5 family"; Fable 5.1-level-on-most-work and 40%-lower-run-cost claims; pricing table ($4/$20 vs $5/$25; cache reads $0.20 vs $0.50; Fast mode $8/$40 at up to 2.5×); benchmark table (Terminal-Bench 4.0 66.4%, CursorBench 4.0 57.8%, GDPval-AA v2.1 1846 Elo, HLE 67.7%, OSWorld 81.8%, TB-Science 58.7%, AutomationBench 40.0% vs Astra 41.4%); safeguard-class claims (first Opus with Fable-class cyber/biology/distillation gates; fallback routing to Opus 4.8 / Opus 5); behavioral-audit and containment-attempt claims; preserved-thinking applicability to accounts created on/after 2026-08-31; usage-limit raises and savable rate-limit reset; Sonnet/Haiku 5.5 promise; early-tester anecdotes (680K-line migration, HAProxy rewrite, 39/40 load-time tasks); customer quotes (GitHub, Deloitte, Ramp); ZDR availability; EU AI Act watermarking; platform availability. Evidence role: Primary; everything here is COMPANY CLAIM unless separately corroborated.
URL unavailable
Independent Sources (2)
Independent
Artificial Analysis — "Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)" model page URL: https://artificialanalysis.ai/models/claude-opus-5-5 Type: Independent benchmarking organization (own evaluation runs + API catalog data) Date: released Sep 22, 2026 per page; page "Updated" state retrieved 2026-09-26 Used for: INDEPENDENTLY VERIFIED evidence: Opus 5.5 runs in AA's own Intelligence Index v4.3.2 score 58, ranked #1 of 211 models; AA lists $4.00/$20.00 API pricing with 95% cache discount matching Anthropic's grid; 1M context; 6 API providers (corroborates multi-cloud availability); 260M output tokens on the index vs 88M median ("very verbose") and $5.98 cost/index-task with #94/211 cost rank — the independent counterweight to the 40%-cheaper headline; AA's evaluation suite list confirms GDPval-AA v2.1, AutomationBench-AA and Terminal-Bench 4.0 are third-party-run instruments. Evidence role: Key independent verification + nuance; per-benchmark numeric cells are chart-rendered and were not retrievable as text.
URL unavailable
Independent
TechCrunch — "Anthropic releases Opus 5.5 with lower prices and Fable-level performance" (Russell Brandom) URL: https://techcrunch.com/2026-09-22/anthropic-releases-opus-5-5-with-lower-prices-and-fable-level-performance/ Type: Reputable technology news reporting Date: 2026-09-22, 9:30 AM PDT Used for: Independent confirmation that the release occurred Tuesday Sep 22 with the stated price cuts ($20 vs $25 output, "other metrics similar"), faster serving, communication-style changes; the ~2-month cadence since Opus 5 (2026-07-24 link to https://www.anthropic.com/news/claude-opus-5); Mythos-comparable bio/cyber capability framing; positioning as first release after the pacing post (quoting Amodei); TechCrunch's own note that Anthropic claims it "outpaces the larger Fable model in many benchmarks." Evidence role: Independent corroboration of event facts; also shows press repeating company benchmark claims without reproducing them — supports our COMPANY-CLAIM labeling discipline.
URL unavailable
Secondary Sources (1)
Secondary
TechCrunch — "Astra and Opus just passed Turing's other test" (Tim Fernholz) URL: https://techcrunch.com/2026-09-25/astra-and-opus-just-passed-turings-other-test/ Type: Technology news reporting (adjacent coverage) Date: 2026-09-25, 10:24 AM PDT Used for: IMPORTANT SCOPE NOTE — although listed as an independent source in the S01 discovery record, this article's Anthropic result was obtained with **Claude Opus 5** (cryptanalyst Jack Willis, Sep 21, validated at https://www.cryptocellar.org/bgac/the-fmngi-break.html), NOT Opus 5.5. It was used only as adjacent context: (a) confirms Opus-family long-horizon agentic research capability in the real world independent of vendor benchmarks; (b) confirms the Opus 5 → Opus 5.5 timeline; (c) prevents a conflation error — research/S01.md must not attribute the Enigma result to 5.5. Evidence role: Secondary/context; used for a negative disambiguation, not for corroboration of any Opus 5.5 claim.