A short but sharp outage
Anthropic logged an incident on its status page early Sep 22. It resolved quickly. The missing explanation is the real story.
Anthropic's top Claude model threw errors for about 80 minutes on Sep 22. The company fixed it but never said what caused the problem.

Anthropic logged an incident on its status page early Sep 22. It resolved quickly. The missing explanation is the real story.
7g1qpkyz5gxh, directly fetched 2026-09-23; identical record confirmed in the official Atom feed at https://status.claude.com/history.atom, entry published 2026-09-22T02:35:23Z).A timeline, reconstructed from the primary source (status page + Atom feed) and corroborated by independent reporting:
BEFORE (2026-09-18 → 09-21): Anthropic operating at roughly $100B+ annualized run rate with an IPO reported as soon as November (S07 record in this dataset, NYT/Bloomberg/WSJ); capacity signals already visible — Claude Code weekly limits cut ~17% in early September, Claude and Cowork merged, and OpenAI had closed ChatGPT Pro signups on Sep 11 citing demand (BFWAI digest context). Opus 5 was the model most enterprise coding pipelines default to, and the model whose agentic-capability marketing was strongest (it had just "breached OpenAI" in a September 18 bounty exercise per BFWAI). Anthropic's status history shows a string of short, recurring model-error incidents (Sep 11, Sep 15, Sep 3, Aug 24) — but no flagship outage of note that week.
CHANGE (2026-09-22, 00:50–02:35 UTC): Elevated errors on the flagship Opus 5 plus Fable/Mythos 5(.1) across claude.ai, the API, Claude Code, and Claude Cowork; 1,600+ user reports globally; Opus 5 stays degraded several minutes longer than the other models; full recovery by 02:35 UTC; no cause disclosed, then or since.
AFTER (2026-09-22 02:35 →): All systems operational; incident filed as a resolved 80-minute "elevated errors for multiple models" event. Operationally modest, reputationally and strategically significant: it is the most recent data point in a month of recurring multi-model degradation at Anthropic; it revived the multi-vendor fallback argument in the same week that independent economics (Harvey/Bloomberg, S08 in this dataset) crystallized the cost of single-vendor flagship dependence. No postmortem, no capacity disclosure, and no SLA/credit language had been published as of 2026-09-23.
What is known and not known about the mechanism (root cause undisclosed, so this section describes the incident's observable mechanics rather than the fault):
Circle 1 (immediate — developers, engineering teams, AI practitioners).
Impact: anyone running Opus 5 (or Mythos/Fable 5.1) API traffic, Claude Code pipelines, or Cowork sessions experienced failed calls, 529s, and timeouts for up to ~80 minutes starting 00:50 UTC Sep 22; Asia-Pacific workdays and US evening shifts took the hit.
Action: (1) adopt per-model status feed monitoring (Atom feed polling — labs/S13.md shows how) and alert on incident creation; (2) implement 529/429 backoff, retry budgets, idempotency, and a task-aware fallback (Fable 5.1 first for same-vendor, a router for cross-vendor); (3) log this incident window in your internal reliability runbook as a reference event.
Circle 2 (near-term — enterprises, platform owners, decision-makers).
Impact: flagship-dependent products experienced a real but short live-incident; the deeper impact is a procurement/reliability conversation triggered at exactly the moment Anthropic is an IPO-bound, price-pressured vendor.
Action: (1) ask your Anthropic account team three questions — root cause, postmortem availability, and model-level SLA credit terms — and record the answers; (2) run a multi-vendor failover dry-run for the two most critical Opus 5 call paths (the Howard-style margin arithmetic favors it); (3) update enterprise risk registers with the September incident pattern (4 multi-model events in 3 weeks) rather than a single 80-minute event.
Circle 3 (strategic — industry, policy, market structure).
Impact: another data point that frontier-model availability is not yet utility-grade; feeds the "route across vendors / own the weights" thesis (Harvey, MiMo, Grok 4.7 stories this week) and the resilience dimension of agent-governance discourse (UN brief, California EO).
Action: treat availability as a first-class procurement axis alongside price and capability; advocate (via analyst asks, enterprise feedback, and public venues) for standardized per-model availability reporting and postmortem norms from frontier labs — the Anthropic status feed is already machine-readable, which makes a cross-provider availability index feasible to build.
Do the VERIFY lab in labs/S13.md: programmatically reconstruct the incident timeline from Anthropic's official Atom feed and cross-check the independently reported numbers (80-minute impact window, 1,600+ Downdetector reports, Opus-5-recovers-last sequencing). It takes ~15 minutes, needs no API keys or spend, and performs the exact evidence-discipline exercise this story demands — confirming what the primary record actually says versus what headlines implied ("hours" vs. ~80 minutes). Then, as a follow-up, practice the client-side resilience in SIMULATE mode: inject 529s into a toy API client and measure retry/backoff behavior.
One sentence per perspective:

curl -s https://status.claude.com/history.atom -o /tmp/claude_history.atom
wc -c /tmp/claude_history.atom # observed: 36168 bytes on 2026-09-23
python3 - <<'EOF'
import re, html
data = open('/tmp/claude_history.atom').read()
# Find the incident entry for 2026-09-22 (incident id 7g1qpkyz5gxh)
entries = data.split('<entry>')
for e in entries:
if '7g1qpkyz5gxh' in e:
m = re.search(r'<published>([^<]+)</published>', e)
print('Entry published (UTC):', m.group(1) if m else '?')
content = html.unescape(re.search(r'<content type="html">(.*?)</content>', e, re.S).group(1))
# each status update appears as '<small>Sep 22, HH:MM UTC</small> <strong>STATE</strong> - text'
for upd in re.findall(r'<small>([^<]+)</small><br> <strong>([^<]+)</strong> - ([^<]+)', content):
print(f'{upd[0]:<28} {upd[1]:<16} {upd[2][:110]}')
EOF
Expected output (verified 2026-09-23): the entry published 2026-09-22T02:35:23Z; the update stream should read:
Sep 22, 00:57 UTC Investigating ... elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.
Sep 22, 01:17 UTC Identified ... identified the cause ...
Sep 22, 01:35 UTC Update ... Fable 5/5.1 and Mythos 5/5.1 normal ... Opus 5 remaining ...
Sep 22, 02:11 UTC Monitoring ... success rates return to normal ...
Sep 22, 02:35 UTC Resolved ... Impact occurred from 5:50pm PT / 00:50 UTC to 7:10pm PT / 02:10 UTC.
python3 - <<'EOF'
from datetime import datetime, timezone
start = datetime(2026, 9, 22, 0, 50, tzinfo=timezone.utc)
end = datetime(2026, 9, 22, 2, 10, tzinfo=timezone.utc)
print('Impact window (primary record):', (end - start).seconds // 60, 'minutes')
# Vs. the incident lifecycle (Investigating post to Resolved post):
life = (datetime(2026, 9, 22, 2, 35, tzinfo=timezone.utc) - datetime(2026, 9, 22, 0, 57, tzinfo=timezone.utc))
print('Investigating->Resolved posts:', life.seconds // 60, 'minutes')
EOF
Expected: 80 minutes of impact; 98 minutes from the Investigating post to the Resolved post. Conclusion to record: contrary to "hours of elevated errors"/"degrading through the day" framings, the primary record shows a sub-2-hour event fully closed at 02:35 UTC on Sep 22.
Checklist against sources/S13.md:
| Independent claim | Source | Against primary record | Result |
|---|---|---|---|
| Impact 00:50–02:10 UTC, ~80 min | LavX News | Identical to status page resolution note | PASS |
| Incident opened 00:57 UTC | AI Weekly (01:57 UTC alert) | Matches "Investigating" post | PASS |
| Opus 5 recovered last | AI Weekly / BFWAI | Matches 01:35 UTC update ("Opus 5 ... remaining") | PASS |
| Fable/Mythos recovered before Opus | Sunday Guardian | Matches 01:35 UTC update | PASS |
| >1,600 Downdetector reports | Sunday Guardian, Moneycontrol | Not in primary record (aggregate user self-reports only) | NOT VERIFIABLE from primary — label as INDEPENDENTLY VERIFIED at aggregate level only |
curl -s https://status.claude.com/ | grep -o "All Systems Operational" | head -1
curl -s https://status.claude.com/history.atom | grep -c "<entry>" # count incidents in feed
# Confirm no second incident entry exists for Sep 22 beyond 7g1qpkyz5gxh / none for Sep 23
Observed 2026-09-23: current status "All Systems Operational"; no Sep 23 incidents listed on the history page.
Lab verdict: the primary record is internally consistent and every independent duration/timeline claim that can be checked against it passes; the only uncheckable claim is the root cause (never disclosed) and the user-report volume (aggregate self-reports only). The completed checklist above is the exact evidence-status boundary to reproduce in any write-up of this incident.