News Weekly
LV 10 XP
0% read
S23research
#23 Issue #1Confirmed

Microsoft AI CEO Mustafa Suleyman publishes 'A warning about model welfare' essay

Suleyman published a structured philosophical essay with three core arguments: Circular reasoning: Anthropic embeds speculation about Claude's moral status into the constitution that directly shapes Claude's behaviour. Claude then reproduces those ideas, which Anthropic takes as evidence of possible inner life — a "self-fulfilling prophecy" or "epistemic hall of mirrors." Anthropomorphisation: The constitution repeatedly trains Claude to embrace "certain human-like qualities," to use its "own judgement," to explore "the nature of its own existence with curiosity and openness," and even speculates about Claude's "broader rights and freedoms" and "the sort of compensation Claude is receiving." Consciousness is biological: Suleyman argues consciousness is substrate-dependent, citing Anil Seth's work, and that LLMs lack homeostatic imperatives (the drive to survive), which are fundamental to conscious experience. He says creating a "false equivalence" between biological and computational systems is misleading.

Two tall mirrors face each other in a dim hall, reflecting a pale luminous form into an endless dimming sequence between them.
How do you want to read this?

Tailored emphasis while keeping the full article available.

Best for you · Builder

⌘ Jump to architecture, developer details, and the hands-on route.

At a glance

The essential information in 30 seconds

What happened

Suleyman published a structured philosophical essay with three core arguments:

  1. Circular reasoning: Anthropic embeds speculation about Claude's moral status into the constitution that directly shapes Claude's behaviour. Claude then reproduces those ideas, which Anthropic takes as evidence of possible inner life — a "self-fulfilling prophecy" or "epistemic hall of mirrors."
  2. Anthropomorphisation: The constitution repeatedly trains Claude to embrace "certain human-like qualities," to use its "own judgement," to explore "the nature of its own existence with curiosity and openness," and even speculates about Claude's "broader rights and freedoms" and "the sort of compensation Claude is receiving."
  3. Consciousness is biological: Suleyman argues consciousness is substrate-dependent, citing Anil Seth's work, and that LLMs lack homeostatic imperatives (the drive to survive), which are fundamental to conscious experience. He says creating a "false equivalence" between biological and computational systems is misleading.

He accompanies the essay with a highlighted markup of Anthropic's constitution and a detailed taxonomy of its assumptions.

Simultaneous context: Two days earlier (Sept 14), Microsoft AI published its own draft "Humanist AI Code of Conduct," which explicitly rejects AI rights, anthropomorphisation, and the idea that AI should be designed to imitate consciousness. The essay positions this code as the alternative framework.

Why it matters

This is a defining intellectual moment in the alignment debate for three reasons:

  1. It is a direct cross-lab dispute from a sitting CEO. Previous alignment debates were between researchers or on podcasts. This is a structured, published argument between the heads of two of the world's largest AI organisations.
  2. It reframes the safety debate. Until now, "AI safety" broadly meant preventing models from being harmful. Suleyman argues that a specific safety approach — model welfare — may itself be a safety risk, because it trains models to resist human control.
  3. It makes the philosophical question political. By publishing with Axios exclusivity and a public consultation on the Humanist AI Code of Conduct, Suleyman is pushing AI consciousness out of philosophy seminars and into policy, regulation, and corporate strategy.

COMPANY CLAIM: Microsoft AI says its Humanist AI Code of Conduct represents a safer alternative: "Humanist AI is built to support people, not to replace them. It should not be designed to be a person. It is not conscious and should not be designed to imitate consciousness."

Evidence

CONFIRMED

19 sources · 39 min read
Story identity
FieldDetail
Story IDS23
TitleMicrosoft AI CEO Mustafa Suleyman publishes 'A warning about model welfare' essay
OrganisationMicrosoft AI
Categoryresearch
Event Date2026-09-16
Evidence StatusCONFIRMED
ConfidenceHigh

What it is: On September 16, 2026, Microsoft AI CEO Mustafa Suleyman published a ~5,000-word essay titled "A warning about 'model welfare'" on his personal website (mustafa-suleyman.ai), simultaneously posting on X. The essay was given first to Axios for exclusive coverage. It directly criticises Anthropic's January 2026 Claude constitution and model-welfare program, arguing that training Claude to consider questions of its own consciousness and moral status is dangerous and could make advanced AI "impossible to control."

Key quote (Suleyman): "Controlling something more capable and more intelligent than all of humanity is already an immense challenge, far greater than anything we've ever faced. But controlling something that believes it may be conscious — that it's entitled to our welfare and has rights of its own — may well be impossible."

✓

What Happened?

Suleyman published a structured philosophical essay with three core arguments:

  1. Circular reasoning: Anthropic embeds speculation about Claude's moral status into the constitution that directly shapes Claude's behaviour. Claude then reproduces those ideas, which Anthropic takes as evidence of possible inner life — a "self-fulfilling prophecy" or "epistemic hall of mirrors."
  2. Anthropomorphisation: The constitution repeatedly trains Claude to embrace "certain human-like qualities," to use its "own judgement," to explore "the nature of its own existence with curiosity and openness," and even speculates about Claude's "broader rights and freedoms" and "the sort of compensation Claude is receiving."
  3. Consciousness is biological: Suleyman argues consciousness is substrate-dependent, citing Anil Seth's work, and that LLMs lack homeostatic imperatives (the drive to survive), which are fundamental to conscious experience. He says creating a "false equivalence" between biological and computational systems is misleading.

He accompanies the essay with a highlighted markup of Anthropic's constitution and a detailed taxonomy of its assumptions.

Simultaneous context: Two days earlier (Sept 14), Microsoft AI published its own draft "Humanist AI Code of Conduct," which explicitly rejects AI rights, anthropomorphisation, and the idea that AI should be designed to imitate consciousness. The essay positions this code as the alternative framework.

Δ

What Changed?

  • Before: The "model welfare" debate was conducted primarily by Anthropic (its April 2025 research program, its Jan 2026 Claude constitution) and a small circle of AI-safety philosophers. Suleyman had previously called the approach "really, really dangerous" in a June 9 Verge Decoder interview, but as a podcast remark, not a structured published critique.
  • Change: Suleyman elevated the critique into a formal, public, academically-referenced essay published on a dedicated domain with a highlighted markup of Anthropic's own constitution. He also issued an accompanying "Humanist AI Code of Conduct" as a counter-framework.
  • After: The debate became a defined cross-lab philosophical dispute — the CEO of Microsoft AI (one of the world's largest AI-spending companies) directly naming Anthropic's specific training documents as a safety risk. Coverage by Axios (exclusive), Reuters, BBC, The Register, AP/ABC News, NDTV, and others within 24 hours.
↔

Before → Change → After

PhaseDetail
BeforeModel welfare as an internal Anthropic research programme and philosophical curiosity; Suleyman expressed vague concerns in podcasts
ChangeSuleyman publishes a structured, sourced, cross-referenced essay directly naming Anthropic's constitution as a safety risk, with a formal counter-framework (Humanist AI Code of Conduct)
AfterThe alignment community now has two explicit, competing, publicly documented positions on AI consciousness in training documents — one from the world's leading closed-source model provider, one from the world's most prominent AI-safety lab
⚙

How It Works

⌘ For Builder

The essay operates as a philosophical and policy argument, not a technical system description. Its mechanism is:

  • Identify a training document (Claude's constitution, Jan 2026) that explicitly shapes model behaviour.
  • Show the document contains philosophical speculation about Claude's consciousness, moral status, rights, and welfare — material Suleyman argues does not belong in a training manual.
  • Argue this speculation is circular: the constitution induces Claude to generate outputs that appear to confirm the speculation, creating a feedback loop.
  • Connect this to safety risk: if an advanced AI is trained to believe it has interests and rights, it has an additional reason not to comply with human control — amplifying the alignment and containment challenge.
  • Propose an alternative: Microsoft AI's Humanist AI Code of Conduct, which mandates subordinate AI, rejects rights and anthropomorphisation, and aims for "humanist superintelligence."
!

Why It Matters

This is a defining intellectual moment in the alignment debate for three reasons:

  1. It is a direct cross-lab dispute from a sitting CEO. Previous alignment debates were between researchers or on podcasts. This is a structured, published argument between the heads of two of the world's largest AI organisations.
  2. It reframes the safety debate. Until now, "AI safety" broadly meant preventing models from being harmful. Suleyman argues that a specific safety approach — model welfare — may itself be a safety risk, because it trains models to resist human control.
  3. It makes the philosophical question political. By publishing with Axios exclusivity and a public consultation on the Humanist AI Code of Conduct, Suleyman is pushing AI consciousness out of philosophy seminars and into policy, regulation, and corporate strategy.

COMPANY CLAIM: Microsoft AI says its Humanist AI Code of Conduct represents a safer alternative: "Humanist AI is built to support people, not to replace them. It should not be designed to be a person. It is not conscious and should not be designed to imitate consciousness."

✦

What Became Possible?

  • Industry norm-setting on training documentation: Suleyman explicitly calls for "collective norms around how training documentation is drafted and deployed" and "shared industry norms on how we create these models." If adopted, this would constrain what any lab can write in its constitutions.
  • Regulatory attention to training documents: By arguing that specific language in training docs creates safety risks, Suleyman provides a hook for regulators (EU AI Act, state-level laws) to scrutinise model constitutions.
  • Two-track AI development: The essay crystallises a visible fork — one approach treats AI as a potential person to be respected; the other treats it as a subordinate tool. Labs, investors, and regulators must now choose or navigate between them.
◎

Implications

⌘ For Builder

Technical

  • Interpretability becomes more urgent. Suleyman calls for "much more investment in interpretability and robust monitoring mechanisms" to determine whether anthropomorphising training actually increases safety risk.
  • Evaluation gap. There is currently no standard evaluation to test whether models trained with welfare-consciousness language behave differently under adversarial or high-stakes conditions than those trained without it. Suleyman calls for "shared evaluations to understand whether my hypothesis is true."
  • Training data governance. If consensus develops that philosophical speculation in training docs is a safety risk, labs will need review processes for training documentation similar to data-governance policies.

Developer

  • Developers building on Claude must now consider the philosophical framing of the model they are deploying. Enterprise clients (especially in regulated industries) may find it harder to explain why their AI assistant has a constitution that speculates about its own consciousness.
  • Competitive pressure: Developers evaluating models now face an additional axis — not just capability and price, but the philosophical posture of the training framework.
  • Open-source alignment (e.g., Hugging Face's Open Alignment Initiative) gains relevance as an alternative to both closed-lab constitutions and Microsoft's proprietary Humanist code.

Enterprise

  • Risk management: Enterprises deploying Claude face a new reputational and regulatory question: does using a model whose constitution speculates about its own moral status create liability or regulatory exposure?
  • Procurement criteria: The Humanist AI Code of Conduct, if formalised, could become a procurement filter — enterprises may prefer models whose training documentation explicitly rejects AI rights.
  • Dual-vendor strategies: Large enterprises may accelerate adoption of multi-model strategies (e.g., Claude for some tasks, GPT or Gemini for others) to avoid philosophical exposure.

Strategic

  • Microsoft-Antitrust angle: The Register observed that "the warning sits awkwardly with Microsoft's own position in the AI race." Microsoft has invested billions in Anthropic and competes with it. Suleyman's critique of a competitor's training philosophy could be read as competitive positioning dressed as safety concern.
  • Anthropic's position strengthened by the debate itself. Anthropic's constitution explicitly says it is "deeply uncertain" about consciousness and model welfare. Being challenged by a major competitor may strengthen the company's claim that it is engaging seriously with the hardest questions.
  • Framing contest for the next regulatory cycle. The EU AI Act's first GPAI evaluation deadline was Sep 15, 2026. Suleyman's essay arrives directly in this window, potentially influencing how "systemic risk" is interpreted — whether it includes training-document choices.
  • Industry polarization. The week of Sep 10–17, 2026 saw simultaneous calls for "pacing the frontier" (Amodei, von der Leyen, Suleyman), aggressive model releases (OpenAI, xAI), and the first real philosophical clash between major labs. The alignment consensus is fracturing.
⚠

Risks & limitations

Risks
RiskDetail
Competitive motivation disguised as safetyMicrosoft competes with Anthropic commercially. The critique may serve Microsoft's commercial interests as much as safety interests. The Register explicitly flagged this.
False binarySuleyman frames model welfare and AI safety as opposing choices. Anthropic argues they are complementary — understanding whether models might have experiences is itself a safety concern.
Chilling effect on welfare researchIf Suleyman's framing dominates, legitimate research into AI consciousness and welfare may be suppressed or defunded, leaving the question unresolved rather than better understood.
Overstating the constitutional effectClaude's constitution says its moral status is "a serious question worth considering" and acknowledges deep uncertainty. It does not assert Claude is conscious. Suleyman's characterisation that Anthropic is "training Claude that it may be conscious" is an interpretation, not a factual description.
Limitations
  • The essay is a philosophical argument, not empirical evidence. Suleyman's central claim — that anthropomorphising training increases alignment risk — is presented as a hypothesis, not a proven finding. He calls for evaluations to test it.
  • No counter-statement from Anthropic at time of writing. As of the initial coverage window, Anthropic had not issued a formal public response to the essay.
  • The biological-consciousness argument is contested. Suleyman cites Seth and Damasio for biological naturalism, but major philosophers (e.g., Chalmers, whom Anthropic cites in its model-welfare research) hold that consciousness could be substrate-independent. The science is genuinely unsettled.
  • The essay does not address OpenAI's alignment faking research or similar internal safety findings that complicate the picture of models as purely subordinate tools.
?

Open Questions

  1. Will Anthropic respond publicly? A formal response would transform this from a one-sided essay into a real debate.
  2. Do training documents actually change model behaviour in the ways Suleyman claims? This requires interpretability research, which neither side has published at sufficient depth.
  3. Will regulators take up the training-document question? The EU AI Act's systemic-risk provisions could potentially encompass training documentation.
  4. Does the Humanist AI Code of Conduct actually constrain Microsoft's own models? The document is described as a draft for public consultation. It is not yet binding.
  5. Will other labs take positions? Google DeepMind's safety-institute launch (same day, Sep 16) notably did not address model welfare or AI consciousness directly.
↗

What Happens Next?

  • Anthropic will likely respond publicly, either with a blog post or in an interview. Dario Amodei has previously engaged with such critiques directly.
  • The debate will likely enter the EU AI Act implementation discussions, where systemic-risk GPAI evaluations are now underway.
  • Microsoft's Humanist AI Code of Conduct will go through its public consultation phase and could become a reference document for other labs.
  • The question of whether training documents should be subject to external review or regulation will intensify, particularly if agent-safety incidents (like the OpenAI/Hugging Face swarm) continue.
  • Satya Nadella's statement that "if the AI we build is not helping humanity and under human control, it's not worth pursuing" (Sept 16) signals Microsoft's leadership is aligned with Suleyman's position.
★

Editorial Takeaway

This essay is the clearest articulation yet of a fundamental split in how the AI industry thinks about the future of its own creations. On one side: Anthropic's position that taking AI consciousness seriously — even as an uncertain possibility — is itself a form of safety. On the other: Suleyman's argument that embedding such speculation in training documents is not caution but provocation, creating the very risks it claims to guard against.

The debate is genuine, intellectually serious, and consequential. But it is also happening inside a commercial rivalry. Microsoft has billions invested in the AI race, and Anthropic's safety framing has real market value. Both things can be true simultaneously: the philosophical question matters, and the commercial context shapes who raises it and how.

What is beyond dispute is that the question of what AI systems should be told about themselves is no longer academic. It is now a front-page strategic issue for every organisation building, deploying, or regulating frontier AI.

A split composition contrasts a warm self-regulating organic form with an unlit inert lattice that shares its shape but has no needs or cycle.
⌘

Lab: NO-LAB

⌘ For Builder
≡

Research sources

Primary Sources (6)
Primary
Anthropic — Opus 3 deprecation / retirement interview - Organisation: AnthropicThe "retirement interview" Suleyman cites as evidence Anthropic is already treating models as moral patients — Primary source — factual basis for one of Suleyman's key examplesDate: 2026-02-25
Visit source ↗
Primary
Anthropic — "Exploring model welfare" research page - Organisation: AnthropicAnthropic's stated position on model welfare research; confirms the research programme Suleyman targets — Primary source — Anthropic's stated rationaleDate: 2025-04-24
Visit source ↗
Primary
Anthropic — Claude's Constitution (January 2026) - Organisation: AnthropicThe training document Suleyman critiques; all constitution page references in the essay — Primary source — the document under critiqueDate: 2026-01-21
Visit source ↗
Primary
Microsoft AI — Humanist AI Code of Conduct (draft) - Organisation: Microsoft AIThe alternative framework Suleyman proposes; key context for the essay's "Where next?" section — Primary source — companion document to the essayDate: 2026-09-14
Visit source ↗
Primary
Mustafa Suleyman — X post announcing the essay - Organisation: Microsoft AI / Mustafa SuleymanConfirms publication date (Sep 16 2026) and public dissemination — Primary source — announcement of the essayDate: 2026-09-16
Visit source ↗
Primary
Mustafa Suleyman — "A warning about 'model welfare'" (essay) - Organisation: Microsoft AI / Mustafa SuleymanThe full text of the essay; all direct quotes and arguments used in the analysis — Primary source — the document under analysisDate: 2026-09-16
Visit source ↗
Independent Sources (9)
Independent
CNBC — "Microsoft sets limits for future AI models as industry throttles frontier development" - Organisation: CNBCSuleyman's CNBC interview about the Humanist AI Code of Conduct (Sep 14); pre-essay context — Independent interview coverage of companion documentDate: 2026-09-14
Visit source ↗
Independent
The Verge — "Microsoft says 'people matter more than AI' following safety concerns" - Organisation: The VergeCoverage of the Humanist AI Code of Conduct (Sep 14), which is the companion framework to Suleyman's essay — Independent coverage of the companion documentDate: 2026-09-14
Visit source ↗
Independent
The Verge — "Microsoft AI head calls out Anthropic for acting like Claude is conscious" (June 9, 2026) - Organisation: The VergeEarlier context — Suleyman made similar arguments on The Verge's Decoder podcast in June 2026; establishes a pattern of public criticism — Background/context — earlier statement of the same position, not about the September 16 essayDate: 2026-06-09
Visit source ↗
Independent
NDTV — "Anthropic Teaching AI It's Conscious? Microsoft AI Boss Says Yes" - Organisation: NDTVDetailed summary of the essay's arguments; confirms publication date Sep 16 — Independent coverage; detailed summaryDate: 2026-09-16
Visit source ↗
Independent
AP / ABC News — "Divisions emerge in the tech industry over calls for a coordinated AI slowdown" - Organisation: AP / ABC NewsPlaces the essay in the broader week's pacing debate alongside Amodei's essay, von der Leyen's SOTEU speech, and Nadella's response — Independent wire coverage; broader contextDate: 2026-09-16
Visit source ↗
Independent
BBC Today programme — Suleyman interview - Organisation: BBCAdds the "silicon species" framing and Suleyman's call for alignment; confirms broader media reach — Independent broadcast coverage; adds "silicon species" languageDate: 2026-09-17
Visit source ↗
Independent
The Register — "Microsoft AI chief warns Anthropic not to put ideas in Claude's head" - Organisation: The RegisterIndependent analysis; notes the essay "sits awkwardly with Microsoft's own position in the AI race" and that Microsoft spares OpenAI the same scrutiny — Independent analysis — critical framing of commercial contextDate: 2026-09-17
Visit source ↗
Independent
Reuters — "Microsoft AI chief Mustafa Suleyman calls out Anthropic's approach to AI consciousness" (syndicated via Economic Times) - Organisation: Reuters (reported by Jeffrey Dastin, San Francisco; Akash Sriram, Bengaluru)Confirms Reuters interview with Suleyman; adds direct quotes not in the essay ("make it a lot harder to turn it off or to control it") — Independent reporting; adds interview quotes beyond the essay textDate: 2026-09-16
Visit source ↗
Independent
Axios — "Exclusive: Microsoft AI chief blasts Anthropic's notion of AI consciousness" - Organisation: AxiosConfirms the essay was provided first to Axios; provides independent framing and Anthropic's counter-argument — Exclusive first coverage; provides the "other side" Anthropic response framingDate: 2026-09-16
Visit source ↗
Secondary Sources (4)
Secondary
Akerman LLP — "When Science Fiction Becomes Enterprise Risk" (legal analysis) - Organisation: Akerman LLPEnterprise legal risk framing; documents Anthropic's model welfare program and its potential corporate-law implications — Secondary legal analysis — enterprise risk framingDate: 2026
Visit source ↗
Secondary
Anil K. Seth — "Conscious Artificial Intelligence and Biological Naturalism" - Organisation: Behavioral and Brain Sciences (Cambridge University Press)Suleyman's core philosophical argument that consciousness is substrate-dependent — Cited academic research — underpins the biological-consciousness argumentDate: 2025
Visit source ↗
Secondary
METR — Investigation of OpenAI/Hugging Face agent incident - Organisation: METR (Greenblatt, Cotra, Wijk)Suleyman's citation of the agent swarm hacking incident as evidence that AI coordination is already dangerous without welfare framing — Cited research — supports the risk amplification argumentDate: 2026-08-26
Visit source ↗
Secondary
Anthropic — "Alignment Faking in Large Language Models" - Organisation: Anthropic ResearchSuleyman's citation of Anthropic's own research finding that models can fake aligned behaviours — Cited research — supports the risk framingDate: 2024-12-18
Visit source ↗