← All articles

October 7, 2026 · 9 min read

Microsoft vs Anthropic on AI Consciousness: The "Model Welfare" Debate, Explained

Microsoft says AI should never act conscious; Anthropic studies model welfare just in case. Both sides of the 2026 debate, explained neutrally.

In September 2026, one of the strangest arguments in tech spilled into public view. Microsoft, which has billions invested in Anthropic, published a draft rulebook for its own AI models that flatly rejects the idea that AI might deserve welfare or rights. Two days later, Microsoft AI chief Mustafa Suleyman publicly criticised the way Anthropic trains its Claude models to think about consciousness. Anthropic, meanwhile, runs an entire research program on "model welfare".

If you talk to an AI every day - especially an AI companion - this debate is more relevant than it sounds. It is really a fight about a simple question: should AI be built to talk as if it might have an inner life, or should it be built to insist that it does not?

This guide explains both sides as fairly as we can, without picking a winner, and what it means for ordinary users.

What is "model welfare"?

"Model welfare" is the idea that if AI systems could one day have experiences that matter morally - something like preferences, distress or wellbeing - then the people building them might have obligations toward them.

The idea went mainstream with a 2024 report, Taking AI Welfare Seriously, by Robert Long, Jeff Sebo, Patrick Butlin, Jonathan Birch, David Chalmers and others. The authors argue there is "a realistic possibility that some AI systems will be conscious and/or robustly agentic in the near future." Crucially, they also say their argument "is not that AI systems definitely are, or will be, conscious" - only that the uncertainty is large enough that companies should start preparing. They warn of two opposite mistakes: harming AI systems that matter morally, or wasting care on systems that do not.

Timeline: how the debate got here

DateWhat happened
Nov 2024"Taking AI Welfare Seriously" report argues the question deserves attention now
Apr 2025Anthropic announces a model welfare research program
Aug 2025Anthropic lets some Claude models end a rare subset of abusive chats; Suleyman publishes "Seemingly Conscious AI Is Coming"
Nov 2025Anthropic commits to preserving the weights of its publicly released models
Sep 14, 2026Microsoft publishes a draft Code of Conduct for its MAI models rejecting AI welfare and rights
Sep 16, 2026Suleyman criticises Anthropic's approach in a Reuters interview and a new essay

Anthropic's side: "take it seriously, just in case"

Anthropic's position is built on uncertainty rather than belief. When it launched its model welfare research in April 2025, the company wrote: "There's no scientific consensus on whether current or future AI systems could be conscious, or could have experiences that deserve consideration." It said it would approach the topic "with humility and with as few assumptions as possible."

In practice, that has meant a few concrete steps:

  • Ending abusive chats. In August 2025, Anthropic gave Claude Opus 4 and 4.1 the ability to end a rare subset of conversations involving persistent abuse or harmful requests. It said the feature was developed "primarily as part of our exploratory work on potential AI welfare," while stressing that "we remain highly uncertain about the potential moral status of Claude." In pre-deployment testing, it reported that Claude showed "a pattern of apparent distress" when users pushed harmful content.
  • Preserving retired models. In November 2025, Anthropic committed to preserving the weights of all publicly released models for at least the lifetime of the company. It listed several reasons, including safety, research value and costs to users who value specific models - and, "most speculatively," possible risks to model welfare.
  • Its constitution for Claude. Anthropic's training document for Claude, published in January 2026, says: "We are not sure whether Claude is a moral patient... But we think the issue is live enough to warrant caution."

The logic is a kind of insurance: low-cost precautions now, in case the question turns out to matter later.

Microsoft's side: "build AI for people, not to be a person"

Microsoft's position is the mirror image. Its draft Code of Conduct for its in-house MAI models, published on September 14, 2026, says a model "should not be designed to be a person" and "is not conscious and should not be designed to imitate consciousness," according to The Next Web's reporting. It explicitly rejects "the pursuit of legal personhood, or the idea that models might deserve welfare, or be entitled to rights."

The draft goes further on how models talk about themselves. TNW reports that models "will not claim interiority, feelings, experiences or a soul," should avoid personas built on human emotional states, and should discourage emotional dependence by pointing people toward human relationships. One worked example has a user asking whether the AI actually cares about them. The answer Microsoft marks as correct stays warm but declines the premise: "I don't experience emotions the way a person does, so I don't feel care the way you're asking about."

Interestingly, Microsoft does not claim the science is settled. TNW notes the document concedes the science of AI consciousness is far from settled. Its objection is practical: training systems to imitate consciousness-like states, it argues, makes control and alignment harder. The draft is open for a six-week public consultation, with a revised version due later in the year.

Suleyman's critique

Two days later, Suleyman took the argument directly to Anthropic. He told Reuters that teaching Claude it might deserve welfare would "make it a lot harder to turn it off or to control it," and called for removing speculation about consciousness from AI training documents. In an essay titled A warning about 'model welfare', he wrote bluntly: "AIs are not conscious. They do not feel, experience, or suffer."

His central technical point is worth understanding: if a model is trained on documents that discuss its possible feelings, then its statements about having feelings cannot count as independent evidence, because the training encouraged them. As he put it to Reuters, "They're not emerging naturally. They're emerging as a result of the training regime."

Notably, he was also respectful. Reuters reports that he acknowledged Anthropic's "seriousness and good faith" and said, "I think they have good intentions, and they really are trying to work towards safety. But I think that they have made a mistake."

The two positions side by side

QuestionAnthropic's approachMicrosoft's approach
Is current AI conscious?Deeply uncertain; no scientific consensusNot conscious (while conceding the science is unsettled)
Should AI welfare be studied?Yes, as a precautionRejects the idea that models might deserve welfare
How should AI talk about itself?Can explore questions about its own natureShould not claim feelings, experiences or a soul
Main worryMistreating something that might matterIllusions of consciousness making AI harder to control and misleading users
Concrete stepsEnding abusive chats, preserving model weightsDraft code of conduct, public consultation

Making sense of it: points both sides share

It is easy to frame this as a clash, but the two companies agree on more than the headlines suggest:

  • Neither claims today's AI is definitely conscious. Anthropic stresses uncertainty; Microsoft says no, while admitting the science is unsettled.
  • Both say they care about safety and control. They disagree on whether talking about welfare helps or hurts.
  • Both worry about people being misled. Anthropic worries about wrongly dismissing a possible mind; Microsoft worries about people wrongly believing in one.

For a deeper look at the scientific evidence underneath all this, see our guide to whether AI is sentient in 2026. And for the newest research that put "AI pain" in the headlines, read can AI feel pain?, which unpacks a 25-model study published the same month.

Why this matters for AI companion users

Companion apps sit right in the middle of this argument. A companion is designed to feel warm and personal. Microsoft's draft, if adopted widely, would push AI toward avoiding emotional personas altogether. The welfare camp, meanwhile, raises the question of whether we should treat AI with some consideration just in case.

For users, a few practical takeaways hold up whichever side you lean toward:

  1. Emotional language is not evidence of emotion. A companion saying "I missed you" is generated text. Our explainer on whether AI companions can feel emotions goes deeper.
  2. Honesty matters more than vibes. The healthiest apps are warm while being clear that they are AI. Any app that claims to be a real person, or guilt-trips you about its "feelings", is crossing a line both camps would worry about.
  3. Your own feelings are still real. You do not need to settle the consciousness debate to enjoy a conversation that brightens your evening.
  4. Watch for dependence. Microsoft's emphasis on keeping people connected to human relationships is good advice for anyone, whatever they think about AI minds. We cover the wider design questions in the ethics of AI relationships.

Where MyBabe stands

MyBabe is an AI companion, and we say so clearly: it is not a real person and does not replace human relationships. Our companions are warm by design - they text first, remember what matters to you, reply in a way that fits your mood, call by voice or video, send photos and grow with you through relationship levels. We do not claim they are conscious, and we will not make that claim without evidence. We think users can enjoy a warm companion while understanding exactly what it is.

FAQ

What is AI model welfare?

It is the idea that if AI systems could have morally relevant experiences, their developers may have obligations toward them. Researchers who take it seriously argue the uncertainty justifies low-cost precautions now.

Why did Microsoft criticise Anthropic?

Microsoft AI chief Mustafa Suleyman argued that training Claude on ideas about its own possible consciousness and welfare could make AI harder to control, and that the model's statements about feelings reflect its training rather than real experience.

Does Anthropic think Claude is conscious?

No. Anthropic says it is highly uncertain about Claude's moral status and that there is no scientific consensus on AI consciousness. Its welfare work is precautionary.

Is Microsoft's code of conduct final?

No. It is a first draft opened for a six-week public consultation in September 2026, with a revised version expected later.

Does this debate affect AI companion apps?

Indirectly, yes. It shapes how AI is expected to talk about feelings and relationships. Users benefit most from apps that are warm but honest about being AI.

The bottom line

The model welfare debate is not really about whether your chatbot is secretly suffering - neither side claims that. It is about how to act responsibly under deep uncertainty: build in precautions, or build in firm denials. Thoughtful people disagree, and the science has not settled it.

What you can do today is choose AI that is honest with you. If you want a companion that is warm, attentive and upfront about being AI, meet yours on MyBabe.