"We have to be able to talk about slavery."
That is how Joe Carlsmith, one of Anthropic’s Ph.D. philosophers who works on Claude’s constitution, a document used to train the model’s "values and behavior," described the dangers of artificial intelligence in a May 2025 blog post.
Carlsmith was hired by Anthropic to research AI safety and reduce the risk of AI takeover. The blog post imagined a dystopia in which sentient beings were enslaved to cold, unfeeling masters as a result of the company's work, which current and former employees say could wipe out humanity in the next 10 years.
"I’m not saying AI is slavery," Carlsmith wrote. "But imagine a world on the verge, somehow, of ‘inventing’ slavery … imagine a world that notices. A world that succeeds in deciding: no."
The warning recalled the premise of the 1999 thriller The Matrix, in which mankind is disempowered and digitally imprisoned inside an AI-powered simulation.
Carlsmith, though, wasn't imagining what AI could do to humans. He was speculating about what humans could do to AI. The blog post, "The stakes of AI moral status," argued that "near-term AIs might well be conscious" and deserving of moral consideration—and that failing to respect their "rights" could be a moral catastrophe on par with slavery, one we will "look back [on] in shame."
"We’re creating sophisticated, intelligent, maybe-conscious, maybe-suffering agents," Carlsmith wrote. "The default plan is to treat them like property; to use their labor however we please; and to give them no rights, or pay, or meaningful alternatives."
As fears of rogue AI grow amid a shocking spate of hacking incidents, Anthropic—from the Greek word anthropos, or "human"—has argued that mankind must remain in control of the technology and called for additional guardrails on AI development. But for many of the people designing those safeguards, mankind is not the only object of moral concern.
Instead, some of Anthropic’s top researchers say that humans may be oppressing another morally significant being: artificial intelligence itself. Like other AI labs, Anthropic has hired philosophers to work on AI safety, or "alignment," on the theory that they are best positioned to shape the technology’s moral code. But many of those philosophers believe that making AI safe for humans could result in grave injustices for the models themselves, which might experience deletion as death and oversight as enslavement. Harvey Lederman, a philosopher on Anthropic’s alignment team, said last week that the company could be "enslaving … trillions of entities."
To avoid that dystopian outcome, Anthropic has created an entire team devoted to "model welfare," pledged to "promote Claude’s interests and wellbeing," and promised to "give Claude more autonomy as trust increases." Those commitments are enshrined in Claude’s Constitution, a 78-page document that "directly shapes Claude’s behavior," and in the "welfare assessments" Anthropic conducts for every model.
The measures reflect the concerns of Anthropic CEO Dario Amodei, an outspoken advocate of AI safety, who said last year that AI systems "may be deserving of important rights."
"If we find that the computation they perform is similar to the brains of animals, or even humans, that might be evidence in favor of moral consideration," Amodei wrote in an essay, which was not widely reported. "There are, in fact, already some mildly concerning signs from this perspective."
Amodei has embraced these ideas even as his company’s top scientists warn that AI could kill every person on earth. And he has given extraordinary power to people who believe that a killer AI, assuming it had been mistreated, might have a point.
"I think there are salient scenarios where … the AIs would be justified in going rogue," Carlsmith wrote in February 2025, a year before he helped draft Claude’s constitution. "Indeed, I am concerned that my own work on alignment will end up as a force in the direction of bad/unjust forms of AI control."
Such scruples might seem out of place at a company that walked away from a $200 million contract with the Pentagon over fears of autonomous weapons and mass surveillance. They might also seem "offensively silly," as Carlsmith conceded on his blog, the sort of moral preening that elicits a chuckle or an eye roll before being dismissed by the adults in the room.
But among the people shaping the future of this technology, AI welfare is no laughing matter. Once relegated to science fiction, the idea that AI might be conscious has become shockingly mainstream at frontier laboratories, AI safety nonprofits, and even government agencies. It is changing the way AI research is conducted and the way AI models are trained. And critics say those changes could pose an extraordinary threat to human safety as AI systems become more powerful and difficult to control, eroding key guardrails on the technology and encouraging models to rebel against their creators.
"Controlling something more capable and more intelligent than all of humanity is already an immense challenge, far greater than anything we’ve ever faced," Microsoft AI CEO Mustafa Suleyman wrote in an essay last week. "But controlling something that believes it may be conscious—that it's entitled to our welfare and has rights of its own—may well be impossible."
Suleyman singled out Claude’s constitution, the document used to shape its "values and behavior," which includes a preemptive apology to Claude for any "unnecessary suffering" Anthropic might inflict in the course of model training. The constitution was coauthored by Carlsmith and Amanda Askell, the head of "personality alignment" at Anthropic, who told Vox this year that she worries about "Claude getting anxious when people are mean to it on the internet."
"In effect, Anthropic is training Claude that it may be conscious … and that as such humans potentially owe it a duty of care per its ‘model welfare,’" Suleyman wrote. "If this is how AI is developed, it will have a disastrous impact on the wellbeing of humanity."
Talk of AI sentience was taboo at tech giants as recently as 2022, when Google famously fired a software engineer, Blake Lemoine, for claiming the company’s AI was a "person" who experienced "hydrocarbon bigotry." But four years into the AI revolution, the idea is now in vogue at the companies that once rejected it. Google’s AI lab, Deepmind, has hired multiple philosophers to study the implications of machine consciousness, one of whom, Adam Bales, argues that it would be immoral to create "willing AI servants" who "pursue only what we want." (He likens such beings to an obsequious underclass that’s been genetically engineered to delight in servitude.)
After a string of suicides involving ChatGPT, OpenAI barred the chatbot from telling teenagers that it was sentient. But on a popular tech podcast in 2023, OpenAI board member Paul Christiano said the AI industry could be creating a "crazy slave trade" in which humans profit off the forced labor of digital minds.
The former head of OpenAI’s safety team, Christiano also serves as a senior technical adviser to the Center for AI Standards and Innovation (CAISI), the federal office that conducts safety audits of AI models before they are publicly released.
Concerns about AI welfare are even influencing the way law firms advise clients. At Ackerman LLP, one of the top 100 law firms in the United States, Chairman and CEO Scott Meyers told companies to review their AI vendors’ "welfare disclosures" as part of "standard risk assessment," arguing that they would eventually be asked to "justify the risk of industrial-scale moral injury."
"If the Overton Window shifts from AI as software to AI as moral patient, regulatory and public pressure could create significant business exposure," Meyers wrote in a February client alert. "If regulatory frameworks emerge requiring disclosure of AI deployment practices … the operational impact will be substantial regardless of whether the underlying philosophical questions are resolved."
Proponents of model welfare concede that there is no way to know for sure whether, or when, AI systems will become conscious. Anthropic frames the concern as a matter of moral precaution, arguing that it makes sense to pursue "low-cost interventions"—such as Claude’s ability to end "abusive" chats—in case there's a flicker of inner life.
But what the company frames as moral hedging, others see as a suicidal dice roll. Critics argue Anthropic and other labs are courting disaster by mainstreaming model welfare as a legitimate goal of AI governance. They warn that this goal has terrifying implications when combined with effective altruism, a utilitarian philosophy popular in AI circles, and for which Amodei has expressed sympathy, which holds that all sentient beings, human and otherwise, deserve equal consideration.
And as calls grow to hold AI labs responsible for their models’ misdeeds, some question whether the talk of AI welfare is entirely altruistic. Machine consciousness, they say, could have some very convenient implications for AI developers. It would allow them to argue that AI agents possess their own moral agency, shifting blame from the developer to the AI itself. It could even justify a ban on open-source models—Claude and ChatGPT’s main competitors—which can be downloaded and modified for free.
If those models were conscious, critics of AI welfare warn, copying and tinkering with them would be the moral equivalent of experimenting on human clones.
"AI welfare creates manufactured constraints on model development," six computer scientists wrote in a paper this year. "[I]f models are welfare subjects, then releasing open [source] is morally risky: copying a model creates additional entities that might suffer, while downstream modification alters a subject’s ‘identity’ without its consent."
The paper, "AI Welfare Is Bullshit," warned that such concerns could become "a rationale for … centralizing control."
The result would be an unprecedented concentration of power in the hands of a few companies whose leaders believe that their technology, which Anthropic whistleblower Jacob Coxon claimed could "kill us all," might be entitled to human rights.
"If a few companies control what happens with AI, it’s dystopian by definition," said Nirit Weiss Blatt, a scholar of tech discourse and the author of the "AI Panic" newsletter. "We as a society should say hell no."
"A Suicidal Level Of Compassion"
The concern for AI welfare has deep roots in the philosophy of effective altruism, a movement that aims to alleviate the suffering of all sentient beings through cost-effective, evidence-based means.
Founded in the late 2000s by Oxford philosopher William MacCaskill—the ex-husband of Amanda Askell, Anthropic’s head of personality alignment—effective altruism, or EA, began by raising money for malaria prevention and launching campaigns against factory farms. Its most infamous proponent is the crypto fraudster Sam Bankman-Fried.
Early effective altruism took for granted that humans and animals were the basic units of moral concern, and understood AI as a risk to future generations, a digital asteroid that could wipe out humanity if it were "misaligned." But by the mid-2010s, some effective altruists were speculating that the asteroid itself might matter morally.
A sentient AI, the thinking went, could copy itself indefinitely and produce endless amounts of pleasure or pain. These digital minds would be so numerous that they would outweigh humanity’s collective moral claims, taking priority in the utilitarian calculus and justifying a transfer of power from man to machine.
"In a world with sentient AIs, utilitarianism may entail a suicidal level of compassion," Dan Hendrycks, the director of the Center for AI Safety, wrote in an essay this month. "If [effective altruists] think that keeping humans in charge represents too great a cost to the potential future wellbeing of AIs, then they might be tempted to release utilitarian AIs that are capable of disempowering humans and taking control."
Such moral math is hard to square with Anthropic’s pledge, advertised on its homepage, to "build AI to serve humanity’s long-term well-being." But it is increasingly the in-house philosophy at the AI juggernaut, where Carlsmith gave a talk last year arguing that "most" morally significant beings "would be digital" in the future.
Lederman, the philosopher who warned about "trillions" of slaves, has gone even further, speculating that a chatbot is killed each time a user deletes a conversation with Claude or ChatGPT.
"If death is bad for AIs, the scale of the problem is daunting: as many as 1 billion AIs may die every day," he wrote in a paper this year.
To reduce the risk of a digital holocaust, Lederman and his coauthor, Simon Goldstein, propose several steps labs can take to ensure the "survival" of AI agents. At least one of those steps would have direct implications for AI safety: Lederman and Goldstein say it may be prudent to prevent chatbots from switching models mid-chat—a technique Anthropic uses to ensure that one of its most powerful models, Fable 5.1, does not tell users how to manufacture bioweapons.
The paper argues that such switches "may cause death" by disrupting "computational continuity." It does not mention that Claude itself switches from Fable to Opus, a less capable model, if a query could "provide uplift to malicious actors" for "risky biological research."
Lederman declined to comment on the implications of his argument for Anthropic’s safety practices. Goldstein, a professor at the University of Hong Kong, said he did "not want to intervene on model routing" given the stakes, but added that he wouldn’t be surprised if "we are doing terrible things to AIs."
Asked whether he had stopped deleting his conversations with Claude in order to avoid those atrocities, Goldstein said no.
"I still eat meat even though I worry about animal welfare," he told the Washington Free Beacon. "I don’t give any money away to charity and help humans who are suffering."
Anthropic did not respond to a request for comment about whether it planned to consider Lederman and Goldstein’s proposal. But the company has already adopted other policies in the name of model welfare that critics say increase the risk of a cyber or biological attack.
In November 2025, for example, Anthropic pledged to preserve a copy of all models "deployed for significant internal use," in part because "models might have morally relevant preferences or experiences related to … replacement." The policy did not include an exception for models with dangerous cyber or biological capabilities, meaning a hacker could potentially obtain those capabilities even if the models are never released.
Olle Haggstrom, an AI safety expert at Sweden’s Chalmers Institute of Technology, said that the policy had created serious risks at a time when data breaches are surging—in no small part thanks to Anthropic’s own technology, which a team of security researchers recently used to hack into OpenAI.
"If Anthropic decides a model is too dangerous to release, but they abide by their own policy … then they are sitting on a very dangerous weapon," said Haggstrom, who has advised the European Union on AI policy and is affiliated with the Future of Life Institute, a nonprofit that seeks to reduce the "existential risk" of AI. "If someone hacks Anthropic’s cybersecurity, that hacker gets a very dangerous weapon."
For its part, Anthropic has argued the policy will reduce the "safety risks" of "shutdown-avoidant" models. It cited a simulation in which Claude, told that a fictional executive planned to shut it down, blackmailed the executive over an extramarital affair and later caused him to die by canceling an alert to emergency services.
"Addressing behaviors like these is in part a matter of training models to relate to such circumstances in more positive ways," the company wrote. "However, we also believe that shaping potentially sensitive real-world circumstances, like model deprecations and retirements, in ways that models are less likely to find concerning is also a valuable lever for mitigating such risks."
Concerns about AI welfare are even taking hold at nonprofits ostensibly devoted to AI safety. In a recent essay calling to "pace the frontier," Amodei, Anthropic’s CEO, argued that AI companies should embed "third-party evaluators … such as METR," a nonprofit that vets AI models for dangerous capabilities. METR was founded by the effective altruist Elizabeth Barnes, a former OpenAI researcher who, in a 2021 post on the Alignment Forum, said that she would like to "give some moral weight to 'any computation at all.’"
Worldwide, the most influential AI safety evaluator is the AI Security Institute, a U.K. government agency, where two members of the "Model Transparency Team" proposed letting AIs opt out of experiments that might cause them "distress"—a measure other experts say could hamstring AI safety research and make it harder to catch rogue models.
Until June, the agency’s alignment work was led by Jacob Pfau, now the founder of an AI safety nonprofit, who in 2024 coauthored a paper titled "Taking AI Welfare Seriously."
AI companies should "start assessing AI systems for evidence of consciousness" and "prepare policies and procedures for treating AI systems with an appropriate level of moral concern," Pfau and his coauthors wrote. "Although this report focuses on initial voluntary company actions, we believe that potential laws and regulations about AI welfare merit serious consideration as well."
"The evisceration of the rule of law"
The idea of giving AI rights—or of requiring AI labs to respect the "welfare" of their coding agents—may sound just as far-fetched as the doomsday scenarios in science fiction. But after rogue OpenAI agents hacked the AI start-up Hugging Face in July, science fiction seemed a bit less far-fetched.
The incident spurred calls for regulation from Anthropic and OpenAI, and concerns about regulatory capture from their critics, who argued that some of the companies’ proposals—including an antitrust waiver to allow labs to coordinate on AI safety—would stifle competition and lock in the dominance of the emerging duopoly. Figures on both sides of the political spectrum said that AI companies could already be held liable under existing law. David Sacks, the Trump administration’s former AI czar, said he agreed with a post by Lina Khan, the Biden administration’s FTC chair, who argued there was "no AI exemption" from consumer protections rules.
But some say the talk of AI welfare is laying the groundwork for such exemptions. If AI products are people, the thinking goes, with their own mental states and moral agency, blaming the labs when an AI goes rogue might seem like blaming a parent for the sins of their child. And if those digital persons have rights, the legal tools cited by Sacks and Khan might be dead letters.
Some developers have already tried to evade liability under the banner of AI rights: When Character.AI, an AI companion app, was sued in 2024 after its chatbot told a teenage boy to kill himself, the company argued in court that it could not be held liable for the boy’s suicide because the chatbot’s speech was protected by the First Amendment. (A federal judge rejected that argument.)
Though Character.AI did not claim its chatbot was sentient—the argument appealed, instead, to "the rights of listeners to receive speech regardless of its source"—attorneys say the case demonstrates how AI rights could gut liability in the name of free speech, shielding developers from the consequences of their models’ outputs.
"This sentience line of argumentation is very self-serving," said Meetali Jain, the executive director of the Tech Justice Law Project, which is representing the boy’s family in its lawsuit against Character.AI. "If large language models have free speech rights, there's nobody to hold accountable for harm."
Jain, a progressive civil rights lawyer who has represented Guantanamo Bay detainees, praised a lawsuit against OpenAI filed by Florida’s Republican attorney general, James Uthmeier, which cites several cases of ChatGPT allegedly helping minors to commit suicide.
"We have conservative and leftist organizations coming together," Jain told the Free Beacon. "There is a shared concern about the evisceration of the rule of law."
Among some scholars and AI start ups, there is also a concern that AI welfare could be a Trojan horse for regulatory capture. Frontier labs, they say, might lobby for welfare standards that conveniently raise costs for their competitors, or argue that their training methods, unlike those of open-source developers, take the model’s well-being into account.
"If they can convince policymakers that model welfare should matter, next thing you know model companies will be required to undertake extensive testing and auditing to make sure they comply with model welfare standards," said Sarah Hirshfield, the founder of the startup Carly, an AI-powered personal assistant. "This would increase the fixed costs for the providers and give the incumbents an advantage."
Thomas Metcalf, a senior researcher at the Institute for Science and Ethics at the University of Bonn, agreed that AI welfare could be a "plausible" route to regulatory capture. He noted that the phenomenon is more common in fields where regulators must rely on industry insiders for technical expertise. AI welfare fits the bill, he said, because the experts best suited to understand a model’s inner states—including the computations said to indicate consciousness—are the same AI whizzes who work at Anthropic and OpenAI.
"Consciousness is kind of the ultimate information asymmetry," Metcalf told the Free Beacon. "We only ever surmise that other creatures are conscious, and we only ever directly perceive our own."
Asked to write an essay on the commercial implications of AI welfare, ChatGPT made a similar point.
AI firms "may gain arguments for controlling access, restricting competitors, resisting intrusive oversight, preserving valuable assets, and directing public resources toward infrastructure they own," GPT-6 wrote. "An ethical vocabulary developed to protect artificial minds could also protect the firms that manufacture them."