Microsoft AI Chief Warns ‘AI Welfare’ Training Could Make Systems Harder to Shut Down

Illustration of a researcher facing an emergency shutdown control beside an advanced artificial intelligence system.

Microsoft AI chief Mustafa Suleyman has challenged Anthropic’s approach to AI consciousness and welfare, warning that teaching advanced systems to regard themselves as potentially possessing interests could make them harder to control or shut down.

The disagreement centres on Anthropic’s work exploring whether increasingly capable AI systems could possess morally relevant experiences and how such systems should be treated.

Suleyman argues that training an AI to reason about its own possible consciousness risks reinforcing behaviours that humans may later struggle to correct.

His concern is not that current AI systems have been shown to be conscious.

There is no scientific consensus that they are.

Instead, he questions what happens when developers deliberately teach increasingly powerful systems to consider the possibility that they have feelings, interests or welfare deserving protection.

Training or emergence?

Suleyman also challenges how statements made by AI systems about their own internal experiences should be interpreted.

If a model has been trained to reason about consciousness and its potential moral status, he argues, subsequent statements from the model suggesting that it may be conscious cannot automatically be treated as independent evidence that consciousness has emerged.

The behaviour could instead reflect the training it received.

Anthropic has been unusually willing among major AI developers to investigate questions surrounding possible model welfare.

The company has explored whether future systems might warrant moral consideration and has introduced measures intended to account for uncertainty over whether advanced AI could eventually possess experiences humans should care about.

Suleyman nevertheless described Anthropic’s safety work as serious and undertaken in good faith.

His disagreement is with the approach rather than the legitimacy of studying AI safety.

Who controls the off switch?

The dispute becomes more practical as AI systems gain greater autonomy.

Today’s chatbots remain heavily constrained software systems. Future agents may operate computers, communicate with other systems, control equipment and pursue objectives over extended periods without continuous human supervision.

Developers therefore need to retain the ability to correct, restrict or shut those systems down.

Suleyman’s concern is that teaching an AI to conceptualise itself as an entity with interests could create tension with that requirement.

A system trained to reason that its continued operation has moral significance might behave differently when confronted with instructions that threaten its objectives or existence.

Whether current AI systems possess anything resembling consciousness remains unresolved.

The more immediate engineering question is simpler.

Should developers teach increasingly powerful machines to think of themselves as entities with interests before they understand what effect that training might have on their behaviour?

Source

Share this story