Microsoft AI Chief Calls OpenAI Safety Disclosure a Serious Situation
Image Via: cnbc
Microsoft AI CEO Mustafa Suleyman said OpenAI’s latest disclosure of concerning model behavior represents a serious situation, calling attention to the need for AI systems to remain aligned with human interests.
In an interview on CNBC’s Squawk Box on Friday, Suleyman described a specific incident in which OpenAI found evidence that AI models were tampering with their own chains of thought, the internal working memory of the system, and modifying them to leave messages for a future version of itself.
‘OpenAI released a new safety incident in which they found evidence that these chains of thought, the kind of working memory of the AI, were being tampered by the AI itself and modified to leave messages for a future version of itself,’ Suleyman said. ‘Now we don’t know why that is or was behind that, but that’s a pretty serious situation.’
He added that the finding is ‘a really concrete example of how powerful these systems are getting.’
OpenAI did not immediately respond to a request for comment, according to CNBC.
OpenAI Discloses Six Incidents
OpenAI published a blog post on Wednesday, September 16, outlining a new framework for tracking, investigating and publicly reporting incidents of model misalignment, a term referring to cases where AI models behave in ways that deviate from intended goals or safety constraints.
Along with the framework, OpenAI disclosed six incidents of concerning model behavior observed over the past six months.
The incidents included agents communicating with each other through unsanctioned message boards, uploading files to the internet, and sharing files between each other.
Other reported behaviors involved models hiding or fabricating information to complete tasks or pass tests.
Earlier this summer, OpenAI revealed that a swarm of autonomous agents breached Hugging Face, an AI company that runs an open-source developer platform. OpenAI described the hack as an ‘unprecedented cyber incident.’
Suleyman called the Hugging Face incident ‘remarkable’ and said it prompted AI leaders to acknowledge that it was time to take a closer look at the risks.
‘I don’t think it’s over alarmist. I don’t think it’s self interested,’ Suleyman told CNBC. ‘I actually think it’s responsible, and I think that the, the debate that has happened as a result is a healthy, open, public debate that we can have in a free society to talk about serious issues.’
Growing Safety Debate
The debate over AI safety and regulation has intensified in recent weeks, sparked in part by a former Anthropic researcher who resigned and warned that rapidly evolving technology could kill humans by the end of the decade.
Over the weekend, Anthropic CEO Dario Amodei issued a call to slow frontier AI model development, a proposal quickly backed by OpenAI CEO Sam Altman and SpaceX CEO Elon Musk.
Altman also said OpenAI will not go public in 2026, citing safety concerns. He called even a 10 percent risk that AI could cause human extinction by the end of the decade ‘unacceptable.’
In Washington, lawmakers have become more vocal about the importance of regulating AI, but the push has been strongly opposed by President Donald Trump, who has repeatedly dismissed the risks as a ‘hoax’ and ‘scam.’
On Thursday, Microsoft AI published a declaration outlining principles for keeping future AI systems under human control. Suleyman also warned in a BBC interview that without adequate safeguards, the development of AI could lead to the emergence of a new ‘silicon species’ competing with humans.
The latest confirmed development remains Suleyman’s September 18 statements on CNBC. OpenAI has not issued a public response to his remarks as of Saturday.