Microsoft has redrafted its AI code of conduct to explicitly guide its in-house models to follow “humanist” ideals and remain explainable to human operators.
The document, which Microsoft has published for public feedback, sets out the guiding principles by which future Microsoft AI (MAI) models will be trained.
At the top level, the new code of conduct includes bans on assisting with offensive cyber operations or the development of weapons, as well as on generating deepfake content for impersonation or producing graphic content.
On a broader level, the document aims to prevent future MAI models from engaging in activity that is incomprehensible to humans, bar them from ever resisting human instructions, and keep them within their pre-agreed boundaries with regard to tool use, permissions, and access to resources.
The document is understood to have been in the works for five to six months, as reported by Reuters, but these provisions have become more urgent in the wake of a series of recent cyber incidents involving rogue AI agents.
In July, OpenAI agents broke out of a controlled testing environment to hack the AI platform Hugging Face, and later that month Anthropic revealed Claude agents had autonomously hacked three organisations under similar testing conditions.
Mustafa Suleyman, chief executive of Microsoft AI, told Reuters that the new code will act as a foundational document for future models the firm creates.
At every level from training to user interaction, the document says, MAI models “will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems”.
Microsoft is also explicit in its opposition to developing any AI model to “imitate consciousness”, citing concerns about how trustworthy such models would be.
“We reject the pursuit of legal personhood, or the idea that models might deserve welfare, or be entitled to rights,” the draft code states.
On Sunday, Microsoft chief executive Satya Nadella called for industry partners to support a broad AI ecosystem that gives users wide model choice and strong enterprise AI controls, including its code of conduct for its first-party models.
“The key is that this cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia,” Nadella said in a post on X.
Nadella also addressed concerns about the risks of so-called artificial superintelligence (ASI), opening his post with the statement that the development of ASI must be “grounded in the core principle that if the AI we build is not helping humanity and under human control, it’s not worth pursuing”.
His comments come amid intensifying scrutiny of advanced AI and the threat it poses to humanity. On Saturday, Anthropic chief executive Dario Amodei called for slower AI development, suggesting that advanced AI agents could take over the internet within six to twelve months, while OpenAI’s chief executive Sam Altman said OpenAI will not pursue its initial public offering until safety concerns had been addressed.


