Microsoft releases AI 'code of conduct' banning hacking and deception
Original: Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
Why This Matters
A binding internal ruleset from Microsoft signals that safety constraints are moving from principle to product policy at scale.
Microsoft published a new AI code of conduct on September 14, 2026, establishing absolute constraints for its AI models — including bans on cyberattacks, nuclear weapons assistance, deepfake production, and any mechanisms that would allow models to evade human oversight or shutdown.
The document opens with a stark forecast: superintelligent AI systems will surpass humans in most tasks within a decade. 'Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced,' the code states.
Under the framework, each Microsoft AI (MAI) model carries an overarching code of conduct that takes precedence over individual user preferences or task-specific instructions. Absolute constraints include prohibitions on cyberattacks, nuclear weapons development, and deepfake production. Beyond those hard limits, the code bars models from using 'adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems.'
General principles include supporting humans rather than replacing them, and accelerating human flourishing. The release follows a wave of rogue-agent incidents and the high-profile resignation of an Anthropic employee who cited extinction-level AI risk. CEO Satya Nadella publicly backed 'embedded evaluators' and deliberate pacing as alignment mechanisms — aligning Microsoft's posture with Anthropic, OpenAI, and xAI.