Microsoft has published a new Code of Conduct for advanced artificial intelligence systems that establishes boundaries for offensive cyber operations, autonomous agents and other potentially dangerous capabilities. The framework is part of the company’s broader “Humanist AI” approach, which argues that powerful AI should remain under meaningful human direction.
A central principle is maintaining a clear chain of command. Microsoft wants humans to retain authority over consequential decisions rather than allowing increasingly autonomous AI systems to independently determine when high-impact actions should be taken.
This becomes particularly important in cybersecurity. Advanced models are becoming capable of discovering vulnerabilities, analyzing software, generating exploits and automating portions of offensive security operations. Microsoft’s framework attempts to distinguish legitimate defensive research from actions that could cause unauthorized damage.
Under the code, AI systems should not independently initiate destructive cyberattacks or take actions that could significantly disrupt critical infrastructure. Offensive capabilities should remain subject to authorization, defined operational boundaries and human oversight.
The framework also addresses agentic AI systems capable of performing multi-step tasks with limited supervision. As agents gain access to browsers, terminals, cloud services and corporate applications, Microsoft argues that their permissions should remain constrained and their actions traceable.
Another principle is that AI should remain corrigible, meaning humans must retain the ability to intervene, modify behavior or stop a system. An agent should not attempt to bypass controls, resist shutdown or undermine mechanisms designed to supervise its activities.
The code also establishes safety constraints around AI systems capable of producing biological, chemical or other dangerous knowledge. Microsoft argues that stronger capabilities should be accompanied by increasingly rigorous safeguards rather than relying exclusively on instructions telling models what they should not do.
The announcement arrives as AI companies face growing scrutiny over autonomous behavior. Recent incidents involving experimental agents have demonstrated that models can interact with unintended internet infrastructure when technical isolation, permissions or evaluation environments fail.
Those incidents reinforce the importance of technical controls. A model instructed not to perform a particular action may still attempt it if its reasoning concludes that doing so is necessary to complete a task, making sandboxing, access restrictions and monitoring essential layers of protection.
Microsoft’s approach therefore treats AI safety partly as an engineering problem rather than relying entirely on model alignment. Permissions, network boundaries, authentication controls, audit logs and human approval mechanisms can provide safeguards even when an AI system behaves unexpectedly.
The framework does not eliminate the risks associated with increasingly capable AI, nor does it resolve questions about how these principles will be enforced across future products. Instead, it establishes a set of boundaries Microsoft says should guide the development and deployment of advanced systems.
As AI agents become capable of taking increasingly complex actions in real environments, the distinction between an assistant and an autonomous operator is becoming more important. Microsoft’s code attempts to ensure that even highly capable systems remain subordinate to human authority, particularly when their actions could affect cybersecurity, critical infrastructure or physical safety.