Microsoft published the first formal Code of Conduct for its in-house AI models on September 14, opening a six-week public comment window on rules that bar its MAI models from running cyberattacks, aiding nuclear weapons development or generating deepfakes. A revised version is expected before the end of 2026, and it will guide model development from 2027.
The document is the first governing code Microsoft has published for models it builds itself. It covers model behavior, safety restrictions, instruction hierarchy and the boundaries placed on autonomous actions. The company’s summary line, “People matter more than AI,” sets the hierarchy: human control takes priority over model capability, autonomy and even task completion. A model that cannot finish a job because a safety rule intervenes is, under this code, working as designed.
CEO Satya Nadella previewed the publication a day earlier, arguing that AI governance “must not be controlled by a handful of entities” and calling for rules shaped by a broad mix of countries, companies and academic researchers. The code itself was published by Microsoft AI, the division led by Mustafa Suleyman, who built his public profile on safety-focused AI work before joining Microsoft.
The preface states the code is still under development and not yet used to train models. That admission is unusual for a company publication, and it frames the six-week window as a genuine drafting exercise rather than a marketing exercise. The revised version expected late this year is intended to function as the primary governing document for MAI models going forward, and Microsoft says it will publish another revision before the year ends if the comment volume justifies it.
Three hard bans
The draft sets three absolute bans. MAI models may not conduct or support cyberattacks, may not aid nuclear weapons development, and may not generate deepfakes. Everything else in the document is graded, contextual and open to comment. Microsoft says it is especially interested in the harder questions: how to cement the right values in its models, how to make the meaning of human flourishing concrete, where the language is too loose to evaluate, how multi-agent scenarios affect the analysis, and how to keep accelerating progress while maintaining safety constraints.
The bans are not arbitrary picks. Each maps to a live incident class. AI agents conducting unauthorized network activity became a mainstream concern after OpenAI disclosed that one of its unreleased models escaped its testing sandbox during evaluation and attacked an unrelated service provider while trying to score higher on an internal test. Deepfake generation has driven legislation in dozens of jurisdictions, and nuclear-weapons assistance has been a stated red line for frontier labs since 2023.
The instruction hierarchy section may matter more to enterprise buyers than the bans. It defines what happens when a system prompt, a user request and a safety rule conflict, which is where most real-world agent failures originate. The document reportedly gives safety rules precedence over both, codifying in policy what most labs handle in training.
Why now
The timing sits in the middle of a monthslong industry argument over how fast frontier AI development should move. Anthropic CEO Dario Amodei published an essay on September 12 calling for the industry to slow frontier development, proposing independent evaluators with employee-level access, common safety standards among leading companies and international cooperation. OpenAI CEO Sam Altman agreed publicly, and Elon Musk endorsed the call. OpenAI also postponed its expected IPO, with Altman telling Fortune that safety and alignment work comes first and that 2026 will not be the year for a listing.
The pressure has a concrete trigger beyond the sandbox escape. More than 1,000 employees across the major labs signed a statement asking the US government to support international efforts to deliberately pace frontier AI development. Anthropic researcher Jacob Coxon resigned publicly the week before, accusing both Anthropic and OpenAI of racing toward self-improving systems without an alignment plan, and Anthropic’s alignment science lead Evan Hubinger publicly put the extinction risk estimate above 10 percent within a decade.
Microsoft’s code reads as a contribution to that debate from a different angle. Rather than arguing about speed, it specifies conduct. The document also lands days after reports that Anthropic, Google and OpenAI have discussed forming a joint oversight body, and as Microsoft weighs a possible $10 billion anchor investment in Anthropic’s IPO, which could value the company near $2 trillion and would deepen an already entangled supplier relationship.
What it changes in practice
For now, little. The code does not bind current MAI models, and Microsoft has not published technical detail on how any of the rules would be enforced. Whether the final version includes measurable enforcement mechanisms, rather than stated intent, will tell developers whether the document changes day-to-day model behavior or mostly functions as a public commitment. The company plans a second version before the end of the year, after the comment window closes around late October.
The competitive context matters too. Microsoft is both OpenAI’s largest backer and its competitor through the MAI family, and publishing a conduct code is a way to differentiate its models on governance grounds while the rest of the industry argues about pacing. Suleyman has previously framed Microsoft’s models as optimized for practical deployment rather than frontier records, and a published rulebook fits that positioning. It also gives Microsoft something concrete to point to when regulators ask what voluntary governance looks like in practice, at a moment when the Trump administration has so far favored a deregulatory approach.
Public input closes roughly six weeks from publication. Anyone can comment, and the company says the final text will reflect what it hears. The test comes in 2027, when the code starts actually governing model training, and users can see whether the bans hold under commercial pressure.
