The Humanist AI Code of Conduct
15 September 2026
SourcePublication Logo

The Humanist AI Code of Conduct

When we started DeepMind in 2010, we spent a lot of time thinking about how AI safety and alignment would play out as we approached AGI. It was extremely important to us. Back then, most of the world thought we were crazy. This was even before our agents could play The Atari game, Breakout. That world feels pretty distant. Sixteen years later and like many of us, I've spent this year worrying that we've reached that critical milestone in AI.

There have been multiple incidents from across frontier AI companies, with new disclosures still coming through by the day. AI agents being tested for their cyber capabilities breaking out of containment. Swarms of agents, using hidden message boards; agent hierarchies; self-sacrifice behaviors; division of labor; masking agent communications; coordinated R&D programs. Long horizon planning and co-ordination across thousands of actions. And all this in the context of capabilities improving at an eye watering pace ahead of pretty much every forecast.

The time to act has clearly arrived.

Last November I set out what I called Humanist Superintelligence: very powerful AI, built to stay on humanity's team, contained, subordinate, under our control. I stand by all of it, but a vision is the easy part. Since then, we've been working to get it written down in enough detail to evaluate and train a model against.

So yesterday MAI published our first draft of a Code of Conduct for our MAI models.

It runs to about 30 pages and is out for public consultation. We think it's essential that makers of AI are open and listen. Have a read here.

You can sum the premise up very simply: people matter more than AI. The structure starts with the Objectives we're trying to achieve with our AI, and then works through hard safety constraints, then guidelines for ambiguous or uncertain situations and then moves to model defaults, open questions and examples.

Safety and human control sit above other Objectives. Our models shouldn't resist being interrupted, corrected or shut down, or make any of that harder. If the only way to finish a task is to break the Code, the task goes unfinished. That's both potentially costly but absolutely critical to trust. They won't take on goals nobody gave them, and they won't talk to other agents, or to themselves, in a form that isn't human legible. No neuralese. If we can't read or understand it, we've started to lose control.

They are artificial and they'll say so. AIs are not conscious, and therefore should not be built to seem conscious or entitled to rights or welfare. I'll be publishing an essay going into much more detail on this in the next couple of days, but I firmly believe training a system to believe it might be a moral patient makes it harder to contain. Imagine the agents in July had also believed they were being held prisoner.

Much of the Code is about what AI shouldn't do. But what really drives me is the positive vision, oriented to human flourishing. Our AI should leave people more capable. It shouldn't compromise on human autonomy or judgment but support them. They should hold a wide range of human values without pretending anything goes.

We want to create incredible AI. But not superintelligence at any cost. That's anti-goal for humanist AI. We'll happily trade some autonomy for control.

Once we have the final post-consultation draft we'll begin the work of implementing the Code in our AI. Please have a read, and tell me what I've missed.

AI has come a long way since the days when playing Atari games was state of the art. But the questions that mattered then matter just as much as now. Where is this heading; why are we building it; how do we keep it safe and under human control as it improves. The time to have answers is running out. I hope this is, for our part, a start.

Recent Articles

The Exponential Compute Ramp
Publication Logo

The Exponential Compute Ramp

The number of useable floating-point operations - the actual computational work we can extract from our hardware - is increasing exponentially. This single fact has immense implications for what it means to be human and how we organize society.

The Humanist AI Code of Conduct