A draft five-stage framework from the MACH Alliance Agent Adoption & Operations Working Group. September 2026 Draft.
Most of the enterprises we talk to have an AI governance policy. The harder thing to produce is the line between that policy and a particular agent: which of its rules apply to the agent a delivery team put into production last month, what the platform enforces on that agent, and where the evidence sits.
That gap is structural. AI governance gets funded as an enterprise program: a single standard, covering every way AI is used across the organization, including which doesn’t always include agents. Agreement across that many owners is slow to reach, and agents keep going live while it is being negotiated. Delivery teams ship into the gap, often on shared service accounts, and reconstructing afterwards what an agent decided, and on whose authority, is expensive.
The MACH Alliance Agent Adoption & Operations Working Group has published the Agent Governance and Assurance Framework, a September 2026 draft that starts from the opposite end. One team can govern one agent in a sprint, and the company-wide standards accumulate out of that work instead of preceding it.
Who it is for
CIOs and enterprise architects being asked to approve agents faster than their governance function can assess them, and the delivery teams who need something concrete to complete.
Sooner or later somebody asks five questions about an agent you own. Internal audit asks them, or a regulator, or a customer’s security review, or your board.
- What can this agent do with no person involved?
- Which data can it read, and which systems can it change?
- Who approved it, and on what date?
- What happens when it goes wrong?
- Can you prove what it did?
There is no shortage of documents setting out AI principles. This one names the documents that answer those five questions, and specifies how they connect.
What it is
Five stages. Each one checks the stage before it.
Stage | What you write | What it does |
|---|---|---|
1. Agent design | Agent Design Document | Records what one agent is for, what it can touch, and what it may do alone |
2. Policy | Risk Classifications and the Classification Test | Sets out which kinds of agent you build, and places this one in a kind |
3. Platform controls | Platform Control Levels, and a Platform and Sign-off Record | Holds the agent to the limits its design declares, and gets it signed before it goes live |
4. Live operations | Monitoring Levels, and a Monitoring and Incident Preparedness Record | Keeps your evals running against the live agent, and says what happens when one fails |
5. Audit and assurance | Audit Log Rules, and an Audit and Assurance Record | Keeps proof of what the agent did, and proof that somebody is still reading the evals |
The connections carry the weight. A rule in your policy holds only if some agent’s design applies it, the platform enforces it on that agent, a named person has accepted the risk that is left, live behavior confirms the description, and the records prove it. Break any one of those links and everything above it is an assertion. The framework also expects revision: each stage produces findings against the stage before it, which is why every artifact carries a version history.
What it helps you do
Get one agent into production with evidence you can put in front of an auditor, inside a quarter, without waiting for a central program.
The first pass is small by design. Classifying one agent requires exactly one risk classification, not a scheme covering the entire organization. You score your platform against two capability ladders and record Level 1 if that is where you are; most teams running their first agent are. You run whatever evals you already have against the configuration you actually deployed, model version included, since results from any other configuration describe that configuration and not the one you shipped.
A named person then signs off on the agent: what it does, its classification, the platform it runs on, what the evals found, and the risk that is left. The sign-off record does two things a policy document usually does not. It gives a shortfall somewhere to live, so a control gap is accepted by name with a closing date.
Classification works as a test rather than a judgment call, for the same reason. Eight questions, each of which sets a minimum classification, and the agent lands on the highest minimum any single answer produces.
Plain words, on purpose
Risk and compliance writing carries a large vocabulary, and agent governance is inheriting all of it. A control that the people applying it cannot parse does not get applied. So the draft uses the plain term: the risk that is left for residual risk, how far the damage spreads for blast radius, an agent login for a non-human identity. The glossary maps each one to its formal equivalent, so the same document still works in a room with an auditor. Terms already clear to the people doing the work stay as they are, including least privilege and eval.
Go and argue with it
We published the structure now so that it can be challenged while changing it is still cheap.
Read the five stages for the reasoning. Then take one of the forms, complete it for an agent you actually run, and tell us where it failed you. We need to know which questions cannot be answered and which ones are missing, whether five stages are the right five, and whether something sits in the wrong place. Open an issue on any part of it.
Members of the working group will be at MACH X: Amsterdam on September 29 and 30. We would be glad to meet you there.

The Agent Adoption & Operations Working Group is part of the MACH Alliance Agent Ecosystem initiative. Its earlier three-part position paper, What Changed, argues that frameworks are the organizational mechanism for scaling the mental model shift agentic systems require.
