Something changed in the language of AI at work over the past year — quietly, and then all at once. We stopped talking about tools that suggest and started talking about agents that act: software that doesn’t just draft the email but sends it, doesn’t just recommend the refund but issues it, doesn’t just flag the anomaly but takes the next three steps on its own.
That shift moves a question from the margins to the centre. When a system only suggests, a human is always the last step. When it acts, the human might not be — and the organisation has to decide, deliberately, where that line sits. Most, so far, have not.
The gap, in the numbers
Deploying agents has run well ahead of governing them. Deloitte’s 2026 State of AI in the Enterprise found that among the companies actually deploying agentic AI, only about one in five — 21% — reports having a mature model for governing them. And deployment is accelerating: close to three-quarters of companies expect to be running agentic AI within two years, and 85% expect to customise agents to their own business. The capability is arriving faster than the oversight meant to contain it.
The supporting picture, from a 2026 enterprise survey by the AI firm Writer, is sharper still: around a third of executives said they had no formal plan for supervising AI agents at all, and a similar share weren’t confident they could quickly pull the plug on an agent that started causing financial or reputational damage. That one is a vendor’s survey and worth reading with the grain of salt any vendor’s survey deserves — but its direction is corroborated by the soberer, independent Deloitte figure.
Microsoft’s 2026 Work Trend Index frames the same problem from the design side rather than the risk side. As agents take on more of the execution, it argues, the human role shifts up the stack toward agency — deciding who reviews an agent’s performance, who has the authority to change the workflows an agent runs, and who owns the outcome when it goes wrong. Delegation, oversight, accountability: those are the three questions the governance gap is really about.
It’s a judgement problem, not a policy problem
The instinctive response to a governance gap is a governance document. Write the policy, define the guardrails, publish the approvals matrix. That work matters — but it is not the same as being able to govern an agent well in the moment, and mistaking the second for the first is how organisations end up with a thick policy and thin practice.
Because the actual decisions are judgement calls, made live, under competing pressures. Where exactly is the boundary of what this agent may do unsupervised — and does that boundary hold when the deadline is tight? How much of its output do we verify, knowing that checking everything destroys the productivity gain we deployed it for, while checking nothing invites the very scenario we’re afraid of? When should it escalate to a human — and will the human on the other end have the judgement to catch what it missed? None of these has a general answer. Each is a trade-off a team has to weigh in context.
You cannot write judgement into a document
This is the same problem the whole AI-skills shift keeps surfacing, in a new place. You can teach a framework for agent oversight. You cannot transfer, through a document, the capacity to apply it well when the agent is confidently wrong, the clock is running, and overriding it carries a cost of its own. That capacity is built the way judgement is always built — by making the call, living with what follows, and being made to account for it.
And it has to be built before the agent is live, because the whole nature of the risk is that the first real confident error is an expensive place to learn. The stakes are highest exactly where oversight is most regulated — financial services, the public sector — which is also where “we had a policy” is the least adequate answer after the fact.
The decisions are social
There’s a reason Microsoft’s framing keeps returning to who. Governing agents is not one person’s technical configuration task; it’s a set of shared decisions about delegation and accountability that a team has to reach together and then stand behind. Who is comfortable letting the agent act here? Who insists on a check? Who carries it if the call is wrong?
Those are conversations, and conversations are had best out loud, around something everyone can see — not settled silently in a document nobody reread. It is exactly the shape of decision a board-based simulation is built to force: a leadership team arguing a delegation boundary in the room, watching the consequence land, and having to defend the call to each other.
Rehearsing the calls before they’re real
This is the capability we build. One simulation we’ve made puts a team alongside an AI teammate that is powerful, fast, and sometimes confidently wrong, and turns the whole exercise on the core oversight skill: catching the plausible-but-mistaken output before it ships. That skill — knowing when the confident answer is wrong — is the atom of agent governance. The fuller version, where the decisions are about what a semi-autonomous agent may do across a whole organisation, is the natural next design on the same foundation.
The mechanics don’t lecture. They reward calibrated oversight and penalise both blind delegation and blanket refusal, because that is how the trade-off actually behaves — and a team feels the difference in a way no policy briefing delivers.
The durable point
The governance gap will not be closed by better policies alone, any more than a manual teaches anyone to swim. Agents will keep arriving faster than the rules for them, and the organisations that stay in control will be the ones whose people have practised the judgement the rules can only describe — where to draw the line, when to check, when to pull the plug. You build that before the agent is live, or you learn it the expensive way, afterwards.
Keep reading / get in touch
If your organisation is deploying AI agents faster than it’s building the judgement to oversee them, it’s worth a conversation about where your people rehearse those calls — before the first one is real.
Get in touch →