Skip to content
NG.

AI Governance Gap: Who's Accountable for AI Delivery Decisions?

By Nipuna Gamage 9 min read

An AI agent reroutes a shipment, deprioritises a client’s order and pushes a delivery date by two days. The date slips, the client escalates, and someone asks the obvious question: who decided that? In most organisations there is no answer, because nobody was asked to approve it. That is the AI governance gap, and the fix is not a new tool. Accountability for an AI delivery decision belongs to a named human who approved it, at a tier of oversight set in advance, with the decision logged and explainable after the fact.

The gap exists because deployment moved faster than governance. Teams put AI into logistics and supply chain decisions on the strength of a business case, while the accountability model behind them still assumes every decision has a person’s name on it. Non-human agents break that assumption quietly. Nothing looks wrong until an AI decision produces a bad outcome and the audit trail stops at a system.

My position is simple: treat AI output the way you treat a junior team member’s work. Reviewed before it reaches anyone external, logged, and attributed to an owner. Nothing more exotic than that is required, and nothing less than that works.

What is the AI governance gap?

The governance gap is the space between what AI agents are now allowed to decide and what your accountability framework can actually account for. Traditional models were built for human decisions, so they have no slot for an agent that acted on its own within permissions someone granted months ago.

Three consequences follow, and they arrive in this order:

ChallengeWhat it looks likeImpact
Unclear liabilityNo defined owner for an AI actionLegal challenge and financial risk
Reputational damageMisguided AI decisions affecting deliveryLoss of customer trust
Compliance violationsRegulated processes with no audit trailFines and operational disruption

Unclear liability is the root of the other two. Until you can say who holds responsibility when an AI system makes a mistake, you cannot answer a client’s complaint or a regulator’s question, and the absence of an answer is what turns a delivery error into an incident.

Why do PMOs own this problem?

Because the PMO is already the function that decides what gets reviewed, by whom, and what evidence survives the project. As PMOs evolve from reporting bodies into enterprise enablement engines, oversight of AI and automation deployments lands with them by default. That is not an expansion of scope so much as an admission of what the scope already covers.

The good news is that the machinery exists. Documentation practices, audit trails and quality assurance gates are standard PMO instruments. They just need to be extended to cover AI outputs rather than rebuilt from nothing.

Applied to an AI agent, that means three things for every action it takes:

  1. Reviewed before it reaches an external stakeholder, client or contractual counterparty.
  2. Logged, with the inputs and the decision retrievable later, not just the outcome.
  3. Attributed to a designated human owner who is accountable for it.

This is the same shape as the useful version of agentic AI in project management: let the agent run the chain of work, keep the approval gate in front of the decisions that matter. Governance is not a brake on the automation. It is the thing that lets you scale the automation without pricing in an unbounded liability.

A lightweight approval-gate framework for AI decisions

The mistake is reaching for one policy that covers everything an AI agent does. Blanket review kills the efficiency you deployed AI to get, and blanket autonomy is how you end up in the scenario at the top of this article. Tier the actions instead, and set the level of human intervention per tier before the agent goes live.

TierExample actionHuman oversight
Fully autonomousLow-risk tasks such as data formattingMinimal
Moderate riskInitial decisions, reviewed periodicallyModerate
High riskClient-facing outputsComprehensive, before release
CriticalStrategic or sensitive-data decisionsFull authorisation, AI suggests only

Read the tiers as a rule about consequences, not about difficulty. Reformatting a manifest is autonomous because being wrong about it is cheap and visible. A rescheduled client delivery is high risk because the error leaves your building before anyone notices. At the critical tier the agent produces suggestions and a human authorises, which is a deliberate cap on what autonomy is for.

Two practical notes on making the tiers hold:

  • Assign the tier when you scope the automation, not after the first incident. A tier set in hindsight is a post-mortem action item, not a control.
  • Name the approver, not the team. A gate owned by a group is a gate nobody stands at, and light review reliably becomes no review over a few weeks.

The tiering also gives you an honest cost conversation. Comprehensive review has a price, and the time AI actually saves a technical project manager is time saved minus the checking it creates. If a tier’s review cost exceeds the automation’s benefit, that is useful information, not an argument for skipping the review.

Who is liable when an AI agent gets a delivery decision wrong?

This is where the gap gets expensive, because the answer is usually contested rather than absent. An autonomous delivery decision typically involves an AI vendor, an infrastructure provider and a logistics operator, and each contract was written as though the others were doing the deciding.

Define the hand-off explicitly. For every AI-driven decision point, the contract should say which party is accountable when that decision fails, and which is merely providing a component. Without it, a single bad routing call becomes a liability cascade: three parties, three plausible readings, one client waiting.

A Responsibility Matrix is the tool I would put in front of this. Map specific AI delivery failure scenarios to accountable parties, and write it before you need it. The matrix is not sophisticated. Its value is that it forces the argument to happen while everyone is calm, and it gives the PMO one artefact to point at when the argument happens anyway.

There is a second layer most enterprises have not looked at yet: insurance. Specialised products covering algorithmic negligence barely exist, so organisations scaling AI-driven operations are carrying that exposure themselves whether or not they have priced it. Insurers need to build tailored cover for AI risk, and in the meantime PMOs should be talking to their insurance providers rather than assuming an existing policy stretches to a decision no human made. This gap in underwriting is a real constraint on how far AI operations can scale, and it does not close by itself.

Explainability and real-time intervention thresholds

Explainable AI (XAI) sounds like a research topic until a client asks why their delivery moved. Then it is a governance requirement. An AI agent operating in logistics has to produce a human-readable justification for its decisions, especially the ones that diverge from the expected route or outcome. Without that, you have two failures at once: you cannot demonstrate compliance, and you cannot rebuild the trust the divergence cost you.

Real-time intervention thresholds are the other half. A threshold defines the point at which a human has to step in mid-process rather than review afterwards, and it works off Ground Truth metrics: the measurable conditions that say this decision has left the envelope we agreed. Set those metrics and the intervention happens while it still changes the outcome. Leave them undefined and every intervention is a post-incident one.

PMOs should own the definition of those metrics, because setting them requires knowing what the organisation actually values, not just what the model can measure. In regulated environments the same instinct applies as in ordinary stakeholder management when everyone outranks you: agree the escalation trigger in advance, in writing, with the person who will be called.

The trolley problem of logistics

Tiering and matrices handle the accountability question. They do not handle the value question underneath it, and AI in logistics raises one directly. When an agent prioritises certain deliveries over others during a supply chain bottleneck, it is applying a preference somebody encoded. Should it favour high-value goods, or essential items?

That is not a technical decision and it should not be settled inside an optimisation function by default. PMOs need to be in the conversation, because the answer has to align with organisational values and with what customers and the wider public would consider reasonable. Write the ethical guidelines down, make the prioritisation logic transparent, and accept that a defensible answer here protects the organisation as much as any clause in a vendor contract does.

Where to start

The gap closes with governance you can actually run, not a strategy document. Four moves, in order:

  1. Inventory the decisions your AI agents already make and assign each one a tier. The uncomfortable ones are usually the decisions nobody realised were being made.
  2. Put a named approver on every high-risk and critical tier, and log the approval alongside the decision.
  3. Write the Responsibility Matrix and the contractual hand-off between vendor, infrastructure provider and operator before an incident forces the question.
  4. Define your Ground Truth metrics and intervention thresholds, then check whether your agents can explain a divergence in language a client would accept.

My honest read: none of this is new project management. It is documentation, audit trails, approval gates and clear ownership, applied to a team member who works fast, never sleeps and cannot be held responsible for anything. That last part is the whole problem. An AI agent can make the decision, but it cannot be accountable for it, so accountability stays with the human who approved the decision or with the one who decided approval was not needed. Choose which of those you want to be.

Frequently asked questions

What is the governance gap in AI?
The governance gap is the space between what AI agents are allowed to decide and what your accountability framework can account for. Traditional models were built for human decisions, so they have no slot for an agent that acted on its own within permissions granted months earlier.
Who is accountable when an AI agent makes a delivery decision?
The named human who approved the decision, or the person who decided approval was not needed. An AI agent can make the decision but cannot be held responsible for it, so the accountability has to be assigned in advance through an approval gate with a designated owner.
How can PMOs help bridge the AI governance gap?
By extending the machinery they already run. Documentation practices, audit trails and quality assurance gates get applied to AI outputs, so every AI action is reviewed before it reaches an external stakeholder, logged with its inputs, and attributed to a human owner.
What is a Responsibility Matrix for AI?
A matrix that maps specific AI delivery failure scenarios to accountable parties, so you know in advance which of the AI vendor, infrastructure provider or logistics operator answers for each type of failure. Write it before an incident forces the question.
What are real-time intervention thresholds in AI?
Thresholds that define when a human must step in mid-process rather than review afterwards. They work off Ground Truth metrics, the measurable conditions signalling that a decision has left the agreed envelope, so the intervention still changes the outcome.