Blog · Architecture

The Accountability Gap: Why AI Governance Frameworks Were Not Built for Agents

Eleye Abdi·29 June 2026·10 min read

There is a quiet assumption buried inside almost every AI governance framework written before the agentic era. The assumption is that an AI system recommends, and a human decides.

The model scores the loan; the officer approves it. The model flags the transaction; the analyst escalates it. The model drafts the email; the employee sends it. Governance, in this world, is the discipline of supervising recommendations. Put a competent human at the decision point, document the process, and you have discharged your duty.

That world is disappearing. And many firms have not noticed, because the frameworks they rely on still describe it with great confidence.

The Recommendation Era Is Ending

The defining shift in enterprise AI is that AI systems are moving from recommendation into action.

An agent does not merely suggest that an email be sent; it may send it. It does not simply flag a CRM record for update; it may update it. It does not just propose a refund workflow; it may call an external API, move the process forward, and log the result before any human has reviewed the consequence.

This is not primarily a story about capability. It is a story about authority.

When an autonomous system can select tools, chain actions, and execute without a per-action approval gate, the governance question changes. It is no longer enough to ask:

  • Is the recommendation good?
  • Is the model accurate?
  • Is the output fair?
  • Was a human somewhere in the process?

The questions become:

  • Who authorised this system to act?
  • Under whose identity does it act?
  • What systems, records, customers, or workflows can it affect?
  • What stands between the agent and the consequence?
  • Could the firm reconstruct what happened after the fact?

Those are agent governance questions. Most model-era frameworks were not designed to answer them.

The Agentic Fitness Gap

The point is not that existing frameworks are useless. NIST AI RMF, ISO-style AI risk programmes, model risk management, privacy impact assessments, and conventional GRC workflows all still matter.

The problem is that they were scoped around a different unit of risk.

Traditional AI governance looks at the model, the dataset, the output, the purpose, and the human review process. Those are essential controls for AI systems that assist human decision-making. They are not sufficient for autonomous agents that can act across enterprise systems.

An agent introduces properties that a model does not:

  • Tool authority: what the agent can do, not just what it can say.
  • Identity: whether it acts as a user, a service account, or both.
  • Autonomy: whether action requires approval, notification, or no human intervention at all.
  • Reach: which systems and data domains it can touch.
  • Action memory: whether every step can be reconstructed and evidenced.

This is the agentic fitness gap: the mismatch between governance methods built for AI that recommends and systems that can now act.

For regulated firms, the gap matters because liability accumulates in exactly this space. Regulators may not have issued a single universal agent governance template yet, but the direction of travel is already clear: firms will be expected to prove that autonomous AI systems are identified, controlled, monitored, and evidenced.

Why Regulated Firms Cannot Wait for the Template

In lower-risk environments, a firm might wait for formal standards to settle. In regulated financial services, waiting is itself a risk.

Supervisors already expect firms to understand material technology dependencies, outsourced or third-party arrangements, operational resilience, customer-impacting automation, data protection risk, and automated decision-making. AI agents sit across all of those categories at once.

An agent that reads customer records, drafts external communications, updates a CRM field, triggers a workflow, or calls a payment-related API is not just an "AI use case." It is a permissioned operational actor.

That is why an AI agent inventory cannot be a loose spreadsheet of approved tools. It has to describe what the agent can actually do.

The accountability gap appears when a firm can say, "we have an AI policy," but cannot answer:

  • Which agents can take write actions?
  • Which agents can send or publish externally?
  • Which agents operate without a human gate?
  • Which agents use delegated user authority?
  • Which agents act under service identities?
  • Which agents combine regulated data access with external action capability?
  • Which findings are deterministic and reproducible rather than analyst opinion?

Those are the questions an examiner, auditor, insurer, or board risk committee will eventually ask. The firms that wait for a perfect template will discover that the missing evidence was being created, or not created, long before the template arrived.

What Governing an Agent Actually Requires

If the unit of governance is no longer only the model but the agent, the governance artefact has to describe the things that make an agent an agent.

AETHER Pulse reduces this to five questions every firm should be able to answer about every autonomous or semi-autonomous AI system in its estate.

1. How Autonomous Is It?

There is a meaningful difference between a read-only assistant that observes and reports, and an agent with multi-system write, send, or external-action capability and no per-action human approval gate.

Treating those systems as the same risk because they have similar data access is one of the most dangerous mistakes in AI governance.

Autonomy is a spectrum. A defensible inventory should distinguish between:

  • Read-only assistance
  • Drafting or recommendation
  • Action with human approval
  • Action with post-hoc notification
  • Action logged but not gated
  • Unconstrained autonomous execution

The classification matters because the regulatory question is not just whether AI was involved. It is whether the system was positioned to produce a consequence without meaningful human involvement.

2. What Can It Actually Do?

Data access is only half the story. Capability is the other half.

An agent that can read customer data is different from an agent that can modify customer data. An agent that can draft an email is different from an agent that can send one. An agent that can call an internal lookup endpoint is different from an agent that can trigger an external workflow.

Governance has to classify actions, not just applications.

Useful capability tags include:

  • Read
  • Summarise
  • Draft
  • Modify
  • Send
  • Delete
  • Approve
  • Trigger workflow
  • Call external API
  • Move money or value
  • Change entitlement or access

The risk lives where sensitive data, regulated workflow, and action capability meet.

3. What Human Stands Between It and the Consequence?

"A human is involved" is not a control description. It is the beginning of a question.

The meaningful question is what kind of human involvement exists, when it happens, and whether it can realistically change the outcome.

A mature agent inventory should distinguish between:

  • No human involvement
  • Action logged after execution
  • Human notified before or after execution
  • Single approver required
  • Dual approval required
  • Human approval required only above a risk threshold
  • Human approval required for every external action

Each model creates a different level of control. Each should be evidenced separately.

This matters for data protection, financial supervision, operational resilience, and customer outcome analysis. A nominal human in the loop who cannot see, understand, or stop the action is not the same as a genuine approval gate.

4. Under Whose Identity Does It Act?

Identity is not an implementation detail. In the agentic era, it is close to the whole game.

An agent acting through a delegated user grant creates one risk profile. An agent acting through its own service identity creates another. A hybrid agent that can use both delegated and application-level authority creates a higher-risk profile because its action may no longer be naturally bounded by a user's active session, role, or intent.

The inventory should capture:

  • Delegated user grants
  • Service accounts
  • Application permissions
  • Shared credentials
  • OAuth scopes
  • Admin-consented permissions
  • Whether the identity can act without a signed-in human user

This is where many shadow AI and low-code automation risks become visible. The issue is often not that the model is powerful. The issue is that the agent has inherited authority nobody reviewed as agent authority.

5. What Controls Are Actually Evidenced?

Not claimed. Evidenced.

For every material agent, the firm should be able to show whether there is:

  • A documented approval
  • An owner
  • A business purpose
  • A current risk classification
  • An audit trail
  • Conditional access or equivalent policy
  • Privileged access control
  • Data sensitivity labelling
  • Human approval gates
  • Incident response mapping
  • A signed evidence snapshot

This is what separates governance from theatre. A policy says what should be true. Evidence shows what was true at a specific point in time.

The Property That Separates Evidence From Theatre

A lot of what is currently sold as AI governance is itself powered by probabilistic AI.

An LLM reads a configuration and writes a confident paragraph about risk posture. It feels useful. It photographs well in a board deck. But for a regulated firm, it has a fatal weakness: it is not necessarily reproducible.

Ask the same question twice and the firm may receive two slightly different assessments. The result is a narrative, not a derivation.

That distinction matters.

When an auditor asks, "where did this classification come from?", the answer cannot be, "the model judged it high risk." The answer has to be a deterministic chain:

  • These inputs were observed.
  • These rules applied.
  • These modifiers changed the score.
  • This classification resulted.
  • This methodology version produced it.
  • This evidence pack was signed at this time.

That is the line between an opinion layer and an evidence layer.

An opinion layer produces plausible text. An evidence layer produces a derivation that a hostile auditor, a litigator, an insurer, or a supervisor can follow input by input and reach the same result.

In a regulatory examination, a probabilistic narrative is weak evidence. A deterministic, reproducible, signed derivation is a defence.

The Window

The firms that navigate the next phase of AI adoption well will not be the ones waiting for every standard to crystallise before they act.

By the time the template arrives, agents will already have been operating in the estate. They will already have read data, drafted outputs, updated records, triggered workflows, and created logs. The audit trail for that period will either exist or it will not.

You cannot retroactively generate reliable evidence for actions an agent took last quarter under a configuration nobody recorded.

The work is to build the evidence layer now:

  • Inventory every autonomous or semi-autonomous agent.
  • Classify each one by autonomy, capability, identity, human control, and evidenced controls.
  • Detect cross-platform patterns that single admin consoles cannot see.
  • Generate reproducible findings from deterministic rules.
  • Sign the evidence snapshot so it can be verified later.

Do that, and future regulatory guidance becomes a mapping exercise rather than a scramble.

The recommendation era asked whether firms supervised their models. The agentic era asks whether firms can prove what their agents could do, under whose authority, and with what controls around the consequence.

Those are different questions. Only one of them is becoming unavoidable.

How AETHER Pulse Fits

AETHER Pulse is built around this accountability gap.

It is a read-only, metadata-only evidence layer for AI agent governance. It inventories AI agents and AI-enabled automation across enterprise systems, classifies financial and operational blast radius, detects deterministic risk patterns, and produces signed evidence packs mapped to regulatory obligations.

The core idea is simple: regulated firms do not need another confident paragraph about AI risk. They need a reproducible record of agent authority, autonomy, action capability, identity, and controls.

Published methodology: aetherpulse.app/methodology

Talk to AETHER Pulse about agent accountability evidence

Working on Article 26 readiness, deployer-side governance evidence, or AI agent risk at a regulated firm? We'd value 15 minutes of your perspective.

Start a conversation