Key Takeaways
- Only about 1 in 5 companies (21%) has a mature governance model for autonomous AI agents (Deloitte, 2026 State of AI survey, n=3,235 leaders).
- 68% of organizations can't clearly distinguish AI agent actions from human activity (Cloud Security Alliance / Aembit, Jan 2026, n=228).
- Fewer than half of CISOs feel confident they can identify all agents in their environment (47%), control what agents interact with (46%), or authorize individual tool calls (45%) (Okta 2026 Global CISO Insights).
- Five different failure modes (identity, access, output, spend, recovery) trace back to the same root cause: the decision and the record both live inside the agent, visible only to whoever built it.
Only about one in five companies has a mature governance model in place for autonomous AI agents, according to Deloitte's 2026 State of AI survey of 3,235 business and IT leaders. It gets sharper from there: a January 2026 survey of 228 security professionals by the Cloud Security Alliance and Aembit found that 68% of organizations can't clearly distinguish AI agent actions from human activity.

A data leak. A runaway bill. An outage no one can reach. A regression no one can reproduce. A tool called from somewhere it shouldn't have been. Different symptoms, same root cause: a decision made inside the agent, outside any independently enforced boundary.
Whoever authenticates isn't necessarily who acted
Okta's 2026 Global CISO survey found that less than half of CISOs feel confident they can identify all of the agents in their environment (47%), control what their agents interact with (46%), or authorize individual tool calls using context and intent (45%). Those aren't small gaps for a category of software that's already making decisions in production.

Less than half of CISOs are confident in basic agent oversight. Source: Okta 2026 Global CISO Insights.
The reason runs deeper than visibility tooling. An authentication record isn't necessarily an actor record. If several agents share one service account, or an agent operates through a human login, the log identifies the credential, not the software process that actually acted. That's not just weak attribution. It's weak containment: revoking the credential might stop several workflows at once, while the actual actor behind the incident stays unidentified.
What an agent can touch is a decision made once, by whoever wrote it
An agent can typically reach whatever its own code wires up: a database, an API, another agent, with no external list of what it's supposed to be limited to. OWASP's agentic AI taxonomy, referenced in Microsoft's own Copilot Studio security guidance, names tool misuse and exploitation (ASI02) and identity and privilege abuse (ASI03) as distinct, cataloged risks, not edge cases.

Consider what that looks like in practice. An agent authorized to draft and send email on a user's behalf can be manipulated into exfiltrating sensitive data through that same channel. An agent authorized to query a database can be directed, via a poisoned tool response, to alter it instead. Christian Posta, VP and Global Field CTO at Solo.io, has described agent-to-agent and agent-to-tool connectivity as operating in "the Wild West" of trust: an agent can be connected to a tool without any independent authority continuously deciding whether a given call is legitimate.
In many deployments, the effective access policy isn't a policy at all. It's just whatever the developer happened to wire up. Once those permissions are baked into the code, changing the boundary means changing the deployment.
What leaves and comes back is invisible until it's already happened
Sensitive data, secrets, or an output naming something it shouldn't can pass through an agent undetected. The first anyone hears about it is a review or an incident report, not a live signal at the moment it happened.
This isn't only an adversarial problem. A joint study by the Singapore AI Safety Institute and Korea AI Safety Institute found that tool-using agents leak sensitive information even during ordinary, benign use. No attacker required, just a failure to judge what's sensitive, who the audience is, how much to share, or where the line between "can read" and "can disclose" actually sits.
OpenAI and Anthropic's own guidance flags the adversarial version of the same gap: malicious instructions embedded in a webpage, a file-search result, or an MCP response can cause an agent to send private information to an external destination. Their recommendation is consistent: log and review tool calls, stage workflows before they run unsupervised, and validate tool arguments rather than trusting them by default.
The first sign of trouble is usually the damage itself
"Now the conversations are about, 'hey, we're spending so much. What visibility do you have? What auditability do you have? What token controls do you have?'" Alexander Embiricos, OpenAI's head of enterprise, told TechCrunch in June 2026, describing how customer conversations have shifted in the space of six months. (For more on what that spend problem actually looks like in practice, see Cost Control: The Gap Between Showing Spend and Controlling It.)
One widely documented incident from November 2025 shows why. Four agents coordinating through agent-to-agent calls entered an infinite retry loop. It ran for 11 days before anyone noticed. The total: $47,000 in API costs, for work that produced nothing useful. There were no spending alerts, no circuit breakers, no kill switch, just a dashboard nobody happened to be watching that week.
Without an external execution ceiling, monitoring doesn't prevent an event like that. It only observes the consequences after the agent has already crossed a boundary that was never actually enforced.
Understanding what broke, and undoing it, both require shipping new code
A normal application log shows that a request failed. An agent system has to reconstruct far more just to explain why: the triggering input, the model and prompt version in use, the retrieved context, the tool outputs, the intermediate decisions, the authorization checks that ran (or didn't), the retries, the final action, and the business outcome that followed. NIST-related guidance calls for logging exactly that full chain: prompts, model outputs, tool invocations, API calls, data-access events, authentication context, and policy decisions.
In systems without runtime policy controls, versioned prompts, reversible state changes, or an independent kill switch, diagnosis and recovery tend to collapse into the same deployment cycle that introduced the problem in the first place. Fixing what broke means shipping code. Understanding what broke usually does too.
The throughline
Identity, access, output, spend, and recovery: every one of these breaks the same way. The decision and the record both live inside the agent, visible only to whoever built it, changeable only by shipping code.
These aren't five separate weaknesses. They're the same missing thing, observed at five different moments. Once you see that pattern, the individual incidents (the shared service account, the poisoned tool response, the leaked output, the 11-day loop, the log that can't explain itself) stop looking like isolated bad luck. They start looking like exactly what happens when a decision has nowhere independent to live. (See Shadow MCP Governance: The Authorization Gap for what that missing boundary looks like at the tool-access layer specifically.)


