In a highly regulated industry, “NOT SURE” is not a status, it is a critical issue that needs to be fixed.
62% of financial services firms have already deployed AI agents, and 93% of the ones who have, granted those agents some form of real autonomy before they had a way to fully see what that autonomy was doing. That is not a research lab statistic. That is the Cloud Security Alliance’s own survey of 340 cloud, AI, cybersecurity, and risk professionals across the industry, fielded this past winter and published in June. Read it twice. The autonomy came first. The visibility is still catching up.
I have sat in enough CIO and CTO rooms this year to know that number does not surprise anyone in the room. What surprises them is the other one sitting next to it: 1 in 5 of those same firms has already had a known AI-related security incident, and another 1 in 5 is not sure whether it has or not. Not “no.” Not sure. In a regulated industry, “not sure” is not a status. It is an admission that nobody can currently answer the question.
Regulators moved, and moved the wrong thing out of scope
Here is the part that should stop technical leaders mid-scroll, the same way it stops risk officers. In April of this year, the OCC and the Federal Reserve issued the first meaningful rewrite of bank model risk management guidance since 2011. SR 26-2 and OCC Bulletin 2026-13 replaced a framework that had not been touched since the iPhone was four years old. And in the same breath that they modernized it, they wrote generative and agentic AI out of it entirely. The language is direct: these models are “novel and rapidly evolving” and therefore “not within the scope of this guidance.” The agencies have said they intend to issue a request for information on AI model risk at some point in the near future. Near future is doing a lot of work in that sentence.
So the fastest-growing category of risk inside the bank is now the one category with no dedicated supervisory framework attached to it, at the exact moment agent autonomy is accelerating past the industry’s ability to observe it. Deloitte’s own count puts the number of distinct risks that can emerge from autonomous behavior in banking systems above 350. JPMorgan has more than 400 production AI use cases running. Goldman Sachs has agents doing trade accounting and compliance work. Lloyds is projecting nine figures in annual value from agentic AI. This is not a someday conversation. It is Tuesday.
I understand the instinct to read a regulatory gap as a green light. No rule yet reads as no rule broken yet. That reading has the physics backwards. Examiners are not waiting for the RFI to become a framework before they start asking questions in exam season, because governance, authority, and kill switches were already fair game before April, and nothing about the exclusion took those questions off the table. What changed is that the bank now has to answer them without a rulebook to point to. Every agent shipped between now and whenever the RFI resolves into something concrete becomes something a technical leader has to be able to reconstruct after the fact, on demand: who approved this agent’s authority, what could it actually touch, what happened the day it was wrong, and can you produce the trail in an afternoon instead of a quarter.
Why the old playbook doesn’t transfer
The honest read of that 62%, 93%, 1-in-5 number set is not that banks are being reckless. Most of the technology leaders I work with are trying to move carefully inside real pressure to show board-level progress on AI. The problem is structural. A traditional model risk framework, the one SR 11-7 built and SR 26-2 just retired, was designed around a static artifact: a model, validated once, monitored on a fixed schedule. An agent is not that. An agent is closer to an employee whose authority keeps expanding, taking actions, chaining decisions, sometimes invoking other tools or other agents on its own initiative mid-task. Bolting 2011-era model validation paperwork onto that produces a binder that satisfies an audit checklist and misses the actual risk, which lives in what the agent is permitted to do without asking, not in whether its outputs were statistically sound at launch.
That distinction is the whole story hiding inside the 62/93 gap. It is not that most firms have zero governance. It is that the governance many of them have was built for a different kind of system entirely, and nobody has gone back to check whether the muscle still fits.
Below is the loop I have started asking every engineering and platform leader I sit with to trace before they scale a single additional agent past a pilot. It is not complicated. It is just rarely built in from day one, which is exactly why it is hard to reconstruct on day four hundred.

Every agent proposal, execution, and escalation writes into an evidence store that a human, and eventually an examiner, can actually query. Not a quarterly PowerPoint summarizing what happened. A system that already knows.
What “governance built in” has to actually mean
If you are building or buying agentic AI inside a regulated enterprise right now, the bar worth holding yourself to is not “does this satisfy current guidance.” There is not much current guidance to satisfy, by design, for now. The bar is closer to this: if an examiner walked in tomorrow and asked you to trace exactly what one of your agents was authorized to do, what it actually did, and who signed off at each escalation point, could you answer inside a day, not a quarter.
That is the design question I keep coming back to on the delivery side of every AI-DLC engagement I run. The audit trail and the authority boundary get built into the system as it is built, the same sprint the capability ships, not retrofitted once the RFI turns into a rule three firms already got burned waiting for. I have watched engineering organizations treat this as a compliance tax to pay later. The ones ahead of the pack this year treated it as part of the architecture from the first Bolt, the same way test coverage or observability stopped being optional a decade ago.
The regulatory vacuum will not last. SR 26-2 all but promises that. What every technical leader builds while it is open is what the organization will still be standing on when the framework finally lands, and that framework will be written, in no small part, from what regulators observe happening in the vacuum right now.
I have spent this year inside enough of these engineering organizations to believe the firms that win the next exam cycle will not be the ones who waited for the RFI. They will be the ones who built as if the strictest plausible version of it was already law. Where is your organization actually building the accountability loop today, not the version in the deck you show the board, the version an engineer could point to in the codebase? Tell me what you’re seeing. I want to know where the real gap is in your shop, because I promise you it is not where the compliance team thinks it is.
Keep growing
Gunjan



