Bank examiners now ask about AI. Policy documents will not do
Federal bank examiners in the United States have started treating artificial intelligence as a standing line item in routine examinations, according to Syed Ali, chief technology officer for global financial services at Ensono, an IT services firm. He dates the change to June 2026 and says the questions are not soft ones: what technical limits sit on model behaviour, how human review is actually structured, and whether the bank could switch a misbehaving system off.
In written answers to The Datatech Times, Ali sets out what examiners are asking, what a defensible answer looks like in practice, and why he thinks every large bank will soon need a team whose full-time job is watching AI agents work.
"The shift came in June 2026, after the Office of the Comptroller of the Currency (OCC) and the Federal Reserve observed three in four banks not being able to confirm if they could switch off a misbehaving model or how to ensure responsible, accurate use," he says. That, in his account, pushed regulators to make AI review a standing item in every routine exam rather than something reserved for special reviews. Examiners now ask what AI is in production and what decisions it touches, whether models read only approved data, and whether vendors and their subcontractors are held to the same standards.
Controls, not documents
"Examiners will not accept policy documents and employee training as acceptable controls. They will expect well-configured controls and visible permissioning as empirical proof. If limits sit in a document instead of digital permissions, it is not a limit. It is a hope."
The controls he describes include clear definitions of what a system can read and access, whether it proposes or executes actions, and swimlanes marking any influence a system could have on a transaction. "A clear ceiling on what AI can do per transaction and per hour offers empirical evidence of safeguards that examiners will look for in ensuring readiness."
The kill switch, easy to claim and harder to evidence, gets the same treatment. "There is a big misconception that a kill switch is like one big red button that halts activity. In practice, it is a process and chain of events just like any other crisis response scenario. There must be a clear chain of command on which executives and leaders can pull the plug, what triggers the call, and has anyone rehearsed it." Regulators want to know by name who has the authority to enact the plan, he says, and that a system can be stopped without breaking something else. A system that has never been turned off on purpose is one nobody knows can be turned off. "Disaster recovery has maintained similar standards for more than 30 years. Nobody accepts 'we are confident the backup site works' without testing it."
Agent operations
The role Ali describes as missing does not have a settled name. "Call it 'agent operations,' a set of first-line moderators that sit within the business, but report into the risk management function." Model validation, he says, asks roughly once a year whether a model was fit for purpose, and compliance teams care whether a rule was broken; put the new function under compliance and it becomes another form of document review. Risk teams understand exposure and impact before they happen, and hold a real stake in agents behaving as expected.
"These employees would fill the above gaps in current oversight function by owning monitoring of live AI agent behaviour. By watching the model at two in the morning when the agent starts drifting on its four-hundredth task, this team airgaps crises and plays an active role in the success of implementations."
On who fills the job: "To succeed in agent operations, these associates must be able to read what the system is doing without an engineer translating and know processes well enough to tell unusual from wrong." He would recruit from operations and payments, "because the AI can be taught, whereas banking and process judgment are hardest to learn."
Institutions get the balance wrong in both directions, he says, and often within the same building. "An example of too little oversight is when a drafting assistant quietly became a decision maker, because the human approving is clearing 40 items an hour. That is a signature, not oversight. On the other hand, too much oversight looks like a checkpoint on every step, which offers no real benefit and can actually make the system less safe, because nobody is reading item 50 of a 1,000. That is a rubber stamp reported to the board as a control." Oversight should be heavy where a decision is irreversible or touches a customer, he says, and light everywhere else.
Where the frameworks run out
Financial services already has model risk management frameworks, and Ali's view is that the newest of them was written to leave the hardest cases out. "OCC Bulletin 2026-13 replaced SR 11-7 as the default standard in April 2026, and by design explicitly excluded generative and agentic AI guidance, with regulators promising separate guidance for these novel and rapidly evolving tools." The consequence, he says, is that the highest-risk categories of AI sit outside the only formal framework built to govern models, "and given it has yet to be published, a bank can be fully compliant with the new rule and still be exposed to some of the systems it is deploying fastest."
Until that guidance arrives he still sees value in the older model risk discipline. Establishing and maintaining an inventory, ownership and documentation lets a bank challenge its own systems and get into the habit of proving safeguards work rather than asserting it. But the older frameworks miss three things about agents, he says. They assume a model produces an output and a person acts on it, where an independently operating agent needs new standards for who is accountable for its actions. They assume single functions that can be tested in isolation, where agents link workflows together and a failure may live anywhere along the chain. And they validate behaviour once a year, where an agent's behaviour changes with every new piece of context it receives.
In the meantime, he says, examiners in the field are borrowing the NIST generative AI profile and Treasury control objectives as interim standards, which he reads as a signal that supervisors would rather see a bank improvising thoughtfully than waiting for an addendum. No date has been given for the separate guidance.
What a bad day looks like
Asked what a realistic failure looks like, Ali does not reach for a dramatic one. "Without agent oversight, a realistic bad day starts with an agent doing something high in volume but low in value, like fee reversals or hardship cases, but getting one rule subtly wrong. On any single file you open, the decision might look defensible, but doing so 11,000 times over a long weekend creates a real snowball of challenges if you do not find out until Tuesday."
"The cost is not the error in that situation." It is reconstructing the systems, remediating the decisions that were made improperly, and answering an investigator asking when you knew, followed by a consent order. "Having thorough controls in place saves you by catching and flagging an issue in hour two, not on day four."