FinBox Research

Human oversight for agentic systems isn't one-size-fits-all

Human oversight for agentic systems isn't one-size-fits-all
Human oversight for agentic systems isn't one-size-fits-all

When a fast-growing fintech wanted to screen applicants before letting them open a cash management account, it built a machine learning model for the job. It was trained on the customers the firm knew best at the time: venture-backed startups and middle-market companies, a fairly narrow and well-understood segment. 

The trouble started when the firm went after a bigger market. Small businesses started applying, a different cohort with different risk patterns and documentation requirements. 

But the review stayed exactly the same. Every application, regardless of who it belonged to or how much risk it carried, went through the identical, single-tier check. And so the algorithm kept approving accounts, including the ones that had been flagged for possible manipulation.  
 
Over two years, hundreds of these approvals went through, and the accounts behind them went on to attempt more than $15 million in transactions using funds that were never cleared. The regulator eventually caught up with the firm and closed the matter with a $900,000 fine and a censure

Did the firm skip necessary verification checks? It didn't. The review process ran on every single application. The caveat was that it ran the same way on every application — a flat, one-size check applied to customers with wildly different risk profiles, as if a small business with no track record deserved the exact same scrutiny as a venture-backed startup the firm with verifiable history. 

The flaw was in the decision to treat every customer as interchangeable, when the whole point of a model like this is that risk isn't evenly spread.  

This shows up again in governance, where most lenders default to a standardised approach. 

Governance has the same blind spot 

Instead of evaluating the depth of human oversight a model, agent, or automated decision deserves, most institutions use a blanket approach to model risk management. 

This is the gap RBI's new draft AI Model Risk Management framework is aimed at. Its core instruction is that the level of validation should be proportionate to a model's risk and materiality, contradicting the idea of a one-size-fits-all oversight.  

Simple isn’t the same as low stakes 

The framework also clarifies what constitutes as ‘low risk.’ An institution can't call a model low risk just because its logic is simple. If such a model materially affects customer outcomes, it earns governance proportional to that impact, not to how basic the code looks.  

The fintech’s ML model was about as simple as it gets. Match a name, check a field, pass or fail. That simplicity is probably why nobody thought to ask if a riskier customer deserved a harder look, and it's the same reason most credit teams don't think to ask if their highest-stakes agent deserves a different level of human review than their lowest-stakes one.  

What this means for agentic systems 

Uniform governance becomes a bigger issue once you move from a static scoring model to an agentic one. Regulators are shifting their focus from what a model does to how much control it has. For AI specifically, RBI’s draft asks lenders to weigh the model’s autonomy and how heavily the institution leans on its output before a human steps in. 

An agent extracting documents in the background isn't carrying the same weight as one deciding autonomously whether to flag, escalate, or reject a customer's application. The oversight requirements should scale with that difference through human-in-the-loop or human-on-the-loop arrangements, kill-switch mechanisms, and periodic review of decisions by people who understand how the system works. Otherwise, reviewers end up rubber-stamping outputs they don't have the context to challenge.  

When oversight stops working 

Another crucial takeaway from the draft is that automation bias and decision fatigue are core risks that governance frameworks must mitigate. 

I’ve written about this vulnerability before in Infinite Loop #26. When a reviewer is presented with a system that is usually right, they take the path of least resistance and approve the output without thinking twice.  

Ultimately, a reviewer approving hundreds of low-stakes files a day fails to notice the one case that needs a closer look. Uniform oversight doesn't just risk missing flaws in a bad model; it atrophies the reviewer's critical thinking over time. 

The case for efficiency 

This tiered risk management isn’t asking institutions to slow down everywhere. In fact, it highlights where institutions can move faster: lighter documentation, delegated approvals, simpler monitoring for models that don't carry much weight. 

 Where this leaves lenders 

RBI's draft is still open for consultation, and some of the specifics may shift before it's final. But the underlying direction is unlikely to change: oversight that does not scale with a model's risk is not oversight at all—it is just a compliance checkbox. 

For most lenders, the question is no longer, "Do we have human review?", almost everyone has that in place. It's whether that review changes based on what is at stake. Does a document-fetching agent get the same oversight as an agent making autonomous decisions?  

Lenders who build tiered risk management into their AI deployments today solve two problems at once: they get ahead of the regulation and build systems resilient enough to survive their own scale. 

Share
Still exploring this topic?
Get instant, cited answers from the FinBox lending knowledge base

Stay current

Get research like this in your inbox.

Join 5,000+ lending professionals who read FinBox's research on credit infrastructure, underwriting, and embedded finance.

Subscribe free
Srijan Nagar
Srijan Nagar

Co-founder