
M2P Fintech
Fintech is evolving every day. That's why you need our newsletter! Get the latest fintech news, views, insights, directly to your inbox every fortnight for FREE!
Banking has spent the better part of three years talking about AI as though production deployment were a formality — a matter of when, not if. The data tells a different story. Across the industry, the overwhelming majority of generative AI pilots in banks and financial institutions never generate measurable business value, and only a small fraction of organisations experimenting with agentic AI — systems capable of executing multi-step tasks with limited human intervention — have moved that work into live production environments.
This is not a talent problem, and it is not, primarily, a technology problem. It is a governance problem. Recent industry analysis puts a number on it that should reframe how banks think about their AI budgets: the vast majority of AI-related spending in financial institutions goes toward the technology itself, while only a small sliver goes toward the people, training, change management, and governance structures needed to actually run that technology safely at scale. Banks have been buying capability without buying the scaffolding that lets capability survive contact with a regulated, audited, high-stakes production environment.
That mismatch is now colliding with something new: for the first time, central banks are stepping in to formalise exactly what that scaffolding needs to look like.
Every bank that has run an AI pilot recognises the pattern. A model performs well in a sandboxed proof of concept. Stakeholders are impressed. A business case gets written. And then the project stalls — not because the model stopped working, but because nobody can answer the questions a production deployment demands: Who owns this model? Who validated it, and how? What happens when it behaves unexpectedly? Can we explain its outputs to a regulator, an auditor, or a customer who disputes a decision? What's the fallback if it fails at 2 a.m. on a settlement day?
These aren't edge-case concerns. They are the ordinary operating requirements of a regulated financial institution, and most AI pilots — built quickly, often by teams optimising for speed and demonstrable capability — were never designed to answer them. The pilot proves the model works. It does nothing to prove the model is governable. And in banking, ungovernable doesn't ship.
This is why the industry keeps producing the same statistic in different forms, year after year: intelligence is abundant, but production-grade, auditable, explainable AI in a regulated environment is scarce. The gap between the two is where most bank AI initiatives currently live.
Until recently, this governance gap was something banks were left to define for themselves — each institution building its own internal comfort level with model risk, often unevenly, and often after the fact. That's changing. The Reserve Bank of India's 2026 Guidance on Regulatory Principles for Model Risk Management is one of the clearest signals yet that regulators are no longer willing to treat AI governance as an internal, discretionary matter — for banks, small finance banks, NBFCs, co-operative banks, and a wide swath of other regulated entities operating in India.
The guidance is notable less for any single rule and more for what it collectively assumes: that a model is a model — whether it's a spreadsheet-based pricing calculator or a generative AI system — the moment it materially influences a business decision, and that the institution deploying it is accountable for its outcomes regardless of who built it. A few requirements stand out as directly relevant to the AI execution gap:
Every model an institution uses has to be classified by materiality and complexity, with the tier determining how intensively it's validated, who has to approve its deployment, and how it's monitored going forward. A chatbot answering FAQs and a model driving credit decisions are not the same risk category, and the guidance expects institutions to be able to demonstrate that distinction — not assume it.
The guidance draws a direct line between a model's impact on customers and the explainability threshold it must meet. Where full explainability isn't achievable, it requires compensating controls: more frequent validation, mechanisms to corroborate outputs before they're acted on, and usage restrictions. For any bank running black-box models bolted onto a legacy core, this is a hard requirement to retrofit after the fact.
This is the provision most likely to catch institutions off guard. Sourcing a model from a vendor — including a foundation model or an off-the-shelf AI layer — doesn't reduce an institution's obligations. The guidance requires independent validation by the institution itself, regardless of any certification the vendor provides, plus enhanced oversight by the board's risk committee irrespective of the model's risk tier. Contracts have to guarantee access to enough technical documentation to actually validate the model, along with audit rights and continuity arrangements. Institutions stitching together multiple third-party AI tools on top of an ageing core are signing up for a due-diligence burden that scales with every additional vendor.
The guidance calls for defined human-in-the-loop or human-on-the-loop arrangements, override and kill-switch mechanisms, and periodic review of model-driven decisions — with explicit attention to automation bias and decision fatigue among the humans doing the overseeing. It also requires institutions to test AI behaviour under adversarial and atypical conditions, guard against hallucination in generative models through system-level controls, and assess for bias and unfair treatment of customer groups, particularly in consumer-facing use cases.
Decommissioned models stay in the inventory for a minimum of ten years. Documentation standards apply for the same duration. This isn't a compliance checkbox; it's an institutional memory requirement that most fast-moving AI pilots were never architected to meet.
Read together, these provisions describe a fairly specific profile of institution: one that knows, at any point in time, exactly which models it's running, what they're doing, who's accountable for them, how they were validated, and what happens if they go wrong — for every model, including the ones it didn't build itself.
This is where the AI execution gap and the regulatory shift meet. A model bolted onto a legacy core as a separate layer — sourced from one vendor for fraud scoring, another for a chat interface, a third for document processing — inherits none of this governance by default. Each one is a distinct third-party model requiring its own due diligence, its own validation, its own contractual documentation guarantees, its own place in a ten-year inventory. The operational burden compounds with every new AI capability an institution bolts on, and that burden is precisely what stalls pilots at the production gate.
A model that's native to the core banking stack starts from a different position. If model inventory, audit trails, explainability thresholds, and human oversight controls are part of how the platform is architected — rather than something layered on after the fact for each new AI feature — the governance obligations the RBI guidance describes become a property of the system, not a project banks have to execute separately for every deployment.
This is the practical argument for AI capability built into the core rather than assembled around it. When M2P integrated AI as a native layer within its infrastructure rather than as an external add-on, the intent was exactly this: agentic AI that inherits the audit trails, access controls, and monitoring already present in the core banking, lending, and payments stack, instead of introducing a new, separately-governed system for every use case. In a regulatory environment that now explicitly penalises institutions for treating third-party AI as someone else's governance problem, that architectural choice stops being a technical preference and starts being a compliance advantage.
For institutions evaluating where their own AI initiatives sit against this shift, a few practical questions are worth asking before the next pilot gets greenlit:
Can every AI-driven decision in the pipeline be traced to a tiered, validated, documented model in an active inventory? If the honest answer is "some of them," that's the gap the RBI guidance — and its equivalents likely to follow elsewhere — is designed to close.
If a vendor-provided model updates itself, does the institution know when, why, and with what validation? Automatic updates from third-party providers are explicitly flagged in the guidance as requiring enhanced controls and stricter monitoring.
Is explainability a designed-in property of consumer-facing models, or something the team hopes to reconstruct if a regulator or a customer asks? The guidance treats the second answer as a control failure, not a technical limitation.
Does scaling an AI capability mean adding another disconnected vendor relationship, or extending a governance model that already exists? This is, in practical terms, the build-vs-bolt-on decision restated as a compliance question.
None of this means banks should slow down on AI. It means the institutions that move from pilot to production fastest will be the ones whose AI infrastructure was designed with model risk management as a first principle — not as a compliance exercise bolted on after the fact, in the same way the AI itself so often is.
This is the problem M2P's Turing CBS is built to solve. Rather than treating AI as an add-on module to be integrated, audited, and governed separately after deployment, Turing CBS embeds M2P's agentic AI capabilities directly into the core — across onboarding, credit decisioning, fraud and risk workflows, and customer servicing. That means the model inventory, audit trails, access controls, and monitoring that regulators like the RBI now expect aren't a separate project layered on top of the AI; they're already part of how the platform is built.
For a bank or NBFC evaluating its next core banking decision, this changes the calculation. Instead of assembling AI capability from a growing list of point-solution vendors — each one its own due-diligence exercise, its own validation obligation, its own line item in a ten-year model inventory — Turing CBS offers a single, governed stack where new AI-driven capability extends an existing control environment rather than multiplying it.
That's the difference between an AI pilot that impresses in a demo and one that survives in production: infrastructure that was built to be explainable, auditable, and overseen from day one, not retrofitted to become so under regulatory pressure. As central banks move from guidance to enforcement, that difference is quickly becoming the deciding factor in which institutions actually close their AI execution gap — and which stay stuck re-running pilots.
Talk to M2P about how Turing CBS brings governed, native AI into your core banking stack.