FREE-AI and Model Risk Rules: Infrastructure Consequences for Banks
Overview
Indian banks and NBFCs planning AI infrastructure for 2026–28 are building against a regulatory frame that now exists in writing: the RBI’s FREE-AI framework (August 2025) — six pillars, seven “sutras” including explainability and accountability, 26 recommendations — followed in June 2026 by draft Model Risk Management guidance covering AI models in decisioning. Neither mandates specific hardware, but both have direct infrastructure consequences: auditability, model inventories, kill-switch-style controls, incident reporting and indigenous-model encouragement all favour architectures where the regulated entity controls the serving stack. This article translates the regulatory direction into GPU infrastructure decisions.


Key takeaways
- FREE-AI is principles-first, not hardware-prescriptive — but audit, explainability and accountability requirements are far easier to evidence on infrastructure the bank controls.
- The 2026 draft MRM guidance treats AI as models subject to inventory, validation and lifecycle control — meaning every deployed LLM needs versioning, monitoring and a rollback path.
- Compliance multiplies compute: challenger models, shadow deployments, full decision logging and periodic revalidation add 30–60% overhead to naive serving estimates.
- Data locality under DPDP plus RBI outsourcing norms makes in-country GPU serving the default posture for customer-data workloads.
- Start with augmentation workloads (documentation, summarisation, service copilots) where FREE-AI’s risk bar is lower than in credit decisioning.
Reading FREE-AI as an infrastructure document
Three of the six pillars have hardware shadows. Infrastructure: the framework explicitly discusses compute enablement and encourages indigenous financial-sector AI models — a policy tailwind for banks building owned capacity rather than defaulting to offshore APIs. Governance and Assurance: board-approved AI policies, model inventories, audits and incident reporting require that every model’s version, inputs and outputs be reconstructible — trivially arranged on a controlled serving stack, contractually painful through third-party API endpoints. Protection: customer-data safeguards stack with the DPDP Act, pushing inference on personal financial data toward in-country, access-controlled infrastructure, the posture detailed in our BFSI controls and auditability article.
The MRM draft: LLMs become inventoried models
The June 2026 draft guidance extends classical model-risk discipline — inventory, independent validation, performance monitoring, lifecycle management — to AI-based decisioning tools. For an LLM estate this concretely means: pinned model versions with documented lineage (no silent upstream updates — a structural problem with consumer API endpoints), pre-production validation environments, champion/challenger operation for material models, drift monitoring, and demonstrable rollback. Each requirement consumes infrastructure: a validation environment is a second (smaller) GPU deployment; challenger models running in shadow double the serving cost of the decisioning path they check; and comprehensive logging of prompts, outputs and model states turns storage into a compliance line item — plan it with the storage planning guide.
Sizing with the compliance multiplier
A bank sizing GenAI serving naively — peak concurrent users through the arithmetic in the inference sizing checklist — will under-buy. Add: shadow/challenger serving for material models (up to +100% on those paths, typically +20–30% blended), the validation environment (a 2–4 GPU node that also serves development), periodic revalidation and fine-tuning batches (schedulable off-peak on serving hardware), and logging/observability overhead. A realistic planning multiplier is 1.3–1.6× the naive serving estimate. The good news: augmentation workloads — the sensible first wave — are small-model-friendly, and a single 4–8 GPU node covers a mid-size bank’s internal copilots, document summarisation and knowledge retrieval with the multiplier included.
Sequencing under the framework
FREE-AI’s risk-proportionate stance suggests a deployment ladder. Rung one: internal augmentation — policy Q&A over RAG, call summarisation, developer copilots — low customer-harm potential, ideal for building the governance muscle (see the RAG reference architecture). Rung two: customer-facing assistance with human oversight and strict scope. Rung three: decisioning-adjacent uses (credit memo drafting, fraud triage) under full MRM treatment with challengers and validation. Each rung reuses the previous rung’s infrastructure, which is the practical argument for owning a modest, well-governed GPU estate early rather than assembling one under regulatory deadline later. Fraud and risk model hosting specifics are covered in our on-prem fraud and risk article.
Control-to-infrastructure mapping
| Regulatory expectation | Source | Infrastructure consequence |
|---|---|---|
| Model inventory + lineage | MRM draft 2026 | Pinned versions, registry, no silent API updates |
| Explainability / audit reconstruction | FREE-AI sutras | Full prompt/output/version logging; storage tier |
| Independent validation | MRM draft | Separate validation GPU environment |
| Champion/challenger monitoring | MRM practice | +20–30% blended serving capacity |
| Kill-switch / rollback | FREE-AI + MRM | Blue-green serving, instant model swap capability |
| Customer-data protection | FREE-AI + DPDP | In-country, access-controlled inference |
Frequently asked questions
Does FREE-AI require banks to run AI on-premises?
No — it is principles-based. But audit, accountability and data-protection expectations are structurally easier to evidence on infrastructure the bank controls, which is why regulated Indian entities increasingly default to on-prem or in-country serving for customer-data workloads.
Are third-party LLM APIs prohibited?
Not prohibited — but under MRM-style discipline they must be inventoried, version-pinned and auditable, which many consumer API terms cannot guarantee. Banks typically restrict them to non-customer-data uses with contractual controls.
How much extra capacity does compliance add?
Plan 1.3–1.6× the naive serving estimate: challenger/shadow paths, a validation environment, revalidation batches and logging overhead. Under-provisioning here surfaces as pressure to skip controls — the wrong failure mode in front of a regulator.
Where should a bank start its GenAI deployment?
Internal augmentation — RAG-based policy Q&A, summarisation, copilots — where risk is low and the governance stack (logging, versioning, review) can mature before customer-facing or decisioning uses arrive.
What changed with the June 2026 MRM draft?
It extends formal model-risk management — inventory, validation, lifecycle control — explicitly across models including AI tools used in decisions, converting FREE-AI’s principles into examinable expectations. It was open for comment through July 2026; final form may adjust details, so track the final circular.
Ready to deploy?
Talk to an RDP architect about power, cooling and lead time.