Skip to content
Make in India OEM · INR-transparent · Pan-India onsite SLATalk to sales: +91 720 794 8743Sign in

FREE-AI and Model Risk Rules: Infrastructure Consequences for Banks

Updated 15 Jul 2026 · 5 min read

Overview

Indian banks and NBFCs planning AI infrastructure for 2026–28 are building against a regulatory frame that now exists in writing: the RBI’s FREE-AI framework (August 2025) — six pillars, seven “sutras” including explainability and accountability, 26 recommendations — followed in June 2026 by draft Model Risk Management guidance covering AI models in decisioning. Neither mandates specific hardware, but both have direct infrastructure consequences: auditability, model inventories, kill-switch-style controls, incident reporting and indigenous-model encouragement all favour architectures where the regulated entity controls the serving stack. This article translates the regulatory direction into GPU infrastructure decisions.

FREE-AI and Model Risk Rules: Infrastructure Consequences for Banks
What you’ll learn: What FREE-AI’s pillars imply for infrastructure, the model-risk-management requirements that touch serving architecture, why controlled on-prem or in-country deployment simplifies compliance, capacity implications of logging and challenger models, and a control-to-infrastructure mapping table.

Key takeaways

  • FREE-AI is principles-first, not hardware-prescriptive — but audit, explainability and accountability requirements are far easier to evidence on infrastructure the bank controls.
  • The 2026 draft MRM guidance treats AI as models subject to inventory, validation and lifecycle control — meaning every deployed LLM needs versioning, monitoring and a rollback path.
  • Compliance multiplies compute: challenger models, shadow deployments, full decision logging and periodic revalidation add 30–60% overhead to naive serving estimates.
  • Data locality under DPDP plus RBI outsourcing norms makes in-country GPU serving the default posture for customer-data workloads.
  • Start with augmentation workloads (documentation, summarisation, service copilots) where FREE-AI’s risk bar is lower than in credit decisioning.

Reading FREE-AI as an infrastructure document

Three of the six pillars have hardware shadows. Infrastructure: the framework explicitly discusses compute enablement and encourages indigenous financial-sector AI models — a policy tailwind for banks building owned capacity rather than defaulting to offshore APIs. Governance and Assurance: board-approved AI policies, model inventories, audits and incident reporting require that every model’s version, inputs and outputs be reconstructible — trivially arranged on a controlled serving stack, contractually painful through third-party API endpoints. Protection: customer-data safeguards stack with the DPDP Act, pushing inference on personal financial data toward in-country, access-controlled infrastructure, the posture detailed in our BFSI controls and auditability article.

The MRM draft: LLMs become inventoried models

The June 2026 draft guidance extends classical model-risk discipline — inventory, independent validation, performance monitoring, lifecycle management — to AI-based decisioning tools. For an LLM estate this concretely means: pinned model versions with documented lineage (no silent upstream updates — a structural problem with consumer API endpoints), pre-production validation environments, champion/challenger operation for material models, drift monitoring, and demonstrable rollback. Each requirement consumes infrastructure: a validation environment is a second (smaller) GPU deployment; challenger models running in shadow double the serving cost of the decisioning path they check; and comprehensive logging of prompts, outputs and model states turns storage into a compliance line item — plan it with the storage planning guide.

Sizing with the compliance multiplier

A bank sizing GenAI serving naively — peak concurrent users through the arithmetic in the inference sizing checklist — will under-buy. Add: shadow/challenger serving for material models (up to +100% on those paths, typically +20–30% blended), the validation environment (a 2–4 GPU node that also serves development), periodic revalidation and fine-tuning batches (schedulable off-peak on serving hardware), and logging/observability overhead. A realistic planning multiplier is 1.3–1.6× the naive serving estimate. The good news: augmentation workloads — the sensible first wave — are small-model-friendly, and a single 4–8 GPU node covers a mid-size bank’s internal copilots, document summarisation and knowledge retrieval with the multiplier included.

Sequencing under the framework

FREE-AI’s risk-proportionate stance suggests a deployment ladder. Rung one: internal augmentation — policy Q&A over RAG, call summarisation, developer copilots — low customer-harm potential, ideal for building the governance muscle (see the RAG reference architecture). Rung two: customer-facing assistance with human oversight and strict scope. Rung three: decisioning-adjacent uses (credit memo drafting, fraud triage) under full MRM treatment with challengers and validation. Each rung reuses the previous rung’s infrastructure, which is the practical argument for owning a modest, well-governed GPU estate early rather than assembling one under regulatory deadline later. Fraud and risk model hosting specifics are covered in our on-prem fraud and risk article.

Control-to-infrastructure mapping

Regulatory expectation Source Infrastructure consequence
Model inventory + lineage MRM draft 2026 Pinned versions, registry, no silent API updates
Explainability / audit reconstruction FREE-AI sutras Full prompt/output/version logging; storage tier
Independent validation MRM draft Separate validation GPU environment
Champion/challenger monitoring MRM practice +20–30% blended serving capacity
Kill-switch / rollback FREE-AI + MRM Blue-green serving, instant model swap capability
Customer-data protection FREE-AI + DPDP In-country, access-controlled inference

Frequently asked questions

Does FREE-AI require banks to run AI on-premises?

No — it is principles-based. But audit, accountability and data-protection expectations are structurally easier to evidence on infrastructure the bank controls, which is why regulated Indian entities increasingly default to on-prem or in-country serving for customer-data workloads.

Are third-party LLM APIs prohibited?

Not prohibited — but under MRM-style discipline they must be inventoried, version-pinned and auditable, which many consumer API terms cannot guarantee. Banks typically restrict them to non-customer-data uses with contractual controls.

How much extra capacity does compliance add?

Plan 1.3–1.6× the naive serving estimate: challenger/shadow paths, a validation environment, revalidation batches and logging overhead. Under-provisioning here surfaces as pressure to skip controls — the wrong failure mode in front of a regulator.

Where should a bank start its GenAI deployment?

Internal augmentation — RAG-based policy Q&A, summarisation, copilots — where risk is low and the governance stack (logging, versioning, review) can mature before customer-facing or decisioning uses arrive.

What changed with the June 2026 MRM draft?

It extends formal model-risk management — inventory, validation, lifecycle control — explicitly across models including AI tools used in decisions, converting FREE-AI’s principles into examinable expectations. It was open for comment through July 2026; final form may adjust details, so track the final circular.

Download the full blueprint (PDF)

Ready to deploy?

Talk to an RDP architect about power, cooling and lead time.

Request a Quote