Skip to content
Make in India OEM · INR-transparent · Pan-India onsite SLATalk to sales: +91 720 794 8743Sign in

Festive-Peak AI Capacity Planning for Indian E-commerce

Updated 15 Jul 2026 · 5 min read

Overview

Festive-season peaks are the sizing event for Indian e-commerce AI infrastructure: Redseer measured the first 11 days of the 2025 festive events at roughly ₹60,000 crore GMV — about 3.5× business-as-usual — with some 90 million shoppers participating. Every AI surface that touches the shopper (search, recommendations, fraud scoring, conversational support) sees its load multiply in the same window. The planning question for 2026–27 is not whether to scale but how to split capacity between an owned GPU baseline sized for BAU-plus and burst capacity rented for the six to eight peak weeks, without letting p99 latency — and therefore conversion — degrade exactly when traffic is most valuable.

Festive-Peak AI Capacity Planning for Indian E-commerce
What you’ll learn: What festive multipliers do to each AI workload class, the owned-baseline-plus-rented-burst pattern, which models to degrade first under load, load-testing practice for inference tiers, and a peak-readiness table by workload.

Key takeaways

  • Festive traffic runs ~3.5× BAU on GMV and higher on request rates; latency-sensitive AI surfaces need 3–5× serving headroom, not GMV-proportional headroom.
  • Own the baseline, rent the burst: size owned GPUs for ~1.5× BAU utilisation and contract in-country burst capacity for the festive window.
  • Not all models deserve peak capacity — pre-compute more, shrink rerankers, and cache aggressively before adding GPUs.
  • Fraud and payment-risk scoring is the one workload that must not be degraded at peak; give it reserved capacity.
  • Load-test the inference tier with production-shaped traffic (burst patterns, cold caches) at least 6 weeks before the event.

What a 3.5× GMV event does to AI workloads

GMV multipliers understate infrastructure multipliers. Shoppers at festive peaks browse more per purchase, search more speculatively, and abandon and return more — so search queries, ranking calls and fraud scores per rupee of GMV all rise. Tier-2+ cities contributed 60–65% of 2025 festive shoppers per Redseer, arriving on slower networks where server-side latency budgets matter more, not less. A platform that sizes AI serving to GMV growth alone will find its reranker and conversational tiers saturating first, because those carry the highest per-request compute.

The owned-baseline, rented-burst pattern

Owning GPU capacity sized for peak means idling it for ten months; renting everything means paying peak cloud rates year-round. The pattern that works: own an inference baseline that runs at healthy utilisation through BAU (covering search, recommendations and support), then contract burst capacity — in-country, for latency and DPDP reasons — for the festive window. IndiaAI-affiliated and domestic GPU providers quoting near $1/GPU-hour make short-window burst rental economically sane. The decision arithmetic mirrors our inference sizing checklist; the ownership threshold is utilisation across the whole year, not the peak.

Degrade by design: the model waterfall

Before buying a single extra GPU, decide what degrades first. A practical waterfall: widen batch windows and cache TTLs (invisible to users), switch long-tail surfaces from real-time to pre-computed output, swap cross-encoder rerankers for lighter models on non-search surfaces, shorten conversational context windows, and only then queue requests. Two workloads are exempt: payment and fraud scoring, where degraded models directly convert to losses, and cart/checkout-path ranking, where latency visibly cuts conversion. Give both reserved capacity fenced off from burst pools.

Load testing an inference tier properly

Web-tier load tests routinely pass while inference tiers fail, because model serving has different failure modes: KV-cache and embedding-cache cold starts, batch-size cliffs, GPU memory fragmentation after hours of mixed traffic, and autoscaler lag measured in minutes while traffic doubles in seconds. Test with production-shaped bursts (flash-sale spikes at minute granularity), cold caches, and the real model mix — and rehearse the degradation waterfall as a game day, not a document. Teams running RAG-backed support assistants should also verify retrieval-index behaviour under concurrent refresh and query load; see the RAG reference architecture for the moving parts.

After the peak: what the data is worth

The festive window generates a year of training signal in weeks — demand shifts, new-user cohorts from tier-2/3 cities, fraud patterns that only appear at scale. Reserve post-peak GPU time for retraining on this data while it is fresh: demand models especially benefit, as covered in demand forecasting GPU sizing. The owned baseline that looked generous in November earns its keep as a training rig in December.

Peak-readiness by workload

Workload Peak multiplier vs BAU Strategy Degradable?
Search (embed + rerank) 3–5× QPS Owned baseline + burst; batch harder Partially — lighter reranker last resort
Recommendations 3–4× Shift long-tail surfaces to pre-compute Yes — batch more, refresh less
Fraud / payment risk 4–6× (attack traffic rises faster) Reserved capacity, no degradation No
Conversational support 5–10× Burst rental; shorter contexts; deflect to FAQ Yes
Catalog / content generation Front-loaded before event Complete batch jobs pre-peak; pause during Fully deferrable

Frequently asked questions

Should we size owned GPUs for festive peak?

No. Size owned capacity for roughly 1.5× business-as-usual at healthy utilisation and rent in-country burst for the festive window. Peak-sized owned fleets idle most of the year and wreck the ownership economics.

Which AI workload fails first under festive load?

Usually the highest per-request compute: cross-encoder rerankers and conversational LLM endpoints. They saturate before lightweight ranking models, which is why the degradation waterfall targets them first.

What must never be degraded during a sale event?

Fraud and payment-risk scoring — attack traffic scales faster than legitimate traffic at peaks — and checkout-path ranking, where added latency measurably cuts conversion. Both get reserved, fenced capacity.

How early should peak load testing happen?

Production-shaped load tests at least six weeks out, with a game-day rehearsal of the degradation waterfall closer to the event. Inference tiers fail differently from web tiers — test cold caches, batch cliffs and autoscaler lag specifically.

Is burst capacity from foreign cloud regions acceptable?

It works technically but adds latency for Indian shoppers and complicates DPDP posture for behavioural data. In-country burst — increasingly available near $1/GPU-hour — is usually the better default for the festive window.

Ready to deploy?

Talk to an RDP architect about power, cooling and lead time.

Request a Quote
👋 Ask GPU Mart AI — voice & text