Retail & E-commerce
8 articlesConversational Commerce in Indian Languages: GPU Sizing for Voice Retail
Voice and chat commerce in Indian languages runs a three-model pipeline per turn: speech recognition, language model, speech synthesis. Latency is the product requirement, and code-mixed Indian speech is where…
Festive-Peak AI Capacity Planning for Indian E-commerce
Festive events run about 3.5x business-as-usual GMV and multiply AI serving load further. The workable pattern: own a GPU baseline sized near 1.5x BAU, rent in-country burst for the peak…
Hybrid Product Search in 2026: GPU Planning for Semantic and Visual Search
Hybrid product search fuses BM25, vector retrieval and a GPU cross-encoder reranker inside a 50-200 ms budget. One L4/L40S-class GPU covers embedding and reranking for mid-market storefronts; vector memory (about…
Demand Forecasting GPUs: Sizing for Retail and Quick Commerce
Demand forecasting is retrain-dominated: size GPUs for the nightly training window across millions of SKU-location series, not for serving. CPUs suffice below a million series; quick commerce forces 4-8 GPU…
Catalog Content Generation at Scale: GPU Planning for Retail GenAI
Catalog enrichment is now a batch GPU workload chaining LLM copy, VLM tagging and diffusion imagery. Owned GPUs beat generation APIs once utilisation passes 50-70%; a 1-4 GPU server covers…
In-Store Vision AI: Edge GPU Planning for Retail Chains
In-store vision AI runs as a hybrid: edge GPU nodes handle 8-30 camera streams each for real-time theft and shelf alerts, while a central 2-8 GPU server retrains models and…
Agentic Commerce Infrastructure: GPU Planning for AI Shopping Agents
AI shopping agents turn retail inference into sustained, machine-speed API traffic. Merchants need a 7-13B tool-calling LLM tier on L40S/H100-class inference GPUs beside existing ranking, structured feeds first, and India-hosted…
GPU Sizing for E-commerce Recommendation and Personalization (2026)
E-commerce recommendation sizing in 2026 hinges on peak requests per second, a sub-100 ms latency budget, and embedding-table memory. Most Indian retailers need one L4/L40S-class inference server to a small…