What exactly is RDP AI Labs—a consultancy or a facility?
A facility and the engineering practice that runs it, doing two things. It is where RDP engineers its own AI hardware platforms—choosing accelerators, host processors, memory and cache strategy, fabric, cooling and power, then qualifying them into shipping SKUs. And it is where customers put their own workloads on real hardware and leave with measured numbers and a bill of materials sized against them. It is not an AI consulting business; the model and application work is done by specialist partners.
You say you build your own AI hardware. What does that mean today?
It means we are doing platform engineering, not badge engineering. We are building to NVIDIA's MGX modular reference architecture and following OCP Open Rack conventions for rack-scale work, evaluating AMD and Intel host platforms on merit, and designing the memory, storage, network, cooling and power around a target cost per useful token rather than around a datasheet. Every design that clears the bench is assembled, tested and packed at the same Hyderabad plant that has shipped 300,000+ devices over fourteen years, and carries three to five years of onsite support under SLA. Several of these designs are in progress rather than shipping, and we say which is which when you ask.
Why do you keep talking about cost per token instead of price per rack?
Because the rack is bought once and the tokens are paid for every day for the next three to five years. A configuration that looks expensive on the quote can be markedly cheaper per unit of served output once utilisation, cache behaviour, power draw and failure rate are counted—and the reverse is just as common. Cost per token is the only figure that lines those up against a cloud invoice.
What is cache tiering, and why not simply buy more VRAM?
High-bandwidth memory on the accelerator is the fastest and by far the most expensive place to hold a working set, and you cannot buy it independently of the GPU. Long-context and multi-user inference spends much of its life re-reading a KV working set that does not all need to live in HBM. Tiering that set across HBM, host DRAM and an NVMe cache tier—with a policy that decides what sits where—lets a given budget serve more concurrent context. It is a design choice with a measurable hit rate, not a slogan, which is why cache hit rate sits in our instrument set.
Do you build on NVIDIA or AMD? Xeon or EPYC?
On merit, per workload, and we will show you the working. Accelerator choice follows the workload and the software stack it has to run; host processor choice follows lane count, memory bandwidth, power budget and price, not brand loyalty. We build on both Intel and AMD host platforms and design to NVIDIA reference architectures, which is precisely what lets us run the comparison rather than assert an answer.
Do we have to travel, or can we do this remotely?
Either. On-site sessions put your engineers on the bench with ours, which is the faster path when you expect to change the configuration during the session. Scheduled remote access suits teams with an already-scripted workload, distributed engineers, or no appetite for travel. We will recommend one based on how your data has to be handled.
What does a lab engagement cost?
It is a scoped, paid engagement quoted before it starts, priced against the workload, the configurations under test and the length of the session. It is deliberately not an open-ended free demo—that model produces shallow results and a report nobody trusts. Tell us what you are trying to decide and we will scope it.
Does our data or model have to leave our environment?
No. Many sessions run on a shape-matched synthetic dataset or a sanitised subset that reproduces the performance characteristics without carrying real records. Where genuine data is required, we agree handling, retention and deletion in writing beforehand, and remote access lets your team keep control of what is loaded.
What if the benchmark says we need less hardware than we planned?
Then that is what the report says and what we subsequently quote. It happens often enough that we plan for it. A smaller correct configuration that gets renewed and expanded is worth considerably more to us than an oversized one that becomes an internal cautionary tale.
We are an ODM, component vendor or silicon partner. Is there a route in?
Yes. We are actively evaluating accelerators, edge AI modules and system-on-module silicon, memory and cache technologies, fabric, power and cooling components for current and next-generation designs. If you have something you believe belongs in an Indian-built AI platform, send it through the form and mark it as a co-engineering enquiry—the lab, not a procurement queue, reads those.
We already have an SI or AI consultancy engaged. Does that conflict?
Not at all—it is the common case. We benchmark and supply the infrastructure; they own the models and the application. We are frequently asked to run a session on behalf of a partner-led programme so that the infrastructure decision is evidence-based, and we are comfortable working to their test plan.
How is the resulting hardware procured?
Direct enterprise purchase, GeM for government and public-sector buyers, or online through RDP GPU Mart for standard systems. Pricing is quoted in INR with the bill of materials itemised, so procurement can review component by component what is being bought and what is being charged for engineering.