Products

Desktops to data center, all Make in India

14 product categories across compute, AI, and data center. Deployment-ready from our 28,000 sq ft facility.

Download Product Catalog
AI Solutions

Sovereign AI infrastructure

End-to-end AI compute under one sovereign umbrella. Designed here. Manufactured here. Supported here.

Talk to a Solutions Architect
Support

SLA-driven. Not ticket-driven.

Warranty. SLA. On-site service. Account management. Every commitment documented, every response time defined.

Download SLA Commitment
Company

Built on process, not promises

ISO 9001. PLI 2.0. SOP-led manufacturing. The systems behind every device we ship.

Our Story
RDP Technologies Limited company logo
Skip to main content

Home  /  RDP AI Labs

RDP AI Labs

We engineer the machine—then prove it on the workload.

RDP AI Labs is where our AI hardware is designed and where it is measured. We engineer platforms to open and reference architectures for the lowest cost per useful token, build them in our own Hyderabad plant with ODM and component partners, and validate them on real workloads—ours and yours—before anything ships.

Platform engineering and workload validation, in one place, attached to a factory that builds what it specifies.

A GPU rack on the RDP AI Labs bench, under benchmark Schematic of a rack of GPU nodes with a pod of three nodes highlighted as under test, feeding a live telemetry panel showing throughput and utilisation meters. THROUGHPUTPOD UNDER TEST

What the lab does

A lab that specifies, and a bench that proves

The point of doing both is that the team deciding what goes into an RDP system is the team that defends it on the bench.

The lab

Platform engineering

We decide what RDP's AI systems are made of, and why. Accelerator, host platform, memory and cache strategy, storage, fabric, cooling, power and the software baseline—chosen against a written design brief rather than a vendor deck.

See the design brief →

The bench

Workload validation

We measure. Every RDP platform is benchmarked against its own design brief before it ships, on the same harness we make available to customers who want to prove a configuration on their own workload.

See how validation works →

Where this happens. The lab sits inside the new-product development floor of our own plant in Hyderabad, with demo hardware of every technology on this page on site. Parts from a component base of thousands of vendors are benchmarked against full-product reference designs and locked to a SKU—the same bench whether the form factor is a 10 TOPS edge module, an eight-GPU node or a multi-node cluster, with larger rack-scale designs in development. RDP assembles, tests and packs the result under its own brand, and supports it onsite for three to five years against an SLA.

The lab · platform engineering

What a world-class RDP AI SKU has to do

Four targets, written down before a bill of materials exists. A configuration that misses any of them is not a cheaper product—it is a different product, and we say so.

Lowest cost per useful token

Not the lowest sticker price. A system is specified against the work it will actually serve—tokens, images, simulations—at a fully-loaded rupee cost including power and amortisation. A cheaper rack that serves fewer tokens per rupee is the more expensive machine.

Reliable under sustained load

Thermal margin at the intended rack density, not at bench idle. Serviceable parts, documented failure behaviour, and a burn-in that runs long enough to find what a short test misses.

Sovereign-ready by construction

Designed and integrated in India, supportable in India, with an auditable bill of materials. Sovereignty is a supply chain and a service model, not a label applied at the end.

Procurement-ready

Itemised BOM, INR pricing, GeM route available for public-sector buyers, and a specification a technical evaluator can check line by line without a sales call.

The decision framework

How the platform actually gets chosen

Ten layers, ten decisions, each with a reason behind it. Where a layer is also a cost-per-token lever, the deck below takes it apart.

LayerThe decisionWhat drives it
AcceleratorGPU, dedicated inference silicon, or edge moduleWorkload shape and how mature the toolchain is for the part. The fastest part is not the right part if the software is not there. Lever 01 →
Host CPUAMD EPYC or Intel XeonLanes per accelerator, memory channels, NUMA layout — and per-core licensing. Lever 06 →
MemoryCapacity, bandwidth, or tieringWhether the workload is capacity-bound or bandwidth-bound. Lever 03 →
Cache strategyHBM only, or tiered KV-cache offloadContext length and concurrency, and how much of the cache truly has to stay resident. Lever 03 →
StorageNVMe scratch, parallel file system, object tierCheckpoint bandwidth and restart cost, not capacity alone. A GPU waiting on storage is the costliest idle asset in the building.
FabricInfiniBand, RoCE, or a single scale-up domainThe collective pattern, and whether the job leaves one box at all. Lever 07 →
CoolingAir, rear-door heat exchanger, or direct liquidRack density and the building it lands in. The facility usually decides before the datasheet does. Lever 08 →
PowerTopology, phasing, redundancyMeasured sustained draw rather than nameplate, and headroom for the next node. Lever 09 →
OS & stackDistribution, kernel, driver and container baselineCombinations validated to work together, plus a versioning policy so a rebuild in twelve months reproduces the machine we shipped. Lever 04 →
Form factorWorkstation, 2U–8U server, or rack-scaleWhere the workload sits today and how it grows. The growth path is part of the specification, not an afterthought. Lever 10 →
Two levers do not sit in any one row. Precision and scheduling cut across every layer in this table — which is why they are levers in the deck rather than lines in the framework above. The framework itself is re-run as the market moves—new accelerators, dedicated inference silicon, edge modules and open rack designs are judged against it rather than bolted on. Each decision ends the same way: locked into the SKU's reference design, down to the component-vendor part, and defended on the bench before the SKU reaches a price list.

Value engineering

A rack has a price. A token has a cost.

Only the second number decides whether a machine was worth buying — so it is the number we engineer against. It decomposes.

cost per
useful token=the number we drive down

what you pay — these add

CapEx01 02 03 06 07 + IT energy08 09 + Cooling power08 + Support10
tokens/sec01 02 03 04 06 07 08 × utilisation05 × uptime10 × service life10

what you get — these multiply

Eight terms. Ten levers. No orphans.

Every lever below moves one of these terms, and every term has a lever against it. The token is the unit because measured inference is where we set the bar; write frames, concurrent streams, training steps or seats in its place, for your workload, and the fraction does not change. A machine that serves one person is judged in a different currency — latency, privacy, watts at the seat — and it is engineered on the same bench. A decision that does not move a term on this page is not value engineering, it is preference.

Costs add. Output multiplies — the capital line is committed the day the purchase order is signed; energy, cooling and support keep accruing for as long as the machine runs, and the denominator has to be earned every hour after.

Lever 01 of 10Benched on your workload, locked to the SKU

The largest line on the quote deserves the most argument

Silicon is the majority of the bill. The question is never which accelerator is fastest — it is which one, and how many, clears your target at the lowest total cost. Sometimes that is a smaller part. Sometimes it is fewer of a bigger one.

The decision
Class, vendor and count — GPU, dedicated inference silicon or edge module — judged against a measured workload rather than a datasheet.
What we measure
Cost per million tokens at your target latency first, then tokens/sec per accelerator and utilisation at that count.

Three ways to reach one target

8× flagship₹₹₹
4× mid-range
6× mid + cache tiercheapest that clears₹₹

your targetclears the targetdoes not clear

Schematic, not a benchmark. The shape is the point: the right answer is the cheapest configuration that clears the line, and it is frequently not the biggest part on the list.

What we will not do. We do not publish a percentage saving for any of these. Every one of them is a trade whose value depends on your context length, concurrency, model and building — a number lifted from someone else's workload would be worthless to you. We will measure it on yours, and hand you the harness. Read our note on KV-cache offloading →

How we build

Built to open and reference designs, with partners

We are not trying to reinvent the rack. We build to published modular and open architectures, and compose the rest with ODM and component partners already excellent at their layer.

What we build to

Published architectures, not guesswork

  • The NVIDIA MGX modular reference architecture, for GPU server and rack-scale designs
  • The OCP Open Rack direction for AI, for rack-scale power, mechanical and serviceability
  • Published vendor reference architectures for validated component combinations
  • Our own design brief on top, which is what makes the result an RDP product rather than a copy

How we compose

ODM and component partners, RDP integration

  • ODM partners for chassis, rack mechanical and high-volume sub-assembly
  • Component partners for memory, NVMe, network adapters, power and cooling — parts qualified on our bench from a base of thousands of vendors, then locked to the SKU
  • RDP owns the platform design, firmware and driver baseline, integration, validation and support
  • Manufactured, rack-integrated and burned in at our own facility in Hyderabad
Why this is the cost answer. Rack-scale is expensive to invent and comparatively affordable to build well. Standing on open and reference designs removes the most expensive engineering risk, and lets the value engineering go where it actually pays—memory strategy, thermal design, integration quality and lifecycle support.

Design to line

How a brief becomes a shipping SKU

Six gates.

1

Brief

Targets set before parts are chosen.

  • Workload profile
  • Cost-per-token target
  • Power envelope
2

Architect

The decision framework, layer by layer.

  • Platform selection
  • Thermal and power model
  • Draft BOM
3

Prototype

First article on the bench.

  • Bring-up
  • Firmware and driver baseline
  • Golden image
4

Validate

Measured against its own brief.

  • Sustained load and thermal
  • Benchmark report
  • Pass or redesign
5

Pilot

Made repeatable.

  • Factory acceptance test
  • Build documentation
  • Service and spares plan
6

Volume

Into the catalogue.

  • Built in Hyderabad
  • Listed on GPU Mart
  • 3–5 year onsite support, under SLA
Gate four is the one that matters. A platform that misses its design brief on the bench does not get quietly shipped with softer marketing—it goes back to architecture. The benchmark that gates it is the same one a customer can later run on their own workload.

The bench · workload validation

The same bench, open to your workload

Everything the lab uses to prove an RDP platform is available to prove a configuration against yours—before you commit capital to scale. On-site in India, or over scheduled remote access.

On-site with our engineers

Your team works alongside ours on the bench. Best when you expect to change the configuration mid-session, want to see the thermals yourself, or need a decision in the room rather than a report three weeks later.

Scheduled remote access

Booked access to lab hardware so your engineers run their own harness on their own terms, without travel. Best when the workload is already scripted or your people are distributed.

What you bring

A representative workload, a dataset or a shape-matched substitute if the real data cannot leave your estate, the target you are trying to hit, and the constraints you already know—power per rack, cooling, floor space, timeline.

What you leave with

A written benchmark report with the method published alongside every number, the test harness so your engineers can reproduce and challenge it, and a bill of materials sized against what was measured.

Commercials, stated plainly. A customer lab session is scoped, priced and quoted before it starts. If a smaller configuration suffices, the report says so. And the benchmark does not stop at the bench: the run you accept here is re-run as part of factory acceptance, so you are not asked to trust that the production cluster behaves like the prototype—it is demonstrated on the same harness before it ships.

The instrument set

What a session settles, and how it is measured

Each of these is a question with a measurable answer, and an expensive one to get wrong after the hardware has landed. One instrument set serves both sides of the lab: what we run against a candidate design is what your session reports. The harness ships with the results—an unreproducible number is marketing.

  • Throughput on your model
  • Multi-node scaling
  • Inference latency
  • Cache and storage strategy
  • Fabric under collectives
  • Thermal and power
  • Right-sizing
  • Configuration A against configuration B
MetricWhat it actually tells youWhy it decides the outcome
What it costs
Cost per million tokens lower is betterFully-loaded rupee cost of served output, including power, cooling and amortisation.The number a CFO can compare against a cloud invoice. It is the design brief in one figure.
Cache hit rate and reuse higher is betterHow much of the KV working set is served from the fast tier rather than recomputed.The lever behind memory tiering. It converts directly into tokens per rupee on long-context work.
What it produces
Tokens / sec / GPU higher is betterServing throughput normalised to hardware at a stated batch size and context length.A necessary normalisation when comparing like hardware at a stated batch size and context length. Cost per million tokens is what settles comparisons across different hardware.
Model FLOPs utilisation higher is betterHow much of the silicon you are paying for is doing useful mathematics, as a fraction of dense peak at the training precision.A widely used efficiency number in training. A low figure means you are about to buy capacity you will not use.
Scaling efficiency higher is betterThroughput retained going from one node to N.Exposes fabric and topology mistakes that no single-node test will ever reveal.
Collective bandwidth higher is betterMeasured all-reduce and all-gather bandwidth across the fabric.The ceiling on distributed training whenever the job is communication-bound, and the one rarely tested before purchase.
Storage throughput higher is betterWhether the data path sustains the accelerators or starves them.Starved accelerators cap every number above this row.
What the user feels
Time to first token lower is betterHow long a user waits before anything appears, at target concurrency.Decides whether an internal assistant gets adopted or quietly abandoned.
p99 latency lower is betterThe tail experience under load, not the average.Averages hide failures. Users and SLAs live in the tail.
Whether it holds up
Checkpoint time lower is betterWrite and restore duration at your model size.Determines what a failure costs you and how often you can afford to save.
Job failure rate lower is betterInterruptions and restarts per thousand GPU-hours.Sets the effective capacity of the cluster, as opposed to the nameplate capacity.
Sustained power and thermal margin more margin is betterDraw per rack under continuous load, and the headroom left before throttling.Separates a burst benchmark from a machine that holds its numbers on day ninety.
What we will not do. We do not present another customer's benchmark as a forecast of yours, we do not quote a platform number we have not measured ourselves, and we do not publish performance figures, certifications or compliance claims we cannot evidence. Where a number in a report is a design assumption rather than a measurement, it is labelled as one.

Starting points

The shapes we design to

Reference designs, sized to you. Each stands on a configuration we have already benchmarked and locked at the component level; the sizing—nodes, memory, storage, fabric—is what your workload settles on the bench. Full bills of materials live on the data centre and AI infrastructure page.

Development & Teaching Cluster

For: universities, R&D groups, platform teams

  • ComputeGPU workstations or a compact GPU server, sized per concurrent user
  • FabricStandard Ethernet; single-node or small shared pool
  • StorageLocal NVMe scratch plus shared project storage
  • FacilityAir cooled, standard power — lab or comms room
  • ProvesPer-user throughput, concurrency limits, software stack fit
Reference configuration

Enterprise Inference Pod

For: enterprises serving internal AI to production users

  • ComputeMulti-GPU inference nodes sized to concurrency and context length
  • FabricHigh-bandwidth east-west; redundant north-south
  • MemoryTiered working set—HBM, host DRAM and NVMe cache, sized and measured together
  • FacilityHigher-density racks; air or rear-door heat exchanger
  • ProvesTTFT, p99 latency at concurrency, cost per million tokens

Multi-Node Training Pod

For: AI teams, neoclouds, national and research programmes

  • ComputeMulti-node GPU servers scaling to rack level
  • FabricLow-latency, rail-aware compute fabric with separate storage path
  • StorageParallel file system sized to checkpoint bandwidth, not just capacity
  • FacilityHigh-density power, liquid cooling options, structural review
  • ProvesScaling efficiency, collective bandwidth, checkpoint cost, thermal ceiling
How to read these. Every configuration is finally sized against measured results, data volume, concurrency, power envelope and growth plan. Standard systems can be bought online at RDP GPU Mart; anything cluster-scale is quoted after a lab session and discovery.

Honest scope

Where RDP stops and your partners begin

RDP is a hardware OEM. We design, build and support the machine and the rack around it. We are not the right people to write your retrieval pipeline.

RDP delivers

The machine, and everything around it

  • Platform selection and value engineering—accelerator, host CPU, memory and cache strategy
  • System and rack design to published reference and open-rack architectures
  • Thermal, power and mechanical engineering, including liquid cooling
  • ODM and component co-engineering, qualification and supply-chain design
  • Workload benchmarking and validation in the lab
  • Sizing, capacity planning and itemised bills of materials
  • Manufacturing, rack integration, cabling design, burn-in and factory acceptance
  • Driver, firmware and container baselines; golden images; OS qualification
  • Onsite deployment, commissioning, lifecycle support and spares held in India

Specialist partners deliver

The model and application layer

  • Model selection, fine-tuning and evaluation
  • Retrieval pipelines, data engineering and vector search design
  • Application, copilot and agent development
  • MLOps platform build and ongoing model operations
  • Domain data labelling, governance and compliance programmes

We work alongside system integrators and AI consultancies rather than competing with them, and we benchmark on behalf of partner engagements already under way.

Why we draw the line here. We draw it here for a specific reason: an infrastructure vendor that also claims to build your models is asking you to take two very different kinds of risk on one supplier. We would rather name the boundary than blur it.

Why RDP AI Labs

One design team, one factory, one contract—all of it in India

Fourteen years designing and manufacturing computing hardware in India. AI Labs is where the next generation of it is engineered and proven—one design team, one factory, one contract.

We engineer the platform, not just resell it

Every layer of the platform is a decision we take and defend with measurements. We design to the NVIDIA MGX modular reference architecture and follow OCP Open Rack conventions so the work composes with the rest of the industry instead of stranding you.

A lab you can actually get to

No overseas benchmarking queue, no shipping your data across borders to get a number, no waiting on a regional lab slot that belongs to a larger account.

We manufacture what you benchmark

A 28,000 sq ft facility in Hyderabad, 300,000+ devices shipped, 1M+ end users. The system on the bench is not a loaner from a distributor—it is the thing we build or are qualifying to build, so what you validate is what you receive.

Certified platforms, stated accurately

Platform certifications and partner-programme memberships are listed on request, with the certificate or catalogue listing behind each one. We publish what we hold and do not imply what we do not.

In-country, in writing

Designed and built in India, supported by an Indian entity, with spares and engineering in-country. For public-sector and regulated buyers that is not a marketing line—it is who holds the design, the build record and the phone number.

Procurement that fits how India buys

Itemised bills of materials, rupee pricing, GeM route for public-sector buyers, and standard systems purchasable online at RDP GPU Mart. Finance and procurement can see line by line what they are approving.

Questions we get asked

Before you talk to the lab

What exactly is RDP AI Labs—a consultancy or a facility?
A facility and the engineering practice that runs it, doing two things. It is where RDP engineers its own AI hardware platforms—choosing accelerators, host processors, memory and cache strategy, fabric, cooling and power, then qualifying them into shipping SKUs. And it is where customers put their own workloads on real hardware and leave with measured numbers and a bill of materials sized against them. It is not an AI consulting business; the model and application work is done by specialist partners.
You say you build your own AI hardware. What does that mean today?
It means we are doing platform engineering, not badge engineering. We are building to NVIDIA's MGX modular reference architecture and following OCP Open Rack conventions for rack-scale work, evaluating AMD and Intel host platforms on merit, and designing the memory, storage, network, cooling and power around a target cost per useful token rather than around a datasheet. Every design that clears the bench is assembled, tested and packed at the same Hyderabad plant that has shipped 300,000+ devices over fourteen years, and carries three to five years of onsite support under SLA. Several of these designs are in progress rather than shipping, and we say which is which when you ask.
Why do you keep talking about cost per token instead of price per rack?
Because the rack is bought once and the tokens are paid for every day for the next three to five years. A configuration that looks expensive on the quote can be markedly cheaper per unit of served output once utilisation, cache behaviour, power draw and failure rate are counted—and the reverse is just as common. Cost per token is the only figure that lines those up against a cloud invoice.
What is cache tiering, and why not simply buy more VRAM?
High-bandwidth memory on the accelerator is the fastest and by far the most expensive place to hold a working set, and you cannot buy it independently of the GPU. Long-context and multi-user inference spends much of its life re-reading a KV working set that does not all need to live in HBM. Tiering that set across HBM, host DRAM and an NVMe cache tier—with a policy that decides what sits where—lets a given budget serve more concurrent context. It is a design choice with a measurable hit rate, not a slogan, which is why cache hit rate sits in our instrument set.
Do you build on NVIDIA or AMD? Xeon or EPYC?
On merit, per workload, and we will show you the working. Accelerator choice follows the workload and the software stack it has to run; host processor choice follows lane count, memory bandwidth, power budget and price, not brand loyalty. We build on both Intel and AMD host platforms and design to NVIDIA reference architectures, which is precisely what lets us run the comparison rather than assert an answer.
Do we have to travel, or can we do this remotely?
Either. On-site sessions put your engineers on the bench with ours, which is the faster path when you expect to change the configuration during the session. Scheduled remote access suits teams with an already-scripted workload, distributed engineers, or no appetite for travel. We will recommend one based on how your data has to be handled.
What does a lab engagement cost?
It is a scoped, paid engagement quoted before it starts, priced against the workload, the configurations under test and the length of the session. It is deliberately not an open-ended free demo—that model produces shallow results and a report nobody trusts. Tell us what you are trying to decide and we will scope it.
Does our data or model have to leave our environment?
No. Many sessions run on a shape-matched synthetic dataset or a sanitised subset that reproduces the performance characteristics without carrying real records. Where genuine data is required, we agree handling, retention and deletion in writing beforehand, and remote access lets your team keep control of what is loaded.
What if the benchmark says we need less hardware than we planned?
Then that is what the report says and what we subsequently quote. It happens often enough that we plan for it. A smaller correct configuration that gets renewed and expanded is worth considerably more to us than an oversized one that becomes an internal cautionary tale.
We are an ODM, component vendor or silicon partner. Is there a route in?
Yes. We are actively evaluating accelerators, edge AI modules and system-on-module silicon, memory and cache technologies, fabric, power and cooling components for current and next-generation designs. If you have something you believe belongs in an Indian-built AI platform, send it through the form and mark it as a co-engineering enquiry—the lab, not a procurement queue, reads those.
We already have an SI or AI consultancy engaged. Does that conflict?
Not at all—it is the common case. We benchmark and supply the infrastructure; they own the models and the application. We are frequently asked to run a session on behalf of a partner-led programme so that the infrastructure decision is evidence-based, and we are comfortable working to their test plan.
How is the resulting hardware procured?
Direct enterprise purchase, GeM for government and public-sector buyers, or online through RDP GPU Mart for standard systems. Pricing is quoted in INR with the bill of materials itemised, so procurement can review component by component what is being bought and what is being charged for engineering.

Talk to the lab

Tell us what you need to decide

A short brief is enough—a workload to benchmark, a configuration to argue out, a component to evaluate. An engineer reads it and replies with a plan, or the questions to write one.

Please do not include confidential data, credentials or personal information about third parties in this field.

We use the details you provide only to respond to this enquiry and do not share them outside RDP. An engineer picks these up during Indian business hours.

RDP AI Labs

Engineer the machine. Then prove it on the workload.

One brief—the lowest cost per useful token, held reliably, sovereign-ready. Bring us a workload and we will measure it against exactly that.