All-Flash NVMe Design for AI: TLC vs QLC, DWPD, and EDSFF Form Factors
Overview
The flash layer of an AI storage tier comes down to four decisions: NAND type (TLC for write-heavy tiers, QLC for capacity tiers), endurance class (1 DWPD mainstream, 3 DWPD for checkpoint-heavy writes, roughly 0.3-0.6 DWPD for QLC), form factor (U.2 remains the compatibility default while E1.S and E3.S EDSFF drives win on density and cooling), and thermals (PCIe Gen5 drives can draw 20-25 W each and will throttle in a badly designed chassis). Get the write-workload maths right first; everything else follows from it.


Key takeaways
- Checkpoint-heavy training tiers are the main endurance risk: a tier absorbing multiple full-model checkpoints per hour can genuinely need 3 DWPD-class TLC, while dataset-serving tiers that are written once and read for months run happily on QLC.
- Enterprise TLC drives typically ship in 1 DWPD (read-intensive) and 3 DWPD (mixed-use) variants; QLC commonly rates around 0.3-0.6 DWPD but reaches 30-60+ TB per drive.
- E1.S suits dense 1U compute-adjacent flash; E3.S is becoming the 2U storage-server default and per vendor materials offers 2-5x the drive density of 2.5-inch U.2 bays.
- PCIe Gen5 NVMe drives deliver up to ~14 GB/s each but dissipate 20 W or more under load – airflow design and slot spacing are part of the storage architecture.
- Do the endurance arithmetic (TB written per day divided by drive capacity) per tier; buying 3 DWPD everywhere wastes budget, buying QLC for a checkpoint tier wastes drives.
Start from the write workload, not the datasheet
AI storage tiers see very different write patterns. Dataset tiers are written once during ingest and preprocessing, then read repeatedly – almost pure read-intensive. Checkpoint tiers absorb periodic multi-terabyte sequential bursts; a 70B-parameter model checkpoints roughly 1 TB of optimizer state and weights, and frontier-scale jobs write far more, as covered in checkpoint storage at frontier scale. Scratch and shuffle tiers see constant mixed I/O. Compute daily terabytes written per tier, divide by deployed capacity, and you have your required DWPD with a margin. Most architects find only the checkpoint and scratch tiers justify mixed-use endurance.
TLC vs QLC: where each belongs
TLC (3 bits per cell) remains the enterprise workhorse: higher endurance, consistent write latency, and sustained write bandwidth that does not collapse when an SLC cache fills. QLC (4 bits per cell) trades endurance and sustained-write behaviour for density and cost per TB – drives of 30 TB to 60 TB and beyond (Solidigm’s P5336 line reaches 61.44 TB, with 122 TB-class drives announced by several vendors) built mainly on QLC. The honest rule: QLC is excellent for read-mostly data that is written sequentially in large blocks – exactly how modern QLC-aware platforms lay out data – and wrong for small random writes, which multiply write amplification against an already-low endurance budget. SNIA’s endurance white paper is a good grounding in how DWPD ratings are actually derived from JEDEC workloads.
Reading DWPD honestly
DWPD (drive writes per day over the warranty period, usually 5 years) is a marketing-normalised view of TBW. Three things to check beyond the headline: the workload assumption (ratings assume the JEDEC enterprise mix; pure sequential checkpoint writes stress drives less than the rating implies, small random writes more), the capacity effect (a 30 TB drive at 0.6 DWPD absorbs more daily terabytes than a 3.84 TB drive at 3 DWPD), and the sustained-write cliff (a QLC drive’s write bandwidth after cache exhaustion can be a small fraction of its burst figure – demand sustained numbers, not peak). Vendor-quoted endurance is a floor for warranty, not a prediction of failure; monitor media wear via SMART and NVMe log pages rather than assuming.
Form factors: U.2, E1.S, E3.S
U.2 (2.5-inch, 15 mm) remains the broadest-compatibility choice with the deepest spares market – a real consideration for Indian procurement where replacement lead times matter. The EDSFF family is where new platforms are heading: E1.S (a ruler-style stick, 5.9-25 mm thicknesses) targets dense 1U servers and compute-node flash; E3.S targets 2U storage servers and all-flash arrays, with vendor materials citing 2-5x the density of equivalent U.2 bays plus connector and airflow designs built for Gen5 and Gen6 power levels. New enterprise drive families – Micron’s Gen6 9650 line, for example, ships in E1.S and E3.S – increasingly treat EDSFF as primary and U.2 as legacy. For a fresh 2026 build, specify E3.S for storage servers unless your chosen platform or spares strategy dictates U.2.
Thermals and power are storage design too
A Gen5 enterprise NVMe drive can dissipate 20-25 W under sustained load; 24 of them in a 2U chassis is a 500 W thermal problem before controllers and NICs. Thermal throttling shows up as mysterious bandwidth sag during long checkpoint writes – exactly when you need the drives most. Check the chassis vendor’s supported drive thermal classes, prefer EDSFF where airflow was designed in rather than retrofitted, and include drive-slot inlet temperature in monitoring. In Indian data centres running warmer ambient setpoints to save cooling power, derate more conservatively; a tier sized for rated bandwidth at 25 C inlet behaves differently at 35 C.
Drive class summary
| Tier | NAND / endurance class | Typical capacities | Form factor | Watch for |
|---|---|---|---|---|
| Checkpoint / burst write | TLC, 3 DWPD mixed-use | 3.2-25.6 TB | E3.S or U.2 | Sustained write bandwidth, power-loss protection |
| Hot dataset serving | TLC, 1 DWPD read-intensive | 7.68-30.72 TB | E3.S / E1.S | Read latency consistency at high queue depth |
| Capacity / warm data | QLC, ~0.3-0.6 DWPD | 30-61+ TB | E3.S / U.2 | Write-cliff behaviour, large-block-only writes |
| Node-local scratch | TLC, 1-3 DWPD | 1.92-7.68 TB | E1.S / M.2 / U.2 | Thermals in 1U compute nodes |
Procurement reality, 2026
Enterprise NAND remains on allocation through 2026, with QLC capacity drives among the most constrained SKUs; lead times and pricing are covered in the 2026 memory and NAND squeeze. Practical consequences: lock drive SKUs early in the design, qualify a second vendor per tier, and treat capacity headroom as something you buy deliberately rather than assume. For the wider tier design these drives slot into, see tiering architecture for AI data and AI factory storage planning.
Frequently asked questions
Is QLC safe for AI training data?
Yes, for the read-mostly dataset tier – written once in large sequential blocks, then read for months. It is the wrong choice for checkpoint or scratch tiers with heavy or small-block writes.
How much endurance do checkpoints really consume?
Compute it: checkpoint size x frequency x replicas per day, divided by tier capacity. A 1 TB checkpoint every 30 minutes is ~48 TB/day; on a 200 TB tier that is 0.24 DWPD – fine on 1 DWPD TLC. Shrink the tier or raise the frequency and the answer changes.
Should I buy U.2 or EDSFF in 2026?
For new storage servers, E3.S where the platform supports it – density, cooling, and future drive availability all favour EDSFF. U.2 remains sensible for spares continuity with an existing estate.
Do Gen5 drives need heatsinks?
Enterprise U.2/EDSFF drives rely on chassis airflow rather than per-drive heatsinks, but they do need the airflow the chassis vendor specifies. Verify supported drive power classes and monitor drive temperatures under sustained load.
What is power-loss protection and do I need it?
Capacitor-backed protection that flushes in-flight writes on power failure – standard on enterprise drives, absent on client drives. For any tier holding checkpoints or metadata, treat it as mandatory and avoid client-class drives entirely.
Ready to deploy?
Talk to an RDP architect about power, cooling and lead time.