Cabling and Optics for AI Fabrics: DAC, AOC, 800G Transceivers and Link Budgets
Overview
The physical layer is the least glamorous part of an AI fabric and the most common source of intermittent cluster degradation. An 8-GPU node at 800G per GPU needs eight fabric links; a 512-GPU cluster runs thousands of cables and transceivers, and a single marginal optic among them can flap under load, trigger priority flow control storms and stall a training job that spans every node. Choosing correctly between DAC copper, active cables and pluggable optics – and enforcing link-budget margin, validation soak tests and labelling discipline – is cheaper than any amount of post-incident debugging.


Key takeaways
- Passive 800G DAC is limited to roughly 1.5-2 m by channel loss, but wins on power (near zero), latency and cost – use it wherever the reach allows, typically intra-rack.
- AEC (active electrical) extends copper to ~3-7 m; AOC suits fixed 3-30 m runs; pluggable transceivers over structured fibre are the norm for leaf-spine spans and anything serviceable.
- An 800G optical link budget is consumed by fibre loss plus every connector; dirty or excess connectors erode the margin that keeps pre-FEC BER inside correctable range.
- Marginal optics rarely fail cleanly: under full training load, temperature and supply shifts push BER past FEC limits, causing link flaps, PFC storms and cluster-wide stalls.
- Rising FEC corrected-codeword counts predict failures days ahead – per-link telemetry, burn-in soak tests, strict labelling and 1-2 percent spares are the operating discipline.
The interconnect ladder: DAC, AEC, AOC, transceivers
Four media classes cover an AI fabric. Passive DAC (direct-attach copper) is a twinax cable with no electronics: near-zero power, near-zero added latency, lowest cost – but at 800G the loss budget caps it around 1.5-2 m, confining it to intra-rack links such as GPU node to top-of-rack or rail leaf in adjacent racks. AEC (active electrical cable) adds retimer chips to stretch copper to roughly 3-7 m at a few watts. AOC (active optical cable) fixes transceivers to both ends of a fibre, reaching 30 m or more; it is cheaper than pluggables per link but unserviceable – a failed end means replacing the whole assembly. Pluggable transceivers over structured fibre (typically OSFP modules at 800G, in DR8, 2xDR4, FR4-class variants for 100 m to 2 km) cost the most and burn 13-18 W each, but are field-swappable and reusable across generations of cable plant. The pattern that works: copper inside the rack, fibre between racks, pluggables wherever a link must be serviceable.
800G optics in 2026: DSP, LPO and CPO
Conventional 800G modules carry a DSP retimer that cleans up the electrical signal – robust, interoperable, power-hungry. Linear pluggable optics (LPO) remove the DSP and lean on the switch SerDes, roughly halving module power to around 10 W and cutting latency, at the price of tighter host-and-module matching. Co-packaged optics (CPO) move the optical engine onto the switch package itself; vendor roadmaps – including Quantum-X and Spectrum-X Photonics systems slated through the second half of 2026 – claim large power and reliability gains, as surveyed by MapYourTech’s CPO status review. For clusters bought in 2026, DSP pluggables remain the safe default, LPO is worth qualifying for power-constrained rows, and CPO is a switch-purchase decision, not a cabling one. Fabric-level context is in AI Fabric Fundamentals.
Link budgets: where the margin goes
Every optical link has a budget: transmitter launch power minus receiver sensitivity, typically a handful of dB for short-reach 800G parts. Fibre itself consumes little over data-centre distances; connectors and splices consume the rest, at roughly 0.2-0.5 dB per mated pair – more when end-faces are dirty. A patch-panel-heavy path with four connector pairs can eat half the budget before the fibre matters. The budget interacts with FEC: 800G links run at raw bit-error rates that only forward error correction makes usable, and the margin between operating BER and the FEC correction ceiling is the real health metric. Design rules that hold up in practice: keep at least 1.5-2 dB of unallocated margin, minimise mated pairs per path, and treat every connector insertion as a cleaning event – inspect-and-clean is the single highest-value habit in optical operations.
Why marginal optics degrade clusters intermittently
A marginal optic passes acceptance testing at idle and fails only under load. Full-power training raises ambient and supply stress; laser bias and receiver characteristics drift; pre-FEC BER climbs past the correction ceiling and the link flaps for a few seconds. In an RDMA fabric that is not a local event: PFC pause frames propagate, congestion spreads to healthy links, NCCL collectives time out, and a job spanning the whole cluster stalls or restarts from checkpoint. Because the trigger is thermal and load-dependent, the symptom appears as unexplained periodic slowdowns rather than a hard fault – among the most expensive failure modes to chase, as optical vendors’ AI-training analyses describe. The defence is telemetry: FEC corrected-codeword counters, DDM optical power and temperature per module. A rising FEC correction trend on one link predicts failure days in advance; swapping that module in a maintenance window costs minutes, not training days. This telemetry belongs in the same day-2 stack as Day-2 Operations for Rack-Scale AI.
Labelling and validation discipline
Large fabrics are debuggable only if the physical layer is documented. Minimum discipline: every cable labelled at both ends with a scheme encoding rack, node, NIC index and leaf port (rail designs depend on NIC k landing on leaf k – see the rail-optimised topology deep dive); a machine-readable cable database reconciled against LLDP after any change; polarity and end-face inspection on every fibre link at install; and a 24-72 hour burn-in soak under synthetic RDMA load with zero-flap acceptance criteria before a fabric carries production training. Hold 1-2 percent of each transceiver and cable type as on-site spares – relevant in India, where replacement optics can carry multi-week import lead times, and where higher ambient temperatures and dust make end-face contamination and thermal margin genuinely tighter than data-sheet conditions assume.
Media options at 800G compared
| Media | Practical reach | Power per end | Cost class | Failure and service profile |
|---|---|---|---|---|
| Passive DAC | ~1.5-2 m | ~0 W | Lowest | Rarely fails; replace whole cable |
| AEC (active copper) | ~3-7 m | ~3-6 W | Low-mid | Retimer failure; replace assembly |
| AOC | ~3-30 m+ | ~8-12 W | Mid | Unserviceable ends; replace whole run |
| DSP pluggable + fibre | 100 m-2 km | ~13-18 W | High | Field-swappable; fibre plant reusable |
| LPO pluggable | ~100-500 m | ~9-10 W | High | Swappable; needs host qualification |
Frequently asked questions
Should I use DAC or optics inside a GPU rack?
DAC wherever reach allows. At 800G that means roughly 1.5-2 m, which covers most intra-rack node-to-leaf runs. The savings are real: near-zero power across hundreds of links, lower latency, lower cost and a failure rate well below active components.
Why did my cluster slow down every afternoon without any hard failure?
That pattern – load- and temperature-correlated degradation – is the classic signature of a marginal optic. Check per-link FEC corrected-codeword trends and module DDM temperature and power; a single link flapping briefly under peak load can stall collectives cluster-wide via PFC back-pressure.
Are third-party transceivers safe in AI fabrics?
Quality third-party optics are widely used and can meet identical specifications, but the burden of qualification shifts to you: insist on per-unit test data, run the same burn-in soak as for OEM parts, and confirm your switch vendor’s support stance. In high-count fabrics, consistency across a batch matters more than brand.
How much link-budget margin should I keep?
Keep at least 1.5-2 dB unallocated after accounting for fibre loss and roughly 0.2-0.5 dB per mated connector pair. Margin absorbs ageing, temperature drift and minor contamination – the difference between a link that degrades gracefully and one that flaps under load.
What acceptance test should a new fabric pass before production?
A burn-in soak of 24-72 hours under synthetic RDMA traffic at full line rate, with acceptance criteria of zero link flaps and flat FEC correction trends, plus a reconciliation of the cable database against LLDP neighbour data. Fabrics that pass this rarely produce mystery incidents in their first months.
Ready to deploy?
Talk to an RDP architect about power, cooling and lead time.