Signal Stack

B2B technology signals above the noise.

AI Infrastructure · 4 min read

ESUN Ethernet for Scale-Up Networking Takes On NVLink

OCP's ESUN effort aims to give Ethernet the lossless, low-latency behavior scale-up AI fabrics need, competing with Nvidia's NVLink and Broadcom's SUE track. The architecture rationale is public; per-port bandwidth and interoperability data are not.

ESUN Ethernet for Scale-Up Networking is the Open Compute Project’s open alternative to Nvidia’s proprietary NVLink, and the numbers explain why hyperscalers are pushing it: NVLink 5.0 already delivers up to 1.8 Tbps of bidirectional bandwidth per GPU, and Broadcom’s Tomahawk Ultra silicon for the related Scale-up Ethernet (SUE) track claims 51.2 Tbps of switch throughput.

Quick take

ESUN is an Open Compute Project workstream defining Ethernet for AI scale-up fabrics, positioned as an open alternative to Nvidia’s proprietary NVLink.

Broadcom’s Tomahawk Ultra ASIC, built for the related SUE track, is rated at 51.2 Tbps of switch throughput.

The specification targets Ethernet’s IP header overhead and Priority Flow Control, which do not fit tightly synchronized AI traffic.

No document in this evidence set states a shipping date or a per-GPU bandwidth figure for ESUN itself, so it cannot yet be benchmarked against NVLink 5.0 directly.

What Changed: An Open Ethernet Answer to NVLink

The Open Compute Project has formalized ESUN alongside a related SUE-T workstream, aiming to give hyperscalers a rack-level scale-up fabric built on Ethernet instead of a single vendor’s proprietary interconnect.

Broadcom is the most visible commercial backer, seeding the Scale-up Ethernet specification and already shipping Tomahawk Ultra hardware that supports ultra-low forwarding latency alongside its 51.2 Tbps throughput figure.

The competitive framing matters because clusters are scaling toward 1,024 accelerators per pod, a density at which the interconnect layer, not the accelerator itself, increasingly decides utilization.

Synopsys’ technical explainer describes ESUN’s goal in narrower terms than a straight NVLink clone: evolving Ethernet into a lossless, low-latency, deterministic transport built around how AI accelerators synchronize step by step.

Approach Type Disclosed capacity figure
NVLink 5.0 Proprietary (Nvidia) 1.8 Tbps bidirectional bandwidth per GPU
Tomahawk Ultra (SUE) Open, Broadcom-led 51.2 Tbps switch throughput

The Technical Gap ESUN Is Closing

Standard Ethernet was built as a best-effort fabric that tolerates occasional loss and jitter, an assumption that breaks down in AI training, where a single late packet in a collective operation can stall an all-reduce and idle the cluster.

Semiconductor Engineering’s account of the specification identifies three structural limits in standard Ethernet for this workload: IP headers that add 28–48 bytes of overhead per packet, no link-level error recovery beyond forward error correction, and Priority Flow Control that pauses an entire link rather than isolating the congested traffic class.

That third limit, coarse congestion control, produces head-of-line blocking exactly when tightly coupled AI traffic can least absorb it, which is the specific failure mode ESUN’s design work is aimed at.

NADDOD’s infrastructure overview explains why the distinction exists at all: scale-up networks carry high-frequency, memory-semantic traffic such as tensor and expert parallelism, while scale-out networks carry the more tolerant pipeline and data-parallel traffic that commodity Ethernet RDMA already handles.

Scale-up links in that framing need sub-microsecond latency because GPUs cycle at over 1GHz, less than 1ns per cycle, which is the performance bar ESUN has to clear to matter at the rack level rather than only at the pod boundary.

Where the Public Evidence Runs Out

None of the documents in this evidence set publish a bandwidth-per-GPU or per-port figure for ESUN itself, so it cannot yet be compared directly against NVLink 5.0’s 1.8 Tbps or against Tomahawk Ultra’s 51.2 Tbps switch throughput.

The evidence also does not establish a ratified release date, a certification program, or interoperability results for ESUN hardware, which means procurement teams are currently evaluating a specification track, not a shipping product line.

A separate and easily confused effort, the Optical Compute Interconnect Multi-Source Agreement formed by AMD, Broadcom, Meta, Microsoft, Nvidia, and OpenAI, addresses co-packaged optics for scale-up links rather than the Ethernet protocol layer ESUN defines; its first generation specifies four wavelengths at 50 Gbps per channel, delivering 200 Gbps per direction per fiber, with the roadmap scaling to 1.6 Tbps per fiber per direction.

That optics roadmap explicitly leaves open how future bandwidth increases get manufactured, wavelength count versus symbol-rate escalation, a physical-layer question distinct from ESUN’s protocol-layer header and congestion work, and buyers should not treat the two initiatives as substitutes for each other.

What This Means for Buyers

Choose ESUN-track Ethernet when the priority is avoiding single-vendor lock-in on scale-up silicon and preserving the option to mix accelerators and switch silicon from multiple suppliers, since that flexibility is the stated rationale behind the open-standard push.

Avoid committing a near-term deployment to ESUN specifically until a vendor publishes per-port bandwidth, latency, and interoperability numbers comparable to what Nvidia already discloses for NVLink 5.0, because the current public record covers architecture intent rather than measured performance.

The scale of what is being contested is visible in AMD’s competing 72-GPU Helios rack, which the vendor says reaches 260TB/s of scale-up bandwidth and 43TB/s of scale-out bandwidth, a reminder that the interconnect decision is being made rack by rack, not switch by switch alone.

That AMD figure is a scale-up bandwidth claim for a specific rack design, not a benchmark of ESUN or SUE hardware, and the evidence does not indicate which fabric technology the Helios rack’s scale-up domain runs on.

The Metric to Verify Before Committing

Before including ESUN-based switching in a procurement plan, verify the published per-port bandwidth and measured tail latency under a realistic all-reduce workload, since tail latency, not average throughput, is what the evidence identifies as the actual failure mode standard Ethernet struggles with.

If a vendor cannot produce that number alongside a stated interoperability test result, the lower-risk near-term path is NVLink or a proprietary scale-up fabric, with ESUN reevaluated once OCP publishes conformance data rather than architecture descriptions.