Cornelis and Delos take on Nvidia's NVLink with new interconnects

Nvidia's NVLink is the scale-up fabric that lets multiple GPUs act as a single, giant accelerator, and its reach has grown from eight GPUs, to 72, to 576 today. As Nvidia extends that lead through its newer NVLink Fusion technology, rival networking vendors are scrambling to build alternatives. Right now, competing protocols such as AMD's Ultra Accelerator Link (UALink) are mostly tunneled over standard Ethernet switches: AMD, for instance, uses Broadcom's 102.4 Tbps Tomahawk 6-based switches, connecting to custom I/O dies on its MI455X accelerators. Purpose-built UALink switches and physical interconnects remain elusive. At the AI Infra Summit this week, two startups set out to change that: Cornelis Networks and Delos Data each unveiled new scale-up networking hardware and standards of their own.
On Monday, HPC-focused vendor Cornelis Networks introduced the Active Compute Fabric (ACF), an open architecture for scale-up and scale-out networking that builds programmable compute directly into the fabric, similar in spirit to the SHARP support inside Nvidia's NVSwitch ASICs, which already offloads collective operations to the switch to free up GPU compute. Cornelis, spun out of Intel in 2020, originally built its Omni-Path technology as a scale-out interconnect for supercomputers such as Trinity and Lynx, and is now also chasing scale-up networking, alongside its soon-to-launch 800 Gbps CN6000-series switches and NICs. ACF is meant to extend the goals of protocols like UALink and Ethernet for Scale-Up Networking (ESUN) into a shared baseline for in-network compute across vendors, not just Cornelis's own CN-series hardware, and the company names specific functions it wants to accelerate: offloading the KV cache that stores model state across sessions, dispatching experts in mixture-of-experts models, offloading message-passing-interface traffic for HPC workloads, and checkpointing training jobs within the fabric itself for faster failure recovery. Cornelis argues the payoff is large: in a hypothetical 100,000-GPU system, its own modeling of public data suggests roughly half of all GPU hours are spent waiting for data, worth about $1.68 billion a year in wasted capacity and 500 GWh of power, though this is Cornelis's own estimate and the source cautions readers to take it with a grain of salt. Alongside the launch, Cornelis announced about $205 million in funding to bring its next generation of scale-up and scale-out products to market.
Delos Data, founded by former execs from Barefoot Networks and Intel, took a different approach, focusing on physical hardware under a lineup it calls Nonstop AI. Its data interface comes in three planned form factors. The first is an I/O die capable of more than 30 Tbps of aggregate bandwidth, about 8 TB/s in either direction, more than double the 3.6 TB/s cap on Nvidia's and AMD's latest accelerators; how competitive that turns out to be depends on when Delos's chiplets actually ship, since the dies have to be co-designed directly into a partner's accelerator package. The second is near-packaged optics (NPO), offering a 10-plus Tbps data interface, equivalent to 2.5 TB/s of bidirectional bandwidth. Today's rack-scale systems, such as Nvidia's 72-GPU NVL72, still rely on copper interconnects because of power constraints, but copper stops being viable as clusters grow into row-scale setups of 576 and, eventually, 1,152 GPUs; Delos's NPO modules keep copper inside the rack and add user-serviceable optics between racks, so a failed optical module does not take out the whole accelerator. The third form factor is a 400-plus Gbps NIC for scale-out networks such as training clusters, front-end access, and storage. All three are protocol-agnostic, working with UALink, ESUN, or other schemes, unlike Nvidia's NVLink Fusion, which still requires customers to buy Nvidia's own NVSwitches for scale-up networking. The hardware builds on Delos's existing compute reference design and its Nonstop AI software, which monitors these fabrics and reroutes traffic automatically if a link fails. Delos said Tuesday that it has now raised more than $100 million in total funding, with backing from Matrix, Playground, and Socratic Partners, among others.
Key facts
- Cornelis Networks unveiled the Active Compute Fabric (ACF), an open architecture that builds programmable in-network compute into scale-up and scale-out fabrics, alongside its upcoming 800 Gbps CN6000-series switches and NICs, and announced about $205 million in new funding.
- Delos Data, founded by former Barefoot Networks and Intel execs, launched a protocol-agnostic Nonstop AI hardware lineup: an I/O die with more than 30 Tbps of aggregate bandwidth (about 8 TB/s per direction, more than double Nvidia's and AMD's 3.6 TB/s cap), near-packaged optics at 10-plus Tbps, and a 400-plus Gbps NIC; it has now raised more than $100 million in total.
- Both target Nvidia's NVLink, which today links up to 576 GPUs into one accelerator, up from 72 and originally eight; AMD currently tunnels the rival UALink protocol over Ethernet instead, via Broadcom's 102.4 Tbps Tomahawk 6-based switches on its MI455X.
- Cornelis models that a hypothetical 100,000-GPU system wastes roughly half its GPU hours, about $1.68 billion a year and 500 GWh of power, waiting on data, though the source cautions the estimate is Cornelis's own and should be taken with a grain of salt.
- Unlike Nvidia's NVLink Fusion, which still requires customers to buy Nvidia's NVSwitches for scale-up networking, Delos's chiplets are protocol-agnostic and work with UALink, ESUN, or other schemes.
Why it matters
NVLink is why Nvidia can sell not just chips but entire clusters that behave like one giant GPU, today scaling to 576 GPUs behaving as a single accelerator, up from 72 and originally eight. That scale-up fabric is a major source of Nvidia's lock-in: buying into NVLink Fusion still means buying Nvidia's own NVSwitches for scale-up networking. Cornelis and Delos are answering with real silicon and an open standard rather than just tunneling UALink over Ethernet, which is how AMD does it today. Two funded startups shipping purpose-built hardware and architecture, not just protocol specifications, is a sign the alternative-interconnect market is moving past the roadmap stage.
Who it affects
Hyperscalers and chip designers who want to build large AI accelerator clusters without paying Nvidia's premium, or without relying on today's workaround of tunneling UALink over standard Ethernet switches the way AMD does with Broadcom's Tomahawk 6 gear on its MI455X. HPC and supercomputing buyers are a natural audience too: Cornelis's underlying Omni-Path technology, spun out of Intel in 2020, already runs in supercomputers including Trinity and Lynx. It also affects the investors backing this race: Cornelis took in about $205 million, and Delos's more than $100 million came from Matrix, Playground, and Socratic Partners, among others.
How to use it
Neither company is shipping yet. The source calls Cornelis's 800 Gbps CN6000-series switches and NICs 'imminent'; Delos gives no ship date for its I/O die, NPO optics, or NIC, and the I/O die specifically needs to be co-designed directly into a partner's accelerator package before anyone could use it. No pricing has been disclosed for any of the hardware on either side. For now, the practical choice is which side to back: chip designers who would rather spend their effort and capital on the AI parts of their accelerators, which the article says applies to most hyperscalers, can look at licensing Delos's protocol-agnostic chiplet designs, which work with UALink, ESUN, or other schemes, or track Cornelis's open ACF standard, built to work across vendors' hardware rather than being tied to Cornelis's own CN-series parts.
How solid is it
Both companies bring real pedigree. Cornelis is an Intel spinout with interconnect technology already deployed in supercomputers; Delos was founded by former Barefoot Networks and Intel execs. Both have closed real funding: about $205 million for Cornelis, more than $100 million in total for Delos. But Cornelis's headline efficiency numbers, that a 100,000-GPU system wastes roughly half its GPU hours and about $1.68 billion a year waiting on data, come from Cornelis's own modeling of public data, not an independent benchmark, and the source cautions readers to take the figure with a grain of salt. No named spokesperson is quoted for Cornelis, only referred to as 'the company'.
Risks and caveats
Purpose-built UALink switches and interconnects are still largely absent industry-wide, and it is not stated whether Cornelis's ACF, or UALink and ESUN generally, have been adopted by any customer, or whether any of this new hardware has actually shipped. Delos's bandwidth edge over Nvidia and AMD is only theoretical until a chipmaker integrates its I/O die into a shipping accelerator package, a process the company itself says depends heavily on co-design. No pricing is available for any of the hardware described. And a better, more open fabric still has to unseat an incumbent: Nvidia's NVLink Fusion already runs in deployed clusters today.
“In a 100,000 GPU system, Cornelis modeling of public data shows that roughly half of all GPU hours are spent waiting for data, worth about $1.68 billion a year in wasted capacity and 500 GWh of power.”
— Cornelis