Signature program

Built for GPUs. Lossless. Non-blocking.

Rail-optimized 400G and 800G Ethernet fabrics with RoCEv2, designed around your training and inference workloads — PFC, ECN and QoS tuned on the twin, NCCL benchmarks run at handover, and the fabric delivered as code.

What's included

01

Workload-to-topology sizing

GPU count, collective patterns and model sizes translated into rail-optimized leaf-spine with the right oversubscription — usually none.

02

Lossless Ethernet design

RoCEv2 with PFC, ECN and DCQCN parameters, buffer profiles and QoS classes tuned per platform and proven on the twin.

03

Platform selection

Arista 7060X/7800R, NVIDIA Spectrum-X, Juniper QFX or Cisco Nexus 9000 — chosen for the workload, with 800G and 1.6T-ready optics.

04

Front-end, storage and management networks

Separate fabrics for storage (GPUDirect), in-band management and out-of-band, each as code.

05

Validation

NCCL all-reduce and all-to-all benchmarks across the full cluster, failure injection, and a performance baseline you keep.

06

Operations

Telemetry for ECN marks, PFC pauses and buffer occupancy in the live portal; optional AIOps operations.

How an engagement runs

01

Sizing

Workload interviews and cluster plan. One to two weeks.

02

Design

Topology, QoS, optics, cabling, as code; twin validation.

03

Build

Rack-and-stack, cabling certification, bring-up.

04

Benchmark and hand over

NCCL results against targets, documentation, repository.

Questions we get asked

InfiniBand or Ethernet?

We build Ethernet fabrics. For most clusters under a few thousand GPUs, a well-tuned RoCEv2 fabric reaches the same job performance with simpler operations and multi-vendor choice.

Can you tune a fabric someone else built?

Yes. Most under-performing training clusters trace back to QoS and congestion-control settings we can measure and fix without new hardware.

Request a quote

Tell us about the project; a senior engineer responds the same business day.

Related
Reviewed by a senior engineer, not a sales queue.

GPUs waiting on the network?

Talk to an engineer