Early-warning bottleneck profiler for GPU training nodes: GPU + RDMA fabric telemetry with job-level classification
-
Updated
Jun 13, 2026 - Python
Early-warning bottleneck profiler for GPU training nodes: GPU + RDMA fabric telemetry with job-level classification
KAI Data Center Builder
Performance metrics for AI / ML cluster
A simple experiment applying compressible flow principles to soften distributed gradient communication stalls.
Ultra-Ethernet compliance test architecture eliminating expensive HBM memory buffers. Design utilizes FPGAs, PRBS payloads, state hash tables, and virtual RDMA.
A distributed hardware-software co-design fabric for MoE models (DeepSeek-V3, Mixtral) that eradicates NCCL All-to-All communication stalls via distributed RoCEv2 RDMA virtual address MUX and JAX/XLA SPMD sharding.
Software examples for RDMA senders and receivers using RDMA WRITE and RDMA READ
Packet-level simulator (ns-3 + RoCEv2/DCQCN/PFC/ECN) comparing fat-tree vs rail-optimized topologies for AI training fabrics. Validated against NCCL on real GPUs.
PowerShell-based toolkit for Storage Spaces Direct strict bundle deployment, validation, and operational checks in Windows Server 2025 environments.
Add a description, image, and links to the rocev2 topic page so that developers can more easily learn about it.
To associate your repository with the rocev2 topic, visit your repo's landing page and select "manage topics."