NVIDIA Mellanox 920-9B110-00FH-0D0 in Practice: Optimizing Low-Latency RDMA Fabrics for HPC and AI Clusters
August 25, 2026
NVIDIA Mellanox 920-9B110-00FH-0D0 in Practice: Optimizing Low-Latency RDMA Fabrics for HPC and AI Clusters
Background & The Challenge: When HDR Performance Meets Budget Reality
A mid-sized AI research laboratory recently faced a familiar challenge: their existing 100Gb/s EDR InfiniBand fabric was becoming a bottleneck for their expanding GPU cluster, which had grown from 128 to 512 NVIDIA A100 nodes. Training times for large-scale transformer models had increased by over 40% due to network congestion during all-reduce operations, and the team needed an upgrade that could deliver substantial performance gains without the capital expenditure required for a full 400Gb/s NDR deployment.
The lab evaluated several options and identified 200Gb/s HDR as the ideal balance of performance and cost. However, they needed a switch that could provide high port density, advanced congestion control, and seamless integration with their existing HDR-compatible cables and transceivers. This is precisely where the NVIDIA Mellanox 920-9B110-00FH-0D0 entered the picture—a switch purpose-built for environments where HDR provides ample bandwidth without overprovisioning.
Solution & Deployment: A Cost-Optimized HDR Fabric
The lab deployed the 920-9B110-00FH-0D0 InfiniBand switch OPN (Ordering Part Number) as the cornerstone of their new leaf-spine architecture. Each compute rack received one MQM8790-HS2F-based switch—the silicon platform underlying the 920-9B110-00FH-0D0 MQM8790-HS2F 200Gb/s HDR—providing 40 ports of 200Gb/s connectivity to GPU servers via OSFP-to-OSFP direct-attach cables. At the spine layer, four additional 920-9B110-00FH-0D0 switches interconnected all leaf switches, creating a non-blocking fat-tree fabric with full bisection bandwidth.
What made this deployment particularly effective was the switch's integrated SHARPv2 (Scalable Hierarchical Aggregation and Reduction Protocol) technology, which offloaded collective communication operations directly onto the switch fabric. In practice, this meant that all-reduce operations, which previously consumed significant host CPU cycles and network bandwidth, were now completed in a single pass across the fabric—dramatically reducing job completion times.
According to the 920-9B110-00FH-0D0 datasheet, the switch delivers sub-200ns cut-through latency, which aligned closely with the lab's performance requirements. The engineering team also confirmed that the 920-9B110-00FH-0D0 compatible ecosystem included all the cable types and optics they had already standardized on, eliminating the need for costly recabling efforts.
Results & Measurable Gains: From Bottleneck to Enabler
After the deployment, the lab measured improvements across several critical dimensions:
- Job completion time reduction: Training iterations for a 13B-parameter model decreased by 35%, with all-reduce latency dropping from 12μs to under 5μs for typical collective sizes.
- Network utilization: Average fabric utilization increased from 58% to 83% without triggering congestion, thanks to the switch's adaptive routing and congestion control mechanisms.
- Power and cooling efficiency: Compared to the previous EDR fabric, the new HDR infrastructure reduced per-port power consumption by 30%, allowing the lab to add more compute nodes without exceeding their data center power budget.
From a total cost of ownership perspective, the 920-9B110-00FH-0D0 price per port was significantly lower than NDR alternatives, enabling the lab to upgrade their entire fabric within their existing budget. The 920-9B110-00FH-0D0 specifications also highlighted the switch's dual-redundant power supplies and hot-swappable fan modules, which contributed to the lab's goal of 99.999% availability for their production training jobs.
Network engineers noted that the 920-9B110-00FH-0D0 InfiniBand switch OPN solution integrated seamlessly with NVIDIA's Unified Fabric Manager (UFM), providing centralized visibility into fabric health and automated alerting. This significantly reduced the time spent on troubleshooting and proactive maintenance, freeing up the team to focus on optimizing their training pipelines rather than managing the network.
Summary & Outlook: The Right-Sized HDR Foundation
The NVIDIA Mellanox 920-9B110-00FH-0D0 proved to be the right choice for this research lab, delivering the performance of 200Gb/s HDR at a cost structure that made business sense. By leveraging the switch's 40-port density, in-network computing capabilities, and robust management features, the lab transformed their network from a performance bottleneck into a scalable foundation for future growth.
Looking ahead, the lab plans to scale their cluster to 1,024 nodes over the next 18 months. With the 920-9B110-00FH-0D0's support for larger topologies through multiple spine tiers, they have confidence that their network can grow alongside their compute capacity. For organizations evaluating their HDR infrastructure options, the 920-9B110-00FH-0D0 for sale through NVIDIA's channel partners offers a compelling, production-validated solution that balances performance, density, and cost.

