How to Deploy Dual-plane Network Architecture_ H3C Unleashes Ten-Thousand-GPU Computing Power with End-to-Network Synergy and Simplified Delivery

2026-08-19 3 min read
Topics:

    The scale evolution of large model ten-thousand-GPU clusters is pushing traditional single-plane networks to their physical limits. Due to the limited port density of a single switch, traditional single-plane networking has to extend vertically when facing the ultra-large capacity access of ten-thousand-GPU scale, stacking into a complex 3-layer Clos topology. The increase in network layers inevitably leads to more hops, significantly raising latency. To break this physical limitation, the dual-plane architecture emerged.

    What is Dual-plane Networking?

    Simply put, dual-plane networking completely decouples the AI computing center network physically, splitting it into Plane 1 and Plane 2—two completely independent and non-interfering Spine-Leaf networks. Each AI server connects across both planes through Bonding (NIC bonding) technology, bundling two physical network ports into a single aggregated link at the system level.

    Taking the H3C S9827-128DH AI computing switch as an example, the 400G ports on both the NIC side and the switch downlink are split in half into two 200G links for precise docking. The entire network is divided into 15 Groups, and with only a two-layer topology, it can natively support a dual-plane cluster with 15,360 200G lossless network ports.

    The physical isolation + high-density cross-connection design brings two major technical guarantees to AI computing centers:

    Avoiding single-point failures on the access side: When one side of the access Leaf switch goes down or one side of the access fiber is interrupted, the server Bonding and network-side routing instantly interlock, and traffic is immediately switched to the other healthy physical plane for forwarding, ensuring uninterrupted training and inference.

    Flat topology reduces construction costs: Single-plane two-layer networking cannot support ten-thousand-GPU scale; beyond ten-thousand GPUs, three layers are required. Dual-plane breaks through the two-layer limit and can natively carry ten-thousand-GPU clusters. Taking H3C S9827-128DH as an example, it directly saves the procurement cost of an entire layer of switching equipment and optical modules.

    Real Traffic Trace Under Dual-plane Architecture

    During daily operation and sudden failures, the traffic flow of the dual-plane network presents the following real trajectories:

    Scenario 1: End-side Single Point Disconnection (SE1 and Plane 1-Leaf1 Link Failure)

    When a 200G fiber link from server SE1 to Leaf1 in Plane 1 suddenly fails:

    SE1 to SE65: The SE1 end-side Bonding instantly senses that the Plane 1 link is unavailable, and switches the outbound traffic to Plane 2-Leaf1 for forwarding.

    SE65 to SE1: Return traffic is split. One path is forwarded directly back to SE1 from Plane 2-Leaf1; the other path follows the usual route through Plane 1. After reaching Plane 1-Leaf, due to the downlink disconnection, the routing has to cross planes through the Spine inter-cross link to the Spine of Plane 2, and then deliver to SE1 via Plane 2-Leaf1.

    Scenario 2: Access Layer Single Point Paralysis (Plane 1, Group1 Leaf1 Failure)

    When the entire Plane 1 Leaf switch inside Group1 suddenly goes down completely:

    Same-group server mutual access (SE2 in Group1 to SE1 connected under the failed Leaf): Since Plane 1-Leaf1 is down, all traffic is redirected through Leaf1 in Plane 2, Group1 for forwarding.

    Cross-group server mutual access (e.g., Server-128 in Group1 to Server-1 connected under the failed Leaf): Under normal conditions, Bonding dual links each share half the throughput. After the failure occurs, Server-128's traffic through Plane 2 is unaffected and still completes forwarding within Plane 2; while traffic originally going up through Plane 1, after reaching the Plane 1-Spine layer, due to the device sensing the target Leaf failure, the traffic crosses to Plane 2-Spine through the physical cross-link between Spines, then down to Plane 2-Leaf1, and finally delivered to Server-1.

    End-to-Network Synergy Breaks Through Ten-Thousand-GPU Deployment Challenges

    Addressing the performance optimization space under normal dual-plane operation and the high-density deployment threshold of ten-thousand GPUs, the H3C solution achieves software-hardware synergy breakthroughs through AD-DC AI Computing Edition, focusing on efficiency release and simplified delivery:

    ·Dynamic Traffic Distribution, More Balanced Bonding Traffic

    H3C's core breakthrough lies in achieving end-to-network synergy through AD-DC AI Computing Edition, cutting the traffic scheduling granularity finer. As the control brain of end-to-network synergy, AD-DC AI Computing Edition builds the basic scheduling framework through the Bond mode on the server end side and the NIC polling distribution algorithm; during normal training, the system evenly sprays AI traffic across all healthy physical channels. This fundamentally overcomes the common traffic imbalance problem of dual-plane architecture in daily state, eliminates channel idleness and unnecessary congestion, allowing the ten-thousand-GPU network to truly run at the most efficient level, fully releasing the physical bandwidth limit.

    ·Automated Operations, Ultra-simple Network-wide Deployment

    Although dual-plane high-density cross-connection networking can improve overall network performance, if the configuration of tens of thousands of interfaces, Bond ports, VLANs, and ACLs on the network side and end side relies on manual command line input one by one, it is not only extremely inefficient but also can easily cause hidden dangers due to configuration errors. H3C achieves automated deployment through AD-DC AI Computing Edition, turning the originally complex work into automatic completion:

    Network-side automation (including S-MLAG simplified deployment): Under ten-thousand-GPU high-density networking, facing the complex S-MLAG high-reliability architecture on the network side, users do not need to manually write complex protocol configurations. Simply check the network model on the AD-DC AI Computing Edition controller, and the system will automatically complete complex configuration delivery based on the entire network planning. The entire online process is completely visualized, greatly reducing human error.

    End-side intelligent optimization (NIC batch configuration): Automation is not limited to the network side, but can extend downward to the server end. The platform can obtain server link information in real time, automatically and intelligently recommend Bond member interfaces on the end side, achieving highly reliable access with end-to-network synergy. At the same time, the system can automatically achieve multi-tenant security isolation combined with VLAN + ACL, completely eliminating traditional inefficient manual configuration.

    In the evolution of large model ten-thousand-GPU clusters toward dual-plane architecture, H3C builds a high-density lossless physical foundation with S9827 series and other AI computing switches, and innovates through end-to-network synergy with AD-DC AI Computing Edition. On one hand, it fully releases the physical bandwidth limit of ten-thousand-GPU clusters through dynamic traffic distribution; on the other hand, it breaks through the deployment threshold of high-density networking through full-process automation. This deterministic network solution that deeply integrates hardcore high-density hardware with simplified end-to-network control not only confirms H3C's software-hardware integration strength in the computing power era, but also provides the most solid continuity guarantee for long-term, large-scale distributed training tasks of large models.

    You may also like

    AI Compute Gridlock? Congestion Fine-Tuning Finds the Culprit in Seconds

    2026-08-25
    When network congestion stalls AI training and leaves expensive GPUs idle, the root cause may not be too much traffic on too few links, but rather a rigid bandwidth allocation mechanism.

    H3C Completes Industry-First Full-Scenario 800G AI Computing Network Interoperability Test

    2026-08-19
    As AI computing moves into large-scale deployment, high-speed networks have become the foundation that keeps AI clusters running efficiently and ties distributed compute into one fabric.

    How to Deploy Dual-plane Network Architecture_ H3C Unleashes Ten-Thousand-GPU Computing Power with End-to-Network Synergy and Simplified Delivery

    2026-08-19
    As AI computing moves into large-scale deployment, high-speed networks have become the foundation that keeps AI clusters running efficiently and ties distributed compute into one fabric.

    Laying the Foundation for the Token Era

    2026-08-10
    The curtain has officially risen on the Token Economy. Today, the efficiency of token production, circulation, and application has emerged as a transformative force driving industrial revolution and societal progress.